Finding the Missing Link for Big Biomedical Data
It has been argued that big data will enable efficiencies and accountability in health care. However, to date, other industries have been far more successful at obtaining value from large-scale integration and analysis of heterogeneous data sources. What these industries have figured out is that big data becomes transformative when disparate data sets can be linked at the individual person level. In contrast, big biomedical data are scattered across institutions and intentionally isolated to protect patient privacy. Both technical and social challenges to linking these data must be addressed before big biomedical data can have their full influence on health care. It is this linkage challenge that we address in this Viewpoint.
Political campaigns, government, and businesses use big data to learn everything possible about their constituents or customers, and then apply advanced computation to hone strategy. The 2012 Obama campaign identified, approached, and influenced swing voters using data fused from Facebook, census, voter lists, and active outreach. The National Security Agency employs massive data on individuals from phone and Internet companies to identify terrorists. Google personalizes search results with the user’s web history and geographic context. In all these examples, the key has been to go beyond aggregate data and link information to individual people. Knowing that there are many swing voters in a zip code is helpful, but contacting those specific individuals may help to win an election.
Linking big data will enable physicians and researchers to test new hypotheses and identify areas of possible intervention. For example, do grocery shopping patterns obtained from stores in various areas predict rates of obesity and type 2 diabetes in public health databases? Does level of exercise recorded by home monitoring devices correlate with response rates of cholesterol-lowering drugs, as measured by continued refills at the pharmacy? Does increased physical distance from patients’ homes to hospitals and pharmacies affect utilization of health care and result in distinct patterns in claims data? To what extent do patients’ Facebook friends influence lifestyle choices and compliance with medical treatments? It is unknown whether these types of correlative inferences will really be found in big data and how physicians would use that information. However, being able to link data at the patient level is a prerequisite to exploring the possibilities.
The first challenge in using big biomedical data effectively is to identify what the potential sources of health care information are and to determine the value of linking these together. The Figure presents a potential way of approaching this problem by organizing data sets along different dimensions of “bigness.” Although some big data, such as electronic health records (EHRs), provide depth by including multiple types of data (eg, images, notes, etc) about individual patient encounters, others, such as claims data provide longitudinality—a view of a patient’s medical history over an extended period for a narrow range of categories. Linking data adds value when they help fill in the gaps. With this in mind, it becomes easier to see how nontraditional sources of biomedical data outside of the health care system fit into the picture. Social media, credit card purchases, census records, and numerous other types of data, despite varying degrees of quality, can help assemble a holistic view of a patient, and, in particular, shed light on social and environmental factors that may be influencing health.
