Big Data has appeared in our lives to stay. With the evolution of digital technology, we are able to generate a large amount of information every day. For example, when you perform such a daily task that we like so much as using Google search, have you ever thought that with that simple gesture you are generating data that forms a huge set of information? We perform millions of searches on Google. This data is part of Big Data and is used to personalize your online experience.
Big Data is a trending topic in our society and is transforming the way we live and work. But is this phenomenon as recent as we think? What are its origins and evolution until today? And the big question is, are we looking at a new data-driven industrial revolution?
In this article, we will explore the history of Big Data, its application in the healthcare sector, its benefits, and the challenges it will present in the future.
From clay tablets to big data: the evolution of data analysis.

The origins of Big Data go back to ancient times. Civilizations such as Babylonian and Egyptian created the first libraries and recorded information to make decisions. However, it is with the advance of science and technology that data analysis consolidates as a discipline.
Important advances in the history of Big Data:
- 17th century: John Graunt performed the first known data analysis experiment. He studied how bubonic plague would spread in the city.
- 19th century: The term of Business Intelligent appeared, and information began to be used to improve business operations.
- 20th century: With the arrival of computers, large volumes of data are beginning to be stored, but the term Big Data is not yet popular.
- End of the 20th century: Journalist Erik Larson speaks for the first time with the term Big Data in a marketing article published in Harpers Magazine.
- 21st century: Along with the great growth of the Internet of Things (IoT) and the creation of open-source platforms such as Hadoop, Big Data becomes popular and becomes a fundamental technology for companies.
What is Big Data?
Big Data refers to very large and complex data sets that require non-standard IT applications to process these data and treat them appropriately. For example, photos and videos, messages and comments on Facebook generate several hundred terabytes of new data every day. If we add up all the data generated worldwide, it is estimated that by 2025 the total will exceed 181 zettabytes (equivalent to 181 billion gigabytes).

The 5 Vs of Big Data: an ever-expanding universe of data.
The main characteristics of Big Data are:
- Volume: this characteristic is the most obvious because Big Data is characterized by its immense size. The large amount of data generated every day on the Internet, with social networks, IoT sensors, and transactions, among others, contributing to the volume of data generated is growing exponentially.
- Velocity: Big Data requires a very high velocity to generate, store, and process information. To be able to work in real time with the data-generating sources, such as video cameras, a social network live, a blog, etc.
- Variety: data can be recorded in different forms. It can be structured and unstructured data. Structured data are easy to process, such a databases. Unstructured data is more complex and can come from audios, videos, or images that we generate from mobile devices, as well as from social media posts or blog articles.
- Accuracy: the main objective is to obtain reliable data in order to improve our decisions. Erroneous or incomplete data can be the cause of wrong conclusions and costly decisions.
Valorization: the value of Big Data is in its ability to generate knowledge and transform businesses. By analyzing large volumes of data, companies can extract patterns, trends, and business opportunities that would otherwise be invisible.
Data: the new engine of innovation in healthcare.

The commercial and financial sector have been the first to push Big Data, applying it in the industry for profit. But what would happen if we used this technology in healthcare? Let’s imagine a future where diagnosis and treatment are fully personalized, based on analysis of large volumes of data.
Big Data in healthcare offers a unique opportunity to improve medical care, identify disease patterns, develop new treatments, and optimize resource management. However, it is essential to guarantee the quality and privacy of the data to ensure safe and effective medical care.
Source of health information:
It comes from a wide variety of sources, from electronic medical records to data generated by mobile devices and social networks. Thanks to this wealth of information, we can build more accurate predictive models, identify pandemic outbreaks early, and develop new therapies.
Health data sources can be classified as follows:
- Clinical data: electronic medical records, medical images, treatment history.
- Data generated by devices: data from sensors, wearables, medical devices.
- External data: social networks, Internet searches, demographic data.
The 3 pillars of Big Data analytics: discover, predict, and decide.
Here are the 3 types of analysis most commonly used in Big Data:
Descriptive analysis: What has happened? This type of analysis is ideal for taking a look at the past and better understanding the present. As a result, we can identify patterns, trends and relationships between data. For example, we can segment our customers into groups with similar characteristics to adapt our marketing strategies.
Predictive analysis: What will happen? It helps us anticipate the future. Using statistical and machine learning models, we can predict future behavior, such as the probability that a customer will buy a product or that a piece of equipment will fail. This type of analysis is important when you make strategic decisions and want to prevent risks.
Prescriptive analysis: What should we do? Here we go a step further, and this analysis tells us what actions we should take to achieve our goals. If we combine the results of the two previous analyses, we can identify better options and optimize our decision. For example, we can determine the best distribution route for a product.

The data scientist: a profile in high demand.
To apply Big Data, new professional profiles called data scientists are needed.
There are highly specialized profiles that require a broad set of skills, including programming (Python, R), statistics, industry knowledge, problem solving and data visualization. Although finding such a complete profile can be a challenge, the growing supply of specialized training is enabling professionals from different areas to acquire the necessary skills to perform this role. The demand for data scientists will continue to grow in the coming years, driving the digital transformation of organizations.

There are so many different skills that a data scientist must have that looking for one would be like looking for a unicorn: a mythical creature with extraordinary abilities. The truth is that I would like to become a unicorn, but a blue one, not a pink one. Soon, I will start studying a Master’s Degree in Data Science, Big Data and AI. Who else is encouraged to train in this field?
What are Big Data projects like?
The first phase begins with an exploratory phase, characterized by critical thinking and the formulation of questions relevant to the business or research. These questions will guide the search and analysis data. In addition, a good analysis can generate new hypotheses to be developed in the future.
In the second phase, a clear strategy is defined, and the goals are established in line with the general objectives of the organization. A deep understanding of the problem to be solved and the resources available is essential.
Once the strategy is established, the most appropriate technological tools for the project are selected and applied. Among the most common technologies are Hadoop, Spark and various visualization tools. The data analysis process typically includes the following stages: obtaining and storing data from various sources, cleaning and transforming the data to ensure its quality, applying analysis techniques to extract insights and visualizing the results in a clear and brief manner.
The Health Data Lake and the importance of interoperability in healthcare.
Interoperability, i.e. the ability of different systems to exchange and use information jointly, is essential for the development of Big Data in healthcare. In Spain, work is underway on the creation of the Health Data Lake, a centralized repository that will store health data from all the autonomous communities. This project will make possible to apply artificial intelligence techniques and large-scale data analysis to improve patient care, identify disease patterns and develop new therapies. Thanks to interoperability, healthcare professionals will have access to a more complete and updated view of patients’ health, allowing them to make more informed and personalized decisions.
Big Data applications in the healthcare sector
Big Data is currently being developed in many areas of healthcare. Even though the electronic medical record (patient history in electronic format) and the rest of the health record are the necessary basis for data management, it is not enough to take full advantage of them. It is also necessary to apply Big Data to get the benefit of the knowledge obtained through a more in-depth data analysis application.
Which fields can benefit from the application of Big Data techniques?
- Genomics.
This is a major field of Big Data application directly related to bioinformatics. The objective is to collect, store, process and interpret the biological information encoded in the human genome. Thus, thanks to powerful analysis systems and a standardization of sequencing processes, we are able to achieve better results in genome research and sequencing. It can be a revolution thanks to the advances of this technique and its cost reduction.
How can it benefit us? It would be possible to predict more accurately whether or not a person would have a tendency to develop a pathology based on his or her genetic factors; we could thus anticipate its development. We would be moving in the direction of preventive medicine. Through pharmaco-genetics, we could choose personalized and more useful medications for patients.
- Clinical research.
By applying Big Data in this case, the causes of diseases can be determined more quickly and accurately, and better solutions can be found. It increases the quality of scientific documentation, among other benefits.
- Epidemiology.
Another important field where Big Data can make a difference is in the study of epidemics and how best to combat them. It is very useful for predicting the spread of a virus and, thanks to the geolocation of cell phones, knowing where it is spreading. In this way, vaccines can be prepared, and care centers can be set up in the most affected areas or to know where to impose movement restrictions if really necessary.
- Monitoring of chronically ill patients.
Thanks to technological innovation, it is now possible to capture and generate biometric signals of chronically ill patients from home without the need to go to or stay in the hospital. In this way, remote healthcare professionals can monitor the disease and personalize treatment.
- Clinical operative.
It is about hospital or medical center management, optimization of patient management, cost reduction, appropriate allocation of resources, reducing patient wait time, etc.
- Pharmacology.
In this field, Big Data is capable of complementing the information obtained in randomized clinical trials with patients, adding data from the real world, and thus improving the efficacy of a given active ingredient in a drug. With this system, the cost of medical development decreases, opening the door to new business models in the pharmaceutical sector. And if this were to happen, a lower-cost treatment for so-called rare diseases could be achieved, and, in this way, we could improve the health of more people in the population.
Big Data success stories in healthcare in Spain: SMUFIN y SAVANA
Spain is at the vanguard of the application of Big Data in the healthcare sector. Projects such as SMUFIN and SAVANA demonstrate the potential of this technology to transform medical research and patient care. SMUFIN, developed by BSC-CNS, has revolutionized genomic analysis of tumors, while SAVANA offer innovative solutions for extracting value from electronic medical records. These examples illustrate how Big Data can improve diagnosis, personalize treatments, and optimize the management of healthcare systems.
SMUFIN (Somatic Mutations Finder)
The journal Nature Biotechnology has published a scientific article on SMUFIN (Somatic Mutations Finder). It is a new system that uses Big Data to analyze the complete genome of a tumor and locate its mutations in a few hours. It also manages to detect alterations that were hidden even with other methods where supercomputers were used for weeks.
This great research work has been developed by the computational genomics group of the Barcelona Supercomputing Centre-Centro Nacional de Supercomputación (BSC-CNS), led by el Dr. David Torrents, in collaboration with the team of Dr. Elías Campo, jefe del grupo IDIBAPS Oncomorfología funcional humana y experimental and research director of the Clínic. El Instituto de Oncología de la Universidad de Oviedo (IUOPA), the European Molecular Biology Laboratory (EMBL, Heidelberg), and the Centro Nacional de Análisis Genómico (CNAG) have also participated.
If you want more information about this subject, follow the next link: Un nuevo método computacional permite analizar los cambios genéticos de pacientes con cancer en pocas horas | Hospital Clínic Barcelona (clinicbarcelona.org)
SAVANA
It is a company that develops platforms and solutions based on Artificial Intelligence algorithms to help hospitals and healthcare organizations in research and data strategy. They are able to analyze, summarize, and present in a simple format the medical information contained in the set of electronic medical records for use in clinical practice in real time.
If you want to take a look, I leave you the link to their website:

Comparison between both projects.
Both projects take advantage of the potential of Big Data to gain valuable information from large data sets in healthcare.
Both SMUFIN and SAVANA will drive innovation in the healthcare sector through the application of advanced technologies. And their ultimate goal is to improve patient care through more reliable diagnoses (SMUFIN) or better understanding of medical records (SAVANA).
The differences can be seen in the following table:

Challenges and opportunities of Big Data in healthcare: barriers and risks to overcome.
Technological barriers: interoperability between healthcare systems is essential for the success of Big Data in healthcare. However, the lack of common standards and the heterogeneity of the data make it difficult to integrate information from different sources. In addition, the quality of the data collected can be variable, which limits its usefulness for advanced analysis.
Legal and ethical barriers: the protection of personal data in the healthcare field is a priority. Current regulations must evolve to ensure the security and confidentiality of information while enabling research and innovation. It is also necessary to address the ethical challenges associated with the use of personal data, such as informed consent and equity of access to information technologies.
Human resources barriers: the shortage of professionals specializing in Big Data in healthcare is a significant obstacle. Specific training is required to develop the necessary skills to analyze large volumes of complex data and extract relevant insights for clinical decision-making.
The future of healthcare through the lens of Big Data: trends and challenges
Big Data is revolutionizing medicine, personalizing treatments and optimizing resources. But what does the future hold?
Talking about the future of this new technology is so dynamic that what is state-of-the-art today may be standard tomorrow, because at present it is such a young technology that it is not yet fully developed.
We are moving towards a more modern and, above all, more personalized medicine. Technological advances in ICT and in the field of genetics, together with new trends in wellness and health, will lead to the individualized design of therapeutic strategies for each patient.
The application of Big Data in healthcare will lead to an improvement in clinical practice processes and medical research; procedures will be performed more quickly, reducing costs, and, in this way, more patients can be reached; even patients with “rare diseases” will be able to enjoy personalized medication.
There could be a clear trend towards global data, i.e, communication between systems could be standardized so that they can communicate between countries around the world.

Perhaps virtual reality could be promoted; just as it is being developed for the world of video games, it would be useful in the medical world. For example, to train before an operation, being able to see the patient’s body in 3D. It could also be used to treat phobic disorders, post-traumatic stress, or addictions.
The development of 5G Technology could involve an improvement in the use of technological devices used for the control and monitoring of chronic pathologies. They are complemented by mobile applications that are used for sharing the information obtained from sensor measurements (patient data).
Visualization tools are closely linked to the mathematical and statistical models that have been currently being used. Why not develop new forms of data visualization that are more natural to humans? To be able to communicate data information as if you were telling a story, for example.
Machine learning is getting better and better results and is becoming more and more famous because the results are surprising. Machine learning is a form of data analysis where the analytical modeling process is automated through algorithms. Through these algorithms, hidden patterns are being found in the data. They are becoming more and more efficient and faster thanks to advances in technology.
Bibliographic references.
Informe Big Data en Salud Digital.pdf (ontsi.es)
https://espanadigital.gob.es/lineas de actuación/data-lake-sanitario
https://www.lamoncloa.gob.es/serviciosdeprensa/notasprensa/transformacion-digital-y-funcion-publica/Paginas/2024/140624-presupuesto-chttps://espanadigital.gob.es/lineas-de-actuacion/data-lake-sanitariocaa-espacio-datos-salud.aspx
Graphics and images.