← Latest papers
📄 public and global health

Analytical Centralization of Health Expenditure at the National Administrator of Health System Resources: Architecture, Data Quality, and Operational Performance of the ADRES Health System Analytics Platform, Colombia

This paper details the design, technical implementation, and operational performance of ADRES's new Azure/Databricks-based analytical platform, which successfully centralized and standardized Colombia's fragmented health data from over 55,000 providers to process trillions of records, offering transferable lessons for similar middle-income health systems facing large-scale digitalization mandates.

Original authors: Garavito Jimenez, D. A., Bello Angulo, D. E., Mejia Lemus, L. T., Chipatecua, D., Fula, D. D., Perez-Rubiano, S., Martinez, F. L., Bohorquez Pinzon, J. C.

Published 2026-06-10
📖 5 min read🧠 Deep dive

Original authors: Garavito Jimenez, D. A., Bello Angulo, D. E., Mejia Lemus, L. T., Chipatecua, D., Fula, D. D., Perez-Rubiano, S., Martinez, F. L., Bohorquez Pinzon, J. C.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine Colombia's entire health system as a massive, bustling city with over 55,000 different shops (hospitals, clinics, and doctors) all selling services to 52 million residents. For years, the city's accountant (ADRES) had a nightmare: every shop sent their receipts in a different language, on different types of paper, and sometimes with missing pages. The accountant had to manually sort through piles of messy paper, often months late, making it impossible to see the big picture or catch mistakes in real-time.

This paper tells the story of how that accountant built a super-powered digital brain in just one year to solve this chaos. Here is how they did it, explained simply:

1. The Problem: A Tower of Babel

Before 2024, the data was like a library where every book was written in a different alphabet.

  • The Chaos: Some shops sent digital files (XML), others sent spreadsheets, and some sent old-school paper lists.
  • The Glitches: The data was full of errors. One shop might say a city is "Bogotá," while another says "Bogota" (without the accent), or use a code that doesn't match the government's list. Some dates were written as "1900" or "9999" just to mean "no date," which broke the computer's math.
  • The Result: The accountant couldn't add up the bills correctly or spot fraud because the pieces didn't fit together.

2. The Solution: Building a "Smart Factory"

ADRES didn't just buy a new computer; they built a massive, automated factory on the cloud (using Microsoft Azure and Databricks) to process this data. Think of it as a three-stage assembly line:

  • Stage 1: The "Bronze" Bin (Raw Reception):
    Imagine a giant dumpster where all the messy receipts are dumped exactly as they arrive. If a receipt is crumpled, torn, or written in a weird font, it goes here first. Nothing is thrown away; everything is kept safe.
  • Stage 2: The "Silver" Wash (Cleaning & Sorting):
    This is the washing machine. Here, robots (software scripts) clean the data. They fix the "1900" dates, translate the different city codes so they all match, and organize the messy XML files into neat rows and columns. They also act as translators, converting old medical codes (ICD-9) into new ones (ICD-10) so history can be compared with the present.
  • Stage 3: The "Gold" Vault (Ready for Use):
    Now the data is shiny, clean, and organized. It's ready for the accountants and doctors to use. They can ask questions like, "How much did we spend on heart surgeries in Bogotá last year?" and get an answer in seconds.

3. The Superpowers

This new system is incredibly fast and strong:

  • The Scale: It holds over 110 billion records (that's like stacking paper receipts from the floor to the moon and back).
  • The Speed: It can process 220 million invoices in just 3 hours.
  • The Brains: It can look at 100 billion records and answer a question in 10 seconds. To do this, it uses a computer cluster as powerful as having 4,096 brains working at once.
  • The Accessibility: They added a "Conversational AI" (like a smart chatbot). Now, even people who don't know how to code can ask the system questions in plain English, and it will find the answer.

4. What They Actually Found (The Results)

Because they finally had all the data in one place, they could do things that were previously impossible:

  • Spotting the Ghosts: They found 471,480 people who were billed for medical services after they had already died. This was a massive waste of money (2.3 trillion Colombian pesos) that was hidden in the messy data before.
  • Checking the Math: They could instantly compare what insurance companies said they spent against what the hospitals actually billed, catching inconsistencies immediately.
  • Answering the Court: They could quickly respond to legal orders from the Constitutional Court with precise data, something that used to take months.

5. The Lessons for Others

The paper concludes with a few key takeaways for other countries facing similar messes:

  • Don't wait for perfection: You don't need the data to be perfect before you start. Build the system to handle the mess first (the Bronze layer), then clean it later.
  • Governance grows with the machine: You don't need to write every rule before you start. You can build the rules as you go, while the machine is already working.
  • The hard part isn't the computer: The biggest challenge wasn't the technology; it was fixing the "translation" problems between different databases (like making sure "Bogotá" and "Bogota" are treated as the same place).

In short: ADRES turned a chaotic pile of paper and digital noise into a clear, real-time map of the entire country's health spending, allowing them to stop fraud, save money, and make better decisions for the 52 million people they serve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →