Privacy Preserving Disease Prediction Across Multiple Hospitals Using CKKS Based Homomorphic Federated Learning
This paper proposes a CKKS-based homomorphic federated learning framework that enables privacy-preserving disease prediction across multiple hospitals by encrypting model updates to prevent membership inference attacks, though this enhanced security comes at the cost of significant communication overhead and reduced predictive utility compared to plaintext approaches.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world of medicine, a doctor's ability to predict a patient's future health often depends on the breadth of data they can see. Machine learning algorithms, the computer programs that learn from patterns, are becoming powerful tools for spotting risks like hospital readmissions or disease progression. However, these tools are only as good as the information they are fed. The problem is that this information—electronic health records containing sensitive details about patients—is scattered across thousands of hospitals, locked away by strict privacy laws and institutional rules. Hospitals cannot simply share their raw patient files with a central authority to build a better model without violating patient trust and legal regulations.
To solve this, researchers have developed a method called federated learning. Imagine a group of chefs who each have a secret family recipe. Instead of sending their recipes to a central kitchen where they might be stolen or copied, they agree to cook their dishes locally and send only the final taste to a central judge. The judge combines these tastes to create a master recipe, which is then sent back to the chefs for their next attempt. In this digital version, the "recipes" are the data, the "chefs" are the hospitals, and the "taste" is the mathematical update to the computer model. The raw patient data never leaves the hospital. Yet, even this method has a flaw: the "taste" sent back can sometimes reveal too much about the specific ingredients used, allowing a curious observer to guess which patients were in the training set. To fix this, scientists are now exploring a technique called homomorphic encryption, which allows calculations to be performed on data that remains locked in a digital safe, ensuring that even the person doing the combining cannot see the raw numbers.
A team of researchers at the MIT School of Engineering in Pune, India, recently put this concept to the test in a simulated environment designed to mimic a network of ten hospitals. Their goal was to see if they could build a machine learning model to predict whether a patient with diabetes would be readmitted to the hospital within thirty days, all while keeping the patient data completely hidden and the model updates encrypted. They used a specific type of digital lock known as CKKS, which allows for approximate math on encrypted numbers, a necessary feature because medical data often involves decimals and fractions rather than just whole numbers.
The researchers set up a virtual experiment using a public dataset containing records from 130 real hospitals, which they sliced into ten separate chunks to represent ten different institutions. Each virtual hospital trained its own version of the prediction model on its local data. In a standard, unencrypted scenario, these hospitals would send their model updates directly to a central server. In this new experiment, however, the hospitals first locked their updates inside a digital safe using the CKKS method. The central server, acting as the honest-but-curious coordinator, received these locked packages. It performed the math to average them together without ever unlocking the individual contributions. Only after the averaging was complete was the final result sent back to a trusted party to be unlocked and used to update the global model.
The results of this experiment revealed a clear and difficult trade-off between privacy and performance. When the researchers ran the same task without encryption, the system was incredibly effective, correctly identifying readmission risks with a high degree of accuracy. The model learned quickly and produced reliable predictions. However, in this unencrypted state, the system was vulnerable; an attacker could look at the updates and successfully guess whether a specific patient's data had been part of the training, effectively breaking the privacy promise.
When the researchers switched on the encryption, the privacy protection worked exactly as intended for the specific attack pathway studied. The system reduced the attacker's ability to guess whether a patient's data was in the training set to the level of random chance, achieving a success rate of exactly fifty percent, which is the same as flipping a coin. This finding is consistent with protection from the evaluated protocol against an honest-but-curious server, though it is not a guarantee of immunity to all privacy attacks or all adversarial conditions. The hospital data remained truly private, and the central server saw nothing but locked numbers.
However, this privacy came at a significant cost. The encrypted system was much slower and required far more data to be sent over the network. The size of the data packages sent from each hospital grew by a factor of twenty-four, meaning the communication traffic jumped from about 3 megabytes to nearly 38 megabytes for every round of training. Furthermore, the predictive power of the model dropped noticeably. While the unencrypted model achieved a high score for correctly identifying readmissions, the encrypted version struggled, with its ability to distinguish between patients who would and would not return falling to a level that, while better than random guessing, was far less useful for clinical decision-making. The model also took longer to stabilize, showing more erratic behavior as it tried to learn from the locked data.
The study concludes that while it is technically possible to build a privacy-preserving system for multi-hospital disease prediction using this method, it is not yet a perfect solution. The technology successfully stops the specific privacy leaks the researchers tested, but it does so by making the system much heavier and less accurate. The researchers found that for this specific setup, the loss in predictive quality and the massive increase in data transmission requirements are substantial hurdles. They suggest that for this approach to be useful in the real world, future work must focus on reducing the size of the encrypted data and improving the model's ability to learn effectively despite the digital locks. Until then, the choice between a highly accurate model that risks privacy and a perfectly private model that is less accurate remains a difficult balance for healthcare institutions to strike.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.