KRAKEN: A provenance-tracked knowledge graph for multiomic and wellness research
KRAKEN is a scalable, provenance-tracked knowledge graph that addresses the underrepresentation of multiomic and wellness data by integrating diverse biomedical sources into a modular, 15-million-node system featuring standardized semantics, built-in analytical tools, and multi-modal interfaces for both human and AI-driven research.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the vast landscape of modern biology, scientists have long sought a way to connect the dots between the microscopic machinery of our cells and the broader picture of human health. For decades, research has focused heavily on understanding how diseases develop and how existing drugs might be repurposed to fight them. This approach has built powerful maps of biological relationships, but these maps often leave out the quieter, everyday aspects of wellness. They frequently miss the detailed chemical signatures of our metabolism, the specific genetic markers that influence our traits, and the complex data that defines how old our bodies feel compared to our actual age. Without a complete picture that includes these wellness-focused details, it remains difficult to see the full story of how our biology functions in a healthy state, not just when it is sick.
A new project called KRAKEN aims to fill this gap by creating a massive, interconnected map that bridges the divide between disease research and general wellness. The researchers behind this work realized that while existing maps were excellent for finding cures, they were incomplete for understanding the whole person. To solve this, they built a system that gathers information from many different sources, including specialized databases for lipids, genetic scores, and standardized health measurements. This new map does not just list facts; it links them together in a way that allows computers to trace connections across different types of data, from the molecules in our blood to the genetic codes in our DNA. The result is a tool that helps researchers and clinicians see how various factors of health and wellness are related, offering a more holistic view of human biology than was previously available in a single, unified system.
The core of this achievement is a knowledge graph that acts as a central hub for diverse biological information. The team combined data from several large, existing networks with specialized sources that focus on wellness, such as catalogs of polygenic scores and measures of biological age. By weaving these threads together, they created a structure containing approximately 15 million nodes, which represent individual pieces of information, and roughly 113 million edges, which represent the relationships between them. This network covers 62 different types of entities, ranging from specific chemicals and genes to broader health concepts. To ensure that this massive collection of data speaks a common language, the researchers used a standard semantic model that aligns with the broader efforts of the National Institutes of Health to make biomedical data easier to share and understand.
What makes this system particularly useful is its flexibility and efficiency. The researchers developed a build process that is lightweight enough to run on standard computing hardware, requiring less than 48 gigabytes of memory at its peak. This design allows users to rebuild the entire graph quickly or to select only specific parts of the data that are relevant to their particular area of interest. Whether a scientist is studying a specific metabolic pathway or looking at broad wellness trends, the system can be tailored to focus on that domain without needing to process the entire dataset. This modularity ensures that the tool remains accessible and practical for a wide range of research questions, rather than being a static, unwieldy archive.
Beyond simply storing information, KRAKEN includes a suite of tools designed to help users explore the data. These tools allow for multi-hop reasoning, which means the system can follow a chain of connections to find indirect relationships between two seemingly unrelated concepts. Users can also perform searches using text, vectors, or a combination of both to find specific entities or patterns within the graph. The platform supports enrichment analyses, which help identify whether certain groups of data appear together more often than would be expected by chance. All of these capabilities are accessible through an interactive web interface and a standard application programming interface, making the data available to both human researchers and automated systems.
The project also introduces a Model Context Protocol server, a feature that allows artificial intelligence agents and large language models to interact directly with the graph. This capability enables these advanced systems to consume the structured data and use it to answer complex questions or generate insights without needing to be retrained on the specific dataset. By making the data available in this way, the researchers are opening the door for new types of automated analysis that can leverage the full depth of the wellness and multiomic information contained within the graph. The entire system is freely available to the public, allowing anyone with an internet connection to explore these connections and contribute to the growing understanding of human health and wellness.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.