dictyBase 2026: a reimplementation of the Dictyostelium model organism database on minimal infrastructure
The paper describes the 2026 reimplementation of dictyBase as a self-contained, static-data-driven platform that consolidates genomic information, the complete Dicty Stock Center catalog, and an author-in-the-loop curation pipeline to serve the *Dictyostelium discoideum* research community.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the soil, a single-celled organism known as Dictyostelium discoideum lives a quiet, solitary life, eating bacteria and growing. But when its food runs out, this tiny amoeba undergoes a dramatic transformation. It signals to its neighbors, and thousands of them gather together, moving as a single unit to form a multicellular structure that resembles a tiny fruiting body. This simple life cycle makes the organism a powerful tool for scientists studying how cells communicate, move, and specialize. Because many of its genes have direct counterparts in humans, researchers use it to investigate complex conditions ranging from cancer to neurodegenerative diseases. For decades, the global community of scientists working with this organism has relied on a central digital library called dictyBase to find gene names, experimental protocols, and biological data. However, maintaining such a library has traditionally required expensive computer servers and specialized technical teams, a burden that has become difficult for smaller scientific communities to carry as funding becomes uncertain.
A team of researchers has now rebuilt this essential resource from the ground up, creating a version that runs on a fraction of the cost and complexity of its predecessors. Instead of relying on a massive, constantly running database server that requires a dedicated engineering crew to keep it alive, the new dictyBase operates as a self-contained service built from static data files. Imagine a library where the books are not stored on shelves that need constant reorganization, but are instead pre-printed, perfectly indexed copies that can be handed out instantly without needing a librarian to fetch them from a back room. The new system takes data from trusted global sources, organizes it into these static files, and serves them directly to users. This approach means the entire database can run on a single, modest computer, eliminating the need for complex, distributed software networks. The result is a robust, full-featured website that remains accessible even if the community's budget shrinks, ensuring that critical biological data does not disappear due to technical or financial hurdles.
The rebuilt database serves as a comprehensive hub for the 13,892 protein-coding genes of the Dictyostelium genome. Each gene has its own dedicated page that brings together information that was previously scattered across different tools. A researcher can view a gene's function, its connection to human diseases, how it behaves during the organism's development, and even see a predicted 3D model of its protein structure, all in one place. The site includes a complete catalog of the Dicty Stock Center, listing 7,055 distinct strains and 1,265 plasmids that scientists can order directly through the website. It also hosts a vast collection of genetic data from 20 different related genomes, allowing for comparisons between the standard laboratory strain and wild isolates found in nature. This capability helps scientists distinguish between genes that are highly conserved and essential for life, and those that vary widely between different populations.
To keep the information current without overwhelming the small team of curators, the researchers introduced a new way of handling scientific literature. When a new paper is published, a language model—a type of artificial intelligence—reads the text and drafts a summary of the gene annotations and findings. The authors of the paper then receive a private link to review and correct these drafts, ensuring accuracy before a human curator gives the final approval. This "author-in-the-loop" system speeds up the process of adding new knowledge while maintaining the high standards of expert review. The system carefully distinguishes between information that has been verified by experts and data that is still awaiting review, using clear visual badges so users know exactly how much confidence to place in any given piece of information.
Beyond simply storing data, the new platform offers a suite of tools designed to help scientists design their own experiments. Users can search for genes based on specific traits, such as how they are expressed during development or whether they are linked to a human disease. The site includes tools to design genetic experiments, such as creating guides for editing the genome or designing primers for DNA amplification, all tailored to the unique genetic makeup of the organism. For those studying the organism's movement and development, the database hosts thousands of video recordings of mutant strains, allowing researchers to watch how specific genetic changes affect the formation of the fruiting body. The site also includes an educational module with interactive diagrams and laboratory protocols, making the science accessible to students and teachers.
This project demonstrates that a full-featured scientific database does not need to be a massive, expensive infrastructure project to be effective. By shifting to a model based on static files and simple, self-contained services, the team has created a system that is easier to maintain, cheaper to run, and more resilient to funding changes. The entire source code and the scripts used to build the data are openly available, allowing other communities to inspect, reuse, or adapt the system for their own needs. The database is freely available to anyone with an internet connection, requiring no registration, and stands as a practical example of how modern tools can help preserve and expand access to vital scientific knowledge without the heavy overhead of traditional systems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.