Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
This paper introduces DGEval, the first benchmark for assessing Large Language Models on the International Maritime Dangerous Goods (IMDG) Code, revealing that while models can outperform humans on multiple-choice questions and structured lookups, their unreliability in safety-critical areas like stowage and segregation necessitates continued human oversight before deployment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The ocean is a vast, powerful force, and moving goods across it requires a strict set of rules to keep ships, crews, and the environment safe. When companies ship hazardous materials like chemicals, batteries, or flammable liquids, they must follow a massive international rulebook called the International Maritime Dangerous Goods Code. This document is not a simple list; it is a complex web of instructions that dictates how every single item must be classified, packed, labeled, and placed on a ship. Getting it wrong can lead to fires, explosions, or toxic releases that are nearly impossible to fix once a ship is at sea. For decades, the people responsible for these decisions have relied on their own training and careful manual checks to ensure compliance. Today, however, many of these professionals are turning to a new kind of tool: artificial intelligence systems known as large language models. These computer programs can read and answer questions in human language, offering the promise of instant access to the thousands of pages of regulations. But because a mistake in this field can cost lives, the question is not just whether these tools are smart, but whether they are safe enough to trust with the most critical decisions.
A team of researchers set out to answer this question by building a rigorous test specifically designed for maritime safety. They created a benchmark called DGEval, which acts as a final exam for artificial intelligence. The test was built from real-world training materials used by shipping professionals and the official list of dangerous goods. It contains nearly 1,700 questions covering everything from identifying the correct name for a chemical to figuring out where a specific container can be placed on a ship without causing a disaster. The researchers evaluated thirteen different artificial intelligence models from six major technology companies. They tested these models in various ways, asking them to answer multiple-choice questions, write out free-text explanations, look up specific data, and identify exactly which part of the rulebook contained the answer. They also tested whether giving the models access to the internet to search for information helped them perform better.
The results revealed a clear divide between what these models can do and what they cannot do. The best-performing model, a large system from Google, was able to answer multiple-choice questions correctly more often than the average trained human professional. It excelled at retrieving straightforward facts, such as the official name of a chemical or its basic hazard class. However, the study found that the models were significantly weaker in the areas that matter most for safety. When the questions involved complex operational decisions—specifically how to separate incompatible goods or where to stow them on a vessel—the models struggled. Even the top-performing systems made frequent errors in these high-risk areas. The researchers also discovered that the models were terrible at pointing to the specific section of the rulebook where an answer could be found, a skill that is essential for a human to verify the AI's work.
The study also looked at how different types of questions affected performance. When the models were asked to simply choose the right answer from a list, they did reasonably well. But when asked to generate an answer from scratch without any options to choose from, their accuracy dropped. This suggests that the models are good at recognizing correct information when it is presented to them, but they are less reliable at producing that information on their own. The researchers found that giving the models access to the internet to search for answers helped them significantly when looking up specific data, but it did not help them with complex reasoning tasks. Interestingly, a model that had been specifically trained on maritime regulations did not perform better than a standard, general-purpose model, indicating that simply adding more text about ships to the training data is not enough to solve the problem.
Perhaps the most important finding was that the models were inconsistent. A system might get a simple fact right but fail on a safety-critical rule that could prevent a fire. The researchers noted that the models often refused to answer questions about dangerous substances like explosives or toxic gases, treating legitimate safety inquiries as if they were harmful requests. This inability to distinguish between a safety question and a safety threat means the tools are currently unreliable for the most dangerous parts of the job. The study concluded that while these artificial intelligence tools can be useful for helping professionals find information quickly, they are not yet ready to make safety decisions on their own. The researchers emphasized that human oversight remains essential. A professional must always check the AI's work against the official rulebook before making a final decision. The study is not a one-time verdict but a tool for continuous safety assurance, reminding the industry that as these computer systems evolve, they must be constantly tested to ensure they do not fail when it matters most.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.