Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
Semalith v1.4 is a highly efficient 184M-parameter safety classifier that achieves state-of-the-art prompt-injection detection and zero false positives on benign agentic prompts with 44x fewer parameters than Llama-Guard-3-8B, while uniquely offering simultaneous detection of general harm and financial regulatory compliance in a single inference pass.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling library where millions of people are chatting with a super-smart robot librarian. This librarian, an Artificial Intelligence, can write stories, solve math problems, and answer questions about anything. But like any powerful tool, it has a dark side: sometimes, sneaky users try to trick the robot into breaking its own rules, revealing secret information, or saying mean things. This is called "prompt injection"—it's like whispering a secret code to the librarian to make them ignore the "Do Not Touch" signs.
To keep the library safe, we need a security guard. In the world of AI, these guards are special programs called "safety classifiers." Their job is to read every message before the robot sees it and shout "STOP!" if the message is dangerous. However, most guards are either too slow, too clumsy, or too paranoid. Some guards are so scared of trouble that they stop the librarian from helping with innocent questions, like asking how to reset a password. Others are so focused on simple bad words that they miss clever, sneaky tricks. The big question scientists are asking is: Can we build a guard that is fast, smart enough to spot tricky tricks, and gentle enough to let harmless conversations pass through?
Enter Semalith v1.4, a new security guard designed to solve exactly this problem. Think of it as a highly trained, 184-million-parameter "detective" (a number representing its brain size) that is much smaller and faster than the massive 8-billion-parameter guards currently used by big tech companies. While the giant guards are like heavy, slow-moving tanks that can get confused by subtle tricks, Semalith is a nimble, lightning-fast ninja.
The paper reveals that Semalith v1.4 is a master at spotting prompt injections—those sneaky attempts to trick the AI. In tests, it won every single one of the 7 prompt-injection benchmark categories it was tested on, beating the giant 8-billion-parameter guards by a huge margin. Even more impressively, it did this with 44 times fewer parameters (brain cells), making it incredibly fast. It can check a message in just 11.6 milliseconds, which is faster than the blink of an eye.
But Semalith isn't just about catching tricks; it's also a specialist in the world of money and finance. It was trained to understand specific rules for banks, loans, and insurance (known as BFSI compliance). This allows it to tell the difference between a harmless question about a credit card and a dangerous attempt to commit fraud. Unlike other guards that might panic and block all financial questions, Semalith correctly identified 208 harmless banking prompts without raising a single false alarm (a 0.000% false positive rate).
However, the authors are very honest about what this guard can't do. It's not a magic wand that solves every safety problem. The paper shows that while Semalith is the best at spotting sneaky tricks and financial rules, it sometimes struggles with general "mean" or "toxic" conversations compared to the giant guards. It's like a guard who is a world-class expert at spotting pickpockets and forgeries but isn't as good at catching someone just being rude. The researchers explain that this is a trade-off: by making the guard so good at spotting specific, complex tricks, it became slightly less sensitive to general rudeness. Furthermore, while it dominates the prompt-injection categories, it doesn't catch every single variation perfectly; for instance, on the hardest "stealth" level of one test (Mosscap L8), it caught about 80% of the attacks, leaving room for improvement.
The paper also highlights that this model was built using real-world data mined from the internet, not fake examples made by other computers. This makes it more reliable in the real world. The researchers tested it against 22 different challenges and found that Semalith won 7 out of 7 prompt-injection tests and 11 out of 18 overall. They admit there are six specific weak spots where the guard still needs work, such as understanding long, confusing stories or certain types of persuasion, but they have mapped out exactly how to fix these in future versions.
In short, Semalith v1.4 proves that you don't need a giant, slow brain to be a great safety guard. With the right training and a clever design, a smaller, faster model can be the perfect shield for AI systems, especially when they are helping people with money and banking, all while letting the good stuff through without a hitch.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.