Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
Researchers have developed Semalith v1.4, a 184M-parameter safety classifier that can detect prompt injection, general harm, and financial-services regulatory compliance in a single forward pass. This classifier outperforms Llama-Guard-3-8B on 22 held-out benchmarks, incl…
Intelligence analysis by Llama

Semalith v1.4 is a state-of-the-art safety classifier that can detect prompt injection, general harm, and financial-services regulatory compliance in a single forward pass, outperforming Llama-Guard-3-8B on 22 held-out benchmarks.
Imagine you have a super-smart computer that can understand what people are saying. But sometimes, people might try to trick the computer into doing something bad. Semalith v1.4 is a special tool that helps keep the computer safe by detecting when someone is trying to trick it.
Analysis
A Safety Classifier for the Modern Era
Semalith v1.4 is a 184M-parameter safety classifier that has been designed to address the complex task of detecting prompt injection, general harm, and financial-services regulatory compliance in a single forward pass. This is a significant achievement, as existing open guardrails have been unable to handle these tasks simultaneously.
Why Semalith v1.4 Matters
The development of Semalith v1.4 is significant for deploying large language models in financial-services and agentic settings. These settings require safety classifiers that can handle multiple tasks simultaneously, and Semalith v1.4 is the first classifier to achieve this.
The Road Ahead
While Semalith v1.4 is a significant achievement, there are still several challenges that need to be addressed. The classifier has been shown to be effective on 22 held-out benchmarks, but it is not clear how it will perform in real-world settings. Additionally, the classifier has been shown to have several weak spots, which need to be addressed in future work.
Key points
- Semalith v1.4 is a 184M-parameter safety classifier that can detect prompt injection, general harm, and financial-services regulatory compliance in a single forward pass.
- The classifier outperforms Llama-Guard-3-8B on 22 held-out benchmarks, including 7/7 prompt-injection evaluations, at 44x fewer parameters.
- Semalith v1.4 has been shown to be effective on 22 held-out benchmarks, but it is not clear how it will perform in real-world settings.
- The classifier has been shown to have several weak spots, which need to be addressed in future work.
If Semalith v1.4 is deployed in financial-services and agentic settings, it could lead to significant improvements in safety and compliance. The classifier's ability to detect prompt injection, general harm, and financial-services regulatory compliance in a single forward pass could make it a game-changer for these industries.
However, there are still several challenges that need to be addressed before Semalith v1.4 can be widely deployed. The classifier has been shown to have several weak spots, and it is not clear how it will perform in real-world settings.


