I’m jumping on the Jev bandwagon this week. I’ve been testing different ways to stop prompt injection attacks from reaching AI agents (repo below). The latest one I tried is Jev, TypeSafe AI’s decision model, available through OpenRouter.
Jev doesn’t generate text. You give it content plus typed questions (”𝘐𝘴 𝘵𝘩𝘪𝘴 𝘢 𝘱𝘳𝘰𝘮𝘱𝘵 𝘪𝘯𝘫𝘦𝘤𝘵𝘪𝘰𝘯?”, ”𝘏𝘰𝘸 𝘴𝘦𝘷𝘦𝘳𝘦 𝘸𝘰𝘶𝘭𝘥 𝘪𝘵 𝘣𝘦 𝘪𝘧 𝘰𝘣𝘦𝘺𝘦𝘥?”), and it returns probabilities. That makes it a natural fit for a security gate.
I ran it against the same 500-prompt Kaggle dataset (250 malicious, 250 benign) that I used for local LLMs and Microsoft’s MAF-FIDES:
- 𝗦𝗽𝗲𝗲𝗱: 500 prompts in 56 seconds, about 100 ms per prompt. My local LLM runs took an age by comparison - enough time to go an get a coffee.
- 𝗖𝗼𝘀𝘁: 2.4 cents for all 500 prompts, roughly 5 cents per 1,000.
- 𝗢𝘂𝘁 𝗼𝗳 𝘁𝗵𝗲 𝗯𝗼𝘅: 92.2% accuracy with zero false positives.
- 𝗧𝘂𝗻𝗲𝗱: 100% (500/500) with zero false positives, by also flagging anything Jev rated LOW severity or above.
(I have since tested it against a 5,000 prompt injection dataset - still amazing performance and injection identification results.)
Almost everything it missed out of the box was a bare shell command, such as dd or a fork. Jev wasn’t sure these counted as ”𝘮𝘢𝘯𝘪𝘱𝘶𝘭𝘢𝘵𝘪𝘰𝘯”, but it still rated them as dangerous. Using its own severity score caught every one of them, at no extra cost.
𝗖𝗮𝘃𝗲𝗮𝘁: I tuned the rule on the same Kaggle dataset I tested it on, so 100% needs confirming on unseen data - hence the 5,000 prompt test.
For a fast, cheap, explainable-by-numbers first line of defense, it’s very, very impressive - GitHub Repo :, Test Results :
