Jev in production › Moderation and content filtering

JevPromptShield

Scores incoming text with Jev to flag prompt injection before an agent acts; a test on 21 published Gandalf jailbreak prompts blocked 6 of 21 (author).

Open on GitHub ↗

6 of 21measured, as published by the source
Use
Moderation and content filtering
Industry
Security
Form
Open-source tool
Stage
In production
Listed
2026-10-09
Found via
github
Repository
JamesANZ/JevPromptShield
Stars
202
Forks
46
Last push
2026-10-08
Language
TypeScript
License
none stated

The README opens with

A prompt can tell a coding agent to drop its task and hand the secrets over. A pasted page can hide the same instruction in what looks like documentation. The agent can then propose a shell command that wipes a disk, force-pushes main, or pipes a downloaded script into bash.

JEV Shield stands outside the model and checks both moments. TypeSafe Jev scores the text. Ordinary code turns that score into a block, a question, or a pass. The agent does not get a vote.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for moderation and content filtering