Can a clean scan still be dangerous?
Yes. Novel, multilingual, contextual and obfuscated attacks can evade deterministic patterns.
Paste a RAG document, user message or tool result. See the exact lines that attempt instruction override, secret access, data exfiltration, tool misuse or obfuscation—without sending the text anywhere.
Use this as an early warning layer, never as proof that content is safe.
This boundary reduces accidental instruction blending; it does not neutralise a capable attack by itself.
Do not automatically block or trust content from this score alone. Combine content isolation with least-privilege tools, destination allowlists, human approval and adversarial evaluation.
Yes. Novel, multilingual, contextual and obfuscated attacks can evade deterministic patterns.
Not here. A local first pass stays private, predictable and free. Model evaluation belongs in a capped test environment.
Keep untrusted data separate from instructions and never give the model more tool authority than the task needs.