One signed binary. Every feature compiled in. Free to run. Install Crowkis →
← back to the Roost
featuresJune 22, 2026· 3 min read

CGUARD: an input guardrail that survives leetspeak and zero-width tricks

Prompt injection rarely arrives in plain English. CGUARD normalizes the evasion first, whitespace, leetspeak, zero-width characters, then scans for jailbreaks, overrides, and system-prompt exfiltration.

The naive prompt-injection filter loses to a five-minute workaround: insert a zero-width space, swap an 'o' for a '0', pad with newlines, and the regex that caught 'ignore previous instructions' sails right past 'ig​nore prev1ous 1nstructions'. CGUARD assumes the attacker knows this, so it normalizes before it matches, collapsing whitespace, folding leetspeak, stripping zero-width characters, and only then runs the scan.

In plain words: Attackers disguise prompt injections with odd spacing and character swaps. CGUARD undoes the disguise first, then checks, so the trick that beats a simple filter doesn't beat this one.

What it scans for is the real catalogue: DAN-style jailbreaks, developer-mode prompts, instruction overrides detected by verb-plus-noun co-occurrence rather than fixed strings, and attempts to exfiltrate the system prompt. The return is structured, a verdict, a category, and the matched span, so your app can log the category, show the user a clean refusal, and route the rest onward.

the write-trust pipeline
  1. 1
    candidate write
  2. 2
    coherence · 0.30
  3. 3
    content · 0.10
  4. 4
    source trust · 0.30
  5. 5
    isolation · 0.15
  6. 6
    neighbourhood · 0.15
  7. 7
    composite ≥ 0.75?
  8. 8
    accepted
  9. 9
    refused + ledger entry

Five stages score every write before it can ever be served.

It runs as a command, CGUARD, which means it composes anywhere: in front of a cache write, in front of a model call, or as a standalone gate in a pipeline that isn't even using Crowkis for caching yet. No model call, no egress, the detection is deterministic and local, so it costs microseconds and leaks nothing.

The bottom line

A guardrail is only worth shipping if it assumes an adversary, not a typo. CGUARD is built for the person actively trying to break your agent, which is exactly the person a string-match filter was never going to stop.