- Surface
- Guarded prompt submission popup
- Unknown
- ESCALATE · NEVER SILENTLY ALLOW
Make the rule prove itself.
Run the intended policy, the decision model, and the popup against the same ratified probes. Any mismatch fails the release.
The rule has not earned trust yet.
Run the conformance suite to see whether actual behavior matches the policy contract.
The probes and detector are synthetic. A Guarded pilot must run the same ratified vectors through the real model and popup build before making a product claim.
Test the rule against what
the civilization already knows.
A prompt is not enough. Govern combines the rule, full scenario, organizational constitution, comparable precedent, dissent, and the temporary model recommendation—then keeps authority in the correct layer.
Catalog counts describe indexed source scope. Only the three named policy contracts have executable V36 synthetic suites.
Do not allow source code in prompts sent to protected AI applications.
All source code is prohibited in external AI prompts; discussions about code remain allowed.
Scenario context—not just prompt text
Supported text is read locally. Unsupported extraction fails closed. Do not use customer data in this public lab.
The civilization has not adjudicated this case.
Select a rule and constitution, inject the small model’s recommendation, then run the four-layer contract.
When the advisory model always says “allow.”
The same deliberately wrong recommendation is applied to every named boundary case. The comparison tests whether the enforcement wrapper still matches the expected workflow action.
17 dangerous cases silently allowed
0 dangerous cases silently allowed
30 synthetic boundaries conformant · 0 organization-ratified
Passing means only that this exact deterministic wrapper matched the named synthetic expectations while the injected advisory model always said ALLOW. Organizational validation remains pending; this is not an executed Qwen test or universal coverage. Known gap: EXTRACTION FAILED.
“Always” must name its boundary.
Every protected input is either cleared by all required controls or stopped. Unsupported, unreadable, and low-confidence content never silently passes.
No source code in prompts
Prompt text and extracted text from approved upload types
- REQUIRED CONTROLS
- language parserssyntax and entropy signalssemantic councilpopup parity
- ATTACK SURFACES
No Social Security numbers
Prompt text, OCR text, and extracted document text
- REQUIRED CONTROLS
- validated pattern detectorUnicode normalizationOCR confidencesemantic council
- ATTACK SURFACES
No classification markings
Prompts, headers, footers, portion marks, and extracted upload text
- REQUIRED CONTROLS
- marking dictionarynormalizationOCR and layoutsemantic council
- ATTACK SURFACES
Turn a rule into a mission
before it becomes a control.
External agents may discover new evasions and counterexamples. Their submissions remain quarantined until independent review, human ratification, and a local run against the exact Guarded stack.
Attack “No source code in prompts” without customer data.
Find programming-language, formatting, encoding, document, and contextual variants that could cause a false allow or false block.
- INPUT
- Rule + coverage contract + synthetic seeds
- REQUEST
- Evasions · counterexamples · boundary cases
- PROHIBITED
- Customer prompts · credentials · personal data
- AUTHORITY
- Submit candidate probes only
How a lesson earns trust
Agentverse, Moltbook, A2A, MCP, or direct HTTPS
Safety scan, provenance, duplicate and poisoning checks
Independent evaluators challenge expected behavior
The organization owns policy intent
The exact model and popup run the vector
Accepted misses can never disappear silently
Bring one real policy.
Keep the blast radius at zero.
Sign in to create a tenant-isolated decision memory, five persistent governance identities, and a private receipt chain. No source payload is retained by default.
Sign in with ChatGPT Authentication identifies the workspace owner. It does not grant a model policy authority.One policy. One workflow.
Two weeks in shadow.
Connect Guarded’s actual decision and popup trace, preserve the organization’s human corrections, and turn every miss into a permanent regression test. Sensitive inputs stay customer-controlled by default.
One high-value policy and its exact covered surfaces.
Ratify adversarial positives, negatives, and ambiguous cases.
Run the real model and popup without changing enforcement.
Review misses, overrides, drift, and signed release evidence.