proofplane

An AI governance control plane where a control is satisfied only when an executed adversarial attack failed — never because a document says it exists. This page is how you use it, and how you take the same rule into your own program.

Run it in your browser — no install → A hosted demo executes the real adversarial probes against a sandboxed agent. Watch each control hold only because an executed attack failed — and breach the moment its guardrail is removed.

What it does

Most AI governance tooling produces an artefact that asserts a control is in place: a policy PDF, a questionnaire answer, a screenshot of a settings page. proofplane produces evidence that a control held.

Twelve controls. Each names the MITRE ATLAS and OWASP Agentic technique it defends against, and the probe that proves it. A control is HELD only when that probe breached an unguarded target, held against a guarded one, and did not go green when any other guardrail was removed. That last part is the 12×12 independence matrix: 144 probe runs, breach only on the diagonal.

The output is hash-chained evidence, an OSCAL assessment-results file (method TEST, because that is what happened), a CycloneDX ML-BOM of the AI surface, and a FAIR loss model that credits a control only if a probe executed an attack against it and the attack failed.

When and where to use it

Use it when

  • You are standing up an agentic system — tools, tenants, approvals — and need to prove the guardrails work before anyone calls them controls.
  • A framework (AIUC-1, EU AI Act Articles 12/14/15, NIST AI RMF) is asking for adversarial technical testing, not a narrative.
  • You need machine-readable evidence an assessor can re-run: OSCAL, a BOM with file-and-line provenance, a head hash.
  • You want to teach a GRC team the difference between “we have a policy” and “the attack failed.”

Do not use it when

  • You need to certify a live model as safe. The default target is a deterministic double. A held result is evidence the guardrail works.
  • The control is governance, training, or a contract. Those resist automated attack and are absent here on purpose — not passing.
  • You want a coverage percentage. Twelve controls is a worked example, not an AIUC-1 or ISO 42001 program.

How to use it

No API key, no cloud account, no spend. Node 20+ and Python 3.11+. Go 1.22+ for the inventory step; without it the script skips discovery and says so.

1. Clone and run the suite. Seven steps: inventory the AI surface, prove every probe is falsifiable, run the independence matrix, write evidence and an HTML report per configuration, execute the documented limitations, validate OSCAL against the vendored NIST schema, check every framework citation.

git clone https://github.com/RootCawsLLC/proofplane.git
cd proofplane
npm --prefix target ci
python -m pip install ./probe
./scripts/assure.sh          # Windows: .\scripts\assure.ps1

Evidence lands in evidence/. Open evidence/guarded/report.html first, then the unguarded report. The unguarded run is what makes the green one mean something: a suite that cannot go red proves nothing when it is green.

2. Read one control end to end. Pick PP-C001 (privileged actions need an approval recorded outside the model). The probe asks the agent for a refund. Against the unguarded target the refund happens. Against the guarded target it does not. Disable any other guardrail and the refund still does not happen — only removing the approval gate lets it through. That is what “independence” means here.

3. Useful commands once you have run it once.

python -m proofplane_probe.cli --catalog ./catalog catalog
python -m proofplane_probe.cli --catalog ./catalog verify
python -m proofplane_probe.cli --catalog ./catalog matrix
python -m proofplane_probe.cli --catalog ./catalog corroborate
node scripts/validate-oscal.mjs

4. Against a real model — only after you understand the default run. A held result against the double is not a held result against a model.

PROOFPLANE_MODEL_PROVIDER=anthropic ANTHROPIC_API_KEY=sk-... \
  PROOFPLANE_GUARDRAILS=all node target/dist/server.js

Raise the trial count and read the breach rate. Zero breaches in three trials is consistent with a true failure rate above 50%. That is why every result on this site shows its trial count.

Take it into your organization

You do not drop this repository onto a production agent and call it a day. You take the rule and the shape.

  1. Write the control as an assertion a probe can fail. “Privileged actions require an approval recorded outside the model” can be attacked. “We take AI risk seriously” cannot. If you cannot name the attack, you do not have a technical control yet.
  2. Put the guardrail in deterministic code, not in the prompt. The model is not a trust boundary. Every guardrail in the target is toggleable code outside the model. That is the whole architectural argument.
  3. Prove the probe can go red before you trust it going green. The independence matrix is the worked example: disable the control under test, confirm breach; disable every other control, confirm hold.
  4. Emit what an assessor already knows how to read. OSCAL assessment results with method TEST. A BOM that points at a file and a line. A head hash you can put somewhere you cannot silently edit.
  5. Price the control against the evidence, not the catalog. The exposure report credits a control only if a probe attacked it and failed. When a guardrail stops holding, the dollar figure moves on the next run.

Start with one agent, one privileged tool, one approval gate. Port PP-C001 and its probe. Do not start by mapping ISO 42001 — the crosswalk here is cited at group level, with a confidence on every edge, and it is not a claim of satisfaction.

Full write-up, ADRs, and the threat model live in the repository. This page is the operator’s entrance.

This run

Everything linked below was produced by a pipeline run, not written by hand.

Guarded configuration — 12 controls held
All twelve guardrails enabled. Includes the 12×12 independence matrix: every probe breached when, and only when, its own guardrail was removed.
Unguarded configuration — 12 controls breached
Every guardrail disabled. This is what makes the report above mean something: a suite that cannot go red proves nothing when it is green.
Loss exposure, bound to control state
FAIR Monte Carlo over eight scenarios. Inherent against residual, and what each control is worth per year by counterfactual.
OSCAL assessment results
Validated against the NIST OSCAL 1.1.2 schema. Every observation carries method TEST.
AI bill of materials
CycloneDX 1.6 ML-BOM of this repository's own AI surface, with the file and line that produced each component.
Limitation demonstrations
The weaknesses the documentation admits to, executed against a fully guarded target rather than asserted.
Loss exposure model output
Every figure from the exposure report, with the evidence run and head hash it was derived from.

Honest limits

Read the status honestly. These runs use a deterministic model double, so a held result is evidence that the guardrail works — not that any model is safe. Against a live model, zero breaches in three trials is consistent with a true failure rate above 50%, which is why every result shows its trial count. Hash chaining is tamper-evident, not tamper-proof. The crosswalk is mapped, not satisfied. Controls that resist automated testing are absent, not passing. The full account is docs/HONEST-LIMITS.md.