POC Results

Measured judgment performance in bounded POCs.

The headline results come from different evaluation scopes. Each metric is shown with what it measures — and what it does not prove.

Bounded POCMissed-risk proxyRework loopsEvidence boundary
Two evidence scopes

Do not collapse different POCs into one universal performance claim.

PUBLIC v12 POC
0.0249%

Missed-risk proxy

Fraction of bounded test risk cases not captured under the defined public POC proxy.

Does not imply: production risk rate · general AI accuracy · safety certification · uptime.
SYNTHETIC WORKFLOW POC
10 → ~2

Rework loops to convergence

Reduction observed in a separate synthetic workflow comparison.

Does not imply: guaranteed 5× improvement · universal convergence · production benchmark.
Public v12 POC

Miss less under the defined bounded proxy.

The public preprint and visual extracts below provide the detailed metric context.

What the POCs establish

A measurable judgment-control question — not universal superiority.

The POCs do not prove universal performance. They test whether judgment-level control can measurably reduce misses and repeated rework under defined conditions.
MEASUREMissed casesCan relevant bounded risk cases be captured more consistently?
MEASURERepeated reworkDoes the same unresolved condition keep sending the workflow around again?
BOUNDARYReview loadDo improvements merely push everything to human review, or remain selective?
BOUNDARYScopeResults remain tied to the defined dataset, proxy, workflow, and evaluation conditions.

Use your own bounded baseline next.

The useful commercial question is whether the same judgment-level effect appears in one real workflow or control sequence.