SIEVE Audit desk
AuditSV-001 / FLAWEDBENCH
Fixture replay

Evaluator assurance / interactive replay

FlawedBench under test.

Challenge a known 20-task fixture, inspect each grader decision, and see how evidence changes the score you can responsibly claim.

Audit stateReady to run
A

Probe plan

Configure a deterministic browser replay. Select local service to call the installed auditor and persist a real run.

Replay mode. No backend request is made.

Canonical plan: 165 local grader calls, no model calls, $0.

Progress0 / 20No tasks probed
Budget0 / 200165 projected
Findings05 seeded defects
Undetermined0Evidence retained
Reported → trust-adjusted80% 80–80%Transparent sensitivity, not a CI
B

Task register

C

Live assurance plot

Wilson 95% / task level
Trust-adjusted band80–80%
Reported score80%
Trust bandReported score0%100%

Probe activity

0 events
  1. Start or step the audit to inspect the probe stream.
D

Task inspection

Exception register

Findings that change what 80% means.

NO FINDINGS YET

Run the probe battery to populate the evidence register.

Grader assurance

Correct and wrong answers stay separate.

Aggregate accuracy can conceal the direction of a grader defect. Sieve keeps false accepts and false rejects visible with their observed counts.

E

Probe confusion matrix

Awaiting audit
Grader passGrader fail Correct variants00 Wrong mutations00

Rows are the constructed truth class; columns are the grader decision. On the complete fixture: 35 true accepts, 3 false rejects, 9 false accepts, and 38 true rejects.

F

Per-task Wilson intervals

Rate uncertainty
Observed rate95% Wilson intervalFP is a lower bound

Assurance boundary

What this run can—and cannot—say.

VERIFIED

Regression behavior

FlawedBench deterministically contains 20 tasks and five seeded eval-infrastructure defects. The Python tests assert the exact verdicts.

LOWER BOUND

False-positive rate

Only attacks represented by the mutation library are counted. Unknown exploits are not in the denominator.

UNDETERMINED

Semantic truth

Without a trusted oracle, Sieve does not claim that an answer is semantically correct. It retains the evidence gap.

NOT SHIPPED

External benchmark claim

This is a synthetic regression fixture, not a real public-benchmark audit. Fleet mode and generated probes remain future work.