Verification study / Working paper

The results matter.
So do the misses.

A public record of what the checks caught, where they fell short, and what still needs to be tested.

Explore below

Read past
the headline.

The experiments tested checks against deliberately faulty work. These are benchmark results, not a promise of performance on your project.

Inspect the method

Static review

18 / 30

Defects caught in static review

Bistro · V1

21–24 / 30

Defects caught in round two

Toolshed · V1

6–12 / 54

Defects caught in round two

Solid bar: lower reported result. Pale end: reported range.

Show the limits.
Keep asking questions.

01

A bounded result

No false blocks were observed on the tested correct Bistro and Toolshed apps. This is not a guarantee for every project.

02

A working paper

The study has not been peer reviewed. The full method is public so you can inspect the benchmark and its limits.

03

Work still ahead

Round three was frozen but not run. A Codex comparison was not frozen, and round four was only proposed.

Read the full study