Execution evidence

Read 7 transcripts
Passed
4 of 4
Model calls
8
Models
3
Cost to run
$0.0090

Wilson lower bound 51%, which is the pass rate this many runs can actually support. It rises as more runs agree.

One of these runs is against a held-out case: same problem, same declared checks, an input not published so a prompt cannot be written to fit it. It counts towards the numbers above and its transcript is not shown.

Judged against the test cases declared on this problem, which were written before these outputs were seen. Open a transcript to read exactly what was sent, what came back, and which assertion decided it.

Version History

1 version
  1. 1
    Initial versionlatestDev

New versions are created automatically whenever you edit the system prompt, user template, model, or parameters.