Back to prompt

Execution transcripts

Every recorded model call for Extract to the end of the input, then judge the addressee

Runs
1
Model calls
3
Models
1
Cost to run
$0.0033

One run is not listed here

This prompt was also run against a held-out case: the same problem with the same declared checks, but an input that is not published. Publishing it would let a prompt be written to fit the exact test it is judged on, which is the one thing a test cannot survive. Its result is counted in the totals on the prompt page.

Models used: gemini-3.5-flash-lite

gemini-3.5-flash-lite14 Aug 2026, 00:59 UTC · 3 calls · $0.004315