Real results from Entry's own agent harness -- the same system prompt, tools, and gateway-routed models every real chat uses. No cherry-picked transcripts: every task is graded by an independent script, not the model itself.
HumanEval subset (humaneval-subset-20-v1), 20 fixed tasks from the canonical OpenAI HumanEval dataset. Last run completed 2026-09-18. Cost shown is the real gateway-metered spend for running this subset, not a per-token rate.