econ-eval
How much do you give up by using a cheap model on the work an economist actually does? Frontier and cheap LLMs scored on trade arithmetic, finance formulas, pipeline code, refereeing, and policy writing in English and Bangla.
DatasetHarness and tasks on GitHub
Leaderboard A: 20 tasks, 12 models, one judge
Five samples per task, 100 completions per model. Objective tracks are graded deterministically; the judged tracks were graded by Claude Opus 5 in one pass over all twelve models. The whisker is a 95% bootstrap confidence interval over tasks.
Score with 95% confidence interval
Mean over 20 tasks of each task's mean over 5 samples. Hover a point for the head-to-head against Opus 4.8.
Quality against cost per correct answer
Cost is API list price times measured tokens; log scale. Opus 4.8 and Muse Spark ran through CLI harnesses that inflate their input tokens, so their cost is an upper bound.
Table
Per track
Coding and quantitative are saturated at this task count. The ranking comes from the writing and reasoning tracks; the two Bangla writing tasks and one Balassa RCA task carry most of the discrimination.
Leaderboard B: 50 tasks, 5 models
The 20 tasks above plus 30 daily-work tasks (bilateral shares, tariff-weighted averages, SRISK, Eisenberg-Noe clearing, CET1, term premium, HS normalisation, referee comments, an agent brief, an X post, a recruiter reply, a plain-Bangla explainer). Five frontier-tier models, mostly through subscription CLIs.
scores.csv.
- Mixed judges. The judge for this run was GPT-6 Astra, itself a contestant. DeepSeek V4.1 Flash and Muse Spark keep their Opus-5-judged rows from Leaderboard A on the 20 original tasks, so their judged scores mix two judges.
- Incomplete Fable row. Claude Fable 5.1 covers of 50 tasks; the run hit the CLI credit wall. Its sign test uses only the tasks it completed.
- No price for two models. GPT-6 Astra and Gemini 3.8 Flash are served only through subscriptions, so no cost is shown for them.
Score with 95% confidence interval
Mean over tasks completed of each task's mean over 5 samples. Hover for the head-to-head against GPT-6 Astra.