econ-eval

How much do you give up by using a cheap model on the work an economist actually does? Frontier and cheap LLMs scored on trade arithmetic, finance formulas, pipeline code, refereeing, and policy writing in English and Bangla.

Models evaluated
Tasks
Graded completions
Transcripts published

Leaderboard A: 20 tasks, 12 models, one judge

Five samples per task, 100 completions per model. Objective tracks are graded deterministically; the judged tracks were graded by Claude Opus 5 in one pass over all twelve models. The whisker is a 95% bootstrap confidence interval over tasks.

Score with 95% confidence interval

Mean over 20 tasks of each task's mean over 5 samples. Hover a point for the head-to-head against Opus 4.8.

Quality against cost per correct answer

Cost is API list price times measured tokens; log scale. Opus 4.8 and Muse Spark ran through CLI harnesses that inflate their input tokens, so their cost is an upper bound.

Table

Per track

Coding and quantitative are saturated at this task count. The ranking comes from the writing and reasoning tracks; the two Bangla writing tasks and one Balassa RCA task carry most of the discrimination.

Leaderboard B: 50 tasks, 5 models

The 20 tasks above plus 30 daily-work tasks (bilateral shares, tariff-weighted averages, SRISK, Eisenberg-Noe clearing, CET1, term premium, HS normalisation, referee comments, an agent brief, an X post, a recruiter reply, a plain-Bangla explainer). Five frontier-tier models, mostly through subscription CLIs.

Read with three caveats, all visible in the dataset's scores.csv.
  1. Mixed judges. The judge for this run was GPT-6 Astra, itself a contestant. DeepSeek V4.1 Flash and Muse Spark keep their Opus-5-judged rows from Leaderboard A on the 20 original tasks, so their judged scores mix two judges.
  2. Incomplete Fable row. Claude Fable 5.1 covers of 50 tasks; the run hit the CLI credit wall. Its sign test uses only the tasks it completed.
  3. No price for two models. GPT-6 Astra and Gemini 3.8 Flash are served only through subscriptions, so no cost is shown for them.

Score with 95% confidence interval

Mean over tasks completed of each task's mean over 5 samples. Hover for the head-to-head against GPT-6 Astra.

Table

Per track