KANG

HEDGE-BENCH 1.0 · FINANCIAL-REASONING BENCHMARK

Can an AI agent reason like a hedge-fund analyst, graded deterministically against the moves the experts actually make?

Kang is Hedge-Bench 1.0, a financial-reasoning benchmark of 102 on-the-job tasks grounded in the explicit reasoning traces of professional hedge-fund analysts. Each task casts an agent as an analyst with a corpus of primary documents (SEC filings, earnings calls, financials, press) and an open-ended theme, then grades its reasoning deterministically against verified expert moves. Everyone is building finance agents; no one can measure their reasoning. Existing benchmarks terminate in checkable answers and grade factual recall; Kang grades the argument itself: conviction, rigour, synthesis, and grounding.

  • Realistic: on-the-job tasks from working hedge-fund analysts
  • Deterministic: open-ended reasoning graded against verified expert moves
  • Reasoning-first: argument and judgement, not recall or retrieval
  • Uncontaminated: proprietary traces no model saw in pretraining
  • Dense 0-4 rubric scored on grounding, coverage, and synthesis

Three numbers that define the scope of Hedge-Bench 1.0.

Benchmark Tasks

102

5,112 curated · 20,448 sub-tasks

Evaluation Trials

6,528

8 frontier models · 8 trials each

Best pass@1

< 16%

far from saturated

What earns credit. Four requirements govern every graded answer.

Requirement 01

Take a Position · Conviction

  • Commit to a clear view rather than surveying both sides
  • A hedged non-answer earns no credit
  • Judgement is the skill being measured

Requirement 02

Engage Counter-Evidence · Rigour

  • Confront the strongest opposing data in the corpus
  • Not just the evidence that supports the thesis
  • Graded on coverage of the required moves

Requirement 03

Reconcile Conflicts · Synthesis

  • Fuse opposing data points into one unified conclusion
  • An explicit reconciliation sentence is required
  • Reserved for the top score of 4.0

Requirement 04

Cite Every Claim · Grounding

  • Every factual claim cites its source file inline
  • Ungrounded facts are discarded by the grader
  • A move resting on a flagged fact is forfeited

The judge

A three-check judge turns an open-ended answer into a deterministic 0-4 score.

Check 01

Grounding

  • Extract every factual claim from the answer
  • Verify each against its cited source; flag the unverifiable
  • A move resting on a flagged fact is forfeited

Check 02

Coverage

  • Per theme, which lettered required moves are hit
  • Matched by concept, not vocabulary
  • A theme is covered when grounded moves reach its threshold

Check 03

Synthesis

  • One explicit reconciliation of opposing data
  • Unlocks the top score of 4.0
  • Pass = a perfect 4.0 / 4.0

Eight frontier models were run across all 102 environments, 8 trials each, for 6,528 evaluation trials (macro-averaged, pass@1).

The best model, Claude-Sonnet-4.6, captures under half the rubric at 1.92 / 4.0 and earns a perfect score on roughly 1 in 6 attempts. Judgement-heavy, forward-looking topics are hardest, with valuation and growth topping the difficulty chart. Quality is not reliability: GPT-5.5 reaches about 88% of Sonnet's quality at less than half the hallucination rate.

Fig. 1: Mean dense score (0-4) across all 102 environments. The best model captures under half the rubric; the field spans 1.92 down to 0.75.

Difficulty by topic category, and theme vs move coverage per model.

Fig. 2: Mean dense score by topic category. Valuation (1.61) and Growth (1.48) are hardest; Risk (1.23) is most tractable.
Fig. 3: Theme coverage vs move coverage (%) per model. No model clears 57% on either axis.
1.92 BEST DENSE SCORE / 4.0
< 16% BEST PASS@1 1 in 6
8 MODELS TESTED 6,528 trials
1.92 / 4.0 Claude-Sonnet-4.6 · best mean dense score

Models

8

Environments

102

Trials

6,528

Best Dense

1.92/4

Loading...
Model Provider Dense 0-4 Themes Moves

Existing finance benchmarks grade checkable answers or process. Kang is the one grading open-ended reasoning deterministically.

The landscape

FinQA / TAT-QA
checkable answers
FinanceBench
factual recall
FAB v2
process-graded
Kang
open-ended reasoning

The moat

Analyst network
perpetual IP licences
Grading
deterministic, open-ended
Tasks
uncontaminated
Backlog
5,112 curated

Need a hard reasoning eval or expert-aligned preference data?

Ethara builds held-out, reproducible agent benchmarks and datasets for frontier model labs and financial institutions.

Get in touch