Requirement 01
Take a Position · Conviction
- Commit to a clear view rather than surveying both sides
- A hedged non-answer earns no credit
- Judgement is the skill being measured
Can an AI agent reason like a hedge-fund analyst, graded deterministically against the moves the experts actually make?
01 · Overview
Kang is Hedge-Bench 1.0, a financial-reasoning benchmark of 102 on-the-job tasks grounded in the explicit reasoning traces of professional hedge-fund analysts. Each task casts an agent as an analyst with a corpus of primary documents (SEC filings, earnings calls, financials, press) and an open-ended theme, then grades its reasoning deterministically against verified expert moves. Everyone is building finance agents; no one can measure their reasoning. Existing benchmarks terminate in checkable answers and grade factual recall; Kang grades the argument itself: conviction, rigour, synthesis, and grounding.
02 · Key metrics
Three numbers that define the scope of Hedge-Bench 1.0.
Benchmark Tasks
102
5,112 curated · 20,448 sub-tasks
Evaluation Trials
6,528
8 frontier models · 8 trials each
Best pass@1
< 16%
far from saturated
03 · Methodology
What earns credit. Four requirements govern every graded answer.
Requirement 01
Requirement 02
Requirement 03
Requirement 04
04 · Grading pipeline
The judge
Check 01
Check 02
Check 03
05 · Results
Eight frontier models were run across all 102 environments, 8 trials each, for 6,528 evaluation trials (macro-averaged, pass@1).
The best model, Claude-Sonnet-4.6, captures under half the rubric at 1.92 / 4.0 and earns a perfect score on roughly 1 in 6 attempts. Judgement-heavy, forward-looking topics are hardest, with valuation and growth topping the difficulty chart. Quality is not reliability: GPT-5.5 reaches about 88% of Sonnet's quality at less than half the hallucination rate.
Difficulty by topic category, and theme vs move coverage per model.
06 · Model Leaderboard
Models
8
Environments
102
Trials
6,528
Best Dense
1.92/4
| Model | Provider | Dense 0-4 | Themes | Moves |
|---|
07 · Where Kang sits
Existing finance benchmarks grade checkable answers or process. Kang is the one grading open-ended reasoning deterministically.
The landscape
The moat
08 · Resources
GitHub
kang-samples
github.com/EtharaOrion/kang-samples
Hugging Face
kang
huggingface.co/datasets/ethara/kang
Ethara builds held-out, reproducible agent benchmarks and datasets for frontier model labs and financial institutions.