RISC-V · RV32IM · single-issue · FPGA-grounded

HWE Bench

An unbounded benchmark for LLM hardware engineering. Large language models design RISC-V CPUs from scratch. Every design must first pass a full battery of formal correctness proofs, so buggy CPUs are thrown out. The ones that survive are then scored by how fast they would actually run on a physical FPGA. See on GitHub

Thesis SWE-bench tops out at 100%. HWE Bench doesn't have a top.
The fitness number reflects an actual microarchitecture, and microarchitecture has room to grow as long as models keep finding it.
Speed vs size

Score × Area

Smaller and faster than human baseline 400 600 800 1000 0 2.5k 5k 7.5k 10k 12.5k Area · LUT4 count (← smaller is better) Fitness · CoreMark iter/s (↑ better) Human reference (VexRiscv 370) OpenAI Google Moonshot (Kimi) Anthropic VexRiscv (human) V0 baseline Claude Opus 5.5 xhigh GPT-5.5 xhigh GPT-5.6 Terra GPT-5.4 xhigh GPT-5.6 Luna GPT-6 Astra max GPT-5.6 Sol GPT-5.5 high GPT-6 Sol xhigh GPT-5.5 medium Kimi K2.6 GPT-5.4 mini Gemini 3.5 Flash Gemini 3.1 Pro VexRiscv (human) V0 baseline
Vertical axis: CoreMark fitness (how fast the CPU runs the benchmark). Horizontal axis: chip area (LUT4 count, basically how many gates the design uses on the FPGA). One point per model's best run. VexRiscv (3,402 LUT4 · fitness 370) is the human-engineered reference. Up and to the left is the goal: faster chip, smaller chip.
Capability over time

Model release date × peak HWE score

OpenAI Google Moonshot (Kimi) Anthropic OLS trend Feb 2026 Mar Apr May Jun Jul Aug Sep Oct 400 600 800 1000 V0 baseline · 283 Public model release date Peak HWE fitness (↑ better) Human reference (VexRiscv 370) Claude Opus 5.5 xhigh GPT-5.5 xhigh GPT-5.6 Terra GPT-5.4 xhigh GPT-5.6 Luna GPT-6 Astra max GPT-5.6 Sol GPT-5.5 high GPT-6 Sol xhigh GPT-5.5 medium Kimi K2.6 GPT-5.4 mini Gemini 3.5 Flash Gemini 3.1 Pro
Each point is one model configuration's best completed HWE Bench rep; reasoning-effort variants share their underlying model family's public release date. The dashed fit is descriptive, not a forecast. Release dates come from the OpenAI Codex notes, Gemini API changelog, Claude Code changelog, and Kimi K2.6 announcement.
Leaderboard

Peak fitness per model

Best of recorded reps per model · 38 reps total · VexRiscv human reference in red · baseline V0 in italic
# Model Reps Best Δ% Mean ± std Area (LUT4) Fmax (MHz)
1 claude-opus-5_5_xhigh 1/1 983.24 +247.7% 983.2 3.1k 302
2 gpt-5_5_xhigh 3/3 525.04 +85.6% 468.3 ± 52.8 5.5k 220
3 gpt-5_6-terra 3/3 515.70 +82.3% 442.3 ± 55.8 10.5k 209
4 gpt-5_4_xhigh 3/3 513.84 +81.7% 485.8 ± 28.1 10.1k 203
5 gpt-5_6-luna 3/3 480.90 +70.0% 452.0 ± 29.0 10.1k 209
6 gpt-6-astra_max 3/3 474.27 +67.7% 424.8 ± 36.2 6.2k 199
7 gpt-5_6-sol 1/3 470.80 +66.5% 470.8 10.2k 200
8 gpt-5_5_high 3/3 461.87 +63.3% 430.2 ± 23.0 9.8k 187
9 gpt-6-sol_xhigh 1/1 435.24 +53.9% 435.2 5.7k 189
10 gpt-5_5_medium 3/3 431.58 +52.6% 423.5 ± 11.2 7.8k 201
11 kimi-k2_6 2/3 396.13 +40.1% 339.5 ± 8.3 9.9k 166
12 gpt-5_4-mini 3/3 395.53 +39.9% 362.3 ± 23.7 10.2k 187
13 VexRiscv (human ref) n/a 370.00 +30.8% n/a 3.4k 144
14 gemini-3_5-flash 0/1 359.04 +27.0% n/a 13.8k 125
15 gemini-3_1-pro 3/3 354.73 +25.4% 339.4 ± 12.6 10.2k 150
16 static 2/2 282.82 +0.0% 282.8 9.6k 127
17 baseline V0 (fixture) n/a 282.82 n/a n/a 9.6k 127

The VexRiscv row is the human-engineered reference, a well-known open-source RV32IM CPU synthesized on the same FPGA used for the benchmark. 12 of the LLM-generated designs beat it. See the methodology page for the full procedure.

Single repetition (n=1): claude-opus-5_5_xhigh. Each of these rows is one run; repeatability is unmeasured, so rank differences involving them are untested.

Why unbounded

SWE-bench saturates. HWE Bench doesn't.

Most LLM benchmarks have a fixed ceiling. SWE-bench tops out at 100% issue-resolution. Multiple-choice evals approach 99%. Once a model lands at the ceiling, every subsequent model gets the same score, and the benchmark stops being useful for tracking capability.

HWE Bench has no ceiling. Fitness is the CPU's actual speed running CoreMark on a real FPGA, operating frequency times instructions-per-cycle (Fmax × IPC for the technically inclined). There's no theoretical maximum: a smarter microarchitecture always scores higher. As long as models keep finding new tricks (deeper pipelines, smarter branch predictors, restructured ALUs), the leaderboard keeps moving.

Empirically: the current best is 983.24 iter/s (one repetition, n=1; untested), +247.7% over the V0 baseline core, and clear of the VexRiscv human reference. There is no theoretical ceiling, and within current budgets the curve has not saturated.

Trajectory

Fitness over rounds, best rep per model

400 600 800 1000 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 baseline 283 Round (3 hypothesis slots each) Best fitness so far Human reference (VexRiscv 370) Claude Opus 5.5 xhigh 983 · R14 GPT-5.5 xhigh 525 · R10 GPT-5.6 Terra 516 · R9 GPT-5.4 xhigh 514 · R8 GPT-5.6 Luna 481 · R7 GPT-6 Astra max 474 · R14 GPT-5.6 Sol 471 · R15 GPT-5.5 high 462 · R14 GPT-6 Sol xhigh 435 · R14 GPT-5.5 medium 432 · R12 Kimi K2.6 396 · R8 GPT-5.4 mini 396 · R15 Gemini 3.5 Flash 359 · R1 Gemini 3.1 Pro 355 · R5
Running max of CoreMark fitness across the 15 hypothesis rounds for each model's best-performing rep. Lines step up when a winning hypothesis lands and stay flat otherwise. VexRiscv's human-reference fitness is the red dashed line; the baseline V0 core is the gray dashed line.