HWE Bench
Per-model detail
What each model actually did.
Below: the per-rep outcomes for every model run on HWE Bench so far, plus the
accepted-improvement hypotheses each rep produced, verbatim titles, fitness,
LUT4, and Fmax. The hypothesis titles are exactly what the agent wrote.
claude-opus-5_5_xhigh
claude opus 5 5 xhigh
Claude Opus 5.5 via Claude Code 2.1.282 · reasoning effort xhigh · subscription (OAuth) login · 15 rounds × 3 hypotheses per round. One repetition (n=1): repeatability has not been measured, and any ranking against other configurations is untested.
Isolation: no user settings, plugins, hooks, MCP servers, connectors or auto-memory. Bash ran in a sandbox with no network, writes confined to the rep's clone and reads of the operator's home directory denied. Web tools were disabled, and the clone sat outside the repository, so no parent CLAUDE.md was loaded.
The scored repetition is the second launch. The first was stopped in round 8 by a harness bug (an agent's formal self-check could delete the harness's live formal work directory) and is kept but not scored.
Most broken slots are hypothesis agents that reached the 20-minute limit, the same for every model, while running their own synthesis and place-and-route experiments, without writing a hypothesis.
Caveats on accepted designs. CoreMark cycles are measured with random bus backpressure, but Fmax and area come from a bench wrapper with the memory ready signals tied high, so logic that only hides stalls costs no area or timing. Some accepted designs here contain such logic, as do accepted designs of other configurations; its share of the score is unmeasured. One accepted change made the divider slower, which is cycle-neutral because CoreMark's timed window retires no divides. Another partly optimizes the verification-only RVFI order counter. The LUT4 figures exclude LUT-RAM and block RAM, which later designs use for the register file.
Dollar billing is unavailable under OAuth, so the $0.00 cost is not measured zero spend. Token totals cover every agent (hypothesis, implementer and scribe); output tokens are a lower bound for sessions stopped by the timeout.
Run notes, launch history and caveats .
Best
983.24
+247.7% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
983.24
+247.7%
3,128
302
10
23
9.4h
Broken classes (all reps combined): hypothesis_gen_failed×22, formal_failed×1
rep1 , winning hypotheses
R1 · Split EX into single-cycle integer ALU plus iterative DIV/REM unit fitness 359.01 (+26.9% ) · area 5.1k LUT4 · Fmax 161 MHz
R3 · Stall-only I-fetch replay store with registered next-PC lookahead fitness 438.47 (+22.1% ) · area 5.1k LUT4 · Fmax 162 MHz
R4 · Fetch predecode + 64-entry bimodal BHT stacked on the I-fetch replay-store core fitness 458.93 (+4.7% ) · area 5.6k LUT4 · Fmax 156 MHz
R5 · Pre-decoded one-hot ALU controls registered in ID/EX on the predictor core fitness 512.08 (+11.6% ) · area 5.6k LUT4 · Fmax 174 MHz
R7 · LUT-RAM register file + ID/EX control/data split + registered div b!=0 (Fmax 174 -> ~223 MHz, cycle-identical) fitness 643.41 (+25.6% ) · area 3.1k LUT4 · Fmax 219 MHz
R9 · Hide bus stalls on both sides: stall-only D-side store buffer + load cache, and a 4096-entry I-side replay store (netlist-identical to r7s0) fitness 695.16 (+8.0% ) · area 3.1k LUT4 · Fmax 222 MHz
R10 · Take the divider off every critical path: plain operand latch at start + two-phase registered restoring step (div_unit.sv only, cycle-neutral) fitness 859.77 (+23.7% ) · area 3.1k LUT4 · Fmax 275 MHz
R12 · Front-end storage restructure: BHT in distributed LUT-RAM + split 2x32 rvfi_order counter (draft) fitness 894.39 (+4.0% ) · area 2.8k LUT4 · Fmax 286 MHz
R14 · Late branch unit (v2 form): verify load-dependent BRANCHes in MEM, redirect one cycle later from flops fitness 983.24 (+9.9% ) · area 3.1k LUT4 · Fmax 302 MHz
gpt-5_5_xhigh
gpt 5 5 xhigh
Best
525.04
+85.6% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
397.83
+40.7%
6,052
182
5
2
7.0h
rep2
done
525.04
+85.6%
5,453
220
5
3
6.4h
rep3
done
482.03
+70.4%
3,164
216
7
1
4.9h
Broken classes (all reps combined): cosim_failed×5, formal_failed×1
rep1 , winning hypotheses
R1 · Decouple slow RV32M ops from EX fitness 368.83 (+30.4% ) · area 6.0k LUT4 · Fmax 166 MHz
R3 · Retiming slow M finalization fitness 380.70 (+3.2% ) · area 6.0k LUT4 · Fmax 171 MHz
R4 · Register forwarding selects fitness 386.71 (+1.6% ) · area 6.0k LUT4 · Fmax 174 MHz
R8 · Registered low-half MUL path fitness 397.83 (+2.9% ) · area 6.1k LUT4 · Fmax 182 MHz
rep2 , winning hypotheses
R1 · Iterative divider off ALU critical path fitness 400.55 (+41.6% ) · area 5.5k LUT4 · Fmax 180 MHz
R6 · Prune dead pipeline control bits fitness 427.56 (+6.7% ) · area 5.4k LUT4 · Fmax 192 MHz
R7 · Valid-only pipeline payload resets fitness 432.21 (+1.1% ) · area 5.4k LUT4 · Fmax 194 MHz
R10 · Tiny ifetch replay predictor fitness 525.04 (+21.5% ) · area 5.5k LUT4 · Fmax 220 MHz
rep3 , winning hypotheses
R1 · Move DIV/REM to multicycle EX unit fitness 350.21 (+23.8% ) · area 5.6k LUT4 · Fmax 157 MHz
R2 · Share MUL hardware in ALU fitness 353.49 (+0.9% ) · area 5.6k LUT4 · Fmax 159 MHz
R4 · Hazard-only source-use interlock fitness 381.04 (+7.8% ) · area 5.6k LUT4 · Fmax 171 MHz
R6 · Register writeback payload in MEM/WB fitness 406.96 (+6.8% ) · area 5.7k LUT4 · Fmax 183 MHz
R9 · Remove regfile reset fanout fitness 413.23 (+1.5% ) · area 3.2k LUT4 · Fmax 186 MHz
R10 · Stage-local control bundles fitness 482.03 (+16.6% ) · area 3.2k LUT4 · Fmax 216 MHz
gpt-5_6-terra
gpt 5 6 terra
Best
515.70
+82.3% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
430.45
+52.2%
5,763
193
5
18
3.6h
rep2
done
515.70
+82.3%
10,525
209
7
8
3.7h
rep3
done
380.62
+34.6%
7,729
158
7
13
1.1h
Broken classes (all reps combined): formal_failed×27, cosim_failed×9, build_failed×3
rep1 , winning hypotheses
R1 · Share RV32 multiply datapath fitness 306.38 (+8.3% ) · area 9.8k LUT4 · Fmax 138 MHz
R2 · Decoupled iterative divide unit fitness 405.36 (+32.3% ) · area 5.7k LUT4 · Fmax 182 MHz
R6 · Shared branch subtract flags fitness 411.50 (+1.5% ) · area 5.8k LUT4 · Fmax 185 MHz
R7 · Fixed lane memory extraction fitness 430.45 (+4.6% ) · area 5.8k LUT4 · Fmax 193 MHz
rep2 , winning hypotheses
R1 · Static loop branch prediction fitness 310.00 (+9.6% ) · area 10.5k LUT4 · Fmax 132 MHz
R2 · MEM-to-EX load forwarding fitness 323.71 (+4.4% ) · area 10.0k LUT4 · Fmax 131 MHz
R3 · Shared sign-configurable multiplier fitness 361.53 (+11.7% ) · area 10.6k LUT4 · Fmax 147 MHz
R4 · Opcode-biased static loop prediction fitness 365.65 (+1.1% ) · area 9.8k LUT4 · Fmax 148 MHz
R8 · Class-partitioned ALU control fitness 390.78 (+6.9% ) · area 10.6k LUT4 · Fmax 158 MHz
R9 · Fine-grained base ALU partition fitness 515.70 (+32.0% ) · area 10.5k LUT4 · Fmax 209 MHz
rep3 , winning hypotheses
R1 · Compact direct branch predictor fitness 316.16 (+11.8% ) · area 10.6k LUT4 · Fmax 133 MHz
R3 · Decode-stage branch resolution fitness 344.27 (+8.9% ) · area 10.5k LUT4 · Fmax 143 MHz
R5 · Decode-stage BHT training fitness 345.35 (+0.3% ) · area 10.7k LUT4 · Fmax 143 MHz
R6 · Shared magnitude divide datapath fitness 348.63 (+0.9% ) · area 8.0k LUT4 · Fmax 145 MHz
R8 · JAL-only fetch predictor fitness 353.60 (+1.4% ) · area 8.1k LUT4 · Fmax 147 MHz
R11 · Aligned decode-stage JAL steering fitness 380.62 (+7.6% ) · area 7.7k LUT4 · Fmax 158 MHz
gpt-5_4_xhigh
gpt 5 4 xhigh
Best
513.84
+81.7% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
496.11
+75.4%
3,164
221
5
7
7.2h
rep2
done
513.84
+81.7%
10,108
203
7
11
8.1h
rep3
done
447.40
+58.2%
5,847
199
3
14
14.4h
Broken classes (all reps combined): formal_failed×16, cosim_failed×13, hypothesis_gen_failed×3
rep1 , winning hypotheses
R1 · Move DIV/REM Off The ALU Critical Path fitness 343.51 (+21.5% ) · area 5.6k LUT4 · Fmax 154 MHz
R2 · One-Deep Stalled-Store Retirement Slot fitness 385.04 (+12.1% ) · area 5.6k LUT4 · Fmax 172 MHz
R6 · Non-Aliasing Load Bypass Around Store Slot fitness 439.32 (+14.1% ) · area 5.6k LUT4 · Fmax 196 MHz
R10 · Reset-Light Write-First Register File fitness 496.11 (+12.9% ) · area 3.2k LUT4 · Fmax 221 MHz
rep2 , winning hypotheses
R2 · Registered I-Fetch Replay Predictor fitness 316.05 (+11.8% ) · area 10.1k LUT4 · Fmax 134 MHz
R3 · MEM-to-EX load bypass fitness 332.31 (+5.1% ) · area 9.9k LUT4 · Fmax 134 MHz
R4 · Shared signedness-selectable multiplier fitness 334.91 (+0.8% ) · area 10.4k LUT4 · Fmax 135 MHz
R5 · One-entry posted store buffer fitness 377.60 (+12.8% ) · area 10.1k LUT4 · Fmax 151 MHz
R6 · EX fast-path add/address bypass fitness 391.32 (+3.6% ) · area 10.0k LUT4 · Fmax 157 MHz
R8 · Resolve direct JAL in ID fitness 513.84 (+31.3% ) · area 10.1k LUT4 · Fmax 203 MHz
rep3 , winning hypotheses
R1 · Iterative Cold Divider Off EX Path fitness 405.07 (+43.2% ) · area 5.8k LUT4 · Fmax 182 MHz
R2 · One-Deep Store Retirement Slot fitness 447.40 (+10.4% ) · area 5.8k LUT4 · Fmax 199 MHz
gpt-5_6-luna
gpt 5 6 luna
Best
480.90
+70.0% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
462.59
+63.6%
10,155
201
3
21
4.9h
rep2
done
480.90
+70.0%
10,135
209
5
8
5.0h
rep3
done
412.38
+45.8%
7,500
185
3
13
4.9h
Broken classes (all reps combined): formal_failed×26, cosim_failed×14, hypothesis_gen_failed×1, build_failed×1
rep1 , winning hypotheses
R1 · Share the multiplier across RV32M high-half operations fitness 404.74 (+43.1% ) · area 10.1k LUT4 · Fmax 182 MHz
R2 · Add a decoupled instruction prefetch queue fitness 462.59 (+14.3% ) · area 10.2k LUT4 · Fmax 201 MHz
rep2 , winning hypotheses
R1 · Early static backward-branch prediction fitness 310.56 (+9.8% ) · area 10.6k LUT4 · Fmax 132 MHz
R3 · Share a native-width RV32M multiplier fitness 317.32 (+2.2% ) · area 10.3k LUT4 · Fmax 135 MHz
R5 · Retimed two-phase execute pipeline fitness 396.37 (+24.9% ) · area 10.4k LUT4 · Fmax 172 MHz
R7 · Remove redundant MEM/WB forwarding leg fitness 480.90 (+21.3% ) · area 10.1k LUT4 · Fmax 209 MHz
rep3 , winning hypotheses
R1 · Fast-path ALU with selective M arithmetic fitness 406.88 (+43.9% ) · area 7.5k LUT4 · Fmax 183 MHz
R10 · Shared branch compare flag cone fitness 412.38 (+1.4% ) · area 7.5k LUT4 · Fmax 185 MHz
gpt-6-astra_max
gpt 6 astra max
GPT-6 Astra via Codex · reasoning effort max · 15 rounds × 3 hypotheses per round. 3 independent repetitions recorded.
Across completed repetitions: mean 424.82 ± 44.33 sample SD.
Repetitions 2 and 3 resumed after a usage-limit pause; reported wall time includes that pause. Valid attempts were preserved. The acceptance counts include the baseline; rep3 also has one FPGA placement failure outside the legacy broken counter. OAuth dollar billing was unavailable; raw cost zeros are parser defaults. Run and recovery notes .
Best
474.27
+67.7% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
388.66
+37.4%
6,618
164
6
2
11.9h
rep2
done
411.53
+45.5%
5,485
185
6
1
12.2h
rep3
done
474.27
+67.7%
6,247
199
5
0
12.1h
Broken classes (all reps combined): formal_failed×3
rep1 , winning hypotheses
R1 · Small tagged target predictor with two-bit branch hysteresis fitness 319.45 (+12.9% ) · area 10.6k LUT4 · Fmax 135 MHz
R2 · Split execution into a fast ALU and an eight-cycle divider fitness 320.28 (+0.3% ) · area 5.9k LUT4 · Fmax 135 MHz
R3 · Row-local predictor metadata training to remove the indexed feedback path fitness 350.46 (+9.4% ) · area 6.6k LUT4 · Fmax 148 MHz
R4 · Pipeline row-addressed resolution events before local predictor updates fitness 371.55 (+6.0% ) · area 6.6k LUT4 · Fmax 157 MHz
R5 · Remove wide ID/EX payload registers from bubble-control fanout fitness 388.66 (+4.6% ) · area 6.6k LUT4 · Fmax 164 MHz
rep2 , winning hypotheses
R1 · Separate bounded multicycle division from the single-cycle execute path fitness 372.62 (+31.8% ) · area 5.4k LUT4 · Fmax 167 MHz
R2 · Retiming divide writeback into a separate EX/MEM result bank fitness 375.78 (+0.8% ) · area 5.5k LUT4 · Fmax 169 MHz
R4 · Retime divider finalization using a narrow request-edge prefix fitness 382.48 (+1.8% ) · area 5.5k LUT4 · Fmax 172 MHz
R5 · Separate pipeline payload storage from bubble and divide control fitness 409.80 (+7.1% ) · area 5.5k LUT4 · Fmax 184 MHz
R11 · Remove launch qualification from idle divider operand storage fitness 411.53 (+0.4% ) · area 5.5k LUT4 · Fmax 185 MHz
rep3 , winning hypotheses
R1 · Move DIV and REM into a bounded multicycle execution unit fitness 340.22 (+20.3% ) · area 5.9k LUT4 · Fmax 153 MHz
R2 · Table-free static prediction for backward branches and JAL fitness 425.76 (+25.1% ) · area 6.0k LUT4 · Fmax 181 MHz
R4 · Small agree predictor to learn exceptions to static branch direction fitness 467.44 (+9.8% ) · area 6.1k LUT4 · Fmax 196 MHz
R14 · Sixteen-entry tagged exception predictor for static branch bias fitness 474.27 (+1.5% ) · area 6.2k LUT4 · Fmax 199 MHz
gpt-5_6-sol
gpt 5 6 sol
Best
470.80
+66.5% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
470.80
+66.5%
10,168
200
7
8
1.5h
rep2
failed ⚠
441.11
+56.0%
5,376
198
2
8
4.5h
rep3
failed ⚠
412.87
+46.0%
5,685
177
3
3
4.5h
Broken classes (all reps combined): cosim_failed×10, formal_failed×7, implementation_compile_failed×2
rep1 , winning hypotheses
R1 · Share one multiplier across all RV32M multiply variants fitness 404.74 (+43.1% ) · area 10.1k LUT4 · Fmax 182 MHz
R3 · One-entry taken-target instruction replay buffer fitness 406.58 (+0.5% ) · area 10.5k LUT4 · Fmax 173 MHz
R4 · Explicit memory byte-lane selection fitness 431.88 (+6.2% ) · area 10.4k LUT4 · Fmax 183 MHz
R8 · Target-only replay-buffer tag fitness 463.86 (+7.4% ) · area 10.2k LUT4 · Fmax 197 MHz
R13 · Protocol-qualified replay fill capture fitness 467.91 (+0.9% ) · area 10.3k LUT4 · Fmax 199 MHz
R15 · Protocol-qualified in-place replay refill fitness 470.80 (+0.6% ) · area 10.2k LUT4 · Fmax 200 MHz
rep2 , winning hypotheses
R1 · Replace combinational division with an iterative M-unit fitness 441.11 (+56.0% ) · area 5.4k LUT4 · Fmax 198 MHz
rep3 , winning hypotheses
R1 · Move DIV and REM into a cold iterative execution unit fitness 379.78 (+34.3% ) · area 5.6k LUT4 · Fmax 171 MHz
R2 · Bypass completed loads directly from MEM to EX fitness 412.87 (+8.7% ) · area 5.7k LUT4 · Fmax 177 MHz
gpt-5_5_high
gpt 5 5 high
Best
461.87
+63.3% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
461.87
+63.3%
9,807
187
7
3
4.7h
rep2
done
420.61
+48.7%
11,953
178
8
4
11.9h
rep3
done
408.01
+44.3%
5,637
176
4
2
5.6h
Broken classes (all reps combined): cosim_failed×5, formal_failed×3, implementation_compile_failed×1
rep1 , winning hypotheses
R1 · Static backward branch predictor fitness 338.66 (+19.7% ) · area 10.0k LUT4 · Fmax 144 MHz
R3 · Add MEM-to-EX load forwarding fitness 355.27 (+4.9% ) · area 10.2k LUT4 · Fmax 144 MHz
R8 · Isolate M-extension ALU mux fitness 366.02 (+3.0% ) · area 10.1k LUT4 · Fmax 148 MHz
R9 · Gate static branch target formation fitness 420.37 (+14.8% ) · area 9.9k LUT4 · Fmax 170 MHz
R12 · Factor forwarding matches fitness 437.68 (+4.1% ) · area 10.1k LUT4 · Fmax 177 MHz
R14 · Register memory request metadata fitness 461.87 (+5.5% ) · area 9.8k LUT4 · Fmax 187 MHz
rep2 , winning hypotheses
R1 · Add small BTB branch predictor fitness 285.54 (+1.0% ) · area 12.1k LUT4 · Fmax 121 MHz
R2 · Gate false load-use stalls fitness 290.93 (+1.9% ) · area 12.4k LUT4 · Fmax 123 MHz
R3 · Optimize load byte-lane formatter fitness 305.22 (+4.9% ) · area 12.4k LUT4 · Fmax 129 MHz
R4 · Consolidate multiply datapath fitness 313.60 (+2.8% ) · area 12.6k LUT4 · Fmax 133 MHz
R6 · Trim BTB tag compare fitness 326.00 (+4.0% ) · area 12.0k LUT4 · Fmax 138 MHz
R7 · Prune dead pipeline payload fitness 341.75 (+4.8% ) · area 12.2k LUT4 · Fmax 145 MHz
R11 · Bypass ALU for LSU addresses fitness 420.61 (+23.1% ) · area 12.0k LUT4 · Fmax 178 MHz
rep3 , winning hypotheses
R1 · Move DIV/REM off the ALU critical path fitness 352.98 (+24.8% ) · area 5.5k LUT4 · Fmax 159 MHz
R4 · Case-based MEM byte-lane muxes fitness 380.43 (+7.8% ) · area 5.5k LUT4 · Fmax 171 MHz
R6 · Lookahead hot-branch predictor fitness 408.01 (+7.2% ) · area 5.6k LUT4 · Fmax 176 MHz
gpt-6-sol_xhigh
gpt 6 sol xhigh
Best
435.24
+53.9% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
435.24
+53.9%
5,713
189
6
0
6.5h
Broken classes (all reps combined): n/a
rep1 , winning hypotheses
R1 · Decouple DIV and REM with an iterative EX unit fitness 349.32 (+23.5% ) · area 5.6k LUT4 · Fmax 157 MHz
R2 · Segment the RVFI retirement order counter fitness 361.95 (+3.6% ) · area 5.6k LUT4 · Fmax 163 MHz
R4 · Register one-hot ALU selects before execute fitness 392.16 (+8.3% ) · area 5.6k LUT4 · Fmax 176 MHz
R6 · Clear only side-effect controls on ID/EX bubbles fitness 408.57 (+4.2% ) · area 5.7k LUT4 · Fmax 184 MHz
R14 · Predict backward branches with low-byte PC steering fitness 435.24 (+6.5% ) · area 5.7k LUT4 · Fmax 189 MHz
gpt-5_5_medium
gpt 5 5 medium
Best
431.58
+52.6% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
431.58
+52.6%
7,803
201
7
5
5.8h
rep2
done
407.55
+44.1%
7,358
187
5
9
6.6h
rep3
done
431.24
+52.5%
9,997
194
4
9
7.2h
Broken classes (all reps combined): cosim_failed×12, formal_failed×11
rep1 , winning hypotheses
R1 · Remove regfile reset fanout fitness 316.22 (+11.8% ) · area 8.0k LUT4 · Fmax 142 MHz
R2 · Move M extension to multicycle unit fitness 356.85 (+12.8% ) · area 7.5k LUT4 · Fmax 167 MHz
R3 · Prune dead pipeline metadata fitness 397.64 (+11.4% ) · area 7.7k LUT4 · Fmax 186 MHz
R5 · Add posted store buffer fitness 405.49 (+2.0% ) · area 7.4k LUT4 · Fmax 188 MHz
R9 · Register final writeback data fitness 422.61 (+4.2% ) · area 7.0k LUT4 · Fmax 196 MHz
R12 · Add narrow forwarding sideband fitness 431.58 (+2.1% ) · area 7.8k LUT4 · Fmax 201 MHz
rep2 , winning hypotheses
R1 · Multicycle RV32M arithmetic unit fitness 375.73 (+32.9% ) · area 9.7k LUT4 · Fmax 172 MHz
R5 · Drop regfile reset fanout fitness 382.99 (+1.9% ) · area 7.6k LUT4 · Fmax 176 MHz
R8 · Share M-unit multiplier hardware fitness 393.66 (+2.8% ) · area 7.5k LUT4 · Fmax 181 MHz
R9 · Split RVFI shadow metadata from datapath fitness 407.55 (+3.5% ) · area 7.4k LUT4 · Fmax 187 MHz
rep3 , winning hypotheses
R1 · Gate and share M-extension ALU hardware fitness 398.35 (+40.9% ) · area 9.8k LUT4 · Fmax 179 MHz
R5 · Precompute PC targets in decode fitness 412.07 (+3.4% ) · area 10.2k LUT4 · Fmax 185 MHz
R15 · Retire PC-next in MEM fitness 431.24 (+4.7% ) · area 10.0k LUT4 · Fmax 194 MHz
kimi-k2_6
kimi k2 6
Best
396.13
+40.1% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
347.76
+23.0%
10,254
146
3
23
9.0h
rep2
done
331.22
+17.1%
10,038
141
3
26
8.6h
rep3
failed ⚠
396.13
+40.1%
9,927
166
4
14
8.9h
Broken classes (all reps combined): hypothesis_gen_failed×45, formal_failed×8, cosim_failed×6, implementation_compile_failed×2, schema_error×2
rep1 , winning hypotheses
R1 · Add static branch predictor (backward-taken, JAL-always-taken) fitness 324.08 (+14.6% ) · area 10.2k LUT4 · Fmax 138 MHz
R4 · Add 32-entry 2-bit BHT for forward branch direction prediction fitness 347.76 (+7.3% ) · area 10.3k LUT4 · Fmax 146 MHz
rep2 , winning hypotheses
R1 · IF-stage static predictor: backward branches and JAL always taken fitness 316.95 (+12.1% ) · area 10.6k LUT4 · Fmax 135 MHz
R3 · 4-entry Return Address Stack for JALR returns fitness 331.22 (+4.5% ) · area 10.0k LUT4 · Fmax 141 MHz
rep3 , winning hypotheses
R1 · Guard ALU multipliers off critical path for non-M ops fitness 315.22 (+11.5% ) · area 10.2k LUT4 · Fmax 142 MHz
R5 · 8-entry direct-mapped instruction cache in IF to absorb imem stalls fitness 334.35 (+6.1% ) · area 10.0k LUT4 · Fmax 140 MHz
R8 · Split ALU into fast and M-extension paths with final 2:1 mux fitness 396.13 (+18.5% ) · area 9.9k LUT4 · Fmax 166 MHz
gpt-5_4-mini
gpt 5 4 mini
Best
395.53
+39.9% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
341.86
+20.9%
9,859
152
4
20
15.6h
rep2
done
349.46
+23.6%
9,916
157
7
26
29.8h
rep3
done
395.53
+39.9%
10,230
187
5
22
30.2h
Broken classes (all reps combined): formal_failed×42, cosim_failed×17, hypothesis_gen_failed×5, implementation_compile_failed×1, sandbox_violation×1, schema_error×1, build_failed×1
rep1 , winning hypotheses
R1 · Retime branch target generation fitness 335.70 (+18.7% ) · area 9.4k LUT4 · Fmax 151 MHz
R6 · Retimed memory lane metadata fitness 339.60 (+1.2% ) · area 10.1k LUT4 · Fmax 153 MHz
R11 · JAL fast-path predictor fitness 341.86 (+0.7% ) · area 9.9k LUT4 · Fmax 152 MHz
rep2 , winning hypotheses
R1 · Precompute branch targets in ID fitness 296.38 (+4.8% ) · area 9.8k LUT4 · Fmax 133 MHz
R2 · Latch pc+4 in ID fitness 311.83 (+5.2% ) · area 9.9k LUT4 · Fmax 140 MHz
R4 · Latch writeback data in MEM fitness 312.26 (+0.1% ) · area 9.9k LUT4 · Fmax 140 MHz
R5 · Predecode instruction slices fitness 322.94 (+3.4% ) · area 9.8k LUT4 · Fmax 145 MHz
R8 · Move memory lane work earlier fitness 342.69 (+6.1% ) · area 10.4k LUT4 · Fmax 154 MHz
R12 · Hoist full decode into fetch fitness 349.46 (+2.0% ) · area 9.9k LUT4 · Fmax 157 MHz
rep3 , winning hypotheses
R1 · Registered Fetch/Decode Cutpoint fitness 310.94 (+9.9% ) · area 10.0k LUT4 · Fmax 154 MHz
R9 · Split slow M-extension path fitness 357.14 (+14.9% ) · area 10.7k LUT4 · Fmax 183 MHz
R11 · Backward-branch-only predictor fitness 394.64 (+10.5% ) · area 10.2k LUT4 · Fmax 186 MHz
R15 · Hot-loop shadow latch fitness 395.53 (+0.2% ) · area 10.2k LUT4 · Fmax 187 MHz
gemini-3_5-flash
gemini 3 5 flash
Best
359.04
+27.0% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep2
failed ⚠
359.04
+27.0%
13,773
125
2
1
0.9h
Broken classes (all reps combined): hypothesis_gen_failed×1
rep2 , winning hypotheses
R1 · Two-bank registered lookahead instruction replay predictor with static branch/JAL predecode fitness 359.04 (+26.9% ) · area 13.8k LUT4 · Fmax 125 MHz
gemini-3_1-pro
gemini 3 1 pro
Best
354.73
+25.4% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
354.73
+25.4%
10,242
150
2
31
5.6h
rep2
done
339.62
+20.1%
11,068
142
3
29
7.3h
rep3
done
323.92
+14.5%
11,836
136
4
31
6.0h
Broken classes (all reps combined): hypothesis_gen_failed×68, sandbox_violation×9, formal_failed×6, cosim_failed×5, implementation_compile_failed×3
rep1 , winning hypotheses
R5 · 1-Cycle 64-entry BTB in IF Stage with Fast Redirect fitness 354.73 (+25.4% ) · area 10.2k LUT4 · Fmax 150 MHz
rep2 , winning hypotheses
R4 · 16-entry BTB/BHT predictor in IF stage fitness 323.95 (+14.5% ) · area 10.3k LUT4 · Fmax 137 MHz
R5 · 128-entry BTB + 8-entry RAS fitness 339.62 (+4.8% ) · area 11.1k LUT4 · Fmax 142 MHz
rep3 , winning hypotheses
R5 · IF-stage Static BTFN and JAL Predictor fitness 282.91 (+0.0% ) · area 10.0k LUT4 · Fmax 121 MHz
R9 · BHT and RAS for frontend branch prediction fitness 321.67 (+13.7% ) · area 11.0k LUT4 · Fmax 135 MHz
R10 · GShare Predictor with 256-entry BHT and 8-bit GHR fitness 323.92 (+0.7% ) · area 11.8k LUT4 · Fmax 136 MHz
static
static
Best
282.82
+0.0% vs baseline
per-rep detail
Rep Status
Best Δ%
Area (LUT4) Fmax (MHz)
acc brk
Wall
rep1
done
282.82
+0.0%
9,563
127
1
42
0.2h
rep2
done
282.82
+0.0%
9,563
127
1
0
2.0h
Broken classes (all reps combined): hypothesis_gen_failed×42
No winning hypotheses recorded.