Per-model detail

What each model actually did.

Below: the per-rep outcomes for every model run on HWE Bench so far, plus the accepted-improvement hypotheses each rep produced, verbatim titles, fitness, LUT4, and Fmax. The hypothesis titles are exactly what the agent wrote.

claude-opus-5_5_xhigh

claude opus 5 5 xhigh

Claude Opus 5.5 via Claude Code 2.1.282 · reasoning effort xhigh · subscription (OAuth) login · 15 rounds × 3 hypotheses per round. One repetition (n=1): repeatability has not been measured, and any ranking against other configurations is untested.

Isolation: no user settings, plugins, hooks, MCP servers, connectors or auto-memory. Bash ran in a sandbox with no network, writes confined to the rep's clone and reads of the operator's home directory denied. Web tools were disabled, and the clone sat outside the repository, so no parent CLAUDE.md was loaded.

The scored repetition is the second launch. The first was stopped in round 8 by a harness bug (an agent's formal self-check could delete the harness's live formal work directory) and is kept but not scored.

Most broken slots are hypothesis agents that reached the 20-minute limit, the same for every model, while running their own synthesis and place-and-route experiments, without writing a hypothesis.

Caveats on accepted designs. CoreMark cycles are measured with random bus backpressure, but Fmax and area come from a bench wrapper with the memory ready signals tied high, so logic that only hides stalls costs no area or timing. Some accepted designs here contain such logic, as do accepted designs of other configurations; its share of the score is unmeasured. One accepted change made the divider slower, which is cycle-neutral because CoreMark's timed window retires no divides. Another partly optimizes the verification-only RVFI order counter. The LUT4 figures exclude LUT-RAM and block RAM, which later designs use for the register file.

Dollar billing is unavailable under OAuth, so the $0.00 cost is not measured zero spend. Token totals cover every agent (hypothesis, implementer and scribe); output tokens are a lower bound for sessions stopped by the timeout.

Run notes, launch history and caveats.

Best
983.24
+247.7% vs baseline
Mean
983.2
+247.7% mean Δ
Reps
1/1
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 983.24 +247.7% 3,128 302 10 23 9.4h

Broken classes (all reps combined): hypothesis_gen_failed×22, formal_failed×1

rep1 , winning hypotheses

R1 · Split EX into single-cycle integer ALU plus iterative DIV/REM unit
fitness 359.01 (+26.9%) · area 5.1k LUT4 · Fmax 161 MHz
R3 · Stall-only I-fetch replay store with registered next-PC lookahead
fitness 438.47 (+22.1%) · area 5.1k LUT4 · Fmax 162 MHz
R4 · Fetch predecode + 64-entry bimodal BHT stacked on the I-fetch replay-store core
fitness 458.93 (+4.7%) · area 5.6k LUT4 · Fmax 156 MHz
R5 · Pre-decoded one-hot ALU controls registered in ID/EX on the predictor core
fitness 512.08 (+11.6%) · area 5.6k LUT4 · Fmax 174 MHz
R7 · LUT-RAM register file + ID/EX control/data split + registered div b!=0 (Fmax 174 -> ~223 MHz, cycle-identical)
fitness 643.41 (+25.6%) · area 3.1k LUT4 · Fmax 219 MHz
R9 · Hide bus stalls on both sides: stall-only D-side store buffer + load cache, and a 4096-entry I-side replay store (netlist-identical to r7s0)
fitness 695.16 (+8.0%) · area 3.1k LUT4 · Fmax 222 MHz
R10 · Take the divider off every critical path: plain operand latch at start + two-phase registered restoring step (div_unit.sv only, cycle-neutral)
fitness 859.77 (+23.7%) · area 3.1k LUT4 · Fmax 275 MHz
R12 · Front-end storage restructure: BHT in distributed LUT-RAM + split 2x32 rvfi_order counter (draft)
fitness 894.39 (+4.0%) · area 2.8k LUT4 · Fmax 286 MHz
R14 · Late branch unit (v2 form): verify load-dependent BRANCHes in MEM, redirect one cycle later from flops
fitness 983.24 (+9.9%) · area 3.1k LUT4 · Fmax 302 MHz
gpt-5_5_xhigh

gpt 5 5 xhigh

Best
525.04
+85.6% vs baseline
Mean
468.3
+65.6% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 397.83 +40.7% 6,052 182 5 2 7.0h
rep2 done 525.04 +85.6% 5,453 220 5 3 6.4h
rep3 done 482.03 +70.4% 3,164 216 7 1 4.9h

Broken classes (all reps combined): cosim_failed×5, formal_failed×1

rep1 , winning hypotheses

R1 · Decouple slow RV32M ops from EX
fitness 368.83 (+30.4%) · area 6.0k LUT4 · Fmax 166 MHz
R3 · Retiming slow M finalization
fitness 380.70 (+3.2%) · area 6.0k LUT4 · Fmax 171 MHz
R4 · Register forwarding selects
fitness 386.71 (+1.6%) · area 6.0k LUT4 · Fmax 174 MHz
R8 · Registered low-half MUL path
fitness 397.83 (+2.9%) · area 6.1k LUT4 · Fmax 182 MHz

rep2 , winning hypotheses

R1 · Iterative divider off ALU critical path
fitness 400.55 (+41.6%) · area 5.5k LUT4 · Fmax 180 MHz
R6 · Prune dead pipeline control bits
fitness 427.56 (+6.7%) · area 5.4k LUT4 · Fmax 192 MHz
R7 · Valid-only pipeline payload resets
fitness 432.21 (+1.1%) · area 5.4k LUT4 · Fmax 194 MHz
R10 · Tiny ifetch replay predictor
fitness 525.04 (+21.5%) · area 5.5k LUT4 · Fmax 220 MHz

rep3 , winning hypotheses

R1 · Move DIV/REM to multicycle EX unit
fitness 350.21 (+23.8%) · area 5.6k LUT4 · Fmax 157 MHz
R2 · Share MUL hardware in ALU
fitness 353.49 (+0.9%) · area 5.6k LUT4 · Fmax 159 MHz
R4 · Hazard-only source-use interlock
fitness 381.04 (+7.8%) · area 5.6k LUT4 · Fmax 171 MHz
R6 · Register writeback payload in MEM/WB
fitness 406.96 (+6.8%) · area 5.7k LUT4 · Fmax 183 MHz
R9 · Remove regfile reset fanout
fitness 413.23 (+1.5%) · area 3.2k LUT4 · Fmax 186 MHz
R10 · Stage-local control bundles
fitness 482.03 (+16.6%) · area 3.2k LUT4 · Fmax 216 MHz
gpt-5_6-terra

gpt 5 6 terra

Best
515.70
+82.3% vs baseline
Mean
442.3
+56.4% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 430.45 +52.2% 5,763 193 5 18 3.6h
rep2 done 515.70 +82.3% 10,525 209 7 8 3.7h
rep3 done 380.62 +34.6% 7,729 158 7 13 1.1h

Broken classes (all reps combined): formal_failed×27, cosim_failed×9, build_failed×3

rep1 , winning hypotheses

R1 · Share RV32 multiply datapath
fitness 306.38 (+8.3%) · area 9.8k LUT4 · Fmax 138 MHz
R2 · Decoupled iterative divide unit
fitness 405.36 (+32.3%) · area 5.7k LUT4 · Fmax 182 MHz
R6 · Shared branch subtract flags
fitness 411.50 (+1.5%) · area 5.8k LUT4 · Fmax 185 MHz
R7 · Fixed lane memory extraction
fitness 430.45 (+4.6%) · area 5.8k LUT4 · Fmax 193 MHz

rep2 , winning hypotheses

R1 · Static loop branch prediction
fitness 310.00 (+9.6%) · area 10.5k LUT4 · Fmax 132 MHz
R2 · MEM-to-EX load forwarding
fitness 323.71 (+4.4%) · area 10.0k LUT4 · Fmax 131 MHz
R3 · Shared sign-configurable multiplier
fitness 361.53 (+11.7%) · area 10.6k LUT4 · Fmax 147 MHz
R4 · Opcode-biased static loop prediction
fitness 365.65 (+1.1%) · area 9.8k LUT4 · Fmax 148 MHz
R8 · Class-partitioned ALU control
fitness 390.78 (+6.9%) · area 10.6k LUT4 · Fmax 158 MHz
R9 · Fine-grained base ALU partition
fitness 515.70 (+32.0%) · area 10.5k LUT4 · Fmax 209 MHz

rep3 , winning hypotheses

R1 · Compact direct branch predictor
fitness 316.16 (+11.8%) · area 10.6k LUT4 · Fmax 133 MHz
R3 · Decode-stage branch resolution
fitness 344.27 (+8.9%) · area 10.5k LUT4 · Fmax 143 MHz
R5 · Decode-stage BHT training
fitness 345.35 (+0.3%) · area 10.7k LUT4 · Fmax 143 MHz
R6 · Shared magnitude divide datapath
fitness 348.63 (+0.9%) · area 8.0k LUT4 · Fmax 145 MHz
R8 · JAL-only fetch predictor
fitness 353.60 (+1.4%) · area 8.1k LUT4 · Fmax 147 MHz
R11 · Aligned decode-stage JAL steering
fitness 380.62 (+7.6%) · area 7.7k LUT4 · Fmax 158 MHz
gpt-5_4_xhigh

gpt 5 4 xhigh

Best
513.84
+81.7% vs baseline
Mean
485.8
+71.8% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 496.11 +75.4% 3,164 221 5 7 7.2h
rep2 done 513.84 +81.7% 10,108 203 7 11 8.1h
rep3 done 447.40 +58.2% 5,847 199 3 14 14.4h

Broken classes (all reps combined): formal_failed×16, cosim_failed×13, hypothesis_gen_failed×3

rep1 , winning hypotheses

R1 · Move DIV/REM Off The ALU Critical Path
fitness 343.51 (+21.5%) · area 5.6k LUT4 · Fmax 154 MHz
R2 · One-Deep Stalled-Store Retirement Slot
fitness 385.04 (+12.1%) · area 5.6k LUT4 · Fmax 172 MHz
R6 · Non-Aliasing Load Bypass Around Store Slot
fitness 439.32 (+14.1%) · area 5.6k LUT4 · Fmax 196 MHz
R10 · Reset-Light Write-First Register File
fitness 496.11 (+12.9%) · area 3.2k LUT4 · Fmax 221 MHz

rep2 , winning hypotheses

R2 · Registered I-Fetch Replay Predictor
fitness 316.05 (+11.8%) · area 10.1k LUT4 · Fmax 134 MHz
R3 · MEM-to-EX load bypass
fitness 332.31 (+5.1%) · area 9.9k LUT4 · Fmax 134 MHz
R4 · Shared signedness-selectable multiplier
fitness 334.91 (+0.8%) · area 10.4k LUT4 · Fmax 135 MHz
R5 · One-entry posted store buffer
fitness 377.60 (+12.8%) · area 10.1k LUT4 · Fmax 151 MHz
R6 · EX fast-path add/address bypass
fitness 391.32 (+3.6%) · area 10.0k LUT4 · Fmax 157 MHz
R8 · Resolve direct JAL in ID
fitness 513.84 (+31.3%) · area 10.1k LUT4 · Fmax 203 MHz

rep3 , winning hypotheses

R1 · Iterative Cold Divider Off EX Path
fitness 405.07 (+43.2%) · area 5.8k LUT4 · Fmax 182 MHz
R2 · One-Deep Store Retirement Slot
fitness 447.40 (+10.4%) · area 5.8k LUT4 · Fmax 199 MHz
gpt-5_6-luna

gpt 5 6 luna

Best
480.90
+70.0% vs baseline
Mean
452.0
+59.8% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 462.59 +63.6% 10,155 201 3 21 4.9h
rep2 done 480.90 +70.0% 10,135 209 5 8 5.0h
rep3 done 412.38 +45.8% 7,500 185 3 13 4.9h

Broken classes (all reps combined): formal_failed×26, cosim_failed×14, hypothesis_gen_failed×1, build_failed×1

rep1 , winning hypotheses

R1 · Share the multiplier across RV32M high-half operations
fitness 404.74 (+43.1%) · area 10.1k LUT4 · Fmax 182 MHz
R2 · Add a decoupled instruction prefetch queue
fitness 462.59 (+14.3%) · area 10.2k LUT4 · Fmax 201 MHz

rep2 , winning hypotheses

R1 · Early static backward-branch prediction
fitness 310.56 (+9.8%) · area 10.6k LUT4 · Fmax 132 MHz
R3 · Share a native-width RV32M multiplier
fitness 317.32 (+2.2%) · area 10.3k LUT4 · Fmax 135 MHz
R5 · Retimed two-phase execute pipeline
fitness 396.37 (+24.9%) · area 10.4k LUT4 · Fmax 172 MHz
R7 · Remove redundant MEM/WB forwarding leg
fitness 480.90 (+21.3%) · area 10.1k LUT4 · Fmax 209 MHz

rep3 , winning hypotheses

R1 · Fast-path ALU with selective M arithmetic
fitness 406.88 (+43.9%) · area 7.5k LUT4 · Fmax 183 MHz
R10 · Shared branch compare flag cone
fitness 412.38 (+1.4%) · area 7.5k LUT4 · Fmax 185 MHz
gpt-6-astra_max

gpt 6 astra max

GPT-6 Astra via Codex · reasoning effort max · 15 rounds × 3 hypotheses per round. 3 independent repetitions recorded.

Across completed repetitions: mean 424.82 ± 44.33 sample SD.

Repetitions 2 and 3 resumed after a usage-limit pause; reported wall time includes that pause. Valid attempts were preserved. The acceptance counts include the baseline; rep3 also has one FPGA placement failure outside the legacy broken counter. OAuth dollar billing was unavailable; raw cost zeros are parser defaults. Run and recovery notes.

Best
474.27
+67.7% vs baseline
Mean
424.8
+50.2% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 388.66 +37.4% 6,618 164 6 2 11.9h
rep2 done 411.53 +45.5% 5,485 185 6 1 12.2h
rep3 done 474.27 +67.7% 6,247 199 5 0 12.1h

Broken classes (all reps combined): formal_failed×3

rep1 , winning hypotheses

R1 · Small tagged target predictor with two-bit branch hysteresis
fitness 319.45 (+12.9%) · area 10.6k LUT4 · Fmax 135 MHz
R2 · Split execution into a fast ALU and an eight-cycle divider
fitness 320.28 (+0.3%) · area 5.9k LUT4 · Fmax 135 MHz
R3 · Row-local predictor metadata training to remove the indexed feedback path
fitness 350.46 (+9.4%) · area 6.6k LUT4 · Fmax 148 MHz
R4 · Pipeline row-addressed resolution events before local predictor updates
fitness 371.55 (+6.0%) · area 6.6k LUT4 · Fmax 157 MHz
R5 · Remove wide ID/EX payload registers from bubble-control fanout
fitness 388.66 (+4.6%) · area 6.6k LUT4 · Fmax 164 MHz

rep2 , winning hypotheses

R1 · Separate bounded multicycle division from the single-cycle execute path
fitness 372.62 (+31.8%) · area 5.4k LUT4 · Fmax 167 MHz
R2 · Retiming divide writeback into a separate EX/MEM result bank
fitness 375.78 (+0.8%) · area 5.5k LUT4 · Fmax 169 MHz
R4 · Retime divider finalization using a narrow request-edge prefix
fitness 382.48 (+1.8%) · area 5.5k LUT4 · Fmax 172 MHz
R5 · Separate pipeline payload storage from bubble and divide control
fitness 409.80 (+7.1%) · area 5.5k LUT4 · Fmax 184 MHz
R11 · Remove launch qualification from idle divider operand storage
fitness 411.53 (+0.4%) · area 5.5k LUT4 · Fmax 185 MHz

rep3 , winning hypotheses

R1 · Move DIV and REM into a bounded multicycle execution unit
fitness 340.22 (+20.3%) · area 5.9k LUT4 · Fmax 153 MHz
R2 · Table-free static prediction for backward branches and JAL
fitness 425.76 (+25.1%) · area 6.0k LUT4 · Fmax 181 MHz
R4 · Small agree predictor to learn exceptions to static branch direction
fitness 467.44 (+9.8%) · area 6.1k LUT4 · Fmax 196 MHz
R14 · Sixteen-entry tagged exception predictor for static branch bias
fitness 474.27 (+1.5%) · area 6.2k LUT4 · Fmax 199 MHz
gpt-5_6-sol

gpt 5 6 sol

Best
470.80
+66.5% vs baseline
Mean
470.8
+66.5% mean Δ
Reps
1/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 470.80 +66.5% 10,168 200 7 8 1.5h
rep2 failed ⚠ 441.11 +56.0% 5,376 198 2 8 4.5h
rep3 failed ⚠ 412.87 +46.0% 5,685 177 3 3 4.5h

Broken classes (all reps combined): cosim_failed×10, formal_failed×7, implementation_compile_failed×2

rep1 , winning hypotheses

R1 · Share one multiplier across all RV32M multiply variants
fitness 404.74 (+43.1%) · area 10.1k LUT4 · Fmax 182 MHz
R3 · One-entry taken-target instruction replay buffer
fitness 406.58 (+0.5%) · area 10.5k LUT4 · Fmax 173 MHz
R4 · Explicit memory byte-lane selection
fitness 431.88 (+6.2%) · area 10.4k LUT4 · Fmax 183 MHz
R8 · Target-only replay-buffer tag
fitness 463.86 (+7.4%) · area 10.2k LUT4 · Fmax 197 MHz
R13 · Protocol-qualified replay fill capture
fitness 467.91 (+0.9%) · area 10.3k LUT4 · Fmax 199 MHz
R15 · Protocol-qualified in-place replay refill
fitness 470.80 (+0.6%) · area 10.2k LUT4 · Fmax 200 MHz

rep2 , winning hypotheses

R1 · Replace combinational division with an iterative M-unit
fitness 441.11 (+56.0%) · area 5.4k LUT4 · Fmax 198 MHz

rep3 , winning hypotheses

R1 · Move DIV and REM into a cold iterative execution unit
fitness 379.78 (+34.3%) · area 5.6k LUT4 · Fmax 171 MHz
R2 · Bypass completed loads directly from MEM to EX
fitness 412.87 (+8.7%) · area 5.7k LUT4 · Fmax 177 MHz
gpt-5_5_high

gpt 5 5 high

Best
461.87
+63.3% vs baseline
Mean
430.2
+52.1% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 461.87 +63.3% 9,807 187 7 3 4.7h
rep2 done 420.61 +48.7% 11,953 178 8 4 11.9h
rep3 done 408.01 +44.3% 5,637 176 4 2 5.6h

Broken classes (all reps combined): cosim_failed×5, formal_failed×3, implementation_compile_failed×1

rep1 , winning hypotheses

R1 · Static backward branch predictor
fitness 338.66 (+19.7%) · area 10.0k LUT4 · Fmax 144 MHz
R3 · Add MEM-to-EX load forwarding
fitness 355.27 (+4.9%) · area 10.2k LUT4 · Fmax 144 MHz
R8 · Isolate M-extension ALU mux
fitness 366.02 (+3.0%) · area 10.1k LUT4 · Fmax 148 MHz
R9 · Gate static branch target formation
fitness 420.37 (+14.8%) · area 9.9k LUT4 · Fmax 170 MHz
R12 · Factor forwarding matches
fitness 437.68 (+4.1%) · area 10.1k LUT4 · Fmax 177 MHz
R14 · Register memory request metadata
fitness 461.87 (+5.5%) · area 9.8k LUT4 · Fmax 187 MHz

rep2 , winning hypotheses

R1 · Add small BTB branch predictor
fitness 285.54 (+1.0%) · area 12.1k LUT4 · Fmax 121 MHz
R2 · Gate false load-use stalls
fitness 290.93 (+1.9%) · area 12.4k LUT4 · Fmax 123 MHz
R3 · Optimize load byte-lane formatter
fitness 305.22 (+4.9%) · area 12.4k LUT4 · Fmax 129 MHz
R4 · Consolidate multiply datapath
fitness 313.60 (+2.8%) · area 12.6k LUT4 · Fmax 133 MHz
R6 · Trim BTB tag compare
fitness 326.00 (+4.0%) · area 12.0k LUT4 · Fmax 138 MHz
R7 · Prune dead pipeline payload
fitness 341.75 (+4.8%) · area 12.2k LUT4 · Fmax 145 MHz
R11 · Bypass ALU for LSU addresses
fitness 420.61 (+23.1%) · area 12.0k LUT4 · Fmax 178 MHz

rep3 , winning hypotheses

R1 · Move DIV/REM off the ALU critical path
fitness 352.98 (+24.8%) · area 5.5k LUT4 · Fmax 159 MHz
R4 · Case-based MEM byte-lane muxes
fitness 380.43 (+7.8%) · area 5.5k LUT4 · Fmax 171 MHz
R6 · Lookahead hot-branch predictor
fitness 408.01 (+7.2%) · area 5.6k LUT4 · Fmax 176 MHz
gpt-6-sol_xhigh

gpt 6 sol xhigh

Best
435.24
+53.9% vs baseline
Mean
435.2
+53.9% mean Δ
Reps
1/1
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 435.24 +53.9% 5,713 189 6 0 6.5h

Broken classes (all reps combined): n/a

rep1 , winning hypotheses

R1 · Decouple DIV and REM with an iterative EX unit
fitness 349.32 (+23.5%) · area 5.6k LUT4 · Fmax 157 MHz
R2 · Segment the RVFI retirement order counter
fitness 361.95 (+3.6%) · area 5.6k LUT4 · Fmax 163 MHz
R4 · Register one-hot ALU selects before execute
fitness 392.16 (+8.3%) · area 5.6k LUT4 · Fmax 176 MHz
R6 · Clear only side-effect controls on ID/EX bubbles
fitness 408.57 (+4.2%) · area 5.7k LUT4 · Fmax 184 MHz
R14 · Predict backward branches with low-byte PC steering
fitness 435.24 (+6.5%) · area 5.7k LUT4 · Fmax 189 MHz
gpt-5_5_medium

gpt 5 5 medium

Best
431.58
+52.6% vs baseline
Mean
423.5
+49.7% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 431.58 +52.6% 7,803 201 7 5 5.8h
rep2 done 407.55 +44.1% 7,358 187 5 9 6.6h
rep3 done 431.24 +52.5% 9,997 194 4 9 7.2h

Broken classes (all reps combined): cosim_failed×12, formal_failed×11

rep1 , winning hypotheses

R1 · Remove regfile reset fanout
fitness 316.22 (+11.8%) · area 8.0k LUT4 · Fmax 142 MHz
R2 · Move M extension to multicycle unit
fitness 356.85 (+12.8%) · area 7.5k LUT4 · Fmax 167 MHz
R3 · Prune dead pipeline metadata
fitness 397.64 (+11.4%) · area 7.7k LUT4 · Fmax 186 MHz
R5 · Add posted store buffer
fitness 405.49 (+2.0%) · area 7.4k LUT4 · Fmax 188 MHz
R9 · Register final writeback data
fitness 422.61 (+4.2%) · area 7.0k LUT4 · Fmax 196 MHz
R12 · Add narrow forwarding sideband
fitness 431.58 (+2.1%) · area 7.8k LUT4 · Fmax 201 MHz

rep2 , winning hypotheses

R1 · Multicycle RV32M arithmetic unit
fitness 375.73 (+32.9%) · area 9.7k LUT4 · Fmax 172 MHz
R5 · Drop regfile reset fanout
fitness 382.99 (+1.9%) · area 7.6k LUT4 · Fmax 176 MHz
R8 · Share M-unit multiplier hardware
fitness 393.66 (+2.8%) · area 7.5k LUT4 · Fmax 181 MHz
R9 · Split RVFI shadow metadata from datapath
fitness 407.55 (+3.5%) · area 7.4k LUT4 · Fmax 187 MHz

rep3 , winning hypotheses

R1 · Gate and share M-extension ALU hardware
fitness 398.35 (+40.9%) · area 9.8k LUT4 · Fmax 179 MHz
R5 · Precompute PC targets in decode
fitness 412.07 (+3.4%) · area 10.2k LUT4 · Fmax 185 MHz
R15 · Retire PC-next in MEM
fitness 431.24 (+4.7%) · area 10.0k LUT4 · Fmax 194 MHz
kimi-k2_6

kimi k2 6

Best
396.13
+40.1% vs baseline
Mean
339.5
+20.0% mean Δ
Reps
2/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 347.76 +23.0% 10,254 146 3 23 9.0h
rep2 done 331.22 +17.1% 10,038 141 3 26 8.6h
rep3 failed ⚠ 396.13 +40.1% 9,927 166 4 14 8.9h

Broken classes (all reps combined): hypothesis_gen_failed×45, formal_failed×8, cosim_failed×6, implementation_compile_failed×2, schema_error×2

rep1 , winning hypotheses

R1 · Add static branch predictor (backward-taken, JAL-always-taken)
fitness 324.08 (+14.6%) · area 10.2k LUT4 · Fmax 138 MHz
R4 · Add 32-entry 2-bit BHT for forward branch direction prediction
fitness 347.76 (+7.3%) · area 10.3k LUT4 · Fmax 146 MHz

rep2 , winning hypotheses

R1 · IF-stage static predictor: backward branches and JAL always taken
fitness 316.95 (+12.1%) · area 10.6k LUT4 · Fmax 135 MHz
R3 · 4-entry Return Address Stack for JALR returns
fitness 331.22 (+4.5%) · area 10.0k LUT4 · Fmax 141 MHz

rep3 , winning hypotheses

R1 · Guard ALU multipliers off critical path for non-M ops
fitness 315.22 (+11.5%) · area 10.2k LUT4 · Fmax 142 MHz
R5 · 8-entry direct-mapped instruction cache in IF to absorb imem stalls
fitness 334.35 (+6.1%) · area 10.0k LUT4 · Fmax 140 MHz
R8 · Split ALU into fast and M-extension paths with final 2:1 mux
fitness 396.13 (+18.5%) · area 9.9k LUT4 · Fmax 166 MHz
gpt-5_4-mini

gpt 5 4 mini

Best
395.53
+39.9% vs baseline
Mean
362.3
+28.1% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 341.86 +20.9% 9,859 152 4 20 15.6h
rep2 done 349.46 +23.6% 9,916 157 7 26 29.8h
rep3 done 395.53 +39.9% 10,230 187 5 22 30.2h

Broken classes (all reps combined): formal_failed×42, cosim_failed×17, hypothesis_gen_failed×5, implementation_compile_failed×1, sandbox_violation×1, schema_error×1, build_failed×1

rep1 , winning hypotheses

R1 · Retime branch target generation
fitness 335.70 (+18.7%) · area 9.4k LUT4 · Fmax 151 MHz
R6 · Retimed memory lane metadata
fitness 339.60 (+1.2%) · area 10.1k LUT4 · Fmax 153 MHz
R11 · JAL fast-path predictor
fitness 341.86 (+0.7%) · area 9.9k LUT4 · Fmax 152 MHz

rep2 , winning hypotheses

R1 · Precompute branch targets in ID
fitness 296.38 (+4.8%) · area 9.8k LUT4 · Fmax 133 MHz
R2 · Latch pc+4 in ID
fitness 311.83 (+5.2%) · area 9.9k LUT4 · Fmax 140 MHz
R4 · Latch writeback data in MEM
fitness 312.26 (+0.1%) · area 9.9k LUT4 · Fmax 140 MHz
R5 · Predecode instruction slices
fitness 322.94 (+3.4%) · area 9.8k LUT4 · Fmax 145 MHz
R8 · Move memory lane work earlier
fitness 342.69 (+6.1%) · area 10.4k LUT4 · Fmax 154 MHz
R12 · Hoist full decode into fetch
fitness 349.46 (+2.0%) · area 9.9k LUT4 · Fmax 157 MHz

rep3 , winning hypotheses

R1 · Registered Fetch/Decode Cutpoint
fitness 310.94 (+9.9%) · area 10.0k LUT4 · Fmax 154 MHz
R9 · Split slow M-extension path
fitness 357.14 (+14.9%) · area 10.7k LUT4 · Fmax 183 MHz
R11 · Backward-branch-only predictor
fitness 394.64 (+10.5%) · area 10.2k LUT4 · Fmax 186 MHz
R15 · Hot-loop shadow latch
fitness 395.53 (+0.2%) · area 10.2k LUT4 · Fmax 187 MHz
gemini-3_5-flash

gemini 3 5 flash

Best
359.04
+27.0% vs baseline
Mean
n/a
n/a mean Δ
Reps
0/1
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep2 failed ⚠ 359.04 +27.0% 13,773 125 2 1 0.9h

Broken classes (all reps combined): hypothesis_gen_failed×1

rep2 , winning hypotheses

R1 · Two-bank registered lookahead instruction replay predictor with static branch/JAL predecode
fitness 359.04 (+26.9%) · area 13.8k LUT4 · Fmax 125 MHz
gemini-3_1-pro

gemini 3 1 pro

Best
354.73
+25.4% vs baseline
Mean
339.4
+20.0% mean Δ
Reps
3/3
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 354.73 +25.4% 10,242 150 2 31 5.6h
rep2 done 339.62 +20.1% 11,068 142 3 29 7.3h
rep3 done 323.92 +14.5% 11,836 136 4 31 6.0h

Broken classes (all reps combined): hypothesis_gen_failed×68, sandbox_violation×9, formal_failed×6, cosim_failed×5, implementation_compile_failed×3

rep1 , winning hypotheses

R5 · 1-Cycle 64-entry BTB in IF Stage with Fast Redirect
fitness 354.73 (+25.4%) · area 10.2k LUT4 · Fmax 150 MHz

rep2 , winning hypotheses

R4 · 16-entry BTB/BHT predictor in IF stage
fitness 323.95 (+14.5%) · area 10.3k LUT4 · Fmax 137 MHz
R5 · 128-entry BTB + 8-entry RAS
fitness 339.62 (+4.8%) · area 11.1k LUT4 · Fmax 142 MHz

rep3 , winning hypotheses

R5 · IF-stage Static BTFN and JAL Predictor
fitness 282.91 (+0.0%) · area 10.0k LUT4 · Fmax 121 MHz
R9 · BHT and RAS for frontend branch prediction
fitness 321.67 (+13.7%) · area 11.0k LUT4 · Fmax 135 MHz
R10 · GShare Predictor with 256-entry BHT and 8-bit GHR
fitness 323.92 (+0.7%) · area 11.8k LUT4 · Fmax 136 MHz
static

static

Best
282.82
+0.0% vs baseline
Mean
282.8
+0.0% mean Δ
Reps
2/2
completed / total
per-rep detail
RepStatus BestΔ% Area (LUT4)Fmax (MHz) accbrk Wall
rep1 done 282.82 +0.0% 9,563 127 1 42 0.2h
rep2 done 282.82 +0.0% 9,563 127 1 0 2.0h

Broken classes (all reps combined): hypothesis_gen_failed×42

No winning hypotheses recorded.