Evalstate Qwen GEPA Report

Run: gepa-evalstate-qwen-overlay-c4-full-20260616T172947Z. Updated 2026-06-17. Dataset: evalstate/openclaw-git-labels. Runtime: localpager-agent over vLLM Qwen.

GEPA Search Result

Best Pareto score
0.6979
delta +0.1237 vs seed
Seed Pareto score
0.5742
v10 overlay seed
Selected iterations
32
31 actual proposal texts
Full-val candidates
20
best candidate 16
Metric calls
1452
target 1440, final eval completed in flight
Structural failures
0
clean final heldout runs

The search gain on Pareto is much larger than the heldout gain, so treat this as useful prompt evidence, not a final production proof. The heldout still improves, mainly by increasing recall and exact matches.

Scores Across Proposals

This is the GEPA search trajectory for the latest run. The blue line is the full Pareto score for accepted proposals, and the green line is the best score seen so far. Best candidate: proposal 27, score 0.6979.

Open full proposal graphs

Heldout Comparison: GEPA Best vs v10 Seed

MetricGEPA bestv10 seedDelta
Mean score0.78820.7686+0.0196
Precision0.87310.8889-0.0158
Recall0.81250.7778+0.0347
Micro-F10.84170.8296+0.0121
Exact matches5046+4
False positives1714+3
False negatives2732-5
Over-label total64+2
Mean predicted labels1.71791.6154+0.1026
Rows7878same

This does not look like a random-more-labels shortcut: mean predicted labels rose by about 0.10 per row, exact matches improved by 4, and precision stayed high. The tradeoff is real, though: GEPA bought recall with 3 extra false positives and 2 extra over-labels.

Whole 330 Comparison

MetricGEPA best repairedv10 seed repairedDelta
GEPA mean score0.73500.7307+0.0043
Micro-F10.82060.8231-0.0025
Precision0.82460.8344-0.0098
Recall0.81670.8120+0.0047
Exact match0.54240.5242+0.0182
False positives110102+8
False negatives116119-3

On the whole 330 rows, GEPA-best narrowly improves the GEPA-style mean score and exact match, but v10 keeps a slightly better micro-F1 because GEPA-best trades 3 fewer false negatives for 8 more false positives.

Run Config

Modelnvidia/Qwen3.6-35B-A3B-NVFP4
Runtimelocalpager-agent via scripts/localpager-classifier
Endpointhttp://127.0.0.1:8000/v1
Concurrency4
Thinkingmedium
Max output tokens8192
Splitsfeedback300 train, pareto60 Pareto, bench78 heldout
Prompt setupv10 overlay-only scaffold seeded from localpager-openclaw-routing-v10-overlay-seed.md
Taxonomyopenclaw-routing-topics.v2.json
Budgetmax_metric_calls=1440, max_candidate_proposals=32, reflection_minibatch_size=4
Reflection LMCodexReflectionLM