HotpotQA demo · UI reconstruction, original reported results retained
01 / CONTEXT & CONTRIBUTION
The problem. My part in it.
Background
Prompt and skill fixes did not prevent recurring failures.
Developers still had to diagnose and correct each run.
More memory added duplicate instructions and retrieval noise.
Optimization had to account for quality, cost and complexity.
Objective
Reuse the reasons for success and failure across tasks.
Retain workflow changes supported by evaluation.
My role
CPO / direction, evaluation criteria and experiment design.
Contributed to implementation, experiments and product decisions.
02 / SYSTEM DESIGN
How it connects.
Processing / data flowIteration / feedbackService / environment boundary
01
A measured change.
Fix the evaluation data and criteria before modifying a workflow. Compare improvements and regressions under the same conditions, then retain or roll back the change.
02
Knowledge that travels.
Connect online knowledge generation with offline optimization. Feed the reasons for success and failure back into the shared knowledge structure.
Optimization workspace · UI reconstruction based on the demo
03 / RESULTS & SCOPE
The outcome.
76.55Internal benchmark aggregate
+23.88Points above baseline
4Evaluation tasks
Internal workflow evaluation
Baseline
52.67
MIPROv2
59.71
TextGrad
61.08
Feedback Descent
66.56
GEPA
69.52
Platform I contributed to
76.55
May 2026 internal evaluation. Team result across HotpotQA, IFBench, HoVer and PUPA under the same conditions. The separate HotpotQA demo above uses a different evaluation scope.