← All work
AGENTS / EVALUATION2026

Agent
Optimization

Turn execution traces into reusable knowledge.

AI PRODUCT / CPO

Case study by

HotpotQA demo · UI reconstruction, original reported results retained
HotpotQA demo · UI reconstruction, original reported results retained

The problem.
My part in it.

Background

  • Prompt and skill fixes did not prevent recurring failures.
  • Developers still had to diagnose and correct each run.
  • More memory added duplicate instructions and retrieval noise.
  • Optimization had to account for quality, cost and complexity.

Objective

  • Reuse the reasons for success and failure across tasks.
  • Retain workflow changes supported by evaluation.

My role

  • CPO / direction, evaluation criteria and experiment design.
  • Contributed to implementation, experiments and product decisions.

How it connects.

Agent learning & optimization architectureSession experience → structured knowledge → contextual execution → measured improvementAgent learning & optimization architectureSession experience → structured knowledge → contextual execution → measured improvementLAYER 01 / SESSION-BASED WISDOM GENERATIONLAYER 02 / WISDOM CURATIONSHARED KNOWLEDGELAYER 03 / EVALUATION-DRIVEN OPTIMIZATIONAgent sessions — Attempts / errors / feedbackAgent sessionsAttempts / errors / feedbackExtract & structure — Execution records → reusable insightsExtract & structureExecution records → reusableinsightsSkills & strategies — Reusable conditions + actionsSkills & strategiesReusable conditions + actionsTask + project context — Retrieve relevant knowledge / Resolve relationshipsTask + project contextRetrieve relevant knowledgeResolve relationshipsExecution plan — Workflow → step → skill / Context-specific guidanceExecution planWorkflow → step → skillContext-specific guidanceKnowledge Graph DB — Skills / strategies / trajectoriesKnowledge Graph DBSkills / strategies / trajectoriesScaffold / Adapt — PRD + data + workflowScaffold / AdaptPRD + data + workflowEvaluate → repair — Prompts / code / orchestrationEvaluate → repairPrompts / code / orchestrationValidate & learn — Keep / rollback / generalizeValidate & learnKeep / rollback / generalizeFeedback from execution becomes the next reusable asset.
Processing / data flowIteration / feedbackService / environment boundary
01

A measured change.

Fix the evaluation data and criteria before modifying a workflow. Compare improvements and regressions under the same conditions, then retain or roll back the change.

02

Knowledge that travels.

Connect online knowledge generation with offline optimization. Feed the reasons for success and failure back into the shared knowledge structure.

Optimization workspace · UI reconstruction based on the demo

The outcome.

76.55Internal benchmark aggregate
+23.88Points above baseline
4Evaluation tasks

Internal workflow evaluation

Baseline
52.67
MIPROv2
59.71
TextGrad
61.08
Feedback Descent
66.56
GEPA
69.52
Platform I contributed to
76.55

May 2026 internal evaluation. Team result across HotpotQA, IFBench, HoVer and PUPA under the same conditions. The separate HotpotQA demo above uses a different evaluation scope.

NEXT PROJECT / 02AI Nutritionist