Product Talk · Free post · Building AI products · Discovery & customer research

4 New Evals and 16 Experiment Variants to Fix 1 Customer Complaint

Teresa TorresSep 16, 202620 min
SourceProduct Talk
KindFree post
PublishedSep 16, 2026
Originalproducttalk.org ↗
N:

Teresa Torres describes a three-week investigation prompted by a beta customer's complaint that an AI-generated opportunity solution tree had too many flat, unstructured opportunities under one parent. She built code-based and LLM-as-a-judge evals to measure the problem, calibrated judges against her own labels, and discovered that upstream errors (poorly framed key moments and parent-restating opportunities) confounded the measurement. Across sixteen prompt and model variants, fixing one error type tended to worsen the other, since constraining how parents are framed made the agent reluctant to add sub-groupings. The breakthrough came from moving a deterministic tree-shape check into an audit loop so the agent could self-correct, trading one failure mode against the other. The piece is a practical case study in building reliable AI product features through evals, calibration, and orchestration.

01Key takeaways

  • Calibrate any LLM-as-a-judge against your own labeled examples and check both specificity and recall before trusting its numbers.
  • Upstream errors can corrupt downstream measurements, so identify and measure them before judging the final output.
  • When two error categories pull against each other, prompt tweaks alone may not resolve them; consider orchestration and self-correction loops.
  • Use cheap deterministic checks, like shape or count metrics, as guardrails inside an audit loop instead of adding another LLM call.
  • Work through a fix in measured variants and track every error category at once so gains in one area don't hide regressions in another.
  • A rule of thumb like three to seven key moments gives a quick structural sanity check on generated outputs.

02Key sections

Why the complaint mattered
A customer's offhand request to 'clean up' a flat branch led the author to ask why the agent failed to add structure at its source rather than adding a cleanup button.
Measuring the error with two evals
A deterministic tree-shape check counted children per parent, while an LLM judge looked for missed sub-groupings, with calibration revealing the judge's specificity was poor.
Upstream errors confound the measurement
Poorly framed key moments and parents that restated their children confused the judge, so the author added two new evals to catch those upstream failures first.
Prompt variants and the tradeoff loop
Sixteen variants, including a model upgrade, showed that improving parent framing reduced new parents while encouraging sub-groupings increased poorly framed parents.
Moving from prompts to orchestration
The fix came from an audit-and-retry loop using deterministic tree-shape checks, which let the agent correct its own missed groupings, followed by tuning to curb over-eager single-child parents.

03From the post

“"That is a lot of first layer opportunities there. I wish I could click on this card and be like, 'Clean this up.'" This was a Vistaly customer reviewing her first AI-generated opportunity solution tree. Up until this point, the feedback was extremely positive. But”

“I wish I could click on this card and be like, 'Clean this up.'”Teresa Torres · Product Talk
“It was the first time one of my calibrated judges didn't hold up against production data.”Teresa Torres · Product Talk
“The benefit of a well-structured opportunity solution tree is to help you know what to work on next.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.