By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres introduces a practical test harness for running AI product evals. Manually running every input through a prompt, saving outputs, scoring them against evals, and repeating for each prompt variant is tedious, so she built a tool to automate the loop. The post walks through a worked example using a fictional houseplant app that turns interview transcripts into stories and status labels, a workflow prone to fabricated quotes, invented facts, and wrong labels. It shows how the three-step process of inspecting outputs, counting errors with evals, and running a baseline-versus-variant experiment works in practice. The key point is that a change is only an improvement if it reduces every error type, not just one. The full tool and repo are gated to paying members, so the free post gives the framing and walkthrough without the code.
01Key takeaways
- Automate the repetitive loop of running inputs, saving outputs, and scoring evals so you can compare prompt variants quickly.
- Read outputs side by side with their source material before counting errors, so you know what failures actually look like.
- Give each error type its own eval so failures can be counted and tracked separately.
- Treat a prompt change as an improvement only if it lowers every error rate, not just the one you targeted.
- Validate LLM-as-a-judge evals against a hand-labeled calibration set, since the judge is also an LLM that can be wrong.
- Keep failure examples readable, so you can see why a score changed and whether your eval measures what you intend.
02Key sections
- The practical problem
- Testing prompts by hand means repeatedly running inputs, saving outputs, and scoring each eval for every variant, which quickly becomes tedious. Torres built a harness to automate that loop.
- A worked example
- A fictional houseplant app turns interview transcripts into stories and status labels. The synthetic transcripts include messy details, like a sister paying for a subscription, designed to expose typical LLM errors.
- Three steps from the guide
- The process is to inspect outputs by reading them against sources, count each error type with its own eval, then run an experiment comparing a baseline and a variant prompt.
- Using the harness on your own workflow
- Each eval is a modular function that the harness calls, so new error checks can be added without changing the core tool. Experiments are defined in simple config files.
- Access
- The harness and worked example are restricted to paying members, so the free post ends with a subscription prompt.
03From the post
“My in-depth AI evals guide explains what evals are and how to build them. This post gives you the tool I use to run them. Here's the practical problem. To get a baseline, you have to run every input through your prompt, save each output, run every”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.