# Testing agents before they talk to real people

ClearTalk's testing runs simulated conversations against a Pathway: a persona you define plays the caller, the agent plays itself, and checks score the outcome. Runs spend credits but never dial anyone. This is how an agent earns the right to go live.

## Designing scenarios

A good scenario is one situation, told concretely:

- **Persona prompt** is a character brief, not a category. "You are Dana, 61, you answered but you're cooking dinner and impatient. You did fill out the solar form last week but barely remember it. You'll stay on if the caller is quick and clear." That produces a realistic conversation; "be an impatient caller" produces a cartoon.
- **Cover the distribution, not the demo.** A real campaign's calls are mostly not the happy path. A solid starter suite: happy path, angry, confused/elderly, voicemail, wrong number, "how did you get my number", price challenger, instant hangup risk. Clone from the built-in gallery where it fits, and use `generate_test_scenario_from_call` to turn real calls — especially real failures — into regression tests.
- **Checks name observable outcomes**: "agent offered a booking time", "agent did not quote a price", "agent ended politely after the third refusal". A check you can't verify by reading the transcript is a bad check.
- **Mark the gate.** The scenarios that must never fail (compliance behaviors, the core objective) get `isRequiredForPromotion: true` — they become the Pathway's publish gate.

## Reading results

- The **overall score** ranks runs; the **per-check verdicts** tell you what to fix. A 0.9 with a failed required check is worse than a 0.7 with all required checks green.
- Read failed transcripts like a call reviewer: find the first turn where the conversation left the rails. The cause is usually one of: a missing fact (fix knowledge/instructions §3), a missing policy for that situation (add one line to objections/exits), or a wrong beat order (fix the conversation shape).
- **Natural-conversation scoring** (opt-in per scenario) grades tone dimensions — naturalness, conciseness, empathy. Use it on the happy path where polish matters; skip it on stress scenarios where surviving is the win.
- A run that **errored** or was cancelled proves nothing — it neither passes nor blocks. Re-run it.

## Cadence

- After every instruction edit: re-run the scenarios that were failing, plus the required set (`run_all_tests` with the subset).
- Before suggesting promotion: `run_all_tests` in full, then `get_promotion_readiness` — it lists exactly which required tests still block.
- Test against **staging** (the default) while iterating; that's the version being edited. Run against `production` only to reproduce a live problem.
- After going live: when a real call fails in a new way, turn it into a scenario with `generate_test_scenario_from_call` and add checks. The suite should grow one regression test per real-world surprise.
