Analytics & metrics

Read the performance summary, KPIs, and eval summary on a run's Analytics tab

A run test’s Analytics tab rolls every call in the run into three reads: a performance summary, a set of KPI cards, and an eval summary. None of them replace opening an individual call, covered in Calls & transcripts; they’re for judging the run as a whole before you drop into one conversation.

Read the performance summary

One read is the performance summary, Test Run Performance Metrics, three cards:

  • Pass Rate: the share of calls in the run that came back passing. A number here below what you’d expect from a stable agent is the first sign something regressed
  • Total Test Runs: how many times this run test has been executed. One card, not per-call, so it tells you how many attempts you have to compare, not how any single attempt went
  • Latest Fail Rate: the fail rate of the most recent execution specifically. Read it next to Pass Rate to tell a run that’s always been shaky from one that just got worse

Alongside the cards, Top Performing Scenarios lists each scenario with a score chip: red under 5, orange under 7, green at 7 and above. A scenario sitting in red is where the agent is struggling hardest, and it’s the one worth opening first.

Read the KPIs

The KPI cards differ by channel, because a voice call and a chat conversation are measured differently.

A voice run shows CSAT (the header label for the call’s overall score), Agent Latency in milliseconds, the agent’s words-per-minute pace, how quickly the agent stops talking when the caller cuts in, average turn count, and a talk ratio comparing how much of the call the agent spent talking against the caller. A low CSAT alongside high latency or a lopsided talk ratio usually points at the same root cause: the agent is talking too much, or too slowly, to keep the caller satisfied.

A chat run shows CSAT again, average latency in milliseconds, average turn count, and three token counts: total, input, and output. Token counts matter here beyond cost. A run whose output tokens climb without a matching lift in CSAT is spending more per reply without the conversation actually getting better.

Every field behind these cards, including the ones not surfaced as KPI cards, is listed in Call metrics.

Read the eval summary

The eval summary is one card per eval attached to the run, graphing how that eval scored across every call. This is where you catch an eval that’s consistently weak across the whole run rather than failing on one unlucky call, and it’s the signal that tells you which eval to chase into individual transcripts. Eval types and how each one is configured belong to Evaluation; this tab only shows the scores it produced.

Note

Rerunning a call, from Calls & transcripts, updates the eval summary and KPIs the next time you load this tab.

Dive deeper

Was this page helpful?

Questions & Discussion