How We Rate AI Speech Coaches: Inside the Speech Agent Benchmark

Cover image for How We Rate AI Speech Coaches: Inside the Speech Agent Benchmark

An open rubric for evaluating AI models on real speech coaching work, built on the Wellspoken Index and modeled on Harvey's Legal Agent Benchmark. Includes the framework, scoring method, and initial results across nine frontier models.

Written byDaniel Park
Published

Summary: We rate AI speech coaches by running them on the Speech Agent Benchmark, a 24-task rubric built on the Wellspoken Index that grades each model on real speech coaching work, like filler breakdowns and pacing diagnoses, instead of single-shot chat replies. It is modeled on Harvey's Legal Agent Benchmark and grades deliverables against expert rubric criteria.

We rate AI speech coaches by running them on the Speech Agent Benchmark (SAB), a 24-task rubric built on the Wellspoken Index that grades each model on real speech coaching work instead of single-shot chat replies. Each task gives the model a recording, a transcript, and a coaching brief, and asks it to produce the same deliverable a human coach would produce: a filler breakdown, a pacing diagnosis, a structure critique, a recommended next drill. The deliverable is graded against expert-written rubric criteria, the same way Harvey grades models on legal work in their Legal Agent Benchmark.

We built SAB because generic chatbot benchmarks tell us almost nothing about speech coaching. A model that scores well on SWE-bench or MMLU can still miss a filler word, miscount pace, or recommend a drill that contradicts the recording. Speech coaching is a long-horizon task with a specific work product, and the field has lacked a shared way to measure it. SAB is our attempt to fix that.

This post covers the framework, a worked example, the scoring method, initial results across nine frontier models, the behavioral patterns that correlate with strong coaching work, and the cost and latency trade-offs that shape what is actually deployable. The full rubric and task set will be open-sourced later this year.

Why We Built an Agent Benchmark

At Wellspoken, we have spent the past two years building AI speech coaches that work inside real recordings. This forced us to think carefully about how to measure model performance on real speech coaching tasks, where the work product is a structured deliverable a speaker can act on.

To date, there has not been a benchmark that illustrates agent progress for long-horizon speech coaching work. Existing evaluations, including SpeechBench, ETS SpeechRater, and our own earlier work on the Wellspoken Index, have graded short-horizon signals: count the fillers, score the pace, rate the pronunciation. Each of those is a measurement. None of them tests whether a model can produce a complete coaching deliverable that a speaker would actually use to practice.

In coding, agent benchmarks have served as an important leading indicator of agent capability. Scores on SWE-Bench Pro, SWE-Bench Verified, and Terminal-Bench 2.0 reflected a step-function improvement around the same time our research team started to feel the shift in practice. The same pattern is now extending beyond coding. Benchmarks such as GDPval, OSWorld-Verified, BrowseComp, MCP Atlas, and Harvey's LAB have made progress legible across knowledge work, computer use, web research, tool use, and legal work.

SAB is intended to provide this same legible index to speech coaching companies, model labs, and researchers. Understanding where agents can do all, some, or none of a coaching task helps companies measure the return on AI investment, identify which dimensions of speaking can be delegated to an agent, and keep the judgment calls, the client relationship, and the live adjustment in human hands.

What Is the Speech Agent Benchmark?

The Speech Agent Benchmark is an open rubric for evaluating AI models on real speech coaching work. Each task gives the model a recording, a transcript, and a coaching brief, and asks it to produce a deliverable a human coach would produce. The deliverable is graded against expert-written rubric criteria, the same structure Harvey uses for legal work.

Three design choices make SAB useful for speech coaching specifically.

It scores the work product, not the chat reply. A coaching deliverable is a structured artifact: a filler timeline, a pacing diagnosis, a structure critique with a recommended next drill. Grading the artifact instead of the prose filters out models that sound good and coach badly.

It is built on the Wellspoken Index dimensions. The Wellspoken Index scores a recording across six observable dimensions: Structure, Conciseness, Confidence, Pronunciation, Filler Rate, and Pace. SAB uses the same six dimensions as the rubric axes, so a model is graded on whether its coaching correctly targets the dimension that is actually costing the speaker.

It uses a strict all-pass standard. A task only counts as resolved if every rubric criterion passes. Partial credit is reported separately as the criterion pass rate, the same way Harvey reports both numbers. The strict number is the one that matters for real coaching, because a coaching report with one wrong filler count is wrong, even if the other ninety criteria are right.

A Worked Example: The Interview Answer Task

To make the structure concrete, consider one task from the Filler Rate category.

The scenario is a senior software engineer preparing for a staff-level promotion interview at a Series B SaaS company. The engineer records a four-minute answer to the prompt "Tell me about a time you disagreed with your manager." The recording, the transcript, and the coaching brief are given to the model.

The brief asks the model to review the recording, identify every filler word by type, map each filler to its syntactic position, assess the speaker's filler rate against the benchmark for practiced speakers, and recommend the next drill. The required output is a review-ready coaching report. It must include a filler timeline with timestamps, a by-type breakdown separating "uh" from "um," a rate-per-minute figure, and a recommended next drill that targets the dimension actually costing the speaker.

The rubric for this task contains 58 criteria covering six coaching moves. The criteria range from straightforward checks, such as whether the report separates "uh" from "um" (a finding from Clark and Fox Tree's 2002 work in Cognition that "uh" signals a minor delay and "um" signals a major delay), to checks on more detailed work product, such as whether the recommended drill targets the weakest dimension rather than the most obvious one, and whether the report avoids the common error of recommending "slow down" as a filler fix.

A task is marked complete only if every criterion passes. A coaching report that catches 57 of 58 rubric items is one criterion short of shippable. The missing criterion could be the one that sends the speaker to the wrong drill for three weeks.

"All-pass grading reflects how real speech coaching is reviewed in practice. A coaching deliverable with one wrong filler count is wrong, even if the other fifty-seven criteria are right, because the speaker may build a practice plan around the wrong number."

The Six Dimensions, Reframed as Agent Tasks

The Wellspoken Index scores a speaker. SAB scores the model that scores the speaker. The six dimensions become six categories of agent task, and each category has four task types, for 24 tasks total.

DimensionWhat the model has to doTask types
StructureIdentify whether the first sentence contains the actual answer, flag run-ups, recommend a frameworkTopic sentence detection, run-up flagging, supporting material assessment, framework recommendation
ConcisenessCount hedging words, flag repetition, recommend cutsHedge detection, repetition flagging, redundancy scoring, compression recommendation
ConfidenceDetect vocal fry, flag up-talk, score projectionVocal fry detection, up-talk detection, breath support assessment, confidence projection scoring
PronunciationScore articulation precision, flag dropped endings, assess intelligibilityEnding completion check, articulation precision scoring, intelligibility rating, accent vs clarity separation
Filler RateCount fillers by type, map to syntactic boundaries, recommend the silent pause swapFiller type breakdown, boundary mapping, filler timeline, replacement drill recommendation
PaceCompute words per minute, flag rushing, recommend pause placementWords-per-minute computation, rushing detection, pause placement coaching, pace variation assessment

The six dimensions of the Speech Agent Benchmark arranged as a hexagonal rubric

The 24 tasks cover the work a human coach does in a single session. A model that aces SWE-bench can still fail the filler type breakdown task, because that task requires counting "uh" and "um" as separate words with different meanings, which is a finding from Clark and Fox Tree's 2002 work in Cognition. The benchmark exists because general intelligence and speech coaching intelligence are different things.

How We Score

Each task is graded against a rubric written by speech coaches on our team. A typical filler detection task has around 60 criteria: did the model count the total fillers correctly, did it separate "uh" from "um," did it map each filler to its syntactic position, did it recommend the silent pause as the replacement, did it avoid the common error of recommending "slow down" as a filler fix.

We report two numbers, the same way Harvey does.

ScoreWhat it measuresWhy it matters
Criterion pass rateThe share of individual rubric criteria the deliverable satisfiesTells you which model gets the most coaching moves right in aggregate
All-pass rateThe share of tasks where every criterion passes, no partial creditTells you which model produces a deliverable a coach would actually ship without editing

The all-pass rate is the honest number. A coaching deliverable with one wrong filler count is wrong, even if the other fifty-nine criteria are right, because the client may build a practice plan around the wrong number. The criterion pass rate is useful for tracking where the field is improving. The all-pass rate is useful for tracking whether the field is ready.

What We Found So Far

We baselined nine frontier models on the SAB holdout set in July 2026. The results mirror Harvey's pattern in one important way: criterion pass rates cluster high, and all-pass rates stay low.

Figure 1. Overall leaders on SAB, ordered by all-pass percentage.

ModelCriterion pass rateAll-pass rate
Claude Opus 5 (max)91.4%21.8%
Claude Fable 5 (max, with fallback)90.9%19.6%
GPT-5.6 Sol (max)89.7%14.2%
Kimi K3 (max)88.5%13.1%
Claude Opus 4.8 (max)87.6%11.4%
Muse Spark 1.2 (xhigh)86.3%9.7%
Gemini 3.6 Flash (high)84.1%6.8%
GLM-5.2 (max)82.9%4.6%
Gemini 3.7 Flash (high)80.4%2.9%

Three findings stand out.

Speech coaching is far from saturated by frontier models. Under the strict all-pass standard, the best model fully resolves 21.8% of tasks end to end. The frontier of model intelligence can produce a complete coaching deliverable on roughly one in five tasks. Read against the strict review standard real coaching demands, this describes a frontier that is improving fast and is not yet ready.

Intelligence is not evenly distributed across dimensions. Improvements in model capabilities are not evenly distributed across the six Wellspoken Index dimensions. Models continue to demonstrate jagged intelligence for specialized coaching work, with high variance across task categories. On SAB, the leaderboard shifts substantially across dimensions, and no single model leads every one. Claude Opus 5 leads on Structure and Conciseness. Kimi K3 leads on Filler Rate and Pace. GPT-5.6 Sol leads on Confidence. Different model families bring different priors to speech coaching, and those priors map onto which dimensions each family can currently complete.

General intelligence and speech coaching intelligence are not the same model. Gemini 3.7 Flash scores above 80% on most general benchmarks and lands near the bottom on SAB. The gap is in the domain-specific work, the same way Harvey found that GPT-5.6 Sol drafts polished legal memos that miss the actual legal instructions. A model that writes clean prose can still miscount fillers.

"No single model is a silver bullet for speech coaching today. Maximizing agent performance on a real coaching workload requires understanding which model family best matches the dimension at hand. The strongest production deployments will be multi-model from the start."

Behavioral Analysis

The way an agent works through a coaching task is as important to understand as the raw task success score. The reinforcement-learning and multi-agent systems literature has long studied agents qualitatively as well as quantitatively, and the same lens is useful here.

Every SAB run produces an agent trace that captures how the model sequenced its actions to produce the final deliverable. The agents we baselined have a fixed action space with five actions: Read (open the recording or transcript), Search (query the transcript for a pattern), Score (compute a metric from the recording), Write (produce or extend the coaching deliverable), and Validate (run an explicit check of the draft against the brief).

We analyzed how those actions combine into behaviors and how those behaviors correlate with outcomes. Five emergent behaviors appeared consistently and meaningfully impacted results.

Listening to the full recording before scoring. Agents that process the recording end to end before producing a score, rather than scoring from the transcript alone, see a 0.7 point average improvement in all-pass score. The signal is that prosodic cues, like pause placement and intonation, are lost in the transcript and matter for coaching.

Cross-checking filler counts against the transcript. Agents that re-query the transcript for filler candidates after the first pass, rather than trusting the initial count, see a 0.9 point average improvement. The behavior catches the common error of conflating "uh" with "um."

Recommending a drill that targets the weakest dimension. Agents that recommend a drill targeting the dimension with the lowest Wellspoken Index score, rather than the most obvious one, see a 1.4 point average improvement. The strongest positive pattern is targeting the coaching to where the speaker actually needs it.

Citing specific timestamps in the feedback. Agents that ground every coaching point in a specific timestamp from the recording, rather than giving generic advice, see a 0.6 point average improvement. Timestamps make the feedback verifiable, and verifiable feedback is what a speaker trusts enough to act on.

Avoiding generic advice. Agents that produce generic coaching like "speak more slowly" or "be more confident," without grounding it in the recording, see a 1.1 point average decrease in all-pass score. Generic advice is the failure mode that most separates polished chat output from shippable coaching work.

The strongest positive signal is targeting the coaching to where the speaker actually needs it. The strongest negative signal is generic advice that could apply to any speaker.

"These patterns look less like arbitrary model quirks and more like recognizable markers of strong coaching work: listening to the recording before scoring, cross-checking the count, targeting the weakest dimension, grounding every point in a timestamp, and avoiding generic advice that could apply to anyone."

Cost and Latency

For production agents, single-dimensional benchmarks indexed only on quality do not capture the full complexity of deployment. We ran additional evaluations that take into account cost, both in dollars and in wall-clock time.

Figure 2. Per-task cost and latency across the nine baselined models.

ModelCost per taskWall time per task
Claude Opus 5 (max)$0.8418.2s
Claude Fable 5 (max)$0.7116.4s
GPT-5.6 Sol (max)$0.5811.7s
Kimi K3 (max)$0.4214.1s
Claude Opus 4.8 (max)$0.4915.9s
Muse Spark 1.2 (xhigh)$0.319.8s
Gemini 3.6 Flash (high)$0.197.2s
GLM-5.2 (max)$0.248.6s
Gemini 3.7 Flash (high)$0.114.3s

The spread is wide on both axes. Claude Opus 5, the highest-performing model by all-pass score, is the most expensive and slowest configuration in the set, at about $0.84 per task and roughly 18 seconds of wall-clock time. Gemini 3.7 Flash is roughly 8x cheaper and 4x faster, and lands near the bottom on the all-pass leaderboard. Production deployments need more than a high benchmark score. They need a high score inside a budget a speaker will tolerate: the cents they will spend per task, and the seconds they will wait for a coaching report. Both budgets shape what is actually deployable, and on SAB both bend sharply across the model frontier.

"Production deployments need more than a high benchmark score. They need a high score inside a budget a speaker will tolerate: the cents they will spend per task, and the seconds they will wait for a coaching report."

What This Means for Picking an AI Speech Coach

Three practical takeaways for anyone choosing an AI speech coach in 2026.

Ask the company which model it uses and how it evaluates that model. A speech coaching product that does not publish its evaluation methodology is a black box. The companies that take this seriously will tell you what they measure and how they score it. The ones that dodge the question usually have nothing to publish.

Prefer products that score the work product, not the chat reply. A model that sounds good in chat can still miscount fillers or recommend a drill that contradicts the recording. The product should show you the structured deliverable, the rubric it uses, and the score breakdown, the same way Wellspoken shows the Wellspoken Index dimension scores.

Treat the all-pass rate as the honest number. A product that reports only a criterion pass rate will look stronger than it is. The all-pass rate is the one that tells you whether the coaching deliverable is correct end to end, which is the only number that matters for real practice.

Next Steps

Closing the gap between today's frontier and reliable speech-agent performance is going to require sustained research across multiple fronts: how agents handle long-horizon coaching tasks, how they internalize the domain knowledge that transfers across the six Wellspoken Index dimensions, and how they deliver that performance inside the cost and latency budgets production deployments actually require. Each of these is its own research direction, and progress on each will compound with progress on the others.

Over time, behavioral analysis will be the tool that makes those directions interpretable to both AI researchers and speech coaches. We cannot improve what we cannot interpret, and on long-horizon coaching work the right unit of measurement is the trajectory as much as the final score. SAB is one of the instruments we are using to do that work.

To facilitate these research directions, SAB is going to keep growing on three fronts.

The benchmark itself. Richer task families covering more coaching scenarios, adjacent professional communication workflows beyond speech coaching, and longer-context recordings. We are partnering with Artificial Analysis to scale up SAB evaluation and publish a regularly-updated leaderboard as new model launches arrive. The leaderboard will be a living record of where the frontier sits on speech coaching work, refreshed as the field moves.

The research program around it. Each trend in this post opens a research direction. We are starting to run those threads with partners to better measure how agents handle long-horizon coaching tasks, study how domain knowledge transfers across the six dimensions, and how to bake cost and latency improvements into agents alongside quality.

Collaboration with AI labs and model providers. We will be working with model providers whose agents we baselined to understand and inform what changes at the model layer move the needle on speech-agent performance.

Key Takeaway

The Speech Agent Benchmark is our open rubric for evaluating AI models on real speech coaching work. It runs each model on 24 tasks across the six Wellspoken Index dimensions, grades the deliverable against expert rubric criteria, and reports both a criterion pass rate and a strict all-pass rate. The early results show the field is improving fast and is not yet ready: the best model fully resolves about one in five coaching tasks. The strongest positive behavior is targeting the coaching to where the speaker actually needs it. The strongest negative behavior is generic advice that could apply to anyone. The benchmark will open-source later this year so other speech coaching companies, model labs, and researchers can run it themselves.

FAQs

Why build a speech coaching benchmark instead of using SWE-bench or MMLU?

SWE-bench measures whether a model can fix a GitHub issue. MMLU measures whether a model can answer multiple-choice questions. Speech coaching is a long-horizon task with a specific work product: a structured deliverable a human coach would produce. General intelligence and speech coaching intelligence are different things, and a model that scores well on generic benchmarks can still miscount fillers or recommend the wrong drill. SAB exists to measure the work speech coaches actually do.

How is the Speech Agent Benchmark different from the Wellspoken Index?

The Wellspoken Index scores a speaker across six dimensions of speaking quality. SAB scores the AI model that scores the speaker. The six dimensions of the Index become the six rubric axes of SAB, and the 24 task types are the agent versions of what a human coach does in a session. The Index measures the speaker. SAB measures the model that measures the speaker.

Will the Speech Agent Benchmark be open-source?

Yes. The rubric criteria, task set, and scoring methodology will be published later this year so other speech coaching companies, model labs, and researchers can run SAB themselves. The version released internally today is the same framework we use to evaluate models before shipping them in the Wellspoken app.

Which AI model is best for speech coaching today?

The current internal leader on the SAB all-pass rate is Claude Opus 5 at 21.8%, with Claude Fable 5 close behind at 19.6%. The criterion pass rate leader is also Claude Opus 5 at 91.4%. The honest reading is that the best model fully resolves about one in five coaching tasks, which means the field is improving fast and is not yet ready. Model choice matters less than the harness and the rubric, which is why we are publishing the methodology alongside the scores.

Can an AI speech coach replace a human coach?

No. The best AI model fully resolves about one in five coaching tasks on the strict all-pass standard. AI speech coaching today is a first-pass tool that accelerates the diagnosis a human coach would do anyway: identifying the dimension costing the speaker, recommending the next drill, and producing a structured deliverable the speaker can act on. The judgment calls, the client relationship, and the live adjustment still belong to a human coach. AI extends what a coach can do. It does not replace what a coach is.

How do you know an AI speech coach is actually helping?

Record yourself answering an unprepared question today, use the AI coach for two weeks, and answer the same question again. Holding the prompt constant is what makes the two recordings comparable. The signals worth watching for in ordinary conversation: people stop asking you to repeat yourself, and you stop noticing the moment where you lost the thread mid-sentence. The Wellspoken app scores each session on the Wellspoken Index and keeps the history, so the comparison across sessions is measured rather than judged by ear.

Daniel Park