The most common questions people ask about interview speech are comparative: how long should my answers be, do I ramble more than other people, and how many filler words is normal. These are empirical questions, and they are usually answered with convention rather than data. The frequently repeated guidance that an interview answer should run one to two minutes has no published corpus behind it.
Research on spoken disfluency has largely been conducted on conversational corpora, task-oriented dialogue, and recorded public speaking.[1][2][6] Comparatively little descriptive data exists on how people speak in a practice interview, where the speaker knows the interlocutor is not a real employer but the evaluative framing is retained.
This study reports descriptive statistics from a corpus of AI-led mock interview sessions, so that the comparative questions above can be answered with a distribution instead of a rule of thumb. It makes no causal claims and tests no hypothesis.
Method
Corpus
The corpus consists of mock interview sessions conducted through Wellspoken, a communication training application, in which a voice agent conducts a role-appropriate interview and the participant answers aloud. Sessions are self-initiated by users for practice; no session was solicited or designed for this study.
A random sample of 1,000 sessions was drawn from completed sessions containing at least four conversational turns. The sample was drawn programmatically using a uniform random sampling operator over the qualifying population, without stratification.
| Parameter | Value |
|---|---|
| Sessions sampled | 1,000 |
| Distinct participants | 848 |
| Candidate answers analysed | 16,928 |
| Collection period | 29 October 2025 to 19 September 2026 |
| Median session duration | 394 seconds |
| Inclusion criterion | Completed session, at least 4 turns |
Measures
Each conversational turn is labelled by speaker role. Only participant turns were analysed. For each turn, word count was computed by whitespace tokenisation. For each session, filled-pause and discourse-marker tokens were counted by case-insensitive whole-word matching over the concatenated participant text, using the token set um, uh, er, hm, like, you know, I mean, basically, actually, kind of, sort of. Rates are expressed per 100 words, following the convention used in conversational disfluency research.[1]
Responses to the opening question were isolated by matching agent turns against the patterns tell me about yourself, tell me a bit about yourself, and walk me through your background/resume, then taking the first participant turn that followed.
Limitations
These limitations are material and should be read before the results.
- Transcript-derived measurement. Counts come from automatic speech recognition output. Recognisers vary in whether they retain filled pauses, and some normalise them away. The filler rates below are therefore a lower bound, and the um/uh figures in particular should not be compared directly with hand-annotated corpora.
- Lexical ambiguity. Like, actually and basically have non-filler senses that whole-word matching cannot separate. Reported counts for those tokens overstate filler use by an unmeasured amount. Um and uh do not have this problem.
- Self-selected participants. Users of a speech-training application are not a random sample of job candidates. They are people who already believe their speaking needs work, which plausibly biases the corpus toward more disfluent speakers.
- Practice, not performance. The interlocutor is a voice agent. Evidence that disfluency rates shift with audience and setting[6] implies these figures need not transfer to interviews with human interviewers and real stakes.
- No outcome data. Nothing here links answer length or disfluency to hiring outcomes, and no such inference should be drawn.
- Single platform. One application, one agent design, one population.
- No longitudinal comparison is possible. Monthly medians of the measured filler rate across the full qualifying population move between 0.49 and 3.72 per 100 words, with step changes that align with transcription pipeline changes rather than with any plausible behavioural shift. Filler counts derived from automatic transcripts are therefore comparable within a fixed pipeline and not across time. This is reported because it bears directly on a live question in the field: several 2026 surveys report that AI use is making spontaneous speech harder,[7] and corpora of this kind are the obvious place to look for confirmation. Transcript-derived disfluency counts cannot supply it unless the recognition pipeline is held constant and documented.
Results
Answer length is bimodal, and the mean describes almost nobody
The median candidate answer ran 17 words. The mean ran 53.6. A gap of that size between median and mean indicates strong right skew, which the percentiles confirm.
| Statistic | Words |
|---|---|
| 25th percentile | 5 |
| Median | 17 |
| 75th percentile | 72 |
| 90th percentile | 147 |
| 99th percentile | 386 |
| Longest single answer | 2,010 |
52.2% of all answers ran under 20 words. 9.6% ran over 150 words, and 2.0% ran over 300. The corpus is not a population of people who talk too much, nor of people who say too little. It is both, often within the same session: the median longest answer in a session was 111.5 words, against that session-level median of 17.
At a conversational speaking rate, a 300-word answer takes roughly two minutes of uninterrupted speech, and the longest answer recorded here would take approximately thirteen.
The opening question produces the longest answers
The opening question was identified in 877 of the 1,000 sessions.
| Statistic | Words |
|---|---|
| Median | 49 |
| 75th percentile | 103 |
| 90th percentile | 188 |
| Proportion over 150 words | 14.6% |
The median response to tell me about yourself was roughly three times the median answer elsewhere in the same corpus. Nearly one in seven ran past 150 words.
Participants start clipped, then expand
The first participant turn in a session had a median of 5 words. Every subsequent turn had a median of 20. The opening turn is typically a greeting or a short acknowledgement rather than a substantive answer, and substantive speech begins only after the agent has asked its first real question.
Filled-pause rates
| Statistic | Fillers per 100 words |
|---|---|
| Median | 2.10 |
| Mean | 3.00 |
| 75th percentile | 4.56 |
| 90th percentile | 7.55 |
| Sessions with zero detected fillers | 21.1% |
Token frequency across the corpus:
| Token | Occurrences |
|---|---|
| uh | 12,109 |
| um | 12,042 |
| like | 4,870 |
| you know | 2,899 |
| actually | 1,412 |
| kind of | 913 |
| basically | 365 |
| sort of | 329 |
| I mean | 314 |
| hm | 311 |
| er | 29 |
Uh and um occur at almost identical frequency, together accounting for 71% of all detected tokens. Clark and Fox Tree argue that the two are not interchangeable, with uh signalling a shorter anticipated delay than um.[2] This corpus cannot test that claim, since it lacks the pause-duration measurements the distinction requires, but the near-parity is itself notable given that the two are often treated as a single category in applied advice.
The distribution again matters more than the centre. A fifth of sessions contained no detected filler at all, while the top decile ran above 7.55 per 100 words, more than three times the median.
Reference table
The figures below are provided so that an individual measurement can be located within this distribution. They describe this corpus under the limitations stated above and are not norms for interview speech in general.
| Measure | 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|---|
| Words per answer | 5 | 17 | 72 | 147 |
| Fillers per 100 words | 0.00 | 2.10 | 4.56 | 7.55 |
| Response to the opening question (words) | — | 49 | 103 | 188 |
A speaker whose answers routinely exceed 150 words is in the top tenth of this corpus by length. A speaker measuring above roughly 7.5 fillers per 100 words is in the top tenth by disfluency. A speaker at 2 per 100 words is at the median, which is to say unremarkable.
Discussion
The descriptive pattern with the clearest practical reading is the bimodality of answer length. Advice aimed at interview candidates is typically framed as a single correction, either to stop rambling or to stop giving one-line answers. In this corpus both behaviours are common, and the median session contains an answer roughly six times longer than that session's typical answer. The problem, at least in practice conditions, is better described as inconsistent calibration to the question than as a stable speaking trait.
The opening-question result is consistent with that reading. Tell me about yourself is an unusually open prompt with no implicit length constraint, and the answers it produced were three times the corpus median.
On disfluency, the headline caution is methodological. A measured median of 2.10 per 100 words sits below rates reported for conversational corpora,[1] but the comparison is not valid: automatic transcription discards an unknown share of filled pauses, and this figure is a floor rather than an estimate. What the data support is relative rather than absolute description: the spread across speakers is wide, uh and um dominate the token set, and one in five sessions shows no detectable filler.
The broader literature suggests caution before treating filler frequency alone as a deficiency measure. Filled pauses carry information for listeners about upcoming delay and speaker uncertainty,[2][3] their frequency varies systematically with speaker age, gender and personality,[4] and rates shift with the public or private character of the setting.[6] Frequency counts have nonetheless been used as an evaluative measure in professional and scientific speaking contexts.[5]
Data availability
Aggregate statistics are reported in full above. The underlying transcripts contain personal information volunteered by participants during practice and are not released.
Notes on reuse
This page reports original descriptive data and may be cited. Figures should be reported with the sample definition attached: 1,000 randomly sampled AI mock interview sessions from 848 participants, October 2025 to September 2026, measured from automatic transcripts.