Filler Word Frequency in Professional Speech: Norms, Contexts, and Perceptual Thresholds

A review of the peer-reviewed research on how often filler words occur in spontaneous speech, what drives their production, and how their use is perceived across professional and evaluative contexts.

By Mia Torres, Staff Writer, Wellspoken

Introduction

Filler words, non-lexical sounds such as "uh" and "um" and short lexical phrases such as "like," "you know," and "I mean," are a near-universal feature of unscripted speech. Linguists distinguish these from other forms of speech disfluency, such as repetitions and false starts, because fillers appear to serve a communicative function rather than representing pure error. [1] Despite their ubiquity, filler words carry a specific social cost in professional contexts: listeners consistently rate speakers who use more fillers as less competent, less confident, and less prepared, even when the underlying content of the speech is identical. [2] At the same time, experimental work on how listeners actually process fillers in real time complicates any simple story that fillers are purely harmful: some research finds that a filled pause can measurably speed up a listener's recognition of the word that follows it, suggesting fillers do real cognitive work for an audience even as they damage a speaker's perceived competence. [11] This article reviews the peer-reviewed research on filler word frequency, the mechanisms behind filler production, how filler rates vary by speaking context, and the perceptual thresholds at which filler use begins to measurably affect how a speaker is judged.

What Counts as a Filler

Disfluency researchers generally separate fillers into two categories. The first is the filled pause, most commonly "uh" and "um," which linguist Herbert Clark and psychologist Jean Fox Tree argued in a widely cited 2002 study should be treated as conventional English words in their own right rather than meaningless noise. [1] Under their account, "uh" signals to a listener that a minor, short delay is coming, while "um" signals a longer one, and speakers select between them based on how much trouble they anticipate having with what comes next. The second category is the discourse marker, phrases such as "like," "you know," "I mean," and "well," which linguist Gunnel Tottie has argued function as pragmatic markers of planning rather than simple hesitation, appearing more frequently in contexts that demand more real-time cognitive work from the speaker, such as narrative storytelling or unrehearsed task explanation. [3] Tottie's classification builds on Deborah Schiffrin's foundational 1987 study of discourse markers, which analyzed words such as "oh," "well," "and," and "but" as devices that provide structural and interpretive coordinates for a conversation rather than semantic content in their own right, establishing the broader framework within which later researchers situate filler-specific markers like "like" and "you know." [13]

Research using factor analysis on naturalistic speech samples has found that filled pauses and discourse markers behave as statistically distinct categories with different demographic patterns. A 2014 study led by Charlyn Laserna, using audio samples collected unobtrusively over several days via wearable recorders, found that filled pause use ("uh," "um") was consistent across gender and age, while discourse marker use ("like," "you know," "I mean") was more common among women, younger speakers, and speakers who scored higher on measures of conscientiousness. [4] The distinction matters for professional speech coaching because it implies that filled pauses and discourse markers likely have different underlying causes and may respond to different interventions.

Baseline Filler and Disfluency Rates

Establishing a single number for "normal" filler use is complicated by wide variation across studies, populations, and methodologies. One frequently cited estimate, from a large corpus study of task-oriented conversation by Heather Bortfeld and colleagues, put the overall disfluency rate (a broader category that includes fillers alongside repetitions and repairs) at approximately six disfluencies per hundred words in ordinary conversation. [5] The same study found that disfluency rates were not fixed: they varied measurably by the speaker's age, by the relationship between speakers (married partners versus strangers), by the difficulty of the topic being discussed, and by the speaker's role in the conversation (the person giving directions produced more disfluencies than the person receiving them). [5] Linguist and journalist Michael Erard, surveying the psycholinguistic literature on verbal blunders more broadly, has described ordinary conversational speech as containing roughly one disfluency for every ten words, a considerably higher estimate that reflects a broader definition of disfluency and a more informal speaking context than Bortfeld's task-based corpus. [6]

This variation across studies is itself an important finding: there is no single universal "normal" filler rate. Rate depends heavily on how a study defines disfluency, what kind of speech is sampled, and how spontaneous or planned that speech is. Prepared, scripted, or rehearsed speech, the kind typically delivered in a formal presentation, produces measurably fewer disfluencies than spontaneous conversation, because much of what drives filler production is real-time planning under time pressure. [7]

What Drives Filler Production

The dominant explanation for why speakers produce fillers in the first place centers on planning difficulty. In a 1994 experiment, social psychologist Nicholas Christenfeld had participants describe maze routes that varied in how many alternate paths were available. Mazes with more possible routes produced significantly more filled pauses from the participants describing them, supporting the hypothesis that fillers emerge when a speaker faces a more complex decision about what to say next. [8] Notably, Christenfeld's participants still produced filled pauses even while describing the simplest mazes, suggesting that decision complexity is only one contributor and that the underlying rhythm and structure of spontaneous speech production plays a role independent of any specific cognitive load. [8]

This planning-difficulty account is consistent with subsequent research on how filled pauses function for listeners. Susan Brennan and Maurice Williams found that listeners use a speaker's filled pauses as a cue to that speaker's own confidence in what they are about to say: a speaker who fills a pause before answering a trivia question is judged by listeners as less certain of that answer than a speaker who answers fluently, and these listener judgments track reasonably well with speakers' own private ratings of how confident they felt. [9] In other words, filled pauses are not simply noise that listeners filter out. They carry real, extractable information about the speaker's own cognitive state, which is part of why they influence how a listener judges the speaker.

The Temporal Delay Hypothesis

A separate line of research asks a narrower question: once a filler has been produced, what effect does it have on the listener's processing of the word that comes immediately after it. Jean Fox Tree found in a 2001 study of English and Dutch listeners that hearing "uh" measurably speeds up a listener's recognition of the word that follows it, while "um" has neither a clear benefit nor cost, a distinction consistent with "uh" signaling a short upcoming delay and "um" signaling a longer one. [10] Martin Corley and Robert Hartsuiker, in a 2011 study published in PLoS ONE, tested competing explanations for this finding. [11] Their experiments supported what they termed the temporal delay hypothesis: the benefit does not come from the filler acting as a special linguistic signal, but simply from the extra time the filler buys before the target word arrives, time the listener's auditory system can use to finish processing what came before and prepare for what comes next. [11] This finding matters for how filler use should be understood in a professional context. It indicates that the two effects of a filler, the social judgment cost documented in perception research and the auditory processing benefit documented in this study, operate on entirely different mechanisms and are not in tension with each other. A filler can measurably help a specific listener parse a specific sentence in the moment while simultaneously damaging that same listener's overall impression of the speaker's competence.

Neuroscience evidence supports the claim that listeners register disfluency automatically, below the level of conscious judgment. Lucy MacGregor and colleagues used event-related potentials, a measure of the brain's electrical response to a stimulus, to compare how listeners' brains responded to disfluent repeated words against the same words heard in fluent speech. [7] Repeated words following a disfluency produced a measurable, early neural response (occurring within 100 to 400 milliseconds of the word) that was absent when the same words were heard fluently, indicating that the brain registers the presence of a disfluency very quickly and automatically, well before a listener could form any deliberate judgment about the speaker. [7] This body of work reinforces a point that runs through the perceptual research discussed below: disfluency is not a stylistic detail that listeners might or might not notice depending on how closely they are paying attention. It is registered by the listener's speech-processing system as a matter of course.

Perceptual Consequences in Professional and Evaluative Contexts

The research on how filler use affects listener judgment is most developed in contexts where a speaker is being formally evaluated. A 2022 study published in Advances in Physiology Education examined filler use specifically in scientific and academic presentations and found that excessive filler use measurably reduces both a speaker's perceived credibility and their audience's comprehension of the material being presented, independent of the technical accuracy of the content. [2] The study's authors attributed excessive filler use in presentation contexts primarily to speaker nervousness, insufficient rehearsal time, and difficulty retrieving infrequently used technical vocabulary under time pressure, and recommended structural interventions, such as breaking content into smaller rehearsed chunks and deliberately building in silent pauses, as more effective than simply instructing speakers to "stop saying um." [2]

Job interviews represent a second well-studied evaluative context, although the body of directly comparable peer-reviewed research here is thinner than for conversational or presentation speech, and the literature is more mixed. Several published studies report that interview candidates who use fewer fillers tend to be rated as more competent and hireable by evaluators, and that filler use in a candidate's response can measurably shift a rater's assessment of that candidate's communication ability. At the same time, research on how listeners judge fluency more broadly cautions that filler presence alone does not fully predict a fluency judgment: a study by Hans Rutger Bosker and colleagues comparing native and non-native listeners' fluency ratings found that speech rate was the strongest predictor of perceived fluency, with pause-related measures, including filler use, contributing but carrying less statistical weight. [12] This suggests that filler frequency functions as one input among several into how listeners judge a speaker in a high-stakes evaluative setting, not as a single determining factor on its own.

Toward a Professional Threshold

Popular communication-skills writing often cites a specific numeric threshold, a certain number of filler words per minute past which perceived competence is said to drop sharply. The peer-reviewed literature reviewed for this article does not support a single, well-established number of this kind. What the research does establish is a qualitative and reasonably consistent pattern rather than a precise cutoff: filler use has a measurable, negative effect on perceived credibility and comprehension once it becomes noticeable enough to draw a listener's attention away from the content of the message, and that threshold for noticeability itself shifts with context, audience expertise, and how much the speaker's fillers cluster together rather than spread evenly through the speech. [2] Establishing a precise, generalizable per-minute threshold would require a dedicated psychophysical study designed for that purpose specifically, which does not currently appear to exist in the published literature.

Context-Dependent Variation

Filler rates are not static across speaking contexts, they shift with the format and stakes of the communication. The clearest peer-reviewed evidence for this comes from conversational research: Bortfeld and colleagues' corpus study found disfluency rates varying systematically with the speaker's conversational role and the topic's difficulty, which implies that more demanding or higher-stakes communicative tasks tend to produce more disfluency, not less, when speech remains spontaneous. [5] This runs somewhat counter to a common assumption that people simply "clean up" their speech under pressure. What appears to happen instead is a tradeoff: formal, prepared contexts (a rehearsed presentation, a scripted announcement) reduce filler rates because the content itself is planned in advance, while formal but unscripted contexts (a job interview, an unrehearsed question-and-answer session) can retain or even increase filler rates because the speaker is doing real-time planning under evaluative pressure, the exact condition Christenfeld's maze experiment identified as a driver of filled-pause production. [8]

Second-language speakers add a further layer of context-dependence. Bosker and colleagues' comparison of native and non-native listeners found that both groups relied primarily on speech rate when judging fluency, but that non-native listeners weighted pause-related cues, including filler use, somewhat differently than native listeners did, indicating that the same filler rate can be perceived differently depending on who is doing the listening, not only who is doing the speaking. [12] This has a practical implication for professional settings with a mixed native and non-native audience: a filler rate that reads as unremarkable to one listener may be weighted more heavily by another, which argues against treating any single filler-rate number as a universal threshold across audiences.

Academic and technical presentation contexts show a further source of context-dependence tied to subject matter itself. Stanley Schachter and colleagues, extending Christenfeld's planning-difficulty framework, compared filled-pause rates across university lecturers in the humanities, social sciences, and natural sciences, and found that humanities lecturers produced measurably more filled pauses than lecturers in the more formally structured natural sciences. [14] The authors attributed this difference to how constrained a discipline's vocabulary and argument structure are: fields with more rigid, well-defined terminology and argument patterns leave a speaker fewer live choices to hesitate over at any given moment, while fields that require more open-ended, discursive explanation leave more such choices, and therefore more opportunities for a filled pause. This finding reinforces the planning-difficulty account of filler production discussed above and extends it from a single laboratory task to naturalistic academic speech.

Filler perception itself is not entirely fixed, a further complication for any single professional threshold. A 2022 study by Minna Kirjavainen and colleagues found that listeners' perception and even production of "um" shifted measurably after repeated exposure to speech containing high rates of filled pauses, suggesting that what counts as a noticeable or excessive filler rate for a given listener is not a fixed property of the sound itself but is calibrated, at least in part, by that listener's recent listening experience. [15] This has a direct implication for workplace contexts where the same listener, a manager or interviewer, evaluates many speakers in succession: exposure to one heavily disfluent speaker could plausibly shift how that listener perceives the filler rate of the speaker who follows.

Practical and Training Implications

The research reviewed above points toward a specific, evidence-based approach to reducing filler use in professional contexts, one that differs from the common advice to simply notice and suppress individual fillers as they happen. Seals and Coppock, writing specifically about filler reduction in scientific and academic presentations, identify inadequate preparation time and difficulty retrieving infrequently used vocabulary under time pressure as primary drivers of excessive filler use, and recommend structural changes to how a talk is prepared and delivered rather than in-the-moment self-monitoring alone. [2] Their recommended interventions include breaking prepared content into smaller, separately rehearsed segments, deliberately increasing preparation time for material with unfamiliar or technical vocabulary, and practicing the substitution of silent pauses for verbal fillers so that the pause itself, rather than a filled sound, absorbs the moment of planning difficulty. [2]

This approach is consistent with the planning-difficulty account of filler production more broadly. If fillers emerge primarily because a speaker is doing real-time cognitive work, reducing the amount of real-time work required, through preparation, rehearsal, and familiarity with the material, addresses the underlying cause more directly than instructing a speaker to simply avoid saying "um." The temporal delay research described above suggests one further implication: because a silent pause and a filled pause can serve a similar function for the listener, buying processing time before a difficult word, training that substitutes silence for filled sound is asking a speaker to change the form of a pause they were likely already about to take, rather than eliminate the underlying pause altogether. [11]

Conclusion

Filler words are a well-documented, functionally meaningful feature of spontaneous speech rather than simple error. Their frequency varies considerably across individuals, contexts, and speech types, with prepared and rehearsed speech producing measurably fewer fillers than spontaneous conversation or unscripted evaluative speech such as job interviews. The research is consistent, however, on the perceptual consequences of filler use in professional and evaluative settings: listeners reliably use filler frequency as one signal, among several, in judging a speaker's confidence, competence, and preparedness, and excessive filler use measurably reduces both perceived credibility and audience comprehension in presentation contexts specifically. The mechanisms driving filler production, primarily real-time planning demand, help explain why filler rates rise in exactly the high-stakes, unscripted situations where the perceptual cost of using them is highest.

FAQs

How often do people use filler words in professional speech?

There is no single agreed-upon rate. Corpus research on ordinary conversation puts overall disfluency, including fillers, repetitions, and repairs, at roughly six per hundred words, while broader surveys of the psycholinguistic literature describe informal speech as containing closer to one disfluency in ten words. Rehearsed, prepared speech such as a formal presentation produces measurably fewer fillers than spontaneous conversation.

Is there a specific number of filler words per minute that hurts credibility?

The peer-reviewed literature does not establish a single validated per-minute threshold. What is well established is that excessive filler use measurably reduces perceived credibility and audience comprehension once it becomes noticeable, with that threshold shifting by context, audience, and how the fillers are distributed through the speech, rather than a fixed number applying universally.

Are filler words a sign of poor communication skill?

Not inherently. Filler production is primarily driven by real-time planning demand rather than by a general communication deficit, and some research finds that filled pauses can even measurably speed up a listener's processing of the word that follows. The perceptual cost of fillers comes from how listeners judge them socially, not from evidence that they impair the underlying message.

Mia Torres