Introduction
Speech is rarely continuous. Spontaneous talk is interrupted by brief stretches of silence and by vocalized interruptions such as "uh" and "um," collectively studied under the broader label of disfluency. Within this category, researchers in psycholinguistics and speech science draw a consistent distinction between filled pauses, hesitations occupied by a vocalization, and silent pauses, hesitations marked by an absence of sound. Although both are often lumped together in casual usage as signs of hesitation, the two phenomena have been shown to differ in their production, their timing relative to speech planning, and how listeners interpret them [1][2][3]. This article reviews the research basis for that distinction, along with what is known about how listeners perceive silence during speech, how speakers' own sense of their pausing compares with objective measurement, and what evidence exists for training-based techniques that substitute silent pauses for filled ones.
Filled Pauses: Distinguishing "Uh" and "Um"
The vocalizations transcribed as "uh" and "um" (and their cognates in other languages) have been studied as a distinct class of speech sound since at least the mid-twentieth century. Maclay and Osgood's early corpus study of spontaneous English speech grouped these vocalizations together with repeats, false starts, and other interruptions under the umbrella term "hesitation phenomena," treating them as symptoms of the cognitive work involved in planning an utterance in real time [2]. Boomer's subsequent work linked the timing of such hesitations to units of grammatical encoding, arguing that pausing patterns track the structure of the phrase being planned rather than occurring at random [3]. Goldman-Eisler's program of research extended this line of inquiry into a broader account of pausing as a window onto the cognitive processes underlying spontaneous speech production [4].
The most influential modern treatment of the specific distinction between "uh" and "um" comes from Clark and Fox Tree's 2002 analysis of several corpora of spontaneous English speech. Clark and Fox Tree argued that "uh" and "um" are conventional English words, interjections with their own phonological form, rather than meaningless noise or involuntary leakage. On their account, speakers use these interjections to signal to a listener that a delay in speaking is coming, and to project, at least roughly, how long that delay is likely to be: "uh" is used to announce a minor, relatively short upcoming delay, while "um" is used to announce a major, comparatively longer one [1]. This is a claim specifically about what the filler communicates regarding the pause that follows it, not a claim that "uh" and "um" themselves differ in duration as spoken sounds, nor a general claim about listener attitudes toward either word. Clark and Fox Tree's evidence for this delay-signaling account came primarily from patterns in where the two forms occur in an utterance and from their association with the length of the silence that follows them, drawn from naturally occurring speech rather than laboratory manipulation.
A separate but related line of work has asked what listeners do with this information once they hear it. Fox Tree's earlier experimental study of speech comprehension examined how the presence of "uh" and "um" affects listeners' processing of the words that follow, treating the interjections as potential cues that shape a listener's expectations about upcoming material rather than as simple noise to be filtered out [5]. Brennan and Schober's study of listener processing during spontaneous self-repairs, in which a speaker interrupts a word to correct or replace it, similarly found that the disfluency itself, including the interruption and any accompanying filled pause, carries information that listeners can use, and that naturally disfluent renditions of a repair were processed differently than versions with the disfluency edited out [6]. Together, this body of work supports treating filled pauses as a structured signal with a communicative function, distinct from silence, rather than as an undifferentiated symptom of speech breakdown.
Silent Pauses and How Listeners Interpret Them
Silent pauses have their own, partly separate research history. Rochester's widely cited review of pause research organized the functions that have been proposed for silence in spontaneous speech, distinguishing pauses that appear to be tied to the cognitive demands of planning what to say next from pauses that fall at points of syntactic or discourse structure, such as the boundary between clauses or topics [7]. This structural function of silence is central to Swerts's study of Dutch spontaneous monologues, which found that pauses, including filled pauses, cluster around major discourse boundaries more than around minor ones, and that the acoustic properties of a pause differ depending on how strong a boundary it marks [8]. A related study by Swerts examined prosodic cues, including pausing, at discourse boundaries of differing strength and found that these cues, including pause duration, scale with how substantial a division listeners perceive between one part of a discourse and the next [9]. Taken together, this research indicates that silence in speech is not perceived by listeners as a single undifferentiated category; its interpretation depends on where it falls and what other cues, such as intonation, accompany it.
Whether a given silence is heard as a meaningful boundary or as a sign of speaker difficulty has also been studied directly in conversational settings. Roberts, Francis, and Morgan examined how the combination of inter-turn silence and prosodic features shapes listeners' perception that something has gone wrong, or that there is "trouble," in an ongoing conversation, indicating that pause duration alone does not determine how a silence is interpreted; the surrounding prosodic context substantially shapes whether a gap reads as a normal turn transition or as a sign of interactional difficulty [10]. In a related vein, Susca and Healey asked listeners to describe their own reactions to speech samples spanning a continuum from fluent to markedly disfluent, and found that listeners' perceptions and affective reactions shifted in identifiable ways as disfluency increased, rather than only registering the presence or absence of any single interruption type [11]. These findings collectively suggest that listeners treat pausing, whether filled or silent, as part of a broader interpretive judgment about a speaker's fluency and about what is happening in the interaction, rather than reacting to raw silence duration in isolation.
Speaker Self-Perception Versus Measured Pause Duration
A separate question from how listeners perceive pauses is how accurately speakers perceive their own pausing. Popular writing on public speaking sometimes asserts that speakers dramatically overestimate how long their own pauses last, occasionally attaching a specific multiplier, such as a claim that speakers perceive their silences as three to four times longer than they actually are. No peer-reviewed primary source establishing a specific multiplier or percentage of this kind was identified in the course of researching this article, and that figure should be treated as unverified rather than as an established research finding.
What is better established, though addressing a related rather than identical question, is that speakers' subjective experience of their own performance during moments of anxiety or self-monitoring can diverge from how that performance is actually perceived by an audience. Savitsky and Gilovich's research on public speaking anxiety describes an "illusion of transparency," in which anxious speakers systematically overestimate how visible their nervousness is to onlookers, and found that making speakers aware of this gap between self-perception and audience perception could reduce anxiety and improve some aspects of performance [12]. This finding concerns the perceived visibility of nervousness in general, not the perceived duration of a pause specifically, and it should not be read as direct evidence for any particular pause-duration multiplier. It does, however, indicate that a documented gap exists between a speaker's internal, anxiety-inflected self-assessment and an external observer's assessment during public speaking, which is broadly consistent with the qualitative direction of the popular "perception gap" claim about pauses, namely that a self-monitoring speaker tends to experience a moment of silence as more salient, and by extension potentially longer or more conspicuous, than an outside listener or an objective timer would register it. Establishing whether that qualitative tendency holds specifically for pause duration, and if so by how much, would require a study directly comparing speakers' duration estimates with synchronized objective timing, which was not located during this review.
Converting Filled Pauses to Silent Pauses: Evidence from Behavioral Training
The idea that a filled pause can be deliberately replaced with a silent one is not only a piece of public speaking folk advice; it has a basis in applied behavior analysis research on habit change. Mancuso and Miltenberger applied a simplified habit reversal procedure, a behavioral technique originally developed for tics and other repetitive habits, to filled pauses ("uh," "um," and other verbal fillers, including overused "like") produced during public speaking. The procedure combined awareness training, in which participants learned to notice their own filled pauses, with competing response training, in which participants practiced an alternative behavior to perform whenever a filled pause was about to occur. All six participants in the study showed an immediate reduction in filled pauses following training [14]. Bördlein and Sander later conducted a partial replication of this habit reversal procedure with a shortened training protocol and a smaller sample of undergraduate students, and again observed reductions in filled pauses that were sustained at follow-up, although the researchers noted that some participants had already begun reducing their filled pauses during the baseline period, which limited how cleanly the study could isolate the effect of training itself [15].
A related applied approach comes from communication training research rather than behavior analysis specifically. Ramos Salazar described a classroom technique, the "VP card," used in public speaking instruction to help students notice and reduce vocalized pauses, giving students concrete, low-cost feedback during practice speeches [16]. Although the mechanisms differ somewhat between the habit-reversal studies and classroom feedback techniques like the VP card, both approaches share the underlying premise that filled pausing is a modifiable habit rather than a fixed trait, and that structured awareness of the behavior, paired with practice at an alternative response, can reduce its frequency. None of the studies located frame the alternative response explicitly as "silence" in as many words in the sense of a specifically measured silent-pause substitution; the published descriptions center on reducing filled pauses through awareness and a trained competing response, with a pause in speaking being the practical result of withholding the filler. This is a meaningful caveat: the research literature supports filled-pause reduction as a trainable outcome, but the more specific framing of "substituting a silent pause for a filled pause" as a named, separately measured technique is better described as consistent with, rather than explicitly established by, the available studies.
Measuring and Categorizing Pauses in Speech Research
Because filled and silent pauses are conceptually distinct but occur within the same stream of speech, pause research has had to develop explicit conventions for identifying, timing, and classifying them. Duez's comparative study of pausing across interviews, prepared political speeches, and other speech styles treated silent and filled ("non-silent") pauses as separate but comparably measurable categories, coding their frequency and duration separately across the different speech contexts examined [17]. This kind of dual coding, in which the same passage of speech is marked for both silent gaps and filled interruptions, is a standard feature of pause research methodology, since a passage can contain either type, both, or neither, and conflating the two categories obscures differences in what each is doing functionally.
A recurring methodological problem in this line of research is deciding what minimum duration of silence should count as a "pause" at all, as opposed to an ordinary articulatory gap between sounds, such as the momentary closure involved in producing certain consonants. Hieke, Kowal, and O'Connell examined this problem directly, arguing that some silences researchers had been coding as meaningful "articulatory" pauses were better explained as ordinary phonetic byproducts of certain sound sequences rather than as pauses reflecting planning or hesitation, and cautioning against treating every measurable gap in the acoustic signal as equivalent evidence of a pause [18]. Reflecting this concern, pause studies commonly apply a minimum duration threshold, often somewhere in the range of a few hundred milliseconds, below which a silent interval is treated as an artifact of articulation rather than counted as a pause; the exact threshold varies across studies and research traditions rather than following one single fixed convention. O'Connell and Kowal's later overview of pause research summarized the broader set of decisions this kind of coding involves, including where to place pause boundaries, how to distinguish pauses tied to word searching from pauses tied to syntactic planning, and how to handle pauses that combine both silence and vocalization, such as a silence immediately preceding or following an "uh" [19]. These methodological choices matter for interpreting any given study's results, since studies that define or threshold pauses differently are not always directly comparable to one another.
FAQs
Are "uh" and "um" the same thing, just different words for the same pause?
Research treats them as related but distinct. Clark and Fox Tree's analysis of spontaneous speech argued that speakers use "uh" and "um" as conventional interjections that announce a coming delay in speech and give a rough indication of how long that delay is likely to be, with "uh" associated with shorter delays and "um" with longer ones, rather than treating the two as interchangeable filler sounds [1].
Do listeners always perceive silence as a sign that a speaker is struggling?
No. Studies of pause perception indicate that listeners' interpretation of silence depends heavily on context, including where the silence falls relative to discourse structure and what prosodic cues, such as intonation, accompany it. Research on discourse boundaries found that pausing patterns differ depending on how major a structural boundary is being marked, and research on conversational turn-taking found that the sense that something has gone wrong during a silence depends on co-occurring prosodic cues rather than on the length of the gap by itself [9][10].
Is there solid research showing speakers can train themselves to pause silently instead of using filler words?
There is applied research specifically on reducing filled pauses through habit reversal training, a behavioral technique combining awareness training with practicing an alternative response, which produced measurable reductions in filled pause frequency in a small study and in a subsequent partial replication [14][15]. This research establishes that filled-pause frequency is a modifiable behavior; it does not, however, provide a precise, independently measured account of silent-pause substitution as a distinct outcome separate from the general reduction in filled pauses.