Speaker Credibility and Verbal Disfluency: A Review of Perception Research in Professional Contexts

A synthesis of experimental and observational research, from 1964 to the present, on how listeners infer a speaker's competence, intelligence, and credibility from filled pauses and other speech disfluencies.

By Margaret Whitaker, Senior Writer, Wellspoken

Introduction

Public speakers, courtroom witnesses, job candidates, and broadcasters are routinely judged not only by what they say but by how smoothly they say it. Speech disfluencies, the interruptions, restarts, filled pauses such as "um" and "uh," and hesitations that punctuate almost all spontaneous talk, are a normal feature of unscripted language production. Yet listeners do not treat them as neutral noise. A substantial body of research in communication studies, social psychology, and psycholinguistics has examined whether, how much, and under what conditions disfluent speech lowers a speaker's perceived credibility, competence, and authority. This article synthesizes that literature, tracing it from early source-credibility experiments in the 1960s through recent work using synthetic speech and large online samples, to describe the mechanism by which listeners form these judgments, the professional settings in which the effect has been documented, and the limits of what is currently known. Because the underlying finding concerns how an audience evaluates a speaker in real time, this literature is also frequently referenced in discussions of public speaking apprehension: it identifies a specific, externally observable channel, the rate and placement of disfluencies, through which a speaker's momentary hesitation could plausibly translate into a negative audience judgment, independent of the prepared content of a talk.

How Listeners Infer Competence from Disfluency Rate

Theoretical accounts of the disfluency-credibility link generally start from a "fluency principle": the idea that the ease or difficulty with which a listener processes incoming speech is itself used as information about the speaker and the message, independent of the semantic content [1][2]. Communication researchers Marko Dragojevic and Howard Giles, who developed this account primarily to explain listener reactions to foreign-accented speech, argue that the underlying mechanism is not specific to accent: any property of a speech signal that increases the listener's real-time processing effort, whether unfamiliar pronunciation, background noise, or frequent hesitation, can trigger the same downstream inference [1][2]. When speech is halting, marked by frequent filled pauses, repetitions, or false starts, listeners experience greater processing effort, and that subjective difficulty tends to be attributed to qualities of the speaker, such as low confidence or weak command of the subject, rather than correctly attributed to the surface form of the utterance [1][2].

This account rests on two separate empirical claims: first, that disfluency rate is not random noise but is systematically produced in relation to the cognitive demands on a speaker, and second, that listeners register and use it as a cue.

On the production side, filled pauses increase when a speaker faces more lexical or conceptual choices. In an early experimental test, Christenfeld had participants describe simple and complex maze routes and found that mazes with more possible paths produced more "um"-type filled pauses, evidence that hesitation tracks the number of options being computed in real time rather than occurring at random, though speakers still produced filled pauses even for the simplest mazes, indicating that choice among options is only one contributor to the phenomenon [3]. A related finding, based on naturalistic observation of university lecturers across ten academic disciplines, showed that instructors in the humanities, whose subject matter involves more open-ended choices among concepts and modes of expression, produced more filled pauses per lecture than instructors in the natural sciences, with social-science instructors falling in between [4]. Anxiety and self-monitoring also modulate disfluency rate: in a field study conducted with bar patrons, the number of filled pauses produced during a short interview declined as estimated blood alcohol level rose, a pattern the authors interpreted as evidence that filled pauses partly reflect a speaker's self-conscious monitoring of their own speech, monitoring that alcohol is thought to suppress [5].

On the listener side, several studies show that disfluencies and related prosodic cues function as legible signals of a speaker's internal state. Brennan and Williams played listeners a set of spontaneous verbal responses and asked them to judge how likely each response was to be correct, a measure the authors term the "feeling of another's knowing." They found that this judgment tracked the presence of filled pauses, rising intonation, and longer response latencies in the speaker's delivery: answers delivered with rising intonation or after a longer delay were rated as less likely to be correct, even when the underlying accuracy of the speaker's answer was held constant across conditions. The effect also depended on the type of utterance, longer latencies raised confidence ratings for non-answers but lowered them for answers, indicating that listeners do not apply a single blanket rule but interpret disfluency cues in combination with what kind of utterance is being produced [6]. Complementary work on the placement of disfluencies within an utterance found that a disfluency preceding a noun phrase increases listeners' expectation that the upcoming referent will be new or relatively unpredictable information, showing that listeners parse disfluencies as meaningful discourse signals rather than discarding them as production noise [7].

The Direction and Rough Size of the Credibility Gap

The earliest and most direct experimental test of the disfluency-credibility relationship was conducted by Miller and Hewgill, who manipulated the frequency of nonfluencies, repetitions and vocalized pauses, in otherwise identical taped speeches and had listeners rate the speaker on standard source-credibility scales. As the rate of nonfluency rose, listeners' ratings of the speaker's credibility fell [8]. Sereno and Hawkins replicated this basic pattern with a similar manipulated-nonfluency design and found that heavier nonfluency also lowered listeners' attitude toward the speech topic itself, indicating that the effect is not confined to judgments of the speaker but can extend to evaluation of the content being delivered [9]. Around the same period, Addington's research on vocal delivery more broadly, examining variables such as speech rate, pitch variety, and vocal quality rather than disfluency specifically, found that several dimensions of a message's vocal delivery shift listeners' ratings of a speaker's competence and sociability, situating disfluency within a wider family of paralinguistic cues that listeners draw on when forming credibility judgments [10].

More recent work using synthesized speech has isolated the effect of filled-pause placement from confounds such as a speaker's natural voice or personal history. Kirkland and colleagues independently manipulated filled-pause location, pitch, and speech rate in synthetic utterances and found that inserting a filled pause, particularly mid-utterance rather than at the start of an utterance, reliably lowered listeners' ratings of the speaker's confidence, with the size of that effect comparable to the effect produced by changes in pitch or speech rate [11].

Across this literature, the direction of the effect, more disfluency corresponding to lower perceived credibility or confidence, is consistently replicated from the 1960s through recent synthetic-speech studies. However, no single, universally agreed percentage or magnitude for the size of the "credibility gap" between fluent and disfluent speakers could be verified in this literature. Individual studies report their results on different scales, some using continuous 0-100 sliders, others using multi-point Likert-type credibility inventories, and others using standardized statistical effect sizes, which are not directly interchangeable with one another and cannot be collapsed into a single cross-study percentage without further meta-analytic work of the kind that has not been identified for the verbal-disfluency literature specifically. The size of the effect also varies by disfluency type, delivery channel, and rating scale used, and it is more accurately described as a small-to-moderate but consistently replicated effect than as a fixed figure.

Fillers, Perceived Intelligence, and Perceived Authority

A related line of research asks not just whether disfluency lowers global "credibility" ratings but specifically whether it changes attributions of intelligence, competence, and authority.

In a courtroom-simulation paradigm, Erickson, Lind, Johnson, and O'Barr manipulated witness testimony to be delivered in a "powerful" style, with few hedges, hesitation forms, or intensifiers, or a "powerless" style, with frequent hedges, hesitation forms including filled pauses, and rising "questioning" intonation, while holding the substantive content of the testimony constant. Listeners rated witnesses who used the powerless style, of which hesitation disfluency was one defining component, as less credible than witnesses who used the powerful style, though the size of this effect was larger when the listener and the witness were the same sex than when they were of opposite sexes, indicating that the credibility penalty for a hesitant, hedge-heavy speech style is not entirely independent of who is listening [12]. Engstrom extended this logic to broadcast journalism, inserting varying numbers of nonfluencies into otherwise identical newscast scripts delivered on audiotape and on video. As the number of nonfluencies rose, listeners' ratings of the newscaster's competence and dynamism dropped significantly, an effect that in the audiotaped condition reached statistical significance for a script containing nine nonfluencies, though ratings of the newscaster's trustworthiness were comparatively unaffected, suggesting that disfluency damages perceived competence more strongly than perceived honesty [13].

A separate, more recent strand of work examines a related but distinct construct: the fluency of the auditory signal itself, for example, clear versus distorted or "tinny" audio quality, rather than the fluency of word choice. Using large preregistered online samples, Walter-Terrill, Ongchoco, and Scholl found that listeners who heard a personal statement played through simulated poor-quality audio rated the speaker as less hireable, less credible, and less intelligent than listeners who heard the identical statement in clear audio, even though word-for-word comprehension of the message did not differ meaningfully between conditions [14]. Because that study manipulated microphone and signal quality rather than the presence of verbal fillers or hesitations, its findings speak to a broader "processing fluency" mechanism that is related to, but conceptually distinct from, the verbal-disfluency research reviewed elsewhere in this article. The authors frame the result as evidence that superficial properties of a speech signal, whether extrinsic sound quality or, by extension, intrinsic verbal disfluency, can bias higher-level social judgment through a shared processing-ease pathway.

Effects Across Professional Contexts

The disfluency-credibility relationship has been examined across several distinct professional settings, with broadly convergent but not identical results.

Broadcast journalism shows the pattern described above: Engstrom's newscast study found that competence and dynamism ratings of a newscaster fell as scripted nonfluencies increased, an effect that held for both audio-only and televised presentation of the same script [13].

In legal settings, beyond the powerful- and powerless-speech-style research conducted in simulated courtrooms [12], a study of witness testimony presented through a simulated videoconferencing "virtual court" system found that lowering the audio quality of a witness's recorded statement, again a signal-fluency rather than verbal-disfluency manipulation, led mock jurors to rate the witness as less credible and to weigh the witness's evidence less heavily in verdict-relevant judgments [15]. Considered together, this pair of legal-context studies suggests that both the verbal form of testimony and the acoustic clarity with which it is delivered can independently affect how credible a witness is judged to be.

In business and hiring contexts, DeGroot and Motowidlo videotaped structured employment interviews with working managers, in one sample from a utility company and in a second, independent sample from a news-publishing company, and had trained coders rate vocal cues including pitch, pitch variability, speech rate, and pausing. A composite of these vocal cues correlated significantly with supervisors' later job-performance ratings in both samples and with interviewers' in-the-moment judgments of the candidate in the second sample, and statistical mediation tests indicated that part of this relationship ran through interviewers' attributed trust and credibility toward the candidate: candidates whose vocal delivery, including their pausing pattern, created a stronger impression of credibility tended also to be rated as stronger performers [16]. A more recent interview study that isolated filler words specifically found a more limited effect. Removing the filler words "um" and "ah" from an otherwise identical interview transcript did not significantly change observers' ability to detect a candidate's anxiety, nor did it change their performance ratings; a candidate portraying high anxiety was rated as more anxious and lower-performing whether filler words and other vocal cues were present or absent, suggesting that in some interview contexts other cues can carry more evaluative weight than filler-word rate alone [17].

In science communication and academic presentation, Seals and Coppock, writing for an audience of physiology educators, reviewed evidence that excessive filler use in scientific talks reduces both the presenter's perceived credibility and audience comprehension of the material, and outlined common causes of filler use, including time pressure and inadequate rehearsal [18]. In a related but signal-quality-focused study, Newman and Schwarz found that research findings presented in a recording with poor audio quality were judged by listeners to be less important, and the presenting researcher less competent, than identical findings presented with clear audio, again isolating a processing-fluency effect adjacent to, but distinct from, verbal disfluency [19]. Separately, the naturalistic observation of university lecturers described earlier found that filled-pause rate varies systematically by academic discipline [4], a finding about the production of disfluency rather than its perception, but one that suggests any perception-based credibility penalty could interact with a lecturer's field in ways not yet directly tested by the perception studies reviewed here.

No peer-reviewed study specifically testing the effect of filler words on perceived physician credibility in clinical encounters was identified for this review. This remains a comparative gap in the published literature relative to the courtroom, broadcast, and business-interview settings summarized above.

Does Listener Identity Change the Effect?

Whether the disfluency-credibility relationship differs according to characteristics of the listener, such as the listener's age, profession, or cultural and linguistic background, is far less thoroughly documented than the basic effect itself. Most of the classic experiments described above drew their listener samples from relatively homogeneous populations, typically undergraduate students at North American universities, and were not designed to test listener-side moderators such as age or occupation.

The clearest documented moderator in the broader literature concerns not the listener's demographic background as such, but the relationship between listener and speaker language background. Lev-Ari and Keysar found that native English-speaking listeners rated simple factual statements as less true when the statements were read aloud by a non-native speaker with a foreign accent than when read by a native speaker, even though the speakers were merely reciting sentences composed by someone else. The size of the effect tracked listeners' subjective processing difficulty in understanding the accented speech, and it was reduced, though not eliminated, when listeners were explicitly told that any difficulty was attributable to the accent rather than to the speaker's own unreliability [20]. A broader meta-analysis by Fuertes and colleagues, statistically pooling results across many individual accent-perception studies, similarly found that speakers whose accents are perceived as non-standard or foreign are consistently rated lower on status- and competence-related dimensions by listeners than speakers with a standard or native accent, an effect distinct from verbal disfluency but understood, under the fluency-principle account developed by Dragojevic and Giles, as operating through an overlapping listener mechanism centered on processing ease rather than through explicit prejudice against a particular group [21][1][2].

Because accent and verbal disfluency are conceptually different properties of speech, these findings should be read as evidence that listener judgments are broadly sensitive to processing difficulty arising from multiple sources, rather than as direct evidence that disfluency itself is judged differently by listeners of different linguistic backgrounds. No study identified for this review directly compared how listeners of different ages, professions, or cultural backgrounds weigh a speaker's filler rate specifically, so claims about such differences cannot currently be supported on the basis of the disfluency-perception literature reviewed here.

Boundary Conditions and Counter-Evidence

Not all disfluency research supports a simple "more fillers, lower credibility" reading, and a full review needs to account for work that complicates that picture.

Several psycholinguistic studies find that filled pauses can carry processing benefits for listeners rather than only costs. Fox Tree found that hearing the filler "uh" sped up listeners' subsequent word recognition, functioning as a cue that a short delay was coming, while the filler "um" showed neither a measurable benefit nor a cost, evidence that listeners treat different filler types as distinct, informative signals rather than as uniformly negative noise [22]. Brennan and Schober found that when a speaker's utterance was interrupted mid-word and then repaired, listeners identified the intended target object more quickly, and no less accurately, when the interruption included a filled pause than when it did not, indicating that a filled pause can actively help a listener process a self-correction in real time [23]. And as noted above, at least one applied interview study found that removing filler words from an otherwise identical transcript did not meaningfully change observers' judgments of a candidate's anxiety or performance, suggesting that the credibility penalty associated with disfluency is not uniform across evaluative contexts and can be outweighed by other cues [17].

Taken together, this body of work suggests that the disfluency-credibility relationship is real and has been repeatedly replicated in controlled experiments dating back to the 1960s, but that it is also context-dependent. It interacts with delivery channel, disfluency type and placement, the listener's task, and the presence of competing cues, rather than operating as a fixed, universal penalty on every disfluent speaker in every setting.

Summary

Across six decades of research spanning classical source-credibility experiments, courtroom-simulation studies, broadcast-journalism research, employment-interview studies, and recent large-sample and synthetic-speech work, listeners consistently rate more disfluent speakers as less competent, less credible, or less confident than more fluent speakers delivering the same content. Current theory attributes this to a general "fluency principle," in which the difficulty of processing a speech signal is misread as a property of the speaker rather than of the signal itself. The effect has been documented most directly in courtroom, broadcast-journalism, business-interview, and academic or scientific-presentation contexts, and a closely related literature on signal-level, rather than verbal, fluency, and on speaker accent, shows a similar pattern operating through what researchers argue is an overlapping cognitive mechanism. At the same time, several studies show that specific fillers can carry real-time processing benefits for listeners, and the credibility penalty for disfluency is not observed in every applied context that has been tested. Direct evidence on whether the size of the effect varies systematically with a listener's age, profession, or cultural background remains limited in the published literature reviewed here.

FAQs

Does saying "um" or "uh" actually make a speaker seem less credible?

A substantial body of experimental research, beginning with source-credibility studies in the 1960s and continuing through recent work using synthetic and digitally manipulated speech, consistently finds that listeners rate speakers who produce more filled pauses and other nonfluencies as less competent, dynamic, or credible than otherwise identical fluent speakers. The size of this effect varies across studies, disfluency types, and delivery contexts, and it is not fixed at any single, universally agreed figure.

Why do listeners judge disfluent speech more harshly?

Researchers generally explain the effect through a "fluency principle": listeners use the ease or difficulty of processing incoming speech as an informal cue to the speaker's underlying qualities. Speech that is harder to process, whether because of frequent hesitations, an unfamiliar accent, or poor audio quality, tends to be misattributed to the speaker being less confident, competent, or trustworthy, even when the actual content of the message is unaffected.

Are filler words always bad for a speaker's credibility?

No. Several studies find that certain fillers and hesitation cues can help listeners process speech, for example by signaling that a short delay or a self-correction is coming, and at least one applied interview study found that removing filler words from a transcript did not change how observers judged a candidate's anxiety or performance. The disfluency-credibility effect is a consistent pattern across many studies rather than a fixed rule that applies identically in every speaking context.

Margaret Whitaker