Screening a job application has always rested on an assumption that was never stated because it never needed stating: that producing a competent written application required some of the competence the job required. A well-argued cover letter was evidence of a person who could argue well in writing.
Generative writing tools severed that link. The cost of producing a polished, role-specific, articulate application fell to roughly zero, and did so for every applicant simultaneously.
What happened to volume
Reported figures describe a sharp increase in application volume over the period in which these tools became widely available. Applications per open role rose from approximately 116 in 2022 to 244 in 2025.[1]
The mechanism is not mysterious. When the marginal cost of one more application approaches zero, the rational number of applications to send rises accordingly, and the applicant pool for any given role expands without the number of qualified applicants changing.
The resulting problem for employers is not that applications are worse. Many are better written than what preceded them. The problem is that written quality no longer discriminates, because it no longer costs the applicant anything to produce.
Where the signal moved
When one channel stops carrying information, screening moves to whichever channel still does. Trade reporting through 2026 describes employers adding verification steps, including interview software that flags tab-switching mid-answer as a possible sign of a candidate reading generated text from a second screen.[2]
The underlying logic is worth stating plainly, because it explains why live speech specifically became the fallback:
- Written submissions are asynchronous and editable. Any asynchronous, editable channel can be fully delegated to a tool, which is precisely what makes it cheap.
- Live unscripted speech is neither. Real-time conversation, with follow-up questions and interruption, is currently expensive to fake convincingly.
- Therefore verification concentrates on the un-editable channel. Not because speech is a better predictor of job performance than writing, but because it is harder to outsource.
That third point deserves emphasis, since the shift is often described as a rediscovery of the value of communication skills. It is better described as a search for a channel that still costs the candidate something.
What this asks of candidates
The consequence is a redistribution of who is advantaged. A hiring process weighted toward written submissions rewards a set of skills that tools can now supply. A process weighted toward live unscripted answers rewards real-time verbal organisation, which they cannot.
Survey evidence suggests the population being asked to do this is not confident about it. In a May 2026 survey of 1,002 employed US respondents, 51% said AI use had made spontaneous conversation feel more difficult, rising to 66% among Gen Z workers, and 44% said they sometimes freeze in person because they cannot review or edit their words first.[3]
Those figures are self-reported perceptions rather than measured ability. But they describe a workforce being asked to compete on the exact dimension it reports feeling worst about, at the moment that dimension became the deciding one.
What is being measured, and whether it should be
If unscripted speech is now carrying screening weight, the question of what interviewers should read from it becomes consequential rather than academic.
Two findings from the disfluency literature bear directly on this. Filled pauses such as uh and um are not noise in the signal: they carry information about upcoming delay and speaker planning, and listeners use them.[4] And disfluency rates vary systematically with speaker age, relationship to the listener, topic, and conversational role.[5] A rate measured in an interview is a property of that situation, not a stable trait of the candidate.
Descriptive corpus data shows how wide the normal range is. Across 16,928 answers from 1,000 mock interview sessions, the median answer ran 17 words while the 90th percentile ran 147, and measured filler rates ranged from zero in about a fifth of sessions to above 7.5 per 100 words in the top decile.[6] Any evaluator treating a single observed rate as a signal of candidate quality is reading a number with substantial situational variance.
The unresolved problem
Verbal verification solves the employer's immediate problem and creates two others.
It advantages fluency independent of competence. The candidate who organises thoughts quickly out loud is not necessarily the one who does the job better. Written screening had a version of this bias too, but it was at least partly trainable by the applicant, and it did not disadvantage non-native speakers in the same way live conversation does.
It is a temporary equilibrium. Real-time voice generation is improving. The property that makes live speech a usable verification channel, namely that it is expensive to fake, is a fact about the current state of the technology rather than a durable feature of speech.
The stable version of this problem is not "how do we detect AI assistance." It is what employers were trying to learn from a cover letter in the first place, and which channel actually carries that information now.