Back

The Candidates Your Voice Screener Mishears

Voice AI screening helps candidates who look bad on a written form, but every score depends on a transcript, and speech recognition still makes more errors for some accents than others. Here's how that turns into quiet rejections, and a fifty-call test to find out whether it's happening in your pipeline.

AI Implementation5 min read
The Candidates Your Voice Screener Mishears

Voice AI screening is having a good year. Agencies that used to text applicants a link to a form now have software phone every one of them, and I'd like to think we helped push that along. So it's worth being clear about the weak spot, because there is one, and it isn't the part people usually ask me about.

Most of the questions I get are about the model: what it scores, whether it's biased, who makes the final call. Far fewer people ask about the step before any of that, where the audio gets turned into text. Everything the model judges is the transcript. If the transcript is wrong, the model is judging something the candidate never said.

The transcript is the candidate, as far as the model knows

In 2020 a team at Stanford ran the same set of recorded interviews through the speech recognition systems from Amazon, Apple, Google, IBM and Microsoft. The average word error rate was 0.35 for Black speakers and 0.19 for white speakers [1]. Put plainly, about one word in three came out wrong for one group and one in five for the other. These systems have improved since, and I'd expect the gap to have narrowed. I wouldn't assume it's gone. Similar gaps show up for non-native speakers and strong regional accents, and in frontline hiring that describes a big share of your applicants.

Here's what that looks like on a real screen. A candidate for a warehouse job is asked about equipment. She says she's run a reach truck and a counterbalance for four years. The transcript reads "a rich truck and a counter balance for for years." Anyone listening would understand her fine. A model reading the text might score it fine too, or it might decide the answer was vague about experience and mark her down a little. Nobody sees this happen. The score just comes back lower.

Why nobody notices

Errors like this don't look like bias in any report you'd normally run. The candidate wasn't rejected for her accent. She was rejected for an answer that, on paper, didn't clearly mention the equipment. If you audit pass rates by group you might spot a gap, but you'd probably put it down to experience or fit, because that's what the notes say.

It's also the exact problem voice was supposed to fix. A few weeks ago we wrote about the English test hidden in frontline job applications, where people who explain things perfectly well out loud get filtered out by a text form. I still think a conversation is the better first screen for most frontline roles. But if the speech-to-text step works worse for the same people, you haven't removed the filter. You've moved it somewhere harder to see, because now there's no form to point at.

What I'd check before trusting any voice screen

You don't need a research team for this. Pull fifty calls from a recent role, picked to cover the range of accents in your applicant pool, and have someone listen to each one while reading the transcript. Mark every place where a word that mattered was wrong: a number of years, a piece of equipment, a shift, a yes that turned into a no. Ignore the harmless slips. You only care about errors that could move a score.

Then look at where they cluster. If one group's transcripts carry most of the errors that matter, you have your answer, and it's much better to find it yourself than in a complaint. California's rules on automated hiring decisions already expect you to be able to explain a screening outcome years after the fact, and "the transcript said so" won't hold up if the transcript was wrong.

Two habits help whatever you find. Make sure the screen asks again when an answer doesn't make sense, instead of scoring a garbled answer as a weak one. And keep the audio close to the transcript, so a recruiter who sees a strange answer can play the ten seconds that matter before anyone gets rejected.

How Asendia AI handles this

Asendia is a voice-first AI recruiter. It phones every applicant, usually within hours of them applying and at any hour, 24/7, and runs the same structured screen for the role every time. When an answer is vague or doesn't fit the question, it asks again in different words rather than moving on. That catches a lot of what a bad transcription would otherwise turn into a bad score.

Every call is transcribed, and the answers, a short summary and the transcript go into the ATS you already use. Your team writes the questions and the pass criteria, so where a job really needs spoken English you can score it as its own item and keep it from dragging down everything else. Most of the agencies using Asendia came to us to get through hundreds of applicants a week without adding recruiters. Even so, I'd rather you test us on this than take my word for it. If our transcripts are worse for the people you hire, you should know, and so should we.

Final Word

Voice screening fixed a real problem for a lot of candidates who were never going to look good on a form. It can bring a quieter version of that problem back if nobody checks how well the software hears people. The model gets most of the scrutiny, but the transcript is what it actually reads.

This week, take ten screens from your busiest role, choose the candidates whose accents are least like your recruiters', and read each transcript while you listen to the call. If you find a no that was really a yes, you know where to start.

Ready to transform your hiring strategy? Schedule a Demo with our founders today!: https://asendia.ai/talk-to-founders

Badis Zormati

Co-Founder, Asendia AI

Ready to transform your hiring strategy?

Schedule a Demo

Keep reading