Looking for a job?

Start here

AI voicenote rating: Building a responsible and scalable solution

AI voicenote rating: Building a responsible and scalable solution


Published

Feb 2, 2026

4

min

Across 12,991 interviews conducted through JOBJACK in 2025, 18.5% were unsuccessful because the candidate’s communication abilities did not meet the role's requirements. That's ahead of every other reason on the list, including under-qualification, insufficient experience and culture misfit.


Poor communication is the number one reason store-level candidates are rejected at the interview stage.

Because this problem only showed up at the interview, hiring managers had no way to see how someone communicated until they were already sitting across the table. By then, hours had already gone into shortlisting, scheduling and interviewing a candidate who would ultimately not meet the requirements. Across a high-volume, decentralised hiring operation, that adds up. We estimate over 1,000 work hours were lost this way in 2025 alone.

We needed a solution that enabled communication capture before the interview and offered a robust rating option that supported recruitment at scale.

To solve this challenge for our customers, we asked the question:

Can AI assess verbal competence in a way that's scalable and objective without losing the nuance a human hiring manager would pick up on?


Building a solution that holds up at scale

The first step was building voice capture functionality on the application.

For jobs that had a specific communication ability requirement, candidates were prompted with a specific scenario and recorded their answer with a concise voice note. With this baseline functionality, hiring managers were able to individually listen to candidates’ responses. Being able to listen to a candidate’s voice ability before the interview was crucial, but it increased hiring managers’ admin at scale.

Adding an AI rating system was the next step to help the functionality scale, but also to make it unbiased and standardised. But before we let AI anywhere near scoring, we had to know the scale itself was sound.

We built a five-dimension model, scoring every voice note 0-10 across two families: knowledge (comprehension, vocabulary, grammar) and delivery (fluency, pronunciation).

We tested this on 1,076 applicants and the results held up well. Reliability was strong, and the scale behaved consistently across gender, age, and the ethnic groups we had enough data to test properly. Where our sample sizes were too small to draw conclusions, we stated that fact rather than stretching the data. A fair instrument isn't just one that shows no bias; it's honest about what it hasn't been tested on yet.

With the scale validated, the real question became whether AI could apply it as well as a trained human rater.

Testing the AI scoring system

We brought in three independent experts, two Masters graduates in English and a PhD candidate. We had them rate 100 voice notes using the same five-dimension scale.

The honest answer: it's a mixed picture, but a useful one.

On total score, agreement between the AI and the human panel was moderate. But that headline number hid something more interesting once we broke it down by dimension:

  • The AI was consistently harsher than the experts on the knowledge side of communication (vocabulary and grammar), scoring roughly 1 to 1.6 points lower.

  • On the delivery side (fluency and pronunciation), it was more lenient. The AI's own written comments on individual scores backed this up. Where it praised fluency, its delivery scores ran about 2 points higher than its knowledge scores on the same response. The pattern showed up in both the numbers and the AI's own reasoning.

In plain terms: the AI currently rewards how someone sounds more than what they actually say. This finding is exactly the kind of thing you only catch by checking an AI rater against real experts rather than trusting the score on its own.


Our consensus on AI: A supplement, not a replacement

Ultimately, we built the tool to support a hiring manager's judgement, not override it.

By starting with a native build to address the problem and then using AI to scale while upholding transparency, we’ve kept our responsibility to equip the user with all the facts.

Every AI-scored voice note shows the reasoning behind the rating, and the hiring manager can rate it themselves if they disagree, with that feedback feeding back into how we improve the model. It's worth remembering that the three human experts we tested against didn't agree with each other perfectly either, with scores varying by up to 2 points on the same recordings.

Humans aren't a perfect baseline either. But a transparent, correctable AI can get closer to fair over time, provided it stays open to being checked.

We're now running this live with select customers, where a voice evaluation isn't a required step in the application itself, and using that real-world data to keep improving the model.


The early signal is encouraging. Across more than 65,000 voice notes assessed so far, candidates who score well are about 50% more likely to reach the interview stage and 80% more likely to be appointed than everyone else. That's a correlation worth taking seriously, not proof the tool is perfect, but a strong indication that what it's measuring actually matters to hiring outcomes.

It's a working tool that's already saving hiring managers time by giving them insight into a candidate's communication ability before they ever get to the interview. But because we know it currently leans toward rewarding delivery over substance, we're actively working to correct that as we collect more data.

Building this kind of tool responsibly means publishing the uncomfortable findings alongside the encouraging ones. That's the balance we're committed to holding as we keep refining it.