← Блог
·8 хв читання·Olexandr Dmytruk, CEO Lumina AI

AI HR voice agent: $0.54 per interview at scale, with the failure logs that actually matter

We deployed a voice-based AI HR pulse system for a 91-person beauty academy in Ukraine. Operational cost per full 9-question interview: $0.54. Pickup rate: 30-40%. The technical wins (multi-layer endpointing, anti-hallucination guards) and what nobody warns you about (Telegram-bot 'team morale' fatigue).

What the system does

A 91-person beauty academy with 5 locations was running quarterly stay-interviews manually: 22 hours of phone calls per cycle, all on one HR coordinator. By the fourth hour, voice gets gravelly, attention drops, the last 30% of interviews are noticeably weaker.

We built "Lina" — an AI voice agent that calls each team member, runs a structured 9-question interview in Ukrainian, transcribes, scores, surfaces red flags to a Telegram Mini App that the CEO and HR check in 5 minutes instead of reading 22 hours of notes.

Live since 28.04.2026, ~12% of team interviewed in the first 12 days.

Per-interview cost: $0.54

Validated on 10 complete 9-question interviews:

ComponentCost per interview
Vapi.ai voice orchestration (~5 min call)$0.18
ElevenLabs Charlotte UA TTS (multilingual_v2)$0.21
Deepgram nova-3 STT (UA, endpointing 500ms)$0.04
Gemini 2.5 Flash analyzer + summary$0.08
Postgres write + Telegram alert~$0.03
Total$0.54

For a 91-person team running quarterly: 91 × 4 × $0.54 = $197/year in operational vendor costs.

Compare to the alternative: 22 hours × 4 quarters × $15/hour HR coordinator time = $1,320/year of avoided labor. Net annual saving: ~$1,100, plus the qualitative gain of consistent interview quality.

The build cost itself ($3,500-8,000 depending on scope) is not justified by labor savings alone. The real value is systemic team-pulse intelligence that's impossible to maintain manually.

What worked technically

Multi-layer endpointing (5-tier pause discrimination)

Default endpointing in voice agents waits a fixed time (usually 600-800ms) before assuming the user finished speaking. That's wrong: humans pause for different reasons.

We built 5 tiers:

  • 0.3 seconds for short answers ("так" / "ні" / numbers)
  • 0.8s for short factual responses
  • 1.5s for thoughtful sentences
  • 2.5s for emotional or considered answers
  • 3.5s for filler-with-thought ("ну, тобто, мені здається...")

The model picks tier dynamically based on what's been said and how. Net effect: 80% reduction in "agent talks over user" complaints in the first 50 calls vs default.

Anti-hallucination guard on the analyzer

Critical failure mode of LLM analyzers: when the input transcript is too short or noisy, the model invents plausible-sounding scores. We caught this in early testing — a 3-turn conversation getting a "comfort: 7/10, eNPS: 8, ready-to-leave: 0/10" verdict despite the user barely saying anything.

The fix: hard pre-check before the analyzer runs.

  • Transcript < 500 characters → analyzer returns null, not a hallucination.
  • Fewer than 3 user turns → same.
  • These go to a "incomplete" bucket for human review instead of contaminating the dataset.

Pre-call SMS warmup

Cold call pickup rate without warmup: ~15%. With 25-second pre-call SMS ("Hi NAME — Lina from the academy will call in a moment for our quarterly check-in"): pickup rate ~30-40%. Doubled.

What didn't work (and matters more)

Pickup rate is 30-40%, not 80%

Voice marketing materials always show 80%+ engagement. Reality: half of any team list won't pick up on first attempt, regardless of warmup. This is infrastructure-side (BYO Zadarma SIP routing) and cultural (Ukrainians are conditioned to ignore unknown numbers due to spam).

The honest answer: budget for 2-3 retry attempts per interviewee. For a 91-person quarterly cycle, that's 200-280 actual call attempts to get 91 completed interviews.

TTS short-brand pronunciation

Charlotte UA read the academy's Latin-letter brand name as one phonetic word. We had to spell it phonetically in the prompt.

This is a generalizable lesson: any 3-5 character brand or acronym needs phonetic spelling for Ukrainian TTS. We hardcode these in the system prompt now.

Vocative case > instrumental

Initial prompt: "Привіт, мене звати Ліна" → triggered native-speaker filter as "robotic." Fixed with proper Ukrainian vocative: "Доброго дня, я Ліна, дзвоню від Академії..." — felt human. This is dialect-level tuning that no English-language voice agent vendor handles.

What this means for buyers

Don't believe vendor pickup-rate claims. Assume 30-40% on cold lists, 60%+ on warm relationships. Plan accordingly.

Localization is not "translate the prompt". It's vocative case, brand pronunciation, dialect-appropriate filler patterns, regional cultural norms. Most Western voice-AI vendors ship "ukrainian language pack" but get the dialect wrong.

Per-interview cost is dominated by TTS, not LLM. Pick voice provider carefully. Charlotte UA at $0.21/call vs cheaper alternatives at $0.05 — the quality gap is worth 4x the cost for HR-grade interactions.

Anti-hallucination is non-negotiable. Demand to see the analyzer's failure-mode handling. If your vendor can't explain how they prevent silent score fabrication, walk away.

Where this is going

Phase 2 (planned): voice agent for outbound warm-lead reactivation in retail (a beauty distributor with 700+ partners who opened the storefront but didn't order). Same architecture, different prompt. Same $0.54/interaction cost ceiling.

The interesting unlock here isn't the voice tech — it's the operational discipline. Multi-layer endpointing, anti-hallucination guards, vocative-case prompts — these are not LLM-side features. They're things you only learn from production debugging across hundreds of real conversations.

If you're evaluating voice AI vendors and the demos sound suspiciously perfect: ask about the failure logs. Ask about pickup rates from actual deployments. Ask what they had to hardcode after launch.

The real moat is in the things they had to fix that you'd never think to ask about.

#AI voice agent#HR automation#voice AI pricing#Vapi.ai cost#ElevenLabs voice agent#AI HR pulse#stay interview automation

Більше про AI-автоматизацію — в нашому каналі

Практичні кейси, гайди по n8n і AI-агентах, добірки інструментів — щотижня в @ai_prosto_ua.

Підписатись на канал →