Cekura
The lowest-friction way into voice-agent QA — $0.25 a test minute, no sales call, an MCP server and Claude Skills on every plan — but it grades transcripts rather than audio, its testing is state-blind, and most of the public evidence for how good its evaluations are was written by Cekura.
PRICING
| Pay as you go | $0 to start | $0.25/voice testing minute, $0.025/chat reply, $0.05/monitored call; 10 concurrent calls; 1 seat free, $30/mo per extra seat; 30-day retention |
| Startup | $500/mo | 10,000 credits ≈ 2,000 test minutes; 50 concurrent calls; 10 seats; signed BAA & DPA; 90-day retention |
| Enterprise | Custom | Annual; red teaming, load testing, vendor bake-offs, VPC/on-prem, SSO/SCIM, audit logs, forward-deployed engineers |
The pricing model changed during 2026. Coverage from February to June 2026 describes a credit-based Developer plan at $30/user/month (750 credits, roughly 5 credits per test minute); the live page on 16/08/2026 is per-minute pay-as-you-go with a $500 Startup tier. HIPAA and GDPR are now listed as included on every plan, where a March 2026 comparison described them as Enterprise-only. Re-check before budgeting.
checked 2026-08-16 · vendor pricing page
Element scores
Strengths
The cheapest credible entry point in voice-agent QA: $0.25 a test minute, 300 free credits, no card and no sales call, with an MCP server and Claude Skills on every plan and unlimited Python metrics free. Coverage is wide for a company founded in 2024 — simulation, red teaming across jailbreak, bias, toxicity and PII leakage, production call monitoring, cross-vendor bake-offs, OpenTelemetry spans across LLM, TTS, STT and tool calls, and Insights clustering failing evals into a ranked list of root causes rather than a pile of logs. Integrations cover Retell, Vapi, ElevenLabs, LiveKit, Pipecat, SIP, Twilio and Plivo.
Honest dings
Evaluation is transcript-primary, so the audio failures that make a voice agent feel wrong need explicit configuration to catch. Testing is state-blind — it cannot set a pre-condition or verify a post-call state change — and there is no CLI. The one cross-platform human benchmark (arXiv 2511.04133, Evalion-co-authored) puts the non-leading platforms at 0.73 F1 against the leader's 0.92. A May 2026 independent roundup notes that a large share of the public evidence base, including the compliance claims, is Cekura-authored. Pricing changed shape twice in six months.
Sources (10) — every claim traceable
Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).
- cekura.ai/pricing — Live pricing 16/08/2026: $0.25/voice minute, $0.025/chat reply, $0.05/monitored call, Startup $500/mo, seat and retention limits; MCP server, Claude Skills, HIPAA and GDPR listed on every plan (accessed 2026-08-16)
- docs.cekura.ai/llms.txt — Official docs index: LLM Judge and Python metrics, rubrics, Conditional Actions, Mock Tools, dynamic variables, test profiles, load and infrastructure testing, integrations (Retell, Vapi, ElevenLabs, LiveKit, Pipecat, SIP, Twilio, Plivo) (accessed 2026-08-16)
- ycombinator.com/companies/cekura-ai — YC Fall 2024, three IIT Bombay founders, 75+ customers claim, and the Red Teaming launch post detailing jailbreak, bias, toxicity and PII-leakage categories plus Red Teaming as a Service (accessed 2026-08-16)
- economictimes.indiatimes.com/tech/startups/yc-backe… — Independent reporting of the $2.4M seed and cofounder Sidhant Kabra's on-record 75+ customer figure and Bangalore office plan (03/07/2025) (accessed 2026-08-16)
- businessinsider.com/read-pitch-deck-y-combinator-st… — Independent reporting of the seed round, the March 2025 rebrand from Vocera, founder backgrounds and team size at the time (seven employees) (accessed 2026-08-16)
- speechmatics.com/company/articles-and-news/de-risk-… — Independent 11-platform roundup (26/05/2026): flags Cekura's SOC 2 / HIPAA / GDPR claims as vendor-authored, notes a large share of the public evidence base is Cekura-written, and records the then-current credit pricing and integration list (accessed 2026-08-16)
- arxiv.org/abs/2511.04133 — Testing the Testers (Andrés et al., 06/11/2025) — the only published cross-platform human benchmark of voice AI testing platforms, 21,600 human judgments; top platform 0.92 F1 evaluation quality against 0.73 for the others. Co-authored by Evalion, a competitor, so read the ranking with that in mind (accessed 2026-08-16)
- coval.ai/blog/coval-vs-cekura — Competitor-authored comparison (Coval, 17/03/2026). Used only for the points it concedes against its own interest — Cekura was first in the category to ship an MCP server — and for the dated state-blind and no-CLI observations (accessed 2026-08-16)
- linkedin.com/company/cekuraai — Official June 2026 product update: OpenTelemetry spans for LLM, TTS, STT and tool calls, Insights root-cause clustering, Cekura Academy, Custom Skills for the Cekura Agent (accessed 2026-08-16)
- sigmamind.ai/blog/ai-voice-agent-testing-platform-b… — Independent buyer's guide: 30+ language simulation, frequency-based load testing and the 10-concurrent-call developer cap, and the observation that Cekura is transcript-centric by default (accessed 2026-08-16)