Coval
The most complete voice-agent evaluation stack we have scored: audio-native metrics, OpenTelemetry tracing, a hosted MCP server and an A2A connection type, all from a 10-person company that has raised $31M and grades other people's agents without publishing any validation of its own judges.
PRICING
| Starter | $100/mo | 100 simulation min/mo, 1,000 monitored calls/mo, 30-day trace retention, 5 concurrent simulations, 1 project, 10 metrics per call, 10 personas, unlimited seats. API, CLI, MCP and skills included. SOC 2 Type II, HIPAA, GDPR. |
| Growth | $500/mo | 1,000 simulation min/mo, 10,000 monitored calls/mo, 90-day retention, 25 concurrent simulations, 5 projects, unlimited metrics, 50 personas, advanced voice models, RBAC, 30-day audit logs. |
| Enterprise | From $4,500/mo | Custom volumes and retention, fine-tuned metrics, custom or BYO voice models, SAML SSO, SCIM, private/VPC deployment, data residency, BYO storage, SLA up to 99.99%, white-label reports, named support engineer. |
Published list pricing, verified on the vendor's live pricing page. EUR figures converted at the ECB euro reference rate of 06/08/2026 (EUR/USD 1.1542) and rounded; Coval bills in USD.
checked 2026-08-07 · vendor pricing page
Element scores
Strengths
Coval treats voice as a distinct engineering problem rather than text with audio attached. Audio LLM judges score the recording itself, and self-hosted models measure audio sentiment, word error rate and timbre drift, which catches a mid-call TTS failover that a transcript cannot show. Simulation ships 27 voices, 10 languages and 20 acoustic environments, plus DTMF and IVR handling. The developer surface is unusually complete for a 10-person company: REST API, CLI, typed TypeScript and Python clients from public OpenAPI specs, GitHub Actions, an npx wizard that auto-instruments Pipecat, LiveKit and Vapi, a hosted OAuth MCP server with 19 tools, and a Chat A2A connection type. Tracing is OpenTelemetry-native and imports from Langfuse, Arize Phoenix and LangSmith in one connect, so teams keep the observability stack they already run. Compliance is not tiered: SOC 2 Type II, HIPAA and GDPR from the $100 plan up.
Honest dings
Coval sells verdicts on other people's agents and publishes no accuracy validation of its own judges. There are no public agreement rates against human reviewers, no judge model catalogue and no documentation of how human-review feedback retrains the judge, which leaves the buyer trusting an unaudited auditor. Entry pricing is steep for the volume: $100/month buys 100 simulation minutes, and overage runs $0.40 per minute on Starter, so a serious pre-launch test sweep lands on the $500 plan quickly and Enterprise starts at $4,500/month. Trace retention is 30 days on Starter, which is short for regression archaeology. The company is 10 people with $31M raised and 60-plus enterprise customers, so support depth and roadmap are concentrated risks. Sofia and Benchmarks are both flagged beta.
Sources
Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).
- coval.ai/pricing — Official pricing page: Starter $100, Growth $500, Enterprise from $4,500, full feature comparison table, overage rates, retention, concurrency, compliance by tier (accessed 2026-08-07)
- coval.ai/products — Official product page: Simulate/Observe/Review loop, 27 voices, 10 languages, 20 environments, CI/CD, OpenTelemetry-native, SOC 2 Type II, GDPR in progress, vendor-agnostic positioning (accessed 2026-08-07)
- docs.coval.ai/llms.txt — Official docs index: supported agent connection types incl. Chat A2A (JSON-RPC), OpenAI Realtime, Gemini Live, Pipecat, LiveKit, SMS; full metric, persona, test-set, benchmark and trace surface (accessed 2026-08-07)
- docs.coval.ai/mcp/overview — Official docs: hosted MCP server at mcp.coval.dev, Streamable HTTP with OAuth, 19 tools across 7 categories, read/write tool separation, consult_sofia read-only (accessed 2026-08-07)
- docs.coval.ai/concepts/metrics/types/llm-judge — Official docs: seven LLM-judge metric variants incl. audio-native judges, optional judge model selection, trace context injected alongside transcript (accessed 2026-08-07)
- docs.coval.ai/concepts/metrics/types/ml-model — Official docs: self-hosted ML metrics — audio sentiment, timbre drift via speaker embeddings, transcription error (WER) (accessed 2026-08-07)
- docs.coval.ai/concepts/simulations/traces — Official docs: OpenTelemetry instrumentation paths, npx @coval/wizard for Pipecat/LiveKit/Vapi, import from Langfuse, Arize Phoenix, LangSmith; Transition Hotspots and Trace Search (accessed 2026-08-07)
- docs.coval.ai/sofia/overview — Official docs: Sofia in-app agent, beta status, confirmation cards for edits/launches/destructive actions, organization scoping, read-only over the MCP connector (accessed 2026-08-07)
- docs.coval.ai/concepts/benchmarks/overview — Official docs: STT/TTS provider benchmarks on your own data, WER and Time-to-First-Audio, provider catalogue (Deepgram, OpenAI, AssemblyAI, ElevenLabs, Google, Azure, Amazon), beta flag and hard limits (accessed 2026-08-07)
- docs.coval.ai/concepts/test-sets/create — Official docs: Test Set Generator drafts cases from a description with typed attributes; knowledge_base_entries types (web_url, plain_text, json, zendesk, shelf, file) (accessed 2026-08-07)
- prnewswire.com/news-releases/coval-raises-28-millio… — Official announcement, 24/06/2026: $28M Series A led by Norwest, $31M total since 2024, 60+ enterprises, Zoom and Deepgram named, claimed 30x manual QA reduction (accessed 2026-08-07)
- ycombinator.com/companies/coval — Independent (YC directory): founded 2024, Summer 2024 batch, team size 10, San Francisco, founder Brooke Hopkins ex-Waymo eval infrastructure (accessed 2026-08-07)
- techcrunch.com/2025/01/23/coval-evaluates-ai-voice-… — Independent reporting, 23/01/2025: $3.3M seed led by MaC Venture Capital, Waymo simulation lineage, early positioning (accessed 2026-08-07)
- techfundingnews.com/coval-28m-series-a-voice-ai-tes… — Independent reporting, 06/2026: 10x year-over-year revenue growth, ARR and headcount not disclosed, regulated-industry customer base (accessed 2026-08-07)
- braintrust.dev/articles/best-voice-agent-evaluation… — Independent competitor comparison: positions Coval as simulation and CI/CD regression specialist, notes the voice-only tradeoff versus general eval platforms (accessed 2026-08-07)
- hamming.ai/resources/hamming-vs-coval — Independent competitor claim (adversarial, treat accordingly): asserts Coval's public materials do not demonstrate voicemail/answering-machine simulation or repeat-caller memory testing (accessed 2026-08-07)
- ecb.europa.eu/stats/policy_and_exchange_rates/euro_… — ECB euro reference exchange rate 06/08/2026: EUR/USD 1.1542, used for the EUR conversions above (accessed 2026-08-07)