RXed AI News

AI to the bone.
@RXed_EU
SCORE 7/1018/20 elements Audited 2026-08-07 · RXed table v1.0

Coval

“The voice AI evaluation platform for agents at scale. Simulate, observe, and review all in one place.” — the vendor’s own words

The most complete voice-agent evaluation stack we have scored: audio-native metrics, OpenTelemetry tracing, a hosted MCP server and an A2A connection type, all from a 10-person company that has raised $31M and grades other people's agents without publishing any validation of its own judges.

SpecialistAutomation & AgentsVoice & SpeechSecurity & ComplianceFreemium
Vendor
Coval (coval.dev), founded 2024 by Brooke Hopkins, previously eval infrastructure lead at Waymo · www.coval.ai
Origin
US — San Francisco
Pricing
Starter $100/mo · Growth $500/mo · Enterprise From $4,500/mo
Users (official only)
Not disclosed as a user count. Coval states more than 60 enterprises use the platform, including Zoom and Deepgram, and reports 10x year-over-year revenue growth and tens of millions of evaluations run. Team size 10 per its Y Combinator profile. (source, 2026-06-24)
Starter$100/mo100 simulation min/mo, 1,000 monitored calls/mo, 30-day trace retention, 5 concurrent simulations, 1 project, 10 metrics per call, 10 personas, unlimited seats. API, CLI, MCP and skills included. SOC 2 Type II, HIPAA, GDPR.
Growth$500/mo1,000 simulation min/mo, 10,000 monitored calls/mo, 90-day retention, 25 concurrent simulations, 5 projects, unlimited metrics, 50 personas, advanced voice models, RBAC, 30-day audit logs.
EnterpriseFrom $4,500/moCustom volumes and retention, fine-tuned metrics, custom or BYO voice models, SAML SSO, SCIM, private/VPC deployment, data residency, BYO storage, SLA up to 99.99%, white-label reports, named support engineer.

Published list pricing, verified on the vendor's live pricing page. EUR figures converted at the ECB euro reference rate of 06/08/2026 (EUR/USD 1.1542) and rounded; Coval bills in USD.

checked 2026-08-07 · vendor pricing page

Element scores

Reactive
Retrieval & Memory
Orchestration
Validation
Models
Primitives
Pr7.5
Prompts
Em5
Embeddings
Cx7
Context
Tr9
Tracing
Lg6
LLM
Compositions
Fc6.5
Function calling
Vx
Vector store
Rg5
RAG
Gr6.5
Guardrails
Mm8.5
Multimodal
Deployment
Ag6.5
Agents
Ft5.5
Fine-tuning
Fw8
Frameworks & harnesses
Ev9
Evaluations
Sm6
Small models
Emerging
Ma5
Multi-agent
Sy7.5
Synthetic data
Pc9
Protocols
In5.5
Interpretability
Th
Thinking models
Tap or hover any element to see why it got that score.

Strengths

Coval treats voice as a distinct engineering problem rather than text with audio attached. Audio LLM judges score the recording itself, and self-hosted models measure audio sentiment, word error rate and timbre drift, which catches a mid-call TTS failover that a transcript cannot show. Simulation ships 27 voices, 10 languages and 20 acoustic environments, plus DTMF and IVR handling. The developer surface is unusually complete for a 10-person company: REST API, CLI, typed TypeScript and Python clients from public OpenAPI specs, GitHub Actions, an npx wizard that auto-instruments Pipecat, LiveKit and Vapi, a hosted OAuth MCP server with 19 tools, and a Chat A2A connection type. Tracing is OpenTelemetry-native and imports from Langfuse, Arize Phoenix and LangSmith in one connect, so teams keep the observability stack they already run. Compliance is not tiered: SOC 2 Type II, HIPAA and GDPR from the $100 plan up.

Honest dings

Coval sells verdicts on other people's agents and publishes no accuracy validation of its own judges. There are no public agreement rates against human reviewers, no judge model catalogue and no documentation of how human-review feedback retrains the judge, which leaves the buyer trusting an unaudited auditor. Entry pricing is steep for the volume: $100/month buys 100 simulation minutes, and overage runs $0.40 per minute on Starter, so a serious pre-launch test sweep lands on the $500 plan quickly and Enterprise starts at $4,500/month. Trace retention is 30 days on Starter, which is short for regression archaeology. The company is 10 people with $31M raised and 60-plus enterprise customers, so support depth and roadmap are concentrated risks. Sofia and Benchmarks are both flagged beta.

Best for: Teams already running a voice agent in production on Vapi, LiveKit, Pipecat, Retell or an in-house stack, who need regression testing in CI and audio-level failure evidence before a release, and who can justify $500 a month or more.
Visit coval.ai Prices and details change — this passport is re-verified at least quarterly.

Sources

Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).

  • coval.ai/pricing — Official pricing page: Starter $100, Growth $500, Enterprise from $4,500, full feature comparison table, overage rates, retention, concurrency, compliance by tier (accessed 2026-08-07)
  • coval.ai/products — Official product page: Simulate/Observe/Review loop, 27 voices, 10 languages, 20 environments, CI/CD, OpenTelemetry-native, SOC 2 Type II, GDPR in progress, vendor-agnostic positioning (accessed 2026-08-07)
  • docs.coval.ai/llms.txt — Official docs index: supported agent connection types incl. Chat A2A (JSON-RPC), OpenAI Realtime, Gemini Live, Pipecat, LiveKit, SMS; full metric, persona, test-set, benchmark and trace surface (accessed 2026-08-07)
  • docs.coval.ai/mcp/overview — Official docs: hosted MCP server at mcp.coval.dev, Streamable HTTP with OAuth, 19 tools across 7 categories, read/write tool separation, consult_sofia read-only (accessed 2026-08-07)
  • docs.coval.ai/concepts/metrics/types/llm-judge — Official docs: seven LLM-judge metric variants incl. audio-native judges, optional judge model selection, trace context injected alongside transcript (accessed 2026-08-07)
  • docs.coval.ai/concepts/metrics/types/ml-model — Official docs: self-hosted ML metrics — audio sentiment, timbre drift via speaker embeddings, transcription error (WER) (accessed 2026-08-07)
  • docs.coval.ai/concepts/simulations/traces — Official docs: OpenTelemetry instrumentation paths, npx @coval/wizard for Pipecat/LiveKit/Vapi, import from Langfuse, Arize Phoenix, LangSmith; Transition Hotspots and Trace Search (accessed 2026-08-07)
  • docs.coval.ai/sofia/overview — Official docs: Sofia in-app agent, beta status, confirmation cards for edits/launches/destructive actions, organization scoping, read-only over the MCP connector (accessed 2026-08-07)
  • docs.coval.ai/concepts/benchmarks/overview — Official docs: STT/TTS provider benchmarks on your own data, WER and Time-to-First-Audio, provider catalogue (Deepgram, OpenAI, AssemblyAI, ElevenLabs, Google, Azure, Amazon), beta flag and hard limits (accessed 2026-08-07)
  • docs.coval.ai/concepts/test-sets/create — Official docs: Test Set Generator drafts cases from a description with typed attributes; knowledge_base_entries types (web_url, plain_text, json, zendesk, shelf, file) (accessed 2026-08-07)
  • prnewswire.com/news-releases/coval-raises-28-millio… — Official announcement, 24/06/2026: $28M Series A led by Norwest, $31M total since 2024, 60+ enterprises, Zoom and Deepgram named, claimed 30x manual QA reduction (accessed 2026-08-07)
  • ycombinator.com/companies/coval — Independent (YC directory): founded 2024, Summer 2024 batch, team size 10, San Francisco, founder Brooke Hopkins ex-Waymo eval infrastructure (accessed 2026-08-07)
  • techcrunch.com/2025/01/23/coval-evaluates-ai-voice-… — Independent reporting, 23/01/2025: $3.3M seed led by MaC Venture Capital, Waymo simulation lineage, early positioning (accessed 2026-08-07)
  • techfundingnews.com/coval-28m-series-a-voice-ai-tes… — Independent reporting, 06/2026: 10x year-over-year revenue growth, ARR and headcount not disclosed, regulated-industry customer base (accessed 2026-08-07)
  • braintrust.dev/articles/best-voice-agent-evaluation… — Independent competitor comparison: positions Coval as simulation and CI/CD regression specialist, notes the voice-only tradeoff versus general eval platforms (accessed 2026-08-07)
  • hamming.ai/resources/hamming-vs-coval — Independent competitor claim (adversarial, treat accordingly): asserts Coval's public materials do not demonstrate voicemail/answering-machine simulation or repeat-caller memory testing (accessed 2026-08-07)
  • ecb.europa.eu/stats/policy_and_exchange_rates/euro_… — ECB euro reference exchange rate 06/08/2026: EUR/USD 1.1542, used for the EUR conversions above (accessed 2026-08-07)