RXed AI News

AI to the bone.
@RXed_EU
Audited 2026-08-11 · RXed table v1.0

Hamming AI

Visit hamming.ai
“The trust platform for voice and chat agents. Test before launch. Monitor in production. Red-team risky behavior before it reaches customers.” — the vendor’s own words

The deepest voice-agent testing loop we have scored on tool-call contracts, DTMF and IVR behaviour, and turning a production failure back into a regression test, sold by a company that will not publish a single price and grades other people's agents on a judge model it refuses to name.

Best for: Teams running a phone-based voice agent in production on Vapi, Retell, LiveKit, Pipecat, ElevenLabs or Bland, where the agent navigates an IVR, calls booking or CRM tools with real side effects, and sits under HIPAA or PCI. Not for a team that needs to price the tool before talking to sales, or one that wants to audit the judge before trusting its verdicts.
Scope18/20
Quality6/10
Where the quality sits
6Reactive
5Retrieval & Memory
6Orchestration
8Validation
6Models
SpecialistAutomation & AgentsVoice & SpeechSecurity & CompliancePaid
Vendor
Hamming AI, founded 2024 by Sumanyu Sharma (CEO) and Marius Buleandra (CTO), both ex-Tesla; Y Combinator Summer 2024 · hamming.ai
Origin
US — San Francisco
Pricing
Startup Contact us · Agency Contact us · Enterprise Contact us
Users (official only)
Not disclosed as a user count. Hamming states 10K+ agents monitored, 4M+ calls tested and 10M+ minutes protected, and names Bland Labs, Booked AI, Grove AI, Synthpop, Maven AGI, Basata, Anima, Serve, Opalite Health and mdhub in customer testimonials and case studies. $3.8M seed led by Mischief announced 18/12/2024; aggregators put total funding at $4.3M-$4.5M and headcount at 17-22 in mid-2026. (source, 2026-08-11)
StartupContact usAutomated voice agent testing, automated call analytics, trust and safety reports, custom scoring templates, 7-days-a-week support, direct support from the founders. No published price, no self-serve signup.
AgencyContact usStartup capabilities plus multi-client management and priority support, aimed at agencies running several voice-agent projects.
EnterpriseContact usStartup plan plus SOC 2 and HIPAA compliance, support SLAs and a dedicated support engineer. Vendor pages cite a 99.9% uptime SLA and response times as fast as 10 minutes for critical production issues.

No EUR figure can be given because Hamming publishes no numeric price at any tier. Every plan is demo-gated. Independent reviews and competitor comparisons corroborate the absence of public pricing, and the vendor's FAQ confirms figures are 'shared during your demo call'.

checked 2026-08-11 · vendor pricing page

Element scores

Reactive
Retrieval & Memory
Orchestration
Validation
Models
Primitives
Pr6
Prompts
Em4
Embeddings
Cx7
Context
Tr9
Tracing
Lg5
LLM
Compositions
Fc8
Function calling
Vx
Vector store
Rg4
RAG
Gr8
Guardrails
Mm9
Multimodal
Deployment
Ag6
Agents
Ft3
Fine-tuning
Fw8
Frameworks & harnesses
Ev9
Evaluations
Sm3
Small models
Emerging
Ma5
Multi-agent
Sy8
Synthetic data
Pc6
Protocols
In6
Interpretability
Th
Thinking models
Tap or hover any element to see why it got that score.

Strengths

Hamming is strongest at the parts of voice QA that only show up once an agent is doing real work. The tool-call contract testing is the deepest we have scored: name, argument schema, call order, idempotency key, timeout and result all get asserted, and sandbox mode runs booking and CRM side effects against fixtures so a test suite cannot write to production. DTMF, IVR trees and voicemail are first-class rather than bolted on, which matters because most enterprise voice agents sit behind a phone tree. Audio is judged directly instead of through the transcript, across accents, background noise, barge-in and long silences, so a mid-call TTS failure is visible. Tracing is OpenTelemetry-native with ASR, LLM, tool and TTS spans correlated to recordings, and retention runs from 7 days to unlimited rather than the fixed tiers rivals impose. The loop closes: any production call converts to a replayable regression test in one click, and CI gating through GitHub Actions or Jenkins blocks the prompt change that caused it. Setup is genuinely fast, with one-click imports for LiveKit, Pipecat, ElevenLabs, Retell, Vapi, Bland and Synthflow, and a first test report claimed inside 10 minutes. Compliance is real: SOC 2 Type II since December 2025, HIPAA with a BAA, PII and PHI redaction at ingestion, EU and UK data residency.

Honest dings

You cannot find out what it costs. Every tier says 'Contact us', there is no self-serve signup, no free tier and no published unit rate, so budget planning starts with a sales call and independent reviewers name this as the most common complaint. The judge is a black box: no model named, no version, no catalogue, no per-metric selection. Hamming does publish 95-96% agreement with human evaluators, which is more than most rivals disclose, but there is no sample size, no dataset, no rater count and no chance-corrected statistic behind it, so the number cannot be checked. Competitive claims are self-refereed, including a 'won about 90% of the time' head-to-head record that only Hamming can see. Compliance is Enterprise-tier only on the pricing page, which pushes any regulated buyer straight into the top plan. There is no public developer documentation site, so the REST API and the MCP server are evaluated by reading marketing pages, and the MCP has no published endpoint, tool list or auth model despite a customer describing it in production. The site contradicts itself on language coverage, claiming 65+ in two places and 49 in two others. The company is roughly 17-22 people on about $4.3M raised in a crowded, well-funded field, and RAG scoring appears to be 2024 content still hosted after the pivot rather than a maintained feature.

Prices and details change — this passport is re-verified at least quarterly.
Sources (24) — every claim traceable

Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).

  • hamming.ai/pricing — Official pricing page: Agency, Startup and Enterprise tiers all 'Contact us', per-tier inclusions, SOC 2 and HIPAA placed on Enterprise only (accessed 2026-08-11)
  • hamming.ai — Official homepage: self-description, 10K+ agents monitored, 50K+ concurrent test calls, 50+ metrics, red-team suite, DTMF and IVR emulation, one-click prod-to-test, 95%+ simulation fidelity claim, named customer testimonials (accessed 2026-08-11)
  • hamming.ai/faqs — Official FAQ: volume-based rather than per-seat pricing with unlimited users, pricing disclosed only on a demo call, 65+ languages, ASR/TTS/VAD/barge-in terminology reference (accessed 2026-08-11)
  • hamming.ai/case-studies — Official case studies: Grove AI, Bland Labs, Basata, Synthpop and Maven AGI; the 95-96% agreement-with-human-evaluators figure; Synthpop's use of the MCP server to stand up evals agentically (accessed 2026-08-11)
  • hamming.ai/enterprise — Official enterprise page: SOC 2 Type II, HIPAA BAA, SSO and RBAC, customer-managed keys, retention from 7 days to unlimited, audit trails exported to SIEM, EU and UK data residency (accessed 2026-08-11)
  • hamming.ai/resources/voice-agent-workflow-testing-r… — Official runbook, 24/05/2026: tool-call assertion ledger (name, args, order, idempotency, timeout, result) mapped to OpenAI, LiveKit, Vapi and Retell webhook surfaces; methodology stated as analysis of 4M+ production calls across 10K+ agents (accessed 2026-08-11)
  • hamming.ai/resources/voice-agent-sandbox-testing-to… — Official docs, 31/05/2026: sandbox, mock and live testing modes for booking and CRM side effects without production writes (accessed 2026-08-11)
  • hamming.ai/resources/opentelemetry-voice-agents-tra… — Official guide, 25/02/2026: OTLP ingestion, W3C trace context, span model across call, turn, STT, LLM, tool, TTS and webhook (accessed 2026-08-11)
  • hamming.ai/resources/pii-redaction-voice-agents-com… — Official guide, 03/02/2026: redaction at ingestion across transcripts, audio, logs and traces; 12-15 entity types typical in finance and healthcare; synthetic PII generated to validate the pipeline (accessed 2026-08-11)
  • hamming.ai/blog/soc-compliance-for-voice-ai — Official announcement, 08/12/2025: SOC 2 Type II certification achieved, audit report available on request (accessed 2026-08-11)
  • hamming.ai/resources/why-hamming-ai-is-the-best-voi… — Official page, 03/09/2025: the only public statement on judge model quality, contrasting 'cheaper evaluation models' without naming Hamming's own (accessed 2026-08-11)
  • hamming.ai/resources/why-engineering-teams-choose-h… — Official page, 23/12/2025: REST endpoint list (agent import, scenario generation, test runs, results, call ingestion), Python and Node SDKs, CI/CD gating on pass rates (accessed 2026-08-11)
  • hamming.ai/resources/hamming-vs-coval — Vendor-authored comparison, 18/05/2026 (self-serving, used only for checkable product facts): MCP server, voicemail and repeat-caller memory testing, authenticated recording ingestion from Twilio, GCS and S3, and the self-refereed 'won about 90% of the time' bake-off claim (accessed 2026-08-11)
  • hamming.ai/integrations — Official integrations page: one-click imports for Vapi, Retell, LiveKit, Pipecat, ElevenLabs and Synthflow, each with auto-generated scenarios and 50+ metrics (accessed 2026-08-11)
  • hamming.ai/resources/multilingual-voice-agent-testi… — Official page, 15/11/2025: '49 languages' with per-language word error rate benchmarks and code-switching detail, contradicting the 65+ figure on the homepage and FAQ (accessed 2026-08-11)
  • hamming.ai/resources/rag-debugging — Official post, 16/04/2024: retrieval precision, recall and hallucination attribution scoring, published before the pivot to voice QA and not referenced by any current product page (accessed 2026-08-11)
  • businesswire.com/news/home/20241218104943/en/Hammin… — Official announcement via wire, 18/12/2024: $3.8M seed led by Mischief, founder backgrounds at Tesla and Citizen (accessed 2026-08-11)
  • ycombinator.com/companies/hamming-ai — Independent directory (YC-hosted): Summer 2024 batch, San Francisco, founder profiles, company-authored one-liner and aggregate usage stats (accessed 2026-08-11)
  • news.ycombinator.com/item?id=41257369 — Independent venue (Hacker News), 15/08/2024: Launch HN thread, founders' own description of the LLM-judge design returning classification plus written reasoning, and the earlier 'usage plus seats' pricing model (accessed 2026-08-11)
  • pypi.org/project/livekit-plugins-hamming — Independent package registry: livekit-plugins-hamming v1.6.7, Apache-2.0, exports LiveKit AgentSession call-review payloads to Hamming (accessed 2026-08-11)
  • pypi.org/project/hamming-sdk — Independent package registry: Python SDK published as hamming-sdk, confirming an SDK exists outside marketing copy (accessed 2026-08-11)
  • coval.ai/blog/hamming-vs-cekura — Competitor-published comparison, 20/03/2026 (adversarial, treat accordingly): independently corroborates that Hamming publishes no pricing and offers no self-serve tier (accessed 2026-08-11)
  • cekura.ai/discover/programmatic-voice-agent-testing… — Competitor-published comparison, 25/07/2026 (adversarial): describes Hamming as the most API-complete rival platform with native GitHub Actions and Jenkins hooks, and notes its concurrency limits are not published and its entry point is sales-gated (accessed 2026-08-11)
  • hackernoon.com/best-voice-agent-evaluation-and-test… — Independent roundup, 10/08/2026: positions Hamming among Coval, Cekura, Maxim and Roark, with automated call simulation at volume as its core strength (accessed 2026-08-11)