Bluejay
Ten people and a $4M seed, and the feature list already reads like a mature QA platform: MCP, OpenTelemetry ingestion, OWASP-mapped red teaming, CI gates and a free tier. Buy it for the testing, not for the company's balance sheet.
PRICING
| Pay-as-you-go | $0/mo + usage | $25 free credits, up to 25 concurrent simulations, 14-day retention, fastest credit burn |
| Growth | $500/mo | 100 concurrent, ~1,500-1,600 simulation minutes, 13,000 monitoring minutes, signed BAA and DPA, standard RBAC, 30-day retention |
| Scale | $1,000/mo | 200 concurrent plus load testing, ~4,000-4,300 simulation minutes, 34,000 monitoring minutes, 90-day retention, lowest cost per minute |
| Enterprise | Custom | Redline BAA and DPA, SSO/SAML with SCIM, custom RBAC, uptime SLA up to 99.9%, dedicated engineer |
Overage credits burn at 1.8x on Growth and 1.5x on Scale, so the headline price is a floor, not a cap. Pay-as-you-go has no overage because it is prepaid. EUR figures are converted at roughly 0.87 and are indicative only.
checked 2026-08-24 · vendor pricing page
Element scores
Strengths
The evaluation stack is the deepest part and it is genuinely deep: 70+ metrics plus unlimited custom ones, Metrics Lab to align the judge to your own reviewers, OWASP-mapped red teaming with a content-safety taxonomy, load testing to 200 concurrent, and production replays that re-run real failures until the agent passes. The developer surface is unusually complete for a company this young, with Bluejay as Code, SDKs, a CLI, webhooks, GitHub Actions gates, a Claude Code skill and an MCP server on every plan including the free one. OpenTelemetry ingestion means it slots next to Langfuse or OpenInference instead of replacing them.
Honest dings
Bluejay grades other people's agents and publishes nothing about the quality of its own judges, and it names no model anywhere, so a silent judge swap moves your scores with no changelog to point at. The company is 10 people on a $4M seed round, founded in 2025, which is real vendor risk for something you wire into a CI gate. Retention is short at the bottom (14 days on pay-as-you-go), overage credits burn at 1.5-1.8x above the included minutes, and no A2A support is documented.
Sources (9) — every claim traceable
Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).
- getbluejay.ai/pricing — Official pricing and full feature matrix: four tiers, concurrency limits, simulation and monitoring minutes, 14/30/90-day retention, overage burn rates, SOC 2 Type II on all plans, BAA/DPA from Growth, SSO/SAML with SCIM and 99.9% SLA on Enterprise, Bluejay MCP on every plan (accessed 2026-08-24)
- docs.getbluejay.ai — Official docs introduction: positioning as the test-monitor-improve layer for conversational AI agents, agent development lifecycle, custom metrics and production monitoring (accessed 2026-08-24)
- docs.getbluejay.ai/llms.txt — Official docs index, used as the capability inventory: Digital Humans and CSV bulk upload, custom metrics with dynamic variables, Metrics Lab, Scenario Builder, red teaming with attack catalog and OWASP/content-safety mapping, Self-Improve, Bluejay as Code plus Claude Code skill, GitHub Actions, webhooks, uptime monitoring, follow-up SMS grading (accessed 2026-08-24)
- docs.getbluejay.ai/key-concepts/traces/overview — Official: accepts any OpenTelemetry-conformant traces including OpenInference, Langfuse, OpenLLMetry and LiveKit, linked to production evaluations and simulation results (accessed 2026-08-24)
- docs.getbluejay.ai/simulation-integrations/telephony — Official: inbound and outbound telephony simulation, Bluejay-provisioned numbers, HTTP trigger endpoints for dialer integration, concurrency differences between test directions (accessed 2026-08-24)
- getbluejay.ai/platform — Official platform page: multichannel simulations across voice, chat and text, production replays, load testing and red teaming, audio-level metrics (latency, interruption count, word error rate) alongside custom metrics (accessed 2026-08-24)
- ycombinator.com/companies/bluejay — Official YC profile: founded 2025, Spring 2025 batch, San Francisco, team size 10, founders ex-AWS Bedrock and ex-Microsoft Copilot, launch post describing Mimic and Skywatch (accessed 2026-08-24)
- businessinsider.com/bluejay-ai-startup-amazon-micro… — Independent: $4M seed led by Floodgate with YC, Peak XV and Homebrew participating, announced August 2025; confirms San Francisco base and synthetic-customer approach (accessed 2026-08-24)
- every.io/blog-post/end-to-end-testing-for-voice-age… — Independent founder interview, 11/06/2026: 500+ simulation variables, replays of production failures, Fortune 50 and multinational customers, teams moving from fortnightly to near-daily releases (accessed 2026-08-24)