RXed AI News

AI to the bone.
Audited 2026-08-24 · RXed table v1.0

Bluejay

Visit getbluejay.ai
“Test, monitor, and improve voice & chat AI agents” — the vendor’s own words

Ten people and a $4M seed, and the feature list already reads like a mature QA platform: MCP, OpenTelemetry ingestion, OWASP-mapped red teaming, CI gates and a free tier. Buy it for the testing, not for the company's balance sheet.

Best for: Teams shipping voice or chat agents who want regression gates in CI and audio-level evaluation without buying an enterprise contract, and who can live with a 10-person vendor.
Scope17/20
Quality7/10
Where the quality sits
7Reactive
6Retrieval & Memory
7Orchestration
8Validation
7Models
SpecialistAutomation & AgentsVoice & SpeechVoice AgentsSecurity & ComplianceFreemium
Vendor
Bluejay · getbluejay.ai
Origin
US — San Francisco
Pricing
Pay-as-you-go $0/mo + usage · Growth $500/mo · Scale $1,000/mo · Enterprise Custom
Users (official only)
Not disclosed as a user count. Bluejay says enterprises, startups and developers use the platform and cites Fortune 50 customers; team size 10 per its Y Combinator profile, founded 2025, $4M seed led by Floodgate. (source, 2026-08-24)
Pay-as-you-go$0/mo + usage$25 free credits, up to 25 concurrent simulations, 14-day retention, fastest credit burn
Growth$500/mo100 concurrent, ~1,500-1,600 simulation minutes, 13,000 monitoring minutes, signed BAA and DPA, standard RBAC, 30-day retention
Scale$1,000/mo200 concurrent plus load testing, ~4,000-4,300 simulation minutes, 34,000 monitoring minutes, 90-day retention, lowest cost per minute
EnterpriseCustomRedline BAA and DPA, SSO/SAML with SCIM, custom RBAC, uptime SLA up to 99.9%, dedicated engineer

Overage credits burn at 1.8x on Growth and 1.5x on Scale, so the headline price is a floor, not a cap. Pay-as-you-go has no overage because it is prepaid. EUR figures are converted at roughly 0.87 and are indicative only.

checked 2026-08-24 · vendor pricing page

Element scores

Reactive
Retrieval & Memory
Orchestration
Validation
Models
Primitives
Pr8
Prompts
Em5
Embeddings
Cx7
Context
Tr9
Tracing
Lg6
LLM
Compositions
Fc7
Function calling
Vx
Vector store
Rg5
RAG
Gr7
Guardrails
Mm9
Multimodal
Deployment
Ag7
Agents
Ft5
Fine-tuning
Fw9
Frameworks & harnesses
Ev9
Evaluations
Sm
Small models
Emerging
Ma6
Multi-agent
Sy8
Synthetic data
Pc8
Protocols
In7
Interpretability
Th
Thinking models
Tap or hover any element to see why it got that score.

Strengths

The evaluation stack is the deepest part and it is genuinely deep: 70+ metrics plus unlimited custom ones, Metrics Lab to align the judge to your own reviewers, OWASP-mapped red teaming with a content-safety taxonomy, load testing to 200 concurrent, and production replays that re-run real failures until the agent passes. The developer surface is unusually complete for a company this young, with Bluejay as Code, SDKs, a CLI, webhooks, GitHub Actions gates, a Claude Code skill and an MCP server on every plan including the free one. OpenTelemetry ingestion means it slots next to Langfuse or OpenInference instead of replacing them.

Honest dings

Bluejay grades other people's agents and publishes nothing about the quality of its own judges, and it names no model anywhere, so a silent judge swap moves your scores with no changelog to point at. The company is 10 people on a $4M seed round, founded in 2025, which is real vendor risk for something you wire into a CI gate. Retention is short at the bottom (14 days on pay-as-you-go), overage credits burn at 1.5-1.8x above the included minutes, and no A2A support is documented.

Prices and details change — this passport is re-verified at least quarterly.
Sources (9) — every claim traceable

Every audit lists the research it rests on — transparency and traceability are the product. Tools evolve: each audit is a snapshot of its audit date, and re-audits supersede older versions (kept below for reference).

  • getbluejay.ai/pricing — Official pricing and full feature matrix: four tiers, concurrency limits, simulation and monitoring minutes, 14/30/90-day retention, overage burn rates, SOC 2 Type II on all plans, BAA/DPA from Growth, SSO/SAML with SCIM and 99.9% SLA on Enterprise, Bluejay MCP on every plan (accessed 2026-08-24)
  • docs.getbluejay.ai — Official docs introduction: positioning as the test-monitor-improve layer for conversational AI agents, agent development lifecycle, custom metrics and production monitoring (accessed 2026-08-24)
  • docs.getbluejay.ai/llms.txt — Official docs index, used as the capability inventory: Digital Humans and CSV bulk upload, custom metrics with dynamic variables, Metrics Lab, Scenario Builder, red teaming with attack catalog and OWASP/content-safety mapping, Self-Improve, Bluejay as Code plus Claude Code skill, GitHub Actions, webhooks, uptime monitoring, follow-up SMS grading (accessed 2026-08-24)
  • docs.getbluejay.ai/key-concepts/traces/overview — Official: accepts any OpenTelemetry-conformant traces including OpenInference, Langfuse, OpenLLMetry and LiveKit, linked to production evaluations and simulation results (accessed 2026-08-24)
  • docs.getbluejay.ai/simulation-integrations/telephony — Official: inbound and outbound telephony simulation, Bluejay-provisioned numbers, HTTP trigger endpoints for dialer integration, concurrency differences between test directions (accessed 2026-08-24)
  • getbluejay.ai/platform — Official platform page: multichannel simulations across voice, chat and text, production replays, load testing and red teaming, audio-level metrics (latency, interruption count, word error rate) alongside custom metrics (accessed 2026-08-24)
  • ycombinator.com/companies/bluejay — Official YC profile: founded 2025, Spring 2025 batch, San Francisco, team size 10, founders ex-AWS Bedrock and ex-Microsoft Copilot, launch post describing Mimic and Skywatch (accessed 2026-08-24)
  • businessinsider.com/bluejay-ai-startup-amazon-micro… — Independent: $4M seed led by Floodgate with YC, Peak XV and Homebrew participating, announced August 2025; confirms San Francisco base and synthetic-customer approach (accessed 2026-08-24)
  • every.io/blog-post/end-to-end-testing-for-voice-age… — Independent founder interview, 11/06/2026: 500+ simulation variables, replays of production failures, Fortune 50 and multinational customers, teams moving from fortnightly to near-daily releases (accessed 2026-08-24)