LLM
The engine itself: which model runs, whether you can choose it, what it costs.
The LLM is the engine every other element in the table feeds, steers, connects or checks. In an audit, Lg does not measure who trained the model — it measures the model surface you get: frontier-class quality, real choice between providers, and honesty about what is answering your prompt. A tool that resells someone else's frontier model can score 9; a tool that hides which model it uses, and swaps it without telling you, cannot.
You need this when
- The output quality of your product is the model's output quality — writing, coding, analysis, support answers.
- Your workload has tiers: high-volume simple tasks plus a hard minority where being wrong is expensive.
- Costs scale with usage, so a per-token price difference of 10x turns into a real line on your P&L.
You can skip it when
- The LLM is a garnish on a workflow product (scheduling, invoicing, CRM) — score the workflow elements instead and treat Lg as a tiebreaker.
- You have one narrow, repetitive task that a small or open-weight model already handles — measure it with evaluations rather than shopping for a bigger engine.
- You are tempted to switch tools purely for a newer model — instead check whether your current vendor lets you select models, since that removes the reason to migrate.
The long version — open when you want the depth
What it is — in one coffee-break
The LLM — large language model — is the engine everything else in this table orchestrates: the thing that actually reads, reasons and writes. Every other element exists to feed it (context, embeddings, retrieval), steer it (prompts, guardrails), connect it (function calling, protocols) or check it (tracing, evaluations). Choosing which engine to run, at what price, is the most-revisited decision in any AI stack — and in 2026 it's a buyer's market.
When you actually need it (and when you don't)
The real question isn't whether you need an LLM — if you're reading this track, something in your business already touches one. The question is which class of model for which job. Our Model Pricing table groups them the way buyers should think: frontier models (the judgment tier — $5–10 per million input tokens) for work where being wrong is expensive; heavy models at a fraction of that for daily production; medium and light models for high-volume simple tasks at prices down to cents per million tokens.
The pattern that saves the most money: match the model to the task, not to the marketing. A €0.10 model classifying support emails all day plus a frontier model reviewing the hard 5% beats one expensive model doing everything — this site runs on exactly that principle, with local open-weight models doing the routine work. And because model quality has commoditized faster than anything else in the stack (a theme we've covered in the deep-dives), the switching cost you should actually fear isn't the model — it's everything you build around it.
How to recognize good vs bad implementations
In the audits, the LLM element (Lg) measures the model surface a tool gives you: quality, choice, and transparency about what's running. The 9-club — ChatGPT, Claude, Gemini, DeepSeek — all pair frontier-class engines with real model choice. More interesting are the tools built ON models: Cursor scores Lg 9 because it lets you pick from every major provider, while some assistants score far lower for hiding which model answers and swapping it without notice. Two tells for any tool: can you choose (or at least see) the model, and does the vendor publish what a task costs?
What this costs
This is the best-documented cost in AI — every serious vendor publishes per-token prices, and we track all of them, sourced and dated, on the Model Pricing tab, including a calculator for your own workload. Two rules of thumb: output tokens cost 3–6× input tokens, and "thinking" models bill their reasoning as output — so a reasoning model can cost several times its sticker price on hard problems. Open-weight models flip the equation: zero per-token cost, but you pay in hardware and operations.
Where to see it scored
Lg scores across the table: Gemini (9), Cursor (9), Suno (8.5 — the engine is the product), Meta Llama (5.5). For head-to-head engine choices, use the AI Versus comparator.
Flashcards
Check yourself
1. What does a high Lg score tell you about a tool?
2. Cursor does not train a model, yet scores Lg 9. Why?
3. Of the 54 tools scored on Lg, the lowest is Freepik (Magnific) at 3. What usually drives a score that low?
4. Which pricing rule of thumb is correct?
5. You process 100k support emails a month with a hard 5% needing judgment. Cheapest sound setup?
Cheat sheet
- Two tells: can you see or pick the model, and are per-task costs published?
- Lg scores quality of access, not ownership — a reseller can legitimately score 9.
- Frontier tier runs $5-10 per million input tokens; light models cost cents.
- Output tokens cost 3-6x input; reasoning models bill their thinking as output.
- Cheap model on the routine 95%, frontier on the hard 5% — the biggest single saving.
- Swapping models is easy; the lock-in is everything you built around the model.
Who actually does this well
| Best on this element | Score | Why it scored that |
|---|---|---|
| GitHub Copilot | 9.5 | The broadest model access on the market: GPT-5.6 Sol/Terra/Luna, Claude Opus 5 and Fable 5, Gemini 3.6 Flash, Grok 4.5, Kimi K2.7 Code and MAI-Code-1-Flash, with a published retire |
| ChatGPT / OpenAI Platform | 9 | GPT-5.x is frontier-grade across chat, code, agents and Computer Use. |
| Claude / Anthropic | 9 | Claude Fable 5 (GA 09/06/2026) is a Mythos-class frontier model; independent reporting puts Anthropic at the top of enterprise LLM API share. |
| Cursor | 9 | Frontier multi-vendor access (GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro, Grok 4.3) plus in-house Composer 2.5 tuned for fast codebase work — quality-of-access is excellent and not l |
| DeepSeek | 9 | Artificial Analysis places V3.2 (Reasoning) level with Kimi K2 Thinking and ahead of Grok 4 and Claude Sonnet 4.5 (Thinking); V4 adds 1M context — frontier quality at open-weight p |
| Gemini / Google | 9 | Gemini 3.5 Flash is frontier-grade for agents and coding (and 4x faster than peers per Google), with 3.1 Pro and Deep Think above it — a genuine frontier lineup. |
| Decagon | 8.5 | Multi-provider by design (OpenAI, Anthropic, Gemini) layered with an increasingly dominant in-house fine-tuned stack — 80%+ of Decagon's own inference traffic as of March 2026. |
| Genspark | 8.5 | No proprietary frontier model; instead orchestrates 80+ third-party models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4.20, DeepSeek and more) with automatic routing to the be |
And the other end of the same column:
| Weakest | Score | Why it scored that |
|---|---|---|
| Cognism | 4 | Third-party models sit behind AI Search and Research by Cognism AI, unnamed and unselectable by the customer. |
| Scrut Automation | 4 | Nothing current is named. The only model disclosure I could find is a 2023 launch note saying ScrutGPT leverages LLMs and OpenAI. Three years on, the AI security page runs long on |
| Freepik (Magnific) | 3 | The core product runs on third-party generative-media models, not a chat LLM; the Agents/Assistant layer exposes a thin NLP planning surface for workflow orchestration rather than |
| Guild | 3 | No model is named, versioned or selectable in any public Guild documentation. The language model is an assist layer on a human-delivered service, not the core of the product. |
Scored on 136 of 144 audited tools. Every score links to the full audit and its reasoning.