A firm that sells judgement has to show its judgement in public. Research with numbers, playbooks with steps, templates you can actually use, and comparisons that name the trade-offs — including ours.
How much accuracy production agents lose between English and the world's major languages — measured across real workflows, with per-language scorecards.
The categories of failure we actually find in production-bound agents, ranked by how often each survives internal testing.
What an expert-led live session catches that automated screening does not — with pass-rate data from our own gates.
How to scope, gate and land a forward-deployed engineering mission — the definition of done, the evidence gates, the first two-week ship.
Scoping an independent adversarial test: what to include, who must be in the room, and how findings should be ranked and reproduced.
A step-by-step route for the AI pilot that will not ship: the data path, the accuracy bar, the exception path, the owner.
The one-page scope we use on every FDE mission: workflow, data, definition of done, evidence gates, named owners. Use it with any vendor — including one who is not us.
A fill-in framework for defining what good looks like before you build: thresholds, harm taxonomy, exception paths, review cadence.
Twenty questions that separate firms that ship from firms that present — evaluation practice, drift ownership, exit terms, audit trails.
The full working detail behind our assurance engagements — published here rather than crowding the services page.
We write the accuracy bar with you, build a domain-specific evaluation set your team could not assemble alone, run it against your system, and report where it falls short — with named expert reviewers behind every judgement, not a crowd of anonymous raters.
A team whose job is to break the system, not launch it. Prompt injection, data exfiltration, tool misuse, jailbreaks, and the domain-specific harms that only a specialist in your field would think to try. Findings ranked by severity with reproduction steps.
The technical substrate regulators and auditors actually ask for: model and system inventory, risk classification, evaluation evidence, logging and traceability, and documentation assembled so your legal team is arguing from a file rather than from memory.
Your agent works in English. The languages your customers speak deserve the same evidence — accuracy does not survive translation intact. Native-speaker domain reviewers, per-language scorecards, and a clear statement of where you should not yet deploy.
The two checklists we use most, published in full. Copy them, adapt them, argue with them — they are the checklists we run before we bill anyone.
Including us. Take these into your next vendor call — the answers separate firms that operate production systems from firms that sell decks about them.
A representative extract from an evaluation report — structure real, figures illustrative. Every engagement produces one, and it is yours whether or not you proceed with us.
Signed by the named reviewers who ran it. "Conditional" verdicts ship with a written fix and a re-test date — a system is not "mostly ready."
Plain-language definitions of the vocabulary that fills AI proposals — so you can read your next vendor document without a translator.
A senior engineer embedded inside a client's team, systems and rituals — building to production in the client's stack rather than delivering from outside. Invented at Palantir; now the delivery model of choice for applied AI.
A structured test of an AI system against a defined set of cases and a written bar. The difference between an eval and a demo: an eval is designed to find the failures, and its result is a number.
Adversarial testing by people whose job is to break the system — prompt injection, data exfiltration, tool misuse, and the domain-specific harms only a specialist would try. Findings are ranked and reproducible.
The slow degradation of a system that passed its launch evaluation — as models change, data shifts and prompts accumulate edits. The reason a one-time test expires, and monitoring accuracy matters more than monitoring uptime.
A milestone that clears only on demonstrated results — a passing eval, a live integration, a signed finding — never on a slide or a status call. If the gate does not clear, the next milestone does not bill.
Expert human rulings on AI output — preference, critique and acceptance decisions from credentialed specialists. The scarcest input in applied AI, and the one that decides whether a system is safe in a regulated domain.
A workflow in which defined classes of AI output route to a named person before they take effect. Meaningful only when the exception path is specific, staffed and tested — otherwise it is a diagram, not a control.
The technical constraints around an AI system: input filtering, output validation, tool permissions, spend limits, escalation triggers. Guardrails are what let a system fail safely instead of publicly.
An AI system that takes actions — calling tools, writing to systems, triggering workflows — rather than only producing text. The action is what makes evaluation, guardrails and operations non-optional.
Three ways to buy AI delivery capacity, compared on ownership, incentive alignment, time-to-value and what happens after go-live.
When your own team's testing is enough, and when independence is the entire point — with the failure modes each approach misses.
Filling an AI capability gap: permanent search, embedded specialists, or an outcome-scoped pod — costs, risks, and the honest break-even.
One email, twice a month: what we shipped, what broke, what we measured, and what it means for people who run real operations. No hype, unsubscribe anytime.