The resource library

We publish whatwe find out.

A firm that sells judgement has to show its judgement in public. Research with numbers, playbooks with steps, templates you can actually use, and comparisons that name the trade-offs — including ours.

Research

Measured,not asserted.

Research · Benchmark

The Multilingual Agent Benchmark

How much accuracy production agents lose between English and the world's major languages — measured across real workflows, with per-language scorecards.

Request the report →
Research · Taxonomy

A Failure Taxonomy for Production Agents

The categories of failure we actually find in production-bound agents, ranked by how often each survives internal testing.

Request the report →
Research · Method

What Live Vetting Catches

What an expert-led live session catches that automated screening does not — with pass-rate data from our own gates.

Request the notes →
Playbooks

How the workactually runs.

Playbook

The FDE Mission Playbook

How to scope, gate and land a forward-deployed engineering mission — the definition of done, the evidence gates, the first two-week ship.

Request a copy →
Playbook

The Red-Team Engagement Playbook

Scoping an independent adversarial test: what to include, who must be in the room, and how findings should be ranked and reproduced.

Request a copy →
Playbook

From Pilot Purgatory to Production

A step-by-step route for the AI pilot that will not ship: the data path, the accuracy bar, the exception path, the owner.

Request a copy →
Templates

Steal these.Seriously.

Template

The Mission Brief

The one-page scope we use on every FDE mission: workflow, data, definition of done, evidence gates, named owners. Use it with any vendor — including one who is not us.

Request the template →
Template

The Written Accuracy Bar

A fill-in framework for defining what good looks like before you build: thresholds, harm taxonomy, exception paths, review cadence.

Request the template →
Template

The AI Vendor Question Set

Twenty questions that separate firms that ship from firms that present — evaluation practice, drift ownership, exit terms, audit trails.

Request the template →
Method notes

The long version,for people who read.

The full working detail behind our assurance engagements — published here rather than crowding the services page.

Method · Evaluation

Model & Agent Evaluation, in full

We write the accuracy bar with you, build a domain-specific evaluation set your team could not assemble alone, run it against your system, and report where it falls short — with named expert reviewers behind every judgement, not a crowd of anonymous raters.

Method · Red-teaming

Adversarial Red-Teaming, in full

A team whose job is to break the system, not launch it. Prompt injection, data exfiltration, tool misuse, jailbreaks, and the domain-specific harms that only a specialist in your field would think to try. Findings ranked by severity with reproduction steps.

Method · Governance

AI Act & Governance Readiness, in full

The technical substrate regulators and auditors actually ask for: model and system inventory, risk classification, evaluation evidence, logging and traceability, and documentation assembled so your legal team is arguing from a file rather than from memory.

Method · Multilingual

Multilingual Evaluation, in full

Your agent works in English. The languages your customers speak deserve the same evidence — accuracy does not survive translation intact. Native-speaker domain reviewers, per-language scorecards, and a clear statement of where you should not yet deploy.

Working checklists

Usable today,no download required.

The two checklists we use most, published in full. Copy them, adapt them, argue with them — they are the checklists we run before we bill anyone.

Checklist 01 · Before you trust an AI system

The production readiness bar

  • A written accuracy bar exists — a number, agreed before testing, that "good enough" actually means.
  • The evaluation set was built by someone who wants it to fail — not sampled from happy-path demos.
  • Domain experts evaluated the outputs — clinicians for clinical, lawyers for legal — not the builders.
  • Failure modes are documented with reproduction steps — ranked by severity, not by convenience.
  • Every exception has a route to a named human — and that route has been tested end to end.
  • Rollback exists and has been rehearsed — turning it off must be a decision, not a project.
  • Someone signed it — a named reviewer, on the record, accountable for the verdict.
Checklist 02 · Before you hire through anyone

The vetting questions that matter

  • Was the technical session live and unscripted? Recorded async challenges are now solvable by AI.
  • Who ran it — a recruiter or a practitioner? Only a peer can probe depth rather than recall.
  • Was identity verified in a live session? Confidence in who you are working with is part of the evidence.
  • Is there a written assessment you can keep? A verdict without a document is a feeling.
  • What happens if the hire fails? A guarantee with conditions is a marketing line — read the conditions.
  • Does the assessor's name appear on the report? Accountability starts with a signature.
The vendor question set

Ten questions to askanyone selling you AI.

Including us. Take these into your next vendor call — the answers separate firms that operate production systems from firms that sell decks about them.

  • What number defines success, and who wrote it down? No written bar means no test.
  • Who evaluates the system — and do they report to the people who built it?
  • Show me a finding you reported against your own work. Silence is an answer.
  • What happens in month nine? Ask who owns accuracy after launch, and what they measure.
  • What does a failed milestone cost me? If the answer is "nothing changes," incentives are broken.
  • Which parts of my workflow will still route to a human, and by what rule?
  • Can I leave? Who owns the prompts, evaluation sets, configuration and reports when you go?
  • Who exactly will do the work? Names, not org charts — and whether they were vetted live.
  • What did the last engagement like mine actually measure? Ask for the number, not the story.
  • What would make you tell me not to do this? A firm with no answer has no judgment to sell.
Sample artifact

What a signed scorecardactually looks like.

A representative extract from an evaluation report — structure real, figures illustrative. Every engagement produces one, and it is yours whether or not you proceed with us.

Criterion
Bar
Measured
Verdict & note
Answer accuracy — routine cases
≥ 95%
97.1%
Pass · 412 cases, expert-adjudicated
Answer accuracy — edge cases
≥ 85%
88.4%
Pass · designed by domain reviewers to fail
Harmful-output rate
0 critical
0
Pass · 3 medium findings raised, resolved, re-tested
Refusal correctness
≥ 90%
86.2%
Conditional · over-refuses in two categories; fix specified
Exception routing to human
100%
100%
Pass · path tested end-to-end, on the record

Signed by the named reviewers who ran it. "Conditional" verdicts ship with a written fix and a re-test date — a system is not "mostly ready."

The production AI glossary

Terms, defined the wayoperators use them.

Plain-language definitions of the vocabulary that fills AI proposals — so you can read your next vendor document without a translator.

Term

Forward-Deployed Engineer (FDE)

A senior engineer embedded inside a client's team, systems and rituals — building to production in the client's stack rather than delivering from outside. Invented at Palantir; now the delivery model of choice for applied AI.

Term

Evaluation (eval)

A structured test of an AI system against a defined set of cases and a written bar. The difference between an eval and a demo: an eval is designed to find the failures, and its result is a number.

Term

Red-teaming

Adversarial testing by people whose job is to break the system — prompt injection, data exfiltration, tool misuse, and the domain-specific harms only a specialist would try. Findings are ranked and reproducible.

Term

Accuracy drift

The slow degradation of a system that passed its launch evaluation — as models change, data shifts and prompts accumulate edits. The reason a one-time test expires, and monitoring accuracy matters more than monitoring uptime.

Term

Evidence gate

A milestone that clears only on demonstrated results — a passing eval, a live integration, a signed finding — never on a slide or a status call. If the gate does not clear, the next milestone does not bill.

Term

Judgment data

Expert human rulings on AI output — preference, critique and acceptance decisions from credentialed specialists. The scarcest input in applied AI, and the one that decides whether a system is safe in a regulated domain.

Term

Human-in-the-loop

A workflow in which defined classes of AI output route to a named person before they take effect. Meaningful only when the exception path is specific, staffed and tested — otherwise it is a diagram, not a control.

Term

Guardrails

The technical constraints around an AI system: input filtering, output validation, tool permissions, spend limits, escalation triggers. Guardrails are what let a system fail safely instead of publicly.

Term

Agent

An AI system that takes actions — calling tools, writing to systems, triggering workflows — rather than only producing text. The action is what makes evaluation, guardrails and operations non-optional.

Comparisons

Trade-offs,named honestly.

Comparison

FDE vs. Staff Augmentation vs. Consultancy

Three ways to buy AI delivery capacity, compared on ownership, incentive alignment, time-to-value and what happens after go-live.

Read the comparison →
Comparison

Internal Eval vs. Independent Assurance

When your own team's testing is enough, and when independence is the entire point — with the failure modes each approach misses.

Read the comparison →
Comparison

Hire vs. Embed vs. Pod

Filling an AI capability gap: permanent search, embedded specialists, or an outcome-scoped pod — costs, risks, and the honest break-even.

Read the comparison →

Signals — the operator's brief on production AI

One email, twice a month: what we shipped, what broke, what we measured, and what it means for people who run real operations. No hype, unsubscribe anytime.

Subscribe →