Services

Three practices.One loop.

We build it, we prove it, we keep it true — Deploy, Assure, Operate — powered by a network of verified experts. Every engagement has a written scope, a named owner and an exit condition, and a thirty-minute working session in front of it that produces a written range before you commit to anything.

Talk to us → Engagement index
The three practices 01 · Deploy 02 · Assure 03 · Operate The expert network Engagement index How we vet Principles FAQ
The three practices

Build it. Prove it.Keep it true.

One loop, three doors in. Enter wherever it hurts — a workflow that will not ship, a system nobody dares to trust, or an agent that decayed after launch.

All three practices run on the Expert Network — engineers, clinicians, lawyers, accountants and linguists who passed six vetting gates — and are counted on Q-Base, the platform we also sell as a product. Front door: a thirty-minute working session, and where useful a two-week readiness sprint that ends in a mission brief, ready to execute.

Practice 01 · Deploy

Put one thing into productionbefore you plan the next ten.

AI fails as a programme and succeeds as a single workflow. We take the one with the clearest economics, build it properly — data, evaluation, guardrails, integration — and then stay on to run it, because month nine is when everyone else has already left.

The flagship · how Offer 01 is built

Mission Pods

The forward-deployed engineer (FDE) — a senior AI engineer embedded inside your team, your stack and your rituals — is the role Palantir invented and the AI labs now hire by the hundreds. We supply it as a service: the FDE builds inside your stack and your rituals; a delivery lead owns scope, stakeholders and the evidence gate. Mission-scoped with a written definition of done, milestones that clear only on evidence, and progress you watch live on a shared delivery dashboard.

Mission-scopedEvidence-gatedUS/EU overlap + overnight progressStand down anytime after the mission
Solo

Embedded FDE

One senior AI engineer, embedded full-time. LLM integrations, agents, RAG, MCP servers, evaluation harnesses — shipped to production.

Monthly · mission-scoped
Pod

FDE Pod

FDE + delivery lead with one mission and a written exit. The unit that turns "we bought AI" into a workflow that runs.

Monthly · mission-scoped

First production ship inside two weeks — deliberately small, deliberately real. If a milestone does not clear its evidence gate, the next one does not bill.

Inside a mission · 4–8 weeks

Agent Launch Sprint

One production workflow end to end: built, integrated, tested against a written accuracy bar, shipped behind guardrails with an exception path to a human, handed over with runbooks.

Range named on the first call
Inside a mission · 4–10 weeks

Data-Ready Engagement

The work that decides whether anything above it survives: sources connected, unstructured content prepared, retrieval built and tuned, quality gates measured against a stated threshold.

Range named on the first call

Front door for new clients: a two-week AI Readiness Sprint — workflows and data assessed, opportunities ranked, ending in an actionable mission brief. Sprint fees are credited against a mission booked within sixty days.

Practice 02 · Assure

Somebody has to tryto break it first.

Independent evaluation is now a regulatory expectation in Europe and a board expectation everywhere else. The strongest evidence comes from qualified reviewers outside the build team — independent, credentialed, and human.

Shape 01 · 3–5 weeks

Model & Agent Evaluation

We write the accuracy bar with you, build a domain evaluation set your team could not assemble alone, run it, and report where the system falls short — expert-reviewed, in 20+ languages where needed.

Fixed scope · range named on the first call
Shape 02 · 3–6 weeks

Adversarial Red-Teaming

A team whose job is to break the system: prompt injection, data exfiltration, tool misuse, and the domain-specific harms only a specialist would try. Findings ranked, with reproduction steps.

Fixed scope · range named on the first call
Shape 03 · 4–8 weeks

AI Act & Governance Readiness

The substrate regulators actually ask for: system inventory, risk classification, evaluation evidence, logging and documentation — assembled so legal argues from a file, not from memory.

Fixed scope · range named on the first call
Shape 04 · 4–6 weeks

Multilingual Evaluation

Native-speaker domain reviewers, per-language scorecards, and a clear statement of where you should not yet deploy — because accuracy does not survive translation intact.

Fixed scope · range named on the first call

Any of these can roll into Continuous Assurance — the quarterly programme above — once the first baseline is set. Boards, regulators and enterprise customers increasingly ask for evidence that testing recurs, not that it happened once before launch.

Practice 03 · Operate — the recurring engine

Launches are moments.Systems are forever.

Models get swapped, prompts get edited, data drifts, and the audience for evidence — boards, regulators, enterprise customers — keeps asking for this quarter's proof, not last year's. These two programmes exist because AI is never finished.

Rolling programmesMonthly or quarterly cadence · named owner · cancel at cycle end
Continuous Assurance · quarterly

The standing evidence pack

Quarterly re-evaluation and adversarial testing against your evolving system, run by named experts with no stake in your roadmap. Every cycle updates a signed evidence pack — the artifact your board cites, your regulator requests, and your enterprise customers ask for in procurement.

  • Quarterly eval + red-team cycles against the written bar
  • Findings ranked, reproducible, signed by named reviewers
  • Evidence pack maintained as a living document
Managed Agent Operations · monthly

Monitoring for truth, not uptime

Your cloud dashboard says the agent is up. It does not say the agent is right. We watch accuracy drift, cost per task, exception volume and tool-use failures — with human escalation paths, monthly re-evaluation, and a named owner who still knows your account in month nine.

  • Accuracy, cost and exception monitoring with human escalation
  • Monthly re-evaluation against the agreed bar; model routing as economics shift
  • New automation scoped each quarter from what operations taught us

Both programmes are where an engagement naturally lands: a mission ships the system, assurance proves it, and the retainers keep both true. Ranges are named on the first call, in writing.

The Expert Network · powering all three practices

Specialists you could nothave found alone.

Senior AI roles take three to five times longer to fill than ordinary software roles, and verifying real, working expertise has never mattered more. Both challenges are answered the same way: a smaller, deeply vetted pool, with a human at the final gate.

Per engagement

Verified Expert Panels

Clinicians, lawyers, accountants, engineers, linguists and scientists assembled to evaluate AI output — for model training, evaluation, preference data or acceptance testing. Credential-verified, paid at specialist rates, and briefed to challenge the system rather than confirm it.

Credential-verifiedDomain-matched20+ languages
Per candidate · 48 hours

Verified Technical Assessment

Assessment as a standalone service, for candidates you sourced yourself. A senior specialist in the same discipline leads a live, unscripted session and issues a signed competency report — the one gate AI assistance cannot sit behind.

Live · unscriptedPractitioner-ledSigned report
Follow-on · monthly

Embedded Experts

Vetted engineers, data and AI specialists working inside your team's rituals and tools — matched in days, productive in week one, scaled either direction as the work changes. Knowledge transfers by default, so you are never dependent on us to change a workflow.

Days to matchInside your teamScale either way
Follow-on · per placement

AI & FDE Search

Specialist permanent search for the roles the market cannot supply: AI engineers, forward-deployed engineers, ML platform and applied research hires. Structured scorecards, a live expert session on every finalist, and a signed competency report you keep whether or not you hire.

Specialist searchSigned reportsReplacement guarantee

We are deliberately smaller than the marketplaces. Every person in the network has passed a live session with a senior specialist in their own discipline — a process that does not scale to millions and is not meant to. If you need ten thousand annotators, we are the wrong firm.

The engagement index

Every way to engage,on one page.

No engagement here is open-ended. Each has a written scope, a typical length, and an exit condition — and you get a written range on the first call, before you commit to anything.

Engagement
What you get
Typical length
PRACTICE 01 · DEPLOY
AI Readiness Sprint
Two weeks: workflows and data assessed, opportunities ranked, ending in a mission brief — the front door, ready to execute.
2 weeks
Embedded FDE
Senior forward-deployed AI engineer, full-time inside your team, shipping to production.
Monthly · rolling
FDE Pod
FDE + delivery lead, mission-scoped with written exit criteria and evidence-gated milestones.
6–12 weeks per mission
PRACTICE 02 · ASSURE
Model & Agent Evaluation
Custom domain evaluation set, expert-reviewed run, baseline report against a written bar.
3–5 weeks
Adversarial Red-Teaming
Independent adversarial testing with severity-ranked, reproducible findings.
3–6 weeks
Governance Readiness
Inventory, risk classification, evaluation evidence, logging and documentation pack.
4–8 weeks
Continuous Assurance
Quarterly re-evaluation and red-teaming against an evolving system.
Rolling
PRACTICE 03 · OPERATE
Managed Agent Operations
Monitoring, monthly re-evaluation, model routing, quarterly new automation, named owner.
Rolling
THE EXPERT NETWORK
Verified Expert Panel
Credential-verified domain experts for training, evaluation or acceptance testing.
Per engagement
Verified Technical Assessment
Live, unscripted session led by a senior expert; signed competency report you keep.
48 hours

Agent Launch and Data-Ready work run inside missions; multilingual evaluation runs inside any assurance engagement. Readiness-sprint fees are credited against a mission booked within sixty days; assessment fees against a placement. Third-party model and licence costs pass through at cost, itemised.

How we vet

Six gates, and the last oneis a human being.

Credentials and automated screens tell only part of the story — real expertise shows in live conversation. That is why every gate in our process is designed to surface genuine, working knowledge, and why the final one is human.

So our final technical gate is deliberately expensive and deliberately human: a senior practitioner in the same discipline, leading a live, unscripted session. It is deliberately unscalable — and it produces a signed competency report the expert carries into every engagement after.

Assess a candidate you sourced →
A live expert-led assessment sessionGate 04 · live, on camera
Gate 01 · Source

Curated intake

Specialists enter by referral, verified track record and targeted search. There is no open signup, and we do not buy lists.

Gate 02 · Screen

Structured assessment

Role-specific technical assessment scored against a rubric agreed with you in advance — not a conversation and a gut feel.

Gate 03 · Craft

Real-work review

Portfolio and production work reviewed by a senior specialist in the same discipline, not by a recruiter with a checklist.

Gate 04 · Live

Live, unscripted session

Unscripted, practitioner-led, probing depth rather than recall. Identity verified. This is the gate AI assistance cannot sit behind, and the one we will not skip for speed.

→ Signed competency report issued
Gate 05 · Validate

Human sign-off

An experienced consultant validates every match against your actual context before it reaches you. No automated shortlists, ever.

Gate 06 · Measure

Performance feeds back

Delivery outcomes are counted — utilisation, milestones, findings quality — and inform every future match.

How an engagement runs

Named owners.Written exit criteria.

If a step does not clear its criterion, we do not start the next one and you do not pay for the next one. That rule is in the contract, not just on this page.

Week 0
01 · Scope

Working session

The system you are unsure about, or the role you cannot fill. Thirty minutes, and a price range before you leave the call.

Week 1
02 · Define

The bar, in writing

What good looks like, what failure looks like, who is qualified to decide, and what evidence you need at the end.

Week 1–2
03 · Assemble

The right people

Experts sourced and verified, or engineers matched. Named individuals, with their assessment reports attached.

Week 2–5
04 · Execute

The actual work

Evaluation run, system built, or role filled — against the criterion agreed in week one, tracked live on Q-Base — the same delivery platform we sell as a product.

Week 5–6
05 · Report

Findings and handover

An honest report including what we could not establish. Runbooks and evidence packs you keep and can hand to an auditor.

Ongoing
06 · Sustain

Re-test & operate

Quarterly re-evaluation, managed operations, or a replacement guarantee — depending on which practice you engaged.

How we work

Principles we do not trade away.

Humans decide

AI does the volume — sourcing, screening, drafting, reconciling. People make the calls. Every meaningful recommendation and every finding carries a named human sign-off.

Independence is the product

An evaluation designed for a system you built is a rehearsal, not a test. Where we built the thing, we say so at the top of the report — and sometimes tell you to hire someone else.

Counted, not claimed

We report what happened — candidates assessed, findings raised, runs completed. We will not model business value from assumptions you never set.

Embedded, not outsourced

Our people work inside your rituals and tools, and knowledge transfers by default. You should be able to run it without us. That is the test we set ourselves.

Questions

Frequently asked.

Do the three practices have to be bought together?+
No, and most clients start with one. They do compound — the network supplies the experts who run evaluations, evaluations tell us what is safe to deploy, and operations feeds the next mission — but we will not make you take one to get another.
Can you red-team something we built ourselves?+
That is the ideal case. Independence is where assurance work gets its value, and we have no stake in the design decisions we are testing. Where we did build the system, we disclose it at the top of the report.
How do you handle our data and our candidates' data?+
Your tenant, your region, a DPA before work starts, least-privilege access reviewed at each stage, and evaluation data destroyed or returned on a schedule you set. Where a workflow touches regulated data we agree the handling model in writing first. Certification status and security documentation are shared with your team under NDA.
Who owns what we build and what you find?+
You own the workflows, prompts, configuration, evaluation sets and reports. We retain our own reusable tooling and methods. Exports on demand, no exit penalty, and we will help you migrate if you leave — leaving has to be possible for staying to mean anything.
One call, three practices

Bring the system you doubt,or the role you cannot fill.

Thirty minutes. A scoped answer, a price range, and an honest view — including that you should wait, if that is what we think.

Talk to us → Back to overview