AI Implementation Engagement
How we build and integrate AI systems end to end, with the evaluation set, guardrails and rollback path treated as part of the build.
| How it is bought | Fixed project |
| Duration | 6 to 12 weeks, scoped before we start |
| Written for | Teams with one identified use case and an appetite to run it in production |
88% of agent pilots never reach production, and the most-cited blocker is evaluation rather than model quality. Most pilots were never built to be shippable.
What you receive
Named artifacts, not activities. If one of these is not delivered, the engagement is not complete.
- A working agent in your environment, against one use case, in production
- The evaluation set, built before the agent and agreed with the process owner
- The guardrail set: allow-list, value ceiling, blast-radius ceiling, reversal path, and a trust boundary on retrieved content
- Logging that would survive an investigation: inputs, outputs, model and prompt version, who was affected
- A runbook, and a handover session with whoever will operate it
How it runs
- Weeks 1 to 2: the evaluation set and the boundary. Nothing is built until 'good enough' is defined in numbers.
- Weeks 3 to 8: build, measure against the evaluation set, iterate.
- Final weeks: production hardening, logging, runbook, handover.
What this deliberately does not include
Published with the same prominence as the deliverables. A scope with no stated exclusions is a scope that will be argued about later, and the argument always happens at the worst possible moment.
- A second use case. One at a time is deliberate, because the second is much cheaper once the first has taught you what your data is really like.
- Model training. This builds on existing models unless we have agreed otherwise.
- Ongoing operation, unless you add AI Automation below.
How you know it is finished
It passes the agreed evaluation set, it is running in production, and your team can operate it without us.
What you need ready
A named process owner, access to a non-production environment, and historical examples of the work the agent will do.
Every engagement starts with a written scope confirmation before any invoice is raised. Nothing here is a click-to-buy: we confirm what you need, agree the scope in writing, then invoice. If the scoping conversation shows that a smaller engagement, or none at all, is the right answer, we say so before anything is signed.
Request this engagement Talk it through first See the guardrails demonstrated
Building an AI feature is the short part. Making it dependable enough that you can leave it running is the work, and it is where most implementations stop short.
We build the whole thing: the feature, the evaluation set that proves it still works next month, the guardrails that decide what it refuses, and the rollback path that has actually been tested.
What "end to end" means here#
| Included from day one | Why it is not a later phase |
|---|---|
| The held-out evaluation set | Without it, "it seems better" is the only available verdict, and quality drifts unnoticed |
| Guardrails as testable rules | Written as code, not as intentions in a prompt |
| Retrieval grounded in your data | With your permissions preserved, not flattened |
| Cost per successful outcome | A feature that delights and loses money at scale is a failure discovered late |
| The autonomy boundary | Exactly what it may do without a person, written down per feature |
| A rehearsed rollback | Timed, not described. The number decides how boldly anyone can move afterwards |
| Disclosure and marking | Required in the EU since 2 August 2026 where a person interacts directly with AI |
How the build runs#
A person decides the architecture. Agents write the change and its tests in the same pass. Every gate after that is real and allowed to stop the change: code review, the test suite, the security scan, then CI/CD.
The order is deliberate. Review runs before the tests, because a review after a green suite only reviews what the tests failed to notice. The security scan runs before deployment, because a finding in production is an incident and the same finding ten minutes earlier is a task.
What you own at the end#
Everything. The code, the evaluation set, the model cards, the guardrail specifications and the runbook. No component of the delivery depends on us continuing to be involved, and we will say plainly which parts your team can maintain and which will need a skill you do not currently have.
What we will not do#
- Ship without an evaluation set. It is the most valuable asset in an AI product and the one most often treated as an afterthought.
- Widen what a system does unsupervised because a model improved. That is a decision a person records, not a consequence of an upgrade.
- Promise a capability we have not built before without saying so. If something is new to us, you will hear that during the proposal rather than during the project.
How we work, in public#
The engineering model and the AI product discipline we use are published, including the KPIs, the release procedures and the templates for model cards and autonomy boundaries.
We also run our own AI products, which means these practices exist because the alternative was found out on live systems rather than because they read well.
Getting started#
The useful first conversation is about one specific process and what a wrong answer would cost. Get in touch.
What else is coming for AI Implementation
The engagement Ready
What it produces and how it runs.