Service 01

AI product engineering

We build language-model features into products that already have users — retrieval over your own data, agents that take real actions, extraction and classification pipelines.

Book a technical call

Overview

The interesting part of an AI feature is not the prompt. It is everything around it.

Evaluation suites, so you can tell whether a change actually helped rather than guessing from three examples. Fallbacks for when the model is slow, wrong, or down. Cost ceilings enforced in code rather than discovered on an invoice. Observability that shows you what users really asked for, not what you assumed they would.

We build that scaffolding on the first pass. Retrofitting it after launch is how AI features quietly get switched off six weeks later — and it is the single most common reason a promising demo never becomes a product.

What we build

  • Retrieval over your own documents, databases and internal APIs
  • Agentic workflows with tool use and human approval steps
  • Structured extraction and classification pipelines over messy input
  • Evaluation harnesses and regression suites that gate model changes
  • Streaming interfaces in React 19, with honest loading and error states
  • Prompt and model version control, with a rollback that takes seconds
  • Token budgets, rate limits and caching, enforced server-side
  • Provider abstraction, so changing model is a config change and not a rewrite

What we commit to

Written into the engagement, not implied.

Streaming from the first sprint
Sprint 1
Evaluation suite must pass at release
Gated
Cost ceiling agreed in writing
In writing
Deterministic fallback when the model is down
Fallback
How it runs

Three phases, and you can stop after the first.

The assessment is priced and delivered on its own. If the answer is that you should not build this, you have paid for one phase and saved the other two.

  1. Assess

    We take the workflow apart and find where a model earns its cost and where it does not. You get a written scope, an evaluation plan and a cost model — yours to keep, and legible to someone who has to approve the budget.

  2. Build

    Two-week increments against the evaluation suite agreed in phase one. Every release either moves the score or explains why it did not. The scaffolding ships with the feature, not after it.

  3. Operate or hand over

    We keep running it under an agreed service level, or we hand it over with runbooks, dashboards and a trained team. Both are real options — we do not make the handover deliberately painful.

The uncomfortable part

A good half of the AI work we are asked to quote should be a database query, a rules engine, or a better form.

We will say so in the assessment, at the same price. It costs us a build and saves you a system you would have had to keep feeding.

  • A chatbot over content nobody reads
  • Summarising records a person already skims in four seconds
  • Natural-language search across six well-labelled filters
  • Anything where a wrong answer is expensive and unreviewed
FAQ

About this service.

Questions about this service, kept separate from the ones answered on the home page.

Not necessarily. We scope the data boundary first — what leaves your systems, what stays, and what can run on models hosted inside the EU or on your own infrastructure. That decision is made before any code is written, not after.

The evaluation suite tells you which. Model upgrades run against it before they reach users, and a regression means the old version stays. This is the main reason we insist on building it in phase two rather than later.

Tell us what you are trying to build.

Two sentences about the product and the constraint. We reply within one working day — with a plan, or an honest referral.

Start a conversation