Crafted by machines, Perfected by humans.
We design and build the software underneath the model — data paths, evaluation, deploy, and the interface a person actually uses. One continuous stroke from problem to product.
Four practices, one delivery model.
AI engineering
Retrieval, agents, and evaluation harnesses. We measure the model before we ship the feature.
ExploreSoftware development
Typed services, boring infrastructure, and test suites a new engineer can read on day one.
ExploreProduct engineering
Interfaces for systems that are partly probabilistic. Design and code in the same repository.
ExplorePlatform & automation
The runtime our clients keep: pipelines, observability, cost controls, and human review in the loop.
ExploreSoftware that bends to the human.
Intelligence that keeps its shape. Six commitments we write into the statement of work, not the pitch deck.
Engineering quality
Review, types, and CI from commit one. No prototype quietly becomes production.
AI expertise
Eval sets before features. We report the failure rate we did not fix yet.
Speed with a brake
Two-week increments, each one deployable. Velocity is measured in releases, not tickets.
Reliability
Budgets, alerts, and runbooks handed over with the code. On-call is documented.
Human-centred systems
A person can always see why the system decided, and overrule it.
Technical depth
Nine-year median tenure. The engineer in the kickoff is the engineer on the deploy.
One continuous stroke, seven turns.
-
01
Discover
Constraints, data, and the real workflow.
-
02
Define
Scope, success metric, eval set.
-
03
Design
Interface and system in one pass.
-
04
Build
Two-week deployable increments.
-
05
Test
Regression, load, and eval gates.
-
06
Deploy
Staged rollout with rollback.
-
07
Improve
The point where the system learns.
Named mechanisms, specific numbers.
Atlas Freight — dispatch that explains itself
A routing assistant that proposes, shows its reasoning trace, and waits for a dispatcher. Eleven months in production across four depots.
Vero Health — triage notes in 40 seconds
Clinician-reviewed summarisation with a full audit trail on every generated line.
Northbank — reconciliation without the spreadsheet
Nightly matching across seven ledgers, with exceptions routed to a named analyst.
The stack we actually run.
- Retrieval & agents
- pgvector · LangGraph
- Evaluation
- Braintrust · custom harness
- Services
- Python · Go · TypeScript
- Data
- Postgres · dbt · Kafka
- Runtime
- K8s · Terraform · OTel
Eval scores from the harness we hand over with the code. Published weekly, including the regressions.
Read the research logWhat they say afterwards.
They shipped the boring parts first. Six weeks in we had a deploy pipeline and a failing test we understood.
The first document was an evaluation plan, not a proposal. That told us everything.
Questions we get in the first call.
How do engagements start?
A 45-minute technical call, then a two-week paid discovery that ends with a scope sketch, an eval plan, and a fixed price for the first increment.
Do you work with our existing team?
Usually. We pair with in-house engineers and hand over ownership deliberately — runbooks, on-call docs, and a walkthrough per subsystem.
What if the model is not good enough?
We find that out in discovery, on your data, against an eval set you agree with. If the accuracy is not there we say so before you commit to a build.
Who owns the code?
You do, from the first commit, in your repository. No licensed runtime, no vendor lock, no per-seat platform fee.
What does it cost?
Discovery is fixed. Builds run as two-week increments priced per team; a typical first release lands between six and twelve weeks.
Bring us the part that keeps breaking.
A 45-minute technical call with the engineer who would lead the build. You leave with a scope sketch either way.