All outcomes

Home energy surveys · UK

Mack, for CMS Surveyors

An AI operator went live inside the business in May, and by August the operations director was building his own automations on it without me.

Before

Eleven people, ten mailboxes, and a CRM that only told you anything if somebody remembered to open it. Work moved by whoever happened to notice it. The operations director spent his day being the routing layer between an inbox, a calendar and a job board, which is not what a co-founder is for.

After

An operator that reads every mailbox, triages what arrives, drafts the replies, writes the CRM notes and reports on the business, then tells one person what is actually parked on them. It has run continuously since 6 May. The team now writes their own automations against it.

emails triaged since going live in May
10,715
agent sessions in the last seven days
3,409
standing automations the team now runs
40

The problem

CMS Surveyors is eleven people running home energy surveys and retrofit installs across the UK. The work itself was fine. The coordination around it was the problem: ten mailboxes, a CRM called Reonic that held the truth only when someone remembered to update it, and a director of operations who had become the human message bus between all of it. Nothing was broken enough to name, which is the hard kind. There was no single failing process to fix, just a hundred small handoffs that each cost someone four minutes and only existed because no system was watching.

What I built

I forked Kern, the multi-agent system I run my own company on, into their environment: their Supabase, their Railway, their Google Workspace, their interface. Not a demo tenant and not my infrastructure. It reads all ten mailboxes on a reconciler that sweeps every ten minutes, classifies each message by tier, drafts what it can and escalates what it should not touch. It writes back into Reonic. It runs the reporting. The design constraint that mattered most was restraint: consequential actions get classified and held for approval rather than fired, because an ops assistant that emails a customer the wrong thing once is an ops assistant nobody trusts again. Every action it takes is logged as a row naming who it went to, so the question "did that actually happen" has an answer that is not "probably".

What it looks like

Mack dashboard: an ask bar, a summary of what happened since the last visit, a ranked queue of items needing a decision, and the day’s calendar.Data substituted
The screen the operations director opens. Above the fold it answers two questions in this order: what happened while I was away, and what is waiting on me. That order is not a design preference, it came out of a week of his real session logs, which showed his most repeated question was verification rather than instruction. Items in the queue are actionable where they sit, so approving a drafted email does not mean navigating anywhere.
The board: install pipeline counts, jobs in progress, jobs awaiting grid approval, and a panel of standing approvals with toggles.Data substituted
Two halves of one idea. On the left, the work in flight, read live out of their CRM so nobody has to remember to open it. On the right, standing approvals: the specific things Mack has been given permission to stop asking about. An assistant that emails the wrong customer once is an assistant nobody trusts again, so consequential actions are held by default and the client widens that permission themselves, one rule at a time. The last rule still reads "anything over £2,000, always ask".
An agent run inspector showing the instruction, the model’s reasoning, five expandable tool calls with arguments and durations, and the final answer.Data substituted
Every run is inspectable end to end: the instruction, the reasoning, each tool call with its arguments and how long it took, and the answer. This one was asked to chase four outstanding grid connection applications. It read the board, checked what had already been sent to each, drafted three chases, and deliberately did not draft the fourth: two emails had already gone unanswered there, so it put a phone call on the list instead.
Automations list: forty scheduled routines with health status, next run time and weekly cost, with four flagged as needing attention at the top.Data substituted
Forty standing automations, most of them written by the client rather than by me. Each shows its schedule, its health and what it costs to run per week, and the four that need attention are pulled to the top rather than left to be discovered. This screen is the actual outcome of the engagement. The capability moved to them.

Shots marked data substituted are the real interface with invented content in place of the client’s. Every name, address, sum of money and email on those screens is made up. The layout, the components and the behaviour are exactly what the client uses.

What happened

Live since 6 May, 98 days and counting. 28,515 agent sessions, 3,409 of them in the last seven days alone. 10,715 emails triaged since the end of May with 728 drafted replies. In one recent eight-day window it logged 1,117 autonomous actions, 650 emails sent and 297 CRM notes written, every one naming its recipient. The outcome I care about most is not in that list. Four weeks after go-live I rebuilt the interface off a week of the operations director's real sessions, which showed three things I had guessed wrong: his most repeated question was verification rather than instruction, work stalled because he could not see what was waiting on him, and about 85% of his input was voice, not typing. The rebuild answered those three. He now runs 40 standing automations, most of which he wrote himself, and asks for the ability to write more.

Claude Agent SDKSupabase / PostgresRailwayNext.js on VercelGoogle Workspace APIsReonic CRMpgvector

Provenance: Every figure measured 2026-08-11 against the live production Supabase project: ki_sessions, mack_pa_processed, mack_action_log, scheduled_tasks and profiles. Interface findings from a review of one week of real session logs, July 2026.

Want the same thing done to your operation?

Ten working days, a fixed fee, and one automation live before we finish. If there is nothing worth automating, I will say so in the findings.

Book a call