Insights

What we've learned building these systems.

Working notes on the problems that come up in every multi-entity engagement — consolidation, cost, and data isolation. Written for the person who has to make the decision, with the numbers we actually publish.

Operations · 9 min read

How to consolidate reporting across outlets on different POS systems

A method for getting one set of numbers out of outlets on different POS systems: ingestion, reconciliation against close reports, item normalisation.

Commercials · 7 min read

What an AI operating system actually costs to run

The three lines behind the monthly running cost of a bespoke AI operating system — inference, infrastructure, integration — and which grow with outlets.

Architecture · 8 min read

Multi-entity data isolation: a filter is not a boundary

Why filtering by entity_id in application code is not isolation, what database-boundary enforcement looks like, and how group reporting survives.

Adoption · 8 min read

Why most AI agent pilots never reach production

Survey data puts agent pilot failure at 88%. The blockers are evaluation, governance and reliability — and the first one is what actually kills projects.

Commercials · 7 min read

The 2026 model tiers, and how to route work between them

Published August 2026 model prices, and the routing rule that decides a bill: a reasoning tier for judgement, a cheap fast tier for everything at volume.

Architecture · 8 min read

MCP went stateless: what the 2026-07-28 spec changes

The largest MCP revision since launch: a stateless core, versioned extensions and OAuth-aligned authorization. What it changes for connecting POS, ERP and HR.

Operations · 7 min read

Where back-office agents pay back, and where they don't

Invoice processing time falls 70–90% and close accelerates 30–50%, yet finance agents take 8.9 months to pay back against 3.4 for sales. Why, and what to do.

Tooling · 8 min read

The AI tool stack for a multi-entity operator

What to actually use in 2026, layer by layer: model tiers, MCP connectors, agent platforms and the data foundation underneath — plus what not to buy yet.

Buying · 8 min read

The questions to ask an AI vendor before you sign

Twelve questions that separate a vendor who has shipped from one who has demoed — ownership, running cost, failure handling, and life after handover.

Operations · 8 min read

Is that outlet actually profitable, or just busy?

Why per-outlet P&Ls disagree with reality: shared-cost allocation, comparing outlets of different sizes, and separating real growth from more traffic.

Controls · 7 min read

What happens when an AI agent gets it wrong

Agents fail. What decides the damage is reversibility, hold-back queues, run records and which actions a system is never allowed to take unsupervised.

Commercials · 7 min read

Another hire, an outsourcer, or an agent?

Comparing a headcount, an outsourced team and an AI agent by cost shape, supervision load and what happens to the work when volume moves.

Architecture · 8 min read

RAG, fine-tuning or long context: which one your problem needs

The three ways to put your own data in front of a model, what each one actually changes, and the question that decides between them.

Governance · 7 min read

Where your data actually goes when you use an LLM

What happens to business data sent to a model provider — retention, training, residency, sub-processors — and which questions get you a straight answer.

Security · 8 min read

Prompt injection: the attack that has no clean fix

Why prompt injection cannot be patched away, what makes agents with tool access genuinely dangerous, and the architectural controls that work.

Commercials · 7 min read

Build or buy: how to decide on an AI system

A decision rule for AI build-versus-buy that survives contact with reality, and the two failure modes that cost the most on each side.

Governance · 7 min read

AI governance that fits an organisation without a compliance department

How to govern AI use without a compliance function: an inventory, a risk tier, four rules people follow, and the reviews that actually catch things.

Architecture · 8 min read

When a small local model beats a frontier API

The cases where a small self-hosted model genuinely beats a frontier API — and the hidden costs that make self-hosting more expensive than it looks.

Commercials · 7 min read

How to measure whether AI is actually paying back

Why AI ROI measurement usually fails, the baseline nobody captures in time, and a measurement design that survives a finance review.

Architecture · 7 min read

Context engineering: the part that decides whether it works

Why the wording of a prompt matters less than what surrounds it, and the engineering decisions that actually determine an AI system's reliability.

Operations · 7 min read

Voice AI on the phone: what works and what still doesn't

What voice AI genuinely handles on inbound calls, where it still fails, and why the hard part is the booking rather than the conversation.

Commercials · 7 min read

Can agents replace software seats? Sometimes, and not how you think

Where replacing per-seat software with agents genuinely works, where it quietly costs more, and the licensing question to ask before you start.

Architecture · 7 min read

Multi-agent systems: when they help and when they hurt

The cases where splitting work across multiple agents genuinely helps, the failure modes it introduces, and a simpler thing to try first.

Operations · 8 min read

How to tell whether your AI system actually works

Building an evaluation set that concludes: where test cases come from, what to measure, and why systems need re-testing after every model change.

Buying · 8 min read

What an AI solution actually is, and how to tell a real one from a wrapper

The term covers four very different products, priced alike. How to tell model access from a wrapper from an operated system, and what to ask before you sign.

Buying · 9 min read

GPT-6 Astra vs Claude Fable 5.1, read as an operator rather than a scoreboard

Both flagship indices are dead heats and headline prices are identical, so cache rate, context economics and retirement terms decide the build instead.

Evaluation · 8 min read

Why benchmark scores don't predict how a model behaves in production

Published benchmarks are saturated, contaminated and scaffolding-sensitive. What they can still tell you, and the three conditions under which they predict.

Buying · 9 min read

AI agent vs chatbot vs workflow automation: who decides the next step

The distinction is who decides the next step: a person, a flowchart you drew, or the model at runtime. Pick the least autonomy that solves the problem.

Operations · 8 min read

n8n, Make and Zapier: what every comparison table leaves out

Integration counts dominate these comparisons and decide almost nothing. Billing unit, exception handling and token spend are what set the real bill.

Start here

Live in 12 weeks.
Proven before it scales.

We pilot on one entity and validate against your own reports before anything goes wider. See the full deployment approach →