Insights
What we've learned building these systems.
Working notes on the problems that come up in every multi-entity engagement — consolidation, cost, and data isolation. Written for the person who has to make the decision, with the numbers we actually publish.
Operations · 9 min read
How to consolidate reporting across outlets on different POS systems
A method for getting one set of numbers out of outlets on different POS systems: ingestion, reconciliation against close reports, item normalisation.
Commercials · 7 min read
What an AI operating system actually costs to run
The three lines behind the monthly running cost of a bespoke AI operating system — inference, infrastructure, integration — and which grow with outlets.
Architecture · 8 min read
Multi-entity data isolation: a filter is not a boundary
Why filtering by entity_id in application code is not isolation, what database-boundary enforcement looks like, and how group reporting survives.
Adoption · 8 min read
Why most AI agent pilots never reach production
Survey data puts agent pilot failure at 88%. The blockers are evaluation, governance and reliability — and the first one is what actually kills projects.
Commercials · 7 min read
The 2026 model tiers, and how to route work between them
Published August 2026 model prices, and the routing rule that decides a bill: a reasoning tier for judgement, a cheap fast tier for everything at volume.
Architecture · 8 min read
MCP went stateless: what the 2026-07-28 spec changes
The largest MCP revision since launch: a stateless core, versioned extensions and OAuth-aligned authorization. What it changes for connecting POS, ERP and HR.
Operations · 7 min read
Where back-office agents pay back, and where they don't
Invoice processing time falls 70–90% and close accelerates 30–50%, yet finance agents take 8.9 months to pay back against 3.4 for sales. Why, and what to do.
Tooling · 8 min read
The AI tool stack for a multi-entity operator
What to actually use in 2026, layer by layer: model tiers, MCP connectors, agent platforms and the data foundation underneath — plus what not to buy yet.
Buying · 8 min read
The questions to ask an AI vendor before you sign
Twelve questions that separate a vendor who has shipped from one who has demoed — ownership, running cost, failure handling, and life after handover.
Operations · 8 min read
Is that outlet actually profitable, or just busy?
Why per-outlet P&Ls disagree with reality: shared-cost allocation, comparing outlets of different sizes, and separating real growth from more traffic.
Controls · 7 min read
What happens when an AI agent gets it wrong
Agents fail. What decides the damage is reversibility, hold-back queues, run records and which actions a system is never allowed to take unsupervised.
Commercials · 7 min read
Another hire, an outsourcer, or an agent?
Comparing a headcount, an outsourced team and an AI agent by cost shape, supervision load and what happens to the work when volume moves.
Architecture · 8 min read
RAG, fine-tuning or long context: which one your problem needs
The three ways to put your own data in front of a model, what each one actually changes, and the question that decides between them.
Governance · 7 min read
Where your data actually goes when you use an LLM
What happens to business data sent to a model provider — retention, training, residency, sub-processors — and which questions get you a straight answer.
Security · 8 min read
Prompt injection: the attack that has no clean fix
Why prompt injection cannot be patched away, what makes agents with tool access genuinely dangerous, and the architectural controls that work.
Commercials · 7 min read
Build or buy: how to decide on an AI system
A decision rule for AI build-versus-buy that survives contact with reality, and the two failure modes that cost the most on each side.
Governance · 7 min read
AI governance that fits an organisation without a compliance department
How to govern AI use without a compliance function: an inventory, a risk tier, four rules people follow, and the reviews that actually catch things.
Architecture · 8 min read
When a small local model beats a frontier API
The cases where a small self-hosted model genuinely beats a frontier API — and the hidden costs that make self-hosting more expensive than it looks.
Commercials · 7 min read
How to measure whether AI is actually paying back
Why AI ROI measurement usually fails, the baseline nobody captures in time, and a measurement design that survives a finance review.
Architecture · 7 min read
Context engineering: the part that decides whether it works
Why the wording of a prompt matters less than what surrounds it, and the engineering decisions that actually determine an AI system's reliability.
Operations · 7 min read
Voice AI on the phone: what works and what still doesn't
What voice AI genuinely handles on inbound calls, where it still fails, and why the hard part is the booking rather than the conversation.
Commercials · 7 min read
Can agents replace software seats? Sometimes, and not how you think
Where replacing per-seat software with agents genuinely works, where it quietly costs more, and the licensing question to ask before you start.
Architecture · 7 min read
Multi-agent systems: when they help and when they hurt
The cases where splitting work across multiple agents genuinely helps, the failure modes it introduces, and a simpler thing to try first.
Operations · 8 min read
How to tell whether your AI system actually works
Building an evaluation set that concludes: where test cases come from, what to measure, and why systems need re-testing after every model change.
Buying · 8 min read
What an AI solution actually is, and how to tell a real one from a wrapper
The term covers four very different products, priced alike. How to tell model access from a wrapper from an operated system, and what to ask before you sign.
Buying · 9 min read
GPT-6 Astra vs Claude Fable 5.1, read as an operator rather than a scoreboard
Both flagship indices are dead heats and headline prices are identical, so cache rate, context economics and retirement terms decide the build instead.
Evaluation · 8 min read
Why benchmark scores don't predict how a model behaves in production
Published benchmarks are saturated, contaminated and scaffolding-sensitive. What they can still tell you, and the three conditions under which they predict.
Buying · 9 min read
AI agent vs chatbot vs workflow automation: who decides the next step
The distinction is who decides the next step: a person, a flowchart you drew, or the model at runtime. Pick the least autonomy that solves the problem.
Operations · 8 min read
n8n, Make and Zapier: what every comparison table leaves out
Integration counts dominate these comparisons and decide almost nothing. Billing unit, exception handling and token spend are what set the real bill.
Start here
Live in 12 weeks.
Proven before it scales.
We pilot on one entity and validate against your own reports before anything goes wider. See the full deployment approach →