Perspectives · Practice · Operating models

The gap between knowing and doing looks the same on my own systems.

I run a production-grade lab — Kubernetes with GPU scheduling, self-hosted model serving, a fleet of agents with persistent memory, the monitoring stack you'd expect. I hold it to the standards I sell: GitOps, secrets management, zero-trust access, runbooks. And I can report, from direct experience, that every failure pattern I diagnose in enterprise platforms shows up in my own racks first.

The dashboards knew about a degrading disk days before I acted on it. A certificate expired that three separate systems had visibility into. I once rebuilt a cluster and lost a week of operational knowledge that lived nowhere but the old cluster's history — the exact "institutional memory resets" problem I'd been describing to clients that same month. The gap between knowing something and doing something about it doesn't care that I'm one engineer instead of a Fortune 500. It's structural. It appears anywhere signals outnumber attention.

Every failure pattern I diagnose for clients shows up in my own racks first. That's not embarrassing. That's the point of the lab.

What dogfooding actually buys

There's a credibility argument for running your own infrastructure — "I use what I recommend" — but that's the least of it. The real value is that operating a system end to end, alone, makes the structural problems impossible to misattribute. In an enterprise, the insight-to-action gap hides behind teams: the monitoring team saw it, the platform team owned it, the app team felt it, and the retro assigns the miss to a communication problem. In a lab with one operator, there's nobody to route the blame through. If the signal didn't become action, the system's design failed — full stop. You learn what actually closes gaps, because you can't fix anything with a meeting.

What closes them, in my experience: making the desired state declarative, so remediation is a diff instead of a decision; wiring signals to the place work happens rather than to a dashboard nobody patrols; and giving the system memory that survives rebuilds — decisions, incidents, and context stored outside any single cluster's lifespan. Every one of those translates directly to client platforms. None of them came from a whitepaper. They came from being the only person around when the disk finally died.

The honest test

So here's a question I'd put to any team selling or buying intelligence systems: where does your own operation still have the gap? If the answer is "nowhere," the diagnosis hasn't been run honestly. The gap between knowing and doing is never fully closed — it's managed, instrumented, and narrowed. The practitioners worth hiring are the ones who can show you where it still bites them, and what they built to make it bite less. Distrust anyone whose own house is described as finished.

This is the work I do.

Bounded proofs on real data, agent platforms your organization owns, and the operating-model design that makes them stick — delivered end to end, corp-to-corp through Mazo Cloud Group LLC.

Start an engagement