Perspectives · Data engineering · AI architecture

Stop asking language models to count.

By most industry estimates, around 80% of what an enterprise knows lives in unstructured form — survey answers, support tickets, emails, call notes, documents. Every one of those organizations has, by now, pointed a language model at that pile and asked it a business question. And every one of them has discovered the same failure, usually in front of an executive: ask an LLM an aggregate question — how many customers who match these three criteria said this specific thing — and you get a confident, fluent, wrong number.

This isn't a model-quality problem, and the next model won't fix it. Language models interpret; they don't compute. Retrieval-augmented generation doesn't fix it either — RAG hands the model relevant passages, but no amount of relevant passages lets a next-token predictor count, filter, and aggregate across a million records. The teams that keep throwing bigger models and fancier retrieval at aggregate questions are solving the wrong layer.

The model interprets. The database counts. Production insight systems live at the seam between the two.

Use the model where meaning lives, the database where math lives

The architecture that actually works is a hybrid, and it's older than it looks: it's an ingestion pipeline wearing an AI costume. Spend the model's judgment once, at ingestion — turning each messy record into typed, named attributes. A free-text answer becomes preferences__likes_red = true; a rambling support ticket becomes a sentiment, a product, a severity. Store those attribute values in a flexible structure (entity–attribute–value works well precisely because the schema keeps growing), with every value linked back to the raw source so nothing loses its provenance.

Then, at query time, use the model three more times — each for a narrow act of interpretation, never for arithmetic. First it reads the natural-language question and picks which attributes matter, so the working schema stays small enough to reason about. Second, against a dynamic view built from just those columns, it writes SQL. The database — the thing that has been good at counting for fifty years — executes it. Third, the model narrates the result back in plain language, with the numbers coming from the engine, not the imagination.

The guardrails are the product

What separates this from a demo is everything around the model calls. Constrained, schema-validated outputs — the model isn't asked to "respond helpfully," it's forced to emit values that type-check against the attribute schema, so an invalid answer is structurally impossible rather than merely discouraged. A retry loop that feeds failures back with context: when generated SQL breaks, the error goes back into the prompt and the system corrects itself instead of paging a human. Scoped, disposable views, so the model only ever sees the handful of columns a question needs — small enough to reason about, cheap enough to throw away. And an audit trail of every prompt, output, and retry, because in any organization worth consulting for, "the AI said so" is not an acceptable lineage for a number in a board deck.

Notice what that list is: type checking, error handling, least privilege, logging. The reliability of an AI insight system comes almost entirely from the traditional engineering wrapped around the model — which is why data engineers, not prompt artisans, are the people who ship these.

This is a platform capability, not a project

Built once, this pattern generalizes across everything unstructured the business holds: the same extract-to-attributes, store-with-provenance, interpret-then-compute loop works for tickets, transcripts, and documents — the only genuinely new work per source is entity resolution, deciding who or what each record is about. That's why I treat it as a platform capability alongside the pipelines it feeds: it composes directly with the RAG, memory, and agent layers in the enterprise AI stack, and it's the layer that lets every one of them answer with numbers instead of vibes. I run this same discipline in my own lab — agents that extract structure from operational noise and store it queryably — because the pattern is the same at every scale: the model finds the meaning, the database does the math, and everything in between is engineered, logged, and owned.

This is the work I do.

Bounded proofs on real data, agent platforms your organization owns, and the operating-model design that makes them stick — delivered end to end, corp-to-corp through Mazo Cloud Group LLC.

Start an engagement