Skip to content

field notes

Insights

War stories from production data work. What broke, what we measured, and what we'd tell you over coffee before you build it yourself.

ai · call centers · quality assurance2026-08-16

Your QA Team Hears 3 Percent of Your Calls

Call centers assure quality by sampling 1 to 3 percent of calls. AI review changes the denominator: every call scored, every score backed by a quoted line.

read →

engineering · data platforms · reliability2026-08-15

Production-Grade Is a Checklist, Not a Compliment

What separates a system a business runs on from a demo that works when everyone behaves: invariants in the database, money in fixed precision, and failure that announces itself.

read →

machine learning · insurance · explainability2026-08-14

We Chose Logistic Regression on Purpose

For an insurance fraud detection system, we picked the boring model: ~0.90 accuracy, ~0.80 AUROC, and every flag explainable to the person who has to act on it.

read →

data quality · analytics · call centers2026-08-10

How a Bad Join Created a 5x Reporting Error

A call center believed carriers rejected 61 percent of its calls. The real number was 11 to 12. The culprit: one join that silently dropped most rows.

read →

machine learning · causal measurement · e-commerce2026-08-09

Measuring Marketing Lift With a Permanent Control Group

Retention tools claim credit for revenue that would have arrived anyway. The fix is old-fashioned science: a permanent randomized holdout that is never emailed.

read →

llm · nlp · hospitality2026-08-08

What It Takes to Label 67,962 Reviews With an LLM

Sentence-level labeling across an 84-aspect ontology, 26 ontology versions, and a model bake-off that kept 95 percent of quality at a quarter of the cost.

read →