- The measurable half becomes the scoreboard. I read it as the whole.
- My AI refused to fix the contradiction. That refusal is the only reason I found it.
- A stated risk with no instrument is a risk you have quietly agreed to stop watching.
I don’t run a law practice. I build the AI operating systems that run other people’s — AI-enablement for legal, full stop. Which means the systems I put under real pressure are my own, and the failures I find in them are mine to report. This one is the most useful mistake I have made in months, and it is not a technical mistake at all. It is a measurement mistake, of exactly the kind a practice makes about itself. Notes nineteen through twenty-one.
19. I graded the half that was cheap to measure
I keep a written decision log, because a decision you cannot reconstruct is a decision you will make again. Earlier this summer I recorded a change of direction on a major piece of work, and I wrote down exactly why: the experiment could not test its own riskiest assumption. I even named the assumption. It was about writing — whether a switchover could run under real daily use without losing a single piece of work. Some time later I recorded the opposite verdict on the same project: keep going. That second decision graded the rollout question — did the two installations move across cleanly — and every check passed, so the project passed. The following morning the system made sixteen small corrections, confirmed each one, and then watched a routine refresh silently undo all sixteen. The assumption I had written down as the one that mattered was still false, and had been the whole time.
What happened is worth naming plainly: I graded the assumption that was cheap to measure instead of the one that could kill the project. Not through laziness. The rollout question had checks, numbers, and a clean green readout, because I had built the instrument for it. The write-correctness question had none of those, because I never built one. So the scoreboard showed the half that could be scored, and I read the scoreboard as the state of the project. This is the textbook version of the mistake, and knowing its name did not save me from it — it was sitting in my own log, in my own words, the entire time.
20. My AI refused to fix it, and that is why I found out
The reason I found out at all is almost absurd. A routine check flagged one leftover contradiction and described it as a citation problem: a document citing a decision that a later decision had displaced. Bureaucratic. Easy to wave off. I nearly did. Underneath that dull description was the whole thing above — two recorded decisions that could not both be true. The checker surfaced it and then stopped, because its rule is that it acts only where the answer is mechanically derivable — a date comparison, a file check, a lookup — and declines anything that needs a written judgment about what I meant.
If it had been helpful, none of this would have surfaced. A reasonable assistant would have quietly reconciled the two records, the report would have gone green, and I would have carried on with a false picture. So the refusal is not a limitation to be engineered away; it is the routing. It is how a question reaches a human at the moment human judgment is actually required. Four properties make it work: the problem is detected mechanically and never by an opinion about what my prose meant; it surfaces before I have formed the first view of the day on top of the bad record; it refuses to invent the resolution; and — the one that carries the others — it acts only where the answer is genuinely derivable. Drop that last one and the first three buy you nothing, because an assistant that helpfully resolves ambiguity converts an open question into a confident wrong answer.
21. A stated risk with no instrument is a risk you have stopped watching
Take this out of software and into a practice, because that is where it actually costs money. Nearly every firm measures the client-satisfaction survey. Almost none measures the thing that actually loses clients — the call that went unreturned for nine days, the matter that sat untouched while something louder got attention, the client who never said a word and simply did not come back. The survey is not measured because it matters most. It is measured because it is the thing that can be measured, and over time the measurable half quietly becomes the whole scoreboard.
The discipline I would hand any solo is boring and costs nothing: when you write down the thing that would sink this, write down how you would know — in the same sitting. Not later, not in the abstract. One sentence naming what you would have to look at, and how often, to catch it going wrong. If you cannot write that sentence, you have not identified a risk you are managing; you have identified one you are hoping about. And be suspicious of any assistant, human or machine, that smooths over every inconsistency it finds — it is removing the only signal that was ever going to reach you. Prefer the one that says: these two things disagree, and I am not going to guess which is right.
These notes aren’t a practice diary — they’re a record of what I keep finding while building AI systems for legal work, written for the solo and small-firm attorneys those systems are for. This one you can act on tonight: open whatever you use to track your practice, find the risk you would call most serious, and see whether anything in your week would actually tell you it was going wrong.
Field Notes are written for licensed attorneys and are not legal advice. Mike Moss is a Utah-admitted attorney doing AI-enablement work — not operating a law practice.