An essay on why AI rollouts fail in firms of 20–100

The knowledge that runs
your company is stored
in about four people.

Not in a system. Not in the manual. In their heads — and it leaves the building every evening. You already know this. What you may not know is why every attempt you've made to write it down has died within a quarter, and what has to change structurally for the next one to survive.

Where it actually lives Today
The shared drive · ~15% of it and most of that is out of date
Each cluster is one person's operating knowledge — the pricing exception they always make, the vendor they never trust, the phrasing that wins. None of it is written anywhere. The band underneath is what you'd find if you opened the shared drive right now.
The reason it keeps failing

It was never a discipline problem

The usual diagnosis is that people are busy or careless, so the usual fixes are a mandate, a template, and a recurring reminder. All three fail, in that order, at every company that tries them.

They fail because the incentive runs backwards. The person who has the knowledge pays the entire cost of writing it down, and the benefit goes to someone else — often someone who hasn't been hired yet. That's not laziness. That's a rational response to a tax with no rebate.

Who pays, and who benefits The broken incentive
Your most experienced person Pays · today · in billable hours writing it down · 40 minutes they don't have cost flows this way A hire who doesn't exist yet Benefits · in fourteen months no system built on this ever survives contact with a busy week
Stated structurally

Any system that requires people to do a second, separate job in order to capture what they learned during the first job will decay to the rate at which people voluntarily do unpaid work. That rate isn't zero — but it's close enough that you should plan as if it were.

So the question isn't how to get people to document more. It's whether the writing-down can stop being a separate act at all.

What changes

Beside the work, or through it

Most companies adopting AI hand people a chat window. Someone copies something out of a system, pastes it in, gets an answer, pastes it back. The assistant never touches the business — it's an articulate stranger you consult on your lunch break. Nothing it learns is kept, because there's nothing connecting it to anything.

Two arrangements Animated
The chat window never connects Copy out. Paste in. Paste back. Nothing is retained. Knowledge captured: none Connected to the work the assistant reads and writes · both directions
The difference looks like a convenience feature and isn't. Being inside the systems is what makes the second direction possible — and the second direction is the whole argument.

Connected, the work happens through the assistant instead of beside it. Not "ask the AI about the proposal" — draft the proposal in the place where the client history, the pricing rules, and the last four proposals already are.

The mechanism

The layer has to run
in both directions

Everyone designs the pull. Ask a question, get an answer grounded in how the company works. That's the demo, that's what gets bought, and on its own it goes stale exactly the way your wiki did.

Nobody designs the push. But when an experienced person does real work through the assistant, they leave evidence — the exception they made, the threshold they applied, the vendor who let them down, the phrasing that landed. That material is produced whether or not anyone intends to produce it. It costs nothing extra, which means there's nothing to skip.

Deposit and extract The loop · animated
The work getting done quotes · proposals · client email · exceptions · handoffs Your operating knowledge the rules, thresholds, exceptions and standards maintained · versioned · owned by someone Deposit what the work reveals automatic · zero extra effort Extract what the person needs grounded in how you work the turn that makes it compound deposits come from many people extractions go to many people and they are not the same people
That last line is the entire value. Your estimator's judgment call gets deposited on Tuesday and shows up in a new hire's draft on Thursday. That transfer used to require the two of them to be in a room together, and it used to only happen by luck.

What your people get

  • Answers grounded in how this company works, not the industry in general
  • Recurring documents produced to your standard instead of from scratch
  • Rules and thresholds applied without having to remember them
  • A new hire operating near month-six competence in week two

What the company gets

  • The exception your best person made, and the reasoning behind it
  • Decision rules that existed only in somebody's head
  • Which vendor, which route, which phrasing actually worked
  • Knowledge that stays when the person doesn't

The second column is what pays for the whole thing, and it's invisible on day one. That's the hard part — and the reason most implementations stop at the first column and quietly decay. How the deposit physically travels is further down; it's simpler than it sounds.

The question every owner asks

Nobody writes
a document

Two assumptions come up immediately, and both are wrong in ways worth being direct about.

It isn't one place everyone works out of. Ask sixty semi-autonomous people to change where they work and adoption is over before it starts. The assistant attaches to the tools they already use — their inbox, their CRM, their files. You meet them where they are.

And it isn't knowledge-transfer assignments. That's the chore, and the chore is the thing that already died. If the plan involves anyone being asked to write up what they know, you've rebuilt the wiki that failed.

Four mechanisms actually move knowledge out of heads. They run at different times and they're worth telling apart.

Four ways in, one gate Nothing enters unreviewed
Review queue everything arrives as a candidate owner + backup 3–5 hrs / week Your operating knowledge
The gate is the part people skip. Every deposit arrives as a proposal, never as truth — because capture is mechanical and deciding what's actually true is judgment. One person confirms, edits or rejects. That queue is the few hours a week, and it's why nobody else needs write access.

1 · Structured interviews

Front-loaded. Where most of the first version comes from.

There's a technique to this and it's the difference between weeks of work and nothing. Experts genuinely cannot tell you what they know — the rule became automatic years ago, so they no longer experience it as a rule. Ask "what are our pricing rules" and you get platitudes everyone already agrees with.

Ask instead: "walk me through the last four quotes you sent." The rules fall out of the cases, including the ones they'd have sworn they didn't have. Cases, never abstractions. That's the whole trick, and it's why this is real labor rather than a form.

2 · Mining what already exists

Front-loaded. A shortcut, not a source.

The last fifty proposals. Two hundred client emails on one topic. The patterns are already in there, and reading them produces a draft of the rules in a fraction of the time it takes to ask.

The critical qualifier: this generates candidates. Old files contain the current policy and the one from 2019 that nobody deleted, and nothing in the folder says which is which. Every mined rule goes back to a human for confirmation. Skip that step and you've automated the propagation of stale truth.

3 · The moment of correction

Ongoing. This is the one that makes it compound.

Someone asks for a draft, gets something ninety percent right, and fixes the last ten. That correction is the deposit — precise, attached to a real case, and free, because they were fixing it anyway.

The agent asks one question about the change, in the same conversation, and a sentence answers it. No document, no form, no Friday roundup. This is the answer to "what keeps it alive after you leave."

4 · The exception moment

Ongoing. Targeted at the highest-value moments.

When somebody escalates, approves a deviation, waives something, or overrides the default — that is an undocumented rule surfacing in real time. It's the densest knowledge your company produces and it's almost never captured, because by the time anyone thinks to ask, the reasoning has evaporated.

You ask in the moment, while the reasoning is still loaded, or you don't get it at all.

The governing rule

Capture is automatic; acceptance never is. Every one of the four routes above ends in the same queue, and a named reviewer decides what becomes true. That distinction is what separates a layer that gets better from one that fills up with sixty contradictory opinions — which is a different failure than going stale, and not a better one.

What the third one actually looks like

People find this one least intuitive, so here's a single instance end to end. Two things worth knowing before you read it.

It's a conversation, not an interface. There's no special screen watching someone edit and popping up a prompt. The agent asks, in the same chat where the work is already happening, because its standing instructions tell it to.

Which means it captures what people tell it, not what they do elsewhere. A correction made half an hour later in Word or Outlook is invisible to it. The deposit comes out of the back-and-forth of getting the work right — which is also, conveniently, where the reasoning actually lives.

The correction moment Step through it
Nobody opened a wiki, wrote a paragraph, or was reminded to document anything. A rule that lived in one person's head for six years is now queued for review — and the person who held it typed nine words she was going to type anyway to get her letter right.
And it won't fire every time

The ask is an instruction, not a mechanism — the agent will usually remember, not always. That's acceptable here and unacceptable elsewhere, and the difference is worth being precise about, because it's the line most AI pitches blur.

Three kinds of obligation

Not everything a system does carries the same duty. Sorting the work into these three categories before anything gets built is what decides where an instruction is sufficient and where you need a mechanism.

Discovery

Probabilistic capture is fine.

Noticing a rule, catching an exception, learning which vendor is unreliable. Missing one costs you nothing permanent because rules recur — the same correction surfaces again next week. This is where the deposit loop lives.

Operating guidance

Grounded generation, then a human reads it.

Drafts, summaries, comparisons, client correspondence. The layer makes these accurate far more often, and a person still signs off. Most of the day-to-day value sits here, and so does most of the time saved.

Control

Deterministic check plus required approval.

Regulatory disclosures, prohibited language, mandatory approvals, records retention. A single miss is the whole problem, so these cannot rest on the model choosing to comply. They need software that runs regardless, and a named person who signs.

The deposit loop is squarely in the first column. Nothing in the third column should ever be described in the same language — and if anyone tells you their assistant "checks for" a compliance requirement, the question worth asking is whether that check is a mechanism or a request.

Why "we already tried this" is true and still not an argument

Twenty-four months,
three arrangements

A manual written once is at its most accurate the day it's finished. That's the trap, and it's why the objection is usually delivered as proof that this doesn't work.

Below is the same two years under three arrangements: no written layer at all, a layer written once and maintained by hand, and a layer fed by the work as it happens. What's being measured is the share of real questions it answers well enough that somebody can act without checking first.

How much of it is still true Illustrative · draws on scroll
Fed by the work Written once, maintained by hand No written layer % of questions answered well enough to act on 100 50 0 Crossover around month nine Month 0 Month 12 Month 24
M3 · price change M5 · someone leaves M11 · two new hires M18 · whoever wrote it moves on
Read the pink line first. It starts highest — a freshly written manual really is accurate — and it's overtaken before month nine as ordinary change piles up underneath it. The teal line starts lower because it hasn't seen much work yet, and climbs for exactly the same reason it started low. Grey is what most companies actually have: not wrong so much as never specific enough to act on. These are illustrative shapes, not measured data — the argument is the direction of each line, not the numbers.
Design

One shared layer,
many personal ones

Plenty of companies this size don't run on employees at all. Licensed agents, producers, advisors, contractors — people with their own books of business who can't simply be told what tools to use. It changes the design in a way that's easy to get wrong and expensive to get wrong.

You cannot quietly absorb someone's client relationships into a company asset. Beyond whatever their agreement says, the moment people suspect it's happening they stop being candid — and candor is the only input the deposit loop has.

The split, and the boundary between Structure
Personal · private · thin Promotion gate on purpose · with their knowledge · never silently The company layer — shared and governed how the firm works · compliance rules · service standards · vendors · procedures everyone reads · changes go through the reviewer
Sixty people with edit access to shared ground truth produces sixty versions of the truth inside a quarter, which is the same as having none. Read widely, write narrowly.

The company layer

Shared and governed. How the firm works, what the rules are, the standards, the vendors, the procedures everyone follows. People read from it; changes route through a named reviewer.

The personal layer

Theirs. Their clients, their territory, their voice, their live work. Private by default, and deliberately thin — every minute someone spends maintaining it is a minute they resent.

The boundary

Things move upward only on purpose and with the person's knowledge. This is the rule that makes the whole arrangement survivable, and it's worth writing down before anything gets built.

How it's actually wired

Almost nobody touches
the machinery

Two constraints shape the whole build. Nobody producing work should ever see a repository or learn a developer tool. And nobody producing work should ever be able to edit shared truth directly, because sixty people with write access to the same files produces sixty versions of the truth inside a quarter.

Worth being precise here, because it's the difference between a policy and a fact: the second constraint isn't something we promise to be careful about — the shared library simply cannot be edited by the people it's delivered to. And where something genuinely must be present on every machine, it can be made impossible to remove. Those are properties of how the tooling is distributed, not house rules somebody has to enforce.

Both are satisfied by a single rule: producers only ever create new files. They never edit existing ones. Everything else follows from that.

two halves, and only one of them is a folder you have to set up brokerage-knowledge/ │ ├── wiki/ ← delivered to everyone · nobody can edit it │ ├── advertising/ │ │ ├── adre-disclosure.md │ │ └── fair-housing-language.md │ ├── transaction/ │ ├── vendors/ │ └── listing-voice.md │ └── inbox/ ← a shared folder · one new file per deposit ├── 2026-08-02-dana-expedite-fee.md ├── 2026-08-02-marco-hoa-signage.md └── 2026-08-03-dana-photo-turnaround.md

New files never collide with each other, which quietly removes every problem you'd otherwise have with sixty people writing at once. No locking, no conflicts, no last-write-wins data loss. The append-only inbox is the entire trick.

Append-only applies to the inbox, not to the knowledge itself. Worth stating plainly, because it's an easy thing to over-apply. The wiki has to be fully editable by the reviewer — rules get corrected, superseded, and sometimes deleted outright, and a body of knowledge that can only ever be added to stops being an operating manual and becomes an archive. Version control gives you the history; the current page still has to say what's true today and nothing else.

The round trip of one deposit Animated
Sixty producers a desktop agent, in the tools they already use no git · no accounts · ever wiki/ the published layer delivered to every machine cannot be edited by producers inbox/ proposals, one file each a shared folder, outside the record append-only owner + backup reviews · decides commits · publishes 3–5 hrs/wk read one new file per correction — never an edit accepted changes publish back to everyone
A producer reads from the left half and drops proposals into the right half. That is the entire extent of their involvement — everything past the inbox is invisible to them and always will be. The left half arrives on its own; the right half is the one piece of plumbing that has to be set up.

A producer corrects something

Mid-task, in the tools they already use. Their agent writes a small proposal file into the inbox — what the rule seems to be, the case it came from, which page it would change, who and when. The producer clicked once.

The owner works the inbox

Not by hand-editing files. They read the proposals and make decisions — accept, revise, reject. Their agent does the mechanics: applying accepted rules to the right pages, updating cross-references, writing one clean, attributed commit per proposal.

Leadership reviews at the change level

Because each accepted proposal became its own coherent commit, there's a real diff to look at — what changed, why, sourced to a specific case. That's a review anyone can do. A pile of loose edits in a shared folder is not.

It reaches everyone, automatically

Accepting the change is the release — there is no separate step where somebody remembers to distribute it. Every producer has the new version by their next task. A rule that lived in one person's head is now applied by sixty people, and nobody attended a meeting about it.

Rejected proposals stay out of the record

Because the inbox sits outside the repository, only reviewed and accepted material is ever committed. A proposal that contained a client's details, or was simply wrong, is deleted rather than preserved — it never enters the permanent knowledge record. Version history is excellent at remembering and very bad at forgetting, so keeping the gate in front of it rather than behind it is the whole point.

To be precise, since this is a privacy claim: that means no permanent trace in the knowledge repository. Ordinary application and security logs are a separate matter, governed by your existing retention policy, and are part of the data-flow document rather than something this design changes.

Two roles, two different tools

Producers get a desktop agent with the firm's library already installed — no repository, no accounts, nothing to set up, and nothing they could break. The reviewers get a developer-grade agent that can read history, write commits and manage the repository. Same knowledge, two entirely different surfaces, matched to who's actually sitting there.

Two people hold that second role, not one. A primary who works the queue and a backup who can cover a holiday, an illness, or a resignation. One reviewer is a single point of failure sitting directly on top of the thing you paid to build.

You don't need any of this in month one

The first version of the inbox is a person noticing something and writing it down, the way your operations lead already handles everything else. Automating it is worth doing once the volume proves it's worth doing — which is a number you get from running a small pilot, not from guessing in advance. Building the machinery first is how implementations end up with excellent plumbing and nothing flowing through it.

Made concrete

What it looks like
in a real company

Take a firm of sixty licensed agents in a single city — a brokerage, though the shape is identical at an insurance agency, a small law firm, an accounting practice, or an agency of any kind. Semi-autonomous producers, regulated written output, and a handful of people who know everything.

0
producers
Each with their own clients, their own habits, and their own private version of the rules.
0
people who know
The broker, the transaction coordinator, and two agents who've been there a decade.
0
of it written down
There's a policy manual. It covers what the state requires and almost nothing about how the work is actually done.

Goes in the company layer

  • Advertising and disclosure rules the state actually enforces
  • Language that must never appear in a listing, and why
  • The firm's presentation standard and quality bar
  • Which inspector, lender and title company to send people to
  • Neighborhood-level knowledge — the well and septic questions, the HOA that always delays
  • What happens at each milestone of a transaction, and who's responsible

Stays in the personal layer

  • Their farm area and past client roster
  • Live listings and buyers they're working
  • How they write — the voice their clients recognize
  • Their niche, whether that's first-time buyers or investment property
The part that surprises owners

In a firm like this the productivity gain mostly belongs to the producers, not the company — they're independent, and they keep what they earn. So the reason the firm pays isn't speed. It's that the knowledge stops walking out the door, quality stops varying by whoever picked up the phone, and there's a defensible answer when someone asks why a document said what it said. Those are owner problems, and no amount of individual productivity fixes them.

brokerages insurance agencies small law firms accounting practices wealth management agencies staffing firms specialty contractors
The four that come up every time

Questions owners ask
before anything else

Every one of these should be settled in writing before a single system is connected — whoever you end up doing this with. The answers below are the ones I'd argue for; the point is less that you agree with them than that somebody has committed to an answer before the work starts.

Where does our information go?

This is the right first question, particularly if you're in a regulated field where client information carries obligations beyond your own preferences.

The principle: your operating knowledge is plain readable text, in a repository your company owns, that stays legible if every tool involved disappears tomorrow. What gets connected, what's excluded, retention, and whether anything can be used to train anything are settled explicitly in the agreement before connection — not assumed, and not left to a default setting somebody didn't read.

Being exact about one thing a technical reader will ask: the knowledge is yours and sits in your repository, but delivering it to sixty people and connecting to your other systems does route through the vendor's infrastructure rather than your local network. That's a normal arrangement and it is not the same as "nothing leaves the building" — so it belongs in the data-flow document, in writing, rather than in a reassuring sentence.

If you have an existing security review or a compliance officer, they belong in that conversation from the start rather than at the end.

What do we own if we stop?

All of it, unconditionally. The extracted knowledge is the company's own knowledge — nobody helping you build it has a defensible claim to it, and anyone who wants one is telling you something.

There's a practical reason beyond the ethical one. The entire input to this process is people being candid about how they actually work. Any ambiguity about who ends up owning that shuts the candor down, and the candor is the whole product.

Plain text, your repository, readable without anyone's tooling. If it stops, you keep a documented company.

What happens when it's wrong?

It will be. The relevant question isn't whether an AI system makes mistakes — it's whether yours are visible and fixable, or invisible and repeated.

A grounded layer changes the failure mode. When an answer is wrong you can see which written rule produced it, correct that rule once, and it's right for everyone from then on. Compare that to a chat window, where the same mistake is regenerated fresh every time and there's nothing to fix.

Constraints on prohibited language and required disclosures reduce risk substantially. They do not eliminate it, and they don't replace your review. Anyone telling you otherwise is selling you something.

Do I have to make people use it?

If you have to, it isn't working — and with independent producers you couldn't anyway.

Adoption is the honest measure of whether this is worth continuing, which is why it's tracked from the first week rather than reported at the end. Every design decision that adds friction for the person doing the work is a tax paid against it.

The pattern that works: start with a small mixed group including at least one person everyone else watches. If they use it visibly, the rest follow. If they don't, that's the finding, and it's cheaper to learn it in week three than in month nine.

Before you get excited

What this doesn't solve

Worth saying plainly, because the version of this argument that skips it is a sales pitch rather than a description.

Capture isn't curation

Recording what happened is mechanical. Deciding which of two contradictory things is now true is judgment, and no amount of automatic capture makes it. Somebody at the company has to own that — three to five hours a week, usually the person everyone already asks when they don't know the rule. Skip it and the layer fills with contradictions instead of going stale. Different failure, not a better one.

The first pass is real work

Getting what's in your people's heads into usable form is interview work — sitting with the experienced ones and asking the questions that surface rules they don't know they're following. Weeks, not days. Anyone describing this as a software install is selling you the demo.

When this isn't worth attempting

Your core systems can't be connected to — everything downstream degrades back to a chat window. Or nobody can be freed for a few hours a week to own what's true. Or the goal is to reduce headcount, in which case the people whose knowledge you need will read the room correctly and tell you nothing useful.

Where this ends up

Start small, or not at all

The conclusion I keep arriving at is that nobody should begin with a rollout. The first thing worth knowing is whether there's anything in your people's heads worth extracting — and the only way to find out is to extract some of it. A handful of people across experience levels, three workflows they do every week, a few weeks of work.

That's also the honest test, and it's why I'd argue for it even though it makes the idea look smaller. If a real sample doesn't produce something the team recognises as valuable, the full version won't either — and you've spent a month finding that out instead of a year.

Most of this comes from doing the same work at much larger scale, where the failure modes are identical and only the budgets differ. I write these up mainly to think them through properly. If you're wrestling with the same problem I'd like to hear how it's going — including the parts where you think I've got it wrong.

Read the long version → More notes →
Neal Meinke

Senior Lead Application Developer, working on enterprise AI assistants at Fortune 500 scale. Ten years in conversational AI, six of them at enterprise scale. This piece is that same method thought through for companies where one person still knows everything — a problem I find more interesting than the size of the company suggests.

These notes are how I work things out. Comments and disagreement welcome.

Email · Notes · Back to the site · LinkedIn