AI agents perform best when your engineering docs work like durable async handoffs, not tribal memory.
01 THE PROBLEM
Async engineering documentation is the practice of writing decisions, constraints, interfaces, and operating context so clearly that another person can continue the work without a meeting. The failure mode is simple: AI agents are being asked to act like teammates in systems that were never documented well enough for remote human teammates.
That gap shows up fast.
A coding agent opens your repo, sees 1,800 files, three half-maintained READMEs, a stale architecture diagram, and a Notion page last updated before your last two reorgs. It can still generate code. It just cannot reliably generate the right code.
This is not a model quality problem first. It is a context quality problem first.
When leaders say, “The agent is inconsistent,” what they often mean is: the system around the agent is inconsistent. Naming conventions drift. Ownership is implicit. Tool contracts live in Slack. Business rules are buried in ticket comments. Runbooks explain what to click, not why the edge case exists. The agent then does exactly what a new engineer would do in the same environment: infer, guess, and occasionally break something expensive.
The timeline is measured in weeks, not quarters.
In the first two weeks of rolling out coding agents, teams usually see a burst of local productivity: tests added faster, boilerplate generated, refactors proposed. By weeks four to eight, a different pattern emerges. The easy wins are gone. The agent starts touching workflows where the missing context matters: auth, billing, provisioning, data contracts, privacy, and incident-prone paths. That is where vague documentation becomes operational risk.
For a 20–200 person company, the consequence is not abstract. It lands as review fatigue, longer validation cycles, rising revert rates, and staff engineers becoming human context routers. Instead of removing interruptions, the agent creates a new class of interruption: “Can someone verify whether this is how we actually do tenant isolation in the EU path?”
Human teams already know this pattern.
GitLab built a large part of its operating model around handbook-first, asynchronous communication precisely because distributed execution breaks when knowledge is trapped in meetings or people’s heads. Stripe’s engineering culture has long emphasized clear API contracts and operational rigor because systems scale poorly when assumptions stay implicit. AI agents intensify the same requirement. They do not “figure out” missing intent. They fill in missing intent with probability.
That is the core problem: AI agents need the kind of documentation habits that strong async engineering teams already rely on, because both are solving the same handoff problem. One handoff is between humans across time zones. The other is between humans and machines across incomplete context windows.
If you would not trust a new senior engineer to ship from your current docs without three onboarding calls, you should not trust an agent either.
02 WHY IT HAPPENS
The structural reason is that most engineering documentation was designed for memory refresh, not execution.
That distinction matters.
A memory-refresh doc assumes the reader already knows the codebase, the company, and the unwritten rules. It exists to jog recall: service names, rough architecture, maybe a setup command or two. It helps insiders remember. It does not help a new actor act.
Execution-grade documentation is different. It tells a reader what the system is for, where boundaries are, what invariants cannot be broken, which edge cases already cost pain, and how to verify a change safely. It is the difference between “here’s our billing service” and “this service may never retry card-capture calls without an idempotency key because duplicate capture caused customer-impacting incidents in 2023.”
Most teams have far less execution-grade documentation than they think.
There are three structural reasons.
First, incentives favor shipping code over preserving context.
DORA’s research, published through Google Cloud’s State of DevOps reporting, consistently ties documentation, developer experience, and internal knowledge accessibility to software delivery performance. But few teams staff or measure documentation with the same seriousness they apply to CI reliability or infrastructure costs. Engineers get rewarded for merged code, not for reducing ambiguity for future readers. Documentation becomes a cleanup task, then a neglected asset class.
Second, synchronous culture masks documentation weakness.
A team can function surprisingly well with mediocre docs when the default operating mode is high-bandwidth conversation. Sit nearby, ask in Slack, jump on Zoom, grab the staff engineer. The organization becomes a living retrieval system. This works until scale, turnover, distribution, or urgency breaks the loop.
Linear is a useful point of comparison here. The company has spoken publicly about valuing sharp product and engineering communication, short feedback loops, and written clarity as part of staying fast with a lean team. Lean teams survive by reducing coordination drag. Agents create the same pressure. Every undocumented decision increases retrieval cost.
Third, AI tooling creates false confidence.
Modern coding agents are strong enough to succeed on shallow tasks, which leads leaders to underestimate how brittle they become under weak context. A model that can write a migration script or add tests to a utility package appears capable of feature work. Then it reaches areas where hidden assumptions matter and quietly degrades.
This mirrors what experienced engineers know about onboarding. New hires often look productive in week one because setup tasks and isolated tickets are tractable. The real signal appears when they touch shared abstractions, legacy behavior, or politically sensitive workflows. AI agents hit the same wall faster because they have no lived memory and no institutional caution unless you encode it.
There is also an architectural reason specific to agents: context windows are finite, retrieval quality is uneven, and repositories are noisy.
Even large-context models still require selection. The system has to decide what to read. If your repo and docs do not make authoritative context legible, the retrieval layer surfaces whatever is adjacent, popular, or textually similar. That is not the same as surfacing what is normatively correct.
A stale ADR can outrank a current convention if your system never marked it obsolete. A verbose design doc can dominate over the one-page runbook that actually captures current production safeguards. A generated SDK file can distract retrieval from the hand-maintained API contract. The result is not random failure. It is deterministic misguidance.
Cloudflare’s engineering writing is instructive because they routinely document architecture, edge behavior, operational mechanics, and product constraints in precise, concrete terms. That style exists for humans, but it also happens to be agent-readable. The docs encode decision boundaries, not just descriptions. That is what agents need.
The deeper point is this: good async documentation and good AI context are the same design problem.
Both require you to answer five questions with precision:
- What is the source of truth?
- What can change freely, and what is invariant?
- Who owns this boundary?
- What failure has happened before?
- How should someone validate the next change?
If your documentation does not answer those for a human working asynchronously, it will not answer them for an agent either.
03 WHAT MOST GET WRONG
The most common misdiagnosis is thinking the solution is “more documentation.”
It is not.
Most organizations already have enough text. What they lack is authoritative, scoped, maintained documentation attached to real workflows. Adding more pages to Confluence or Notion usually increases ambiguity because it multiplies possible truths.
The second common mistake is treating AI agent documentation as a separate category.
You see this in teams that start writing “agent instructions” before fixing the underlying engineering docs. They create prompt files, coding rules, tool descriptions, and model-specific hints. That can help at the margin. But if your source material is weak, you are just adding a nicer wrapper around uncertainty.
The agent then follows beautifully written instructions into a swamp.
A third mistake is overinvesting in repository summaries and underinvesting in interfaces and constraints.
Repo maps are useful. README cleanups are useful. But agents usually fail at the seams: API behavior, state transitions, migration rules, security boundaries, and operator expectations. That is where the hidden cost sits. A polished top-level overview does not compensate for an undocumented “never call this endpoint in parallel” invariant.
A fourth mistake is assuming code is enough.
This is the oldest engineer instinct and still one of the most expensive. “The code is the documentation” only works when intent is obvious from implementation and consequences of being wrong are low. In distributed systems, payments, permissions, and data pipelines, intent is not obvious. Code tells you what the system does. It often does not tell you what it must never do, what legacy behavior is preserved for contractual reasons, or which ugly branch exists because a major customer depends on it.
GitHub’s own engineering and product work around Copilot has underscored a version of this reality: AI-assisted development is strongest when grounded in repository context, developer workflow, and review systems. Suggestion quality without surrounding engineering discipline is not enough. The tool assists; it does not replace source-of-truth engineering practices.
The fifth mistake is trying to solve this with a giant one-time documentation push.
This fails for the same reason massive test rewrites fail. The documentation drifts before the cleanup finishes. Nobody owns freshness. The team experiences docs as a periodic pain project, not part of delivery. Three months later, the new “AI docs initiative” looks a lot like the old wiki graveyard.
There is a recognizable failure pattern here from incident management.
The Google SRE Book is clear that reliability comes from engineered systems, not heroics. Postmortems, runbooks, and service ownership are valuable because they reduce repeated cognitive load under pressure. Teams that skip those artifacts become dependent on experienced operators. AI agents create a similar dependency trap: if only two senior engineers can tell whether generated work is safe, the organization did not gain leverage. It moved work into a more fragile review bottleneck.
A concrete analogy is the Knight Capital trading incident in 2012. The root cause was not “documentation quality” in isolation, but it remains a canonical example of operational ambiguity around software deployment, old code paths, and insufficiently controlled change process. The lesson for agent adoption is direct: when implicit system state and hidden behavior are allowed to persist, automation amplifies the blast radius.
What most teams get wrong, then, is not effort. It is target selection.
They document what is easy to write, not what is expensive to misunderstand.
04 THE FRAMEWORK
The approach that works is to treat documentation for AI agents the same way strong asynchronous teams treat documentation for distributed execution: as an operational interface, not a knowledge archive.
That means building a layered system with clear ownership, explicit freshness rules, and narrow formats tied to real decisions.
Here is the framework.
1. Start with failure-prone surfaces, not broad coverage
Do not begin by “documenting the whole platform.” Start with the places where wrong changes are costly.
For most Series A–C software companies, that list is usually six surfaces:
- Authentication and authorization
- Billing and money movement
- Data model and migration rules
- External API contracts
- Provisioning and environment setup
- Incident-prone production workflows
These are where agents produce the highest downside when context is incomplete.
Create a ranked map of these surfaces using two inputs:
- incident history from the last 6–12 months
- review hotspots where staff engineers repeatedly intervene
If a subsystem generated more than three significant review escalations or at least one Sev-2 or worse incident in the last two quarters, it belongs near the front of the queue. That threshold is practical because it identifies systems where ambiguity is already expensive.
This is the same prioritization logic high-performing platform teams use for reliability work: go where cognitive load and error rates are concentrated.
2. Write execution docs, not descriptive docs
Each high-risk surface should have one execution-grade doc with the same five sections:
- Purpose: what this system is responsible for
- Invariants: what must remain true after any change
- Failure history: what has gone wrong before
- Change protocol: how to safely modify it
- Validation: what evidence proves the change is safe
That format forces the most useful context into the open.
Example: an auth service doc should not just list components. It should say things like:
- Tenant membership checks must occur server-side before resource access.
- JWT parsing alone is insufficient for authorization.
- SCIM group sync can lag by up to N minutes; changes must tolerate stale group state.
- Never rely on client role claims for admin workflows.
- Required verification includes regression tests on role downgrade, token refresh, and cross-tenant access attempts.
This is what async-ready teams do instinctively. They encode decision boundaries.
Stripe is a useful mental model. Their public engineering writing on APIs, idempotency, and reliability consistently emphasizes explicit contract behavior. Agents benefit from the same clarity because they need normative guidance, not just code visibility.
3. Move source-of-truth docs into version control
If the docs govern how code should change, they should live where code changes happen.
That does not mean every piece of documentation belongs in the repo. Company strategy docs, broad product specs, and research notes can stay in Notion or a handbook. But anything that defines engineering behavior for a service, tool, or boundary should be versioned next to the implementation or in a docs repo with review rules tied to the same ownership model.
This matters for three reasons.
First, freshness improves because documentation can change in the same pull request as code.
Second, review improves because domain owners can inspect whether the change updates both implementation and context.
Third, retrieval improves because agents can access docs as part of the codebase context instead of relying on a brittle bridge into another system.
HashiCorp’s long-standing investment in docs-as-product is relevant here. Their tools succeed partly because operational behavior is documented with precision, versioning, and strong information architecture. The principle scales down: if a change affects runtime behavior, its documentation should be change-managed with similar discipline.
A simple rule works well:
- if violating the doc could cause a bug, security issue, or failed deploy, keep it in version control
4. Adopt three doc types only
Most teams create too many formats. Limit the system to three durable document types.
A. System Contract
One page per service or domain boundary. Contents:- purpose
- inputs/outputs
- invariants
- dependencies
- owner
- related runbooks
- last materially reviewed date
B. Change Guide
A short operational doc for common modification paths. Contents:- typical tasks
- common edge cases
- test expectations
- rollback notes
- anti-patterns
C. Incident Memory
A distilled postmortem appendix focused on future execution. Contents:- trigger
- hidden assumption that failed
- guardrail added
- what future changes must account for
This is enough.
Anything else tends to become a narrative layer that nobody maintains.
Netflix’s engineering culture around paved roads and operational clarity illustrates the broader principle: standardize the interfaces engineers use repeatedly, because free-form process does not scale. Agents benefit from the same standardization because predictable formats are easier to retrieve and follow.
5. Define freshness with an explicit SLA
Most doc systems fail because “keep it up to date” is not an operating rule.
Set a freshness policy.
A practical one for growth-stage companies:
- Tier 1 docs: reviewed every 90 days
- Tier 2 docs: reviewed every 180 days
- incident memory: updated within 5 business days of postmortem completion
Tier 1 should include any system touching auth, billing, customer data, infrastructure provisioning, or public APIs.
This is not arbitrary. Ninety days is short enough to catch ownership drift and product changes before the docs become fiction, but long enough not to overwhelm leads. It also aligns with the cadence many engineering orgs already use for roadmap and OKR reviews.
Track freshness visibly.
A stale doc badge is better than a silent stale doc. Silent staleness is what hurts agents most because they cannot infer whether a page is obsolete unless you say so.
6. Add machine-readable context where it matters
Natural-language docs are necessary but insufficient for agents in some areas.
For critical systems, add a small amount of structured metadata:
- owner team
- criticality tier
- service dependencies
- last reviewed date
- approved deployment path
- prohibited operations
- required tests before merge
This can live in YAML, frontmatter, CODEOWNERS-linked metadata, or an internal registry.
The point is not to over-formalize. The point is to improve retrieval and policy enforcement.
For example:
- if a file under `/billing/` changes, the agent should be able to discover that integration tests and finance reconciliation checks are mandatory
- if a migration touches PII tables, the agent should surface privacy review requirements
- if a service is Tier 1, generated changes should trigger stricter review routing
Cloudflare, GitHub, and Shopify all operate at scales where explicit ownership, service metadata, and deployment controls are not optional. You do not need their size to benefit from the same discipline. You only need enough complexity that hidden rules are slowing review or creating defects.
7. Tie docs to review gates
Documentation only becomes real when the delivery system enforces it.
Add lightweight review checks:
- PR template asks whether a system contract or change guide should be updated
- CODEOWNERS routes doc changes to domain maintainers
- CI validates required metadata fields on Tier 1 docs
- release checklist references affected execution docs
- postmortem template includes “which docs should change?”
This is where most initiatives either become habit or die.
The goal is not bureaucracy. The goal is reducing avoidable ambiguity at merge time instead of after deployment.
A useful benchmark here comes from DORA’s four key metrics: deployment frequency, lead time for changes, change failure rate, and time to restore service. Better docs should improve at least two of these within one or two quarters if the intervention is working. Specifically:
- lead time should fall because reviewers spend less time reconstructing system context
- change failure rate should fall because invariants are explicit before merge
If your documentation work does not improve a delivery metric within six months, it is probably too detached from actual engineering flow.
8. Document “why not,” not just “how to”
This is the most undervalued habit for agent-readiness.
Agents are competent at finding plausible implementation paths. What they cannot infer reliably is why certain plausible paths are forbidden.
Every high-risk doc should include anti-patterns:
- do not batch these writes because ordering matters for downstream reconciliation
- do not retry this webhook blindly because the upstream provider is not idempotent
- do not denormalize this field without updating the analytics export contract
- do not add customer-facing flags here because this config is cached for 30 minutes at the edge
Figma’s engineering work has often highlighted the complexity of multiplayer systems and performance-sensitive architecture. In systems like that, anti-pattern knowledge is often more valuable than happy-path explanation. The same is true for AI agents. Negative constraints prevent the most expensive category of “looks right, breaks reality” output.
9. Design for narrow autonomy
The winning pattern is not “let the agent roam the whole codebase.”
The winning pattern is scoped autonomy with strong local context.
Give agents:
- a bounded task
- the relevant contract docs
- the nearest change guide
- the last incident memory for that surface
- a defined validation routine
That setup makes the agent look much smarter because the environment is smarter.
This is the same lesson platform teams learned from paved-road design. Engineers move faster when the safe path is legible. Agents do too.
There is a tradeoff here.
Broader autonomy reduces human setup time but increases validation cost and tail risk. Narrow autonomy requires better documentation and task framing up front but produces changes that are easier to trust. In regulated or reliability-sensitive systems, narrow autonomy wins almost every time.
10. Measure review load, not just agent output
Most organizations evaluate AI coding tools using output metrics:
- lines changed
- tickets closed
- draft PRs opened
- time to first patch
Those are weak indicators.
The better signal is review burden:
- how many review cycles did the PR require?
- how often did a senior engineer have to provide hidden context?
- how many generated changes were reverted?
- how many production issues traced back to undocumented assumptions?
If the agent writes code faster but increases senior review load by 30%, you did not gain capacity. You shifted effort upstream.
A practical dashboard for quarter one:
- PRs with agent assistance as % of total
- median review rounds for agent-assisted PRs vs non-agent PRs
- revert rate for agent-assisted PRs
- stale Tier 1 docs as %
- time-to-merge for changes touching documented vs undocumented critical surfaces
That is the operating view a CTO can actually use.
05 STRATEGIC TAKEAWAY
AI agents are not forcing a new documentation philosophy; they are exposing whether your engineering organization ever built an async-capable one. If you apply this discipline, agents become useful force multipliers on bounded work within one quarter because the context they need is explicit, current, and reviewable. If you do not, the likely outcome is familiar: local productivity demos, followed by rising staff review load, inconsistent changes in critical systems, and a quiet loss of trust that stalls adoption. The CTO decision this quarter is not “which agent should we buy?” It is whether the organization will fund source-of-truth documentation as part of software delivery, the same way it funds CI, observability, and incident response.
06 IMPLEMENTATION ANGLE
Start with one workflow, not a platform-wide mandate. Pick a high-friction surface like billing integrations, RBAC, or data migrations. Create one system contract, one change guide, and one incident memory doc for that surface. Put them in version control. Then require that any PR touching that area either references them or updates them.
For tooling, keep it boring. GitHub or GitLab PR templates, CODEOWNERS, markdown frontmatter, and a simple doc freshness check in CI are enough to start. If you already use an internal developer portal like Backstage, attach contract docs and ownership metadata there too. The important move is not fancy retrieval. It is making authoritative context discoverable and maintained. related topic
At the team level, assign a directly responsible individual for each Tier 1 surface. Not “the platform team.” A named engineer or tech lead. In scaling orgs, this is where Amplify can help engineering teams scale: not by replacing technical leadership, but by helping install repeatable ownership, documentation, and review patterns before complexity compounds.



