A handbook for software teams

Software is moving faster than you can review it.

Agents can write, test, review, and ship software at a pace that changes the old way of working. The harder problem is building the system around them: clear intent, useful context, good evidence, safe boundaries, and feedback from what actually happens in production. This handbook is about how teams are figuring that out.

01 Why this exists

PostHogPostHog went from 1,441 PRs in January to 4,725 in June. Over that same stretch, the percentage of PRs created by AI agents grew from around 20% to over 70%. The interesting question is no longer just how to make agents write code. It is how to build an engineering system that can keep up with that pace.

The teams figuring this out are changing the whole pipeline: how production is observed, how intent and context reach an agent, how a change is understood before it is merged, how evidence is gathered, and how outcomes feed back into the next decision. Humans move toward the things that still need judgment: what's actually worth building, what is safe enough to ship, how to keep systems reliable.

Most teams already have the raw tools: GitHub, CI, observability, browser tests, deployment platforms, and agent tooling. The harder problem is connecting them into a system that stays useful as the software, models, and workflows change.

There is another problem. Faster development can also mean faster accumulation of mess: stale tests, duplicated abstractions, outdated documentation, weak conventions, and architectural drift. The more we ship, the more important this maintenance becomes.

I am writing this by studying teams who are figuring this out: reading their engineering work, following what they are shipping, talking to engineers directly, and testing the patterns myself. The goal is to understand what actually works, what doesn't, and why.

02 How it's organized

The whole book maps to six stages.

The loop is simple. The hard part is building each stage well enough that the next one can rely on it.

  1. 1

    OBSERVE

    Know what is actually happening: errors, usage, traces, replays, deployments, and failures.

  2. 2

    UNDERSTAND

    Know what the change is supposed to do, what it touches, and what context the agent needs before it acts.

  3. 3

    CHANGE

    Let humans or agents make the change with the right tools, constraints, context, and durable execution.

  4. 4

    VERIFY

    Gather evidence that the intended behavior still holds. A green test suite is only one piece of that evidence.

  5. 5

    SHIP

    Move changes through safe boundaries with approvals, progressive delivery, rollback, and clear ownership.

  6. 6

    LEARN

    Turn production outcomes, failures, and human decisions into better context, evaluations, checks, and workflows.

  7. Then back to Observe. The point is to make the next cycle better.

03 What's inside

The full table of contents

Twenty-two chapters across six parts. The structure might evolve as I learn from teams doing this in production. Each chapter is grounded in practical work, research, or both.

Part I: The New Engineering System

  1. 01
    When Writing Code Stops Being the Bottleneck

    How the engineering workflow changes when implementation becomes less of a bottleneck, what happens when humans can no longer review every change, and where human judgment becomes more important.

  2. 02
    The Engineering Loop

    Observe, Understand, Change, Verify, Ship, Learn, and why the stages have to work as one system.

  3. 03
    What Does Correct Actually Mean?

    Intent, specifications, acceptance criteria, user journeys, behavioral contracts, invariants, and definitions of done.

  4. 04
    The Cost of Moving Fast

    Review fatigue, AI-generated mess, architectural drift, stale tests and docs, duplicated abstractions, and the case for continuous maintenance.

Part II: Give the System the Right Context

  1. 05
    Make the Codebase Legible to Agents

    Repository knowledge, architecture, domain context, history, ownership, retrieval, skills, and keeping context fresh.

  2. 06
    Understand the Blast Radius

    Connect a change to dependencies, user journeys, ownership, historical patterns, and production behavior before it is made.

  3. 07
    Give Agents a View of Production

    Errors, logs, traces, analytics, session replay, feature flags, deployments, and the signals that help explain what users actually experience.

  4. 08
    The Agent Harness

    Tools, skills, memory, permissions, sandboxes, state, checkpoints, retries, and the environment around the model.

Part III: Change With Guardrails

  1. 09
    Before You Give an Agent Write Access

    Repository conventions, contracts, test environments, secrets, permissions, and the boundaries that make agent work safe.

  2. 10
    Long-Running Agents

    Durable execution, resumability, partial failures, checkpoints, artifacts, handoffs, and what changes when work outlives a single session.

  3. 11
    What Counts as Evidence?

    Why a green check is only one signal, and how tests, browser runs, traces, replays, static checks, and human judgment fit together.

  4. 12
    Verify What Changed

    Risk-based verification, test-impact analysis, targeted browser checks, historical regressions, and evidence proportional to the change.

Part IV: Trust, Autonomy & Shipping

  1. 13
    Who Should Decide?

    What agents can decide alone, what needs approval, and how teams increase autonomy without losing accountability.

  2. 14
    Independent Verification

    How to keep an agent from becoming its own oracle, with external checks, evidence quality, and failure-aware review.

  3. 15
    Human Review Without the Noise

    Review routing, ownership, alert fatigue, AI review quality, and preserving human attention for decisions that need it.

  4. 16
    Ship Safely

    Preview environments, CI/CD, progressive delivery, rollback, auditability, and the boundary between merge and production.

Part V: Learn, Evaluate, Improve

  1. 17
    Turn Production Failures Into Learning

    Use incidents, user behavior, replays, and escaped defects to create regression knowledge and better future verification.

  2. 18
    Build the Evaluation Loop

    Offline evals, production-derived evals, traces, golden cases, human judgments, false positives, and evaluating the evaluator.

  3. 19
    Keep the System From Decaying

    Continuously maintain tests, docs, abstractions, architecture, agent skills, and the rules that keep the system legible.

  4. 20
    Measure Whether It Actually Works

    Escaped defects, detection time, regression catch rate, evidence quality, intervention rate, cost, and safe engineering velocity.

Part VI: Put It Into Practice

  1. 21
    Connect the Tools You Already Have

    Practical setups for GitHub, CI, preview environments, Playwright, Sentry, PostHog, Slack, agent tools, and the feedback loops between them.

  2. 22
    Case Studies, Playbooks & Research

    Close reads of teams doing this in production, practical workflows you can copy, and a small research shelf kept deliberately high signal.

04 Is this for you

This is for you if

  • You ship multiple times a week and a green checkmark doesn't fully reassure you anymore.
  • You've started letting AI agents make meaningful changes and you're figuring out what context, evidence, and review they actually need.
  • You want to understand the engineering systems emerging around agents, not another list of AI tools.
  • You're interested in making software improve its own checks and workflows from what happens in production.

Probably not if

  • You're looking for a plug-and-play tool. This is about the engineering system around the tools.
  • You want a generic AI-agent tutorial covering prompts, models, or frameworks.
  • You need a finished manual right now. I'm still learning and writing this as the field moves.
Call for contributions

How does your team ship code when agents write a lot of it?

This handbook isn't built on theories or based on one person's opinions. It's informed by software teams figuring this out in production right now. So I'm looking for concrete workflows, failures, internal tools, and lessons from teams figuring this out in prod. If you've built something that works (or discovered where your workflow breaks), I want to learn from your experiences and feature it.

05 Who's putting this together

Sourav

Sourav

Product Engineer

Hi, I'm Sourav, a product engineer. I've spent the last few years building for small teams (Paragraph, Pimlico, Gallery, RabbitHole) and working on my own things.

Right now, I'm building BeenThere, a minimal travel platform. To move faster as a solo developer, I started relying on agents to write code. I quickly learned that generating code is the easy part. The harder problem is building the system around the agent: giving it the right context, understanding what a change can affect, verifying the result, and learning from what happens after shipping.

I'm putting this handbook together because I needed it. I'm studying engineering work from teams already doing this in production, talking to engineers directly, and reading research on agentic software engineering, evaluation, observability, reliability, and self-healing systems. I want to understand what really works in practice.

These are the teams with useful public work or workflows worth studying. My role is to test the patterns hands-on, see what actually works, and organize what I learn. If your team is figuring this out in production, share how your team ships.