Glossary of AI-era software delivery
43 terms, defined plainly — including the unglamorous ones teams say daily and never write down. Opinionated where the honest answer is an opinion.
For the practice itself, start with spec-driven development.
The practice
- Spec-driven developmentalso: SDD
Building software where the specification is the durable artefact and the code is downstream of it. The document is precise enough that a competent engineer — or agent — could implement from it and land somewhere you would recognise.
Not a return to waterfall: the loop is short and the specs are small. What changed is that generating code got cheap, so agreeing what to build became the expensive part.
- Spec
One unit of work, carrying two documents: requirements (what and why) and implementation (how). Everything else — tasks, approvals, code refs, pull requests — hangs off it.
- Requirements documentalso: product spec, PRD
The what-and-why half. Problem, who has it, what changes when it ships, what you are deliberately not building, and the hard rules.
Length is a feature. A reviewer who has to read 1,400 words to find the decision will skim, and skimming is how bad specs get approved.
- Implementation documentalso: design doc, technical spec
The how half: approach, data model, API surface, failure modes, test strategy, slices, rollout. Written after reading the actual code, not from imagination.
- Approval gate
A point a spec cannot pass until a named human approves it. The approval is bound to the document version they read — edit the document and the approval dies.
Without the binding, an approval is a timestamp on a document that has since changed, which is worse than no approval because it looks like one.
- Risk tier
How much scrutiny a spec needs. Low-risk work takes one approver; work that can cost money or break trust takes two distinct ones.
- Slice
A piece of implementation small enough to be wrong cheaply — two days or less, one reviewable pull request. Agents will happily generate a week of work in one pass, which nobody reviews properly.
- Outcome review
The stage after Released: did the thing you shipped do what the spec claimed? Most teams never close this loop, which is why the same mistake gets specified twice.
- Audit trail
The append-only record of who moved a spec, who approved which version, and when. The thing that answers “who decided this?” six months later, when nobody remembers the meeting.
Code and review
- Code refalso: linked code, code anchor
A GitHub permalink attached to a spec — repo, path, line range, commit — with a note saying why those lines matter. Rendered as source inline beside the document.
- Permalink pinning
Linking code at a commit SHA rather than a branch, so the link keeps showing the code the decision was made against instead of whatever the branch drifted into.
- Blast radius
Code a change could break but does not touch. Marking it explicitly tells a reviewer where to look for second-order damage without reading the whole diff.
- Draft PR
A pull request opened to explore rather than to merge: stubs, a shape, enough to discover the constraint nobody remembered. It informs the plan and should never be treated as evidence the work is done.
- Prompt-to-PR
The pattern where a request goes to an agent and comes back as a pull request. Fast, and it skips the two questions that matter: was this the right thing to build, and did anyone agree to it?
Agents and protocols
- MCPalso: Model Context Protocol
An open protocol that lets an AI agent read and write real systems — files, APIs, trackers — instead of being told the state of the world in every prompt.
Practically: the difference between an assistant you brief from scratch each time and one that can look things up and write back.
- MCP server
A service exposing tools an agent may call. A well-designed one is defined as much by what it omits — RelayWork has no approve tool at all, so no misconfiguration can let a machine sign off its own work.
- Tool scope
Whether a token may read or also write, and which tools it can reach. Scoped, revocable tokens are what make agent access auditable rather than a shared password.
- Machine attribution
Recording an automated write as the token that made it rather than as a person. Without it, an agent’s edit silently borrows a human’s authority.
- Agent
A model given tools and a goal, running a loop rather than answering once. The useful distinction from a chatbot is not intelligence — it is that an agent changes state in systems you care about.
- Vibe coding
Building by describing what you want and accepting what comes back without reading it closely. Genuinely fine for a prototype. The failure mode is shipping it, then discovering nobody — human or model — knows why the code does what it does.
- Context engineering
Deciding what an agent should be able to see. Increasingly the real skill: a model with the approved spec, the review notes and the relevant code will beat a better model working from a paragraph.
- Context window
How much text a model can consider at once. Treat it as a budget you spend deliberately, not a bucket to fill — a window packed with marginal material produces worse answers than a curated one.
- Human in the loop
A design where a person must act before something proceeds. Worth distinguishing from human-on-the-loop, where a person merely watches and can intervene — the second is what most products actually ship while claiming the first.
- Guardrail
A constraint that holds regardless of what the model decides. A prompt asking an agent not to do something is not a guardrail; removing the capability is.
- House template
The fixed shape documents must follow in a workspace, enforced on submission. It is what stops agent output from being 1,400 words of fluent, unreviewable prose.
Engineering shorthand
- Dogfoodingalso: eat your own dog food
Using your own product to do real work, before your customers have to. Not a demo account — the actual work of the company, with the actual consequences of it being broken.
- Backfill
A one-off pass that fills in data the new code expects but the old data lacks — running the new logic across everything that already exists.
The word doubles as a hiring term (backfilling a departed role), which is why it confuses people in mixed company. In engineering it almost always means the data run.
- Drift
When two things that should agree quietly stop agreeing — the doc and the code, the counter and reality, staging and production.
Drift is rarely a bug you can point at. It accumulates, then surfaces as “that number has been wrong for months”.
- Idempotent
Safe to run twice. The second run changes nothing beyond what the first did, which is what makes retries, replays and backfills survivable.
- Soft delete
Marking a record deleted rather than removing it, so an accident is recoverable. The cost is that every query has to remember to exclude it, and one that forgets shows users their own deleted data.
- Toil
Manual work that scales with usage and teaches you nothing — re-running the same job, copying the same numbers. Worth automating not to save the minutes but because nobody does it reliably at 6pm on a Friday.
- Yak shaving
The chain of unrelated prerequisites you end up doing before the actual task. Four levels deep and you are upgrading a build tool to fix a typo.
- Bikeshedding
Spending the review on the trivial and legible part while the consequential part passes unexamined. Naming gets forty comments; the data model gets none.
- Flaky test
A test that fails sometimes with no code change. Worse than a failing test, because it teaches the team to ignore red.
- Feature flag
A switch that decides at runtime whether new code path is live, so shipping and releasing stop being the same event. The debt is real: every flag is a branch you must eventually delete.
- Canary
Releasing to a small slice first and watching before continuing. Only useful if someone is actually watching, and if the metric would move within the window.
- Tech debt
A shortcut taken knowingly, with interest paid in future slowness. The word is abused to mean “code I dislike” — the useful test is whether someone chose it and what they chose it for.
- Spike
A time-boxed investigation whose output is an answer, not code you keep. Anything built during one is throwaway by definition, and treating it otherwise is how prototypes reach production.
- Postmortemalso: incident review
The write-up after something broke: what happened, what it cost, what would have prevented it. Blameless in the sense that the target is the system, not the person who typed the command.
Words teams use and rarely define
- Ship
To put a change in front of real users. Worth pinning down, since “shipped” variously means merged, deployed behind a flag, or actually released — and those are three different truths in a standup.
- WIP limit
A cap on how much work is open at once. Its purpose is not tidiness — it is forcing the team to finish something before starting something, which is the only reliable cure for a board full of 80%-done work.
- Scope creep
Work growing after it was agreed, one reasonable addition at a time. A written “what we are not building” section is the cheapest defence, because it makes each addition a visible decision rather than a drift.
- Source of truth
The one place a fact is authoritative, from which copies derive. Two sources of truth is a euphemism for none, and it is usually how a team describes the state it is trying to leave.