2026-09-18 #architecture #capabilities #drift #distributed-systems

3,005 words · 16 min read

Capability first, distributed

I have argued elsewhere that a capability should be the primary artifact of a service, and that endpoints, tools, and clients are projections of it. That argument stood at the edge of one service and looked outward. This essay is about what capability first has to contend with once there are many services calling each other.

A capability is a named unit of business function. It has typed inputs, typed outputs, a closed set of failures, one line of intent, and an owner, the part of the business that performs it.

Above is the first essay’s definition with the intent line added. Every time this essay says capability, it means all of that.

In a microservice estate a capability’s producer and its consumers are different deployables, owned by different teams, released on different days. Whatever “capability first” means, it has to mean something under those conditions, or it is a tidy way to organise a codebase and nothing more. The nuances are all about time. Who holds the promise, when they learned it, and what happens when the promise moves while they are not looking.

There are two ways the promise moves. It moves between a producer and its consumers, on different clocks. And it moves between the layers that advertise it, the service, the bounded context, and the organisation, each publishing a list of what it has, each one step further from the code that makes the list true. This essay is about both, and about what it takes to maintain a capability list under them. The clearest way to see either is to start where neither exists.

One container#

A monolith is one process, one codebase, one deployment. Inside a typed one, a capability is a method. generateStatement(account, period, format, delivery), throwing AccountClosed or PeriodInvalid, in a package called reporting. That method is the capability, every part of it. It is also a symbol, and symbols come with a property nobody notices until it is gone: every use of the capability and the capability itself are in the same place, under the same compiler, changed by the same commit.

the monolithone process, one clockpackage reportinggenerateStatement(…)throws AccountClosed, PeriodInvalidpackage billingchargeStatement()package notificationssendStatement()package portalshowStatement()every use and the capability itself: one place, one compiler, one commit
The monolith: the capability and every use of it in one place, under one compiler, changed by one commit. It has no list, because nobody needs one.

Rename the method and every caller follows. Add a checked failure and every caller that does not handle it stops compiling. Remove a parameter and the build tells you, in the same minute, exactly who depended on it. The monolith has one lifecycle. When it deploys, the promise a capability makes and every understanding of that promise deploy together, because they are the same artifact.

Notice what the monolith does not have. It has no list of its capabilities. Nobody needs one. The compiler holds the whole estate in its head, and “what can this system do” is answered by reading the service layer.

Where the symbol stops#

Then the monolith is split into services, for reasons that are all good. Teams need to deploy without waiting for each other. Parts need to scale differently. Ownership needs edges. The capability generateStatement moves into a statements service, in the reporting context, and the code that called it moves into a billing service, a notifications service, and a customer portal.

Something happened at that moment that is rarely said plainly. The capability stopped being a symbol. On the producer side it is still a method. On every consumer side it is now a client, a copy of the producer’s shape, written or generated at some point in the past, held in a different codebase, deployed on a different day. The one artifact became several, and the several are held by different people on different clocks.

That is the whole change. Not the network, not the latency, not the serialisation. Those are costs, and teams budget for them. What changed is that a promise and every understanding of it stopped deploying together.

the monolithone process, one compiler, one clockstatements servicedeployed Tuesdaybilling servicedeployed 3 weeks agonotifications servicedeployed in springcustomer portal servicedeployed last yeargenerateStatement(…)the promisegenerateStatementa client, generated at a momentgenerateStatementa client, generated at a momentgenerateStatementa client, generated at a moment
The split. One artifact becomes several: the promise stays in the statements service, and every caller now holds a client generated from it at a moment, deployed on a clock of its own.

Three containers, three clocks, three advertisers#

After the split a capability lives in three containers at once. Each one changes for its own reasons on its own clock, and each one advertises what it has.

The service is the first container, the one that implements it. A service is built, deployed, scaled, refactored, renamed, split, and retired, on a clock of days. Everything about the implementation is free to change, and should be. What the service owes the layers above is the promise, unchanged.

The bounded context is the second container, the one that owns it. Reporting is a part of the business with its own vocabulary, and generate_statement is a word in that vocabulary. A context’s language forms, settles, and grows, and occasionally the context splits or merges with another, on a clock of years. What the capability becomes here is a name with an owner and a place among siblings, and its failures have to mean what the context means by them. What the context owes the layer above is stable names.

The organisation is the third container, the one that depends on it. Billing depends on reporting’s statements without sharing reporting’s vocabulary. A partner depends on them through a public API. An agent depends on them through a tool. The organisation reorganises, acquires, divests, and changes strategy, on a clock of years to decades. What the capability becomes here is a qualified entry that strangers depend on, projected to edges, partners, and agents. What must be stable here is the qualified name and the shape the strangers hold, and nothing else.

The rule these three produce is simple to state.

What may change freely at one layer is exactly what must be held stable by the layer above it. Implementation is free at the service. The name and the failures are held at the context. The qualified name and the shape consumers hold are held at the organisation.

the organisationreorganises over decadescontext: reportingits language settles over yearsholds stable: the name, the failuresservice: statementsdeploys in daysgenerate_statement(account, period, format, delivery) → runfetch_statements(account, period) → statementsfree to change: everything about the implementationdepends on it: billing, without reporting's vocabulary · a partner, through an API · an agent, through a toolholds stable: the qualified name, and the shape strangers hold
Three containers, three clocks. Service changes in days and is free to. Context settles over years and holds the name and the failures. The organisation moves over decades and holds the qualified name and the shape strangers depend on.

The package the estate never had#

There is a shorter way to say why the monolith is easy. A package is a driftless container, and a bounded context after the split is the same idea with the driftlessness removed.

A package gives you four things, and one compiler enforces all of them at one moment. A namespace, so names are unique inside it and qualified by it from outside. Membership decided by location, so a class is in the package because its file is and nobody keeps a list. Visibility decided by the owner, so the package says what leaves it and the compiler refuses a caller that reaches past that. And one clock, so a change to what the package exports is a change every caller sees at once.

A bounded context in a distributed estate has the namespace as a convention, the membership as an entry in a list the context maintains by hand, the visibility as a norm someone once wrote down, and no clock at all, because its members are separate deployables. Every property the package enforced, the context asks people to honour. The estate has nothing that provides a driftless notion of a package, and that sentence is the whole gap.

The analogy holds structurally and breaks temporally, and it is worth keeping the break honest. A package’s members deploy together. A context’s never will. So the context cannot borrow the package’s mechanism. It has to be given the package’s properties another way, and that is the question the rest of this essay circles.

a packageone compilernamespacenames qualified by itmembershipby locationvisibilitythe owner marks exportsone clockevery caller sees the changea context, after the splitno compilernamespacea conventionmembershipa list kept by handvisibilitya norm, once written downmany clocksone per deployable
A package and a split context give the same four things. The compiler enforces all four for the package. The context asks people to honour them, and has no clock at all.

One of those properties deserves its own sentence, because it changes who decides. In a package the owner marks what is public. In most estates every endpoint is public because nobody had a way to say otherwise, and “internal” means “we hope nobody found it.” A context should decide which of its capabilities are published beyond it, which are visible to its own services only, and which stay inside one service, and that decision should be enforced the way a module’s exports are enforced, not hoped for. Eric Evans gave this the names it still carries, in Domain-Driven Design more than twenty years ago. The open host service is “a protocol that gives access to your subsystem as a set of services,” what a context offers outward. The published language is the well-documented shared language it speaks there. The concept and the names have been his since 2003. Neither ever had a compiler.

Drift, the first way: producer and consumer#

Now put the clocks together and watch what happens.

The statements service deploys on Tuesday. In that deployment, generate_statement gained a failure, delivery_unavailable, because the postal integration now has an outage mode. It is a small, correct change. The reporting team’s tests pass. Their deployment is green.

Billing deployed three weeks ago. Its client for generate_statement was generated against the shape that existed then. Billing’s understanding of the promise has two failures in it. Reporting’s promise now has three. Nobody is wrong, and nobody is told. The first time the postal service has an outage, billing receives a failure it has no branch for, and whatever billing does with an unexpected outcome is what happens to the customer.

statements servicedeployed Tuesdaygenerate_statementfails:account_closedperiod_invaliddelivery_unavailablea small, correct change; tests greenbilling serviceclient generated 3 weeks agogenerate_statementthe client handles:account_closedperiod_invalidno branchcallsdelivery_unavailablewhatever billing does nowis what happens to the customer
Drift, the first way. On Tuesday the promise gains a failure. The client billing holds does not. Nobody is wrong, nobody is told, and the first outage arrives at a branch that does not exist.

That is drift, and it is worth stating exactly, because the word gets used loosely.

Drift is not a bug in either service. It is the producer’s promise and the consumer’s understanding of that promise changing on different clocks, with nothing that holds them together.

The monolith held them together by being one artifact. The split removed that, and nothing replaced it.

It is worth separating this from a different drift that gets the same name. An endpoint can drift without the capability changing at all. A parameter moves from the query string to a form field. A date changes format. A field is renamed in the JSON while the shape underneath stays what it was. Callers break, and the promise did not move an inch. That is costume drift, and it is the drift the industry has built its tooling to catch. A specification describes the costume, a contract test exercises the costume, a schema registry diffs the costume, and each of them can catch the parameter that moved. None of them would have said anything about delivery_unavailable unless someone had modelled the failure set in the costume, and almost nobody does. A generic error object with a free-text code looks identical on Tuesday and three weeks before, and even a specification that lists the codes would have reported the change to whoever reads specification diffs, which is not billing. The tooling is aimed at the drift you can see from outside. The drift that reaches the customer is the one inside the promise.

Every mechanism teams add after the split is an attempt to replace what the monolith had, and each replaces a part. Versioned paths let the old understanding keep working, for a while, if the producer remembers to keep the old version alive. Contract tests catch the disagreement, after the fact, for the pairs that wrote one. A published specification lets the consumer regenerate its understanding, if the consumer notices the specification changed. Release notes tell humans, who tell other humans. All of these are people and processes standing where a compiler used to be.

And drift compounds with the clocks. A service changes in days, a context’s language in years, and a partner integration is built once and forgotten. The faster the producer’s clock and the slower the consumer’s, the wider the gap between the promise and the understanding grows, and the consumers on the slowest clocks, the partners and the agents, are exactly the ones nobody remembers to tell.

Drift, the second way: the advertising chain#

The first drift runs sideways, between a producer and its consumers. The second runs upward, between the layers that advertise the capability, and it is the one that quietly breaks every catalogue an organisation has ever kept.

A service advertises what it has. If its list is derived from its code, read from the methods, their types, and their failures at build time, the list cannot drift from the service. It is true by construction and it is the one place in the chain where that is so.

A bounded context advertises what its services have. That list is gathered, and the moment gathering is a person or a periodic job, the list is a snapshot. Reporting’s list says what reporting’s services published the last time anyone gathered them. Add a service, rename a capability, retire one, and the list is wrong until someone notices. Most contexts keep this list as a page in a wiki or a catalogue entry maintained by hand, which is why most contexts’ lists are wrong.

The organisation advertises what its contexts have. That list is gathered from the contexts’ lists, one more hop, one more moment, and it drifts further than any of them because it is furthest from the code. This is the enterprise capability map, redrawn by hand once a year from interviews, and the reason it is always describing last year’s estate is that nothing between the code and the map is derived.

list · statementsderived at build: truefetch_statementsgenerate_statementwithdraw_runrenamed today: cancel_run → withdraw_runlist · reportingby hand, as of last monthfetch_statementsgenerate_statementcancel_runlist · organisationby hand, as of last yearreporting.fetch_statementsreporting.generate_statementreporting.cancel_runa person, a periodic job: a snapshota snapshot of a snapshotbilling serviceclient generated 3 weeks agogenerate_statementthe client: an understanding formed at a momentnothing between the list and the client: drift, unseen
Drift, the second way. Only the service's list is derived and true by construction. The context's list and the organisation's list are gathered by hand at a moment, and a rename at the bottom reaches neither.

And every service is also a consumer at each of those levels. It takes a client from another service, generated at a moment. It takes its understanding of a neighbouring context from that context’s list, read at a moment. It takes its picture of what the organisation can do from the organisation’s list, read at a moment. Three hops, three moments, three chances for the understanding it holds to have been true once and not now.

Maintaining a capability list in a distributed estate is therefore not one job. It is a chain of advertisements, each of which is only as true as the way it was produced, and a chain of consumptions, each of which is only as current as the day it was formed. The practical questions are all about that chain. Who gathers a context’s list, and when. What happens when two services in one context advertise the same name. What a context does when it renames itself, and how the organisation’s list and every consumer of the old name find out. Whether an edge that presents a resource to partners is reading a context’s list or a snapshot of one. None of these questions exist in the monolith, and all of them are the daily work of anyone who has tried to keep a catalogue honest.

List is the start. A drift-free list is the answer.#

An obvious response to both drifts is to write the capabilities down. Every service publishes its list. I have argued for that list before and I still would. It gives a newcomer something to read before the routes. It gives an agent something to call. It makes the capability visible where before only its costume was.

But look at what the list is. It is the producer’s promise, written down, at a moment. It is one more artifact on the producer’s clock. If billing reads the list on the day it generates its client, and reporting changes the list three weeks later, billing has a stale list and a stale client, and the list changed nothing about the outage. And the list is only true at the level where it was derived. The moment it is gathered into a context’s list or an organisation’s list by hand, it is a snapshot of a snapshot, and it drifts the second way as surely as the client drifts the first. The list makes drift visible to anyone who looks. It does not make drift impossible, and nobody looks on the day that matters.

So the property worth having is not a list of capabilities. It is a list that cannot drift, at every level of the chain, with every holder bound to it. A capability whose promise cannot change without every holder of that promise knowing before the change ships, the way every caller in the monolith knew, in the same minute, because the build told them. That is what the monolith had. That is what the split lost. Everything else about services, the independence, the scaling, the ownership, was worth keeping. This one thing was not worth losing, and we lost it without deciding to.

What it would take#

The shape of the answer is not in doubt. What is in doubt is whether it can be produced in practice.

What would it take for a capability to keep the monolith’s property across service boundaries?

The promise and every understanding of it would have to share a source again, so that a change to one is a change to all. Every advertisement above the service, the context’s list and the organisation’s list, would have to be derived from the level below it rather than gathered by hand, so that the chain is true all the way up. Someone or something would have to know who holds the promise, at every hop, so that “every caller” is a known set and not a hope. A change to the promise would have to be refused, not reported, while a holder of the old understanding is still live, because reporting after the fact is what we have now and it is not enough. And a context would have to be able to say what leaves it, with that refused rather than hoped for, the way a module’s exports are.

list · statementsderived at buildfetch_statementsgenerate_statementwithdraw_runthe promise: one sourcelist · reportingderived from the services' listsfetch_statementsgenerate_statementwithdraw_runlist · organisationderived from the contexts' listsreporting.fetch_statementsreporting.generate_statementreporting.withdraw_runderived: true by constructionderived: true by constructionbilling servicebound to the listgenerate_statementthe same source as the promisea change to the promise, while billingstill holds the old one: refused, not reported
The goal: every hop derived, every holder bound to the same source as the promise, and a change refused, not reported, while a holder of the old understanding is live.

Every one of those is a practical problem, not a conceptual one. Deriving the service’s list is the easy part. It is a build step, not a research problem. Deriving the context’s list means the gathering has to be an act the system performs, with rules for what happens when two services advertise one name, and nobody gathers by hand. Deriving the organisation’s list means the same one level up, and surviving a context that renames itself. Knowing every holder means the clients have to come from the list, not be written against it. Refusing a change means someone has to accept that a deploy can be stopped by a consumer they have never met. Some of this can be done with discipline, in a small estate, by people who talk to each other. I have never seen it survive the third team. Past that, it has to be built.