Capability first, distributed
I have argued elsewhere that a capability should be the primary artifact of a service, and that endpoints, tools, and clients are projections of it. That argument stood at the edge of one service and looked outward. This essay is about what capability first has to contend with once there are many services calling each other.
A capability is a named unit of business function. It has typed inputs, typed outputs, a closed set of failures, one line of intent, and an owner, the part of the business that performs it.
Above is the first essay’s definition with the intent line added. Every time this essay says capability, it means all of that.
In a microservice estate a capability’s producer and its consumers are different deployables, owned by different teams, released on different days. Whatever “capability first” means, it has to mean something under those conditions, or it is a tidy way to organise a codebase and nothing more. The nuances are all about time. Who holds the promise, when they learned it, and what happens when the promise moves while they are not looking.
There are two ways the promise moves. It moves between a producer and its consumers, on different clocks. And it moves between the layers that advertise it, the service, the bounded context, and the organisation, each publishing a list of what it has, each one step further from the code that makes the list true. This essay is about both, and about what it takes to maintain a capability list under them. The clearest way to see either is to start where neither exists.
One container#
A monolith is one process, one codebase, one deployment. Inside a typed one, a capability is a method. generateStatement(account, period, format, delivery), throwing AccountClosed or PeriodInvalid, in a package called reporting. That method is the capability, every part of it. It is also a symbol, and symbols come with a property nobody notices until it is gone: every use of the capability and the capability itself are in the same place, under the same compiler, changed by the same commit.
Rename the method and every caller follows. Add a checked failure and every caller that does not handle it stops compiling. Remove a parameter and the build tells you, in the same minute, exactly who depended on it. The monolith has one lifecycle. When it deploys, the promise a capability makes and every understanding of that promise deploy together, because they are the same artifact.
Notice what the monolith does not have. It has no list of its capabilities. Nobody needs one. The compiler holds the whole estate in its head, and “what can this system do” is answered by reading the service layer.
Where the symbol stops#
Then the monolith is split into services, for reasons that are all good. Teams need to deploy without waiting for each other. Parts need to scale differently. Ownership needs edges. The capability generateStatement moves into a statements service, in the reporting context, and the code that called it moves into a billing service, a notifications service, and a customer portal.
Something happened at that moment that is rarely said plainly. The capability stopped being a symbol. On the producer side it is still a method. On every consumer side it is now a client, a copy of the producer’s shape, written or generated at some point in the past, held in a different codebase, deployed on a different day. The one artifact became several, and the several are held by different people on different clocks.
That is the whole change. Not the network, not the latency, not the serialisation. Those are costs, and teams budget for them. What changed is that a promise and every understanding of it stopped deploying together.
Three containers, three clocks, three advertisers#
After the split a capability lives in three containers at once. Each one changes for its own reasons on its own clock, and each one advertises what it has.
The service is the first container, the one that implements it. A service is built, deployed, scaled, refactored, renamed, split, and retired, on a clock of days. Everything about the implementation is free to change, and should be. What the service owes the layers above is the promise, unchanged.
The bounded context is the second container, the one that owns it. Reporting is a part of the business with its own vocabulary, and generate_statement is a word in that vocabulary. A context’s language forms, settles, and grows, and occasionally the context splits or merges with another, on a clock of years. What the capability becomes here is a name with an owner and a place among siblings, and its failures have to mean what the context means by them. What the context owes the layer above is stable names.
The organisation is the third container, the one that depends on it. Billing depends on reporting’s statements without sharing reporting’s vocabulary. A partner depends on them through a public API. An agent depends on them through a tool. The organisation reorganises, acquires, divests, and changes strategy, on a clock of years to decades. What the capability becomes here is a qualified entry that strangers depend on, projected to edges, partners, and agents. What must be stable here is the qualified name and the shape the strangers hold, and nothing else.
The rule these three produce is simple to state.
What may change freely at one layer is exactly what must be held stable by the layer above it. Implementation is free at the service. The name and the failures are held at the context. The qualified name and the shape consumers hold are held at the organisation.
The package the estate never had#
There is a shorter way to say why the monolith is easy. A package is a driftless container, and a bounded context after the split is the same idea with the driftlessness removed.
A package gives you four things, and one compiler enforces all of them at one moment. A namespace, so names are unique inside it and qualified by it from outside. Membership decided by location, so a class is in the package because its file is and nobody keeps a list. Visibility decided by the owner, so the package says what leaves it and the compiler refuses a caller that reaches past that. And one clock, so a change to what the package exports is a change every caller sees at once.
A bounded context in a distributed estate has the namespace as a convention, the membership as an entry in a list the context maintains by hand, the visibility as a norm someone once wrote down, and no clock at all, because its members are separate deployables. Every property the package enforced, the context asks people to honour. The estate has nothing that provides a driftless notion of a package, and that sentence is the whole gap.
The analogy holds structurally and breaks temporally, and it is worth keeping the break honest. A package’s members deploy together. A context’s never will. So the context cannot borrow the package’s mechanism. It has to be given the package’s properties another way, and that is the question the rest of this essay circles.
One of those properties deserves its own sentence, because it changes who decides. In a package the owner marks what is public. In most estates every endpoint is public because nobody had a way to say otherwise, and “internal” means “we hope nobody found it.” A context should decide which of its capabilities are published beyond it, which are visible to its own services only, and which stay inside one service, and that decision should be enforced the way a module’s exports are enforced, not hoped for. Eric Evans gave this the names it still carries, in Domain-Driven Design more than twenty years ago. The open host service is “a protocol that gives access to your subsystem as a set of services,” what a context offers outward. The published language is the well-documented shared language it speaks there. The concept and the names have been his since 2003. Neither ever had a compiler.
Drift, the first way: producer and consumer#
Now put the clocks together and watch what happens.
The statements service deploys on Tuesday. In that deployment, generate_statement gained a failure, delivery_unavailable, because the postal integration now has an outage mode. It is a small, correct change. The reporting team’s tests pass. Their deployment is green.
Billing deployed three weeks ago. Its client for generate_statement was generated against the shape that existed then. Billing’s understanding of the promise has two failures in it. Reporting’s promise now has three. Nobody is wrong, and nobody is told. The first time the postal service has an outage, billing receives a failure it has no branch for, and whatever billing does with an unexpected outcome is what happens to the customer.
That is drift, and it is worth stating exactly, because the word gets used loosely.
Drift is not a bug in either service. It is the producer’s promise and the consumer’s understanding of that promise changing on different clocks, with nothing that holds them together.
The monolith held them together by being one artifact. The split removed that, and nothing replaced it.
It is worth separating this from a different drift that gets the same name. An endpoint can drift without the capability changing at all. A parameter moves from the query string to a form field. A date changes format. A field is renamed in the JSON while the shape underneath stays what it was. Callers break, and the promise did not move an inch. That is costume drift, and it is the drift the industry has built its tooling to catch. A specification describes the costume, a contract test exercises the costume, a schema registry diffs the costume, and each of them can catch the parameter that moved. None of them would have said anything about delivery_unavailable unless someone had modelled the failure set in the costume, and almost nobody does. A generic error object with a free-text code looks identical on Tuesday and three weeks before, and even a specification that lists the codes would have reported the change to whoever reads specification diffs, which is not billing. The tooling is aimed at the drift you can see from outside. The drift that reaches the customer is the one inside the promise.
Every mechanism teams add after the split is an attempt to replace what the monolith had, and each replaces a part. Versioned paths let the old understanding keep working, for a while, if the producer remembers to keep the old version alive. Contract tests catch the disagreement, after the fact, for the pairs that wrote one. A published specification lets the consumer regenerate its understanding, if the consumer notices the specification changed. Release notes tell humans, who tell other humans. All of these are people and processes standing where a compiler used to be.
And drift compounds with the clocks. A service changes in days, a context’s language in years, and a partner integration is built once and forgotten. The faster the producer’s clock and the slower the consumer’s, the wider the gap between the promise and the understanding grows, and the consumers on the slowest clocks, the partners and the agents, are exactly the ones nobody remembers to tell.
Drift, the second way: the advertising chain#
The first drift runs sideways, between a producer and its consumers. The second runs upward, between the layers that advertise the capability, and it is the one that quietly breaks every catalogue an organisation has ever kept.
A service advertises what it has. If its list is derived from its code, read from the methods, their types, and their failures at build time, the list cannot drift from the service. It is true by construction and it is the one place in the chain where that is so.
A bounded context advertises what its services have. That list is gathered, and the moment gathering is a person or a periodic job, the list is a snapshot. Reporting’s list says what reporting’s services published the last time anyone gathered them. Add a service, rename a capability, retire one, and the list is wrong until someone notices. Most contexts keep this list as a page in a wiki or a catalogue entry maintained by hand, which is why most contexts’ lists are wrong.
The organisation advertises what its contexts have. That list is gathered from the contexts’ lists, one more hop, one more moment, and it drifts further than any of them because it is furthest from the code. This is the enterprise capability map, redrawn by hand once a year from interviews, and the reason it is always describing last year’s estate is that nothing between the code and the map is derived.
And every service is also a consumer at each of those levels. It takes a client from another service, generated at a moment. It takes its understanding of a neighbouring context from that context’s list, read at a moment. It takes its picture of what the organisation can do from the organisation’s list, read at a moment. Three hops, three moments, three chances for the understanding it holds to have been true once and not now.
Maintaining a capability list in a distributed estate is therefore not one job. It is a chain of advertisements, each of which is only as true as the way it was produced, and a chain of consumptions, each of which is only as current as the day it was formed. The practical questions are all about that chain. Who gathers a context’s list, and when. What happens when two services in one context advertise the same name. What a context does when it renames itself, and how the organisation’s list and every consumer of the old name find out. Whether an edge that presents a resource to partners is reading a context’s list or a snapshot of one. None of these questions exist in the monolith, and all of them are the daily work of anyone who has tried to keep a catalogue honest.
List is the start. A drift-free list is the answer.#
An obvious response to both drifts is to write the capabilities down. Every service publishes its list. I have argued for that list before and I still would. It gives a newcomer something to read before the routes. It gives an agent something to call. It makes the capability visible where before only its costume was.
But look at what the list is. It is the producer’s promise, written down, at a moment. It is one more artifact on the producer’s clock. If billing reads the list on the day it generates its client, and reporting changes the list three weeks later, billing has a stale list and a stale client, and the list changed nothing about the outage. And the list is only true at the level where it was derived. The moment it is gathered into a context’s list or an organisation’s list by hand, it is a snapshot of a snapshot, and it drifts the second way as surely as the client drifts the first. The list makes drift visible to anyone who looks. It does not make drift impossible, and nobody looks on the day that matters.
So the property worth having is not a list of capabilities. It is a list that cannot drift, at every level of the chain, with every holder bound to it. A capability whose promise cannot change without every holder of that promise knowing before the change ships, the way every caller in the monolith knew, in the same minute, because the build told them. That is what the monolith had. That is what the split lost. Everything else about services, the independence, the scaling, the ownership, was worth keeping. This one thing was not worth losing, and we lost it without deciding to.
What it would take#
The shape of the answer is not in doubt. What is in doubt is whether it can be produced in practice.
What would it take for a capability to keep the monolith’s property across service boundaries?
The promise and every understanding of it would have to share a source again, so that a change to one is a change to all. Every advertisement above the service, the context’s list and the organisation’s list, would have to be derived from the level below it rather than gathered by hand, so that the chain is true all the way up. Someone or something would have to know who holds the promise, at every hop, so that “every caller” is a known set and not a hope. A change to the promise would have to be refused, not reported, while a holder of the old understanding is still live, because reporting after the fact is what we have now and it is not enough. And a context would have to be able to say what leaves it, with that refused rather than hoped for, the way a module’s exports are.
Every one of those is a practical problem, not a conceptual one. Deriving the service’s list is the easy part. It is a build step, not a research problem. Deriving the context’s list means the gathering has to be an act the system performs, with rules for what happens when two services advertise one name, and nobody gathers by hand. Deriving the organisation’s list means the same one level up, and surviving a context that renames itself. Knowing every holder means the clients have to come from the list, not be written against it. Refusing a change means someone has to accept that a deploy can be stopped by a consumer they have never met. Some of this can be done with discipline, in a small estate, by people who talk to each other. I have never seen it survive the third team. Past that, it has to be built.