Operations

Find the bottleneck, not the symptom

The loudest problem is rarely the real one. The discipline of tracing work backward to the single constraint everything else waits on.

Most operational problems present as a symptom: a delayed launch, a quality issue, a frustrated customer. The instinct is to fix the symptom. The discipline is to trace the work backward until you find the single bottleneck that all the symptoms point to. Fix the bottleneck, the symptoms resolve. Fix the symptoms, they keep coming back.

The underlying law, borrowed from the theory of constraints and still undefeated: a system's throughput is set by exactly one constraint at a time. Improving anything that is not the constraint improves nothing except the local team's dashboard; the work just arrives at the real bottleneck faster and queues there deeper. This is why so much process improvement produces motion without movement, and why finding the constraint is worth more than ten optimisations placed anywhere else.

How to find it

Walk the process backward from the symptom. At each step, ask: what is this step waiting for? Where does the work sit longest? Where is the queue longest? The bottleneck is the step everything else feeds into and almost always the step that takes longest to clear.

Make the walk empirical, not anecdotal: take five recent items end to end and timestamp every state change, then compute touch time versus wait time per step. The step with the fattest wait column is the constraint, and it is very often not where the organisation believes it is (the loudest team is usually downstream of the bottleneck, drowning in its output starvation, not causing it). One afternoon of timestamps beats a month of opinions.

Why teams miss it

Bottlenecks are usually held by people who are already working flat out, which makes them look like high performers, not constraints. The team protects them by routing work around them, which hides the bottleneck instead of fixing it. A protected bottleneck is still a bottleneck, and it caps the throughput of the entire system.

There is also an emotional reason: naming the bottleneck feels like blaming a person, so diagnoses stay vague. The reframe that unlocks honest analysis: the constraint is a position, not a performance. The person at the bottleneck is usually the most capable person in the flow, which is exactly how they ended up with everything routing through them. Saying "the constraint is review" indicts the design that put one person in every critical path, not the reviewer.

What to actually do

Reduce demand on the bottleneck, add capacity to it, or redesign the process so it is no longer in the critical path. Anything else is rearranging the symptoms.

In that order, deliberately. Demand first, because it is free: a surprising fraction of what queues at any constraint does not need to be there at all (approvals that could be thresholds, reviews that could be samples, requests that a written FAQ would absorb). Then exploitation: protect the constraint's hours, feed it prepared work, strip everything from its plate that anyone else could do. Only then add capacity, because capacity added before the demand cleanup just gets consumed by the same junk volume. And expect the sequel: fix one constraint and the bottleneck moves somewhere new. That is not failure; that is the system getting faster, one constraint at a time. The discipline is re-running the trace each quarter instead of assuming last quarter's answer still holds.

Takeaways

What to do with this

Related

Keep reading

Put the playbook to work.

Cafiyn Lens tells you which market is worth the effort, and Cafiyn FlyWheel runs the acquisition loop against it. Two products, one shared Blueprint, from $14.99/mo.