Pushes, Bearings, and Mass
Every company wants a flywheel. Almost nobody says which part of one they are building.
The metaphor belongs to Jim Collins, who introduced it in Good to Great in 2001. What he got right is the physics: momentum comes from many consistent pushes, none of them heroic, and the wheel eventually turns under its own weight. Where his version stops being useful is that his flywheel is one undifferentiated wheel, and the only thing you can do to it is push.
A real flywheel is a machine: a heavy wheel on an axle, with one hand on it. That is a good enough picture of a company, and it has exactly three parts you can spend an hour on. You can turn the wheel, you can improve its bearings, or you can add mass to its rim. The parts respond differently to the same hour, which is the detail the borrowed metaphor drops.
An hour on a push turns the wheel once. Ship the feature, run the launch, write the spec. Well-aimed, often the most valuable thing available that week, and when it is done the wheel is exactly where you left it.
An hour on the bearings improves the machine instead of moving it. The monorepo, CI, the design system, observability. Bearings add no energy of their own. They decide how fast the wheel bleeds off speed between one push and the next.
An hour spent adding mass to the rim builds something that stores momentum: an evaluation suite for an AI pipeline, the labeled data that improves a recommendation model every week. Mass is what keeps the wheel turning after the hand comes off.
All three are legitimate, and being honest about which one you are doing is most of the discipline. The mistake is calling a push an investment. Only the third is the flywheel everybody says they want, and it is the one almost nobody is actually building.
Seized Bearings Look Like Bad Luck
The complaint I hear most often sounds like this: we shipped all year, and January feels like starting from scratch. Teams read that as a missing flywheel. Usually it is simpler than that: seized bearings.
The version of this I have lived through was a design system. We had fifty screens, shipped over six months, each with its own hand-rolled spacing, colors and layout code. Every one of those screens had been a perfectly good push. Then a rebranding arrived, and the estimate for changing the colors and typography across all fifty came back so large that we stopped estimating. The cost of the pushes only became visible when we had to change all of them at once.
So we put in a design system and moved the existing screens onto it. The same change now costs nothing: colors and typography live in design tokens, and every screen updates the moment a token does. A screen built by someone who has since left is something a new joiner can safely modify on a Tuesday afternoon.
Nothing about the original pushes changed. What changed is how much we can still get out of them, which is a bearings question and stays open long after the push is done.
This is also where the word flywheel gets misapplied. A design system makes work easier, and easier is not the same as compounding. Friction has a floor: you can remove it once, and once it is gone the returns stop arriving. The fortieth screen is dramatically cheaper than the fourth. The four hundredth is not cheaper than the fortieth.
Mass has no floor. Every case you add to an evaluation suite sits on top of every case already in it. A team that reports its reuse gains as compounding is setting an expectation the second half of the year cannot meet.
You also cannot half-install a bearing. A monorepo migration that is sixty percent done is worse than never starting, because you are paying the full migration tax and collecting none of the relief. Mass can be added a kilogram at a time; bearings pay out at the end or not at all.
So pick the date the migration is finished, and defend it against whatever urgent thing arrives in week three. Let nothing new join the list until the current one is done. Half-finished foundations are usually invisible in the plan, because every individual decision to pause one looked reasonable at the time.
What Counts as Mass
Not everything you build stores momentum, and the difference has almost nothing to do with how good it is. The test is whether the work touches it on its own.
The clearest case I have lived through is a release process. Every release meant manual regression testing, driven by spreadsheets of who needed to check what. The spreadsheets were not neglected; they were maintained by people who cared, which made the weakness hard to see. The work was tedious but tolerable, right up until one of those people was out, and the whole release stalled with them.
Then we automated the regression suite. It runs with nobody in the loop and surfaces only when it actually needs a human. Same knowledge, written by the same people. Only one of the two versions still worked on the day someone was away.
An asset counts if the work reaches it without anyone deciding to reach for it, and it does not count if it only functions when a person remembers it exists. A document nobody opens, a skill no agent loads, a metric nobody reads: all of it is weight and none of it is momentum. The check is whether anything downstream would notice the artifact’s absence within a week.
A Heavier Wheel Is Harder to Start
The reason these assets so rarely get begun is honest rather than lazy. A heavier wheel takes more to get moving, and adding mass genuinely slows you down for months, right up until it is the only thing carrying you.
Most of that cost is coordination. To build an evaluation suite you first have to get people to agree on what a good output is. That is slower and more political than shipping the feature would have been.
It feels like overhead around the real work, but the agreement about what good means is the asset. You pay once, in meetings, for a judgment that afterwards executes for free. That is also why the first kilogram is the expensive one: the second asset inherits an argument that has already been had.
The Loop Nobody Closes
Mass only accumulates if each pass through the work feeds the next one, and often the loop does not close for structural reasons.
Take a recommendation model. One team generates the recommendations, another serves them in the product, users click or they do not, and a fourth group builds the dashboards that score how well the recommendations did. Every piece has an owner and is run well.
The step where the scored result changes what gets generated next belongs to nobody. It sits between two teams rather than inside either, so it appears on no plan, and what should have been a mechanism becomes a slide in a quarterly review.
That is why the evaluation corpus everybody agrees would be valuable never appears bottom-up. Somebody standing above all the owners has to hand the closing step to a person by name.
Where This Bites
The most common failure has nothing to do with taking this model too seriously. It is a team that only ever pushes. Every quarter is features, nobody is allowed near the bearings, and the debt compounds until a change that should take a day takes three weeks. The wheel is barely moving while everyone shoves as hard as they can. Pushes are the only thing that turns the wheel, and pushes alone will eventually stop it.
The opposite failure comes from taking the model too enthusiastically: the startup that builds the platform, the abstraction layer and the internal tooling before there is a product anyone wants, and ends up with a beautifully engineered thing that has never turned. Mass is only momentum once something is moving. Both mistakes look like diligence from the inside, which is the whole reason to have a name for each part.
It is worth being straight about the third part: mass is the one you can do without. Plenty of good companies never deliberately add any, and as long as the wheel is sound and someone keeps pushing, they run well for years. Nobody fails for want of a flywheel. It is closer to what separates a good company from a great one, which is a real thing to want and a bad thing to treat as an emergency.
Which One You Need
The model earns its keep as a diagnosis, and the symptom usually tells you which part to reach for.
If everything runs smoothly and the only real problem is that there is more work than there are people, you probably need mass. Build the thing that keeps turning when nobody is pushing. If nothing is technically difficult and yet everything is slow, if the estimates keep coming back larger than the work sounds, the bearings need attention. Pushing harder just produces heat.
And if the system is genuinely good and very little is coming out of it, you are not pushing enough. That one is worth sitting with. The missing pushes are often not engineering at all: getting the product in front of people, finding more users, collecting the data the model needs. Basic, unglamorous work that a strong engineering organization can spend years being too sophisticated to do.
The model does suggest one thing to do on Monday. Take the current plan and label every line on it as a push, a bearing, or mass, then look at what the split actually is. The exercise takes ten minutes and is usually uncomfortable, because most of what the plan calls investment turns out to be pushes with better names. Most weeks the split will come back nearly all push, and most weeks that is correct.
The label does not change the plan. It changes what you expect back, and that turns out to matter more, because the plans that go wrong are rarely the ones with the wrong tasks on them. They are the ones where three different things, with three different shapes of return, were all filed under investment, so nobody noticed that the wheel had been pushed hard all year and was turning no faster than when the year began.
Viktor Stojanov
Head of Engineering at Babbel, writing about engineering leadership in the AI-native era. Builds small Flutter products under Stojanov Ventures.