In CD Office Hours episode 9, we discussed the compliance ratchet and how accelerating software development without the right practices can lead to an irreversible increase in controls. In this episode, we look at a classic software problem: the lure of the software rewrite and the inevitable hazards such endeavors bring.
When the cost to maintain the software and add new features becomes burdensome, a decision is often made to rewrite it rather than improve it “in place”. This is sometimes a valid approach, but more often it’s based on an exaggeration of how bad the code is and an understatement of the work involved.
In some cases, the qualities of the codebase that a team dislikes stem from an architectural decision they don’t understand, either because there’s no record of it or because the records aren’t put to use. In many cases, the layers of change applied to the code over time have made it messy and hard to comprehend.
When considering a big-bang rewrite, it is practically impossible to determine whether the rewrite is required or will provide a return on investment. That’s why big-bang rewrites have such a high failure rate and should be avoided.
Watch the episode
Legacy or heritage
The term “legacy” has become associated with old and unusable software, but more often, legacy software is the long-running backbone of an organization. It captures vast institutional knowledge and endures for so long because it works. In many cases, software inevitably becomes legacy software, as the alternative is that it fails.
That’s why we prefer to use the term “heritage” software, as it’s an artifact of special value to the organization. The code for heritage software has captured not just the essence of the solution, but hundreds or even thousands of edge cases and enhancements based on years of learning.
The hazards of the rewrite
Most codebases have parts that are untidy or don’t make sense. There’s a tipping point that makes us want to rewrite the whole thing, but this is often the wrong reaction. It’s very rare for the cost of maintenance to exceed the cost of wholesale replacement.
When teams consider a big-bang rewrite, they tend to estimate the task based on the obvious use cases. That means the new software is naturally cleaner, because it doesn’t handle any of the less obvious problems. When the new software version goes live, the team begins rediscovering the many edge cases that must be handled.
By the time these edge cases have been incorporated into the software, it has started to look complicated once more, because the complexity came from the domain rather than a careless implementation. The team may find more graceful ways to handle edge cases, but by this time, the rewrite will have vastly exceeded all estimates of effort and expense.
When the rewrite is complete, the organization will effectively be “back to square one”, which means they will have fallen behind the competition. That means the anticipated cost of maintenance and new features must be substantially lower to allow the software to catch up and pass competitor offerings.
Essentially, wholesale rewrites should be avoided; a different approach is needed.
Architectural and team strategies
The alternative to the big-bang rewrite is to identify a specific component to replace. You might take a calculation routine and move it out of a monolith into a service, creating the missing tests and refining the routines as you extract it. The result is a new, highly maintainable component and a slightly smaller monolith that calls it rather than doing the work itself.
This is the strangler fig pattern, named for the plants that cheat the competition for sunlight by growing high up on a host tree, eventually sending down their own roots and often remaining after the host tree has died and rotted away. In the same way, the heritage software can continue to support users while an increasing number of components are rewritten and deleted from the monolith.
You can use both domain-driven design and team topologies to inform your gradual rewrite. If you can identify a seam that would carve off a meaningful part of the domain and have a team that can own the domain and the component that serves it, you can create a meaningful interface for the software and improve the organization’s communication structure.
Combining the strangler fig pattern with architectural and team design results in teams with deep knowledge of their specific areas and the ability to move with relative independence. The number of teams you have can be instructive when considering how far you want to break down a monolith into cohesive components.
Wholesale rewrites invite disaster, but working in small steps and considering organizational design as well as architectural design is the path to success.









