Available for Hire
ZB

Memuat...

Back to Blog
DevOps 6 min read · 1100 words Featured

Auditing Legacy Code Before Deciding on a Rewrite

Rewriting an old system is tempting and almost always costs more than expected. How to examine legacy code in a structured way so the decision rests on evidence rather than irritation.

#legacy #audit #security #refactoring

Every developer who inherits an old system has felt it: one hour into reading the code, the urge to rewrite it from scratch arrives. That feeling is almost always wrong, but sometimes it is right. The difficulty is that telling the two apart requires evidence, not impressions.

This is about examining a legacy system in a structured way, so your recommendation can survive contact with the person holding the budget.

Why a Rewrite Usually Loses

A messy old system is ugly, but it has one advantage that gets underestimated: it has already survived years of real cases. Every strange condition that ever occurred has a handler somewhere in there, often in the form of an if that makes no sense until you know the story behind it.

A rewrite starts from zero, including zero experience. You will rediscover every one of those edge cases one at a time, in production, with real users as your testers.

So the starting question is not “should we rewrite this”, it is “what is actually broken, how badly, and can it be fixed incrementally”.

Start with Security, Not Aesthetics

Ugly code makes development slow. A security hole leaks data. Both are problems, but only one can end an organisation overnight.

Look first at the things whose consequences cannot be undone:

SQL injection. Look for queries assembled through string concatenation. In legacy PHP the pattern usually looks like this:

$q = "SELECT * FROM users WHERE email = '" . $_POST['email'] . "'";

A single one of these is enough to read the entire database.

Password storage. Search for md5( or sha1( near anything called password. Neither is a password hashing function, and both can be broken with off-the-shelf tables. The correct choices are bcrypt, scrypt, or Argon2.

Credentials in the repository. Search the Git history, not just the current state. A secret that ever entered history can still be recovered even though the file is gone.

Authorisation that only exists in the interface. Hiding a delete button from ordinary users is not access control. Check whether the endpoint itself verifies the role, or whether it merely relies on users not knowing the address.

File uploads. Are extensions validated? Are files stored in a directory that can execute them? An image upload that turns out to be able to run PHP is a shortcut straight to the server.

For each finding, record three things: which file it is in, what the impact is if exploited, and how hard it is to exploit. Without all three, your finding reads as a complaint.

Measure, Do Not Just Feel

“The code is a mess” can neither be disputed nor proven. Numbers can.

A few measures that are cheap to obtain and easy to explain to non-technical people:

  • Lines per file, and how many files exceed a thousand
  • Code duplication, what percentage of the repository is essentially a copy
  • Branch depth, a function with five levels of nested if is a warning sign
  • Test coverage, and if it is zero, say zero
  • Number of dependencies that are unmaintained or carry known vulnerabilities
  • How long it takes from changing one line to seeing the result

That last one is often the most persuasive. If changing a single label requires a forty-minute manual deploy, that is a cost you can express per month.

Tools such as PHPStan, SonarQube, or npm audit can supply some of these numbers without you reading a single file.

Map What Is Actually Used

Legacy systems are almost always smaller than they look. Many menus have not been clicked by anyone in a long time.

Before estimating effort, find out which parts are alive. Server access logs are the cheapest source. Analytics helps if it exists. If neither is available, ask the users directly: which features do you use every day, every month, and never.

The results are often surprising. The module that looks most frightening in the code turns out to be used twice a year, while the one used daily is straightforward. That changes priorities dramatically.

Write the Findings for Two Readers

An audit report only a programmer can read will stop at a programmer’s desk. The person who decides the budget is usually not a programmer.

Structure it in two layers. A summary up front covering the current state, the biggest risks, and the options with estimated cost. Technical detail behind it, complete with file locations and reproduction steps, for whoever will do the work.

For risks, avoid adjectives. Not “security is weak”, but “all user data including phone numbers can be retrieved by anyone without logging in, through a single request to page X”.

Three Options, Not Two

The question “fix or rewrite” is a false choice. There is nearly always a third path.

Fix in place. Appropriate when the architecture still makes sense and the problems are localised. Cheapest and least risky.

Strangle incrementally. Build the new system alongside, move one module at a time, and put a proxy in front so users do not notice the difference. Slower, but the old system keeps running throughout and every step can be reversed.

Full rewrite. Only makes sense when the technology is genuinely unsupported, or the business requirements have changed so completely that the old system answers the wrong question.

Present all three with their time estimates and risks. Letting stakeholders choose between three options is far more likely to produce a decision than presenting a single demand.

Whatever You Decide, Put the Net Up First

If the outcome is incremental repair, do not touch the code yet. Put in place the things that make change safe:

  1. Backups whose restore path has actually been tested
  2. Characterisation tests on the most critical flows, simply locking in current behaviour as it is
  3. Error monitoring, so you learn something broke before a user phones in
  4. A staging environment whose data resembles production

Characterisation tests feel odd because you are writing tests for behaviour that may be wrong. But the goal is not to prove the code correct, it is to make sure your change did not alter anything you did not intend to alter.

Closing

A good audit turns the conversation from taste into evidence. Once there is a concrete list of risks, numbers that can be compared, and three options with their consequences, the decision usually becomes obvious on its own. And not infrequently the conclusion is not a rewrite at all, but three weeks of focused repair that clears eighty percent of the complaints.

Share this article:

Enjoyed this article?

0 reactions