How long does it usually take to investigate a user registration failure?
If it is only a form validation error, the cause may be found quickly. But in a real project that has been running for some time, gone through multiple changes, and been handed over between different people, an ordinary-looking registration issue can eventually lead all the way to the database, third-party authentication, legacy user data, the permissions system, and even the entire deployment process.
That was exactly what happened during this project takeover. The original task was not even to “fix the authentication system,” but simply to determine whether one of two seemingly duplicate registration pages could be safely removed.
On the surface, it was a registration bug. The real problem could involve the entire user identity chain.
This article is based on an anonymised real project case. To protect the client’s privacy, the client name, specific business information, service provider names, and sensitive code have been removed, and some technical details have been simplified.
TL;DR: What Did This Bug Eventually Reveal?
The problem initially seemed very simple: both registration pages reported that registration had succeeded, but users created through one of the pages could not be found in the admin panel. Further investigation revealed that the two pages were actually using completely different user systems—one wrote users into the local database, while the other directly created accounts in a third-party authentication service.
More importantly, this was not the only inconsistency. Legacy code, historical users, permission checks, and the admin system relied on different user identifiers, while the project also lacked a separate testing environment. As a result, a small registration problem eventually exposed technical debt across the entire identity chain.
| What appeared to be the problem | What further investigation revealed |
|---|---|
| Two duplicate registration pages | They were actually connected to two different user systems |
| Registration succeeded, but the user could not be found in the admin panel | The third-party identity existed, but the local user might not |
| Some users could not log in | New and legacy users used different identity mappings |
| Permission issues occurred occasionally | Different modules used different User IDs |
| Everything worked locally | Production contained historical data and different configurations |
| The old page seemed safe to delete | Deleting it could prevent a group of legacy users from logging in |
The most typical part of this case was that each individual module still appeared to work, but once they were combined, nobody could clearly explain what was supposed to happen during a complete registration flow. This type of problem is often harder to investigate than a function that throws an obvious exception, because the real issue is not that “one part is broken,” but that multiple parts have gradually become inconsistent.
At First, It Was Just Two Registration Pages That Looked the Same
When I first took over the project, it had already been running in production for some time, but the overall system was not stable. Registration, login, and admin management occasionally had problems, and there were also some obvious legacy features left in the codebase.
The original developer told me that the project had two registration pages that seemed to do the same thing and asked me to confirm whether one could be removed. Based on the task description, this did not even sound like a bug. It looked more like routine code cleanup.
But for a system that already has real users, I would not delete one page simply because two pages “look the same.” This is especially true for registration, login, payments, and permissions, where similar UI does not necessarily mean the underlying data flow is also the same.
So I created two test accounts, one through each registration page. Both pages redirected normally and displayed a successful registration message, with no obvious errors.
I then opened the admin system and checked the two users I had just created. Only one of the accounts appeared. The other account had supposedly been registered successfully, yet it could not be found at all.
At that point, the question changed from “Which page can be deleted?” to “Where exactly was this user registered?”
Behind the Two Registration Pages Were Actually Two User Systems
After checking the APIs, database, and authentication code, the problem quickly became clear. Although both pages were called “Register” and collected almost the same information, they did not call the same registration flow.
One page created a user record in the project’s own database. The other directly called a third-party authentication service and created an Authentication Identity in an external system. In other words, the product appeared to have one registration system, but in reality it maintained two different answers to the question of “Who is this user?”
| Registration entry | Where the user was created | Primary identity | Directly visible in admin? |
| Registration Page A | Local business database | Local User ID | Yes |
| Registration Page B | Third-party authentication service | External Identity ID | Not necessarily |
| Some legacy flows | Mixed | Mapping / Legacy ID | Depends on legacy logic |
This meant that an account being successfully authenticated by the third-party service did not necessarily mean that the project’s own business database contained that user. Conversely, having a User Record in the local database did not guarantee that the account could log in correctly through the external authentication service.
The dangerous part was not the use of third-party authentication itself. The real issue was that the system did not clearly define which side was the Source of Truth for identity. As long as that boundary remained unclear, permissions, admin management, and business data could easily begin depending on different identity sources.

When the System Has Two User IDs, the Problem Starts to Spread
As the investigation continued, I found that the two identity systems were not completely isolated from each other. The project already contained a considerable amount of legacy compatibility code that attempted to connect the Local User ID with the Third-party Identity ID.
Some features read local users, some APIs directly trusted the identity returned by the third-party authentication service, and some legacy code searched both sides depending on the situation. The system therefore ended up with multiple answers to the same question: “Who is this user?”
At that point, one user could exist in many different states. For example, the third-party account might have been created successfully while the local user did not exist, or both records might exist but the mapping between them might be incorrect.
The symptoms also became increasingly confusing. Some users registered successfully but could not log in, some could be identified on the frontend but could not be found in the admin panel, and some accounts even received different permission results in different parts of the system.
| Identity state | What the user might experience |
| Third-party account exists, local user does not | Authentication succeeds, but entering the application fails |
| Local user exists, third-party account does not | Visible in admin, but unable to log in normally |
| Both exist, but the mapping is incorrect | User is connected to the wrong or incomplete business identity |
| New and legacy IDs are mixed | Different modules produce inconsistent permission results |
| Historical data is missing newer fields | Legacy users work, but newer features fail |
On the surface, these issues belonged to four different areas: registration, login, admin management, and permissions. In reality, they were different symptoms of the same underlying architectural problem.
If the system does not have one clear answer to “Who is this user?”, then registration, login, permissions, and business data are all difficult to keep stable.
The Real Problem Was Not Third-Party Authentication, but Unclear System Boundaries
It is easy to misunderstand this situation and conclude that the project should not have used third-party authentication at all. That is not the case. Modern software commonly uses services such as Auth0, Amazon Cognito, Firebase Authentication, Keycloak, and other identity platforms.
Delegating password handling, secure authentication, OAuth, and similar capabilities to specialised services can reduce risk. What really needs to be designed clearly is the relationship between Authentication and the Application User.
A clearer architecture normally separates the two responsibilities. The third-party identity platform is responsible for proving “who this person is and whether they have successfully logged in,” while the project’s own database stores “who this person is within our business system, what permissions they have, and what business data they are associated with.”
The two sides can then be connected through a stable and unique Identity ID. Even if the authentication platform and the business database are separate systems, developers can still clearly understand whether the relationship is one-to-one or follows another explicitly defined model.
The real problem begins when some code assumes that “the third-party user is the user,” while other code assumes that “the database record is the user,” and more and more compatibility logic is added to connect the two designs. As historical data accumulates, this kind of system usually becomes harder to maintain rather than naturally improving over time.
During the broader project review, I also found that authentication logic was not concentrated in one clearly defined module. It was scattered across frontend pages, APIs, database operations, and legacy compatibility code, while some early debugging entry points could still affect the production system even though they were no longer part of the primary flow.
At the same time, the available logging was not sufficient to reconstruct an identity request from beginning to end. When a login failed, it was difficult to quickly determine whether third-party authentication had failed, the local user did not exist, the Identity Mapping was incorrect, permissions were missing, or a piece of Legacy Logic had been triggered unexpectedly.
This is why a seemingly simple bug can become deeper the further you investigate it. The most time-consuming part is often not fixing one line of code, but first reconstructing how the system actually works today.
Without a Testing Environment, Hidden Problems Reached Production More Easily
Beyond the identity problem, the investigation also exposed another major risk: the project did not have a separate Staging Environment. Developers normally made changes locally, confirmed that the feature appeared to work, merged the code directly into main, and then automatically deployed it to production.
This approach does not mean every Deployment will immediately fail. The real issue is that code working locally only proves that it works with the developer’s current data, environment variables, and configuration. It does not prove that the behaviour will be identical in the real production environment.
Production already contained a large amount of historical user data, and those users may have come from different versions of the registration flow. Some records had gone through migrations, some had been created by legacy code, and others may even have been modified manually.
As a result, a feature that worked perfectly for a newly created test account could immediately fail when it encountered a real user created several months earlier. Historical data had effectively become part of the system’s behaviour, and a local development environment could not naturally reproduce all of those states.
The most dangerous part of this process was that many changes received their first complete validation only after they reached production. As long as the page did not fail immediately, problems hidden in legacy data, third-party services, permission mappings, and environment configuration could remain there until a real user eventually triggered them.
For a system already in production, the purpose of Staging is not simply to provide “one more deployment environment.” Its real value is to provide a production-like space where developers can run the complete registration, login, admin, permissions, and data migration flows without affecting real customers.
Why “The Page Says Success” Is Far From Enough
Many MVPs and early-stage projects use a very direct way to validate features. Fill in the form, click the button, and if the page does not show an error, the feature is considered complete.
For a completely isolated feature, that can sometimes be enough. But registration, payments, permissions, file processing, Webhooks, and third-party synchronisation are usually much more than “one successful request.”
A seemingly ordinary registration process may actually follow a chain like this:
Registration Form → Backend API → Identity Provider → Local User → Identity Mapping → Permissions → Login → Admin
If any one of these steps fails, the overall process may only be partially completed. More importantly, if an earlier operation succeeds while a later operation fails without proper rollback or compensation, the system can be left in a long-term partially completed state.

Suppose a user submits a registration form and the system first calls a third-party Identity Provider to create the account. Once that operation succeeds, the third-party platform has already permanently stored a new identity.
The system then needs to create a local User Record, but this step fails because of a database constraint, network error, or code problem. If the application does not have proper error handling and compensation logic, the registration has effectively succeeded only halfway.
The frontend may still display Registration Successful based on the earlier successful operation. The user therefore believes the account has been created, but when they try to log in, the business system cannot find the corresponding local record.
From the UI perspective, it may simply look like “the page does not work after registration.” From an engineering perspective, it has become a typical Cross-system Data Consistency problem.
A success response can prove that one operation succeeded. It does not necessarily prove that the entire business process succeeded.
Why Not Just Delete One of the Registration Pages?
Once we discovered that the two pages used different systems, the most obvious solution seemed simple: identify the correct registration page and delete the other one.
If this had been a new project with no real users, that might really have taken only a few minutes. But a production system with accumulated historical data cannot be changed based only on “how things should work from today onward.”
The old registration flow may already have created a group of real users. If those accounts exist only in the third-party authentication system, while the new login flow begins depending entirely on the local User Record, then simply deleting the old logic could prevent all of those users from logging in the next time they return.
On the other hand, if both systems are kept permanently for compatibility with historical users, new accounts may continue entering different data paths. What began as a historical legacy problem would then continue creating new legacy problems.
The real questions were therefore not “Which page should stay?” but the deeper issues below:
| Question that needed to be answered | Why it mattered |
| Which system is the Source of Truth for identity? | Ensures there is one final answer to “Who is this user?” |
| How should the third-party Identity map to the local User? | Prevents data drift between the two systems |
| How should historical users be migrated? | Avoids fixing new users while breaking legacy users |
| How should orphaned accounts be handled? | Cleans up data that exists on only one side |
| Which stable ID should permissions use? | Prevents different modules from making different decisions |
| Can the Migration be run repeatedly? | Makes it easier to repair abnormal historical data |
| How should the identity chain be logged? | Makes future problems easier to locate |
These questions were clearly beyond the scope of a single registration page. That is why adding another if statement or another database lookup to the legacy logic could delay the problem, but would not truly solve it.
Final Decision: Redefine the Identity Boundaries Instead of Adding More Patches
After multiple discussions, the team ultimately decided not to keep adding compatibility logic to the existing authentication system. We began redefining the relationships between users, Authentication, Authorization, and business data, and gradually replaced the old implementation through a new frontend and backend architecture.
The goal was not to “over-engineer” a small project. In fact, the real objective was the opposite: reduce the number of special cases in the system and make sure each core question had one clear answer.
Who the user is should come from a clearly defined identity source. How the user logs in should follow a traceable authentication chain, and where permissions come from should also rely on a unified and stable data model.
If these basic boundaries remain undefined, every new feature may require developers to understand or duplicate more legacy compatibility logic. In the end, what slows development down is often not growing business complexity, but the fact that developers become increasingly afraid to change old code.
This is also common when taking over an existing software project. The client may only see a broken button, a user who cannot log in, or one missing record in the admin system, but those surface symptoms do not necessarily reveal the true scope of the problem.
What engineers need to determine is whether the Bug is an isolated implementation error or a symptom of a deeper system problem. If it is a single-point error, a local fix is obviously the most reasonable approach. But if the underlying data boundaries have been inconsistent for a long time, adding more patches may simply create additional risk.
| Bug type | Common approach |
| Form validation error | Local fix |
| Single API implementation error | Code fix + testing |
| Missing data migration | Data repair + Migration |
| Third-party integration failure | Retry, compensation, and monitoring |
| Multiple identity models coexist | Redefine identity boundaries |
| Large amounts of interdependent legacy compatibility logic | Phased refactoring |
Whether a system should be refactored should not be decided by whether the code “looks old.” A more practical question is whether the current architecture can still reliably answer the most basic business questions, and whether the risk of continuing to add features has become greater than the cost of reorganising the core boundaries.
This project also reinforced another point: many of the hardest Production Bugs are not caused by one obviously incorrect line of code. More often, they come from multiple systems, multiple versions, and multiple sets of data gradually becoming inconsistent over a long period of time.
The code may be inconsistent with the database, the local database may be inconsistent with third-party services, and the development environment may be inconsistent with production. Even the state displayed to users on the page may not match the state actually stored by the system.
More importantly, these inconsistencies do not necessarily cause the system to fail immediately. Most users may continue using the product normally, allowing the problem to remain hidden for months until one user happens to meet the specific conditions needed to trigger it.
Technical debt does not always throw an error on the day it is created. Sometimes it simply gets stored in the system, waiting for a real user to trigger it later.
So when a production system develops a “user registration failure,” the registration button is rarely the only thing worth checking. The registration API, third-party identity service, local user table, Identity Mapping, permission data, historical users, logs, environment configuration, and Deployment Process may all be part of the problem.
This is why software maintenance and project takeovers can sometimes begin with what appears to be one small Bug and eventually involve much broader system design. It is not because engineers deliberately make the issue more complicated, but because the symptom seen by the user and the actual source of the problem may be separated by many layers.
For a production system, fixing the current Bug is obviously important. But it is even more important to understand why it reached production, why it was able to remain there for so long, and whether other similar inconsistencies are still waiting to be triggered.
What you see is one failed registration.
What the engineer may need to investigate is the entire identity chain.

