Recently, while developing SnowLead, I deliberately conducted an experiment that was close to “pure Vibe Coding.” Normally when I build projects myself, I first assess the scale of the requirements, data volume, failure modes, and ongoing maintenance costs before deciding whether to introduce new abstractions and infrastructure. This time I did the opposite: I deliberately put myself in the position of a non-technical founder or small-business owner, simply told Claude Code what I wanted to achieve, and let it read the project, design the solution, create the code, and keep improving it on its own.
The result was interesting. A data-import process in SnowLead that was originally very simple — I estimated the core code could be finished in about 500 lines — was expanded by Claude Code into nearly 10,000 lines. Thread pools, task scheduling, Retry, exponential backoff, concurrency control, graceful shutdown, state management, and a large amount of Manager, Executor, Policy, and configuration code all appeared.
None of these techniques are wrong in themselves. The real question is whether the current project actually needs them. The risk of Vibe Coding is not only whether the AI will write incorrect code. A more subtle problem is that the AI can easily turn a simple requirement into an engineering system that far exceeds the current scale of the business.
This experience made me realize that one of the biggest risks of AI Coding is not necessarily “poorly written code,” but that complexity has become too cheap to generate. In the past, if a developer wanted to add a thread pool, task scheduling, Retry, a state machine, and various abstractions, they actually had to spend time designing and implementing them. Now a Coding Agent can generate all of these things very quickly, so “we might need it in the future” easily becomes a reason to keep adding more code.
The Requirements for SnowLead Were Originally Very Straightforward
The module that went wrong this time was the data-source import module of SnowLead. In the future SnowLead will need to connect to company data from different sources, and the fields, formats, and calling methods of those sources may be completely different. To avoid having to rewrite large amounts of Java business code every time a new source is added, I wanted to put the concrete scraping and transformation logic into Lua scripts, leaving Java mainly responsible for the runtime environment, database access, and basic process control.
The core flow is actually very simple: external data source → Lua scraping script → temporary data table → Lua transformation script → formal business table. The scraping script is responsible for obtaining data from the external source and first saving the raw results. The transformation script then unifies the data from different sources into SnowLead’s own business structure — company name, address, contact information, and other fields — and finally writes it into the formal business table.
The real problems that needed to be solved in the first stage were also quite limited: Was the data fetched? Was the temporary data saved? Did the transformation succeed? Was the formal data written? If any step failed, could the reason be seen and the process re-executed? As long as these questions are answered, the first version already has practical value.
Key reminder
What the first version really needs to solve is simply moving data stably from the external source into the formal business structure. A short flow with clear responsibilities, where every step can be observed and troubleshot independently, is the state that early-stage products should pursue.

But the AI Quickly Started Solving Problems That Had Not Yet Occurred
Claude Code’s thinking was not obviously wrong. When it saw an external API it began to consider request latency and failure cases; when it saw task execution it began to consider asynchrony and concurrency; when it saw the possibility of connecting multiple data sources in the future it began to consider task isolation, scheduling, and scalability.
So a thread pool appeared. Once the thread pool appeared, it naturally became necessary to consider task submission, queues, thread count, cancellation, shutdown, and exception handling. External requests might fail, so Retry and exponential backoff appeared; once Retry was introduced, error types, retry counts, and task states also needed to be managed. Further on, task execution, script running, data scraping, state recording, exception handling, and lifecycle gradually formed their own components and abstractions.
The original flow was: scrape → temporary save → transform → formal save. In the end it gradually became: scheduling → task management → Executor → thread pool → Worker → Retry → state management → Lua Runner → temporary data → transformation pipeline → formal data. With every additional layer the AI could give a reason that sounded reasonable.
The real question is not whether those reasons make sense, but: Why do these problems have to be solved today?
“We Might Need It Later” Is the Phrase Most Likely to Make a System Bloat
In software development, “we might need it later” is very common. We might need higher concurrency later, we might connect more data sources later, we might need automatic task recovery later, we might need horizontal scaling later, we might need more complex scheduling later. All of these things can of course happen.
But if a project is designed for every possible future situation, any simple system can ultimately become infinitely complex. For a mature large-scale platform, building certain capabilities in advance is reasonable because the business scale, user volume, and growth path are already relatively clear. For an early-stage product the situation is usually completely different.
Today’s product direction may change in three months; the extension interfaces designed today may never have a second implementation; the high-concurrency capacity reserved today may not be used for years. In that case, so-called “designing for the future” is very likely just pulling today’s maintenance costs forward.
“We might need it in the future” does not equal “it is worth implementing now.” The truly important engineering judgment is not whether more edge cases can be imagined, but whether one can judge which problems are worth solving today and which should wait until real pressure appears.
The Most Dangerous Part Is That These Pieces of Code Often Look Very Professional
If the AI writes an obvious bug, it is relatively easy to discover. Test failures, interface errors, database write exceptions, or a page that will not open at least give clear feedback. Even someone who cannot write code can know that the system has a problem.
Over-engineering is completely different. It may run perfectly normally, even very stably. You open the thread-pool configuration and see nothing obviously wrong; you look at the Retry logic and it matches common engineering practice; task states, exception handling, and graceful shutdown all look complete. So a strange situation easily arises: every individual part has no obvious error, yet the whole system has already gone off course.
Because the real question is not “Is a thread pool a good thing?” but “Does this project currently need to bear the complexity that a thread pool brings?” Likewise, message queues, caches, complex abstractions, distributed locks, and various fault-tolerance mechanisms are not problems in themselves. The key is always whether they correspond to real business pressure that currently exists.
| Dimension | Simple implementation (~500 lines) | Over-engineered (~10,000 lines) |
|---|---|---|
| Core flow | Scrape → temporary save → transform → formal save | Scheduling → task management → thread pool → Retry → state machine → multi-layer abstractions |
| Troubleshooting difficulty | Linear path, locate in minutes | Fault tree expands, requires understanding multiple layers of components |
| Maintenance cost | Low, easy for newcomers / AI to take over | High, large amount of context needed before any change |
| Suitable scenarios | Early validation, small data volume | Already-validated high-concurrency production environments |
Turning 500 Lines into 10,000 Lines Does Not Add Functionality — It Adds Maintenance Cost
If AI writing code has almost no marginal cost, it is easy to fall into the illusion that writing a bit more does not matter. But the cost of code is never incurred only at the moment it is first generated. Later it still needs to be read, modified, tested, upgraded, debugged, and handed over. The more complex a module is, the more background must be understood for every subsequent change.
Suppose one day SnowLead has a very simple problem: why did today’s scraped data not enter the formal business table? If the system remains simple, the troubleshooting path is clear: first look at the scraping result, then the temporary data, then the transformation, and finally the formal write. But if a complete task system already exists in the middle, the problem immediately expands. Was the task submitted? Did the Worker execute? Is the thread pool blocked? Is it currently Retrying? Has the state been updated? Did Lua actually run? Was the exception wrapped at some layer? Did the transaction actually commit?
An originally linear path turns into a fault tree. That is the real cost of turning 500 lines into 10,000 lines. It is not hard-disk space or the size of the Git repository, but: how many things does a developer need to understand before they dare to safely change this piece of code? For a small team this is especially important, because the person who takes over in the future may not be the original author, but another developer, an outsourcing team, or even another AI Agent.
Why Is Vibe Coding Especially Prone to Making Small Projects “Big”?
Coding Agents are very good at continuing to expand requirements. They proactively read the project, fill in edge cases, discover potential risks, and then conveniently improve the architecture. This ability is valuable in itself, but without clear constraints it can easily lead to scope creep.
You only say “help me make a data import,” and when the AI sees an external API it thinks of failure; thinking of failure it thinks of Retry; thinking of task execution it thinks of concurrency; thinking of concurrency it thinks of thread management. Each of these inferences looks fine on its own. What is truly missing is another set of information: How much data does this system process per day right now? How many users are there? How many data sources? Has the product been running stably for years, or is it just starting validation? Does the team have dozens of engineers, or only one or two people maintaining it?
These background facts directly determine whether the same technique is mature design or unnecessary burden. If the AI is not actively given scale and complexity boundaries, it can easily continue designing according to an imagined system that is larger than the real business.
AI Writing Code Should Also “Go from Thin to Thick, Then from Thick to Thin”
There used to be a saying about reading books: first read the book from thin to thick, then from thick to thin. The first time you read, you constantly encounter new concepts, background, and details, so a book that was originally thin becomes thicker and thicker in your mind. Once you truly understand it, you begin to know which content is core, which is only expansion, and which can be merged, and finally you read it thin again.
I increasingly feel that AI writing code should go through a similar process. Vibe Coding is well suited to completing the first step of “from thin to thick.” At the beginning there may be only one sentence of requirement: obtain company data from an external data source and transform it into a format that SnowLead can use. The AI can quickly expand it into code, supplementing data structures, error handling, tests, and various implementation details. The first version that used to take a developer a long time to build can now be generated very quickly.
The problem is that many people stop here. The code is generated, the tests pass, the page opens, so it is deployed directly. But the real software-engineering work may precisely begin from this point. The second step should be to read the code from thick back to thin. Re-examine the entire implementation and judge which code truly serves the current business and which is only for theoretical future expansion; which exception handling is genuinely worth keeping and which infrastructure was only added by the AI according to “best practices.”
Taking the SnowLead case as an example, Claude Code first expanded a simple requirement to nearly 10,000 lines. That does not mean all 10,000 lines have no value. There may be reasonable data structures, logging, tests, exception handling, and real problems I had not considered at the beginning. What should really be done is a second-pass Review. Take out the thread pool, complex task management, premature abstractions, and mechanisms that are currently completely unused one by one and ask: Does it solve a problem that already exists, or a future problem imagined by the AI? If it is truly needed, keep it. If not, delete it. The code that remains may be more complete than the originally imagined 500 lines, but there is absolutely no need for it to approach 10,000 lines.

In the AI Era, What May Be Truly Scarce Is Not Generating Code, but Deleting Code
In the past, over-engineering at least had a natural resistance: writing these things took time. If a developer had to create dozens of classes, a whole thread-management system, a state system, and Retry mechanisms for a small feature, they would at least reconsider whether it was worth it because of the workload. That resistance is now disappearing rapidly. AI can generate large amounts of code in a very short time, and the surface quality may even be quite good.
So a new problem has appeared in software development: when writing code becomes cheaper and cheaper, who decides which code should not be left at all? That is also why I do not think Vibe Coding itself is the wrong direction. On the contrary, it is well suited for rapid exploration, building the first version, and validating ideas, and it can also help developers discover problems they had not originally considered. The real danger is treating “generation finished” as “engineering finished.”
A healthier workflow should be: first Vibe Coding to write the requirement thick; then do a Review to read the system thin. AI is responsible for rapidly expanding possibilities; humans are responsible for judging which complexity is truly worth entering the product. At this point the goal of Code Review also changes. In the past we mainly checked for bugs, security issues, performance problems, and style issues. Now we should add one more question: Is there code here that simply does not need to exist? Sometimes the most valuable result of a good code review is not adding another component, but discovering that these three layers can actually all be deleted.
Conclusion: AI Can Write Code Thick, but It Will Not Automatically Help You Read the System Thin
The data-import module of SnowLead this time ultimately brought me back to the original question. What it really needed to accomplish had never changed: scrape data → temporary save → transform → formal save. Claude Code can expand this requirement very quickly, considering various failure cases, expansion methods, and engineering details. This ability itself is very valuable. But generation should not be the final step.
After AI has written the code from thin to thick, someone still needs to re-examine it, delete complexity that has no real basis, merge repeated abstractions, and put infrastructure that may be needed in the future but is completely unused today back into the future. What remains in the end is not necessarily the least code. It should be: code that is just enough to solve the current problem.
In the past, the hard part of writing software was getting things built. In the AI era, another increasingly important ability is knowing what should not be left behind. So this experiment has not made me reject Vibe Coding. On the contrary, I will still use it to generate the first version quickly. The workflow just needs to be a bit more complete: first write the code thick, then read the system thin. AI can be responsible for rapidly expanding possibilities. Ultimately deciding which complexity is worth entering the product remains one of the most important judgments in software engineering.

