A founder gets in touch. He has already built the thing.
Not a slide deck, not a wireframe—a working enterprise application. Dashboards, chat, documents, timesheets, projects, test management, and an approvals flow. He built it himself, first on one coding tool, then a second when he hit the first one’s ceiling. He has a developer helping part-time. He has a customer waiting to use it. He has four or five more people who have said they will test it.
And he is stuck.
Not stuck in the way founders used to be stuck. He isn’t waiting on a quote, hiring a development team, or starting a six-month build. The product mostly works. He can demo it. But something tells him it is not ready for a real customer.
When he says it out loud, it comes out as a question about time: Could someone finish this in two months? But that is not really the question. The real question is harder: What, exactly, is left between a working prototype and production-ready software?
We have had this conversation enough times now that the shape of it is familiar. What follows is what we have learned to look for—and what we think the practical answer is.
Let us be clear about something, because the sceptics get this wrong.
What these founders have built is genuinely impressive, and it is genuinely theirs. A domain expert using modern coding tools can now get surprisingly far in months, alone, at night, around a day job. That is not a toy. It represents a real shift in how software products can get started.
It works because the person building often knows exactly what they want. They have lived the problem for years. They do not need someone to explain what a resource plan looks like—they have suffered through bad ones. When they describe a screen, they describe it precisely, and modern development tools can turn that knowledge into working functionality quickly.
So the first 80% is not luck. It is domain knowledge finally meeting tools that can keep up with it.
The trouble is what that 80% is made of.
When you build primarily by conversation, you tend to build screen by screen. You ask for a dashboard, you get a dashboard. You ask for a timesheet module, you get a timesheet module. Each feature may work when you look at it individually, but each one may have been created with limited understanding of the decisions made elsewhere in the application.
What you can end up with is not yet a system. It is a collection of features sharing a database.
And the last 20% is not simply 20% more features. It is the work of turning that collection into reliable, secure, testable, and maintainable software.
This distinction matters because development tools are already widely used. Stack Overflow’s 2025 Developer Survey found that 84% of respondents were using or planning to use development tools with generative capabilities, while 46% distrusted the accuracy of their output. The most common frustration, reported by 66%, was getting solutions that were almost right but not quite.
The speed is real. So is the need for verification.
These are the patterns that show up again and again. None of them are particularly exotic, but most are invisible during a polished product demo.
Two screens both answer the same question: “How many leads came from this channel?”
They were built in different sessions, so each got its own query. One excludes archived records; the other doesn’t. For months nobody noticed because nobody put the two numbers side by side.
Then someone does.
We have seen a dashboard say 15 and the list behind it open 8. The cause was a single date boundary interpreted two different ways. This is one reason production readiness depends on more than whether individual screens work—the underlying business rules need to remain consistent across the application.
A date filter takes 2026-10-05 and reads it as midnight. Everything that happened later on the fifth disappears from the report.
Nothing crashes, and nothing necessarily logs an obvious error. The number is simply wrong. These are dangerous bugs because the interface can still look completely normal.
A prompt asks the user to summarise what happened before they can mark a task complete. It was built on the screen where the founder tested it.
The same action on another screen posts straight through with no prompt. Now half the records contain the required field and half do not, and nobody can immediately explain why.
Navigation, feature access, and data visibility may each have their own mechanism because they were added at different times. There can now be two systems for roles, and they disagree.
Someone is granted access but still cannot see the screen because the permission landed in one system while the navigation reads another. This becomes much more serious when the product moves from a founder’s test account to multiple real users, teams, or client organizations.
A change was applied manually in one environment and through a migration in another. The migration later fails where it lacks permission to run, and a log entry is created that nobody notices.
The two databases slowly move apart. Now a problem occurs in production but cannot be reproduced reliably in staging.
This is exactly the kind of issue that requires architecture, database, and deployment review—not another interface fix.
This is the big one. When every change starts as a conversation, automated tests may never become part of the process unless someone deliberately adds them.
A fix to one screen can therefore break another. The founder in our opening call said it himself, almost in passing: you touch one thing somewhere, it touches something else.
He was right. What he lacked was a reliable way to know when that happened.
This is where structured automated testing vs manual testing becomes important. Regression testing can repeatedly check known workflows after changes, while manual and exploratory testing remain useful for situations that require human judgment.
Six months later, the founder cannot remember why a particular rule exists. The reasoning may have lived only in an earlier conversation and was never recorded anywhere else.
The code still tells you what happens. It may not tell you why.
The next developer now has to infer the original business decision from the implementation. That is technical debt before the product has even properly launched.
A useful way to understand the last 20% is to compare what matters during a demo with what matters after customers arrive.
| Area | Working prototype | Production-ready product |
| Features | Main workflows work | Workflows work together consistently |
| Architecture | Enough to make the product run | Dependencies and data flows are understood |
| Business logic | May exist in several places | Shared rules have a clear source of truth |
| Validation | Works on tested paths | Consistent across every entry point |
| Permissions | Basic access control | Roles, features, and data visibility agree |
| Database | Works with current data | Migrations and environments remain consistent |
| Testing | Founder checks key screens | Critical workflows have repeatable test coverage |
| Deployment | New version can be deployed | Releases are repeatable and verifiable |
| Monitoring | Problems are found manually | Important failures can be detected and investigated |
| Documentation | Code explains implementation | Important decisions and reasons are recorded |
| Customer readiness | Good enough to demonstrate | Reliable enough for real workflows and real data |
This is why estimating the final stage by counting unfinished screens usually gives the wrong answer. The remaining work is horizontal—it cuts across the whole product.
Security is another area founders should not leave until after launch.
Recent research into software produced with coding agents shows why functional testing alone is not enough. One 2025 benchmark evaluated 200 real-world software engineering tasks selected around vulnerable implementation patterns. In one tested configuration, 61% of solutions were functionally correct, while only 10.5% were considered secure.
That does not mean every vibe-coded application is insecure. It shows something more useful: functional correctness and security are different tests.
A login can work and still be insecure. An API can return the correct data and still expose records it should not. A file upload can succeed and still accept something dangerous.
Before real customer data enters the product, the review should cover authentication, authorization, secrets, input validation, APIs, file handling, tenant isolation where applicable, logging, backups, and third-party dependencies.
For a larger SaaS or enterprise product, these areas should form part of the wider web application development and production-readiness process rather than being treated as a final checkbox.
None of this means stop using modern coding tools to build. We build this way ourselves. The answer is not to throw away the speed that helped create the product.
The answer is to add a handful of engineering disciplines that make fast-built software hold together.
Six rules do most of the work.
Before touching a line, trace how the data actually flows—which screen writes it, which service reads it, which database table stores it, which scheduled job changes it, and what other screens depend on the same value.
Searching the code and finding nothing is not evidence that nothing is there. Sometimes it simply means you searched for the wrong name or looked at the wrong layer. Most damage comes from a confident edit made before this step.
“The code is deployed” is not the same as “the feature works.”
Query the database. Open the actual page. Run the workflow. Check the API response. Confirm that the built file contains the intended change.
We have watched a correct fix sit in a deployed file and do nothing because another rule further down the stylesheet quietly overrode it. Nothing is finished until it is confirmed where the result actually lands.
When you discover that a date boundary is wrong in one place, search for the same pattern throughout the application.
If the underlying problem exists in four places, fix the class of problem rather than only the reported instance. Otherwise, the bug does not disappear—it comes back later under a different ticket.
Anything that can affect a customer, payment, permission, or important business record deserves an appropriate review point. Do not review only the code; review the outcome.
If a generated customer message requires approval, send it to an approvals queue and assign responsibility to a named person. The same principle applies beyond generated content: production systems need clear ownership for actions that carry business risk.
Restarting a service during a scheduled job window, changing a shared constant that three features read, editing a permission that also controls navigation, or changing a database field used by reports can all have wider effects.
Before changing something, ask: What else depends on this?
That question costs seconds and can save hours.
The code already says what it does. What future developers need is the reason behind important decisions.
Why does this user type have different access? Why is this calculation based on creation date rather than payment date? Why is this integration retried three times?
Put that reasoning in the commit, ticket, comment, or handover documentation. This is one of the cheapest ways to make a product easier to maintain.
Notice that none of these rules requires a particular framework or programming language. They are engineering habits, and they are a large part of the difference between software that demos well and software that continues working when real customers start using it.
Before estimating “how long until launch?”, review the product systematically. The assessment should cover the areas that can affect a real customer once the application moves beyond a controlled demo:
For existing business-critical products, these concerns continue after launch. Application Management Services can cover monitoring, incident resolution, performance, security, deployment support, and ongoing application improvement rather than treating launch as the end of development.
There is a second failure mode, and it has nothing to do with code.
The founder in our opening call had built a product before—a different product, years ago, in a different market. It went wide. He took on more customers than he could support, every one of them wanted something slightly different, and he spent his life adjusting.
He said it plainly: he would not make that mistake twice.
That instinct is correct.
When you finally have something that works, the pull is to show everyone. Four people are waiting to test. Each one feels like validation, but every additional early customer increases the surface you have to support while the product is still changing.
Each customer will use it a little differently, creating bugs, requests, and edge cases faster than a small team may be able to absorb them.
One customer. Tested properly. Until it genuinely works. Then the second.
Two things make that first engagement more manageable.
Not as an apology—as context. They may encounter issues, and you will be available to investigate them quickly.
A first customer who understands the stage of the product can provide much more useful feedback than someone expecting mature software from day one.
Ask of every module: Does the first customer need this to receive the core value on day one?
Often, the answer is no.
Finance, resource planning, advanced reporting, and secondary administration modules are common places where scope expands because they connect to many other workflows. Ship a lighter version or leave them out of the first release.
Scope that stays cut is what protects the launch date.
The purpose of MVP development is not to fit every future requirement into version one. It is to build enough of the product to validate the core workflow, gather real feedback, and decide what should come next.
A founder tests whether the product behaves as intended. A customer tests whether it survives reality.
Those are different things.
A real customer may:
That is why the first customer should be treated as a controlled learning stage rather than simply the first sale.
The goal is not to avoid every issue. The goal is to discover issues with one manageable source of feedback, fix the underlying causes, add regression coverage, and make the product stronger before the next customer arrives.
If you are the founder in this story and you are about to bring someone in, these are the questions worth asking them. The answers tell you more than a rate card does.
| Ask them | What a good answer sounds like |
| How will you work out what this app actually does before you change it? | They describe tracing architecture and data flow, not simply reading screens. |
| How will you know your fix worked? | They verify the database, API, live page, or other final destination—not only the deployment log. |
| You find the same bug in three places. What do you do? | They identify the common cause and address every affected path. |
| How will you stop future changes from breaking existing features? | They discuss regression coverage and critical customer journeys. |
| How will you review security before customer data arrives? | They discuss authentication, authorization, input handling, secrets, and application-specific risks. |
| How will I know what you changed in six months? | They document the reason behind important decisions, not just the changed files. |
| How will you decide what is ready for the first customer? | They define measurable acceptance criteria for the agreed MVP. |
And there is one question for yourself that matters more than any of the above:
What does finished actually look like?
Not simply “it works.” Define which modules are included, which workflows must work end to end, what integrations are required, what security checks must pass, what performance is acceptable, whose data will be used for testing, and who decides whether each critical workflow is ready.
Most troubled engagements do not fail because everyone involved is unreasonable. They fail because nobody wrote this down, and two reasonable people meant different things when they said finished.
A founder who has built a working product has already completed something difficult. The goal should not be to throw that work away and rebuild automatically.
The next step is to understand it.
Trace the architecture. Identify duplicated business logic. Review the database. Fix inconsistent validation and permissions. Add tests around the workflows that matter. Check security. Stabilize deployment. Define the first-customer scope. Document the decisions future developers will otherwise have to guess.
This is where experienced enterprise application development practices become useful. Architecture, testing, security, integration, deployment, and ongoing maintenance become more important once software moves from a founder-controlled environment into real business operations.
You built the part that required knowing what should exist. The last 20% requires a different discipline: making sure it continues to work when somebody other than you depends on it.
It is less visible than building another feature.
It is also the difference between a demo and a product.
And it is worth getting right.
FAQs
A vibe-coded product is software built mainly by describing requirements to coding tools rather than writing every part of the code manually. It can help founders build working products quickly, but the code still needs proper testing, security checks, and architecture review before real customers depend on it.
Not always. A product may work well in a demo but still have issues with permissions, data validation, security, database migrations, integrations, or regression testing. A production-readiness review helps identify these gaps before launch.
Review the application architecture, database, business logic, permissions, security, integrations, deployment process, performance, and critical customer workflows. Automated and manual testing should also cover the features customers will use most.
There is no fixed timeline because it depends on the product’s size, code quality, integrations, security requirements, and existing test coverage. The best approach is to audit the application first and estimate the work based on actual technical gaps rather than the number of unfinished features.
Usually, not automatically. If the core product works, start by reviewing the existing architecture and code to identify what can be kept, refactored, tested, or secured. A full rebuild should only be considered when the existing foundation creates serious technical, security, or scalability problems.
About the Author
Latest Blog