ShippingAIMVPsinSixWeeks:APracticalEngagementStructure
Devika Sundaram· Director of Product· September 10, 2025· 7 min read
Why six weeks is the right constraint
Most AI MVP engagements fail not because the model underperforms, but because the scope never stops growing. Six weeks is short enough that a team cannot quietly redefine the goal twice, and long enough to build something a real user can evaluate honestly. We treat the deadline as a forcing function: it pushes every decision toward the smallest version of the product that still proves the underlying bet is worth making.
The constraint also changes how stakeholders behave. When a project is open-ended, everyone adds one more requirement because there is always more time later. When the clock is visible and shared, requests get triaged against a single question: does this change whether we learn something in week six? Most requests fail that test, and that is the point.
Week 1: Scoping around one measurable outcome
We spend the first week refusing to write code. Instead we pin down one outcome metric the MVP has to move, such as time-to-resolution on a support ticket or accuracy on a document classification task, and we write down the current baseline before any AI is involved. Without a baseline, an eventual demo that looks impressive tells you nothing about whether it is actually better than what existed before.
This is also when we decide what the MVP will explicitly not do. We write a short list of adjacent features that are out of scope, put it in the kickoff document, and refer back to it every time a new idea surfaces during the build. It is the cheapest scope-creep insurance we have found.
Eval-first development, not vibes-first development
Before we write the first prompt or fine-tune a single model, we build a small evaluation set: 30 to 80 real or realistic examples with known-good answers, pulled from the client's actual data wherever possible. Every subsequent change to the prompt, retrieval pipeline, or model choice gets scored against this set. This sounds like overhead in a six-week project, but it is what keeps the team from mistaking a good-looking demo for a working system.
The eval set also becomes the shared language between engineering and the client. Instead of debating whether an output feels right, we can point to a score: 71% pass rate on the golden set this week versus 58% last week. That number is what actually earns the next round of scope, not a screenshot.
Weeks 2 through 5: build against the harness, not the demo
With the eval harness in place, the middle of the engagement is iterative: change one variable, re-run the eval, keep what improves the score, revert what does not. This looks slower than freestyle prompt engineering, but it prevents the common failure mode where a team spends four weeks polishing a UI around a model that silently regressed after the third prompt tweak.
We also deliberately build the ugliest possible interface first, usually a plain internal tool with a table and a text box, and only invest in UI polish once the underlying pipeline is stable against the eval set. Polishing an interface around an unstable core is one of the most common ways MVP budgets get wasted.
Guarding against scope creep without saying no to everything
Saying no to every new idea kills stakeholder trust as fast as saying yes to all of them. Our rule is that any new request gets logged in a backlog visible to everyone, dated, and explicitly deferred rather than silently dropped. This keeps people from repeating requests to get attention, because they can see it is already tracked for the next phase.
What done looks like at week six
A finished MVP under this model is not a polished product; it is a working system with a documented eval score, a clear list of what it does and does not do, and a specific recommendation for the next phase based on evidence rather than optimism. Roughly a third of the MVPs we build this way lead directly into a second phase; the rest tell the client something equally valuable, that the original bet needed to change before more money went into it.