AI strategy · 7 min read
Why AI pilots stall, and how a 2-week sprint gets them moving
Most AI pilots don't fail on the model. They stall on scope, data and the missing definition of 'good'. Here is a practical way to get from demo to decision in two weeks.

On this page
Every company has one: an AI pilot that looked great in a demo, then sat for months while people debated whether to fund it. In our experience the model is rarely the problem. The pilot stalls because nobody can answer three simple questions.
The three questions every stalled pilot can't answer
A theme like "AI for customer service" can't be tested. A workflow can.
Without a test set and a target, every review becomes a debate about examples.
Finance can't approve a number nobody has estimated.
If you can't answer all three for your pilot today, it's probably stalled for one of the reasons below.
Five reasons pilots stall
1. The scope is a theme, not a workflow
"Use AI for customer service" is a direction. "Draft replies to refund emails and pull the order number from the attachment" is a workflow. Only the second can be built, measured and handed to a team.
2. The demo ran on clean sample data
Real inputs are scanned PDFs, spreadsheets with creative column names and messages in two languages. A pilot that never touched them has proved very little.
3. There's no definition of "good"
Without a small labeled test set and an agreed target, reviews turn into anecdotes. One person remembers the impressive answer, another remembers the wrong one, and nothing gets decided.
4. Nobody priced production
Token costs, latency, hosting, monitoring and the human review queue are unknown. Finance can't approve a budget line nobody can explain.
5. It doesn't fit into the real process
The output lands in a separate tool that nobody opens. Adoption dies quietly, and the pilot gets blamed.
Demo vs. production: what actually changes
| Typical demo | Production-ready pilot | |
|---|---|---|
| Data | Hand-picked, clean examples | Real inputs, including the ugly ones |
| Quality | "It looked right" | Accuracy on a labeled test set, with known failure cases |
| Cost | Unknown | Cost per item, at expected volume |
| Speed | One request at a time | Latency and throughput measured |
| Risk | Not discussed | Human review where errors are costly |
| Fit | Separate tool | Output lands where the team already works |
The 2-week AI Sprint
We run one sprint on one workflow with your real data. The goal is not a prettier demo. It's a decision.
Swipe to see the full diagram →
Days 1–2: Frame the workflow
Pick the workflow with the clearest value. Write down the input, the output, who uses it and what "correct" means. Agree on the target the business needs, not the one the model can hit.
Days 3–4: Build the test set
Collect 50 to 200 real examples, including the difficult ones. Label the expected answers with the people who do the job today. This test set becomes the referee for every decision that follows.
Days 5–8: Prototype and compare
Build the simplest pipeline that could work: extraction, retrieval or an agent, plus rule-based checks. Try two or three models and compare them on the same test set, including smaller, cheaper ones.
Days 9–10: Measure and price
Run the test set. Report accuracy, where it fails, speed and cost per item. Estimate what running it in production would cost, including the human review step.
What you get at the end
Running on your own data, not a sample set.
Accuracy on your test set, failure cases, speed and cost per item.
How it fits your systems, with the risks and the review points marked.
Timeline, team and running costs, ready for a go/no-go decision.
Reading the results: an example
The evaluation report is what changes the conversation. Here is the kind of summary it contains, for an invoice-extraction workflow.
| Measure | Result | What it means |
|---|---|---|
| Fields extracted correctly | 94% of 150 test invoices | Good enough to automate, with review on flagged items |
| Main failure case | Handwritten totals on scanned copies | Route scans with handwriting to a person |
| Average processing time | 6 seconds per invoice | Fast enough for same-day processing |
| Model cost | A few cents per invoice | A small fraction of the manual cost |
The discussion moves from "Is AI ready?" to "Is 94% good enough, and what happens with the other 6%?" That's a question a team can answer.
When a sprint is not the right move
A sprint works best when the workflow is repetitive, the data is accessible and someone owns the outcome. It's the wrong tool when:
- There's no access to real data yet. Solve that first.
- Nobody owns the result. A sprint without a business owner produces a report nobody acts on.
- The task happens a few times a month. The effort to automate it may never pay back.
Where we've applied this
We've used this pattern across very different domains: logistics order intake, digital publishing pipelines, performance and feedback systems and large-scale media monitoring. The domains differ, but the reasons pilots stall are the same.
Frequently asked questions
How much data do we need for a sprint?
Usually 50 to 200 real examples of the workflow are enough to build a meaningful test set. Quality and variety matter more than volume.
Do we need to choose a model provider first?
No. Comparing two or three models on your own test set is part of the sprint, so the choice is based on your results, not on benchmarks.
What happens after the sprint?
You get a costed build plan. If the answer is "go", the next step is usually a 6 to 12 week build that takes the workflow to production with monitoring in place.
Keep reading

How smaller companies win with AI: lessons from Arabyati's award-winning EdTech platform
Smaller companies don't need bigger budgets to compete with AI. They need focus. Here is how Arabyati used AI to improve its product and its operations, and what other growing companies can take from it.
September 28, 2026
Turning anonymous employee feedback into decisions with an LLM
People finally speak up, and leadership gets hundreds of comments nobody has time to read. Here is how we turned anonymous feedback into themes and actions with an LLM, and why the rewards engine next to it uses rules instead.
September 28, 2026