Across four products of our own, the AI model was never the thing that went wrong. What went wrong was the cost of running it, the permissions screen before anyone saw it, the absence of a signal telling us whether it mattered, and the half of a marketplace nobody wants to build. Every one of those is a decision made before a line of model code is written.
Agencies show you wins. This is the other list, and it is more useful, because each of these mistakes now shapes a question we ask clients on the first call.
What we expected to be hard, against what actually was
| Product | What we braced for | What actually cost us |
|---|---|---|
| Decornoa | Would the generated rooms look good enough? | Cost per render, queueing under load, and judging quality at volume |
| Crodo | Voice accuracy and screen understanding | The macOS permission wall standing between install and first use |
| TurboType | Shipping something focused, fast | Having no price, and therefore no signal about what it was worth |
| BookMyPodcastStudio | Building booking, search and listings | Supply. A marketplace with no studios is a search box |
Read down the right-hand column and notice what is missing: model quality. Four products, four different domains, and not one of them was decided by whether the AI was clever enough.
Decornoa: the model was the cheap part
Decornoa turns a photograph of a room into AI-generated redesigns. Going in, the anxiety was aesthetic: would the output look like a designer made it?
That worry resolved quickly. The problems that persisted were operational. Image generation costs real money per render and takes real seconds, which means three things at once: a bill that scales with enthusiasm rather than revenue, a queue that forms exactly when a burst of visitors arrives, and a quality problem that is not "is this good" but "how do we know it is still good across thousands of renders without looking at all of them".
Those are engineering and economics problems wearing an AI costume. They are also why our AI quotes now spend more schedule on the plumbing around the model than on the model, and why we insist on prototyping the unit economics before the interface. The full arithmetic of that habit is in what an AI agent costs to run.
Crodo: the first five minutes are the product
Crodo is a macOS voice assistant. It listens, it can see what is on your screen, and it takes actions in tools like Gmail, Calendar and Slack. Every one of those capabilities requires the operating system to ask the user for permission first: microphone, screen recording, accessibility.
So the honest shape of the product is that a person downloads it and is immediately presented with a series of system dialogs asking for the most alarming-sounding permissions macOS offers, before the app has demonstrated a single useful thing. We built the assistant first and treated that sequence as setup. It is not setup. It is the product's first impression, and for a meaningful share of people it is also the last one.
The lesson generalised: onboarding is a feature with a design, a budget and a test plan, not the week before launch. We now schedule it early, and on client projects the first thing we ask about a new AI feature is what a user must agree to before it can help them.

TurboType: shipping fast is the easy half
TurboType is typing practice on real code, free, in eight languages. It is the product we point to when a client asks how quickly a genuinely focused thing can ship, and that reputation is deserved: one platform, one flow, no payments, no accounts to speak of.
What we did not think hard enough about is that removing the price also removed the signal. A free tool collects usage and no evidence of value. People arrive, type, and leave, and none of it tells you whether anyone would have paid, which of the eight languages mattered, or what to build next. Speed to launch was real; speed to learning was not, because we had not decided what we wanted to learn.
This is the mistake most likely to be repeated by a founder reading this. Shipping in weeks is a genuine advantage and we still recommend it, but an MVP without a defined question is just a small product. Define the number that has to move before you start, which is the same discipline the AI readiness self-test applies to a business case.
BookMyPodcastStudio: a marketplace is two products
BookMyPodcastStudio lets people find and book podcast studios: listings, search, comparison, booking. Every one of those words describes the demand side, and the demand side is the half that is fun to build and easy to demo.
The half that decides whether a marketplace exists is supply. Studios have to be found, contacted, persuaded, onboarded, verified and kept accurate, and almost none of that is software. We estimated the product and under-estimated the business, which is a distinct and more expensive kind of error.
It is why our marketplace quotes now run longer than founders expect, and why the first thing we ask is which side you already have. A client who arrives with fifty suppliers already signed is building a different, cheaper product than a client who arrives with a design.
The pattern, stated plainly
Four products, four industries, and one shape underneath all of them: the risk was never in the part we were best at. We are good at building software, so the software was fine. The failures clustered in the decisions surrounding it, the ones easy to postpone because they feel like business rather than engineering.
That is the entire argument for hiring a team that runs its own products. Not that we write better code because of it, but that we have paid for these specific mistakes with our own money, and now ask about them on the first call instead of the fourth month. All four are listed, with links, on our products page.
What it changed about how we quote
Three concrete habits came out of this list, and you will meet all of them if you work with us:
- We prototype the economics, not only the output. Before an AI feature is approved, we produce a cost per case from real usage, because the answer being good is necessary and nowhere near sufficient.
- Onboarding gets scheduled time. Permissions, empty states and the first five minutes are designed and tested, not squeezed into the last sprint.
- We ask what you want to learn. If a build cannot name the number it exists to move, it is not an MVP, and we will say so before quoting it. That conversation is what the AI readiness audit formalises.
Questions people ask
Do you still run all of these?
The four on this page are live and you can open them now, which is the point of listing them rather than describing them. We have built others that we no longer run; products end, and a page claiming otherwise would fail the first check anyone sensible performs.
Is the AI model really never the problem?
Rarely, and less every year. Models improved faster than anything around them, so the differentiator moved to retrieval quality, cost control, evaluation and interface. Where the model genuinely is the constraint, it usually shows up as a data problem instead, which is the first pillar of our readiness assessment.
What would you do differently if you started one today?
Decide the learning goal and the unit economics in week one, and design the permission or sign-up sequence before the feature it protects. Both are cheap at the start and expensive to retrofit, which is the entire content of this article.
Does running your own products make you more expensive?
No, and it makes our estimates hold better, which is worth more. Knowing where a project actually goes wrong means fewer change requests, which is the cost that surprises people. The ranges themselves are in the app development cost guide.
The honest limit
Four products is a small sample, all of them ours, none of them at the scale where a different set of problems begins. Nothing here proves a general law about AI products; it describes what four attempts taught one team, which is exactly as much as we would want a reader to take from it.
What we would say with confidence is narrower: if you are planning an AI product and your project plan spends most of its worry on model quality, the worry is in the wrong place. Move it to cost per user, to the first five minutes, and to whichever side of your market you do not yet have.
