What we build with large language models
Generative AI is easy to demo and hard to ship. A demo answers three questions well. A product has to answer ten thousand, show where each answer came from, refuse what it should not answer, and cost less to run than the person it helps.
The most requested build by far is some version of chat over your own documents: contracts, manuals, policies, tickets or product data, asked in plain language and answered with the source shown. The second is a copilot inside a tool your team already uses. Both are the same machinery underneath.
How retrieval actually works, and why it fails
The technique is retrieval-augmented generation, or RAG: find the relevant parts of your documents first, then ask the model to answer using only those. It is what makes an assistant answer from your refund policy instead of inventing one.
When these projects disappoint, the model is almost never the culprit. Retrieval returned the wrong passages and the model answered from them faithfully. The quality levers sit earlier in the pipeline, and none of them is exotic:
| Lever | Why it decides the outcome |
|---|---|
| Chunking | Split a policy mid-clause and no query can ever retrieve the whole rule |
| Document quality | RAG surfaces the true state of your documentation, contradictions included |
| Hybrid search | Customers search order numbers and product codes, which embeddings handle badly |
| Reranking | A second small model ordering candidates, often the largest quality gain per rupee |
| Honest fallback | Saying "I do not know, here is a person" preserves trust that a fluent guess spends |
Most of a RAG schedule is that list. The model integration is roughly a day. The fuller argument, including when to fine-tune instead, is in RAG vs fine-tuning.
Permissions are part of the retrieval problem
An internal assistant that answers from company documents will happily tell a new joiner what is in the salary review file, unless permissions are enforced at retrieval time rather than bolted on afterwards.
We filter the index by the asking user's access before the model sees anything, so the assistant cannot reveal what that person could not already open. This is the single most common thing missing from internal AI pilots we are asked to review, and it is far cheaper to build in than to retrofit.
How an LLM app is built
We have shipped this for our own products. Crodo answers questions about whatever is on your screen and acts in your apps, which is retrieval over live context rather than a document store. Building it taught us more about search, latency and cost than any brief could.
Measuring quality, because it drifts
An LLM app does not fail loudly. It gets quietly worse after a model update or a prompt change nobody thought was risky, and nothing in your monitoring goes red.
The defence is an evaluation set: real questions with known correct answers, scored automatically on every change, with a threshold that blocks a release when quality drops. We build one before writing the app, not after, and hand it over with the code so your team can keep using it. It is also what makes "is it good enough to launch" a number rather than an argument.
Our LLM app development process
- Discovery, week 1. Which questions, which documents, which users, and what they are allowed to see. We collect a sample and write the first fifty test questions with you.
- Prototype, weeks 2 to 3. Search and answers on your real documents, scored against those questions. We compare two or three models on quality, speed and cost.
- Build, weeks 4 to 10. Permissions, the interface inside your tools, cost controls, monitoring and weekly demos.
- Run. Quality dashboards, a monthly review of wrong answers, and improvements on a retainer or a full handover.