Skip to content
AppWizards

02 / Custom LLM & Generative AI Apps

Custom LLM and generative AI app development, tested on your data.

AI assistants, chat over your documents and copilots inside your tools, built on large language models and tested against real examples before launch.

6 to 10 weeks · fixed price after a one-week discovery

Fig. 05 / Crodo, answers about what is on your screenOUR PRODUCT
INVoice and screen capture
LLMLanguage model with screen context and tools
OUTAnswer, note, or action in Gmail, Calendar, Slack

What we build

What you get: specific deliverables, not a list of capabilities.

01Chat over your documents (RAG)Ask questions across contracts, manuals, tickets or policies and get answers with the source shown. Works with your permissions, so people only see what they are allowed to.
02AI assistants with memoryCustomer-facing or internal assistants that remember the conversation and the customer, and hand over to a person cleanly when needed.
03Copilots inside your toolsSummaries, drafts and suggestions inside the CRM, helpdesk or admin panel your team already uses, instead of another tab to open.
04Document and content generationProposals, reports, product descriptions and emails written from your data in your tone of voice, and reviewed by a person before they go out.
05Testing and quality checksA set of real questions with known correct answers, checked automatically every time the app changes, so quality never drops without you knowing.

What we build with large language models

Generative AI is easy to demo and hard to ship. A demo answers three questions well. A product has to answer ten thousand, show where each answer came from, refuse what it should not answer, and cost less to run than the person it helps.

The most requested build by far is some version of chat over your own documents: contracts, manuals, policies, tickets or product data, asked in plain language and answered with the source shown. The second is a copilot inside a tool your team already uses. Both are the same machinery underneath.

How retrieval actually works, and why it fails

The technique is retrieval-augmented generation, or RAG: find the relevant parts of your documents first, then ask the model to answer using only those. It is what makes an assistant answer from your refund policy instead of inventing one.

When these projects disappoint, the model is almost never the culprit. Retrieval returned the wrong passages and the model answered from them faithfully. The quality levers sit earlier in the pipeline, and none of them is exotic:

LeverWhy it decides the outcome
ChunkingSplit a policy mid-clause and no query can ever retrieve the whole rule
Document qualityRAG surfaces the true state of your documentation, contradictions included
Hybrid searchCustomers search order numbers and product codes, which embeddings handle badly
RerankingA second small model ordering candidates, often the largest quality gain per rupee
Honest fallbackSaying "I do not know, here is a person" preserves trust that a fluent guess spends

Most of a RAG schedule is that list. The model integration is roughly a day. The fuller argument, including when to fine-tune instead, is in RAG vs fine-tuning.

Permissions are part of the retrieval problem

An internal assistant that answers from company documents will happily tell a new joiner what is in the salary review file, unless permissions are enforced at retrieval time rather than bolted on afterwards.

We filter the index by the asking user's access before the model sees anything, so the assistant cannot reveal what that person could not already open. This is the single most common thing missing from internal AI pilots we are asked to review, and it is far cheaper to build in than to retrofit.

How an LLM app is built

Fig. 05 / Crodo, answers about what is on your screenOUR PRODUCT
INVoice and screen capture
LLMLanguage model with screen context and tools
OUTAnswer, note, or action in Gmail, Calendar, Slack

We have shipped this for our own products. Crodo answers questions about whatever is on your screen and acts in your apps, which is retrieval over live context rather than a document store. Building it taught us more about search, latency and cost than any brief could.

Measuring quality, because it drifts

An LLM app does not fail loudly. It gets quietly worse after a model update or a prompt change nobody thought was risky, and nothing in your monitoring goes red.

The defence is an evaluation set: real questions with known correct answers, scored automatically on every change, with a threshold that blocks a release when quality drops. We build one before writing the app, not after, and hand it over with the code so your team can keep using it. It is also what makes "is it good enough to launch" a number rather than an argument.

Our LLM app development process

  1. Discovery, week 1. Which questions, which documents, which users, and what they are allowed to see. We collect a sample and write the first fifty test questions with you.
  2. Prototype, weeks 2 to 3. Search and answers on your real documents, scored against those questions. We compare two or three models on quality, speed and cost.
  3. Build, weeks 4 to 10. Permissions, the interface inside your tools, cost controls, monitoring and weekly demos.
  4. Run. Quality dashboards, a monthly review of wrong answers, and improvements on a retainer or a full handover.

Frequently asked questions

Common questions about Custom LLM & Generative AI Apps.

What is an LLM app?

An LLM app is software built on a large language model such as GPT, Claude or Gemini. The model reads and writes text; the app around it adds your data, your rules, a user interface and checks on quality. Chat over documents, AI assistants and writing tools are all LLM apps.

What is RAG and why does it matter?

RAG (retrieval-augmented generation) means the app first finds the relevant parts of your own documents and then asks the model to answer using only those. It is how an AI assistant answers from your policies or manuals instead of guessing, and it lets every answer show its source.

How do you deal with wrong or made-up answers?

Answers show their sources, low-confidence questions go to a person, and we measure the rate of wrong answers on a test set before and after every change. You see that number and you decide what is acceptable.

Which AI models do you use?

Whichever fits the job: OpenAI, Anthropic and Google models through their APIs, or open models such as Llama, Mistral and Qwen hosted privately when your data must stay in a particular place or costs need to be lower. We compare candidates on your data before choosing.

Will our data be used to train someone else's model?

No. We use API terms that exclude training, or host the model ourselves. Your documents stay in your cloud or ours under NDA, personal data is removed before it reaches a third-party model, and we never train on client data.

What does an LLM app cost to run each month?

We estimate usage and hosting from the prototype's real traffic and design to keep it down, with caching, smaller models for simple steps and batching. As a reference point, chat over documents at 50,000 questions a month is usually tens of dollars in model usage plus the search index; the cost that surprises people is regenerating answers that could have been cached.

Can you build a chatbot for our website?

Yes, and the useful version is grounded in your own content rather than a generic bot. It answers from your help pages, product data and policies, shows the source of each answer, says it does not know when retrieval finds nothing, and hands over to a person cleanly. That last behaviour is what separates a useful assistant from one people learn to ignore.

Why do RAG projects fail?

Almost always on retrieval, not on the model. The passages returned were wrong, so the model faithfully answered from the wrong thing. The levers are chunking along a document's real structure, cleaning contradictory or outdated content, combining keyword search with semantic search so order numbers and product codes still work, and adding a reranking step. We wrote this up in detail in our RAG versus fine-tuning article.

Next step

Tell us what you want to build. We will tell you what it costs and how long it takes.

A free 30-minute call with an engineer, not a salesperson. You leave with a clear plan, a price range and an honest opinion on whether AI is the right tool for the job.