— Practice / RAG

RAG development.

Retrieval-augmented generation lets a model answer from your own documents, with citations, instead of guessing. We build production RAG pipelines on Cloudflare Workers. Chunking, embeddings, the vector store, and generation are each tuned to your corpus, which is where tutorials stop.

— Why us

Grounded, with citations.

The AI search on this site is a working RAG pipeline. It answers from our own pages and links the sources. Our open-source RAG template deploys in five minutes, and client builds follow the same pattern: no LangChain, no framework, and each primitive maps to one service you can reason about.

Deploy your own in five minutes: git clone https://github.com/setkernel/cf-rag-template

— What we build

The whole pipeline.

Chunking and ingestion

Unglamorous, and it decides answer quality. We split your content so retrieval returns the right context, and keep it in sync as the source changes.

Embeddings

Turning your text into vectors with a model matched to the corpus and the budget, on Cloudflare Workers AI or OpenAI.

Vector search

Cloudflare Vectorize or another vector store, with hybrid retrieval where it helps, so the model gets the passages that actually answer the question.

Grounded generation with citations

The model answers strictly from retrieved context and cites the source, with a system prompt that refuses to invent. When it is wrong, the cause is usually a missing document, which you can fix.

Evals

Retrieval and answer quality measured in the pipeline, so a change that quietly degrades results fails the build instead of shipping.

The right surface

A search box, an API, an agent tool, or an MCP server. The pipeline goes wherever it fits, and it stays cheap to run because Workers bills CPU time and not time spent waiting on the model.

— How it works

The engagement.

Same five-step method as every SetKernel build (Brief, Architect, Sprint, Ship, Operate), each with a written artefact you review. We start from a short written brief: the questions the system should answer, the documents it should answer from, and what a good answer looks like. You get a scoped price and a fit / no-fit answer within one business day.

— Where we work

Atlantic Canada, and worldwide.

We are a complete technology partner in Halifax, Nova Scotia, Canada, and we work remotely with teams well beyond the region. A RAG pipeline lives in the cloud, so where your team sits does not change the build.

— Questions

Before you write.

What is RAG, in one sentence?

Retrieval-augmented generation retrieves the most relevant passages from your own content and hands them to a language model as context, so the answer is grounded in your documents and can cite them, instead of relying on whatever the model happened to memorise.

RAG or fine-tuning: which do I need?

Usually RAG. Fine-tuning changes how a model writes. RAG changes what it knows, updates the moment your documents do, and gives you citations to check. Most "the model should know our stuff" problems are retrieval problems. If yours is the exception, we will say so.

How do you stop it from making things up?

The system prompt instructs the model to answer only from retrieved context and to say so when the answer is not there. Retrieval is tuned so the right passages surface, and evals catch regressions. A wrong answer then traces back to a missing or mis-chunked document, which is something you can fix.

Can we see your work first?

Yes. The AI search on this site is a live RAG pipeline with citations. Read the open-source cf-rag-template and its companion essay on building RAG on Cloudflare without LangChain. Working code says more than a pitch.

How do we start?

Send a short written brief: the questions to answer, the documents to answer from, the deadline. We reply in writing within one business day with fit / no-fit and, if fit, a scope and price. No discovery call before the brief.

— Engage

Have a pile of documents a model should be able to answer from?

Tell us in two paragraphs: the questions, the documents, what a good answer looks like. We reply in writing within one business day.

Esc