As part of my internship at Valentius Kryptix, I set out to build something that ties together four of the core AI skills we’ve been developing: working with an LLM API, engineering effective prompts, building a genuine retrieval pipeline, and wrapping all of it into a real, deployed web application. The result is a Retrieval-Augmented Generation (RAG) assistant — a chatbot that doesn’t just answer questions from its own general training, but actually retrieves and reasons over a specific set of documents before responding.
You can try it live here: https://chatbot-with-rag-zeta.vercel.app
What I Built
At its core, the app lets a user chat with an AI assistant that is “grounded” in a document set. Instead of the classic problem where a chatbot confidently makes something up, this assistant only answers using content it has actually retrieved from real source documents — and clearly says “I don’t know” when the answer isn’t present in that context.
To make this genuinely testable (and not just a wrapper around a general-purpose model), I added a feature that shows exactly which documents the assistant is currently grounded in, along with a live upload option: anyone, including a reviewer, can upload their own .txt file on the spot and immediately start asking it questions specific to that new content. This was an important addition after early feedback — it’s one thing to claim retrieval works, and another to let someone verify it in real time with their own data.
How It Works Under the Hood
The pipeline follows the standard shape of a RAG system, built from scratch rather than through a heavier framework, so I could understand each moving part:
- Chunking — source documents are split into overlapping text chunks, since embedding an entire document at once loses the specificity needed for accurate retrieval.
- Embedding — each chunk is converted into a vector using Google’s Gemini embedding model, capturing its semantic meaning as a set of numbers.
- Retrieval — when a user asks a question, the question itself is embedded, and cosine similarity is used to find the chunks whose meaning is closest to the question.
- Generation — those top-matching chunks are inserted into a carefully engineered prompt, and Gemini’s chat model generates an answer using only that retrieved context.
The whole thing runs as a Next.js app deployed on Vercel, with the LLM calls handled through free-tier API routes — no dedicated backend server or paid infrastructure required.

Key Learnings
1. Retrieval quality matters more than model choice. Early on, I noticed that even a capable model gives weak answers if the wrong chunks get retrieved. Tuning chunk size and overlap, and returning enough (but not too many) chunks, made a bigger difference to answer quality than tweaking the prompt itself.
2. Prompt engineering is about constraints, not cleverness. The most valuable part of my system prompt wasn’t a clever instruction — it was a strict constraint: answer only from the provided context, and explicitly say when the answer isn’t there. This one rule is what separates a real RAG app from a chatbot that just sounds confident regardless of whether it’s actually right.
3. “Deployed” isn’t the same as “production-ready.” Getting an app live on Vercel took minutes. Getting it to fail gracefully — handling a missing API key, a failed embedding call, or an empty vector store without crashing the UI — took real thought. Loading states and visible error handling turned out to be just as important as the AI logic itself.
4. Transparency builds trust in AI tools. Adding the “documents currently loaded” panel seemed like a small UI addition, but it fundamentally changes how much a user can trust the system. An AI assistant that shows its sources — and lets you test it with your own data — is far more credible than one that just answers.



Leave a Reply
You must be logged in to post a comment.