Learn CodeNotes from sessions with Claude Code, my teacher
Lesson 05

Building a RAG from scratch

2026-10-07 · about 6 minutes
AIRAGembeddingsPHPOpenAI

What we built#

A private web page where you type a question and an AI answers using only the content you gave it: first an FAQ we wrote together, later any PDF, Word file or web page. Under every answer it shows which passages it used and how closely each one matched.

This technique has a name: RAG, short for Retrieval-Augmented Generation.

  • Retrieval: find the pieces of your content that relate to the question.
  • Augmented: add those pieces to the instructions we send the AI.
  • Generation: the AI writes the answer from them.

Why not just ask ChatGPT?

A general AI model knows a lot about the world but nothing about your business, and when it doesn't know, it may guess confidently. RAG fixes both: we hand it the right facts at the moment of the question and tell it to answer only from them. If the facts aren't there, it says so.

The big idea: meaning as numbers#

Computers can't compare meaning directly, so we convert text into a list of numbers called an embedding. We used OpenAI's text-embedding-3-small, which turns any text into 1,536 numbers.

The magic property: texts with similar meaning get similar lists of numbers, even when they use different words.

Text Similar meaning to "Do I need to code?"
"No coding experience is required to join." ✅ very close
"The beginner modules start from the basics." ✅ close
"ContentGuard checks similarity to the source." ❌ far

Measuring "closeness": cosine similarity#

Think of each embedding as an arrow pointing in a direction in a 1,536-dimension space. Two texts that mean similar things point in nearly the same direction. We measure the angle between arrows with cosine similarity:

  • 1.0 = same direction (same meaning)
  • 0 = unrelated
  • In practice, good matches for us scored about 0.5–0.8.

The calculation is just multiply the two lists number by number and add it all up (the dot product). OpenAI's embeddings are already "unit length", so the dot product is the cosine similarity.

$score = 0.0;
foreach ($passageVector as $i => $v) {
    $score += $v * $questionVector[$i];   // 1,536 multiplications
}

The two stages#

BUILD (once, when content changes)        ASK (every question)
──────────────────────────────────         ─────────────────────────────
FAQ / PDF / web page                       visitor's question
   │ split into passages                      │ embed the question
   ▼                                          ▼
passages ──► embed each ──► save        compare with every saved passage
                         (index file)        │ keep the top 5 (score ≥ 0.20)
                                             ▼
                                   AI model answers from those 5 only
                                             │
                                             ▼
                                answer + "used these passages"

Stage 1: building the index#

  1. Write the source. We drafted 24 questions and answers about the business, using only facts we could verify. Unknowns (prices, contact details) were deliberately left out and listed for later: an AI can't invent what isn't there.
  2. Split into passages ("chunking"). One FAQ entry fits in one passage. For long documents we later split into ~800-character pieces (lesson 07).
  3. Embed every passage with one API call (you can send many texts at once).
  4. Save the passages and their numbers in a file: the index.

Stage 2: answering a question#

  1. Embed the question (one small API call).
  2. Score every passage against it with cosine similarity.
  3. Keep the top 5 that score at least 0.20 (below that it's noise).
  4. Send the AI model a message like this:
You are the assistant for "SgAiSolution FAQ".
Answer using ONLY the passages below. Keep answers short.
If the passages don't contain the answer, say the knowledge base
doesn't cover that yet. Never invent prices, dates or contact details.

Passages:
[FAQ] Q: Do I need coding experience? A: No...
...
  1. Return the answer plus the passages and their scores, so the reader can see where it came from.

The instruction does half the work

"Answer ONLY from the passages, and say so when they don't cover it" is what stops the model guessing. We tested it with "What is the capital of France?": no passages matched, so it politely declined, which is exactly right for a business assistant.

Where each piece runs (and why)#

Part Runs on Why
Chat page (demo.html) Visitor's browser Just the interface
Answer endpoint (demo-api.php) The web server Holds the API key, does retrieval and calls the model
API key + index A private server folder Never in a web page: anyone can read a web page's code

Why PHP?

The website runs on shared hosting, and shared hosting runs PHP out of the box. Choosing the language the server already speaks meant nothing new to install.

Keeping it private#

  • Password: the server asks for a username and password before showing the page (HTTP Basic Auth, configured in a file called .htaccess, see lesson 08).
  • Not in Google: a noindex tag and header tell search engines not to list it.
  • Secrets locked away: the key and index live in a folder the web server refuses to serve (we tested: it returns 403 Forbidden).

What went wrong (and how we found it)#

\"Call to undefined function mb_strlen()\"

The phone's PHP was missing the text extension (mbstring) that the server has. Instead of depending on it, we wrote a tiny fallback: use the multibyte version if it exists, otherwise the plain one. Lesson: code that runs in two places should not assume both places are identical.

The code changed but the answers didn't

After editing, test answers still used the old wording. The files on disk were correct, so something was serving a stale copy: PHP's opcache caches compiled code and checks each file's "last modified" time. Android's storage doesn't update that time, so PHP never noticed the edits. Fix: turn the cache off for local testing. Lesson: when the code and the behaviour disagree, look for a cache.

Knowledge base saved as an empty file

Saving with a file lock (LOCK_EX) silently failed on the phone's storage. We switched to a safer pattern used in professional systems: write to a temporary file, then rename it into place. A rename is all-or-nothing, so a crash can never leave a half-written file.

Key takeaways#

  • RAG = find the relevant facts, then ask the model to answer from them.
  • Embeddings turn meaning into numbers; cosine similarity measures how close two meanings are.
  • The prompt must forbid guessing and allow "I don't know".
  • Keys and data stay on the server, never in the browser.
  • Show your sources: it builds trust and makes debugging easy.

Quick quiz#

1. What does each letter in RAG stand for?

Retrieval (find relevant passages), Augmented (add them to the prompt), Generation (the model writes the answer).

2. Two passages score 0.72 and 0.18 against a question. Which is used?

Only the 0.72 one. We ignore anything under 0.20 as unrelated noise.

3. Why can't the API key go inside demo.html?

Everything in a web page is downloadable by the visitor. Anyone could copy the key and spend your money. It must stay on the server.

4. Why did we leave prices out of the FAQ?

We didn't know them. If a fact isn't in the source, the assistant correctly says it doesn't know instead of inventing one.

Try it yourself#

Rewrite one FAQ answer

Change one answer in your FAQ (e.g. add a real price), rebuild the index, and ask the question again. Watch the answer change, and notice the match score barely moves, because the question still means the same thing.