One builder, many knowledge bases
From one FAQ to anything#
The first assistant only knew our FAQ. The next step was a builder: a page where you name a knowledge base, add files or links, and get a new assistant. Any content, not just ours.
Step 1: getting the text out#
Every source has to become plain text first. Each type needs a different tool:
| Source | How the text is extracted | Where |
|---|---|---|
| pdf.js (Mozilla's PDF reader) reads the text layer page by page | Your browser | |
| Word (.docx) | mammoth.js reads the document's XML inside the .docx zip | Your browser |
| TXT / Markdown / CSV | Read as-is | Your browser |
| Web page (URL) | The server downloads it and strips the HTML | The server |
Why read PDFs in the browser?
PHP on shared hosting has no good built-in PDF reader, but browsers can run pdf.js. Doing it on the visitor's device also keeps big files off the server: only the text is sent.
Cleaning a web page#
A web page is mostly menus, scripts and footers. We:
- Remove
<script>,<style>,<nav>,<footer>,<form>… - Prefer the
<main>or<article>part if the page has one. - Turn block tags (headings, paragraphs, list items) into line breaks, so questions and answers stay on separate lines.
- Add a space after links and buttons, otherwise side-by-side tabs glue together ("Accessing a loungeBringing guests").
Safety when the server fetches a URL: SSRF
If a server fetches any address it's given, an attacker could ask it to fetch internal addresses (like localhost, the machine itself, or the special "link-local" address clouds use for instance metadata) and leak secrets. This is called Server-Side Request Forgery (SSRF). Our fetcher only accepts http/https, looks up the host's IP, and refuses private, loopback and reserved ranges, including after redirects. We tested all four blocked cases.
Step 2: chunking with overlap#
Long text is split into passages of about 800 characters, broken at paragraph and sentence boundaries. Each passage repeats the last ~150 characters of the previous one.
Passage 1: [..................answer begins | ...]
Passage 2: [...answer begins | answer ends.....]
└── overlap ──┘
Why overlap?
Without it, an answer that straddles a boundary gets cut in half, and neither passage alone contains it. Overlap means every sentence appears whole in at least one passage.
Passages starting mid-word
Our first version cut the overlap at exactly 150 characters, so a passage could start with "…ay be eligible". Harmless for search, ugly to read. Fix: move the overlap start forward to the next space so it begins on a whole word.
Why 800 characters? Small enough that a match is precise (the passage is mostly about one thing), big enough to hold a full question and answer. It's a tunable setting, not a law.
Step 3: the knowledge base file#
Each knowledge base is one JSON file on the server:
{
"title": "Priority Pass Lounge FAQ",
"model": "text-embedding-3-small",
"samples": ["Are guests allowed?", "..."],
"sources": [{ "name": "Airport Lounge Access FAQ", "type": "Web", "text": "full text..." }],
"chunks": [{ "source": "...", "text": "passage...", "e": "base64 of 1,536 float32 numbers" }]
}
Storing vectors compactly
1,536 numbers written as JSON text take about 11 KB. Packed as raw 4-byte floats and base64-encoded, they take about 6 KB: almost half the size, and faster to load.
No link back to the website#
A key decision: the knowledge base stores the extracted text, not the URL. The assistant never goes back to the original site, so if the page changes or disappears, your copy still works. (Using someone else's content publicly still needs care: summarise or link, don't republish.)
Sample questions per knowledge base#
Each base gets five sample questions written by the AI from its own content, refreshed whenever a source is added. A Priority Pass base suggests "Are guests allowed?", not questions about a different business.
Step 4: taking it with you#
Three downloads per knowledge base:
| Download | What it is | Use it for |
|---|---|---|
.md |
All the extracted text, readable | Keep, edit, reuse |
.json |
Everything incl. embeddings | Re-import later without paying to re-embed |
.html |
A self-contained offline assistant | Open on any PC by double-clicking |
The offline assistant: two kinds of search#
- Keyword search (no internet, no key) using BM25, the classic search-engine formula. It scores passages by how often your words appear, giving extra weight to rare words and adjusting for passage length.
- AI answers (needs internet + your key): the same RAG as the website, but run from the file: embed the question, compare against the embedded vectors, ask the model.
BM25 vs embeddings
BM25 matches words: "guests" finds "guests". Embeddings match meaning: "can my kids come in?" finds the passage about guests even with no shared words. Having both means the file is useful even offline.
CORS: can a file talk to an API?
Browsers block a page from calling another website's API unless that API allows it (CORS, Cross-Origin Resource Sharing). A file opened from disk has origin null. Before building the AI mode we checked that OpenAI answers Access-Control-Allow-Origin: null, which it does. Your key stays in that browser's local storage and is never written into the file.
Key takeaways#
- Every source becomes plain text first; pick the right extractor per format.
- A server that fetches URLs must block private addresses (SSRF).
- Chunk with overlap so no answer is split in half.
- Store the knowledge, not the link: your copy survives the source changing.
- Keyword search (BM25) works offline; embeddings understand meaning. Both are useful.
Quick quiz#
1. Why can a scanned PDF produce no text?
It's just pictures of pages, with no text layer. It would need OCR (optical character recognition) first.
2. What does the 150-character overlap protect against?
An answer being split across two passages so neither contains it whole.
3. Which search mode finds "can my kids come in?" → passage about guests?
The embedding (AI) mode, because it matches meaning, not words.
Try it yourself#
Build one from a page you care about
Pick an FAQ page you often visit, build a knowledge base from it, download the .html file, turn off your Wi-Fi, and use offline search. Then turn Wi-Fi back on and ask the same thing in your own words with AI mode.