All writing
Last updated Sep 1, 20269 min read

Your AI assistant is only as good as what it can retrieve

Semantic searchEmbeddingsAI

In short

Semantic search finds notes by meaning rather than exact words. Each note is converted into an embedding — a vector of numbers positioned so that related ideas sit near each other — and a query is converted the same way, then compared by cosine similarity. This is why searching "how does login work" can surface a note titled "JWT refresh flow" that shares no keywords. For a personal knowledge base the practical rules are: chunk notes into sections rather than embedding whole documents, combine semantic search with keyword search since exact terms still matter, and always show citations so a wrong answer is verifiable rather than merely convincing.

Almost every AI note-taking tool demos the same way. Someone types a question, and a well-formatted answer appears. It looks like magic and it usually is not — because the hard part is not writing the answer. Language models have been good at that for years.

The hard part is finding the right four paragraphs out of your two thousand notes and putting them in front of the model. Get that wrong and you have built a very confident liar with access to your journal.

This is a retrieval problem, and it is worth understanding, because whether a tool is useful or annoying is decided almost entirely here.

Keyword search and its failure mode

Traditional search matches words. You type auth, it finds documents containing auth. Postgres full-text search adds stemming and ranking, so running matches run, and this is genuinely good technology — fast, predictable, cheap, and when you remember the exact term you want, unbeatable.

The failure is specific: it cannot match on meaning. You search how do users log in and you have a note titled JWT refresh flow. Zero shared keywords. Zero results. The note is sitting right there, it is the exact note you want, and the search engine has no way to know.

For a personal knowledge base this failure is common rather than rare, because you are searching for things you half-remember. If you remembered the exact words you would not need to search — you would just open the note.

What embeddings do

An embedding model reads text and returns a list of numbers — a vector, typically several hundred to a couple thousand of them. The useful property is how those vectors are arranged: text with similar meaning produces vectors that are close together in that space, even when the words differ.

So how do users log in and JWT refresh flow land near each other, because the model has learned from a great deal of text that these concepts are related. Search becomes geometry. Embed the query, find the nearest note vectors, return them ranked by distance. The usual distance measure is cosine similarity, which measures the angle between two vectors and ignores their length — you get a score from -1 to 1, and in practice anything above about 0.75 is worth looking at.

That is the whole idea. Everything else is engineering.

Chunking is where it goes wrong

Here is the mistake that quietly ruins most implementations: embedding entire documents.

A vector is a fixed size regardless of input length. Embed four words and you get a vector that means those four words. Embed a 3,000-word note covering database schema, auth, deployment, and a rant about your CI provider, and you get a vector that means approximately the average of all of that — which is to say, nothing in particular. It sits in a vague middle region of the space, moderately close to everything and genuinely close to nothing. It will rank mediocre-ly for every query and win none of them.

The fix is chunking: split notes into sections and embed each one. Now the auth section has its own vector that is sharply about auth.

The tradeoffs are real and there is no universally correct answer. Chunk too small and you lose context — a paragraph that says "this approach did not work" is useless without knowing which approach. Chunk too large and you are back to averaging. A few things that hold up in practice:

  • Split on structure, not character count. Markdown headings are a gift here. The author already told you where the topics change; use that instead of guessing at 500-character boundaries.
  • Overlap slightly. Carrying a sentence or two across the boundary keeps ideas that straddle a heading from being cut in half.
  • Prepend the note title to each chunk. Cheap, and it stops a chunk from losing all sense of what document it belongs to.

Semantic search alone is not enough

Having argued for embeddings, the honest position is that you want both.

Semantic search is bad at exact matches. Search for an error code, a function name, a person's name, or a specific ticket number, and embeddings will hand you things that are thematically related — which is precisely wrong. You wanted that string. Keyword search nails it instantly.

So: run both, then merge. The standard approach is reciprocal rank fusion, which sounds fancier than it is — score each result by its rank position in each list rather than its raw score, then add the scores. It sidesteps the problem that a cosine similarity of 0.82 and a BM25 score of 4.7 are not on comparable scales.

The result is a search that handles how does login work and ERR_JWT_EXPIRED equally well, which is what you actually need, because you search for both kinds of thing.

One caveat I only learned by measuring it: fusion is not free. When we built an eval for this — twenty questions over a corpus seeded with deliberately confusable notes — equal-weight fusion scored worse than keyword search alone. The reason turned out to be boring and important: the semantic leg in that harness was a bag-of-words fingerprint standing in for a real embedding model, so we were fusing a lexical ranker with another lexical ranker. No independent signal, just noise. Down-weighting the semantic side recovered parity but never beat it.

The lesson is not "hybrid retrieval is overrated." It is that fusion only pays when the two rankers fail differently, and you cannot know whether yours do without measuring. If you take one thing from this section, take that: build the eval before you tune the weights.

Citations are not a nicety

Once retrieval works, you can feed the top chunks to a model and get an answer. This is the part that looks impressive in a demo, and it is also where the tool earns or loses your trust permanently.

An answer without sources is unfalsifiable. It is fluent, it is plausible, and you have no way to check it short of manually searching for what you just asked. If it is subtly wrong — a date off by a month, two decisions conflated, a detail imported from the model's training rather than your notes — you will not catch it. You will act on it.

An answer with citations is a different object. Every claim links to the note it came from. You skim the sources, you see whether they support the claim, you move on in about four seconds. And critically, the failure modes become visible: if the answer cites nothing relevant, you can tell the retrieval missed rather than assuming the answer is right.

There is a second-order effect worth naming. Once answers carry citations, you start noticing which of your notes are load-bearing — which ones keep getting cited. That is real signal about what you actually know versus what you merely wrote down once.

What this looks like when it works

The honest version of a good day with this setup:

You are six months past a decision and cannot remember why you made it. You ask, in ordinary words, why did we not use Firebase. Retrieval finds a chunk from a note you wrote in March, a comment on a task, and a paragraph from a daily entry the week you were deciding. The answer summarizes the tradeoff and cites all three. You click the March note, confirm it, and get back to work.

Total elapsed time: fifteen seconds. Without it: either ten minutes of searching, or — far more likely — you re-derive the decision from scratch and possibly reach a different conclusion than your past self did, for no reason other than that you could not find what you already knew.

That second outcome is the expensive one, and it is invisible. Nobody logs the hours lost to re-deriving. It just feels like work.

Practical notes if you are building this

A few things that are not obvious until you hit them:

Re-embed on edit, not on read. Embeddings go stale when the text changes. Do it on write, in the background. Users will not wait.

Store the model version alongside the vector. You will change embedding models eventually, and vectors from different models are not comparable. Without a version column you will have a silently broken index and a very confusing afternoon.

pgvector is usually enough. If you are already running Postgres, an HNSW index over a vector column will comfortably serve a personal knowledge base — we are talking thousands to hundreds of thousands of chunks, not billions. A dedicated vector database is a second system to operate, and at this scale it buys you very little.

Show the scores while developing. Displaying the similarity score next to each result during development is the fastest way to build intuition about what your thresholds should be, and to notice when chunking has gone wrong.

The point

Retrieval quality is the product. A brilliant model over bad retrieval produces confident nonsense. An ordinary model over good retrieval produces something that feels like a colleague who has actually read your notes.

Most of the work — and most of the value — is in the unglamorous half: how you chunk, how you combine ranking signals, and whether you show your sources.

Frequently asked questions

What is semantic search?

Semantic search finds notes by meaning rather than exact words. Each note is converted into an embedding — a vector of numbers positioned so that related ideas sit near each other. A query is converted the same way, then compared by cosine similarity. This is why searching 'how does login work' can surface a note titled 'JWT refresh flow' that shares no keywords.

Why does keyword search fail for personal knowledge bases?

Traditional keyword search matches words, not meaning. You search 'how do users log in' and have a note titled 'JWT refresh flow' — zero shared keywords, zero results. For a personal knowledge base this failure is common because you're searching for things you half-remember. If you remembered the exact words, you wouldn't need to search.

What is chunking and why does it matter for embeddings?

Chunking is splitting notes into sections before embedding them. A vector is a fixed size regardless of input length — embed a 3,000-word note covering four topics and you get a vector that means 'approximately the average of all of that,' which is nothing in particular. Split on markdown headings, overlap slightly across boundaries, and prepend the note title to each chunk.

Should I use semantic search or keyword search?

Use both. Semantic search is bad at exact matches (error codes, function names, ticket numbers), while keyword search nails those instantly. The standard approach is reciprocal rank fusion: score each result by its rank position in each list, then merge. This handles 'how does login work' and 'ERR_JWT_EXPIRED' equally well.

Why are citations important in AI search?

An answer without sources is unfalsifiable — it's fluent and plausible, but you have no way to check it. An answer with citations lets you verify every claim in seconds. Citations also create a second-order effect: you start noticing which notes are load-bearing (which ones keep getting cited), giving you signal about what you actually know.

Keep reading

Stop re-deriving what you already knew.

EngineerOS indexes every note and task as you write, then answers questions with citations back to the source.