Your AI assistant is only as good as what it can retrieve
In short
Semantic search finds notes by meaning rather than exact words. Each note is converted into an embedding — a vector of numbers positioned so that related ideas sit near each other — and a query is converted the same way, then compared by cosine similarity. This is why searching "how does login work" can surface a note titled "JWT refresh flow" that shares no keywords. For a personal knowledge base the practical rules are: chunk notes into sections rather than embedding whole documents, combine semantic search with keyword search since exact terms still matter, and always show citations so a wrong answer is verifiable rather than merely convincing.
Almost every AI note-taking tool demos the same way. Someone types a question, and a well-formatted answer appears. It looks like magic and it usually is not — because the hard part is not writing the answer. Language models have been good at that for years.
The hard part is finding the right four paragraphs out of your two thousand notes and putting them in front of the model. Get that wrong and you have built a very confident liar with access to your journal.
This is a retrieval problem, and it is worth understanding, because whether a tool is useful or annoying is decided almost entirely here.
Keyword search and its failure mode
Traditional search matches words. You type auth, it finds documents containing auth. Postgres full-text search adds stemming and ranking, so running matches run, and this is genuinely good technology — fast, predictable, cheap, and when you remember the exact term you want, unbeatable.
The failure is specific: it cannot match on meaning. You search how do users log in and you have a note titled JWT refresh flow. Zero shared keywords. Zero results. The note is sitting right there, it is the exact note you want, and the search engine has no way to know.
For a personal knowledge base this failure is common rather than rare, because you are searching for things you half-remember. If you remembered the exact words you would not need to search — you would just open the note.
What embeddings do
An embedding model reads text and returns a list of numbers — a vector, typically several hundred to a couple thousand of them. The useful property is how those vectors are arranged: text with similar meaning produces vectors that are close together in that space, even when the words differ.
So how do users log in and JWT refresh flow land near each other, because the model has learned from a great deal of text that these concepts are related. Search becomes geometry. Embed the query, find the nearest note vectors, return them ranked by distance. The usual distance measure is cosine similarity, which measures the angle between two vectors and ignores their length — you get a score from -1 to 1, and in practice anything above about 0.75 is worth looking at.
That is the whole idea. Everything else is engineering.
Chunking is where it goes wrong
Here is the mistake that quietly ruins most implementations: embedding entire documents.
A vector is a fixed size regardless of input length. Embed four words and you get a vector that means those four words. Embed a 3,000-word note covering database schema, auth, deployment, and a rant about your CI provider, and you get a vector that means approximately the average of all of that — which is to say, nothing in particular. It sits in a vague middle region of the space, moderately close to everything and genuinely close to nothing. It will rank mediocre-ly for every query and win none of them.
The fix is chunking: split notes into sections and embed each one. Now the auth section has its own vector that is sharply about auth.
The tradeoffs are real and there is no universally correct answer. Chunk too small and you lose context — a paragraph that says "this approach did not work" is useless without knowing which approach. Chunk too large and you are back to averaging. A few things that hold up in practice:
- Split on structure, not character count. Markdown headings are a gift here. The author already told you where the topics change; use that instead of guessing at 500-character boundaries.
- Overlap slightly. Carrying a sentence or two across the boundary keeps ideas that straddle a heading from being cut in half.
- Prepend the note title to each chunk. Cheap, and it stops a chunk from losing all sense of what document it belongs to.
Semantic search alone is not enough
Having argued for embeddings, the honest position is that you want both.
Semantic search is bad at exact matches. Search for an error code, a function name, a person's name, or a specific ticket number, and embeddings will hand you things that are thematically related — which is precisely wrong. You wanted that string. Keyword search nails it instantly.
So: run both, then merge. The standard approach is reciprocal rank fusion, which sounds fancier than it is — score each result by its rank position in each list rather than its raw score, then add the scores. It sidesteps the problem that a cosine similarity of 0.82 and a BM25 score of 4.7 are not on comparable scales.
The result is a search that handles how does login work and ERR_JWT_EXPIRED equally well, which is what you actually need, because you search for both kinds of thing.
One caveat I only learned by measuring it: fusion is not free. When we built an eval for this — twenty questions over a corpus seeded with deliberately confusable notes — equal-weight fusion scored worse than keyword search alone. The reason turned out to be boring and important: the semantic leg in that harness was a bag-of-words fingerprint standing in for a real embedding model, so we were fusing a lexical ranker with another lexical ranker. No independent signal, just noise. Down-weighting the semantic side recovered parity but never beat it.
The lesson is not "hybrid retrieval is overrated." It is that fusion only pays when the two rankers fail differently, and you cannot know whether yours do without measuring. If you take one thing from this section, take that: build the eval before you tune the weights.
Citations are not a nicety
Once retrieval works, you can feed the top chunks to a model and get an answer. This is the part that looks impressive in a demo, and it is also where the tool earns or loses your trust permanently.
An answer without sources is unfalsifiable. It is fluent, it is plausible, and you have no way to check it short of manually searching for what you just asked. If it is subtly wrong — a date off by a month, two decisions conflated, a detail imported from the model's training rather than your notes — you will not catch it. You will act on it.
An answer with citations is a different object. Every claim links to the note it came from. You skim the sources, you see whether they support the claim, you move on in about four seconds. And critically, the failure modes become visible: if the answer cites nothing relevant, you can tell the retrieval missed rather than assuming the answer is right.
There is a second-order effect worth naming. Once answers carry citations, you start noticing which of your notes are load-bearing — which ones keep getting cited. That is real signal about what you actually know versus what you merely wrote down once.
What this looks like when it works
The honest version of a good day with this setup:
You are six months past a decision and cannot remember why you made it. You ask, in ordinary words, why did we not use Firebase. Retrieval finds a chunk from a note you wrote in March, a comment on a task, and a paragraph from a daily entry the week you were deciding. The answer summarizes the tradeoff and cites all three. You click the March note, confirm it, and get back to work.
Total elapsed time: fifteen seconds. Without it: either ten minutes of searching, or — far more likely — you re-derive the decision from scratch and possibly reach a different conclusion than your past self did, for no reason other than that you could not find what you already knew.
That second outcome is the expensive one, and it is invisible. Nobody logs the hours lost to re-deriving. It just feels like work.
Practical notes if you are building this
A few things that are not obvious until you hit them:
Re-embed on edit, not on read. Embeddings go stale when the text changes. Do it on write, in the background. Users will not wait.
Store the model version alongside the vector. You will change embedding models eventually, and vectors from different models are not comparable. Without a version column you will have a silently broken index and a very confusing afternoon.
pgvector is usually enough. If you are already running Postgres, an HNSW index over a vector column will comfortably serve a personal knowledge base — we are talking thousands to hundreds of thousands of chunks, not billions. A dedicated vector database is a second system to operate, and at this scale it buys you very little.
Show the scores while developing. Displaying the similarity score next to each result during development is the fastest way to build intuition about what your thresholds should be, and to notice when chunking has gone wrong.
The point
Retrieval quality is the product. A brilliant model over bad retrieval produces confident nonsense. An ordinary model over good retrieval produces something that feels like a colleague who has actually read your notes.
Most of the work — and most of the value — is in the unglamorous half: how you chunk, how you combine ranking signals, and whether you show your sources.