Day 3 of Ask TanStack Query: the manual is now a stack of cards
Yesterday the docs were 270 files on my computer. Today they are 905 searchable cards in a database. The cards are real, and a lot of them are cut in the wrong place.
The short version
You cannot hand an assistant a whole manual for every question. Each page has to be cut into small pieces, and each piece needs a way to be found later.
Today the program did that cut, turned every piece into a list of numbers that stand for its meaning, and stored the result. It is the step between the pile of docs from Day 2 and an assistant that can look things up.
905 cards from 270 pages
The database is one table called chunks. Each row is one card: which page it came from, the title, the text, and 1,536 numbers that represent the meaning of that text.
A second copy of the text is kept for exact-word search, the way Cmd+F works. That part waits until a later day.
The table is locked. A public key in a browser cannot read it. Only the server, using a secret key, can.
The cutter is dumb on purpose
The first cutter does not look for sentences or headings. It takes about 1,500 characters, then starts the next card 200 characters early so a sentence on the edge appears on both cards. That overlap is the only clever part.
def chunk_fixed(text, size=1500, overlap=200):
if not text:
return []
chunks = []
start = 0
while True:
chunks.append(text[start : start + size])
if start + size >= len(text):
break
start += size - overlap
return chunksStarting this simple leaves a "before" version. Later, a smarter cutter can be measured against it.
What a bad card looks like
Ten random cards were enough to see the problem.
One card from a page called InfiniteQueryObserverOptions begins ether errors should be thrown. The word was Whether. The cut landed in the middle of it.
Another, from TimeoutManager, begins meoutManager } from '@tanstack/query-core'. That is the middle of a code sample, not the start of an explanation.
The save worked. The cards are the messy part. A character count does not know where a word or a code block ends.
What I set up on Day 3
Day 3 was six commits.
- The table.
sql/01_schema.sqlcreateschunks, turns on meaning search, and locks the table. - The database library. Python can now talk to Supabase.
- Shared setup.
app/clients.pyholds the OpenAI client and the database client so every later script uses the same ones. - The cutter.
ingest/chunking.pyslices each page into fixed-size cards. - The save.
ingest/embed.pycut 270 pages into 905 cards, asked OpenAI for the meaning-numbers 50 cards at a time, and stored them. - The journal. Two bad cards are written down, so the next cutter has something to beat.
Frequently asked questions
What is a chunk?
One small piece of a docs page, stored as its own row. The assistant will search these pieces instead of reading all 270 pages for every question.
What is an embedding?
A list of 1,536 numbers that stands for the meaning of a chunk. Chunks about the same idea get similar numbers, so a question can find the nearby cards.
Why are some cards cut in the middle of a word?
The cutter counts characters. It does not look for the end of a sentence, a heading, or a code block. That is a known flaw, kept on purpose so a better cutter can be compared with it later.
Why keep the bad cards?
They are the "before" picture. If a later change cuts on headings instead, these examples are how you tell whether the cards got better.
Next up
Next, a question goes in and an answer comes back from these cards, with the page it came from. The bad cuts stay in the journal until the cutter gets smarter.
Related: Day 2 of Ask TanStack Query: I went and got the books and the quiz
