It looks simple in class, here I go deeper
Articles on programming, mathematics and AI. Not the how, but the why.
Naive, hybrid and agentic RAG: where each one fails
An assistant with a search engine answers "no", with two citations, and gets it wrong. Why naive RAG fails, what hybrid search fixes and what it can't, and how agentic RAG lets the model decide what the code used to decide: whether to search, for what, where, and when to stop.
Read the articleWhy AdamW for training language models, and what is changing
Of all the optimisation methods that exist, language models are trained with the simplest one, and with one specific variant of it: AdamW. Why the only information you can afford at that scale is the gradient, how every improvement since buys back on credit the curvature Newton had for free, and what has started to change in 2026.
Read the articleThe evolution of the tokeniser: from words to bytes
The Spanish sentence "El niño juega en el jardín" cost twelve tokens in 2019 and costs seven today, the same as its English translation. The story of why: from the word tokeniser to BPE, WordPiece, Unigram and the byte floor.
Read the articleSign in to be notified when a new article goes up