10 posts found
A model does not look anything up and does not decide what is true. It estimates which token comes next. Almost everythi...
Models do not see characters or words. They see tokens — and once you know how text becomes tokens, several odd behaviou...
An embedding turns text into coordinates, and nearby coordinates mean related meaning. That single idea is what makes se...
Bigger context windows did not give models memory. They gave you a larger envelope to fill on every single request — and...
Hallucination is not a glitch that a better model will one day remove. It is what generation does when it has nothing to...
Memory is not a feature you enable. It is a set of decisions about what to keep, what to summarise and what to let go —...
Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...
Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...
If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...
Per-token pricing looks trivial until you multiply by retries, conversation history and the context you resend on every...