15 posts found
A model does not look anything up and does not decide what is true. It estimates which token comes next. Almost everythi...
Models do not see characters or words. They see tokens — and once you know how text becomes tokens, several odd behaviou...
Turning temperature down does not make a model more accurate. It makes it more repeatable — and confusing the two is how...
An embedding turns text into coordinates, and nearby coordinates mean related meaning. That single idea is what makes se...
Teams reach for fine-tuning when they need context, and for prompting when they need behaviour. Knowing which problem ea...
Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...
Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...
If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...
Streaming makes a slow response feel fast. It also means rendering text that is syntactically incomplete, and handling a...
Every AI feature meets a 429 eventually. Whether that is a blip or an outage depends on retry logic written before you n...
You cannot unit test "is this a good answer", but you can build a set of real cases with known-good outputs. An afternoo...
Per-token pricing looks trivial until you multiply by retries, conversation history and the context you resend on every...