From our blog

Check out our latest news and updates.

What a large language model actually predicts
What a large language model actually predicts
12/09/2026 — [email protected]

A model does not look anything up and does not decide what is true. It estimates which token comes next. Almost everythi...

Tokens, not words: how a model reads your text
Tokens, not words: how a model reads your text
10/09/2026 — [email protected]

Models do not see characters or words. They see tokens — and once you know how text becomes tokens, several odd behaviou...

Why temperature changes the answer, not the knowledge
Why temperature changes the answer, not the knowledge
08/09/2026 — [email protected]

Turning temperature down does not make a model more accurate. It makes it more repeatable — and confusing the two is how...

Embeddings explained without the linear algebra
Embeddings explained without the linear algebra
06/09/2026 — [email protected]

An embedding turns text into coordinates, and nearby coordinates mean related meaning. That single idea is what makes se...

Training, fine-tuning and prompting are three different tools
Training, fine-tuning and prompting are three different tools
04/09/2026 — [email protected]

Teams reach for fine-tuning when they need context, and for prompting when they need behaviour. Knowing which problem ea...

RAG in one page: retrieve, rank, answer
RAG in one page: retrieve, rank, answer
03/08/2026 — [email protected]

Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...

Chunking is the part of RAG nobody tunes
Chunking is the part of RAG nobody tunes
01/08/2026 — [email protected]

Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...

Prompt caching: the cheapest speedup you are not using
Prompt caching: the cheapest speedup you are not using
30/07/2026 — [email protected]

If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...

Streaming responses without breaking your UI
Streaming responses without breaking your UI
28/07/2026 — [email protected]

Streaming makes a slow response feel fast. It also means rendering text that is syntactically incomplete, and handling a...

Rate limits, retries and the backoff you actually need
Rate limits, retries and the backoff you actually need
26/07/2026 — [email protected]

Every AI feature meets a 429 eventually. Whether that is a blip or an outage depends on retry logic written before you n...