From our blog

Check out our latest news and updates.

Why temperature changes the answer, not the knowledge
Why temperature changes the answer, not the knowledge
08/09/2026 — [email protected]

Turning temperature down does not make a model more accurate. It makes it more repeatable — and confusing the two is how...

Embeddings explained without the linear algebra
Embeddings explained without the linear algebra
06/09/2026 — [email protected]

An embedding turns text into coordinates, and nearby coordinates mean related meaning. That single idea is what makes se...

The context window is a budget, not a memory
The context window is a budget, not a memory
02/09/2026 — [email protected]

Bigger context windows did not give models memory. They gave you a larger envelope to fill on every single request — and...

Why models hallucinate, and what actually reduces it
Why models hallucinate, and what actually reduces it
31/08/2026 — [email protected]

Hallucination is not a glitch that a better model will one day remove. It is what generation does when it has nothing to...

Choosing between a large model and a small one
Choosing between a large model and a small one
25/08/2026 — [email protected]

Most production traffic does not need the largest model available. Routing by task rather than defaulting to the top of...

Why agents fail on long tasks
Why agents fail on long tasks
19/08/2026 — [email protected]

An agent that handles five steps beautifully can fall apart at twenty. The reason is rarely reasoning — it is that the c...

Giving an agent memory without giving it amnesia
Giving an agent memory without giving it amnesia
17/08/2026 — [email protected]

Memory is not a feature you enable. It is a set of decisions about what to keep, what to summarise and what to let go —...

When a workflow beats an agent
When a workflow beats an agent
15/08/2026 — [email protected]

If you already know the steps, do not ask a model to rediscover them on every request. A fixed pipeline with model calls...

Reviewing AI-generated UI: a checklist
Reviewing AI-generated UI: a checklist
05/08/2026 — [email protected]

Generated interfaces fail in a predictable set of places. Knowing which ones turns review from a vague unease into a lis...

RAG in one page: retrieve, rank, answer
RAG in one page: retrieve, rank, answer
03/08/2026 — [email protected]

Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...

Chunking is the part of RAG nobody tunes
Chunking is the part of RAG nobody tunes
01/08/2026 — [email protected]

Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...

Prompt caching: the cheapest speedup you are not using
Prompt caching: the cheapest speedup you are not using
30/07/2026 — [email protected]

If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...