From our blog

Check out our latest news and updates.

Embeddings explained without the linear algebra
Embeddings explained without the linear algebra
06/09/2026 — [email protected]

An embedding turns text into coordinates, and nearby coordinates mean related meaning. That single idea is what makes se...

The context window is a budget, not a memory
The context window is a budget, not a memory
02/09/2026 — [email protected]

Bigger context windows did not give models memory. They gave you a larger envelope to fill on every single request — and...

Why agents fail on long tasks
Why agents fail on long tasks
19/08/2026 — [email protected]

An agent that handles five steps beautifully can fall apart at twenty. The reason is rarely reasoning — it is that the c...

Giving an agent memory without giving it amnesia
Giving an agent memory without giving it amnesia
17/08/2026 — [email protected]

Memory is not a feature you enable. It is a set of decisions about what to keep, what to summarise and what to let go —...

RAG in one page: retrieve, rank, answer
RAG in one page: retrieve, rank, answer
03/08/2026 — [email protected]

Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...

Chunking is the part of RAG nobody tunes
Chunking is the part of RAG nobody tunes
01/08/2026 — [email protected]

Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...

Prompt caching: the cheapest speedup you are not using
Prompt caching: the cheapest speedup you are not using
30/07/2026 — [email protected]

If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...