From our blog

Check out our latest news and updates.

An agent is a loop, not a personality
An agent is a loop, not a personality
23/08/2026 — [email protected]

Strip away the branding and an agent is a small loop: call the model, run the tool it asked for, feed the result back, r...

Tool calling is an API contract the model can read
Tool calling is an API contract the model can read
21/08/2026 — [email protected]

A tool definition is documentation with a schema attached. Write it for a capable colleague who has never seen your syst...

Why agents fail on long tasks
Why agents fail on long tasks
19/08/2026 — [email protected]

An agent that handles five steps beautifully can fall apart at twenty. The reason is rarely reasoning — it is that the c...

Giving an agent memory without giving it amnesia
Giving an agent memory without giving it amnesia
17/08/2026 — [email protected]

Memory is not a feature you enable. It is a set of decisions about what to keep, what to summarise and what to let go —...

When a workflow beats an agent
When a workflow beats an agent
15/08/2026 — [email protected]

If you already know the steps, do not ask a model to rediscover them on every request. A fixed pipeline with model calls...

RAG in one page: retrieve, rank, answer
RAG in one page: retrieve, rank, answer
03/08/2026 — [email protected]

Retrieval-augmented generation is three steps and a lot of tuning. The architecture is simple; the quality lives almost...

Chunking is the part of RAG nobody tunes
Chunking is the part of RAG nobody tunes
01/08/2026 — [email protected]

Teams spend weeks on rerankers and leave chunking at the default. It is usually the other way round that pays — how you...

Prompt caching: the cheapest speedup you are not using
Prompt caching: the cheapest speedup you are not using
30/07/2026 — [email protected]

If every request begins with the same two thousand tokens of instructions, you are paying full price to resend them. Ord...

Streaming responses without breaking your UI
Streaming responses without breaking your UI
28/07/2026 — [email protected]

Streaming makes a slow response feel fast. It also means rendering text that is syntactically incomplete, and handling a...

Rate limits, retries and the backoff you actually need
Rate limits, retries and the backoff you actually need
26/07/2026 — [email protected]

Every AI feature meets a 429 eventually. Whether that is a blip or an outage depends on retry logic written before you n...

How to evaluate an AI feature before you ship it
How to evaluate an AI feature before you ship it
24/07/2026 — [email protected]

You cannot unit test "is this a good answer", but you can build a set of real cases with known-good outputs. An afternoo...

The cost model of an AI feature
The cost model of an AI feature
22/07/2026 — [email protected]

Per-token pricing looks trivial until you multiply by retries, conversation history and the context you resend on every...