Insights

Notes from building agents in production.

What we learn shipping agentic systems for real operations — including the parts that went badly. Written by the engineers doing the work, not a content team.

Deflection is not resolution

Support dashboards reward the wrong number. Why we rebuilt a client's metric before we rebuilt their agent — and what changed when we did.

The eval suite is the deliverable

Clients ask for an agent. What actually makes the agent survive contact with production is the test suite around it. A practical guide to building one from historical data.

Fewer topics, more authority

The counterintuitive design lesson from a year of support agents: narrowing scope while widening permissions beats the reverse, every time.

When the model changes underneath you

Provider upgrades are not backwards-compatible in the ways that matter. How we structure regression testing so a version bump is a Tuesday, not an incident.

Retrieval quality is a product decision

Chunking strategy, re-ranking, and the "I don't know" threshold are not infrastructure details. They determine whether anyone trusts the answer.

The use cases we talk clients out of

Four patterns that look like ideal agent candidates and consistently are not — and the question we ask in discovery that surfaces them early.

Guardrails are an architecture, not a prompt

Asking a model nicely to stay in bounds is not a control. Where real constraints live in an agentic system, and how to reason about blast radius.

Article pages are being migrated — titles above reflect our current writing queue. For an early copy of any piece, email contact@flashcycle.ai.

Start the cycle

Rather talk it through?

Most of what we write started as a question a client asked on a discovery call. Bring yours.