Work

Outcomes we can point to.

Every system below is in production and still running. We report the numbers our clients measured, not the ones that would look best on a slide — including where the honest answer is "a human still reviews this".

40+ agents shipped 98% average eval pass rate Client-owned IP
Fintech

Autonomous reconciliation agent

Build · 7 weeks Shipped 2025 Claude · LangGraph · MCP Full autonomy, review above threshold

The problem

A finance team was closing its books eleven days into every month. Six systems held overlapping ledger data, none agreed exactly, and the rules for what counted as a real discrepancy lived in people's heads.

What we built

An agent that pulls all six sources, normalises them, and works the differences the way an analyst would — checking timing, currency and counterparty quirks. It writes an audit note for every item it clears.

Why it worked

The eval suite was built from two years of historical closes, so we could prove the agent agreed with the analysts before it replaced them. Where it disagreed, the analysts were sometimes wrong.

92%manual work removed
3.4×faster monthly close
11 → 3days to close
SaaS · Support

Tier-1 support that actually resolves

Build · 6 weeks Shipped 2025 Claude · RAG · Zendesk API Resolve, refund, escalate

The problem

The client already had a chatbot. It deflected 40% of tickets in the sense that customers gave up — a win on the dashboard, a loss in churn. Real resolution needed account state and an action.

What we built

A support agent with retrieval over the full docs and scoped write access to billing. It can refund below a threshold, reset a config, or change a plan — and hands off the moment sentiment turns.

Why it worked

We measured resolution, not deflection. That reframing changed the design: fewer topics with real authority beat every topic with none.

71%full auto-resolution
-58%first response time
+22CSAT points
Healthcare ops

Prior-authorisation workflow engine

Build · 8 weeks Shipped 2024 Claude · multi-agent · HIPAA-aligned Drafts only — human signs off

The problem

Prior-auth packets took a clinical team about forty minutes each, mostly spent locating evidence across the chart and formatting it to each payer's requirements. Volume was outgrowing headcount.

What we built

Three agents: one extracts and cites the clinical evidence, one checks the packet against that payer's criteria, one formats and routes it. A clinician reviews and signs every packet.

Why it worked

We said from the first workshop that full autonomy was the wrong goal. The value was removing forty minutes of assembly, not two minutes of judgement — which also made compliance review straightforward.

4.1khours saved per month
99.2%accuracy at review
40 → 6minutes per packet
Logistics

Exception triage for freight operations

Discovery → Build · 9 weeks Shipped 2024 Claude · RAG · internal APIs Read and draft, no dispatch

The problem

A freight operator's exception queue — delays, damage claims, customs holds — was handled by whoever was free. The same exception type got a different answer depending on who picked it up.

What we built

An agent that classifies each exception, retrieves the relevant contract terms and past precedent, and drafts the response with its reasoning attached. Dispatch stays a human decision.

Why it worked

Discovery found the real problem was consistency, not speed. That changed what we built — and the client killed two other candidate use cases in the same two weeks.

-64%time to first response
2.8×queue throughput
~0variance between handlers
Testimonials

Trusted by the teams who shipped with us.

“Flash Cycle delivered an agent that genuinely runs a chunk of our operation. The eval discipline is what made our board comfortable putting it in production.”
BM BilalDirector of Engineering, Grayphite
“We'd been burned by AI demos before. This was the first team that talked like engineers and shipped like one. Six weeks, in production, measurable ROI.”
MMassaki HatanoFounder & CEO, Marvis
Start the cycle

Your operation, on this list.

Tell us where the work piles up. We'll tell you whether an agent genuinely helps — and roughly what it would take to find out.