← Insights

Why we are rebuilding our practice around agents

Seven years of putting models into production taught us where the hours really go. Agents finally let us automate that part of the work, not just the prediction.

  • Agents
  • Strategy

For most of our history, the model was the easy part. We built forecasts, risk scores and classifiers that worked well in a notebook. The hard part was everything around them: collecting the inputs, cleaning them, moving the output into the system where someone could act on it, and handling the cases the model was never meant to see.

That surrounding work was usually done by people. A planner copied the forecast into a schedule. An analyst looked up the address the matcher could not resolve. A coordinator called a phone line to confirm what a document said. The model saved time in the middle of a process that stayed mostly manual at both ends.

What changed

Two things changed in the past year.

First, models became reliable enough to use tools. When Anthropic released Claude 4 in May, its Opus and Sonnet models scored 72.5% and 72.7% on SWE-bench Verified, a benchmark built from real software tasks. The same announcement made Claude Code generally available. A model that can read a repository, run a test and fix what broke can also read a case file, query a system and write a result back.

Second, the industry started to agree on how to build these systems. Anthropic's Building effective agents separates workflows, where models and tools follow predefined code paths, from agents, where the model directs its own process. Its advice is to find the simplest solution possible and only add complexity when needed. That matches what we learned shipping ML: most value comes from boring, well-instrumented pipelines.

What we are doing about it

We are rebuilding our practice around one idea: automate the work around the decision, not just the decision. In practice that means:

  • Agents that take a case from intake to done, with a person handling the exceptions.
  • Voice agents for the phone calls that still hold up back-office work.
  • Document agents that read what arrives as PDF or fax and turn it into structured data.
  • Evaluation sets built from real cases, so "it works" is a number, not a feeling.

We are not starting from zero. The habits that kept our models alive in production (versioned data, monitoring, rollback, knowledge transfer) are exactly what agents need.

A healthy dose of caution

Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. We think that is right, and it is why we start every engagement by agreeing what "correct" means and what a case should cost.

Over the next months we will write here about what we are building and what we are reading. Short notes, about twice a month.

Sources

  1. Introducing Claude 4, Anthropic, May 22, 2025.
  2. Building effective agents, Anthropic, December 19, 2024.
  3. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Gartner, June 25, 2025.