Insights
Notes from the agentic shift.
What is changing in AI, insurance and healthcare operations, and what it takes to put it into production. Short reads, about twice a month.
RSS feedYour agent's test bench is wired to the real world
Anthropic's test agents filed a false police tip and real web forms. For back-office agents, containment has to start with the first evaluation run.
Read the articleDon't rent your agents
Gartner predicts 70% of enterprises will abandon agentic AI built by vendor forward-deployed engineers by 2028. With costs falling 13x a year, owning the capability is the whole game.
Read · 2 minThe AI rulebook timelines moved. Your controls should not
The EU delayed its high-risk AI deadlines, Colorado replaced its AI Act, and insurance regulators are piloting AI evaluation tools. Why we keep building to the stricter standard anyway.
Read · 2 min87% want a human. Design voice agents with an exit
A new Gartner survey says customers accept AI service when a human is reachable. The same principle makes back-office voice agents safer, cheaper and easier to trust.
Read · 2 minStateless MCP and a summer of new models: what changed for builders
A breaking MCP release, a new generation of frontier models and price-performance that keeps falling. Our notes on what to change in production systems, and what to leave alone.
Read · 2 minWhat production agent teams actually measure
In LangChain's survey of 1,340 practitioners, 57% have agents in production and quality is the top barrier, yet only about half run offline evals. Closing that gap is the cheapest quality win available.
Read · 2 minCanada's AI strategy and the 12 percent
Ottawa's new national AI strategy puts adoption by small and mid-sized businesses on the agenda, alongside new privacy legislation. What a Toronto AI shop thinks Canadian firms should do next.
Read · 2 minPrior auth reform, one year later
The insurer pledge turns one, physicians are not convinced, CMS wants e-prior auth for drugs, and voice models just got another step better. Where automation fits in the meantime.
Read · 2 minStore labour forecasting, then and now
We built store-level demand and labour forecasts for automotive-service retailers. Foundation models and new research on scheduling and churn change what we would build today.
Read · 2 minAgentic AI is at the peak of the hype cycle. Good.
Gartner puts agentic AI at the Peak of Inflated Expectations, while the AI Index shows agents still fail about one attempt in three. That gap is exactly where production engineering earns its keep.
Read · 2 minDo document agents still need an OCR pipeline?
New research suggests multimodal models extract business documents about as well from the image alone as with OCR. What that changes for appeals, claims and onboarding packs.
Read · 2 minReading every survey answer: free-text analysis then and now
In 2020 we built NLP to code workplace survey comments into themes. The same job today costs a fraction as much, runs in hours, and explains itself. A retrospective.
Read · 2 minThe cost-per-resolution reality check for customer service AI
Gartner now expects generative AI cost per resolution to pass $3 by 2030, and half of companies that cut service staff to rehire. The lesson is not "less AI." It is "aim it better."
Read · 2 minPrior authorization in 2026: the rules that just took effect
January 1 brought new decision timeframes, denial reasons and public metrics, plus a Medicare model that uses AI for prior auth review. What operations teams should automate first.
Read · 2 minEvals are the product
An agent without an evaluation set is a demo. How we build test sets from real cases, choose between pass@k and pass^k, and decide when an agent is ready to ship.
Read · 2 minMCP goes to a foundation. Why that matters for enterprise integration
The Model Context Protocol now lives at the Linux Foundation, backed by Anthropic, OpenAI and Block. For enterprises, the question shifts from "which model?" to "which systems can agents safely reach?"
Read · 2 minWhat "65% automatable" underwriting looks like inside an MGA
Accenture says up to 65% of underwriting hours can be automated or augmented. In a small commercial MGA, most of those hours are spent assembling evidence, not deciding.
Read · 2 minOSFI E-23 is final. What it means for AI at Canadian insurers
Canada's model risk guideline now explicitly covers AI and machine learning, including third-party models. Teams have until May 2027, and an agent counts.
Read · 2 minFrom crawler scripts to research agents: a 2020 project, rebuilt
We once built a crawler per website to collect property data for insurance risk models. With computer-use models and agent SDKs, most of that brittle code becomes one research agent.
Read · 2 minVoice agents grew up this summer
Speech-to-speech models, native phone calling and tool use in one API change what a voice agent can be. Here is what it means for teams that still run on phone calls.
Read · 2 minThe 95% problem: what the MIT GenAI report gets right
The headline says most generative AI pilots fail. The detail says something more useful: back-office automation pays, and a learning gap, not the model, is what stalls the rest.
Read · 2 minPhone calls to insurers are the most automatable hour in healthcare admin
Physicians report about 13 hours a week on prior authorization, and much of the surrounding work still happens on hold. Voice agents are finally good enough to take those calls.
Read · 2 minOver 40% of agentic projects will be cancelled. How not to be one of them
Gartner and IBM point at the same three failure modes: cost, unclear value and weak controls. Each one has a design answer you can apply before the first line of code.
Read · 2 minWhy we are rebuilding our practice around agents
Seven years of putting models into production taught us where the hours really go. Agents finally let us automate that part of the work, not just the prediction.
Read · 3 min