← Insights

From crawler scripts to research agents: a 2020 project, rebuilt

We once built a crawler per website to collect property data for insurance risk models. With computer-use models and agent SDKs, most of that brittle code becomes one research agent.

  • Insurance
  • Agents
  • Retrospective

In 2020 we built data acquisition pipelines for a national insurance data provider. The goal was simple to state: collect public information about properties, listings and vacant land, match it to addresses, and feed it into risk models.

The work was not simple. Each source needed its own crawler and parser. When a site changed its layout, the parser broke, usually quietly. Address matching needed hand-written rules and a person to review the cases the rules could not settle. PDFs were their own project. The models at the end of the pipeline were the easy part.

What has changed in the last month alone

Two releases make this kind of work look very different.

  • Anthropic's Claude Sonnet 4.5 scored 61.4% on OSWorld, a benchmark of real computer tasks, up from 42.2% for Sonnet 4 four months earlier. It shipped with the Claude Agent SDK, "the same infrastructure behind Claude Code," for building agents beyond coding.
  • Google introduced the Gemini 2.5 Computer Use model, a specialised model for controlling browsers and user interfaces, in public preview through the Gemini API and Vertex AI.

How we would build it today

Instead of one crawler per site, we would build one research agent with a narrow job:

  1. Take an address or parcel as input.
  2. Search and browse approved sources, reading pages the way a person would, including ones that changed last week.
  3. Extract a fixed schema: building type, year, use, listing status, evidence links.
  4. Attach the evidence for each field, and a confidence.
  5. Send low-confidence records to a reviewer, with the pages already open.

The deterministic parts stay deterministic. Where a source has a stable API or a clean file, a script is still cheaper and more reliable than a model, and we keep it. The agent handles the long tail that used to consume most of the maintenance.

What does not change

The discipline. Allowed sources are listed, not discovered. Every record carries its evidence. A sample is checked by a person every week, and the error rate is a number on a dashboard. Before, we measured crawler uptime; now we measure extraction accuracy against a labelled set.

The result is the same data, with far less code to keep alive.

Sources

  1. Introducing Claude Sonnet 4.5, Anthropic, September 29, 2025.
  2. Introducing the Gemini 2.5 Computer Use model, Google, October 7, 2025.