Reading every survey answer: free-text analysis then and now
In 2020 we built NLP to code workplace survey comments into themes. The same job today costs a fraction as much, runs in hours, and explains itself. A retrospective.
Between 2020 and 2022 we worked with an HR technology company on workplace inclusion and employee experience surveys. The multiple-choice questions were easy. The open-ended ones were where the insight was, and where the work piled up.
A typical project, especially during the pandemic, came with thousands of free-text answers to questions like "what would make you feel safer returning to the office?" The client needed themes, counts per theme, examples, and differences between groups, fast enough to act on.
How we did it then
The pipeline we built at the time was solid ML for its day: text cleaning, embeddings, clustering to propose themes, a classifier trained on a hand-labelled sample to assign answers to themes, and a review pass to fix what the classifier got wrong. It worked, and it was a big step up from coding answers in a spreadsheet.
But every new survey meant a new labelled sample, because every survey asked different questions. The labelling was the bottleneck.
How we do it now
The same job today looks like this:
- A language model reads a sample and proposes a codebook: themes, definitions, inclusion rules and example quotes.
- A person edits the codebook. This is where the analyst's judgment goes, and it takes an hour instead of a week.
- The model codes every answer against the codebook and returns the theme, a short rationale and a confidence.
- A random sample and all low-confidence answers go to a reviewer. Their agreement rate with the model is reported with the results.
No training set per survey. The rationale for each label makes the output easy to audit, which matters when results go to leadership.
Why the economics changed
The cost of running a model over every answer collapsed. Stanford's 2025 AI Index found the inference cost of a system performing at the level of GPT-3.5 dropped over 280-fold between November 2022 and October 2024. Epoch AI estimated that the price to reach GPT-4's performance on one benchmark fell 40x per year, with declines across milestones ranging from 9x to 900x per year.
At those prices, "code every answer, explain every label" is no longer a luxury. It is the default.
What stays the same
The codebook is still a human decision. Agreement with a reviewer is still the metric. And the people who answered the survey still deserve to have their words read carefully, which is easier to promise when a model reads all of them.
Sources
- The 2025 AI Index Report, Stanford HAI, April 2025.
- LLM inference prices have fallen rapidly but unequally across tasks, Epoch AI, March 12, 2025.