Back to the blog
Blog/AI and human validation: how to analyse discourse on social media without trusting the model blindly
2 min read

September 28, 2026

AI and human validation: how to analyse discourse on social media without trusting the model blindly

A study classifies 403 headlines with a language model and checks the result against human reviewers. The order in which it does so is the lesson.

Classifying thousands of messages with a language model is easy and cheap. Knowing whether it did it well is another matter. A study published as a preprint on 23 September by independent researcher Shahan Ahmed is a good example, small and reproducible, of how to answer that second question.

What it does

It analyses how seven English-language outlets in Bangladesh headlined the country's 2026 measles outbreak. It gathers 403 headlines and asks a language model to classify each one on two dimensions: sentiment and stance, the latter with four closed categories.

It then compares what the model says with what human reviewers say on 153 headlines. Agreement is measured with Cohen's kappa, which runs from 0 (the agreement you would get by chance) to 1 (total agreement):

  • Stance: 0.89.
  • Sentiment: 0.75.

What it finds

Negative headlines went from 56% to 88% as the outbreak progressed, and those amplifying risk from 44% to 84%. Blame, by contrast, appeared rarely, in around 9% of headlines. And in 86% of those cases it was attributed to systemic failures, not to named political figures.

The detail that interests us

Press headlines are not social media, but the method is the one needed to analyse discourse on TikTok or Instagram. And one figure deserves attention: people and the model agree more on stance (0.89) than on sentiment (0.75). Sentiment is more ambiguous than it seems. That is why a "positive sentiment" percentage with no validation behind it deserves little trust.

How we read it at Trendia

The order matters:

  1. Closed categories, defined before classifying.
  2. A sample annotated by hand by people.
  3. Measure the agreement between the people and the model.
  4. Only then, apply AI to the whole set.

This is the approach we work with. In the project with the UOC on mental health on TikTok there was a taxonomy of 15 labels, 350 videos annotated by hand as a reference and, after that, the classification of the full sample, more than 2,200 videos.

AI lets you reach thousands of pieces. Human validation is what lets you trust the result.

You can see the project with the UOC or how we work with research teams.

Source: Threat Amplified, Blame Restrained: LLM-Assisted Media Framing Analysis of the 2026 Bangladesh Measles Outbreak, arXiv, 23 September 2026.