Social media data for research

With human validation at every phase.

We turn TikTok and Instagram videos and comments into data ready for research. With official access to Meta and TikTok. The annotated corpus, the data and the report belong to your team.

OutreachTestimonyHumourTrendExpert advicePersonal story

What you can research with social media data

Questions this process can answer. Yours can be another one.

  • A conversation

    How a conversation about a topic is built, who enters it, and which formats carry it.

  • A taxonomy

    Which categories appear in a corpus, and how they change across periods or accounts.

  • A discourse

    Social media discourse analysis. Which emotion, narrative and music travel with it.

  • A comparison

    How the same topic compares across two moments, or two communities.

Social media content analysis methodology, in 8 steps

A process shared with your team and validated at every phase. Each result can be traced and described in a paper.

  1. Diagram of a research question and its scope

    Hypothesis and scope

    We define the questions, state the hypotheses, and agree the limits, the sample, and the success criteria.

  2. Diagram of a branching taxonomy

    Taxonomy

    We turn the questions into categories and multilabels. They are reviewed before a single data point is classified.

  3. Data flowing through a validation check

    Extraction check

    A pilot checks sources, fields, coverage, and quality. The extraction method is validated before it scales.

  4. Videos and comments flowing into a database

    Extraction

    We collect the full corpus and normalise it so videos, text, interactions, and metadata share one structure.

  5. A mesh representing model training

    Models

    When the question needs it, we train models to detect patterns, classify content, or produce variables that do not exist in a public dataset.

  6. Manual selection of a reference set

    Golden set

    We annotate between 15% and 20% of the data by hand. That set trains, checks, and measures performance.

  7. Labels spreading from a sample to the full set

    Inference

    We apply the validated system to the rest of the corpus and leave a consistent base, ready to analyse.

  8. Validated data turning into analysis

    Validation and analysis

    We present the method and the metrics. Then come the analyses, the predictions, and the interpretation.

Where the TikTok and Instagram data comes from

The TikTok and Instagram research data comes from sources you can explain to an ethics board and describe in a paper.

The UOC project: a dataset of TikTok videos on mental health

A closed project, with figures that can be shown.

Universitat Oberta de Catalunya

Mental health on TikTok

UOC needed to understand what mental-health content circulates on TikTok. Trendia turned thousands of videos and comments into a structured, validated database ready for analysis. Trendia delivered the validated corpus; the conclusions belong to the UOC team.

videos analysed
2,200+
comments
67,000+
multilabels in the taxonomy
15
hypotheses validated
10
videos annotated by hand
350
to have the data ready
14 days
Question
How is the mental-health conversation built on TikTok?
Method
Multilabel taxonomy, a hand-annotated golden set, and inference across the full sample.
Deliverable
Validated data for analysing narratives, formats, and patterns.

What is delivered: annotated corpus, data and report

The team keeps the material to defend the result and keep working.

  • Illustrative

    Annotated corpus

    The hand-reviewed set that supports training and validation.

    Data
  • Illustrative

    Dataset

    A clean, normalised base, enriched with the study's variables.

    Data
  • Illustrative

    Report

    The answer to the hypotheses, with the method and the limits beside it.

    Results

Scope changes the detail. Traceability and the handoff do not.

Social media data for research, frequently asked questions

Where does the data come from?

From public sources and official access to Meta and TikTok. Sources, fields and coverage are documented and travel with the deliverable.

Does it fit our ethics board and GDPR?

We work with public data and a documented protocol that can be presented to the board. If your institution has its own requirements, we review them before we start.

Can I publish with this data?

Yes. The annotated corpus, the data and the conclusions belong to your team. We document how the corpus was built so the method can be described in the paper.

Who defines the taxonomy?

It is defined with your team from the hypotheses and reviewed before a single data point is classified.

How long does a project take?

It depends on the scope. In the UOC project, the data was ready in 14 days.

Is there a published price?

No. The budget depends on the question, the corpus and the deliverables.

Which question do you need answered?

Tell us the question, the corpus you have in mind, and the deadline. We will say whether the process fits.

Propose your research question