Tiarne Hawkins

AI training data

Models are only as good as the data behind them.

Eight years building training data, human data operations and evaluation programmes for the labs and platforms defining frontier AI — and for enterprises that need their own data to be an advantage rather than a liability.

AI training data for

OpenAI logo
Google DeepMind logo
Meta logo
Microsoft logo
AWS logo
Alexa logo
Siri logo
Cohere logo
TikTok logo
xAI logo
OpenAI logo
Google DeepMind logo
Meta logo
Microsoft logo
AWS logo
Alexa logo
Siri logo
Cohere logo
TikTok logo
xAI logo

What I work on.

01

Data strategy

Decide what data you actually need before you buy any of it.

Model goals translated into a data plan: domains, languages, modalities, volumes and the order to build them in — so spend follows capability, not vendor appetite.

  • Data roadmaps
  • Build vs licence vs collect
  • Coverage gap analysis
02

Sourcing & licensing

Data you can defend in a room full of lawyers.

Provenance, consent, rights and territory handled up front, with vendor selection and commercial structures that hold up as models and markets change.

  • Provenance and consent
  • Vendor selection
  • Commercial structures
03

Human data operations

Expert humans, at scale, without quality collapse.

Annotation and expert-contributor programmes designed end to end: recruiting, guidelines, calibration, pay and throughput, run across time zones and languages.

  • Expert networks
  • Guidelines and calibration
  • Throughput and cost models
04

RLHF & preference data

Preference data that moves the model in the direction you meant.

Instruction, preference and reasoning data programmes with rubric design, inter-rater reliability and the feedback loops that keep signal clean as models improve.

  • Rubric design
  • Preference and reasoning data
  • Rater reliability
05

Evaluation & benchmarks

Know whether the model got better, not just different.

Task suites, human evaluation harnesses and domain benchmarks tied to the decisions the model is trusted with — plus the reporting leadership can act on.

  • Human eval harnesses
  • Domain benchmarks
  • Release reporting
06

Safety & red teaming data

Find the failures before your users do.

Adversarial and safety data programmes for models and agents: taxonomies, attack coverage, escalation paths and the evidence trail your risk function will ask for.

  • Harm taxonomies
  • Adversarial data
  • Agentic red teaming
07

Agentic AI data

Teach agents how to reason, act and recover.

Data for systems that do more than generate an answer — they plan, use tools, take actions and pursue goals across workflows. Trajectory design, tool-use data, multi-step task data, agent-to-agent interactions, human escalation, failure recovery and long-horizon evaluation.

  • Agent trajectories
  • Tool-use & actions
  • Multi-agent workflows
08

Synthetic & simulation data

Create the data the real world hasn’t given you yet.

Generate targeted data for sparse, sensitive, dangerous or hard-to-capture scenarios — designed for coverage, not simply volume. Synthetic training and evaluation data, personas, scenario generation, edge cases, rare events and simulated environments for models and agents.

  • Synthetic generation
  • Personas & scenarios
  • Edge-case coverage

Data is where trust starts.

Every safety claim, every evaluation, every agent you let act on its own traces back to the data it learned from and the humans who shaped it. Get that layer right and the rest of the AI programme becomes defensible.