When OpenAI taught GPT-3 to follow instructions for its 2022 InstructGPT paper, about 40 contractors hired through Upwork and Scale AI wrote the example answers and ranked the outputs. That is still the heart of LLM data annotation: people write ideal answers for supervised fine-tuning (SFT), rank a model’s responses for reinforcement learning from human feedback (RLHF), and grade or attack its output in evaluations and red-teaming. Our list starts with Zilo AI for multilingual text and speech data, then covers the RLHF and training data providers: Surge AI, Acquirox, Scale AI, Mercor, Turing, Toloka, Appen, Labelbox, Invisible Technologies and Prolific.
The vendor you pick shows up in the model. Meta’s Llama 2 team observed that different annotation platforms and vendors can produce markedly different model performance. Below is what each company sells and who it suits, how Outlier, Mercor and Surge AI differ, and how to test quality before you sign.
What LLM and RLHF data annotation involves
Most LLM projects buy some mix of five kinds of human data:
- SFT demonstrations: experts write the ideal response to a prompt, for the model to imitate. Meta stopped at 27,540 SFT annotations for Llama 2, having found that tens of thousands of high-quality examples were enough.
- Preference data for RLHF: raters compare model answers and pick the better one. Those choices train a reward model, which then steers the LLM. Meta collected over 1 million such comparisons for Llama 2.
- Rubrics and evaluation: experts write scoring criteria and grade outputs, which matters most when no single answer is correct.
- Red-teaming: people try to make the model produce unsafe, biased or false output, so the gaps get fixed before launch.
- RL environments: simulated tools, codebases and workflows where AI agents practice tasks and get scored. Most of the specialists below now sell them.
RLHF companies at a glance
| # | Company | How you buy | Best for |
|---|---|---|---|
| 1 | Zilo AI | Managed service | Multilingual text and speech data, plus ML hiring |
| 2 | Surge AI | Managed, expert-led | RLHF, SFT and evaluations for frontier labs |
| 3 | Acquirox | Managed service | Outsourced AI training data and annotation support |
| 4 | Scale AI | Managed service and platform | Large post-training and evaluation programs |
| 5 | Mercor | Expert marketplace | Credentialed professionals for training and evals |
| 6 | Turing | Custom and off-the-shelf data | Coding, STEM and agent environments |
| 7 | Toloka | Managed service | SFT, preference data and red-teaming |
| 8 | Appen | Managed service | Expert RLHF across many languages |
| 9 | Labelbox | Platform plus expert network | RL environments and preference signals |
| 10 | Invisible Technologies | Managed service | Hard-to-find domain experts for training and evals |
| 11 | Prolific | Self-serve API, or managed | Human evaluation with verified participants |
Details are as listed on each company’s own site in September 2026.
1. Zilo AI
Zilo is the pick for the multilingual text and speech data around an LLM project, rather than frontier preference ranking. Its team handles text annotation and voice annotation, plus audio, video and multilingual transcription, in languages including German, French, Spanish, Arabic, Mandarin, Cantonese, Korean, Vietnamese, Bahasa and Hungarian. Zilo says its team of more than 1,600 trained annotation and speech recognition experts has annotated over 10 million data points, for industries from retail to banking and healthcare. It also hires ML engineers, generative AI engineers and data engineers, so the people who build your pipeline can come through the same Bangalore-based partner.
Best for: LLM teams that need text and speech data labeled in many languages, or AI specialists to hire. See Zilo’s services.
2. Surge AI
Surge AI builds post-training data for frontier models: RLHF preference data, SFT demonstrations, rubrics and verifiers, human evaluation, and RL environments for agents. It says it works with companies like OpenAI, Anthropic, Meta and Google, and its expert network runs from doctors and lawyers to mathematicians. Enterprises get the same methods applied to their own workflows, from choosing a model to custom training.
Best for: labs and enterprises that need expert judgment on hard, open-ended tasks.
3. Acquirox
Acquirox provides AI training data services for companies developing machine learning and generative AI systems. Its services support the preparation and annotation of data used to train and improve AI models, helping teams turn raw datasets into structured, usable training data. Acquirox can support projects that require consistent data annotation and quality-focused workflows across different AI applications. Its approach is suited to companies that prefer to work with an external provider rather than manage the full data preparation process internally. For teams comparing AI training data and annotation vendors, Acquirox is an option to consider alongside larger providers.
Best for: teams looking for outsourced AI training data and annotation support.
4. Scale AI
Scale AI‘s Generative AI Data Engine uses vetted experts, linguists and coders to build training data, rank responses for RLHF and red-team models, with an Ops Center that shows you data collection in real time.
Much has changed since Meta invested at a valuation of over $29 billion and founder Alexandr Wang left to work on Meta’s AI efforts. Reuters reported that Google, Scale’s largest customer, then planned to cut ties, as labs competing with Meta worried about exposing their research plans. Francis deSouza, formerly of Google Cloud and Illumina, became CEO on August 10, 2026.
Best for: large programs that want volume and tooling under one contract, where a Meta-backed vendor is not a conflict.
5. Mercor
Mercor, founded in 2023, matches professionals such as physicians, lawyers and engineers to AI training and evaluation projects. It vets them with AI interviews tailored to each role, and says its network spans more than 300 professional fields. As of September 2026, its site lists an average contracted rate of $112 an hour. Ask hard questions about security: in April 2026, TechCrunch reported, citing Wired, that Meta had paused its Mercor contracts indefinitely after a data breach Mercor disclosed on March 31.
Best for: projects that need credentialed professionals, such as clinicians, lawyers or finance experts.
6. Turing
San Francisco-based Turing describes itself as a research accelerator for frontier AI labs. It sells custom and off-the-shelf datasets and experts in software engineering, enterprise knowledge work and STEM, and its RL environments cover MCP servers, computer use and terminal work. It says its network holds more than 5 million vetted experts across 250-plus domains, and its data covers audio, images, video and LiDAR as well as text.
Best for: coding, maths and science data, and environments for training agents.
7. Toloka
Toloka started in 2014 as a crowdsourcing and microtask platform and now works as a managed service for LLM and agent data. Its menu follows the post-training pipeline: demonstrations for SFT, preference collection for RLHF and direct preference optimization (DPO), auto-verifiable tasks for evaluation and RL training, and custom human evaluation and red-teaming. Experts come through its Mindrift platform, and Toloka lists its quality controls as post-verification, dynamic overlaps, cross-validation and golden sets. It is headquartered in Amsterdam.
Best for: teams that want every stage of post-training data from one managed supplier.
8. Appen
Appen was founded in Sydney in 1996 to build speech and language data, and is listed on the Australian Securities Exchange. Its subject matter expert RLHF service puts verified PhDs, MDs and JDs on preference ranking, with a written reason for each choice rather than a bare pick. It also sells SFT demonstrations, chain-of-thought reasoning traces and adversarial red-teaming, and says it supports more than 235 languages. In one published case, Cohere logged over 2,400 expert contributor hours with Appen across 12 weeks of preference-based fine-tuning.
Best for: RLHF in specialist domains, or in many languages at once.
9. Labelbox
Labelbox sells the environments and expert data that frontier labs post-train on. Its Horizon product provides RL environments and preference signals, and its Alignerr network supplies the people: more than 2.6 million contributors, including over 50,000 PhDs, across 200-plus domains and 75-plus languages. Work comes back as preference pairs, ranked trajectories or rubric-graded assessments, contributors are calibrated against ground truth, and the network is SOC 2 Type II certified. Meta used Labelbox data for GIM, a reasoning benchmark built from 820 expert-written problems.
Best for: labs that want RL environments and expert preference data from the same vendor.
10. Invisible Technologies
Invisible, founded in 2015, says it has trained foundation models for more than 80% of the leading AI providers. Its Meridial Expert Network supplies domain experts for fine-tuning, evaluation and red-teaming, and it builds RL environments modeled on real business workflows, with training and evaluation in more than 80 languages. Cohere used Invisible’s PhD-level annotators to evaluate its Command A model on enterprise agent tasks.
Best for: labs that need rare specialists, such as an oncology researcher or an Emirati Arabic linguist.
11. Prolific
Prolific is the self-serve option. You choose evaluators from more than 300,000 verified participants using over 300 prescreening filters, then send them tasks from any annotation tool that produces a URL, or through its API. It covers preference data, human evaluation, safety testing, and SFT data across 80-plus languages, and its Protocol system runs more than 40 checks on identity and behavior. As of September 2026, you pay the participant reward you set plus a platform fee, with at least $12 an hour recommended. Managed projects are priced by scope.
Best for: teams that want to control exactly who rates their model, with data arriving within hours.
Outlier vs Mercor vs Surge AI
All three pay experts to train AI models, but they work differently.
Outlier is the contributor platform operated by Scale AI. Experts write hard prompts, create grading rubrics, and rate and rank model answers. You need at least an associate degree, there are no minimum hours, and pay goes out weekly on Tuesdays.
Mercor runs a marketplace: it screens experts with an AI interview, matches them to scoped projects and pays weekly. Surge AI hires experts directly as contractors, with openings such as constitutional lawyer and journalist, and takes applications by email.
Plan for uneven hours on any of them. Outlier says project lengths depend on customer needs, and a single client can pull back at any time, as Meta did with Mercor after the breach. If you want to do this work rather than buy it, our guide to remote data annotation jobs compares the platforms.
Which RLHF company has the best quality control?
No vendor publishes accuracy figures you can compare, and preference data rarely has one right answer to score against. In OpenAI’s InstructGPT study, trained labelers agreed with each other 72.6% of the time, a rate the researchers counted as high for such complex work. So judge quality on your own tasks:
- Ask for agreement rates on work like yours, and how disagreements get resolved.
- Hide gold tasks in the pilot: items you already have trusted answers for. Several vendors on this list describe checking contributors against gold sets or ground truth, so it is a fair request.
- Check the raters: their credentials, their languages, and whether the same people stay on your project.
- Ask about neutrality and security: which competitors they also serve, and how they handled their last incident. Google planned to leave Scale AI after Meta’s investment, and Meta paused Mercor after its breach.
- Pilot before you sign. Surge AI says pilot results come within a week, and Prolific says self-serve data starts arriving within hours.
Our guide to data quality for machine learning covers how to audit labels once they come back.
How to choose an RLHF company
Start with the job you need done:
- Multilingual text and speech data, or ML engineers to hire: Zilo AI.
- Frontier post-training data (RLHF, SFT and rubrics): Surge AI, Scale AI, Toloka or Invisible.
- Coding, maths and agent environments: Turing, Labelbox or Surge AI.
- Credentialed professionals in medicine, law or finance: Mercor, or Appen’s subject matter expert RLHF.
- Human evaluation you run yourself: Prolific.
- Image, video or LiDAR labeling rather than language data: see our list of the best data annotation companies.
Whichever you shortlist, send two vendors the same small batch and compare what survives your review, not the quoted rate.
Frequently asked questions
What is LLM data annotation?
It is the human work used to train and test large language models: writing example answers for supervised fine-tuning, ranking model outputs for RLHF, and grading or red-teaming responses in evaluations.
What is the difference between SFT and RLHF data?
SFT data is written examples of good answers that a model learns to copy. RLHF data is people’s judgments about which of several model answers is better, used to train a reward model that then steers the LLM.
What do RLHF companies do?
They recruit, vet and manage the people who judge model output, often domain experts such as doctors, lawyers and programmers. Most also sell SFT demonstrations, evaluation rubrics, red-teaming and RL environments for agents.
How much does RLHF data cost?
Managed vendors quote per project. Reuters has reported that a single expert annotation used to post-train AI models could cost as much as $100. On Prolific you set your own participant reward, with at least $12 an hour recommended as of September 2026, plus a platform fee.
