AI
Indika/
Back to Media
RLHF · Jul 2026 · 6 min read

Why generic labeling is no longer enough for frontier labs

Eighteen months ago, most frontier labs could get by with a large pool of general-purpose annotators. Today, that approach quietly breaks down — and the failure mode is subtle enough that most teams don't notice until an eval score plateaus.

Data infrastructure

When we started building RLHF programs for frontier labs in 2024, the brief was simple: get enough human raters to rank enough model outputs, and preference-tuning would do the rest. That worked — for a while. Base model capability was rising faster than the nuance of the tasks being evaluated, so a competent generalist rater could usually tell a good response from a bad one.

The plateau nobody talks about

Somewhere around the transition from GPT-4-class models to the current generation, we started seeing a specific failure pattern across client evals: preference-ranking agreement between generalist raters would sit in the high 80s, but domain experts reviewing the same responses would disagree with the generalist consensus 15–20% of the time — and disagree with each other far less. That gap is the signal. It means the remaining errors in a model's outputs are no longer "obviously wrong" — they're wrong in ways that require domain judgment to catch.

A generalist rater can tell you a customer-support response is unhelpful. They usually can't tell you that a clinical intake response skipped a red-flag symptom, that a contract clause subtly shifts liability in the drafter's favor, or that a generated financial model uses a discount rate that doesn't match the stated risk profile. Those are the failure modes that matter now.

What changed in our own programs

We restructured our RLHF delivery around three changes, in order of impact:

1. Rubrics co-written with domain experts, not just prompt engineers. A rubric that says "penalize unhelpful responses" doesn't help a rater catch a missed drug interaction. We now build rubrics with the same specialists who'll apply them — physicians, attorneys, CFAs — so the criteria reflect real domain failure modes, not generic helpfulness heuristics.

2. Tiered review, not flat crowds. Every task now runs through a first pass by trained generalists (catching the obvious cases cheaply) and a second pass by domain specialists on anything flagged as ambiguous or high-stakes. This keeps cost close to generalist-only while catching the errors that actually move eval scores.

3. Consensus scoring with adjudication, not majority vote. When specialist reviewers disagree, we don't average — we route to a senior adjudicator and log the reasoning. That adjudication log has become one of the more valuable artifacts we hand back to clients, because it doubles as documentation for why a preference decision was made.

What this means if you're scoping a program

If your evals have plateaued and you're still running a flat generalist pool, the fix usually isn't more raters — it's the right raters on the right slice of tasks. We typically see teams get more lift from routing 10–15% of ambiguous cases to domain specialists than from doubling generalist throughput.

This is exactly the kind of engagement our Human feedback & evaluation program is built for — and it's why our expert network exists as a distinct capability rather than a generalist-only labeling pool.

Want to talk through your eval setup?

We'll scope a pilot around your specific model stage and use case.

Talk to our RLHF team →
AI
Indika/

The data foundation enterprises trust to build reliable AI, since 2021.

Solutions
Industries
Company
For experts
Trust
ISO 27001 & 9001SOC 2GDPR compliant
© 2026 Indika AI. All rights reserved.
PrivacyTerms