AI
Indika/
Back to Media
Whitepaper · 24 pages · 2026

The State of Human Feedback Data, 2026

A survey of 43 AI teams — 21 frontier labs, 22 enterprises — on how RLHF and evaluation budgets are shifting toward domain-specialist review, and what separates programs that keep improving from programs that plateau.

Fieldwork: Jan–May 2026Sample: n=43 AI teamsMethod: structured interview + program data

Executive summary

Between January and May 2026, we surveyed 43 AI teams on how their human-feedback programs have changed since 2024. The headline finding: the share of RLHF budget spent on specialist review, versus generalist raters, rose from an average of 12% in 2024 to 34% in 2026. Teams that made that shift earlier report meaningfully better eval-plateau outcomes, faster iteration cycles, and fewer post-launch regressions traced back to human-feedback quality.

This report walks through four findings from that survey, the reasoning behind each, and what we'd recommend if you're scoping or re-scoping a human-feedback program this year.

Finding 1 — Specialist-review spend has nearly tripled

We asked respondents what share of their human-feedback budget went to domain-specialist reviewers (vs. trained generalist raters) in 2024 and again in 2026.

12%
2024
19%
2025
34%
2026
Fig. 1 — Share of RLHF budget allocated to domain-specialist review, 2024–2026 (n=43)

The trend holds across both frontier labs and enterprises, though frontier labs moved first — their 2025 figure (24%) already exceeded where enterprises sat in 2026 (19% at the start of the year, rising through it).

Finding 2 — Generalist-only programs plateau nearly 3x more often

We asked teams whether they'd experienced a measurable eval plateau (defined as less than 2% improvement across three consecutive fine-tuning cycles) within six months of a human-feedback program's launch.

Generalist-only programs68%
Tiered generalist + specialist programs24%
Fig. 2 — Share of programs reporting an eval plateau within 6 months, by review structure

The 44-point gap is the largest effect size in the entire survey. Respondents running tiered review consistently described the same mechanism: generalist raters catch the "obviously wrong" 80–85% of cases cheaply, while specialists catch the ambiguous, domain-judgment cases that are exactly the ones blocking further improvement once a model has already cleared the easy bar.

Finding 3 — Domain disagreement rate is a leading indicator, not a lagging one

Teams that track specialist-vs-generalist disagreement rate as an explicit ongoing metric caught quality regressions an average of 5.2 weeks earlier than teams that only monitored raw agreement or downstream eval scores.

Week 1Week 4Week 8Week 12
— Disagreement-rate signal- - Downstream eval score
Fig. 3 — Illustrative regression-detection lag: disagreement-rate monitoring vs. eval-score monitoring alone

Finding 4 — Rubric co-authorship beats rubric length

Programs where domain experts co-authored the evaluation rubric — rather than having it handed to them by ML engineers — reported 19% higher reviewer-to-reviewer agreement, regardless of how detailed the rubric was. Rubric length itself showed no significant correlation with agreement once co-authorship was controlled for.

76%
Agreement — rubric written by ML engineers only
95%
Agreement — rubric co-authored with domain experts

Also notable: red-teaming is shifting from checklist to simulation

58% of frontier-lab respondents now use goal-based, multi-turn red-teaming rather than static prompt lists — up from 21% in 2024. Respondents attributed this shift to diminishing returns from public jailbreak lists, which model providers now patch within weeks of disclosure, versus persistent value from adversarial roleplay that adapts across turns.

Methodology

Respondents were sourced from Indika's existing client base and outreach to AI teams not currently working with us, to reduce selection bias toward our own recommended practices. Interviews were structured (45–60 minutes) and supplemented with anonymized program metrics where respondents consented to share them (31 of 43 teams). All figures in this report are aggregated; no individual team's data is identifiable.

Recommendations

If your team is still running flat generalist review, the data suggests the highest-leverage change isn't more raters — it's three concrete moves: route the top 10–15% most-disagreed-upon cases to domain specialists, co-author your rubric with the specialists who'll apply it rather than handing them an ML-engineer draft, and track disagreement rate explicitly rather than relying on downstream eval scores as your only signal. Teams in our survey who made these three changes together saw them pay for themselves within one quarter through reduced re-training cycles.

Full respondent-level methodology and the complete question set are available on request — reach out to our RLHF team for the extended dataset.

Want to benchmark your own program?

We'll walk through how your current setup compares to this year's data.

Talk to our RLHF team →
AI
Indika/

The data foundation enterprises trust to build reliable AI, since 2021.

Solutions
Industries
Company
For experts
Trust
ISO 27001 & 9001SOC 2GDPR compliant
© 2026 Indika AI. All rights reserved.
PrivacyTerms