A survey of 43 AI teams — 21 frontier labs, 22 enterprises — on how RLHF and evaluation budgets are shifting toward domain-specialist review, and what separates programs that keep improving from programs that plateau.
Between January and May 2026, we surveyed 43 AI teams on how their human-feedback programs have changed since 2024. The headline finding: the share of RLHF budget spent on specialist review, versus generalist raters, rose from an average of 12% in 2024 to 34% in 2026. Teams that made that shift earlier report meaningfully better eval-plateau outcomes, faster iteration cycles, and fewer post-launch regressions traced back to human-feedback quality.
This report walks through four findings from that survey, the reasoning behind each, and what we'd recommend if you're scoping or re-scoping a human-feedback program this year.
We asked respondents what share of their human-feedback budget went to domain-specialist reviewers (vs. trained generalist raters) in 2024 and again in 2026.
The trend holds across both frontier labs and enterprises, though frontier labs moved first — their 2025 figure (24%) already exceeded where enterprises sat in 2026 (19% at the start of the year, rising through it).
We asked teams whether they'd experienced a measurable eval plateau (defined as less than 2% improvement across three consecutive fine-tuning cycles) within six months of a human-feedback program's launch.
The 44-point gap is the largest effect size in the entire survey. Respondents running tiered review consistently described the same mechanism: generalist raters catch the "obviously wrong" 80–85% of cases cheaply, while specialists catch the ambiguous, domain-judgment cases that are exactly the ones blocking further improvement once a model has already cleared the easy bar.
Teams that track specialist-vs-generalist disagreement rate as an explicit ongoing metric caught quality regressions an average of 5.2 weeks earlier than teams that only monitored raw agreement or downstream eval scores.
Programs where domain experts co-authored the evaluation rubric — rather than having it handed to them by ML engineers — reported 19% higher reviewer-to-reviewer agreement, regardless of how detailed the rubric was. Rubric length itself showed no significant correlation with agreement once co-authorship was controlled for.
58% of frontier-lab respondents now use goal-based, multi-turn red-teaming rather than static prompt lists — up from 21% in 2024. Respondents attributed this shift to diminishing returns from public jailbreak lists, which model providers now patch within weeks of disclosure, versus persistent value from adversarial roleplay that adapts across turns.
Respondents were sourced from Indika's existing client base and outreach to AI teams not currently working with us, to reduce selection bias toward our own recommended practices. Interviews were structured (45–60 minutes) and supplemented with anonymized program metrics where respondents consented to share them (31 of 43 teams). All figures in this report are aggregated; no individual team's data is identifiable.
If your team is still running flat generalist review, the data suggests the highest-leverage change isn't more raters — it's three concrete moves: route the top 10–15% most-disagreed-upon cases to domain specialists, co-author your rubric with the specialists who'll apply it rather than handing them an ML-engineer draft, and track disagreement rate explicitly rather than relying on downstream eval scores as your only signal. Teams in our survey who made these three changes together saw them pay for themselves within one quarter through reduced re-training cycles.
Full respondent-level methodology and the complete question set are available on request — reach out to our RLHF team for the extended dataset.
We'll walk through how your current setup compares to this year's data.
Talk to our RLHF team →The data foundation enterprises trust to build reliable AI, since 2021.