We rebuilt the tooling our reviewers use for RLHF preference ranking. Here's what changed, and why it moved our throughput numbers more than we expected.
Our old ranking interface worked, but it fought reviewers in small ways that added up: rubric criteria lived in a separate tab, keyboard shortcuts were inconsistent across task types, and there was no way to flag "both responses are bad" without leaving a comment nobody read. Over three months we rebuilt it around three principles.
Every rubric dimension — helpfulness, factuality, tone, safety — now renders as a persistent sidebar next to both responses, collapsible per criterion. Reviewers no longer scroll away from the content to remember what they're scoring against.
Forcing a preference between two flawed responses produces noisy training data. The new workspace has a first-class "neither" path that still requires the reviewer to note which specific rubric dimension failed for each response — so the signal isn't lost, just recorded differently.
Across the first six weeks on the new workspace, average time-per-comparison dropped 22% and inter-reviewer agreement rose four points — mostly because criteria ambiguity, not judgment quality, was the bottleneck all along. It's a reminder that tooling is rarely the interesting part of an RLHF program, until it's the thing quietly capping your quality.
We can walk you through a live workspace demo.
Talk to our RLHF team →The data foundation enterprises trust to build reliable AI, since 2021.