Screening brief
What this video covers
Surge’s Andrew Mauboussin discusses practical aspects of collecting preference data for reinforcement learning from human feedback (RLHF). He outlines production challenges in specifying tasks, gathering responses, and applying quality controls within Surge’s end-to-end RLHF data collection product.
The talk emphasizes risks tied to low-quality RLHF data and reviews technical and operational strategies Surge uses to mitigate those risks while delivering feedback datasets to ML teams at organizations such as Anthropic and OpenAI. The speaker frames the content from his experience building Surge’s systems and workflows for scalable human feedback collection.
Summary and topic guide by HumanData.TV, based on the original publisher’s description and our editorial catalogue. This is not a transcript or independent verification of the speaker’s claims.


