Paper2022
Red Teaming Language Models to Reduce Harms
Ganguli et al.
Reports on systematically attacking a company's own language models with thousands of red-team conversations, finding that larger and RLHF-trained models are not automatically harder to provoke into harmful output.
33 pageslink checked 17 Sept 2026FreeIntermediate