PKU-SafeRLHF-orpo-72k

<span style="color: red;">Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful.</span>

original PKU-SafeRLHF datasets (click 🔗 for more details)

what's the advantage of this train dataset over the original one ?

  • standard chosen/rejected format of preference datasets : make 'chosen' and 'rejected' according to 'better_response_id'
  • only one file : merge three train datasets(Alpaca-7B、Alpaca2-7B、Alpaca3-8B) into one file