跳到主要导航 跳到搜索 跳到主要内容

CLHA: A Simple yet Effective Contrastive Learning Framework for Human Alignment

  • Feiteng Fang
  • , Liang Zhu
  • , Min Yang*
  • , Xi Feng
  • , Jinchang Hou
  • , Qixuan Zhao
  • , Chengming Li
  • , Xiping Hu
  • , Ruifeng Xu
  • *此作品的通讯作者
  • University of Science and Technology of China
  • Shenzhen Institute of Advanced Technology
  • Southern University of Science and Technology
  • Shenzhen MSU-BIT University
  • Harbin Institute of Technology Shenzhen

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Reinforcement learning from human feedback (RLHF) is a crucial technique in aligning large language models (LLMs) with human preferences, ensuring these LLMs behave in beneficial and comprehensible ways to users. However, a longstanding challenge in human alignment techniques based on reinforcement learning lies in their inherent complexity and difficulty in training. To address this challenge, we present a simple yet effective Contrastive Learning Framework for Human Alignment (CLHA) to align LLMs with human preferences directly. CLHA employs a novel rescoring strategy to evaluate the noise within the data by considering its inherent quality and dynamically adjusting the training process. Simultaneously, CLHA utilizes pairwise contrastive loss and adaptive supervised fine-tuning loss to adaptively modify the likelihood of generating responses, ensuring enhanced alignment with human preferences. Using advanced methods, CLHA surpasses other algorithms, showcasing superior performance in terms of reward model scores, automatic evaluations, and human assessments on the widely used “Helpful and Harmless” dataset. For reproducibility, we release our code and data at: https://github.com/calubkk/CLHA.

源语言英语
主期刊名2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings
编辑Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
出版商European Language Resources Association (ELRA)
3325-3334
页数10
ISBN(电子版)9782493814104
出版状态已出版 - 2024
已对外发布
活动Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024 - Hybrid, Torino, 意大利
期限: 20 5月 202425 5月 2024

出版系列

姓名2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings

会议

会议Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024
国家/地区意大利
Hybrid, Torino
时期20/05/2425/05/24

指纹

探究 'CLHA: A Simple yet Effective Contrastive Learning Framework for Human Alignment' 的科研主题。它们共同构成独一无二的指纹。

引用此