Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
fix: correct role of the beta hyperparameter on the DPO loss (#818)
Increasing beta leads to less divergence between the new model and the reference model.
A
Andreas Yin committed
8e170312fec72450ab41d9232a143a2677afe5a4
Parent: 32965e0
Committed by GitHub <noreply@github.com>
on 9/13/2025, 1:21:38 AM