PyTorch patterns for implementing preference optimization losses (DPO, SimPO, etc.) for LLM training.