Marketplaces / wshobson/agents / llm-finetuning/preference-optimization
llm-finetuning/preference-optimization
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.
skill · safety 96/100 (warn) · @ 367cb6a
Install in kendex: kendex add --skill llm-finetuning/preference-optimization after subscribing to wshobson/agents.
This package carries no rendered README.