Marketplaces / wshobson/agents / llm-finetuning/preference-optimization
llm-finetuning/preference-optimization
Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.
skill · @ 46891e7
Supported tools: all tools
Install in kendex: kendex add --skill llm-finetuning/preference-optimization after subscribing to wshobson/agents.
Open in app
No app yet? Download kendex.
This package carries no rendered README.