kendex.ai

Marketplaces / wshobson/agents / llm-finetuning/preference-optimization

llm-finetuning/preference-optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

skill · safety 96/100 (warn) · @ 367cb6a

Install in kendex: kendex add --skill llm-finetuning/preference-optimization after subscribing to wshobson/agents.

This package carries no rendered README.