Login
Download
Skill UI
Browse and discover
11317+
curated skills
All
Development
Artificial Intelligence
Design & Creative
Product & Business
Data Science
Marketing
Soft Skills
Productivity
Engineering
Languages
Search
Preference Optimization
, found
2
results
Default
Newest
Most Downloaded
Simple Preference Optimization for LLM Alignment
simpo-training
Orchestra-Research/AI-Research-SKILLs
113
SimPO (Simple Preference Optimization) is a state-of-the-art, reference-free method designed for aligning Large Language Models (LLMs) using human preference data. It is an efficient alternative to DPO and PPO, notably outperforming DPO without requiring a separate reference model. It is ideal for practitioners who need faster, simpler, and more resource-efficient fine-tuning for preference alignment.
View Details
Train and Fine-Tune LLMs Using TRL
trl-training
sickn33/antigravity-awesome-skills
316
This skill provides expert capabilities for training and fine-tuning transformer language models using the TRL (Transformers Reinforcement Learning) library. It supports state-of-the-art post-training methods, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and Reward Model training. Use this skill to align and customize foundation models for advanced tasks via CLI commands.
View Details
1
Language
简体中文
English