Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
trl-lib
's Collections
Preference datasets
Stepwise supervision datasets
Prompt-completion datasets
Prompt-only datasets
Unpaired preference datasets
Comparing DPO with IPO and KTO
Online-DPO
Preference datasets
updated
1 day ago
Upvote
-
trl-lib/hh-rlhf-helpful-base
Viewer
•
Updated
1 day ago
•
46.2k
•
60
trl-lib/lm-human-preferences-descriptiveness
Viewer
•
Updated
1 day ago
•
6.26k
•
32
•
1
trl-lib/lm-human-preferences-sentiment
Viewer
•
Updated
1 day ago
•
6.26k
•
28
trl-lib/rlaif-v
Viewer
•
Updated
1 day ago
•
83.1k
•
136
•
3
trl-lib/tldr-preference
Viewer
•
Updated
1 day ago
•
179k
•
258
trl-lib/ultrafeedback_binarized
Viewer
•
Updated
Sep 12, 2024
•
63.1k
•
4.96k
•
6
Upvote
-
Share collection
View history
Collection guide
Browse collections