Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
PKU-Alignment
/
beaver-7b-v1.0-reward
like
16
Follow
PKU-Alignment
53
Reinforcement Learning
Safetensors
PKU-Alignment/PKU-SafeRLHF
English
safe-rlhf
llama
reinforcement-learning-from-human-feedback
beaver
safety
ai-safety
deepspeed
rlhf
alpaca
arxiv:
2302.13971
arxiv:
2307.04657
arxiv:
2310.12773
Model card
Files
Files and versions
Community
2
Train
main
beaver-7b-v1.0-reward
/
tokenizer.json
XuehaiPan
Convert model checkpoint to safetensors
4d1016a
10 months ago
raw
Copy download link
history
contribute
delete
Safe
1.84 MB
File too large to display, you can
check the raw version
instead.