The following hyperparameters were used during training:
learning_rate: 5e-07
eta: 1000
per_device_train_batch_size: 8
gradient_accumulation_steps: 1
seed: 42
distributed_type: deepspeed_zero3
num_devices: 8
optimizer: RMSProp
lr_scheduler_type: linear
lr_scheduler_warmup_ratio: 0.1
num_train_epochs: 6.0 (stop at epoch=1.0)
Model Citation
@misc{wu2024self,
title={Self-Play Preference Optimization for Language Model Alignment},
author={Wu, Yue and Sun, Zhiqing and Yuan, Huizhuo and Ji, Kaixuan and Yang, Yiming and Gu, Quanquan},
year={2024},
eprint={2405.00675},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Runs of QuantFactory Llama-3-Instruct-8B-SPPO-Iter3-GGUF on huggingface.co
1.2K
Total runs
0
24-hour runs
51
3-day runs
-59
7-day runs
678
30-day runs
More Information About Llama-3-Instruct-8B-SPPO-Iter3-GGUF huggingface.co Model
More Llama-3-Instruct-8B-SPPO-Iter3-GGUF license Visit here:
Llama-3-Instruct-8B-SPPO-Iter3-GGUF huggingface.co is an AI model on huggingface.co that provides Llama-3-Instruct-8B-SPPO-Iter3-GGUF's model effect (), which can be used instantly with this QuantFactory Llama-3-Instruct-8B-SPPO-Iter3-GGUF model. huggingface.co supports a free trial of the Llama-3-Instruct-8B-SPPO-Iter3-GGUF model, and also provides paid use of the Llama-3-Instruct-8B-SPPO-Iter3-GGUF. Support call Llama-3-Instruct-8B-SPPO-Iter3-GGUF model through api, including Node.js, Python, http.
Llama-3-Instruct-8B-SPPO-Iter3-GGUF huggingface.co is an online trial and call api platform, which integrates Llama-3-Instruct-8B-SPPO-Iter3-GGUF's modeling effects, including api services, and provides a free online trial of Llama-3-Instruct-8B-SPPO-Iter3-GGUF, you can try Llama-3-Instruct-8B-SPPO-Iter3-GGUF online for free by clicking the link below.
QuantFactory Llama-3-Instruct-8B-SPPO-Iter3-GGUF online free url in huggingface.co:
Llama-3-Instruct-8B-SPPO-Iter3-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Llama-3-Instruct-8B-SPPO-Iter3-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3-Instruct-8B-SPPO-Iter3-GGUF install, users can directly use Llama-3-Instruct-8B-SPPO-Iter3-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3-Instruct-8B-SPPO-Iter3-GGUF install url in huggingface.co: