CausalLM / 14B-DPO-alpha

huggingface.co
Total runs: 261
24-hour runs: -6
7-day runs: -7
30-day runs: -4.3K
Model's Last Updated: February 11 2025
text-generation

Introduction of 14B-DPO-alpha

Model Details of 14B-DPO-alpha

Sorry, it's no longer available on Hugging Face. Please reach out to those who have already downloaded it. If you have a copy, please refrain from re-uploading it to Hugging Face.

Due to repeated conflicts with HF and what we perceive as their repeated misuse of the "Contributor Covenant Code of Conduct," we have lost confidence in the platform and decided to temporarily suspend all new download access requests. It appears to us that HF's original intention has been abandoned in pursuit of commercialization, and they no longer prioritize the well-being of the community.

Demo:

For details, please refer to the version without DPO training: CausalLM/14B .

Model MT-Bench
GPT-4 8.99
GPT-3.5-Turbo 7.94
Zephyr-7b-β (Overfitting) 7.34
Zephyr-7b-α 6.88
CausalLM/14B-DPO-α 7.618868
CausalLM/7B-DPO-α 7.038125

Dec 3, 2023 Rank #1 non-base model, of its size on 🤗 Open LLM Leaderboard, outperforms ALL ~13B chat models.

image/png

It should be noted that this is not a version that continues training on CausalLM/14B & 7B, but rather an optimized version that has undergone DPO training concurrently on a previous training branch, and some detailed parameters may have changed. You will still need to download the full model.

The beta branch will soon be released, employing some aggressive approaches that might be detrimental in certain tasks, in order to achieve better alignment with human preferences, aiming to meet or exceed the GPT-3.5 benchmarks. Stay tuned.

Disclaimer: Please note that the model was trained on unfiltered internet data. Since we do not have the capacity to vet all of it, there may be a substantial amount of objectionable content, pornography, violence, and offensive language present that we are unable to remove. Therefore, you will still need to complete your own checks on the model's safety and filter keywords in the output. Due to computational resource constraints, we are presently unable to implement RLHF for the model's ethics and safety, nor training on SFT samples that refuse to answer certain questions for restrictive fine-tuning.

更多详情,请参见未经DPO训练的版本: CausalLM/14B

需要注意的是,这并不是在 CausalLM/14B & 7B 上继续训练的版本,而是在之前的训练分支上同时进行了 DPO 训练的优化版本,一些细节参数可能发生了变化。 您仍然需要下载完整模型。

很快将会发布beta分支,采用了一些可能不利于某些任务的激进方法,以实现更好地符合人类偏好以接近和超过GPT-3.5基准。敬请期待。

免责声明:请注意,模型是在未经过滤的互联网数据上进行训练的。由于我们无法审核所有数据,可能会出现大量不良内容、色情、暴力和冒犯性语言,我们无法删除这些内容。因此,您仍然需要对模型的安全性进行自己的检查,并对输出中的关键词进行过滤。由于计算资源的限制,我们目前无法为模型的伦理和安全实施RLHF,也无法对拒绝回答某些问题的SFT样本进行训练以进行限制性微调。

Runs of CausalLM 14B-DPO-alpha on huggingface.co

261
Total runs
-6
24-hour runs
-7
3-day runs
-7
7-day runs
-4.3K
30-day runs

More Information About 14B-DPO-alpha huggingface.co Model

More 14B-DPO-alpha license Visit here:

https://choosealicense.com/licenses/wtfpl

14B-DPO-alpha huggingface.co

14B-DPO-alpha huggingface.co is an AI model on huggingface.co that provides 14B-DPO-alpha's model effect (), which can be used instantly with this CausalLM 14B-DPO-alpha model. huggingface.co supports a free trial of the 14B-DPO-alpha model, and also provides paid use of the 14B-DPO-alpha. Support call 14B-DPO-alpha model through api, including Node.js, Python, http.

14B-DPO-alpha huggingface.co Url

https://huggingface.co/CausalLM/14B-DPO-alpha

CausalLM 14B-DPO-alpha online free

14B-DPO-alpha huggingface.co is an online trial and call api platform, which integrates 14B-DPO-alpha's modeling effects, including api services, and provides a free online trial of 14B-DPO-alpha, you can try 14B-DPO-alpha online for free by clicking the link below.

CausalLM 14B-DPO-alpha online free url in huggingface.co:

https://huggingface.co/CausalLM/14B-DPO-alpha

14B-DPO-alpha install

14B-DPO-alpha is an open source model from GitHub that offers a free installation service, and any user can find 14B-DPO-alpha on GitHub to install. At the same time, huggingface.co provides the effect of 14B-DPO-alpha install, users can directly use 14B-DPO-alpha installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

14B-DPO-alpha install url in huggingface.co:

https://huggingface.co/CausalLM/14B-DPO-alpha

Url of 14B-DPO-alpha

14B-DPO-alpha huggingface.co Url

Provider of 14B-DPO-alpha huggingface.co

CausalLM
ORGANIZATIONS

Other API from CausalLM

huggingface.co

Total runs: 8.6K
Run Growth: 415
Growth Rate: 4.80%
Updated:May 25 2024
huggingface.co

Total runs: 689
Run Growth: 466
Growth Rate: 67.63%
Updated:February 11 2025
huggingface.co

Total runs: 291
Run Growth: 82
Growth Rate: 28.18%
Updated:December 10 2023
huggingface.co

Total runs: 221
Run Growth: 85
Growth Rate: 38.46%
Updated:February 11 2025
huggingface.co

Total runs: 121
Run Growth: 83
Growth Rate: 68.60%
Updated:February 15 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 28 2024