There are some issues with the model weights in terms of precision. In the next version update, we will roll back some progress and retrain to fix these issues as soon as possible.
Please note:
Do not use "accelerated inference frameworks" like
VLLM
temporarily. Instead, use Transformers for inference. Otherwise, due to precision issues, the output quality will be significantly degraded. If you need faster inference, you can consider using the q8_0 quantization (faster and better than bf16 vllm for this model only) with llama.cpp temporarily or wait for the official version.
To be fixed in the upcoming next version update.
no repetition_penalty!
Please do not use wikitext for quantization calibration because all wikitext have been re-aligned on synthetic dataset, and its distribution differs significantly from the original wikitext.
MT-Bench: 8.5
Some contamination detection if you want to check:
34b-beta huggingface.co is an AI model on huggingface.co that provides 34b-beta's model effect (), which can be used instantly with this CausalLM 34b-beta model. huggingface.co supports a free trial of the 34b-beta model, and also provides paid use of the 34b-beta. Support call 34b-beta model through api, including Node.js, Python, http.
34b-beta huggingface.co is an online trial and call api platform, which integrates 34b-beta's modeling effects, including api services, and provides a free online trial of 34b-beta, you can try 34b-beta online for free by clicking the link below.
CausalLM 34b-beta online free url in huggingface.co:
34b-beta is an open source model from GitHub that offers a free installation service, and any user can find 34b-beta on GitHub to install. At the same time, huggingface.co provides the effect of 34b-beta install, users can directly use 34b-beta installed effect in huggingface.co for debugging and trial. It also supports api for free installation.