This model is the gptq-quantized version of
Salamandra-7b
for speculative decoding.
The model weights are quantized from FP16 to W4A16 (4-bit weights and FP16 activations) using the
GPTQ
algorithm.
Inferencing with this model can be done using
VLLM
.
Salamandra is a highly multilingual model pre-trained from scratch that comes in three different
sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants,
promoted and financed by the Government of Catalonia through the
Aina Project
and the
Ministerio para la Transformación Digital y de la Función Pública
- Funded by EU – NextGenerationEU
within the framework of
ILENIA Project
with reference 2022/TL22/00215337.
This model card corresponds to the gptq-quantized version of Salamandra-7b for speculative decoding.
The entire Salamandra family is released under a permissive
Apache 2.0 license
.
How to Use
The following example code works under
Python 3.9.16
,
vllm==0.6.3.post1
,
torch==2.4.0
and
torchvision==0.19.0
, though it should run on
any current version of the libraries. This is an example of how to create a text completion using the model:
For further information, please send an email to
[email protected]
.
Acknowledgements
We appreciate the collaboration with IBM in this work.
Specifically, the IBM team created gptq-quantized version of the Salamandra-7b model for speculative decoding released here.
Disclaimer
Be aware that the model may contain biases or other unintended distortions.
When third parties deploy systems or provide services based on this model, or use the model themselves,
they bear the responsibility for mitigating any associated risks and ensuring compliance with applicable
regulations, including those governing the use of Artificial Intelligence.
Barcelona Supercomputing Center and International Business Machines shall
not be held liable for any outcomes resulting from third-party use.
salamandra-7b-base-gptq huggingface.co is an AI model on huggingface.co that provides salamandra-7b-base-gptq's model effect (), which can be used instantly with this BSC-LT salamandra-7b-base-gptq model. huggingface.co supports a free trial of the salamandra-7b-base-gptq model, and also provides paid use of the salamandra-7b-base-gptq. Support call salamandra-7b-base-gptq model through api, including Node.js, Python, http.
salamandra-7b-base-gptq huggingface.co is an online trial and call api platform, which integrates salamandra-7b-base-gptq's modeling effects, including api services, and provides a free online trial of salamandra-7b-base-gptq, you can try salamandra-7b-base-gptq online for free by clicking the link below.
BSC-LT salamandra-7b-base-gptq online free url in huggingface.co:
salamandra-7b-base-gptq is an open source model from GitHub that offers a free installation service, and any user can find salamandra-7b-base-gptq on GitHub to install. At the same time, huggingface.co provides the effect of salamandra-7b-base-gptq install, users can directly use salamandra-7b-base-gptq installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
salamandra-7b-base-gptq install url in huggingface.co: