solidrust / gemma-2-9b-it-AWQ

huggingface.co
Total runs: 64.8K
24-hour runs: 0
7-day runs: 0
30-day runs: 63.9K
Model's Last Updated: September 03 2024
text-generation

Introduction of gemma-2-9b-it-AWQ

Model Details of gemma-2-9b-it-AWQ

google/gemma-2-9b-it AWQ

How to use
Install the necessary packages
pip install --upgrade autoawq autoawq-kernels
Example Python code
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer, TextStreamer

model_path = "solidrust/gemma-2-9b-it-AWQ"
system_message = "You are gemma-2-9b-it, incarnated as a powerful AI. You were created by google."

# Load model
model = AutoAWQForCausalLM.from_quantized(model_path,
                                          fuse_layers=True)
tokenizer = AutoTokenizer.from_pretrained(model_path,
                                          trust_remote_code=True)
streamer = TextStreamer(tokenizer,
                        skip_prompt=True,
                        skip_special_tokens=True)

# Convert prompt to tokens
prompt_template = """\
<|im_start|>system
{system_message}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant"""

prompt = "You're standing on the surface of the Earth. "\
        "You walk one mile south, one mile west and one mile north. "\
        "You end up exactly where you started. Where are you?"

tokens = tokenizer(prompt_template.format(system_message=system_message,prompt=prompt),
                  return_tensors='pt').input_ids.cuda()

# Generate output
generation_output = model.generate(tokens,
                                  streamer=streamer,
                                  max_new_tokens=512)
About AWQ

AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings.

AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead.

It is supported by:

Runs of solidrust gemma-2-9b-it-AWQ on huggingface.co

64.8K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
63.9K
30-day runs

More Information About gemma-2-9b-it-AWQ huggingface.co Model

gemma-2-9b-it-AWQ huggingface.co

gemma-2-9b-it-AWQ huggingface.co is an AI model on huggingface.co that provides gemma-2-9b-it-AWQ's model effect (), which can be used instantly with this solidrust gemma-2-9b-it-AWQ model. huggingface.co supports a free trial of the gemma-2-9b-it-AWQ model, and also provides paid use of the gemma-2-9b-it-AWQ. Support call gemma-2-9b-it-AWQ model through api, including Node.js, Python, http.

gemma-2-9b-it-AWQ huggingface.co Url

https://huggingface.co/solidrust/gemma-2-9b-it-AWQ

solidrust gemma-2-9b-it-AWQ online free

gemma-2-9b-it-AWQ huggingface.co is an online trial and call api platform, which integrates gemma-2-9b-it-AWQ's modeling effects, including api services, and provides a free online trial of gemma-2-9b-it-AWQ, you can try gemma-2-9b-it-AWQ online for free by clicking the link below.

solidrust gemma-2-9b-it-AWQ online free url in huggingface.co:

https://huggingface.co/solidrust/gemma-2-9b-it-AWQ

gemma-2-9b-it-AWQ install

gemma-2-9b-it-AWQ is an open source model from GitHub that offers a free installation service, and any user can find gemma-2-9b-it-AWQ on GitHub to install. At the same time, huggingface.co provides the effect of gemma-2-9b-it-AWQ install, users can directly use gemma-2-9b-it-AWQ installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

gemma-2-9b-it-AWQ install url in huggingface.co:

https://huggingface.co/solidrust/gemma-2-9b-it-AWQ

Url of gemma-2-9b-it-AWQ

gemma-2-9b-it-AWQ huggingface.co Url

Provider of gemma-2-9b-it-AWQ huggingface.co

solidrust
ORGANIZATIONS

Other API from solidrust