This model was quantized with
GPTQ
and saved in the Marlin format for efficient 4-bit inference. Marlin is a highly optimized inference kernel for 4 bit models.
Inference
Install
nm-vllm
for fast inference and low memory-usage:
pip install nm-vllm[sparse]
Run in a Python pipeline for local inference:
from transformers import AutoTokenizer
from vllm import LLM, SamplingParams
model_id = "neuralmagic/OpenHermes-2.5-Mistral-7B-marlin"
model = LLM(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
messages = [
{"role": "user", "content": "What is synthetic data in machine learning?"},
]
formatted_prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
sampling_params = SamplingParams(max_tokens=200)
outputs = model.generate(formatted_prompt, sampling_params=sampling_params)
print(outputs[0].outputs[0].text)
"""Synthetic data is data that has been artificially created or modified to serve the needs of machine learning and data analysis tasks. It can be generated either through title methods like stochastic simulations or through processes of data augmentation that take original data and modify/manipulate it to create new samples. Synthetic data is often used in machine learning when the available amount of real-world data is insufficient or in cases where the creation of real-world data can be dangerous, costly, or time-consuming."""
Quantization
For details on how this model was quantized and converted to marlin format, run the
quantization/apply_gptq_save_marlin.py
script:
Runs of neuralmagic OpenHermes-2.5-Mistral-7B-marlin on huggingface.co
667
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About OpenHermes-2.5-Mistral-7B-marlin huggingface.co Model
OpenHermes-2.5-Mistral-7B-marlin huggingface.co
OpenHermes-2.5-Mistral-7B-marlin huggingface.co is an AI model on huggingface.co that provides OpenHermes-2.5-Mistral-7B-marlin's model effect (), which can be used instantly with this neuralmagic OpenHermes-2.5-Mistral-7B-marlin model. huggingface.co supports a free trial of the OpenHermes-2.5-Mistral-7B-marlin model, and also provides paid use of the OpenHermes-2.5-Mistral-7B-marlin. Support call OpenHermes-2.5-Mistral-7B-marlin model through api, including Node.js, Python, http.
OpenHermes-2.5-Mistral-7B-marlin huggingface.co is an online trial and call api platform, which integrates OpenHermes-2.5-Mistral-7B-marlin's modeling effects, including api services, and provides a free online trial of OpenHermes-2.5-Mistral-7B-marlin, you can try OpenHermes-2.5-Mistral-7B-marlin online for free by clicking the link below.
neuralmagic OpenHermes-2.5-Mistral-7B-marlin online free url in huggingface.co:
OpenHermes-2.5-Mistral-7B-marlin is an open source model from GitHub that offers a free installation service, and any user can find OpenHermes-2.5-Mistral-7B-marlin on GitHub to install. At the same time, huggingface.co provides the effect of OpenHermes-2.5-Mistral-7B-marlin install, users can directly use OpenHermes-2.5-Mistral-7B-marlin installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
OpenHermes-2.5-Mistral-7B-marlin install url in huggingface.co: