Qwen2 orca_mini_v7_7b is trained with various SFT Datasets
Passionate about Generative AI? I help companies to privately train and deploy custom LLM/MLLM affordably. For startups, I can even assist with securing GPU grants to get you started. Let's chat!
By providing proper credit and attribution, you are granted permission to use this model as a foundational base for further Full fine tuning, DPO, PPO or ORPO tuning and any kind of Merges.
I actively encourage users to customize and enhance the model according to their specific needs, as this version is designed to be a comprehensive general model.
Dive in and innovate!
Evaluation
Coming Soon..
Example Usage
Here is the ChatML prompt format
<|im_start|>system
You are Orca Mini, a helpful AI assistant.<|im_end|>
<|im_start|>user
Hello Orca Mini, what can you do for me?<|im_end|>
<|im_start|>assistant
Below shows a code example on how to use this model
from transformers import AutoModel, AutoTokenizer
model_slug = "pankajmathur/orca_mini_v7_7b"
model = AutoModel.from_pretrained(model_slug)
tokenizer = AutoTokenizer.from_pretrained(model_slug)
messages = [
{"role": "system", "content": "You are Orca Mini, a helpful AI assistant."},
{"role": "user", "content": "Hello Orca Mini, what can you do for me?"}
]
gen_input = tokenizer.apply_chat_template(messages, return_tensors="pt")
model.generate(**gen_input)
To handle extensive inputs exceeding 32,768 tokens, we utilize
YARN
, a technique for enhancing model length extrapolation, ensuring optimal performance on lengthy texts.
For deployment, we recommend using vLLM. You can enable the long-context capabilities by following these steps:
Install vLLM
: You can install vLLM by running the following command.
Configure Model Settings
: After downloading the model weights, modify the
config.json
file by including the below snippet:
{"architectures":["Qwen2ForCausalLM"],// ..."vocab_size":152064,// adding the following snippets"rope_scaling":{"factor":4.0,"original_max_position_embeddings":32768,"type":"yarn"}}
This snippet enable YARN to support longer contexts.
Model Deployment
: Utilize vLLM to deploy your model. For instance, you can set up an openAI-like server using the command:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{ "model": "pankajmathur/orca_mini_v7_7b", "messages": [ {"role": "system", "content": "You are Orca Mini, a helpful AI assistant."}, {"role": "user", "content": "Hello Orca Mini, what can you do for me?"} ] }'
Note
: Presently, vLLM only supports static YARN, which means the scaling factor remains constant regardless of input length,
potentially impacting performance on shorter texts
. We advise adding the
rope_scaling
configuration only when processing long contexts is required.
Runs of emplitude rubywork on huggingface.co
6
Total runs
0
24-hour runs
0
3-day runs
1
7-day runs
0
30-day runs
More Information About rubywork huggingface.co Model
rubywork huggingface.co is an AI model on huggingface.co that provides rubywork's model effect (), which can be used instantly with this emplitude rubywork model. huggingface.co supports a free trial of the rubywork model, and also provides paid use of the rubywork. Support call rubywork model through api, including Node.js, Python, http.
rubywork huggingface.co is an online trial and call api platform, which integrates rubywork's modeling effects, including api services, and provides a free online trial of rubywork, you can try rubywork online for free by clicking the link below.
emplitude rubywork online free url in huggingface.co:
rubywork is an open source model from GitHub that offers a free installation service, and any user can find rubywork on GitHub to install. At the same time, huggingface.co provides the effect of rubywork install, users can directly use rubywork installed effect in huggingface.co for debugging and trial. It also supports api for free installation.