Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
The DictaLM-2.0-Instruct Large Language Model (LLM) is an instruct fine-tuned version of the
DictaLM-2.0
generative model using a variety of conversation datasets.
This model contains the GPTQ 4-bit quantized version of the instruct-tuned model designed for chat
DictaLM-2.0-Instruct
.
You can view and access the full collection of base/instruct unquantized/quantized versions of
DictaLM-2.0
here
.
Instruction format
In order to leverage instruction fine-tuning, your prompt should be surrounded by
[INST]
and
[/INST]
tokens. The very first instruction should begin with a begin of sentence id. The next instructions should not. The assistant generation will be ended by the end-of-sentence token id.
E.g.
text = """<s>[INST] איזה רוטב אהוב עליך? [/INST]
טוב, אני די מחבב כמה טיפות מיץ לימון סחוט טרי. זה מוסיף בדיוק את הכמות הנכונה של טעם חמצמץ לכל מה שאני מבשל במטבח!</s>[INST] האם יש לך מתכונים למיונז? [/INST]"
This format is available as a
chat template
via the
apply_chat_template()
method:
Example Code
Running this code requires less than 5GB of GPU VRAM.
from transformers import AutoModelForCausalLM, AutoTokenizer
device = "cuda"# the device to load the model onto
model = AutoModelForCausalLM.from_pretrained("dicta-il/dictalm2.0-instruct-GPTQ", device_map=device)
tokenizer = AutoTokenizer.from_pretrained("dicta-il/dictalm2.0-instruct-GPTQ")
messages = [
{"role": "user", "content": "איזה רוטב אהוב עליך?"},
{"role": "assistant", "content": "טוב, אני די מחבב כמה טיפות מיץ לימון סחוט טרי. זה מוסיף בדיוק את הכמות הנכונה של טעם חמצמץ לכל מה שאני מבשל במטבח!"},
{"role": "user", "content": "האם יש לך מתכונים למיונז?"}
]
encoded = tokenizer.apply_chat_template(messages, return_tensors="pt").to(device)
generated_ids = model.generate(encoded, max_new_tokens=50, do_sample=True)
decoded = tokenizer.batch_decode(generated_ids)
print(decoded[0])
# <s> [INST] איזה רוטב אהוב עליך? [/INST]# טוב, אני די מחבב כמה טיפות מיץ לימון סחוט טרי. זה מוסיף בדיוק את הכמות הנכונה של טעם חמצמץ לכל מה שאני מבשל במטבח!</s> [INST] האם יש לך מתכונים למיונז? [/INST]# בטח, הנה מתכון קל מאוד למיונז ביתי:# # מרכיבים:# - 2 ביצים גדולות# - 1 כף חרדל דיז'ון# - 2 כפות# (it stopped early because we set max_new_tokens=50)
Model Architecture
DictaLM-2.0-Instruct follows the
Zephyr-7B-beta
recipe for fine-tuning an instruct model, with an extended instruct dataset for Hebrew.
Limitations
The DictaLM 2.0 Instruct model is a demonstration that the base model can be fine-tuned to achieve compelling performance.
It does not have any moderation mechanisms. We're looking forward to engaging with the community on ways to
make the model finely respect guardrails, allowing for deployment in environments requiring moderated outputs.
Citation
If you use this model, please cite:
@misc{shmidman2024adaptingllmshebrewunveiling,
title={Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities},
author={Shaltiel Shmidman and Avi Shmidman and Amir DN Cohen and Moshe Koppel},
year={2024},
eprint={2407.07080},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.07080},
}
Runs of dicta-il dictalm2.0-instruct-GPTQ on huggingface.co
12
Total runs
-1
24-hour runs
-2
3-day runs
0
7-day runs
4
30-day runs
More Information About dictalm2.0-instruct-GPTQ huggingface.co Model
dictalm2.0-instruct-GPTQ huggingface.co is an AI model on huggingface.co that provides dictalm2.0-instruct-GPTQ's model effect (), which can be used instantly with this dicta-il dictalm2.0-instruct-GPTQ model. huggingface.co supports a free trial of the dictalm2.0-instruct-GPTQ model, and also provides paid use of the dictalm2.0-instruct-GPTQ. Support call dictalm2.0-instruct-GPTQ model through api, including Node.js, Python, http.
dictalm2.0-instruct-GPTQ huggingface.co is an online trial and call api platform, which integrates dictalm2.0-instruct-GPTQ's modeling effects, including api services, and provides a free online trial of dictalm2.0-instruct-GPTQ, you can try dictalm2.0-instruct-GPTQ online for free by clicking the link below.
dicta-il dictalm2.0-instruct-GPTQ online free url in huggingface.co:
dictalm2.0-instruct-GPTQ is an open source model from GitHub that offers a free installation service, and any user can find dictalm2.0-instruct-GPTQ on GitHub to install. At the same time, huggingface.co provides the effect of dictalm2.0-instruct-GPTQ install, users can directly use dictalm2.0-instruct-GPTQ installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
dictalm2.0-instruct-GPTQ install url in huggingface.co: