Quantization methods compatible with latest llama.cpp from June 6, commit 2d43387.
Provided Files
Name
Quant method
Bits
Size
Max RAM required, no GPU offloading
Use case
nous-hermes-13b.ggmlv3.q2_K.bin
q2_K
2
5.43 GB
7.93 GB
New k-quant method. Uses GGML_TYPE_Q4_K for the attention.vw and feed_forward.w2 tensors, GGML_TYPE_Q2_K for the other tensors.
nous-hermes-13b.ggmlv3.q3_K_L.bin
q3_K_L
3
6.87 GB
9.37 GB
New k-quant method. Uses GGML_TYPE_Q5_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else GGML_TYPE_Q3_K
nous-hermes-13b.ggmlv3.q3_K_M.bin
q3_K_M
3
6.25 GB
8.75 GB
New k-quant method. Uses GGML_TYPE_Q4_K for the attention.wv, attention.wo, and feed_forward.w2 tensors, else GGML_TYPE_Q3_K
nous-hermes-13b.ggmlv3.q3_K_S.bin
q3_K_S
3
5.59 GB
8.09 GB
New k-quant method. Uses GGML_TYPE_Q3_K for all tensors
nous-hermes-13b.ggmlv3.q4_0.bin
q4_0
4
7.32 GB
9.82 GB
Original llama.cpp quant method, 4-bit.
nous-hermes-13b.ggmlv3.q4_1.bin
q4_1
4
8.14 GB
10.64 GB
Original llama.cpp quant method, 4-bit. Higher accuracy than q4_0 but not as high as q5_0. However has quicker inference than q5 models.
nous-hermes-13b.ggmlv3.q4_K_M.bin
q4_K_M
4
7.82 GB
10.32 GB
New k-quant method. Uses GGML_TYPE_Q6_K for half of the attention.wv and feed_forward.w2 tensors, else GGML_TYPE_Q4_K
nous-hermes-13b.ggmlv3.q4_K_S.bin
q4_K_S
4
7.32 GB
9.82 GB
New k-quant method. Uses GGML_TYPE_Q4_K for all tensors
nous-hermes-13b.ggmlv3.q5_0.bin
q5_0
5
8.95 GB
11.45 GB
Original llama.cpp quant method, 5-bit. Higher accuracy, higher resource usage and slower inference.
nous-hermes-13b.ggmlv3.q5_1.bin
q5_1
5
9.76 GB
12.26 GB
Original llama.cpp quant method, 5-bit. Even higher accuracy, resource usage and slower inference.
nous-hermes-13b.ggmlv3.q5_K_M.bin
q5_K_M
5
9.21 GB
11.71 GB
New k-quant method. Uses GGML_TYPE_Q6_K for half of the attention.wv and feed_forward.w2 tensors, else GGML_TYPE_Q5_K
nous-hermes-13b.ggmlv3.q5_K_S.bin
q5_K_S
5
8.95 GB
11.45 GB
New k-quant method. Uses GGML_TYPE_Q5_K for all tensors
nous-hermes-13b.ggmlv3.q6_K.bin
q6_K
6
10.68 GB
13.18 GB
New k-quant method. Uses GGML_TYPE_Q8_K - 6-bit quantization - for all tensors
nous-hermes-13b.ggmlv3.q8_0.bin
q8_0
8
13.83 GB
16.33 GB
Original llama.cpp quant method, 8-bit. Almost indistinguishable from float16. High resource use and slow. Not recommended for most users.
Model Card: Nous-Hermes-13b
Model Description
Nous-Hermes-13b is a state-of-the-art language model fine-tuned on over 300,000 instructions. This model was fine-tuned by Nous Research, with Teknium and Karan4D leading the fine tuning process and dataset curation, Redmond AI sponsoring the compute, and several other contributors. The result is an enhanced Llama 13b model that rivals GPT-3.5-turbo in performance across a variety of tasks.
This model stands out for its long responses, low hallucination rate, and absence of OpenAI censorship mechanisms. The fine-tuning process was performed with a 2000 sequence length on an 8x a100 80GB DGX machine for over 50 hours.
Model Training
The model was trained almost entirely on synthetic GPT-4 outputs. This includes data from diverse sources such as GPTeacher, the general, roleplay v1&2, code instruct datasets, Nous Instruct & PDACTL (unpublished), CodeAlpaca, Evol_Instruct Uncensored, GPT4-LLM, and Unnatural Instructions.
Additional data inputs came from Camel-AI's Biology/Physics/Chemistry and Math Datasets, Airoboros' GPT-4 Dataset, and more from CodeAlpaca. The total volume of data encompassed over 300,000 instructions.
Collaborators
The model fine-tuning and the datasets were a collaboration of efforts and resources between Teknium, Karan4D, Nous Research, Huemin Art, and Redmond AI.
Huge shoutout and acknowledgement is deserved for all the dataset creators who generously share their datasets openly.
Special mention goes to @winglian, @erhartford, and @main_horse for assisting in some of the training issues.
Among the contributors of datasets, GPTeacher was made available by Teknium, Wizard LM by nlpxucan, and the Nous Research Instruct Dataset was provided by Karan4D and HueminArt.
The GPT4-LLM and Unnatural Instructions were provided by Microsoft, Airoboros dataset by jondurbin, Camel-AI datasets are from Camel-AI, and CodeAlpaca dataset by Sahil 2801.
If anyone was left out, please open a thread in the community tab.
Runs of localmodels Nous-Hermes-13B-ggml on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About Nous-Hermes-13B-ggml huggingface.co Model
Nous-Hermes-13B-ggml huggingface.co
Nous-Hermes-13B-ggml huggingface.co is an AI model on huggingface.co that provides Nous-Hermes-13B-ggml's model effect (), which can be used instantly with this localmodels Nous-Hermes-13B-ggml model. huggingface.co supports a free trial of the Nous-Hermes-13B-ggml model, and also provides paid use of the Nous-Hermes-13B-ggml. Support call Nous-Hermes-13B-ggml model through api, including Node.js, Python, http.
Nous-Hermes-13B-ggml huggingface.co is an online trial and call api platform, which integrates Nous-Hermes-13B-ggml's modeling effects, including api services, and provides a free online trial of Nous-Hermes-13B-ggml, you can try Nous-Hermes-13B-ggml online for free by clicking the link below.
localmodels Nous-Hermes-13B-ggml online free url in huggingface.co:
Nous-Hermes-13B-ggml is an open source model from GitHub that offers a free installation service, and any user can find Nous-Hermes-13B-ggml on GitHub to install. At the same time, huggingface.co provides the effect of Nous-Hermes-13B-ggml install, users can directly use Nous-Hermes-13B-ggml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Nous-Hermes-13B-ggml install url in huggingface.co: