The following parameters were reduced from the original model:
Parameter
Original
Tiny
num_hidden_layers
16
4
hidden_size
2048
2048
intermediate_size
8192
8192
num_attention_heads
32
32
num_key_value_heads
8
8
Checkpoint Structure
This model uses a single
model.safetensors
file containing all weights. The checkpoint structure is identical to the original model, with the standard Llama architecture tensors:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("inference-optimization/Llama-3.2-0.5B-Instruct", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("inference-optimization/Llama-3.2-0.5B-Instruct")
input_ids = tokenizer("According to all known laws", return_tensors="pt").input_ids.to(model.device)
output = model.generate(input_ids, max_new_tokens=20)
print(tokenizer.decode(output[0]))
Validation
Success: 1.0247299671173096 <= 10.0
==================================================
Generating sample text:
According to all known laws of aviation, there is no way a bee should be able to fly
==================================================
Creation Process
This model was created using the llm-compressor
create-tiny-model
claude skill:
Inspected the original model configuration to identify key parameters
Created a tiny version by reducing
num_hidden_layers
from 16 to 4
Fine-tuned the model on a toy dataset (famous copypastas) to validate learning capability
Achieved target perplexity of ~1.02 on the validation text
Validated checkpoint structure matches the original model format
Confirmed successful loading and inference
Notes
This model was fine-tuned on a small corpus of internet copypastas to ensure it can learn effectively
The model maintains the same Llama 3.2 architecture (including RoPE parameters) as the base model, just with fewer layers
Due to the reduced layer count, this model has approximately 25% of the original model's parameters
This is intended for development and testing purposes, not production use
Runs of inference-optimization Llama-3.2-0.5B-Instruct on huggingface.co
16.7K
Total runs
0
24-hour runs
-290
3-day runs
-1.2K
7-day runs
-1.9K
30-day runs
More Information About Llama-3.2-0.5B-Instruct huggingface.co Model
Llama-3.2-0.5B-Instruct huggingface.co is an AI model on huggingface.co that provides Llama-3.2-0.5B-Instruct's model effect (), which can be used instantly with this inference-optimization Llama-3.2-0.5B-Instruct model. huggingface.co supports a free trial of the Llama-3.2-0.5B-Instruct model, and also provides paid use of the Llama-3.2-0.5B-Instruct. Support call Llama-3.2-0.5B-Instruct model through api, including Node.js, Python, http.
Llama-3.2-0.5B-Instruct huggingface.co is an online trial and call api platform, which integrates Llama-3.2-0.5B-Instruct's modeling effects, including api services, and provides a free online trial of Llama-3.2-0.5B-Instruct, you can try Llama-3.2-0.5B-Instruct online for free by clicking the link below.
inference-optimization Llama-3.2-0.5B-Instruct online free url in huggingface.co:
Llama-3.2-0.5B-Instruct is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.2-0.5B-Instruct on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.2-0.5B-Instruct install, users can directly use Llama-3.2-0.5B-Instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3.2-0.5B-Instruct install url in huggingface.co: