Here’s the complete, refined code for patching the weights:
# Import required librariesfrom transformers import AutoProcessor, AutoTokenizer, AutoModelForImageTextToText, AutoModelForCausalLM
# Load the 11B Vision-Instruct model
processor = AutoProcessor.from_pretrained("meta-llama/Llama-3.2-11B-Vision-Instruct")
model = AutoModelForImageTextToText.from_pretrained("meta-llama/Llama-3.2-11B-Vision-Instruct")
# Load the 8B text-only model
s_tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
s_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
# Prepare input text for testing
input_text = "Write me a poem about Machine Learning."
input_ids = s_tokenizer(input_text, return_tensors="pt")
# Test the original 8B model
outputs = s_model.generate(**input_ids, do_sample=False, max_new_tokens=10)
print("8B Model Output:", s_tokenizer.decode(outputs[0]))
# Patch weights from the 11B model into the 8B model
model_weight = model.state_dict()
s_model_dict = s_model.state_dict()
skip_layer = 0# Track skipped layersfor key in s_model_dict.keys():
if"layers."in key:
layer_idx = int(key.split("layers.")[1].split(".")[0]) # Extract layer indextry:
s_model_dict[key] = model_weight[
"language_model." + key.replace(f"layers.{layer_idx}.", f"layers.{layer_idx + skip_layer}.")
]
except KeyError:
skip_layer += 1
s_model_dict[key] = model_weight[
"language_model." + key.replace(f"layers.{layer_idx}.", f"layers.{layer_idx + skip_layer}.")
]
else:
s_model_dict[key] = model_weight["language_model." + key]
# Test the patched 8B model
outputs = s_model.generate(**input_ids, do_sample=False, max_new_tokens=10)
print("Patched 8B Model Output:", s_tokenizer.decode(outputs[0]))
# Test the original 11B model
outputs = model.generate(**input_ids, do_sample=False, max_new_tokens=10)
print("11B Model Output:", s_tokenizer.decode(outputs[0]))
Example Outputs
Prompt:
"Write me a poem about Machine Learning."
Outputs:
8B Model Output (Before Patching):
<|begin_of_text|>Write me a poem about Machine Learning.
Artificial minds, born from code,
Learning
Patched 8B Model Output:
<|begin_of_text|>Write me a poem about Machine Learning.
In silicon halls, where data reigns
11B Model Output:
<|begin_of_text|>Write me a poem about Machine Learning.
In silicon halls, where data reigns
Model Details
Model Description
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
Developed by:
[More Information Needed]
Funded by [optional]:
[More Information Needed]
Shared by [optional]:
[More Information Needed]
Model type:
[More Information Needed]
Language(s) (NLP):
[More Information Needed]
License:
[More Information Needed]
Finetuned from model [optional]:
[More Information Needed]
Model Sources [optional]
Repository:
[More Information Needed]
Paper [optional]:
[More Information Needed]
Demo [optional]:
[More Information Needed]
Uses
Direct Use
[More Information Needed]
Downstream Use [optional]
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Bias, Risks, and Limitations
[More Information Needed]
Recommendations
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Runs of voidful Llama-3.2-8B-Instruct on huggingface.co
260
Total runs
0
24-hour runs
8
3-day runs
-45
7-day runs
-31
30-day runs
More Information About Llama-3.2-8B-Instruct huggingface.co Model
Llama-3.2-8B-Instruct huggingface.co
Llama-3.2-8B-Instruct huggingface.co is an AI model on huggingface.co that provides Llama-3.2-8B-Instruct's model effect (), which can be used instantly with this voidful Llama-3.2-8B-Instruct model. huggingface.co supports a free trial of the Llama-3.2-8B-Instruct model, and also provides paid use of the Llama-3.2-8B-Instruct. Support call Llama-3.2-8B-Instruct model through api, including Node.js, Python, http.
Llama-3.2-8B-Instruct huggingface.co is an online trial and call api platform, which integrates Llama-3.2-8B-Instruct's modeling effects, including api services, and provides a free online trial of Llama-3.2-8B-Instruct, you can try Llama-3.2-8B-Instruct online for free by clicking the link below.
voidful Llama-3.2-8B-Instruct online free url in huggingface.co:
Llama-3.2-8B-Instruct is an open source model from GitHub that offers a free installation service, and any user can find Llama-3.2-8B-Instruct on GitHub to install. At the same time, huggingface.co provides the effect of Llama-3.2-8B-Instruct install, users can directly use Llama-3.2-8B-Instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Llama-3.2-8B-Instruct install url in huggingface.co: