Reasoning Llama model series fine-tuned on microsoft/orca-math-word-problems-200k using GRPO(Group Relative Policy Optimization) reinforcement learning technique.
Base model: meta-llama/Llama-3.1-8B-Instruct
Parameters
learning_rate = 5e-6,
adam_beta1 = 0.9,
adam_beta2 = 0.99,
weight_decay = 0.1,
warmup_ratio = 0.1,
lr_scheduler_type = "cosine",
optim = "paged_adamw_8bit",
Suggested system prompt for reasoning
Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>
Do not forget <reasoning></reasoning><answer></answer> tags.
Support:
If you find this work useful, you can support me!
Runs of suayptalha ThinkerLlama-8B-v1 on huggingface.co
2
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About ThinkerLlama-8B-v1 huggingface.co Model
ThinkerLlama-8B-v1 huggingface.co is an AI model on huggingface.co that provides ThinkerLlama-8B-v1's model effect (), which can be used instantly with this suayptalha ThinkerLlama-8B-v1 model. huggingface.co supports a free trial of the ThinkerLlama-8B-v1 model, and also provides paid use of the ThinkerLlama-8B-v1. Support call ThinkerLlama-8B-v1 model through api, including Node.js, Python, http.
ThinkerLlama-8B-v1 huggingface.co is an online trial and call api platform, which integrates ThinkerLlama-8B-v1's modeling effects, including api services, and provides a free online trial of ThinkerLlama-8B-v1, you can try ThinkerLlama-8B-v1 online for free by clicking the link below.
suayptalha ThinkerLlama-8B-v1 online free url in huggingface.co:
ThinkerLlama-8B-v1 is an open source model from GitHub that offers a free installation service, and any user can find ThinkerLlama-8B-v1 on GitHub to install. At the same time, huggingface.co provides the effect of ThinkerLlama-8B-v1 install, users can directly use ThinkerLlama-8B-v1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.