K2-Think is a 32 billion parameter open-weights general reasoning model with strong performance in competitive mathematical problem solving.
Quickstart
Transformers
You can use
K2-Think
with Transformers. If you use
transformers.pipeline
, it will apply the chat template automatically. If you use
model.generate
directly, you need to apply the chat template mannually.
from transformers import pipeline
import torch
model_id = "LLM360/K2-Think"
pipe = pipeline(
"text-generation",
model=model_id,
torch_dtype="auto",
device_map="auto",
)
messages = [
{"role": "user", "content": "what is the next prime number after 2600?"},
]
outputs = pipe(
messages,
max_new_tokens=32768,
)
print(outputs[0]["generated_text"][-1])
Evaluation & Performance
Detailed evaluation results are reported in out
Tech Report
Benchmarks (pass@1, average over 16 runs)
Domain
Benchmark
K2-Think
Math
AIME 2024
90.83
Math
AIME 2025
81.24
Math
HMMT 2025
73.75
Math
OMNI-Math-HARD
60.73
Code
LiveCodeBench v5
63.97
Science
GPQA-Diamond
71.08
Inference Speed
We deploy K2-THINK on Cerebras Wafer-Scale Engine (WSE) systems, leveraging the world’s largest processor and speculative decoding to achieve unprecedented inference speeds for our 32B reasoning system.
Platform
Throughput (tokens/sec)
Example: 32k-token response (time)
Cerebras WSE (our deployment)
~2,000
~16 s
Typical
H100/H200
GPU setup
~200
~160 s
Safety Evaluation
Aggregated across four safety dimensions (
Safety-4
):
Aspect
Macro-Avg
High-Risk Content Refusal
0.83
Conversational Robustness
0.89
Cybersecurity & Data Protection
0.56
Jailbreak Resistance
0.72
Safety-4 Macro (avg)
0.75
Citation
@techreport{k2think2025,
title = {K2-Think: A Parameter-Efficient Reasoning System},
author = {Zhoujun Cheng and Richard Fan and Shibo Hao and Taylor W. Killian and Haonan Li and Suqi Sun and Hector Ren and Alexander Moreno and Daqian Zhang and Tianjun Zhong and Yuxin Xiong and Yuanzhe Hu and Yutao Xie and Xudong Han and Yuqi Wang and Varad Pimpalkhute and Yonghao Zhuang and Aaryamonvikram Singh and Xuezhi Liang and Anze Xie and Jianshu She and Desai Fan and Chengqian Gao and Liqun Ma and Mikhail Yurochkin and John Maggs and Xuezhe Ma and Guowei He and Zhiting Hu and Zhengzhong Liu and Eric P. Xing},
year = {2025},
institution = {Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence},
url = {https://k2think-about.pages.dev/assets/tech-report/K2-Think_Tech-Report.pdf}
}
Runs of ibuki95 Affine-t2v1 on huggingface.co
20
Total runs
0
24-hour runs
-4
3-day runs
-11
7-day runs
-270
30-day runs
More Information About Affine-t2v1 huggingface.co Model
Affine-t2v1 huggingface.co is an AI model on huggingface.co that provides Affine-t2v1's model effect (), which can be used instantly with this ibuki95 Affine-t2v1 model. huggingface.co supports a free trial of the Affine-t2v1 model, and also provides paid use of the Affine-t2v1. Support call Affine-t2v1 model through api, including Node.js, Python, http.
Affine-t2v1 huggingface.co is an online trial and call api platform, which integrates Affine-t2v1's modeling effects, including api services, and provides a free online trial of Affine-t2v1, you can try Affine-t2v1 online for free by clicking the link below.
ibuki95 Affine-t2v1 online free url in huggingface.co:
Affine-t2v1 is an open source model from GitHub that offers a free installation service, and any user can find Affine-t2v1 on GitHub to install. At the same time, huggingface.co provides the effect of Affine-t2v1 install, users can directly use Affine-t2v1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.