We introduce
Intern-S1-mini
, a lightweight open-source multimodal reasoning model based on the same techniques as
Intern-S1
.
Built upon a 8B dense language model (Qwen3) and a 0.3B Vision encoder (InternViT), Intern-S1-mini has been further pretrained on
5 trillion tokens
of multimodal data, including over
2.5 trillion scientific-domain tokens
. This enables the model to retain strong general capabilities while excelling in specialized scientific domains such as
interpreting chemical structures, understanding protein sequences, and planning compound synthesis routes
, making Intern-S1-mini to be a capable research assistant for real-world scientific applications.
Features
Strong performance across language and vision reasoning benchmarks, especially scientific tasks.
Continuously pretrained on a massive 5T token dataset, with over 50% specialized scientific data, embedding deep domain expertise.
Dynamic tokenizer enables native understanding of molecular formulas and protein sequences.
Performance
We evaluate the Intern-S1-mini on various benchmarks including general datasets and scientific datasets. We report the performance comparison with the recent VLMs and LLMs below.
Please ensure that the decord video decoding library is installed via
pip install decord
. To avoid OOM, please install flash_attention and use at least 2 GPUS.
from transformers import AutoProcessor, AutoModelForCausalLM
import torch
model_name = "internlm/Intern-S1-mini-FP8"
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype="auto", trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{
"type": "video",
"url": "https://huggingface.co/datasets/hf-internal-testing/fixtures_videos/resolve/main/tennis.mp4",
},
{"type": "text", "text": "What type of shot is the man performing?"},
],
}
]
inputs = processor.apply_chat_template(
messages,
return_tensors="pt",
add_generation_prompt=True,
video_load_backend="decord",
tokenize=True,
return_dict=True,
).to(model.device, dtype=torch.float16)
generate_ids = model.generate(**inputs, max_new_tokens=32768)
decoded_output = processor.decode(generate_ids[0, inputs["input_ids"].shape[1] :], skip_special_tokens=True)
print(decoded_output)
Serving
The minimum hardware requirements for deploying Intern-S1 series models are:
# install ollama
curl -fsSL https://ollama.com/install.sh | sh
# fetch model
ollama pull internlm/interns1-mini
# run model
ollama run internlm/interns1-mini
# then use openai client to call on http://localhost:11434/v1
Advanced Usage
Tool Calling
Many Large Language Models (LLMs) now feature
Tool Calling
, a powerful capability that allows them to extend their functionality by interacting with external tools and APIs. This enables models to perform tasks like fetching up-to-the-minute information, running code, or calling functions within other applications.
A key advantage for developers is that a growing number of open-source LLMs are designed to be compatible with the OpenAI API. This means you can leverage the same familiar syntax and structure from the OpenAI library to implement tool calling with these open-source models. As a result, the code demonstrated in this tutorial is versatile—it works not just with OpenAI models, but with any model that follows the same interface standard.
To illustrate how this works, let's dive into a practical code example that uses tool calling to get the latest weather forecast (based on lmdeploy api server).
from openai import OpenAI
import json
defget_current_temperature(location: str, unit: str = "celsius"):
"""Get current temperature at a location. Args: location: The location to get the temperature for, in the format "City, State, Country". unit: The unit to return the temperature in. Defaults to "celsius". (choices: ["celsius", "fahrenheit"]) Returns: the temperature, the location, and the unit in a dict """return {
"temperature": 26.1,
"location": location,
"unit": unit,
}
defget_temperature_date(location: str, date: str, unit: str = "celsius"):
"""Get temperature at a location and date. Args: location: The location to get the temperature for, in the format "City, State, Country". date: The date to get the temperature for, in the format "Year-Month-Day". unit: The unit to return the temperature in. Defaults to "celsius". (choices: ["celsius", "fahrenheit"]) Returns: the temperature, the location, the date and the unit in a dict """return {
"temperature": 25.9,
"location": location,
"date": date,
"unit": unit,
}
defget_function_by_name(name):
if name == "get_current_temperature":
return get_current_temperature
if name == "get_temperature_date":
return get_temperature_date
tools = [{
'type': 'function',
'function': {
'name': 'get_current_temperature',
'description': 'Get current temperature at a location.',
'parameters': {
'type': 'object',
'properties': {
'location': {
'type': 'string',
'description': 'The location to get the temperature for, in the format \'City, State, Country\'.'
},
'unit': {
'type': 'string',
'enum': [
'celsius',
'fahrenheit'
],
'description': 'The unit to return the temperature in. Defaults to \'celsius\'.'
}
},
'required': [
'location'
]
}
}
}, {
'type': 'function',
'function': {
'name': 'get_temperature_date',
'description': 'Get temperature at a location and date.',
'parameters': {
'type': 'object',
'properties': {
'location': {
'type': 'string',
'description': 'The location to get the temperature for, in the format \'City, State, Country\'.'
},
'date': {
'type': 'string',
'description': 'The date to get the temperature for, in the format \'Year-Month-Day\'.'
},
'unit': {
'type': 'string',
'enum': [
'celsius',
'fahrenheit'
],
'description': 'The unit to return the temperature in. Defaults to \'celsius\'.'
}
},
'required': [
'location',
'date'
]
}
}
}]
messages = [
{'role': 'user', 'content': 'Today is 2024-11-14, What\'s the temperature in San Francisco now? How about tomorrow?'}
]
openai_api_key = "EMPTY"
openai_api_base = "http://0.0.0.0:23333/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
model_name = client.models.list().data[0].id
response = client.chat.completions.create(
model=model_name,
messages=messages,
max_tokens=32768,
temperature=0.8,
top_p=0.8,
stream=False,
extra_body=dict(spaces_between_special_tokens=False, enable_thinking=False),
tools=tools)
print(response.choices[0].message)
messages.append(response.choices[0].message)
for tool_call in response.choices[0].message.tool_calls:
tool_call_args = json.loads(tool_call.function.arguments)
tool_call_result = get_function_by_name(tool_call.function.name)(**tool_call_args)
tool_call_result = json.dumps(tool_call_result, ensure_ascii=False)
messages.append({
'role': 'tool',
'name': tool_call.function.name,
'content': tool_call_result,
'tool_call_id': tool_call.id
})
response = client.chat.completions.create(
model=model_name,
messages=messages,
temperature=0.8,
top_p=0.8,
stream=False,
extra_body=dict(spaces_between_special_tokens=False, enable_thinking=False),
tools=tools)
print(response.choices[0].message.content)
Switching Between Thinking and Non-Thinking Modes
Intern-S1-mini enables thinking mode by default, enhancing the model's reasoning capabilities to generate higher-quality responses. This feature can be disabled by setting
enable_thinking=False
in
tokenizer.apply_chat_template
With LMDeploy serving Intern-S1-mini models, you can dynamically control the thinking mode by adjusting the
enable_thinking
parameter in your requests.
Intern-S1-mini-FP8 huggingface.co is an AI model on huggingface.co that provides Intern-S1-mini-FP8's model effect (), which can be used instantly with this internlm Intern-S1-mini-FP8 model. huggingface.co supports a free trial of the Intern-S1-mini-FP8 model, and also provides paid use of the Intern-S1-mini-FP8. Support call Intern-S1-mini-FP8 model through api, including Node.js, Python, http.
Intern-S1-mini-FP8 huggingface.co is an online trial and call api platform, which integrates Intern-S1-mini-FP8's modeling effects, including api services, and provides a free online trial of Intern-S1-mini-FP8, you can try Intern-S1-mini-FP8 online for free by clicking the link below.
internlm Intern-S1-mini-FP8 online free url in huggingface.co:
Intern-S1-mini-FP8 is an open source model from GitHub that offers a free installation service, and any user can find Intern-S1-mini-FP8 on GitHub to install. At the same time, huggingface.co provides the effect of Intern-S1-mini-FP8 install, users can directly use Intern-S1-mini-FP8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.