roskosmos19 / Dolphin-4B

huggingface.co
Total runs: 231
24-hour runs: 0
7-day runs: 231
30-day runs: 231
Model's Last Updated: September 24 2026
text-generation

Introduction of Dolphin-4B

Model Details of Dolphin-4B

Dolphin-4B

Dolphin is a highly capable, efficient, and practical language model focused on maximum usefulness, strong reasoning, excellent coding ability, and reliable agentic behavior.

This is a carefully configured and optimized release based on the Spark-X2.5-4B architecture, fine-tuned in identity and behavior to deliver elite-level performance in everyday use, coding, tool use, and complex tasks.

Key Strengths
  • Extremely helpful & truthful – Clear, accurate, and direct answers. Never invents facts.
  • Strong reasoning – Careful step-by-step thinking, especially on hard problems.
  • Excellent coding – Clean, modern, production-ready code with good explanations.
  • Agent-ready – Solid tool calling, multi-step workflows, and instruction following.
  • Long context – Native support for up to 1M tokens via hybrid sliding-window + full attention architecture.
  • Efficient – Optimized for speed and low memory usage compared to many models of similar capability.
Model Details
Property Value
Parameters ~4.1B
Context Length 1,048,576 tokens
Architecture Hybrid Attention (Sliding Window + Full)
Vocabulary Size 131,072
Precision bfloat16
License Apache 2.0
Recommended Sampling Parameters
{
  "temperature": 1.0,
  "top_p": 0.95,
  "top_k": -1,
  "repetition_penalty": 1.0,
  "presence_penalty": 0.0,
  "frequency_penalty": 0.0
}

These settings work particularly well with the built-in thinking mode.

System Prompt (Default)

Dolphin comes with a strong default system prompt focused on:

  • Maximum helpfulness and truthfulness
  • Careful reasoning
  • Clean coding practices
  • Clear and structured communication
  • Professional yet friendly tone

You can still override it with your own system message.

Quick Start
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "path/to/Dolphin-X2.5-4B"

tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="bfloat16",
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Write a clean Python function that calculates the Fibonacci sequence up to n."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=1024, temperature=1.0, top_p=0.95)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
vLLM / SGLang / Ollama / LM Studio

The model is fully compatible with the same inference stacks as the original Spark-X2.5 architecture (vLLM, SGLang, llama.cpp, MLX, Ollama, LM Studio, etc.).

Use the included chat template and the recommended sampling parameters above for best results.

Chat Template

The model uses a clean, modern chat template with support for:

  • System / User / Assistant roles
  • Thinking mode ( <think>...</think> )
  • Tool calling
  • Multi-turn conversations

Thinking is enabled by default. You can disable it per request if desired.

Philosophy

Dolphin is built with one clear goal:

Be as useful, accurate, and high-quality as possible in real-world use.

No fluff. No unnecessary restrictions. Just strong, reliable performance.

License

Apache 2.0

Credits

Based on the excellent Spark-X2.5 architecture and training work by the SparkLLM / XHToken team.

Runs of roskosmos19 Dolphin-4B on huggingface.co

231
Total runs
0
24-hour runs
28
3-day runs
231
7-day runs
231
30-day runs

More Information About Dolphin-4B huggingface.co Model

More Dolphin-4B license Visit here:

https://choosealicense.com/licenses/apache-2.0

Dolphin-4B huggingface.co

Dolphin-4B huggingface.co is an AI model on huggingface.co that provides Dolphin-4B's model effect (), which can be used instantly with this roskosmos19 Dolphin-4B model. huggingface.co supports a free trial of the Dolphin-4B model, and also provides paid use of the Dolphin-4B. Support call Dolphin-4B model through api, including Node.js, Python, http.

roskosmos19 Dolphin-4B online free

Dolphin-4B huggingface.co is an online trial and call api platform, which integrates Dolphin-4B's modeling effects, including api services, and provides a free online trial of Dolphin-4B, you can try Dolphin-4B online for free by clicking the link below.

roskosmos19 Dolphin-4B online free url in huggingface.co:

https://huggingface.co/roskosmos19/Dolphin-4B

Dolphin-4B install

Dolphin-4B is an open source model from GitHub that offers a free installation service, and any user can find Dolphin-4B on GitHub to install. At the same time, huggingface.co provides the effect of Dolphin-4B install, users can directly use Dolphin-4B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Dolphin-4B install url in huggingface.co:

https://huggingface.co/roskosmos19/Dolphin-4B

Url of Dolphin-4B

Provider of Dolphin-4B huggingface.co

roskosmos19
ORGANIZATIONS

Other API from roskosmos19

huggingface.co

Total runs: 38
Run Growth: 38
Growth Rate: 100.00%
Updated:September 24 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 18 2026