FlameF0X / i3-200m-v2

huggingface.co
Total runs: 70
24-hour runs: 0
7-day runs: 14
30-day runs: 70
Model's Last Updated: December 03 2025
text-generation

Introduction of i3-200m-v2

Model Details of i3-200m-v2

i3-200M - Hybrid Architecture Language Model

Gemini_Generated_Image_ore10zore10zore1

Model Description

The i3-200M Model (aka Redherring) is a novel hybrid architecture combining convolutional/recurrent layers with full attention layers for efficient language modeling. This architecture uniquely blends RWKV-style time-mixing with Mamba state-space dynamics in the early layers, followed by standard multi-head attention in deeper layers.

To use the model try it here .

Model Statistics
  • Total Parameters : ~169.85M
  • Architecture : 10 Hybrid (RWKV-Mamba) + 6 Full Attention Layers = 16 Total Layers
  • Vocabulary Size : 32,000 tokens
  • Hidden Dimension (d_model) : 512
  • Attention Heads : 16
  • State Dimension (d_state) : 32
  • Max Sequence Length : 256
  • Tokenization : BPE
Architecture Breakdown
Layers 1-10:  RWKV-Mamba Hybrid Blocks (Recurrent/Conv)
              ├─ RWKVMambaHybrid (Time-mixing + State-space)
              └─ Feed-Forward Network (4x expansion)

Layers 11-16: Full Attention Blocks
              ├─ Multi-Head Attention (16 heads)
              └─ Feed-Forward Network (4x expansion)
Comparison with i3-80M
Feature i3-22M i3-80M i3-200m (This Model)
Parameters 22.6M 82.77M 169.85M
Architecture 24 Hybrid Layers 10 Hybrid + 6 Attention Layers 10 Hybrid + 6 Attention Layers
Hidden Dimension 512 512 512
Vocabulary Size 4,466 35,560 32,000
Training Dataset TinyChat only TinyStories + TinyChat + HQ Sentences TinyStories + TinyChat + HQ Sentences + Wikitext
Total Tokens ~1M conversations ~3M+ tokens N/A
Final Loss ~2.0 ~2.0 1.6
Final Perplexity 7.29-9.70 7.29-10.0 5.2
Training Time ~17 hours ~2-4 hours ~1-2 hours
Attention Layers None (Pure Hybrid) 6 Full Attention Layers 6 Full Attention Layers
Key Improvements Over i3-80M

To Be added

Key Features
  1. Hybrid Architecture : Combines the efficiency of recurrent/convolutional processing with the power of attention

    • Early layers use RWKV-Mamba hybrid for efficient sequence processing
    • Later layers use full multi-head attention for complex pattern recognition
  2. Memory-Optimized Training :

    • Streaming vocabulary building (no full text storage)
    • Vocabulary caching (build once, reuse)
    • Efficient chunk frequency counting
    • Automatic memory cleanup
  3. Multi-Dataset Pre-training : Trained on diverse text sources for robust language understanding

    • TinyStories: Narrative and storytelling
    • TinyChat: Conversational dynamics
    • High-Quality English Sentences: Linguistic diversity
    • wikitext
  4. BPE Tokenization :

    • Total tokens processed: N/A
    • Handles unknown tokens gracefully with , , , token
Training Details
Training Configuration
  • Datasets :
    • agentlans/high-quality-english-sentences
    • roneneldan/TinyStories
    • starhopp3r/TinyChat
    • wikitext
  • Training Steps : 250 iterations
  • Batch Size : 4 (with gradient accumulation support)
  • Learning Rate : 4e-4 (with warmup and cosine decay)
  • Optimizer : AdamW with gradient clipping (max norm: 1.0)
  • Hardware : NVIDIA P100 (16GB VRAM)
  • Training Time : ~1-2 hours
  • Framework : PyTorch
Training Dynamics
  • GPU Utilization : Stable at ~N/A% during training
  • GPU Memory : ~20% allocated (~4GB / 12GB) (Math is not mathing??)
  • Power Usage : ~50W average
  • Throughput : ~300 tokens/sec
Performance Metrics
Metric Initial Final
Training Loss ~10.0 1.6
Perplexity ~4000+ 5.2
Technical Innovations
  1. RWKV-Mamba Hybrid Recurrence : Combines RWKV's time-mixing with Mamba's state-space dynamics

    • Linear complexity for long sequences
    • Efficient recurrent processing
    • State-space modeling for temporal dependencies
  2. Hierarchical Processing :

    • Lower layers focus on local patterns (conv/recurrent)
    • Upper layers capture global dependencies (attention)
  3. Memory Efficiency :

    • Streaming tokenization during vocab building
    • No full dataset storage in RAM
    • Automatic cleanup of intermediate data
Model Files
  • pytorch_model.bin : Model weights
  • config.json : Model configuration
  • tokenizer.json : Tokenizer vocabulary
Limitations
  • Trained on English text only
  • Limited to 256 token context window
  • May require fine-tuning for specific downstream tasks
  • Conversational style influenced by TinyChat dataset
Model Series
  • i3-22M - Original model with pure hybrid architecture
  • i3-80M - Scaled version with attention layers and multi-dataset training
  • i3-200M (This model)
Citation
@article{mamba,
  title={Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
  author={Gu, Albert and Dao, Tri},
  journal={arXiv preprint arXiv:2312.00752},
  year={2023}
}
@article{RWKV,
  title={RWKV: Reinventing RNNs for the Transformer Era},
  author={Peng, Bo and others},
  journal={arXiv preprint arXiv:2305.13048},
  year={2023}
}

Runs of FlameF0X i3-200m-v2 on huggingface.co

70
Total runs
0
24-hour runs
1
3-day runs
14
7-day runs
70
30-day runs

More Information About i3-200m-v2 huggingface.co Model

More i3-200m-v2 license Visit here:

https://choosealicense.com/licenses/apache-2.0

i3-200m-v2 huggingface.co

i3-200m-v2 huggingface.co is an AI model on huggingface.co that provides i3-200m-v2's model effect (), which can be used instantly with this FlameF0X i3-200m-v2 model. huggingface.co supports a free trial of the i3-200m-v2 model, and also provides paid use of the i3-200m-v2. Support call i3-200m-v2 model through api, including Node.js, Python, http.

FlameF0X i3-200m-v2 online free

i3-200m-v2 huggingface.co is an online trial and call api platform, which integrates i3-200m-v2's modeling effects, including api services, and provides a free online trial of i3-200m-v2, you can try i3-200m-v2 online for free by clicking the link below.

FlameF0X i3-200m-v2 online free url in huggingface.co:

https://huggingface.co/FlameF0X/i3-200m-v2

i3-200m-v2 install

i3-200m-v2 is an open source model from GitHub that offers a free installation service, and any user can find i3-200m-v2 on GitHub to install. At the same time, huggingface.co provides the effect of i3-200m-v2 install, users can directly use i3-200m-v2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

i3-200m-v2 install url in huggingface.co:

https://huggingface.co/FlameF0X/i3-200m-v2

Url of i3-200m-v2

i3-200m-v2 huggingface.co Url

Provider of i3-200m-v2 huggingface.co

FlameF0X
ORGANIZATIONS

Other API from FlameF0X

huggingface.co

Total runs: 164
Run Growth: 138
Growth Rate: 84.15%
Updated:July 17 2026
huggingface.co

Total runs: 144
Run Growth: 49
Growth Rate: 34.03%
Updated:May 26 2026
huggingface.co

Total runs: 114
Run Growth: -76
Growth Rate: -66.67%
Updated:December 01 2025
huggingface.co

Total runs: 112
Run Growth: 16
Growth Rate: 14.29%
Updated:October 31 2025
huggingface.co

Total runs: 109
Run Growth: 0
Growth Rate: 0.00%
Updated:July 25 2026
huggingface.co

Total runs: 62
Run Growth: 12
Growth Rate: 19.35%
Updated:November 29 2025
huggingface.co

Total runs: 50
Run Growth: 47
Growth Rate: 94.00%
Updated:June 29 2025
huggingface.co

Total runs: 44
Run Growth: 11
Growth Rate: 25.00%
Updated:November 25 2025
huggingface.co

Total runs: 39
Run Growth: -135
Growth Rate: -346.15%
Updated:October 23 2025
huggingface.co

Total runs: 36
Run Growth: 20
Growth Rate: 55.56%
Updated:February 23 2026
huggingface.co

Total runs: 30
Run Growth: -254
Growth Rate: -846.67%
Updated:May 20 2026
huggingface.co

Total runs: 21
Run Growth: -10
Growth Rate: -47.62%
Updated:October 17 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:December 19 2025
huggingface.co

Total runs: 14
Run Growth: -272
Growth Rate: -1942.86%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 06 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 29 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 31 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 17 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 07 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 31 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 19 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 25 2026