FlameF0X / i3-200m

huggingface.co
Total runs: 44
24-hour runs: 0
7-day runs: 11
30-day runs: 11
Model's Last Updated: November 25 2025
text-generation

Introduction of i3-200m

Model Details of i3-200m

Gemini_Generated_Image_ore10zore10zore1

Model Description

i3-200m (codename: i3-redherring) is an enhanced hybrid language model that combines GRU-Mamba recurrence with multi-pattern attention mechanisms. The model features a unique 16-layer architecture (10 hybrid layers + 6 attention layers) designed for efficient sequence modeling with advanced memory optimization techniques.

  • Developed by: B. Daniel (Me)
  • Model type: Hybrid Recurrent-Attention Language Model
  • Language(s): English
  • Architecture: Enhanced i3 with GRU-Mamba Hybrid + Multi-Pattern Attention
  • Parameters: ~200M (512 dimensions, 16 layers, 16 attention heads)
Model Sources
Uses
Direct Use

The model can be used for:

  • Text generation and completion
  • Conversational AI applications
  • Creative writing assistance
  • Educational content generation (trained on TinyStories dataset)
Downstream Use

The model can be fine-tuned for:

  • Domain-specific text generation
  • Dialog systems
  • Story writing applications
  • Chat applications
Out-of-Scope Use

This model should not be used for:

  • High-stakes decision making without human oversight
  • Generation of harmful, biased, or misleading content
  • Applications requiring perfect factual accuracy
  • Medical, legal, or financial advice
Bias, Risks, and Limitations
Known Limitations
  • Context Length: Limited to 256 tokens per sequence
  • Training Data: Primarily trained on simplified English datasets (TinyStories, TinyChat, high-quality sentences)
  • Scale: At 200M parameters, the model has limited capacity compared to larger language models
  • Vocabulary: Uses a character-chunk tokenization strategy (2-3 character chunks) which may not be optimal for all languages
Bias Considerations

The model inherits biases from its training data:

  • TinyStories dataset may contain simplified narratives with limited cultural diversity
  • Training data is English-only, limiting multilingual capabilities
  • May reflect biases present in the source datasets
Safety Recommendations
  • Always review generated content before use in production
  • Implement content filtering for sensitive applications
  • Monitor for potential bias in downstream applications
  • Use human oversight for critical applications
Training Details
Training Data

The model was trained on a combination of three datasets:

  1. agentlans/high-quality-english-sentences - Curated high-quality English sentences
  2. roneneldan/TinyStories - Short stories written in simple language
  3. starhopp3r/TinyChat - Conversational text data

Total training tokens: 1,288,126,684

Training Procedure
Preprocessing
  • Tokenization: Custom ChunkTokenizer using variable 2-3 character chunks
  • Vocabulary Size: [Specify from training - appears to be dynamic based on dataset]
  • Special Tokens: <UNK> for unknown tokens
  • Sequence Length: 256 tokens
Training Hyperparameters
  • Training regime: Mixed precision (FP16/BF16 where supported)
  • Batch Size: 4 (micro-batch) × 4 (gradient accumulation) = 16 effective batch size
  • Sequence Length: 256 tokens
  • Learning Rate: 3e-4 with cosine decay and warmup
  • Warmup Steps: 100 iterations
  • Total Iterations: 300 (configurable up to 5000)
  • Optimizer: AdamW
  • Gradient Clipping: 1.0
  • Progressive Sparsity: Up to 30% sparsification (warmup: 1000 steps)
Architecture Details

Layers:

  • 10 Enhanced Hybrid Blocks (GRU-Mamba)

    • GRU-style gating (reset, update, candidate gates)
    • Mamba state-space dynamics (64-dimensional state space)
    • FFN with GELU activation (4x expansion)
  • 6 Enhanced Attention Blocks

    • Multi-pattern attention (sliding window, dilated, chunked)
    • Dynamic pattern routing with learned weights
    • Sparse Mixture of Experts FFN (4 experts, top-2 routing)

Key Features:

  • GRU-Mamba Hybrid Recurrence: Combines GRU gating with state-space models for efficient sequence modeling
  • Multi-Pattern Attention:
    • Sliding window (64 tokens)
    • Dilated causal (dilation factor: 2)
    • Chunked (32 token chunks)
  • Sparse MoE FFN: 4 experts with top-2 routing for computational efficiency
  • Progressive Sparsity: Gradual pruning of attention heads during training
  • Memory Optimization: Gradient checkpointing and mixed precision training
Training Infrastructure
  • Hardware: CUDA-compatible GPU (details based on availability)
  • Training Time: ~4 hours for 300 iterations (varies by hardware)
  • Memory Usage: ~15% GPU memory allocation (with optimization)
  • Framework: PyTorch with mixed precision training
Evaluation
Testing Data, Factors & Metrics
Testing Data

Evaluation performed on held-out samples from the training datasets.

Metrics

The model tracks multiple perplexity metrics:

  • Current Perplexity: Instantaneous perplexity per iteration
  • Smoothed Perplexity: Exponentially smoothed with α=0.1
  • Windowed Perplexity: 100-iteration rolling average
  • Token-Weighted Perplexity: Weighted by sequence length
  • Harmonic Mean Perplexity: Better metric for averaging
Results

Based on the training visualization provided:

Final Performance (at iteration ~300):

  • Training Loss: ~4.0
  • Current Perplexity: ~55-60
  • Smoothed Perplexity: ~60
  • Best Perplexity Achieved: ~48
  • Token-Weighted Perplexity: ~64

Training Dynamics:

  • Strong initial convergence (perplexity dropped from ~45,000 to ~1,000 in first 10 iterations)
  • Stable training after iteration 100
  • Learning rate follows cosine decay schedule
  • Training throughput: ~197,000 tokens/second
  • GPU utilization: ~20-30% (well-optimized)
  • GPU temperature: Stable at ~37°C

Model Efficiency:

  • Sparsity: 0% (not yet activated in warmup phase at 300 iterations)
  • Memory allocated: ~2.7GB (15% of available)
  • Gradient norm: Stable throughout training
  • No memory errors or instability
Environmental Impact
  • Hardware Type: NVIDIA GPU (CUDA-compatible)
  • Hours used: ~4 hours for 300 iterations
  • Cloud Provider: Kaggle
  • Compute Region: Unknown
  • Carbon Emitted: N/A

Efficiency Measures:

  • Gradient checkpointing reduces memory footprint
  • Mixed precision training (FP16/BF16)
  • Progressive sparsity for reduced computation
  • Efficient multi-pattern attention mechanisms
Technical Specifications
Model Architecture and Objective
Enhancedi3Model(
  vocab_size=dynamic,
  d_model=512,
  n_heads=16,
  n_layers=16,
  max_seq_len=256,
  d_state=64
)

Components:

  • Embedding layer (vocab_size × 512)
  • Positional embedding (256 × 512)
  • 10 Enhanced Hybrid Blocks (GRU-Mamba + FFN/MoE)
  • 6 Enhanced Attention Blocks (Multi-Pattern + FFN/MoE)
  • Layer normalization
  • Output projection (512 × vocab_size)
Compute Infrastructure
  • Hardware: CUDA-compatible GPU
  • Software: PyTorch, WandB for experiment tracking
  • Optimization: Mixed precision, gradient checkpointing, gradient accumulation
Citation

BibTeX:

@misc{i3-200m-redherring,
  title={Project i3-RedHerring: Efficient Sequence Modeling via Hybrid GRU-Mamba and Multi-Pattern Attention Architectures},
  author={B. Daniel},
  year={2025},
  url={https://github.com/FlameF0X/i3-papers}
}
@article{mamba,
  title={Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
  author={Gu, Albert and Dao, Tri},
  journal={arXiv preprint arXiv:2312.00752},
  year={2023}
}

Runs of FlameF0X i3-200m on huggingface.co

44
Total runs
0
24-hour runs
11
3-day runs
11
7-day runs
11
30-day runs

More Information About i3-200m huggingface.co Model

More i3-200m license Visit here:

https://choosealicense.com/licenses/apache-2.0

i3-200m huggingface.co

i3-200m huggingface.co is an AI model on huggingface.co that provides i3-200m's model effect (), which can be used instantly with this FlameF0X i3-200m model. huggingface.co supports a free trial of the i3-200m model, and also provides paid use of the i3-200m. Support call i3-200m model through api, including Node.js, Python, http.

FlameF0X i3-200m online free

i3-200m huggingface.co is an online trial and call api platform, which integrates i3-200m's modeling effects, including api services, and provides a free online trial of i3-200m, you can try i3-200m online for free by clicking the link below.

FlameF0X i3-200m online free url in huggingface.co:

https://huggingface.co/FlameF0X/i3-200m

i3-200m install

i3-200m is an open source model from GitHub that offers a free installation service, and any user can find i3-200m on GitHub to install. At the same time, huggingface.co provides the effect of i3-200m install, users can directly use i3-200m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

i3-200m install url in huggingface.co:

https://huggingface.co/FlameF0X/i3-200m

Url of i3-200m

Provider of i3-200m huggingface.co

FlameF0X
ORGANIZATIONS

Other API from FlameF0X

huggingface.co

Total runs: 164
Run Growth: 138
Growth Rate: 84.15%
Updated:July 17 2026
huggingface.co

Total runs: 144
Run Growth: 49
Growth Rate: 34.03%
Updated:May 26 2026
huggingface.co

Total runs: 114
Run Growth: -76
Growth Rate: -66.67%
Updated:December 01 2025
huggingface.co

Total runs: 112
Run Growth: 16
Growth Rate: 14.29%
Updated:October 31 2025
huggingface.co

Total runs: 109
Run Growth: 0
Growth Rate: 0.00%
Updated:July 25 2026
huggingface.co

Total runs: 70
Run Growth: 70
Growth Rate: 100.00%
Updated:December 03 2025
huggingface.co

Total runs: 62
Run Growth: 12
Growth Rate: 19.35%
Updated:November 29 2025
huggingface.co

Total runs: 50
Run Growth: 47
Growth Rate: 94.00%
Updated:June 29 2025
huggingface.co

Total runs: 39
Run Growth: -135
Growth Rate: -346.15%
Updated:October 23 2025
huggingface.co

Total runs: 36
Run Growth: 20
Growth Rate: 55.56%
Updated:February 23 2026
huggingface.co

Total runs: 30
Run Growth: -254
Growth Rate: -846.67%
Updated:May 20 2026
huggingface.co

Total runs: 21
Run Growth: -10
Growth Rate: -47.62%
Updated:October 17 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:December 19 2025
huggingface.co

Total runs: 14
Run Growth: -272
Growth Rate: -1942.86%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 06 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 29 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 31 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 17 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 07 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 31 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 19 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 25 2026