Model Description
i3-200m (codename: i3-redherring) is an enhanced hybrid language model that combines GRU-Mamba recurrence with multi-pattern attention mechanisms. The model features a unique 16-layer architecture (10 hybrid layers + 6 attention layers) designed for efficient sequence modeling with advanced memory optimization techniques.
Developed by:
B. Daniel (Me)
Model type:
Hybrid Recurrent-Attention Language Model
Language(s):
English
Architecture:
Enhanced i3 with GRU-Mamba Hybrid + Multi-Pattern Attention
Parameters:
~200M (512 dimensions, 16 layers, 16 attention heads)
Model Sources
Uses
Direct Use
The model can be used for:
Text generation and completion
Conversational AI applications
Creative writing assistance
Educational content generation (trained on TinyStories dataset)
Downstream Use
The model can be fine-tuned for:
Domain-specific text generation
Dialog systems
Story writing applications
Chat applications
Out-of-Scope Use
This model should not be used for:
High-stakes decision making without human oversight
Generation of harmful, biased, or misleading content
Applications requiring perfect factual accuracy
Medical, legal, or financial advice
Bias, Risks, and Limitations
Known Limitations
Context Length:
Limited to 256 tokens per sequence
Training Data:
Primarily trained on simplified English datasets (TinyStories, TinyChat, high-quality sentences)
Scale:
At 200M parameters, the model has limited capacity compared to larger language models
Vocabulary:
Uses a character-chunk tokenization strategy (2-3 character chunks) which may not be optimal for all languages
Bias Considerations
The model inherits biases from its training data:
TinyStories dataset may contain simplified narratives with limited cultural diversity
Training data is English-only, limiting multilingual capabilities
May reflect biases present in the source datasets
Safety Recommendations
Always review generated content before use in production
Implement content filtering for sensitive applications
Monitor for potential bias in downstream applications
Use human oversight for critical applications
Training Details
Training Data
The model was trained on a combination of three datasets:
agentlans/high-quality-english-sentences
- Curated high-quality English sentences
roneneldan/TinyStories
- Short stories written in simple language
starhopp3r/TinyChat
- Conversational text data
Total training tokens: 1,288,126,684
Training Procedure
Preprocessing
Tokenization:
Custom ChunkTokenizer using variable 2-3 character chunks
Vocabulary Size:
[Specify from training - appears to be dynamic based on dataset]
Special Tokens:
<UNK>
for unknown tokens
Sequence Length:
256 tokens
Training Hyperparameters
Training regime:
Mixed precision (FP16/BF16 where supported)
Batch Size:
4 (micro-batch) × 4 (gradient accumulation) = 16 effective batch size
Sequence Length:
256 tokens
Learning Rate:
3e-4 with cosine decay and warmup
Warmup Steps:
100 iterations
Total Iterations:
300 (configurable up to 5000)
Optimizer:
AdamW
Gradient Clipping:
1.0
Progressive Sparsity:
Up to 30% sparsification (warmup: 1000 steps)
Architecture Details
Layers:
Key Features:
GRU-Mamba Hybrid Recurrence:
Combines GRU gating with state-space models for efficient sequence modeling
Multi-Pattern Attention:
Sliding window (64 tokens)
Dilated causal (dilation factor: 2)
Chunked (32 token chunks)
Sparse MoE FFN:
4 experts with top-2 routing for computational efficiency
Progressive Sparsity:
Gradual pruning of attention heads during training
Memory Optimization:
Gradient checkpointing and mixed precision training
Training Infrastructure
Hardware:
CUDA-compatible GPU (details based on availability)
Training Time:
~4 hours for 300 iterations (varies by hardware)
Memory Usage:
~15% GPU memory allocation (with optimization)
Framework:
PyTorch with mixed precision training
Evaluation
Testing Data, Factors & Metrics
Testing Data
Evaluation performed on held-out samples from the training datasets.
Metrics
The model tracks multiple perplexity metrics:
Current Perplexity:
Instantaneous perplexity per iteration
Smoothed Perplexity:
Exponentially smoothed with α=0.1
Windowed Perplexity:
100-iteration rolling average
Token-Weighted Perplexity:
Weighted by sequence length
Harmonic Mean Perplexity:
Better metric for averaging
Results
Based on the training visualization provided:
Final Performance (at iteration ~300):
Training Loss: ~4.0
Current Perplexity: ~55-60
Smoothed Perplexity: ~60
Best Perplexity Achieved: ~48
Token-Weighted Perplexity: ~64
Training Dynamics:
Strong initial convergence (perplexity dropped from ~45,000 to ~1,000 in first 10 iterations)
Stable training after iteration 100
Learning rate follows cosine decay schedule
Training throughput: ~197,000 tokens/second
GPU utilization: ~20-30% (well-optimized)
GPU temperature: Stable at ~37°C
Model Efficiency:
Sparsity: 0% (not yet activated in warmup phase at 300 iterations)
Memory allocated: ~2.7GB (15% of available)
Gradient norm: Stable throughout training
No memory errors or instability
Environmental Impact
Hardware Type:
NVIDIA GPU (CUDA-compatible)
Hours used:
~4 hours for 300 iterations
Cloud Provider:
Kaggle
Compute Region:
Unknown
Carbon Emitted:
N/A
Efficiency Measures:
Gradient checkpointing reduces memory footprint
Mixed precision training (FP16/BF16)
Progressive sparsity for reduced computation
Efficient multi-pattern attention mechanisms
Technical Specifications
Model Architecture and Objective
Enhancedi3Model(
vocab_size=dynamic,
d_model=512,
n_heads=16,
n_layers=16,
max_seq_len=256,
d_state=64
)
Components:
Embedding layer (vocab_size × 512)
Positional embedding (256 × 512)
10 Enhanced Hybrid Blocks (GRU-Mamba + FFN/MoE)
6 Enhanced Attention Blocks (Multi-Pattern + FFN/MoE)
Layer normalization
Output projection (512 × vocab_size)
Compute Infrastructure
Hardware:
CUDA-compatible GPU
Software:
PyTorch, WandB for experiment tracking
Optimization:
Mixed precision, gradient checkpointing, gradient accumulation
Citation
BibTeX:
@misc{i3-200m-redherring,
title={Project i3-RedHerring: Efficient Sequence Modeling via Hybrid GRU-Mamba and Multi-Pattern Attention Architectures},
author={B. Daniel},
year={2025},
url={https://github.com/FlameF0X/i3-papers}
}
@article{mamba,
title={Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
author={Gu, Albert and Dao, Tri},
journal={arXiv preprint arXiv:2312.00752},
year={2023}
}