Lyon28 / caca-65M-untrained

huggingface.co
Total runs: 16
24-hour runs: 0
7-day runs: -4
30-day runs: 5
Model's Last Updated: July 19 2026
text-generation

Introduction of caca-65M-untrained

Model Details of caca-65M-untrained

caca-65M

๐Ÿค– caca-65M

Arsitektur Transformer Modern dengan Fitur Canggih

License Python 3.8+ PyTorch Transformers Model Type Parameters Status

65,012,096 parameters โ€ข 65.01M โ€ข 77 layers โ€ข 2,048 tokens

๐Ÿ“š Documentation โ€ข ๐Ÿ’ป Usage โ€ข โš™๏ธ Configuration โ€ข ๐Ÿ”ฌ Architecture


โš ๏ธ PENTING: Model Belum Dilatih (Untrained)
โš ๏ธ PERHATIAN : Ini adalah model yang belum melalui proses training . Bobot model masih dalam kondisi random initialization . Output yang dihasilkan akan tidak bermakna dan acak .

Status Model:

  • ๐Ÿ”ด Belum dilatih - Bobot masih random (Kaiming/Xavier init)
  • ๐ŸŸก Untuk riset & eksperimen - Arsitektur sudah siap, tinggal train
  • ๐ŸŸข Production-ready architecture - Teruji dan optimal

Widget di atas hanya menunjukkan format input yang diharapkan . Setelah model dilatih dengan dataset yang tepat, format yang sama akan menghasilkan output berkualitas tinggi.

๐ŸŽฏ Apa yang Bisa Dilakukan?
โœ… Bisa โŒ Belum Bisa
Load model architecture Generate teks bermakna
Test forward pass Menjawab pertanyaan
Measure memory & speed Reasoning & understanding
Start training Production deployment
Fine-tuning experiments Real-world applications

๐Ÿ“‹ Deskripsi

Caca adalah arsitektur Large Language Model (LLM) generasi terbaru yang menggabungkan berbagai teknik state-of-the-art dalam deep learning. Model ini dirancang dengan fokus pada efisiensi komputasi , skalabilitas , dan performa tinggi .

๐Ÿ“– Tentang Project Caca

Caca adalah eksperimen open-source Indonesian LLM yang dibuat dari nol secara individual dan bertahap. Bukan kompetitor siapa-siapa, cuma pengen eksplorasi apa yang bisa dilakukan dengan budget terbatas, passion unlimited, dan mindset collaborative.

Kalau berguna buat orang lain, alhamdulillah. Kalau enggak, ya tetap fun kok. Ini proyek eksplorasi, jadi kalau gagal ya bagian dari proses belajar. Kalau berhasil, itu bonus.

โ€” Lyon , Creator

๐ŸŒŸ Mengapa Caca?
  1. ๐Ÿ‡ฎ๐Ÿ‡ฉ Fokus pada Bahasa Indonesia - Dirancang dengan mempertimbangkan karakteristik bahasa Indonesia
  2. โšก Efisiensi Tinggi - GQA & Flash Attention untuk inferensi 3-5x lebih cepat
  3. ๐Ÿ’พ Memory Efficient - Hemat 50% memory untuk KV cache
  4. ๐Ÿ”ง Modular & Extensible - Mudah dikustomisasi untuk berbagai use case
  5. ๐ŸŒ Bilingual - Support optimal untuk Indonesia & English
๐ŸŽฏ Keunggulan vs Model Lain
Fitur Caca caca-65M LLaMA-2 65.01M GPT-3 65.01M
Attention Type GQA GQA MHA
Position Encoding RoPE + ALiBI RoPE Learned
Activation SwiGLU SwiGLU GELU
Flash Attention โœ… v2 โœ… v1/v2 โŒ
Long Context Sliding Window + Sink โœ… Limited
MoE Support โœ… Optional โŒ โŒ
Multimodal โœ… Optional โŒ โŒ
Quantization 4/8-bit 4/8-bit Limited

๐ŸŽฏ Use Cases & Applications
โœ… Cocok Untuk

๐Ÿ”ฌ Research & Development

  • Eksperimen arsitektur transformer
  • Ablation studies
  • Novel training techniques
  • Architecture search

๐Ÿ“š Academic & Education

  • Thesis & research papers
  • Teaching materials
  • Student projects
  • LLM internals understanding

๐Ÿš€ Base Model for Fine-tuning

  • Task-specific models
  • Domain adaptation
  • Instruction tuning
  • RLHF experiments

๐Ÿ’ก Prototyping

  • Proof of concept
  • Feature testing
  • A/B testing architectures
  • Benchmark comparisons
โŒ Tidak Cocok Untuk
  • ๐Ÿšซ Production Applications - Model belum dilatih, output random
  • ๐Ÿšซ Real-world Deployment - Perlu training & safety alignment dulu
  • ๐Ÿšซ Safety-critical Systems - Tidak ada safety guardrails
  • ๐Ÿšซ Direct User-facing Apps - Output tidak dapat diprediksi
  • ๐Ÿšซ Commercial Use (as-is) - Harus dilatih terlebih dahulu

๐Ÿ“Š Spesifikasi Model
Parameter Value Parameter Value
Total Parameters 65,012,096 Vocab Size 16,000
Hidden Size 256 Intermediate Size 704
Num Layers 77 Attention Heads 4
KV Heads (GQA) 2 Head Dimension 64
Max Context Length 2,048 RoPE Base (ฮธ) 10,000
Model Size (FP16) 0.13 GB Formatted Size 65.01M

๐ŸŽฏ Core Features
๐Ÿ” Klik untuk expand/collapse
  • โœ… Grouped Query Attention (GQA) - Efisiensi memori dan komputasi superior

    • Query heads: 4
    • KV heads: 2
    • Ratio: 2:1 (hemat ~50% memory KV cache)
    • Benefit : Inferensi lebih cepat dengan memory footprint lebih kecil
  • โœ… Rotary Position Embeddings (RoPE) - Generalisasi konteks panjang lebih baik

    • Theta (ฮธ): 10,000
    • Support extrapolation untuk konteks > training length
    • Benefit : Performa stabil pada sequence length yang belum pernah dilihat saat training
  • โœ… RMSNorm - Normalisasi lebih stabil dan ~50% lebih cepat dari LayerNorm

    • Epsilon: 1e-06
    • Benefit : Training lebih stabil, inference lebih cepat, gradient flow lebih baik
  • โœ… SwiGLU Activation - Performa 10-15% lebih baik dari ReLU/GELU

    • Intermediate size: 704 (2.8x hidden)
    • Benefit : Kapasitas model lebih besar tanpa menambah parameter signifikan
  • โœ… Flash Attention 2 - Akselerasi hingga 3x dengan memory efficiency

    • Otomatis aktif jika tersedia CUDA device
    • IO-aware algorithm untuk minimal HBM access
    • Benefit : Training & inference jauh lebih cepat, support batch size lebih besar
๐Ÿ”ฅ Advanced Features
๐ŸŽฏ Mekanisme Attention
  • โšก Flash Attention v2 - Algoritma IO-aware yang 3x lebih cepat dari attention standar
  • ๐Ÿ”‘ Grouped Query Attention (GQA) - 4 Query heads : 2 KV heads
    • Rasio kompresi: 2:1 (hemat ~50% memory KV cache)
  • ๐Ÿš€ xFormers Support - Fallback memory-efficient attention
  • ๐ŸŽฏ PyTorch SDPA - Native scaled dot product attention
๐Ÿ“ Position Encodings
  • ๐Ÿ”„ RoPE (Rotary Position Embeddings) - Base frequency ฮธ=10,000
    • Generalisasi lebih baik untuk sequence panjang dibanding absolute PE
๐ŸŽ“ Optimisasi Training
  • ๐Ÿ’พ Gradient Checkpointing - Trade compute for memory (support model hingga 100B+ params)
  • ๐ŸŽฏ Mixed Precision Training - Support FP16, BF16, dan TF32
  • ๐Ÿ“‰ Dropout Regularization
    • Hidden dropout: 0.1
    • Attention dropout: 0.0
    • Residual dropout: 0.1
๐Ÿ“ฆ Dukungan Quantization
  • 4๏ธโƒฃ 4-bit Quantization - NF4 & FP4 via bitsandbytes
    • Memory reduction: ~ 75% (4GB โ†’ 1GB)
    • Accuracy loss: <2% pada kebanyakan tasks
    • Support double quantization untuk kompresi maksimal
  • 8๏ธโƒฃ 8-bit Quantization - LLM.int8() dengan outlier handling
    • Memory reduction: ~ 50% (4GB โ†’ 2GB)
    • Accuracy loss: <1%
  • ๐Ÿ”„ Dynamic Quantization - Runtime quantization tanpa calibration
๐Ÿ”ฌ Advanced Features
  • ๐Ÿ“Š Automatic Mixed Precision (AMP) - Dynamic loss scaling
  • ๐ŸŽฏ Gradient Clipping - Stabilitas training dengan max norm clipping
  • ๐Ÿ“ˆ Learning Rate Scheduling - Support cosine, linear, warmup
  • ๐Ÿ’ก Smart Memory Management - Auto cache clearing & monitoring
  • ๐Ÿ” Metrics Tracking - Real-time perplexity, loss, gradient norms
  • ๐Ÿ›ก๏ธ NaN/Inf Detection - Automatic recovery dari numerical instability

๐Ÿ’พ Kebutuhan Memory
Training Requirements
Configuration Model Weights + Optimizer States Total Training
FP32 (AdamW) 0.26 GB +0.78 GB 1.04 GB
Mixed Precision 0.13 GB +0.91 GB 1.04 GB
+ Gradient Checkpointing Menghemat ~30-50% activation memory ~0.62 GB
Inference Requirements
Precision Model Size KV Cache (2K ctx) Total Memory Memory Saving
FP16 / BF16 0.13 GB 0.08 GB 0.21 GB Baseline
INT8 0.07 GB 0.08 GB 0.15 GB ~50% โ†“
INT4 (NF4) 0.03 GB 0.08 GB 0.11 GB ~75% โ†“

๐Ÿ’ก Note : KV cache bertambah secara linear dengan panjang sequence. Untuk context 8K, kalikan nilai KV cache dengan 4.

Performance Estimates
Metric Value Notes
FLOPs per Token 130,024,192 Forward pass only
TFLOPs per Token 0.0001 โ‰ˆ 6ร— untuk backward
Bandwidth (FP16) 0.13 GB/token Memory bandwidth requirement

๐Ÿ“ Struktur Arsitektur Lengkap
๐Ÿ” Klik untuk lihat detail arsitektur
CacaForCausalLM (65.01M)
โ”‚
โ”œโ”€ Embedding: 16,000 ร— 256
โ”‚
โ”œโ”€ Transformer Layers (77x)
โ”‚  โ”œโ”€ RMSNorm
โ”‚  โ”œโ”€ Attention (GQA)
โ”‚  โ”‚  โ”œโ”€ Q: 4 heads ร— 64 dim
โ”‚  โ”‚  โ”œโ”€ KV: 2 heads ร— 64 dim
โ”‚  โ”‚  โ”œโ”€ RoPE (ฮธ=10,000)
โ”‚  โ”‚  โ””โ”€ Flash Attention v2
โ”‚  โ”œโ”€ Residual
โ”‚  โ”œโ”€ RMSNorm
โ”‚  โ”œโ”€ FFN (SwiGLU)
โ”‚  โ”‚  โ”œโ”€ Gate: 256 โ†’ 704
โ”‚  โ”‚  โ”œโ”€ Up: 256 โ†’ 704
โ”‚  โ”‚  โ””โ”€ Down: 704 โ†’ 256
โ”‚  โ””โ”€ Residual
โ”‚
โ”œโ”€ Final RMSNorm
โ””โ”€ LM Head: 256 โ†’ 16,000

โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
๐Ÿ“Š PARAMETER BREAKDOWN:
โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
Embeddings:                 4,096,000 (  6.3%)
Transformer Layers:        56,770,560 ( 87.3%)
  โ”œโ”€ Attention:            15,138,816
  โ””โ”€ FFN:                  41,631,744
Final Norm:                       256 (  0.0%)
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
TOTAL:                     65,012,096 (100.0%)
โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

Key Design Decisions:

  1. GQA over MHA : Hemat 50% KV cache memory dengan minimal accuracy loss
  2. SwiGLU over GELU : ~10% better performance pada language modeling
  3. RMSNorm over LayerNorm : Lebih cepat & stabil, tanpa bias term
  4. RoPE over Learned : Better extrapolation untuk sequence length > training
  5. No Bias in Linear : Mengikuti modern LLM best practices (LLaMA-style)

๐Ÿ“š Dokumentasi
๐Ÿ“ฆ Instalasi Dependencies
# Core dependencies (REQUIRED)
pip install torch>=2.0.0 transformers>=4.35.0 accelerate safetensors

# Optional: Untuk performa maksimal
pip install flash-attn --no-build-isolation  # Flash Attention 2 (3x speedup)
pip install xformers                          # Memory efficient attention
pip install bitsandbytes                      # 4/8-bit quantization

# Optional: Untuk monitoring & profiling
pip install tensorboard wandb               # Training monitoring
pip install gputil psutil                   # Resource monitoring

Compatibility Matrix:

Component Version Note
Python 3.8 - 3.11 3.11 recommended
PyTorch โ‰ฅ 2.0.0 2.1+ untuk SDPA optimal
CUDA 11.8 / 12.1 Untuk Flash Attention
Transformers โ‰ฅ 4.35.0 Untuk AutoModel support
Cara Penggunaan
1๏ธโƒฃ Basic Loading
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
import torch

# Load configuration
config = AutoConfig.from_pretrained(
    "Lyon28/caca-65M-untrained",
    trust_remote_code=True
)

# Load model (FP16 untuk efisiensi)
model = AutoModelForCausalLM.from_pretrained(
    "Lyon28/caca-65M-untrained",
    config=config,
    trust_remote_code=True,
    torch_dtype=torch.float16,
    device_map="auto"  # Automatic device placement
)

# Model ini UNTRAINED - butuh training dulu!
print(f"Model loaded: {model.num_parameters():,} parameters")
print("โš ๏ธ  Model ini belum dilatih dan belum bisa digunakan untuk inference")
2๏ธโƒฃ Quantized Loading (4-bit/8-bit)
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True
)

# Load model dengan quantization
model = AutoModelForCausalLM.from_pretrained(
    "Lyon28/caca-65M-untrained",
    trust_remote_code=True,
    quantization_config=bnb_config,
    device_map="auto"
)

print(f"Memory footprint: ~0.03GB (4-bit)")
3๏ธโƒฃ Training Setup
from transformers import TrainingArguments, Trainer

# Training configuration
training_args = TrainingArguments(
    output_dir="./output",
    per_device_train_batch_size=1,
    gradient_accumulation_steps=16,
    learning_rate=2e-4,
    max_steps=10000,
    lr_scheduler_type="cosine",
    warmup_steps=500,
    logging_steps=10,
    save_steps=500,
    fp16=True,  # Mixed precision
    gradient_checkpointing=True,  # Memory efficient
)

# Initialize trainer
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
)

# Start training
trainer.train()
Advanced Usage
Gradient Checkpointing (Memory Efficient)
model.gradient_checkpointing_enable()
print("โœ… Gradient checkpointing enabled - saves ~40% memory")
Custom Training Loop
from torch.optim import AdamW
from torch.cuda.amp import autocast, GradScaler

optimizer = AdamW(model.parameters(), lr=2e-4)
scaler = GradScaler()

for batch in dataloader:
    # Mixed precision forward
    with autocast(dtype=torch.bfloat16):
        outputs = model(**batch)
        loss = outputs.loss

    # Backward with gradient scaling
    scaler.scale(loss).backward()
    scaler.step(optimizer)
    scaler.update()
    optimizer.zero_grad()
Multi-GPU Training (DDP)
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel

# Initialize process group
dist.init_process_group(backend="nccl")

# Wrap model
model = DistributedDataParallel(
    model,
    device_ids=[local_rank],
    find_unused_parameters=False
)

โš™๏ธ Konfigurasi Detail
Full Configuration JSON
{
  "architectures": ["CacaForCausalLM"],
  "model_type": "caca",
  "vocab_size": 16000,
  "hidden_size": 256,
  "intermediate_size": 704,
  "num_hidden_layers": 77,
  "num_attention_heads": 4,
  "num_key_value_heads": 2,
  "head_dim": 64,
  "max_position_embeddings": 2048,
  "rope_theta": 10000,
  "rms_norm_eps": 1e-06,
  "use_cache": true,
  "use_qk_norm": true,
  "use_flash_attn": true,
  "attention_dropout": 0.0,
  "hidden_dropout": 0.1,
  "torch_dtype": "float16"
}
Custom Configuration
from transformers import AutoConfig

# Load dan modifikasi config
config = AutoConfig.from_pretrained("Lyon28/caca-65M-untrained")

# Custom modifications
config.max_position_embeddings = 16384  # Extend context
config.rope_scaling = {"type": "linear", "factor": 2.0}
config.use_flash_attn = True
config.hidden_dropout = 0.05

# Save custom config
config.save_pretrained("./custom_config")

๐Ÿ”ฌ Arsitektur
Layer Structure
Input Tokens
    โ†“
Embedding Layer (16,000 โ†’ 256)
    โ†“
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Decoder Block ร— 77                  โ”‚
โ”‚                                     โ”‚
โ”‚  โ”Œโ”€ RMSNorm                        โ”‚
โ”‚  โ”œโ”€ Multi-Head Attention (GQA)     โ”‚
โ”‚  โ”‚  - Flash Attention v2           โ”‚
โ”‚  โ”‚  - 4 Query heads, 2 KV heads       โ”‚
โ”‚  โ”‚  - RoPE position encoding       โ”‚
โ”‚  โ”œโ”€ Residual Connection            โ”‚
โ”‚  โ”‚                                  โ”‚
โ”‚  โ”œโ”€ RMSNorm                        โ”‚
โ”‚  โ”œโ”€ Feed-Forward Network (SwiGLU)  โ”‚
โ”‚  โ”‚  - Gate: 256 โ†’ 704     โ”‚
โ”‚  โ”‚  - Up:   256 โ†’ 704     โ”‚
โ”‚  โ”‚  - Down: 704 โ†’ 256     โ”‚
โ”‚  โ””โ”€ Residual Connection            โ”‚
โ”‚                                     โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
    โ†“
RMSNorm (Final)
    โ†“
LM Head (256 โ†’ 16,000)
    โ†“
Output Logits
Attention Mechanism (GQA)
Query:  [4 heads ร— 64 dim] = 256
Key:    [2 heads ร— 64 dim] = 128
Value:  [2 heads ร— 64 dim] = 128

Grouped Query Attention:
- Setiap 2 query heads berbagi 1 KV head
- Memory KV cache: 50% lebih kecil dari Multi-Head Attention
- Kualitas mendekati MHA, speed mendekati MQA
Feed-Forward Network (SwiGLU)
FFN(x) = (SiLU(xW_gate) โŠ™ xW_up) W_down

Where:
- W_gate: 256 ร— 704
- W_up:   256 ร— 704
- W_down: 704 ร— 256
- SiLU(x) = x ยท sigmoid(x)
- โŠ™ = element-wise multiplication
๐Ÿ’ฌ Format Chat & Prompt Engineering
๐Ÿ“ Chat Template

Model mendukung format chat standar untuk conversational AI:

# Format chat template bawaan
chat_template = """
{% for message in messages %}
{% if message['role'] == 'system' %}
System: {{ message['content'] }}

{% elif message['role'] == 'user' %}
User: {{ message['content'] }}

{% elif message['role'] == 'assistant' %}
Assistant: {{ message['content'] }}

{% endif %}
{% endfor %}
{% if add_generation_prompt %}Assistant:{% endif %}
"""

# Contoh penggunaan
messages = [
    {"role": "system", "content": "Kamu adalah asisten AI yang membantu dan ramah."},
    {"role": "user", "content": "Jelaskan tentang fotosintesis"},
    {"role": "assistant", "content": "Fotosintesis adalah proses di mana tumbuhan mengubah cahaya matahari menjadi energi kimia..."},
    {"role": "user", "content": "Apa manfaatnya bagi manusia?"},
]

# Apply template
formatted = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

print(formatted)
# Output:
# System: Kamu adalah asisten AI yang membantu dan ramah.
#
# User: Jelaskan tentang fotosintesis
# Assistant: Fotosintesis adalah proses di mana tumbuhan...
# User: Apa manfaatnya bagi manusia?
# Assistant:

๐ŸŽฏ Use Cases

Model ini dirancang untuk berbagai aplikasi NLP setelah melalui proses training:

Text Generation
  • โœ๏ธ Creative writing & storytelling
  • ๐Ÿ“ฐ Article generation
  • ๐Ÿ’ฌ Conversational AI
  • ๐Ÿ”„ Text completion
Language Understanding
  • ๐Ÿ“Š Text classification
  • ๐Ÿท๏ธ Named Entity Recognition (NER)
  • โ“ Question Answering
  • ๐Ÿ“ Summarization
Code Generation
  • ๐Ÿ’ป Code completion
  • ๐Ÿ› Bug fixing suggestions
  • ๐Ÿ“š Documentation generation
  • ๐Ÿ”„ Code translation
Multilingual Tasks
  • ๐ŸŒ Translation (ID โ†” EN)
  • ๐Ÿ—ฃ๏ธ Cross-lingual understanding
  • ๐ŸŒ Multilingual classification

๐Ÿ“ˆ Benchmark & Evaluation

โš ๏ธ Model belum melalui evaluasi karena status untrained

Setelah training, model akan dievaluasi pada:

Indonesian Benchmarks
  • IndoNLU : Comprehensive Indonesian NLU tasks
  • IndoQA : Indonesian Question Answering
  • IndoSum : Summarization
  • IndoNER : Named Entity Recognition
Multilingual Benchmarks
  • MMLU : Massive Multitask Language Understanding
  • HellaSwag : Common sense reasoning
  • ARC : Science QA
  • TruthfulQA : Truthfulness evaluation
Generation Quality
  • Perplexity : Language modeling quality
  • BLEU/ROUGE : Translation & summarization
  • Human Evaluation : Fluency, coherence, factuality

๐Ÿ› ๏ธ Development & Training Tips
Optimal Batch Size
# Rule of thumb untuk 65.01M model
# GPU Memory โ†’ Batch size per device

if gpu_memory >= 80:  # A100 80GB
    batch_size = 246
    gradient_accumulation = 1
elif gpu_memory >= 40:  # A100 40GB
    batch_size = 123
    gradient_accumulation = 1
elif gpu_memory >= 24:  # RTX 3090/4090
    batch_size = 1
    gradient_accumulation = 1

# Effective batch size = batch_size ร— gradient_accumulation ร— num_gpus
Learning Rate Scheduling
# Recommended untuk 65.01M model
learning_rate = 0.0005  # Base LR
warmup_ratio = 0.05  # 5% of total steps
lr_scheduler = "cosine"  # atau "linear"

# Learning rate scaling rule:
# LR โˆ sqrt(batch_size)
# Untuk batch size 256: LR = 0.0005
# Untuk batch size 512: LR = 7.07e-04
Gradient Clipping
# Prevent gradient explosion
max_grad_norm = 1.0  # Clip at 1.0

# Monitor gradients
from torch.nn.utils import clip_grad_norm_

grad_norm = clip_grad_norm_(model.parameters(), max_grad_norm)
if grad_norm > 10.0:
    print(f"โš ๏ธ High gradient norm: {grad_norm:.2f}")
Training Stability
# Tips untuk stable training:

1. **Warmup**: Mulai dengan LR rendah
2. **Gradient Checkpointing**: Kurangi memory footprint
3. **Mixed Precision**: Gunakan BF16 jika tersedia (lebih stable dari FP16)
4. **Batch Size**: Start small, increase gradually
5. **Monitor**: Track loss, perplexity, gradient norms

๐Ÿ”ง Troubleshooting
Out of Memory (OOM)
# Solusi OOM saat training:

โœ… 1. Enable gradient checkpointing
model.gradient_checkpointing_enable()

โœ… 2. Reduce batch size
per_device_train_batch_size = 1

โœ… 3. Increase gradient accumulation
gradient_accumulation_steps = 32

โœ… 4. Use quantization
load_in_8bit = True  # atau load_in_4bit

โœ… 5. Reduce sequence length
max_length = 2048  # Start dengan ini

โœ… 6. CPU offloading (jika perlu)
device_map = "auto"
offload_folder = "offload"
Slow Training
# Optimasi kecepatan training:

โœ… 1. Flash Attention
config.use_flash_attn = True  # 2-3x speedup

โœ… 2. Compile model (PyTorch 2.0+)
model = torch.compile(model, mode="reduce-overhead")

โœ… 3. DataLoader optimization
dataloader = DataLoader(
    dataset,
    batch_size=batch_size,
    num_workers=4,  # Parallel data loading
    pin_memory=True,  # Faster GPU transfer
    prefetch_factor=2
)

โœ… 4. Mixed precision
use_fp16 = True  # atau bf16

โœ… 5. Optimize communication (multi-GPU)
find_unused_parameters = False
gradient_as_bucket_view = True
NaN Loss
# Jika loss menjadi NaN:

โœ… 1. Reduce learning rate
learning_rate = learning_rate * 0.1

โœ… 2. Check gradient norms
clip_grad_norm_(model.parameters(), 1.0)

โœ… 3. Use BF16 instead of FP16
torch_dtype = torch.bfloat16  # Lebih stable

โœ… 4. Add epsilon to RMSNorm
rms_norm_eps = 1e-5  # Increase jika perlu

โœ… 5. Check data
# Pastikan tidak ada inf/nan di dataset
assert not torch.isnan(input_ids).any()
assert not torch.isinf(attention_mask).any()

๐Ÿšซ Prohibited Uses

Model ini TIDAK BOLEH digunakan untuk:

  • ๐Ÿšซ Harmful content generation (violence, self-harm, illegal acts)
  • ๐Ÿšซ Misinformation/disinformation campaigns
  • ๐Ÿšซ Harassment or hate speech
  • ๐Ÿšซ Impersonation or identity theft
  • ๐Ÿšซ Child safety violations (CSAM, grooming, exploitation)
  • ๐Ÿšซ Privacy violations (doxxing, stalking, surveillance abuse)
  • ๐Ÿšซ Malicious code generation (malware, exploits, etc)
  • ๐Ÿšซ Spam or manipulation (fake reviews, astroturfing)
  • ๐Ÿšซ Medical/legal advice (tanpa disclaimer & expert review)
  • ๐Ÿšซ Financial fraud (scams, market manipulation)

Violation consequences: Model access revocation + legal action jika applicable


๐Ÿ“œ License & Citation
๐Ÿ“„ License

Model ini dirilis di bawah Apache License 2.0

โœ… Anda BEBAS untuk:

  • โœ”๏ธ Gunakan secara komersial
  • โœ”๏ธ Modifikasi sesuka hati
  • โœ”๏ธ Distribusi ulang
  • โœ”๏ธ Patent use
  • โœ”๏ธ Private use

โš ๏ธ Dengan syarat:

  • ๐Ÿ“„ Include license & copyright notice
  • ๐Ÿ“ State changes yang dibuat
  • ๐Ÿ“‹ Disclaimer of warranty

โŒ Tanpa jaminan apapun (use at your own risk)

Full license text : Apache-2.0

๐Ÿ“– Citation

Jika Anda menggunakan model ini dalam penelitian, mohon sitasi:

@misc{cacacaca65m,
  author = {Lyon},
  title = {Caca-caca-65M: Modern Transformer Architecture with Grouped Query Attention},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Model Hub},
  howpublished = {\url{https://huggingface.co/Lyon28/caca-65M-untrained}},
  note = {Untrained model with 65,012,096 parameters}
}

APA Style:

Lyon. (2026). Caca-caca-65M: Modern Transformer Architecture with Grouped
Query Attention [Untrained model]. Hugging Face.
https://huggingface.co/Lyon28/caca-65M-untrained

MLA Style:

Lyon. "Caca-caca-65M: Modern Transformer Architecture with Grouped Query Attention."
Hugging Face, 2026, huggingface.co/Lyon28/caca-65M-untrained.

๐Ÿ™ Acknowledgments

Model ini berdiri di pundak para raksasa! Terima kasih kepada:

๐Ÿ›๏ธ Klik untuk daftar lengkap acknowledgments
๐Ÿ—๏ธ Core Architecture
  • LLaMA/LLaMA 2 (Meta AI, 2023) - Decoder-only architecture, RMSNorm, SwiGLU
  • GPT-3 (OpenAI, 2020) - Transformer language modeling paradigm
  • PaLM (Google, 2022) - SwiGLU activation insights
๐ŸŽฏ Attention Mechanisms
๐Ÿ“ Position Encodings
๐ŸชŸ Long Context & Efficiency
๐Ÿง  Mixture of Experts
  • Mixtral 8x7B (Albert Jiang et al., Mistral AI, 2024)
  • Switch Transformers (William Fedus et al., Google, 2021)
  • GLaM (Nan Du et al., Google, 2021) - Generalist Language Model
  • Expert Choice Routing (Yanqi Zhou et al., Google, 2022)
    • Better load balancing
๐ŸŽ“ Training Optimizations
๐Ÿ“ฆ Quantization
๐ŸŽจ Multimodal
๐Ÿ› ๏ธ Normalization & Activations
๐Ÿ”ง Tools & Frameworks
  • ๐Ÿค— Hugging Face - Transformers, Accelerate, PEFT
    • Making NLP accessible to everyone
  • PyTorch - Deep learning framework
    • Facebook AI Research team
  • Safetensors - Secure serialization
    • Hugging Face team
  • DeepSpeed - Distributed training
    • Microsoft Research
  • Flash Attention Implementation - Tri Dao & team
๐Ÿ‡ฎ๐Ÿ‡ฉ Indonesian NLP Community

Special thanks to Indonesian NLP researchers & practitioners yang telah membangun foundation untuk Indonesian language AI.


๐Ÿ“„ License

Model ini dirilis di bawah Apache License 2.0 .

Ketentuan Penggunaan:
  • โœ… Bebas digunakan untuk keperluan komersial dan non-komersial
  • โœ… Modifikasi diperbolehkan
  • โœ… Distribusi diperbolehkan dengan attribution
  • โš ๏ธ No Warranty - model disediakan "as is"
  • ๐Ÿ“ Attribution Required - sertakan copyright notice

Lihat LICENSE untuk detail lengkap.


๐Ÿค Contributing

Kami sangat terbuka untuk kontribusi! Berikut cara Anda bisa berkontribusi:

Training & Fine-tuning
  • ๐ŸŽ“ Train model ini dengan dataset Anda
  • ๐Ÿ“Š Share benchmark results
  • ๐Ÿ”ฌ Experiment dengan hyperparameters
Code & Architecture
  • ๐Ÿ› Report bugs atau issues
  • ๐Ÿ’ก Suggest improvements
  • ๐Ÿ”ง Submit pull requests
Documentation
  • ๐Ÿ“š Improve documentation
  • ๐ŸŒ Add translations
  • โœ๏ธ Write tutorials & guides
Dataset & Evaluation
  • ๐Ÿ“ Contribute training data
  • ๐Ÿงช Create evaluation benchmarks
  • ๐ŸŽฏ Share fine-tuned versions

๐Ÿ‘ฅ Team & Acknowledgments
Core Team
  • LyonPoy - Architecture design & implementation
Special Thanks
  • ๐Ÿค— Hugging Face - Infrastructure & community
  • โšก FlashAttention Team - Efficient attention implementation
  • ๐Ÿง  Anthropic, Google, Meta - Research inspirations
Community

Terima kasih kepada komunitas open-source yang telah berkontribusi pada:

  • Transformers library
  • PyTorch framework
  • Datasets & evaluation tools

๐Ÿ“ž Contact & Support
Community

๐ŸŒŸ Star History
๐Ÿ’ Dibuat dengan โค๏ธ untuk Komunitas AI Indonesia
Caca Logo
Terima kasih telah menggunakan Caca!

Jika model ini berguna, jangan lupa โญ repository kami!

โญ
Star Repo
Show your support
๐Ÿ”—
Share
Tell your friends
๐Ÿ’ฌ
Join Discussion
Ask questions
๐Ÿค
Contribute
Make it better
๐Ÿš€ Happy Training! ๐Ÿš€

Model ini menunggu untuk dilatih dan menjadi foundation untuk aplikasi AI Anda.

๐Ÿ“ฅ Download Model โ€ข ๐Ÿ“– Read Docs โ€ข ๐Ÿ’ฌ Join Community


๐Ÿ“Š Model Statistics
Parameters Status License Architecture Type Context
๐ŸŽจ Daily Inspiration
Daily Quote

๐Ÿ“ˆ Quick Stats
Metric Value
๐Ÿ’Ž Total Parameters 65,012,096
๐Ÿ—๏ธ Layers 77
๐ŸŽฏ Attention Heads 4
๐Ÿ“– Max Context 2,048 tokens
๐Ÿ’พ Size (FP16) 0.13 GB
๐Ÿ’พ Size (INT4) 0.03 GB

Model ini adalah bagian dari Caca Project - Open source initiative untuk membangun Indonesian LLM ecosystem.
Created with ๐Ÿ’ป by @Lyon28 | Licensed under Apache 2.0 | Built with ๐Ÿค— HuggingFace



๐ŸŒŸ "Dari nol, untuk semua" ๐ŸŒŸ

Last updated: january 2026


Built with โค๏ธ by Caca Transformers Team
Powered by ๐Ÿค— Transformers โ€ข โšก PyTorch โ€ข ๐Ÿ”ฅ Flash Attention

Runs of Lyon28 caca-65M-untrained on huggingface.co

16
Total runs
0
24-hour runs
0
3-day runs
-4
7-day runs
5
30-day runs

More Information About caca-65M-untrained huggingface.co Model

More caca-65M-untrained license Visit here:

https://choosealicense.com/licenses/apache-2.0

caca-65M-untrained huggingface.co

caca-65M-untrained huggingface.co is an AI model on huggingface.co that provides caca-65M-untrained's model effect (), which can be used instantly with this Lyon28 caca-65M-untrained model. huggingface.co supports a free trial of the caca-65M-untrained model, and also provides paid use of the caca-65M-untrained. Support call caca-65M-untrained model through api, including Node.js, Python, http.

caca-65M-untrained huggingface.co Url

https://huggingface.co/Lyon28/caca-65M-untrained

Lyon28 caca-65M-untrained online free

caca-65M-untrained huggingface.co is an online trial and call api platform, which integrates caca-65M-untrained's modeling effects, including api services, and provides a free online trial of caca-65M-untrained, you can try caca-65M-untrained online for free by clicking the link below.

Lyon28 caca-65M-untrained online free url in huggingface.co:

https://huggingface.co/Lyon28/caca-65M-untrained

caca-65M-untrained install

caca-65M-untrained is an open source model from GitHub that offers a free installation service, and any user can find caca-65M-untrained on GitHub to install. At the same time, huggingface.co provides the effect of caca-65M-untrained install, users can directly use caca-65M-untrained installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

caca-65M-untrained install url in huggingface.co:

https://huggingface.co/Lyon28/caca-65M-untrained

Url of caca-65M-untrained

caca-65M-untrained huggingface.co Url

Provider of caca-65M-untrained huggingface.co

Lyon28
ORGANIZATIONS

Other API from Lyon28

huggingface.co

Total runs: 449
Run Growth: 449
Growth Rate: 100.00%
Updated:December 11 2025
huggingface.co

Total runs: 255
Run Growth: 255
Growth Rate: 100.00%
Updated:December 11 2025
huggingface.co

Total runs: 202
Run Growth: 202
Growth Rate: 100.00%
Updated:December 11 2025
huggingface.co

Total runs: 125
Run Growth: 125
Growth Rate: 100.00%
Updated:December 09 2025
huggingface.co

Total runs: 123
Run Growth: 123
Growth Rate: 100.00%
Updated:December 09 2025
huggingface.co

Total runs: 40
Run Growth: 40
Growth Rate: 100.00%
Updated:December 09 2025
huggingface.co

Total runs: 27
Run Growth: 27
Growth Rate: 100.00%
Updated:December 09 2025
huggingface.co

Total runs: 8
Run Growth: 0
Growth Rate: 0.00%
Updated:July 08 2025
huggingface.co

Total runs: 7
Run Growth: -2
Growth Rate: -28.57%
Updated:July 11 2025
huggingface.co

Total runs: 6
Run Growth: -10
Growth Rate: -166.67%
Updated:August 14 2025
huggingface.co

Total runs: 6
Run Growth: -7
Growth Rate: -116.67%
Updated:August 12 2025
huggingface.co

Total runs: 6
Run Growth: 3
Growth Rate: 50.00%
Updated:July 01 2025
huggingface.co

Total runs: 5
Run Growth: 5
Growth Rate: 100.00%
Updated:May 19 2025
huggingface.co

Total runs: 5
Run Growth: 2
Growth Rate: 40.00%
Updated:January 06 2026
huggingface.co

Total runs: 2
Run Growth: 2
Growth Rate: 100.00%
Updated:May 19 2025
huggingface.co

Total runs: 1
Run Growth: -2
Growth Rate: -200.00%
Updated:June 28 2025
huggingface.co

Total runs: 1
Run Growth: -2
Growth Rate: -200.00%
Updated:July 01 2025
huggingface.co

Total runs: 1
Run Growth: 1
Growth Rate: 100.00%
Updated:May 19 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 05 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 26 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 03 2025