FlameF0X / FWKV-ROSA

huggingface.co
Total runs: 109
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: July 25 2026
text-generation

Introduction of FWKV-ROSA

Model Details of FWKV-ROSA

FWKV-ROSA --- Read more here

A 56M‑parameter , from‑scratch recurrent language model that combines a per‑channel leaky integrator (FWKV) with the RWKV‑8 ROSA copy‑signal mechanism. It is a research experiment designed to explore how far purely recurrent architectures can go on small‑scale, curated conversational data.

The model was chat‑tuned on HuggingFaceH4/ultrachat_200k and uses a simple two‑role template:

<|user|> Your message
<|assistant|> Model reply<|endoftext|>
Model Description
  • Architecture: 14 stacked FWKV blocks. Each block replaces the standard attention with a fixed, data‑independent decayed accumulator: stateₜ = W·stateₜ₋₁ + kₜ·vₜ , where W = clamp(sigmoid(w), min=0.1) . The recurrence is computed exactly via a vectorised parallel scan (no approximations).
  • ROSA (Rapid Online Suffix Automaton): A parameter‑free, causal predictor that injects the token that historically followed the longest matching suffix of the current context. The ROSA signal is embedded and added to the input representation of each token.
  • Factorised embedding/head: 128‑dimensional embedding space, projected to a 512‑dimensional model space, with tied weights.
  • Context length: 1024 tokens. No positional embeddings are used, making the context window “free” in terms of parameters.
  • Tokenizer: GPT‑2 tokenizer extended with the special tokens <|user|> and <|assistant|> .
Uses
Direct Use

FWKV-ROSA is intended for research on efficient language models and for educational demonstrations of recurrent architectures. You can chat with it in a multi‑turn setting using the template above.

Out‑of‑Scope Use
  • This model is not suitable for any production or safety‑critical application.
  • It has not been aligned with RLHF or other safety methods and may generate inappropriate or harmful content.
  • The limited size and training data mean it cannot be relied upon for factual knowledge or reasoning.
Bias, Risks, and Limitations
  • Trained on a relatively small synthetic dataset, the model can produce repetitive or nonsensical output.
  • The ROSA copy mechanism may occasionally copy large chunks of the user’s prompt verbatim.
  • Biases present in the original UltraChat data are likely reflected in the model’s responses.
How to Get Started

The model relies on a custom architecture. To load it, you must provide the modeling_fwkv.py file (found in the repository) and trust the remote code:

from transformers import AutoTokenizer, AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "FlameF0X/FWKV-ROSA",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("FlameF0X/FWKV-ROSA")

Then format your prompts exactly with the chat tokens:

device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device).eval()

prompt = "<|user|> What is the capital of France?\n<|assistant|>"
input_ids = tokenizer.encode(prompt)
# ROSA IDs must be computed – you can import `rosa` from modeling_fwkv
from modeling_fwkv import rosa
rosa_ids = torch.tensor([rosa(input_ids)], device=device)
out = model(input_ids=torch.tensor([input_ids], device=device),
            rosa_ids=rosa_ids, use_cache=True)
# Continue autoregressive sampling…

For a fully working chat demo, see the Gradio app provided in the repository.

Training Details
Dataset
  • Name: UltraChat 200k
  • Splits: train_sft (20,000 examples), test_sft (1,000 examples)
  • Format: multi‑turn conversations; only assistant tokens contribute to the loss.
Training Procedure
Hyperparameter Value
Architecture 14 FWKV blocks, d_model=512, d_emb=128
FFN multiplier 4
WKV decay floor 0.1
Batch size (per GPU) 8
Gradient accumulation 4
Effective batch size 32
Learning rate 3×10⁻⁴ (cosine schedule)
Weight decay 0.1
Gradient clipping 1.0
Optimizer AdamW (fused)
Precision bfloat16 mixed
Epochs 2
Hardware 1× NVIDIA T4 (15 GB)
Training time ~2 hours
Speed ~80,000 tokens/second

Gradient checkpointing was enabled to fit the 1024‑token sequences in memory.

Evaluation
Metric Value
Validation loss 4.216
Validation perplexity 67.78

The perplexity is relatively high due to the small model size and limited training data. It is comparable to other similarly‑sized recurrent LMs on UltraChat.

Environmental Impact

The training ran for about 2 hours on a single NVIDIA T4 GPU (maximum power draw ~70 W), resulting in an estimated 0.14 kWh of electricity consumption and approximately 0.06 kg CO₂eq (assuming a grid carbon intensity of 0.4 kg/kWh). This is a negligible footprint.

Technical Specifications
  • Model type: Recurrent neural network (linear RNN)
  • Parameters: 56.2 million
    • Backbone (FWKV blocks + embeddings): ~50M
    • ROSA embedding: ~6.2M
  • Checkpoint format: PyTorch safetensors
  • Required files in the repo:
    • config.json
    • model.safetensors (or pytorch_model.bin )
    • modeling_fwkv.py
    • tokenizer.json / vocab.json / merges.txt
  • Auto‑mapping: The config.json includes "auto_map": { "AutoModelForCausalLM": "modeling_fwkv.FWKVLanguageModel" } , so loading with trust_remote_code=True will automatically locate the correct class.
Citation

If you use FWKV-ROSA in your research, please cite it as:

@misc{fwkv-rosa,
  author = {FlameF0X},
  title = {FWKV-ROSA: A 56M Recurrent Chat LM with RWKV-style Decay and ROSA Copy Signal},
  year = {2025},
  howpublished = {\url{https://huggingface.co/FlameF0X/FWKV-ROSA}},
}
Additional Information

This model was built as an experiment to test the combination of a simple leaky integrator with the ROSA copy signal on a small, clean conversational dataset. It demonstrates that a pure linear RNN can learn to produce coherent multi‑turn dialogue without any attention mechanisms. Feedback and contributions are welcome!

Runs of FlameF0X FWKV-ROSA on huggingface.co

109
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About FWKV-ROSA huggingface.co Model

More FWKV-ROSA license Visit here:

https://choosealicense.com/licenses/apache-2.0

FWKV-ROSA huggingface.co

FWKV-ROSA huggingface.co is an AI model on huggingface.co that provides FWKV-ROSA's model effect (), which can be used instantly with this FlameF0X FWKV-ROSA model. huggingface.co supports a free trial of the FWKV-ROSA model, and also provides paid use of the FWKV-ROSA. Support call FWKV-ROSA model through api, including Node.js, Python, http.

FlameF0X FWKV-ROSA online free

FWKV-ROSA huggingface.co is an online trial and call api platform, which integrates FWKV-ROSA's modeling effects, including api services, and provides a free online trial of FWKV-ROSA, you can try FWKV-ROSA online for free by clicking the link below.

FlameF0X FWKV-ROSA online free url in huggingface.co:

https://huggingface.co/FlameF0X/FWKV-ROSA

FWKV-ROSA install

FWKV-ROSA is an open source model from GitHub that offers a free installation service, and any user can find FWKV-ROSA on GitHub to install. At the same time, huggingface.co provides the effect of FWKV-ROSA install, users can directly use FWKV-ROSA installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

FWKV-ROSA install url in huggingface.co:

https://huggingface.co/FlameF0X/FWKV-ROSA

Url of FWKV-ROSA

FWKV-ROSA huggingface.co Url

Provider of FWKV-ROSA huggingface.co

FlameF0X
ORGANIZATIONS

Other API from FlameF0X

huggingface.co

Total runs: 164
Run Growth: 138
Growth Rate: 84.15%
Updated:July 17 2026
huggingface.co

Total runs: 144
Run Growth: 49
Growth Rate: 34.03%
Updated:May 26 2026
huggingface.co

Total runs: 114
Run Growth: -76
Growth Rate: -66.67%
Updated:December 01 2025
huggingface.co

Total runs: 112
Run Growth: 16
Growth Rate: 14.29%
Updated:October 31 2025
huggingface.co

Total runs: 70
Run Growth: 70
Growth Rate: 100.00%
Updated:December 03 2025
huggingface.co

Total runs: 62
Run Growth: 12
Growth Rate: 19.35%
Updated:November 29 2025
huggingface.co

Total runs: 50
Run Growth: 47
Growth Rate: 94.00%
Updated:June 29 2025
huggingface.co

Total runs: 44
Run Growth: 11
Growth Rate: 25.00%
Updated:November 25 2025
huggingface.co

Total runs: 39
Run Growth: -135
Growth Rate: -346.15%
Updated:October 23 2025
huggingface.co

Total runs: 36
Run Growth: 20
Growth Rate: 55.56%
Updated:February 23 2026
huggingface.co

Total runs: 30
Run Growth: -254
Growth Rate: -846.67%
Updated:May 20 2026
huggingface.co

Total runs: 21
Run Growth: -10
Growth Rate: -47.62%
Updated:October 17 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:December 19 2025
huggingface.co

Total runs: 14
Run Growth: -272
Growth Rate: -1942.86%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 06 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 29 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 31 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 17 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 07 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 15 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 31 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 19 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 25 2026