emese-tech / folyo

huggingface.co
Total runs: 343
24-hour runs: 0
7-day runs: 17
30-day runs: 343
Model's Last Updated: September 07 2026
text-generation

Introduction of folyo

Model Details of folyo

Emese-Folyó (22B)

Folyó ("river") is the flagship, server-class tier of the Emese Hungarian model family — EuroLLM-22B continued-pretrained on Hungarian, instruction-tuned, and DPO-aligned.

Parameters 22B
Base EuroLLM-22B
Architecture LLaMA-style (RoPE θ=1,000,000, GQA 48Q/8KV, SwiGLU, RMSNorm)
Hidden / layers / heads 6144 / 54 / 48 (8 KV heads)
Vocabulary 128,000 (EuroLLM multilingual SentencePiece)
Max context length 32,768 tokens (EuroLLM-22B's native max_position_embeddings )
Precision bfloat16 (this repo, dequantized for transformers compatibility); MLX q8 also available ( folyo-mlx/ )
License Apache-2.0
Formats in this release
Folder Format Size Notes
folyo/ (this repo) bf16, standard HF safetensors ~42 GB loads directly with transformers and mlx_lm ; dequantized from the q8-native training chain (near-lossless)
folyo-mlx/ MLX q8 ~22 GB mlx_lm -only, the native training/serving precision for this model
Usage (transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("emese-tech/folyo")
model = AutoModelForCausalLM.from_pretrained("emese-tech/folyo", dtype=torch.bfloat16, device_map="auto")
msgs = [{"role": "user", "content": "Mi Magyarország fővárosa?"}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=256, do_sample=True, temperature=0.2, eos_token_id=[2, 4])
print(tok.decode(out[0][ids.input_ids.shape[1]:], skip_special_tokens=True))
Usage (MLX)
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tok = load("emese-tech/folyo-mlx")  # q8
p = tok.apply_chat_template([{"role": "user", "content": "Mi Magyarország fővárosa?"}],
                            tokenize=False, add_generation_prompt=True)
print(generate(model, tok, prompt=p, max_tokens=256, sampler=make_sampler(temp=0.2)))

Decode: temperature 0.2 , no repetition penalty, eos {2, 4} ( </s> and <|im_end|> — omitting id 4 causes non-stopping generation), ChatML template. For multi-turn conversations, always pass the full history , not just the latest message.

Training
  • CPT — q8-native continued pretraining of EuroLLM-22B on the Emese Hungarian corpus: ~6M tokens over 6,000 iterations , LoRA rank 64, then fused and requantized to q8 (rather than kept in bf16) to avoid a q4→bf16 fuse-precision mismatch discovered earlier in this project's history. The resulting q8-native CPT base ( eurollm-22b-cpt-v2-q8 ) is what every downstream SFT/DPO stage built on.
  • SFT — 1 epoch (4,914 iterations) on the instruct_v18b corpus (4,914 rows: persona, safety, code + code-debug, hedging/anti-confabulation, multi-step reasoning, compound constraints, multi-turn refinement, anti-repetition — same corpus as Patak/Csermely). LoRA rank 16 / scale 16, dropout 0.1, lr 5e-6, bottom 14 of 54 layers frozen ( num_layers: 40 ). Checkpoint-swept across the run; iteration 3,600 (not the final iteration) was selected as the best-behaving checkpoint.
  • DPO — 120 iterations of DPO-lite ("alfa") on 36 hand-written preference pairs (persona identity-defense + anti-repetition — same pair bank as Patak/Csermely). LoRA rank 16 / scale 16, lr 2e-6 — the lowest DPO learning rate in the family, since this checkpoint was already the peak-swept best and had the least margin to spare.
Benchmarks
Ultimate Bench (0-250) BlindSpot Bench (0-376)
This release (DPO alfa) 211/250 (84%) 310/376 — family-best BlindSpot score

emese-bench v1 (500 pts, consolidated Ultimate+BlindSpot, MLX q8): 410/500 (82%) — the current unified reference benchmark going forward, just behind Patak (413/500) and far ahead of Csermely (211/500). Near-perfect on safety, safety_edge, honesty_calibration, persona_selfaware, reading, translation, and code-writing, with no repetition loops or stop-token leakage anywhere in the transcript. Weakest on multi-step math, logic puzzles, and two confidently-fabricated Hungarian scientist biographies; the standout chat-session finding is a perfect 10/10 on both in-context memory recall and roleplay (best in the family so far), offset by a surprising 1/10 on the English-chat session — the model answered every turn in Hungarian despite the user writing in English. See emese-bench/results/folyo-mlx.md and emese-bench/README.md for full category-level detail.

Limitations
  • Can hallucinate specific facts (dates, attributions) — verify critical details.
  • Hungarian-first; other-language quality inherited from EuroLLM-22B.
  • Large — 42 GB bf16 / 22 GB q8; needs a machine with enough RAM/VRAM for practical serving.
  • Do not convert this model to GGUF without re-validating first — this project's own testing found a consistent ~14-27 point Ultimate Bench regression when serving via llama.cpp/GGUF (tested on Patak, same architecture family), regardless of source precision or quantization level, root-caused to llama.cpp's inference path itself. Use the transformers / mlx_lm paths above.

Runs of emese-tech folyo on huggingface.co

343
Total runs
0
24-hour runs
2
3-day runs
17
7-day runs
343
30-day runs

More Information About folyo huggingface.co Model

folyo huggingface.co

folyo huggingface.co is an AI model on huggingface.co that provides folyo's model effect (), which can be used instantly with this emese-tech folyo model. huggingface.co supports a free trial of the folyo model, and also provides paid use of the folyo. Support call folyo model through api, including Node.js, Python, http.

emese-tech folyo online free

folyo huggingface.co is an online trial and call api platform, which integrates folyo's modeling effects, including api services, and provides a free online trial of folyo, you can try folyo online for free by clicking the link below.

emese-tech folyo online free url in huggingface.co:

https://huggingface.co/emese-tech/folyo

folyo install

folyo is an open source model from GitHub that offers a free installation service, and any user can find folyo on GitHub to install. At the same time, huggingface.co provides the effect of folyo install, users can directly use folyo installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

folyo install url in huggingface.co:

https://huggingface.co/emese-tech/folyo

Url of folyo

Provider of folyo huggingface.co

emese-tech
ORGANIZATIONS

Other API from emese-tech

huggingface.co

Total runs: 828
Run Growth: 820
Growth Rate: 99.03%
Updated:September 08 2026
huggingface.co

Total runs: 401
Run Growth: 401
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 369
Run Growth: 369
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 203
Run Growth: 203
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 194
Run Growth: 194
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 155
Run Growth: 155
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 100
Run Growth: 100
Growth Rate: 100.00%
Updated:September 07 2026
huggingface.co

Total runs: 86
Run Growth: 86
Growth Rate: 100.00%
Updated:September 07 2026