Lambent / Qwen3-4B-Base-Continued-GRPO-B

huggingface.co
Total runs: 9
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: January 01 2026

Introduction of Qwen3-4B-Base-Continued-GRPO-B

Model Details of Qwen3-4B-Base-Continued-GRPO-B

Experimental GRPO "continued pretraining" - rewarding the model for completions that resembled the target data (in varying complex ways). Reward is calculated differently for creative text and code.

Trained for 1034 steps on a 3090, rank 128 QLoRA with alpha 256. Learning rate 1e-6 seemed ideal.

For this one, added LLM as judge to the reward functions; gpt-4o-mini for decent reward with logprobs for continuous and gemini-3-flash-preview for difficult-to-fool binary bonus:

class RewardModel:
    """Multi-domain reward model with LLM judge for creative domains."""

    def __init__(self, device: str = "cuda"):
        self.device = device
        print("Loading reward model components...")
        self.semantic_model = SentenceTransformer('all-MiniLM-L6-v2', device=device)
        self.chrf = CHRF(word_order=2)
        self.rouge = rouge_scorer.RougeScorer(['rougeL'], use_stemmer=True)

    def compute_reward(
        self,
        prediction: str,
        reference: str,
        reward_type: str,
        prefix: str = None,
    ) -> float:
        if not prediction or not reference:
            return 0.0

        if reward_type == "llm_judge":
            # Base reward (GT-anchored)
            embs = self.semantic_model.encode([reference, prediction], convert_to_tensor=True)
            sem_score = torch.nn.functional.cosine_similarity(embs[0:1], embs[1:2]).item()
            chrf_score = self.chrf.sentence_score(prediction, [reference]).score / 100.0
            base_reward = 0.4 * sem_score + 0.6 * chrf_score

            # LLM judge bonus
            if prefix is not None:
                llm_reward, flash_bonus = get_llm_judge_reward(prefix, prediction, reference)
            else:
                llm_reward, flash_bonus = 0.5, 0.0

            # Multiplicative: base * (1 + llm_bonus + flash_bonus)
            # This ensures LLM bonus scales with GT similarity
            return base_reward * (1 + 0.3 * llm_reward + 0.2 * flash_bonus)

        elif reward_type == "creative":
            # Original creative reward (fallback)
            embs = self.semantic_model.encode([reference, prediction], convert_to_tensor=True)
            sem_score = torch.nn.functional.cosine_similarity(embs[0:1], embs[1:2]).item()
            chrf_score = self.chrf.sentence_score(prediction, [reference]).score / 100.0
            return 0.4 * sem_score + 0.6 * chrf_score

        elif reward_type == "hybrid":
            embs = self.semantic_model.encode([reference, prediction], convert_to_tensor=True)
            sem_score = torch.nn.functional.cosine_similarity(embs[0:1], embs[1:2]).item()
            rouge_result = self.rouge.score(reference, prediction)
            rouge_l = rouge_result['rougeL'].fmeasure

            ref_len = len(reference.split())
            pred_len = len(prediction.split())
            if ref_len > 0:
                len_ratio = pred_len / ref_len
                length_penalty = max(0.0, 1.0 - abs(1.0 - len_ratio))
            else:
                length_penalty = 1.0 if pred_len == 0 else 0.0

            return 0.6 * sem_score + 0.3 * rouge_l + 0.1 * length_penalty

        elif reward_type == "levenshtein":
            max_len = max(len(prediction), len(reference))
            if max_len == 0:
                return 1.0
            dist = Levenshtein.distance(prediction, reference)
            return max(0.0, 1.0 - (dist / max_len))

        else:
            raise ValueError(f"Unknown reward type: {reward_type}")

Quick capability diagnostics compared to original (most rapid subset to test both on):

Task Metric Base Trained Delta
arc_easy acc 0.7891 0.7925 +0.43%
arc_easy acc_norm 0.7609 0.7647 +0.50%
lambada_openai acc 0.6912 0.6971 +0.85%
lambada_openai perplexity 4.2433 4.0663 -4.2% ↓
openbookqa acc 0.3160 0.3180 +0.63%
openbookqa acc_norm 0.4100 0.4080 -0.49%
piqa acc 0.7797 0.7824 +0.35%
piqa acc_norm 0.7807 0.7797 -0.13%

Runs of Lambent Qwen3-4B-Base-Continued-GRPO-B on huggingface.co

9
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About Qwen3-4B-Base-Continued-GRPO-B huggingface.co Model

More Qwen3-4B-Base-Continued-GRPO-B license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-4B-Base-Continued-GRPO-B huggingface.co

Qwen3-4B-Base-Continued-GRPO-B huggingface.co is an AI model on huggingface.co that provides Qwen3-4B-Base-Continued-GRPO-B's model effect (), which can be used instantly with this Lambent Qwen3-4B-Base-Continued-GRPO-B model. huggingface.co supports a free trial of the Qwen3-4B-Base-Continued-GRPO-B model, and also provides paid use of the Qwen3-4B-Base-Continued-GRPO-B. Support call Qwen3-4B-Base-Continued-GRPO-B model through api, including Node.js, Python, http.

Qwen3-4B-Base-Continued-GRPO-B huggingface.co Url

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-B

Lambent Qwen3-4B-Base-Continued-GRPO-B online free

Qwen3-4B-Base-Continued-GRPO-B huggingface.co is an online trial and call api platform, which integrates Qwen3-4B-Base-Continued-GRPO-B's modeling effects, including api services, and provides a free online trial of Qwen3-4B-Base-Continued-GRPO-B, you can try Qwen3-4B-Base-Continued-GRPO-B online for free by clicking the link below.

Lambent Qwen3-4B-Base-Continued-GRPO-B online free url in huggingface.co:

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-B

Qwen3-4B-Base-Continued-GRPO-B install

Qwen3-4B-Base-Continued-GRPO-B is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-4B-Base-Continued-GRPO-B on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-4B-Base-Continued-GRPO-B install, users can directly use Qwen3-4B-Base-Continued-GRPO-B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-4B-Base-Continued-GRPO-B install url in huggingface.co:

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-B

Url of Qwen3-4B-Base-Continued-GRPO-B

Qwen3-4B-Base-Continued-GRPO-B huggingface.co Url

Provider of Qwen3-4B-Base-Continued-GRPO-B huggingface.co

Lambent
ORGANIZATIONS

Other API from Lambent

huggingface.co

Total runs: 40
Run Growth: 18
Growth Rate: 45.00%
Updated:March 17 2026
huggingface.co

Total runs: 26
Run Growth: 12
Growth Rate: 46.15%
Updated:March 26 2026
huggingface.co

Total runs: 12
Run Growth: -16
Growth Rate: -133.33%
Updated:March 09 2026