Lambent / Qwen3-4B-Base-Continued-GRPO-Wave

huggingface.co
Total runs: 30
24-hour runs: 0
7-day runs: 4
30-day runs: 4
Model's Last Updated: February 04 2026
text-generation

Introduction of Qwen3-4B-Base-Continued-GRPO-Wave

Model Details of Qwen3-4B-Base-Continued-GRPO-Wave

Tested out GRPO training on domain-specific adapters and then using a WAVE merge in a very ... uh, scattershot way.

It worked out apparently smarter than when I tried to trim down to the "best" adapters of the exploration, though, so scattershot it is.

I'm not sure if they're a "better" base model here; domain-wise they may have lost out slightly on other domains and improved mainly on Python code?

The lm-eval diagnostic tasks here look promising though.

Task Metric Qwen3-4B-Base GRPO-Merge Δ Base GRPO-Wave Δ Base Δ Merge
arc_easy acc 0.7891 0.7870 -0.27% 0.7912 +0.27% +0.53%
arc_easy acc_norm 0.7609 0.7605 -0.05% 0.7643 +0.45% +0.50%
lambada_openai acc 0.6912 0.6984 +1.04% 0.7006 +1.36% +0.31%
lambada_openai perplexity ↓ 4.2433 4.0490 -4.58% 3.9616 -6.64% -2.16%
openbookqa acc 0.3160 0.3180 +0.63% 0.3180 +0.63% ±0.00%
openbookqa acc_norm 0.4100 0.4120 +0.49% 0.4100 ±0.00% -0.49%
piqa acc 0.7797 0.7807 +0.13% 0.7813 +0.21% +0.08%
piqa acc_norm 0.7807 0.7807 ±0.00% 0.7813 +0.08% +0.08%

This is a merge of pre-trained language models created using mergekit .

Merge Details
Merge Method

This model was merged using the WAVE merge method using Qwen/Qwen3-4B-Base as a base.

Models Merged

The following models were included in the merge:

Configuration

The following YAML configuration was used to produce this model:

models:
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-python-creative
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-ao3-minilm
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-aware-5e6
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-aware-test
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-merge-llm-judge-ep2
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-bbc-qwen
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO+../rlvr-envs/grpo-ao3-qwen
  - model: Lambent/Qwen3-4B-Base-Continued-GRPO-Merge
merge_method: wave
base_model: Qwen/Qwen3-4B-Base
parameters:
  synergy: 0.5  # 0.0 to 1.0. Higher = keep more "controversial" high-variance parameters
  entropy: 0.1  # Adds slight noise to break ties/prevent overfitting
dtype: bfloat16
tokenizer_source: Lambent/Qwen3-4B-Base-Continued-GRPO-Merge

Runs of Lambent Qwen3-4B-Base-Continued-GRPO-Wave on huggingface.co

30
Total runs
0
24-hour runs
0
3-day runs
4
7-day runs
4
30-day runs

More Information About Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co Model

More Qwen3-4B-Base-Continued-GRPO-Wave license Visit here:

https://choosealicense.com/licenses/apache-2.0

Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co

Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co is an AI model on huggingface.co that provides Qwen3-4B-Base-Continued-GRPO-Wave's model effect (), which can be used instantly with this Lambent Qwen3-4B-Base-Continued-GRPO-Wave model. huggingface.co supports a free trial of the Qwen3-4B-Base-Continued-GRPO-Wave model, and also provides paid use of the Qwen3-4B-Base-Continued-GRPO-Wave. Support call Qwen3-4B-Base-Continued-GRPO-Wave model through api, including Node.js, Python, http.

Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co Url

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-Wave

Lambent Qwen3-4B-Base-Continued-GRPO-Wave online free

Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co is an online trial and call api platform, which integrates Qwen3-4B-Base-Continued-GRPO-Wave's modeling effects, including api services, and provides a free online trial of Qwen3-4B-Base-Continued-GRPO-Wave, you can try Qwen3-4B-Base-Continued-GRPO-Wave online for free by clicking the link below.

Lambent Qwen3-4B-Base-Continued-GRPO-Wave online free url in huggingface.co:

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-Wave

Qwen3-4B-Base-Continued-GRPO-Wave install

Qwen3-4B-Base-Continued-GRPO-Wave is an open source model from GitHub that offers a free installation service, and any user can find Qwen3-4B-Base-Continued-GRPO-Wave on GitHub to install. At the same time, huggingface.co provides the effect of Qwen3-4B-Base-Continued-GRPO-Wave install, users can directly use Qwen3-4B-Base-Continued-GRPO-Wave installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Qwen3-4B-Base-Continued-GRPO-Wave install url in huggingface.co:

https://huggingface.co/Lambent/Qwen3-4B-Base-Continued-GRPO-Wave

Url of Qwen3-4B-Base-Continued-GRPO-Wave

Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co Url

Provider of Qwen3-4B-Base-Continued-GRPO-Wave huggingface.co

Lambent
ORGANIZATIONS

Other API from Lambent

huggingface.co

Total runs: 40
Run Growth: 18
Growth Rate: 45.00%
Updated:March 17 2026
huggingface.co

Total runs: 26
Run Growth: 12
Growth Rate: 46.15%
Updated:March 26 2026
huggingface.co

Total runs: 19
Run Growth: -15
Growth Rate: -78.95%
Updated:March 07 2024
huggingface.co

Total runs: 12
Run Growth: -16
Growth Rate: -133.33%
Updated:March 09 2026