High-Throughput Asynchronous Reinforcement Learning from Human Feedback (RLHF) with In-VRAM Tensor-Native Rewards & Second-Moment Off-Policy Control (M2PO / GRPO)
Modern Reinforcement Learning from Human Feedback (RLHF) for Large Language Models (LLMs)—including Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO)—faces two critical engineering bottlenecks:
The CPU-GPU Memory Wall (SerDes Overhead)
:
Rollout generates token sequences on the GPU. Standard reward computation then:
Copies generated token IDs across the PCIe bus to CPU memory (
.cpu()
).
Decodes IDs into UTF-8 strings (
tokenizer.decode()
).
Runs Python string matching, regular expressions, or rule-based scoring on the host CPU.
Converts scalar scores back into PyTorch tensors and transfers them across PCIe back into GPU memory (
.cuda()
).
In high-throughput generation regimes (batch size ≥ 64, sequence length ≥ 1024), CPU serialization and PCIe roundtrips introduce severe throughput degradation, consuming up to 30–50% of the entire pipeline duration.
The Synchronous Lockstep Barrier (GPU Underutilization)
:
In synchronous PPO, rollout generation and trainer parameter optimization run in strict lockstep:
While the trainer runs backpropagation, rollout GPU workers sit completely idle. Conversely, while rollout workers generate tokens autoregressively, training GPUs idle waiting for batches. This lockstep barrier causes severe GPU idle time ("bubble overhead"), frequently exceeding 40–60% of total cluster compute time.
AsyncTensorRLHF
eliminates both synchronization bottlenecks through architectural disaggregation:
Zero-Copy In-VRAM Tensor-Native Rewards
:
All reward computations are executed entirely within GPU memory on
torch.Tensor
structures using parallel 1D sliding-window convolutions (
.unfold()
) and tensor operations. No CPU string decoding, no UTF-8 serialization, and zero host-device bus transfers occur during reward assignment.
Asynchronous Continuous Rollout with Second-Moment Staleness Control (M2PO)
:
Rollout workers continuously generate responses into a non-blocking, thread-safe experience replay buffer. The trainer continuously samples from the buffer and optimizes the policy. To handle the resulting off-policy divergence $\theta - \theta_{\text{old}}$, the framework incorporates:
Dynamic staleness eviction: Experiences with age $\tau = v_{\text{current}} - v_{\text{data}} \gt \tau_{\text{max}}$ are immediately discarded.
M2PO Second-Moment Trust Region Loss: Dynamically bounds the second moment of the importance weight $\mathbb{E}[r(\theta)^2]$, preventing policy collapse under asynchronous drift.
Group-Aware Buffers for GRPO: Standardizes advantage estimates across groups of candidate generations per prompt.
3.1 Policy Gradient under Asynchronous Staleness (τ)
In a distributed asynchronous RLHF pipeline, an experience tuple $(x, y, r, \log \pi_{\theta_{\text{old}}}(y \mid x))$ collected at policy version $\theta_{\text{old}}$ is consumed by the trainer at parameter version $\theta_{\text{current}}$, where staleness is defined as:
τ
=
version
(
θ
current
)
−
version
(
θ
old
)
≥
0
The policy gradient under importance sampling is:
g
(
θ
)
=
E
(
x
,
y
)
∼
D
[
π
θ
old
(
y
∣
x
)
∇
θ
π
θ
(
y
∣
x
)
A
π
θ
old
(
x
,
y
)
]
When staleness $\tau \gt 0$, the importance sampling weight $r_t(\theta) = \frac{\pi_\theta(y_t \mid x, y_{\lt t})}{\pi_{\theta_{\text{old}}}(y_t \mid x, y_{\lt t})}$ exhibits high variance:
Var
y
∼
π
θ
old
[
r
t
(
θ
)]
≈
exp
(
D
χ
2
(
π
θ
∥
π
θ
old
)
)
−
1
If $\tau$ grows without constraint, standard PPO clipping $\text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon)$ saturates, causing vanishing gradient updates on fresh tokens and destructive updates on stale outliers.
3.2 Proximal Policy Optimization (PPO)
AsyncTensorRLHF implements clipped PPO with per-token importance weighting:
L
PPO
(
θ
)
=
−
B
⋅
L
1
b
=
1
∑
B
t
=
1
∑
L
min
(
r
b
,
t
(
θ
)
A
b
,
t
,
clip
(
r
b
,
t
(
θ
)
,
1
−
ϵ
,
1
+
ϵ
)
A
b
,
t
)
where the per-token importance weight ratio is:
r
b
,
t
(
θ
)
=
exp
(
lo
g
π
θ
(
y
b
,
t
∣
x
b
,
y
b
,
<
t
)
−
lo
g
π
θ
old
(
y
b
,
t
∣
x
b
,
y
b
,
<
t
)
)
3.3 Second-Moment Trust Region Optimization (M2PO)
To guarantee stability under asynchronous rollout where $\tau \in [1, \tau_{\text{max}}]$, AsyncTensorRLHF incorporates
M2PO
(Second-Moment Trust Region Policy Optimization). M2PO constrains the empirical second moment of the importance weight:
M
2
=
B
⋅
L
1
b
=
1
∑
B
t
=
1
∑
L
r
b
,
t
(
θ
)
2
Tokens whose importance weight violates the second-moment threshold $r_{b,t}(\theta)^2 \ge \gamma_{\text{threshold}}$ are masked:
m
b
,
t
=
I
(
r
b
,
t
(
θ
)
2
<
γ
threshold
)
L
M2PO
(
θ
)
=
−
max
(
1
,
∑
b
=
1
B
∑
t
=
1
L
m
b
,
t
)
∑
b
=
1
B
∑
t
=
1
L
m
b
,
t
⋅
min
(
r
b
,
t
(
θ
)
A
b
,
t
,
clip
(
r
b
,
t
(
θ
)
,
1
−
ϵ
,
1
+
ϵ
)
A
b
,
t
)
This eliminates destructive gradient spikes caused by stale off-policy rollouts without stalling generation.
3.4 Group Relative Policy Optimization (GRPO)
For mathematical, programmatic, and structured reasoning tasks (e.g. DeepSeek-Math, DeepSeek-R1), AsyncTensorRLHF implements
GRPO
. GRPO foregoes a learned critic model and instead normalizes advantages within a group of $G$ responses generated for the identical prompt $x$:
μ
g
=
G
1
i
=
1
∑
G
R
g
,
i
,
σ
g
=
G
1
i
=
1
∑
G
(
R
g
,
i
−
μ
g
)
2
+
ϵ
eps
A
^
g
,
i
=
σ
g
R
g
,
i
−
μ
g
The GRPO objective is:
L
GRPO
(
θ
)
=
−
B
⋅
G
⋅
L
1
b
=
1
∑
B
i
=
1
∑
G
t
=
1
∑
L
min
(
r
b
,
i
,
t
(
θ
)
A
^
b
,
i
,
clip
(
r
b
,
i
,
t
(
θ
)
,
1
−
ϵ
,
1
+
ϵ
)
A
^
b
,
i
)
When all responses in a group receive identical rewards (e.g., all correct $R_i=1$ or all wrong $R_i=0$), $\sigma_g \to 0$. AsyncTensorRLHF's implementation adds numerical smoothing ($\epsilon = 10^{-8}$) to ensure $\hat{A}_{g,i} \to 0$ without
NaN
or
Inf
divergence.
The reward engine executes token-level subsequence matching completely on GPU tensors:
deftensor_native_reward(
generated_ids: torch.Tensor, answer_patterns: List[torch.Tensor], eos_token_id: int, device: str = "cpu",
) -> torch.Tensor:
# 1. Trims sequences at first EOS occurrence using argmax over mask
eos_mask = generated_ids == eos_token_id
first_eos = torch.where(
eos_mask.any(dim=1),
eos_mask.int().argmax(dim=1),
torch.full((B,), L, device=device_obj, dtype=torch.long),
)
# 2. Extracts sliding window views via .unfold(dimension, size, step)for i inrange(B):
seq = generated_ids[i, : first_eos[i]]
pattern = answer_patterns[i]
pat_len = pattern.shape[0]
if pat_len == 0or pat_len > seq.shape[0]:
continue
windows = seq.unfold(0, pat_len, 1)
match = (windows == pattern).all(dim=1).any()
rewards[i] = 1.0ifmatchelse0.0return rewards
Key Optimizations
:
Zero memory allocation for substrings;
.unfold()
creates lightweight strided tensor views.
Truncates sequences at the exact first EOS boundary, ignoring post-EOS artifact tokens.
Fully compatible with
torch.compile(mode="reduce-overhead")
for kernel fusion.
5.2 Experience Replay Subsystems (
src/buffer/
)
BoundedReplayBuffer
: Thread-safe FIFO queue backed by
queue.Queue
with non-blocking
.push(exp)
and
.sample(batch_size)
. Automatically evicts the oldest item when capacity is exceeded.
VersionedReplayBuffer
: Tracks policy staleness. When an experience is pushed:
if exp.policy_version < self.current_version - self.max_staleness:
return# Stale: silently evicted without wasting trainer compute
GroupAwareReplayBuffer
: Buffers $G$ completions per prompt ID. When the $G$-th completion arrives, advantages are computed in-place and the atomic
GroupBufferEntry
is moved to the ready queue.
5.3 Asynchronous Rollout Engines (
src/rollout/
)
StubEngine
: Pure-Python mock generating uniform tokens with
asyncio.sleep(0.001)
to test concurrent coroutine interleaving without GPU dependencies.
HFEngine
: Real autoregressive PyTorch/Transformers engine running directly on CUDA GPUs. Samples token distributions via
torch.multinomial
and extracts exact log-probabilities in a single forward pass.
VLLMEngineWrapper
: Production adapter. Automatically routes generation requests to
vllm.AsyncLLMEngine
if installed; otherwise falls back gracefully to
HFEngine
or
StubEngine
.
5.4 Distributed Trainer Workers (
src/trainer/
)
TrainerWorker
automatically detects available hardware:
Dynamically allocates models on
cuda
when
torch.cuda.is_available()
is True; falls back cleanly to
cpu
.
Implements
step(batch_size)
: samples experiences from the shared buffer actor, constructs padded policy and advantage tensors, computes PPO/M2PO/GRPO loss, backpropagates gradients, and calls
optimizer.step()
.
5.5 Orchestration & Version Management (
src/orchestrator/
)
VersionManager
: Thread-safe policy version counter with atomic
.bump()
and
.staleness(exp_version)
.
Orchestrator
: Asynchronous control loop that dispatches prompts to rollout workers, steps trainer actors, and periodically triggers
weight_sync_fn(new_version)
to push updated weights to inference workers.
6. Installation & Environment Setup
Requirements
Python 3.10, 3.11, 3.12, 3.13, or 3.14
PyTorch $\ge 2.2.0$ (CUDA 12.1+ recommended for GPU)
NumPy, PyYAML, PyTest
Clone and Install
git clone https://github.com/Hooshaai/AsyncTensorRLHF.git
cd AsyncTensorRLHF
pip install -r requirements.txt
7. Verification & Benchmarking
7.1 Running the 41-Test Comprehensive Suite
Run all unit, edge-case, and end-to-end integration tests:
Q: Can I run AsyncTensorRLHF without Ray?
A:
Yes. All components (
RolloutWorker
,
TrainerWorker
,
PromptQueue
,
ReplayBufferActor
) automatically detect if Ray is installed. When Ray is absent, they execute as standard high-performance Python classes using
asyncio
and
threading
.
Q: Does it work on single-GPU or laptop setups?
A:
Yes. The framework was benchmarked and validated on a single NVIDIA GeForce RTX 4070 Laptop GPU running Windows 11 with PyTorch 2.6.0+cu124, achieving a complete pipeline footprint of just 17 MB VRAM.
Q: What happens if vLLM is not installed?
A:
VLLMEngineWrapper
automatically falls back to
HFEngine
(which runs autoregressive inference using native PyTorch/Transformers on CUDA or CPU) or
StubEngine
(for testing).
Q: How does M2PO prevent training collapse with stale data?
A:
Stale data produces outlier importance ratios $r_t(\theta) \gg 1$. M2PO tracks the second moment $\mathbb{E}[r_t(\theta)^2]$ across tokens and masks out elements exceeding the
m2_threshold
, bounding gradient variance.
11. Research Paper & BibTeX Citation
A complete 6-page research paper detailing the theory, algorithm, proofs, and empirical evaluation of AsyncTensorRLHF is available:
If you use AsyncTensorRLHF or our benchmarks in your academic research or production deployment, please cite:
@article{majlesi2026asynctensorrlhf,
title = {AsyncTensorRLHF: High-Throughput Asynchronous RLHF with In-VRAM Tensor-Native Rewards},
author = {Majlesi, Taha},
journal = {arXiv preprint arXiv:2603.XXXXX},
year = {2026},
url = {https://github.com/Hooshaai/AsyncTensorRLHF}
}
12. License
This project is licensed under the
Apache License, Version 2.0
. You may freely use, modify, distribute, and commercialize this software according to the terms specified in the
LICENSE
file.
Runs of tahamajs AsyncTensorRLHF on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About AsyncTensorRLHF huggingface.co Model
AsyncTensorRLHF huggingface.co is an AI model on huggingface.co that provides AsyncTensorRLHF's model effect (), which can be used instantly with this tahamajs AsyncTensorRLHF model. huggingface.co supports a free trial of the AsyncTensorRLHF model, and also provides paid use of the AsyncTensorRLHF. Support call AsyncTensorRLHF model through api, including Node.js, Python, http.
AsyncTensorRLHF huggingface.co is an online trial and call api platform, which integrates AsyncTensorRLHF's modeling effects, including api services, and provides a free online trial of AsyncTensorRLHF, you can try AsyncTensorRLHF online for free by clicking the link below.
tahamajs AsyncTensorRLHF online free url in huggingface.co:
AsyncTensorRLHF is an open source model from GitHub that offers a free installation service, and any user can find AsyncTensorRLHF on GitHub to install. At the same time, huggingface.co provides the effect of AsyncTensorRLHF install, users can directly use AsyncTensorRLHF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.