If we only trained agents through imitation - copying what humans demonstrated - they would inherit all the limitations of those demonstrations. Agents would never improve beyond the ceiling of the training data. But the real world demands more. An agent navigating a 50-step workflow needs feedback that guides it toward success, not just copies of averagely-good traces.
This is where Reinforcement Learning (RL) comes in. Unlike supervised learning which asks "can you copy this?", reinforcement learning asks "can you maximize success over time?" This fundamental shift unlocks agents that continuously improve.
Code repo: RL_Agents
Notebook: post01_rlhf_rlvr_foundations.ipynb
Consider training a software development agent. You could collect 10,000 transcripts of expert programmers and have your agent learn to copy their style. But here's the problem: once your agent encounters a situation slightly different from those transcripts, it has no principled way to decide what to do. It learned patterns, not principles. And critically, it never learned from failure, because the training data only contained successful examples (or at least, the annotator's best judgment of what successful looks like).
For short, well-defined tasks this might work. For long workflows with many decisions, it breaks down immediately. The agent needs outcome feedback - a signal that says "that sequence of actions worked, or it didn't" - so it can propagate learning from failures and reinforce successes.
RLHF (Reinforcement Learning from Human Feedback) is the technique that gave modern large language models (LLMs) their instruction-following ability. It emerged as the key breakthrough for making GPT-style models useful for real-world tasks. Here's the full pipeline:
Take a base model (usually just supervised fine-tuned on conversations). Generate two candidate responses to the same prompt. Show both to human raters and ask: "Which response is better?" Better could mean more helpful, more honest, safer, more detailed - whatever your values are.
Collect thousands of these pairwise comparisons. Each comparison is a data point: prompt → (response A, response B) → "A is better" (or "B is better", or "tie").
Now you have a classification dataset. Use it to train a reward model (RM) - a neural network that learns to predict what humans prefer. Feed it a prompt + response, it outputs a scalar reward score. This model essentially learns to approximate the implicit human preference function.
The code below simulates the same three RLHF stages you would run in production, but with synthetic data so results are instant.
from __future__ import annotations
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
class PreferenceBandit:
"""
4-arm bandit representing four response-quality tiers.
arm 0 = low quality (true reward ≈ 0.20)
arm 1 = medium-low (true reward ≈ 0.45)
arm 2 = medium-high (true reward ≈ 0.72)
arm 3 = high quality (true reward ≈ 0.92)
Pulls return true_reward + Gaussian noise, so the reward model must
learn from imperfect signal — exactly as in real RLHF pipelines.
"""
TRUE_REWARDS = np.array([0.20, 0.45, 0.72, 0.92])
N_ARMS = 4
def __init__(self, noise: float = 0.05, seed: int = 7) -> None:
self.noise = noise
self.rng = np.random.default_rng(seed)
def pull(self, arm: int) -> float:
r = self.TRUE_REWARDS[arm] + self.rng.normal(0.0, self.noise)
return float(np.clip(r, 0.0, 1.0))
def compare(self, arm_a: int, arm_b: int) -> tuple[float, float, int]:
"""Pull both arms; return (reward_a, reward_b, preferred).
preferred=1 means arm_a was better — mirrors human-labelling."""
ra, rb = self.pull(arm_a), self.pull(arm_b)
return ra, rb, int(ra >= rb)
def collect_bandit_preferences(
bandit: PreferenceBandit,
n_pairs: int = 80,
seed: int = 3,
) -> pd.DataFrame:
"""Collect preference pairs by randomly sampling and comparing two arms."""
rng = np.random.default_rng(seed)
rows = []
for _ in range(n_pairs):
a, b = int(rng.choice(bandit.N_ARMS)), int(rng.choice(bandit.N_ARMS))
while b == a:
b = int(rng.choice(bandit.N_ARMS))
ra, rb, pref = bandit.compare(a, b)
rows.append({
"arm_a": a, "arm_b": b,
"reward_a": round(ra, 3), "reward_b": round(rb, 3),
"reward_diff": round(ra - rb, 3),
"preferred": pref,
})
return pd.DataFrame(rows)
def train_reward_model(df: pd.DataFrame) -> tuple[LogisticRegression, float]:
"""Fit logistic regression on preference pairs.
Feature: reward_diff. Label: preferred (1 = arm_a was better)."""
X = df[["reward_diff"]].values
y = df["preferred"].values
model = LogisticRegression(max_iter=300)
model.fit(X, y)
return model, float(model.score(X, y))
# Run the preference collection and RM training
bandit = PreferenceBandit()
df = collect_bandit_preferences(bandit, n_pairs=80)
rm, accuracy = train_reward_model(df)
print(f"Preference pairs collected: {len(df)}")
print(df.head())
print(f"\nRM Training Accuracy: {accuracy:.1%}")
The RM learns to score responses on whatever implicit quality signal the human raters were using. In the example above, it learns that responses with higher quality differences should be preferred. In a real system, this is trained on tens of thousands of examples from actual human raters.
Now use this learned reward model to fine-tune your original base model via PPO (Proximal Policy Optimization) - a type of reinforcement learning algorithm. The training loop is:
The result: your model learns to generate responses that the RM predicts are high-quality. Since the RM was trained to predict human preference, the model becomes better at following human intentions.
RLHF was a massive breakthrough. It's the foundation of ChatGPT, Claude, and other instruction-tuned LLMs. But it hits hard walls for agentic tasks:
Imagine training an agent to debug software at scale using RLHF. You'd need to hire hundreds of experienced engineers to judge whether the agent's debugging approach was correct on each attempt. The cost is astronomical. The inconsistency is terrible. It's simply not feasible.
RLVR (Reinforcement Learning on Verifiable Rewards) removes the human-in-the-loop bottleneck entirely. Instead of asking humans "which response is better?", you ask verifiers "did this response pass an objective test?"
For many agent tasks, correctness is objective:
The verifier is a simple program - a test suite, a validator function, a diff checker - not a human. And here's the magic: you can generate unlimited amounts of supervision by simply creating more problems programmatically.
Here's the complete RLHF training loop — the policy learns to pick high-quality responses guided by the reward model:
class SoftmaxPolicy:
"""
Softmax policy over N arms. pi(a) = softmax(logits)[a]
Updated via REINFORCE: logits += lr * (R - baseline) * (1[a] - pi)
"""
def __init__(self, n_arms: int, lr: float = 0.08, seed: int = 42) -> None:
self.logits = np.zeros(n_arms)
self.lr = lr
self.rng = np.random.default_rng(seed)
def probs(self) -> np.ndarray:
e = np.exp(self.logits - self.logits.max())
return e / e.sum()
def sample(self) -> int:
return int(self.rng.choice(len(self.logits), p=self.probs()))
def update(self, arm: int, reward: float, baseline: float = 0.0) -> None:
"""Single REINFORCE gradient step."""
advantage = reward - baseline
grad = -self.probs()
grad[arm] += 1.0 # nabla log pi(arm)
self.logits += self.lr * advantage * grad
def train_rlhf_loop(
bandit: PreferenceBandit,
reward_model: LogisticRegression,
n_episodes: int = 300,
seed: int = 1,
) -> pd.DataFrame:
"""Full REINFORCE loop using reward-model scores as the training signal.
Episode: sample arm -> pull bandit -> RM scores pull -> REINFORCE update."""
policy = SoftmaxPolicy(n_arms=bandit.N_ARMS, lr=0.08, seed=seed)
baseline = 0.5
alpha_b = 0.05
history = []
for ep in range(n_episodes):
arm = policy.sample()
raw = bandit.pull(arm)
mean_r = float(bandit.TRUE_REWARDS.mean())
diff = np.array([[raw - mean_r]])
rm_score = float(reward_model.predict_proba(diff)[0, 1])
policy.update(arm, rm_score, baseline)
baseline += alpha_b * (rm_score - baseline)
history.append({
"episode": ep + 1,
"arm": arm,
"raw_reward": round(raw, 3),
"rm_score": round(rm_score, 3),
"baseline": round(baseline, 3),
"best_arm_prob": round(float(policy.probs()[-1]), 3),
})
return pd.DataFrame(history)
# Run the full RLHF training loop
history_df = train_rlhf_loop(bandit, rm, n_episodes=300)
print(f"Final best-arm probability: {history_df['best_arm_prob'].iloc[-1]:.1%}")
print(history_df.tail(5)[["episode", "arm", "rm_score", "best_arm_prob"]])
In this example, all four verifiers pass, giving rewards of 1 each. In a real training loop, you'd generate many problems (1000s or 10,000s), run the agent on each, collect verifier signals, and accumulate a reward dataset. Then use GRPO or PPO to fine-tune the model.
This is exactly what powered DeepSeek-R1, OpenAI's o1, and similar reasoning-focused models. The models learned to generate chain-of-thought reasoning, check the answer, and propagate the reward signal backward through all the reasoning steps. The result: agents that can reason through multi-step problems effectively.
RLVR is transformative, but it requires objective correctness checks. What if your task doesn't have a clear verifier?
For these tasks, RLHF remains necessary - you need human feedback. But the more you can decompose your agent's task into verifiable subtasks, the more you can use RLVR instead. This becomes a key design pattern: structure agent tasks to maximize the verifiable components.
| Dimension | RLHF | RLVR |
|---|---|---|
| Reward source | Human raters | Automated verifier |
| Scaling cost | Linear with volume (hire more raters) | Sublinear (write problems, run verifier) |
| Consistency | Variable (humans disagree) | Perfect (verifier is deterministic) |
| Applicable domains | Any task | Tasks with objective correctness |
| Training speed | Slow (waiting for human annotations) | Fast (automated rewards) |
The shift from RLHF to RLVR is reshaping the entire field. Here's why:
The frontier agentic models being built now use RLVR as a core component. Where human feedback is needed (for tasks without clear verifiers), they layer in RLAIF (Reinforcement Learning from AI Feedback - using a strong model to rate outputs instead of humans), which we'll cover later.
To recap: RLHF was the breakthrough that made LLMs instruction-following. But for agents performing long-horizon, verifiable tasks, RLVR is the better approach. It removes the human bottleneck and enables scalable training with objective reward signals.
In the next post, we'll dive deep into the verifier design patterns used in practice - how to build verifiers for code, math, SQL, and tool use. We'll also explore what happens when verifiers are imperfect, and how to handle edge cases.
All code examples above are from the RL_Agents repository:
To run locally:
# Clone the repo
git clone https://github.com/Pulkit12dhingra/RL_Agents
cd RL_Agents
# Sync dependencies with uv
uv sync
# Run the notebook
jupyter notebook notebooks/post01_rlhf_rlvr_foundations.ipynb