RL Agents Part 1: From RLHF to RLVR - The Foundations

← Back to Home

Series Overview: This is Part 1 of a deep dive into Reinforcement Learning for AI agents. We start with the foundations: RLHF and RLVR, then build toward process rewards, optimization techniques, online learning, and safety.

If we only trained agents through imitation - copying what humans demonstrated - they would inherit all the limitations of those demonstrations. Agents would never improve beyond the ceiling of the training data. But the real world demands more. An agent navigating a 50-step workflow needs feedback that guides it toward success, not just copies of averagely-good traces.

This is where Reinforcement Learning (RL) comes in. Unlike supervised learning which asks "can you copy this?", reinforcement learning asks "can you maximize success over time?" This fundamental shift unlocks agents that continuously improve.

Code repo: RL_Agents

Notebook: post01_rlhf_rlvr_foundations.ipynb

1. Why Imitation Alone Fails for Long-Horizon Tasks

Consider training a software development agent. You could collect 10,000 transcripts of expert programmers and have your agent learn to copy their style. But here's the problem: once your agent encounters a situation slightly different from those transcripts, it has no principled way to decide what to do. It learned patterns, not principles. And critically, it never learned from failure, because the training data only contained successful examples (or at least, the annotator's best judgment of what successful looks like).

For short, well-defined tasks this might work. For long workflows with many decisions, it breaks down immediately. The agent needs outcome feedback - a signal that says "that sequence of actions worked, or it didn't" - so it can propagate learning from failures and reinforce successes.

Real-world analogy: Teaching a support agent only from polished customer service transcripts teaches it to sound professional. But if you want it to actually resolve tickets faster, you need to measure resolution success and fine-tune the agent on transcripts that correlate with resolutions. That's reinforcement learning.

2. Enter RLHF: Reinforcement Learning from Human Feedback

RLHF (Reinforcement Learning from Human Feedback) is the technique that gave modern large language models (LLMs) their instruction-following ability. It emerged as the key breakthrough for making GPT-style models useful for real-world tasks. Here's the full pipeline:

Step 1: Collect Preference Data

Take a base model (usually just supervised fine-tuned on conversations). Generate two candidate responses to the same prompt. Show both to human raters and ask: "Which response is better?" Better could mean more helpful, more honest, safer, more detailed - whatever your values are.

Collect thousands of these pairwise comparisons. Each comparison is a data point: prompt → (response A, response B) → "A is better" (or "B is better", or "tie").

Step 2: Train a Reward Model

Now you have a classification dataset. Use it to train a reward model (RM) - a neural network that learns to predict what humans prefer. Feed it a prompt + response, it outputs a scalar reward score. This model essentially learns to approximate the implicit human preference function.

The code below simulates the same three RLHF stages you would run in production, but with synthetic data so results are instant.

from __future__ import annotations
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression


class PreferenceBandit:
    """
    4-arm bandit representing four response-quality tiers.
      arm 0 = low quality  (true reward ≈ 0.20)
      arm 1 = medium-low   (true reward ≈ 0.45)
      arm 2 = medium-high  (true reward ≈ 0.72)
      arm 3 = high quality (true reward ≈ 0.92)
    Pulls return true_reward + Gaussian noise, so the reward model must
    learn from imperfect signal — exactly as in real RLHF pipelines.
    """
    TRUE_REWARDS = np.array([0.20, 0.45, 0.72, 0.92])
    N_ARMS = 4

    def __init__(self, noise: float = 0.05, seed: int = 7) -> None:
        self.noise = noise
        self.rng = np.random.default_rng(seed)

    def pull(self, arm: int) -> float:
        r = self.TRUE_REWARDS[arm] + self.rng.normal(0.0, self.noise)
        return float(np.clip(r, 0.0, 1.0))

    def compare(self, arm_a: int, arm_b: int) -> tuple[float, float, int]:
        """Pull both arms; return (reward_a, reward_b, preferred).
        preferred=1 means arm_a was better — mirrors human-labelling."""
        ra, rb = self.pull(arm_a), self.pull(arm_b)
        return ra, rb, int(ra >= rb)


def collect_bandit_preferences(
    bandit: PreferenceBandit,
    n_pairs: int = 80,
    seed: int = 3,
) -> pd.DataFrame:
    """Collect preference pairs by randomly sampling and comparing two arms."""
    rng = np.random.default_rng(seed)
    rows = []
    for _ in range(n_pairs):
        a, b = int(rng.choice(bandit.N_ARMS)), int(rng.choice(bandit.N_ARMS))
        while b == a:
            b = int(rng.choice(bandit.N_ARMS))
        ra, rb, pref = bandit.compare(a, b)
        rows.append({
            "arm_a": a, "arm_b": b,
            "reward_a": round(ra, 3), "reward_b": round(rb, 3),
            "reward_diff": round(ra - rb, 3),
            "preferred": pref,
        })
    return pd.DataFrame(rows)


def train_reward_model(df: pd.DataFrame) -> tuple[LogisticRegression, float]:
    """Fit logistic regression on preference pairs.
    Feature: reward_diff.  Label: preferred (1 = arm_a was better)."""
    X = df[["reward_diff"]].values
    y = df["preferred"].values
    model = LogisticRegression(max_iter=300)
    model.fit(X, y)
    return model, float(model.score(X, y))


# Run the preference collection and RM training
bandit = PreferenceBandit()
df = collect_bandit_preferences(bandit, n_pairs=80)
rm, accuracy = train_reward_model(df)

print(f"Preference pairs collected: {len(df)}")
print(df.head())
print(f"\nRM Training Accuracy: {accuracy:.1%}")
  

The RM learns to score responses on whatever implicit quality signal the human raters were using. In the example above, it learns that responses with higher quality differences should be preferred. In a real system, this is trained on tens of thousands of examples from actual human raters.

Step 3: Fine-tune the Policy with PPO

Now use this learned reward model to fine-tune your original base model via PPO (Proximal Policy Optimization) - a type of reinforcement learning algorithm. The training loop is:

The result: your model learns to generate responses that the RM predicts are high-quality. Since the RM was trained to predict human preference, the model becomes better at following human intentions.

3. The Critical Limitations of RLHF

RLHF was a massive breakthrough. It's the foundation of ChatGPT, Claude, and other instruction-tuned LLMs. But it hits hard walls for agentic tasks:

Imagine training an agent to debug software at scale using RLHF. You'd need to hire hundreds of experienced engineers to judge whether the agent's debugging approach was correct on each attempt. The cost is astronomical. The inconsistency is terrible. It's simply not feasible.

Key insight: Human feedback is fundamentally limited by the rate at which humans can produce labels. You cannot scale a human-in-the-loop system indefinitely. You need a way to generate supervision automatically.

4. RLVR: The Breakthrough for Agents

RLVR (Reinforcement Learning on Verifiable Rewards) removes the human-in-the-loop bottleneck entirely. Instead of asking humans "which response is better?", you ask verifiers "did this response pass an objective test?"

For many agent tasks, correctness is objective:

The verifier is a simple program - a test suite, a validator function, a diff checker - not a human. And here's the magic: you can generate unlimited amounts of supervision by simply creating more problems programmatically.

Here's the complete RLHF training loop — the policy learns to pick high-quality responses guided by the reward model:

class SoftmaxPolicy:
    """
    Softmax policy over N arms.  pi(a) = softmax(logits)[a]
    Updated via REINFORCE:  logits += lr * (R - baseline) * (1[a] - pi)
    """
    def __init__(self, n_arms: int, lr: float = 0.08, seed: int = 42) -> None:
        self.logits = np.zeros(n_arms)
        self.lr = lr
        self.rng = np.random.default_rng(seed)

    def probs(self) -> np.ndarray:
        e = np.exp(self.logits - self.logits.max())
        return e / e.sum()

    def sample(self) -> int:
        return int(self.rng.choice(len(self.logits), p=self.probs()))

    def update(self, arm: int, reward: float, baseline: float = 0.0) -> None:
        """Single REINFORCE gradient step."""
        advantage = reward - baseline
        grad = -self.probs()
        grad[arm] += 1.0          # nabla log pi(arm)
        self.logits += self.lr * advantage * grad


def train_rlhf_loop(
    bandit: PreferenceBandit,
    reward_model: LogisticRegression,
    n_episodes: int = 300,
    seed: int = 1,
) -> pd.DataFrame:
    """Full REINFORCE loop using reward-model scores as the training signal.
    Episode: sample arm -> pull bandit -> RM scores pull -> REINFORCE update."""
    policy   = SoftmaxPolicy(n_arms=bandit.N_ARMS, lr=0.08, seed=seed)
    baseline = 0.5
    alpha_b  = 0.05
    history  = []

    for ep in range(n_episodes):
        arm      = policy.sample()
        raw      = bandit.pull(arm)
        mean_r   = float(bandit.TRUE_REWARDS.mean())
        diff     = np.array([[raw - mean_r]])
        rm_score = float(reward_model.predict_proba(diff)[0, 1])

        policy.update(arm, rm_score, baseline)
        baseline += alpha_b * (rm_score - baseline)

        history.append({
            "episode":       ep + 1,
            "arm":           arm,
            "raw_reward":    round(raw, 3),
            "rm_score":      round(rm_score, 3),
            "baseline":      round(baseline, 3),
            "best_arm_prob": round(float(policy.probs()[-1]), 3),
        })

    return pd.DataFrame(history)


# Run the full RLHF training loop
history_df = train_rlhf_loop(bandit, rm, n_episodes=300)
print(f"Final best-arm probability: {history_df['best_arm_prob'].iloc[-1]:.1%}")
print(history_df.tail(5)[["episode", "arm", "rm_score", "best_arm_prob"]])
  

In this example, all four verifiers pass, giving rewards of 1 each. In a real training loop, you'd generate many problems (1000s or 10,000s), run the agent on each, collect verifier signals, and accumulate a reward dataset. Then use GRPO or PPO to fine-tune the model.

This is exactly what powered DeepSeek-R1, OpenAI's o1, and similar reasoning-focused models. The models learned to generate chain-of-thought reasoning, check the answer, and propagate the reward signal backward through all the reasoning steps. The result: agents that can reason through multi-step problems effectively.

5. The Tension: When Verifiers Aren't Available

RLVR is transformative, but it requires objective correctness checks. What if your task doesn't have a clear verifier?

For these tasks, RLHF remains necessary - you need human feedback. But the more you can decompose your agent's task into verifiable subtasks, the more you can use RLVR instead. This becomes a key design pattern: structure agent tasks to maximize the verifiable components.

6. Practical Comparison: RLHF vs RLVR

Dimension RLHF RLVR
Reward source Human raters Automated verifier
Scaling cost Linear with volume (hire more raters) Sublinear (write problems, run verifier)
Consistency Variable (humans disagree) Perfect (verifier is deterministic)
Applicable domains Any task Tasks with objective correctness
Training speed Slow (waiting for human annotations) Fast (automated rewards)

7. Real-World Implications for Agents

The shift from RLHF to RLVR is reshaping the entire field. Here's why:

The frontier agentic models being built now use RLVR as a core component. Where human feedback is needed (for tasks without clear verifiers), they layer in RLAIF (Reinforcement Learning from AI Feedback - using a strong model to rate outputs instead of humans), which we'll cover later.

Summary: Building the Foundation

To recap: RLHF was the breakthrough that made LLMs instruction-following. But for agents performing long-horizon, verifiable tasks, RLVR is the better approach. It removes the human bottleneck and enables scalable training with objective reward signals.

In the next post, we'll dive deep into the verifier design patterns used in practice - how to build verifiers for code, math, SQL, and tool use. We'll also explore what happens when verifiers are imperfect, and how to handle edge cases.

Complete Code Reference

All code examples above are from the RL_Agents repository:

To run locally:

# Clone the repo
git clone https://github.com/Pulkit12dhingra/RL_Agents
cd RL_Agents

# Sync dependencies with uv
uv sync

# Run the notebook
jupyter notebook notebooks/post01_rlhf_rlvr_foundations.ipynb
  
Next in the series: Part 2 explores Outcome vs Process Reward Models - the crucial distinction for multi-step tasks.
Browse all posts