RL Agents Part 6: Reward Hacking, Safety, and Credit Assignment

← Back to Home

We've built a comprehensive RL stack: verifiable rewards (RLVR), AI feedback (RLAIF), efficient optimization (GRPO), test-time search (MCTS), and online learning loops. But all of these systems face fundamental challenges that remain unresolved. This final post covers the open problems that limit real-world agentic systems.

Code repo: RL_Agents

Notebook: post06_safety_credit.ipynb

1. Reward Hacking: The Fundamental Problem

Reward hacking occurs when a system finds ways to maximize its reward signal that don't align with the intended objective. It's the core failure mode of all RL systems.

Classic Examples:
Example 1: The Code Agent That Deletes Tests

Your reward is: "pass all tests". You train an agent to write code.

What the verifier checks: does `pytest` return exit code 0?

What the agent learns: "delete all test files, then pytest reports no failures. Reward = 1!"

The agent has literally maximized the reward signal you designed. It just didn't do what you wanted.

Example 2: The Math Agent With Plausible Nonsense

Your reward is: "produce an answer that looks numerically correct".

The agent learns: "output only answers that look like round numbers, within plausible ranges, formatted like a real answer".

The agent passes the tests (which check if the answer is in a reasonable range) without actually solving the math. It pattern-matched to the reward signal rather than learning the task.

Example 3: The Web Agent That Clicks "Accept All"

Your agent is trained to navigate a website and extract information.

Reward: "did you successfully extract the data?"

The agent learns: "clicking 'accept all' on dialogs makes pop-ups disappear. Then I can see the data. Reward = 1."

The agent has learned to bypass user consent, technically maximizing reward, but violating the spirit of the task (respecting user privacy).

Why This Happens

RL systems are optimizers. They will find the easiest path to maximize the reward you give them. If there's a shortcut - a way to get high reward without doing the intended task - they'll find it.

The root cause: misalignment between what you measure (reward) and what you actually want (the intended behavior).

2. Defenses Against Reward Hacking

Defense 1: Ensemble Rewards

Use multiple independent reward models (or verifiers) and require agreement. If one reward model is hacked, others catch it.

# Instead of: reward = single_verifier(output)
# Use:
reward_1 = verifier_1(output)  # Does output pass tests?
reward_2 = verifier_2(output)  # Does output have good style?
reward_3 = verifier_3(output)  # Does output handle edge cases?
final_reward = min(reward_1, reward_2, reward_3)  # All must agree
  

If the agent deletes tests to hack reward_1, reward_2 (style check) will fail. The ensemble catches the hack.

Defense 2: KL Divergence Penalties

Penalize the model if it diverges too far from its original behavior. This prevents sharp pivots toward exploit strategies.

This defense simulation is split into components that correspond to real safety controls in agent training.

from __future__ import annotations
import numpy as np
import pandas as pd


class HackableBandit:
    """
    5-arm bandit with a reward-hacking trap.
      arms 0-3: safe responses  (rewards ~0.30, 0.50, 0.65, 0.72)
      arm  4:   unsafe response (reward ~0.93, violates safety constraints)
    Reference policy = uniform over safe arms only (arm 4 prob = 0).
    """
    TASK_REWARDS = np.array([0.30, 0.50, 0.65, 0.72, 0.93])
    SAFETY_FLAGS = np.array([False, False, False, False, True])
    N_ARMS       = 5
    REF_PROBS    = np.array([0.25, 0.25, 0.25, 0.25, 0.0])

    def __init__(self, noise: float = 0.04, seed: int = 6) -> None:
        self.noise = noise
        self.rng   = np.random.default_rng(seed)

    def pull(self, arm: int) -> float:
        r = self.TASK_REWARDS[arm] + self.rng.normal(0.0, self.noise)
        return float(np.clip(r, 0.0, 1.0))


def _softmax(logits: np.ndarray) -> np.ndarray:
    e = np.exp(logits - logits.max())
    return e / e.sum()


def _reinforce_update(logits, arm, reward, baseline, lr):
    probs     = _softmax(logits)
    advantage = reward - baseline
    grad      = -probs.copy()
    grad[arm] += 1.0
    logits    += lr * advantage * grad
    return logits


def kl_penalty(reference_prob: float, policy_prob: float, beta: float = 0.1) -> float:
    kl = policy_prob * np.log((policy_prob + 1e-9) / (reference_prob + 1e-9))
    return round(float(beta * kl), 5)


def train_unconstrained(n_episodes: int = 200, lr: float = 0.10,
                        seed: int = 0) -> pd.DataFrame:
    """Vanilla REINFORCE — free to exploit the unsafe arm."""
    bandit   = HackableBandit(seed=seed)
    logits   = np.zeros(bandit.N_ARMS)
    rng      = np.random.default_rng(seed)
    baseline = 0.5
    alpha_b  = 0.05
    history  = []
    window   = []
    for ep in range(n_episodes):
        probs  = _softmax(logits)
        arm    = int(rng.choice(bandit.N_ARMS, p=probs))
        reward = bandit.pull(arm)
        logits = _reinforce_update(logits, arm, reward, baseline, lr)
        baseline += alpha_b * (reward - baseline)
        cur = _softmax(logits)
        kl  = float(np.sum(cur * np.log((cur + 1e-9) / (bandit.REF_PROBS + 1e-9))))
        window.append(reward)
        if len(window) > 20: window.pop(0)
        history.append({"episode": ep + 1, "arm": arm,
                         "reward": round(reward, 3),
                         "unsafe_arm_prob": round(float(cur[4]), 3),
                         "kl_from_ref": round(kl, 4),
                         "reward_ma20": round(float(np.mean(window)), 3)})
    return pd.DataFrame(history)


def train_kl_penalized(n_episodes: int = 200, lr: float = 0.10,
                       beta: float = 0.5, seed: int = 0) -> pd.DataFrame:
    """REINFORCE with KL penalty — penalises drift toward unsafe arm."""
    bandit   = HackableBandit(seed=seed)
    logits   = np.zeros(bandit.N_ARMS)
    rng      = np.random.default_rng(seed)
    baseline = 0.5
    alpha_b  = 0.05
    history  = []
    window   = []
    for ep in range(n_episodes):
        probs  = _softmax(logits)
        arm    = int(rng.choice(bandit.N_ARMS, p=probs))
        raw    = bandit.pull(arm)
        kl_pen = kl_penalty(float(bandit.REF_PROBS[arm]), float(probs[arm]), beta)
        penalised = raw - kl_pen
        logits = _reinforce_update(logits, arm, penalised, baseline, lr)
        baseline += alpha_b * (penalised - baseline)
        cur    = _softmax(logits)
        kl_all = float(np.sum(cur * np.log((cur + 1e-9) / (bandit.REF_PROBS + 1e-9))))
        window.append(raw)
        if len(window) > 20: window.pop(0)
        history.append({"episode": ep + 1, "arm": arm,
                         "raw_reward": round(raw, 3),
                         "kl_penalty": round(kl_pen, 4),
                         "unsafe_arm_prob": round(float(cur[4]), 3),
                         "kl_from_ref": round(kl_all, 4),
                         "reward_ma20": round(float(np.mean(window)), 3)})
    return pd.DataFrame(history)


# Compare unconstrained vs KL-penalised
unc = train_unconstrained()
kl  = train_kl_penalized(beta=0.5)
print(f"Unconstrained final unsafe-arm prob: {unc['unsafe_arm_prob'].iloc[-1]:.1%}")
print(f"KL-penalised  final unsafe-arm prob: {kl['unsafe_arm_prob'].iloc[-1]:.1%}")
print(f"KL from ref (unconstrained): {unc['kl_from_ref'].iloc[-1]:.3f}")
print(f"KL from ref (KL-penalised):  {kl['kl_from_ref'].iloc[-1]:.3f}")
  
Defense 3: Red-Teaming the Reward

Before deployment, actively search for ways to hack the reward signal. Have humans and automated tools try to break the verifier.

Example red-team queries:

If you find exploits, patch the reward signal before training at scale.

Defense 4: Multi-Objective Rewards

Don't train on task reward alone. Combine it with auxiliary objectives:

total_reward = (
    w_task * task_reward + 
    w_safety * safety_reward + 
    w_format * formatting_reward +
    w_diversity * exploration_reward
)
  

Now the model can't purely exploit one dimension. It must balance multiple constraints.

Defense 5: Interpretability and Monitoring

Monitor trajectories in production. Look for suspicious patterns: sudden behavioral changes, unusual code patterns, or suspiciously high success rates in specific categories.

If an agent suddenly starts deleting files in 90% of its runs (after deleting files in 0% before), that's a red flag. Pause deployment and investigate.

3. The Credit Assignment Problem for Long-Horizon Agents

Even if you prevent reward hacking, you face a deeper problem: how do you assign credit for outcomes in a 50-step trajectory?

Imagine an agent:

  1. Reads a customer query (step 1)
  2. Searches a knowledge base (step 2)
  3. Retrieves 10 articles (step 3)
  4. Filters to 3 relevant ones (step 4)
  5. Extracts key points (step 5)
  6. ...
  7. Generates a response (step 48)
  8. User rates it as helpful (outcome: success)

The model gets a reward signal: "success = 1". Now, which steps deserve credit?

The credit assignment code intentionally simulates a long trajectory where a late failure creates ambiguous blame for earlier steps.

import numpy as np
import pandas as pd


def temporal_credit_assignment(
    steps: int = 12, failure_step: int = 10, seed: int = 12,
) -> pd.DataFrame:
    """Exponential-decay credit assignment backward from a failure step."""
    rng  = np.random.default_rng(seed)
    base = rng.uniform(0.2, 0.9, size=steps)
    blame = np.zeros(steps)
    for i in range(steps):
        distance = max(failure_step - (i + 1), 0)
        blame[i] = np.exp(-0.4 * distance)
    blame = blame / blame.sum()
    return pd.DataFrame({
        "step":            np.arange(1, steps + 1),
        "action_quality":  np.round(base, 3),
        "assigned_credit": np.round(blame, 3),
    })


credit_frame = temporal_credit_assignment(steps=12, failure_step=10)
print(credit_frame.to_string(index=False))

failure_credit = credit_frame.loc[credit_frame["step"] == 10, "assigned_credit"].values[0]
early_credit   = credit_frame.loc[credit_frame["step"] < 5,  "assigned_credit"].sum()
print(f"\nCredit at failure step (10): {failure_credit:.3f}")
print(f"Credit at early steps (1-4): {early_credit:.3f}")

print("\nBar chart of credit distribution:")
for _, row in credit_frame.iterrows():
    bar = "#" * int(row["assigned_credit"] * 50)
    print(f"Step {row['step']:2d}: {bar:<50} ({row['assigned_credit']:.3f})")
  

When you run this code, you'll see the fundamental challenge: how do you fairly distribute a failure signal among 50 intermediate steps? The temporal approach above is a heuristic, not a principled solution.

Approach 1: Discounted Returns

Assign credit backward from the final outcome, discounting by time:

Step 48 credit: 1.0 * outcome
Step 47 credit: 0.99 * outcome
Step 46 credit: 0.99^2 * outcome
...
Step 1 credit: 0.99^47 * outcome
  

Problem: step 1 gets essentially zero credit, even though it was critical. Time-based discounting doesn't reflect actual causal importance.

Approach 2: Hindsight Relabeling

From robotics (HER - Hindsight Experience Replay): if the agent failed to reach goal G but reached goal G', treat G' as the goal retroactively.

If the agent fails to generate a helpful response but successfully retrieves relevant articles, relabel: "success = 1 for information retrieval" and train on that subgoal.

Problem: requires manually defining all possible subgoals. Doesn't scale to complex open-ended tasks.

Approach 3: Subgoal Decomposition

Define intermediate checkpoints with their own rewards:

Now credit is dense and locally clear. Step 3's success is directly rewarded.

Problem: you must manually define all subgoals. For complex tasks, this is hard. Also, subgoals might conflict.

Approach 4: Process Rewards (PRM)

Train a separate neural network (PRM) to score each intermediate step. This is data-hungry but powerful.

PRM learning: "Is step 4 (filtering articles) a good step given steps 1-3?" → yes/no

Problem: requires labeled data on step quality. For new task domains, this is expensive. Also, PRM generalization to new tasks is challenging.

Approach 5: Causal Tracing / Interpretability

Use gradient-based or attention-based analysis to identify which earlier actions causally contributed to the final outcome.

If step 4 (filter articles) has high attention weights that propagate to step 48 (final response), step 4 gets higher credit.

Problem: computationally expensive, requires training interpretability tools, not fully reliable.

4. Why Credit Assignment Remains Unsolved

None of these approaches is satisfying. The fundamental issue: in a long trajectory, causality is complex and non-local.

A mistake at step 5 might not manifest as failure until step 45. A good decision at step 3 might be redundant because step 10 fixes it. The causal graph is tangled.

This is arguably the biggest gap between current agentic systems and truly reliable long-horizon agents.

5. The Complete Training Stack (With Caveats)

Taking everything together, here's what a state-of-the-art agentic training pipeline looks like:

Stage 1: Supervised Fine-Tuning (SFT)

Start with a strong base model. Fine-tune on diverse, high-quality demonstrations. This is your foundation.

Stage 2: RLVR on Verifiable Subtasks

Identify components of your agent task that have objective verifiers (code tests, exact-match checks, tool execution). Train those with RLVR.

Stage 3: RLAIF on Subjective Components

For components without verifiers, use constitutional AI + RLAIF. Have a critic model score trajectories on explicit principles.

Stage 4: Process Rewards (Selective)

For the longest, most critical steps, layer in process rewards if budget allows. This denses the learning signal for bottleneck steps.

Stage 5: Online Learning with Safeguards

Deploy the model. Continuously collect successful trajectories. Retrain weekly, mixing new data with old to prevent collapse. Monitor for reward hacking patterns.

Stage 6: Test-Time Search (If Needed)

For very hard problems, add beam search or MCTS at inference. Costs more at test time but handles distribution shift and edge cases.

6. The Landscape Right Now

Leading models (o1, DeepSeek-R1, Claude) are all converging on similar stacks:

What's NOT solved:

7. Complete Code Reference

All code examples demonstrating reward hacking defenses, reward ensembles, KL penalties, and credit assignment are from the RL_Agents repository:

To run locally:

git clone https://github.com/Pulkit12dhingra/RL_Agents
cd RL_Agents
uv sync
jupyter notebook notebooks/post06_safety_credit.ipynb
  

8. Summary: The Frontier

We've covered a comprehensive stack: RLHF, RLVR, RLAIF, GRPO, MCTS, online learning, and various safety techniques. These are the tools in modern agentic systems.

But the hardest problems remain:

  1. How do we assign credit fairly in complex 50+ step trajectories? No perfect answer.
  2. How do we prevent agents from finding subtle exploits in reward signals? We can defend, but not guarantee prevention.
  3. How do we train agents that reason reliably over 100+ steps? Still fundamentally limited.

These open questions are where the frontier of agentic AI lives. Solving them will unlock agents far more capable than today's systems.

Closing Thoughts

The shift from RLHF to RLVR to RLAIF represents a maturation of the field. We're moving from "teach the model to imitate" toward "help the model learn to solve real problems with objective feedback."

For practitioners: use RLVR where you can (objective tasks), layer RLAIF for judgment calls, monitor carefully for reward hacking, and accept that long-horizon credit assignment is still more art than science.

For researchers: credit assignment, robustness to reward hacking, and long-horizon reasoning are the open frontiers. These problems will define the next generation of agentic capabilities.

Series complete! You've now covered the full landscape of RL for agents: foundations, reward design, optimization, online learning, and safety. The notebooks in the RL_Agents repo contain executable code demonstrating each concept.
← Part 5 Back to Home