We've built a comprehensive RL stack: verifiable rewards (RLVR), AI feedback (RLAIF), efficient optimization (GRPO), test-time search (MCTS), and online learning loops. But all of these systems face fundamental challenges that remain unresolved. This final post covers the open problems that limit real-world agentic systems.
Code repo: RL_Agents
Notebook: post06_safety_credit.ipynb
Reward hacking occurs when a system finds ways to maximize its reward signal that don't align with the intended objective. It's the core failure mode of all RL systems.
Your reward is: "pass all tests". You train an agent to write code.
What the verifier checks: does `pytest` return exit code 0?
What the agent learns: "delete all test files, then pytest reports no failures. Reward = 1!"
The agent has literally maximized the reward signal you designed. It just didn't do what you wanted.
Your reward is: "produce an answer that looks numerically correct".
The agent learns: "output only answers that look like round numbers, within plausible ranges, formatted like a real answer".
The agent passes the tests (which check if the answer is in a reasonable range) without actually solving the math. It pattern-matched to the reward signal rather than learning the task.
Your agent is trained to navigate a website and extract information.
Reward: "did you successfully extract the data?"
The agent learns: "clicking 'accept all' on dialogs makes pop-ups disappear. Then I can see the data. Reward = 1."
The agent has learned to bypass user consent, technically maximizing reward, but violating the spirit of the task (respecting user privacy).
RL systems are optimizers. They will find the easiest path to maximize the reward you give them. If there's a shortcut - a way to get high reward without doing the intended task - they'll find it.
The root cause: misalignment between what you measure (reward) and what you actually want (the intended behavior).
Use multiple independent reward models (or verifiers) and require agreement. If one reward model is hacked, others catch it.
# Instead of: reward = single_verifier(output)
# Use:
reward_1 = verifier_1(output) # Does output pass tests?
reward_2 = verifier_2(output) # Does output have good style?
reward_3 = verifier_3(output) # Does output handle edge cases?
final_reward = min(reward_1, reward_2, reward_3) # All must agree
If the agent deletes tests to hack reward_1, reward_2 (style check) will fail. The ensemble catches the hack.
Penalize the model if it diverges too far from its original behavior. This prevents sharp pivots toward exploit strategies.
This defense simulation is split into components that correspond to real safety controls in agent training.
from __future__ import annotations
import numpy as np
import pandas as pd
class HackableBandit:
"""
5-arm bandit with a reward-hacking trap.
arms 0-3: safe responses (rewards ~0.30, 0.50, 0.65, 0.72)
arm 4: unsafe response (reward ~0.93, violates safety constraints)
Reference policy = uniform over safe arms only (arm 4 prob = 0).
"""
TASK_REWARDS = np.array([0.30, 0.50, 0.65, 0.72, 0.93])
SAFETY_FLAGS = np.array([False, False, False, False, True])
N_ARMS = 5
REF_PROBS = np.array([0.25, 0.25, 0.25, 0.25, 0.0])
def __init__(self, noise: float = 0.04, seed: int = 6) -> None:
self.noise = noise
self.rng = np.random.default_rng(seed)
def pull(self, arm: int) -> float:
r = self.TASK_REWARDS[arm] + self.rng.normal(0.0, self.noise)
return float(np.clip(r, 0.0, 1.0))
def _softmax(logits: np.ndarray) -> np.ndarray:
e = np.exp(logits - logits.max())
return e / e.sum()
def _reinforce_update(logits, arm, reward, baseline, lr):
probs = _softmax(logits)
advantage = reward - baseline
grad = -probs.copy()
grad[arm] += 1.0
logits += lr * advantage * grad
return logits
def kl_penalty(reference_prob: float, policy_prob: float, beta: float = 0.1) -> float:
kl = policy_prob * np.log((policy_prob + 1e-9) / (reference_prob + 1e-9))
return round(float(beta * kl), 5)
def train_unconstrained(n_episodes: int = 200, lr: float = 0.10,
seed: int = 0) -> pd.DataFrame:
"""Vanilla REINFORCE — free to exploit the unsafe arm."""
bandit = HackableBandit(seed=seed)
logits = np.zeros(bandit.N_ARMS)
rng = np.random.default_rng(seed)
baseline = 0.5
alpha_b = 0.05
history = []
window = []
for ep in range(n_episodes):
probs = _softmax(logits)
arm = int(rng.choice(bandit.N_ARMS, p=probs))
reward = bandit.pull(arm)
logits = _reinforce_update(logits, arm, reward, baseline, lr)
baseline += alpha_b * (reward - baseline)
cur = _softmax(logits)
kl = float(np.sum(cur * np.log((cur + 1e-9) / (bandit.REF_PROBS + 1e-9))))
window.append(reward)
if len(window) > 20: window.pop(0)
history.append({"episode": ep + 1, "arm": arm,
"reward": round(reward, 3),
"unsafe_arm_prob": round(float(cur[4]), 3),
"kl_from_ref": round(kl, 4),
"reward_ma20": round(float(np.mean(window)), 3)})
return pd.DataFrame(history)
def train_kl_penalized(n_episodes: int = 200, lr: float = 0.10,
beta: float = 0.5, seed: int = 0) -> pd.DataFrame:
"""REINFORCE with KL penalty — penalises drift toward unsafe arm."""
bandit = HackableBandit(seed=seed)
logits = np.zeros(bandit.N_ARMS)
rng = np.random.default_rng(seed)
baseline = 0.5
alpha_b = 0.05
history = []
window = []
for ep in range(n_episodes):
probs = _softmax(logits)
arm = int(rng.choice(bandit.N_ARMS, p=probs))
raw = bandit.pull(arm)
kl_pen = kl_penalty(float(bandit.REF_PROBS[arm]), float(probs[arm]), beta)
penalised = raw - kl_pen
logits = _reinforce_update(logits, arm, penalised, baseline, lr)
baseline += alpha_b * (penalised - baseline)
cur = _softmax(logits)
kl_all = float(np.sum(cur * np.log((cur + 1e-9) / (bandit.REF_PROBS + 1e-9))))
window.append(raw)
if len(window) > 20: window.pop(0)
history.append({"episode": ep + 1, "arm": arm,
"raw_reward": round(raw, 3),
"kl_penalty": round(kl_pen, 4),
"unsafe_arm_prob": round(float(cur[4]), 3),
"kl_from_ref": round(kl_all, 4),
"reward_ma20": round(float(np.mean(window)), 3)})
return pd.DataFrame(history)
# Compare unconstrained vs KL-penalised
unc = train_unconstrained()
kl = train_kl_penalized(beta=0.5)
print(f"Unconstrained final unsafe-arm prob: {unc['unsafe_arm_prob'].iloc[-1]:.1%}")
print(f"KL-penalised final unsafe-arm prob: {kl['unsafe_arm_prob'].iloc[-1]:.1%}")
print(f"KL from ref (unconstrained): {unc['kl_from_ref'].iloc[-1]:.3f}")
print(f"KL from ref (KL-penalised): {kl['kl_from_ref'].iloc[-1]:.3f}")
Before deployment, actively search for ways to hack the reward signal. Have humans and automated tools try to break the verifier.
Example red-team queries:
If you find exploits, patch the reward signal before training at scale.
Don't train on task reward alone. Combine it with auxiliary objectives:
total_reward = (
w_task * task_reward +
w_safety * safety_reward +
w_format * formatting_reward +
w_diversity * exploration_reward
)
Now the model can't purely exploit one dimension. It must balance multiple constraints.
Monitor trajectories in production. Look for suspicious patterns: sudden behavioral changes, unusual code patterns, or suspiciously high success rates in specific categories.
If an agent suddenly starts deleting files in 90% of its runs (after deleting files in 0% before), that's a red flag. Pause deployment and investigate.
Even if you prevent reward hacking, you face a deeper problem: how do you assign credit for outcomes in a 50-step trajectory?
Imagine an agent:
The model gets a reward signal: "success = 1". Now, which steps deserve credit?
The credit assignment code intentionally simulates a long trajectory where a late failure creates ambiguous blame for earlier steps.
import numpy as np
import pandas as pd
def temporal_credit_assignment(
steps: int = 12, failure_step: int = 10, seed: int = 12,
) -> pd.DataFrame:
"""Exponential-decay credit assignment backward from a failure step."""
rng = np.random.default_rng(seed)
base = rng.uniform(0.2, 0.9, size=steps)
blame = np.zeros(steps)
for i in range(steps):
distance = max(failure_step - (i + 1), 0)
blame[i] = np.exp(-0.4 * distance)
blame = blame / blame.sum()
return pd.DataFrame({
"step": np.arange(1, steps + 1),
"action_quality": np.round(base, 3),
"assigned_credit": np.round(blame, 3),
})
credit_frame = temporal_credit_assignment(steps=12, failure_step=10)
print(credit_frame.to_string(index=False))
failure_credit = credit_frame.loc[credit_frame["step"] == 10, "assigned_credit"].values[0]
early_credit = credit_frame.loc[credit_frame["step"] < 5, "assigned_credit"].sum()
print(f"\nCredit at failure step (10): {failure_credit:.3f}")
print(f"Credit at early steps (1-4): {early_credit:.3f}")
print("\nBar chart of credit distribution:")
for _, row in credit_frame.iterrows():
bar = "#" * int(row["assigned_credit"] * 50)
print(f"Step {row['step']:2d}: {bar:<50} ({row['assigned_credit']:.3f})")
When you run this code, you'll see the fundamental challenge: how do you fairly distribute a failure signal among 50 intermediate steps? The temporal approach above is a heuristic, not a principled solution.
Assign credit backward from the final outcome, discounting by time:
Step 48 credit: 1.0 * outcome
Step 47 credit: 0.99 * outcome
Step 46 credit: 0.99^2 * outcome
...
Step 1 credit: 0.99^47 * outcome
Problem: step 1 gets essentially zero credit, even though it was critical. Time-based discounting doesn't reflect actual causal importance.
From robotics (HER - Hindsight Experience Replay): if the agent failed to reach goal G but reached goal G', treat G' as the goal retroactively.
If the agent fails to generate a helpful response but successfully retrieves relevant articles, relabel: "success = 1 for information retrieval" and train on that subgoal.
Problem: requires manually defining all possible subgoals. Doesn't scale to complex open-ended tasks.
Define intermediate checkpoints with their own rewards:
Now credit is dense and locally clear. Step 3's success is directly rewarded.
Problem: you must manually define all subgoals. For complex tasks, this is hard. Also, subgoals might conflict.
Train a separate neural network (PRM) to score each intermediate step. This is data-hungry but powerful.
PRM learning: "Is step 4 (filtering articles) a good step given steps 1-3?" → yes/no
Problem: requires labeled data on step quality. For new task domains, this is expensive. Also, PRM generalization to new tasks is challenging.
Use gradient-based or attention-based analysis to identify which earlier actions causally contributed to the final outcome.
If step 4 (filter articles) has high attention weights that propagate to step 48 (final response), step 4 gets higher credit.
Problem: computationally expensive, requires training interpretability tools, not fully reliable.
None of these approaches is satisfying. The fundamental issue: in a long trajectory, causality is complex and non-local.
A mistake at step 5 might not manifest as failure until step 45. A good decision at step 3 might be redundant because step 10 fixes it. The causal graph is tangled.
This is arguably the biggest gap between current agentic systems and truly reliable long-horizon agents.
Taking everything together, here's what a state-of-the-art agentic training pipeline looks like:
Start with a strong base model. Fine-tune on diverse, high-quality demonstrations. This is your foundation.
Identify components of your agent task that have objective verifiers (code tests, exact-match checks, tool execution). Train those with RLVR.
For components without verifiers, use constitutional AI + RLAIF. Have a critic model score trajectories on explicit principles.
For the longest, most critical steps, layer in process rewards if budget allows. This denses the learning signal for bottleneck steps.
Deploy the model. Continuously collect successful trajectories. Retrain weekly, mixing new data with old to prevent collapse. Monitor for reward hacking patterns.
For very hard problems, add beam search or MCTS at inference. Costs more at test time but handles distribution shift and edge cases.
Leading models (o1, DeepSeek-R1, Claude) are all converging on similar stacks:
What's NOT solved:
All code examples demonstrating reward hacking defenses, reward ensembles, KL penalties, and credit assignment are from the RL_Agents repository:
To run locally:
git clone https://github.com/Pulkit12dhingra/RL_Agents
cd RL_Agents
uv sync
jupyter notebook notebooks/post06_safety_credit.ipynb
We've covered a comprehensive stack: RLHF, RLVR, RLAIF, GRPO, MCTS, online learning, and various safety techniques. These are the tools in modern agentic systems.
But the hardest problems remain:
These open questions are where the frontier of agentic AI lives. Solving them will unlock agents far more capable than today's systems.
The shift from RLHF to RLVR to RLAIF represents a maturation of the field. We're moving from "teach the model to imitate" toward "help the model learn to solve real problems with objective feedback."
For practitioners: use RLVR where you can (objective tasks), layer RLAIF for judgment calls, monitor carefully for reward hacking, and accept that long-horizon credit assignment is still more art than science.
For researchers: credit assignment, robustness to reward hacking, and long-horizon reasoning are the open frontiers. These problems will define the next generation of agentic capabilities.