Now that we understand RLVR, we hit the next critical question: when do you give the reward signal? At the very end of the task? Or at each intermediate step? This distinction matters enormously for long-horizon agent training.
Code repo: RL_Agents
Notebook: post03_orm_vs_prm.ipynb
There are fundamentally two ways to structure reward signals during agent trajectories:
An ORM evaluates the final output and gives a single reward (or score) for the entire trajectory. For example:
Advantages of ORM:
Disadvantages of ORM:
A PRM evaluates the trajectory at each step and provides feedback throughout. For example:
Advantages of PRM:
Disadvantages of PRM:
Let's say an agent is tasked with fixing a broken Python script. It can take up to 20 actions: read the error, hypothesize a bug, modify the code, run tests, iterate, etc.
Problem: if tests fail, the agent has no idea which of the 20 actions was wrong. Was it action 3? Action 15? Was the final modification unnecessary and bad? The model must figure this out entirely from the one binary signal at the end. This is extremely sample-inefficient.
Now the agent gets step-level feedback. If action 15 was a bad code change, the agent sees immediately that it failed the step-level verifier. It learns to avoid that pattern. This is much more efficient.
The research on this is clear: for multi-step reasoning and agent tasks, PRMs significantly outperform ORMs. Here's why:
The tradeoff is real: PRM requires defining intermediate evaluation criteria, and it's more computationally expensive. But for agents that must reliably solve multi-step problems, it's non-negotiable.
| Task Characteristics | ORM Viable? | PRM Needed? |
|---|---|---|
| Single decision (translate, summarize) | ✓ Yes | No |
| Few steps (3-5), clear goal | ✓ Maybe | Helps, but ORM works |
| Multi-step (10-20), complex reasoning | ✗ No | ✓ Strongly recommended |
| Long horizon (40-100 steps) | ✗ No way | ✓ Essential |
| Intermediate states hard to evaluate | ✓ Acceptable | Difficult to implement |
The simulation below mirrors a real multi-step agent episode and shows exactly where each reward architecture injects learning signal.
from dataclasses import dataclass
import sqlite3
import numpy as np
import pandas as pd
@dataclass
class VerifierResult:
task_name: str
passed: bool
reward: int
details: str
def verify_math(expected: int, predicted: int) -> VerifierResult:
passed = expected == predicted
return VerifierResult("math", passed, int(passed),
f"expected={expected}, predicted={predicted}")
def verify_code() -> VerifierResult:
def is_even(n: int) -> bool:
return n % 2 == 0
tests = [is_even(2), not is_even(3), is_even(8)]
passed = all(tests)
return VerifierResult("code", passed, int(passed), f"tests={tests}")
def verify_sql() -> VerifierResult:
conn = sqlite3.connect(":memory:")
cur = conn.cursor()
cur.execute("CREATE TABLE orders(customer TEXT, amount INTEGER)")
cur.executemany("INSERT INTO orders VALUES(?, ?)",
[("alice", 120), ("bob", 40), ("alice", 80)])
cur.execute(
"SELECT customer, SUM(amount) total "
"FROM orders GROUP BY customer ORDER BY total DESC"
)
rows = cur.fetchall()
conn.close()
expected = [("alice", 200), ("bob", 40)]
passed = rows == expected
return VerifierResult("sql", passed, int(passed), f"rows={rows}")
def run_verifier_batch() -> pd.DataFrame:
results = [verify_math(expected=64, predicted=64), verify_code(), verify_sql()]
frame = pd.DataFrame({
"task": [r.task_name for r in results],
"passed": [r.passed for r in results],
"reward": [r.reward for r in results],
"details": [r.details for r in results],
})
frame["pass_rate"] = frame["passed"].expanding().mean().round(2)
return frame
# ── Real RLVR: REINFORCE on integer addition tasks ───────────────────────────
class MathTaskEnv:
"""Integer addition environment. Agent learns a+b from binary reward only."""
def __init__(self, seed: int = 17) -> None:
self.rng = np.random.default_rng(seed)
self.a = self.b = self.answer = 0
def reset(self) -> tuple[int, int]:
self.a = int(self.rng.integers(1, 6))
self.b = int(self.rng.integers(1, 6))
self.answer = self.a + self.b
return self.a, self.b
def step(self, predicted: int) -> int:
return verify_math(self.answer, predicted).reward
class GaussianPolicy:
"""Learns to always pick offset=0 (the correct answer) via REINFORCE."""
OFFSETS = [-2, -1, 0, 1, 2]
N = 5
def __init__(self, lr: float = 0.20, seed: int = 7) -> None:
self.logits = np.array([-0.2, -0.1, 0.0, -0.1, -0.2])
self.lr = lr
self.rng = np.random.default_rng(seed)
self._last_idx = 0
self.bias = 0.0
def _probs(self) -> np.ndarray:
e = np.exp(self.logits - self.logits.max())
return e / e.sum()
def sample(self, a: int, b: int) -> int:
probs = self._probs()
self._last_idx = int(self.rng.choice(self.N, p=probs))
return a + b + self.OFFSETS[self._last_idx]
def update(self, reward: float, baseline: float) -> None:
advantage = reward - baseline
grad = -self._probs()
grad[self._last_idx] += 1.0
self.logits += self.lr * advantage * grad
self.bias = round(float(self.logits[2] - self.logits.mean()), 3)
def train_rlvr_loop(n_episodes: int = 200) -> pd.DataFrame:
"""REINFORCE on math tasks. Binary verifier is the sole reward signal."""
env = MathTaskEnv()
policy = GaussianPolicy()
baseline = 0.3
alpha_b = 0.05
history = []
window = []
for ep in range(n_episodes):
a, b = env.reset()
predicted = policy.sample(a, b)
reward = env.step(predicted)
policy.update(float(reward), baseline)
baseline += alpha_b * (float(reward) - baseline)
window.append(reward)
if len(window) > 20:
window.pop(0)
history.append({
"episode": ep + 1,
"a": a, "b": b,
"predicted": predicted,
"correct": a + b,
"reward": reward,
"bias": policy.bias,
"pass_rate_20": round(float(np.mean(window)), 3),
})
return pd.DataFrame(history)
# Run verifiers, then the RLVR training loop
print(run_verifier_batch().to_string(index=False))
print()
df = train_rlvr_loop(n_episodes=200)
print(f"Final 20-ep pass rate: {df['pass_rate_20'].iloc[-1]:.1%}")
print(df.tail(5)[["episode", "predicted", "correct", "reward", "pass_rate_20"]])
The verifiers are pure Python functions that return a binary reward with no human involvement. The RLVR training loop shows the policy improving from random (~30% pass rate) to near-perfect accuracy purely from binary verifier signals — no ground truth is ever exposed to the policy directly.
The code above is adapted from the RL_Agents repository module:
If you commit to PRM, you need step-level verifiers. Here are examples:
The key insight: you need domain-specific verifiers for each task type. But the payoff is enormous: your agents learn much faster and solve much harder problems.
In OpenAI's o1 and DeepSeek-R1 research, models trained with process rewards (step-level feedback) dramatically outperform models trained with only outcome rewards (final answer only):
The reason: PRMs provide much richer learning signal, reducing the sample complexity of learning long-horizon policies.
For short, well-defined tasks: ORM is sufficient and cheaper.
For long-horizon agent tasks: PRM is nearly mandatory. The dense feedback enables much more reliable learning.
For deployed agents in the wild: The best-performing systems use PRM selectively - evaluating steps that matter most and skipping steps where feedback is obvious.
In the next post, we'll explore the optimization algorithms that actually use these rewards - PPO, GRPO, and why newer approaches are often more efficient than classic policy gradient methods.