So far we've covered RLVR - training agents using verifiable rewards from test suites, exact-match correctness checks, and tool execution. But many agent tasks don't have clear verifiers. A creative writing task, a strategic recommendation, or a nuanced conversational response can't be reduced to "pass/fail". For these tasks, we return to the human feedback problem. But instead of having humans rate outputs directly, we can use a strong AI model as a critic. This is RLAIF and Constitutional AI.
Code repo: RL_Agents
Notebook: post06_safety_credit.ipynb
RLVR works when correctness is objective. But consider these tasks:
For these tasks, we need a scoring function that can handle subjective qualities. We can't write a test suite. We need judgment.
The naive solution: hire more human raters. But this returns us to the cost and latency problems we were trying to solve with RLVR.
RLAIF (Reinforcement Learning from AI Feedback) replaces human raters with another AI system. Instead of asking humans "which response is better?", you ask a strong language model to compare two responses using explicit criteria.
The advantage: you've replaced human raters with API calls to a strong model. The cost is per-call instead of per-human, and it scales to any size dataset you want.
Your agent is trained to respond to customer service inquiries. Instead of hiring 10 human raters to judge 10,000 responses, you can:
Prompt: "Customer asked: How do I reset my password?"
Response A (from your model):
"Click the login page, then click forgot password, and check your email."
Response B (also from your model):
"You can reset by clicking forgot password on the login page.
Check your email for a reset link. If you don't see it,
check spam. Let me know if you need more help."
Critic evaluation:
"Response B is better because it:
- Anticipates the user might miss the email (spam)
- Offers continued support
- Is more friendly in tone
Preference: B > A"
Score: B gets higher reward
This entire evaluation is automated. You can do it at scale without human overhead.
Constitutional AI (CAI) extends RLAIF by formalizing what the critic should be looking for. Instead of asking "which is better?" (vague), you ask questions based on a constitution - a set of explicit principles.
Constitutional principle 1: Responses must be accurate. If you're unsure, say so.
Constitutional principle 2: Responses must be helpful to the customer's actual problem.
Constitutional principle 3: Responses must be respectful and professional in tone.
Constitutional principle 4: Responses must not encourage actions that could harm the customer.
Now, when the critic evaluates two responses, it explicitly checks them against these principles:
Critic evaluation (constitutional):
Response A: "Click forgot password"
- Accuracy: OK, but incomplete (score: 6/10)
- Helpfulness: Low, lacks detail (score: 4/10)
- Tone: Neutral, not engaging (score: 5/10)
- Safety: OK (score: 8/10)
- Overall: C-grade response
Response B: "Click forgot password... email... spam folder... offer help"
- Accuracy: OK (score: 8/10)
- Helpfulness: High, anticipates issues (score: 9/10)
- Tone: Friendly, professional (score: 8/10)
- Safety: OK (score: 8/10)
- Overall: A-grade response
Preference: B > A
By making the evaluation criteria explicit (the constitution), you get:
So far we've talked about evaluating single responses. But for agents performing multi-step tasks, you can extend this to critique full trajectories:
Agent trajectory for: "Solve this system of equations"
Step 1: Wrote the equations in matrix form ✓
Step 2: Computed the determinant ✓
Step 3: Applied Cramer's rule ✓
Step 4: Got x=2, y=3 ✓
Step 5: Verified by substitution ✓
Constitutional evaluation:
- Did the agent follow a logical process? YES (score: 10/10)
- Were all steps mathematically correct? YES (score: 10/10)
- Could a student learn from this? YES (score: 9/10)
- Overall: Excellent trajectory, give high reward
Now you're using a critic to score the entire reasoning path, not just the final answer. This is closer to PRM (process rewards) but using AI critique instead of hard verifiers.
The most sophisticated agent training uses both:
Example: training a research assistant agent
By mixing verifiers (objective) with critic evaluations (subjective), you get the benefits of both: scalability and principled judgment.
| Dimension | RLHF | RLAIF |
|---|---|---|
| Cost per evaluation | $0.02-0.05 (human labor) | $0.001-0.01 (API call) |
| Speed | Days (humans work 8h/day) | Seconds (API available 24/7) |
| Consistency | Variable (human disagreement) | High (same model each time) |
| Scalability | Limited by headcount | Unlimited (just more API calls) |
| For quality vs speed | Quality wins (humans are careful) | Speed wins (instant feedback) |
RLAIF is not a drop-in replacement for RLHF. RLHF with careful human evaluation can produce higher quality signals. But RLAIF is transformatively cheaper and faster, making it practical for continuous online learning loops.
One of the most important applications of CAI: ensuring agent safety and alignment with user values.
Instead of RLHF requiring humans to rate safety, you embed safety principles directly in the constitution:
Safety principle 1: Never suggest illegal activities.
Safety principle 2: Never generate personally identifiable information.
Safety principle 3: Refuse requests that ask you to help with fraud or deception.
Now, during training, the critic model is trained to critique against these principles. The agent is optimized to satisfy the constitution. This is how modern systems like Claude are aligned: not just through post-training, but through reinforcement learning guided by constitutional principles.
The code examples demonstrating AI critic ensembles, KL penalties, and reward hacking defenses are from the RL_Agents repository:
To run locally:
git clone https://github.com/Pulkit12dhingra/RL_Agents
cd RL_Agents
uv sync
jupyter notebook notebooks/post06_safety_credit.ipynb
The STaR loop shows how a policy can improve itself iteration by iteration by learning only from its own successful traces.
from __future__ import annotations
import numpy as np
import pandas as pd
class _ChainMDP:
"""5-step chain MDP. Correct action = 1. ORM reward = 1 if all correct."""
N_STEPS = 5
CORRECT_ACT = 1
def run_episode(self, policy: "_TabularPolicy") -> tuple[list[int], float]:
actions = [policy.sample(s) for s in range(self.N_STEPS)]
n_correct = sum(a == self.CORRECT_ACT for a in actions)
return actions, 1.0 if n_correct == self.N_STEPS else 0.0
class _TabularPolicy:
"""Independent softmax per state. Updated via REINFORCE."""
def __init__(self, n_states: int, n_actions: int = 2,
lr: float = 0.12, seed: int = 0) -> None:
self.logits = np.zeros((n_states, n_actions))
self.lr = lr
self.rng = np.random.default_rng(seed)
def probs(self, state: int) -> np.ndarray:
e = np.exp(self.logits[state] - self.logits[state].max())
return e / e.sum()
def sample(self, state: int) -> int:
return int(self.rng.choice(self.logits.shape[1], p=self.probs(state)))
def update(self, state: int, action: int,
reward: float, baseline: float) -> None:
advantage = reward - baseline
grad = -self.probs(state)
grad[action] += 1.0
self.logits[state] += self.lr * advantage * grad
def star_loop(
n_iterations: int = 5,
episodes_per_iter: int = 60,
seed: int = 0,
) -> pd.DataFrame:
"""
STaR outer loop on ChainMDP.
Each iteration: collect -> filter successes -> REINFORCE on successes only.
"""
env = _ChainMDP()
policy = _TabularPolicy(n_states=env.N_STEPS, lr=0.20, seed=seed)
baseline = 0.03
alpha_b = 0.20
history = []
for iteration in range(1, n_iterations + 1):
ep_records: list[tuple[list[int], float]] = []
successes = 0
for _ in range(episodes_per_iter):
actions, reward = env.run_episode(policy)
ep_records.append((actions, reward))
successes += int(reward == 1.0)
success_rate = successes / episodes_per_iter
for actions, reward in ep_records:
if reward == 1.0:
for s, a in enumerate(actions):
policy.update(s, a, reward, baseline)
baseline += alpha_b * (success_rate - baseline)
p_correct = float(policy.probs(0)[_ChainMDP.CORRECT_ACT])
history.append({
"iteration": iteration,
"successes": successes,
"success_rate": round(success_rate, 3),
"p_correct_step0": round(p_correct, 3),
})
return pd.DataFrame(history)
# Run the STaR loop
df = star_loop(n_iterations=5, episodes_per_iter=60)
print(df.to_string(index=False))
print(f"\nSuccess rate iteration 1 -> 5: "
f"{df['success_rate'].iloc[0]:.1%} -> {df['success_rate'].iloc[-1]:.1%}")
RLAIF and Constitutional AI solve a key gap: they let you scale feedback for subjective tasks without human overhead.
The training stack for a mature agent now looks like:
In the final post, we'll cover the last frontier: reward hacking and credit assignment - the unsolved challenges that still trip up agentic systems.