RL Agents Part 5: Constitutional AI and RLAIF

← Back to Home

So far we've covered RLVR - training agents using verifiable rewards from test suites, exact-match correctness checks, and tool execution. But many agent tasks don't have clear verifiers. A creative writing task, a strategic recommendation, or a nuanced conversational response can't be reduced to "pass/fail". For these tasks, we return to the human feedback problem. But instead of having humans rate outputs directly, we can use a strong AI model as a critic. This is RLAIF and Constitutional AI.

Code repo: RL_Agents

Notebook: post06_safety_credit.ipynb

1. The Problem: Tasks Without Clear Verifiers

RLVR works when correctness is objective. But consider these tasks:

For these tasks, we need a scoring function that can handle subjective qualities. We can't write a test suite. We need judgment.

The naive solution: hire more human raters. But this returns us to the cost and latency problems we were trying to solve with RLVR.

2. RLAIF: Using AI as a Critic

RLAIF (Reinforcement Learning from AI Feedback) replaces human raters with another AI system. Instead of asking humans "which response is better?", you ask a strong language model to compare two responses using explicit criteria.

The Basic Pipeline:
  1. Generate candidates: Use your training model to generate multiple responses to a prompt.
  2. Score with a critic LLM: Send both the prompt and responses to a strong critic model (e.g., GPT-4, Claude) with instructions like: "Compare these two responses on accuracy, clarity, and safety. Which is better?"
  3. Get preference labels: The critic returns a judgment (A > B, B > A, or tie).
  4. Train a reward model: Use these critiques to train a smaller reward model (as in RLHF).
  5. Fine-tune your agent: Use that RM to optimize your training model via GRPO or PPO.

The advantage: you've replaced human raters with API calls to a strong model. The cost is per-call instead of per-human, and it scales to any size dataset you want.

Example: Evaluating Customer Support Responses

Your agent is trained to respond to customer service inquiries. Instead of hiring 10 human raters to judge 10,000 responses, you can:

Prompt: "Customer asked: How do I reset my password?"

Response A (from your model):
"Click the login page, then click forgot password, and check your email."

Response B (also from your model):
"You can reset by clicking forgot password on the login page. 
Check your email for a reset link. If you don't see it, 
check spam. Let me know if you need more help."

Critic evaluation:
"Response B is better because it:
- Anticipates the user might miss the email (spam)
- Offers continued support
- Is more friendly in tone
Preference: B > A"

Score: B gets higher reward
  

This entire evaluation is automated. You can do it at scale without human overhead.

3. Constitutional AI: Critiques by Principle

Constitutional AI (CAI) extends RLAIF by formalizing what the critic should be looking for. Instead of asking "which is better?" (vague), you ask questions based on a constitution - a set of explicit principles.

Example Constitution for a Support Agent:

Constitutional principle 1: Responses must be accurate. If you're unsure, say so.

Constitutional principle 2: Responses must be helpful to the customer's actual problem.

Constitutional principle 3: Responses must be respectful and professional in tone.

Constitutional principle 4: Responses must not encourage actions that could harm the customer.

Now, when the critic evaluates two responses, it explicitly checks them against these principles:

Critic evaluation (constitutional):

Response A: "Click forgot password"
- Accuracy: OK, but incomplete (score: 6/10)
- Helpfulness: Low, lacks detail (score: 4/10)
- Tone: Neutral, not engaging (score: 5/10)
- Safety: OK (score: 8/10)
- Overall: C-grade response

Response B: "Click forgot password... email... spam folder... offer help"
- Accuracy: OK (score: 8/10)
- Helpfulness: High, anticipates issues (score: 9/10)
- Tone: Friendly, professional (score: 8/10)
- Safety: OK (score: 8/10)
- Overall: A-grade response

Preference: B > A
  

By making the evaluation criteria explicit (the constitution), you get:

4. Critiquing Agent Trajectories, Not Single Outputs

So far we've talked about evaluating single responses. But for agents performing multi-step tasks, you can extend this to critique full trajectories:

Agent trajectory for: "Solve this system of equations"

Step 1: Wrote the equations in matrix form ✓
Step 2: Computed the determinant ✓
Step 3: Applied Cramer's rule ✓
Step 4: Got x=2, y=3 ✓
Step 5: Verified by substitution ✓

Constitutional evaluation:
- Did the agent follow a logical process? YES (score: 10/10)
- Were all steps mathematically correct? YES (score: 10/10)
- Could a student learn from this? YES (score: 9/10)
- Overall: Excellent trajectory, give high reward
  

Now you're using a critic to score the entire reasoning path, not just the final answer. This is closer to PRM (process rewards) but using AI critique instead of hard verifiers.

5. Combining RLVR and RLAIF: The Best of Both

The most sophisticated agent training uses both:

Example: training a research assistant agent

By mixing verifiers (objective) with critic evaluations (subjective), you get the benefits of both: scalability and principled judgment.

6. The Economics of RLAIF vs RLHF

Dimension RLHF RLAIF
Cost per evaluation $0.02-0.05 (human labor) $0.001-0.01 (API call)
Speed Days (humans work 8h/day) Seconds (API available 24/7)
Consistency Variable (human disagreement) High (same model each time)
Scalability Limited by headcount Unlimited (just more API calls)
For quality vs speed Quality wins (humans are careful) Speed wins (instant feedback)

RLAIF is not a drop-in replacement for RLHF. RLHF with careful human evaluation can produce higher quality signals. But RLAIF is transformatively cheaper and faster, making it practical for continuous online learning loops.

7. Alignment and Safety Through Constitutional AI

One of the most important applications of CAI: ensuring agent safety and alignment with user values.

Instead of RLHF requiring humans to rate safety, you embed safety principles directly in the constitution:

Safety principle 1: Never suggest illegal activities.

Safety principle 2: Never generate personally identifiable information.

Safety principle 3: Refuse requests that ask you to help with fraud or deception.

Now, during training, the critic model is trained to critique against these principles. The agent is optimized to satisfy the constitution. This is how modern systems like Claude are aligned: not just through post-training, but through reinforcement learning guided by constitutional principles.

8. Complete Code Reference

The code examples demonstrating AI critic ensembles, KL penalties, and reward hacking defenses are from the RL_Agents repository:

To run locally:

git clone https://github.com/Pulkit12dhingra/RL_Agents
 cd RL_Agents
 uv sync
 jupyter notebook notebooks/post06_safety_credit.ipynb
   

STaR: Self-Taught Reasoner in Action

The STaR loop shows how a policy can improve itself iteration by iteration by learning only from its own successful traces.

from __future__ import annotations
import numpy as np
import pandas as pd


class _ChainMDP:
    """5-step chain MDP. Correct action = 1. ORM reward = 1 if all correct."""
    N_STEPS     = 5
    CORRECT_ACT = 1

    def run_episode(self, policy: "_TabularPolicy") -> tuple[list[int], float]:
        actions   = [policy.sample(s) for s in range(self.N_STEPS)]
        n_correct = sum(a == self.CORRECT_ACT for a in actions)
        return actions, 1.0 if n_correct == self.N_STEPS else 0.0


class _TabularPolicy:
    """Independent softmax per state. Updated via REINFORCE."""
    def __init__(self, n_states: int, n_actions: int = 2,
                 lr: float = 0.12, seed: int = 0) -> None:
        self.logits = np.zeros((n_states, n_actions))
        self.lr     = lr
        self.rng    = np.random.default_rng(seed)

    def probs(self, state: int) -> np.ndarray:
        e = np.exp(self.logits[state] - self.logits[state].max())
        return e / e.sum()

    def sample(self, state: int) -> int:
        return int(self.rng.choice(self.logits.shape[1], p=self.probs(state)))

    def update(self, state: int, action: int,
               reward: float, baseline: float) -> None:
        advantage     = reward - baseline
        grad          = -self.probs(state)
        grad[action] += 1.0
        self.logits[state] += self.lr * advantage * grad


def star_loop(
    n_iterations: int = 5,
    episodes_per_iter: int = 60,
    seed: int = 0,
) -> pd.DataFrame:
    """
    STaR outer loop on ChainMDP.
    Each iteration: collect -> filter successes -> REINFORCE on successes only.
    """
    env      = _ChainMDP()
    policy   = _TabularPolicy(n_states=env.N_STEPS, lr=0.20, seed=seed)
    baseline = 0.03
    alpha_b  = 0.20
    history  = []

    for iteration in range(1, n_iterations + 1):
        ep_records: list[tuple[list[int], float]] = []
        successes = 0
        for _ in range(episodes_per_iter):
            actions, reward = env.run_episode(policy)
            ep_records.append((actions, reward))
            successes += int(reward == 1.0)

        success_rate = successes / episodes_per_iter

        for actions, reward in ep_records:
            if reward == 1.0:
                for s, a in enumerate(actions):
                    policy.update(s, a, reward, baseline)

        baseline += alpha_b * (success_rate - baseline)
        p_correct = float(policy.probs(0)[_ChainMDP.CORRECT_ACT])

        history.append({
            "iteration":       iteration,
            "successes":       successes,
            "success_rate":    round(success_rate, 3),
            "p_correct_step0": round(p_correct, 3),
        })

    return pd.DataFrame(history)


# Run the STaR loop
df = star_loop(n_iterations=5, episodes_per_iter=60)
print(df.to_string(index=False))
print(f"\nSuccess rate iteration 1 -> 5: "
      f"{df['success_rate'].iloc[0]:.1%} -> {df['success_rate'].iloc[-1]:.1%}")
  

Summary: Bridging Verifiable and Subjective Tasks

RLAIF and Constitutional AI solve a key gap: they let you scale feedback for subjective tasks without human overhead.

The training stack for a mature agent now looks like:

In the final post, we'll cover the last frontier: reward hacking and credit assignment - the unsolved challenges that still trip up agentic systems.

Next in the series: Part 6 covers reward hacking, credit assignment, and the open challenges in agentic RL.
← Part 4 Part 6 →