# Common Pitfalls in RL Post-Training ## Table of Contents - Log-Probability Pitfalls - Advantage Estimation Pitfalls - Generation and Decoding Pitfalls - Reward Function Pitfalls - Gradient and Optimization Pitfalls - Configuration Pitfalls --- ## Log-Probability Pitfalls ### Sign Errors in Log-Space **Pattern:** A subtraction or negation in log-space is reversed, so quantities that should be non-positive come out positive (or vice versa). The classic offenders are log-softmax, the DPO log-ratio `log pi(chosen) - log pi(rejected)`, and KL divergence `E[log p - log q]`. **Symptoms:** - A quantity you know must be bounded (e.g. log-prob, KL) violates its bound - Loss still decreases but the policy moves *away* from high-reward completions - Ratios in PPO-style updates have the wrong scale **How to detect:** - Bound checks (e.g. `log_probs.max() <= 0`) are the fastest signal - Compare your manual implementation against the framework reference (`F.log_softmax`, `F.kl_div`) on a small deterministic input - Spot-check the sign on a near-one-hot input where the answer is predictable **Why it happens:** Log-space identities are easy to write down in either direction. When variable names don't encode the polarity (`a - b` vs `b - a` looks symmetric in code review), the reversed form passes linting and tests that only check shape. ### Incorrect Gathering Along the Vocabulary Dimension **Pattern:** `gather` is called on the wrong axis, or the index tensor is not unsqueezed to match rank. The call still runs, but indexes into the sequence dimension instead of vocab. **Symptoms:** Values are plausible-looking but uncorrelated with the intended tokens. Training is slow or unstable with no obvious bug. **How to detect:** Compare against the explicit `log_softmax(...).gather(-1, idx.unsqueeze(-1)).squeeze(-1)` on a tiny input. ### Precision Loss in Half-Precision **Pattern:** `logsumexp` computed in float16 or bfloat16 overflows or loses precision when logits span a large range. **Symptoms:** Sporadic NaNs, especially on longer sequences or with sharper distributions after a few training steps. **How to detect:** Cast logits to float32 before the logsumexp step and compare; differences > 0.01 indicate precision issues. --- ## Advantage Estimation Pitfalls ### Epsilon Dominates the Denominator **Pattern:** The numerical-stability epsilon added to `std` is large enough to wash out the reward-normalized signal. Epsilon should be a rounding guard for `std == 0`, not a real additive term. **Common ways this happens:** - Typo in scientific notation (positive exponent where a negative was intended) - Copy-pasting a constant from an unrelated context (gradient clipping threshold, softmax temperature, reward clipping bound) - Hyperparameter sweep that includes `epsilon=1.0` as a "neutral" default **Symptoms:** - Advantages are near-zero in magnitude even though rewards clearly vary across the group - Flat loss curve — the policy gradient sees no signal - Model is functionally unchanged after many steps **How to detect:** ```python # 1. Sanity-check the constant itself assert epsilon < 1.0 and epsilon > 0, f"epsilon={epsilon} out of plausible range" # 2. Sanity-check the output when rewards are obviously non-constant rewards = torch.tensor([1.0, 0.0, 0.5, 0.0]) advantages = compute_advantages(rewards) # whatever your pipeline calls assert advantages.abs().max() > 0.1, "advantages vanishing despite varied rewards" ``` ### Wrong Grouping Axis **Pattern:** `rewards.view(-1, G)` is used with the wrong `G`, or the reshape mixes different prompts into the same group. The `mean` and `std` then correspond to no well-defined baseline. **Symptoms:** Advantages in a group don't sum to roughly zero. Training noisier than expected. ### Broadcast Omitted **Pattern:** Per-group statistics are computed correctly (`[B]`) but not broadcast back to the `[B*G]` completion axis before subtraction. **Symptoms:** Shape error, or silent broadcasting that scrambles the group-to-completion mapping. --- ## Generation and Decoding Pitfalls ### Reasoning-Model Completion Cases Reasoning-tuned models can emit completions in several shapes, and a decoder sitting between the model and the reward function has to decide how each maps to the **completion** text the reward function will score. Handling reasoning-model outputs can be convoluted (e.g. TRL's `decode_and_strip_padding`, vLLM's `ReasoningParser`). | # | Shape | Intended return | |---|-------|-----------------| | 1 | Plain text, no reasoning markers — a direct answer, no prelude | The text unchanged (after padding strip) | | 2 | Completed reasoning followed by answer (e.g. `...answer`) | Text after the closing marker, stripped — the answer portion only | | 3 | Incomplete reasoning block (e.g. `...` with no close) | Empty string — no scoreable answer was produced | | 4 | Reasoning without a separable answer | Derivative of 1-3: `...` with nothing after → empty (case 2's split-and-strip); unclosed `...` → empty (case 3); reasoning-sounding prose with no markers → unchanged (case 1) | | 5 | Nested or repeated reasoning markers | `split("", 1)[-1]` keeps everything after the *first* closing marker — e.g. `amidbend` → `"midbend"`. Inner markers are not re-stripped | Three distinct behaviors underlie all five cases: **pass-through**, **extract after the closing marker**, and **empty**. Cases 1-3 each map cleanly to one behavior; cases 4 and 5 are composites whose output falls out of how the implementation handles the three primary cases. ### Padding Token Contamination **Pattern:** Padding tokens are not stripped, so the reward function scores text that includes runs of `` or `<|endoftext|>`. **Symptoms:** Rewards systematically lower than a manual spot-check would predict. Decoded text visibly contains pad markers. ### Special-Token Over-Removal **Pattern:** `skip_special_tokens=True` strips tokens that were registered as special but are semantically load-bearing (custom answer tags added to the vocabulary, domain-specific delimiters). **Symptoms:** The reward function's regex or parser can't find the format it expects, even though the model is producing it. --- ## Reward Function Pitfalls ### Reward/Decoding Format Mismatch **Pattern:** The reward function expects a specific shape (e.g. `...`, a last-line number, a JSON object) but the decoded text has been modified upstream so that shape no longer appears. **Symptoms:** Reward is pinned to its minimum value even for completions that look correct to a human reader. ### Scalar vs. Tensor Return Type **Pattern:** Reward function returns a Python `float` or `list[float]`, but the trainer expects a tensor of a specific shape/dtype. **Symptoms:** Runtime type error, or silent broadcasting that produces per-token rewards when per-sequence was intended. ### Sparse Binary Rewards with Small Groups **Pattern:** Reward is `{0, 1}` with no shaping, combined with small `G`. Most groups are all-zero or all-one, so advantages collapse to zero. **Symptoms:** Learning is slow or absent even though the reward function is correct. Increasing `G`, adding reward shaping, or adding a small continuous component resolves it. --- ## Gradient and Optimization Pitfalls ### Reference Model Not Frozen **Pattern:** The reference model (used for the KL term and, in some setups, `log_pi_old`) is constructed by reference or copy but its parameters are left with `requires_grad=True`. The optimizer then updates it alongside the policy. **Symptoms:** - KL divergence stays pinned near zero across training — the "reference" drifts with the policy - The intended regularization vanishes; training diverges or produces degenerate outputs - Memory usage is higher than expected because reference-model gradients are materialized **How to detect:** ```python assert all(not p.requires_grad for p in ref_model.parameters()), \ "reference model has trainable parameters" # And confirm the reference model is a distinct object, not an alias: assert ref_model is not policy_model ``` ### Missing `.detach()` on Reference Outputs **Pattern:** Reference model outputs are used inside the loss without `.detach()` or a `torch.no_grad()` context. Gradients flow through the reference graph even though its parameters are frozen, inflating memory and potentially interacting badly with accumulated graphs across steps. **Symptoms:** Out-of-memory errors that scale with sequence length more than expected; slower-than-expected backward passes. **How to detect:** Wrap reference computation in `with torch.no_grad():` and confirm loss is identical (it should be, if detach was intended). ### In-Place Ops Breaking Autograd **Pattern:** Masking or padding via in-place assignment (`tensor[mask] = 0`, `tensor.mul_`) on a tensor that's part of the autograd graph. **Symptoms:** `RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation`. Usually surfaces on the first `.backward()` call. **How to detect:** Run with `torch.autograd.set_detect_anomaly(True)` to localize the offending op. ### Loss Not Connected to Policy Parameters **Pattern:** A computation in the loss path uses `.item()`, `.tolist()`, or a numpy round-trip that breaks the gradient graph. The loss is numerically correct but `loss.backward()` produces zero gradients for the policy. **Symptoms:** Gradient norm is zero or near-zero. Loss is non-trivial but the policy doesn't change. **How to detect:** ```python total_norm = sum(p.grad.norm().item() ** 2 for p in model.parameters() if p.grad is not None) ** 0.5 assert total_norm > 1e-8, "policy gradients are vanishing" ``` --- ## Configuration Pitfalls ### SFT-Scale Learning Rate **Pattern:** A learning rate appropriate for supervised fine-tuning (e.g. 1e-5 to 5e-5) is reused for RL post-training, where stable updates typically require 1e-7 to 5e-6. **Symptoms:** Loss and KL oscillate wildly. Model outputs degrade into gibberish within a handful of steps. **How to detect:** Log KL per step. If it exceeds typical bounds (e.g. > 1.0) within the first few steps, LR is likely too high. ### KL Coefficient Dominates the Loss **Pattern:** `beta` is large enough that the KL term swamps the policy gradient term. The optimizer minimizes loss by keeping the policy identical to the reference. **Symptoms:** Policy outputs are indistinguishable from the reference model. Advantage-weighted loss contribution is small compared to the KL contribution. **How to detect:** Log the two loss components separately. If the KL term is consistently ≥ 10× the policy-gradient term, `beta` is too large. ### Clip Range Too Narrow **Pattern:** The PPO-style clip epsilon is tight enough (e.g. 0.01-0.05) that almost every ratio is clipped, so most samples contribute zero gradient. **Symptoms:** Training proceeds but barely progresses. Fraction-of-clipped-samples is near 1.0. ### `num_generations` Too Small **Pattern:** `G ≤ 2`. With `G=1`, `std` is undefined; with `G=2`, `std` is extremely noisy and reward-normalized advantages are near-random. **Symptoms:** Training is unstable even with correct epsilon, reward, and log-prob code.