RL · Reinforcement Learning
Training that optimizes a model against a reward signal rather than labeled data; post-training RL mixes training and inference and loads clusters in spiky patterns.
Current numbers
~80%of wall-clock spent on rollout generation in agentic/reasoning RL post-training
10K–100K+tokens per RL trajectory for reasoning/agentic tasks — the rollout that dominates cost
2.5xwall-clock speedup of variance-controlled async RL vs synchronous at equal accuracy (~42h vs ~105h)