Post-training
The fine-tuning and alignment stages (SFT, RLHF) after pretraining that shape a model's behavior and usefulness.
Current numbers
~80%of wall-clock spent on rollout generation in agentic/reasoning RL post-training
The fine-tuning and alignment stages (SFT, RLHF) after pretraining that shape a model's behavior and usefulness.