Chain-of-thought · CoT
Prompting or training a model to reason in explicit intermediate steps, trading more tokens for better answers.
Current numbers
~70%of compute consumed by rollouts at 16K-token generation length (RLVR long-CoT)
Prompting or training a model to reason in explicit intermediate steps, trading more tokens for better answers.