Weights alone for four assumed equal 70B PPO-RLHF models: actor, reference, reward and critic
560 GBderived
| Value kind | derived — Derived values are guide calculations; their result is only as fixed as the stated inputs, scope, and method. |
|---|---|
| Scope | Four equal 70B models at two bytes per weight; weights only. Excludes gradients, optimizer state, activations, buffers and per-device placement. Actual model sizes, sharing and residency depend on the algorithm. |
| As of | 2026-09 |
| Source | NVIDIA NeMo Framework 25.02, PPO model roles; Guide calculation assumes four equal 70B models at two bytes per weight. Version vintage retained; no exact publication day asserted. · PPO Training: actor, reference, reward and critic. Guide assumes four equal 70B parameter counts and two bytes/weight; no device allocation is inferred. |
| Derivation | 4 × 70 × 10^9 parameters × 2 bytes/parameter = 560 × 10^9 bytes = 560 GB, decimal. Model count, equal sizes and precision are exact teaching operands. |
| Review | checking…review by 2027-03-05 · standard cadence |
| Recorded changes | last 2026-09-16 · 3 revisions tracked |
| Claim id | just-to-hold-weights-for-a-70b-ppo-rlhf-stack |
Where the guide uses it
← Full numbers register — every date-stamped figure in the guide, with revision history.