BF16 · Brain Float 16
A 16-bit floating-point format with a wide exponent range, a common default for stable AI training.
Current numbers
~14 bytes/paramstated BF16 weights + FP32 master and two FP32 Adam moments; unique serialized state, before implementation metadata or replication
40 GiB aggregate; 5 GiB/rank at ideal TP=8Llama 3.1 70B logical BF16 KV at exactly 131,072 tokens; ideal eight-way head sharding, before overhead
504 KiB/token (516.096 kB/token)Llama 3.1 405B aggregate logical BF16 KV per retained token, before layout or replication