FP16 · 16-bit floating point
A half-precision number format; it is the usual reference rung of the precision ladder, since each step down from it roughly doubles a chip's headline throughput.
Current numbers
3.5x / 1.8xNVFP4 memory reduction vs FP16 / vs FP8
~$2.50/M tokmarket-average self-hosted inference cost, fell ~$10→~$2.50 in a year; worked example ~$1.90/M (8xH100, Llama-70B FP16)
~$1.90/M tokself-hosted inference (8x H100 @ ~$19.20/hr, Llama-70B FP16); market avg fell ~$10 → ~$2.50/M in a year