Q: How does Triton Inference Server's dynamic batcher aggregate asynchronous individual inference requests, and how do you calculate optimal values for `max_batch_size`, `max_queue_delay_microseconds`, and `instance_group` counts under a 50ms p99 SLA?
Tuning NVIDIA Triton Inference Server dynamic batching algorithms, queue delay windows, and instance group concurrency to maximize GPU throughput while satisfying strict p99 latency SLAs.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Configure Dynamic Batching and Queue Delays
In the model configuration (`config.pbtxt`), declare `dynamic_batching`. Set `max_batch_size` based on GPU memory limits and model latency scaling curves. Configure `max_queue_delay_microseconds`: this specifies how long Triton will hold an incoming request to aggregate additional queries into the current batch before dispatching execution to the GPU engine.
# config.pbtxt
name: "resnet50_tensorrt"
platform: "tensorrt_plan"
max_batch_size: 64
dynamic_batching {
preferred_batch_size: [ 8, 16, 32, 64 ]
max_queue_delay_microseconds: 5000
}
Tune Instance Groups for Concurrent Model Execution
A single GPU can host multiple execution instances of the same model concurrently. If a model does not saturate the GPU's Streaming Multiprocessors (SMs) or has CPU-bound pre/post-processing steps, configuring `instance_group [ { count: 2, kind: KIND_GPU } ]` allows Triton to run two batches simultaneously, increasing GPU utilization.
instance_group [
{
count: 2
kind: KIND_GPU
gpus: [ 0 ]
}
]
Benchmark Optimal Parameters via Triton Model Analyzer
Never guess batching parameters. Execute `model-analyzer profile` across ranges of batch sizes (1 to 64), concurrency levels (1 to 128), and queue delays (1000us to 10000us). Model Analyzer sweeps configurations against a synthetic client load, automatically identifying the exact pareto-optimal configuration that maximizes inferences/sec while keeping latency under the specified SLA constraint.
# Model Analyzer run
model-analyzer profile \
--model-repository /models \
--profile-models resnet50_tensorrt \
--latency-budget 50 \
--output-model-repository-path /analyzed_models
- Individual client requests underutilize GPU compute cores. Triton dynamic batching combines them into optimal batches on the fly.
- We tune `max_queue_delay_microseconds` to 5ms: Triton waits up to 5 milliseconds to form full batches, amortizing kernel launch overhead.
- By profiling with Model Analyzer and configuring multiple instance groups per GPU, we increased throughput by 4.2x while keeping p99 latency well under our 50ms SLA.