Q: How do you choose between Triton Model Ensembles and Business Logic Scripting (BLS) when chaining tokenizers, embedding models, LLMs, and post-processing rankers, and how does BLS eliminate inter-process serialization overhead?
Designing multi-stage machine learning inference pipelines in Triton Inference Server using zero-copy Model Ensembles and dynamic Python Business Logic Scripting (BLS).
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Understand Model Ensembles and Zero-Copy Tensor Routing
Triton Model Ensembles define a static directed acyclic graph (DAG) in `config.pbtxt`. Tensors flow directly between model outputs and inputs in-memory without leaving Triton or passing through Python runtimes. Ensembles are ideal for fixed, deterministic workflows (e.g., Image Preprocess -> TensorRT CNN -> Softmax Classification) and offer maximum possible throughput with zero serialization overhead.
# Ensemble config.pbtxt routing snippet
platform: "ensemble"
ensemble_scheduling {
step [
{
model_name: "preprocess"
model_version: -1
input_map { key: "RAW_IMAGE" value: "IMAGE" }
output_map { key: "PREPROCESSED_IMAGE" value: "PREPROCESSED" }
},
{
model_name: "resnet50_trt"
model_version: -1
input_map { key: "INPUT" value: "PREPROCESSED" }
output_map { key: "OUTPUT" value: "PROBABILITIES" }
}
]
}
Implement Business Logic Scripting (BLS) for Dynamic Workflows
When pipelines require dynamic conditional logic (e.g., branching based on classifier confidence, looping during iterative token generation, or making external vector DB calls), Model Ensembles are insufficient. Triton BLS runs inside the Python backend and allows executing sub-models asynchronously (`pb_utils.InferenceRequest`) while keeping tensors in shared memory.
# BLS model.py snippet
import triton_python_backend_utils as pb_utils
class TritonPythonModel:
def execute(self, requests):
responses = []
for request in requests:
# Execute embedding sub-model
infer_request = pb_utils.InferenceRequest(model_name="embedder", inputs=[...])
infer_response = infer_request.exec()
# Branching logic in Python
if confidence > 0.8:
# Call fast model
else:
# Call large model
responses.append(pb_utils.InferenceResponse(...))
return responses
Optimize Shared Memory and IPC Between BLS and GPU Engines
By default, Triton Python backend runs in a separate process and communicates via POSIX shared memory. For high-throughput pipelines, configure Triton system shared memory or CUDA shared memory, allowing BLS Python scripts to pass GPU tensor pointers directly to TensorRT/C++ models without copying data back and forth between host RAM and GPU VRAM.
# config.pbtxt for Python backend enabling GPU shared memory
parameters: {
key: "force_gpu_memory_type"
value: { string_value: "GPU" }
}
- Chaining multi-step models through REST microservices adds tens of milliseconds of HTTP latency and JSON serialization overhead.
- Triton Ensembles route tensors directly between models in-memory with zero CPU overhead for static DAGs.
- For dynamic pipelines requiring conditional branching or custom logic, we use BLS (Business Logic Scripting) with CUDA shared memory, passing GPU memory pointers directly between models.