Q: Running 100 hyperparameter tuning trials sequentially or via brute-force grid search consumes hundreds of thousands of dollars in GPU time. How do you architect Kubeflow Katib with Bayesian Optimization, Median Stopping early-termination rules, and elastic spot GPU node pools?
Operating Kubeflow Katib for automated large-scale hyperparameter optimization on Kubernetes, implementing Bayesian search algorithms, Median Stopping early termination, and elastic GPU pool scheduling.
Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Define Katib Experiment CRD with Bayesian Optimization
Declare a Katib `Experiment` custom resource. Select advanced search algorithms such as Bayesian Optimization (`algorithmName: bayesianoptimization`) or Hyperband (`algorithmName: hyperband`). These algorithms model the objective function probabilistically, choosing subsequent hyperparameter trials that maximize expected improvement rather than sampling randomly.
apiVersion: kubeflow.org/v1beta1
kind: Experiment
metadata:
name: bert-finetune-hpo
spec:
algorithm:
algorithmName: bayesianoptimization
parallelTrialCount: 8
maxTrialCount: 64
maxFailedTrialCount: 3
objective:
type: maximize
goal: 0.95
objectiveMetricName: val_accuracy
Configure Early Stopping via Median Stopping Rule
Configure early stopping rules (`earlyStopping: { algorithmName: medianstop }`). The Katib controller monitors intermediate validation metric reports emitted by active trials. If a trial's performance at epoch $E$ is worse than the median performance of previous trials at the same epoch, Katib immediately terminates the trial pod, freeing the GPU for the next experiment.
earlyStopping:
algorithmName: medianstop
algorithmSettings:
- name: min_trials_resource
value: "4"
- name: start_step
value: "5"
Target Elastic Spot GPU Node Pools via Trial Pod Templates
Configure the trial template to schedule onto elastic spot GPU node pools managed by Karpenter. Because Katib trials are stateless and independent, spot interruptions only terminate individual trials (which Katib automatically reschedules), enabling teams to run large-scale HPO sweeps at 70% lower cloud cost.
trialTemplate:
primaryContainerName: training-container
trialParameters:
- name: learningRate
reference: lr
trialSpec:
apiVersion: batch/v1
kind: Job
spec:
template:
spec:
nodeSelector:
karpenter.sh/capacity-type: spot
- Brute-force grid search wastes massive GPU budgets letting failing experiments run to completion.
- We deploy Kubeflow Katib with Bayesian optimization to intelligently discover optimal parameters in 75% fewer trials.
- By enabling Median Stopping rules and running trials on elastic Spot GPU pools, underperforming experiments are terminated in the first 5 epochs, cutting our HPO bill by over 80%.