Q: A new firmware build pipeline starts failing during artifact uploads. S3 shows no error, Jenkins logs are incomplete. Where do you start?
Triage runbook for resolving high-value firmware build pipeline upload failures where AWS S3 shows no error logs and Jenkins output is truncated.
Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Check Linux Kernel dmesg for OOM-Killer Execution
When Jenkins logs truncate abruptly without a stack trace, the process was killed violently by `SIGKILL` (signal 9) rather than exiting gracefully. Check `dmesg` or system logs on the build node for `Out of memory: Killed process`.
# Check for OOM killer termination on Jenkins agent
dmesg -T | grep -i -E 'oom[-_]killer|killed process'
# If on Kubernetes runner:
kubectl describe pod <jenkins-agent-pod> | grep -i 'OOMKilled'
Inspect Agent Ephemeral Disk Space & Temp Inodes
The AWS CLI and SDKs buffer large multi-part uploads in the local temporary directory (`/tmp`) before splitting into 5 MB chunks. If the agent disk hits 100%, the process crashes instantly.
df -h /tmp
df -i /tmp
Enable Verbose AWS CLI Debug Telemetry
Re-run the upload command with `--debug`. This exposes exact TCP socket timeouts, SSL handshake failures, and IAM STS token expiration during long-running uploads.
aws s3 cp firmware.bin s3://tesla-firmware-builds/ --debug > upload_debug.log 2>&1
Tune Multi-Part Upload Concurrency and Chunk Size
Adjust AWS S3 CLI configurations to handle massive firmware binaries safely: Increase `multipart_chunksize` to `64MB` and tune `max_concurrent_requests` to avoid saturating agent memory and file descriptors.
aws configure set default.s3.multipart_threshold 64MB
aws configure set default.s3.multipart_chunksize 64MB
aws configure set default.s3.max_concurrent_requests 10
- Check dmesg -T on the agent for Linux OOM-killer terminating the upload process with SIGKILL.
- Verify local agent /tmp disk space and inode availability.
- Execute the upload with the aws s3 --debug flag to uncover socket resets or expired STS tokens.
- Tune AWS S3 multipart_chunksize to 64MB and control concurrency to reduce memory overhead.