⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 89 of 98 in FinOps & System Design
Staff Data Platform Architect / SRE System Design Serverless & Data Lake Architecture System Design

Q: Your mobile applications and IoT devices write millions of tiny (2 KB) raw JSON files directly into an S3 data lake every hour. Downstream Athena and Presto analytics queries grind to a halt due to the 'small files problem' (scanning 100 million tiny S3 objects). How do you design an event-driven serverless ingestion and compaction architecture that consolidates small files into 128 MB Snappy-compressed Parquet files in near-real-time?

Engineering a high-throughput, serverless event-driven data lake ingestion pipeline processing millions of raw JSON events per hour with automated small-file Parquet compaction and schema enforcement.

#System Design #Data Lake #AWS S3 #EventBridge #SQS #Parquet #Compaction
🎙️ Candidate Opening & Architectural Context
"The 'small files problem' kills data lake query performance and generates massive S3 GET request costs. We designed a serverless, event-driven compaction pipeline utilizing AWS SQS, EventBridge, and AWS Lambda / Glue Iceberg compaction."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Buffer Micro-Batches via Amazon Kinesis Firehose / S3 Event Notifications

Prevent millions of individual 2 KB direct S3 writes at the edge:

  • Edge Buffering: Routed incoming API event streams through Amazon Kinesis Firehose.
  • Initial Buffer Thresholds: Configured buffer hints: Interval: 300s or Size: 5MB, concatenating raw events into interim staging files in s3://data-lake-raw/landing/.
  • EventBridge Routing: S3 ObjectCreated events publish notifications to an Amazon EventBridge bus.
Pro Tip: Buffering micro-batches at the ingestion perimeter immediately eliminates 90% of raw S3 PUT request costs.
2️⃣

Enforce Optimal Hive Partitioning & Apache Iceberg Table Formats

Structure data layouts for maximum query scan pruning:

  • Partition Scheme: Partitioned data by year, month, day, and event type: /events/event_type=orders/year=2026/month=10/day=06/.
  • Apache Iceberg Metadata: Adopted Apache Iceberg table format, providing ACID transactional commits, hidden partitioning, and automated file manifest pruning.
Pro Tip: Apache Iceberg replaces fragile Hive directory scans with deterministic metadata manifests, speeding up query planning by 10x.
3️⃣

Deploy Serverless Compactor Worker (Lambda / AWS Glue)

Compact interim small files into optimized 128 MB Parquet blocks:

  • SQS Batch Queue: EventBridge routes landing notifications to an Amazon SQS FIFO queue grouped by partition path.
  • Compactor Execution: When a partition accumulates 50 MB of files or 15 minutes elapsed, worker reads small raw files, converts JSON to Apache Parquet with Snappy compression, and writes a single 128 MB Parquet file to /curated/.
  • Columnar Statistics: Computes min/max dictionary bounds per Parquet column, enabling Athena query pushdown.
Pro Tip: Parquet compaction reduces storage size by 85% compared to raw JSON and accelerates Athena/Presto query execution speeds by 20x.
4️⃣

Atomically Commit Manifests & Clean Staging via S3 Lifecycle Rules

Guarantee zero duplicate reads and automate disk reclamation:

  • Atomic Swap: Iceberg commits the new compacted Parquet files into the table metadata snapshot in an atomic transaction; readers instantly see compacted data with zero query lockouts.
  • S3 Lifecycle Expiration: Configured S3 Lifecycle rule automatically purging raw landing files older than 3 days: Expiration: Days: 3.
  • Athena Query Benchmarks: Analytical query runtimes dropped from 4 minutes down to 3.8 seconds, while S3 API request costs plummeted by 92%.
Pro Tip: Atomic metadata commits guarantee that analytical queries reading the table never encounter duplicate rows or partial read states during compaction.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Solving the data lake small files problem requires ingestion buffering, Apache Iceberg table formats, SQS-triggered serverless Parquet compaction to 128 MB chunks, and atomic metadata snapshot commits."
⚡ 60-Second Elevator Pitch Talking Points
  • Buffer incoming streams in Kinesis Firehose to eliminate millions of direct 2 KB S3 PUTs.
  • Organize data using Apache Iceberg table formats for ACID transactions and metadata pruning.
  • Compact landing files into 128 MB Snappy Parquet files using SQS-triggered serverless workers.
  • Achieve 20x faster Athena queries and slash S3 API costs by over 90%.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →