Q: Your mobile applications and IoT devices write millions of tiny (2 KB) raw JSON files directly into an S3 data lake every hour. Downstream Athena and Presto analytics queries grind to a halt due to the 'small files problem' (scanning 100 million tiny S3 objects). How do you design an event-driven serverless ingestion and compaction architecture that consolidates small files into 128 MB Snappy-compressed Parquet files in near-real-time?
Engineering a high-throughput, serverless event-driven data lake ingestion pipeline processing millions of raw JSON events per hour with automated small-file Parquet compaction and schema enforcement.
Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Buffer Micro-Batches via Amazon Kinesis Firehose / S3 Event Notifications
Prevent millions of individual 2 KB direct S3 writes at the edge:
- Edge Buffering: Routed incoming API event streams through Amazon Kinesis Firehose.
- Initial Buffer Thresholds: Configured buffer hints:
Interval: 300sorSize: 5MB, concatenating raw events into interim staging files ins3://data-lake-raw/landing/. - EventBridge Routing: S3 ObjectCreated events publish notifications to an Amazon EventBridge bus.
Enforce Optimal Hive Partitioning & Apache Iceberg Table Formats
Structure data layouts for maximum query scan pruning:
- Partition Scheme: Partitioned data by year, month, day, and event type:
/events/event_type=orders/year=2026/month=10/day=06/. - Apache Iceberg Metadata: Adopted Apache Iceberg table format, providing ACID transactional commits, hidden partitioning, and automated file manifest pruning.
Deploy Serverless Compactor Worker (Lambda / AWS Glue)
Compact interim small files into optimized 128 MB Parquet blocks:
- SQS Batch Queue: EventBridge routes landing notifications to an Amazon SQS FIFO queue grouped by partition path.
- Compactor Execution: When a partition accumulates 50 MB of files or 15 minutes elapsed, worker reads small raw files, converts JSON to Apache Parquet with Snappy compression, and writes a single 128 MB Parquet file to
/curated/. - Columnar Statistics: Computes min/max dictionary bounds per Parquet column, enabling Athena query pushdown.
Atomically Commit Manifests & Clean Staging via S3 Lifecycle Rules
Guarantee zero duplicate reads and automate disk reclamation:
- Atomic Swap: Iceberg commits the new compacted Parquet files into the table metadata snapshot in an atomic transaction; readers instantly see compacted data with zero query lockouts.
- S3 Lifecycle Expiration: Configured S3 Lifecycle rule automatically purging raw landing files older than 3 days:
Expiration: Days: 3. - Athena Query Benchmarks: Analytical query runtimes dropped from 4 minutes down to 3.8 seconds, while S3 API request costs plummeted by 92%.
- Buffer incoming streams in Kinesis Firehose to eliminate millions of direct 2 KB S3 PUTs.
- Organize data using Apache Iceberg table formats for ACID transactions and metadata pruning.
- Compact landing files into 128 MB Snappy Parquet files using SQS-triggered serverless workers.
- Achieve 20x faster Athena queries and slash S3 API costs by over 90%.