⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All FinOps & System Design Interview Questions Scenario 95 of 98 in FinOps & System Design
Staff SRE / Financial Systems Architect System Design Financial Systems & High-Volume Metering System Design

Q: Your enterprise monetizes API calls (LLM tokens, API requests, data transfer). Every customer request must be metered, aggregated, and billed accurately. If the metering pipeline drops events, company revenue is lost; if it double-counts events, customers are illegally overcharged. How do you design an audit-compliant usage metering architecture handling 200,000 events/sec with guaranteed exactly-once billing accuracy?

Architectural design for a high-throughput, audit-compliant API usage metering and usage-based billing aggregation engine processing 200,000 API calls/sec with exactly-once aggregation in ClickHouse and Stripe integration.

#System Design #Usage Metering #Billing Engine #Kafka #ClickHouse #Stripe #FinTech
🎙️ Candidate Opening & Architectural Context
"Usage-based billing requires mathematical precision: dropping events causes revenue leakage, while duplicate billing triggers lawsuits and regulatory fines. We engineered a financial-grade usage metering platform utilizing Apache Kafka idempotent producers, ClickHouse deduplication engines, and Stripe/Lago billing integration."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? The Linux Foundation's FinOps Certified Practitioner (FOCP) Program covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Emit Cryptographically Signed Metering Events at API Gateway Edge

Capture consumption events at the ingress gateway without adding latency:

  • Edge Metering Filter: Envoy / API Gateway generates an immutable metering event upon request completion: { 'event_id': 'uuidv7', 'tenant_id': 'org_99', 'metric': 'llm_tokens', 'quantity': 450, 'timestamp': 1762391029 }.
  • Asynchronous Non-Blocking Emission: Metering events flush to local memory buffers and stream asynchronously over gRPC, adding < 0.1ms to client response latency.
Pro Tip: Generating unique, time-ordered UUIDv7 event IDs at the edge provides the cryptographic anchor required for downstream deduplication.
2️⃣

Buffer & Order Metering Events via Idempotent Apache Kafka

Prevent event loss during downstream database maintenance or traffic surges:

  • Idempotent Producers: Configured Kafka producers with enable.idempotence=true and acks=all.
  • Partition Key: Partitioned Kafka topics by tenant_id to ensure all events for a given customer are processed sequentially.
  • Durability: Retains 7 days of raw unaggregated metering events, allowing complete historical reprocessing if billing algorithms are updated.
Pro Tip: Idempotent Kafka producers guarantee that network retries between gateway pods and Kafka brokers never create duplicate messages.
3️⃣

Execute Exactly-Once Aggregation in ClickHouse ReplacingMergeTree

Deduplicate incoming events and compute real-time hourly and monthly billing totals:

  • Deduplication Engine: Raw events write to ClickHouse table using ReplacingMergeTree(event_id) ordered by (tenant_id, metric, event_id).
  • Materialized Views: Configured Materialized Views automatically aggregating consumption into 1-hour windows: SUM(quantity) GROUP BY tenant_id, metric, toStartOfHour(timestamp).
  • Sub-Second Invoicing Queries: Aggregating billions of events into invoice line items executes in < 80ms.
Pro Tip: ClickHouse ReplacingMergeTree automatically discards duplicate events with identical event_ids, providing deterministic exactly-once semantics.
4️⃣

Synchronize Aggregated Usage with Billing APIs (Stripe Metered Billing)

Push reconciled billing units to external payment gateways safely:

  • Hourly Reconciliation Worker: Background worker reads confirmed hourly aggregates and calls Stripe Usage Records API: stripe.subscriptionItems.createUsageRecord(itemId, { quantity, timestamp, action: 'set' }).
  • Action=Set Idempotency: Using action: 'set' with explicit timestamps guarantees that retrying the Stripe API call never increments the customer's bill twice.
  • Audit Reconciliation: Nightly ledger reconciliation matches raw gateway logs against Stripe invoices with 100.000% mathematical parity.
Pro Tip: Using action='set' instead of action='increment' ensures that payment gateway API retries are completely idempotent.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Usage-based API metering requires UUIDv7 event IDs emitted at the edge, idempotent Kafka buffering, ClickHouse ReplacingMergeTree for exactly-once deduplication, and idempotent Stripe synchronization."
⚡ 60-Second Elevator Pitch Talking Points
  • Emit immutable UUIDv7 metering events at the API Gateway edge with zero client latency overhead.
  • Buffer events in Kafka with acks=all and idempotence enabled, partitioned by tenant_id.
  • Deduplicate and aggregate billions of events in ClickHouse ReplacingMergeTree tables.
  • Push reconciled usage to Stripe using idempotent action='set' calls, guaranteeing zero billing discrepancies.
Advertisement
Want more FinOps & System Design scenarios?
Explore our complete collection of scenario-based FinOps & System Design interview runbooks.
Browse All FinOps & System Design Questions →