⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
← Back to All Platform Engineering & IDP Interview Questions Scenario 45 of 50 in Platform Engineering & IDP
Staff Platform Engineer Platform Engineering Backstage & IDP Backstage
🎯 Target Role / Context: Staff Platform Architect leading Developer Portal and enterprise Backstage adoption.

Q: How do you architect and scale the Backstage Software Catalog across an enterprise with thousands of services split between hundreds of individual repositories and giant monorepos, while overcoming GitHub API rate limits and stale catalog entities?

Engineering a high-performance, real-time Backstage Software Catalog capable of indexing tens of thousands of microservices across hundreds of repositories and giant monorepos without hitting API rate limits.

#Backstage #Software Catalog #Monorepo #IDP #DevEx #Platform Engineering
🎙️ Candidate Opening & Architectural Context
"As Backstage expands to thousands of services, relying on naive catalog discovery processors that scan git repositories every few minutes quickly exhausts GitHub/GitLab API rate limits and creates stale catalog data. Platform teams must adopt event-driven ingestion and optimized catalog providers."
Advertisement
⚡ Recommended Practice Lab

Want to master this scenario in a live sandbox? KodeKloud's CKA & CKAD Hands-On Certification Track covers this exact problem with hands-on terminal drills.

🛠️ Production Runbook & Step-by-Step Resolution

1

Migrate from Discovery Processors to Event-Driven Entity Providers

Replace polling-based catalog processors (catalog-backend-module-github) with event-driven Entity Providers (GithubEntityProvider) paired with GitHub Webhooks. Changes to catalog-info.yaml trigger immediate delta ingestion rather than periodic full-tree scans.

// catalog.ts
const githubProvider = GithubEntityProvider.fromConfig(env.config, {
  logger: env.logger,
  schedule: env.scheduler.createScheduledTaskRunner({
    frequency: { minutes: 120 },
    timeout: { minutes: 15 },
  }),
});
builder.addEntityProvider(githubProvider);
2

Handle Monorepos via Path Pattern Matching and Custom Ingestion

For large monorepos containing hundreds of services, avoid single giant catalog files. Structure nested directories with individual catalog-info.yaml files and register wildcards (target: 'https://github.com/org/monorepo/blob/main/**/catalog-info.yaml'), leveraging monorepo sparse checkouts or tree API caches in the backend.

# app-config.yaml
catalog:
  locations:
    - type: url
      target: https://github.com/org/monorepo/blob/main/**/catalog-info.yaml
      rules:
        - allow: [Component, System, API]
Advertisement
3

Implement Enterprise Database Backend and Read-Replicas

Replace default SQLite with a high-availability PostgreSQL cluster with read-replicas. Configure catalog processing queues and batching, caching GitHub API responses in Redis to withstand burst rate limits during enterprise-wide pipeline runs.

database:
  client: pg
  connection:
    host: postgres-cluster.internal
    port: 5432
    user: backstage
    password: ${POSTGRES_PASSWORD}
    ssl: { rejectUnauthorized: true }
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Scaling Backstage requires transitioning from polling processors to webhook-driven Entity Providers, backing the catalog with high-availability PostgreSQL and Redis caching, and handling monorepos with sparse path indexing to avoid GitHub API throttling."
⚡ 60-Second Elevator Pitch Talking Points
  • Polling 2,000+ repos every 5 minutes will crash Backstage and exhaust GitHub API quotas.
  • We converted our catalog architecture to event-driven GitHub Entity Providers triggered by webhooks, backed by an HA PostgreSQL cluster.
  • Monorepos are indexed using sparse path globbing, cutting catalog update latency from 30 minutes to under 5 seconds.
Advertisement
Want more Platform Engineering & IDP scenarios?
Explore our complete collection of scenario-based Platform Engineering & IDP interview runbooks.
Browse All Platform Engineering & IDP Questions →