Q: Your monorepo has grown to 25 GB and 10 years of history. New hires take 45 minutes to clone, IDE indexing is slow, and most engineers only need ~5% of the tree. What Git-side techniques would you use, and where do they fall short?
Three orthogonal techniques solve three different problems:
🛠️ Production Runbook & Step-by-Step Resolution
Initial Diagnostics & Root Cause Analysis
Three orthogonal techniques solve three different problems:
- Partial clone — defers downloading file blobs until they're actually accessed:
- Sparse-checkout — limits which paths exist in your working tree:
- Git LFS for large binaries — design files, ML model artifacts, video. LFS stores a small pointer file in Git and the actual blob on a separate object store, fetched on demand. Adopt this *going forward*; converting historical large blobs requires
git lfs migrate importplus a force-push, which rewrites history. - Partial clone makes operations like
git log -p,git blame, and bisect across many files trigger on-demand fetches that can be slow or fail offline. Engineers on flaky networks suffer.
Remediation & Permanent Safeguards
Fetches all commits and trees but no file contents. When you checkout or log -p a path, Git lazily fetches just those blobs. Cuts initial clone from 25 GB to a few hundred MB. Requires a server that supports it (GitHub, GitLab, Azure DevOps, Gitea all do). You still have the full history, but only the directories you listed materialize on disk. IDE indexing now sees 5% of the tree. Combine with partial clone for maximum effect (git clone --filter=blob:none --sparse ). Where they fall short: The pragmatic recipe for most teams: sparse + partial clone via a git clone wrapper script for new hires, LFS adopted for any binary > a few MB, and a "no committing build artifacts" lint in CI.
git clone --filter=blob:none <url>
- Sparse-checkout doesn't help operations that *traverse* the repo —
git grepover a sparse checkout misses files you didn't materialize, which can be surprising. Cone mode mitigates but doesn't eliminate this. - LFS adds operational dependency on the LFS server, increases hosting cost, and storage isn't free or fast for huge blobs. It's not a silver bullet — for truly large binary pipelines (e.g. game asset workflows), Perforce or a content-addressed object store is sometimes a better fit than Git.
- None of these address the *root* problem on Git itself: extremely deep histories on a few hot files (think: lockfiles touched by everyone) can still slow
blameandlog. At true Google/Microsoft scale you eventually outgrow vanilla Git and look at VFS for Git, Scalar, or Piper-style virtual filesystems.
- Partial clone — defers downloading file blobs until they're actually accessed:
- Sparse-checkout — limits which paths exist in your working tree:
- Git LFS for large binaries — design files, ML model artifacts, video. LFS stores a small pointer ...