⚡ ~/naveed Interview Prep
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 998+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Staff SRE / Principal Architect [L3] Git Pull updates from an incremental bundle: Staff SRE Scenario [L3]

Q: Explain how Git stores data internally. What are blobs, trees, commits, and tags at the object level?

Git is fundamentally a content-addressable filesystem. Every piece of data is stored as an object identified by its SHA-1 (or SHA-256) ha...

#Git #Pull updates from an incremental bundle: #L3 #Version Control #Collaboration
🎙️ Candidate Opening & Architectural Context
""During a major release branch cut, we encountered this exact scenario and used Git internals to recover cleanly. The interviewer is testing: Deep understanding of Git's content-addressable object model.. I structure my answer around systematic triage first, root cause analysis second, and permanent remediation third.""
Advertisement

🛠️ Production Runbook & Step-by-Step Resolution

1️⃣

Initial Diagnostics & Root Cause Analysis

Git is fundamentally a content-addressable filesystem. Every piece of data is stored as an object identified by its SHA-1 (or SHA-256) hash. There are four object types:

  • Blob — stores file content (just the raw bytes, no filename or permissions). Two files with identical content share the same blob object, regardless of filename or location. This is how Git deduplicates storage.
  • Tree — represents a directory. Contains entries mapping filenames → blob SHAs (for files) or other tree SHAs (for subdirectories), along with file mode (permissions). A tree is a snapshot of a directory at a point in time.
  • Commit — points to one tree (the root tree of the project at that moment) and zero or more parent commits. Contains author, committer, timestamp, and message. The first commit has no parent; merge commits have two or more parents. The commit's SHA is a hash of all this — changing anything (message, author, parent, tree) produces a different SHA.
2️⃣

Remediation & Permanent Safeguards

How a commit represents a full snapshot: Each commit stores a complete snapshot, not a diff. Git computes diffs on the fly by comparing two commits' trees. This is why checkout is fast (just materialize one tree) and why branches are cheap (a branch is a 41-byte file containing a commit SHA). Packfiles: for efficiency, Git periodically packs loose objects into packfiles (.pack + .idx), using delta compression (storing diffs between similar blobs) to reduce disk usage. git gc triggers this. Packing is a storage optimization — the logical model is still immutable, content-addressed objects. --- ## 🟣 Diffing, Logging & Code Archaeology

commit abc123
  └── tree def456 (root directory)
       ├── blob 111aaa  README.md
       ├── tree 222bbb  src/
       │    ├── blob 333ccc  main.py
       │    └── blob 444ddd  utils.py
       └── tree 555eee  tests/
            └── blob 666fff  test_main.py
  • Tag (annotated) — points to a commit (or any object) and adds tagger identity, date, message, and optional signature.
💡 The Senior SRE Gold Nugget (Key Architectural Takeaway)
"Pro-Tip: Blob — stores file content (just the raw bytes, no filename or permissions). Two files with identical content share the same blob ."
⚡ 60-Second Elevator Pitch Talking Points
  • Blob — stores file content (just the raw bytes, no filename or permissions). Two files with ident...
  • Tree — represents a directory. Contains entries mapping filenames → blob SHAs (for files) or othe...
  • Commit — points to one tree (the root tree of the project at that moment) and zero or more parent...
Advertisement
Want more Git scenarios?
Explore our complete collection of scenario-based Git interview runbooks.
Browse All Git Questions →

📚 Related Production Scenarios in Git