Q: Your CMS and machine learning model serving workloads running on GKE require hundreds of pods across multiple zones to mount the exact same shared file storage volume with ReadWriteMany (RWX) access. Standard Persistent Disks only support ReadWriteOnce (RWO). How do you architect and manage scalable shared storage using GCP Filestore Enterprise?
Production guide for provisioning multi-writer ReadWriteMany (RWX) storage in GKE using Google Cloud Filestore Enterprise, tuning NFS client mount options, and automating point-in-time snapshot replication.
Want to master this scenario in a live sandbox? Stephane Maarek's AWS Certified DevOps Engineer Professional Masterclass on Udemy covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Enable Filestore CSI Driver on GKE Cluster
Activate the managed Google Cloud Filestore CSI driver to enable dynamic persistent volume provisioning:
- Cluster Update: Executed
gcloud container clusters update prod-cluster --update-addons=GcpFilestoreCsiDriver=ENABLED. - Verify Controller: Confirmed
filestore.csi.storage.gke.ioDaemonSet pods were running and healthy on all cluster nodes.
Define Multi-Zone Filestore Enterprise StorageClass
Configure dynamic provisioning for high-availability multi-zone shared NFS file systems:
- StorageClass YAML: Created StorageClass with
tier: enterpriseandnetwork: prod-vpc. - Enterprise SLA: Filestore Enterprise provides 99.99% availability with synchronous replication across three availability zones in the region.
Deploy ReadWriteMany (RWX) PVC & Tune NFS Mount Options
Bind application pods to the shared file share with optimized client-side caching and locking parameters:
- PVC Declaration: Declared PersistentVolumeClaim with
accessModes: [ReadWriteMany]andstorage: 1Ti. - Mount Tuning: Tuned mount options in StorageClass:
mountOptions: ['hard', 'nfsvers=4.1', 'rsize=1048576', 'wsize=1048576', 'timeo=600']to prevent IO freezing during brief network blips.
Automate Volume Snapshots & Point-in-Time Disaster Recovery
Establish continuous automated backup schedules using Kubernetes VolumeSnapshot CRDs:
- VolumeSnapshotClass: Configured VolumeSnapshotClass backed by
filestore.csi.storage.gke.iodriver. - Automated CronJob: Deployed a Kubernetes CronJob creating daily snapshots retaining 30 days of point-in-time recovery points.
- Rapid Restore Verification: Validated restoring a 1 TB snapshot into a new test PVC in under 90 seconds without production downtime.
- Enable the native GKE Filestore CSI driver to dynamically provision shared storage.
- Deploy Filestore Enterprise StorageClass with 99.99% multi-zone synchronous replication.
- Configure ReadWriteMany PVCs with optimized NFS 4.1 buffer and timeout mount options.
- Automate point-in-time VolumeSnapshots for rapid recovery against accidental file corruption.