Staff SRE
Reliability and infrastructure for large-scale AI training and inference systems. GPU cluster provisioning and lifecycle management, platform security hardening, and OS fleet upgrades across heterogeneous hardware.
Progressed from SRE to Staff SRE, working across the core infrastructure stack. Owned reliability for critical platform services at global scale. Led datacenter build-outs, mass provisioning of bare-metal fleets, and DC shuffle operations. Owned OS and kernel upgrade programs including the Hadoop data platform. SRE coverage for the database tier — MySQL, PostgreSQL, and Vertica. Configuration management with Puppet, Python automation for on-call toil reduction, security hardening, and SLO frameworks.
Grew from Operations Engineer to System Architect. Large-scale hosting and web infrastructure — systems automation, capacity planning, datacenter operations, and infrastructure-as-code.
Full work history on LinkedIn
Whether you want to talk reliability engineering, share ideas, or just say hello — my inbox is open.
web@saigopal.com