Sai Gopal

Staff SRE

Where I've worked

AI Research Lab2024 – Present
Member of Technical Staff

Reliability and infrastructure for large-scale AI training and inference systems. GPU cluster provisioning and lifecycle management, platform security hardening, and OS fleet upgrades across heterogeneous hardware.

Large Social Platform2018 – Present · ~8 yrs
Staff Site Reliability Engineer → Staff

Progressed from SRE to Staff SRE, working across the core infrastructure stack. Owned reliability for critical platform services at global scale. Led datacenter build-outs, mass provisioning of bare-metal fleets, and DC shuffle operations. Owned OS and kernel upgrade programs including the Hadoop data platform. SRE coverage for the database tier — MySQL, PostgreSQL, and Vertica. Configuration management with Puppet, Python automation for on-call toil reduction, security hardening, and SLO frameworks.

Internet Services Company2014 – 2018 · ~4 yrs
System Architect → Architect

Grew from Operations Engineer to System Architect. Large-scale hosting and web infrastructure — systems automation, capacity planning, datacenter operations, and infrastructure-as-code.

Full work history on LinkedIn

Let's connect

Whether you want to talk reliability engineering, share ideas, or just say hello — my inbox is open.

web@saigopal.com