Site Reliability Engineer (MTS) GovCloud 24x7
Software Engineering
McLean, VA, USA · Burlington, MA, USA · Denver, CO, USA
The Experience
Join our team and contribute to the operational excellence of the Salesforce GovCloud! Are you passionate about ensuring the reliability and performance of mission-critical cloud services? Salesforce is seeking a talented Site Reliability Engineer to join our dynamic team, supporting our GovCloud environment. As a key member of our Site Reliability organization, you'll play a vital role in maintaining 99.99% uptime for customer-facing services, proactively addressing issues, and ensuring the security of our data. We foster a collaborative and innovative culture, where you'll work alongside skilled engineers to solve complex problems and drive continuous improvement.
Please Note: This position requires a successful background investigation and the ability to obtain and maintain a specific level of U.S. government background clearance. Details will be provided during the interview process.
Shift Requirements: This role involves shift work, including night shifts, as part of a 24/7 support team. We provide a rotating schedule and ensure adequate compensation for shift differentials.
GovCloud Incident Response (GIR) maintains critical infrastructure stability through incident management, automated alerting, collaborative smart-hands support, thoughtful retrospectives, and sustainable long-term remediation strategy.
What You'll Actually Be Doing
- Maintain system reliability and high performance across customer-facing services by supporting foundational infrastructure health.
- Lead incident response efforts during major operational events (e.g., Sev0/Sev1) and collaborate in post-incident reviews to drive problem resolution.
- Populate and participate in Root Cause Analyses (RCAs) and partner with Global Solutions teams to implement preventative solutions.
- Ensure Site Reliability engineering activities align with organizational compliance, security standards, and operational guidelines.
- Demonstrate passion for collaborative problem-solving and addressing complex technical challenges across cross-functional teams.
- Mentor and collaborate with teammates to explore emerging technologies, foster continuous learning, and support professional growth.
- Adapt to dynamic, fast-paced operational needs while prioritizing competing tasks effectively.
- Work to automate detection and resolution of recurring issues in the production environment.
- Collaborate on improving existing workflows to streamline engineering practices and minimize operations and engineering toil.
You're Our Person If...
- You're a U.S. citizen (U.S. born or naturalized) who does not hold dual citizenship, and you agree to complete a Minimum Background Investigation (MBI) for a Moderate Public Trust position with the U.S. federal government or other clearances as deemed appropriate for the role.
- You have a related technical degree.
- You have demonstrated experience with enterprise-scale internet service operations or systems engineering.
- You have expertise in TCP/IP related technologies (networking protocols, network programming, etc.).
- You have command-line expertise and administration knowledge of Unix variants (Linux/Solaris/BSD) systems (e.g., Red Hat Enterprise Linux and Solaris).
- You have solid comprehension of security monitoring systems and infrastructure administration.
- You have effective written and verbal communication skills with a focus on collaborative teamwork.
- You're familiar with Incident Management principles and IT service operation frameworks.
- You can support a 24/7 high-availability operational environment, including shift and on-call rotations as needed.
- You have experience provisioning, operating, and running AWS/C2S based infrastructure and systems.
- You're proficient in scripting or programming languages such as Python, Go, or similar technologies.
- You have an open mindset toward leveraging AI technologies to enhance engineering productivity and continuous learning across technical domains.
Even Better If...
- You have prior Chef/Puppet or automated deployment experience.
- You have prior Jenkins/Bamboo/Spinnaker pipeline execution experience.
- You have experience supporting and maintaining monitoring and alert systems.
- You have experience supporting and maintaining Java applications.
- You have hands-on experience configuring and running AWS, using the CLI/SDKs.
- You hold certifications in Linux+, RedHat, and AWS.
- You have experience supporting and leading Kubernetes based applications and services.
- You've taken part in blameless retrospectives, learning from incidents, and conducting post-incident investigations, including incident analysis as well as performance evaluations of responders.
- You have working knowledge of and interest in resilience engineering, including concepts such as Safety II — looking at how things go right instead of how things go wrong, being proactive instead of reactive, and investigating complex sociotechnical systems.
- You're familiar with Agile process and DevOps.
- You have experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows.
- You have advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready.
Must be a U.S. Citizen operating on U.S. Soil with ability to meet customer and government screening standards applicable to this role, including a Criminal Justice Information Services screening with fingerprint scan. Due to the citizenship requirement for this role, which supports U.S. federal, state, and/or local government customers, citizenship will be verified through two of the following REAL ID Act documents: U.S. Passport, Passport Card, REAL Driver's License, Global Entry Card, U.S. Government CAC/PIV.