Site Reliability Engineer
- Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.
- Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.
- Participates in team stand-ups on projects, incidents, and daily priorities.
- Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.
- Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.
- Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.
Must have:
- Architecture, design patterns, reliability, and scaling of new and existing systems.
- Incident command experience — driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.
- Observability built from the ground up — defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.
- Linux kernel internals — scheduler, memory allocation, driver subsystems.
- High-quality code in at least one language (Python, Go, or similar).
- System-level debugging — kdump, kernel panic analysis.
- IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
- TCP/IP and network programming.
- Distributed storage systems — object, block, and/or file storage paradigms.
- Strong communication skills.
Nice to have:
- Hardware and GPU troubleshooting.
- OVN/OVS-based networking stack exposure.
-
Seniority level
Mid-Senior level -
Employment type
Contract -
Job function
Engineering and Information Technology -
Industries
IT Services and IT Consulting
Referrals increase your chances of interviewing at Neurealm by 2x
See who you knowGet notified about new Site Reliability Engineer jobs in Sunnyvale, CA.
Sign in to create job alertSimilar jobs
People also viewed
-
Software Engineer, Site Reliability
Software Engineer, Site Reliability
-
Site Reliability Engineer
Site Reliability Engineer
-
Platform Engineer — Infra / Reliability Specialist
Platform Engineer — Infra / Reliability Specialist
-
Reliability Engineer
Reliability Engineer
-
Technology, DevOps/Site Reliability Engineer
Technology, DevOps/Site Reliability Engineer
-
Site Reliability Engineer-Americas & EMEA Tech
Site Reliability Engineer-Americas & EMEA Tech
-
Infrastructure engineer
Infrastructure engineer
-
Staff Site Reliability Engineer, Ads
Staff Site Reliability Engineer, Ads
-
Infrastructure Engineer
Infrastructure Engineer
-
Site Reliability Engineer
Site Reliability Engineer
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content