Site Reliability Engineer
Runloop AI
San Francisco, CA
See who Runloop AI has hired for this role
See who Runloop AI has hired for this role
Runloop AI provided pay range
This range is provided by Runloop AI. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.
Base pay range
$150,000.00/yr - $250,000.00/yr
About Runloop
Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.
The Role
We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.
Responsibilities
Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.
Compensation Range: $150K - $250K
Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.
The Role
We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.
Responsibilities
- Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds
- Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users
- Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
- Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
- Participate in an on-call rotation to support our production systems
- Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
- Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
- Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
- Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
- Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
- Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
- 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
- Strong programming skills in languages like Python or Go
- Deep expertise in containerization technologies such as Docker and Kubernetes
- Experience with cloud infrastructure and tools like Terraform and/or Pulumi
- Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
- A solid understanding of networking, security, and Linux systems administration
- Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
- Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
- Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
- Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance
- Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery
- Competitive salary and equity
- Comprehensive health, dental, and vision insurance for employee and dependents
- Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
- Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
- Onsite 4 days a week in San Francisco; Optional 1 day a week remote
Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.
Compensation Range: $150K - $250K
-
Seniority level
Mid-Senior level -
Employment type
Full-time -
Job function
Engineering and Information Technology -
Industries
Technology, Information and Internet
Referrals increase your chances of interviewing at Runloop AI by 2x
See who you knowSimilar jobs
People also viewed
-
Software Engineer, Site Reliability
Software Engineer, Site Reliability
-
Site Reliability Engineer
Site Reliability Engineer
-
Platform Engineer — Infra / Reliability Specialist
Platform Engineer — Infra / Reliability Specialist
-
Reliability Engineer
Reliability Engineer
-
Technology, DevOps/Site Reliability Engineer
Technology, DevOps/Site Reliability Engineer
-
Site Reliability Engineer-Americas & EMEA Tech
Site Reliability Engineer-Americas & EMEA Tech
-
Infrastructure engineer
Infrastructure engineer
-
Staff Site Reliability Engineer, Ads
Staff Site Reliability Engineer, Ads
-
Infrastructure Engineer
Infrastructure Engineer
-
Site Reliability Engineer
Site Reliability Engineer
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content