About Runlayer
AI is transforming how every company operates, but most enterprises are stuck. They want to move fast with AI Agents, tools, and workflows, but they can't do it safely. We're fixing that.
Our team built AI Actions for OpenAI, shipped Zapier Agents to millions of users, and launched the first remote MCP server with Anthropic. We helped establish the protocol, and now we're building the platform enterprises need to actually put AI to work.
Runlayer is one platform for MCPs, Skills, and Agents: purpose-built security, fine-grained governance, and complete observability so organizations can go all-in on AI across the entire company without the risk. We just raised a $30M Series A led by Felicis, with participation from Khosla Ventures, bringing our total raised to $42M. Already trusted by Gusto, Instacart, Opendoor, dbt Labs, and Decagon.
About the Role
As our Site Reliability Engineer, you'll own the reliability, performance, and scalability of Runlayer's infrastructure as we grow to serve enterprise customers across multiple deployment models: multi-tenant SaaS, single-tenant SaaS, and BYOC.
Why You'll Thrive Here
Impact: Build the infrastructure foundation for the enterprise AI platform, directly enabling AI adoption at scale.
Excellence: Work closely with founders and a small, senior engineering team shipping fast in a high-growth environment.
Ownership: Own reliability end-to-end, from database performance to incident response to CI/CD pipelines.
What You'll Do
Own reliability and performance of our cloud infrastructure across AWS (ECS, Aurora, CloudWatch) and GCP
Manage and optimize Kubernetes clusters and container orchestration
Drive database reliability engineering, including performance tuning and scaling
Build and maintain CI/CD pipelines for rapid, safe deployments
Run incident response and on-call rotations
Partner with product engineers to design scalable, resilient systems
What We're Looking For
Experience deploying and supporting on-prem / BYOC environments. We are looking for engineers who have experienced the challenges of scaling BYOC operations, and know what they would/wouldn’t do again
Background at a B2B company serving enterprise customers, ideally building infrastructure/platform products
Strong AWS experience, particularly ECS, Aurora, Kinesis
Networking: VPC peering, Transit Gateway, PrivateLink, security groups, DNS, TLS termination
CI/CD pipeline ownership and incident response experience
Bonus Qualifications
What We Offer
We provide a competitive package designed to attract and retain top talent who can work effectively with enterprise customers.
Competitive salary and equity — compensation that reflects your expertise and customer-facing responsibilities.
Paid time off — paid vacation, paid sick leave, and paid parental leave.
Professional development — budget for conferences, courses, and certifications in AI, enterprise software, and customer success.
Top-tier equipment — your choice of laptop and accessories to create your ideal work environment.
Health benefits — comprehensive health, dental, and vision coverage.
Customer interaction opportunities — work directly with innovative companies and see the immediate impact of your work.
Not quite the right fit? Reach out to careers@runlayer.com with details about your experience and interests.