Garner Health is a healthcare technology startup focused on reducing employee health insurance costs for large companies by transforming the U.S. healthcare system using proprietary clinical metrics and AI.
About the role:
We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role is part of the Platform Engineering team and involves defining and upholding SLOs, leading incident response, and driving automation and standards to enable faster and more reliable shipping by engineers. The role emphasizes automation-first approaches using AI tools to convert manual operational work into monitored, hands-free processes.
Responsibilities:
— Own end-to-end reliability, performance, and resilience of cloud environments (AWS, Kubernetes), including AI/ML workloads
— Define, measure, and uphold SLOs across critical services
— Lead incident response, participate in on-call rotation, conduct root cause analysis, and review infrastructure changes
— Build and maintain monitoring, alerting, and observability systems
— Translate scaling requirements into automated infrastructure-as-code deliverables using Terraform
— Identify and implement cost-efficiency and performance improvements across the stack
— Reduce operational toil and tech debt through AI-driven automation
— Build and maintain deployment and observability standards to empower engineering teams
— Ensure infrastructure and operations comply with security and HIPAA requirements
Requirements:
— 4+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering roles
— Deep expertise with Kubernetes and Terraform in cloud-first environments (AWS preferred)
— Strong experience with production observability: defining SLOs, monitoring, alerting, incident response, and blameless post-incident reviews
— Strong software engineering skills in Python or Go applied to infrastructure automation; experience with Kubernetes APIs is a plus
— Experience optimizing cloud cost and performance across compute, storage, and networking
— Experience supporting AI/ML or data-intensive workloads in production is a plus
— Experience in security-conscious or regulated environments (HIPAA, SOC 2) is a plus
— Familiarity or strong motivation to use AI tools (e.g., Claude) for engineering and operations workflows
— Commitment to high performance, accountability, and authentic feedback
Technologies used:
AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab
Additional information:
— Remote position with occasional travel to NYC headquarters
— No visa sponsorship available
— Compensation range: $191,000 - $226,000 per year plus equity and benefits including flexible PTO, medical/dental/vision plans, 401(k) with company match, flexible spending accounts, and Teladoc Health
— Equal Employment Opportunity employer committed to diversity and accommodations for disabilities
Контакты работодателя доступны по кнопке «Откликнуться» после входа.