— Define service-level objectives for production services. — Build and maintain useful reliability dashboards.
About the Project
The platform supports medical voice services and workflow automation for healthcare systems in the US. The role focuses on production reliability: monitoring, incident response, SLOs, deployment safety, recovery, and uptime evidence.
MUST HAVE
— Strong production experience in SRE / DevOps / Backend Operations
— Hands-on experience with monitoring, logging, alerting, and tracing
— Real experience with Incident Response / Incident Management
— Clear responsibility for uptime, SLA / SLO / service reliability
— Experience defining or working with SLOs and production dashboards
— Strong cloud infrastructure background
— Experience with containers and deployment automation
— Experience with Infrastructure as Code (IaC)
— Strong CI/CD knowledge
— Python or comparable scripting skills for operational troubleshooting and automation
— Experience diagnosing and resolving production failures
— Understanding of backup, recovery, and capacity testing
— Ability to communicate clearly during incidents and provide concise written updates
Responsibilities
— Define service-level objectives for production services
— Build and maintain useful reliability dashboards
— Instrument metrics, logs, and traces
— Monitor workflow, voice, and third-party service dependencies
— Lead incident response during agreed coverage hours
— Write clear post-incident reviews
— Improve deployment safety with rollback and staged releases
— Maintain change records and release processes
— Test capacity, backups, and disaster recovery
— Coordinate handover and escalation with engineering teams
— Provide accurate uptime and reliability evidence
— Support continuous improvement of platform resilience
Strong Advantage
— Healthcare or another regulated production environment
— WebRTC / real-time media / voice infrastructure
— LLM-backed services
— Experience handling third-party provider outages
— Grafana / Prometheus / Datadog or similar observability tools
— Experience with customer-facing incident communication
Working Hours - Important
The team works with US-based stakeholders, and calls are expected during Arizona time (MST, UTC-7). Candidates should be comfortable with the possibility of regular communication and incident-related coverage during US / Arizona business hours.
What We Offer
— Small-company feel within a fast-growing international environment
— Friendly, collaborative, and mission-driven team
— 25 calendar days of vacation + 5 additional paid sick days
— Medical insurance
— Corporate English courses
— Corporate events and team-building activities
— Support with professional certifications
— Reimbursement for professional courses and training
— Long-term international projects
— Opportunities for professional and technical growth
Контакты работодателя доступны по кнопке «Откликнуться» после входа.
59/100
Средняя вакансия
Оценка от Job Hunters AI
Senior DevOps / Site Reliability Engineer; компания не названа; зарплата не указана; понятные задачи; современный стек.
Из чего сложилась оценка. Нажмите, чтобы узнать подробнее
Компания не названа
The employer is not named anywhere in the text or fields, which reduces transparency and trust for candidates.
Из объявления
🔥 Senior DevOps / Site Reliability Engineer - Healthcare / Voice / Observability … We are currently looking for a Middle+ / Senior DevOps / Site Reliability Engineer to join an international healthcare project focused on workflow and voice services for US health-system use. … About the Project