Líder Técnico Azure
Hace 2 días
Guadalajara Jal, Guadalajara (municipio); Estado de Jalisco, México
CLOUDSUFI
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Role: Lead SRE Engineer
Location:
Guadalajara, Jalisco, Mexico
Experience:
6–10 years
Key Responsibilities
Lead Site Reliability Engineering initiatives across enterprise production environments. Drive reliability, availability, scalability, and performance improvements. Lead P1/P2 incident response, RCA, and postmortems. Define and improve SLI, SLO, and Error Budget practices. Build and maintain observability, monitoring, dashboards, and alerts. Automate operational processes using Python/Bash and APIs. Manage Infrastructure as Code using Terraform. Support AWS cloud infrastructure, including Kubernetes/EKS, ECS/Fargate, Lambda, RDS/Aurora, SQS/SNS, and related services. Mentor SRE/DevOps engineers and collaborate with development, security, platform, and architecture teams. Contribute to AI-enabled reliability and AIOps initiatives. Must-Have Experience Observability / Monitoring – Mandatory Hands-on experience with Datadog or an equivalent APM, logging, and monitoring platform , such as New Relic, Dynatrace, Grafana/Prometheus, Splunk, ELK/OpenSearch, CloudWatch, or AppDynamics. Experience should include: APM Metrics and logs Distributed tracing Dashboards and monitoring Alerting Production troubleshooting Alerting / Incident Management – Mandatory Hands-on experience with PagerDuty or an equivalent alerting/paging platform , such as Opsgenie, ServiceNow, Splunk On-Call, xMatters, or Grafana Alerting. Experience should include:
Location:
Guadalajara, Jalisco, Mexico
Experience:
6–10 years
Key Responsibilities
Lead Site Reliability Engineering initiatives across enterprise production environments. Drive reliability, availability, scalability, and performance improvements. Lead P1/P2 incident response, RCA, and postmortems. Define and improve SLI, SLO, and Error Budget practices. Build and maintain observability, monitoring, dashboards, and alerts. Automate operational processes using Python/Bash and APIs. Manage Infrastructure as Code using Terraform. Support AWS cloud infrastructure, including Kubernetes/EKS, ECS/Fargate, Lambda, RDS/Aurora, SQS/SNS, and related services. Mentor SRE/DevOps engineers and collaborate with development, security, platform, and architecture teams. Contribute to AI-enabled reliability and AIOps initiatives. Must-Have Experience Observability / Monitoring – Mandatory Hands-on experience with Datadog or an equivalent APM, logging, and monitoring platform , such as New Relic, Dynatrace, Grafana/Prometheus, Splunk, ELK/OpenSearch, CloudWatch, or AppDynamics. Experience should include: APM Metrics and logs Distributed tracing Dashboards and monitoring Alerting Production troubleshooting Alerting / Incident Management – Mandatory Hands-on experience with PagerDuty or an equivalent alerting/paging platform , such as Opsgenie, ServiceNow, Splunk On-Call, xMatters, or Grafana Alerting. Experience should include: