Site Reliability Engineer

Hace 13 horas

Guadalajara, Sonora, México CLOUDSUFI Jornada completa
Job Title: Site Reliability Engineer

Experience:
1–3 Years

Location:
Guadalajara, Mexico Must have
- SRE : Datadog(or any apm, logging and monitoring tools) + Pagerduty(or any alerting tool)

About the Role
CLOUDSUFI is looking for a Junior SRE Engineer to join an AI-driven reliability engineering team supporting a regulated FinTech platform running on AWS. You will work alongside senior SREs and architects to operate, automate, and improve production reliability using AI-assisted SRE, observability, cloud automation, and modern DevOps practices.

Key Responsibilities
• Support incident response, triage, RCA, and postmortems using AI-assisted tools.
• Build and maintain Datadog monitors, dashboards, and SLOs using Terraform.
• Improve alert quality, reduce noise, and support monitor-hygiene initiatives.
• Support AWS services including ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, and Step Functions.
• Identify reliability, capacity, performance, and cloud-cost anomalies.
• Develop automation using Python, Bash, and REST APIs.
• Contribute to Terraform, CI/CD, runbooks, and automated remediation.
• Support AI-agent governance, reliability reviews, and production-readiness activities.
• Collaborate with SRE, Platform, Security, and Engineering teams. What We're Looking For
• 1–3 years of experience in SRE, DevOps, Cloud, Platform, or Production Support.
• Foundational knowledge of AWS, Linux, Docker/Kubernetes, Terraform, and CI/CD.
• Exposure to Datadog, Prometheus/Grafana, CloudWatch, ELK, or similar.
• Basic understanding of SLI/SLO, error budgets, incident management, and observability.
• Working knowledge of Python or Bash and REST APIs.
• Strong troubleshooting, analytical, and communication skills.
• Interest in AI/AIOps, automation, reliability engineering, and cloud technologies.
• Willingness to participate in a shadow on-call rotation and learn production operations. Preferred Certifications
• AWS Cloud Practitioner / Associate
• Terraform Associate
• Datadog Fundamentals
• KCNA / CKA
• AI/ML or AIOps certification/exposure Why Join CLOUDSUFI? Gain hands-on experience with large-scale AWS, AI-driven SRE, observability, automation, FinOps, CI/CD, and self-healing systems while working closely with experienced SREs and architects. Reports To: Lead SRE / Solution Architect – Reliability Engineering :::