Junior SRE Engineer
Hace 9 horas
Guadalajara, Sonora, México
CLOUDSUFI
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Al continuar, aceptas nuestros Términos & Política de Privacidad.
Junior SRE Engineer – AI-Driven SRE & Cloud SRE
Location:
Country Club, Guadalajara, Jalisco, Mexico 44610
Experience:
1–3 Years
Education:
BTech / BE / MCA / MSc Computer Science Reporting To: Lead SRE / Solution Architect – Reliability Engineering
About the Role
CLOUDSUFI is looking for a Junior SRE Engineer to join an AI-driven Site Reliability Engineering team supporting a regulated enterprise platform in Guadalajara. This role is ideal for an early-career SRE, DevOps, Cloud, or Production Support Engineer who wants hands-on exposure to AWS, Datadog, Terraform, Kubernetes, Python/Bash automation, CI/CD, incident management, and AI-driven reliability engineering. You will work closely with senior SRE engineers and architects to support production reliability, observability, automation, cloud operations, and AI-enabled incident response.
Key Responsibilities
1. AI-Augmented Incident Response & RCA
• Support incident response as a secondary/shadow responder under senior SRE guidance.
• Analyze metrics, logs, traces, and deployment history to help identify incident causes.
• Support AI-driven incident detection, correlation, summarization, and RCA tools.
• Assist with postmortems and tracking corrective and preventive actions. 2. Observability & Alert Management
• Build and maintain Datadog monitors, dashboards, and SLOs using Terraform.
• Support anomaly detection, outlier detection, forecast monitors, composite alerts, and dynamic thresholds.
• Help reduce alert noise and improve monitor quality.
• Conduct monitor-hygiene and coverage-gap reviews.
• Support SLI/SLO and error-budget reviews for critical services. 3. AWS Cloud Reliability
• Support AWS workloads including ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, and Step Functions.
• Monitor capacity, performance, saturation, and cloud-cost trends.
• Assist with cloud observability and cost governance.
• Identify unusual or high-cost telemetry and infrastructure patterns and escalate findings. 4. AI Governance & Reliability
• Participate in reliability and architecture reviews under senior guidance.
• Learn how risks such as prompt injection, authorization gaps, secrets exposure, and autonomous-agent blast radius are assessed.
• Support production-readiness, change-management, and audit requirements.
• Assist in tracking AI-risk and reliability remediation activities. 5. DevOps, CI/CD & Automation
• Support reliability and security controls within CI/CD pipelines.
• Assist with AI-enabled change-risk, configuration, dependency, and infrastructure-drift checks.
• Develop automation using Python and Bash against AWS, Datadog, GitHub, PagerDuty, and Jira APIs.
• Contribute to Terraform modules and pull requests.
• Help create runbooks and progressively automate operational procedures. 6. Reliability Workflow Support
• Support multiple reliability initiatives across SRE and Security teams.
• Maintain operational documentation and workflow tracking.
• Contribute to a reliability-first, automation-focused engineering culture. Required / Relevant Technical Skills
• 1–3 years of experience in SRE, DevOps, Cloud, Platform Engineering, or Production Support.
• Exposure to AWS or another major cloud platform.
• Knowledge of monitoring/observability concepts: metrics, logs, traces, and APM.
• Exposure to Datadog, Grafana, Prometheus, New Relic, CloudWatch, ELK, or OpenSearch.
• Basic understanding of SLI, SLO, error budgets, and alert management.
• Knowledge of Docker and Kubernetes/ECS.
• Beginner-to-intermediate Terraform, Ansible, or CloudFormation experience.
• Basic Linux and networking troubleshooting skills.
• Understanding of DNS, TLS, load balancing, timeouts, and retries.
• Some hands-on experience with Python, Bash, or similar scripting.
• Understanding of REST APIs and JSON.
• Experience with Git and pull-request workflows.
• Exposure to CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or ArgoCD.
• Familiarity with incident management, escalation processes, and PagerDuty/Opsgenie.
• Interest in AI/LLM-based operational tools and AIOps. Core Competencies
• SRE fundamentals: SLI, SLO, error budgets, and toil reduction
• Observability and monitoring
• AIOps and anomaly detection
• Incident response and RCA
• AWS cloud infrastructure
• Docker, ECS, and Kubernetes
• Terraform / Infrastructure as Code
• CI/CD and release engineering
• Python/Bash automation
• Cloud cost and reliability engineering
• Resilience patterns
• Disaster recovery and failure-management concepts
• Runbook automation and self-healing
• DevSecOps fundamentals
• AI-agent risk and governance fundamentals Preferred Certifications Certifications are preferred but not mandatory:
• AWS Certified Cloud Practitioner or Associate
• HashiCorp Terraform Associate
• Datadog Fundamentals or equivalent
• CKA or KCNA
• AI/ML or AIOps-related certification/exposure Ideal Candidate We are looking for someone who:
• Has 1–3 years of hands-on technical experience.
• Enjoys troubleshooting problems using data and evidence.
• Is interested in cloud infrastructure and reliability engineering.
• Wants to grow in AWS, Terraform, Datadog, Kubernetes, and automation.
• Is curious about applying AI to SRE and operational workflows.
• Understands the importance of production discipline and incident escalation.
• Is comfortable working in a US shift and participating in a shadow on-call rotation.
• Communicates clearly and works well with senior engineers and architects.
Location:
Country Club, Guadalajara, Jalisco, Mexico 44610
Location:
Country Club, Guadalajara, Jalisco, Mexico 44610
Experience:
1–3 Years
Education:
BTech / BE / MCA / MSc Computer Science Reporting To: Lead SRE / Solution Architect – Reliability Engineering
About the Role
CLOUDSUFI is looking for a Junior SRE Engineer to join an AI-driven Site Reliability Engineering team supporting a regulated enterprise platform in Guadalajara. This role is ideal for an early-career SRE, DevOps, Cloud, or Production Support Engineer who wants hands-on exposure to AWS, Datadog, Terraform, Kubernetes, Python/Bash automation, CI/CD, incident management, and AI-driven reliability engineering. You will work closely with senior SRE engineers and architects to support production reliability, observability, automation, cloud operations, and AI-enabled incident response.
Key Responsibilities
1. AI-Augmented Incident Response & RCA
• Support incident response as a secondary/shadow responder under senior SRE guidance.
• Analyze metrics, logs, traces, and deployment history to help identify incident causes.
• Support AI-driven incident detection, correlation, summarization, and RCA tools.
• Assist with postmortems and tracking corrective and preventive actions. 2. Observability & Alert Management
• Build and maintain Datadog monitors, dashboards, and SLOs using Terraform.
• Support anomaly detection, outlier detection, forecast monitors, composite alerts, and dynamic thresholds.
• Help reduce alert noise and improve monitor quality.
• Conduct monitor-hygiene and coverage-gap reviews.
• Support SLI/SLO and error-budget reviews for critical services. 3. AWS Cloud Reliability
• Support AWS workloads including ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, and Step Functions.
• Monitor capacity, performance, saturation, and cloud-cost trends.
• Assist with cloud observability and cost governance.
• Identify unusual or high-cost telemetry and infrastructure patterns and escalate findings. 4. AI Governance & Reliability
• Participate in reliability and architecture reviews under senior guidance.
• Learn how risks such as prompt injection, authorization gaps, secrets exposure, and autonomous-agent blast radius are assessed.
• Support production-readiness, change-management, and audit requirements.
• Assist in tracking AI-risk and reliability remediation activities. 5. DevOps, CI/CD & Automation
• Support reliability and security controls within CI/CD pipelines.
• Assist with AI-enabled change-risk, configuration, dependency, and infrastructure-drift checks.
• Develop automation using Python and Bash against AWS, Datadog, GitHub, PagerDuty, and Jira APIs.
• Contribute to Terraform modules and pull requests.
• Help create runbooks and progressively automate operational procedures. 6. Reliability Workflow Support
• Support multiple reliability initiatives across SRE and Security teams.
• Maintain operational documentation and workflow tracking.
• Contribute to a reliability-first, automation-focused engineering culture. Required / Relevant Technical Skills
• 1–3 years of experience in SRE, DevOps, Cloud, Platform Engineering, or Production Support.
• Exposure to AWS or another major cloud platform.
• Knowledge of monitoring/observability concepts: metrics, logs, traces, and APM.
• Exposure to Datadog, Grafana, Prometheus, New Relic, CloudWatch, ELK, or OpenSearch.
• Basic understanding of SLI, SLO, error budgets, and alert management.
• Knowledge of Docker and Kubernetes/ECS.
• Beginner-to-intermediate Terraform, Ansible, or CloudFormation experience.
• Basic Linux and networking troubleshooting skills.
• Understanding of DNS, TLS, load balancing, timeouts, and retries.
• Some hands-on experience with Python, Bash, or similar scripting.
• Understanding of REST APIs and JSON.
• Experience with Git and pull-request workflows.
• Exposure to CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or ArgoCD.
• Familiarity with incident management, escalation processes, and PagerDuty/Opsgenie.
• Interest in AI/LLM-based operational tools and AIOps. Core Competencies
• SRE fundamentals: SLI, SLO, error budgets, and toil reduction
• Observability and monitoring
• AIOps and anomaly detection
• Incident response and RCA
• AWS cloud infrastructure
• Docker, ECS, and Kubernetes
• Terraform / Infrastructure as Code
• CI/CD and release engineering
• Python/Bash automation
• Cloud cost and reliability engineering
• Resilience patterns
• Disaster recovery and failure-management concepts
• Runbook automation and self-healing
• DevSecOps fundamentals
• AI-agent risk and governance fundamentals Preferred Certifications Certifications are preferred but not mandatory:
• AWS Certified Cloud Practitioner or Associate
• HashiCorp Terraform Associate
• Datadog Fundamentals or equivalent
• CKA or KCNA
• AI/ML or AIOps-related certification/exposure Ideal Candidate We are looking for someone who:
• Has 1–3 years of hands-on technical experience.
• Enjoys troubleshooting problems using data and evidence.
• Is interested in cloud infrastructure and reliability engineering.
• Wants to grow in AWS, Terraform, Datadog, Kubernetes, and automation.
• Is curious about applying AI to SRE and operational workflows.
• Understands the importance of production discipline and incident escalation.
• Is comfortable working in a US shift and participating in a shadow on-call rotation.
• Communicates clearly and works well with senior engineers and architects.
Location:
Country Club, Guadalajara, Jalisco, Mexico 44610