Senior Site Reliability Engineer
hace 2 semanas
Join our team as a **Senior Site Reliability Engineer** focused on delivering advanced support for critical Azure-based systems.**Responsibilities**- Troubleshoot and resolve complex incidents to maintain system uptime- Ensure reliability and performance of Azure-based enterprise infrastructure- Implement observability, monitoring, and logging solutions- Automate infrastructure provisioning and deployment using Terraform and scripting- Optimize system performance and uptime through proactive monitoring and alerting- Collaborate with cross-functional teams to improve service reliability- Conduct root cause analysis and postmortems for incident management- Manage deployment pipelines in Azure DevOps for secure and scalable workflows- Develop and maintain automation scripts for routine tasks and incident recovery- Enhance monitoring frameworks with tools like Prometheus and Grafana- React quickly to incidents to avoid SLA degradation- Integrate monitoring data from Azure and AWS environments- Support continuous improvement of service reliability and observability practices- Document technical processes and incident reports- Participate in Agile team activities and prioritize competing tasks**Requirements**:- Minimum 3 years of experience in site reliability engineering or related DevOps roles- Hands-on experience with Azure services, including AKS, Azure Monitor, Application Insights, Log Analytics, Cosmos DB, and PostgreSQL- Strong expertise in Azure DevOps and Terraform for infrastructure automation- Proficient scripting skills in Bash, PowerShell, and Python- Experience with monitoring and observability tools such as Prometheus and Grafana- Solid background in incident management and ITSM processes with root cause analysis capabilities- Ability to troubleshoot and debug complex technical issues in real-time- Experience working in fast-paced Agile environments- Strong verbal and written communication skills for collaboration and reporting- Proactive approach to setting alerts and preventing SLA degradation- Experience with cloud infrastructure scaling and security best practices- Knowledge of Kubernetes administration and orchestration- Ability to collaborate effectively with cross-functional teams- English language proficiency at B2 level or above**Nice to have**- Hands-on experience with AWS services including EKS, RDS, CloudWatch, and X-Ray- Familiarity with distributed logging pipelines and incident automation tools- Knowledge of advanced Kubernetes use cases for scaling and network configurations- Certifications such as Microsoft Azure Administrator or AWS Certified DevOps Engineer- Experience with observability tools like OpenSearch for AWS workloads**We offer**- Career plan and real growth opportunities- Unlimited access to LinkedIn learning solutions- International Mobility Plan within 25 countries- Constant training, mentoring, online corporate courses, eLearning and more- English classes with a certified teacher- Support for employee’s initiatives (Algorithms club, toastmasters, agile club and more)- Enjoyable working environment (Gaming room, napping area, amenities, events, sport teams and more)- Flexible work schedule and dress code- Collaborate in a multicultural environment and share best practices from around the globe- Hired directly by EPAM & 100% under payroll- Law benefits (IMSS, INFONAVIT, 25% vacation bonus)- Major medical expenses insurance: Life, Major medical expenses with dental & visual coverage (for the employee and direct family members)- 13 % employee savings fund, capped to the law limit- Grocery coupons- 30 days December bonus- Employee Stock Purchase Plan- 12 vacations days plus 4 floating days- Official Mexican holidays, plus 5 extra holidays (Maundry Thursday and Friday, November 2nd, December 24th & 31st)- Monthly non-taxable amount for the electricity and internet billsEPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
-
Site Reliability Engineer
hace 1 semana
Desde casa, México thegetch mexico A tiempo completo**Función: Site Reliability Engineer****Aperturas: más de 10 contrataciones****Ubicación: - any city with TCS Office presence (Queretaro, Guadalajara, Mexico City or Monterrey)****Salario:- 25-33 USD/hr****Comunicación en inglés: avanzado****Experiência: 4+ años****Responsabilidades de Site Reliability Engineer**:Reúna y analice métricas de sistemas...
-
Senior Site Reliability Engineer
hace 1 semana
Desde casa, México EPAM Systems, Inc. A tiempo completoWe are seeking an experienced **Senior Site Reliability Engineer**to join our team. As a key member of the Reliability Tooling team, you will be responsible for writing and reviewing code, contributing to critical technical decisions, and mentoring engineers within your squad. This role requires a deep understanding of SRE principles and best practices, as...
-
Senior Site Reliability Engineer
hace 2 semanas
Desde casa, México EPAM Systems, Inc. A tiempo completoJoin our team as a **Senior Site Reliability Engineer** focused on delivering advanced support for critical Azure-based systems. **Responsibilities** - Troubleshoot and resolve complex incidents to maintain system uptime - Ensure reliability and performance of Azure-based enterprise infrastructure - Implement observability, monitoring, and logging...
-
Senior Site Reliability Engineer
hace 3 semanas
Desde casa, México EPAM Systems A tiempo completo**DESCRIPTION**:Join EPAM as a **Senior Site Reliability Engineer specializing in AWS!**In this role, you'll ensure fleet services reliability and availability under the SRE model.If you have a good track record of highly scalable, distributed systems projects and previous experience working as an SRE, we'd love to hear from you.EPAM is a leading global...
-
Senior Azure Site Reliability Engineer
hace 2 semanas
Desde casa, México Pinnacle A tiempo completo**Job Title**: Senior Azure Site Reliability Engineer**Reports** **To**: Azure Site Reliability Lead**About us**:Welcome to Pinnacle, the ultimate destination for sports enthusiasts seeking an exhilarating sportsbook and gaming experience! Established in 1998, we have solidified our position as one of the globe's foremost licensed online gaming companies....
-
Senior Azure Site Reliability Engineer
hace 2 semanas
Desde casa, México Pinnacle A tiempo completo**Job Title**: Senior Azure Site Reliability Engineer **Reports** **To**: Azure Site Reliability Lead **About us**: Welcome to Pinnacle, the ultimate destination for sports enthusiasts seeking an exhilarating sportsbook and gaming experience! Established in 1998, we have solidified our position as one of the globe's foremost licensed online gaming...
-
Site Reliability Engineer
hace 16 horas
Desde casa, México Luxoft A tiempo completo**Project description**: Do you like to work with existing and new software product development teams? This position is to instrument end-to-end observability and visibility for business-critical systems with log ingestion, metrics, and traces. You will function as a site reliability engineer (SRE) that will collaborate with product teams, infrastructure...
-
Site Reliability Engineer
hace 4 días
Desde casa, México Right Balance A tiempo completo**Overview** We're looking for a Site Reliability Engineer. Headquartered in Los Angeles, California, Right Balance provides top-tier technology talent for innovative companies in the US. We’re in the top 50 companies to watch in LA. **Engagement Details** Our client is a USA-based company producing video solutions with the mission to advance scientific...
-
Site Reliability Engineer
hace 2 días
Desde casa, México Right Balance A tiempo completo**Overview**We're looking for a Site Reliability Engineer. Headquartered in Los Angeles, California, Right Balance provides top-tier technology talent for innovative companies in the US. We’re in the top 50 companies to watch in LA.**Engagement Details**Our client is a USA-based company producing video solutions with the mission to advance scientific...
-
Lead Site Reliability Engineer
hace 3 semanas
Desde casa, México Tekshapers Inc A tiempo completo**Position : Lead Site Reliability Engineer****Location : Remote****Duration : Contract**- Lead and mentor a team of SREs to ensure operational excellence and maximize the reliability and availability of client systems.- Minimum 10 years of work experience in DevOps/SRE, including leadership roles.- Architect and design highly scalable and available...