Site Reliability Engineer

Hace 5 días

apodaca, nuevo león, México Qualcomm Jornada completa
## \nCompany:\n\nQUALCOMM SEMICONDUCTORES Y SISTEMAS AVANZADOS DE BAJA CALIFORNIA\n\n## Job Area:\n\nEngineering Group, Engineering Group \u003e Software Engineering\n\nGeneral Summary:\n\nCloud Infrastructure \u0026 Infrastructure as Code\n\n
* Design, build, and manage cloud infrastructure with a primary focus on AWS, integrated with OpenStack environments\n
* Build and maintain Infrastructure as Code using:\n
* Terraform\n
* Ansible\n
* Kubernetes (manifests / Helm)\n\n\n
* Design infrastructure solutions for:\n
* Scalability\n
* High availability\n
* Performance\n
* Reliability\n
* Cost efficiency\n\n\n
* Implement redundancy, failover, and disaster\u2011recovery patterns across services and regions\n
* Perform capacity planning based on performance metrics, usage trends, and utilization data\n\n\n\nKubernetes \u0026 Platform Reliability\n\n
* Operate and scale production Kubernetes clusters in large\u2011scale environments\n
* Partner with development and QA teams to:\n
* Improve system reliability and resiliency\n
* Automate scalability and availability mechanisms\n\n\n
* Apply SRE principles including:\n
* Service reliability ownership\n
* Proactive failure prevention\n
* Continuous improvement of operational processes\n\n\n
* Support microservices\u2011based and distributed system architectures\n\n\n\nCI/CD, Automation \u0026 Operational Excellence\n\n
* Manage and evolve CI/CD pipelines (e.g., Jenkins)\n
* Automate infrastructure provisioning, configuration, and lifecycle management\n
* Write, maintain, and improve runbooks for operational processes\n
* Build automation to reduce manual intervention and operational toil\n
* Plan and execute infrastructure upgrades and maintenance activities\n
* Proactively identify and address technical and infrastructure debt\n\n\n\nData Platforms \u0026 Streaming Systems\n\n
* Operate, tune, and scale data and streaming platforms, including:\n
* Kafka, Zookeeper\n
* NiFi\n
* Elasticsearch\n
* MySQL, Vertica\n\n\n
* Diagnose and resolve performance and stability issues across data pipelines\n
* Ensure data platform reliability, throughput, and resilience at scale\n\n\n\nAI\u2011Assisted SRE \u0026 Intelligent Automation\n\n
* Design and maintain knowledge\u2011driven automated runbooks and operational bots\n
* Develop AI\u2011assisted operational workflows, including:\n
* Incident analysis and summarization\n
* Intelligent diagnostics and remediation suggestions\n
* Automation of repetitive operational decision\u2011making\n\n\n
* Work with LLM\u2011based agent frameworks (e.g., Claude Agent SDK or similar):\n
* Integrate agents with logs, metrics, monitoring, and internal tools\n
* Implement guard\u2011railed, controlled\u2011action automation for production use\n\n\n
* Research and propose new concepts, tools, and AI\u2011driven approaches to improve reliability and efficiency\n\n\n\nMonitoring, Reliability \u0026 Incident Management\n\n
* Design and operate monitoring and observability systems using:\n
* Prometheus\n
* Grafana\n
* ELK stack\n\n\n
* Improve alert quality, signal\u2011to\u2011noise ratio, and troubleshooting efficiency\n
* Lead incident response activities, root cause analysis, and post\u2011incident reviews\n
* Support software engineers in debugging complex production issues across distributed systems\n
* Embed reliability, automation, and operational readiness into system design\n\n\n\nExperience Required\n\n
* Extensive experience operating large\u2011scale distributed cloud systems\n
* Hands\u2011on experience with AWS in production environments\n
* Direct experience working with OpenStack\n
* Strong Linux background in large\u2011scale SaaS or production systems\n
* Ability to:\n
* Maintain and improve existing mission\u2011critical systems\n
* Prioritize and systematically reduce technical and infrastructure debt\n\n\n
* Strong understanding of designing for operational excellence, not just greenfield solutions\n\n\n\nRequired Skills\n\n
* Programming: Strong experience with Python and/or Go\n
* Cloud \u0026 IaC: Terraform, Ansible, CloudFormation or equivalent\n
* Containers: Kubernetes (production experience)\n
* CI/CD: Jenkins and modern CI/CD practices\n
* Data \u0026 Streaming: Kafka, NiFi, Elasticsearch, MySQL, Vertica, Zookeeper\n
* Observability: Prometheus, Grafana, ELK\n
* Infrastructure: Nginx, Linux internals\n
* AI / Automation (advantage):\n
* Experience integrating AI or LLMs into operational workflows\n
* Familiarity with agent\u2011based automation concepts\n\n\n\nExperience Guidelines\n\n3+ years in:\n\n
* overall experience managing infrastructure\n
* Linux administration in large\u2011scale environments\n
* operating production systems on AWS and/or OpenStack\n
* managing Kubernetes in production\n
* using infrastructure as code\n
* working with CI/CD systems\n\n\n\nMinimum Qualificatio