Principal SRE – Cloud Automation
hace 7 días
Job Description
This role requires a SRE mindset combined with AI/ML expertise and strong application engineering skills across public and private cloud environments.
Responsibilities
Key Responsibilities
- End-to-end service ownership: design for telemetry, security, resiliency, scalability, and performance; lead sizing/architecture; drive service health reviews and process simplification.
- Incident management and prevention: lead postmortems/RCAs, coordinate fixes, define repair items, and implement data-driven prevention and continuous improvement.
- AI/ML and GenAI delivery: design and integrate solutions with LLMs, RAG, agentic workflows, and conversational AI; build low-latency model serving and retraining pipelines.
- Application engineering: develop performant microservices for distributed, containerized, cloud-native systems.
- Automation: eliminate toil by automating operational workflows, recovery procedures, code delivery, and configuration management; build internal tools and reusable scripts/services to accelerate delivery and reduce errors.
- Observability: define and implement monitoring, logging, alerting, and tracing strategies; establish SLOs/SLIs/error budgets; improve diagnostics and performance visibility for rapid triage.
- Cross-functional collaboration: partner with product, operations, and data teams to translate requirements into secure, scalable solutions; communicate effectively with technical and non-technical stakeholders.
Minimum Qualifications
- BS/MS in Computer Science or related field; 10+ years of software engineering in cloud environments.
- Strong in distributed systems/microservices using java / python; SQL/data modeling; python for AI/automation.
- SRE/DevOps expertise: systems and networking fundamentals, application security, observability, performance analysis, and incident response.
- Proven SDLC excellence: code quality, reviews, version control, CI/CD, testing, and release engineering.
- Excellent written and verbal communication; English fluency.
Preferred/Technical Skills
- AI/ML/GenAI: experience with foundational models, RAG, agentic architectures; model deployment, optimization, monitoring, and retraining.
- Cloud and containers: experience with containerization, orchestration, and resilient, fault-tolerant microservices.
- Observability: hands-on experience designing dashboards, alerts, traces, logs, and metrics; defining SLOs/SLIs and error budgets; on-call readiness and runbook quality.
- Operations: performance tuning across java / python and SQL for large-scale enterprise applications; strong Linux/Unix expertise; capacity planning and reliability reviews.
- Automation and scripting: proficiency in scripting to automate operational workflows, build tooling, and CI/CD tasks (e.g., shell scripting, python, configuration-as-code, task runners).
- Familiarity with enterprise ERP applications and standard DevOps tooling and practices.
Qualifications
Career Level - IC4
About Us
As a world leader in cloud solutions, Oracle uses tomorrow's technology to tackle today's challenges. We've partnered with industry-leaders in almost every sector—and continue to thrive after 40+ years of change by operating with integrity.
We know that true innovation starts when everyone is empowered to contribute. That's why we're committed to growing an inclusive workforce that promotes opportunities for all.
Oracle careers open the door to global opportunities where work-life balance flourishes. We offer competitive benefits based on parity and consistency and support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation- or by calling in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
-
Zapopan, Jalisco, México Oracle A tiempo completoDescriptionOwn and scale mission-critical ERP/SaaS services while building intelligent, cloud-native capabilities. This role requires a SRE mindset combined with AI/ML expertise and strong application engineering skills across public and private cloud environments.ResponsibilitiesKey Responsibilities- End-to-end service ownership: design for telemetry,...
-
Site Reliability Developer 4
hace 2 semanas
Zapopan, Jalisco, México Oracle A tiempo completoWe are looking for a skilled and motivated Cloud Region Build Site Reliability Engineer (SRE) to join our Oracle Cloud Infrastructure Region Build team. In this role, you will be responsible for building, deploying, and maintaining compute cloud infrastructure services across multiple regions to ensure high availability, scalability, and performance. You...
-
Site Reliability Developer 4
hace 5 días
Zapopan, Jalisco, México Oracle A tiempo completoDescriptionWe are looking for a skilled and motivated Cloud Region Build Site Reliability Engineer (SRE) to join our Oracle Cloud Infrastructure Region Build team. In this role, you will be responsible for building, deploying, and maintaining compute cloud infrastructure services across multiple regions to ensure high availability, scalability, and...
-
Principal Site Reliability Developer
hace 5 días
Zapopan, Jalisco, México Oracle A tiempo completoDescriptionWork with the Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, focusing...
-
Blockchain Senior DevOps/SRE
hace 2 días
Zapopan, Jalisco, México Oracle A tiempo completoJob DescriptionOracle Blockchain Platform is an enterprise-grade, pre-assembled platform, ready for blockchain application development and deployment in Oracle public cloud and other infrastructures.At Oracle Development Center in Zapopan we have built a team ground up working on various cloud service platforms that are running on Oracle's next-generation...
-
Zapopan, Jalisco, México Oracle A tiempo completoOracle & Data Core (Mandatory): The candidate must demonstrate strong expertise in Oracle Database 19c and above, with advanced proficiency in SQL—covering CTEs, analytic functions, and window functions. Intermediate to advanced PL/SQL skills are required for developing packages, jobs, and robust exception handling. Hands-on experience with Materialized...
-
Site Reliability Developer 3
hace 2 semanas
Zapopan, Jalisco, México Oracle A tiempo completoDescriptionSolve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems....
-
Site Reliability Developer 4
hace 2 semanas
Zapopan, Jalisco, México Oracle A tiempo completoJob DescriptionWe are looking for a skilled and motivated Cloud Region Build Site Reliability Engineer (SRE) to join our Oracle Cloud Infrastructure Region Build team. In this role, you will be responsible for building, deploying, and maintaining compute cloud infrastructure services across multiple regions to ensure high availability, scalability, and...
-
Principal Site Reliability Developer
hace 5 días
Zapopan, Jalisco, México Oracle A tiempo completoJob DescriptionWork with an elite team to provide Oracle Database Administration support for customer production systems in the Oracle Cloud, with the opportunity to work on the latest Oracle database releases and features as part of the cloud first strategy. Provide DBA operational support with a high degree of customer service, technical expertise, and...
-
Site Reliability Developer 3
hace 2 semanas
Zapopan, Jalisco, México Oracle A tiempo completoDescription"The Oracle Cloud Infrastructure (OCI) team can provide you the opportunity to build and operate a suite of massive scale, integrated cloud services in a broadly distributed, multi-tenant cloud environment. OCI is committed to providing the best in cloud products that meet the needs of our customers who are tackling some of the world's biggest...