Cloud Platform Technical Lead ID92209

Hace 2 días

Jalisco, México AgileEngine, LLC. Jornada completa

5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
We are looking for a Cloud Platform Technical Lead to own the reliability and orchestration of an enterprise data platform in a regulated healthcare environment.
6+ years of professional experience in Cloud Engineering, Platform Engineering, DevOps or Site Reliability Engineering
~ Deep hands‐on expertise with AWS production environments , including cloud architecture, security, networking fundamentals, access management, monitoring, capacity, and operational troubleshooting.
~ Advanced experience with Kubernetes and Amazon EKS , including workload deployment, cluster and application troubleshooting, observability, scaling, upgrades, access, and reliability.
~ Hands‐on experience operating Argo Workflows or a comparable orchestration platform supporting production data workloads.
~ Strong experience with Infrastructure as Code, preferably Terraform , and with source‐controlled configuration, CI/CD, release automation, and rollback practices.
~ Strong experience designing and operating observability, logging, monitoring, alerting, and incident‐routing solutions for distributed production platforms.
~ Demonstrated ability to lead major incidents and cross‐functional troubleshooting across infrastructure, applications, data pipelines, and analytics layers.
~ Experience operating production platforms with defined service levels, escalation paths, runbooks, change controls, release processes, and on‐call responsibilities.
~ Strong working knowledge of modern data platforms, including Snowflake, S3‐based data lakes, SQL, dbt, managed ingestion tools such as Fivetran or HVR, and custom data pipelines .
~ Understanding of data quality, freshness, lineage, schema evolution, pipeline dependencies, backfills, and recovery procedures.
~ Proficiency in Python, Shell, Bash, or comparable languages for automation and operational tooling.
~ Strong written and verbal English communication skills.
~ Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed senior escalation and on‐call rotation .

Experience with data observability platforms such as SYNQ and operational tooling such as Splunk, PagerDuty, Opsgenie, or comparable solutions.
Experience implementing dependency‐aware alerting, selective auto‐remediation, impact analysis, or other advanced reliability practices.
Own the technical direction and end‐to‐end reliability of the managed Data Platform across AWS, Amazon EKS, Kubernetes, Argo Workflows, Snowflake, S3, Fivetran and custom ingestion pipelines, and Tableau dependencies.
Establish and evolve cloud architecture principles, Infrastructure as Code standards, CI/CD and release practices, observability patterns, operational controls, and platform engineering priorities.
Provide technical leadership for AWS infrastructure, Kubernetes and EKS operations, Argo Workflows, deployment automation, secrets and access management, platform capacity, and production reliability.
Guide cross‐platform decisions involving Snowflake, dbt, Fivetran, custom ingestion, S3 data‐lake operations, workflow orchestration, data quality, and downstream analytics dependencies.
Lead complex cross‐platform troubleshooting and L3 escalation, coordinating engineers when incidents span infrastructure, ingestion, orchestration, Snowflake, and Tableau.
Direct technical response during major incidents, including impact assessment, stakeholder communication, recovery strategy, root‐cause analysis, post‐incident review, and preventive actions.
Ensure that changes and releases include appropriate technical review, dependency analysis, validation, rollback planning, traceability, and coordination across affected platform components.
Define and monitor reliability indicators such as availability, ingestion success, data freshness, workflow reliability, deployment outcomes, alert quality, mean time to restore service, and recurring incident patterns.
Lead improvements in automation, observability, auto‐remediation, platform resilience, deployment safety, performance, security, capacity management, and cloud cost efficiency.
Partner with Security, Governance, Analytics, IT, and platform stakeholders to ensure least‐privilege access, secrets management, auditability, controlled changes, and appropriate handling of regulated data.
Mentor Senior and Middle‐level engineers, delegate ownership effectively, improve team practices, and build consistent technical capability across the distributed team.
Maintain alignment between the LatAm and India coverage windows through clear ownership, escalation paths,