Senior MLOps Engineer
Hace 7 horas
Mexico, chihuahua
EPAM Systems, Inc.
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Al continuar, aceptas nuestros Términos & Política de Privacidad.
We are seeking a Senior MLOps Engineer to join an MVP engagement with a major AAA game publisher, building a test intelligence platform for two game franchises in parallel. A core design principle is full re-derivability and model lineage from day one — every run must be replayable from its stored configuration version and feed read positions. The signal catalog feeds a scoring strategy engine with versioned configurations, and calibration sweeps over historical data produce suggested weight updates surfaced directly in the Settings View.
This role ensures the ML and signal components are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.
ResponsibilitiesOwn Signal Catalogue operations: signal refresh orchestration triggered by feed read-position advances, grain translation between per-test, per-area, and per-run signal families, and provenance capture across all 8 signalsOperate the Semantic Vector Index versioning: coordinate with the Senior AI Developer on model and dimension stamp conventions; design and execute the controlled reindex path when the enterprise AI gateway model changesDesign, implement, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runnerEnforce holdout patch-set discipline, configuration sweep over route limits, thresholds, weights, and Composition setting; produce catch-rate vs. scope tables per candidate configuration and publish winning configurations as suggested-weight proposals into the Settings ViewLead Model Generation & Experimentation: systematic experimentation framework over scoring strategy configurations, tracking which signal weights and route combinations yield the best catch-rate vs. scope trade-offMaintain model lineage across configuration versions for both franchisesImplement the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channelsManage exploration cadence support: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendationsContribute to operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex proceduresRequirements3+ years of experience in MLOps or ML platform engineeringExpertise in ML model lifecycle management, including versioning, configuration management, and rollbackBackground in signal computation pipeline design, covering scheduled refresh, provenance capture, and grain translationProficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurementKnowledge of as-of temporal data systems or back-test harness design and operationSkills in Python and SQL for ML pipeline automationCompetency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift countersCapability to collaborate with data engineers and AI developers on feature alignmentQualifications in documenting calibration procedures, signal definitions, and operational runbooksFamiliarity with Spec Driven DevelopmentEnglish proficiency at an Upper-Intermediate level (B2) or higherNice to haveUnderstanding of MLflow, Kubeflow, or equivalent experiment tracking platformsFamiliarity with Snowflake ML or SnowparkShowcase of embedding model versioning and controlled reindex orchestrationSkills in pgvector or vector store operational managementBackground in gaming domain or QA toolingWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
This role ensures the ML and signal components are production-ready, reproducible, and improvable over time, forming the foundation of the system's long-term value as franchise history accumulates and models are refined.
ResponsibilitiesOwn Signal Catalogue operations: signal refresh orchestration triggered by feed read-position advances, grain translation between per-test, per-area, and per-run signal families, and provenance capture across all 8 signalsOperate the Semantic Vector Index versioning: coordinate with the Senior AI Developer on model and dimension stamp conventions; design and execute the controlled reindex path when the enterprise AI gateway model changesDesign, implement, and own the Back-test & Calibration Harness: as-of temporal filtering across all record families, replay runner, look-ahead spot audit, and configuration sweep runnerEnforce holdout patch-set discipline, configuration sweep over route limits, thresholds, weights, and Composition setting; produce catch-rate vs. scope tables per candidate configuration and publish winning configurations as suggested-weight proposals into the Settings ViewLead Model Generation & Experimentation: systematic experimentation framework over scoring strategy configurations, tracking which signal weights and route combinations yield the best catch-rate vs. scope trade-offMaintain model lineage across configuration versions for both franchisesImplement the MLOps Monitor and Data-Health Monitor: catch-rate floor monitoring, run-behaviour drift counters (per-run candidate volumes per route, score distribution vs. usual range), data-health telemetry across all ingestion channelsManage exploration cadence support: unbiased random-sample injection with provenance ensuring exploration entries are never counted as model recommendationsContribute to operator runbook sections covering signal refresh, calibration campaigns, model generation runs, and embedding reindex proceduresRequirements3+ years of experience in MLOps or ML platform engineeringExpertise in ML model lifecycle management, including versioning, configuration management, and rollbackBackground in signal computation pipeline design, covering scheduled refresh, provenance capture, and grain translationProficiency in calibration methodology: holdout discipline, configuration sweep design, and catch-rate vs. scope measurementKnowledge of as-of temporal data systems or back-test harness design and operationSkills in Python and SQL for ML pipeline automationCompetency in model monitoring, including drift detection, catch-rate floor monitoring, and run-behavior drift countersCapability to collaborate with data engineers and AI developers on feature alignmentQualifications in documenting calibration procedures, signal definitions, and operational runbooksFamiliarity with Spec Driven DevelopmentEnglish proficiency at an Upper-Intermediate level (B2) or higherNice to haveUnderstanding of MLflow, Kubeflow, or equivalent experiment tracking platformsFamiliarity with Snowflake ML or SnowparkShowcase of embedding model versioning and controlled reindex orchestrationSkills in pgvector or vector store operational managementBackground in gaming domain or QA toolingWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.