Senior Data Engineer

Hace 3 días

México Deel Jornada completa $2 - $3 Por obra

Who We Are Is What We Do.

Deel is the all-in-one payroll and HR platform for global teams. Our vision is to unlock global opportunity for every person, team, and business. Built for the way the world works today, Deel combines HRIS, payroll, compliance, benefits, performance, and equipment management into one seamless platform. With AI-powered tools and a fully owned payroll infrastructure, Deel supports every worker type in 150+ countries—helping businesses scale smarter, faster, and more compliantly.

Who We Are Is What We Do.

Deel is the all-in-one payroll and HR platform for global teams. Our vision is to unlock global opportunity for every person, team, and business. Built for the way the world works today, Deel combines HRIS, payroll, compliance, benefits, performance, and equipment management into one seamless platform. With AI-powered tools and a fully owned payroll infrastructure, Deel supports every worker type in 150+ countries—helping businesses scale smarter, faster, and more compliantly.

Among the largest globally distributed companies in the world, our team of 7,000 spans more than 100 countries, speaks 74 languages, and brings a connected and dynamic culture that drives continuous learning and innovation for our customers.

Why should you be part of our success story?

As the fastest-growing Software as a Service (SaaS) company in history, Deel is transforming how global talent connects with world-class companies – breaking down borders that have traditionally limited both hiring and career opportunities. We're not just building software; we're creating the infrastructure for the future of work, enabling a more diverse and inclusive global economy. In 2024 alone, we paid $11.2 billion to workers in nearly 100 currencies and provided healthcare and benefits to workers in 109 countries—ensuring people get paid and protected, no matter where they are.

Our momentum is reflected in our achievements and customer satisfaction: CNBC Disruptor 50, Forbes Cloud 100, Deloitte Fast 500, and repeated recognition on Y Combinator's top companies list—all while maintaining a 4.83 average rating from 15,000 reviews across G2, Trustpilot, Captera, Apple and Google.

Your experience at Deel will be a career accelerator. At the forefront of the global work revolution, you'll tackle complex challenges that impact millions of people's working lives. With our momentum—backed by a $17.3 billion valuation and $1 B in Annual Recurring Revenue (ARR) in just over five years—you'll drive meaningful impact while building expertise that makes you a sought-after leader in the transformation of global work.

Our Data Platform team is rebuilding how Deel turns operational data into data products. We're moving from a batch, trigger-based system to a streaming architecture built on change data capture (CDC), Spark Structured Streaming, and Apache Iceberg.

The goal is to deliver business-critical data models and transformations in under 5 minutes, make their cost visible, and let engineering teams onboard their own models. This platform is the foundation for Deel's customer analytics and audit log features.

You’ll own this platform end to end. You'll design the streaming jobs, run the compute clusters they live on, keep the lakehouse tables healthy, and make the whole system easier and cheaper to operate every week.

What You’ll Do

  • Design, build, and harden Spark Structured Streaming applications that process CDC events into near-real-time data models. This includes stateful processing, slowly changing dimensions (SCD), late and partial records, and safe restarts.
  • Run and scale our AWS EMR clusters. You'll own instance sizing, managed scaling, multi-cluster deployment, driver and executor memory tuning, and consolidation of low-traffic workloads to cut cost.
  • Own the health of our Apache Iceberg tables. You'll design partition specs and run compaction, snapshot expiry, and delete-file cleanup at scale, including fast-lane maintenance for high-churn tables.
  • Onboard new business domains to the streaming platform. You'll configure CDC sources and destinations in Estuary Flow, set up Kafka-compatible topics, verify backfills in Snowflake, and coordinate with source-database owners.
  • Debug production issues across the stack and fix the root cause, not the symptom. Examples include connection drops, offset regressions, token expiry, schema mismatches, and write conflicts.
  • Build observability and tooling that shorten time to diagnosis. Examples include shipping Spark driver and executor logs to Grafana Loki, watchdog and recovery steps, and AI-assisted runbooks.
  • Improve deployment and developer experience through CI/CD pipelines, configuration structure, and guardrails that let other teams ship data models safely on their own.
  • Handle ad hoc ingestion/backfill requests and es