Sr. Data Engineer

Hace 23 horas

México YemPover Inc Jornada completa

•

Experience:
8+ years in data engineering, with a strong focus on large-scale, distributed, cloud-native systems.
• Snowflake depth (primary): Proven, hands-on production experience with Snowflake — warehouse tuning and cost optimization, RBAC and data governance, Streams/Tasks, Snowpipe, Time Travel, and Snowpark. You can reason about query profiles and micro-partitioning, not just write SQL.
• Core languages: Expert Python and strong, advanced SQL; PySpark/Snowpark a plus.
• Transformation & orchestration: Production experience with dbt and a workflow orchestrator (Airflow strongly preferred), designing for portability across engines.
• AWS ecosystem: Solid hands-on experience with S3, Glue, Step Functions, IAM, and event services (SNS/SQS).
• Streaming: Track record building streaming/CDC pipelines with Kinesis or Kafka and tools like Debezium, Fivetran, or Airbyte.
• Data quality: Demonstrated ownership of automated testing, data profiling, reconciliation, and pipeline validation — you own the QA of your own pipelines.
• Ways of working: Strong documentation habits (playbooks, technical specs), a bias toward reproducibility and version control, and a clear ownership mindset.
• Top desirable — Automotive domain: Direct experience with automotive repair-order data, dealership fixed-operations, or DMS feeds is a top differentiator for this role.
• Top desirable — Databricks: Hands-on Databricks experience (Spark, Delta Lake, Unity Catalog, Workflows) used alongside Snowflake. We value portability across engines and may extend the stack over time, so multi-platform depth is highly valued.
• Certifications: SnowPro Core or Advanced (Data Engineer), Databricks Certified Data Engineer Professional, or AWS Certified Data Engineer.
• Governance & lineage: Exposure to cataloging/lineage tooling (OpenMetadata) and modern CI/CD for data (tested, peer-reviewed deployments).

Key Responsibilities
1. Snowflake Warehouse Engineering
• Design and maintain governed, well-modeled schemas in Snowflake using dimensional and Medallion (bronze/silver/gold) patterns, applying dbt for modular, version-controlled transformations.
• Optimize Snowflake for performance and cost — right-size virtual warehouses, manage clustering keys and micro-partition pruning, tune query profiles, and control credit consumption through resource monitors and warehouse scaling policies.
• Leverage native Snowflake features such as Streams and Tasks (or external orchestration), Snowpipe / Snowpipe Streaming for continuous ingestion, Dynamic Tables, Time Travel, Zero-Copy Cloning, and Secure Data Sharing.
• Implement governance with role-based access control (RBAC), row-access and masking policies, and object tagging for PII and sensitive automotive/customer data. 2. Pipeline Development & AWS Data Lake Engineering
• Build and orchestrate complex ETL/ELT pipelines through a central control plane (e.g. Airflow with astronomer-cosmos) that runs consistently across AWS, Snowflake, and Databricks, favoring portable, engine-agnostic designs over transformations locked into any single platform's native constructs.
• Manage the data lake on AWS S3 as the primary landing zone, optimizing open storage formats (Iceberg, Parquet, Delta) and integrating them with Snowflake external tables and Iceberg tables.
• Build decoupled, event-driven architectures using SNS and SQS for high-throughput messaging between data services. 3. Real-Time Data Streaming & Ingestion
• Develop low-latency ingestion using AWS Kinesis or Kafka, feeding operational analytics through Snowpipe Streaming and Kafka connectors.
• Implement Change Data Capture (CDC) with Debezium, Fivetran, or Airbyte to keep Snowflake in near-real-time sync with source systems windowed on update timestamps. 4. Data Quality, Reconciliation & QA Ownership
• Own end-to-end data quality by embedding automated tests directly into ETL/ELT flows — using portable frameworks such as dbt tests and Great Expectations that apply uniformly whether transformations run on AWS, Snowflake, or Databricks — rather than treating QA as a separate downstream step.
• Build cross-engine reconciliation between source systems and Snowflake — control totals, partition/row hashing, and drill-down on discrepancies — to guarantee completeness and correctness at scale.
• Enforce data contracts and schema-evolution guidelines, and stand up proactive observability and alerting to catch data drift and pipeline anomalies before they reach users. 5. Engineering for ML/AI
• Engineer ML-ready datasets and feature pipelines to support the Data Science team, managing feature freshness and lineage.
• Operationalize AI workflows using Snowflake Cortex, Snowpark for Python, and AWS Bedrock where appropriate. 6. Technical Leadership & Collaboration
• Mentor engineers on coding standards, SQL and Snowflake optimization, dbt modeling, and Python development, and lead code reviews.
• Partner with architects, product, and analytics teams to translate designs into reliable, documented, production code, and help drive SDLC and CI/CD maturity across the data platform.