Ingeniero de datos
Hace 1 día
Mexico
NTT DATA Europe & Latam
Jornada completa
Gratis con email o Google
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Gratis con email o Google
Location:
LATAM
- 100% remote Contract Length: 12 months Contract Rate: 20 USD per hour NTT DATA is seeking a driven Data Engineer (Denodo) to identify and cultivate new business opportunities within the CPG industry, with a focus on our client in the Bottle and Brewing segment. This is a performance-based role centered on relationship development and market access, designed for a professional with established networks in the CPG space who is ready to open doors and create meaningful introductions for our sales and delivery teams.
Responsibilities:
- Design, build and maintain the Denodo semantic layer, translating business requirements into governed, reusable views and data products published to the Data Catalog and Data Marketplace.
- Optimise query performance end to end through delegation, caching, summary views and model refactoring, so that reports and dashboards meet agreed response time targets.
- Integrate new sources (AWS, Oracle, Redshift, APIs and enterprise systems) into the virtualization layer, defining connection, security and refresh strategies for each.
- Build and evolve the gold layer: curated dimensional models, certified metrics and aggregates that serve as the single point of consumption for analytics teams.
- Support the visualization teams as the technical counterpart for data access: model adjustments, aggregate design and troubleshooting of slow or failing reports.
- Develop and maintain ingestion and transformation pipelines with Informatica and AWS Glue where physical processing is required upstream of the virtual layer.
- Apply and document development standards, naming conventions, modelling patterns and promotion procedures across Dev/QA/Prod environments.
- Provide L2/L3 support for the virtualization platform: incident analysis, root-cause investigation, cache and workload tuning, and continuous-improvement backlog.
- Collaborate with data governance and security teams to enforce access policies, lineage capture and metadata quality on every published asset.
- Contribute to technical estimation, solution design reviews and knowledge transfer to other engineers. Required
Experience:
- Logical modelling and layering: base views, derived views (join, union, selection, projection, aggregation, flatten) and interface views; disciplined separation between source, integration and business/semantic layers; naming conventions, reusability and version control of metadata.
- Semantic layer for data products: business-friendly models, certified KPIs and metrics, association and hierarchy definition, field-level descriptions and business metadata so that consumers can self-serve without re-modelling.
- Advanced SQL: complex queries, stored procedures, custom functions.
- Query optimization (core requirement): cost-based optimizer, statistics management, query pushdown and delegation to Redshift/Oracle/S3, branch pruning, join strategy selection (merge, hash, nested), partitioned unions, and execution-trace analysis to diagnose and remediate slow queries.
- Acceleration and caching: full, partial and incremental cache strategies, cache refresh scheduling and invalidation, summary views and smart query acceleration, and MPP acceleration over object storage.
- Visualization performance: tuning of the consumption layer for Power BI, Tableau and similar tools: live vs. import trade-offs, aggregate awareness, result set reduction, concurrency and workload management to keep dashboard response times within agreed SLAs.
- Publication and consumption: Data Catalog and Data Marketplace publication, tagging and categorization. Nice-to-Have:
- Informatica: IDMC (Cloud Data Integration, Cloud Mass Ingestion) or PowerCenter, with mapping and taskflow development, pushdown optimization, and data quality rules feeding the curated layers.
- AWS Glue: PySpark jobs, crawlers and Data Catalog management, and Glue workflows for ingestion and transformation into S3/Redshift.
- Orchestration and DataOps: AWS Step Functions or equivalent schedulers, Git-based workflows, CI/CD pipelines, code review and environment promotion practices.
- Python and SQL for automation, validation and data reconciliation. Tools & Technologies:
- Medallion architecture: clear understanding of bronze, silver and gold responsibilities, and where virtualization complements or replaces physical persistence at each stage.
- Gold layer construction: dimensional and star-schema modelling, conformed dimensions, slowly changing dimensions, aggregates and business-rule implementation for consumption-ready models.
- Data products: definition of data contracts, ownership, quality expectations, SLAs and versioning; documentation and lifecycle of published assets.
- Working knowledge of Data Mesh and Lakehouse principles and of data governance practices (lineage, cataloguing, stewardship). Education & Certifications:
- Denodo Platform 9 Certified Developer Associate or Professional
- Denodo Platform 9 Certified Administrator Assoc