Data Engineer
Aviso de fuente externaen Elios AI
Data Engineer Location: Remote (Mexico) | Type: Contract | Experience: 5+ years | Pay: Up to $50/hr About the RoleWe're hiring two Data Engineers to build and operate the data layer of an...
Salario
No especificado
Ubicación
Mexico City, México
Tipo de empleo
No especificado
Modalidad
No especificado
Data Engineer
Mexico City, México
Descripción del empleo
Data Engineer
Location: Remote (Mexico) | Type: Contract | Experience: 5+ years | Pay: Up to $50/hr
About the RoleWe're hiring two Data Engineers to build and operate the data layer of an enrichment tool running on an enterprise client's Databricks lakehouse. The design work is already done. There's an accepted solution design, 13 data contracts, and 13 ADRs. What's left is standing up 26 governed tables across core, ledger, and serving layers and making them hold under real production load.
You'll be working through a digital product studio backed by a global strategy consulting parent, embedded with their engineers on a program for one of their largest enterprise accounts. The client teams sit in Mexico, so the day-to-day runs in Spanish: standups, design discussions, and documentation.
If you'd rather implement against written contracts than relitigate schema decisions every sprint, this is that kind of program. The decisions are made. The question is whether the pipeline runs clean, the quality gates hold, and the numbers reconcile.
What You'll DoBuild production PySpark jobs on Databricks: scheduled workflows, Delta merge semantics, Unity Catalog schemas and permissions, deployment through Asset Bundles or job YAMLStand up a net-new governed schema alongside existing platform tables, following medallion patterns across core, ledger, and serving layersOwn governed file intake end to end, including landing, validation, versioning, checksums, quality gates, and the rejection and replay paths when a file failsIntegrate with external systems over REST, both consuming and publishing, as the planning exchange moves off files and onto direct API in both directionsDesign the batch computation layer: proportional disaggregation through the cascade engine, precomputed aggregates for serving, accuracy and bias scoring, materialized measurement tablesImplement the dual store pattern, with a PostgreSQL operational database publishing to Delta at cycle close, plus append-only ledger and audit designIngest model outputs from the data science pipeline (SHAP driver attribution, conformal prediction intervals) into governed tablesPut row-level security and Entra ID based access patterns in place at the database layerSet up monitoring and alerting so batch failures and data quality breaks surface before the client finds themWork to the defined grain, keys, quality rules, and ownership in each data contract, and push back when something in the design doesn't survive contact with real data
QualificationsCore Data Engineering5+ years building production data pipelines, with real ownership of what runs on a scheduleDatabricks and PySpark at production level, not notebook-only exposureDelta Lake merge semantics, Unity Catalog schemas and permissions, and deployment via Databricks Asset Bundles or job YAMLMedallion architecture experience, ideally building a new governed schema next to tables you don't controlPostgreSQL past the ORM, including schema design for an operational store that publishes downstreamGovernance and IntegrationBuilt file intake that had to survive bad inputs: validation, versioning, checksums, rejection handling, replayREST integration with authentication on both the consuming and publishing sideRow-level security and Entra ID or equivalent identity-based access patterns in the data layerComfortable working to data contracts with defined grain, keys, quality rules, and named ownershipMonitoring and alerting for batch and data quality jobsLanguageFluent Spanish, spoken and written, at a level that supports technical discussion and documentationProfessional English for collaboration with the US-based studio teamNice to HaveExperience with batch computation design: disaggregation, precomputed aggregates, accuracy or bias scoringExposure to consuming data science outputs (feature attribution, prediction intervals) into governed tablesConsulting or agency background delivering directly to enterprise stakeholders
Why Join UsThe studio calls its engineers crafters, and the bar shows up in the details: end-to-end ownership, short delivery cycles, and no long approval chains. You'll have a clear design to build against and the autonomy to decide how it gets built.
This is long-term contract work on a platform that other teams will build on top of, with people who care whether it holds up six months from now.
¿Es tuya esta vacante?
Reclámala gratis y recibe candidatos con video en CazVid.