Senior Data Engineer
Ref Id: HREZ-16365
Published Date: Aug 26, 2026
Job location: Remote, Latin America
General information
Ref Id: HREZ-16365
Published Date: Aug 26, 2026
Job location: Remote, Latin America
Description
HATCHWORKS AI

Data Engineer (Senior)

Lakehouse / Delta / Enterprise Integration  ·  Remote / LatAm  ·  Colombia, Costa Rica, Brazil  ·  6+ Yrs Exp.  ·  English Required  ·  Full-Time

ABOUT HATCHWORKS AI

HatchWorks AI is an enterprise AI consulting firm. We design, build, and operate production-grade AI and software systems for clients across regulated and high-growth industries. Our teams pair senior architects and engineers with AI-assisted delivery methods to move clients from discovery to execution faster, without sacrificing rigor.
 

THE ROLE

There's no AI without data. There's no data without engineering. You know this — and you've built the pipelines that prove it.
We're looking for Data Engineers at the Senior level to build and maintain the data foundation for production AI systems. You will design and run pipelines from enterprise source systems (ERP, planning, CRM, regulatory documents) into curated gold-layer data products on a modern Lakehouse — the substrate every model, every optimizer, every rule depends on.
This is not an analytics role. This is data infrastructure for AI in production — high-throughput, schema-disciplined, observable, and trusted by downstream engineers who can't afford garbage in.
If you want your pipelines to be the load-bearing layer of an AI platform — this is the role.
 

WHAT YOU'LL ACTUALLY DO

  • Build and maintain Lakehouse pipelines — bronze → silver → gold — with Delta Lake (or equivalent) as the production format.
  • Integrate with enterprise source systems — ERP platforms, planning systems, CRM, operational sources — through sanctioned, approved integration patterns. Read-first is the default; write-back is an explicit design decision, not an assumption.
  • Ingest semi-structured documents (PDF, spreadsheet, structured exports) into the gold layer — substrate for downstream knowledge-engineering and rule authoring.
  • Design schema discipline for AI consumption — typed, versioned, contract-tested, drift-detected.
  • Engineer structured feedback-capture pipelines — user override and correction events become the labeled-data source for downstream model retuning. The loop only closes if the data is well-engineered.
  • Harden the integration surface — service-account management, rate-limit configuration, security review work that is part of production cutover, not an afterthought.
  • Implement data quality gates — Great Expectations, dbt tests, or equivalent — that block bad data before it reaches a model.
  • Optimize for cost and performance — partitioning, Z-ordering, compaction, caching strategies. Production economics matter.
  • Partner with ML, optimization, and knowledge engineers — they're your customers. The data they get is the data you ship.
  • Pair with client engineers embedded in the team — ownership transfer is a delivery requirement.
  • Instrument observability — pipeline SLAs, freshness monitors, lineage tracking, anomaly detection.
  • Contribute to HatchWorks AI's data engineering practice — Lakehouse patterns, AI-ready data products, federated architectures.
 

THE TECHNICAL BAR

You will be expected to execute hands-on technical work from day one. The requirements below reflect the actual skills needed to deliver outcomes for enterprise clients.
Must-Haves
  • Python + SQL — 5+ years (Senior), production-grade
  • PySpark — Production experience on Lakehouse platforms
  • Delta Lake / Iceberg — Schema evolution, time travel, merge patterns at scale
  • Pipeline Orchestration — Airflow, Lakehouse-native workflows, or equivalent
  • Data Quality — Great Expectations, dbt tests, or production-grade validation frameworks
  • Cloud — Production depth in a major cloud platform
  • Enterprise Integration — Senior: ERP integration in production. Mid: willingness to learn.
  • Security Patterns — Service-account hardening, IAM, rate-limiting in cloud-hosted Lakehouse
  • CI/CD — GitLab/Git, dbt CI, Lakehouse deployment automation
 

CORE TECH STACK

Full Stack
  • Languages — Python  ·  SQL  ·  PySpark
  • Lakehouse — Databricks  ·  Delta Lake  ·  Unity Catalog  ·  Iceberg (a plus)
  • Orchestration — Apache Airflow  ·  Lakehouse-native workflows  ·  Prefect
  • Quality / Testing — Great Expectations  ·  dbt  ·  pytest
  • Integration — Enterprise ERP / planning systems  ·  REST  ·  Kafka  ·  CDC patterns
  • Cloud — Major cloud platforms hosting modern Lakehouse
  • DevOps — Docker  ·  CI/CD  ·  GitLab  ·  Git
 

NICE TO HAVE

  • Manufacturing, logistics, or industrial-operations industry experience
  • Knowledge graphs / RDF — integrating structured + unstructured data
  • Streaming / CDC at production scale (Kafka, Debezium)
  • Cost optimization on Lakehouse platforms at scale
  • dbt — for analytics engineering and metric layer
  • Data contracts and event-driven schema patterns
 

WORKING CONTEXT

This is an active, client-embedded delivery engagement. You will ship production pipelines from week one.
All team members operate within client timezone hours. LatAm candidates are strongly preferred.
We hire data engineers at both Senior and Mid-Level. The application process is the same — we calibrate based on what you've shipped, not on years claimed.
 

THE HUMAN BAR

Pipeline skill is necessary. Customer empathy is the differentiator.
The data engineers who thrive here treat ML and optimization engineers as their customers. The data those engineers receive is the product. Quality, schema, lineage, freshness — these are not back-office concerns. They are the thing.
  • AI-aware — you understand what downstream models need. You design schemas with feature engineering in mind, not just BI reporting.
  • Quality discipline — bad data caught at ingest is cheap. Bad data caught in a user's recommendation is expensive. You catch it early.
  • Cost-conscious — production Lakehouse runs at scale. You think about partitioning, compaction, and compute spend as design choices, not afterthoughts.
  • Delivery ownership — you scope, you build, you ship, you monitor. Pipelines are products. You own them.

WHAT MAKES THIS DIFFERENT

Most data engineering roles are sized to BI dashboards or ETL into a warehouse. This one is sized to AI in production. The pipelines you build feed forecasting models, optimization engines, and knowledge graphs — every one of them sensitive to schema drift, freshness, and quality.
You will be embedded with a major global enterprise client — operating in supply chain, logistics, manufacturing, or adjacent operational industries — building production AI systems that decision-makers actually use.
You will be part of a unified delivery team — your work isn't a back-office function. Your customers are ML and optimization engineers two desks away, and the consequences of a broken pipeline are visible the same day.
 

YOU'LL WORK INSIDE A PROPRIETARY AI DELIVERY FRAMEWORK

Most teams using AI tools today are flying blind — they feel faster but can't prove it. We built GenDD (Generative Driven Development) to change that, and you'll be working inside it from day one.
GenDD is a proprietary methodology and toolset built around one belief: great engineers need full context to move fast without breaking things. So we built the tools to give them exactly that.
Zero Ramp-Up
Our tooling surfaces full project context — architecture, data flows, history — so you hit the ground running on every engagement.
Proven Value
GenDD instruments every agent interaction, commit, and deploy — turning your work into clear, client-ready evidence of real impact.
All-In Team
Everyone here uses agents seriously. You're not evangelizing — you're exploring. We learn fast, share everything, and enjoy the work.
This is the environment technologists thrive in: structured methodology, tooling that solves real problems, teammates who move at pace, and clients who care about the results.
 


< Back to job list
Senior Data Engineer application
Name *
Email *
Phone number *
Location *
LinkedIn URL *
Resume *
SUBMIT
CANCEL
Powered byhireEZ