This one is closed
Live roles like this one
-
C
8h ago
Toptal LATAM, Canada, USA
-
B
6h ago
Staff Software Engineer, Enterprise Data Warehouse
HOME DEPOT U.S.A., INC. United States $120k - $190k/yr
-
B
8h ago
Software Engineering, Senior Advisor
Peraton United States $135k - $216k/yr
-
B
1d ago
Capital One Shopping Tech, Data Engineer 5 (Remote- Eligible)
Capital One United States $209k - $239k/yr
See every "Data Engineer" role →
Get new “Data Engineer” roles by email
One email a day with what is new in "Data Engineer". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 76/100, which is a B. It lost the most ground on remote clarity. See the breakdown
- Pay transparency 25 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Freshness 12 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Role specificity 6 / 10 Whether the listing is tagged well enough to tell what the role actually is.
- Corroboration 5 / 10 Whether more than one source carries this listing.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Senior Full Time
This role sits at the intersection of engineering, analytics, and platform operations. You will write production-grade Python and SQL, enforce data contracts, and help Protective treat data as a product with named owners and real consumers. You will be embedded on a delivery pod that owns its data products end to end, rather than servicing tickets from a queue.
On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description.
Key Responsibilities
- • Design, develop, and maintain production data pipelines on Databricks using Python, SQL, Apache Spark, and Delta Lake.
• Build Bronze-layer ingestion that reliably captures data from APIs, relational databases, flat files, cloud storage, and SaaS platforms — using dlt (dltHub) and Databricks-native ingestion where each fits — including incremental loading, pagination, watermarking, state management, and replay after failure.
• Develop Silver-layer transformations in dbt and Python over Delta Lake that cleanse, standardize, type, deduplicate, validate, conform, and enrich data so that it is reusable across domains. A meaningful share of this role is making messy source data trustworthy.
• Create Gold-layer data products: dimensional models, slowly changing dimensions, fact and bridge tables, aggregates, and serving tables aligned to how consumers actually query.
• Produce and maintain the curated datasets ML engineering trains and serves models from — feature and training tables that are versioned and reproducible, not one-off extracts.
• Author and maintain data contracts using the Open Data Contract Standard (ODCS) — schema with real semantics, named owner, known consumers, quality rules, and freshness expectations — and assess backward compatibility before every change.
• Implement data quality as code: uniqueness and not-null on keys at minimum, plus referential, accepted-value, freshness, and custom business-rule tests, surfaced to producers and consumers rather than buried in logs.
• Orchestrate ingestion and transformation as assets in Dagster, deployed to Dagster Cloud, and operate what you build across development, branch, and production deployments — schedules and sensors, asset dependencies, backfills, and run observability.
• Apply governance through Unity Catalog — catalogs, schemas, external locations, grants, row- and column-level security, and lineage — and handle credentials through Azure Key Vault rather than in code.
• Implement incremental and merge-based processing with Delta Lake (MERGE, schema evolution, time travel, OPTIMIZE) and tune Spark jobs, table layouts, and compute for performance and cost.
• Troubleshoot production failures, data-quality issues, source-system changes, and late-arriving or duplicate data — including backfills and recovery — and take part in the pod’s on-call rotation for the pipelines it owns, with root-cause analysis that closes the gap rather than reopening the ticket.
• Build and maintain CI/CD for data assets in Azure DevOps — automated tests and CI checks on dlt, dbt, and Dagster changes, promotion from development through branch deployments to production, and releases that are repeatable and auditable.
• Instrument what you own for observability: freshness, volume, quality, latency, and cost, with alerting tied to the SLAs and SLOs your contract commits to instead of depending on someone noticing.
• Work inside the platform’s control expectations — least-privilege access, secrets in Azure Key Vault, change management through pull request and pipeline, and audit evidence that falls out of the deployment path rather than being reconstructed later.
• Participate in code review and document architecture, runbooks, and data products so others can discover, trust, and reuse them.
• Work with data architects, analysts, product owners, and business stakeholders to translate requirements into maintainable data solutions.
Qualifications
Required Qualifications
• Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
• 3+ years building and supporting production data pipelines in a cloud data platform environment.
• Strong hands-on Python and SQL. Both are used daily and neither substitutes for the other.
• Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake tables, MERGE, and incremental load patterns.
• Practical understanding of medallion / multi-layer lakehouse design, and the judgment to say what belongs in Bronze versus Silver versus Gold.
• Experience ingesting data from APIs, relational databases, files, or SaaS applications, including the incremental and state-management problems that come with it.
• Working knowledge of dimensional modeling — grain, keys, facts and dimensions, slowly changing dimensions — and of ELT design patterns and data quality practice.
• Experience with orchestration and scheduling using Dagster, Databricks Workflows, Airflow, Azure Data Factory, or similar.
• Git-based source control, pull request review, automated testing, and CI/CD as normal practice — Azure DevOps or comparable.
• Experience troubleshooting production data failures, performance bottlenecks, and source-system changes.
• Experience with pipeline monitoring and alerting, and a working understanding of what a freshness or quality SLA means once real consumers depend on it.
• Ability to explain technical designs and trade-offs to both technical and non-technical partners.
Preferred Qualifications
• Databricks certification (Data Engineer Associate or Professional) or equivalent demonstrated depth.
• Unity Catalog experience: catalogs, schemas, volumes, external locations, storage credentials, permissions, and lineage.
• dbt on Databricks, or another transformation framework used alongside Spark.
• Python-based modeling frameworks over Delta Lake, and experience implementing Type 2 history, surrogate keys, and merge strategies in code.
• Experience with a declarative Python ingestion framework such as dlt (dltHub), Airbyte, Meltano, or Fivetran.
• Dagster experience specifically, including assets, asset checks, sensors, schedules, and branch deployments.
Originally posted on Himalayas
Apply for this role Opens jobs.lever.co — verified as the employer's own application page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 20 Sep 2026 Himalayas first sighting
Seen on 1 board over 0 days.