Day to day:
The Data Engineer works as a hands-on individual contributor, building and maintaining production data pipelines using Python, PySpark, Spark SQL, and Databricks. The role focuses approximately 60% on developing new pipelines and onboarding data from various enterprise sources and 40% on supporting and maintaining existing Databricks pipelines and environments.
The Data Engineer develops scalable, reliable data solutions using Databricks Catalog, Lakeflow Declarative Pipelines/Delta Live Tables, Databricks SQL, and Delta Lake, with a strong understanding of MERGE semantics, Change Data Feed/CDC, schema evolution, dimensional modeling, SCD Type 1 and Type 2, and idempotent incremental processing. They deploy Python scripts and Asset Bundles through Jenkins and Git-based CI/CD processes and follow test-driven development practices to maintain reliable, well-tested production code.
The Data Engineer works with Azure services including ADLS Gen2, Azure Data Factory, and Key Vault, troubleshoots production issues, and performs ongoing environment and pipeline support. They collaborate closely with technical and executive stakeholders, communicate clearly, and independently drive engineering solutions from development through deployment. Experience with Claude and familiarity with agentic frameworks are valuable for supporting emerging AI-driven engineering initiatives.