Data Engineering Lead (Databricks / PySpark)
Apply for this position
All fields marked * are required
Key Responsibilities
• Own end-to-end technical design of lakehouse pipelines and frameworks: ingestion patterns, medallion layer standards, metadata-driven frameworks, reusable libraries, and naming/coding conventions.
• Lead and grow a squad of data engineers (typically 5–10): task breakdown and estimation, sprint planning with the Scrum Master/EM, code reviews, pairing, and performance feedback.
• Translate business and architecture requirements into technical designs and delivery plans; present and defend design decisions and trade-offs to QBE architects and platform owners.
• Set and enforce engineering quality gates: PR review standards, unit/integration test coverage for pipelines, data quality SLAs, CI/CD promotion criteria, and documentation.
• Own non-functional outcomes — performance, cost optimisation (cluster/DBU governance), security (Unity Catalog permissions, PII handling), reliability, and observability of the platform.
• Manage technical risk: dependency tracking, proof-of-concepts for new patterns (streaming, DLT, Asset Bundles), remediation plans for tech debt, and production incident command for critical issues.
• Coordinate across workstreams — Data Modelers, BDAs, testing, governance — to keep specs, models, code, and test coverage in lockstep; run design authority sessions.
• For the Onshore Lead: front-door for client stakeholders — requirement workshops, steering-committee technical inputs, escalation handling, and onshore–offshore handshake quality.
• For the Offshore Lead: run offshore ceremonies, own sprint delivery and status, ensure overlap-hours coverage, and manage the offshore–onshore work packaging.
Must-Have Skills & Experience
• 12+ years in data engineering with 5+ years leading teams delivering on Spark platforms; deep, current hands-on Databricks + PySpark (this is a coding lead, not a pure people-manager role).
• Proven architecture-level command of the lakehouse/medallion pattern, Delta Lake, Unity Catalog, and Azure data services (ADLS Gen2, ADF, Key Vault, networking basics for Databricks).
• Track record designing metadata/config-driven ingestion and transformation frameworks used by multiple teams.
• Advanced Spark performance engineering and cost governance at platform scale.
• Strong SDLC leadership: Git branching strategy, CI/CD (Azure DevOps/GitHub Actions), environment strategy, release management.
• Insurance domain experience — P&C strongly preferred (policy, claims, premium, reinsurance, actuarial data flows).
• Experience running distributed onshore/offshore delivery models with measurable quality outcomes.
Good-to-Have
• DLT, Databricks Asset Bundles, Terraform for Databricks/Azure.
• Streaming architectures (Auto Loader, Structured Streaming, Kafka/Event Hubs).
• Migration experience (on-prem/legacy ETL → Databricks); Informatica/DataStage/SSIS conversion.
• Databricks Professional certification; Azure Solutions Architect (AZ-305) or Data Engineer (DP-203/DP-700).
Qualifications
• Bachelor’s/Master’s in Computer Science, Engineering, or related field.
Professional & Communication Skills
• Executive-ready communication — can hold the room with client architects and translate for business stakeholders.
• Decision-making with documented trade-offs; comfortable saying no with alternatives.
• Mentoring culture-builder; raises the squad’s bar rather than becoming the bottleneck.
Required