Back to all jobs
Cosette Network

Data Engineering Lead (Databricks / PySpark)

Bengaluru, Noida, Pune, Gurgaon, Mumbaihybrid12.0 - 15.0 years2 openings

Apply for this position

All fields marked * are required

Name & contact details are extracted from your resume automatically.

Job Description

Key Responsibilities

•   Own end-to-end technical design of lakehouse pipelines and frameworks: ingestion patterns, medallion layer standards, metadata-driven frameworks, reusable libraries, and naming/coding conventions.

•   Lead and grow a squad of data engineers (typically 5–10): task breakdown and estimation, sprint planning with the Scrum Master/EM, code reviews, pairing, and performance feedback.

•   Translate business and architecture requirements into technical designs and delivery plans; present and defend design decisions and trade-offs to QBE architects and platform owners.

•   Set and enforce engineering quality gates: PR review standards, unit/integration test coverage for pipelines, data quality SLAs, CI/CD promotion criteria, and documentation.

•   Own non-functional outcomes — performance, cost optimisation (cluster/DBU governance), security (Unity Catalog permissions, PII handling), reliability, and observability of the platform.

•   Manage technical risk: dependency tracking, proof-of-concepts for new patterns (streaming, DLT, Asset Bundles), remediation plans for tech debt, and production incident command for critical issues.

•   Coordinate across workstreams — Data Modelers, BDAs, testing, governance — to keep specs, models, code, and test coverage in lockstep; run design authority sessions.

•   For the Onshore Lead: front-door for client stakeholders — requirement workshops, steering-committee technical inputs, escalation handling, and onshore–offshore handshake quality.

•   For the Offshore Lead: run offshore ceremonies, own sprint delivery and status, ensure overlap-hours coverage, and manage the offshore–onshore work packaging.

Must-Have Skills & Experience

•   12+ years in data engineering with 5+ years leading teams delivering on Spark platforms; deep, current hands-on Databricks + PySpark (this is a coding lead, not a pure people-manager role).

•   Proven architecture-level command of the lakehouse/medallion pattern, Delta Lake, Unity Catalog, and Azure data services (ADLS Gen2, ADF, Key Vault, networking basics for Databricks).

•   Track record designing metadata/config-driven ingestion and transformation frameworks used by multiple teams.

•   Advanced Spark performance engineering and cost governance at platform scale.

•   Strong SDLC leadership: Git branching strategy, CI/CD (Azure DevOps/GitHub Actions), environment strategy, release management.

•   Insurance domain experience — P&C strongly preferred (policy, claims, premium, reinsurance, actuarial data flows).

•   Experience running distributed onshore/offshore delivery models with measurable quality outcomes.

Good-to-Have

•   DLT, Databricks Asset Bundles, Terraform for Databricks/Azure.

•   Streaming architectures (Auto Loader, Structured Streaming, Kafka/Event Hubs).

•   Migration experience (on-prem/legacy ETL → Databricks); Informatica/DataStage/SSIS conversion.

•   Databricks Professional certification; Azure Solutions Architect (AZ-305) or Data Engineer (DP-203/DP-700).

Qualifications

•   Bachelor’s/Master’s in Computer Science, Engineering, or related field.

Professional & Communication Skills

•   Executive-ready communication — can hold the room with client architects and translate for business stakeholders.

•   Decision-making with documented trade-offs; comfortable saying no with alternatives.

•   Mentoring culture-builder; raises the squad’s bar rather than becoming the bottleneck.

Skills

Required

Databricks and PySpark; Advanced SQLmedallion/lakehouse; Data vault 2.0
Data Engineering Lead (Databricks / PySpark) at Cosette Network | Talynce Jobs