Back to all jobs
Cosette Network

Senior Data Engineer (Databricks / PySpark)

Bengaluru, Noida, Pune, Mumbai, Gurgaonremote9.0 - 12.0 years7 openings

Apply for this position

All fields marked * are required

Name & contact details are extracted from your resume automatically.

Job Description

Work Location: Remote (India)

  • Candidates should be flexible to work from an EXL office whenever business requires. While the role is primarily remote, it is not a permanent work-from-home position.

    Key Responsibilities

•   Design, build, and productionise batch and incremental ETL/ELT pipelines using Databricks (notebooks, Jobs/Workflows, DLT), PySpark, and Spark SQL across bronze/silver/gold Delta Lake layers.

•   Implement the medallion architecture with Delta Lake features — ACID merges/upserts (SCD1/SCD2), schema evolution/enforcement, time travel, OPTIMIZE/Z-ORDER, vacuuming — and manage data with Unity Catalog.

•   Orchestrate workloads via Databricks Workflows and/or Azure Data Factory; parameterise pipelines, implement restartability, checkpointing, and idempotent re-runs.

•   Tune Spark for performance and cost: partitioning strategy, join/broadcast optimisation, skew handling, caching, cluster sizing/autoscaling, Photon usage, and job-level cost monitoring.

•   Build data quality checks (expectations, reconciliation against source, row/measure-level controls) and implement failure alerting, logging, and lineage-friendly design.

•   Ingest from diverse sources — RDBMS (SQL Server/Oracle), files (CSV/Parquet/JSON/XML), APIs, and streaming/CDC feeds — into ADLS Gen2 landing zones.

•   Contribute to CI/CD for data: Git-based development, pull-request reviews, Azure DevOps/GitHub Actions pipelines, environment promotion, and infrastructure/config as code.

•   Collaborate with Data Modelers, BDAs, testers, and onshore leads in Agile ceremonies; convert mapping specifications into robust, reviewed code; mentor junior engineers.

•   Provide L3 production support for owned pipelines: triage incidents, perform root-cause analysis, and drive permanent fixes within SLAs.

Must-Have Skills & Experience

•   9–12 years overall; strong recent hands-on Databricks and PySpark delivery experience (3+ years Databricks preferred).

•   Expert-level Spark (DataFrame API, Spark SQL, window functions, UDF trade-offs) and advanced Python for data engineering.

•   Advanced SQL — complex transformations, performance tuning, analytical/window queries on large volumes.

•   Delta Lake internals and medallion/lakehouse architecture in production.

•   Azure data stack: ADLS Gen2, Azure Data Factory, Key Vault, and Databricks administration basics (clusters, pools, jobs, secrets).

•   Data warehousing concepts: star schemas, fact/dimension loading patterns, SCD handling, reconciliation.

•   Git-based SDLC with CI/CD exposure (Azure DevOps or GitHub).

•   Insurance or financial-services data experience — ideally P&C (policy/claims/premium/reinsurance structures).

Good-to-Have

•   Delta Live Tables, Unity Catalog governance, Databricks Asset Bundles.

•   Streaming (Structured Streaming, Auto Loader, Event Hubs/Kafka) and CDC tools.

•   dbt on Databricks; Airflow; Terraform.

•   Guidewire, Duck Creek, or London-market source systems; Snowflake or Synapse coexistence.

•   Databricks Data Engineer Associate/Professional certification.

Qualifications

•   Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

Professional & Communication Skills

•   Clear written and spoken English; comfortable presenting design decisions to onshore architects and client stakeholders.

•   Proven offshore-delivery discipline: crisp status reporting, proactive risk flagging, dependable overlap-hours availability.

•   Ownership mindset — drives issues to closure without follow-up.

Skills

Required

Databricks and PySpark; Advanced SQLmedallion/lakehouse; Data vault 2.0