Senior Data Engineer (Databricks / PySpark)
Apply for this position
All fields marked * are required
Work Location: Remote (India)
Candidates should be flexible to work from an EXL office whenever business requires. While the role is primarily remote, it is not a permanent work-from-home position.
Key Responsibilities
• Design, build, and productionise batch and incremental ETL/ELT pipelines using Databricks (notebooks, Jobs/Workflows, DLT), PySpark, and Spark SQL across bronze/silver/gold Delta Lake layers.
• Implement the medallion architecture with Delta Lake features — ACID merges/upserts (SCD1/SCD2), schema evolution/enforcement, time travel, OPTIMIZE/Z-ORDER, vacuuming — and manage data with Unity Catalog.
• Orchestrate workloads via Databricks Workflows and/or Azure Data Factory; parameterise pipelines, implement restartability, checkpointing, and idempotent re-runs.
• Tune Spark for performance and cost: partitioning strategy, join/broadcast optimisation, skew handling, caching, cluster sizing/autoscaling, Photon usage, and job-level cost monitoring.
• Build data quality checks (expectations, reconciliation against source, row/measure-level controls) and implement failure alerting, logging, and lineage-friendly design.
• Ingest from diverse sources — RDBMS (SQL Server/Oracle), files (CSV/Parquet/JSON/XML), APIs, and streaming/CDC feeds — into ADLS Gen2 landing zones.
• Contribute to CI/CD for data: Git-based development, pull-request reviews, Azure DevOps/GitHub Actions pipelines, environment promotion, and infrastructure/config as code.
• Collaborate with Data Modelers, BDAs, testers, and onshore leads in Agile ceremonies; convert mapping specifications into robust, reviewed code; mentor junior engineers.
• Provide L3 production support for owned pipelines: triage incidents, perform root-cause analysis, and drive permanent fixes within SLAs.
Must-Have Skills & Experience
• 9–12 years overall; strong recent hands-on Databricks and PySpark delivery experience (3+ years Databricks preferred).
• Expert-level Spark (DataFrame API, Spark SQL, window functions, UDF trade-offs) and advanced Python for data engineering.
• Advanced SQL — complex transformations, performance tuning, analytical/window queries on large volumes.
• Delta Lake internals and medallion/lakehouse architecture in production.
• Azure data stack: ADLS Gen2, Azure Data Factory, Key Vault, and Databricks administration basics (clusters, pools, jobs, secrets).
• Data warehousing concepts: star schemas, fact/dimension loading patterns, SCD handling, reconciliation.
• Git-based SDLC with CI/CD exposure (Azure DevOps or GitHub).
• Insurance or financial-services data experience — ideally P&C (policy/claims/premium/reinsurance structures).
Good-to-Have
• Delta Live Tables, Unity Catalog governance, Databricks Asset Bundles.
• Streaming (Structured Streaming, Auto Loader, Event Hubs/Kafka) and CDC tools.
• dbt on Databricks; Airflow; Terraform.
• Guidewire, Duck Creek, or London-market source systems; Snowflake or Synapse coexistence.
• Databricks Data Engineer Associate/Professional certification.
Qualifications
• Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
Professional & Communication Skills
• Clear written and spoken English; comfortable presenting design decisions to onshore architects and client stakeholders.
• Proven offshore-delivery discipline: crisp status reporting, proactive risk flagging, dependable overlap-hours availability.
• Ownership mindset — drives issues to closure without follow-up.
Required