All fields marked * are required
JD: Spark + Scala+ Python +Github Copilot
Location : Only Bangalore
• Design, develop, deploy, and support enterprise-scale Data Lake and distributed processing solutions using Spark, PySpark, Scala, Python, Hive, Spark SQL, and Airflow.
• Build modular, testable, reusable ETL pipelines across HDFS, Hive, Parquet, and Cloudera CDP on-premises environments.
• Apply partitioning, bucketing, schema evolution, data quality, reconciliation, error handling, monitoring, recovery, and security practices.
• Optimize workloads through expertise in joins, shuffles, serialization, caching, partition sizing, file formats, query plans, and resource utilization.
• Implement data modeling, CI/CD, Git, GitLab, Jenkins, Maven/SBT, release, and operational practices.
• Troubleshoot logs, failures, performance bottlenecks, and production incidents through effective root-cause analysis.
• Demonstrate strong conceptual depth, coding discipline, analytical thinking, and attention to detail.
• Communicate clearly, challenge assumptions, learn rapidly, and convert ambiguity into pragmatic outcomes.
• Own customer-facing delivery with a Forward Deployed Engineer mindset: adaptable, hands-on, action-oriented, and accountable.
• Use Claude, Cursor, and Copilot responsibly for coding, testing, automation, and documentation.
Required
All fields marked * are required
JD: Spark + Scala+ Python +Github Copilot
Location : Only Bangalore
• Design, develop, deploy, and support enterprise-scale Data Lake and distributed processing solutions using Spark, PySpark, Scala, Python, Hive, Spark SQL, and Airflow.
• Build modular, testable, reusable ETL pipelines across HDFS, Hive, Parquet, and Cloudera CDP on-premises environments.
• Apply partitioning, bucketing, schema evolution, data quality, reconciliation, error handling, monitoring, recovery, and security practices.
• Optimize workloads through expertise in joins, shuffles, serialization, caching, partition sizing, file formats, query plans, and resource utilization.
• Implement data modeling, CI/CD, Git, GitLab, Jenkins, Maven/SBT, release, and operational practices.
• Troubleshoot logs, failures, performance bottlenecks, and production incidents through effective root-cause analysis.
• Demonstrate strong conceptual depth, coding discipline, analytical thinking, and attention to detail.
• Communicate clearly, challenge assumptions, learn rapidly, and convert ambiguity into pragmatic outcomes.
• Own customer-facing delivery with a Forward Deployed Engineer mindset: adaptable, hands-on, action-oriented, and accountable.
• Use Claude, Cursor, and Copilot responsibly for coding, testing, automation, and documentation.
Required