All fields marked * are required
Senior Data Engineer – PyTorch & ML Data Platforms
Location: Bengaluru
Work Mode: Work from Office – 5 Days a Week
Experience: 5+ Years
Employment Type: Full-Time
CTC: ₹22–35 LPA
Job Overview
We are looking for a highly skilled Senior Data Engineer with strong experience in building production-grade data pipelines and hands-on expertise in PyTorch and ML data platforms.
The ideal candidate will work on data ingestion, transformation, feature engineering, training data management, ML workflows, and infrastructure supporting machine learning and AI applications. This is a hands-on engineering role requiring strong coding, data engineering, and ML systems experience.
Key Responsibilities
Design, build, and maintain scalable batch and streaming data pipelines for model training and inference.
Build and optimize PyTorch training and inference workflows.
Develop custom PyTorch Datasets and DataLoaders.
Work with distributed training, mixed precision, checkpointing, and reproducible ML workflows.
Design and implement feature engineering pipelines and feature store solutions.
Ensure training and serving data consistency and prevent data leakage.
Curate, version, validate, and maintain training datasets.
Implement data quality checks, data validation, and drift detection.
Deploy and serve ML models using TorchServe, ONNX Runtime, Triton, or similar technologies.
Build data infrastructure for AI and GenAI applications, including embedding pipelines, vector stores, chunking strategies, and evaluation datasets.
Implement pipeline observability, data lineage, data contracts, monitoring, alerting, and model performance tracking.
Work with engineering teams to design and deliver end-to-end data and AI solutions.
Review code and contribute to engineering best practices and technical standards.
Mentor and support other engineers on data and ML engineering practices.
Mandatory Skills & Experience
5+ years of professional experience in Data Engineering, ML Engineering, or a related field.
Strong production-level Python programming experience.
Ability to write clean, tested, maintainable, and scalable Python code.
Strong hands-on experience with PyTorch.
Experience building and training ML models using PyTorch.
Experience developing custom Datasets and DataLoaders.
Experience debugging training workflows and taking at least one ML model into production.
Strong SQL skills.
Solid understanding of data modelling, including dimensional modelling.
Experience with partitioning, indexing, and query optimization.
Production experience with distributed data processing technologies such as:
Apache Spark
Ray
Dask
or equivalent
Experience with workflow orchestration tools such as:
Apache Airflow
Dagster
Prefect
or similar
Experience handling backfills, idempotency, failure recovery, and production pipeline reliability.
Experience with at least one cloud platform:
AWS
GCP
Azure
Experience with cloud data services and platforms such as S3/GCS, Glue, Dataproc, EMR, Redshift, BigQuery, Snowflake, Databricks, or equivalent.
Experience with Docker and containerization.
Strong understanding of Git-based development workflows.
Experience implementing or working with CI/CD pipelines.
Strong communication and documentation skills.
Preferred Skills
Experience with LLM and GenAI systems.
Knowledge of embeddings, vector databases, and RAG pipelines.
Experience with vector databases such as:
pgvector
Pinecone
Weaviate
Qdrant
Experience with LLM evaluation frameworks or evaluation datasets.
Experience with Hugging Face Transformers.
Experience with fine-tuning techniques such as LoRA or QLoRA.
Experience with MLOps tools such as MLflow, Weights & Biases, Kubeflow, SageMaker, or Vertex AI.
Experience with Kubernetes.
Experience with Infrastructure as Code using Terraform or similar tools.
Experience with streaming technologies such as Kafka, Kinesis, Flink, or Spark Structured Streaming.
Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
Experience with dbt and modern analytics engineering practices.
Knowledge of GPU performance optimization, quantization, or ML inference cost optimization.
Ideal Candidate
The ideal candidate should have a strong combination of Data Engineering and Machine Learning Engineering expertise. They should be comfortable building production-grade data platforms, working hands-on with PyTorch, optimizing ML workflows, and solving complex data and ML infrastructure challenges.
Location: Bengaluru
Work Mode: Work from Office – 5 Days a Week
Experience: 5+ Years
CTC: ₹22–35 LPA
Required
Preferred
All fields marked * are required
Senior Data Engineer – PyTorch & ML Data Platforms
Location: Bengaluru
Work Mode: Work from Office – 5 Days a Week
Experience: 5+ Years
Employment Type: Full-Time
CTC: ₹22–35 LPA
Job Overview
We are looking for a highly skilled Senior Data Engineer with strong experience in building production-grade data pipelines and hands-on expertise in PyTorch and ML data platforms.
The ideal candidate will work on data ingestion, transformation, feature engineering, training data management, ML workflows, and infrastructure supporting machine learning and AI applications. This is a hands-on engineering role requiring strong coding, data engineering, and ML systems experience.
Key Responsibilities
Design, build, and maintain scalable batch and streaming data pipelines for model training and inference.
Build and optimize PyTorch training and inference workflows.
Develop custom PyTorch Datasets and DataLoaders.
Work with distributed training, mixed precision, checkpointing, and reproducible ML workflows.
Design and implement feature engineering pipelines and feature store solutions.
Ensure training and serving data consistency and prevent data leakage.
Curate, version, validate, and maintain training datasets.
Implement data quality checks, data validation, and drift detection.
Deploy and serve ML models using TorchServe, ONNX Runtime, Triton, or similar technologies.
Build data infrastructure for AI and GenAI applications, including embedding pipelines, vector stores, chunking strategies, and evaluation datasets.
Implement pipeline observability, data lineage, data contracts, monitoring, alerting, and model performance tracking.
Work with engineering teams to design and deliver end-to-end data and AI solutions.
Review code and contribute to engineering best practices and technical standards.
Mentor and support other engineers on data and ML engineering practices.
Mandatory Skills & Experience
5+ years of professional experience in Data Engineering, ML Engineering, or a related field.
Strong production-level Python programming experience.
Ability to write clean, tested, maintainable, and scalable Python code.
Strong hands-on experience with PyTorch.
Experience building and training ML models using PyTorch.
Experience developing custom Datasets and DataLoaders.
Experience debugging training workflows and taking at least one ML model into production.
Strong SQL skills.
Solid understanding of data modelling, including dimensional modelling.
Experience with partitioning, indexing, and query optimization.
Production experience with distributed data processing technologies such as:
Apache Spark
Ray
Dask
or equivalent
Experience with workflow orchestration tools such as:
Apache Airflow
Dagster
Prefect
or similar
Experience handling backfills, idempotency, failure recovery, and production pipeline reliability.
Experience with at least one cloud platform:
AWS
GCP
Azure
Experience with cloud data services and platforms such as S3/GCS, Glue, Dataproc, EMR, Redshift, BigQuery, Snowflake, Databricks, or equivalent.
Experience with Docker and containerization.
Strong understanding of Git-based development workflows.
Experience implementing or working with CI/CD pipelines.
Strong communication and documentation skills.
Preferred Skills
Experience with LLM and GenAI systems.
Knowledge of embeddings, vector databases, and RAG pipelines.
Experience with vector databases such as:
pgvector
Pinecone
Weaviate
Qdrant
Experience with LLM evaluation frameworks or evaluation datasets.
Experience with Hugging Face Transformers.
Experience with fine-tuning techniques such as LoRA or QLoRA.
Experience with MLOps tools such as MLflow, Weights & Biases, Kubeflow, SageMaker, or Vertex AI.
Experience with Kubernetes.
Experience with Infrastructure as Code using Terraform or similar tools.
Experience with streaming technologies such as Kafka, Kinesis, Flink, or Spark Structured Streaming.
Experience with Delta Lake, Apache Iceberg, or Apache Hudi.
Experience with dbt and modern analytics engineering practices.
Knowledge of GPU performance optimization, quantization, or ML inference cost optimization.
Ideal Candidate
The ideal candidate should have a strong combination of Data Engineering and Machine Learning Engineering expertise. They should be comfortable building production-grade data platforms, working hands-on with PyTorch, optimizing ML workflows, and solving complex data and ML infrastructure challenges.
Location: Bengaluru
Work Mode: Work from Office – 5 Days a Week
Experience: 5+ Years
CTC: ₹22–35 LPA
Required
Preferred