EXL
Senior Databricks Engineer
Pune, Maharashtra, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Senior
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 23 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
- Description
- Databricks Platform Engineering ● Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments. ● Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage. ● Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance. ● Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager). ● Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management. Delta Lake & Lakehouse Architecture ● Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM). ● Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization. ● Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations. ● Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing. ● Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL) ● Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake. ● Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables. ● Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning. ● Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity. ● Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines. MLflow & AI/ML Workloads ● Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks. ● Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances. ● Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features. ● Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI). ● Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints. Cloud Integration & DevOps ● Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS). ● Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI. ● Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources. ● Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte). ● Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools. Governance, Security & Optimization ● Implement row-level security, column masking, and dynamic data views using Unity Catalog policies. ● Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations. ● Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement. ● Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog. ● Document architecture decisions, runbooks, and operational guides for Databricks workloads.
- Responsibilities
- Databricks Platform Engineering ● Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments. ● Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage. ● Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance. ● Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager). ● Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management. Delta Lake & Lakehouse Architecture ● Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM). ● Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization. ● Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations. ● Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing. ● Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL) ● Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake. ● Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables. ● Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning. ● Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity. ● Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines. MLflow & AI/ML Workloads ● Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks. ● Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances. ● Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features. ● Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI). ● Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints. Cloud Integration & DevOps ● Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS). ● Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI. ● Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources. ● Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte). ● Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools. Governance, Security & Optimization ● Implement row-level security, column masking, and dynamic data views using Unity Catalog policies. ● Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations. ● Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement. ● Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog. ● Document architecture decisions, runbooks, and operational guides for Databricks workloads.
- Qualifications
- Databricks Platform Engineering ● Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments. ● Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage. ● Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance. ● Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager). ● Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management. Delta Lake & Lakehouse Architecture ● Design and implement Delta Lake tables with
Similar jobs
- Senior Azure Databricks EngineerTata Consultancy Services · IndiaFirst seen 5d ago
- Senior Databricks EngineerDATABEAT · Hyderabad, IndiaFirst seen 9d ago
- Senior Databricks Engineer, Apache Spark and AWS - Vice PresidentCiti · Pune Maharashtra IndiaFirst seen 5d ago
- Azure Databricks EngineerCapgemini · Mumbai (ex Bombay), Pune, Bangalore, Chennai (ex Madras), INFirst seen 6d ago
- Lead Databricks Engineer (Senior Software Engineer II)The Nielsen Company · Mumbai, Maharashtra, IndiaFirst seen 7d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job