hirly

Infosys

Databricks, Pyspark

Chennai, India

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Infosys first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
IN
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Join a high-impact data engineering team where you’ll shape modern analytics platforms and help turn raw data into trusted, actionable insights. In this role, you’ll lead the design and delivery of scalable data pipelines on Databricks using PySpark, partnering closely with analysts, data scientists, and platform teams to enable faster decision-making across the business. You’ll bring strong engineering discipline—clean code, performance tuning, and reliable operations—while guiding best practices and mentoring teammates. If you enjoy solving complex data challenges, optimizing distributed workloads, and building systems that are resilient, secure, and easy to evolve, this is a great opportunity to drive meaningful outcomes in a collaborative, growth-focused environment.

Responsibilities

Key Responsibilities:

Lead the development of end-to-end data pipelines on Databricks using PySpark for batch and incremental processing.

Design scalable data models and curated datasets to support analytics and downstream consumption.

Write and optimize advanced SQL for transformations, validations, and performance-critical queries.

Implement robust data quality checks, reconciliation logic, and monitoring to ensure trusted datasets.

Tune Spark jobs for performance and cost efficiency (partitioning, caching, file formats, cluster sizing).

Establish coding standards, reusable frameworks, and review practices to improve maintainability.

Collaborate with stakeholders to translate requirements into technical designs and delivery plans.

Troubleshoot production issues, perform root-cause analysis, and drive preventive improvements.

Mentor team members and provide technical guidance across design, implementation, and optimization.

Minimum Qualifications:

BTECH, MTECH, MCA, or MSC in Computer Science, Information Technology, or a related field.

6–8 years of overall experience in data engineering or large-scale data processing roles.

Strong hands-on experience with PySpark for distributed data processing and transformation logic.

Strong hands-on experience with Databricks for building, running, and managing data workloads.

Proficiency in Advanced SQL including complex joins, window functions, and query optimization.

Experience building reliable pipelines with strong focus on data quality, performance, and stability

Technical requirements

Technology->Analytics - Solutions->SQL Server - Analytics

Technology->Big Data - Data Processing->PySpark

Technology->Data Engineering->Databricks

Additional responsibilities

Preferred Qualifications:

Experience designing lakehouse-style architectures and organizing curated layers for analytics readiness.

Strong experience with Spark optimization techniques and handling large-scale datasets efficiently.

Ability to build reusable PySpark utilities/frameworks for ingestion, transformation, and validation patterns.

Experience with orchestration and scheduling approaches for dependable pipeline execution and recovery.

Proven track record of technical leadership: mentoring, conducting reviews, and driving engineering best practices.

Education

MCA,MSc,MTech,Bachelor of Engineering,BTech

Original posting on Infosys's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Databricks, Pyspark – Infosys · Chennai | hirly.me