hirly

Infosys

PySpark Developer

Bangalore, India

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Country
IN
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

Responsibilities

Key Responsibilities

Develop and maintain data pipelines using PySpark

Process and analyze large-scale datasets in distributed environments

Design and implement ETL/ELT workflows

Optimize Spark jobs for performance and scalability

Work with data stored in HDFS, Hive, or cloud storage (S3, ADLS)

Collaborate with data engineers, analysts, and business teams

Ensure data quality, integrity, and governance

Debug and troubleshoot data processing issues

Automate workflows using scheduling tools (Airflow, Oozie, etc.)

Write clean, scalable, and efficient code

Required Skills & Qualifications

Technical Skills

Strong proficiency in Python and PySpark

Good experience with Apache Spark (RDDs, DataFrames, Spark SQL)

Knowledge of Hadoop ecosystem (HDFS, Hive)

Experience in ETL pipeline development

Familiarity with SQL and database concepts

Experience with data formats (Parquet, ORC, JSON, CSV)

Basic understanding of distributed computing concepts

Exposure to version control tools (Git)

Technical requirements

Technology->Big Data - Data Processing->PySpark

Education

MCA,MSc,MTech,Bachelor of Engineering,BCA,BSc,BTech

Original posting on Infosys's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job