Infosys
Pyspark
Bangalore, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About the job:
Join a collaborative data engineering team where your work directly powers reliable analytics and smarter business decisions. In this role, you’ll build and optimize scalable data processing pipelines using PySpark and Spark, working closely with engineers, analysts, and stakeholders to turn raw data into trusted, high-quality datasets. You’ll be encouraged to take ownership, suggest improvements, and contribute to a culture that values clean engineering, performance, and continuous learning. If you enjoy solving data challenges, tuning distributed jobs, and delivering dependable solutions in a fast-moving environment, this is a great opportunity to grow your impact while working with modern big data technologies.
Responsibilities
Key Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for batch and/or incremental processing.
Build and optimize Apache Spark jobs with focus on performance, partitioning strategy, caching, and efficient transformations/actions.
Perform data cleansing, validation, and reconciliation to ensure accuracy, completeness, and consistency of datasets.
Collaborate with cross-functional teams to understand requirements and translate them into robust data processing solutions.
Troubleshoot pipeline failures, analyze logs, identify bottlenecks, and implement fixes to improve reliability and throughput.
Write clean, maintainable code with reusable components and clear documentation for pipelines and data flows.
Support deployment and operationalization of Spark workloads, including monitoring and basic production support activities.
Contribute to code reviews and follow engineering best practices to improve quality and maintainability.
Minimum Qualifications:
Education: BTECH, MTECH, MCA, MSC.
2–3 years of experience in data engineering or large-scale data processing roles.
Strong hands-on experience with PySpark for building data pipelines and transformations.
Working knowledge of Apache Spark concepts such as RDD/DataFrame, joins, shuffles, and performance considerations.
Ability to debug Spark applications and resolve data/job issues effectively.
Technical requirements
Good to have skills:
SQL, Hadoop, Hive, Kafka, Airflow
Additional responsibilities
Preferred Qualifications:
Experience optimizing Spark workloads (tuning partitions, managing skew, memory/executor settings) for performance and cost efficiency.
Exposure to building end-to-end data pipelines with strong data quality checks and automated validations.
Familiarity with distributed processing patterns and designing reusable PySpark modules for scalable development.
Experience collaborating in agile teams, participating in code reviews, and improving engineering standards for data pipelines.
Education
MCA,MSc,MTech,Bachelor of Engineering,BTech
Similar jobs
- Spark or Pyspark & Scala & Snowflake/ Databrick DeveloperIqvia · Bangalore, IndiaFirst seen 2d ago
- ADF (Azure Data Factory)DatabricksPysparkInfosys Limited · Bangalore, Karnataka, IndiaFirst seen 2d ago
- Engineer (A2 DES, Databricks, PySpark, Python)KPMG · Bangalore, Karnataka, IndiaFirst seen 5d ago
- Azure Databricks, PysparkCognizant · Chennai, Tamil Nadu, IndiaFirst seen yesterday
- PysparkScalaSQLHiveAirflow (Data analyst)LTM · Hyderabad, Telangana, IndiaFirst seen 2d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job