Infosys
Pyspark
Bangalore, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Join a fast-paced, collaborative team where data powers smarter decisions and better customer experiences. In this role, you’ll work hands-on with large-scale datasets to build reliable, high-performing data processing solutions using PySpark and Spark. You’ll partner closely with engineers, analysts, and stakeholders to understand business needs, translate them into scalable pipelines, and continuously improve data quality and performance. If you enjoy solving complex data challenges, optimizing distributed workloads, and taking ownership from design to delivery, this is a great opportunity to grow your impact. You’ll be encouraged to share ideas, learn from peers, and contribute to a culture that values clarity, craftsmanship, and continuous improvement.
Responsibilities
Key Responsibilities
Design, develop, and maintain scalable batch data pipelines using PySpark and Apache Spark for large datasets.
Perform data ingestion, transformation, and enrichment while ensuring accuracy, completeness, and consistency of outputs.
Optimize Spark jobs for performance (partitioning, caching, joins, shuffles) and improve runtime efficiency and resource utilization.
Implement robust error handling, logging, and monitoring to ensure reliable pipeline execution and faster issue resolution.
Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver well-documented solutions.
Conduct code reviews, follow engineering best practices, and contribute to reusable components and standards.
Troubleshoot production issues, perform root-cause analysis, and drive corrective and preventive actions.
Technical requirements
Primary skills:Technology->Big Data - Data Processing->PySpark
Additional responsibilities
Minimum Qualifications:
Bachelor’s degree (or equivalent) in Engineering/Computer Science/IT or related field.
3–5 years of experience in data engineering or big data development roles.
Strong hands-on experience with PySpark and Apache Spark for building data processing workflows.
Solid understanding of distributed data processing concepts and performance tuning fundamentals.
Ability to translate business requirements into technical implementations and deliver within timelines.
Preferred Qualifications:
Experience building end-to-end Spark applications including job orchestration, dependency management, and production support readiness.
Strong data transformation skills with a focus on data quality checks, reconciliation, and pipeline reliability.
Exposure to designing modular, reusable Spark components and maintaining clean, maintainable codebases.
Familiarity with structured and semi-structured data formats and efficient processing patterns in Spark.
Proven ability to collaborate effectively across teams, communicate clearly, and contribute to continuous improvement initiatives.
Good to have skills:
Spark SQL, Delta Lake, Databricks, Airflow, Hadoop
Education
MCA,MSc,MTech,Bachelor of Engineering,BTech
Similar jobs
- ADF (Azure Data Factory)DatabricksPysparkInfosys Limited · Bangalore, Karnataka, IndiaFirst seen 3d ago
- Engineer (A2 DES, Databricks, PySpark, Python)KPMG · Bangalore, Karnataka, IndiaFirst seen 6d ago
- Azure Databricks, PysparkCognizant · Chennai, Tamil Nadu, IndiaFirst seen 2d ago
- PysparkScalaSQLHiveAirflow (Data analyst)LTM · Hyderabad, Telangana, IndiaFirst seen 3d ago
- Databricks, Pyspark, Kafka/NiFiLTM · Hyderabad, Telangana, IndiaFirst seen 3d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job