hirly
Likely filled

This role has closed. Infosys has taken the posting down.

hirly last saw it live on 30 September 2026. See similar open roles below, or browse all Data Engineer jobs in Bengaluru.

Infosys

Data Engineer - DaAI

Bangalore, India

This one has closed. See which open jobs fit you. Free.

Upload your resume and hirly scores it against open Data Engineer jobs in Bengaluru, then shows your best matches and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Country
IN
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting.

the posting

As a Data Engineer, you will help build the data foundation for our agentic AI platform. You will work with senior data architects, AI/ML engineers, and platform engineers to implement data ingestion, transformation, profiling, enrichment, validation, and preparation pipelines across structured and unstructured enterprise data sources.

This is a hands-on engineering role for someone who enjoys working with real-world enterprise data, building reliable pipelines, writing robust Python and SQL, and helping convert raw enterprise information into AI-ready data assets.

Responsibilities

Build and maintain data ingestion pipelines for structured enterprise systems such as ERP, CRM, billing, finance, HR, OSS/BSS, ServiceNow, Salesforce, SAP, Oracle, databases, and APIs.

Build pipelines for unstructured and semi-structured data sources such as documents, emails, logs, transcripts, PDFs, spreadsheets, and media metadata.

Develop ETL/ELT workflows using Python, SQL, PySpark, Apache Spark, Airflow, dbt, Dagster, cloud-native services, or equivalent technologies.

Support data profiling routines to identify missing values, duplicates, inconsistent formats, incomplete master data, schema changes, and conflicting records.

Implement data quality checks using frameworks such as Great Expectations, dbt tests, AWS Glue DataBrew, custom validation scripts, or equivalent tools.

Support data labelling, contextualization, harmonization, enrichment, and classification workflows required for AI agent configuration.

Prepare data outputs for downstream AI consumption, including embeddings, metadata, semantic tags, graph-ready datasets, and retrieval-ready document chunks.

Technical requirements

Working knowledge of data pipeline development using PySpark, Apache Spark, Airflow, dbt, Dagster, or equivalent technologies.

Experience working with structured data from databases, APIs, enterprise applications, data lakes, warehouses, or lakehouse platforms.

Exposure to cloud data platforms such as Databricks, Snowflake, BigQuery, Azure Data Lake, AWS S3, Google Cloud Storage, or equivalent platforms.

Understanding of data modelling, schema design, joins, keys, relationships, data validation, and data quality concepts.

Practical experience with data profiling, cleansing, transformation, and reconciliation.

Familiarity with Git, CI/CD basics, unit testing, and production-grade engineering practices.

Education

Bachelor of Engineering

Original posting on Infosys's site ↗

Browse similar roles