hirly

Thermofisher

Data Engineer (ETL, Python, SQL)

Taguig City, Philippines

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Thermofisher first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Country
PH
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Work Schedule

Standard (Mon-Fri)

Environmental Conditions

Office

Job Description

Summarized Purpose:

We are offering an opportunity for a Mid-Level Data Engineer to design, build, test, tune, and support production data pipelines using PySpark, Python, advanced SQL, AWS data services, secure data handling practices, and AI-assisted data engineering capabilities.

Education/Experience:

Bachelor's degree or equivalent in Computer Science, Information Technology, Data Engineering, or related field

3-5 years of experience in data engineering, ETL development, SQL, AWS data platforms, or production data pipeline support

Major Job Responsibilities:

Develop, test, tune, and maintain ETL and data pipelines using PySpark, Python, SQL, and AWS services

Support ingestion and transformation of flat files, relational databases, APIs, data warehouses, and enterprise data sources

Collaborate with business analysts, data architects, QA, DevOps, and senior engineers to implement source-to-target mappings and data solutions

Implement CDC, incremental load design, idempotent pipeline processing, and data reconciliation patterns for reliable data movement

Maintain technical documentation, mapping specifications, data catalog updates, runbooks, automated tests, and release support materials

Knowledge, Skills, and Abilities:

Hands-on experience with PySpark, Python, advanced SQL, ETL best practices, data modeling, and large-scale data processing

Deep knowledge of Redshift performance tuning including distribution keys, sort keys, compression encoding, Spectrum, materialized views, WLM, vacuum, and analyze

Strong knowledge of Athena optimization including partition pruning, file formats, compression, schema evolution, and cost-efficient query design

Strong understanding of DynamoDB data modeling, access-pattern-based design, capacity planning, GSIs/LSIs, TTL, Streams, and performance tuning

Exposure to secure PHI/PII handling including encryption, access controls, auditability, retention, masking, and de-identification where applicable

Strong analytical, troubleshooting, documentation, communication, and cross-functional collaboration skills

Must Have Skills:

PySpark, Python, advanced SQL, ETL development, and data pipeline implementation experience

AWS data services experience including S3, Glue, Lambda, Step Functions, ECS, DynamoDB, Redshift, PostgreSQL, SQL Server, and Athena integration

Flat-file ingestion, source-to-target mapping, transformation logic, CDC, incremental loads, idempotent processing, reconciliation, and data quality checks

CI/CD, GitHub workflows, automated testing, and release management for data pipelines and database changes

Problem-solving, production support, debugging, documentation, and Agile delivery skills

Good to Have Skills:

Exposure to AI-assisted mapping automation and use of LLMs for data cleaning, data quality checks, transformation logic, or documentation

Familiarity with RAG patterns, embeddings, vector databases, semantic search, or AI-enabled data discovery solutions

Understanding of healthcare data standards such as HL7, FHIR, CCD, claims data, EMR extracts, clinical trial data, and patient de-identification

Familiarity with infrastructure as code such as Terraform or CloudFormation, plus Databricks, Snowflake, streaming, observability, or DevOps practices

Working Hours:

Philippines: 08:00 PM to 05:00 AM PHT

Original posting on Thermofisher's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job