Resilientco
Sr Platform/ Infrastructure Engineer
Argentina · Brazil · Chile · Paraguay · Uruguay · Colombia · Mexico · Costa Rica
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Senior
- Countries
- AR, BR, CL, PY, UY, CO, MX, CR
- Work mode
- Remote-friendly
- First seen by hirly
- 11 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.
You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.
Responsibilities
Design, deploy, and maintain production Kubernetes clusters and related services.
Build and maintain automation and tooling using Python to support platform operations.
Integrate and operate Prometheus for monitoring, alerting, and observability.
Deploy and manage Ceph storage solutions for distributed workloads.
Support platform modernization initiatives and migrate services to cloud-native patterns.
Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
Document platform designs, runbooks, and operational procedures.
Participate in on-call rotations and incident response to maintain platform availability.
Requirements
5+ years of experience in platform, infrastructure, or site reliability engineering roles.
Proven experience deploying and operating Kubernetes in production.
Strong Python skills for automation, tooling, and operational scripts.
Experience implementing and operating Prometheus-based monitoring and alerting.
Hands-on experience with Ceph or similar distributed storage systems.
Cloud experience with AWS and Azure (designing, deploying, and operating services).
Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
Experience collaborating across teams to deliver platform improvements and migrations.
Nice to Have
Experience with OpenSearch.
Proficiency with Bash scripting.
Familiarity with Java-based services.
Experience with Fluent Bit for log collection.
Experience working with PostgreSQL.
Similar jobs
- Senior Data Architect - Databricks PlatformCaylent · ARGENTINAFirst seen 23d agoremote
- Senior Backend Engineer (Backend Platform)Feverup · ArgentinaFirst seen 26d ago
- Solutions Architect - Platform, Cloud Infrastructure and Security SunnyData · ArgentinaFirst seen 2d agoremote
- Senior Power Platform Developer - EY Global Delivery ServicesEY · CABA, B, ARFirst seen 3d ago
- Senior Consultant - OneStream Platform - EY GDSEY · CABA, B, ARFirst seen 3d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job