Bloomreach
Senior Site Reliability Engineer for Fuse Team
Czechia
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Stated salary
- CZK 1,260,000 – CZK 1,572,000 per year
- Country
- CZ
- Work mode
- Remote-friendly
- First seen by hirly
- 10 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Bloomreach is building the world’s premier agentic platform for personalization .We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey.
We're taking autonomous search mainstream, making product discovery more intuitive and conversational for customers, and more profitable for businesses.
We’re making conversational shopping a reality, connecting every shopper with tailored guidance and product expertise — available on demand, at every touchpoint in their journey.
We're designing the future of autonomous marketing , taking the work out of workflows, and reclaiming the creative, strategic, and customer-first work marketers were always meant to do.
And we're building all of that on the intelligence of a single AI engine — Loomi — so that personalization isn't only autonomous…it's also consistent.From retail to financial services, hospitality to gaming, businesses use Bloomreach to drive higher growth and lasting loyalty. We power personalization for more than 1,400 global brands, including American Eagle, Sonepar, and Pandora.
Become a Senior SRE for Bloomreach!
Join the Fuse team — the team responsible for the item data management capabilities that connect Bloomreach Data Hub with Marketing, Search, Recommendations, and emerging Loomi agent use cases.
Fuse owns and evolves the systems behind Data Hub item collections : ingesting items data, transforming and validating it, managing data schemas and lifecycle, and distributing data reliably to downstream Bloomreach products. Item collections provide a unified source of data that can be used across all Bloomreach products.
Our current areas of focus include:
Unified items data pipelines : processing records into structured items and keeping data synchronized with Marketing and Search destinations.
Catalog APIs and lifecycle management : customer-facing and internal APIs, catalog creation and naming, schemas, destinations, migrations, and backward-compatible evolution.
Scalable storage and indexing : operating and improving systems built on PostgreSQL, Bigtable, Elasticsearch, Solr.
Reliable jobs execution : submission, queueing, execution, progress reporting, retries, cancellation, rate limiting, and operational tooling.
Cross-product capabilities : catalog data triggers, multi-dimensional data support, custom item types, catalog data enrichment, recommendations, and semantic catalog profiles for agentic use cases.
As a Senior SRE, you will be the team’s reliability and operability leader . You will work alongside backend engineers, embedded QA, Product, and Engineering Management to make complex product-data systems observable, scalable, safe to release, and straightforward to operate.
Fuse embraces AI-assisted engineering. We expect engineers to use modern coding agents thoughtfully to accelerate investigation, development, testing, documentation, and operational work while retaining full ownership of correctness, security, and production outcomes.
Working from one of our Central European offices (Bratislava, Prague, or Brno), or remotely (Czechia, Slovakia) on a full-time basis, you’ll become a core part of the Engineering organization.
What challenge awaits you?
As a P3 Senior SRE at Bloomreach, you are an independent reliability professional who can turn ambiguous operational problems into measurable improvements and lead initiatives end-to-end with minimal day-to-day guidance.
Your challenge will be to make Fuse’s distributed data platform dependable across the complete data path:
customer or integration → Data Hub API → records and transformations → items → asynchronous jobs execution engine → storage and indexes → Marketing, Search, Recommendations, and Loomi consumers
Your responsibilities
a. Platform reliability and observability
Own and improve the reliability posture of Fuse services, workers, APIs, queues, storage systems, and destination synchronization pipelines.
Establish meaningful SLIs, SLOs, and error budgets for customer-facing APIs, asynchronous jobs, catalog data freshness, destination synchronization, and indexing.
Build end-to-end observability across Data Hub item collections, from API request and job submission through processing, persistence, indexing, and downstream delivery.
Ensure engineers can trace a workspace, item collection, catalog, or job across services without manually correlating disconnected logs and database records.
Create and maintain actionable dashboards, alerts, and service health views using Grafana, Prometheus-compatible metrics, OpenTelemetry, PagerDuty, and GCP tooling.
Detect missing, stalled, duplicated, or inconsistent processing before customers or downstream teams report it.
Improve capacity planning and autoscaling using workload telemetry, queue depth, processing throughput, latency, memory usage, storage growth, and customer-level traffic patterns.
Reduce noisy alerts and replace symptom-based monitoring with signals tied to customer impact.
b. Reliability of catalog storage and indexing
Improve the availability, scalability, and operability of catalog data across PostgreSQL/Cloud SQL, Bigtable, Elasticsearch, GCS, Kafka, and related storage systems.
Support catalog placement, routing, index lifecycle, shard management, safe migration, and recovery across multiple Elasticsearch clusters.
Develop safeguards for full replacements, delta updates, deletions, schema changes, destination changes, and catalog reindexing.
Define and automate data-consistency checks between source records, transformed items, job state, Bigtable, Elasticsearch, and downstream destinations.
Help establish practical platform limits and quotas for catalog size, API traffic, job concurrency, queue depth, payload size, and expensive operations.
Partner with engineers on performance testing for large catalogs and high-throughput customer workloads.
c. Infrastructure, deployments, and release safety
Own and evolve Kubernetes configuration and operational infrastructure for Fuse components.
Improve deployment automation, progressive rollout, rollback, and validation across development and production environments.
Make coordinated releases safer when changes span app/app , Fuse workers, Kubernetes configuration, and PostgreSQL migrations.
Automate operational procedures that currently depend on manual commands, one-off scripts, or specialist knowledge.
Maintain CI/CD pipelines with tests, linters, dependency management, security checks, image publication, and release verification.
Create reusable tooling for local development, ephemeral environments, end-to-end testing, load testing, and production diagnosis.
Ensure runbooks remain executable and are validated through exercises rather than existing only as documentation.
d. Incident management and L3 support
Participate in and help improve the Fuse L3/on-call rotation .
Lead incident investigation, mitigation, stakeholder communication, and follow-up for Fuse-owned systems.
Use logs, metrics, traces, database state, queue state, and Kubernetes signals to diagnose failures across distributed workflows.
Build safe operational tools for common support activities such as job tracing, queue inspection, rate-limit diagnosis, catalog health checks, and index recovery.
Facilitate blameless incident reviews and ensure resulting actions address root causes rather than only immediate symptoms.
Improve the handoff between customer support, L2, Fuse L3, Infrastructure, and dependent engineering teams.
Reduce recurring support demand by turning incident knowledge into safeguards, automation, tests, dashboards, and clear documentation.
e. Security, isolation, and compliance
Help Fuse meet Bloomreach security and compliance requirements, including ISO and SOC 2 controls.
Enforce least-privilege access, workload identity, service-
Similar jobs
- Principal Site Reliability EngineerIBM · Brno, Czech RepublicFirst seen today
- Staff Site Reliability EngineerSentinellabs · Prague, Czech RepublicFirst seen todayremote
- Staff Site Reliability EngineerSentinellabs · Czech RepublicFirst seen todayremote
- Site Reliability EngineerThales · PrahaFirst seen yesterday
- Site Reliability EngineerFtmo · Prague officeFirst seen 9d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job