Rakuten
Site Reliability Engineer
RSIN_Bangalore
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 6 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Job Description:
Designation: Site Reliability Engineer
Experience: 5-10 years
Location: Bengaluru, India
Why should you choose us?
Rakuten Symphony is reimagining telecom, changing supply chain norms and disrupting outmoded thinking that threatens the industry’s pursuit of rapid innovation and growth. Based on proven modern infrastructure practices, its open interface platforms make it possible to launch and operate advanced mobile services in a fraction of the time and cost of conventional
approaches, with no compromise to network quality or security. Rakuten Symphony has operations in Japan, the United States, Singapore, India, South Korea, Europe, and the Middle East Africa region. For more information, visit: https://symphony.rakuten.com
Building on the technology Rakuten used to launch Japan’s newest mobile network, we are taking our mobile offering global.
To support our ambitions to provide an innovative cloud-native telco platform for our customers, Rakuten Symphony is looking to recruit and develop top talent from around the globe. We are looking for individuals to join our team across all functional areas of our business – from sales to engineering, support functions to product development. Let’s build the future of mobile telecommunications together!
About Rakuten Group, Inc. (TSE: 4755) is a global leader in internet services that empower individuals, communities, businesses and society. Founded in Tokyo in 1997 as an online marketplace, Rakuten has expanded to offer services in e-commerce, fintech, digital content and communications to approximately 1.9 billion members around the world. The Rakuten Group has over 30,000 employees, and operations in 30 countries and regions. For more
information visit https://global.rakuten.com/corp/.
- ■ Role Overview
- L2 SRE supporting day-to-day operations of RCP and BMM escalated from L1 SRE.
L2 SRE responsible for advanced incident resolution, operations, and automation development for RCP and BMM. Acts as escalation point for L1 SRE and drives proactive improvements to RCP/BMM stability.
Responsibilities include network monitoring, fault management (Basic troubleshooting), ticket management, alarm monitoring, performance monitoring, configuration management, new node handover and transfer to operation (HOTO).
- ■ Work Conditions
- Work Location: Rakuten Symphony India (Bangalore)
Work Time: IST (India)
Holiday Calendar: As per Rakuten India (RSI) Holiday Calendar (India)
Human Skills
Should have interpersonal, verbal, and written communication.
Should have the ability to work in a team.
Should be ready for working in 24*7 support (24*7 Shift)
Proactive, punctual, and ready to learn new skills.
Can make prompt and accurate status reportingWork Experience
Experience in providing infrastructure support for applications, testing third-party applications, and operating container platforms/application is required
Experience in leading solution implementation and designing strategies for Cloud applications' deployment is also a must.
Previous experience in the telecommunications industry is preferred.
Technical
[Minimum Technical Qualifications]
Experience in private cloud (on-premises private cloud) design, deployment, and operations
Strong Kubernetes expertise with CKA certification or equivalent
Virtualization technologies (VMware, KVM)
Deep Linux knowledge including performance tuning and troubleshooting
Proven track record in incident management and root cause analysis
[Preferred / Value-added Skill Details]
Have experience in Telecommunication domain.
Technical leadership and communication skills
Basic knowledge of database operations (SQL)
Experience with cloud and telecom networking fundamentals
Development skill sets, Automation technologies in Python, Bash, ansible.
[Soft Skills / Expectations]
Good communication and coordination skills
Ability to follow processes and work in 24/7 support environments
Willingness to learn and keep improving
Basic analytical and problem-solving mindset
L2 Tasks
[BMM operations & management]
Manage and operate the BMM and take necessary action for seamless running of application
Upgrade BMM application
Develop/modify python or shell script to automate repetitive RCP cluster operations executed as BMM workflows
[RCP operation support]
In response to user requests, update, create, or delete inventory data stored in external inventory and managed by BMM
Support RCP SRE for BMM workflow execution for RCP cluster operation like cluster building, node commissioning, and node decommissioning
[L1 SRE support]
Support the L1 team , including assistance with issues, escalations, and operational guidance, while ensuring overall BMM service stability and continuity
Transfer knowledges of routine works to L1 SRE
[Change & Operational Activities]
Raise Change Requests (CRs) and execute it for resolving troubles.
Participate in huddle calls to report incidents and CR updates
[Common]
Analyze, understand, and clearly explain complex technical concepts to development teams (as L3 support) with detailed and actionable requirements
Lead and support solution implementation , adopting modern cloud-native architecture, Agile practices, and DevOps delivery models.
Provide 24×7×365 operational support in 8-hour shifts based on a mutually agreed roster plan.
Manage operational projects and track issues through the ticketing system , including change requests and service requests.
Perform incident management, including RCA, corrective actions, and service restoration in coordination with internal teams and vendors.
Actively participate in ERT calls and support critical incident resolution.
Prepare and maintain MOPs, operational documentation, and internal Wiki for knowledge sharing and standardization.
Provide infrastructure support for applications running on fully virtualized and containerized cloud-native mobile networks.
Operate, manage, and maintain container platforms , including Kubernetes-based environments.
Create and manage Kubernetes clusters using ROBIN , including initial setup, lifecycle management, and decommissioning.
Perform Kubernetes and RCP cluster upgrades , rebuilds, and maintenance activities following approved procedures.
Provide deployment and post-deployment support to application teams.
Analyze, understand, and clearly explain complex technical concepts to development teams with detailed and actionable requirements.
Lead and support solution implementation , adopting modern cloud-native architecture, Agile practices, and DevOps delivery models.
Provide 24×7×365 operational support in 8-hour shifts based on a mutually agreed roster plan.
Manage operational projects and track issues through the ticketing system, including change requests and service requests.
Perform incident management , including RCA, corrective actions, and service restoration in coordination with internal teams and vendors.
Actively participate in ERT calls and support critical incident resolution.
Prepare and maintain MOPs, operational documentation, and internal Wiki for knowledge sharing and standardization.
Conduct testing and validation of changes prior to production implementation to ensure stability and reliability.
Develop automation scripts to support monitoring, routine operational tasks, and efficiency improvements.
Support the L1 team , including assistance with hardware-related issues, escalations, and operational guidance, while ensuring overall cluster health and hygiene.
RAKUTEN SHUGI PRINCIPLES:
Our worldwide practices describe specific behaviours that make Rakuten unique and united
across the world. We expect Rakuten employees to model these 5 Shugi Principles of Success.
Always improve, always advance. Only be satisfied with complete success - Kaizen.
Be passionately professional. Take an uncompromising approach to your work and be determine
Listed on hirly, a job board. hirly is not the employer: Rakuten is hiring for this role.
Similar jobs
- Site Reliability EngineerRyansg · Mumbai - IndiaFirst seen today
- Site Reliability Engineer, AVPRbs · 3 LocationsFirst seen today
- Site Reliability Engineer (II)Ncr · CHENNAI, INDFirst seen today
- Site Reliability Engineer ICmegroup · Bangalore - Bagmane TridibFirst seen today
- Site Reliability EngineerOkta · Bengaluru, IndiaFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job