掲載日 ・ 2026/08/04
楽天グループ株式会社
楽天グループ株式会社:1035521 Lead Site Reliability Engineer (Lead SRE) – Incentive Platform Department (INPD)
非公開
東京都
楽天グループ
インターネットサービス(EC、メディア、アプリ)
ネットワークエンジニア
会社名
楽天グループ株式会社
会社概要
未来を信じ、より良い明日を創っていく。
イノベーションを通じて、人々と社会をエンパワーメントする。私たちは、そんな想いを大切に世界の人々に喜びと楽しさを届けます。
楽天は、E コマース、FinTech、デジタルコンテンツ、通信など、70 を超えるサービスを展開し、世界10 億以上のユーザーに利用されています。
これら様々なサービスを、楽天会員を中心としたメンバーシップを軸に有機的に結び付け、他にはない独自の「楽天エコシステム」を形成しています。ダイバーシティ推進は、楽天にとって最優先の企業戦略のひとつです。従業員の出身は70カ国・地域以上。世界中からユニークで多様な文化的背景や視点を持つ優秀な人材が集まり、イノベーションの原動力になっています。社内カフェテリアにはベジタリアン、ハラル対応のメニューを用意。礼拝所(Prayer room)もあります。
また、仕事と育児の両立支援や、障がい者雇用・活躍促進も積極的に推進。社内のLGBT(※1)当事者やアライ(※2)に対して、情報共有やサポート体制の強化も進めています。誰もが自分らしく力を最大限発揮して働ける。それが楽天のダイバーシティです。
70を超えるサービスを提供し、世界30カ国にサービス展開拠点を持ち、従業員の出身国・地域数は100を超え、オープンポジション制度を活用して多様なキャリアを描くことができる点も魅力です。
フレックスタイム制度、事情に応じたリモートワークの活用が可能です。本社には託児所やフィットネスジム、三食無料で利用可能なカフェテリアが併設されるなど、社員を支える環境が整備されています。
ポジション
1035521 Lead Site Reliability Engineer (Lead SRE) - Incentive Platform Department (INPD)
仕事内容
Position:
Position Details
As a Lead Site Reliability Engineer (Lead SRE), you will set the technical direction for the stable operation and continuous improvement of our mission-critical services. You will define and drive initiatives to enhance service reliability, scalability, and performance across the team, spanning incident response, automation, monitoring, capacity planning, and observability strategy. You will serve as the primary technical authority within the SRE team, guiding and upskilling other engineers while working closely with product development, infrastructure, and security teams.
Responsibilities
- Service Quality Definition & Achievement: Define Service Level Objectives (SLOs) and Service Level Agreements (SLAs). Lead the planning and execution of improvement activities to achieve them. Drive the adoption and operation of Error Budgets across the team.
- Performance & Latency Improvement: Identify bottlenecks in service performance and latency. Lead the team in proposing and implementing solutions, setting technical standards for performance work.
- Incident Management & Troubleshooting: Act as incident commander during production outages, leading rapid restoration efforts. Drive Root Cause Analysis (RCA) processes and the implementation of systemic preventative measures.
- Operational Efficiency & Automation: Champion the automation of operational processes to reduce toil. Architect scalable operational frameworks and establish best practices for the team.
- Technical Leadership & Mentorship: Provide technical guidance and mentorship to SRE team members. Conduct technical design reviews, define engineering standards, and contribute to the overall skill development of the team.
- Cross-functional Collaboration: Lead collaboration with product development teams, infrastructure teams, security teams, and other relevant departments. Foster a DevOps culture and drive alignment on reliability goals across the organization.
- On-call: Participate in and help shape the 24/7 on-call rotation, including refining escalation paths and runbooks.
求める経験・スキル
Mandatory Qualifications:
- Bachelor's degree in Computer Science or related field, or equivalent practical experience.
- More than 5 years of hands-on experience in SRE, infrastructure engineering, or a related field, with demonstrated technical leadership experience.
- Experience building and operating production systems in public cloud (AWS, GCP, Azure, etc.) or private cloud environments.
- Extensive experience designing, building, operating, and scaling Kubernetes environments.
- Deep knowledge and hands-on experience building and operating modern monitoring, alerting, and logging tools (e.g., Prometheus, Grafana, ELK Stack, Datadog).
- In-depth knowledge of UNIX-like operating system internals and/or networking.
- Deep knowledge of IP network systems and protocols (TCP/IP, HTTP, etc.) and hands-on troubleshooting experience.
- Experience building automated workflows using CI/CD tools (e.g., Jenkins, CircleCI, GitLab, CI/CD).
- Experience developing operational automation tools and scripts using scripting languages such as Shell, Python, etc.
- Proven track record of leading production incident handling end-to-end (detection, triage, short-term / long-term fix, root cause analysis).
- Experience in system performance tuning and capacity planning.
- Proficiency with Git and GitHub for version control and collaboration.
- Strong communication, negotiation, and collaboration skills to articulate complex technical issues and align with internal and external stakeholders.
Desired Qualifications:
- Experience developing or maintaining GCP environments (e.g., GKE, Cloud Run, BigQuery, Cloud Monitoring, IAM).
- Experience in web application development.
- Deep knowledge and practical experience in observability, and a strong drive to improve services leveraging SLIs/SLOs.
- Experience implementing and operating error budgets, or a proven track record in toil reduction initiatives.
- Experience driving cross-team or org-wide reliability improvements (e.g., defining standards, leading postmortem culture).
- Experience working with cross-cultural global teams in different locations.