僱用形式
Senior Site Reliability Engineer
招聘已結束
刊登於 09-09-2026
月薪 Monthly
Job Summary
Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen , you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.
Key Responsibilities
- Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
- Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
- Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
- Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
- Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.
Requirements
Qualifications & Requirements
- Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
- Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux , networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
- Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
- Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
- SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
- Location & Communication: Based in Hong Kong or Shenzhen . Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.
Leadership & Operational Excellence
- Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
- Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
- Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
- Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.
Benefits
- Competitive Package
- A dynamic and innovative team
- Collaborative, inclusive working environment
薪酬
月薪 Monthly
提防求職陷阱
謹慎選用可靠的求職招聘平台,小心提防求職陷阱及虛假招聘騙局。
在申請工作前先了解清楚僱主的公司組織檔案。
任何時候切勿向他人提供重要個人資料,包括身份證、銀行戶口、信用卡資料。
了解工作內容和薪酬是否符合市場現實;不要相信「無需專業和技能」、「工作內容輕鬆簡單」,而又「人工高、福利好」,這種工作並不可能存在。
HKESE 致力保障求職者用戶,嚴格審核每則招聘廣告。如你發現此招聘項目有任何問題,請聯絡通知我們。
你可能感興趣的工作






