Site Reliability Engineer / Devops — Retail Engineering

Shanghai Posted 9/11/2026 Leaves the board in 7 days
ATS Keyword Match See which key terms from this posting your resume already has. Free.

You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines. You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.

Minimum Qualifications
Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience.
7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications.
Strong technical grasp of Open Source technologies designed for large-scale data processing.
Proven expertise in designing, analyzing, and troubleshooting complex distributed systems.
Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar).
Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.).
Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.

Preferred Qualifications
In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry).
Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ).
Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR).
Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.

Share your thoughts