This role requires your to operate, monitor, and triage applications in the production and non-production environments, ensuring reliability, scalability and performance across its low-latency, active/active cloud infrastructure.
The role will requires you to architect, build, and own robust CI/CD pipelines and automation systems, from design through long-term maintenance.
Minimum Qualifications
5 years of hands-on experience in infrastructure engineering, DevOps, or Site Reliability engineering.
Proficiency in at least one scripting or programming language such as Python, Go or Java.
Experience operating and supporting applications across cloud infrastructure ( eg, AliCoud, AWS and GCP).
Experience with Infrastructure as Code tools such as Terraform, Pulumi, and Helm.
Experience building and maintaining CI/CD pipelines.
Good command on Linux, networking concepts (TLS/SSL, DNS, Load Balancers, etc.,) and troubleshooting skills in large scale environments
Exposure to AI-assisted or AI-infused automation in DevOps workflows (e.g., using LLM-based tools, copilots, or intelligent agents for scripting, incident triage, or pipeline optimization).
Hands on experience in one or more databases (Relational / NoSQL like Oracle, MongoDB). Experience managing large Databases with very high transaction rates. Knowledge of Oracle Database architecture and replication.
Hands-on experience with Splunk, Prometheus, and Grafana.
BS or MS in Engineering, Computer Science, or equivalent education/experience.
Preferred Qualifications
8 years of hands-on experience in infrastructure, DevOps, or software engineering, with a proven track record of independently owning complex systems or initiatives from start to finish.
Extensive experience architecting and operating multi-region, high-availability infrastructure across AliCloud, AWS and GCP, with ownership of cloud cost, resilience, and scaling strategy.
Proven ownership of build/release systems and CI/CD pipelines at scale, including designing pipeline architecture, improving reliability, and reducing release friction.
Proven experience in Terraform, Pulumi, and Helm, including designing reusable modules and adopting organizational-level IaC standards.
Proven ability to operate with minimal direction: scoping ambiguous problems, proposing architecture, executing, and being accountable for outcomes.
Hands-on experience designing and building AI-powered automation for DevOps, including LLM-powered runbooks, autonomous remediation agents, anomaly detection pipelines, and AI-assisted code/infra review tools, with a proven track record of team adoption.