Middle/Senior Site Reliability Engineer

Middle/Senior Site Reliability Engineer

LOCATION: Hanoi/HCMC

JOB DESCRIPTION 

  • Own the reliability, availability and resilience of the audax platform across production environments for multiple bank clients
  • Design and operate cloud and hybrid infrastructure using infrastructure-as-code, supporting deployments on AWS, Azure, GCP or within client-managed environments
  • Build standardised, repeatable deployment blueprints so new bank clients can be onboarded quickly and safely
  • Define and track SLIs, SLOs and error budgets that align with client SLAs and banking-grade availability targets
  • Build end-to-end observability across microservices, integrations and third-party dependencies (core banking, card processors, eKYC providers, payment networks)
  • Lead incident response and blameless post-incident reviews, and communicate clearly with internal teams and client stakeholders during incidents
  • Design and regularly test disaster recovery, backup and business continuity plans to meet regulatory RTO and RPO requirements
  • Participate in the on-call rotation and reduce operational toil through automation
  • Work with security and compliance teams on access control, secrets management, encryption, vulnerability management and audit evidence
  • Support client and regulator technology due diligence, including infrastructure documentation and control assessments

REQUIRED SKILLS AND EXPERIENCE 

  • +3 years of experience in SRE, DevOps, platform or infrastructure engineering
  • Strong hands-on experience with at least one major cloud provider (AWS, Azure or GCP); multi-cloud exposure is a plus
  • Production experience with Kubernetes and containerised microservices
  • Proficiency with infrastructure-as-code (Terraform, Pulumi or similar) and configuration management
  • Experience with observability tooling such as Prometheus, Grafana, OpenTelemetry, Datadog, ELK or Splunk
  • Solid Linux fundamentals and networking knowledge (DNS, TCP/IP, load balancing, TLS, VPN and private connectivity)
  • Scripting or programming skills in at least one language (Python, Go, Bash or similar)
  • Experience with CI/CD and GitOps tooling (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
  • A calm, structured approach to troubleshooting and incident management
  • Clear written and spoken English, and comfort working with stakeholders across different countries

Nice to Have:

  • Experience in banking, payments, or other highly regulated financial environments
  • Familiarity with regulatory and security frameworks such as MAS TRM Guidelines, PCI DSS, ISO 27001, SOC 2, or central bank IT requirements in APAC or the Middle East
  • Experience deploying software into client-owned or on-premises bank environments with strict change management
  • Experience with hybrid connectivity between cloud platforms and bank data centres
  • Operating event streaming and messaging systems (Kafka, RabbitMQ or similar)
  • Database operations experience (PostgreSQL, Oracle, MongoDB, replication, backup and recovery)
  • Service mesh, API gateway or zero-trust networking experience
  • Cloud cost optimisation and multi-tenant platform design
  • Relevant certifications (AWS, Azure, GCP, CKA, CKS, CISSP)

BENEFITS 

  • Salary: up to 50M
  • Probation salary is 100% of the official salary 
  • 13th-month salary and annual performance review 
  • Bonus for special occasions each year (Labor Day, National Day, Solar New Year, Lunar New Year) 
  • Project bonus 
  • Employee’s professional certification and training allowances subject to company regulations 
  • Health Care Insurance  
  • Annual Health Assessment 
  • Social, health and unemployment insurance following government policy 
  • Enjoy company summer trips and other team-building activities held monthly and quarterly 
  • Professional, creative and dynamic working environment  
  • Work five days per week with flexible check-in time 

CONTACT