- Kubernetes Platform Management: Maintain and optimize Kubernetes clusters via Rancher and Harbor registry; implement and streamline GitOps workflows using ArgoCD;
- Infrastructure as Code & Automation: Provision, manage, and automate infrastructure using Terraform, Terragrunt, Ansible, and Atlantis;
- Observability & Monitoring: Deploy, operate, and optimize centralized logging, metrics, and alerting pipelines built on Prometheus, AlertManager, Loki, Mimir, Grafana, and Graylog;
- Security & Secret Management: Enforce infrastructure security and manage credentials securely using HashiCorp Vault, SOPS, and git-crypt;
- MLOps Infrastructure: Build custom authentication middleware for MLFlow and develop MLFlow model synchronization tools in Golang;
- System Operations & Support: Monitor system health, troubleshoot production issues, and ensure high availability for dependent services
- Ability to program (structured and OOP) using one or more high-level languages, such as Python, Golang;
- Experience with dynamic resource management frameworks (Kubernetes, Nomad, Yarn);
- Experience manage infrastructure as code (Terraform,..);
- Experience with source version control (git, svn...), as well as configuration management (Ansible, Puppet, Salt stack...);
- Experience with distributed storage technologies such as NFS, HDFS, Ceph and Amazon S3;
- Proactive approach to identifying problems, performance bottlenecks, and areas for improvement;
Preferred skills and qualifications:
- Previous success in technical engineering;
- Coding experience beyond simple scripts.
We are seeking a System Engineer to operate, automate, and enhance the reliability of large-scale infrastructure, directly supporting core products and services including Zalo, Zalo AI, Kiki, etc
We are seeking a System Engineer to operate, automate, and enhance the reliability of large-scale infrastructure, directly supporting core products and services including Zalo, Zalo AI, Kiki, etc