**Key Responsibilities**
**1\. DataOps / MLOps Infrastructure Operations \& Scaling**
- Operate, monitor, and scale streaming/orchestration systems such as Kafka and Airflow, as well as analytical databases such as ClickHouse or equivalent, running on Kubernetes.
- Ensure the reliability and availability of data pipelines and AI/ML inference services, including workloads running on GPU nodes.
- Design and improve observability across the data/ML infrastructure, including metrics, logging, alerting, SLOs, and service health monitoring.
**2\. Infrastructure Standardization – IaC / GitOps**
- Migrate manually managed or fragmented infrastructure components to Terraform and GitOps (ArgoCD).
- Design secure and reliable CI/CD pipelines for data and ML services, ensuring consistency, traceability, and auditability.
**3\. Data \& AI Infrastructure Security**
- Implement IAM least\-privilege, network policies, and network segmentation across data/ML infrastructure.
- Implement and maintain proper secrets management, including secure storage and regular credential rotation.
- Work with Security and Engineering teams to identify and remediate CVEs across infrastructure components, API gateways, authentication, and related systems.
**Must\-Have Qualifications**
- 5\+ years of experience in SRE, DevOps, or Platform Engineering.
- Proven experience designing and operating production\-grade, highly available infrastructure for data or ML workloads, including multi\-AZ/HA, disaster recovery, and capacity planning.
- Strong hands\-on experience with Kubernetes and cloud/infrastructure platforms.
- Experience operating data infrastructure such as Kafka, Airflow, ClickHouse, or equivalent technologies.
- Proven experience acting as an Incident Lead for high\-severity production incidents.
- Hands\-on experience defining and operating SLI/SLO, error budgets, monitoring, alerting, and postmortem processes.
- Strong understanding of infrastructure security, including IAM, network security, secrets management, and vulnerability remediation.
- Ability to design systems and infrastructure rather than simply operate existing runbooks.
- Strong ownership, troubleshooting, and cross\-functional collaboration skills.
**Ideal Candidate: A Senior SRE / Platform Engineer with 5 years of hands\-on experience as SRE / Platform Engineering cho Data Platform, ML Platform hoặc AI Platform.**
**Apply CV via [email protected] or contact for other information via phone/zalo: 0376796319**