**Senior Director, MLOps**
Work Location: HN/HCM
**Job Overview**
We are seeking a Senior Director of MLOps to lead the strategy, architecture, engineering, and
operational excellence of enterprise\-scale platforms and capabilities that enable AI/ML models
to move reliably and efficiently from experimentation to production.
**Key Responsibilities**
*Strategy \& Technology Leadership*
- Define and execute the enterprise MLOps strategy, target architecture, and technology roadmap aligned with the organization's AI and technology strategy.
- Own the end\-to\-end technical capabilities required to industrialize AI/ML solutions, from experimentation and training through deployment, serving, monitoring, continuous improvement, and retirement.
- Establish enterprise\-wide architecture principles, engineering standards, reference architectures, and best practices for production\-grade AI/ML systems.
- Drive a platform and self\-service engineering approach, enabling Data Scientists, ML
- Engineers, and AI Engineers to build and operate models with greater speed, consistency, and autonomy.
*Platform Engineering \& Automation*
- Lead the design and evolution of scalable platforms and engineering capabilities supporting machine learning, deep learning, and Generative AI workloads.
- Drive automation across the ML delivery lifecycle, including experimentation, data and feature processing, training, validation, deployment, monitoring, retraining, and decommissioning.
- Establish standardized CI/CD/CT practices for AI/ML, enabling repeatable, traceable, and controlled promotion of models across environments.
- Enable reusable and governed data and feature capabilities that improve consistency between model development and production inference while reducing duplicated engineering effort.
*Production Reliability \& Operational Excellence*
- Establish engineering practices that ensure AI/ML services operate with high levels of availability, scalability, performance, resilience, and recoverability.
- Define and govern appropriate SLIs, SLOs, and operational standards for production AI/ML systems.
- Drive end\-to\-end observability across models, data, infrastructure, and serving layers, including model quality, drift, feature/data freshness, latency, throughput, availability, resource utilization, and cost.
- Establish proactive mechanisms for detecting model degradation, data or feature anomalies, training\-serving inconsistencies, and production performance issues.
*Governance, Security \& Cost Management*
- Embed governance, security, auditability, and compliance controls throughout the ML lifecycle without creating unnecessary friction for engineering teams.
- Ensure traceability across data, features, code, experiments, models, deployments, and predictions where required.
- Partner with Security, Risk, Architecture, and AI Governance teams to establish appropriate controls and guardrails for production AI.
**Qualifications**
- Bachelor's or Master's degree in Computer Science, Software Engineering, Data/AI, or related fields.
- Extensive experience in MLOps, ML Platform Engineering, AI Infrastructure, Cloud Engineering, or DevOps/SRE, with large\-scale production leadership.
- Proven track record in leading engineering teams and enterprise\-scale AI/ML platforms.
- Expertise in end\-to\-end ML lifecycle, including pipelines, feature engineering, model registry, serving, observability, lineage, and governance.
- Strong knowledge of cloud\-native architecture, Kubernetes, distributed systems, and data platforms.
- Experience with AWS, Infrastructure as Code, GitOps, and container orchestration.
- Knowledge of Kafka, Spark, Airflow, Data Lakehouse, streaming, and distributed data processing.
- Knowledge of PyTorch, TensorFlow, and GenAI/LLM technologies.
- Expertise in reliability, security, observability, performance, and cost optimization at scale.
**Apply via:**
**Mail**
For more information, please contact the Talent Acquisition Department – VinSmartFuture