**About VinSmart Future**
VinSmart Future (VSF) is Vingroup's technology company, formed by merging the Group's entire technology ecosystem. As a core driver of Vingroup's future growth, VSF is AI\-first \- with artificial intelligence as the foundation of everything we build. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high\-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
As the "Digital Core" of the Group, we develop foundational platforms, data infrastructure, and AI capabilities that power Vingroup's operations at scale \- from a unified SuperApp and enterprise systems to proprietary AI models, ride\-hailing technology, digital healthcare, travel, loyalty, and customer service. Our work is embedded across every major business vertical of Vingroup \- driving efficiency, innovation, and growth at scale.
**Senior Director, MLOps**
● Work Location:
○ Hochiminh: Vincom Dong Khoi, HCM
○ Hanoi: Technopark Tower, Gia Lam, Hanoi
**Job Overview**
- We are seeking a Senior Director of MLOps to lead the strategy, architecture, engineering, and operational excellence of enterprise\-scale platforms and capabilities that enable AI/ML models to move reliably and efficiently from experimentation to production.
- This role is accountable for establishing a scalable, secure, and highly automated ML engineering ecosystem covering data and feature enablement, model development and deployment, lifecycle governance, production reliability, and observability.
- The successful candidate will combine strong technical leadership with organizational and strategic capabilities, working closely with AI, Data, Engineering, Cloud Infrastructure, Architecture, Security, and business teams to accelerate AI adoption while ensuring reliability, governance, security, and cost efficiency at scale.
**Key Responsibilities**
**Strategy \& Technology Leadership**
- Define and execute the enterprise MLOps strategy, target architecture, and technology roadmap aligned with the organization's AI and technology strategy.
- Own the end\-to\-end technical capabilities required to industrialize AI/ML solutions, from experimentation and training through deployment, serving, monitoring, continuous improvement, and retirement.
- Establish enterprise\-wide architecture principles, engineering standards, reference architectures, and best practices for production\-grade AI/ML systems.
- Drive a platform and self\-service engineering approach, enabling Data Scientists, ML Engineers, and AI Engineers to build and operate models with greater speed, consistency, and autonomy.
- Lead technology evaluation and strategic build\-versus\-buy decisions across the MLOps ecosystem.
- Anticipate emerging AI/ML infrastructure trends and translate them into practical technology strategies and investment priorities.
**Platform Engineering \& Automation**
- Lead the design and evolution of scalable platforms and engineering capabilities supporting machine learning, deep learning, and Generative AI workloads.
- Drive automation across the ML delivery lifecycle, including experimentation, data and feature processing, training, validation, deployment, monitoring, retraining, and decommissioning.
- Establish standardized CI/CD/CT practices for AI/ML, enabling repeatable, traceable, and controlled promotion of models across environments.
- Enable reusable and governed data and feature capabilities that improve consistency between model development and production inference while reducing duplicated engineering effort.
- Establish capabilities for experiment tracking, artifact management, model versioning, lineage, registry, approval workflows, and reproducibility.
- Enable deployment strategies such as canary releases, shadow deployment, A/B testing, champion/challenger, automated rollback, and controlled model promotion.
- Ensure platforms support both traditional ML and large\-scale GenAI/LLM workloads, including distributed training, fine\-tuning, batch inference, real\-time inference, and high\-performance model serving.
**Production Reliability \& Operational Excellence**
- Establish engineering practices that ensure AI/ML services operate with high levels of availability, scalability, performance, resilience, and recoverability.
- Define and govern appropriate SLIs, SLOs, and operational standards for production AI/ML systems.
- Drive end\-to\-end observability across models, data, infrastructure, and serving layers, including model quality, drift, feature/data freshness, latency, throughput, availability, resource utilization, and cost.
- Establish proactive mechanisms for detecting model degradation, data or feature anomalies, training\-serving inconsistencies, and production performance issues.
- Lead the development of incident management, root\-cause analysis, disaster recovery, rollback, and post\-incident improvement practices for AI/ML systems.
- Ensure effective capacity and performance management for compute\-intensive AI workloads, particularly GPU\-based training and inference.
- Drive continuous improvement in platform reliability, deployment efficiency, operational automation, and engineering productivity.
**Governance, Security \& Cost Management**
- Embed governance, security, auditability, and compliance controls throughout the ML lifecycle without creating unnecessary friction for engineering teams.
- Ensure traceability across data, features, code, experiments, models, deployments, and predictions where required.
- Partner with Security, Risk, Architecture, and AI Governance teams to establish appropriate controls and guardrails for production AI.
- Establish standards for model validation, production readiness, approval, access control, lineage, versioning, and lifecycle management.
- Drive FinOps practices for AI/ML, optimizing infrastructure utilization and economics across training, inference, storage, and GPU workloads.
- Own capacity planning and provide input into strategic infrastructure and technology investments supporting future AI growth.
**Leadership \& Organization Development**
- Build, lead, and develop high\-performing engineering teams responsible for enterprise MLOps capabilities.
- Define organizational capabilities, engineering roles, competency frameworks, and career development paths required to support the organization's AI ambitions.
- Recruit, coach, and develop engineering leaders and senior technical talent.
- Foster a culture of engineering excellence, automation, reliability, ownership, and continuous improvement.
- Establish effective operating models and ways of working across AI, Data Science, Data Engineering, Software Engineering, Cloud/Infrastructure, Architecture, Security, and Risk teams.
- Own strategic roadmap, resource planning, budget, technology investments, and vendor/partner management within the assigned scope.
- Act as a senior technology leader and trusted advisor on MLOps, AI infrastructure, production AI, and ML engineering practices.
**Qualifications**
- Bachelor's or Master's degree in Computer Science, Software Engineering, Data/AI, or related fields.
- Extensive experience in MLOps, ML Platform Engineering, AI Infrastructure, Cloud Engineering, or DevOps/SRE, with large\-scale production leadership.
- Proven track record in leading engineering teams and enterprise\-scale AI/ML platforms.
- Expertise in end\-to\-end ML lifecycle, including pipelines, feature engineering, model registry, serving, observability, lineage, and governance.
- Strong knowledge of cloud\-native architecture, Kubernetes, distributed systems, and data platforms.
- Experience with AWS, Infrastructure as Code, GitOps, and container orchestration.
- Knowledge of Kafka, Spark, Airflow, Data Lakehouse, streaming, and distributed data processing.
- Knowledge of PyTorch, TensorFlow, and GenAI/LLM technologies.
- Expertise in reliability, security, observability, performance, and cost optimization at scale.
- Strong strategic and architectural thinking with a platform/product mindset.
- Ability to drive enterprise technology decisions and balance speed, scalability, reliability, security, governance, and cost.
- Proven leadership of senior engineering and technical teams.
**Preferred Qualifications**
- Experience building or leading enterprise\-scale MLOps/AI platforms.
- Experience with self\-service and internal developer platforms.
- Knowledge of GenAI/LLM, RAG, vector databases, distributed inference, and GPU optimization.
- Knowledge of Responsible AI, AI Security, Model Risk Management, and AI Governance.
- Experience with large\-scale cloud/GPU infrastructure and AI FinOps.
- AWS Professional/Specialty, CKA, or equivalent certifications are a plus.
**Why You’ll Love Working Here:**
- Flexible working hours and attendance policy (Work from Home on working Saturdays).
- Attractive compensation and bonus packages, highly competitive in the market.
- Exclusive employee benefits across the Group’s ecosystem in accordance with company policies. • Opportunity to work on large\-scale and strategic technology projects.
- Professional technology environment with leading scientists, experts, and engineers from top technology companies in Vietnam and around the world.
- Exclusive mentoring programs from the Group and Company leadership team.
- Full statutory insurance coverage in accordance with Vietnamese Labor Law (Social Insurance, Health Insurance, Unemployment Insurance), along with private healthcare insurance based on job grade and annual health check\-ups at reputable hospitals and healthcare centers nationwide.
- Participation in internal activities, team\-building programs, and annual company events.