**About VinSmart Future**
VinSmart Future (VSF) is the leading technology company within the Vingroup Corporation, formed by the merger of the group's entire technology ecosystem, including VinApp, VinIT, VinBigdata, and other tech units. As a core driver of Vingroup's future growth, VSF is at the forefront of technological development, with artificial intelligence (AI) as its foundation. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high\-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
**AI Model Optimization Manager**
**● Work Location:**
HCMC: Vincom Dong Khoi, District 1\.
**Key Responsibilities:**
**Fine\-tuning \& Model Customization**
- Perform fine\-tuning of open\-source foundation models (e.g., Llama, Mistral, Qwen) using techniques such as Supervised Fine\-Tuning (SFT), LoRA, and QLoRA to optimize model performance for enterprise\-specific use cases.
- Customize and adapt Large Language Models (LLMs) to meet business requirements, domain knowledge, and operational constraints.
**Model Optimization**
- Apply model compression and optimization techniques, including Quantization (AWQ, GPTQ, GGUF), Pruning, and Draft Model architectures (Speculative Decoding), to reduce inference latency, improve throughput, and lower operational costs.
- Optimize model deployment performance across various hardware environments. Alignment \& Reinforcement Learning Pipelines
- Design, implement, and manage advanced alignment and reinforcement learning pipelines, including RLHF (Reinforcement Learning from Human Feedback), RLAIF (Reinforcement Learning from AI Feedback), DPO (Direct Preference Optimization), and PPO\-based training workflows.
- Ensure model behavior aligns with business objectives, safety requirements, and user expectations.
**Large\-Scale Frameworks \& Infrastructure**
- Work extensively with high\-performance training and inference frameworks such as NVIDIA NeMo, vLLM, TensorRT\-LLM, DeepSpeed, or equivalent technologies.
- Leverage distributed training and inference techniques to improve scalability, efficiency, and resource utilization.
**LLM Evaluation \& Validation**
- Design, develop, and maintain comprehensive LLM evaluation pipelines to assess model quality before production deployment.
- Utilize industry\-standard benchmarks (e.g., MMLU, GSM8K) and advanced evaluation methodologies such as Ragas and LLM\-as\-a\-Judge frameworks.
- Establish evaluation metrics and processes to measure model performance, reliability, safety, and business effectiveness in real\-world environments.
**Programming Languages**
**Preferred Qualifications**
**Hands\-on Industry Experience**
- Minimum 4 years of hands\-on experience in NLP and Generative AI, with proven expertise in training, fine\-tuning, and aligning Large Language Models (LLMs).
- Strong practical experience in developing and deploying LLM\-based solutions for real\-world applications.
**Tools \& Framework Expertise**
- Deep expertise in the Hugging Face ecosystem, NVIDIA NeMo, vLLM/SGLang, PyTorch, and distributed training/resource management tools.
- Proven ability to manage large\-scale model training and inference workloads efficiently.
**Optimization Mindset**
- Strong understanding of Transformer architectures and modern LLM internals.
- Deep knowledge of GPU architecture, CUDA, and memory optimization techniques.
- Ability to identify and resolve training and deployment bottlenecks related to computational resources and system performance.
**Evaluation \& Data Engineering Skills**
- Solid Data Engineering capabilities for dataset collection, cleaning, preprocessing, and preparation for fine\-tuning workflows.
- Strong understanding of production\-grade LLM evaluation metrics and methodologies, including hallucination detection, accuracy measurement, toxicity assessment, and model safety evaluation.
- Experience building scalable evaluation and monitoring systems for production environments.
**Nice\-to\-Have Qualifications**
- Experience with low\-level hardware optimization, including custom CUDA kernel development.
- Proven track record of deploying LLM\-powered products in production environments serving large\-scale user bases.
- Experience with large\-scale AI infrastructure, model serving platforms, and performance engineering for enterprise AI systems.
**Why You’ll Love Working Here:**
- Flexible working hours and attendance policy (Work from Home on working Saturdays).
- Attractive compensation and bonus packages, highly competitive in the market.
- Exclusive employee benefits across the Group’s ecosystem in accordance with company
policies.
- Opportunity to work on large\-scale and strategic technology projects.
- Professional technology environment with leading scientists, experts, and engineers from top
technology companies in Vietnam and around the world.
- Free access to learning platforms such as Udemy, Coursera, and O’Reilly; internal
workshops; sponsorship for professional certifications; and exclusive mentoring programs
from the Group and Company leadership team.
- Full statutory insurance coverage in accordance with Vietnamese Labor Law (Social
Insurance, Health Insurance, Unemployment Insurance), along with private healthcare
insurance based on job grade and annual health check\-ups at reputable hospitals and
healthcare centers nationwide.
- Participation in internal activities, team\-building programs, and annual company events.