- Build and optimize data pipelines for diverse use cases (data mining, audience targeting, analytics) with strict compliance to data privacy regulations;
- Engineer reliable pipelines to synchronize data across databases, powering batch + stream processing, large-scale analytics, and low-latency online serving;
- Collaborate with Data Science, Policy, System Operation, and Research teams to understand their data needs and deliver solutions;
- Establish solid design and engineering best practices for both technical and non-technical partners.
Technical skills:
- 5+ years hands-on in the Big Data ecosystem (Spark, Kafka, ClickHouse, Cassandra or similar);
- Open table formats (Apache Iceberg / Delta / Hudi) with an open catalog (Unity / Polaris / Glue); storage & compute optimization (compression, partitioning, bucketing, ORC/Parquet, compaction, small-file handling);
- Reliable pipelines syncing data across databases; CDC (Debezium), Kafka; unified batch + stream processing with exactly-once/idempotency, schema evolution, backfill;
- Medallion Lakehouse; dimensional modeling / SCD / data vault; ELT (dbt) + semantic/metric layer;
- Data contracts, testing, SLA/SLO, data observability & lineage;
- Build the self-serve, AI-backed AutoEDA / analytics platform enabling stakeholders to mine data without engineering overhead (serving DS/ML — model training not required);
- API & Platform: High-throughput APIs (gRPC, REST) on Kubernetes; latency & scalability optimization;
- Airflow, Docker/containerization, CI/CD, observability (Prometheus, Grafana, OpenTelemetry), canary/blue-green deployment;
- Strong Scala or Python for high-performance data processing; SQL;
- Compliance with Vietnam's Personal Data Protection Law 2025 (Law 91/2025/QH15) & Decree 356/2025/NĐ-CP; GDPR for cross-border data;
Leadership & Stakeholder collaboration:
- Mentor, code review, build engineering culture;
- Collaborate with DS / BI / Product / Partner teams under a data-as-product mindset.
Nice to have:
- Strong plus: HDFS / object-storage administration: backup, cleanup, archiving, capacity planning;
- Centralized Feature Store (4,000+ features): point-in-time correctness, low-latency online serving, online/offline consistency, batch + real-time ingestion;
- AutoML & MLOps: auto-training, evaluation, prediction at scale; model registry & serving; end-to-end ML pipeline understanding;
- Fundamental knowledge of data science workflows and ML integration;
- Experience with DMP/CDP and vector databases for AI/RAG use cases.
Join the Audience Platform team at Zalo Group and architect the data backbone for one of Vietnam's largest digital ecosystems. With nearly 80 million Monthly Active Users and TBs of data flowing through our systems daily, we're looking for a Data Engineer to build robust pipelines that power large-scale data mining and give decision-makers seamless access to actionable insights.
Scope of work:
You won't just move data - you'll build the platforms that make data mining effortless. You will develop a centralized Feature Store and self-served, AI-backed automated analytics (AutoEDA) serving the Ads, Fintech, and VAS business lines. Your mission is to bridge massive datasets and business impact through thoroughly architected, highly optimized, rock-solid data infrastructure.
Join the Audience Platform team at Zalo Group and architect the data backbone for one of Vietnam's largest digital ecosystems. With nearly 80 million Monthly Active Users and TBs of data flowing through our systems da...