- Build and operate the online serving data pipeline end-to-end, from raw log ingestion and feature extraction to model scoring and low-latency serving for Deep Learning models that drive ads distribution;
- Collaborate with Data Scientists to assess new requirements, surface data and resource constraints, propose alternatives, and proactively source new data to improve model quality;
- Connect and process data across teams, including dev and product, to enable shared use and analytics;
- Optimize large-scale storage and feature serving to sustain high-throughput query and aggregation cycles;
- Set up monitoring, alerting, and dashboards so pipeline issues are caught early and stakeholders can track data quality, model performance, and system health;
- Establish solid design and engineering best practices for both technical and non-technical stakeholders.
- 4+ years of working experience in Big Data and production data pipelines;
- Strong command of Python for high-performance data processing; Scala is a plus;
- Proficiency with orchestration and deployment workflows: Airflow, Docker/containerization, CI/CD, and Kubernetes (K8s);
- Proven ability to optimize storage and query performance on columnar stores (e.g., ClickHouse) through partitioning, indexing, and compression;
- Experience building dashboards and observability with tools such as Streamlit or Metabase;
- Solid skills in data ingestion, transformation, and analysis, and in synchronizing data across diverse databases for batch and stream processing;
- A proactive, research-driven mindset toward the global technology landscape and a strong willingness to adopt emerging technologies, AI-driven tools, and AI-native development in the workflow;
- Intellectual curiosity, effective communication, and strong stakeholder collaboration skills.
Nice to have
- Working knowledge of data science / ML workflows and the model lifecycle (training, evaluation, serving), and how data engineering decisions affect model accuracy;
- Domain knowledge of digital ads or experience in the advertising field;
- Experience deploying and serving LLMs in production (e.g., inference optimization, scalable serving infrastructure, latency/throughput tuning).
Data Engineers in the Adtima team build and operate the data backbone of our ads ecosystem, turning raw logs into training-ready features and production-grade model scores, and working closely with Data Scientists to make experimentation fast and reliable. They also support product teams with data for analytics tasks.
Data Engineers in the Adtima team build and operate the data backbone of our ads ecosystem, turning raw logs into training-ready features and production-grade model scores, and working closely with Data Scientists to ...