Tóm tắt
JOB DESCRIPTION
We are looking for a Python AI Engineer to join us in building and advancing our AI Indexing & Search Pipeline. In this role, you will be responsible for designing and implementing end-to-end pipelines for data crawling, content processing, vector embedding generation, and indexing to power high-performance semantic search capabilities.
Pipeline Development & Core Architecture:
Develop, optimize, and maintain the Python-based Batch AI Indexer system.
Build and operate an end-to-end data pipeline:
Content Crawling $\rightarrow$ Data Processing $\rightarrow$ Text Chunking $\rightarrow$ Vector Embedding $\rightarrow$ Elasticsearch Indexing.
Build and optimize Semantic Search / Vector Search capabilities.
Data Crawling & Processing:
Crawl and collect content from diverse web sources (Homepage, Naver Blog, and other relevant web data sources).
Optimize text chunking strategies and data preprocessing tailored for language models.
AI Models & Search Engine Integration:
Utilize OpenAI Embeddings for content vectorization.
Integrate and manage Elasticsearch 8.x for vector storage and search (working extensively with 1536-dimensional dense_vector and the Korean Nori Analyzer).
Utilize FAISS for local vector search experiments and rapid prototyping.
Batch, State & Reliability Management:
Design and implement mechanisms for Incremental Indexing and Full Re-indexing.
Manage incremental batch states using Redis.
Track batch execution history and metadata using PostgreSQL.
Implement robust retry mechanisms and error-handling logic to ensure high system reliability.
Scheduling, Deployment & Operations:
Configure automated job scheduling using APScheduler or Linux Crontab.
Deploy and operate batch applications on Linux environments using systemd and Python virtual environments (venv).
Perform logging, monitoring, and troubleshooting for batch job operations.
Collaborate closely with Backend, AI, and DevOps team members to deploy and scale system operations.
REQUIRED SKILLS AND EXPERIENCE
Core Experience & Technical Skills:
2–4+ years of strong hands-on development experience in Python.
Proficiency in web scraping and crawling libraries (e.g., BeautifulSoup, Scrapy, Playwright, or Selenium).
Practical experience with Elasticsearch 8.x (specifically Vector Search, dense_vector, kNN search, and custom analyzer configurations like Nori).
Deep understanding and hands-on experience with Embedding models (e.g., OpenAI Embeddings, HuggingFace) and text chunking/preprocessing techniques.
Databases & System Architecture:
Solid skills in Redis (for caching/state management) and PostgreSQL (or equivalent relational databases).
Strong knowledge of Batch Processing architectures, state management, idempotency, and error handling within data pipelines.
Experience with FAISS or other vector search libraries for local experimentation.
DevOps & Environment:
Proficient in Linux environments, application deployment via systemd, and Python virtual environments.
Experience with scheduling tools (crontab, APScheduler).
Strong command of Git, logging frameworks, and system troubleshooting.
Nice-to-Have Qualifications:
Experience with Korean Natural Language Processing (Korean NLP).
Experience with LLM orchestration frameworks such as LangChain or LlamaIndex.
Hands-on experience with Docker, Kubernetes, or CI/CD pipelines.
Đang tải…
Nguồn: careers GEM Corporation
Global Software Development & AI Solutions
Hanoi & Ho Chi Minh City · Sync · 37 job
Chưa có referrer tại công ty này
Yêu cầu referral vẫn được Admin xử lý thủ công.
| Job | Địa điểm | Thời gian | Hành động |
|---|---|---|---|
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam | ||
Hanoi, Vietnam | Hanoi, Vietnam |
37 job
37 job
Tóm tắt
JOB DESCRIPTION
We are looking for a Python AI Engineer to join us in building and advancing our AI Indexing & Search Pipeline. In this role, you will be responsible for designing and implementing end-to-end pipelines for data crawling, content processing, vector embedding generation, and indexing to power high-performance semantic search capabilities.
Pipeline Development & Core Architecture:
Develop, optimize, and maintain the Python-based Batch AI Indexer system.
Build and operate an end-to-end data pipeline:
Content Crawling $\rightarrow$ Data Processing $\rightarrow$ Text Chunking $\rightarrow$ Vector Embedding $\rightarrow$ Elasticsearch Indexing.
Build and optimize Semantic Search / Vector Search capabilities.
Data Crawling & Processing:
Crawl and collect content from diverse web sources (Homepage, Naver Blog, and other relevant web data sources).
Optimize text chunking strategies and data preprocessing tailored for language models.
AI Models & Search Engine Integration:
Utilize OpenAI Embeddings for content vectorization.
Integrate and manage Elasticsearch 8.x for vector storage and search (working extensively with 1536-dimensional dense_vector and the Korean Nori Analyzer).
Utilize FAISS for local vector search experiments and rapid prototyping.
Batch, State & Reliability Management:
Design and implement mechanisms for Incremental Indexing and Full Re-indexing.
Manage incremental batch states using Redis.
Track batch execution history and metadata using PostgreSQL.
Implement robust retry mechanisms and error-handling logic to ensure high system reliability.
Scheduling, Deployment & Operations:
Configure automated job scheduling using APScheduler or Linux Crontab.
Deploy and operate batch applications on Linux environments using systemd and Python virtual environments (venv).
Perform logging, monitoring, and troubleshooting for batch job operations.
Collaborate closely with Backend, AI, and DevOps team members to deploy and scale system operations.
REQUIRED SKILLS AND EXPERIENCE
Core Experience & Technical Skills:
2–4+ years of strong hands-on development experience in Python.
Proficiency in web scraping and crawling libraries (e.g., BeautifulSoup, Scrapy, Playwright, or Selenium).
Practical experience with Elasticsearch 8.x (specifically Vector Search, dense_vector, kNN search, and custom analyzer configurations like Nori).
Deep understanding and hands-on experience with Embedding models (e.g., OpenAI Embeddings, HuggingFace) and text chunking/preprocessing techniques.
Databases & System Architecture:
Solid skills in Redis (for caching/state management) and PostgreSQL (or equivalent relational databases).
Strong knowledge of Batch Processing architectures, state management, idempotency, and error handling within data pipelines.
Experience with FAISS or other vector search libraries for local experimentation.
DevOps & Environment:
Proficient in Linux environments, application deployment via systemd, and Python virtual environments.
Experience with scheduling tools (crontab, APScheduler).
Strong command of Git, logging frameworks, and system troubleshooting.
Nice-to-Have Qualifications:
Experience with Korean Natural Language Processing (Korean NLP).
Experience with LLM orchestration frameworks such as LangChain or LlamaIndex.
Hands-on experience with Docker, Kubernetes, or CI/CD pipelines.
Đang tải…
Nguồn: careers GEM Corporation