JOB DESCRIPTION
We are looking for a Python AI Engineer to join us in building and advancing our AI Indexing & Search Pipeline. In this role, you will be responsible for designing and implementing end-to-end pipelines for data crawling, content processing, vector embedding generation, and indexing to power high-performance semantic search capabilities.
-
Pipeline Development & Core Architecture:
-
Develop, optimize, and maintain the Python-based Batch AI Indexer system.
-
Build and operate an end-to-end data pipeline:
Content Crawling$\rightarrow$Data Processing$\rightarrow$Text Chunking$\rightarrow$Vector Embedding$\rightarrow$Elasticsearch Indexing. -
Build and optimize Semantic Search / Vector Search capabilities.
-
-
Data Crawling & Processing:
-
Crawl and collect content from diverse web sources (Homepage, Naver Blog, and other relevant web data sources).
-
Optimize text chunking strategies and data preprocessing tailored for language models.
-
-
AI Models & Search Engine Integration:
-
Utilize OpenAI Embeddings for content vectorization.
-
Integrate and manage Elasticsearch 8.x for vector storage and search (working extensively with 1536-dimensional
dense_vectorand the Korean Nori Analyzer). -
Utilize FAISS for local vector search experiments and rapid prototyping.
-
-
Batch, State & Reliability Management:
-
Design and implement mechanisms for Incremental Indexing and Full Re-indexing.
-
Manage incremental batch states using Redis.
-
Track batch execution history and metadata using PostgreSQL.
-
Implement robust retry mechanisms and error-handling logic to ensure high system reliability.
-
-
Scheduling, Deployment & Operations:
-
Configure automated job scheduling using APScheduler or Linux Crontab.
-
Deploy and operate batch applications on Linux environments using
systemdand Python virtual environments (venv). -
Perform logging, monitoring, and troubleshooting for batch job operations.
-
Collaborate closely with Backend, AI, and DevOps team members to deploy and scale system operations.
-
REQUIRED SKILLS AND EXPERIENCE
-
Core Experience & Technical Skills:
-
2–4+ years of strong hands-on development experience in Python.
-
Proficiency in web scraping and crawling libraries (e.g., BeautifulSoup, Scrapy, Playwright, or Selenium).
-
Practical experience with Elasticsearch 8.x (specifically Vector Search,
dense_vector, kNN search, and custom analyzer configurations like Nori). -
Deep understanding and hands-on experience with Embedding models (e.g., OpenAI Embeddings, HuggingFace) and text chunking/preprocessing techniques.
-
-
Databases & System Architecture:
-
Solid skills in Redis (for caching/state management) and PostgreSQL (or equivalent relational databases).
-
Strong knowledge of Batch Processing architectures, state management, idempotency, and error handling within data pipelines.
-
Experience with FAISS or other vector search libraries for local experimentation.
-
-
DevOps & Environment:
-
Proficient in Linux environments, application deployment via
systemd, and Python virtual environments. -
Experience with scheduling tools (
crontab,APScheduler). -
Strong command of Git, logging frameworks, and system troubleshooting.
-
-
Nice-to-Have Qualifications:
-
Experience with Korean Natural Language Processing (Korean NLP).
-
Experience with LLM orchestration frameworks such as LangChain or LlamaIndex.
-
Hands-on experience with Docker, Kubernetes, or CI/CD pipelines.
-

