We're looking for an engineer with hands-on experience building and evaluating GenAI services - from RAG and agentic reasoning systems to production-grade LLM deployments. You'll work closely with Frontend and Backend teams to bring AI agents into real products, with a strong focus on reliability, safety, and shipping working prototypes fast.
Hard Skills:
- Practical experience developing and evaluating GenAI services, including RAG systems, understanding of the "ReAct" (Reasoning + Acting) paradigm and Agentic RAG, LLM API integration, and prompt/context engineering, as well as training, fine-tuning, and deploying ML models in production environments.
- Knowledge of ML/GenAI frameworks: LangGraph or LangChain, PyTorch / TensorFlow, Hugging Face, OpenAI/Anthropic SDK.
- Practical commercial experience working with AI models via API (Gemini, Anthropic) and self-hosted models (Llama 3, Mistral, Mixtral), with an understanding of model differences based on functional/non-functional requirements (FR/NFR).
- Deep understanding of how LLMs interact with external APIs via Function Calling.
- Experience with vector databases, semantic search methods, and principles of database structuring and cleaning.
- Practical experience with at least one cloud platform (AWS, GCP, or Azure).
- Proficiency in Python and understanding of asynchronous programming.
- Understanding of the AI model lifecycle: monitoring, versioning, and quality evaluation (RAGAS, DeepEval), with hands-on experience using these tools.
- Experience with Guardrails: setting hard constraints on conversation topics and agent actions, filters that automatically strip personal data before sending requests to external LLMs, and the ability to build output filters that fact-check generated responses before they're displayed.
- Deterministic Logic Integration — running AI agents on strict schemas to prevent the model from "making things up."
- A plus: knowledge of automated testing approaches for evaluating responses across large datasets to measure hallucination rates before MVP launch.
- Knowledge of Human-in-the-loop mechanisms, ensuring agents cannot execute actions without final user verification.
- Ability to design memory systems that store context from a client's previous conversations and operations for personalization (Long-term Memory & User Context).
Soft Skills:
- Ability to clearly communicate complex technical concepts and mentor team members.
- Ability to quickly test and evaluate new libraries and approaches.
- Analytical problem solving - debugging complex "black boxes" and understanding why an agent behaves unpredictably.
- Focus on delivering a working prototype rather than a perfect research paper.
- Close collaboration with Frontend and Backend developers to seamlessly integrate AI agents into the required environment.