Senior MLOps Engineer (#5860)

Azerbaijan, Europe, Georgia, Kazakhstan, Turkey, Ukraine
Work type:
Office/Remote
Technical Level:
Senior
Job Category:
Software Development

Client Overview:
Our client is an Azerbaijani telecommunications company, the largest mobile network operator in Azerbaijan. The main products are: Fixed telephony, Mobile telephony, Internet services, Wireless broadband, and Value-added services.

Project Objectives:
The primary goal is to accelerate the client’s Data & AI initiatives via a secure, hybrid cloud foundation on AWS while systematically modernizing the IT estate as part of the cloud migration.

Key Project Objectives include:

  • Cloud Foundation & Landing Zone: Deploy target hybrid network architectures, establishing a secure Landing Zone and hybrid Data/AI platforms on AWS.
  • Security, Compliance & Governance: Operationalize on-prem tokenization (achieving zero raw PII in the cloud), resolve policy blockers to include AWS in the ISMS, and establish a Cloud Center of Excellence (CCoE) to govern Cloud adoption.
  • AI Chatbot & Voicebot Design & Implementation: Develop and operationalize a flagship Customer Care Chatbot and Voicebot as the first hybrid-setup consumer.

Responsibilities:

  • Build, operationalize, and automate end-to-end MLOps pipelines using Amazon SageMaker Pipelines and MLflow for experiment tracking, model versioning, and registry lifecycle management.
  • Design, deploy, and manage production SageMaker inference endpoints (real-time, serverless, and batch) and Amazon Bedrock API integrations for LLM/SLM deployment with cost controls and latency optimization (Bedrock API Gatekeeper).
  • Implement AgentOps / LLMOps frameworks (AgentCore, Bedrock Guardrails, Promptfoo) to manage multi-agent orchestration, prompt evaluation, safety guardrails, and RAG retrieval pipelines.
  • Operationalize real-time STT / TTS (Speech-to-Text / Text-to-Speech) voicebot pipelines and low-latency speech inference on hybrid/cloud GPU node pools for the flagship Customer Care Voicebot.
  • Optimize specialized GPU node pools (NVIDIA A100/L40S / EC2 GPU instance types) for Azerbaijani SLM/LLM model training, fine-tuning, and scalable inference workloads.
  • Establish automated CI/CD for Machine Learning using GitLab CI/CD pipelines and Infrastructure-as-Code (Terraform or AWS CDK) to enforce security-gated MLOps promotion workflows (from SageMaker Canvas/Sandbox to production).
  • Integrate data de-identification, Format Preserving Encryption (FPE), and tokenization wrappers into ML data pipelines to ensure zero raw PII enters AWS cloud environments during model training and inference.
  • Set up telemetry, performance monitoring, model drift detection, and cost anomaly alerting for AI/ML workloads using Amazon CloudWatch, Splunk, and FinOps spend control frameworks.
  • Collaborate with Data Engineering, AI Architects, and Cloud Teams to integrate vector storage/retrieval (RAG), Apache Spark/EMR-on-EKS runtimes, and local tokenization databases.
  • Author technical MLOps runbooks, model deployment procedures, governance documentation, and disaster recovery playbooks.

Requirements:

  • 4+ years of hands-on experience in MLOps, DataOps, or Platform Engineering with a primary focus on enterprise Amazon SageMaker (Pipelines, Feature Store, Model Registry, Endpoints).
  • Proven experience deploying and operating Generative AI, LLM/SLM models, and Amazon Bedrock services alongside agentic frameworks and RAG pipelines.
  • Hands-on expertise with MLflow for experiment tracking, model registry, and lifecycle management.
  • Solid experience in GPU optimization and orchestration (NVIDIA A100/L40S, AWS EC2 GPU instances) for model training, fine-tuning, and low-latency real-time inference (STT/TTS voice pipelines).
  • Proficient in building CI/CD for Machine Learning (GitLab CI/CD, GitHub Actions) and Infrastructure-as-Code (Terraform or AWS CDK).
  • Practical knowledge of LLMOps / AgentOps tools and methodologies (AgentCore, prompt evaluations, Bedrock Guardrails, vector databases for RAG).
  • Strong understanding of data security, privacy, and tokenization (FPE, handling sensitive/PII data within ML pipelines).
  • Proficient in Python, PySpark, Docker, and Kubernetes/EKS fundamentals for containerized ML workloads.

Nice-to-Have Skills:

  • AWS Certified Machine Learning – Specialty certification.
  • AWS Certified Solutions Architect – Associate/Professional or AWS Certified DevOps Engineer – Professional.
  • Experience in telecom domain AI/ML applications, low-latency real-time voice/chat processing (ASR/TTS), or hybrid cloud data sovereignty architectures.
  • Experience with EMR-on-EKS, Starburst/Athena, or Apache Iceberg data lake integrations.

Soft Skills & Team Fit:

  • Strong critical thinking, problem-solving, and analytical skills.
  • Excellent communication and collaboration skills to work closely with cross-functional teams (Data Engineering, AI/GenAI Engineers, Security, Cloud/Infrastructure).
  • Results-oriented, proactive mindset with strong ownership of deliverables within an Agile / Scrum framework.
  • Upper-Intermediate+ English level (written and spoken).

What we propose:

  • Opportunity to lead critical, high-impact Data & AI platform delivery for a major telecommunications operator.
  • Hands-on work with modern MLOps and GenAI stack (Amazon SageMaker, Amazon Bedrock, MLflow, AgentCore, STT/TTS voicebot pipelines).
  • Flexible remote work options with structured, predictable collaboration within a well-balanced team.

We offer*:

  • Flexible working format - remote, office-based or flexible
  • A competitive salary and good compensation package
  • Personalized career growth
  • Professional development tools (mentorship program, tech talks and trainings, centers of excellence, and more)
  • Active tech communities with regular knowledge sharing
  • Education reimbursement
  • Memorable anniversary presents
  • Corporate events and team buildings
  • Other location-specific benefits

*not applicable for freelancers

×

Easy apply

    or
    Refer a friend