Description:
In this role, you'll build and operationalize machine learning and AI capabilities that from idea to production and deliver measurable product value. You'll develop data move pipelines, training and inference workflows, model-serving services, APIs, and evaluation frameworks for use cases such as recommendation, forecasting, anomaly detection, classification, NLP, semantic search, and generative AI.
Your work will include feature engineering, experiment design, model tuning, offline and online validation, and integration of models into scalable product architectures. You'll help improve LLM-based workflows, including prompt design, retrieval-augmented generation, vector search, guardrails, and response quality evaluation. You’ll also strengthen the engineering backbone around AI through CI/CD, monitoring, observability, testing, and model lifecycle automation so solutions are reliable, cost-aware, secure, and ready for enterprise scale.
What You Bring
- You bring strong programming skills in Python, along with working knowledge of Java or Go, for building production-grade services and APIs
- You have a solid understanding of machine learning fundamentals, including supervised and unsupervised methods such as classification, regression, clustering, ranking, and recommendation systems
- You have hands-on experience with deep learning frameworks (e.g., PyTorch or TensorFlow) for model training, fine-tuning, and inference
- You demonstrate strong capabilities in data preparation, feature engineering, data validation, and model evaluation using appropriate offline and online metrics
- You have experience building, deploying, and integrating ML models into production systems through batch, real-time, or streaming pipelines
- You are familiar with generative AI concepts, including LLMs, embeddings, vector databases, prompt engineering, and retrieval-augmented generation, and how to apply them in practical use cases
- You bring working knowledge of MLOps and modern data infrastructure, including experiment tracking, model versioning, CI/CD, and tools such as Spark, Kafka, Airflow, and feature stores
- You have experience operating ML systems in production, including monitoring for drift, latency, accuracy, cost, bias, and performing debugging and failure analysis to ensure reliability and business impact
About You
- You have 1-3+ years of experience in machine learning engineering, software engineering, or a related field, with a track record of deploying models into production
- You are passionate about building reliable, scalable ML systems and take ownership of delivering end-to-end solutions
- You balance experimentation with engineering rigor, making thoughtful trade-offs to ensure models are both innovative and production-ready
- You are a collaborative problem-solver who thrives in ambiguous environments and is motivated by delivering measurable business impact