MLOps Engineer
Evlo AI · ·
About the Role
About The Role The MLOps Engineer owns the infrastructure and automation that move machine learning models from experimentation into reliable, observable production services. The role spans training pipelines, model registries, deployment workflows, feature and data validation, and operational monitoring across cloud environments. You will work with ML engineers, data scientists, platform engineers, and application teams to make model delivery repeatable and safe. The team is building systems where deployment speed, inference latency, reproducibility, security, and model quality all matter. Key Responsibilities Build and maintain CI/CD pipelines for ML using GitHub Actions, GitLab CI, or Jenkins, integrating automated testing, model validation, and release approvals Deploy and operate training and inference workloads on AWS, GCP, or Azure using Docker, Kubernetes, and infrastructure-as-code tools such as Terraform Develop reproducible training and orchestration workflows with Python, Airflow, Kubeflow, Argo Workflows, or equivalent platforms Implement model lifecycle management with tools such as MLflow, including experiment tracking, artifact versioning, model registries, and promotion workflows Instrument production systems to monitor latency, throughput, resource utilization, data drift, model performance, and service reliability using Prometheus, Grafana, Datadog, or equivalent tools Improve platform reliability through autoscaling, rollback strategies, canary deployments, secrets management, incident response, and disaster recovery procedures Partner with engineering and data science teams to establish standards for testing, reproducibility, governance, documentation, and secure access to ML infrastructure What We Are Looking For 3–8 years of experience in MLOps, ML platform engineering, DevOps, software engineering, or a closely related discipline, including experience supporting production machine learning systems Strong Python and Linux skills, with the ability to build maintainable automation, services, command-line tools, and operational workflows Hands-on experience with Docker, Kubernetes, CI/CD, cloud infrastructure, and infrastructure-as-code using Terraform, Pulumi, or equivalent tools Experience with ML lifecycle and orchestration technologies such as MLflow, Kubeflow, Airflow, SageMaker, Vertex AI, or Azure ML Solid understanding of production ML systems, including model serving, feature and data validation, experiment reproducibility, model versioning, monitoring, and rollback procedures Bachelor’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent practical experience Bonus: Experience with Ray, Feast, Argo CD, Spark, GPU workloads, service mesh technologies, or regulated environments; familiarity with LLM serving and evaluation infrastructure is a plus Show more Show less
Ready to apply?
Takes you directly to Evlo AI's application page
About Evlo AI
Get similar jobs in your inbox
Weekly digest of AI engineering roles matched to your stack. Free forever.
Subscribe — Free