LLM / GenAI Engineer
Evlo AI · ·
About the Role
About The Role The LLM / GenAI Engineer will design, build, and deploy production AI systems, including retrieval-augmented generation pipelines, tool-using agents, model fine-tuning workflows, and evaluation infrastructure. The role spans experimentation and backend engineering, with a focus on reliability, latency, cost, and measurable model quality. Working with applied scientists, platform engineers, and product teams, the role will turn emerging foundation-model capabilities into secure, observable services used in real customer workflows. The position is remote, with a preference for candidates based in Austin, TX. Key Responsibilities Design and implement production RAG and agentic workflows using Python, LangChain, LlamaIndex, or equivalent frameworks Build ingestion, chunking, embedding, reranking, and retrieval pipelines backed by vector stores such as pgvector, Pinecone, Weaviate, or Milvus Develop LLM evaluation systems covering offline benchmarks, groundedness, factuality, safety, latency, cost, and regression testing Fine-tune and optimize open-source language models using supervised fine-tuning, LoRA, QLoRA, quantization, and distributed training techniques Deploy model-powered services through containerized APIs and cloud infrastructure using Docker, Kubernetes, AWS, GCP, or Azure Instrument production systems with tracing, logging, feedback collection, and monitoring for quality degradation, drift, latency, and token usage Partner with software engineers and applied scientists on architecture reviews, data strategy, experimentation, code quality, and production incident response What We Are Looking For 3–8 years of experience in software engineering, machine learning engineering, or applied AI, including at least 1 year delivering LLM or GenAI systems to production Advanced Python skills with experience building asynchronous services, REST or gRPC APIs, testing frameworks, and maintainable production code Strong understanding of transformer-based language models, embeddings, tokenization, context windows, prompting, structured generation, and inference tradeoffs Hands-on experience with RAG architecture, vector databases, hybrid search, reranking, document processing, and retrieval-quality measurement Experience with at least one major cloud platform and production deployment tooling, including Docker, Kubernetes, CI/CD, and observability systems Bachelor’s or master’s degree in computer science, engineering, mathematics, or a related technical field, or equivalent practical experience Bonus: Experience with distributed GPU training or inference, open-source model serving, multimodal models, guardrails, MLflow, vLLM, TensorRT-LLM, or enterprise AI security Show more Show less
Ready to apply?
Takes you directly to Evlo AI's application page
About Evlo AI
Get similar jobs in your inbox
Weekly digest of AI engineering roles matched to your stack. Free forever.
Subscribe — Free