LLM / GenAI Engineer
Evlo AI · ·
Tech Stack Required
About the Role
About The Role The LLM / GenAI Engineer will build production AI systems that combine foundation models, retrieval, structured data, and backend services. The role covers the full lifecycle of generative AI features, including RAG pipelines, agentic workflows, model adaptation, evaluation, and deployment across cloud infrastructure. The engineer will work with applied scientists, platform engineers, and product teams to turn ambiguous use cases into reliable systems with measurable gains in quality, latency, cost, and safety. This role is based in New York, NY and is remote. Key Responsibilities Design and deploy RAG and agentic workflows using Python, LangChain, LlamaIndex, or custom orchestration services Build ingestion, chunking, embedding, retrieval, reranking, and citation pipelines using vector stores such as pgvector, Pinecone, Weaviate, or OpenSearch Develop LLM evaluation systems covering offline benchmarks, human review, LLM-as-judge workflows, regression testing, and production quality monitoring Implement model adaptation workflows, including prompt optimization, supervised fine-tuning, and parameter-efficient methods such as LoRA and QLoRA Integrate hosted and open-weight models through APIs and inference platforms, optimizing throughput, latency, context usage, and per-request cost Productionize AI services with Docker, Kubernetes, CI/CD, observability, and cloud infrastructure across AWS, GCP, or Azure Partner with security and product stakeholders to define guardrails for privacy, prompt injection, hallucination, access control, and responsible model behavior What We Are Looking For 3–8 years of experience in software engineering, machine learning engineering, or applied AI, including at least 2 years building LLM or generative AI systems Strong Python skills and experience developing production services with REST APIs, asynchronous programming, testing, and relational or document databases Hands-on experience designing RAG systems, including embedding models, hybrid search, vector databases, reranking, metadata filtering, and retrieval evaluation Proficiency with one or more LLM frameworks or platforms, such as LangChain, LlamaIndex, Hugging Face Transformers, vLLM, OpenAI APIs, or equivalent tools Practical understanding of fine-tuning, prompt engineering, tokenization, context windows, decoding strategies, model evaluation, and inference optimization Experience deploying AI workloads with Docker and cloud services; familiarity with Kubernetes, Terraform, monitoring, tracing, and experiment or model versioning Bachelor’s or master’s degree in computer science, machine learning, artificial intelligence, or a related technical field; Bonus: experience with multimodal models, speech systems, GPU optimization, distributed training, knowledge graphs, or AI safety and red-teaming Show more Show less
Ready to apply?
Takes you directly to Evlo AI's application page
About Evlo AI
Get similar jobs in your inbox
Weekly digest of AI engineering roles matched to your stack. Free forever.
Subscribe — Free