Artificial Intelligence Specialist
Stelvio Inc. · ·
Tech Stack Required
About the Role
Principal AI/ML Engineer – LLM & Inference Systems Location: Austin, TX Employment Type: Full-time The Opportunity Stelvio is supporting an Austin-based AI company seeking a highly capable Principal-level AI/ML Engineer to design, develop, optimize and deploy production-grade AI systems. This is a hands-on engineering position for someone who has worked directly with models - not solely integrated commercial LLM APIs. The successful candidate will take ownership across model experimentation, fine-tuning, evaluation, inference optimization and production deployment. The role operates within a technically complex B2B environment where accuracy, reliability, explainability and reproducibility are essential. Responsibilities Design, train, fine-tune and evaluate large language models and other machine-learning models. Develop domain-specific models using techniques such as LoRA, QLoRA, PEFT, quantization and model compression. Build and improve production RAG, semantic-retrieval and agentic AI systems. Create rigorous evaluation frameworks covering model accuracy, hallucinations, retrieval quality, robustness and failure modes. Optimize model inference across latency, throughput, GPU utilization, reliability and cost. Build scalable Python services and pipelines supporting model training, serving and monitoring. Deploy models into production using modern MLOps, containerization and cloud infrastructure. Investigate performance issues and make informed decisions across model quality, infrastructure requirements and commercial constraints. Work autonomously across research, engineering and deployment while collaborating with technical and business stakeholders. Contribute to the architecture and continued development of proprietary AI products and models. Core Requirements Five or more years of relevant AI/ML engineering, applied-science or model-development experience. Strong Python programming and software-engineering fundamentals. Hands-on experience developing, adapting or fine-tuning LLMs or foundation models. Production experience with PyTorch, Hugging Face Transformers or comparable frameworks. Experience building and deploying RAG, retrieval or agentic AI systems. Strong understanding of model evaluation, experimentation and performance measurement. Experience deploying and supporting models within production environments. Knowledge of MLOps, CI/CD, model monitoring, versioning and reproducible releases. Ability to independently take technically ambiguous problems from experimentation through production. Candidates with slightly fewer than five years may still be considered where they can demonstrate exceptional depth in model development, fine-tuning, evaluation or inference optimization. Applications should clearly describe the models or AI systems you have personally built, fine-tuned, evaluated, optimized or deployed. Show more Show less
Ready to apply?
Takes you directly to Stelvio Inc.'s application page
About Stelvio Inc.
Get similar jobs in your inbox
Weekly digest of AI engineering roles matched to your stack. Free forever.
Subscribe — Free