Gen AI Inferencing Engineer
Skills
About this role
Look for: an ML platform/infra engineer, not primarily an app developer.
search for "vLLM," "Triton," "model serving," "MLOps" in resumes
• Background running models in production at scale — deployment via vLLM or Triton Inference Server, containerized (Docker/K8s), with real throughput/latency tuning experience
• Strong MLOps chops: CI/CD for ML pipelines, fine-tuning workflows, inference framework internals
• Comfortable owning infrastructure other data science teams build on top of (shared tooling, not one-off notebooks)
• RAG knowledge is a plus but secondary — they should be stronger on "how do I serve this efficiently" than "how do I build the retrieval logic"
• Good fit: someone from an ML platform, SRE-for-ML, or MLOps background who's touched GenAI serving specifically (not just traditional ML model serving)