Remote
AI Inference Engineer
About this role
To define and build the inference serving strategy at scale, the founding AI Inference Engineer will work remotely, focusing on designing the serving stack and optimizing model-level strategies while collaborating closely with CUDA and GPU engineering teams. Key responsibilities Define the inference serving strategy and architecture from first principles Design and build the serving stack for high-throughput, latency-sensitive inference workloads Own the model-level optimisation strategy, partnering with CUDA/GPU engineers for integration Required qualifications 4+ years of experience building or operating large-scale inference serving systems Deep, hands-on experience with inference serving frameworks and optimisation techniques Strong systems thinking with the ability to reason about the full request-response path Proven track record of making high-stakes architecture calls Comfort operating without a playbook in an early-stage function
Source listing: virtualvocations_main