Remote
Applied AI Engineer, Inference
About this role
π Description Build benchmarking workflows to measure latency, throughput, quality regressions, and cost. Benchmark inference stack against realistic workloads and baselines to identify gaps. Profile model-serving across frameworks, runtimes, and hardware to find bottlenecks. Drive optimizations for customer workloads, tuning serving configs and validating changes. Design and run experiments on model-serving techniques like quantization and caching.
Partner with inference engineers to productionize improvements and tests. π― Requirements 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work. Strong Python programming and comfort working in production engineering environments. Experience running empirical evaluations, benchmarks, or experiments and translating results into concrete engineering decisions.
Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks. Understanding of practical tradeoffs involved in latency, throughput, batching, GPU utilization, quantization, and quality regression analysis. Ability to work across model, systems, and product boundaries and stay focused on outcomes that matter for customers. π Benefits Medical, dental, and vision insurance - 100% paid for by CoreWeave Company-paid Life Insurance Voluntary supplemental life insurance Short and long-term disability insurance Flexible Spending Account Health Savings Account
Source listing: empllo_remote