Remote
Engineering Manager, GPU Infrastructure
About this role
Leading the planning and deployment of large-scale AI infrastructure environments, the full-time salaried Engineering Manager, GPU Infrastructure will oversee the integration and operational readiness of GPU clusters for AI training and high-performance computing, working remotely within the Pacific Time Zone. Key Responsibilities Manage the deployment and integration of GPU-based compute platforms and oversee end-to-end AI cluster deployment initiatives Lead operational validation of high-performance GPU interconnects and coordinate with network engineering teams for performance optimization Drive infrastructure automation initiatives for cluster provisioning and lifecycle management, establishing repeatable deployment methodologies Required Qualifications Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field (or equivalent experience) 10+ years of infrastructure engineering or datacenter deployment experience 5+ years of experience leading deployment or operations teams for large-scale AI or GPU infrastructure Hands-on experience with deploying and operating large GPU clusters in enterprise or hyperscale environments Strong expertise in Canonical MaaS, data storage platforms, InfiniBand and Ethernet GPU fabrics, and Linux systems administration
Source listing: virtualvocations_main