Cluster Design

External job listingat Blue Signal Search

Cluster DesignLocation: On-site, San Francisco, CA A rapidly growing AI infrastructure company is seeking a highly technical infrastructure expert to lead the design, deployment, and opti...

External source - not verified1 hour agoOpen until: Dec 29, 2026

Salary

Not provided

Location

San Francisco, United States

Employment type

Not provided

Workplace

Not provided

Cluster Design

San Francisco, United States

Job description

Cluster DesignLocation: On-site, San Francisco, CA A rapidly growing AI infrastructure company is seeking a highly technical infrastructure expert to lead the design, deployment, and optimization of large-scale GPU computing environments. This is an opportunity to have a direct impact on the architecture powering advanced AI workloads while working alongside an experienced engineering team building from the ground up. If you thrive on solving complex infrastructure challenges, enjoy working directly with cutting-edge hardware, and prefer remaining deeply technical rather than moving into management, this role offers exceptional ownership and influence. What You'll DoArchitect and deploy high performance GPU compute clusters supporting production AI and machine learning workloads.Design, configure, and optimize high speed networking environments using InfiniBand and or RoCEv2 technologies.Drive performance tuning across the complete GPU software stack, including CUDA, ROCm, drivers, firmware, NCCL, RCCL, and system benchmarking.Troubleshoot infrastructure, networking, and software bottlenecks affecting large scale distributed compute environments.Develop scalable infrastructure standards, deployment methodologies, and operational best practices.Collaborate directly with customers to understand workload requirements and deliver optimized infrastructure solutions.Lead technical investigations during production incidents and provide expert guidance through issue resolution.Create operational documentation, implementation procedures, and performance validation processes for future deployments. Required QualificationsProven hands-on experience designing, deploying, and operating production-scale GPU clusters.Deep expertise implementing InfiniBand and or RoCEv2 networking fabrics.Strong experience working with CUDA and or ROCm environments, including firmware management, driver integration, NCCL or RCCL optimization, and performance analysis.Demonstrated success supporting customer facing production infrastructure.Excellent troubleshooting skills spanning compute, networking, storage, and distributed systems.Strong communication skills with the ability to explain complex technical concepts to both engineering teams and customers.Preference for remaining an individual contributor focused on technical excellence rather than pursuing people management responsibilities.Preferred ExperienceBackground supporting AI, HPC, or accelerated computing environments.Experience evaluating system performance through benchmarking and workload optimization.Familiarity with distributed infrastructure operations and production support methodologies.Passion for building reliable, scalable infrastructure that directly impacts customer success. About Blue Signal:Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services.

Is this your job posting?

Claim it for free and receive video applications on CazVid.

Similar jobs