Performance Engineer (GPU)
Anthropic
- As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models
- You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency
- Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures
- Your work will span the entire stack—from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization
- Co-design attention mechanisms and algorithms for next-generation hardware architectures
- Develop custom kernels for emerging quantization formats and mixed-precision techniques
- Design distributed communication strategies for multi-node GPU clusters
- Optimize end-to-end training and inference pipelines for frontier language models
- Build performance modeling frameworks to predict and optimize GPU utilization
- Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
- Create resilient systems for planet-scale distributed training infrastructure
- Profile and eliminate performance bottlenecks in production serving infrastructure
- Partner with hardware vendors to influence future accelerator capabilities and software stacks
Benefits
- Comprehensive health, dental, and vision insurance for you and your dependents
- Inclusive fertility benefits via Carrot Fertility
- 22 weeks of paid parental leave
- Flexible paid time off and absence policies
- Mental health support for you and your dependents
- Competitive salary and equity packages
- Optional equity donation matching at a 1:1 ratio, up to 25% of your equity grant
- Retirement plans with competitive matching
- Life and income protection plans
- $500/month flexible wellness and time saver stipend
- Commuter benefits
- Annual education stipend
- Home office stipends
- Relocation support for those moving for Anthropic
- Daily meals and snacks in the office
Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineersHave deep experience with GPU programming and optimization at scaleCare about the societal impacts of your workCan navigate complex systems from hardware interfaces to high-level ML frameworksAre impact-driven, passionate about delivering measurable performance breakthroughsEnjoy collaborative problem-solving and pair programmingThrive in ambiguous environments where you define the path forwardWant to work on state-of-the-art language models with real-world impactEducation requirements: We require at least a Bachelor's degree in a related field or equivalent experienceGPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimizationML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operatorsPerformance Engineering: Kernel fusion, memory bandwidth optimization, profiling with NsightDistributed Systems: NCCL, NVLink, collective communication, model parallelismLow-Precision: INT8/FP8 quantization, mixed-precision techniquesProduction Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration
#J-18808-LjbffrReference: WJ-766_22188279