Skip to main contentSkip to navigationSkip to search

Staff Software Engineer, Training (Bay Area / Paris / Remote)

Genesis AI
Bay Area
Principal
Remote
Posted 11 days ago
Expires Dec 8, 2025

AI Technologies

PyTorch
ML

About the Role

WHAT YOU'LL DO - Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels - Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization - Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks - Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking - Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures WHAT YOU'LL BRING - Deep experience in distributed systems, ML infrastruc

Requirements

    About Genesis AI

    No company description available.