Member of Technical Staff — Training at RadixArk
RadixArk is hiring a Member of Technical Staff — Training in Palo Alto, CA, US. On-site. Pay: USD 200k-400k/yr.
About RadixArk
RadixArk is an infrastructure-first deep-tech company that builds large-scale inference and training systems for the AI community. Its mission is to make frontier-level AI infrastructure open and accessible. The company develops SGLang, an open engine for serving AI models, and Miles, a large-scale reinforcement learning (RL) framework. RadixArk was founded by AI infrastructure engineers with experience at xAI and NVIDIA.
Member of Technical Staff — Training job description
About the Role
In This Role, You Will
- Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
- Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
- Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
- Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
- Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
- Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.
Minimum Qualifications
- 3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
- Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
- Experience in at least two of the following areas:
- Performance, efficiency, and scalability of multi-GPU, multi-node workloads
- Numerical correctness or low precision
- Stability, reliability, or fault tolerance
- Post-training algorithm recipes and orchestration infrastructure for large training runs
- Multimodal training or inference, including vision-language models and multimodal generation
- Agent infrastructure, including sandboxes, harnesses, and eval systems
- Building and maintaining open-source projects widely adopted in industry and academia
Preferred Qualifications
- Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
- Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
- Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
- GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
- Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
- Top-tier publications in ML systems or other systems fields.
Even if you don't meet every qualification above, we encourage you to apply — we care most about demonstrated ability to build and reason about large-scale systems.
About RadixArk
Compensation
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Apply now
Applications go straight to RadixArk. We never sit between you and the employer.
Apply now ↗

