SECTION I · THE BRIEF

Brief #04708Updated 15 JUN 2026PALO ALTO, CAGreenhouseACCEL

Employbl Dossier

Member of Technical Staff — Training

RadixArk focuses on developing infrastructure for AI inference and training systems. It builds tools to make frontier-level AI more efficient, accessible, and cost-effective.

Location: Palo Alto, CA
Company size: 10–50
Posted: 2w ago
Via: Greenhouse

Section II · Premium ProfileMembers only

01Comp band & equity packageLocked
02Seniority & experience requirementsLocked
03Interview process & rubricLocked
04Hiring manager & team contextLocked
05Growth trajectory in this roleLocked
06Offer & decision timelineLocked

Start free trial Sign in

7-day free trial · $25/mo · cancel anytime

Member of Technical Staff — Training - RadixArk

View Company Profile

Job Title

Member of Technical Staff — Training

Job Location

Palo Alto, CA

Job Listing URL

https://job-boards.greenhouse.io/radixark/jobs/4134907009

Job Description

About the Role

RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models.

You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, accuracy and reliability across 10k, or 100k+ of GPUs. This role sits at the intersection of ML, systems, and performance engineering.

Your work will directly impact how next-generation AI models are trained and scaled.

This is a deeply technical, high-impact role for engineers who enjoy solving hard systems problems at extreme scale.

Requirements

3+ years of experience in ML systems, or large-scale training infrastructure
Experience building or operating large-scale agentic post-training systems.
Experience working on training / inference correctness or other precision-related problem
Experience debugging performance and stability issues in large post-training jobs
Experience improving training or inference efficiency.

Strong Plus

Experience training 100+ billion-parameter models
Experience with train / inference optimization for large-scale RL or other production workload.
Familiarity with training stacks (e.g. Megatron-LM, FSDP, torchtitan, etc.) and inference stack (e.g. SGLang, vLLM, etc.)
Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL, AReaL, etc.)
Experience with RDMA, InfiniBand, NVLink, NCCL/RCCL, or high-speed GPU interconnects
Contributions to ML systems open-source projects
Experience with checkpointing, fault recovery, and elastic training.
Experience building infrastructure for agentic post-training, such as async rollout pipelines, sandbox, or harness system.

Responsibilities

Contribute to open-source large-scale post-training infrastructure Miles, and inference system SGLang.
Optimize throughput, scalability, and hardware efficiency
Improve reliability and fault tolerance for long-running training jobs
Develop training frameworks and infrastructure tooling
Collaborate with model researchers to support frontier experiments
Debug and resolve cross-layer performance bottlenecks
Build observability systems for training performance and reliability
Drive capacity planning and cluster utilization strategies
Contribute to long-term training infrastructure architecture

About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework).

We're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training.

Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

We're backed by well-known infrastructure investors and partner with Nvidia, Google, AWS, and frontier AI labs.

Join us in building infrastructure that gives real leverage back to the AI community.

Compensation

We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

View Job Listing

Get the Saturday tech briefing

New company profiles, funding moves, and who’s hiring across the market — every Saturday morning.

Where this role is based

Palo Alto, CA

Loading map…

RadixArk Headquarters Location

San Francisco, CA

View company profile

RadixArk Company Size

Between 10 - 50 employees

RadixArk Founded Year

2025

RadixArk Total Amount Raised

$100,000,000

RadixArk Funding Rounds

View funding details

Seed
$100,000,000 USD
Jan 21, 2026

RadixArk's Investors

Accel

Spark Capital

RadixArk's Industries

Software Companies

Member of Technical Staff — Training

Member of Technical Staff — Training - RadixArk

About the Role

Requirements

Strong Plus

Responsibilities

About RadixArk

Compensation

Equal Opportunity

Get the Saturday tech briefing

Where this role is based

RadixArk Headquarters Location

RadixArk Company Size

RadixArk Founded Year

RadixArk Total Amount Raised

RadixArk Funding Rounds

RadixArk's Investors

RadixArk's Industries

💸 Investors

📚 Popular Blog Posts