Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands.
We're building at the absolute frontier of AI-powered video creation and next-generation creative tools. Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like.
Optimize the training clusters: distributed training at scale - NCCL tuning, InfiniBand/RoCE fabric health, topology-aware scheduling and gang placement, GPU/network throughput, fast checkpointing, job preemption and recovery. Make every training run use the hardware it paid for.
Own Talos / Sidero Omni cluster lifecycleacross the GPU fleet: node bootstrap and upgrades, GPU drivers / NVIDIA GPU Operator / DCGM on an immutable OS, zero-downtime rollouts.
Operate the multi-provider GPU fleet: capacity planning across Nebius regions and bare-metal RTX Pro pools, hardware incident escalation to providers, node lifecycle (NotReady triage, XID errors, driver upgrades).
Own inference autoscaling: KEDA-driven, Kafka-queue-based scaling of GPU consumers; GPU-aware scheduling; warm pools and cold-start reduction; supply/demand tuning of our in-house autoscaler (higgscaler).
GitOps everything: ArgoCD multi-cluster (10+ clusters from one repo), Helm, Terraform (HCP). No hand-labeled nodes, no console drift - if it's not in git, it doesn't exist.
Observability & SLOs: VictoriaMetrics/Logs/Traces, Prometheus, DCGM exporters; honest dashboards for utilization, training throughput, queue latency, cost per generation.
GPU efficiency as a discipline: hunt idle allocations, capacity/demand mismatches, starved queues - our AIOps platform (Mycelium) files these findings automatically; you close the loop with real fixes.
Partner with ML engineers on training runs and model-serving rollouts (runtimes, batching, memory sizing) and with the core team on AWS EKS (Karpenter, Istio, Bottlerocket, gVisor sandboxes).
Our stack
Talos Linux, EKS, NVIDIA GPU Operator, DCGM, CUDA, NCCL, InfiniBand/RoCE · ArgoCD, Helm, Terraform Cloud · KEDA, Karpenter, Kafka · VictoriaMetrics/Logs/Traces, Prometheus, Grafana · Istio, Cloudflare · Python/Go.
You have
3+ years running production Kubernetes as SRE/Platform/MLOps, includingGPU workloads.
Hands-on distributed training operations: NCCL, high-speed interconnects (InfiniBand/RoCE), multi-node job scheduling, checkpointing strategies - and the habit of measuring throughput before and after every change.
Bare-metal Kubernetes experience: Talos or similar immutable-OS setups; node lifecycle without a cloud safety net.
The NVIDIA stack: drivers, container toolkit, GPU Operator, DCGM metrics; you can debug "GPU visible but not allocatable" at 3am.
GitOps fluency (ArgoCD/Flux + Helm) and Terraform; strong Linux and networking (multi-cluster, VPN/TGW topologies).
Queue-based autoscaling (KEDA/HPA) and enough Kafka to reason about consumer lag.
Python or Go for automation; you write things down.