Why join Upscale AI
Upscale AI is building the high-performance infrastructure powering the next generation of artificial intelligence. Backed by over $300M in funding and rapid global adoption, we are scaling systems designed for the world’s most demanding AI workloads.
We focus on first-principles engineering across silicon, systems, and networking—where performance, scale, and execution are critical. Our team is talent-dense and high-performing. We value ownership, technical rigor, and speed, and we offer the opportunity to work on foundational problems with immediate, real-world impact.
If you’re looking to do high-impact work, move fast, and help define the infrastructure behind the future of AI—Upscale AI is where you can produce meaningful work at the frontier—and operate at a high standard.
About the role
AI training and inference clusters live or die on the network fabric. ScaleUp builds the software that programs our high-performance Ethernet switch silicon — SDK, SAI, simulation, and the CI that keeps that stack shippable.
We’re looking for a DevOps engineer at the intersection of AI infrastructure, data-center networking, and silicon-aware software: reliable pipelines for an ASIC SDK, fast feedback for developers, and release paths worthy of cloud- and AI-scale deployment.
You’ll partner with SDK, SAI, and QA engineers, and work in GitHub + Jira day to day. Success means CI is trusted, simulation and software gates are clear, and infrastructure never blocks the next generation of AI networking features.
Responsibility
Own CI/CD for the ScaleUp stack that powers AI/HPC Ethernet switching (build, unit/integration, gating, artifacts, release promotion).
Scale pipelines across multi-repo dependencies: switch SDK, SAI adapters, shared test frameworks, and switch simulation / model targets.
Improve build caching, shared runners, and nightlies so large C/C++ and Python SDK builds stay fast and predictable.
Own release / promote workflows (branch policies, artifact publish, pre-gate vs post-gate steps) so SDK drops are repeatable.
Shorten commit → green for teams building programmable data-plane, QoS, ACL, RoCE-era fabrics, and AI-workload traffic patterns.
Add and maintain quality gates: lint, unit tests, model smoke, coverage where useful, overnight soak — high signal, low flake.
Manage artifacts and caches (CI artifacts, object storage as needed) with clear retention and failure handling.
Operate Linux build/test farms for C/C++ SDKs and Python harnesses (toolchains, containers, lab/sim hosts).
Improve pipeline observability (dashboards, alerts, runbooks, postmortems).
Secure secrets and access for GitHub Actions, artifact stores, and internal services.
Partner with QA on PR / nightly / soak jobs that validate end-to-end switch behavior before customer and cloud deployments.
Publish reusable pipeline templates so new ScaleUp / AI-networking repos onboard quickly.
Jira: keep engineering work traceable — epics/stories/bugs for CI and infra, link PRs and releases to tickets, support sprint and release planning with SDK/SAI/QA, tighten ticket hygiene (status, components, labels) so blockers and CI debt are visible.
Qualification
Strong CI/CD experience (GitHub Actions and/or Jenkins; GitLab CI also fine) in multi-repo environments.
Deep Linux fluency: shells, packaging, toolchains, debugging native builds and shared-library / SDK load paths.
Automation in Python and bash; enough CMake/Make to unblock switch-SDK CI.
Containers (Docker) and self-hosted or cloud runners at meaningful scale.
Proven work reducing flaky tests and improving CI signal-to-noise.
Hands-on Jira (or equivalent): workflows, boards, linking commits/PRs to issues, release/version fields — comfortable driving process with developers, not only “keeping the lights on.”
Clear communication with hardware-adjacent software teams.
Excitement about AI data centers, Ethernet fabrics, and silicon-software co-design — not only generic cloud DevOps.
Nice to have
Background in networking ASICs, switch SDKs, SAI, DPDK, or NIC/SmartNIC software.
Exposure to AI/ML cluster networking (GPU fabrics, RoCE/RDMA, congestion control, telemetry).
Simulation / hardware-in-the-loop CI, pytest at scale, junit/Allure-style reporting.
Build-cache and large-monorepo / multi-repo release patterns; backport or branch-gating workflows.
Static analysis / lint gates in CI (e.g. cppcheck, language linters).
IaC (Terraform/Ansible) and cloud runners (AWS/GCP).
Atlassian suite beyond Jira (Confluence runbooks, Jira + GitHub automation).
Release hygiene: versioning, artifacts, SBOM/signing, branch protection / merge gates.
Why ScaleUp
Your pipelines sit under real AI networking product software — not a side CRUD service.
You influence how fast we ship switch SDK + SAI features that AI clusters depend on.
Small team, high ownership: changes land and developers feel them the next day.