About Us
Are you ready to build the future of supply chain? At Gather AI, we're not just creating software; we're pioneering a new era of warehouse intelligence. We've developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining "on-time, in full" delivery.
If you're looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We're leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time.
About the Team
You’ll be embedded within our Engineering organization, working alongside the full stack and WMS integration teams. This group is responsible for the production systems that power our warehouse intelligence platform - the APIs, data pipelines, and application layers that customers and internal teams depend on every day. The team is moving from growth-era foundations toward enterprise-grade reliability, and is looking for a technical anchor to lead that transition.
About the Role
We are seeking a Principal Engineer (Distributed Systems), Full Stack to own technical clarity across our full stack and WMS integration teams. You will stabilize and de-risk our data and platform layer by eliminating scalability bottlenecks, retiring compliance and reliability debt, and establishing production-grade observability, database governance, and performance standards.
This is an architectural leadership role - not a DevOps or feature-delivery position. You will shift the organization from reactive firefighting to predictable, durable execution, setting technical direction that allows the team to build faster and safer as we scale.
What You’ll Do
- Set technical direction across backend, frontend, platform, and infrastructure - aligning teams on a coherent, enterprise-grade architecture and holding those boundaries without becoming a bottleneck
- Eliminate scalability and reliability risks in production, including Postgres upgrades, data-lifecycle controls, latency-critical API optimization, and reducing single-threaded dependencies
- Establish and operate production-grade observability, runbooks, RCA processes, and incident-response practices that replace firefighting with predictable, repeatable execution
- Drive the platform toward multi-regional architecture - including a read-write replica strategy - and evolve the data model and schema to support AI/ML workloads at scale
- Mentor senior engineers across the full stack and WMS integration teams; define shared terminology, coding standards, and automation frameworks that raise the bar for the whole organization
- Collaborate cross-functionally with Product, DevOps, QA, and Autonomy/ML teams to ensure platform primitives and data contracts support evolving product and intelligence workloads
What You’ll Need
- 10+ years of hands-on engineering experience, with demonstrated ownership of complex, stateful production systems at scale - where reliability, performance, and correctness materially impacted the business
- Deep, production-scale expertise with PostgreSQL (upgrades, replication tradeoffs, query tuning, schema evolution) and strong SQL fundamentals; hands-on experience with Node.js and/or Python
- Cloud-native fluency on at least one major provider (Azure preferred), plus working experience with Kubernetes and Docker for running and scaling production workloads
- Production-grade observability (logs, metrics, traces), CI/CD, and SRE fundamentals (SLIs, SLOs, incident response), paired with strong opinions on testing and validation as system design concerns
- Architectural leadership experience: the ability to set technical direction, align cross-functional teams, and evolve mature systems - with precise written communication and documentation skills
Nice to Have
- Domain experience in logistics, warehouse systems, robotics-adjacent platforms, or similar physical-world problem spaces
- Experience with multi-regional architectures, advanced replication strategies, or formalized data-governance programs
- Familiarity with supporting AI/ML workloads at the infrastructure and data layer
- Prior experience driving organization-wide mentorship, engineering standards, or terminology alignment initiatives