SECTION I · THE BRIEF
Brief #04526Updated 22 AUG 2026REMOTEYc
Employbl Company Profile

Backend Performance Engineer

LiteLLM is hiring a Backend Performance Engineer in Remote. Full brief — comp band, interview process, and hiring team — available to Employbl members.

Location
Remote
Company size
1–10
Posted
4d ago
Via
Yc
Section II · Premium ProfileMembers only
  • 01Comp band & equity packageLocked
  • 02Seniority & experience requirementsLocked
  • 03Interview process & rubricLocked
  • 04Hiring manager & team contextLocked
  • 05Growth trajectory in this roleLocked
  • 06Offer & decision timelineLocked

7-day free trial · $25/mo · cancel anytime

LiteLLM logo

Backend Performance Engineer · LiteLLM

View company profile
Job title
Backend Performance Engineer
Job location
San Francisco, CA, US / Remote (US)
Job description
### **TLDR** LiteLLM is an **open-source LLM Gateway with 28K+ stars on GitHub** and trusted by companies like **NASA, Rocket Money, Samsara, Lemonade, and Adobe.** We’re rapidly expanding and seeking a performance engineer to help scale the platform to handle 5K RPS (Requests per second). We’re based in San Francisco. ### **What is LiteLLM** LiteLLM provides an **open source Python SDK and Python FastAPI Server that allows calling 100+ LLM APIs (Bedrock, Azure, OpenAI, VertexAI, Cohere, Anthropic) in the OpenAI format** We just hit **$2.5M ARR** and have raised a **$1.6M seed round from Y Combinator, Gravity Fund and Pioneer Fund.** You can find more information on our [**website**](https://www.litellm.ai/), [**Github**](https://github.com/BerriAI/litellm) and [**Technical Documentation.**](https://docs.litellm.ai/docs/) ### About the Role We're hiring a Python performance engineer to own maximizing throughput, minimizing latency and ensuring our platform is reliable in production. **Roadmap for Performance Engineer:** * By end of this year our RPS and latency overhead should be at parity with industry benchmarks. Cover stream + non-stream for /chat/completions, /completions, /embeddings, /realtime, /audio/transcriptions * Reduce e2e overhead latency for cache misses. Currently at 100ms-500ms - ensure we meet industry standards. * Reduce e2e overhead latency for cache hits - ensure we meet industry benchmarks. * Ensure overhead latency scales well when other components are added to the platform - e.g Redis, Redis Cluster, DB, Non-Admin Virtual Keys * Ensure overhead latency scales well with payload size - 1MB prompt with streaming should be sub 100ms * Address customer specific and pipeline specific latency issues. * e.g. Enterprise customers reporting high overhead - this person should be able to debug these issues, get on support calls and help address any environment specific settings. * Address paying customer memory leaks * Enterprise clients have ongoing memory leaks that need resolution * Longer term - should add coverage over new endpoints - /realtime, /audio/transcriptions/, /audio/speech
View job listing ↗
The Saturday Briefing

Get the Saturday tech briefing

New company profiles, funding moves, and who’s hiring across the market — every Saturday morning.

LiteLLM headquarters

Company size

110 employees

Founded

2023

View company profile ↗