SECTION I · THE BRIEF
Brief #00617Updated 26 AUG 2026REMOTEGreenhouseSV ANGEL
Employbl Company Profile

LLM Inference Engineer

NEAR is the first AI app development platform where businesses can build their internal applications without coding. We conduct research in areas of deep learning and program synthesis in order to understand intent of…

Location
Remote
Company size
1–10
Posted
4w ago
Via
Greenhouse
Section II · Full ProfileFree with an account
  • 01Comp band & equity packageLocked
  • 02Seniority & experience requirementsLocked
  • 03Interview process & rubricLocked
  • 04Hiring manager & team contextLocked
  • 05Growth trajectory in this roleLocked
  • 06Offer & decision timelineLocked

Free account · no card · 2 minutes

LLM Inference Engineer

NEAR· San Francisco or RemoteView company profile


Job title
LLM Inference Engineer
Job location
San Francisco or Remote
Job description

Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.

View job listing ↗
The Saturday Briefing

Get the Saturday tech briefing

New company profiles, funding moves, and who’s hiring across the market — every Saturday morning.

NEAR headquarters

Gardiner, ME

Company size

1–10 employees

Founded

2017

View company profile ↗