Company
Pramaana Labs is a frontier AI lab building the verification layer for AI. Founded in 2025 and headquartered in Palo Alto, the company turns complex human knowledge, including tax codes, clinical guidelines, legal rules, and safety constraints, into machine-checkable logic, so every answer can be traced, challenged, and proved. Pramaana was founded by a team out of Google, DeepMind, and Glean, with researchers and engineers from CMU, Berkeley, Meta, Stanford and Google.
Pramaana's architecture pairs foundation models trained to formalize and reason with a symbolic world model encoded in the Lean proof language. Regulatory, statutory, policy, and scientific text is converted into formal representations, and the system returns proof artifacts that domain experts can inspect. Accountability is built into the architecture rather than added afterward through policy or prompting.
The role
You'll be a Research Scientist on the AI Team - the group responsible for the models underneath our prover and solver This is a hands-on research role: you'll own experiments end-to-end, from the idea through the training run, the debugging, the eval, and the honest write-up of what did and didn't work.
What you'll actually work on
Post-training and RL for proof search. Designing reward functions using the Lean4 compiler, and training models for theorem proving and autoformalisation.
Distillation and capability transfer. Closing the gap between a large teacher and a model small so as to run a search loop economically and in a time efficient manner.
Autoformalization. Getting models to translate natural-language statements into formal representations that typecheck and mean what the source meant.
Inference-time search: In domains where rules evolve, such as tax and law, provers need to search and adapt at inference time. Building efficient inference-time search infrastructure is therefore critical to building fast, capable provers.
What we're looking for
We hire for depth over pedigree. We care about what you've actually built and understood, judged on the work itself - not where you did it or how many things you've touched.
Technical depth. You understand your own work from first principles. You know why you made the choices you made, not just what you did and that holds up under questioning rather than getting vague.
You build and ship. You can implement, train, debug, and scale your own ideas end-to-end, including the infra reality of pre-training and RL loops. You don't need someone else to make the idea real.
Rigor. You care instinctively about sound claims, verification, and reproducibility. We come from a world where correctness is the whole game; overclaiming is disqualifying here in a way it isn't everywhere.
Agency. You drive toward outcomes and follow through. You don't wait to be told the next step, and you don't let things quietly stall.
Perseverance. You've gone deep on something over a long stretch — and stayed with it after the first three ideas failed.
Alignment. You want to work on this problem. We'd rather hire someone convinced this is the most important thing they could be doing than someone excellent who's shopping.