We are a small team building environments where an agent's competence on real financial work can be measured rather than asserted. That requires people who are unusually careful about what correct means.
Every role here is in San Francisco. We work in person. We are deliberately small, and we expect people to own a surface end to end rather than a ticket queue.
How to apply
There is no form. Email matt@standardthinking.com and describe the hardest problems you have attempted and solved — what made them hard, what you tried that did not work, and how you knew when you were done. Name the role you are interested in, or don't; if the work is good we will find the right shape.
We read everything. Depth beats polish, and a specific account of one genuinely difficult problem is worth more to us than a complete résumé.
Open roles
Member of Technical Staff, Environments
Environments · Full-time · San Francisco
Build the environments themselves: the execution layer that runs a financial model exactly as the tools it was authored in would, and the task layer that turns real analyst work into something an agent can be scored on.
What we look for
- Systems engineering at the level where correctness is measured against a reference implementation, not a spec
- Comfort with calculation graphs, dependency resolution, and convergence behaviour
- A high tolerance for the edge cases that only appear at scale
Member of Technical Staff, Reinforcement Learning
Research · Full-time · San Francisco
Own post-training against our environments — reward design, credit assignment over long horizons, and the training loops that turn a verifiable task distribution into capability that transfers.
What we look for
- Hands-on RL post-training experience on large models, not just familiarity with the literature
- Judgement about which reward signals are load-bearing and which are gameable
- Interest in long-horizon, sparse-reward settings with real ground truth
Research Engineer, Evaluation
Research · Full-time · San Francisco
Decide what counts as correct. Build the graders, the tie-out checks and the contamination controls that let us claim a number and defend it.
What we look for
- A track record of building evaluations that survived contact with people trying to beat them
- Statistical care about variance, sample size and what a benchmark actually shows
- Suspicion of any metric that only ever improves
Environment Author, Finance
Environments · Full-time · San Francisco
You have built the models. Now encode that work: what the task actually is, what a correct answer looks like, where the shortcuts are, and which mistakes a first-year makes that a VP catches.
What we look for
- Two or more years in investment banking, private equity, a hedge fund or a market maker
- Fluency in the artifacts — three-statement, LBO, merger, and the diligence work around them
- Willingness to write the answer key as carefully as the question
Infrastructure Engineer
Platform · Full-time · San Francisco
Run environments at the scale training demands: isolation, determinism, and throughput measured in millions of episodes rather than requests.
What we look for
- Distributed systems and container orchestration under real load
- Instinct for where determinism leaks and how to pin it down
- Care about cost per episode as a first-class metric
Forward-Deployed Engineer
Partnerships · Full-time · San Francisco
Sit with frontier labs and enterprise partners, translate their objectives into environments and evaluations, and bring back what the work actually requires.
What we look for
- Strong engineering plus the judgement to be in the room with a customer
- Ability to move between a training run and a trading floor in the same day
- Directness about what will and will not work
Nothing here fits, but you think you should be working on this? Write to us anyway, at matt@standardthinking.com.