standard thinking

Careers

The work is the hard part.

We are a small team building environments where an agent's competence on real financial work can be measured rather than asserted. That requires people who are unusually careful about what correct means.

Every role here is in San Francisco. We work in person. We are deliberately small, and we expect people to own a surface end to end rather than a ticket queue.

How to apply

There is no form. Email matt@standardthinking.com and describe the hardest problems you have attempted and solved — what made them hard, what you tried that did not work, and how you knew when you were done. Name the role you are interested in, or don't; if the work is good we will find the right shape.

We read everything. Depth beats polish, and a specific account of one genuinely difficult problem is worth more to us than a complete résumé.

Open roles

Member of Technical Staff, Environments

Environments · Full-time · San Francisco

Build the environments themselves: the execution layer that runs a financial model exactly as the tools it was authored in would, and the task layer that turns real analyst work into something an agent can be scored on.

What we look for

  • Systems engineering at the level where correctness is measured against a reference implementation, not a spec
  • Comfort with calculation graphs, dependency resolution, and convergence behaviour
  • A high tolerance for the edge cases that only appear at scale

Member of Technical Staff, Reinforcement Learning

Research · Full-time · San Francisco

Own post-training against our environments — reward design, credit assignment over long horizons, and the training loops that turn a verifiable task distribution into capability that transfers.

What we look for

  • Hands-on RL post-training experience on large models, not just familiarity with the literature
  • Judgement about which reward signals are load-bearing and which are gameable
  • Interest in long-horizon, sparse-reward settings with real ground truth

Research Engineer, Evaluation

Research · Full-time · San Francisco

Decide what counts as correct. Build the graders, the tie-out checks and the contamination controls that let us claim a number and defend it.

What we look for

  • A track record of building evaluations that survived contact with people trying to beat them
  • Statistical care about variance, sample size and what a benchmark actually shows
  • Suspicion of any metric that only ever improves

Environment Author, Finance

Environments · Full-time · San Francisco

You have built the models. Now encode that work: what the task actually is, what a correct answer looks like, where the shortcuts are, and which mistakes a first-year makes that a VP catches.

What we look for

  • Two or more years in investment banking, private equity, a hedge fund or a market maker
  • Fluency in the artifacts — three-statement, LBO, merger, and the diligence work around them
  • Willingness to write the answer key as carefully as the question

Infrastructure Engineer

Platform · Full-time · San Francisco

Run environments at the scale training demands: isolation, determinism, and throughput measured in millions of episodes rather than requests.

What we look for

  • Distributed systems and container orchestration under real load
  • Instinct for where determinism leaks and how to pin it down
  • Care about cost per episode as a first-class metric

Forward-Deployed Engineer

Partnerships · Full-time · San Francisco

Sit with frontier labs and enterprise partners, translate their objectives into environments and evaluations, and bring back what the work actually requires.

What we look for

  • Strong engineering plus the judgement to be in the room with a customer
  • Ability to move between a training run and a trading floor in the same day
  • Directness about what will and will not work

Nothing here fits, but you think you should be working on this? Write to us anyway, at matt@standardthinking.com.