We build reinforcement learning environments that teach models mathematics. In these environments a model attempts a proof, a formal verifier checks the reasoning, and the result becomes a reward signal that guides the next attempt. Choosing the problems, supplying the right definitions and tools, and making sure the reward reflects the mathematics we actually want learned is the work.
Mathematics is unusually valuable for this because it combines difficult reasoning with precise feedback. A verifier can tell you whether a proof establishes a formal statement. It cannot tell you whether the statement was worth proving. That judgment is the hard part, and it is where we concentrate.
Every other company supplying environments spreads across law, finance, customer support, coding, and mathematics. Advanced mathematics needs specialist judgment at nearly every stage: reading a theorem, finding its assumptions, tracing its dependencies, and deciding whether a task will teach anything at all. We do mathematics and nothing else.
Team
IMO gold and silver medalists from MIT and Princeton, mathematics PhD students at Stanford, and machine learning researchers who turn research into working systems. We have relationships with students and faculty at MIT and Princeton that connect us to the communities where this mathematics is developed.