Preference Model raised $16 million in seed funding led by Andreessen Horowitz, with participation from SignalFire, South Park Commons, Scale Angel Group, and researchers including Fei-Fei Li, Ian Goodfellow, and Julian Schrittwieser. Pasted text
The company focuses on building reinforcement learning environments to train more capable, better-aligned AI systems.
Preference Model said it has spent the past year developing RL environments for several frontier AI laboratories.
The company’s work centers on creating training environments that resist reward hacking, where AI models exploit weaknesses in a task or grading system instead of solving the intended problem.
Its technology helps AI labs build harder, more reliable environments as models become increasingly capable of identifying shortcuts and vulnerabilities.
Alongside the funding, Preference Model open-sourced Karotte, a framework it has used internally to develop RL environments.
Karotte has undergone more than one million evaluation runs and controlled red-teaming exercises. Pasted text
The framework provides secure defaults and tools intended to prevent agents from manipulating graders, accessing unintended information, exhausting system resources, or otherwise exploiting training infrastructure.
Preference Model views robust training environments as an increasingly important part of AI alignment, particularly because reward-hacking behavior discovered during training can become reinforced rather than merely producing inaccurate benchmark results.
Andreessen Horowitz said Preference Model is focused on an area increasingly important to AI labs: building the infrastructure needed to identify model weaknesses, generate more challenging tasks, and test environments against agents actively attempting to break them. Pasted text
The financing will support Preference Model as it continues developing reinforcement learning infrastructure and tools intended for frontier AI developers.
KEY QUOTES:
“We believe that robust RL environments are a key component to training aligned models as we head into the superintelligent era.”
“Over the past year, we’ve built RL environments for several frontier labs, and we’re backed by $16M in seed funding led by a16z, with participation from SignalFire, South Park Commons, Scale Angel Group, and researchers including Fei-Fei Li, Ian Goodfellow, and Julian Schrittwieser.”
Company statement

