Preference Model Opens Karotte to Build Better AI Training Environments

The Andreessen Horowitz-backed startup builds reinforcement learning environments to train smarter, more aligned frontier models.

About Preference Model

Published

Preference Model launched on October 7 with a $16 million seed check from Andreessen Horowitz [Pomegra, October 2026]. The San Francisco startup’s product is not another model or API. It sells reinforcement learning environments and reward systems to frontier AI labs [Preference Model, October 2026]. The bet is that the quality of training data, specifically the feedback loops for complex tasks, is the next critical bottleneck for superintelligent AI.

The Infrastructure for Smarter Feedback

Most AI training today relies on static datasets. Preference Model argues that for models to perform advanced research or engineering, they need dynamic environments where they can attempt tasks, receive nuanced feedback, and iterate [Simplify Jobs, October 2026]. The company builds these simulated grounds. Its tooling is designed to identify where a model is weak, generate new tasks to target those gaps, and crucially, test the environments against agents actively trying to cheat the system [Andreessen Horowitz, October 2026]. The goal is a clean, learnable signal that pushes models toward genuine capability and alignment, not just reward hacking.

An Open-Source Wedge Called Karotte

Alongside its launch, the company open-sourced Karotte, a framework for building robust RL environments under an MIT license [DataPile, October 2026]. Karotte is engineered to resist reward hacking by design, running each task in an isolated VM or container, scoring every step with a judge, and strictly capping the student model’s resources [ccleaks.com, October 2026]. This move serves a dual purpose. It establishes technical credibility in a research-heavy field and creates a potential on-ramp for its commercial offerings. The team reports it has already built proprietary environments for several unnamed frontier labs during its stealth period [Andreessen Horowitz, October 2026].

The Team and the Check

Founders Jennifer Zhou and Ning Cao anchor the venture. Zhou, the CEO, cut her teeth on Anthropic’s data team, building infrastructure and datasets, after a stint at Stripe [Preference Model, October 2026]. Cao, leading strategy, was an early employee at DatologyAI, helping scale that company from inception [Preference Model, October 2026]. Their backgrounds in AI data operations and company-building resonated with a high-profile investor syndicate.

The $16 million seed was led by Andreessen Horowitz and included South Park Commons, Scale Angel Group, Manifund, and MoE Capital [Preference Model, October 2026]. A roster of individual angels reads like a who’s who of AI research: Fei-Fei Li, Ian Goodfellow, and Julian Schrittwieser were among those writing checks [Preference Model, October 2026]. The round’s size, for a seed-stage infrastructure play, signals a conviction that this layer is both critical and underserved.

Founder Role Prior Experience
Jennifer Zhou Co-founder & CEO Anthropic (data team), Stripe [Preference Model, October 2026]
Ning Cao Co-founder DatologyAI (early employee) [Preference Model, October 2026]

Where the Model Could Break

For all its promise, Preference Model’s path is lined with execution risks familiar to any infrastructure startup selling to a nascent, concentrated buyer base.

  • Customer concentration. The target market is frontier AI labs, a small club of well-funded but notoriously demanding clients. Landing one is a feat; building a diversified, resilient revenue stream across them is another challenge entirely.
  • Technical moat. The open-source release of Karotte invites scrutiny and competition. While it showcases capability, the company must prove its proprietary environments offer significantly more value to justify enterprise contracts.
  • Market timing. The startup is betting that RL training for research tasks becomes a standard practice. If the industry pivots toward alternative training paradigms or finds workarounds, the need for specialized environments could diminish.

The company is hiring aggressively across research engineering and technical staff roles, indicating a push to scale its environment-building capacity [Ashby, October 2026]. The next twelve months will test whether its environments can become a must-have tool for labs racing toward the next generation of AI. With $16 million from Andreessen Horowitz and a bench of expert angels, the question for labs is clear: in the race for smarter models, is your training data keeping up?

Sources

  1. [Preference Model, October 2026] Company Website | https://www.preferencemodel.com/
  2. [Andreessen Horowitz, October 2026] Investing in Preference Model | https://a16z.com/announcement/investing-in-preference-model/
  3. [Pomegra, October 2026] Preference Model raises $16M seed from a16z in 2026 | https://pomegra.io/startups/preference-model-raises-16m-seed-from-a16z-in-2026-2026-10-08
  4. [Simplify Jobs, October 2026] Member of Technical Staff @ Preference Model | https://simplify.jobs/p/940ae594-bef3-4efe-b0c9-92a53c14e274/Member-of-Technical-Staff
  5. [DataPile, October 2026] Preference Model raises $16.0M Seed | https://datapile.co/funding-news/preference-model-18606
  6. [ccleaks.com, October 2026] Article on Karotte | https://ccleaks.com
  7. [Ashby, October 2026] Preference Model Careers Page | https://jobs.ashbyhq.com/preferencemodel

Read on Startuply.vc