Preference Model
Develops reinforcement learning environments and reward systems for training frontier AI models.
Website: https://www.preferencemodel.com/
Cover Block
From the public record
| Field | Value |
|---|---|
| Name | Preference Model |
| Tagline | Develops reinforcement learning environments and reward systems for training frontier AI models. |
| Headquarters | San Francisco, United States [Preference Model, October 2026] |
| Founded | 2024 [Crunchbase] |
| Stage | Seed [Pomegra, October 2026] |
| Business model | B2B |
| Industry | Deeptech |
| Technology | AI / Machine Learning |
| Geography | North America |
| Growth profile | Venture Scale |
| Founding team | Co-Founders (2), Jennifer Zhou and Ning Cao [Preference Model, October 2026] |
| Funding label | Seed |
| Total disclosed funding | ~$16,000,000 [Pomegra, October 2026] |
Links
From the public record
- Website: https://www.preferencemodel.com/
The Short Version
PUBLIC Preference Model builds reinforcement learning environments and reward systems for frontier AI model training, and it merits attention now because it emerged from stealth alongside a $16 million seed round led by Andreessen Horowitz on October 7, 2026, with a product thesis tied to one of the harder bottlenecks in post-training, namely generating clean, failure-resistant feedback loops for advanced models [Andreessen Horowitz, October 2026] [Preference Model, October 2026] [Pomegra, October 2026]. The company is based in San Francisco and was founded in 2024 by Jennifer Zhou and Ning Cao, according to public company materials and profile data, though the public record is still thin in the way it often is for newly launched technical infrastructure companies [Preference Model, October 2026] [Crunchbase].
The product claim is fairly specific. Preference Model says it develops environments where models perform research and engineering tasks, receive realistic feedback, and improve through repeated rollouts, while a16z adds that the tooling is meant to identify model weaknesses, generate targeted tasks, and test whether agents can exploit the environment itself [Simplify Jobs, October 2026] [Andreessen Horowitz, October 2026]. That framing matters because the differentiation appears to rest less on model architecture and more on the training substrate: the company also open-sourced Karotte, described as a framework for building RL environments intended to resist reward hacking, with supporting detail available through its GitHub presence and technical materials [DataPile, October 2026] [GitHub] [Preference Model, Karotte Technical Deep Dive].
The founding team reads as relevant to the problem set, even if much of the detail remains company-supplied. Preference Model says Zhou previously worked on Anthropic's data team and earlier at Stripe, while public profile evidence supports her prior Anthropic tenure and current role at Preference Model; LinkedIn also supports Cao's co-founder status and prior time as an early employee at DatologyAI [Preference Model, October 2026] [Yutori, May 2026] [LinkedIn]. For an investor, the encouraging signal is domain adjacency: both founders appear to come from data, infrastructure, or AI workflow settings rather than from a purely financial or academic angle [Preference Model, October 2026] [LinkedIn] [RocketReach].
The business model is B2B, and the likely buyer set is frontier AI labs, which aligns with the company's statement that it has built RL environments for leading or several frontier labs, although no customers have been named publicly and no revenue metrics are disclosed [Preference Model, October 2026] [Andreessen Horowitz, October 2026] [Pomegra, October 2026]. The financing is straightforward on the surface: multiple public sources point to a $16 million seed round in October 2026, led by Andreessen Horowitz, with additional institutional and angel backers listed by the company [Pomegra, October 2026] [startups.gallery, October 2026] [Preference Model, October 2026].
Over the next 12 to 18 months, the main questions are commercial proof and product durability. Public hiring activity suggests the company is still building out core research and engineering capacity, and the real test will be whether it can convert early work with frontier labs into repeatable deployments, defensible tooling, and evidence that its environments remain useful as model capabilities improve [Ashby, October 2026] [Wesleyan Career Center] [Andreessen Horowitz, October 2026].
Single-source, plausible -- Core funding and launch facts are corroborated by Andreessen Horowitz, Preference Model, and Pomegra, but several team and product details remain company-supplied or partially corroborated.
Taxonomy Snapshot
| Axis | Value |
|---|---|
| Stage | Seed |
| Business Model | B2B |
| Industry / Vertical | Deeptech |
| Technology Type | AI / Machine Learning |
| Geography | North America |
| Growth Profile | Venture Scale |
| Founding Team | Co-Founders (2) |
| Funding | Seed, total disclosed approximately $16,000,000 |
The Company in Brief
PUBLIC
Preference Model surfaced publicly in October 2026 with a narrowly defined pitch: building reinforcement learning environments and reward systems for training frontier AI models [Preference Model, October 2026]. The company is based in San Francisco and is listed on Crunchbase as founded in 2024, with Jennifer Zhou and Ning Cao identified as co-founders [Crunchbase] [Preference Model, October 2026].
The public milestone sequence is short but clear. Crunchbase places the company’s formation in 2024, while the company website shows it operating by October 2026 with a San Francisco presence, named founders, and a disclosed investor roster that includes Andreessen Horowitz, South Park Commons, Scale Angel Group, Manifund, MoE Capital, and several individual backers [Crunchbase] [Preference Model, October 2026]. On the same public launch window, the company also presented Karotte as an open-source framework for RL environments built to reduce reward hacking risk [Preference Model, October 2026].
The legal entity is not established in the cited public record used for this section. What is established is a young San Francisco startup, founded in 2024 and publicly launched in 2026, positioning itself as infrastructure for frontier-model training rather than a general-purpose application layer [Crunchbase] [Preference Model, October 2026].
Single-source, plausible -- Core company identity is corroborated by Crunchbase and the company website, but the legal entity is not confirmed in the cited sources.
What They Have Built
MIXED
The product story is narrow but technically legible. Preference Model says it builds reinforcement learning environments and reward systems for training more capable, better-aligned frontier AI models, with a particular focus on ML research and engineering tasks [Preference Model, October 2026]. Public descriptions from Andreessen Horowitz are directionally consistent: the company develops tooling to identify model weaknesses, generate tasks that target those gaps, and test environments against agents trying to exploit or break them [Andreessen Horowitz, October 2026]. That matters because the claimed differentiation is not a general model layer, but the environment design and evaluation loop around post-training.
The most concrete public artifact is Karotte, an open-source framework that Preference Model released on October 7, 2026 under the MIT license, according to contemporaneous coverage [ccleaks.com, October 2026]. GitHub and the company’s technical write-up describe Karotte as a framework for building RL environments that resist reward hacking, with the company arguing that the system prevents whole classes of exploits by construction [GitHub]; [Preference Model, Karotte Technical Deep Dive]. A separate technical description says Karotte runs tasks inside a VM or container, scores each step with a judge, and constrains memory, processes, and files, which gives outside readers a clearer picture of how Preference Model is thinking about evaluation integrity, even if the claim is not independently benchmarked against named alternatives [ccleaks.com, October 2026].
The stack itself is only partly visible in public. Job postings and recruiting pages indicate active hiring around research, post-training, RL environments, and automated ML research engineering, which supports the view that the company is building infrastructure close to the training loop rather than a consumer-facing application [Ashby, October 2026]; [Simplify Jobs, October 2026]; [Wesleyan Career Center]. Any deeper read on orchestration, model serving, or deployment architecture would be speculative from current materials, aside from the containerized or VM-based environment approach described for Karotte (inferred from job postings) [Ashby, October 2026]; [ccleaks.com, October 2026].
Single-source, plausible -- Core product claims are corroborated across company materials, a16z, recruiting pages, and GitHub, but several technical implementation details rely on single-source or company-adjacent reporting.
Market Size and Demand
PUBLIC
The market matters now because frontier-model developers are shifting from pretraining scale to post-training performance, and that raises the value of environments, evaluators, and reward systems that can produce reliable feedback on harder tasks [Andreessen Horowitz, October 2026] [Preference Model, October 2026].
Public evidence here is thin on formal market sizing, so the market has to be framed through adjacent categories rather than a clean, company-specific TAM. Preference Model is positioned around reinforcement learning environments for AI research and engineering workflows, with stated buyers in frontier AI labs rather than broad enterprise teams [Preference Model, October 2026] [Simplify Jobs, October 2026]. That places it at the intersection of model evaluation infrastructure, synthetic or generated training data, and tooling for post-training and alignment, but none of the captured sources provide a named third-party TAM, SAM, or SOM for that exact wedge [Andreessen Horowitz, October 2026] [Pomegra, October 2026].
The demand signal that does show up consistently is qualitative: as models become more capable, labs need harder tasks, cleaner reward signals, and environments that are resistant to reward hacking rather than easy to exploit [Andreessen Horowitz, October 2026] [GitHub]. Preference Model's open-source Karotte framing is useful as a market clue because it points to a specific pain point, namely building RL environments that do not collapse when agents find shortcuts in the reward function [Preference Model, Karotte Technical Deep Dive] [DataPile, October 2026]. If that problem persists across labs, the spend pool is less likely to look like generic annotation and more likely to resemble specialized infrastructure for evaluation, environment design, and automated task generation [Andreessen Horowitz, October 2026].
Adjacent markets are easier to identify than the core market itself. The nearest substitutes appear to be internal tooling at major AI labs, broader model evaluation stacks, synthetic-data pipelines, and human-feedback or labeling vendors extending into RL and post-training workflows [Simplify Jobs, October 2026] [Andreessen Horowitz, October 2026]. The practical question is whether buyers treat environment construction as a strategic internal capability or as infrastructure they are willing to source from a specialist vendor. The available public record does not resolve that yet, but Preference Model's stated work with several frontier labs suggests there is at least some willingness to externalize part of the stack [a16z.news] [Pomegra, October 2026].
Macro and regulatory forces cut both ways. On the supportive side, greater scrutiny of model reliability and safety should increase interest in testbeds that surface failure modes before deployment, especially in research and engineering settings where agents can exploit weak reward functions [Andreessen Horowitz, October 2026] [Preference Model, Karotte Technical Deep Dive]. On the constraining side, this is still a narrow buyer universe with concentrated budget authority inside a small number of advanced-model labs, which can slow vendor formation even when technical need is real [Pomegra, October 2026] [Preference Model, October 2026]. That combination usually produces a market that can matter strategically before it looks large in conventional software-category terms.
| Market lens | What public sources support | Evidence |
|---|---|---|
| Core wedge | RL environments and reward systems for frontier-model training | Preference Model describes itself as building RL environments for capable, aligned models [Preference Model, October 2026] |
| Buyer set | Frontier AI labs | Recruiting and investor materials point to frontier labs as the target customer base [Simplify Jobs, October 2026] [Andreessen Horowitz, October 2026] |
| Demand driver | Need to identify model weaknesses and create harder tasks | a16z describes tooling for finding weaknesses, generating tasks, and testing against exploitative agents [Andreessen Horowitz, October 2026] |
| Adjacent categories | Evaluation infrastructure, synthetic training data, alignment tooling | Inference from the company's described workflow and product scope, grounded in cited product descriptions [Preference Model, October 2026] [DataPile, October 2026] |
The table shows why this is best read as an emerging infrastructure niche rather than a mature software category. Public sources support the problem definition and buyer profile more clearly than they support market size, which means market conviction still rests on whether post-training and evaluation budgets consolidate into a durable vendor category.
Single-source, plausible -- Section relies on company materials and a16z's investment note, with partial corroboration from recruiting and funding coverage; no independent third-party market report or exact TAM data was identified.
Who Else Is Fighting for This
Competitive Map
MIXED Preference Model is positioning itself upstream of application-layer AI tooling, selling the environments and reward design used to train frontier models rather than the models or end-user software themselves [Preference Model, October 2026] [Andreessen Horowitz, October 2026].
That creates a competitive set that is easier to describe by segment than by a clean list of named peers, because the source set does not identify direct startup competitors by name. On one side sit internal teams at frontier labs, which can build proprietary post-training environments in-house if they believe environment quality is strategic and if they have the engineering depth to maintain them [Andreessen Horowitz, October 2026]. On another sit adjacent infrastructure vendors that support model development workflows more broadly, even if the public record here does not show they offer Preference Model's specific combination of RL environments, reward design, and reward-hacking resistance [Built In] [Simplify Jobs, October 2026]. A third substitute is open-source tooling, most concretely Karotte itself, which Preference Model released under the MIT license on October 7, 2026, potentially lowering barriers for labs to replicate parts of the workflow without buying a full-stack external solution [ccleaks.com, October 2026] [GitHub].
The most credible edge visible in public materials is talent adjacency to frontier-model training and a narrow product focus on environment hardness. Jennifer Zhou's prior work at Anthropic on data infrastructure, tokenizers, and datasets is company-stated and partially corroborated by an external profile showing she was at Anthropic from August 2022 to November 2024, while Ning Cao's DatologyAI background is corroborated by LinkedIn [Preference Model, October 2026] [Yutori, May 2026] [LinkedIn]. Andreessen Horowitz's investment note also points to tooling for finding model weaknesses, generating targeted tasks, and testing environments against agents trying to break them, which suggests the company is not just packaging annotation or eval software but trying to own a more difficult layer of post-training infrastructure [Andreessen Horowitz, October 2026]. That edge may be durable if frontier labs prefer specialist vendors with scarce alignment and RL-systems talent, but it is also perishable because the same customers are sophisticated enough to internalize the workflow once patterns standardize.
The main exposure is distribution and customer concentration risk, even if the public record stops short of naming accounts. Preference Model says it has built RL environments for several frontier labs, but has not disclosed customer names, deployment scale, or revenue, which makes it hard to judge whether it is becoming embedded infrastructure or still operating as a high-end research contractor [a16z.news] [Pomegra, October 2026]. The open-source release of Karotte cuts both ways: it can establish technical credibility and a de facto standard, but it can also give larger labs and adjacent tooling vendors a starting point to absorb parts of the stack without depending on Preference Model as a vendor [DataPile, October 2026] [GitHub].
Over the next 18 months, the most plausible competitive scenario is not a broad land grab across enterprise AI, but a narrower race to become the preferred external environment layer for a small number of frontier labs. Preference Model is the clearest public winner if those labs continue to outsource specialized RL-environment construction and value a vendor focused on reward-hacking resistance and ML research tasks [Andreessen Horowitz, October 2026] [Preference Model, Karotte Technical Deep Dive]. The more likely loser if that assumption breaks is not a named startup in this source set, but the external-vendor category itself: if leading labs decide environment design is too core to training performance and safety to leave outside, in-house teams become the effective winner and specialist providers face a tighter lane [Andreessen Horowitz, October 2026].
Single-source, plausible -- Core positioning and product claims are corroborated by company materials and Andreessen Horowitz, but the section cannot name direct competitors from the provided public source set and several competitive inferences rest on absence of disclosed customer and revenue data.
Opportunity
PUBLIC The prize here is unusually large: if reinforcement learning environments become a gating layer for frontier model progress, Preference Model could end up supplying a core part of the training stack to the labs spending the most aggressively on post-training and alignment infrastructure [Andreessen Horowitz, October 2026] [Preference Model, October 2026].
The headline opportunity is not generic AI tooling. It is the chance to become a specialized infrastructure provider for a narrow but high-value bottleneck: creating environments and reward systems that let frontier models practice research and engineering work, receive realistic feedback, and improve through repeated rollouts [Simplify Jobs, October 2026]. That outcome looks reachable, not merely aspirational, because the public record already shows three useful signals in the same direction: the company says it has built RL environments for leading or several frontier labs, Andreessen Horowitz framed the product as tooling that finds model weaknesses and stress-tests environments against exploitation, and the team open-sourced Karotte, a framework intended to make RL environments more resistant to reward hacking [a16z.news] [Pomegra, October 2026] [Andreessen Horowitz, October 2026] [GitHub] [DataPile, October 2026]. In plain terms, the company appears to be working on a real training problem that matters more as models get better at gaming simplistic evaluations.
A conservative upside case is that Preference Model becomes the outside supplier labs use when they need harder environments faster than they can build them internally. A more expansive one is that it becomes the default framework layer for environment design, testing, and reward integrity across research organizations, combining proprietary services with an open-source standard in Karotte [GitHub] [Preference Model, Karotte Technical Deep Dive]. That second path is harder, but the open-source release gives the company at least a visible wedge into developer workflows rather than relying only on closed enterprise selling [ccleaks.com, October 2026] [DataPile, October 2026].
| Scenario | What happens | Catalyst | Why it's plausible |
|---|---|---|---|
| Frontier-lab infrastructure vendor | Preference Model becomes a repeat supplier of RL environments and reward systems to a small set of top AI labs, with revenue concentrated in high-value technical engagements | One or two named lab relationships or deeper public endorsements from existing backers and customers | The company and a16z both say it has already built environments for frontier labs, and the product is described in terms of concrete research workflows rather than broad automation claims [Preference Model, October 2026] [Andreessen Horowitz, October 2026] [Pomegra, October 2026] |
| Open-source standard to enterprise conversion | Karotte becomes a widely adopted framework for building harder-to-game RL environments, then pulls paid demand for hosted tooling, testing, or custom environment design | Sustained GitHub adoption, external contributors, or visible integrations around Karotte | The company has already open-sourced Karotte under the MIT license, and public descriptions position it around a specific pain point, reward hacking resistance, that should matter across advanced-agent training efforts [ccleaks.com, October 2026] [GitHub] [Preference Model, Karotte Technical Deep Dive] |
| Post-training control plane | Preference Model expands from environment creation into a broader layer for identifying model weaknesses, generating targeted tasks, and adversarially testing agents before deployment | Product broadening from environment design into continuous eval and red-team loops | Andreessen Horowitz's investment note describes exactly that workflow, which suggests the initial product surface may already sit adjacent to a larger control-plane opportunity in frontier model development [Andreessen Horowitz, October 2026] |
The compounding logic is straightforward if execution holds. Every environment built for a demanding lab should teach the company where frontier models fail, which tasks produce clean learning signals, and which reward schemes break under pressure; those learnings can improve future environments, make Karotte more useful, and strengthen the company's ability to identify and patch reward-hacking pathways [Andreessen Horowitz, October 2026] [Preference Model, Karotte Technical Deep Dive] [GitHub]. The public hiring footprint also points in that direction: roles spanning research, post-training, research engineering, and RL environment review suggest the company is building a repeatable production system around environment creation rather than a one-off consulting effort [Ashby, October 2026] [Simplify Jobs, October 2026] [Wesleyan Career Center].
The size of the win is easiest to frame at the category level rather than through current company metrics, which are still sparse. If Preference Model were to become a meaningful infrastructure layer inside frontier model training, the relevant value creation could resemble other picks-and-shovels companies that sit near model development rather than end-user applications; in that scenario, a multi-billion-dollar outcome is conceivable (scenario, not a forecast), because the buyers are few but well-funded, the technical switching costs could be high, and the product appears tied to a mission-critical training function [Andreessen Horowitz, October 2026] [Pomegra, October 2026]. That remains conditional on public proof of customer depth, but the early ingredients, technical specificity, an a16z-led $16 million seed, named investors with deep AI credibility, and an open-source wedge, are enough to treat the upside as concrete rather than theoretical [Pomegra, October 2026] [Preference Model, October 2026] [Wilson Sonsini].
Single-source, plausible -- Section relies on one independent lead investor source, one funding/news aggregator, open-source repository evidence, and company materials; major upside claims remain analytical and customer names are not publicly confirmed.
Sources
From the public record
[Preference Model, October 2026] Preference Model | https://www.preferencemodel.com/
[Pomegra, October 2026] Preference Model raises $16M seed from a16z in 2026 | https://pomegra.io/startups/preference-model-raises-16m-seed-from-a16z-in-2026-2026-10-08
[Andreessen Horowitz, October 2026] Investing in Preference Model | https://a16z.com/announcement/investing-in-preference-model/
[Simplify Jobs, October 2026] Member of Technical Staff @ Preference Model | https://simplify.jobs/p/940ae594-bef3-4efe-b0c9-92a53c14e274/Member-of-Technical-Staff
[DataPile, October 2026] Preference Model raises $16.0M Seed | https://datapile.co/funding-news/preference-model-18606
Articles about Preference Model
- Preference Model Opens Karotte to Build Better AI Training Environments — The Andreessen Horowitz-backed startup builds reinforcement learning environments to train smarter, more aligned frontier models.