Selene
The world’s most accurate LLM-as-a-Judge
Website: https://atla-ai.com
Cover Block
| Field | Value |
|---|---|
| Name | Selene (developed by Atla AI) |
| Tagline | "The world's most accurate LLM-as-a-Judge" [Y Combinator] |
| Headquarters | London, United Kingdom |
| Founded | 2023 |
| Stage | Seed |
| Business Model | API / Developer Platform |
| Industry | AI Evaluation / Developer Infrastructure |
| Technology Type | AI / Machine Learning (fine-tuned LLM evaluators) |
| Growth Profile | Venture Scale |
| Funding Label | Seed |
| Total Disclosed | ~$5,000,000 [Crunchbase] |
Links
- Website: https://www.atla-ai.com/
- LinkedIn: https://www.linkedin.com/company/atla-ai
- Y Combinator profile: https://www.ycombinator.com/companies/atla
- Hugging Face (Selene 1 Mini): https://huggingface.co/blog/AtlaAI/selene-1-mini
- Crunchbase: https://www.crunchbase.com/organization/atla-037f
The Short Version
Selene, the flagship product of London-based Atla AI, is a family of fine-tuned language models purpose-built to evaluate the outputs of other language models, a category commonly called "LLM-as-a-Judge." The company was founded in 2023 by Maurice Burger and Roman Engeler and went through Y Combinator before raising a seed round of approximately $5 million led by Creandum on December 7, 2023 [Crunchbase]. Its core technical bet is that a small, specialized evaluator model can grade the quality of generative outputs more accurately than a general-purpose frontier model called via API, and the team's published benchmarks for Selene 1 Mini, a fine-tuned Llama-3.1-8B variant, claim to outperform GPT-4o on RewardBench, EvalBiasBench and Auto-J [Atla AI]. Commercially, Selene is offered via API and positioned to sit inside the development and observability loops of teams shipping LLM-powered products, putting Atla in direct competition with Patronus AI, Galileo and Braintrust.
Data Accuracy: GREEN -- Confirmed by Crunchbase, Y Combinator and Atla AI's own technical posts.
The Company in Brief
Atla AI, the legal entity behind the Selene product line, was founded in 2023 and is headquartered in London. The company entered Y Combinator and used the program's launch channel to introduce Selene as "the world's most accurate LLM-as-a-Judge," pitched as a drop-in evaluator that works with frontier models including GPT-4o and Claude 3.5 Sonnet [Y Combinator]. Cofounders Maurice Burger and Roman Engeler, the latter serving as CTO, lead a team that LinkedIn lists at roughly ten people [LinkedIn] [Y Combinator].
The company's first publicly traceable financing event is a Seed round closed on December 7, 2023 for approximately $5 million, led by Creandum with participation from Y Combinator [Crunchbase]. The most visible product milestone since then is the release of Selene 1 Mini, a fine-tuned Llama-3.1-8B model, which Atla published both as a research artifact (an arXiv preprint titled "Atla Selene Mini: A General Purpose Evaluation Model") and as a hosted API offering [arXiv / Atla] [Atla AI]. Hugging Face hosts the model weights and an accompanying blog post positioning Selene 1 Mini as "the best small language model-as-a-judge" [Hugging Face / Atla AI].
Data Accuracy: GREEN -- Confirmed by Crunchbase, Y Combinator, LinkedIn and Atla's own publications.
What They Have Built
Selene is a family of evaluator models designed to score, critique or rank the outputs of other language models against criteria such as factuality, helpfulness, safety and adherence to instructions. According to Atla, Selene 1 Mini is an 8-billion-parameter model fine-tuned from Meta's Llama-3.1-8B, and it is offered both as open weights on Hugging Face and as a hosted API [Hugging Face / Atla AI] [Atla AI]. Atla's own benchmark write-up reports that Selene 1 Mini "outperform[s] top small models including GPT-4o mini on average performance across 11 benchmarks for evaluations" [Atla AI], and a separate post claims it "beats models several times its size, outperforming GPT-4o on RewardBench, EvalBiasBench, and Auto-J" [Atla AI]. Independent commentary from Galtea, an evaluation-tooling vendor, describes Selene-1-Mini-Llama-3.1-8B as consistently outperforming other compact models on accuracy in their internal review [Galtea].
Functionally, the product slots into the developer workflow at two points. The first is offline evaluation: a team running regression tests on prompt or model changes can call Selene to score outputs and detect quality drift before shipping. The second is online evaluation: a team running an LLM in production can route a sample of traffic through Selene to monitor for hallucinations, bias or policy violations. The Y Combinator launch page emphasizes drop-in compatibility with major model providers, framing Selene as model-agnostic infrastructure [Y Combinator].
Data Accuracy: YELLOW -- Model claims are sourced primarily from Atla's own publications and the arXiv preprint, with partial third-party corroboration from Galtea.
Market Research and Opportunity
LLM evaluation has gone from a research afterthought to a budgeted line item in roughly eighteen months, and Selene is one of a small number of companies built specifically to own that line item. As enterprises move generative AI features from prototype into customer-facing production, the absence of reliable, automated quality measurement has become the limiting factor on shipping velocity. Human review does not scale, and using a frontier model as the judge is expensive and, according to Atla's published benchmarks, less accurate than a purpose-built evaluator on tasks such as RewardBench and Auto-J [Atla AI] [arXiv / Atla].
As an analogous frame of reference, the broader AI observability and MLOps category has attracted material venture capital into companies such as Galileo and Patronus AI, both directly named as Selene competitors, suggesting institutional investors are already underwriting the thesis that evaluation is a standalone budget. Demand drivers surfaced in the cited research include the rapid proliferation of agentic and multi-step LLM applications, the emergence of model-routing and model-switching strategies that require a neutral evaluator, and growing regulatory attention on AI safety and bias.
| Metric | Value |
|---|---|
| Atla Seed round | ~$5,000,000 |
| Selene 1 Mini parameter count | 8 billion |
| Public benchmarks on which Selene 1 Mini reportedly beats GPT-4o | 3 (RewardBench, EvalBiasBench, Auto-J) |
Data Accuracy: YELLOW -- Demand drivers and competitor identities are corroborated by multiple sources.
Who Else Is Fighting for This
Selene is positioned as the accuracy-first specialist in a young category where the alternatives are either broader observability platforms or the brute-force option of using a frontier model as a judge.
| Company | Positioning | Stage / Funding | Notable Differentiator |
|---|---|---|---|
| Selene (Atla AI) | Specialized small-model LLM judge, API + open weights | Seed, ~$5M | Claims SOTA accuracy at 8B parameters, beating GPT-4o on three named benchmarks |
| Patronus AI | LLM evaluation and guardrails platform | Venture-backed | Enterprise-oriented evaluation suite with focus on safety and compliance |
| Galileo | AI evaluation and observability for generative apps | Venture-backed | Broader observability surface bundled with evaluation |
| Braintrust | Eval and experimentation platform for LLM developers | Venture-backed | Developer-workflow-first, strong on prompt iteration |
Data Accuracy: YELLOW -- Competitor identities corroborated by Y Combinator competitive set.
Opportunity
If Atla executes, Selene becomes the default evaluator that sits between every production LLM application and the model serving it. The single largest outcome Selene could plausibly become is the standard third-party judge layer for generative AI in production. The cited evidence makes this reachable rather than aspirational for two reasons. First, the technical artifact already exists in a form that customers can audit: an 8-billion-parameter open-weights model with a published preprint and benchmark wins on RewardBench, EvalBiasBench and Auto-J [Atla AI] [arXiv / Atla]. Second, the category is being underwritten by serious capital across multiple competitors (Patronus AI, Galileo, Braintrust), which is the market signaling that evaluation is a standalone budget line.
Data Accuracy: YELLOW -- Scenarios are constructed from cited product evidence and named competitors.
Sources
- [Y Combinator] Launch YC: Selene - The World's Most Accurate LLM-as-a-Judge | https://www.ycombinator.com/launches/Mu5-selene-the-world-s-most-accurate-llm-as-a-judge
- [Y Combinator] Atla: The improvement engine for AI agents | https://www.ycombinator.com/companies/atla
- [Crunchbase] atla - Crunchbase Company Profile & Funding | https://www.crunchbase.com/organization/atla-037f
- [Crunchbase] Seed Round - atla - 2023-12-07 | https://www.crunchbase.com/funding_round/atla-037f-seed--a00aef2a
- [LinkedIn] atla company page | https://www.linkedin.com/company/atla-ai
- [Atla AI] Selene Mini: SOTA 8B LLM Judge, now available via API | https://atla-ai.com/post/selene-mini-api
- [Atla AI] Selene 1 Mini: the best small language model-as-a-judge | https://www.atla-ai.com/post/selene-1-mini
- [Hugging Face / Atla AI] Selene 1 Mini: the best small language model-as-a-judge | https://huggingface.co/blog/AtlaAI/selene-1-mini
- [arXiv / Atla] Atla Selene Mini: A General Purpose Evaluation Model | https://arxiv.org/html/2501.17195v1
- [Galtea] Exploring state-of-the-art LLMs as Judges | https://galtea.ai/blog/exploring-state-of-the-art-llms-as-judges
- [Fondo] Atla Launches Selene: The World's Most Accurate LLM-as-a-Judge | https://www.tryfondo.com/blog/atla-launches-selene
- [Toolify] Selene 1 Alternatives in 2026 | https://www.toolify.ai/alternative/selene-1
Articles about Selene
- Selene Is Betting an 8-Billion-Parameter Judge Can Grade GPT-4o's Homework — The London YC startup, backed by Creandum's $5M seed, wants to be the scoring layer every AI agent runs through before shipping.