Selene

The world’s most accurate LLM-as-a-Judge

Website: https://atla-ai.com

Cover Block

Field Value
Name Selene (developed by Atla AI)
Tagline "The world's most accurate LLM-as-a-Judge" [Y Combinator]
Headquarters London, United Kingdom
Founded 2023
Stage Seed
Business Model API / Developer Platform
Industry AI Evaluation / Developer Infrastructure
Technology Type AI / Machine Learning (fine-tuned LLM evaluators)
Growth Profile Venture Scale
Funding Label Seed
Total Disclosed ~$5,000,000 [Crunchbase]

Links

The Short Version

Selene, the flagship product of London-based Atla AI, is a family of fine-tuned language models purpose-built to evaluate the outputs of other language models, a category commonly called "LLM-as-a-Judge." The company was founded in 2023 by Maurice Burger and Roman Engeler and went through Y Combinator before raising a seed round of approximately $5 million led by Creandum on December 7, 2023 [Crunchbase]. Its core technical bet is that a small, specialized evaluator model can grade the quality of generative outputs more accurately than a general-purpose frontier model called via API, and the team's published benchmarks for Selene 1 Mini, a fine-tuned Llama-3.1-8B variant, claim to outperform GPT-4o on RewardBench, EvalBiasBench and Auto-J [Atla AI]. Commercially, Selene is offered via API and positioned to sit inside the development and observability loops of teams shipping LLM-powered products, putting Atla in direct competition with Patronus AI, Galileo and Braintrust.

Data Accuracy: GREEN -- Confirmed by Crunchbase, Y Combinator and Atla AI's own technical posts.

The Company in Brief

Atla AI, the legal entity behind the Selene product line, was founded in 2023 and is headquartered in London. The company entered Y Combinator and used the program's launch channel to introduce Selene as "the world's most accurate LLM-as-a-Judge," pitched as a drop-in evaluator that works with frontier models including GPT-4o and Claude 3.5 Sonnet [Y Combinator]. Cofounders Maurice Burger and Roman Engeler, the latter serving as CTO, lead a team that LinkedIn lists at roughly ten people [LinkedIn] [Y Combinator].

The company's first publicly traceable financing event is a Seed round closed on December 7, 2023 for approximately $5 million, led by Creandum with participation from Y Combinator [Crunchbase]. The most visible product milestone since then is the release of Selene 1 Mini, a fine-tuned Llama-3.1-8B model, which Atla published both as a research artifact (an arXiv preprint titled "Atla Selene Mini: A General Purpose Evaluation Model") and as a hosted API offering [arXiv / Atla] [Atla AI]. Hugging Face hosts the model weights and an accompanying blog post positioning Selene 1 Mini as "the best small language model-as-a-judge" [Hugging Face / Atla AI].

Data Accuracy: GREEN -- Confirmed by Crunchbase, Y Combinator, LinkedIn and Atla's own publications.

What They Have Built

Selene is a family of evaluator models designed to score, critique or rank the outputs of other language models against criteria such as factuality, helpfulness, safety and adherence to instructions. According to Atla, Selene 1 Mini is an 8-billion-parameter model fine-tuned from Meta's Llama-3.1-8B, and it is offered both as open weights on Hugging Face and as a hosted API [Hugging Face / Atla AI] [Atla AI]. Atla's own benchmark write-up reports that Selene 1 Mini "outperform[s] top small models including GPT-4o mini on average performance across 11 benchmarks for evaluations" [Atla AI], and a separate post claims it "beats models several times its size, outperforming GPT-4o on RewardBench, EvalBiasBench, and Auto-J" [Atla AI]. Independent commentary from Galtea, an evaluation-tooling vendor, describes Selene-1-Mini-Llama-3.1-8B as consistently outperforming other compact models on accuracy in their internal review [Galtea].

Functionally, the product slots into the developer workflow at two points. The first is offline evaluation: a team running regression tests on prompt or model changes can call Selene to score outputs and detect quality drift before shipping. The second is online evaluation: a team running an LLM in production can route a sample of traffic through Selene to monitor for hallucinations, bias or policy violations. The Y Combinator launch page emphasizes drop-in compatibility with major model providers, framing Selene as model-agnostic infrastructure [Y Combinator].

Data Accuracy: YELLOW -- Model claims are sourced primarily from Atla's own publications and the arXiv preprint, with partial third-party corroboration from Galtea.

Market Research and Opportunity

LLM evaluation has gone from a research afterthought to a budgeted line item in roughly eighteen months, and Selene is one of a small number of companies built specifically to own that line item. As enterprises move generative AI features from prototype into customer-facing production, the absence of reliable, automated quality measurement has become the limiting factor on shipping velocity. Human review does not scale, and using a frontier model as the judge is expensive and, according to Atla's published benchmarks, less accurate than a purpose-built evaluator on tasks such as RewardBench and Auto-J [Atla AI] [arXiv / Atla].

As an analogous frame of reference, the broader AI observability and MLOps category has attracted material venture capital into companies such as Galileo and Patronus AI, both directly named as Selene competitors, suggesting institutional investors are already underwriting the thesis that evaluation is a standalone budget. Demand drivers surfaced in the cited research include the rapid proliferation of agentic and multi-step LLM applications, the emergence of model-routing and model-switching strategies that require a neutral evaluator, and growing regulatory attention on AI safety and bias.

Metric Value
Atla Seed round ~$5,000,000
Selene 1 Mini parameter count 8 billion
Public benchmarks on which Selene 1 Mini reportedly beats GPT-4o 3 (RewardBench, EvalBiasBench, Auto-J)

Data Accuracy: YELLOW -- Demand drivers and competitor identities are corroborated by multiple sources.

Who Else Is Fighting for This

Selene is positioned as the accuracy-first specialist in a young category where the alternatives are either broader observability platforms or the brute-force option of using a frontier model as a judge.

Company Positioning Stage / Funding Notable Differentiator
Selene (Atla AI) Specialized small-model LLM judge, API + open weights Seed, ~$5M Claims SOTA accuracy at 8B parameters, beating GPT-4o on three named benchmarks
Patronus AI LLM evaluation and guardrails platform Venture-backed Enterprise-oriented evaluation suite with focus on safety and compliance
Galileo AI evaluation and observability for generative apps Venture-backed Broader observability surface bundled with evaluation
Braintrust Eval and experimentation platform for LLM developers Venture-backed Developer-workflow-first, strong on prompt iteration

Data Accuracy: YELLOW -- Competitor identities corroborated by Y Combinator competitive set.

Opportunity

If Atla executes, Selene becomes the default evaluator that sits between every production LLM application and the model serving it. The single largest outcome Selene could plausibly become is the standard third-party judge layer for generative AI in production. The cited evidence makes this reachable rather than aspirational for two reasons. First, the technical artifact already exists in a form that customers can audit: an 8-billion-parameter open-weights model with a published preprint and benchmark wins on RewardBench, EvalBiasBench and Auto-J [Atla AI] [arXiv / Atla]. Second, the category is being underwritten by serious capital across multiple competitors (Patronus AI, Galileo, Braintrust), which is the market signaling that evaluation is a standalone budget line.

Data Accuracy: YELLOW -- Scenarios are constructed from cited product evidence and named competitors.

Sources

  1. [Y Combinator] Launch YC: Selene - The World's Most Accurate LLM-as-a-Judge | https://www.ycombinator.com/launches/Mu5-selene-the-world-s-most-accurate-llm-as-a-judge
  2. [Y Combinator] Atla: The improvement engine for AI agents | https://www.ycombinator.com/companies/atla
  3. [Crunchbase] atla - Crunchbase Company Profile & Funding | https://www.crunchbase.com/organization/atla-037f
  4. [Crunchbase] Seed Round - atla - 2023-12-07 | https://www.crunchbase.com/funding_round/atla-037f-seed--a00aef2a
  5. [LinkedIn] atla company page | https://www.linkedin.com/company/atla-ai
  6. [Atla AI] Selene Mini: SOTA 8B LLM Judge, now available via API | https://atla-ai.com/post/selene-mini-api
  7. [Atla AI] Selene 1 Mini: the best small language model-as-a-judge | https://www.atla-ai.com/post/selene-1-mini
  8. [Hugging Face / Atla AI] Selene 1 Mini: the best small language model-as-a-judge | https://huggingface.co/blog/AtlaAI/selene-1-mini
  9. [arXiv / Atla] Atla Selene Mini: A General Purpose Evaluation Model | https://arxiv.org/html/2501.17195v1
  10. [Galtea] Exploring state-of-the-art LLMs as Judges | https://galtea.ai/blog/exploring-state-of-the-art-llms-as-judges
  11. [Fondo] Atla Launches Selene: The World's Most Accurate LLM-as-a-Judge | https://www.tryfondo.com/blog/atla-launches-selene
  12. [Toolify] Selene 1 Alternatives in 2026 | https://www.toolify.ai/alternative/selene-1

Articles about Selene

View on Startuply.vc