StarVLA

An open-source, modular codebase for Vision-Language-Action (VLA) model development in embodied AI.

Website: https://starvla.github.io/

Cover Block

Name StarVLA
Tagline An open-source, modular codebase for Vision-Language-Action (VLA) model development in embodied AI.
Stage Other
Business Model Open Source / Commercial
Industry Deeptech
Technology AI / Machine Learning
Geography Global / Remote-First
Growth Profile Other
Founding Team Academic Spinout

Data Accuracy: GREEN -- Confirmed by project homepage and repository.

What an Investor Needs First

StarVLA is an open-source research codebase that standardizes the development of Vision-Language-Action models for embodied AI [arXiv, April 2026]. The project's value proposition is a modular, Lego-like framework designed to unify training, evaluation, and deployment workflows [starVLA, retrieved 2026]. Originating from a collaboration between researchers at the Hong Kong University of Science and Technology (HKUST) and the open-source community, it operates as an academic spinout [36Kr, May 2026]. Its core differentiation lies in providing a unified, reproducible platform where researchers can plug in different model backbones and action heads, a capability demonstrated by benchmark results showing a 98.8% average success rate on standardized LIBERO tasks [arXiv, April 2026]. The project's traction is measured by its 1.5k GitHub stars and active maintenance by a community of contributors, including key individuals like Jinhui Ye [Jinhui Ye's Homepage, retrieved 2026].

Data Accuracy: YELLOW -- Core product claims and benchmark results are documented in arXiv preprints and the project's homepage; team and funding details are less directly verified.

Taxonomy Snapshot

Axis Classification
Stage Other
Business Model Open Source / Commercial
Industry / Vertical Deeptech
Technology Type AI / Machine Learning
Geography Global / Remote-First
Growth Profile Other
Founding Team Academic Spinout

Inside the Company

StarVLA presents as a community-driven, open-source research project. The project's public identity is anchored to its technical contributions and academic affiliations, with no formal company formation or headquarters disclosed. The most concrete organizational attribution points to the Hong Kong University of Science and Technology (HKUST) and a broader open-source community team [36Kr, May 2026].

Key development milestones are documented through research papers and code releases. The project was formally introduced with the April 2026 publication of the "StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing" paper on arXiv [arXiv, April 2026]. This was followed by a technical follow-up, "StarVLA-α: Reducing Complexity in Vision-Language-Action Systems," which reported benchmark results including a 98.8% average success rate on the LIBERO benchmark [arXiv, April 2026]. Public coverage in May 2026 from outlets like 36Kr and QQ News framed the release as a significant open-source contribution aimed at unifying VLA research paradigms [36Kr, May 2026][QQ News, May 2026].

Individual contributors are identifiable. Jinhui Ye is listed as a main code contributor [Jinhui Ye's Homepage]. The project's Zenodo record also lists multiple authors with HKUST and Microsoft Open Source affiliations [Zenodo, January 2026]. The project remains actively maintained, with a training efficiency report released in January 2026 and ongoing updates to its GitHub repository, which has accumulated over 1.5k stars [GitHub cheng-haha/starvla_m, retrieved 2026][starVLA, retrieved 2026].

Data Accuracy: YELLOW -- Key attributions are confirmed by multiple press reports, but foundational corporate details are absent from public records.

Under the Hood

StarVLA's core offering is a modular, open-source codebase designed to accelerate research into Vision-Language-Action models. The project's primary value is a unified framework that standardizes the process of developing and benchmarking VLA systems [arXiv, April 2026].

The architecture is designed for component swapping, described as a "Lego-like" system where researchers can mix and match backbone vision-language models, action heads, datasets, and training recipes [starVLA, retrieved 2026]. This modularity is operationalized through several pre-configured variants: StarVLA-FAST, StarVLA-OFT, StarVLA-PI, and StarVLA-GR00T [starVLA, retrieved 2026]. The framework supports reproducible evaluation across major embodied AI benchmarks, including LIBERO, RoboCasa, and BEHAVIOR, and uses a unified WebSocket interface to bridge policy control from simulation to real robots [starVLA, retrieved 2026].

Performance claims center on benchmark results published in accompanying research papers. The StarVLA-α variant is reported to achieve a 98.8% average success rate on the LIBERO benchmark and up to 53.8% success on more challenging dual-arm and humanoid tasks [arXiv, April 2026]. A separate analysis notes that reinforcement learning techniques within the framework can push simpler imitation learning models from approximately 70% to over 98% success on the same benchmark [EmergentMind, retrieved 2026]. The project's GitHub repository shows over 2.5k stars and 174 forks [starVLA, retrieved 2026] [Changsheng Lu's Homepage, retrieved 2026].

Data Accuracy: YELLOW -- Technical specifications and benchmark results are sourced from the project's own documentation and arXiv pre-prints.

Market Research

The market for embodied AI development tools is gaining momentum as robotics research shifts from isolated prototypes to scalable, language-driven systems. The global market for AI in robotics was valued at $14.6 billion in 2024 and is projected to reach $84.1 billion by 2032 [Precedence Research, 2024]. The market for AI software development platforms is forecast to reach $115 billion by 2027 [Gartner, 2024].

Metric Value
AI in Robotics (2024) $14.6B
AI in Robotics (2032 est.) $84.1B
AI Software Dev Platforms (2027 est.) $115B

Data Accuracy: YELLOW -- Market sizing is drawn from third-party analyst reports for analogous sectors; direct TAM for VLA frameworks is not yet published.

Competition and Substitutes

StarVLA enters the embodied AI research ecosystem as an open-source integrator and benchmark platform. A comparison of key open-source frameworks in the VLA research space shows the project's current position.

Company Positioning Notable Differentiator
StarVLA Modular, Lego-like codebase unifying training, inference, and benchmarks. Unified interface for plug-and-play model swapping and reproducible benchmarking.
OpenVLA-OFT Open-source implementation of a VLA using the OFT method. Focuses on a specific, efficient fine-tuning technique.
π₀ (PiZero) Generalist embodied AI model from Google DeepMind. Represents the state-of-the-art frontier in generalist agent capabilities.
GR00T NVIDIA's foundation model project for humanoid robotics. Backed by proprietary compute and robotics simulation assets.
InternVLA-M1 VLA model from Shanghai AI Laboratory. Often used as a performance benchmark or backbone.

Data Accuracy: YELLOW -- Competitor analysis is based on project documentation and technical literature.

Opportunity

The opportunity for StarVLA is to become the foundational software layer for developing embodied AI. The project's modular architecture addresses the engineering overhead required to reproduce, compare, and extend complex VLA models [arXiv, April 2026]. By unifying training recipes, evaluation benchmarks, and deployment interfaces, StarVLA offers a path to becoming the default starting point for academic labs and corporate R&D teams.

Data Accuracy: YELLOW -- Opportunity scenarios are extrapolated from the project's technical design and early traction.

Sources

  1. [arXiv, April 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://arxiv.org/abs/2604.05014
  2. [starVLA, retrieved 2026] starVLA | Agile Lego-like Embodied AI Development | https://starvla.github.io/
  3. [36Kr, May 2026] Unify VLA Paradigm: HKUST Open-Sources StarVLA Lego-Style Architecture, Drastically Reducing Reproduction Cost | https://eu.36kr.com/en/p/3764865889125128
  4. [Jinhui Ye's Homepage, retrieved 2026] Jinhui Ye's Homepage | https://jinhuiye.github.io/
  5. [arXiv, April 2026] StarVLA-α: Reducing Complexity in Vision-Language-Action Systems | https://arxiv.org/abs/2604.11757
  6. [EmergentMind, retrieved 2026] StarVLA: Modular Codebase for Vision-Language-Action Models | https://www.emergentmind.com/news/2024-10-14-starvla-modular-codebase-for-vision-language-action-models
  7. [Changsheng Lu's Homepage, retrieved 2026] Changsheng Lu's Homepage | 卢长胜的主页 | https://alanlusun.github.io/
  8. [GitHub cheng-haha/starvla_m, retrieved 2026] StarVLA Training Efficiency Report & Training Curves | https://github.com/cheng-haha/starvla_m
  9. [Zenodo, January 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://zenodo.org/records/18264214
  10. [QQ News, May 2026] VLA的PyTorch时刻已至!港科大联手社区开源StarVLA | https://news.qq.com/rain/a/20260509A03AI500?suid=&media_id=
  11. [Papers.cool, retrieved 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://papers.cool/arxiv/2604.05014
  12. [Precedence Research, 2024] AI in Robotics Market Size, Share, Growth Report 2032 | https://www.precedenceresearch.com/ai-in-robotics-market
  13. [Gartner, 2024] Gartner Forecasts Worldwide Artificial Intelligence Software Market to Reach $115 Billion by 2027 | https://www.gartner.com/en/newsroom/press-releases/2024-07-10-gartner-forecasts-worldwide-artificial-intelligence-software-market-to-reach-115-billion-by-2027

Articles about StarVLA

View on Startuply.vc