StarVLA
An open-source, modular codebase for Vision-Language-Action (VLA) model development in embodied AI.
Website: https://starvla.github.io/
Cover Block
| Name | StarVLA |
| Tagline | An open-source, modular codebase for Vision-Language-Action (VLA) model development in embodied AI. |
| Stage | Other |
| Business Model | Open Source / Commercial |
| Industry | Deeptech |
| Technology | AI / Machine Learning |
| Geography | Global / Remote-First |
| Growth Profile | Other |
| Founding Team | Academic Spinout |
- Website: https://starvla.github.io/
- GitHub: https://github.com/StarVLA
Data Accuracy: GREEN -- Confirmed by project homepage and repository.
What an Investor Needs First
StarVLA is an open-source research codebase that standardizes the development of Vision-Language-Action models for embodied AI [arXiv, April 2026]. The project's value proposition is a modular, Lego-like framework designed to unify training, evaluation, and deployment workflows [starVLA, retrieved 2026]. Originating from a collaboration between researchers at the Hong Kong University of Science and Technology (HKUST) and the open-source community, it operates as an academic spinout [36Kr, May 2026]. Its core differentiation lies in providing a unified, reproducible platform where researchers can plug in different model backbones and action heads, a capability demonstrated by benchmark results showing a 98.8% average success rate on standardized LIBERO tasks [arXiv, April 2026]. The project's traction is measured by its 1.5k GitHub stars and active maintenance by a community of contributors, including key individuals like Jinhui Ye [Jinhui Ye's Homepage, retrieved 2026].
Data Accuracy: YELLOW -- Core product claims and benchmark results are documented in arXiv preprints and the project's homepage; team and funding details are less directly verified.
Taxonomy Snapshot
| Axis | Classification |
|---|---|
| Stage | Other |
| Business Model | Open Source / Commercial |
| Industry / Vertical | Deeptech |
| Technology Type | AI / Machine Learning |
| Geography | Global / Remote-First |
| Growth Profile | Other |
| Founding Team | Academic Spinout |
Inside the Company
StarVLA presents as a community-driven, open-source research project. The project's public identity is anchored to its technical contributions and academic affiliations, with no formal company formation or headquarters disclosed. The most concrete organizational attribution points to the Hong Kong University of Science and Technology (HKUST) and a broader open-source community team [36Kr, May 2026].
Key development milestones are documented through research papers and code releases. The project was formally introduced with the April 2026 publication of the "StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing" paper on arXiv [arXiv, April 2026]. This was followed by a technical follow-up, "StarVLA-α: Reducing Complexity in Vision-Language-Action Systems," which reported benchmark results including a 98.8% average success rate on the LIBERO benchmark [arXiv, April 2026]. Public coverage in May 2026 from outlets like 36Kr and QQ News framed the release as a significant open-source contribution aimed at unifying VLA research paradigms [36Kr, May 2026][QQ News, May 2026].
Individual contributors are identifiable. Jinhui Ye is listed as a main code contributor [Jinhui Ye's Homepage]. The project's Zenodo record also lists multiple authors with HKUST and Microsoft Open Source affiliations [Zenodo, January 2026]. The project remains actively maintained, with a training efficiency report released in January 2026 and ongoing updates to its GitHub repository, which has accumulated over 1.5k stars [GitHub cheng-haha/starvla_m, retrieved 2026][starVLA, retrieved 2026].
Data Accuracy: YELLOW -- Key attributions are confirmed by multiple press reports, but foundational corporate details are absent from public records.
Under the Hood
StarVLA's core offering is a modular, open-source codebase designed to accelerate research into Vision-Language-Action models. The project's primary value is a unified framework that standardizes the process of developing and benchmarking VLA systems [arXiv, April 2026].
The architecture is designed for component swapping, described as a "Lego-like" system where researchers can mix and match backbone vision-language models, action heads, datasets, and training recipes [starVLA, retrieved 2026]. This modularity is operationalized through several pre-configured variants: StarVLA-FAST, StarVLA-OFT, StarVLA-PI, and StarVLA-GR00T [starVLA, retrieved 2026]. The framework supports reproducible evaluation across major embodied AI benchmarks, including LIBERO, RoboCasa, and BEHAVIOR, and uses a unified WebSocket interface to bridge policy control from simulation to real robots [starVLA, retrieved 2026].
Performance claims center on benchmark results published in accompanying research papers. The StarVLA-α variant is reported to achieve a 98.8% average success rate on the LIBERO benchmark and up to 53.8% success on more challenging dual-arm and humanoid tasks [arXiv, April 2026]. A separate analysis notes that reinforcement learning techniques within the framework can push simpler imitation learning models from approximately 70% to over 98% success on the same benchmark [EmergentMind, retrieved 2026]. The project's GitHub repository shows over 2.5k stars and 174 forks [starVLA, retrieved 2026] [Changsheng Lu's Homepage, retrieved 2026].
Data Accuracy: YELLOW -- Technical specifications and benchmark results are sourced from the project's own documentation and arXiv pre-prints.
Market Research
The market for embodied AI development tools is gaining momentum as robotics research shifts from isolated prototypes to scalable, language-driven systems. The global market for AI in robotics was valued at $14.6 billion in 2024 and is projected to reach $84.1 billion by 2032 [Precedence Research, 2024]. The market for AI software development platforms is forecast to reach $115 billion by 2027 [Gartner, 2024].
| Metric | Value |
|---|---|
| AI in Robotics (2024) | $14.6B |
| AI in Robotics (2032 est.) | $84.1B |
| AI Software Dev Platforms (2027 est.) | $115B |
Data Accuracy: YELLOW -- Market sizing is drawn from third-party analyst reports for analogous sectors; direct TAM for VLA frameworks is not yet published.
Competition and Substitutes
StarVLA enters the embodied AI research ecosystem as an open-source integrator and benchmark platform. A comparison of key open-source frameworks in the VLA research space shows the project's current position.
| Company | Positioning | Notable Differentiator |
|---|---|---|
| StarVLA | Modular, Lego-like codebase unifying training, inference, and benchmarks. | Unified interface for plug-and-play model swapping and reproducible benchmarking. |
| OpenVLA-OFT | Open-source implementation of a VLA using the OFT method. | Focuses on a specific, efficient fine-tuning technique. |
| π₀ (PiZero) | Generalist embodied AI model from Google DeepMind. | Represents the state-of-the-art frontier in generalist agent capabilities. |
| GR00T | NVIDIA's foundation model project for humanoid robotics. | Backed by proprietary compute and robotics simulation assets. |
| InternVLA-M1 | VLA model from Shanghai AI Laboratory. | Often used as a performance benchmark or backbone. |
Data Accuracy: YELLOW -- Competitor analysis is based on project documentation and technical literature.
Opportunity
The opportunity for StarVLA is to become the foundational software layer for developing embodied AI. The project's modular architecture addresses the engineering overhead required to reproduce, compare, and extend complex VLA models [arXiv, April 2026]. By unifying training recipes, evaluation benchmarks, and deployment interfaces, StarVLA offers a path to becoming the default starting point for academic labs and corporate R&D teams.
Data Accuracy: YELLOW -- Opportunity scenarios are extrapolated from the project's technical design and early traction.
Sources
- [arXiv, April 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://arxiv.org/abs/2604.05014
- [starVLA, retrieved 2026] starVLA | Agile Lego-like Embodied AI Development | https://starvla.github.io/
- [36Kr, May 2026] Unify VLA Paradigm: HKUST Open-Sources StarVLA Lego-Style Architecture, Drastically Reducing Reproduction Cost | https://eu.36kr.com/en/p/3764865889125128
- [Jinhui Ye's Homepage, retrieved 2026] Jinhui Ye's Homepage | https://jinhuiye.github.io/
- [arXiv, April 2026] StarVLA-α: Reducing Complexity in Vision-Language-Action Systems | https://arxiv.org/abs/2604.11757
- [EmergentMind, retrieved 2026] StarVLA: Modular Codebase for Vision-Language-Action Models | https://www.emergentmind.com/news/2024-10-14-starvla-modular-codebase-for-vision-language-action-models
- [Changsheng Lu's Homepage, retrieved 2026] Changsheng Lu's Homepage | 卢长胜的主页 | https://alanlusun.github.io/
- [GitHub cheng-haha/starvla_m, retrieved 2026] StarVLA Training Efficiency Report & Training Curves | https://github.com/cheng-haha/starvla_m
- [Zenodo, January 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://zenodo.org/records/18264214
- [QQ News, May 2026] VLA的PyTorch时刻已至!港科大联手社区开源StarVLA | https://news.qq.com/rain/a/20260509A03AI500?suid=&media_id=
- [Papers.cool, retrieved 2026] StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing | https://papers.cool/arxiv/2604.05014
- [Precedence Research, 2024] AI in Robotics Market Size, Share, Growth Report 2032 | https://www.precedenceresearch.com/ai-in-robotics-market
- [Gartner, 2024] Gartner Forecasts Worldwide Artificial Intelligence Software Market to Reach $115 Billion by 2027 | https://www.gartner.com/en/newsroom/press-releases/2024-07-10-gartner-forecasts-worldwide-artificial-intelligence-software-market-to-reach-115-billion-by-2027
Articles about StarVLA
- StarVLA's Lego-Like Codebase Is Assembling the Embodied AI Lab — The open-source platform from HKUST and Microsoft researchers has hit 98.8% success on benchmark tasks and unified a fragmented research field.