Beam's Open-Source AI Runtime Launches GPU Containers in Under a Second

The Y Combinator-backed platform is betting that pay-per-second billing and self-hostability can carve a niche in the crowded AI infrastructure market.

About Beam

Published

The most expensive part of running an AI model is the time it spends idle. Beam, an open-source infrastructure platform, is built on the premise that you should only pay for the milliseconds of GPU time you actually use. Its technical wedge is a serverless runtime that can spin up a GPU-backed container in under a second, offering a Pythonic interface and the option to run the whole stack on your own hardware [Y Combinator, 2024].

The Pay-Per-Second Wedge

Beam's core offering is a managed cloud for AI workloads, but its differentiation is granular. Where traditional cloud providers bill for GPU instances by the hour, Beam charges for compute in one-second increments [Perplexity Sonar Pro Brief, 2024]. For developers running sporadic inference jobs, background task queues, or experimental sandboxes, this can translate to meaningful cost savings. The platform abstracts the underlying infrastructure behind a simple SDK.

Metric Value
Inference Endpoints 1 product surface
Task Queues 1 product surface
Sandboxes 1 product surface

The product surfaces are straightforward: high-performance inference endpoints, managed task queues for large-scale workloads, and secure, sandboxed code execution environments [Beam, 2024].

Backing and Early Traction

Founded in 2022 by Eli Mernit and Luke Lombardi, Beam was part of Y Combinator's W22 batch. The founders previously built Slai, a company focused on smart decision-making tools [TechCrunch, 2026]. The investor list includes notable angels from Snyk and GitHub, alongside institutional backing from Tiger Global Management and Sequoia Scout [Perplexity Sonar Pro Brief, 2024].

Early customers named by the company include Coca-Cola, Magellan AI, Shippabo, and Stratum, using Beam for serverless inference, sandboxes, and background jobs [Perplexity Sonar Pro Brief, 2024].

The Competitive Landscape

Beam operates in a space crowded with well-funded alternatives. It positions itself as an open-source alternative to Modal [Perplexity Sonar Pro Brief, 2024]. The competitive set also includes specialized GPU cloud providers like CoreWeave and broader infrastructure tools like Spheron.

  • Cost transparency. Pay-per-second billing is a clear, quantitative differentiator for workloads with spiky or unpredictable demand.
  • Deployment control. The open-source Beta9 runtime allows for complete self-hosting, addressing data sovereignty and vendor lock-in concerns [Beam, 2026].
  • Speed as a feature. Sub-second container launch times are critical for interactive applications like AI chatbots or agentic workflows.

Technical Breakdown and Scale Considerations

Under the hood, Beam's challenge is orchestrating cold starts efficiently. Launching a container with a loaded GPU model in under a second requires significant optimization at the container orchestration and driver level. The pay-per-second model depends on efficient bin-packing of workloads across a finite GPU fleet to maintain profitability. While the platform supports custom model inference, it is not currently positioned for large-scale distributed training jobs.

The Road Ahead

The next twelve months will be about proving that the early wedge can drive sustainable growth. Key milestones to watch include a potential Series A round to scale the cloud infrastructure, the expansion of the runtime to support more languages, and the landing of a flagship enterprise deal. The company is actively hiring for growth and engineering roles [Y Combinator, 2024]. For Beam, the path forward is not about out-featuring the giants, but about being indispensable for the specific, granular workloads where its economics and performance uniquely align.

Read on Startuply.vc