Blaxel is joining Baseten to build the future of agentic infrastructure

Baseten Blog ·

A source describing Blaxel's technical infrastructure for autonomous agents—covering microVM sandbox performance, production networking design, inference colocation requirements, and foundational infrastructure design principles—based on statements attributed to Blaxel. Read 4 viewpoints with supporting evidence and source links.

Amir Haghighat, Tuhin Srivastava, Paul Sinai

Understand this piece

4 key points

Synthesis

  1. MicroVM sandboxes with 25ms suspend/resume

    Blaxel built isolated microVM sandboxes for agents, enabling safe code execution and state persistence. They suspend and resume in 25 milliseconds—up to 5× faster than other sandbox products—and can remain idle for months at near-zero cost while retaining state.

    Supporting evidence 1

    Original excerpt

    Every agent gets its own microVM. We made them extremely fast: suspend and resume in 25 milliseconds, up to 5x faster than other sandbox products. A sandbox can sit idle for months at close to zero cost and come back before the model finishes its next sentence, so state persists without paying for compute you aren't using.

    Amir Haghighat, Tuhin Srivastava, Paul Sinai · Paragraph 6

    Context

    Sandboxes came first because nothing else works without them. Agents write and run code, call tools, and carry state across long tasks, and you can't do that safely in a shared process.

    Read in source context →
  2. Rebuilt networking layer with isolation and access controls

    Blaxel rebuilt its networking layer to connect agents with tools, MCP servers, APIs, and other agents—designed specifically to meet production teams’ requirements for isolation and access controls.

    Supporting evidence 1

    Original excerpt

    And agents have to talk to things. Tools, MCP servers, APIs, other agents. We rebuilt a networking layer that connects all of these pieces with the isolation and access controls a production team will actually sign off on.

    Amir Haghighat, Tuhin Srivastava, Paul Sinai · Paragraph 9

    Read in source context →
  3. Inference must sit adjacent to compute, storage, and network

    The Blaxel team observed that inference for autonomous agents is no longer a remote service call: it must be colocated with compute, storage, and networking to avoid compounding latency on every agent loop.

    Supporting evidence 1

    Original excerpt

    The primitive we didn't own was inference. We've seen open-weight & custom models become core to how autonomous agents get deployed in production. Inference is no longer a service you call from another data center. It has to sit next to compute, storage, and network; or latency compounds on every loop.

    Amir Haghighat, Tuhin Srivastava, Paul Sinai · Paragraph 10

    Read in source context →
  4. First-class consideration of performance, reliability, security, and ownership

    Both Blaxel and Baseten built their infrastructure from the ground up—prioritizing performance, reliability, security, scalability, flexibility, cost-efficiency, and customer visibility, control, and ownership—rather than taking the fastest path to market.

    Supporting evidence 1

    Original excerpt

    Both companies built from the ground up for what it would take to truly scale the workload to our customers’ requirements, rather than taking the fastest path to market. We had both decided that performance, reliability, scalability, security, flexibility, and cost-efficiency needed to be first-class considerations. And we each believed customers need maximum visibility, control, and ownership.

    Amir Haghighat, Tuhin Srivastava, Paul Sinai · Paragraph 16

    Read in source context →

Key passages4

Attributed passages with the context to verify them. Open the original text to check the source.

infrastructure design philosophy

First-class consideration of performance, reliability, security, and ownership

Original excerpt

Both companies built from the ground up for what it would take to truly scale the workload to our customers’ requirements, rather than taking the fastest path to market. We had both decided that performance, reliability, scalability, security, flexibility, and cost-efficiency needed to be first-class considerations. And we each believed customers need maximum visibility, control, and ownership.
agent sandbox performance

MicroVM sandboxes with 25ms suspend/resume

Original excerpt

Every agent gets its own microVM. We made them extremely fast: suspend and resume in 25 milliseconds, up to 5x faster than other sandbox products. A sandbox can sit idle for months at close to zero cost and come back before the model finishes its next sentence, so state persists without paying for compute you aren't using.
Context

Sandboxes came first because nothing else works without them. Agents write and run code, call tools, and carry state across long tasks, and you can't do that safely in a shared process.

inference colocation for agents

Inference must sit adjacent to compute, storage, and network

Original excerpt

The primitive we didn't own was inference. We've seen open-weight & custom models become core to how autonomous agents get deployed in production. Inference is no longer a service you call from another data center. It has to sit next to compute, storage, and network; or latency compounds on every loop.

Source & methodology

These viewpoints are linked to their original sources. Paraphrases are labeled and are not verbatim quotes.

Open transcript or source material (opens in a new tab)Report an issue