Skip to main content
Accessibility
← Back to feed
Official announcementNVIDIA Developer Blog

NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories

Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users,... Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agent

NVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories

Traditional cloud infrastructure was designed for predictable, general-purpose workloads and standard interfaces. Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security. NVIDIA is introducing Scale-In network infrastructure, the fifth pillar of NVIDIA AI networking, bringing purpose-built acceleration to secure, manage, and operate agentic AI factories.

Scale-In evolves north-south networks into a coordinated infrastructure domain for the AI factory. Powered by NVIDIA BlueField-4 and NVIDIA DOCA and connected over NVIDIA Spectrum-X Ethernet, Scale-In accelerates the services that secure the full AI stack and move application, data, and storage traffic across the AI factory. Dedicated, host-independent processing keeps these infrastructure services off host CPUs, helping prevent security, data access, and operations from becoming bottlenecks as AI compute scales. The result is a more secure and efficient shared AI infrastructure with consistent access to AI services and data as demand grows.

This post explores the BlueField-4 architecture and explains how Scale-In delivers secure access, high-performance data movement, tenant isolation, simpler operations, and more predictable performance for agentic AI factories.

Scale-In accelerates north-south AI factory infrastructure[](#scale-in_accelerates_north-south_ai_factory_infrastructure)

AI factories have distinct infrastructure requirements at different scales:

  • Scale-Up: NVIDIA NVLink unites GPUs as a coherent accelerator.
  • Scale-Out: NVIDIA Spectrum-X Ethernet and NVIDIA Quantum InfiniBand connect servers across an AI factory.
  • Scale-Across: NVIDIA Spectrum-XGS Ethernet connects distributed AI factories.
  • Context Memory: NVIDIA CMX, built on the NVIDIA STX modular foundation for AI-native storage, provides shared KV-cache storage within the factory.
  • Scale-In: NVIDIA BlueField-4, NVIDIA DOCA, and NVIDIA Spectrum-X Ethernet accelerate the access, security, data movement, and infrastructure operations surrounding AI compute.

North-south networks provide the access path into and out of a data center, connecting users, applications, data sources, storage systems, and services to individual systems. Traditional cloud data centers built these networks around software-defined infrastructure, composability, and elasticity, so resources and access could be provisioned and scaled as demand changed.

Agentic AI raises the demand on this infrastructure. Software-defined networking, composability, and elasticity remain essential, but they are no longer sufficient. AI factories bring together massive accelerated compute with growing numbers of users, agents, applications, enterprise data sources, and storage systems, all interacting continuously and at scale.

The infrastructure must be accelerated and co-designed as part of the AI factory. Security, multi-tenant networking, data and storage access, and infrastructure operations can’t rely solely on software running on general-purpose host CPUs or operate as independently managed layers. These functions must work together as a unified infrastructure domain, giving operators consistent control over access, data movement, security, provisioning, and observability across the AI factory.

Scale-In addresses these requirements by evolving north-south access into a unified, accelerated infrastructure domain across the AI factory. BlueField-4 provides host-independent acceleration and offload, DOCA provides a unified model for programming and operating infrastructure services, and Spectrum-X Ethernet provides high-performance Ethernet connectivity across the Scale-In access path, including external storage and data sources. Together, they give operators consistent control over access, data movement, security, provisioning, and observability across the infrastructure that supports accelerated compute and data storage.

Figure 1. Scale-Up, Scale-Out, and Scale-Across expand compute connectivity, and Scale-In connects, secures, provisions, and observes the infrastructure surrounding the compute domain

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC
News summary

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC

Anthropic News

Anthropic pushes into physical world with new standard to help AI agents operate machines CNBC

AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery
Official announcement

AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery

AWS What’s New

AWS Elastic Disaster Recovery (AWS DRS) now offers Recovery Plans, a capability that automates the sequential launch of multi-server applications during recovery and drills. Instead of launching servers one at a time and tracking dependencies manually, you define the recovery seq

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
Official announcement

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules

NVIDIA Developer Blog

The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
Official announcement

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit

NVIDIA Developer Blog

Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and ac

Samsung Introduces New Odyssey Lineup for Fast-Paced Gaming at Gamescom 2026
Official announcement

Samsung Introduces New Odyssey Lineup for Fast-Paced Gaming at Gamescom 2026

Samsung Newsroom

Samsung Electronics today announced its 2027 Odyssey gaming monitor lineup at Gamescom 2026, the world’s largest gaming event, being held in Cologne, Germany from Aug. 26-30. The new lineup introduces multiple Odyssey models that feature world-first innovations, empowering player

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
Official announcement

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache

Sebastiaan Neuteboom

Big Pineapple , the platform behind 1.1.1.1 , Gateway DNS , DNS Firewall , AS112 , and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across