AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver. For AI inference workloads, this makes application-level performance per watt the key metric for measuring AI factory efficiency.
Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS
AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available... AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but h

Not every megawatt translates to revenue-generating compute. Power distribution, cooling, networking, storage, backup, and facility overhead take a share of the power before it reaches a GPU. Static rack provisioning exacerbates this: an outdated approach to data center design allocates the maximum power draw per rack to meet worst-case peak demand, even though real workloads have different power needs and may leave some portion of that maximum power unused. Operators reserve additional capacity for failures, operational flexibility, and expansion.
In one representative power-budget view examined by NVIDIA, about 60% of delivered site power is allocated to compute for AI output.
NVIDIA DSX MaxLPS is a suite of chip, thermal, system, and software technologies that maximizes AI factory throughput within a fixed power budget. MaxLPS stands for Maximum Land Power Shell, the site-level constraints that define an AI factory: land, utility power, and the physical shell holding power, cooling, networking, and compute infrastructure.
MaxLPS is designed to optimize these three layers:
- Dynamic power allocation: Continuously monitors and allocates unused power headroom to GPUs
- Advanced performance per watt techniques: Software power optimization techniques that improve job-level performance at a fixed power budget
- 45° C thermal efficiency and site design: Cuts cooling overhead through warm-water liquid cooling, improving power usage effectiveness (PUE) to convert directly into more compute within the same fixed envelope
Why static rack provisioning strands power[](#why_static_rack_provisioning_strands_power)
Traditional data center power planning reserves enough power as though every rack could draw its specified maximum power simultaneously. That protects the facility against peak demand, but it treats each rack as an isolated power island. An isolated rack provisioned with excess power cannot lend that unused power to a neighbor that could turn it into tokens.
At AI factory scale, facility overhead, rack losses, and operational inefficiencies during failures, restarts, and checkpointing reduce the power available to the AI load. Separately, static rack provisioning can strand headroom within the power allocated to racks. Power reserved for one rack’s peak demand may sit unused while another rack could use it. Dynamic power allocation targets this reclaimable rack-level headroom.
Figure 1 shows a power-budget waterfall for a 100 MW AI factory. Of the 100 MW grid input, 20 MW is allocated to facility overhead, 10 MW to rack losses, and 10 MW is unavailable for AI load because of operational inefficiency during failures, restarts, and checkpointing. This leaves 60 MW available for AI load. Each deduction is expressed as a share of the original grid input.
Figure 1. Illustrative power-budget waterfall for a 100 MW AI factory
Related stories

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC
Anthropic News
Anthropic pushes into physical world with new standard to help AI agents operate machines CNBC
AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery
AWS What’s New
AWS Elastic Disaster Recovery (AWS DRS) now offers Recovery Plans, a capability that automates the sequential launch of multi-server applications during recovery and drills. Instead of launching servers one at a time and tracking dependencies manually, you define the recovery seq

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
NVIDIA Developer Blog
The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
NVIDIA Developer Blog
Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and ac

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
Sebastiaan Neuteboom
Big Pineapple , the platform behind 1.1.1.1 , Gateway DNS , DNS Firewall , AS112 , and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across

Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs
Lauro Ojeda
Summary This post is for technical decision makers evaluating where to run production PostgreSQL workloads. It compares two valid operating models—self-managed PostgreSQL and a managed database service—through business and operational outcomes: control, engineering capacity, resi
