<p>Amazon SageMaker HyperPod now enhances support for Ray with built-in observability, resilient training, accelerated inference and managed development environments. Ray is a popular open-source framework for scaling AI workloads on a unified compute layer, from data processing and distributed training to reinforcement learning and model serving. Running Ray on Kubernetes at production scale can be an operational burden: job hangs, low GPU utilization from static team allocations, and multi-step observability setup. Also, lack of interactive development environment means every code change needs another job submission and familiarity with kubectl.</p>
<p>HyperPod now brings easier development, resilient training, and accelerated inference to Ray. Data scientists create, edit, monitor, and delete Ray clusters from a web-based interface in Amazon SageMaker Studio, then attach JupyterLab, Code Editor, or a local IDE to a running Ray cluster and iterate interactively against cluster-scale compute. A multi-node Ray cluster behaves like a local development environment, so you test each change immediately, without waiting for a new job to queue and start. For Observability, HyperPod provisions Grafana dashboards with metrics in Amazon Managed Service for Prometheus and allows one-click access to the Ray Dashboard through a secure browser link, giving you visibility into your workloads from the first run. For training at scale, HyperPod node auto recovery and hung job detection handle GPU faults, job hangs, loss spikes, and degraded throughput. Tiered checkpointing restores state from cluster memory to maximize goodput, and task governance improves compute utilization through quotas, priorities, and preemption. Together, these keep your long training runs progressing through failures and maximize the useful work done per GPU-hour. For inference with Ray Serve, a tiered KV cache reuses cached prefixes to reduce time to first token, and you can deploy Amazon SageMaker JumpStart models directly.</p>
<p>Open-source Ray code runs unchanged and you can either adopt the purpose-built experience in SageMaker Studio or take individual capabilities to integrate into your own ML platform.</p>
<p>Ray support is available for HyperPod clusters orchestrated by Amazon EKS, in AWS Regions where SageMaker HyperPod is supported. To learn more, see the <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-ray.html" target="_blank" rel="noopener noreferrer">SageMaker HyperPod documentation</a>, and explore the <a href="https://d1dpyy0tl92esj.cloudfront.net/overview" target="_blank" rel="noopener noreferrer">interactive demo</a>.</p>
Amazon SageMaker HyperPod enhances support for Ray
Amazon SageMaker HyperPod now enhances support for Ray with built-in observability, resilient training, accelerated inference and managed development environments. Ray is a popular open-source framework for scaling AI workloads on a unified compute layer, from data processing and

Pixabay (free commercial use)
Related stories

Anthropic pushes into physical world with new standard to help AI agents operate machines - CNBC
Anthropic News
Anthropic pushes into physical world with new standard to help AI agents operate machines CNBC
AWS Elastic Disaster Recovery introduces Recovery Plans for orchestrated application recovery
AWS What’s New
AWS Elastic Disaster Recovery (AWS DRS) now offers Recovery Plans, a capability that automates the sequential launch of multi-server applications during recovery and drills. Instead of launching servers one at a time and tracking dependencies manually, you define the recovery seq

Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
NVIDIA Developer Blog
The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,... The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales

How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
NVIDIA Developer Blog
Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and accessible interfaces to the... Atomistic simulation requires three things: knowledge of the science, compute-efficient implementation of simulations, and ac

How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
Sebastiaan Neuteboom
Big Pineapple , the platform behind 1.1.1.1 , Gateway DNS , DNS Firewall , AS112 , and several other Cloudflare DNS services, stores over 250 billion DNS cache entries at any given time. At that scale, wasting a single byte per entry costs more than 250 gigabytes of memory across

Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs
Lauro Ojeda
Summary This post is for technical decision makers evaluating where to run production PostgreSQL workloads. It compares two valid operating models—self-managed PostgreSQL and a managed database service—through business and operational outcomes: control, engineering capacity, resi
