Skip to main content
← Back to feed
Official announcementSnowflake Blog

Per-User Quotas for AI: Govern Individual Spend in Snowflake

Per-user quotas are now generally available in Snowflake. Set individual credit limits on AI Functions, Cortex Agents, CoWork and more — with automatic blocking, notifications and custom actions to govern AI spend at the person level.

Per-User Quotas for AI: Govern Individual Spend in Snowflake

Every enterprise executive faces the same fundamental tension when deploying AI: how to give teams the freedom to innovate while ensuring a single runaway process doesn't result in a five-figure AI bill.

When AI access is unrestricted, a single developer triggering an agent loop or an analyst running an unoptimized prompt across a million-row table can unexpectedly drain a department's monthly budget in a single afternoon. The traditional solution — restricting access to a small group of approved users — protects margins but can stall business transformation.

Per-user quotas, now generally available in Snowflake, are designed to eliminate this trade-off.

Instead of an enterprise having to choose between financial risk and slow innovation, per-user quotas provide automated, individual-level spend guardrails. Every user gets the access they need, while finance and platform teams gain the confidence that costs remain predictable, controlled and scalable.

Traditional cost controls (such as aggregate team budgets) monitor total spending for an entire department or warehouse. While effective for macrolevel tracking, aggregate budgets have a critical blind spot: They show you that a team overspent, but not who did it or why until after the bill arrives.

Per-user quotas flip this model from reactive tracking to proactive protection:

One runaway script or prompt drains a team's entire monthly budget.

Isolate and restrict overages with granular daily and monthly limits.

Access is restricted to select power users to prevent surprise costs.

Universal self-service provides access across the entire organization.

FinOps spends hours hunting down user attribution after a cost spike.

Automated block enforcement and user notifications make it easy.

Users are kept in the dark until access is manually revoked.

Proactive alerts nudge users to self-correct before hitting limits.

A key enabler for scaling AI safely is block enforcement, which allows a quota to act as an automated “circuit breaker” when usage reaches an overload point. This allows for:

Targeted containment: If a user hits their daily or monthly limit on a governed AI service, their access to that specific service is automatically paused until the next cycle.

Limited collateral damage: Restricting one user's AI access never disrupts their standard warehouse compute or impacts anyone else on the team. Access lifts automatically when the new cycle begins — no manual IT support tickets required.

Fast enforcements: Tracking against users spend happens in minutes including for in flight requests and queries which means overspend is caught in a timely manner.

Per-user quotas deliver value across every tier of the business:

Per-user quotas offer granular coverage across both platform compute and AI services, allowing organizations to calibrate limits based on specific risk profiles. For instance, the quotas provide coverage for the following services:

Warehouse compute: Track query execution credits at the individual level.

Cortex AI Functions: Govern Snowflake Cortex AI’s SQL functions (such as AI_COMPLETE, AI_SUMMARIZE and AI_TRANSLATE).

Cortex Agents and CoWork: Control operational credits used across agent workflows and collaborative business tools.

Snowflake CoCo: Manage developer assistant credit usage across Snowsight UI, CLI and desktop environments.

Because platform compute and AI features use different credit structures, limits are set independently. You can set conservative caps on high-velocity AI functions while providing generous limits for standard daily queries.

Per-user quotas are part of a complete, multilayered cost governance framework within Snowflake that includes:

Snowflake Budgets: Provide macrolevel visibility and aggregate tracking across teams, projects and cost centers.

Per-user quotas: Provide microlevel enforcement and individual guardrails to prevent outlier spend.

Migrating the GitHub Copilot runtime to Rust, using Copilot
Official announcement

Migrating the GitHub Copilot runtime to Rust, using Copilot

GitHub BlogStephen Toub

The GitHub Copilot CLI , GitHub Copilot app , and GitHub Copilot SDK are all backed by the Copilot agent runtime, an agentic harness that can be embedded into applications and services. It was originally written in TypeScript on Node.js and the V8 JavaScript engine for what is no

How to Use AI Agents to Prepare 3D Scenes for Simulation
Official announcement

How to Use AI Agents to Prepare 3D Scenes for Simulation

NVIDIA Developer Blog

Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in... Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
Official announcement

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA Developer Blog

AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through... AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answe

"Regex for Rows": Simplifying Pattern Detection in SQL with MATCH_RECOGNIZE
Official announcement

"Regex for Rows": Simplifying Pattern Detection in SQL with MATCH_RECOGNIZE

Databricks Blog

Imagine you work in cybersecurity and you have a table that tracks login attempts...

Translating CUDA Tile Operations from Python to Rust Using Agentic AI
Official announcement

Translating CUDA Tile Operations from Python to Rust Using Agentic AI

NVIDIA Developer Blog

cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to... cuTile Rust () is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Exte

Threads Introduces Parental Supervision for Teens in APAC
Official announcement

Threads Introduces Parental Supervision for Teens in APAC

Meta NewsroomFacebook

As part of our ongoing commitment to providing parents with helpful tools to support their teens across Meta’s apps, we’re bringing parental supervision to Threads. Starting this week in Asia Pacific countries, parents and guardians in Family Center will have visibility and contr