Connect with us
Photo courtesy Red Hat.

Product of the Day

Red Hat puts AI on a tighter leash

An update to the company’s enterprise software platform adds model checks, GPU controls and agent safeguards, writes GUIDO RUSSO.

Red Hat, a leading provider of open-source technology, has revealed significant updates across its AI portfolio. These include the release of Red Hat AI 3.5, an update to the company’s enterprise software platform designed to run, scale and secure AI models and agents across hybrid cloud environments.

“As enterprise teams move past early experimentation and pilot successes, IT and platform engineering leaders face the challenge of running AI with the same operational rigor as mission-critical infrastructure,” says Red Hat. “By providing the scalable foundation required to control, secure and observe these workloads across the hybrid cloud, Red Hat AI 3.5 bridges the gap between isolated AI pilots and a fully governed enterprise architecture.”

Red Hat AI 3.5 adds new safety, observability and infrastructure management features aimed at supporting AI deployments across hybrid environments.

The release includes EvalHub, which allows organisations to evaluate models before deployment using safety benchmarks and create regulatory compliance certifications. New observability dashboards provide metrics on inference health, GPU utilisation and AI model performance. Non-admin users can also access dashboards showing per-user token consumption and distributed inference workloads.

Red Hat AI 3.5 also adds multi-tenancy capabilities for AI service providers and use cases that require greater separation between workloads. Priority-aware serving supports multiple tenants sharing GPU infrastructure, while organisations requiring stronger isolation can run Red Hat AI on Red Hat OpenShift hosted control planes deployed on Red Hat OpenShift Virtualisation.

Hosted control planes provide each tenant with a dedicated cluster control plane while sharing the underlying hardware. AI workloads can also run in Red Hat OpenShift Virtualisation virtual machines, providing VM-level isolation across shared GPU infrastructure. Red Hat says this configuration allows infrastructure providers to manage and upgrade the underlying environment from a central point.

The update also expands tools for developing and managing AI agents. AutoRAG connects enterprise data repositories to agentic applications and adds multilingual document support, conversational testing and contextual retrieval. A visual pipeline can be used to test retrieval-augmented generation configurations before deployment.

Agent templates provide pre-configured implementations for use cases including code review, document processing and research workflows. Inference-Time Scaling can adjust the amount of compute allocated according to the complexity of a query, with the aim of managing GPU usage more efficiently.

Joe Fernandes, Red Hat VP and GM for the AI business unit. Photo supplied.

Joe Fernandes, Red Hat VP and GM for the AI business unit, says: “The conversation has moved from getting AI into production to running it at scale as trusted enterprise infrastructure, which requires safety evidence, governed agents, cost attribution and multi-tenancy. With Red Hat AI 3.5, we are delivering the operational controls, verifiable trust and agentic foundations IT leaders need to run AI as a safe, controlled and accountable enterprise AI architecture across the hybrid cloud.”

Red Hat provides the following key takeaways:

  • Verifiable pre-deployment safety and evaluation: Evaluated catalogue models feature built-in Garak benchmark scores, while the general availability of EvalHub automates safety and auditable compliance reporting for custom models, RAG and agents.
  • Shared GPU control for multi-tenant inference: Fair-share GPU scheduling manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing to protect real-time inference and allows background workloads to use available capacity.
  • Agent APIs and gateway security: General availability support for the Responses API and built-in RAG provides a unified open-source interface for multi-turn agent conversations, reinforced by integrated NeMo Guardrails that intercept malicious tool calls.
  • Enterprise data grounding and efficient reasoning: AutoRAG with pgvector support, native AutoML, and Inference-Time Scaling (ITS) allow models to adapt compute usage dynamically based on query difficulty.
  • Built-in observability and MaaS showback: Delivers per-user token metering, performance dashboards for models and agents, MLflow visual agentic tracing, and GPU utilisation dashboards for clear operational and usage transparency.
  • Pre-built agent templates for faster development: AI Hub introduces agent templates and starter kits with pre-configured reference implementations for common enterprise patterns,  including code review, document processing, and research workflows.

Red Hat AI 3.5 is now generally available. Red Hat AI 3.5 is also now available as part of Red Hat AI Factory with NVIDIA.

Subscribe to our free newsletter
To Top