System Administrators#
Deploy and run AMD Instinct GPUs on bare metal, in containers, and across clusters.
Bare metal#
Instinct GPU Driver
Install and configure the GPU, including logging and error codes.
GPU Partitioning
Split compute units and memory to partition a single GPU.
AMD SMI
Unified user-space tool to manage and monitor GPUs and drivers.
ROCm Validation Suite
System validation and hardware diagnostics.
Customer Acceptance Guide
Configure, validate, benchmark, and baseline Instinct GPUs.
Cluster Validation Suite
Test scripts that validate AMD AI clusters end to end.
Containers & orchestration#
GPU Operator
Deploy and manage Instinct GPUs in Kubernetes clusters.
Network Operator
Simplify AMD AINICs in Kubernetes environments.
Device Plugin
Register AMD GPUs with a Kubernetes container cluster.
Device Metrics Exporter
Prometheus-format GPU metrics for HPC and AI environments.
AMD Container Toolkit
Integrate Instinct GPUs with Docker and container runtimes.
Spur
AI-native job scheduler, drop-in compatible with Slurm, with GPU-first scheduling and Raft-based state.
Cluster, cloud & virtualization#
Enterprise AI
Tools to manage enterprise AI infrastructure at scale.
Omnistat
Profile GPU resource utilization across the cluster.
Cluster Networking
Optimize the network for Instinct GPU applications.
Instinct on Azure
Get started with AMD Instinct GPUs on Microsoft Azure.
Virtualization Driver
Explore the virtualization driver for Instinct GPUs.
AMD SMI for Virtualization
Manage and monitor virtualization-enabled AMD GPUs.