AMD Instinct™ · Data Center GPU
Systems & Infrastructure Documentation#
Everything you need to deploy, validate, and operate AMD Instinct™ Data Center GPUs at scale — drivers, orchestration, cluster management, and acceptance testing for HPC and AI. For API and software-stack reference, see the ROCm documentation.
Start here
System Administrators
Deploy and run AMD Instinct GPUs on bare metal, in containers, and across clusters. These guides are the most frequently updated content on this site.
Bare metal
Instinct GPU Driver
Install and configure the GPU, including logging and error codes.
GPU Partitioning
Split compute units and memory to partition a single GPU.
AMD SMI
Unified user-space tool to manage and monitor GPUs and drivers.
ROCm Validation Suite
System validation and hardware diagnostics.
Customer Acceptance Guide
Configure, validate, benchmark, and baseline Instinct GPUs.
Cluster Validation Suite
Test scripts that validate AMD AI clusters end to end.
Containers & orchestration
GPU Operator
Deploy and manage Instinct GPUs in Kubernetes clusters.
Network Operator
Simplify AMD AINICs in Kubernetes environments.
Device Plugin
Register AMD GPUs with a Kubernetes container cluster.
Device Metrics Exporter
Prometheus-format GPU metrics for HPC and AI environments.
AMD Container Toolkit
Integrate Instinct GPUs with Docker and container runtimes.
Spur
AI-native job scheduler, drop-in compatible with Slurm, with GPU-first scheduling and Raft-based state.
Cluster, cloud & virtualization
Enterprise AI
Tools to manage enterprise AI infrastructure at scale.
Omnistat
Profile GPU resource utilization across the cluster.
Cluster Networking
Optimize the network for Instinct GPU applications.
Instinct on Azure
Get started with AMD Instinct GPUs on Microsoft Azure.
Virtualization Driver
Explore the virtualization driver for Instinct GPUs.
AMD SMI for Virtualization
Manage and monitor virtualization-enabled AMD GPUs.
Common Reference
Architecture, programming models, and technical collateral that span every deployment.
Instinct Micro-architecture
Hardware details for MI350, MI300, MI200, and MI100 accelerators.
AMD SMI API Reference
Full AMD SMI documentation covering all use cases.
HIP C++
Learn the HIP programming model.
OpenMP
Explore the OpenMP programming model.
Technical Information Portal
NDA technical documentation and design collateral. Login required.