AMD Instinct MI100#
The AMD Instinct™ MI100 is a data-center compute PCIe-form-factor GPU. This document provides MI100-specific prerequisites, health checks, validation steps, and performance acceptance criteria.
Overview#
The AMD Instinct MI100 introduces the first-generation CDNA architecture in a standard full-height, full-length, dual-slot PCIe® add-in card aimed at HPC and accelerated computing workloads. Each MI100 provides 120 compute units with Matrix Core technology, 32 GB of HBM2 memory at up to 1.2 TB/s, and AMD Infinity Fabric™ link support for direct GPU-to-GPU connectivity in 2- and 4-GPU hive configurations. The card is passively cooled with a 300 W TDP and supports PCIe® Gen4 host connectivity.
The MI100 is built on the CDNA architecture (gfx908) with 120 compute units and 32 GB of HBM2 memory per GPU. The MI100 Infinity Fabric™ topology tops out at 4 GPUs per hive, so the validation reference configuration for this document is a single 4-GPU MI100 hive with Infinity Fabric™ bridges providing direct GPU-to-GPU connectivity across all peers. Larger deployments (for example, dual-socket servers with two 4-GPU hives for 8 MI100s total) are common; in those systems, cross-hive traffic traverses the host PCIe fabric and the per-hive criteria below apply to each hive independently.
System requirements#
Operating system support#
For the most up-to-date information on supported operating systems and distributions, refer to the official ROCm documentation:
ROCm System Requirements - Supported Distributions
Note
ROCm docs is the single source of truth for supported versions, distribution compatibility, and required dependencies for the ROCm toolkit.
For general NUMA and OS-level tuning that applies to all AMD Instinct hosts, see OS tuning. MI100-specific BIOS settings are in the BIOS settings section below.
GPU identification#
All MI100 GPUs (PCI vendor:device 1002:738c) should appear in lspci output:
sudo lspci -d 1002:738c
Expected output example (4-GPU MI100 hive):
1d:00.0 Display controller: Advanced Micro Devices, Inc. [AMD/ATI] Arcturus GL-XL [Instinct MI100] (rev 01)
20:00.0 Display controller: Advanced Micro Devices, Inc. [AMD/ATI] Arcturus GL-XL [Instinct MI100] (rev 01)
23:00.0 Display controller: Advanced Micro Devices, Inc. [AMD/ATI] Arcturus GL-XL [Instinct MI100] (rev 01)
26:00.0 Display controller: Advanced Micro Devices, Inc. [AMD/ATI] Arcturus GL-XL [Instinct MI100] (rev 01)
BIOS settings#
For maximum MI100 GPU performance on systems with AMD EPYC™ 7002-series processors (codename “Rome”) and AMI System BIOS, the following BIOS settings have been validated. These settings should be set as default values in the system BIOS. Analogous settings for other non-AMI BIOS providers can be configured similarly. On systems with Intel processors, some settings may not apply.
Note
The BIOS setting locations and names may vary by hardware and BIOS vendor. Consult your system documentation for details.
BIOS setting location |
Parameter |
Value |
Comments |
|---|---|---|---|
Advanced / PCI subsystem settings |
Above 4G decoding |
Enabled |
GPU large BAR support. |
AMD CBS / CPU common options |
Global C-state control |
Auto |
Global C-states. |
AMD CBS / CPU common options |
CCD/Core/Thread enablement |
Accept |
May be necessary to enable the SMT control menu. |
AMD CBS / CPU common options / performance |
SMT control |
Disable |
Set to Auto if the primary application is not compute-bound. |
AMD CBS / DF common options / memory addressing |
NUMA nodes per socket |
NPS 1, 2, or 4 |
See NPS memory configuration below. |
AMD CBS / DF common options / memory addressing |
Memory interleaving |
Auto |
Depends on NPS setting. |
AMD CBS / DF common options / link |
4-link xGMI max speed |
18 Gbps |
Set to highest rate supported by the CPU. |
AMD CBS / DF common options / link |
3-link xGMI max speed |
18 Gbps |
Set to highest rate supported by the CPU. |
AMD CBS / NBIO common options |
IOMMU |
Disabled |
|
AMD CBS / NBIO common options |
PCIe ten bit tag support |
Enabled |
|
AMD CBS / NBIO common options |
Preferred IO |
Manual |
|
AMD CBS / NBIO common options |
Preferred IO bus |
Use |
|
AMD CBS / NBIO common options |
Enhanced Preferred IO mode |
Enabled |
|
AMD CBS / NBIO common options / SMU common options |
Determinism control |
Manual |
|
AMD CBS / NBIO common options / SMU common options |
Determinism slider |
Power |
|
AMD CBS / NBIO common options / SMU common options |
cTDP control |
Manual |
|
AMD CBS / NBIO common options / SMU common options |
cTDP |
240 |
Value in watts. |
AMD CBS / NBIO common options / SMU common options |
Package power limit control |
Manual |
|
AMD CBS / NBIO common options / SMU common options |
Package power limit |
240 |
Value in watts. |
AMD CBS / NBIO common options / SMU common options |
xGMI link width control |
Manual |
|
AMD CBS / NBIO common options / SMU common options |
xGMI force link width |
2 |
0: x2 / 1: x8 / 2: x16 |
AMD CBS / NBIO common options / SMU common options |
xGMI force link width control |
Force |
|
AMD CBS / NBIO common options / SMU common options |
APBDIS |
1 |
Disable DF P-states. |
AMD CBS / NBIO common options / SMU common options |
DF C-states |
Auto |
|
AMD CBS / NBIO common options / SMU common options |
Fixed SOC P-state |
P0 |
|
AMD CBS / UMC common options / DDR4 common options |
Enforce POR |
Accept |
|
AMD CBS / UMC common options / DDR4 common options / Enforce POR |
Overclock |
Enabled |
|
AMD CBS / UMC common options / DDR4 common options / Enforce POR |
Memory clock speed |
1600 MHz |
Set to max speed if using 3200 MHz DIMMs. |
AMD CBS / UMC common options / DDR4 common options / DRAM controller configuration / DRAM power options |
Power down enable |
Disabled |
RAM power down. |
AMD CBS / security |
TSME |
Disabled |
Memory encryption. |
NBIO link clock frequency#
The NBIOs (4× per AMD EPYC™ processor) are the serializers/deserializers (SerDes) that convert and prepare I/O signals for the processor’s 128 external I/O interface lanes (32 per NBIO). The NBIO link clock frequency (LCLK) controls the speed of the internal bus connecting NBIO silicon to the data fabric. All data between the processor and PCIe lanes flows through the data fabric at this frequency, so it must be forced to the maximum for optimal PCIe performance.
For AMD EPYC™ 7002-series processors, this cannot be set via BIOS alone. The AMD-IOPM-UTIL must be run at every server boot to disable Dynamic Power Management for all PCIe root complexes and NBIOs and lock them into the highest performance mode. See AMD-IOPM-UTIL below.
NPS memory configuration#
For the number of NUMA nodes per socket (NPS), follow the guidance of the HPC Tuning Guide for AMD EPYC™ 7002 Series Processors for the optimal host configuration.
NPS1: Bidirectional copy bandwidth between host memory and GPU memory may be up to ~16% higher than with NPS4. Recommended for applications not optimized for NUMA locality.
NPS4: Recommended for memory-bandwidth-sensitive applications using MPI.
AMD-IOPM-UTIL#
The AMD I/O Power Management Utility (AMD-IOPM-UTIL) disables Dynamic Power Management (DPM) for all PCIe root complexes on AMD EPYC™ 7002-series processors and locks the NBIO logic into the highest performance operational mode. This is required for MI100 systems to ensure maximum LCLK frequency.
Disabling I/O DPM reduces latency and improves throughput for low-bandwidth PCIe messages, benefiting InfiniBand NICs, GPUs, and other bursty PCIe devices.
Note
The effects of AMD-IOPM-UTIL do not persist across reboots. No firmware settings need to change when using this utility. The “Preferred IO” and “Enhanced Preferred IO” BIOS settings should remain enabled.
The recommended method is to create a one-shot systemd service unit that runs the utility at boot. The installer packages from the AMD I/O Power Management Utility page create and enable this service unit automatically. The service runs in one-shot mode, so systemctl status will show inactive after a successful run — this is expected. A failed status means the utility did not run correctly.
To undo the effects, disable the service with systemctl disable and reboot.
Acceptance criteria#
The MI100 system acceptance process validates that the platform is correctly configured, stable, and performing to expectations. Follow the sequence: Prerequisites → Basic Health Checks → System Validation → Performance Benchmarks.
System acceptance process#
Prerequisites validation - Ensure all system requirements and dependencies are met
Basic health checks - Verify hardware detection and basic system health
System validation - Conduct comprehensive stress testing and qualification
Performance benchmarks - Validate compute, memory, and interconnect performance
The system is accepted when all criteria below are successfully validated.
Prerequisites validation#
Ensure all system requirements are met before proceeding with validation. See the Prerequisites documentation and System setup for more details.
✅ Supported operating system version installed
✅ Compatible ROCm version installed
✅ BIOS configured per MI100 BIOS settings (EPYC 7002-specific table above)
✅ Required kernel parameters present:
pci=realloc=off,pci=bfsort,iommu=pt, andamd_iommu=on(orintel_iommu=onon Intel hosts) — see Kernel Parameters✅ Minimum 256G system memory available
✅ Latest applicable firmware applied consistently across nodes
✅ ROCm Validation Suite (RVS) installed
Basic health checks#
These checks ensure fundamental system health and proper GPU detection. For detailed procedures, see Health Checks.
Test |
Command |
Pass/Fail criteria |
|---|---|---|
|
Pass: OS version listed in compatibility matrix |
|
|
Pass: Contains |
|
|
Pass: Null |
|
|
Pass: ≥ 256G |
|
|
Pass: 4 MI100 GPUs found (per hive) |
|
|
Pass: Speed 16GT/s, width |
|
|
Pass: Idle metrics as specified |
|
|
Pass: Null |
System validation#
Comprehensive validation ensures system stability under load. For detailed procedures, see System Validation.
Test |
Command |
Pass/Fail criteria |
|---|---|---|
|
Pass: All GPUs listed with no errors |
|
|
Pass: |
|
|
Pass: |
|
|
Pass: All tests passed; bandwidth ≥ 800 GB/s per GPU |
|
|
Pass: All distances and bandwidths displayed |
|
|
Pass: All actions true |
|
|
Pass: |
Note
The reference configuration for this document is a single 4-GPU MI100 hive with AMD Infinity Fabric™ bridges installed, so intra-hive PBQT and TransferBench numbers reflect XGMI throughput. On systems without bridges, P2P traffic traverses the host PCIe fabric and these thresholds will not be met.
Performance benchmarks#
Performance validation ensures the system meets MI100 specifications. For detailed procedures, see Performance Benchmarking.
TransferBench a2aPass: ≥ 270 GB/s aggregate
TransferBench p2pTest |
Pass criteria |
|---|---|
UniDir |
≥ 30 GB/s |
BiDir |
≥ 57 GB/s |
build/all_reduce_perf -b 8 -e 8G -f 2 -g 4Pass: ≥ 72 GB/s busbw (peak, at 8 GiB message size)
rocblas-bench (see code block below)rocblas-bench -f gemm \
-r s -m 4000 -n 4000 -k 4000 \
--lda 4000 --ldb 4000 --ldc 4000 \
--transposeA N --transposeB T
Pass: ≥ 28 TFLOPS per GPU
mpiexec -n 4 wrapper.shKernel |
Threshold (MB/s) |
|---|---|
Copy |
≥ 940,000 |
Mul |
≥ 940,000 |
Add |
≥ 910,000 |
Triad |
≥ 910,000 |
Dot |
≥ 950,000 |