AMD Instinct MI350P 144GB Specs, Benchmarks & Pricing
The AMD Instinct MI350P is the PCIe variant of AMD's CDNA 4 architecture AI and HPC datacenter accelerator family, announced on May 7, 2026. It is not a binned or cut-down MI350X: AMD built a smaller purpose-designed CDNA 4 package for it, pairing a single I/O die with 4 Accelerator Complex Dies (XCDs) instead of the MI350X's 2 I/O dies carrying 8 XCDs. That yields 128 compute units, 8,192 stream processors, and 512 Matrix Cores at a 2,200 MHz peak engine clock, with 73 billion transistors and 128 MB of LLC. The MI350P delivers 72 TFLOPS of FP32 vector and matrix performance, 36 TFLOPS of FP64 vector performance, 1.15 PFLOPS of dense FP16/BF16 matrix performance (2.3 PFLOPS with 2:4 structured sparsity), 2.3 POPS of dense INT8 matrix performance, 2.3 PFLOPS of dense OCP-FP8 matrix performance, and 4.6 PFLOPS of MXFP4 microscaling performance. It carries 144 GB of HBM3E memory via a 4,096-bit interface at 4 TB/s (4,000 GB/s) of peak bandwidth. Unlike the OAM-based MI350X, the MI350P uses a full-height full-length dual-slot PCIe 5.0 x16 card form factor with passive air cooling at 600W maximum TBP (450W configurable), designed for drop-in deployment in existing enterprise rack servers without liquid cooling or a Universal Base Board. AMD does not expose the GPU-to-GPU Infinity Fabric links on the MI350P, so multi-GPU scaling goes over the PCIe Gen5 x16 link.
- Release Date: May 7, 2026
- GPU Architecture: CDNA 4
- Hardware-Accelerated GEMM Operations:FP64 FP32 TF32 BF16 FP16 FP8 FP4 INT8 INT4
- CUDA Compute Capability : n/a
Strengths
- Excellent FP32 compute performance (top 23% of GPUs)
- Excellent FP16 compute performance (top 6% of GPUs)
Specifications for AMD Instinct MI350P
| Specification | Performance Ranking |
|---|---|
| FP32 TFLOPs | |
| FP16 TFLOPs | |
| Tensor Core Count | |
| Memory Capacity (GB) | |
| Memory Bandwidth (GB/s) | |
| Int8 TOPs | |
| FP8 TFLOPs | |
| FP4 TFLOPs |
Real-time AMD Instinct MI350P GPU Prices
Compare Price/Performance to other GPUs
Compare AMD Instinct MI350P to Another GPU
Price History
AMD Instinct MI350P Price History
References
- https://www.amd.com/en/products/accelerators/instinct/mi350/mi350p.html
- https://www.amd.com/en/products/accelerators/instinct/mi350.html
- https://www.amd.com/en/blogs/2026/amd-instinct-mi350p-pcie-gpus-run-enterprise-ai-on-your.html
- https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/white-papers/amd-cdna-4-architecture-whitepaper.pdf
- https://www.exxactcorp.com/blog/news/amd-instinct-mi350p
- https://www.servethehome.com/amd-intros-instinct-mi350p-accelerator-cdna-4-comes-to-pcie-cards/
- https://www.storagereview.com/news/amd-instinct-mi350p-enterprise-pcie-ai-inference-returns-to-standard-servers
Notes
- fp32TFLOPS of 72 represents peak FP32 vector/matrix performance per AMD official product page (amd.com/en/products/accelerators/instinct/mi350/mi350p.html). AMD lists FP32 matrix performance equal to FP32 vector at 72 TFLOPS, consistent with the CDNA 4 parity seen on the MI350X.
- fp16TFLOPS of 2300 represents peak FP16 matrix performance with 2:4 structured sparsity (2.3 PFLOPs) per AMD official product page. Dense (non-sparse) FP16 matrix performance is 1150 TFLOPS (1.15 PFLOPs). Sparse value used per agent policy (FP16 Matrix SPARSE), following the same convention as the MI350X file.
- int8TOPS of 2300 represents non-sparse (dense) INT8 matrix performance per AMD official product page (listed as "Peak INT8 Matrix Performance: 2.3 POPs"). With 2:4 structured sparsity, INT8 performance is 4600 TOPS (4.6 POPs). Dense value used for consistent cross-vendor comparison.
- fp8TFLOPS of 2300 represents non-sparse (dense) OCP-FP8 (E5M2, E4M3) matrix performance per AMD official product page. With 2:4 structured sparsity, FP8 performance is 4600 TFLOPS. Dense value used for consistent cross-vendor comparison.
- fp4TFLOPS of 4600 (4.6 PFLOPS) represents MXFP4 matrix performance per AMD official product page. AMD publishes FP4/MXFP4 as a single microscaling-format peak rate (no separate dense-vs-sparse split is listed, unlike the FP16/BF16/INT8/OCP-FP8 rows which each list an explicit "with Structured Sparsity" variant). The 4.6 PFLOPs figure is the dense rate: it is exactly 2x dense OCP-FP8 (2.3 PFLOPs) and 4x dense FP16/BF16 (1.15 PFLOPs), the expected dense precision scaling for CDNA 4 matrix cores.
- Additional precision performance per AMD official product page: FP64 vector 36 TFLOPS; FP64 matrix 36 TFLOPS; FP32 vector 72 TFLOPS; FP32 matrix 72 TFLOPS; FP16 vector 72 TFLOPS; FP16/BF16 matrix dense 1.15 PFLOPs; FP16/BF16 matrix sparse 2.3 PFLOPs; OCP-FP8 dense 2.3 PFLOPs; OCP-FP8 sparse 4.6 PFLOPs; INT8 dense 2.3 POPs; INT8 sparse 4.6 POPs; MXFP8 2.3 PFLOPs; MXFP6 4.6 PFLOPs; MXFP4 4.6 PFLOPs. Note: MXFP6 is not representable in the schema's supportedHardwareOperations enum but CDNA 4 supports it natively, and the AMD announcement blog calls out "native support for lower-precision MXFP6 and MXFP4".
- tensorCoreCount of 512 represents AMD Matrix Cores per AMD official product page listing "Matrix Cores: 512". CDNA 4 has 4 Matrix Cores per compute unit x 128 compute units = 512. AMD uses the term "Matrix Cores" rather than "Tensor Cores".
- memoryBandwidthGBs of 4000 represents peak theoretical bandwidth of 4 TB/s from 144 GB HBM3E via 4096-bit memory interface per AMD official product page.
- Architecture: the MI350P is a purpose-built smaller CDNA 4 package, NOT a salvaged/binned MI350X. It pairs a single I/O die with 4 Accelerator Complex Dies (XCDs, 32 CUs each), versus the MI350X's 2 I/O dies carrying 8 XCDs, for 128 compute units and 8,192 stream processors total. Sources: ServeTheHome (servethehome.com/amd-intros-instinct-mi350p-accelerator-cdna-4-comes-to-pcie-cards) - "AMD is not using salvaged MI350X chips for this product. Instead, they are building a smaller chip especially for use on the MI350P"; StorageReview (storagereview.com/news/amd-instinct-mi350p-enterprise-pcie-ai-inference-returns-to-standard-servers) - "The MI350P is not a binned MI350X. AMD designed a smaller chip for it." Corroborated by the transistor count: 73 billion per AMD official product page versus 185 billion for the MI350X, i.e. below half, which a partially-disabled MI350X package would not show. Total 128 MB LLC and 2,200 MHz peak engine clock per AMD official product page.
- Form factor: AMD's official product page lists "PCIe Add-in Card", "PCIe 5.0 x16", "Passive" cooling, and "600W TBP (Max); 450W TBP configurable". The more specific full-height full-length (FHFL) dual-slot physical description comes from ServeTheHome ("a very standard and by-the-books full height full length (FHFL) dual-slot card") and StorageReview ("a dual-slot, full-height, full-length design"), not from the AMD page. Likewise, the absence of GPU-to-GPU Infinity Fabric is not stated on the AMD product page - ServeTheHome reports "AMD is not exposing its GPU-to-GPU Infinity Fabric links in any way on the MI350P. As a result, multi-card setups are limited to just using the PCIe bus (PCIe Gen5 x16)", and StorageReview concurs that "all collective communications go through the PCIe Gen5 x16 (128 GB/s) link".
- TF32 and INT4 are excluded from supportedHardwareOperations. CDNA 4 supports TF32 only through software emulation (same as MI350X), not native hardware acceleration. INT4 is not listed as a hardware-supported precision on the AMD official product page.
- MSRP is null because AMD does not publish official MSRP for datacenter GPU accelerators. The MI350P is sold through OEM and channel system partnerships rather than as a standalone retail card, and no verified OEM list prices were available at time of research. Re-checked 2026-09-17: HPE lists the card as an orderable server option ("AMD Instinct MI350P 144GB PCIe Accelerator for HPE", product ID S8A40C) but publishes no price, and no second OEM list price (Dell, Lenovo, Cisco GPL) was found, so no estimated MSRP is recorded rather than extrapolating from a single unpriced listing.