AI hardware timeline

54 entries10 manufacturers

Entries by year

Release history

Manufacturers
Selection criteria

54 entries

2026

5 entries

A 1,024-card SuperPoD is publicly demonstrated with unified addressing and high-speed links.

Public demonstrationCompute systems
Source

Eighth-generation TPUs split into training-oriented 8t and inference-oriented 8i.

AnnouncementCustom accelerators
Source

950PR launches in the Atlas 350 accelerator card for recommendation and LLM inference.

AvailabilityCustom accelerators
Source

NVIDIA Vera Rubin

NVIDIAMilestone

Rubin GPUs, Vera CPUs, and new interconnects form a new AI compute platform.

AnnouncementData-center GPUs
Source

2025

8 entries

Third-generation Trainium becomes available through Trn3 UltraServers for training and serving.

AvailabilityCustom accelerators
Source

GB10 and unified memory move into a desktop system for local model development.

AvailabilityLocal & edge
Source

CDNA 4 expands low-precision compute and HBM3e capacity for generative AI.

AnnouncementData-center GPUs
Source

Google TPU · Ironwood

GoogleMilestone

Seventh-generation TPUs target large-scale AI serving with closer chip and system integration.

AnnouncementCustom accelerators
Source

B300 and GB300 platforms expand memory and inference compute for longer reasoning workloads.

AnnouncementData-center GPUs
Source

Apple M3 Ultra

AppleMilestone

Larger unified-memory configurations increase capacity for running large models locally.

AnnouncementLocal & edge
Source

GeForce RTX 5090

NVIDIAMilestone

The Blackwell consumer flagship expands to 32 GB of GDDR7 for local AI and graphics workloads.

AnnouncementLocal & edge
Source

2024

8 entries

A developer kit combines a software upgrade and lower pricing for edge generative AI.

AvailabilityLocal & edge
Source

Trainium2 becomes available through Trn2, expanding custom compute for training and inference.

AvailabilityCustom accelerators
Source

The MI300 line gains HBM3e for greater generative-AI memory capacity and bandwidth.

AnnouncementData-center GPUs
Source

Sixth-generation TPUs expand compute, HBM, and interconnect capacity for generative models.

AnnouncementCustom accelerators
Source

Meta’s custom accelerator serves production ranking and recommendation models in its data centers.

AnnouncementCustom accelerators
Source

Gaudi 3 combines matrix compute, HBM, and Ethernet links for large-model training and inference.

AnnouncementCustom accelerators
Source

Blackwell combines a new GPU generation with Grace CPUs and NVLink systems for model training and inference.

AnnouncementData-center GPUs
Source

Cerebras WSE-3 / CS-3

CerebrasMilestone

The third wafer-scale engine integrates a large compute array and on-chip memory on one wafer.

AnnouncementCustom accelerators
Source

2023

7 entries

CDNA 3 and 192 GB of HBM3 expand large-model capacity per accelerator.

AvailabilityData-center GPUs
Source

A fifth-generation TPU focused on large-model training is announced with AI Hypercomputer.

AnnouncementCustom accelerators
Source

Hopper gains HBM3e to increase memory capacity and bandwidth for large-model inference.

AnnouncementData-center GPUs
Source

The cost-focused fifth-generation TPU supports both training and inference.

AnnouncementCustom accelerators
Source

Powers the Spark all-in-one system introduced by iFLYTEK and Huawei for enterprise LLM deployment.

Deployment recordCustom accelerators
Source

Second-generation inference chips power EC2 Inf2 for generative-AI deployment.

AvailabilityCustom accelerators
Source

NVIDIA L4

NVIDIA

A low-power Ada GPU succeeds T4 for generative AI and video inference.

AnnouncementData-center GPUs
Source

2022

5 entries

First-generation Trainium becomes generally available through EC2 Trn1 for model training.

AvailabilityCustom accelerators
Source

Ada and 24 GB of memory support local inference, image generation, and model development.

AnnouncementLocal & edge
Source

The second Gaudi generation moves to 7 nm while retaining Ethernet-based training scale-out.

AnnouncementCustom accelerators
Source

NVIDIA H100

NVIDIAMilestone

Hopper introduces the Transformer Engine and FP8 for large-model training and inference.

AnnouncementData-center GPUs
Source

A wafer-on-wafer power-delivery design advances Graphcore’s distributed IPU architecture.

AvailabilityCustom accelerators
Source

2021

3 entries

CDNA 2 uses a multi-die design to expand memory capacity for training and scientific computing.

AnnouncementData-center GPUs
Source

Fourth-generation TPUs expand pod-scale model training.

AnnouncementCustom accelerators
Source

The second wafer-scale engine uses a 7 nm process to expand compute and on-chip memory.

AnnouncementCustom accelerators
Source

2020

5 entries

AMD Instinct MI100

AMDMilestone

The first CDNA architecture targets compute with ROCm support for AI and HPC.

AnnouncementData-center GPUs
Source

Apple M1

AppleMilestone

Unified memory and the Neural Engine come to the Mac as a platform for local machine learning.

AnnouncementLocal & edge
Source

24 GB of memory and Ampere Tensor Cores expand capacity for local model experiments.

AnnouncementLocal & edge
Source

Second-generation IPUs use distributed on-chip memory and parallel execution for machine learning.

AnnouncementCustom accelerators
Source

NVIDIA A100

NVIDIAMilestone

Ampere unifies training and inference and introduces Multi-Instance GPU.

AvailabilityData-center GPUs
Source

2019

4 entries

First-generation Inferentia becomes available for machine-learning inference through EC2 Inf1.

AvailabilityCustom accelerators
Source

Ascend 910

HuaweiMilestone

Ascend 910 launches as a dedicated processor for neural-network training.

AnnouncementCustom accelerators
Source

Cerebras WSE

CerebrasMilestone

The first wafer-scale engine integrates compute, memory, and interconnect on one large chip.

AnnouncementCustom accelerators
Source

Habana Gaudi

IntelMilestone

An Ethernet-connected architecture targets scalable neural-network training.

AnnouncementCustom accelerators
Source

2018

4 entries

Ascend 310 / 910

HuaweiMilestone

Huawei announces the Ascend chip family and Da Vinci architecture for inference and training.

AnnouncementCustom accelerators
Source

NVIDIA Tesla T4

NVIDIAMilestone

Turing Tensor Cores bring low-precision compute to data-center inference.

AnnouncementData-center GPUs
Source

Third-generation TPUs expand training scale with liquid-cooled pods.

AnnouncementCustom accelerators
Source

2017

2 entries

Cloud TPU · v2

GoogleMilestone

Second-generation TPUs add training support and a route to external access through Cloud TPU.

AnnouncementCustom accelerators
Source

Tesla V100

NVIDIAMilestone

Volta introduces Tensor Cores to accelerate matrix operations for deep learning.

AnnouncementData-center GPUs
Source

2016

2 entries

Google TPU

GoogleMilestone

Google describes its custom tensor processor for neural-network inference.

AnnouncementCustom accelerators
Source

Tesla P100

NVIDIAMilestone

Pascal combines HBM2 and NVLink for deep learning and multi-GPU computing.

AnnouncementData-center GPUs
Source

2014

1 entry

Tesla K80

NVIDIA

A dual-GPU accelerator expands memory and throughput for machine learning and scientific computing.

AnnouncementData-center GPUs
Source

All 54 entries shown

Search names, organizations, or keywords across all timelines.