AI technology timeline

69 entries32 organizations / projects

Entries by year

Release history

Organizations / projects
Selection criteria

69 entries

2026

1 entry

FlashAttention-4

FlashAttention

Algorithm and kernel-pipeline co-design adapts attention to Blackwell’s compute and bandwidth characteristics.

Research paperInference & deployment
Source

2025

3 entries

AMD previews ROCm 7 with updated software support for Instinct accelerators.

Public previewFrameworks & compute
Source

An open-source framework coordinates distributed model serving across GPUs.

AnnouncementInference & deployment
Source

A rebuilt inference core simplifies scheduling and improves prefix caching and execution.

Public previewInference & deployment
Source

2024

4 entries

FlashAttention-3

FlashAttention

Attention kernels are redesigned around Hopper GPUs’ asynchronous execution and low-precision features.

Research paperInference & deployment
Source

MLA · DeepSeek-V2

DeepSeekMilestone

Low-rank joint compression reduces the KV-cache requirements of attention.

Research paperMethods
Source

Packages isolated code execution into an SDK so agents can run generated programs and inspect their results.

Software releaseContainers & sandboxes
Source

GRPO · DeepSeekMath

DeepSeekMilestone

Group-relative rewards estimate advantages without a separate value model.

Research paperMethods
Source

2023

12 entries

SGLang

SGLangMilestone

A structured generation language and runtime reuse prefix computation and cached state.

Research paperInference & deployment
Source

MLX

Apple

An array framework brings automatic differentiation and unified-memory execution to Apple silicon.

Software releaseFrameworks & compute
Source

Mamba

State SpacesMilestone

Selective state-space models offer an alternative approach to sequence modeling.

Research paperMethods
Source

An inference library for large language models on NVIDIA GPUs becomes publicly available.

Software releaseInference & deployment
Source

FlashAttention-2

FlashAttention

Revised work partitioning improves GPU parallelism for attention computation.

Research paperInference & deployment
Source

Paged KV-cache management improves memory utilization in language-model serving.

Software releaseInference & deployment
Source

AWQ

MIT

Activation statistics protect salient weights for low-bit model deployment.

Research paperInference & deployment
Source

DPO

StanfordMilestone

Direct optimization on preference data simplifies language-model alignment.

Research paperMethods
Source

QLoRA

U. WashingtonMilestone

Low-bit quantization and low-rank adaptation reduce memory requirements for model fine-tuning.

Research paperMethods
Source

PyTorch 2.0

PyTorchMilestone

torch.compile adds compilation-based acceleration to the established PyTorch workflow.

Software releaseFrameworks & compute
Source

llama.cpp

ggmlMilestone

A C/C++ implementation runs LLaMA locally with quantized inference.

Software releaseInference & deployment
Source

2022

8 entries

Speculative Decoding

GoogleMilestone

A small model drafts tokens and a target model verifies them, accelerating generation while preserving the target distribution.

Research paperInference & deployment
Source

GPTQ

ISTA

Approximate second-order information supports post-training weight quantization for lower-memory inference.

Research paperInference & deployment
Source

Flow Matching

MetaMilestone

Vector fields along conditional probability paths train continuous-flow generative models.

Research paperMethods
Source

ReAct

GoogleMilestone

Reasoning steps and tool actions are interleaved so language models can adapt their plans to environment feedback.

Research paperMethods
Source

FlashAttention

FlashAttentionMilestone

Reduced data movement between GPU memory and on-chip storage accelerates exact attention.

Research paperInference & deployment
Source

Multiple sampled reasoning paths are aggregated to improve answer consistency.

Research paperMethods
Source

Intermediate reasoning steps in examples improve multi-step language-model tasks.

Research paperMethods
Source

2021

5 entries

Latent Diffusion

CompVisMilestone

Diffusion in a compressed latent space reduces the cost of high-resolution image generation.

Research paperMethods
Source

Triton 1.0

OpenAIMilestone

A Python-like language and compiler simplify the development of optimized GPU kernels.

Software releaseFrameworks & compute
Source

LoRA

MicrosoftMilestone

Frozen base weights and trainable low-rank matrices reduce the cost of model adaptation.

Research paperMethods
Source

RoPE · RoFormer

ZhuiyiMilestone

Rotations encode position within attention computation.

Research paperMethods
Source

CLIP

OpenAIMilestone

Image–text contrastive pretraining enables visual concepts to be recognized from natural-language descriptions.

AnnouncementMethods
Source

2020

6 entries

Vision Transformer

GoogleMilestone

Images become sequences of patches for Transformer-based visual learning.

Research paperMethods
Source

CANN 3.0

Huawei

An updated Ascend software stack covers operator development, model compilation, and execution.

Software releaseFrameworks & compute
Source

DDPM

UC BerkeleyMilestone

Iterative denoising provides a training approach for diffusion-based image generation.

Research paperMethods
Source

RAG

MetaMilestone

Document retrieval supplies external knowledge to a text-generation model.

Research paperMethods
Source

DeepSpeed · ZeRO

MicrosoftMilestone

Partitioning optimizer state and other training data reduces distributed-training memory use.

Software releaseFrameworks & compute
Source

Neural Scaling Laws

OpenAIMilestone

A systematic study relates language-model performance to model size, data, and compute.

Research paperMethods
Source

2019

3 entries

Hugging Face Transformers

Hugging FaceMilestone

A unified API makes Transformer architectures and pretrained models accessible for research and applications.

Research paperFrameworks & compute
Source

Megatron-LM

NVIDIAMilestone

Intra-layer model parallelism distributes large Transformer training across GPUs.

Research paperFrameworks & compute
Source

Cloud Hypervisor v0.1.0

Cloud Hypervisor

A Rust-based virtual machine monitor provides a compact VM layer for modern cloud workloads.

Public previewContainers & sandboxes
Source

2018

6 entries

JAX

Google

NumPy-style programming combines with automatic differentiation and accelerator compilation.

Software releaseFrameworks & compute
Source

ONNX Runtime

Microsoft

A cross-platform runtime connects model formats to execution backends.

Software releaseInference & deployment
Source

AWS releases lightweight microVM technology to isolate short-lived, multi-tenant workloads with separate kernels.

AnnouncementContainers & sandboxes
Source

Subword models train directly on raw text without language-specific word segmentation.

Research paperFrameworks & compute
Source

Kata Containers 1.0

Kata Containers

A container runtime uses lightweight VMs to give containers or pods their own guest kernel.

Software releaseContainers & sandboxes
Source

gVisor

Google

Google introduces a user-space kernel sandbox that adds isolation for untrusted code in containers.

AnnouncementContainers & sandboxes
Source

2017

6 entries

ONNX

Microsoft

Microsoft and Facebook introduce a common model representation for framework interoperability.

AnnouncementInference & deployment
Source

PPO

OpenAI

Constrained policy updates offer a practical approach to stable reinforcement learning.

Research paperMethods
Source

Transformer

GoogleMilestone

Google researchers introduce the Transformer, an attention-based sequence model, in Attention Is All You Need.

Research paperMethods
Source

Sparsely-Gated MoE

GoogleMilestone

Sparse routing selects a small subset of experts to increase capacity while limiting computation.

Research paperMethods
Source

PyTorch

PyTorchMilestone

Dynamic computation graphs and a Python workflow support deep-learning research.

Software releaseFrameworks & compute
Source

2016

1 entry

Layer Normalization

U. TorontoMilestone

Normalization statistics are computed within each example rather than across a batch.

Research paperMethods
Source

2015

4 entries

ResNet

MicrosoftMilestone

Residual connections make very deep networks easier to optimize and train.

Research paperMethods
Source

TensorFlow

GoogleMilestone

Google releases its machine-learning framework as open-source software.

Software releaseFrameworks & compute
Source

BPE · Subword Tokenization

U. EdinburghMilestone

Byte-pair encoding is applied to subword segmentation to address unknown words.

Research paperMethods
Source

Knowledge Distillation

GoogleMilestone

A teacher model’s output distribution trains a smaller student model.

Research paperMethods
Source

2014

5 entries

Adam

U. AmsterdamMilestone

First- and second-moment gradient estimates adapt optimization step sizes.

Research paperMethods
Source

Sequence to Sequence

GoogleMilestone

An encoder and decoder map variable-length input sequences to variable-length outputs.

Research paperMethods
Source

cuDNN

NVIDIAMilestone

A GPU deep-learning primitive library provides optimized operations for frameworks.

AnnouncementFrameworks & compute
Source

Additive Attention

U. MontréalMilestone

A translation model learns to attend to different input positions as it generates each word.

Research paperMethods
Source

GAN

U. MontréalMilestone

Adversarial training between a generator and a discriminator learns a data distribution.

Research paperMethods
Source

2013

3 entries

Variational Autoencoder · VAE

U. AmsterdamMilestone

Variational inference and reparameterization learn a latent representation that can be sampled.

Research paperMethods
Source

Docker

Docker

Image-based application packaging makes reproducible environments easier to build, share, and run.

AnnouncementContainers & sandboxes
Source

word2vec

GoogleMilestone

Efficient word-vector learning represents relationships in a continuous space.

Research paperMethods
Source

2012

1 entry

AlexNet

U. TorontoMilestone

A deep convolutional network trained on GPUs advances ImageNet image classification.

Research paperMethods
Source

2007

1 entry

CUDA SDK

NVIDIAMilestone

A public beta of the CUDA toolkit and SDK brings C-based programming to GPU computing.

Public previewFrameworks & compute
Source

All 69 entries shown

Search names, organizations, or keywords across all timelines.