Algorithm and kernel-pipeline co-design adapts attention to Blackwell’s compute and bandwidth characteristics.
AI technology timeline
Entries by year
Release history
2026
1 entry2025
3 entries2024
4 entriesAttention kernels are redesigned around Hopper GPUs’ asynchronous execution and low-precision features.
Packages isolated code execution into an SDK so agents can run generated programs and inspect their results.
Group-relative rewards estimate advantages without a separate value model.
2023
12 entriesAn inference library for large language models on NVIDIA GPUs becomes publicly available.
Revised work partitioning improves GPU parallelism for attention computation.
Paged KV-cache management improves memory utilization in language-model serving.
Groups of query heads share key-value heads to balance quality and KV-cache cost.
torch.compile adds compilation-based acceleration to the established PyTorch workflow.
2022
8 entriesA small model drafts tokens and a target model verifies them, accelerating generation while preserving the target distribution.
Vector fields along conditional probability paths train continuous-flow generative models.
Reduced data movement between GPU memory and on-chip storage accelerates exact attention.
A fixed training-compute budget motivates a different balance of model size and data.
Multiple sampled reasoning paths are aggregated to improve answer consistency.
Intermediate reasoning steps in examples improve multi-step language-model tasks.
2021
5 entriesDiffusion in a compressed latent space reduces the cost of high-resolution image generation.
A Python-like language and compiler simplify the development of optimized GPU kernels.
2020
6 entriesPartitioning optimizer state and other training data reduces distributed-training memory use.
A systematic study relates language-model performance to model size, data, and compute.
2019
3 entriesA unified API makes Transformer architectures and pretrained models accessible for research and applications.
A Rust-based virtual machine monitor provides a compact VM layer for modern cloud workloads.
2018
6 entriesAWS releases lightweight microVM technology to isolate short-lived, multi-tenant workloads with separate kernels.
Subword models train directly on raw text without language-specific word segmentation.
A container runtime uses lightweight VMs to give containers or pods their own guest kernel.
2017
6 entriesHuman comparisons train a reward signal that guides reinforcement learning.
Google researchers introduce the Transformer, an attention-based sequence model, in Attention Is All You Need.
Sparse routing selects a small subset of experts to increase capacity while limiting computation.
2016
1 entryNormalization statistics are computed within each example rather than across a batch.
2015
4 entriesByte-pair encoding is applied to subword segmentation to address unknown words.
2014
5 entriesAn encoder and decoder map variable-length input sequences to variable-length outputs.
A translation model learns to attend to different input positions as it generates each word.
2013
3 entriesVariational inference and reparameterization learn a latent representation that can be sampled.
2012
1 entry2007
1 entryAll 69 entries shown