Selection criteria
Model Atlas selects public events across models, agents, hardware, and AI technology. Every entry must qualify as an established classic, a representative technical development, or an event with verifiable industry or research influence. Reliable dates and sources are required in all cases.
- Technical significance
- Prioritize influential foundational methods, widely used platforms, and major generations. Recent entries need a clear change in capabilities, architecture, or use, rather than attention alone.
- Release significance
- Include debuts, major generations, and substantial capability updates. Group same-day family variants, reasoning settings, and platform copies.
- Verifiable sources
- Each entry needs a verifiable event date and source. Public evaluations, papers, technical documentation, and authoritative reporting are accepted, with cross-checks for ambiguity. Company recognition or a leaderboard listing alone does not qualify.
Organization order
Organizations on the model timeline follow the Artificial Analysis Intelligence Index, using each organization’s highest-scoring configuration in the snapshot’s Current list. Displayed ties retain source order; unscored organizations appear last. The representative need not be the newest release.
Scores marked * are AA estimates. The order is a manually maintained snapshot, without combining scores from other evaluations. Other timelines retain curation order for organizations and projects.
Artificial Analysis · Leaderboard checked
Agents and harnesses
The Agent timeline covers coding assistants, general-purpose task tools, execution frameworks, and related protocols. Entries must execute multi-step tasks or provide essential capabilities such as tool use and state management.
Agent harness
The tools and control system around a model: tool calls, context, session state, and the task loop. It may also handle permissions, filesystem access, and coordination between agents.
Models and tools are recorded separately. Product debuts, public previews, framework announcements, and major versions retain their respective event dates.
AI hardware
The hardware timeline covers data-center GPUs, specialized accelerators, compute systems, and local or edge devices. It prioritizes major architectures, representative chip generations, and deployment platforms rather than every vendor product.
Announcements, hardware demonstrations, availability, and deployment records are labeled separately. Cloud-service and SuperPoD events do not imply a chip launch or mass-production date. Future roadmaps and planned availability remain in related notes, rather than being counted as completed launches.
Hardware is assessed through architectural contributions, model-specific deployments, training or inference services, and research use. New architecture announcements can qualify on documented technical significance; availability and adoption are described separately. Roadmaps, theoretical peaks, and planned orders are labeled as such.
Specifications and prices
Compute chips identify precision, dense or sparse arithmetic, and the exact configuration. TFLOPS / PFLOPS express floating-point rates. TOPS means trillions of operations per second and requires a precision such as INT8 or FP4. Theoretical peaks do not equal actual tokens/s. Per-chip, board, and system figures are not directly comparable. Launch-time estimates are labeled.
Memory, bandwidth, architecture, and power apply to the named model and configuration. Per-chip, dual-chip board, system, and rack figures are distinct, as are HBM, on-chip SRAM, and system unified memory.
Verified prices are labeled as US starting prices at announcement and provide historical context. Peak compute depends on precision, sparsity, and configuration, so these figures are not combined into a ranking. MLPerf results provide workload- and system-specific performance evidence.
AI technology
Research milestones and AI engineering share one timeline, grouped into methods, frameworks and compute, inference and deployment, and containers and sandboxes. Papers document how methods emerge; engineering projects show how they enter training and deployment workflows.
Selection focuses on influential methods, project debuts, and architectural changes. It does not attempt to list every paper, patch, or integration. Coding agents, harnesses, and tool protocols remain on the Agent timeline.
Filters group entries by publisher, project, or representative research institution for browsing, rather than replacing full author attribution. Transformer shares one canonical record across the LLM and Technology timelines, with sources accessible from either page.
Containers & sandboxes
Selected projects have documented ties to AI engineering: Docker packages environments; Kata connects VM isolation to container runtimes; Cloud Hypervisor and Firecracker provide the VM layer. gVisor uses a user-space kernel, while E2B provides a code-execution API. Original releases and later agent adoption are distinguished. Related components such as containerd and runc do not each receive a separate entry.
Dates and specifications
Event dates and counts
Papers use their first public date; models and tools use the official release, preview, or weight-release date. Details explain date differences. Foundational research such as Transformer is included as a research-paper entry.
Charts count the releases included on this site. Select a year for the monthly distribution, then a month to filter entries. Organization, keyword, and other filters apply together. Months after the research cutoff are shown with a dash.
Where a source confirms only a year, no month or day is invented. These entries count toward annual totals but are excluded from monthly bins, with a note on the chart. Each page counts its own entries; cross-listed events are not added into a global unique-release total.
Model types
Models are classified by input modality as Text or Multimodal. Multimodal means the model directly accepts non-text inputs such as images, audio, or video. External OCR, transcription, or vision tools do not change the underlying model’s classification.
Classification follows the recorded version or release scope. A family containing multimodal variants is classified as Multimodal, with variant differences explained in the details. Later product upgrades do not automatically alter earlier snapshots. General use, coding, reasoning, and agentic capabilities overlap and remain descriptive tags.
Open and closed source
This site uses a broad classification for browsing. Models with official weights, or an explicitly documented public base version, are labeled open source. Tools with public core code or protocol specifications are also included. Entries available only as hosted products or APIs are labeled closed source.
Access is recorded for the specific version at the check date. A later weight release does not change the original launch date. The open-source category includes open weights; applicable licenses, research approval, or commercial restrictions are explained in the details.
Context and specifications
Specifications come from official documentation, model cards, and papers and apply to the named version. Context, input, and output use tokens; input and output limits are not additive. Native and configured extended windows are listed separately. Unverified fields are omitted.
Parameter counts use M (million), B (billion), and T (trillion). MoE models distinguish total from active parameters. An agent tool’s usable window depends on its model and runtime configuration.
API pricing
Prices use the published rate per million tokens, separating uncached text input from text output and identifying the currency, model, and check date. Labels identify the priced variant and its standard tier; other variants and long-context tiers are available in the details.
Caching, batch processing, tools, and audio or video follow each service’s separate pricing rules. Open weights do not imply free hosted inference. A documented price for a historical model does not imply current availability.
The annual chart and year navigation omit years without entries, with an asterisk note; monthly charts retain all months. Counts reflect this curated collection, not all industry releases. Year-only records count annually but are not assigned to a month.
Evaluations and sources
Public leaderboards help identify candidates and verify evaluation results. Dates, specifications, and deployment records may come from official material, papers, authoritative reporting, or credible industry research; event and publication dates are distinguished. Licenses and API rates primarily follow license files and service documentation, with announced prices distinguished from availability.
Model scores
Cards show reviewed AA Intelligence Index and Arena Text Overall (Style Control) scores. AA aggregates benchmark tasks; Arena measures text-chat preferences. Open a chip for the exact model, reasoning configuration, source, and snapshot date. The scales are not interchangeable.
Only unambiguous version matches are linked. Family-entry labels identify the evaluated variant; unavailable or ambiguous results are omitted. An asterisk marks an AA estimate or an Arena preliminary result, explained in the details. Scores are snapshots, not launch-time or live results.
Model evaluations
Agent evaluations
Task sets, model versions, reasoning budgets, and environments differ across evaluations, so scores are not directly interchangeable. Agent results depend on the model, harness, and tool configuration together.
Hardware evaluations
Maintenance
Data is manually reviewed, with quality taking priority over coverage. Routine revisions, duplicate listings, renaming alone, and marketing events do not receive separate entries. Candidates without sufficient evidence are omitted. Each entry retains its sources and relevant date notes.
This is a static website. Search, filters, and charts run in the browser. The home page initially follows browser language preferences. Each language has its own URLs; manual choices are remembered, and a language specified in a shared link takes priority.
Release records, organization order, specifications, prices, and evaluations retain their check dates. Structured public facts, source links, and selection decisions are saved in local snapshots for reuse, with follow-up research focused on new or changed information. Scores and prices are not synchronized in real time.
Release coverage through