Gemini 4 Argon
A new multimodal model for complex coding, knowledge work, and cyber defense, initially available to invited teams.
SourceModels, agents, hardware and technology, in context.
20072026327 events
A new multimodal model for complex coding, knowledge work, and cyber defense, initially available to invited teams.
SourceImproves coding, computer use, and professional work, with task delegation available within a request.
SourceThe second Claude 5.5 model supports coding, document processing, and office tasks.
SourceOpus targets complex knowledge work, coding, and judgment-intensive tasks.
SourceTwo GPT-6 models address different speed and cost requirements.
SourceMiMo combines multimodal understanding with large-scale reinforcement learning.
SourceA flagship preview targets long-running engineering, knowledge work, and research.
SourceText, images, audio, and video share one multimodal model.
SourceCowork and chat merge into one Claude experience for documents and ongoing tasks.
SourceKimi Code gains a stronger coding and agent preview.
SourceA new architecture adds native vision and more efficient agent reasoning.
SourceGPT combines reasoning, computer use, and multi-step professional work.
SourceFlash improves sustained coding and professional reasoning.
SourceMuse improves long-running agents, coding, and continued collaboration.
SourceImproves engineering projects, multi-tool orchestration, chart reasoning, and document parsing.
SourceFable and Mythos improve coding, knowledge work, and research.
SourceTencent opens a flagship preview for long-context and complex work.
SourceNative vision and hybrid linear/sparse attention support coding, browser, and GUI tasks.
SourceAn efficient multimodal model supports coding and high-volume agent use.
SourcePlugins compose models, tools, sessions, sandboxes and execution loops, with a recorded task trajectory.
SourceV4 Pro reaches general availability with stronger agents and Responses API support.
SourceFlash improves multi-step coding, knowledge work, and web development.
SourceGrok improves sustained work across agents, interactive apps, and vision.
SourceQwen opens 2.4T-A95B weights, followed by a 27B model.
SourceA native vision-language flagship supports a one-million-token context.
SourceOpus upgrades engineering, knowledge work, and computer use.
SourceMI400 GPUs and Helios rack systems extend AMD large-scale AI infrastructure.
SourceFlash improves reasoning efficiency alongside a lighter Flash-Lite model.
SourceA 1,024-card SuperPoD is publicly demonstrated with unified addressing and high-speed links.
SourceSol, Terra, and Luna address different knowledge-work and coding needs.
SourceMuse improves coding, computer use, and multimodal understanding.
SourceTencent improves long-running agents and product integration.
SourceSonnet expands planning, browser use, terminal work, and sustained execution.
SourceKimi improves long-running coding, tool calls, and efficient thinking.
SourceFable and Mythos target demanding coding and knowledge work.
SourceSparse attention adds native image and video understanding.
SourceStep improves efficient reasoning and practical agent execution.
SourceOpus improves tool use, complex judgments, and collaboration.
SourceA public preview targets sustained coding and remote agent tasks.
SourceQwen improves agentic coding, office work, and long-running execution.
SourceGoogle I/O introduces a Flash update focused on speed and agent tasks.
SourcePro and Flash previews introduce million-token context and open weights.
SourceGPT improves tasks spanning multiple tools and applications.
SourceMiMo upgrades complex tasks, multimodal work, and inference efficiency.
SourceEighth-generation TPUs split into training-oriented 8t and inference-oriented 8i.
SourceA flagship preview strengthens coding and complex task execution.
SourceOpus improves difficult engineering and visual understanding.
SourceMeta introduces a new family combining multimodality, reasoning, and tools.
Source950PR launches in the Atlas 350 accelerator card for recommendation and LLM inference.
SourceMiMo expands flagship reasoning and multimodal understanding.
SourceSoftware engineering and professional work gain stronger agent collaboration.
SourceOne model combines reasoning, vision, and coding with adjustable thinking.
SourceHybrid experts and long context support multi-agent work.
SourceReasoning, coding, and native computer use target professional work.
SourceAlgorithm and kernel-pipeline co-design adapts attention to Blackwell’s compute and bandwidth characteristics.
SourceA persistent agent combines memory, reusable skills and tools to carry out tasks locally or on a server.
SourceGemini strengthens reasoning for scientific, engineering, and complex tasks.
SourceSonnet improves coding, visual interaction, planning, and knowledge work.
SourceGLM moves from individual coding tasks toward broader software engineering.
SourceA smaller model previews rapid, interactive coding.
SourceReinforcement learning targets practical coding, search, office work, and tools.
SourceOpus expands complex codebase analysis and long-running agent tasks.
SourceCoding and general reasoning combine in longer engineering workflows.
SourceA desktop app organizes parallel coding agents and code review.
SourceSparse mixture-of-experts modeling balances speed with agent capabilities.
SourceA lightweight model extends the GLM-4.7 family.
SourceClaude Code-style execution expands to files, documents, and office work.
SourceRubin GPUs, Vera CPUs, and new interconnects form a new AI compute platform.
SourceMiniMax improves programming across languages and practical office tasks.
SourceGemini 3 reasoning reaches a faster, more efficient Flash tier.
SourceXiaomi opens an efficient model for reasoning, coding, and agent tasks.
SourceNVIDIA announces Nemotron 3 and releases Nano first.
SourceHandles software tasks asynchronously in isolated environments, retaining context from project history and review.
SourceThird-generation Trainium becomes available through Trn3 UltraServers for training and serving.
SourceMistral expands its open-weight multimodal model family.
SourceConfigurable extended thinking supports reasoning, coding, and agent workflows.
SourceDeepSeek advances efficient attention, reasoning, and agentic tool use.
SourceA minimal coding harness emphasizes tools, context control, and session management.
SourceOpus improves coding, computer use, and long-running work.
SourceAgents work across the editor, terminal and browser, presenting artifacts that document progress and verification.
SourceGoogle introduces a new generation of multimodal reasoning models.
SourceKiro reaches general availability and introduces a CLI for specification-driven development in the terminal.
SourceKimi extends the K2 family with deeper reasoning and tool-assisted work.
SourceAn open-source coding agent connects to different models to read files, edit code and use project tools.
SourceMiniMax opens a model designed around coding and agent workflows.
SourceOrganizes agents around roles, tasks and flows as its open-source core reaches general availability.
SourceInstructions, scripts, and resources package reusable, on-demand agent expertise.
SourceHandles research, analysis, presentations and website development, including full-stack web applications.
SourceA smaller Claude model improves coding and computer use at lower latency.
SourceGB10 and unified memory move into a desktop system for local model development.
SourceClaude Code SDK becomes Claude Agent SDK, expanding its focus to general-purpose agents.
SourceSonnet advances coding, computer use, and sustained agentic tasks.
SourceHybrid attention and highly sparse activation target efficient long-context processing.
SourcePlans and executes development tasks with repository context; Quest mode supports longer delegated work.
SourceThinking and non-thinking modes are unified with stronger agent capabilities.
SourceOpus improves coding, practical reasoning, and agent performance.
SourceOpenAI releases two open-weight reasoning models.
SourceA reusable harness combines planning, a filesystem, and subagents for longer tasks.
SourceAn open-source command-line coding agent released alongside Qwen3-Coder for repository work and task execution.
SourceAn open mixture-of-experts model targets agentic coding and long tasks.
SourceResearch and action come together through a virtual computer, browser, and terminal.
SourceConnects requirements, coding, browser validation and deployment within one development task.
SourceGoogle releases an open-source terminal agent for coding, troubleshooting, and tools.
SourceHybrid attention explores efficient reasoning and long-context work.
SourceCDNA 4 expands low-precision compute and HBM3e capacity for generative AI.
SourceAMD previews ROCm 7 with updated software support for Instinct accelerators.
SourceA major R1 update improves reasoning, coding, structured output, and tool calls.
SourceOpus 4 and Sonnet 4 advance coding and longer-running work.
SourceAssigned issues become background coding tasks and reviewable pull requests.
SourceCloud sandboxes support parallel engineering tasks and reviewable code changes.
SourceDefine an agent with a model, instructions and tools, letting the model choose the next steps.
SourceAn enterprise model balances multimodal capability, coding, and deployment cost.
SourceAn open-source terminal coding agent works directly with local code and tools.
SourceReasoning models learn to combine search, code, and visual tools.
Source384 Ascend 910C processors are connected into a unified compute system.
SourceAn open protocol supports agent discovery, task delegation and status updates across vendors.
SourceTools for building, orchestrating, evaluating and deploying agents, including multi-agent applications.
SourceSeventh-generation TPUs target large-scale AI serving with closer chip and system integration.
SourceThinking becomes central to complex reasoning, coding, and multimodal tasks.
SourceTencent introduces a deep-reasoning model trained with reinforcement learning.
SourceB300 and GB300 platforms expand memory and inference compute for longer reasoning workloads.
SourceAn open-source framework coordinates distributed model serving across GPUs.
SourceAgent orchestration adds handoffs, guardrails, and tracing around model and tool calls.
SourceLarger unified-memory configurations increase capacity for running large models locally.
SourceA terminal agent reads code, edits files, runs tests, and uses command-line tools.
SourceOne model combines fast responses with extended thinking.
SourcexAI describes reasoning capabilities and the DeepSearch experience.
SourceMulti-step web research analyzes sources and produces cited reports.
SourceEfficient reasoning targets mathematics, science, and programming.
SourceDeepSeek releases a reasoning model and distilled variants with downloadable weights.
SourceThe Blackwell consumer flagship expands to 32 GB of GDDR7 for local AI and graphics workloads.
SourceModels express actions as code and use tools and execution environments to complete multi-step tasks.
SourceOpen weights expand large-scale mixture-of-experts modeling.
SourceA developer kit combines a software upgrade and lower pricing for edge generative AI.
SourceAn experimental Flash model combines multimodality with native tool use.
SourceTrainium2 becomes available through Trn2, expanding custom compute for training and inference.
SourceAmazon introduces text and multimodal models across several cost tiers.
SourceAn open protocol standardizes connections between models, data, and tools.
SourceComposer gains an early agent that retrieves context and uses the terminal.
SourceAn agentic editor combines project context with multi-step code changes.
SourceThe MI300 line gains HBM3e for greater generative-AI memory capacity and bandwidth.
SourceClaude Dev becomes Cline, with streamed edits, mid-task feedback and broader model support.
SourceNatural-language app building combines setup, coding, and execution.
SourceReasoning-time computation becomes a new route to stronger task performance.
SourceA 123B model expands coding, multilingual, and long-context capabilities.
SourceAn execution platform connects coding agents to files, terminals, and browsers.
SourceA 405B model extends the family with longer context, multilingual support, and tool use.
SourceA lower-cost text and vision model expands lightweight applications.
SourceAttention kernels are redesigned around Hopper GPUs’ asynchronous execution and low-precision features.
SourceSonnet improves coding and visual understanding alongside the introduction of Artifacts.
SourceSixth-generation TPUs expand compute, HBM, and interconnect capacity for generative models.
SourceLow-rank joint compression reduces the KV-cache requirements of attention.
SourceA purpose-built computer interface lets agents navigate code, edit files, and run tests.
SourceMixture-of-experts and latent attention improve training and inference efficiency.
SourcePackages isolated code execution into an SDK so agents can run generated programs and inspect their results.
SourceMeta’s custom accelerator serves production ranking and recommendation models in its data centers.
SourceGaudi 3 combines matrix compute, HBM, and Ethernet links for large-model training and inference.
SourceBlackwell combines a new GPU generation with Grace CPUs and NVLink systems for model training and inference.
SourceThe third wafer-scale engine integrates a large compute array and on-chip memory on one wafer.
SourceGemini 1.5 Pro demonstrates million-token context processing.
SourceGroup-relative rewards estimate advantages without a separate value model.
SourceA sparse mixture-of-experts model expands efficient open-weight inference.
SourceCDNA 3 and 192 GB of HBM3 expand large-model capacity per accelerator.
SourceA fifth-generation TPU focused on large-model training is announced with AI Hypercomputer.
SourceGoogle introduces the multimodal Gemini family: Ultra, Pro, and Nano.
SourceHopper gains HBM3e to increase memory capacity and bandwidth for large-model inference.
SourceA developer preview expands context to 128K and advances vision and tool use.
SourceAn inference library for large language models on NVIDIA GPUs becomes publicly available.
SourceAn efficient 7B model is released under Apache 2.0.
SourceThe cost-focused fifth-generation TPU supports both training and inference.
SourcePowers the Spark all-in-one system introduced by iFLYTEK and Huawei for enterprise LLM deployment.
SourcePaged KV-cache management improves memory utilization in language-model serving.
SourceGroups of query heads share key-value heads to balance quality and KV-cache cost.
SourceAn early autonomous agent uses repeated model calls and tools, adding Selenium-based web browsing in this release.
SourceSecond-generation inference chips power EC2 Inf2 for generative-AI deployment.
Sourcetorch.compile adds compilation-based acceleration to the established PyTorch workflow.
SourceSynthetic instruction data adapts LLaMA into a small instruction-following model.
SourceA GPT-3.5-based research preview supports multi-turn conversation.
SourceA small model drafts tokens and a target model verifies them, accelerating generation while preserving the target distribution.
SourceFirst-generation Trainium becomes generally available through EC2 Trn1 for model training.
SourceVector fields along conditional probability paths train continuous-flow generative models.
SourceReasoning steps and tool actions are interleaved so language models can adapt their plans to environment feedback.
SourceAda and 24 GB of memory support local inference, image generation, and model development.
SourceReduced data movement between GPU memory and on-chip storage accelerates exact attention.
SourceThe second Gaudi generation moves to 7 nm while retaining Ethernet-based training scale-out.
SourceA fixed training-compute budget motivates a different balance of model size and data.
SourceHopper introduces the Transformer Engine and FP8 for large-model training and inference.
SourceMultiple sampled reasoning paths are aggregated to improve answer consistency.
SourceA wafer-on-wafer power-delivery design advances Graphcore’s distributed IPU architecture.
SourceIntermediate reasoning steps in examples improve multi-step language-model tasks.
SourceReinforcement learning from human feedback improves instruction following and alignment with user intent.
SourceDiffusion in a compressed latent space reduces the cost of high-resolution image generation.
SourceCDNA 2 uses a multi-die design to expand memory capacity for training and scientific computing.
SourceA Python-like language and compiler simplify the development of optimized GPU kernels.
SourceFourth-generation TPUs expand pod-scale model training.
SourceThe second wafer-scale engine uses a 7 nm process to expand compute and on-chip memory.
SourceRotations encode position within attention computation.
SourceThe first CDNA architecture targets compute with ROCm support for AI and HPC.
SourceImages become sequences of patches for Transformer-based visual learning.
Source24 GB of memory and Ampere Tensor Cores expand capacity for local model experiments.
SourceSecond-generation IPUs use distributed on-chip memory and parallel execution for machine learning.
SourcePartitioning optimizer state and other training data reduces distributed-training memory use.
SourceA systematic study relates language-model performance to model size, data, and compute.
SourceFirst-generation Inferentia becomes available for machine-learning inference through EC2 Inf1.
SourceA unified API makes Transformer architectures and pretrained models accessible for research and applications.
SourceIntra-layer model parallelism distributes large Transformer training across GPUs.
SourceAscend 910 launches as a dedicated processor for neural-network training.
SourceThe first wafer-scale engine integrates compute, memory, and interconnect on one large chip.
SourceA Rust-based virtual machine monitor provides a compact VM layer for modern cloud workloads.
SourceAn Ethernet-connected architecture targets scalable neural-network training.
SourceA cross-platform runtime connects model formats to execution backends.
SourceAWS releases lightweight microVM technology to isolate short-lived, multi-tenant workloads with separate kernels.
Source7 nm Vega accelerators introduce PCIe 4.0 for deep learning and HPC.
SourceHuawei announces the Ascend chip family and Da Vinci architecture for inference and training.
SourceTuring Tensor Cores bring low-precision compute to data-center inference.
SourceSubword models train directly on raw text without language-specific word segmentation.
SourceA container runtime uses lightweight VMs to give containers or pods their own guest kernel.
SourceThird-generation TPUs expand training scale with liquid-cooled pods.
SourceHuman comparisons train a reward signal that guides reinforcement learning.
SourceGoogle researchers introduce the Transformer, an attention-based sequence model, in Attention Is All You Need.
SourceSecond-generation TPUs add training support and a route to external access through Cloud TPU.
SourceVolta introduces Tensor Cores to accelerate matrix operations for deep learning.
SourceSparse routing selects a small subset of experts to increase capacity while limiting computation.
SourceNormalization statistics are computed within each example rather than across a batch.
SourceGoogle describes its custom tensor processor for neural-network inference.
SourcePascal combines HBM2 and NVLink for deep learning and multi-GPU computing.
SourceGoogle releases its machine-learning framework as open-source software.
SourceByte-pair encoding is applied to subword segmentation to address unknown words.
SourceA teacher model’s output distribution trains a smaller student model.
SourceAn encoder and decoder map variable-length input sequences to variable-length outputs.
SourceA translation model learns to attend to different input positions as it generates each word.
SourceVariational inference and reparameterization learn a latent representation that can be sampled.
Source