AI capability timeline (2012-2026) - by branch (LLMs are one of them)
(AI-chronology Tab. Organized by branch, because "AI" is not one arc - each branch (vision, generative image, language, reinforcement learning, science) has its own capability curve. LLMs are one of them.)
Vision (the first superhuman branch)
- 2012 AlexNet -> 2014 GANs -> 2015 ResNet (surpasses human ImageNet top-5) -> object detection/segmentation -> by ~2018 vision benchmarks near-saturated. Vision was superhuman years before language.
Generative image / video (a separate branch)
- 2014 GANs -> 2020 diffusion (DDPM) -> 2022 DALL-E 2 / Stable Diffusion / Midjourney (text-to-image goes mainstream) -> 2024-26 video (Sora-class). Distinct from LLMs; later fused into multimodal systems.
Language -> multimodal (the branch now dominant)
- 2018 BERT / GPT -> 2019 GPT-2 -> 2020 GPT-3 (few-shot, scaling laws) -> 2022 ChatGPT (mass adoption) -> 2023 GPT-4 / Claude / Gemini (multimodal) -> 2024-26 "reasoning" models (o-series, DeepSeek R1) + long context + tool use. The capability jumps are real; the "reasoning" / "AGI" labels are provider framing (graded).
Reinforcement learning (feeds two arcs)
- 2013 DQN -> 2016 AlphaGo -> 2017 AlphaZero -> 2019 AlphaStar/OpenAI Five. RL then re-entered the language branch twice: RLHF/RLAIF (aligning LLMs, 2022+) and RL-on-verifiable-rewards (the 2025-26 "reasoning" training).
Agents (the 2024-2026 shift)
- Tool-use -> computer-use / browser agents / coding agents -> multi-step autonomy. Capability is uneven and benchmark-gamed; long-horizon reliability remains the open problem (autonomy claims are reported, not settled).
Science (furthest ahead, and not an LLM)
- AlphaFold (2020) -> AlphaFold3 (2024) -> biology / materials / weather-prediction models. The domains where non-LLM AI is most clearly superhuman - the taxonomy's living proof.
The efficiency inflection (2025-2026)
- DeepSeek R1 + the open-weight wave showed frontier-adjacent capability at a fraction of the cost - bending the arc from pure scale toward efficiency (fin-ai-efficiency-counter-thesis).
Honest limits
Releases, dates, and benchmarks are fact; "reasoning"/"agentic"/"AGI-adjacent" are marketing-laden labels (graded); capability claims are provider-reported unless independently benchmarked; benchmark-gaming is real. The point of the by-branch view: the capability frontier is plural, and equating "AI progress" with "LLM releases" mis-measures it.
Sources: model releases + benchmark milestones across branches. Cross-refs: ImageNet_AlexNet, Diffusion_Models, Transformer, Multimodal_AI, Agentic_AI, Reinforcement_Learning, AlphaGo_AlphaFold, DeepSeek, Deep_Learning.
← Research index · structured data: spec-ai-capability-timeline.json · spec-ai-capability-timeline.md