Content
<p align="center">
<img alt="capabilities" src="https://img.shields.io/badge/capabilities-18-7c3aed?style=flat-square&labelColor=0f172a"/>
<img alt="tools" src="https://img.shields.io/badge/tools-85%2B-0ea5e9?style=flat-square&labelColor=0f172a"/>
<img alt="license" src="https://img.shields.io/badge/license-CC0--1.0-22c55e?style=flat-square&labelColor=0f172a"/>
<img alt="last%20verified" src="https://img.shields.io/badge/last%20verified-2026--05--21-f59e0b?style=flat-square&labelColor=0f172a"/>
</p>
<h2 align="center">ai-capability-atlas · what AI can actually do, honestly</h2>
<p align="center">
<em>One honest pick per capability — open 🐙 or closed ☁️ — plus a Goal → Stack Cookbook<br/>graded by what really ships, and idea lenses on what's buildable today.</em>
</p>
---
Stop bookmarking "awesome AI" link dumps. **One honest pick per capability** (open 🐙 or closed ☁️), a straight **verdict on which actually wins**, and a **[Goal → Stack Cookbook](#-goal--stack-cookbook)** of goal → stack recipes graded by what really ships — plus idea lenses on what's buildable, sci-fi-but-possible, and ripe to rebuild 10× better.
**Last verified:** 2026-05-21
> **🍳 New: a [Goal → Stack Cookbook](#-goal--stack-cookbook).** The map tells you the
> tool per *capability*; the cookbook tells you the **stack** per *goal*. [Jump to it ↓](#-goal--stack-cookbook)
## ⚖️ Open vs Closed — which should you reach for?
*The honest verdict per capability. **✅** open is there · **➗** depends (cost / privacy / quality tradeoff) · **❌** closed still wins. Full per-capability calls live in each domain's ⚖️ line below.*
| Capability | Leading closed (☁️) | Leading open (🐙) | Which to reach for? |
|---|---|---|---|
| [Chat & reasoning](#-frontier-llms-chat--reasoning) | ChatGPT · Claude | DeepSeek · Llama | ➗ open is close; closed still leads the hardest reasoning |
| [Image generation](#-image-generation) | Midjourney · Gemini | SD / Flux (ComfyUI) | ➗ closed for ease, open for control & cost |
| [Video](#-video-generation) | Veo · Kling | Wan 2.2 · HunyuanVideo | ❌ closed still wins |
| [Music (full songs)](#-music--audio) | Suno · Udio | YuE | ❌ closed still wins |
| [Speech-to-text](#-speech) | Deepgram | Whisper | ✅ open is there |
| [Voice cloning / TTS](#-speech) | ElevenLabs | GPT-SoVITS | ➗ open is excellent; closed is easier |
| [Coding](#-code--coding-assistants) | Cursor · Claude Code | OpenHands · aider | ➗ closed UX, open autonomy |
| [Deep research](#-deep-research--ai-search) | Perplexity | local-deep-research | ➗ depends on sources + control |
| [Document AI / OCR](#-document-ai) | Google Document AI | PaddleOCR | ✅ open is there |
> **How to read this.** Every domain lists leading tools by one bar: *most capable at
> its actual job, usable today, with real adoption — not demos or vaporware.* `🐙` =
> open-source repo (Signal shows GitHub ⭐); `☁️` = closed/hosted product (Signal shows
> access: free / freemium / paid / API). The two signals aren't comparable on purpose —
> stars prove open adoption; the access tag tells you how to get a closed tool. No links are
> sponsored or affiliate; inclusion is not endorsement-for-pay. Star counts and access tags
> drift — re-verify against the source before citing, and see the *Last verified* date above.
> **Running the open ones.** Most *generative* repos (video, music, image, voice) want an
> NVIDIA GPU with enough VRAM — they won't run well on CPU or integrated graphics. *Analysis,
> agent, and RAG* tools usually run on CPU or lean on a hosted model API. No GPU? Reach for a
> hosted GPU (Replicate / fal / HF Inference) instead of running locally.
---
## 🍳 Goal → Stack Cookbook
*The map above answers "which tool per **capability**." The [**Cookbook**](RECIPES.md)
answers "what **stack** for my **goal**" — with an honest effectiveness verdict and a
chooser that routes by what you care about (free / local / quality / easy). The top picks:*
| Goal | ⭐ Default pick | Effectiveness |
|---|---|---|
| Make a music video from a prompt | Suno → Kling → CapCut | 🟢 Battle-tested |
| Make an expensive-looking UI/UX website | v0 / Lovable | 🟡 Emerging |
| Voice-cloned audiobook from any PDF | docling → GPT-SoVITS | 🟢 Battle-tested |
| Turn any website into clean data for an LLM | firecrawl | 🟢 Battle-tested |
| Chat with my PDFs and docs | ragflow / NotebookLM | 🟢 Battle-tested |
| Run a local "ChatGPT", fully offline | Ollama + DeepSeek/Qwen | 🟢 Battle-tested |
**[→ Full cookbook: recipes, sibling stacks, and the decision guide](RECIPES.md)**
---
## Contents
**🍳 Build something** · [Goal → Stack Cookbook](RECIPES.md) — pick a goal, get the stack
**Foundations & assistants**
[🧠 Frontier LLMs](#-frontier-llms-chat--reasoning) · [🔬 Deep Research & AI Search](#-deep-research--ai-search) · [🧑💻 Code & Coding Assistants](#-code--coding-assistants) · [🖥️ Computer-Use](#-computer-use--browser-agents) · [🤖 Agent Frameworks](#-agent-frameworks)
**Generative & media**
[🎬 Video](#-video-generation) · [🎵 Music & Audio](#-music--audio) · [🗣️ Speech](#-speech) · [🖼️ Image](#-image-generation) · [🧊 3D & Rendering](#-3d--neural-rendering) · [🧑🎤 Talking Heads](#-talking-heads)
**Understanding & data**
[👁️ Vision](#-vision--segmentation) · [📄 Document AI](#-document-ai) · [📚 RAG](#-rag--knowledge) · [🗂️ Knowledge Work](#-knowledge-work) · [🌍 Geospatial & Maps](#-geospatial--maps)
**Embodied & physical**
[🦾 Robotics](#-robotics--embodied-ai) · [🚗 Autonomous Driving](#-autonomous-driving)
**Ideas & lenses**
[🧪 Frontier](#-frontier--emerging) · [🔭 Watch List](#-watch-list) · [🧩 Theoretically Possible](#-theoretically-possible) · [🛸 Sci-Fi but Possible](#-sci-fi-but-technically-possible) · [🔁 Ripe for a Redo](#-ripe-for-a-10x-redo) · [🛡️ Runs-Local](#-runs-local--privacy-first) · [🛑 Still Can't](#-still-cant-hype-check) · [🚧 Unlocks When X](#-unlocks-when-x-lands) · [💸 Cost Collapse](#-cost-collapse) · [🧱 Primitives](#-core-lego-brick-primitives)
---
## 🧠 Frontier LLMs (Chat & Reasoning)
*Talk to a top model — reason, write, analyze, code.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [ChatGPT](https://chatgpt.com) | ☁️ | freemium | The default frontier assistant; strong reasoning + tools |
| [Claude](https://claude.ai) | ☁️ | freemium | Frontier reasoning, long context, strong at code + analysis |
| [Gemini](https://gemini.google.com) | ☁️ | freemium | Frontier multimodal; deep Google integration |
| [deepseek-ai/DeepSeek-V3](https://github.com/deepseek-ai/DeepSeek-V3) | 🐙 | 103.6k⭐ | Open-weights frontier-class reasoning you can self-host |
| [QwenLM/Qwen3](https://github.com/QwenLM/Qwen3) | 🐙 | 27.2k⭐ | Strong open-weights family, many sizes, easy to run |
| [meta-llama/llama-models](https://github.com/meta-llama/llama-models) | 🐙 | 7.6k⭐ | The reference open-weights family |
**⚖️ Open vs closed:** ➗ open weights (DeepSeek/Qwen/Llama) are close on many tasks; ❌ closed still leads the hardest reasoning and agentic benchmarks.
## 🔬 Deep Research & AI Search
*Autonomous, multi-source, cited research — and answer engines.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Perplexity](https://www.perplexity.ai) | ☁️ | freemium | Cited answer engine; fast multi-source synthesis |
| [ChatGPT / Gemini deep research](https://chatgpt.com) | ☁️ | freemium | Long-running agentic research with citations |
| [Alibaba-NLP/DeepResearch](https://github.com/Alibaba-NLP/DeepResearch) | 🐙 | 18.9k⭐ | Tongyi's leading open deep-research agent; plans, searches, synthesizes |
| [dzhng/deep-research](https://github.com/dzhng/deep-research) | 🐙 | 19.0k⭐ | The simplest deep-research agent — refines its own direction; great to learn from |
| [LearningCircuit/local-deep-research](https://github.com/LearningCircuit/local-deep-research) | 🐙 | 7.9k⭐ | ~95% SimpleQA fully local; arXiv/PubMed/your docs, encrypted |
**⚖️ Open vs closed:** ➗ closed answer engines (Perplexity) win for instant cited answers; open agents win for deep, local, controllable research over your own sources.
## 🎬 Video Generation
*Text→video, image→video, animated shorts.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Google Veo](https://gemini.google.com) | ☁️ | freemium | Frontier text→video with native audio |
| [Runway](https://runwayml.com) | ☁️ | freemium | Pro creative video suite (gen + edit) |
| [Kling](https://klingai.com) | ☁️ | freemium | Strong, fast image→video |
| [Seedance (Dreamina)](https://dreamina.capcut.com) | ☁️ | freemium | ByteDance multimodal video — native audio, physics-accurate motion |
| [Wan-Video/Wan2.2](https://github.com/Wan-Video/Wan2.2) | 🐙 | 15.8k⭐ | MoE open video — the open cinematic quality leader (text & image→video) |
| [Tencent-Hunyuan/HunyuanVideo](https://github.com/Tencent-Hunyuan/HunyuanVideo) | 🐙 | 12.1k⭐ | 13B open model rivaling closed systems on cinematic realism |
| [zai-org/CogVideo](https://github.com/zai-org/CogVideo) | 🐙 | 12.7k⭐ | CogVideoX — the most-used open Sora-style text & image → video model |
| [Lightricks/LTX-Video](https://github.com/Lightricks/LTX-Video) | 🐙 | 10.3k⭐ | Fast DiT video generation that runs on consumer GPUs |
**⚖️ Open vs closed:** ❌ closed still wins — Veo/Kling lead on coherence, length, and native audio; open (Wan 2.2 / HunyuanVideo) is closing fast on quality but wants more VRAM.
## 🎵 Music & Audio
*Generate music, songs, and sound.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Suno](https://suno.com) | ☁️ | freemium | Standout full-song generation (vocals + backing) from a prompt |
| [Udio](https://www.udio.com) | ☁️ | freemium | High-fidelity music generation with fine vocal control |
| [multimodal-art-projection/YuE](https://github.com/multimodal-art-projection/YuE) | 🐙 | 6.2k⭐ | Open full-*song* generation — self-hostable, an open Suno |
| [facebookresearch/audiocraft](https://github.com/facebookresearch/audiocraft) | 🐙 | 23.3k⭐ | MusicGen + EnCodec — controllable music from text + melody |
| [deezer/spleeter](https://github.com/deezer/spleeter) | 🐙 | 28.2k⭐ | Split any song into stems (vocals / drums / bass / other) |
**⚖️ Open vs closed:** ❌ closed still wins on full songs (Suno/Udio); 🐙 open is fine for stems and controllable backing tracks.
## 🗣️ Speech
*TTS, voice cloning, realtime, ASR.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [ElevenLabs](https://elevenlabs.io) | ☁️ | freemium | Leading voice cloning + expressive TTS |
| [Deepgram](https://deepgram.com) | ☁️ | API | Fast, accurate hosted speech-to-text |
| [openai/whisper](https://github.com/openai/whisper) | 🐙 | 99.9k⭐ | The reference speech-to-text; robust across languages/accents |
| [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) | 🐙 | 57.6k⭐ | Clone a usable voice from ~1 minute of audio |
| [index-tts/index-tts](https://github.com/index-tts/index-tts) | 🐙 | 20.6k⭐ | Industrial zero-shot, controllable, cross-lingual TTS |
**⚖️ Open vs closed:** ✅ open is there for transcription (Whisper); ➗ for voice cloning open (GPT-SoVITS) is excellent, ElevenLabs is just easier and more polished.
## 🖼️ Image Generation
*Create and edit images with diffusion.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Midjourney](https://www.midjourney.com) | ☁️ | paid | Polished general image generation, minimal effort |
| [Gemini (Nano Banana)](https://gemini.google.com) | ☁️ | freemium | Frontier image generation + conversational editing |
| [GPT-Image / DALL·E](https://platform.openai.com) | ☁️ | API | Strong prompt-following image gen via API |
| [black-forest-labs/flux](https://github.com/black-forest-labs/flux) | 🐙 | 25.6k⭐ | FLUX — frontier open image model (FLUX.2); top prompt adherence + text rendering |
| [AUTOMATIC1111/stable-diffusion-webui](https://github.com/AUTOMATIC1111/stable-diffusion-webui) | 🐙 | 163.2k⭐ | The canonical local Stable Diffusion playground |
| [comfyanonymous/ComfyUI](https://github.com/comfyanonymous/ComfyUI) | 🐙 | 113.8k⭐ | Node-based diffusion workflow GUI — the power-user standard |
| [huggingface/diffusers](https://github.com/huggingface/diffusers) | 🐙 | 33.7k⭐ | SOTA diffusion for image/video/audio — the library everything builds on |
**⚖️ Open vs closed:** ➗ closed (Midjourney/Gemini) for effortless quality; 🐙 open (SD/Flux via ComfyUI) for control, fine-tuning, and zero per-image cost.
## 🧊 3D & Neural Rendering
*Reconstruct and render 3D from images — NeRF & Gaussian splatting.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [graphdeco-inria/gaussian-splatting](https://github.com/graphdeco-inria/gaussian-splatting) | 🐙 | 22.0k⭐ | The reference 3D Gaussian Splatting implementation — real-time radiance fields |
| [nerfstudio-project/nerfstudio](https://github.com/nerfstudio-project/nerfstudio) | 🐙 | 11.6k⭐ | Collaboration-friendly studio for NeRFs and splats |
| [playcanvas/supersplat](https://github.com/playcanvas/supersplat) | 🐙 | 8.5k⭐ | Browser-based 3D Gaussian-splat editor |
**⚖️ Open vs closed:** ✅ open dominates; no real closed contender — the research and tooling are open-source-led.
## 🧑🎤 Talking Heads
*Digital humans, avatars, and lip-sync.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [HeyGen](https://www.heygen.com) | ☁️ | freemium | Polished talking-avatar video from text, many languages |
| [OpenTalker/SadTalker](https://github.com/OpenTalker/SadTalker) | 🐙 | 13.8k⭐ | One photo + audio → talking-face video |
| [lipku/LiveTalking](https://github.com/lipku/LiveTalking) | 🐙 | 7.7k⭐ | Real-time streaming digital human (lip-sync + NeRF) |
| [Open-LLM-VTuber/Open-LLM-VTuber](https://github.com/Open-LLM-VTuber/Open-LLM-VTuber) | 🐙 | 7.8k⭐ | Hands-free voice chat with a local Live2D avatar |
**⚖️ Open vs closed:** ➗ closed (HeyGen/Synthesia) leads polished presenter avatars; 🐙 open wins for real-time, custom, and offline use.
## 🖥️ Computer-Use & Browser Agents
*AI that operates real software and the web.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [ChatGPT Agent](https://chatgpt.com) | ☁️ | freemium | OpenAI's browsing/computer-use agent (formerly Operator) |
| [Manus](https://manus.im) | ☁️ | freemium | General autonomous task agent that drives tools + web |
| [browser-use/browser-use](https://github.com/browser-use/browser-use) | 🐙 | 94.9k⭐ | Let an LLM drive a real browser to do tasks |
| [firecrawl/firecrawl](https://github.com/firecrawl/firecrawl) | 🐙 | 122.4k⭐ | Turn any website into clean LLM-ready markdown for agents |
| [trycua/cua](https://github.com/trycua/cua) | 🐙 | 17.0k⭐ | Sandboxed full-desktop control (macOS/Linux/Windows) for agents |
**⚖️ Open vs closed:** ➗ closed agents are more reliable out of the box; 🐙 open gives you sandboxing, scripting, and no per-task fees.
## 🧑💻 Code & Coding Assistants
*AI that writes, fixes, and ships software — editors and autonomous agents.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Cursor](https://cursor.com) | ☁️ | freemium | The popular AI-native code editor |
| [GitHub Copilot](https://github.com/features/copilot) | ☁️ | freemium | In-editor completion + chat across IDEs |
| [Claude Code](https://www.anthropic.com/claude-code) | ☁️ | paid | Terminal-native agentic coding |
| [OpenHands/OpenHands](https://github.com/OpenHands/OpenHands) | 🐙 | 74.3k⭐ | Full AI-driven development agent — edits, runs, and tests code |
| [cline/cline](https://github.com/cline/cline) | 🐙 | 62.1k⭐ | Autonomous coding agent in your editor, plan + act |
| [Aider-AI/aider](https://github.com/Aider-AI/aider) | 🐙 | 45.1k⭐ | AI pair programming in your terminal, git-native |
| [SWE-agent/SWE-agent](https://github.com/SWE-agent/SWE-agent) | 🐙 | 19.3k⭐ | Takes a GitHub issue and auto-fixes it with the LM of your choice |
**⚖️ Open vs closed:** ➗ closed assistants (Cursor/Copilot/Claude Code) lead daily editor UX; 🐙 open agents (OpenHands/aider) win for autonomous, self-hosted, scriptable runs.
## 📄 Document AI
*PDF/image → structured, LLM-ready data.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Google Document AI](https://cloud.google.com/document-ai) | ☁️ | API | Managed parsing/extraction with form + table models |
| [AWS Textract](https://aws.amazon.com/textract) | ☁️ | API | Hosted OCR + form/table extraction at scale |
| [PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR) | 🐙 | 78.3k⭐ | 100+ language OCR; PDF/image → structured data for LLMs |
| [docling-project/docling](https://github.com/docling-project/docling) | 🐙 | 60.1k⭐ | Any document → gen-AI-ready markdown, tables intact |
| [Unstructured-IO/unstructured](https://github.com/Unstructured-IO/unstructured) | 🐙 | 14.7k⭐ | ETL that turns messy docs into clean RAG chunks |
**⚖️ Open vs closed:** ✅ open is there — PaddleOCR/docling rival the cloud APIs; reach for closed only for managed scale + SLAs.
## 📚 RAG & Knowledge
*Chat with your data.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Vertex AI Search](https://cloud.google.com/enterprise-search) | ☁️ | API | Managed retrieval + grounding, zero-ops RAG |
| [infiniflow/ragflow](https://github.com/infiniflow/ragflow) | 🐙 | 80.9k⭐ | Leading open RAG engine fused with agent capabilities |
| [vanna-ai/vanna](https://github.com/vanna-ai/vanna) | 🐙 | 23.5k⭐ | Chat with your SQL database — agentic text-to-SQL |
| [khoj-ai/khoj](https://github.com/khoj-ai/khoj) | 🐙 | 34.6k⭐ | Self-hostable "second brain" over your docs + web |
**⚖️ Open vs closed:** ✅ open is there — ragflow/khoj are strong; reach for closed (Vertex) only when you want managed, zero-ops retrieval.
## 🗂️ Knowledge Work
*Turn your docs, meetings, and notes into answers and artifacts.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [NotebookLM](https://notebooklm.google.com) | ☁️ | free | Notes-Q&A grounded in your sources; audio overviews |
| [Granola](https://www.granola.ai) | ☁️ | freemium | Meeting capture → structured notes, no bot in the call |
| [Gamma](https://gamma.app) | ☁️ | freemium | Deck/doc generation from a prompt |
| [Notion AI](https://www.notion.so/product/ai) | ☁️ | paid | In-doc AI: write, summarize, ask across your workspace |
**⚖️ Open vs closed:** ❌ closed leads knowledge-work UX; the open peers are general RAG tools (see [RAG](#-rag--knowledge)), not polished products.
## 🌍 Geospatial & Maps
*Satellite/aerial imagery and "Google Maps"-adjacent AI.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Google Earth Engine](https://earthengine.google.com) | ☁️ | freemium | Planetary-scale satellite analysis in the cloud |
| [opengeos/geoai](https://github.com/opengeos/geoai) | 🐙 | 3.0k⭐ | Segment buildings/roads/features straight off maps & imagery |
| [microsoft/torchgeo](https://github.com/microsoft/torchgeo) | 🐙 | 4.0k⭐ | PyTorch datasets/models for satellite & aerial imagery |
| [obss/sahi](https://github.com/obss/sahi) | 🐙 | 5.3k⭐ | Sliced inference for tiny objects in huge satellite/drone images |
**⚖️ Open vs closed:** ➗ open tooling is strong for analysis; reach for closed (Earth Engine) for planetary-scale compute you don't host.
## 👁️ Vision & Segmentation
*Segment, detect, and understand anything in images & video.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [Google Cloud Vision](https://cloud.google.com/vision) | ☁️ | API | Managed detection, OCR, labels, faces |
| [facebookresearch/sam3](https://github.com/facebookresearch/sam3) | 🐙 | 9.9k⭐ | Latest Segment Anything — segment any object from a prompt |
| [IDEA-Research/Grounded-Segment-Anything](https://github.com/IDEA-Research/Grounded-Segment-Anything) | 🐙 | 17.6k⭐ | Text → detect → segment → edit/generate anything |
| [ultralytics/ultralytics](https://github.com/ultralytics/ultralytics) | 🐙 | 57.4k⭐ | YOLO — real-time detection, segmentation, pose, tracking; the default |
**⚖️ Open vs closed:** ✅ open dominates — SAM/YOLO are the standard; closed cloud vision is for managed convenience, not capability.
## 🤖 Agent Frameworks
*Build multi-agent systems and run models locally.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [langchain-ai/langchain](https://github.com/langchain-ai/langchain) | 🐙 | 137.3k⭐ | The agent-engineering platform; LangGraph for stateful agents |
| [langgenius/dify](https://github.com/langgenius/dify) | 🐙 | 142.1k⭐ | Visual builder for agentic workflows + RAG, low/no-code |
| [ollama/ollama](https://github.com/ollama/ollama) | 🐙 | 171.8k⭐ | One command to run open models locally |
**⚖️ Open vs closed:** ✅ open dominates — closed agent platforms (OpenAI Assistants, Vertex Agent Builder) exist but trade portability for lock-in.
## 🦾 Robotics & Embodied AI
*Teach machines to act in the physical world.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [huggingface/lerobot](https://github.com/huggingface/lerobot) | 🐙 | 24.2k⭐ | End-to-end learning for robotics (datasets, policies, real arms) |
| [haosulab/ManiSkill](https://github.com/haosulab/ManiSkill) | 🐙 | 2.9k⭐ | GPU-parallelized robotics simulator + manipulation benchmark |
| [dora-rs/dora](https://github.com/dora-rs/dora) | 🐙 | 3.8k⭐ | Low-latency dataflow middleware for building robotic apps |
**⚖️ Open vs closed:** ✅ open dominates the tooling; general-purpose robot foundation models remain mostly closed research (see [Watch List](#-watch-list)).
## 🚗 Autonomous Driving
*Open self-driving stacks and simulators.*
| Tool | Type | Signal | What's cool |
|---|---|---|---|
| [commaai/openpilot](https://github.com/commaai/openpilot) | 🐙 | 61.0k⭐ | Driver-assistance OS that upgrades 300+ production cars |
| [carla-simulator/carla](https://github.com/carla-simulator/carla) | 🐙 | 14.0k⭐ | The open simulator for autonomous-driving research |
| [autowarefoundation/autoware](https://github.com/autowarefoundation/autoware) | 🐙 | 11.6k⭐ | The leading open-source full AV software stack |
**⚖️ Open vs closed:** ✅ open dominates what's public; full production AV stacks (Waymo, Tesla FSD) are closed and not a tool you can run.
## 🧪 Frontier & Emerging
*Admitted on novelty of capability, not star count — the genuinely new stuff. All 🐙 open.*
| Repo | ⭐ | Why it's novel |
|---|---|---|
| [SakanaAI/AI-Scientist](https://github.com/SakanaAI/AI-Scientist) | 13.7k | Fully automated, open-ended scientific discovery: ideates, runs experiments, writes the paper |
| [SakanaAI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2) | 6.3k | v2 — workshop-level automated discovery via agentic tree search |
| [Genesis-Embodied-AI/Genesis](https://github.com/Genesis-Embodied-AI/Genesis) | 28.8k | A generative physics "world" to train robots in — describe a task, it builds the simulation |
| [danijar/dreamerv3](https://github.com/danijar/dreamerv3) | 3.3k | World models — learns its environment and plans by "imagining" outcomes |
| [RosettaCommons/RFdiffusion](https://github.com/RosettaCommons/RFdiffusion) | 2.9k | Generative protein *design* — diffusion that invents new protein structures from scratch |
| [jwohlwend/boltz](https://github.com/jwohlwend/boltz) | 4.0k | Open AlphaFold3-class biomolecular structure + binding prediction |
| [TransformerLensOrg/TransformerLens](https://github.com/TransformerLensOrg/TransformerLens) | 3.4k | Mechanistic interpretability — crack open a model and see what each part computes |
## 🔭 Watch List
*Real capabilities that don't yet have a strong, runnable open repo. Parked so they aren't lost — promote when one matures.*
| Capability | Why it's not here yet | Closest signal |
|---|---|---|
| Playable neural game engines | "Generate a game you can play" (Genie / Oasis-class) is mostly closed | Dreamer / Genesis cover world-models for *control* (see Frontier), not playable generation |
| AI-native maps ("AI + Google Maps") | Shows up as a tool *inside* agent frameworks, not a standalone model | geocoding / routing live in agent tool-calls |
| General-purpose robot foundation models | One policy for any robot/task (π0 / RT-X-style) is mostly closed research | LeRobot hosts some open policies (see Robotics) |
| Full-duplex real-time voice | Model-level interrupt-and-respond leads in closed (GPT/Gemini realtime) | STT→LLM→TTS stacks (Pipecat / LiveKit) approximate it |
| Deepfake detection | No strong *maintained* open anchor; aging challenge solutions + paper lists | [selimsef/dfdc_deepfake_challenge](https://github.com/selimsef/dfdc_deepfake_challenge) (~0.9k) |
| AI chip / hardware design | High value, locked inside labs and EDA vendors | — |
| Large-scale autonomous materials discovery | GNoME-scale closed-loop discovery remains lab-bound | open ML interatomic potentials (MACE / CHGNet) exist |
## 🧩 Theoretically Possible
*Effectuation, not goal-chasing: treat this page as a stock of **means** and ask "what becomes possible?" Each row is a buildable combination — cheap to prototype, every piece available today, and not yet an obvious off-the-shelf product. Speculative, not endorsements. Off-page inputs (your photos, a drone feed) are noted in plain text.*
| What you could build | Stitch together | Why it's newly possible |
|---|---|---|
| **Spoken weather briefing** | GraphCast + Whisper + a TTS (IndexTTS) | Ask "what's my week look like?" out loud, get a spoken local forecast from a model that beats traditional numerical weather prediction. |
| **One-prompt music video** | YuE + CogVideo / LTX-Video + SadTalker | Text prompt → original song → matching video → a lip-synced performer, end to end, runnable locally. |
| **Signed ↔ spoken translator** | a sign-recognition model + SeamlessM4T + a TTS | Recognize signs from video, translate, speak them — bridging a weak open domain (sign) with a strong one (translation). |
| **Avatar research explainer** | a deep-research agent + voice clone (GPT-SoVITS) + LiveTalking | Ask a question; a real-time talking-head explains the *cited* answer in your own voice. |
| **Local autonomous engineer** | OpenHands + llama.cpp + a knowledge-graph memory | A coding agent that runs fully offline and remembers your repo across sessions. |
| **Drone disaster mapper** | SAHI + SAM + TorchGeo + *drone feed* | Detect tiny objects in huge frames, track them across video, and georeference damage live. |
| **Self-driving materials lab** | AI-Scientist + ML interatomic potentials | An autonomous loop that proposes candidate materials and simulates their properties — no wet lab. |
| **Memory documentary maker** | *your photo folder* + voice clone + CogVideo + a music model | Turn a folder of memories into a narrated, scored short film. |
| **Voiced tabletop DM** | Ollama (local LLM) + a music model + a TTS + SadTalker | An RPG narrator that talks, scores the scene, and gives NPCs faces. |
| **Podcast → study pack** | Whisper + a deep-research agent + docling | Drop a talk or lecture; get cited notes, a summary, and a quiz. |
| **Voice-cloned audiobook** | docling + GPT-SoVITS | Any PDF, read aloud in any voice you give one minute of. |
| **Search your whole life** | khoj + Whisper | One semantic search across everything you've written, said, or saved. |
## 🛸 Sci-Fi, But Technically Possible
*Wilder, multi-system, and hard to integrate — but every component already exists in open source or active research. The gap is engineering and data, not physics. Honest difficulty is in the third column; a few raise real ethics questions, flagged inline.*
| What you could build | The pieces | The hard part / how far off |
|---|---|---|
| **Imagination-to-video** | EEG decoding + a video diffusion model + a world model | fMRI-to-image is published; EEG is far cruder. Real, but low-fidelity for years. |
| **Persistent agent town** | swarms of browser-use / OpenHands agents + shared memory + a world model | Coordinating many agents without drift is unsolved; shared memory + a sim make it conceivable. |
| **Robot apprentice from YouTube** | LeRobot + SAM + a vision-language-action model + web video | Learning manipulation from third-person video is active research; sim-to-real is the wall. |
| **A simulator of your life** | lifelogging (camera/logs, off-page) + a world model | "What happens if I…" answered by a model of *your* environment. Data capture is the bottleneck. |
| **Closed-loop autonomous lab** | AI-Scientist-v2 + ML potentials + lab robotics (off-page) | The software loop exists; the wet-lab robotics integration is the missing, expensive half. |
| **Interactive memory twin** ⚠️ | voice clone + talking-head + RAG over someone's writings | Buildable now — and ethically loaded. Consent and likeness rights are the real blockers, not the tech. |
| **Game from a sentence** | a neural world model (Genie / Oasis-class, mostly closed) + diffusion | Playable worlds generated live; the strong models aren't open yet (see Watch List). |
## 🔁 Ripe for a 10x Redo
*Old categories where one newly-cheap or newly-open capability resets the bar. Not new ideas — newly **better** ones.*
| Category | The old way | What's now possible |
|---|---|---|
| **Audiobook narration** | Hire narrators, or settle for robotic TTS | Clone a voice from one minute (GPT-SoVITS / IndexTTS) → any book, any voice, free. |
| **Subtitles & captions** | Manual transcription or clunky ASR | Whisper + speaker diarization + SeamlessM4T → accurate multilingual subs cheaply. |
| **Weather apps** | Wrap a national-forecast API | Run GraphCast yourself — SOTA 10-day forecasts you control. |
| **3D scanning / photogrammetry** | Expensive rigs, slow pipelines | Gaussian splatting (nerfstudio) from a phone video. |
| **Stock photos & b-roll** | License from a library | Generate the exact shot with diffusers / CogVideo. |
| **GIS map digitizing** | Humans tracing buildings off imagery | geoai / SAM auto-segment features straight from satellite. |
| **Language learning** | Scripted lessons, fixed dialogues | Local LLM + voice clone + TTS → an infinite, personalized conversation partner. |
| **Enterprise / doc search** | Keyword search over a file share | RAG (ragflow) → semantic search that actually finds it. |
| **Indie music production** | Studio time and session players | YuE generates stems and full songs from a prompt. |
## 🛡️ Runs-Local / Privacy-First
*Capabilities that run entirely on your machine — no data leaves, no API key, works offline. All 🐙 open, pulled from the domains above.*
- **Local LLMs** — Ollama, llama.cpp, DeepSeek / Qwen / Llama weights
- **Local research & RAG** — local-deep-research, khoj, ragflow
- **Local speech** — Whisper / faster-whisper (CPU-OK), GPT-SoVITS
- **Local vision** — SAM, YOLO (Ultralytics)
- **Local image** — Stable Diffusion / Flux via ComfyUI
*What you trade to keep it local: you give up the closed leader's quality and UX ceiling (Suno for songs, Veo/Kling for video, GPT-4o/Gemini for the hardest reasoning) in exchange for privacy, zero per-use cost, and offline use. The gap is **smallest** for transcription, RAG, OCR, and detection — and **largest** for video and full-song music.*
## 🛑 Still Can't (Hype Check)
*Honest limits — things people assume AI nails but it doesn't, yet. A "what AI can do" list owes you the inverse. Split into "closed can, open can't yet" vs "nobody can yet."*
| The claim | The reality |
|---|---|
| "Open models match the leading closed everywhere" | **Closed can, open can't yet** — open is close on many tasks, still behind on the hardest reasoning and agentic benchmarks. |
| "Generate a feature-length film from a prompt" | **Nobody can yet** — a few coherent minutes is the frontier (open or closed); character and scene continuity break down fast. |
| "Agents can run my business autonomously" | **Nobody can yet** — agents drift over long horizons; reliable multi-step autonomy is unsolved. Keep a human in the loop. |
| "A robot can learn any task from a video" | **Nobody can yet** — manipulation from third-person video is active research; sim-to-real is the wall. |
| "Decode imagined speech from a headband" | **Nobody can yet** — non-invasive EEG is crude; high-fidelity decoding still needs invasive BCIs. |
| "Hallucination-free RAG" | **Nobody can yet** — retrieval reduces hallucination; it doesn't eliminate it. Cite and verify. |
## 🚧 Unlocks When X Lands
*Builds blocked on ONE missing piece maturing. Watch the [Watch List](#-watch-list); when a gate opens, these go from sci-fi to weekend.*
| Build | Blocked on | Unlocks when |
|---|---|---|
| Game from a sentence | An open world model (Genie / Oasis-class) | A strong open neural game engine ships |
| Truly local generative video | Cheap consumer GPUs / efficient models | Video models run well on 8-12 GB VRAM |
| Thought-to-text on a headband | Non-invasive BCI fidelity | EEG decoding crosses usable accuracy |
| Drop-in AI maps | An open AI-native maps stack | Routing / geocoding / places get a strong OSS model |
## 💸 Cost Collapse
*Where the bill went to ~zero. ([Ripe for a Redo](#-ripe-for-a-10x-redo) is about better **quality**; this is about lower **price** — what you used to pay specialists, vendors, or closed SaaS for.)*
| Was (the closed bill) | Now (self-host open) |
|---|---|
| Pro voiceover / ElevenLabs minutes | Clone a voice from one minute (GPT-SoVITS) |
| Suno / Udio subscription | YuE generates stems + full songs locally |
| Per-minute transcription SaaS | Whisper, local and free |
| Document OCR vendors (Textract/DocAI) | PaddleOCR + an LLM → structured data |
| LLM API spend at scale | Self-host DeepSeek / Qwen / Llama with Ollama / llama.cpp |
| Translation / localization agencies | SeamlessM4T, speech↔text across ~100 languages |
| 3D-scanning rigs | Gaussian splatting from a phone video |
## 🧱 Core Lego-Brick Primitives
*The repos that show up in the most combinations above — learn these first; everything else snaps onto them. All 🐙 open.*
| Primitive | Go-to repo | Snaps into |
|---|---|---|
| Speech → text | Whisper / faster-whisper | research, meetings, study packs, voice agents |
| Text → speech | IndexTTS / GPT-SoVITS | audiobooks, avatars, briefings, agents |
| Segment anything | SAM | vision, geospatial, robotics |
| Image / video gen | diffusers | image, video, b-roll |
| Run a model locally | Ollama / llama.cpp | every offline LLM build |
| Drive software | browser-use | agents, scraping, automation |
---
## Notes
- **Open vs closed is a coarse split on purpose.** "Open" spans open-weights models, open-source
apps, and source-available code; "closed" spans hosted SaaS, API-only models, and consumer apps.
The 🐙/☁️ badge is a quick orientation, not a license classification — check each project's terms.
- **Verdicts are curation, held to a bar** (most capable at its job, usable today, real adoption),
not benchmarks. Disagree? Open a PR with a one-line rationale — see [CONTRIBUTING](CONTRIBUTING.md).
- **No links are sponsored or affiliate.** Inclusion is editorial, never paid.
- This is a **curation-taste map, not a comprehensive index** — one tight, hand-picked set per
capability plus the idea lenses, not breadth competition with the big "awesome AI" lists.
- Star counts and closed-tool access tags drift. Re-verify before citing; update the *Last verified*
date when you do.
## Contributing
Additions welcome — see [CONTRIBUTING.md](CONTRIBUTING.md) for the inclusion bar
(the leading pick per capability, open or closed; the closed-tool source + no-affiliate rules;
novelty routed to Frontier).
## License
The curation — selection, descriptions, verdicts, and structure — is dedicated to the public
domain under [CC0-1.0](LICENSE). Each linked repository or product keeps its own license and terms.
---
**Sibling work** — part of Mike Ilog's GitHub portfolio. The [AI-evaluation engineering portfolio](https://github.com/Mike-E-Log#now-public-eval-coded) is the eval-craft side; this repo is the AI-capability-map side.
Profile: [github.com/Mike-E-Log](https://github.com/Mike-E-Log) · website: [mikeilog.com](https://mikeilog.com)
Connection Info
You Might Also Like
Train-in-Silence
The first Task-Aware MCP server and automated VRAM calculator for LLM...
stacklit
108,000 lines of code. 4,000 tokens of index. One command makes any repo...
AppClaw
AI-powered mobile automation agent — describe what you want in plain...
pdf-mcp
Production-ready MCP server for PDF processing with intelligent caching....
kotadb
Local-only code intelligence API for AI developer workflows (Bun +...
gemini-api-docs-mcp
A remote HTTP MCP server for searching Google Gemini API documentation.