Series · 6 of 6 parts · Advanced
How Vision Models Work
From CLIP to diffusion to video: how machines learned to see, draw, and read the visual world.
What you’ll understand
By the end you understand the architectures behind image understanding and image generation, and can reason about their failure modes.
Episodes
01How CLIP Taught AI to See and Read at OnceContrastive learning on 400 million image-text pairs: the foundation under modern visual AI.12 min→02Inside LLaVA: Bolting Eyes Onto a Language ModelHow open vision-language models connect an image encoder to an LLM, layer by layer.12 min→03Diffusion: How Image Generators Actually DrawNoise, denoising, and guidance: the process behind every generated image you've seen.10 min→04Video Models: Why Motion Is So Much HarderTemporal consistency, world models, and what current video generators actually learned.10 min→05Document AI: How Models Read PDFs, Tables, and ReceiptsOCR to layout understanding: the unglamorous vision task every business needs.13 min→06Multimodal Agents: When Vision Meets ActionScreenshots, UIs, and the computer-use agents that see what you see.12 min→