register · ep 07 · jul 17
SEASON02 SHIPPED06 UPCOMING01 VENUEYOUTUBE LIVE

SHOW US YOUR [ AGENT ] SKILLS

[ vanishing / gradients ] × PyMC LABS
EXCEL WORLD CHAMPIONSHIPS × EUROVISION

Long-time Python, ML, AI and data builders show how they're actually using agents today. Not vibe coding. Not demos. The real workflows: agent skills, harnesses, voice-memo memory, background reviewers, from people whose software you've been using for years.

SEASON
02
EPISODES
06
BUILDERS
22
SKILLS · WORKFLOWS
61 · 87
NEXT EPISODE JUL 17 · 2026 friday · live on youtube

EP 07 · COMING UP

Greg Ceccarelli and Han-Chung Lee on what they are building with AI agents
HOSTS HUGO BOWNE-ANDERSON · THOMAS WIECKI
REGISTER ▶
EP 06
· JUL 2026 · 02 GUESTS · FULL EPISODE

AGENTS IN THE DATA STACK

Running a company on agent-readable context, handwritten AGENTS.md files, telemetry agents can inspect, an always-on personal agent, and workflow code with human checkpoints.

EP 06 · AGENTS IN THE DATA STACK - company context, feedback systems, long agent turns, and workflows with human checkpoints
▣ HIGHLIGHT REEL
01
Agents read a directory describing the whole company: accounting, legal, customers, contracts. They found contracted customers who were never billed. · Matt Rocklin
00:14:23
02
Decide what evidence counts before the work starts, walk away for an hour or two, then review the evidence instead of the diff. · Matt Rocklin
00:22:34
03
Skylar's personal agent, always on and running on a Mac Mini, runs an 80-person tech community, answers email, and keeps working while he is away. · Skylar Payne
00:45:36
04
Procedures become Python workflow code with typed outputs; ask() puts the human's decision inside the program. · Skylar Payne
01:03:48
05
Skylar's skill for writing new agent workflows: agent steps, artifacts, parallel work, and human review checkpoints. · Skylar Payne
01:04:49
EP 05
· JUN 2026 · 03 GUESTS · FULL EPISODE

COPILOTS & CODING AGENTS

Rook, Raw2Draft, MCut, Conductor, private skills, and the 90/10 handoff.

EP 05 · COPILOTS & CODING AGENTS - portable agents, editable artifacts, video timelines, and personal tools
▣ HIGHLIGHT REEL
01
Rook follows John from Obsidian to Wikipedia to the grocery-store thought experiment, borrowing local affordances at each stop. · John Berryman
00:18:39
02
A page-local skill: ask Rook to find the original BERT model-size passage and highlight the evidence in place. · John Berryman
00:24:58
03
Matt augments the Notion MCP with his own page-formatting taste: callouts, toggles, tables, and color. · Matt Palmer
01:41:31
04
Matt's revision skill pulls in Williams and Bizup so the agent edits prose against a real writing tradition. · Matt Palmer
01:42:56
05
MCut exposes tracks, transcript, timestamps, and timeline state so an agent can work against the same video session Matt edits. · Matt Palmer
00:54:20
06
Push the model toward cuts, plainness, and critique before the human adds taste. · Isaac Flath
01:52:02
07
Let the agent get the artifact most of the way there, then keep the rendered thing editable for the human final pass. · Isaac Flath
01:25:52
EP 04
· MAY 2026 · 03 GUESTS · FULL EPISODE

HOW TO EVALUATE AGENTIC WORKFLOWS

Skill scepticism, plan review, implementation review, agentic search, and hidden holdout tests.

EP 04 · HOW TO EVALUATE AGENTIC WORKFLOWS - skill scepticism, review loops, and hidden holdout tests
▣ HIGHLIGHT REEL
01
Read public skills like code: check provenance, maintenance, and constraints, then fork the idea instead of importing a shortcut. · Hamel Husain
00:22:32
02
Turn cheap experimentation into checkpoints: plan first, review red and yellow flags, implement, then review the code before trust. · Chris Fonnesbeck
01:05:53
03
Give the agent room to mutate a search ranker, but keep the final score hidden so improvements have to survive a real holdout. · Doug Turnbull
01:41:07
EP 03
· MAY 2026 · 06 GUESTS · RUNTIME 3H 18M

FROM SKILLS TO AGENT HARNESSES

Research memory, local boxes, debug panes, live notebooks, video generation, and code repair.

EP 03 · FROM SKILLS TO AGENT HARNESSES - research memory, debug surfaces, notebooks, video, and code repair
▣ HIGHLIGHT REEL
01
A narrow audit pass for the failure mode agents love: broad try blocks and exception handlers that make bad code look green. · Matthew Honnibal
00:12:09
02
Instead of hunting today's bugs, write incident reports for the failures a future reasonable edit could cause. · Matthew Honnibal
00:14:10
03
Deliberately break the code, one mutation at a time, to find the bugs your test suite would let through. · Matthew Honnibal
00:14:10
04
here-now skill
Collapse publishing into one instruction: the agent ships HTML pages and small sites to live URLs without a GitHub Pages detour. · Eleanor Berger
00:45:55
05
Drive Anki through its local API with confirmation checks, so an agent can maintain flashcards without silently mutating memory. · Eleanor Berger
00:49:46
06
Give coding agents a design language: fewer generic panels, more interfaces that look like someone meant it. · Eleanor Berger
00:50:02
07
Asked once from a phone; the agent invented browser login, transcript fetching, caching, and secret gist summaries. · Eleanor Berger
00:52:57
08
Use failed threads as harness training data: trace missteps to instructions, then delete or sharpen the rule that caused them. · Nicolay Gerold
01:59:04
09
Record a few minutes of audio; let the skill carry video taste, timing rules, frame checks, and avatar compositing. · Alan Nichol
02:46:00
10
research skill
Turn trusted sources into a durable research wiki, so future agents query accumulated context instead of starting over. · Paul Iusztin
02:19:52
11
Run a personal agent on a separate Mac mini, with Discord as the front door and autonomy earned inside a hard perimeter. · Eleanor Berger
00:47:50
EP 02
· MAY 2026 · 04 GUESTS · RUNTIME 2H 14M

BUILDING AGENTS THAT IMPROVE THE WORKFLOW

Prompt refinement, eval-driven charts, human-in-the-loop EDA, and local-first inference.

EP 02 · BUILDING AGENTS THAT IMPROVE THE WORKFLOW - prompt refinement, eval-driven charts, EDA, and local-first inference
▣ HIGHLIGHT REEL
01
Interview intent first, then generate risky variants and score them against a rubric written before the run. · Hilary Mason
01:01:00
02
Drop the agent inside a live Marimo kernel, so plots, widgets, markdown, and corrections happen in one reactive notebook. · Eric Ma
00:11:57
03
agentic-eda workflow
The human chooses the next scientific question; the agent renders evidence fast enough to keep exploration one plot at a time. · Eric Ma
00:23:27
04
Turn every failed chart eval into a library feature, so the package cannot regress on a case it already learned. · Bryan Bischof
01:25:11
05
Schedule three bad-idea personas to pitch, critique, and write moonshot docs no product roadmap would allow. · Hilary Mason
01:14:20
06
Default local: fast Qwen on a laptop, private workflows, offline flights, and cloud calls only for named exceptions. · Tomasz Tunguz
02:07:42
EP 01
· APR 2026 · 03 GUESTS · RUNTIME 1H 32M

THE AGENTIC SOFTWARE FACTORY

RoboRev, agent memory, personal commands, and LLM-as-judge chart checks.

EP 01 · THE AGENTIC SOFTWARE FACTORY - RoboRev, agent memory, personal commands, and LLM-as-judge checks
▣ HIGHLIGHT REEL
01
explain skill
When ten agents are running, each one explains the change like a colleague, not a diff bot. · Jeremiah Lowin
00:46:14
02
A tiny etiquette layer for OSS maintenance: reject clearly, sound human, and stop wrapping "no" in fake praise. · Jeremiah Lowin
00:54:08
03
ship-it skill
One phrase, one override: in Jeremiah's world "ship it" means open the PR, never merge it. · Jeremiah Lowin
00:54:52
04
Run search, chart variants, linting, and an LLM-as-judge Tufte test until the graphic actually carries the story. · Randy Olson
01:12:37
05
The show's own production skill: turn guest photos into pixel art, animate them, and feed the retro livestream system. · Show Us Your Agent Skills
ep 01
06
Scale agentic engineering with commits every turn, RoboRev reading every line, and a review queue agents must drain. · Wes McKinney
00:27:14
07
second-brain workflow
Feed daily voice memos into editable agent memory, turning personal context into a substrate future sessions can use. · Jeremiah Lowin
00:35:50

GUEST DOSSIERS

one page per builder: segment video, field notes, artifacts, and the workflow they showed
Matt Rocklin MATT ROCKLINEP 06the whole company as context, AGENTS.md, and feedback systems Skylar Payne SKYLAR PAYNEEP 06an always-on personal agent, executable workflows, and ask() checkpoints John Berryman JOHN BERRYMANEP 05Rook, page-local skills, and agents that follow you Isaac Flath ISAAC FLATHEP 05Raw2Draft, editable artifacts, and the 90/10 handoff Matt Palmer MATT PALMEREP 05agent-edited video timelines and personal tools that last Hamel Husain HAMEL HUSAINEP 04skill scepticism, internal APIs, and constraints over prose Chris Fonnesbeck CHRIS FONNESBECKEP 04review loops for Bayesian modeling and agent-written code Doug Turnbull DOUG TURNBULLEP 04agentic search, hidden validation, and anti-overfit evals Paul Iusztin PAUL IUSZTINEP 03research skills, trusted sources, and a writing knowledge base Eleanor Berger ELEANOR BERGEREP 03Hermes, fnord, local boxes, and agent boundaries Alan Nichol ALAN NICHOLEP 03programmatic video, Remotion, and command-line production Vincent Warmerdam VINCENT WARMERDAMEP 03marimo pair, widgets, and notebooks as shared state Nicolay Gerold NICOLAY GEROLDEP 03debug surfaces, thread postmortems, and reusable context Matthew Honnibal MATTHEW HONNIBALEP 03try-except repair, mutation testing, and self-review loops Hilary Mason HILARY MASONEP 02prompt refinement, weekly gremlins, and taste loops Bryan Bischof BRYAN BISCHOFEP 02eval-driven charts, ratchets, and chart judgment Eric Ma ERIC MAEP 02agentic EDA, marimo, and human-in-the-loop plots Tomasz Tunguz TOMASZ TUNGUZEP 02local-first agents, earnings analysis, and fast local models Wes McKinney WES MCKINNEYEP 01RoboRev, long-running agents, and agentic engineering Jeremiah Lowin JEREMIAH LOWINEP 01personal software, explain, ship-it, and GitHub replies Randy Olson RANDY OLSONEP 01Tufte checks, chart workflows, and LLM-as-judge review

SELECTED SKILLS & WORKFLOWS

a selection from the companion repo · hugobowne/show-us-your-agent-skills ↗

A selection of skills and workflows from the streams, packaged into the companion repo.

$ npx skills add https://github.com/hugobowne/show-us-your-agent-skills
skill
explain
EP 01
Agent narrates what it just did, like a teammate handing off.
Jeremiah Lowin · Prefect / FastMCP ↗ youtube · 00:46:14
skill
Replies to GitHub contributors in your voice. No "great work, but rejected" sandwiches.
Jeremiah Lowin · Prefect / FastMCP ↗ youtube · 00:54:08
skill
ship-it
EP 01
Re-trains "ship it" to mean open a PR, not merge.
Jeremiah Lowin · Prefect / FastMCP ↗ youtube · 00:54:52
skill
One-line idea → Tufte-style chart with an LLM-as-judge verifier loop.
Randy Olson · Goodeye Labs · r/dataisbeautiful ↗ youtube · 01:12:37
skill
Turns guest headshots into 8-bit pixel-art video clips for livestream intros and cutaways.
Show Us Your Agent Skills ↗ ep 01 on youtube
skill
Interview intent, ask for three variations at different magnitudes, score against a rubric.
Hilary Mason · Hidden Door ↗ youtube · 01:01:00
skill
Agent drives a reactive Marimo notebook through a bash bridge into the Python kernel.
Eric Ma · Moderna ↗ youtube · 00:11:57
workflow
Human-in-the-loop EDA. Every claim backed by an artifact.
Eric Ma · Moderna ↗ youtube · 00:23:27
workflow
Build an agent-facing chart library by generalising eval failures into features; the package can never regress on an eval it once passed.
Bryan Bischof · Theory Ventures ↗ youtube · 01:25:11
workflow
Three agent personas pull from a bad-ideas backlog, pitch and critique each other, and write design docs for moonshots no roadmap would schedule.
Hilary Mason · Hidden Door ↗ youtube · 01:14:20

YOUR HOSTS

two builders · two hosts · on a mission to find out what people at the top of the game are actually doing
Hugo Bowne-Anderson
HUGO
BOWNE-ANDERSON
host · vanishing gradients
AI builder, consultant, educator of 6+ million students; ex-Yale, ex-Max Planck.
Thomas Wiecki
THOMAS
WIECKI
host · pymc labs
Co-creator of PyMC. Founder of PyMC Labs. Has built Bayesian models for hedge funds, Fortune 500s, and indie tinkerers for over a decade.
DON'T MISS THE NEXT
EPISODE.

Episode announcements, the workflows we cut for time, and what long-time builders are actually doing with agents. Free, no spam.

▶ SUBSCRIBE ON SUBSTACK ↗

Doing something weird with agents? Nominate yourself or someone else ↗