Testing & regression-checking Claude Code agent extensions before you ship them
How to prove the skill, subagent, hook, plugin, or MCP server you authored actually works before your team installs it — and where, in 2026, the tooling still runs out.
Talks and workshops — for technical and non-technical audiences.
How to prove the skill, subagent, hook, plugin, or MCP server you authored actually works before your team installs it — and where, in 2026, the tooling still runs out.
Complete technical journey from Gaussian noise schedules through score networks and CLIP conditioning to production Midjourney v6 craft — with a post-mortem on why hands failed and how the same architectural fixes inform smarter prompting.
A layered defense guide for text-to-SQL pipelines: schema grounding, read-only enforcement, query validation, pgvector semantic search, and the silent correctness failures that survive all of them.
Complete blueprint for running a live head-to-head AI coding tool competition: task spec, team playbook, scoring rubric, debrief script, and a published comparison table — covering Claude Code, Cursor, Copilot, Codex CLI, and local-model alternatives.
A 90–120 min expert session blueprint covering the full RAG production stack — from embedding and naive failure modes through hybrid search, reranking, hallucination guardrails, and evaluation.
A complete session blueprint for expert developers covering golden datasets, LLM-as-judge validation, CI delta gates, and a live regression demo that catches what vibe checks miss.
Everything to run an evals session for a client-shipping consultancy: the methodology that actually matters, the 2026 tooling shake-up, CI gates, agent/RAG specifics, and a runnable 2h lab.
A session blueprint for teaching devs to build an MCP server: a demo-led live build of a zero-auth TypeScript query server, with the protocol, debugging, security, deployment, and the RAG question all scoped in.
The harness is the performance variable — seven design schools, 12 frameworks, 30-point benchmark spreads, and the evaluation discipline to run a fair comparison.
A runnable playbook for a 90-min AI-assisted TDD workshop targeting senior developers: exercise catalog, Codespaces scaffolding, facilitation roles, and honest evidence on what AI actually delivers in 2026.
Expert session blueprint: description-as-interface, progressive disclosure across all three extension layers, the client-side MCP primitives builders skip, tool-poisoning as the live security thread, and the token/cost mechanics that make these choices matter at team scale.
Session blueprint for an expert-audience deep-dive on extending Claude Code through MCP, Skills, Plugins, and the marketplaces around them — with a comparison table, live-demo recipes, and the trust-boundary callbacks a third-in-series session needs to earn its slot.
A facilitator's full playbook for a 2-3 h vibe coding workshop: shortlist of laymen-shippable apps, a Lovable/Bolt stack costed in EUR for shared wifi, a ChatGPT-refine → Lovable-build prompting drill, and a take-home path that survives the two-week tail.
Keep 'ProbLLMs: Why You Can't Trust the Robot', cut the developer-jargon half of the deck, and graft in four 2026 angles (voice-clone scams, kids + AI companions, AI-injected ads, deepfake nudes) that show up in actual public-concern polls.