I had this idea to use D3 Word Cloud from a prior notebook and have it ‘drive’ the indexing into a video hosting an interview with Mike. I’m already thinking of a future variation that dynamically leverages the API into HowTo100M.
Developed in about six hours with Gemini. Probably would have taken me three times as long without AI-assist and I would not have documented/commented it as much.
https://observablehq.com/@ambassadors/visualizing-bostocks-insights-with-a-word-cloud
Update: added a second AI-built pipeline + a live comparison toggle
Since posting this, I worked with the Observable notebook assistant to push the project further — turns out there’s more than one reasonable way to teach a word cloud what actually matters in a transcript.
The original word-extraction logic (built with Gemini) ranks words by simple frequency, filtered through one large hand-tuned stop-word list. It works well, but I was curious whether a different AI-assisted approach could tackle the same problem more elegantly — so I asked, and we ended up building a genuinely different second pipeline instead of just patching the first one:
Instead of one big stop-word list, candidates are filtered using precise part-of-speech tags (via Compromise NLP), backed by two tiny, purpose-specific exception lists — one for short technical acronyms (“D3”, “AI”), one for unavoidable spoken filler words (“like”, “you know”).
Instead of ranking by raw frequency, common words are ranked by semantic similarity to an embedding of the entire talk, computed in-browser with a small sentence-transformer model (Xenova/all-MiniLM-L6-v2 via transformers.js) — no API key, no server calls.
Person/company names get their own separate ranking path, since a name doesn’t “semantically” resemble a talk’s topic even when it’s clearly important to keep.
You can now flip between both pipelines live with a “Words Source” toggle sitting right above the word cloud.
Honestly, the most interesting part wasn’t the final result — it’s that both approaches had real, concrete failure modes that only surfaced once we actually tested them against the real transcript:
The frequency-based pipeline let a meaningless, mistagged word dominate the cloud purely through repetition (one word was said 74 times and meant nothing).
The semantic-similarity pipeline systematically under-ranked real names, since a guest’s name isn’t “about” the topic of the talk even though it’s obviously relevant.
Neither AI assistant got it right on the first try — both mistakes were only caught by generating the output and actually looking at it, not by trusting the design going in. That, more than the code itself, was the real lesson here.
Update 2: This turns out to be an excellent case study in Human/AI collaboration. After viewing the video again, I concluded that the main theme of Mike’s interview was not being captured in the word cloud. The ‘Acknowledgments & Collaboration’ section of the notebook discusses this ‘gap’ and how it was addressed with ‘developer’ value-added content to the AI logic. Ironically, the VERY TOPIC of the video itself!
Updated notebook (same link): https://observablehq.com/@ambassadors/visualizing-bostocks-insights-with-a-word-cloud
