My First Observable 2.0 Notebook

I had this idea to use D3 Word Cloud from a prior notebook and have it ‘drive’ the indexing into a video hosting an interview with Mike. I’m already thinking of a future variation that dynamically leverages the API into HowTo100M.

Developed in about six hours with Gemini. Probably would have taken me three times as long without AI-assist and I would not have documented/commented it as much.

https://observablehq.com/@ambassadors/visualizing-bostocks-insights-with-a-word-cloud

Update: added a second AI-built pipeline + a live comparison toggle

Since posting this, I worked with the Observable notebook assistant to push the project further — turns out there’s more than one reasonable way to teach a word cloud what actually matters in a transcript.

The original word-extraction logic (built with Gemini) ranks words by simple frequency, filtered through one large hand-tuned stop-word list. It works well, but I was curious whether a different AI-assisted approach could tackle the same problem more elegantly — so I asked, and we ended up building a genuinely different second pipeline instead of just patching the first one:

Instead of one big stop-word list, candidates are filtered using precise part-of-speech tags (via Compromise NLP), backed by two tiny, purpose-specific exception lists — one for short technical acronyms (“D3”, “AI”), one for unavoidable spoken filler words (“like”, “you know”).

Instead of ranking by raw frequency, common words are ranked by semantic similarity to an embedding of the entire talk, computed in-browser with a small sentence-transformer model (Xenova/all-MiniLM-L6-v2 via transformers.js) — no API key, no server calls.

Person/company names get their own separate ranking path, since a name doesn’t “semantically” resemble a talk’s topic even when it’s clearly important to keep.

You can now flip between both pipelines live with a “Words Source” toggle sitting right above the word cloud.

Honestly, the most interesting part wasn’t the final result — it’s that both approaches had real, concrete failure modes that only surfaced once we actually tested them against the real transcript:

The frequency-based pipeline let a meaningless, mistagged word dominate the cloud purely through repetition (one word was said 74 times and meant nothing).

The semantic-similarity pipeline systematically under-ranked real names, since a guest’s name isn’t “about” the topic of the talk even though it’s obviously relevant.

Neither AI assistant got it right on the first try — both mistakes were only caught by generating the output and actually looking at it, not by trusting the design going in. That, more than the code itself, was the real lesson here.

Update 2: This turns out to be an excellent case study in Human/AI collaboration. After viewing the video again, I concluded that the main theme of Mike’s interview was not being captured in the word cloud. The ‘Acknowledgments & Collaboration’ section of the notebook discusses this ‘gap’ and how it was addressed with ‘developer’ value-added content to the AI logic. Ironically, the VERY TOPIC of the video itself!

Updated notebook (same link): https://observablehq.com/@ambassadors/visualizing-bostocks-insights-with-a-word-cloud

The interlinking of topics and video-scrubbing is cool! Honestly I never watched the whole video at the time it came out and I found myself exploring more of it just playing around with this.

It makes me wonder about a version where, instead of a word cloud, you have a full-width video and a topic timeline below it. You could give each topic its own row and highlight parts of the timeline that cover it, or you could… idk.

(Like, when I see the word cloud, I wish that instead the x and y coordinates of the words were encoding meaningful information.)

This is what Observable AI Assistant came back with when I asked it for some creative visualizations:

Great question — you’ve already built most of the hard infrastructure (timestamped cues, POS-tagged nouns, frequency counts, and a full sentence-embedding pipeline via derivedDocEmbedding/embedBatch/cosine), so a lot of these could reuse existing cells rather than starting from scratch. Here’s a tour organized by what kind of story each one tells, from data you already have vs. data you’d need to derive.

Here’s a quick prototype in an unlisted forked notebook: https://observablehq.com/d/e09f5c9b83ce5112

See more from the conversation below:

1. Time-shaped visualizations (the video as an X-axis)

Concept-density ridgeline / joyplot — my favorite, given the notebook’s own narrative. You already have the machinery to detect that “interpretability,” “verifiability,” and “trust” are discussed in a concentrated stretch around 50:58. Generalize that: bin the transcript into ~2-minute windows, and for each of your top N words/themes compute mention density (or cosine similarity of that window’s chunk embedding to each theme’s own embedding). Stack small areaY ridges, one per theme, so viewers can literally see where in the 69 minutes each topic clusters — clicking anywhere on a ridge seeks there. This turns your “human override” story into a visual: the interpretability/verifiability/trust ridge would show a visible spike right at 50:58 that no frequency-only chart would surface.

“Story shape” sentiment/energy arc — a single line (à la Kurt Vonnegut’s story-shape lectures) plotting some derived signal per time-chunk — sentiment polarity, speaking pace (words/minute from parsedTranscript cue lengths ÷ duration), or even embedding-drift (cosine distance between consecutive chunk embeddings, i.e. “how much is the topic changing right now”). Peaks/valleys become natural chapter markers; clicking a point seeks.

Mention barcode / rug timeline — like a genome browser or a Spotify lyrics scrubber: one horizontal row per selected word, with a tick mark (tickX or thin rect) at every timestamp in its startTimes. Stack several important words’ rows and you get an at-a-glance “when do these ideas cluster together” view — great complement to the cloud, since the cloud discards when entirely.

2. Relationship / network visualizations

Force-directed co-occurrence graph — nodes = your existing word list (reuse count/freq for radius), edges = “appeared in the same transcript cue” or “within N seconds of each other,” weight = co-occurrence count. This is very D3-idiomatic (d3.forceSimulation) and reveals clusters of ideas that travel together rather than just individually-frequent words. Click a node to seek to its first/next mention (same interaction model as the cloud); click an edge to seek to the moment two ideas were mentioned together.

Arc diagram over the timeline — draw the transcript as a single horizontal line (start → end of video), then draw small arcs above it connecting every pair of repeated mentions of the same word/theme, colored by word. This is a lovely visual complement to your ridgeline: instead of density, it emphasizes recurrence — you’d literally see “trust” arc back and forth across the whole hour, versus a word that’s said three times in a row and never again.

3. Semantic-space visualizations (you already compute the embeddings!)

2D “semantic map” scatterplot — you already embed every candidate word via embedBatch. Project those ~75+ vectors into 2D (a small D3-friendly PCA, or import umap-js), plot as a dot/custom D3 scatter, color by cluster (k-means over the same vectors). Words that mean similar things sit near each other spatially — a genuinely different lens than font-size-by-relevance. Click a dot to seek, same as today.

Dendrogram of vocabulary — hierarchical-cluster the same word embeddings (agglomerative clustering by cosine distance) and render as a D3 tree/dendrogram. Great as a companion “table of contents”: collapsing/expanding branches could reveal how granular subtopics nest inside your bigger themes.

Sunburst / icicle of topic → subtopic → word — if you cluster embeddings at two levels (broad theme, then sub-cluster within it), a radial sunburst (d3.partition + arcs) gives a drill-down control: click the outer “AI coding agents” wedge, it flies to that region of video; click an inner wedge for a specific sub-idea.

4. Radial/time-encoded variants of the cloud itself

Radial timeline cloud — right now d3-cloud’s spiral placement is essentially decorative; position encodes nothing. A fun reinterpretation: place each word at an angle proportional to its first (or average) timestamp around a circle (like a 69-minute clock face), and radius encoding importance (similarity/frequency). You’d get all the visual appeal of a word cloud, but suddenly “when” is legible too — words from the video’s back half cluster on one side of the circle.

5. Flow visualizations

Sankey of topic transitions — cluster time-chunks into a small number of topic labels (k-means on chunk embeddings, same derivedDocEmbedding machinery), then treat consecutive chunks as a sequence and draw a Sankey (d3-sankey) of topic → topic transitions over the interview. This answers “does the conversation wander and return to earlier topics, or move linearly?” — a structural question the word cloud can’t answer at all.