# word2vec

**URL:** <https://talk.observablehq.com/t/word2vec/5378>\
**Category:** Show and tell\
**Created:** [July 26, 2021, 9:43pm UTC](https://talk.observablehq.com/t/word2vec/5378 "2021-07-26T21:43:29Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![tomlarkworthy](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/tomlarkworthy/32/5940_2.png) [@tomlarkworthy](https://talk.observablehq.com/u/tomlarkworthy)\
**Post date:** [July 26, 2021, 9:43pm UTC](https://talk.observablehq.com/t/word2vec/5378/1 "2021-07-26T21:43:29Z")

</div>

Sometimes I want to analyse text but its kind of a pain. I thought word2vec would be useful but its not so easy to find online… so I created this [word2Vec / Endpoint Services / Observable](https://observablehq.com/@endpointservices/word2vec) its the full 3GB 2.7M word word2vec accessible via API so the notebook can sample from it without having to host it.

---

<div class="post-metadata">

**Author:** ![cornhundred](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/cornhundred/32/1705_2.png) [@cornhundred](https://talk.observablehq.com/u/cornhundred)\
**Post date:** [July 27, 2021, 12:44pm UTC](https://talk.observablehq.com/t/word2vec/5378/2 "2021-07-27T12:44:18Z")

</div>

This is cool thanks for sharing!! I want to try to combine this with on the fly clustering and heatmap visualization from this notebook [Python (Pyodide) on Observable running Clustergrammer2](https://talk.observablehq.com/t/python-pyodide-on-observable-running-clustergrammer2/5363)

---

<div class="post-metadata">

**Author:** ![tomlarkworthy](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/tomlarkworthy/32/5940_2.png) [@tomlarkworthy](https://talk.observablehq.com/u/tomlarkworthy)\
**Post date:** [July 27, 2021, 1:00pm UTC](https://talk.observablehq.com/t/word2vec/5378/3 "2021-07-27T13:00:17Z")

</div>

yeah I saw that! It would go neatly together I think too.

@mootari also found a quite nice off-the-shelf clustering + word2vec application [https://wikipedia2vec.github.io/demo/](https://wikipedia2vec.github.io/demo/) which provides some intuition on how the words should cluster.

an issue with the google word2vec is its full of garbage so it’s better to start with a clean corpus of words, lookup their vectors then cluster.

---

<div class="post-metadata">

**Author:** ![dkirkby](https://yyz2.discourse-cdn.com/flex030/user_avatar/talk.observablehq.com/dkirkby/32/3775_2.png) [@dkirkby](https://talk.observablehq.com/u/dkirkby)\
**Post date:** [December 31, 2021, 9:50pm UTC](https://talk.observablehq.com/t/word2vec/5378/4 "2021-12-31T21:50:11Z")

</div>

For a fun application of word embedding vectors, I implemented an AI assistant for the game CodeNames in this notebook [Cipher Words / David Kirkby / Observable](https://observablehq.com/@dkirkby/cipher-words)

(this is using 100-dimensional [GloVe vectors pre-trained on Wikipedia](https://nlp.stanford.edu/projects/glove/)).
