跳到正文
原文
Google AI:DEV 作者专属(RSS)· Irakoze Deborah·· 4 小时前AI 评分40

不用词典,如何教计算机理解词义:Word2Vec 与词嵌入入门

Teaching a Computer Word Meaning Without a Dictionary

AI 导读

词嵌入把单词表示为一串数字向量,让上下文相似的词(如 cat 与 dog)在向量空间中彼此靠近,而无需任何词典或人工标注。文章用纯 Python 演示:one-hot 编码下 cat 与 dog、car 的相似度均为 0,而基于上下文计数的向量让 cat 与 dog 余弦相似度达 0.75、cat 与 car 仅 0.41。

正文

When I started learning Natural Language Processing (NLP), one question bugged me: how does a computer understand words?

The honest answer is that it doesn't. Computers only understand numbers. So before a model can work with language, we have to turn words into numbers, and the way we do it makes a huge difference. That is what word embeddings are about.

In this post I'll explain what they are, show tiny examples you can run in seconds with plain Python, and look at how Word2Vec learns them.

What is a word embedding?

A word embedding is a word written as a list of numbers (a vector):

cat → [0.21, 0.85, -0.13, 0.42, ...]

Nobody types these numbers in by hand. They are learned from text. The big idea is that words used in similar contexts get similar numbers.

The cat eats fish.
The dog eats meat.

"cat" and "dog" both eat things, so their vectors end up close together. A machine-learning model can then treat them as related instead of as complete strangers. That is why embeddings matter: the model can reuse what it knows about "cat" when it meets "dog".

The old way: one-hot encoding

The simplest approach gives each word a vector with a single 1 and the rest 0. Let's test how "related" these vectors are. The function below multiplies matching positions and adds them up (a dot product):

cat = [1, 0, 0]
dog = [0, 1, 0]
car = [0, 0, 1]

def similarity(a, b):
    return sum(x * y for x, y in zip(a, b))

print("cat vs dog:", similarity(cat, dog))
print("cat vs car:", similarity(cat, car))
print("cat vs cat:", similarity(cat, cat))

Output:

cat vs dog: 0
cat vs car: 0
cat vs cat: 1

A word is only similar to itself. To one-hot encoding, "cat" is as unrelated to "dog" as it is to "car". There is also a practical problem: with a 100,000-word vocabulary, every vector would have 100,000 slots, almost all of them zeros. Embeddings fix both issues by using short, dense vectors (usually 50 to 300 numbers) where distance carries meaning.

Why context is so powerful

Before any machine learning, let's see the core idea. This code collects the words that appear in the same sentence as each word:

sentences = [
    "the cat eats fish",
    "the cat drinks milk",
    "the dog eats meat",
    "the dog drinks water",
    "the car needs fuel",
]

context = {}
for sentence in sentences:
    words = sentence.split()
    for w in words:
        context.setdefault(w, set()).update(words)
        context[w].discard(w)

print("cat & dog share:", sorted(context["cat"] & context["dog"]))
print("cat & car share:", sorted(context["cat"] & context["car"]))

Output:

cat & dog share: ['drinks', 'eats', 'the']
cat & car share: ['the']

"cat" and "dog" share three context words, while "cat" and "car" share only one. Just by looking at neighbours, we discovered that cats are closer to dogs than to cars.

Word2Vec in 30 seconds

Word2Vec takes this same idea and scales it up. It slides a small window over text and learns from the words nearby. It comes in two flavours:

  • Skip-gram: given a center word, predict its neighbours. (cat → the, eats, fish)
  • CBOW (Continuous Bag of Words): given the neighbours, predict the center word. (the ___ eats fish → cat)

While the model gets better at these predictions, it adjusts its internal vectors. The useful by-product is that words appearing in similar contexts end up with similar vectors. With millions of sentences, famous results like king − man + woman ≈ queen start to appear.

Build a tiny embedding yourself

Real Word2Vec libraries are heavy, so let's build the core idea in plain Python. Each word becomes a vector of counts: how often every other word appears in a sentence with it. Then we compare vectors with cosine similarity (1 means same direction, 0 means nothing in common).

sentences = [
    "the cat eats fish",
    "the cat drinks milk",
    "the dog eats meat",
    "the dog drinks water",
    "the car needs fuel",
]

vocab = sorted({w for s in sentences for w in s.split()})

def vector(word):
    counts = {v: 0 for v in vocab}
    for s in sentences:
        words = s.split()
        if word in words:
            for w in words:
                if w != word:
                    counts[w] += 1
    return [counts[v] for v in vocab]

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / ((sum(x * x for x in a) ** 0.5) * (sum(y * y for y in b) ** 0.5))

print("vocab:", vocab)
print("cat vector:", vector("cat"))
print("cat vs dog:", round(cosine(vector("cat"), vector("dog")), 2))
print("cat vs car:", round(cosine(vector("cat"), vector("car")), 2))

Output:

vocab: ['car', 'cat', 'dog', 'drinks', 'eats', 'fish', 'fuel', 'meat', 'milk', 'needs', 'the', 'water']
cat vector: [0, 0, 0, 1, 1, 1, 0, 0, 1, 0, 2, 0]
cat vs dog: 0.75
cat vs car: 0.41

Unlike one-hot encoding, "cat" is now much closer to "dog" (0.75) than to "car" (0.41), and nobody told the program what any word means. Word2Vec goes further: instead of raw counts, it learns short dense vectors by predicting words from their neighbours, on millions of sentences.

What surprised me

The model is never told what a word means. There is no dictionary, no labels, and no "cat = animal". It only sees which words sit near which, and relationships appear on their own. I assumed a computer would need every word explained to it, so this was the most unintuitive part of the topic for me, and the most exciting once it clicked.

What I struggled with

I tried to install gensim, the popular Word2Vec library, and it failed on Windows with a "Microsoft Visual C++ 14.0 is required" build error. Instead of fighting the installer, I built a small count-based version in plain Python to show the same idea. It was a good reminder that understanding the concept matters more than the library.

Your turn: challenges

  • [ ] Add 3 more sentences about animals. Does cat vs dog go up?
  • [ ] Add a sentence like "the car needs oil". How does cat vs car change?
  • [ ] Try vector("fish") vs vector("milk"). Are they similar?
  • [ ] On a machine where it installs, try pip install gensim and compare with Word2Vec.

Share your numbers in the comments, I'd love to compare!

Ideas to explore next

  • GloVe and FastText: how do they differ from Word2Vec?
  • Famous analogies like king − man + woman ≈ queen
  • Plotting word vectors in 2D with PCA

Sources

来源:Google AI:DEV 作者专属(RSS) · dev.to