用 Word2Vec 入门词嵌入:CBOW 与 Skip-gram 原理及 gensim 示例
My First Steps into Word Embeddings with Word2Vec
Word2Vec 通过上下文学习词向量,把词表示为数字向量,让上下文相似的词获得相近向量。它有两种训练方式:CBOW 用周围词预测缺失词,Skip-gram 用单个词预测周围词。文中用 gensim 的 Word2Vec 在 5 句示例语料上演示训练,并指出训练数据太少时结果可能不可靠。
Explain what word embeddings are and why they matter: so, Word embeddings are numerical representations of words. Each word is represented by a list of numbers called a vector.
For example, imagine representing two words using simplified vectors:
happy = [0.8, 0.7, 0.2]
joyful = [0.7, 0.8, 0.3]
These numbers are only illustrative, not real trained embeddings.
And why they matter :Word embeddings are useful because many NLP tasks require a computer to process language beyond simply identifying individual words.
- Cover at least one embedding method which is Word2Vec.so,Word2Vec is a popular technique used in Natural Language Processing (NLP) to learn word embeddings. It converts words into numerical vectors by learning from the contexts in which words appear in a large collection of text.
Word2Vec is a method that teaches a computer to understand relationships between words by learning from the words that appear around them.
- How does it work?
. It reads sentences
For example: “I love learning Python.”
Word2Vec looks at the words and how they appear together.
How does it work?
. It learns from context
If it sees “I love learning Python” and “I enjoy learning Python,” it notices that love and enjoy appear in similar contexts.
. It converts words into numbers
During training, Word2Vec learns a numerical vector for each word. Words used in similar contexts may end up with similar vectors.
. It learns relationships
After training, we can use the vectors to find words that are similar in context.
Two ways Word2Vec learns
CBOW: Uses surrounding words to predict a missing word.
Skip-gram: Uses one word to predict the surrounding words.
Sentences → Context → Training → Word vectors → Word relationships`
Word2Vec does not understand language exactly like a human. It learns patterns from the text it receives. If its training data is too small, its results may not be reliable.
. Include at least one concrete example: a code snippet, a similarity result, an analogy, or a
visualisation.
<>
from gensim.models import Word2Vec
Example training sentences
sentences = [
["i", "love", "learning", "python"],
["i", "enjoy", "learning", "python"],
["i", "love", "learning", "mathematics"],
["python", "is", "interesting"],
["mathematics", "is", "interesting"]
]
4.Share something you learned, found surprising, or struggled with :
What I Learned and Found Surprising
One thing I learned is that computers can learn relationships between words by studying the words that appear around them. I found it surprising that Word2Vec can convert words into numerical vectors and use those vectors to identify words that are used in similar contexts.
At first, I struggled to understand how words could be represented by numbers and how those numbers could show relationships between words. After learning about CBOW and Skip-gram, I understood that Word2Vec learns from context: CBOW predicts a word from its surrounding words, while Skip-gram predicts surrounding words from a given word.
This experience helped me understand how Natural Language Processing (NLP) allows computers to work with human language. I also learned that the quality of the results depends on the amount and quality of the training data.
Learning about Word2Vec has helped me understand an important concept in Natural Language Processing. I now understand how a model can learn word vectors by studying surrounding words and how these vectors can be used to explore similarities between words.
来源:Google AI:DEV 作者专属(RSS) · dev.to