---
title: "Chinese-Word-Vectors vs wikipedia2vec"
type: "comparison"
canonical_url: "https://www.graphcanon.com/compare/embedding-chinese-word-vectors-vs-wikipedia2vec-wikipedia2vec"
tools: ["embedding-chinese-word-vectors", "wikipedia2vec-wikipedia2vec"]
---

# Chinese-Word-Vectors vs wikipedia2vec

*GraphCanon updated Aug 22, 2026*

## Verdict

Pick Chinese-Word-Vectors if chinese-Word-Vectors offers over 100 pre-trained Chinese word vectors for various NLP tasks; pick wikipedia2vec if a Python-based tool for generating embeddings derived from Wikipedia content.

[Chinese-Word-Vectors](https://github.com/Embedding/Chinese-Word-Vectors) reports 12k GitHub stars, 2.3k forks, and 60 open issues, last pushed Oct 30, 2023. [wikipedia2vec](http://wikipedia2vec.github.io/) has 971 stars, 100 forks, and 8 open issues, last pushed May 3, 2024. Figures are from public GitHub metadata via [Chinese-Word-Vectors's repository](https://github.com/Embedding/Chinese-Word-Vectors) and [wikipedia2vec's repository](https://github.com/wikipedia2vec/wikipedia2vec).

| | [Chinese-Word-Vectors](/tools/embedding-chinese-word-vectors.md) | [wikipedia2vec](/tools/wikipedia2vec-wikipedia2vec.md) |
| --- | --- | --- |
| Tagline | 上百种预训练中文词向量 | A tool for learning vector representations of words and entities from Wikipedia |
| Stars | 12,227 | 971 |
| Forks | 2,323 | 100 |
| Open issues | 60 | 8 |
| Language | Python | Python |
| Adopt for | Chinese-Word-Vectors offers over 100 pre-trained Chinese word vectors for various NLP tasks. | A Python-based tool for generating embeddings derived from Wikipedia content. |
| Persona | - | - |
| Runtime | - | - |
| License | Apache-2.0 | Other |
| Categories | Data & Retrieval | Vector Databases |

## Trust and health

_Sourced signals - not a safety guarantee. No winner column._

| | [Chinese-Word-Vectors](/tools/embedding-chinese-word-vectors.md) | [wikipedia2vec](/tools/wikipedia2vec-wikipedia2vec.md) |
| --- | --- | --- |
| Days since push | 1026d | 840d |
| Open issues (now) | 60 | 8 |
| Stars delta | -3 (30d) | +4 (30d) |
| Owner type | User | Organization |
| Full report | [trust report](/tools/embedding-chinese-word-vectors/trust.md) | [trust report](/tools/wikipedia2vec-wikipedia2vec/trust.md) |

## Decision facts: Chinese-Word-Vectors

- **Adopt for:** Chinese-Word-Vectors offers over 100 pre-trained Chinese word vectors for various NLP tasks.

## Decision facts: wikipedia2vec

- **Adopt for:** A Python-based tool for generating embeddings derived from Wikipedia content.

## Choose when

### Choose Chinese-Word-Vectors if…

- License: Chinese-Word-Vectors is Apache-2.0, wikipedia2vec is Other.
- Tags unique to Chinese-Word-Vectors: chinese, embedding, word-embeddings.
- Also covers Data & Retrieval.
- Use when you specifically require a wide variety (over 100) of pre-trained Chinese word vectors tailored to different aspects of natural language processing within your project.

### Choose wikipedia2vec if…

- License: wikipedia2vec is Other, Chinese-Word-Vectors is Apache-2.0.
- Tags unique to wikipedia2vec: embeddings, natural-language-processing, nlp, python.
- Also covers Vector Databases.
- You need to generate word and entity embeddings based on extensive Wikipedia data

## When NOT to use Chinese-Word-Vectors

- Avoid using this tool if your project requires fine-tuning on a very specific domain that is not well-represented among the existing 100+ models provided.
- This collection may not be optimal if you are working strictly within English or other languages, as it focuses primarily on Chinese word vectors.

## When NOT to use wikipedia2vec

- Your dataset doesn't intersect with or benefit from Wikipedia content
- You require real-time updating capabilities that exceed static Wikipedia dumps

## Common questions

### What is the difference between Chinese-Word-Vectors and wikipedia2vec?

Chinese-Word-Vectors: 上百种预训练中文词向量. wikipedia2vec: A tool for learning vector representations of words and entities from Wikipedia. See the comparison table for live GitHub stats and shared categories.

### When should I choose Chinese-Word-Vectors over wikipedia2vec?

Choose Chinese-Word-Vectors over wikipedia2vec when License: Chinese-Word-Vectors is Apache-2.0, wikipedia2vec is Other; Tags unique to Chinese-Word-Vectors: chinese, embedding, word-embeddings; Also covers Data & Retrieval; Use when you specifically require a wide variety (over 100) of pre-trained Chinese word vectors tailored to different aspects of natural language processing within your project.

### When should I choose wikipedia2vec over Chinese-Word-Vectors?

Choose wikipedia2vec over Chinese-Word-Vectors when License: wikipedia2vec is Other, Chinese-Word-Vectors is Apache-2.0; Tags unique to wikipedia2vec: embeddings, natural-language-processing, nlp, python; Also covers Vector Databases; You need to generate word and entity embeddings based on extensive Wikipedia data.

### When should I avoid Chinese-Word-Vectors?

Avoid using this tool if your project requires fine-tuning on a very specific domain that is not well-represented among the existing 100+ models provided. This collection may not be optimal if you are working strictly within English or other languages, as it focuses primarily on Chinese word vectors.

### When should I avoid wikipedia2vec?

Your dataset doesn't intersect with or benefit from Wikipedia content You require real-time updating capabilities that exceed static Wikipedia dumps

### Is Chinese-Word-Vectors or wikipedia2vec more popular on GitHub?

Chinese-Word-Vectors has more GitHub stars (12,227 vs 971). Stars measure visibility, not whether either tool fits your constraints.

### Are Chinese-Word-Vectors and wikipedia2vec open source?

Yes - both are open-source projects on GitHub (Chinese-Word-Vectors: Apache-2.0, wikipedia2vec: Other).

### Where can I find alternatives to Chinese-Word-Vectors or wikipedia2vec?

GraphCanon lists graph-backed alternatives at [Chinese-Word-Vectors alternatives](/tools/embedding-chinese-word-vectors/alternatives) and [wikipedia2vec alternatives](/tools/wikipedia2vec-wikipedia2vec/alternatives) ([Chinese-Word-Vectors markdown twin](/tools/embedding-chinese-word-vectors/alternatives.md), [wikipedia2vec markdown twin](/tools/wikipedia2vec-wikipedia2vec/alternatives.md)), ranked by typed relationship edges rather than popularity votes.

### Is there a machine-readable version of this comparison?

Yes. The markdown twin at [this comparison](/compare/embedding-chinese-word-vectors-vs-wikipedia2vec-wikipedia2vec.md) mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.

### Which is better maintained, Chinese-Word-Vectors or wikipedia2vec?

Chinese-Word-Vectors: Dormant. wikipedia2vec: Dormant. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.

### Where are the full trust reports for Chinese-Word-Vectors and wikipedia2vec?

GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: [Chinese-Word-Vectors trust report](/tools/embedding-chinese-word-vectors/trust); [wikipedia2vec trust report](/tools/wikipedia2vec-wikipedia2vec/trust).

---

**Machine-readable endpoints**

- JSON: [`/api/graphcanon/graph?tool=embedding-chinese-word-vectors`](/api/graphcanon/graph?tool=embedding-chinese-word-vectors)
- LLM index: [/llms.txt](/llms.txt)
- Full corpus: [/llms-full.txt](/llms-full.txt)

_GraphCanon - The knowledge graph for AI development. https://www.graphcanon.com/_
