{"data":{"slug":"embedding-chinese-word-vectors","name":"Chinese-Word-Vectors","tagline":"上百种预训练中文词向量","github_url":"https://github.com/Embedding/Chinese-Word-Vectors","owner":"Embedding","repo":"Chinese-Word-Vectors","owner_avatar_url":"https://avatars.githubusercontent.com/u/31317254?v=4","primary_language":"Python","stars":12227,"forks":2323,"topics":["chinese","chinese-word-segmentation","embedding","embeddings","vectors-trained","word-embeddings"],"archived":false,"github_pushed_at":"2023-10-30T14:44:50+00:00","maintenance_label":"Dormant","stars_delta_30d":-3,"url":"https://www.graphcanon.com/tools/embedding-chinese-word-vectors","markdown_url":"https://www.graphcanon.com/tools/embedding-chinese-word-vectors.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/embedding-chinese-word-vectors","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=embedding-chinese-word-vectors","description":"100+ Chinese Word Vectors 上百种预训练中文词向量 ","homepage_url":null,"license":"Apache-2.0","open_issues":60,"watchers":279,"ai_summary":"提供超过100种预先训练的中文词向量，适用于各种自然语言处理任务。","readme_excerpt":"# Chinese Word Vectors 中文词向量\n[中文](https://github.com/Embedding/Chinese-Word-Vectors/blob/master/README_zh.md)\n\nThis project provides 100+ Chinese Word Vectors (embeddings) trained with different **representations** (dense and sparse), **context features** (word, ngram, character, and more), and **corpora**. One can easily obtain pre-trained vectors with different properties and use them for downstream tasks. \n\nMoreover, we provide a Chinese analogical reasoning dataset **CA8** and an evaluation toolkit for users to evaluate the quality of their word vectors.\n\n## Reference\nPlease cite the paper, if using these embeddings and CA8 dataset.\n\nShen Li, Zhe Zhao, Renfen Hu, Wensi Li, Tao Liu, Xiaoyong Du, <a href=\"http://aclweb.org/anthology/P18-2023\"><em>Analogical Reasoning on Chinese Morphological and Semantic Relations</em></a>, ACL 2018.\n\n```\n@InProceedings{P18-2023,\n  author =  \"Li, Shen\n    and Zhao, Zhe\n    and Hu, Renfen\n    and Li, Wensi\n    and Liu, Tao\n    and Du, Xiaoyong\",\n  title =   \"Analogical Reasoning on Chinese Morphological and Semantic Relations\",\n  booktitle =   \"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)\",\n  year =  \"2018\",\n  publisher =   \"Association for Computational Linguistics\",\n  pages =   \"138--143\",\n  location =  \"Melbourne, Australia\",\n  url =   \"http://aclweb.org/anthology/P18-2023\"\n}\n```\n\n&nbsp;\n\nA detailed analysis of the relation between the intrinsic and extrinsic evaluations of Chinese word embeddings is shown in the paper:\n\nYuanyuan Qiu, Hongzheng Li, Shen Li, Yingdi Jiang, Renfen Hu, Lijiao Yang. <a href=\"http://www.cips-cl.org/static/anthology/CCL-2018/CCL-18-086.pdf\"><em>Revisiting Correlations between Intrinsic and Extrinsic Evaluations of Word Embeddings</em></a>. Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data. Springer, Cham, 2018. 209-221. (CCL & NLP-NABD 2018 Best Paper)\n\n```\n@incollection{qiu2018revisiting,\n  title={Revisiting Correlations between Intrinsic and Extrinsic Evaluations of Word Embeddings},\n  author={Qiu, Yuanyuan and Li, Hongzheng and Li, Shen and Jiang, Yingdi and Hu, Renfen and Yang, Lijiao},\n  booktitle={Chinese Computational Linguistics and Natural Language Processing Based on Naturally Annotated Big Data},\n  pages={209--221},\n  year={2018},\n  publisher={Springer}\n}\n```\n\n## Format\nThe pre-trained vector files are in text format. Each line contains a word and its vector. Each value is separated by space. The first line records the meta information: the first number indicates the number of words in the file and the second indicates the dimension size. \n\nBesides dense word vectors (trained with SGNS), we also provide sparse vectors (trained with PPMI). They are in the same format with liblinear, where the number before \" : \" denotes dimension index and the number after the \" : \" denotes the value. \n\n## Pre-trained Chinese Word Vectors\n\n### Basic Settings\n\n<table align=\"center\">\n  <tr align=\"center\">\n    <td><b>Window Size</b></td>\n    <td><b>Dynamic Window</b></td>\n    <td><b>Sub-sampling</b></td>\n    <td><b>Low-Frequency Word</b></td>\n    <td><b>Iteration</b></td>\n    <td><b>Negative Sampling<sup>*</sup></b></td>\n  </tr>\n  <tr align=\"center\">\n    <td>5</td>\n    <td>Yes</td>\n    <td>1e-5</td>\n    <td>10</td>\n    <td>5</td>\n    <td>5</td>\n  </tr>\n</table>\n\n<sup>\\*</sup>Only for SGNS.\n\n### Various Domains\n\nChinese Word Vectors trained with different representations, context features, and corpora.\n\n<table align=\"center\">\n    <tr align=\"center\">\n        <td colspan=\"5\"><b>Word2vec / Skip-Gram with Negative Sampling (SGNS)</b></td>\n    </tr>\n    <tr align=\"center\">\n        <td rowspan=\"2\">Corpus</td>\n        <td colspan=\"4\">Context Features</td>\n    </tr>\n    <tr  align=\"center\">\n      <td>Word</td>\n      <td>Word + Ngram</td>\n      <td>Word + Character</td>\n      <td>Word + Character + Ngram</td>\n    </tr>\n    <tr  align=\"center\">\n      <td>Baidu Enc","github_created_at":"2018-01-09T09:48:49+00:00","created_at":"2026-07-11T11:28:18.843432+00:00","updated_at":"2026-08-22T00:01:23.304772+00:00","categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"}],"tags":[{"slug":"chinese","name":"chinese"},{"slug":"embedding","name":"embedding"},{"slug":"word-embeddings","name":"word-embeddings"}],"trust":{"provenance":{"is_fork":false,"github_id":116797311,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-22T00:01:22.549Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":1026,"last_release_at":null,"stars_delta_30d":-3,"open_issues_delta_30d":0},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:28:20.188Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-22T00:01:22.999Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-22T00:01:22.999Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-22T00:01:22.999Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["Use when you specifically require a wide variety (over 100) of pre-trained Chinese word vectors tailored to different aspects of natural language processing within your project.","Choose if you are working in Python and seek compatibility with Apache-2.0 licensed tools, ensuring flexibility for both commercial and personal projects."],"when_not_to_use":["Avoid using this tool if your project requires fine-tuning on a very specific domain that is not well-represented among the existing 100+ models provided.","This collection may not be optimal if you are working strictly within English or other languages, as it focuses primarily on Chinese word vectors."],"source":"enrich:decision_facts","observed_at":"2026-07-11T15:28:51.371Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Chinese-Word-Vectors offers over 100 pre-trained Chinese word vectors for various NLP tasks."}]}}