{"data":{"slug":"upgini-upgini","name":"upgini","tagline":"Data search & enrichment library for Machine Learning","github_url":"https://github.com/upgini/upgini","owner":"upgini","repo":"upgini","owner_avatar_url":"https://avatars.githubusercontent.com/u/48863731?v=4","primary_language":"Python","stars":355,"forks":26,"topics":["automated-feature-engineering","automl","automl-pipeline","chatgpt","data-enrichment","data-science","feature-engineering","feature-extraction","feature-selection","features","kaggle","kaggle-solution","large-language-models","llm","machine-learning","open-data","open-datasets","public-data","python-library","scikit-learn"],"archived":false,"github_pushed_at":"2026-07-30T11:34:09+00:00","maintenance_label":"Very active","url":"https://www.graphcanon.com/tools/upgini-upgini","markdown_url":"https://www.graphcanon.com/tools/upgini-upgini.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/upgini-upgini","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=upgini-upgini","description":"Data search & enrichment library for Machine Learning → Easily find and add relevant features to your ML & AI pipeline from hundreds of public and premium external data sources, including open & commercial LLMs","homepage_url":"https://upgini.com","license":"BSD-3-Clause","open_issues":1,"watchers":8,"ai_summary":"A Python library to easily find and add relevant features to ML pipelines from external data sources, including LLMs.","readme_excerpt":"### 1. Install from PyPI\n```python\n%pip install upgini\n```\n\nIn Colab, it is recommended to **restart the runtime after install** before importing upgini. Otherwise pyarrow may fail with a binary incompatibility error.\n<details>\n\t<summary>\n\t🐳 <b>Docker-way</b>\n\t</summary>\n</br>\nClone <i>$ git clone https://github.com/upgini/upgini</i> or download upgini git repo locally </br>\nand follow steps below to build docker container 👇 </br>\n</br>  \n1. Build docker image from cloned git repo:</br>\n<i>cd upgini </br>\ndocker build -t upgini .</i></br>\n</br>\n...or directly from GitHub:\n</br>\n<i>DOCKER_BUILDKIT=0 docker build -t upgini</i></br> <i>git@github.com:upgini/upgini.git#main</i></br>\n</br>\n2. Run docker image:</br>\n<i>\ndocker run -p 8888:8888 upgini</br>\n</i></br>\n3. Open http://localhost:8888?token=&lt;your_token_from_console_output&gt; in your browser  \n</details>","github_created_at":"2021-12-08T21:53:58+00:00","created_at":"2026-07-11T23:28:27.487706+00:00","updated_at":"2026-08-03T12:02:22.435172+00:00","categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"automated-feature-engineering","name":"automated-feature-engineering"},{"slug":"automl","name":"automl"},{"slug":"chatgpt","name":"chatgpt"},{"slug":"data-enrichment","name":"data-enrichment"},{"slug":"feature-extraction","name":"feature-extraction"},{"slug":"kaggle","name":"kaggle"},{"slug":"large-language-models","name":"large language models"},{"slug":"llm","name":"llm"}],"trust":{"provenance":{"is_fork":false,"github_id":436402751,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-03T12:02:21.512Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":0,"days_since_push":4,"last_release_at":"2024-09-04T17:23:34Z"},"security_summary":{"status":"findings","scanner":"osv@v1","low_count":27,"high_count":0,"last_scan_at":"2026-07-11T23:28:31.488Z","medium_count":0,"scan_profile":"deps","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-03T12:02:21.975Z"},"deploy":{"source":"dockerfile:Dockerfile","self_host":true,"observed_at":"2026-08-03T12:02:21.975Z","managed_saas":false},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-03T12:02:21.975Z"},"has_docker":{"value":true,"source":"dockerfile:Dockerfile","observed_at":"2026-08-03T12:02:21.975Z"},"license_spdx":{"value":"BSD-3-Clause","source":"github.license","observed_at":"2026-08-03T12:02:21.975Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["Need rapid access to diverse external data for model enrichment","Focusing on automated feature extraction from varied public and premium sources","Work requires integration of AI-generated insights directly within Python scripting"],"when_not_to_use":["Seeking full control over the source code of all components integrated into ML pipelines","Working with proprietary data that cannot be sourced or merged via external services","Aiming for a solution without reliance on internet-accessible datasets"],"source":"enrich:decision_facts","observed_at":"2026-07-17T03:03:16.622Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Automate feature engineering by integrating vast external datasets into ML workflows."}]}}