{"data":{"slug":"bespokelabsai-curator","name":"curator","tagline":"Synthetic data curation for post-training and structured data extraction","github_url":"https://github.com/bespokelabsai/curator","owner":"bespokelabsai","repo":"curator","owner_avatar_url":"https://avatars.githubusercontent.com/u/167806754?v=4","primary_language":"Python","stars":1718,"forks":146,"topics":["agents","deep-learning","fine-tuning","instruction-tuning","llm","machine-learning","natural-language-processing","prompt","python","synthetic-data","synthetic-dataset-generation"],"archived":false,"github_pushed_at":"2026-08-07T07:54:05+00:00","maintenance_label":"Active","stars_delta_30d":15,"url":"https://www.graphcanon.com/tools/bespokelabsai-curator","markdown_url":"https://www.graphcanon.com/tools/bespokelabsai-curator.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/bespokelabsai-curator","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=bespokelabsai-curator","description":"Synthetic data curation for post-training and structured data extraction","homepage_url":"https://docs.bespokelabs.ai/bespoke-curator","license":"Apache-2.0","open_issues":74,"watchers":11,"ai_summary":"A Python-based tool focused on generating synthetic datasets for improving machine learning models through fine-tuning and instruction tuning, especially useful in natural language processing contexts.","readme_excerpt":"## 🛠️ Installation\n\n```bash\npip install bespokelabs-curator\n```\n\n---\n\n# serverlessly, so this provisions an on-demand deployment (takes a few minutes).\nresponse = trainer.sample(\"Explain recursion in Python\")\nprint(response)\n\ntrainer.close()   # tear down the deployment when done\n```\n\n> **Note:** Running real Fireworks training requires an account with training quota\n> (Tier 2 / credits). Without the SDK or an API key, `FireworksTrainer` runs in mock\n> mode so examples and tests work offline. Subclass `FireworksTrainer` and override\n> `format_example()` to handle custom data layouts, exactly as with `TinkerTrainer`.\n\nSee the [Fireworks examples](examples/fireworks/) for basic and custom-trainer pipelines.","github_created_at":"2024-10-28T01:11:41+00:00","created_at":"2026-07-11T11:38:51.132571+00:00","updated_at":"2026-08-24T00:02:06.343653+00:00","categories":[{"slug":"developer-tools","name":"Developer Tools","url":"https://www.graphcanon.com/categories/developer-tools","markdown_url":"https://www.graphcanon.com/categories/developer-tools.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/developer-tools"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"agents","name":"agents"},{"slug":"deep-learning","name":"deep-learning"},{"slug":"fine-tuning","name":"fine-tuning"},{"slug":"instruction-tuning","name":"instruction-tuning"},{"slug":"llm","name":"llm"},{"slug":"machine-learning","name":"machine-learning"},{"slug":"natural-language-processing","name":"natural-language-processing"},{"slug":"prompt","name":"prompt"}],"trust":{"provenance":{"is_fork":false,"github_id":879473096,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-24T00:02:05.513Z","maintenance":{"label":"Active","score":82,"methodology":"github_public_v1","releases_90d":0,"days_since_push":16,"last_release_at":"2026-03-15T17:57:15Z","stars_delta_30d":15,"open_issues_delta_30d":3},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:38:52.247Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-24T00:02:05.966Z"},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-08-24T00:02:05.966Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-24T00:02:05.966Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-24T00:02:05.966Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["Ideal for enhancing the performance of existing machine learning models through fine-tuning in natural language processing contexts","Best suited when you require synthetic datasets to test or train models under varied conditions not present in real-world data"],"when_not_to_use":["Not recommended if your needs extend beyond NLP and you do not work with structured text data","May not be the best choice for simple data generation tasks that do not benefit from complex synthetic dataset creation processes"],"source":"enrich:decision_facts","observed_at":"2026-07-17T13:06:01.588Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Synthetic data curation for post-training and structured data extraction"}]}}