{"data":{"slug":"tolitius-cupel","name":"cupel","tagline":"discovery tool for evaluating LLM performance","github_url":"https://github.com/tolitius/cupel","owner":"tolitius","repo":"cupel","owner_avatar_url":"https://avatars.githubusercontent.com/u/136575?v=4","primary_language":"Python","stars":64,"forks":0,"topics":["llm","llm-evaluation","local-llm"],"archived":false,"github_pushed_at":"2026-08-31T04:05:52+00:00","maintenance_label":"Active","stars_delta_30d":13,"url":"https://www.graphcanon.com/tools/tolitius-cupel","markdown_url":"https://www.graphcanon.com/tools/tolitius-cupel.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/tolitius-cupel","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=tolitius-cupel","description":"discover LLMs punching above their weight","homepage_url":"https://cupel.run","license":"Apache-2.0","open_issues":2,"watchers":0,"ai_summary":"Cupel is a JavaScript-based toolkit for discovering and evaluating the capabilities of various LLMs through configurable prompt generation, scoring mechanisms, multi-turn dialogues, and local inference server discovery.","readme_excerpt":"## install\n\n```bash\ncurl -fsSL https://cupel.run/install | bash\n```\nor\n```\npip install cupel\n```\nthe UI is bundled in the package\n\n---\n\n## quick start\n\n```bash\ncupel\n```\n\nopens a browser at `localhost:8042`\n\n<img src=\"doc/first-time.png\" alt=\"first time cupel is started\" width=\"76%\">\n\nships with example data (8 models scored by Claude Opus 4.6 on 8 prompts) — the dashboard is populated on first launch\n\n- **LLM-assisted authoring** — describe what you want to test, an LLM drafts the prompt and 0–3 rubric\n- **local + cloud** — oMLX, Ollama, LM Studio, SGLang, OpenRouter, Anthropic, OpenAI\n- **configurable judge** — any model can score responses on a 0–3 rubric with reasoning\n- **thinking model support** — separates `<think>` blocks from answers, only judges the response\n- **multi-turn + tool calling** — multi-step conversations with injected tool results\n- **speed tracking** — tok/s and response times per model\n- **auto-discovery** — probes known ports for local inference servers\n\n---\n\n## license\n\nCopyright © 2026 tolitius\n\nDistributed under the [Apache 2.0](LICENSE) License.","github_created_at":"2026-04-08T07:57:39+00:00","created_at":"2026-07-15T10:40:43.215679+00:00","updated_at":"2026-09-20T04:25:16.926055+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"inference-servers-discovery","name":"inference-servers-discovery"},{"slug":"llm-evaluation","name":"llm-evaluation"},{"slug":"local-llm","name":"local-llm"},{"slug":"multi-turn-dialogue","name":"multi-turn-dialogue"},{"slug":"prompt-generation","name":"prompt-generation"},{"slug":"scoring-mechanisms","name":"scoring-mechanisms"}],"trust":{"provenance":{"is_fork":false,"github_id":1204659735,"owner_type":"User","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-09-10T06:00:20.126Z","maintenance":{"label":"Active","score":82,"methodology":"github_public_v1","releases_90d":0,"days_since_push":10,"last_release_at":null,"stars_delta_30d":13,"open_issues_delta_30d":0},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-15T10:40:44.770Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-09-10T06:00:20.656Z"},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-09-10T06:00:20.656Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-09-10T06:00:20.656Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-09-10T06:00:20.656Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When aiming to evaluate LLMs on local servers due to its auto-discovery feature for known ports of inference servers","For generating initial test scenarios where an LLM can draft prompts and rubrics based on input descriptions"],"when_not_to_use":["If you require a solution that supports a non-JavaScript runtime environment, as Cupel is JavaScript-exclusive","When you need a tool without UI capabilities since Cupel's UI is bundled in the package and may not suit headless operations"],"source":"enrich:decision_facts","observed_at":"2026-07-17T05:24:03.412Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Cupel is a JavaScript-based toolkit for discovering and evaluating the performance of large language models using configurable prompts, scoring mechanisms, multi-turn dialogues, and local inference server discovery."}]}}