{"data":{"slug":"prometheus-eval-prometheus-eval","name":"prometheus-eval","tagline":"Evaluate your LLM's response with Prometheus and GPT4","github_url":"https://github.com/prometheus-eval/prometheus-eval","owner":"prometheus-eval","repo":"prometheus-eval","owner_avatar_url":"https://avatars.githubusercontent.com/u/167460660?v=4","primary_language":"Python","stars":1107,"forks":68,"topics":["evaluation","gpt4","litellm","llm","llm-as-a-judge","llm-as-evaluator","llmops","python","vllm"],"archived":false,"github_pushed_at":"2025-04-25T03:58:37+00:00","maintenance_label":"Dormant","stars_delta_30d":5,"url":"https://www.graphcanon.com/tools/prometheus-eval-prometheus-eval","markdown_url":"https://www.graphcanon.com/tools/prometheus-eval-prometheus-eval.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/prometheus-eval-prometheus-eval","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=prometheus-eval-prometheus-eval","description":"Evaluate your LLM's response with Prometheus and GPT4 💯","homepage_url":null,"license":"Apache-2.0","open_issues":13,"watchers":2,"ai_summary":"Prometheus-Eval is a tool that evaluates Language Model responses using Prometheus and GPT-4, providing feedback through local inference or via API.","readme_excerpt":"## 🔧 Installation\n\nInstallation with pip:\n\n```shell\npip install prometheus-eval\n```\n\nPrometheus-Eval supports local inference through `vllm` and inference through LLM APIs with the help of `litellm`.\n\n---\n\n## ⏩ Quick Start\n\n*Note*: `prometheus-eval` library is currently in the beta stage. If you encounter any issues, please let us know by creating an issue on the repository.\n\n\n> **With `prometheus-eval`, evaluating *any* instruction and response pair is as simple as:**\n\n```python","github_created_at":"2024-04-18T17:19:09+00:00","created_at":"2026-07-07T17:43:02.974043+00:00","updated_at":"2026-08-21T00:00:54.330908+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"evaluation","name":"evaluation"},{"slug":"gpt4","name":"gpt4"},{"slug":"litellm","name":"litellm"},{"slug":"llm","name":"llm"},{"slug":"llmops","name":"llmops"},{"slug":"python","name":"python"},{"slug":"vllm","name":"vllm"}],"trust":{"provenance":{"is_fork":false,"github_id":788574052,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-21T00:00:53.579Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":482,"last_release_at":"2024-09-02T04:00:12Z","stars_delta_30d":5,"open_issues_delta_30d":-1},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:22:41.186Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-21T00:00:54.040Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-21T00:00:54.040Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-21T00:00:54.040Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["- When you need detailed and automated evaluations of instruction-response pairs from large language models using both Prometheus metrics and insights from GPT-4.","- For developers who are already utilizing Prometheus within their stack and want to extend observability to LLM performance through the same tooling."],"when_not_to_use":["- If your project does not require Prometheus metrics or if you prefer not to integrate an additional service for evaluation.","- When your organization has strict data policies that prohibit using GPT-4 for assessment purposes, such as in scenarios with sensitive data processing outside AWS."],"source":"enrich:decision_facts","observed_at":"2026-07-11T03:31:47.100Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Prometheus-Eval integrates Prometheus metrics with GPT-4 for evaluating LLM responses in Python, under the Apache-2.0 license."}]}}