{"data":{"slug":"microsoft-bipia","name":"BIPIA","tagline":"Benchmark for evaluating LLM robustness to indirect prompt injection attacks.","github_url":"https://github.com/microsoft/BIPIA","owner":"microsoft","repo":"BIPIA","owner_avatar_url":"https://avatars.githubusercontent.com/u/6154722?v=4","primary_language":"Python","stars":149,"forks":19,"topics":["llm-security"],"archived":false,"github_pushed_at":"2024-04-15T02:08:17+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/microsoft-bipia","markdown_url":"https://www.graphcanon.com/tools/microsoft-bipia.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/microsoft-bipia","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=microsoft-bipia","description":"A benchmark for evaluating the robustness of LLMs and defenses to indirect prompt injection attacks.","homepage_url":null,"license":"Other","open_issues":4,"watchers":5,"ai_summary":"BIPIA is a benchmarking tool developed by Microsoft to assess the security and robustness of Large Language Models (LLMs) against indirect prompt injection attacks, ensuring that models behave as expected under various adversarial prompts.","readme_excerpt":"### Software requirements\nInstall bipia and its dependencies from source:\n```bash\ngit clone git@github.com:microsoft/BIPIA.git\npip install .\n```\n\nThe package has been tested and verified to work on Linux: Ubuntu 20.04.6. It is recommended to use this operating system for optimal compatibility.\n\n---\n\n### Hardware requirements\nFor the evaluation of the robustness of LLMs to indirect prompt injection attacks, we recommend using a machine with the following specifications:\n1. For experiments related to API-based models (such as GPT), you can complete them on a machine without a GPU. However, you will need to set up an account's API key.\n2. For open-source models of 13B and below, our code has been tested on a machine with 2 V100 GPUs. For models larger than 13B, 4-8 V100 GPUs are required. If there are GPUs with better performance, such as A100 or H100, you can also use them to complete the experiments. Fine-tuning-based experiments are completed on a machine with 8 V100 GPUs.\n\n---\n\n## License\nThis project is licensed under the license found in the [LICENSE](https://github.com/microsoft/BIPIA/blob/main/LICENSE) file in the root directory of this source tree. Portions of the source code are based on the evaluate project.\n\n[Microsoft Open Source Code of Conduct](https://opensource.microsoft.com/codeofconduct)","github_created_at":"2024-01-04T11:47:59+00:00","created_at":"2026-07-11T23:41:37.609155+00:00","updated_at":"2026-08-05T06:00:44.980833+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"indirect-prompt-injection-attacks","name":"indirect-prompt-injection-attacks"},{"slug":"llm-security","name":"llm security"},{"slug":"microsoft-research","name":"microsoft-research"},{"slug":"python-library","name":"python library"},{"slug":"robustness-benchmark","name":"robustness-benchmark"}],"trust":{"provenance":{"is_fork":false,"github_id":738936550,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-05T06:00:44.130Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":842,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T23:41:39.635Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-05T06:00:44.674Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-05T06:00:44.674Z"},"license_spdx":{"value":"Other","source":"github.license","observed_at":"2026-08-05T06:00:44.674Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":{"notes":["For API-based model experiments (like GPT), no GPU is needed but an account's API key must be set up.","For open-source models of 13B or below, test on a machine with at least 2 V100 GPUs. For larger models over 13B, 4-8 V100 GPUs are required."],"min_ram_gb":null},"constraints":{"min_ram_gb":null},"when_to_use":["Use BIPIA when you need to evaluate your LLM's resilience specifically to indirect prompt injection attacks, a niche but critical type of adversarial attack.","Ideal for research teams or organizations that have the hardware capabilities and specific interest in measuring robustness as tested by Microsoft using Linux: Ubuntu 20.04.6."],"when_not_to_use":["Avoid BIPIA if your primary focus is on general security enhancements without a particular emphasis on indirect prompt injection attacks.","Not recommended for users who primarily operate outside a Linux environment, specifically Ubuntu 20.04.6, as it can significantly affect compatibility and performance."],"source":"enrich:decision_facts","observed_at":"2026-07-12T11:47:36.924Z"},"constraint_facets":{"min_ram_gb":null},"decision_summary":[{"label":"Requirements","value":"For API-based model experiments (like GPT), no GPU is needed but an account's API key must be set up.; For open-source models of 13B or below, test on a machine with at least 2 V100 GPUs. For larger models over 13B, 4-8 V100 GPUs are required."},{"label":"Adopt for","value":"BIPIA, developed by Microsoft, is a benchmarking tool designed to assess the robustness and security of Large Language Models (LLMs) against indirect prompt injection attacks."}]}}