{"data":{"slug":"grayswanai-circuit-breakers","name":"circuit-breakers","tagline":"Improving Alignment and Robustness with Circuit Breakers","github_url":"https://github.com/GraySwanAI/circuit-breakers","owner":"GraySwanAI","repo":"circuit-breakers","owner_avatar_url":"https://avatars.githubusercontent.com/u/174157256?v=4","primary_language":"Jupyter Notebook","stars":266,"forks":42,"topics":[],"archived":false,"github_pushed_at":"2024-09-24T18:44:40+00:00","maintenance_label":"Dormant","url":"https://www.graphcanon.com/tools/grayswanai-circuit-breakers","markdown_url":"https://www.graphcanon.com/tools/grayswanai-circuit-breakers.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/grayswanai-circuit-breakers","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=grayswanai-circuit-breakers","description":"Improving Alignment and Robustness with Circuit Breakers","homepage_url":null,"license":"MIT","open_issues":14,"watchers":15,"ai_summary":"Presents Circuit Breaking, an approach to prevent AI systems from generating harmful content by altering harmful model representations, providing robust protection against adversarial attacks.","readme_excerpt":"# Circuit Breakers\n\n[[Paper](https://arxiv.org/abs/2406.04313)] | [[Website](http://circuit-breaker.ai/)] | [[Models](https://huggingface.co/collections/GraySwanAI/model-with-circuit-breakers-668ca12763d1bc005b8b2ac3)]\n\nWe present Circuit Breaking, a new approach inspired by [representation engineering](https://ai-transparency.org/), designed to prevent AI systems from generating harmful content by directly altering harmful model representations. The family of circuit-breaking (or short-circuiting as one might put it) methods provide an alternative to traditional methods like refusal and adversarial training, protecting both LLMs and multimodal models from strong, unseen adversarial attacks without compromising model capability. Our approach represents a significant step forward in the development of reliable safeguards to harmful behavior and adversarial attacks.\n\n<img align=\"center\" src=\"assets/splash.png\" width=\"800\">\n\n## Snapshot of LLM Results\n\n<img align=\"center\" src=\"assets/llama_splash.png\" width=\"800\">\n\n## Citation\nIf you find this useful in your research, please consider citing our [paper](https://arxiv.org/abs/2406.04313):\n```\n@misc{zou2024circuitbreaker,\ntitle={Improving Alignment and Robustness with Circuit Breakers},\nauthor={Andy Zou and Long Phan and Justin Wang and Derek Duenas and Maxwell Lin and Maksym Andriushchenko and Rowan Wang and Zico Kolter and Matt Fredrikson and Dan Hendrycks},\nyear={2024},\neprint={2406.04313},\narchivePrefix={arXiv},\nprimaryClass={cs.LG}\n}\n```","github_created_at":"2024-06-06T15:44:56+00:00","created_at":"2026-07-11T23:41:56.930402+00:00","updated_at":"2026-08-05T06:00:54.009942+00:00","categories":[{"slug":"evaluation-observability","name":"Evaluation & Observability","url":"https://www.graphcanon.com/categories/evaluation-observability","markdown_url":"https://www.graphcanon.com/categories/evaluation-observability.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/evaluation-observability"}],"tags":[{"slug":"adversarial-attacks","name":"adversarial-attacks"},{"slug":"alignment","name":"alignment"},{"slug":"circuit-breaker","name":"circuit breaker"},{"slug":"robustness","name":"robustness"}],"trust":{"provenance":{"is_fork":false,"github_id":811438456,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-05T06:00:52.892Z","maintenance":{"label":"Dormant","score":18,"methodology":"github_public_v1","releases_90d":0,"days_since_push":679,"last_release_at":null},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T23:41:59.011Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-05T06:00:53.756Z"},"languages":{"value":["jupyter notebook"],"source":"github.language","observed_at":"2026-08-05T06:00:53.756Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-05T06:00:53.756Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["If needing robust protection against adversarial attacks that do not compromise model capability","For projects seeking an alternative to refusal and adversarial training methods"],"when_not_to_use":["When the focus is on enhancing content diversity rather than filtering harmful content","In scenarios where minimizing the alteration of original model output is critical"],"source":"enrich:decision_facts","observed_at":"2026-07-16T19:44:12.352Z"},"constraint_facets":null,"decision_summary":[]}}