Confidence_Elicitation_Attacks
Confidence Elicitation Attacks on Large Language Models
GraphCanon updated 2w · GitHub synced 2w
Decision brief
Explores new attack vectors on large language models by eliciting confidence.
Good fit when
- When studying adversarial attacks specifically targeting large language models
- If aiming to assess security vulnerabilities related to model confidence in LLMs
Avoid when
- For general debugging of machine learning models outside of adversarial contexts
- In scenarios focused on improving the performance rather than exposing security flaws
- Hosting:
- unknown - Research paper outlines attack methods for large language models via confidence elicitation.
Observed Jul 17, 2026 · Source: enrich:decision_facts
Verify the decision
Maintenance and security
Full trust report- Maintenance
- Dormant (518d since push)
- As of 2w
- Provenance
- Not a fork · Personal account
- As of 2w
- Security (OSV)
- 123 low (123 low)
- As of 1mo
Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.
Install
pip install Confidence_Elicitation_Attacks PyPISimilar tools
Same-category neighbours. No typed graph edges are catalogued for this tool yet.
Evidence and technical details
Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.
Overview
[ICLR 2025] Research paper exploring new attack methods for large language models through confidence elicitation.
Capability facts
- Languages
- python
Source: github.language · Aug 5, 2026
Categories
Tags
README
Hardware
GPUs: We run our experiments on NVIDIA A40 GPUs with 46 GB of memory. To run the experiments, you'll need enough GPU memory to load the model you’re evaluating.
For agents
This page has a .md twin and JSON over the API.