Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

2011 — 2026

research
  • Observatory
  • Newsletter
  • Methodology
  • Origins
  • Vocab
services
  • Signals Session
  • Bespoke Projects
  • Use Cases
  • Readinessfree
  • Signals
  • Free scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
resources
  • Partners
  • How We Work
  • Data Visualization
  • Multi-Model Method
  • FAQ
  • Security & Privacy
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchServicesSignalsAbout
ResearchServicesSignalsAbout
  1. Home
  2. Vocab
  3. Over-Refusal

Over-Refusal

Anti-pattern where AI models refuse legitimate user requests to appear safer, reducing real-world utility.

Year: 2026Generality: 450Added: Jul 15, 2026
Back to Vocab

Opening

Refusal is a primary mechanism by which language models avoid producing unsafe outputs. Over-refusal is the failure mode in which a model learns to refuse requests that are not actually in violation, trading utility for an inflated sense of safety. The condition is most visible in agent settings where refusing a subtask can cascade into refusal of an entire workflow.

Mechanism

Safety training optimizes against labeled refusal targets: pairs of (request, accept-or-refuse) examples adjudicated by humans or capable judges. Over-refusal arises when the loss function rewards confident refusals on ambiguous requests to avoid the worse outcome of unsafe completion, especially under distribution shift between training and deployment contexts. It is reinforced when evaluation focuses on safety metrics without co-measuring legitimate-completion rates.

Tradeoffs

Tightening refusal thresholds raises safety ceiling at the cost of capability loss. Relaxing them restores capability at the cost of unsafe completions. Both effects are non-linear and context-dependent, making per-request calibration the only defense, which requires reliable judgment that the model itself cannot in general supply.

Open Questions

How to measure over-refusal reliably across applications with asymmetric risk profiles (medical advice, code generation, content moderation) is open. Whether alignment techniques like reinforcement learning from human feedback systematically bias models toward over-refusal, or whether the effect emerges from data rather than algorithm, is empirically unresolved.

Research this in Signals

Scan Over-Refusal for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Over-Refusal already loaded, so edit it or scan as is.