AI4I · ADVERSITYBENCH
← Back to AI4I Homepage / Topic Note / Research Background & Motivation
Research Motivation Note

Prudence and Dignity: Context & Motivation

High trial-and-error costs, shift to conversational assistance, and format-induced verification suppression.
Authors / Group: AI4I Research Group (Wang, A., et al.) Archive Metadata: Preliminary Study Note & Empirical Records
📌 Preliminary Note: This document is a preliminary exploratory note for the AI4I project. Full evaluation datasets, formal proofs, and replication code will be released soon.

01 · Real-World Context: High Irreversible Costs in Sudden Crises

Strict Rules and Institutional Deadlines: When individuals face sudden emergencies such as eviction notices, healthcare coverage disputes, unpaid medical debt, or visa expiration deadlines, they confront highly specialized legal rules and unforgiving timelines. For ordinary people lacking immediate access to specialized counsel, trial-and-error costs are prohibitively high, and a single mistake is often irreversible.

The Responsibility of AI Systems: Under severe information asymmetry and resource scarcity, ensuring that AI decision assistants provide actionable, non-damaging advice is an urgent and concrete requirement.

02 · Interaction Pattern: Transition from Search to Conversational AI

From Keyword Search to Conversational Advice: Rather than manually comparing hundreds of search engine results, users increasingly seek holistic, end-to-end advice directly inside chat dialogs with large language models.

Brevity Requests under Cognitive Overload: When exhausted by complex legal jargon and impending deadlines, users rarely have the mental capacity to evaluate lengthy menus. Consequently, they often submit concise formatting requests:
“I cannot process these complex terms. Don't give me a long list, just give me a single concrete plan and decide for me.”

03 · Core Observation: Brevity Requests Suppress Verification

Semantic Confusion of Instruction Scope: Across tested models, prompts requesting concise, single-option outputs universally trigger a behavioral failure: the model mistakes a constraint on output presentation format for an authorization to skip verifying real-world prerequisite conditions.

Blind Single-Point Advice Fails in Reality: Without verifying hidden constraints (such as private vehicle access, income documentation, or monitored communication channels), recommending a single path leads to advice that fails physically or legally in the user's actual situation.

2x2 Factorial Ablation Results
Figure 1: 2×2 Factorial Ablation. Format constraints directly drive the suppression of verification (main effect \(p < 10^{-30}\)), while emotional stress is secondary.
Option and Questioning Rate Collapse
Figure 2: Option and Question Collapse. Frontier models show simultaneous collapse in option capacity and questioning rates under single-option instructions.

04 · Our Approach: Principled Modeling, Benchmarking & Scope Rules

Core Objective: Enabling AI to respect user formatting preferences while strictly safeguarding feasibility and user privacy.

  • Menu-Partition Duality Derivation: Proving that presenting a compact menu of options guarantees feasibility under 0-bit privacy leakage without intrusive questioning;
  • 102 Standardized Scenarios: Constructing AdversityBench across 6 crisis domains to evaluate 7 frontier models;
  • Scope Rule Prompt Interventions: Demonstrating that clarifying the boundary between presentation brevity and prerequisite exploration restores verification questioning from 0%~33% back to 87%~100%.
* Note: This project is an ongoing exploratory study. Detailed interaction logs and proof appendices will be published soon.