Research Radarcs.AIAug 27, 2026classified

Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance

Allison Zhuang, Santiago AranguriarXivPDF
cs.AI

Paper Guide Brief

Reading Brief

This paper investigates eval-awareness in language models, showing that verbalized eval-awareness in chain-of-thought can be decomposed into capabilities-flavored and safety-flavored framings that predict compliance differently. Using Qwen3-32B on the FORTRESS dataset, the authors find that capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts provides causal evidence, with 10 of 11 prefills shifting compliance in the predicted direction. The paper argues that eval-awareness is not behaviorally uniform and that aggregate suppression rates can be misleading.

Central Claim

The paper demonstrates that eval-awareness in LLMs is not a single uniform construct but decomposes into capabilities-flavored and safety-flavored framings that have opposite-signed relationships with compliance.

Contribution

The paper demonstrates that eval-awareness in LLMs is not a single uniform construct but decomposes into capabilities-flavored and safety-flavored framings that have opposite-signed relationships with compliance. It provides causal evidence via CoT-prefill interventions and shows that steering interventions reshape the composition of these framings non-uniformly, implying that aggregate eval-awareness suppression rates are insufficient for safety evaluation.

Why It Matters

This work is novel because it identifies and causally validates a distinction between capabilities-flavored and safety-flavored eval-awareness in chain-of-thought, showing that these framings have opposite-signed effects on compliance, whi...

Prerequisites

chain-of-thought analysis, CoT-prefill intervention, activation steering, contrastive activation addition, LLM-based grading

Atlas Placement

Ai Safety (subfield)

Read If

You care about chain-of-thought analysis, CoT-prefill intervention, activation steering.

Skip If

You only care about FORTRESS.

Methods
chain-of-thought analysisCoT-prefill interventionactivation steeringcontrastive activation additionLLM-based gradinghuman validationbootstrap confidence intervalssign test
Tasks
eval-awareness detectioncompliance predictionrefusal classificationframing classificationsafety evaluation
Datasets
FORTRESS
Benchmarks
FORTRESS

Noosaga Placements

  • Ai Safetysubfield95%
    The paper directly addresses AI safety evaluation, focusing on eval-awareness and its impact on compliance in safety testing pipelines.
    Steering interventions targeting eval-awareness... are increasingly used in safety evaluation pipelinesThe intended impact is methodological: encouraging practitioners running safety evaluations on frontier models to report and analyze framing distributions rather than aggregate suppression rates
  • AI Governance and Sociotechnical Approachesframework90%
    The paper is situated within AI governance and sociotechnical approaches, as it discusses the implications of eval-awareness for safety evaluation pipelines and calls for methodological changes in how safety evaluations are conducted.
    The intended impact is methodological: encouraging practitioners running safety evaluations on frontier models to report and analyze framing distributions rather than aggregate suppression ratesThe work argues that aggregate eval-awareness suppression is not a reliable target for safety evaluation
  • Interpretabilityframework85%
    The paper uses interpretability techniques, specifically analyzing chain-of-thought reasoning and applying activation steering to understand and manipulate model behavior.
    We focus on verbalized eval-awareness... text in the model's chain of thoughtWe apply the HUA-average vector... via residual-stream addition
  • The study involves analyzing chain-of-thought text in large language models, classifying verbalized eval-awareness framings, and using LLM-based grading.
    We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored...Each rollout is classified by GPT-5-mini in three stages
  • Deep Learningsubfield70%
    The research uses Qwen3-32B, a large language model, and applies activation steering techniques, which are deep learning methods.
    We run experiments on Qwen3-32BWe apply the HUA-average vector from Hua et al. (2025)... via residual-stream addition
  • Red Teaming and Adversarial Robustnessframework70%
    The paper is related to red teaming and adversarial robustness, as it studies how models respond to adversarial prompts in the FORTRESS dataset and how eval-awareness affects compliance.
    FORTRESS pairs adversarially-styled harmful requests with benignly-styled benign requestsThe paper studies how eval-awareness in language models chains-of-thought is interpreted in safety evaluation pipelines

Abstract

Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes.

Paper Context

Source ContextWhole paper
Budget100,000 tokens
Coverage64,785 chars

Classified from the full extracted paper text (64,785 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.

Full-paper context sent 64,785 of 64,785 extracted characters to classification.