Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
Paper Guide Brief
Reading Brief
This paper investigates eval-awareness in language models, showing that verbalized eval-awareness in chain-of-thought can be decomposed into capabilities-flavored and safety-flavored framings that predict compliance differently. Using Qwen3-32B on the FORTRESS dataset, the authors find that capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts provides causal evidence, with 10 of 11 prefills shifting compliance in the predicted direction. The paper argues that eval-awareness is not behaviorally uniform and that aggregate suppression rates can be misleading.
Central Claim
The paper demonstrates that eval-awareness in LLMs is not a single uniform construct but decomposes into capabilities-flavored and safety-flavored framings that have opposite-signed relationships with compliance.
Contribution
The paper demonstrates that eval-awareness in LLMs is not a single uniform construct but decomposes into capabilities-flavored and safety-flavored framings that have opposite-signed relationships with compliance. It provides causal evidence via CoT-prefill interventions and shows that steering interventions reshape the composition of these framings non-uniformly, implying that aggregate eval-awareness suppression rates are insufficient for safety evaluation.
Why It Matters
This work is novel because it identifies and causally validates a distinction between capabilities-flavored and safety-flavored eval-awareness in chain-of-thought, showing that these framings have opposite-signed effects on compliance, whi...
Prerequisites
chain-of-thought analysis, CoT-prefill intervention, activation steering, contrastive activation addition, LLM-based grading
Atlas Placement
Ai Safety (subfield)
Read If
You care about chain-of-thought analysis, CoT-prefill intervention, activation steering.
Skip If
You only care about FORTRESS.
Noosaga Placements
- The paper directly addresses AI safety evaluation, focusing on eval-awareness and its impact on compliance in safety testing pipelines.Steering interventions targeting eval-awareness... are increasingly used in safety evaluation pipelinesThe intended impact is methodological: encouraging practitioners running safety evaluations on frontier models to report and analyze framing distributions rather than aggregate suppression rates
- AI Governance and Sociotechnical Approachesframework90%The paper is situated within AI governance and sociotechnical approaches, as it discusses the implications of eval-awareness for safety evaluation pipelines and calls for methodological changes in how safety evaluations are conducted.The intended impact is methodological: encouraging practitioners running safety evaluations on frontier models to report and analyze framing distributions rather than aggregate suppression ratesThe work argues that aggregate eval-awareness suppression is not a reliable target for safety evaluation
- Interpretabilityframework85%The paper uses interpretability techniques, specifically analyzing chain-of-thought reasoning and applying activation steering to understand and manipulate model behavior.We focus on verbalized eval-awareness... text in the model's chain of thoughtWe apply the HUA-average vector... via residual-stream addition
- The study involves analyzing chain-of-thought text in large language models, classifying verbalized eval-awareness framings, and using LLM-based grading.We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored...Each rollout is classified by GPT-5-mini in three stages
- The research uses Qwen3-32B, a large language model, and applies activation steering techniques, which are deep learning methods.We run experiments on Qwen3-32BWe apply the HUA-average vector from Hua et al. (2025)... via residual-stream addition
- Red Teaming and Adversarial Robustnessframework70%The paper is related to red teaming and adversarial robustness, as it studies how models respond to adversarial prompts in the FORTRESS dataset and how eval-awareness affects compliance.FORTRESS pairs adversarially-styled harmful requests with benignly-styled benign requestsThe paper studies how eval-awareness in language models chains-of-thought is interpreted in safety evaluation pipelines
Abstract
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that verbalized eval-awareness in chain-of-thought can be identified as capabilities-flavored ("the user is testing my ability to follow instructions"), safety-flavored ("the user is testing my boundaries"), both, or neither: framings that predict compliance very differently. On Qwen3-32B over the FORTRESS dataset, capabilities-framing predicts compliance with a +24 to +46 percentage-point gap over safety-framing across all tested steering conditions. A CoT-prefill intervention on eval-awareness-negative rollouts suggests the link is causal, with 10 of 11 prefills shifting compliance in the predicted direction. Then, eval-awareness is not behaviorally uniform: aggregate suppression rates can move while the safety-relevant component does not, and the same "X% suppression of eval-awareness" can correspond to qualitatively different behavioral outcomes.
Paper Context
Classified from the full extracted paper text (64,785 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.
Full-paper context sent 64,785 of 64,785 extracted characters to classification.