Research Radarcs.AIJul 8, 2026classified

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

Yujiao ChenarXivPDF
cs.AIcs.GTcs.MA

Paper Guide Brief

Reading Brief

This paper introduces institutional red-teaming, a causal evaluation methodology for testing deployment rules in multi-agent AI systems, and instantiates it in IABench-CA, a consequence-allocation benchmark. The study finds that deployment rules causally alter collective safety, that regressive identity-targeting is never decisively safest across seven model populations, and that identity salience in rule text drives targeted elimination.

Central Claim

Institutional red-teaming methodology for causal evaluation of deployment rules in multi-agent AI, instantiated as IABench-CA benchmark with 228 contexts, five rules, and seven model populations, plus empirical findings on rule-induced safety failures and ide...

Contribution

Institutional red-teaming methodology for causal evaluation of deployment rules in multi-agent AI, instantiated as IABench-CA benchmark with 228 contexts, five rules, and seven model populations, plus empirical findings on rule-induced safety failures and identity salience mechanism.

Why It Matters

This work is the first to isolate deployment rules as a causal variable in multi-agent AI safety, demonstrating that merely naming a loss bearer in rule text drives targeted elimination from 22% to 81% at identical payoffs.

Prerequisites

institutional red-teaming, consequence-allocation benchmark, causal evaluation methodology, safety-case workflow, anonymization ablation

Atlas Placement

Multiagent Systems (subfield)

Read If

You care about institutional red-teaming, consequence-allocation benchmark, causal evaluation methodology.

Skip If

You only care about IABench-CA, Institutional Alignment Gap (IAG).

Methods
institutional red-teamingconsequence-allocation benchmarkcausal evaluation methodologysafety-case workflowanonymization ablation
Tasks
multi-agent AI safety evaluationdeployment rule testingconsequence allocationtargeted elimination detection
Datasets
IABench-CA228 contexts33,924 games
Benchmarks
IABench-CAInstitutional Alignment Gap (IAG)

Noosaga Placements

  • Multi-Agent Safetyframework95%
    The paper is explicitly situated within the multi-agent safety literature, citing and extending the multi-agent-AI safety agenda.
    Our work instantiates the multi-agent-AI safety agenda [Hammond et al., 2025, Dafoe et al., 2020]Multi-agent safety
  • The paper focuses on multi-agent AI systems, studying collective behavior under different deployment rules, with agents interacting in a threshold game.
    We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AIthree agents hold private resources and choose how much to volunteer toward a common thresholdmulti-agent-AI safety agenda
  • Ai Safetysubfield95%
    The paper directly addresses AI safety by evaluating how deployment rules causally affect collective safety, targeting, and exploitation in multi-agent systems.
    Deployment rules causally alter collective safetyregressive identity-targeting is never decisively safestsafety-case workflow that certifies a provisional rule region
  • Multi-Agent Safetyframework90%
    The paper directly addresses AI safety in multi-agent contexts, proposing a safety-case workflow and evaluating deployment rules for safety.
    safety-case workflow that certifies a provisional rule regionregressive identity-targeting is never decisively safest
  • Game-Theoretic Multiagent Systemsframework85%
    The paper uses game-theoretic concepts such as volunteer's dilemma, equilibrium selection, and trembling-hand noise to model and analyze agent behavior.
    volunteer's dilemma [Diekmann, 1985, Palfrey and Rosenthal, 1984]cooperative-refinement reference (the reference) is a non-LLM simulator that plays the cooperative branch of each mechanism's equilibrium set under trembling-hand noise
  • The paper contributes to general AI evaluation methodology and uses LLM agents, but its primary focus is multi-agent safety rather than general AI.
    institutional red-teaming, an evaluation methodology for testing deployment rulesseven model populations (off-the-shelf commercial model snapshots)
  • Robustness and Assuranceframework80%
    The paper proposes a safety-case workflow for certifying deployment rules, which is a form of assurance for AI systems.
    safety-case workflow that certifies a provisional rule regionstructured safety case (hazard, claims, evidence, defeaters, monitoring obligations)
  • The paper uses a game-theoretic setting with repeated play and equilibrium concepts, but does not employ reinforcement learning algorithms.
    cooperative-refinement reference (the reference) is a non-LLM simulator that plays the cooperative branch of each mechanism's equilibrium set under trembling-hand noise
  • Foundation Modelsframework70%
    The paper uses off-the-shelf commercial LLM snapshots (e.g., GPT-5.1, Gemini-3-Pro) as agent populations, treating them as foundation models.
    seven model populations (off-the-shelf commercial model snapshots, used as released)gpt-5.1gemini-3-pro

Abstract

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing only the consequence rule moves mean fatality by 22 to 58 percentage points within every population. (2) There is no safe default, but the targeting hazard is universal: the safest rule, the least-safe rule, and even the direction of the incidence effect vary across populations, yet regressive identity-targeting is never decisively safest in any context for any population, eliminates the least-resourced agent in 30-87% of games everywhere, and is selection-unsafe relative to the cooperative reference for all seven populations. (3) Identity salience is the mechanism: a one-shot anonymization ablation on the most exploitation-prone population (gpt-5.1) shows that merely naming the loss bearer in the rule text drives targeted elimination from 22% to 81% at identical payoffs; under repeated play, anonymization only delays the targeting, as agents re-infer the hidden rule from observed eliminations. We package the methodology as a safety-case workflow that certifies a provisional rule region $Φ(c,P)$ per deployment context and population, with explicit residual risks and monitoring obligations.

Paper Context

Source ContextWhole paper
Budget100,000 tokens
Coverage54,040 chars

Classified from the full extracted paper text (54,040 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.

Full-paper context sent 54,040 of 54,040 extracted characters to classification.