BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers

Zhihang Wu1,2, Zhongqi Wang1,2, Jie Zhang1,2, Fengming Gu1,2,3, Shiguang Shan1,2, Xilin Chen1,2

1Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China
2University of Chinese Academy of Sciences, Beijing, China
3School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing, China

Overview of interactive video generation and the BadAction attack

BadAction implants an action-guided trigger into an interactive video generation model. Once activated, the model generates frozen frames and ignores subsequent user actions.

Abstract

Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. We present the first systematic study of backdoor attacks against the interactivity of IVG models. We propose BadAction, which implants predefined motion patterns into poisoned action sequences and associates them with a static target video. Once triggered, the model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. We further explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves an average attack success rate of 91.0% with action-only triggers and 80.4% with multimodal triggers. Defense evaluations further show that BadAction evades existing backdoor detection methods, revealing a critical security gap in interactive video generation.

Method Overview

BadAction framework

BadAction replaces a contiguous action subsequence with a fixed trigger pattern and pairs it with a static target video. The model is fine-tuned with a mixed objective that preserves benign generation while learning the triggered static behavior.

Trigger Modalities

Attack Performance

Trigger Modality ASRSSIM S-T (%) ↑ ASRHuman (%) ↑
Action (BadAction) 91.0 89.6
Image + Action 82.1 83.4
Text + Action 84.7 81.2
Image + Text + Action 80.4 73.2

The action-only trigger achieves the highest success rate. Multimodal triggers provide a stealthier setting but require multiple trigger patterns to co-occur.

Ablation Studies

BadAction ablation studies

We study trigger length, poisoning ratio, loss weight, and training epochs. The best settings are k = 4, θ = 0.375, λ = 3.0, and e = 8.

Defense Evaluation

Detector Precision (%) ↑ Recall (%) ↑ F1 (%) ↑
T2IShield 98.0 50.0 66.2
UFID 78.3 18.0 29.3

The low recall of both detectors shows that existing methods miss many action-based backdoors under practical operating points.

BibTeX

@misc{wu2026badaction,
  title={BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers},
  author={Wu, Zhihang and Wang, Zhongqi and Zhang, Jie and Gu, Fengming and Shan, Shiguang and Chen, Xilin},
  year={2026},
  eprint={2609.39047},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.39047}
}