Skip to main navigation Skip to search Skip to main content

Attacking hard-label large vision–language models with model-sensitive adversarial patch designs

  • Nian Ai
  • , Guangke Chen
  • , Xiaowen Cai
  • , Zhongliang Guo*
  • , Daizong Liu*
  • , Pan Zhou
  • , Ognjen Arandjelović
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

Large vision–language models (LVLMs) have demonstrated impressive capabilities in handling multi-modal downstream tasks, gaining increasing popularity. However, recent studies show that LVLMs are susceptible to both intentional and inadvertent attacks. Existing attackers ideally optimize adversarial perturbations with backpropagated gradients from LVLMs, thus limiting their scalability in practical scenarios as real-world LVLM applications will not provide any LVLM’s gradient or details. Motivated by this research gap and counter-practical phenomenon, we propose a novel hard-label attack method for LVLMs, named HardPatch, to generate visual adversarial patches by solely querying the model. Our method provides deeper insights into how to investigate the vulnerability of LVLMs in local visual regions and generate corresponding adversarial substitution under the practical yet challenging hard-label setting. Specifically, we first split each image into uniform patches and mask each of them to individually assess their sensitivity to the LVLM model. Then, according to the descending order of sensitive scores, we iteratively select the most vulnerable patch to initialize noise and estimate gradients with further additive random noises for optimization. In this manner, multiple patches are perturbed until the altered image satisfies the adversarial condition. Extensive LVLM models and datasets are evaluated to demonstrate the adversarial nature of HardPatch. Our empirical observations suggest that with appropriate patch substitution and optimization, HardPatch can craft effective adversarial images to attack hard-label LVLMs.
Original languageEnglish
Article number114261
JournalPattern Recognition
Volume180
Issue numberC
Early online date16 Jun 2026
DOIs
Publication statusE-pub ahead of print - 16 Jun 2026

Fingerprint

Dive into the research topics of 'Attacking hard-label large vision–language models with model-sensitive adversarial patch designs'. Together they form a unique fingerprint.

Cite this