Students are increasingly using Large Language Models (LLMs) to answer multiple-choice questions (MCQs) without engaging with the underlying concepts, which undermines the validity of traditional assessments. Instead of trying to detect or restrict LLM usage, an alternative strategy is redesigning MCQs so that relying on LLMs becomes ineffective. In this paper, we propose GradQuiz, a framework that generates adversarial distractors targeting the decision-making process of LLMs. GradQuiz leverages gradient-based signals from a target LLM to perturb semantically influential entities in an MCQ. This process generates plausible distractors, which are then refined to ensure grammatical and semantic coherence. We evaluated GradQuiz on two MCQ benchmarks, i.e., OpenTriviaQA and Massive Multitask Language Understanding (MMLU), across multiple LLM families, including both open-weight and proprietary black-box models. Results show that GradQuiz consistently reduces LLM accuracy in answering MCQs. In particular, when applied to gemma-3-27b-it, it reduces accuracy by 67.26% compared to the original quizzes and by 51.07% compared to the strongest competing adversarial distractor approach on the MMLU dataset, while it achieves accuracy reductions of 49.31% and 42.92%, respectively, on the OpenTriviaQA dataset. Human evaluations confirm that the generated distractors preserve pedagogical coherence and relevance, while a controlled study on students shows that GradQuiz does not increase MCQ difficulty for learners who do not rely on LLMs to answer quizzes.

Leveraging Adversarial Attacks to Generate Multiple-Choice Quizzes Robust Against Large Language Models / Bonifazi, G., Buratti, C., Marchetti, M., Parlapiano, F., Traini, D., Ursino, D., Virgili, L.. - In: ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY. - ISSN 2157-6912. - (2026). [Epub ahead of print] [10.1145/3803797]

Leveraging Adversarial Attacks to Generate Multiple-Choice Quizzes Robust Against Large Language Models

G. Bonifazi
Primo
;
C. Buratti;M. Marchetti;F. Parlapiano;D. Traini;D. Ursino;L. Virgili
Ultimo
2026-01-01

Abstract

Students are increasingly using Large Language Models (LLMs) to answer multiple-choice questions (MCQs) without engaging with the underlying concepts, which undermines the validity of traditional assessments. Instead of trying to detect or restrict LLM usage, an alternative strategy is redesigning MCQs so that relying on LLMs becomes ineffective. In this paper, we propose GradQuiz, a framework that generates adversarial distractors targeting the decision-making process of LLMs. GradQuiz leverages gradient-based signals from a target LLM to perturb semantically influential entities in an MCQ. This process generates plausible distractors, which are then refined to ensure grammatical and semantic coherence. We evaluated GradQuiz on two MCQ benchmarks, i.e., OpenTriviaQA and Massive Multitask Language Understanding (MMLU), across multiple LLM families, including both open-weight and proprietary black-box models. Results show that GradQuiz consistently reduces LLM accuracy in answering MCQs. In particular, when applied to gemma-3-27b-it, it reduces accuracy by 67.26% compared to the original quizzes and by 51.07% compared to the strongest competing adversarial distractor approach on the MMLU dataset, while it achieves accuracy reductions of 49.31% and 42.92%, respectively, on the OpenTriviaQA dataset. Human evaluations confirm that the generated distractors preserve pedagogical coherence and relevance, while a controlled study on students shows that GradQuiz does not increase MCQ difficulty for learners who do not rely on LLMs to answer quizzes.
2026
Computing methodologies, Natural language processing, Machine learning, Applied computing, Large Language Models, Educational Assessment, Gradient-Based Perturbation, Education, Multiple-Choice Quizzes, Distractor Generation
File in questo prodotto:
File Dimensione Formato  
Bonifazi_Leveraging-Adversarial-Attacks-Generate_2026.pdf

Solo gestori archivio

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza d'uso: Tutti i diritti riservati
Dimensione 489.54 kB
Formato Adobe PDF
489.54 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11566/354212
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact