Workshop on Secure and Trustworthy AI
7 September 2026 in Naples, Italy
co-located with the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases

Keynote

Battista Biggio, University of Cagliari

Battista Biggio (MSc 2006, PhD 2010) is a Professor of Computer Engineering at the University of Cagliari, Italy, and research co-director of AI Security at the sAIfer lab (www.saiferlab.ai). He has been attacking machine-learning (ML) models well before adversarial examples were discovered, in the context of cybersecurity applications such as spam filtering, malware detection, web security, and biometric recognition (PRJ 2018). His team was the first to formalize attacks on ML models as optimization problems and to demonstrate gradient-based evasion (ECML-PKDD 2013) and poisoning (ICML 2012) attacks on ML algorithms, playing a leading role in the establishment and advancement of this research field. His seminal paper on “Poisoning Attacks against Support Vector Machines” won the 2022 ICML Test of Time Award. His work on “Wild Patterns” won the 2021 Best Paper Award and Pattern Recognition Medal from Elsevier Pattern Recognition. Prof. Biggio has managed several industrial, national, and EU-funded projects, and regularly serves as Area Chair for top-tier conferences in machine learning and computer security, such as NeurIPS and the IEEE Symposium on Security and Privacy. He is an Associate Editor-in-Chief of Pattern Recognition and chaired IAPR TC1 (2016-2020). He is a Fellow of IEEE, IAPR, and AAIS, a Senior Member of ACM, and a member of ELLIS.

Programme

The following times are in UTC+0.

2:00–2:05 PM Opening and Welcome
2:05–3:05 PM Keynote
Wild Patterns: A Twenty-Year Arms Race between Attacks and Defenses in AI
Battista Biggio, University of Cagliari
Machine learning systems are everywhere — filtering spam, detecting malware, and now writing code and making decisions on our behalf. But these systems can be fooled, sometimes surprisingly easily. An image can be slightly modified to make a classifier see a different object than what is depicted. A few crafted inputs slipped into training data can quietly corrupt a model's behavior. For twenty years, security researchers and attackers have been locked in a back-and-forth: every new defense inspires a cleverer attack, and every clever attack inspires a new defense. In this talk, I'll walk through that history — starting with the earliest attacks on spam and malware detectors, through the discovery of "adversarial examples" that fool image classifiers, and up to the brand-new challenges posed by large language models and AI agents. Along the way, I'll explain the core ideas of adversarial machine learning in plain terms: what makes a model vulnerable, how attackers find its weak points, and why building a truly robust defense has proved so hard. I'll also get a bit critical: why do so many proposed defenses look great on paper but fail in practice? Part of the answer lies in how we evaluate these systems — we often lack rigorous, scalable ways to stress-test models under adversarial or unusual conditions, and we don't have good tools to catch evaluation errors, biased datasets, or models that are right for the wrong reasons. I'll share some of our lab's recent work tackling these problems and try to convince you that the path to trustworthy AI isn't a single, unbreakable model, but rather a carefully engineered system with a defense-in-depth approach.
3:05–4:00 PM Session: Explainability, Trustworthiness and AI Security
Enhancing global explainability for phishing email detection via global additive surrogate models
Authors: Luu Minh Thong Tran, Thien Dang, Xuxian Lu, Viet A. Pham, Howard Kangdrew
Beyond Accuracy: A Trustworthiness Evaluation Methodology for Lab-Generated Endpoint Security Datasets
Authors: Utsav Belbase.
A Cost-Aware Explainability Framework for Closed-Source Large Language Models
Authors: Gabriele Cerizza, Nicola Alboré, Manuela Bazzarelli, Alessandro Castelnovo, Leone De Marco, Luca Puggini.
Unstable Action Boundaries in Lethal Drone Agents
Authors: Jacques Dumora, André-Louis Rochet, Simon Roburin.
4:00–4:30PM Coffee Break
4:30-5:50 Session: Adversarial Robustness, Verification and AI Safety
Adversarially Guided Diffusion for LiDAR Range Image Synthesis
Authors: Stavros Bouras, Antonios Makris, Alexandros Gkillas, Aris Lalos, Konstantinos Tserpes
Contour vs. Interior: A Region-Aware Analysis of Adversarial Vulnerabilities in Object Detection
Authors: Wasaif Alsolami, Raul Santos-Rodriguez, Zahraa S Abdallah, James Pope.
Runtime Plausibility Verification through Semantic Consistency
Authors: C. Sieberichs, Simon Geerkens, Thomas Waschulzik, Ramesh Visvanathan, Alexander Braun.
Exploring Solver-Level Warmstarting for Neural Network Verification
Authors: Annelot Bosman, Minghao Liu, Marta Kwiatkowska, Holger Hoos, Jan van Rijn.
Scale Matters: Purification Failures in Certified Object Detection
Authors: Anant Thunuguntla, Prasad Tadepalli, Giuseppe Raffa.
5:50–6:00PM Closing remarks

Call for Papers

Important Dates

  • Paper submission deadline: June 5th June 15th, 2026 (AoE, UTC-12)
  • Acceptance notification: June 27th June 30th, 2026 (AoE, UTC-12)
  • Camera ready due: To be announced
  • Workshop day: September 7th, 2026

Overview

The increasing adoption of Artificial Intelligence (AI) technologies in critical infrastructure and decision-making processes has made AI-driven components a foundational part of complex software, cyber-physical, and socio-technical systems (e.g., malware detection, fraud detection, autonomous driving, and biometric systems). In these settings, AI outputs directly influence automated and human-in-the-loop decisions, making failures consequential beyond technical performance and raising fundamental concerns regarding the security, robustness, transparency, and trustworthiness of AI systems. While machine learning has demonstrated remarkable performance across a wide range of applications, a growing body of research has shown that AI systems are inherently vulnerable. Adversarial manipulation can compromise not only model predictions but also other critical properties of AI systems, exposing organizations and individuals to significant risks. Vulnerabilities such as adversarial examples, neural backdoors, bias, privacy leakage, and lack of transparency can undermine safety, reliability, and public trust, particularly in security- and safety-critical environments.

This workshop addresses the challenge of holistically securing AI systems beyond an accuracy-centric perspective. It focuses on vulnerabilities and defense strategies across the full AI lifecycle, including adversarial learning, security-critical AI applications, and the role of auxiliary components that support model deployment, interpretation, and human oversight. In particular, the workshop emphasizes the security implications of mechanisms such as explainability, uncertainty estimation, and system-level constraints, considering them as integral parts of AI systems rather than isolated add-ons.

Scope of Papers

STAI welcomes both research papers reporting results from mature work and recently published work, as well as more speculative papers describing new ideas or preliminary exploratory work. Papers reporting industry experiences and case studies will also be encouraged. Submissions are accepted in two formats:

  • Regular research papers with 12 to 16 pages, including references. Research papers must be original, not published previously, and not submitted concurrently elsewhere to be published in the proceedings.
  • Short research statements of at most 6 pages, including references. Research statements aim at fostering discussion and collaboration. They may review previously published research or outline emerging ideas. Papers based on recently published work will not be considered for publication in the proceedings.

Topics of Interest

Topics of interest include but are not limited to:

  • Trustworthy and secure-by-design training and AI pipelines
  • Adversarial machine learning
  • Evasion, poisoning, backdoor, jailbreak, physical-world, and supply-chain attacks
  • Prompt injection and security risks in foundation and generative models
  • Model extraction, inversion, membership inference, and other privacy attacks
  • Attacks on explanations, uncertainty estimation, confidence calibration, and sustainability constraints
  • Robustness and defenses against adversarial and system-level attacks
  • System-level robustness assessment beyond predictive accuracy
  • Trustworthiness evaluation metrics and holistic AI security benchmarks
  • Explainability of machine learning and deep learning models
  • Explainable AI for the explanation of security AI-based systems
  • Explainable AI to improve the accuracy of AI models
  • Explainable AI to improve the robustness of AI models against malicious attacks
  • Attacks on explainability methods and explanation manipulation
  • Privacy and information leakage through interpretability mechanisms
  • Privacy-preserving learning and differential privacy under adversarial settings
  • Human-in-the-loop security and adversarial decision manipulation
  • Security and trustworthiness of agentic and autonomous AI systems
  • Applications of AI to improve security in safety-critical applications (e.g., cybersecurity, fraud detection, biometrics, autonomous systems)
  • Artificial Intelligence for cyber threat detection (e.g., in malware detection, intrusion detection, spam detection)
  • Data-centric security, including poisoning detection, secure data curation, and lifecycle protection

Submission Guidelines

All submissions should be made in PDF using the Microsoft CMT and must adhere to the Springer LNCS style. Templates are available here. Tentatively, all regular workshop papers will be published in an LNCS proceedings volume (to be defined). At a minimum, a proceedings volume will be edited and published online.

Submissions must not substantially overlap with papers that have been published or that are simultaneously submitted to a journal or conference with proceedings. Also, authors should refer to their previous work in the third person. Accepted papers will be published as Springer LNCS proceedings. One author of each accepted paper is required to attend the workshop and present the paper for it to be included in the proceedings.

All accepted submissions must be presented at the workshop. One author of each accepted paper is required to attend the workshop and present the paper for it to be included in the proceedings.

Submission link: https://cmt3.research.microsoft.com/ECMLPKDDWT2026/Track/34/Submission/Create

Committee

Workshop Chairs

Program Committee

  • Alessandro Erba, Karlsruhe Institute of Technology, Germany
  • Andrea Ponte, University of Genova, Italy
  • Angelo Impedovo, NSA Italia s.r.l, Italy
  • Angelo Sotgiu, University of Cagliari, Italy
  • Annalisa Appice, University of Bari, Italy
  • Antonio Pecchia, University of Sannio, Italy
  • Cristian Manca, University of Cagliari, Italy
  • Daniele Ghiani, University of Cagliari, Italy
  • Dario Lazzaro, University of Genova, Italy
  • Donato Malerba, University of Bari, Italy
  • Gianluca Zaza, University of Bari, Italy
  • Gianluca Capozzi, Karlsruhe Institute of Technology, Germany
  • Giulio Rossolini, Scuola Superiore Sant'Anna, Italy
  • Hubert Baniecki, University of Warsaw, Poland
  • Igor Maljkovic, University of Genova, Italy
  • Jonathan Evertz, CISPA Helmholtz Center for Information Security, Germany
  • Lea Schönherr, CISPA Helmholtz Center for Information Security, Germany
  • Lorenzo Cazzaro, University of Luxembourg, Luxembourg
  • Luca Melis, University of Cagliari, Italy
  • Luca Minnei, University of Cagliari, Italy
  • Maria Rosaria Briglia, Sapienza University of Rome, Italy
  • Mohamed Djilani, University of Luxembourg, Luxembourg
  • Srishti Gupta, CISPA Helmholtz Center for Information Security, Germany
  • Thorsten Eisenhofer, CISPA Helmholtz Center for Information Security, Germany
  • Tommaso Zoppi, University of Florence, Italy
  • Vera Rimmer, DistriNet, KU Leuven, Belgium
  • Xinran Zheng, UCL London, England