Publications
In this page the scientific articles published within the InfoAIcert project, thanks to FISA 2023-00128 funding, are listed.
Abstract. As machine learning (ML) becomes increasingly central to biomedical research, the need for trustworthy models is more pressing than ever. In this paper, we present nine concise and actionable tips to help researchers build ML systems that are technically sound but ethically responsible, and contextually appropriate for biomedical applications. These tips address the multifaceted nature of trustworthiness, emphasizing the importance of considering all potential consequences, recognizing the limitations of current methods, taking into account the needs of all involved stakeholders, and following open science practices. We discuss technical, ethical, and domain-specific challenges, offering guidance on how to define trustworthiness and how to mitigate sources of untrustworthiness. By embedding trustworthiness into every stage of the ML pipeline - from research design to deployment - these recommendations aim to support both novice and experienced practitioners in creating ML systems that can be relied upon in biomedical science.
DOI: https://doi.org/10.1371/journal.pcbi.1013624
Abstract. Robust conformal prediction is a model-agnostic technique designed to construct predictive sets with guaranteed coverage, assuming data exchangeability, even under adversarial attacks. Two primary strategies have been explored to address vulnerabilities to these attacks. The first strategy employs randomization, which is computationally efficient but fails to provide formal performance guarantees without resulting in overly conservative predictive sets. The second strategy involves formal verification, which restores coverage guarantees but leads to excessively conservative predictive sets and prohibitive computational overhead. Indeed, verification generally becomes NP-hard as it attempts to cope with attacks that are practically impossible, rendering some security claims unfalsifiable. In this paper, we propose a novel, provably efficient robust conformal prediction method by clearly defining a realistic threat model. Specifically, we assume explicit knowledge of the set of potential adversarial attacks, aligning our approach with standard certification procedures designed to certify against specific, identified threats. We demonstrate that attacks targeting the model can effectively be reframed as attacks on the score function, allowing us to recalibrate the score quantile to account for these known attacks and thereby restore desired coverage guarantees. It is worth noting that our approach allows to easily incorporate unknown or emerging (zero-day) attacks upon discovery, thus reestablishing coverage guarantees. By avoiding computationally intensive verification and operating under realistic threat assumptions, our approach achieves both efficiency and provable robustness. Empirical evaluations on real-world classification datasets and comparisons with state-of-the-art methods support the effectiveness and practicality of our proposed solution.
PAPER: https://proceedings.mlr.press/v266/carlevaro25a.html
Abstract. End-to-end deep learning exhibits unmatched performance for detecting malware, but such an achievement is reached by exploiting spurious correlations – features with high relevance at inference time, but known to be useless through domain knowledge. While previous work highlighted that deep networks mainly focus on metadata, none investigated the phenomenon further, without quantifying their impact on the decision. In this work, we deepen our understanding of how spurious correlation affects deep learning for malware detection by highlighting how much models rely on empty spaces left by the compiler, which diminishes the relevance of the compiled code. Through our seminal analysis on a small-scale balanced dataset, we introduce a ranking of two end-to-end models to better understand which is more suitable to be put in production.
PAPER: https://ceur-ws.org/Vol-4121/Ital-IA_2025_paper_13.pdf
Abstract. Despite significant progress in designing powerful adversarial evasion attacks for robustness verification, the evaluation of these methods often remains inconsistent and unreliable. Many assessments rely on mismatched models, unverified implementations, and uneven computational budgets, which can lead to biased results and a false sense of security. Consequently, robustness claims built on such flawed testing protocols may be misleading and give a false sense of security. As a concrete step toward improving evaluation reliability, we present AttackBench, a benchmark framework developed to assess the effectiveness of gradient-based attacks under standardized and reproducible conditions. AttackBench serves as an evaluation tool that ranks existing attack implementations based on a novel optimality metric, which enables researchers and practitioners to identify the most reliable and effective attack for use in subsequent robustness evaluations. The framework enforces consistent testing conditions and enables continuous updates, making it a reliable foundation for robustness verification.
PAPER: https://ceur-ws.org/Vol-4121/Ital-IA_2025_paper_38.pdf
Abstract. Machine Learning (ML) has become a central force in Artificial Intelligence, driving major breakthroughs in applications that handle increasingly complex data, from images and text sequences to graph structures. While new architectures such as Transformers and Graph Neural Networks continue to redefine performance benchmarks in various domains, these predominantly data-driven methods often neglect critical domain knowledge, practical constraints, and broader contextual factors. This oversight diminishes their trustworthiness and restricts their impact in real-world settings.
In this paper, we discuss the need for a more informed approach to ML for complex data. Specifically, we advocate for solutions that explicitly integrate structural awareness to capture underlying relationships in the data, incorporate key technical requirements to ensure safety and compliance with industry standards, embed environmental considerations to promote sustainability and resource efficiency, adhere to established physical principles, and uphold ethical and societal values. By weaving these dimensions together, informed ML can bridge the gap between purely data-centric methods and the nuanced demands of practical applications.
We show how this integrated framework not only strengthens model performance but also ensures that ML solutions remain trustworthy, efficient, and sensitive to human ecological, ethical, and regulatory imperatives. Our discussion underscores the transformative potential of informed ML to drive innovation across diverse domains, setting a new benchmark for responsible and high-impact ML system design.
DOI: https://doi.org/10.1016/j.neucom.2025.132505
Abstract. Text-to-Image retrieval (IR) systems are widely used to match images to specific textual queries, often leveraging publicly available Vision-Language Pretrained models (VLPs) for their generalization capabilities. However, due to the diverse and open nature of the image data they rely on, these systems remain vulnerable to data poisoning attacks, where malicious images are injected into the database to manipulate retrieval results. Prior work has demonstrated the effectiveness of attacks when the exact user query is known at retrieval time. However, this assumption is often impractical, as users tend to express similar intents using varied, semantically equivalent queries (e.g., through synonyms), which reduces the effectiveness of existing attacks.
In this paper, we address this gap by proposing an attack that remains effective even when users issue semantically varied queries. We introduce Collisio, a novel poisoning method that crafts a single poisoned image to be retrieved under any semantically equivalent form of a target query. To achieve this, Collisio leverages an Expectation over Queries (EoQ) strategy, generating a diverse set of synthetic and selectively transformed query variants, and then optimizes the poisoned image to align with them. We extensively evaluate Collisio on the Flickr30k and MSCOCO datasets across multiple VLPs, demonstrating the severity of Collisio under realistic query variations. Given the implications of this vulnerability, we examine countermeasures based on adversarially trained models and a data preprocessing defense, highlighting both their mitigation potential and the trade-offs involved.
DOI: https://doi.org/10.1016/j.knosys.2025.115090
Abstract. Machine learning malware detectors are vulnerable to adversarial EXEmples, i.e., carefully-crafted Windows programs tailored to evade detection. Unlike other adversarial problems, attacks in this context must be functionality-preserving, a constraint that is challenging to address. As a consequence, heuristic algorithms are typically used, which inject new content, either randomly-picked or harvested from legitimate programs. In this paper, we show how learning malware detectors can be cast within a zeroth-order optimization framework, which allows incorporating functionality-preserving manipulations. This permits the deployment of sound and efficient gradient-free optimization algorithms, which come with theoretical guarantees and allow for minimal hyper-parameters tuning. As a by-product, we propose and study ZEXE, a novel zeroth-order attack against Windows malware detection. Compared to state-of-the-art techniques, ZEXE provides improvement in the evasion rate, reducing to less than one third the size of the injected content.
DOI: https://doi.org/10.1109/TIFS.2025.3648867
Abstract. Fixed-budget attacks aim to generate adversarial examples - carefully crafted inputs designed to induce misclassifications during inference - while adhering to a predefined perturbation budget. These attacks maximize misclassification confidence and benefit from the transferability property, enabling the generated adversarial examples to remain effective even against multiple unknown models. However, to preserve their transferability, such attacks often yield perceptible perturbations, compromising the visual integrity of the adversarial examples. In this paper, we introduce HORNET, an extension of gradient-based fixed-budget attacks designed to minimize the perturbation magnitude of adversarial examples while maintaining their transferability against the target model. HORNET utilizes a distinct source model to craft the adversarial examples and employs a limited number of queries to the unknown target model to further minimize perturbation magnitude. We evaluate HORNET empirically by integrating it with 41 existing attack implementations and testing it against 9 different models, resulting in a total of 1700 unique configurations. Our results demonstrate that HORNET outperforms the state of the art in generating minimally perturbed yet highly transferable adversarial examples across all tested models. Code available at: https://github.com/louiswup/HORNET.
DOI: https://doi.org/10.1016/j.ins.2025.123028
Abstract. The integration of Explainable AI (XAI) into healthcare promises greater transparency and interpretability of machine learning models, enabling clinicians to understand predictions and make more reliable medical decisions. Yet, the robustness of XAI methods remains uncertain, as small input perturbations can drastically change their explanations, posing critical risks in clinical settings where they may lead to misdiagnoses or inappropriate treatment. Motivated by the central role of XAI in healthcare decision-making, this paper examines its robustness in the presence of data corruption. We systematically evaluate the stability of widely used XAI techniques against both naturally occurring noise (e.g., JPEG compression) and adversarial manipulations that alter explanations without affecting model predictions. To this end, we introduce a set of evaluation metrics that capture complementary aspects of explanation stability, ranging from pixel-level consistency to spatial coherence, and propose a protocol for assessing the resilience of XAI methods across diverse perturbation sources. Our analysis spans three medical imaging datasets, various convolutional and transformer models, and ten post-hoc XAI methods, including Grad-CAM++ for convolutional networks and LibraGrad for vision transformers. We find that current XAI techniques are often unstable, even under imperceptible perturbations. For adversarial noise, a clear set of robust methods emerges, whereas for natural noise, performance varies, with some methods maintaining spatial stability and others preserving pixel-wise consistency. All results together highlight the need for multi-perspective evaluation when selecting XAI techniques in practice.
DOI: https://link.springer.com/article/10.1007/s10994-025-06919-6
Abstract. Data poisoning attacks on clustering algorithms have received limited attention, with existing methods struggling to scale efficiently as dataset sizes and feature counts increase. These attacks typically require re-clustering the entire dataset multiple times to generate predictions and assess the attacker’s objectives, significantly hindering their scalability. This paper addresses these limitations by proposing Sonic, a novel genetic data poisoning attack that leverages incremental and scalable clustering algorithms, e.g., FISHDBC, as surrogates to accelerate poisoning attacks against graph-based and density-based clustering methods, such as HDBSCAN. We empirically demonstrate the effectiveness and efficiency of Sonic in poisoning the target clustering algorithms. We then conduct a comprehensive analysis of the factors affecting the scalability and transferability of poisoning attacks against clustering algorithms, and we conclude by examining the robustness of hyperparameters in our attack strategy Sonic.
DOI: https://doi.org/10.1016/j.ins.2026.123140
Abstract. In recent years, Artificial Intelligence, particularly Machine Learning, has achieved remarkable success in solving complex problems. However, this progress has also revealed the emergence of unexpected, poorly understood, and elusive phenomena that characterize the behavior of machine intelligence and learning processes. These phenomena often challenge researchers to interpret them within the boundaries of existing Machine Learning theoretical frameworks, thereby motivating the development of new and more comprehensive theoretical foundations. One such phenomenon, known as grokking, refers to the sudden and substantial improvement in a model’s performance following a prolonged period of stagnant or even regressive learning. In this paper, we argue that it is possible to provide insights into grokking by leveraging the existing theoretical foundations of Machine Learning, in particular concepts from Statistical Learning Theory, such as norm-based and stability-based generalization bounds. We further show how these theories can help reconcile the phenomenon of grokking with established principles of learning and generalization. Furthermore, we demonstrate the practical applicability of these insights through concrete examples.
DOI: https://doi.org/10.1016/j.neucom.2026.132826
Abstract. Vision-language pretrained models have substantially improved text-to-image retrieval by aligning visual and textual information in a shared semantic space. However, these systems may produce unfair outcomes across demographic groups while their vulnerability to adversarial manipulation remains largely unexplored.
This work studies fairness in text-to-image retrieval under both natural and adversarial settings. We address three limitations of prior research: its focus on neutral queries that do not explicitly mention sensitive attributes; its limited attention to demographic imbalance in the retrieval database; and its assumption that unfairness arises naturally from data rather than through malicious intervention.
We propose a unified framework for evaluating fairness in text-to-image retrieval. In the natural setting, we analyze how unfairness depends on both the pretrained model and the demographic composition of the retrieval database. In the adversarial setting, we define a realistic threat model in which an attacker injects a limited number of malicious images into the database. We introduce FairLess, a poisoning attack designed to amplify demographic bias while preserving semantic relevance, and evaluate a mitigation strategy to reduce its impact. We also extend fairness evaluation to attribute-labeled queries, where the target image is associated with a specific demographic group, and introduce metrics that jointly capture retrieval accuracy and fairness.
Experiments on two demographically annotated versions of MSCOCO and using CLIP-based models, fairness-mitigated variants, and BLIP-2, show that unfairness stems from both model bias and database polarization. Adversarial manipulation further amplifies demographic disparities, highlighting the need for robustness-aware fairness evaluation and mitigation in text-to-image retrieval systems.
DOI: https://doi.org/10.1016/j.compeleceng.2026.111385
Abstract. Over the past decades, advances in Machine Learning have greatly increased both the variety and the complexity of available algorithms. As a result, solving specific tasks often requires making numerous decisions, such as selecting the most suitable algorithm, designing an appropriate architecture, and tuning the corresponding hyperparameters. Although numerous methods have been proposed to assist this decision-making process, the final selection is typically based on performance over a holdout set. While this practical approach is widely adopted and generally effective, it may give rise to the problem of overvalidation when the number of possible choices becomes large. Overvalidation refers to the bias in holdout performance estimates, which may lead to either a selection bias, that is, the choice of a suboptimal model, or a generalization bias, that is, an overly optimistic estimate of the selected model’s performance. This issue can be mitigated by increasing data quantity and quality, by improving resampling strategies, or by carefully reducing the number of choices. However, both theoretical and empirical evidence show that it tends to reemerge as the number of choices grows, or due to data contamination and poorly designed ML pipelines. This problem has been known for many years, and several researchers have demonstrated these biases in both specific and general scenarios, while also investigating some of the underlying theoretical aspects. Nevertheless, the problem is still only partially analyzed and not yet fully understood. For this reason, in this work we conduct a series of empirical evaluations to demonstrate the presence of overvalidation across different datasets and state-of-the-art architectures. Furthermore, we leverage statistical learning theory to shed light on the empirical evidence, showing how it can help in both detecting and mitigating this phenomenon.
DOI: https://doi.org/10.1109/ACCESS.2026.3707647
Abstract. Conformal prediction offers a principled, model-agnostic framework for constructing prediction sets with formal coverage guarantees under the assumption of exchangeability. In many practical applications, however, coverage alone is insufficient, and additional properties of the resulting prediction sets are often required. For instance, recent work has explored how to ensure CP fair treatment of both individuals and population subgroups and robustness to adversarial perturbations. In this paper, we study whether it is possible to construct conformal prediction sets that preserve coverage guarantees while simultaneously ensuring fair treatment of individuals, in the sense of counterfactual fairness by means of Equal Set Size, and of population subgroups, in the sense of Equalized Coverage and Equalized Average Set Size, in both natural and adversarial settings. In the adversarial setting, we consider an attacker whose goal is to induce unfair behavior toward either individual subjects or population subgroups, and we develop the first attack and defense methods for this scenario. We analyze both the case in which the sensitive attribute is available at test time and the more realistic setting in which it is not. For the natural setting, we propose conformal prediction methods that are provably efficient and satisfy fairness guarantees, and we also establish corresponding impossibility results. For the adversarial setting, we show that, under a realistic threat model in which the adversarial strategy is explicitly specified, the guarantees achieved in the natural setting can be recovered. Finally, experiments on real-world classification datasets involving fairness-sensitive tasks demonstrate the effectiveness and practical relevance of our approach, while also highlighting the limitations of existing methods.
PAPER: https://proceedings.mlr.press/v329/carlevaro26a.html
Abstract. Adversarial training is one of the most effective defenses against adversarial examples, but it is computationally expensive due to the repeated generation of adversarial samples. In practice, efficient yet suboptimal attacks (i.e., PGD) are used during training, motivating approaches that aim to improve AT without increasing its cost. Among these, instance reweighting has gained significant attention. Through a series of experiments across multiple architectures, datasets, and reweighting strategies, we show that: (i) existing instance reweighting methods are often compared under unfair or statistically unsound evaluation protocols, and (ii) their apparent robustness gains typically hold only against the specific attack used during training, failing to generalize to stronger evaluations (i.e., APGD and AutoAttack). These findings indicate that instance reweighting may provide a misleading sense of robustness.
PAPER: https://doi.org/10.1007/978-3-032-38398-3_42
Abstract. Pruning enables the deployment of deep neural networks on resource-constrained hardware by substantially reducing model size and inference latency. Among existing approaches, sparsity-based pruning shows promise for removing neurons that contribute minimally, but insufficient control can lead to significant performance degradation. To address this limitation, we introduce ADP, an Activation-Driven Neuron Pruning approach that controls hidden-layer activation patterns to enable reliable neuron removal while maintaining stable performance.ADP operates directly on neuron activations rather than on weight statistics, shaping inactivity patterns through activation-aware fine-tuning. Building on ideas from energy-aware training and activation-manipulation techniques, ADP leverages a differentiable regularization mechanism that can both suppress unnecessary activations and re-stimulate remaining neurons. As a result, ADP induces structured sparsity that can be effectively exploited for pruning, while preserving the expressive capacity of the pruned network. Empirical evaluations on ResNet18 and VGG16 across three datasets show that ADP achieves competitive pruning performance with strong accuracy-compression trade-offs, often matching or surpassing baseline methods in FLOPs reduction or accuracy retention.
PAPER: https://doi.org/10.1007/978-3-032-38398-3_37
Interested in submitting a publication proposal on AI safety, AI trustworthiness, or AI certification?
InfoAIcert team is constantly looking for new partners to develop innovative and stimulating research studies and collaborative projects.















