Author : Ritaben M. Marwada, Nilesh Modi
Date of Publication : August 2026
Abstract: The increasing sophistication of AI-generated phishing attacks has rendered traditional signature-based web security defenses inadequate, as 40% of these attacks now target businesses and achieve a 60% victim engagement rate (Generative NLP Models for Mitigating Phishing Attacks, n.d.). This paper presents a transparent, multi-layered defense framework that integrates a fine-tuned RoBERTa transformer with the SHAP interpretability engine to provide both high-accuracy phishing classification and forensic-grade explanations for each prediction. The proposed methodology leverages adversarial training using LLM-generated synthetic samples to bolster resilience against evolving, unknown threats (Figueiredo et al., 2024; Generative NLP Models for Mitigating Phishing Attacks, n.d.; Kulal et al., 2025). Experimental validation demonstrates that the proposed hybrid framework achieves 98.4% accuracy on the PhreshPhish benchmark while reducing false positives by 12.7% compared to standard transformer baselines, all while maintaining the transparency required for GDPR-compliant privacy auditing (Cadet et al., 2024; Lim et al., 2025; Shendkar et al., 2024).
Reference :