1 IT Security Department, Central Bank of Egypt, Egypt
2 Chief Expert of Cybersecurity, National Telecommunications Regulatory Authority (NTRA), Egypt
3 Arab Academy for Science, Technology and Maritime Transportation, Egypt
*Corresponding author:Mayar Khaled, IT Security Department, Central Bank of Egypt, Egypt
Submission: June 06, 2026:Published: August 27, 2026
ISSN: 2577-2007Volume6 Issue 2
Deep learning-based detection of SQL Injection (SQLi) and Cross-Site Scripting (XSS) vulnerabilities is significantly constrained by the limitations of existing datasets, which commonly separate malicious payloads from the vulnerable source code they exploit. This separation prevents models from learning the semantic and causal relationship between attack patterns and insecure code structures. To address this limitation, this study introduces a systematic six-stage pipeline for constructing xss_sqli_augmented. csv, a paired code-payload dataset containing 13,234 balanced samples enriched with 43 multidimensional features. The primary contribution of this work is the dataset construction methodology and feature engineering framework, rather than the development of a new detection architecture. To enhance coverage of modern attack scenarios, a context-aware hybrid synthesis approach based on GPT- 4o is employed to generate obfuscated samples, with distributional consistency validated using twosample Kolmogorov–Smirnov testing (mean p = 0.41 across 43 features). To ensure robust evaluation, SMOTE-based balancing is performed exclusively within training folds to prevent data leakage. Using an LSTM model as a benchmark classifier, the proposed paired representation achieves an F1-score of 0.989 [95% CI: 0.984–0.994], significantly outperforming payload-only representations (McNemar χ²=14.3, p<0.001). Statistical analysis using the Wilcoxon signed-rank test confirms the superiority of the paired approach across ten evaluated models (W=55, p<0.001). Furthermore, cross-dataset evaluation on an independent OWASP-based corpus demonstrates strong generalization with an F1-score of 0.967 [95% CI: 0.958–0.976]. The proposed dataset, feature extraction framework, and complete pipeline are publicly released to support reproducible research in intelligent web vulnerability detection.
Keywords:Web application Security; Dataset construction methodology; Multi-dimensional feature engineering; LLM-augmented synthesis; SQL injection; Cross-site scripting; Vulnerability detection
a Creative Commons Attribution 4.0 International License. Based on a work at www.crimsonpublishers.com.
Best viewed in