METHOD OF CONTEXT-DRIVEN DYNAMIC SOFTWARE SECURITY TESTING USING LARGE LANGUAGE MODEL-BASED AGENTS

Authors

DOI:

https://doi.org/10.31673/2409-7292.2026.035907

Abstract

The advancement of artificial intelligence methods, in particular large language models and the agent-based
systems built upon them, opens new opportunities for improving automated software security testing. Among the existing
methods, dynamic application security testing (DAST) is the most promising direction for such improvement, since it
performs end-to-end verification of a real running application and reflects its actual exploitable behaviour. At the same
time, classical DAST tools construct attack scenarios through stochastic enumeration based on predefined wordlists,
which results in incomplete coverage and considerable resource consumption. The aim of this work is to develop a method
of context-driven dynamic software security testing based on agents built upon large language models, which increases
the completeness of vulnerability detection and reduces the resource cost of scanning. The proposed method uses the
OpenAPI specification as a source of semantic context, on the basis of which agents generate a set of end-to-end (E2E)
security tests that act as a deterministic source of traffic and steer the dynamic scan through real business processes instead
of random crawling. The method provides two improvements: extended coverage of vulnerabilities that are unreachable
for signature-based scanners, and a reduced volume of redundant requests, since only the generated traffic is scanned. A
vulnerability coverage model, the workflow, and the test scenario generation strategy are described, and metrics for evaluating the effectiveness of the method are defined. The results indicate the feasibility of the agent-based approach for
extending coverage and improving the resource efficiency of dynamic testing, and outline the direction of further
experimental research.
Keywords: dynamic application security testing, large language models, agent-based systems, OpenAPI, BOLA,
IDOR, test automation.

References
1. Qadir, S., Waheed, E., Khanum, A., & Jehan, S. (2025). Comparative evaluation of approaches & tools for
effective security testing of Web applications. PeerJ Computer Science, 11, e2821. https://doi.org/10.7717/peerj-cs.2821.
2. Xu, H., Wang, S., Li, N., Wang, K., Zhao, Y., Chen, K., Yu, T., Liu, Y., & Wang, H. (2025). Large language
models for cyber security: A systematic literature review. ACM Transactions on Software Engineering and Methodology.
https://doi.org/10.1145/3769676.
3. Eriksson, B., Pellegrino, G., & Sabelfeld, A. (2021). Black Widow: Blackbox data-driven web scanning. In
2021 IEEE Symposium on Security and Privacy (SP) (pp. 1125–1142). IEEE. https://doi.org/10.1109/ SP40001.
2021.00022.
4. Xue, J., He, Y., Xiao, Y., Chen, Y., Wang, Z., & Lu, H. (2025). From heuristics to intelligence: Evolving blackbox scanners for modern web applications. In 2025 IEEE 10th International Conference on Data Science in Cyberspace
(DSC) (pp. 247–254). IEEE. https://doi.org/10.1109/DSC67331.2025.00039.
5. OWASP Foundation. (2023). API1:2023 Broken object level authorization. OWASP API Security Top 10 –
2023. https://owasp.org/API-Security/editions/2023/en/0xa1-broken-object-level-authorization/.
6. Alazmi, S., & De Leon, D. C. (2022). A systematic literature review on the characteristics and effectiveness of
web application vulnerability scanners. IEEE Access, 10, 33200–33219. https://doi.org/10.1109/ACCESS.2022.3161522.
7. Millar, S., Podgurskii, D., Kuykendall, D., Martínez del Rincón, J., & Miller, P. (2022). Optimising vulnerability
triage in DAST with deep learning. In Proceedings of the 15th ACM Workshop on Artificial Intelligence and Security
(pp. 137–147). ACM. https://doi.org/10.1145/3560830.3563724.
8. Khare, A., Dutta, S., Li, Z., Solko-Breslin, A., Alur, R., & Naik, M. (2025). Understanding the effectiveness of
large language models in detecting security vulnerabilities. In 2025 IEEE Conference on Software Testing, Verification
and Validation (ICST) (pp. 103–114). IEEE. https://doi.org/10.1109/ICST62969.2025.10988968.
9. Gnieciak, D., & Szandała, T. (2025). Large language models versus static code analysis tools: A systematic
benchmark for vulnerability detection. IEEE Access, 13, 198410–198422. https://doi.org/10.1109/ ACCESS.2025.
3635168.
10. Happe, A., & Cito, J. (2023). Getting pwn’d by AI: Penetration testing with large language models. In
Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of
Software Engineering (pp. 2082–2086). ACM. https://doi.org/10.1145/3611643.3613083.
11. Deng, G., Liu, Y., Mayoral-Vilches, V., Liu, P., Li, Y., Xu, Y., Zhang, T., Liu, Y., Pinzger, M., & Rass, S.
(2024). PentestGPT: Evaluating and harnessing large language models for automated penetration testing. In Proceedings
of the 33rd USENIX Security Symposium (pp. 847–864). USENIX Association. https://www.usenix.org/conference/
usenixsecurity24/presentation/deng.
12. Shen, X., Wang, L., Li, Z., Chen, Y., Zhao, W., Sun, D., Wang, J., & Ruan, W. (2025). PentestAgent:
Incorporating LLM agents to automated penetration testing. In Proceedings of the 20th ACM Asia Conference on
Computer and Communications Security (pp. 375–391). ACM. https://doi.org/10.1145/3708821.3733882.
13. Barabanov, A., Dergunov, D., Makrushin, D., & Teplov, A. (2022). Automatic detection of access control
vulnerabilities via API specification processing. arXiv. https://doi.org/10.48550/arXiv.2201.10833.
14. Santos Filho, A., Rodríguez, R. J., & Feitosa, E. L. (2025). Automated broken object-level authorization attack
detection in REST APIs through OpenAPI to colored Petri nets transformation. International Journal of Information
Security, 24(2), 83. https://doi.org/10.1007/s10207-024-00970-5.
15. Huang, Y., Shi, C., Lu, J., Li, H., Meng, H., & Li, L. (2024). Detecting broken object-level authorization
vulnerabilities in database-backed applications. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and
Communications Security (pp. 2934–2948). ACM. https://doi.org/10.1145/3658644.3690227.
16. Pasca, E. M., Delinschi, D., Erdei, R., & Matei, O. (2025). LLM-driven, self-improving framework for security
test automation: Leveraging Karate DSL for augmented API resilience. IEEE Access, 13, 56861–56886.
https://doi.org/10.1109/ACCESS.2025.3554960.
17. Pasca, E. M., Erdei, R., Delinschi, D., & Matei, O. (2024). Enhancing API security testing against BOLA and
authentication vulnerabilities through an LLM-enhanced framework. In 19th International Conference on Soft Computing
Models in Industrial and Environmental Applications (SOCO 2024) (Lecture Notes in Networks and Systems, pp. 231–
240). Springer. https://doi.org/10.1007/978-3-031-75010-6_23.
18. OpenAPI Initiative. (2024). OpenAPI Specification v3.1.1. https://spec.openapis.org/oas/v3.1.1.html.
19. ZAP (Zed Attack Proxy) [Програмне забезпечення]. (2025). Checkmarx. https://www.zaproxy.org/.

Published

2026-09-15

Issue

Section

Articles