LLM Copilot for Privacy-Respecting Cyber Incident Response
- Authors
-
-
Pramod Prakash
Author
-
- Keywords:
- Incident Response, Large Language Models, Privacy, Cybersecurity, Automation, Benchmark
- Abstract
-
Security operations centers face challenges in managing complex cyber incidents while protecting sensitive forensic data. This paper introduces a multi-agent LLM framework for automating incident response through collaborative planning, execution, analysis, and reflection. We develop a benchmark comprising 130 subtasks across 12 cyber-range scenarios mapped to NIST Incident Response stages (Detection, Response, Recovery), with ten difficulty levels. The system is evaluated using six LLM variants (GPT-4, GPT-4o, Claude-3.5-Sonnet, Llama3-70B, DeepSeek-V3, and GPT-o1). Results demonstrate that the multi-agent architecture achieves sub-task completion rates of 72.3--86.9%, outperforming baselines by 13.9--21.5 percentage points, with privacy leakage reduced from 16.7% to 3.3%. Failure analysis reveals strategic formulation errors account for 44% of failures, followed by command generation (19%) and evidence omission (18%).
- References
- Downloads
- Published
- 2026-08-22
- Issue
- Vol. 1 No. 4 (2026)
- Section
- Articles
- License
-
Copyright (c) 2026 International Journal of Intelligent Systems and Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.
