arXiv Artificial Intelligence

APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry

APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry

Quick summary

arXiv:2609.38954v1 Announce Type: cross Abstract: Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or sampling. We introduce APTInvestBench, a benchmark for evaluating cross-telemetry robustness in autonomous APT investigation. It comprises 370 cases across seven SOC-inspired conditions, derived from 56 report-informed attack reconstruct

Key takeaways

  • arXiv:2609.38954v1 Announce Type: cross Abstract: Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response.
  • Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or sampling.
  • We introduce APTInvestBench, a benchmark for evaluating cross-telemetry robustness in autonomous APT investigation.

Why it matters

“APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗