科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Journal of Optical Communications and Networking2026-02-20· Computer science

AutoONBench: a benchmark for large language model agents in autonomous optical networks

Yihao Zhang, Qizhi Qiu, Jiaping Wu, Xiaomin Liu, Weisheng Hu, Qunbi Zhuge

原始摘要(英文原文)· Original abstract
The integration of large language models (LLMs) into autonomous optical networks (AONs) promises to revolutionize network management. However, the advancement of this field is currently hindered by the lack of a standardized evaluation framework. To bridge this gap, we introduce AutoONBench v1.0, the inaugural version of a comprehensive benchmark designed to assess LLM-based agentic systems within the optical network domain. AutoONBench constructs an evaluation environment incorporating a field-trial dataset, a neural-network-based digital twin (DT), domain-specific operational tools, and documents. It encompasses five task categories covering the complete network lifecycle: service management, network maintenance, failure handling, physical-layer modeling, and network optimization. We further propose a hybrid evaluation methodology that combines quantitative metrics with an LLM-as-a-judge mechanism to provide a multidimensional performance assessment. Extensive evaluations of modern commercial LLMs reveal several problems. While current agents demonstrate proficiency in following standard operating procedures for routine tasks, they exhibit significant limitations in context-aware execution, complex context retrieval, and numerical analysis. AutoONBench establishes a baseline for future research and facilitates the industrial deployment of agentic AON systems.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AutoONBench: a benchmark for large language model agents in autonomous optical networks — 科研速览 Science Skim