Ahmed Twabi, Yepeng Ding, Tohru Kondo
Tool-augmented large language model agents are increasingly proposed for network configuration, but routing protocols differ in the control-plane state each commanded router can observe. This difference creates a specific problem for multi-agent orchestration: agents may coordinate more, yet still fail when correct verification depends on peer- or remote-router evidence. We study this interaction through 350 controlled runs on RIP, OSPF, and BGP tasks implemented with FRRouting and Containerlab, comparing a single-agent baseline with multi-agent orchestration patterns across language models. Protocol-centric trace metrics, including spatial coverage, coordination tax, and cross-router verification gap, are combined with intent-property scores and model-balanced bootstrap analysis. The results show that observability explains performance more clearly than orchestration patterns: multi-agent templates trail the baseline on local RIP feedback, show only small and uncertain gains on single-area OSPF troubleshooting, and remain near zero on stricter multi-area OSPF and BGP tasks where peer-side verification gaps are often complete. The main contribution is therefore a protocol-centered account of when agentic orchestration helps, when it adds coordination cost, and why current architectures face a cross-router verification ceiling.