Vivek Gujar, Advait Gujar, Ashwani Kumar Rathore
Designing AI-enabled video surveillance systems has become far more challenging than simply deciding the number of cameras or estimating storage. Once AI inference is performed on cameras or edge devices, every design decision influences several others. For example, increasing the pixel density needed for facial recognition affects lens selection, bitrate, storage capacity, and network bandwidth, while adding more AI analytics changes edge-computing requirements and software licensing. As a result, conventional CCTV planning methods are no longer sufficient for modern AI deployments. This paper presents the architecture of IndoAI’s integrated planning toolkit, which follows a three-stage workflow ie Design, Size, and Verify, using nine specialised calculators for camera placement, infrastructure sizing, and deployment validation. Alongside these tools, an Appization-based AI Agent License Sizing and ROI Predictor estimates software licensing requirements and deployment economics. Building on recent advances in Vision Language Models (VLMs) and platform economics, we propose a VLM as the natural-language interface to this planning framework. Users can describe surveillance requirements using text, photographs or floor plans instead of manually entering technical parameters. The VLM converts these inputs into structured data for the calculator suite. We also present two systems under development: a VLM-driven video intelligence pipeline that analyses surveillance footage to generate timestamped event detections, evaluated using the UCF-Crime dataset and a VLM-assisted Bill of Quantities generator that produces annotated camera layouts together with storage, bandwidth and licensing estimates for deployment planning.