GeoGTM: Vision-Language Geospatial Intelligence for Street-Level Telecom Growth

Authors: Junchi Ren, Chengyu Zhou, Tao He
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Grounded multimodal agent \and Telecom prospecting \and Deliverability-aware reasoning \and Evidence citation

Abstract

Street-scale telecom growth requires jointly reasoning over heterogeneous evidence, including geospatial context (buildings, communities, and business districts), network capability (coverage and capacity), and customer-side signals (service usage and support tickets). Existing LLM-based sales assistants are often ungrounded: they ignore deliverability constraints, lack verifiable evidence, and fail to produce actionable plans that can be executed by field teams. In this paper, we introduce StreetCopilot, a grounded vision-language agent for street-level prospecting and service planning. StreetCopilot integrates (i) geospatial visual cues (e.g., street-view/remote-sensing building context and POIs), (ii) structured telecom signals (coverage maps, traffic KPIs, product portfolios), and (iii) unstructured operational text (tickets and visit notes) via a retrieval-augmented reasoning pipeline. To ensure actionability, we propose a deliverability-aware constraint module that verifies whether recommended bundles (network, wireless coverage, industrial devices, and cloud services) are feasible under local coverage and resource conditions, and a citation-grounded generation mechanism that attaches evidence snippets to each recommendation for auditability. We further present a street-scale closed-loop evaluation protocol that measures not only recommendation accuracy but also plan feasibility, evidence faithfulness, and end-to-end business outcomes (lead acceptance and conversion). Experiments on a real-world deployment in a city subregion demonstrate that StreetCopilot substantially improves prospect ranking quality and proposal drafting efficiency while maintaining high feasibility and evidence faithfulness, shedding light on grounded multimodal agents for real-world decision-making.
📄 View Full Paper (PDF) 📋 Show Citation