SGAFormer: Skeleton-Graph and Agent-guided Transformer network for monocular video-based 3D human pose estimation

Authors: Min Li, Lina Du, Yuanyao Lu, Ning Wang, Zixuan Xu
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Monocular 3D human pose estimation, Skeleton-constrained graph attention, Transformer, Agent token, Spatio-temporal modeling.

Abstract

Monocular video-based 3D human pose estimation remains challenging due to depth ambiguity, noisy 2D observations, and complex spatio-temporal depend-encies. Existing Transformer-based methods can capture global relationships, but they often lack explicit skeletal topology constraints and may introduce re-dundant dense temporal interactions. To address these issues, this paper pro-poses a Skeleton-Graph and Agent-guided Transformer network (SGAFormer) for skeleton-aware and temporally coherent 3D human pose estimation from 2D pose sequences. In spatial modeling, the proposed Skeleton-constrained Adaptive Graph-order Attention Module (SAGA) introduces skeleton-constrained dynamic graph attention and multi-order graph propagation to cap-ture joint self-information, direct skeletal connections, and indirect structural dependencies. By adaptively fusing different graph-order branches, SAGA en-hances dynamic joint interaction modeling under physical skeleton constraints. In temporal modeling, the proposed Agent-token Global-local Temporal At-tention Module (AGTA) reorganizes dense frame-wise interactions through agent tokens, while incorporating temporal positional bias and depthwise sepa-rable temporal convolution to enhance global-local motion representation. Ex-periments on Human3.6M show that SGAFormer remains competitive under detected 2D keypoint inputs, achieves superior temporal consistency, and ob-tains strong upper-bound performance with ground-truth 2D keypoints, demon-strating the effectiveness of the proposed spatial-temporal modeling frame-work.
📄 View Full Paper (PDF) 📋 Show Citation