Diagnosing Video Foundation Models for Single-Signer RGB-Only Auslan-Daily News Translation

Authors: Maoyang Li, Prashan Premaratne, Peter Vial
Conference: ICIC 2026 Posters, Toronto, Canada, July 22-26, 2026
Pages: -
Keywords: Sign Language Translation, Video Foundation Models, Low-Resource Auslan Translation.

Abstract

Sign language recognition (SLR) through computer vision could be seen as one of the most challenging tasks associated with AI and computer vision. Many attempts in the world have still not shown reasonable progress in this approach. Recent advancements in video foundation models have made it possible to improve sign language translation (SLT) performance by adopting stronger visual encoders, especially in low-resource settings. In this paper, we examine this assumption on the Auslan-Daily News split under a single-signer, RGB-only setting. We systematically compare three representative pipeline families, namely I3D, Image Swin-T, and VideoMAE, under a unified evaluation protocol. VideoMAE achieves the best performance but still remains clearly lower than the official benchmark. To better understand this gap, we further analyze model outputs, test sensitivity to temporal order and visual content, and compare several VideoMAE-based variants. Our results show that the limitation cannot be explained by backbone choice alone. The main difficulty lies in how long video sequences are compressed and then carried into the translation stage under limited supervision. And future progress in low-resource Australian Sign Language (Auslan) translation will depend not only on stronger visual encoders, but also on better temporal modeling and more reliable visual-to-text transfer.
📄 View Full Paper (PDF) 📋 Show Citation