Precise finger articulation in sign language video generation
Journal
Enterprise Information Systems
ISSN
1751-7575
1751-7583
Date Issued
2026-02-18
Author(s)
Hsiao, Po-Hsuan
Yu, Jia-Xuan
Chen, Kuan-Lun
Fan Chiang, Po-Hsuan
Chien, Jen-Tzung
Wu, Hsin-Te
Abstract
Sign language relies on coordinated facial expressions and hand gestures, posing challenges for generative video systems that typically emphasize facial or body realism alone. This paper presents a virtual sign language video generation framework based on Stable Video Diffusion (SVD) that produces full-length signing animations from a single portrait. Temporal consistency is improved via frame refinement, while SignPose and PoseNet enhance motion accuracy and detail. A hand-region enhancement strategy further improves finger reconstruction. Experimental results demonstrate clearer gestures, more natural facial expressions, and realistic, comprehensible virtual sign language videos.
Subjects
convolutional neural networks
deep learning
generative AI
sign language
Stable Video Diffusion
Publisher
Taylor and Francis Ltd.
Type
journal article
