A frame-level video annotation tool for dynamic gestures and apply in Vietnamese sign language
Keywords
DOI:
https://doi.org/10.54939/1859-1043.j.mst.112.2026.167-175Abstract
Human Action Recognition (HAR) in video is essential for human–computer interaction, particularly in sign language and smart device control. However, model performance depends heavily on accurate temporal annotation of dynamic gestures. This study proposes a frame-level video annotation tool for Vietnamese sign language, enabling precise temporal segmentation and structured JSON export for deep learning applications. A dataset of 15 dynamic gesture classes is also constructed using a multi-view acquisition setup. Annotation quality is evaluated using Temporal IoU, Cohen’s Kappa, and annotation time. Results show high inter-annotator agreement at 0.9470 and 0.9888, respectively, demonstrating the effectiveness of the proposed tool for reliable and efficient gesture annotation.