Enhancing Sign Language Recognition Accuracy Through EfficientFormer-L1 Transfer Learning
ID:84
Submission ID:475 View Protection:ATTENDEE
Updated Time:2026-07-22 16:09:47 Hits:18
Online
Start Time:2026-07-30 15:25 (Asia/Kolkata)
Duration:15min
Session:[S4] Computer Vision and Pattern Recognition » [S4-2] Computer Vision and Pattern Recognition
Video
No Permission
Presentation File
Tips: The file permissions under this presentation are only for participants. You have not logged in yet and cannot view it temporarily.
Abstract
This research presents a deep learning-based
approach for sign language recognition using the
EfficientFormer-L1 architecture, a hybrid model combining the
efficiency of convolutional networks with the representational
power of Transformers. The system was trained and evaluated
on the Sign Language MNIST dataset, which contains 27,455
training images and 7,172 testing images representing 25 classes
of the American Sign Language (ASL) alphabet. The images
were preprocessed and resized to 224×224 pixels, and transfer
learning was applied using the pre-trained EfficientFormer-L1
backbone, fine-tuned for classification. The model was
optimized using the Adam optimizer with a learning rate
scheduler to enhance convergence stability. Experimental
results demonstrated a strong capability of EfficientFormer-L1
in learning discriminative gesture features, achieving a final test
accuracy of 99.9-100% while maintaining computational
efficiency due to its lightweight design. This study highlights the
effectiveness of Transformer-based hybrid models for real-time
sign language recognition and provides a robust foundation for
assistive communication technologies.
approach for sign language recognition using the
EfficientFormer-L1 architecture, a hybrid model combining the
efficiency of convolutional networks with the representational
power of Transformers. The system was trained and evaluated
on the Sign Language MNIST dataset, which contains 27,455
training images and 7,172 testing images representing 25 classes
of the American Sign Language (ASL) alphabet. The images
were preprocessed and resized to 224×224 pixels, and transfer
learning was applied using the pre-trained EfficientFormer-L1
backbone, fine-tuned for classification. The model was
optimized using the Adam optimizer with a learning rate
scheduler to enhance convergence stability. Experimental
results demonstrated a strong capability of EfficientFormer-L1
in learning discriminative gesture features, achieving a final test
accuracy of 99.9-100% while maintaining computational
efficiency due to its lightweight design. This study highlights the
effectiveness of Transformer-based hybrid models for real-time
sign language recognition and provides a robust foundation for
assistive communication technologies.
Keywords
EfficientFormer-L1,Sign Language,Transfer Learning,Deaf and DumbDeaf and Dumb,,Deep Learning.
Speaker
Comment submit