Speech Recognition for Lampung LanguageUsing Wav2Vec 2.0
DOI:
https://doi.org/10.62201/fxnqm295Keywords:
Lampung Language , Speech Recognition, Low-resource LanguageAbstract
Indonesia is home to more than 700 regional languages, many of which are threatened by declining numbers of speakers. One of these languages is Lampung, a low-resource language that is experiencing language shift and faces the risk of extinction. Automatic Speech Recognition (ASR) can support language preservation by enabling speech-to-text conversion; however, developing ASR systems for Lampung remains challenging due to the limited availability of annotated speech data. This study investigates the adaptation of self-supervised speech models for low-resource Lampung ASR through fine-tuning of the Wav2Vec 2.0 Base (W2V2-B) and Wav2Vec 2.0 Large (W2V2-L) architectures. A Lampung speech corpus consisting of approximately 10 hours of transcribed speech and 17,682 audio recordings was used for training and evaluation. The dataset was divided into training and testing subsets, and model performance was assessed using Word Error Rate (WER). The main contribution of this work is the establishment of a benchmark for Lampung ASR using state-of-the-art self-supervised learning models and an empirical comparison of W2V2 architectures in a low-resource language setting. Experimental results show that W2V2-B achieved a WER of 36.23%, while W2V2-L achieved a WER of 36.30%. These findings demonstrate the feasibility of applying self-supervised speech representation learning to low-resource Lampung languages and provide a foundation for future research on multilingual transfer learning and data augmentation for Lampung ASR.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Proceeding of International Conference on Digital, Social, and Science

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.







