Shi-Yan Weng

2021

pdf abs
A Preliminary Study on Environmental Sound Classification Leveraging Large-Scale Pretrained Model and Semi-Supervised Learning
You-Sheng Tsao | Tien-Hong Lo | Jiun-Ting Li | Shi-Yan Weng | Berlin Chen
Proceedings of the 33rd Conference on Computational Linguistics and Speech Processing (ROCLING 2021)

With the widespread commercialization of smart devices, research on environmental sound classification has gained more and more attention in recent years. In this paper, we set out to make effective use of large-scale audio pretrained model and semi-supervised model training paradigm for environmental sound classification. To this end, an environmental sound classification method is first put forward, whose component model is built on top a large-scale audio pretrained model. Further, to simulate a low-resource sound classification setting where only limited supervised examples are made available, we instantiate the notion of transfer learning with a recently proposed training algorithm (namely, FixMatch) and a data augmentation method (namely, SpecAugment) to achieve the goal of semi-supervised model training. Experiments conducted on bench-mark dataset UrbanSound8K reveal that our classification method can lead to an accuracy improvement of 2.4% in relation to a current baseline method.

pdf bib
The NTNU Taiwanese ASR System for Formosa Speech Recognition Challenge 2020
Fu-An Chao | Tien-Hong Lo | Shi-Yan Weng | Shih-Hsuan Chiu | Yao-Ting Sung | Berlin Chen
International Journal of Computational Linguistics & Chinese Language Processing, Volume 26, Number 1, June 2021

Shi-Yan Weng

2021

2019

Co-authors

Venues