《计算机技术与发展杂志》发表论文赏析

低资源青岛方言语音识别方法研究

来源:计算机技术与发展杂志2024年第04期北京时间:

作者:相紫涵;谷潇;饶崇郅;渐令

单位:中国石油大学(华东) 经济管理学院,山东 青岛 266580 Author(s): XIANG Zi-han;GU Xiao;RAO Chong-zhi;JIAN Ling School of Economics and Management,China University of Petroleum ( East China) ,Qingdao 266580,China 关键词: 语音识别;端到端;低资源;数据增强;青岛方言 Keywords: speech recognition;end-to-end;low resource;data augmentation;Qingdao dialect 分类号: TP391. 42;TN912. 34 DOI: 10. 3969 / j. issn. 1673-629X. 2024. 04. 022 摘要: 方言识别是语音识别的重要研究方向,常见的语音识别系统是基于标准语言训练的,导致其方言识别效果不佳。鉴于此,该文选择青岛方言作为应用案例开展方言语音识别研究。 为解决方言语料匮乏、训练深度网络模型困难导致识别准确率受限等问题,提出应用数据增强方法,搭建基于改进 Conformer 的方言语音识别模型。 首先,收集多源语音数据构建方言小型语料库;其次,采用数据增强技术扩充训练数据,以解决语料匮乏问题;最后,为了更好地提取信息,改进Conformer 模型的降采样结构,引入膨胀卷积和 Mish 激活函数,实现语音到文本的直接映射。 实验结果表明,提出的改进降采样模块的端到端模型结合数据增强方法后字错率可达 25. 96% ,能有效实现低资源条件下的方言识别。 Abstract: Dialect recognition is an important research direction in automatic speech recognition. Common speech recognition systems arebased on standard language training,which results in poor performance in dialect recognition. In view of this,we choose Qingdao dialectas an application case for dialect speech recognition research. In order to solve?the problems of lack of dialect corpus and difficulty intraining deep network model,which lead to limited recognition accuracy, we propose to apply data augmentation method?and build adialect speech recognition model based on improved Conformer. Firstly,multi-source speech data is collected to construct a small-scaledialect corpus. Secondly,data augmentation techniques are applied to expand the training data to address the problem of data scarcity. Finally,in order to better extract information,the down -sampling structure of the Conformer model is improved,and dilated convolutionand Mish activation function are introduced to realize the direct mapping from speech to text. Experimental results show that the charactererror rate of the end-to-end model with improved down-sampling module combined with data augmentation method can reach 25. 96% ,which can effectively realize dialect recognition under low resource conditions.

摘要:方言识别是语音识别的重要研究方向,常见的语音识别系统是基于标准语言训练的,导致其方言识别效果不佳。鉴于此,该文选择青岛方言作为应用案例开展方言语音识别研究。 为解决方言语料匮乏、训练深度网络模型困难导致识别准确率受限等问题,提出应用数据增强方法,搭建基于改进 Conformer 的方言语音识别模型。 首先,收集多源语音数据构建方言小型语料库;其次,采用数据增强技术扩充训练数据,以解决语料匮乏问题;最后,为了更好地提取信息,改进Conformer 模型的降采样结构,引入膨胀卷积和 Mish 激活函数,实现语音到文本的直接映射。 实验结果表明,提出的改进降采样模块的端到端模型结合数据增强方法后字错率可达 25. 96% ,能有效实现低资源条件下的方言识别。

关键词:语音识别;端到端;低资源;数据增强;青岛方言

填文献完整题目 获取完整文献

填写需求
联系方式
注:学术顾问会在1小时内联系您,请留意!