Jinhua Peng

dblp:47/8637 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LooM: Learning-Based Multipath Scheduling for Out-of-Order Mitigation in Mobile Networks
abstract
Multipath transmission offers bandwidth aggregation capabilities for mobile networks. However, path heterogeneity and user mobility often lead to increased packet out-oforder (OFO) rate, causing buffer blocking, reduced throughput, and degraded transmission quality. To mitigate the OFO effect in multipath transmission, this paper proposes LooM, a learning-based multipath scheduler. LooM is designed to optimize throughput and OFO rate, employing a learning-based scheduling strategy to achieve the optimal packet scheduling under fluctuating paths. Particularly, LooM employs singleround scheduling as its basic unit, calculating the number of OFO packets across scheduling units to dynamically set path blocking delay, thereby adjusting packet transmission order to ensure in-order delivery. Simulation results show that, compared to traditional scheduling algorithms, LooM reduces the OFO rate by 16% while maintaining high throughput and achieving a lower packet loss rate.
Mingyuan Liu 0001, Jinhua Peng, Nan Cheng 0001, Wei Quan 0001
ICC4
2023 Scalable Identity-Oriented Speech Retrieval
abstract
With the prevalence of voice devices in our daily life, speech data is accumulated at an unprecedented speed, forming an invaluable database for security surveillance and financial risk management. In these applications, a key task is given a querying speech snippet to retrieve all speech snippets that are uttered by the same speaker as the querying one, namely Identity-Oriented Speech Retrieval (IO-SR). In this paper, we propose an accuracy and scalable system for IO-SR, which seamlessly integrates speaker modeling and deep indexing techniques. Evaluations on an industrial dataset containing millions of speech snippets show that our system achieves superior performance compared with the state-of-the-art methods.
Chaotao Chen, Di Jiang 0004, Jinhua Peng, Rongzhong Lian, Yawen Li 0001, Chen Zhang 0013, Lei Chen 0002, Lixin Fan
IEEE Trans. Knowl. Data Eng.3
2021 A Health-friendly Speaker Verification System Supporting Mask Wearing
abstract
We demonstrate a health-friendly speaker verification system for voice-based identity verification on mobile devices. The system is built upon a speech processing module, a ResNet-based local acoustic feature extractor and a multi-head attention-based embedding layer, and is optimized under an additive margin softmax loss for discriminative speaker verification. It is shown that the system achieves superior performance no matter whether there is mask wearing or not. This characteristic is important for speaker verification services operating in regions affected by the raging coronavirus pneumonia. With this demonstration, the audience will have an in-depth experience of how the accuracy of bio-metric verification and the personal health are simultaneously ensured. We wish that this demonstration would boost the development of next-generation bio-metric verification technologies.
Chaotao Chen, Di Jiang 0004, Jinhua Peng, Rongzhong Lian, Chen Zhang 0013, Qian Xu 0005, Lixin Fan, Qiang Yang 0001
AAAI3
2021 Familia: A Configurable Topic Modeling Framework for Industrial Text Engineering
Di Jiang 0004, Yuanfeng Song, Rongzhong Lian, Siqi Bao, Jinhua Peng, Huang He, Hua Wu 0003, Chen Zhang 0013, Lei Chen 0002
DASFAA (3)5
2021 A GDPR-compliant Ecosystem for Speech Recognition with Transfer, Federated, and Evolutionary Learning
abstract
Automatic Speech Recognition (ASR) is playing a vital role in a wide range of real-world applications. However, Commercial ASR solutions are typically “one-size-fits-all” products and clients are inevitably faced with the risk of severe performance degradation in field test. Meanwhile, with new data regulations such as the European Union’s General Data Protection Regulation (GDPR) coming into force, ASR vendors, which traditionally utilize the speech training data in a centralized approach, are becoming increasingly helpless to solve this problem, since accessing clients’ speech data is prohibited. Here, we show that by seamlessly integrating three machine learning paradigms (i.e., T ransfer learning, F ederated learning, and E volutionary learning (TFE)), we can successfully build a win-win ecosystem for ASR clients and vendors and solve all the aforementioned problems plaguing them. Through large-scale quantitative experiments, we show that with TFE, the clients can enjoy far better ASR solutions than the “one-size-fits-all” counterpart, and the vendors can exploit the abundance of clients’ data to effectively refine their own ASR products.
Di Jiang 0004, Conghui Tan, Jinhua Peng, Chaotao Chen, Xueyang Wu 0001, Yuanfeng Song, Yongxin Tong, Chang Liu 0069, Qian Xu 0005, Qiang Yang 0001
ACM Trans. Intell. Syst. Technol.3
2021 Industrial Federated Topic Modeling
abstract
Probabilistic topic modeling has been applied in a variety of industrial applications. Training a high-quality model usually requires a massive amount of data to provide comprehensive co-occurrence information for the model to learn. However, industrial data such as medical or financial records are often proprietary or sensitive, which precludes uploading to data centers. Hence, training topic models in industrial scenarios using conventional approaches faces a dilemma: A party (i.e., a company or institute) has to either tolerate data scarcity or sacrifice data privacy. In this article, we propose a framework named Industrial Federated Topic Modeling (iFTM), in which multiple parties collaboratively train a high-quality topic model by simultaneously alleviating data scarcity and maintaining immunity to privacy adversaries. iFTM is inspired by federated learning, supports two representative topic models (i.e., Latent Dirichlet Allocation and SentenceLDA) in industrial applications, and consists of novel techniques such as private Metropolis-Hastings, topic-wise normalization, and heterogeneous model integration. We conduct quantitative evaluations to verify the effectiveness of iFTM and deploy iFTM in two real-life applications to demonstrate its utility. Experimental results verify iFTM’s superiority over conventional topic modeling.
Di Jiang 0004, Yongxin Tong, Yuanfeng Song, Xueyang Wu 0001, Jinhua Peng, Rongzhong Lian, Qian Xu 0005, Qiang Yang 0001
ACM Trans. Intell. Syst. Technol.6
2020 Federated Acoustic Model Optimization for Automatic Speech Recognition
Conghui Tan, Di Jiang 0004, Huaxiao Mo, Jinhua Peng, Yongxin Tong, Chaotao Chen, Rongzhong Lian, Yuanfeng Song, Qian Xu 0005
DASFAA (3)4
2020 A De Novo Divide-and-Merge Paradigm for Acoustic Model Optimization in Automatic Speech Recognition
abstract
Due to the rising awareness of privacy protection and the voluminous scale of speech data, it is becoming infeasible for Automatic Speech Recognition (ASR) system developers to train the acoustic model with complete data as before. In this paper, we propose a novel Divide-and-Merge paradigm to solve salient problems plaguing the ASR field. In the Divide phase, multiple acoustic models are trained based upon different subsets of the complete speech data, while in the Merge phase two novel algorithms are utilized to generate a high-quality acoustic model based upon those trained on data subsets. We first propose the Genetic Merge Algorithm (GMA), which is a highly specialized algorithm for optimizing acoustic models but suffers from low efficiency. We further propose the SGD-Based Optimizational Merge Algorithm (SOMA), which effectively alleviates the efficiency bottleneck of GMA and maintains superior performance. Extensive experiments on public data show that the proposed methods can significantly outperform the state-of-the-art.
Conghui Tan, Di Jiang 0004, Jinhua Peng, Xueyang Wu 0001, Qian Xu 0005, Qiang Yang 0001
IJCAI3
2019 Generating Multiple Diverse Responses with Multi-Mapping and Posterior Mapping Selection
abstract
In human conversation an input post is open to multiple potential responses, which is typically regarded as a one-to-many problem. Promising approaches mainly incorporate multiple latent mechanisms to build the one-to-many relationship. However, without accurate selection of the latent mechanism corresponding to the target response during training, these methods suffer from a rough optimization of latent mechanisms. In this paper, we propose a multi-mapping mechanism to better capture the one-to-many relationship, where multiple mapping modules are employed as latent mechanisms to model the semantic mappings from an input post to its diverse responses. For accurate optimization of latent mechanisms, a posterior mapping selection module is designed to select the corresponding mapping module according to the target response for further optimization. We also introduce an auxiliary matching loss to facilitate the optimization of posterior mapping selection. Empirical results demonstrate the superiority of our model in generating multiple diverse and informative responses over the state-of-the-art methods.
Chaotao Chen, Jinhua Peng, Fan Wang 0021, Jun Xu 0027, Hua Wu 0003
IJCAI2
2019 Learning to Select Knowledge for Response Generation in Dialog Systems
abstract
End-to-end neural models for intelligent dialogue systems suffer from the problem of generating uninformative responses. Various methods were proposed to generate more informative responses by leveraging external knowledge. However, few previous work has focused on selecting appropriate knowledge in the learning process. The inappropriate selection of knowledge could prohibit the model from learning to make full use of the knowledge. Motivated by this, we propose an end-to-end neural model which employs a novel knowledge selection mechanism where both prior and posterior distributions over knowledge are used to facilitate knowledge selection. Specifically, a posterior distribution over knowledge is inferred from both utterances and responses, and it ensures the appropriate selection of knowledge during the training process. Meanwhile, a prior distribution, which is inferred from utterances only, is used to approximate the posterior distribution so that appropriate knowledge can be selected even without responses during the inference process. Compared with the previous work, our model can better incorporate appropriate knowledge in response generation. Experiments on both automatic and human evaluation verify the superiority of our model over previous baselines.
Rongzhong Lian, Fan Wang 0021, Jinhua Peng, Hua Wu 0003
IJCAI4