Yizhou Lu

dblp:00/2235 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
11since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Exploratory Combinatorial Optimization Problem Solving via Gauge Transformation
abstract
The combinatorial optimization problems (COPs) over graph are of great significance both in theory and practice, covering a wide range of scenarios in daily life and industrial production. Recent years, reinforcement learning (RL) based models have emerged as a promising direction, which treat solving the COPs as a heuristic learning problem. However, current finite-horizon Markov Decision Process (MDP) based RL models are not allowed to explore adquately for improving solutions at test time, which may be necessary given the complexity of NP-hard optimization tasks. Some recent attempts solve this issue by focusing on reward design and state feature engineering, which are tedious and ad-hoc. To address this challenge, we introduce a physics-inspired technique called gauge transformation (GT), which is highly effective in enabling RL agents to explore and continuously enhance solution quality during testing. GT seamlessly transforms any state within the MDP back to its initial state, allowing the RL agent to continue exploration within the transformed space. Empirically, we demonstrate that traditional RL models equipped with the GT technique achieve the SOTA performance on the MaxCut problem. Moreover, GT is exclusively applied during testing and does not alter the training phase of the model. It can be readily integrated into existing RL models, providing a pathway for more effective exploration in solving the COPs.
Tianle Pu, Changjun Fan, Mutian Shen, Yizhou Lu, Zohar Nussinov
ICDM4
2023 ClusterSeg: A crowd cluster pinpointed nucleus segmentation framework with cross-modality datasets
Jing Ke, Yizhou Lu, Yiqing Shen 0003, Junchao Zhu, Yijin Zhou, Jinghan Huang 0002, Jieteng Yao, Xiaoyao Liang, Yi Guo 0001, Zhonghua Wei, Fusong Jiang, Dinggang Shen
Medical Image Anal.2
2023 Artifact Detection and Restoration in Histology Images With Stain-Style and Structural Preservation
abstract
The artifacts in histology images may encumber the accurate interpretation of medical information and cause misdiagnosis. Accordingly, prepending manual quality control of artifacts considerably decreases the degree of automation. To close this gap, we propose a methodical pre-processing framework to detect and restore artifacts, which minimizes their impact on downstream AI diagnostic tasks. First, the artifact recognition network AR-Classifier first differentiates common artifacts from normal tissues, e.g., tissue folds, marking dye, tattoo pigment, spot, and out-of-focus, and also catalogs artifact patches by their restorability. Then, the succeeding artifact restoration network AR-CycleGAN performs de-artifact processing where stain styles and tissue structures can be maximally retained. We construct a benchmark for performance evaluation, curated from both clinically collected WSIs and public datasets of colorectal and breast cancer. The functional structures are compared with state-of-the-art methods, and also comprehensively evaluated by multiple metrics across multiple tasks, including artifact classification, artifact restoration, downstream diagnostic tasks of tumor classification and nuclei segmentation. The proposed system allows full automation of deep learning based histology image analysis without human intervention. Moreover, the structure-independent characteristic enables its processing with various artifact subtypes. The source code and data in this research are available at https://github.com/yunboer/AR-classifier-and-AR-CycleGAN.
Jing Ke, Kai Liu 0034, Yuxiang Sun 0004, Yuying Xue, Jiaxuan Huang, Yizhou Lu, Yaobing Chen, Xiaodan Han, Yiqing Shen 0003, Dinggang Shen
IEEE Trans. Medical Imaging6
2022 Language Adaptive Cross-Lingual Speech Representation Learning with Sparse Sharing Sub-Networks
abstract
Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across multiple languages. However, standard XLSR model suffers from language interference problem due to the lack of language specific modeling ability. In this work, we investigate language adaptive training on XLSR models. More importantly, we propose a novel language adaptive pretraining approach based on sparse sharing sub-networks. It makes room for language specific modeling by pruning out unimportant parameters for each language, without requiring any manually designed language specific component. After pruning, each language only maintains a sparse sub-network, while the sub-networks are partially shared with each other. Experimental results on a downstream multilingual speech recognition task show that our proposed method significantly outperforms baseline XLSR models on both high resource and low resource languages. Besides, our proposed method consistently outperforms other adaptation methods and requires fewer parameters.
Yizhou Lu, Mingkun Huang, Xinghua Qu, Pengfei Wei 0001, Zejun Ma 0001
ICASSP1
2022 Time-Varying Group Formation-Containment Tracking Control for General Linear Multiagent Systems With Unknown Inputs
abstract
Time-varying group formation-containment tracking problems for general linear multiagent systems with unknown control input are investigated. Agents are classified into tracking leaders, formation leaders, and followers and assigned in groups. Tracking leaders with unknown control inputs provide unpredictable trajectories as macroscopic moving references. Formation leaders accomplish desired subformations while following the trails of tracking leaders. At the same time, followers converge into different convex hulls spanned by formation leaders. First, formation-containment tracking protocols are designed with neighboring relative information and effects of unknown input of tracking leaders. Then, the design of group division is analyzed by adjusting the properties in Laplacian matrices, which represent interaction relationships. An algorithm to determine the parameters in control protocols is proposed, and the formation tracking feasible constraint is presented. Next, it is proved that the general linear multiagent system can achieve time-varying group formation-containment control effectively with errors uniformly asymptotically converging to zero under designed protocols. Finally, a numerical simulation is given to verify the effectiveness of obtained theoretical results.
Yizhou Lu, Xiwang Dong, Qingdong Li, Jinhu Lü 0001, Zhang Ren
IEEE Trans. Cybern.1
2021 Cluster Image Patches with Multiple Mutual Information in Unlabelled Whole-Slide Image
abstract
The massive annotation workload has always hindered the progress towards an automatic analysis of gigapixel whole-slide images. Histologically, individual patches from a constrained spatial region may share rich phenotypic information, where the morphological correlations have the potential to be mined for a grouping or clustering task. In this paper, we propose a clustering technique to extract multiple mutual information from histology images without prior domain knowledge. Specifically, our framework automatically localizes morphologically homogeneous patches within an extended solution space. Our novelty is an expanse and the pattern with which invariant information can be learnt, in contrast to the current literature of feature generation or parametric transformation within an individual patch. Additionally, structure-independent, the model may be applicable to any backbone convolutional neural network architectures. The empirical validation on The Cancer Genome Atlas (TCGA) datasets illustrates an observable margin of patch-level classification accuracy in comparison with state-of-the-art unsupervised approaches.
Yiqing Shen 0003, Yizhou Lu, Yulin Luo, Jing Ke
BIBM2
2021 AISpeech-SJTU ASR System for the Accented English Speech Recognition Challenge
abstract
This paper describes the AISpeech-SJTU ASR system for the Interspeech-2020 Accented English Speech Recognition Challenge (AESRC). This task is challenging due to the diversity of pronunciation accuracy, intonation speed and pronunciation of some syllables. All participants were restricted to develop their systems based on the speech and text corpora provided by the organizer. To work around the data-scarcity problem, data augmentation was first explored including noise simulation, SpecAugment, speed perturbation and TTS simulation. Moreover, SOTA CNN-transformer-based joint CTC-attention system was built and accent adaptation was proposed to train an accent robust system. Finally, the first-pass recognition hypotheses generated from CTC head were rescored by forward, backward LSTM-LM and the attention head. Our system with the best configuration achieves second place in the challenge, resulting in a word error rate (WER) of 4.00% on dev set and 4.47% WER on test set, while WER on test set of the top-performing, second runner-up and official baseline systems are 4.06%, 4.52%, 8.29%, respectively.
Tian Tan 0002, Yizhou Lu, Rao Ma, Sen Zhu, Yanmin Qian
ICASSP2
2021 The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods
abstract
The variety of accents has posed a big challenge to speech recognition. The Accented English Speech Recognition Challenge (AESRC2020) is designed for providing a common testbed and promoting accent-related research. Two tracks are set in the challenge – English accent recognition (track 1) and accented English speech recognition (track 2). A set of 160 hours of accented English speech collected from 8 countries is released with labels as the training set. Another 20 hours of speech without labels is later released as the test set, including two unseen accents from another two countries used to test the model generalization ability in track 2. We also provide baseline systems for the participants. This paper first reviews the released dataset, track setups, baselines and then summarizes the challenge results and major techniques used in the submissions.
Xian Shi, Fan Yu 0002, Yizhou Lu, Yuhao Liang, Qiangze Feng, Daliang Wang, Yanmin Qian, Lei Xie 0001
ICASSP3
2021 Towards Data Selection on TTS Data for Children's Speech Recognition
abstract
Although great progress has been made on automatic speech recognition (ASR) systems, children’s speech recognition still remains a challenging task. General ASR systems for children’s speech suffer from the lack of corpora and mismatch between children’s and adults’ speech. Efforts have been made to reduce such mismatch by applying normalization methods to generate modified adults’ speech for ASR training. However, modified adults’ data can reflect the characteristics of children’s speech to a very limited extent. In this work, we adopt text-to-speech data augmentation to improve the performance of children’s speech recognition system. We find that the children’s TTS model generates speech with inconsistent quality due to children’s substandard pronunciations of phonemes, and the ASR system suffers when trained with these additional synthesized data. To solve this problem, we propose data selection strategies on the TTS augmented data, and the effectiveness of the synthesized data can be substantially boosted for children’s ASR modeling. We show that the speaker embedding similarity based data selection strategy can obtain the best position: relative 14.0% and 14.7% CER reduction for child conversation and child reading test set respectively compared to the baseline model trained on real data.
Wei Wang 0010, Zhikai Zhou, Yizhou Lu, Chenpeng Du, Yanmin Qian
ICASSP3
2021 Layer-Wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
abstract
Accent variability has posed a huge challenge to automatic speech recognition~(ASR) modeling. Although one-hot accent vector based adaptation systems are commonly used, they require prior knowledge about the target accent and cannot handle unseen accents. Furthermore, simply concatenating accent embeddings does not make good use of accent knowledge, which has limited improvements. In this work, we aim to tackle these problems with a novel layer-wise adaptation structure injected into the E2E ASR model encoder. The adapter layer encodes an arbitrary accent in the accent space and assists the ASR model in recognizing accented speech. Given an utterance, the adaptation structure extracts the corresponding accent information and transforms the input acoustic feature into an accent-related feature through the linear combination of all accent bases. We further explore the injection position of the adaptation layer, the number of accent bases, and different types of accent bases to achieve better accent adaptation. Experimental results show that the proposed adaptation structure brings 12\% and 10\% relative word error rate~(WER) reduction on the AESRC2020 accent dataset and the Librispeech dataset, respectively, compared to the baseline.
Xun Gong 0005, Yizhou Lu, Zhikai Zhou, Yanmin Qian
Interspeech2
2021 Data Augmentation for end-to-end Code-Switching Speech Recognition
abstract
Training a code-switching end-to-end automatic speech recognition (ASR) model normally requires a large amount of data, while code-switching data is often limited. In this paper, three novel approaches are proposed for code-switching data augmentation. Specifically, they are audio splicing with the existing code-switching data, and TTS with new code-switching texts generated by word translation or word insertion. Our experiments on 200 hours Mandarin-English code-switching dataset show that all the three proposed approaches yield significant improvements on code-switching ASR individually. Moreover, all the proposed approaches can be combined with recent popular SpecAugment, and an addition gain can be obtained. WER is significantly reduced by relative 24.0% compared to the system without any data augmentation, and still relative 13.0% gain compared to the system with only SpecAugment.
Chenpeng Du, Yizhou Lu, Yanmin Qian
SLT3
2020 Bi-Encoder Transformer Network for Mandarin-English Code-Switching Speech Recognition Using Mixture of Experts
Yizhou Lu, Mingkun Huang, Yanmin Qian
INTERSPEECH1
2020 Modular End-to-End Automatic Speech Recognition Framework for Acoustic-to-Word Model
abstract
End-to-end (E2E) systems have played a more and more important role in automatic speech recognition (ASR) and achieved great performance. However, E2E systems recognize output word sequences directly with the input acoustic feature, which can only be trained on limited acoustic data. The extra text data is widely used to improve the results of traditional artificial neural network-hidden Markov model (ANN-HMM) hybrid systems. The involving of extra text data to standard E2E ASR systems may break the E2E property during decoding. In this paper, a novel modular E2E ASR system is proposed. The modular E2E ASR system consists of two parts: an acoustic-to-phoneme (A2P) model and a phoneme-to-word (P2W) model. The A2P model is trained on acoustic data, while extra data including large scale text data can be used to train the P2W model. This additional data enables the modular E2E ASR system to model not only the acoustic part but also the language part. During the decoding phase, the two models will be integrated and act as a standard acoustic-to-word (A2W) model. In other words, the proposed modular E2E ASR system can be easily trained with extra text data and decoded in the same way as a standard E2E ASR system. Experimental results on the Switchboard corpus show that the modular E2E model achieves better word error rate (WER) than standard A2W models.
Qi Liu 0018, Zhehuai Chen, Mingkun Huang, Yizhou Lu, Kai Yu 0004
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Exploring Model Units and Training Strategies for End-to-End Speech Recognition
abstract
In this work, we explore end-to-end speech recognition models (CTC, RNN-Transducer and attention-based models) with different model units (character, wordpiece and word) and various training strategies. We show that wordpiece unit outperforms character unit for all end-to-end systems on the Switchboard Hub5'00 benchmark. To improve the performance of end-to-end systems, we propose a multi-stage pretraining strategy, which gives 25.0% and 18.0% relative improvements over training from scratch for attention and RNN-T models respectively with wordpiece units. We achieve state-of-the-art performance on the Switchboard+Fisher-2000h task, outperforming all prior work. Together with other training strategies such as label smoothing and data augmentation, we achieve 5.9%/12.1% WER on the Switch-board/CallHome test set without using any external language models. This is a new performance milestone for a single end-to-end system, and it is also much better than the previous published best hybrid system, which is 6.7%/12.5% on each set individually.
Mingkun Huang, Yizhou Lu, Yanmin Qian, Kai Yu 0004
ASRU2
2004 Link fusion: a unified link analysis framework for multi-type interrelated data objects
abstract
Web link analysis has proven to be a significant enhancement for quality based web search. Most existing links can be classified into two categories: intra-type links (e.g., web hyperlinks), which represent the relationship of data objects within a homogeneous data type (web pages), and inter-type links (e.g., user browsing log) which represent the relationship of data objects across different data types (users and web pages). Unfortunately, most link analysis research only considers one type of link. In this paper, we propose a unified link analysis framework, called "link fusion", which considers both the inter- and intra- type link structure among multiple-type inter-related data objects and brings order to objects in each data type at the same time. The PageRank and HITS algorithms are shown to be special cases of our unified link analysis framework. Experiments on an instantiation of the framework that makes use of the user data and web pages extracted from a proxy log show that our proposed algorithm could improve the search effectiveness over the HITS and DirectHit algorithms by 24.6% and 38.2% respectively.
Wensi Xi, Benyu Zhang, Zheng Chen 0001, Yizhou Lu, Shuicheng Yan, Wei-Ying Ma, Edward A. Fox
WWW4