VLDB 2026 Research / reviewers in the wild / expert
Zhiyuan Tang
dblp:72/7546
· DBLP profile ↗
23ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAD model reconstruction of spherical-jointed lattice structures from point clouds with discrete planar convex hulls
Mulin Yu, Chunjiang Wang, Zhiyuan Tang, Xiangdong Sun |
Comput. Aided Des. | 5 |
| 2026 | CalliRehab: Supporting Motor-Cognitive Recovery in Post-Stroke Rehabilitation Through AR-Enhanced Multisensory Calligraphy TherapyabstractIntegrating motor and cognitive training is crucial for stroke patients. However, interactive technologies for assisting such synergistic trainings are underdeveloped. This article presents the design and evaluation of CalliRehab, an AR-enhanced multisensory tool for facilitating the coordinated training of upper limb motor skills and cognitive functions through calligraphy therapy tasks. The effects of CalliRehab were examined through a within-subject experiment involving 21 stroke patients. Results indicate that, compared to the non-feedback baseline, multimodal feedback improved the task accuracy with a significant reduction in error counts (p < 0.001). Participants’ cognitive load has been remarkably lowered in the NASA task load index (p < 0.001), whereas the intrinsic motivation (p < 0.001) and the technology acceptance (p < 0.001) have been increased significantly. This study confirms that AR-enhanced multisensory training tools can effectively support integrating motor and cognitive training in stroke patients, particularly in enhancing task quality and training experiences. Jing Qu 0001, Linxin Du, Zhiyuan Tang, Lingguo Bu, Xipei Ren |
Int. J. Hum. Comput. Interact. | 4 |
| 2026 | A Two-Stage Data-Driven Topology Identification in Three-Phase Distribution NetworksabstractTopology identification lays out the essential foundation for the operation monitoring and management of distribution networks. In this article, a novel two-stage data-driven topology identification approach is proposed for unbalanced three-phase distribution networks utilizing the measurements of smart meters. In the first stage, the phase sequence of each bus is recovered sequentially using the proposed phase identification method, where the similarity criteria are employed to reduce the influence of line impedance on voltage correlation. In the second stage, based on the phase identification results, the buses with relatively low active power injections are grouped into different clusters. By regarding each cluster as an aggregated node, the simplified system topology and the local topology of each cluster (i.e., each aggregated node) are identified sequentially using the ridge regression method. The full-scale topology of system is obtained by integrating the simplified system topology and the local topology of each cluster. Moreover, to further improve the identification efficiency, a novel optimal input design approach is proposed to select rich-information data from historical records. Various case studies are conducted to demonstrate the effectiveness and advantages of the proposed topology identification approach. Wenjie Xiong, Zhiyuan Tang, Hongjun Gao, Youbo Liu, Ao Qiao, Junyong Liu |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Full-text Error Correction for Chinese Speech Recognition with Large Language ModelabstractLarge Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper investigates the effectiveness of LLMs for error correction in full-text generated by ASR systems from longer speech recordings, such as transcripts from podcasts, news broadcasts, and meetings. First, we develop a Chinese dataset for full-text error correction, named ChFT, utilizing a pipeline that involves text-to-speech synthesis, ASR, and error-correction pair extractor. This dataset enables us to correct errors across contexts, including both full-text and segment, and to address a broader range of error types, such as punctuation restoration and inverse text normalization, thus making the correction process comprehensive. Second, we fine-tune a pre-trained LLM on the constructed dataset using a diverse set of prompts and target formats, and evaluate its performance on full-text error correction. Specifically, we design prompts based on full-text and segment, considering various output formats, such as directly corrected text and JSON-based error-correction pairs. Through various test settings, including homogeneous, up-to-date, and hard test sets, we find that the finetuned LLMs perform well in the full-text setting with different prompts, each presenting its own strengths and weaknesses. This establishes a promising baseline for further research. The dataset is available on the website1. Zhiyuan Tang, Shen Huang, Shidong Shang |
ICASSP | 1 |
| 2025 | Match Made with Matrix Completion: Efficient Offline and Online Learning in Matching MarketsabstractOnline matching markets face increasing needs to accurately learn the matching qualities between demand and supply for effective design of matching policies. However, the growing diversity of participants introduces a high-dimensional challenge in practice, as there are a substantial number of unknown matching rewards and learning all rewards requires a large amount of data. We leverage a natural low-rank matrix structure of the matching rewards in these two-sided markets, and propose to utilize matrix completion (specifically the nuclear norm regularization approach) to accelerate the reward learning process with only a small amount of offline data. A key challenge in our setting is that the matrix entries are observed with matching interference, distinct from the independent sampling assumed in existing matrix completion literature. We propose a new proof technique and prove a near-optimal average accuracy guarantee with improved dependence on the matrix dimensions. Furthermore, to guide matching decisions, we develop a novel "double-enhancement" procedure that refines the nuclear norm regularized estimates and further provides near-optimal entry-wise estimations. Our paper makes the first investigation into adopting matrix completion techniques for matching problems. We also extend our approach to online learning settings for optimal matching and stable matching by incorporating matrix completion in multi-armed bandit algorithms. We present improved regret bounds in matrix dimensions through reduced costs during the exploration phase. Finally, we demonstrate the practical value of our methods using both synthetic data and real data of labor markets. Zhiyuan Tang, Wanning Chen, Kan Xu |
EC | 1 |
| 2024 | Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
Zhiyuan Tang, Shen Huang, Shidong Shang |
INTERSPEECH | 1 |
| 2024 | Research on AI Energy Management of Ultra-Speed Magnetic Levitation Flywheel Energy Storage
Shi Xiao, Zhiyuan Tang |
TENCON | 2 |
| 2024 | Multiagent Soft Actor-Critic Learning for Distributed ESS Enabled Robust Voltage Regulation of Active Distribution GridsabstractIn this article, a novel data-driven robust voltage regulation method employing the multiagent soft actor–critic algorithm for photovoltaic-rich distribution grids considering storage lifetime and topology flexibility is proposed. In the proposed scheme, the active and reactive power from distributed energy storage system (ESS) are coordinated to deliver effective voltage support. To account for the long-term influence of ESS behavior on its lifetime, the life costs associated with the energy throughput are firstly formulated into the reward function of the Markov game-based voltage regulation model. Then, the topology status is represented by continuous variables transformed via Gumbel-softmax and embedded into the local observation of ESS agents for being aware of topology variations due to operational reconfiguration. In addition, to enhance the robustness of the voltage regulation method against imperfect measurements, the designed state space incorporates solely partially observed information from the entire distribution networks. Numerical simulations on IEEE 69-bus and IEEE 141-bus test systems confirm the outperforming of the proposed method over the previously implemented voltage regulation approaches. Yongdong Chen, Youbo Liu, Zhiyuan Tang, Gao Qiu, Junyong Liu |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Topology-Transferable Physics-Guided Graph Neural Network for Real-Time Optimal Power FlowabstractLarger-scale stochastic power systems urge the development of real-time alternating current optimal power flow, artificial intelligence (AI) thus becomes an alternative. However, traditional AI only imitates experiences, and cannot follow in-depth physics. This may cause an undesired nongeneralizability and topology intractability. To address this issue, a physics-guided graph neutral network (PG-GNN) is proposed. The PG-GNN firstly capture the physical constraints by a dual Lagrangian. Besides, the branch features of power grids are fully exploited to allow the PG-GNN to master tremendous topological patterns. To further manage the out-of-distribution topology, stability property of the PG-GNN is proved, then upon this evidence, an online transfer learning is proposed to allow the PG-GNN to fast master the unexpected topology. Numerical tests on benchmarks show that, the proposed method holds well topology-transferability, enables near or even better solutions than conventional optimizer, but merits much more than 100 times efficiency. Gao Qiu, Junyong Liu, Youbo Liu, Tingjian Liu, Zhiyuan Tang, Lijie Ding, Yue Shui, Kai Liu 0012 |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | Semantic Data Augmentation for End-to-End Mandarin Speech RecognitionabstractEnd-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications.During the training of end-to-end ASR, data augmentation is a quite effective technique for regularizing the neural networks.This paper proposes a novel data augmentation technique based on semantic transposition of the transcriptions via syntax rules for end-to-end Mandarin ASR.Specifically, we first segment the transcriptions based on part-of-speech tags.Then transposition strategies, such as placing the object in front of the subject or swapping the subject and the object, are applied on the segmented sentences.Finally, the acoustic features corresponding to the transposed transcription are reassembled based on the audio-to-text forced-alignment produced by a pre-trained ASR system.The combination of original data and augmented one is used for training a new ASR system.The experiments are conducted on the Transformer[1] and Conformer[2] based ASR.The results show that the proposed method can give consistent performance gain to the system.Augmentation related issues, such as comparison of different strategies and ratios for data combination are also investigated. Zhiyuan Tang, Hengxin Yin, Shuaijiang Zhao, Xiaoning Lei, Xiangang Li |
Interspeech | 2 |
| 2021 | Can We Trust Deep Speech Prior?abstractRecently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) architecture. Compared to conventional approaches that represent clean speech by shallow models such as Gaussians with a low-rank covariance, the new approach employs deep generative models to represent the clean speech, which often provides a better prior. Despite the clear advantage in theory, we argue that deep priors must be used with much caution, since the likelihood produced by a deep generative model does not always coincide with the speech quality. We designed a comprehensive study on this issue and demonstrated that based on deep speech priors, a reasonable SE performance can be achieved, but the results might be suboptimal. A careful analysis showed that this problem is deeply rooted in the disharmony between the flexibility of deep generative models and the nature of the maximum-likelihood (ML) training. Ying Shi 0001, Zhiyuan Tang, Lantian Li, Dong Wang 0013, Jiqing Han 0001 |
SLT | 3 |
| 2020 | ASR-Free Pronunciation AssessmentabstractMost of the pronunciation assessment methods are based on local features derived from automatic speech recognition (ASR), e.g., the Goodness of Pronunciation (GOP) score. In this paper, we investigate an ASR-free scoring approach that is derived from the marginal distribution of raw speech signals. The hypothesis is that even if we have no knowledge of the language (so cannot recognize the phones/words), we can still tell how good a pronunciation is, by comparatively listening to some speech data from the target language. Our analysis shows that this new scoring approach provides an interesting correction for the phone-competition problem of GOP. Experimental results on the ERJ dataset demonstrated that combining the ASR-free score and GOP can achieve better performance than the GOP baseline. Sitong Cheng, Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng |
INTERSPEECH | 4 |
| 2019 | Gaussian-constrained Training for Speaker VerificationabstractNeural models, in particular the d-vector and x-vector architectures, have produced state-of-the-art performance on many speaker verification tasks. However, two potential problems of these neural models deserve more investigation. Firstly, both models suffer from `information leak', which means that some parameters participating in model training will be discarded during inference, i.e, the layers that are used as the classifier. Secondly, these models do not regulate the distribution of the derived speaker vectors. This `unconstrained distribution' may degrade the performance of the subsequent scoring component, e.g., PLDA. This paper proposes a Gaussian-constrained training approach that (1) discards the parametric classifier, and (2) enforces the distribution of the derived speaker vectors to be Gaussian. Our experiments on the VoxCeleb and SITW databases demonstrated that this new training approach produced more representative and regular speaker embeddings, leading to consistent performance improvement. Lantian Li, Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013 |
ICASSP | 2 |
| 2018 | Full-Info Training for Deep Speaker Feature LearningabstractIn recent studies, it has shown that speaker patterns can be learned from very short speech segments (e.g., 0.3 seconds) by a carefully designed convolutional & time-delay deep neural network (CT-DNN) model. By enforcing the model to discriminate the speakers in the training data, frame-level speaker features can be derived from the last hidden layer. In spite of its good performance, a potential problem of the present model is that it involves a parametric classifier, i.e., the last affine layer, which may consume some discriminative knowledge, thus leading to `information leak' for the feature learning. This paper presents a full-info training approach that discards the parametric classifier and enforces all the discriminative knowledge learned by the feature net. Our experiments on the Fisher database demonstrate that this new training scheme can produce more coherent features, leading to consistent and notable performance improvement on the speaker verification task. Lantian Li, Zhiyuan Tang, Dong Wang 0013, Thomas Fang Zheng |
ICASSP | 2 |
| 2018 | Deep Factorization for Speech SignalabstractVarious informative factors mixed in speech signals, leading to great difficulty when decoding any of the factors. An intuitive idea is to factorize each speech frame into individual informative factors, though it turns out to be highly difficult. Recently, we found that speaker traits, which were assumed to be long-term distributional properties, are actually short-time patterns, and can be learned by a carefully designed deep neural network (DNN). This discovery motivated a cascade deep factorization (CDF) framework that will be presented in this paper. The proposed framework infers speech factors in a sequential way, where factors previously inferred are used as conditional variables when inferring other factors. We will show that this approach can effectively factorize speech signals, and using these factors, the original speech spectrum can be recovered with a high accuracy. This factorization and reconstruction approach provides potential values for many speech processing tasks, e.g., speaker recognition and emotion recognition, as will be demonstrated in the paper. Lantian Li, Dong Wang 0013, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Thomas Fang Zheng |
ICASSP | 5 |
| 2018 | Human and Machine Speaker Recognition Based on Short Trivial EventsabstractHuman speech often has events that we will call trivial events, e.g., cough, laugh and sniff. Compared to regular speech, these trivial events are usually short and variable, thus generally regarded as not speaker discriminative and so are largely ignored by present speaker recognition research. However, these trivial events are highly valuable in some particular circumstances such as forensic examination, as they are less subjected to intentional change, so can be used to discover the genuine speaker from disguised speech. In this paper, we collect a trivial event speech database that involves 75 speakers and 6 types of events, and report preliminary speaker recognition results on this database, by both human listeners and machines. Particularly, the deep feature learning technique recently proposed by our group is utilized to analyze and recognize the trivial events, leading to acceptable equal error rates (EERs) ranging from 5% to 15% despite the extremely short durations (0.2-0.5 seconds) of these events. Comparing different types of events, `hmm' seems more speaker discriminative. Xiaofei Kang, Lantian Li, Zhiyuan Tang, Haisheng Dai, Dong Wang 0013 |
ICASSP | 5 |
| 2018 | A 2MHz Constant-Frequency AOT V2 Buck Converter with Adaptive Dead Time Control for Data CentersabstractGaN-based Adaptive-On-Time (AOT) V2buck converter is getting more attention for the applications as data center power supply because of its fast transient response and high efficiency. However, in high-frequency and high step-down Buck converters, the frequency deviation and the dead time power loss become serious problems. In this paper, a GaN-based AOT V2buck converter that adopts a high-accuracy on-time generator and successive feedback regulation with capacitive-scaling dead time control is presented. Designed by using CSMC high voltage 0.25 um process, the controller IC can enable a 100 W (5 V/20 A) buck converter to operate at power switch frequency of 2 MHz and adaptively adjust the dead time under various conditions. With the proposed on-time and dead time control, the frequency variation is only 4% and reverse conduction length of GaN devices is less than 2 ns. The input range of this converter is from 12 V to 60 V and the peak efficiency is 90% at 48 V typical input. Zhiyuan Tang, Shengpeng Tang, Jianxiong Xi, Lenian He, Kexu Sun |
IECON | 1 |
| 2018 | Phonetic Temporal Neural Model for Language IdentificationabstractDeep neural models, particularly the long short-term memory recurrent neural network (LSTM-RNN) model, have shown great potential for language identification (LID). However, the use of phonetic information has been largely overlooked by most existing neural LID methods, although this information has been used very successfully in conventional phonetic LID systems. We present a phonetic temporal neural model for LID, which is an LSTM-RNN LID system that accepts phonetic features produced by a phone-discriminative DNN as the input, rather than raw acoustic features. This new model is similar to traditional phonetic LID methods, but the phonetic knowledge here is much richer: It is at the frame level and involves compacted information of all phones. Our experiments conducted on the Babel database and the AP16-OLR database demonstrate that the temporal phonetic neural approach is very effective, and significantly outperforms existing acoustic neural models. It also outperforms the conventional i-vector approach on short utterances and in noisy conditions. Zhiyuan Tang, Dong Wang 0013, Yixiang Chen 0003, Lantian Li, Andrew Abel |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Memory visualization for gated recurrent neural networks in speech recognitionabstractRecurrent neural networks (RNNs) have shown clear superiority in sequence modeling, particularly the ones with gated units, such as long short-term memory (LSTM) and gated recurrent unit (GRU). However, the dynamic properties behind the remarkable performance remain unclear in many applications, e.g., automatic speech recognition (ASR). This paper employs visualization techniques to study the behavior of LSTM and GRU when performing speech recognition tasks. Our experiments show some interesting patterns in the gated memory, and some of them have inspired simple yet effective modifications on the network structure. We report two of such modifications: (1) lazy cell update in LSTM, and (2) shortcut connections for residual learning. Both modifications lead to more comprehensible and powerful networks. Zhiyuan Tang, Ying Shi 0001, Dong Wang 0013, Yang Feng 0004, Shiyue Zhang 0001 |
ICASSP | 1 |
| 2017 | Deep Speaker Feature Learning for Text-Independent Speaker VerificationabstractRecently deep neural networks (DNNs) have been used to learn speaker features.However, the quality of the learned features is not sufficiently good, so a complex back-end model, either neural or probabilistic, has to be used to address the residual uncertainty when applied to speaker verification, just as with raw features.This paper presents a convolutional timedelay deep neural network structure (CT-DNN) for speaker feature learning.Our experimental results on the Fisher database demonstrated that this CT-DNN can produce highquality speaker features: even with a single feature (0.3 seconds including the context), the EER can be as low as 7.68%.This effectively confirmed that the speaker trait is largely a deterministic short-time property rather than a long-time distributional pattern, and therefore can be extracted from just dozens of frames. Lantian Li, Yixiang Chen 0003, Ying Shi 0001, Zhiyuan Tang, Dong Wang 0013 |
INTERSPEECH | 4 |
| 2017 | Collaborative Joint Training With Multitask Recurrent Model for Speech and Speaker RecognitionabstractAutomatic speech and speaker recognition are traditionally treated as two independent tasks and are studied separately. The human brain in contrast deciphers the linguistic content, and the speaker traits from the speech in a collaborative manner. This key observation motivates the work presented in this paper. A collaborative joint training approach based on multitask recurrent neural network models is proposed, where the output of one task is backpropagated to the other tasks. This is a general framework for learning collaborative tasks and fits well with the goal of joint learning of automatic speech and speaker recognition. Through a comprehensive study, it is shown that the multitask recurrent neural net models deliver improved performance on both automatic speech and speaker recognition tasks as compared to single-task systems. The strength of such multitask collaborative learning is analyzed, and the impact of various training configurations is investigated. Zhiyuan Tang, Lantian Li, Dong Wang 0013, Ravichander Vipperla |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Recurrent neural network training with dark knowledge transferabstractRecurrent neural networks (RNNs), particularly long short-term memory (LSTM), have gained much attention in automatic speech recognition (ASR). Although some successful stories have been reported, training RNNs remains highly challenging, especially with limited training data. Recent research found that a well-trained model can be used as a teacher to train other child models, by using the predictions generated by the teacher model as supervision. This knowledge transfer learning has been employed to train simple neural nets with a complex one, so that the final performance can reach a level that is infeasible to obtain by regular training. In this paper, we employ the knowledge transfer learning approach to train RNNs (precisely LSTM) using a deep neural network (DNN) model as the teacher. This is different from most of the existing research on knowledge transfer learning, since the teacher (DNN) is assumed to be weaker than the child (RNN); however, our experiments on an ASR task showed that it works fairly well: without applying any tricks on the learning scheme, this approach can train RNNs successfully even with limited training data. Zhiyuan Tang, Dong Wang 0013, Zhiyong Zhang 0001 |
ICASSP | 1 |
| 2005 | An innovative scheme for model consistency in collaborative design environmentabstractIn distributed collaborative design environments, product models and virtual studios should keep consistency to create sense of sharing the world and collaboration among physically distributed users. Model consistency is usually achieved by maintaining dynamic shared states. To make those events in physical locations virtually "seen" by the other participants, networks communications are indispensable. Because of limited network bandwidths and CPU resources in distributed and collaborative design environments, to decrease communication traffic is a critical issue. In this paper, we present a "near-sight" mode, which resorts to a method of using AOI, area of interest, to reduce efficiently communications to maintain dynamic shared states. Chenhui Yang, Zhiyuan Tang, Bucai Ye |
CSCWD (1) | 2 |