VLDB 2026 Research / reviewers in the wild / expert
Yuhao Liang
dblp:286/0779
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Novel Single-Stage Single-Phase Transformerless Grid-Connected Photovoltaic InverterabstractThis paper proposes a novel single-stage single-phase transformerless topology based on a buck-boost converter for grid-connected photovoltaic (PV) inverters. The proposed inverter has a wide input voltage range, low total harmonic distortion of current, and effective suppression of common-mode leakage current. The single-input structure ensures that it does not suffer from energy imbalance problems. The new inverter utilizes dead beat control and refines the method of controlling the inverter when the input energy is insufficient. The simulation results verify that the proposed grid-connected PV inverter maintains high grid-connected power quality both during normal operation under conventional conditions and when operating under discontinuous conduction mode (DCM) during energy insufficiency. The simulation models the operating mode of the inverter under abnormal environments, such as load changes and grid voltage transients, which proves that the proposed inverter under dead beat control has a better dynamic performance. Yuhao Liang, Weimin Wu 0001, Houqing Wang, Frede Blaabjerg, Mohamed Orabi |
IECON | 1 |
| 2024 | Secure Source Identification Scheme for Revocable Instruction Sharing in Vehicle PlatoonabstractThe secure transmission of instructions among vehicles in a platoon is one of the most essential needs for a vehicle platoon. Despite the existence of cryptographic methods to securely share instructions, instruction sharing is still subject to forgery, tampering, and denial-of-service attacks. Therefore, it is urgent to find a solution to perform data source identification to filter out irrelevant information (not instructions) while ensuring the authenticity of encrypted instructions is urgent to address. In addition, immediate revocation of credentials is also a crucial requirement for a vehicle platoon when an authorized vehicle member misbehaves. In this paper, we propose the first Secure Source Identification Scheme for Revocable Instruction Sharing (SI-RIS) to securely simultaneously achieve bilateral fine-grained access control, data source identification, immediate vehicle user revocation, and efficient encryption in vehicle platoons. Specifically, our SI-RIS solution supports fine-grained access control for both the sender and receiver over the encrypted instructions. As a result, only authorized correspondents are able to access the commands. Furthermore, upon identification of malicious members in the platoon, our SI-RIS provides an efficient direct vehicle user revocation mechanism capable of immediate revocation credentials without affecting other vehicles. We prove the security of our SI-RIS via rigorous mathematical security proof. Moreover, performance evaluation and comparisons illustrate the feasibility and practicability of SI-RIS for vehicle platoon. Yanan Zhao 0002, Haiyang Yu 0002, Yuhao Liang, Alessandro Brighente, Mauro Conti, Jianfei Sun, Yilong Ren |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | A Sanitizable Access Control With Policy-Protection for Vehicular Social NetworksabstractAs an emerging field of communication, Vehicular Social Networks (VSNs) can reduce traffic congestion while enhancing road safety by sharing data among groups of commuters. In VSNs, Vehicular Cloud Server (VCS) based data sharing technology with encrypted primitives allows local users to outsource encrypted data for reducing the storage burden on the user side and sharing data without location restrictions. However, existing data encryption solutions that have been applied in VSNs environments still encounter weaknesses in efficiency, security, or privacy due to the following problems: (1) lack of effective access policies for flexible authorizing ciphertext to multiple data users; (2) data breaches caused by malicious data publishers; (3) necessity in hiding the private information of receivers. To date, no such solution has been available that securely enables one-to-many user authorization with privacy protection, while greatly resisting malicious data publishers. We propose a Sanitizable Access Control System with Policy-protection (SASP) for VSNs in this paper. Our SASP enables a sanitizer to test and sanitize encrypted data to defend against malicious data publishers, ensuring that the plaintext can only be recovered if an authorized user has a valid key. Furthermore, in our SASP system, the access policy is separated into attribute names and attribute values. Wherein, the attribute values contain a lot of private information, which is hidden in the ciphertext to guarantee data users’ privacy. Rigorous security analysis and performance evaluations demonstrate the practicality of SASP for VSNs. Yanan Zhao 0002, Haiyang Yu 0002, Yuhao Liang, Mauro Conti, Wael Bazzi, Yilong Ren |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | BA-MoE: Boundary-Aware Mixture-of-Experts Adapter for Code-Switching Speech RecognitionabstractMixture-of-experts based models, which use language experts to extract language-specific representations effectively, have been well applied in code-switching automatic speech recognition. However, there is still substantial space to improve as similar pronunciation across languages may result in ineffective multi-language modeling and inaccurate language boundary estimation. To eliminate these drawbacks, we propose a cross-layer language adapter and a boundary-aware training method, namely Boundary-Aware Mixture-of-Experts (BA-MoE). Specifically, we introduce language-specific adapters to separate language-specific representations and a unified gating layer to fuse representations within each encoder layer. Second, we compute language adaptation loss of the mean output of each language-specific adapter to improve the adapter module’s language-specific representation learning. Besides, we utilize a boundary-aware predictor to learn boundary representations for dealing with language boundary confusion. Our approach achieves significant performance improvement, reducing the mixture error rate by 16.55% compared to the baseline on the ASRU 2019 Mandarin-English code-switching challenge dataset. Peikun Chen, Fan Yu 0002, Yuhao Liang, Hongfei Xue, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001 |
ASRU | 3 |
| 2023 | Sa-Paraformer: Non-Autoregressive End-To-End Speaker-Attributed ASRabstractJoint modeling of multi-speaker ASR and speaker diarization has recently shown promising results in speaker-attributed automatic speech recognition (SA-ASR). Although being able to obtain state-of-the-art (SOTA) performance, most of the studies are based on an autoregressive (AR) decoder which generates tokens one-by-one and results in a large real-time factor (RTF). To speed up inference, we introduce a recently proposed non-autoregressive model Paraformer as an acoustic model in the SA-ASR model. Paraformer uses a single-step decoder to enable parallel generation, obtaining comparable performance to the SOTA AR transformer models. Besides, we propose a speaker-filling strategy to reduce speaker identification errors and adopt an inter-CTC strategy to enhance the encoder’s ability in acoustic modeling. Experiments on the AliMeeting corpus show that our model outperforms the cascaded SA-ASR model by a 6.1% relative speaker-dependent character error rate (SD-CER) reduction on the test set. Moreover, our model achieves a comparable SD-CER of 34.8% with only 1/10 RTF compared with the SOTA joint AR SA-ASR model. Yangze Li, Fan Yu 0002, Yuhao Liang, Mohan Shi, Zhihao Du, Shiliang Zhang, Lei Xie 0001 |
ASRU | 3 |
| 2023 | The Second Multi-Channel Multi-Party Meeting Transcription Challenge (M2MeT 2.0): A Benchmark for Speaker-Attributed ASRabstractWith the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SAASR), which directly addresses the practical and challenging problem of “who spoke what at when” at typical meeting scenario. We particularly established two sub-tracks. The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. The open training condition sub-track, which allows for the use of all available data and models without limitation. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR. Yuhao Liang, Mohan Shi, Fan Yu 0002, Yangze Li, Shiliang Zhang, Zhihao Du, Qian Chen 0003, Lei Xie 0001, Yanmin Qian, Jian Wu 0027, Zhuo Chen 0006, Kong-Aik Lee, Zhijie Yan, Hui Bu |
ASRU | 1 |
| 2023 | BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR
Yuhao Liang, Fan Yu 0002, Yangze Li, Shiliang Zhang, Qian Chen 0003, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2023 | Identity-Based Broadcast Signcryption Scheme for Vehicular Platoon CommunicationabstractVehicular platooning is emerging as a promising method that can enhance road utilization, alleviate traffic congestion, and even decrease energy expenditure by shortening the distance between vehicles in the platoon. In platooning, a platoon leader (PL) is required to communicate with platoon members (PMs) to issue instructions. In this way, the PMs are simply expected to follow the command and enjoy their free time. However, when the same instruction is sent to different PMs in the form of single-hop unicast, the ciphertext will go up as the quantity of PMs in the platoon resulting in a dramatic increase in transmission time. Moreover, transmission instructions without security and authentication guarantees are easily intercepted, forged, or deleted. Therefore, in this article, we propose broadcast signcryption scheme for platoon communication (BSPC), a broadcast signcryption scheme for platoon communication based on identity. In our BSPC, with only one signcryption, the PL can generate a public verifiable ciphertext with fixed-length for PMs, while employing broadcast functionality to deliver the ciphertext to multiple PMs at once. In such a manner, the confidentiality, integrity, and authentication of transmitted commands are ensured, such that malicious vehicles cannot manipulate data and unauthorized vehicles have no way to access it. Besides, the instructions' transmit time is greatly reduced. Our BSPC proposal is proved to be secure and unforgeable through rigorous analysis. Furthermore, simulation results demonstrate our BSPC is practical and efficient for platoon communication. Yanan Zhao 0002, Yuhao Liang, Haiyang Yu 0002, Yilong Ren |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | MFCCA:Multi-Frame Cross-Channel Attention for Multi-Speaker ASR in Multi-Party Meeting ScenarioabstractRecently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations between sequences of different channels or exploiting fine-grained channel-wise information effectively at each time step. Considering the delay of microphone array receiving sound, we propose a multi-frame cross-channel attention, which models cross-channel information between adjacent frames to exploit the complementarity of both frame-wise and channel-wise knowledge. Besides, we also propose a multi-layer convolutional mechanism to fuse the multi -channel output and a channel masking strategy to combat the channel number mismatch problem between training and inference. Experiments on the AliMeeting, a real-world corpus, reveal that our proposed model outperforms single-channel model by 31.7% and 37.0% CER reduction on Eval and Test sets. Moreover, with comparable model parameters and training data, our proposed model achieves a new SOTA performance on the AliMeeting corpus, as compared with the top ranking systems in the ICASSP2022 M2MeT challenge, a recently held multi-channel multi-speaker ASR challenge. Fan Yu 0002, Shiliang Zhang, Yuhao Liang, Zhihao Du, Yuxiao Lin, Lei Xie 0001 |
SLT | 4 |
| 2021 | Boundary and Context Aware Training for CIF-Based Non-Autoregressive End-to-End ASRabstractContinuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with competitive performance compared with other NAR methods. However, such an alignment learning strategy may suffer from an erroneous acoustic boundary estimation, severely hindering the convergence speed as well as the system performance. In this paper, we propose a boundary and context aware training approach for CIF based NAR models. Firstly, the connectionist temporal classification (CTC) spike information is utilized to guide the learning of acoustic boundaries in the CIF. Besides, an additional contextual decoder is introduced behind the CIF decoder, aiming to capture the linguistic dependencies within a sentence. Finally, we adopt a recently proposed Conformer architecture to improve the capacity of acoustic modeling. Experiments on the open-source Mandarin AISHELL-1 corpus show that the proposed method achieves a comparable character error rates (CERs) of 4.9% with only 1/24 latency compared with a state-of-the-art autoregressive (AR) Conformer model. Futhermore, when evaluating on an internal 7500 hours Mandarin corpus, our model still outperforms other NAR methods and even reaches the AR Conformer model on a challenging real-world noisy test set. Fan Yu 0002, Haoneng Luo, Yuhao Liang, Zhuoyuan Yao, Lei Xie 0001, Yingying Gao, Leijing Hou, Shilei Zhang |
ASRU | 4 |
| 2021 | The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and MethodsabstractThe variety of accents has posed a big challenge to speech recognition. The Accented English Speech Recognition Challenge (AESRC2020) is designed for providing a common testbed and promoting accent-related research. Two tracks are set in the challenge – English accent recognition (track 1) and accented English speech recognition (track 2). A set of 160 hours of accented English speech collected from 8 countries is released with labels as the training set. Another 20 hours of speech without labels is later released as the test set, including two unseen accents from another two countries used to test the model generalization ability in track 2. We also provide baseline systems for the participants. This paper first reviews the released dataset, track setups, baselines and then summarizes the challenge results and major techniques used in the submissions. Xian Shi, Fan Yu 0002, Yizhou Lu, Yuhao Liang, Qiangze Feng, Daliang Wang, Yanmin Qian, Lei Xie 0001 |
ICASSP | 4 |