EDBT 2026 Demo / reviewers in the wild / expert
Yu Liao
dblp:61/6559
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 57% Multimedia analysis and retrieval · 43% | |
| Artificial intelligence
1 paper |
Vision and language · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Integrated circuit design · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language pretraining |
0.8 | 1 | 2024 | Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method · ACM Multimedia 2024 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.8 | 1 | 2024 | Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method · ACM Multimedia 2024 |
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval |
0.8 | 1 | 2024 | Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval Method · ACM Multimedia 2024 |
Visualization and visual analytics › information visualization
composite visualization |
0.7 | 1 | 2023 | Revisiting the Design Patterns of Composite Visualizations · IEEE Trans. Vis. Comput. Graph. 2023 |
Visualization and visual analytics › visualization design
design patterns |
0.7 | 1 | 2023 | Revisiting the Design Patterns of Composite Visualizations · IEEE Trans. Vis. Comput. Graph. 2023 |
Visualization and visual analytics
visualization design |
0.7 | 1 | 2023 | Revisiting the Design Patterns of Composite Visualizations · IEEE Trans. Vis. Comput. Graph. 2023 |
Integrated circuit design
analog and mixed-signal circuits |
0.2 | 1 | 2013 | A 65 mW fully integrated UHF-band CMMB tuner in 65 nm CMOS process · Sci. China Inf. Sci. 2013 |
Integrated circuit design › semiconductor device fabrication
CMOS technology |
0.0 | 1 | 2013 | A 65 mW fully integrated UHF-band CMMB tuner in 65 nm CMOS process · Sci. China Inf. Sci. 2013 |
Methods — techniques the papers use, named apart from their topics
multimodal interaction · 1.5key local selection · 1.5corpus analysis · 0.7co-occurrence analysis · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Out-of-Domain Noise for Unsupervised Domain Adaptation in Speech EnhancementabstractWhen there’s a mismatch between the training and test domains, supervised speech enhancement (SE) models trained on synthetic paired noisy-clean data often struggle in real-world scenarios, highlighting the industry’s strong demand for unsupervised training and domain adaptation methods. In this study, we introduce PHA-ReMixIT, a novel approach for leveraging out-of-domain (OOD) noise signals to enhance unsupervised domain adaptation in SE. Our method builds upon the state-of-the-art ReMixIT by introducing a paired unsupervised remixing technique, which augments the diversity of target domain training data with OOD noise signals. We further propose a heterogeneous noise invariant training to align the OOD augmented noisy mixtures with their paired heterogeneous counterparts, encouraging the model to output cleaner speech. Additionally, an adaptive focal weighting mechanism is also introduced to dynamically emphasize the data importance of both in-domain and OOD noisy mixtures during model adaptation. Experiments on CHiME-7 unsupervised domain adaptation for conversational speech enhancement (UDASE) task demonstrate that PHA-ReMixIT significantly outperforms the ReMixIT baseline, boosting SE performance on both real and synthesized test sets. Yu Liao, Haixin Guan, Yanhua Long |
ICASSP | 1 |
| 2024 | GeneFormer: Learned Gene Compression using Transformer-Based Context ModelingabstractThe development of gene sequencing technology sparks an explosive growth of gene data. Thus, the storage of gene data has become an important issue. Recently, researchers begin to investigate deep learning-based gene data compression, which outperforms general traditional methods. In this paper, we propose a transformer-based gene compression method named GeneFormer. Specifically, we first introduce a modified transformer encoder with latent array to eliminate the dependency of the nucleotide sequence. Then, we design a multi-level-grouping method to accelerate and improve the compression process. Experimental results on real-world datasets show that our method achieves significantly better compression ratio compared with state-of-the-art method, and the decoding speed is significantly faster than all existing learning-based gene compression methods. We will release our code on github once the paper is accepted. Zhanbei Cui, Tongda Xu, Yu Liao, Yan Wang 0105 |
ICASSP | 4 |
| 2024 | Selection and Reconstruction of Key Locals: A Novel Specific Domain Image-Text Retrieval MethodabstractIn recent years, Vision-Language Pre-training (VLP) models have demonstrated rich prior knowledge for multimodal alignment, prompting investigations into their application in Specific Domain Image-Text Retrieval(SDITR) such as Text-Image Person Re-identification (TIReID) and Remote Sensing Image-Text Retrieval (RSITR). Due to the unique data characteristics in specific scenarios, the primary challenge is to leverage discriminative fine-grained local information for improved mapping of images and text into a shared space. Current approaches interact with all multimodal local features for alignment, implicitly focusing on discriminative local information to distinguish data differences, which may bring noise and uncertainty. Furthermore, their VLP feature extractors like CLIP often focus on instance-level representations, potentially reducing the discriminability of fine-grained local features. To alleviate these issues, we propose an Explicit Key Local information Selection and Reconstruction Framework (EKLSR), which explicitly selects key local information to enhance feature representation. Specifically, we introduce a Key Local information Selection and Fusion (KLSF) that utilizes hidden knowledge from the VLP model to select interpretably and fuse key local information. Secondly, we employ Key Local segment Reconstruction (KLR) based on multimodal interaction to reconstruct the key local segments of images (text), significantly enriching their discriminative information and enhancing both inter-modal and intra-modal interaction alignment. To demonstrate the effectiveness of our approach, we conducted experiments on five datasets across TIReID and RSITR. Notably, our EKLSR model achieves state-of-the-art performance on two RSITR datasets. Yu Liao, Rui Yang 0038, Jianwei Tao, Bai Liu 0002, Zhipeng Hu, Shuang Wang 0001, Zeng Zhao |
ACM Multimedia | 1 |
| 2024 | Continual learning for cross-modal image-text retrieval based on domain-selective attention
Rui Yang 0038, Shuang Wang 0001, Yu Gu 0015, Jihui Wang, Yingzhi Sun, Yu Liao, Licheng Jiao |
Pattern Recognit. | 7 |
| 2024 | Multi-View Feature Fusion and Visual Prompt for Remote Sensing Image CaptioningabstractRemote sensing image (RSI) captioning is a vision-language multimodal task concentrating on both image comprehension and sentence generation. Several studies suggest that encoder–decoder-based methods have achieved success in RSI captioning. However, existing encoder–decoder-based methods may not fully explore image representations for RSI captioning and suffer from a lack of additional prompt information for sentence generation. In this article, a novel multi-view feature fusion and prompt (MVP)-based model is proposed to obtain better RSI representations and enhance language model performance in RSI captioning. Specifically, we design an attention-based feature fusion module to dynamically fuse multi-view visual features, which are extracted from the fine-tuned vision-language pretraining (VLP) model and the vision-task pretraining (VP) model. Then, a flexible visual prefix mapping module is proposed to transform images into visual prefixes, providing semantic information for the subsequent sentence generation. Finally, a BERT-based caption generator is applied to generate accurate descriptions based on the fused visual features and the visual prefixes, which are both outputs from our designed modules. Extensive experiments are conducted on three well-known benchmark datasets, demonstrating that our method achieves state-of-the-art (SOTA) performance. The relevant code is available athttps://github.com/QiaoLing-Lin/MVP. Shuang Wang 0001, Qiaoling Lin, Xiutiao Ye, Yu Liao, Dou Quan, ZhongQian Jin, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge DistillationabstractWith the increasing development of remote sensing (RS) technology, remote sensing cross-modal image-text retrieval (RSCMITR) task has gradually attracted wide attention. At present, the large-scale pre-training model is brilliant in the field of natural images cross-modal retrieval, but the current RSCMITR models do not focus on it, resulting in less retrieval performance improvement. This paper proposes a lightweight network structure based on large-scale pre-training model and knowledge distillation, designing a lightweight model based on separable convolution and text convolution. Knowledge distillation technology is used to make the Light model learn the hidden knowledge of large-scale model CLIP-RS, which realizes fast and accurate retrieval. The proposed method achieves state-of-the-art performance on four commonly used RSCMITR datasets. Yu Liao, Rui Yang 0038, Hantong Xing, Dou Quan, Shuang Wang 0001, Biao Hou |
IGARSS | 1 |
| 2023 | Revisiting the Design Patterns of Composite VisualizationsabstractComposite visualization is a popular design strategy that represents complex datasets by integrating multiple visualizations in a meaningful and aesthetic layout, such as juxtaposition, overlay, and nesting. With this strategy, numerous novel designs have been proposed in visualization publications to accomplish various visual analytic tasks. However, there is a lack of understanding of design patterns of composite visualization, thus failing to provide holistic design space and concrete examples for practical use. In this article, we opted to revisit the composite visualizations in IEEE VIS publications and answered what and how visualizations of different types are composed together. To achieve this, we first constructed a corpus of composite visualizations from the publications and analyzed common practices, such as the pattern distributions and co-occurrence of visualization types. From the analysis, we obtained insights into different design patterns on the utilities and their potential pros and cons. Furthermore, we discussed usage scenarios of our taxonomy and corpus and how future research on visualization composition can be conducted on the basis of this study. Dazhen Deng, Weiwei Cui 0001, Xiyu Meng, Mengye Xu, Yu Liao, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | A Transformer-Based Cross-Modal Image-Text Retrieval Method using Feature Decoupling and ReconstructionabstractWith the increasing application of remote sensing technology, the task of cross-modal retrieval of remote sensing images (CMRRS) has gradually attracted widespread attention. Ex-isting methods often completely map the features of different modalities to a shared space and do not decouple between the modal-invariant information and modal-heterogeneous in-formation, which leads to redundant information in feature mapping and usually gets sub-optimal retrieval performance. This paper proposes a Transformer-based CMRRS method using feature decoupling and reconstruction (TBFDR) to solve this problem. TBFDR achieves state-of-the-art performance in remote sensing image-text retrieval task on Sydney-Captions dataset. Yingzhi Sun, Yu Liao, Rui Yang 0038, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 3 |
| 2021 | Cross-Modal Feature Fusion Retrieval for Remote Sensing Image-Voice RetrievalabstractWith the increasing popularity of remote sensing technology applications, some emergency scenarios require rapid retrieval of remote sensing images, such as earthquake rescue, etc. Due to the high efficiency of voice input, researchers have focused on cross-modal remote sensing image-voice retrieval methods. However, these methods have two major drawbacks: speech input lacks discrimination and the intra-modal semantic information is under used. To address these drawbacks, we propose a novel cross-modal feature fusion retrieval model. Our model provides a more optimized cross-modal common feature space than previous models and thus optimizes the retrieval performance. First, our model adds the extra textual keyword information to the audio feature for remote sensing image retrieval. Second, it introduces inter-modality adversarial learning and intra-modality semantic discrimination into the remote sensing image-voice retrieval task. We conducted experiments on two datasets modified from the UCM-Captions dataset and the Remote Sensing Image Caption Dataset. The experimental results show that our model outperforms state-of-the-art models in this task. Rui Yang 0038, Yu Gu 0015, Yu Liao, Yingzhi Sun, Shuang Wang 0001, Biao Hou, Licheng Jiao |
IGARSS | 3 |
| 2013 | A 65 mW fully integrated UHF-band CMMB tuner in 65 nm CMOS process
Junhua Liu 0001, Chen Li 0014, Long Chen 0009, Congyin Shi, Xuankai Weng, Yixiao Wang 0001, Yu Liao, Le Ye, Huailin Liao, Ru Huang 0001 |
Sci. China Inf. Sci. | 8 |
| 2004 | A Simple parallel dual code decoding algorithm for convolutional codes with high throughput and low latencyabstractA highly parallel maximum a posteriori (MAP) decoding algorithm is proposed for high code rate, R = (n-1)/ n , ordinary (nonpunctured) convolutional codes using trellises of reciprocal dual convolutional codes. The advantages of this approach include a substantial reduction of decoding latency and decoding complexity, and a substantial increase of decoding throughput. Applying the proposed parallel decoding algorithm to a class of serial concatenation codes that consist of high rate ordinary convolutional codes, over additive white Gaussian noise (AWGN) channels, good bit error rate (BER) and block error rate performance, comparable to that of turbo codes and low density parity check (LDPC) codes, can be obtained with smaller overall decoding complexity than that of LDPC codes. Yu Liao, John C. Kieffer |
ISIT | 1 |