Xiyan Liu

dblp:116/7319 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HSA: An Efficient Sparse CNN Accelerator Based on Kernel-Aware Hybrid Pruning
abstract
The deployment of large-scale convolutional neural networks (CNNs) on hardware often results in reduced execution efficiency and increased hardware overhead. Prior works have addressed this challenge by employing pruning techniques to reduce parameters and computational loads. Among these techniques, fine-grainedN:Mpruning methods achieve high sparsity rates but at the cost of generating substantial indexing overhead and potential workload imbalance. In contrast, pattern-based pruning methods reduce indexing complexity and achieve workload balancing but suffer from low sparsity. To overcome these limitations, this article proposes a kernel-aware hybrid pruning (KAHP) method that simultaneously attains high sparsity, workload balancing, and reduced indexing overhead. Moreover, to accelerate the inference of the models pruned by the KAHP method, we design an efficient accelerator based on the systolic array architecture. Experimental results demonstrate that, compared to theN:Mpruning method, the KAHP method achieves workload balancing and reduces the number of indices by up to 88.7% on ResNet-20. Compared to pattern-based pruning methods, the KAHP method achieves up to$31.8\times $sparsity. The proposed accelerator achieves up to 93.18 GOPS/W power efficiency and 0.721 GOPS/DSP computation efficiency.
Xiyan Liu, Qiang Liu 0011
IEEE Trans. Very Large Scale Integr. Syst.1
2025 LDMapNet-U: An End-to-End System for City-Scale Lane-Level Map Updating
abstract
An up-to-date city-scale lane-level map is an indispensable infrastructure and a key enabling technology for ensuring the safety and user experience of autonomous driving systems. In industrial scenarios, reliance on manual annotation for map updates creates a critical bottleneck. Lane-level updates require precise change information and must ensure consistency with adjacent data while adhering to strict standards. Traditional methods utilize a three-stage approach -- construction, change detection, and updating -- which often necessitates manual verification due to accuracy limitations. This results in labor-intensive processes and hampers timely updates. To address these challenges, we propose LDMapNet-U, which implements a new end-to-end paradigm for city-scale lane-level map updating. By reconceptualizing the update task as an end-to-end map generation process grounded in historical map data, we introduce a paradigm shift in map updating that simultaneously generates vectorized maps and change information. To achieve this, a Prior-Map Encoding (PME) module is introduced to effectively encode historical maps, serving as a critical reference for detecting changes. Additionally, we incorporate a novel Instance Change Prediction (ICP) module that learns to predict associations with historical maps. Consequently, LDMapNet-U simultaneously achieves vectorized map element generation and change detection. To demonstrate the superiority and effectiveness of LDMapNet-U, extensive experiments are conducted using large-scale real-world datasets. In addition, LDMapNet-U has been successfully deployed in production at Baidu Maps since April 2024, supporting lane-level map updating for over 360 cities and significantly shortening the update cycle from quarterly to weekly, thereby enhancing the timeliness and accuracy of lane-level map. The nationwide, high-frequency city-scale lane-level map has been instrumental in the development of the lane-level navigation product serving hundreds of millions of users, while also integrating into the autonomous driving systems of several leading vehicle companies.
Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0088, Chenting Gong, Xiao Tan 0001, Jizhou Huang, Mengmeng Yang 0001, Diange Yang
KDD (1)3
2024 CDUMA: An Adaptive Approach for Mitigating Confounder for MCQA
abstract
Multiple-choice question answering (MCQA) requires the model to select the correct answer from a set of candidate options when given a passage and a question. Previous research has achieved promising results with the assistance of Pre-trained Language Models(PrLMs). However, it has been observed that these models heavily rely on text similarity for inference. In this study, we approach the MCQA task from a causal perspective, treating text similarity as a confounder and establishing a structural causal model(SCM) for the MCQA task. To mitigate the impact of this confounder, we propose a novel Causal Dual Multi-head Co-Attention (CDUMA) model. CDUMA reduces the influence of text similarity between questions and options on model inference and enhances the model’s generalization ability during routine testing. Experimental results on two widely used datasets demonstrate the effectiveness of our approach.
Chenji Lu, Ge Bai, Xiyan Liu, Ruifang Liu
ICASSP5
2024 CausalME: Balancing bi-modalities in Visual Question Answering
abstract
Mitigating linguistic bias and attaining modal equilibrium in Visual Question Answering (VQA) tasks constitute a pivotal concern. Previous work has mainly focused on data augmentation or a uni-modal approach, which is insufficient to fully utilize bi-modal information. In this work, we propose a new causal modal equilibrium framework CausalME, addressing the issue from a causal perspective. CausalME utilizes a question-only branch to capture the linguistic bias of the textual modality and mitigate its causal effect with a newly designed adaptive paradigm. Additionally, CausalME employs counterfactual generation to enhance the causal effect of visual modality. By optimizing the objective function of the entire VQA model, CausalME balances the causal effects of bi-modalities and explicitly guides the model to align text and image information. We conducted extensive experiments and the results show that CausalME brings significant improvements and achieves competitive performance on the bias-sensitive VQA-CP v2 dataset.
Chenji Lu, Ge Bai, Xiyan Liu, Zerong Zeng, Ruifang Liu
ICASSP5
2024 DuMapNet: An End-to-End Vectorization System for City-Scale Lane-Level Map Generation
abstract
Generating city-scale lane-level maps faces significant challenges due to the intricate urban environments, such as blurred or absent lane markings. Additionally, a standard lane-level map requires a comprehensive organization of lane groupings, encompassing lane direction, style, boundary, and topology, yet has not been thoroughly examined in prior research. These obstacles result in labor-intensive human annotation and high maintenance costs. This paper overcomes these limitations and presents an industrial-grade solution named DuMapNet that outputs standardized, vectorized map elements and their topology in an end-to-end paradigm. To this end, we propose a group-wise lane prediction (GLP) system that outputs vectorized results of lane groups by meticulously tailoring a transformer-based network. Meanwhile, to enhance generalization in challenging scenarios, such as road wear and occlusions, as well as to improve global consistency, a contextual prompts encoder (CPE) module is proposed, which leverages the predicted results of spatial neighborhoods as contextual information. Extensive experiments conducted on large-scale real-world datasets demonstrate the superiority and effectiveness of DuMapNet. Additionally, DuMapNet has already been deployed in production at Baidu Maps since June 2023, supporting lane-level map generation tasks for over 360 cities while bringing a 95% reduction in costs. This demonstrates that DuMapNet serves as a practical and cost-effective industrial solution for city-scale lane-level map generation.
Deguo Xia, Weiming Zhang 0006, Xiyan Liu, Wei Zhang 0114, Chenting Gong, Jizhou Huang, Mengmeng Yang 0001, Diange Yang
KDD3
2024 Few-shot video object segmentation with prototype evolution
Binjie Mao, Xiyan Liu, Linsu Shi, Jiazhong Yu, Shiming Xiang
Neural Comput. Appl.2
2024 Singularity Analysis and Solutions for the Origami Transmission Mechanism of Fast-Moving Untethered Insect-Scale Robot
abstract
Designing insect-scale robots with high mobility is becoming an essential challenge in the field of robotics research. Among the methods for fabricating the transmission mechanism of the insect-scale robot, the smart composite microstructure (SCM) method is getting more and more attention. This method can construct compact and functional miniature origami mechanisms through planarized fabrication and folding assembly processes. Our previous work has proposed an untethered robot S$^{\text{2}}$worm equipped with a novel 2-DoF origami transmission mechanism. The S$^{\text{2}}$worm is fabricated through SCM and holds a top speed of 27.4 cm/s. In this work, we propose a novel strategy for designing the insect-scale robot with high mobility, that is, applying Grassmann–Cayley Algebra (GCA) to avoid the singularity of the transmission mechanism. The experimental results prove that the singularity of the previous work has been solved. The new robot prototype S$^{\text{2}}$worm-G weighs 4.71 g, scales 4.0 cm, achieves a top speed of 75.0 cm/s and a relative speed of 18.8 bodylength/s. To the best of our knowledge, the 2-DoF origami transmission mechanism is the first parallel mechanism designed for the insect-scale robot and the singularity of the mechanism is found and solved here. The experimental results prove that the refined S$^{\text{2}}$worm-G robot is one of the best insect-scale robots for its size, mass, and mobility.
Yide Liu, Bo Feng 0006, Tianlun Cheng, Xiyan Liu, Shaoxing Qu, Wei Yang 0058
IEEE Trans. Robotics5
2023 Bilateral Memory Consolidation for Continual Learning
abstract
Humans are proficient at continuously acquiring and integrating new knowledge. By contrast, deep models forget catastrophically, especially when tackling highly long task sequences. Inspired by the way our brains constantly rewrite and consolidate past recollections, we propose a novel Bilateral Memory Consolidation (BiMeCo) framework that focuses on enhancing memory interaction capabilities. Specifically, BiMeCo explicitly decouples model parameters into short-term memory module and long-term memory module, responsible for representation ability of the model and generalization over all learned tasks, respectively. BiMeCo encourages dynamic interactions between two memory modules by knowledge distillation and momentum-based updating for forming generic knowledge to prevent forgetting. The proposed BiMeCo is parameter-efficient and can be integrated into existing methods seamlessly. Extensive experiments on challenging benchmarks show that BiMeCo significantly improves the performance of existing continual learning methods. For example, combined with the state-of-the-art method CwD [55], BiMeCo brings in significant gains of around 2% to 6% while using 2x fewer parameters on CIFAR-100 under ResNet-18.
Xing Nie, Shixiong Xu, Xiyan Liu, Gaofeng Meng, Chunlei Huo, Shiming Xiang
CVPR3
2023 An Interpretable Model Using Evidence Information for Multi-Hop Question Answering Over Long Texts
abstract
Machine Reading Comprehension (MRC) is a challenging task in natural language understanding, especially multi-hop question answering (QA) in long texts. One of the challenges in multi-hop QA requires models to produce interpretable answers based on evidence that is selected from a given long text. Based on the Retriever-Reader architecture, existing work tackles this problem by using different methods to exploit various evidence information. To better use evidence information, we propose a loss function considering answer groups, which improves the reasoning ability of the reader in the Retriever-Reader architecture. Besides, we introduce the relevance constraint factor containing evidence information to improve the reader’s ability of locating key sentences. Evaluated on the HotpotQA dataset, the proposed methods achieve improvement, demonstrating the effectiveness of our methods and the importance of evidence information.
Yanyi Chen, Ruifang Liu, Xiyan Liu, Yidong Shi, Ge Bai
ICASSP3
2023 Narrow Down Before Selection: A Dynamic Exclusion Model for Multiple-Choice QA
abstract
Multiple-choice question answering (MCQA) is a challenging task that requires selecting the correct answer from a set of options based on a given question. There is a trend to use pre-trained encoder-decoder models to solve MCQA. Previous works concentrate on the decoder and adopt the generated text to enhance model performance. However, few studies have optimized the use of encoders for the characteristics of MCQA. In this work, we propose a dynamic exclusion model for MCQA named ExcMC, which mimics human thinking in selection. It dynamically eliminates several incorrect options to optimize the encoder usage. ExcMC outperforms existing comparable works on two widely-used MCQA datasets, demonstrating the effectiveness of our model.
Xiyan Liu, Yidong Shi, Ruifang Liu, Ge Bai, Yanyi Chen
ICASSP1
2023 Progressive generation of 3D point clouds with hierarchical consistency
Xiyan Liu, Jizhou Huang, Deguo Xia, Jianzhong Yang
Pattern Recognit.2
2022 DuARUS: Automatic Geo-object Change Detection with Street-view Imagery for Updating Road Database at Baidu Maps
abstract
As the core foundation of web mapping, each geographic object (geo-object), such as a traffic sign, plays a vital role in navigation and intelligent driving. Determining how to obtain the latest high-precision geo-object information is a classic topic in updating road databases. Benefiting from the cost-effective attribute and availability of the positioning equipment and camera, the vision-based update pattern is becoming increasingly popular in the industry. Generally speaking, the road database update mainly includes three phases: geo-object recognition, localization, and change detection. Previous change detection strategies are mainly performed by comparing the historical road information (i.e., geo-object type and position) with the new geographic data of geo-objects collected from the street-view imagery. However, limited by the localization precision of the positioning equipment and the discriminative power of the vanilla differential-based method, the accuracy, recall, and efficiency of previous systems for geo-object change detection are greatly impaired. In addition, the artificially prescribed production standards make the geo-object position in the map data deviate from its position in the real world, as well as some geo-objects do not need to be updated (e.g., temporary speed limit), which further yields many false-positive detections and significantly increases the labor costs of existing systems. To address these challenges, we propose a novel framework called DuARUS for automatic geo-object change detection with street-view imagery. In this paper, we mainly focus on automatic geo-object localization and change detection. Specifically, for geo-object localization, we propose a two-stage, integrated localization algorithm based on image matching and monocular depth estimation. Furthermore, to achieve automatic change detection, vision-based representation learning and scene understanding strategies are introduced to build a large-scale geo-object semantic map, which can provide sufficient multimodal information support for change detection. Based on such artful modeling, we recast the complicated, labor-based change detection problem as a vanilla binary classification task, which is a robust and efficient strategy that contributes to resolving this problem. By combining these operations, we construct an industrial-grade, fully automatic production system for road database updates. Extensive experiments conducted on large-scale, real-world datasets from Baidu Maps demonstrate the superiority and effectiveness of the system. Moreover, this system has already been deployed in production at Baidu Maps since July 2020, handling 96% of automatic road database updates. DuARUS improves the annual update mileage from millions to tens of millions, and it achieves weekly updates.
Deguo Xia, Jizhou Huang, Jianzhong Yang, Xiyan Liu, Haifeng Wang 0001
CIKM4
2022 DuTraffic: Live Traffic Condition Prediction with Trajectory Data and Street Views at Baidu Maps
abstract
The task of live traffic condition prediction, which aims at predicting live traffic conditions (i.e., fast, slow, and congested) based on traffic information on roads, plays a vital role in intelligent transportation systems, such as navigation, route planning, and ride-hailing services. Existing solutions have adopted aggregated trajectory data to generate traffic estimates, which inevitably suffer from GPS drift caused by cluttered urban road scenarios. In addition, the trajectory information alone is insufficient to provide evidence for sudden traffic situations and perception of street-wise elements. To alleviate these problems, in this paper, we present DuTraffic, which is a robust and production-ready solution for live traffic condition prediction by taking both trajectory data and street views into account. Specifically, the vision-based detection and segmentation modules are developed to forecast traffic flow by using street views. Then, we propose a spatial-temporal-based module, TRST-Net, to learn the latent trajectory representation. Finally, a bilinear model is introduced to mix these two representations and then predicts live traffic conditions with trajectory data and street views in a mutually complementary manner. The task is recast as a multi-task learning problem, which could benefit from the strong representation of latent space manifold modeling. Extensive experiments conducted on large-scale, real-world datasets from Baidu Maps demonstrate the superiority and effectiveness of DuTraffic. In addition, DuTraffic has already been deployed in production at Baidu Maps since December 2020, handling tens of millions of requests every day. This demonstrates that DuTraffic is a practical and robust industrial solution for live traffic condition prediction.
Deguo Xia, Xiyan Liu, Wei Zhang 0088, Chengzhou Li, Weiming Zhang 0006, Jizhou Huang, Haifeng Wang 0001
CIKM2
2022 Components Regulated Generation of Handwritten Chinese Text-lines in Arbitrary Length
abstract
Generating readable images of handwritten Chinese text-lines is very challenging due to complicated topological structures in Chinese. To address this problem, we propose a components regulated model named HCT-GAN to generate the entire lines of Chinese handwriting from text-line labels. Specifically, HCT-GAN is designed as a CGAN-based architecture that additionally integrates a Chinese text encoder (CTE), a sequence recognition module(SRM), and a spatial perception module (SPM). Compared with the one-hot embedding, CTE learns the latent content representation by reusing the structure and component embedding shared among the Chinese characters. SRM provides sequence-level constraints to the generated images. SPM can adaptively constrain the spatial correlation between the generated components, which facilitates the modeling of characters with complicated topological structures. Benefiting from such artful modeling, our model suffices to generate images of handwritten Chinese text-lines in arbitrary length. Extensive experimental results demonstrate that our model achieves state-of-the-art performance in handwritten Chinese lines generation.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICPR2
2022 Decoupled Representation Learning for Character Glyph Synthesis
abstract
Character glyph synthesis is still an open challenging problem, which involves two related aspects,i.e., font style transfer and content consistency. In this paper, we propose a novel model named FontGAN, which integrates the character structure stylization, de-stylization and texture transfer into a unified framework. Specifically, we decouple character images into style representation and content representation, which offers fine-grained control of these two types of variables, thus improving the quality of the generated results. To effectively capture the style information, a style consistency module (SCM) is introduced. Technically, SCM exploits category-guided Kullback-Leibler divergence to explicitly model the style representation into different prior distributions. In this way, our model is capable of implementing transformations between multiple domains in one framework. In addition, we propose content prior module (CPM) to provide content prior for the model to guide the content encoding process and alleviates the problem of stroke deficiency during structure de-stylization. Benefiting from the idea of decoupling and regrouping, our FontGAN suffices to achieve many-to-many translation tasks for glyph structure. Experimental results demonstrate that the proposed FontGAN achieves the state-of-the-art performance in character glyph synthesis.
Xiyan Liu, Gaofeng Meng, Jianlong Chang, Ruiguang Hu, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.1
2021 Handwritten Text Generation via Disentangled Representations
abstract
Automatically generating handwritten text images is a challenging task due to the diverse handwriting styles and the irregular writing in natural scenes. In this paper, we propose an effective generative model called HTG-GAN to synthesize handwritten text images from latent prior. Unlike single-character synthesis, our method is capable of generating images of sequence characters with arbitrary length, which pays more attention to the structural relationship between characters. We model the structural relationship as the style representation to avoid explicitly modeling the stroke layout. Specifically, the text image is disentangled into style representation and content representation, where the style representation is mapped into Gaussian distribution and the content representation is embedded using character index. In this way, our model can generate new handwritten text images with specified contents and various styles to perform data augmentation, thereby boosting handwritten text recognition (HTR). Experimental results show that our method achieves state-of-the-art performance in handwritten text generation.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
IEEE Signal Process. Lett.1
2020 Geometric rectification of document images using adversarial gated unwarping network
Xiyan Liu, Gaofeng Meng, Bin Fan 0001, Shiming Xiang, Chunhong Pan
Pattern Recognit.1
2019 Scene text detection and recognition with advances in deep learning: a survey
Xiyan Liu, Gaofeng Meng, Chunhong Pan
Int. J. Document Anal. Recognit.1
2018 Semantic Image Synthesis via Conditional Cycle-Generative Adversarial Networks
abstract
Traditional approaches for semantic image synthesis mainly focus on text descriptions while ignoring the related structures and attributes in the original images. Therefore, some critical information, e.g., the style, backgrounds, objects shapes and pose, is missed in the generated images. In this paper, we propose a novel framework called Conditional Cycle-Generative Adversarial Network (CCGAN) to address this issue. Our model can generate photo-realistic images conditioned on the given text descriptions, while maintaining the attributes of the original images. The framework mainly consists of two coupled conditional adversarial networks, which are able to learn a desirable image mapping that can keep the structures and attributes in the images. We introduce a conditional cycle consistency loss to prevent the contradiction between two generators. This loss allows the generated images to retain most of the features of the original image, so as to improve the stability of network training. Moreover, benefiting from the mechanism of circular training, the proposed networks can learn the semantic information of the text much accurately. Experiments on Caltech-UCSD Bird dataset and Oxford-102 flower dataset demonstrate that the proposed method significantly outperforms the existing methods in terms of image details reconstruction and semantic information expression.
Xiyan Liu, Gaofeng Meng, Shiming Xiang, Chunhong Pan
ICPR1
2012 Unsupervised Feature Selection for Multi-cluster Data via Smooth Distributed Score
Furui Liu, Xiyan Liu
ICIC (3)2