Pei Yu

dblp:33/674 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-8848-0584ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 FairScene: Learning Class-Disentangled 2D/3D Representations for Semantic Scene Completion
abstract
Semantic Scene Completion (SSC) aims to predict the semantic occupancy of each voxel within a 3D scene using sensor data, a critical task for autonomous driving and robotics. Despite recent progress, camera-based SSC remains challenging due to various difficulties, including voxel class imbalance, occlusion, and depth ambiguity. This paper introduces FairScene, a novel approach that learns class-disentangled 2D/3D representations to improve SSC. By ensuring balanced representations across classes, FairScene mitigates the dominance of majority classes and promotes fairer voxel categorization. Additionally, FairScene explicitly models spatial dependencies between different classes through a novel inter-class occupancy reasoning mechanism. Such explicit modeling helps alleviate occlusion and depth ambiguities in SSC. To address the scarcity of SSC training data, we propose OccMix, a novel augmentation strategy that generalizes MixUp from 2D to 2.5D and 3D metric spaces while maintaining geometric consistency. Extensive quantitative and qualitative experiments demonstrate that FairScene out-performs prior methods on both the SemanticKITTI and SSCBench-KITTI-360 benchmarks. The code is available at https://github.com/DianJJ/FairScene.
Dian Jia, Pei Yu, Wei Tang 0016
WACV2
2026 Modeling and Learning Multiple Hypotheses for Monocular 3D Object Detection
abstract
Detecting objects in 3D space using a monocular image is inherently a highly ill-posed problem: multiple plausible 3D bounding boxes can explain the same 2D observation of an object. Existing approaches typically follow a single-point prediction paradigm, failing to capture this multimodal nature and often regressing to an implausible mean solution. This paper introduces MonoMH, a novel multi-hypothesis framework for monocular 3D object detection. By explicitly modeling and learning the multimodal distribution of plausible 3D object configurations, MonoMH not only significantly improves detection performance but also provides richer information to support downstream decision-making. MonoMH introduces three key innovations: (1) a novel multi-hypothesis predictor that leverages spatially diverse features across different windows within an RoI to generate a rich variety of hypotheses without increasing model complexity; (2) a new multi-hypothesis learning approach that derives diverse and relevant hypotheses from single-modal ground truth by integrating uncertainty modeling with Best-of-Many learning; and (3) a hypothesis filtering mechanism that enhances detection capability by dynamically retaining a variable number of plausible hypotheses based on each object’s uncertainty. Experimental results demonstrate the effectiveness of our approach. Notably, MonoMH achieves 29.12/20.88/17.93 Car AP3D(easy/mod./hard) on the KITTI test set, significantly outperforming previous state-of-the-art methods. The code can be found at https://github.com/HyeonjeongPark37/MonoMH.
Hyeonjeong Park, Peixi Xiong, Pei Yu, Wei Tang 0016
WACV3
2026 Enhancing interpretation of clinical disease-associated copy number variations from multiple sequencing strategies with CNVSeeker
abstract
MOTIVATION: DNA copy number variations (CNVs) exert a profound impact on major genetic disorders in humans. Although multiple sequencing technologies have become the first line of molecular diagnosis for CNVs, existing tools are unable to resolve the pathogenicity of CNVs directly from raw sequencing data. RESULTS: We developed CNVSeeker, a one-stop and easy-to-use pipeline that provides comprehensive analysis from raw sequencing data to variant interpretation reports, and supports multiple types of sequencing data including short-read data such as whole genome sequencing data and whole exome sequencing data, and long-read sequencing data from Pacific Biosciences HiFi platform or Oxford Nanopore Technologies platform. Through extensive benchmarking, CNVSeeker demonstrated comparable enhancement over the state-of-the-art methods for CNV calling. Moreover, CNVSeeker enables significantly precise variant classification with an accuracy of ∼87%. By applying CNVSeeker to 1946 individuals with autism spectrum disorder (ASD), a total of 133 ASD-associated CNVs in 122 patients were identified, yielding a diagnostic yield of ∼6.3%. Additionally, we have also provided a user-friendly webserver for intuitive visualization of results. This study highlights the potential of CNVSeeker to benefit clinicians and geneticists with limited bioinformatic skill by aiding them interpret CNVs directly from various types of raw sequencing data for auxiliary disease diagnosis. AVAILABILITY AND IMPLEMENTATION: The web server is freely available at https://genemed.tech/cnvseeker and the open-source code can be found at https://github.com/lovelycatZ/CNVSeeker.
Xudong Xiang, Xinxin Mao, Tengfei Luo, Chenbin Liu, Bozhao Li, Pei Yu, Dai Wu, Yixiao Zhu, Guihu Zhao, Jinchen Li
Bioinform.6
2026 Clairaudience: a lightweight attentional residual neural network with data augmentation and feature fusion for underwater acoustic target recognition
abstract
Abstract Marine engineering has boomed and many deep learning-based methods have been proposed for underwater acoustic target recognition. However, most of these methods are dedicated to develop more complex convolutional neural networks to achieve better performance. This results in these models being unable to be deployed to low cost and miniaturized automatic underwater vehicles. A novel lightweight attentional residual neural network with data augmentation and feature fusion is proposed in this paper. Mel Frequency Cepstral Coefficient (MFCC), delta-MFCC and delta–delta MFCC features are extracted in the time dimension for fusion to obtain the fusion feature. The SpecAugment data augmentation is also used to enhance the randomness and diversity of features by masking in time and frequency dimensions randomly. Shuffle attention in the residual blocks is introduced to enhance the representation of features. The lightweight model is evaluated and compared by using several metrics on ShipsEar and DeepShip datasets. The proposed lightweight model only requires 1.628 M parameters for the trained model. This work shows that the proposed method requires small memory storage, while it achieved comparative performance.
Jing Li 0057, Yucheng Han, Lili Zhang 0014, Wei Wei 0053, Pei Yu, Hongxin Tan
Comput. J.7
2026 LSOD-DETR: a lightweight small object detection model based on real-time detection transformer
Lili Zhang 0014, Wenshuo Han, Ke Zhang 0033, Ruiyang Xiao, Jing Li 0057, Wei Wei 0053, Pei Yu, Hongxin Tan
J. Supercomput.9
2025 Improving Multilingual Sign Language Translation with Automatically Clustered Language Family Information
abstract
Sign Language Translation (SLT) bridges the communication gap between deaf and hearing individuals by converting sign language videos into spoken language texts. While most SLT research has focused on bilingual translation models, the recent surge in interest has led to the exploration of Multilingual Sign Language Translation (MSLT). However, MSLT presents unique challenges due to the diversity of sign languages across nations. This diversity can lead to cross-linguistic conflicts and hinder translation accuracy. To use the similarity of actions and semantics between sign languages to alleviate conflict, we propose a novel approach that leverages sign language families to improve MSLT performance. Sign languages were clustered into families automatically based on their Language distribution in the MSLT network. We compare the results of our proposed family clustering method with the analysis conducted by sign language linguists and then train dedicated translation models for each family in the many-to-one translation scenario. Our experiments on the SP-10 dataset demonstrate that our approach can achieve a balance between translation accuracy and computational cost by regulating the number of language families.
Ruiquan Zhang, Pei Yu, Yidong Chen 0001
COLING3
2025 Learning Partonomic 3D Reconstruction from Image Collections
abstract
Reconstructing the 3D shape of an object from a single-view image is a fundamental task in computer vision. Recent advances in differentiable rendering have enabled 3D reconstruction from image collections using only 2D annotations. However, these methods mainly focus on whole-object reconstruction and overlook object partonomy, which is essential for intelligent agents interacting with physical environments. This paper aims at learning partonomic 3D reconstruction from collections of images with only 2D annotations. Our goal is not only to reconstruct the shape of an object from a single-view image but also to decompose the shape into meaningful semantic parts. To handle the expanded solution space and frequent part occlusions in single-view images, we introduce a novel approach that represents, parses, and learns the structural compositionality of 3D objects. This approach comprises: (1) a compact and expressive compositional representation of object geometry, achieved through disentangled modeling of large shape variations, constituent parts, and detailed part deformations as multi-granularity neural fields; (2) a part transformer that recovers precise partonomic geometry and handles occlusions, through effective part-to-pixel grounding and part-to-part relational modeling; and (3) a 2D-supervised learning method that jointly learns the compositional representation and part transformer, by bridging object shape and parts, image synthesis, and differentiable rendering. Extensive experiments on ShapeNetPart, PartNet, and CUB-200-2011 demonstrate the effectiveness of our approach on both overall and partonomic reconstruction. Code, models, and data are avaliable at https://github.com/XiaoqianRuan1/Partonomic_Reconstruction.
Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park, Peixi Xiong, Wei Tang 0016
CVPR2
2025 TSD-DETR: A lightweight real-time detection transformer of traffic sign detection for long-range perception of autonomous driving
Lili Zhang 0014, Yucheng Han, Jing Li 0057, Wei Wei 0053, Hongxin Tan, Pei Yu, Ke Zhang 0033
Eng. Appl. Artif. Intell.7
2025 Improving few-shot object detection via mislabeling mitigation
Pei Yu, Gaocai Wang, Man Wu, Lili Wen
Neurocomputing1
2025 Underwater acoustic target recognition based on multi-scale feature and CRDNet
Jing Li 0057, Lili Zhang 0014, Wei Wei 0053, Pei Yu, Hongxin Tan
J. Supercomput.7
2025 Traffic environmental protection edge computing: a monitoring algorithm and system of truck black smoke emission in complex scene
Lili Zhang 0014, Yucheng Han, Ke Zhang 0033, Jing Li 0057, Wei Wei 0053, Hongxin Tan, Pei Yu
J. Supercomput.8
2025 Driving risks from light pollution: an improved YOLOv8 detection network for high beam vehicle image recognition
Lili Zhang 0014, Ke Zhang 0033, Wei Wei 0053, Jing Li 0057, Hongxin Tan, Pei Yu, Yucheng Han
J. Supercomput.7
2024 Improving Non-Autoregressive Sign Language Translation with Random Ordering Progressive Prediction Pretraining
abstract
Recently, the Non-AutoRegressive (NAR) decoding mechanism, effectively reducing the inference latency of text generation, has been applied to Sign Language Translation (SLT). Typically, the current best NAR SLT model using a Curriculum-based Non-autoregressive Decoder (CND) outperforms AutoRegressive (AR) baselines in speed and performance. Although it has been proven that AutoRegressive Pre-trained Language Models (AR-PLMs) further boost the performance of AR SLT models, combining NAR Pretrained Language Models (NAR-PLMs) with NAR SLT model remains challenge due to (1) existing NAR-PLMs’ inability to model token dependencies between decoder layers, crucial for NAR SLT models using CND; (2) the modality gap between the decoder’s inputs of the NAR-PLMs and NAR SLT models. To address these, we propose a Random Ordering Progressive Prediction Pre-training task for NAR SLT models using CND, enabling the decoder to predict target sequences in diverse orderings and enhancing the modeling of target token dependencies between layers. Moreover, we propose a CTC-enhanced Soft Copy method to incorporate target-side information in the decoder’s inputs, alleviating the modality gap. Experimental results on PHOENIX-2014T and CSL-Daily demonstrate that our model consistently outperforms all strong baselines and achieves competitive performance with AR SLT models equipped with AR-PLMs.
Pei Yu, Changhao Lai, Biao Fu, Yidong Chen 0001
ECAI1
2024 An Explicit Multi-Modal Fusion Method for Sign Language Translation
abstract
Sign Language Translation (SLT) aims to convert sign language videos into corresponding spoken text sequences. However, the inherent modality gap between sign language video and text hinders the development of SLT. Motivated by the linguistic consistency between gloss1and text, we propose EMF-SLT, an Explicit Multi-modal Fusion method for Sign Language Translation to mitigate the modality gap with the help of gloss. Specifically, EMF-SLT first leverages a vector quantizer and a fusion module to align and fuse sign language and gloss features, respectively, resulting in more informative multi-modal features for the decoder. Then, a multi-task mutual learning framework is introduced to regularize the output predictions from different modalities, which ensures the consistency of outputs across modalities and encourages different modalities to learn from each other. Experiments on two SLT benchmarks and further analyses show that our method achieves significant improvements over the baselines and effectively alleviates the modality gap.
Biao Fu, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP3
2023 A Token-Level Contrastive Framework for Sign Language Translation
abstract
Sign Language Translation (SLT) is a promising technology to bridge the communication gap between the deaf and the hearing people. Recently, researchers have adopted Neural Machine Translation (NMT) methods, which usually require large-scale corpus for training, to achieve SLT. However, the publicly available SLT corpus is very limited, which causes the collapse of the token representations and the inaccuracy of the generated tokens. To alleviate this issue, we propose Con-SLT, a novel token-level Contrastive learning framework for Sign Language Translation , which learns effective token representations by incorporating token-level contrastive learning into the SLT decoding process. Concretely, ConSLT treats each token and its counterpart generated by different dropout masks as positive pairs during decoding, and then randomly samples K tokens in the vocabulary that are not in the current sentence to construct negative examples. We conduct comprehensive experiments on two benchmarks (PHOENIX14T and CSL-Daily) for both end-to-end and cascaded settings. The experimental results demonstrate that ConSLT can achieve better translation quality than the strong baselines1.
Biao Fu, Peigen Ye, Pei Yu, Xiaodong Shi, Yidong Chen 0001
ICASSP4
2023 Efficient Sign Language Translation with a Curriculum-based Non-autoregressive Decoder
abstract
Most existing studies on Sign Language Translation (SLT) employ AutoRegressive Decoding Mechanism (AR-DM) to generate target sentences. However, the main disadvantage of the AR-DM is high inference latency. To address this problem, we introduce Non-AutoRegressive Decoding Mechanism (NAR-DM) into SLT, which generates the whole sentence at once. Meanwhile, to improve its decoding ability, we integrate the advantages of curriculum learning and NAR-DM and propose a Curriculum-based NAR Decoder (CND). Specifically, the lower layers of the CND are expected to predict simple tokens that could be predicted correctly using source-side information solely. Meanwhile, the upper layers could predict complex tokens based on the lower layers' predictions. Therefore, our CND significantly reduces the model's inference latency while maintaining its competitive performance. Moreover, to further boost the performance of our CND, we propose a mutual learning framework, containing two decoders, i.e., an AR decoder and our CND. We jointly train the two decoders and minimize the KL divergence between their outputs, which enables our CND to learn the forward sequential knowledge from the strengthened AR decoder. Experimental results on PHOENIX2014T and CSL-Daily demonstrate that our model consistently outperforms all competitive baselines and achieves 7.92/8.02× speed-up compared to the AR SLT model respectively. Our source code is available at https://github.com/yp20000921/CND.
Pei Yu, Biao Fu, Yidong Chen 0001
IJCAI1
2023 Exploring Effective Inter-Encoder Semantic Interaction for Document-Level Relation Extraction
abstract
In document-level relation extraction (RE), the models are required to correctly predict implicit relations in documents via relational reasoning. To this end, many graph-based methods have been proposed for this task. Despite their success, these methods still suffer from several drawbacks: 1) their interaction between document encoder and graph encoder is usually unidirectional and insufficient; 2) their graph encoders often fail to capture the global context of nodes in document graph. In this paper, we propose a document-level RE model with a Graph-Transformer Network (GTN). The GTN includes two core sublayers: 1) the graph-attention sublayer that simultaneously models global and local contexts of nodes in the document graph; 2) the cross-attention sublayer, enabling GTN to capture the non-entity clue information from the document encoder. Furthermore, we introduce two auxiliary training tasks to enhance the bidirectional semantic interaction between the document encoder and GTN: 1) the graph node reconstruction that can effectively train our cross-attention sublayer to enhance the semantic transition from the document encoder to GTN; 2) the structure-aware adversarial knowledge distillation, by which we can effectively transfer the structural information of GTN to the document encoder. Experimental results on four benchmark datasets prove the effectiveness of our model. Our source code is available at https://github.com/DeepLearnXMU/DocRE-BSI.
Zijun Min, Jinsong Su, Pei Yu, Ante Wang, Yidong Chen 0001
IJCAI4
2022 A Unified Transferable Model for ML-Enhanced DBMS
Ziniu Wu, Pei Yu, Peilun Yang, Yuxing Han 0002, Yaliang Li, Defu Lian, Kai Zeng 0002, Jingren Zhou 0001
CIDR2
2022 Should All Proposals Be Treated Equally in Object Detection?
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Pei Yu, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos
ECCV (25)6
2022 SC-UDA: Style and Content Gaps aware Unsupervised Domain Adaptation for Object Detection
abstract
Current state-of-the-art object detectors can have significant performance drop when deployed in the wild due to domain gaps with training data. Unsupervised Domain Adaptation (UDA) is a promising approach to adapt detectors for new domains/environments without any expensive label cost. Previous mainstream UDA works for object detection usually focused on image-level and/or feature-level adaptation by using adversarial learning methods. In this work, we show that such adversarial-based methods can only reduce domain style gap, but cannot address the domain content gap that is also important for object detectors. To overcome this limitation, we propose the SC-UDA framework to concurrently reduce both gaps: We propose fine-grained domain style transfer to reduce the style gaps with finer image details preserved for detecting small objects; Then we leverage the pseudo label-based self-training to reduce content gaps; To address pseudo label error accumulation during self-training, novel optimizations are proposed, including uncertainty-based pseudo labeling and imbalanced mini-batch sampling strategy. Experiment results show that our approach consistently outperforms prior state-of-the-art methods (up to 8.6%, 2.7% and 2.5% mAP on three UDA benchmarks).
Fuxun Yu, Di Wang 0003, Yinpeng Chen, Nikolaos Karianakis, Pei Yu, Dimitrios Lymberopoulos, Sidi Lu, Weisong Shi, Xiang Chen 0010
WACV6
2018 Deeply Learned Compositional Models for Human Pose Estimation
Wei Tang 0016, Pei Yu, Ying Wu 0001
ECCV (3)2
2017 Human Action Segmentation using 3D Fully Convolutional Network
Pei Yu, Jiang Wang 0001, Ying Wu 0001
BMVC1
2017 Towards a Unified Compositional Model for Visual Pattern Modeling
abstract
Compositional models represent visual patterns as hierarchies of meaningful and reusable parts. They are attractive to vision modeling due to their ability to decompose complex patterns into simpler ones and resolve the lowlevel ambiguities in high-level image interpretations. However, current compositional models separate structure and part discovery from parameter estimation, which generally leads to suboptimal learning and fitting of the model. Moreover, the commonly adopted latent structural learning is not scalable for deep architectures. To address these difficult issues for compositional models, this paper quests for a unified framework for compositional pattern modeling, inference and learning. Represented by And-or graphs (AOGs), it jointly models the compositional structure, parts, features, and composition/sub-configuration relationships. We show that the inference algorithm of the proposed framework is equivalent to a feed-forward network. Thus, all the parameters can be learned efficiently via the highly-scalable back-propagation (BP) in an end-to-end fashion. We validate the model via the task of handwritten digit recognition. By visualizing the processes of bottom-up composition and top-down parsing, we show that our model is fully interpretable, being able to learn the hierarchical compositions from visual primitives to visual patterns at increasingly higher levels. We apply this new compositional model to natural scene character recognition and generic object detection. Experimental results have demonstrated its effectiveness.
Wei Tang 0016, Pei Yu, Jiahuan Zhou, Ying Wu 0001
ICCV2
2017 Efficient Online Local Metric Adaptation via Negative Samples for Person Re-identification
abstract
Many existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs is very demanding in practice. In contrast to these methods, this paper advocates a different paradigm: part of the learning can be performed online but with nominal costs, so as to achieve online metric adaptation for different input probes. A major challenge here is that no positive training pairs are available for the probe anymore. By only exploiting easily-available negative samples, we propose a novel solution to achieve local metric adaptation effectively and efficiently. For each probe at the test time, it learns a strictly positive semi-definite dedicated local metric. Comparing to offline global metric learning, its computational cost is negligible. The insight of this new method is that the local hard negative samples can actually provide tight constraints to fine tune the metric locally. This new local metric adaptation method is generally applicable, as it can be used on top of any global metric to enhance its performance. In addition, this paper gives in-depth theoretical analysis and justification of the new method. We prove that our new method guarantees the reduction of the classification error asymptotically, and prove that it actually learns the optimal local metric to best approximate the asymptotic case by a finite number of training data. Extensive experiments and comparative studies on almost all major benchmarks (VIPeR, QMUL GRID, CUHK Campus, CUHK03 and Market-1501) have confirmed the effectiveness and superiority of our method.
Jiahuan Zhou, Pei Yu, Wei Tang 0016, Ying Wu 0001
ICCV2
2016 Learning Reconstruction-Based Remote Gaze Estimation
abstract
It is a challenging problem to accurately estimate gazes from low-resolution eye images that do not provide fine and detailed features for eyes. Existing methods attempt to establish the mapping between the visual appearance space to the gaze space. Different from the direct regression approach, the reconstruction-based approach represents appearance and gaze via local linear reconstruction in their own spaces. A common treatment is to use the same local reconstruction in the two spaces, i.e., the reconstruction weights in the appearance space are transferred to the gaze space for gaze reconstruction. However, this questionable treatment is taken for granted but has never been justified, leading to significant errors in gaze estimation. This paper is focused on the study of this fundamental issue. It shows that the distance metric in the appearance space needs to be adjusted, before the same reconstruction can be used. A novel method is proposed to learn the metric, such that the affinity structure of the appearance space under this new metric is as close as possible to the affinity structure of the gaze space under the normal Euclidean metric. Furthermore, the local affinity structure invariance is utilized to further regularize the solution to the reconstruction weights, so as to obtain a more robust and accurate solution. Effectiveness of the proposed method is validated and demonstrated through extensive experiments on different subjects.
Pei Yu, Jiahuan Zhou, Ying Wu 0001
CVPR1
2012 Inter-cell cooperation aided dynamic base station switching for energy efficient cellular networks
abstract
In this paper, a cell zooming based dynamic base stations (BS) switching scheme is conceived for energy saving in cellular networks. A cell zooming algorithm is developed for leveraging the BS's coverage with inter-cell cooperation and downlink power control. Meanwhile, we design a feasible working strategy for the cell zooming scheme to maximizing the energy saving. Simulations are provided to demonstrate the advantages of the proposed scheme.
Pei Yu, Qinghai Yang, Fenglin Fu, Kyung Sup Kwak
APCC1
2008 Globally exponentially attractive sets of the family of Lorenz systems
Xiaoxin Liao, Yuli Fu 0001, Shengli Xie 0001, Pei Yu
Sci. China Ser. F Inf. Sci.4
2006 Absolute Stability of Hopfield Neural Network
Xiaoxin Liao, Fei Xu 0007, Pei Yu
ISNN (1)3
2003 A matching pursuit technique for computing the simplest normal forms of vector fields
Pei Yu, Yuan Yuan 0008
J. Symb. Comput.1