Qingqing Yan

dblp:133/6875 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Representation learning for skeleton-based action recognition from a causal perspective
Mengxian Hu, Liuyi Wang, Chenpeng Yao, Qingqing Yan
Knowl. Based Syst.4
2026 A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation
Liuyi Wang, Yongrui Qin, Haojie Dai, Jingwei Yang 0002, Qingqing Yan
IEEE Trans Autom. Sci. Eng.9
2026 Efficient Text-Driven Motion Generation via Latent Consistency Training
abstract
Text-driven human motion generation based on diffusion strategies establishes a reliable foundation for multimodal applications in human–computer interactions. However, existing advances face significant efficiency challenges due to the substantial computational overhead of iteratively solving for nonlinear reverse diffusion trajectories during the inference phase. To this end, we propose the motion latent consistency training (MLCT) framework, which precomputes reverse diffusion trajectories from raw data in the training phase and enables few-step or single-step inference via self-consistency constraints in the inference phase. Specifically, a motion autoencoder with quantization constraints is first proposed for constructing concise and bounded solution distributions for motion diffusion processes. Subsequently, a classifier-free guidance (CFG) format is constructed via an additional unconditional loss function to accomplish the precomputation of conditional diffusion trajectories in the training phase. Finally, a clustering guidance module based on theK-nearest-neighbor (KNN) algorithm is developed for the chain-conduction optimization mechanism of self-consistency constraints, which provides additional references of solution distributions at a small query cost. By combining these enhancements, we achieve stable and consistency training in nonpixel modality and latent representation spaces. Benchmark experiments demonstrate that our method significantly outperforms traditional consistency distillation methods with reduced training cost and enhances the consistency model to perform comparably to state-of-the-art (SOTA) models with lower inference costs.
Mengxian Hu, Qingqing Yan, Shu Li 0005
IEEE Trans. Syst. Man Cybern. Syst.4
2025 A multilevel attention network with sub-instructions for continuous vision-and-language navigation
Liuyi Wang, Qingqing Yan
Appl. Intell.4
2024 Multimodal Evolutionary Encoder for Continuous Vision-Language Navigation
abstract
Can multimodal encoder evolve when facing increasingly tough circumstances? Our work investigates this possibility in the context of continuous vision-language navigation (continuous VLN), which aims to navigate robots under linguistic supervision and visual feedback. We propose a multimodal evolutionary encoder (MEE) comprising a unified multimodal encoder architecture and an evolutionary pre-training strategy. The unified multimodal encoder unifies rich modalities, including depth and sub-instruction, to enhance the solid understanding of environments and tasks. It also effectively utilizes monocular observation, reducing the reliance on panoramic vision. The evolutionary pre-training strategy exposes the encoder to increasingly unfamiliar data domains and difficult objectives. The multi-stage adaption helps the encoder establish robust intra- and inter-modality connections and improve its generalization to unfamiliar environments. To achieve such evolution, we collect a large-scale multi-stage dataset with specialized objectives, addressing the absence of suitable continuous VLN pre-training. Evaluation on VLN-CE demonstrates the superiority of MEE over other direct action-predicting methods. Furthermore, we deploy MEE in real scenes using self-developed service robots, showcasing its effectiveness and potential for real-world applications. Our code and dataset are available at https://github.com/RavenKiller/MEE.
Liuyi Wang, Shu Li 0005, Qingqing Yan
IROS5
2024 SNF-Feat: Semantic-Guided Negative-Sample-Free Representation Learning for Local Feature Extraction
abstract
Local feature extraction constitutes a foundational module crucial for numerous downstream tasks of computer vision. Its primary challenge lies in the generation of discriminative feature representations. Prior methodologies have employed contrastive learning within their pipelines, yet have encountered limitations stemming from inherent conflicts within their training data, including the ambiguity of negative samples and the distortion of positive samples. In this study, we propose a semantic-guided negative-sample-free method for local feature learning, denoted as SNF-Feat. Our framework entails dense patch-level representation learning without reliance on negative samples, aiming to ensure that descriptors derived from transformed views of the same local area exhibit predictive capability towards each other. To assess the impact of positive sample distortion, we harness high-level semantic information to derive point-wise loss weights. Furthermore, we establish a self-supervised feature learning paradigm that extends our utilization of datasets. Experimental results demonstrate the superior performance of our method across a range of typical datasets and tasks in comparison to state-of-the-art approaches.
Qingqing Yan, Mengxian Hu
IROS2
2024 PASTS: Progress-aware spatio-temporal transformer speaker for vision-and-language navigation
Liuyi Wang, Shu Li 0005, Qingqing Yan, Huiyi Chen
Eng. Appl. Artif. Intell.5
2024 Learning Depth Representation From RGB-D Videos by Time-Aware Contrastive Pre-Training
abstract
Existing end-to-end depth representation in embodied AI is often task-specific and lacks the benefits of emerging pre-training paradigm due to limited datasets and training techniques for RGB-D videos. To address the challenge of obtaining robust and generalized depth representation for embodied AI, we introduce a unified RGB-D video dataset (UniRGBD) and a novel time-aware contrastive (TAC) pre-training approach. UniRGBD addresses the scarcity of large-scale depth pre-training datasets by providing a comprehensive collection of data from diverse sources in a unified format, enabling convenient data loading and accommodating various data domains. We also design an RGB-Depth alignment evaluation procedure and introduce a novel Near-K accuracy metric to assess the scene understanding capability of the depth encoder. Then, the TAC pre-training approach fills the gap in depth pre-training methods suitable for RGB-D videos by leveraging the intrinsic similarity between temporally proximate frames. TAC incorporates a soft label design that acts as valid label noise, enhancing the depth semantic extraction and promoting diverse and generalized knowledge acquisition. Furthermore, the adjustments in perspective between temporally proximate frames facilitate the extraction of invariant and comprehensive features, enhancing the robustness of the learned depth representation. Additionally, the inclusion of temporal information stabilizes training gradients and enables spatio-temporal depth perception. Comprehensive evaluation of RGB-Depth alignment demonstrates the superiority of our approach over state-of-the-art methods. We also conduct uncertainty analysis and a novel zero-shot experiment to validate the robustness and generalization of the TAC approach. Moreover, our TAC pre-training demonstrates significant performance improvements in various embodied AI tasks, providing compelling evidence of its efficacy across diverse domains.
Liuyi Wang, Ronghao Dang, Shu Li 0005, Qingqing Yan
IEEE Trans. Circuits Syst. Video Technol.5
2024 DR-Block: Convolutional Dense Reparameterization for CNN Generalization Free Improvement
abstract
As an emerging and popular technique for boosting CNNs, structural reparameterization (SR) decouples the training and inference structures to alter the training dynamics and achieve cost-free improvement of a given network. Existing SR methods often prioritize network expressiveness enhancement but have yet to investigate approaches to mitigate significant bias and non-robustness of model prediction due to over-reliance on training data distribution and image noise. To this end, inspired by the effective strength of implicit regularization on the problem, this paper introduces an extra balanced implicit regularization mechanism into SR techniques to enhance the generalization of a given network for the first time. Specifically, we propose a novel SR module named DR-Block, which is used to complicate each convolutional layer of a given CNN during training. It draws on the advantages of deep matrix factorization with the regularization effect and further improves singular value dynamics by introducing batch normalization and dense connections to alleviate network degradation. At inference time, DR-Block can be equivalently reparameterized back into a single convolution for deployment. Furthermore, we empirically demonstrate the role of each design in DR-Block and explicitly reveal its inherent mechanism, which lies in enhancing the movement of large singular values while countering the attenuation of small ones. This helps enhance the interpretability of SR techniques. Experiments illustrate that DR-Block is an impressive alternative for a regular convolution layer of any structure and outperforms the existing SR methods in improving mainstream network architectures on various visual tasks. The code is available athttps://github.com/qyan0131/DRBlock.
Qingqing Yan, Shu Li 0005, Mengxian Hu
IEEE Trans. Circuits Syst. Video Technol.1
2024 NDNet: Spacewise Multiscale Representation Learning via Neighbor Decoupling for Real-Time Driving Scene Parsing
abstract
As a safety-critical application, autonomous driving requires high-quality semantic segmentation and real-time performance for deployment. Existing method commonly suffers from information loss and massive computational burden due to high-resolution input-output and multiscale learning scheme, which runs counter to the real-time requirements. In contrast to channelwise information modeling commonly adopted by modern networks, in this article, we propose a novel real-time driving scene parsing framework named NDNet from a novel perspective of spacewise neighbor decoupling (ND) and neighbor coupling (NC). We first define and implement the reversible operations called ND and NC, which realize lossless resolution conversion for complementary thumbnails sampling and collation to facilitate spatial modeling. Based on ND and NC, we further propose three modules, namely, local capturer and global dependence builder (LCGB), spacewise multiscale feature extractor (SMFE), and high-resolution semantic generator (HSG), which form the whole pipeline of NDNet. The LCGB serves as a stem block to preprocess the large-scale input for fast but lossless resolution reduction and extract initial features with global context. Then the SMFE is used for dense feature extraction and can obtain rich multiscale features in spatial dimension with less computational overhead. As for high-resolution semantic output, the HSG is designed for fast resolution reconstruction and adaptive semantic confusion amending. Experiments show the superiority of the proposed method. NDNet achieves the state-of-the-art performance on the Cityscapes dataset which reports 76.47% mIoU at 240+ frames/s and 78.8% mIoU at 150+ frames/s on the benchmark. Codes are available at https://github.com/LiShuTJ/NDNet.
Shu Li 0005, Qingqing Yan, Deming Wang
IEEE Trans. Neural Networks Learn. Syst.2
2023 FDLNet: Boosting Real-time Semantic Segmentation by Image-size Convolution via Frequency Domain Learning
abstract
This paper proposes a novel real-time semantic segmentation network via frequency domain learning, called FDLNet, which revisits the segmentation task from two critical perspectives: spatial structure description and multilevel feature fusion. We first devise an image-size convolution (IS-Conv) as a global frequency-domain learning operator to capture long-range dependency in a single shot. To model spatial structure information, we construct the global structure representation path (GSRP) based on IS-Conv, which learns a unified edge-region representation with affordable complexity. For efficient and lightweight multi-level feature fusion, we propose the factorized stereoscopic attention (FSA) module, which alleviates semantic confusion and reduces feature redundancy by introducing level-wise attention before channel and spatial attention. Combining the above modules, we propose a concise semantic segmentation framework named FDLNet. We experimentally demonstrate the effectiveness and superiority of the proposed method. FDLNet achieves state-of-the-art performance on the Cityscapes, which reports 76.32% mIoU at 150+ FPS and 79.0% mIoU at 41+ FPS. The code is available at https://github.com/qyan0131/FDLNet.
Qingqing Yan, Shu Li 0005, Ming Liu 0001
ICRA1
2023 PUW-Feat: A Progressive and Unified Method for Weakly Supervised Local Feature Learning
abstract
Local feature extraction is a fundamental module in computer vision. Weakly supervised learning methods have more convenience of collecting datasets, but they have still not get proper trade-off among training costs, accuracy and speed. In this paper, we propose PUW-Feat, a Progressive and Unified Weakly supervised learnable local Feature extractor. We design a progressive describe-then-detect learning pipeline to save training costs, which partly decouples the training process yet ensures its consistency by sharing the loss function structure. We build a unified keypoint location training framework which can predict keypoint locations by a learnable network branch to avoid slow post-process, thus we increase speed while keep accuracy. Our method achieves the best balance on training costs, accuracy and real-time performance in experiments on different tasks.
Qingqing Yan
SMC2
2023 Age of Information Minimization for Short-Packet Communications RSMA in Satellite-based IoT
abstract
This paper aims to minimize the age of information (AoI) of downlink rate-splitting multiple access (RSMA) in satellite-based Internet of Things (S-IoT) network over shadowed-Rician fading channels, where a satellite multicasts with multiple user equipments (UEs) by timely transmitting short-packet status updates. First, the expressions for block error rate (BLER) and average AoI (AAoI) are derived in a closed-form for short-packet communications with finite blocklength bound. Then, we formulate an AAoI minimization problem based on the theoretical derivations for the downlink RSMA S-IoT network, and design an age-optimal stationary power allocation (ASPA) scheme to solve the problem by utilizing the particle swarm optimization (PSO) algorithm. We further propose an age-optimal dynamic power allocation (ADPA) scheme based on the Markov decision process (MDP), and solve it by two deep reinforcement learning (DRL) algorithms. Monte Carlo simulations verify the accuracy of our derivations of BLER and AAoI, and also show that our ADPA scheme outperforms the related schemes.
Qingqing Yan, Jian Jiao 0001, Yasong Wang, Lirong An, Rongxing Lu, Qinyu Zhang 0001
VTC Fall1
2023 A Geometric Knowledge Oriented Single-Frame 2D-to-3D Human Absolute Pose Estimation Method
abstract
As a critical part of the 3D human pose estimation (HPE), establishing the 2D-to-3D lifting mapping is limited by depth ambiguity. Most current works generally lack the quantitative analysis of the relative depth expression and the depth ambiguity error expression in lifting mapping, resulting in low prediction efficiency and poor interpretability. To this end, this paper mines and leverages prior geometric knowledge of these expressions based on the pinhole imaging principle, decoupling the 2D-to-3D lifting mapping and simplifying the model training. Specifically, this paper proposes a prior geometric knowledge oriented pose estimation model with two-branch transformer architectures, explicitly introducing high-dimensional prior geometric features to improve model efficiency and interpretability. It converts the regression of spatial coordinates into the prediction of spatial direction vectors between joints to generate multiple feasible solutions further alleviate the depth ambiguity. Moreover, this paper raises a novel non-learning-based absolute depth estimation algorithm based on prior geometric relationship decoupling from relative depth expression for the first time. It establishes multiple independent depth mapping from non-root nodes to the root node to calculate the absolute depth candidate, which is parameter-free, plug-and-play, and interpretable. Experiments show that the proposed pose estimation model achieves state-of-the-art performance on Human 3.6M and MPI-INF-3DHP benchmarks with lower parameters and faster inference speed, and the proposed absolute depth estimation algorithm achieves similar performance to traditional methods without any network parameters. The source code are available athttps://github.com/Humengxian/GKONet.
Mengxian Hu, Shu Li 0005, Qingqing Yan, Qin Fang
IEEE Trans. Circuits Syst. Video Technol.4
2022 HoloSeg: An Efficient Holographic Segmentation Network for Real-time Scene Parsing
abstract
Real-time semantic segmentation is a crucial but challenging dense prediction task for scene parsing. However, the existing CNN-based methods commonly bias the model in favor of speed-boosting compromising spatial resolution due to business requirements and hardware constrains, which impedes the high-accuracy segmentation result. To address the dilemma, we provide a novel Holographic Segmentation Network (HoloSeg), which presents a strong ability of comprehensive information preservation and extraction, and achieves a better trade-off between speed and accuracy. We first design a Lossless Sample Pair (LSP) without any stride for early spatial preservation and later resolution recovery while modeling long-range context dependence. Then, we propose Distributed Pyramid Learning (DPL) to efficiently extract multiscale features and saves a lot of computation. Finally, we propose Resolution Fusion and Restoration (RFR) to fuse multi-level semantic representations across stages and generate output without decoder. Without bells and whistles, HoloSeg achieves state-of-the-art performance on the Cityscapes benchmark which reports 76.24% mIoU at 231 FPS. Code is available online: https://github.com/LiShuTJ/HoloSeg.
Shu Li 0005, Qingqing Yan, Ming Liu 0001
ICRA2
2022 Perceptual quality assessment for no-reference image via optimization-based meta-learning
Longsheng Wei, Qingqing Yan, Wei Liu 0005, Dapeng Luo
Inf. Sci.2
2022 RoboSeg: Real-Time Semantic Segmentation on Computationally Constrained Robots
abstract
Real-time and high-performance segmentation is a crucial but challenging perception task for computationally constrained robots, such as the humanoid NAO robot used in the RoboCup Soccer Standard Platform League. However, most existing convolutional neural network (CNN)-based models for semantic segmentation suffer from massive computational costs, which prevents them from being applied to performing real-time inference with a NAO. In this article, we first publish meticulously annotated datasets for training and evaluating semantic segmentation models. Then, we propose a fast downsampling module that downsamples the image while maintaining the spatial information and a novel dense learning module that learns high-level semantic information while recovering the spatial details. Based on these operations, by using a multiscale fusion method to recover the resolution, we propose a more efficient and real-time segmentation model called RoboSeg primarily aimed at offering better speed and accuracy tradeoffs. Finally, to accommodate practical engineering applications, we offer a promising deployment guideline for the CNN model describing how to deploy it on computational resource-limited robots and achieve real-time performance. The experimental results show that the RoboSeg exceeds the state-of-the-art networks in RoboCup scene segmentation: we attain a mean IoU of 87.35% and a pixel accuracy of 96.88% on our dataset using a model that contains only 0.29M parameters and performs just 0.73 GFLOPs. Under the proposed deployment strategies, the network can run at above 30 FPS on NAO robots with downsampled frames.
Qingqing Yan, Shu Li 0005, Ming Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Real-Time Lightweight CNN in Robots with Very Limited Computational Resources: Detecting Ball in NAO
Qingqing Yan
ICVS1