Qiong Liu 0006

dblp:20/718-6 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0002-6410-615XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Walk in Others' Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective Transformation
abstract
Visual perspective-taking, an ability to envision others’ perspectives from a single self-perspective, is vital in human-robot interactions. Thus, we introduce a human-centric visual grounding task and a dataset to evaluate this ability. Recent advances in vision-language models (VLMs) have shown potential for inferring others’ perspectives, yet are insensitive to information differences induced by slight perspective changes. To address this problem, we propose a top-view enhanced perspective transformation (TEP) method, which decomposes the transition from robot to human perspectives through an abstract top-view representation. It unifies perspectives and facilitates the capture of information differences from diverse perspectives. Experimental results show that TEP improves performance by up to 18%, exhibits perspective-taking abilities across various perspectives, and generalizes effectively to robotic and dynamic scenarios.
Yuqi Bu, Xin Wu 0003, Zirui Zhao, Yi Cai 0001, David Hsu, Qiong Liu 0006
ACL (1)6
2025 Highly realistic synthetic dataset for pixel-level DensePose estimation via diffusion model
Jiaxiao Wen, Tao Chu, Qiong Liu 0006
Pattern Recognit.3
2025 Error-Aware Generative Reasoning for Zero-Shot Visual Grounding
abstract
Zero-shot visual grounding is the task of identifying and localizing an object in an image based on a referring expression without task-specific training. Existing methods employ heuristic rules to step-by-step perform visual perception for visual grounding. Despite their remarkable performance, there are still two limitations. First, such a rule-based manner struggles with expressions that are not covered by predefined rules. Second, existing methods lack a mechanism for identifying and correcting visual perceptual errors of incomplete information, resulting in cascading errors caused by reasoning based on incomplete visual perception results. In this article, we propose an Error-Aware Generative Reasoning (EAGR) method for zero-shot visual grounding. To address the limited adaptability of existing methods, a reasoning chain generator is presented, which prompts LLMs to dynamically generate reasoning chains for specific referring expressions. This generative manner eliminates the reliance on human-written heuristic rules. To mitigate visual perceptual errors of incomplete information, an error-aware mechanism is presented to elicit LLMs to identify these errors and explore correction strategies. Experimental results on four benchmarks show that EAGR outperforms state-of-the-art zero-shot methods by up to 10% and an average of 7%.
Yuqi Bu, Xin Wu 0003, Yi Cai 0001, Qiong Liu 0006, Tao Wang 0036, Qingbao Huang
IEEE Trans. Multim.4
2024 Depth Decoupling for Bottom-Up Multi-Person 3D Pose Estimation
Zhaokun Li, Qiong Liu 0006
PRCV (11)2
2024 CoNPL: Consistency training framework with noise-aware pseudo labeling for dense pose estimation
Jiaxiao Wen, Tao Chu, Junyao Sun, Qiong Liu 0006
Image Vis. Comput.4
2024 FDNet: Feature decoupling for single-stage pose estimation in complex scenes
Qiong Liu 0006
J. Vis. Commun. Image Represent.2
2024 Discriminative features enhancement for low-altitude UAV object detection
Shuqin Huang, Shasha Ren, Qiong Liu 0006
Pattern Recognit.4
2024 Devil in the details: Delving into accurate quality scoring for DensePose
abstract
How to score the quality of the network output is an essential but long-neglected problem in DensePose, which dramatically limits the potential of the existing methods. To fill the blank in the quality estimation of DensePose, we conduct rigorous experiments to clarify the key factors that accurately reflect the quality of DensePose results. We find that the accurate results already exist in the candidate pool but are mistakenly removed due to the inappropriate quality scores. To solve this problem, we proposed DensePose Scoring RCNN (DS RCNN), a simple and comprehensive quality estimation framework to learn the calibrated quality score and select high-quality results from the pool. DS RCNN introduces a quality scoring module (QSM) and a quality perception module (QPM) into the existing high-performance pipeline. The QSM scores the quality of DensePose results by fusing diverse quality information, and the QPM enhances the ability of quality perception by extracting instance-aware quality features guided by the predicted IUV maps. Benefiting from the superiority of QSM and QPM, DS RCNN outperforms baselines by up to 4.8 AP on the DensePose-COCO dataset.
Junyao Sun, Qiong Liu 0006
Pattern Recognit.2
2024 Small Target Augmentation for Urban Remote Sensing Image Real-Time Segmentation
abstract
Urban remote sensing (URS) image segmentation is very important for many applications from automotive navigation to infrastructure monitoring, and urban management. There are numerous small targets in URS image due to a large shooting field of view. However, the existing learning-based real-time image segmentation methods are not strong enough to handle small target and edge segmentation problems well, resulting in small targets that are easy to be missed, and blurred target edges. To segment the small targets and edges more accurately in real time, we propose a fast URS image segmentation method based on a multi-layer pixel attention mechanism (MPAM). We improve the performance and efficiency of the URS image segmentation from two perspectives: model and data. Specifically, to enhance semantic detailed information such as small targets and edges, we design mask-guided edge and small target feature enhancement modules in the real-time segmentation network. In addition, we propose a small target data enhancement method which uses an interpolation algorithm to amplify small targets in URS images, in order to improve the efficiency of existing URS data. The experimental results on Vaihingen, Potsdam, and DLRSD datasets show that the segmentation accuracy of our method reaches 86.74% mIoU, which is better than the state-of-the-art algorithms STDC, CFNet, and UNetFormer.
Shasha Ren, Qiong Liu 0006
IEEE Trans. Intell. Transp. Syst.2
2023 NRPose: Towards noise resistance for multi-person pose estimation
Jianhang He, Junyao Sun, Qiong Liu 0006, Shaowu Peng
Pattern Recognit.3
2023 PoiseNet: Dealing With Data Imbalance in DensePose
abstract
Data imbalance, a foundational problem in machine learning, has received little attention in DensePose and has become one of the main obstacles in front of existing methods. We reveal two imbalances in DensePose: inter-surface imbalance and intra-surface imbalance. First, the human body parts can be of various sizes, making 3D surfaces contain different numbers of annotations. Classifiers trained on such unbalanced data suffer a decline on surfaces with fewer annotations. Second, annotations within each surface are not uniformly distributed. Regressors trained on such uneven data suffer a decline in 3D surface areas with sparse annotations. To solve these imbalances, we propose PoiseNet which integrates adaptive equalization loss (AEQL) and block balanced localization (BBL). Specifically, to address the inter-surface imbalance, AEQL adaptively reweights surfaces based on the number of annotations and the classification scores. BBL alleviates the intra-surface imbalance by unevenly blocking each surface according to annotation distribution. Experimental results on the DensePose-COCO dataset show that our PoiseNet surpasses baselines by up to 1.4 AP.
Junyao Sun, Jingkai Zhou, Qiong Liu 0006
IEEE Trans. Circuits Syst. Video Technol.3
2023 Scene-Text Oriented Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attributes and neglect scene text. Since scene text has the functions of object identification and disambiguation, it is naturally and frequently used to refer to objects. However, existing methods do not explicitly recognize text in images and fail to align scene text mentioned in expressions with the text shown in images, resulting in object localization errors. This article takes the first step toward addressing these limitations. First, we introduce a new task called scene-text oriented referring expression comprehension, which aims to align visual cues and textual semantics of scene text with referring expressions and visual contents. Second, we propose a scene text awareness network that can bridge the gap between texts from two modalities by grounding visual representations of expression-correlated scene texts. Specifically, we propose a correlated text extraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address the fixed alignment problem and absent alignment problem. These modules ensure that the proposed method focuses on local regions that are most relevant to scene text, thus mitigating the misalignment of scene text with irrelevant regions. Third, to conduct quantitative evaluations, we establish a new benchmark dataset called RefText. Experimental results demonstrate that the proposed method can effectively comprehend scene-text oriented referring expressions and achieves excellent performance.
Yuqi Bu, Liuwu Li, Jiayuan Xie, Qiong Liu 0006, Yi Cai 0001, Qingbao Huang, Qing Li 0001
IEEE Trans. Multim.4
2022 Bridging the Gap between Expression and Scene Text for Referring Expression Comprehension (Student Abstract)
abstract
Referring expression comprehension aims at grounding the object in an image referred to by the expression. Scene text that serves as an identifier has a natural advantage in referring to objects. However, existing methods only consider the text in the expression, but ignore the text in the image, leading to a mismatch. In this paper, we propose a novel model that can recognize the scene text. We assign the extracted scene text to its corresponding visual region and ground the target object guided by expression. Experimental results on two benchmarks demonstrate the effectiveness of our model.
Yuqi Bu, Jiayuan Xie, Liuwu Li, Qiong Liu 0006, Yi Cai 0001
AAAI4
2022 Complementarity-aware cross-modal feature fusion network for RGB-T semantic segmentation
Tao Chu, Qiong Liu 0006
Pattern Recognit.3
2021 Ground Plane Context Aggregation Network for Day-and-Night on Vehicular Pedestrian Detection
abstract
Ground plane context is an essential semantic information in on-road pedestrian detection task. Due to viewpoint geometry constraints, pedestrians only appear in certain regions of the image, which is close to the horizon area of the ground plane. As a result, the lacking to ground plane context information may cause pedestrian detection system suffering from severe false alarm (i.e. high false positive (FP) rate). For Advanced Driver Assistance System (ADAS), high FP rate not only distracts the driver, but also causes frequent unexpected braking to damage the vehicle’s hardware. In this paper, a novel pedestrian detection method called ground plane context aggregation network (GPCAnet) is proposed, which integrates ground plane context information into deep learning based detector to drastically reduce the FP rate. The proposed GPCAnet consists of two modules: i) a ground area predication (GAP) branch is appended on top of convolutional feature map of the backbone network, in parallel with existing branches, for region proposal, classification and bounding box regression; ii) based on GAP, a ground-region proposal network (GRPN) is designed to filter FP cases in order to reduce computations. To evaluate the effectiveness of proposed GPCAnet, experiments on day and night on-road pedestrian detection are performed on both visible and far infrared pedestrian detection datasets, e.g. Caltech and SCUT. Experimental results show that GPCAnet achieves better performance than state-of-the-art methods, while drastically reducing FP rate in pedestrian detection.
Zhewei Xu, Chi-Man Vong, Chi-Chong Wong, Qiong Liu 0006
IEEE Trans. Intell. Transp. Syst.4
2020 Novel up-scale feature aggregation for object detection in aerial images
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006
Neurocomputing5
2020 SCNet: Scale-aware coupling-structure network for efficient video object detection
Fengchao Wang, Zhewei Xu, Chi-Man Vong, Qiong Liu 0006
Neurocomputing5
2020 Tracking objects with partial occlusion by background alignment
Feng Wu 0004, Chi-Man Vong, Qiong Liu 0006
Neurocomputing3
2019 Scale adaptive image cropping for UAV object detection
Jingkai Zhou, Chi-Man Vong, Qiong Liu 0006, Zhenyu Wang 0001
Neurocomputing3
2018 Object tracking via Online Multiple Instance Learning with reliable components
Feng Wu 0004, Shaowu Peng, Jingkai Zhou, Qiong Liu 0006, Xiaojia Xie
Comput. Vis. Image Underst.4
2016 Online Multi-Instance Multi-Label learning for protein function prediction
abstract
Protein function prediction is a challenging and essential research problem in the field of computational biology. Conventionally, a protein consists of a number of structural domains and performs multiple function. By representing proteins, domains and functions by bags as well as instances and classes respectively, we are able to model the protein function prediction task as the Multi-Instance Multi-Label (MIML) learning problem. Existing MIML algorithms mainly focus on batch setting where training examples are available before learning. Such offline paradigm works well in simulation, but it may be not feasible for real-world online applications where data comes one by one or chunk by chunk. In this paper, we investigate the protein function prediction problem under a new learning framework, called Online Multi-Instance Multi-Label (OMIML) learning, where MIML protein examples arrive sequentially in an online setting, and develop two OMIML algorithms (OMIML-I and OMIML-B) to make predictions for the incoming data. In the proposed OMIML algorithms, variable-length features are constructed to represent the MIML protein examples based on an incremental vocabulary mechanism. In particular, the incremental vocabularies that OMIML-I and OMIML-B are based on consist of instances and bags, respectively. Then we seek an online prediction for each new arrived protein example by incorporating the constructed features into an online multi-label learning algorithm which is constructed by introducing an artificial label into an online multi-label ranking model. We evaluate the algorithms on the protein dataset consisting of seven real-world organisms. Experimental results have demonstrated the effectiveness of the proposed OMIML algorithms for protein function prediction.
Feng Wu 0004, Qiong Liu 0006, Tianyong Hao, Xiaojun Chen 0006, Qingyao Wu
BIBM2