Qingyong Hu

dblp:192/3850 · DBLP profile ↗
← Back
53ranked-venue papers
8as first author
47since 2021 · last 2026
0000-0003-0337-9207ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 25 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration
abstract
Large language models (LLMs) have achieved remarkable success in various natural language processing tasks, yet they remain prone to generating factually incorrect outputs—known as "hallucinations". While recent approaches have shown promise for hallucination detection by repeatedly sampling from LLMs and quantifying the semantic inconsistency among the generated responses, they rely on fixed sampling budgets that fail to adapt to query complexity, resulting in computational inefficiency. We propose an Adaptive Bayesian Estimation framework for Semantic Entropy with Guided Semantic Exploration, which dynamically adjusts sampling requirements based on observed uncertainty. Our approach employs a hierarchical Bayesian framework to model the semantic distribution, enabling dynamic control of sampling iterations through variance-based thresholds that terminate generation once sufficient certainty is achieved. We also develop a perturbation-based importance sampling strategy to systematically explore the semantic space. Extensive experiments on four QA datasets demonstrate that our method achieves superior hallucination detection performance with significant efficiency gains. In low-budget scenarios, our approach requires about 50% fewer samples to achieve comparable detection performance to existing methods, while delivers an average AUROC improvement of 12.6% under the same sampling budget.
Qiyao Sun, Xingming Li 0001, Xixiang He, Ao Cheng, Xuanyu Ji, Hailun Lu, Runke Huang, Qingyong Hu
AAAI8
2026 When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools
Xingming Li 0001, Runke Huang, Yanan Bao, Yuye Jin, Yuru Jiao, Qingyong Hu
AIED (1)6
2025 Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
abstract
Periodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) offer promising potential to effectively capture and understand their complex nature. However, current MLLMs struggle with periodic tasks due to limitations in: 1) lack of temporal modelling and 2) conflict between short and long periods. This paper introduces Period-LLM, a multimodal large language model designed to enhance the performance of periodic tasks across various modalities, and constructs a benchmark of various difficulty for evaluating the cross-modal periodic capabilities of large models. Specially, We adopt an "Easy to Hard Generalization" paradigm, starting with relatively simple text-based tasks and progressing to more complex visual and multimodal tasks, ensuring that the model gradually builds robust periodic reasoning capabilities. Additionally, we propose a "Resisting Logical Oblivion" optimization strategy to maintain periodic reasoning abilities during semantic alignment. Extensive experiments demonstrate the superiority of the proposed Period-LLM over existing MLLMs in periodic tasks. The code is available at https: //github.com/keke-nice/Period-LLM.
Yuting Zhang 0008, Hao Lu 0009, Qingyong Hu, Yin Wang 0004, Kaishen Yuan, Xin Liu 0012, Kaishun Wu
CVPR3
2025 ToolFiVe: Enhancing Tool-Augmented LLMs via Tool Filtering and Verification
abstract
Tool-augmented Large Language Models (LLMs) provide a robust theoretical foundation for AI agents, with the generation of reasoning plans being a crucial stage. Previous methods for generating reasoning plans primarily rely on In-Context Learning (ICL) or Supervised Fine-Tuning (SFT). However, methods based on ICL struggle with accurately utilizing tools, while those based on SFT face challenges in adapting to new toolsets and tasks. To address these issues, we introduce ToolFiVe, a general, plug-and-play, self-correction-based framework for leveraging specialized tools in compositional reasoning tasks. ToolFiVe redesigns the process of reasoning plan generation and integrates the tool-execution stage. Specifically, in the reasoning plan generation stage, ToolFiVe employs a filtering module to exclude task-irrelevant tools, creating a candidate toolset. ToolFiVe then constructs prompts using the candidate toolset and iteratively generates and refines the reasoning plan. Subsequently, a verification module evaluates the completeness of the reasoning plan, providing feedback for the self-correction loop. Ultimately, the reasoning plan developed during this stage guides the execution of the tools. Extensive experiments demonstrate that ToolFiVe outperforms other state-of-the-art (SOTA) methods, highlighting the significance of reasoning plan generation for general tool-augmented LLMs.
Hailun Lu, Xingming Li 0001, Xuanyu Ji, Zhigang Kan, Qingyong Hu
ICASSP5
2025 $U^2$ Frame: A Unified and Unsupervised Learning Framework for LiDAR-Based Loop Closing
abstract
Loop closing is critically important in Simultaneous Localization and Mapping (SLAM) due to its ability to correct accumulated localization errors. However, existing methods are hindered by the difficulty of acquiring pose labels and the unreliability of ground truth data. In this paper, we propose$U^{2}$Frame, a unified LiDAR-based loop closing framework that handles both loop closure detection and relative pose estimation without any ground truth training data. Specifically, the natural temporal-spatial correlation in point cloud sequences is first leveraged to supervise the network training, where near scans are treated as positives and vice versa as negatives. A new neural architecture is then constructed to jointly learn highly discriminative local and global features for loop closure detection. Additionally, an effective candidate verification module that exploits high-order geometric information is presented to further filter out false loop closures and estimate precise poses. We extensively evaluate$U^{2}$Frame on multiple datasets according to two tasks derived from loop closing: loop closure detection and loop pose estimation. Comparative experiments demonstrate that our method outperforms existing state-of-the-art supervised techniques and has a strong generalization ability across unseen scenarios. Our code is released at https://github.com/yxin-zhang/U2Frame.
Sheng Ao, Ye Zhang 0037, Qingyong Hu, Tao Chang, Yulan Guo
ICRA5
2025 mmTremor: Practical Tremor Monitoring for Parkinson's Disease and Essential Tremor in Daily Life
abstract
Tremor, a prevalent symptom in various neurological disorders, significantly impacts patients' quality of life. Regular and precise tremor monitoring is essential for optimizing treatment effectiveness. Existing at-home monitoring solutions have limitations in daily life due to low adherence, privacy concerns, or inevitable interference from confounding body components and overlapped motions. In this paper, we propose mmTremor, the first privacy-preserving contactless system to achieve practical tremor detection during activities of daily living (ADL) in real-world settings. We design a mmWave-depth fusion tracking algorithm to handle the high interference from confounding body components and a multimodal spatiotemporal deep learning pipeline to effectively exploit latent tremor information. Additionally, a spatial contrastive unsupervised adaptation is proposed for better discriminability and adaptation capability to unseen domains. To fully assess mmTremor in real-world scenarios, we collect a diverse dataset of 28 patients and 9 healthy subject simulations in over 20 distinct environments, including offices, hospitals, and homes. Extensive evaluations demonstrate that mmTremor achieves a high macro-F1 of 0.877 in tremor detection, showcasing the potential as a transformative contactless solution for tremor monitoring in daily life. The dataset will be open-source to facilitate future research.
Qingyong Hu, Jinjian Wang, Zirui Huang, Guihua Li, Qian Zhang 0001
MobiCom1
2025 PhysDrive: A Multimodal Remote Physiological Measurement Dataset for In-vehicle Driver Monitoring
abstract
Robust and unobtrusive in-vehicle physiological monitoring is crucial for ensuring driving safety and user experience. While remote physiological measurement (RPM) offers a promising non-invasive solution, its translation to real-world driving scenarios is critically constrained by the scarcity of comprehensive datasets. Existing resources are often limited in scale, modality diversity, the breadth of biometric annotations, and the range of captured conditions, thereby omitting inherent real-world challenges in driving. Here, we present PhysDrive, the first large-scale multimodal dataset for contactless in-vehicle physiological sensing with dedicated consideration of various modality settings and driving factors. PhysDrive collects data from 48 drivers, including synchronized RGB, near-infrared camera, and raw mmWave radar data, accompanied by six synchronized ground truths (ECG, BVP, Respiration, HR, RR, and SpO2). It covers a wide spectrum of naturalistic driving conditions, including driver motions, dynamic natural light, vehicle types, and road conditions. We extensively evaluate both signal‑processing and deep‑learning methods on PhysDrive, establishing a comprehensive benchmark across all modalities, and release full open‑source code with compatibility for mainstream public toolboxes. We envision PhysDrive will serve as a foundational resource and accelerate research on multimodal driver monitoring and smart‑cockpit systems.
Jiyao Wang 0002, Xiao Yang 0025, Qingyong Hu, Jack Tang, Dengbo He, Ying-Cong Chen, Kaishun Wu
NeurIPS3
2025 WSSIC-Net: Weakly-Supervised Semantic Instance Completion of 3D Point Cloud Scenes
abstract
Semantic instance completion aims to recover the complete 3D shapes of foreground objects together with their labels from a partial 2.5D scan of a scene. Previous works have relied on full supervision, which requires ground-truth annotations, in the form of bounding boxes and complete 3D objects. This has greatly limited their real-world application because the acquisition of ground-truth data is very costly and time-consuming. To address this bottleneck, we propose a Weakly-Supervised Semantic Instance Completion Network (WSSIC-Net), which learns real-world partial point cloud object completion without requiring the ground truth of complete 3D objects. Instead, WSSIC-Net leverages 3D ground-truth bounding boxes, partial objects of a raw scene, and unpaired synthetic 3D point clouds. More specifically, a 3D detector is used to encode partial point clouds into proposal features, which are then fed into two branches. The first branch uses fully supervised box prediction based on proposal features. The second branch, hereinafter called instance completion, leverages the proposal features as partial object features to achieve weakly-supervised instance completion. A Generative Adversarial Network (GAN) completes the partial features of the 2.5D foreground objects of real-world scenes using only unpaired but semantically-consistent complete synthetic point clouds. In our experiments, we demonstrate that the fully-supervised 3D detection and the weakly-supervised instance completion complement one another. The qualitative and quantitative evaluations on the ScanNet v2 dataset demonstrate that the proposed "weakly-supervised" approach consistently achieves comparable performance to the state-of-the-art "fully supervised" methods.
Zhiheng Fu, Yulan Guo, Minglin Chen, Qingyong Hu, Hamid Laga, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Image Process.4
2025 Ubicon-BP: Towards Ubiquitous, Contactless Blood Pressure Detection Using Smartphone
abstract
Blood pressure (BP) is a critical physiological parameter closely associated with severe diseases such as heart failure and kidney damage. Current methods either require additional or dedicated hardware, or closing touching to the devices, causing discomfort and inconvenience. Therefore, a convenient, contactless BP measurement solution is highly desired. In this work, we present Ubicon-BP, a ubiquitous, device-free, and contactless BP detection application. Ubicon-BP calculates BP based on the pulse transit time (PTT), a key feature that is medically proven correlated with BP. However, using smartphone sensors to contactless calculate PTT is non-trivial since it requires a micro-second level precision for cardiac event detection. To address this issue, we propose leveraging the acoustic sensors in smartphone to detect vibrations caused by heart valve movements, as well as camera sensors to measure finger pulses. To accurately measure heartbeat signal that are susceptible to motion, we first improve the sensing granularity of acoustic signals and then introduce the IQ-MVED model to eliminate motion interference. Furthermore, when recovering pulse signals from video signals, issues such as poor generalization performance arise. Consequently, we propose the TS-CAN and meta-learning models to obtain personalized pulse signals. Finally, we transform the extracted time-frequency features from the recovered heartbeats and pulse signals to the corresponding BP. Comprehensive testing involving 50 subjects reveal a standard deviation error of$ 4.27 \; \text{mmHg}$for diastolic pressure and$ 6.36 \; \text{mmHg}$for systolic pressure, respectively.
Yuan Wu 0007, Shoudu Bai, Qingyong Hu, Bo Wang 0047, Xinrong Hu, Yanjiao Chen
IEEE Trans. Mob. Comput.3
2025 Adaptive Sparse Memory Networks for Efficient and Robust Video Object Segmentation
abstract
Recently, memory-based networks have achieved promising performance for video object segmentation (VOS). However, existing methods still suffer from unsatisfactory segmentation accuracy and inferior efficiency. The reasons are mainly twofold: 1) during memory construction, the inflexible memory storage mechanism results in a weak discriminative ability for similar appearances in complex scenarios, leading to video-level temporal redundancy, and 2) during memory reading, matching robustness and memory retrieval accuracy decrease as the number of video frames increases. To address these challenges, we propose an adaptive sparse memory network (ASM) that efficiently and effectively performs VOS by sparsely leveraging previous guidance while attending to key information. Specifically, we design an adaptive sparse memory constructor (ASMC) to adaptively memorize informative past frames according to dynamic temporal changes in video frames. Furthermore, we introduce an attentive local memory reader (ALMR) to quickly retrieve relevant information using a subset of memory, thereby reducing frame-level redundant computation and noise in a simpler and more convenient manner. To prevent key features from being discarded by the subset of memory, we further propose a novel attentive local feature aggregation (ALFA) module, which preserves useful cues by selectively aggregating discriminative spatial dependence from adjacent frames, thereby effectively increasing the receptive field of each memory frame. Extensive experiments demonstrate that our model achieves state-of-the-art performance with real-time speed on six popular VOS benchmarks. Furthermore, our ASM can be applied to existing memory-based methods as generic plugins to achieve significant performance improvements. More importantly, our method exhibits robustness in handling sparse videos with low frame rates.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Qingyong Hu, Yulan Guo
IEEE Trans. Neural Networks Learn. Syst.5
2024 Learning Coupled Dictionaries from Unpaired Data for Image Super-Resolution
abstract
The difficulty of acquiring high-resolution (HR) and low-resolution (LR) image pairs in real scenarios limits the performance of existing learning-based image super-resolution (SR) methods in the real world. To conduct training on real-world unpaired data, current methods focus on synthesizing pseudo LR images to associate unpaired images. However, the realness and diversity of pseudo LR images are vulnerable due to the large image space. In this paper, we cir-cumvent the difficulty of image generation and propose an alternative to build the connection between unpaired images in a compact proxy space. Specifically, we first construct coupled HR and LR dictionaries, and then encode HR and LR images into a common latent code space using these dictionaries. In addition, we develop an autoencoder-based framework to couple these dictionaries during optimization by reconstructing input HR and LR images. The coupled dictionaries enable our method to employ a shal-low network architecture with only 18 layers to achieve efficient image SR. Extensive experiments show that our method (DictSR) can effectively model the LR-to-HR mapping in coupled dictionaries and produces state-of-the-art performance on benchmark datasets.
Longguang Wang, Juncheng Li 0003, Yingqian Wang 0002, Qingyong Hu, Yulan Guo
CVPR4
2024 ACRF: Compressing Explicit Neural Radiance Fields via Attribute Compression
abstract
In this work, we study the problem of explicit NeRF compression. Through analyzing recent explicit NeRF models, we reformulate the task of explicit NeRF compression as 3D data compression. We further introduce our NeRF compression framework, Attributed Compression of Radiance Field (ACRF), which focuses on the compression of the explicit neural 3D representation. The neural 3D structure is pruned and converted to points with features, which are further encoded using importance-guided feature encoding. Furthermore, we employ an importance-prioritized entropy model to estimate the probability distribution of transform coefficients, which are then entropy coded with an arithmetic coder using the predicted distribution. Within this framework, we present two models, ACRF and ACRF-F, to strike a balance between compression performance and encoding time budget. Our experiments, which include both synthetic and real-world datasets such as Synthetic-NeRF and Tanks&Temples, demonstrate the superior performance of our proposed algorithm.
Guangchi Fang, Qingyong Hu, Longguang Wang, Yulan Guo
ICLR2
2024 Lighten CARAFE: Dynamic Lightweight Upsampling with Guided Reassemble Kernels
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yinghui Gao, Ping Zhong 0001
ICPR (4)2
2024 Content-Aware Feature Upsampling for Voxel-Based 3D Semantic Segmentation
Ruigang Fu, Qingyong Hu, Ping Zhong 0001
ICPR (30)3
2024 ICPR 2024 Competition on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Han Wang 0049, Furui Chen, Silei Liu, Xiaomin Huang, Shining Wang, Ying Li 0017, Peng Wang 0015, Shiyong Peng, Xiaokai Bi, Renbin Zou, Wenjing Deng, Zhen Cui 0001
ICPR (34)3
2024 Cross-shaped Separated Spatial-Temporal UNet Transformer For Accurate Channel Prediction
abstract
Accurate channel estimation is crucial for the performance gains of massive multiple-input multiple-output (mMIMO) technologies. However, it is bandwidth-unfriendly to estimate large channel matrix frequently to combat the time-varying wireless channel. Deep learning-based channel prediction has emerged to exploit the temporal relationships between historical and future channels to address the bandwidth-accuracy trade-off. Existing methods with convolutional or recurrent neural networks suffer from their intrinsic limitations, including restricted receptive fields and propagation errors. Therefore, we propose a Transformer-based model, CS3T-UNet tailored for mMIMO channel prediction. Specifically, we combine the cross-shaped spatial attention with a group-wise temporal attention scheme to capture the dependencies across spatial and temporal domains, respectively, and introduce the shortcut paths to well-aggregate multi-resolution representations. Thus, CS3T-UNet can globally capture the complex spatial-temporal relationship and predict multiple steps in parallel, which can meet the requirement of channel coherence time. Extensive experiments demonstrate that the prediction performance of CS3T-UNet surpasses the best baseline by at most 6.86 dB with a smaller computation cost on two channel conditions.
Hua Kang, Qingyong Hu, Huangxun Chen, Qianyi Huang, Qian Zhang 0001
INFOCOM2
2024 Rethinking Dual-Stream Super-Resolution Semantic Learning in Medical Image Segmentation
abstract
Image segmentation is fundamental task for medical image analysis, whose accuracy is improved by the development of neural networks. However, the existing algorithms that achieve high-resolution performance require high-resolution input, resulting in substantial computational expenses and limiting their applicability in the medical field. Several studies have proposed dual-stream learning frameworks incorporating a super-resolution task as auxiliary. In this paper, we rethink these frameworks and reveal that the feature similarity between tasks is insufficient to constrain vessels or lesion segmentation in the medical field, due to their small proportion in the image. To address this issue, we propose a DS2F (Dual-Stream Shared Feature) framework, including a Shared Feature Extraction Module (SFEM). Specifically, we present Multi-Scale Cross Gate (MSCG) utilizing multi-scale features as a novel example of SFEM. Then we define a proxy task and proxy loss to enable the features focus on the targets based on the assumption that a limited set of shared features between tasks is helpful for their performance. Extensive experiments on six publicly available datasets across three different scenarios are conducted to verify the effectiveness of our framework. Furthermore, various ablation studies are conducted to demonstrate the significance of our DS2F.
Zhongxi Qiu, Xiaoshan Chen, Dan Zeng 0002, Qingyong Hu, Jiang Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 BSTS: A Weakly-Supervised Method for Semantic Learning of 3D Point Clouds
abstract
Point cloud semantic understanding with fewer point-wise annotations is an ongoing challenge that has yet to be fully addressed in the literature. Although previous approaches have achieved some success with weak supervision, our research reveals that even basic bounding box annotations and subcloud-level tags can provide valuable information for point cloud semantic segmentation. We propose a framework using Bounding boxes and Subcloud-level Tags for Semantic Segmentation, named BSTS. Our method explores local topological structures and geometric priors within and outside bounding boxes to produce reliable pseudo labels. Once bounding boxes of instances are provided for a point cloud, raw points can be divided into three categories: potential foreground points, ambiguous points, and clear background points. To ensure the reliability of the pseudo labels derived from weak supervision, we utilized an Attention-based Self-Training (AST) pipeline and the Point Class Activation Maps (PCAMs) technique. Subsequently, the segmentation network is trained using the generated pseudo labels. Experiments are conducted on two widely used large-scale benchmarks, including S3DIS and ScanNet. Our method achieves competitive semantic performance with the fully-supervised counterpart via low-cost bounding box annotations and subcloud-level tags.
Yan Liu 0043, Qingyong Hu, Yulan Guo
IEEE Trans. Circuits Syst. Video Technol.2
2024 PDANet: Point Distribution Awareness for 3-D Object Detection From LiDAR Point Clouds
abstract
Recent deep learning models for 3-D object detection (3DOD) based on LiDAR point clouds have exhibited promising performance in the detection of on-road objects. However, their performance is often constrained when detecting small objects like pedestrians and cyclists. Due to occlusions and their inherent volumes, the spatial and semantic information of small objects is relatively limited. Therefore, we propose a point distribution awareness network (PDANet) that emphasizes the structural distribution of objects, to enhance the extraction of positional, geometric, and semantic features. It comprises two modules: a multifeature adaptive perception (MFAP) module, enhancing local feature perception, and a semantic-aware downsampling module, enhancing the selection of foreground points. These modules can be used in a plug-and-play manner for point downsampling feature perception. They are then integrated as a branch into representative point-based (IA-SSD), point-voxel-based (PV-RCNN), and voxel-based (Centerpoint) detectors, resulting in PDA-SSD, PDA-PV-RCNN, and PDA-CenterPoint models. Through experiments on the ONCE dataset, the PDANet brought up the mean average precision (mAP) of all three types of baseline models by 3.03%, 4.60%, and 6.15%, respectively. Further experiments on the KITTI dataset illustrate an improvement in average precision (AP) of small objects up to 7.41%.
Dianyu Yu, Qingyong Hu, Wenxia Dai
IEEE Trans. Geosci. Remote. Sens.3
2024 Divide and Conquer: Improving Multi-Camera 3D Perception With 2D Semantic-Depth Priors and Input-Dependent Queries
abstract
3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that accurately estimating both semantic and 3D scene layouts are crucial for this task, existing techniques often neglect the synergistic effects of semantic and depth cues, leading to the occurrence of classification and position estimation errors. Additionally, the input-independent nature of initial queries also limits the learning capacity of Transformer-based models. To tackle these challenges, we propose an input-aware Transformer framework that leverages Semantics and Depth as priors (named SDTR). Our approach involves the use of an S-D Encoder that explicitly models semantic and depth priors, thereby disentangling the learning process of object categorization and position estimation. Moreover, we introduce a Prior-guided Query Builder that incorporates the semantic prior into the initial queries of the Transformer, resulting in more effective input-aware queries. Extensive experiments on the nuScenes and Lyft benchmarks demonstrate the state-of-the-art performance of our method in both 3D object detection and BEV segmentation tasks.
Qingyong Hu, Yongquan Chen, Rui Huang 0001
IEEE Trans. Image Process.2
2023 BUFFER: Balancing Accuracy, Efficiency, and Generalizability in Point Cloud Registration
abstract
An ideal point cloud registration framework should have superior accuracy, acceptable efficiency, and strong generalizability: However, this is highly challenging since existing registration techniques are either not accurate enough, far from efficient, or generalized poorly. It remains an open question that how to achieve a satisfying balance between this three key elements. In this paper, we propose BUFFER, a point cloud registration method for balancing accuracy, efficiency, and generalizability. The key to our approach is to take advantage of both point-wise and patch-wise techniques, while overcoming the inherent drawbacks simultaneously. Different from a simple combination of existing methods, each component of our network has been carefully crafted to tackle specific issues. Specifically, a Point-wise Learner is first introduced to enhance computational efficiency by predicting keypoints and improving the representation capacity of features by estimating point orientations, a Patch-wise Embedder which leverages a lightweight local feature learner is then deployed to extract efficient and general patch features. Additionally, an Inliers Generator which combines simple neural layers and general features is presented to search inlier correspondences. Extensive experiments on real-world scenarios demonstrate that our method achieves the best of both worlds in accuracy, efficiency, and generalization. In particular, our method not only reaches the highest success rate on unseen domains, but also is almost 30 times faster than the strong baselines specializing in generalization. Code is available at https://github.com/aosheng1996/BUFFER.
Sheng Ao, Qingyong Hu, Hanyun Wang, Kai Xu 0004, Yulan Guo
CVPR2
2023 Backdoor Defense via Deconfounded Representation Learning
abstract
Deep neural networks (DNNs) are recently shown to be vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by injecting a few poisoned examples into the training dataset. While extensive efforts have been made to detect and remove backdoors from backdoored DNNs, it is still not clear whether a backdoor-free clean model can be directly obtained from poisoned datasets. In this paper, we first construct a causal graph to model the generation process of poisoned data and find that the backdoor attack acts as the confounder, which brings spurious associations between the input images and target labels, making the model predictions less reliable. Inspired by the causal understanding, we propose the Causality-inspired Backdoor Defense (CBD), to learn deconfounded representations for reliable classification. Specifically, a backdoored model is intentionally trained to capture the confounding effects. The other clean model dedicates to capturing the desired causal effects by minimizing the mutual information with the confounding representations from the backdoored model and employing a sample-wise re-weighting scheme. Extensive experiments on multiple benchmark datasets against 6 state-of-the-art attacks verify that our proposed defense method is effective in reducing backdoor threats while maintaining high accuracy in predicting benign samples. Further analysis shows that CBD can also resist potential adaptive attacks. The code is available at https://github.com/zaixizhang/CBD.
Zaixi Zhang, Qi Liu 0003, Zhicai Wang, Zepu Lu, Qingyong Hu
CVPR5
2023 VPPT: Visual Pre-Trained Prompt Tuning Framework for Few-Shot Image Classification
abstract
Large-scale pre-trained transformers have recently achieved remarkable success in several computer vision tasks. However, it remains highly challenging to fully fine-tune models for downstream tasks, due to the expensive computational and storage cost. Recently, Parameter-Efficient Tuning (PETuning) techniques, e.g., Visual Prompt Tuning (VPT), have significantly reduced the computation cost by inserting lightweight prompt modules including prompt tokens or adapter layers, into the pre-trained models and tuning these prompt modules with a small number of trainable parameters, while keeping the transformer backbone freeze. Although encouraging results were achieved, existing PETuning methods cannot perform well under the few-shot learning settings (i.e., extremely limited training data, with only 1 or 2 shots per class), due to the scarce supervision signal. To this end, we first empirically identify the poor performance is mainly due to the inappropriate way of initializing prompt modules, which has also been verified in the pre-trained language models. Next, we propose a Visual Pre-trained Prompt Tuning framework (VPPT), which pre-trains the prompt modules first and then leverages the pre-trained modules along with the pre-trained transformer backbone to perform prompt tuning on downstream tasks. Extensive experiments show that our VPPT framework achieves 16.08% average accuracy absolute improvement under 1 shot setting on five fine-grained visual classification datasets, compared with the previous PETuning techniques, e.g., VPT, in few-shot image classification.
Zhao Song 0011, Ke Yang 0004, Naiyang Guan, Peng Qiao, Qingyong Hu
ICASSP6
2023 CSI-StripeFormer: Exploiting Stripe Features for CSI Compression in Massive MIMO System
abstract
The massive MIMO gain for wireless communication has been greatly hindered by the feedback overhead of channel state information (CSI) growing linearly with the number of antennas. Recent efforts leverage the DNN-based encoder-decoder framework to exploit correlations within the CSI matrix for better CSI compression. However, existing works have not fully exploited the unique features of CSI, resulting in an unsatisfactory performance under high compression ratios and sensitivity to multipath effects. Instead of treating CSI as common 2D matrices like images, we reveal the intrinsic stripe-based correlation across the CSI matrix. Driven by this insight, we propose CSI-StripeFormer, a stripe-aware encoder-decoder framework to exploit the unique stripe feature for better CSI compression. We design a lightweight encoder with asymmetric convolution kernels to capture various shape features. We further incorporate novel designs tailored for stripe features, including a novel hierarchical Transformer backbone in the decoder and a hybrid attention mechanism to extract and fuse correlations in angular and delay domains. Our evaluation results show that our system achieves an over 7dB channel reconstruction gain under a high compression ratio of 64 in multipath-rich scenarios, significantly superior to current state-of-the-art approaches. This gain can be further improved to 17dB given the extended embedded dimension of our backbone.
Qingyong Hu, Hua Kang, Huangxun Chen, Qianyi Huang, Qian Zhang 0001
INFOCOM1
2023 RIScan: RIS-aided Multi-user Indoor Localization Using COTS Wi-Fi
abstract
Multi-user indoor localization is considered to be one of the most useful wireless applications. Low latency and high robustness to dynamic interference from surrounding people are essential requirements for multi-user localization. However, state-of-the-art (SOTA) indoor localization systems cannot satisfy both requirements at the same time. In this paper, we propose RIScan, a Reconfigurable Intelligent Surface (RIS)-aided localization system that can achieve both low latency and high reliability. We leverage RIS to perform Wi-Fi beam scanning so all clients can figure out their direction in a single scan. However, compared with traditional AP-based systems, the introduction of RIS creates a more complicated signal superposition at the receiver, preventing clients from directly obtaining target beams for direction derivation and localization. To overcome this challenge, we fully utilize the reconfigurability of RIS to endow target beams with distinguishing features, so that RIScan can extract stable and accurate direction information from complex and dynamic environments. RIScan is implemented in the real system with our own developed 16 × 16 RIS prototype and COTS Wi-Fi devices. Extensive experiments show that RIScan achieves a median localization error of 47cm and 71cm in static and dynamic environments with only two RIS anchors. Compared to the SOTA methods, RIScan reduces the localization latency by more than an order of magnitude.
Chenggao Li, Qianyi Huang, Yandao Huang, Qingyong Hu, Huangxun Chen, Qian Zhang 0001
SenSys5
2023 PDAssess: A Privacy-preserving Free-speech based Parkinson's Disease Daily Assessment System
abstract
In-time disease assessment is essential to better customize the medication scheme and improve the quality of life for chronic diseases like Parkinson's disease (PD). Toward the inconvenience problem in current clinical assessment practice, mobile sensing solutions based on detecting Parkinson's vocal changes are proposed. However, current solutions either can only achieve binary disease detection task or require patients to perform specific speaking tasks, which is not effective and practical for disease stage assessment in daily scenario. Moreover, most of existing solutions do not take speech privacy into consideration. In this work, we present PDAssess, a free speech-based daily assessment system that can perform 4-stage Parkinson's disease assessment in a privacy-preserving manner. We observe that current solutions did not fully leverage the rich information embedded in free speech due to the linguistic content variations, and therefore leverage a pre-trained automatic speech recognition (ASR) model to achieve a content variation-aware feature-extraction. In order to distinguish subtle stage-wise differences, we design a novel attention-based neural network architecture with a customized loss function for disease assessment task. Towards the potential privacy leakage problem, we design a Split Learning-based framework with pseudo-labeling and local domain adversarial training to better preserve speech content privacy. We collaborate with a medical center and evaluate the performance of PDAssess on real-world speech data collected from 50 PD subjects and 50 healthy subjects. The evaluation result shows that PDAssess can perform 4-stage PD assessment with an average person-wise F1 score of 89.1% and voice sample-wise F1 score of 75.1%.
Baichen Yang, Qingyong Hu, Wentao Xie 0001, Qian Zhang 0001
SenSys2
2023 V2P-SSD: Single-Stage 3-D Object Detection With Voxel-to-Point Transformation
abstract
We study the problem of efficient object detection in 3-D point clouds with the voxel-point framework. Considering a large number of redundant and dense proposals are usually generated for small-sized objects during inference in voxel-based single-stage detectors, the existing detectors usually introduce extra subnetworks to filter and further refine the redundancy proposals. Albeit feasible, the computational and memory cost also increase during inference. In this letter, we introduce a novel voxel-to-point 3-D detector, termed V2P-SSD, which is a novel and lightweight pipeline that jointly integrates the voxel backbone and point head together in a single-stage framework. Different from dense predictions in feature maps, voxels related to objects in our framework are sampled with a fixed number and then transformed into points. Consequently, the point head is used to dynamically generate object proposals. Our voxel-to-point detection paradigm demonstrates a significant precision improvement on small-sized objects without introducing extra memory footprints. Extensive experiments conducted on KITTI and ONCE benchmarks validate the superiority of our method.
Yifan Zhang 0030, Qingyong Hu, Ke Xu 0013, Jianwei Wan, Yulan Guo
IEEE Geosci. Remote. Sens. Lett.2
2023 You Only Train Once: Learning General and Distinctive 3D Local Descriptors
abstract
Extracting distinctive, robust, and general 3D local features is essential to downstream tasks such as point cloud registration. However, existing methods either rely on noise-sensitive handcrafted features, or depend on rotation-variant neural architectures. It remains challenging to learn robust and general local feature descriptors for surface matching. In this paper, we propose a new, simple yet effective neural network, termed SpinNet, to extract local surface descriptors which are rotation-invariant whilst sufficiently distinctive and general. A Spatial Point Transformer is first introduced to embed the input local surface into an elaborate cylindrical representation (SO(2) rotation-equivariant), further enabling end-to-end optimization of the entire framework. A Neural Feature Extractor, composed of point-based and 3D cylindrical convolutional layers, is then presented to learn representative and general geometric patterns. An invariant layer is finally used to generate rotation-invariant feature descriptors. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities.
Sheng Ao, Yulan Guo, Qingyong Hu, Bo Yang 0027, Andrew Markham, Zengping Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 RoReg: Pairwise Point Cloud Registration With Oriented Descriptors and Local Rotations
abstract
We present RoReg, a novel point cloud registration framework that fully exploits oriented descriptors and estimated local rotations in the whole registration pipeline. Previous methods mainly focus on extracting rotation-invariant descriptors for registration but unanimously neglect the orientations of descriptors. In this paper, we show that the oriented descriptors and the estimated local rotations are very useful in the whole registration pipeline, including feature description, feature detection, feature matching, and transformation estimation. Consequently, we design a novel oriented descriptor RoReg-Desc and apply RoReg-Desc to estimate the local rotations. Such estimated local rotations enable us to develop a rotation-guided detector, a rotation coherence matcher, and a one-shot-estimation RANSAC, all of which greatly improve the registration performance. Extensive experiments demonstrate that RoReg achieves state-of-the-art performance on the widely-used 3DMatch and 3DLoMatch datasets, and also generalizes well to the outdoor ETH dataset. In particular, we also provide in-depth analysis on each component of RoReg, validating the improvements brought by oriented descriptors and the estimated local rotations. Source code and supplementary material are available at https://github.com/HpWang-whu/RoReg.
Haiping Wang 0004, Yuan Liu 0025, Qingyong Hu, Bing Wang 0013, Zhen Dong 0005, Yulan Guo, Wenping Wang 0001, Bisheng Yang
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 CenterTube: Tracking Multiple 3D Objects With 4D Tubelets in Dynamic Point Clouds
abstract
3D Multi-Object Tracking (MOT) in dynamic point cloud sequences is a fundamental research problem for several downstream tasks such as motion planning and action recognition. Existing methods usually rely on the traditional tracking-by-detection (TBD) paradigm, which performs the tracking based on the results achieved by dedicated detectors. However, this two-stage framework usually cannot sufficiently exploit spatial-temporal information and end-to-end optimization, leading to sub-optimal tracking performance, especially when the object is partially or completely occluded. In this paper, we propose a joint detection and tracking framework namedCenterTubefor dynamic point cloud sequences. The key to our approach is to formulate the problem of multiple object trajectory predictions as 4D tubelet detections. In particular, the proposed CenterTube is composed of three head branches, including a center branch, a regression branch, and a movement branch for the estimation of object center, object size, instance movement, and frame interval, respectively. Additionally, a Tube BEV-IoU (TB-IoU) is also presented to link the generated clip-level tubelets and form the final tracks. Extensive experiments conducted on the KITTI-MOT and nuScenes datasets demonstrate that our model achieves competitive performances even if no ready-made detection results is adopted.
Hao Liu 0061, Yanni Ma, Qingyong Hu, Yulan Guo
IEEE Trans. Multim.3
2022 STPLS3D: A Large-Scale Synthetic and Real Aerial Photogrammetry 3D Point Cloud Dataset
Meida Chen, Qingyong Hu, Zifan Yu, Hugues Thomas, Andrew Feng, Kyle McCullough, Fengbo Ren, Lucio Soibelman
BMVC2
2022 3DAC: Learning Attribute Compression for Point Clouds
abstract
We study the problem of attribute compression for large-scale unstructured 3D point clouds. Through an in-depth exploration of the relationships between different encoding steps and different attribute channels, we introduce a deep compression network, termed 3DAC, to explicitly compress the attributes of 3D point clouds and reduce storage usage in this paper. Specifically, the point cloud attributes such as color and reflectance are firstly converted to transform coefficients. We then propose a deep entropy model to model the probabilities of these coefficients by considering information hidden in attribute transforms and previous encoded attributes. Finally, the estimated probabilities are used to further compress these transform coefficients to a final attributes bitstream. Extensive experiments conducted on both indoor and outdoor large-scale open point cloud datasets, including ScanNet and SemanticKITTI, demonstrated the superior compression rates and reconstruction quality of the proposed method.
Guangchi Fang, Qingyong Hu, Hanyun Wang, Yiling Xu, Yulan Guo
CVPR2
2022 Not All Points Are Equal: Learning Highly Efficient Point-based Detectors for 3D LiDAR Point Clouds
abstract
We study the problem of efficient object detection of 3D LiDAR point clouds. To reduce the memory and computational cost, existing point-based pipelines usually adopt task-agnostic random sampling or farthest point sampling to progressively downsample input point clouds, despite the fact that not all points are equally important to the task of object detection. In particular, the foreground points are inherently more important than background points for object detectors. Motivated by this, we propose a highly-efficient single-stage point-based 3D detector in this paper, termed IA-SSD. The key of our approach is to exploit two learnable, task-oriented, instance-aware downsampling strategies to hierarchically select the foreground points belonging to objects of interest. Additionally, we also introduce a contextual centroid perception module to further estimate precise instance centers. Finally, we build our IA-SSD following the encoder-only architecture for efficiency. Extensive experiments conducted on several large-scale detection benchmarks demonstrate the competitive performance of our IA-SSD. Thanks to the low memory footprint and a high degree of parallelism, it achieves a superior speed of 80+ frames-per-second on the KITTI dataset with a single RTX2080Ti GPU. The code is available at https://github.com/yifanzhang713/IA-SSD.
Yifan Zhang 0030, Qingyong Hu, Guoquan Xu, Yanxin Ma, Jianwei Wan, Yulan Guo
CVPR2
2022 No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with Static Models by Fitting Feature-level Space-time Surfaces
abstract
Scene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspondences. To capture 3D motions without explicitly tracking correspondences, we propose a kinematics-inspired neural network (Kinet) by generalizing the kinematic concept of ST-surfaces to the feature space. By unrolling the normal solver of ST-surfaces in the feature space, Kinet implicitly encodes feature-level dynamics and gains advantages from the use of mature back-bones for static point cloud processing. With only minor changes in network structures and low computing overhead, it is painless to jointly train and deploy our framework with a given static model. Experiments on NvGesture, SHREC'17, MSRAction-3D, and NTU-RGBD demonstrate its efficacy in performance, efficiency in both the number of parameters and computational complexity, as well as its versatility to various static backbones. Noticeably, Kinet achieves the accuracy of 93.27% on MSRAction-3D with only 3.20M parameters and 10.35G FLOPS. The code is available at https://github.com/jx-zhong-for-academic-purpose/Kinet.
Jia-Xing Zhong, Kaichen Zhou, Qingyong Hu, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham
CVPR3
2022 Meta-sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds
Ta Ying Cheng, Qingyong Hu, Qian Xie 0001, Agathoniki Trigoni, Andrew Markham
ECCV (2)2
2022 SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds
Qingyong Hu, Bo Yang 0027, Guangchi Fang, Yulan Guo, Ales Leonardis, Agathoniki Trigoni, Andrew Markham
ECCV (27)1
2022 DevNet: Self-supervised Monocular Depth Learning via Density Volume Construction
Kaichen Zhou, Lanqing Hong, Changhao Chen, Hang Xu 0004, Chaoqiang Ye, Qingyong Hu, Zhenguo Li
ECCV (39)6
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR3
2022 Hierarchical Graph Transformer with Adaptive Node Sampling
abstract
The Transformer architecture has achieved remarkable success in a number of domains including natural language processing and computer vision. However, when it comes to graph-structured data, transformers have not achieved competitive performance, especially on large graphs. In this paper, we identify the main deficiencies of current graph transformers: (1) Existing node sampling strategies in Graph Transformers are agnostic to the graph characteristics and the training process. (2) Most sampling strategies only focus on local neighbors and neglect the long-range dependencies in the graph. We conduct experimental investigations on synthetic datasets to show that existing sampling strategies are sub-optimal. To tackle the aforementioned problems, we formulate the optimization strategies of node sampling in Graph Transformer as an adversary bandit problem, where the rewards are related to the attention weights and can vary in the training procedure. Meanwhile, we propose a hierarchical attention scheme with graph coarsening to capture the long-range interactions while reducing computational complexity. Finally, we conduct extensive experiments on real-world datasets to demonstrate the superiority of our method over existing graph transformers and popular GNNs.
Zaixi Zhang, Qi Liu 0003, Qingyong Hu, Cheekong Lee
NeurIPS3
2022 SensatUrban: Learning Semantics from Urban-Scale Photogrammetric Point Clouds
Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham
Int. J. Comput. Vis.1
2022 Continuous Mapping Convolution for Large-Scale Point Clouds Semantic Segmentation
abstract
In this letter, we introduce MappingConvSeg, a continuous convolution network for semantic segmentation of large-scale point clouds. In particular, a conceptually simple, end-to-end learnable, and continuous convolution operator is proposed for learning spatial correlation of unstructured 3-D point clouds. For each local point set, the unstructured point features are first mapped onto a series of learned kernel points based on the spatial relationship, and the continuous convolution is then applied to capture specific local geometrical patterns. Taking the proposed mapping convolution operation as the building block, a hierarchical network is then built for large-scale point cloud semantic segmentation. Experimental results conducted on two public benchmarks, including Toronto-3D and Stanford large-scale 3-D Indoor Spaces (S3DIS) dataset, demonstrate the superiority of the proposed method.
Kunping Yan, Qingyong Hu, Hanyun Wang, Li Li 0100, Song Ji
IEEE Geosci. Remote. Sens. Lett.2
2022 Learning Semantic Segmentation of Large-Scale Point Clouds With Random Sampling
abstract
We study the problem of efficient semantic segmentation of large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Comparative experiments show that our RandLA-Net can process 1 million points in a single pass up to 200× faster than existing approaches. Moreover, extensive experiments on five large-scale point cloud datasets, including Semantic3D, SemanticKITTI, Toronto3D, NPM3D and S3DIS, demonstrate the state-of-the-art semantic segmentation performance of our RandLA-Net.
Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Detecting and Tracking Small and Dense Moving Objects in Satellite Videos: A Benchmark
abstract
Satellite video cameras can provide continuous observation for a large-scale area, which is important for many remote sensing applications. However, achieving moving object detection and tracking in satellite videos remains challenging due to the insufficient appearance information of objects and lack of high-quality datasets. In this article, we first build a large-scale satellite video dataset with rich annotations for the task of moving object detection and tracking. This dataset is collected by the Jilin-1 satellite constellation and composed of 47 high-quality videos with 1 646 038 instances of interest for object detection and 3711 trajectories for object tracking. We then introduce a motion modeling baseline to improve the detection rate and reduce false alarms based on accumulative multiframe differencing and robust matrix completion. Finally, we establish the first public benchmark for moving object detection and tracking in satellite videos and extensively evaluate the performance of several representative approaches on our dataset. Comprehensive experimental analyses and insightful conclusions are also provided. The dataset is available athttps://github.com/QingyongHu/VISO.
Qingyong Hu, Hao Liu 0061, Feng Zhang 0046, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo
IEEE Trans. Geosci. Remote. Sens.2
2021 SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration
abstract
Extracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor representative. In this paper, we introduce a new, yet conceptually simple, neural architecture, termed SpinNet, to extract local features which are rotationally invariant whilst sufficiently informative to enable accurate registration. A Spatial Point Transformer is first introduced to map the input local surface into a carefully designed cylindrical space, enabling end-to-end optimization with SO(2) equivariant representation. A Neural Feature Extractor which leverages the powerful point-based and 3D cylindrical convolutional neural layers is then utilized to derive a compact and representative descriptor for matching. Extensive experiments on both indoor and outdoor datasets demonstrate that SpinNet outperforms existing state-of-the-art techniques by a large margin. More critically, it has the best generalization ability across unseen scenarios with different sensor modalities. The code is available at https://github.com/QingyongHu/SpinNet.
Sheng Ao, Qingyong Hu, Bo Yang 0027, Andrew Markham, Yulan Guo
CVPR2
2021 Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
abstract
An essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly available datasets are either in relative small spatial scales or have limited semantic annotations due to the expensive cost of data acquisition and data annotation, which severely limits the development of fine-grained semantic understanding in the context of 3D point clouds. In this paper, we present an urban-scale photogrammetric point cloud dataset with nearly three billion richly annotated points, which is three times the number of labeled points than the existing largest photogrammetric point cloud dataset. Our dataset consists of large areas from three UK cities, covering about 7.6 km2of the city landscape. In the dataset, each 3D point is labeled as one of 13 semantic classes. We extensively evaluate the performance of state-of-the-art algorithms on our dataset and provide a comprehensive analysis of the results. In particular, we identify several key challenges towards urban-scale point cloud understanding. The dataset is available at https://github.com/QingyongHu/SensatUrban.
Qingyong Hu, Bo Yang 0027, Sheikh Khalid, Agathoniki Trigoni, Andrew Markham
CVPR1
2021 Cgan-Net: Class-Guided Asymmetric Non-Local Network for Real-Time Semantic Segmentation
abstract
By introducing various non-local blocks to capture the long-range dependencies, remarkable progress has been achieved in semantic segmentation recently. However, the improvement in segmentation accuracy usually comes at the price of significant reductions in network efficiency, as non-local block usually requires expensive computation and memory cost for dense pixel-to-pixel correlation. In this paper, we introduce a Class-Guided Asymmetric Non-local Network (CGAN-Net) to enhance the class-discriminability in learned feature map, while maintaining real-time efficiency. The key to our approach is to calculate the dense similarity matrix in coarse semantic prediction maps, instead of the high-dimensional latent feature map. This is not only computationally and memory efficient, but helps to learn query-dependent global context. Experiments conducted on Cityscape and CamVid demonstrate the compelling performance of our CGAN-Net. In particular, our network achieves 76.8% mean IoU on the Cityscapes test set with a speed of 38 FPS for 1024×2048 images on a single Tesla V100 GPU.
Qingyong Hu, Jun-Gang Yang, Yulan Guo
ICASSP2
2021 Deep Learning for 3D Point Clouds: A Survey
abstract
Point cloud learning has lately attracted increasing attention due to its wide applications in many areas, such as computer vision, autonomous driving, and robotics. As a dominating technique in AI, deep learning has been successfully used to solve various 2D vision problems. However, deep learning on point clouds is still in its infancy due to the unique challenges faced by the processing of point clouds with deep neural networks. Recently, deep learning on point clouds has become even thriving, with numerous methods being proposed to address different problems in this area. To stimulate future research, this paper presents a comprehensive review of recent progress in deep learning methods for point clouds. It covers three major tasks, including 3D shape classification, 3D object detection and tracking, and 3D point cloud segmentation. It also presents comparative results on several publicly available datasets, together with insightful observations and inspiring future research directions.
Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu 0061, Li Liu 0002, Mohammed Bennamoun
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao
BMVC2
2020 RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
abstract
We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we introduce RandLA-Net, an efficient and lightweight neural architecture to directly infer per-point semantics for large-scale point clouds. The key to our approach is to use random point sampling instead of more complex point selection approaches. Although remarkably computation and memory efficient, random sampling can discard key features by chance. To overcome this, we introduce a novel local feature aggregation module to progressively increase the receptive field for each 3D point, thereby effectively preserving geometric details. Extensive experiments show that our RandLA-Net can process 1 million points in a single pass with up to 200x faster than existing approaches. Moreover, our RandLA-Net clearly surpasses state-of-the-art approaches for semantic segmentation on two large-scale benchmarks Semantic3D and SemanticKITTI.
Qingyong Hu, Bo Yang 0027, Linhai Xie, Stefano Rosa, Yulan Guo, Zhihua Wang 0005, Agathoniki Trigoni, Andrew Markham
CVPR1
2019 Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds
abstract
We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cloud, while simultaneously predicting a point-level mask for each instance. It consists of a backbone network followed by two parallel network branches for 1) bounding box regression and 2) point mask prediction. 3D-BoNet is single-stage, anchor-free and end-to-end trainable. Moreover, it is remarkably computationally efficient as, unlike existing approaches, it does not require any post-processing steps such as non-maximum suppression, feature sampling, clustering or voting. Extensive experiments show that our approach surpasses existing work on both ScanNet and S3DIS datasets while being approximately 10x more computationally efficient. Comprehensive ablation studies demonstrate the effectiveness of our design.
Bo Yang 0027, Ronald Clark, Qingyong Hu, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni
NeurIPS4
2018 Long-term Object Tracking with Instance Specific Proposals
abstract
Correlation filter based trackers have been extensively investigated for their superior efficiency and fairly good robustness. However, it remains challenging to achieve longterm tracking when the object is under occlusion and severe deformation. In this paper, we propose a tracker named Complementary Learners with Instance-specific Proposals (CLIP). The CLIP tracker consists of three main components, including a translation filter, a scale filter, and an error correction module. Complementary features are incorporated into the translation filter to cope with illumination changes and deformation, and an adaptive updating mechanism is proposed to prevent model corruption. The translation filter aims to provide an excellent real-time inference. Furthermore, the error correction module is activated to correct the localization error by an instance-specific proposal generator, especially when the target suffers from dramatic appearance changes. Experimental results on the OTB, Temple-Color 128 and UAV20L datasets demonstrate that the CLIP tracker performs favorably against existing competitive trackers in term of accuracy and robustness. Moreover, our proposed CLIP tracker runs at the speed of 33 fps on the OTB. It is highly suitable for real-time applications.
Hao Liu 0061, Qingyong Hu, Yulan Guo
ICPR2
2018 Semi-Online Multiple Object Tracking Using Graphical Tracklet Association
abstract
Online multiple object tracking (MOT) is highly challenging when multiple objects have similar appearance or under long occlusion. In this letter, we propose a semi-online MOT method using online discriminative appearance learning and tracklet association with a sliding window. We connect similar detections of neighboring frames in a temporal window, and improve the performance of appearance feature by online discriminative appearance learning. Then, tracklet association is performed by minimizing a subgraph decomposition cost. Occlusions and missing detections are recovered after tracklet stitching. Our method has been tested on two public datasets. Experimental results have demonstrated the significant performance improvement of our method. Specifically, the proposed method is improved by 8.31% and 12.38% in terms of Multiple Object Tracking Accuracy and Multiple Object Tracking Precision, respectively, as compared to the baseline.
Yulan Guo, Xing Tang 0003, Qingyong Hu, Wei An 0003
IEEE Signal Process. Lett.4
2017 Correlation Filter Tracking: Beyond an Open-loop System
Qingyong Hu, Yulan Guo, Yunjin Chen, Wei An 0003
BMVC1