EDBT 2026 Demo / reviewers in the wild / expert
Qingchao Chen
dblp:123/9213
· DBLP profile ↗
54ranked-venue papers
5as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 28 since 2021Artificial intelligence and machine learning · 31 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward brain magnetic resonance imaging analysis intelligence: A review of federated learning and visual foundation models
Qingchao Chen |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Large-Scale Pre-Trained Models Empowering Phrase Generalization in Temporal Sentence Localization
Yang Liu 0105, Minghang Zheng, Qingchao Chen, Shaogang Gong, Yuxin Peng 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Contactless OSAHS Respiration and Sleep Stage Classification Using Monitored Snoring Events With Heterogeneous Multiscale Selective Distillation for Internet of Medical Things
Shaoxing Zhang, Yang Liu 0105, Qingchao Chen |
IEEE Internet Things J. | 4 |
| 2026 | Confidence-Aware Pseudo-Label Self-Correction for Weakly Supervised Visual GroundingabstractWeakly supervised visual grounding aims to locate a region in an image based on an input query sentence, without access to the mapping between image regions and queries during training. Current methods treat spatial grounding as an object retrieval task, relying on cross-modal similarity scores for proposal selection. However, they fail to address model overfitting caused by unreliable cross-modal similarity scores. To overcome this, we first propose the Confidence-aware Pseudo-label Learning (CPL) framework. CPL first generates diverse pseudo queries for region proposals, and then establishes reliable associations for model training based on the uni-modal similarity score. Secondly, we propose a cross-modal verification module based on the pretrained vision-language model to verify associations. However, the verification module is isolated from the grounding model, so it can only assess associations in a static manner, but not correct the suspicious ones. Finally, we introduce CPL++ to make two-fold improvements. For one thing, we upgrade the verification process based on the model's grounding loss value to identify suspicious associations dynamically and selectively leverage them in the training. For another, we propose a self-supervised association correction module to rectify suspicious associations, thereby mitigating the risk of error propagation. Experimental results on five datasets demonstrate the superiority of our approach. Yang Liu 0105, Zijing Zhao 0004, Qingchao Chen, Yuxin Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Joint Class-level and Instance-level Relationship Modeling for Novel Class DiscoveryabstractNovel class discovery(NCD) aims to cluster the unlabeled data with the help of a labeled set containing different but related classes. The key to solving NCD is the knowledge transfer between labeled and unlabeled sets.Since NCD requires that known classes and unknown classes are related, it is significant to explore class-level relationships between known and unknown for more effective knowledge transfer. However, most existing methods either facilitate knowledge transfer by learning a shared representation space or by modeling coarse-grained or asymmetric relationships between known and unknown, neglecting class-level relationships. To tackle these challenges, we propose a symmetric class-to-class relationship modeling and knowledge transfer method, achieving bidirectional knowledge transfer at class-level. Considering that class-level modeling often overlooks the subtle distinctions between samples, we propose pairwise similarity-based relationship modeling and consistency constraint for instance-level knowledge transfer. Extensive experiments on CIFAR100 and three fine-grained datasets demonstrate that our method achieves significant performance improvements compared to state-of-the-art methods. Jiaying Zhou, Qingchao Chen |
AAAI | 2 |
| 2025 | Radar2ECG: Multi-Scale Bottleneck Fusion and Cross-modal Semantic Distillation for Conditional Electrocardiogram Generation from Radar Heart SoundabstractThe field of conditional Electrocardiogram(ECG) generation focuses on generating specified ECGs under given conditions for medical purposes. Existing methods are typically based on conditions of simple inputs like text or lead types. However, they struggle to handle the complexity of radar heart sound signals due to the lack of effective feature extraction, which hinders capturing the intricate waveform correlations between radar heart sounds and ECGs. Considering that radar-detected heart sound signals are contactless, the application is of essential value in a real-world deployment like sleep scenarios. Moreover, no prior approaches have addressed this specific task. To tackle this challenge, we propose a novel multi-scale feature fusion network framework, Radar2ECG. This model leverages pre-trained autoencoders for heart sound and ECG signals, aligning and integrating multi-layer features through a bottleneck structure to enhance receptive fields and reduce redundant features, thereby capturing the correlations between heart sounds and ECGs. Finally, we employ knowledge distillation to transfer knowledge from the ECG decoder to the heart sound decoder. We present three anomaly type datasets and extensive experiments conducted on both normal and abnormal datasets demonstrate that our method outperforms existing models in both accuracy and robustness. The multi-scale feature fusion significantly improves performance, showcasing strong potential in ECG generation and heart sound anomaly detection tasks. Jinye Li, Aidong Men, Yang Liu 0105, Pengda Han, Qingchao Chen |
ICASSP | 5 |
| 2025 | Open-Vocabulary Hoi Detection With Interaction-Aware Prompt and Concept Calibration
Ting Lei 0001, Shaofeng Yin, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 3 |
| 2025 | TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge Transferring
Ting Lei 0001, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 5 |
| 2025 | CubeDN: Real-Time Drone Detection in 3D Space from Dual mmWave Radar CubesabstractAs drone use has become more widespread, there is a critical need to ensure safety and security. A key element of this is robust and accurate drone detection and localization. While cameras and other optical sensors like LiDAR are commonly used for object detection, their performance degrades under adverse lighting and environmental conditions. Therefore, this has generated interest in finding more reliable alternatives, such as millimeter-wave (mmWave) radar. Recent research on mmWave radar object detection has predominantly focused on 2D detection of road users. Although these systems demonstrate excellent performance for 2D problems, they lack the sensing capability to measure elevation, which is essential for 3D drone detection. To address this gap, we propose CubeDN, a single-stage end-to-end radar object detection network specifically designed for flying drones. CubeDN overcomes challenges such as poor elevation resolution by utilizing a dual radar configuration and a novel deep learning pipeline. It simultaneously detects, localizes, and classifies drones of two sizes, achieving decimeter-level tracking accuracy at closer ranges with overall 95% average precision (AP) and 85% average recall (AR). Furthermore, CubeDN completes data processing and inference at 10Hz, making it highly suitable for practical applications. Fangzhan Shi, Xijia Wei, Qingchao Chen, Kevin Chetty, Simon J. Julier |
ICRA | 4 |
| 2025 | BayeSMM: Robust Deep Combined Computing Tackling Heavy-Tailed Distribution in Medical Images
Yuanye Liu, Ruoxuan Zhen, Shangqi Gao, Xinzhe Luo, Qingchao Chen, Xiahai Zhuang |
MICCAI (13) | 6 |
| 2025 | Investigating Domain Gaps for Indoor 3D Object DetectionabstractAs a fundamental task for indoor scene understanding, 3D object detection has been extensively studied, and the accuracy on indoor point cloud data has been substantially improved. However, existing researches have been conducted on limited datasets, where the training and testing sets share the same distribution. In this paper, we consider the task of adapting indoor 3D object detectors from one dataset to another, presenting a comprehensive benchmark with ScanNet, SUN RGB-D and 3D Front datasets, as well as our newly proposed large-scale datasets ProcTHOR-OD and ProcFront generated by a 3D simulator. Since indoor point cloud datasets are collected and constructed in different ways, the object detectors are likely to overfit to specific factors within each dataset, such as point cloud quality, bounding box layout and instance features. We conduct experiments across datasets on different adaptation scenarios including synthetic-to-real adaptation, point cloud quality adaptation, layout adaptation and instance feature adaptation, analyzing the impact of different domain gaps on 3D object detectors. We also introduce several approaches to improve adaptation performances, providing baselines for domain adaptive indoor 3D object detection, hoping that future works may propose detectors with stronger generalization ability across domains. Our project homepage can be found in https://jeremyzhao1998.github.io/DAVoteNet-release/. Zijing Zhao 0004, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ACM Multimedia | 3 |
| 2025 | Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Jiayi Gao, Changcheng Hua, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ACM Multimedia | 3 |
| 2025 | Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
Wentao Mo, Qingchao Chen, Yuxin Peng 0001, Siyuan Huang 0001, Yang Liu 0105 |
ACM Multimedia | 2 |
| 2025 | IR-Based Sleep Monitoring: Movement, Respiration, and Sleep Staging With a Structure-Aware Cross-Modal EEG Knowledge DistillationabstractIt is inevitably crucial to classify the sleep stage for assessing sleep quality and diagnosing related diseases. However, existing automated diagnosis methods mostly adopt the goldstandard Electroencephalogram (EEG) or other sensing signals from the PolySomnoGraphy (PSG) machine in hospital, which are expensive, importable, and therefore unsuitable for point-ofcare monitoring at home. To enable the sleep stage monitoring at home, in this paper, we integrate the external visual information from infrared (IR) videos with internal bio-electrical knowledge from EEG signals, and propose a novel IR-based sleep monitoring system in IoT scenarios: which can non-contact monitor and predict body movement, respiration rate, and sleep staging tasks based solely on IR video. It is different from previous video classification and multi-modal analysis systems, mainly in that (1) the temporal duration of the IR video is relatively long, reaching 10 hours per night. (2) the semantic gap between the EEG signal and IR video is disparate and much larger than conventional cross-modal data in multimedia analysis such as video and audio. To establish a solid cross-modal benchmark in sleep monitoring, we develop a new dataset (S3IE), which is a large-scale dataset including synchronized IR videos and EEG signals. We also propose a novel cross-modal knowledge distillation method, namely the structure-aware contrastive distillation (SACD) to distill the EEG bio-electrical knowledge to IR visual features. Our SACD achieves the SOTA performances on both our S3IE and other cross-modal distillation benchmarks. We expect to raise more attention and promote more developments for the sleep and, more importantly, the cross-modal distillation from clinical signal/media to conventional media in IoT scenarios. Shaoxing Zhang, Yang Liu 0105, Qingchao Chen |
IEEE Internet Things J. | 4 |
| 2025 | scDD: scRNA-seq dataset distillation in latent codes with single-step conditional diffusion generator
Qingchao Chen |
Knowl. Based Syst. | 4 |
| 2025 | MERIT: Multi-view evidential learning for reliable and interpretable liver fibrosis staging
Yuanye Liu, Zheyao Gao, Nannan Shi, Fuping Wu, Qingchao Chen, Xiahai Zhuang |
Medical Image Anal. | 6 |
| 2025 | Selection, Ensemble, and Adaptation: Advancing Multi-Source-Free Domain Adaptation via Architecture ZooabstractConventional Multi-Source Free Domain Adaptation (MSFDA) assumes that each source domain provides a single source model, and all source models adopt a uniform architecture. This paper introduces Zoo-MSFDA, a more general setting that allows each source domain to offer a zoo of multiple source models with different architectures. While it enriches the source knowledge, Zoo-MSFDA risks being dominated by suboptimal/harmful models. To address this issue, we theoretically analyze the model selection problem in Zoo-MSFDA, and introduce two principles: transferability principle and diversity principle. Recognizing the challenge of measuring transferability, we subsequently propose a novel Source-Free Unsupervised Transferability Estimation (SUTE). It enables assessing and comparing transferability across multiple source models with different architectures under domain shift, without requiring target labels and source data. Based on above, we introduce a Selection, Ensemble, and Adaptation (SEA) framework to address Zoo-MSFDA, which consists of: 1) source models selection based on the proposed principles and SUTE; 2) ensemble construction based on SUTE-estimated transferability; 3) target-domain adaptation of the ensemble model. Evaluations demonstrate that our SEA framework, with the introduced Zoo-MSFDA setting, significantly improves adaptation performance in 2D image classification tasks. Additionally, our SUTE achieves state-of-the-art performance in transferability estimation. Jiangbo Pei, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Incorporating Pre-Training Data Matters in Unsupervised Domain AdaptationabstractIn deep learning, initializing models with pre-trained weights has become the de facto practice for various downstream tasks. Many unsupervised domain adaptation (UDA) methods typically adopt a backbone pre-trained on ImageNet, and focus on reducing the source-target domain discrepancy. However, the impact of pre-training on adaptation received little attention. In this study, we delve into UDA from the novel perspective of pre-training. We first demonstrate the impact of pre-training by analyzing the dynamic distribution discrepancies between pre-training data domain and the source/ target domain during adaptation. Then, we reveal that the target error also stems from the pre-training in the following two factors: 1) empirically, target error arises from the gradually degenerative pre-trained knowledge during adaptation; 2) theoretically, the error bound depends on difference between the gradient of loss function, i.e., on the target domain and pre-training data domain. To address these two issues, we redefine UDA as a three-domain problem, i.e., source domain, target domain, and pre-training data domain; then we propose a novel framework, named TriDA. We maintain the pre-trained knowledge and improve the error bound by incorporating pre-training data into adaptation for both vanilla UDA and source-free UDA scenarios. For efficiency, we introduce a selection strategy for pre-training data, and offer a solution with synthesized images when pre-training data is unavailable during adaptation. Notably, TriDA is effective even with a small amount of pre-training or synthesized images, and seamlessly complements the two scenario UDA methods, demonstrating state-of-the-art performance across multiple benchmarks. We hope our work provides new insights for better understanding and application of domain adaptation. Yinsong Xu 0002, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Semantic-Guided Novel Category DiscoveryabstractThe Novel Category Discovery problem aims to cluster an unlabeled set with the help of a labeled set consisting of disjoint but related classes. However, existing models treat class names as discrete one-hot labels and ignore the semantic understanding of these classes. In this paper, we propose a new setting named Semantic-guided Novel Category Discovery (SNCD), which requires the model to not only cluster the unlabeled images but also semantically recognize these images based on a set of their class names. The first challenge we confront pertains to effectively leveraging the class names of unlabeled images, given the inherent gap between the visual and linguistic domains. To address this issue, we incorporate a semantic-aware recognition mechanism. This is achieved by constructing dynamic class-wise visual prototypes as well as a semantic similarity matrix that enables the projection of visual features into the semantic space. The second challenge originates from the granularity disparity between the classification and clustering tasks. To deal with this, we develop a semantic-aware clustering process to facilitate the exchange of knowledge between the two tasks. Through extensive experiments, we demonstrate the mutual benefits of the recognition and clustering tasks, which can be jointly optimized. Experimental results on multiple datasets confirm the effectiveness of our proposed method. Our code is available at https://github.com/wang-weishuai/Semantic-guided-NCD. Weishuai Wang, Ting Lei 0001, Qingchao Chen, Yang Liu 0105 |
AAAI | 3 |
| 2024 | Novel Class Discovery in Chest X-rays via Paired Images and TextabstractNovel class discover(NCD) aims to identify new classes undefined during model training phase with the help of knowledge of known classes. Many methods have been proposed and notably boosted performance of NCD in natural images. However, there has been no work done in discovering new classes based on medical images and disease categories, which is crucial for understanding and diagnosing specific diseases. Moreover, most of the existing methods only utilize information from image modality and use labels as the only supervisory information. In this paper, we propose a multi-modal novel class discovery method based on paired images and text, inspired by the low classification accuracy of chest X-ray images and the relatively higher accuracy of the paired text. Specifically, we first pretrain the image encoder and text encoder with multi-modal contrastive learning on the entire dataset and then we generate pseudo-labels separately on the image branch and text branch. We utilize intra-modal consistency to assess the quality of pseudo-labels and adjust the weights of the pseudo-labels from both branches to generate the ultimate pseudo-labels for training. Experiments on eight subset splits of MIMIC-CXR-JPG dataset show that our method improves the clustering performance of unlabeled classes by about 10% on average compared to state-of-the-art methods. Code is available at: https://github.com/zzzzzzzzjy/MMNCD-main. Jiaying Zhou, Yang Liu 0105, Qingchao Chen |
AAAI | 3 |
| 2024 | OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationabstractDynamic Scene Graph Generation (DSGG) focuses on identifying visual relationships within the spatial-temporal domain of videos. Conventional approaches often employ multi-stage pipelines, which typically consist of object detection, temporal association, and multi-relation classification. However, these methods exhibit inherent limitations due to the separation of multiple stages, and independent optimization of these sub-problems may yield sub-optimal solutions. To remedy these limitations, we propose a one-stage end-to-end framework, termed OED, which streamlines the DSGG pipeline. This framework reformulates the task as a set prediction problem and leverages pairwise features to represent each subject-object pair within the scene graph. Moreover, another challenge of DSGG is capturing temporal dependencies, we introduce a Progressively Refined Module (PRM) for aggregating temporal context without the constraints of additional trackers or handcrafted trajectories, enabling end-to-end optimization of the network. Extensive experiments conducted on the Action Genome benchmark demonstrate the effectiveness of our design. The code and models are available at https://github.com/guanw-pku/OED. Qingchao Chen, Yang Liu 0105 |
CVPR | 3 |
| 2024 | Training-Free Video Temporal Grounding Using Large-Scale Pre-trained Models
Minghang Zheng, Xinhao Cai, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ECCV (82) | 3 |
| 2024 | Semantic-Aware Human Object Interaction Image GenerationabstractRecent text-to-image generative models have demonstrated remarkable abilities in generating realistic images. Despite their great success, these models struggle to generate high-fidelity images with prompts oriented toward human-object interaction (HOI). The difficulty in HOI generation arises from two aspects. Firstly, the complexity and diversity of human poses challenge plausible human generation. Furthermore, untrustworthy generation of interaction boundary regions may lead to deficiency in HOI semantics. To tackle the problems, we propose a Semantic-Aware HOI generation framework SA-HOI . It utilizes human pose quality and interaction boundary region information as guidance for denoising process, thereby encouraging refinement in these regions to produce more reasonable HOI images. Based on it, we establish an iterative inversion and image refinement pipeline to continually enhance generation quality. Further, we introduce a comprehensive benchmark for HOI generation, which comprises a dataset involving diverse and fine-grained HOI categories, along with multiple custom-tailored evaluation metrics for HOI generation. Experiments demonstrate that our method significantly improves generation quality under both HOI-specific and conventional image evaluation metrics. The code is available at https://github.com/XZPKU/SA-HOI.git Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ICML | 2 |
| 2024 | 3D Vision and Language Pretraining with Large-Scale Synthetic Data
Dejie Yang, Wentao Mo, Qingchao Chen, Siyuan Huang 0001, Yang Liu 0105 |
IJCAI | 4 |
| 2024 | Poisson Ordinal Network for Gleason Group Estimation Using Bi-Parametric MRI
Yinsong Xu 0002, Ziyi Shen, Iani J. M. B. Gayo, Natasha Thorley, Shonit Punwani, Aidong Men, Dean C. Barratt, Qingchao Chen, Yipeng Hu |
MICCAI (5) | 9 |
| 2024 | ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual GroundingabstractVisual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objects within multiple-instance distractions (multiple objects of the same category as the target) remains a significant challenge. Existing methods demonstrate a significant performance drop when there are multiple distractions in an image, indicating an insufficient understanding of the fine-grained semantics and spatial relationships between objects. In this paper, we propose a novel approach, the Relation and Semantic-sensitive Visual Grounding (ResVG) model, to address this issue. Firstly, we enhance the model's understanding of fine-grained semantics by injecting semantic prior information derived from text queries into the model. This is achieved by leveraging text-to-image generation models to produce images representing the semantic attributes of target objects described in queries. Secondly, we tackle the lack of training samples with multiple distractions by introducing a relation-sensitive data augmentation method. This method generates additional training data by synthesizing images containing multiple objects of the same category and pseudo queries based on their spatial relationships. The proposed ReSVG model significantly improves the model's ability to comprehend both object semantics and spatial relations, leading to enhanced performance in visual grounding tasks, particularly in scenarios with multiple-instance distractions. We conduct extensive experiments to validate the effectiveness of our methods on five datasets. Code is available at https://github.com/minghangz/ResVG. Minghang Zheng, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
ACM Multimedia | 3 |
| 2024 | IoT-V2E: An Uncertainty-Aware Cross-Modal Hashing Retrieval Between Infrared-Videos and EEGs for Automated Sleep State AnalysisabstractEstimating and monitoring the sleep states at home using ubiquitous infrared (IR) visual camera sensors is an essential healthcare problem. Currently, the common challenge of using IoT sensors to predict sleep stages is the “semantic gap” between the IoT sensory signals and the medical signals, where fewer correlations between IoT sensory signals and the sleep stage labels are observed. To bridge this gap, we propose a novel systematic and methodological IoT design (IoT-V2E) to retrieve the most similar electroencephalogram signal representations in a database given an IR visual query for sleep-related analysis. Specifically, we make the following specific contributions: 1) we collect a cross-modal retrieval data set, including the IR sensory signals and the synchronized Polysomnography signals with sleep stage ground-truth annotations; 2) we propose a novel uncertainty-aware hashing retrieval method, presenting superior performances, sufficient interpretability, and high memory efficiency; 3) our method achieves the state-of-the-art sleep stage retrieval results and provides the uncertainty for each query in the inference; and 4) most importantly, our system is evaluated to be able to assist the physicians not only in diagnosing sleep-related diseases but also finding the subjects with the most similar sleep patterns. Our project is available athttps://github.com/SPIresearch/IoT-V2E. Aidong Men, Yang Liu 0105, Ziming Yao, Shaoxing Zhang, Qingchao Chen |
IEEE Internet Things J. | 7 |
| 2024 | Radar Can See and Hear as Well: A New Multimodal Benchmark Based on Radar SensingabstractRadar technology has emerged as a pivotal component for various applications within the Internet of Things (IoT). To promote the understanding and integration of radar sensing in developing multi-modal applications, we introduce the RAdar Can sEe and heaR (RACER) dataset. This dataset encompasses synchronized radar sensing, audio, and visual data. Radar, with its capability to detect vocal cord vibrations and lip movements, addresses scenarios where conventional microphone and camera setups may falter, such as through-wall or non-line-of-sight sensing. Specifically, the radar discerns and characterizes human lip and vocal cord movements in the range-Doppler domain. We employ deep neural networks to capture the inherent relationships among radar signatures, audio vocal sound, and visual lip movements during human pronunciations. We evaluate the performances of radar sensing using experiments on speech classification, cross-modality retrieval among audio, video, and radar, and cross-modality distillation from video or audio to radar. We summarize the findings and the limitations of using radar sensing in speech-related multi-modal analysis applications. Our codes are available at: https://github.com/SPIresearch/RACER. Yinsong Xu 0002, Qingchao Chen |
IEEE Internet Things J. | 2 |
| 2024 | Evidential Multi-Source-Free Unsupervised Domain AdaptationabstractMulti-Source-Free Unsupervised Domain Adaptation (MSFUDA) requires aggregating knowledge from multiple source models and adapting it to the target domain. Two challenges remain: 1) suboptimal coarse-grained (domain-level) aggregation of multiple source models, and 2) risky semantics propagation based on local structures. In this article, we propose an evidential learning method for MSFUDA, where we formulate two uncertainties, i.e. Evidential Prediction Uncertainty (EPU) and Evidential Adjacency-Consistent Uncertainty (EAU), respectively for addressing the two challenges. The former, EPU, captures the uncertainty of a sample fitted to a source model, which can suggest the preferences of target samples for different source models. Based on this, we develop an EPU-Based Multi-Source Aggregation module to achieve fine-grained, instance-level source knowledge aggregation. The latter, EAU, provides a robust measure of consistency among adjacent samples in the target domain. Utilizing this, we develop an EAU-Guided Local Structure Mining module to ensure the trustworthy propagation of semantics. The two modules are integrated into the Evidential Aggregation and Adaptation Framework (EAAF), and we demonstrated that this framework achieves state-of-the-art performances on three MSFUDA benchmarks. Jiangbo Pei, Aidong Men, Yang Liu 0105, Xiahai Zhuang, Qingchao Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | EviPrompt: A Training-Free Evidential Prompt Generation Method for Adapting Segment Anything Model in Medical ImagesabstractMedical image segmentation is a critical task in clinical applications. Recently, the Segment Anything Model (SAM) has demonstrated potential for natural image segmentation. However, the requirement for expert labour to provide prompts, and the domain gap between natural and medical images pose significant obstacles in adapting SAM to medical images. To overcome these challenges, this paper introduces a novel prompt generation method named EviPrompt. The proposed method requires only a single reference image-annotation pair, making it a training-free solution that significantly reduces the need for extensive labelling and computational resources. First, prompts are automatically generated based on the similarity between features of the reference and target images, and evidential learning is introduced to improve reliability. Then, to mitigate the impact of the domain gap, committee voting and inference-guided in-context learning are employed, generating prompts primarily based on human prior knowledge and reducing reliance on extracted semantic information. EviPrompt represents an efficient and robust approach to medical image segmentation. We evaluate it across a broad range of tasks and modalities, confirming its efficacy. The source code is available at https://github.com/SPIresearch/EviPrompt. Yinsong Xu 0002, Jiaqi Tang 0012, Aidong Men, Qingchao Chen |
IEEE Trans. Image Process. | 4 |
| 2024 | Query-Adaptive Late Fusion for Hierarchical Fine-Grained Video-Text RetrievalabstractRecently, a hierarchical fine-grained fusion mechanism has been proved effective in cross-modal retrieval between videos and texts. Generally, the hierarchical fine-grained semantic representations (video-text semantic matching is decomposed into three levels including global-event representation matching, action-relation representation matching, and local-entity representation matching) to be fused can work well by themselves for the query. However, in real-world scenarios and applications, existing methods failed to adaptively estimate the effectiveness of multiple levels of the semantic representations for a given query in advance of multilevel fusion, resulting in a worse performance than expected. As a result, it is extremely essential to identify the effectiveness of hierarchical semantic representations in a query-adaptive manner. To this end, this article proposes an effective query-adaptive multilevel fusion (QAMF) model based on manipulating multiple similarity scores between the hierarchical visual and text representations. First, we decompose video-side and text-side representations into hierarchical semantic representations consisting of global-event level, action-relation level, and local-entity level, respectively. Then, the multilevel representation of the video-text pair is aligned to calculate the similarity score for each level. Meanwhile, the sorted similarity score curves of the good semantic representation are different from the inferior ones, which exhibit a "cliff" shape and gradually decline (see Fig. fig1 as an example). Finally, we leverage the Gaussian decay function to fit the tail of the score curve and calculate the area under the normalized sorted similarity curve as the indicator of semantic representation effectiveness, namely, the area of good semantic representation is small, and vice versa. Extensive experiments on three public benchmark video-text datasets have demonstrated that our method consistently outperforms the state-of-the-art (SoTA). A simple demo of QAMF will soon be publicly available on our homepage: https://github.com/Lab-ANT. Wentao Ma 0003, Qingchao Chen, Fang Liu 0002, Tongqing Zhou, Zhiping Cai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Phrase-Level Temporal Relationship Mining for Temporal Sentence LocalizationabstractIn this paper, we address the problem of video temporal sentence localization, which aims to localize a target moment from videos according to a given language query. We observe that existing models suffer from a sheer performance drop when dealing with simple phrases contained in the sentence. It reveals the limitation that existing models only capture the annotation bias of the datasets but lack sufficient understanding of the semantic phrases in the query. To address this problem, we propose a phrase-level Temporal Relationship Mining (TRM) framework employing the temporal relationship relevant to the phrase and the whole sentence to have a better understanding of each semantic entity in the sentence. Specifically, we use phrase-level predictions to refine the sentence-level prediction, and use Multiple Instance Learning to improve the quality of phrase-level predictions. We also exploit the consistency and exclusiveness constraints of phrase-level and sentence-level predictions to regularize the training process, thus alleviating the ambiguity of each phrase prediction. The proposed approach sheds light on how machines can understand detailed phrases in a sentence and their compositions in their generality rather than learning the annotation biases. Experiments on the ActivityNet Captions and Charades-STA datasets show the effectiveness of our method on both phrase and sentence temporal localization and enable better model interpretability and generalization when dealing with unseen compositions of seen concepts. Code can be found at https://github.com/minghangz/TRM. Minghang Zheng, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
AAAI | 3 |
| 2023 | Efficient Adaptive Human-Object Interaction Detection with Concept-guided MemoryabstractHuman Object Interaction (HOI) detection aims to localize and infer the relationships between a human and an object. Arguably, training supervised models for this task from scratch presents challenges due to the performance drop over rare classes and the high computational cost and time required to handle long-tailed distributions of HOIs in complex HOI scenes in realistic settings. This observation motivates us to design an HOI detector that can be trained even with long-tailed labeled data and can leverage existing knowledge from pre-trained models. Inspired by the powerful generalization ability of the large Vision-Language Models (VLM) on classification and retrieval tasks, we propose an efficient Adaptive HOI Detector with Concept-guided Memory (ADA-CM). ADA-CM has two operating modes. The first mode makes it tunable without learning new parameters in a training-free paradigm. Its second mode incorporates an instance-aware adapter mechanism that can further efficiently boost performance if updating a lightweight set of parameters can be afforded. Our proposed method achieves competitive results with state-of-the-art on the HICO-DET and V-COCO datasets with much less training time. Code can be found at https://github.com/ltttpku/ADA-CM. Ting Lei 0001, Fabian Caba Heilbron, Qingchao Chen, Hailin Jin, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 3 |
| 2023 | Confidence-aware Pseudo-label Learning for Weakly Supervised Visual GroundingabstractVisual grounding aims at localizing the target object in image which is most related to the given free-form natural language query. As labeling the position of target object is labor-intensive, the weakly supervised methods, where only image-sentence annotations are required during model training have recently received increasing attention. Most of the existing weakly-supervised methods first generate region proposals via pre-trained object detectors and then employ either cross-modal similarity score or reconstruction loss as the criteria to select proposal from them. However, due to the cross-modal heterogeneous gap, these method often suffer from high confidence spurious association and model prone to error propagation. In this paper, we propose Confidence-aware Pseudo-label Learning (CPL) to overcome the above limitations. Specifically, we first adopt both the uni-modal and cross-modal pre-trained models and propose conditional prompt engineering to automatically generate multiple ‘descriptive, realistic and diverse’ pseudo language queries for each region proposal, and then establish reliable cross-modal association for model training based on the uni-modal similarity score (between pseudo and real text queries). Secondly, we propose a confidence-aware pseudo label verification module which reduces the amount of noise encountered in the training process and the risk of error propagation. Experiments on five widely used datasets validate the efficacy of our proposed components and demonstrate state-of-the-art performance. Code can be found at https://github.com/zjh31/CPL.git Yang Liu 0105, Qingchao Chen, Yuxin Peng 0001 |
ICCV | 3 |
| 2023 | Masked Retraining Teacher-Student Framework for Domain Adaptive Object DetectionabstractDomain adaptive Object Detection (DAOD) leverages a labeled domain (source) to learn an object detector generalizing to a novel domain without annotation (target). Recent advances use a teacher-student framework, i.e., a student model is supervised by the pseudo labels from a teacher model. Though great success, they suffer from the limited number of pseudo boxes with incorrect predictions caused by the domain shift, misleading the student model to get sub-optimal results. To mitigate this problem, we propose Masked Retraining Teacher-student framework (MRT) which leverages masked autoencoder and selective retraining mechanism on detection transformer. Specifically, we present a customized design of masked autoencoder branch, masking the multi-scale feature maps of target images and reconstructing features by the encoder of the student model and an auxiliary decoder. This helps the student model capture target domain characteristics and become a more data-efficient learner to gain knowledge from the limited number of pseudo boxes. Furthermore, we adopt selective retraining mechanism, periodically re-initializing certain parts of the student parameters with masked autoencoder refined weights to allow the model to jump out of the local optimum biased to the incorrect pseudo labels. Experimental results on three DAOD benchmarks demonstrate the effectiveness of our method. Code can be found at https://github.com/JeremyZhao1998/MRT-release. Zijing Zhao 0004, Sitong Wei, Qingchao Chen, Dehui Li, Yuxin Peng 0001, Yang Liu 0105 |
ICCV | 3 |
| 2023 | Using Multimodal Contrastive Knowledge Distillation for Video-Text RetrievalabstractCross-modal retrieval aims to enable a flexible bi-directional retrieval experience across different modalities (e.g., searching for videos with texts). Many existing efforts tend to learn a common semantic representation embedding space in which items of different modalities can be directly compared, wherein the positive global representations of video-text pairs are pulled close while the negative ones are pushed apart via pair-wise ranking loss. However, such a vanilla loss would unfortunately yield ambiguous feature embeddings for texts of different videos, causing inaccurate cross-modal matching and unreliable retrievals. Toward this end, we propose a multimodal contrastive knowledge distillation method for instance video-text retrieval, called MCKD, by adaptively using the general knowledge of self-supervised model (teacher) to calibrate mixed boundaries. Specifically, the teacher model is tailored for robust (less-ambiguous) visual-text joint semantic space by maximizing mutual information of co-occurred modalities during multimodal contrastive learning. This robust and structural inter-instance knowledge is then distilled, with the help of explicit discrimination loss, to a student model for improved matching performance. Extensive experiments on four public benchmark video-text datasets (MSR-VTT, TGIF, VATEX, and Youtube2Text) demonstrate that our MCKD can achieve at most 8.8%, 6.4%, 5.9%, and 5.3% improvement in text-to-video performance by the$\text{R}\text{@}1$metric, compared with 14 SoTA baselines. Wentao Ma 0003, Qingchao Chen, Tongqing Zhou, Shan Zhao 0002, Zhiping Cai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Uncertainty-Induced Transferability Representation for Source-Free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation (SFUDA) aims to learn a target domain model using unlabeled target data and the knowledge of a well-trained source domain model. Most previous SFUDA works focus on inferring semantics of target data based on the source knowledge. Without measuring the transferability of the source knowledge, these methods insufficiently exploit the source knowledge, and fail to identify the reliability of the inferred target semantics. However, existing transferability measurements require either source data or target labels, which are infeasible in SFUDA. To this end, firstly, we propose a novel Uncertainty-induced Transferability Representation (UTR), which leverages uncertainty as the tool to analyse the channel-wise transferability of the source encoder in the absence of the source data and target labels. The domain-level UTR unravels how transferable the encoder channels are to the target domain and the instance-level UTR characterizes the reliability of the inferred target semantics. Secondly, based on the UTR, we propose a novel Calibrated Adaption Framework (CAF) for SFUDA, including i) the source knowledge calibration module that guides the target model to learn the transferable source knowledge and discard the non-transferable one, and ii) the target semantics calibration module that calibrates the unreliable semantics. With the help of the calibrated source knowledge and the target semantics, the model adapts to the target domain safely and ultimately better. We verified the effectiveness of our method using experimental results and demonstrated that the proposed method achieves state-of-the-art performances on the three SFUDA benchmarks. Code is available at https://github.com/SPIresearch/UTR. Jiangbo Pei, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
IEEE Trans. Image Process. | 6 |
| 2022 | Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningabstractVideo moment localization aims at localizing the video segments which are most related to the given free-form natural language query. The weakly supervised setting, where only video level description is available during training, is getting more and more attention due to its lower annotation cost. Prior weakly supervised methods mainly use sliding windows to generate temporal proposals, which are independent of video content and low quality, and train the model to distinguish matched video-query pairs and unmatched ones collected from different videos, while neglecting what the model needs is to distinguish the unaligned segments within the video. In this work, we propose a novel weakly supervised solution by introducing Contrastive Negative sample Mining (CNM). Specifically, we use a learnable Gaussian mask to generate positive samples, highlighting the video frames most related to the query, and consider other frames of the video and the whole video as easy and hard negative samples respectively. We then train our network with the Intra-Video Contrastive loss to make our positive and negative samples more discriminative. Our method has two advantages: (1) Our proposal generation process with a learnable Gaussian mask is more efficient and makes our positive sample higher quality. (2) The more difficult intra-video negative samples enable our model to distinguish highly confusing scenes. Experiments on two datasets show the effectiveness of our method. Code can be found at https://github.com/minghangz/cnm. Minghang Zheng, Yanjie Huang, Qingchao Chen, Yang Liu 0105 |
AAAI | 3 |
| 2022 | Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningabstractTemporal sentence grounding aims to detect the most salient moment corresponding to the natural language query from untrimmed videos. As labeling the temporal boundaries is labor-intensive and subjective, the weakly- supervised methods have recently received increasing attention. Most of the existing weakly-supervised methods gen-erate the proposals by sliding windows, which are content- independent and of low quality. Moreover, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video. In this paper, we propose Contrastive Proposal Learning(CPL) to overcome the above limitations. Specifi-cally, we use multiple learnable Gaussian functions to gen-erate both positive and negative proposals within the same video that can characterize the multiple events in a long video. Then, we propose a controllable easy to hard neg-ative proposal mining strategy to collect negative samples within the same video, which can ease the model opti-mization and enables CPL to distinguish highly confusing scenes. The experiments show that our method achieves state-of-the-art performance on Charades-STA and Activi-tyNet Captions datasets. The code and models are available at https://github.com/minghangz/cpl. Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 0001, Yang Liu 0105 |
CVPR | 3 |
| 2022 | Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal FusionabstractIn multimodal video event localization, we usually leverage feature fusion across different axes, such as the modality and temporal axes, for better context. To reduce the costs of detailed annotations, recent solutions explore weakly supervised settings. However, we observe that when feature fusion meets weakly supervised localization, problems can occur. It may cause "feature cross-interference", which produces a smearing effect on the localization result and can’t be effectively supervised with conventional multiple instance learning loss. We verify it quantitatively on the audio-visual video parsing (AVVP) task, and propose a cross-instance data-augmentation framework, which can preserve the benefits of feature fusion while providing explicit feedbacks for feature cross-interference. We show that our method can enhance performance of existing models on two weakly supervised audio-visual localization tasks, i.e. AVVP and AVE. Yonggang Zhu, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Qingchao Chen |
ICASSP | 6 |
| 2022 | Delving into the Continuous Domain AdaptationabstractExisting domain adaptation methods assume that domain discrepancies are caused by a few discrete attributes and variations, e.g., art, real, painting, quickdraw, etc. We argue that this is not realistic as it is implausible to define the real-world datasets using a few discrete attributes. Therefore, we propose to investigate a new problem namely the Continuous Domain Adaptation (CDA) through the lens where infinite domains are formed by continuously varying attributes. Leveraging knowledge of two labeled source domains and several observed unlabeled target domains data, the objective of CDA is to learn a generalized model for whole data distribution with the continuous attribute. Besides the contributions of formulating a new problem, we also propose a novel approach as a strong CDA baseline. To be specific, firstly we propose a novel alternating training strategy to reduce discrepancies among multiple domains meanwhile generalize to unseen target domains. Secondly, we propose a continuity constraint when estimating the cross-domain divergence measurement. Finally, to decouple the discrepancy from the mini-batch size, we design a domain-specific queue to maintain the global view of the source domain that further boosts the adaptation performances. Our method is proven to achieve the state-of-the-art in CDA problem using extensive experiments. The code is available at https://github.com/SPIresearch/CDA. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
ACM Multimedia | 5 |
| 2022 | Toward a perceptive pretraining framework for Audio-Visual Video Parsing
Jianning Wu, Zhuqing Jiang, Qingchao Chen, Shiping Wen 0001, Aidong Men, Haiying Wang 0005 |
Inf. Sci. | 3 |
| 2021 | Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalabstractWhen can we expect a text-video retrieval system to work effectively on datasets that differ from its training domain? In this work, we investigate this question through the lens of unsupervised domain adaptation in which the objective is to match natural language queries and video content in the presence of domain shift at query-time. Such systems have significant practical applications since they are capable generalising to new data sources without requiring corresponding text annotations. We make the following contributions: (1) We propose the UDAVR (Unsupervised Domain Adaptation for Video Retrieval) benchmark and employ it to study the performance of text-video retrieval in the presence of domain shift. (2) We propose Concept-Aware-Pseudo-Query (CAPQ), a method for learning discriminative and transferable features that bridge these cross-domain discrepancies to enable effective target domain retrieval using source domain supervision. (3) We show that CAPQ outperforms alternative domain adaptation strategies on UDAVR. Qingchao Chen, Yang Liu 0105, Samuel Albanie |
AAAI | 1 |
| 2021 | Adaptive Cross-Modal Prototypes for Cross-Domain Visual-Language RetrievalabstractIn this paper, we study the task of visual-text retrieval in the highly practical setting in which labelled visual data with paired text descriptions are available in one domain (the "source"), but only unlabelled visual data (without text descriptions) are available in the domain of interest (the "target"). We propose the Adaptive Cross-Modal Prototypes framework which seeks to enable target domain retrieval by learning cross-modal visual-text representations while minimising both uni-modal and cross-modal distribution shift across the source and target domains. Our approach is built upon two key ideas: first, we encode the inductive bias that the learned cross-modal representations should be compositional with respect to concepts in each modality—this is achieved through clustering pretrained uni-modal features across each domain and designing a careful regularisation scheme to preserve the resulting structure. Second, we employ mutual information maximisation between cross-modal representations in the source and target domains during learning—this provides a mechanism that preserves commonalities between the domains while discarding signal in each that cannot be inferred from the other. We showcase our approach for the task of cross-domain visual-text retrieval, outperforming existing approaches for both images and videos. Yang Liu 0105, Qingchao Chen, Samuel Albanie |
CVPR | 2 |
| 2020 | Structure-Aware Feature Fusion for Unsupervised Domain AdaptationabstractUnsupervised domain Adaptation (UDA) aims to learn and transfer generalized features from a labelled source domain to a target domain without any annotations. Existing methods only aligning high-level representation but without exploiting the complex multi-class structure and local spatial structure. This is problematic as 1) the model is prone to negative transfer when the features from different classes are misaligned; 2) missing the local spatial structure poses a major obstacle in performing the fine-grained feature alignment. In this paper, we integrate the valuable information conveyed in classifier prediction and local feature maps into global feature representation and then perform a single mini-max game to make it domain invariant. In this way, the domain-invariant feature not only describes the holistic representation of the original image but also preserves mode-structure and fine-grained spatial structural information. The feature integration is achieved by estimating and maximizing the mutual information (MI) among the global feature, local feature and classifier prediction simultaneously. As the MI is hard to measure directly in high-dimension spaces, we adopt a new objective function that implicitly maximizes the MI via an effective sampling strategy and a discriminator design. Our STructure-Aware Feature Fusion (STAFF) network achieves the state-of-the-art performances in various UDA datasets. Qingchao Chen, Yang Liu 0105 |
AAAI | 1 |
| 2020 | Amplifying Key Cues for Human-Object-Interaction Detection
Yang Liu 0105, Qingchao Chen, Andrew Zisserman |
ECCV (14) | 2 |
| 2020 | Longitudinal Image Registration with Temporal-Order and Subject-Specificity Discrimination
Qianye Yang, Yunguan Fu, Francesco Giganti, Nooshin Ghavami, Qingchao Chen, J. Alison Noble, Tom Vercauteren, Dean C. Barratt, Yipeng Hu |
MICCAI (3) | 5 |
| 2018 | Dictionary Learning Inspired Deep Network for Scene RecognitionabstractScene recognition remains one of the most challenging problems in image understanding. With the help of fully connected layers (FCL) and rectified linear units (ReLu), deep networks can extract the moderately sparse and discriminative feature representation required for scene recognition. However, few methods consider exploiting a sparsity model for learning the feature representation in order to provide enhanced discriminative capability. In this paper, we replace the conventional FCL and ReLu with a new dictionary learning layer, that is composed of a finite number of recurrent units to simultaneously enhance the sparse representation and discriminative abilities of features via the determination of optimal dictionaries. In addition, with the help of the structure of the dictionary, we propose a new label discriminative regressor to boost the discrimination ability. We also propose new constraints to prevent overfitting by incorporating the advantage of the Mahalanobis and Euclidean distances to balance the recognition accuracy and generalization performance. Our proposed approach is evaluated using various scene datasets and shows superior performance to many state-of-the-art approaches. Yang Liu 0105, Qingchao Chen, Wei Chen 0016, Ian J. Wassell |
AAAI | 2 |
| 2018 | Re-Weighted Adversarial Adaptation Network for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to transfer domain knowledge from existing well-defined tasks to new ones where labels are unavailable. In the real-world applications, as the domain (task) discrepancies are usually uncontrollable, it is significantly motivated to match the feature distributions even if the domain discrepancies are disparate. Additionally, as no label is available in the target domain, how to successfully adapt the classifier from the source to the target domain still remains an open question. In this paper, we propose the Re-weighted Adversarial Adaptation Network (RAAN) to reduce the feature distribution divergence and adapt the classifier when domain discrepancies are disparate. Specifically, to alleviate the need of common supports in matching the feature distribution, we choose to minimize optimal transport (OT) based Earth-Mover (EM) distance and reformulate it to a minimax objective function. Utilizing this, RAAN can be trained in an end-to-end and adversarial manner. To further adapt the classifier, we propose to match the label distribution and embed it into the adversarial training. Finally, after extensive evaluation of our method using UDA datasets of varying difficulty, RAAN achieved the state-of-the-art results and outperformed other methods by a large margin when the domain shifts are disparate. Qingchao Chen, Yang Liu 0105, Ian J. Wassell, Kevin Chetty |
CVPR | 1 |
| 2017 | Deep network for image super-resolution with a dictionary learning layerabstractThe aim of single image super-resolution (SR) is to generate a high-resolution (HR) image from a low-resolution (LR) observable image. In this paper, we address this task by integrating sparse coding and dictionary learning schemes into an end-to-end deep architecture. More specifically, we propose a new non-linear dictionary learning layer composed of a finite number of recurrent units to solve the sparse codes and also to yield the relevant gradients to update the dictionary. In addition, we present a new deep network architecture using the proposed non-linear layers, where two separate parallel dictionaries are adopted to represent the LR and HR images respectively. The whole network is optimized by back propagation, constraining not only reconstruction errors between the restored and the ground truth HR images but also between the sparse codes of the LR and HR image pairs. Various datasets are used to evaluate the performance of the proposed approach and it is shown to outperform many state-of-the-art single image super-resolution algorithms. Yang Liu 0105, Qingchao Chen, Ian J. Wassell |
ICIP | 2 |
| 2016 | Support Discrimination Dictionary Learning for Image Classification
Yang Liu 0105, Wei Chen 0016, Qingchao Chen, Ian J. Wassell |
ECCV (2) | 3 |
| 2016 | Activity recognition based on micro-Doppler signature with in-home Wi-FiabstractDevice free activity recognition and monitoring has become a promising research area with increasing public interest in pattern of life monitoring and chronic health conditions. This paper proposes a novel framework for in-home Wi-Fi signal-based activity recognition in e-healthcare applications using passive micro-Doppler (m-D) signature classification. The framework includes signal modeling, Doppler extraction and m-D classification. A data collection campaign was designed to verify the framework where six m-D signatures corresponding to typical daily activities are sucessfully detected and classified using our software defined radio (SDR) demo system. Analysis of the data focussed on potential discriminative characteristics, such as maximum Doppler frequency and time duration of activity. Finally, a sparsity induced classifier is applied for adaptting the method in healthcare application scenarios and the results are compared with those from the well-known Support Vector Machine (SVM) method. Qingchao Chen, Bo Tan 0003, Kevin Chetty, Karl Woodbridge |
HealthCom | 1 |
| 2015 | Indoor target tracking using high doppler resolution passive Wi-Fi radarabstractThis paper describes two Doppler only indoor passive Wi-Fi tracking methods based on high Doppler resolution passive radar. Two filters are investigated in this paper, the extended Kalman filter and the sequential importance resampling (SIR) particle filter. Experimental results for these two tracking filters are presented using results from software defined passive Wi-Fi radar using a standard 802.11 access point as an illuminator. The experimental results show that the SIR particle filter performs well using Wi-Fi signals for indoor tracking with a high degree of accuracy. Proposals for simplifying the SIR particle and application to multiple target tracking are also discussed. Qingchao Chen, Bo Tan 0003, Karl Woodbridge, Kevin Chetty |
ICASSP | 1 |
| 2012 | RSS-Based Node Localization in the Existence of Moving ObstructionsabstractIn the context of wireless sensor networks, a node's location must be known for its data to be meaningful in many cases. Received signal strength (RSS)-based localization has been widely used because of low complexity and easy deployment. This paper proposes a novel method to localize nodes in the presence of randomly moving obstructions. We introduce background learning to reduce interferences caused by moving obstructions such as people or other objects. Based on our experimental results, each link of data is modeled as a mixture of Gaussians (MoG) and its parameters are updated by background learning. In this way, we can reduce the interferences of moving obstructions from obtained RSS measurements. Then we use least-square (LS) cooperative localization algorithm to implement node localization and the experimental results show good performance. Bo Yang 0007, Aidong Men, Qingchao Chen |
VTC Fall | 5 |