Donglian Qi

dblp:48/8863 · DBLP profile ↗
← Back
39ranked-venue papers
1as first author
29since 2021 · last 2025
0000-0002-6535-2221ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ROSE: Remove Objects with Side Effects in Videos
abstract
Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as supervision. This paper presents \method, termed \textbf{R}emove \textbf{O}bjects with \textbf{S}ide \textbf{E}ffects, a framework that systematically studies the object's effects on environment, which can be categorized into five common cases: shadows, reflections, light, translucency and mirror. Given the challenges of curating paired videos exhibiting the aforementioned effects, we leverage a 3D rendering engine for synthetic data generation. We carefully construct a fully-automatic pipeline for data preparation, which simulates a large-scale paired dataset with diverse scenes, objects, shooting angles, and camera trajectories. ROSE is implemented as an video inpainting model built on diffusion transformer. To localize all object-correlated areas, the entire video is fed into the model for reference-based erasing. Moreover, additional supervision is introduced to explicitly predict the areas affected by side effects, which can be revealed through the differential mask between the paired videos. To fully investigate the model performance on various side effect removal, we presents a new benchmark, dubbed ROSE-Bench, incorporating both common scenarios and the five special side effects for comprehensive evaluation. Experimental results demonstrate that \method achieves superior performance compared to existing video object erasing models and generalizes well to real-world video scenarios.
Chenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao, Hantang Liu, Yunfeng Yan, Donglian Qi, Xi Chen 0119, Hengshuang Zhao
NeurIPS7
2025 MDRN: Multi-domain representation network for unsupervised domain generalization
abstract
Abstract In deep neural networks, performance can degrade when test data distributions differ from training data. Unsupervised Domain Generalization (UDG) aims to improve generalization across unseen domains by leveraging multiple source domains without supervision. Traditional methods focus on extracting domain‐invariant features, potentially at the expense of feature space integrity and generalization potential. We presents a Multi‐Domain Representation Network (MDRN) for unsupervised multi‐domain learning. MDRN innovates by disentangling and preserving both domain‐invariant and domain‐specific features through an unsupervised cross‐domain reconstruction task. It employs content encoders for domain‐invariant features and multi‐domain style encoders for domain‐specific characteristics. By merging these features based on domain similarity, MDRN constructs a comprehensive feature space that enhances image reconstruction across domains. Additionally, MDRN integrates domain‐specific classifiers, which learn domain classification and provide weighted fusion of domain‐specific features. This design facilitates effective inter‐domain distance measurement and feature integration. Experiments on PACS and DomainNet show MDRN's superior performance over existing state‐of‐the‐art UDG approaches, highlighting its effectiveness in handling distribution shifts between source and target domains.
Yangyang Zhong, Yunfeng Yan, Pengxin Luo, Weizhen He, Yiheng Deng, Donglian Qi
IET Image Process.6
2025 Adept: Annotation-denoising auxiliary tasks with discrete cosine transform map and keypoint for human-centric pretraining
Weizhen He, Yunfeng Yan, Shixiang Tang, Yiheng Deng, Yangyang Zhong, Pengxin Luo, Donglian Qi
Neurocomputing7
2025 Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-Identification
abstract
Recently, person re-identification (ReID) has witnessed fast development due to its broad practical applications and proposed various settings, e.g., traditional ReID, clothes-changing ReID, and visible-infrared ReID. However, current studies primarily focus on single specific tasks, which limits model applicability in real-world scenarios. This paper aims to address this issue by introducing a novel instruct-ReID task that unifies 6 existing ReID tasks in one model and retrieves images based on provided visual or textual instructions. Instruct-ReID is the first exploration of a general ReID setting, where 6 existing ReID tasks can be viewed as special cases by assigning different instructions. To facilitate research in this new instruct-ReID task, we propose a large-scale OmniReID++ benchmark equipped with diverse data and comprehensive evaluation methods, e.g., task-specific and task-free evaluation settings. In the task-specific evaluation setting, gallery sets are categorized according to specific ReID tasks. We propose a novel baseline model, IRM, with an adaptive triplet loss to handle various retrieval tasks within a unified framework. For task-free evaluation setting, where target person images are retrieved from task-agnostic gallery sets, we further propose a new method called IRM++ with novel memory bank-assisted learning. Extensive evaluations of IRM and IRM++ on OmniReID++ benchmark demonstrate the superiority of our proposed methods, achieving state-of-the-art performance on 10 test sets.
Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu 0006, Yizhou Wang 0007, Lei Bai 0001, Qingsong Xie, Rui Zhao 0001, Donglian Qi, Wanli Ouyang, Shixiang Tang
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 Wavelet-Driven Spatiotemporal Predictive Learning: Bridging Frequency and Time Variations
abstract
Spatiotemporal predictive learning is a paradigm that empowers models to learn spatial and temporal patterns by predicting future frames from past frames in an unsupervised manner. This method typically uses recurrent units to capture long-term dependencies, but these units often come with high computational costs and limited performance in real-world scenes. This paper presents an innovative Wavelet-based SpatioTemporal (WaST) framework, which extracts and adaptively controls both low and high-frequency components at image and feature levels via 3D discrete wavelet transform for faster processing while maintaining high-quality predictions. We propose a Time-Frequency Aware Translator uniquely crafted to efficiently learn short- and long-range spatiotemporal information by individually modeling spatial frequency and temporal variations. Meanwhile, we design a wavelet-domain High-Frequency Focal Loss that effectively supervises high-frequency variations. Extensive experiments across various real-world scenarios, such as driving scene prediction, traffic flow prediction, human motion capture, and weather forecasting, demonstrate that our proposed WaST achieves state-of-the-art performance over various spatiotemporal prediction methods.
Xuesong Nie, Yunfeng Yan, Siyuan Li 0002, Cheng Tan 0012, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Stan Z. Li, Donglian Qi
AAAI9
2024 Expert-Guided Model Cultivation: CoTeaching to Resolve Abstruseness and Enhance Learning Performance
Feng Zhou 0011, Zhidong Li, Yang Wang 0002, Donglian Qi, Shuming Li
ADMA (2)5
2024 Instruct-ReID: A Multi-Purpose Person Re-Identification Task with Instructions
abstract
Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a new instruct-ReID task that requires the model to retrieve images according to the given image or language instructions. Our instruct-ReID is a more general ReID setting, where existing 6 ReID tasks can be viewed as special cases by designing different instructions. We propose a large-scale OmniReID benchmark and an adaptive triplet loss as a baseline method to facilitate research in this new setting. Experimental results show that the proposed multi-purpose ReID model, trained on our OmniReID benchmark without finetuning, can improve +0.5%, +0.6%, +7.7% mAP on Market1501, MSMT17, CUHK03 for traditional ReID, +6.4%, +7.1%, +11.2% mAP on PRCC, VC-Clothes, LTCC for clothes-changing ReID, +11.7% mAP on COCAS+ real2 for clothes template based clothes-changing ReID when using only RGB images, +24.9% mAP on COCAS+ real2 for our newly defined language-instructed ReID, +4.3% on LLCM for visible-infrared ReID, +2.6% on CUHK-PEDES for text-to-image ReID. The datasets, the model, and code are available at https://github.com/hwz-zju/Instruct-ReID.
Weizhen He, Yiheng Deng, Shixiang Tang, Qihao Chen, Qingsong Xie, Yizhou Wang 0007, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Wanli Ouyang, Donglian Qi, Yunfeng Yan
CVPR11
2024 PredToken: Predicting Unknown Tokens and Beyond with Coarse-to-Fine Iterative Decoding
abstract
Predictive learning models, which aim to predict future frames based on past observations, are crucial to constructing world models. These models need to maintain low-level consistency and capture high-level dynamics in unannotated spatiotemporal data. Transitioning from frame-wise to token-wise prediction presents a viable strategy for addressing these needs. How to improve token representation and optimize token decoding presents significant challenges. This paper introduces PredToken, a novel predictive framework that addresses these issues by decoupling space-time tokens into distinct components for iterative cascaded decoding. Concretely, we first design a “decomposition, quantization, and reconstruction” schema based on VQGAN to improve the token representation. This scheme disentangles low- and high-frequency representations and employs a dimension-aware quantization model, allowing more low-level details to be preserved. Building on this, we present a “coarse-to-fine iterative decoding” method. It leverages dynamic soft decoding to refine coarse tokens and static soft decoding for fine tokens, enabling more high-level dynamics to be captured. These designs make Pred-Token produce high-quality predictions. Extensive experiments demonstrate the superiority of our method on various real-world spatiotemporal predictive benchmarks. Furthermore, PredToken can also be extended to other visual generative tasks to yield realistic outcomes.
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
CVPR6
2024 SAMP: Adapting Segment Anything Model for Pose Estimation
abstract
Segment Anything Model (SAM) exhibits superior performance for segmentation. Many follow-up works explore adapting this powerful model to specific domains. However, those works mainly focus on different sub-tasks of segmentation. The cross-task generalization ability of SAM is still not explored. In this paper, we propose SAMP (SAM for Pose), which makes the first attempt to adapt SAM for pose estimation. We observe that SAM could segment different human parts with specific prompts, proving that it contains the knowledge to understand the human structure. Considering that localizing keypoints requires fine-grained perceptual capabilities, we design a Detail-aware Adapter (DA-Adapter), which complements the features of the SAM encoder with multi-scale feature fusion and multi-level supervision. Experimental results demonstrate that SAMP achieves novel state-of-the-art against previously specifically designed pose estimation methods. Specifically, with ViT-B backbone, SAMP achieves 78.1% AP on the COCO val2017, 77.1% AP on the COCO test-dev2017, and 70.5% AP on the CrowdPose dataset.
Zhihang Zhu, Yunfeng Yan, Haoyuan Jin, Xuesong Nie, Donglian Qi, Xi Chen 0119
ICME6
2024 Object-Level Pseudo-3D Lifting for Distance-Aware Tracking
abstract
Multi-object tracking (MOT) is a pivotal task for media interpretation, where reliable motion and appearance cues are essential for cross-frame identity preservation. However, limited by the inherent perspective properties of 2D space, the crowd density and frequent occlusions in real-world scenes expose the fragility of these cues. We observe the natural advantage of objects being well-separated in high-dimensional space and propose a novel 2D MOT framework, "Detecting-Lifting-Tracking'' (DLT). Initially, a pre-trained detector is employed to capture 2D object information. Secondly, we introduce a Mamba Distance Estimator to obtain the distances of objects to a monocular camera with temporal consistency, achieving object-level pseudo-3D lifting. Finally, we thoroughly explore distance-aware tracking via pseudo-3D information. Specifically, we introduce a Score-Distance Hierarchical Matching and Short-Long Terms Association to enhance accurate and robust association capability. Even without appearance cues, our DLT achieves state-of-the-art performance on MOT17, MOT20, and DanceTrack, demonstrating its potential to address occlusion challenges.
Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
ACM Multimedia6
2024 Triplet Attention Transformer for Spatiotemporal Predictive Learning
abstract
Spatiotemporal predictive learning offers a self-supervised learning paradigm that enables models to learn both spatial and temporal patterns by predicting future sequences based on historical sequences. Mainstream methods are dominated by recurrent units, yet they are limited by their lack of parallelization and often underperform in real-world scenarios. To improve prediction quality while maintaining computational efficiency, we propose an innovative triplet attention transformer designed to capture both inter-frame dynamics and intra-frame static features. Specifically, the model incorporates the Triplet Attention Module (TAM), which replaces traditional recurrent units by exploring self-attention mechanisms in temporal, spatial, and channel dimensions. In this configuration: (i) temporal tokens contain abstract representations of inter-frame, facilitating the capture of inherent temporal dependencies; (ii) spatial and channel attention combine to refine the intra-frame representation by performing fine-grained interactions across spatial and channel dimensions. Alternating temporal, spatial, and channel-level attention allows our approach to learn more complex short-and long-range spatiotemporal dependencies. Extensive experiments demonstrate performance surpassing existing recurrent-based and recurrent-free methods, achieving state-of-the-art under multi-scenario examination including moving object trajectory prediction, traffic flow prediction, driving scene prediction, and human motion capture.
Xuesong Nie, Xi Chen 0119, Haoyuan Jin, Zhihang Zhu, Yunfeng Yan, Donglian Qi
WACV6
2024 ADIR: Advanced domain-invariant representation via decoupling learning and information bottleneck
abstract
Abstract The discrepancy in data distribution between training and testing scenarios, as well as the inductive bias of convolutional neural networks towards image styles, reduces the model's generalization ability. Many unsupervised domain generalization methods based on feature decoupling suffer from an initial neglect of explicit decoupling of content and style features, resulting in content features that still contain considerable redundant information, thereby restricting improvements in generalization capability. To tackle this problem, this paper optimizes the learning process of domain‐invariant (content) features into an information compression issue, minimizing redundancy in content features. Furthermore, to enhance decoupled learning, this paper introduces innovative cross‐domain loss functions and image reconstruction modules that explicitly decouple and merge content and style across different domains. Extensive experiments demonstrate the method's significant enhancements over recent cutting‐edge approaches.
Yangyang Zhong, Yunfeng Yan, Pengxin Luo, Donglian Qi
IET Image Process.5
2024 ScopeViT: Scale-Aware Vision Transformer
Xuesong Nie, Haoyuan Jin, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
Pattern Recognit.6
2024 AHOR: Online Multi-Object Tracking With Authenticity Hierarchizing and Occlusion Recovery
abstract
Despite extensive exploration of more powerful multi-object tracking (MOT) frameworks, the impact of frequent occlusion has remained a formidable challenge. In this work, we present a novel MOT framework with Authenticity Hierarchizing and Occlusion Recovery (AHOR), that strikingly handles occlusion and demonstrates superior precision and adaptability. Specifically, through an in-depth analysis of the classical tracking-by-detection (TBD) paradigm, we fully upgrade three aspects. Firstly, we propose an Existence Score that provides a more accurate depiction of detection authenticity under occlusion, enhancing the effectiveness and robustness of the hierarchical association. Secondly, we present an ingeniously devised pre-processing method in conjunction with a Recovery Intersection over Union (RIoU) for location similarity measurement, addressing the adverse effects of occlusion-induced disparity between visible and true object regions. Lastly, we introduce an Occluded Person Re-identification Module (ODReID) that extracts appearance features from the restricted visible region, overcoming the critical dependence on object quality. Results of extensive experiments demonstrate that our AHOR achieves state-of-the-art performance on MOT17, MOT20, DanceTrack, and VisDrone test sets.
Haoyuan Jin, Xuesong Nie, Yunfeng Yan, Xi Chen 0119, Zhihang Zhu, Donglian Qi
IEEE Trans. Circuits Syst. Video Technol.6
2024 Ultrasound Communication Using the Nonlinearity Effect of Microphone Circuits in Smart Devices
abstract
Acoustic communication has become a research focus without requiring extra hardware and facilitates numerous near-field applications such as mobile payment. To communicate, existing researchers use either an audible frequency band or an inaudible one. The former gains a high throughput but endures being audible, which can be annoying to users. The latter, although inaudible, falls short in throughput due to the available (near) ultrasonic bandwidth. In this article, we achieve both high speed and inaudibility for acoustic communication by utilizing the nonlinearity effect on microphones. We theoretically prove the maximum throughput of inaudible acoustic communication by modulating an audible signal onto an ultrasonic band. Then, we design and implementUltraComm, which utilizes a specially designed OFDM scheme. The scheme takes into account the characteristics of the nonlinear speaker-to-microphone channel, aiming to mitigate the effects of signal distortion. We evaluateUltraCommon different mobile devices and achieve throughput as high as 16.24 kbps.
Xiaoyu Ji 0001, Donglian Qi, Wenyuan Xu 0001
ACM Trans. Sens. Networks4
2024 MLG-NCS: Multimodal Local-Global Neuromorphic Computing System for Affective Video Content Analysis
abstract
Despite neuromorphic computing (NC) technologies offer tremendous potential in executing computationally intensive tasks with high efficiency and low latency, most of existing methods are still difficult to achieve software-comparable accuracy. To address this challenge, we develop a multimodal local–global NC system (MLG-NCS) that can capture local characteristics and exchange global cross-modal information sufficiently. Specifically, a high-density memristor crossbar array is prepared to perform efficient parallel in-memory operations, serving as the fundamental component of the proposed MLG-NCS. To facilitate understanding of the proposed MLG-NCS design, the local feature representation module, the global cross-modal interaction module, and the output module are designed. The experimental results show that the proposed system has advantages in classification accuracy (ranked top three), time consumption (approximately ten times speed up), and latency (about 1.2–15.3 times faster), enabling good inter-related tradeoffs between latency, efficiency, and accuracy. This study is expected to promote the revolution and development of next-generation computing system, which takes a firm step toward artificial general intelligence (AGI).
Xiaoyue Ji, Zhekang Dong, Guangdong Zhou, Chun Sing Lai, Donglian Qi
IEEE Trans. Syst. Man Cybern. Syst.5
2023 UniHCP: A Unified Model for Human-Centric Perceptions
abstract
Human-centric perceptions (e.g., pose estimation, human parsing, pedestrian detection, person re-identification, etc.) play a key role in industrial applications of visual models. While specific human-centric tasks have their own relevant semantic aspect to focus on, they also share the same underlying semantic structure of the human body. However, few works have attempted to exploit such homogeneity and design a general-propose model for human-centric tasks. In this work, we revisit a broad range of human-centric tasks and unify them in a minimalist manner. We propose UniHCP, a Unified Model for Human-Centric Perceptions, which unifies a wide range of human-centric tasks in a simplified end-to-end manner with the plain vision transformer architecture. With large-scale joint training on 33 human-centric datasets, UniHCP can outperform strong baselines on several in-domain and downstream tasks by direct evaluation. When adapted to a specific task, UniHCP achieves new SOTAs on a wide range of human-centric tasks, e.g., 69.8 mIoU on CIHP for human parsing, 86.18 mA on PA100K for attribute prediction, 90.3 mAP on Market1501 for ReID, and 85.8 JI on CrowdHuman for pedestrian detection, performing better than specialized models tailored for each task. The code and pretrained model are available at https://github.com/OpenGVLab/UniHCP.
Yuanzheng Ci, Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Lei Bai 0001, Feng Zhu 0006, Rui Zhao 0001, Fengwei Yu, Donglian Qi, Wanli Ouyang
CVPR9
2023 Unsupervised Prompt Tuning for Text-Driven Object Detection
abstract
Grounded language-image pre-trained models have shown strong zero-shot generalization to various downstream object detection tasks. Despite their promising performance, the models rely heavily on the laborious prompt engineering. Existing works typically address this problem by tuning text prompts using downstream training data in a few-shot or fully supervised manner. However, a rarely studied problem is to optimize text prompts without using any annotations. In this paper, we delve into this problem and propose an Unsupervised Prompt Tuning framework for text-driven object detection, which is composed of two novel mean teaching mechanisms. In conventional mean teaching, the quality of pseudo boxes is expected to optimize better as the training goes on, but there is still a risk of overfitting noisy pseudo boxes. To mitigate this problem, 1) we propose Nested Mean Teaching, which adopts nested-annotation to supervise teacher-student mutual learning in a bi-level optimization manner; 2) we propose Dual Complementary Teaching, which employs an offline pre-trained teacher and an online mean teacher via data-augmentation-based complementary labeling so as to ensure learning without accumulating confirmation bias. By integrating these two mechanisms, the proposed unsupervised prompt tuning framework achieves significant performance improvement on extensive object detection datasets.
Weizhen He, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Donglian Qi, Yueting Zhuang
ICCV7
2023 A Brain-Inspired Hierarchical Interactive In-Memory Computing System and Its Application in Video Sentiment Analysis
abstract
Video sentiment analysis can effectively establish the relationship between the emotion state and the multimodal information, while still suffer from intensive computation and low efficiency, due to the von Neumann computing architecture. Here, we present a brain-inspired hierarchical interactive in-memory computing (IMC) system, which can efficiently solve ‘von Neumann bottleneck’, enabling cross-modal interactions and semantic gap elimination. First, a 1T1M synapse array is fabricated using cost-effective, highly stable, flexible, and eco-friendly carbon materials, offering efficient analog multiply-accumulate operations. To illustrate the complexity of the proposed brain-inspired hierarchical interactive IMC system, three modules are proposed: 1) unimodal extraction module, 2) hierarchical interactive module, 3) output module. Furthermore, the proposed system is validated by applying it to video sentiment analysis. The experimental results demonstrate that the proposed system outperforms the existing state-of-the-art methods with high computational efficiency and good robustness. This work opens up a new way to achieve the deep integration of nanomaterials, deep learning, and modern electronics into IMC.
Xiaoyue Ji, Zhekang Dong, Yifeng Han, Chun Sing Lai, Donglian Qi
IEEE Trans. Circuits Syst. Video Technol.5
2023 Distributed Self-Triggered Control for Frequency Restoration and Active Power Sharing in Islanded Microgrids
abstract
Distributed event-triggered secondary control in microgrids has been widely investigated to improve system efficiency. But most of them are based on consecutive triggering condition monitors, which would in turn increase the computation burden of the system. To this end, this article presents distributed self-triggered algorithmic solutions to the frequency restoration control and active power sharing control of islanded microgrids. Different from event-triggered control schemes, in our self-triggered solutions, each distributed generator is equipped with a local algorithm that enables it to pre-compute the next triggering time instant according to the states at the previous one. Our starting point is to design a triggering condition with a novel estimate error. Then, the next triggering time instant is determined by solving a quadratic equation established based on the triggering condition, rather than monitoring the triggering condition consecutively. Theoretical analysis and simulation results show that the proposed distributed self-triggered secondary controllers can highly reduce the communication and computation costs simultaneously.
Keng-Weng Lao, Donglian Qi, Hongxun Hui, Yunfeng Yan
IEEE Trans. Ind. Informatics3
2023 NIM-Nets: Noise-Aware Incomplete Multi-View Learning Networks
abstract
Data in real world are usually characterized in multiple views, including different types of features or different modalities. Multi-view learning has been popular in the past decades and achieved significant improvements. In this paper, we investigate three challenging problems in the field of incomplete multi-view representation learning, namely, i) how to reduce the influences produced by missing views in multi-view dataset, ii) how to learn a consistent and informative representation among different views and iii) how to alleviate the impacts of the inherent noise in multi-view data caused by high-dimensional features or varied quality for different data points. To address these challenges, we integrate these three tasks into a problem and propose a novel framework termed Noise-aware Incomplete Multi-view Learning Networks (NIM-Nets). NIM-Nets fully utilize incomplete data from different views to produce a multi-view shared representation which is consistent, informative and robust to noise. We model the inherent noise in data by defining the distribution $\Gamma $ and assuming that each observation in the incomplete dataset is sampled from the distribution $\Gamma $ . To the best of our knowledge, this is the first work to unify learning the consistent and informative representation, alleviating the impacts of noise in data and handling the view-missing patterns in multi-view learning into a framework. We also first give a definition of robustness and completeness for incomplete multi-view representation learning. Based on NIM-Nets, we present joint optimization models for classification and clustering, respectively. Extensive experiments on different datasets demonstrate the effectiveness of our method over the existing work based on classification and clustering tasks in terms of different metrics.
Yalan Qin, Chuan Qin 0001, Xinpeng Zhang 0001, Donglian Qi, Guorui Feng
IEEE Trans. Image Process.4
2022 FocalClick: Towards Practical Interactive Image Segmentation
abstract
Interactive segmentation allows users to extract target masks by making positive/negative clicks. Although explored by many previous works, there is still a gap between academic approaches and industrial needs: first, existing models are not efficient enough to work on low-power devices; second, they perform poorly when used to refine preexisting masks as they could not avoid destroying the correct part. FocalClick solves both issues at once by predicting and updating the mask in localized areas. For higher efficiency, we decompose the slow prediction on the entire image into two fast inferences on small crops: a coarse segmentation on the Target Crop, and a local refinement on the Focus Crop. To make the model work with preexisting masks, we formulate a sub-task termed Inter-active Mask Correction, and propose Progressive Merge as the solution. Progressive Merge exploits morphological information to decide where to preserve and where to update, enabling users to refine any preexisting mask effectively. FocalClick achieves competitive results against SOTA methods with significantly smaller FLOPs. It also shows significant superiority when making corrections on preexisting masks. Code and data will be released at github.com/XavierCHEN34/ClickSEG
Xi Chen 0119, Zhiyan Zhao, Manni Duan, Donglian Qi, Hengshuang Zhao
CVPR5
2022 Revisiting the Transferability of Supervised Pretraining: an MLP Perspective
abstract
The pretrain-finetune paradigm is a classical pipeline in visual learning. Recent progress on unsupervised pretraining methods shows superior transfer performance to their supervised counterparts. This paper revisits this phenomenon and sheds new light on understanding the transferability gap between unsupervised and supervised pretraining from a multilayer perceptron (MLP) perspective. While previous works [6], [8], [17] focus on the effectiveness of MLP on unsupervised image classification where pretraining and evaluation are conducted on the same dataset, we reveal that the MLP projector is also the key factor to better transferability of unsupervised pretraining methods than supervised pretraining methods. Based on this observation, we attempt to close the transferability gap between supervised and unsupervised pretraining by adding an MLP projector before the classifier in supervised pretraining. Our analysis indicates that the MLP projector can help retain intra-class variation of visual features, decrease the feature distribution distance between pretraining and evaluation datasets, and reduce feature redundancy. Extensive experiments on public benchmarks demonstrate that the added MLP projector significantly boosts the transferability of supervised pretraining, e.g. +7.2% top-1 accuracy on the concept generalization task, +5.8% top-1 accuracy for linear evaluation on 12 -domain classification tasks, and +0.8% AP on COCO object detection task, making supervised pretraining comparable or even better than unsupervised pretraining.
Yizhou Wang 0007, Shixiang Tang, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Donglian Qi, Wanli Ouyang
CVPR6
2022 Learning Domain Adaptive Object Detection with Probabilistic Teacher
abstract
Self-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts.
Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu
ICML8
2022 Unsupervised Object Detection Pretraining with Joint Object Priors Generation and Detector Learning
abstract
Unsupervised pretraining methods for object detection aim to learn object discrimination and localization ability from large amounts of images. Typically, recent works design pretext tasks that supervise the detector to predict the defined object priors. They normally leverage heuristic methods to produce object priors, \emph{e.g.,} selective search, which separates the prior generation and detector learning and leads to sub-optimal solutions. In this work, we propose a novel object detection pretraining framework that could generate object priors and learn detectors jointly by generating accurate object priors from the model itself. Specifically, region priors are extracted by attention maps from the encoder, which highlights foregrounds. Instance priors are the selected high-quality output bounding boxes of the detection decoder. By assuming objects as instances in the foreground, we can generate object priors with both region and instance priors. Moreover, our object priors are jointly refined along with the detector optimization. With better object priors as supervision, the model could achieve better detection capability, which in turn promotes the object priors generation. Our method improves the competitive approaches by \textbf{+1.3 AP}, \textbf{+1.7 AP} in 1\% and 10\% COCO low-data regimes object detection.
Yizhou Wang 0007, Meilin Chen, Shixiang Tang, Feng Zhu 0006, Haiyang Yang, Lei Bai 0001, Rui Zhao 0001, Yunfeng Yan, Donglian Qi, Wanli Ouyang
NeurIPS9
2022 iNL: Implicit non-local network
Yifeng Han, Xi Chen 0119, Songjie Zhang, Donglian Qi
Neurocomputing4
2021 Online Pseudo Label Generation by Hierarchical Cluster Dynamics for Adaptive Person Re-identification
abstract
Adaptive person re-identification (adaptive ReID) targets at transferring learned knowledge from the labeled source domain to the unlabeled target domain. Pseudo-label-based methods that alternatively generate pseudo labels and optimize the training model have demonstrated great effectiveness in this field. However, the generated pseudo labels are inaccurate and cannot reflect the true semantic meaning of the unlabeled samples. We consider such inaccuracy stems from both the lagged update of the pseudo labels as well as the simple criterion of the employed clustering method. To tackle the problem, we propose an online pseudo label generation by hierarchical cluster dynamics for adaptive ReID. In particular, hierarchical label banks are constructed for all the samples in the dataset, and we update the pseudo labels of the sample in each coming mini-batch, performing the model optimization and the label generation simultaneously. A new hierarchical cluster dynamics is built for the label update, where cluster merge and cluster split are driven by a possibility computed by the label propagation. Our method can achieve better pseudo labels and higher reid accuracy. Extensive experiments on Market-to-Duke, Duke-to-Market, MSMT-to-Market, MSMT-to-Duke, Market-to-MSMT, and Duke-to-MSMT verify the effectiveness of our proposed method.
Shixiang Tang, Guolong Teng, Yixiao Ge, Kaijian Liu, Harry Qin, Donglian Qi, Dapeng Chen
ICCV7
2021 Neuromorphic extreme learning machines with bimodal memristive synapses
Zhekang Dong, Chun Sing Lai, Donglian Qi, Mingyu Gao 0002, Shukai Duan 0001
Neurocomputing4
2021 Distributed Dynamic Averaging Tracking Without Rate Measurements
abstract
This article introduces a dynamic average tracking algorithm for multiagent systems over weight-balanced direct graphs. Toward the end, we propose a proportional-integral distributive observer for each system with exponential convergence. In particular, the proposed algorithm only requires mild initialization conditions, does not rely on the time derivatives of reference signals and can ensure average consensus tracking with bounded steady-state error. The simulation results verified the performance of the proposed scheme.
Donglian Qi, Chaoyong Li
IEEE Trans. Syst. Man Cybern. Syst.2
2020 State-Aware Tracker for Real-Time Video Object Segmentation
abstract
In this work, we address the task of semi-supervised video object segmentation (VOS) and explore how to make efficient use of video property to tackle the challenge of semi-supervision. We propose a novel pipeline called State-Aware Tracker (SAT), which can produce accurate segmentation results with real-time speed. For higher efficiency, SAT takes advantage of the inter-frame consistency and deals with each target object as a tracklet. For more stable and robust performance over video sequences, SAT gets awareness for each state and makes self-adaptation via two feedback loops. One loop assists SAT in generating more stable tracklets. The other loop helps to construct a more robust and holistic target representation. SAT achieves a promising result of 72.3% J&F mean with 39 FPS on DAVIS 2017-Val dataset, which shows a decent trade-off between efficiency and accuracy.
Xi Chen 0119, Zuoxin Li, Ye Yuan 0012, Gang Yu 0002, Jianxin Shen, Donglian Qi
CVPR6
2020 Intraday Residential Demand Response Scheme Based on Peer-to-Peer Energy Trading
abstract
The intermittency introduced by the increasing integration of distributed renewable energy sources is challenging the efficient operation of residential distribution systems. A promising solution to tackle this challenge is the implementation of residential demand response through responsive household appliances such as heat pumps, refrigeration devices, and energy storage units. In this article, a peer-to-peer energy trading platform among residential houses is proposed to coordinate demand response schemes and level off potential generation/consumption disturbances in the hour-ahead intraday context. First, the day-ahead and intraday energy management models for residential houses are established considering the characteristics of responsive household appliances and energy storages. The discomfort and possible economic losses for performing demand responses are quantified with respect to the risk preferences of residential customers. The peer-to-peer energy trading platform is developed and a double-auction mechanism employed to promote the collaborative demand response schemes in the face of disturbances. An optimal bidding strategy of residential houses is also proposed. The feasibility of the proposed models and bidding strategy are verified through case studies. It is also illustrated that the residential demand response schemes and intraday peer-to-peer energy trading are effective in managing the uncertainties of load demand and renewable generation.
Donglian Qi, Fushuan Wen
IEEE Trans. Ind. Informatics2
2019 UltraComm: High-Speed and Inaudible Acoustic Communication
Xiaoyu Ji 0001, Donglian Qi, Wenyuan Xu 0001
QSHINE4
2019 Vision-based crater and rock detection using a cascade decision forest
abstract
Both crater and rock detection are components of the autonomous landing and hazard avoidance technology (ALHAT) sensor suite, as craters and rocks represent the majority of landing hazards. Furthermore, places with scientific values are very probable next to craters and rocks. Unsupervised approaches, which potentially use the pattern recognition techniques of ring threshold finding, perform quickly; however, they suffer from handling small craters. The supervised pattern recognition method is more powerful but is time‐consuming. To address these issues, here, a simultaneous multi‐size crater and rock detection algorithm is studied. The authors propose a new supervised machine‐learning framework using a cascade decision forest. Sliding windows are utilised in order to search basic features, and a multi‐grained cascade structure is introduced to enhance the framework's ability to learn the representations of the features. The training time of the proposed algorithm on a PC is comparable to that of deep neural networks, and the efficiency is enhanced for a large‐scale database. The outputs of the simulation verify the effectiveness and validity of the introduced technique.
Yunfeng Yan, Donglian Qi, Chaoyong Li
IET Comput. Vis.2
2018 A general memristor-based pulse coupled neural network with variable linking coefficient for multi-focus image fusion
Zhekang Dong, Chun Sing Lai, Donglian Qi, Zhao Xu 0002, Chaoyong Li, Shukai Duan 0001
Neurocomputing3
2017 A FO-ADRC Based Neutral-Point Potential Balancing for Three-Level Inverter
abstract
The neutral-point balancing control of neutral-point-clamped (NPC) three-level inverter is concerned. To solve this problem, two steps are considered. First, a new state-space model of DC bus voltage is established to extract the fluctuation of neutral-point potential (NPP) as a special disturbance. Then, an advanced voltage control strategy is designed based on fractional-order active disturbance rejection control (FO-ADRC) to suppress NPP. The developed FO-ADRC can not only compensate the fluctuation of NPP without a complex process of parameter turning, but enhance the robustness of the NPC inverter. Experiments are performed on a 10kW prototype of three-phase three-level inverter, and the efficiency is verified.
Zhenming Li, Guoyue Zhang, Donglian Qi
ICPADS5
2017 NIPAD: a non-invasive power-based anomaly detection scheme for programmable logic controllers
abstract
Industrial control systems (ICSs) are widely used in critical infrastructures, making them popular targets for attacks to cause catastrophic physical damage. As one of the most critical components in ICSs, the programmable logic controller (PLC) controls the actuators directly. A PLC executing a malicious program can cause significant property loss or even casualties. The number of attacks targeted at PLCs has increased noticeably over the last few years, exposing the vulnerability of the PLC and the importance of PLC protection. Unfortunately, PLCs cannot be protected by traditional intrusion detection systems or antivirus software. Thus, an effective method for PLC protection is yet to be designed. Motivated by these concerns, we propose a non-invasive powerbased anomaly detection scheme for PLCs. The basic idea is to detect malicious software execution in a PLC through analyzing its power consumption, which is measured by inserting a shunt resistor in series with the CPU in a PLC while it is executing instructions. To analyze the power measurements, we extract a discriminative feature set from the power trace, and then train a long short-term memory (LSTM) neural network with the features of normal samples to predict the next time step of a normal sample. Finally, an abnormal sample is identified through comparing the predicted sample and the actual sample. The advantages of our method are that it requires no software modification on the original system and is able to detect unknown attacks effectively. The method is evaluated on a lab testbed, and for a trojan attack whose difference from the normal program is around 0.63%, the detection accuracy reaches 99.83%.
Yujun Xiao, Wenyuan Xu 0001, Zhenhua Jia, Donglian Qi
Frontiers Inf. Technol. Electron. Eng.5
2015 A new game model for distributed optimization problems with directed communication topologies
Jianliang Zhang, Donglian Qi, Guangzhou Zhao
Neurocomputing2
2011 Optimal H∞ fusion filters for a class of discrete-time intelligent systems with time delays and missing measurement
Meiqin Liu 0001, Donglian Qi, Senlin Zhang, Meikang Qiu, Shiyou Zheng
Neurocomputing2
2010 Exponential H∞ synchronization of general discrete-time chaotic neural networks with or without time delays
abstract
This brief studies exponential H(infinity) synchronization of a class of general discrete-time chaotic neural networks with external disturbance. On the basis of the drive-response concept and H(infinity) control theory, and using Lyapunov-Krasovskii (or Lyapunov) functional, state feedback controllers are established to not only guarantee exponential stable synchronization between two general chaotic neural networks with or without time delays, but also reduce the effect of external disturbance on the synchronization error to a minimal H(infinity) norm constraint. The proposed controllers can be obtained by solving the convex optimization problems represented by linear matrix inequalities. Most discrete-time chaotic systems with or without time delays, such as Hopfield neural networks, cellular neural networks, bidirectional associative memory networks, recurrent multilayer perceptrons, Cohen-Grossberg neural networks, Chua's circuits, etc., can be transformed into this general chaotic neural network to be H(infinity) synchronization controller designed in a unified way. Finally, some illustrated examples with their simulations have been utilized to demonstrate the effectiveness of the proposed methods.
Donglian Qi, Meiqin Liu 0001, Meikang Qiu, Senlin Zhang
IEEE Trans. Neural Networks1