EDBT 2026 Demo / reviewers in the wild / expert
Qiang Wang 0015
dblp:64/5630-15
· DBLP profile ↗
30ranked-venue papers
10as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Consistency Representation Learning for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) exhibits a contradictory relationship between intra-domain discrimination and inter-domain gaps when learning from continuous data. Intra-domain discrimination focuses on individual nuances (i.e., clothing type, accessories,etc.), while inter-domain gaps emphasize domain consistency. Achieving a trade-off between maximizing intra-domain discrimination and minimizing inter-domain gaps is a crucial challenge for improving LReID performance. Most existing methods strive to reduce inter-domain gaps through knowledge distillation to maintain domain consistency. However, they often ignore intra-domain discrimination. To address this challenge, we propose a novel domain consistency representation learning (DCR) model that explores global and attribute-wise representations as a bridge to balance intra-domain discrimination and inter-domain gaps. At the intra-domain level, we explore the complementary relationship between global and attribute-wise representations to improve discrimination among similar identities. Excessive learning intra-domain discrimination can lead to catastrophic forgetting. We further develop an attribute-oriented anti-forgetting (AF) strategy that explores attribute-wise representations to enhance inter-domain consistency, and propose a knowledge consolidation (KC) strategy to facilitate knowledge transfer. Extensive experiments show that our DCR achieves superior performance compared to state-of-the-art LReID methods. Our code is available at https://github.com/LiuShiBen/DCR. Shiben Liu, Huijie Fan, Qiang Wang 0015, Weihong Ren, Yandong Tang, Yang Cong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PGFormer: Prompt guide network for underwater image enhancementabstractUnderwater images are often influenced by light scattering and refraction, which leads to color deviation and poor quality. The enhancement of underwater images is significant for high-level semantic learning but also challenging. In this paper, we introduce PGFormer, a novel underwater image enhancement network that leverages prompt priors by integrating global and local prior information to improve underwater image quality. PGFormer comprises a global local enhancement module (GLEM) and a prompt-guided forward feedback network (PGFN). The GLEM extracts robust feature information through global and local feature modulation, whereas PGFN introduces prompt information into the local optimization process to further enhance local expression and refinement. Extensive experiments on various underwater datasets show that our method outperforms existing state-of-the-art techniques in terms of both visual quality and quantitative performance. Xin Luan, Huijie Fan, Qiang Wang 0015, Yandong Tang |
CEC | 3 |
| 2025 | Joint Attention Mechanism and Multi-task Learning for Weakly Supervised Skin Image SegmentationabstractIn recent years, the application of deep learning technology in the field of medical image segmentation has become increasingly mature and has made certain progress. Applying this technique requires the use of a large number of medical image datasets with pixel-level annotations, however, the cost of pixel-level annotation of medical images is high. For this situation, we propose a weakly supervised medical image semantic segmentation model based on attention mechanism and multi-task learning. The attention mechanism can be used to focus on the information that is more critical to the current task in a large number of input information, and multi-task learning can share information among related tasks to promote network learning. As a result, the model can achieve better segmentation performance by using only image-level groundtruth labels. Since semantic segmentation task and saliency detection task are dense pixel prediction tasks,and class activation map (CAM) extracted from classification task is also used to represent pixel-level semantic information, it shows that these two tasks are highly related to the semantic segmentation task. Based on this premise,We propose the module of finding similarity from attention(SFA), which is used to calculate the similarity between multi-task feature maps to learn the correlation between tasks, and use similarity to optimize the prediction results of specific tasks. In addition, in order to extensively explore context relationships in dense pixel prediction tasks to achieve more accurate prediction results, we also propose a global context module (GCM) that pays attention to the global information of dense prediction tasks, captures long-distance dependencies in images, and improves the segmentation performance of the network. We did a lot of experiments, our method achieves 68.56% and 62.32% Miou on ISBI2016 and ISIC2017 datasets, respectively, and F1 -score is 78.25% on PH2 data set, significantly outperforming several recent state-of-the-art weakly supervised semantic segmentation methods. Yujianing Wang, Qiang Wang 0015, Huijie Fan |
CEC | 2 |
| 2025 | All-Day Multi-Camera Multi-Target TrackingabstractThe capability of tracking objects in low-light environments like nighttime is crucial for numerous real-world applications. However, previous Multi-Camera Multi-Target(MCMT) tracking methods are primarily focused on tracking during daytime with favorable lighting, overlooking the challenge posed by low-light conditions. The main difficulty of tracking under low-light condition is the lack of detailed visible appearance features. To address this issue, we incorporate the infrared modality into MCMT tracking framework to provide more useful information. We constructed the first Multi-modality (RGBT) Multi-camera Multi-target tracking dataset named M3Track, which contains sequences captured in low-light environments, laying a solid foundation for all-day multi-camera tracking. Based on the proposed dataset, we propose All-Day Multi-Camera Multi-Target tracking network, termed as ADM-CMT. Specifically, we propose an All-Day Mamba Fusion(ADMF) module to fuse information from different modalities adaptively. Within ADMF, the Lighting Guidance Model(LGM) extracts lighting relevant information to guide the fusion process. Furthermore, the Nearby Target Collection(NTC) strategy is designed to enhance tracking accuracy by leveraging information derived from surrounding objects of targets. Experiments conducted on M3Track demonstrate that ADMCMT exhibits strong generalization across different lighting conditions. The code will be released at https://github.com/QTRACKY/ADMCMT. Huijie Fan, Yihao Zhen, Tinghui Zhao, Baojie Fan, Qiang Wang 0015 |
CVPR | 6 |
| 2025 | FMambaIR: A Hybrid State-Space Model and Frequency Domain for Image RestorationabstractWith the development of deep learning, impressive progress has been made in the field of image restoration. The existing methods mainly rely on CNN and Transformer to obtain multi-scale feature information. However, these methods rarely integrate frequency domain information effectively during feature extraction, limiting their performance in image restoration. Additionally, few have combined Mamba with the Fourier domain for image restoration, which limits Mamba’s ability to perceive global degradation in the frequency domain. Therefore, we propose a new image restoration model called FMambaIR, which utilizes the complementarity between frequency and Mamba for image restoration. The core of FMambaIR is the F-Mamba block, which combines Fourier transform and Mamba for global degradation perception modeling. Specifically, F-Mamba adopts a dual branch complementary structure, including spatial Mamba branches and Fourier frequency domain global modeling. Mamba models the long-range dependencies of the entire image features, and the frequency branch utilizes Fourier to extract global degraded features from the image. Finally, we use a forward feedback network to integrate local information, which is beneficial for improving the recovery details. We comprehensively evaluate FMambaIR on several image restoration tasks, including underwater image enhancement, remote sensing image dehazing, and low-light image enhancement. The experimental results demonstrate that FMambaIR not only achieves superior performance compared to state-of-the-art methods but also significantly reduces computational complexity. Our code is available at https://github.com/mickoluan/FMambaIR. Xin Luan, Huijie Fan, Qiang Wang 0015, Shiben Liu, Xiaofeng Li 0001, Yandong Tang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Diverse Representations Embedding for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to continuously learn from sequential data streams, enabling cross-camera matching of individuals over time. A critical challenge in LReID lies in balancing the preservation of previously acquired knowledge with the incremental acquisition of new information, due to task-level gaps and limited representation capacity. Conventional methods relying on CNN backbones struggle to fully capture the diverse perspectives of each instance, leading to suboptimal model performance. To tackle these limitations, we propose a diverse representation embedding (DRE) framework that balances preserving old knowledge with adapting to new information. Specifically, our DRE incorporates a robust Transformer-based backbone that utilizes maximum embedding (ME) and multiple class tokens to generate overlapping representations for each instance. To further enhance the model's representation capacity, we design an adaptive constraint module (ACM), which performs integration and discrimination operations on overlapping representations to yield diverse yet diverse representations. Furthermore, we propose two strategies: knowledge update (KU) and knowledge preservation (KP), implemented within the adjustment and learner models, respectively. The KU strategy enhances the learner model's ability to adapt to new information by leveraging prior knowledge from the adjustment model. The KP strategy ensures the retention of historical knowledge while maintaining the model's adaptability. Extensive experiments validate that our DRE surpasses state-of-the-art approaches across large-scale, occluded, and holistic datasets, demonstrating significant performance gains. Our code is available at https://github.com/LiuShiBen/DRE. Shiben Liu, Huijie Fan, Qiang Wang 0015, Xi'ai Chen, Zhi Han, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Residual Denoising Diffusion ModelsabstractWe propose residual denoising diffusion models (RDDM), a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models, initially uninterpretable for image restoration, into a unified and interpretable model for both image generation and restoration by introducing residuals. Specifically, our residual diffusion represents directional diffusion from the target image to the degraded input image and explicitly guides the reverse generation process for image restoration, while noise diffusion represents random perturbations in the diffusion process. The residual prioritizes certainty, while the noise emphasizes diversity, enabling RDDM to effectively unify tasks with varying certainty or diversity requirements, such as image generation and restoration. We demonstrate that our sampling process is consistent with that of DDPM and DDIM through coefficient transformation, and propose a partially path-independent generation process to better understand the reverse process. Notably, our RDDM enables a generic UNet, trained with only an L1 loss and a batch size of 1, to compete with state-of-the-art image restoration methods. We provide code and pre-trained models to encourage further exploration, application, and development of our innovative framework (https://github.com/nachifurlRDDM). Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Yandong Tang, Liangqiong Qu |
CVPR | 2 |
| 2024 | Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionabstractRecently, Vision-Language Model (VLM) has greatly ad-vanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g., a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However, such approaches, only encoding the action-specific text prompts in vocabulary level, may suffer from learning ambiguity without exploring the fine-grained clues from the perspective of interaction context. In this paper, we propose a novel method to discover Syntactic Interaction Clues for HOI detection (SICHOI) by using VLM. Specifically, we first investigate what are the essen-tial elements for an interaction context, and then establish a syntactic interaction bank from three levels: spatial relationship, action-oriented posture and situational condition. Further, to align visual features with the syntactic interaction bank, we adopt a multi-view extractor to jointly aggre-gate visual features from instance, interaction, and image levels accordingly. In addition, we also introduce a dual cross-attention decoder to perform context propagation be-tween text knowledge and visual features, thereby enhancing the HOI detection. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on HICO-DET and V-COCO. Jinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen, Qiang Wang 0015, Zhi Han, Honghai Liu 0001 |
CVPR | 5 |
| 2024 | Feature distillation and guide network for unsupervised underwater image enhancement
Xin Luan, Qiang Wang 0015, Huijie Fan, Xiai Chen, Zhi Han, Yandong Tang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Unsupervised person re-identification based on adaptive information supplementation and foreground enhancementabstractAbstract Unsupervised person re‐identification has attracted vital interest because of its ability to protect privacy, significantly lower the expense of manual annotation, and eliminate the need for data labels. General unsupervised methods train the network only through global features, which causes the fine‐grained information contained in local features to be ignored in the recognition process, resulting in large amounts of label noise and affecting the recognition accuracy. Moreover, more robust pedestrian features can also improve the accuracy of clustering and enable unsupervised person re‐identification to obtain better results. To address these issues, first, a dual‐branch structure was proposed, which separately obtains the global features of the pedestrian and the local features by dividing the global features into a few equal sections. Then, an adaptive information supplementation (AIS) method based on the k‐nearest neighbor algorithm is designed to ascertain each local feature's relevance to the global features, calculating adaptive weight scores for information supplementation. Finally, these weight scores are used to reallocate the weights of the global features in each part, acquiring features that contain more pedestrian information during the representation learning process. These better features are used to reduce label noise to obtain more accurate pseudo‐labels. Second, an adaptive foreground enhancement module (AFEM) was proposed and inserted before clustering to increase the robustness of pedestrian features, which increases the precision of the pseudo‐labels that are produced after clustering. Experiments on Market‐1501, DukeMTMC‐reID, and MSMT17 demonstrate that the proposed method achieves better results than state‐of‐the‐art methods in fully unsupervised person re‐identification tasks. Qiang Wang 0015, Huijie Fan, Shengpeng Fu, Yandong Tang |
IET Image Process. | 1 |
| 2024 | Wavelet-pixel domain progressive fusion network for underwater image enhancement
Shiben Liu, Huijie Fan, Qiang Wang 0015, Zhi Han, Yandong Tang |
Knowl. Based Syst. | 3 |
| 2024 | Skip Connection Aggregation Transformer for Occluded Person ReidentificationabstractThe occlusion problem is a significant challenge for person reidentification. Recently, transformer-based methods have been introduced to solve the occlusion problem and achieve performance improvements. However, the existing methods only apply the features of the last transformer layer and fail to consider the alignment of visible body parts. They also ignore fine-grained local features. Thus, they usually suffer from misalignment in occluded image matching. We observe that features from the high layers of the transformer focus on classification information and global features, while those from the middle layers pay more attention to pedestrians. We think that making full use of the features of different layers will facilitate alignment and then will promote reidentification accuracy. Therefore, we propose a novel skip connection aggregation transformer (SCAT) network by utilizing features from different transformer layers to increase the diversity of features and align visible body parts in occluded images. The diverse features include the following: first, features of the middle layer, which focus on the pedestrian in nonoccluded regions and favor alignment, second, features of high layers, which focus on global information, third fine-grained local features, which are obtained by the part pooling encoder and the fusion reconstruction module. The part pooling encoder and the fusion reconstruction module are proposed to obtain part-based local features and fused local features, respectively. The experimental results on the occluded, partial, and holistic benchmarks demonstrate that our method can significantly promote the accuracy of occluded person reidentification. Huijie Fan, Qiang Wang 0015, Sheng-Peng Fu, Yandong Tang |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | QueryTrack: Joint-Modality Query Fusion Network for RGBT TrackingabstractExisting RGB-Thermal trackers usually treat intra-modal feature extraction and inter-modal feature fusion as two separate processes, therefore the mutual promotion of extraction and fusion is neglected. Then, the complementary advantages of RGB-T fusion are not fully exploited, and the independent feature extraction is not adaptive to modal quality fluctuation during tracking. To address the limitations, we design a joint-modality query fusion network, in which the intra-modal feature extraction and the inter-modal fusion are coupled together and promote each other via joint-modality queries. The queries are initialized based on the multimodal features of the current frame, making the subsequent fusion adaptive to modal quality fluctuation during tracking. Then the joint-modality query fusion (JQF) utilizes the queries to interact with RGB-T features, allowing the intra-modal enhancement and the inter-modal interactions to be unified for mutual promotion. In this way, JQF can distinguish and enhance the complementary modality features, while filtering out redundant information. For real-time tracking, we propose regional cross-attention for cross-modal interactions to reduce computational cost. Our end-to-end tracker sets a new state-of-the-art performance on multiple RGBT tracking benchmarks including LasHeR, VTUAV, RGBT234 and GTOT, while running at a real-time speed. Huijie Fan, Zhencheng Yu, Qiang Wang 0015, Baojie Fan, Yandong Tang |
IEEE Trans. Image Process. | 3 |
| 2024 | A Shadow Imaging Bilinear Model and Three-Branch Residual Network for Shadow RemovalabstractThe current shadow removal pipeline relies on the detected shadow masks, which have limitations for penumbras and tiny shadows, and results in an excessively long pipeline. To address these issues, we propose a shadow imaging bilinear model and design a novel three-branch residual (TBR) network for shadow removal. Our bilinear model reveals the single-image shadow removal process and can explain why simply increasing the brightness of shadow areas cannot remove shadows without artifacts. We considerably shorten the shadow removal pipeline by modeling illumination compensation and developing a single-stage shadow removal network without additional detection and refinement networks. Specifically, our network consists of three task branches, i.e., shadow image reconstruction, shadow matte estimation, and shadow removal. To merge these three branches and enhance the shadow removal branch, we design a model-based TBR module. Multiple TBR modules are cascaded to generate an intensive information flow and facilitate feature integration among the three branches. Thus, our network ensures the fidelity of nonshadow areas and restores the light intensity of shadow areas through three-branch collaboration. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The model and code are available at https://github.com/nachifur/TBRNet. Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Jiandong Tian, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Underwater image enhancement via a channel-wise transmission estimation networkabstractAbstract Underwater image enhancement for image processing and underwater robotic vision have recently attracted much academic attention. However, in most existing methods, underwater image enhancement is completed with a simple assumption: the attenuation coefficients are unified across the color channels. This assumption leads to unstable and visually unpleasing enhancement results. Moreover, these methods cannot be successfully applied to explore relatively independent transmissions from multiple color channels with complimentary feature information. To address these challenges, a novel channel‐wise transmission estimation network (CTEN) is proposed, which aims to pioneer the exploration of the transmission difference across the color channels in an underwater scene. Specifically, a color‐specific correction module is proposed to automatically quantify the transmission ability of multiple color channels in the underwater environment. Furthermore, a channel‐wise transmission estimation module is designed to simultaneously explore the relative independence of multi‐color channels and estimate the medium transmissions for each color channel, which represents the attenuation degree of different color radiances after reflecting in the water. Then, a novel residual strategy is introduced to integrate these two modules to complete the underwater enhancement. Using the model, the authors are able to provide an answer as to why channel‐wise transmission estimation are better than single transmission estimation and establish a generalization theory to show the effect of the independent transmission estimation model for each color channel. Experiments on several underwater image datasets verify the superiority of the proposed CTEN model. Qiang Wang 0015, Huijie Fan |
IET Image Process. | 1 |
| 2023 | Region Selective Fusion Network for Robust RGB-T TrackingabstractRGB-T tracking utilizes thermal infrared images as a complement to visible light images in order to perform more robust visual tracking in various scenarios. However, the highly aligned RGB-T image pairs introduces redundant information, the modal quality fluctuation during tracking also brings unreliable information. Existing RGB-T trackers usually use channelwise multi-modal feature fusion in which the low-quality features degrades the fused features and causes trackers to drift. In this work, we propose a region selective fusion network that first evaluates each image region by cross-modal and cross-region modeling, then removes low-quality redundant region features to alleviate the negative effects caused by unreliable information in multi-modal fusion. Besides, the region removal scheme brings a efficiency boost as redundant features are removed progressively, this enables the tracker to run at a high tracking speed.Extensive experiments show that the proposed tracker achieves competitive performance with a real-time tracking speed on multiple RGB-T tracking benchmarks including LasHeR, RGBT234 and GTOT. Zhencheng Yu, Huijie Fan, Qiang Wang 0015, Ziwan Li, Yandong Tang |
IEEE Signal Process. Lett. | 3 |
| 2023 | A Decoupled Multi-Task Network for Shadow RemovalabstractShadow removal, which aims to restore the illumination in shadow regions, is challenging due to the diversity of shadows in terms of location, intensity, shape, and size. Different from most multi-task methods, which design elaborate multi-branch or multi-stage structures for better shadow removal, we introduce feature decomposition to learn better feature representations. Specifically, we propose a single-stage and decoupled multi-task network (DMTN) to explicitly learn the decomposed features for shadow removal, shadow matte estimation, and shadow image reconstruction. First, we propose several coarse-to-fine semi-convolution (SMC) modules to capture features sufficient for joint learning of these three tasks. Second, we design a theoretically supported feature decoupling layer to explicitly decouple the learned features into shadow image features and shadow matte features via weight reassignment. Last, these features are converted to a target shadow-free image, affiliated shadow matte, and shadow image, supervised by multi-task joint loss functions. With multi-task collaboration, DMTN effectively recovers the illumination in shadow areas while ensuring the fidelity of non-shadow areas. Experimental results show that DMTN competes favorably with state-of-the-art multi-branch/multi-stage shadow removal methods, while maintaining the simplicity of single-stage methods. We have released our code to encourage future exploration in powerful feature representation for shadow removalhttps://github.com/nachifur/DMTN Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Liangqiong Qu, Yandong Tang |
IEEE Trans. Multim. | 2 |
| 2022 | Towards collaborative appearance and semantic adaptation for medical image segmentation
Qiang Wang 0015, Yingkui Du, Huijie Fan |
Neurocomputing | 1 |
| 2022 | Data Poisoning Attacks on Federated Machine LearningabstractFederated machine learning which enables resource-constrained node devices (e.g., Internet of Things (IoT) devices and smartphones) to establish a knowledge-shared model while keeping the raw data local, could provide privacy preservation, and economic benefit by designing an effective communication protocol. However, this communication protocol can be adopted by attackers to launch data poisoning attacks for different nodes, which has been shown as a big threat to most machine learning models. Therefore, we in this article intend to study the model vulnerability of federated machine learning, and even on IoT systems. To be specific, we here attempt to attacking a popular federated multitask learning framework, which uses a general multitask learning framework to handle statistical challenges in the federated learning setting. The problem of calculating optimal poisoning attacks on federated multitask learning is formulated as a bilevel program, which is adaptive to the arbitrary selection oftargetnodes andsource attackingnodes. We then propose a novel systems-aware optimization method, called as attack on federated learning (AT2FL), to efficiently derive the implicit gradients for poisoned data, and further attain optimal attack strategies in the federated machine learning. This is an earlier work, to our knowledge, that explores attacking federated machine learning via data poisoning. Finally, experiments on several real-world data sets demonstrate that when the attackers directly poison thetargetnodes or indirectly poison the related nodes via using the communication protocol, the federated multitask learning model is sensitive to both poisoning attacks. Gan Sun, Yang Cong, Jiahua Dong 0001, Qiang Wang 0015, Lingjuan Lyu, Ji Liu 0002 |
IEEE Internet Things J. | 4 |
| 2022 | APAN: Across-Scale Progressive Attention Network for Single Image DerainingabstractRecent single image deraining works have achieved significant improvement using convolutional neural networks. However, the rain streaks in the rain image share similar patterns with its multi-scale versions, which are not fully exploited in recent works. In this paper, we propose anAcross-scaleProgressiveAttentionNetwork (i.e.,APAN) to explore the multi-scale collaborative representation for single image deraining. Specifically, we represent each rainy image via a multi-scale module. An across-scale attention module is then used to capture long-range feature correspondences from multi-scale features, which can model the rain streaks at an enlarging feature dimension. Afterwards, we construct a pyramid structure and further predict the rain streak progressively, which also guides the across-scale attention module to refine the feature representation from coarse to fine. The proposed model exploits self-similarity of features via an across-scale attention between different scales, which can well model the rain streak with long-range information. Experiments on several datasets show that our model achieves significant improvement compared with most state-of-the-art deraining models. Qiang Wang 0015, Gan Sun, Huijie Fan, Yandong Tang |
IEEE Signal Process. Lett. | 1 |
| 2022 | Continuous Multi-View Human Action RecognitionabstractHuman action recognition which recognizes human actions in a video is a fundamental task in computer vision field. Although multiple existing methods with single-view or multi-view have been presented for human action recognition, these recognition approaches cannot be extended into new action recognition or action classification tasks, as well as discover underlying correlations among different views. To tackle the above problem, this paper proposes a new lifelong multi-view subspace learning framework for continuous human action recognition, which could exploit the complementary information amongst different views from a lifelong learning perspective. More specifically, a set of view-specific libraries is established to gradually store the useful information within multiple views. As a new action recognition task comes, we decompose the model parameters into a set of embedded parameters over view-specific libraries. A latent representation subspace is constructed via encouraging it to be close to different view-specific libraries, which can leverage the high-order correlations among different views and further avoid partial information for action recognition task. Meanwhile, we propose to employ an alternating direction strategy to optimize our proposed method. Empirical studies on real-world multi-view action recognition datasets have shown that our proposed framework attains the superior recognition performance and saves the computational time when continually learning new action recognition tasks. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Qianqian Wang 0001, Zhengming Ding |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | PFDN: Pyramid Feature Decoupling Network for Single Image DerainingabstractRestoring images degraded by rain has attracted more academic attention since rain streaks could reduce the visibility of outdoor scenes. However, most existing deraining methods attempt to remove rain while recovering details in a unified framework, which is an ideal and contradictory target in the image deraining task. Moreover, the relative independence of rain streak features and background features is usually ignored in the feature domain. To tackle these challenges above, we propose an effective Pyramid Feature Decoupling Network (i.e., PFDN) for single image deraining, which could accomplish image deraining and details recovery with the corresponding features. Specifically, the input rainy image features are extracted via a recurrent pyramid module, where the features for the rainy image are divided into two parts, i.e., rain-relevant and rain-irrelevant features. Afterwards, we introduce a novel rain streak removal network for rain-relevant features and remove the rain streak from the rainy image by estimating the rain streak information. Benefiting from lateral outputs, we propose an attention module to enhance the rain-irrelevant features, which could generate spatially accurate and contextually reliable details for image recovery. For better disentanglement, we also enforce multiple causality losses at the pyramid features to encourage the decoupling of rain-relevant and rain-irrelevant features from the high to shallow layers. Extensive experiments demonstrate that our module can well model the rain-relevant information over the domain of the feature. Our framework empowered by PFDN modules significantly outperforms the state-of-the-art methods on single image deraining with multiple widely-used benchmarks, and also shows superiority in the fully-supervised domain. Qiang Wang 0015, Gan Sun, Jiahua Dong 0001, Yulun Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Recurrent Generative Adversarial Network for Face CompletionabstractMost recently-proposed face completion algorithms use high-level features extracted from convolutional neural networks (CNNs) to recover semantic texture content. Although the completed face is natural-looking, the synthesized content still lacks lots of high-frequency details, since the high-level features cannot supply sufficient spatial information for details recovery. To tackle this limitation, in this paper, we propose aRecurrentGenerativeAdversarialNetwork (RGAN) for face completion. Unlike previous algorithms, RGAN can take full advantage of multi-level features, and further provide advanced representations from multiple perspectives, which can well restore spatial information and details in face completion. Specifically, our RGAN model is composed of a CompletionNet and a DisctiminationNet, where the CompletionNet consists of two deep CNNs and a recurrent neural network (RNN). The first deep CNN is presented to learn the internal regulations of a masked image and represent it with multi-level features. The RNN model then exploits the relationships among the multi-level features and transfers these features in another domain, which can be used to complete the face image. Benefiting from bidirectional short links, another CNN is used to fuse multi-level features transferred from RNN and reconstruct the face image in different scales. Meanwhile, two context discrimination networks in the DisctiminationNet are adopted to ensure the completed image consistency globally and locally. Experimental results on benchmark datasets demonstrate qualitatively and quantitatively that our model performs better than the state-of-the-art face completion models, and simultaneously generates realistic image content and high-frequency details. The code will be released available soon. Qiang Wang 0015, Huijie Fan, Gan Sun, Weihong Ren, Yandong Tang |
IEEE Trans. Multim. | 1 |
| 2020 | Dually Connected Deraining Net Using Pixel-Wise AttentionabstractRecent single image deraining methods either use a recurrent mechanism to gradually learn the mapping between clear images and rainy images, or focus on designing various loss functions to supervise the learning process. In this letter, we propose a dually connected deraining net using pixel-wise attention, for single image rain removal. Specifically, the deraining net adopts an encoder-decoder net as a backbone, which can effectively learn a residual rain-streaks map by jointly using skip sum connection and skip concatenation connection. The dual connections enable the deraining net to promote information flow between layers, and thus can allow it to discriminate and localize the rain streaks. To preserve image details, the decoded features are weighted by the learnable pixel-wise attention for adaptively recalibrating their responses. Experimental results on synthetic datasets demonstrate that the proposed model outperforms the recent state-of-the-art deraining methods. Weihong Ren, Jiandong Tian, Qiang Wang 0015, Yandong Tang |
IEEE Signal Process. Lett. | 3 |
| 2019 | Weighted locality collaborative representation based on sparse subspace
Huaxiang Zhang 0001, Lei Zhu 0002, Wenbo Wan, Zhenhua Wang 0004, Qiang Wang 0015, Peilian Guo, Jiande Sun 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2019 | Laplacian pyramid adversarial network for face completion
Qiang Wang 0015, Huijie Fan, Gan Sun, Yang Cong, Yandong Tang |
Pattern Recognit. | 1 |
| 2019 | Deeply Supervised Face Completion With Multi-Context Generative Adversarial NetworkabstractRecent face completion works have achieved significant improvement using generative adversarial networks (GANs). There are still two important issues in this challenging task: first, semantic understanding; and second, high-frequency details prediction. In this letter, we propose a unified model by introducing multi-context structures within GANs. Our model, named multi-context generative adversarial networks (MCGAN), automatically learns the hierarchical appearances of a corrupted image and predicted the missing regions from different perspectives. In this model, semantic understanding and high-frequency details are both taken into account and modeled with two parallel networks, respectively. While one learns the semantic understanding of the input face image at a high level, the other extracts low-level features for high-frequency details prediction. Our MCGAN takes full advantage of multi-scale features learned from two complementary networks and generates semantically new pixels for the missing region with fine details. Extensive quantitative and qualitative experiments on benchmark datasets show that the proposed model outperforms several state-of-the-art models. Qiang Wang 0015, Huijie Fan, Yandong Tang |
IEEE Signal Process. Lett. | 1 |
| 2018 | Online Low-Rank Metric Learning via Parallel Coordinate Descent Methodabstract11The corresponding author is Prof. Yang Cong. This work is supported by Nature Science Foundation of China under Grant (61722311, U1613214, 61533015) and CAS-Youth Innovation Promotion Association Scholarship (2012163)Recently, many machine learning problems rely on a valuable tool: metric learning. However, in many applications, large-scale applications embedded in high-dimensional feature space may induce both computation and storage requirements to grow quadratically. In order to tackle these challenges, in this paper, we intend to establish a robust metric learning formulation with the expectation that online metric learning and parallel optimization can solve large-scale and high-dimensional data efficiently, respectively. Specifically, based on the matrix factorization strategy, the first step aims to learn a similarity function in the objective formulation for similarity measurement; in the second step, we derive a variational trace norm to promote low-rankness on the transformation matrix. After converting this variational regularization into its separable form, for the model optimization, we present an parallel block coordinate descent method to learn the optimal metric parameters, which can handle the high-dimensional data in an efficient way. Crucially, our method shares the efficiency and flexibility of block coordinate descent method, and it is also guaranteed to converge to the optimal solution. Finally, we evaluate our approach by analyzing scene categorization dataset with tens of thousands of dimensions, and the experimental results show the effectiveness of our proposed model. Gan Sun, Yang Cong, Qiang Wang 0015, Xiaowei Xu 0001 |
ICPR | 3 |
| 2018 | Joint graph regularization based modality-dependent cross-media retrieval
Jihong Yan, Huaxiang Zhang 0001, Jiande Sun 0001, Qiang Wang 0015, Peilian Guo, Lili Meng, Wenbo Wan |
Multim. Tools Appl. | 4 |
| 2017 | Large receptive field convolutional neural network for image super-resolutionabstractThis paper presents a new approach to Single Image Super Resolution (SISR), based upon Convolutional Neural Network (CNN). Although the SISR is ill-posed which can be seen as finding a non-linear mapping from a low to high-dimensional space. Deep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration and non-linear mapping problems. We consider the single image Super-Resolution (SR) problem as convolution operators and develop a CNN to capture the characteristics of Low-Resolution (LR) input image. We find that increasing the receptive field shows the improvement in accuracy. Our solution is to establish the connection between traditional optimization-based schemes and neural network architectures. In the paper a novel, separable structure is introduced as a reliable support for robust convolution against artifacts. Our proposed method performs better than existing methods in terms of accuracy and visual improvements in our results are easily noticeable. Qiang Wang 0015, Huijie Fan, Yang Cong, Yandong Tang |
ICIP | 1 |