Yandong Tang

dblp:29/1922 · DBLP profile ↗
← Back
83ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0003-3805-7654ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 23 since 2021Artificial intelligence and machine learning · 35 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment
abstract
We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM's hidden states with visual representations from external visual foundational models via a global visual alignment loss and a hybrid token, . This token enforces dual constraints: local next-token prediction and global semantic distillation, enabling LLMs to implicitly learn spatial and contextual coherence while retaining their original autoregressive paradigm. Extensive experiments validate ARRA's plug-and-play versatility. When training T2I LLMs from scratch, ARRA reduces FID by 16.6% (ImageNet), 12.0% (LAION-COCO) for autoregressive LLMs like LlamaGen, without modifying original architecture and inference mechanism. For training from text-generation-only LLMs, ARRA reduces FID by 25.5% (MIMIC-CXR), 8.8% (DeepEyeNet) for advanced LLMs like Chameleon. For domain adaptation, ARRA aligns general-purpose LLMs with specialized models (e.g., BioMedCLIP), achieving an 18.6% FID reduction over direct fine-tuning on medical imaging (MIMIC-CXR). These results demonstrate that training objective redesign, rather than architectural modifications, can resolve cross-modal global coherence challenges. ARRA offers a complementary paradigm for advancing autoregressive models.
Jiawei Liu 0003, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu
AAAI6
2026 Domain Consistency Representation Learning for Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) exhibits a contradictory relationship between intra-domain discrimination and inter-domain gaps when learning from continuous data. Intra-domain discrimination focuses on individual nuances (i.e., clothing type, accessories,etc.), while inter-domain gaps emphasize domain consistency. Achieving a trade-off between maximizing intra-domain discrimination and minimizing inter-domain gaps is a crucial challenge for improving LReID performance. Most existing methods strive to reduce inter-domain gaps through knowledge distillation to maintain domain consistency. However, they often ignore intra-domain discrimination. To address this challenge, we propose a novel domain consistency representation learning (DCR) model that explores global and attribute-wise representations as a bridge to balance intra-domain discrimination and inter-domain gaps. At the intra-domain level, we explore the complementary relationship between global and attribute-wise representations to improve discrimination among similar identities. Excessive learning intra-domain discrimination can lead to catastrophic forgetting. We further develop an attribute-oriented anti-forgetting (AF) strategy that explores attribute-wise representations to enhance inter-domain consistency, and propose a knowledge consolidation (KC) strategy to facilitate knowledge transfer. Extensive experiments show that our DCR achieves superior performance compared to state-of-the-art LReID methods. Our code is available at https://github.com/LiuShiBen/DCR.
Shiben Liu, Huijie Fan, Qiang Wang 0015, Weihong Ren, Yandong Tang, Yang Cong
IEEE Trans. Circuits Syst. Video Technol.5
2026 Lightweight Temporal-Frequency Perception Sparse State Space Models for Unified Image Restoration
abstract
Unified image restoration has become a fundamental issue in image processing. State space models have demonstrated significant potential in image restoration. However, their multi-directional scanning mechanism may introduce computational and feature redundancy, failing to satisfy lightweight deployment requirements. Furthermore, state space models have limitations in perceiving local detail features. To address this, we propose a lightweight channel-adaptive temporal-frequency sparse state space model for unified image restoration. This model enhances the local detail perception capability of the state space model using frequency domain features and simplifies the complexity of the network by sparse mechanisms. Specifically, we designed a U-shaped image restoration deep network based on the channel-adaptive temporal-frequency sparse state space module. This module consists of a temporal-domain dynamic sparse visual state space module and a frequency-domain sparse wavelet detail enhancement module in parallel, and uses a channel shuffling operation to realize temporal-frequency feature fusion. The dynamic sparse state space module uses a top-k mechanism to sparsify features across different scan paths for computational efficiency. The frequency-domain sparse wavelet detail enhancement module utilizes wavelet transformation and convolution operations to extract and enhance details in different directions, and then uses a top-k mechanism to perform sparse processing. Moreover, we introduce a degradation semantic perception module at the end of the encoder to guide the restoration network to adaptively learn the semantics of different degradation types, thereby realizing unified image restoration in complex outdoor environments. Extensive experimental results demonstrate that our method significantly outperforms 31 baseline methods in five complex weather and illumination degradation image restoration tasks while maintaining the lowest parameters and FLOPs.
Pengyue Li, Yinke Dou, Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.6
2026 DVG-Diffusion: Dual-View-Guided Diffusion Model for CT Reconstruction From X-Rays
abstract
Directly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D CT volume. In this work, we facilitate complex 2D X-ray image to 3D CT mapping by incorporating new view synthesis, and reduce the learning difficulty through view-guided feature alignment. Specifically, we propose a dual-view guided diffusion model (DVG-Diffusion), which couples a real input X-ray view and a synthesized new X-ray view to jointly guide CT reconstruction. First, a novel view parameter-guided encoder captures features from X-rays that are spatially aligned with CT. Next, we concatenate the extracted dual-view features as conditions for the latent diffusion model to learn and refine the CT latent representation. Finally, the CT latent representation is decoded into a CT volume in pixel space. By incorporating view parameter guided encoding and dual-view guided CT reconstruction, our DVG-Diffusion can achieve an effective balance between high fidelity and perceptual quality for CT reconstruction. Experimental results demonstrate our method outperforms state-of-the-art methods. Based on experiments, the comprehensive analysis and discussions for views and reconstruction are also presented. The model and code are available at https://github.com/xiexing0916/DVG-Diffusion.
Jiawei Liu 0003, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu
IEEE Trans. Image Process.5
2026 CAD-Mesher: A Convenient, Accurate, Dense Mesh-Based Mapping Module in SLAM for Dynamic Environments
abstract
Most LiDAR odometry and SLAM systems construct maps in point clouds, which are discrete and sparse when zoomed in, making them not directly suitable for navigation. Mesh maps represent a dense and continuous map format with low memory consumption, which can approximate complex structures with simple elements, attracting significant attention of researchers in recent years. However, most existing methods operate under a static environment assumption. In effect, moving objects cause ghosting, degrading the quality of meshing. To address these issues, we propose a plug-and-play meshing module adapting to dynamic environments, which can easily integrate with various LiDAR odometry to generally improve the pose estimation accuracy of odometry. In our meshing module, a novel two-stage coarse-to-fine dynamic removal method is designed to effectively filter dynamic objects, generating consistent, accurate, and dense mesh maps. To the best of our knowledge, this is the first mesh construction method with explicit dynamic removal. Additionally, sliding window-based keyframe aggregation and adaptive downsampling strategies are used to ensure the uniformity of point cloud, benefiting for Gaussian process in mesh construction. We evaluate the localization and mapping accuracy on six publicly available datasets. Extensive experiments demonstrate the superiority of our method compared with the state-of-the-art algorithms. The code and introduction video are publicly available at https://yaepiii.github.io/CAD-Mesher/.
Yanpeng Jia, Fengkui Cao, Ting Wang 0018, Yandong Tang, Shiliang Shao, Lianqing Liu
IEEE Trans. Multim.4
2025 PGFormer: Prompt guide network for underwater image enhancement
abstract
Underwater images are often influenced by light scattering and refraction, which leads to color deviation and poor quality. The enhancement of underwater images is significant for high-level semantic learning but also challenging. In this paper, we introduce PGFormer, a novel underwater image enhancement network that leverages prompt priors by integrating global and local prior information to improve underwater image quality. PGFormer comprises a global local enhancement module (GLEM) and a prompt-guided forward feedback network (PGFN). The GLEM extracts robust feature information through global and local feature modulation, whereas PGFN introduces prompt information into the local optimization process to further enhance local expression and refinement. Extensive experiments on various underwater datasets show that our method outperforms existing state-of-the-art techniques in terms of both visual quality and quantitative performance.
Xin Luan, Huijie Fan, Qiang Wang 0015, Yandong Tang
CEC5
2025 Multimodal prompt state space models for unified adverse weather removal
Pengyue Li, Jiandong Tian, Yandong Tang
Eng. Appl. Artif. Intell.4
2025 Rope-net: deep convolutional neural network via robust principal component analysis
Baichen Liu, Zhi Han, Xi'ai Chen, Yandong Tang
Mach. Learn.5
2025 A Subspace-Based Method for Facial Image Editing
abstract
In the realm of computational social systems, the ability to edit facial attributes accurately plays a crucial role in enhancing user experience on social media platforms and virtual environments. However, we face significant challenges in isolated attribute manipulation and balancing the tradeoff between editing fidelity and facial identity preservation. Here, this article presents a novel approach to constructing an orthogonal decomposition subspace, enabling precise editing control over individual attributes with minimal impact on others and maintaining identity consistency. We introduce an adaptive weight modulation (AWM) method and a maximum slope truncation (MST) formula. The AWM method, founded on a sufficient convergent criterion, performs singular value decomposition to yield subspace parameters that preserve rich facial knowledge within the generative model, facilitating high-quality facial generation with reduced parameterization. This empowers meaningful semantic interpretation of attributes, supporting diverse editing tasks such as pose, age, and eyewear adjustments. The MST formula rigorously defines the editing bounds to effectively navigate the tradeoff between editing depth and identity retention. We also propose a guideline for deciphering the specific meanings of unsupervised semantics, potentially advancing interpretability in social behavioral studies. An accompanying web application, available athttps://github.com/mickoluan/GreenLimeSia, has been developed, granting users the freedom to perform tailored facial edits. Extensive experimental results show we pave the way for more personalized and authentic interactions within computational social platforms.
MengChu Zhou, Xin Luan, Liang Qi 0001, Yandong Tang, Zhi Han
IEEE Trans. Comput. Soc. Syst.5
2025 FMambaIR: A Hybrid State-Space Model and Frequency Domain for Image Restoration
abstract
With the development of deep learning, impressive progress has been made in the field of image restoration. The existing methods mainly rely on CNN and Transformer to obtain multi-scale feature information. However, these methods rarely integrate frequency domain information effectively during feature extraction, limiting their performance in image restoration. Additionally, few have combined Mamba with the Fourier domain for image restoration, which limits Mamba’s ability to perceive global degradation in the frequency domain. Therefore, we propose a new image restoration model called FMambaIR, which utilizes the complementarity between frequency and Mamba for image restoration. The core of FMambaIR is the F-Mamba block, which combines Fourier transform and Mamba for global degradation perception modeling. Specifically, F-Mamba adopts a dual branch complementary structure, including spatial Mamba branches and Fourier frequency domain global modeling. Mamba models the long-range dependencies of the entire image features, and the frequency branch utilizes Fourier to extract global degraded features from the image. Finally, we use a forward feedback network to integrate local information, which is beneficial for improving the recovery details. We comprehensively evaluate FMambaIR on several image restoration tasks, including underwater image enhancement, remote sensing image dehazing, and low-light image enhancement. The experimental results demonstrate that FMambaIR not only achieves superior performance compared to state-of-the-art methods but also significantly reduces computational complexity. Our code is available at https://github.com/mickoluan/FMambaIR.
Xin Luan, Huijie Fan, Qiang Wang 0015, Shiben Liu, Xiaofeng Li 0001, Yandong Tang
IEEE Trans. Geosci. Remote. Sens.7
2025 Diverse Representations Embedding for Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) aims to continuously learn from sequential data streams, enabling cross-camera matching of individuals over time. A critical challenge in LReID lies in balancing the preservation of previously acquired knowledge with the incremental acquisition of new information, due to task-level gaps and limited representation capacity. Conventional methods relying on CNN backbones struggle to fully capture the diverse perspectives of each instance, leading to suboptimal model performance. To tackle these limitations, we propose a diverse representation embedding (DRE) framework that balances preserving old knowledge with adapting to new information. Specifically, our DRE incorporates a robust Transformer-based backbone that utilizes maximum embedding (ME) and multiple class tokens to generate overlapping representations for each instance. To further enhance the model's representation capacity, we design an adaptive constraint module (ACM), which performs integration and discrimination operations on overlapping representations to yield diverse yet diverse representations. Furthermore, we propose two strategies: knowledge update (KU) and knowledge preservation (KP), implemented within the adjustment and learner models, respectively. The KU strategy enhances the learner model's ability to adapt to new information by leveraging prior knowledge from the adjustment model. The KP strategy ensures the retention of historical knowledge while maintaining the model's adaptability. Extensive experiments validate that our DRE surpasses state-of-the-art approaches across large-scale, occluded, and holistic datasets, demonstrating significant performance gains. Our code is available at https://github.com/LiuShiBen/DRE.
Shiben Liu, Huijie Fan, Qiang Wang 0015, Xi'ai Chen, Zhi Han, Yandong Tang
IEEE Trans. Neural Networks Learn. Syst.6
2024 Residual Denoising Diffusion Models
abstract
We propose residual denoising diffusion models (RDDM), a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models, initially uninterpretable for image restoration, into a unified and interpretable model for both image generation and restoration by introducing residuals. Specifically, our residual diffusion represents directional diffusion from the target image to the degraded input image and explicitly guides the reverse generation process for image restoration, while noise diffusion represents random perturbations in the diffusion process. The residual prioritizes certainty, while the noise emphasizes diversity, enabling RDDM to effectively unify tasks with varying certainty or diversity requirements, such as image generation and restoration. We demonstrate that our sampling process is consistent with that of DDPM and DDIM through coefficient transformation, and propose a partially path-independent generation process to better understand the reverse process. Notably, our RDDM enables a generic UNet, trained with only an L1 loss and a batch size of 1, to compete with state-of-the-art image restoration methods. We provide code and pre-trained models to encourage further exploration, application, and development of our innovative framework (https://github.com/nachifurlRDDM).
Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Yandong Tang, Liangqiong Qu
CVPR5
2024 Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open World
abstract
Universal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value.
Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han
ACM Multimedia5
2024 Feature distillation and guide network for unsupervised underwater image enhancement
Xin Luan, Qiang Wang 0015, Huijie Fan, Xiai Chen, Zhi Han, Yandong Tang
Eng. Appl. Artif. Intell.6
2024 Unsupervised person re-identification based on adaptive information supplementation and foreground enhancement
abstract
Abstract Unsupervised person re‐identification has attracted vital interest because of its ability to protect privacy, significantly lower the expense of manual annotation, and eliminate the need for data labels. General unsupervised methods train the network only through global features, which causes the fine‐grained information contained in local features to be ignored in the recognition process, resulting in large amounts of label noise and affecting the recognition accuracy. Moreover, more robust pedestrian features can also improve the accuracy of clustering and enable unsupervised person re‐identification to obtain better results. To address these issues, first, a dual‐branch structure was proposed, which separately obtains the global features of the pedestrian and the local features by dividing the global features into a few equal sections. Then, an adaptive information supplementation (AIS) method based on the k‐nearest neighbor algorithm is designed to ascertain each local feature's relevance to the global features, calculating adaptive weight scores for information supplementation. Finally, these weight scores are used to reallocate the weights of the global features in each part, acquiring features that contain more pedestrian information during the representation learning process. These better features are used to reduce label noise to obtain more accurate pseudo‐labels. Second, an adaptive foreground enhancement module (AFEM) was proposed and inserted before clustering to increase the robustness of pedestrian features, which increases the precision of the pseudo‐labels that are produced after clustering. Experiments on Market‐1501, DukeMTMC‐reID, and MSMT17 demonstrate that the proposed method achieves better results than state‐of‐the‐art methods in fully unsupervised person re‐identification tasks.
Qiang Wang 0015, Huijie Fan, Shengpeng Fu, Yandong Tang
IET Image Process.5
2024 CCR: Facial Image Editing with Continuity, Consistency and Reversibility
Xin Luan, Huidi Jia, Zhi Han, Xiaofeng Li 0001, Yandong Tang
Int. J. Comput. Vis.6
2024 Wavelet-pixel domain progressive fusion network for underwater image enhancement
Shiben Liu, Huijie Fan, Qiang Wang 0015, Zhi Han, Yandong Tang
Knowl. Based Syst.6
2024 Skip Connection Aggregation Transformer for Occluded Person Reidentification
abstract
The occlusion problem is a significant challenge for person reidentification. Recently, transformer-based methods have been introduced to solve the occlusion problem and achieve performance improvements. However, the existing methods only apply the features of the last transformer layer and fail to consider the alignment of visible body parts. They also ignore fine-grained local features. Thus, they usually suffer from misalignment in occluded image matching. We observe that features from the high layers of the transformer focus on classification information and global features, while those from the middle layers pay more attention to pedestrians. We think that making full use of the features of different layers will facilitate alignment and then will promote reidentification accuracy. Therefore, we propose a novel skip connection aggregation transformer (SCAT) network by utilizing features from different transformer layers to increase the diversity of features and align visible body parts in occluded images. The diverse features include the following: first, features of the middle layer, which focus on the pedestrian in nonoccluded regions and favor alignment, second, features of high layers, which focus on global information, third fine-grained local features, which are obtained by the part pooling encoder and the fusion reconstruction module. The part pooling encoder and the fusion reconstruction module are proposed to obtain part-based local features and fused local features, respectively. The experimental results on the occluded, partial, and holistic benchmarks demonstrate that our method can significantly promote the accuracy of occluded person reidentification.
Huijie Fan, Qiang Wang 0015, Sheng-Peng Fu, Yandong Tang
IEEE Trans. Ind. Informatics5
2024 QueryTrack: Joint-Modality Query Fusion Network for RGBT Tracking
abstract
Existing RGB-Thermal trackers usually treat intra-modal feature extraction and inter-modal feature fusion as two separate processes, therefore the mutual promotion of extraction and fusion is neglected. Then, the complementary advantages of RGB-T fusion are not fully exploited, and the independent feature extraction is not adaptive to modal quality fluctuation during tracking. To address the limitations, we design a joint-modality query fusion network, in which the intra-modal feature extraction and the inter-modal fusion are coupled together and promote each other via joint-modality queries. The queries are initialized based on the multimodal features of the current frame, making the subsequent fusion adaptive to modal quality fluctuation during tracking. Then the joint-modality query fusion (JQF) utilizes the queries to interact with RGB-T features, allowing the intra-modal enhancement and the inter-modal interactions to be unified for mutual promotion. In this way, JQF can distinguish and enhance the complementary modality features, while filtering out redundant information. For real-time tracking, we propose regional cross-attention for cross-modal interactions to reduce computational cost. Our end-to-end tracker sets a new state-of-the-art performance on multiple RGBT tracking benchmarks including LasHeR, VTUAV, RGBT234 and GTOT, while running at a real-time speed.
Huijie Fan, Zhencheng Yu, Qiang Wang 0015, Baojie Fan, Yandong Tang
IEEE Trans. Image Process.5
2024 Online Video Sparse Noise Removing via Nonlocal Robust PCA
abstract
Online schemes and nonlocal similarity are two effective approaches for strengthening robust principal component analysis (RPCA) techniques in video denoising. However, their limitations are also evident. The online scheme is usually highly efficient but lacks consideration of regional appearance information, thus it cannot effectively handle videos with complex dynamics such as object movements. On the other hand, nonlocal similarity is used to better utilize regional information but incurs a heavy computational cost. Moreover, these two techniques are incompatible and challenging to work together. To overcome this barrier and harness the advantages of both approaches, this paper proposes a novel online nonlocal RPCA method. 1) A clustering based nonlocal strategy (ClusNonlocal) is adopted, which not only greatly reduces the computation cost, but also forms low-dimensional subspaces for online processing; 2) a new weighted RPCA model is proposed, which regards samples with different importances and improves the performance of subspace pursuit and video recovery; 3) a multi-level subspace updating scheme and weighted projection method is proposed, which keeps the performance of online video data processing at a high level at all time. A series of video denoising experiments are carried out to demonstrate the overall advantages of our procedure over several other ones, in terms of both visual quality and running speed.
Zhi Han, Huijie Fan, Yandong Tang, Yao Wang 0003
IEEE Trans. Multim.5
2024 A Shadow Imaging Bilinear Model and Three-Branch Residual Network for Shadow Removal
abstract
The current shadow removal pipeline relies on the detected shadow masks, which have limitations for penumbras and tiny shadows, and results in an excessively long pipeline. To address these issues, we propose a shadow imaging bilinear model and design a novel three-branch residual (TBR) network for shadow removal. Our bilinear model reveals the single-image shadow removal process and can explain why simply increasing the brightness of shadow areas cannot remove shadows without artifacts. We considerably shorten the shadow removal pipeline by modeling illumination compensation and developing a single-stage shadow removal network without additional detection and refinement networks. Specifically, our network consists of three task branches, i.e., shadow image reconstruction, shadow matte estimation, and shadow removal. To merge these three branches and enhance the shadow removal branch, we design a model-based TBR module. Multiple TBR modules are cascaded to generate an intensive information flow and facilitate feature integration among the three branches. Thus, our network ensures the fidelity of nonshadow areas and restores the light intensity of shadow areas through three-branch collaboration. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The model and code are available at https://github.com/nachifur/TBRNet.
Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Jiandong Tian, Yandong Tang
IEEE Trans. Neural Networks Learn. Syst.5
2023 Neural network equivalent model for highly efficient massive data classification
Siquan Yu, Zhi Han, Yandong Tang, Chengdong Wu 0001
Sci. China Inf. Sci.3
2023 Progressive feature-aware recurrent net for low-light image enhancement
Pengyue Li, Xiai Chen, Jiandong Tian, Yandong Tang
Signal Process. Image Commun.4
2023 Region Selective Fusion Network for Robust RGB-T Tracking
abstract
RGB-T tracking utilizes thermal infrared images as a complement to visible light images in order to perform more robust visual tracking in various scenarios. However, the highly aligned RGB-T image pairs introduces redundant information, the modal quality fluctuation during tracking also brings unreliable information. Existing RGB-T trackers usually use channelwise multi-modal feature fusion in which the low-quality features degrades the fused features and causes trackers to drift. In this work, we propose a region selective fusion network that first evaluates each image region by cross-modal and cross-region modeling, then removes low-quality redundant region features to alleviate the negative effects caused by unreliable information in multi-modal fusion. Besides, the region removal scheme brings a efficiency boost as redundant features are removed progressively, this enables the tracker to run at a high tracking speed.Extensive experiments show that the proposed tracker achieves competitive performance with a real-time tracking speed on multiple RGB-T tracking benchmarks including LasHeR, RGBT234 and GTOT.
Zhencheng Yu, Huijie Fan, Qiang Wang 0015, Ziwan Li, Yandong Tang
IEEE Signal Process. Lett.5
2023 A Decoupled Multi-Task Network for Shadow Removal
abstract
Shadow removal, which aims to restore the illumination in shadow regions, is challenging due to the diversity of shadows in terms of location, intensity, shape, and size. Different from most multi-task methods, which design elaborate multi-branch or multi-stage structures for better shadow removal, we introduce feature decomposition to learn better feature representations. Specifically, we propose a single-stage and decoupled multi-task network (DMTN) to explicitly learn the decomposed features for shadow removal, shadow matte estimation, and shadow image reconstruction. First, we propose several coarse-to-fine semi-convolution (SMC) modules to capture features sufficient for joint learning of these three tasks. Second, we design a theoretically supported feature decoupling layer to explicitly decouple the learned features into shadow image features and shadow matte features via weight reassignment. Last, these features are converted to a target shadow-free image, affiliated shadow matte, and shadow image, supervised by multi-task joint loss functions. With multi-task collaboration, DMTN effectively recovers the illumination in shadow areas while ensuring the fidelity of non-shadow areas. Experimental results show that DMTN competes favorably with state-of-the-art multi-branch/multi-stage shadow removal methods, while maintaining the simplicity of single-stage methods. We have released our code to encourage future exploration in powerful feature representation for shadow removalhttps://github.com/nachifur/DMTN
Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Liangqiong Qu, Yandong Tang
IEEE Trans. Multim.6
2022 SemanticGAN: Facial Image Editing with Semantic to Realize Consistency
Xin Luan, Huijie Fan, Yandong Tang
PRCV (3)4
2022 A novel compact design of convolutional layers with spatial transformation towards lower-rank representation for image classification
abstract
Convolutional neural networks (CNNs) usually come with numerous parameters and thus are not convenient for some situations, such as when the storage space is limited. Low-rank decomposition is one effective way for network compression or compaction. However, the current methods are far from theoretical optimal compression performance because the low-rankness of the commonly trained convolution filter sets is limited because of the versatility of convolution filters. We propose a novel compact design for convolutional layers with spatial transformation for achieving a much lower-rank form. The convolution filters in our design are generated using a predefined Tucker product form, followed by learnable individual spatial transformations on each filter. The low-rank (Tucker) part lowers the parameter capacity while the transformation part enhances the feature representation capacity. We validate our proposed approach on an image classification task. Our approach focuses on compressing parameters while also improving accuracy. We perform experiments on the MNIST, CIFAR10, CIFAR100, and ImageNet datasets. On the ImageNet dataset, our approach outperforms low-rank based state-of-the-arts by 2% to 6% in top-1 validation accuracy. Furthermore, our approach outperforms a series of low-rank-based state-of-the-arts on various datasets. The experiments validate the efficacy of our proposed method. Our code is available at https://github.com/liubc17/low_rank_compact_transformed.
Baichen Liu, Zhi Han, Xiai Chen, Wenming Shao, Huidi Jia, Yandong Tang
Knowl. Based Syst.7
2022 APAN: Across-Scale Progressive Attention Network for Single Image Deraining
abstract
Recent single image deraining works have achieved significant improvement using convolutional neural networks. However, the rain streaks in the rain image share similar patterns with its multi-scale versions, which are not fully exploited in recent works. In this paper, we propose anAcross-scaleProgressiveAttentionNetwork (i.e.,APAN) to explore the multi-scale collaborative representation for single image deraining. Specifically, we represent each rainy image via a multi-scale module. An across-scale attention module is then used to capture long-range feature correspondences from multi-scale features, which can model the rain streaks at an enlarging feature dimension. Afterwards, we construct a pyramid structure and further predict the rain streak progressively, which also guides the across-scale attention module to refine the feature representation from coarse to fine. The proposed model exploits self-similarity of features via an across-scale attention between different scales, which can well model the rain streak with long-range information. Experiments on several datasets show that our model achieves significant improvement compared with most state-of-the-art deraining models.
Qiang Wang 0015, Gan Sun, Huijie Fan, Yandong Tang
IEEE Signal Process. Lett.5
2022 Effective Tensor Completion via Element-Wise Weighted Low-Rank Tensor Train With Overlapping Ket Augmentation
abstract
Tensor completion methods based on the tensor train (TT) have the issues of inaccurate weight assignment and ineffective tensor augmentation pre-processing. In this work, we propose a novel tensor completion approach via the element-wise weighted technique. Accordingly, a novel formulation for tensor completion and an effective optimization algorithm, called tensor completion by parallel weighted matrix factorization via tensor train (TWMac-TT), is proposed. In addition, we specifically consider the recovery quality of edge elements from adjacent blocks. Different from traditional reshaping and ket augmentation, we utilize a new tensor augmentation technique called overlapping ket augmentation, which can further avoid blocking artifacts. We then conduct extensive performance evaluations on synthetic data and several real image data sets. Our experimental results demonstrate that the proposed algorithm TWMac-TT outperforms several other competing tensor completion methods. The code is available athttps://github.com/yzcv/TWMac-TT-OKA
Yang Zhang 0073, Yao Wang 0003, Zhi Han, Xiai Chen, Yandong Tang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Dual Aligned Siamese Dense Regression Tracker
abstract
Anchor or anchor-free based Siamese trackers have achieved the astonishing advancement. However, their parallel regression and classification branches lack the tracked target information link and interaction, and the corresponding independent optimization maybe lead to task-misalignment, such as the reliable classification prediction with imprecisely localization and vice versa. To address this problem, we develop a general Siamese dense regression tracker (SDRT) with both task and feature alignments. It consists of two cooperative and mutual-guidance core branches: dense local regression with RepPoint representation, the global and local multi-classifier fusion with aligned features. They complement and boost each other to constrain the results with well-localized followed to also be well-classified. Specifically, a dense local regression with RepPoint representation, directly estimates and averages multiple dense local bounding box offsets for accurate localization. And then, the refined bounding boxes can be used to learn the global and local affine alignment features for reliable multi-classifier fusion. The classified scores in turn guide the assigned positive bounding boxes for the regression task. The mutual guidance operations can bridge the connection between classification and regression substantially, since the assigned labels of one task depend on the prediction quality of the other task. The proposed tracking module is general, and it can boost both the anchor or anchor-free based Siamese trackers to some extent. The extensive tracking comparisons on six tracking benchmarks verify its favorable and competitive performance over states-of-the-arts tracking modules.
Baojie Fan, Hui Zhang 0023, Yang Cong, Yandong Tang, Huijie Fan, Jiandong Tian
IEEE Trans. Image Process.4
2022 Discriminative Siamese Complementary Tracker With Flexible Update
abstract
The offline generative Siamese trackers are equipped with the pre-defined anchors and the fixed target template. They overlook the target-background discriminative information, and lack the flexible target-specific update strategy. To overcome above drawbacks, we propose an adaptive and discriminative Siamese complementary tracking network with flexible update scheme. It consists of three collaborate subnetworks: anchor-free Siamese attention classification and regression subnetwork, online discriminative learning with multi-attention and multi-peak suppression, classifier guided template update subnetwork. All of them are interdependent and complementary to enhance each other for accurate target location. Specifically, an anchor-free multi-attention Siamese tracking subnetwork directly classifies the corresponding image patches with reliability assessment, and cascaded regresses the bounding boxes to progressively refine the predicting accuracy. Its evaluation is flexible and general with both proposal and anchor free in per-pixel prediction manner. Then, we integrate an online discriminative classifier optimizing module as a complementary subnetwork. It introduces spatial-temporal attention mechanism to fully explore multi-view multi-scale target-specific features, and evaluates multi-peak suppression to obtain a single centered peak response map. Its classified results can be fused with Siamese classification branch for accurate target location. Finally, the template update subnetwork is guided by the online discriminative classification scores. Extensive experiments on recent tracking datasets verify its top-ranked tracking accuracy and robustness against some state-of-the-art trackers.
Baojie Fan, Jiandong Tian, Yan Peng 0001, Yandong Tang
IEEE Trans. Multim.4
2021 Temporal pyramid attention-based spatiotemporal fusion model for Parkinson's disease diagnosis from gait data
abstract
Abstract Parkinson's disease (PD) is currently an ongoing challenge in daily clinical medicine. To reduce diagnosis time and arduousness and even assess PD levels, a temporal pyramid attention‐based spatiotemporal (PAST) fusion model for diagnosis of PD is produced by using gait data from ground reaction forces. This model is innovative in two aspects. First, by using the temporal pyramid attention module, multiscale temporal attention is obtained from raw sequences. Second, 1D convolutional neural network and bidirectional long short‐term memory layers are used together to learn spatial fusion features from multiple channels in the spatial domain to obtain multichannel, multiscale fusion features. Experiments are performed on the PhysioBank data set, and the results show that the proposed PAST model outperforms other state‐of‐the‐art methods on classification results. This model can assist in the diagnosis and treatment of PD by using gait data.
Xiaomin Pei, Huijie Fan, Yandong Tang
IET Signal Process.3
2021 Dynamic and reliable subtask tracker with general schatten p-norm regularization
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang
Pattern Recognit.4
2021 Structured and Consistent Multi-Layer Multi-Kernel Subtask Correction Filter Tracker
abstract
Some multi-task correlation filter trackers achieve the top-ranked performance in terms of accuracy and robustness. However, they directly fuse multiple types of features into a single kernel space. This operation fails to fully explore the discriminative strength and diversity of different features, and also ignores the structured correspondence of different tasks. To solve these issues, we propose a structured multi-kernel subtask correlation filter tracker with temporal-spatial consistency, which enjoys the merits of both layered multi-kernel subtask learning and structured correlation filter. Specifically, we firstly assign one kernel space to each channel feature. Multi-channel features correspond to multi-kernel spaces to boost their powerful discriminability. And then, we divide the target into multi-layer patches with different sizes, and regard the correlation filter trace of each patch with one channel feature as a subtask. In the following, we incorporate globally and locally structured correlation filters into a unified multi-kernel subtask particle tracking framework. The global and local subtasks complement and enhance each other with similar motion model. The proposed tracker not only exploits the cooperation and complementarity of layered multi-kernel subtask correlation filters, but also mines the underlying geometric structure of global subtasks, and the inner spatial locality correspondences of local subtasks inside the target. This operation is achieved by dual group sparsity regularized terms with mixed-norm lp,q, which decomposes the multi-kernel subtask filter matrix into two collaborative components. They correspond to the adaptive filter feature selection and outlier subtask detection, respectively. Besides, the developed tracking model maintains the temporal coherence and spatial consistency of multi-layer subtask filters via the smooth regularizer. Finally, the tracking formulation is optimized by the accelerated proximal gradient approach (APG). Encouraging analyses on six benchmark datasets, verify the favorable effectiveness and robustness of our method against state-of-the-art trackers.
Baojie Fan, Yang Cong, Yandong Tang, Jiandong Tian, Chenliang Xu
IEEE Trans. Circuits Syst. Video Technol.3
2021 Deep Retinex Network for Single Image Dehazing
abstract
In this paper, we propose a retinex-based decomposition model for a hazy image and a novel end-to-end image dehazing network. In the model, the illumination of the hazy image is decomposed into natural illumination for the haze-free image and residual illumination caused by haze. Based on this model, we design a deep retinex dehazing network (RDN) to jointly estimate the residual illumination map and the haze-free image. Our RDN consists of a multiscale residual dense network for estimating the residual illumination map and a U-Net with channel and spatial attention mechanisms for image dehazing. The multiscale residual dense network can simultaneously capture global contextual information from small-scale receptive fields and local detailed information from large-scale receptive fields to precisely estimate the residual illumination map caused by haze. In the dehazing U-Net, we apply the channel and spatial attention mechanisms in the skip connection of the U-Net to achieve a trade-off between overdehazing and underdehazing by automatically adjusting the channel-wise and pixel-wise attention weights. Compared with scattering model-based networks, fully data-driven networks, and prior-based dehazing methods, our RDN can avoid the errors associated with the simplified scattering model and provide better generalization ability with no dependence on prior information. Extensive experiments show the superiority of the RDN to various state-of-the-art methods.
Pengyue Li, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001
IEEE Trans. Image Process.3
2021 Tracking-by-Counting: Using Network Flows on Crowd Density Maps for Tracking Multiple Targets
abstract
State-of-the-art multi-object tracking (MOT) methods follow the tracking-by-detection paradigm, where object trajectories are obtained by associating per-frame outputs of object detectors. In crowded scenes, however, detectors often fail to obtain accurate detections due to heavy occlusions and high crowd density. In this paper, we propose a new MOT paradigm, tracking-by-counting, tailored for crowded scenes. Using crowd density maps, we jointly model detection, counting, and tracking of multiple targets as a network flow program, which simultaneously finds the global optimal detections and trajectories of multiple targets over the whole video. This is in contrast to prior MOT methods that either ignore the crowd density and thus are prone to errors in crowded scenes, or rely on a suboptimal two-step process using heuristic density-aware point-tracks for matching targets. Our approach yields promising results on public benchmarks of various domains including people tracking, cell tracking, and fish tracking.
Weihong Ren, Xinchao Wang, Jiandong Tian, Yandong Tang, Antoni B. Chan
IEEE Trans. Image Process.4
2021 Recurrent Generative Adversarial Network for Face Completion
abstract
Most recently-proposed face completion algorithms use high-level features extracted from convolutional neural networks (CNNs) to recover semantic texture content. Although the completed face is natural-looking, the synthesized content still lacks lots of high-frequency details, since the high-level features cannot supply sufficient spatial information for details recovery. To tackle this limitation, in this paper, we propose aRecurrentGenerativeAdversarialNetwork (RGAN) for face completion. Unlike previous algorithms, RGAN can take full advantage of multi-level features, and further provide advanced representations from multiple perspectives, which can well restore spatial information and details in face completion. Specifically, our RGAN model is composed of a CompletionNet and a DisctiminationNet, where the CompletionNet consists of two deep CNNs and a recurrent neural network (RNN). The first deep CNN is presented to learn the internal regulations of a masked image and represent it with multi-level features. The RNN model then exploits the relationships among the multi-level features and transfers these features in another domain, which can be used to complete the face image. Benefiting from bidirectional short links, another CNN is used to fuse multi-level features transferred from RNN and reconstruct the face image in different scales. Meanwhile, two context discrimination networks in the DisctiminationNet are adopted to ensure the completed image consistency globally and locally. Experimental results on benchmark datasets demonstrate qualitatively and quantitatively that our model performs better than the state-of-the-art face completion models, and simultaneously generates realistic image content and high-frequency details. The code will be released available soon.
Qiang Wang 0015, Huijie Fan, Gan Sun, Weihong Ren, Yandong Tang
IEEE Trans. Multim.5
2021 Robust 3-D Object Recognition via View-Specific Constraint
abstract
Three-dimensional (3-D) object recognition task focuses on detecting the objects of a scene and estimating their 6-DOF pose via effective feature extraction methods. Most recent feature extraction methods are based on the deep neural networks and show good performances. However, these methods require rendering engine to assist in generating a large amount of training data, which need much time to converge and further lead to the block in a rapid industrial production line. Besides, for the common hand-crafted features, the lack of discriminant feature-points amongst various texture-less and surface-smooth objects can cause ambiguity in the process of feature-points matching. To address these challenges above, a hand-crafted 3-D feature descriptor with center offset and pose annotations is proposed in this article, which is called view-specific local projection statistics (VSLPSs). By relying on these annotations as seeds, a voting strategy is then used to transform the feature-points matching problem into the problem of voting an optimal model-view in the 6-DOF space. In this way, the ambiguity of feature-points matching caused by poor feature discrimination is eliminated. To the end, various experiments on three public datasets and our built 3-D bin-picking dataset demonstrate that our proposed VSLPS method performs well in comparison with the state-of-the-art.
Hongsen Liu, Yang Cong, Gan Sun, Yandong Tang
IEEE Trans. Syst. Man Cybern. Syst.4
2021 Low-rank decomposition on transformed feature maps domain for image denoising
Qiong Luo 0003, Baichen Liu, Yang Zhang 0073, Zhi Han, Yandong Tang
Vis. Comput.5
2020 Dually Connected Deraining Net Using Pixel-Wise Attention
abstract
Recent single image deraining methods either use a recurrent mechanism to gradually learn the mapping between clear images and rainy images, or focus on designing various loss functions to supervise the learning process. In this letter, we propose a dually connected deraining net using pixel-wise attention, for single image rain removal. Specifically, the deraining net adopts an encoder-decoder net as a backbone, which can effectively learn a residual rain-streaks map by jointly using skip sum connection and skip concatenation connection. The dual connections enable the deraining net to promote information flow between layers, and thus can allow it to discriminate and localize the rain streaks. To preserve image details, the decoded features are weighted by the learnable pixel-wise attention for adaptively recalibrating their responses. Experimental results on synthetic datasets demonstrate that the proposed model outperforms the recent state-of-the-art deraining methods.
Weihong Ren, Jiandong Tian, Qiang Wang 0015, Yandong Tang
IEEE Signal Process. Lett.4
2020 Reliable Multi-Kernel Subtask Graph Correlation Tracker
abstract
Many astonishing correlation filter trackers pay limited concentration on the tracking reliability and locating accuracy. To solve the issues, we propose a reliable and accurate cross correlation particle filter tracker via graph regularized multi-kernel multi-subtask learning. Specifically, multiple non-linear kernels are assigned to multi-channel features with reliable feature selection. Each kernel space corresponds to one type of reliable and discriminative features. Then, we define the trace of each target subregion with one feature as a single view, and their multi-view cooperations and interdependencies are exploited to jointly learn multi-kernel subtask cross correlation particle filters, and make them complement and boost each other. The learned filters consist of two complementary parts: weighted combination of base kernels and reliable integration of base filters. The former is associated to feature reliability with importance map, and the weighted information reflects different tracking contribution to accurate location. The second part is to find the reliable target subtasks via the response map, to exclude the distractive subtasks or backgrounds. Besides, the proposed tracker constructs the Laplacian graph regularization via cross similarity of different subtasks, which not only exploits the intrinsic structure among subtasks, and preserves their spatial layout structure, but also maintains the temporal-spatial consistency of subtasks. Comprehensive experiments on five datasets demonstrate its remarkable and competitive performance against state-of-the-art methods.
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.4
2019 Stacked dense networks for single-image snow removal
Pengyue Li, Mengshen Yun, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001
Neurocomputing4
2019 Efficient 3D object recognition via geometric information preservation
Hongsen Liu, Yang Cong, Chenguang Yang 0001, Yandong Tang
Pattern Recognit.4
2019 Laplacian pyramid adversarial network for face completion
Qiang Wang 0015, Huijie Fan, Gan Sun, Yang Cong, Yandong Tang
Pattern Recognit.5
2019 Deeply Supervised Face Completion With Multi-Context Generative Adversarial Network
abstract
Recent face completion works have achieved significant improvement using generative adversarial networks (GANs). There are still two important issues in this challenging task: first, semantic understanding; and second, high-frequency details prediction. In this letter, we propose a unified model by introducing multi-context structures within GANs. Our model, named multi-context generative adversarial networks (MCGAN), automatically learns the hierarchical appearances of a corrupted image and predicted the missing regions from different perspectives. In this model, semantic understanding and high-frequency details are both taken into account and modeled with two parallel networks, respectively. While one learns the semantic understanding of the input face image at a high level, the other extracts low-level features for high-frequency details prediction. Our MCGAN takes full advantage of multi-scale features learned from two complementary networks and generates semantically new pixels for the missing region with fine details. Extensive quantitative and qualitative experiments on benchmark datasets show that the proposed model outperforms several state-of-the-art models.
Qiang Wang 0015, Huijie Fan, Yandong Tang
IEEE Signal Process. Lett.4
2018 Fusing Crowd Density Maps and Visual Object Trackers for People Tracking in Crowd Scenes
abstract
While visual tracking has been greatly improved over the recent years, crowd scenes remain particularly challenging for people tracking due to heavy occlusions, high crowd density, and significant appearance variation. To address these challenges, we first design a Sparse Kernelized Correlation Filter (S-KCF) to suppress target response variations caused by occlusions and illumination changes, and spurious responses due to similar distractor objects. We then propose a people tracking framework that fuses the S-KCF response map with an estimated crowd density map using a convolutional neural network (CNN), yielding a refined response map. To train the fusion CNN, we propose a two-stage strategy to gradually optimize the parameters. The first stage is to train a preliminary model in batch mode with image patches selected around the targets, and the second stage is to fine-tune the preliminary model using the real frame-by-frame tracking process. Our density fusion framework can significantly improves people tracking in crowd scenes, and can also be combined with other trackers to improve the tracking performance. We validate our framework on two crowd video datasets.
Weihong Ren, Yandong Tang, Antoni B. Chan
CVPR3
2018 Robust video denoising with sparse and dense noise modelings
Guiping Shen, Zhi Han, Xiai Chen, Yandong Tang
Sci. China Inf. Sci.4
2018 Evaluation of shadow features
abstract
Shadow features such as colour ratio, texture, and chromaticity have proved to be quite effective in shadow detection. Many shadow detection methods have been proposed on the basis of different features. However, previous works for shadow detection mainly focus on designing an effective classifier for existing shadow features, but pay less attention on the analysis of shadow features themselves. The majority of studies simply report the final shadow detection results rather than make an evaluation on each feature. Readers often do not know which features are more effective or whether these shadow features are complementary. The following problems are still unsolved: the robustness of each feature, which feature plays the most important role in a detection method, and what is the best performance that current features can reach. The purpose of this study is to answer these questions, and the authors hope that this study can offer guidance for future shadow detection algorithms via the evaluation of frequently used shadow features. Several useful and interesting conclusions are obtained after conducting extensive comparison experiments on a large dataset.
Liangqiong Qu, Jiandong Tian, Huijie Fan, Yandong Tang
IET Comput. Vis.5
2018 Methods and datasets on semantic segmentation: A review
Hongshan Yu, Zhengeng Yang, Yaonan Wang 0001, Wei Sun 0028, Mingui Sun, Yandong Tang
Neurocomputing7
2018 Structured and weighted multi-task low rank tracker
Baojie Fan, Xiaomao Li, Yang Cong, Yandong Tang
Pattern Recognit.4
2018 Dual-Graph Regularized Discriminative Multitask Tracker
abstract
Multitask and low-rank learning methods have attracted increasing attention for visual tracking. However, most trackers only focus on learning appearance subspace basis or the sparse low rankness of representation and, thus, do not make full use of the structure information among and inside target candidates (or samples). In this paper, we propose a dual-graph regularized discriminative low-rank learning for a multitask tracker, which integrates the discriminative subspace and intrinsic geometric structures among tasks. By constructing dual-graph regulations from two views of multitask observation, the developed model not only exploits the intrinsic relationship among tasks, and preserves the spatial layout structure among the local patches inside each candidate, but also learns the salient features of the target samples. This operation has the benefit of having good target representation and improving the performance of the tracker. Moreover, our developed tracker is a collaborate multitask tracking model and learns the discriminative subspace with adaptive dimension and optimal classifier simultaneously. Then, a collaborate metric is developed to find the best candidate, which integrates both classification reliability and representation accuracy. Encouraging experimental results on a large set of public video sequences justify that our tracker performs favorably against many other state-of-the-art trackers.
Baojie Fan, Yang Cong, Yandong Tang
IEEE Trans. Multim.3
2018 Snowflake Removal for Videos via Global and Local Low-Rank Decomposition
abstract
Falling snow not only blocks human vision, but also significantly degrades the effectiveness of computer vision systems in outdoor environment. In this paper, we aim to remove snowflakes in videos by using the global and local low-rank property of snowflake-removed scenes. The stationary background and the mixture of moving foreground as well as falling snowflake are extracted via the global low-rank matrix decomposition. Some snowflake features, such as its color and size, are used to separate out the snowflakes from other moving objects. Then, the mean absolute difference based patch matching is applied to align every same moving object over frames to grab its low-rank structure. As such, the falling snowflake in front of moving objects can be removed via the local low-rank decomposition. Finally, the snowflake removed videos are generated by pasting moving foreground to stationary backgrounds. Experiments show that our method can remove snowflakes effectively and outperforms the comparison methods.
Jiandong Tian, Zhi Han, Weihong Ren, Xiai Chen, Yandong Tang
IEEE Trans. Multim.5
2018 A Generalized Model for Robust Tensor Factorization With Noise Modeling by Mixture of Gaussians
abstract
The low-rank tensor factorization (LRTF) technique has received increasing attention in many computer vision applications. Compared with the traditional matrix factorization technique, it can better preserve the intrinsic structure information and thus has a better low-dimensional subspace recovery performance. Basically, the desired low-rank tensor is recovered by minimizing the least square loss between the input data and its factorized representation. Since the least square loss is most optimal when the noise follows a Gaussian distribution, -norm-based methods are designed to deal with outliers. Unfortunately, they may lose their effectiveness when dealing with real data, which are often contaminated by complex noise. In this paper, we consider integrating the noise modeling technique into a generalized weighted LRTF (GWLRTF) procedure. This procedure treats the original issue as an LRTF problem and models the noise using a mixture of Gaussians (MoG), a procedure called MoG GWLRTF. To extend the applicability of the model, two typical tensor factorization operations, i.e., CANDECOMP/PARAFAC factorization and Tucker factorization, are incorporated into the LRTF procedure. Its parameters are updated under the expectation-maximization framework. Extensive experiments indicate the respective advantages of these two versions of MoG GWLRTF in various applications and also demonstrate their effectiveness compared with other competing methods.
Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Lin Lin 0007, Yandong Tang
IEEE Trans. Neural Networks Learn. Syst.7
2017 DeshadowNet: A Multi-context Embedding Deep Network for Shadow Removal
abstract
Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a multi-context architecture, where the output shadow matte is predicted by embedding information from three different perspectives. The first global network extracts shadow features from a global view. Two levels of features are derived from the global network and transferred to two parallel networks. While one extracts the appearance of the input image, the other one involves semantic understanding for final prediction. These two complementary networks generate multi-context features to obtain the shadow matte with fine local details. To evaluate the performance of the proposed method, we construct the first large scale benchmark with 3088 image pairs. Extensive experiments on two publicly available benchmarks and our large-scale benchmark show that the proposed method performs favorably against several state-of-the-art methods.
Liangqiong Qu, Jiandong Tian, Shengfeng He, Yandong Tang, Rynson W. H. Lau
CVPR4
2017 Video Desnowing and Deraining Based on Matrix Decomposition
abstract
The existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snowflakes/rain streaks. In this paper, we propose a model based on matrix decomposition for video desnowing and deraining to solve the problems mentioned above. We divide snowflakes/rain streaks into two categories: sparse ones and dense ones. With background fluctuations and optical flow information, the detection of moving objects and sparse snowflakes/rain streaks is formulated as a multi-label Markov Random Fields (MRFs). As for dense snowflakes/rain streaks, they are considered to obey Gaussian distribution. The snowflakes/rain streaks, including sparse ones and dense ones, in scene backgrounds are removed by low-rank representation of the backgrounds. Meanwhile, a group sparsity term in our model is designed to filter snow/rain pixels within the moving objects. Experimental results show that our proposed model performs better than the state-of-the-art methods for snow and rain removal.
Weihong Ren, Jiandong Tian, Zhi Han, Antoni B. Chan, Yandong Tang
CVPR5
2017 Tensor RPCA by Bayesian CP Factorization with Complex Noise
abstract
The RPCA model has achieved good performances in various applications. However, two defects limit its effectiveness. Firstly, it is designed for dealing with data in matrix form, which fails to exploit the structure information of higher order tensor data in some pratical situations. Secondly, it adopts L1-norm to tackle noise part which makes it only valid for sparse noise. In this paper, we propose a tensor RPCA model based on CP decomposition and model data noise by Mixture of Gaussians (MoG). The use of tensor structure to raw data allows us to make full use of the inherent structure priors, and MoG is a general approximator to any blends of consecutive distributions, which makes our approach capable of regaining the low dimensional linear subspace from a wide range of noises or their mixture. The model is solved by a new proposed algorithm inferred under a variational Bayesian framework. The superiority of our approach over the existing state-of-the-art approaches is demonstrated by extensive experiments on both of synthetic and real data.
Qiong Luo 0003, Zhi Han, Xiai Chen, Yao Wang 0003, Deyu Meng, Yandong Tang
ICCV7
2017 Large receptive field convolutional neural network for image super-resolution
abstract
This paper presents a new approach to Single Image Super Resolution (SISR), based upon Convolutional Neural Network (CNN). Although the SISR is ill-posed which can be seen as finding a non-linear mapping from a low to high-dimensional space. Deep learning techniques have been successfully applied in many areas of computer vision, including low-level image restoration and non-linear mapping problems. We consider the single image Super-Resolution (SR) problem as convolution operators and develop a CNN to capture the characteristics of Low-Resolution (LR) input image. We find that increasing the receptive field shows the improvement in accuracy. Our solution is to establish the connection between traditional optimization-based schemes and neural network architectures. In the paper a novel, separable structure is introduced as a reliable support for robust convolution against artifacts. Our proposed method performs better than existing methods in terms of accuracy and visual improvements in our results are easily noticeable.
Qiang Wang 0015, Huijie Fan, Yang Cong, Yandong Tang
ICIP4
2017 Deep learning of directional truncated signed distance function for robust 3D object recognition
abstract
In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression tasks. To achieve this, we first adopt the Truncated Signed Distance Function (TSDF) to encode the point cloud and extract low compact discriminative feature via unsupervised deep learning network. This approach can not only eliminate the dense scale sampling for offline model training but also reduce the distortion by mapping the 3D shape to the 2D plane and overcome the dependence on color cues. Then, we train a Hough forests to achieve multi-object detection and 6-DoF pose estimation simultaneously. In addition, we propose a robust multilevel verification strategy that effectively reduces the unpredictability of output values which occurs in the hough regression module. Experiments on public datasets demonstrate that our approach provides effective results comparable to the state-of-the-arts.
Hongsen Liu, Yang Cong, Shuai Wang 0003, Huijie Fan, Dongying Tian, Yandong Tang
IROS6
2017 Single image dehazing by latent region-segmentation based transmission estimation and weighted L 1-norm regularisation
abstract
Image dehazing is a useful technique which can eliminate the bad effect of haze on images and enhance the performances of image/video processing algorithms in the hazy weather. In this study, a single image dehazing method is proposed. The authors estimate the initial transmission properly based on latent region‐segmentation and refine the estimated initial transmission by an objective function with a novel weighted L 1 ‐norm regularisation term. The half‐quadratic splitting minimisation method is employed to solve this optimisation problem. They also define an evaluation function to estimate the reliable global atmospheric light. With the refined transmission map and atmospheric light they recover the haze‐free image by the haze imaging model. The authors’ method is compared with three state‐of‐the‐art methods and is also validated by two image quality assessment methods. The comparative experimental results and evaluations demonstrate that their method can recover comparable and even better results with clear details, low contrast loss and high contrast in most cases.
Tong Cui, Jiandong Tian, Ende Wang, Yandong Tang
IET Image Process.4
2017 A New Intrinsic-Lighting Color Space for Daytime Outdoor Images
abstract
Extracting or separating intrinsic information and illumination from natural images is crucial for better solving computer vision tasks. In this paper, we present a new illumination-based color space, the IL (intrinsic information and lighting level) space. Its first two channels represent 2D intrinsic information, and the third channel is for lighting levels. The IL color space has a one-to-one correspondence with the RGB color space. One valuable benefit of the IL color space is that illumination-related processing can be realized by directly operating on the lighting channel. As an example, based on the extracted lighting channel, we propose a new algorithm to estimate the intrinsic lighting level of an image such that the shadow-free color image and relighting series are obtained. In contrast to the existing color spaces for display or printing, the IL color space intuitively shows the information of reflectance and lighting levels for colors separately.
Zhi Han, Jiandong Tian, Liangqiong Qu, Yandong Tang
IEEE Trans. Image Process.4
2017 RGBD Salient Object Detection via Deep Fusion
abstract
Numerous efforts have been made to design various low-level saliency cues for RGBD saliency detection, such as color and depth contrast features as well as background and color compactness priors. However, how these low-level saliency cues interact with each other and how they can be effectively incorporated to generate a master saliency map remain challenging problems. In this paper, we design a new convolutional neural network (CNN) to automatically learn the interaction mechanism for RGBD salient object detection. In contrast to existing works, in which raw image pixels are fed directly to the CNN, the proposed method takes advantage of the knowledge obtained in traditional saliency detection by adopting various flexible and interpretable saliency feature vectors as inputs. This guides the CNN to learn a combination of existing features to predict saliency more effectively, which presents a less complex problem than operating on the pixels directly. We then integrate a superpixel-based Laplacian propagation framework with the trained CNN to extract a spatially consistent saliency map by exploiting the intrinsic structure of the input image. Extensive quantitative and qualitative experimental evaluations on three data sets demonstrate that the proposed method consistently outperforms the state-of-the-art methods.
Liangqiong Qu, Shengfeng He, Jiawei Zhang 0002, Jiandong Tian, Yandong Tang, Qingxiong Yang
IEEE Trans. Image Process.5
2017 Specular Reflection Separation With Color-Lines Constraint
abstract
According to dichromatic reflection model, the previous methods of specular reflection separation in image processing often separate specular reflection from a single image using patch-based priors. Due to lack of global information, these methods often cannot completely separate the specular component of an image and are incline to degrade image textures. In this paper, we derive a global color-lines constraint from dichromatic reflection model to effectively recover specular and diffuse reflection. Our key observation is from that each image pixel lies along a color line in normalized RGB space and the different color lines representing distinct diffuse chromaticities intersect at one point, namely, the illumination chromaticity. For pixels along the same color line, they spread over the entire image and their distances to the illumination chromaticity reflect the amount of specular reflection components. With global (non-local) information from these color lines, our method can effectively separate specular and diffuse reflection components in a pixelwise way for a single image, and it is suitable for real-time applications. Our experimental results on synthetic and real images show that our method performs better than the state-of-the-art methods to separate specular reflection.
Weihong Ren, Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.3
2017 Multi-Class Latent Concept Pooling for Computer-Aided Endoscopy Diagnosis
abstract
Successful computer-aided diagnosis systems typically rely on training datasets containing sufficient and richly annotated images. However, detailed image annotation is often time consuming and subjective, especially for medical images, which becomes the bottleneck for the collection of large datasets and then building computer-aided diagnosis systems. In this article, we design a novel computer-aided endoscopy diagnosis system to deal with the multi-classification problem of electronic endoscopy medical records (EEMRs) containing sets of frames, while labels of EEMRs can be mined from the corresponding text records using an automatic text-matching strategy without human special labeling. With unambiguous EEMR labels and ambiguous frame labels, we propose a simple but effective pooling scheme called Multi-class Latent Concept Pooling, which learns a codebook from EEMRs with different classes step by step and encodes EEMRs based on a soft weighting strategy. In our method, a computer-aided diagnosis system can be extended to new unseen classes with ease and applied to the standard single-instance classification problem even though detailed annotated images are unavailable. In order to validate our system, we collect 1,889 EEMRs with more than 59K frames and successfully mine labels for 348 of them. The experimental results show that our proposed system significantly outperforms the state-of-the-art methods. Moreover, we apply the learned latent concept codebook to detect the abnormalities in endoscopy images and compare it with a supervised learning classifier, and the evaluation shows that our codebook learning method can effectively extract the true prototypes related to different classes from the ambiguous data.
Shuai Wang 0003, Yang Cong, Huijie Fan, Baojie Fan, Lianqing Liu, Yunsheng Yang, Yandong Tang, Huaici Zhao
ACM Trans. Multim. Comput. Commun. Appl.7
2016 Robust Tensor Factorization with Unknown Noise
abstract
Because of the limitations of matrix factorization, such as losing spatial structure information, the concept of tensor factorization has been applied for the recovery of a low dimensional subspace from high dimensional visual data. Generally, the recovery is achieved by minimizing the loss function between the observed data and the factorization representation. Under different assumptions of the noise distribution, the loss functions are in various forms, like L1 and L2 norms. However, real data are often corrupted by noise with an unknown distribution. Then any specific form of loss function for one specific kind of noise often fails to tackle such real data with unknown noise. In this paper, we propose a tensor factorization algorithm to model the noise as a Mixture of Gaussians (MoG). As MoG has the ability of universally approximating any hybrids of continuous distributions, our algorithm can effectively recover the low dimensional subspace from various forms of noisy observations. The parameters of MoG are estimated under the EM framework and through a new developed algorithm of weighted low-rank tensor factorization (WLRTF). The effectiveness of our algorithm are substantiated by extensive experiments on both of synthetic data and real image data.
Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Yandong Tang
CVPR6
2016 Structured low rank tracker with smoothed regularization
abstract
In this paper, we propose a structured low rank learning algorithm with smoothed regularization for robust object tracking, under particle filter framework. Specifically, the relationships among the particles are exploited with structured low rank regularization term, and simultaneously handle the outlier using a group sparsity regularization. The label information from training data is incorporated into the tracking objective function as the classification error term and idea coding regularization term respectively. By the smoothed regularization, the developed structured low rank learning based tracker can be efficiently solved by iterative reweighed least squares algorithm(IRLS), and avoids svd operation. Moreover, the collaborate normalized metric is developed to find the best candidate. Compared with some state-of-the-art tracking methods on 50 challenging sequences, the proposed algorithms perform well in terms of accuracy, robustness.
Baojie Fan, Yang Cong, Xiaomao Li, Yandong Tang
VCIP4
2016 Scalable gastroscopic video summarization via similar-inhibition dictionary selection
Shuai Wang 0003, Yang Cong, Jun Cao 0002, Yunsheng Yang, Yandong Tang, Huaici Zhao
Artif. Intell. Medicine5
2016 Nonconvex plus quadratic penalized low-rank and sparse decomposition for noisy image alignment
Xiai Chen, Zhi Han, Yao Wang 0003, Yandong Tang
Sci. China Inf. Sci.4
2016 New spectrum ratio properties and features for shadow detection
Jiandong Tian, Xiaojun Qi 0001, Liangqiong Qu, Yandong Tang
Pattern Recognit.4
2016 Blind Deconvolution With Nonlocal Similarity and l0 Sparsity for Noisy Image
abstract
The blind image deconvolution techniques with sparsity prior in gradient domain are sensitive to noise, even a small amount of noise. To address this problem, in this letter, we propose a novel blind deconvolution model that combines low-rank property, nonlocal similarity, and l0sparsity prior. Low-rank property makes the proposed deblurring model robust to image noise. The joint utilization of nonlocal similarity and l0sparsity prior has improved the accuracy of blur kernel estimation and restores the fine image details. A numerical method is also given to solve the proposed problem. Experimental results on synthetic and real data show that our algorithm performs better against with the state-of-the-art methods for both noise and noise-free images.
Weihong Ren, Jiandong Tian, Yandong Tang
IEEE Signal Process. Lett.3
2015 Computer aided endoscope diagnosis via weakly labeled data mining
abstract
In comparison to most computer aided endoscope diagnosis methods using pixel-wise groundtruth by physicians manually, it is easy to get lots of endoscope images with corresponding diagnostic reports. In this paper, we intend to mine pixel-wise label information from these reports with weak frame-level labels automatically. To achieve this, we formulate our computer aided diagnosis problem as a Multiple Instance Learning (MIL) issue, where we represent each image as superpixels. Each image and each superpixel is cast as bag and instance, respectively. We then evaluate and select the most positive instances from positive bags automatically which helps us transform the frame-level classification problem into a standard supervised learning problem. In the experiment, we build a new gastroscopic image dataset with more than 3000 weakly labeled images, and ours outperforms the state-of-the-art methods, which verifies the effectiveness of our model.
Shuai Wang 0003, Yang Cong, Huijie Fan, Yunsheng Yang, Yandong Tang, Huaici Zhao
ICIP5
2015 Real-time one-dimensional motion estimation and its application in computer vision
Yang Cong, Haifeng Gong, Yandong Tang, Shuzhi Sam Ge, Jiebo Luo 0001
Mach. Vis. Appl.3
2015 Object detection based on scale-invariant partial shape matching
Huijie Fan, Yang Cong, Yandong Tang
Mach. Vis. Appl.3
2013 A new optimal seam finding method based on tensor analysis for automatic panorama construction
Yandong Tang
Pattern Recognit. Lett.3
2013 Video Anomaly Search in Crowded Scenes via Spatio-Temporal Motion Context
abstract
Video anomaly detection plays a critical role for intelligent video surveillance. We present an abnormal video event detection system that considers both spatial and temporal contexts. To characterize the video, we first perform the spatio-temporal video segmentation and then propose a new region-based descriptor called “Motion Context,” to describe both motion and appearance information of the spatio-temporal segment. For anomaly measurements, we formulate the abnormal event detection as a matching problem, which is more robust than statistic model-based methods, especially when the training dataset is of limited size. For each testing spatio-temporal segment, we search for its best match in the training dataset, and determine how normal it is using a dynamic threshold. To speed up the search process, compact random projections are also adopted. Experiments on the benchmark dataset and comparisons with the state-of-the-art methods validate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Yandong Tang
IEEE Trans. Inf. Forensics Secur.3
2012 Object tracking via online metric learning
abstract
By considering visual tracking as a similarity matching problem, we propose a self-supervised tracking method that incorporates adaptive metric learning and semi-supervised learning into the framework of object tracking. For object representation, the spatial-pyramid structure is applied by fusing both the shape and texture cues as descriptors. A metric learner is adaptively trained online to best distinguish the foreground object and background, and a new bi-linear graph is defined accordingly to propagate the label of each sample. Then high-confident samples are collected to self-update the model to handle large-scale issue. Experiments on the benchmark dataset and comparisons with the state-of-the-art methods validate the advantages of our algorithm.
Yang Cong, Junsong Yuan 0001, Yandong Tang
ICIP3
2012 Self-closed partial shape descriptor for shape retrieval
abstract
We propose a discriminative partial-based algorithm for shape recognition and retrieval. A key distinction of our approach is that we use pairwise geometric relations between contour fragments containing important and salient shape information to establish self-closed partial descriptor (SCPD), it can capture similar local parts in matching shape contours and meanwhile overcome part occlusion and distortion. We establish local coordinate system for each fragment to make sure SCPD is invariant to RST (rotation, scaling, and translation) transformation. In the matching stage, a scale approximation scheme is used to get rid of invalid matches. We experiment on MPEG7 shape database, and experimental results illustrate that our algorithm performs well on shape retrieval.
Huijie Fan, Yang Cong, Yandong Tang
ICIP3
2012 Active drift correction template tracking algorithm
abstract
This paper presents a novel active drift correction template tracking algorithm. Compared to Matthews' algorithm in [8], the proposed algorithm achieves synchronously object tracking and drift correction, and save half running time. For the template drift problem during long sequential object tracking, we introduce the active drift correction term into inverse compositional affine image alignment algorithm. This operation can avoid the template drift before it occurs, or reduce the drift after it happens. The total energy function consists of two terms: the tracking term and the active drift correction term. By minimizing the total energy function with the steepest descent algorithm, the proposed algorithm can decrease the accumulative tracking error, and prevent the drift during the tracking process effectively. Various object tracking experiments show that our method has super performance than the passive drift correction algorithm in [8].
Baojie Fan, Yingkui Du, Yang Cong, Yandong Tang
ICIP4
2011 Linearity of each channel pixel values from a surface in and out of shadows and its applications
abstract
Shadows, the common phenomena in most outdoor scenes, are illuminated by diffuse skylight whereas shaded from direct sunlight. Generally shadows take place in sunny weather when the spectral power distributions (SPD) of sunlight, skylight, and daylight show strong regularity: they principally vary with sun angles. In this paper, we first deduce that the pixel values of a surface illuminated by skylight (in shadow region) and by daylight (in non-shadow region) have a linear relationship, and the linearity is independent of surface reflectance and holds in each color channel. We then use six simulated images that contain 1995 surfaces and two real captured images to test the linearity. The results validate the linearity. Based on the deduced linear relationship, we develop three shadow processing applications include intrinsic image deriving, shadow verification, and shadow removal. The results of the applications demonstrate that the linear relationship have practical values.
Jiandong Tian, Yandong Tang
CVPR2
2011 A robust template tracking algorithm with weighted active drift correction
Baojie Fan, Yingkui Du, Yandong Tang
Pattern Recognit. Lett.5
2010 Skew detection in document images based on rectangular active contour
Huijie Fan, Yandong Tang
Int. J. Document Anal. Recognit.3
2009 Flow mosaicking: Real-time pedestrian counting without scene-specific learning
abstract
In this paper, we present a novel algorithm based on flow velocity field estimation to count the number of pedestrians across a detection line or inside a specified region. We regard pedestrians across the line as fluid flow, and design a novel model to estimate the flow velocity field. By integrating over time, the dynamic mosaics are constructed to count the number of pixels and edges passed through the line. Consequentially, the number of pedestrians can be estimated by quadratic regression, with the number of weighted pixels and edges as input. The regressors are learned off line from several camera tilt angles, and have taken the calibration information into account.We use tilt-angle-specific learning to ensure direct deployment and avoid overfitting while the commonly used scene-specific learning scheme needs on-site annotation and always trends to overfitting. Experiments on a variety of videos verified that the proposed method can give accurate estimation under different camera setup in real-time.
Yang Cong, Haifeng Gong, Song-Chun Zhu, Yandong Tang
CVPR4
2009 Tricolor Attenuation Model for Shadow Detection
abstract
Shadows, the common phenomena in most outdoor scenes, bring many problems in image processing and computer vision. In this paper, we present a novel method focusing on extracting shadows from a single outdoor image. The proposed tricolor attenuation model (TAM) that describe the attenuation relationship between shadow and its nonshadow background is derived based on image formation theory. The parameters of the TAM are fixed by using the spectral power distribution (SPD) of daylight and skylight, which are estimated according to Planck's blackbody irradiance law. Based on the TAM, a multistep shadow detection algorithm is proposed to extract shadows. Compared with previous methods, the algorithm can be applied to process single images gotten in real complex scenes without prior knowledge. The experimental results validate the performance of the model.
Jiandong Tian, Yandong Tang
IEEE Trans. Image Process.3
2008 Lunar terrain reconstruction using PDEs
abstract
Based on the geometry features of lunar terrain, this paper treats lunar terrain reconstruction as a surface reconstruction problem. We define an energy functional model consisting of local energy term and smooth energy term for lunar terrain reconstruction. The solution to minimize the functional (by partial differential equations) is defined as the optimal surface. In the smooth energy term, we design a vector field of depth discontinuousness likelihood (VFDDL) to control the direction and degree of smoothing. Experiments indicate that accurate VFDDL can lead to an exact reconstructed surface. Thus, VFDDL transfers 3D terrain reconstruction into a 2D image processing problem. An innovative method is proposed to estimate VFDDL, using image local and statistical features. Experiments verify our method and show a good performance in terrain reconstruction.
Ji Liu 0002, Yang Cong, Xiaomao Li, Yuechao Wang, Yandong Tang, Chuan Zhou 0010
ICIP5