EDBT 2026 Demo / reviewers in the wild / expert
Tao Lu 0001
dblp:03/5189-1
· DBLP profile ↗
109ranked-venue papers
18as first author
66since 2021 · last 2026
0000-0001-8117-2012ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 11 first-author · 28 since 2021Artificial intelligence and machine learning · 35 · 3 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 11 since 2021Systems, architecture and hardware · 9 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Few-shot 3D point cloud segmentation via dynamic multi-scale sparse attention with adaptive gated context enhancement
Yilin Chen 0001, Tao Lu 0001, Yiqi Wu, Hui Li 0128, Lu Zou |
Expert Syst. Appl. | 3 |
| 2026 | Marker and dynamic geometry aware transformer for robust point cloud registration
Yilin Chen 0001, Qinjie Zheng, Tao Lu 0001, Lu Zou, Xiantao Cai, Xiangyun Liao |
Expert Syst. Appl. | 3 |
| 2026 | Task-aware dynamic routing network for cross-domain few-shot learning
Yanan Li 0006, Haoyang Ye, Huabing Zhou, Tao Lu 0001, Hao Lu 0003 |
Neurocomputing | 4 |
| 2026 | Adaptive spatial feature extraction and graphical feature awareness for robust point cloud registration
Yilin Chen 0001, Yang Mei, Tao Lu 0001, Lu Zou, Xiangyun Liao, Fazhi He |
Neural Networks | 3 |
| 2026 | From forgotten to pan-sharpening
Jiaming Wang 0001, Yansong Lin, Chuanxi Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001 |
Pattern Recognit. | 7 |
| 2026 | NSBRNet: Non-Local Spatio-Temporal Bidirectional Recurrent Network for Satellite Video Super-ResolutionabstractIn recent years, intelligent processing of satellite videos has emerged as a significant research focus within the field of remote sensing, driven by the growing demand for enhanced spatial resolution. This need has led to increased interest in satellite video super-resolution (SVSR) algorithms, which aim to improve the quality of satellite imagery. However, many existing SVSR methods tend to neglect the global dependencies among frames in satellite videos, resulting in an incomplete utilization of spatio-temporal feature information. To tackle this issue, we propose a novel non-local spatio-temporal bidirectional recurrent network specifically designed for SVSR applications. Our approach employs a gate-guided deformable alignment module that effectively enhances feature alignment and fusion using a dynamic gating mechanism. This allows the network to adaptively focus on relevant features during the reconstruction process. Furthermore, we introduce a non-local spatio-temporal fusion module that integrates both temporal and spatial relationships over long sequences of frames, ensuring a comprehensive extraction of feature information. Through extensive experiments, our proposed method demonstrates superior performance compared to state-of-the-art SVSR techniques in terms of reconstruction quality. Additionally, it demonstrates outstanding performance in downstream satellite video applications, showcasing its potential in satellite video processing tasks. The source code is publicly available at https://github.com/Yu-Wang-0801/NSBRNet. Yu Wang 0140, Xiaolong Zuo, Tao Lu 0001, Jiaming Wang 0001, Yuankun Wang, Siyuan Wang 0011, Zhizheng Zhang 0009, Xiaojin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Matching While Perceiving: Enhance Image Feature Matching with Applicable Semantic AmalgamationabstractImage feature matching is a cardinal problem in computer vision, aiming to establish accurate correspondences between two-view images. Existing methods are constrained by the performance of feature extractors and struggle to capture local information affected by sparse texture or occlusions. Recognizing that human eyes consider not only similar local geometric features but also high-level semantic information of scene objects when matching images, this paper introduces SemaGlue. This novel algorithm perceives and incorporates semantic information into the matching process. In contrast to recent approaches that leverage semantic consistency to narrow the scope of matching areas, SemaGlue achieves semantic amalgamation with the designed Semantic-Aware Fusion (SAF) Block by injecting abundant semantic features from the pre-trained segmentation model. Moreover, the Cross-Domain Alignment (CDA) Block is proposed to address domain alignment issues, bridging the gaps between semantic and geometric domains to ensure applicable semantic amalgamation. Extensive experiments demonstrate that SemaGlue outperforms state-of-the-art methods across various applications such as homography estimation, relative pose estimation, and visual localization. Zhenjie Zhu, Zizhuo Li, Tao Lu 0001, Jiayi Ma 0001 |
AAAI | 4 |
| 2025 | End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph Generation
Yanduo Zhang, Tao Lu 0001, Huiqin Zhang, Jiayi Ma 0001, Huabing Zhou |
ICCV | 3 |
| 2025 | DiffuseDoc: Document geometric rectification via diffusion model
Wenfei Xiong, Huabing Zhou, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Contour-texture preservation transformer for face super-resolution
Ziyi Wu 0001, Yanduo Zhang, Tao Lu 0001, Kanghui Zhao, Jiaming Wang 0001 |
Neurocomputing | 3 |
| 2025 | GDPS: A general distillation architecture for end-to-end person search
Shichang Fu, Tao Lu 0001, Jiaming Wang 0001, Jiayi Cai, Kui Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Lightweight remote sensing super-resolution with multi-scale graph attention network
Yu Wang 0140, Tao Lu 0001, Xiao Huang 0003, Jiaming Wang 0001, Zhizheng Zhang 0009, Xiaolong Zuo |
Pattern Recognit. | 3 |
| 2025 | CMANet: A CNN-Mamba aggregation network for face super-resolution
Ziyi Wu 0001, Tao Lu 0001, Yanduo Zhang, Xiaoyu Chai |
Pattern Recognit. | 2 |
| 2025 | Frequency Decoupling Fusion for Image Super-Resolution of Remote SensingabstractBenefiting from the excellent global expression ability, transformer-based image super-resolution (SR) has made significant progress. However, the existing transformer-based SR methods still have the problem of high-frequency information reconstruction loss when processing remote sensing images due to their wide imaging range, rich high-frequency information and large differences, which affects the characterization ability of the transformer. In addition, the high computational overhead is unacceptable. To alleviate the above problems, we consider the remote sensing image SR from the perspective of the frequency domain. Specifically, we propose an efficient frequency decoupling-fusion remote sensing image SR framework, which is called FDFNet. In particular, we consider that when a large amount of previous work was carried out to extract features in the spatial domain, it was very easy to lose the high-frequency information in the original image. Therefore, we first introduce a frequency decoupling block (FDB), which decouples the image into low-frequency and high-frequency components, processes high-frequency and low-frequency information respectively in a divide-and-conquer manner, and restores high-frequency details before delving deeper. Furthermore, we notice that spatial self-attention is a low-pass filter that tends to have global perception and to demonstrate limitations in reconstructing high-frequency details. Therefore, we meticulously designed a parallel frequency-aware transformer module (PFTM) to extract spatial frequency attention and channel transposition attention, which enables our model to focus more on local texture details to restore high-frequency details. A large number of experimental results on multiple public datasets show that our FDFNet outperforms the state-of-the-art SR methods in quantitative metrics and visual quality, and achieves a balance between performance and efficiency within a limited computing budget. Kanghui Zhao, Tao Lu 0001, Jiaming Wang 0001, Yu Wang 0140, Yuanzhi Wang, Yanduo Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Multi-Stage Statistical Texture-Guided GAN for Tilted Face FrontalizationabstractExisting pose-invariant face recognition mainly focuses on frontal or profile, whereas high-pitch angle face recognition, prevalent under surveillance videos, has yet to be investigated. More importantly, tilted faces significantly differ from frontal or profile faces in the potential feature space due to self-occlusion, thus seriously affecting key feature extraction for face recognition. In this paper, we asymptotically reshape challenging high-pitch angle faces into a series of small-angle approximate frontal faces and exploit a statistical approach to learn texture features to ensure accurate facial component generation. In particular, we design a statistical texture-guided GAN for tilted face frontalization (STG-GAN) consisting of three main components. First, the face encoder extracts shallow features, followed by the face statistical texture modeling module that learns multi-scale face texture features based on the statistical distributions of the shallow features. Then, the face decoder performs feature deformation guided by the face statistical texture features while highlighting the pose-invariant face discriminative information. With the addition of multi-scale content loss, identity loss and adversarial loss, we further develop a pose contrastive loss of potential spatial features to constrain pose consistency and make its face frontalization process more reliable. On this basis, we propose a divide-and-conquer strategy, using STG-GAN to progressively synthesize faces with small pitch angles in multiple stages to achieve frontalization gradually. A unified end-to-end training across multiple stages facilitates the generation of numerous intermediate results to achieve a reasonable approximation of the ground truth. Extensive qualitative and quantitative experiments on multiple-face datasets demonstrate the superiority of our approach. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Chao Liang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | A dual-archive niche with two-stage directed differential evolution for multimodal multi-objective optimization
Yilin Chen 0001, Tao Lu 0001, Yiqi Wu, Xiangyun Liao, Qiong Wang 0001 |
J. Supercomput. | 3 |
| 2025 | A multi-scale large kernel attention with U-Net for medical image registration
Yilin Chen 0001, Tao Lu 0001, Lu Zou, Xiangyun Liao |
J. Supercomput. | 3 |
| 2025 | Hidden Inverted Specific-Class Distance Measure for Nominal AttributesabstractThe inverted specific-class distance measure (ISCDM) is a popular distance metric that uses conditional probability term to calculate the distance between two nominal attribute values, but the reliability of the conditional probability term is limited by the attribute independence assumption, which leads to the suboptimal performance in applications involving sophisticated attribute dependencies. To obtain more accurate conditional probability estimation, in this study, we derive an enhanced ISCDM by leveraging structure extension to alleviate the unrealistic attribute assumption. We denominate the resulting model as the hidden inverted specific-class distance measure (HISCDM). In HISCDM, the structure extension scheme of hidden naive bayes is adopted to find the weighted dependence relationships between attributes, and then is incorporated into the conditional probability estimation. The comprehensive experimental results demonstrate that our proposed HISCDM significantly outperforms all other methods used for comparison in terms of classification accuracy. Fang Gong, Tao Lu 0001, Kuayue Liu |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Rethinking the Role of Panchromatic Images in Pan-SharpeningabstractRecent pan-sharpening methods have predominantly utilized techniques tailored for natural image scenes, often overlooking the unique features arising from non-overlapping spectral responses. In light of this, we have reevaluated the utility of panchromatic (PAN) images and introduced a theory anchored in the spectral response of satellite sensors. This posits that a PAN image is effectively a linear weighted summation of individual bands from its corresponding multi-spectral (MS) image, offset by an error map. We developed a deep unmixing network termed “DUN” that integrates an unmixing network, a fusion mechanism, and a distinctive mutual information contrastive loss function. Notably, the unmixing network is adept at decomposing a PAN image into its MS counterpart and error map. Further, the demixed image alongside the low-resolution MS image is channeled into the fusion network for pan-sharpening. Recognizing the challenges of achieving robust supervised learning directly from the unmixing phase, we have innovated a mutual information contrastive learning loss function, ensuring enhanced separation and minimizing overlap during the unmixing process. Preliminary experiments underscore both the quantitative and qualitative prowess of the proposed method. Jiaming Wang 0001, Xitong Chen, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140, Tao Lu 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Double-Graph Representation With Relational Enhancement for Emotion-Cause Pair ExtractionabstractThe emotion-cause pair extraction (ECPE) task is to simultaneously extract emotions and causes as pairs (EC-pairs) from documents, which is important for natural language processing. Previous research tackled this task via a two-step approach, which first predicts separately the emotion and cause clauses, and then pairs them up by using a binary classifier. However, such a two-step approach may suffer from the possible propagation of errors, and it neglects the interaction between emotions and causes. In this article, an end-to-end double-graph method with relational enhancement (DGRE) is proposed to stimulate two relationship modes among clauses, i.e., semantic dependence and logical dependence. First, two united graph encoders are established to embed the semantic dependence into the representation of clauses and pairs. The first encoder is built on graph attention networks (GATs) for clause-level representation, the result of which is used by a relational graph convolutional network (RGCN) for the refinement of pair-level representation. Aiming to enhance the fitting ability of logical dependence, the emotion-type classification task is introduced into the multitask learning framework of GATs, which can effectively distinguish the logical relations between clauses according to their emotion types. Moreover, seven types of dependence relations have been designed for the node connections in RGCN, which emphasize the contextual interaction and clustering among neighboring nodes. Experiments on a benchmark Chinese corpus demonstrate that the proposed DGRE approach could effectively establish the communication mechanism between clauses and pairs from multiple perspectives, and comparisons with state-of-the-art (SOTA) models well validate its effectiveness. Zhe Chen 0029, Vasile Palade, Tao Lu 0001, Junchi Zhang, Yanduo Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | A Robust Mutual-Reinforcing Framework for 3D Multi-Modal Medical Image Fusion Based on Visual-Semantic ConsistencyabstractThis work proposes a robust 3D medical image fusion framework to establish a mutual-reinforcing mechanism between visual fusion and lesion segmentation, achieving their double improvement. Specifically, we explore the consistency between vision and semantics by sharing feature fusion modules. Through the coupled optimization of the visual fusion loss and the lesion segmentation loss, visual-related and semantic-related features will be pulled into the same domain, effectively promoting accuracy improvement in a mutual-reinforcing manner. Further, we establish the robustness guarantees by constructing a two-level refinement constraint in the process of feature extraction and reconstruction. Benefiting from full consideration for common degradations in medical images, our framework can not only provide clear visual fusion results for doctor's observation, but also enhance the defense ability of lesion segmentation against these negatives. Extensive evaluations of visual fusion and lesion segmentation scenarios demonstrate the advantages of our method in terms of accuracy and robustness. Moreover, our proposed framework is generic, which can be well-compatible with existing lesion segmentation algorithms and improve their performance. The code is publicly available at https://github.com/HaoZhang1018/RMR-Fusion. Hao Zhang 0073, Xuhui Zuo, Huabing Zhou, Tao Lu 0001, Jiayi Ma 0001 |
AAAI | 4 |
| 2024 | Adaptive Cross-Spatial Sensing Network for Change Detection
Liyuan Jin, Yanduo Zhang, Tao Lu 0001, Jiaming Wang 0001 |
PRCV (13) | 3 |
| 2024 | A Deep Error Removal Network for Pan-SharpeningabstractThe phenomenon of nonoverlapping spectral responses is an inevitable but an overlooked problem in the deep-learning-based panchromatic (PAN) and multispectral (MS) images’ fusion task, which will introduce some error information from the PAN image. In light of this, we construct a novel prior model based on spectral response theory and develop a model-based pan-sharpening network. Specifically, we extract the initial error map from the PAN and interpolate the MS image as the initial pan-sharpened result. Then, two optimization problems regularized by the deep prior are formulated to update the error map and pan-sharpened image. By alternately optimizing the above subtasks, error information is gradually separated from PAN images and the lost texture information in MS images is gradually restored, which can effectively alleviate the negative impact of low coupling information from PAN and MS images. Plenty of experimental results on different kinds of satellite datasets demonstrate that the proposed method shows a better balance between interpretability and lightweight structure. The proposed method will be open-sourced inhttps://github.com/jiaming-wang/DERN. Jiaming Wang 0001, Tao Lu 0001, Xiao Huang 0003, Ruiqian Zhang, Dongyue Luo |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Remote Sensing Pan-Sharpening via Cross-Spectral-Spatial Fusion NetworkabstractPan-sharpening is a technique used to create high-resolution multispectral (HRMS) images by merging low-resolution multispectral (LRMS) images with corresponding high-resolution panchromatic (PAN) images. Despite achieving state-of-the-art performance, existing panchromatic sharpening networks based on deep-learning (DL) methods suffer from spectral distortion and insufficient spatial texture enhancement. To address this challenge, this letter introduces a novel cross-spectral–spatial fusion network (CSSFN) for pan-sharpening remote sensing images. The network utilizes a cross-spectral–spatial attention block (CSSAB) to extract both the spectral information of the MS branch and the spatial information of the PAN branch. The spectral and spatial feature representations of remote sensing images are then progressively enhanced, which improves the fusion process and generates multispectral images with high spatial resolution. Our network outperforms other pan-sharpening methods on two publicly available datasets, as demonstrated by extensive experiments, yielding state-of-the-art results. Yu Wang 0140, Tao Lu 0001, Jiaming Wang 0001, Gui Cheng, Xiaolong Zuo, Chaoya Dang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Pan-sharpening via intrinsic decomposition knowledge distillation
Jiaming Wang 0001, Xiao Huang 0003, Ruiqian Zhang, Xitong Chen, Tao Lu 0001 |
Pattern Recognit. | 6 |
| 2024 | Hyper-Laplacian Prior for Remote Sensing Image Super-ResolutionabstractImage explicit prior has made breakthrough progress in the super-resolution (SR) due to the additional supervisory information provided. However, existing explicit prior-guided SR methods directly use the Gaussian gradient or Laplacian gradient prior, which cannot fit the gradient distribution of remote sensing images. Through the statistics of gradient probability density distribution of the remote sensing image dataset, we found that the hyper-Laplacian prior can fit the heavy-tailed distribution better, which aroused us to use the hyper-Laplacian before facilitating the SR reconstruction. We propose a novel hyper-Laplacian prior SR method for remote sensing images in this manuscript. Specifically, our model consists of three components: rough reconstruction subnetwork (RRS), hyper-Laplacian prior subnetwork (HPS), and image refinement enhancement subnetwork (RES). In the RRS, we reconstruct low-resolution (LR) images into rough SR images by a set of resblocks. In the HPS, we first introduce the hyper-Laplacian prior for LR images to provide an additional texture. Hereafter, we set up a prior loss which imposes a second-order supervision on the SR image. Like the previous image space loss function, it helps the model to gather the geometric structure of the image. Finally, the outputs of the RRS and HPS are fused and then fed to the RES for high-quality image reconstruction. Numerous studies of SR reconstruction and segmentation on UCMerced, PatternNet, and OpenBayes datasets confirm that our method is superior compared to state-of-the-art methods. Kanghui Zhao, Tao Lu 0001, Jiaming Wang 0001, Yanduo Zhang, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Learning to Hallucinate Face in the DarkabstractFace hallucination in low-light environments is an extremely challenging task due to the significant loss of facial structure and facial texture information. Although cascading image relighting and face hallucination tasks is a feasible strategy, simply cascading these two tasks does not achieve satisfactory results because they do not fit into each other naturally. In this article, we propose a novel duplex fusing-embedding learning approach to tackle this challenge in low-light environments. The core of the proposed approach is the duplexity of feature fusion and embedding between relighting and hallucination tasks. In the feature fusion phase, the shallow features from two tasks are bidirectionally fused and activated into a consistent feature space. In the feature embedding phase, the fused features from the previous iteration are fed back and bidirectionally embedded into the deep features of two tasks in the current iteration so that they can learn feature representations that consistently represent both tasks, thereby boosting the performance of relighting and hallucination to generate photorealistic HR face images. Experimental results show that the proposed approach allows current face hallucination methods to learn to hallucinate face in the dark. Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Zixiang Xiong |
IEEE Trans. Multim. | 2 |
| 2024 | Rethinking Prior-Guided Face Super-Resolution: A New Paradigm With Facial Component PriorabstractRecently, facial priors (e.g., facial parsing maps and facial landmarks) have been widely employed in prior-guided face super-resolution (FSR) because it provides the location of facial components and facial structure information, and helps predict the missing high-frequency (HF) information. However, most existing approaches suffer from two shortcomings: 1) the extracted facial priors are inaccurate since they are extracted from low-resolution (LR) or low-quality super-resolved (SR) face images and 2) they only consider embedding facial priors into the reconstruction process from LR to SR face images, thus failing to explore facial priors to generate LR face image. In this article, we propose a novel pre-prior guided approach that extracts facial prior information from original high-resolution (HR) face images and embeds them into LR ones to obtain HF information-rich LR face images, thereby improving the performance of face reconstruction. Specifically, a novel component hybrid method is proposed, which fuses HR facial components and LR facial background to generate new LR face images (namely, LRmix) via facial parsing maps extracted from HR face images. Furthermore, we design a component hybrid network (CHNet) that learns the LR to LRmix mapping function to ensure that the LRmix can be obtained from LR face images in testing and real-world datasets. Experimental results show that our proposed scheme significantly improves the reconstruction performance for FSR. Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Junjun Jiang, Zhongyuan Wang 0001, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Omniscient Video Super-Resolution with Explicit-Implicit AlignmentabstractWhen considering the temporal relationships, most previous video super-resolution (VSR) methods follow the iterative or recurrent framework. The iterative framework adopts neighboring low-resolution (LR) frames from a sliding window, while the recurrent framework utilizes the output generated in the previous SR procedure. The hybrid framework combines them but still cannot fully leverage the temporal relationships. Meanwhile, the existing methods are limited in the receptive field of the optical flow or lack semantic constrains on motion information. In this work, we propose an omniscient framework to fully explore the temporal relationships in the video, which encompasses both LR frames and SR outputs from the past, present, and future. The omniscient framework is more generic because the iterative, recurrent, and hybrid frameworks can be regarded as its special cases. Besides, when addressing the motion information, most previous VSR methods adopt the explicit motion estimation and compensation, while many recent methods turn to implicit alignment. In implicit alignment methods, because basic non-local means suffers from heavy computational costs, we improve it by capturing the non-local correlations in a relatively local manner to reduce the complexity. Moreover, we integrate the explicit and implicit methods into an explicit-implicit alignment module to better utilize motion information. We have conducted extensive experiments on public datasets, which show that our method is superior over the state-of-the-art methods in objective metrics, subjective visual quality, and complexity. In particular, on datasets of Vid4 and UDM10, our method improves PSNR by 0.19 dB, 0.49 dB against the most advanced method BasicVSR++, respectively. Peng Yi 0002, Zhongyuan Wang 0002, Laigan Luo, Kui Jiang, Zheng He 0001, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super ResolutionabstractRecently, face super-resolution has achieved pleasing performance. Numerous works have shown that texture features and structural information play a crucial role for super-resolution reconstruction. However, effective co-learning of both has been limiting the performance improvement of existing state-of-the-art methods. Therefore, we focus on the texture and structure of images in this paper, and design a two-branch network containing a texture network (T-Net) and a structure network (S-Net) to jointly explore texture and structure information for co-learning. T-Net serves as the backbone network to learn both texture and structure information for reconstruction, while the S-Net serves as the auxiliary network that can effectively exploit the multi-scale information of the T-Net encoder to recover the structure. To better facilitate the co-learning of the two branches, two co-learning modules deal with the information flow interaction between the encoder and decoder of the two branches, respectively, thus explicitly guiding the structure-aware image reconstruction. Additionally, a dense feature enhancement module investigates the channel and spatial correlation of features and enhances the representation capability of the network. Extensive experiments on the CelebA and Helen datasets show that our proposed approach outperforms state-of-the-art methods. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008 |
ICASSP | 3 |
| 2023 | Remote Sensing Image Super-Resolution via Multiscale Enhancement NetworkabstractIn recent years, remote sensing images have attracted a lot of attention because of their special value. However, images acquired by satellite sensors are usually low-resolution (LR), so remote sensing images are much more difficult to infer high-frequency details from compared with ordinary digital images, which means they cannot meet the needs of certain downstream tasks. In this letter, we propose a multiscale enhancement network (MEN), which uses multiscale features of remote sensing images to enhance the network’s reconstruction capability. Specifically, the network extracts the coarse features of LR remote sensing images using convolutional layers. Then, these features are fed into the multiscale enhancement module (MEM) proposed by this network, which uses a combination of convolutional layers with multiple convolutional kernel sizes to refine the extraction of multiscale features, and finally, the final reconstructed image is generated by the reconstruction module. Extensive experiments show that MEN achieves significant reconstruction advantages in both objective and subjective aspects. Yu Wang 0140, Tao Lu 0001, Changzhi Wu, Jiaming Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Self-attention learning network for face super-resolution
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Jiaming Wang 0001, Zixiang Xiong |
Neural Networks | 3 |
| 2023 | PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001 |
Pattern Recognit. | 6 |
| 2023 | Implicit space pose consistent transfer network for deep face verification
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Zhen Han 0002 |
Pattern Recognit. Lett. | 3 |
| 2023 | FaceFormer: Aggregating Global and Local Representation for Face HallucinationabstractRecently, face hallucination methods either feed whole face image into convolutional neural networks (CNNs) or utilize extra facial priors (e.g., facial parsing maps and landmarks) to focus on global facial structure and constrain facial texture generation. However, the limited receptive fields of CNNs and inaccurate facial priors will reduce the naturalness and fidelity of restored face. In this paper, we propose a FaceFormer that aggregates global representation of Transformers and local representation of CNNs to maintain the consistency of facial structure while restoring local facial details. The reason for this design is that the Transformer can capture global facial information by exploiting the long-distance visual relation modeling, while the local modeling capability of CNNs can recover fine-grained facial details. Therefore, aggregating these two independent representations can help to maximize their merits and reconstruct high-quality and high-fidelity face images. Experimental results of face reconstruction and recognition verify that the proposed FaceFormer significantly outperforms current state-of-the-arts. Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Joint Segmentation and Identification Feature Learning for Occlusion Face RecognitionabstractThe existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification. Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | Degrade Is Upgrade: Learning Degradation for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and color in two steps. Inspired by the color image formulation (diffuse illumination color plus environment illumination color), we first estimate the degradation from low-light inputs to simulate the distortion of environment illumination color, and then refine the content to recover the loss of diffuse illumination color. To this end, we propose a novel Degradation-to-Refinement Generation Network (DRGN). Its distinctive features can be summarized as 1) A novel two-step generation network for degradation learning and content refinement. It is not only superior to one-step methods, but also capable of synthesizing sufficient paired samples to benefit the model training; 2) A multi-resolution fusion network to represent the target information (degradation or contents) in a multi-scale cooperative manner, which is more effective to address the complex unmixing problems. Extensive experiments on both the enhancement task and the joint detection task have verified the effectiveness and efficiency of our proposed method, surpassing the SOTA by 1.59dB on average and 3.18\% in mAP on the ExDark dataset. The code will be available soon. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Chen Chen 0001, Peng Yi 0002, Tao Lu 0001, Chia-Wen Lin |
AAAI | 6 |
| 2022 | Video Face Recognition Using Neural Aggregation Networks with Mutual Relational LearningabstractVideo face recognition benefits profoundly from deep convolutional neural networks (CNNs), which learn robust feature embeddings. However, due to their fixed geometric structures, CNNs are inherently limited in modeling the significant variations from the angle, pose, occlusion and other factors of face images. In this paper, a neural aggregation network based on mutual relation learning is proposed for video face recognition. First, Intra-frame Relational Learning network (Intra-Net) is introduced, which models the interdependencies between the re-gional components of individual features and develops relevance between fine-grained features. Such processing can determine the region of interest adaptively according to the quality of the input face image to achieve the extraction of valuable information. Secondly, we introduce Inter-frame Relational Learning Network (Inter-Net), which considers the most significant appearance representation in the overall structure of the face image to cor-relate the complementarity of features between frames. Finally, information aggregation is performed by combining Inter-Net and Intra-Net. Joint optimization of the two branches allows our model to effectively exploit the complementary information between them to improve the aggregation capability. We validate the effectiveness of our model for video face recognition, proving its superiority over state-of-the-art methods on two benchmark datasets. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008 |
ICTAI | 3 |
| 2022 | Deep locally linear embedding network
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Xitong Chen |
Inf. Sci. | 4 |
| 2022 | Structure-Texture Parallel Embedding for Remote Sensing Image Super-ResolutionabstractThe structure and texture of images are crucial for remote sensing image super-resolution. Generative adversarial networks (GANs) recover image details through adversarial training. However, the recovered images always have structural distortions on the one hand, and GANs are difficult to train on the other hand. In addition, some methods assist reconstruction by introducing prior information of the image, but this brings additional computational cost. To address this issue, we propose a novel structure-texture parallel embedding (SPE) method for super-resolution (SR) of remote sensing images. Our method does not require additional image priors to reconstruct high-quality images. Specifically, we use the global structure information and local texture information of the image in the ascending space to guide the reconstruction result of the image. Firstly, we design a structure preserving block (SPB) to extract global structural features in the ascending space of the image, so as to obtain global structure information for a priori representation. Then, we design a local texture attention module (LTAM) to restore richer texture details. We have conducted lots of experiments on Draper public dataset. Experimental results show that our proposed method not only achieves a better trade-off between computational cost and performance, but also outperforms the existing several SR methods in terms of objective index evaluation and subjective visual effects. Tao Lu 0001, Kanghui Zhao, Yuntao Wu, Zhongyuan Wang 0001, Yanduo Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | SSCAN: A Spatial-Spectral Cross Attention Network for Hyperspectral Image DenoisingabstractHyperspectral images (HSIs) have been widely used in a variety of applications thanks to the rich spectral information they are able to provide. Among all HSI processing tasks, HSI denoising is a crucial step. Recent years have seen great progress in deep learning-based image denoising methods. However, existing efforts tend to ignore the correlations between adjacent spectral bands, leading to problems such as spectral distortion and blurred edges in denoised results. In this study, we propose a novel HSI denoising network, termed spectral–spatial cross attention network (SSCAN), that combines group convolutions and attention modules. Specifically, we use a group convolution with a spatial attention module to facilitate feature extraction by directing models’ attention to bandwise important features. We also propose a spectral–spatial attention block (SSAB) to effectively exploit the spatial and spectral information in HSIs. In addition, we adopt residual learning operations with skip connections to ensure training stability. The experimental results indicate that the proposed SSCAN outperforms several state-of-the-art HSI denoising algorithms. Xiao Huang 0003, Jiaming Wang 0001, Tao Lu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Deep representation learning for face hallucination
Tao Lu 0001, Yu Wang 0140, Ruobo Xu, Wei Liu 0123, Wenhua Fang, Yanduo Zhang |
Multim. Tools Appl. | 1 |
| 2022 | A Progressive Fusion Generative Adversarial Network for Realistic and Consistent Video Super-ResolutionabstractHow to effectively fuse temporal information from consecutive frames remains to be a non-trivial problem in video super-resolution (SR), since most existing fusion strategies (direct fusion, slow fusion, or 3D convolution) either fail to make full use of temporal information or cost too much calculation. To this end, we propose a novel progressive fusion network for video SR, in which frames are processed in a way of progressive separation and fusion for the thorough utilization of spatio-temporal information. We particularly incorporate multi-scale structure and hybrid convolutions into the network to capture a wide range of dependencies. We further propose a non-local operation to extract long-range spatio-temporal correlations directly, taking place of traditional motion estimation and motion compensation (ME&MC). This design relieves the complicated ME&MC algorithms, but enjoys better performance than various ME&MC schemes. Finally, we improve generative adversarial training for video SR to avoid temporal artifacts such as flickering and ghosting. In particular, we propose a frame variation loss with a single-sequence training method to generate more realistic and temporally consistent videos. Extensive experiments on public datasets show the superiority of our method over state-of-the-art methods in terms of performance and complexity. Our code is available at https://github.com/psychopa4/MSHPFNL. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006 |
Pattern Recognit. | 3 |
| 2022 | Few-Shot Semantic Segmentation via Frequency Guided Neural NetworkabstractPrototype learning is extensively used in few-shot semantic segmentation due to its excellent capability of semantic information extraction and effective prevention of overfitting. The previous prototype based methods ignore the frequency discrepancy inside the object, thereby leading to semantic confusion of the object. In this paper, we propose a frequency guided network (FGNet) which explicitly models the semantic information of different frequencies and precisely guides the semantic alignment of the object. Specifically, the proposed FGNet consists of two modules: a frequency separation module (FSM) and a multi-guided feature enrichment module (MG-FEM) to complete the multi-frequency semantic information extraction and alignment, respectively. Experiments on PASCAL-$5^{i}$dataset show that our FGNet achieves mIoU score of 61.2% in 1-shot which surpasses the state-of-the-art methods. Xiya Rao, Tao Lu 0001, Zhongyuan Wang 0001, Yanduo Zhang |
IEEE Signal Process. Lett. | 2 |
| 2022 | Classifying Facial Regions for Face HallucinationabstractRecently, convolutional neural networks (CNNs) have dominated the face hallucination task due to their powerful feature representation capability. However, most of them simply use the same weights to treat different facial regions without considering the reconstruction difficulty of different facial regions, resulting in the component regions (e.g., eyes, nose, mouth) of the reconstructed faces tending to be blurred. In this paper, we propose a novel facial region classification network (FRCN) to address this problem. The proposed method first divides the input low-resolution (LR) facial image into several patch blocks, then classifies them into three categories according to their reconstruction difficulty, and finally inputs the three types of patch blocks into three networks with different weights for reconstruction and combining, thereby recovering high-quality high-resolution (HR) facial image. Experimental results show that FRCN can remarkably improve face reconstruction's performance. Yiyao Wang, Tao Lu 0001, Yuanzhi Wang, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | A Dual-Path Fusion Network for Pan-SharpeningabstractMost existing deep learning-based pan-sharpening methods own several widely recognized issues, such as spectral distortion and insufficient spatial texture enhancement. To address these challenges in pan-sharpening, we propose a novel dual-path fusion network (DPFN). The proposed DPFN includes two major components: 1) the global subnetwork (GSN) and 2) the local subnetwork (LSN). In particular, GSN aims to search similar image blocks in panchromatic (PAN) space and multispectral (MS) space and exploits HR textural information from the PAN space and spectral information from the MS space for the fine representation of pan-sharpened MS features by employing a cross nonlocal block. Meanwhile, the proposed LSN based on a high-pass modification block (HMB) is designed to learn the high-pass information, aiming to enhance bandwise spatial information from MS images. HMB forces the fused image to obtain high-frequency details from PAN images. Moreover, to facilitate the generation of visually appealing pan-sharpened images, we propose a perceptual loss function and further optimize the model based on high-level features in the near-infrared space. Experiments demonstrate the superior performance of the proposed method quantitatively and qualitatively compared to existing state-of-the-art pan-sharpening methods. The source code is available athttps://github.com/jiaming-wang/DPFN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Pan-Sharpening via Deep Locally Linear Embedding Residual NetworkabstractThe goal of pan-sharpening tasks is to fuse panchromatic (PAN) images and low-spatial-resolution (LR) multispectral (MS) images for the purpose of aggregating texture and spectral information. Although traditional embedding-based pan-sharpening methods achieve competitive results, they are limited by the shallow network and not suitable for large-scale datasets. In this study, we design a novel multiscale locally linear embedding residual network (LLERN) that consists of two phases: the spectral preservation phase and the structural preservation phase. As the pretreatment of the structural preservation network, the spectral preservation network aims to upscale the LR MS image while retaining spectral information. The proposed locally linear embedding residual block (LLERB) in the structural preservation phase can search for similar sparse patches from the PAN image space and embed the corresponding local geometric relationship into the residual space to enhance the MS image. Extensive experiments suggest that the proposed LLERN outperforms state-of-the-art methods from visual and quantitative perspectives, and confirm the assumption that LR image patches and residual image patches in a local region share a similar manifold structure, which can be used to guide deep-learning modeling with improved interpretability. The source code is available athttps://github.com/jiaming-wang/LLERN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Gui Cheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | From Artifact Removal to Super-ResolutionabstractDeep-learning-based super-resolution methods have been extensively studied and achieved significant performance with deep convolutional neural networks. However, the results still suffer from the ringing effect, especially in satellite image super-resolution tasks, due to the loss of image details in the satellite degradation process. In this paper, we build a novel satellite super-resolution framework by decomposing a high-resolution image into three components, i.e., low-resolution, artifact, and high-frequency information. Specifically, we propose an artifact removal network with a self-adaption difference convolution (SDC) to fully exploit the structure prior in the low-resolution image and predict the artifact map. Considering that the artifact map and the high-frequency map share a similar pattern, we introduce the supervised structure correction block (SSC) that establishes a bridge between the high-frequency generation process and the artifact removal process. Experimental results on satellite images demonstrate that the proposed method owns an improved tradeoff between the performance and the computational cost compared to existing state-of-the-art satellite and natural super-resolution methods. The source code is available at https://github.com/jiaming-wang/ARSRN. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dual-Path Deep Fusion Network for Face Image HallucinationabstractAlong with the performance improvement of deep-learning-based face hallucination methods, various face priors (facial shape, facial landmark heatmaps, or parsing maps) have been used to describe holistic and partial facial features, making the cost of generating super-resolved face images expensive and laborious. To deal with this problem, we present a simple yet effective dual-path deep fusion network (DPDFN) for face image super-resolution (SR) without requiring additional face prior, which learns the global facial shape and local facial components through two individual branches. The proposed DPDFN is composed of three components: a global memory subnetwork (GMN), a local reinforcement subnetwork (LRN), and a fusion and reconstruction module (FRM). In particular, GMN characterize the holistic facial shape by employing recurrent dense residual learning to excavate wide-range context across spatial series. Meanwhile, LRN is committed to learning local facial components, which focuses on the patch-wise mapping relations between low-resolution (LR) and high-resolution (HR) space on local regions rather than the entire image. Furthermore, by aggregating the global and local facial information from the preceding dual-path subnetworks, FRM can generate the corresponding high-quality face image. Experimental results of face hallucination on public face data sets and face recognition on real-world data sets (VGGface and SCFace) show the superiority both on visual effect and objective indicators over the previous state-of-the-art methods. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Tao Lu 0001, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Omniscient Video Super-ResolutionabstractMost recent video super-resolution (SR) methods either adopt an iterative manner to deal with low-resolution (LR) frames from a temporally sliding window, or leverage the previously estimated SR output to help reconstruct the current frame recurrently. A few studies try to combine these two structures to form a hybrid framework but have failed to give full play to it. In this paper, we propose an omniscient framework to not only utilize the preceding SR output, but also leverage the SR outputs from the present and future. The omniscient framework is more generic because the iterative, recurrent and hybrid frameworks can be regarded as its special cases. The proposed omniscient framework enables a generator to behave better than its counterparts under other frameworks. Abundant experiments on public datasets show that our method is superior to the state-of-the-art methods in objective metrics, subjective visual effects and complexity. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICCV | 5 |
| 2021 | Pan-Sharpening Via High-Pass Modification Convolutional Neural NetworkabstractMost existing deep learning-based pan-sharpening methods have several widely recognized issues, such as spectral distortion and insufficient spatial texture enhancement, we propose a novel pan-sharpening convolutional neural network based on a high-pass modification b lock. Different from existing methods, the proposed block is designed to learn the high-pass information, leading to enhance spatial information in each band of the multi-spectral-resolution images. To facilitate the generation of visually appealing pan-sharpened images, we propose a perceptual loss function and further optimize the model based on high-level features in the near-infrared space. Experiments demonstrate the superior performance of the proposed method compared to the state-of the-art pan-sharpening methods, both quantitatively and qualitatively. The proposed model is open-sourced at https://github.com/jiaming-wang/HMB. Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
ICIP | 4 |
| 2021 | Face Super-Resolution Through Dual-Identity ConstraintabstractRecently, most existing face SR methods only focus on generating pleasant texture details, even artifacts. Identity information is an important high-level face attribute, which is often ignored in the low-level super-resolution (SR) task. In view of this, we propose a dual-identity constraint dual-loop network (DIDnet), which employs identity information to constrain the SR model. First, the proposed framework consists of two closed-loop networks: one of the networks is used for generating high resolution (HR) images for exploring identity-preserving in HR feature space and the other one can learn degradation process for utilizing low resolution (LR) identity information. Furthermore, we integrate dual-identity constraints together for rendering characteristic facial images. Extensive experimental results are conducted on face databases and real-world data, which confirmed that the proposed DIDnet consistently and significantly improves both objective and subjective facial image reconstruction performances. Fangfang Cheng, Tao Lu 0001, Yu Wang 0140, Yanduo Zhang |
ICME | 2 |
| 2021 | Unsupervised Remoting Sensing Super-Resolution via Migration Image PriorabstractRecently, satellites with high temporal resolution have fostered wide attention in various practical applications. Due to limitations of bandwidth and hardware cost, however, the spatial resolution of such satellites is considerably low, largely limiting their potentials in scenarios that require spatially explicit information. To improve image resolution, numerous approaches based on training low-high resolution pairs have been proposed to address the super-resolution (SR) task. De-spite their success, however, low/high spatial resolution pairs are usually difficult to obtain in satellites with a high temporal resolution, making such approaches in SR impractical to use. In this paper, we proposed a new unsupervised learning framework, called "MIP", which achieves SR tasks without low/high resolution image pairs. First, random noise maps are fed into a designed generative adversarial network (GAN) for reconstruction. Then, the proposed method converts the reference image to latent space as the migration image prior. Finally, we update the input noise via an implicit method, and further transfer the texture and structured information from the reference image. Extensive experimental results on the Draper dataset show that MIP achieves significant improvements over state-of-the-art methods both quantitatively and qualitatively. The proposed MIP is open-sourced at https://github.com/jiaming-wang/MIP. Jiaming Wang 0001, Tao Lu 0001, Xiao Huang 0003, Ruiqian Zhang, Yu Wang 0140 |
ICME | 3 |
| 2021 | Face Hallucination via Split-Attention in Split-Attention NetworkabstractRecently, convolutional neural networks (CNNs) have been widely employed to promote the face hallucination due to the ability to predict high-frequency details from a large number of samples. However, most of them fail to take into account the overall facial profile and fine texture details simultaneously, resulting in reduced naturalness and fidelity of the reconstructed face, and further impairing the performance of downstream tasks (e.g., face detection, facial recognition). To tackle this issue, we propose a novel external-internal split attention group (ESAG), which encompasses two paths responsible for facial structure information and facial texture details, respectively. By fusing the features from these two paths, the consistency of facial structure and the fidelity of facial details are strengthened at the same time. Then, we propose a split-attention in split-attention network (SISN) to reconstruct photorealistic high-resolution facial images by cascading several ESAGs. Experimental results on face hallucination and face recognition unveil that the proposed method not only significantly improves the clarity of hallucinated faces, but also encourages the subsequent face recognition performance substantially. Codes have been released at https://github.com/mdswyz/SISN-Face-Hallucination. Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Yu Wang 0140, Wei Liu 0123, Zhongyuan Wang 0001, Junjun Jiang |
ACM Multimedia | 1 |
| 2021 | An effective weighted vector median filter for impulse noise reduction based on minimizing the degree of aggregationabstractAbstract Impulse noise is regarded as an outlier in the local window of an image. To detect noise, many proposed methods are based on aggregated distance, including spatially weighted aggregated distance, n nearest neighbour distance, local density, and angle‐weighted quaternion aggregated distance. However, these methods ignore the weight of each pixel or have limited adaptability. This study introduces the concept of degree of aggregation and proposes a weighting method to obtain the weight vector of the pixels by minimizing the degree of aggregation. The weight vector obtained gives larger components on the signal pixels than on the noisy pixels. Then it is fused with the aggregated distance to form a weighted aggregated distance that can reasonably characterise the noise and signal. The weighted aggregated distance, along with an adaptive segmentation method, can effectively detect the noise. To further enhance the effect of noise detection and removal, an adaptive selection strategy is incorporated to reduce the noise density in the local window. At last, noisy pixels detected are replaced with the weighted channel combination optimization values. The experimental results exhibit the validity of the proposed method by showing better performance in terms of both objective criteria and visual effects. Tongwei Lu, Feng Min, Tao Lu 0001 |
IET Image Process. | 4 |
| 2021 | Cross-task feature alignment for seeing pedestrians in the dark
Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Wenhua Fang, Yuntao Wu, Zhongyuan Wang 0001 |
Neurocomputing | 2 |
| 2021 | Spatial-temporal pooling for action recognition in videos
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Xianwei Lv 0002 |
Neurocomputing | 4 |
| 2021 | Internal and external spatial-temporal constraints for person reidentification
Jiaming Wang 0001, Tao Lu 0001, Ruiqian Zhang, Xiao Huang 0003, Xianwei Lv 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Discriminative metric learning for face verification using enhanced Siamese neural network
Tao Lu 0001, Wenhua Fang, Yanduo Zhang |
Multim. Tools Appl. | 1 |
| 2021 | Enhanced image prior for unsupervised remoting sensing super-resolution
Jiaming Wang 0001, Xiao Huang 0003, Tao Lu 0001, Ruiqian Zhang, Jiayi Ma 0001 |
Neural Networks | 4 |
| 2021 | Single Image Super-Resolution via Multi-Scale Information Polymerization NetworkabstractRecently, the performances of deep convolution neural networks (CNNs)-based single-image super-resolution (SISR) have been significantly improved. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper networks and ignore the potential relationship between multi-scale features, leading to the limited representation ability of the reconstructed network. To address this problem, we propose a new multi-scale information polymerization network (MIPN). Specifically, we propose a multi-scale information polymerization block (MIPB), which uses convolution layers of different convolution kernel sizes to extract multi-scale image features, and effectively polymerizate the extracted features together to obtain fine image features. Moreover, we also propose a shallow residual block in MIPB. Compared with the traditional convolution layer, this proposed block can effectively extract image features without increasing the number of parameters. Extensive experiments show that the proposed method performs better than several state-of-the-art methods in quantitative and visual quality indicators. Tao Lu 0001, Yu Wang 0140, Jiaming Wang 0001, Wei Liu 0123, Yanduo Zhang |
IEEE Signal Process. Lett. | 1 |
| 2021 | Seeing in the Dark by Component-GANabstractRecently, Retinex theory based low-light image enhancement (LLIE) algorithms have achieved impressive results in controlled environment. However, the majority of deep learning based LLIE algorithms leverage relighting by enhancing the illumination components that directly determines the image brightness, regretfully, they ignore the information of reflectance components, which may cause problems such as image noise and color distortion in reconstructed images. To tackle this problem, in this letter, we propose a component enhancement network based on Generative Adversarial Network (Component-GAN) for recovering clear images from low-light ones. Specifically, the network is composed of the decomposition part for dividing the paired low/normal-light images into illumination components and reflectance components, and the enhancement part for generating high-quality images. It is worth to note that we provide two branches of component enhancement network, which are parallel to improve the two components simultaneously. Hereby, we treat the reconstruction part as the generative network and adopt discriminative network to boost image reconstruction performance. Through extensive experiments, the proposed approach outperforms some state-of-the-art LLIE methods in terms of visual and subjective qualities. Ning Rao, Tao Lu 0001, Yanduo Zhang, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Decomposition Makes Better Rain Removal: An Improved Attention-Guided Deraining NetworkabstractRain streaks in the air show diverse characteristics with different shapes, directions, densities, even the complex overlapped phenomenon, causing great challenges for the deraining task. Recently, deep learning based image deraining methods have been extensively investigated due to their excellent performance. However, most of the existing algorithms still have limitations in removing rain streaks while preserving rich textural details under complicated rain conditions. To this end, we propose to decompose rain streaks into multiple rain layers and individually estimate each of them along the network stages to cope with the increasing abstracts. To better characterize rain layers, an improved non-local block is designed to exploit the self-similarity of rain information by learning the holistic spatial feature correlations while reducing the calculation complexity. Moreover, a mixed attention mechanism is applied to guide the fusion of rain layers by focusing on the local and global overlaps among these rain layers. Extensive experiments on both synthetic rainy/rain-haze/raindrop datasets, real-world samples, the haze, and low-light scenarios show substantial improvements both on quantitative indicators and visual effects over the current state-of-the-art technologies. The source code is available athttps://github.com/kuihua/IADN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zhen Han 0002, Tao Lu 0001, Baojin Huang, Junjun Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Robust Feature Matching for Remote Sensing Image Registration via Linear Adaptive FilteringabstractAs a fundamental and critical task in feature-based remote sensing image registration, feature matching refers to establishing reliable point correspondences from two images of the same scene. In this article, we propose a simple yet efficient method termed linear adaptive filtering (LAF) for both rigid and nonrigid feature matching of remote sensing images and apply it to the image registration task. Our algorithm starts with establishing putative feature correspondences based on local descriptors and then focuses on removing outliers using geometrical consistency priori together with filtering and denoising theory. Specifically, we first grid the correspondence space into several nonoverlapping cells and calculate a typical motion vector for each one. Subsequently, we remove false matches by checking the consistency between each putative match and the typical motion vector in the corresponding cell, which is achieved by a Gaussian kernel convolution operation. By refining the typical motion vector in an iterative manner, we further introduce a progressive strategy based on the coarse-to-fine theory to promote the matching accuracy gradually. In addition, an adaptive parameter setting strategy and posterior probability estimation based on the expectation-maximization algorithm enhance the robustness of our method to different data. Most importantly, our method is quite efficient where the gridding strategy enables it to achieve linear time complexity. Consequently, some sparse point-based tasks may inspire from our method when they are achieved by deep learning techniques. Extensive feature matching and image registration experiments on several remote sensing data sets demonstrate the superiority of our approach over the state of the art. Xingyu Jiang 0005, Jiayi Ma 0001, Aoxiang Fan, Haiping Xu, Geng Lin, Tao Lu 0001, Xin Tian 0006 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Image Defogging Quality Assessment: Real-World Database and MethodabstractFog removal from an image is an active research topic in computer vision. However, current literature is weak in the following two areas which in many ways are hindering progress for developing defogging algorithms. First, there is no true real-world and naturally occurring foggy image datasets suitable for developing defogging models. Second, there is no suitable mathematically simple and easy to use image quality assessment (IQA) methods for evaluating the visual quality of defogged images. We address these two aspects in this paper. We first introduce a new foggy image dataset called multiple real-world foggy image dataset (MRFID). MRFID contains foggy and clear images of 200 outdoor scenes. For each scene, one clear image and 4 foggy images of different densities defined as slightly foggy, moderately foggy, highly foggy, and extremely foggy, are manually selected from images taken from these scenes over the course of one calendar year. We then process the foggy images of MRFID using 16 defogging methods to obtain 12,800 defogged images (DFIs) and perform a comprehensive subjective evaluation of the visual quality of the DFIs. Through collecting the mean opinion score (MOS) of 120 subjects and evaluating a variety of fog-relevant image features, we have developed a new Fog-relevant Feature based SIMilarity index (FRFSIM) for assessing the visual quality of DFIs. We present extensive experimental results to show that our new visual quality assessment measure, the FRFSIM, is more consistent with the MOS than other IQA methods and is therefore more suitable for evaluating defogged images than other state-of-the-art IQA methods. Our dataset and relevant code are available at http://www.vistalab.ac.cn/MRFID-for-defogging/. Wei Liu 0123, Fei Zhou 0001, Tao Lu 0001, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 3 |
| 2020 | Masked Face Recognition with Identification AssociationabstractIn the crime scene, criminals often consciously conceal their facial identity through face-masked disguise, which poses a huge challenge to identity recognition. Existing disguised face recognition techniques aiming for light even slight occlusions are completely invalid for face-masked identification. To this end, this paper proposes a masked face recognition method based on person re-identification association, which converts the masked face recognition problem into an association uncovering problem between the masked face and the appearing faces of the same person. Based on the characteristics that person re-identification technique does not rely solely on facial information, it first takes advantages of re-identification to establish the association between face-masked pedestrians and face-unveiled pedestrians. It further provides an effective face image quality assessment to select the most identifiable faces for subsequent recognition from a variety of appearing candidate faces. Finally, the selected high-quality recognizable faces are used to replace masked faces for identification. The comparison experiments with the existing disguise face recognition methods show its superiority in terms of accuracy. Zhongyuan Wang 0001, Zheng He 0001, Nanxi Wang, Xin Tian 0006, Tao Lu 0001 |
ICTAI | 6 |
| 2020 | Face Super-Resolution by Learning Multi-view Texture Compensation
Yu Wang 0140, Tao Lu 0001, Ruobo Xu, Yanduo Zhang |
MMM (2) | 2 |
| 2020 | Global-local fusion network for face super-resolution
Tao Lu 0001, Jiaming Wang 0001, Junjun Jiang, Yanduo Zhang |
Neurocomputing | 1 |
| 2020 | Face super-resolution via nonlinear adaptive representation
Tao Lu 0001, Kangli Zeng, Shenming Qu, Yanduo Zhang |
Neural Comput. Appl. | 1 |
| 2019 | GAN-Based Multi-level Mapping Network for Satellite Imagery Super-ResolutionabstractAlthough many deep-learning-based image super-resolution (SR) methods have been proposed, most of them assume that all hierarchical features share the unified mapping equations. They ignore the differences between mapping equations at different feature levels, and create an average effect of mapping prediction, thus poorly building the mapping relations between low resolution (LR) and high resolution (HR) spaces. In this paper, we propose a multi-level mapping framework along with the adversarial learning strategy, namely MMGAN, for satellite imageries SR reconstruction. We also construct a feature extraction and tuning block (FETB) for fine feature expression. In particular, a novel two-dimension dense unit (DU) and a mapping attention unit (MAU) are constructed for building multi-level mappings in different stages. With our strategies, an HR image is reconstructed directly from the input image using multi-level mappings. Extensive experiments on Kaggle Open Source Dataset and Jilin-1 video satellite images exhibit superior reconstruction performance when compared with the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang, Guangcheng Wang, Zhen Han 0002, Tao Lu 0001 |
ICME | 7 |
| 2019 | Uav Image Mosaic Based on Non-Rigid Matching and Bundle AdjustmentabstractThis study introduces a robust method for panoramic unmanned aerial vehicle (UAV) image mosaic. The traditional automatic panoramic image stitching method requires the camera to carefully rotate the optical center to obtain an image, but the image used for mosaic in reality cannot easily achieve this ideal state. In particular, remote sensing images obtained by UAVs do not satisfy such a situation. The images may not be on a plane yet, and several of them may even have non-rigid changes. Therefore, the classical method of UAV image stitching is expected to produce poor results. To this end, we improve the traditional stitching method to overcome the abovementioned challenges. Specifically, a non-rigid matching algorithm is introduced to the system to make it suitable for remote sensing images. We perform bundle adjustments using a new strategy to make the mosaic system suitable for UAV images. Experimental results show that our method is more robust than the traditional method. Linbo Luo 0002, Jun Chen 0019, Tao Lu 0001, Yong Wang 0036 |
IGARSS | 4 |
| 2019 | Identifying Users by Asynchronous Mobility TrajectoriesabstractWith the popularity of location-based services and applications, a large amount of mobility data has been generated. Identity recognition through mobile trajectory information, especially asynchronous trajectory data has arisen great concerns in social security prevention and control. This paper advocates an identification resolution method based on the most frequently distributed TOP-N regions regarding user trajectories. This method first finds TOP-N regions whose trajectory points are most frequently distributed so as to reduce the computational complexity. It then combines probabilistic deviation and angle cosine to calculate TOP-N region similarity between two trajectories to identify the same user. We conducted extensive experiments on two real GPS trajectory datasets GeoLife and Cabspotting and comprehensively discussed the experimental results. The experimental results show that this method is substantially effective and efficiency for user identification. Mengjun Qi, Zhongyuan Wang 0001, Zheng He 0001, Tao Lu 0001 |
IGARSS | 4 |
| 2019 | Face super-resolution via bilayer contextual representation
Kangli Zeng, Tao Lu 0001, Xuefeng Liang, Kai Li 0005, Yanduo Zhang |
Signal Process. Image Commun. | 2 |
| 2019 | Edge-Enhanced GAN for Remote Sensing Image SuperresolutionabstractThe current superresolution (SR) methods based on deep learning have shown remarkable comparative advantages but remain unsatisfactory in recovering the high-frequency edge details of the images in noise-contaminated imaging conditions, e.g., remote sensing satellite imaging. In this paper, we propose a generative adversarial network (GAN)-based edge-enhancement network (EEGAN) for robust satellite image SR reconstruction along with the adversarial learning strategy that is insensitive to noise. In particular, EEGAN consists of two main subnetworks: an ultradense subnetwork (UDSN) and an edge-enhancement subnetwork (EESN). In UDSN, a group of 2-D dense blocks is assembled for feature extraction and to obtain an intermediate high-resolution result that looks sharp but is eroded with artifacts and noises as previous GAN-based methods do. Then, EESN is constructed to extract and enhance the image contours by purifying the noise-contaminated components with mask processing. The recovered intermediate image and enhanced edges can be combined to generate the result that enjoys high credibility and clear contents. Extensive experiments on Kaggle Open Source Data set, Jilin-1 video satellite images, and Digitalglobe show superior reconstruction performance compared to the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Tao Lu 0001, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Multi-Memory Convolutional Neural Network for Video Super-ResolutionabstractVideo super-resolution (SR) is focused on reconstructing high-resolution (HR) frames from consecutive lowresolution (LR) frames. Most previous video SR methods based on convolutional neural network (CNN) use a direct connection and single-memory module within the network, and they thus fail to make full use of spatio-temporal complementary information from LR observed frames. To fully exploit spatio-temporal correlations between adjacent LR frames and reveal more realistic details, this paper proposes a multi-memory convolutional neural network (MMCNN) for video SR, cascading an optical flow network and an image-reconstruction network. A serial of residual blocks engaged in utilizing intra-frame spatial correlations are proposed for feature extraction and reconstruction. Particularly, instead of using single-memory module, we embed convolutional long short-term memory (ConvLSTM) into the residual block, thus form a multi-memory residual block to progressively extract and retain inter-frame temporal correlations between consecutive LR frames. We conduct extensive experiments on numerous testing datasets with respect to different scaling factors. Our proposed MMCNN shows superiority over the state-of-the-art methods in terms of PSNR and visual quality and surpasses the best counterpart method 1 dB at most. The code and datasets are available at https://github.com/psychopa4/MMCNN. Zhongyuan Wang 0001, Peng Yi 0002, Kui Jiang, Junjun Jiang, Zhen Han 0002, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | SMIM: Superpixel Mutual Information Measurement for Image Quality Assessment
Jiaming Wang 0001, Tao Lu 0001, Yanduo Zhang |
ICA3PP (2) | 2 |
| 2018 | Contextual-Field Supported Iterative Representation for Face Hallucination
Kangli Zeng, Tao Lu 0001, Yanduo Zhang, Li Peng 0003, Shenming Qu |
ICA3PP (3) | 2 |
| 2018 | Facial Shape and Expression Transfer via Non-rigid Image Deformation
Huabing Zhou, Shiqiang Ren, Yuyu Kuang, Yanduo Zhang, Wei Zhang 0259, Tao Lu 0001, Hanwen Chen, Deng Chen |
ICA3PP (3) | 7 |
| 2018 | Feature Matching Based on Top K Rank SimilarityabstractFeature matching plays a key component in many computer vision and pattern recognition tasks. Observing that the spatial neighborhood relationship (representing the topological structures of an image scene) is generally well preserved between two feature points of an image pair, some mismatch removing methods based on maintaining the local neighborhood structures of the potential true matches have been proposed. How to define the local neighborhood structure is an issue of vital importance. In this paper, we propose a robust and efficient method, called Top$K$Rank Preservation (Top-KRP), for mismatch removal from given putative point set matching correspondences. Instead of preserving the intersection of neighbors, TopKRP aims at preserving the top$K$rank of two feature points. The developed approach is validated on numerous challenging real image pairs for general feature matching, and the experimental results demonstrate that it outperforms several state-of-the-art feature matching methods, especially in case of a large number of mismatches. Junjun Jiang, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
ICASSP | 3 |
| 2018 | Face Hallucination Using Manifold-Regularized Group Locality-Constrained RepresentationabstractSparsity and locality regularizations are successfully applied to face hallucination algorithms to ameliorate their ill-posed nature. However, most of patch-based face hallucination approaches only consider the manifold structure of single patch, thus resulting in unstable solution for image reconstruction. In this paper, we propose a novel face hallucination, termed manifold-regularized group locality-constrained representation (MGLR), in order to exploit the multiple manifold structures rooted in grouped self-similarly patches. Specifically, we first group similar patches to form a matrix which contains the recurrent non-local patches. Then graph regularization term is formulated to represent the group manifolds for better reconstruction quality. Taking advantages of grouped self-similar patches, MGLR can offer stable sparse solution to take advantage of the the accurate prior for super-resolution reconstruction. Experimental results on LFW database and CMU real-world images demonstrate the superiority of the proposed method over some state-of-the-art face methods both in terms of subjective and objective qualities. Tao Lu 0001, Kangli Zeng, Junjun Jiang, Yanduo Zhang, Zhongyuan Wang 0001, Huabing Zhou |
ICIP | 1 |
| 2018 | Robust and efficient face recognition via low-rank supported extreme learning machine
Tao Lu 0001, Yingjie Guan, Yanduo Zhang, Shenming Qu, Zixiang Xiong |
Multim. Tools Appl. | 1 |
| 2017 | Face hallucination using region-based deep convolutional networksabstractMost deep learning based face hallucinations exploit random patch prior from training samples, then to learn the mapping functions between low-resolution (LR) and high-resolution (HR) images, and achieve satisfactory reconstruction performance. However, most of them do not take into account the prior information on facial structure, which is pivotal for face hallucination. Different from random patch prior based deep learning approaches, in this paper, we utilize facial structural prior and develop a simple yet powerful face hallucination, named region-based deep convolutional networks (RDCN). Firstly, we divide facial image into several regions of interest, then to train multiple parallel subnetworks of these regions for exacting better structure priors, finally HR output is reconstructed by stitching facial parts. Experiments on the FEI database demonstrate that the proposed region-based convolution networks outperform other state-of-the-art, including recently proposed deep learning based approaches, both in subjective and objective reconstruction qualities. Tao Lu 0001, Hao Wang 0237, Zixiang Xiong, Junjun Jiang, Yanduo Zhang, Huabing Zhou, Zhongyuan Wang 0001 |
ICIP | 1 |
| 2017 | Non-rigid image deformation algorithm based on MRLS-TPSabstractIn this paper, we propose a novel closed-form transformation estimation method based on moving regularized least squares optimization with thin-plate spline (MRLS-TPS) for non-rigid image deformation. The method takes the user-controlled point-offset-vectors as the input data, and estimates the spatial transformation about the two control point sets for each pixel. To achieve a realistic deformation, we formulates the transformation estimation as a vector-field interpolation problem by a moving regularized least squares method. Unlike MLS, the mapping function is modeled by a non-rigid function thin-plate spline with regularization technique, such that the deformation can satisfy both global linear affine motion and local non-rigid warping. We derive a closed-form solution of the transformation and achieve a fast implementation. In addition, the proposed method can give a wonderful user experience, fast and convenient manipulating. Extensive experiments on real images demonstrated the proposed method outperforms other state-of-the-art methods and the commercial software Adobe PhotoShop CS 6, especially in case of flexible object motion. Huabing Zhou, Yuyu Kuang, Zhenghong Yu, Shiqiang Ren, Anna Dai, Yanduo Zhang, Tao Lu 0001, Jiayi Ma 0001 |
ICIP | 7 |
| 2017 | DLML: Deep linear mappings learning for face super-resolution with nonlocal-patchabstractLearning-based face super-resolution approaches rely on representative dictionary as self-similarity prior from training samples to estimate the relationship between the low-resolution (LR) and high-resolution (HR) image patches. The most popular approaches, learn mapping function directly from LR patches to HR ones but neglects the multi-layered nature of image degradation process (resolution down-sampling) which means observed LR images are gradually formed from HR version to lower resolution ones. In this paper, we present a novel deep linear mappings learning framework for face super-resolution to learn the complex relationship between LR features and HR ones by alternately updating multi-layered embedding dictionaries and linear mapping matrices instead of directly mapping. Furthermore, in contrast to existing position based studies that only use local patch for self-similarity prior, we develop a feature-induced nonlocal dictionary pair embedding method to support hierarchical multiple linear mappings learning. With coarse-to-fine nature of deep learning architecture, cascaded incremental linear mappings matrices can be used to exploit the complex relationship between LR and HR images. Experimental results demonstrate that such framework outperforms state-of-the-art (including both general super-resolution approaches and face super-resolution approaches) on FEI face database. Tao Lu 0001, Lanlan Pan, Junjun Jiang, Yanduo Zhang, Zixiang Xiong |
ICME | 1 |
| 2017 | Face hallucination using deep collaborative representation for local and non-local patchesabstractPatch-based face hallucination algorithms utilize either local patches (e.g., position-patch approaches) or nonlocal patches (e.g., dictionary-learning approaches) to exploit self-similarity prior from training samples. Although they yield decent results, solo source patches limit their performance due to not fully taking self-similarity prior from both local and nonlocal ones. In order to overcome this shortcoming, we propose a novel and efficient deep collaborative representation (DCR) based approach, to exploit both local and nonlocal self-similarity patches, for boosting face hallucination performance. First we learn a feature-inducing dictionary pair to represent local and nonlocal self-similarity prior, then deep (multiple-layer) representation weights and corresponding support dictionaries are iteratively updated to exploit accurate prior from coarse to fine. Finally, the high resolution (HR) output are optimized layer by layer. Experimental results outperform some state-of-the-art (e.g. Convolutional Neural Network based deep learning approach) which verify the validity of the proposed approach. Tao Lu 0001, Lanlan Pan, Hao Wang 0237, Yanduo Zhang, Zixiang Xiong |
ISCAS | 1 |
| 2017 | Low-rank constrained collaborative representation for robust face recognitionabstractRecently, sparse representation based classifiers (SRC) and collaborative representation based classifiers (CRC) have been shown to give very good performance under controlled scenarios. However, in practical applications, face recognition often encounters variations in illumination, expression, noise and occlusion, which cause severe performance degradation (due to the outliers in testing). In this paper, we present a novel robust face recognition algorithm based on class-wise low-rank constrained collaborative representations. We impose a low-rank constraint on the representation coefficient matrix to discriminate against outliers. The resulting low-rank constrained collaborative representation based classifier (LCRC) jointly minimizes the class-wise reconstruction error and rank of coefficient matrix. Experiments show that LCRC outperforms popular classifiers such as SRC, CRC, SVM, PROCRC on the AR, CMU PIE and LFW databases. Tao Lu 0001, Yingjie Guan, Deng Chen, Zixiang Xiong |
MMSP | 1 |
| 2017 | Single Image Super-Resolution via Locally Regularized Anchored Neighborhood Regression and Nonlocal MeansabstractThe goal of learning-based image super resolution (SR) is to generate a plausible and visually pleasing high-resolution (HR) image from a given low-resolution (LR) input. The SR problem is severely underconstrained, and it has to rely on examples or some strong image priors to reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e., projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. Encouraged by recent developments in image prior modeling, where the state-of-the-art algorithms are formed with nonlocal self-similarity and local geometry priors, we seek an SR algorithm of similar nature that will incorporate these two priors into the learning from LR space to HR space. The nonlocal self-similarity prior takes advantage of the redundancy of similar patches in natural images, while the local geometry prior of the data space can be used to regularize the modeling of the nonlinear relationship between LR and HR spaces. Based on the above two considerations, we first apply the local geometry prior to regularize the patch representation, and then utilize the nonlocal means filter to improve the super-resolved outcome. Experimental results verify the effectiveness of the proposed algorithm compared with the state-of-the-art SR methods. Junjun Jiang, Chen Chen 0001, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | L1-L1 norms for face super-resolution with mixed Gaussian-impulse noiseabstractIn real world surveillance application, the captured faces are often low resolution (LR) and corrupted by mixed Gaussian-impulse noise during the acquisition and transmission processes. In this paper, we propose an effective patch-based face super-resolution method to reconstruct a high resolution (HR) face image given an LR observation that is corrupted by mixed Gaussian-impulse noise. To represent the corrupted image patches, a sparse regularization combined with an l\ data fitting term is proposed. In the proposed model, both the patch reconstruction term and the regularization term are in the l\ norm form. As a result, the model is called norms. In addition, since image pixels have nonnegative intensities, we further add a nonnegative constraint to the patch representation model. Experimental results demonstrate that the proposed norms based method can achieve superior face super-resolution performance over several state-of-the-art approaches based on the objective results in terms of P-SNR, as well as the visual perceptual quality. Junjun Jiang, Zhongyuan Wang 0001, Chen Chen 0001, Tao Lu 0001 |
ICASSP | 4 |
| 2016 | Face hallucination via locality-constrained low-rank representationabstractFace hallucination (FH) based on sparse representation (SR) and locality-constrained representation (LCR) gives reasonably good performance. However, neither SR-nor LCR-based methods make full use of the structure information in the training data. On the other hand, low-rank representation (LRR) has been utilized to cluster samples into their respective classes by exploiting low-rank structures of the data. In this paper, we propose a locality-constrained low-rank representation (LCLRR) method to take advantage of both LCR and LRR for FH. LCLRR first enforces a low-rank constraint on choosing the dictionary atoms that belong to a subspace that correspond to the same cluster, it then imposes a locality constraint on selecting atoms that are in the vicinity of test samples. Experiments show that LCLRR outperforms both SR- and LCR-based methods on subjectively and objectively, proving that exploiting the structure information in the training data is feasible in face hallucination. Tao Lu 0001, Zixiang Xiong, Yongjing Wan |
ICASSP | 1 |
| 2016 | Very Low-Resolution Face Recognition via Semi-Coupled Locality-Constrained RepresentationabstractRecognition tasks in very low-resolution (VLR) images are more challenging than those in high-resolution (HR) due to lack of adequate discriminative information. Previous VLR and HR coupled learning scheme limits both the representation and discriminative ability of features. In this work, we propose a semi-coupled locality-constrained representation (SLR) approach to learn the discriminative representations and the mapping relationship between VLR and HR features simultaneously. Both VLR and HR local manifold geometries are coded during representation, while the learned mapping function improves the manifold consistency by transforming VLR features to HR ones. Finally, the resolutionrobust features are fed into a sparse representation based classifier (SRC) to predict the face labels. The proposed algorithm gives better performance than many state-of-the-art VLR recognition algorithms. Tao Lu 0001, Yanduo Zhang, Zixiang Xiong |
ICPADS | 1 |
| 2016 | Adaptive boosting for image denoising: Beyond low-rank representation and sparse codingabstractIn the past decade, much progress has been made in image denoising due to the use of low-rank representation and sparse coding. In the meanwhile, state-of-the-art algorithms also rely on an iteration step to boost the denoising performance. However, the boosting step is fixed or non-adaptive. In this work, we perform rank-1 based fixed-point analysis, then, guided by our analysis, we develop the first adaptive boosting (AB) algorithm, whose convergence is guaranteed. Preliminary results on the same image dataset show that AB uniformly outperforms existing denoising algorithms on every image and at each noise level, with more gains at higher noise levels. Tao Lu 0001, Zixiang Xiong |
ICPR | 2 |
| 2016 | Efficient low-rank supported extreme learning machine for robust face recognitionabstractRecently, deep learning based face recognition algorithms have achieved great success in recognition performance. However, designing and training complex learning models suffer from time and labor efficiency. In this paper, we propose a novel three-layer low-rank supported extreme learning machine (LSELM) algorithm to take advantage of both robust feature representation and fast classification for efficient recognition. Every given probe sample is first clustered into a sub-class spanned by linear representation. With this sub-class, low-rank and robust features that are insensitive to disguise, noise, variant expression or illumination are recovered. These discriminative features are then coded to support a forward neural network for efficient prediction. Experimental results show that LSELM is on par with other deep learning based face recognition algorithms in recognition performance but has less time complexity on both AR and extend Yale-B datasets. Yingjie Guan, Tao Lu 0001, Yanduo Zhang, Zixiang Xiong |
VCIP | 2 |
| 2016 | Smooth sparse representation for noise robust face super-resolutionabstractFace super-resolution has attracted much attention in recent years. Many algorithms have been proposed. Among them, sparse representation based face super-resolution approaches are able to achieve competitive performance. However, these sparse representation based approaches only perform well under the condition that the input is noiseless or has small noise. When the input is corrupted by large noise, the reconstruction weights of the input LR patches using sparse representation based approaches will be seriously unstable, thus leading to poor reconstruction results. To this end, in this paper, we propose a novel sparse representation based face super-resolution approach that incorporates a smooth prior to enforce similar training patches having similar sparse coding coefficients. Specifically, we introduce the fused Lasso to the least squares representation of the input LR image in order to obtain a stable sparse representation, especially when the noise level of the input LR image is high. Experiments are carried out on the benchmark FEI face dataset. Visual and quantitative comparisons show that the proposed face super-resolution method achieves comparable performance to the state-of-the-art methods under noiseless condition, and yields superior super-resolution results when the input LR face image is contaminated by strong noise. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Zhongyuan Wang 0001, Tao Lu 0001 |
VCIP | 5 |
| 2015 | Locally regularized Anchored Neighborhood Regression for fast Super-ResolutionabstractThe goal of learning-based image Super-Resolution (SR) is to generate a plausible and visually pleasing High-Resolution (HR) image from a given Low-Resolution (LR) input. The problem is dramatically under-constrained, which relies on examples or some strong image priors to better reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e. projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. One recently proposed method, Anchored Neighborhood Regression (ANR) [1], provides state-of-the-art quality performance and is very fast. In this paper, we propose an improved variant of ANR, namely Locally regularized Anchored Neighborhood Regression (LANR), which utilizes the locality-constrained regression in place of the ridge regression in ANR. LANR assigns different freedom for each neighbor dictionary atom according to its correlation to the input LR patch, thus the learned projection matrices are much more flexible. Experimental results demonstrate that the proposed algorithm performs efficiently and effectively over state-of-the-art methods, e.g., 0.1–0.4 dB in term of PSNR better than ANR. Junjun Jiang, Jican Fu, Tao Lu 0001, Ruimin Hu, Zhongyuan Wang 0001 |
ICME | 3 |
| 2015 | Trilateral constrained sparse representation for Kinect depth hole filling
Zhongyuan Wang 0001, Shizheng Wang, Tao Lu 0001 |
Pattern Recognit. Lett. | 4 |
| 2014 | Efficient single image super-resolution via graph-constrained least squares regression
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001 |
Multim. Tools Appl. | 4 |
| 2013 | Coupled-layer neighbor embedding for surveillance face hallucinationabstractAs the face image captured by a surveillance camera is typically very low-resolution (LR), blurred and noisy, traditional neighbor embedding method considers only one manifold (the LR image manifold) and fails very often to reliably estimate the intention geometrical structure. In this paper, we introduce the notion of neighbor embedding from the LR image manifold and the high-resolution (HR) one simultaneously and propose a novel neighbor embedding model, termed the coupled-layer neighbor embedding (CLNE), for surveillance face hallucination. CLNE differs substantially from other neighbor embedding models in that the former has two layers: the LR layer and the the HR layer. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights; the HR layer in this model is a set of HR training patches that guide the K-nearest neighbor (K-NN) searching and geometrically constrain the reconstruction weights. By this coupled constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust neighbor embedding through the significant degradation process. Indeed, the experimental results confirm that our method outperforms the related state-of-the-art methods by having better objective values as well as better visual results. Junjun Jiang, Ruimin Hu, Liang Chen 0026, Zhen Han 0002, Tao Lu 0001, Jun Chen 0001 |
ICIP | 5 |
| 2013 | Locality-constraint iterative neighbor embedding for face hallucinationabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the neighbor embedding based super-resolution algorithms try to preserve the local geometry of the patch manifold for the reconstructed HR patch manifold. However, due to “one-to-many” mappings between LR and HR images, the neighborhood relationship of the LR patch manifold can't reflect the inherent data structure. In this paper, we explore the data structure by both considering the LR patch and HR patch manifolds instead of only considering one manifold (LR patch manifold). By incorporating the position prior of face and local geometry of HR patch manifold, we propose an improved neighbor embedding method to face hallucination, namely locality-constraint iterative neighbor embedding (LINE), in which we iteratively update the K-nearest neighbors (K-NN) and reconstruction weights based on the result (the hallucinated HR patch) from previous iteration, giving rise to improved performance compared with traditional neighbor embedding algorithms. Experimental results with application to face hallucination on simulated LR face images and real world ones demonstrate the effectiveness of the proposed method. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Tao Lu 0001, Jun Chen 0001 |
ICME | 5 |
| 2013 | Robust super-resolution for face images via principle component sparse representation and least squares regressionabstractFace image super-resolution (SR) reconstruction is the problem of inducing a high-resolution (HR) face image from a low-resolution (LR) one. Traditional face SR methods are either sensitive to noise, i.e., local patch based technologies, or lacking facial details, i.e., global face reconstruction, thus could not achieve a satisfying result. In order to overcome these problems, we propose in this paper a novel face SR method. Taking full advantages of Principle Component analysis and Sparse Representation (PCSR), it aims to obtain an accurate and noise robust representation, transforming the image patch to the principle component sparse feature space (PC-SFS). Moreover, in PC-SFS, we try to learn a mapping function between the LR image patches and HR ones through Least Squares Regression. Given a LR patch, we first transform it to the LR PC-SFS by PCSR to obtain the robust and accurate representation, and then project the representation to the HR PC-SFS thus get the target HR patch. Experiments on the frontal faces SR in noise conditions demonstrate our method outperforms state of the art. Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang |
ISCAS | 1 |
| 2013 | From local representation to global face hallucination: A novel super-resolution method by nonnegative feature transformationabstractMost of global face hallucination methods treat the face as a whole, ignoring the fact that the face is composed by part-based organs. Therefore, the results obtained by these methods always lack of detailed information. Nonnegative matrix factorization (NMF) based face hallucination method is properly used to enhance the detailed information. Usually, NMF basis is only learnt from high-resolution (HR) samples, leading to over-smooth output and lack of high frequency details. In order to solve this problem, we propose a simple but novel face hallucination method using nonnegative feature transformation by two-step framework. In particular, we learn the NMF basis from low-resolution (LR) and HR samples separately, and then transform the local representation feature of input into the global representation subspaces, keeping the weights into the HR samples space for output. Furthermore, the maximum a posteriori (MAP) method is used to estimate a better output. Experiments show that the hallucinated face of the proposed method is not only more high-frequency details, but also has better performance than many state-of-art algorithms. Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang, Yanduo Zhang |
VCIP | 1 |
| 2012 | A super-resolution method for low-quality face image through RBF-PLS regression and neighbor embeddingabstractIn this paper, a new two-step method is proposed to infer a high-quality and high-resolution (HR) face image from a low-quality and low-resolution (LR) observation based on training samples in the database. First, a global face image is reconstructed based on the non-linear relationship between LR and HR face images, which is established according to radial basis function and partial least squares (RBF-PLS) regression. Based on the reconstructed global face patches manifold (formed by the image patches at the same position of all global face images), whose local geometry is more consistent with that of original HR face patches manifold than noisy LR one is, the Neighbor Embedding is applied to induce the target HR face image by preserving the similar local geometry between global face patches manifold and the original HR face patches manifold. A comparison of some state-of-the-art methods shows the superiority of our method, and experiments also demonstrate the effectiveness both under simulation and real conditions. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ICASSP | 4 |
| 2012 | Graph discriminant analysis on multi-manifold (GDAMM): A novel super-resolution method for face recognitionabstractHow to efficiently recognize low-resolution (LR) probe images of one face recognition system, in which high-resolution (HR) gallery of faces is enrolled, is still an open problem. In this paper, we develop a novel super-resolution method, namely Graph Discriminant Analysis on Multi-Manifold (GDAMM), to super-resolved the HR version of a LR probe image and then perform matching at the resolution of the HR gallery. Unlike classical super-resolution approaches considering only the data fidelity, GDAMM takes the advantages of both manifold learning and discriminant analysis to integrate the data constraint and discriminant constraint, seeking the mapping between LR images and HR ones. In the reconstructed HR image space, faces of one person in the same manifold are close and those in different manifolds are far apart. Experiments on Extended Yale-B database and AR face database demonstrate that the learned discriminant information is essential for improving recognition accuracy. Through the contrastive experiment, the results (recognition rates) indicate that the proposed GDAMM method can greatly surpass classical super-resolution approaches, even outperforming the ideal case of having probe images of HR gallery by a big margin (nearly 9% on Extended Yale-B database and 8% on AR face database). Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001 |
ICIP | 5 |
| 2012 | Efficient Single Image Super-Resolution via Graph EmbeddingabstractWe explore in this paper efficient algorithmic solutions to single image super-resolution (SR). We propose the GESR, namely Graph Embedding Super-Resolution, to super-resolve a high-resolution (HR) image from a single low-resolution (LR) observation. The basic idea of GESR is to learn a projection matrix mapping the LR image patch to the HR image patch space while preserving the intrinsic geometrical structure of original HR image patch manifold. While GESR resembles other manifold learning-based SR methods in persevering the local geometric structure of HR and LR image patch manifold, the innovation of GESR lies in that it preserves the intrinsic geometrical structure of original HR image patch manifold rather than LR image patch manifold, which may be contaminated because of image degeneration (e.g., blurring, down-sampling and noise). Experiments on benchmark test images show that GESR can achieve very competitive performance as Neighbor Embedding based SR (NESR) and Sparse representation based SR (SSR). Beyond subjective and objective evaluation, all experiments show that GESR is much faster than both NESR and SSR. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001 |
ICME | 5 |
| 2012 | Position-Patch Based Face Hallucination via Locality-Constrained RepresentationabstractInstead of using probabilistic graph based or manifold learning based models, some approaches based on position-patch have been proposed for face hallucination recently. In order to obtain the optimal weights for face hallucination, they represent image patches through those patches at the same position of training face images by employing least square estimation or convex optimization. However, they can hope neither to provide unbiased solutions nor to satisfy locality conditions, thus the obtained patch representation is not the best. In this paper, a simpler but more effective representation scheme- Locality-constrained Representation (LcR) has been developed, compared with the Least Square Representation (LSR) and Sparse Representation (SR). It imposes a locality constraint onto the least square inversion problem to reach sparsity and locality simultaneously. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art face hallucination approaches. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ICME | 4 |
| 2012 | Face hallucination via K-selection mean constrained sparse representation
Kebin Huang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Junjun Jiang |
ICPR | 4 |
| 2012 | Surveillance face hallucination via variable selection and manifold learningabstractIn this paper, we propose a new two-step face hallucination method to induce a high-resolution (HR) face image from a low-resolution (LR) observation. Especially for low-quality surveillance face image, an RBF-PLS based variable selection method is presented for the reconstruction of global face image. Further more, in order to compensate for the reconstruction errors, which are lost high frequency detailed face features, the Neighbor Embedding (NE) based residue face hallucination algorithm is used. Compared with current methods, the proposed RBF-PLS based method can generate a global face more similar to the original face and less sensitive to noise, moreover, the NE algorithm can reduce the reconstruction errors caused by misalignment on the basis of a carefully designed search strategy. Experiments show the superiority of the proposed method compared with some state-of-the-art approaches and the efficacy both in simulation and real surveillance condition. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ISCAS | 4 |
| 2012 | Face image super-resolution via nearest feature lineabstractIn this paper, we propose a manifold learning based algorithm using 'Nearest Feature Line - NFL' to hallucinate high-resolution face image. According to the fact that existing NFL can effectively characterize the geometrical proportions to the face samples, we propose using NFL metric to define the neighborhood relations between face samples. Our algorithm can solve the problem that traditional method cannot effectively reveal the similar local geometry between high-resolution and low-resolution face manifolds under the condition that the training sample size is small. Moreover, in order to enhance the representation capacity of available face samples and reduce the computational complexity, we select neighborhood samples for each input LR image. Experimental results demonstrate that our algorithm can generates clearer local feature details, and the PSNR is 1.4 dB higher than that of the best manifold learning based method reported so far. Zhen Han 0002, Junjun Jiang, Ruimin Hu, Tao Lu 0001, Kebin Huang |
ACM Multimedia | 4 |
| 2010 | Global Face Super Resolution and Contour Region Constraints
Chengdong Lan, Ruimin Hu, Tao Lu 0001, Ding Luo, Zhen Han 0002 |
ISNN (2) | 3 |