VLDB 2026 Research / reviewers in the wild / expert
Zhen Han 0002
dblp:62/302-2
· DBLP profile ↗
83ranked-venue papers
1as first author
35since 2021 · last 2026
0000-0002-1862-4781ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 1 first-author · 19 since 2021Artificial intelligence and machine learning · 24 · 17 since 2021Systems, architecture and hardware · 4Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and DatasetabstractElectrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its motion patterns, leading to distinctive spatio-temporal characteristics across smoky laparoscopic videos. However, existing desmoking methods fail to account for such smoke-type-specific distinctions. Therefore, we propose the first Smoke-Type-Aware Laparoscopic Video Desmoking Network (STANet) by introducing two smoke types: Diffusion Smoke and Ambient Smoke. Specifically, a smoke mask segmentation sub-network is designed to jointly conduct smoke mask and smoke type predictions based on the attention-weighted mask aggregation, while a smokeless video reconstruction sub-network is proposed to perform specially desmoking on smoky features guided by two types of smoke mask. To address the entanglement challenges of two smoke types, we further embed a coarse-to-fine disentanglement module into the mask segmentation sub-network, which yields more accurate disentangled masks through the smoke-type-aware cross attention between non-entangled and entangled regions. In addition, we also construct the first large-scale synthetic video desmoking dataset with smoke type annotations. Extensive experiments demonstrate that our method not only outperforms state-of-the-art approaches in quality evaluations, but also exhibits superior generalization across multiple downstream surgical tasks. Qifan Liang, Zhen Han 0002, Xihao Wang, Zhongyuan Wang 0001, Bin Mei |
AAAI | 3 |
| 2026 | An angle-guided bidirectional feature transformation network for multi-frame tilt-angle face recognition
Wenqin Song, Xihao Wang, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | IDRetracor: Towards Visual Forensics against Malicious Face SwappingabstractThe deepfake-based face swapping technique poses significant risks to personal identity security. Although many detection methods have been proposed to counter malicious face swapping, they typically provide only binary labels (Fake/Real), lacking reliable and interpretable evidence. To address this limitation, we introduce a novel task called face retracing, which aims to visually trace back the original target face from a given fake one through inverse mapping. This task is based on the observation that current face swapping methods are neither flawless nor entirely random, leaving recoverable traces of the original identity. To this end, we propose IDRetracor, a model designed to recover arbitrary original target identities from fake faces generated by various face swapping techniques. Specifically, we first employ a mapping resolver to estimate the possible solution space of the original face for inverse mapping. Then, we introduce Mapping-Aware Convolutions (MACs), which consist of multiple dynamically combined kernels guided by the mapping resolver to adaptively handle diverse face swapping patterns. Extensive experiments demonstrate that IDRetracor achieves strong performance in retracing original faces, validated by both quantitative metrics and qualitative assessments. Jikang Cheng, Jiaxin Ai, Zhen Han 0002, Chao Liang 0001, Qin Zou 0001, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Link-based Contrastive Learning for One-Shot Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) transfers knowledge from a labeled source domain to an unlabeled target domain via distribution alignment. However, in real-world scenarios like public safety or access control, obtaining sufficient source data is challenging, limiting existing UDA methods. This paper investigates a realistic but rarely studied problem called one-shot unsupervised domain adaptation (OSUDA), where only one source example per category is available. OSUDA faces dual challenges in feature learning and domain alignment due to the extreme source data scarcity. To address these, we propose link-based contrastive learning (LCL), a simple yet effective approach for OSUDA. LCL leverages in-domain links to learn discriminative features from abundant unlabeled target data and cross-domain links to achieve precise domain alignment with only one source sample per category. Extensive experiments on four domain adaptation benchmarks (VisDA-2017, Office-31, Office-Home, and DomainNet) demonstrate LCL’s effectiveness under the OSUDA setting. Additionally, we construct a real-world OSUDA surveillance face recognition dataset, where LCL consistently improves recognition performance across various face recognition methods. Yue Zhang 0104, Mingyue Bin, Zhongyuan Wang 0001, Zhen Han 0002, Chao Liang 0001 |
CVPR | 5 |
| 2025 | Multimodal Re-Ranking for Heterogeneous Face Re-IdentificationabstractHeterogeneous face re-identification (Re-ID), aiming to match low-quality faces captured by disjoint visible light (VIS) and near-infrared (NIR) cameras, has become a critical application in video surveillance. However, the domain discrepancy between the NIR-VIS faces degrades the Re-ID performance. To solve this problem, this paper proposes a multimodal re-ranking method including two stages. Firstly, we utilize the VIS-NIR face bi-directional modality transformation based on the positive and negative samples separate training strategy to reduce domain discrepancy and generate the multimodal ranking lists of face Re-ID with complementarities. Secondly, we propose linear and nonlinear multimodal ranking lists fusion strategies based on single-modal and multi-modal k-reciprocal nearest neighbors (K-RNNs) to obtain a more accurate fused ranking list for face Re-ID. Extensive experiments on heterogeneous face datasets demonstrate the superior performance of our method over existing methods. Wenqin Song, Jiawei Zhang 0002, Zhen Han 0002, Yunfeng Xue, Xihao Wang, Zhongyuan Wang 0001 |
ICIP | 4 |
| 2025 | Adversarial intensity awareness for robust object detection
Jikang Cheng, Baojin Huang, Zhen Han 0002, Zhongyuan Wang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Image deraining via dual-level contextual information associated learning for autonomous driving
Bin Yang 0026, Zhen Han 0002, Zheng Wang 0007 |
Knowl. Based Syst. | 4 |
| 2025 | Multi-Stage Statistical Texture-Guided GAN for Tilted Face FrontalizationabstractExisting pose-invariant face recognition mainly focuses on frontal or profile, whereas high-pitch angle face recognition, prevalent under surveillance videos, has yet to be investigated. More importantly, tilted faces significantly differ from frontal or profile faces in the potential feature space due to self-occlusion, thus seriously affecting key feature extraction for face recognition. In this paper, we asymptotically reshape challenging high-pitch angle faces into a series of small-angle approximate frontal faces and exploit a statistical approach to learn texture features to ensure accurate facial component generation. In particular, we design a statistical texture-guided GAN for tilted face frontalization (STG-GAN) consisting of three main components. First, the face encoder extracts shallow features, followed by the face statistical texture modeling module that learns multi-scale face texture features based on the statistical distributions of the shallow features. Then, the face decoder performs feature deformation guided by the face statistical texture features while highlighting the pose-invariant face discriminative information. With the addition of multi-scale content loss, identity loss and adversarial loss, we further develop a pose contrastive loss of potential spatial features to constrain pose consistency and make its face frontalization process more reliable. On this basis, we propose a divide-and-conquer strategy, using STG-GAN to progressively synthesize faces with small pitch angles in multiple stages to achieve frontalization gradually. A unified end-to-end training across multiple stages facilitates the generation of numerous intermediate results to achieve a reasonable approximation of the ground truth. Extensive qualitative and quantitative experiments on multiple-face datasets demonstrate the superiority of our approach. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Chao Liang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | DiffusionMOT: A Diffusion-Based Multiple Object TrackerabstractRecently, researchers have introduced diffusion models into multiple object tracking (MOT) tasks. However, existing diffusion-based MOT methods, such as DiffusionTrack, have significant limitations, including frequent ID switching, reduced performance when tracking nonlinear motion objects, and long inference time. To this end, we propose a more effective diffusion-based multiple object tracker named DiffusionMOT. In particular, we propose a mixed intersection over union (IoU) and Re-Identification (ReID) method for trajectory matching, which effectively reduces incorrect matches. Meanwhile, we propose a secondary calibration method for trajectory boxes, improving the accuracy of the generated detection boxes. Moreover, we introduce the parallel sampling technique from the field of image generation into object tracking and propose a parallel sampling module to enhance the model's inference speed while maintaining tracking accuracy. Furthermore, we design a pair-based two-stage matching (PTM) pipeline to more effectively utilize potential detection information. Extensive experiments on several public MOT benchmarks, including DanceTrack, SportsMOT, MOT20, and MOT17, demonstrate that our approach achieves state-of-the-art (SOTA) performance. The code and models are available at https://github.com/sad123-yx/DiffusionMOT. Yaxuan Hu 0001, Jie Hua 0005, Zhen Han 0002, Hua Zou 0002, Gang Wu 0010, Zhongyuan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Unconventional Face Adversarial Attack
Baojin Huang, Zhen Han 0002, Dengshi Li |
ICANN (2) | 3 |
| 2024 | Dual-Domain Multi-Model GAN Fingerprint Restoration for Compressed Fake Face AttributionabstractRecent advances in GAN fingerprint have shown increasing success in fake face attribution. However, the fake faces are usually compressed during network transmission, which causes the degradation of GAN fingerprint and the decrease of attribution accuracy. To this issue, a dual-domain multi-model GAN fingerprint restoration method for compressed fake face attribution is proposed in this paper. Firstly, considering that image-domain and fingerprint-domain are directly and indirectly affected by compression respectively, we propose a dual-domain parallel restoration architecture that enhances GAN fingerprint using direct image-domain and indirect fingerprint-domain restoration, thereby improving attribution performance by mining the cross-domain complementarity. Secondly, since real and fake GAN-speciffc restoration models can describe GAN fingerprint from different aspects, we first enhance GAN fingerprint by multiple restoration models, and then improve attribution performance by exploiting the cross-model complementarity through the multi-model restoration fusion strategy. Experiments demonstrate the superiority of our method under different compression qualities. Chengxiang Fan, Aohong Shen, Zhen Han 0002, Cai Tong, Zhongyuan Wang 0001, Dekang Yi |
ICME | 3 |
| 2024 | Unlabeled Data Assistant: Improving Mask Robustness for Face RecognitionabstractThe existing masked face recognition algorithms almost tend to adopt synthetic masked face datasets for training. However, these models are limited as they rely on existing mask augmentation methods, which contain few mask patterns and cannot simulate shadows and textures in realistic scenes. To overcome this limitation, we propose a semi-supervised face recognition framework to fully exploit unlabeled real masked face samples, improving the mask robustness of the recognition model. More specifically, unlike the original face embedding network, we design a part-aware network to explore multi-region face representation based on the face structure. In this way, we obtain multiple face sub-embeddings, which correspond to different regions of the face, including the upper half, the lower half and the whole. Crucially, we use the norm of the sub-embedding to represent the activation state of the facial region features. For the input unlabeled masked face image, we restrict the sub-embedding norm of its lower half to weaken the face feature representation of the occluded area. For normal face samples, their partial features are kept activated by maintaining the sub-embedding norm, which guides the deep network does not ignore the available information. Moreover, we employ the margin-based recognition loss for normal samples to ensure that the model is sufficiently discriminative for normal facial features. Extensive experimental results on both normal and real masked face datasets show that our approach significantly outperforms the state-of-the-arts. Code is available at https://github.com/Baojin-Huang/UFace. Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Zhen Han 0002, Chao Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Few-Shot Face Sketch-to-Photo Synthesis via Global-Local Asymmetric Image-to-Image TranslationabstractFace sketch-to-photo synthesis is widely used in law enforcement and digital entertainment, which can be achieved by Image-to-Image (I2I) translation. Traditional I2I translation algorithms usually regard the bidirectional translation of two image domains as two symmetric processes, so the two translation networks adopt the same structure. However, due to the scarcity of face sketches and the abundance of face photos, the sketch-to-photo and photo-to-sketch processes are asymmetric. Considering this issue, we propose a few-shot face sketch-to-photo synthesis model based on asymmetric I2I translation, where the sketch-to-photo process uses a feature-embedded generating network, while the photo-to-sketch process uses a style transfer network. On this basis, a three-stage asymmetric training strategy with style transfer as the trigger is proposed to optimize the proposed model by utilizing the advantage that the style transfer network only needs few-shot face sketches for training. Additionally, we discover that stylistic differences between the global and local sketch faces lead to inconsistencies between the global and local sketch-to-photo processes. Thus, a dual branch of the global face and local face is adopted in the sketch-to-photo synthesis model to learn the specific transformation processes for global structure and local details. Finally, the high-quality synthetic face photo can be generated through the global-local face fusion sub-network. Extensive experimental results demonstrate that the proposed Global-Local Asymmetric (GLAS) I2I translation algorithm compared to SOTA methods, at least improves FSIM by 0.0126, and reduces LPIPS (alex), LPIPS (squeeze), and LPIPS (vgg) by 0.0610, 0.0883, and 0.0719, respectively. Qifan Liang, Zhen Han 0002, Wenjun Mai, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Cross-Modal Face Super-Resolution Based on Quasi-Siamese Domain Transfer Fusion NetworkabstractIn this paper, we propose a Cross-Modal Face Super-Resolution (CMFSR) method to construct high-resolution (HR) facial images from low-resolution (LR) cross-modal facial images captured respectively by disjoint visible light (VIS) and near-infrared (NIR) cameras. Due to the coupling of modality transformation and information fusion, CMFSR is more difficult to obtain HR reconstructed results compared with traditional super-resolution. To solve this problem, a Quasi-Siamese Domain Transfer Fusion Network (QSDTFN) for CMFSR is proposed in this paper, whose two branches transfer two LR face modality to HR face modality by domain transfer respectively. Different from two completely independent branches in the traditional pseudo-siamese network, only the HR-to-LR face transfer processes of the two branches in our quasi-siamese network are independent, while the LR-to-HR face transfer processes are coupled. This coupled module called the Adaptive Weighted Domain Transfer Fusion Module (AWDTFM) disentangles the modality and identity information in the two LR faces, thus achieving modality transformation and identity information fusion simultaneously. In order to strengthen the optimization on the process of CMFSR, this method further introduces the backward QSDTFN to form a higher-level bidirectional structure with the forward QSDTFN, and specifically designs two types of losses: intra-network loss and inter-network loss, to constrain the modality and identity consistencies within one QSDTFN and between two QSDTFNs respectively. The experimental results on the challenging LR cross-modal face datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Jiaxing Wen, Aohong Shen, Zhen Han 0002, Zhongyuan Wang 0001, Liang Chen 0026 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Learning Degradation for Real-World Face Super-Resolution
Jun Chen 0001, Dongshu Xu, Chao Liang 0001, Zhen Han 0002 |
CGI | 6 |
| 2023 | LSA3D: Lightweight Separate Asynchronous 3D Convolutional Neural Network for Gait Recognition
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Jinsheng Xiao, Zhen Han 0002 |
ICANN (10) | 5 |
| 2023 | Multi-frame Tilt-angle Face Recognition Using Fusion Re-ranking
Wenqin Song, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001 |
ICANN (2) | 2 |
| 2023 | Deepfake Face Provenance for Proactive ForensicsabstractMalicious deepfake face not only violates the privacy of personal identities, but also confuses the public and causes huge social harm. The current deepfake detection only stays at the level of distinguishing between true and false, but cannot trace the original genuine face corresponding to the fake face, that is, it does not have the ability to trace the source of evidence. The deepfake countermeasure technology for judicial forensics urgently calls for deepfake inversion. This paper pioneers an interesting question about face deepfake, active forensics that "know what it is and how it happened". Given that deepfake faces do not completely discard the features of original faces, especially facial expressions and poses, we argue that original faces can be approximately speculated from their deepfake counterparts. Correspondingly, we design a disentangling reversing network that decouples latent space features of deepfake faces under the supervision of real-fake face pair samples to infer original faces in reverse. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002, Qin Zou 0001 |
ICIP | 4 |
| 2023 | DeepReversion: Reversely Inferring the Original Face from the DeepFake FaceabstractDeepfake techniques can generate realistic fake images and videos. Malicious fake facial images quickly spread through the Internet, posing a potential threat to personal privacy and judicial forensics. However, the defense methods against deepfake proposed so far mainly focus on the discrimination of authenticity, but cannot identify the true source of the forged face, i.e., the original genuine face corresponding to the face-swapped fake face. This paper poses an interesting issue for face deepfake, which is the proactive forensics of “knowing what and knowing how”. In view of the fact that the fake face exhibits high similarity with the original face, especially the facial expression and pose, we argue that the original face can be approximately estimated from the deepfake counterpart. Accordingly, we advocate a deep-learning-based face inversion approach, so-called DeepReversion, which learns the inverse mapping from the deepfake face to the original face. Based on UNet, we design a specific end-to-end DeepReversion network, and conduct comprehensive experiments on public deepfake datasets. The experimental results show that the speculated face is highly consistent with the original face in terms of visual effects, PSNR, SSIM and similarity given by face recognizers. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002 |
IJCNN | 4 |
| 2023 | PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001 |
Pattern Recognit. | 5 |
| 2023 | HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002 |
Pattern Recognit. | 6 |
| 2023 | Implicit space pose consistent transfer network for deep face verification
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Zhen Han 0002 |
Pattern Recognit. Lett. | 5 |
| 2023 | Stylized image denoising via noise style transfer and Quasi Siamese network
Jikang Cheng, Zhen Han 0002, Zhongyuan Wang 0001 |
Signal Process. Image Commun. | 2 |
| 2023 | Joint Segmentation and Identification Feature Learning for Occlusion Face RecognitionabstractThe existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification. Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Multi-Scale Hybrid Fusion Network for Single Image DerainingabstractDeep learning models have been able to generate rain-free images effectively, but the extension of these methods to complex rain conditions where rain streaks show various blurring degrees, shapes, and densities has remained an open problem. Among the major challenges are the capacity to encode the rain streaks and the sheer difficulty of learning multi-scale context features that preserve both global color coherence and exactness of detail. To address the first problem, we design a non-local fusion module (NFM) and an attention fusion module (AFM), and construct the multi-level pyramids' architecture to explore the local and global correlations of rain information from the rain image pyramid. More specifically, we apply the non-local operation to fully exploit the self-similarity of rain streaks and perform the fusion of multi-scale features along the image pyramid. To address the latter challenge, we additionally design a residual learning branch that is capable of adaptively bridging the gaps (e.g., texture and color information) between the predicted rain-free image and the clean background via a hybrid embedding representation. Extensive results have demonstrated that our proposed method is able to generate much better rain-free images on several benchmark datasets than the state-of-the-art algorithms. Moreover, we conduct the joint evaluation experiments with respect to deraining performance and the detection/segmentation accuracy to further verify the effectiveness of our deraining method for downstream vision tasks/applications. The source code is available at https://github.com/kuihua/MSHFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Guangcheng Wang, Zhen Han 0002, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Local Eyebrow Feature Attention Network for Masked Face RecognitionabstractDuring the COVID-19 coronavirus epidemic, wearing masks has become increasingly popular. Traditional occlusion face recognition algorithms are almost ineffective for such heavy mask occlusion. Therefore, it is urgent to improve the recognition performance of the existing face recognition technology on masked faces. Due to the limited visible feature points of the masked face image relative to the normal face image, we have to exploit the identification potential of eyebrow (referring to eyes and brows) features. This article proposes a local eyebrow feature attention network for masked face recognition, which consists of feature extraction, eyebrow region pooling, and feature fusion. To highlight the eyebrow region, we first use the eyebrow region pooling to separate the local features of eyebrows from the learned overall facial features. We then make full use of the symmetry of left and right eyebrows to emphasize their discriminant ability, due to the inadequate fine information of the low-resolution eyebrows. In particular, in view of the symmetrical similarity between eyebrow pairs and the subordinate relationship between facial components and the whole, we propose a feature fusion model based on graph convolutional network (GCN) to learn the feature association structure of eye features, brow features, and global facial features. We construct the benchmark datasets for masked face recognition to validate our approach, including real-world masked face recognition dataset (RMFRD) and synthetic masked face recognition dataset (SMFRD). Extensive experimental results on both public datasets and our built masked face datasets show that our approach significantly outperforms the state-of-the-arts. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Zhen Han 0002, Kui Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Face Super-Resolution with Better Semantics and More Efficient Guidance
Jun Chen 0001, Zheng Wang 0007, Chao Liang 0001, Zhen Han 0002, Chia-Wen Lin |
CGI | 5 |
| 2022 | Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006 |
Pattern Recognit. | 6 |
| 2022 | Real-Time Video Deraining via Global Motion Compensation and Hybrid Multi-Scale Temporal CorrelationsabstractThe current video deraining algorithms mainly use adjacent frames to optimize the target frame information. However, they only consider the inter-frame temporal correlations of a uniform scale between frames, ignoring the inter-frame temporal correlations of different scales. In addition, the high computational cost is another drawback of the current video deraining algorithms. To this end, we propose a novel aggregation network that explores the inter-frame multi-scale temporal correlations for video deraining with the small computational cost. First, we construct a hybrid multi-scale feature extraction structure in the network to increase the receptive field of multi-scale features. For similar rain streaks at adjacent frames with different scales, a hybrid multi-scale residual block (HMSRB) is proposed to explore the complementary and redundant information at the temporal dimension to characterize the target frame. At the same time, we introduce an improved global context module (GCM) to avoid the complex motion estimation and motion compensation (ME&MC) operation as in previous video deraining approaches, while reducing the calculation complexity. Finally, a fusion block is utilized to adaptively merge the extracted features. Experiments demonstrate that our proposed network is proved to be more efficient and effective than the existing algorithms. Jun Chen 0001, Zhen Han 0002, Qikui Zhu, Weijian Ruan |
IEEE Signal Process. Lett. | 3 |
| 2021 | When Face Recognition Meets Occlusion: A New BenchmarkabstractThe existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001 |
ICASSP | 6 |
| 2021 | A Tilt-Angle Face Dataset And Its ValidationabstractSince the surveillance cameras are usually mounted at a high position to overlook targets, tilt-angle faces on overhead view are common in the public video surveillance environment. Face recognition approaches based on deep learning models have achieved excellent performance, but there remains a large gap for the overlooking surveillance scenarios. The results of face recognition depend not only on the structure of the model, but also on the completeness and diversity of the training samples. The existing multi-pose face datasets do not cover complete top-view face samples, and the models trained by them thus cannot provide satisfactory accuracy. To this end, this paper pioneers a multi-view tilt-angle face dataset (TFD), which is collected with an elaborately devised overhead capture equipment. TFD contains 11,124 face images from 927 subjects, covering a variety of tilt angles on the overhead view. To verify the validity of the constructed dataset, we further conduct comprehensive face detection and recognition experiments using the corresponding models trained by WiderFace, Webface and our TFD, respectively. Experimental results show that our TFD substantially promotes the face detection and recognition accuracy under the top-view situation. TFD is available at https://github.com/huang1204510135/D FD. Nanxi Wang, Zhongyuan Wang 0001, Zheng He 0001, Baojin Huang, Liguo Zhou, Zhen Han 0002 |
ICIP | 6 |
| 2021 | Adaptive Texture Distillation Network for Image Hybrid Super-ResolutionabstractTo save the transmission bandwidth of high-resolution (HR) images, we can send down-sampled low-resolution (LR) images and reconstruct them using super-resolution (SR) technology at the receiving end. However, image down-sampling by a large factor results in the loss of many spatial details. Instead, we use a combination of spatial down-sampling by a small factor and gray-level quantization to obtain the low hybrid-resolution images. Although the small down-sampling factor makes images retain more spatial details and real textures, the gray-level quantization introduces fake textures. Obviously, the real textures should be enhanced, and the fake textures should be eliminated. To address this issue, we propose a lightweight Adaptive Texture Distillation Network (ATDN) for image hybrid super-resolution. Our model uses the texture enhancement block (TEB) and the texture smoothing block (TSB) to handle real and fake textures in different ways. Considering that the mixing proportions of two kinds of textures in low hybrid-resolution images vary with regions, we specifically use a cascaded weight branch to adaptively adjust the weights of real and fake textures. Experiments reveal that our model can effectively deal with the mixing problem of real and fake textures, and our method can achieve superior performance to other lightweight methods. Chunlei Liu 0006, Zhen Han 0002, Jiaxing Wen, Zhongyuan Wang 0001, Weiping Tu |
IJCNN | 2 |
| 2021 | "One-Shot" Super-Resolution via Backward Style Transfer for Fast High-Resolution Style TransferabstractOwing to the excellent visual quality of results, Gatys et al.'s Neural Style Transfer (NST) online algorithm is regarded as the gold-standard in the community of NST, but this algorithm is quite time-consuming especially for high-resolution (HR) image. In this letter, we propose “One-Shot” super-resolution (SR) for fast high-resolution style transfer. We first generate a low-resolution (LR) stylized image by NST, and then use “One-Shot” super-resolution to restore the HR stylized image by learning the mapping relations between HR-LR stylized images from HR-LR style images. However, due to the style loss is not eliminated, there are some subtle but important fine-grained style differences between LR stylized and style images. These differences lead to the poor visual quality of SR results. To reduce the style differences further, we adjust the texture of LR style image to approach LR stylized image by backward style transfer. The result of backward style transfer will be treated as the LR part of the “One-Shot” example pair, which leads to a better SR. The experimental results show that with good visual quality, our method reduces the time consumption by 81.6%. Especially in a specific application scenario of fixed style image and changed content image, our method reduces the time consumption by 89.3%. Jikang Cheng, Zhen Han 0002, Zhongyuan Wang 0001, Liang Chen 0026 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Decomposition Makes Better Rain Removal: An Improved Attention-Guided Deraining NetworkabstractRain streaks in the air show diverse characteristics with different shapes, directions, densities, even the complex overlapped phenomenon, causing great challenges for the deraining task. Recently, deep learning based image deraining methods have been extensively investigated due to their excellent performance. However, most of the existing algorithms still have limitations in removing rain streaks while preserving rich textural details under complicated rain conditions. To this end, we propose to decompose rain streaks into multiple rain layers and individually estimate each of them along the network stages to cope with the increasing abstracts. To better characterize rain layers, an improved non-local block is designed to exploit the self-similarity of rain information by learning the holistic spatial feature correlations while reducing the calculation complexity. Moreover, a mixed attention mechanism is applied to guide the fusion of rain layers by focusing on the local and global overlaps among these rain layers. Extensive experiments on both synthetic rainy/rain-haze/raindrop datasets, real-world samples, the haze, and low-light scenarios show substantial improvements both on quantitative indicators and visual effects over the current state-of-the-art technologies. The source code is available athttps://github.com/kuihua/IADN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zhen Han 0002, Tao Lu 0001, Baojin Huang, Junjun Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Multi-Stage Degradation Homogenization for Super-Resolution of Face Images With Extreme DegradationsabstractFace Super-Resolution (FSR) aims to infer High-Resolution (HR) face images from the captured Low-Resolution (LR) face image with the assistance of external information. Existing FSR methods are less effective for the LR face images captured with serious low-quality since the huge imaging/degradation gap caused by the different imaging scenarios (i.e., the complex practical imaging scenario that generates test LR images, the simple manual imaging degradation that generates the training LR images) is not considered in these algorithms. In this paper, we propose an image homogenization strategy via re-expression to solve this problem. In contrast to existing methods, we propose a homogenization projection in LR space and HR space as compensation for the classical LR/HR projection to formulate the FSR in a multi-stage framework. We then develop a re-expression process to bridge the gap between the complex degradation and the simple degradation, which can remove the heterogeneous factors such as serious noise and blur. To further improve the accuracy of the homogenization, we extract the image patch set that is invariant to degradation changes as Robust Neighbor Resources (RNR), with which these two homogenization projections re-express the input LR images and the initial inferred HR images successively. Both quantitative and qualitative results on the public datasets demonstrate the effectiveness of the proposed algorithm against the state-of-the-art methods. Liang Chen 0026, Jinshan Pan, Junjun Jiang, Jiawei Zhang 0002, Zhen Han 0002, Linchao Bao |
IEEE Trans. Image Process. | 5 |
| 2020 | Image Super-Resolution Using Residual Global Context NetworkabstractRecent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the fact that natural images have strong internal data repetition which requires the network to capture long-range dependencies between pixels and considering the interdepen-dencies between channels can better exploit the input information of the network. In addition, although past studies have proved that deep convolutional neural network benefit the performance of image super-resolution, it also means that the network needs more memory consumption and higher computational complexity. To solve these problem,we introduce Global Context block (GCB) and design a comparative shallow network called Residual Global Context Networks (RGC-N). It achieves a better trade-off between the amount of parameter and the quality of image reconstruction. Extensive experiments demonstrate that the proposed method is superior to the state-of-the-art methods. Kuangye Liu, Zhen Han 0002, Junkui Chen, Chunlei Liu 0006, Jun Chen 0001, Zhongyuan Wang 0001 |
ICASSP | 2 |
| 2020 | Low-quality watermarked face inpainting with discriminative residual learningabstractMost existing image inpainting methods assume that the location of the repair area (watermark) is known, but this assumption does not always hold. In addition, the actual watermarked face is in a compressed low-quality form, which is very disadvantageous to the repair due to compression distortion effects. To address these issues, this paper proposes a low-quality watermarked face inpainting method based on joint residual learning with cooperative discriminant network. We first employ residual learning based global inpainting and facial features based local inpainting to render clean and clear faces under unknown watermark positions. Because the repair process may distort the genuine face, we further propose a discriminative constraint network to maintain the fidelity of repaired faces. Experimentally, the average PSNR of inpainted face images is increased by 4.16dB, and the average SSIM is increased by 0.08. TPR is improved by 16.96% when FPR is 10% in face verification. Zheng He 0001, Xueli Wei, Kangli Zeng, Zhen Han 0002, Qin Zou 0001, Zhongyuan Wang 0001 |
MMAsia | 4 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 4 |
| 2020 | Ultra-dense GAN for satellite imagery super-resolution
Zhongyuan Wang 0001, Kui Jiang, Peng Yi 0002, Zhen Han 0002, Zheng He 0001 |
Neurocomputing | 4 |
| 2020 | Modeling and Optimizing of the Multi-Layer Nearest Neighbor Network for Face Image Super-ResolutionabstractIn this paper, we propose a face super-resolution (FSR) method to handle the decreasing face recognition rate caused by low-quality images. To better model the input images, we build a nearest neighbor network (NNN) which consists of nodes and paths by introducing the second-layer nearest neighbors (SLNNs), where the paths of the network represent the distance between nodes. As the SLNN is trained in the high-resolution (HR) space and is exponentially supplementary to the traditional first-layer nearest neighbors (FLNNs), the neighbor inadequacy problem can be effectively solved by enriching the neighbor candidate set via NNN. Furthermore, we solve the NNN for the optimal weights of neighbors. Finally, we fuse the refined weights and neighbors for better reconstruction results. The effectiveness of this fusion strategy is validated by both quantitative and qualitative experimental results. The extensive experimental results on the public face datasets and real-world challenging low-resolution (LR) images demonstrate that the proposed method performs favorably against the state-of-the-art methods. Liang Chen 0026, Jinshan Pan, Ruimin Hu, Zhen Han 0002, Chao Liang 0001, Yi Wu 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | S3D: Scalable Pedestrian Detection via Score Scale Surface DiscriminationabstractPedestrian detection has remained an important research topic in both the computer vision and multimedia communities because of its importance in practical applications, such as driving assistance and video surveillance. Existing methods compare the response score with a fixed threshold to determine whether a candidate region contains pedestrians and produce dissatisfactory results that contain either missed detections or false detections, which are difficult to balance. This situation has a serious impact under the condition of variable scale. This paper investigates the functional relationship between the scores and scales of pedestrians. By designing experiments with multiple scales, we have found a discriminant surface in the score scale space. Pedestrians can be distinguished at various scale levels according to their locations on the discriminant surface. The proposed approach is evaluated using four challenging pedestrian detection datasets, including Caltech, INRIA, ETH, and KITTI, and the superior experimental results are achieved when compared with baseline methods. Xiao Wang 0029, Chao Liang 0001, Chen Chen 0001, Jun Chen 0001, Zheng Wang 0007, Zhen Han 0002, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Rain Streak Removal via Multi-scale Mixture Exponential Power ModelabstractRain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GMM compromises the ability of the model in fitting real noise, such as sparse noise, which is exactly the characteristic of the rain streaks. In this paper, a novel model named Mixture Exponential Power Model (MEPM) is exploited. It sets multiple Laplace noise components and expands the representation capability for the sparse noise. Moreover, considering that the rain streaks in a video occur in different distances from the camera, we encode rain streaks into Multi-scale Mixture Exponential Power Model. The model is opti-mized by expectation-maximization (EM) algorithm and La-grange multiplier strategy. Experiments are implemented on synthetic and real rain videos and verify the superiority of the proposed method, compared with state-of-the-art methods. Jun Chen 0001, Zhen Han 0002, Mingfu Xiong, Chao Liang 0001, Zhongyuan Wang 0001 |
ICASSP | 3 |
| 2019 | GAN-Based Multi-level Mapping Network for Satellite Imagery Super-ResolutionabstractAlthough many deep-learning-based image super-resolution (SR) methods have been proposed, most of them assume that all hierarchical features share the unified mapping equations. They ignore the differences between mapping equations at different feature levels, and create an average effect of mapping prediction, thus poorly building the mapping relations between low resolution (LR) and high resolution (HR) spaces. In this paper, we propose a multi-level mapping framework along with the adversarial learning strategy, namely MMGAN, for satellite imageries SR reconstruction. We also construct a feature extraction and tuning block (FETB) for fine feature expression. In particular, a novel two-dimension dense unit (DU) and a mapping attention unit (MAU) are constructed for building multi-level mappings in different stages. With our strategies, an HR image is reconstructed directly from the input image using multi-level mappings. Extensive experiments on Kaggle Open Source Dataset and Jilin-1 video satellite images exhibit superior reconstruction performance when compared with the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang, Guangcheng Wang, Zhen Han 0002, Tao Lu 0001 |
ICME | 6 |
| 2019 | Multi-Memory Convolutional Neural Network for Video Super-ResolutionabstractVideo super-resolution (SR) is focused on reconstructing high-resolution (HR) frames from consecutive lowresolution (LR) frames. Most previous video SR methods based on convolutional neural network (CNN) use a direct connection and single-memory module within the network, and they thus fail to make full use of spatio-temporal complementary information from LR observed frames. To fully exploit spatio-temporal correlations between adjacent LR frames and reveal more realistic details, this paper proposes a multi-memory convolutional neural network (MMCNN) for video SR, cascading an optical flow network and an image-reconstruction network. A serial of residual blocks engaged in utilizing intra-frame spatial correlations are proposed for feature extraction and reconstruction. Particularly, instead of using single-memory module, we embed convolutional long short-term memory (ConvLSTM) into the residual block, thus form a multi-memory residual block to progressively extract and retain inter-frame temporal correlations between consecutive LR frames. We conduct extensive experiments on numerous testing datasets with respect to different scaling factors. Our proposed MMCNN shows superiority over the state-of-the-art methods in terms of PSNR and visual quality and surpasses the best counterpart method 1 dB at most. The code and datasets are available at https://github.com/psychopa4/MMCNN. Zhongyuan Wang 0001, Peng Yi 0002, Kui Jiang, Junjun Jiang, Zhen Han 0002, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | A Novel Frontal Facial Synthesis Algorithm Based on Individual Residual Face
Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001 |
MMM (2) | 3 |
| 2017 | A joint learning based Face Super Resolution approach via contextual topological structureabstractFace Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are with low quality, the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we introduce the contextual topological structure of target patch to improve the consistency. The contextual topological structure consists of the target patch as well as its adjacent patches, we explore the relationship between them based on statistical probability and apply the relationship for joint learning progress of mapping from LR to HR. By incorporating the contextual topological structure, the robustness to noise of approach is increased as well as the LR/HR consistency. The effectiveness of proposed method is verified both quantitatively and qualitatively. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001 |
ICASSP | 3 |
| 2017 | Efficient mode decision for noisy video transcodingabstractIn many practical transcoding applications, such as video surveillance, the source videos are often contaminated by noise. The presence of noise not only results in poor compression efficiency and visual quality, but also imposes an adverse effect on the performance of subsequent video analysis tasks. Thereby it is very necessary to denoise the video. In this paper, we propose an efficient mode decision method for H.264 noisy video transcoding through analysing the effect of noise on mode decision. In the algorithm, we use the information available from previously decoded MBs to decide which modes can be overpassed with little loss to the rate-distortion performance. Experimental results show that our method saves the computational complexity nearly 65%, without noticeable rate-distortion(R-D) loss in comparison with the “Full-coding” (cascaded transcoding) method. Anna Zhang, Zhen Han 0002, Zhongyuan Wang 0001 |
ICASSP | 2 |
| 2017 | Face super resolution based on parent patch prior for VLQ scenarios
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Qing Li 0001, Zheng Lu 0002 |
Multim. Tools Appl. | 3 |
| 2017 | A novel face super resolution approach for noisy images using contour feature and standard deviation prior
Liang Chen 0026, Ruimin Hu, Chao Liang 0001, Qing Li 0001, Zhen Han 0002 |
Multim. Tools Appl. | 5 |
| 2017 | HRM graph constrained dictionary learning for face image super-resolution
Kebin Huang, Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
Multim. Tools Appl. | 4 |
| 2016 | Fast video enhancement transcodingabstractIn this paper, we pose a new problem of video enhancement transcoding, which converts the compressed dark video into compressed normal-lighting one. Distinct statistics of dark and normal videos result in quite different coding modes, which thus enforces latent constraints on mode conversion during transcoding. Following this idea, we propose a fast mode decision algorithm to speed up computation while maintaining rate-distortion (RD) performance. Experimental results show that our method saves the computational complexity nearly 70%, without noticeable RD loss in comparison with the cascaded decoder-encoder approach. Kefan Shen, Zhongyuan Wang 0001, Zhen Han 0002 |
ICIP | 3 |
| 2016 | Face Super Resolution for VLQ facial images via parent patch matchingabstractFace Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are Very Low Quality(VLQ), the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we use the target patch as well as the surrounding pixels, which we called parent patch, to represent the target patch. By incorporating the peripheral information, the parent patch is much more robust to noise in the LR and HR consistency learning. The effectiveness of proposed method is verified both quantitatively and qualitatively. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001, Zheng Lu 0002 |
IJCNN | 3 |
| 2016 | Face Image Super-Resolution Through Improved Neighbor Embedding
Kebin Huang, Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
MMM (1) | 4 |
| 2016 | Heteroskedasticity tuned mixed-norm sparse regularization for face hallucination
Zhongyuan Wang 0001, Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
Multim. Tools Appl. | 4 |
| 2016 | Facial Image Hallucination Through Coupled-Layer Neighbor EmbeddingabstractAs the facial image captured by a low-cost camera is typically very low resolution (LR), blurring, and noisy, traditional neighbor-embedding-based facial image hallucination methods from one single manifold (i.e., the LR image manifold) fail to reliably estimate the intention geometrical structure, consequently leading to a bias to the image reconstruction result. In this paper, we introduce the notion of neighbor embedding (NE) from the LR and the high-resolution (HR) image manifolds simultaneously and propose a novel NE model, termed the coupled-layer NE (CLNE), for facial image hallucination. CLNE differs substantially from other NE models in that it has two layers: the LR and the HR layers. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights of the LR patches; the HR layer is the intrinsic geometry that can geometrically constrain the reconstruction weights. With this coupled-constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust NE through iteratively updating the LR patch reconstruction weights and the estimated HR patch. The experimental results in simulation and real conditions confirm that the proposed method outperforms the related state-of-the-art methods in both quantitative and visual comparisons. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Coupled Discriminant Multi-Manifold Analysis with Application to Low-Resolution Face Recognition
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Liang Chen 0026, Jun Chen 0001 |
MMM (1) | 3 |
| 2014 | Gabor-based patch covariance matrix for face sketch synthesisabstractIn this paper, we propose a novel face sketch/photo synthesis method by utilizing Gabor-based Patch Covariance Matrix (GPCM) as face descriptor, a.k.a. symmetric positive definite matrix, which lie on a Riemannian manifold. In particular, both pixel locations and Gabor coefficients of one patch are employed to form the covariance matrix. In this way, the sketch/photo can be then transformed from the pixel space to the Riemannian manifold space. With the aid of the recently introduced Stein kernel theory, we advance to perform Regularized Least Square Representation (RLSR) in Stein space. Based on the assumption that the Stein divergence manifold of photo/sketch patch and the sketch/photo share the same topology, a new sketch/photo patch of the same position can be synthesized by keeping the weights and replacing the photo/sketch training image patches with the corresponding sketch/photo ones. Experimental results demonstrate the superiority of the proposed method. Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
ICIP | 4 |
| 2014 | Efficient learning based face hallucination approach via facial standard deviation priorabstractMost state-of-the-art face hallucination approaches suffer from complicated learning patterns and highly intensive computation, which will lead to low efficiency and considerable computing resources. Therefore, how to restore real face image quickly and efficiently is still an important issue in this field. To solve or partially solve the problem, this paper proposed a novel facial standard deviation prior based approach which can provide superior results with high efficiency for real face images. The high frequency information of test image will be enhanced via a facial specific sharpening operator which is obtained through the learning of standard deviation correspondence of training set. Experiments in simulation and real world images verified the effectiveness of proposed approach, and the distinct advantage on runtime and resource requirement of proposed approach. Liang Chen 0026, Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
ISCAS | 4 |
| 2014 | Noise robust face hallucination employing Gaussian-Laplacian mixture model
Zhongyuan Wang 0001, Zhen Han 0002, Ruimin Hu, Junjun Jiang |
Neurocomputing | 2 |
| 2014 | Efficient single image super-resolution via graph-constrained least squares regression
Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001 |
Multim. Tools Appl. | 3 |
| 2014 | Face image super-resolution through locality-induced support regression
Junjun Jiang, Ruimin Hu, Chao Liang 0001, Zhen Han 0002, Chunjie Zhang 0001 |
Signal Process. | 4 |
| 2014 | Face Super-Resolution via Multilayer Locality-Constrained Iterative Neighbor Embedding and Intermediate Dictionary LearningabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) manifolds are locally isometric, the neighbor embedding super-resolution algorithms try to preserve the geometry (reconstruction weights) of the LR space for the reconstructed HR space, but neglect the geometry of the original HR space. Due to the degradation process of the LR image (e.g., noisy, blurred, and down-sampled), the neighborhood relationship of the LR space cannot reflect the truth. To this end, this paper proposes a coarse-to-fine face super-resolution approach via a multilayer locality-constrained iterative neighbor embedding technique, which intends to represent the input LR patch while preserving the geometry of original HR space. In particular, we iteratively update the LR patch representation and the estimated HR patch, and meanwhile an intermediate dictionary learning scheme is employed to bridge the LR manifold and original HR manifold. The proposed method can faithfully capture the intrinsic image degradation shift and enhance the consistency between the reconstructed HR manifold and the original HR manifold. Experiments with application to face super-resolution on the CAS-PEAL-R1 database and real-world images demonstrate the power of the proposed algorithm. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 4 |
| 2014 | Noise Robust Face Hallucination via Locality-Constrained RepresentationabstractRecently, position-patch based approaches have been proposed to replace the probabilistic graph-based or manifold learning-based models for face hallucination. In order to obtain the optimal weights of face hallucination, these approaches represent one image patch through other patches at the same position of training faces by employing least square estimation or sparse coding. However, they cannot provide unbiased approximations or satisfy rational priors, thus the obtained representation is not satisfactory. In this paper, we propose a simpler yet more effective scheme called Locality-constrained Representation (LcR). Compared with Least Square Representation (LSR) and Sparse Representation (SR), our scheme incorporates a locality constraint into the least square inversion problem to maintain locality and sparsity simultaneously. Our scheme is capable of capturing the non-linear manifold structure of image patch samples while exploiting the sparse property of the redundant data representation. Moreover, when the locality constraint is satisfied, face hallucination is robust to noise, a property that is desirable for video surveillance applications. A statistical analysis of the properties of LcR is given together with experimental results on some public face databases and surveillance images to show the superiority of our proposed scheme over state-of-the-art face hallucination approaches. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
IEEE Trans. Multim. | 4 |
| 2013 | LBP-Guided Depth Image FilterabstractThe multi-view video plus depth (MVD) format has been put forward for the call for proposals in free view video (FVV) and 3DTV. Since representing the 3D scene geometry, depth maps are used for synthesizing virtual views. However, compression artifacts of the depth images always lead to geometry distortions in synthesized views. By exploiting LBP features of the corresponding color samples, we propose a novel local binary pattern (LBP) guided depth filter which enables the local neighborhood samples those are in the same object of the current pixel to be filtering input. In recognition of its ability for describing the object edges, the LBP operator is used to calculate the weighted values of the local depth pixels for the depth-map filter. Furthermore, the filter is incorporated into the framework of H.264/MVC as an in-loop filter. The experimental results demonstrate that the proposed approach offers 0.45dB and 0.66dB average PSNR gains in terms of video rendering quality and depth coding efficiency, as well as significant subjective improvement in rendering views. Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
DCC | 5 |
| 2013 | Manifold regularized sparse support regression for single image super-resolutionabstractIn this paper, we present a novel single image super-resolution method. To simultaneously improve the resolution and perceptual image quality, we bring forward a practical solution combining manifold regularization and sparse support regression. The main contribution of this paper is twofold. Firstly, a mapping function from low resolution (LR) patches to high-resolution (HR) patches will be learned by a local regression algorithm called sparse support regression, which can be constructed from the support bases of the LR-HR dictionary. Secondly, we propose to preserve the geometrical structure of the image patch dictionary, which is critical for reducing the artifacts and obtaining better visual quality. Experimental results demonstrate that the proposed method produces high quality results both quantitatively and perceptually. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Shi Dong 0004 |
ICASSP | 4 |
| 2013 | A joint learning based face hallucination approach for low quality face imageabstractThis paper describes a novel method for single-image super-resolution (SR) based on a neighbor embedding technique which uses coupled feature spaces under surveillance scenarios. For surveillance face images, traditional neighbor embedding SR approaches could not offer counterintuitive results because consistency between high resolution images and low resolution images is destroyed by serious noise which caused by environmental impact factors and large distance between the camera and objects. In order to reinforce the consistency, we extend the learning space from single to a coupled feature space that combine image intensity feature and contour model. The contour model describes facial contour information as images generated from original low resolution ones. Simulation experiments show that this proposed approach could provide competitive results in simulation experiments in subjective and objective quality. Even in surveillance scenario the proposed method outperforms the traditional methods. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Junjun Jiang |
ICIP | 3 |
| 2013 | Coupled-layer neighbor embedding for surveillance face hallucinationabstractAs the face image captured by a surveillance camera is typically very low-resolution (LR), blurred and noisy, traditional neighbor embedding method considers only one manifold (the LR image manifold) and fails very often to reliably estimate the intention geometrical structure. In this paper, we introduce the notion of neighbor embedding from the LR image manifold and the high-resolution (HR) one simultaneously and propose a novel neighbor embedding model, termed the coupled-layer neighbor embedding (CLNE), for surveillance face hallucination. CLNE differs substantially from other neighbor embedding models in that the former has two layers: the LR layer and the the HR layer. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights; the HR layer in this model is a set of HR training patches that guide the K-nearest neighbor (K-NN) searching and geometrically constrain the reconstruction weights. By this coupled constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust neighbor embedding through the significant degradation process. Indeed, the experimental results confirm that our method outperforms the related state-of-the-art methods by having better objective values as well as better visual results. Junjun Jiang, Ruimin Hu, Liang Chen 0026, Zhen Han 0002, Tao Lu 0001, Jun Chen 0001 |
ICIP | 4 |
| 2013 | Locality-constraint iterative neighbor embedding for face hallucinationabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the neighbor embedding based super-resolution algorithms try to preserve the local geometry of the patch manifold for the reconstructed HR patch manifold. However, due to “one-to-many” mappings between LR and HR images, the neighborhood relationship of the LR patch manifold can't reflect the inherent data structure. In this paper, we explore the data structure by both considering the LR patch and HR patch manifolds instead of only considering one manifold (LR patch manifold). By incorporating the position prior of face and local geometry of HR patch manifold, we propose an improved neighbor embedding method to face hallucination, namely locality-constraint iterative neighbor embedding (LINE), in which we iteratively update the K-nearest neighbors (K-NN) and reconstruction weights based on the result (the hallucinated HR patch) from previous iteration, giving rise to improved performance compared with traditional neighbor embedding algorithms. Experimental results with application to face hallucination on simulated LR face images and real world ones demonstrate the effectiveness of the proposed method. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Tao Lu 0001, Jun Chen 0001 |
ICME | 3 |
| 2013 | Support-driven sparse coding for face hallucinationabstractBy incorporating the prior of positions, position patch based face hallucination methods can produce high-quality results and save computation time. Given a low-resolution face image, the key issue of these methods is how to encode the input low-resolution patch. However, due to stability and accuracy issues, the coding approaches proposed so far are not satisfactory. In this paper, we present a novel sparse coding method via exploiting the support information on the coding coefficients. In particular, the support information is characterized by the locality of the image patch manifold, which has been shown to be critical in data representation and analysis. According to the distances between the input patch and bases in the dictionary, we first assign different weights to the coding coefficients and then obtain the coding coefficients by solving a weighted sparse problem. Our proposed method exploits the non-linear manifold structure of patch samples and the sparse property of the redundant data, leading to stable and accurate representation. Experiments on commonly used databases demonstrate that our method outperforms state of the art. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong, Zhen Han 0002 |
ISCAS | 5 |
| 2013 | Robust super-resolution for face images via principle component sparse representation and least squares regressionabstractFace image super-resolution (SR) reconstruction is the problem of inducing a high-resolution (HR) face image from a low-resolution (LR) one. Traditional face SR methods are either sensitive to noise, i.e., local patch based technologies, or lacking facial details, i.e., global face reconstruction, thus could not achieve a satisfying result. In order to overcome these problems, we propose in this paper a novel face SR method. Taking full advantages of Principle Component analysis and Sparse Representation (PCSR), it aims to obtain an accurate and noise robust representation, transforming the image patch to the principle component sparse feature space (PC-SFS). Moreover, in PC-SFS, we try to learn a mapping function between the LR image patches and HR ones through Least Squares Regression. Given a LR patch, we first transform it to the LR PC-SFS by PCSR to obtain the robust and accurate representation, and then project the representation to the HR PC-SFS thus get the target HR patch. Experiments on the frontal faces SR in noise conditions demonstrate our method outperforms state of the art. Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang |
ISCAS | 3 |
| 2013 | From local representation to global face hallucination: A novel super-resolution method by nonnegative feature transformationabstractMost of global face hallucination methods treat the face as a whole, ignoring the fact that the face is composed by part-based organs. Therefore, the results obtained by these methods always lack of detailed information. Nonnegative matrix factorization (NMF) based face hallucination method is properly used to enhance the detailed information. Usually, NMF basis is only learnt from high-resolution (HR) samples, leading to over-smooth output and lack of high frequency details. In order to solve this problem, we propose a simple but novel face hallucination method using nonnegative feature transformation by two-step framework. In particular, we learn the NMF basis from low-resolution (LR) and HR samples separately, and then transform the local representation feature of input into the global representation subspaces, keeping the weights into the HR samples space for output. Furthermore, the maximum a posteriori (MAP) method is used to estimate a better output. Experiments show that the hallucinated face of the proposed method is not only more high-frequency details, but also has better performance than many state-of-art algorithms. Tao Lu 0001, Ruimin Hu, Zhen Han 0002, Junjun Jiang, Yanduo Zhang |
VCIP | 3 |
| 2012 | A super-resolution method for low-quality face image through RBF-PLS regression and neighbor embeddingabstractIn this paper, a new two-step method is proposed to infer a high-quality and high-resolution (HR) face image from a low-quality and low-resolution (LR) observation based on training samples in the database. First, a global face image is reconstructed based on the non-linear relationship between LR and HR face images, which is established according to radial basis function and partial least squares (RBF-PLS) regression. Based on the reconstructed global face patches manifold (formed by the image patches at the same position of all global face images), whose local geometry is more consistent with that of original HR face patches manifold than noisy LR one is, the Neighbor Embedding is applied to induce the target HR face image by preserving the similar local geometry between global face patches manifold and the original HR face patches manifold. A comparison of some state-of-the-art methods shows the superiority of our method, and experiments also demonstrate the effectiveness both under simulation and real conditions. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ICASSP | 3 |
| 2012 | Graph discriminant analysis on multi-manifold (GDAMM): A novel super-resolution method for face recognitionabstractHow to efficiently recognize low-resolution (LR) probe images of one face recognition system, in which high-resolution (HR) gallery of faces is enrolled, is still an open problem. In this paper, we develop a novel super-resolution method, namely Graph Discriminant Analysis on Multi-Manifold (GDAMM), to super-resolved the HR version of a LR probe image and then perform matching at the resolution of the HR gallery. Unlike classical super-resolution approaches considering only the data fidelity, GDAMM takes the advantages of both manifold learning and discriminant analysis to integrate the data constraint and discriminant constraint, seeking the mapping between LR images and HR ones. In the reconstructed HR image space, faces of one person in the same manifold are close and those in different manifolds are far apart. Experiments on Extended Yale-B database and AR face database demonstrate that the learned discriminant information is essential for improving recognition accuracy. Through the contrastive experiment, the results (recognition rates) indicate that the proposed GDAMM method can greatly surpass classical super-resolution approaches, even outperforming the ideal case of having probe images of HR gallery by a big margin (nearly 9% on Extended Yale-B database and 8% on AR face database). Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001 |
ICIP | 3 |
| 2012 | Efficient Single Image Super-Resolution via Graph EmbeddingabstractWe explore in this paper efficient algorithmic solutions to single image super-resolution (SR). We propose the GESR, namely Graph Embedding Super-Resolution, to super-resolve a high-resolution (HR) image from a single low-resolution (LR) observation. The basic idea of GESR is to learn a projection matrix mapping the LR image patch to the HR image patch space while preserving the intrinsic geometrical structure of original HR image patch manifold. While GESR resembles other manifold learning-based SR methods in persevering the local geometric structure of HR and LR image patch manifold, the innovation of GESR lies in that it preserves the intrinsic geometrical structure of original HR image patch manifold rather than LR image patch manifold, which may be contaminated because of image degeneration (e.g., blurring, down-sampling and noise). Experiments on benchmark test images show that GESR can achieve very competitive performance as Neighbor Embedding based SR (NESR) and Sparse representation based SR (SSR). Beyond subjective and objective evaluation, all experiments show that GESR is much faster than both NESR and SSR. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Kebin Huang, Tao Lu 0001 |
ICME | 3 |
| 2012 | Position-Patch Based Face Hallucination via Locality-Constrained RepresentationabstractInstead of using probabilistic graph based or manifold learning based models, some approaches based on position-patch have been proposed for face hallucination recently. In order to obtain the optimal weights for face hallucination, they represent image patches through those patches at the same position of training face images by employing least square estimation or convex optimization. However, they can hope neither to provide unbiased solutions nor to satisfy locality conditions, thus the obtained patch representation is not the best. In this paper, a simpler but more effective representation scheme- Locality-constrained Representation (LcR) has been developed, compared with the Least Square Representation (LSR) and Sparse Representation (SR). It imposes a locality constraint onto the least square inversion problem to reach sparsity and locality simultaneously. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art face hallucination approaches. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ICME | 3 |
| 2012 | Improvements of dynamic texture synthesis for video coding
Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
ICPR | 4 |
| 2012 | Face hallucination via K-selection mean constrained sparse representation
Kebin Huang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Junjun Jiang |
ICPR | 3 |
| 2012 | Surveillance face hallucination via variable selection and manifold learningabstractIn this paper, we propose a new two-step face hallucination method to induce a high-resolution (HR) face image from a low-resolution (LR) observation. Especially for low-quality surveillance face image, an RBF-PLS based variable selection method is presented for the reconstruction of global face image. Further more, in order to compensate for the reconstruction errors, which are lost high frequency detailed face features, the Neighbor Embedding (NE) based residue face hallucination algorithm is used. Compared with current methods, the proposed RBF-PLS based method can generate a global face more similar to the original face and less sensitive to noise, moreover, the NE algorithm can reduce the reconstruction errors caused by misalignment on the basis of a carefully designed search strategy. Experiments show the superiority of the proposed method compared with some state-of-the-art approaches and the efficacy both in simulation and real surveillance condition. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Tao Lu 0001, Kebin Huang |
ISCAS | 3 |
| 2012 | Face image super-resolution via nearest feature lineabstractIn this paper, we propose a manifold learning based algorithm using 'Nearest Feature Line - NFL' to hallucinate high-resolution face image. According to the fact that existing NFL can effectively characterize the geometrical proportions to the face samples, we propose using NFL metric to define the neighborhood relations between face samples. Our algorithm can solve the problem that traditional method cannot effectively reveal the similar local geometry between high-resolution and low-resolution face manifolds under the condition that the training sample size is small. Moreover, in order to enhance the representation capacity of available face samples and reduce the computational complexity, we select neighborhood samples for each input LR image. Experimental results demonstrate that our algorithm can generates clearer local feature details, and the PSNR is 1.4 dB higher than that of the best manifold learning based method reported so far. Zhen Han 0002, Junjun Jiang, Ruimin Hu, Tao Lu 0001, Kebin Huang |
ACM Multimedia | 1 |
| 2010 | A face super-resolution approach using shape semantic mode regularizationabstractIn actual imaging environment, a variety of factors have an impact on the quality of images, which leads to pixel distortion and aliasing. The traditional face super-resolution algorithm only uses the difference of image pixel values as similarity criterion, which degrades similarity and identification of reconstructed facial images. Image semantic information with human understanding, especially structural information, is robust to the degraded pixel values. In this paper, we propose a face super-resolution approach using shape semantic model. This method describes the facial shape as a series of fiducial points on facial image. And shape semantic information of input image is obtained manually. Then a shape semantic regularization is added to the original objective function. The steepest descent method is used to obtain the unified coefficient. Experimental results demonstrate that the proposed method outperforms the traditional schemes significantly both in subjective and objective quality. Chengdong Lan, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001 |
ICIP | 3 |
| 2010 | Global Face Super Resolution and Contour Region Constraints
Chengdong Lan, Ruimin Hu, Tao Lu 0001, Ding Luo, Zhen Han 0002 |
ISNN (2) | 5 |
| 2010 | Face hallucination with shape parameters projection constraintabstractIn real surveillance scenarios, a variety of factors have an impact on the quality of images, which leads to pixel distortion and aliasing. Traditional face super-resolution algorithms only use the difference of image pixel values as similarity criterion, which degrades similarity and identification of reconstructed facial images. Image semantic information with human understanding, especially structural data of shapes, is robust to the degraded images. In this paper, we propose a face hallucination with shape parameters projection constraint. This method uses a parameter model to represent face shapes, and shape information of input image is introduced to improving the quality of reconstructed image. The shape model regularization is first added to original objective function. Then shape parameters are projected into the domain of image parameters by a linear regression model. Finally, the gradient descent method is used to obtain the unified parameters. Experimental results demonstrate the proposed method outperforms the traditional schemes significantly both in subjective and objective quality. Chengdong Lan, Ruimin Hu, Kebin Huang, Zhen Han 0002 |
ACM Multimedia | 4 |
| 2007 | A Novel Intra/Inter Mode Decision Algorithm for H.264/AVC Based on Spatio-temporal Correlation
Qiong Liu 0001, Shengfeng Ye, Ruimin Hu, Zhen Han 0002 |
MMM (1) | 4 |