EDBT 2026 Demo / reviewers in the wild / expert
Zhongyuan Wang 0001
dblp:84/6394-1
· DBLP profile ↗
250ranked-venue papers
10as first author
144since 2021 · last 2026
0000-0002-9796-488XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 160 · 6 first-author · 74 since 2021Artificial intelligence and machine learning · 88 · 4 first-author · 71 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 13 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Computer networks · 9 · 9 since 2021Security and privacy · 3 · 3 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object TrackerabstractLow-frame-rate (LFR) Multi-Object Tracking (MOT) is crucial for efficient tracking on edge devices, as it significantly reduces computational and storage demands. However, existing trackers struggle in LFR settings due to large temporal gaps, extreme appearance changes, and motion non-linearity. While Graph Neural Network (GNN)-based trackers are effective at associating objects across these gaps, most operate offline, which prevents their use for online tracking. To address these limitations, we propose GLoMOT, a novel online GNN-based Low-Frame-Rate Multi-Object Tracker designed for robust performance in LFR videos. To bridge the large temporal gaps, we introduce a Dynamic Node Buffer Pool. This acts as a long-term memory, caching the states of absent objects to enable their robust re-association. To tackle extreme motion uncertainty, we propose an adaptive context-aware module that dynamically adjusts the weights of positional and appearance features, generating more robust features for predicting node connections. Furthermore, we propose a pseudo-depth feature calculation method. This provides the GNN with critical geometric context, which helps resolve spatial ambiguity arising from occlusions. Extensive experiments on several public MOT benchmarks, including DanceTrack, MOT17, and VisDrone, demonstrate GLoMOT's effectiveness and superiority, particularly in challenging Low-Frame-Rate conditions. Yaxuan Hu 0001, Jie Hua 0005, Gang Wu 0010, Yuhong Yang 0001, Atsushi Suzuki 0002, Zhongyuan Wang 0001 |
AAAI | 6 |
| 2026 | Rethinking Surgical Smoke: A Smoke-Type-Aware Laparoscopic Video Desmoking Method and DatasetabstractElectrocautery or lasers will inevitably generate surgical smoke, which hinders the visual guidance of laparoscopic videos for surgical procedures. The surgical smoke can be classified into different types based on its motion patterns, leading to distinctive spatio-temporal characteristics across smoky laparoscopic videos. However, existing desmoking methods fail to account for such smoke-type-specific distinctions. Therefore, we propose the first Smoke-Type-Aware Laparoscopic Video Desmoking Network (STANet) by introducing two smoke types: Diffusion Smoke and Ambient Smoke. Specifically, a smoke mask segmentation sub-network is designed to jointly conduct smoke mask and smoke type predictions based on the attention-weighted mask aggregation, while a smokeless video reconstruction sub-network is proposed to perform specially desmoking on smoky features guided by two types of smoke mask. To address the entanglement challenges of two smoke types, we further embed a coarse-to-fine disentanglement module into the mask segmentation sub-network, which yields more accurate disentangled masks through the smoke-type-aware cross attention between non-entangled and entangled regions. In addition, we also construct the first large-scale synthetic video desmoking dataset with smoke type annotations. Extensive experiments demonstrate that our method not only outperforms state-of-the-art approaches in quality evaluations, but also exhibits superior generalization across multiple downstream surgical tasks. Qifan Liang, Zhen Han 0002, Xihao Wang, Zhongyuan Wang 0001, Bin Mei |
AAAI | 5 |
| 2026 | DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery LocalizationabstractTemporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries, sparse forgeries, and limited long-range modeling. We propose DeformTrace, which enhances SSMs with deformable dynamics and relay mechanisms to address these challenges. Specifically, Deformable Self-SSM (DS-SSM) introduces dynamic receptive fields into SSMs for precise temporal localization. To further enhance its capacity for temporal reasoning and mitigate long-range decay, a Relay Token Mechanism is integrated into DS-SSM. Besides, Deformable Cross-SSM (DC-SSM) partitions the global state space into query-specific subspaces, reducing non-forgery information accumulation and boosting sensitivity to sparse forgeries. These components are integrated into a hybrid architecture that combines the global modeling of Transformers with the efficiency of SSMs. Extensive experiments show that DeformTrace achieves state-of-the-art performance with fewer parameters, faster inference, and stronger robustness. Suting Wang, Yuanming Zheng, Junqi Yang, Yangxu Liao, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001 |
AAAI | 8 |
| 2026 | An angle-guided bidirectional feature transformation network for multi-frame tilt-angle face recognition
Wenqin Song, Xihao Wang, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
Zhengqian Wu, Zhixian Liu, Aodong Chen, Jingyang Zhang, Ruizhe Li 0004, Hanlin Ge, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | A cyclic diffusion framework for structure-authentic and annotation-disentangled anomaly generation
Linchun Wu, Qin Zou 0001, Xianbiao Qi, Zhongyuan Wang 0001, Qingquan Li 0001 |
Neurocomputing | 4 |
| 2026 | Tortoise plastron versus adulterants: identification and comparative study using image recognition technologyabstractCompared with botanical medicines, the intelligent identification of animal-derived medicines has developed relatively slowly and presents greater challenges due to species diversity, substantial morphological variations after processing, and the prevalence of adulteration. Tortoise plastron is a representative animal-derived medicine whose subtle morphological differences complicate reliable authentication. This study aimed to establish a standardized image-based framework to achieve accurate and reproducible identification of tortoise plastron and its common adulterants. An RGB image dataset covering Chinemys reevesii , Mauremys mutica , Ocadia sinensis , Malayemys subtrijuga , and Trachemys scripta elegans was constructed. Images were enhanced through geometric transformations, color jittering, and Gaussian blurring to simulate diverse acquisition conditions. Segmentation was performed using the SAM2 model to remove background noise and extract core regions, and classification was conducted using YOLO11 models of three scales (m, l, x). Training and validation were carried out on datasets with and without segmentation-based augmentation, each repeated three times, with average performance recorded. Statistical analysis included comparison of model performance metrics across datasets and model scales. The combined preprocessing–segmentation–classification workflow effectively captured both global and fine-grained features of tortoise plastron. YOLO11m with segmentation-based augmentation achieved the most balanced performance across accuracy and robustness. Eight technical modules, including attention-enhanced feature extraction and multi-scale pooling, contributed to improved classification precision. A practical recognition application was developed to facilitate user-friendly deployment. This study established a comprehensive digital framework for the objective identification of tortoise plastron and adulterants, transforming subjective trait-based evaluation into quantitative image analysis. The integration of advanced segmentation and multi-scale feature fusion provides a transferable paradigm for the intelligent identification of animal-derived medicines, with potential to enhance quality control and authenticity assurance in traditional Chinese medicine. Haoyu Tu, Xiaoshun Wang, Zifang Wu, Xinyue Zhou, Jiaxin Zou, Yaodong Ping, Wentao Sheng, Lei Wang 0084, Pengfei Jin, Hankun Hu, Zhongyuan Wang 0001 |
Mach. Vis. Appl. | 14 |
| 2026 | Prompt-guided Modality Completion for cardiac pathology segmentation
Donggen Fang, Yajie Chen, Yuliang Gu, Lingyi Yu, Zhongyuan Wang 0001, Bo Du 0001, Lianming Wu, Yongchao Xu |
Pattern Recognit. | 5 |
| 2026 | MSTDNet: Multi-scale traffic object detection network with smooth information perception
Jie Hua 0005, Zhongyuan Wang 0001, Hua Zou 0002, Gang Wu 0010, Jiayi Ma 0001 |
Pattern Recognit. | 2 |
| 2026 | Physics-inspired pseudo anomaly generation and prototype feature guidance for 3D anomaly detection
Jian Ning, Qin Zou 0001, Linchun Wu, Yuanhao Yue, Kunmo Li, Shoubin Chen, Zhongyuan Wang 0001 |
Pattern Recognit. | 7 |
| 2026 | STKPS-Net: Spatio-Temporal Key Patch Selection Network for Few Shot Anomalous Action RecognitionabstractFor providing timely warnings and preventing potential damages, it is crucial to detect anomalous actions that threaten public safety through surveillance cameras. Compared to normal actions, anomalous actions often occupy only a small portion of surveillance videos and exhibit more complex manifestations in terms of time and space. Considering that normal action recognition methods fail to highlight crucial information from small-sized patches, we propose the Spatio-temporal Key Patch Selection Network (STKPS-Net). It includes a spatially adaptive key patch selection module to select small but informative patches, and a long-short feature map spatio-temporal relation module to capture dynamic changes in anomalous actions. Additionally, a spatio-temporal refined loss is introduced to enhance fine-grained feature learning. Experimental results on the HMDB51, Kinetics, and UCF-Crime v2 datasets show that our STKPS-Net achieves state-of-the-art performance in few-shot anomalous action recognition, outperforming the most competitive methods by 1.2% on the anomalous action dataset UCF-Crime v2. Jinsheng Xiao, Ruidi Chen, Xingyu Gao 0001, Hailong Shi, Zhongyuan Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | IDRetracor: Towards Visual Forensics against Malicious Face SwappingabstractThe deepfake-based face swapping technique poses significant risks to personal identity security. Although many detection methods have been proposed to counter malicious face swapping, they typically provide only binary labels (Fake/Real), lacking reliable and interpretable evidence. To address this limitation, we introduce a novel task called face retracing, which aims to visually trace back the original target face from a given fake one through inverse mapping. This task is based on the observation that current face swapping methods are neither flawless nor entirely random, leaving recoverable traces of the original identity. To this end, we propose IDRetracor, a model designed to recover arbitrary original target identities from fake faces generated by various face swapping techniques. Specifically, we first employ a mapping resolver to estimate the possible solution space of the original face for inverse mapping. Then, we introduce Mapping-Aware Convolutions (MACs), which consist of multiple dynamically combined kernels guided by the mapping resolver to adaptively handle diverse face swapping patterns. Extensive experiments demonstrate that IDRetracor achieves strong performance in retracing original faces, validated by both quantitative metrics and qualitative assessments. Jikang Cheng, Jiaxin Ai, Zhen Han 0002, Chao Liang 0001, Qin Zou 0001, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story VideosabstractVideo question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos, the most significant feature of story videos is storylines, which are composed of complex interactions and long-range evolvement of core story topics including characters, actions and locations. Understanding these topics requires models to possess DVU capability. However, existing DVU datasets rarely organize questions according to these story topics, making them difficult to comprehensively assess VideoQA models' DVU capability of complex storylines. Additionally, the question quantity and video length of these dataset are limited by high labor costs of handcrafted dataset building method. In this paper, we devise a large language model based multi-agent collaboration framework, StoryMind, to automatically generate a new large-scale DVU dataset. The dataset, FriendsQA, derived from the renowned sitcom Friends with an average episode length of 1,358 seconds, contains 44.6K questions evenly distributed across 14 fine-grained topics. Finally, We conduct comprehensive experiments on 10 state-of-the-art VideoQA models using the FriendsQA dataset. Zhengqian Wu, Ruizhe Li 0004, Zijun Xu 0001, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
AAAI | 4 |
| 2025 | Multi-Shape Matching with Cycle Consistency Basis via Functional MapsabstractMulti-shape matching is a central problem in various applications of computer vision and graphics, where cycle consistency constraints play a pivotal role. For this issue, we propose a novel and efficient approach that models multi-shapes as directed graphs for two-stage optimization, i.e., optimizing pairwise correspondence accuracy using landmarks, and refining matching consistency through cycle consistency basis. Specifically, we utilize local mapping distortion to identify landmarks and extract the dimension of the functional space, which is then used to upsample in the spectral domain, thereby producing smoother results. Next, to optimize the consistency of correspondences, we introduce the cycle consistency basis, which succinctly describes all consistent cycles in the collection. We then propose cycle consistency refinement, which resolves inconsistencies in cycles efficiently via the alternating direction method of multipliers. Our approach simultaneously balances the accuracy and consistency of multi-shape matching, achieving lower correspondence errors. Extensive experiments on several public datasets demonstrate the superiority of our approach over current state-of-the-art methods. Tianwei Ye, Huabing Zhou, Zhongyuan Wang 0001, Jiayi Ma 0001 |
AAAI | 4 |
| 2025 | Cross-Modal Stealth: A Coarse-to-Fine Attack Framework for RGB-T TrackerabstractCurrent research on adversarial attacks mainly focuses on RGB trackers, with no existing methods for attacking RGB-T cross-modal trackers. To fill this gap and overcome its challenges, we propose a progressive adversarial patch generation framework and achieve cross-modal stealth. On the one hand, we design a coarse-to-fine architecture grounded in the latent space to progressively and precisely uncover the vulnerabilities of RGB-T trackers. On the other hand, we introduce a correlation-breaking loss that disrupts the modal coupling within trackers, spanning from the pixel to the semantic level. These two design elements ensure that the proposed method can overcome the obstacles posed by cross-modal information complementarity in implementing attacks. Furthermore, to enhance the reliable application of the adversarial patches in real world, we develop a point tracking-based reprojection strategy that effectively mitigates performance degradation caused by multi-angle distortion during imaging. Extensive experiments demonstrate the superiority of our method. Xinyu Xiang, Qinglong Yan, Hao Zhang 0073, Jianfeng Ding, Han Xu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
AAAI | 6 |
| 2025 | OODML: Whole Slide Image Classification Meets Online Pseudo-Supervision and Dynamic Mutual LearningabstractBag-label-based multi-instance learning (MIL) has demonstrated significant performance in whole slide image (WSI) analysis, particularly in pseudo-label-based learning schemes. However, due to inaccurate feature representation and interference, existing MIL methods often yield unreliable pseudo-labels, which spawn undesired predictions. To address these issues, we propose an Online Pseudo-Supervision and Dynamic Mutual Learning (OODML) framework that enhances pseudo-label generation and feature representation while exploring their mutual learning to improve bag-level prediction. Specifically, we design an Adaptive Memory Bank (AMB) to collect the most informative components of the current WSI. We also introduce a Self-Progressive Feature Fusion (SPFF) module that integrates label-related historical information from the AMB with current semantic variations, thereby enhancing the representation of pseudo-bag tokens. Furthermore, we propose a Decision Revision Pseudo-Label (DRPL) generation scheme to explore intrinsic connections between pseudo-bag representations and bag-label predictions, resulting in more reliable pseudo-label generation. To alleviate redundant and ambiguous representations, the class-wise prior of pseudo-label prediction is borrowed to facilitate label-related feature learning and to update the AMB, forming a mutual refinement between feature representation and pseudo-label generation. Additionally, a Dynamic Decision-Making (DDM) module is developed to harmonize explicit and implicit representations of bag information for more robust decision-making. Extensive experiments on four datasets demonstrate that our OODML surpasses the state-of-the-art by 3.3% and 6.9% on the CAMELYON16 and TCGA Lung datasets. Kui Jiang, Hongxun Yao, Yi Xiao 0003, Zhongyuan Wang 0001 |
AAAI | 5 |
| 2025 | EndoCADx: A Real-Time LVLM-Based CADx System for Multimodal Diagnosis of Gastric Lesions in White-Light EndoscopyabstractAccurate characterization and detailed documentation of gastric lesions during endoscopy are critical for early diagnosis and effective patient management. However, conventional computer-aided detection (CAD) systems primarily focus on lesion detection and lack the ability to generate comprehensive semantic descriptions, which limits their clinical utility. To address this gap, we have developed EndoCADx, a real-time computer-aided diagnosis (CADx) system integrating the Qwen2-VL large vision-language model. It was specifically designed for the real-time detection and description of gastric focal lesions in white-light endoscopic images. To fine-tune and evaluate the model, we constructed a domain-specific multimodal dataset comprising 7,543 expert-annotated image-text pairs across five common lesion types, named EndoGastro-7k. EndoCADx achieved a lesion detection accuracy of 91.6% and an accuracy in describing five key lesion characteristics of 84.7%, outperforming several state-of-the-art large vision-language models (LVLMs). Clinical evaluations further demonstrated its practical utility, with 86% of physicians endorsing its diagnostic completeness and 90% expressing a willingness to integrate it into routine practice. EndoCADx represents the first LVLM-based real-time CADx system tailored to gastric lesion analysis. It offers enhanced diagnostic accuracy, richer lesion interpretation, and improved support for endoscopic decision-making. Tingshun Xiong, Changhong Xiao, Ruiqing Jiang, Bin Mei, Lianlian Wu, Zhongyuan Wang 0001 |
BIBM | 9 |
| 2025 | Stacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery DetectionabstractThe rapid advancement of face forgery techniques has introduced a growing variety of forgeries. Incremental Face Forgery Detection (IFFD), involving gradually adding new forgery data to fine-tune the previously trained model, has been introduced as a promising strategy to deal with evolving forgery methods. However, a naively trained IFFD model is prone to catastrophic forgetting when new forgeries are integrated, as treating all forgeries as a single “Fake” class in the Real/Fake classification can cause different forgery types overriding one another, thereby resulting in the forgetting of unique characteristics from earlier tasks and limiting the model’s effectiveness in learning forgery specificity and generality. In this paper, we propose to stack the latent feature distributions of previous and new tasks brick by brick, i.e., achieving aligned feature isolation. In this manner, we aim to preserve learned forgery information and accumulate new knowledge by minimizing distribution overriding, thereby mitigating catastrophic forgetting. To achieve this, we first introduce Sparse Uniform Replay (SUR) to obtain the representative subsets that could be treated as the uniformly sparse versions of the previous global distributions. We then propose a Latent-space Incremental Detector (LID) that leverages SUR data to isolate and align distributions. For evaluation, we construct a more advanced and comprehensive benchmark tailored for IFFD. The leading experimental results validate the superiority of our method. Code is available at https://github.com/beautyremain/SUR-LID . Jikang Cheng, Zhiyuan Yan 0002, Ying Zhang 0021, Jiaxin Ai, Qin Zou 0001, Chen Li 0031, Zhongyuan Wang 0001 |
CVPR | 8 |
| 2025 | Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GameabstractMulti-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of multi-exit neural networks, we identify a long-neglected flaw in the evaluation of previous studies: simply using a fixed set of exits for attack may lead to an overestimation of their defense capacity. Based on this finding, our work explores the following three key aspects in the adversarial robustness of multi-exit neural networks: (1) we discover that a mismatch of the network exits used by the attacker and defender is responsible for the overestimated robustness of previous defense methods; (2) by finding the best strategy in a two-player zero-sum game, we propose AIMER as an improved evaluation scheme to measure the intrinsic robustness of multi-exit neural networks; (3) going further, we introduce NEED defense method under the evaluation of AIMER that can optimize the defender’s strategy by finding a Nash equilibrium of the game. Experiments over 3 datasets, 7 architectures, 7 attacks and 4 baselines show that AIMER evaluates the robustness 13.52% lower than previous methods under AutoAttack, while the robust performance of NEED surpasses single-exit networks of the same backbones by 5.58% maximally. Keyizhi Xu, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
CVPR | 4 |
| 2025 | Link-based Contrastive Learning for One-Shot Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) transfers knowledge from a labeled source domain to an unlabeled target domain via distribution alignment. However, in real-world scenarios like public safety or access control, obtaining sufficient source data is challenging, limiting existing UDA methods. This paper investigates a realistic but rarely studied problem called one-shot unsupervised domain adaptation (OSUDA), where only one source example per category is available. OSUDA faces dual challenges in feature learning and domain alignment due to the extreme source data scarcity. To address these, we propose link-based contrastive learning (LCL), a simple yet effective approach for OSUDA. LCL leverages in-domain links to learn discriminative features from abundant unlabeled target data and cross-domain links to achieve precise domain alignment with only one source sample per category. Extensive experiments on four domain adaptation benchmarks (VisDA-2017, Office-31, Office-Home, and DomainNet) demonstrate LCL’s effectiveness under the OSUDA setting. Additionally, we construct a real-world OSUDA surveillance face recognition dataset, where LCL consistently improves recognition performance across various face recognition methods. Yue Zhang 0104, Mingyue Bin, Zhongyuan Wang 0001, Zhen Han 0002, Chao Liang 0001 |
CVPR | 4 |
| 2025 | ACRL-10K: A Dataset for Air Conditioner Refrigerant Leak Smoke DetectionabstractIn this paper, we introduce a new dataset for air conditioner refrigerant leak smoke detection, called ACRL-10K. The dataset is designed to develop algorithms for detecting refrigerant leak smoke faults during air conditioner recycling. It contains a total of 10,724 images covering three common scenarios of air conditioner refrigerant leak: loading port, refrigerant extraction, and disassembly. All images are sourced from 656 video segments of real air conditioner recycling scenarios, captured by surveillance cameras deployed on an environmental company’s recycling production line. The annotations for the refrigerant leak smoke in the images are performed by industry experts to ensure high accuracy and consistency. To the best of our knowledge, ACRL-10K is the first dataset specifically designed for refrigerant leak smoke detection. Based on the ACRL-10K dataset, we present the performance of mainstream object detectors, including the YOLO series, as a baseline to conduct the following work: 1) preliminarily summarize the challenges of using the dataset for refrigerant leak smoke detection; 2) show the detection results of benchmark methods; and 3) make a comparison to identify the strengths and weaknesses of the baseline algorithms. In practice, the ACRL-10K dataset would hopefully advance research and applications in refrigerant leak early warning during air conditioner dismantling and recycling processes. Zhongyuan Wang 0001, Jie Hua 0005, Jinbi Liang, Zengmin Xu |
ICASSP | 2 |
| 2025 | ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges
Jiaxin Ai, Zhaopan Xu, Fanrui Zhang, Zizhen Li, Yukang Feng, Baojin Huang, Zhongyuan Wang 0001, Kaipeng Zhang |
ICCV | 10 |
| 2025 | Multimodal Re-Ranking for Heterogeneous Face Re-IdentificationabstractHeterogeneous face re-identification (Re-ID), aiming to match low-quality faces captured by disjoint visible light (VIS) and near-infrared (NIR) cameras, has become a critical application in video surveillance. However, the domain discrepancy between the NIR-VIS faces degrades the Re-ID performance. To solve this problem, this paper proposes a multimodal re-ranking method including two stages. Firstly, we utilize the VIS-NIR face bi-directional modality transformation based on the positive and negative samples separate training strategy to reduce domain discrepancy and generate the multimodal ranking lists of face Re-ID with complementarities. Secondly, we propose linear and nonlinear multimodal ranking lists fusion strategies based on single-modal and multi-modal k-reciprocal nearest neighbors (K-RNNs) to obtain a more accurate fused ranking list for face Re-ID. Extensive experiments on heterogeneous face datasets demonstrate the superior performance of our method over existing methods. Wenqin Song, Jiawei Zhang 0002, Zhen Han 0002, Yunfeng Xue, Xihao Wang, Zhongyuan Wang 0001 |
ICIP | 7 |
| 2025 | PGD-N2L: A Parameter-Guided Disentanglement Approach for Normal-To-Lombard Speech ConversionabstractThe Normal-To-Lombard (N2L) speech conversion can effectively improve speech intelligibility in noisy communication scenarios and serve as a data augmentation tool for various speech-related algorithms. However, existing N2L methods did not aim to disentangle the Lombard effect from other speech attributes, leading to incomplete conversions. In this paper, we propose a Parameter-Guided Disentanglement approach for N2L speech conversion (PGD-N2L) which decomposes speech into linguistic content, speaker identity, and Lombard effect. To extract disentangled linguistic content, we propose a DeLomb-Based content encoder. To extract disentangled speaker identity and Lombard effect, we propose a style encoder that combines a fine-tuned speaker encoder and a learnable Lombard encoder to form a personalized style embedding. Furthermore, an En-Lomb-Based injection module is designed to accurately integrate the target Lombard effect and speaker identity into the linguistic content based on personalized style embedding, ensuring complete Lombard conversion. Experimental results demonstrate that our proposed method outperforms existing N2L models in speech intelligibility, acoustic similarity, and speech quality. Ablation studies confirm that the fine-tuned speaker encoder and the De-Lomb block effectively improve speech intelligibility and acoustic similarity, while the En-Lomb block enables the converted speech to more closely match the target Lombard speech. Hongyang Chen 0004, Yuhong Yang 0001, Xinmeng Xu, Weiping Tu, Zhongyuan Wang 0001, Cedar Lin |
ICME | 6 |
| 2025 | An Improved CenterNet2 Model for Long-Arm Engineering Vehicle Object DetectionabstractBecause the power grid system is easily damaged by external forces from construction vehicles, performing early warning of construction vehicles through object detection is of practical significance for protecting the safety of the power grid. However, accurate detection of engineering vehicles under complex and diverse scenarios is a challenging task. When the existing advanced anchor-free probabilistic two-stage detector CenterNet2 is directly applied to this task, the effect is unsatisfactory. The essential reason is that construction vehicles such as excavators, tower cranes, and truck cranes are different from ordinary objects. They usually have stretchable long robotic arms, resulting in serious deviation of the center point of the bounding box and excessive coverage of the background area by the detection bounding box. Therefore, this paper has made beneficial improvements to two-stage CenterNet2 model. Firstly, it proposes an adaptive positive sample selection algorithm based on Gaussian kernel function to improve the insufficient and unreasonable positive and negative sample sampling in the first-stage regional candidate network. Secondly, it furthers proposes a joint prediction of location accuracy and classification confidence with the IoU-aware category labels, which improves the bounding box ranking during the second-stage network and prevents some potential prediction results being mistakenly filtered out in the post-processing stage. The experimental results on our self-built engineering vehicle dataset indicate that the proposed method effectively alleviates the above problems and achieves promising results in actual engineering vehicle detection. Jie Hua 0005, Zhongyuan Wang 0001, Feng Tian 0006 |
IJCNN | 3 |
| 2025 | Cross-Modal Integrative Feature Network for Sketch-based 3D Shape RetrievalabstractThis paper proposes a novel neural network architecture dubbed Cross-Modal Integrative Feature Network (CMIFN) to address three challenges on sketch-based 3D shape retrieval. Firstly, existing methods, like those based on multi-view CNNs, mostly capture surface visual features, ignoring internal geometry features. CMIFN integrates both multi-view and geometry features of 3D objects, consequently extracting a comprehensive global feature. Secondly, existing methods often manipulate sketches to enhance them, which may introduce superfluous data. Utilising an attention mechanism, CMIFN keeps redundancy in check while achieving a more accurate sketch representation. Thirdly, existing methods often compare the distance between sketches and 3D shapes in the same feature space without considering their inherent differences, which can lead to sub-optimal retrieval results. CMIFN introduces a modality-weighted classifier module, which assigns different weights to features from different modalities, creating a shared feature space to minimize the gap between similar objects across modalities thus increase the retrieval accuracy. Our comprehensive experiments have demonstrated CMIFN’s state-of-the-art performance on benchmark datasets. Xiaoheng Li, Feng Tian 0006, Jinyuan Jia 0002, Zhongyuan Wang 0001 |
IJCNN | 7 |
| 2025 | Quantum Interference-Inspired Who-What-Where Composite-Semantics Instance Search for Story VideosabstractThe Who-What-Where (3W) composite-semantics video Instance Search (INS) task aims to find video shots about a person doing an action in a location. The state-of-the-art (SOTA) methods decompose 3W INS into three 2W INS, i.e., who-what, what-where and where-who semantic correlation modeling, and directly multiply three 2W INS results to produce the final 3W INS result. Obviously, overlapping semantics exist among the above 2Ws, e.g., who-what and what-where share the action component. The semantic overlap indicates that the 2Ws are mutually interdependent rather than independent. According to probability theory, the product of interdependent variables cannot be directly multiplied to obtain an accurate result, and such a direct product would yield a suboptimal outcome. This interdependence exerts diverse influences on the 3W INS results. For instance, fusing two 2W INS results ''Dr. Kelleher-provide medical guidance'' and ''provide medical guidance-in the hospital'', ''provide medical guidance'' is a pivotal connection, of positively enhancing the rationality of both person and location. Conversely, while both ''Ross-lifts heavy objects'' and ''lift heavy objects-Ross'' are individually coherent, combining them by overlapping the shared element ''Ross'' creates a conflict between the hazardous setting and strenuous labor, ultimately undermining the overall plausibility. Inspired by quantum interference theory, we propose a Quantum Interference Partial Decomposition (QIPD) method to model the diverse influences of semantic overlap from 2W to 3W INS. Specifically, QIPD incorporates two core modules, i.e., semantic interference and temporal interference. The former derives the 3W amplitude by converting 2W samples into amplitudes and phases and performing interference, while the latter sets the current shot's phase as baseline, amplifying the influence of adjacent shots while attenuating distant shots. Extensive evaluations on three large-scale 3W INS datasets demonstrate that QIPD outperforms SOTA baselines. Zijun Xu 0001, Chunjie Zhang 0001, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
ACM Multimedia | 4 |
| 2025 | Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation LearningabstractTemporal forgery in multimedia-where audio or video streams are subtly manipulated-poses critical challenges for content authenticity verification. While video-level detection has advanced, Temporal Forgery Localization (TFL) remains underexplored, often limited by weak audio-visual modeling and reliance on non-learnable post-processing. To address these challenges, we propose RegQAV, a Register-enhanced Query-based Audio-Visual framework for TFL. RegQAV exploits pretrained foundation models to capture fine-grained audio-visual correspondences and learnable registers are introduced to mitigate the model's tendency to overly focus on a limited set of temporal features. A query-based localization strategy enables end-to-end optimization without post-processing. We also introduce a Modality Fusion Adapter (MFA) for effective multi-scale integration of audio-visual data, a Deepfake Queries Generation (DQG) module for efficient query initialization, and a Poisson Count-Based Approach to dynamically predict the number of forgeries. Experiments on LAV-DF and AV-Deepfake1M show that RegQAV achieves state-of-the-art performance with fewer parameters, faster inference, and stronger generalization. This work offers significant potential for real-time deepfake detection and other multimedia verification applications. The code is available at https://github.com/zxd3099/RegQAV. Suting Wang, Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001 |
ACM Multimedia | 6 |
| 2025 | A Robust 3D CNN with Pyramidal Attention for Spatiotemporal Gait RecognitionabstractGait recognition has become an increasingly important biometric technique for identifying individuals from a distance without requiring their active cooperation. Since gait involves a sequence of motion patterns, effectively capturing temporal dynamics is essential for accurate recognition. Traditional methods that extract temporal features independently and fuse them at a later stage often fail to model the continuity and interdependence of motion across frames. To overcome this limitation, we propose a novel three-dimensional convolutional architecture named Robust Spatiotemporal 3D Convolutional Neural Network (RST3D), which jointly captures spatial and temporal correlations throughout gait sequences. The proposed architecture incorporates a comprehensive 3D convolutional block that operates along the temporal, height, and width dimensions, enabling the network to learn more expressive and coherent spatiotemporal representations. In addition, we introduce a Temporal Pyramidal Attention (TPA) block to enhance the network’s ability to model temporal dependencies by capturing discriminative motion patterns across multiple temporal scales. We evaluate our method on four large-scale gait recognition datasets: CASIA-B, OUMVLP, GREW, and Gait3D. Experimental results show that our approach consistently achieves superior performance compared to existing 3D CNN-based methods, particularly under challenging conditions such as view variation, clothing changes, and occlusion. Jianyu Chen 0008, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Zengmin Xu, Gang Wu 0010, Zhongyuan Wang 0001 |
MMAsia | 7 |
| 2025 | Multi-Modal Gait Recognition via Collaborative Feature Learning from Silhouettes and Skeletons
Jianyu Chen 0008, Zhongyuan Wang 0001, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Gang Wu 0010 |
PRCV (15) | 2 |
| 2025 | Efficient Extended Neighborhoods Dynamic Selection Re-Ranking for Person Re-IdentificationabstractPerson re-identification (re-ID) is a challenging retrieval task that requires matching a person’s captured images across non-overlapping camera views, with re-ranking being a critical step for improving accuracy. The k-nearest neighbors relationship is commonly used to determine the rank results by selecting only k fixed pedestrian image values for distance calculations. This operation, however, generates additional distance errors due to changes in the appearance of pedestrians. This paper addresses the above issue by proposing a simple but effective Extended Neighborhood Dynamic Selection (ENDS), distance to optimize the performance of ReID reranking. The number of images selected is then distributed over an interval. A limit of upper and lower is placed on the selected number to ensure that it is neither too low nor too high. The automatic selection of adjacent images is achieved using this method. This distance is determined by combining ENDS distances with Jaccard distances. It is the core principle of this method that, instead of using fixed values, the choice of the number of images to be included in each neighborhood should be made automatically. It also allows the removal of images that are dissimilar in favour of those that are more representative. Experimental results demonstrate the novel method of ranking by using Market-1501 and DukeMTMC’s reID dataset. In this paper, we propose a method that increases the Market-1501 mAP/Rank1 by 29.4%/12.9% while DukeMTMC-reID reranking by 35%/20.7%. Chao Wang 0084, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Mithun Mukherjee 0001 |
SMC | 2 |
| 2025 | Multi-receptive field interaction network for shape from polarization
Yini Peng, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
Sci. China Inf. Sci. | 4 |
| 2025 | Adversarial intensity awareness for robust object detection
Jikang Cheng, Baojin Huang, Zhen Han 0002, Zhongyuan Wang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2025 | Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques
Junqi Yang, Yuhong Yang 0001, Weiping Tu, Zhongyuan Wang 0001 |
Neural Networks | 5 |
| 2025 | Unsupervised learning non-uniform face enhancement under physics-guided model of illumination decoupling
Zhongyuan Wang 0001, Qiong Liu 0001, You Yang 0002, Zhenyu Shu |
Pattern Recognit. | 3 |
| 2025 | Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong |
Pattern Recognit. | 2 |
| 2025 | ED4: Explicit Data-Level Debiasing for Deepfake DetectionabstractLearning intrinsic bias from limited data has been considered the main reason for the failure of deepfake detection with generalizability. Apart from the discovered content and specific-forgery bias, we reveal a novel spatial bias, where detectors inertly anticipate observing structural forgery clues appearing at the image center, also can lead to the poor generalization of existing methods. We present ED4, a simple and effective strategy, to address aforementioned biases explicitly at the data level in a unified framework rather than implicit disentanglement via network design. In particular, we develop ClockMix to produce facial structure preserved mixtures with arbitrary samples, which allows the detector to learn from an exponentially extended data distribution with much more diverse identities, backgrounds, local manipulation traces, and the co-occurrence of multiple forgery artifacts. We further propose the Adversarial Spatial Consistency Module (AdvSCM) to prevent extracting features with spatial bias, which adversarially generates spatial-inconsistent images and constrains their extracted feature to be consistent. As a model-agnostic debiasing strategy, ED4 is plug-and-play: it can be integrated with various deepfake detectors to obtain significant benefits. We conduct extensive experiments to demonstrate its effectiveness and superiority over existing deepfake detection approaches. Code is available at https://github.com/beautyremain/ED4. Jikang Cheng, Ying Zhang 0021, Qin Zou 0001, Zhiyuan Yan 0002, Chao Liang 0001, Zhongyuan Wang 0001, Chen Li 0031 |
IEEE Trans. Image Process. | 6 |
| 2025 | Who, What, and Where: Composite-Semantics Instance Search for Story VideosabstractWho, What and Where (3W)are the three core elements of storytelling, and accurately identifying the 3W semantics is critical to understanding the story in a video. This paper studies the 3W composite-semantics video Instance Search (INS) problem, which aims to find video shots about a specific person doing a concrete action in a particular location. The popular Complete-Decomposition (CD) methods divide a composite-semantics query into multiple single-semantics queries, which are likely to yield inaccurate or incomplete retrieval results due to neglecting important semantic correlations. Recent Non-Decomposition (ND) methods utilize Vision Language Model (VLM) to directly measure the similarity between textual query and video content. However, the accuracy is limited by VLM's immature capability to recognize fine-grained objects. To address the above challenges, we propose a video structure-aware Partial-Decomposition (PD) method. Its core idea is to partially decompose the 3W INS problem into three semantic-correlated 2W INS problems i.e., person-action INS, action-location INS, and location-person INS. Thereafter, we respectively model the correlations between pairs of semantics at frames, shots and scenes of story videos. With the help of the spatial consistency and temporal continuity contained in the unique hierarchical structure of story videos, we can finally obtain identity-matching, logic-consistent, and content-coherent 3W INS results. To validate the effectiveness of the proposed method, we specifically build three large-scale 3W INS datasets based on three TV series Eastenders, Friends and The Big Bang Theory, totally comprising over 670K video shots spanning 700 hours. Extensive experiments show that the proposed PD method surpasses the current state-of-the-art CD and ND methods for 3W INS in story videos. Ankang Lu, Zhengqian Wu, Zhongyuan Wang 0001, Chao Liang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Multi-Stage Statistical Texture-Guided GAN for Tilted Face FrontalizationabstractExisting pose-invariant face recognition mainly focuses on frontal or profile, whereas high-pitch angle face recognition, prevalent under surveillance videos, has yet to be investigated. More importantly, tilted faces significantly differ from frontal or profile faces in the potential feature space due to self-occlusion, thus seriously affecting key feature extraction for face recognition. In this paper, we asymptotically reshape challenging high-pitch angle faces into a series of small-angle approximate frontal faces and exploit a statistical approach to learn texture features to ensure accurate facial component generation. In particular, we design a statistical texture-guided GAN for tilted face frontalization (STG-GAN) consisting of three main components. First, the face encoder extracts shallow features, followed by the face statistical texture modeling module that learns multi-scale face texture features based on the statistical distributions of the shallow features. Then, the face decoder performs feature deformation guided by the face statistical texture features while highlighting the pose-invariant face discriminative information. With the addition of multi-scale content loss, identity loss and adversarial loss, we further develop a pose contrastive loss of potential spatial features to constrain pose consistency and make its face frontalization process more reliable. On this basis, we propose a divide-and-conquer strategy, using STG-GAN to progressively synthesize faces with small pitch angles in multiple stages to achieve frontalization gradually. A unified end-to-end training across multiple stages facilitates the generation of numerous intermediate results to achieve a reasonable approximation of the ground truth. Extensive qualitative and quantitative experiments on multiple-face datasets demonstrate the superiority of our approach. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Chao Liang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | ReID-FSAI: Person Re-Identification Network Fused With Semantic and Attribute InformationabstractPersonal belongings information (e.g., backpacks and reticules) and attribute descriptions (e.g., gender and age) provide critical discriminative cues for person re-identification (Re-ID) tasks. However, existing Re-ID algorithms leveraging additional semantic models often fail to accurately recognize personal belongings and suffer from noisy attribute predictions derived from global or local features, as they inadequately exploit attribute correlations. To address these challenges, we propose a novel person re-identification network, ReID-FSAI, which fuses personal belongings information and attribute descriptions from isolated semantic regions. ReID-FSAI integrates personal belongings areas identified through feature clustering with semantic parsing results from an auxiliary semantic model. By treating the generated semantic regions as body labels, our network refines global features into precise semantic features and accurately predicts attribute information from these regions. Furthermore, ReID-FSAI employs a reweighting model to enhance the confidence in specific attributes, improving attribute prediction accuracy. By combining predictions of attributes and personal belongings with global features, our approach significantly improves the representation ability of pedestrians. Experimental evaluations on the Market-1501 and DukeMTMC-reID datasets demonstrate that ReID-FSAI achieves superior performance in both person re-ID and attribute prediction, surpassing state-of-the-art methods. Jinsheng Xiao, Qiuze Yu, Zhongyuan Wang 0001, Yuan-Fang Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | DALFace: Dynamic Association Learning for Face RecognitionabstractFace recognition owes its success to the availability of large-scale training data. Recent adaptive margin-based loss functions pay more attention to hard (misclassified) samples, resulting in more discriminative face embeddings. However, large-scale datasets inevitably include open-set noise samples, which are usually mistaken for hard samples by mining-based methods and thus mislead the training of the model. In this work, we redefine hard samples and further design a dynamic association learning strategy for mining hard samples while ignoring noise. We argue that the difficulty of recognizing a sample depends on both identity-related and objective factors. On one hand, intrinsic attributes such as facial structure and face shape inherently influence the ease of identity recognition. On the other hand, external factors, including pose, occlusion, and resolution, directly affect the recognizability of a sample. Particularly in the case of noise samples, although they pose challenges for the deep network similar to hard samples, should not be regarded as hard samples. To this end, we propose an associated prototype learning method to achieve an approximation of face identity difficulty by exploring the fitting trends of identity prototype. Furthermore, we design a dynamic sample learning method to distinguish noise samples from hard samples by observing the distance fluctuation from the class center during sample learning. All observations are integrated into the loss function through adaptive margins and sample weights. Extensive experiments and visualizations on several datasets demonstrate that our method significantly outperforms state-of-the-art counterparts. Baojin Huang, Guangcheng Wang, Kui Jiang, Zhongyuan Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Adaptive Clustering and Weighted Regularization Contrastive Learning Framework for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID) has recently gained significant attention from researchers. ReID matches images of the same person from different camera views in various scenes without any labels. Existing clustering methods primarily rely on a fixed threshold (the maximum distance between sample points and clustering centroids) and overlook the importance of adjusting this threshold during continuous model optimization. This mismatch between clustering thresholds and inter- or intra-class spacing reduces clustering accuracy. To address this issue, this study proposes an Adaptive Clustering and Weighted Regularization Contrastive Learning (ACWRCL) framework for unsupervised person ReID. The ACWRCL framework comprises two main components: (1) the Clustering Threshold Adaptive Adjustment (CTAA) module, and (2) the Weighted Regularization Contrastive Learning (WRCL) module. The CTAA module dynamically adjusts the clustering threshold to align with model optimization, ensuring that the threshold remains within an appropriate range to prevent under- or over-robustness in the clustering model. The WRCL module uses the similarity ratio between the query sample and the clustering centroid relative to the overall similarity of all samples with the same labels as the query sample. This ratio is used as the weight in the loss function to penalize incorrect clustering and improve pseudo-label generation accuracy. Extensive experiments on public ReID datasets—Market-1501, MSMT17, Veri776, CUHK03, and PersonX—demonstrate the effectiveness of the proposed method. Mingfu Xiong, Kaikang Hu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | DiffusionMOT: A Diffusion-Based Multiple Object TrackerabstractRecently, researchers have introduced diffusion models into multiple object tracking (MOT) tasks. However, existing diffusion-based MOT methods, such as DiffusionTrack, have significant limitations, including frequent ID switching, reduced performance when tracking nonlinear motion objects, and long inference time. To this end, we propose a more effective diffusion-based multiple object tracker named DiffusionMOT. In particular, we propose a mixed intersection over union (IoU) and Re-Identification (ReID) method for trajectory matching, which effectively reduces incorrect matches. Meanwhile, we propose a secondary calibration method for trajectory boxes, improving the accuracy of the generated detection boxes. Moreover, we introduce the parallel sampling technique from the field of image generation into object tracking and propose a parallel sampling module to enhance the model's inference speed while maintaining tracking accuracy. Furthermore, we design a pair-based two-stage matching (PTM) pipeline to more effectively utilize potential detection information. Extensive experiments on several public MOT benchmarks, including DanceTrack, SportsMOT, MOT20, and MOT17, demonstrate that our approach achieves state-of-the-art (SOTA) performance. The code and models are available at https://github.com/sad123-yx/DiffusionMOT. Yaxuan Hu 0001, Jie Hua 0005, Zhen Han 0002, Hua Zou 0002, Gang Wu 0010, Zhongyuan Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Optimal Illumination Distance Metrics for Person Re-Identification in Complex Lighting ConditionsabstractPerson re-identification is extensively applied in public security and surveillance. However, environmental factors like time and location often lead to varying lighting conditions in captured pedestrian images, significantly impacting identification accuracy. Current approaches mitigate this issue through lighting transformation techniques, aiming to normalize images to a standard lighting condition for consistent person re-identification results. Yet, these methods overlook the fact that different content may hold distinct identification values under diverse lighting conditions. To address this, we conducted an analysis on the identification distance between images of the same or different pedestrians under pre-defined lighting conditions. From this analysis, we introduce the concept of optimal lighting: a condition where the distance between image pairs is minimized compared to other lighting scenarios. We propose utilizing this optimal lighting distance in the image retrieval process for final ranking. Our study, validated on synthetic datasets Market-IA and Duke-IA, demonstrates that optimal lighting is independent of image texture information. Each image pair exhibits a unique optimal lighting, yet consistently shows a minimum distance value. Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Dual-Domain Multi-Model GAN Fingerprint Restoration for Compressed Fake Face AttributionabstractRecent advances in GAN fingerprint have shown increasing success in fake face attribution. However, the fake faces are usually compressed during network transmission, which causes the degradation of GAN fingerprint and the decrease of attribution accuracy. To this issue, a dual-domain multi-model GAN fingerprint restoration method for compressed fake face attribution is proposed in this paper. Firstly, considering that image-domain and fingerprint-domain are directly and indirectly affected by compression respectively, we propose a dual-domain parallel restoration architecture that enhances GAN fingerprint using direct image-domain and indirect fingerprint-domain restoration, thereby improving attribution performance by mining the cross-domain complementarity. Secondly, since real and fake GAN-speciffc restoration models can describe GAN fingerprint from different aspects, we first enhance GAN fingerprint by multiple restoration models, and then improve attribution performance by exploiting the cross-model complementarity through the multi-model restoration fusion strategy. Experiments demonstrate the superiority of our method under different compression qualities. Chengxiang Fan, Aohong Shen, Zhen Han 0002, Cai Tong, Zhongyuan Wang 0001, Dekang Yi |
ICME | 5 |
| 2024 | XFusion: Cross-Attention Transformer for Multi-focus Image Fusion
Shouxi Zhao, Qin Zou 0001, Chi Chen 0002, Zhongyuan Wang 0001 |
ICONIP (7) | 5 |
| 2024 | Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
Hongyang Chen 0004, Yuhong Yang 0001, Zhongyuan Wang 0001, Weiping Tu, Haojun Ai, Cedar Lin |
INTERSPEECH | 3 |
| 2024 | Refining Intraocular Lens Power Calculation: A Multi-modal Framework Using Cross-Layer Attention and Effective Channel Attention
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (1) | 3 |
| 2024 | TalkSee: Interactive Video Retrieval Engine Using Large Language Model
Guihe Gu, Zhengqian Wu, Jiangshan He, Zhongyuan Wang 0001, Chao Liang 0001 |
MMM (4) | 5 |
| 2024 | Can We Leave Deepfake Data Behind in Training Deepfake Detector?abstractThe generalization ability of deepfake detectors is vital for their applications in real-world scenarios. One effective solution to enhance this ability is to train the models with manually-blended data, which we termed ''blendfake'', encouraging models to learn generic forgery artifacts like blending boundary. Interestingly, current SoTA methods utilize blendfake $\textit{without}$ incorporating any deepfake data in their training process. This is likely because previous empirical observations suggest that vanilla hybrid training (VHT), which combines deepfake and blendfake data, results in inferior performance to methods using only blendfake data (so-called "1+1<2"). Therefore, a critical question arises: Can we leave deepfake behind and rely solely on blendfake data to train an effective deepfake detector? Intuitively, as deepfakes also contain additional informative forgery clues ($\textit{e.g.,}$ deep generative artifacts), excluding all deepfake data in training deepfake detectors seems counter-intuitive. In this paper, we rethink the role of blendfake in detecting deepfakes and formulate the process from "real to blendfake to deepfake" to be a $\textit{progressive transition}$. Specifically, blendfake and deepfake can be explicitly delineated as the oriented pivot anchors between "real-to-fake" transitions. The accumulation of forgery information should be oriented and progressively increasing during this transition process. To this end, we propose an $\underline{O}$riented $\underline{P}$rogressive $\underline{R}$egularizor (OPR) to establish the constraints that compel the distribution of anchors to be discretely arranged. Furthermore, we introduce feature bridging to facilitate the smooth transition between adjacent anchors. Extensive experiments confirm that our design allows leveraging forgery information from both blendfake and deepfake effectively and comprehensively. Code is available at https://github.com/beautyremain/ProDet. Jikang Cheng, Zhiyuan Yan 0002, Ying Zhang 0021, Yuhao Luo 0002, Zhongyuan Wang 0001, Chen Li 0031 |
NeurIPS | 5 |
| 2024 | SF-Gait: Two-Stage Temporal Compression Network for Learning Gait Micro-Motions and Cycle Patterns
Yuanhao Yue, Yunhe Wang 0011, Laixiang Shi, Zhongyuan Wang 0001, Qin Zou 0001 |
PRCV (15) | 4 |
| 2024 | Optimal Illumination Distance Metrics for Person Re-identification
Chao Wang 0084, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001, Wen Zhou 0029 |
PRICAI (4) | 2 |
| 2024 | Bidirectional scale-aware upsampling network for arbitrary-scale video super-resolution
Laigan Luo, Benshun Yi, Zhongyuan Wang 0001, Zheng He 0001 |
Image Vis. Comput. | 3 |
| 2024 | Domain generalized person reidentification based on skewness regularity of higher-order statistics
Mingfu Xiong, Ruimin Hu, Zhongyuan Wang 0001, Javier Del Ser, Khan Muhammad 0001, Zixiang Xiong |
Knowl. Based Syst. | 4 |
| 2024 | Multi-scale motion contrastive learning for self-supervised skeleton-based action recognition
Yushan Wu, Zengmin Xu, Mengwei Yuan, Tianchi Tang, Ruxing Meng, Zhongyuan Wang 0001 |
Multim. Syst. | 6 |
| 2024 | Efficient lightweight network for video super-resolution
Laigan Luo, Benshun Yi, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001 |
Neural Comput. Appl. | 3 |
| 2024 | Re-decoupling the classification branch in object detectors for few-class scenes
Jie Hua 0005, Zhongyuan Wang 0001, Qin Zou 0001, Jinsheng Xiao, Xin Tian 0006 |
Pattern Recognit. | 2 |
| 2024 | Deep motion estimation through adversarial learning for gait recognition
Yuanhao Yue, Laixiang Shi, Long Chen 0005, Zhongyuan Wang 0001, Qin Zou 0001 |
Pattern Recognit. Lett. | 5 |
| 2024 | Unlabeled Data Assistant: Improving Mask Robustness for Face RecognitionabstractThe existing masked face recognition algorithms almost tend to adopt synthetic masked face datasets for training. However, these models are limited as they rely on existing mask augmentation methods, which contain few mask patterns and cannot simulate shadows and textures in realistic scenes. To overcome this limitation, we propose a semi-supervised face recognition framework to fully exploit unlabeled real masked face samples, improving the mask robustness of the recognition model. More specifically, unlike the original face embedding network, we design a part-aware network to explore multi-region face representation based on the face structure. In this way, we obtain multiple face sub-embeddings, which correspond to different regions of the face, including the upper half, the lower half and the whole. Crucially, we use the norm of the sub-embedding to represent the activation state of the facial region features. For the input unlabeled masked face image, we restrict the sub-embedding norm of its lower half to weaken the face feature representation of the occluded area. For normal face samples, their partial features are kept activated by maintaining the sub-embedding norm, which guides the deep network does not ignore the available information. Moreover, we employ the margin-based recognition loss for normal samples to ensure that the model is sufficiently discriminative for normal facial features. Extensive experimental results on both normal and real masked face datasets show that our approach significantly outperforms the state-of-the-arts. Code is available at https://github.com/Baojin-Huang/UFace. Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Zhen Han 0002, Chao Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Toward Robust Adversarial Purification for Face Recognition Under Intensity-Unknown AttacksabstractRecent years have witnessed dramatic progress in adversarial attacks, which can easily mislead face recognition systems via the injection of imperceptible perturbations on the input image. Many defense methods have been proposed to mitigate the detrimental impact of adversarial attacks, including adversarial purification which intends to reconstruct clean images through a generative model. This paper studies a more practical and challenging problem: how to defend face recognition systems against intensity-unknown or even intensity-varying adversarial attacks? We attempt to crack this tough nut from the dimensionality of input resolutions. Looking into the performance of purification methods with various input resolutions, we reveal a phenomenon that, higher-resolution input images help better defend against weaker attacks, while lower-resolution ones are naturally defensive against stronger attacks. It inspires us to design an adaptive purification framework under intensity-unknown attacks, dubbed adversarial Intensity-guided Multi-scale Attention (IMA). Via the aggregation of information from different resolution scales and flexible adjustment according to an estimation of adversarial intensity, it leverages the respective advantages of different scales and constructs a robust ensemble against intensity-unknown attacks. We validate the superiority of IMA by defending against both face obfuscation and impersonation of 9 typical attack algorithms under gray-box, white-box and black-box evaluation, outperforming state-of-the-art defense methods on LFW and YTF datasets. Keyizhi Xu, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Deep Hashing Network With Hybrid Attention and Adaptive Weighting for Image RetrievalabstractDue to the low computational cost of Hamming distance, hashing-based image retrieval has been universally acknowledged. Therefore, it is becoming increasingly important to quickly generate high-precision hash codes (also hash features) from images. However, the existing deep hashing methods are vulnerable to image content variations; that is, it is difficult to generate stable and consistent hash codes for similar images. In addition, generating hash codes of different lengths requires retraining the model, which is expensive in training time. To address these problems, this paper proposes a deep hashing network (DHN) with a hybrid attention mechanism and adaptive weighting (HAAW) learning. It mainly consists of a feature extraction module, feature refinement module, classification layer, hash layer and an adaptive weight layer. In particular, the hybrid attention mechanism combines bottom-up pixel saliency and top-down semantic constraints, in which the former is achieved through channel and spatial attention (CSA) and the latter is supervised by classification labels. In this way, it encourages the network to focus on dominant semantic features without being disturbed by irrelevant objects so that semantically similar images can be mapped to approximate hash codes. We further propose an adaptive weighting learning algorithm to generate weights for each bit of the hash code generated by the deep network. Then, we directly generate shorter hash codes from the available long hash code according to the importance of bits represented by the weights. This avoids retraining the network for learning hash codes of different lengths. Extensive experiments on public CIFAR-10, NUS_WIDE and ImageNet datasets show that our method has achieved substantial improvements over the counterparts in terms of precision and speed. Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu |
IEEE Trans. Multim. | 2 |
| 2024 | Rethinking Prior-Guided Face Super-Resolution: A New Paradigm With Facial Component PriorabstractRecently, facial priors (e.g., facial parsing maps and facial landmarks) have been widely employed in prior-guided face super-resolution (FSR) because it provides the location of facial components and facial structure information, and helps predict the missing high-frequency (HF) information. However, most existing approaches suffer from two shortcomings: 1) the extracted facial priors are inaccurate since they are extracted from low-resolution (LR) or low-quality super-resolved (SR) face images and 2) they only consider embedding facial priors into the reconstruction process from LR to SR face images, thus failing to explore facial priors to generate LR face image. In this article, we propose a novel pre-prior guided approach that extracts facial prior information from original high-resolution (HR) face images and embeds them into LR ones to obtain HF information-rich LR face images, thereby improving the performance of face reconstruction. Specifically, a novel component hybrid method is proposed, which fuses HR facial components and LR facial background to generate new LR face images (namely, LRmix) via facial parsing maps extracted from HR face images. Furthermore, we design a component hybrid network (CHNet) that learns the LR to LRmix mapping function to ensure that the LRmix can be obtained from LR face images in testing and real-world datasets. Experimental results show that our proposed scheme significantly improves the reconstruction performance for FSR. Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Junjun Jiang, Zhongyuan Wang 0001, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Interpretable Model-Driven Deep Network for Hyperspectral, Multispectral, and Panchromatic Image FusionabstractSimultaneously fusing hyperspectral (HS), multispectral (MS), and panchromatic (PAN) images brings a new paradigm to generate a high-resolution HS (HRHS) image. In this study, we propose an interpretable model-driven deep network for HS, MS, and PAN image fusion, called HMPNet. We first propose a new fusion model that utilizes a deep before describing the complicated relationship between the HRHS and PAN images owing to their large resolution difference. Consequently, the difficulty of traditional model-based approaches in designing suitable hand-crafted priors can be alleviated because this deep prior is learned from data. We further solve the optimization problem of this fusion model based on the proximal gradient descent (PGD) algorithm, achieved by a series of iterative steps. By unrolling these iterative steps into several network modules, we finally obtain the HMPNet. Therefore, all parameters besides the deep prior are learned in the deep network, simplifying the selection of optimal parameters in the fusion and achieving a favorable equilibrium between the spatial and spectral qualities. Meanwhile, all modules contained in the HMPNet have explainable physical meanings, which can improve its generalization capability. In the experiment, we exhibit the advantages of the HMPNet over other state-of-the-art methods from the aspects of visual comparison and quantitative analysis, where a series of simulated as well as real datasets are utilized for validation. Xin Tian 0006, Kun Li 0025, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Person-action Instance Search in Story Videos: An Experimental StudyabstractPerson-Action instance search (P-A INS) aims to retrieve the instances of a specific person doing a specific action, which appears in the 2019–2021 INS tasks of the world-famous TREC Video Retrieval Evaluation (TRECVID). Most of the top-ranking solutions can be summarized with a Division-Fusion-Optimization (DFO) framework, in which person and action recognition scores are obtained separately, then fused, and, optionally, further optimized to generate the final ranking. However, TRECVID only evaluates the final ranking results, ignoring the effects of intermediate steps and their implementation methods. We argue that conducting the fine-grained evaluations of intermediate steps of DFO framework will (1) provide a quantitative analysis of the different methods’ performance in intermediate steps; (2) find out better design choices that contribute to improving retrieval performance; and (3) inspire new ideas for future research from the limitation analysis of current techniques. Particularly, we propose an indirect evaluation method motivated by the leave-one-out strategy, which finds an optimal solution surpassing the champion teams in 2020–2021 INS tasks. Moreover, to validate the generalizability and robustness of the proposed solution under various scenarios, we specifically construct a new large-scale P-A INS dataset and conduct comparative experiments with both the leading NIST TRECVID INS solution and the state-of-the-art P-A INS method. Finally, we discuss the limitations of our evaluation work and suggest future research directions. Yanrui Niu, Chao Liang 0001, Ankang Lu, Baojin Huang, Zhongyuan Wang 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2024 | An Image Arbitrary-Scale Super-Resolution Network Using Frequency-domain InformationabstractImage super-resolution (SR) is a technique to recover lost high-frequency information in low-resolution (LR) images. Since spatial-domain information has been widely exploited, there is a new trend to involve frequency-domain information in SR tasks. Besides, image SR is typically application-oriented and various computer vision tasks call for image arbitrary magnification. Therefore, in this article, we study image features in the frequency domain to design a novel image arbitrary-scale SR network. First, we statistically analyze LR-HR image pairs of several datasets under different scale factors and find that the high-frequency spectra of different images under different scale factors suffer from different degrees of degradation, but the valid low-frequency spectra tend to be retained within a certain distribution range. Then, based on this finding, we devise an adaptive scale-aware feature division mechanism using deep reinforcement learning, which can accurately and adaptively divide the frequency spectrum into the low-frequency part to be retained and the high-frequency one to be recovered. Finally, we design a scale-aware feature recovery module to capture and fuse multi-level features for reconstructing the high-frequency spectrum at arbitrary scale factors. Extensive experiments on public datasets show the superiority of our method compared with state-of-the-art methods. Yinbo Yu, Zhongyuan Wang 0001, Ruimin Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Few-Shot Face Sketch-to-Photo Synthesis via Global-Local Asymmetric Image-to-Image TranslationabstractFace sketch-to-photo synthesis is widely used in law enforcement and digital entertainment, which can be achieved by Image-to-Image (I2I) translation. Traditional I2I translation algorithms usually regard the bidirectional translation of two image domains as two symmetric processes, so the two translation networks adopt the same structure. However, due to the scarcity of face sketches and the abundance of face photos, the sketch-to-photo and photo-to-sketch processes are asymmetric. Considering this issue, we propose a few-shot face sketch-to-photo synthesis model based on asymmetric I2I translation, where the sketch-to-photo process uses a feature-embedded generating network, while the photo-to-sketch process uses a style transfer network. On this basis, a three-stage asymmetric training strategy with style transfer as the trigger is proposed to optimize the proposed model by utilizing the advantage that the style transfer network only needs few-shot face sketches for training. Additionally, we discover that stylistic differences between the global and local sketch faces lead to inconsistencies between the global and local sketch-to-photo processes. Thus, a dual branch of the global face and local face is adopted in the sketch-to-photo synthesis model to learn the specific transformation processes for global structure and local details. Finally, the high-quality synthetic face photo can be generated through the global-local face fusion sub-network. Extensive experimental results demonstrate that the proposed Global-Local Asymmetric (GLAS) I2I translation algorithm compared to SOTA methods, at least improves FSIM by 0.0126, and reduces LPIPS (alex), LPIPS (squeeze), and LPIPS (vgg) by 0.0610, 0.0883, and 0.0719, respectively. Qifan Liang, Zhen Han 0002, Wenjun Mai, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Cross-Modal Face Super-Resolution Based on Quasi-Siamese Domain Transfer Fusion NetworkabstractIn this paper, we propose a Cross-Modal Face Super-Resolution (CMFSR) method to construct high-resolution (HR) facial images from low-resolution (LR) cross-modal facial images captured respectively by disjoint visible light (VIS) and near-infrared (NIR) cameras. Due to the coupling of modality transformation and information fusion, CMFSR is more difficult to obtain HR reconstructed results compared with traditional super-resolution. To solve this problem, a Quasi-Siamese Domain Transfer Fusion Network (QSDTFN) for CMFSR is proposed in this paper, whose two branches transfer two LR face modality to HR face modality by domain transfer respectively. Different from two completely independent branches in the traditional pseudo-siamese network, only the HR-to-LR face transfer processes of the two branches in our quasi-siamese network are independent, while the LR-to-HR face transfer processes are coupled. This coupled module called the Adaptive Weighted Domain Transfer Fusion Module (AWDTFM) disentangles the modality and identity information in the two LR faces, thus achieving modality transformation and identity information fusion simultaneously. In order to strengthen the optimization on the process of CMFSR, this method further introduces the backward QSDTFN to form a higher-level bidirectional structure with the forward QSDTFN, and specifically designs two types of losses: intra-network loss and inter-network loss, to constrain the modality and identity consistencies within one QSDTFN and between two QSDTFNs respectively. The experimental results on the challenging LR cross-modal face datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods. Jiaxing Wen, Aohong Shen, Zhen Han 0002, Zhongyuan Wang 0001, Liang Chen 0026 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Inter-camera Identity Discrimination for Unsupervised Person Re-identificationabstractUnsupervised person re-identification (Re-ID) has garnered significant attention because of its data-friendly nature, as it does not require labeled data. Existing approaches primarily address this challenge by employing feature-clustering techniques to generate pseudo-labels. In addition, camera-proxy-based methods have emerged because of their impressive ability to cluster sample identities. However, these methods often blur the distinctions between individuals within inter-camera views, which is crucial for effective person re-ID. To address this issue, this study introduces an inter-camera-identity-difference-based contrastive learning framework for unsupervised person Re-ID. The proposed framework comprises two key components: (1) a different sample cross-view close-range penalty module and (2) the same sample cross-view long-range constraint module. The former aims at penalizing excessive similarity among different subjects across inter-camera views, whereas the latter mitigates the challenge of excessive dissimilarity among the same subject across camera views. To validate the performance of our method, we conducted extensive experiments on three existing person Re-ID datasets (Market-1501, MSMT17, and PersonX). The results demonstrate the effectiveness of the proposed method, which shows a promising performance. The code is available at https://github.com/hooldylan/IIDCL . Mingfu Xiong, Kaikang Hu, Zhihan Lyu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Joint Distortion Restoration and Quality Feature Learning for No-reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) methods, inspired by the free energy principle, improve the accuracy of image quality prediction by simulating the human brain’s repair process for distorted images. However, existing methods use separate optimization schemes for distortion restoration and quality prediction, which undermines the accurate mapping of feature representations to quality scores. To address this issue, we propose a joint restoration and quality feature learning NR-IQA (RQFL-IQA) method to jointly tackle distortion image restoration and quality prediction within a unified framework. To accurately establish the quality reconstruction relationship between distorted and restored images, a hybrid loss function based on pixel-wise and structure-wise representations is used to improve the restoration capability of the image restoration network. The proposed RQFL-IQA exploits rich labels, including restored images and quality scores, to enable the model to learn more discriminative features and establish a more accurate mapping from feature representation to quality scores. In addition, to avoid the impact of poor restoration on quality prediction, we propose a module with a cleaning function to reweight the fusion of restored and primitive features to achieve more perceptual consistency in feature fusion. Experimental results on public IQA datasets show that the proposed RQFL-IQA is superior over existing methods. Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Zixiang Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Auxiliary Information Guided Self-attention for Image Quality AssessmentabstractImage quality assessment (IQA) is an important problem in computer vision with many applications. We propose a transformer-based multi-task learning framework for the IQA task. Two subtasks: constructing an auxiliary information error map and completing image quality prediction, are jointly optimized using a shared feature extractor. We use visual transformers (ViT) as a feature extractor for feature extraction and guide ViT to focus on image quality-related features by building auxiliary information error map subtask. In particular, we propose a fusion network that includes a channel focus module. Unlike the fusion methods commonly used in previous IQA methods, we use the fusion network, including the channel attention module, to fuse the auxiliary information error map features with the image features, which facilitates the model to mine the image quality features for more accurate image quality assessment. And by jointly optimizing the two subtasks, ViT focuses more on extracting image quality features and building a more precise mapping from feature representation to quality score. With slight adjustments to the model, our approach can be used in both no-reference (NR) and full-reference (FR) IQA environments. We evaluate the proposed method in multiple IQA databases, showing better performance than state-of-the-art FR and NR IQA methods. Jifan Yang, Zhongyuan Wang 0001, Guangcheng Wang, Baojin Huang, Yuhong Yang 0001, Weiping Tu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Implicit Identity Driven Deepfake Face Swapping DetectionabstractIn this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit identity and implicit identity, which respectively corresponds to the identity of the source face and target face during face swapping. Note that the explicit identities of faces can be extracted by regular face recognizers. Particularly, the implicit identity of real face is consistent with the its explicit identity. Thus the difference between explicit and implicit identity of face facilitates face swapping detection. Following this idea, we propose a novel implicit identity driven framework for face swapping detection. Specifically, we design an explicit identity contrast (EIC) loss and an implicit identity exploration (IIE) loss, which supervises a CNN backbone to embed face images into the implicit identity space. Under the guidance of EIC, real samples are pulled closer to their explicit identities, while fake samples are pushed away from their explicit identities. More-over, IIE is derived from the margin-based classification loss function, which encourages the fake faces with known target identities to enjoy intra-class compactness and inter-class diversity. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art counterparts. Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Jiaxin Ai, Qin Zou 0001, Qian Wang 0002, Dengpan Ye |
CVPR | 2 |
| 2023 | LSTFE-Net: Long Short-Term Feature Enhancement Network for Video Small Object DetectionabstractVideo small object detection is a difficult task due to the lack of object information. Recent methods focus on adding more temporal information to obtain more potent high-level features, which often fail to specify the most vital information for small objects, resulting in insufficient or inappropriate features. Since information from frames at different positions contributes differently to small objects, it is not ideal to assume that using one universal method will extract proper features. We find that context information from the long-term frame and temporal information from the short-term frame are two useful cues for video small object detection. To fully utilize these two cues, we propose a long short-term feature enhancement network (LSTFE-Net) for video small object detection. First, we develop a plug-and-play spatiotemporal feature alignment module to create temporal correspondences between the short-term and current frames. Then, we propose a frame selection module to select the long-term frame that can provide the most additional context information. Finally, we propose a long short-term feature aggregation module to fuse long short-term features. Compared to other state-of-the-art methods, our LSTFE-Net achieves 4.4% absolute boosts in AP on the FL-Drones dataset. More details can be found at https://github.com/xiaojs18/LSTFE-Net. Jinsheng Xiao, Yuanxu Wu, Yunhua Chen, Zhongyuan Wang 0001, Jiayi Ma 0001 |
CVPR | 5 |
| 2023 | LSA3D: Lightweight Separate Asynchronous 3D Convolutional Neural Network for Gait Recognition
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Jinsheng Xiao, Zhen Han 0002 |
ICANN (10) | 2 |
| 2023 | Multi-frame Tilt-angle Face Recognition Using Fusion Re-ranking
Wenqin Song, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001 |
ICANN (2) | 4 |
| 2023 | Continuous Learning for Blind Image Quality Assessment with Contrastive TransformerabstractMost existing blind image quality assessment (BIQA) models focus on improving performance on existing datasets and are weak in adapting to unknown distortion or degradation types. In this paper, we propose a Transformer-based BIQA contrastive continual learning approach to improve model transfer performance. The basic idea is that the model continuously learns from the IQA data stream, integrating new knowledge from the current dataset. At the same time, limited access to previous data using a limited memory budget prevents forgetting the knowledge gained from the dataset of old tasks. We design an attentional contrastive learning strategy based on the Transformer architecture with a designed attentional focus contrastive loss to rebalance the contrastive learning between the new and old tasks, which can consolidate the previously learned representations. In addition, we used a structure similar to the cumulative classifier, balancing the learning of the current task quality score with the quality scores of all observed tasks. Extensive experiments demonstrate the feasibility of the proposed continuous learning approach compared to the standard training techniques of BIQA. Jifan Yang, Zhongyuan Wang 0001, Baojin Huang |
ICASSP | 2 |
| 2023 | Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super ResolutionabstractRecently, face super-resolution has achieved pleasing performance. Numerous works have shown that texture features and structural information play a crucial role for super-resolution reconstruction. However, effective co-learning of both has been limiting the performance improvement of existing state-of-the-art methods. Therefore, we focus on the texture and structure of images in this paper, and design a two-branch network containing a texture network (T-Net) and a structure network (S-Net) to jointly explore texture and structure information for co-learning. T-Net serves as the backbone network to learn both texture and structure information for reconstruction, while the S-Net serves as the auxiliary network that can effectively exploit the multi-scale information of the T-Net encoder to recover the structure. To better facilitate the co-learning of the two branches, two co-learning modules deal with the information flow interaction between the encoder and decoder of the two branches, respectively, thus explicitly guiding the structure-aware image reconstruction. Additionally, a dense feature enhancement module investigates the channel and spatial correlation of features and enhances the representation capability of the network. Extensive experiments on the CelebA and Helen datasets show that our proposed approach outperforms state-of-the-art methods. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008 |
ICASSP | 2 |
| 2023 | Deepfake Face Provenance for Proactive ForensicsabstractMalicious deepfake face not only violates the privacy of personal identities, but also confuses the public and causes huge social harm. The current deepfake detection only stays at the level of distinguishing between true and false, but cannot trace the original genuine face corresponding to the fake face, that is, it does not have the ability to trace the source of evidence. The deepfake countermeasure technology for judicial forensics urgently calls for deepfake inversion. This paper pioneers an interesting question about face deepfake, active forensics that "know what it is and how it happened". Given that deepfake faces do not completely discard the features of original faces, especially facial expressions and poses, we argue that original faces can be approximately speculated from their deepfake counterparts. Correspondingly, we design a disentangling reversing network that decouples latent space features of deepfake faces under the supervision of real-fake face pair samples to infer original faces in reverse. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002, Qin Zou 0001 |
ICIP | 2 |
| 2023 | Promoting adversarial transferability with enhanced loss flatnessabstractCarefully crafted small perturbations, when added to an image, can mislead the deep neural networks to give wrong outputs. Such mischievous images are called adversarial examples. Transfer-based black-box attacks use a surrogate white-box model to generate adversarial examples which can be transferred and attack black-box models with little known information. We propose to increase the transferability of adversarial examples by smoothing the geometric surface of loss function at the adversarial example point. By looking ahead the optimization path for a few steps, we define a future geometric vicinity using the integration of neighbourhood of those predicted data points. By sampling in this area and using the summation of gradients at those sampled data points for optimization, our method avoids local fluctuation of loss function. Experiments on ImageNet validation dataset show that our method outperforms state-of-the-art attacks by a large margin. Zhongyuan Wang 0001, Jikang Cheng, Chao Liang 0001 |
ICME | 2 |
| 2023 | Who, What and Where: Composite-semantic Instance Search for Story VideosabstractThis paper studies Who-What-Where (3W) composite-semantic video instance search (INS) problem, which aims to find a specific person doing a queried action in a particular place. Mainstream approaches adopt a complete decomposition strategy, which divides a composite-semantic query into multiple single-semantic queries. However, due to the lack of necessary correlation analysis among constituent semantics, these methods cannot always generate identity-matching and semantics-consistent 3W INS results. To address the above challenges, we propose a partial decomposition scheme with action as the link. Specifically, we selectively split the 3W INS as person-action INS and action-location INS. The former ensures the retrieved person and action share the same identity by modeling their relative spatial positions at the frame level, while the latter improves the semantic consistency between action and location with a cross-semantic attention mechanism at the shot level. Particularly, we build a large-scale 3W INS dataset, containing over 470k video shots, on basis of NIST TRECVID 2016-2021 INS tasks and verify the effectiveness of the proposed method with both quantitative and qualitative experiments. Chao Liang 0001, Zhongyuan Wang 0001 |
ICME | 3 |
| 2023 | DeepReversion: Reversely Inferring the Original Face from the DeepFake FaceabstractDeepfake techniques can generate realistic fake images and videos. Malicious fake facial images quickly spread through the Internet, posing a potential threat to personal privacy and judicial forensics. However, the defense methods against deepfake proposed so far mainly focus on the discrimination of authenticity, but cannot identify the true source of the forged face, i.e., the original genuine face corresponding to the face-swapped fake face. This paper poses an interesting issue for face deepfake, which is the proactive forensics of “knowing what and knowing how”. In view of the fact that the fake face exhibits high similarity with the original face, especially the facial expression and pose, we argue that the original face can be approximately estimated from the deepfake counterpart. Accordingly, we advocate a deep-learning-based face inversion approach, so-called DeepReversion, which learns the inverse mapping from the deepfake face to the original face. Based on UNet, we design a specific end-to-end DeepReversion network, and conduct comprehensive experiments on public deepfake datasets. The experimental results show that the speculated face is highly consistent with the original face in terms of visual effects, PSNR, SSIM and similarity given by face recognizers. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002 |
IJCNN | 2 |
| 2023 | Cerebral Thrombus Segmentation in CT Angiography using Refinement Segmentation Network with Context PoolingabstractCerebral thrombus segmentation from Computed Tomography Angiography (CTA) images is significant for accurate diagnosis in acute ischemic stroke. Automatic medical image segmentation can help doctors improve diagnosis efficiency in the clinic. In the past years, convolutional neural networks (CNNs) have been widely applied to automatic medical image segmentation and achieved impressive performance. However, since most existing methods fail to make full use of the contextual information of targets, the segmentation results are far from usability for low-contrast and tiny cerebral thrombosis. To tackle this problem, we proposed a Context Pooling Module (CPM) to take full advantages of local and global context information. In addition, segmentation results of present methods are still inaccurate in the marginal area. Thus, we designed a refinement network to reduce deviations in the coarse segmentation results. We applied our approach to thrombus segmentation in CTA images and tumor segmentation in Magnetic Resonance (MR) images. The experimental results showed that the proposed method outperforms the original U-Net and other state-of-the-art methods for thrombus segmentation, and also achieves competitive results on tumor segmentation. Jinbi Liang, Zhongyuan Wang 0001, Chuang Nie, Zhiming Kang, Bin Mei |
IJCNN | 4 |
| 2023 | A Spatio-Temporal Identity Verification Method for Person-Action Instance Search in Movies
Yanrui Niu, Jingyao Yang, Chao Liang 0001, Baojin Huang, Zhongyuan Wang 0001 |
MMM (1) | 5 |
| 2023 | An end-to-end network for co-saliency detection in one single image
Yuanhao Yue, Qin Zou 0001, Hongkai Yu, Qian Wang 0002, Zhongyuan Wang 0001, Song Wang 0002 |
Sci. China Inf. Sci. | 5 |
| 2023 | Tiny object detection with context enhancement and feature purification
Jinsheng Xiao, Haowen Guo, Jian Zhou 0011, Qiuze Yu, Yunhua Chen, Zhongyuan Wang 0001 |
Expert Syst. Appl. | 7 |
| 2023 | GaitAMR: Cross-view gait recognition via aggregated multi-feature representation
Jianyu Chen 0008, Zhongyuan Wang 0001, Caixia Zheng, Kangli Zeng, Qin Zou 0001, Laizhong Cui |
Inf. Sci. | 2 |
| 2023 | Face enhancement and hallucination in the wild
Ruimin Hu, Zhongyuan Wang 0001 |
Neural Comput. Appl. | 3 |
| 2023 | Depth map guided triplet network for deepfake face detection
Buyun Liang 0002, Zhongyuan Wang 0001, Baojin Huang, Qin Zou 0001, Qian Wang 0002 |
Neural Networks | 2 |
| 2023 | Self-attention learning network for face super-resolution
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Jiaming Wang 0001, Zixiang Xiong |
Neural Networks | 2 |
| 2023 | Single-channel Multi-speakers Speech Separation Based on Isolated Speech Segments
Shanfa Ke, Zhongyuan Wang 0001, Ruimin Hu, Xiaochen Wang 0001 |
Neural Process. Lett. | 2 |
| 2023 | PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001 |
Pattern Recognit. | 2 |
| 2023 | Variational Bayesian deep network for blind Poisson denoising
Hao Liang 0008, Rui Liu 0041, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
Pattern Recognit. | 3 |
| 2023 | HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002 |
Pattern Recognit. | 2 |
| 2023 | A Dual Self-Attention mechanism for vehicle re-Identification
Wenqian Zhu, Zhongyuan Wang 0001, Xiaochen Wang 0001, Ruimin Hu, Huikai Liu, Chao Wang 0084, Dengshi Li |
Pattern Recognit. | 2 |
| 2023 | Implicit space pose consistent transfer network for deep face verification
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Zhen Han 0002 |
Pattern Recognit. Lett. | 2 |
| 2023 | Stylized image denoising via noise style transfer and Quasi Siamese network
Jikang Cheng, Zhen Han 0002, Zhongyuan Wang 0001 |
Signal Process. Image Commun. | 3 |
| 2023 | FaceFormer: Aggregating Global and Local Representation for Face HallucinationabstractRecently, face hallucination methods either feed whole face image into convolutional neural networks (CNNs) or utilize extra facial priors (e.g., facial parsing maps and landmarks) to focus on global facial structure and constrain facial texture generation. However, the limited receptive fields of CNNs and inaccurate facial priors will reduce the naturalness and fidelity of restored face. In this paper, we propose a FaceFormer that aggregates global representation of Transformers and local representation of CNNs to maintain the consistency of facial structure while restoring local facial details. The reason for this design is that the Transformer can capture global facial information by exploiting the long-distance visual relation modeling, while the local modeling capability of CNNs can recover fine-grained facial details. Therefore, aggregating these two independent representations can help to maximize their merits and reconstruct high-quality and high-fidelity face images. Experimental results of face reconstruction and recognition verify that the proposed FaceFormer significantly outperforms current state-of-the-arts. Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Coarse-to-Fine Cross-Domain Learning Fusion Network for PansharpeningabstractDeep learning (DL) based pansharpening methods have shown great advantages in fusing multispectral (MS) and panchromatic (PAN) images to obtain a high-resolution MS image in remote sensing applications. However, most DL methods have low generalization capability that will cause severe spatial or spectral distortions, especially when a large distribution gap exists between training data from a source domain and testing data from another target domain. To overcome this problem, we propose a coarse-to-fine adaption learning fusion network for pansharpening. We first learn the priori mapping relationships between MS and PAN images in the source domain through the coarse-fusion network, which combines the advantages of UNet and Transformer architectures that helps to explore texture information of different characteristics. To generate a clear fusion result with good preservation of spatial and spectral information in the target domain, the fine-fusion network is further proposed to adjust the spatial and spectral information of the coarse-fusion image in an unsupervised learning manner based on the target-specific knowledge. Therefore, the generalization capability can be effectively improved because the general mapping relationship from the source domain and the specific target knowledge from the target domain are both considered. Experiments on simulated and real datasets are conducted to demonstrate the superiority of our proposed method over other state-of-the-art DL methods in terms of visual quality and quantitative analysis. Chengjie Ke, Wei Zhang 0259, Zhongyuan Wang 0001, Jiayi Ma 0001, Xin Tian 0006 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Joint Segmentation and Identification Feature Learning for Occlusion Face RecognitionabstractThe existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification. Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multi-Scale Hybrid Fusion Network for Single Image DerainingabstractDeep learning models have been able to generate rain-free images effectively, but the extension of these methods to complex rain conditions where rain streaks show various blurring degrees, shapes, and densities has remained an open problem. Among the major challenges are the capacity to encode the rain streaks and the sheer difficulty of learning multi-scale context features that preserve both global color coherence and exactness of detail. To address the first problem, we design a non-local fusion module (NFM) and an attention fusion module (AFM), and construct the multi-level pyramids' architecture to explore the local and global correlations of rain information from the rain image pyramid. More specifically, we apply the non-local operation to fully exploit the self-similarity of rain streaks and perform the fusion of multi-scale features along the image pyramid. To address the latter challenge, we additionally design a residual learning branch that is capable of adaptively bridging the gaps (e.g., texture and color information) between the predicted rain-free image and the clean background via a hybrid embedding representation. Extensive results have demonstrated that our proposed method is able to generate much better rain-free images on several benchmark datasets than the state-of-the-art algorithms. Moreover, we conduct the joint evaluation experiments with respect to deraining performance and the detection/segmentation accuracy to further verify the effectiveness of our deraining method for downstream vision tasks/applications. The source code is available at https://github.com/kuihua/MSHFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Guangcheng Wang, Zhen Han 0002, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Local Eyebrow Feature Attention Network for Masked Face RecognitionabstractDuring the COVID-19 coronavirus epidemic, wearing masks has become increasingly popular. Traditional occlusion face recognition algorithms are almost ineffective for such heavy mask occlusion. Therefore, it is urgent to improve the recognition performance of the existing face recognition technology on masked faces. Due to the limited visible feature points of the masked face image relative to the normal face image, we have to exploit the identification potential of eyebrow (referring to eyes and brows) features. This article proposes a local eyebrow feature attention network for masked face recognition, which consists of feature extraction, eyebrow region pooling, and feature fusion. To highlight the eyebrow region, we first use the eyebrow region pooling to separate the local features of eyebrows from the learned overall facial features. We then make full use of the symmetry of left and right eyebrows to emphasize their discriminant ability, due to the inadequate fine information of the low-resolution eyebrows. In particular, in view of the symmetrical similarity between eyebrow pairs and the subordinate relationship between facial components and the whole, we propose a feature fusion model based on graph convolutional network (GCN) to learn the feature association structure of eye features, brow features, and global facial features. We construct the benchmark datasets for masked face recognition to validate our approach, including real-world masked face recognition dataset (RMFRD) and synthetic masked face recognition dataset (SMFRD). Extensive experimental results on both public datasets and our built masked face datasets show that our approach significantly outperforms the state-of-the-arts. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Zhen Han 0002, Kui Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Degrade Is Upgrade: Learning Degradation for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve an image's visibility while keeping its visual naturalness. Different from existing methods, which tend to accomplish the relighting task directly, we investigate the intrinsic degradation and relight the low-light image while refining the details and color in two steps. Inspired by the color image formulation (diffuse illumination color plus environment illumination color), we first estimate the degradation from low-light inputs to simulate the distortion of environment illumination color, and then refine the content to recover the loss of diffuse illumination color. To this end, we propose a novel Degradation-to-Refinement Generation Network (DRGN). Its distinctive features can be summarized as 1) A novel two-step generation network for degradation learning and content refinement. It is not only superior to one-step methods, but also capable of synthesizing sufficient paired samples to benefit the model training; 2) A multi-resolution fusion network to represent the target information (degradation or contents) in a multi-scale cooperative manner, which is more effective to address the complex unmixing problems. Extensive experiments on both the enhancement task and the joint detection task have verified the effectiveness and efficiency of our proposed method, surpassing the SOTA by 1.59dB on average and 3.18\% in mAP on the ExDark dataset. The code will be available soon. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Chen Chen 0001, Peng Yi 0002, Tao Lu 0001, Chia-Wen Lin |
AAAI | 2 |
| 2022 | Ranking Aggregation with Interactive Feedback for Collaborative Person Re-identification
Chao Liang 0001, Yue Zhang 0104, Zhongyuan Wang 0001, Chunjie Zhang 0001 |
BMVC | 4 |
| 2022 | Deepfake Video Detection Exploiting Binocular Synchronization
Zhongyuan Wang 0001, Guangcheng Wang, Qin Zou 0001 |
ICANN (3) | 2 |
| 2022 | ECL: Exclusive Curriculum Learning for Video Super-ResolutionabstractVideo super-resolution (VSR) problem has gained a soaring development along with deep learning methods. However, the further progress requires the blessing of more complex architectures. Unlike them, this paper promotes VSR performance from a new perspective of sample difficulty. We propose an exclusive curriculum learning strategy for VSR, which can improve the representation power without noticeable computation increment. Specifically, this paper memorizes the performance track of every sample and calculate a customized weight for each sample according to it. In this way, the model can automatically concentrate on the easy samples first and gradually focus on the hard ones. Experimental analysis on training process and benchmark datasets demonstrate that our method can substantially boost the performance with a superior convergence speed and a limited number of parameters. Sicheng Hu, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001, Jinsheng Xiao, Jing Xiao 0004 |
ICME | 2 |
| 2022 | Video Face Recognition Using Neural Aggregation Networks with Mutual Relational LearningabstractVideo face recognition benefits profoundly from deep convolutional neural networks (CNNs), which learn robust feature embeddings. However, due to their fixed geometric structures, CNNs are inherently limited in modeling the significant variations from the angle, pose, occlusion and other factors of face images. In this paper, a neural aggregation network based on mutual relation learning is proposed for video face recognition. First, Intra-frame Relational Learning network (Intra-Net) is introduced, which models the interdependencies between the re-gional components of individual features and develops relevance between fine-grained features. Such processing can determine the region of interest adaptively according to the quality of the input face image to achieve the extraction of valuable information. Secondly, we introduce Inter-frame Relational Learning Network (Inter-Net), which considers the most significant appearance representation in the overall structure of the face image to cor-relate the complementarity of features between frames. Finally, information aggregation is performed by combining Inter-Net and Intra-Net. Joint optimization of the two branches allows our model to effectively exploit the complementary information between them to improve the aggregation capability. We validate the effectiveness of our model for video face recognition, proving its superiority over state-of-the-art methods on two benchmark datasets. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008 |
ICTAI | 2 |
| 2022 | DANet: Image Deraining via Dynamic Association LearningabstractRain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end, we propose a dynamic associated network (DANet) to achieve the association learning between rain streak removal and background recovery. There are two key aspects to fulfill the association learning: 1) DANet unveils the latent association knowledge between rain streak prediction and background texture recovery, and leverages it as an extra prior via an associated learning module (ALM) to promote the texture recovery. 2) DANet introduces the parametric association constraint for enhancing the compatibility of deraining model with background reconstruction, enabling it to be automatically learned from the training data. Moreover, we observe that the sampled rainy image enjoys the similar distribution to the original one. We thus propose to learn the rain distribution at the sampling space, and exploit super-resolution to reconstruct high-frequency background details for computation and memory reduction. Our proposed DANet achieves the approximate deraining performance to the state-of-the-art MPRNet but only requires 52.6\% and 23\% inference time and computational cost, respectively. Kui Jiang, Zhongyuan Wang 0001, Zheng Wang 0007, Peng Yi 0002, Junjun Jiang, Jinsheng Xiao, Chia-Wen Lin |
IJCAI | 2 |
| 2022 | Long-Tailed Multi-label Retinal Diseases Recognition via Relational Learning and Knowledge Distillation
Qian Zhou 0001, Hua Zou 0002, Zhongyuan Wang 0001 |
MICCAI (2) | 3 |
| 2022 | Event-guided Video Clip Generation from Blurry ImagesabstractDynamic and active pixel vision sensors (DAVIS) can simultaneously produce streams of asynchronous events captured by the dynamic vision sensor (DVS) and intensity frames from the active pixel sensor (APS). Event sequences show high temporal resolution and high dynamic range, while intensity images easily suffer from motion blur due to the low frame rate of APS. In this paper, we present an end-to-end convolutional neural network based method under the local and global constraints of events to restore clear, sharp intensity frames through collaborative learning from a blurry image and its associated event streams. Specifically, we first learn a function of the relationship between the sharp intensity frame and the corresponding blurry image with its event data. Then we propose a generation module to realize it with a supervision module to constrain the restoration in the motion process. We also capture the first realistic dataset with paired blurry frame/events and sharp frames by synchronizing a DAVIS camera and a high-speed camera. Experimental results show that our method can reconstruct high-quality sharp video clips, and outperform the state-of-the-art on both simulated and real-world data. Tsuyoshi Takatani, Zhongyuan Wang 0001, Ying Fu 0001, Yinqiang Zheng |
ACM Multimedia | 3 |
| 2022 | Magic ELF: Image Deraining Meets Association Learning and TransformerabstractConvolutional neural network (CNN) and Transformer have achieved great success in multimedia applications. However, little effort has been made to effectively and efficiently harmonize these two architectures to satisfy image deraining. This paper aims to unify these two architectures to take advantage of their learning merits for image deraining. In particular, the local connectivity and translation equivariance of CNN and the global aggregation ability of self-attention (SA) in Transformer are fully exploited for specific local context and global structure representations. Based on the observation that rain distribution reveals the degradation location and degree, we introduce degradation prior to help background recovery and accordingly present the association refinement deraining scheme. A novel multi-input attention module (MAM) is proposed to associate rain perturbation removal and background recovery. Moreover, we equip our model with effective depth-wise separable convolutions to learn the specific feature representations and trade off computational complexity. Extensive experiments show that our proposed method (dubbed as ELF) outperforms the state-of-the-art approach (MPRNet) by 0.25 dB on average, but only accounts for 11.7% and 42.1% of its computational cost and parameters. Kui Jiang, Zhongyuan Wang 0001, Chen Chen 0001, Zheng Wang 0007, Laizhong Cui, Chia-Wen Lin |
ACM Multimedia | 2 |
| 2022 | Face hallucination based on degradation analysis for robust manifold
Ruimin Hu, Zheng He 0001, Chao Liang 0001, Zhongyuan Wang 0001 |
Neurocomputing | 5 |
| 2022 | Facial expressions recognition with multi-region divided attention networks for smart education cloud applications
Yifei Guo, Mingfu Xiong, Zhongyuan Wang 0001, Xinrong Hu, Mohammad Hijji |
Neurocomputing | 4 |
| 2022 | Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006 |
Inf. Sci. | 2 |
| 2022 | Structure-Texture Parallel Embedding for Remote Sensing Image Super-ResolutionabstractThe structure and texture of images are crucial for remote sensing image super-resolution. Generative adversarial networks (GANs) recover image details through adversarial training. However, the recovered images always have structural distortions on the one hand, and GANs are difficult to train on the other hand. In addition, some methods assist reconstruction by introducing prior information of the image, but this brings additional computational cost. To address this issue, we propose a novel structure-texture parallel embedding (SPE) method for super-resolution (SR) of remote sensing images. Our method does not require additional image priors to reconstruct high-quality images. Specifically, we use the global structure information and local texture information of the image in the ascending space to guide the reconstruction result of the image. Firstly, we design a structure preserving block (SPB) to extract global structural features in the ascending space of the image, so as to obtain global structure information for a priori representation. Then, we design a local texture attention module (LTAM) to restore richer texture details. We have conducted lots of experiments on Draper public dataset. Experimental results show that our proposed method not only achieves a better trade-off between computational cost and performance, but also outperforms the existing several SR methods in terms of objective index evaluation and subjective visual effects. Tao Lu 0001, Kanghui Zhao, Yuntao Wu, Zhongyuan Wang 0001, Yanduo Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Spatiotemporal two-stream LSTM network for unsupervised video summarization
Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong |
Multim. Tools Appl. | 3 |
| 2022 | Generative adversarial network with hybrid attention and compromised normalization for multi-scene image conversion
Jinsheng Xiao, Shuhao Zhang 0008, Yuntao Yao, Zhongyuan Wang 0001, Yongqin Zhang, Yuan-Fang Wang |
Neural Comput. Appl. | 4 |
| 2022 | A Progressive Fusion Generative Adversarial Network for Realistic and Consistent Video Super-ResolutionabstractHow to effectively fuse temporal information from consecutive frames remains to be a non-trivial problem in video super-resolution (SR), since most existing fusion strategies (direct fusion, slow fusion, or 3D convolution) either fail to make full use of temporal information or cost too much calculation. To this end, we propose a novel progressive fusion network for video SR, in which frames are processed in a way of progressive separation and fusion for the thorough utilization of spatio-temporal information. We particularly incorporate multi-scale structure and hybrid convolutions into the network to capture a wide range of dependencies. We further propose a non-local operation to extract long-range spatio-temporal correlations directly, taking place of traditional motion estimation and motion compensation (ME&MC). This design relieves the complicated ME&MC algorithms, but enjoys better performance than various ME&MC schemes. Finally, we improve generative adversarial training for video SR to avoid temporal artifacts such as flickering and ghosting. In particular, we propose a frame variation loss with a single-sequence training method to generate more realistic and temporally consistent videos. Extensive experiments on public datasets show the superiority of our method over state-of-the-art methods in terms of performance and complexity. Our code is available at https://github.com/psychopa4/MSHPFNL. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | COVID-19 contact tracking by group activity trajectory recovery over camera networks
Chao Wang 0084, Xiaochen Wang 0001, Zhongyuan Wang 0001, Wenqian Zhu, Ruimin Hu |
Pattern Recognit. | 3 |
| 2022 | Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006 |
Pattern Recognit. | 2 |
| 2022 | Rethinking Lightweight: Multiple Angle Strategy for Efficient Video Action RecognitionabstractVideo action recognition task involves modeling spatiotemporal information, and efficiency is critical to capture spatiotemporal dependencies in the video. Most existing models rely on optical flow information to capture the dynamic visual tempos between consecutive video frames. Although impressive performance can be achieved by combining optical flow with RGB, the time-consuming nature of optical flow computation cannot be ignored. Moreover, 3D CNN has successfully modeled spatiotemporal information, yet the enormous computational volume is unsuitable for real-time action recognition. In this letter, we propose a novel lightweight video feature extraction strategy that achieves better recognition performance with lower FLOPs. In particular, we perform convolution on the video cube from three orthogonal angles to learn its appearance and motion features. Compared with the computational volume of 3D CNN, our proposed method is more economical and thus meets the lightweight requirements. Extensive experimental results on public Something Something-V1$\&$V2 and Diving48 datasets show our approach achieves the state-of-the-art performance. Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Zheng He 0001, Zixiang Xiong |
IEEE Signal Process. Lett. | 2 |
| 2022 | Few-Shot Semantic Segmentation via Frequency Guided Neural NetworkabstractPrototype learning is extensively used in few-shot semantic segmentation due to its excellent capability of semantic information extraction and effective prevention of overfitting. The previous prototype based methods ignore the frequency discrepancy inside the object, thereby leading to semantic confusion of the object. In this paper, we propose a frequency guided network (FGNet) which explicitly models the semantic information of different frequencies and precisely guides the semantic alignment of the object. Specifically, the proposed FGNet consists of two modules: a frequency separation module (FSM) and a multi-guided feature enrichment module (MG-FEM) to complete the multi-frequency semantic information extraction and alignment, respectively. Experiments on PASCAL-$5^{i}$dataset show that our FGNet achieves mIoU score of 61.2% in 1-shot which surpasses the state-of-the-art methods. Xiya Rao, Tao Lu 0001, Zhongyuan Wang 0001, Yanduo Zhang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Classifying Facial Regions for Face HallucinationabstractRecently, convolutional neural networks (CNNs) have dominated the face hallucination task due to their powerful feature representation capability. However, most of them simply use the same weights to treat different facial regions without considering the reconstruction difficulty of different facial regions, resulting in the component regions (e.g., eyes, nose, mouth) of the reconstructed faces tending to be blurred. In this paper, we propose a novel facial region classification network (FRCN) to address this problem. The proposed method first divides the input low-resolution (LR) facial image into several patch blocks, then classifies them into three categories according to their reconstruction difficulty, and finally inputs the three types of patch blocks into three networks with different weights for reconstruction and combining, thereby recovering high-quality high-resolution (HR) facial image. Experimental results show that FRCN can remarkably improve face reconstruction's performance. Yiyao Wang, Tao Lu 0001, Yuanzhi Wang, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Reference-Free DIBR-Synthesized Video Quality Metric in Spatial and Temporal DomainsabstractDepth image-based rendering (DIBR) techniques play an important role in free viewpoint videos (FVVs), which have a wide range of applications including immersive entertainment, remote monitoring, education, etc. FVVs are usually synthesized by DIBR techniques in a “blind” environment (without a reference video). Thus, an effective reference-free synthesized video quality assessment (VQA) metric is vital. At present, many image quality assessment (IQA) algorithms for DIBR-synthesized images have been proposed, but limited researches have been concerned about the quality assessment of DIBR-synthesized videos. To this end, this paper proposes a novel reference-free VQA method for synthesized videos, which operates in Spatial and Temporal Domains, dubbed as STD. The design fundamental of the proposed STD metric considers the effects of two major distortions introduced by DIBR techniques on the visual quality of synthesized videos. First, considering the geometric distortion introduced by DIBR technologies can increase high-frequency contents of the synthesized frame, the influence of the geometric distortion on the visual quality of a synthesized video can be effectively evaluated by estimating high-frequency energies of each synthesized frame in spatial domain. Second, temporal inconsistency caused by DIBR techniques brings the temporal flicker distortion, which is one of the most annoying artifacts in DIBR-synthesized videos. In temporal domain, we quantify temporal inconsistency by measuring motion differences between consecutive frames. Specifically, optical flow method is first used to estimate the motion field between adjacent frames. Then, we calculate the structural similarity of adjacent optical flow fields and further adopt the structural similarity value to weight the pixel differences of adjacent optical flow fields. Experiments show that the above two features are able to well perceive the visual quality of DIBR-synthesized videos. Furthermore, since the two features are extracted from spatial and temporal domains, respectively, we integrate them using a linear weighting strategy to obtain our STD metric, which proves advantageous over two components and the competing state-of-the-art I/VQA methods. The source code is available athttps://github.com/wgc-vsfm/DIBR-video-quality-assessment. Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Kui Jiang, Zheng He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | NWPU-Captions Dataset and MLCA-Net for Remote Sensing Image CaptioningabstractRecently, the burgeoning demands for captioning-related applications have inspired great endeavors in the remote sensing community. However, current benchmark datasets are deficient in data volume, category variety, and description richness, which hinders the advancement of new remote sensing image captioning approaches, especially those based on deep learning. To overcome this limitation, we present a larger and more challenging benchmark dataset, termed NWPU-Captions. NWPU-Captions contains 157,500 sentences, with all 31,500 images annotated manually by 7 experienced volunteers. The superiority of NWPU-Captions over current publicly available benchmark datasets not only lies in its much larger scale but also in its wider coverage of complex scenes and the richness and variety of describing vocabularies. Further, a novel encoder-decoder architecture, multi-level and contextual attention network (MLCA-Net), is proposed. MLCA-Net employs a multi-level attention module to adaptively aggregate image features of specific spatial regions and scales and introduces a contextual attention module to explore the latent context hidden in remote sensing images. MLCA-Net improves the flexibility and diversity of the generated captions while keeping their accuracy and conciseness by exploring the properties of scale variations and semantic ambiguity. Finally, the effectiveness, robustness, and generalization of MLCA-Net are proved through extensive experiments on existing datasets and NWPU-Captions. Qimin Cheng, Yuzhuo Zhou, Huanying Li, Zhongyuan Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Road and Car Extraction Using UAV Images via Efficient Dual Contextual Parsing NetworkabstractThe rapid development and commercialization of Unmanned Aerial Vehicle (UAV) technology has made it possible to conduct urban traffic information extraction using UAV images. However, the large variations of targets in urban environments, complex foregrounds and backgrounds in cities, and severe tree and shadow occlusions pose great challenges in car and road extraction using UAV images. In this study, we propose a lightweight, Efficient Dual Contextual Parsing Network (EDCPNet) to address the above issues. The proposed EDCP module in EDCPNet is mainly composed of spatial contextual parsing (SCP) and channel contextual parsing (CCP), which can effectively acquire rich contextual features in both spatial and channel dimensions, adaptively recalibrate the attention weights, perceive the salient features of targets in images, and suppress the importance of irrelevant elements. It thus leads to improved performance and adaptability that facilitate the practical applications of large-scale urban traffic monitoring in complex urban scenes. We conduct experiments on two benchmark datasets (UAVid and UDD) by comparing the proposed EDCPNet with six other competing methods, i.e., U-Net, PSPNet, Deelabv3+, SegNet, ESNet, and ERFNet, and validate the effectiveness of the proposed EDCP module via extensive ablation studies. The results suggest that the proposed network outperforms all competing methods in car and road extraction from UAV images with a balanced computational cost. Its great performance and low computational demand (with only 2.37M model parameters) facilitate its deployment on edge computing devices with memory constraints. Gui Cheng, Xiao Huang 0003, Zhongyuan Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | VP-Net: An Interpretable Deep Network for Variational PansharpeningabstractIn this study, we propose an interpretable deep network for variational pansharpening (VP), named VP-Net. Different from traditional priors using linear operators, such as the gradient, we construct a prior based on the similarity between panchromatic (PAN) and high-resolution multispectral (HRMS) images by a nonlinear operator that can be learned through a deep network. Considering the spectral difference of various satellite multispectral (MS) imaging platforms, we specifically seek the aforementioned similarity from the PAN image and the intensity of the HRMS image to reduce the spectral distortion. Based on this prior, we propose a novel VP model by further incorporating a data fidelity term from the low-resolution MS image. Specifically, we build the VP-Net by unrolling the variable splitting method for an optimal solution to this model. Consequently, all modules in VP-Net have clear physical meanings and strong generalization capabilities. Meanwhile, all parameters and the aforementioned nonlinear operator are learned in VP-Net, avoiding the difficulty of selecting optimal handcrafted parameters in traditional methods. Therefore, VP-Net not only achieves an optimal balance between spatial and spectral qualities but also has a strong generalization capability across different types of training and testing data. In the experiment, we first demonstrate the superiority of the proposed method over the current state of the arts in terms of both visual effect and quantitative analysis on different satellite datasets. Moreover, we carry out an normalized difference vegetation index (NDVI) experiment to demonstrate its potential in remote sensing. Xin Tian 0006, Kun Li 0025, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | HyperFusion: A Computational Approach for Hyperspectral, Multispectral, and Panchromatic Image FusionabstractFusing hyperspectral image (HSI) and multispectral image (MSI) of high spatial resolution is typically utilized to obtain HSIs of high spatial resolution. However, the spatial quality of most existing methods is unsatisfactory due to the limited spatial resolution of an MSI. To further improve the spatial resolution of the fused HSI while keeping the spectral information well, we propose a new computational paradigm, named HyperFusion, which simultaneously fuses HSI, MSI, and panchromatic (PAN) image. To achieve this goal, we first establish two data fidelity terms based on a physical observation that HSI and MSI can be treated as degraded versions of the fused HSI. Consequently, the spatial and spectral information from HSI and MSI can be well preserved. To efficiently transfer the spatial details of PAN into the fused HSI while keeping the spectral information well, we further construct a prior constraint from PAN based on the structural similarity. Meanwhile, we impose another low-rank prior constraint on the coefficient matrix to accurately describe the latent characteristics of the HSI with high spatial resolution. By incorporating the aforementioned data fidelity terms and prior constraints, we finally formulate the objective as an optimization problem and utilize the alternative direction multiplier method to solve it efficiently. Comprehensive experiments on simulated and real datasets are carried out to demonstrate the superiority of HyperFusion over other state of the arts in terms of visual quality and quantitative analysis. We also adopt a simulated experiment of vegetation coverage index analysis to verify the effectiveness of HyperFusion in remote sensing applications. Xin Tian 0006, Wei Zhang 0259, Yuerong Chen, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Real-Time and Accurate UAV Pedestrian Detection for Social Distancing Monitoring in COVID-19 PandemicabstractCoronavirus Disease 2019 (COVID-19) is a highly infectious virus that has created a health crisis for people all over the world. Social distancing has proved to be an effective non-pharmaceutical measure to slow down the spread of COVID-19. As unmanned aerial vehicle (UAV) is a flexible mobile platform, it is a promising option to use UAV for social distance monitoring. Therefore, we propose a lightweight pedestrian detection network to accurately detect pedestrians by human head detection in real-time and then calculate the social distancing between pedestrians on UAV images. In particular, our network follows the PeleeNet as backbone and further incorporates the multi-scale features and spatial attention to enhance the features of small objects, like human heads. The experimental results on Merge-Head dataset show that our method achieves 92.22% AP (average precision) and 76 FPS (frames per second), outperforming YOLOv3 models and SSD models and enabling real-time detection in actual applications. The ablation experiments also indicate that multi-scale feature and spatial attention significantly contribute the performance of pedestrian detection. The test results on UAV-Head dataset show that our method can also achieve high precision pedestrian detection on UAV images with 88.5% AP and 75 FPS. In addition, we have conducted a precision calibration test to obtain the transformation matrix from images (vertical images and tilted images) to real-world coordinate. Based on the accurate pedestrian detection and the transformation matrix, the social distancing monitoring between individuals is reliably achieved. Gui Cheng, Jiayi Ma 0001, Zhongyuan Wang 0001, Jiaming Wang 0001, DeRen Li |
IEEE Trans. Multim. | 4 |
| 2022 | Dual-Path Deep Fusion Network for Face Image HallucinationabstractAlong with the performance improvement of deep-learning-based face hallucination methods, various face priors (facial shape, facial landmark heatmaps, or parsing maps) have been used to describe holistic and partial facial features, making the cost of generating super-resolved face images expensive and laborious. To deal with this problem, we present a simple yet effective dual-path deep fusion network (DPDFN) for face image super-resolution (SR) without requiring additional face prior, which learns the global facial shape and local facial components through two individual branches. The proposed DPDFN is composed of three components: a global memory subnetwork (GMN), a local reinforcement subnetwork (LRN), and a fusion and reconstruction module (FRM). In particular, GMN characterize the holistic facial shape by employing recurrent dense residual learning to excavate wide-range context across spatial series. Meanwhile, LRN is committed to learning local facial components, which focuses on the patch-wise mapping relations between low-resolution (LR) and high-resolution (HR) space on local regions rather than the entire image. Furthermore, by aggregating the global and local facial information from the preceding dual-path subnetworks, FRM can generate the corresponding high-quality face image. Experimental results of face hallucination on public face data sets and face recognition on real-world data sets (VGGface and SCFace) show the superiority both on visual effect and objective indicators over the previous state-of-the-art methods. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Tao Lu 0001, Junjun Jiang, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | When Face Recognition Meets Occlusion: A New BenchmarkabstractThe existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC. Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001 |
ICASSP | 2 |
| 2021 | A Triplet Appearance Parsing Network for Person Re-IdentificationabstractAs one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of person representation for cluttered backgrounds, are difficult to adapt the challenges of the harsh real-world scenarios. This study develops a Triplet person Appearances Parsing Framework (TAPF) which eliminates the surrounding interference factors of bounding boxes for person re-identification. The framework consists of a triplet person parsing network and an integration mechanism for person local and global appearance information. Concretely, the triplet parsing network includes a channel parsing module, a position parsing module and a color parsing module, which are used to extract the person channel parsing descriptor, regional descriptor and color perception descriptor, respectively. Then, a local and global flatten gaussian operations are performed to integrate the person appearance parsing descriptors to obtain more discriminative features for the person representation. The experimental results have been conducted to validate our proposed algorithm can achieve a better performance for person re-identification on several public datasets, i.e., VIPeR and Market-1501, respectively. Mingfu Xiong, Zhongyuan Wang 0001, Ruhan He, Xinrong Hu, Xiao Qin 0001, Jia Chen 0012 |
ICASSP | 2 |
| 2021 | Omniscient Video Super-ResolutionabstractMost recent video super-resolution (SR) methods either adopt an iterative manner to deal with low-resolution (LR) frames from a temporally sliding window, or leverage the previously estimated SR output to help reconstruct the current frame recurrently. A few studies try to combine these two structures to form a hybrid framework but have failed to give full play to it. In this paper, we propose an omniscient framework to not only utilize the preceding SR output, but also leverage the SR outputs from the present and future. The omniscient framework is more generic because the iterative, recurrent and hybrid frameworks can be regarded as its special cases. The proposed omniscient framework enables a generator to behave better than its counterparts under other frameworks. Abundant experiments on public datasets show that our method is superior to the state-of-the-art methods in objective metrics, subjective visual effects and complexity. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Tao Lu 0001, Xin Tian 0006, Jiayi Ma 0001 |
ICCV | 2 |
| 2021 | PCNET: Progressive Coupled Network for Real-Time Image DerainingabstractImage deraining is an effective solution to avoid performance drop of vision-oriented tasks in rainy weather. Most existing image deraining approaches either fail to produce satisfactory restoration results or cost too much computation. In this paper, we propose a low-complexity and high-performance coupled representation module (CRM), designed to learn the joint features of rain-free contents and rain information as well as their blending correlations. To promote the computation efficiency, we employ depth-wise separable convolutions, and construct CRM in an asymmetric U-shaped architecture to reduce model parameters and memory footprint. Our final model–PCNet achieves the progressive separation of rain-free contents and rain streaks using cascaded residual learning. Extensive experiments are conducted to evaluate the efficacy of the proposed PCNet on several synthetic and real-world rain datasets. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zheng Wang 0007, Chia-Wen Lin |
ICIP | 2 |
| 2021 | A Tilt-Angle Face Dataset And Its ValidationabstractSince the surveillance cameras are usually mounted at a high position to overlook targets, tilt-angle faces on overhead view are common in the public video surveillance environment. Face recognition approaches based on deep learning models have achieved excellent performance, but there remains a large gap for the overlooking surveillance scenarios. The results of face recognition depend not only on the structure of the model, but also on the completeness and diversity of the training samples. The existing multi-pose face datasets do not cover complete top-view face samples, and the models trained by them thus cannot provide satisfactory accuracy. To this end, this paper pioneers a multi-view tilt-angle face dataset (TFD), which is collected with an elaborately devised overhead capture equipment. TFD contains 11,124 face images from 927 subjects, covering a variety of tilt angles on the overhead view. To verify the validity of the constructed dataset, we further conduct comprehensive face detection and recognition experiments using the corresponding models trained by WiderFace, Webface and our TFD, respectively. Experimental results show that our TFD substantially promotes the face detection and recognition accuracy under the top-view situation. TFD is available at https://github.com/huang1204510135/D FD. Nanxi Wang, Zhongyuan Wang 0001, Zheng He 0001, Baojin Huang, Liguo Zhou, Zhen Han 0002 |
ICIP | 2 |
| 2021 | Variation-Net: Interpretable Variation-Inspired Deep Network for PansharpeningabstractIn this study, we propose Variation-net, an interpretable variation-inspired deep network for pansharpening, which aims to fuse panchromatic (PAN) and multispectral (MS) images for a high-resolution MS image. We first construct a novel variational pan-sharpening model with clear physical meanings. As the relationship between the PAN and MS images in the real situation is complex and nonlinear, we explore the similarity between PAN and MS images from the sparsity of nonlinear transforms in this variational pansharpening model. As a result, spatial details can be accurately transferred from PAN image to MS image. Furthermore, we build the Variation-net by unrolling the iterative shrinkage-thresholding algorithm to solve the proposed variational pansharpening model. Therefore, all modules in Variation-net have clear physical meanings and are easily observed, leading to good generalization capability. Meanwhile, nonlinear transforms and other parameters in the variational pansharpening model are learned end–to–end. The experiments demonstrate that Variation-net outperforms the state-of-the-art methods from the aspects of visual effect and objective quality analysis. Kun Li 0025, Wei Zhang 0259, Xin Tian 0006, Jiayi Ma 0001, Huabing Zhou, Zhongyuan Wang 0001 |
ICME | 6 |
| 2021 | Cross-View Gait Recognition Based on Feature FusionabstractCompared to face recognition, gait recognition is one of the most promising video biometric recognition technologies given that gait images can be readily captured at a distance and gait characteristics are robust to appearance camouflage. A lot of existing gait recognition methods aim at a single scene such as fixed cameras, but the recognition accuracy will decrease sharply if the viewpoints are changed. In this paper, we improve the existing methods and propose a cross-view gait recognition method based on feature fusion. Firstly, a multi-scale feature fusion module is proposed to extract the features of gait sequences with different granularities. Then, a dual-path structure is introduced to learn global appearance features and fine-grained local features, respectively. The features of two paths are gradually merged as the network deepens to obtain the complementary information. In the last feature mapping stage, the Generalized-Mean pooling is used to favour discriminative representation. Extensive experiments on the public dataset CASIA-B show that our method can achieve state-of-the-art recognition performance. Zhongyuan Wang 0001, Jianyu Chen 0008, Baojin Huang |
ICTAI | 2 |
| 2021 | Adaptive Texture Distillation Network for Image Hybrid Super-ResolutionabstractTo save the transmission bandwidth of high-resolution (HR) images, we can send down-sampled low-resolution (LR) images and reconstruct them using super-resolution (SR) technology at the receiving end. However, image down-sampling by a large factor results in the loss of many spatial details. Instead, we use a combination of spatial down-sampling by a small factor and gray-level quantization to obtain the low hybrid-resolution images. Although the small down-sampling factor makes images retain more spatial details and real textures, the gray-level quantization introduces fake textures. Obviously, the real textures should be enhanced, and the fake textures should be eliminated. To address this issue, we propose a lightweight Adaptive Texture Distillation Network (ATDN) for image hybrid super-resolution. Our model uses the texture enhancement block (TEB) and the texture smoothing block (TSB) to handle real and fake textures in different ways. Considering that the mixing proportions of two kinds of textures in low hybrid-resolution images vary with regions, we specifically use a cascaded weight branch to adaptively adjust the weights of real and fake textures. Experiments reveal that our model can effectively deal with the mixing problem of real and fake textures, and our method can achieve superior performance to other lightweight methods. Chunlei Liu 0006, Zhen Han 0002, Jiaxing Wen, Zhongyuan Wang 0001, Weiping Tu |
IJCNN | 5 |
| 2021 | Face Hallucination via Split-Attention in Split-Attention NetworkabstractRecently, convolutional neural networks (CNNs) have been widely employed to promote the face hallucination due to the ability to predict high-frequency details from a large number of samples. However, most of them fail to take into account the overall facial profile and fine texture details simultaneously, resulting in reduced naturalness and fidelity of the reconstructed face, and further impairing the performance of downstream tasks (e.g., face detection, facial recognition). To tackle this issue, we propose a novel external-internal split attention group (ESAG), which encompasses two paths responsible for facial structure information and facial texture details, respectively. By fusing the features from these two paths, the consistency of facial structure and the fidelity of facial details are strengthened at the same time. Then, we propose a split-attention in split-attention network (SISN) to reconstruct photorealistic high-resolution facial images by cascading several ESAGs. Experimental results on face hallucination and face recognition unveil that the proposed method not only significantly improves the clarity of hallucinated faces, but also encourages the subsequent face recognition performance substantially. Codes have been released at https://github.com/mdswyz/SISN-Face-Hallucination. Tao Lu 0001, Yuanzhi Wang, Yanduo Zhang, Yu Wang 0140, Wei Liu 0123, Zhongyuan Wang 0001, Junjun Jiang |
ACM Multimedia | 6 |
| 2021 | Metric Learning for Anti-Compression Facial Forgery DetectionabstractDetecting facial forgery images and videos is an increasingly important topic in multimedia forensics. As forgery images and videos are usually compressed into different formats such as JPEG and H264 when circulating on the Internet, existing forgery-detection methods trained on uncompressed data often suffer from significant performance degradation in identifying them. To solve this problem, we propose a novel anti-compression facial forgery detection framework, which learns a compression-insensitive embedding feature space utilizing both original and compressed forgeries. Specifically, our approach consists of three ideas: (i) extracting compression-insensitive features from both uncompressed and compressed forgeries using an adversarial learning strategy; (ii) learning a robust partition by constructing a metric loss that can reduce the distance of the paired original and compressed images in the embedding space; (iii) improving the accuracy of tampered localization with an attention-transfer module. Experimental results demonstrate that, the proposed method is highly effective in handling both compressed and uncompressed facial forgery images. Shenhao Cao, Qin Zou 0001, Xiuqing Mao, Dengpan Ye, Zhongyuan Wang 0001 |
ACM Multimedia | 5 |
| 2021 | Cross-task feature alignment for seeing pedestrians in the dark
Yuanzhi Wang, Tao Lu 0001, Yanduo Zhang, Wenhua Fang, Yuntao Wu, Zhongyuan Wang 0001 |
Neurocomputing | 6 |
| 2021 | Silicone mask face anti-spoofing detection based on visual saliency and facial motion
Guangcheng Wang, Zhongyuan Wang 0001, Kui Jiang, Baojin Huang, Zheng He 0001, Ruimin Hu |
Neurocomputing | 2 |
| 2021 | "One-Shot" Super-Resolution via Backward Style Transfer for Fast High-Resolution Style TransferabstractOwing to the excellent visual quality of results, Gatys et al.'s Neural Style Transfer (NST) online algorithm is regarded as the gold-standard in the community of NST, but this algorithm is quite time-consuming especially for high-resolution (HR) image. In this letter, we propose “One-Shot” super-resolution (SR) for fast high-resolution style transfer. We first generate a low-resolution (LR) stylized image by NST, and then use “One-Shot” super-resolution to restore the HR stylized image by learning the mapping relations between HR-LR stylized images from HR-LR style images. However, due to the style loss is not eliminated, there are some subtle but important fine-grained style differences between LR stylized and style images. These differences lead to the poor visual quality of SR results. To reduce the style differences further, we adjust the texture of LR style image to approach LR stylized image by backward style transfer. The result of backward style transfer will be treated as the LR part of the “One-Shot” example pair, which leads to a better SR. The experimental results show that with good visual quality, our method reduces the time consumption by 81.6%. Especially in a specific application scenario of fixed style image and changed content image, our method reduces the time consumption by 89.3%. Jikang Cheng, Zhen Han 0002, Zhongyuan Wang 0001, Liang Chen 0026 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Seeing in the Dark by Component-GANabstractRecently, Retinex theory based low-light image enhancement (LLIE) algorithms have achieved impressive results in controlled environment. However, the majority of deep learning based LLIE algorithms leverage relighting by enhancing the illumination components that directly determines the image brightness, regretfully, they ignore the information of reflectance components, which may cause problems such as image noise and color distortion in reconstructed images. To tackle this problem, in this letter, we propose a component enhancement network based on Generative Adversarial Network (Component-GAN) for recovering clear images from low-light ones. Specifically, the network is composed of the decomposition part for dividing the paired low/normal-light images into illumination components and reflectance components, and the enhancement part for generating high-quality images. It is worth to note that we provide two branches of component enhancement network, which are parallel to improve the two components simultaneously. Hereby, we treat the reconstruction part as the generative network and adopt discriminative network to boost image reconstruction performance. Through extensive experiments, the proposed approach outperforms some state-of-the-art LLIE methods in terms of visual and subjective qualities. Ning Rao, Tao Lu 0001, Yanduo Zhang, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2021 | Decomposition Makes Better Rain Removal: An Improved Attention-Guided Deraining NetworkabstractRain streaks in the air show diverse characteristics with different shapes, directions, densities, even the complex overlapped phenomenon, causing great challenges for the deraining task. Recently, deep learning based image deraining methods have been extensively investigated due to their excellent performance. However, most of the existing algorithms still have limitations in removing rain streaks while preserving rich textural details under complicated rain conditions. To this end, we propose to decompose rain streaks into multiple rain layers and individually estimate each of them along the network stages to cope with the increasing abstracts. To better characterize rain layers, an improved non-local block is designed to exploit the self-similarity of rain information by learning the holistic spatial feature correlations while reducing the calculation complexity. Moreover, a mixed attention mechanism is applied to guide the fusion of rain layers by focusing on the local and global overlaps among these rain layers. Extensive experiments on both synthetic rainy/rain-haze/raindrop datasets, real-world samples, the haze, and low-light scenarios show substantial improvements both on quantitative indicators and visual effects over the current state-of-the-art technologies. The source code is available athttps://github.com/kuihua/IADN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zhen Han 0002, Tao Lu 0001, Baojin Huang, Junjun Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Rain-Free and Residue Hand-in-Hand: A Progressive Coupled Network for Real-Time Image DerainingabstractRainy weather is a challenge for many vision-oriented tasks (e.g., object detection and segmentation), which causes performance degradation. Image deraining is an effective solution to avoid performance drop of downstream vision tasks. However, most existing deraining methods either fail to produce satisfactory restoration results or cost too much computation. In this work, considering both effectiveness and efficiency of image deraining, we propose a progressive coupled network (PCNet) to well separate rain streaks while preserving rain-free details. To this end, we investigate the blending correlations between them and particularly devise a novel coupled representation module (CRM) to learn the joint features and the blending correlations. By cascading multiple CRMs, PCNet extracts the hierarchical features of multi-scale rain streaks, and separates the rain-free content and rain streaks progressively. To promote computation efficiency, we employ depth-wise separable convolutions and a U-shaped structure, and construct CRM in an asymmetric architecture to reduce model parameters and memory footprint. Extensive experiments are conducted to evaluate the efficacy of the proposed PCNet in two aspects: (1) image deraining on several synthetic and real-world rain datasets and (2) joint image deraining and downstream vision tasks (e.g., object detection and segmentation). Furthermore, we show that the proposed CRM can be easily adopted to similar image restoration tasks including image dehazing and low-light enhancement with competitive performance. The source code is available at https://github.com/kuijiang0802/PCNet. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zheng Wang 0007, Xiao Wang 0029, Junjun Jiang, Chia-Wen Lin |
IEEE Trans. Image Process. | 2 |
| 2020 | Multi-Scale Progressive Fusion Network for Single Image DerainingabstractRain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary information for rain streak representation. In this work, we explore the multi-scale collaborative representation for rain streaks from the perspective of input image scales and hierarchical deep features in a unified framework, termed multi-scale progressive fusion network (MSPFN) for single image rain streak removal. For the similar rain streaks at different positions, we employ recurrent calculation to capture the global texture, thus allowing to explore the complementary and redundant information at the spatial dimension to characterize target rain streaks. Besides, we construct multi-scale pyramid structure, and further introduce the attention mechanism to guide the fine fusion of these correlated information from different scales. This multi-scale progressive fusion strategy not only promotes the cooperative representation, but also boosts the end-to-end training. Our proposed method is extensively evaluated on several benchmark datasets and achieves the state-of-the-art results. Moreover, we conduct experiments on joint deraining, detection, and segmentation tasks, and inspire a new research direction of vision task driven image deraining. The source code is available at https://github.com/kuihua/MSPFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Baojin Huang, Yimin Luo, Jiayi Ma 0001, Junjun Jiang |
CVPR | 2 |
| 2020 | Cartoon-Texture Decomposition-Based Variational PansharpeningabstractPansharpening is widely used to increase the spatial resolution of a multispectral (MS) image by fusing with a panchromatic (PAN) image that has high-spatial resolution and the same scene. In this paper, the similarities of MS and PAN images in cartoon-texture space are exploited. The cartoon and texture components of an image always contain the global structure information and the locally-patterned information, respectively. Therefore, the global and local spatial details (i.e., high-order information) could be preserved well in the fused high-spatial resolution MS image after leveraging the similarities of these images. Pansharpening is formulated as the optimization problem with respect to the cartoon-texture similarities between the MS and the PAN images in this work based on the aforementioned observation. Specifically, cartoon similarity is determined through gradient sparsity and formulated as a total variation term, whereas texture similarity is described according to the low-rank property. The alternative direction multiplier method is used to solve the optimization problem. In the experiment, the Gaofen-1 satellite dataset is used to compare the proposed method with other classical pansharpening methods. Experimental results demonstrate that our method outperforms the comparison methods in terms of visual and quantitative qualities. Yuerong Chen, Mengliang Zhang, Zhongyuan Wang 0001, Xin Tian 0006 |
ICASSP | 4 |
| 2020 | Attention-Guided Deraining Network Via Stage-Wise LearningabstractDue to diverse rain shapes, directions, densities as well as different distances to cameras, rain streaks in the air are interweaved and overlapped. However, most existing deraining methods are inherently oblivious this phenomenon and tend to learn a single rain streak layer to simulate this complex distribution, consequently failing to restore high-quality rain-free images. To solve this problem, along with the stage-wise learning, we propose a novel attention-guided deraining network (ADN) for rain streak removal. Specially, we decompose the rain streaks into multiple rain streak layers, and individually model them along the stages of the network to match the increasing abstracts. Moreover, the attention mechanism is utilized to guide the fusion of these rain streak layers by handling the overlaps between them. Extensive experiments on several benchmark datasets and real-world scenarios show substantial improvements both on quantitative indicators and visual effects over the current top-performing methods. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Yuhong Yang 0001, Xin Tian 0006, Junjun Jiang |
ICASSP | 2 |
| 2020 | Image Super-Resolution Using Residual Global Context NetworkabstractRecent studies have showed that convolutional neural networks (CNN) can effectively improve the performance of single image super-resolution (SR). However, previous methods rarely considered long-range dependencies between pixels and channel-wise interdependencies at the same time. They ignores the fact that natural images have strong internal data repetition which requires the network to capture long-range dependencies between pixels and considering the interdepen-dencies between channels can better exploit the input information of the network. In addition, although past studies have proved that deep convolutional neural network benefit the performance of image super-resolution, it also means that the network needs more memory consumption and higher computational complexity. To solve these problem,we introduce Global Context block (GCB) and design a comparative shallow network called Residual Global Context Networks (RGC-N). It achieves a better trade-off between the amount of parameter and the quality of image reconstruction. Extensive experiments demonstrate that the proposed method is superior to the state-of-the-art methods. Kuangye Liu, Zhen Han 0002, Junkui Chen, Chunlei Liu 0006, Jun Chen 0001, Zhongyuan Wang 0001 |
ICASSP | 6 |
| 2020 | Fusionndvi: A Novel Fusion Method for NDVI in Remote SensingabstractNormalized difference vegetation index (NDVI) is widely utilized to examine vegetation coverage and estimate crop yield. To obtain a high-resolution (HR) NDVI, fusion techniques, which first generates a HR multispectral (MS) image by fusing a low-resolution (LR) MS image and a HR panchromatic image, and then calculates the HR NDVI based on the fused HR MS image, are utilized in previous studies. A HR vegetation index calculated on the basis of HR panchromatic image could provide HR spatial resolution, and this vegetation index has a spatial structure that is similar to that of NDVI. Therefore, this similarity is investigated to construct a novel method called FusionNDVI to improve the fusion performance in this study. The fusion problem is formulated to minimize a least square fitting error term and a nonlocal gradient sparsity regularization term. The fitting term is used to limit the difference between the fused HR NDVI and the LR NDVI, whereas the regularizer enforces a similar nonlocal spatial structure in the fused NDVI and the HR vegetation index. An efficient solving algorithm based on the augmented Lagrangian method of multipliers is derived. The superiority of the proposed FusionNDVI method over the state-of-the-art ones is verified via simulations. Mengliang Zhang, Ziping Zhao 0002, Yuerong Chen, Zhongyuan Wang 0001, Xin Tian 0006 |
ICASSP | 4 |
| 2020 | Robust CBCT Reconstruction Based On Low-Rank Tensor Decomposition And Total Variation RegularizationabstractCone-beam computerized tomography (CBCT) has been widely used in numerous clinical applications. To reduce the effects of X-ray on patients, a low radiation dose is always recommended in CBCT. However, noise will seriously degrade image quality under a low dose condition because the intensity of the signal is relatively low. In this study, we propose to use the Huber loss function as a data fidelity term in CBCT reconstruction, making the reconstruction robust to impulse noise under low radiation dose condition. Furthermore, a low-rank tensor property is adopted as the prior term. Such property is helpful in recovering the missing structure information caused by impulse noise. The proposed CBCT reconstruction model is formulated by further integrating a 3D total variation term for reducing Gaussian noise. An alternative direction multiplier method is adopted to solve the optimization problem. Experiments on simulated and real data show that the proposed model outperforms existing CBCT reconstruction algorithms. Xin Tian 0006, Zhongyuan Wang 0001 |
ICIP | 5 |
| 2020 | Tell The Truth From The Front: Anti-Disguise Vehicle Re-IdentificationabstractRecent efforts have been increasingly made on vehicle reidentification (re-ID), which has huge contributions to intelligent transportation and criminal investigation. However, most existing methods heavily rely on the color and texture features of vehicles to discern their identities, which turn invalid under adversarial social security occasions where vehicles' color and style are always tampered or forged by crime suspects. In this paper, we propose a local feature preservation method to learn the structure-aware features from the position distribution of individual local regions within vehicle front window area, which appears more robust and discriminative upon disguise. We further develop a two-branch deep convolutional network framework to integrate the structure-aware features with vehicle model features for vehicle Re-ID. The experimental results on datasets VehicleID and Vehicle-1M show that our end-to-end framework achieves promising performance and outperforms the state-of-the-art methods proposed so far. Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao |
ICME | 3 |
| 2020 | Masked Face Recognition with Identification AssociationabstractIn the crime scene, criminals often consciously conceal their facial identity through face-masked disguise, which poses a huge challenge to identity recognition. Existing disguised face recognition techniques aiming for light even slight occlusions are completely invalid for face-masked identification. To this end, this paper proposes a masked face recognition method based on person re-identification association, which converts the masked face recognition problem into an association uncovering problem between the masked face and the appearing faces of the same person. Based on the characteristics that person re-identification technique does not rely solely on facial information, it first takes advantages of re-identification to establish the association between face-masked pedestrians and face-unveiled pedestrians. It further provides an effective face image quality assessment to select the most identifiable faces for subsequent recognition from a variety of appearing candidate faces. Finally, the selected high-quality recognizable faces are used to replace masked faces for identification. The comparison experiments with the existing disguise face recognition methods show its superiority in terms of accuracy. Zhongyuan Wang 0001, Zheng He 0001, Nanxi Wang, Xin Tian 0006, Tao Lu 0001 |
ICTAI | 2 |
| 2020 | Lightweight Progressive Residual Clique Network for Image Super-ResolutionabstractDeeper and wider convolutional neural networks (CNN) hava been widely applied to the single image super-resolution (SR) task for its appealing performance. However, enormous parametric memory footprint hinders its real-time application on mobile devices, especially in the energy-sensitive environment. In this work, we take both the reconstruction performance and efficiency into consideration and propose a lightweight progressive residual clique network (PRCN) for image SR. PRCN is built on the two-stage residual channel separation block (RCSB) and long-skip connections. First, we divide the input into four channel groups to differently learn texture details, immediately followed by a primary fusion to establish cross-channel correspondence in the first stage. Then we perform a further fusion on the outputs of the first stage to constitute a clique for the refinement in the second stage. Meanwhile, we employ SENet to improve the outputs of the second stage with the separate features of the first stage. This design not only enforces the correlation across channels, but also allows fewer densely connected blocks. Experimental results on public datasets show that PRCN outperforms state-of-the-art methods in terms of performance and complexity. Baojin Huang, Zheng He 0001, Zhongyuan Wang 0001, Kui Jiang, Guangcheng Wang |
ICTAI | 3 |
| 2020 | Story segmentation for news broadcast based on primary captionabstractIn the information explosion era, people only want to access the news information that they are interested in. News broadcast story segmentation is strongly needed, which is an essential basis for personalized delivery and short video. The existing advanced story boundary segmentation methods utilize semantic similarity of subtitles, thus entailing complex semantic computation. The title texts of news broadcast programs include headline (or primary) captions, dialogue captions and the channel logo, while the same story clips only render one primary caption in most news broadcast. Inspired by this fact, we propose a simple method for story segmentation based on the primary caption, which combines YOLOv3 based primary caption extraction and preliminary location of boundaries. In particular, we introduce mean hash to achieve the fast and reliable comparison for detected small-size primary caption blocks. We further incorporate scene recognition to exact the preliminary boundaries, because the primary captions always appear later than the story boundary. Experimental results on two Chinese news broadcast datasets show that our method enjoys high accuracy in terms of R, P and F1-measures. Heling Chen, Zhongyuan Wang 0001, Yingjiao Pei, Baojin Huang, Weiping Tu |
MMAsia | 2 |
| 2020 | Low-quality watermarked face inpainting with discriminative residual learningabstractMost existing image inpainting methods assume that the location of the repair area (watermark) is known, but this assumption does not always hold. In addition, the actual watermarked face is in a compressed low-quality form, which is very disadvantageous to the repair due to compression distortion effects. To address these issues, this paper proposes a low-quality watermarked face inpainting method based on joint residual learning with cooperative discriminant network. We first employ residual learning based global inpainting and facial features based local inpainting to render clean and clear faces under unknown watermark positions. Because the repair process may distort the genuine face, we further propose a discriminative constraint network to maintain the fidelity of repaired faces. Experimentally, the average PSNR of inpainted face images is increased by 4.16dB, and the average SSIM is increased by 0.08. TPR is improved by 16.96% when FPR is 10% in face verification. Zheng He 0001, Xueli Wei, Kangli Zeng, Zhen Han 0002, Qin Zou 0001, Zhongyuan Wang 0001 |
MMAsia | 6 |
| 2020 | Video scene detection based on link prediction using graph convolution networkabstractWith the development of the Internet, multimedia data grows by an exponential level. The demand for video organization, summarization and retrieval has been increasing where scene detection plays an essential role. Existing shot clustering algorithms for scene detection usually treat temporal shot sequence as unconstrained data. The graph based scene detection methods can locate the scene boundaries by taking the temporal relation among shots into account, while most of them only rely on low-level features to determine whether the connected shot pairs are similar or not. The optimized algorithms considering temporal sequence of shots or combining multi-modal features will bring parameter trouble and computational burden. In this paper, we propose a novel temporal clustering method based on graph convolution network and the link transitivity of shot nodes, without involving complicated steps and prior parameter setting such as the number of clusters. In particular, the graph convolution network is used to predict the link possibility of node pairs that are close in temporal sequence. The shots are then clustered into scene segments by merging all possible links. Experimental results on BBC and OVSD datasets show that our approach is more robust and effective than the comparison methods in terms of F1-score. Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu |
MMAsia | 2 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 6 |
| 2020 | Ultra-dense GAN for satellite imagery super-resolution
Zhongyuan Wang 0001, Kui Jiang, Peng Yi 0002, Zhen Han 0002, Zheng He 0001 |
Neurocomputing | 1 |
| 2020 | Learning latent geometric consistency for 6D object pose estimation in heavily cluttered scenes
Qingnan Li, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Single Channel multi-speaker speech Separation based on quantized ratio mask and residual network
Shanfa Ke, Ruimin Hu, Xiaochen Wang 0001, Tingzhao Wu, Zhongyuan Wang 0001 |
Multim. Tools Appl. | 6 |
| 2020 | Hierarchical dense recursive network for image super-resolution
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang |
Pattern Recognit. | 2 |
| 2020 | Saliency-Aware Convolution Neural Network for Ship Detection in Surveillance VideoabstractReal-time detection of inshore ships plays an essential role in the efficient monitoring and management of maritime traffic and transportation for port management. Current ship detection methods which are mainly based on remote sensing images or radar images hardly meet real-time requirement due to the timeliness of image acquisition. In this paper, we propose to use visual images captured by an on-land surveillance camera network to achieve real-time detection. However, due to the complex background of visual images and the diversity of ship categories, the existing convolution neural network (CNN) based methods are either inaccurate or slow. To achieve high detection accuracy and real-time performance simultaneously, we propose a saliency-aware CNN framework for ship detection, comprising comprehensive ship discriminative features, such as deep feature, saliency map, and coastline prior. This model uses CNN to predict the category and the position of ships and uses the global contrast based salient region detection to correct the location. We also extract coastline information and respectively incorporate it into CNN and saliency detection to obtain more accurate ship locations. We implement our model on Darknet under CUDA 8.0 and CUDNN V5 and use a real-world visual image dataset for training and evaluation. The experimental results show that our model outperforms representative counterparts (Faster R-CNN, SSD, and YOLOv2) in terms of accuracy and speed. Linggang Wang, Zhongyuan Wang 0001, Wan Du |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Long-Term Background Redundancy Reduction for Earth Observatory Video CodingabstractHuge earth observatory video data (EOVD) and the limited transmission bandwidth from satellites to terrestrial devices pose serious challenges to compression efficiency for satellite video. In this work, we deeply explore long-term background redundancy caused by periodical satellite revisit, and make utilization of long-term background reference (LTBR) to design a high-efficiency coding method specific for EOVD. Firstly, we use data of Google Earth as prior background knowledge and construct LTBR from it. Then, two novel prediction methods are proposed to make full use of prior information of LTBR, including color and definition correction based inter prediction with LTBR (CDIP-LTBR), structure and texture constrained intra prediction with LTBR (STIP-LTBR). Lastly, the proposed two new prediction schemes are integrated into one unified coding framework along with HEVC to achieve improved coding performance, and an improved RDO method is designed for additional prediction modes selection problem in the framework to obtain a higher prediction efficiency. Extensive experiments on real-world EOVD show that the proposed coding scheme exhibit significant improvement over HEVC and H.264, in terms of BD-Rate, BD-PSNR and rate distortion comparison. Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004, Shin'ichi Satoh 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Multi-Temporal Ultra Dense Memory Network for Video Super-ResolutionabstractVideo super-resolution (SR) aims to reconstruct the corresponding high-resolution (HR) frames from consecutive low-resolution (LR) frames. It is crucial for video SR to harness both inter-frame temporal correlations and intra-frame spatial correlations among frames. Previous video SR methods based on convolutional neural network (CNN) mostly adopt a single-channel structure and a single memory module, so they are unable to fully exploit inter-frame temporal correlations specific for video. To this end, this paper proposes a multi-temporal ultra-dense memory (MTUDM) network for video super-resolution. Particularly, we embed convolutional long-short-term memory (ConvLSTM) into ultra-dense residual block (UDRB) to construct an ultra-dense memory block (UDMB) for extracting and retaining spatio-temporal correlations. This design also reduces the layer depth by expanding the width, thus avoiding training difficulties, such as gradient exploding and vanishing under a large model. We further adopt multi-temporal information fusion (MTIF) strategy to merge the extracted temporal feature maps in consecutive frames, improving the accuracy without requiring much extra computational cost. The experimental results on extensive public datasets demonstrate that our method outperforms the state-of-the-art methods by a large margin. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Blind Quality Metric of DIBR-Synthesized Images in the Discrete Wavelet Transform DomainabstractFree viewpoint video (FVV) has received considerable attention owing to its widespread applications in several areas such as immersive entertainment, remote surveillance and distanced education. Since FVV images are synthesized via a depth image-based rendering (DIBR) procedure in the "blind" environment (without reference images), a real-time and reliable blind quality assessment metric is urgently required. However, the existing image quality assessment metrics are insensitive to the geometric distortions engendered by DIBR. In this research, a novel blind method of DIBR-synthesized images is proposed based on measuring geometric distortion, global sharpness and image complexity. First, a DIBR-synthesized image is decomposed into wavelet subbands by using discrete wavelet transform. Then, the Canny operator is employed to detect the edges of the binarized low-frequency subband and high-frequency subbands. The edge similarities between the binarized low-frequency subband and high-frequency subbands are further computed to quantify geometric distortions in DIBR-synthesized images. Second, the log-energies of wavelet subbands are calculated to evaluate global sharpness in DIBR-synthesized images. Third, a hybrid filter combining the autoregressive and bilateral filters is adopted to compute image complexity. Finally, the overall quality score is derived to normalize geometric distortion and global sharpness by the image complexity. Experiments show that our proposed quality method is superior to the competing reference-free state-of-the-art DIBR-synthesized image quality models. Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Leida Li, Zhifang Xia, Lifang Wu |
IEEE Trans. Image Process. | 2 |
| 2020 | ATMFN: Adaptive-Threshold-Based Multi-Model Fusion Network for Compressed Face HallucinationabstractAlthough tremendous strides have been recently made in face hallucination, exiting methods based on a single deep learning framework can hardly satisfactorily provide fine facial features from tiny faces under complex degradation. This article advocates an adaptive-threshold-based multi-model fusion network (ATMFN) for compressed face hallucination, which unifies different deep learning models to take advantages of their respective learning merits. First of all, we construct CNN-, GAN- and RNN-based underlying super-resolvers to produce candidate SR results. Further, the attention subnetwork is proposed to learn the individual fusion weight matrices capturing the most informative components of the candidate SR faces. Particularly, the hyper-parameters of the fusion matrices and the underlying networks are optimized together in an end-to-end manner to drive them for collaborative learning. Finally, a threshold-based fusion and reconstruction module is employed to exploit the candidates' complementarity and thus generate high-quality face images. Extensive experiments on benchmark face datasets and real-world samples show that our model outperforms the state-of-the-art SR methods in terms of quantitative indicators and visual effects. The code and configurations are released at https://github.com/kuihua/ATMFN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Ke Gu 0001, Junjun Jiang |
IEEE Trans. Multim. | 2 |
| 2019 | Long Term Background Reference Based Satellite Video CodingabstractVideo transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video coding method based on long term background reference (LTBR) to eliminate redundancy caused by periodical revisit. Firstly, data of Google Earth is used to provide prior information for establishing LTBR. Then a novel intra prediction method guided by pixels' cluster information from LTBR is introduced. Experiments demonstrate that our method outperforms HEVC and H.264 , in terms of rate-distortion, BD-PSNR and BD-Rate performance. Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004 |
ICASSP | 3 |
| 2019 | Multisource Surveillance Video Coding by Exploiting 3D and 2D KnolwedgeabstractThe rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy of pose parameter of the objects, thus leading to large prediction residuals. In this paper, a novel surveillance video coding utilizing 3D and 2D knowledge is proposed. On the one hand, we generate a knowledge based reference frame from 3D models of the objects and incorporate it into the block based coding framework to remove global redundancy while improve the robustness to pose errors. On the other hand, 2D knowledge in the form of visual appearances of the objects in the previously encoded frames is employed to rectify the knowledge based reference frame for local redundancy removal. Experimental results demonstrate the effectiveness of our proposed method against HEVC and the knowledge based coding method. Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001 |
ICASSP | 4 |
| 2019 | Rain Streak Removal via Multi-scale Mixture Exponential Power ModelabstractRain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GMM compromises the ability of the model in fitting real noise, such as sparse noise, which is exactly the characteristic of the rain streaks. In this paper, a novel model named Mixture Exponential Power Model (MEPM) is exploited. It sets multiple Laplace noise components and expands the representation capability for the sparse noise. Moreover, considering that the rain streaks in a video occur in different distances from the camera, we encode rain streaks into Multi-scale Mixture Exponential Power Model. The model is opti-mized by expectation-maximization (EM) algorithm and La-grange multiplier strategy. Experiments are implemented on synthetic and real rain videos and verify the superiority of the proposed method, compared with state-of-the-art methods. Jun Chen 0001, Zhen Han 0002, Mingfu Xiong, Chao Liang 0001, Zhongyuan Wang 0001 |
ICASSP | 7 |
| 2019 | Blind Quality Assessment for 3D-synthesized Images by Measuring Geometric Distortions and Image ComplexityabstractFree viewpoint video (FVV), owing to its comprehensive applications in immersive entertainment, remote surveillance and distanced education, has received extensive attention and been regarded as a new important direction of video technology development. Depth image-based rendering (DIBR) technologies are employed to synthesize FVV images in the "blind" environment. Therefore, a real-time reliable blind quality assessment metric is urgently required. However, existing stste-of-art quality assessment methods are limited to estimate geometric distortions generated by DIBR. In this research, a novel blind quality metric, measuring Geometric Distortions and Image Complexity (GDIC), is proposed for DIBR-synthesized images. Firstly, a DIBR-synthesized image is decomposed into wavelet subbands by using discrete wavelet transform. Then, we adopt canny operator to capture the edge of wavelet subbands and compute the edge similarity between low-frequency subband and highfrequency subbands. The edge similarity is used to quantify geometric distortions in DIBR-synthesized images. Secondly, a hybrid filter combining the autoregressive and bilateral filter is adopted to compute image complexity. Finally, the overall quality score is calculated by normalizing geometric distortions via image complexity. Experiments show that our proposed GDIC is superior to prevailing image quality assessment metrics, which were intended for natural and DIBR-synthesized images. Guangcheng Wang, Zhongyuan Wang 0001, Ke Gu 0001, Zhifang Xia |
ICASSP | 2 |
| 2019 | Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsabstractMost previous fusion strategies either fail to fully utilize temporal information or cost too much time, and how to effectively fuse temporal information from consecutive frames plays an important role in video super-resolution (SR). In this study, we propose a novel progressive fusion network for video SR, which is designed to make better use of spatio-temporal information and is proved to be more efficient and effective than the existing direct fusion, slow fusion or 3D convolution strategies. Under this progressive fusion framework, we further introduce an improved non-local operation to avoid the complex motion estimation and motion compensation (ME&MC) procedures as in previous video SR approaches. Extensive experiments on public datasets demonstrate that our method surpasses state-of-the-art with 0.96 dB in average, and runs about 3 times faster, while requires only about half of the parameters. Peng Yi 0002, Zhongyuan Wang 0001, Kui Jiang, Junjun Jiang, Jiayi Ma 0001 |
ICCV | 2 |
| 2019 | GAN-Based Multi-level Mapping Network for Satellite Imagery Super-ResolutionabstractAlthough many deep-learning-based image super-resolution (SR) methods have been proposed, most of them assume that all hierarchical features share the unified mapping equations. They ignore the differences between mapping equations at different feature levels, and create an average effect of mapping prediction, thus poorly building the mapping relations between low resolution (LR) and high resolution (HR) spaces. In this paper, we propose a multi-level mapping framework along with the adversarial learning strategy, namely MMGAN, for satellite imageries SR reconstruction. We also construct a feature extraction and tuning block (FETB) for fine feature expression. In particular, a novel two-dimension dense unit (DU) and a mapping attention unit (MAU) are constructed for building multi-level mappings in different stages. With our strategies, an HR image is reconstructed directly from the input image using multi-level mappings. Extensive experiments on Kaggle Open Source Dataset and Jilin-1 video satellite images exhibit superior reconstruction performance when compared with the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang, Guangcheng Wang, Zhen Han 0002, Tao Lu 0001 |
ICME | 2 |
| 2019 | Multi-speakers Speech Separation Based on Modified Attractor Points Estimation and GMM ClusteringabstractIn this paper, a new attractor points estimation method for DANet algorithm used in single channel multi-speaker speech separation has been proposed. A prerequisite is that there must be separate segments of each source in the mixture. This condition is met in the actual situation because the source signal is not overlapping at any time. With this prerequisite, an isolated source segments extracted from the mixture is converted to the embedding space. With the embedding of isolated source segments, a more accurately attractor point for each source will be created, due to it does not contain components of other sources. In addition, a gaussian mixture model(GMM) clustering method instead of K-means clustering method were used at run time. The experiment demonstrated that the proposed method gets a better separation performance than state of the art method up to 1.04dB in SDR. Shanfa Ke, Ruimin Hu, Tingzhao Wu, Xiaochen Wang 0001, Zhongyuan Wang 0001 |
ICME | 6 |
| 2019 | Identifying Users by Asynchronous Mobility TrajectoriesabstractWith the popularity of location-based services and applications, a large amount of mobility data has been generated. Identity recognition through mobile trajectory information, especially asynchronous trajectory data has arisen great concerns in social security prevention and control. This paper advocates an identification resolution method based on the most frequently distributed TOP-N regions regarding user trajectories. This method first finds TOP-N regions whose trajectory points are most frequently distributed so as to reduce the computational complexity. It then combines probabilistic deviation and angle cosine to calculate TOP-N region similarity between two trajectories to identify the same user. We conducted extensive experiments on two real GPS trajectory datasets GeoLife and Cabspotting and comprehensively discussed the experimental results. The experimental results show that this method is substantially effective and efficiency for user identification. Mengjun Qi, Zhongyuan Wang 0001, Zheng He 0001, Tao Lu 0001 |
IGARSS | 2 |
| 2019 | Deep Structural Feature Learning: Re-Identification of simailar vehicles In Structure-Aware Map SpaceabstractVehicle re-identification (re-ID) has received more attention in recent years as a significant work, making huge contribution to the intelligent video surveillance. The complex intra-class and inter-class variation of vehicle images bring huge challenges for vehicle re-ID, especially for the similar vehicle re-ID. In this paper we focus on an interesting and challenging problem, vehicle re-ID of the same/similar model. Previous works mainly focus on extracting global features using deep models, ignoring the individual loa-cal regions in vehicle front window, such as decorations and stickers attached to the windshield, that can be more discriminative for vehicle re-ID. Instead of directly embedding these regions to learn their features, we propose a Regional Structure-Aware model (RSA) to learn structure-aware cues with the position distribution of individual local regions in vehicle front window area, constructing a FW structural map space. In this map sapce, deep models are able to learn more robust and discriminative spatial structure-aware features to improve the performance for vehicle re-ID of the same/similar model. We evaluate our method on a large-scale vehicle re-ID dataset Vehicle-1M. The experimental results show that our method can achieve promising performance and outperforms several recent state-of-the-art approaches. Wenqian Zhu, Ruimin Hu, Zhongyuan Wang 0001, Dengshi Li, Xiyue Gao |
MMAsia | 3 |
| 2019 | Multisource surveillance video coding with synthetic reference frame
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Multisource surveillance video data coding with hierarchical knowledge library
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Liang Xu 0010, Zhongyuan Wang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Multiscale Locality and Rank Preservation for Robust Feature Matching of Remote Sensing ImagesabstractAs a fundamental and important task in many applications of remote sensing and photogrammetry, feature matching tries to seek correspondences between the two feature sets extracted from an image pair of the same object or scene. This paper focuses on eliminating mismatches from a set of putative feature correspondences constructed according to the similarity of existing well-designed feature descriptors. Considering the stable local topological relationship of the potential true correspondences, we propose a simple yet efficient method named multiscale Top K Rank Preservation (mTopKRP) for robust feature matching. To this end, we first search the K-nearest neighbors of each feature point and generate a ranking list accordingly. Then we design a metric based on the weighted Spearman's footrule distance to describe the similarity of two ranking lists specifically for the matching problem. We build a mathematical optimization model and derive its closed-form solution, enabling our method to establish reliable correspondences in linearithmic time complexity, which requires only tens of milliseconds to handle over 1000 putative matches. We also introduce a multiscale strategy for neighborhood construction, which increases the robustness of our method and can deal with different types of degradation, even when the image pair suffers from a large scale change, rotation, nonrigid deformation, or a large number of mismatches. Extensive experiments on several representative remote sensing image data sets demonstrate the superiority of our method over state of the art. Xingyu Jiang 0005, Junjun Jiang, Aoxiang Fan, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Edge-Enhanced GAN for Remote Sensing Image SuperresolutionabstractThe current superresolution (SR) methods based on deep learning have shown remarkable comparative advantages but remain unsatisfactory in recovering the high-frequency edge details of the images in noise-contaminated imaging conditions, e.g., remote sensing satellite imaging. In this paper, we propose a generative adversarial network (GAN)-based edge-enhancement network (EEGAN) for robust satellite image SR reconstruction along with the adversarial learning strategy that is insensitive to noise. In particular, EEGAN consists of two main subnetworks: an ultradense subnetwork (UDSN) and an edge-enhancement subnetwork (EESN). In UDSN, a group of 2-D dense blocks is assembled for feature extraction and to obtain an intermediate high-resolution result that looks sharp but is eroded with artifacts and noises as previous GAN-based methods do. Then, EESN is constructed to extract and enhance the image contours by purifying the noise-contaminated components with mask processing. The recovered intermediate image and enhanced edges can be combined to generate the result that enjoys high credibility and clear contents. Extensive experiments on Kaggle Open Source Data set, Jilin-1 video satellite images, and Digitalglobe show superior reconstruction performance compared to the state-of-the-art SR approaches. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Guangcheng Wang, Tao Lu 0001, Junjun Jiang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Multi-Memory Convolutional Neural Network for Video Super-ResolutionabstractVideo super-resolution (SR) is focused on reconstructing high-resolution (HR) frames from consecutive lowresolution (LR) frames. Most previous video SR methods based on convolutional neural network (CNN) use a direct connection and single-memory module within the network, and they thus fail to make full use of spatio-temporal complementary information from LR observed frames. To fully exploit spatio-temporal correlations between adjacent LR frames and reveal more realistic details, this paper proposes a multi-memory convolutional neural network (MMCNN) for video SR, cascading an optical flow network and an image-reconstruction network. A serial of residual blocks engaged in utilizing intra-frame spatial correlations are proposed for feature extraction and reconstruction. Particularly, instead of using single-memory module, we embed convolutional long short-term memory (ConvLSTM) into the residual block, thus form a multi-memory residual block to progressively extract and retain inter-frame temporal correlations between consecutive LR frames. We conduct extensive experiments on numerous testing datasets with respect to different scaling factors. Our proposed MMCNN shows superiority over the state-of-the-art methods in terms of PSNR and visual quality and surpasses the best counterpart method 1 dB at most. The code and datasets are available at https://github.com/psychopa4/MMCNN. Zhongyuan Wang 0001, Peng Yi 0002, Kui Jiang, Junjun Jiang, Zhen Han 0002, Tao Lu 0001, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Separability and Compactness Network for Image Recognition and SuperresolutionabstractConvolutional neural networks (CNNs) have wide applications in pattern recognition and image processing. Despite recent advances, much remains to be done for CNNs to learn a better representation of image samples. Therefore, constant optimizations should be provided on CNNs. To achieve a good performance on classification, intuitively, samples' interclass separability, or intraclass compactness should be simultaneously maximized. Accordingly, in this paper, we propose a new network, named separability and compactness network (SCNet) to rectify this problem. SCNet minimizes the softmax loss and the distance between features of samples from the same class under a jointly supervised framework, resulting in simultaneous maximization of interclass separability and intraclass compactness of samples. Furthermore, considering the convenience and the efficiency of the cosine similarity in face recognition tasks, we incorporate it into SCNet's distance metric to enable sample features from the same class to line up in the same direction and those from different classes to have a large angle of separation. We apply SCNet to three different tasks: visual classification, face recognition, and image superresolution. Experiments on both public data sets and real-world satellite images validate the effectiveness of our SCNet. Liguo Zhou, Zhongyuan Wang 0001, Yimin Luo, Zixiang Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Feature Matching Based on Top K Rank SimilarityabstractFeature matching plays a key component in many computer vision and pattern recognition tasks. Observing that the spatial neighborhood relationship (representing the topological structures of an image scene) is generally well preserved between two feature points of an image pair, some mismatch removing methods based on maintaining the local neighborhood structures of the potential true matches have been proposed. How to define the local neighborhood structure is an issue of vital importance. In this paper, we propose a robust and efficient method, called Top$K$Rank Preservation (Top-KRP), for mismatch removal from given putative point set matching correspondences. Instead of preserving the intersection of neighbors, TopKRP aims at preserving the top$K$rank of two feature points. The developed approach is validated on numerous challenging real image pairs for general feature matching, and the experimental results demonstrate that it outperforms several state-of-the-art feature matching methods, especially in case of a large number of mismatches. Junjun Jiang, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
ICASSP | 4 |
| 2018 | Edge-Aware Context Encoder for Image InpaintingabstractWe present Edge-aware Context Encoder (E-CE): an image inpainting model which takes scene structure and context into account. Unlike previous CE which predicts the missing regions using context from entire image, E-CE learns to recover the texture according to edge structures, attempting to avoid context blending across boundaries. In our approach, edges are extracted from the masked image, and completed by a full-convolutional network. The completed edge map together with the original masked image are then input into the modified CE network to predict the missing region. The experiments demonstrate that E-CE can generate images with better shapes and structures than CE. Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001 |
ICASSP | 4 |
| 2018 | Improving Convolutional Neural Networks Via Compacting FeaturesabstractConvolutional neural networks (CNNs) have shown great advantages in computer vision fields, and loss functions are of great significance to their gradient descent algorithms. Softmax loss, a combination of cross-entropy loss and Softmax function, is the most commonly used one for CNNs. Hence, it can continuously increase the discernibility of sample features in classification tasks. Intuitively, to promote the discrimination of CNNs, the learned features are desirable when the inter-class separability and intra-class compactness are maximized simultaneously. Since Softmax loss hardly motivates this inter-class separability and intra-class compactness simultaneously and explicitly, we propose a new method to achieve this simultaneous maximization. This method minimizes the distance between features of homogeneous samples along with Softmax loss and thus improves CNNs' performance on vision-related tasks. Experiments on both visual classification and face verification datasets validate the effectiveness and advantages of our method. Liguo Zhou, Yimin Luo, Zhongyuan Wang 0001 |
ICASSP | 5 |
| 2018 | Face Hallucination Using Manifold-Regularized Group Locality-Constrained RepresentationabstractSparsity and locality regularizations are successfully applied to face hallucination algorithms to ameliorate their ill-posed nature. However, most of patch-based face hallucination approaches only consider the manifold structure of single patch, thus resulting in unstable solution for image reconstruction. In this paper, we propose a novel face hallucination, termed manifold-regularized group locality-constrained representation (MGLR), in order to exploit the multiple manifold structures rooted in grouped self-similarly patches. Specifically, we first group similar patches to form a matrix which contains the recurrent non-local patches. Then graph regularization term is formulated to represent the group manifolds for better reconstruction quality. Taking advantages of grouped self-similar patches, MGLR can offer stable sparse solution to take advantage of the the accurate prior for super-resolution reconstruction. Experimental results on LFW database and CMU real-world images demonstrate the superiority of the proposed method over some state-of-the-art face methods both in terms of subjective and objective qualities. Tao Lu 0001, Kangli Zeng, Junjun Jiang, Yanduo Zhang, Zhongyuan Wang 0001, Huabing Zhou |
ICIP | 6 |
| 2018 | Two-Level Segment-Based Bitrate Control for Live ABR Streaming
Jing Xiao 0004, Gen Zhan, Xu Wang 0015, Zhongyuan Wang 0001 |
MMM (1) | 5 |
| 2018 | A Novel Frontal Facial Synthesis Algorithm Based on Individual Residual Face
Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001 |
MMM (2) | 4 |
| 2018 | Coarse-to-Fine Image Super-Resolution Using Convolutional Neural Networks
Liguo Zhou, Zhongyuan Wang 0001, Yimin Luo |
MMM (2) | 2 |
| 2018 | A Progressively Enhanced Network for Video Satellite Imagery SuperresolutionabstractDeep convolutional neural networks (CNNs) have been extensively applied to image or video processing and analysis tasks. For single-image superresolution (SR) processing, previous CNN-based methods have led to significant improvements, when compared to the shallow learning-based methods. However, these CNN-based algorithms with simply direct or skip connections are not suitable for satellite imagery SR because of complex imaging conditions and unknown degradation process. More importantly, they ignore the extraction and utilization of the structural information in satellite images, which is very unfavorable for video satellite imagery SR with such characteristics as small ground targets, weak textures, and over-compression distortion. To this end, this letter proposes a novel progressively enhanced network for satellite image SR called PECNN, which is composed of a pretraining CNN-based network and an enhanced dense connection network. The pretraining part is used to extract the low-level feature maps and reconstructs a basic high-resolution image from the low-resolution input. In particular, we propose a transition unit to obtain the structural information from the base output. Then, the obtained structural information and the extracted low-level feature maps are transmitted to the enhanced network for further extraction to enforce the feature expression. Finally, a residual image with enhanced fine details obtained from the dense connection network is used to enrich the basic image for the ultimate SR output. Experiments on real-world Jilin-1 video satellite images and Kaggle Open Source Dataset show that the proposed PECNN outperforms the state-of-the-art methods both in visual effects and quantitative metrics. Code is available at https://github.com/kuihua/PECNN. Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Junjun Jiang |
IEEE Signal Process. Lett. | 2 |
| 2018 | Virtual Background Reference Frame Based Satellite Video CodingabstractVideo transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellites calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video compression method to eliminate long-term redundancy among multiple periodically revisited videos, based on virtual background reference frame (VBRF) obtained from Google Earth data. First, we make full use of Google Earth data to create VBRF for representing the constant ground background. Then, we encode all I frames by referring to VBRF and performing interprediction. Experiments demonstrate that our method outperforms H.264 and HEVC, in terms of rate distortion, bjøntegaard delta peak signal-to-noise rate (BD-PSNR), and BD-Rate performance. Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Smart Monitoring Cameras Driven Intelligent Processing to Big Surveillance Video DataabstractVideo surveillance system has become a critical part in the security and protection system of modem cities, since smart monitoring cameras equipped with intelligent video analytics techniques can monitor and pre-alarm abnormal behaviors or events. However, with the expansion of the surveillance network, massive surveillance video data poses huge challenges to the analytics, storage and retrieval in the Big Data era. This paper presents a novel intelligent processing and utilization solution to big surveillance video data based on the event detection and alarming messages from front-end smart cameras. The method includes three parts: the intelligent pre-alarming for abnormal events, smart storage for surveillance video and rapid retrieval for evidence videos, which fully explores the temporal-spatial association analysis with respect to the abnormal events in different monitoring sites. Experimental results reveal that our proposed approach can reliably pre-alarm security risk events, substantially reduce storage space of recorded video and significantly speed up the evidence video retrieval associated with specific suspects. Zhongyuan Wang 0001 |
IEEE Trans. Big Data | 3 |
| 2018 | SuperPCA: A Superpixelwise PCA Approach for Unsupervised Feature Extraction of Hyperspectral ImageryabstractAs an unsupervised dimensionality reduction method, the principal component analysis (PCA) has been widely considered as an efficient and effective preprocessing step for hyperspectral image (HSI) processing and analysis tasks. It takes each band as a whole and globally extracts the most representative bands. However, different homogeneous regions correspond to different objects, whose spectral features are diverse. Therefore, it is inappropriate to carry out dimensionality reduction through a unified projection for an entire HSI. In this paper, a simple but very effective superpixelwise PCA (SuperPCA) approach is proposed to learn the intrinsic low-dimensional features of HSIs. In contrast to classical PCA models, the SuperPCA has four main properties: 1) unlike the traditional PCA method based on a whole image, the SuperPCA takes into account the diversity in different homogeneous regions, that is, different regions should have different projections; 2) most of the conventional feature extraction models cannot directly use the spatial information of HSIs, while the SuperPCA is able to incorporate the spatial context information into the unsupervised dimensionality reduction by superpixel segmentation; 3) since the regions obtained by superpixel segmentation have homogeneity, the SuperPCA can extract potential low-dimensional features even under noise; and 4) although the SuperPCA is an unsupervised method, it can achieve a competitive performance when compared with supervised approaches. The resulting features are discriminative, compact, and noise-resistant, leading to an improved HSI classification performance. Experiments on three public data sets demonstrate that the SuperPCA model significantly outperforms the conventional PCA-based dimensionality reduction baselines for HSI classification, and some state-of-the-art feature extraction approaches. The MATLAB source code is available at https://github.com/junjun-jiang/SuperPCA. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Zhongyuan Wang 0001, Zhihua Cai, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | SeaShips: A Large-Scale Precisely Annotated Dataset for Ship DetectionabstractIn this paper, we introduce a new large-scale dataset of ships, called SeaShips, which is designed for training and evaluating ship object detection algorithms. The dataset currently consists of 31 455 images and covers six common ship types (ore carrier, bulk cargo carrier, general cargo ship, container ship, fishing boat, and passenger ship). All of the images are from about 10 080 real-world video segments, which are acquired by the monitoring cameras in a deployed coastline video surveillance system. They are carefully selected to mostly cover all possible imaging variations, for example, different scales, hull parts, illumination, viewpoints, backgrounds, and occlusions. All images are annotated with ship-type labels and high-precision bounding boxes. Based on the SeaShips dataset, we present the performance of three detectors as a baseline to do the following: 1) elementarily summarize the difficulties of the dataset for ship detection; 2) show detection results for researchers using the dataset; and 3) make a comparison to identify the strengths and weaknesses of the baseline algorithms. In practice, the SeaShips dataset would hopefully advance research and applications on ship detection. Zhongyuan Wang 0001, Wan Du |
IEEE Trans. Multim. | 3 |
| 2017 | Cruise UAV Video Compression Based on Long-Term Wide-Range BackgroundabstractWith the rapid development of Unmanned Aerial Vehicle (UAV), the compression of video data captured by UAV has become a growing critical issue. However, most advanced coding schemes, like H.264 and HEVC, are oriented for common videos and thus cannot afford ideal coding efficiency when applied to UAV platform. Considering the characteristics of UAV video, much more improvement could be imposed onto current coding schemes to make full use of UAV's sensor information. In this paper, we exploit long term redundancy existing in the video data captured by cruise UAV. Firstly, we establish a long-term wide-range background set for reference. Then we separate each frame into new-area part and overlapped part. Lastly, we use GPS information of each frame to get reference from background set and compress two parts individually. In the experiments, by comparing to standard HEVC, our method has given more than 20% reduction in bitrate and meanwhile more than 4% gain in PSNR. Xu Wang 0015, Jing Xiao 0004, Ruimin Hu, Zhongyuan Wang 0001 |
DCC | 4 |
| 2017 | A joint learning based Face Super Resolution approach via contextual topological structureabstractFace Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are with low quality, the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we introduce the contextual topological structure of target patch to improve the consistency. The contextual topological structure consists of the target patch as well as its adjacent patches, we explore the relationship between them based on statistical probability and apply the relationship for joint learning progress of mapping from LR to HR. By incorporating the contextual topological structure, the robustness to noise of approach is increased as well as the LR/HR consistency. The effectiveness of proposed method is verified both quantitatively and qualitatively. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001 |
ICASSP | 4 |
| 2017 | Efficient mode decision for noisy video transcodingabstractIn many practical transcoding applications, such as video surveillance, the source videos are often contaminated by noise. The presence of noise not only results in poor compression efficiency and visual quality, but also imposes an adverse effect on the performance of subsequent video analysis tasks. Thereby it is very necessary to denoise the video. In this paper, we propose an efficient mode decision method for H.264 noisy video transcoding through analysing the effect of noise on mode decision. In the algorithm, we use the information available from previously decoded MBs to decide which modes can be overpassed with little loss to the rate-distortion performance. Experimental results show that our method saves the computational complexity nearly 65%, without noticeable rate-distortion(R-D) loss in comparison with the “Full-coding” (cascaded transcoding) method. Anna Zhang, Zhen Han 0002, Zhongyuan Wang 0001 |
ICASSP | 3 |
| 2017 | Face hallucination using region-based deep convolutional networksabstractMost deep learning based face hallucinations exploit random patch prior from training samples, then to learn the mapping functions between low-resolution (LR) and high-resolution (HR) images, and achieve satisfactory reconstruction performance. However, most of them do not take into account the prior information on facial structure, which is pivotal for face hallucination. Different from random patch prior based deep learning approaches, in this paper, we utilize facial structural prior and develop a simple yet powerful face hallucination, named region-based deep convolutional networks (RDCN). Firstly, we divide facial image into several regions of interest, then to train multiple parallel subnetworks of these regions for exacting better structure priors, finally HR output is reconstructed by stitching facial parts. Experiments on the FEI database demonstrate that the proposed region-based convolution networks outperform other state-of-the-art, including recently proposed deep learning based approaches, both in subjective and objective reconstruction qualities. Tao Lu 0001, Hao Wang 0237, Zixiang Xiong, Junjun Jiang, Yanduo Zhang, Huabing Zhou, Zhongyuan Wang 0001 |
ICIP | 7 |
| 2017 | Non-rigid feature matching for image retrieval using global and local regularizationsabstractIn this paper, we propose a probabilistic method for feature matching of near-duplicate images undergoing non-rigid transformations. We start by creating a set of putative correspondences based on the feature similarity, and then focus on removing outliers from the putative set and estimating the transformation as well. This is formulated as a maximum likelihood estimation of a Bayesian model with latent variables indicating whether matches in the putative set are inliers or outliers. We impose the non-parametric global geometrical constraints on the correspondence using Tikhonov regularizers in a reproducing kernel Hilbert space. We also introduce a local geometrical constraint to preserve local structures among neighboring feature points. The problem is solved by using the Expectation Maximization algorithm, and the closed-form solution of the transformation is derived in the maximization step. Moreover, a fast implementation based on sparse approximation is given which reduces the method computation complexity to linearithmic without performance sacrifice. Extensive experiments on real near-duplicate images for both feature matching and image retrieval demonstrate accurate results of the proposed method which outperforms current state-of-the-art methods, especially in case of severe outliers. Yong Ma 0001, Huabing Zhou, Jun Chen 0001, Jingshu Shi, Zhongyuan Wang 0001 |
ICME | 5 |
| 2017 | Video Satellite Imagery Super Resolution via Convolutional Neural NetworksabstractVideo satellite imagery is a new technique for earth dynamic observation and has a wide range of uses in environmental fields. Despite its capability of dynamic targets' detection, it sustains a serious restriction of the image quality due to the degradation and compression in its imaging process. Hence, the super-resolution (SR) reconstruction on these compressed low-spatial-resolution images is of significance to afterward ground objects recognition and detection tasks. Based on the recent proposed state-of-the-art convolutional neural networks (CNNs) SR methods, we proposed an SR method which could get more precise reconstructed high-spatial-resolution images. Trained with Gaofen-2 satellite images, a robust CNN model specified in satellite image SR is obtained. Experimentally, the reconstruction results on Jilin-1 mission satellite images validate the effectiveness of our method. Yimin Luo, Liguo Zhou, Zhongyuan Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | A sensitive object-oriented approach to big surveillance data compression for social security applications in smart citiesabstractSummary Surveillance has become a fairly common practice with the global boom in “smart cities”. How to efficiently store and manage the vast quantities of surveillance data is a persistent challenge in terms of analyzing social security problems. Developing data compression technology under the analytic requirements of surveillance data is the key to solving the storage problem. Criminal investigation demands the quality preservation of sensitive objects, typically pedestrians, human faces, vehicles, and license plates; however, the analytical value of surveillance data is rapidly lost as the compression ratio increases. In this paper, we propose a sensitive object‐oriented regions of interest‐based coding strategy for preserving the analytical value of surveillance data. In the proposed method, instead of generating a saliency map based on human visual perception, we consider saliency as a set of characteristics important for object detection and recognition. By making this modification, almost all sensitive objects necessary in a criminal investigation are assigned high saliency value rather than only one or two salient regions. Motions in the temporal domain are integrated to place emphasis on moving objects, namely moving sensitive objects, which then gain the highest saliency. Finally, a saliency‐based rate control algorithm embedded in High Efficiency Video Coding is used to maintain the quality of sensitive objects in the encoded video under a fixed bitrate. Experiments were conducted on two analytical indexes: Feature similarity and object detection accuracy. The results showed that by achieving the same feature similarity and object detection accuracy, our method can save 20% and 40% bitrate over High Efficiency Video Coding, respectively, for the storage of big surveillance data. Copyright © 2016 John Wiley & Sons, Ltd. Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021, Jun Xiao 0004, Gen Zhan, Ruimin Hu |
Softw. Pract. Exp. | 2 |
| 2017 | SRLSP: A Face Image Super-Resolution Algorithm Using Smooth Regression With Local Structure PriorabstractThe performance of traditional face recognition systems is sharply reduced when encountered with a low-resolution (LR) probe face image. To obtain much more detailed facial features, some face super-resolution (SR) methods have been proposed in the past decade. The basic idea of a face image SR is to generate a high-resolution (HR) face image from an LR one with the help of a set of training examples. It aims at transcending the limitations of optical imaging systems. In this paper, we regard face image SR as an image interpolation problem for domain-specific images. A missing intensity interpolation method based on smooth regression with a local structure prior (LSP), named SRLSP for short, is presented. In order to interpolate the missing intensities in a target HR image, we assume that face image patches at the same position share similar local structures, and use smooth regression to learn the relationship between LR pixels and missing HR pixels of one position patch. Performance comparison with the state-of-the-art SR algorithms on two public face databases and some real-world images shows the effectiveness of the proposed method for a face image SR in general. In addition, we conduct a face recognition experiment on the extended Yale-B face database based on the super-resolved HR faces. Experimental results clearly validate the advantages of our proposed SR method over the state-of-the-art SR methods in face recognition application. Junjun Jiang, Chen Chen 0001, Jiayi Ma 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu |
IEEE Trans. Multim. | 5 |
| 2017 | Single Image Super-Resolution via Locally Regularized Anchored Neighborhood Regression and Nonlocal MeansabstractThe goal of learning-based image super resolution (SR) is to generate a plausible and visually pleasing high-resolution (HR) image from a given low-resolution (LR) input. The SR problem is severely underconstrained, and it has to rely on examples or some strong image priors to reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e., projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. Encouraged by recent developments in image prior modeling, where the state-of-the-art algorithms are formed with nonlocal self-similarity and local geometry priors, we seek an SR algorithm of similar nature that will incorporate these two priors into the learning from LR space to HR space. The nonlocal self-similarity prior takes advantage of the redundancy of similar patches in natural images, while the local geometry prior of the data space can be used to regularize the modeling of the nonlinear relationship between LR and HR spaces. Based on the above two considerations, we first apply the local geometry prior to regularize the patch representation, and then utilize the nonlocal means filter to improve the super-resolved outcome. Experimental results verify the effectiveness of the proposed algorithm compared with the state-of-the-art SR methods. Junjun Jiang, Chen Chen 0001, Tao Lu 0001, Zhongyuan Wang 0001, Jiayi Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Person Re-Identification via Multiple Coarse-to-Fine Deep MetricsabstractPerson re-identification, aiming to identify images of the same person from various cameras views in different places, has attracted a lot of research interests in the field of artificial intelligence and multimedia. As one of its popular research directions, the metric learning method plays an important role for seeking a proper metric space to generate accurate feature comparison. However, the existing metric learning methods mainly aim to learn an optimal distance metric function through a single metric, making them difficult to consider multiple similar relationships between the samples. To solve this problem, this paper proposes a coarse-to-fine deep metric learning method equipped with multiple different Stacked Auto-Encoder (SAE) networks and classification networks. In the perspective of the human's visual mechanism, the multiple different levels of deep neural networks simulate the information processing of the brain's visual system, which employs different patterns to recognize the character of objects. In addition, a weighted assignment mechanism is presented to handle the different measure manners for final recognition accuracy. The experimental results conducted on two public datasets, i.e., VIPeR and CUHK have shown the prospective performance of the proposed method. Mingfu Xiong, Jun Chen 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu, Chao Liang 0001, Daming Shi 0001 |
ECAI | 4 |
| 2016 | L1-L1 norms for face super-resolution with mixed Gaussian-impulse noiseabstractIn real world surveillance application, the captured faces are often low resolution (LR) and corrupted by mixed Gaussian-impulse noise during the acquisition and transmission processes. In this paper, we propose an effective patch-based face super-resolution method to reconstruct a high resolution (HR) face image given an LR observation that is corrupted by mixed Gaussian-impulse noise. To represent the corrupted image patches, a sparse regularization combined with an l\ data fitting term is proposed. In the proposed model, both the patch reconstruction term and the regularization term are in the l\ norm form. As a result, the model is called norms. In addition, since image pixels have nonnegative intensities, we further add a nonnegative constraint to the patch representation model. Experimental results demonstrate that the proposed norms based method can achieve superior face super-resolution performance over several state-of-the-art approaches based on the objective results in terms of P-SNR, as well as the visual perceptual quality. Junjun Jiang, Zhongyuan Wang 0001, Chen Chen 0001, Tao Lu 0001 |
ICASSP | 2 |
| 2016 | Hyperspectral image denoising based on low-rank representation and superpixel segmentationabstractRecently, low-rank representation (LRR) based methods have been used for hyperspectral image (HSI) denoising, which can simultaneously remove different types of noise: Gaussian noise, impulse noise, dead lines, and so on. However, the LRR based method does not make full use of the spatial information in HSI. In this paper, we integrate the superpixel segmentation (SS) into the LRR, and propose a novel denoising method named SS-LRR. We first use the principle component analysis (PCA) to obtain the first principle component of HSI. Then the superpixel segmentation is adopted to the first principle component of HSI to get homogeneous regions. Finally, we employ the LRR to each homogeneous region of HSI, which enable us to simultaneously remove all the above mentioned mixed noise. Extensive experiments on both simulated and real hyperspectral images demonstrate that the proposed SS-LRR is efficient for HSI denoising. Jiayi Ma 0001, Chang Li 0001, Yong Ma 0001, Zhongyuan Wang 0001 |
ICIP | 4 |
| 2016 | Fast video enhancement transcodingabstractIn this paper, we pose a new problem of video enhancement transcoding, which converts the compressed dark video into compressed normal-lighting one. Distinct statistics of dark and normal videos result in quite different coding modes, which thus enforces latent constraints on mode conversion during transcoding. Following this idea, we propose a fast mode decision algorithm to speed up computation while maintaining rate-distortion (RD) performance. Experimental results show that our method saves the computational complexity nearly 70%, without noticeable RD loss in comparison with the cascaded decoder-encoder approach. Kefan Shen, Zhongyuan Wang 0001, Zhen Han 0002 |
ICIP | 2 |
| 2016 | Depth image in-loop filter via graph cutabstractWith the ability of representing 3D scene geometry, depth maps are used to synthesize virtual views in free view video (FVV) or 3DTV. However, compression artefact of the depth images always lead to seriously geometry distortions in synthesized view, which severely affects the visual perception of 3D display. To solve the problem caused by depth artifact, the bilateral filter based method is presented to alleviate the noise by local weighted sum of neighboring pixels. However, to overcome the unstable feature of local weight of noisy pixels, we propose a novel graph cuts algorithm for depth filter with the constraint of corresponding structure information in color image. As a depth in-loop filter, the filter is incorporated into the framework of H.264/MVC. The proposed approach offers 0.5dB and 0.8dB average PSNR gains in terms of video rendering quality and depth coding efficiency comparing with the state-of-the-art. Liguo Zhou, Zhongyuan Wang 0001, Youming Fu, Jun Chen 0001, Rui Xiang, Shizheng Wang |
ICIP | 2 |
| 2016 | Kinect Depth Holes Filling by Similarity and Position Constrained Sparse Representation
Zhongyuan Wang 0001, Ruolin Ruan |
ICISP | 2 |
| 2016 | Face Super Resolution for VLQ facial images via parent patch matchingabstractFace Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are Very Low Quality(VLQ), the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we use the target patch as well as the surrounding pixels, which we called parent patch, to represent the target patch. By incorporating the peripheral information, the parent patch is much more robust to noise in the LR and HR consistency learning. The effectiveness of proposed method is verified both quantitatively and qualitatively. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001, Zheng Lu 0002 |
IJCNN | 4 |
| 2016 | Spatiotemporal saliency based on location prior modelabstractSaliency detection for images and videos becomes increasingly popular due to its wide applicability. Enormous research efforts have been focused on saliency detection, but it still has some issues in maintaining spatiotemporal consistency of videos and uniformly highlighting entire objects. To address these issues, this paper proposes a superpixel-level spatiotemporal saliency model for saliency detection in videos. To detect salient object, we extract multiple spatiotemporal features combined with intra-consistency motion information preliminarily. Meanwhile, considering inter-consistency of foreground in videos, a set of foreground locations are obtained from previous frames. Then, we introduce foreground-background and local foreground contrast saliency cues of those features using the location prior information of foreground. These two improved contrast saliency cues uniformly highlight the entire object and suppress the background effectively. Finally, we use an interactively dynamic fusion method to integrate the output spatial and temporal saliency maps. The proposed approach is validated on challenging sets of video sequences. Subjective observations and objective evaluations demonstrate that the proposed model achieves a better performance on saliency detection compared with the state-of-the-art spatiotemporal saliency methods. Liuyi Hu, Zhongyuan Wang 0001, Mang Ye, Jing Xiao 0004, Ruimin Hu |
IJCNN | 2 |
| 2016 | Smooth sparse representation for noise robust face super-resolutionabstractFace super-resolution has attracted much attention in recent years. Many algorithms have been proposed. Among them, sparse representation based face super-resolution approaches are able to achieve competitive performance. However, these sparse representation based approaches only perform well under the condition that the input is noiseless or has small noise. When the input is corrupted by large noise, the reconstruction weights of the input LR patches using sparse representation based approaches will be seriously unstable, thus leading to poor reconstruction results. To this end, in this paper, we propose a novel sparse representation based face super-resolution approach that incorporates a smooth prior to enforce similar training patches having similar sparse coding coefficients. Specifically, we introduce the fused Lasso to the least squares representation of the input LR image in order to obtain a stable sparse representation, especially when the noise level of the input LR image is high. Experiments are carried out on the benchmark FEI face dataset. Visual and quantitative comparisons show that the proposed face super-resolution method achieves comparable performance to the state-of-the-art methods under noiseless condition, and yields superior super-resolution results when the input LR face image is contaminated by strong noise. Junjun Jiang, Jiayi Ma 0001, Chen Chen 0001, Zhongyuan Wang 0001, Tao Lu 0001 |
VCIP | 4 |
| 2016 | Sparse unmixing of hyperspectral data based on robust linear mixing modelabstractRecently, sparse unmixing (SU) of hyperspectral data has received particular attention for analyzing remote sensing images. However, most of SU methods are based on the commonly admitted linear mixing model (LMM), which ignores the possible nonlinear effects (i.e. nonlinearity). In this paper, we proposed a new method named robust collaborative sparse regression (RCSR) for hyperspectral unmixing, which is based on the robust LMM (rLMM). The rLMM takes the nonlinearity into consideration, and the nonlinearity is merely treated as outlier, which has the underlying sparse property. The RCSR takes the collaborative sparse property of the abundance and sparsely distributed additive property of the outlier into consideration, which can be formed as a robust joint sparse regression problem. Experiments on synthetic datasets demonstrate that the proposed RCSR is efficient for solving the hyperspectral SU problem compared with other five state-of-the-art algorithms. Chang Li 0001, Yong Ma 0001, Yuan Gao 0015, Zhongyuan Wang 0001, Jiayi Ma 0001 |
VCIP | 4 |
| 2016 | Filling Kinect depth holes via position-guided matrix completion
Zhongyuan Wang 0001, Shizheng Wang, Jing Xiao 0004, Ruimin Hu |
Neurocomputing | 1 |
| 2016 | Heteroskedasticity tuned mixed-norm sparse regularization for face hallucination
Zhongyuan Wang 0001, Ruimin Hu, Junjun Jiang, Zhen Han 0002 |
Multim. Tools Appl. | 1 |
| 2016 | CDMMA: Coupled discriminant multi-manifold analysis for matching low-resolution face images
Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhihua Cai |
Signal Process. | 3 |
| 2016 | Facial Image Hallucination Through Coupled-Layer Neighbor EmbeddingabstractAs the facial image captured by a low-cost camera is typically very low resolution (LR), blurring, and noisy, traditional neighbor-embedding-based facial image hallucination methods from one single manifold (i.e., the LR image manifold) fail to reliably estimate the intention geometrical structure, consequently leading to a bias to the image reconstruction result. In this paper, we introduce the notion of neighbor embedding (NE) from the LR and the high-resolution (HR) image manifolds simultaneously and propose a novel NE model, termed the coupled-layer NE (CLNE), for facial image hallucination. CLNE differs substantially from other NE models in that it has two layers: the LR and the HR layers. The LR layer in this model is the local geometrical structure of the LR patch manifold, which is characterized by the reconstruction weights of the LR patches; the HR layer is the intrinsic geometry that can geometrically constrain the reconstruction weights. With this coupled-constraint paradigm between the adaptation of the LR layer and the HR one, CLNE can achieve a more robust NE through iteratively updating the LR patch reconstruction weights and the estimated HR patch. The experimental results in simulation and real conditions confirm that the proposed method outperforms the related state-of-the-art methods in both quantitative and visual comparisons. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Jiayi Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Knowledge-Based Coding of Objects for Multisource Surveillance Video DataabstractGlobal object redundancy (GOR), as opposed to local spatial/temporal redundancies in a single video clip, is a new form of redundancy common in multisource surveillance video data (MSVD). GOR is induced by the repetition of foreground objects across multiple cameras, and becomes influential as the number of objects increases. Eliminating GOR considerably improves MSVD coding efficiency. In an effort to accomplish this, this study first proposes a knowledge-based representation of objects based on careful analysis of GOR composition. The representation contains a constant part and a variational part: the former is used to represent the common knowledge shared by an object across multiple cameras, while the latter is used to represent local variations on the object's surfaces. Based on the proposed representation, a knowledge-based coding (KBC) method is then proposed in which each foreground object is encoded with a hybrid prediction scheme, where the constant part of the object is generated via global prediction from a model library and the variational part is predicted via local reference frames with pose-based, short-term prediction. Experimental results showed that the KBC method saves more than 39% bits on average for encoding foreground objects in high-resolution video clips (compared to 16% for the entire videos). Applying the proposed coding method to surveillance videos in large spatial and temporal scale allows storage savings at the PB level. Jing Xiao 0004, Ruimin Hu, Yu Chen 0021, Zhongyuan Wang 0001, Zixiang Xiong |
IEEE Trans. Multim. | 5 |
| 2015 | Face hallucination via Cauchy regularized sparse representationabstractIn dictionary-learning-based face hallucination, the testing image is represented as a linear combination of the training samples, and how to obtain the optimal coefficients is the primary issue. Sparse representation (SR) has ever been widely used in face hallucination, however, due to the fact that SR overemphasizes the sparsity, the obtained linear combination coefficients turn out far aggressively sparse, then leading to unsatisfactory hallucinated results. In this paper, we present a moderately sparse prior model for face hallucination problem with the L1 norm penalty in classic SR replaced by a Cauchy penalty term. An iterative optimization is further presented to solve the minimization of Cauchy regularized objective function. The experimental results on public face database demonstrate that our method is much more effective than state-of-the-art methods. Shenming Qu, Ruimin Hu, Zhongyuan Wang 0001, Junjun Jiang |
ICASSP | 4 |
| 2015 | Locally regularized Anchored Neighborhood Regression for fast Super-ResolutionabstractThe goal of learning-based image Super-Resolution (SR) is to generate a plausible and visually pleasing High-Resolution (HR) image from a given Low-Resolution (LR) input. The problem is dramatically under-constrained, which relies on examples or some strong image priors to better reconstruct the missing HR image details. This paper addresses the problem of learning the mapping functions (i.e. projection matrices) between the LR and HR images based on a dictionary of LR and HR examples. One recently proposed method, Anchored Neighborhood Regression (ANR) [1], provides state-of-the-art quality performance and is very fast. In this paper, we propose an improved variant of ANR, namely Locally regularized Anchored Neighborhood Regression (LANR), which utilizes the locality-constrained regression in place of the ridge regression in ANR. LANR assigns different freedom for each neighbor dictionary atom according to its correlation to the input LR patch, thus the learned projection matrices are much more flexible. Experimental results demonstrate that the proposed algorithm performs efficiently and effectively over state-of-the-art methods, e.g., 0.1–0.4 dB in term of PSNR better than ANR. Junjun Jiang, Jican Fu, Tao Lu 0001, Ruimin Hu, Zhongyuan Wang 0001 |
ICME | 5 |
| 2015 | Dictionary based surveillance image compression
Jing-Ya Zhu, Zhongyuan Wang 0001, Shenming Qu |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | 3D hybrid just noticeable distortion modeling for depth image-based rendering
Ruimin Hu, Zhongyuan Wang 0001, Shizheng Wang |
Multim. Tools Appl. | 3 |
| 2015 | Trilateral constrained sparse representation for Kinect depth hole filling
Zhongyuan Wang 0001, Shizheng Wang, Tao Lu 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | Face hallucination via re-identified K-nearest neighbors embeddingabstractBased on locally linear embedding (LLE) manifold learning theory, which assumes that the low-resolution (LR) manifold and high-resolution (HR) manifold spaces share the same local geometry structure, neighbor embedding based super-resolution(SR) methods search K-nearest neighbors(K-NN) of LR patch, then use the counterpart HR patches to estimate HR patch. The primary issue of these methods is how to search the optimal K-NN. However, due to the “one-to-many” mapping between the LR image and HR ones in practice, the neighborhood relationship of the LR patch in LR space is very different with its HR counterpart's. In this paper, we explore a novel and effective re-identified K-NN(RIKNN) method to search neighbors of LR patch by taking into consideration the neighbor information in the HR space. It searches K-NN of LR patch in the LR space and then refines the searching results by re-identifying in the HR space, thus giving rise to accurate K-NN and improvement performance. Experimental results with application to face hallucination demonstrate that our method outperforms state of the art in terms of subjective and objective results and computational complexity. Shenming Qu, Ruimin Hu, Junjun Jiang, Zhongyuan Wang 0001, Jun Chen 0001 |
ICME | 5 |
| 2014 | Cloud Model-Based Dynamic Texture Synthesis for Video CodingabstractThis paper presents a novel cloud model which is designed for the inter prediction coding using virtual frame technique. The virtual frame obtained by typical dynamic texture synthesis methods can have a better prediction result in some regions of the encoding frame than the common reference frames, because the non-linear motion and global illumination change between frames is taken into account. However, there are still many limiting factors for current dynamic texture models, which make the video scenes in virtual frames can't better reflect the motion and changing trend of the real scenes. The prediction performance of virtual frame is to some degree reduced or covered up. The proposed method firstly combined the powerful computing capabilities of cloud platform and the extrapolation process of dynamic texture synthesis for video codec. Results show that this method can be used as a significant complement to the current virtual frame technique. Ruimin Hu, Zhongyuan Wang 0001 |
ICPR | 3 |
| 2014 | Noise robust face hallucination employing Gaussian-Laplacian mixture model
Zhongyuan Wang 0001, Zhen Han 0002, Ruimin Hu, Junjun Jiang |
Neurocomputing | 1 |
| 2014 | Generating algorithm for integer DST radixes in video coding
Zhongyuan Wang 0001, Ruimin Hu |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Fast Synopsis for Moving Objects Using Compressed VideoabstractWith the increasing volume of video data, how to analyze and browse video in a fast and effective way has become an urgent problem in applications. This letter proposes a novel video synopsis method in compressed domain for browsing video captured by static cameras. Synopsis video is a video abstraction, which displays moving objects from different periods simultaneously on the primary background contents of original video. To overcome the low efficiency of traditional video synopsis for compressed video, our method presents a new graph cut algorithm to extract objects tubes and meanwhile gives a fast solution to minimize energy function in compressed domain. Experimental results in H.264 video have demonstrated the high-efficiency of this new video synopsis scheme for massive video browsing. Ruimin Hu, Zhongyuan Wang 0001, Shizheng Wang |
IEEE Signal Process. Lett. | 3 |
| 2014 | Face Hallucination Via Weighted Adaptive Sparse RegularizationabstractSparse representation-based face hallucination approaches proposed so far use fixed ℓ1norm penalty to capture the sparse nature of face images, and thus hardly adapt readily to the statistical variability of underlying images. Additionally, they ignore the influence of spatial distances between the test image and training basis images on optimal reconstruction coefficients. Consequently, they cannot offer a satisfactory performance in practical face hallucination applications. In this paper, we propose a weighted adaptive sparse regularization (WASR) method to promote accuracy, stability and robustness for face hallucination reconstruction, in which a distance-inducing weighted ℓqnorm penalty is imposed on the solution. With the adjustment to shrinkage parameter q , the weighted ℓqpenalty function enables elastic description ability in the sparse domain, leading to more conservative sparsity in an ascending order of q . In particular, WASR with an optimal q > 1 can reasonably represent the less sparse nature of noisy images and thus remarkably boosts noise robust performance in face hallucination. Various experimental results on standard face database as well as real-world images show that our proposed method outperforms state-of-the-art methods in terms of both objective metrics and visual quality. Zhongyuan Wang 0001, Ruimin Hu, Shizheng Wang, Junjun Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Face Super-Resolution via Multilayer Locality-Constrained Iterative Neighbor Embedding and Intermediate Dictionary LearningabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) manifolds are locally isometric, the neighbor embedding super-resolution algorithms try to preserve the geometry (reconstruction weights) of the LR space for the reconstructed HR space, but neglect the geometry of the original HR space. Due to the degradation process of the LR image (e.g., noisy, blurred, and down-sampled), the neighborhood relationship of the LR space cannot reflect the truth. To this end, this paper proposes a coarse-to-fine face super-resolution approach via a multilayer locality-constrained iterative neighbor embedding technique, which intends to represent the input LR patch while preserving the geometry of original HR space. In particular, we iteratively update the LR patch representation and the estimated HR patch, and meanwhile an intermediate dictionary learning scheme is employed to bridge the LR manifold and original HR manifold. The proposed method can faithfully capture the intrinsic image degradation shift and enhance the consistency between the reconstructed HR manifold and the original HR manifold. Experiments with application to face super-resolution on the CAS-PEAL-R1 database and real-world images demonstrate the power of the proposed algorithm. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 3 |
| 2014 | Noise Robust Face Hallucination via Locality-Constrained RepresentationabstractRecently, position-patch based approaches have been proposed to replace the probabilistic graph-based or manifold learning-based models for face hallucination. In order to obtain the optimal weights of face hallucination, these approaches represent one image patch through other patches at the same position of training faces by employing least square estimation or sparse coding. However, they cannot provide unbiased approximations or satisfy rational priors, thus the obtained representation is not satisfactory. In this paper, we propose a simpler yet more effective scheme called Locality-constrained Representation (LcR). Compared with Least Square Representation (LSR) and Sparse Representation (SR), our scheme incorporates a locality constraint into the least square inversion problem to maintain locality and sparsity simultaneously. Our scheme is capable of capturing the non-linear manifold structure of image patch samples while exploiting the sparse property of the redundant data representation. Moreover, when the locality constraint is satisfied, face hallucination is robust to noise, a property that is desirable for video surveillance applications. A statistical analysis of the properties of LcR is given together with experimental results on some public face databases and surveillance images to show the superiority of our proposed scheme over state-of-the-art face hallucination approaches. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
IEEE Trans. Multim. | 3 |
| 2013 | LBP-Guided Depth Image FilterabstractThe multi-view video plus depth (MVD) format has been put forward for the call for proposals in free view video (FVV) and 3DTV. Since representing the 3D scene geometry, depth maps are used for synthesizing virtual views. However, compression artifacts of the depth images always lead to geometry distortions in synthesized views. By exploiting LBP features of the corresponding color samples, we propose a novel local binary pattern (LBP) guided depth filter which enables the local neighborhood samples those are in the same object of the current pixel to be filtering input. In recognition of its ability for describing the object edges, the LBP operator is used to calculate the weighted values of the local depth pixels for the depth-map filter. Furthermore, the filter is incorporated into the framework of H.264/MVC as an in-loop filter. The experimental results demonstrate that the proposed approach offers 0.45dB and 0.66dB average PSNR gains in terms of video rendering quality and depth coding efficiency, as well as significant subjective improvement in rendering views. Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
DCC | 3 |
| 2013 | Manifold regularized sparse support regression for single image super-resolutionabstractIn this paper, we present a novel single image super-resolution method. To simultaneously improve the resolution and perceptual image quality, we bring forward a practical solution combining manifold regularization and sparse support regression. The main contribution of this paper is twofold. Firstly, a mapping function from low resolution (LR) patches to high-resolution (HR) patches will be learned by a local regression algorithm called sparse support regression, which can be constructed from the support bases of the LR-HR dictionary. Secondly, we propose to preserve the geometrical structure of the image patch dictionary, which is critical for reducing the artifacts and obtaining better visual quality. Experimental results demonstrate that the proposed method produces high quality results both quantitatively and perceptually. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002, Shi Dong 0004 |
ICASSP | 3 |
| 2013 | Face hallucination via weighted sparse representationabstractBy incorporating the priors of image positions, position-patch based face hallucination methods can produce high-quality results and save computation time. These methods represent the test image patch as a linear combination of the same position patches in a training dictionary, and the key issue is how to obtain the optimal coefficients. Due to stability and accuracy issues, methods based on least square estimation or sparse representation (SR) proposed so far are not satisfactory. In this paper, we improve existing SR methods by exploiting similarity between the test and training patches. In particular, we impose a similarity constraint (in terms of the distance between the test patch and bases in the dictionary) on the ℓ1minimization regularization term and obtain the coefficients by solving a weighted SR problem. We also provide a new prospective on weighted SR and investigate its robustness to illumination variations. Experiments on commonly used database demonstrate that our method outperforms state of the art. Zhongyuan Wang 0001, Junjun Jiang, Zixiang Xiong, Ruimin Hu |
ICASSP | 1 |
| 2013 | Kinect depth map based enhancement for low light surveillance imageabstractHigh noise level from darkness and low dynamic range are two characteristics of low light surveillance image that severely degrade the visual quality. Traditional low light image enhancement methods merely use the 2D cues without the depth information of the scene. Recently, the depth based image enhancement methods are proposed to enhance the depth perception of the image. However, these depth based methods are focus on the normal light image and only enhance the local depth perception. In this paper, based on the characteristics that the depth map captured by Kinect is less affected by low light condition than color image, we propose a Kinect depth based enhancement algorithm to enlarge the dynamic range and meanwhile to enhance the depth perception for the low light surveillance image. In our algorithm, firstly, the depth level similarity is incorporated into the non-local means denoising to remove the noises while better preserve object edges. Then, the depth aware contrast stretching is performed to enlarge the dynamic range and meanwhile to enhance both globe and local depth perception for low light surveillance image. Experimental results on low light surveillance images show that our proposed algorithm achieves better perceptual quality than previous work. Ruimin Hu, Zhongyuan Wang 0001, Mang Duan |
ICIP | 3 |
| 2013 | Locality-constraint iterative neighbor embedding for face hallucinationabstractBased on the assumption that low-resolution (LR) and high-resolution (HR) patch manifolds are locally isometric, the neighbor embedding based super-resolution algorithms try to preserve the local geometry of the patch manifold for the reconstructed HR patch manifold. However, due to “one-to-many” mappings between LR and HR images, the neighborhood relationship of the LR patch manifold can't reflect the inherent data structure. In this paper, we explore the data structure by both considering the LR patch and HR patch manifolds instead of only considering one manifold (LR patch manifold). By incorporating the position prior of face and local geometry of HR patch manifold, we propose an improved neighbor embedding method to face hallucination, namely locality-constraint iterative neighbor embedding (LINE), in which we iteratively update the K-nearest neighbors (K-NN) and reconstruction weights based on the result (the hallucinated HR patch) from previous iteration, giving rise to improved performance compared with traditional neighbor embedding algorithms. Experimental results with application to face hallucination on simulated LR face images and real world ones demonstrate the effectiveness of the proposed method. Junjun Jiang, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Tao Lu 0001, Jun Chen 0001 |
ICME | 4 |
| 2013 | Face hallucination based on stepwise sparse reconstructionabstractFace hallucination methods based on low-resolution (LR) and high-resolution (HR) dictionary pair scheme infer HR patches by directly reusing coding coefficients trained by LR patches over LR dictionary. This scheme implies that LR and HR patch manifolds share highly similar local geometric structure. However, latest preliminary studies argue that the manifold assumption does not hold well such that face hallucination performance inevitably suffers from inconsistency of coding coefficients between LR and HR patches. In this paper, we are the first to observe that coding coefficients of LR patches are more relevant to latent those of HR patches under conditions of involving small magnifying factor. On the basis of this finding, we suggest a stepwise reconstruction scheme to minimize inconsistency risk in solution space. In particular, this scheme divides face hallucination process into multiple cascaded incremental training-synthesis steps, in which each individual step allows smaller magnifying factor as well as the corresponding intermediate resolution (IR) dictionary rather than merely LR and HR dictionary based learning. Moreover, in order to keep sparse representation (SR) sufficiently sparse while favoring its locality, we introduce a weighted ℓ1/ℓ2mixed norms minimization SR method and formulate a unified framework together with stepwise scheme. Experiments on commonly used face database demonstrate that our framework achieves state-of-the-art results. Zhongyuan Wang 0001, Shizheng Wang, Ruimin Hu |
ICME | 1 |
| 2013 | Support-driven sparse coding for face hallucinationabstractBy incorporating the prior of positions, position patch based face hallucination methods can produce high-quality results and save computation time. Given a low-resolution face image, the key issue of these methods is how to encode the input low-resolution patch. However, due to stability and accuracy issues, the coding approaches proposed so far are not satisfactory. In this paper, we present a novel sparse coding method via exploiting the support information on the coding coefficients. In particular, the support information is characterized by the locality of the image patch manifold, which has been shown to be critical in data representation and analysis. According to the distances between the input patch and bases in the dictionary, we first assign different weights to the coding coefficients and then obtain the coding coefficients by solving a weighted sparse problem. Our proposed method exploits the non-linear manifold structure of patch samples and the sparse property of the redundant data, leading to stable and accurate representation. Experiments on commonly used databases demonstrate that our method outperforms state of the art. Junjun Jiang, Ruimin Hu, Zhongyuan Wang 0001, Zixiang Xiong, Zhen Han 0002 |
ISCAS | 3 |
| 2013 | Color image guided locality regularized representation for Kinect depth holes fillingabstractThe emergence of Microsoft Kinect has attracted the attention not only from consumers but also from researchers in the field of computer vision. It facilitates the possibility to capture the depth map of the scene in real time and with low cost. Nonetheless, due to the limitations of structured light measurements used by Kinect, the captured depth map suffers random depth missing in the occlusion or smooth regions, which affects the accuracy of many Kinect based applications. In order to fill in the holes existing in Kinect depth map, some approaches that adopted color image guided in-painting or joint bilateral filter have been proposed to represent the missing depth pixel by available depth pixels. However, they are not able to obtain the optimal weights, thus the obtained missing depth values are not best. In this paper, we propose a color image guided locality regularized representation (CGLRR) to reconstruct the missing depth pixels by comprehensively determining the optimal weights of the available depth pixels from collocated patches in color image. Experimental results demonstrate that the proposed algorithm can better fill in the holes of depth map both in smooth and edge region than previous works. Ruimin Hu, Zhongyuan Wang 0001, Mang Duan |
VCIP | 3 |
| 2013 | Surveillance video synopsis in the compressed domain for fast video browsing
Shizheng Wang, Zhongyuan Wang 0001, Ruimin Hu |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Background Subtraction With Video CodingabstractThe classic Gaussian mixture model is based on the statistical information of every pixel; it is not robust to light changes. Before analysing every pixel in videos, it must be decoded to raw videos. In this letter, the method combining video coding and the Gaussian mixture model together is proposed. We use intra mode and motion vectors to find the foreground macroblock, then add one overhead flag in the compressed video to indicate it. In the decoder, we just decode possible foreground areas and detect moving objects in these areas. In our experiments, we test this method on two datasets, both of them with unique, dynamic, illumination conditions. Results show that the proposed method is effective to detect moving objects and easily assemble to current automated video surveillance systems. Zhenkun Huang, Ruimin Hu, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2012 | Improvements of dynamic texture synthesis for video coding
Ruimin Hu, Zhongyuan Wang 0001, Zhen Han 0002 |
ICPR | 3 |
| 2012 | Intracoding and Refresh With Compression-Oriented Video Epitomic PriorsabstractIn video compression, intracoding plays an important role in terms of coding efficiency and error resilience and has been an attractive research topic since the standardization of H.264/AVC. In this paper, we propose a high-performance intracoding scheme with the help of epitomic priors. Different from intracoding in H.264/AVC and other video standards, we construct image epitomes as coding priors and use them to generate predictions of intrablocks at the encoder. In addition, we losslessly code and transmit the image epitomes to the decoder. We perform compression-oriented video epitomic analysis and search for the best epitomic priors by using the expectation maximization algorithm. The resulting image epitomes for a video sequence can be viewed as the base layer in spatially scalable video coding. Experiments show that our proposed intracoding scheme improves the state of the art by an average of 0.53 dB in PSNR. Simulations under a packet loss environment also demonstrate that intrarefresh with epitomic priors outperforms random intrarefresh by up to 2 dB, leading to better subjective quality. Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | A Novel Frame Error Concealment Algorithm Based on Dynamic Texture SynthesisabstractDynamic textures are sequences of frames of moving scenes that exhibit certain stationary properties in time, which have great significance in applications such as video production, virtual simulation and virtual walkthroughs. This paper presents an algorithm for dynamic texture extrapolation using for H.264 decoding system. The synthesized frames can be used by the decoder for whole frames loss error concealment. The simulation results show that the proposed whole frames loss error concealment algorithm achieves significant improvement over the motion vector extrapolation method. Ruimin Hu, Dan Mao, Zhongyuan Wang 0001 |
DCC | 4 |
| 2010 | Spatially Scalable Video Coding Based on Hybrid Epitomic ResizingabstractScalable video coding (SVC) is considered as a potentially promising solution to enable the adaptability of video to heterogonous networks and various devices. In spatially scalable video encoder, how to resize the captured high-resolution video to get low-resolution video has great effect on the quality of experience (QoE) in the clients receiving low-resolution video. In this paper, we propose a new resizing algorithm called hybrid epitomic resizing (HER), which can make the resized image preserve the same ‘physical’ resolution with original image by the way of utilizing texture similarity inside image and highlight regions of interest while avoiding potential artifacts. For hybrid epitomic resizing, we also design two new inter-layer prediction methods to eliminate the redundancy between adjacent spatial layers instead of conventional inter-layer prediction. Experimental results show that HER can get resized images with perceptually much better quality and the performance of new inter-layer prediction are comparable to that of conventional inter-layer prediction in H.264 SVC. Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001 |
DCC | 3 |
| 2010 | A face super-resolution approach using shape semantic mode regularizationabstractIn actual imaging environment, a variety of factors have an impact on the quality of images, which leads to pixel distortion and aliasing. The traditional face super-resolution algorithm only uses the difference of image pixel values as similarity criterion, which degrades similarity and identification of reconstructed facial images. Image semantic information with human understanding, especially structural information, is robust to the degraded pixel values. In this paper, we propose a face super-resolution approach using shape semantic model. This method describes the facial shape as a series of fiducial points on facial image. And shape semantic information of input image is obtained manually. Then a shape semantic regularization is added to the original objective function. The steepest descent method is used to obtain the unified coefficient. Experimental results demonstrate that the proposed method outperforms the traditional schemes significantly both in subjective and objective quality. Chengdong Lan, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001 |
ICIP | 4 |
| 2010 | Temporal color Just Noticeable Distortion model and its application for video codingabstractJust Noticeable Distortion (JND), which is utilized to reduce the bit rate without introducing noticeable visual distortion, plays an important role in perceptual image and video processing. For the temporal color JND, it takes into account not only the spatial and luminance HVS properties, but also the temporal and chroma HVS properties. In this paper, we first develop a spatio-temporal model estimating JND for color video by completely incorporating the color CSF, the frequency property of DCT coefficient, the contrast masking effect and the motion property. Then we incorporate the JND model into video encoding system via the residue filtering process. This method can work with any prevalent video coding standards. To demonstrate the effectiveness of the JND model, we has implemented it into the H.264/AVC reference software JM12.4 and the experimental results show that the bit rate can be reduced by average 18.20%, which reflects that our JND model is able to exploit the HVS bounds more aggressively without introducing noticeable visual distortions. Ruimin Hu, Zhongyuan Wang 0001 |
ICME | 4 |
| 2010 | Video coding using dynamic texture synthesisabstractIn video coding system, because of the delay between current picture and the reference pictures, the temporal prediction is not good for the video sequences with nonlinear motion and global illumination change between frames. Dynamic textures are sequences of frames of moving scenes that exhibit certain stationary properties in time, which have great significance in applications such as video production and virtual simulation. This paper presents a new algorithm for dynamic texture extrapolation using for H.264 encoding and decoding system. The synthesized frames are used by the encoder for virtual reference frames choice in inter prediction and decoder for whole frames loss error concealment. The simulation results show that the proposed virtual reference frames choice algorithm improves encoding efficiency compared with the H.264 standard, and the proposed whole frames loss error concealment algorithm achieves significant improvement over the motion vector extrapolation method. Ruimin Hu, Dan Mao, Zhongyuan Wang 0001 |
ICME | 5 |
| 2010 | Intra coding and refresh based on video epitomic analysisabstractIn video coding, intra coding plays an important role in both coding efficiency and error resilience. In this paper, a new intra coding method based on video epitomic analysis is proposed. Image epitome suitable for coding is extracted through video epitomic analysis, and is used to generate the prediction for each block. The mapping between the original image and image epitome as well as image epitome should be coded and transmitted to the decoder side. Due to the independence of adjacent reconstructed blocks, the proposed intra coding method can also be applied to improve error resilience of intra refresh. Experimental results show that, the proposed intra coding method can significantly improve the performance of intra coding by 0.7dB in terms of PSNR. The simulation results under packet loss environment show that the proposed intra refresh can out-perform random intra refresh (RIR) by up to 1dB, and better subjective quality can also be obtained. Qijun Wang, Ruimin Hu, Zhongyuan Wang 0001, Bo Hang |
ICME | 3 |
| 2010 | Inter prediction based on spatio-temporal adaptive localized learning modelabstractInter prediction based on block matching motion estimation is important for video coding. But this method suffers from the additional overhead in data rate representing the motion information that needs to be transmitted to the decoder. To solve this problem, we present an improved implicit motion information inter prediction algorithm for P slice in H.264/AVC based on the spatio-temporal adaptive localized learning (STALL) model. According to 4 × 4 block transform structure in H.264/AVC, we first adaptively choose nine spatial neighbors and nine temporal neighbors, and a localized 3D casual cube is designed as training window. By using these information, the model parameters could be adaptively computed based on the Least Square Prediction (LSP) method. Finally, we add a new inter prediction mode into H.264/AVC standard for P slice. The experimental results show that our algorithm improves encoding efficiency compared with H.264/AVC standard, with relatively increases in complexity. Ruimin Hu, Zhongyuan Wang 0001 |
PCS | 3 |
| 2009 | Improved Object Tracking Algorithm Based on New HSV Color Probability Model
Gang Tian, Ruimin Hu, Zhongyuan Wang 0001, Youming Fu |
ISNN (2) | 3 |