VLDB 2026 Research / reviewers in the wild / expert
Yang Zhao 0002
dblp:50/2082-2
· DBLP profile ↗
81ranked-venue papers
16as first author
53since 2021 · last 2026
0000-0002-4032-8049ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 11 first-author · 39 since 2021Artificial intelligence and machine learning · 32 · 4 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Guided Scene Text Image Super-ResolutionabstractScene text image super-resolution aims to enhance text legibility by recovering high-resolution text images from low-resolution inputs. However, maintaining fine details such as text strokes, edges, and textual accuracy remains challenging, particularly in low-light environments and high-speed motion scenarios, where degradation is more severe. Event cameras, with their high temporal resolution and ability to capture intensity changes, offer a promising solution for restoring lost fine details and mitigating degradation in these challenging conditions. In this paper, we propose EvTSR, the first framework that integrates Event data for scene Text image Super-Resolution. The core of EvTSR is the dual-stream frequency boost (DSFB) mechanism, which separates image features into high- and low-frequency components. High-frequency details like edges and strokes are enhanced using event data via the event-guided high-frequency (EGH) mechanism, while low-frequency components, responsible for global structure, are refined using the Text-Guided Low-frequency (TGL) mechanism with a pre-trained text recognizer, ensuring textual coherence. To further improve cross-modal integration, we introduce the cross-modal fusion (CMF) mechanism, which effectively aligns event and image features, enabling robust information fusion. Extensive experiments demonstrate that EvTSR achieves superior performance over existing methods. Zihan Qi, Zeyu Xiao 0002, Haoyi Zhao, Yang Zhao 0002, Feng Xue 0002, Wei Jia 0001 |
AAAI | 4 |
| 2026 | LSAP-PV: High-Fidelity Palm Vein Image Synthesis via Layered Spectral Absorption Projection-Guided Diffusion ModelabstractPalm vein recognition has emerged as a promising biometric technology, yet its development remains constrained by the scarcity of large-scale publicly available datasets. Several methods of palm vein image generation have been proposed to address this issue. These methods usually focus on the anatomical realism of palm vein patterns, but overlook the biophysical correlation between identities and vein patterns, particularly in simulating identity-specific vein contrast. To tackle this limitation, we propose a novel biophysics-driven synthesis method. Our method constructs a 3D palm vascular tree via established modeling method. Then, a projection model is proposed to map the 3D tree into 2D space to derive palm vein patterns. The projection model is based on skin spectral absorption and simulates the natural attenuation of light passing through the skin using a layer integration method. For different identities, we sample different skin parameters, resulting in varying degrees of attenuation. This method effectively simulates the variation in vein contrast across different identities. Furthermore, we introduce a conditional diffusion model that uses the projected patterns as identity conditions to generate palm vein images. To the best of our knowledge, this is the first palm vein generation method based on the diffusion model. Experimental results demonstrate that our method not only outperforms existing methods, but also enables a recognition model trained on our synthetic data to achieve superior performance compared to a model trained on real-world data at a scale of 2,000 IDs under an open-set protocol with a TAR@FAR=1:1 of 1e-4. Sheng Shang, Chenglong Zhao, Jianlong Jin, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001 |
AAAI | 7 |
| 2026 | Semantic attention and progressive training for image authenticity assessment and tampered region highlighting
Bo Wang 0072, Qi Si, Yang Zhao 0002, Zhong-Qiu Zhao, Zhao Zhang 0001 |
Inf. Sci. | 4 |
| 2026 | Multi-scale spatial diffusion under frequency information-guidance For low-light image enhancement
Jinhan Guan, Bo Wang 0072, Zhao Zhang 0001, Yang Zhao 0002, Haijun Zhang 0002, Yun Yang 0003, Xianming Ye, Meng Wang 0001 |
Neural Networks | 4 |
| 2026 | Parameter-efficient transfer for CLIP-based text-to-person retrieval
Hai Min, Guanghui Zhan, Chunxiao Fan 0002, Yang Zhao 0002, Wei Jia 0001 |
Signal Process. Image Commun. | 4 |
| 2026 | BITMNet: A Degradation-Aware Mixture-of-Experts Framework for Blind Inverse Tone Mapping
Wenyou Zhang, Yang Zhao 0002, Fangxing Zhang, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Event-Based Dynamic Turbulence MitigationabstractAtmospheric turbulence induces coupled spatio-temporal distortions, including blur, geometric deformation, and temporal jitter, which severely degrade image quality. We propose EvTurM, a practical framework leveraging event camera data for dynamic turbulence mitigation with precise motion cues and stable temporal modeling. Leveraging the high temporal resolution and dynamic range of events, EvTurM achieves robust restoration under diverse turbulence conditions. EvTurM comprises two key modules: (1) the event-aware modality enhancement module, which uses event-derived motion to enrich RGB features and recover structural details, and (2) the bidirectional modality calibration module, which jointly aligns RGB and event features in forward and backward propagation to reduce misalignment and enhance temporal consistency. Extensive experiments show EvTurM consistently surpasses existing methods and achieves superior performance. Haoyi Zhao, Zeyu Xiao 0002, Zihan Qi, Yang Zhao 0002, Wei Jia 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Joint Resolution and Rendering Artifacts Removal for Cloud Gaming ImageabstractWith the rapid development of the cloud gaming industry, low-quality rendering and rescaling strategies are commonly employed to mitigate the high costs of cloud-based computation and bandwidth. As a result, client-side images often contain artifacts such as mixed aliasing and resolution distortions, which cannot be effectively handled by current super-resolution models. In response, this paper proposes a cloud gaming image enhancement (CGIE) model to tackle both rendering and resolution degradations. Initially, this paper builds a large dataset by rendering and rescaling paired data with different qualities from collected 3D game scenes. Subsequently, a lightweight dual-branch enhancement network is designed, which consists of a high-frequency branch primarily focused on detail enhancement and a sampling-space branch aimed at enlarging the receptive field and perceiving multi-scale aliasing artifacts. Experimental results demonstrate the superior anti-aliasing and image enhancement performance of the proposed method across various real-world cloud games and even mobile games. The dataset and codes are available at https://github.com/YCheno/CGIE/. Yang Zhao 0002, Yuan Chen 0012, Lin Li 0053, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Learning Dual Modality Interactions for Event-Based Motion DeblurringabstractEvent cameras hold great potential for motion deblurring because they capture motion information with microsecond precision, offering robustness to motion blur. However, the limited interaction between RGB frames and event streams presents a significant challenge, preventing the full utilization of the event cameras' unique advantages. To address this, we proposeDual frame-eventInteraction and introduce a multi-scaleNetwork structure, DuInt-Net. DuInt-Net aims to tackle two key challenges: (1) enhancing the representational and interaction capabilities between RGB frames and event streams, and (2) adaptively selecting richer visual features for improved motion deblurring. We introduce an event-frame joint interaction module that consists of three branches: a base branch, a global awareness attention branch, and a local enhancement attention branch. The base branch processes essential pixel-level features that retain the original structural information. The global branch integrates event data to improve large-scale motion understanding, while the local branch uses large-kernel convolutions to refine fine-grained details in RGB frames. For superior reconstruction performance, we also propose the event-guided multi-scale fusion attention module, which effectively combines local visual information and global frame-event relationships. Extensive experiments demonstrate that DuInt-Net achieves superior performance, both quantitatively and qualitatively, showcasing its superior motion deblurring capabilities. Zeyu Xiao 0002, Zhuoyuan Li 0001, Yang Zhao 0002, Yu Liu 0023, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | PVTree: Realistic and Controllable Palm Vein Generation for Recognition TasksabstractPalm vein recognition is an emerging biometric technology that offers enhanced security and privacy. However, acquiring sufficient palm vein data for training deep learning-based recognition models is challenging due to the high costs of data collection and privacy protection constraints. This has led to a growing interest in generating pseudo-palm vein data using generative models. Existing methods, however, often produce unrealistic palm vein patterns or struggle with controlling identity and style attributes. To address these issues, we propose a novel palm vein generation framework named PVTree. First, the palm vein identity is defined by a complex and authentic 3D palm vascular tree, created using an improved Constrained Constructive Optimization (CCO) algorithm. Second, palm vein patterns of the same identity are generated by projecting the same 3D vascular tree into 2D images from different views and converting them into realistic images using a generative model. As a result, PVTree satisfies the need for both identity consistency and intra-class diversity. Extensive experiments conducted on several publicly available datasets demonstrate that our proposed palm vein generation method surpasses existing methods and achieves a higher TAR@FAR=1e-4 under the 1:1 Open-set protocol. To the best of our knowledge, this is the first time that the performance of a recognition model trained on synthetic palm vein data exceeds that of the recognition model trained on real data, which indicates that palm vein image generation research has a promising future. Sheng Shang, Chenglong Zhao, Jianlong Jin, Rizen Guo, Shouhong Ding, Yunsheng Wu, Yang Zhao 0002, Wei Jia 0001 |
AAAI | 9 |
| 2025 | Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video ParsingabstractThe Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features for intra- and cross-modal temporal interactions. However, each segment may contain multiple events, resulting in semantically mixed holistic features that can lead to semantic interference during intra- or cross-modal interactions: the event semantics of one segment may incorporate semantics of unrelated events from other segments. To address this issue, our method begins with a Class-Aware Feature Decoupling (CAFD) module, which explicitly decouples the semantically mixed features into distinct class-wise features, including multiple event-specific features and a dedicated background feature. The decoupled class-wise features enable our model to selectively aggregate useful semantics for each segment from clearly matched classes contained in other segments, preventing semantic interference from irrelevant classes. Specifically, we further design a Fine-Grained Semantic Enhancement module for encoding intra- and cross-modal relations. It comprises a Segment-wise Event Co-occurrence Modeling (SECM) block and a Local-Global Semantic Fusion (LGSF) block. The SECM exploits inter-class dependencies of concurrent events within the same timestamp with the aid of a novel event co-occurrence loss. The LGSF further enhances the event semantics of each segment by incorporating relevant semantics from more informative global video features. Extensive experiments validate the effectiveness of the proposed modules and loss functions, resulting in a new state-of-the-art parsing performance. Jinxing Zhou, Yang Zhao 0002, Dan Guo 0001 |
AAAI | 3 |
| 2025 | Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion ModelsabstractPalmprint recognition is significantly limited by the lack of large-scale publicly available datasets. Previous methods have adopted Bézier curves to simulate the palm creases, which then serve as input for conditional GANs to generate realistic palmprints. However, without employing real data fine-tuning, the performance of the recognition model trained on these synthetic datasets would drastically decline, indicating a large gap between generated and real palmprints. This is primarily due to the utilization of an inaccurate palm crease representation and challenges in balancing intra-class variation with identity consistency. To address this, we introduce a polynomial-based palm crease representation that provides a new palm crease generation mechanism more closely aligned with the real distribution. We also propose the palm creases conditioned diffusion model with a novel intra-class variation control method. By applying our proposed K-step noise-sharing sampling, we are able to synthesize palmprint datasets with large intra-class variation and high identity consistency. Experimental results show that, for the first time, recognition models trained solely on our synthetic datasets, without any fine-tuning, outperform those trained on real datasets. Furthermore, our approach achieves superior recognition performance as the number of generated identities increases. Jianlong Jin, Chenglong Zhao, Sheng Shang, Jianqing Xu, Shaoming Wang, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001, Yunsheng Wu |
CVPR | 8 |
| 2025 | PsSR: Hybrid Path Selection Mechanism for Efficient Image Super-ResolutionabstractIn practical applications, large images are processed in patches, many of which are simple and smooth, making them suitable for lighter network processing. This paper proposes a hybrid path selection mechanism that enables the dynamic routing of patches through the network, achieving a balance between accuracy and computational efficiency. The proposed method develops a dual-branch architecture, in which the two branches keep the same structural design, but differ in the number of channels. An efficient classification module is integrated for directing patches to the appropriate branches within the model. In addition, within each branch, we introduce an early exit strategy to further reduce the calculation cost. Subsequently, we enhance the restoration capabilities of the lightweight branch through channel and spatial feature distillation. Experimental results demonstrate that our Hybrid Path Selection based Super-Resolution (PsSR) mechanism effectively reduces computational costs while preserving the integrity of the backbone performance. Zhong-Qiu Zhao, Yang Zhao 0002 |
ICASSP | 3 |
| 2025 | High-Fidelity Stereoscopic Image Rain Removal with Texture Integrity and Disparity ConsistencyabstractThis paper tackles the challenge of stereoscopic image rain removal by focusing on enhancing texture integrity and disparity consistency. Existing stereoscopic rain removal techniques often fall short due to 1) disruptions in texture coherence caused by complex rain streaks, and 2) inaccuracies in disparity estimation from inadequate feature fusion. To overcome these limitations, we introduce the StereoIRR method, which incorporates: 1) a Long-range and Cross-view Interaction (LCI) framework that preserves texture integrity by mitigating rain’s adverse effects on stereoscopic features, and 2) a Dual-view Mutual Attention mechanism that ensures disparity consistency by generating precise mutual attention maps for cross-view feature fusion. Our approach not only maintains the integrity of stereoscopic textures but also significantly reduces errors in disparity estimation. Extensive experiments demonstrate that StereoIRR consistently outperforms state-of-the-art monocular and stereoscopic methods on multiple benchmark datasets. Yanyan Wei, Zhao Zhang 0001, Zhong-Qiu Zhao, Yang Zhao 0002, Richang Hong, Yi Yang 0001, Meng Wang 0001 |
ICASSP | 4 |
| 2025 | Unified Adversarial Augmentation for Improving Palmprint Recognition
Jianlong Jin, Chenglong Zhao, Sheng Shang, Yang Zhao 0002, Shouhong Ding, Wei Jia 0001, Yunsheng Wu |
ICCV | 5 |
| 2025 | Contrastive Lie Algebra Learning for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself via contrastive learning has demonstrated great progress towards learning discriminative representation. Yet forcingly dividing highly similar categories at the representation level may over-guide the learned feature space, leading to overfitting in the ultra-FGVC tasks. To this end, this paper introduces CLA-Net, a novel contrastive Lie algebra learning framework to address this fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images at a second-order representation level. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLA-Net demonstrates strong performance on eight publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at: https://github.com/zichengpan/CLA-NET. Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0002, Qin Zhang 0011, Yongsheng Gao 0001 |
ACM Multimedia | 3 |
| 2025 | Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit
Yang Zhao 0002, Xueshang Feng |
ACM Multimedia | 1 |
| 2025 | Multi-Layer Gaussian Splatting for Single-Image Feed-Forward Spatial Scene ReconstructionabstractRecently, 3D Gaussian Splatting (3DGS) has achieved remarkable results in 3D reconstruction and view synthesis tasks. However, single-view feed-forward 3DGS still faces significant challenges. Current state-of-the-art (SOTA) single-view 3DGS methods typically employ a small number of layers (1-2 layers) with Gaussian Splatting (GS) representations at the same resolution as the input image to address the irregularity of GS data. However, such shallow and uniform GS primitive distributions is difficult to represent occluded regions and important spatial details. Inspired by multi-plane images, this paper proposes a Multi-Layer Gaussian Splatting (MLGS) representation, which consists of shallow base GS layers for visible content and multiple occlusion GS layers dedicated to reconstructing occluded regions. The proposed MLGS representation explicitly decouples the learning processes of visible and occluded content while enhancing occlusion prediction through the following components. First, spatial stratification of GS is achieved by estimating the depth distribution range of GS primitives across different layers, forcing GS to learn spatial content reconstruction at different depths. Second, a mask-guided mechanism is proposed to effectively isolate occlusion regions and guide inpainting using spatially context-aware features. Finally, a gated convolution block is designed to dynamically modulate feature fusion to enhance reconstruction fidelity. With separate loss supervision for base and occlusion layers, MLGS enables geometrically plausible scene completion. Experiments on RealEstate10K, KITTI, and NYUv2 datasets demonstrate that the proposed method achieves SOTA performance for single-image spatial scene reconstruction. Shanding Diao, Yang Zhao 0002, Yuan Chen 0012, Zhao Zhang 0001, Wei Jia 0001, Ronggang Wang |
ACM Multimedia | 2 |
| 2025 | Hybrid Scalable Video Coding with Neural Compression and Enhancement for Streaming Media
Yuyao Ye, Yang Zhao 0002, Mengping Gao, Hongbin Cao, Ronggang Wang |
MMM (2) | 3 |
| 2025 | Individual/joint deblurring and low-light image enhancement in one go via unsupervised deblurring paradigm
Suiyi Zhao, Zhao Zhang 0001, Yanyan Wei, Jicong Fan 0001, Yang Zhao 0002, Shuicheng Yan, Meng Wang 0001 |
Sci. China Inf. Sci. | 5 |
| 2025 | When low-light meets flares: Towards Synchronous Flare Removal and Brightness Enhancement
Jiahuan Ren, Zhao Zhang 0001, Suiyi Zhao, Jicong Fan 0001, Zhong-Qiu Zhao, Yang Zhao 0002, Richang Hong, Meng Wang 0001 |
Neural Networks | 6 |
| 2025 | Cut-and-Paste: Subject-driven video editing with attention control
Zhichao Zuo, Zhao Zhang 0001, Yan Luo 0004, Yang Zhao 0002, Haijun Zhang 0002, Yi Yang 0001, Meng Wang 0001 |
Neural Networks | 4 |
| 2025 | Local Texture Pattern Estimation for Image Detail Super-ResolutionabstractIn the image super-resolution (SR) field, recovering missing high-frequency textures has always been an important goal. However, deep SR networks based on pixel-level constraints tend to focus on stable edge details and cannot effectively restore random high-frequency textures. It was not until the emergence of the generative adversarial network (GAN) that GAN-based SR models achieved realistic texture restoration and quickly became the mainstream method for texture SR. However, GAN-based SR models still have some drawbacks, such as relying on a large number of parameters and generating fake textures that are inconsistent with ground truth. Inspired by traditional texture analysis research, this paper proposes a novel SR network based on local texture pattern estimation (LTPE), which can restore fine high-frequency texture details without GAN. A differentiable local texture operator is first designed to extract local texture structures, and a texture enhancement branch is used to predict the high-resolution local texture distribution based on the LTPE. Then, the predicted high-resolution texture structure map can be used as a reference for the texture fusion SR branch to obtain high-quality texture reconstruction. Finally, $L_{1}$L1 loss and Gram loss are simultaneously used to optimize the network. Experimental results demonstrate that the proposed method can effectively recover high-frequency texture without using GAN structures. In addition, the restored high-frequency details are constrained by local texture distribution, thereby reducing significant errors in texture generation. Yang Zhao 0002, Yuan Chen 0012, Nannan Li 0001, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Stereo Vision Conversion from Planar Videos Based on Temporal Multiplane ImagesabstractWith the rapid development of 3D movie and light-field displays, there is a growing demand for stereo videos. However, generating high-quality stereo videos from planar videos remains a challenging task. Traditional depth-image-based rendering techniques struggle to effectively handle the problem of occlusion exposure, which occurs when the occluded contents become visible in other views. Recently, the single-view multiplane images (MPI) representation has shown promising performance for planar video stereoscopy. However, the MPI still lacks real details that are occluded in the current frame, resulting in blurry artifacts in occlusion exposure regions. In fact, planar videos can leverage complementary information from adjacent frames to predict a more complete scene representation for the current frame. Therefore, this paper extends the MPI from still frames to the temporal domain, introducing the temporal MPI (TMPI). By extracting complementary information from adjacent frames based on optical flow guidance, obscured regions in the current frame can be effectively repaired. Additionally, a new module called masked optical flow warping (MOFW) is introduced to improve the propagation of pixels along optical flow trajectories. Experimental results demonstrate that the proposed method can generate high-quality stereoscopic or light-field videos from a single view and reproduce better occluded details than other state-of-the-art (SOTA) methods. https://github.com/Dio3ding/TMPI Shanding Diao, Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
AAAI | 3 |
| 2024 | PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint GenerationabstractThe lack of large-scale data seriously hinders the development of palmprint recognition. Recent approaches address this issue by generating large-scale realistic pseudo palmprints from Bézier curves. However, the significant difference between Bézier curves and real palmprints limits their effectiveness. In this paper, we divide the Bézier-Real difference into creases and texture differences, thus reducing the generation difficulty. We introduce a new palm crease energy (PCE) domain as a bridge from Bézier curves to real palmprints and propose a two-stage generation model. The first stage generates PCE images (realistic creases) from Bézier curves, and the second stage outputs realistic palmprints (realistic texture) with PCE images as input. In addition, we also design a lightweight plug-and-play line feature enhancement block to facilitate domain transfer and improve recognition performance. Extensive experimental results demonstrate that the proposed method surpasses state-of-the-art methods. Under extremely few data settings like 40 IDs (only 2.5% of the total training set), our model achieves a 29% improvement over RPG-Palm and outperforms ArcFace with 100% training set by more than 6% in terms of TAR@FAR=1e-6. Jianlong Jin, Chenglong Zhao, Shouhong Ding, Yang Zhao 0002, Wei Jia 0001 |
AAAI | 8 |
| 2024 | Deep Video Inverse Tone Mapping Based on Temporal CluesabstractInverse tone mapping (ITM) aims to reconstruct high dynamic range (HDR) radiance from low dynamic range (LDR) content. Although many deep image ITM methods can generate impressive results, the field of video ITM is still to be explored. Processing video sequences by image ITM methods may cause temporal inconsistency. Besides, they aren't able to exploit the potentially useful information in the temporal domain. In this paper, we analyze the process of video filming, and then propose a Global Sample and Local Propagate strategy to better find and utilize temporal clues. To better realize the proposed strategy, we design a two-stage pipeline which includes modules named Incremental Clue Aggregation Module and Feature and Clue Propagation Module. They can align andfuseframes effectively under the condition of brightness changes and propagate features and temporal clues to all frames efficiently. Our temporal clues based video ITM method can recover realistic and temporal consistent results with high fidelity in over-exposed regions. Qualitative and quantitative experiments on public datasets show that the proposed method has significant advantages over existing methods. The code is available at https://github.com/ye3why/VITM-TC/. Yuyao Ye, Ning Zhang 0023, Yang Zhao 0002, Hongbin Cao, Ronggang Wang |
CVPR | 3 |
| 2024 | Blind Video Bit-Depth ExpansionabstractWith the rapid development of high-bit-depth display devices, bit-depth expansion (BDE) algorithms that extend low-bit-depth images to high-bit-depth images have received increasing attention. Due to the sensitivity of bit-depth distortions to tiny numerical changes in the least significant bits, the nuanced degradation differences in the training process may lead to varying degradation data distributions, causing the trained models to overfit specific types of degradations. This paper focuses on the problem of blind video BDE, proposing a degradation prediction and embedding framework, and designing a video BDE network based on a recurrent structure and dual-frame alignment fusion. Experimental results demonstrate that the proposed model can outperform some state-of-the-art (SOTA) models in terms of banding artifact removal and color correction, avoiding overfitting to specific degradations and obtaining better generalization ability across multiple datasets. https://github.com/duanpanjun/BVBDE Panjun Duan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
ACM Multimedia | 2 |
| 2024 | Integrating pseudo labeling with contrastive clustering for transformer-based semi-supervised action recognition
Nannan Li 0001, Kan Huang, Qingtian Wu, Yang Zhao 0002 |
Appl. Intell. | 4 |
| 2024 | Regional Traditional Painting Generation Based on Controllable Disentanglement ModelabstractAutomatic generation of painting images is an interesting and difficult task, especially for regional traditional paintings with unique cultural styles while lacking large-scale training sets. In this paper, a hierarchical painting generation method is proposed, which can disentangle the generation of content and style. By mimicking the human painting process, the proposed method introduces multiple content blocks first and gradually generates image contents. In each block, a spatial self-modulation module is proposed to inject local details while preserving the global layout. After the preliminary generation of contents, a series of style blocks are presented to gradually adjust the artistic style. In the style block, an edge-oriented style-modulation module is proposed, which focuses on the lines and edges. In addition, edge adversarial training is used to further improve the quality of generated lines. To train and evaluate the proposed method, we construct datasets for five types of Chinese folk paintings. Experimental results demonstrate that the proposed method can generate high-quality and diverse painting images. More importantly, it can disentangle content and style sufficiently, so that the generation of specific contents or styles can be controlled freely. The datasets and source codes is available at https://github.com/Ritsu-mio/HPGN. Yang Zhao 0002, Huaen Li, Zhao Zhang 0001, Yuan Chen 0012, Qing Liu 0022 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Toward Individual Tone Preference in Underwater Image EnhancementabstractUnderwater images often suffer from severe color distortion due to the challenging imaging environment. Underwater image enhancement (UIE) techniques have been developed to recover clear images, laying the foundation for various underwater research. However, existing UIE methods tend to produce fixed results without considering individual preferences for different color tones. And there is no dataset with ground truth (GT) in different tones. Therefore, we came up with the possibility of using the currently popular multimodal methods to control the color tone of enhanced images. This article proposes a method for generating underwater enhanced images with cold, warm, and normal tones using multimodal information supervision (MM-UIE). First, we leverage the relationship between text prompts and images to supervise the generation of cold or warm images. In addition, we introduce a 6-D color operator, which not only enhances the tone control of underwater images but also serves as a bridge between different tone images. Finally, we also found that multimodal supervision methods can not only control the color tone of underwater images but also improve the quality of underwater image generation. Experimental results demonstrate the superior performance of our method compared to state-of-the-art (SOTA) techniques. Our codes will be publicly available athttps://github.com/perseveranceLX/MM-UIE. Yang Zhao 0002, Kaichen Chi, Zhao Zhang 0001, Wei Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Progressive Stereo Image Dehazing Network via Cross-View Region InteractionabstractStereo image dehazing aims to restore haze-free images by leveraging the complementary information contained in binocular images. Current methods primarily focus on designing image-level modules and pipelines to utilize complementary information between the left and right-view images. However, these image-level cross-view interactions overlook regional differences in haze concentration and stereo image disparity maps. Consequently, we propose a Progressive Stereo Image Dehazing Network via Cross-view Region Interaction, termed PSIDNet, which fully considers the internal characteristics and external manifestation of haze and disparity, and explicitly addresses the stereo image dehazing task by a regional-aware interactive mechanism. Specifically, we divide hazy images into regions and independently interact with left and right-view information at region levels, meaning weights are not shared across regional patches. This approach allows us to treat different regions with different priorities, i.e., concentrate on regional patches with heavier haze concentration and larger disparities, hence enabling more accurate restoration of hazy images. Furthermore, we introduce an effective cross-view region interactive block that extracts information based on the channel dimension of dual views and later adopts matrix multiplication to generate mutual attention maps based on the fused features. Extensive experiments on synthetic and real-scenario datasets demonstrate the efficacy of our method, compared to other related monocular and stereo image dehazing and restoration methods. Our code will be released publicly at https://github.com/Alvin2112/PSIDNet. Junhu Wang, Yanyan Wei, Zhao Zhang 0001, Jicong Fan 0001, Yang Zhao 0002, Yi Yang 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Revisiting the Stack-Based Inverse Tone MappingabstractCurrent stack-based inverse tone mapping (ITM) methods can recover high dynamic range (HDR) radiance by predicting a set of multi-exposure images from a single low dynamic range image. However, there are still some limitations. On the one hand, these methods estimate a fixed number of images (e.g., three exposure-up and three exposure-down), which may introduce unnecessary computational cost or reconstruct incorrect results. On the other hand, they neglect the connections between the up-exposure and down-exposure models and thus fail to fully excavate effective features. In this paper, we revisit the stack-based ITM approaches and propose a novel method to reconstruct HDR radiance from a single image, which only needs to estimate two exposure images. At first, we design the exposure adaptive block that can adaptively adjust the exposure based on the luminance distribution of the input image. Secondly, we devise the cross-model attention block to connect the exposure adjustment models. Thirdly, we propose an end-to-end ITM pipeline by incorporating the multi-exposure fusion model. Furthermore, we propose and open a multi-exposure dataset that indicates the optimal exposure-up/down levels. Experimental results show that the proposed method outperforms some state-of-the-art methods. Ning Zhang 0023, Yuyao Ye, Yang Zhao 0002, Ronggang Wang |
CVPR | 3 |
| 2023 | RPG-Palm: Realistic Pseudo-data Generation for Palmprint RecognitionabstractPalmprint recently shows great potential in recognition applications as it is a privacy-friendly and stable biometric. However, the lack of large-scale public palmprint datasets limits further research and development of palmprint recognition. In this paper, we propose a novel realistic pseudo-palmprint generation (RPG) model to synthesize palmprints with massive identities. We first introduce a conditional modulation generator to improve the intra-class diversity. Then an identity-aware loss is proposed to ensure identity consistency against unpaired training. We further improve the Bézier palm creases generation strategy to guarantee identity independence. Extensive experimental results demonstrate that synthetic pretraining significantly boosts the recognition model performance. For example, our model improves the state-of-the-art BézierPalm by more than 5% and 14% in terms of TAR@FAR=1e-6 under the 1 : 1 and 1 : 3 Open-set protocol. When accessing only 10% of the real training data, our method still outperforms ArcFace with 100% real training data, indicating that we are closer to real-data-free palmprint recognition. Jianlong Jin, Huaen Li, Kai Zhao 0012, Shouhong Ding, Yang Zhao 0002, Wei Jia 0001 |
ICCV | 9 |
| 2023 | Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based ApproachabstractTraditional image codecs prioritize signal fidelity and human perception, often neglecting machine vision tasks. Deep learning approaches have shown promising coding performance by leveraging rich semantic embeddings that can be optimized for both human and machine vision. However, these compact embeddings struggle to represent low-level details like contours and textures, leading to imperfect reconstructions. Additionally, existing learning-based coding tools lack scalability. To address these challenges, this paper presents a content-adaptive diffusion model for scalable image compression. The method encodes accurate texture through a diffusion process, enhancing human perception while preserving important features for machine vision tasks. It employs a Markov palette diffusion model with commonly-used feature extractors and image generators, enabling efficient data compression. By utilizing collaborative texture-semantic feature extraction and pseudo-label generation, the approach accurately learns texture information. A content-adaptive Markov palette diffusion model is then applied to capture both low-level texture and high-level semantic knowledge in a scalable manner. This framework enables elegant compression ratio control by flexibly selecting intermediate diffusion states, eliminating the need for deep learning model re-training at different operating points. Extensive experiments demonstrate the effectiveness of the proposed framework in image reconstruction and downstream machine vision tasks such as object detection, segmentation, and facial landmark detection. It achieves superior perceptual quality scores compared to state-of-the-art methods. Sha Guo, Zhuo Chen 0006, Yang Zhao 0002, Ning Zhang 0023, Ling-Yu Duan |
ACM Multimedia | 3 |
| 2023 | Dynamic Grouped Interaction Network for Low-Light Stereo Image EnhancementabstractLow-Light Stereo Image Enhancement (LLSIE) tackles the challenge of improving the illumination and restoring the details in stereo images. However, existing deep learning-based LLSIE methods trained on high-resolution low-light images often exhibit sub-optimal performance when interacting with information from the left and right views. We find that this is because of: (1) the high computational cost arising from quadratic complexity, which hinders the enhancement model's ability to process high-resolution images; and (2) the limitations of conventional fusion strategies in previous work, which inadequately capture cross-view cues, resulting in weak feature representation and compromised detail recovery. To address these limitations, we propose a novel Dynamic Grouped Interaction Network (DGI-Net) to enhance illumination and recover more details while reducing the computational cost. Specifically, DGI-Net employs the U-Net structure, which effectively mitigates noise during the low-light enhancement. Furthermore, we design a Grouped Stereo Interaction Module (GSIM) with a grouping strategy to efficiently discover cross-view cues while minimizing computations. To dynamically fuse stereo information and fully exploit cross-view correlations, we also introduce a Dynamic Embedding Module (DEM) to establish dynamic connections between inter-view cues and intra-view features, which performs dynamic weight processing on cross-view cues to eliminate noise during fusion. For intra-view processing, we present a Diversity Enhanced Block (DEB) to extract multi-scale features, thereby improving diversity and feature representation. This multi-scale feature extraction also addresses low image contrast in dark lighting conditions. Experimental results demonstrate that DGI-Net outperforms current state-of-the-art methods in low-light stereo image enhancement. Baiang Li, Zhao Zhang 0001, Yang Zhao 0002, Zhong-Qiu Zhao, Haijun Zhang 0002 |
ACM Multimedia | 4 |
| 2023 | Cross-view Resolution and Frame Rate Joint Enhancement for Binocular VideoabstractWith the popular of stereo video and free-viewpoint video, binocular and multi-view video enhancement has attracted increasing attention. Current binocular video enhancement methods mainly focus on stereo super-resolution. In this paper, we tend to discuss a new binocular video resolution and frame-rate enhancement scenario to fully utilize the cross-view complementary information. Specifically, one view is captured with high resolution (HR) and low frame-rate (LFR), while the other viewpoint records low resolution (LR) and high frame-rate (HFR) video. Then, a binocular video joint enhancement network, which adopts dual-branch structure with cross-view guidance, is proposed to jointly reconstruct HR and HFR stereo videos. The proposed framework can reduce the capture, storage, compression, and transmission cost of normal HR and HFR stereo videos. Compared with single-view super-resolution and video frame interpolation techniques, the proposed method can recover more realistic HR details and intermediate motion by using cross-view reference. Experimental results on stereo video datasets demonstrate the effectiveness of the proposed joint resolution and frame-rate enhancement framework. Panda Pan, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Zhao Zhang 0001, Ronggang Wang |
ACM Multimedia | 2 |
| 2023 | Boths: Super Lightweight Network-Enabled Underwater Image EnhancementabstractSince light is scattered and absorbed by water, underwater images have inherent degradation (e.g., hazing, color shift), consequently impeding the development of remotely operated vehicles (ROVs). Toward this end, we propose a novel method, referred to as${B}$est${o}\text{f}$Bo th World${s}$(Boths). With parameters of only 0.0064 M, Boths can be considered a super lightweight neural network for underwater image enhancement. On the whole, it has three levels: structure and detail features; pixel and channel dimensions; high- and low-frequency information. Each of these three levels represents “Best of Both Worlds.” Initially, by interacting with structure and detail features, Boths can focus on these two aspects at the same time. Further, our network can simultaneously consider channel and pixel dimensions through 3-D attention learning, which is more similar to human visual perception. Lastly, the proposed model can focus on high- and low-frequency information, through a novel loss function based on the wavelet transforms. Upon subsequent analysis and evaluation, Boths has shown superior performance compared with state-of-the-art (SOTA) methods. Our models and datasets are publicly available at:https://github.com/perseveranceLX/Boths. Sen Lin 0003, Kaichen Chi, Zhiyong Tao, Yang Zhao 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Hybrid feature enhancement network for few-shot semantic segmentation
Hai Min, Yemao Zhang, Yang Zhao 0002, Wei Jia 0001, Ying-Ke Lei, Chunxiao Fan 0002 |
Pattern Recognit. | 3 |
| 2023 | Fast Blind Decontouring NetworkabstractContouring artifacts usually appear in large and smooth flat areas, which are caused by many widely used processes such as bit-depth expansion, compression, image sharpening and contrast enhancement. Unfortunately, recent decontouring methods were mainly designed for specific and non-blind degradations, which significantly reduces the generalization ability of these methods when applied to complex and various real-world false contours. Therefore, this paper explores the blind decontouring problem by proposing a blind decontouring network (BDCN). Instead of directly training a decontouring network with mixed degradations, the proposed model consists of two independent modules, i.e., a flat region detection module (FDM) and a decontouring module (DCM). The FDM is designed to extract flat region masks robust to various false contours, which can preserve texture details from global smoothing. Then, the task of DCM becomes simply smoothing different contouring artifacts. Both the FDM and DCM are designed with a lightweight architecture and reparameterization strategy. Experimental results on both synthetic and real-world contouring artifacts demonstrate the effectiveness and generalization of the proposed method. Yang Zhao 0002, Wei Jia 0001, Yuan Chen 0012, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Learning Deep Blind Quality Assessment for Cartoon ImagesabstractAlthough the cartoon industry has developed rapidly in recent years, few studies pay special attention to cartoon image quality assessment (IQA). Unfortunately, applying blind natural IQA algorithms directly to cartoons often leads to inconsistent results with subjective visual perception. Hence, this brief proposes a blind cartoon IQA method based on convolutional neural networks (CNNs). Note that training a robust CNN depends on manually labeled training sets. However, for a large number of cartoon images, it is very time-consuming and costly to manually generate enough mean opinion scores (MOSs). Therefore, this brief first proposes a full reference (FR) cartoon IQA metric based on cartoon-texture decomposition and then uses the estimated FR index to guide the no-reference IQA network. Moreover, in order to improve the robustness of the proposed network, a large-scale dataset is established in the training stage, and a stochastic degradation strategy is presented, which randomly implements different degradations with random parameters. Experimental results on both synthetic and real-world cartoon image datasets demonstrate the effectiveness and robustness of the proposed method. Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | An iterative solution for improving the generalization ability of unsupervised skeleton motion retargeting
Shujie Li 0002, Wei Jia 0001, Yang Zhao 0002, Liping Zheng |
Comput. Graph. | 4 |
| 2022 | Cartoon Image Processing: A Survey
Yang Zhao 0002, Diya Ren, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
Int. J. Comput. Vis. | 1 |
| 2022 | EEPNet: An efficient and effective convolutional neural network for palmprint recognition
Wei Jia 0001, Yang Zhao 0002, Shujie Li 0002, Hai Min |
Pattern Recognit. Lett. | 3 |
| 2022 | Rethinking Deinterlacing for Early Interlaced VideosabstractIn recent years, high-definition restoration of early videos have received much attention. Real-world interlaced videos usually contain various degradations mixed with interlacing artifacts, such as noises and compression artifacts. Unfortunately, traditional deinterlacing methods only focus on the inverse process of interlacing scanning, and cannot remove these complex and complicated artifacts. Hence, this paper proposes an image deinterlacing network (DIN), which is specifically designed for joint removal of interlacing mixed with other artifacts. The DIN is composed of two stages,i.e., a cooperative vertical interpolation stage for splitting and fully using the information of adjacent fields, and a field-merging stage to perceive movements and suppress ghost artifacts. Experimental results demonstrate the effectiveness of the proposed DIN on both synthetic and real-world test sets. Yang Zhao 0002, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multiframe Joint Enhancement for Early Interlaced VideosabstractEarly interlaced videos usually contain multiple and interlacing and complex compression artifacts, which significantly reduce the visual quality. Although the high-definition reconstruction technology for early videos has made great progress in recent years, related research on deinterlacing is still lacking. Traditional methods mainly focus on simple interlacing mechanism, and cannot deal with the complex artifacts in real-world early videos. Recent interlaced video reconstruction deep deinterlacing models only focus on single frame, while neglecting important temporal information. Therefore, this paper proposes a multiframe deinterlacing network joint enhancement network for early interlaced videos that consists of three modules, i.e., spatial vertical interpolation module, temporal alignment and fusion module, and final refinement module. The proposed method can effectively remove the complex artifacts in early videos by using temporal redundancy of multi-fields. Experimental results demonstrate that the proposed method can recover high quality results for both synthetic dataset and real-world early interlaced videos. At the same time, the method also won the first place in the MSU Deinterlacer Benchmark. The code is available at: https://github.com/anymyb/MFDIN. Yang Zhao 0002, Yanbo Ma, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
IEEE Trans. Image Process. | 1 |
| 2022 | Audio Matters in Video Super-Resolution by Implicit Semantic GuidanceabstractVideo super-resolution (VSR) aims to use multiple consecutive low-resolution frames to recover the corresponding high-resolution frames. However, existing VSR methods only consider videos as image sequences, ignoring another essential timing informationaudio, while in fact, there is a semantic link between audio and vision, and extensive studies have shown that audio can provide supervisory information in visual networks. Meanwhile, the addition of semantic priors has been proven to be effective in super-resolution (SR) tasks, but a pretrained segmentation network is required to obtain semantic segmentation maps. By contrast, audio as the information contained in the video itself can be directly used. Therefore, in this study, we propose a novel and pluggable multiscale audiovisual fusion (MS-AVF) module to enhance VSR performance by exploiting the relevant audio information, which can be regarded as implicit semantic guidance compared with the kind of explicit segmentation priors. Specifically, we first fuse audiovisual features on the semantic feature maps of different granularities of the target frames, and then through a top-down multiscale fusion approach, feedback high-level semantics to the underlying global visual features layer by layer, thereby providing effective audio implicit semantic guidance for VSR. Experimental results show that audio can further improve the VSR effect. Moreover, by visualizing the learned attention mask, the proposed end-to-end model can automatically learn potential audiovisual semantic links, especially improving the accuracy and effectiveness of the SR of sound sources and their surrounding regions. Meibin Qi, Yang Zhao 0002, Wei Jia 0001, Ronggang Wang |
IEEE Trans. Multim. | 4 |
| 2022 | A Real-Time Semi-Supervised Deep Tone Mapping NetworkabstractTone mapping operators (TMOs) can compress the range of high dynamic range (HDR) images so that they can be displayed normally on the low dynamic range (LDR) devices. Recent TMOs based on deep neural networks can produce impressive results, but there are still some shortcomings. On the one hand, their supervised learning procedure requires a high-quality paired dataset which is hard to be accessed. On the other hand, they are too slow and heavy to meet the needs of practical applications. This paper proposes a real-time deep semi-supervised learning TMO to solve the above problems. The proposed method learns in a semi-supervised manner by combining the adversarial loss, cycle consistency loss, and the pixel-wise loss. The first two can simulate the image distributions in the real world from the unpaired LDR data and the latter can learn the guidance of paired LDR labels. In this way, the proposed method only requires HDR sources, unpaired high-quality LDR images, and a few well tone-mapped HDR-LDR pairs as training data. Furthermore, the proposed method divides tone mapping into luminance mapping and saturation adjustment and then processes them simultaneously. By this strategy, we can reconstruct each component more precisely. Based on the aforementioned improvements, we propose a lightweight tone mapping network that is efficient in tone mapping task (up to 5000x parameters-saving and 27x time-saving compared to the learning-based TMOs). Both quantitative and qualitative results demonstrate that the proposed method performs favorable against state-of-the-art TMOs. Ning Zhang 0023, Yang Zhao 0002, Chao Wang 0037, Ronggang Wang |
IEEE Trans. Multim. | 2 |
| 2021 | Deep Multi-loss Hashing Network for Palmprint Retrieval and RecognitionabstractWith the wide application of biometrics technology, the scale of biometrics databases is increasing rapidly. In this situation, fast retrieval technology is more and more necessary for large-scale biometrics retrieval and recognition. Palmprint recognition is one of the emerging biometrics technologies. However, the research on fast palmprint retrieval algorithm is still preliminary. Hashing is one of the most popular image retrieval technologies due to its fast speed and low storage cost. In this paper, we propose a new deep palmprint hashing method, which integrates classification loss, pairing loss and quantization loss in a unified deep learning framework. Experimental results show that the proposed deep multi-loss hashing method has better performance for palmprint recognition and retrieval than other existing classic hashing methods. Wei Jia 0001, Shuwei Huang, Lunke Fei, Yang Zhao 0002, Hai Min |
IJCB | 5 |
| 2021 | An Acceleration Framework for Super-Resolution Network via Region Difficulty Self-adaption
Zhenfang Guo, Yuyao Ye, Yang Zhao 0002, Ronggang Wang |
MMM (1) | 3 |
| 2021 | Real-time automatic helmet detection of motorcyclists in urban traffic using improved YOLOv5 detectorabstractAbstract In traffic accidents, motorcycle accidents are the main cause of casualties, especially in developing countries. The main cause of fatal injuries in motorcycle accidents is that motorcycle riders or passengers do not wear helmets. In this paper, an automatic helmet detection of motorcyclists method based on deep learning is presented. The method consists of two steps. The first step uses the improved YOLOv5 detector to detect motorcycles (including motorcyclists) from video surveillance. The second step takes the motorcycles detected in the previous step as input and continues to use the improved YOLOv5 detector to detect whether the motorcyclists wear helmets. The improvement of the YOLOv5 detector includes the fusion of triplet attention and the use of soft‐NMS instead of NMS. A new motorcycle helmet dataset (HFUT‐MH) is being proposed, which is larger and more comprehensive than the existing dataset derived from multiple traffic monitoring in Chinese cities. Finally, the proposed method is verified by experiments and compared with other state‐of‐the‐art methods. Our method achieves mAP of 97.7%, F1‐score of 92.7% and frames per second (FPS) of 63, which outperforms other state‐of‐the‐art detection methods. Wei Jia 0001, Shiquan Xu, Yang Zhao 0002, Hai Min, Shujie Li 0002 |
IET Image Process. | 4 |
| 2021 | A survey on dorsal hand vein biometrics
Wei Jia 0001, Bob Zhang 0001, Yang Zhao 0002, Lunke Fei, Wenxiong Kang, Di Huang 0001, Guodong Guo |
Pattern Recognit. | 4 |
| 2021 | Lighter but Efficient Bit-Depth Expansion NetworkabstractWith the development of display technology, bit-depth expansion (BDE) has emerged as a basic process to display low-bit-depth image and video resources on high-bit-depth monitors. Most current BDE methods are based on traditional algorithms, and the few existing methods based on deep neural networks still suffer from loss of pixel-level details or from high computational cost. This paper proposes a lightweight but efficient BDE network that can effectively improve the capacity of shallow network by introducing a residual-block-in-residual-block structure. Furthermore, the proposed network adopts residual network architecture and dilated convolution to balance the preservation of pixel-level information and the expansion of the receptive field. Hence, the proposed method can also totally remove significant artifacts from very low-bit-depth images. Experimental results demonstrate that the proposed method can achieve performance comparable to or even better than that of some state-of-the-art methods while having much lighter architecture and fewer parameters. Yang Zhao 0002, Ronggang Wang, Yuan Chen 0012, Wei Jia 0001, Xiaoping Liu 0003, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Handling Outliers by Robust M-Estimation in Blind Image DeblurringabstractThe major task of traditional motion deblurring methods is to estimate the blur kernel and restore the latent image. In low-light conditions, the pointolite is likely to produce saturated light streaks in captured blurred images. The light streaks are usually double-edged swords—outliers to the deconvolution, but a cue to kernel estimation. In this paper, we propose a novel blind motion deblurring method for blurred images including light streaks. The main idea is to model the non-linear blur caused by outliers as the Huber's M-estimation in blind deconvolution and take the shape of the light streak as a cue to estimate the blur kernel. Specifically, the optimal light streak patch is selected automatically according to the characteristics of light streaks and the blur kernel. This simple yet effective selection strategy solves the problems of false detection of candidate light streaks and optimal light streak in existing methods. Then, the optimal light streak patch is parameterized as a prior and is combined with other regularizers to estimate the blur kernel. Compared with the state-of-the-art kernel estimation methods, the proposed algorithm reduces the influence of outliers on deconvolution and utilizes more information. Thus, the restored image is more accurate. Experimental results on both synthetic and real images demonstrate the high accuracy of our algorithm. Xinxin Zhang 0004, Ronggang Wang, Da Chen 0002, Yang Zhao 0002, Wen Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | A Flexible Recurrent Residual Pyramid Network for Video Frame Interpolation
Haoxian Zhang, Yang Zhao 0002, Ronggang Wang |
ECCV (25) | 2 |
| 2020 | Estimated Exposure Guided Reconstruction Model for Low-Light Image Enhancement
Xiaona Liu, Yang Zhao 0002, Yuan Chen 0012, Wei Jia 0001, Ronggang Wang, Xiaoping Liu 0003 |
PRCV (1) | 2 |
| 2020 | A night-time outdoor data set for low-light enhancementabstractLow light Enhancement has been a hot topic in recent years, and many deep neural network (DNN)-based methods have achieved remarkable performance. However, the rapid development of DNNs also raises the urgent requirement of high-quality training sets, especially supervised night-time data sets. In this paper, we establish a night-time outdoor data set (NOD1) that contains 1214 groups of images. We also generate appropriate and high-quality reference images for each group based on multi-exposure fusion strategy, which not only focuses on dark areas but also provides details for over-exposed areas in low light images. Furthermore, a simple but efficient network is presented as the baseline of NOD. Experimental results on NOD and other data sets show the generalizability and effectiveness of the proposed data set and baseline model. Yudong Zhou, Ronggang Wang, Yang Zhao 0002 |
VCIP | 3 |
| 2020 | Adversarial-learning-based image-to-image transformation: A survey
Yuan Chen 0012, Yang Zhao 0002, Wei Jia 0001, Xiaoping Liu 0003 |
Neurocomputing | 2 |
| 2020 | Blind Quality Assessment for Cartoon ImagesabstractCurrent blind image quality assessment (BIQA) algorithms are mainly designed for natural images. Unfortunately, cartoon and cartoon-like images are quite different from natural images. Hence, recent BIQA methods are not very robust to cartoon images. In this paper, we propose a specific BIQA algorithm designed for cartoon images, which consists of the following terms. First, a cartoon image is divided into edge areas and nonedge areas via a Tchebichef moment (TM)-based process. Second, a multiorder sharpness statistic term is used to measure the quality of the edges, and a sharpness statistic prior model of high-quality (HQ) cartoon images is built. Finally, a local encoding statistic term is adopted to describe the textural complexity in the nonedge areas, and a texture statistic prior model is also established. The experimental results on the cartoon image datasets demonstrate that the proposed method can accurately evaluate the visual quality of cartoon images and is more suitable for cartoon scenarios than some traditional BIQA algorithms. Yuan Chen 0012, Yang Zhao 0002, Shujie Li 0002, Wangmeng Zuo, Wei Jia 0001, Xiaoping Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep tone mapping network in HSV color spaceabstractTone mapping operators can convert high dynamic range (HDR) images to low dynamic range (LDR) images so that we can enjoy the informative contents of HDR images with LDR devices. However, current state-of-the-art tone mapping algorithms mainly focus on the luminance mapping while neglecting the color component. Meanwhile, they often suffer from halo artifacts and over-enhancement. In this paper, we propose a tone mapping network (TMNet) in Hue-Saturation-Value (HSV) color space to obtain better luminance and color mapping. We adopt the improved Wasserstein generative adversarial network (WGAN-GP) as the basic architecture and further introduce several improvements. A meticulously designed loss function is adopted to push tone mapped image to the natural image manifold. What’s more, we create a tone mapped image dataset in which the label images are manually adjusted by photographers. Compared with some state-of-the-art tone mapping methods, the proposed method can achieve better performance in both subjective and objective evaluations. Ning Zhang 0023, Chao Wang 0037, Yang Zhao 0002, Ronggang Wang |
VCIP | 3 |
| 2019 | Bidirectional recurrent autoencoder for 3D skeleton motion data refinement
Shujie Li 0002, Haisheng Zhu, Wenjun Xie, Yang Zhao 0002, Xiaoping Liu 0003 |
Comput. Graph. | 5 |
| 2019 | Deep learning-based methods for person re-identification: A comprehensive review
Di Wu 0030, Si-Jia Zheng, Xiao-Ping Zhang 0002, Chang-an Yuan 0001, Yang Zhao 0002, Yong-Jun Lin, Zhong-Qiu Zhao, Yong-Li Jiang, De-Shuang Huang |
Neurocomputing | 6 |
| 2019 | Deep Reconstruction of Least Significant Bits for Bit-Depth ExpansionabstractBit-depth expansion (BDE) is important for displaying a low bit-depth image in a high bit-depth monitor. Current BDE algorithms often utilize traditional methods to fill the missing least significant bits and suffer from multiple kinds of perceivable artifacts. In this paper, we present a deep residual network-based method for BDE. Based on the different properties of flat and non-flat areas, two channels are proposed to reconstruct these two kinds of areas, respectively. Moreover, a simple yet efficient local adaptive adjustment preprocessing is presented in the flat-area-channel. By combining the benefits of both the traditional debanding strategy and network-based reconstruction, the proposed method can further promote the subjective quality of the flat area. Experimental results on several image sets demonstrate that the proposed BDE network can obtain favorable visual quality as well as decent quantitative performance. Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wangmeng Zuo, Xiaoping Liu 0003, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Speaker-Invariant Training Via Adversarial LearningabstractWe propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based ASR system. We call the scheme speaker-invariant training (SIT). In SIT, a DNN acoustic model and a speaker classifier network are jointly optimized to minimize the senone (tied triphone state) classification loss, and simultaneously mini-maximize the speaker classification loss. A speaker-invariant and senone-discriminative deep feature is learned through this adversarial multi-task learning. With SIT, a canonical DNN acoustic model with significantly reduced variance in its output probabilities is learned with no explicit speaker-independent (SI) transformations or speaker-specific representations used in training or testing. Evaluated on the CHiME-3 dataset, the SIT achieves 4.99% relative word error rate (WER) improvement over the conventional SI acoustic model. With additional unsupervised speaker adaptation, the speaker-adapted (SA) SIT model achieves 4.86% relative WER gain over the SA SI acoustic model. Zhong Meng, Jinyu Li 0001, Zhuo Chen 0006, Yang Zhao 0002, Vadim Mazalov, Yifan Gong 0001, Biing-Hwang Juang |
ICASSP | 4 |
| 2018 | A Hybrid Deep Model for Person Re-Identification
Di Wu 0030, Si-Jia Zheng, Yang Zhao 0002, Chang-an Yuan 0001, Xiao Qin 0005, Yong-Li Jiang, De-Shuang Huang |
ICIC (3) | 4 |
| 2018 | A Simple and Effective Deep Model for Person Re-identification
Si-Jia Zheng, Di Wu 0030, Yang Zhao 0002, Chang-an Yuan 0001, Xiao Qin 0005, De-Shuang Huang |
ICIC (3) | 4 |
| 2018 | An effective local regional model based on salient fitting for image segmentation
Hai Min, Wei Jia 0001, Yang Zhao 0002, Yue-Tong Luo |
Neurocomputing | 4 |
| 2018 | Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001 |
Inf. Sci. | 1 |
| 2018 | A polynomial piecewise constant approximation method based on dual constraint relaxation for segmenting images with intensity inhomogeneity
Hai Min, Wei Jia 0001, Yang Zhao 0002, Yue-Tong Luo |
Pattern Recognit. | 4 |
| 2018 | LATE: A Level-Set Method Based on Local Approximation of Taylor Expansion for Segmenting Intensity Inhomogeneous ImagesabstractIntensity inhomogeneity is common in real-world images and inevitably leads to many difficulties for accurate image segmentation. Numerous level-set methods have been proposed to segment images with intensity inhomogeneity. However, most of these methods are based on linear approximation, such as locally weighted mean, which may cause problems when handling images with severe intensity inhomogeneities. In this paper, we view segmentation of such images as a nonconvex optimization problem, since the intensity variation in such an image follows a nonlinear distribution. Then, we propose a novel level-set method named local approximation of Taylor expansion (LATE), which is a nonlinear approximation method to solve the nonconvex optimization problem. In LATE, we use the statistical information of the local region as a fidelity term and the differentials of intensity inhomogeneity as an adjusting term to model the approximation function. In particular, since the first-order differential is represented by the variation degree of intensity inhomogeneity, LATE can improve the approximation quality and enhance the local intensity contrast of images with severe intensity inhomogeneity. Moreover, LATE solves the optimization of function fitting by relaxing the constraint condition. In addition, LATE can be viewed as a constraint relaxation of classical methods, such as the region-scalable fitting model and the local intensity clustering model. Finally, the level-set energy functional is constructed based on the Taylor expansion approximation. To validate the effectiveness of our method, we conduct thorough experiments on synthetic and real images. Experimental results show that the proposed method clearly outperforms other solutions in comparison. Hai Min, Wei Jia 0001, Yang Zhao 0002, Wangmeng Zuo, Haibin Ling, Yue-Tong Luo |
IEEE Trans. Image Process. | 3 |
| 2017 | Iterative projection reconstruction for fast and efficient image upsampling
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Wenmin Wang 0001, Wen Gao 0001 |
Neurocomputing | 1 |
| 2017 | Palmprint Recognition Based on Complete Direction RepresentationabstractDirection information serves as one of the most important features for palmprint recognition. In the past decade, many effective direction representation (DR)-based methods have been proposed and achieved promising recognition performance. However, due to an incomplete understanding for DR, these methods only extract DR in one direction level and one scale. Hence, they did not fully utilize all potentials of DR. In addition, most researchers only focused on the DR extraction in spatial coding domain, and rarely considered the methods in frequency domain. In this paper, we propose a general framework for DR-based method named complete DR (CDR), which reveals DR by a comprehensive and complete way. Different from traditional methods, CDR emphasizes the use of direction information with strategies of multi-scale, multi-direction level, multi-region, as well as feature selection or learning. This way, CDR subsumes previous methods as special cases. Moreover, thanks to its new insight, CDR can guide the design of new DR-based methods toward better performance. Motived this way, we propose a novel palmprint recognition algorithm in frequency domain. First, we extract CDR using multi-scale modified finite radon transformation. Then, an effective correlation filter, namely, band-limited phase-only correlation, is explored for pattern matching. To remove feature redundancy, the sequential forward selection method is used to select a small number of CDR images. Finally, the matching scores obtained from different selected features are integrated using score-level-fusion. Experiments demonstrate that our method can achieve better recognition accuracy than the other state-of-the-art methods. More importantly, it has fast matching speed, making it quite suitable for the large-scale identification applications. Wei Jia 0001, Bob Zhang 0001, Yihai Zhu, Yang Zhao 0002, Wangmeng Zuo, Haibin Ling |
IEEE Trans. Image Process. | 5 |
| 2016 | Local Quantization Code histogram for texture classification
Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
Neurocomputing | 1 |
| 2016 | Multilevel Modified Finite Radon Transform Network for Image UpsamplingabstractA local line-like feature is the most important discriminate information in the image upsampling scenario. In recent example-based upsampling methods, grayscale and gradient features are often adopted to describe the local patches, but these simple features cannot accurately characterize complex patches. In this paper, we present a feature representation of local edges by means of a multilevel filtering network, namely, multilevel modified finite Radon transform network (MMFRTN). In the proposed MMFRTN, the MFRT is utilized in the filtering layer to extract the local line-like feature; the nonlinear layer is set to be a simple local binary process; for the feature-pooling layer, we concatenate the mapped patches as the feature of local patch. Then, we propose a new example-based upsampling method by means of the MMFRTN feature. Experimental results demonstrate the effectiveness of the proposed method over some state-of-the-art methods. Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | An Intensity-Texture model based level set method for image segmentation
Hai Min, Wei Jia 0001, Yang Zhao 0002, Rong-Xiang Hu, Yue-Tong Luo, Feng Xue 0002 |
Pattern Recognit. | 4 |
| 2015 | High Resolution Local Structure-Constrained Image UpsamplingabstractWith the development of ultra-high-resolution display devices, the visual perception of fine texture details is becoming more and more important. A method of high-quality image upsampling with a low cost is greatly needed. In this paper, we propose a fast and efficient image upsampling method that makes use of high-resolution local structure constraints. The average local difference is used to divide a bicubic-interpolated image into a sharp edge area and a texture area, and these two areas are reconstructed separately with specific constraints. For reconstruction of the sharp edge area, a high-resolution gradient map is estimated as an extra constraint for the recovery of sharp and natural edges; for the reconstruction of the texture area, a high-resolution local texture structure map is estimated as an extra constraint to recover fine texture details. These two reconstructed areas are then combined to obtain the final high-resolution image. The experimental results demonstrated that the proposed method recovered finer pixel-level texture details and obtained top-level objective performance with a low time cost compared with state-of-the-art methods. Yang Zhao 0002, Ronggang Wang, Wenmin Wang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Angular Pattern and Binary Angular Pattern for Shape RetrievalabstractIn this paper, we propose two novel shape descriptors, angular pattern (AP) and binary angular pattern (BAP), and a multiscale integration of them for shape retrieval. Both AP and BAP are intrinsically invariant to scale and rotation. More importantly, being global shape descriptors, the proposed shape descriptors are computationally very efficient, while possessing similar discriminability as state-of-the-art local descriptors. As a result, the proposed approach is attractive for real world shape retrieval applications. The experiments on the widely used MPEG-7 and TARI-1000 data sets demonstrate the effectiveness of the proposed method in comparison with existing methods. Rong-Xiang Hu, Wei Jia 0001, Haibin Ling, Yang Zhao 0002, Jie Gui |
IEEE Trans. Image Process. | 4 |
| 2014 | Histogram of Oriented Lines for Palmprint RecognitionabstractSubspace learning methods are very sensitive to the illumination, translation, and rotation variances in image recognition. Thus, they have not obtained promising performance for palmprint recognition so far. In this paper, we propose a new descriptor of palmprint named histogram of oriented lines (HOL), which is a variant of histogram of oriented gradients (HOG). HOL is not very sensitive to changes of illumination, and has the robustness against small transformations because slight translations and rotations make small histogram value changes. Based on HOL, even some simple subspace learning methods can achieve high recognition rates. Wei Jia 0001, Rong-Xiang Hu, Ying-Ke Lei, Yang Zhao 0002, Jie Gui |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2013 | Completed robust local binary pattern for texture classification
Yang Zhao 0002, Wei Jia 0001, Rong-Xiang Hu, Hai Min |
Neurocomputing | 1 |
| 2012 | An Efficient Multi-scale Overlapped Block LBP Approach for Leaf Image Recognition
Xiao-Ming Ren, Yang Zhao 0002 |
ICIC (2) | 3 |
| 2012 | Perceptually motivated morphological strategies for shape retrieval
Rong-Xiang Hu, Wei Jia 0001, Yang Zhao 0002, Jie Gui |
Pattern Recognit. | 3 |
| 2012 | Completed Local Binary Count for Rotation Invariant Texture ClassificationabstractIn this brief, a novel local descriptor, named local binary count (LBC), is proposed for rotation invariant texture classification. The proposed LBC can extract the local binary grayscale difference information, and totally abandon the local binary structural information. Although the LBC codes do not represent visual microstructure, the statistics of LBC features can represent the local texture effectively. In addition, a completed LBC (CLBC) is also proposed to enhance the performance of texture classification. Experimental results obtained from three databases demonstrate that the proposed CLBC can achieve comparable accurate classification rates with completed local binary pattern. Yang Zhao 0002, De-Shuang Huang, Wei Jia 0001 |
IEEE Trans. Image Process. | 1 |