EDBT 2026 Demo / reviewers in the wild / expert
Qing Zhu 0004
dblp:74/963-4
· DBLP profile ↗
51ranked-venue papers
1as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 23 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UCAMNet: HVI Color Space Based Unsupervised Low-Light Enhancement via Uncertainty Constraint and Attention Mechanism
Jingshuo Guan, Na Qi, Qing Zhu 0004, Liang Chen 0026 |
MMM (2) | 3 |
| 2026 | UQuadCGAN: Uncertainty-driven cycle-consistent GAN with channel-spatial guided attention for low-light image enhancement
Jingshuo Guan, Na Qi, Qing Zhu 0004, Liang Chen 0026 |
Neurocomputing | 3 |
| 2026 | FIB: Find Interaction Behavior on Social Media for Rumor DetectionabstractRumor detection is a key concern on social media, where rich user interactions pose significant challenges to rumor governance. And it also contributes greatly to rumor detection. However, existing multimodal rumor detection methods suffer from the following limitations. First, they overlook the dynamic nature of social interactions and fail to consider the evolving influence of social context during rumor propagation. Second, they rely on manually constructed social patterns to extract interaction features, which introduces subjective bias and lacks adaptability to changing social environments. Finally, they lack effective fusion strategies for integrating multiple modalities, such as text, images, and social context. To address the issues above, we propose a novel approach to find interaction behavior (FIB) on social media for rumor detection. Specifically, we first infer the latent interaction behavior to mitigate the adverse effects of snapshot limitations when leveraging social modalities. Then, instead of manually constructing, we capture interaction patterns by learning the interaction path automatically. Finally, we design a cross-modal fusion method based on a selective state space model to integrate complex relationships across multiple modalities. Extensive experiments on five large-scale datasets show that the proposed FIB model consistently outperforms all the state-of-the-art rumor detection baselines by 0.09%–4.3% on F1 score. Qing Zhu 0004, Yang Xiao 0017 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Hyper-NeuS: Hypernetworks for Neural SDF Implicit Surface Reconstruction by Volume Rendering
Jingkun Li, Na Qi, Qing Zhu 0004 |
MMM (2) | 3 |
| 2025 | Self-supervised Reference-Based Image Super-Resolution with Conditional Diffusion Model
Na Qi, Yezi Li, Qing Zhu 0004 |
MMM (3) | 4 |
| 2025 | SSCDUF: Spatial-Spectral Correlation Transformer Based on Deep Unfolding Framework for Hyperspectral Image Reconstruction
Na Qi, Qing Zhu 0004, Xiumin Lin |
MMM (4) | 3 |
| 2025 | S2TRAT: Image Style Transfer with Similarity Metric-Guided Region Aware Transformer
Na Qi, Yezi Li, Liang Chen 0026, Qing Zhu 0004 |
PRCV (9) | 5 |
| 2025 | GarTrans: Transformer-Based Architecture for Dynamic and Detailed Garment DeformationabstractIn this paper, we introduce GarTrans, a novel graph-learning based method for the task of garment animation. It emphasizes efficiently rendering realistic deformation effects. GarTrans goes beyond existing models by providing improved generalization capabilities, along with the ability to capture fine-scale garment dynamics and details. Our approach begins by constructing a garment graph that comprehensively encodes the dynamic state of the garment, taking into account its shape and topology, as well as the underlying body shape and corresponding motion. We have also designed a structure-augmented transformer (SAT) capable of processing the node information and edges within the graph, enabling the generation of deformation details that are contextually informed. Our model employs a unified optimization scheme that incorporates both supervised and unsupervised loss functions, enabling a robust approach capable of realistically mimicking the behavior of intricate garments. Experimental evaluations show that our method surpasses the existing state-of-the-art in terms of both functional capabilities and visual fidelity, advancing the field of garment animation. Tianxing Li 0002, Zhi Qiao 0006, Zihui Li, Qing Zhu 0004 |
Comput. Vis. Media | 5 |
| 2025 | STRAT: Image style transfer with region-aware transformer
Na Qi, Yezi Li, Qing Zhu 0004 |
Neurocomputing | 4 |
| 2025 | DualFLAT: Dual Flat-Lattice Transformer for domain-specific Chinese named entity recognition
Yinlong Xiao, Zongcheng Ji, Jianqiang Li 0002, Qing Zhu 0004 |
Inf. Process. Manag. | 4 |
| 2025 | Spectrum-Enhanced Graph Attention Network for Garment Mesh DeformationabstractWe present a novel solution for mesh-based deformation simulation from a spectral perspective. Unlike existing approaches that demand separate training for each garment or body type and often struggle to produce rich folds and lifelike dynamics, our method achieves the quality of physics-based simulations while maintaining superior efficiency within a unified model. The key to achieve this lies in the development of a spectrum-enhanced deformation network, a result of in-depth theoretical analysis bridging neural networks and garment deformations. This enhancement compels the network to focus on learning spectral information predominantly within the frequency band associated with intricate deformations. Furthermore, building upon standard blend skinning techniques, we introduce target-aware temporal skinning weights. The weights describe how the underlying human skeleton dynamically affects the mesh vertices according to the garment and body shape, as well as the motion state. We validate our method on various garments, bodies, and motions through extensive ablation studies. Finally, we conduct comparisons to confirm its superiority in generalization, deformation quality, and performance over several state-of-the-art methods. Tianxing Li 0002, Qing Zhu 0004, Liguo Zhang 0001, Takashi Kanai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Frequency-Divided Learning of Fine-Grained Clothing Behavior via Flexible Dynamic GraphsabstractDespite significant advancements in neural simulation techniques for clothing animation, these methods struggle to capture the dynamic details of garments during movement. This limitation restricts their applicability in scenarios where high-quality garment deformation is essential. To address this challenge, we introduce a novel graph learning-based approach to enhance deformation realism through designed mechanisms for mesh information propagation and external optimization strategies during model training. First, we address the issue of over-smoothing common in conventional graph processing techniques by introducing a flexible message-passing method. This approach effectively manages node interactions within the mesh, thereby improving the expressiveness of the model. Furthermore, acknowledging that uniform model supervision typically neglects high-frequency details during optimization, we analyze the spectral properties of clothing meshes. Based on this analysis, we introduce a frequency-division constraint aligned with the characteristics of different frequency bands, which aids in precisely controlling the generation of details. Our model further integrates self-collision and other physics-aware losses, enabling the learning of generalized and fine-grained dynamic deformations. Extensive evaluations and comparisons demonstrate the effectiveness of our approach, showing notable improvements over existing state-of-the-art solutions. Tianxing Li 0002, Takashi Kanai, Qing Zhu 0004 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | SBC-Net: semantic-guided brightness curve estimation network for low-light image enhancement
Shize Wang, Jin Wang 0023, Qing Zhu 0004, Yunhui Shi |
Vis. Comput. | 4 |
| 2024 | UTrCGAN: Uncertainty-Driven Cycle-Consistent Generative Adversarial Network for Low-Light Image EnhancementabstractLow-light image enhancement is a computer vision task that aims to improve the visual perceptual quality of images captured in poorly illuminated scenes. At present, deep learning-based low-light enhancement methods can obtain high-quality enhanced images. However, it does not consider the statistical characteristics of different regions, such as edge, structure, and texture. The uncertainty of image regions is not well characterized and utilized. To address this problem, we propose a novel UnCertainty-driven Cycle-Consistent Generative Adversarial Network (UTrCGAN) to improve the performance of low-light enhancement. UTrCGAN first decomposes the unpaired low/normal-light images into reflectance and illumination components based on the Retinex theory. Then a generative adversarial network guided by uncertainty constraint is proposed to enhance the illumination component, in which the quality of the enhanced image is further improved by the guidance of variance estimation. Experimental results on the widely-used LOL dataset show that UTrCGAN outperforms the state-of-the-art methods in terms of visual quality and quantitative metrics. Jingshuo Guan, Na Qi, Qing Zhu 0004, Liang Chen 0026 |
ICIP | 3 |
| 2024 | Coarse-To-Fine Spatio-Temporal Luminance-Aware Reconstruction For High-Speed Motion SceneabstractThe continuous emission of spike stream offers more significant advantages over traditional fixed low sampling rate cameras. Although many reconstruction methods from spike streams have been proposed, the quality of recovered images remains suboptimal. Coarse-to-fine high-speed motion scene reconstruction reconstructs the spike sequence by dividing the dynamic and static regions. However, issues of limited texture richness and low contrasts are unsolved during the reconstruction of static spike. To address these issues, we propose a Coarse-to-Fine spatio-temporal Luminance-Aware Reconstruction (CFLAR) framework. Specially, we propose an adaptive luminance-aware reconstruction in spatio-temporal domain. To be specific, in the spatial domain, we perform region division and region merging of spike sequences based on luminance information, while non-uniform quantization of luminance information is mainly achieved through binary division and adaptive parameters division. In the temporal domain, we integrate alterable window length into the texture from playback to propose adaptive time shift window reconstruction, which enables to obtain reconstructed images with richer textures and higher contrasts. Experimental results demonstrate that our CFLAR method outperforms state-of-the-art approaches in terms of objective and subjective quality. Zhangke Wang, Na Qi, Wei Xu 0059, Jingzhong Qi, Qing Zhu 0004 |
ICIP | 6 |
| 2024 | SwinGar: Spectrum-Inspired Neural Dynamic Deformation for Free-Swinging GarmentsabstractOur work presents a novel spectrum-inspired learning-based approach for generating clothing deformations with dynamic effects and personalized details. Existing methods in the field of clothing animation are limited to either static behavior or specific network models for individual garments, which hinders their applicability in real-world scenarios where diverse animated garments are required. Our proposed method overcomes these limitations by providing a unified framework that predicts dynamic behavior for different garments with arbitrary topology and looseness, resulting in versatile and realistic deformations. First, we observe that the problem of bias towards low frequency always hampers supervised learning and leads to overly smooth deformations. To address this issue, we introduce a frequency-control strategy from a spectral perspective that enhances the generation of high-frequency details of the deformation. In addition, to make the network highly generalizable and able to learn various clothing deformations effectively, we propose a spectral descriptor to achieve a generalized description of the global shape information. Building on the above strategies, we develop a dynamic clothing deformation estimator that integrates graph attention mechanisms with long short-term memory. The estimator takes as input expressive features from garments and human bodies, allowing it to automatically output continuous deformations for diverse clothing types, independent of mesh topology or vertex count. Finally, we present a neural collision handling method to further enhance the realism of garments. Our experimental results demonstrate the effectiveness of our approach on a variety of free-swinging garments and its superiority over state-of-the-art methods. Tianxing Li 0002, Qing Zhu 0004, Takashi Kanai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Cluster-based two-branch framework for point cloud attribute compression
Longhua Sun, Jin Wang 0023, Qing Zhu 0004, Jiaying Liu 0015, Jiawen Yu |
Vis. Comput. | 3 |
| 2023 | G2CNN: Geometric Prior Based GCNN for Single-View 3D Reconstruction with Loop SubdivisionabstractSingle-view 3D reconstruction is a fundamental operation in computer vision. Although significant progress has been made by learning-based approaches, it remains a challenge that the reconstructed mesh is usually coarse since the geometric prior is ignored. In this paper, we propose a geometric prior based graph convolution neural network model (named G2CNN) for single-view 3D reconstruction with Loop subdivision. G2CNN is a data-driven deep neural network (DNN) with the geometry knowledge. To make the reconstructed results with abundant geometric details, we generate shapes with a coarse-to-fine strategy and utilize the Gaussian curvature loss as a geometric supervision. Furthermore, to produce the physically accurate 3D geometry, the mesh subdivision module is designed with Loop subdivision to exploit the vertex localizations and connectivity, which can refine and smooth the mesh surface. Experimental results on both synthesized data and real data demonstrate the effectiveness of our method in terms of both subjective and objective quality. Na Qi, Wei Xu 0059, Qing Zhu 0004, Shibo Xu, Changxin Pan |
ICASSP | 4 |
| 2023 | Color Guided Depth Map Super-Resolution with Nonlocla Autoregres-Sive ModelingabstractDepth map captured by 3D cameras usually suffers from low resolution and insufficient quality, which limits its applications in real world. Thus, it is an essential task to develop efficient and effective techniques to handle various depth degradations. In this paper, we propose a color guided depth map super-resolution method with nonlocal autoregressive modeling. Considering that textures in depth map demonstrate distinct geometry direction, we exploit the multi-directional dictionary which is effective in recovering subtle structures of depth patches. We further introduce two regularization terms into the sparse representation framework. Firstly, a patch based autoregressive model is introduced to represent the local patterns in a small area. Secondly, inspired by the structure consistence between depth map and color image, we propose a color guided nonlocal similarity to provide nonlocal constraint to the local structures, which is very helpful in preserving local structures and suppressing noise. Experimental results demonstrate the superior of our method compared with state-of-the-art methods. Wei Xu 0059, Na Qi, Qing Zhu 0004, Jingzhong Qi, Longlu Huang, Yuxin Bao |
ICASSP | 3 |
| 2023 | SwinCGH-Net: Enhancing Robustness of Object Detection in Autonomous Driving with Weather Noise via Attention
Shi Cao, Qing Zhu 0004, Wanting Zhu |
ICIC (5) | 2 |
| 2023 | A Lightweight Detail-Fusion Progressive Network for Image Deraining
Siyi Ding, Qing Zhu 0004, Wanting Zhu |
ICIC (5) | 2 |
| 2023 | Octree-Based Temporal-Spatial Context Entropy Model for LiDAR Point Cloud CompressionabstractIt’s difficult to effectively remove redundancy in Li-DAR point clouds due to their extremely sparse and nonuniform distribution. Taking advantage of both octree-based methods and voxel-based schemes, we propose to design an effective temporal-spatial context to compress the sequence octree-structured point cloud data into a more compact bitstream. In this paper, we first build a temporal-spatial multiscale context for the deep learning entropy model. It further utilize the correlation of sequential point cloud data from both the spatial domain and temporal domain. In terms of spatial context, we design a hierarchical dependency in an octree to encode the occupancy information of each non-leaf octree node into a bitstream. We propose to further group the nodes according to their octant which effectively expands the context receptive field. In terms of temporal context, the KNN algorithm is applied to explore the most relative context with the strongest dependency in the temporal domain. Finally, we design a voxel re-localization network to convert the discrete voxels into refined 3D points, which makes up for the coordinate loss in the process of generating an octree. The quantitative evaluation shows that our method outperforms state-of-the-art baselines with saving most bitrate on KITTI Odometry dataset, and achieving the best reconstreuction benefit by the designed refinement module. Longhua Sun, Jin Wang 0023, Yunhui Shi, Qing Zhu 0004, Nam Ling |
VCIP | 4 |
| 2022 | Multispectral Image Denoising via Structural Tensor Sparsity Promoting ModelabstractMultispectral images (MSIs) contain more spectral information than traditional 2D images, which can provide a more accurate representation of objects. MSIs are easily affected by various noises when captured by sensors. In recent years, many MSI denoising methods, especially the Kronecker-basis-representation (KBR) method, have achieved great success. KBR uses tensor representation and decomposition to achieve good MSI denoising performance. However, each full band patch (FBP) group is decomposed in this method so that too many dictionary atoms are generated. In this paper, we propose a structural tensor sparsity promoting (STSP) model for MSI denoising. In order to decrease the number of dictionary atoms, we cluster FBP groups and learn orthogonal dictionaries for each class rather than each FBP group. To improve the denoising performance, the structural similarity among FBP groups are utilized in the STSP model by enforcing nonlocal centralized sparse constraint, where the compromise parameter is statistically and adaptively determined. Experimental results on the the CAVE dataset demonstrate that our model outperforms the state-of-art methods in terms of both objective and subjective quality. Longlu Huang, Na Qi, Qing Zhu 0004 |
MMAsia | 3 |
| 2022 | Progressive GAN-Based Transfer Network for Low-Light Image Enhancement
Na Qi, Qing Zhu 0004, Haoran Ouyang |
MMM (2) | 3 |
| 2022 | SUnet++: Joint Demosaicing and Denoising of Extreme Low-Light Raw Image
Jingzhong Qi, Na Qi, Qing Zhu 0004 |
MMM (2) | 3 |
| 2022 | Enhancing Chinese Medical Named Entity Recognition with Auto-Mined LexiconabstractRecently, lexicon-based Chinese Named Entity Recognition (NER) models have achieved state-of-the-art performance by benefiting from the rich boundary and semantic information contained in the lexicon. However, in the Chinese medical domain, it’s difficult to obtain the medical lexicon related to the target medical corpus. In this paper, we propose a new paradigm, enhancing Chinese medical NER with Auto-mined Lexicon (ALNER), which alleviates the difficulty of obtaining the medical lexicon by designing a data-driven automatic lexicon construction method. We define medical lexicon construction as a high-quality phrase mining task. We perform secondary annotation on the NER annotated data and use the secondary annotated data to train a deep learning-based phrase tagger. Experimental results show that our method can be combined with different lexicon-based Chinese NER models to improve performance and that the method does not require an external medical lexicon. Yinlong Xiao, Jianqiang Li 0002, Qing Zhao 0005, Qing Zhu 0004, Yu-Chih Wei |
SMC | 4 |
| 2022 | Depth Map Super-Resolution Based on Dual Normal-Depth Regularization and Graph Laplacian PriorabstractThe edge information plays a key role in the restoration of a depth map. Most conventional methods assume that the color image and depth map are consistent in edge areas. However, complex texture regions in the color image do not match exactly with edges in the depth map. In this paper, firstly, we point out that in most cases the consistency between normal map and depth map is much higher than that between RGB-D pairs. Then we propose a dual normal-depth regularization term to guide the restoration of depth map, which constrains the edge consistency between normal map and depth map back and forth. Moreover, considering the bimodal characteristic of weight distribution that exists in depth discontinuous areas, a reweighted graph Laplacian regularizer is proposed to promote this bimodal characteristic. And this regularization is incorporated into a unified optimization framework to effectively protect the piece-wise smoothness(PWS) characteristics of depth map. By treating depth image as graph signal, the weight between two nodes is adapted according to its content. The proposed method is tested for both noise-free and noisy cases, and is compared against the state-of-the-art methods on both synthesis and real captured datasets. Extensive experimental results demonstrate the superior performance of our method compared with most state-of-the-art works in terms of both objective and subjective quality evaluations. Specifically, our method is more effective on edge areas and more robust to noises. Jin Wang 0023, Longhua Sun, Ruiqin Xiong, Yunhui Shi, Qing Zhu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Depth Map Super-Resolution via Joint Local Gradient and Nonlocal Structural RegularizationsabstractDepth maps have been widely used in many real world applications, such as human-computer interaction and virtual reality. However, due to the limitation of current depth sensing technology, the captured depth maps usually suffer from low resolution and insufficient quality. In this paper, we propose a depth map super-resolution method via joint local gradient and nonlocal structural regularizations. Depth maps contain mainly smooth areas separated by textures which demonstrate distinct geometry direction characteristic. Motivated by this, we classify depth map patches according to their geometrical directions and learn a compact online dictionary in each class. We further introduce two regularization terms into the sparse representation framework. Firstly, a multi-directional total variation model is proposed to characterize the local patterns in the gradient domain. Secondly, a nonlocal autoregressive model is introduced to provide nonlocal constraint to the local structures, which can effectively restore image details and suppress noise. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior performance for various configurations of magnification factors and datasets. Wei Xu 0059, Qing Zhu 0004, Na Qi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Sparse Representation Based Image Restoration With Denoising PriorabstractAs a powerful statistical signal modeling technique, sparse representation has been widely used in various image restoration (IR) applications. The sparsity-based methods have achieved leading performance in the past few decades. However, in recent years it has been surpassed by other methods, especially the recent deep learning based methods. In this paper, we address the question that whether sparse representation can be competitive again. The way we answer this question is to redesign it with a deep architecture. To be specific, we propose an end-to-end deep architecture that follows the process of the sparse representation based IR. In particular, we learn a sparse convolutional dictionary to replace the traditional dictionary, and a convolutional neural network (CNN) denoising prior to replace the image prior. Through end-to-end training, the parameters in convolutional dictionary and CNN denoiser can be jointly optimized. Experimental results on several representative IR tasks, including image denoising, deblurring and super-resolution, demonstrate that the proposed deep network can achieve superior performance against state-of-the-art model-based and learning-based methods. Wei Xu 0059, Qing Zhu 0004, Na Qi, Dongpan Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Depth Map Super-Resolution By Multi-Direction Dictionary And Joint RegularizationabstractDepth maps acquired by 3D cameras usually suffer from low resolution and insufficient quality, which makes it difficult to be directly used in visual depth perception and 3D reconstruction. To handle this problem, we propose a novel multi-direction dictionary and joint regularization model for high quality depth recovery. To enhance the sparsity, image is divided into classified patches according to the same geometrical direction and a compact dictionary is trained within each class. Then for a patch to be coded, the most relevant dictionary can be selected according to its geometrical direction. We further introduce two regularization terms into the reconstruction model. One is the anisotropic total variation (TV) defined by the local gradients of color/depth pair. The other is the nonlocal similarity to provide nonlocal constraint to the local structure. Experimental results demonstrate that our method outperforms other state-of-the-art methods in terms of both subjective quality and objective quality. Wei Xu 0059, Jin Wang 0023, Longhua Sun, Qing Zhu 0004 |
ICIP | 4 |
| 2021 | Dual Regularization Based Depth Map Super-Resolution with Graph Laplacian PriorabstractThe edge information plays key role in the restoration of depth map. Most conventional methods assume that the RGB-D pairs are consistent in edge areas. In this paper, firstly, we point out that in most cases the consistency between normal map and depth map(N-D pairs) are much higher than that be-tween RGB-D pairs. Then we propose a dual regularization term to guide the restoration of depth map, which constrains the consistency between N-D pairs back and forth. Moreover, a reweighted graph Laplacian prior is incorporated into a unified optimization framework to effectively protect piece-wise smoothness(PWS) characteristics of depth map. By treating depth maps as graph signals, the weight between two nodes is adapted according to its content. Extensive experimental results demonstrate the superior performance of our method compared with other state-of-the-art works in terms of objective and subjective quality evaluations. Longhua Sun, Jin Wang 0023, Ruiqin Xiong, Yunhui Shi, Qing Zhu 0004 |
ICME | 5 |
| 2020 | Light Field Image Compression Using Multi-branch Spatial Transformer Networks Based View SynthesisabstractThe recent years have witnessed the widespread of light field imaging in interactive and immersive visual applications. To record the directional information of the light rays, larger storage space is required by light field images compared with conventional 2D images. Hence, the efficient compression of light field image is highly desired for further applications. In this paper, we propose a novel light field image compression scheme using multi-branch spatial transformer networks based view synthesis. Firstly, a sparse subset of views are selected and are rearranged into a pseudo sequence to be encoded by a video codec at encoder. Then the other unselected views are synthesized based on the similarity between neighboring views with our proposed method at decoder. To better characterize the non-linear relationship between the sub-views, a multi-branch spatial transformer networks (MSTN) is designed to adaptively learn the affine transformations between the neighboring views, which are used to warp the input views to generate accurate approximation of the target views. Moreover, to better obtain the final view by the generated approximation views, the Wasserstein generative adversarial networks(WGAN) is applied with the improved training. Experimental results show the superior compression performance of our scheme compared with the state-of-the-art methods. Jin Wang 0023, Ruiqin Xiong, Qing Zhu 0004 |
DCC | 4 |
| 2020 | Tensor-Based Light Field Denoising By Exploiting Non-Local Similarities Across Multiple ResolutionsabstractLight field is a kind of 4D signal that contains rich information about position and angle of light, which can express the scene more accurately. Light field is easily affected by noise for the hardware sensitivity. This paper utilizes the intrinsic tensor sparsity model and integrates super-resolution(SR) into a unified light field denoising method based on tensor operation. Avoiding vectorization, we make full use of correlation of light field. By exploiting SR method, we avoid sub-pixel mis-alignment in the searching process of similar patch. Experimental results validate that our proposed method outperforms the state-of-art methods in terms of both objective and subjective quality on the HCI light field old dataset. Na Qi, Qing Zhu 0004 |
ICIP | 3 |
| 2020 | See Through Occlusions: Detailed Human Shape Estimation From A Single Image With Occlusionsabstract3D human body shape and pose reconstructing from a single RGB image is a challenging task in the field of computer vision and computer graphics. Since occlusions are prevalent in real application scenarios, it's important to develop 3D human body reconstruction algorithms with occlusions. However, existing methods didn't take this problem into account. In this paper, we present a novel depth estimation Neural Network, named Detailed Human Depth Network(DHDNet), which aims to reconstruct the detailed and completed depth map from a single RGB image contains occlusions of human body. Inspired by the previous works [1], [2], we propose an end-to-end method to obtain the fine detailed 3D human mesh. The proposed method follows a coarse-to-fine refinement scheme. Using the depth information generated from DHDNet, the coarse 3D mesh can recover detailed spatial structure, even the part behind occlusions. We also construct DepthHuman, a 2D in-the-wild human dataset containing over 18000 synthetic human depth maps and corresponding RGB images. Extensive experimental results demonstrate that our approach has significant improvement in 3D mesh reconstruction accuracy on the occluded parts. Jin Wang 0023, Qing Zhu 0004 |
ICIP | 3 |
| 2020 | Generative Image Inpainting Based on Wavelet Transform Attention ModelabstractImage inpainting is a challenging task in image processing and widely applied in many areas such as photo editing. Traditional patch-based methods are not effective to deal with complex or non-repetitive structures. Recently, deep learning-based approaches have shown promising results for image inpainting. However, they usually generate contents with artificial boundaries, distorted structures or blurry textures. To handle this problem, we propose a novel image inpainting method based on wavelet transform attention model (WTAM). The wavelet transform decomposes features into multi-frequency sub-bands for extracting and transmitting deep information, and the attention mechanism enhances the ability of wavelet transform to capture significant detailed information in each level's subband images. Extensive experimental results on multiple datasets (Paris StreetView, CelebA and CelebAMask-HQ) demonstrate that our method can not only synthesize sharp image structures but also generate fine-detailed textures in missing regions, significantly outperforming the state-of-the-art methods. Jin Wang 0023, Qing Zhu 0004 |
ISCAS | 3 |
| 2020 | Image Inpainting Based on Multi-frequency Probabilistic Inference ModelabstractImage inpainting methods usually fail to reconstruct reasonable structure and fine-grained texture simultaneously. This paper handles this problem from a novel perspective of predicting low-frequency semantic structural contents and high-frequency detailed textures respectively, and proposes a multi-frequency probabilistic inference model(MPI model) to predict the multi-frequency information of missing regions by estimating the parametric distribution of multi-frequency features over the corresponding latent spaces. Firstly, in order to extract the information of different frequencies without any interference, wavelet transform is utilized to decompose the input image into low-frequency subband and high-frequency subbands. Furthermore, an MPI model is designed to estimate the underlying multi-frequency distribution of input images. With this model, closer approximation to the true posterior distribution can be constrained and maximum-likelihood assignment can be approximated. Finally, based on the proposed MPI model, a two-path network consisting of inference network(InferenceNet) and generation network(GenerationNet) is trained parallelly to enforce the consistency of global structure and local texture between the generated image and ground truth. We qualitatively and quantitatively compare our method with other state-of-the-art methods on Paris StreetView, CelebA, CelebAMask-HQ and Places2 datasets. The results show the superior performance of our method, especially in the aspects of realistic texture details and semantic structural consistency. Jin Wang 0023, Qingming Huang, Yunhui Shi, Jian-Feng Cai 0001, Qing Zhu 0004 |
ACM Multimedia | 6 |
| 2020 | Transfer non-stationary texture with complex appearanceabstractTexture transfer has been successfully applied in computer vision and computer graphics. Since non-stationary textures are usually complex and anisotropic, it is challenging to transfer these textures by simple supervised method. In this paper, we propose a general solution for non-stationary texture transfer, which can preserve the local structure and visual richness of textures. The inputs of our framework are source texture and semantic annotation pair. We record different semantics as different regions and obtain the color and distribution information from different regions, which is used to guide the the low-level texture transfer algorithm. Specifically, we exploit these local distributions to regularize the texture transfer objective function, which is minimized by iterative search and voting steps. In the search step, we search the nearest neighbor fields of source image to target image through Generalized PatchMatch (GPM) algorithm. In the voting step, we calculate histogram weights and coherence weights for different semantic regions to ensure color accuracy and texture continuity, and to further transfer the textures from the source to the target. By comparing with state-of-the-art algorithms, we demonstrate the effectiveness and superiority of our technique in various non-stationary textures. Na Qi, Qing Zhu 0004 |
MMAsia | 3 |
| 2020 | Two-stage structure aware image inpainting based on generative adversarial networksabstractIn recent years, the image inpainting technology based on deep learning has made remarkable progress, which can better complete the complex image inpainting task compared with traditional methods. However, most of the existing methods can not generate reasonable structure and fine texture details at the same time. To solve this problem, in this paper we propose a two-stage image inpainting method with structure awareness based on Generative Adversarial Networks, which divides the inpainting process into two sub tasks, namely, image structure generation and image content generation. In the former stage, the network generates the structural information of the missing area; while in the latter stage, the network uses this structural information as a prior, and combines the existing texture and color information to complete the image. Extensive experiments are conducted to evaluate the performance of our proposed method on Places2, CelebA and Paris Streetview datasets. The experimental results show the superior performance of the proposed method compared with other state-of-the-art methods qualitatively and quantitatively. Jin Wang 0023, Qing Zhu 0004 |
MMAsia | 4 |
| 2020 | UGNet: Underexposed Images Enhancement Network based on Global Illumination EstimationabstractThis paper proposes a new neural network for enhancing underexposed images. Instead of the decomposition method based on Retinex theory, we introduce smooth dilated convolution to estimate global illumination of the input image, and implement an end-to-end learning network model. Based on this model, we formulate a multi-term loss function that combines content, color, texture and smoothness losses. Our extensive experiments demonstrate that this method is superior to other methods in underexposed image enhancement. It can cover more color details and be applied to various underexposed images robustly. Wenzhe Zhu, Qing Zhu 0004 |
VCIP | 3 |
| 2020 | HDR Image Compression with Convolutional AutoencoderabstractAs one of the next-generation multimedia technology, high dynamic range (HDR) imaging technology has been widely applied. Due to its wider color range, HDR image brings greater compression and storage burden compared with traditional LDR image. To solve this problem, in this paper, a two-layer HDR image compression framework based on convolutional neural networks is proposed. The framework is composed of a base layer which provides backward compatibility with the standard JPEG, and an extension layer based on a convolutional variational autoencoder neural networks and a post-processing module. The autoencoder mainly includes a nonlinear transform encoder, a binarized quantizer and a nonlinear transform decoder. Compared with traditional codecs, the proposed CNN autoencoder is more flexible and can retain more image semantic information, which will improve the quality of decoded HDR image. Moreover, to reduce the compression artifacts and noise of reconstructed HDR image, a post-processing method based on group convolutional neural networks is designed. Experimental results show that our method outperforms JPEG XT profile A, B, C and other methods in terms of HDR-VDP-2 evaluation metric. Meanwhile, our scheme also provides backward compatibility with the standard JPEG. Jin Wang 0023, Ruiqin Xiong, Qing Zhu 0004 |
VCIP | 4 |
| 2020 | Icon Colorization Based On Triple Conditional Generative Adversarial NetworksabstractCurrent automatic colorization systems have many defects such as "contour blur", "color overflow"and "color miscellaneous", especially when they are coloring the images with hollowed-out structure. We propose a model based on triple conditional generative adversarial networks, for generator we provide contour image, colored icon and colorization mask as inputs, our network has three discriminators, structure discriminator is trained to judge if the generated icon has similar contour to the input icon, color discriminator anticipates generated icon and the input icon has the similar color style, the function of mask discriminator is to distinguish whether the output has the similar colorization area to the input mask. For the evaluation, we compared with some existing colorization models, also we made a questionnaire to obtain the evaluation of generated icons from different models. The results showed that our colorization model obtain better results comparing to the other models both in generating hollowed-out and solid structure icons. Qinru Han, Wenzhe Zhu, Qing Zhu 0004 |
VCIP | 3 |
| 2020 | Generative image completion with image-to-image translation
Shuzhen Xu, Qing Zhu 0004, Jin Wang 0023 |
Neural Comput. Appl. | 2 |
| 2020 | Correction to: Generative image completion with image-to-image translation
Shuzhen Xu, Qing Zhu 0004, Jin Wang 0023 |
Neural Comput. Appl. | 2 |
| 2020 | Multi-Direction Dictionary Learning Based Depth Map Super-Resolution With Autoregressive Modelingabstract3D depth cameras have become more and more popular in recent years. However, depth maps captured by these cameras can hardly be used in 3D reconstruction directly because they often suffer from low resolution and blurring depth discontinuities. Super resolution of depth maps is necessary. In depth maps, the edge areas play more important role and demonstrate distinct geometry directions compared with natural images. However, most existing super-resolution methods ignore this fact, and they can not handle depth edges properly. Motivated by this, we propose a compound method that combines multi-direction dictionary sparse representation and autoregressive (AR) models, so that the depth edges are presented precisely at different levels. In the patch level, the depth edge patches with geometry directions are well represented by the pre-trained multi-directional dictionaries. Compared with a universal dictionary, multiple dictionaries trained from different directional patches can represent the directional depth patch much better. In the finer pixel level, we utilize an adaptive AR model to represent the local correlation patterns in small areas. Extensive experimental results on both synthetic and real datasets demonstrate that, the proposed model outperforms state-of-the-art depth map super-resolution methods in terms of both quantitative metrics and subjective visual quality. Jin Wang 0023, Wei Xu 0059, Jian-Feng Cai 0001, Qing Zhu 0004, Yunhui Shi |
IEEE Trans. Multim. | 4 |
| 2019 | Surface Normal Data Guided Depth Recovery with Graph Laplacian RegularizationabstractHigh-quality depth information has been increasingly used in many real-world multimedia applications in recent years. Due to the limitation of depth sensor and sensing technology, actually, the captured depth map usually has low resolution and black holes. In this paper, inspired by the geometric relationship between surface normal of a 3D scene and their distance from camera, we discover that surface normal map can provide more spatial geometric constraints for depth map reconstruction, as depth map is a special image with spatial information, which we called 2.5D image. To exploit this property, we propose a novel surface normal data guided depth recovery method, which uses surface normal data and observed depth value to estimate missing or interpolated depth values. Moreover, to preserve the inherent piecewise smooth characteristic of depth maps, graph Laplacian prior is applied to regularize the inverse problem of depth maps recovery and a graph Laplacian regularizer(GLR) is proposed. Finally, the spatial geometric constraint and graph Laplacian regularization are integrated into a unified optimization framework, which can be efficiently solved by conjugate gradient(CG). Extensive quantitative and qualitative evaluations compared with state-of-the-art schemes show the effectiveness and superiority of our method. Longhua Sun, Jin Wang 0023, Yunhui Shi, Qing Zhu 0004 |
MMAsia | 4 |
| 2019 | CR-U-Net: Cascaded U-Net with Residual Mapping for Liver Segmentation in CT Images*abstractAbdominal computed tomography (CT) is a common modality to detect liver lesions. Liver segmentation in CT scan is important for diagnosis and analysis of liver lesions. However, the accuracy of existing liver segmentation methods is slightly insufficient. In this paper, we propose a liver segmentation architecture named CR-U-Net, which is composed of cascade U-Net combined with residual mapping. We make use of the MDice loss function for training in CR-U-Net, and the second-level of cascade network is deeper than the first-level to extract more detailed image features. Morphological algorithms are utilized as an intermediate-processing step to improve the segmentation accuracy. In addition, we evaluate our proposed CR-U-Net on liver segmentation task under the dataset provided by the 2017 ISBI LiTS Challenge. The experimental result demonstrates that our proposed CR-U-Net can outperform the state-of-the-art methods in term of the performance measures, such as Dice score, VOE, and so on. Na Qi, Qing Zhu 0004 |
VCIP | 3 |
| 2018 | High Dynamic Range Image Compression Based on Visual SaliencyabstractHigh dynamic range (HDR) image has larger luminance range than conventional low dynamic range (LDR) image, which is more consistent with human visual system (HVS). Recently, JPEG committee releases a new HDR image compression standard JPEG XT. It decomposes input HDR image into base layer and extension layer. However, this method doesn't make full use of HVS, causing waste of bits on imperceptible regions to human eyes. In this paper, a visual saliency based HDR image compression scheme is proposed. The saliency map of tone mapped HDR image is first extracted, then is used to guide extension layer encoding. The compression quality is adaptive to the saliency of the coding region of the image. Extensive experimental results show that our method outperforms JPEG XT profile A, B, C, and offers the JPEG compatibility at the same time. Moreover, our method can provide progressive coding of extension layer. Shenda Li, Jin Wang 0023, Qing Zhu 0004 |
PCS | 3 |
| 2018 | Compressively Sensed Multi-View Image Reconstruction Using Joint Optimization ModelingabstractUtilizing both intra and inter views correlation plays a key role to improve compressive sensing reconstruction of multi-view images. For this goal, this paper presents a joint optimization model (JOM) for compressively-sensed multi-view image reconstruction, which jointly optimizes an adaptive disparity compensated residual total variation (ARTV) and a multi-image nonlocal low-rank tensor (MNLRT). To exploit the inter-view correlation efficiently, the ARTV method adaptively forms suitable dynamic image set to help reconstruct the current one. Different from previous work, the MNLRT regularization uses tensor rather than 2D matrix to exploit nonlocal low-rank property, which keeps intrinsic geometrical structures of image patches. An efficient algorithm is further proposed to solve the joint optimization problem via Split-Bregman based technique. Extensive experimental results demonstrate our method outperforms state-of-the-arts algorithms with almost 1.5 dB gain in terms of PSNR, while obtaining dramatically improved visual quality for edge area, especially at low sampling rates. Jin Wang 0023, Qing Zhu 0004 |
VCIP | 3 |
| 2017 | Depth map super-resolution via multiclass dictionary learning with geometrical directionsabstractDepth cameras have gained significant popularity due to their affordable cost in recent years. However, the resolution of depth map captured by these cameras is rather limited, and thus it hardly can be directly used in visual depth perception and 3D reconstruction. In order to handle this problem, we propose a novel multiclass dictionary learning method, in which depth image is divided into classified patches according to their geometrical directions and a sparse dictionary is trained within each class. Different from previous SR works, we build the correspondence between training samples and their corresponding register color image via sparse representation. We further use the adaptive autoregressive model as a reconstruction constraint to preserve smooth regions and sharp edges. Experimental results demonstrate that our method outperforms state-of-the-art methods in depth map super-resolution in terms of both subjective quality and objective quality. Wei Xu 0059, Jin Wang 0023, Qing Zhu 0004 |
VCIP | 3 |
| 2015 | An Effective Method for Gender Classification with Convolutional Neural Networks
Qing Zhu 0004, Xiaoqi Jia |
ICA3PP (2) | 2 |
| 2013 | An Efficient Projective Rectification for Trinocular Stereo VisionabstractIn this paper, we describe an efficient projective rectification method that can be applied to trinocular rectification to improve the visual appearance of stereo vision. This method expends the matrix-based correcting homographic approach which has been proposed recently. In order to implement trinocular rectification, firstly, we obtain triplet images from the target objects through three cameras. The triplet images can be treated as two pairs and the homographies can project different image planes onto new planes. As a result, the rectification becomes the problems of seeking the common plane, and applying transformation or homography to images. Finally, through the proposed solution the trinocular rectification is accomplished. At the same time, Distortion reduction makes the rectified images more realistic. Qing Zhu 0004 |
PDCAT | 1 |