VLDB 2026 Research / reviewers in the wild / expert
Zhuo Su 0001
dblp:02/10578-1
· DBLP profile ↗
70ranked-venue papers
17as first author
34since 2021 · last 2026
0000-0002-6090-0110ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 13 first-author · 26 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NEGS-Avatar: Normal Embedded Gaussians for 2D avatar from monocular video
Zedan Zheng, Yudi Tan, Zhuo Su 0001, Fan Zhou 0001, Baoquan Zhao |
Comput. Graph. | 3 |
| 2026 | Action-aware anchor-based frame selection strategy for action recognition
Fan Zhou 0001, Ge Lin 0002, Zhuo Su 0001 |
Image Vis. Comput. | 5 |
| 2026 | PDDA: Prompt-Driven Domain Adaptation for Real-World Image DehazingabstractDue to the complexity and diversity of practical environments, real-world image dehazing remains an unresolved problem, with one of the key challenges being how to bridge the distribution gap between synthetic and real domains. This paper proposes a Prompt-driven Domain Adaptation (PDDA) framework within the bi-level optimization perspective. Specifically, we introduce hyperparameter optimization-based bi-level modeling: the lower-level optimization emphasizes prior learning within the synthetic domain to stabilize dehazing performance, while the upper-level optimization focuses on enhancing cross-domain adaptability to ensure that the model can generalize across different domains. Given the scarcity of paired real haze images, we train learnable haze prompts by jointly optimizing the text-image similarity between positive/negative prompts and corresponding clear/haze images in the CLIP latent space to more effectively capture real-world haze characteristics. Based on the learned haze prompts, we construct an unsupervised cross-domain loss function that enhances the adaptability to complex real-world scenarios by integrating prompt learning with bi-level optimization strategy. Furthermore, we conduct a comprehensive exploration to uncover the inherent properties of PDDA, including architecture-irrelevant flexibility and domain-agnostic robustness. Extensive experiments across a wide range of benchmark datasets demonstrate that our method achieves both quantitative and qualitative improvements across diverse scenarios, showing robust performance not only in real-world daytime conditions but also exhibiting superior cross-domain adaptation capabilities in nighttime scenarios. Codes are available at https://github.com/YanZhang-zy/PDDA.git. Yan Zhang 0002, Xin Li 0175, Fan Zhou 0001, Zhuo Su 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | CoA: Towards Real Image Dehazing via Compression-and-AdaptationabstractLearning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both efficiency and adaptability to address real image dehazing effectively. This work proposes a Compression-and-Adaptation (CoA) computational flow to tackle these challenges from a divide-and-conquer perspective. First, model compression is performed in the synthetic domain to develop a compact dehazing parameter space, satisfying efficiency demands. Then, a bilevel adaptation in the real domain is introduced to be fearless in unknown real environments by aggregating the synthetic dehazing capabilities during the learning process. Leveraging a succinct design free from additional constraints, our CoA exhibits domain-irrelevant stability and model-agnostic flexibility, effectively bridging the model chasm between synthetic and real domains to further improve its practical utility. Extensive evaluations and analyses underscore the approach's superiority and effectiveness. The code is publicly available at https://github.com/fyxnl/COA. Long Ma 0002, Yan Zhang 0002, Jinyuan Liu 0001, Weimin Wang 0007, Guang-Yong Chen, Chengpei Xu, Zhuo Su 0001 |
CVPR | 8 |
| 2025 | MCSMoG: Multi-Conditional Diffusion for Stylized Motion Generation with Parametric ControlabstractStylized human motion synthesis remains a fundamental challenge in computer animation and graphics, with a wide spectrum of applications spanning gaming, film production, virtual reality, and beyond. While recent advances in text-driven motion generation have shown promise, existing approaches face critical limitations including the inability to maintain consistent trajectory control, the lack of fine-grained stylization intensity adjustment, and inadequate generalization across diverse motion styles. To address these challenges, We introduce MCSMoG, a novel framework for controllable stylized motion synthesis through multi-conditional guidance. First, a new Multi-Conditional Motion Latent Diffusion (MC-MLD) model is proposed to introduce additional trajectory guidance and achieve trajectory decoupling. Second, we develop a Style and Non-Style Feature Fusion Module that dynamically blends motion features through an adjustable parameter, providing control over stylization intensity. Third, we integrate MotionCLIP as our style encoder, enhancing the model’s generalization capability across diverse and unseen motion styles. Extensive experiments conducted on the combined HumanML3D and 100STYLE datasets demonstrate that our approach outperforms state-of-the-art methods, achieving a 4.6% reduction in FID scores and a 4.1% increase in motion diversity. User studies further confirm the superiority of our method in style fidelity, semantic consistency, and motion naturalness. Xinzhu Li, Guanghui Yue 0001, Wei Zhou 0021, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ICME | 6 |
| 2025 | Breaking the Synthetic Barrier: Towards Stable and Generalizable Real-World Image DehazingabstractExisting learning-based dehazing methods perform well on synthetic data but struggle in real scenarios due to the domain gap, causing residual haze and detail loss. To address this, we propose a Multilevel Subspace Distribution Adapter (MSDA) to progressively reduce the feature distribution gap through hierarchical subspace modeling. We also introduce a Dual-Domain Synchronous Optimization (DDSO) strategy that jointly leverages synthetic supervision and adaptation to the real domain in a unified training scheme. Extensive experiments underscore the superiority of our approach and its excellence on no-reference image quality metrics. Zhuo Su 0001, Jufeng Li, Yan Zhang 0002, Xin Li 0175, Fan Zhou 0001 |
ACM Multimedia | 1 |
| 2025 | Part-aware distillation and aggregation network for human parsing
Yuntian Lai, Fan Zhou 0001, Zhuo Su 0001 |
Image Vis. Comput. | 4 |
| 2025 | Single-View Clothed Human Reconstruction With Multi-View Consistency RepresentationabstractFor single-view clothed human reconstruction, the fashionable PIFu-like framework depends on the pixel-aligned feature essentially, while this leads to depth ambiguity and inaccuracy of representation. Additionally, this task faces the inherent problem of the lack of invisible information. To solve these two problems, we propose depth-guided pixel-aligned feature and multi-view consistency prior to constrain representation learning of the single-view reconstruction task. The difference, between the depth values of points and estimated depth map, is used to filter pixel-aligned features. Thus, the image encoder can focus on capturing the feature of the visible part which is more effective in feature representation. The method introduces contrastive learning and masked autoencoder to achieve consistency of SMPL vertex features in each view which helps model to imagine invisible information. The experimental results show that the proposed method enables the feature to represent surface details more efficiently, thus achieves more reasonable and accurate representation learning. The qualitative and quantitative evaluations on the public and commercial datasets show that the proposed method can achieve better performance than previous implicit representation based methods. Zhuo Su 0001, Yudi Tan, Zedan Zheng, Fan Zhou 0001, Baoquan Zhao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Visual Boundary-Guided Pseudo-Labeling for Weakly Supervised 3D Point Cloud Segmentation in Indoor EnvironmentsabstractAccurate segmentation of 3D point clouds in indoor scenes remains a challenging task, often hindered by the labor-intensive nature of data annotation. While weakly supervised learning approaches have shown promise in leveraging partial annotations, they frequently struggle with imbalanced performance between foreground and background elements due to the complex structures and proximity of objects in indoor environments. To address this issue, we propose a novel foreground-aware label enhancement method utilizing visual boundary priors. Our approach projects 3D point clouds onto 2D planes and applies 2D image segmentation to generate pseudo-labels for foreground objects. These labels are subsequently back-projected into 3D space and used to train an initial segmentation model. We further refine this process by incorporating prior knowledge from projected images to filter the predicted labels, followed by model retraining. We introduce this technique as the Foreground Boundary Prior (FBP), a versatile, plug-and-play module designed to enhance various weakly supervised point cloud segmentation methods. We demonstrate the efficacy of our approach on the widely-used 2D-3D-Semantic dataset, employing both random-sample and bounding-box based weak labeling strategies. Our experimental results show significant improvements in segmentation performance across different architectural backbones, highlighting the method's effectiveness and portability. Zhuo Su 0001, Yudi Tan, Boliang Guan, Fan Zhou 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns
Xin Li 0175, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001 |
IJCAI | 5 |
| 2024 | Utilizing Text-Video Relationships: A Text-Driven Multi-modal Fusion Framework for Moment Retrieval and Highlight Detection
Ruomei Wang 0001, Zhuo Su 0001 |
PRCV (10) | 4 |
| 2024 | Video object segmentation via couple streams and feature memoryabstractAbstract In recent years, most video segmentation methods use deep CNN to process the input image, but they did not fully mine the rich intermediate predictions in spatio‐temporal space. And, the segmentation challenges such as occlusion, severe deformation and illumination have not been well solved so far. To alleviate these problems, this paper focuses on constructing multi module network structures that represent multi semantics and proposes a video object segmentation network via coupled‐stream architecture with feature memory mechanism. This network first extracts high‐level semantic features, edge features, long‐term and short‐term stable depth features of the target, and then decode them into the segmentation mask of target. In addition, negative skeleton inhibition and frame interpolation are used to prevent the interference of similar objects and motion blur, respectively. The method has a low GPU memory usage, regardless of the number of object in video. And performs 86.5%and 62.4% in J&F measure on DAVIS 2016 and DAVIS 2017 validation set, without fine‐tuning and online training. Yun Liang 0003, Xinjie Xiao, Shaojian Qiu, Zhuo Su 0001 |
IET Image Process. | 5 |
| 2024 | Advancing Real-World Image Dehazing: Perspective, Modules, and TrainingabstractRestoring high-quality images from degraded hazy observations is a fundamental and essential task in the field of computer vision. While deep models have achieved significant success with synthetic data, their effectiveness in real-world scenarios remains uncertain. To improve adaptability in real-world environments, we construct an entirely new computational framework by making efforts from three key aspects: imaging perspective, structural modules, and training strategies. To simulate the often-overlooked multiple degradation attributes found in real-world hazy images, we develop a new hazy imaging model that encapsulates multiple degraded factors, assisting in bridging the domain gap between synthetic and real-world image spaces. In contrast to existing approaches that primarily address the inverse imaging process, we design a new dehazing network following the "localization-and-removal" pipeline. The degradation localization module aims to assist in network capture discriminative haze-related feature information, and the degradation removal module focuses on eliminating dependencies between features by learning a weighting matrix of training samples, thereby avoiding spurious correlations of extracted features in existing deep methods. We also define a new Gaussian perceptual contrastive loss to further constrain the network to update in the direction of the natural dehazing. Regarding multiple full/no-reference image quality indicators and subjective visual effects on challenging RTTS, URHI, and Fattal real hazy datasets, the proposed method has superior performance and is better than the current state-of-the-art methods. Long Ma 0002, Xiaozhe Meng, Fan Zhou 0001, Risheng Liu, Zhuo Su 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Bridging the Gap Between Haze Scenarios: A Unified Image Dehazing ModelabstractIn real-world scenarios, the haze presents diversity and complexity. However, current dehazing researches usually focus solely on specific categories or the removal of common white haze, frequently lacking the ability to adapt across various unknown haze types. In this study, our emphasis is on constructing a model that shows excellent adaptability across diverse haze conditions. Unlike approaches that solely rely on network structure design to enhance model adaptability, we comprehensively improve dehazing model adaptability from three key aspects: constructing the multitype haze dataset from designed haze degradation models, designing the network architecture, and formulating training strategies suitable for cross-scene generalization. Firstly, to meet the diverse haze training data requirements, we design a multitype haze degradation model to generate more realistic pairs of hazy images. Secondly, to ensure thorough haze removal and natural restoration of texture details in the recovered images, we construct a dual-branch ensemble network framework by leveraging pre-trained clear image prior features and the characteristics of 2D discrete wavelet priors. Finally, to further enhance the adaptability for removing various types of haze, we employ a sample reweighting decorrelation strategy during the network training phase to eliminate dependencies between haze and haze-free background features. Through extensive experiments, our approach shows remarkable performance across diverse haze scenarios. Our method not only outperforms state-of-the-art scene-specific dehazing methods in typical scenarios like daytime and nighttime, but it also excels in handling challenging scenarios such as dusty conditions, and color haze. See more resultshttps://github.com/fyxnl/Image-dehazing-CGID. Zhuo Su 0001, Long Ma 0002, Xin Li 0175, Risheng Liu, Fan Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Subtask Prior-Driven Optimized Mechanism on Joint Video Moment Retrieval and Highlight DetectionabstractJoint video moment retrieval and highlight detection is an emerging and challenging research task. It requires the generation of robust joint task features to satisfy the demands of video moment retrieval and video highlight detection. Moreover, it involves the interaction of multiple modalities. Presently, methods typically focus on the design of distinct enhancement modules and the addition of supplementary input data to improve the solution for joint video moment retrieval and highlight detection. However, they overlook subtask interference during joint training. Joint task learning leverages the correlations and complementarities between tasks, yet it also introduces task interference arising from the differences between tasks. In order to address task interference, we proposes a subtask prior-driven optimized mechanism. The mechanism consists of two stages. In the free stage, we train subtask model to get subtask prior features. In the constrained stage, the joint task model is constrained by the subtask. Besides, we propose a cross adaptive-gated mechanism. It addresses the issue of information loss in cross-modal fusion and filters out redundant information by conducting cross-modal interaction during feature compression and an adaptive gating process. Extensive experimental results exhibit the effectiveness of the subtask prior-driven optimized mechanism and the cross adaptive-gated transformer in joint video moment retrieval and highlight detection. Ruomei Wang 0001, Fan Zhou 0001, Zhuo Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | DMAP: Decoupling-Driven Multi-Level Attribute Parsing for Interpretable Outfit CollocationabstractOutfit collocation requires considering the interrelationship and adaptability among the attributes of component items. However, with the numerous and diverse attributes of fashion items, accurately capturing attribute features and modeling the complex relationships between attributes become the key challenges. To address these challenges, we propose a novel scheme Decoupling-driven Multi-level Attribute Parsing for interpretable outfit collocation. First, we decouple a series of attribute features from the item's visual feature by fully supervised, which can improve the robustness of the model in processing both relevant and irrelevant attributes of items. Furthermore, employing a deep deconvolution neural network with attention mechanisms to reconstruct the decoupled attribute features into a visual image that is close to the original item image. It ensures all attribute features can be combined to contain complete item information. Next, graph attention networks are constructed to parse multi-level attribute compatibility relationships from three perspectives: intra-attribute, inter-attribute, and item integration relationships. Finally, we use multi-layer perceptrons to fuse the score distributions of the three and output the outfit compatibility score. Experiments conducted on the IQON3000 dataset demonstrate that our model outperforms existing state-of-the-art methods and exhibits good interpretability. Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001, Ge Lin 0002 |
IEEE Trans. Multim. | 1 |
| 2023 | High Fidelity Virtual Try-On via Dual Branch Bottleneck Transformer
Xiuxiang Li, Guifeng Zheng, Fan Zhou 0001, Zhuo Su 0001, Ge Lin 0002 |
ICIG (1) | 4 |
| 2023 | Feature Representation for High-resolution Clothed Human ReconstructionabstractAbstract Detailed and accurate feature representation is essential for high‐resolution reconstruction of clothed human. Herein we introduce a unified feature representation for clothed human reconstruction, which can adapt to changeable posture and various clothing details. The whole method can be divided into two parts: the human shape feature representation and the details feature representation. Specifically, we firstly combine the voxel feature learned from semantic voxel with the pixel feature from input image as an implicit representation for human shape. Then, the details feature mixed with the clothed layer feature and the normal feature is used to guide the multi‐layer perceptron to capture geometric surface details. The key difference from existing methods is that we use the clothing semantics to infer clothed layer information, and further restore the layer details with geometric height. We qualitative and quantitative experience results demonstrate that proposed method outperforms existing methods in terms of handling limb swing and clothing details. Our method provides a new solution for clothed human reconstruction with high‐resolution details (style, wrinkles and clothed layers), and has good potential in three‐dimensional virtual try‐on and digital characters. Juncheng Pu, Li Liu 0032, Xiaodong Fu, Zhuo Su 0001, Wei Peng 0004 |
Comput. Graph. Forum | 4 |
| 2023 | OaIF: Occlusion-Aware Implicit Function for Clothed Human Re-constructionabstractAbstract Clothed human re‐construction from a monocular image is challenging due to occlusion, depth‐ambiguity and variations of body poses. Recently, shape representation based on an implicit function, compared to explicit representation such as mesh and voxel, is more capable with complex topology of clothed human. This is mainly achieved by using pixel‐aligned features, facilitating implicit function to capture local details. But such methods utilize an identical feature map for all sampled points to get local features, making their models occlusion‐agnostic in the encoding stage. The decoder, as implicit function, only maps features and does not take occlusion into account explicitly. Thus, these methods fail to generalize well in poses with severe self‐occlusion. To address this, we present OaIF to encode local features conditioned in visibility of SMPL vertices. OaIF projects SMPL vertices onto image plane to obtain image features masked by visibility. Vertices features integrated with geometry information of mesh are then feed into a GAT network to encode jointly. We query hybrid features and occlusion factors for points through cross attention and learn occupancy fields for clothed human. The experiments demonstrate that OaIF achieves more robust and accurate re‐construction than the state of the art on both public datasets and wild images. Yudi Tan, Boliang Guan, Fan Zhou 0001, Zhuo Su 0001 |
Comput. Graph. Forum | 4 |
| 2023 | Boundary-guided part reasoning network for human parsing
Zhuo Su 0001, Huiqiang Guan, Yuntian Lai, Fan Zhou 0001, Yun Liang 0003 |
Neurocomputing | 1 |
| 2023 | Towards real-world haze removal with uncorrelated graph model
Xiaozhe Meng, Fan Zhou 0001, Yun Liang 0003, Zhuo Su 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Attention-adaptive multi-scale feature aggregation dehazing network
Zhuo Su 0001, Fan Zhou 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Transfer Learning Dehazing Network by Gaussian Process Mapping
Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001 |
Signal Process. Image Commun. | 3 |
| 2023 | Real-World Non-Homogeneous Haze Removal by Sliding Self-Attention Wavelet NetworkabstractIn complex natural haze scenes, image haze removal still faces significant challenges in removing non-homogeneous and dense haze. The double complexity of haze distribution, on the one hand, is reflected in the interference of haze to the global image information, and on the other hand, it is reflected in the imbalance of image brightness and color caused by random haze distribution. In natural scenes with prominent edge and texture features, the above problems may cause severe degradation of image quality and performance of various tasks. Numerous studies on network learning show that the effect of haze removal is closely related to haze feature expression. Therefore, to improve the performance of dehazing, this paper proposes a sliding self-attention wavelet network. Specifically, we first design a sliding self-attention module to identify haze regions in images and capture rich haze-related feature information. Then, considering the uneven distribution of haze in images, discrete wavelet transform (DWT) and inverse transform (IDWT) are used for constructing a hierarchical encoder-decoder structure, which can fully use the multi-resolution characteristics of DWT, locally decompose feature maps of different scales, extract low and high-frequency information, and then gradually recover sharp edges and precise texture details from hazy images. Finally, to enable the proposed network to generate more realistic haze-free images on different complex haze scenes, we develop a DWT-based adversarial loss function to constrain the low and high-frequency components of generated images closer to the corresponding clear images. Experimental results on the relevant public benchmark datasets show that the proposed algorithm achieves favorable dehazing performance. Xiaozhe Meng, Fan Zhou 0001, Weisi Lin, Zhuo Su 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Multi-Target Domain Transfer Detector on Bad Weather ConditionsabstractWith the popularization of object detection in fields such as autonomous driving and surveillance, object detection networks are required to be used in more changeable scenarios. However, a multi-target domain adaptive object detection network is a challenging problem. Due to the lack of relevant datasets and the limitation of traditional domain adaptive methods, there is a lack of pertinent research in this field. Traditional single-target domain adaptive methods are often ineffective when used in multi-target domain adaptive problems. To allow the object detection network to be applied to multiple target domains, we propose the “domain transfer module” and the “multi-scale hybrid attention domain alignment module”. At the same time, we synthesized the “BlendedCityspce” dataset for training and testing of a multi-target domain adaptive target object network. The network provided in this paper has a good performance in the multi-target domain adaptation and reaches the state-of-the-art in the single-target domain adaptation. You can find our source code in https://github.com/Chonghuan-Liu/Multi-Target-Domain-Adaptative-Detector. Chonghuan Liu, Zhuo Su 0001, Fan Zhou 0001 |
ICME | 3 |
| 2022 | Unsupervised Domain Adaptation Image Dehazing with Contrastive Nearest-Farthest Subspace DistanceabstractHaze can cause significant changes in the data domain of the image. Due to domain-shift problem, the model trained on the synthetic haze images has a weak or even invalid dehazing ef-fect on the real data. This paper proposes an unsupervised domain adaptation dehazing method based on the nearest-farthest subspace distance in response to this problem. For the middle convolved features, the matrix singular value de-composition is used to obtain the basis vectors of the support subspace. Narrowing the angle between the basis vectors of natural and synthetic haze domains could reduce the differ-ence between the data domains. In subspace, a newly defined distance with a new penalty term named subspace measure distance is employed to constrain the model. Furthermore, to avoid search blindness of the proposed method in the real data training phase, we propose the nearest-farthest subspace distance inspired by contrastive learning. In addition, we use a new training strategy. Finally, the experimental tests on the synthetic and real images prove the effectiveness of the pro-posed method. Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001 |
ICME | 3 |
| 2022 | Cross-Rolling Attention Network for Fashion Landmark DetectionabstractThe variety and deformable characteristics of fashions are the challenges in improving the accuracy of fashion landmark detection. Considering the delicate characteristics, this paper proposes a new method for fashion landmark detection based on the cross-rolling attention module and the feature pyramid network. It imports the cross-rolling attention mechanism into the multi-scale structure to extract and fuse different features with rich context, which is beneficial to capture fine-grained fashion keypoints more accurately and improve detection performance. While constructing the network, we also focus on controlling the network scale to avoid a large increase in parameters and longtime training. Finally, we perform some experiments suggesting that the proposed method achieves superior performance on both DeepFashion and FLD datasets. It demonstrates the effectiveness of the cross-rolling attention module in the proposed method. Fan Zhou 0001, Zhuo Su 0001 |
ICPR | 3 |
| 2022 | Feature Dense Relevance Network for Single Image DehazingabstractExisting learning-based dehazing methods do not fully use non-local information, which makes the restoration of seriously degraded region very tough. We propose a novel dehazing network by defining the Feature Dense Relevance module (FDR) and the Shallow Feature Mapping module (SFM). The FDR is defined based on multi-head attention to construct the dense relationship between different local features in the whole image. It enables the network to restore the degraded local regions by non-local information in complex scenes. In addition, the raw distant skip-connection easily leads to artifacts while it cannot deal with the shallow features effectively. Therefore, we define the SFM by combining the atmospheric scattering model and the distant skip-connection to effectively deal with the shallow features in different scales. It not only maps the degraded textures into clear textures by distant dependence, but also reduces artifacts and color distortions effectively. We introduce contrastive loss and focal frequency loss in the network to obtain a realitic and clear image. The extensive experiments on several synthetic and real-world datasets demonstrate that our network surpasses most of the state-of-the-art methods. Yun Liang 0003, Enze Huang, Zhuo Su 0001, Dong Wang 0041 |
IJCAI | 4 |
| 2022 | Learning compatibility knowledge for outfit recommendation with complementary clothing matching
Ruomei Wang 0001, Zhuo Su 0001 |
Comput. Commun. | 3 |
| 2022 | Attribute-aware heterogeneous graph network for fashion compatibility prediction
Zhouyi Zhou, Zhuo Su 0001, Ruomei Wang 0001 |
Neurocomputing | 2 |
| 2022 | Multi-receptive Field Aggregation Network for single image deraining
Songliang Liang, Xiaozhe Meng, Zhuo Su 0001, Fan Zhou 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Relation-Aware Attribute Network for Fine-Grained Clothing Recognition
Mingjian Yang, Zhuo Su 0001, Fan Zhou 0001 |
ICONIP (6) | 3 |
| 2021 | MVSN: A Multi-view stack network for human parsing
Zhuo Su 0001, Minshi Chen, Enbo Huang, Ge Lin 0002, Fan Zhou 0001 |
Neurocomputing | 1 |
| 2021 | Automatic 3D virtual fitting system based on skeleton driving
Guangyuan Shi, Chengying Gao, Dong Wang 0041, Zhuo Su 0001 |
Vis. Comput. | 4 |
| 2020 | Scale-Aware Rolling Fusion Network for Crowd CountingabstractDue to wide application prospects and various challenges such as large scale variation, inter-occlusion between crowd people and background noise, crowd counting is receiving increasing attention. In this paper, we propose a scale-aware rolling fusion network (SRF-Net) for crowd counting, which focuses on dealing with scale variation in highly congested noisy scenes. SRF-Net is a two-stage architecture that consists of a band-pass stage and a rolling guidance stage. Compared with the existing methods, SRF-Net achieves better results in retaining appropriate multi-level features and capturing multi-scale features, thus improving the quality of density estimation maps in crowded scenarios with large scale variation. We evaluate our method on three popular crowd counting datasets (ShanghaiTech, UCF_CC_50 and UCF-QNRF), and extensive experiments show its outperformance over the state-of-the-art approaches. Chengying Gao, Zhuo Su 0001, Xiangjian He |
ICME | 3 |
| 2020 | Tao: A Trilateral Awareness Operation For Human ParsingabstractHuman parsing, which segments a human image into semantic regions, is a fundamental task in human-centric analysis. Recently, numerous human parsing approaches based on convolutional neural networks (CNNs) have made significant progress. However, these models still suffer the high time-complexity and poor efficiency issues, since they always introduce too many and complex hypothetical priors. Therefore, in this paper, we attempt to solve above issues by mining the deep information hidden in data, such as the scale, spatial information and statistics of the image descriptors, etc. By this way, we propose an effective structure, named the trilateral awareness operation (TAO), which can boost the scale, spatial and fine-grained awareness of a CNN. In practice, it is a simple yet effective architecture without using bells and whistles (such as human pose and edge information), to address human parsing tasks. Comprehensive experiments and corresponding results on three public datasets have demonstrated that the proposed TAO is superior to the state-of-the-art methods. Enbo Huang, Zhuo Su 0001, Fan Zhou 0001 |
ICME | 2 |
| 2020 | Learning rebalanced human parsing model from imbalanced datasets
Enbo Huang, Zhuo Su 0001, Fan Zhou 0001, Ruomei Wang 0001 |
Image Vis. Comput. | 2 |
| 2019 | Image-Based Virtual Try-on Network with Structural CoherenceabstractVirtual try-on system could demonstrate the visual effect of wearing certain clothes, which is in a great demand for the online clothing customers. Due to the neglect of the original person structure, the previous try-on schemes usually encounter with the problems like the loss of body parts, the missing of body details and the deviation of clothing style. In this paper, we propose a novel image-based virtual try-on network, which could maintain the structural consistency between the generated image and the original image by human parsing. Our network consists of three components. Given the original person and the clothing images, the target human parsing maps are generated. Then, the parsing maps are matched with the target clothes to generate the warped clothes. Finally, according to the parsing results, the parts to be replaced the original images are intercepted, and more original information is retained as the input of the network to generate the final results. Experiments on an existing benchmark demonstrate our method maintains the consistency of structure and achieves the state-of-the-art performance. Jiaming Guo, Zhuo Su 0001, Chengying Gao |
ICIP | 3 |
| 2019 | Learning Transmission Filtering Network for Image-Based Pm2.5 EstimationabstractPM2.5 is an important indicator of the severity of air pollution and its level can be predicted through hazy photographs caused by its degradation. Image-based PM2.5 estimation is thus extensively employed in various multimedia applications but is challenging because of its ill-posed property. In this paper, we convert it to the problem of estimating the PM2.5-relevant haze transmission and propose a learning model called the transmission filtering network. Different from most methods that generate a transmission map directly from a hazy image, our model takes the coarse transmission map derived from the dark channel prior as the input. To obtain a transmission map that satisfies the local smoothness constraint without regional boundary degradation, our model performs the edge-preserving smoothing filtering as the refinement on the map. Moreover, we introduce the attention mechanism to the network architecture for more efficient feature extraction and smoothing effects in the transmission estimation. Experimental results prove that our model performs favorably against the state-of-the-art dehazing methods in a variety of hazy scenes. Yinghong Liao, Bin Qiu, Zhuo Su 0001, Ruomei Wang 0001, Xiangjian He |
ICME | 3 |
| 2019 | Residual Magnifier: A Dense Information Flow Network for Super ResolutionabstractRecently, deep learning methods have been successfully applied to single image super-resolution tasks. However, some networks with extreme depth failed to achieve better performance because of the insufficient utilization of the local residual information extracted at each stage. To solve the above question, we propose a Dense Information Flow Network (DIF-Net), which can fully extract and utilize the local residual information at each stage to accomplish a better reconstruction. Specifically, we present a Two-stage Residual Extraction Block (TREB) to extract the shallow and deep local residual information at each stage. The dense connection mechanism is introduced throughout the model and within TREBs to dramatically increase the information flow. Meanwhile this mechanism prevents the shallow features extracted earlier from being diluted. Finally, we propose a lightweight subnet (residual enhancer) to efficiently recycle the overflow residual information from the backbone net for detail enhancement of the residual image. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art methods with relatively-less parameters. Mengcheng Cheng, Zhuo Su 0001, Xiangjian He |
ICME | 4 |
| 2019 | Rain Wiper: An Incremental Randomly Wired Network for Single Image DerainingabstractAbstract Single image rain removal is a challenging ill‐posed problem due to various shapes and densities of rain streaks. We present a novel incremental randomly wired network (IRWN) for single image deraining. Different from previous methods, most structures of modules in IRWN are generated by a stochastic network generator based on the random graph theory, which ease the burden of manual design and further help to characterize more complex rain streaks. To decrease network parameters and extract more details efficiently, the image pyramid is fused via the multi‐scale network structure. An incremental rectified loss is proposed to better remove rain streaks in different rain conditions and recover the texture information of target objects. Extensive experiments on synthetic and real‐world datasets demonstrate that the proposed method outperforms the state‐of‐the‐art methods significantly. In addition, an ablation study is conducted to illustrate the improvements obtained by different modules and loss items in IRWN. Xiangguo Liang, Bin Qiu, Zhuo Su 0001, Chengying Gao, X. Shi, Ruomei Wang 0001 |
Comput. Graph. Forum | 3 |
| 2019 | Learning mean progressive scattering using binomial truncated loss for image dehazingabstractIn this study, the authors propose a novel progressive dehazing network to address the single image haze removal problem based on a new mean progressive scattering model. Different from methods that learn atmosphere light and transmission maps with different networks, these two variables are optimised in a unified network. Following the methodology of traditional prior‐based methods that estimate a coarse transmission map first, a progressive refinement branch in the decoder has been designed to restore the fine‐scale transmission map. To improve the prediction accuracy of the transmission map, a novel binomial truncated loss that assigns weights to error values according to the probabilities of error occurrences has been proposed. An ablation study is conducted to verify the effectiveness of the components in the proposed method. Experiments in the synthetic datasets and real images demonstrate that the proposed method outperforms other state‐of‐the‐art methods. Bin Qiu, Xiwen Liang, Zhuo Su 0001, Ruomei Wang 0001, Fan Zhou 0001 |
IET Image Process. | 3 |
| 2019 | Conditional progressive network for clothing parsingabstractClothing parsing is significant to many clothing applications. Recently, a lot of clothing parsing methods have been presented, which explore the innovation of the parsing pipeline or try to find more specific prior information. Although these methods perform well in some benchmarks, a few challenging problems have not been solved yet, such as the complicated mutual interference among labels. In this study, the authors propose a Conditional Progressive Network to parse clothing in different scales and prevent the mutual interference among labels. The authors’ solution consists of three sub‐networks, including Conditional Parsing Network (CPN), Pose Estimation Network (PEN) and Label Transform Network (LTN). Specifically, the CPN module generates the intermediate parsing result in the form of the multiple progressive stages, which combines with the previous outputs in each stage and the specific prior conditions. The PEN module provides a series of heat maps about the human pose information. The LTN module suppresses the redundant labels to avoid the mutual interference among labels. They demonstrate their solution in parsing the fashion clothing cases on the ATR and the Fashion dataset. In their experiments, their method obtains a better performance than the state‐of‐the‐art methods. Zhuo Su 0001, Jiaming Guo, Gengwei Zhang, Xianghui Luo, Ruomei Wang 0001, Fan Zhou 0001 |
IET Image Process. | 1 |
| 2018 | Trusted Guidance Pyramid Network for Human ParsingabstractHuman parsing, which segments a human-centric image into pixel-wise categorization, has a wide range of applications. However, none of the existing methods can productively solve the issue of label parsing fragmentation due to confused and complicated annotations. In this paper, we propose a novel Trusted Guidance Pyramid Network (TGPNet) to address this limitation. Based on a pyramid architecture, we design a Pyramid Residual Pooling (PRP) module setting at the end of a bottom-up approach to capture both global and local level context. In the top-down approach, we propose a Trusted Guidance Multi-scale Supervision (TGMS) that efficiently integrates and supervises multi-scale contextual information. Furthermore, we present a simple yet powerful Trusted Guidance Framework (TGF) which imposes global-level semantics into parsing results directly without extra ground truth labels in model training. Extensive experiments on two public human parsing benchmarks well demonstrate that our TGPNet has a strong ability in solving label parsing fragmentation problem and has an obtained improvement than other methods. Xianghui Luo, Zhuo Su 0001, Jiaming Guo, Gengwei Zhang, Xiangjian He |
ACM Multimedia | 2 |
| 2018 | 3D medical model low-pass filtering based on non-uniform spectral synthesis
Yihui Guo, Zhuo Su 0001, Shujin Lin, Jiyuan Lu, Xueling Zhong |
Comput. Aided Des. | 2 |
| 2018 | PencilArt: A Chromatic Penciling Style Generation FrameworkabstractAbstract Non‐photorealistic rendering has been an active area of research for decades whereas few of them concentrate on rendering chromatic penciling style. In this paper, we present a framework named as PencilArt for the chromatic penciling style generation from wild photographs. The structural outline and textured map for composing the chromatic pencil drawing are generated, respectively. First, we take advantage of deep neural network to produce the structural outline with proper intensity variation and conciseness. Next, for the textured map, we follow the painting process of artists to adjust the tone of input images to match the luminance histogram and pencil textures of real drawings. Eventually, we evaluate PencilArt via a series of comparisons to previous work, showing that our results better capture the main features of real chromatic pencil drawings and have an improved visual appearance. Chengying Gao, Mengyue Tang, Xiangguo Liang, Zhuo Su 0001, Changqing Zou |
Comput. Graph. Forum | 4 |
| 2018 | An edge-refined vectorized deep colorization model for grayscale-to-color images
Zhuo Su 0001, Xiangguo Liang, Jiaming Guo, Chengying Gao |
Neurocomputing | 1 |
| 2018 | PSI: A probabilistic semantic interpretable framework for fine-grained image rankingabstractImage Ranking is one of the key problems in information science research area. However, most current methods focus on increasing the performance, leaving the semantic gap problem, which refers to the learned ranking models are hard to be understood, remaining intact. Therefore, in this article, we aim at learning an interpretable ranking model to tackle the semantic gap in fine‐grained image ranking. We propose to combine attribute‐based representation and online passive‐aggressive (PA) learning based ranking models to achieve this goal. Besides, considering the highly localized instances in fine‐grained image ranking, we introduce a supervised constrained clustering method to gather class‐balanced training instances for local PA‐based models, and incorporate the learned local models into a unified probabilistic framework. Extensive experiments on the benchmark demonstrate that the proposed framework outperforms state‐of‐the‐art methods in terms of accuracy and speed. Hefeng Wu, Shujin Lin, Zhuo Su 0001 |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2017 | ℒ0 Gradient-Preserving Color TransferabstractAbstract This paper presents a new two‐step color transfer method which includes color mapping and detail preservation. To map source colors to target colors, which are from an image or palette, the proposed similarity‐preserving color mapping algorithm uses the similarities between pixel color and dominant colors as existing algorithms and emphasizes the similarities between source image pixel colors. Detail preservation is performed by an ℒ0 gradient‐preserving algorithm. It relaxes the large gradients of the sparse pixels along color region boundaries and preserves the small gradients of pixels within color regions. The proposed method preserves source image color similarity and image details well. Extensive experiments demonstrate that the proposed approach has achieved a state‐of‐art visual performance. Dong Wang 0041, Changqing Zou, Guiqing Li, Chengying Gao, Zhuo Su 0001 |
Comput. Graph. Forum | 5 |
| 2017 | Maximised self-similarity upsamplerabstractImage self‐similarity property is important to super‐resolution reconstruction. However, how to effectively exploit the self‐similarity information to reconstruct an underlying high‐resolution image is still a challenging problem. The authors propose a novel model for solving the single image upsampling problem with the self‐similarity property. First, the authors construct a statistical prior that requires maximising the similarity between the low‐ and high‐resolution image pairs. Then, the authors develop an alternative Gaussian approximation solver based on the Gaussian mixture model to find the optimal high‐resolution output. To obtain a better performance, the authors summarise some refined implementation skills to raise the reconstruction quality. For demonstration, a series of objective and subjective measurements are used to evaluate the performance of the model. Zhuo Su 0001, Langyu Li |
IET Image Process. | 1 |
| 2017 | Semi-guided bilateral filterabstractThe bilateral filter (BF) is a non‐linear filter that spatially smooths images with awareness of large structures such as edges. The level of smoothness applied to a pixel is constrained by a photometric weight, which can be obtained from the same image to be filtered (in case of the original BF) or from a guided image (in case of the joint/cross BF). In this study, the authors propose a new filter called the semi‐guided BF which is derived from solving a non‐linear constraint least square problem. The proposed filter's photometric weight incorporates information from the image to be filtered and the guided image. They propose a fast implementation of the filter based on layer approximation. They also study the iterative application of the proposed filter and show that the filter can preserve large structures while smoothing out small structures. This makes the proposed filter an efficient and effective tool for structure‐aware image smoothing. Experimental results have demonstrated that performance of the proposed filter is comparable to those of the state‐of‐the‐art algorithms. Ba Thai, Mukhalad Al-nasrawi, Guang Deng, Zhuo Su 0001 |
IET Image Process. | 4 |
| 2017 | Boosting attribute recognition with latent topics by matrix factorizationabstractAttribute‐based approaches have recently attracted much attention in visual recognition tasks. These approaches describe images by using semantic attributes as the mid‐level feature. However, low recognition accuracy becomes the biggest barrier that limits their practical applications. In this paper, we propose a novel framework termed Boosting Attribute Recognition (BAR) for the image recognition task. Our framework stems from matrix factorization, and can explore latent relationships from the aspect of attribute and image simultaneously. Furthermore, to apply our framework in large‐scale visual recognition tasks, we present both offline and online learning implementation of the proposed framework. Extensive experiments on 3 data sets demonstrate that our framework achieves a sound accuracy of attribute recognition. Zhuo Su 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2017 | A data-driven editing framework for automatic 3D garment modeling
Li Liu 0032, Zhuo Su 0001, Xiaodong Fu, Ruomei Wang 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Example-Based Image Colorization Using Locality Consistent Sparse RepresentationabstractImage colorization aims to produce a natural looking color image from a given gray-scale image, which remains a challenging problem. In this paper, we propose a novel example-based image colorization method exploiting a new locality consistent sparse representation. Given a single reference color image, our method automatically colorizes the target gray-scale image by sparse pursuit. For efficiency and robustness, our method operates at the superpixel level. We extract low-level intensity features, mid-level texture features, and high-level semantic features for each superpixel, which are then concatenated to form its descriptor. The collection of feature vectors for all the superpixels from the reference image composes the dictionary. We formulate colorization of target superpixels as a dictionary-based sparse reconstruction problem. Inspired by the observation that superpixels with similar spatial location and/or feature representation are likely to match spatially close regions from the reference image, we further introduce a locality promoting regularization term into the energy formulation, which substantially improves the matching consistency and subsequent colorization results. Target superpixels are colorized based on the chrominance information from the dominant reference superpixels. Finally, to further improve coherence while preserving sharpness, we develop a new edge-preserving filter for chrominance channels with the guidance from the target gray-scale image. To the best of our knowledge, this is the first work on sparse pursuit image colorization from single reference images. Experimental results demonstrate that our colorization method outperforms the state-of-the-art methods, both visually and quantitatively using a user study. Bo Li 0023, Fuchen Zhao, Zhuo Su 0001, Xiangguo Liang, Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 3 |
| 2017 | A Dual-Domain Perceptual Framework for Generating Visual Inconspicuous CounterpartsabstractFor a given image, it is a challenging task to generate its corresponding counterpart with visual inconspicuous modification. The complexity of this problem reasons from the high correlativity between the editing operations and vision perception. Essentially, a significant requirement that should be emphasized is how to make the object modifications hard to be found visually in the generative counterparts. In this article, we propose a novel dual-domain perceptual framework to generate visual inconspicuous counterparts, which applies the perceptual bidirectional similarity metric (PBSM) and appearance similarity metric (ASM) to create the dual-domain perception error minimization model. The candidate targets are yielded by the well-known PatchMatch model with the strokes-based interactions and selective object library. By the dual-perceptual evaluation index, all candidate targets are sorted to select out the best result. For demonstration, a series of objective and subjective measurements are used to evaluate the performance of our framework. Zhuo Su 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2016 | A 3D model perceptual feature metric based on global height field
Yihui Guo, Shujin Lin, Zhuo Su 0001, Ruomei Wang 0001, Yang Kang |
Vis. Comput. | 3 |
| 2015 | Super-resolution by polar Newton-Thiele's rational kernel in centralized sparsity paradigm
Lei He 0002, Jieqing Tan, Zhuo Su 0001, Chengjun Xie |
Signal Process. Image Commun. | 3 |
| 2015 | Corrigendum to "Super-resolution by polar Newton-Thiele's rational kernel in centralized sparsity paradigm" [Signal Processing: Image Communication, 31(2015), pp 86-99]
Lei He 0002, Jieqing Tan, Zhuo Su 0001, Chengjun Xie |
Signal Process. Image Commun. | 3 |
| 2015 | Robust tracking via discriminative sparse feature selection
Jin Zhan, Zhuo Su 0001, Hefeng Wu |
Vis. Comput. | 2 |
| 2014 | High-order wavelet reconstruction for multi-scale edge aware tone mapping
Alessandro Artusi, Zhuo Su 0001, Zongwei Zhang, Dimitris Drikakis |
Comput. Graph. | 2 |
| 2014 | Mesh-based anisotropic cloth deformation for virtual fitting
Li Liu 0032, Ruomei Wang 0001, Zhuo Su 0001, Chengying Gao |
Multim. Tools Appl. | 3 |
| 2014 | Corruptive Artifacts Suppression for Example-Based Color TransferabstractExample-based color transfer is a critical operation in image editing but easily suffers from some corruptive artifacts in the mapping process. In this paper, we propose a novel unified color transfer framework with corruptive artifacts suppression, which performs iterative probabilistic color mapping with self-learning filtering scheme and multiscale detail manipulation scheme in minimizing the normalized Kullback-Leibler distance. First, an iterative probabilistic color mapping is applied to construct the mapping relationship between the reference and target images. Then, a self-learning filtering scheme is applied into the transfer process to prevent from artifacts and extract details. The transferred output and the extracted multi-levels details are integrated by the measurement minimization to yield the final result. Our framework achieves a sound grain suppression, color fidelity and detail appearance seamlessly. For demonstration, a series of objective and subjective measurements are used to evaluate the quality in color transfer. Finally, a few extended applications are implemented to show the applicability of this framework. Zhuo Su 0001, Li Liu 0032, Bo Li 0023 |
IEEE Trans. Multim. | 1 |
| 2013 | Material-aware cloth simulation via constrained geometric deformation
Li Liu 0032, Zhuo Su 0001, Ruomei Wang 0001 |
Comput. Graph. | 2 |
| 2013 | Optimised image retargeting using aesthetic-based cropping and scalingabstractImage retargeting is a critical technique in displaying images on devices with different resolutions. This study presents a new image retargeting algorithm based on aesthetic‐based cropping and scaling. A composite measurement is first constructed under the guidelines of composition aesthetics in photographing. An aesthetic‐based cropping is proposed to yield an optimal candidate retargeted image with maximum aesthetic value computed via a constructed composite measurement. The optimal candidate is uniformly scaled to obtain the retargeted image of target size. Some subjective and objective assessments demonstrate that the proposed scheme significantly improves the aesthetics of retargeted images while preserving the important objects. It also achieves better performance in terms of aesthetics than a number of conventional image retargeting approaches. Yun Liang 0003, Zhuo Su 0001, Chuntao Wang, Dong Wang 0041 |
IET Image Process. | 2 |
| 2013 | Edge-Preserving Texture Suppression Filter Based on Joint Filtering SchemesabstractObtaining a texture-smoothing and edge-preserving filtered output is significant to image decomposition. Although the edge and the texture have salient difference in human vision, automatically distinguishing them is a difficult task, for they have similar intensity difference or gradient response. The state-of-the-art edge-preserving smoothing (EPS) based decomposition approaches are hard to obtain a satisfactory result. We propose a novel edge-preserving texture suppression filter, exploiting the joint bilateral filter as a bridge to achieve the purpose of both properties of texture-smoothing and edge-preserving. We develop the iterative asymmetric sampling and the local linear model to produce the degenerative image to suppress the texture, and apply the edge correction operator to achieve edge-preserving. An efficient accelerating implementation is introduced to improve the performance of filtering response. The experiments demonstrate that our filter produces satisfactory outputs with both properties of texture-smoothing and edge-preserving, while compared with the results of other popular EPS approaches in signal, visual and time analysis. Finally, we extend our filter to a variety of image processing applications. Zhuo Su 0001, Zhengjie Deng, Yun Liang 0003, Zhen Ji |
IEEE Trans. Multim. | 1 |
| 2013 | A novel image decomposition approach and its applications
Zhuo Su 0001, Alessandro Artusi |
Vis. Comput. | 1 |
| 2012 | Online boosted tracking with discriminative feature selection and scale adaptationabstractWe track the object by separating it from the surrounding with an ensemble of boosted classifiers, which are trained in a discriminative feature space that is determined on the fly. Contour refinement and weight thresholding techniques are used to select good examples for training. While tracking, location calibration and scale adaptation are used to improve the tracker's performance. We update the ensemble of weak classifiers online to adapt to appearance changes, and use the positive occupancy ratio to detect occlusion. A center-surround discrepancy measure is presented to evaluate the discriminative power of the current feature space and to invoke re-initialization of feature selection and classifier training if necessary. Experiments on challenging video sequences demonstrate the effectiveness of the proposed approach. Hefeng Wu, Guanbin Li, Zhuo Su 0001 |
ICIP | 3 |
| 2012 | Local color editing using color classification and boundary inpainting
Zhuo Su 0001, Dong Wang 0041 |
ICPR | 1 |
| 2012 | Color transfer based on multiscale gradient-aware decomposition and color distribution mappingabstractAutomatic global color transfer is a challenging problem in image editing. In this paper, we propose a novel color transfer method, which is based on the gradient-aware decomposition and the color distribution mapping. Firstly, a gradient-aware decomposition model is established to separate the target image into the base and detail layers. Then, the colors of each separated base layer are enforced to match those of a given reference image by Pitie's multi-dimensional probability density function transfer method. After that, all mapped base layers are combined with corresponding boosted detail layers to produce the final output. The experiments demonstrate that our method can achieve a visual satisfied result without post-processing gradient correction. Zhuo Su 0001, Daiguo Deng |
ACM Multimedia | 1 |
| 2012 | Patchwise scaling method for content-aware image resizing
Yun Liang 0003, Zhuo Su 0001 |
Signal Process. | 2 |