VLDB 2026 Research / reviewers in the wild / expert
Xiaoyan Zhang 0002
dblp:63/4485-2 · also Xiao-Yan Zhang 0002
· DBLP profile ↗
31ranked-venue papers
14as first author
8since 2021 · last 2025
0000-0002-5446-7568ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VAT: Visibility Aware Transformer for Fine-Grained Clothed Human ReconstructionabstractIn order to reconstruct 3D clothed human with accurate fine-grained details from sparse views, we propose a deep cooperating two-level global to fine-grained reconstruction framework that constructs robust global geometry to guide fine-grained geometry learning. The core of the framework is a novel visibility aware Transformer VAT, which bridges the two-level reconstruction architecture by connecting its global encoder and fine-grained decoder with two pixel-aligned implicit functions, respectively. The global encoder fuses semantic features of multiple views to integrate global geometric features. In the fine-grained decoder, visibility aware attention mechanism is designed to efficiently fuse multi-view and multi-scale features for mining fine-grained geometric features. The global encoder and fine-grained decoder are connected by a global embeding module to form a deep cooperation in the two-level framework, which provides global geometric embedding as a query guidance for calculating visibility aware attention in the fine-grained decoder. In addition, to extract highly aligned multi-scale features for the two-level reconstruction architecture, we design an image feature extractor MSUNet, which establishes strong semantic connections between different scales at minimal cost. Our proposed framework is end-to-end trainable, with all modules jointly optimized. We validate the effectiveness of our framework on public benchmarks, and experimental results demonstrate that our method has significant advantages over state-of-the-art methods in terms of both fine-grained performance and generalization. Xiaoyan Zhang 0002, Zibin Zhu, Sisi Ren, Jianmin Jiang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | BiPR-RL: Portrait relighting via bi-directional consistent deep reinforcement learning
Yukai Song, Guangxin Xu, Xiaoyan Zhang 0002, Zhijun Zhang 0003 |
Comput. Vis. Image Underst. | 3 |
| 2023 | FFFN: Fashion Feature Fusion Network by Co-Attention Model for Fashion RecommendationabstractFashion complementary recommendation has always been crucial in the field of recommendation. In previous work, researchers did not pay attention to the connection and combinability between multi-dimensional image features of fashion items. To effectively utilize the advantages of high-level and low-level features in images, we propose a Fashion Feature Fusion Network (FFFN) to solve the fashion complementary recommended tasks, which extracts and combines the features of different dimensions in the neural network into a fusion feature. Then, we input the fused image features and the category text features of fashion items into the co-attention module to realize the guidance of the text attention information to the image attention information. Experimental results on different datasets show that our model has advantages compared to the state-of-the-art methods. Zhantu Lin, Xiaoyan Zhang 0002 |
ICASSP | 2 |
| 2023 | RepEPnP: Weakly Supervised 3D Human Pose Estimation with EPnP AlgorithmabstractThis paper describes an end-to-end weakly super-vised framework for estimating 3D human pose from a single image. The model is trained by projecting 3D pose to 2D pose for matching ground-truth 2D pose for supervision. To obtain accu-rate projection from 3D pose to 2D pose, a mathematical camera model based on intrinsic and extrinsic camera parameters is used. Specifically, we use EPnP algorithm to estimate extrinsic transformation matrix to transform the estimated 3D pose to be reprojected back to 2D pose. The advantage of this projection is that it requires no training and it is robust to the diversity of training datasets. We further constrain the pose generation using an adversarial generative network, where a transformer is used as the 3D pose generator. Transformer can use self-attention mechanism to establish dependencies between each joint and predict pose based on important joints. Based on our reprojection method, our method achieves competitive results on Human3.6M and MPI-INF-3DHP among weakly supervised methods. The experiments also demonstrate our model's generalization ability for wild images. Huaijing Lai, Zhenhua Tang 0001, Xiaoyan Zhang 0002 |
IJCNN | 3 |
| 2022 | PR-RL: Portrait Relighting Via Deep Reinforcement LearningabstractIn this paper, we propose a portrait relighting method based on deep reinforcement learning (called PR-RL). Our PR-RL model could conduct portrait relighting by sequentially predicting local light editing strokes, and use strokes to conduct dodge and burn operations on the image lightness, simulating image editing by artists using brush strokes. Reinforcement learning with Deep Deterministic Policy Gradient is introduced to design our PR-RL model, defining the action (stroke parameters) in a continuous space, through which a reward can be designed to guide the agent to learn and relight a portrait image like an artist. To optimize the relighting effect, we further enable the reward to be location relevant and hence a coarse-to-fine strategy can be applied to select corresponding actions and maximize the performance of the proposed method. In comparison with the existing efforts, our proposed PR-RL method is locally effective, scale-invariant and interpretable. We apply the proposed method to tasks of portrait relighting based on both SH-lighting and reference images. The experiments show that our PR-RL method outperforms state-of-the-art methods in generating locally effective and interpretable high resolution relighting results for wild portrait images. Xiaoyan Zhang 0002, Yukai Song, Zhuopeng Li, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2021 | Reference-Based Video Colorization With Multi-Scale Semantic Fusion And Temporal AugmentationabstractThe reference-based video colorization method hallucinates a plausible color version for a gray-scale video by referring distributions of possible colors from an input color frame, which has semantic correspondences with the gray-scale frames. The plausibility of colors and the temporal consistency are two significant challenges in this task. In this paper, we propose a novel Generative Adversarial Network (GAN) with a siamese training framework to tackle these challenges. Specifically, the siamese training framework allows us to implement temporal feature augmentation, enhancing temporal consistency. Further, to improve the plausibility of colorization results, we propose a multi-scale fusion module that correlates features of reference frames to source frames accurately. Experiments on various datasets demonstrate that our proposed method performs favorably against the state-of the-art approaches. Xiaoyan Zhang 0002, Xiaogang Xu 0002 |
ICIP | 2 |
| 2021 | Semantic-Aware Video Style Transfer Based on Temporal Consistent Sparse Patch ConstraintabstractThis paper proposes a practical style transfer method to synthesize a temporally smooth video whose style information is semantically consistent with the reference video. Due to the lack of paired videos for training, we extend the structure of CycleGAN with sparse patch and temporal constraints, including a new semantic patch loss and a novel temporal loss. Our approach’s key insights are: (1) the semantically paired sparse patches chosen from synthesized videos and reference frames would promote the semantic meaning of style transfer, the preservation of video content, and the smoothness of results by minimizing the discrepancies between these paired patches. (2) the forward and backward temporal consistency among neighbouring frames can reduce the discontinuity in the synthesized video. Extensive quantitative and qualitative experiments on various metrics demonstrate the superiority of our method over state-of-the-art strategies. Xiaoyan Zhang 0002, Xiaogang Xu 0002 |
ICME | 2 |
| 2021 | Emotion Attention-Aware Collaborative Deep Reinforcement Learning for Image CroppingabstractThis paper proposes a collaborative deep reinforcement learning model for automatic image cropping (called CDRL-IC). By modeling image cropping as a decision-making process of reinforcement learning, our model could generate optimal cropping result in a few moving and zooming steps. An image with good composition is a comprehensive result by considering the relative importance of objects and also the spatial organization of visual elements. Therefore, emotion attention information which indicates the relationship and importance between objects is applied together with contextual information of color image for image cropping. In order to sufficiently use the emotion attention map and the color image, they are processed by two collaborative agents. The two agents make their primary learning separately and then share information through an information interaction module for making joint action prediction. In order to efficiently evaluate the cropping quality in the reward function, weighted Intersection Over Union (WIoU) is designed by integrating emotion attention map in the traditional IoU. Our CDRL-IC model is tested on a variety of datasets for both image cropping and thumbnail generation. The experiments show that our CDRL-IC model outperforms state-of-the-art methods on these benchmark datasets. Xiaoyan Zhang 0002, Zhuopeng Li, Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2020 | Leveraging Title-Abstract Attentive Semantics for Paper RecommendationabstractPaper recommendation is a research topic to provide users with personalized papers of interest. However, most existing approaches equally treat title and abstract as the input to learn the representation of a paper, ignoring their semantic relationship. In this paper, we regard the abstract as a sequence of sentences, and propose a two-level attentive neural network to capture: (1) the ability of each word within a sentence to reflect if it is semantically close to the words within the title. (2) the extent of each sentence in the abstract relative to the title, which is often a good summarization of the abstract document. Specifically, we propose a Long-Short Term Memory (LSTM) network with attention to learn the representation of sentences, and integrate a Gated Recurrent Unit (GRU) network with a memory network to learn the long-term sequential sentence patterns of interacted papers for both user and item (paper) modeling. We conduct extensive experiments on two real datasets, and show that our approach outperforms other state-of-the-art approaches in terms of accuracy. Guibing Guo, Bowei Chen 0004, Xiaoyan Zhang 0002, Zhenhua Dong, Xiuqiang He 0001 |
AAAI | 3 |
| 2020 | Convolutional Attention Model For Restaurant Recommendation With Multi-View Visual FeaturesabstractCurrent recommendation systems usually use image as a type of side information to enhance recommendation systems. However, there are few models to analyze and apply the effectiveness of different categories of images. In order to effectively combine multi-category image information in the recommendation system, we propose a novel deep network model with different categories of images as side information for recommendation, in which we use convolutional attention to analyze the importance of images within different image categories. The convolutional layers of the attention module are used to evaluate the user's visual preferences for images in different categories. The prediction model uses the cosine similarity of user factors and restaurant factors incorporating visual information. Visual features of the restaurants are extracted from multi-category images by a pre-trained neural network. We apply our model to two real-world restaurant recommendation data sets. Experimental results show that the performance of our model is better than models without or with only one category visual information. Haihua Luo, Xiaoyan Zhang 0002, Guibing Guo |
ICIP | 2 |
| 2020 | Multi-view visual Bayesian personalized ranking for restaurant recommendation
Xiaoyan Zhang 0002, Haihua Luo, Bowei Chen 0004, Guibing Guo |
Appl. Intell. | 1 |
| 2019 | An Articulated Structure-aware Network for 3D Human Pose EstimationabstractIn this paper, we propose a new end-to-end articulated structure-aware network to regress 3D joint coordinates from the given 2D joint detections. The proposed method is capable of dealing with hard joints well that usually fail existing methods. Specifically, our framework cascades a refinement network with a basic network for two types of joints, and employs a attention module to simulate a camera projection model. In addition, we propose to use a random enhancement module to intensify the constraints between joints. Experimental results on the Human3.6M and HumanEva databases demonstrate the effectiveness and flexibility of the proposed network, and errors of hard joints and bone lengths are significantly reduced, compared with state-of-the-art approaches. Zhenhua Tang 0001, Xiaoyan Zhang 0002, Junhui Hou |
ACML | 2 |
| 2019 | Collaborative Deep Reinforcement Learning for Image CroppingabstractAn automatic photo composition method based on collaborative deep reinforcement learning(called CDRL-RC) is proposed in this paper. Our method models photo composition as a markov decision-making process by reinforcement learning and generates cropping result through a series of moving and zooming actions. Emotional attention information is added to the composition task, which was trained by eye-tracking datasets to consider the relationship and importance between objects. In order to sufficiently use the emotional attention map and original image for image cropping, they are processed as inputs to two collaborative agents. For the collaborative composition of two agents, we design an information interaction module, which allows inter-agents to exchange information and give advice to each other, and finally predict the action together. In addition, we add attention weight to the traditional IoU to efficiently evaluate the cropping quantity in the reward function. Experiment results show that our CDRL-RC model achieved the state-of-the-art photo composition performance on a variety of datasets. Zhuopeng Li, Xiaoyan Zhang 0002 |
ICME | 2 |
| 2019 | Semantic Prior Guided Face InpaintingabstractFace inpainting is a sub-task of image inpainting designed to repair broken or occluded incomplete portraits. Due to the high complexity of face image details, inpainting on the face is more difficult. At present, face-related tasks often draw on excellent methods from face recognition and face detection, using multitasking to boost its effect. Therefore, this paper proposes to add the face prior knowledge to the existing advanced inpainting model, combined with perceptual loss and SSIM loss to improve the model repair efficiency. A new face inpainting process and algorithm is implemented, and the repair effect is improved. Xiaobo Zhou 0001, Shengjie Zhao 0001, Xiaoyan Zhang 0002 |
MMAsia | 4 |
| 2019 | Deep Reinforcement Learning for Automatic Thumbnail Generation
Zhuopeng Li, Xiaoyan Zhang 0002 |
MMM (2) | 2 |
| 2019 | Deep photographic style transfer guided by semantic correspondence
Xiaoyan Zhang 0002, Xiaole Zhang, Zhijiao Xiao |
Multim. Tools Appl. | 1 |
| 2019 | 3D human pose estimation via human structure-aware fully connected network
Xiaoyan Zhang 0002, Zhenhua Tang 0001, Junhui Hou, Yanbin Hao |
Pattern Recognit. Lett. | 1 |
| 2019 | Pose-Based Composition Improvement for Portrait PhotographsabstractThis paper studies the composition in portrait paintings and develops an algorithm to improve the composition of portrait photographs based on example portrait paintings. A study of portrait paintings shows that the placement of the face and the figure is pose-related. Based on this observation, this paper develops an algorithm to improve the composition of a portrait photograph by learning the placement of the face and the figure from an example portrait painting. This example portrait painting is selected based on the similarity of its figure pose to that of the input photograph. This similarity measure is modeled as a graph matching problem. Finally, space cropping is performed using an optimization function to assign a similar location for each body part of the figure in the photograph with that of the figure in the example portrait painting. The experimental results demonstrate the effectiveness of the proposed method. A user study shows that the proposed pose-based composition improvement is preferred more than rule-based methods and learning-based methods. Xiaoyan Zhang 0002, Zhuopeng Li, Martin Constable, Kap Luk Chan, Zhenhua Tang 0001, Gaoyang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Multi-view Visual Bayesian Personalized Ranking from Implicit FeedbackabstractIn this paper, we propose a new factorization model that combines multi-view visual feature information with the implicit feedback data for prediction and ranking. The visual information is integrated into a collaborative filtering framework. The visual features of images are extracted by using a deep neural network. In order to conduct personalized recommendation better, the multi-view visual features are fused through user related weights. The user related weights reflect the personalized visual preference for items. They are different and independent between users. Experimental results show that our model with multi-view visual information achieves the better performance than models without or with only single-view visual information. Haihua Luo, Xiaoyan Zhang 0002, Bowei Chen 0004, Guibing Guo |
UMAP | 2 |
| 2018 | Haze removal method for natural restoration of images with sky
Yingying Zhu 0001, Gaoyang Tang, Xiaoyan Zhang 0002, Jianmin Jiang, Qi Tian 0001 |
Neurocomputing | 3 |
| 2018 | Transfer of content-aware vignetting effect from paintings to photographs
Xiaoyan Zhang 0002, Martin Constable, Kap Luk Chan |
Multim. Tools Appl. | 1 |
| 2017 | Transfer of vignetting effect from paintings to photographsabstractThis paper addresses the aesthetic enhancement of the vignetting effect in photographs to follow such effect organization in example paintings. The example painting is selected based on a comparison of the contextual geometry of the painting and the photograph. Then an algorithm is developed to transfer the lightness weighting extracted from an example painting to a photograph to create the painter-style vignetting effect. Experiments show that the proposed algorithm can successfully transfer the vignetting effect from an example painting to a photograph. Meanwhile, the vignetting effect is more naturally presented with regard to aesthetic composition comparing with popular software tools and camera models. Xiaoyan Zhang 0002, Martin Constable, Kap Luk Chan |
ICASSP | 1 |
| 2017 | Pose-based composition improvement for portrait photographsabstractThis paper studies the composition of portrait paintings and develops an algorithm to improve the composition of portrait photographs. The study of portrait paintings shows that placement of the face and the figure in portrait paintings is pose-related. Based on this observation, this paper develops an algorithm to improve the composition of a portrait photograph by learning the placement of the face and the figure from an example portrait painting. The example portrait painting is selected based on the similarity of its figure pose to that of the input photograph. This similarity measure is modeled as a graph matching problem. Finally, space cropping is performed using an optimization function. Experimental results and a user study demonstrate that the proposed pose-based improvement is preferred more than rule-based methods. Xiaoyan Zhang 0002, Kap Luk Chan, Martin Constable |
ICASSP | 1 |
| 2016 | A new neural-dynamic control method of position and angular stabilization for autonomous quadrotor UAVsabstractQuadrotor unmanned aerial vehicles (UAVs) have been widely used or have great potential applications in military, entertainment, postal delivery, agriculture for working aloft, photographing, etc. The position and angular stabilization of quadrotor UAVs is very significant and it is a challenging work because of the nonlinear dynamic behavior. In this paper, a neural dynamic method based control system is designed and investigated by combination of Zhang dynamics and gradient dynamic (ZD-GD) methods. Quadrotor UAVs equipped with the ZD-GD controllers can realize position and angular stabilization autonomously. Computer simulation results substantiate the efficiency and accuracy of the proposed neural dynamic method based ZD-GD controllers. Besides, the performance of the controllers can be remarkably advanced. Zhijun Zhang 0003, Jianli Yu, Yuanqing Li 0001, Xiaoyan Zhang 0002 |
FUZZ-IEEE | 4 |
| 2016 | Two-Stage Simulation Method to Improve Facial Soft Tissue Prediction Accuracy for Orthognathic Surgery
Daeseung Kim, Chien-Ming Chang, Dennis Chun-Yu Ho, Xiaoyan Zhang 0002, Shunyao Shen, Peng Yuan 0001, Huaming Mai, Xiaobo Zhou 0001, Jaime Gateno, Michael A. K. Liebschner, James J. Xia |
MICCAI (1) | 4 |
| 2015 | Prediction of facial soft tissue deformations with improved rubin-bodner model after craniomaxillofacial (CMF) surgeryabstractAccurate prediction of the soft tissue deformation is a key issue in craniomaxillofacial (CMF) surgery, which makes it possible to transform a good surgical plan to a successful real surgical outcome. However, it is difficult to simulate the soft tissue reactions caused by CMF surgery according to its nonlinear and anisotropic attributes. In this paper, we originally improved the Rubin-Bodner (RB) model to describe the biomechanical interaction of the soft tissue after CMF surgery, where the elastic relevant parameters are trained by Generalized Regression Neural Network (GRNN) corresponding to different CMF surgical types respectively. Subsequently, finite element model (FEM) is applied to calculate the stress of each node in the RB model. Finally, the statistical Kernel Ridge Regression (KRR) method is implemented to obtain the relationship between the bone displacement and the stress. Therefore, we can predict the soft tissue deformation from the displacement of the facial bone. Cross-validation has been demonstrated and satisfactory performance has been presented. James J. Xia, Xiaoyan Zhang 0002, Xiaobo Zhou 0001 |
ICIP | 3 |
| 2015 | Automated Three-Piece Digital Dental Articulation
Jianfu Li, Flavio Ferraz, Shunyao Shen, Yi-Fang Lo, Xiaoyan Zhang 0002, Peng Yuan 0001, Ken-Chung Chen, Jaime Gateno, Xiaobo Zhou 0001, James J. Xia |
MICCAI (1) | 5 |
| 2014 | Exemplar-Based Portrait Photograph Enhancement as Informed by Portrait PaintingsabstractAbstract This paper proposes an approach to enhance the regional contrasts in snap‐shot style portrait photographs by using pre‐modern portrait paintings as aesthetic exemplars. The example portrait painting is selected based on a comparison of the existing contrast properties of the painting with those of the photograph. The contrast organization in the selected example painting is transferred to the photograph by mapping the inter‐ and intra‐regional contrasts of the regions, such as the face and skin areas of the foreground figure, the non‐face/skin part of the foreground and the background region. A piecewise non‐linear transformation curve is used to achieve the contrast mapping. Finally, the transition boundary between regions is smoothed to achieve the final results. The experimental results and user study demonstrate that, by using this proposed approach, the visual appeal of the portrait photographs is effectively improved, and the face and the figure become more salient. Xiaoyan Zhang 0002, Martin Constable, Kap Luk Chan |
Comput. Graph. Forum | 1 |
| 2014 | Atmospheric Perspective Effect Enhancement of Landscape Photographs Through Depth-Aware Contrast ManipulationabstractThe atmospheric perspective effect is a physical phenomenon relating to the effect that atmosphere has on distant objects, causing them to be lighter and less distinct. The exaggeration of this effect by artists in 2D images increases the illusion of depth thereby making the image more interesting. This paper addresses the enhancement of the atmospheric perspective effect in landscape photographs by the manipulation of depth-aware lightness and saturation contrast values. The form of this manipulation follows the organization of such contrasts in landscape paintings. The objective of this manipulation is based on a statistical study which has clearly shown that the saturation and lightness contrasts inter- and intra- depth planes in paintings are more purposefully organized than those in photographs. This contrast organization in paintings respects the existing contrast relationships within a natural scene guided by the atmospheric perspective effect and also exaggerates them sufficiently with a view to improving the visual appeal of the painting and the illusion of depth within it. The depth-aware lightness and saturation contrasts revealed in landscape paintings guide the mapping of such contrasts in photographs. This contrast mapping is formulated as an optimization problem that simultaneously considers the desired inter-contrast, intra-contrast, and specified gradient constraints. Experimental results demonstrate that by using this proposed approach, both the visual appeal and the illusion of depth in the photographs are effectively improved. Xiaoyan Zhang 0002, Kap Luk Chan, Martin Constable |
IEEE Trans. Multim. | 1 |
| 2012 | Example-based contrast enhancement for portrait photograph
Xiaoyan Zhang 0002, Martin Constable, Kap Luk Chan |
ICPR | 1 |
| 2011 | Aesthetic enhancement of landscape photographs as informed by paintings across depth layersabstractThis paper addresses the aesthetic enhancement of contrasts in saturation and luminance in a landscape photograph to follow the organization of such contrasts in landscape paintings. Different from many existing work on similar topics that mainly emulate the surface characteristics of a painting, this paper presents a technique that transfer the contrast organization in landscape paintings inter and intra four depth layers onto landscape photographs. The contrasts in saturation and luminance revealed in landscape paintings along the scene depth provide the references to guide the mapping of such contrasts in photographs. A novel inter and intra depth layer luminance and saturation contrast mapping algorithm using gradient histogram matching is developed. Experimental results demonstrate the effectiveness of the proposed method. Xiaoyan Zhang 0002, Martin Constable, Kap Luk Chan |
ICIP | 1 |