VLDB 2026 Research / reviewers in the wild / expert
Guochen Xie
dblp:241/5497
· DBLP profile ↗
13ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-6494-6362ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contrastive Learning With Multiple Prototypes for Unsupervised Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptive semantic segmentation aims to transfer knowledge from the annotated source domain to the unlabeled target domain. Recently, self-training methods have gained substantial attention, which leverage high-confidence predictions in the target domain as pseudo labels for supervision. However, limited exploration of intra-class variations across domains, including significant visual differences within each category, has led to misalignment between feature distribution across domains. In this article, we present a unified non-parametric distance-based online clustering method to efficiently maintain multiple centroid-based prototypes within each category subspace instead of one prototype for each category subspace, which enables prototypes to possess the capacity for richer feature representation. Then, considering the variance across different dimensions of a feature representation, we then extend the prototypes from centroid-based ones to distribution-based ones. Specifically, each subspace is modeled using a Gaussian mixture model which includes several anisotropic Gaussian distributions, aimed at prioritizing discriminative dimensions and obtaining a finer measurement of the pixel-to-prototype similarity. Meanwhile, a category-aware feature space is achieved through pixel-to-prototype contrastive learning to ensure the compactness of pixel features in the same subcategory and drive the separation between pixel features of different subcategories. What's more, multi-resolution features are utilized to promote diversity and robustness among intra-class prototypes. Experiments validate the competitiveness of our two prototype-based methods against existing state-of-the-art methods, with a mIoU of 76.8% on GTA$\rightarrow$Cityscapes, 68.4% on Synthia$\rightarrow$Cityscapes, 54.5% on Cityscapes$\rightarrow$DarkZurich and 56.4% on Cityscapes$\rightarrow$ACDC. Notably, our method is able to seamlessly integrate with existing UDA methods. Jun Yu 0001, Guochen Xie, Quansheng Liu, Zhen Kan, Lei Wang 0203, Qiang Ling 0001, Fang Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Adaptive task recommendation based on reinforcement learning in mobile crowd sensing
Guisong Yang, Guochen Xie, Yunhuai Liu |
Appl. Intell. | 2 |
| 2023 | Leveraging the Latent Diffusion Models for Offline Facial Multiple Appropriate Reactions GenerationabstractOffline Multiple Appropriate Facial Reaction Generation (OMAFRG) aims to predict the reaction of different listeners given a speaker, which is useful in the senario of human-computer interaction and social media analysis. In recent years, the Offline Facial Reactions Generation (OFRG) task has been explored in different ways. However, most studies only focus on the deterministic reaction of the listeners. The research of the non-deterministic (i.e. OMAFRG) always lacks of sufficient attention and the results are far from satisfactory. Compared with the deterministic OFRG tasks, the OMAFRG task is closer to the true circumstance but corresponds to higher difficulty for its requirement of modeling stochasticity and context. In this paper, we propose a new model named FRDiff to tackle this issue. Our model is developed based on the diffusion model architecture with some modification to enhance its ability of aggregating the context features. And the inherent property of stochasticity in diffusion model enables our model to generate multiple reactions. We conduct experiments on the datasets provided by the ACM Multimedia REACT2023 and obtain the second place on the board, which demonstrates the effectiveness of our method. Jun Yu 0001, Ji Zhao 0020, Guochen Xie, Fengxin Chen, Minglei Li 0001, Zonghong Dai |
ACM Multimedia | 3 |
| 2023 | A viable framework for semi-supervised learning on realistic dataset
Guochen Xie, Jun Yu 0001, Qiang Ling 0001, Fang Gao 0001 |
Mach. Learn. | 2 |
| 2023 | Seamless Reconstruction of AMSR-E Land Surface Temperature Swath Gaps for China's LandmassabstractAll-weather Land Surface Temperature (LST) derived from passive microwave (PMW) sensors has significant implications for characterization on the physical processes of surface energy and water balance at local through global scales. However, the PMW sensors (e.g., the AMSR-E) suffer from swath gaps, cannot provide completely spatial-gapless observations. The existing Multi-temporal Feature Connection-CNN (MTFC-CNN) method caused obvious traces of ‘gaps’ when the sample number is small or features are not rich. This paper proposes a Sample Optimized-MTFC (SO-MTFC) seamless reconstruction method based on analyzing the periodicity and complementarity of AMSR-E swath gaps. Sample optimization includes two aspects: sample enhancement and single-cycle mask strategy. Taking China’s landmass as the study area, experimental results show that the original AMSR-E LSTs and the reconstructed AMSR-E LSTs are basically connected seamlessly. Validation against with the MODIS LSTs show that the daytime (nighttime) RMSEs of the original and the reconstructed AMSR-E LSTs are 3.87 K (2.57 K) and 4.76 K (2.96 K), respectively; while the corresponding daytime (nighttime) R2are 0.88 (0.94) and 0.73 (0.90), respectively. Validation against with the six in-situ LSTs show that the RMSEs and R2of reconstructed AMSR-E LSTs against in-situ LSTs are almost consistent with those of the original AMSR-E LST. The ablation study proves the effectiveness of the sample optimization. These findings indicated the SO-MTFC achieved a good reconstruction effect. Compared with the MTFC-CNN, the SO-MTFC got higher scores in visually and quantitatively. The SO-MTFC can potentially be implemented with other satellite PMW sensors to produce completely spatial-seamless PMW LST records on a global scale. Xiaohan Huang 0010, Chan Li, Biao Cao, Jie Cheng 0001, Guochen Xie, Penghai Wu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Facial Expression Spotting Based on Optical Flow FeaturesabstractThe purpose of micro expression (ME) and macro expression (MaE) spotting task is to locate the onset and offset frames of MaE and ME clips. Compared with MaEs, MEs are shorter in duration and lower in intensity, which makes MEs harder to be spotted. In this paper, we propose an efficient pipeline based on optical flow features to spot MEs and MaEs. We crop and align the faces and select the eyebrows area, nose area, and mouth area as our regions of interest to exclude the interference of extraneous factors on the face expression representation. Then, we extract the optical flow in these regions and enhance the features of the expressions in the optical flow with low-pass filter and EMD method. Finally, the sliding window method is used to locate the peaks of optical flow features and get the intervals containing MEs or MaEs. We evaluate the performance of our method on the MEGC2022-TestSet including 10 long videos from SAMM and CAS(ME)3 and achieve the first place in the MEGC2022 Challenge. The results prove the effectiveness of our method. Jun Yu 0001, Zhongpeng Cai, Guochen Xie, Peng He 0004 |
ACM Multimedia | 4 |
| 2022 | Micro Expression Generation with Thin-plate Spline Motion Model and Face ParsingabstractMicro-expression generation aims at transfering the expression from the driving videos to the source images, which can be viewed as a motion transfer task. Recently, several works have been proposed to tackle this problem and achieve great performance. However, due to the intrinsic complexity of the face motion and different attributes of face regions, the task still remains challenging. In this paper, we propose an end-to-end unsupervised motion transfer network to tackle this challenge. As the motion of the face is non-rigid, we adopt an effective and flexible thin-plate spline motion estimation method to estimate the optical flow of the face motion. What's more, we find that several faces with eyeglasses show weird deformation in motion transfering. Thus, we introduce face parsing method to pay specific attention to the eyeglasses regions to ensure the reasonability of the deformation. We conduct several experiments on the provided datasets of the ACM MM 2022 micro-expression grand challenge (MEGC2022) and compare our method with several other typical methods. In comparison, our method shows the best performance. We (Team: USTC-IAT-United) also compare our method with other competitors' in MEGC2022, and the expert evaluation results show that our method performs best, which verifies the effectiveness of our method. Our code is available at https://github.com/HowToNameMe/micro-expression Jun Yu 0001, Guochen Xie, Zhongpeng Cai, Peng He 0004, Fang Gao 0001, Qiang Ling 0001 |
ACM Multimedia | 2 |
| 2021 | Deep Kinship Verification and Retrieval Based on Fusion Siamese Neural NetworkabstractAutomatic kinship analysis, which aims to judge the kinship of different individuals, has been widely used in many real world applications such as helping missing persons reunite with their families and social media analysis. In this work, we focus on three practical and challenging tasks related to kinship analysis, i.e., kinship verification, tri-subject kinship verification and kinship retrieval. A deep fusion Siamese neural network is proposed to address these tasks in a flexible and progressive manner. Firstly, we propose a basic deep Siamese neural network for kinship verification to judge the kinship between individuals based on face images. More specifically, the Siamese neural network takes two input face images and then outputs the similarity between them. To improve the performance, a jury system is also introduced for multi-model fusion. Secondly, we integrate two basic deep Siamese neural networks for tri-subject kinship verification(father, mother and child), which is intended to decide whether a child is related to a pair of parents or not. Specifically, the kinship similairty score of the triplet for verification is obtained by weighting the similarity scores of the father-child and mother-child ones. Thirdly, the proposed deep Siamese neural network can be used to quantify the similarity between any two persons. Thus, it is natural and easy to extend its application to the kinship retrieval task by sorting the similarities between the candidates and faces in the database. We conduct experiments on the RFIW2021 dataset, and final results validate the effectiveness of our solution. Jun Yu 0001, Guochen Xie, Xinlong Hao, Zeyu Cui, Zhongpeng Cai |
FG | 2 |
| 2020 | Retrieval of Family Members Using Siamese Neural NetworkabstractRetrieval of family members in the wild aims at finding family members of the given subject in the dataset, which is useful in finding the lost children and analyzing the kinship. However, due to the diversity in age, gender, pose and illumination of the collected data, this task is always challenging. To solve this problem, we propose our solution with deep Siamese neural network. Our solution can be divided into two parts: similarity computation and ranking. In training procedure, the Siamese network firstly takes two candidate images as input and produces two feature vectors. And then, the similarity between the two vectors is computed with several fully connected layers. While in inference procedure, we try another similarity computing method by dropping the followed several fully connected layers and directly computing the cosine similarity of the two feature vectors. After similarity computation, we use the ranking algorithm to merge the similarity scores with the same identity and output the ordered list according to their similarities. To gain further improvement, we try different combinations of backbones, training methods and similarity computing methods. Finally, we submit the best combination as our solution and our team(ustc-nelslip) obtains favorable result in the track3 of the RFIW2020 challenge with the first runner-up, which verifies the effectiveness of our method. Our code is available at: https://github.com/gniknoil/FG2020-kinship. Jun Yu 0001, Guochen Xie, Xinlong Hao |
FG | 2 |
| 2020 | Deep Fusion Siamese Network for Automatic Kinship VerificationabstractAutomatic kinship verification aims to determine whether some individuals belong to the same family. It is of great research significance to help missing persons reunite with their families. In this Work, the challenging problem is progressively addressed in two respects. First, we propose a deep siamese network to quantify the relative similarity between two individuals. When given two input face images, the deep siamese network extracts the features from them and fuses these features by combining and concatenating. Then, the fused features are fed into a fully-connected network to obtain the similarity score between two faces, which is used to verify the kinship. To improve the performance, a jury system is also employed for multi-model fusion. Second, two deep siamese networks are integrated into a deep triplet network for tri-subject (i.e., father, mother and child) kinship verification, which is intended to decide whether a child is related to a pair of parents or not. Specifically, the obtained similarity scores of father-child and mother-child are weighted to generate the parent-child similarity score for kinship verification. Recognizing Families In the Wild (RFIW) is a challenging kinship recognition task with multiple tracks, which is based on Families in the Wild (FIW), a large-scale and comprehensive image database for automatic kinship recognition. The Kinship Verification (track I) and Tri-Subject Verification (track II) are supported during the ongoing RFIW2020 Challenge. Our team (ustc-nelslip) ranked 1st in track II, and 3rd in track L The code is available at https://github.com/gniknoil/FG2020-kinship. Jun Yu 0001, Xinlong Hao, Guochen Xie |
FG | 4 |
| 2020 | Attention Based Beauty Product Retrieval Using Global and Local DescriptorsabstractBeauty product retrieval has drawn more and more attention for its wide application outlook and enormous economic benefits. However, this task is always challenging due to the variation of products, especially the disturbance of clustered background. In this paper, we first introduce attention mechanism into a global image descriptor, i.e., Maximum Activation of Convolutions (MAC), and propose Attention-based MAC (AMAC). With this enhancement, we can suppress the negative effect of background and highlight the foreground in an unsupervised manner. Then, AMAC and local descriptors are ensembled to complementarily increase the performance. Furthermore, we try to finetune multiple retrieval methods on the different datasets and adopt a query expansion strategy to obtain more improvements. Extensive experiments conducted on a dataset containing more the half million beauty products (Perfect-500K) demonstrate the effectiveness of the proposed method. Finally, our team (USTC-NELSLIP) wins the first place on the leaderboard of the 'AI Meets Beauty'Grand Challenge of ACM Multimedia 2020. The code is available at: https://github.com/gniknoil/Perfect500K-Beauty-Product-Retrieval-Challenge. Jun Yu 0001, Guochen Xie, Haonian Xie, Xinlong Hao, Fang Gao 0001, Feng Shuang 0002 |
ACM Multimedia | 2 |
| 2020 | A Deep Learning Approach for Face Hallucination Guided by Facial Boundary ResponsesabstractFace hallucination is a domain-specific super-resolution (SR) problem of learning a mapping between a low-resolution (LR) face image and its corresponding high-resolution (HR) image. Tremendous progress on deep learning has shown exciting potential for a variety of face hallucination tasks. However, most deep-learning–based methods are limited to handle facial appearance information without paying attention to facial structure priors. In this article, we propose an open source 1 Boundary-aware Dual-branch Network (BDN) for face hallucination, which simultaneously extracts face features and estimates facial boundary responses from LR inputs, ultimately fusing them to reconstruct HR results. Specifically, we first upsample LR face images to HR feature maps, and then feed the upsampled HR features into a memory unit and an attention unit synchronously to obtain the refined features and predict facial boundary responses. Next, they are fed into a feature map fusion unit to combine facial appearance and structure information by a spatial attention mechanism. Moreover, we employ a series of stacked units to boost performance before recovering HR face images. Finally, a discriminative network is developed to improve visual quality by introducing adversarial learning strategy. Extensive experiments show that the proposed approach achieves superior face hallucination results against the state-of-the-art ones. Zhaoyu Zhang 0001, Guochen Xie, Jun Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Beauty Product Retrieval Based on Regional Maximum Activation of Convolutions with Generalized AttentionabstractBeauty and Personal care product retrieval has attracted more and more research attention for its value in real life. However, suffering from data variants and complex background, this task has been very challenging. In this paper, we propose a novel Generalized-attention Regional Maximal Activation of Convolutions (GRMAC) descriptor which helps to generate image features for retrieval. This method introduces attention mechanism to reduce the influence of clustered background and highlight the target, and thus contributes to enhancing the effectiveness of features and boosting the retrieval performance. Different from other attention-based methods, our method supports adjusting mask with a hyperparameter p, which is more flexible and accurate in real application. To demonstrate its effectiveness, we conduct experiments on the dataset containing more than half million personal care products (Perfect-500K) and obtain remarkable results. Furthermore, we try to fuse multiple features from different models for more improvements. And finally, our team (USTC_NELSLIP) ranked 1st in the Grand Challenge of AI Meets Beauty in ACM Multimedia 2019 with a MAP score of 0.408614. Our code is available at: https://github.com/gniknoil/Perfect500K-Beauty-and-Personal-Care-Products-Retrieval-Challenge Jun Yu 0001, Guochen Xie, Haonian Xie, Lingyun Yu 0002 |
ACM Multimedia | 2 |