EDBT 2026 Demo / reviewers in the wild / expert
Xiangyun Zhao
dblp:172/9876
· DBLP profile ↗
11ranked-venue papers
7as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Image recognition and object detection · 29% Segmentation and scene understanding · 25% Representation and self-supervised learning · 13% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
0.8 | 2 | 2020 | Object Detection with a Unified Label Space from Multiple Datasets · ECCV (14) 2020 Pseudo Mask Augmented Object Detection · CVPR 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 1 | 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video coding
image compression |
0.8 | 1 | 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video coding › image compression › learned image compression
variable-rate image compression |
0.8 | 1 | 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.6 | 2 | 2021 | Contrastive Learning for Label Efficient Semantic Segmentation · ICCV 2021 Pseudo Mask Augmented Object Detection · CVPR 2018 |
Computer vision › Segmentation and scene understanding
annotation-efficient segmentation |
0.5 | 1 | 2021 | Contrastive Learning for Label Efficient Semantic Segmentation · ICCV 2021 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Contrastive Learning for Label Efficient Semantic Segmentation · ICCV 2021 |
Computer vision › Image recognition and object detection › object detection
few-shot object detection |
0.5 | 1 | 2021 | Morphable Detector for Object Detection on Demand · ICCV 2021 |
Machine learning › Representation and self-supervised learning › contrastive learning › dense contrastive learning
pixel-level contrastive learning |
0.5 | 1 | 2021 | Contrastive Learning for Label Efficient Semantic Segmentation · ICCV 2021 |
Computer vision › Face, body and person analysis
face alignment |
0.4 | 1 | 2019 | Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble · ICCV 2019 |
Machine learning › Transfer learning and domain adaptation
zero-shot learning |
0.4 | 1 | 2019 | Recognizing Part Attributes With Insufficient Data · ICCV 2019 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.3 | 1 | 2018 | Pseudo Mask Augmented Object Detection · CVPR 2018 |
Machine learning › Learning paradigms
multi-task learning |
0.3 | 1 | 2018 | A Modulation Module for Multi-task Learning with Applications in Image Retrieval · ECCV (1) 2018 |
Computer vision › Segmentation and scene understanding › instance segmentation
weakly supervised instance segmentation |
0.3 | 1 | 2018 | Pseudo Mask Augmented Object Detection · CVPR 2018 |
Information retrieval
image retrieval |
0.3 | 1 | 2018 | A Modulation Module for Multi-task Learning with Applications in Image Retrieval · ECCV (1) 2018 |
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition |
0.2 | 1 | 2016 | Peak-Piloted Deep Network for Facial Expression Recognition · ECCV (2) 2016 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
0.2 | 1 | 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image Compression · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset learning |
0.1 | 1 | 2020 | Object Detection with a Unified Label Space from Multiple Datasets · ECCV (14) 2020 |
Computer vision › Segmentation and scene understanding › semantic segmentation › weakly supervised semantic segmentation
pseudo mask generation |
0.1 | 1 | 2018 | Pseudo Mask Augmented Object Detection · CVPR 2018 |
Methods — techniques the papers use, named apart from their topics
invertible activation transformation · 1.5prototype learning · 0.5expectation-maximization · 0.5cross-entropy fine-tuning · 0.5contrastive pre-training · 0.5structured ensemble · 0.4heatmap regression · 0.4deep learning · 0.4concept sharing network · 0.4modulation module · 0.3alternating optimization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Skeleton-Based Gait Recognition Based on Deep Neuro-Fuzzy NetworkabstractGait recognition aims to identify users by their walking patterns. Compared with appearance-based methods, skeleton-based methods exhibit well robustness to cluttered backgrounds, carried items, and clothing variations. However, skeleton extraction faces the wrong human tracking and keypoints missing problems, especially under multiperson scenarios. To address above issues, this article proposes a novel gait recognition method using deep neural network specifically designed for multiperson scenarios. The method consists of individual gait separate module (IGSM) and fuzzy skeleton completion network (FU-SCN). To achieve effective human tracking, IGSM employs root–skeleton keypoints predictions and object keypoint similarity (OKS)-based skeleton calculation to separate individual gait sets when multiple persons exist. In addition, keypoints missing renders human poses estimation fuzzy. We propose FU-SCN, a deep neuro-fuzzy network, to enhances the interpretability of the fuzzy pose estimation via generating fine-grained gait representation. FU-SCN utilizes fuzzy bottleneck structure to extract features on low-dimension keypoints, and multiscale fusion to extract dissimilar relations of human body during walking on each scale. Extensive experiments are conducted on the CASIA-B dataset and our multigait dataset. The results show that our method is one of the SOTA methods and shows outperformance under complex scenarios. Compared with PTSN, PoseMapGait, JointsGait, GaitGraph2, and CycleGait, our method achieves an average accuracy improvement of 53.77%, 42.07%, 25.3%, 13.47%, and 9.5%, respectively, and it keeps low time cost with average 180 ms using edge devices. Jiefan Qiu, Yizhe Jia, Xiangyun Zhao, Hailin Feng, Kai Fang 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | I2C: Invertible Continuous Codec for High-Fidelity Variable-Rate Image CompressionabstractLossy image compression is a fundamental technology in media transmission and storage. Variable-rate approaches have recently gained much attention to avoid the usage of a set of different models for compressing images at different rates. During the media sharing, multiple re-encodings with different rates would be inevitably executed. However, existing Variational Autoencoder (VAE)-based approaches would be readily corrupted in such circumstances, resulting in the occurrence of strong artifacts and the destruction of image fidelity. Based on the theoretical findings of preserving image fidelity via invertible transformation, we aim to tackle the issue of high-fidelity fine variable-rate image compression and thus propose the Invertible Continuous Codec (I2C). We implement the I2C in a mathematical invertible manner with the core Invertible Activation Transformation (IAT) module. I2C is constructed upon a single-rate Invertible Neural Network (INN) based model and the quality level (QLevel) would be fed into the IAT to generate scaling and bias tensors. Extensive experiments demonstrate that the proposed I2C method outperforms state-of-the-art variable-rate image compression methods by a large margin, especially after multiple continuous re-encodings with different rates, while having the ability to obtain a very fine variable-rate control without any performance compromise. Shilv Cai, Zhijun Zhang 0009, Xiangyun Zhao, Jiahuan Zhou, Yuxin Peng 0001, Luxin Yan, Sheng Zhong 0001, Xu Zou 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Contrastive Learning for Label Efficient Semantic SegmentationabstractCollecting labeled data for the task of semantic segmentation is expensive and time-consuming, as it requires dense pixel-level annotations. While recent Convolutional Neural Network (CNN) based semantic segmentation approaches have achieved impressive results by using large amounts of labeled training data, their performance drops significantly as the amount of labeled data decreases. This happens because deep CNNs trained with the de facto cross-entropy loss can easily overfit to small amounts of labeled data. To address this issue, we propose a simple and effective contrastive learning-based training strategy in which we first pretrain the network using a pixel-wise, label-based contrastive loss, and then fine-tune it using the cross-entropy loss. This approach increases intra-class compactness and inter-class separability, thereby resulting in a better pixel classifier. We demonstrate the effectiveness of the proposed training strategy using the Cityscapes and PASCAL VOC 2012 segmentation datasets. Our results show that pretraining with the proposed contrastive loss results in large performance gains (more than 20% absolute improvement in some settings) when the amount of labeled data is limited. In many settings, the proposed contrastive pretraining strategy, which does not use any additional data, is able to match or outperform the widely-used ImageNet pretraining strategy that uses more than a million additional labeled images. Xiangyun Zhao, Raviteja Vemulapalli, Philip Andrew Mansfield, Boqing Gong, Bradley Green, Lior Shapira, Ying Wu 0001 |
ICCV | 1 |
| 2021 | Morphable Detector for Object Detection on DemandabstractMany emerging applications of intelligent robots need to explore and understand new environments, where it is desirable to detect objects of novel classes on the fly with minimum online efforts. This is an object detection on demand (ODOD) task. It is challenging, because it is impossible to annotate a large number of data on the fly, and the embedded systems are usually unable to perform back-propagation which is essential for training. Most existing few-shot detection methods are confronted here as they need extra training. We propose a novel morphable detector (MD), that simply "morphs" some of its changeable parameters online estimated from the few samples, so as to detect novel classes without any extra training. The MD has two sets of parameters, one for the feature embedding and the other for class representation (called "prototypes"). Each class is associated with a hidden prototype to be learned by integrating the visual and semantic embeddings. The learning of the MD is based on the alternate learning of the feature embedding and the prototypes in an EM-like approach which allows the recovery of an unknown prototype from a few samples of a novel class. Once an MD is learned, it is able to use a few samples of a novel class to directly compute its prototype to fulfill the online morphing process. We have shown the superiority of the MD in Pascal [12], COCO [27] and FSOD [13] datasets. Xiangyun Zhao, Xu Zou 0002, Ying Wu 0001 |
ICCV | 1 |
| 2020 | Object Detection with a Unified Label Space from Multiple Datasets
Xiangyun Zhao, Samuel Schulter, Yi-Hsuan Tsai, Manmohan Krishna Chandraker, Ying Wu 0001 |
ECCV (14) | 1 |
| 2019 | Recognizing Part Attributes With Insufficient DataabstractRecognizing the attributes of objects and their parts is central to many computer vision applications. Although great progress has been made to apply object-level recognition, recognizing the attributes of parts remains less applicable since the training data for part attributes recognition is usually scarce especially for internet-scale applications. Furthermore, most existing part attribute recognition methods rely on the part annotations which are more expensive to obtain. In order to solve the data insufficiency problem and get rid of dependence on the part annotation, we introduce a novel Concept Sharing Network (CSN) for part attribute recognition. A great advantage of CSN is its capability of recognizing the part attribute (a combination of part location and appearance pattern) that has insufficient or zero training data, by learning the part location and appearance pattern respectively from the training data that usually mix them in a single label. Extensive experiments on CUB, Celeb A, and a newly proposed human attribute dataset demonstrate the effectiveness of CSN and its advantages over other methods, especially for the attributes with few training samples. Further experiments show that CSN can also perform zero-shot part attribute recognition. Xiangyun Zhao, Yi Yang 0007, Feng Zhou 0002, Xiao Tan 0001, Yuchen Yuan, Sid Ying-Ze Bao, Ying Wu 0001 |
ICCV | 1 |
| 2019 | Learning Robust Facial Landmark Detection via Hierarchical Structured EnsembleabstractHeatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed to tackle this issue, they all heavily rely on manually designed tree structures. The designed hierarchical structure is likely to be completely corrupted due to the missing or inaccurate prediction of landmarks. To the best of our knowledge, in the context of deep learning, no work before has investigated how to automatically model proper structures for facial landmarks, by discovering their inherent relations. In this paper, we propose a novel Hierarchical Structured Landmark Ensemble (HSLE) model for learning robust facial landmark detection, by using it as the structural constraints. Different from existing approaches of manually designing structures, our proposed HSLE model is constructed automatically via discovering the most robust patterns so HSLE has the ability to robustly depict both local and holistic landmark structures simultaneously. Our proposed HSLE can be readily plugged into any existing facial landmark detection baselines for further performance improvement. Extensive experimental results demonstrate our approach significantly outperforms the baseline by a large margin to achieve a state-of-the-art performance. Xu Zou 0002, Sheng Zhong 0001, Luxin Yan, Xiangyun Zhao, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 4 |
| 2018 | Pseudo Mask Augmented Object DetectionabstractIn this work, we present a novel and effective framework to facilitate object detection with the instance-level segmentation information that is only supervised by bounding box annotation. Starting from the joint object detection and instance segmentation network, we propose to recursively estimate the pseudo ground-truth object masks from the instance-level object segmentation network training, and then enhance the detection network with top-down segmentation feedbacks. The pseudo ground truth mask and network parameters are optimized alternatively to mutually benefit each other. To obtain the promising pseudo masks in each iteration, we embed a graphical inference that incorporates the low-level image appearance consistency and the bounding box annotations to refine the segmentation masks predicted by the segmentation network. Our approach progressively improves the object detection performance by incorporating the detailed pixel-wise information learned from the weakly-supervised segmentation network. Extensive evaluation on the detection task in PASCAL VOC 2007 and 2012 [12] verifies that the proposed approach is effective. Xiangyun Zhao, Shuang Liang 0001 |
CVPR | 1 |
| 2018 | A Modulation Module for Multi-task Learning with Applications in Image Retrieval
Xiangyun Zhao, Xiaohui Shen, Xiaodan Liang, Ying Wu 0001 |
ECCV (1) | 1 |
| 2016 | Peak-Piloted Deep Network for Facial Expression Recognition
Xiangyun Zhao, Xiaodan Liang, Luoqi Liu, Teng Li 0001, Yugang Han, Nuno Vasconcelos, Shuicheng Yan |
ECCV (2) | 1 |
| 2015 | Single underwater image enhancement using depth estimation based on blurrinessabstractIn this paper, we propose to use image blurriness to estimate the depth map for underwater image enhancement. It is based on the observation that objects farther from the camera are more blurry for underwater images. Adopting image blurriness with the image formation model (IFM), we can estimate the distance between scene points and the camera and thereby recover and enhance underwater images. Experimental results on enhancing such images in different lighting conditions demonstrate the proposed method performs better than other IFM-based enhancement methods. Yan-Tsung Peng, Xiangyun Zhao, Pamela C. Cosman |
ICIP | 2 |