VLDB 2026 Research / reviewers in the wild / expert
Xiaoliang Zhang 0002
dblp:22/8401-2
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2023
0000-0001-9313-8428ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Contrastive Continuity on Augmentation Stability Rehearsal for Continual Self-Supervised LearningabstractSelf-supervised learning has attracted a lot of attention recently, which is able to learn powerful representations without any manual annotations. However, self-supervised learning needs to develop the ability to continuously learn to cope with a variety of real-world challenges, i.e., Continual Self-Supervised Learning (CSSL). Catastrophic forgetting is a notorious problem in CSSL, where the model tends to forget the learned knowledge. In practice, simple rehearsal or regularization will bring extra negative effects while alleviating catastrophic forgetting in CSSL, e.g., overfitting on the rehearsal samples or hindering the model from encoding fresh information. In order to address catastrophic forgetting without overfitting on the rehearsal samples, we propose Augmentation Stability Rehearsal (ASR) in this paper, which selects the most representative and discriminative samples by estimating the augmentation stability for rehearsal. Meanwhile, we design a matching strategy for ASR to dynamically update the rehearsal buffer. In addition, we further propose Contrastive Continuity on Augmentation Stability Rehearsal (C2ASR) based on ASR. We show that C2ASR is an upper bound of the Information Bottleneck (IB) principle, which suggests that C2ASR essentially preserves as much information shared among seen task streams as possible to prevent catastrophic forgetting and dismisses the redundant information between previous task streams and current task stream to free up the ability to encode fresh information. Our method obtains a great achievement compared with state-of-the-art CSSL methods on a variety of CSSL benchmarks. Haoyang Cheng, Haitao Wen, Xiaoliang Zhang 0002, Heqian Qiu, Lanxiao Wang, Hongliang Li 0001 |
ICCV | 3 |
| 2023 | Disturbed Augmentation Invariance for Unsupervised Visual Representation LearningabstractContrastive learning has gained great prominence recently, which achieves excellent performance by simple augmentation invariance. However, the simple contrastive pairs suffer from lacking of diversity due to the mechanical augmentation strategies. In this paper, we propose Disturbed Augmentation Invariance (DAI for abbreviation), which constructs disturbed contrastive pairs by generating appropriate disturbed views for each augmented view in the feature space to increase the diversity. In practice, we establish a multivariate normal distribution for each augmented view, whose mean is corresponding augmented view and covariance matrix is estimated from its nearest neighbors in the dataset. Then we sample random vectors from this distribution as the disturbed views to construct disturbed contrastive pairs. In order to avoid extra computational cost with the increase of disturbed contrastive pairs, we utilize an upper bound of the trivial disturbed augmentation invariance loss to construct the DAI loss. In addition, we propose Bottleneck version of Disturbed Augmentation Invariance (BDAI for abbreviation) inspired by the Information Bottleneck principle, which further refines the extracted information and learns a compact representation by additionally increasing the variance of the original contrastive pair. In order to make BDAI work effectively, we design a statistical strategy to control the balance between the amount of the information shared by all disturbed contrastive pairs and the compactness of the representation. Our approach gets a consistent improvement over the popular contrastive learning methods on a variety of downstream tasks, e.g. image classification, object detection and instance segmentation. Haoyang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Heqian Qiu, Xiaoliang Zhang 0002, Fanman Meng, Taijin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Unsupervised Visual Representation Learning via Multi-Dimensional Relationship AlignmentabstractRecently, contrastive learning based on augmentation invariance and instance discrimination has made great achievements, owing to its excellent ability to learn beneficial representations without any manual annotations. However, the natural similarity among instances conflicts with instance discrimination which treats each instance as a unique individual. In order to explore the natural relationship among instances and integrate it into contrastive learning, we propose a novel approach in this paper, Relationship Alignment (RA for abbreviation), which forces different augmented views of current batch instances to main a consistent relationship with other instances. In order to perform RA effectively in existing contrastive learning framework, we design an alternating optimization algorithm where the relationship exploration step and alignment step are optimized respectively. In addition, we add an equilibrium constraint for RA to avoid the degenerate solution, and introduce the expansion handler to make it approximately satisfied in practice. In order to better capture the complex relationship among instances, we additionally propose Multi-Dimensional Relationship Alignment (MDRA for abbreviation), which aims to explore the relationship from multiple dimensions. In practice, we decompose the final high-dimensional feature space into a cartesian product of several low-dimensional subspaces and perform RA in each subspace respectively. We validate the effectiveness of our approach on multiple self-supervised learning benchmarks and get consistent improvements compared with current popular contrastive learning methods. On the most commonly used ImageNet linear evaluation protocol, our RA obtains significant improvements over other methods, our MDRA gets further improvements based on RA to achieve the best performance. The source code of our approach will be released soon. Haoyang Cheng, Hongliang Li 0001, Heqian Qiu, Qingbo Wu 0001, Xiaoliang Zhang 0002, Fanman Meng, King Ngi Ngan |
IEEE Trans. Image Process. | 5 |
| 2023 | What Happens in Crowd Scenes: A New Dataset About Crowd Scenes for Image CaptioningabstractMaking machines endowed with eyes and brains to effectively understand and analyze crowd scenes is of paramount importance for building a smart city to serve people. This is of far-reaching significance for the guidance of dense crowds and accident prevention, such as crowding and stampedes. As a typical multimodal scene understanding task, image captioning has always attracted widespread attention. However, crowd scene understanding captioning is rarely studied due to the unobtainability of related datasets. Therefore, it is difficult to know what happens in crowd scenes. In order to fill this research gap, we propose a crowd scenes caption dataset named CrowdCaption which has the advantages of crowd-topic scenes, comprehensive and complex caption descriptions, typical relationships and detailed grounding annotations. The complexity and diversity of the descriptions and the specificity of the crowd scenes make this dataset extremely challenging to most current methods. Thus, we propose a Multi-hierarchical Attribute Guided Crowd Caption Network (MAGC) based on crowd objects, actions, and status (such as position, dress, posture, etc.) aiming to generate crowd-specific detailed descriptions. We conduct extensive experiments on our CrowdCaption dataset, and our proposed method reaches the state-of-the-art (SoTA) performance. We hope the CrowdCaption dataset can assist future studies related to crowd scenes in the multimodal domain. Lanxiao Wang, Hongliang Li 0001, Wenzhe Hu, Xiaoliang Zhang 0002, Heqian Qiu, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Mining Regional Relation from Pixel-wise Annotation for Scene ParsingabstractScene parsing is an important and challenging task in computer vision, which assigns semantic labels to each pixel in the entire scene. Existing scene parsing methods only utilize pixel-wise annotation as the supervision of neural network, thus, some similar categories are easy to be misclassified in the complex scenes without the utilization of regional relation. To tackle these above challenging problems, a Regional Relation Network (RRNet) is proposed in this paper, which aims to boost the scene parsing performance by mining regional relation from pixel-wise annotation. Specifically, the pixel-wise annotation is divided into a lot of fixed regions, so that intra- and inter-regional relation are able to be extracted as the supervision of network. We firstly design an intra-regional relation module to predict category distribution in each fixed region, which is helpful for reducing the misclassification phenomenon in regions. Secondly, an inter-regional relation module is proposed to learn the relationships among each region in scene images. With the guideline of relation information extracted from the ground truth, the network is able to learn more discriminative relation representations. To validate our proposed model, we conduct experiments on three typical datasets, including NYU-depth-v2, PASCAL-Context and ADE20k. The achieved competitive results on all three datasets demonstrate the effectiveness of our method. Zichen Song 0002, Hongliang Li 0001, Heqian Qiu, Xiaoliang Zhang 0002 |
VCIP | 4 |
| 2022 | Real-time panoptic segmentation with relationship between adjacent pixels and boundary prediction
Xiaoliang Zhang 0002, Hongliang Li 0001, Lanxiao Wang, Haoyang Cheng, Heqian Qiu, Wenzhe Hu, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 1 |
| 2022 | Segmenting Beyond the Bounding Box for Instance SegmentationabstractInstance segmentation needs to locate all instances in an image correctly and segment each instance precisely. Currently, the most dominant methods for instance segmentation take object detection as a pre-task. However, they rely on the accuracy of object detection incredibly. If the pre-task cannot predict an accurate bounding box, the performance of instance segmentation will degenerate. In this paper, we present a novel method for instance segmentation to solve this problem, which is calledSegmentingBeyond theBoundingBox (S3B-Net). Our S3B-Net designs a sub-network to help instance segmentation methods based on object detection to segment the part of an instance beyond the bounding box. Specifically, the sub-network first predicts a two-dimensional pixel embedding for each pixel. Then, the Gaussian function is employed to calculate a pixel’s probability belongs to a corresponding instance according to the two-dimensional pixel embedding. Finally, the output of the sub-network combines with the output of instance segmentation based on object detection to generate a more precise instance mask. Our sub-network can easily extend on the existing instance segmentation method based on object detection to segment instance beyond the bounding box. We do our experiments on dominant instance segmentation datasets, such as the COCO dataset and Cityscapes dataset. The results show that our method can achieve 6.8 points gain compared with the baseline Mask R-CNN with ResNet-50-FPN in Cityscapes datasets, and 1.7 points gain with ResNet-101-FPN-DCN in COCO datasets. Our S3B-Net outperforms the previous state-of-the-art instance segmentation method, which proves our method is competitive. The source code of our method will be made available. Xiaoliang Zhang 0002, Hongliang Li 0001, Fanman Meng, Zichen Song 0002, Linfeng Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |