VLDB 2026 Research / reviewers in the wild / expert
Haohang Xu
dblp:254/0948
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-4715-1338ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Token-Wise Feature Caching: Accelerating Diffusion Transformers With Dual Feature CachingabstractDiffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs. As an effective approach for DiT acceleration, feature caching methods are designed to cache the features of DiT in previous timesteps and reuse them in the next timesteps, allowing us to skip the computation in the next timesteps. Among them, token-wise feature caching has been introduced to perform different caching ratios for different tokens in DiTs, aiming to skip the computation for unimportant tokens while still computing the important ones. In this paper, we propose to carefully check the effectiveness in token-wise feature caching with the following two questions: 1) Is it really necessary to compute the so-called "important" tokens in each step? 2) Are so-called important tokens really important? Surprisingly, this paper gives some counter-intuition answers, demonstrating that consistently computing the selected "important tokens" in all steps is not necessary. The selection of the so-called "important tokens" is often ineffective, and even sometimes shows inferior performance than random selection. Based on these observations, this paper introduces dual feature caching referred to as DuCa, which performs aggressive caching strategy and conservative caching strategy iteratively and selects the tokens for computing randomly. Extensive experimental results demonstrate the effectiveness of our method in DiT, PixArt, FLUX, and OpenSora, demonstrating significant improvements than the previous token-wise feature caching. Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | ProReflow: Progressive Reflow with Decomposed VelocityabstractDiffusion models have achieved significant progress in both image and video generation while still suffering from huge computation costs. As an effective solution, rectified flow aims to rectify the diffusion process of diffusion models into a straight line for few-step and even one-step generation. However, in this paper, we suggest that the original training pipeline of reflow is not optimal and introduce two techniques to improve it. Firstly, we introduce progressive reflow, which progressively reflows the diffusion models in local timesteps until the whole diffusion progresses, reducing the difficulty of flow matching. Second, we introduce aligned v-prediction, which highlights the importance of direction matching in flow matching over magnitude matching. Experimental results on SDv1.5 and SDXL demonstrate the effectiveness of our method, for example, conducting on SDv1.5 achieves an FID of 10.70 on MSCOCO2014 validation set with only 4 sampling steps, close to our teacher model (32 DDIM steps, FID = 10.05). Our codes will be released at Github. Lei Ke, Haohang Xu, Xuefei Ning, Yu Li 0022, Haoling Li, Dongsheng Jiang, Yujiu Yang 0001, Linfeng Zhang 0001 |
CVPR | 2 |
| 2024 | Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
Shuangrui Ding, Rui Qian 0001, Haohang Xu, Dahua Lin, Hongkai Xiong |
ECCV (45) | 3 |
| 2023 | Seed the Views: Hierarchical Semantic Alignment for Contrastive Representation LearningabstractSelf-supervised learning based on instance discrimination has shown remarkable progress. In particular, contrastive learning, which regards each image as well as its augmentations as an individual class and tries to distinguish them from all other images, has been verified effective for representation learning. However, conventional contrastive learning does not model the relation between semantically similar samples explicitly. In this paper, we propose a general module that considers the semantic similarity among images. This is achieved by expanding the views generated by a single image to Cross-Samples and Multi-Levels, and modeling the invariance to semantically similar images in a hierarchical way. Specifically, the cross-samples are generated by a data mixing operation, which is constrained within samples that are semantically similar, while the multi-level samples are expanded at the intermediate layers of a network. In this way, the contrastive loss is extended to allow for multiple positives per anchor, and explicitly pulling semantically similar images together at different layers of the network. Our method, termed as CSML, has the ability to integrate multi-level representations across samples in a robust way. CSML is applicable to current contrastive based methods and consistently improves the performance. Notably, using MoCo v2 as an instantiation, CSML achieves 76.6% top-1 accuracy with linear evaluation using ResNet-50 as backbone, 66.7% and 75.1% top-1 accuracy with only 1% and 10% labels, respectively. All these numbers set the new state-of-the-art. The code is available at https://github.com/haohang96/CSML. Haohang Xu, Xiaopeng Zhang 0008, Hao Li 0090, Lingxi Xie, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Semi-Supervised Contrastive Learning With Similarity Co-CalibrationabstractSemi-supervised learning acts as an effective way to leverage massive unlabeled data. In this paper, we propose a novel training strategy, termed asSemi-supervised Contrastive Learning (SsCL), which combines the well-known contrastive loss in self-supervised learning with the cross entropy loss in semi-supervised learning, and jointly optimizes the two objectives in an end-to-end way. The highlight is that different from self-training based semi-supervised learning that conducts prediction and retraining over the same model weights, SsCL interchanges the predictions over the unlabeled data between the two branches, and thus formulates a co-calibration procedure, which we find is beneficial for better prediction and avoids being trapped in local minimum. Towards this goal, the contrastive loss branch models pairwise similarities among samples, using the pseudo labels generated from the cross entropy branch, and in turn calibrates the prediction distribution of the cross entropy branch with the contrastive similarity. We show that SsCL produces more discriminative representation and is beneficial to semi-supervised learning. Notably, on ImageNet with ResNet50 as the backbone, SsCL achieves$\bm {60.2\%}$and$\bm {72.1\%}$top-1 accuracy with 1% and 10% labeled samples respectively, which significantly outperforms the baseline, and is better than previous semi-supervised and self-supervised methods. Yuhang Zhang 0012, Xiaopeng Zhang 0008, Jie Li 0002, Robert C. Qiu, Haohang Xu, Qi Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Motion-aware Contrastive Video Representation Learning via Foreground-background MergingabstractIn light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two augmented views of a video closer, the model however tends to learn the common static background as a shortcut but fails to capture the motion information, a phenomenon dubbed as background bias. Such bias makes the model suffer from weak generalization ability, leading to worse performance on downstream tasks such as action recognition. To alleviate such bias, we propose Foreground-background Merging (FAME) to deliberately compose the moving foreground region of the selected video onto the static background of others. Specifically, without any off-the-shelf detector, we extract the moving fore-ground out of background regions via the frame difference and color statistics, and shuffle the background regions among the videos. By leveraging the semantic consistency between the original clips and the fused ones, the model focuses more on the motion patterns and is debiased from the background shortcut. Extensive experiments demonstrate that FAME can effectively resist background cheating and thus achieve the state-of-the-art performance on downstream tasks across UCF101, HMDB51, and Diving48 datasets. The code and configurations are released at https://github.com/Mark12Ding/FAME. Shuangrui Ding, Maomao Li, Tianyu Yang 0003, Rui Qian 0001, Haohang Xu, Qingyi Chen, Jue Wang 0001, Hongkai Xiong |
CVPR | 5 |
| 2022 | Bag of Instances Aggregation Boosts Self-supervised Distillation
Haohang Xu, Jiemin Fang, Xiaopeng Zhang 0008, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, Qi Tian 0001 |
ICLR | 1 |
| 2022 | $K$K-Shot Contrastive Learning of Visual Features With Multiple Instance AugmentationsabstractIn this paper, we propose the K-Shot Contrastive Learning (KSCL) of visual features by applying multiple augmentations to investigate the sample variations within individual instances. It aims to combine the advantages of inter-instance discrimination by learning discriminative features to distinguish between different instances, as well as intra-instance variations by matching queries against the variants of augmented samples over instances. Particularly, for each instance, it constructs an instance subspace to model the configuration of how the significant factors of variations in K-shot augmentations can be combined to form the variants of augmentations. Given a query, the most relevant variant of instances is then retrieved by projecting the query onto their subspaces to predict the positive instance class. This generalizes the existing contrastive learning that can be viewed as a special one-shot case. An eigenvalue decomposition is performed to configure instance subspaces, and the embedding network can be trained end-to-end through the differentiable subspace configuration. Experiment results demonstrate the proposed K-shot contrastive learning achieves superior performances to the state-of-the-art unsupervised methods. Haohang Xu, Hongkai Xiong, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Auto-Encoding Transformations in Reparameterized Lie Groups for Unsupervised LearningabstractUnsupervised training of deep representations has demonstrated remarkable potentials in mitigating the prohibitive expenses on annotating labeled data recently. Among them is predicting transformations as a pretext task to self-train representations, which has shown great potentials for unsupervised learning. However, existing approaches in this category learn representations by either treating a discrete set of transformations as separate classes, or using the Euclidean distance as the metric to minimize the errors between transformations. None of them has been dedicated to revealing the vital role of the geometry of transformation groups in learning representations. Indeed, an image must continuously transform along the curved manifold of a transformation group rather than through a straight line in the forbidden ambient Euclidean space. This suggests the use of geodesic distance to minimize the errors between the estimated and groundtruth transformations. Particularly, we focus on homographies, a general group of planar transformations containing the Euclidean, similarity and affine transformations as its special cases. To avoid an explicit computing of intractable Riemannian logarithm, we project homographies onto an alternative group of rotation transformations SR(3) with a tractable form of geodesic distance. Experiments demonstrate the proposed approach to Auto-Encoding Transformations exhibits superior performances on a variety of recognition problems. Feng Lin 0009, Haohang Xu, Houqiang Li, Hongkai Xiong, Guo-Jun Qi |
AAAI | 2 |
| 2020 | FedMax: Enabling a Highly-Efficient Federated Learning FrameworkabstractIoT devices produce a wealth of data desired for learning models to empower more intelligent applications. However, such data is often privacy sensitive making data owners reluctant upload their data to a central server for learning purposes. Federated learning provides a promising privacy-preserving learning approach, which decouples the model training from the need of accessing to the sensitive data. However, realizing a deployed, dependable federated learning system faces critical challenges, such as frequent dropouts of learning workers, heterogeneity of workers computation, and limited communication. In this paper, we focus on the systems aspects to advance federated learning and contribute a highly efficient and reliable distributed federated learning framework, FedMax, aiming to tackle these challenges. In designing FedMax, we contribute new techniques in light of the properties of a real federated learning setting, including a relaxed synchronization communication scheme and a similarity-based worker selection approach. We have implemented a prototype of FedMax and evaluated FedMax upon multiple popular machine learning models and datasets, showing that FedMax significantly increases the robustness of a federated learning system, speeds up the convergence rate by 25%, and increases the system efficiency by 50%, in comparison with state-of-the-art approaches. Haohang Xu, Jin Li 0057, Hongkai Xiong, Hui Lu 0001 |
CLOUD | 1 |