VLDB 2026 Research / reviewers in the wild / expert
Chuhui Xue
dblp:223/4745
· DBLP profile ↗
16ranked-venue papers
7as first author
12since 2021 · last 2024
0000-0002-3562-3094ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DragDiffusion: Harnessing Diffusion Models for Interactive Point-Based Image EditingabstractAccurate and controllable image editing is a challenging task that has attracted significant attention recently. Notably, DRAGGAN developed by Pan et al. (2023) [33] is an interactive point-based image editing framework that achieves impressive editing results with pixel-level precision. However, due to its reliance on generative adversarial networks (GANs), its generality is limited by the capacity of pretrained GAN models. In this work, we extend this editing framework to diffusion models and propose a novel approach Dragdiffusion. By harnessing large-scale pretrained diffusion models, we greatly enhance the applicability of interactive point-based editing on both real and diffusion-generated images. Unlike other diffusion-based editing methods that provide guidance on diffusion latents of multiple time steps, our approach achieves efficient yet accurate spatial control by optimizing the latent of only one time step. This novel design is motivated by our observations that UNet features at a specific time step provides sufficient semantic and geometric information to support the drag-based editing. Moreover, we introduce two additional techniques, namely identity-preserving fine-tuning and reference-latent-control, to further preserve the identity of the original image. Lastly, we present a challenging benchmark dataset called DRAGBENCH─ the first benchmark to evaluate the performance of interactive point-based image editing methods. Experiments across a wide range of challenging cases (e.g., images with multiple objects, diverse object categories, various styles, etc.) demonstrate the versatility and generality of Dragdiffusion. Code and the Dragbench dataset: https://github.com/Yujun-Shi/DragDiffusion. Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Hanshu Yan, Vincent Y. F. Tan, Song Bai 0001 |
CVPR | 2 |
| 2024 | Free-ATM: Harnessing Free Attention Masks for Representation Learning on Diffusion-Generated Images
Junhao Zhang 0001, Mutian Xu, Jay Zhangjie Wu, Chuhui Xue, Xiaoguang Han 0001, Song Bai 0001, Zheng Shou 0001 |
ECCV (40) | 4 |
| 2024 | Lowis3D: Language-Driven Open-World Instance-Level 3D Scene UnderstandingabstractOpen-world instance-level scene understanding aims to locate and recognize unseen object categories that are not present in the annotated dataset. This task is challenging because the model needs to both localize novel 3D objects and infer their semantic categories. A key factor for the recent progress in 2D open-world perception is the availability of large-scale image-text pairs from the Internet, which cover a wide range of vocabulary concepts. However, this success is hard to replicate in 3D scenarios due to the scarcity of 3D-text pairs. To address this challenge, we propose to harness pre-trained vision-language (VL) foundation models that encode extensive knowledge from image-text pairs to generate captions for multi-view images of 3D scenes. This allows us to establish explicit associations between 3D shapes and semantic-rich captions. Moreover, to enhance the fine-grained visual-semantic representation learning from captions for object-level categorization, we design hierarchical point-caption association methods to learn semantic-aware embeddings that exploit the 3D geometry between 3D points and multi-view images. In addition, to tackle the localization challenge for novel classes in the open-world setting, we develop debiased instance localization, which involves training object grouping modules on unlabeled data using instance-level pseudo supervision. This significantly improves the generalization capabilities of instance grouping and, thus, the ability to accurately locate novel objects. We conduct extensive experiments on 3D semantic, instance, and panoptic segmentation tasks, covering indoor and outdoor scenes across three datasets. Our method outperforms baseline methods by a significant margin in semantic segmentation (e.g. 34.5% ∼ 65.3%), instance segmentation (e.g. 21.8% ∼ 54.0%), and panoptic segmentation (e.g. 14.7% ∼ 43.3%). Code will be available. Runyu Ding, Jihan Yang, Chuhui Xue, Song Bai 0001, Xiaojuan Qi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Understanding and Mitigating Dimensional Collapse in Federated LearningabstractFederated learning aims to train models collaboratively across different clients without sharing data for privacy considerations. However, one major challenge for this learning paradigm is thedata heterogeneityproblem, which refers to the discrepancies between the local data distributions among various clients. To tackle this problem, we first study how data heterogeneity affects the representations of the globally aggregated models. Interestingly, we find that heterogeneous data results in the global model suffering from severedimensional collapse, in which representations tend to reside in a lower-dimensional space instead of the ambient space. This dimensional collapse phenomenon severely curtails the expressive power of models, leading to significant degradation in the performance. Next, via experiments, we make more observations and posit two reasons that result in this phenomenon: 1) dimensional collapse on local models; 2) the operation of global averaging on local model parameters. In addition, we theoretically analyze the gradient flow dynamics to shed light on how data heterogeneity result in dimensional collapse. To remedy this problem caused by the data heterogeneity, we proposeFedDecorr, a novel method that can effectively mitigate dimensional collapse in federated learning. Specifically,FedDecorrapplies a regularization term during local training that encourages different dimensions of representations to be uncorrelated.FedDecorr, which is implementation-friendly and computationally-efficient, yields consistent improvements over various baselines on five standard benchmark datasets including CIFAR10, CIFAR100, TinyImageNet, Office-Caltech10, and DomainNet. Yujun Shi, Jian Liang 0001, Chuhui Xue, Vincent Y. F. Tan, Song Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | PLA: Language-Driven Open-Vocabulary 3D Scene UnderstandingabstractOpen-vocabulary scene understanding aims to localize and recognize unseen categories beyond the annotated label space. The recent breakthrough of 2D open-vocabulary perception is largely driven by Internet-scale paired image-text data with rich vocabulary concepts. However, this success cannot be directly transferred to 3D scenarios due to the inaccessibility of large-scale 3D-text pairs. To this end, we propose to distill knowledge encoded in pretrained vision-language (VL) foundation models through captioning multi-view images from 3D, which allows explicitly associating 3D and semantic-rich captions. Further, to foster coarse-to-fine visual-semantic representation learning from captions, we design hierarchical 3D-caption pairs, leveraging geometric constraints between 3D scenes and multi-view images. Finally, by employing contrastive learning, the model learns language-aware embeddings that connect 3D and text for open-vocabulary tasks. Our method not only remarkably outperforms baseline methods by 25.8% ~ 44.7% hIoU and 14.5% ~ 50.4% hAP50in open-vocabulary semantic and instance segmentation, but also shows robust transferability on challenging zero-shot domain transfer tasks. See the project website at https://dingry.github.io/projects/PLA. Runyu Ding, Jihan Yang, Chuhui Xue, Song Bai 0001, Xiaojuan Qi 0001 |
CVPR | 3 |
| 2023 | Is Synthetic Data from Generative Models Ready for Image Recognition?
Ruifei He, Shuyang Sun, Xin Yu 0004, Chuhui Xue, Philip Torr 0001, Song Bai 0001, Xiaojuan Qi 0001 |
ICLR | 4 |
| 2023 | Mixed Samples as Probes for Unsupervised Model Selection in Domain AdaptationabstractUnsupervised domain adaptation (UDA) has been widely applied in improving model generalization on unlabeled target data. However, accurately selecting the best UDA model for the target domain is challenging due to the absence of labeled target data and domain distribution shifts. Traditional model selection approaches involve training extra models with source data to estimate the target validation risk. Recent studies propose practical methods that are based on measuring various properties of model predictions on target data. Although effective for some UDA models, these methods often lack stability and may lead to poor selections for other UDA models.
In this paper, we present MixVal, an innovative model selection method that operates solely with unlabeled target data during inference. MixVal leverages mixed target samples with pseudo labels to directly probe the learned target structure by each UDA model. Specifically, MixVal employs two distinct types of probes: the intra-cluster mixed samples for evaluating neighborhood density and the inter-cluster mixed samples for investigating the classification boundary. With this comprehensive probing strategy, MixVal elegantly combines the strengths of two state-of-the-art model selection methods, Entropy and SND. We extensively evaluate MixVal on 11 UDA methods across 4 adaptation settings, including classification and segmentation tasks. Experimental results consistently demonstrate that MixVal achieves state-of-the-art performance and maintains exceptional stability in model selection.
Code is available at \url{https://github.com/LHXXHB/MixVal}. Dapeng Hu, Jian Liang 0001, Jun Hao Liew, Chuhui Xue, Song Bai 0001, Xinchao Wang |
NeurIPS | 4 |
| 2023 | Image-to-Character-to-Word Transformers for Accurate Scene Text RecognitionabstractLeveraging the advances of natural language processing, most recent scene text recognizers adopt an encoder-decoder architecture where text images are first converted to representative features and then a sequence of characters via 'sequential decoding'. However, scene text images suffer from rich noises of different sources such as complex background and geometric distortions which often confuse the decoder and lead to incorrect alignment of visual features at noisy decoding time steps. This paper presents I2C2W, a novel scene text recognition technique that is tolerant to geometric and photometric degradation by decomposing scene text recognition into two inter-connected tasks. The first task focuses on image-to-character (I2C) mapping which detects a set of character candidates from images based on different alignments of visual features in an non-sequential way. The second task tackles character-to-word (C2W) mapping which recognizes scene text by decoding words from the detected character candidates. The direct learning from character semantics (instead of noisy image features) corrects falsely detected character candidates effectively which improves the final text recognition accuracy greatly. Extensive experiments over nine public datasets show that the proposed I2C2W outperforms the state-of-the-art by large margins for challenging scene text datasets with various curvature and perspective distortions. It also achieves very competitive recognition performance over multiple normal scene text datasets. Chuhui Xue, Jiaxing Huang 0001, Shijian Lu, Changhu Wang, Song Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Fourier Document Restoration for Robust Document Dewarping and RecognitionabstractState-of-the-art document dewarping techniques learn to predict 3-dimensional information of documents which are prone to errors while dealing with documents with irregular distortions or large variations in depth. This paper presents FDRNet, a Fourier Document Restoration Network that can restore documents with different distortions and improve document recognition in a reliable and simpler manner. FDRNet focuses on high-frequency components in the Fourier space that capture most structural information but are largely free of degradation in appearance. It dewarps documents by a flexible Thin-Plate Spline transformation which can handle various deformations effectively without requiring deformation annotations in training. These features allow FDRNet to learn from a small amount of simply labeled training images, and the learned model can dewarp documents with complex geometric distortion and recognize the restored texts accurately. To facilitate document restoration research, we create a benchmark dataset consisting of over one thousand camera documents with different types of geometric and photometric distortion. Extensive experiments show that FDRNet outperforms the state-of-the-art by large margins on both dewarping and text recognition tasks. In addition, FDRNet requires a small amount of simply labeled training data and is easy to deploy. The proposed dataset is available at https://sg-vilab.github.io/event/warpdoc/. Chuhui Xue, Zichen Tian, Fangneng Zhan, Shijian Lu, Song Bai 0001 |
CVPR | 1 |
| 2022 | Contextual Text Block Detection Towards Scene Text Understanding
Chuhui Xue, Jiaxing Huang 0001, Shijian Lu, Changhu Wang, Song Bai 0001 |
ECCV (28) | 1 |
| 2022 | Language Matters: A Weakly Supervised Vision-Language Pre-training Approach for Scene Text Detection and Spotting
Chuhui Xue, Shijian Lu, Philip Torr 0001, Song Bai 0001 |
ECCV (28) | 1 |
| 2022 | Detection and rectification of arbitrary shaped scene texts by using text keypoints and links
Chuhui Xue, Shijian Lu, Steven C. H. Hoi |
Pattern Recognit. | 1 |
| 2019 | GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and RecognitionabstractRecent adversarial learning research has achieved very impressive progress for modelling cross-domain data shifts in appearance space but its counterpart in modelling cross-domain shifts in geometry space lags far behind. This paper presents an innovative Geometry-Aware Domain Adaptation Network (GA-DAN) that is capable of modelling cross-domain shifts concurrently in both geometry space and appearance space and realistically converting images across domains with very different characteristics. In the proposed GA-DAN, a novel multi-modal spatial learning structure is designed which can convert a source-domain image into multiple images of different spatial views as in the target domain. A new disentangled cycle-consistency loss is introduced which balances the cycle consistency and greatly improves the concurrent learning in both appearance and geometry spaces. The proposed GA-DAN has been evaluated for the classic scene text detection and recognition tasks, and experiments show that the domain-adapted images achieve superior scene text detection and recognition performance while applied to network training. Fangneng Zhan, Chuhui Xue, Shijian Lu |
ICCV | 2 |
| 2019 | MSR: Multi-Scale Shape Regression for Scene Text DetectionabstractState-of-the-art scene text detection techniques predict quadrilateral boxes that are prone to localization errors while dealing with straight or curved text lines of different orientations and lengths in scenes. This paper presents a novel multi-scale shape regression network (MSR) that is capable of locating text lines of different lengths, shapes and curvatures in scenes. The proposed MSR detects scene texts by predicting dense text boundary points that inherently capture the location and shape of text lines accurately and are also more tolerant to the variation of text line length as compared with the state of the arts using proposals or segmentation. Additionally, the multi-scale network extracts and fuses features at different scales which demonstrates superb tolerance to the text scale variation. Extensive experiments over several public datasets show that the proposed MSR obtains superior detection performance for both curved and straight text lines of different lengths and orientations. Chuhui Xue, Shijian Lu, Wei Zhang 0021 |
IJCAI | 1 |
| 2018 | Accurate Scene Text Detection Through Border Semantics Awareness and Bootstrapping
Chuhui Xue, Shijian Lu, Fangneng Zhan |
ECCV (16) | 1 |
| 2018 | Verisimilar Image Synthesis for Accurate Detection and Recognition of Texts in Scenes
Fangneng Zhan, Shijian Lu, Chuhui Xue |
ECCV (8) | 3 |