EDBT 2026 Demo / reviewers in the wild / expert
Zhangxuan Gu
dblp:243/6953
· DBLP profile ↗
16ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-2102-2693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GUI-G²: Gaussian Reward Modeling for GUI GroundingabstractGraphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks. Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang |
AAAI | 2 |
| 2025 | Efficient Transfer Learning for Video-language Foundation ModelsabstractPre-trained vision-language models provide a robust foundation for efficient transfer learning across various downstream tasks. In the field of video action recognition, mainstream approaches often introduce additional modules to capture temporal information. Although the additional modules increase the capacity of model, enabling it to better capture video-specific inductive biases, existing methods typically introduce a substantial number of new parameters and are prone to catastrophic forgetting of previously acquired generalizable knowledge. In this paper, we propose a parameter-efficient Multi-modal Spatio-Temporal Adapter (MSTA) to enhance the alignment between textual and visual representations, achieving a balance between generalizable knowledge and task-specific adaptation. Furthermore, to mitigate over-fitting and enhance generalizability, we introduce a spatio-temporal description-guided consistency constraint. This constraint involves providing template inputs (e.g., "a video of {cls}") to the trainable language branch and LLM-generated spatio-temporal descriptions to the pre-trained language branch, enforcing output consistency between the branches. This approach reduces overfitting to downstream tasks and enhances the distinguishability of the trainable branch within the spatio-temporal semantic space. We evaluate the effectiveness of our approach across four tasks: zero-shot transfer, few-shot learning, base-to-novel generalization, and fully-supervised learning. Compared to many state-of-the-art methods, our MSTA achieves outstanding performance across all evaluations, while using only 2-7% of the trainable parameters in the original model. Haoxing Chen, Zizheng Huang, Yan Hong 0001, Yanshuo Wang, Zhongcai Lyu, Zhuoer Xu, Jun Lan 0001, Zhangxuan Gu |
CVPR | 8 |
| 2025 | Conditional Prototype Rectification Prompt LearningabstractPre-trained large-scale vision-language models (VLMs) have acquired profound understanding of general visual concepts. Recent advancements in efficient transfer learning (ETL) have shown remarkable success in fine-tuning VLMs within the scenario of limited data, introducing only a few parameters to harness task-specific insights from VLMs. Despite significant progress, current leading ETL methods tend to overfit the narrow distributions of base classes seen during training and encounter two primary challenges: (i) only utilizing uni-modal information to modeling task-specific knowledge; and (ii) using costly and time-consuming methods to supplement knowledge. To address these issues, we propose a Conditional Prototype Rectification Prompt Learning (CPR) method to correct the bias of the base examples and augment limited data in an effective way. Specifically, we alleviate over-fitting on base classes from two aspects. First, each input image acquires knowledge from both textual and visual prototypes and then generates sample-conditional text tokens. Second, we extract utilizable knowledge from unlabeled data to further refine the prototypes. These two strategies mitigate biases that stem from base classes, yielding a more effective classifier. Extensive experiments on 11 benchmark datasets show that our CPR achieves state-of-the-art performance on few-shot classification, base-to-new generalization, and cross-dataset generalization tasks. Our code is available at https://github.com/chenhaoxing/CPR. Haoxing Chen, Zizheng Huang, Yan Hong 0001, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Segment Anything Model Meets Image HarmonizationabstractImage harmonization is a crucial technique in image composition that aims to seamlessly match the background by adjusting the foreground of composite images. Current methods adopt either global-level or pixel-level feature matching. Global-level feature matching ignores the proximity prior, treating foreground and background as separate entities. On the other hand, pixel-level feature matching loses contextual information. Therefore, it is necessary to use the information from semantic maps that describe different objects to guide harmonization. In this paper, we propose Semantic-guided Region-aware Instance Normalization (SRIN) that can utilize the semantic segmentation maps output by a pre-trained Segment Anything Model (SAM) to guide the visual consistency learning of foreground and background features. Abundant experiments demonstrate the superiority of our method for image harmonization over state-of-the-art methods. Haoxing Chen, Zhangxuan Gu, Zhuoer Xu, Jun Lan 0001, Huaxiong Li |
ICASSP | 3 |
| 2024 | Diffusioninst: Diffusion Model for Instance SegmentationabstractDiffusion frameworks have achieved comparable performance with previous state-of-the-art image generation models. This paper proposes DiffusionInst, a novel framework representing instances as vectors and formulates instance segmentation as a noise-to-vector denoising process. The model is trained to reverse the noisy groundtruth mask without any inductive bias from RPN. It takes a randomly generated vector as input and outputs mask with multi-step denoising during inference. Extensive experimental results on COCO and LVIS show that DiffusionInst achieves competitive performance. Our code is available at https://github.com/chenhaoxing/DiffusionInst. Zhangxuan Gu, Haoxing Chen, Zhuoer Xu |
ICASSP | 1 |
| 2024 | PC2: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal RetrievalabstractIn the realm of cross-modal retrieval, seamlessly integrating diverse modalities within multimedia remains a formidable challenge, especially given the complexities introduced by noisy correspondence learning (NCL). Such noise often stems from mismatched data pairs, which is a significant obstacle distinct from traditional noisy labels. This paper introduces Pseudo-Classification based Pseudo-Captioning (PC$^2$) framework to address this challenge. PC$^2$ offers a threefold strategy: firstly, it establishes an auxiliary "pseudo-classification" task that interprets captions as categorical labels, steering the model to learn image-text semantic similarity through a non-contrastive mechanism. Secondly, unlike prevailing margin-based techniques, capitalizing on PC$^2$'s pseudo-classification capability, we generate pseudo-captions to provide more informative and tangible supervision for each mismatched pair. Thirdly, the oscillation of pseudo-classification is borrowed to assistant the correction of correspondence. In addition to technical contributions, we develop a realistic NCL dataset called Noise of Web (NoW), which could be a new powerful NCL benchmark where noise exists naturally. Empirical evaluations of PC$^2$ showcase marked improvements over existing state-of-the-art robust cross-modal retrieval techniques on both simulated and realistic datasets with various NCL settings. The contributed dataset and source code are released at https://github.com/alipay/PC2-NoiseofWeb. Yue Duan, Zhangxuan Gu, Zhenzhe Ying, Lei Qi 0001, Changhua Meng, Yinghuan Shi |
ACM Multimedia | 2 |
| 2023 | Mobile User Interface Element Detection Via Adaptively Prompt TuningabstractRecent object detection approaches rely on pretrained vision-language models for image-text alignment. However, they fail to detect the Mobile User Interface (MUI) element since it contains additional OCR information, which describes its content and function but is often ignored. In this paper, we develop a new MUI element detection dataset named MUI-zh and propose an Adaptively Prompt Tuning (APT) module to take advantage of discriminating OCR information. APT is a lightweight and effective module to jointly optimize category prompts across different modalities. For every element, APT uniformly encodes its visual features and OCR descriptions to dynamically adjust the representation of frozen category prompts. We evaluate the effectiveness of our plug-and-play APT upon several existing CLIP-based detectors for both standard and open-vocabulary MUI element detection. Extensive experiments show that our method achieves considerable improvements on two datasets. The datasets is available at github.com/antmachineintelligence/MUI-zh. Zhangxuan Gu, Zhuoer Xu, Haoxing Chen, Jun Lan 0001, Changhua Meng, Weiqiang Wang 0002 |
CVPR | 1 |
| 2023 | Backpropagation Path Search On Adversarial TransferabilityabstractDeep neural networks are vulnerable to adversarial examples, dictating the imperativeness to test the model’s robustness before deployment. Transfer-based attackers craft adversarial examples against surrogate models and transfer them to victim models deployed in the black-box situation. To enhance the adversarial transferability, structure-based attackers adjust the backpropagation path to avoid the attack from overfitting the surrogate model. However, existing structure-based attackers fail to explore the convolution module in CNNs and modify the backpropagation graph heuristically, leading to limited effectiveness. In this paper, we propose backPropagation pAth Search (PAS), solving the aforementioned two problems. We first propose SkipConv to adjust the backpropagation path of convolution by structural reparameterization. To overcome the drawback of heuristically designed backpropagation paths, we further construct a Directed Acyclic Graph (DAG) search space, utilize one-step approximation for path evaluation and employ Bayesian Optimization to search for the optimal path. We conduct comprehensive experiments in a wide range of transfer settings, showing that PAS improves the attack success rate by a huge margin for both normally trained and defense models. Zhuoer Xu, Zhangxuan Gu, Jianping Zhang 0002, Shiwen Cui, Changhua Meng, Weiqiang Wang 0002 |
ICCV | 2 |
| 2023 | Hierarchical Dynamic Image HarmonizationabstractImage harmonization is a critical task in computer vision, which aims to adjust the foreground to make it compatible with the background. Recent works mainly focus on using global transformations (i.e., normalization and color curve rendering) to achieve visual consistency. However, these models ignore local visual consistency and their huge model sizes limit their harmonization ability on edge devices. In this paper, we propose a hierarchical dynamic network (HDNet) to adapt features from local to global view for better feature transformation in efficient image harmonization. Inspired by the success of various dynamic models, local dynamic (LD) module and mask-aware global dynamic (MGD) module are proposed in this paper. Specifically, LD matches local representations between the foreground and background regions based on semantic similarities, then adaptively adjust every foreground local representation according to the appearance of its K-nearest neighbor background regions. In this way, LD can produce more realistic images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. The MGD effectively applies distinct convolution to the foreground and background, learning the representations of foreground and background regions as well as their correlations to the global harmonization, facilitating local visual consistency for the images much more efficiently. Experimental results demonstrate that the proposed HDNet significantly reduces the total model parameters by more than 80% compared to previous methods, while still attaining state-of-the-art performance on the popular iHarmony4 dataset. Additionally, we introduced a lightweight version of HDNet, i.e., HDNet-lite, which has only 0.65MB parameters, yet it still achieve competitive performance. Our code is avaliable at https://github.com/chenhaoxing/HDNet. Haoxing Chen, Zhangxuan Gu, Jun Lan 0001, Changhua Meng, Weiqiang Wang 0002, Huaxiong Li |
ACM Multimedia | 2 |
| 2023 | DiffUTE: Universal Text Editing Diffusion ModelabstractDiffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we propose a universal self-supervised text editing diffusion model (DiffUTE), which aims to replace or modify words in the source image with another one while maintaining its realistic appearance. Specifically, we build our model on a diffusion model and carefully modify the network structure to enable the model for drawing multilingual characters with the help of glyph and position information. Moreover, we design a self-supervised learning framework to leverage large amounts of web data to improve the representation ability of the model. Experimental results show that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. Our code will be avaliable in \url{https://github.com/chenhaoxing/DiffUTE}. Haoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 0001, Xing Zheng, Changhua Meng, Huijia Zhu, Weiqiang Wang 0002 |
NeurIPS | 3 |
| 2023 | From Pixel to Patch: Synthesize Context-Aware Features for Zero-Shot Semantic SegmentationabstractZero-shot learning (ZSL) has been actively studied for image classification tasks to relieve the burden of annotating image labels. Interestingly, the semantic segmentation task requires more labor-intensive pixel-wise annotation, but zero-shot semantic segmentation has not attracted extensive research interest. Thus, we focus on zero-shot semantic segmentation that aims to segment unseen objects with only category-level semantic representations provided for unseen categories. In this article, we propose a novel context-aware feature generation network (CaGNet) that can synthesize context-aware pixel-wise visual features for unseen categories based on category-level semantic representations and pixel-wise contextual information. The synthesized features are used to fine-tune the classifier to enable segmenting of unseen objects. Furthermore, we extend pixel-wise feature generation and fine-tuning to patch-wise feature generation and fine-tuning, which additionally considers the interpixel relationship. Experimental results on Pascal-VOC, Pascal-context, and COCO-stuff show that our method significantly outperforms the existing zero-shot semantic segmentation methods. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingabstractRecently, various multimodal networks for Visually-Rich Document Understanding(VRDU) have been proposed, showing the promotion of transformers by integrating visual and layout information with the text embeddings. However, most existing approaches utilize the position embeddings to incorporate the sequence information, neglecting the noisy improper reading order obtained by OCR tools. In this paper, we propose a robust layout-aware multimodal network named XYLayoutLM to capture and leverage rich layout information from proper reading orders produced by our Augmented XY Cut. Moreover, a Dilated Conditional Position Encoding module is proposed to deal with the input sequence of variable lengths, and it additionally extracts local layout information from both textual and vi-sual modalities while generating position embeddings. Experiment results show that our XYLayoutLM achieves competitive results on document understanding tasks. Zhangxuan Gu, Changhua Meng, Ke Wang 0042, Jun Lan 0001, Weiqiang Wang 0002, Ming Gu 0011, Liqing Zhang 0001 |
CVPR | 1 |
| 2021 | Hard Pixel Mining for Depth Privileged Semantic SegmentationabstractSemantic segmentation has achieved remarkable progress but remains challenging due to the complex scene, object occlusion, and so on. Some research works have attempted to use extra information such as a depth map to help RGB based semantic segmentation because the depth map could provide complementary geometric cues. However, due to the inaccessibility of depth sensors, depth information is usually unavailable for the test images. In this paper, we leverage only the depth of training images as the privileged information to mine the hard pixels in semantic segmentation, in which depth information is only available for training images but not available for test images. Specifically, we propose a novel Loss Weight Module, which outputs a loss weight map by employing two depth-related measurements of hard pixels: Depth Prediction Error and Depth-aware Segmentation Error. The loss weight map is then applied to segmentation loss, with the goal of learning a more robust model by paying more attention to the hard pixels. Besides, we also explore a curriculum learning strategy based on the loss weight map. Meanwhile, to fully mine the hard pixels on different scales, we apply our loss weight module to multi-scale side outputs. Our hard pixels mining method achieves the state-of-the-art results on three benchmark datasets, and even outperforms the methods which need depth input during testing. Zhangxuan Gu, Li Niu 0002, Haohua Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Context-aware Feature Generation For Zero-shot Semantic SegmentationabstractExisting semantic segmentation models heavily rely on dense pixel-wise annotations. To reduce the annotation pressure, we focus on a challenging task named zero-shot semantic segmentation, which aims to segment unseen objects with zero annotations. This task can be accomplished by transferring knowledge across categories via semantic word embeddings. In this paper, we propose a novel context-aware feature generation method for zero-shot segmentation named CaGNet. In particular, with the observation that a pixel-wise feature highly depends on its contextual information, we insert a contextual module in a segmentation network to capture the pixel-wise contextual information, which guides the process of generating more diverse and context-aware features from semantic word embeddings. Our method achieves state-of-the-art results on three benchmark datasets for zero-shot segmentation. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
ACM Multimedia | 1 |
| 2020 | Multi-mode neural network for human action recognitionabstractVideo data are of two different intrinsic modes, in‐frame and temporal. It is beneficial to incorporate static in‐frame features to acquire dynamic features for video applications. However, some existing methods such as recurrent neural networks do not have a good performance, and some other such as 3D convolutional neural networks (CNNs) are both memory consuming and time consuming. This study proposes an effective framework that takes the advantage of deep learning on the static image feature extraction to tackle the video data. After extracting in‐frame feature vectors using a pretrained deep network, the authors integrate them and form a multi‐mode feature matrix, which preserves the multi‐mode structure and high‐level representation. They propose two models for follow‐up classification. The authors first introduce a temporal CNN, which directly feeds the multi‐mode feature matrix into a CNN. However, they show that characteristics of the multi‐mode features differ significantly in distinct modes. The authors therefore further propose the multi‐mode neural network (MMNN), in which different modes deploy different types of layers. They evaluate their algorithm with the task of human action recognition. The experimental results show that the MMNN achieves a much better performance than the existing long short‐term memory‐based methods and consumes far fewer resources than the existing 3D end‐to‐end models. Haohua Zhao 0001, Weichen Xue, Zhangxuan Gu, Li Niu 0002, Liqing Zhang 0001 |
IET Comput. Vis. | 4 |
| 2019 | Clothes Keypoints Localization and Attribute Recognition via Prior KnowledgeabstractRich clothes datasets and high-quality annotations have driven recent advances in fashion clothes recognition. However, the existing approaches treat clothes as common images, ignoring the prior clothing knowledge such as spatial relations, symmetry, proportions, and key characteristics of clothes. In order to combine the semantic information with the advantages of deep learning, we propose Detection+, a model using the prior symmetric constraint to refine the keypoints located by any backbone detection networks. To deal with uncertainty in labelling clothing, we introduce a new loss to utilize all available data which contain "maybe" labels. Detection+ has reduced about 2.54% Normalized Error in FashionAI dataset and improved 3.2% AP in human keypoints dataset coco2017 compared to the Mask R-CNN baseline. A large number of experimental results show the proposed approach achieves better results in different recognition datasets (resp., FashionAI, and Deepfashion) with about (resp., 2.57% mAP, and 10% recall) improvements. Zhangxuan Gu, Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
ICME | 1 |