Guangzhi Wang

dblp:67/1183 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-4677-1041ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BlobCtrl: Taming Controllable Blob for Element-level Image Editing
abstract
As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-based methods. In this work, we present BlobCtrl, a framework for element-level image editing based on a probabilistic blob-based representation. Treating blobs as visual primitives, BlobCtrl disentangles layout from appearance, affording fine-grained, controllable object-level elements manipulation. Our key contributions are twofold: 1) an in-context dual-branch diffusion model that separates foreground and background processing, incorporating blob representations to explicitly decouple layout and appearance; and 2) a self-supervised disentangle-then-reconstruct training paradigm with an identity-preserving loss function, along with tailored strategies to efficiently leverage blob-image pairs. To foster further research, we introduce BlobData for large-scale training, and BlobBench, a benchmark for systematic evaluation. Experimental results demonstrate that BlobCtrl achieves state-of-the-art performance in a variety of element-level editing tasks—such as object addition, removal, scaling, and replacement—while maintaining computational efficiency.
Yaowei Li 0001, Lingen Li, Zhaoyang Zhang 0004, Xiaoyu Li 0002, Guangzhi Wang, Hongxiang Li 0004, Xiaodong Cun, Ying Shan, Yuexian Zou
SIGGRAPH Asia5
2025 EVD Surgical Guidance With Retro-Reflective Tool Tracking and Spatial Reconstruction Using Head-Mounted Augmented Reality Device
abstract
Augmented Reality (AR) has been proven beneficial to External Ventricular Drain (EVD) surgery by providing in-situ visual guidance during operations. During this procedure, the key challenge is estimating the spatial relationship between pre-operative images and actual patient anatomy accurately and efficiently. Previous works have revealed conflicts between tracking accuracy, workflow efficiency, and non-invasiveness in tracking pipelines. This research fully utilizes the capabilities of Time of Flight (ToF) depth sensors, including retro-reflective tool tracking and dense surface information, to construct a convenient and accurate EVD guiding pipeline. As previous studies have proven significant depth errors in ToF depth sensors, we first evaluated the feasibility of using ToF sensors in surgical guidance by estimating its accuracy under different conditions and corrected this error in our pipeline. Our results show $ \text{7.580}\pm \text{1.488}\,\text{mm}$7.580±1.488mm depth value errors on human skin under HoloLens 2 depth camera, indicating the significance of depth correction. This error was reduced by over 85% using proposed depth correction method on head phantoms in different materials. The corrected depth information can then be utilized to reconstruct the head surface with sub-millimeter accuracy, validated on a series of 3D-printed models and a sheep head. To demonstrate the effectiveness of the proposed framework, we conducted a case study simulating EVD surgery. Five surgeons were involved in this study, each performing nine k-wire insertions on a head phantom under virtual guidance without tracking for surgical tools. The results revealed $ \text{2.09} \pm \text{1.00}\,\text{mm}$2.09±1.00mm translational and $\text{2.97}\pm \text{1.95}^\circ$2.97±1.95∘ orientational guidance accuracy, demonstrating competitive performance with previous research.
Wenqing Yan, Du Liu, Yuxing Yang, Yihao Liu 0004, Zhe Zhao 0005, Hui Ding 0003, Guangzhi Wang
IEEE Trans. Vis. Comput. Graph.9
2024 PELA: Learning Parameter-Efficient Models with Low-Rank Approximation
abstract
Applying a pre-trained large model to downstream tasks is prohibitive under resource-constrained conditions. Recent dominant approaches for addressing efficiency issues involve adding a few learnable parameters to the fIxed backbone model. This strategy, however, leads to more challenges in loading large models for downstream finetuning with limited resources. In this paper, we propose a novel method for increasing the parameter efficiency of pretrained models by introducing an intermediate pre-training stage. To this end, we first employ low-rank approximation to compress the original large model and then devise a feature distillation module and a weight perturbation regularization module. These modules are specifically designed to enhance the low-rank model. In particular, we update only the low-rank model while freezing the backbone parameters during pre-training. This allows for direct and efficient utilization of the low-rank model for downstream finetuning tasks. The proposed method achieves both efficiencies in terms of required parameters and computation time while maintaining comparable results with minimal modifications to the backbone architecture. Specifically, when applied to three vision-only and one vision-language Transformer models, our approach often demonstrates a merely rvO.6 point decrease in performance while reducing the original parameter size by 1/3 to 2/3. We release our code at link.
Guangzhi Wang, Mohan Kankanhalli
CVPR2
2024 SEED-Bench: Benchmarking Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given in-terleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of single image-text inputs, failing to keep up with the strides made in MLLMs. A comprehensive benchmark is imperative for investigating the progress and uncovering the limitations of current MLLMs. In this work, we categorize the capabilities of MLLMs into hierarchical levels from L0to L4based on the modalities they can ac-cept and generate, and propose SEED-Bench, a comprehensive benchmark that evaluates the hierarchical capa-bilities of MLLMs. Specifically, SEED-Bench comprises 24K multiple-choice questions with accurate human annotations, which span 27 dimensions, including the evaluation of both text and image generation. Multiple-choice questions with ground truth options derived from human annotation enable an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 22 prominent open-source MLLMs and summarize valuable observations. By revealing the limitations of existing MLLMs through extensive evaluations, we aim for SEED-Bench to provide insights that will mo-tivate future research toward the goal of General Artificial Intelligence. Dataset and evaluation code are available at https://github.com/AILab-CVC/SEED-Bench.
Bohao Li 0002, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang 0092, Ruimao Zhang, Ying Shan
CVPR4
2024 Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-Robustness
abstract
Human-Object Interaction (HOI) Detection constitutes an important aspect of human-centric scene understanding, which requires precise object detection and interaction recognition. Despite increasing advancement in detection, recognizing subtle and intricate interactions remains challenging. Recent methods have endeavored to leverage the rich semantic representation from pretrained CLIP, yet fail to efficiently capture finer-grained spatial features that are highly informative for interaction discrimination. In this work, instead of solely using representations from CLIP, we fill the gap by proposing a spatial adapter that efficiently utilizes the multi-scale spatial information in the pretrained detector. This leads to a bilateral adaptation that mutually produces complementary features. To further improve interaction recognition under occlusion, which is common in crowded scenarios, we propose an Occluded Part Extrapolation module that guides the model to recover the spatial details from manually occluded feature maps. Moreover, we design a Conditional Contextual Mining module that further mines informative contextual clues from the spatial features via a tailored cross-attention mechanism. Extensive experiments on V-COCO and HICO-DET benchmarks demonstrate that our method significantly outperforms prior art on both standard and zero-shot settings, resulting in new state-of-the-art performance. Additional ablation studies further validate the effectiveness of each component in our method.
Guangzhi Wang, Ziwei Xu 0001, Mohan Kankanhalli
CVPR1
2024 STTAR: Surgical Tool Tracking Using Off-the-Shelf Augmented Reality Head-Mounted Displays
abstract
The use of Augmented Reality (AR) for navigation purposes has shown beneficial in assisting physicians during the performance of surgical procedures. These applications commonly require knowing the pose of surgical tools and patients to provide visual information that surgeons can use during the performance of the task. Existing medical-grade tracking systems use infrared cameras placed inside the Operating Room (OR) to identify retro-reflective markers attached to objects of interest and compute their pose. Some commercially available AR Head-Mounted Displays (HMDs) use similar cameras for self-localization, hand tracking, and estimating the objects' depth. This work presents a framework that uses the built-in cameras of AR HMDs to enable accurate tracking of retro-reflective markers without the need to integrate any additional electronics into the HMD. The proposed framework can simultaneously track multiple tools without having previous knowledge of their geometry and only requires establishing a local network between the headset and a workstation. Our results show that the tracking and detection of the markers can be achieved with an accuracy of$0.09\pm 0.06\ mm$on lateral translation,$0.42 \pm 0.32\ mm$on longitudinal translation and$0.80 \pm 0.39^\circ$for rotations around the vertical axis. Furthermore, to showcase the relevance of the proposed framework, we evaluate the system's performance in the context of surgical procedures. This use case was designed to replicate the scenarios of k-wire insertions in orthopedic procedures. For evaluation, seven surgeons were provided with visual navigation and asked to perform 24 injections using the proposed framework. A second study with ten participants served to investigate the capabilities of the framework in the context of more general scenarios. Results from these studies provided comparable accuracy to those reported in the literature for AR-based navigation procedures.
Alejandro Martin-Gomez, Tianyu Song 0002, Guangzhi Wang, Hui Ding 0003, Nassir Navab, Zhe Zhao 0005, Mehran Armand
IEEE Trans. Vis. Comput. Graph.5
2023 Text to Point Cloud Localization with Relation-Enhanced Transformer
abstract
Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on a text-to-point-cloud cross-modal localization problem. Given a textual query, it aims to identify the described location from city-scale point clouds. The task involves two challenges. 1) In city-scale point clouds, similar ambient instances may exist in several locations. Searching each location in a huge point cloud with only instances as guidance may lead to less discriminative signals and incorrect results. 2) In textual descriptions, the hints are provided separately. In this case, the relations among those hints are not explicitly described, leaving the difficulties of learning relations to the agent itself. To alleviate the two challenges, we propose a unified Relation-Enhanced Transformer (RET) to improve representation discriminability for both point cloud and nature language queries. The core of the proposed RET is a novel Relation-enhanced Self-Attention (RSA) mechanism, which explicitly encodes instance (hint)-wise relations for the two modalities. Moreover, we propose a fine-grained cross-modal matching method to further refine the location predictions in a subsequent instance-hint matching stage. Experimental results on the KITTI360Pose dataset demonstrate that our approach surpasses the previous state-of-the-art method by large margins.
Guangzhi Wang, Hehe Fan, Mohan Kankanhalli
AAAI1
2023 Diffusion MRI data analysis assisted by deep learning synthesized anatomical images (DeepAnat)
abstract
-weighted (T1w) anatomical MRI data, which may be unacquired, corrupted by subject motion or hardware failure, or cannot be accurately co-registered to the diffusion data that are not corrected for susceptibility-induced geometric distortion. To address these challenges, this study proposes to synthesize high-quality T1w anatomical images directly from diffusion data using convolutional neural networks (CNNs) (entitled "DeepAnat"), including a U-Net and a hybrid generative adversarial network (GAN), and perform brain segmentation on synthesized T1w images or assist the co-registration using synthesized T1w images. The quantitative and systematic evaluations using data of 60 young subjects provided by the Human Connectome Project (HCP) show that the synthesized T1w images and results for brain segmentation and comprehensive diffusion analysis tasks are highly similar to those from native T1w data. The brain segmentation accuracy is slightly higher for the U-Net than the GAN. The efficacy of DeepAnat is further validated on a larger dataset of 300 more elderly subjects provided by the UK Biobank. Moreover, the U-Nets trained and validated on the HCP and UK Biobank data are shown to be highly generalizable to the diffusion data from Massachusetts General Hospital Connectome Diffusion Microstructure Dataset (MGH CDMD) acquired with different hardware systems and imaging protocols and therefore can be used directly without retraining or with fine-tuning for further improved performance. Finally, it is quantitatively demonstrated that the alignment between native T1w images and diffusion images uncorrected for geometric distortion assisted by synthesized T1w images substantially improves upon that by directly co-registering the diffusion and T1w images using the data of 20 subjects from MGH CDMD. In summary, our study demonstrates the benefits and practical feasibility of DeepAnat for assisting various diffusion MRI data analyses and supports its use in neuroscientific applications.
Qiuyun Fan, Berkin Bilgic, Guangzhi Wang, Wenchuan Wu 0003, Jonathan R. Polimeni, Karla L. Miller, Susie Yi Huang, Qiyuan Tian
Medical Image Anal.4
2023 Semantic-Aware Triplet Loss for Image Classification
abstract
Successful image classification requires a discriminative representation learning model for images. To approach this idea, deep metric learning (DML), serving as building a basic feature space with a pre-defined metric, has demonstrated compelling performance over the years. DML is often implemented with a carefully crafted loss function, such as the representative triplet loss, which encourages a positive sample to be by a fixed margin closer to the anchor than the negative. Despite its efficacy, the negative samples are treated uniformly, rendering the feature space less informative since different negative samples can be largely different from the anchor. In this work, we, for the first time, propose to exploit the semantic information inherent in discrete class labels as an aid for the triplet loss. Specifically, we build a bi-level negative sampling strategy,i.e., strong negative and weak negative sampling, with the guidance of an external knowledge source, from which rich class semantics can be extracted. With several fine-grained and complementary triplet losses based on this strategy, our method is enhanced with semantic awareness for image classification. In addition, to coordinate with the complicated training dynamics, we devise an ad-hoc Semantic Relation Weighting module, which consistently inspects model states and dynamically adjusts the importance of each triplet loss. It is worth noting that our method is plug-and-play, and we thus test its validity over various backbones and knowledge sources. Both qualitative and quantitative experimental results on benchmark datasets demonstrate the effectiveness of employing semantics for image classification.
Guangzhi Wang, Ziwei Xu 0001, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Multim.1
2022 Chairs Can Be Stood On: Overcoming Object Bias in Human-Object Interaction Detection
Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
ECCV (24)1
2022 Distance Matters in Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection has received considerable attention in the context of scene understanding. Despite the growing progress, we realize existing methods often perform unsatisfactorily on distant interactions, where the leading causes are two-fold: 1) Distant interactions are by nature more difficult to recognize than close ones. A natural scene often involves multiple humans and objects with intricate spatial relations, making the interaction recognition for distant human-object largely affected by complex visual context. 2) Insufficient number of distant interactions in datasets results in under-fitting on these instances. To address these problems, we propose a novel two-stage method for better handling distant interactions in HOI detection. One essential component in our method is a novel Far Near Distance Attention module. It enables information propagation between humans and objects, whereby the spatial distance is skillfully taken into consideration. Besides, we devise a novel Distance-Aware loss function which leads the model to focus more on distant yet rare interactions. We conduct extensive experiments on HICO-DET and V-COCO datasets. The results show that the proposed method surpass existing methods significantly, leading to new state-of-the-art results.
Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
ACM Multimedia1
2022 Relation-Aware Compositional Zero-Shot Learning for Attribute-Object Pair Recognition
abstract
This paper proposes a novel model for recognizing images with composite attribute-object concepts, notably for composite concepts that are unseen during model training. We aim to explore the three key properties required by the task — relation-aware, consistent, and decoupled—to learn rich and robust features for primitive concepts that compose attribute-object pairs. To this end, we propose the Blocked Message Passing Network (BMP-Net). The model consists of two modules. The concept module generates semantically meaningful features for primitive concepts, whereas the visual module extracts visual features for attributes and objects from input images. A message passing mechanism is used in the concept module to capture the relations between primitive concepts. Furthermore, to prevent the model from being biased towards seen composite concepts and reduce the entanglement between attributes and objects, we propose a blocking mechanism that equalizes the information available to the model for both seen and unseen concepts. Extensive experiments and ablation studies on two benchmarks show the efficacy of the proposed model.
Ziwei Xu 0001, Guangzhi Wang, Yongkang Wong, Mohan Kankanhalli
IEEE Trans. Multim.2
2021 Dynamic Knowledge Distillation with Cross-Modality Knowledge Transfer
abstract
Supervised learning for vision tasks has achieved great success be-cause of the advances of deep learning research in many areas, such as high quality datasets, network architectures and regularization methods. In the vanilla deep learning paradigm, training a model for visual tasks is mainly based on the provided training images and annotations. Inspired by human learning with knowledge transfer where information from multiples modalities are considered, we pro-pose to improve visual tasks' performance by introducing explicit knowledge extracted from other modalities. As the first step, we propose to improve image classification performance by introducing linguistic knowledge as additional constraints in model learning. This knowledge is represented as a set of constraints to be jointly utilized with visual knowledge. To coordinate the training dynamic, we propose to imbue our model the ability of dynamic distilling from multiple knowledge sources. This is done via a model agnostic knowledge weighting module which guides the learning process and updates via meta-steps during training. Preliminary experiments on various benchmark datasets validate the efficacy of our method. Our code will be made publicly available to ensure reproducibility.
Guangzhi Wang
ACM Multimedia1
2020 Multi-Source Distilling Domain Adaptation
abstract
Deep neural networks suffer from performance decay when there is domain shift between the labeled source domain and unlabeled target domain, which motivates the research on domain adaptation (DA). Conventional DA methods usually assume that the labeled data is sampled from a single source distribution. However, in practice, labeled data may be collected from multiple sources, while naive application of the single-source DA algorithms may lead to suboptimal solutions. In this paper, we propose a novel multi-source distilling domain adaptation (MDDA) network, which not only considers the different distances among multiple sources and the target, but also investigates the different similarities of the source samples to the target ones. Specifically, the proposed MDDA includes four stages: (1) pre-train the source classifiers separately using the training data from each source; (2) adversarially map the target into the feature space of each source respectively by minimizing the empirical Wasserstein distance between source and target; (3) select the source training samples that are closer to the target to fine-tune the source classifiers; and (4) classify each encoded target feature by corresponding source classifier, and aggregate different predictions using respective domain weight, which corresponds to the discrepancy between each source and target. Extensive experiments are conducted on public DA benchmarks, and the results demonstrate that the proposed MDDA significantly outperforms the state-of-the-art approaches. Our source code is released at: https://github.com/daoyuan98/MDDA.
Sicheng Zhao, Guangzhi Wang, Shanghang Zhang, Yaxian Li, Zhichao Song, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer
AAAI2
2020 Automatic Radiofrequency Ablation Planning for Liver Tumors With Multiple Constraints Based on Set Covering
abstract
Radiofrequency ablation (RFA) is now a widely used minimally invasive treatment method for hepatic tumors. Preoperative planning plays a vital role in RFA therapy. With increasing tumor size, multiple overlapping ablations are needed, which are challenging to optimize while considering clinical constraints. In this paper, we present a new automatic RFA planning method. First, a 2-steps set cover-based model is formulated, which can integrate multiple clinical constraints for optimization of overlapping ablations. To ensure that the planning model can be solved in a reasonable time, a search space reducing strategy is then proposed. We also developed an algorithm for automatic RFA electrode selection, which provides a proper electrode ablation zone for the planning model. The proposed method was evaluated with 20 tumors of varying sizes (0.92 cm3to 28.4 cm3). Results showed that the proposed method can generate clinical feasible RFA plans with a minimum number of RFA electrodes and ablations, complete tumor coverage and minimized ablation of normal tissue.
Libin Liang, Derek W. Cool, Nirmal Kakani, Guangzhi Wang, Hui Ding 0003, Aaron Fenster
IEEE Trans. Medical Imaging4
2019 Development of a Multi-objective Optimized Planning Method for Microwave Liver Tumor Ablation
Libin Liang, Derek W. Cool, Nirmal Kakani, Guangzhi Wang, Hui Ding 0003, Aaron Fenster
MICCAI (5)4
2016 Shape context and projection geometry constrained vasculature matching for 3D reconstruction of coronary artery
Ruoxiu Xiao, Jian Yang 0009, Jingfan Fan, Danni Ai, Guangzhi Wang, Yongtian Wang
Neurocomputing5