VLDB 2026 Research / reviewers in the wild / expert
Guoguang Hua
dblp:240/7280
· DBLP profile ↗
20ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-guided policy network for zero-shot object goal visual navigation
Guoguang Hua, Yaqiong Ding, Yuhuan Chen, Dan Xiang, Wenbin Zou |
Knowl. Based Syst. | 1 |
| 2025 | Class-discriminative domain generalization for semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Rong You, Wenbin Zou, Xia Li 0006 |
Image Vis. Comput. | 4 |
| 2025 | Domain-generalized token linking in vision foundation models for semantic segmentationabstract[S U M M A R Y] Vision Foundation Models (VFMs) achieve remarkable performance compared with traditional methods based on convolutional neural networks and vision transformer networks in Domain-Generalized Semantic Segmentation (DGSS). These VFM-based DGSS methods focus on adopting efficient parameter fine-tuning strategies that use a set of learnable tokens to fine-tune VFMs to the downstream DGSS task, yet struggle to mine domain-invariant information from VFMs since the backbone of VFMs is frozen during the fine-tuning stage. To address this issue, a Domain-Generalized Token Linking (DGTL) approach is proposed to mine domain-invariant information from VFMs for improving the performance in unseen target domains, which contains a Text-guided Dual Token Linking (TDTL) module and a Text-guided Distribution Normalization (TDN) strategy. For the TDTL module, first, a set of learnable tokens is linked to the text embeddings for building the relations between the learnable tokens and text embeddings, which is beneficial for learning domain-invariant tokens since the text embeddings generated from the CLIP model are domain-invariant. Second, the feature-level and mask-level linking strategies are proposed to link the learned domain-invariant tokens to the features and masks to guide the mining of domain-invariant information from the VFM. For the TDN strategy, the pairwise similarity between the predictive masks associated with the learnable tokens and the text embeddings is utilized to explicitly align the semantic distribution of visual features in the learnable tokens with the text embeddings. Extensive experiments demonstrate that the DGTL approach achieves superior performance to recent methods across multiple DGSS benchmarks. The code is released on GitHub: https://github.com/seabearlmx/DGTL . Muxin Liao, Jia-Yang Wang, Hong Deng, Yingqiong Peng, Hua Yin, Guoguang Hua |
Knowl. Based Syst. | 7 |
| 2025 | A global reweighting approach for cross-domain semantic segmentation
Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004 |
Signal Process. Image Commun. | 4 |
| 2025 | Contextual Guidance Network for Real-Time Semantic Segmentation of Autonomous DrivingabstractWith the rise of mobile computing and the increasing demand for real-time applications, the need for efficient and accurate semantic segmentation models has become paramount. However, existing state-of-the-art models are often hindered by heavy computational requirements, rendering them impractical for real-time applications. To tackle this challenge, we introduce the Contextual Guidance Network (CGNet), an efficient and lightweight network designed specifically for real-time semantic segmentation in autonomous driving. CGNet primarily consists of two key components: the Contextual Guidance Module (CGM) and the Triple-Branch Residual Fusion Module (TRFM). The CGM is comprised of the Downsampling Refine Unit (DRU) and the Contextual Guidance Bottleneck (CGB), which are utilized to gather dense contextual information. The DRU functions as a downsampling tool to generate low-resolution images, while the CGB extracts rich contextual information from both spatial and channel dimensions. Additionally, the TRFM utilizes the Residual Fusion Module (RFM) to achieve feature fusion and enhance pixel prediction accuracy. Without bells and whistles, CGNet achieves impressive mean intersection over union (mIoU) scores of 77.11% with 1.00 million parameters at 86.71 frames per second (fps) on the Cityscapes dataset, 72.26% mIoU at 88.62 fps on the CamVid dataset, and 63.32% mIoU on the BDD100K dataset. Extensive experiments demonstrate that CGNet achieves a favorable tradeoff between segmentation accuracy, inference speed and computational cost, making it suitable for autonomous driving systems with limited hardware resources. The source code will be available on GitHub: https://github.com/lv881314/CGNet Muxin Liao, Guoguang Hua, Yuhang Zhang 0011, Wenbin Zou |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Strip and asymmetric aggregation network for unstructured terrain segmentation in wild environments
Shishun Tian, Yuhang Zhang 0011, Muxin Liao, Guoguang Hua, Wenbin Zou |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | PDA: Progressive Domain Adaptation for Semantic Segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Knowl. Based Syst. | 4 |
| 2024 | Considering representation diversity and prediction consistency for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Knowl. Based Syst. | 4 |
| 2024 | Preserving Label-Related Domain-Specific Information for Cross-Domain Semantic SegmentationabstractUnsupervised domain adaptation semantic segmentation (UDASS) methods aim to learn domain-invariant information for alleviating the distribution shift problem between the source and target domains. However, ignoring the learning of domain-specific information that is label-related may limit the class discriminability on the target domain. We argue that a good representation for the UDASS task not only contains domain-invariant information but also preserves label-related domain-specific information. In this paper, a novel frequency spectrum domain adaptation approach via meta-learning (ML-FSDA) is proposed to achieve this goal for improving the class discriminability and generalization ability. ML-FSDA contains a frequency-spectrum meta-learning framework (FMF) and a class-aware domain-specific memory bank (CDMB). Specifically, first, inspired by the observation that the high-frequency component is consistent across different domains while the low-frequency component is much more domain-specific, the FMF aims to respectively learn label-related domain-specific and domain-invariant information from low-frequency and high-frequency images in a unified framework via the meta-learning strategy. Second, the CDMB is designed to preserve the label-related domain-specific information of each class in an external memory bank while the CDMB is updated in every iteration of the meta-training stage. Finally, the CDMB is utilized to embed the label-related domain-specific information into domain-invariant information at the class level during the meta-testing stage to enhance the class discriminability on the target domain. Extensive experiments demonstrate the effectiveness of ML-FSDA on two challenging cross-domain semantic segmentation benchmarks. Notably, for the GTA5 to Cityscapes task and the SYNTHIA to Cityscapes task, the proposed ML-FSDA achieves superior performance with 77.3% mIoU and 68.8% mIoU, respectively. The source code is released at https://github.com/seabearlmx/FSL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Calibration-Based Multi-Prototype Contrastive Learning for Domain Generalization Semantic Segmentation in Traffic ScenesabstractPrototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features for domain generalization semantic segmentation. These methods assume that the prototypes in different domains are invariant. However, the prototypes in different domains have discrepancies as well. First, the prototypes of the same class in different domains may be different. Second, the prototypes of different classes may be similar. To address these issues, a calibration-based multi-prototype contrastive learning (CMPCL) approach is proposed, which contains an uncertainty-guided multi-prototype contrastive learning (UMPCL) and a hard-weighted multi-prototype contrastive learning (HMPCL). Specifically, the UMPCL uses an uncertainty probability matrix, derived from element-wise discrepancies between the prototypes of the same class, to calibrate the weights of prototypes for alleviating the discrepancy between the prototypes of the same class in different domains. The HMPCL uses a hard-weighted matrix that is generated by the similarity between the prototypes of different classes, to calibrate the weights of the hard-aligned prototypes for alleviating the issue of similar prototypes between different classes, with hard-aligned prototypes referring to those exhibiting such similarity. Furthermore, since the learned class-wise domain-invariant features may overfit the prototype in the source domain, multi-prototype contrastive learning is used in the UMPCL and HMPCL to avoid this risk. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on multiple benchmarks of domain generalization semantic segmentation. The source code has been released onhttps://github.com/seabearlmx/CMPCL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | "Where Does the Devil Lie?": Multimodal Multitask Collaborative Revision Network for Trusted Road SegmentationabstractRoad segmentation is an essential component of navigation systems. Although recent advancements in road segmentation, the occurrence of failure segmentations remains inevitable. For safety-critical tasks, e.g., navigation, knowing when and where road segmentation fails is crucial. In this paper, we propose a novel trusted road segmentation architecture, namely Multimodal Multitask Collaborative Revision Network (M2CRN), to improve the trust of road segmentation. Our approach incorporates two strategies to predict and rectify segmentation errors. Firstly, a joint learning framework is devised to generate road segmentation results while estimating failure segmentation masks. Secondly, the road segmentation branch is equipped with an Uncertainty-Aware Revision Module (UARM), which eliminates the error in road segmentation. Additionally, we suppress the response of error regions in the road segmentation branch with an innovative design, called Adaptive Soft Error Suppression (ASES). To validate our methods, extensive experiments are conducted on three benchmark road segmentation datasets. The results demonstrate significant performance improvements with a real-time inference speed of 33.3 FPS, reaffirming the soundness of our revision model. Guoguang Hua, Dalian Zheng, Shishun Tian, Wenbin Zou, Shenglan Liu 0001, Xia Li 0006 |
IEEE Trans. Multim. | 1 |
| 2023 | Need a dog for seeing eye? A Walk Viewpoint Dataset for Freespace Detection in Unstructured EnvironmentsabstractFreespace Detection (FD) is crucial for robust and safe autonomous navigation. However, existing datasets usually concentrate on structure road environments. The FD in unstructured environments, e.g., walk assistance for the visually-impaired, has been rarely investigated. In this paper, We propose a novel dataset called the Walk Viewpoint Dataset (WVD). Different from the previous datasets, we focus on the walk viewpoint, where FD can provide the potential for improving the walking of visually impaired people. The target regions of WVD are annotated with 20 categories by fine-grained labels, which consist of 3,737 images and depth images. Moreover, we propose a new annotation hierarchy, which allows different degrees of complexity and creates opportunities for new training methods. Finally, our study provides the statistical analysis of label characteristics and baseline analysis, which demonstrates its distinction compared to previous datasets. The dataset can be accessed through the project pages: http://www.sensingAI.com.cn. Wenbin Zou, Guoguang Hua, Guangxu Chen, Zaiyue He, Guangli Liu, Huakun Li, Shishun Tian |
ICME | 2 |
| 2023 | Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic SegmentationabstractPrototypical contrastive learning (PCL) has been widely used to learn class-wise domain-invariant features recently. These methods are based on the assumption that the prototypes, which are represented as the central value of the same class in a certain domain, are domain-invariant. Since the prototypes of different domains have discrepancies as well, the class-wise domain-invariant features learned from the source domain by PCL need to be aligned with the prototypes of other domains simultaneously. However, the prototypes of the same class in different domains may be different while the prototypes of different classes may be similar, which may affect the learning of class-wise domain-invariant features. Based on these observations, a calibration-based dual prototypical contrastive learning (CDPCL) approach is proposed to reduce the domain discrepancy between the learned class-wise features and the prototypes of different domains for domain generalization semantic segmentation. It contains an uncertainty-guided PCL (UPCL) and a hard-weighted PCL (HPCL). Since the domain discrepancies of the prototypes of different classes may be different, we propose an uncertainty probability matrix to represent the domain discrepancies of the prototypes of all the classes. The UPCL estimates the uncertainty probability matrix to calibrate the weights of the prototypes during the PCL. Moreover, considering that the prototypes of different classes may be similar in some circumstances, which means these prototypes are hard-aligned, the HPCL is proposed to generate a hard-weighted matrix to calibrate the weights of the hard-aligned prototypes during the PCL. Extensive experiments demonstrate that our approach achieves superior performance over current approaches on domain generalization segmentation tasks. The source code will be released at https://github.com/seabearlmx/CDPCL. Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
ACM Multimedia | 4 |
| 2023 | SPNet: An RGB-D Sequence Progressive Network for Road Semantic SegmentationabstractRoad semantic segmentation is an essential component of autonomous driving and blind navigation. Although many excellent RGB-based road semantic segmentation algorithms have been proposed, these methods may not detect correctly due to the lack of geometric information. Recently, RGB-D road semantic segmentation methods attract more research attention. However, the existing RGB-D methods ignore the impact of unknown noise in sensors. To solve this problem, we propose an RGB-D Sequence Progressive Network (SPNet) for road semantic segmentation. Specifically, we first propose a sequence-based RGB-D feature extractor to alleviate the effect of noise. Then, We propose a multi-modal feature fusion (MMFF) module to enhance the feature representation of multi-modal data by further alleviating the effect of noise. Finally, we propose a semantic flow prediction (SFP) module that aims to align the multi-modal features in the decoder. Extensive experiments are conducted on several challenging datasets, including KITTI and GMRP. Our method achieves an F-score of 97.21% on the KITTI official leaderboard and ranked third in the official leaderboard. Yuhang Zhang 0011, Guoguang Hua, Ruijing Long, Shishun Tian, Wenbin Zou |
MMSP | 3 |
| 2023 | Domain-invariant information aggregation for domain generalization semantic segmentation
Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Guoguang Hua, Wenbin Zou, Xia Li 0006 |
Neurocomputing | 4 |
| 2023 | Learning Shape-Invariant Representation for Generalizable Semantic SegmentationabstractSemantic segmentation assigns a category for each pixel and has achieved great success in a supervised manner. However, it fails to generalize well in new domains due to the domain gap. Domain adaptation is a popular way to solve this issue, but it needs target data and cannot handle unavailable domains. In domain generalization (DG), the model is trained without the target data and DG aims to generalize well in new unavailable domains. Recent works reveal that shape recognition is beneficial for generalization but still lack exploration in semantic segmentation. Meanwhile, the object shapes also exist a discrepancy in different domains, which is often ignored by the existing works. Thus, we propose a Shape-Invariant Learning (SIL) framework to focus on learning shape-invariant representation for better generalization. Specifically, we first define the structural edge, which considers both the object boundary and the inner structure of the object to provide more discrimination cues. Then, a shape perception learning strategy including a texture feature discrepancy reduction loss and a structural feature discrepancy enlargement loss is proposed to enhance the shape perception ability of the model by embedding the structural edge as a shape prior. Finally, we use shape deformation augmentation to generate samples with the same content and different shapes. Essentially, our SIL framework performs implicit shape distribution alignment at the domain-level to learn shape-invariant representation. Extensive experiments show that our SIL framework achieves state-of-the-art performance. Yuhang Zhang 0011, Shishun Tian, Muxin Liao, Guoguang Hua, Wenbin Zou, Chen Xu 0004 |
IEEE Trans. Image Process. | 4 |
| 2023 | Multiple Relational Learning Network for Joint Referring Expression Comprehension and SegmentationabstractMulti-task learning is a successful learning framework which improves the performance of prediction models by leveraging knowledge among related tasks. Referring expression comprehension (REC) and segmentation (RES) are highly relevant tasks, which both are language-guided visual recognition tasks. However, their relations have not yet been fully exploited in previous works. In this paper, a Multiple Relational Learning Network (MRLN) is proposed for multi-task learning of REC and RES. First, a feature-feature interaction learning module is introduced to handle the complicated interactions among features. Moreover, we propose a feature-task dependence learning module, which associates the related features with target tasks. Furthermore, a task-task relationship learning module is designed, which captures the relationships among tasks automatically and guides the REC and RES fine-tuning adaptively. To verify our proposed approach, experiments are conducted on three benchmark datasets, i.e., RefCOCO, RefCOCO+, and RefCOCOg. Extensive experiments demonstrate that the multiple relationships are more appealing since it alleviates the prediction inconsistency issue in multi-task setup. In addition, the experimental results report the significant performance gains of MRLN over most existing methods, i.e., up to 83.46 % for REC and 63.62 % for RES over state-of-the-art methods, which demonstrate the validity and superiority of MRLN. Guoguang Hua, Muxin Liao, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou |
IEEE Trans. Multim. | 1 |
| 2022 | Exploring more concentrated and consistent activation regions for cross-domain semantic segmentation
Muxin Liao, Guoguang Hua, Shishun Tian, Yuhang Zhang 0011, Wenbin Zou, Xia Li 0006 |
Neurocomputing | 2 |
| 2020 | Multipath affinage stacked - hourglass networks for human pose estimation
Guoguang Hua, Shiguang Liu |
Frontiers Comput. Sci. | 1 |
| 2019 | Human Pose Estimation in Video via Structured Space Learning and Halfway Temporal EvaluationabstractHuman pose estimation from image or video is a basic issue in computer graphics and computer vision. The challenge of human pose estimation in video lies in the temporal coherency issue. The temporal consistency in video is the contents' similarity shown in the video frames. In video, temporal consistency maintenance of human pose estimation is to obtain better long-term consistency. Great major methods for the long-term consistency are using the whole video optimization method, which makes very large computation and the absence of consistency before and after the articulated limbs. In this paper, a novel method for the maintenance of temporal consistency is proposed. We maintain the temporal consistency of the video by the structured space learning and halfway temporal evaluation methods. We adopt a three-stage multi-feature deep convolution network framework to generate the initial posture joints position data, and a long-term temporal coherence is propagated to the overall video at each stage. The long-term consistency is more appealing since it produces stable results over larger periods of time. Our method can achieve good temporal consistency and get accurate and stable human pose estimation results. Various experimental results demonstrated the superiority of our method. Shiguang Liu, Yang Li 0067, Guoguang Hua |
IEEE Trans. Circuits Syst. Video Technol. | 3 |