EDBT 2026 Demo / reviewers in the wild / expert
Zonghao Guo
dblp:193/8032
· DBLP profile ↗
19ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-8492-2130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic PyramidabstractVision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic levels. To address this, we present Hierarchical window (Hiwin) transformer as a plug-and-play solution for MLLMs, centered around our inverse semantic pyramid (ISP). Hiwin transformer comprises two key modules: (i) a visual detail injection module, which progressively injects low-level visual details into high-level language-aligned semantics features, thereby constructing an ISP, and (ii) a hierarchical window attention module, which leverages cross-scale windows to condense multi-level semantics from the ISP. Notably, our design achieves an average boost of 3.7% across 14 benchmarks compared with the baseline method, 9.3% on DocVQA for instance. Zonghao Guo, Xuesong Yang, Chi Chen 0005, Yuan Yao 0013, Tat-Seng Chua, Maosong Sun 0001 |
AAAI | 3 |
| 2026 | RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic EvaluationabstractBingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang, Chi Chen, Ling Yao, Chenghu Zhou, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingxian Wu, Yu Zhang 0186, Zonghao Guo, Xingbo Du, Yanghao Li, Chi Chen 0005, Ling Yao, Chenghu Zhou, Maosong Sun 0001 |
ACL (1) | 3 |
| 2025 | XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?abstractThe astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the imagery features ultra-high resolution that incorporates extremely complex semantic relationships. Existing benchmarks usually adopt notably smaller image sizes than real-world RS scenarios, suffer from limited annotation quality, and consider insufficient dimensions of evaluation. To address these issues, we present XLRS-Bench: a comprehensive benchmark for evaluating the perception and reasoning capabilities of MLLMs in ultra-high-resolution RS scenarios. XLRS-Bench boasts the largest average image size (8500×8500) observed thus far, with all evaluation samples meticulously annotated manually, assisted by a novel semi-automatic captioner on ultra-high-resolution RS images. On top of the XLRS-Bench, 16 sub-tasks are defined to evaluate MLLMs’ 10 kinds of perceptual capabilities and 6 kinds of reasoning capabilities, with a primary emphasis on advanced cognitive processes that facilitate real-world decision-making and the capture of spatiotemporal changes. The results of both general and RS-focused MLLMs on XLRS-Bench indicate that further efforts are needed for real-world RS applications. We have open-sourced XLRS-Bench to support further research in developing more powerful MLLMs for remote sensing. Fengxiang Wang 0004, Hongzhen Wang, Zonghao Guo, Di Wang 0023, Yulin Wang 0002, Mingshuo Chen, Long Lan, Wenjing Yang 0002, Jing Zhang 0037, Zhiyuan Liu 0001, Maosong Sun 0001 |
CVPR | 3 |
| 2025 | Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang 0004, Hongzhen Wang, Di Wang 0023, Zonghao Guo, Zhenyu Zhong, Long Lan, Wenjing Yang 0002, Jing Zhang 0037 |
ICCV | 4 |
| 2025 | Video-R1: Reinforcing Video Reasoning in MLLMsabstractInspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly applying RL training with the GRPO algorithm to video reasoning presents two primary challenges: (i) a lack of temporal modeling for video reasoning, and (ii) the scarcity of high-quality video-reasoning data. To address these issues, we first propose the T-GRPO algorithm, which encourages models to utilize temporal information in videos for reasoning. Additionally, instead of relying solely on video data, we incorporate high-quality image-reasoning data into the training process. We have constructed two datasets: Video-R1-CoT-165k for SFT cold start and Video-R1-260k for RL training, both comprising image and video data. Experimental results demonstrate that Video-R1 achieves significant improvements on video reasoning benchmarks such as VideoMMMU and VSI-Bench, as well as on general video benchmarks including MVBench and TempCompass, etc. Notably, Video-R1-7B attains a 37.1\% accuracy on video spatial reasoning benchmark VSI-bench, surpassing the commercial proprietary model GPT-4o. All code, models, and data will be released. Kaituo Feng, Kaixiong Gong, Zonghao Guo, Tianshuo Peng, Junfei Wu, Benyou Wang, Xiangyu Yue 0001 |
NeurIPS | 4 |
| 2025 | GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K ResolutionabstractUltra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce **SuperRS-VQA** (avg. 8,376$\times$8,376) and **HighRS-VQA** (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: *Background Token Pruning* and *Anchored Token Selection*, to reduce the memory footprint while preserving key semantics. Integrating these techniques, we introduce **GeoLLaVA-8K**, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench. Datasets and code were released at https://github.com/MiliLab/GeoLLaVA-8K. Fengxiang Wang 0004, Mingshuo Chen, Di Wang 0023, Haotian Wang 0001, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang 0002, Hongzhen Wang, Wenjing Yang 0002, Bo Du 0001, Jing Zhang 0037 |
NeurIPS | 6 |
| 2025 | Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation
Zonghao Guo, Fang Wan 0001, Mingxiang Liao, Qixiang Ye |
Int. J. Comput. Vis. | 1 |
| 2025 | Adaptive Deformation-Learning and Multiscale-Integrated Network for Remote Sensing Object DetectionabstractModern human productivity and daily life rely on identifying ground objects using remote sensing images (RSIs). Traditional remote sensing object detection (RSOD) techniques lack timeliness and accuracy and fail to meet practical demands. Existing deep learning algorithms face continued challenges when processing RSIs because of the diverse shapes and extensive scale variations of objects, of which a significant proportion are small scale. To address these challenges, we propose the PSWP-DETR, a transformer-based network that leverages adaptive deformation learning and multiscale integration for enhanced object detection in remote sensing. First, we propose PradatorConv (PdConv) to address the significant shape changes of objects because it adaptively learns the horizontal and vertical deformations to perceive the complex geometric features of RSIs. Second, we propose scale-wise differential modules (SDMs), which comprise multiscale convolution (MSC) and edge captor convolution (ECC). SDM integrates features across various scales and captures edge characteristics and local textures. This is advantageous for detecting multiscale objects, tiny objects with limited feature information. Finally, we propose the whale particle optimization (WPO) algorithm for learning rate optimization, which improves convergence speed and accuracy. Experiments using the VisDrone2019-DET, DIOR, and AI-TOD datasets demonstrated that PSWP-DETR achieves the best accuracy benefits, offering significant insights for future RSOD efforts. The source code will be available athttps://github.com/Get1star/PSWP-DETR.git. Xiyu Zhong, Jialei Zhan, Lingtao Zhang, Guoxiong Zhou, Mingyue Liang, Kaitai Yang, Zonghao Guo, Liujun Li |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Hierarchical AttentionShift for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) remains a challenging task when appearance variances across object parts cause semantic inconsistency. In this article, we propose a hierarchical AttentionShift approach, to solve the semantic inconsistency issue through exploiting the hierarchical nature of semantics and the flexibility of key-point representation. The estimation of hierarchical attention is defined upon key-point sets. The representative key points are iteratively estimated spatially and in the feature space to capture the fine-grained semantics and cover the full object extent. Hierarchical AttentionShift is performed at instance, part, and fine-grained levels, optimizing object semantics while promoting the conventional self-attention activation to hierarchical activation with local refinement. Experiments on PASCAL VOC 2012 Aug and MS-COCO 2017 benchmarks show that hierarchical AttentionShift improves the state-of-the-art (SOTA) method by 10.4% and 7.0% upon mean average precision (mAP)50, respectively. When applying hierarchical AttentionShift to the segment anything model (SAM), 9.4% AP improvement on the COCO test-dev is achieved. Hierarchical AttentionShift provides a fresh insight to regularize the self-attention mechanism for fine-grained vision tasks. The code is available at github.com/MingXiangL/AttentionShift. Mingxiang Liao, Fang Wan 0001, Zonghao Guo, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | LLaVA-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images
Zonghao Guo, Ruyi Xu, Yuan Yao 0013, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu 0001, Gao Huang 0001 |
ECCV (83) | 1 |
| 2024 | ControlCap: Controllable Region-Level Captioning
Yuzhong Zhao, Zonghao Guo, Weijia Wu 0001, Chen Gong 0005, Qixiang Ye, Fang Wan 0001 |
ECCV (38) | 3 |
| 2023 | AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance SegmentationabstractPointly supervised instance segmentation (PSIS) learns to segment objects using a single point within the object extent as supervision. Challenged by the non-negligible semantic variance between object parts, however, the single supervision point causes semantic bias and false segmentation. In this study, we propose an AttentionShift method, to solve the semantic bias issue by iteratively decomposing the instance attention map to parts and estimating fine-grained semantics of each part. AttentionShift consists of two modules plugged on the vision transformer backbone: (i) token querying for pointly supervised attention map generation, and (ii) key-point shift, which re-estimates part-based attention maps by key-point filtering in the feature space. These two steps are iteratively performed so that the part-based attention maps are optimized spatially as well as in the feature space to cover full object extent. Experiments on PASCAL VOC and MS COCO 2017 datasets show that AttentionShift respectively improves the state-of-the-art of by 7.7% and 4.8% under [email protected], setting a solid PSIS baseline using vision transformer. Mingxiang Liao, Zonghao Guo, Yuze Wang 0004, Bailan Feng, Fang Wan 0001 |
CVPR | 2 |
| 2023 | Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionabstractModern object detectors have taken the advantages of backbone networks pre-trained on large scale datasets. Except for the backbone networks, however, other components such as the detector head and the feature pyramid network (FPN) remain trained from scratch, which hinders the generalization capacity of detectors. In this study, we propose to integrally migrate pre-trained transformer encoder-decoders (imTED) to a detector, constructing a feature extraction path which is "fully pre-trained" so that detectors’ generalization capacity is maximized. The essential differences between imTED with the baseline detector are twofold: (1) migrating the pre-trained transformer decoder to the detector head while removing the randomly initialized FPN from the feature extraction path; and (2) defining a multi-scale feature modulator (MFM) to enhance scale adaptability. Such designs not only reduce randomly initialized parameters significantly but also unify detector training with representation learning intendedly. Experiments on the MS COCO object detection dataset show that imTED consistently outperforms its counterparts by ~2.4 AP. Without bells and whistles, imTED improves the state-of-the-art of few-shot object detection by up to 7.6 AP. Code is released at https://github.com/LiewFeng/imTED. Feng Liu 0050, Xiaosong Zhang 0004, Zhiliang Peng, Zonghao Guo, Fang Wan 0001, Xiangyang Ji, Qixiang Ye |
ICCV | 4 |
| 2023 | Conformer: Local Features Coupling Global Representations for Recognition and DetectionabstractWith convolution operations, Convolutional Neural Networks (CNNs) are good at extracting local features but experience difficulty to capture global representations. With cascaded self-attention modules, vision transformers can capture long-distance feature dependencies but unfortunately deteriorate local feature details. In this paper, we propose a hybrid network structure, termed Conformer, to take both advantages of convolution operations and self-attention mechanisms for enhanced representation learning. Conformer roots in feature coupling of CNN local features and transformer global representations under different resolutions in an interactive fashion. Conformer adopts a dual structure so that local details and global dependencies are retained to the maximum extent. We also propose a Conformer-based detector (ConformerDet), which learns to predict and refine object proposals, by performing region-level feature coupling in an augmented cross-attention fashion. Experiments on ImageNet and MS COCO datasets validate Conformer's superiority for visual recognition and object detection, demonstrating its potential to be a general backbone network. Zhiliang Peng, Zonghao Guo, Yaowei Wang 0001, Lingxi Xie, Jianbin Jiao, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Bidirectional Feature Globalization for Few-shot Semantic Segmentation of 3D Point Cloud ScenesabstractFew-shot segmentation of point cloud remains a challenging task, as there is no effective way to convert local point cloud information to global representation, which hinders the generalization ability of point features. In this study, we propose a bidirectional feature globalization (BFG) approach, which leverages the similarity measurement between point features and prototype vectors to embed global perception to local point features in a bidirectional fashion. With point-to-prototype globalization (P02PrG), BFG aggregates local point features to prototypes according to similarity weights from dense point features to sparse prototypes. With prototype-to-point globalization (Pr2PoG), the global perception is embedded to local point features based on similarity weights from sparse prototypes to dense point features. The sparse prototypes of each class embedded with global perception are summarized to a single prototype for few-shot 3D segmentation based on the metric learning framework. Extensive experiments on S3DIS and ScanNet demonstrate that BFG significantly outperforms the state-of-the-art methods. Yongqiang Mao, Zonghao Guo, Haowen Guo |
3DV | 2 |
| 2022 | Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects is a challenging problem considering that the receptive field intersection between objects causes spatial feature aliasing. In this paper, we propose a convex-hull feature adaptation (CFA) approach, with the aim to configure convolutional features in accordance with irregular object layouts. CFA roots in the convex-hull feature representation, which defines a set of dynamically sampled feature points guided by the convex intersection over union (CIoU) to bound object extent. CFA pursues optimal feature assignment by constructing convex-hull sets and iteratively splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA defines a systematic way to adapt convolutional features on regular grids to objects of irregular shapes. Experiments on DOTA and SKU110K-R datasets show that CFA achieved new state-of-the-art performance for detecting oriented and densely packed objects. CFA also sets a solid baseline for convex polygon prediction on the MS COCO dataset defined for general object detection. Code is available athttps://github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Xiaosong Zhang 0004, Chang Liu 0047, Xiangyang Ji, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Beyond Bounding-Box: Convex-Hull Feature Adaptation for Oriented and Densely Packed Object DetectionabstractDetecting oriented and densely packed objects remains challenging for spatial feature aliasing caused by the intersection of reception fields between objects. In this paper, we propose a convex-hull feature adaptation (CFA) approach for configuring convolutional features in accordance with oriented and densely packed object layouts. CFA is rooted in convex-hull feature representation, which defines a set of dynamically predicted feature points guided by the convex intersection over union (CIoU) to bound the extent of objects. CFA pursues optimal feature assignment by constructing convex-hull sets and dynamically splitting positive or negative convex-hulls. By simultaneously considering overlapping convex-hulls and objects and penalizing convex-hulls shared by multiple objects, CFA alleviates spatial feature aliasing towards optimal feature adaptation. Experiments on DOTA and SKU110K-R datasets show that CFA significantly outperforms the baseline approach, achieving new state-of-the-art detection performance. Code is available at github.com/SDL-GuoZonghao/BeyondBoundingBox. Zonghao Guo, Chang Liu 0042, Xiaosong Zhang 0004, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 1 |
| 2021 | Long-tailed Distribution AdaptationabstractRecognizing images with long-tailed distributions remains a challenging problem while there lacks an interpretable mechanism to solve this problem. In this study, we formulate Long-tailed recognition as Domain Adaption (LDA), by modeling the long-tailed distribution as an unbalanced domain and the general distribution as a balanced domain. Within the balanced domain, we propose to slack the generalization error bound, which is defined upon the empirical risks of unbalanced and balanced domains and the divergence between them. We propose to jointly optimize empirical risks of the unbalanced and balanced domains and approximate their domain divergence by intra-class and inter-class distances, with the aim to adapt models trained on the long-tailed distribution to general distributions in an interpretable way. Experiments on benchmark datasets for image recognition, object detection, and instance segmentation validate that our LDA approach, beyond its interpretability, achieves state-of-the-art performance. Zhiliang Peng, Zonghao Guo, Xiaosong Zhang 0004, Jianbin Jiao, Qixiang Ye |
ACM Multimedia | 3 |
| 2016 | Optimal control for context-sensitive probabilistic Boolean networks with perturbation using probabilisitic model checkingabstractA context-sensitive probabilistic Boolean network with perturbation (CS-PBNp) closely models gene regulatory networks under external controls that alter the evolution of the networks in a desirable way over a finite time horizon. In this paper, we consider optimal control for a CS-PBNp, proposing an approach, based on a formal verification technique - probabilistic model checking, for finding optimal control policy that minimizes the expected cost over the entire control horizon. To this end, we first present a detailed procedure of modeling a CS-PBNp using the modeling language of a widely used probabilistic model checker PRISM. Furthermore, by analyzing computation of reward-based temporal properties, we provide a reduction approach allowing us to formulate the optimal control problem as minimum reachability reward properties. Based on this result, we incorporate control and state cost information into the PRISM code of a CS-PBNp such that automated model checking a minimum reachability reward property on the code gives the solution to the optimal control problem. Experiment results on an apoptosis network demonstrate the feasibility and effectiveness of our approach. Ou Wei, Zonghao Guo, Yun Niu, Wenyuan Liao |
BIBM | 2 |