VLDB 2026 Research / reviewers in the wild / expert
Xu Zheng 0002
dblp:147/1591-2
· DBLP profile ↗
30ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0003-4008-8951ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 10 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 16 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image ComprehensionabstractSatire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks. Haiwei Xue, Minghao Han, Mingcheng Li, Xiaolu Hou, Dingkang Yang, Lihua Zhang 0002, Xu Zheng 0002 |
AAAI | 8 |
| 2026 | Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsabstractChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng, Yiyu Wang, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou, Yuqian Fu, Bin Ren, Linfeng Zhang, Xuming Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng 0002, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou 0001, Yuqian Fu, Bin Ren 0005, Linfeng Zhang 0001, Xuming Hu |
ACL (1) | 4 |
| 2026 | Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Seigo Ito, Hiroshi Murase |
Int. J. Comput. Vis. | 4 |
| 2026 | Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi |
Int. J. Comput. Vis. | 5 |
| 2026 | Correction: Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi |
Int. J. Comput. Vis. | 5 |
| 2026 | CLIP-to-Seg Distillation for Zero-Shot Semantic Segmentation
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Daisuke Deguchi, Hiroshi Murase |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation
Yihong Cao, Jiaming Zhang 0001, Xu Zheng 0002, Hao Shi 0004, Kunyu Peng, Kailun Yang 0001, Hui Zhang 0023 |
ICCV | 3 |
| 2025 | Reducing Unimodal Bias in Multi-Modal Semantic Segmentation With Multi-Scale Functional Entropy RegularizationabstractFusing and balancing multi-modal inputs from novel sensors for dense prediction tasks, particularly semantic segmentation, is critically important yet remains a significant challenge. One major limitation is the tendency of multi-modal frameworks to over-rely on easily learnable modalities, a phenomenon referred to as unimodal dominance or bias. This issue becomes especially problematic in real-world scenarios where the dominant modality may be unavailable, resulting in severe performance degradation. To this end, we apply a simple but effective plug-and-play regularization term based on functional entropy, which introduces no additional parameters or modules. This term is designed to intuitively balance the contribution of each visual modality to the segmentation results. Specifically, we leverage the log-Sobolev inequality to bound functional entropy using functional-Fisher-information. By maximizing the information contributed by each visual modality, our approach mitigates unimodal dominance and establishes a more balanced and robust segmentation framework. A multi-scale regularization module is proposed to apply our proposed plug-and-play term on high-level features and also segmentation predictions for more balanced multi-modal learning. Extensive experiments on three datasets demonstrate that our proposed method achieves superior performance, i.e., +13.94%, +3.25%, and +3.64%, without introducing any additional parameters. Xu Zheng 0002, Yuanhuiyi Lyu, Lutao Jiang, Danda Pani Paudel, Luc Van Gool, Xuming Hu |
ICCV | 1 |
| 2025 | OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
Ding Zhong, Xu Zheng 0002, Chenfei Liao, Yuanhuiyi Lyu, Jialei Chen 0001, Shengyang Wu, Linfeng Zhang 0001, Xuming Hu |
ICCV | 2 |
| 2025 | RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive LearningabstractRecent text-to-image generative models, e.g., Stable Diffusion V3 and Flux, have achieved notable progress. However, these models are strongly restricted to their limited knowledge, a.k.a., their own fixed parameters, that are trained with closed datasets. This leads to significant hallucinations or distortions when facing fine-grained and unseen novel real-world objects, e.g., the appearance of the Tesla Cybertruck. To this end, we present the first real-object-based retrieval-augmented generation framework (RealRAG), which augments fine-grained and unseen novel object generation by learning and retrieving real-world images to overcome the knowledge gaps of generative models. Specifically, to integrate missing memory for unseen novel object generation, we train a reflective retriever by self-reflective contrastive learning, which injects the generator’s knowledge into the sef-reflective negatives, ensuring that the retrieved augmented images compensate for the model’s missing knowledge. Furthermore, the real-object-based framework integrates fine-grained visual knowledge for the generative models, tackling the distortion problem and improving the realism for fine-grained object generation. Our Real-RAG is superior in its modular application to all types of state-of-the-art text-to-image generative models and also delivers remarkable performance boosts with all of them, such as a gain of 16.18% FID score with the auto-regressive model on the Stanford Car benchmark. Yuanhuiyi Lyu, Xu Zheng 0002, Lutao Jiang, Xin Zou 0001, Huiyu Zhou 0005, Linfeng Zhang 0001, Xuming Hu |
ICML | 2 |
| 2025 | Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language GuidanceabstractThe perception capability of robotic systems relies on the richness of the dataset. Although Segment Anything Model 2 (SAM2), trained on large datasets, demonstrates strong perception potential in perception tasks, its inherent training paradigm prevents it from being suitable for RGB-T tasks. To address these challenges, we propose SHIFNet, a novel SAM2-driven Hybrid Interaction Paradigm that unlocks the potential of SAM2 with linguistic guidance for efficient RGB-Thermal perception. Our framework consists of two key components: (1) Semantic-Aware Cross-modal Fusion (SACF) module that dynamically balances modality contributions through text-guided affinity learning, overcoming SAM2’s inherent RGB bias; (2) Heterogeneous Prompting Decoder (HPD) that enhances global semantic information through a semantic enhancement module and then combined with category embeddings to amplify cross-modal semantic consistency. With 32.27M trainable parameters, SHIFNet achieves state-of-the-art segmentation performance on public benchmarks, reaching 89.8% on PST900 and 67.8% on FMB, respectively. The framework facilitates the adaptation of pre-trained large models to RGB-T segmentation tasks, effectively mitigating the high costs associated with data collection while endowing robotic systems with comprehensive perception capabilities. The source code will be made publicly available at https://github.com/iAsakiT3T/SHIFNet. Zhiyong Li 0001, Xu Zheng 0002, Kailun Yang 0001 |
IROS | 6 |
| 2025 | Omnidirectional Spatial Modeling from Correlated PanoramasabstractOmnidirectional scene understanding is vital for various downstream applications, such as embodied AI, autonomous driving, and immersive environments, yet remains challenging due to geometric distortion and complex spatial relations in 360° imagery. Existing omnidirectional methods achieve scene understanding within a single frame while neglecting cross-frame correlated panoramas. To bridge this gap, we introduce CFpano, the first benchmark dataset dedicated to cross-frame correlated panoramas visual question answering in the holistic 360° scenes. CFpano consists of over 2700 images together with over 8000 question-answer pairs, and the question types include both multiple choice and open-ended VQA. Building upon our CFpano, we further present Pano-R1, a multi-modal large language model (MLLM) fine-tuned with Group Relative Policy Optimization (GRPO) and a set of tailored reward functions for robust and consistent reasoning with cross-frame correlated panoramas. Benchmark experiments with existing MLLMs are conducted with our CFpano. The experimental results demonstrate that Pano-R1achieves state-of-the-art performance across both multiple-choice and open-ended VQA tasks, outperforming strong baselines on all major reasoning categories (+5.37% in overall performance). Our analyses validate the effectiveness of GRPO and establish a new benchmark for panoramic scene understanding. Xinshen Zhang, Tongxi Fu, Xu Zheng 0002 |
MMAsia | 3 |
| 2025 | Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object DetectionabstractCross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need for both visual realism and domain alignment. Existing strategies, such as copy-paste augmentation and text-to-image generation, often fail to preserve the correct object category or produce backgrounds coherent with the target domain, making them non-trivial to apply directly to CD-FSOD. To address these challenges, we propose Domain-RAG, a training-free, retrieval-guided compositional image generation framework tailored for CD-FSOD. Domain-RAG consists of three stages: domain-aware background retrieval, domain-guided background generation, and foreground-background composition. Specifically, the input image is first decomposed into foreground and background regions. We then retrieve semantically and stylistically similar images to guide a generative model in synthesizing a new background, conditioned on both the original and retrieved contexts. Finally, the preserved foreground is composed with the newly generated domain-aligned background to form the generated image. Without requiring any additional supervision or training, Domain-RAG produces high-quality, domain-consistent samples across diverse tasks, including CD-FSOD, remote sensing FSOD, and camouflaged FSOD. Extensive experiments show consistent improvements over strong baselines and establish new state-of-the-art results. Codes will be released upon acceptance.The source code and instructions are available at https://github.com/LiYu0524/Domain-RAG. Yu Li 0007, Xingyu Qiu, Yuqian Fu, Tianwen Qian, Xu Zheng 0002, Danda Pani Paudel, Yanwei Fu 0001, Xuanjing Huang 0001, Luc Van Gool, Yu-Gang Jiang 0001 |
NeurIPS | 6 |
| 2025 | Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionabstractDespite their powerful capabilities, multimodal large language models (MLLMs) suffer from considerable computational overhead due to their reliance on massive visual tokens. Recent studies have explored token pruning to alleviate this problem, which typically uses text-vision cross-attention or [CLS] attention to assess and discard redundant visual tokens. In this work, we identify a critical limitation of such attention-first pruning approaches, i.e., they tend to preserve semantically similar tokens, resulting in pronounced performance drops under high pruning rates. To this end, we propose HoloV, a simple yet effective, plug-and-play visual token pruning framework for efficient inference. Distinct from previous attention-first schemes, HoloV rethinks token retention from a holistic perspective. By adaptively distributing the pruning budget across different spatial crops, HoloV ensures that the retained tokens capture the global visual context rather than isolated salient features. This strategy minimizes representational collapse and maintains task-relevant information even under aggressive pruning. Experimental results demonstrate that our HoloV achieves superior performance across various tasks, MLLM architectures, and pruning ratios compared to SOTA methods. For instance, LLaVA1.5 equipped with HoloV preserves 95.8% of the original performance after pruning 88.9% of visual tokens, achieving superior efficiency-accuracy trade-offs. Xin Zou 0001, Yuanhuiyi Lyu, Xu Zheng 0002, Linfeng Zhang 0001, Xuming Hu |
NeurIPS | 6 |
| 2025 | Edge Priors Image Inpaintig With StyleGAN2abstractABSTRACT Image inpainting represents a fundamental task in computer vision, focusing primarily on the generation of missing content within an image to restore its integrity and aesthetics. Existing GAN‐based approaches often produce content with ambiguity and require a high training difficulties. Moreover, they tend to focus narrowly on damaged regions, leading to edge distortions that hinder generalisation. To address these challenges, we propose an algorithm that consist of two distinct networks. The first network, called Edge‐e4e, is designed for initial image restoration and integrates a pre‐trained StyleGAN2 as the generator to mitigate edge distortions. This network employs an encoder‐StyleGAN2 architecture, where only the encoder part is trained, thereby reducing training costs compared to traditional GAN methods. To resolve ambiguities in the restored content, we incorporate edge information into the damaged regions, guiding the network to generate content that is consistent with the original image. The second network, called Appending network, includes two style‐based encoders and a generator to improve the similarity between the images restored by Edge‐e4e and the original images. Specifically, we subtract the restored images from the input images in the channel dimension to obtain distortion maps, which serve as a prior to refine the restored images from Edge‐e4e. To further enhance the quality of refined images, we propose incorporating plugin and modulate plugin modules for style extraction and fusion. These modules utilise information from the input images and seamlessly integrate it into the style‐based generator. Experimental results demonstrate that our algorithm achieves high‐fidelity restoration and excellent generalisation, with optimal FID and Lpips metrics of 0.0631 and 0.875, respectively. The code is publicly available at: https://github.com/MengZhen‐Chi/Edge‐Pries‐Image‐Inpainting‐with‐StyleGAN2 . Mengzhen Chi, Chong Fu 0001, Xu Zheng 0002, Jialei Chen 0001, Chiu-Wing Sham |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | 360SFUDA++: Towards Source-Free UDA for Panoramic Segmentation by Learning Reliable Category PrototypesabstractIn this paper, we address the challenging source-free unsupervised domain adaptation (SFUDA) for pinhole-to-panoramic semantic segmentation, given only a pinhole image pre-trained model (i.e., source) and unlabeled panoramic images (i.e., target). Tackling this problem is non-trivial due to three critical challenges: 1) semantic mismatches from the distinct Field-of-View (FoV) between domains, 2) style discrepancies inherent in the UDA problem, and 3) inevitable distortion of the panoramic images. To tackle these problems, we propose 360SFUDA++ that effectively extracts knowledge from the source pinhole model with only unlabeled panoramic images and transfers the reliable knowledge to the target panoramic domain. Specifically, we first utilize Tangent Projection (TP) as it has less distortion and meanwhile slits the equirectangular projection (ERP) to patches with fixed FoV projection (FFP) to mimic the pinhole images. Both projections are shown effective in extracting knowledge from the source model. However, as the distinct projections make it less possible to directly transfer knowledge between domains, we then propose Reliable Panoramic Prototype Adaptation Module (RPAM) to transfer knowledge at both prediction and prototype levels. RPAM selects the confident knowledge and integrates panoramic prototypes for reliable knowledge adaptation. Moreover, we introduce Cross-projection Dual Attention Module (CDAM), which better aligns the spatial and channel characteristics across projections at the feature level between domains. Both knowledge extraction and transfer processes are synchronously updated to reach the best performance. Extensive experiments on the synthetic and real-world benchmarks, including outdoor and indoor scenarios, demonstrate that our 360SFUDA++ achieves significantly better performance than prior SFUDA methods. Xu Zheng 0002, Peng Yuan Zhou, Athanasios V. Vasilakos, Lin Wang 0025 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Distilling efficient Vision Transformers from CNNs for semantic segmentation
Xu Zheng 0002, Yunhao Luo 0001, Peng Yuan Zhou, Lin Wang 0025 |
Pattern Recognit. | 1 |
| 2024 | UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them AllabstractWe present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities - image, text, audio, point cloud, thermal, video, and event data. Existing works, e.g., ImageBind [13], treat the image as the central modality and build an image-centered representation space; however, the space may be sub-optimal as it leads to an unbalanced representation space among all modalities. Moreover, the category names are directly used to extract text embeddings for the down-stream tasks, making it hardly possible to represent the se-mantics of multi-modal data. The ‘out-of-the-box’ insight of our UniBind is to make the alignment centers modality-agnostic and further learn a unified and balanced repre-sentation space, empowered by the large language mod-els (LLMs). UniBind is superior in its flexible application to all CLIP-style models and delivers remarkable performance boosts. To make this possible, we 1) construct a knowledge base of text with the help of LLMs and multi-modal LLMs; 2) adaptively build LLM-augmented class-wise embedding centers on top of the knowledge base and encoded visual embeddings; 3) align all the embeddings to the LLM-augmented embedding centers via contrastive learning to achieve a unified and balanced representation space. UniBind shows strong zero-shot recognition performance gains over prior arts by an average of 6.36%. Fi-nally, we achieve new state-of-the-art performance, e.g., a 6.75% gain on ImageNet, on the multi-modal fine-tuning setting while reducing 90% of the learnable parameters. Yuanhuiyi Lyu, Xu Zheng 0002, Jiazhou Zhou, Lin Wang 0025 |
CVPR | 2 |
| 2024 | EventDance: Unsupervised Source-Free Cross-Modal Adaptation for Event-Based Object RecognitionabstractIn this paper, we make the first attempt at achieving the cross-modal (i.e., image-to-events) adaptation for event-based object recognition without accessing any labeled source image data owning to privacy and commercial issues. Tackling this novel problem is non-trivial due to the novelty of event cameras and the distinct modality gap between images and events. In particular, as only the source model is available, a hurdle is how to extract the knowledge from the source model by only using the unlabeled target event data while achieving knowledge transfer. To this end, we propose a novel framework, dubbed Event-Dance for this unsupervised source-free cross-modal adaptation problem. Importantly, inspired by event-to-video reconstruction methods, we propose a reconstruction-based modality bridging (RMB) module, which reconstructs intensity frames from events in a self-supervised manner. This makes it possible to build up the surrogate images to extract the knowledge (i.e., labels) from the source model. We then propose a multi-representation knowledge adaptation (MKA) module that transfers the knowledge to target models learning events with multiple representation types for fully exploring the spatiotemporal information of events. The two modules connecting the source and target models are mutually updated so as to achieve the best performance. Experiments on three benchmark datasets with two adaption settings show that EventDance is on par with prior methods utilizing the source data. Xu Zheng 0002, Lin Wang 0025 |
CVPR | 1 |
| 2024 | Semantics, Distortion, and Style Matter: Towards Source-Free UDA for Panoramic SegmentationabstractThis paper addresses an interesting yet challenging problem-source-free unsupervised domain adaptation (SFUDA) for pinhole-to-panoramic semantic segmentation-given only a pinhole image-trained model (i.e., source) and unlabeled panoramic images (i.e., target). Tackling this problem is nontrivial due to the semantic mismatches, style discrepancies, and inevitable distortion of panoramic images. To this end, we propose a novel method that utilizes Tangent Projection (TP) as it has less distortion and meanwhile slits the equirectangular projection (ERP) with a fixed FoV to mimic the pinhole images. Both projections are shown effective in extracting knowledge from the source model. However, the distinct projection discrepancies between source and target domains impede the direct knowledge transfer; thus, we propose a panoramic prototype adaptation module (PPAM) to integrate panoramic prototypes from the extracted knowledge for adaptation. We then impose the loss constraints on both predictions and prototypes and propose a cross-dual attention module (CDAM) at the feature level to better align the spatial and channel characteristics across the domains and projections. Both knowledge extraction and transfer processes are synchronously updated to reach the best performance. Extensive experiments on the synthetic and real-world benchmarks, including outdoor and indoor scenarios, demonstrate that our method achieves significantly better performance than prior SFUDA methods for pinhole-to-panoramic adaptation. Xu Zheng 0002, Peng Yuan Zhou, Athanasios V. Vasilakos, Lin Wang 0025 |
CVPR | 1 |
| 2024 | ExACT: Language-Guided Conceptual Reasoning and Uncertainty Estimation for Event-Based Action Recognition and MoreabstractEvent cameras have recently been shown beneficial for practical vision tasks, such as action recognition, thanks to their high temporal resolution, power efficiency, and reduced privacy concerns. However, current research is hindered by 1) the difficulty in processing events because of their prolonged duration and dynamic actions with complex and ambiguous semantics and 2) the redundant action depiction of the event frame representation with fixed stacks. We find language naturally conveys abundant semantic information, rendering it stunningly superior in reducing semantic uncertainty. In light of this, we propose ExACT, a novel approach that, for the first time, tackles event-based action recognition from a cross-modal conceptualizing perspective. Our ExACT brings two technical contributions. Firstly, we propose an adaptive fine-grained event (AFE) representation to adaptively filter out the repeated events for the stationary objects while preserving dynamic ones. This subtly enhances the performance of ExACT without extra computational cost. Then, we propose a conceptual reasoning-based uncertainty estimation module, which simulates the recognition process to enrich the semantic representation. In particular, conceptual reasoning builds the temporal relation based on the action semantics, and uncertainty estimation tackles the semantic uncertainty of actions based on the distributional representation. Experiments show that our ExACT achieves superior recognition accuracy of 94.83%(+2.23%), 90.10%(+37.47%) and 67.24% on PAF, HARDVS and our SeAct datasets respectively. Jiazhou Zhou, Xu Zheng 0002, Yuanhuiyi Lyu, Lin Wang 0025 |
CVPR | 2 |
| 2024 | Learning Modality-Agnostic Representation for Semantic Segmentation from Any Modalities
Xu Zheng 0002, Yuanhuiyi Lyu, Lin Wang 0025 |
ECCV (34) | 1 |
| 2024 | Centering the Value of Every Modality: Towards Efficient and Resilient Modality-Agnostic Semantic Segmentation
Xu Zheng 0002, Yuanhuiyi Lyu, Jiazhou Zhou, Lin Wang 0025 |
ECCV (69) | 1 |
| 2024 | EventBind: Learning a Unified Representation to Bind Them All for Event-Based Open-World Understanding
Jiazhou Zhou, Xu Zheng 0002, Yuanhuiyi Lyu, Lin Wang 0025 |
ECCV (70) | 2 |
| 2024 | Transformer-CNN Cohort: Semi-supervised Semantic Segmentation by the Best of Both StudentsabstractThe popular methods for semi-supervised semantic segmentation mostly adopt a unitary network model using convolutional neural networks (CNNs) and enforce consistency of the model’s predictions over perturbations applied to the inputs or model. However, such a learning paradigm suffers from two critical limitations: a) learning the discriminative features for the unlabeled data; b) learning both global and local information from the whole image. In this paper, we propose a novel Semi-supervised Learning (SSL) approach, called Transformer-CNN Cohort (TCC), that consists of two students with one based on the vision transformer (ViT) and the other based on the CNN. Our method subtly incorporates the multi-level consistency regularization on the predictions and the heterogeneous feature spaces via pseudo-labeling for the unlabeled data. First, as the inputs of the ViT student are image patches, the feature maps extracted encode crucial class-wise statistics. To this end, we propose class-aware feature consistency distillation (CFCD) that first leverages the outputs of each student as the pseudo labels and generates class-aware feature (CF) maps for knowledge transfer between the two students. Second, as the ViT student has more uniform representations for all layers, we propose consistency-aware cross distillation (CCD) to transfer knowledge between the pixel-wise predictions from the cohort. We validate the TCC framework on Cityscapes and Pascal VOC 2012 datasets, which outperforms existing SSL methods by a large margin. Project page: https://vlislab22.github.io/TCC/. Xu Zheng 0002, Yunhao Luo 0001, Chong Fu 0001, Kangcheng Liu, Lin Wang 0025 |
ICRA | 1 |
| 2024 | Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event CameraabstractThe ability to detect objects in all lighting (i.e., normal-, over-, and under-exposed) conditions is crucial for real-world applications, such as self-driving. Traditional RGB-based detectors often fail under such varying lighting conditions. Therefore, recent works utilize novel event cameras to supplement or guide the RGB modality; however, these methods typically adopt asymmetric network structures that rely predominantly on the RGB modality, resulting in limited robustness for all-day detection. In this paper, we propose EOLO, a novel object detection framework that achieves robust and efficient all-day detection by fusing both RGB and event modalities. Our EOLO framework is built based on a lightweight spiking neural network (SNN) to efficiently leverage the asynchronous property of events. Buttressed by it, we first introduce an Event Temporal Attention (ETA) module to learn the high temporal information from events while preserving crucial edge information. Secondly, as different modalities exhibit varying levels of importance under diverse lighting conditions, we propose a novel Symmetric RGB-Event Fusion (SREF) module to effectively fuse RGB-Event features without relying on a specific modality, thus ensuring a balanced and adaptive fusion for all-day detection. In addition, to compensate for the lack of paired RGB-Event datasets for all-day training and evaluation, we propose an event synthesis approach based on the randomized optical flow that allows for directly generating the event frame from a single exposure image. We further build two new datasets, E-MSCOCO and E-VOC based on the popular benchmarks MSCOCO and PASCAL VOC. Extensive experiments demonstrate that our EOLO outperforms the state-of-the-art detectors, e.g., RENet [1], by a substantial margin (+3.74% mAP50) in all lighting conditions. Our code and datasets will be available at https://vlislab22.github.io/EOLO/. Jiahang Cao, Xu Zheng 0002, Yuanhuiyi Lyu, Renjing Xu, Lin Wang 0025 |
ICRA | 2 |
| 2024 | Frozen is better than learning: A new design of prototype-based classifier for semantic segmentationabstractSemantic segmentation models comprise an encoder to extract features and a classifier for prediction. However, the learning of the classifier suffers from the ambiguity which is caused by two factors: (1) the weights of a classifier for similar categories may have positive similarities lowing the performance for similar categories, named correlation ambiguity, and (2) the classifier is prone to predict the category with a larger ℓ2 norm and vice versa, termed prior ambiguity. To comedy the issues, we propose Category-Basis Prototype (CBP), frozen and mutually orthogonalized prototypes with equalℓ2 norm. Orthogonalization prevents the prototypes from being similar to each other and the equality decouples the prediction from the ℓ2 norm. To better shape the feature space, we propose Online Centroid Contrastive Loss (OCCL) equipped with centroid and category-level losses. Experiments show that our method yields compelling results over two widely applied benchmarks indicating the effectiveness of our methods. Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Hiroshi Murase |
Pattern Recognit. | 4 |
| 2023 | Both Style and Distortion Matter: Dual-Path Unsupervised Domain Adaptation for Panoramic Semantic SegmentationabstractThe ability of scene understanding has sparked active research for panoramic image semantic segmentation. However, the performance is hampered by distortion of the equirectangular projection (ERP) and a lack of pixel-wise annotations. For this reason, some works treat the ERP and pinhole images equally and transfer knowledge from the pinhole to ERP images via unsupervised domain adaptation (UDA). However, they fail to handle the domain gaps caused by: 1) the inherent differences between camera sensors and captured scenes; 2) the distinct image formats (e.g., ERP and pinhole images). In this paper, we propose a novel yet flexible dual-path UDA framework, DPPASS, taking ERP and tangent projection (TP) images as inputs. To reduce the domain gaps, we propose cross-projection and intra-projection training. The cross-projection training includes tangent-wise feature contrastive training and prediction consistency training. That is, the former formulates the features with the same projection locations as positive examples and vice versa, for the models' awareness of distortion, while the latter ensures the consistency of cross-model predictions between the ERP and TP. Moreover, adversarial intra-projection training is proposed to reduce the inherent gap, between the features of the pinhole images and those of the ERP and TP images, respectively. Importantly, the TP path can be freely removed after training, leading to no additional inference cost. Extensive experiments on two benchmarks show that our DPPASS achieves + 1.06% mIoU increment than the state-of-the-art approaches. https://vlis2022.github.io/cvpr23/DPPASS Xu Zheng 0002, Jinjing Zhu, Yexin Liu, Zidong Cao, Chong Fu 0001, Lin Wang 0025 |
CVPR | 1 |
| 2023 | Look at the Neighbor: Distortion-aware Unsupervised Domain Adaptation for Panoramic Semantic SegmentationabstractEndeavors have been recently made to transfer knowledge from the labeled pinhole image domain to the unlabeled panoramic image domain via Unsupervised Domain Adaptation (UDA). The aim is to tackle the domain gaps caused by the style disparities and distortion problem from the non-uniformly distributed pixels of equirectangular projection (ERP). Previous works typically focus on transferring knowledge based on geometric priors with specially designed multi-branch network architectures. As a result, considerable computational costs are induced, and meanwhile, their generalization abilities are profoundly hindered by the variation of distortion among pixels. In this paper, we find that the pixels’ neighborhood regions of the ERP indeed introduce less distortion. Intuitively, we propose a novel UDA framework that can effectively address the distortion problems for panoramic semantic segmentation. In comparison, our method is simpler, easier to implement, and more computationally efficient. Specifically, we propose distortion-aware attention (DA) capturing the neighboring pixel distribution without using any geometric constraints. Moreover, we propose a class-wise feature aggregation (CFA) module to iteratively update the feature representations with a memory bank. As such, the feature similarity between two domains can be consistently optimized. Extensive experiments show that our method achieves new state-of-the-art performance while remarkably reducing 80% parameters. Xu Zheng 0002, Tianbo Pan, Yunhao Luo 0001, Lin Wang 0025 |
ICCV | 1 |
| 2023 | A Good Student is Cooperative and Reliable: CNN-Transformer Collaborative Learning for Semantic SegmentationabstractIn this paper, we strive to answer the question ‘how to collaboratively learn convolutional neural network (CNN)-based and vision transformer (ViT)-based models by selecting and exchanging the reliable knowledge between them for semantic segmentation?’ Accordingly, we propose an online knowledge distillation (KD) framework that can simultaneously learn compact yet effective CNN-based and ViT-based models with two key technical breakthroughs to take full advantage of CNNs and ViT while compensating their limitations. Firstly, we propose heterogeneous feature distillation (HFD) to improve students’ consistency in low-layer feature space by mimicking heterogeneous features between CNNs and ViT. Secondly, to facilitate the two students to learn reliable knowledge from each other, we propose bidirectional selective distillation (BSD) that can dynamically transfer selective knowledge. This is achieved by 1) region-wise BSD determining the directions of knowledge transferred between the corresponding regions in the feature space and 2) pixel-wise BSD discerning which of the prediction knowledge to be transferred in the logit space. Extensive experiments on three benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art online distillation methods by a large margin, and shows its efficacy in learning collaboratively between ViT-based and CNN-based models. Jinjing Zhu, Yunhao Luo 0001, Xu Zheng 0002, Hao Wang 0005, Lin Wang 0025 |
ICCV | 3 |