EDBT 2026 Demo / reviewers in the wild / expert
Yachao Zhang 0001
dblp:40/10584-1
· DBLP profile ↗
54ranked-venue papers
7as first author
53since 2021 · last 2026
0000-0002-6153-5004ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 39 since 2021Artificial intelligence and machine learning · 35 · 6 first-author · 34 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveabstractOpen-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction capability, specifically this pixel-level multimodal alignment. Although existing methods achieve promising results by leveraging CLIP’s vision-language alignment, they rarely investigate the performance boundaries of CLIP for dense prediction from an interpretability mechanisms perspective. In this work, we systematically investigate CLIP's internal mechanisms and identify a critical phenomenon: analogous to human distraction, CLIP diverts significant attention resources from target regions to irrelevant tokens. Our analysis reveals that these tokens arise from dimension-specific over-activation; filtering them enhances CLIP's dense prediction performance. Consequently, we propose Refocusing CLIP (RF-CLIP), a training-free approach that emulates human distraction-refocusing behavior to redirect attention from distraction tokens back to target regions, thereby refining CLIP's multimodal alignment granularity. Our method achieves SOTA performance on eight benchmarks while maintaining high inference efficiency. Jiahao Li 0003, Yang Lu 0009, Yachao Zhang 0001, Fangyong Wang, Yuan Xie 0006, Yanyun Qu |
AAAI | 3 |
| 2026 | PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentationabstract3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe performance degradation in complex, multi-object scenes that are common in real-world settings, hindering practical deployment. Existing methods face two key challenges in complex, multi-object scenes: inadequate parsing of implicit localization cues critical for disambiguating visually similar objects, and ineffective suppression of dynamic spatial interference from co-occurring objects, resulting in degraded grounding accuracy. To address these challenges, we propose PC-CrossDiff, a unified dual-task framework with a dual-level cross-modal differential attention architecture for 3DREC and 3DRES. Specifically, the framework introduces: (i) Point-Level Differential Attention (PLDA) modules that apply bidirectional differential attention between text and point clouds, adaptively extracting implicit localization cues via learnable weights to improve discriminative representation; (ii) Cluster-Level Differential Attention (CLDA) modules that establish a hierarchical attention mechanism to adaptively enhance localization-relevant spatial relationships while suppressing ambiguous or irrelevant spatial relations through a localization-aware differential attention block. To address the scale disparity and conflicting gradients in joint 3DREC–3DRES training, we propose L_DGTL, a unified loss function that explicitly reduces multi-task crosstalk and enables effective parameter sharing across tasks. Our method achieves state-of-the-art performance on the ScanRefer, NR3D, and SR3D benchmarks. Notably, on the Implicit subsets of ScanRefer, it improves the [email protected] score by +10.16% for the 3DREC task, highlighting its strong ability to parse implicit spatial cues. Wenbin Tan 0001, Jiawen Lin, Fangyong Wang, Yuan Xie 0006, Yachao Zhang 0001, Yanyun Qu |
AAAI | 6 |
| 2026 | BeyondSparse: Facilitating Mamba to Enhance Cross-Domain 3D Semantic Segmentation in Adverse WeatherabstractDomain generalization (DG) and domain adaptation (DA) for 3D semantic segmentation enable the model to maintain high performance while avoiding labor-intensive and time-consuming annotation of target-domain data. However, under adverse weather conditions, the injection of spatial noise will affect the reflectivity of LiDAR point clouds, exacerbate domain distribution discrepancies, and degrade the generalization ability of the model. Current methods mainly rely on sparse convolution-based architecture. Due to its limited receptive field, the model captures varying local geometric information when dealing with point clouds of different sparsities, thereby limiting its transferability. To this end, we propose BeyondSparse, a novel cross-domain 3D semantic segmentation method under adverse weather that incorporates a state-space model into a 3D sparse convolution-based architecture, sequentially modeling all features to learn domain-invariant representations. This method consists of two main components: domain feature decoupling and Mamba-based encoder. The former performs feature disentanglement before sequential modeling, while the latter performs global modeling on voxelized point cloud data. In addition, we introduce a token-style augmentation to capture the intrinsic properties of input data. Extensive experimental results demonstrate that our method outperforms SOTA competitors in both DG and DA tasks, for instance, achieving +4.6% and +0.8% mIoU on ``SynLiDAR to SemanticSTF''. Mingwei Xing, Yachao Zhang 0001, Fangyong Wang, Yanyun Qu |
AAAI | 3 |
| 2026 | xMHashSeg: Cross-modal Hash Learning for Training-free Unsupervised LiDAR Semantic Segmentationabstract3D semantic segmentation serves as a fundamental component in many applications, such as autonomous driving and medical image analysis. Although recent methods have advanced the field, adapting these methods to new environments or object categories without extensive retraining remains a significant challenge. To address this, we introduce xMHashSeg, a novel training-free cross-modal LiDAR semantic segmentation framework. xMHashSeg leverages foundation models and non-parametric network to extract features from 2D images and 3D point clouds, subsequently integrating these features through hash learning. Specifically, We develop point-SANN, a novel self-adaption non-parametric network that can extract robust 3D features from raw point clouds, while 2D features are directly extracted through the foundation model DINOv2. To reconcile inconsistencies across different modals, we introduce a Hash Code Learning Module that projects all information into a common hash space, learning a consistent hash code that enhances feature integration. Additionally, depth maps are utilized as an intermediary form between 2D and 3D data to facilitate convergence during hash code learning. Our experimental results on various multi-modality datasets demonstrate that xMHashSeg outperforms zero-shot learning approaches and achieve performance close to that of unsupervised domain adaptation and test-time adaptation methods, without requiring any annotations or additional training. Jialong Zhang 0002, Yachao Zhang 0001, Jiangming Shi, Fangyong Wang, Yanyun Qu |
AAAI | 2 |
| 2026 | Source Free Domain Adaptation For 3D Cross-modal Semantic Segmentation
Jianshe Duan, Yachao Zhang 0001, Yuehui Qu, Yanyun Qu |
ISCAS | 3 |
| 2026 | Instructing visual feature modeling with semantic guidance for 3D visual grounding
Yachao Zhang 0001, Shiran Bian, Jiahao Li 0003, Jiawen Lin, Fangyong Wang, Yuan Xie 0006, Yanyun Qu |
Pattern Recognit. | 1 |
| 2025 | Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic SegmentationabstractSource-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervisory information. To mitigate this problem, Active Learning (AL) is introduced to combine with SFDA, endeavoring to actively label a small set of the most high-quality target points so that models with satisfactory performance can be obtained at an acceptable cost. Nevertheless, several issues remain unresolved, namely when to query new labels during training, what kind of samples deserve labeling to ensure rich information, and where the labels should be distributed to guarantee diversity. Thus we elaborate OmniQuery to omnibearing address the “When, What, and Where” problems about active points querying in source-free domain adaptation for cross-modal 3D semantic segmentation. The method consists of three main components: Query Decider, Point Ranker, and Budget Slicer. The Query Decider determines the optimal timing to query new points by fitting the validation curves during training. The Point Ranker nominates points for annotation by calculating the ambiguity of neighboring points in the feature space. The Budget Slicer allocates the annotation quota, i.e., labeling percentage of the point cloud, to different semantic regions by utilizing the advanced 2D semantic segmentation capabilities of the Segment Anything Model (SAM). Extensive experiments demonstrate the effectiveness of our proposed method, achieving up to 99.64% of fully supervised performance with only 3% of labels, and consistently outperforming comparison methods across various scenarios. Jianxiang Xie, Yachao Zhang 0001, Zhongchao Shi, Jianping Fan 0007, Yuan Xie 0006, Yanyun Qu |
AAAI | 3 |
| 2025 | AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision RewardabstractRecently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex relationship between textual prompts and desired motion outcomes. To address this, we introduce AToM, a framework that enhances the alignment between generated motion and text prompts by leveraging reward from GPT-4Vision. AToM comprises three main stages: Firstly, we construct a dataset MotionPreferthat pairs three types of event-level textual prompts with generated motions, which cover the integrity, temporal relationship and frequency of motion. Secondly, we design a paradigm that utilizes GPT-4Vision for detailed motion annotation, including visual data formatting, task-specific instructions and scoring rules for each sub-task. Finally, we fine-tune an existing text-to-motion model using reinforcement learning guided by this paradigm. Experimental results demonstrate that AToM significantly improves the event-level alignment quality of text-to-motion generation. Project page is available at https://atom-motion.github.io/. Haonan Han, Xiangzuo Wu, Huan Liao, Zunnan Xu, Zhongyuan Hu, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
CVPR | 7 |
| 2025 | Task-Aware Prompt Gradient Projection for Parameter-Efficient Tuning Federated Class-Incremental Learning
Hualong Ke, Jiangming Shi, Yachao Zhang 0001, Fangyong Wang, Yanyun Qu |
ICCV | 3 |
| 2025 | Multi-Schema Proximity Network for Composed Image Retrieval
Jiangming Shi, Xiangbo Yin, Yeyun Chen, Yachao Zhang 0001, Zhizhong Zhang 0001, Yanyun Qu |
ICCV | 4 |
| 2025 | A Plug-And-Play Physical Motion Restoration Approach for In-The-Wild High-Difficulty MotionsabstractExtracting physically plausible 3D human motion from videos is a critical task. Although existing simulation-based motion imitation methods can enhance the physical quality of daily motions estimated from monocular video capture, extending this capability to high-difficulty motions remains an open challenge. This can be attributed to some flawed motion clips in video-based motion capture results and the inherent complexity in modeling high-difficulty motions. Therefore, sensing the advantage of segmentation in localizing human body, we introduce a mask-based motion correction module (MCM) that leverages motion context and video mask to repair flawed motions, producing imitation-friendly motions; and propose a physics-based motion transfer module (PTM), which employs a pretrain and adapt approach for motion imitation, improving physical plausibility with the ability to handle in-the-wild and challenging motions. Our approach is designed as a plug-and-play module to physically refine the video motion capture results, including high-difficulty in-the-wild motions. Finally, to validate our approach, we collected a challenging in-the-wild test set to establish a benchmark, and our method has demonstrated effectiveness on both the new benchmark and existing public datasets.https://physicalmotionrestoration.github.io Youliang Zhang, Ronghui Li, Yachao Zhang 0001, Liang Pan, Yebin Liu, Xiu Li 0001 |
ICCV | 3 |
| 2025 | Prompt-driven Multi-modal Unsupervised Domain Adaptation for 3D Semantic SegmentationabstractExisting multi-modal unsupervised domain adaptation (MM-UDA) methods focus on feature alignment to minimize distribution differences between source and target domains, but this can distort semantic structures and reduce visual feature discriminability. To solve this, we introduce PromptUDA, a prompt-driven MM-UDA method that leverages vision-language models to enhance visual feature discriminability using multi-modal prompts. PromptUDA comprises three crucial components: Multi-modal Data Preparation (MDP), Multi-modal Collaborative Interaction (MCI), and Cross-modal Cross-domain Prototype Contrastive Learning (CPCL). MDP employs a bidirectional fusion approach to process data, which facilitates better learning of multi-modal prompts. MCI enhances the interaction between domain-invariant prompt information and visual features, generating more discriminative semantic information. CPCL further explores the potential of visual features integrated with prompts, leveraging cross-domain and cross-modal advantages to learn domain-invariant features. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in four domain adaptation scenarios. Mingwei Xing, Yachao Zhang 0001, Yanyun Qu |
ICME | 3 |
| 2025 | Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic SegmentationabstractOpen-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in discriminative modeling of latent unseen category. To address this challenge, existing vision-language model (VLM)-based approaches demonstrate commendable performance through pre-trained multi-modal representations. However, the fundamental mechanisms of latent semantic comprehension remain underexplored, making the bottleneck for OVSS. In this work, we initiate a probing experiment to explore distribution patterns and dynamics of latent semantics in VLMs under inductive learning paradigms. Building on these insights, we propose X-Agent, an innovative OVSS framework employing latent semantic-aware ''agent'' to orchestrate cross-modal attention mechanisms, simultaneously optimizing latent semantic dynamic and amplifying its perceptibility. Extensive benchmark evaluations demonstrate that X-Agent achieves state-of-the-art performance while effectively enhancing the latent semantic saliency. Jiahao Li 0003, Yang Lu 0009, Yachao Zhang 0001, Fangyong Wang, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 3 |
| 2025 | SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Groundingabstract3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for real-world applications since eliminating scene-specific training requirements. However, existing zero-shot methods face challenges of spatial-limited reasoning due to reliance on single-view localization, and contextual omissions or detail degradation. To address these issues, we propose SeqVLM, a novel zero-shot 3DVG framework that leverages multi-view real-world scene images with spatial information for target object reasoning. Specifically, SeqVLM first generates 3D instance proposals via a 3D semantic segmentation network and refines them through semantic filtering, retaining only semantic-relevant candidates. A proposal-guided multi-view projection strategy then projects these candidate proposals onto real scene image sequences, preserving spatial relationships and contextual details in the conversion process of 3D point cloud to images. Furthermore, to mitigate VLM computational overload, we implement a dynamic scheduling mechanism that iteratively processes sequances-query prompts, leveraging VLM's cross-modal reasoning capabilities to identify textually specified objects. Experiments on the ScanRefer and Nr3D benchmarks demonstrate state-of-the-art performance, achieving [email protected] scores of 55.6% and 53.2%, surpassing previous zero-shot methods by 4.0% and 5.2%, respectively, which advance 3DVG toward greater generalization and real-world applicability. Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan 0001, Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 5 |
| 2025 | PLATO-TTA: Prototype-Guided Pseudo-Labeling and Adaptive Tuning for Multi-Modal Test-Time Adaptation of 3D SegmentationabstractMulti-modal test-time adaptation (TTA) for 3D semantic segmentation has increasingly become a research hotspot due to its ability to address label dependency and enable rapid adaptation. Existing methods rely on learnable extra components to mitigate reliability bias, however, learning-based approaches in TTA scenarios often lack sufficient training. Moreover, most existing approaches update only normalization layers in the teacher-student framework, which limits their ability to model domain shifts. To overcome these limitations, we propose PLATO-TTA, a novel multi-modal TTA method for 3D semantic segmentation leveraging the native stability in robust prototypes and adaptive tuning of critical teacher-student parameters. The approach contains three key components: Prototype-Guided Pseudo-Labeling (PGPL), Consistency Based Backtracking (CBB), and Domain Specific Updating (DSU). PGPL reduces reliability bias by constructing pseudo-source domain prototypes and computing modality fusion weights based on domain discrepancies. CBB updates all student model parameters while preventing catastrophic forgetting through a parameter backtracking mechanism. DSU selectively updates the teacher model using only domain-specific parameters from the student model, ensuring rapid adaptation and stable guidance. Extensive experiments demonstrate the effectiveness of PLATO-TTA, bringing a 6.3% gain to the SynthiatoSemanticKITTI scenario with severe reliability bias and significant domain discrepancy, and achieve state-of-the-art performance across various domain adaptation scenarios. Jianxiang Xie, Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 3 |
| 2025 | Segment Concealed Objects With Incomplete SupervisionabstractIncompletely-Supervised Concealed Object Segmentation (ISCOS) involves segmenting objects that seamlessly blend into their surrounding environments, utilizing incompletely annotated data, such as weak and semi-annotations, for model training. This task remains highly challenging due to (1) the limited supervision provided by the incompletely annotated training data, and (2) the difficulty of distinguishing concealed objects from the background, which arises from the intrinsic similarities in concealed scenarios. In this paper, we introduce the first unified method for ISCOS to address these challenges. To tackle the issue of incomplete supervision, we propose a unified mean-teacher framework, SEE, that leverages the vision foundation model, "Segment Anything Model (SAM)", to generate pseudo-labels using coarse masks produced by the teacher model as prompts. To mitigate the effect of low-quality segmentation masks, we introduce a series of strategies for pseudo-label generation, storage, and supervision. These strategies aim to produce informative pseudo-labels, store the best pseudo-labels generated, and select the most reliable components to guide the student model, thereby ensuring robust network training. Additionally, to tackle the issue of intrinsic similarity, we design a hybrid-granularity feature grouping module that groups features at different granularities and aggregates these results. By clustering similar features, this module promotes segmentation coherence, facilitating more complete segmentation for both single-object and multiple-object images. We validate the effectiveness of our approach across multiple ISCOS tasks, and experimental results demonstrate that our method achieves state-of-the-art performance. Furthermore, SEE can serve as a plug-and-play solution, enhancing the performance of existing models. Chunming He, Kai Li 0012, Yachao Zhang 0001, Ziyun Yang, Youwei Pang, Longxiang Tang, Chengyu Fang 0001, Yulun Zhang 0001, Linghe Kong, Xiu Li 0001, Sina Farsiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Attention-driven acoustic properties learning for underwater target ranging
Xiaohui Chu, Hantao Zhou, Yan Zhang 0109, Yachao Zhang 0001, Runze Hu, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001, Rongrong Ji |
Pattern Recognit. | 4 |
| 2025 | Fusion-Then-Distillation: Toward Cross-Modal Positive Distillation for Domain Adaptive 3D Semantic SegmentationabstractIn cross-modal unsupervised domain adaptation, a model trained on source-domain data (e.g., synthetic) is adapted to target-domain data (e.g., real-world) without access to target annotation. Previous methods seek to mutually mimic cross-modal outputs in each domain, which enforces a class probability distribution that is agreeable in different domains. However, they overlook the complementarity brought by the heterogeneous fusion in cross-modal learning. In light of this, we propose a novel fusion-then-distillation (FtD++) method to explore cross-modal positive distillation of the source and target domains for 3D semantic segmentation. FtD++ realizes distribution consistency between outputs not only for 2D images and 3D point clouds but also for source-domain and augment-domain. Specially, our method contains three key ingredients. First, we present a model-agnostic feature fusion module to generate the cross-modal fusion representation for establishing a latent space. In this space, two modalities are enforced maximum correlation and complementarity. Second, the proposed cross-modal positive distillation preserves the complete information of multi-modal input and combines the semantic content of the source domain with the style of the target domain, thereby achieving domain-modality alignment. Finally, cross-modal debiased pseudo-labeling is devised to model the uncertainty of pseudo-labels via a self-training manner. Extensive experiments report state-of-the-art results on several domain adaptive scenarios under unsupervised and semi-supervised settings. Code is available athttps://github.com/Barcaaaa/FtD-PlusPlus Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Cross-Cloud Consistency for Weakly Supervised Point Cloud Semantic SegmentationabstractWeakly supervised point cloud semantic segmentation is an increasingly active topic, because fully supervised learning acquires well-labeled point clouds and entails high costs. The existing weakly supervised methods either need meticulously designed data augmentation for self-supervised learning or ignore the negative effects of learning on pseudolabel noises. In this article, by designing different granularity of cross-cloud structures, we propose a cross-cloud consistency method for weakly supervised point cloud semantic segmentation which forms the expectation-maximum (EM) framework. Benefiting from the cross-cloud constraints, our method allows effective learning alternatively between refining pseudolabels and updating network parameters. Specifically, in E-step, we propose a pseudolabel selecting (PLS) strategy based on cross subcloud consistency, improving the credibility of selected pseudolabels explicitly. In M-step, a cross-scene contrastive regularization enforces cross-scene prototypes with the same label in different scenes to be more similar, while keeping prototypes with different labels to be a clear margin, reducing the noise fitting. Finally, we give some insight into the optimization of our method in the EM theoretical way. The proposed method is evaluated on three challenging datasets, where experimental results demonstrate that our method significantly outperforms state-of-the-art weakly supervised competitors. Our code is available online: https://github.com/Yachao-Zhang/Cross-Cloud-Consistency. Yachao Zhang 0001, Yuxiang Lan, Yuan Xie 0006, Cuihua Li, Yanyun Qu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | UniHead: Unifying Multi-Perception for Detection HeadsabstractThe detection head constitutes a pivotal component within object detectors, tasked with executing both classification and localization functions. Regrettably, the commonly used parallel head often lacks omni perceptual capabilities, such as deformation perception (DP), global perception (GP), and cross-task perception (CTP). Despite numerous methods attempting to enhance these abilities from a single aspect, achieving a comprehensive and unified solution remains a significant challenge. In response to this challenge, we develop an innovative detection head, termed UniHead, to unify three perceptual abilities simultaneously. More precisely, our approach: 1) introduces DP, enabling the model to adaptively sample object features; 2) proposes a dual-axial aggregation transformer (DAT) to adeptly model long-range dependencies, thereby achieving GP; and 3) devises a cross-task interaction transformer (CIT) that facilitates interaction between the classification and localization branches, thus aligning the two tasks. As a plug-and-play method, the proposed UniHead can be conveniently integrated with existing detectors. Extensive experiments on the COCO dataset demonstrate that our UniHead can bring significant improvements to many detectors. For instance, the UniHead can obtain +2.7 AP gains in RetinaNet, +2.9 AP gains in FreeAnchor, and +2.1 AP gains in GFL. The code is available at https://github.com/zht8506/UniHead. Hantao Zhou, Rui Yang 0040, Yachao Zhang 0001, Haoran Duan 0001, Yawen Huang, Runze Hu, Xiu Li 0001, Yefeng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlabstractThis study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing during inference. To address this problem, we suggest using speech-derived multimodal priors to improve gesture generation. We introduce a novel method that separates priors from speech and employs multimodal priors as constraints for generating gestures. Our approach utilizes a chain-like modeling method to generate facial blendshapes, body movements, and hand gestures sequentially. Specifically, we incorporate rhythm cues derived from facial deformation and stylization prior based on speech emotions, into the process of generating gestures. By incorporating multimodal priors, our method improves the quality of generated gestures and eliminate the need for expensive setup preparation during inference. Extensive experiments and user studies confirm that our proposed approach achieves state-of-the-art performance. Zunnan Xu, Yachao Zhang 0001, Ronghui Li, Xiu Li 0001 |
AAAI | 2 |
| 2024 | Cross-Modal Match for Language Conditioned 3D Object GroundingabstractLanguage conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion mechanism or bridging the gap between detection and matching. However, several mismatches are ignored, i.e., mismatch in local visual representation and global sentence representation, and mismatch in visual space and corresponding label word space. In this paper, we propose crossmodal match for 3D grounding from mitigating these mismatches perspective. Specifically, to match local visual features with the global description sentence, we propose BEV (Bird’s-eye-view) based global information embedding module. It projects multiple object proposal features into the BEV and the relations of different objects are accessed by the visual transformer which can model both positions and features with long-range dependencies. To circumvent the mismatch in feature spaces of different modalities, we propose crossmodal consistency learning. It performs cross-modal consistency constraints to convert the visual feature space into the label word feature space resulting in easier matching. Besides, we introduce label distillation loss and global distillation loss to drive these matches learning in a distillation way. We evaluate our method in mainstream evaluation settings on three datasets, and the results demonstrate the effectiveness of the proposed method. Yachao Zhang 0001, Runze Hu, Ronghui Li, Yanyun Qu, Yuan Xie 0006, Xiu Li 0001 |
AAAI | 1 |
| 2024 | Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance PrimitivesabstractWe propose Lodge, a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture, and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations between two diffusion models. The first stage is global diffusion, which focuses on comprehending the coarse-level music-dance correlation and production characteristic dance primitives. In contrast, the second-stage is the local diffusion, which parallelly generates detailed motion sequences under the guidance of the dance primitives and choreographic rules. In addition, we propose a Foot Refine Block to optimize the contact between the feet and the ground, enhancing the physical realism of the motion. Our approach can parallelly generate dance sequences of extremely long length, striking a balance between global choreographic patterns and local motion quality and expressiveness. Extensive experiments validate the efficacy of our method. Code, models, and demonstrative video results are available at: https://li-ronghui.github.io/lodge Ronghui Li, Yuxiang Zhang 0006, Yachao Zhang 0001, Hongwen Zhang 0001, Yan Zhang 0002, Yebin Liu, Xiu Li 0001 |
CVPR | 3 |
| 2024 | Multi-memory Matching for Unsupervised Visible-Infrared Person Re-identification
Jiangming Shi, Xiangbo Yin, Yeyun Chen, Yachao Zhang 0001, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ECCV (18) | 4 |
| 2024 | One-Stage Training Generative Paradigm for Generalized Zero-Shot LearningabstractZero-shot learning image classification aims to identify unseen classes not present during training. Generalized zero-shot learning (GZSL) is more in line with realistic scenarios due to its ability of recognizing both seen and unseen classes. Current GZSL methods mostly utilize generative adversarial networks (GANs) but typically follow a two-stage training: first, training the GAN and then, using its synthetic features to train a classifier, which is limited by isolated optimizations rather than federated. We propose a novel One-stage Training Generative Paradigm that incorporates the classifier as a unique synthetic label generator and builds a three-player game involving a generator, discriminator, and classifier, which ensures a unified optimization objective, eliminating the discrete optimization approach of two-stage methods. We also propose a label-attribute classifier that leverages both labels and attributes, surpassing traditional softmax classifiers that only use labels. Our test results show the effectiveness of the proposed methods. Shiran Bian, Xiaofan Li 0008, Yachao Zhang 0001, Jiayong Zhong, Yanyun Qu |
ICASSP | 3 |
| 2024 | Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable AttributeabstractGenerating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address these issues, we propose Text2Avatar, which can generate realistic-style 3D avatars based on the coupled text prompts. Text2Avatar leverages a discrete codebook as an intermediate feature to establish a connection between text and avatars, enabling the disentanglement of features. Furthermore, to alleviate the scarcity of realistic style 3D human avatar data, we utilize a pre-trained unconditional 3D human avatar generation model to obtain a large amount of 3D avatar pseudo data, which allows Text2Avatar to achieve realistic style generation. Experimental results demonstrate that our method can generate realistic 3D avatars from coupled textual data, which is challenging for other existing methods in this field. Chaoqun Gong, Yuqin Dai, Ronghui Li, Achun Bao, Jun Li 0027, Jian Yang 0003, Yachao Zhang 0001, Xiu Li 0001 |
ICASSP | 7 |
| 2024 | Exploring Multi-Modal Control in Music-Driven Dance GenerationabstractExisting music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal control, including genre control, semantic control, and spatial control. First, we decouple the dance generation network from the dance control network, thereby avoiding the degradation in dance quality when adding additional control information. Second, we design specific control strategies for different control information and integrate them into a unified framework. Experimental results show that the proposed dance generation framework outperforms state-of-the-art methods in terms of motion quality and controllability. Ronghui Li, Yuqin Dai, Yachao Zhang 0001, Jun Li 0027, Jian Yang 0003, Xiu Li 0001 |
ICASSP | 3 |
| 2024 | Strategic Preys Make Acute Predators: Enhancing Camouflaged Object Detectors by Generating Camouflaged ObjectsabstractCamouflaged object detection (COD) is the challenging task of identifying camouflaged objects visually blended into surroundings. Albeit achieving remarkable success, existing COD detectors still struggle to obtain precise results in some challenging cases. To handle this problem, we draw inspiration from the prey-vs-predator game that leads preys to develop better camouflage and predators to acquire more acute vision systems and develop algorithms from both the prey side and the predator side. On the prey side, we propose an adversarial training framework, Camouflageator, which introduces an auxiliary generator to generate more camouflaged objects that are harder for a COD method to detect. Camouflageator trains the generator and detector in an adversarial way such that the enhanced auxiliary generator helps produce a stronger detector. On the predator side, we introduce a novel COD method, called Internal Coherence and Edge Guidance (ICEG), which introduces a camouflaged feature coherence module to excavate the internal coherence of camouflaged objects, striving to obtain more complete segmentation results. Additionally, ICEG proposes a novel edge-guided separated calibration module to remove false predictions to avoid obtaining ambiguous boundaries. Extensive experiments show that ICEG outperforms existing COD detectors and Camouflageator is flexible to improve various COD detectors, including ICEG, which brings state-of-the-art COD performance. Chunming He, Kai Li 0012, Yachao Zhang 0001, Yulun Zhang 0001, Chenyu You, Zhenhua Guo 0001, Xiu Li 0001, Martin Danelljan, Fisher Yu 0001 |
ICLR | 3 |
| 2024 | Source-Free Domain Adaptation for Point Cloud Semantic SegmentationabstractPoint cloud semantic segmentation (PCSS) is fundamental in 3D scene understanding. Domain adaptation methods for PCSS enable transferring the knowledge learned from a labeled source domain to an unlabeled target domain with different data distributions. However, they become inapplicable in privacy-preserving scenarios, where we only can leverage a given source model instead of the source data. In this paper, we propose the first source-free domain adaptation PCSS framework to make a given source model generalize to the target domain well. We found existing self-training based source-free domain adaptation methods inevitably lead to confirmation bias and suffer from serious model degradation on PCSS. Therefore, we devise a novel Teacher-Guide Source-Free (TGSF) framework to conquer the above two challenges, including the bidirectional pseudo-label selection and multi-level target consistency learning. Extensive experiments on three benchmarks verify that TGSF achieves state-of-the-art performance, and even outperforms those methods that access the source data. Jianshe Duan, Yachao Zhang 0001, Yanyun Qu |
ICME | 2 |
| 2024 | Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion PriorsabstractReconstructing 3D objects from a single image guided by pretrained diffusion models has demonstrated promising outcomes. However, due to utilizing the case-agnostic rigid strategy, their generalization ability to arbitrary cases and the 3D consistency of reconstruction are still poor. In this work, we propose Consistent123, a case-aware two-stage method for highly consistent 3D asset reconstruction from one image with both 2D and 3D diffusion priors. In the first stage, Consistent123 utilizes only 3D structural priors for sufficient geometry exploitation, with a CLIP-based case-aware adaptive detection mechanism embedded within this process. In the second stage, 2D texture priors are introduced and progressively take on a dominant guiding role, delicately sculpting the details of the 3D model. Consistent123 aligns more closely with the evolving trends in guidance requirements, adaptively providing adequate 3D geometric initialization and suitable 2D texture refinement for different objects. Consistent123 can obtain highly 3D-consistent reconstruction and exhibits strong generalization ability across various objects. Qualitative and quantitative experiments show that our method significantly outperforms state-of-the-art image-to-3D methods. Yukang Lin, Haonan Han, Chaoqun Gong, Zunnan Xu, Yachao Zhang 0001, Xiu Li 0001 |
ACM Multimedia | 5 |
| 2024 | CLIP2UDA: Making Frozen CLIP Reward Unsupervised Domain Adaptation in 3D Semantic SegmentationabstractMulti-modal Unsupervised Domain Adaptation (MM-UDA) for large-scale 3D semantic segmentation involves adapting 2D and 3D models to a target domain without labels, which significantly reduces the labor-intensive annotations. Existing MM-UDA methods have often attempted to mitigate the domain discrepancy by aligning features between the source and target data. However, this implementation falls short when applied to image perception due to the susceptibility of images to environmental changes compared to point clouds. To mitigate this limitation, in this work, we explore the potentials of an off-the-shelf Contrastive Language-Image Pre-training (CLIP) model with rich whilst heterogeneous knowledge. To make CLIP task-specific, we propose a top-performing method, dubbed CLIP2UDA, which makes frozen CLIP reward unsupervised domain adaptation in 3D semantic segmentation. Specifically, CLIP2UDA alternates between two steps during adaptation: (a) Learning task-specific prompt. 2D features response from the visual encoder are employed to initiate the learning of adaptive text prompt of each domain, and (b) Learning multi-modal domain-invariant representations. These representations interact hierarchically in the shared decoder to obtain unified 2D visual predictions. This enhancement allows for effective alignment between the modality-specific 3D and unified feature space via cross-modal mutual learning. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in several widely-recognized adaptation scenarios. Code is available at: https://github.com/Barcaaaa/CLIP2UDA. Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 3 |
| 2024 | Robust Pseudo-label Learning with Neighbor Relation for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised Visible-Infrared Person Re-identification (USVI-ReID) presents a formidable challenge, which aims to match pedestrian images across visible and infrared modalities without any annotations. Recently, clustered pseudo-label methods have become predominant in USVI-ReID, although the inherent noise in pseudo-labels presents a significant obstacle. Most existing works primarily focus on shielding the model from the harmful effects of noise, neglecting to calibrate noisy pseudo-labels usually associated with hard samples, which will compromise the robustness of the model. To address this issue, we design a Robust Pseudo-label Learning with Neighbor Relation (RPNR) framework for USVI-ReID. To be specific, we first introduce a straightforward yet potent Noisy Pseudo-label Calibration module to correct noisy pseudo-labels. Due to the high intra-class variations, noisy pseudo-labels are difficult to calibrate completely. Therefore, we introduce a Neighbor Relation Learning module to reduce high intra-class variations by modeling potential interactions between all samples. Subsequently, we devise an Optimal Transport Prototype Matching module to establish reliable cross-modality correspondences. On that basis, we design a Memory Hybrid Learning module to jointly learn modality-specific and modality-invariant information. Comprehensive experiments conducted on two widely recognized benchmarks, SYSU-MM01 and RegDB, demonstrate that RPNR outperforms the current state-of-the-art GUR with an average Rank-1 improvement of 10.3%. The code is available at https://github.com/XiangboYin/RPNR. Xiangbo Yin, Jiangming Shi, Yachao Zhang 0001, Yang Lu 0009, Zhizhong Zhang 0001, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 3 |
| 2024 | Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationabstractUnsupervised visible-infrared person re-identification (USVI-ReID) aims to match specified persons in infrared images to visible images without annotations, and vice versa. USVI-ReID is a challenging yet underexplored task. Most existing methods address the USVI-ReID through cluster-based contrastive learning, which simply employs the cluster center to represent an individual. However, the cluster center primarily focuses on commonality, overlooking divergence and variety. To address the problem, we propose a Progressive Contrastive Learning with Hard and Dynamic Prototypes for USVI-ReID. In brief, we generate the hard prototype by selecting the sample with the maximum distance from the cluster center. We reveal that the inclusion of the hard prototype in contrastive loss helps to emphasize divergence. Additionally, instead of rigidly aligning query images to a specific prototype, we generate the dynamic prototype by randomly picking samples within a cluster. The dynamic prototype is used to encourage variety. Finally, we introduce a progressive learning strategy to gradually shift the model's attention towards divergence and variety, avoiding cluster deterioration. Extensive experiments conducted on the publicly available SYSU-MM01 and RegDB datasets validate the effectiveness of the proposed method. Jiangming Shi, Xiangbo Yin, Yachao Zhang 0001, Zhizhong Zhang 0001, Yuan Xie 0001, Yanyun Qu |
NeurIPS | 3 |
| 2024 | UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models Priorabstract3D semantic segmentation using an adapting model trained from a source domain with or without accessing unlabeled target-domain data is the fundamental task in computer vision, containing domain adaptation and domain generalization.
The essence of simultaneously solving cross-domain tasks is to enhance the generalizability of the encoder.
In light of this, we propose a groundbreaking universal method with the help of off-the-shelf Visual Foundation Models (VFMs) to boost the adaptability and generalizability of cross-domain 3D semantic segmentation, dubbed $\textbf{UniDSeg}$.
Our method explores the VFMs prior and how to harness them, aiming to inherit the recognition ability of VFMs.
Specifically, this method introduces layer-wise learnable blocks to the VFMs, which hinges on alternately learning two representations during training: (i) Learning visual prompt. The 3D-to-2D transitional prior and task-shared knowledge is captured from the prompt space, and then (ii) Learning deep query. Spatial Tunability is constructed to the representation of distinct instances driven by prompts in the query space.
Integrating these representations into a cross-modal learning framework, UniDSeg efficiently mitigates the domain gap between 2D and 3D modalities, achieving unified cross-domain 3D semantic segmentation.
Extensive experiments demonstrate the effectiveness of our method across widely recognized tasks and datasets, all achieving superior performance over state-of-the-art methods. Remarkably, UniDSeg achieves 57.5\%/54.4\% mIoU on ``A2D2/sKITTI'' for domain adaptive/generalized tasks. Code is available at https://github.com/Barcaaaa/UniDSeg. Mingwei Xing, Yachao Zhang 0001, Xiaotong Luo, Yuan Xie 0006, Yanyun Qu |
NeurIPS | 3 |
| 2024 | MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space ModelsabstractGesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality.
Recent advancements have utilized the diffusion model to improve gesture synthesis.
However, the high computational complexity of these techniques limits the application in reality.
In this study, we explore the potential of state space models (SSMs).
Direct application of SSMs in gesture synthesis encounters difficulties, which stem primarily from the diverse movement dynamics of various body parts.
The generated gestures may also exhibit unnatural jittering issues.
To address these, we implement a two-stage modeling strategy with discrete motion priors to enhance the quality of gestures.
Built upon the selective scan mechanism, we introduce MambaTalk, which integrates hybrid fusion modules, local and global scans to refine latent space representations.
Subjective and objective experiments demonstrate that our method surpasses the performance of state-of-the-art models. Our project is publicly available at~\url{https://kkakkkka.github.io/MambaTalk/}. Zunnan Xu, Yukang Lin, Haonan Han, Ronghui Li, Yachao Zhang 0001, Xiu Li 0001 |
NeurIPS | 6 |
| 2024 | Perturbed Progressive Learning for Semisupervised Defect SegmentationabstractRecently, with the development of intelligent manufacturing, the demand for surface defect inspection has been increasing. Deep learning has achieved promising results in defect inspection. However, due to the rareness of defect data and the difficulties of pixelwise annotation, the existing supervised defect inspection methods are too inferior to be implemented in practice. To solve the problem of defect segmentation with few labeled data, we propose a simple and efficient method for semisupervised defect segmentation (SSDS), named perturbed progressive learning (PPL). On the one hand, PPL decouples the predictions of student and teacher networks as well as alleviates overfitting on noisy pseudo-labels. On the other hand, PPL encourages consistency across various perturbations in a broader stagewise scope, alleviating drift caused by the noisy pseudo-labels. Specifically, PPL contains two training stages. In the first stage, the teacher network gives the unlabeled data with pseudo-labels that are divided into the easy and hard groups. The labeled data and the unlabeled data in the easy group with their perturbation are both used to train for a better-performing student network. In the second stage, the unlabeled data in the hard group are predicted by the obtained student network, so the refined pseudo-labeled data are enlarged. All the pseudo-labeling data and labeled data with their perturbation are used to retrain the student network, progressively improving the defect feature representation. We build a mobile screen defect dataset (MSDD-3) with three classes of defects. PPL is implemented on MSDD-3 as well as other public datasets. Extensive experimental results demonstrate that PPL significantly surpasses the state-of-the-art methods across all evaluation partition protocols. Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Zongze Wu 0001, Yanyun Qu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Learning All-In Collaborative Multiview Binary Representation for ClusteringabstractMultiview clustering via binary representation has attracted intensive attention due to its effectiveness in handling large-scale multiple view data. However, these kind of clustering approaches usually ignore a very important potential high-order correlation in discrete representation learning. In this article, we propose a novel all-in collaborative multiview binary representation for clustering (AC-MVBC) framework, where multiview collaborative binary representation and clustering structure are learned in a joint manner. Specifically, using a new type of tensor low-rank constraint, the high-order collaborations, i.e., cross-view and inner view collaborations, can be effectively captured in our model. Moreover, by incorporating the Bregman discrepancy, the projective consistency among different views can be guaranteed to achieve a more powerful binary representation. An efficient optimization algorithm is also proposed to solve the objective function with fast convergence empirically. Experimental results on several challenge datasets demonstrate that the proposed method has achieved highly competent performance compared with the state-of-the-art multiview clustering (MVC) methods while maintaining low computational and memory requirements. Yachao Zhang 0001, Yuan Xie 0006, Cuihua Li, Zongze Wu 0001, Yanyun Qu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Weakly Supervised 3D Segmentation via Receptive-Driven Pseudo Label Consistency and Structural ConsistencyabstractAs manual point-wise label is time and labor-intensive for fully supervised large-scale point cloud semantic segmentation, weakly supervised method is increasingly active. However, existing methods fail to generate high-quality pseudo labels effectively, leading to unsatisfactory results. In this paper, we propose a weakly supervised point cloud semantic segmentation framework via receptive-driven pseudo label consistency and structural consistency to mine potential knowledge. Specifically, we propose three consistency contrains: pseudo label consistency among different scales, semantic structure consistency between intra-class features and class-level relation structure consistency between pair-wise categories. Three consistency constraints are jointly used to effectively prepares and utilizes pseudo labels simultaneously for stable training. Finally, extensive experimental results on three challenging datasets demonstrate that our method significantly outperforms state-of-the-art weakly supervised methods and even achieves comparable performance to the fully supervised methods. Yuxiang Lan, Yachao Zhang 0001, Yanyun Qu, Cong Wang 0039, Yuan Xie 0006, Zongze Wu 0001 |
AAAI | 2 |
| 2023 | Camouflaged Object Detection with Feature Decomposition and Edge ReconstructionabstractCamouflaged object detection (COD) aims to address the tough issue of identifying camouflaged objects visually blended into the surrounding backgrounds. COD is a challenging task due to the intrinsic similarity of camouflaged objects with the background, as well as their ambiguous boundaries. Existing approaches to this problem have developed various techniques to mimic the human visual system. Albeit effective in many cases, these methods still struggle when camouflaged objects are so deceptive to the vision system. In this paper, we propose the FEature Decomposition and Edge Reconstruction (FEDER) model for COD. The FEDER model addresses the intrinsic similarity of foreground and background by decomposing the features into different frequency bands using learnable wavelets. It then focuses on the most informative bands to mine subtle cues that differentiate foreground and background. To achieve this, a frequency attention module and a guidance-based feature aggregation module are developed. To combat the ambiguous boundary problem, we propose to learn an auxiliary edge reconstruction task alongside the COD task. We design an ordinary differential equation-inspired edge reconstruction module that generates exact edges. By learning the auxiliary task in conjunction with the COD task, the FEDER model can generate precise prediction maps with accurate object boundaries. Experiments show that our FEDER model significantly outperforms state-of-the-art methods with cheaper computational and memory costs. The code will be available at https://github.com/ChunmingHe/FEDER. Chunming He, Kai Li 0012, Yachao Zhang 0001, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001 |
CVPR | 3 |
| 2023 | Sample-Aware Knowledge Distillation for Long-Tailed LearningabstractImage classification for long-tailed scenarios has attracted more attention because its distribution is more similar to real-world image data. From the perspective of solving imbalance at the sample level, we propose a simple but effective method, named Sample-aware Knowledge Distillation (SAKD), which includes Selective Knowledge Distillation module and Stable Feature Center Learning module. The former conducts knowledge distillation at the sample-level by selecting samples, in which whether the sample needs to be distilled and to what extent is determined by evaluating the teacher network’s predictions for this sample. The latter is used to obtaining the stable feature center and making the feature center free from perturbation by hard samples, then further improving the classification boundary. We conduct extensive experiments on several long-tailed benchmark datasets and these results demonstrate that SAKD is effective. In addition, our SFCL module can be combined with other methods and also improve their performance. Shanshan Zheng, Yachao Zhang 0001, Hongyi Huang, Yanyun Qu |
ICASSP | 2 |
| 2023 | Efficient Converted Spiking Neural Network for 3D and 2D ClassificationabstractSpiking Neural Networks (SNNs) have attracted enormous research interest due to their low-power and biologically plausible nature. Existing ANN-SNN conversion methods can achieve lossless conversion by converting a well-trained Artificial Neural Network (ANN) into an SNN. However, converted SNN requires a large amount of time steps to achieve competitive performance with the well-trained ANN, which means a large latency. In this paper, we propose an efficient unified ANN-SNN conversion method for point cloud classification and image classification to significantly reduce the time step to meet the fast and lossless ANN-SNN transformation. Specifically, we first adaptively adjust the threshold according to the activation state of spiking neurons, ensuring a certain proportion of spiking neurons are activated at each time step to reduce the time for accumulation of membrane potential. Next, we use an adaptive firing mechanism to enlarge the range of spiking output, getting more discrimination features in short time steps. Extensive experimental results on challenging point cloud and image datasets demonstrate that the suggested approach significantly outmatches state-of-the-art ANN-SNN conversion based methods. Yuxiang Lan, Yachao Zhang 0001, Xu Ma 0005, Yanyun Qu, Yun Fu 0001 |
ICCV | 2 |
| 2023 | BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic SegmentationabstractCross-modal Unsupervised Domain Adaptation aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in an unknown domain. However, the training of these methods relies on access to target samples, meaning the trained model only works in a specific target domain. In light of this, we propose cross-modal learning under bird’s-eye view for Domain Generalization (DG) of 3D semantic segmentation, called BEV-DG. DG is more challenging because the model cannot access the target domain during training, meaning it needs to rely on cross-modal learning to alleviate the domain gap. Since 3D semantic segmentation requires the classification of each point, existing cross-modal learning is directly conducted point-to-point, which is sensitive to the misalignment in projections between pixels and points. To this end, our approach aims to optimize domain-irrelevant representation modeling with the aid of cross-modal learning under bird’s-eye view. We propose BEV-based Area-to-area Fusion (BAF) to conduct cross-modal learning under bird’s-eye view, which has a higher fault tolerance for point-level misalignment. Furthermore, to model domain-irrelevant representations, we propose BEV-driven Domain Contrastive Learning (BDCL) with the help of cross-modal learning under bird’s-eye view. We design three domain generalization settings based on three 3D datasets, and BEV-DG significantly outperforms state-of-the-art competitors with tremendous margins in all settings. Miaoyu Li, Yachao Zhang 0001, Xu Ma 0005, Yanyun Qu, Yun Fu 0001 |
ICCV | 2 |
| 2023 | FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationabstractGenerating full-body and multi-genre dance sequences from given music is a challenging task, due to the limitations of existing datasets and the inherent complexity of the fine-grained hand motion and dance genres. To address these problems, we propose FineDance, which contains 14.6 hours of music-dance paired data, with fine-grained hand motions, fine-grained genres (22 dance genres), and accurate posture. To the best of our knowledge, FineDance is the largest music-dance paired dataset with the most dance genres. Additionally, to address monotonous and unnatural hand movements existing in previous methods, we propose a full-body dance generation network, which utilizes the diverse generation capabilities of the diffusion model to solve monotonous problems, and use expert nets to solve unreal problems. To further enhance the genre-matching and long-term stability of generated dances, we propose a Genre&Coherent aware Retrieval Module. Besides, we propose a novel metric named Genre Matching Score to evaluate the genre-matching degree between dance and music. Quantitative and qualitative experiments demonstrate the quality of FineDance, and the state-of-the-art performance of FineNet. The FineDance Dataset and more qualitative samples can be found at website. Ronghui Li, Junfan Zhao, Yachao Zhang 0001, Mingyang Su, Zeping Ren, Yansong Tang, Xiu Li 0001 |
ICCV | 3 |
| 2023 | Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the annotation in single-modality. To reduce the extensive cost of annotation, we explore two practical semi-supervised settings: uni-semi-supervised (annotating only visible images) and bi-semi-supervised (annotating partially in both modalities). These two semi-supervised settings face two challenges due to the large cross-modality discrepancies and the lack of correspondence supervision between visible and infrared images. Thus, it is diffi-cult to generate reliable pseudo-labels and learn modality-invariant features from noise pseudo-labels. In this paper, we propose a dual pseudo-label interactive self-training (DPIS) for these two semi-supervised VI-ReID. Our DPIS integrates two pseudo-labels generated by distinct models into a hybrid pseudo-label for unlabeled data. However, the hybrid pseudo-label still inevitably contains noise. To eliminate the negative effect of noise pseudo-labels, we introduce three modules: noise label penalty (NLP), noise correspondence calibration (NCC), and unreliable anchor learning (UAL). Specifically, NLP penalizes noise labels, NCC calibrates noisy correspondences, and UAL mines the hard-to-discriminate features. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that our DPIS achieves impressive performance under these two semi-supervised settings. Jiangming Shi, Yachao Zhang 0001, Xiangbo Yin, Yuan Xie 0006, Zhizhong Zhang 0001, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu |
ICCV | 2 |
| 2023 | VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot LearningabstractUnlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to the extreme training imbalance. Recently, some feature generation methods introduce metric learning to enhance the discriminability of visual features. Although these methods achieve good results, they focus only on metric learning in the visual feature space to enhance features and ignore the association between the feature space and the semantic space. Since the GZSL method uses semantics as prior knowledge to migrate visual knowledge to unseen classes, the consistency between visual space and semantic space is critical. To this end, we propose relational metric learning which can relate the metrics in the two spaces and make the distribution of the two spaces more consistent. Based on the generation method and relational metric learning, we proposed a novel GZSL method, termed VS-Boost, which can effectively boost the association between vision and semantics. The experimental results demonstrate that our method is effective and achieves significant gains on five benchmark datasets compared with the state-of-the-art methods. Xiaofan Li 0008, Yachao Zhang 0001, Shiran Bian, Yanyun Qu, Yuan Xie 0006, Zhongchao Shi, Jianping Fan 0007 |
IJCAI | 2 |
| 2023 | Cross-modal Unsupervised Domain Adaptation for 3D Semantic Segmentation via Bidirectional Fusion-then-DistillationabstractCross-modal Unsupervised Domain Adaptation (UDA) becomes a research hotspot because it reduces the laborious annotation of target domain samples. Existing methods only mutually mimic the outputs of cross-modality in each domain, which enforces the class probability distribution agreeable in different domains. However, these methods ignore the complementarity brought by the modality fusion representation in cross-modal learning. In this paper, we propose a cross-modal UDA method for 3D semantic segmentation via Bidirectional Fusion-then-Distillation, named BFtD-xMUDA, which explores cross-modal fusion in UDA and realizes distribution consistency between outputs of two domains not only for 2D image and 3D point cloud but also for 2D/3D and fusion. Our method contains three significant components: Model-agnostic Feature Fusion Module (MFFM), Bidirectional Distillation (B-Distill), and Cross-modal Debiased Pseudo-Labeling (xDPL). MFFM is employed to generate cross-modal fusion features for establishing a latent space, which enforces maximum correlation and complementarity between two heterogeneous modalities. B-Distill is introduced to exploit bidirectional knowledge distillation which includes cross-modality and cross-domain fusion distillation, and well-achieving domain-modality alignment. xDPL is designed to model the uncertainty of pseudo-labels by self-training scheme. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in several adaptation scenarios. Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu |
ACM Multimedia | 3 |
| 2023 | Weakly-Supervised Concealed Object Segmentation with SAM-based Pseudo Labeling and Multi-scale Feature GroupingabstractWeakly-Supervised Concealed Object Segmentation (WSCOS) aims to segment objects well blended with surrounding environments using sparsely-annotated data for model training. It remains a challenging task since (1) it is hard to distinguish concealed objects from the background due to the intrinsic similarity and (2) the sparsely-annotated training data only provide weak supervision for model learning. In this paper, we propose a new WSCOS method to address these two challenges. To tackle the intrinsic similarity challenge, we design a multi-scale feature grouping module that first groups features at different granularities and then aggregates these grouping results. By grouping similar features together, it encourages segmentation coherence, helping obtain complete segmentation results for both single and multiple-object images. For the weak supervision challenge, we utilize the recently-proposed vision foundation model, ``Segment Anything Model (SAM)'', and use the provided sparse annotations as prompts to generate segmentation masks, which are used to train the model. To alleviate the impact of low-quality segmentation masks, we further propose a series of strategies, including multi-augmentation result ensemble, entropy-based pixel-level weighting, and entropy-based image-level selection. These strategies help provide more reliable supervision to train the segmentation model. We verify the effectiveness of our method on various WSCOS tasks, and experiments demonstrate that our method achieves state-of-the-art performance on these tasks. Chunming He, Kai Li 0012, Yachao Zhang 0001, Guoxia Xu, Longxiang Tang, Yulun Zhang 0001, Zhenhua Guo 0001, Xiu Li 0001 |
NeurIPS | 3 |
| 2023 | RICH: Robust Implicit Clothed Humans Reconstruction from Multi-scale Spatial Cues
Yukang Lin, Ronghui Li, Kedi Lyu, Yachao Zhang 0001, Xiu Li 0001 |
PRCV (2) | 4 |
| 2023 | An object perception and positioning method via deep perception learning object detectionabstractAbstract One of the fundamental problems when building perception systems for robot is to be able to provide semantic information as well as positioning in three‐dimensional (3D) space. However, two‐dimensional (2D) object detectors only can provide the semantic information and pixel coordinate in 2D space. While, the depth image can reflect the relative distance, and the semantic description of the object is poor. In this article, a novel object perception and positioning method via deep perception learning object detection is proposed. First, the RGB image and depth image are collected through the Kinect, and the depth image is processed to ensure the robustness of the model. Then, the RGB image can obtain the object semantic and pixel location information through an object detector based on deep learning. Finally, the object size measurement and 3D positioning are realized by combining the pixel location and the depth information. As a result, the advantages of very accurate 2D detector and the accurate depth information can be effectively captured in our model. Experimental results demonstrate that our method achieves a high accuracy of size measurement and spatial positioning. Limei Xiao, Yachao Zhang 0001, Weizhe Gao, Dayou Xu, Ce Li 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic SegmentationabstractWith the emergence of multi-modal datasets where LiDAR and camera are synchronized and calibrated, cross-modal Unsupervised Domain Adaptation (UDA) has attracted increasing attention because it reduces the laborious annotation of target domain samples. To alleviate the distribution gap between source and target domains, existing methods conduct feature alignment by using adversarial learning. However, it is well-known to be highly sensitive to hyperparameters and difficult to train. In this paper, we propose a novel model (Dual-Cross) that integrates Cross-Domain Knowledge Distillation (CDKD) and Cross-Modal Knowledge Distillation (CMKD) to mitigate domain shift. Specifically, we design the multi-modal style transfer to convert source image and point cloud to target style. With these synthetic samples as input, we introduce a target-aware teacher network to learn knowledge of the target domain. Then we present dual-cross knowledge distillation when the student is learning on source domain. CDKD constrains teacher and student predictions under same modality to be consistent. It can transfer target-aware knowledge from the teacher to the student, making the student more adaptive to the target domain. CMKD generates hybrid-modal prediction from the teacher predictions and constrains it to be consistent with both 2D and 3D student predictions. It promotes the information interaction between two modalities to make them complement each other. From the evaluation results on various domain adaptation settings, Dual-Cross significantly outperforms both uni-modal and cross-modal state-of-the-art methods. Miaoyu Li, Yachao Zhang 0001, Yuan Xie 0006, Zuodong Gao, Cuihua Li, Zhizhong Zhang 0001, Yanyun Qu |
ACM Multimedia | 2 |
| 2022 | Self-supervised Exclusive Learning for 3D Segmentation with Cross-Modal Unsupervised Domain Adaptationabstract2D-3D unsupervised domain adaptation (UDA) tackles the lack of annotations in a new domain by capitalizing the relationship between 2D and 3D data. Existing methods achieve considerable improvements by performing cross-modality alignment in a modality-agnostic way, failing to exploit modality-specific characteristic for modeling complementarity. In this paper, we present self-supervised exclusive learning for cross-modal semantic segmentation under the UDA scenario, which avoids the prohibitive annotation. Specifically, two self-supervised tasks are designed, named "plane-to-spatial'' and "discrete-to-textured''. The former helps the 2D network branch improve the perception of spatial metrics, and the latter supplements structured texture information for the 3D network branch. In this way, modality-specific exclusive information can be effectively learned, and the complementarity of multi-modality is strengthened, resulting in a robust network to different domains. With the help of the self-supervised tasks supervision, we introduce a mixed domain to enhance the perception of the target domain by mixing the patches of the source and target domain samples. Besides, we propose a domain-category adversarial learning with category-wise discriminators by constructing the category prototypes for learning domain-invariant features. We evaluate our method on various multi-modality domain adaptation settings, where our results significantly outperform both uni-modality and multi-modality state-of-the-art competitors. Yachao Zhang 0001, Miaoyu Li, Yuan Xie 0006, Cuihua Li, Cong Wang 0039, Zhizhong Zhang 0001, Yanyun Qu |
ACM Multimedia | 1 |
| 2021 | Weakly Supervised Semantic Segmentation for Large-Scale Point CloudabstractExisting methods for large-scale point cloud semantic segmentation require expensive, tedious and error-prone manual point-wise annotation. Intuitively, weakly supervised training is a direct solution to reduce the labeling costs. However, for weakly supervised large-scale point cloud semantic segmentation, too few annotations will inevitably lead to ineffective learning of network. We propose an effective weakly supervised method containing two components to solve the above problem. Firstly, we construct a pretext task, \textit{i.e.,} point cloud colorization, with a self-supervised training manner to transfer the learned prior knowledge from a large amount of unlabeled point cloud to a weakly supervised network. In this way, the representation capability of the weakly supervised network can be improved by knowledge from a heterogeneous task. Besides, to generative pseudo label for unlabeled data, a sparse label propagation mechanism is proposed with the help of generated class prototypes, which is used to measure the classification confidence of unlabeled point. Our method is evaluated on large-scale point cloud datasets with different scenarios including indoor and outdoor. The experimental results show the large gain against existing weakly supervised methods and comparable results to fully supervised methods. Yachao Zhang 0001, Yuan Xie 0006, Yanyun Qu, Cuihua Li, Tao Mei 0001 |
AAAI | 1 |
| 2021 | Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic SegmentationabstractLarge-scale point cloud semantic segmentation has wide applications. Current popular researches mainly focus on fully supervised learning which demands expensive and tedious manual point-wise annotation. Weakly supervised learning is an alternative way to avoid this exhausting an-notation. However, for large-scale point clouds with few labeled points, the network is difficult to extract discriminative features for unlabeled points, as well as the regularization of topology between labeled and unlabeled points is usually ignored, resulting in incorrect segmentation results.To address this problem, we propose a perturbed self-distillation (PSD) framework. Specifically, inspired by self-supervised learning, we construct the perturbed branch and enforce the predictive consistency among the perturbed branch and original branch. In this way, the graph topology of the whole point cloud can be effectively established by the introduced auxiliary supervision, such that the in-formation propagation between the labeled and unlabeled points will be realized. Besides point-level supervision, we present a well-integrated context-aware module to explicitly regularize the affinity correlation of labeled points. Therefore, the graph topology of the point cloud can be further refined. The experimental results evaluated on three large-scale datasets show the large gain (3.0% on average) against recent weakly supervised methods and comparable results to some fully supervised methods. Yachao Zhang 0001, Yanyun Qu, Yuan Xie 0006, Zonghao Li, Shanshan Zheng, Cuihua Li |
ICCV | 1 |
| 2018 | 3D Reconstruction of Indoor Scenes via Image Registration
Ce Li 0001, Yachao Zhang 0001, Hao Liu 0060, Yanyun Qu |
Neural Process. Lett. | 3 |