EDBT 2026 Demo / reviewers in the wild / expert
Lechao Cheng
dblp:165/9781
· DBLP profile ↗
71ranked-venue papers
4as first author
68since 2021 · last 2026
0000-0002-7546-9052ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 4 first-author · 48 since 2021Artificial intelligence and machine learning · 40 · 2 first-author · 39 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningabstractOpen-world 3D scene understanding is fundamentally challenging for vision and robotics, due to the constraints of closed-vocabulary supervision and static annotations. To address this, we propose a unified framework for Open-World 3D Scene Graph Generation with Retrieval-Augmented Reasoning, which enables generalizable and interactive 3D scene understanding. Our method integrates vision-language models with retrieval-based reasoning to support multimodal exploration and language-guided interaction. The framework comprises two key components: (1) a dynamic scene graph generation module that detects objects and infers semantic relationships without fixed label sets, and (2) a retrieval-augmented reasoning pipeline that encodes scene graphs into a vector database to support text/image-conditioned queries. We evaluate our method on 3DSSG and Replica benchmarks across four tasks—scene question answering, visual grounding, instance retrieval, and task planning—demonstrating robust generalization and superior performance in diverse environments. Our results highlight the effectiveness of combining open-vocabulary perception with retrieval-based reasoning for scalable 3D scene understanding. Fei Yu 0012, Shengeng Tang, Lechao Cheng |
AAAI | 5 |
| 2026 | RepAttn3D: Re-parameterizing 3D attention with spatiotemporal augmentation for video understanding
Xiusheng Lu, Lechao Cheng, Sicheng Zhao, Ying Zheng 0009, Yongheng Wang, Guiguang Ding, Mingli Song |
Neural Networks | 2 |
| 2026 | High-frequency structure transformer for magnetic resonance image super-resolution
Chaowei Fang, Bolin Fu, De Cheng, Lechao Cheng, Dingwen Zhang |
Pattern Recognit. | 4 |
| 2026 | Dual-Domain Adaptation Networks for Realistic Image Super-ResolutionabstractRealistic image super-resolution (SR) focuses on transforming real-world low-resolution (LR) images into high-resolution (HR) ones, handling more complex degradation patterns than synthetic SR tasks. This is critical for applications like surveillance, medical imaging, and consumer electronics. However, current methods struggle with limited real-world LR-HR data, impacting the learning of basic image features. Pre-trained SR models from large-scale synthetic datasets offer valuable prior knowledge, which can improve generalization, speed up training, and reduce the need for extensive real-world data in realistic SR tasks. In this paper, we introduce a novel approach,Dual-domain Adaptation Networks, which is able to efficiently adapt pre-trained image SR models from simulated to real-world datasets. To achieve this target, we first set up a spatial-domain adaptation strategy through selectively updating parameters of pre-trained models and employing the low-rank adaptation technique to adjust frozen parameters. Recognizing that image super-resolution involves recovering high-frequency components, we further integrate a frequency domain adaptation branch into the adapted model, which combines the spectral data of the input and the spatial-domain backbone's intermediate features to infer HR frequency maps, enhancing the SR result. Experimental evaluations on public realistic image SR benchmarks, including RealSR, D2CRealSR, and DRealSR, demonstrate the superiority of our proposed method over existing state-of-the-art models. Chaowei Fang, Bolin Fu, De Cheng, Lechao Cheng, Guanbin Li |
IEEE Trans. Multim. | 4 |
| 2025 | Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsabstractGenerating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address this, we propose a novel framework, Sign-D2C, that employs a conditional diffusion model to synthesize contextually smooth transition frames, enabling the seamless construction of continuous sign language sequences. Our approach transforms the unsupervised problem of transition frame generation into a supervised training task by simulating the absence of transition frames through random masking of segments in long-duration sign videos. The model learns to predict these masked frames by denoising Gaussian noise, conditioned on the surrounding sign observations, allowing it to handle complex, unstructured transitions. During inference, we apply a linearly interpolating padding strategy that initializes missing frames through interpolation between boundary frames, providing a stable foundation for iterative refinement by the diffusion model. Extensive experiments on the PHOENIX14T, USTCCSL100, and USTC-SLR500 datasets demonstrate the effectiveness of our method in producing continuous, natural sign language videos. Shengeng Tang, Lechao Cheng, Jingjing Wu 0001, Dan Guo 0001, Richang Hong |
CVPR | 3 |
| 2025 | ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and GroundingabstractWe present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4). Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for accurately manipulation detection and grounding. While existing DGM4methods pay rare attention to the cross-modal alignment, hampering the accuracy of manipulation detecting to step further. To remedy this issue, this work targets to advance the semantic alignment learning to promote this task. Particularly, we utilize the off-the-shelf large models to construct paired image-text pairs, especially for the manipulated instances. Subsequently, a cross-modal alignment learning is performed to enhance the semantic alignment. Besides the explicit auxiliary clues, we further design a Manipulation-Guided Cross Attention (MGCA) to provide implicit guidance for augmenting the manipulation perceiving. With the grounding truth available during training, MGCA encourages the model to concentrate more on manipulated components while downplaying normal ones, enhancing the model’s ability to capture manipulations. Extensive experiments are conducted on the DGM4dataset, the results demonstrate that our model can surpass the comparison method with a clear margin. Code will be released at https://github.com/CriliasMiller/ASAP. Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo 0001, Meng Wang 0001 |
CVPR | 3 |
| 2025 | DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) achieve remarkable performance in multimodal tasks but suffer from high computational costs due to the large number of visual tokens.Existing pruning methods either apply after visual tokens enter the LLM or perform pre-pruning based solely on visual attention.Both fail to balance efficiency and semantic alignment, as post-pruning incurs redundant computation, while visual-only pre-pruning overlooks multimodal relevance.To address this limitation, we propose Dual-Cue Pruning (DCP), a novel cross-modal pruning framework that jointly considers textual semantics and visual selfattention.DCP consists of a text-aware computation module, which employs a gradientweighted attention mechanism to enhance textvisual alignment, and an image-aware computation module, which utilizes deep-layer selfattention distributions to retain essential structural information.By integrating both cues, DCP adaptively selects the most informative visual tokens, achieving efficient inference acceleration while maintaining strong task performance.Experimental results show that DCP can retain only 25% of the visual tokens, with a minimal performance degradation of 0.063% on LLaVA-1.5-13B,demonstrating its effectiveness in balancing efficiency and accuracy. Zixun Zhang, Yuting Zeng, Chunzhao Xie, Tongxuan Liu, Lechao Cheng |
EMNLP | 7 |
| 2025 | Efficient Asymmetric Shared Low-Rank Adaptation Based on Selective Scanning Vision Mamba for Medical Imaging AnalysisabstractThe Vision Mamba has garnered significant attention for its efficiency and accuracy. Nonetheless, the potential of parameter-efficient tuning in Vision Mamba for medical image analysis remains under-explored. In this work, we investigate the position-sensitive selective scanning mechanism in Vision Mamba and propose a novel low-rank adaptation strategy tailored for planar traversal dependencies. Specifically, we first reshape input patches into sequences following four distinct traversal paths, enabling selective scanning from multiple spatial perspectives. To adapt these paths efficiently, each Visual State Space module is equipped with a shared down-projection matrix and four learnable low-rank up-projection matrices, one for each traversal direction. By combining the shared initial parameters with the direction-specific low-rank components, our approach provides a flexible yet compact way to fine-tune model weights. To further enhance training efficiency, we introduce an iterative optimization strategy that updates the low-rank parameters and task-specific head layers in alternating steps. This strategy encourages faster convergence and reduced computational overhead, making it particularly valuable for large-scale medical image tasks. Extensive experiments on classification, detection, and segmentation benchmarks confirm that our method surpasses existing parameter-efficient approaches. Ziqiu Dong, Yuetong Luo, Yankong Zhang, Shengeng Tang, Lechao Cheng |
ICIP | 6 |
| 2025 | Exploring Effective Unfolding Covering Prompt Tuning for Vision MambaabstractThe Vision Mamba proposed recently has emerged as a popular architecture solution known for its efficiency in computational resources. However, the core designation of parallelized selective scan operation poses a challenge when performing visual prompt tuning on downstream tasks. The difficulty arises from the conflict between the unordered insertion of tokens in visual prompt tuning and the varying importance of tokens at different positions in vision mamba. To alleviate this issue, we propose an Unfolding Covering prompt tuning strategy that effectively customizes downstream tasks. In this work, we explore several visual prompting strategies to further improve performance with limited data. Exhaustive experiments on general tasks like classification and detection have demonstrated the superiority of our approach. Mingwang Wu, Yuetong Luo, Yankong Zhang, Shengeng Tang, Lechao Cheng |
ICIP | 6 |
| 2025 | STD-FD: Spatio-Temporal Distribution Fitting Deviation for AIGC Forgery IdentificationabstractWith the rise of AIGC technologies, particularly diffusion models, highly realistic fake images that can deceive human visual perception has become feasible. Consequently, various forgery detection methods have emerged. However, existing methods treat the generation process of fake images as either a black-box or an auxiliary tool, offering limited insights into its underlying mechanisms. In this paper, we propose Spatio-Temporal Distribution Fitting Deviation (STD-FD) for AIGC forgery detection, which explores the generative process in detail. By decomposing and reconstructing data within generative diffusion models, initial experiments reveal temporal distribution fitting deviations during the image reconstruction process. These deviations are captured through reconstruction noise maps for each spatial semantic unit, derived via a super-resolution algorithm. Critical discriminative patterns, termed DFactors, are identified through statistical modeling of these deviations. Extensive experiments show that STD-FD effectively captures distribution patterns in AIGC-generated data, demonstrating strong robustness and generalizability while outperforming state-of-the-art (SOTA) methods on major datasets. The source code is available at [this link](https://github.com/HengruiLou/STDFD). Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Jie Lei 0002, Lechao Cheng, Jie Song 0011, Mingli Song, Yijun Bei |
ICML | 6 |
| 2025 | Navigating Semantic Drift in Task-Agnostic Class-Incremental LearningabstractClass-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the gap in feature distribution between novel and existing tasks is primarily driven by differences in mean and covariance moments. Building on this insight, we propose a novel semantic drift calibration method that incorporates mean shift compensation and covariance calibration. Specifically, we calculate each class's mean by averaging its sample embeddings and estimate task shifts using weighted embedding changes based on their proximity to the previous mean, effectively capturing mean shifts for all learned classes with each new task. We also apply Mahalanobis distance constraint for covariance calibration, aligning class-specific embedding covariances between old and current networks to mitigate the covariance shift. Additionally, we integrate a feature-level self-distillation approach to enhance generalization. Comprehensive experiments on commonly used datasets demonstrate the effectiveness of our approach. The source code is available at https://github.com/fwu11/MACIL.git. Fangwen Wu, Lechao Cheng, Shengeng Tang, Chaowei Fang, Dingwen Zhang, Meng Wang 0001 |
ICML | 2 |
| 2025 | Knowledge Swapping via Learning and UnlearningabstractWe introduce Knowledge Swapping, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user-specified information, retaining essential knowledge, and acquiring new knowledge simultaneously. By delving into the analysis of knock-on feature hierarchy, we find that incremental learning typically progresses from low-level representations to higher-level semantics, whereas forgetting tends to occur in the opposite direction—starting from high-level semantics and moving down to low-level features. Building upon this, we propose to benchmark the knowledge swapping task with the strategy of Learning Before Forgetting. Comprehensive experiments on various tasks like image classification, object detection, and semantic segmentation validate the effectiveness of the proposed strategy. The source code is available at https://github.com/xingmingyu123456/KnowledgeSwapping. Mingyu Xing, Lechao Cheng, Shengeng Tang, Yaxiong Wang, Zhun Zhong, Meng Wang 0001 |
ICML | 2 |
| 2025 | Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video EditingabstractText-driven video editing powered by generative diffusion models holds significant promise for applications spanning film production, advertising, and beyond. However, the limited expressiveness of pre-trained word embeddings often restricts nuanced edits, especially when targeting novel concepts with specific attributes. In this work, we present a novel Concept-Augmented Textual Inversion (CATI) framework that flexibly integrates new object information from user-provided concept videos. By fine-tuning only the V (Value) projection in attention via Low-Rank Adaptation (LoRA), our approach preserves the original attention distribution of the diffusion model while efficiently incorporating external concept knowledge. To further stabilize editing results and mitigate the issue of attention dispersion when prompt keywords are modified, we introduce a Dual Prior Supervision (DPS) mechanism. DPS supervises cross-attention between the source and target prompts, preventing undesired changes to non-target areas and improving the fidelity of novel concepts. Extensive evaluations demonstrate that our plug-and-play solution not only maintains spatial and temporal consistency but also outperforms state-of-the-art methods in generating lifelike and stable edited videos. The source code is publicly available at https://guomc9.github.io/STIVE-PAGE/. Mingce Guo, Jingxuan He 0001, Yufei Yin, Zhangye Wang, Shengeng Tang, Lechao Cheng |
IJCAI | 6 |
| 2025 | Self-Classification Enhancement and Correction for Weakly Supervised Object DetectionabstractIn recent years, weakly supervised object detection (WSOD) has attracted much attention due to its low labeling cost. The success of recent WSOD models is often ascribed to the two-stage multi-class classification (MCC) task, i.e., multiple instance learning and online classification refinement. Despite achieving non-trivial progresses, these methods overlook potential classification ambiguities between these two MCC tasks and fail to leverage their unique strengths. In this work, we introduce a novel WSOD framework to ameliorate these two issues. For one thing, we propose a self-classification enhancement module that integrates intra-class binary classification (ICBC) to bridge the gap between the two distinct MCC tasks. The ICBC task enhances the network’s discrimination between positive and mis-located samples in a class-wise manner and forges a mutually reinforcing relationship with the MCC task. For another, we propose a self-classification correction algorithm during inference, which combines the results of both MCC tasks to effectively reduce the mis-classified predictions. Extensive experiments on the prevalent VOC 2007 & 2012 datasets demonstrate the superior performance of our framework. Yufei Yin, Lechao Cheng, Wengang Zhou 0001, Jiajun Deng, Houqiang Li |
IJCAI | 2 |
| 2025 | Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training ApproachabstractMicro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluate traditional Semi-Supervised Learning (SSL) methods to SSMAR and find that these methods tend to overfit on inaccurate pseudo-labels, leading to error accumulation and degraded performance. This issue primarily arises from the common practice of directly using the predictions of classifier as pseudo-labels to train the model. To solve this issue, we propose a novel framework, called Asynchronous Pseudo Labeling and Training (APLT), which explicitly separates the pseudo-labeling process from model training. Specifically, we introduce a semi-supervised clustering method during the offline pseudo-labeling phase to generate more accurate pseudo-labels. Moreover, a self-adaptive thresholding strategy is proposed to dynamically filter noisy labels of different classes. We then build a memory-based prototype classifier based on the filtered pseudo-labels, which is fixed and used to guide the subsequent model training phase. By alternating the two pseudo-labeling and model training phases in an asynchronous manner, the model can not only be learned with more accurate pseudo-labels but also avoid the overfitting issue. Experiments on three MAR datasets show that our APLT largely outperforms state-of-the-art SSL methods. For instance, APLT improves accuracy by 14.5% over FixMatch on the MA-12 dataset when using only 50% labeled data. Code is available at https://github.com/zy-hfut/APLT Yan Zhang 0053, Lechao Cheng, Yaxiong Wang, Zhun Zhong, Meng Wang 0001 |
IJCAI | 2 |
| 2025 | A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Lechao Cheng, Jie Lei 0002, Jie Song 0011, Mingli Song, Yijun Bei |
ACM Multimedia | 5 |
| 2025 | Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsabstractThe detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly reflect real-world manipulation patterns: practical attacks typically maintain semantic consistency across modalities, whereas current datasets artificially disrupt cross-modal alignment, creating easily detectable anomalies. To bridge this gap, we pioneer the detection of semantically-coordinated manipulations where visual edits are systematically paired with semantically consistent textual descriptions. Our approach begins with constructing the first Semantic-Aligned Multimodal Manipulation (SAMM) dataset, generated through a two-stage pipeline: 1) applying state-of-the-art image manipulations, followed by 2) generation of contextually-plausible textual narratives that reinforce the visual deception. Building on this foundation, we propose a Retrieval-Augmented Manipulation Detection and Grounding (RamDG) framework. RamDG commences by harnessing external knowledge repositories to retrieve contextual evidence, which serves as the auxiliary texts and encoded together with the inputs through our image forgery grounding and deep manipulation detection modules to trace all manipulations. Extensive experiments demonstrate our framework significantly outperforms existing methods, achieving 2.06% higher detection accuracy on SAMM compared to state-of-the-art approaches. The dataset and code are publicly available at https://github.com/shen8424/SAMM-RamDG-CAP. Jinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu, Zhun Zhong |
ACM Multimedia | 3 |
| 2025 | Beyond General Alignment: Fine-Grained Entity-Centric Image-Text Matching with Multimodal Attentive ExpertsabstractRecent progress in aligning images with texts has achieved remarkable results, however, existing models tend to serve general queries and often fall short when dealing with detailed query requirements. In this paper, we work towards Entity-centric Image-Text Matching (EITM), a finer-grained image-text matching task that aligns texts and images centered around specific entities. The main challenge in EITM lies in bridging the substantial semantic gap between entity-related information in texts and images, which is more pronounced than in general image-text matching problems. To address this challenge, we adopt CLIP as our foundational model and devise a Multimodal Attentive Experts (MMAE)-based contrastive learning to adapt CLIP into an expert for EITM problem. Particularly, the core of our multimodal attentive experts learning is to generate explanation texts by Large Language Models (LLMs) as bridging clues. In specific, we first employ off-the-shelf LLMs to generate explanatory text. This text, along with the original image and text, is then fed into our Multimodal Attentive Experts module to narrow the semantic gap within a unified semantic space. Upon the enriched feature representations generated by MMAE, we have further developed an effective Gated Integrative Image-text Matching (GI-ITM) strategy. GI-ITM utilizes an adaptive gating mechanism to combine features from MMAE, followed by applying image-text matching constraints to enhance the alignment precision. Our method has been extensively evaluated on three social media news benchmarks: N24News, VisualNews, and GoodNews. The experimental results demonstrate that our approach significantly outperforms competing methods. Our code is available at: https://github.com/wangyxxjtu/ETE. Yaxiong Wang, Lianwei Wu, Lechao Cheng, Zhun Zhong, Yujiao Wu, Meng Wang 0001 |
SIGIR | 3 |
| 2025 | Temporal multi-modal knowledge graph generation for link prediction
Yuandi Li, Hui Ji 0004, Fei Yu 0012, Lechao Cheng, Nan Che |
Neural Networks | 4 |
| 2025 | Unsupervised Pre-Training With Language-Vision Prompts for Low-Data Instance SegmentationabstractIn recent times, following the paradigm of DETR (DEtection TRansformer), query-based end-to-end instance segmentation (QEIS) methods have exhibited superior performance compared to CNN-based models, particularly when trained on large-scale datasets. Nevertheless, the effectiveness of these QEIS methods diminishes significantly when confronted with limited training data. This limitation arises from their reliance on substantial data volumes to effectively train the pivotal queries/kernels that are essential for acquiring localization and shape priors. To address this problem, we propose a novel method for unsupervised pre-training in low-data regimes. Inspired by the recently successful prompting technique, we introduce a new method, Unsupervised Pre-training with Language-Vision Prompts (UPLVP), which improves QEIS models' instance segmentation by bringing language-vision prompts to queries/kernels. Our method consists of three parts: (1) Masks Proposal: Utilizes language-vision models to generate pseudo masks based on unlabeled images. (2) Prompt-Kernel Matching: Converts pseudo masks into prompts and injects the best-matched localization and shape features to their corresponding kernels. (3) Kernel Supervision: Formulates supervision for pre-training at the kernel level to ensure robust learning. With the help of our pre-training method, QEIS models can converge faster and perform better than CNN-based models in low-data regimes. Experimental evaluations conducted on MS COCO, Cityscapes, and CTW1500 datasets indicate that the QEIS models' performance can be significantly improved when pre-trained with our method. Dingwen Zhang, Hao Li 0075, Diqi He, Nian Liu 0002, Lechao Cheng, Jingdong Wang 0001, Junwei Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Weakly Supervised Semantic Segmentation via Alternate Self-Dual TeachingabstractWeakly supervised semantic segmentation (WSSS) is a challenging yet important research field in vision community. In WSSS, the key problem is to generate high-quality pseudo segmentation masks (PSMs). Existing approaches mainly depend on the discriminative object part to generate PSMs, which would inevitably miss object parts or involve surrounding image background, as the learning process is unaware of the full object structure. In fact, both the discriminative object part and the full object structure are critical for deriving of high-quality PSMs. To fully explore these two information cues, we build a novel end-to-end learning framework, alternate self-dual teaching (ASDT), based on a dual-teacher single-student network architecture. The information interaction among different network branches is formulated in the form of knowledge distillation (KD). Unlike the conventional KD, the knowledge of the two teacher models would inevitably be noisy under weak supervision. Inspired by the Pulse Width (PW) modulation, we introduce a PW wave-like selection signal to alleviate the influence of the imperfect knowledge from either teacher model on the KD process. Comprehensive experiments on the PASCAL VOC 2012 and COCO-Stuff 10K demonstrate the effectiveness of the proposed ASDT framework, and new state-of-the-art results are achieved. Dingwen Zhang, Hao Li 0075, Wenyuan Zeng, Chaowei Fang, Lechao Cheng, Ming-Ming Cheng, Junwei Han 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | ODMixer: Fine-Grained Spatial-Temporal MLP for Metro Origin-Destination PredictionabstractMetro Origin-Destination (OD) prediction is a crucial yet challenging spatial-temporal prediction task in urban computing, which aims to accurately forecast cross-station ridership for optimizing metro scheduling and enhancing overall transport efficiency. Analyzing fine-grained and comprehensive relations among stations effectively is imperative for metro OD prediction. However, existing metro OD models either mix information from multiple OD pairs from the station's perspective or exclusively focus on a subset of OD pairs. These approaches may overlook fine-grained relations among OD pairs, leading to difficulties in predicting potential anomalous conditions. To address these challenges, we learn traffic evolution from the perspective of all OD pairs and propose a fine-grained spatialtemporal MLP architecture for metro OD prediction, namely ODMixer. Specifically, our ODMixer has double-branch structure and involves the Channel Mixer, the Multi-view Mixer, and the Bidirectional Trend Learner. The Channel Mixer aims to capture short-term temporal relations among OD pairs, the Multi-view Mixer concentrates on capturing spatial relations from both origin and destination perspectives. To model long-term temporal relations, we introduce the Bidirectional Trend Learner. Extensive experiments on two large-scale metro OD prediction datasets HZMOD and SHMO demonstrate the advantages of our ODMixer. Our code is available at https://github.com/KLatitude/ODMixer Yang Liu 0084, Binglin Chen, Yongsen Zheng, Lechao Cheng, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Mixed Attention and Channel Shift Transformer for Efficient Action RecognitionabstractThe practical use of the Transformer-based methods for processing videos is constrained by the high computing complexity. Although previous approaches adopt the spatiotemporal decomposition of 3D attention to mitigate the issue, they suffer from the drawback of neglecting the majority of visual tokens. This article presents a novel mixed attention operation that subtly fuses the random, spatial, and temporal attention mechanisms. The proposed random attention stochastically samples video tokens in a simple yet effective way, complementing other attention methods. Furthermore, since the attention operation concentrates on learning long-distance relationships, we employ the channel shift operation to encode short-term temporal characteristics. Our model can provide more comprehensive motion representations thanks to the amalgamation of these techniques. Experimental results show that the proposed method produces competitive action recognition results with low computational overhead on both large-scale and small-scale public video datasets. Xiusheng Lu, Yanbin Hao, Lechao Cheng, Sicheng Zhao, Yutao Liu 0002, Mingli Song |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | ViT-Calibrator: Decision Stream Calibration for Vision TransformerabstractA surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing internal model architecture designs that often entail significant trial and error with high burdens. In this work, we propose a new paradigm dubbed Decision Stream Calibration that boosts the performance of general Vision Transformers. To achieve this, we shed light on the information propagation mechanism in the learning procedure by exploring the correlation between different tokens and the relevance coefficient of multiple dimensions. Upon further analysis, it was discovered that 1) the final decision is associated with tokens of foreground targets, while token features of foreground target will be transmitted into the next layer as much as possible, and the useless token features of background area will be eliminated gradually in the forward propagation. 2) Each category is solely associated with specific sparse dimensions in the tokens. Based on the discoveries mentioned above, we designed a two-stage calibration scheme, namely ViT-Calibrator, including token propagation calibration stage and dimension propagation calibration stage. Extensive experiments on commonly used datasets show that the proposed approach can achieve promising results. Zhijie Jia, Lechao Cheng, Yang Gao 0001, Jie Lei 0002, Yijun Bei, Zunlei Feng |
AAAI | 3 |
| 2024 | Progressive Feature Self-Reinforcement for Weakly Supervised Semantic SegmentationabstractCompared to conventional semantic segmentation with pixel-level supervision, weakly supervised semantic segmentation (WSSS) with image-level labels poses the challenge that it commonly focuses on the most discriminative regions, resulting in a disparity between weakly and fully supervision scenarios. A typical manifestation is the diminished precision on object boundaries, leading to deteriorated accuracy of WSSS. To alleviate this issue, we propose to adaptively partition the image content into certain regions (e.g., confident foreground and background) and uncertain regions (e.g., object boundaries and misclassified categories) for separate processing. For uncertain cues, we propose an adaptive masking strategy and seek to recover the local information with self-distilled knowledge. We further assume that confident regions should be robust enough to preserve the global semantics, and introduce a complementary self-distillation method that constrains semantic consistency between confident regions and an augmented view with the same class labels. Extensive experiments conducted on PASCAL VOC 2012 and MS COCO 2014 demonstrate that our proposed single-stage approach for WSSS not only outperforms state-of-the-art counterparts but also surpasses multi-stage methods that trade complexity for accuracy. Jingxuan He 0001, Lechao Cheng, Chaowei Fang, Zunlei Feng, Tingting Mu, Mingli Song |
AAAI | 2 |
| 2024 | Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningabstractCamera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths for every pixel without prioritizing objects, which are precisely the entities requiring detection in the 3D space. To this end, we propose IA-BEV, which integrates image-plane instance awareness into the depth estimation process within a BEV-based detector. First, a category-specific structural priors mining approach is proposed for enhancing the efficacy of monocular depth generation. Besides, a self-boosting learning strategy is further proposed to encourage the model to place more emphasis on challenging objects in computation-expensive temporal stereo matching. Together they provide advanced depth estimation results for high-quality BEV features construction, benefiting the ultimate 3D detection. The proposed method achieves state-of-the-art performances on the challenging nuScenes benchmark, and extensive experimental results demonstrate the effectiveness of our designs. Zequn Jie, Shaoxiang Chen 0001, Lechao Cheng, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
AAAI | 4 |
| 2024 | Open-Vocabulary Video Relation Extraction
Wentao Tian, Zheng Wang 0059, Yuqian Fu, Jingjing Chen 0001, Lechao Cheng |
AAAI | 5 |
| 2024 | GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene UnderstandingabstractApplying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task, i.e., the “label rendering” task, to build semantic NeRFs. However, by rendering semantic/instance labels per pixel without considering the contextual information of the rendered image, these methods usually suffer from unclear boundary segmentation and abnormal segmentation of pixels within an object. To solve this problem, we propose Generalized Perception NeRF (GP-NeRF), a novel pipeline that makes the widely used segmentation model and NeRF work compatibly under a unified framework, for facilitating context-aware 3D scene perception. To accomplish this goal, we introduce transformers to aggregate radiance as well as semantic embedding fields jointly for novel views and facilitate the joint volumetric rendering of both fields. In addition, we propose two self-distillation mechanisms, i.e., the Semantic Distill Loss and the Depth-Guided Semantic Distill Loss, to enhance the discrimination and quality of the semantic field and the maintenance of geometric consistency. In evaluation, as shown in Fig. 1 we conduct experimental comparisons under two perception tasks (i.e. semantic and instance segmentation) using both synthetic and real-world datasets. Notably, our method outperforms SOTA approaches by 6.94%,11.76%, and 8.47% on generalized semantic segmentation, finetuning semantic segmentation, and instance segmentation, respectively. Project. Hao Li 0075, Dingwen Zhang, Yalun Dai, Nian Liu 0002, Lechao Cheng, Jingfeng Li, Jingdong Wang 0001, Junwei Han 0001 |
CVPR | 5 |
| 2024 | 3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing
Haoran Li 0020, Haolin Shi, Yanbin Hao, Yong Liao 0003, Lechao Cheng, Peng Yuan Zhou |
ECCV (62) | 6 |
| 2024 | MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
Lechao Cheng, Weikai Chen 0001, Liang Lin 0004, Guanbin Li |
ECCV (17) | 2 |
| 2024 | Improving Knowledge Distillation via Regularizing Feature Direction and Norm
Lechao Cheng, Manni Duan, Yongheng Wang, Zunlei Feng, Shu Kong |
ECCV (24) | 2 |
| 2024 | Revisiting the Power of Prompt for Visual TuningabstractVisual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks. However, VPT and its variants often encounter challenges like prompt initialization, prompt length, and subpar performance in self-supervised pretraining, hindering successful contextual adaptation. This study commences by exploring the correlation evolvement between prompts and patch tokens during proficient training. Inspired by the observation that the prompt tokens tend to share high mutual information with patch tokens, we propose initializing prompts with downstream token prototypes. The strategic initialization, a stand-in for the previous initialization, substantially improves performance. To refine further, we optimize token construction with a streamlined pipeline that maintains excellent performance with almost no increase in computational expenses compared to VPT. Exhaustive experiments show our proposed approach outperforms existing methods by a remarkable margin. For instance, after MAE pre-training, our method improves accuracy by up to 10%$\sim$30% compared to VPT, and outperforms Full fine-tuning 19 out of 24 cases while using less than 0.4% of learnable parameters. Besides, the experimental results demonstrate the proposed SPT is robust to prompt lengths and scales well with model capacity and training data size. We finally provide an insightful exploration into the amount of target data facilitating the adaptation of pre-trained models to downstream tasks. The code is available at https://github.com/WangYZ1608/Self-Prompt-Tuning. Lechao Cheng, Chaowei Fang, Dingwen Zhang, Manni Duan, Meng Wang 0001 |
ICML | 2 |
| 2024 | LoopGaussian: Creating 3D Cinemagraph with Multi-view Images via Eulerian Motion FieldabstractCinemagraph creates captivating video experience by combining elements of still photography and subtle motion. However, most existing cinemagraph video generation lacks depth information, being restricted within 2-dimensional (2D) image space. We advance cinemagraph from 2D image space to 3-dimensional (3D) space with high quality by proposing LoopGaussian. It is based on 3D Gaussian modeling, taking advantage of the 3D Gaussian Splatting (3D-GS) technique that has significantly improved the field of novel view synthesis. Here is a brief overview of our new approach: It employs 3D-GS to reconstruct 3D Gaussian point clouds from multi-view images of static scenes, where shape regularization is used to prevent blurring or artifacts caused by object deformation. To maintain local continuity between scenes, it then clusters the 3D Gaussian points by the proposed SuperGaussian algorithm using features acquired by an autoencoder tailored for 3D Gaussian. Similarities between clusters are used to derive an Eulerian motion field for describing velocities across the entire scene. The estimated Eulerian motion field drives the movement of the 3D Gaussian points, based on which a 3D Cinemagraph is generated through bidirectional animation. The resulting 3D Cinemagraph exhibits natural and seamlessly loopable dynamics. Experiment results validate the effectiveness of the proposed approach, demonstrating high-quality and visually appealing video generation. Jiyang Li, Lechao Cheng, Zhangye Wang, Tingting Mu, Jingxuan He 0001 |
ACM Multimedia | 2 |
| 2024 | Fire and Smoke Detection with Burning Intensity Representation
Xiaoyi Han, Yanfei Wu, Nan Pu, Zunlei Feng, Qifei Zhang 0001, Yijun Bei, Lechao Cheng |
MMAsia | 7 |
| 2024 | Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution DetectionabstractUnsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for developing a safe real-world machine learning system. Current reconstruction-based method provides a good alternative approach, by measuring the reconstruction error between the input and its corresponding generative counterpart in the pixel/feature space. However, such generative methods face the key dilemma, $i.e.$, improving the reconstruction power of the generative model, while keeping compact representation of the ID data. To address this issue, we propose the diffusion-based layer-wise semantic reconstruction approach for unsupervised OOD detection. The innovation of our approach is that we leverage the diffusion model's intrinsic data reconstruction ability to distinguish ID samples from OOD samples in the latent feature space. Moreover, to set up a comprehensive and discriminative feature representation, we devise a multi-layer semantic feature extraction strategy. Through distorting the extracted features with Gaussian noises and applying the diffusion model for feature reconstruction, the separation of ID and OOD samples is implemented according to the reconstruction errors. Extensive experimental results on multiple benchmarks built upon various datasets demonstrate that our method achieves state-of-the-art performance in terms of detection accuracy and speed. Ying Yang 0020, De Cheng, Chaowei Fang, Yubiao Wang, Changzhe Jiao, Lechao Cheng, Nannan Wang 0001, Xinbo Gao 0001 |
NeurIPS | 6 |
| 2024 | Behavior Capture Based Explainable Engagement Recognition
Yijun Bei, Songyuan Guo, Kewei Gao, Zunlei Feng, Yining Tong, Weimin Cai, Lechao Cheng |
PRCV (10) | 7 |
| 2024 | Benchmarking Multi-Scene Fire and Smoke Detection
Xiaoyi Han, Nan Pu, Zunlei Feng, Yijun Bei, Qifei Zhang 0001, Lechao Cheng |
PRCV (11) | 6 |
| 2024 | JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation
Dengqing Yang, Zhenhua Tang 0001, Jinmeng Wu, Shuo Wang 0008, Lechao Cheng, Yanbin Hao |
PRCV (6) | 5 |
| 2024 | Masked Collaborative Contrast for Weakly Supervised Semantic SegmentationabstractThis study introduces an efficacious approach, Masked Collaborative Contrast (MCC), to highlight semantic regions in weakly supervised semantic segmentation. MCC adroitly draws inspiration from masked image modeling and contrastive learning to devise a novel framework that induces keys to contract toward semantic regions. Unlike prevalent techniques that directly eradicate patch regions in the input image when generating masks, we scrutinize the neighborhood relations of patch tokens by exploring masks considering keys on the affinity matrix. Moreover, we generate positive and negative samples in contrastive learning by utilizing the masked local output and contrasting it with the global output. Elaborate experiments on commonly employed datasets evidences that the proposed MCC mechanism effectively aligns global and local perspectives within the image, attaining impressive performance. The source code is available at https://github.com/fwu11/MCC. Fangwen Wu, Jingxuan He 0001, Yufei Yin, Yanbin Hao, Gang Huang 0004, Lechao Cheng |
WACV | 6 |
| 2024 | Multimodality-guided Visual-Caption Semantic Enhancement
Nan Che, Fei Yu 0012, Lechao Cheng, Yuxuan Wang 0001, Chenrui Liu |
Comput. Vis. Image Underst. | 4 |
| 2024 | Mixed Resolution Network with hierarchical motion modeling for efficient action recognition
Xiusheng Lu, Sicheng Zhao, Lechao Cheng, Ying Zheng 0009, Xueqiao Fan, Mingli Song |
Knowl. Based Syst. | 3 |
| 2024 | Life regression based patch slimming for vision transformers
Tianqi Shi, Lechao Cheng, Zunlei Feng, Mingli Song |
Neural Networks | 5 |
| 2024 | Efficient Unsupervised Video Hashing With Contextual Modeling and Structural ControllingabstractThe most important effect of the video hashing technique is to support fast retrieval, which is benefiting from the high efficiency of binary calculation. Current video hash approaches are thus mainly targeted at learning compact binary codes to represent video content accurately. However, they may overlook the generation efficiency for hash codes, i.e., designing lightweight neural networks. This paper proposes anEfficientUnsupervisedVideoHashing (EUVH)method, which is not only for computing compact hash codes but also for designing a lightweight deep model. Specifically, we present an MLP-based model, where the video tensor is split into several groups and multiple axial contexts are explored to separately refine them in parallel. The axial contexts are referred to as the dynamics aggregated from different axial scales, including long/middle/short-range dependencies. The group operation significantly reduces the computational cost of the MLP backbone. Moreover, to achieve compact video hash codes, three structural losses are utilized. As demonstrated by the experiment, the three structures are highly complementary for approximating the real data structure. We conduct extensive experiments on three benchmark datasets for the unsupervised video hashing task and show the superior trade-off between performance and computational cost of our EUVH to the state of the arts. Jingru Duan, Yanbin Hao, Bin Zhu 0006, Lechao Cheng, Peng Yuan Zhou, Xiang Wang 0010 |
IEEE Trans. Multim. | 4 |
| 2024 | Separating Noisy Samples From Tail Classes for Long-Tailed Image Classification With Label NoiseabstractMost existing methods that cope with noisy labels usually assume that the classwise data distributions are well balanced. They are difficult to deal with the practical scenarios where training samples have imbalanced distributions, since they are not able to differentiate noisy samples from tail classes' clean samples. This article makes an early effort to tackle the image classification task in which the provided labels are noisy and have a long-tailed distribution. To deal with this problem, we propose a new learning paradigm which can screen out noisy samples by matching between inferences on weak and strong data augmentations. A leave-noise-out regularization (LNOR) is further introduced to eliminate the effect of the recognized noisy samples. Besides, we propose a prediction penalty based on the online classwise confidence levels to avoid the bias toward easy classes which are dominated by head classes. Extensive experiments on five datasets including CIFAR-10, CIFAR-100, MNIST, FashionMNIST, and Clothing1M demonstrate that the proposed method outperforms the existing algorithms for learning with long-tailed distribution and label noise. Chaowei Fang, Lechao Cheng, Yining Mao, Dingwen Zhang, Yixiang Fang, Guanbin Li, Huiyan Qi, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Progressive Adapting and Pruning: Domain-Incremental Learning for Saliency PredictionabstractSaliency prediction (SAP) plays a crucial role in simulating the visual perception function of human beings. In practical situations, humans can quickly grasp saliency extraction in new image domains. However, current SAP methods mainly concentrate on training models in single domains, which do not effectively handle diverse content and styles present in real-world images. As a result, it would be of great significance if SAP models could efficiently adjust to new image domains. To this end, this article aims to design SAP models that can imitate the incremental learning ability of human beings on multiple image domains and name domain-incremental saliency prediction (DISAP). To make a tradeoff between preventing the forgetting of historical domains and achieving high performance on new domains, we propose a progressively updated domain incremental encoder. This encoder consists of a domain-sharing branch and a domain-specific branch. The domain-sharing branch includes a feature selection mechanism to preserve crucial parameters after fine-tuning the model on each current domain. The remaining parameters are reserved to absorb knowledge from future domains. Furthermore, to capture the unique characteristics of each domain with relatively low computational overhead, we introduce a lightweight design to construct the domain-specific branch, enabling effective adaptation to new domains. Extensive experiments are conducted on multiple domain-incremental learning settings formed by four saliency prediction datasets, including Salicon, MIT1003, the art subset of CAT2000, and WebSal. The results demonstrate that our method outperforms existing methods significantly. The code is available at https://github.com/KaIi-github/DIL4SAP . Kaihui Yang, Junwei Han 0001, Guangyu Guo 0001, Chaowei Fang, Yingzi Fan, Lechao Cheng, Dingwen Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | ScrollTimes: Tracing the Provenance of Paintings as a Window Into HistoryabstractThe study of cultural artifact provenance, tracing ownership and preservation, holds significant importance in archaeology and art history. Modern technology has advanced this field, yet challenges persist, including recognizing evidence from diverse sources, integrating sociocultural context, and enhancing interactive automation for comprehensive provenance analysis. In collaboration with art historians, we examined the handscroll, a traditional Chinese painting form that provides a rich source of historical data and a unique opportunity to explore history through cultural artifacts. We present a three-tiered methodology encompassing artifact, contextual, and provenance levels, designed to create a "Biography" for handscroll. Our approach incorporates the application of image processing techniques and language models to extract, validate, and augment elements within handscroll using various cultural heritage databases. To facilitate efficient analysis of non-contiguous extracted elements, we have developed a distinctive layout. Additionally, we introduce ScrollTimes, a visual analysis system tailored to support the three-tiered analysis of handscroll, allowing art historians to interactively create biographies tailored to their interests. Validated through case studies and expert interviews, our approach offers a window into history, fostering a holistic understanding of handscroll provenance and historical significance. Wei Zhang 0219, Kamkwai Wong, Yitian Chen 0004, Ailing Jia, Luwei Wang, Jianwei Zhang 0015, Lechao Cheng, Huamin Qu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Adapting Object Size Variance and Class Imbalance for Semi-supervised Object DetectionabstractSemi-supervised object detection (SSOD) attracts extensive research interest due to its great significance in reducing the data annotation effort. Collecting high-quality and category-balanced pseudo labels for unlabeled images is critical to addressing the SSOD problem. However, most of the existing pseudo-labeling-based methods depend on a large and fixed threshold to select high-quality pseudo labels from the predictions of a teacher model. Considering different object classes usually have different detection difficulty levels due to scale variance and data distribution imbalance, conventional pseudo-labeling-based methods are arduous to explore the value of unlabeled data sufficiently. To address these issues, we propose an adaptive pseudo labeling strategy, which can assign thresholds to classes with respect to their “hardness”. This is beneficial for ensuring the high quality of easier classes and increasing the quantity of harder classes simultaneously. Besides, label refinement modules are set up based on box jittering for guaranteeing the localization quality of pseudo labels. To further improve the algorithm’s robustness against scale variance and make the most of pseudo labels, we devise a joint feature-level and prediction-level consistency learning pipeline for transferring the information of the teacher model to the student model. Extensive experiments on COCO and VOC datasets indicate that our method achieves state-of-the-art performance. Especially, it brings mean average precision gains of 2.08 and 1.28 on MS-COCO dataset with 5% and 10% labeled images, respectively. Yuxiang Nie, Chaowei Fang, Lechao Cheng, Liang Lin 0004, Guanbin Li |
AAAI | 3 |
| 2023 | De-biased Teacher: Rethinking IoU Matching for Semi-supervised Object DetectionabstractMost of the recent research in semi-supervised object detection follows the pseudo-labeling paradigm evolved from the semi-supervised image classification task. However, the training paradigm of the two-stage object detector inevitably makes the pseudo-label learning process for unlabeled images full of bias. Specifically, the IoU matching scheme used for selecting and labeling candidate boxes is based on the assumption that the matching source~(ground truth) is accurate enough in terms of the number of objects, object position and object category. Obviously, pseudo-labels generated for unlabeled images cannot satisfy such a strong assumption, which makes the produced training proposals extremely unreliable and thus severely spoil the follow-up training. To de-bias the training proposals generated by the pseudo-label-based IoU matching, we propose a general framework -- De-biased Teacher, which abandons both the IoU matching and pseudo labeling processes by directly generating favorable training proposals for consistency regularization between the weak/strong augmented image pairs. Moreover, a distribution-based refinement scheme is designed to eliminate the scattered class predictions of significantly low values for higher efficiency. Extensive experiments demonstrate that the proposed De-biased Teacher consistently outperforms other state-of-the-art methods on the MS-COCO and PASCAL VOC benchmarks. Source codes are available at https://github.com/wkfdb/De-biased-Teracher. Jingyu Zhuang, Guanbin Li, Chaowei Fang, Lechao Cheng, Liang Lin 0004 |
AAAI | 5 |
| 2023 | Boosting Low-Data Instance Segmentation by Unsupervised Pre-training with Saliency PromptabstractInspired by DETR variants, query-based end-to-end instance segmentation (QEIS) methods have recently outperformed CNN-based models on large-scale datasets. Yet they would lose efficacy when only a small amount of training data is available since it's hard for the crucial queries/kernels to learn localization and shape priors. To this end, this work offers a novel unsupervised pre-training solution for low-data regimes. Inspired by the recent success of the Prompting technique, we introduce a new pre-training method that boosts QEIS models by giving Saliency Prompt for queries/kernels. Our method contains three parts: 1) Saliency Masks Proposal is responsible for generating pseudo masks from unlabeled images based on the saliency mechanism. 2) Prompt-Kernel Matching transfers pseudo masks into prompts and injects the corresponding localization and shape priors to the best-matched kernels. 3) Kernel Supervision is applied to supply supervision at the kernel level for robust learning. From a practical perspective, our pre-training method helps QEIS models achieve a similar convergence speed and comparable performance with CNN-based models in low-data regimes. Experimental results show that our method significantly boosts several QEIS models on three datasets.11Code: https://github.com/lifuguan/saliency.prompt Hao Li 0075, Dingwen Zhang, Nian Liu 0002, Lechao Cheng, Yalun Dai, Xinggang Wang, Junwei Han 0001 |
CVPR | 4 |
| 2023 | Generalization Matters: Loss Minima Flattening via Parameter Hybridization for Efficient Online Knowledge DistillationabstractMost existing online knowledge distillation (OKD) techniques typically require sophisticated modules to produce diverse knowledge for improving students' generalization ability. In this paper, we strive to fully utilize multi-model settings instead of well-designed modules to achieve a distillation effect with excellent generalization performance. Generally, model generalization can be reflected in the flatness of the loss landscape. Since averaging parameters of multiple models can find flatter minima, we are inspired to extend the process to the sampled convex combinations of multi-student models in OKD. Specifically, by linearly weighting students' parameters in each training batch, we construct a Hybrid-Weight Model (HWM) to represent the parameters surrounding involved students. The supervision loss of HWM can estimate the landscape's curvature of the whole region around students to measure the generalization explicitly. Hence we integrate HWM's loss into students' training and propose a novel OKD framework via parameter hybridization (OKDPH) to promote flatter minima and obtain robust solutions. Considering the redundancy of parameters could lead to the collapse of HWM, we further introduce a fusion operation to keep the high similarity of students. Compared to the state-of-the-art (SOTA) OKD methods and SOTA methods of seeking flat minima, our OKDPH achieves higher performance with fewer parameters, benefiting OKD with lightweight and robust characteristics. Our code is publicly available at https://github.com/tianlizhang/OKDPH. Tianli Zhang, Mengqi Xue, Haofei Zhang, Yu Wang 0176, Lechao Cheng, Jie Song 0011, Mingli Song |
CVPR | 6 |
| 2023 | Model Doctor for Diagnosing and Treating Segmentation ErrorabstractDespite the remarkable progress in semantic segmentation tasks with the advancement of deep neural networks, existing U-shaped hierarchical typical segmentation networks still suffer from local misclassification of categories and inaccurate target boundaries. In an effort to alleviate this issue, we propose a Model Doctor for semantic segmentation problems. The Model Doctor is designed to diagnose the aforementioned problems in existing pre-trained models and treat them without introducing additional data, with the goal of refining the parameters to achieve better performance. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method. Code is available at https://github.com/zhijiejia/DoctorForSeg. Zhijie Jia, Kaiwen Hu, Lechao Cheng, Zunlei Feng, Mingli Song |
ICIP | 4 |
| 2023 | Team DETR: Guide Queries as a Professional Team in Detection TransformersabstractRecent proposed DETR variants have made tremendous progress in various scenarios due to their streamlined processes and remarkable performance. However, the learned queries usually explore the global context to generate the final set prediction, resulting in redundant burdens and unfaithful results. More specifically, a query is commonly responsible for objects of different scales and positions, which is a challenge for the query itself, and will cause spatial resource competition among queries. To alleviate this issue, we propose Team DETR, which leverages query collaboration and position constraints to embrace objects of interest more precisely. We also dynamically cater to each query member’s prediction preference, offering the query better scale and spatial priors. In addition, the proposed Team DETR is flexible enough to be adapted to other existing DETR variants without increasing parameters and calculations. Extensive experiments on the COCO dataset showcase that Team DETR achieves remarkable gains, especially for small and large objects. Code is available at https://github.com/horrible-dong/TeamDETR. Linyun Zhou, Lechao Cheng, Zunlei Feng, Mingli Song |
ICIP | 4 |
| 2023 | Text-Guided Mask-Free Local Image RetouchingabstractIn the realm of multi-modality, text-guided image retouching techniques emerged with the advent of deep learning. Most currently available text-guided methods, however, rely on object-level supervision to confine the region of interest that may be updated. This not only makes it more challenging to develop these algorithms but also limits how widely deep learning can be used for image retouching. In this paper, we offer a text-guided mask-free image retouching approach that yields consistent results to address this concern. Specifically, we propose a two-stage mask-free training paradigm tailored for text-guided image retouching tasks. In the first stage, an unified mask is proposed according to the query description, and then several candidate images are generated with the provided mask and the conditional description based on diffusion model. Extensive experiments have shown that our method can produce high-quality images based on spoken language. Zerun Liu, Fan Zhang 0051, Jingxuan He 0001, Zhangye Wang, Lechao Cheng |
ICME | 6 |
| 2023 | SASFormer: Transformers for Sparsely Annotated Semantic SegmentationabstractSemantic segmentation based on sparse annotation has advanced in recent years. It labels only part of each object in the image, leaving the remainder unlabeled. Most of the existing approaches are time-consuming and often necessitate a multi-stage training strategy. In this work, we propose a simple yet effective sparse annotated semantic segmentation framework based on segformer, dubbed SASFormer, that achieves remarkable performance. Specifically, the framework first generates hierarchical patch attention maps, which are then multiplied by the network predictions to produce correlated regions separated by valid labels. Besides, we also introduce the affinity loss to ensure consistency between the features of correlation results and network predictions. Extensive experiments showcase that our proposed approach is superior to existing methods and achieves cutting-edge performance. The source code is available at https://github.com/su-hui-zz/SASFormer. Hui Su, Yue Ye, Lechao Cheng, Mingli Song |
ICME | 4 |
| 2023 | Propheter: Prophetic Teacher Guided Long-Tailed Distribution Learning
Yongcheng Jing, Linyun Zhou, Wenqi Huang 0002, Lechao Cheng, Zunlei Feng, Mingli Song |
ICONIP (4) | 5 |
| 2023 | A privacy-aware visual query approach for location-based data
Ziliang Wu, Erqing Zhang, Zhaosong Huang, Mingliang Xu 0001, Lechao Cheng, Minfeng Zhu 0001, Wei Chen 0001 |
Comput. Graph. | 6 |
| 2023 | Disassembling Convolutional Segmentation Network
Kaiwen Hu, Fangyuan Mao, Xinhui Song, Lechao Cheng, Zunlei Feng, Mingli Song |
Int. J. Comput. Vis. | 5 |
| 2023 | KE-RCNN: Unifying Knowledge-Based Reasoning Into Part-Level Attribute ParsingabstractPart-level attribute parsing is a fundamental but challenging task, which requires the region-level visual understanding to provide explainable details of body parts. Most existing approaches address this problem by adding a regional convolutional neural network (RCNN) with an attribute prediction head to a two-stage detector, in which attributes of body parts are identified from localwise part boxes. However, localwise part boxes with limit visual clues (i.e., part appearance only) lead to unsatisfying parsing results, since attributes of body parts are highly dependent on comprehensive relations among them. In this article, we propose a knowledge-embedded RCNN (KE-RCNN) to identify attributes by leveraging rich knowledge, including implicit knowledge (e.g., the attribute “above-the-hip” for a shirt requires visual/geometry relations of shirt-hip) and explicit knowledge (e.g., the part of “shorts” cannot have the attribute of “hoodie” or “lining”). Specifically, the KE-RCNN consists of two novel components, that is: 1) implicit knowledge-based encoder (IK-En) and 2) explicit knowledge-based decoder (EK-De). The former is designed to enhance part-level representation by encoding part–part relational contexts into part boxes, and the latter one is proposed to decode attributes with a guidance of prior knowledge about part–attribute relations. In this way, the KE-RCNN is plug-and-play, which can be integrated into any two-stage detectors, for example, Attribute-RCNN, Cascade-RCNN, HRNet-based RCNN, and SwinTransformer-based RCNN. Extensive experiments conducted on two challenging benchmarks, for example, Fashionpedia and Kinetics-TPS, demonstrate the effectiveness and generalizability of the KE-RCNN. In particular, it achieves higher improvements over all existing methods, reaching around 3% of${\mathrm{ AP}}^{\mathrm{ all}}_{\rm IoU+F_{1}}$on Fashionpedia and around 4% of${\mathrm{ Acc}}_{p}$on Kinetics-TPS. Code and models are publicly available at:https://github.com/sota-joson/KE-RCNN. Xuanhan Wang, Jingkuan Song, Xiaojia Chen, Lechao Cheng, Lianli Gao, Heng Tao Shen |
IEEE Trans. Cybern. | 4 |
| 2023 | Reliable Mutual Distillation for Medical Image Segmentation Under Imperfect AnnotationsabstractConvolutional neural networks (CNNs) have made enormous progress in medical image segmentation. The learning of CNNs is dependent on a large amount of training data with fine annotations. The workload of data labeling can be significantly relieved via collecting imperfect annotations which only match the underlying ground truths coarsely. However, label noises which are systematically introduced by the annotation protocols, severely hinders the learning of CNN-based segmentation models. Hence, we devise a novel collaborative learning framework in which two segmentation models cooperate to combat label noises in coarse annotations. First, the complementary knowledge of two models is explored by making one model clean training data for the other model. Secondly, to further alleviate the negative impact of label noises and make sufficient usage of the training data, the specific reliable knowledge of each model is distilled into the other model with augmentation-based consistency constraints. A reliability-aware sample selection strategy is incorporated for guaranteeing the quality of the distilled knowledge. Moreover, we employ joint data and model augmentations to expand the usage of reliable knowledge. Extensive experiments on two benchmarks showcase the superiority of our proposed method against existing methods under annotations with different noise levels. For example, our approach can improve existing methods by nearly 3% DSC on the lung lesion segmentation dataset LIDC-IDRI under annotations with 80% noise ratio. Code is available at: https://github.com/Amber-Believe/ReliableMutualDistillation. Chaowei Fang, Lechao Cheng, Zhifan Gao, Chengwei Pan, Zhaohui Zheng 0004, Dingwen Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2023 | From External to Internal: Structuring Image for Text-to-Image Attributes ManipulationabstractManipulating visual attributes of an image through a natural language description, known as text-to-image attributes manipulation (T2AM), is a challenging task. However, existing approaches tend to search the whole image to manipulate the target instance indicated by a description, thus they often fail to locate and manipulate the accurate text-relevant regions, and even disturb the text-irrelevant contents, e.g. texture and background. Meanwhile, the model efficiency needs to be improved. To tackle the above issues, we introduce a novel yet simple GAN-based approach, namelyStructuringImage forManipulating(SIMGAN), to narrow down the optimization areas from external to internal. It consists of two major components: 1)External Structuring(ExST), a pretrained segmentation network, for recognizing and separating the target instances and background from an image; and 2)Internal Structuring(InST) for seeking out and editing the text-relevant attributes of the target instances based on the given description and masked hierarchical image representations from ExST. Specifically, the InST structures target instances from outline to detail by firstly drawing the sketch and colors underpainting of instances with anOutline-Oriented Structuring(OuST), and then enhancing the text-relevant attributes and elaborating on details with aDetail-Oriented Structuring(DeST). Extensive experiments on benchmark datasets demonstrate that our framework significantly outperforms state-of-the-art both quantitatively and qualitatively. Compared with the state-of-the-art method ManiGAN, our approach reduces the training time by 88%, while the inferring time is three times faster. In addition, our approach is easily extended to solve the instance-level image-to-image translation problem, and the results exhibit the versatility and effectiveness of our approach. This code is released inhttps://github.com/qikizh/SIMGAN. Lianli Gao, Qike Zhao, Junchen Zhu, Sitong Su, Lechao Cheng, Lei Zhao 0017 |
IEEE Trans. Multim. | 5 |
| 2022 | Re-Attention Transformer for Weakly Supervised Object Localization
Hui Su, Yue Ye, Mingli Song, Lechao Cheng |
BMVC | 5 |
| 2022 | Compound Batch Normalization for Long-tailed Image ClassificationabstractSignificant progress has been made in learning image classification neural networks under long-tail data distribution using robust training algorithms such as data re-sampling, re-weighting, and margin adjustment. Those methods, however, ignore the impact of data imbalance on feature normalization. The dominance of majority classes (head classes) in estimating statistics and affine parameters causes internal covariate shifts within less-frequent categories to be overlooked. To alleviate this challenge, we propose a compound batch normalization method based on a Gaussian mixture. It can model the feature space more comprehensively and reduce the dominance of head classes. In addition, a moving average-based expectation maximization (EM) algorithm is employed to estimate the statistical parameters of multiple Gaussian distributions. However, the EM algorithm is sensitive to initialization and can easily become stuck in local minima where the multiple Gaussian components continue to focus on majority classes. To tackle this issue, we developed a dual-path learning framework that employs class-aware split feature normalization to diversify the estimated Gaussian distributions, allowing the Gaussian components to fit with training samples of less-frequent classes more comprehensively. Extensive experiments on commonly used datasets demonstrated that the proposed method outperforms existing methods on long-tailed image classification. Lechao Cheng, Chaowei Fang, Dingwen Zhang, Guanbin Li, Gang Huang 0004 |
ACM Multimedia | 1 |
| 2022 | Cross-Modality High-Frequency Transformer for MR Image Super-ResolutionabstractImproving the resolution of magnetic resonance (MR) image data is critical to computer-aided diagnosis and brain function analysis. Higher resolution helps to capture more detailed content, but typically induces to lower signal-to-noise ratio and longer scanning time. To this end, MR image super-resolution has become a widely-interested topic in recent times. Existing works establish extensive deep models with the conventional architectures based on convolutional neural networks (CNN). In this work, to further advance this research field, we make an early effort to build a Transformer-based MR image super-resolution framework, with careful designs on exploring valuable domain prior knowledge. Specifically, we consider two-fold domain priors including the high-frequency structure prior and the inter-modality context prior, and establish a novel Transformer architecture, called Cross-modality high-frequency Transformer (Cohf-T), to introduce such priors into super-resolving the low-resolution (LR) MR images. Experiments on two datasets indicate that Cohf-T achieves new state-of-the-art performance. Chaowei Fang, Dingwen Zhang, Liang Wang 0001, Yulun Zhang 0001, Lechao Cheng, Junwei Han 0001 |
ACM Multimedia | 5 |
| 2022 | Mix-DANN and Dynamic-Modal-Distillation for Video Domain AdaptationabstractVideo domain adaptation is non-trivial due to video is inherently involved with multi-dimensional and multi-modal information. Existing works mainly adopt adversarial learning and self-supervised tasks to align features. Nevertheless, the explicit interaction between source and target in the temporal dimension, as well as the adaptation between modalities, are unexploited. In this paper, we propose Mix-Domain-Adversarial Neural Network and Dynamic-Modal-Distillation (MD-DMD), a novel multi-modal adversarial learning framework for unsupervised video domain adaptation. Our approach incorporates the temporal information between source and target domains, as well as the diversity of adaptability between modalities. On the one hand, for every single modality, we mix the frames from source and target domains to form mix-samples, then let the adversarial-discriminator predict the mix ratio of a mix-sample to further enhance the ability of the model to capture domain-invariant feature representations. On the other hand, we dynamically estimate the adaptability for different modalities during training, then pick the most adaptable modality as a teacher to guide other modalities by knowledge distillation. As a result, modalities are capable of learning transferable knowledge from each other, which leads to more effective adaptation. Experiments on two video domain adaptation benchmarks demonstrate the superiority of our proposed MD-DMD over state-of-the-art methods. Yuehao Yin, Bin Zhu 0006, Jingjing Chen 0001, Lechao Cheng, Yu-Gang Jiang 0001 |
ACM Multimedia | 4 |
| 2022 | Long-term Leap Attention, Short-term Periodic Shift for Video ClassificationabstractVideo transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes T times longer sequence than the latter under the current attention of quadratic complexity (T2N2). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy. Hao Zhang 0047, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo |
ACM Multimedia | 2 |
| 2021 | Visual Boundary Knowledge Translation for Foreground SegmentationabstractWhen confronted with objects of unknown types in an image, humans can effortlessly and precisely tell their visual boundaries. This recognition mechanism and underlying generalization capability seem to contrast to state-of-the-art image segmentation networks that rely on large-scale category-aware annotated training samples. In this paper, we make an attempt towards building models that explicitly account for visual boundary knowledge, in hope to reduce the training effort on segmenting unseen categories. Specifically, we investigate a new task termed as Boundary Knowledge Translation (BKT). Given a set of fully labeled categories, BKT aims to translate the visual boundary knowledge learned from the labeled categories, to a set of novel categories, each of which is provided only a few labeled samples. To this end, we propose a Translation Segmentation Network (Trans-Net), which comprises a segmentation network and two boundary discriminators. The segmentation network, combined with a boundary-aware self-supervised mechanism, is devised to conduct foreground segmentation, while the two discriminators work together in an adversarial manner to ensure an accurate segmentation of the novel categories under light supervision. Exhaustive experiments demonstrate that, with only tens of labeled samples as guidance, Trans-Net achieves close results on par with fully supervised methods. Zunlei Feng, Lechao Cheng, Xinchao Wang, Xiang Wang 0010, Ya Jie Liu, Xiangtong Du, Mingli Song |
AAAI | 2 |
| 2021 | Edge-competing Pathological Liver Vessel Segmentation with Limited LabelsabstractThe microvascular invasion (MVI) is a major prognostic factor in hepatocellular carcinoma, which is one of the malignant tumors with the highest mortality rate. The diagnosis of MVI needs discovering the vessels that contain hepatocellular carcinoma cells and counting their number in each vessel, which depends heavily on experiences of the doctor, is largely subjective and time-consuming. However, there is no algorithm as yet tailored for the MVI detection from pathological images. This paper collects the first pathological liver image dataset containing $522$ whole slide images with labels of vessels, MVI, and hepatocellular carcinoma grades. The first and essential step for the automatic diagnosis of MVI is the accurate segmentation of vessels. The unique characteristics of pathological liver images, such as super-large size, multi-scale vessel, and blurred vessel edges, make the accurate vessel segmentation challenging. Based on the collected dataset, we propose an Edge-competing Vessel Segmentation Network (EVS-Net), which contains a segmentation network and two edge segmentation discriminators. The segmentation network, combined with an edge-aware self-supervision mechanism, is devised to conduct vessel segmentation with limited labeled patches. Meanwhile, two discriminators are introduced to distinguish whether the segmented vessel and background contain residual features in an adversarial manner. In the training stage, two discriminators are devised to compete for the predicted position of edges. Exhaustive experiments demonstrate that, with only limited labeled patches, EVS-Net achieves a close performance of fully supervised methods, which provides a convenient tool for the pathological liver vessel segmentation. Code is publicly available at https://github.com/wang97zh/EVS-Net. Zunlei Feng, Xinchao Wang, Xiuming Zhang, Lechao Cheng, Jie Lei 0002, Mingli Song |
AAAI | 5 |
| 2021 | Boundary Knowledge Translation based Reference Semantic SegmentationabstractGiven a reference object of an unknown type in an image, human observers can effortlessly find the objects of the same category in another image and precisely tell their visual boundaries. Such visual cognition capability of humans seems absent from the current research spectrum of computer vision. Existing segmentation networks, for example, rely on a humongous amount of labeled data, which is laborious and costly to collect and annotate; besides, the performance of segmentation networks tend to downgrade as the number of the category increases. In this paper, we introduce a novel Reference semantic segmentation Network (Ref-Net) to conduct visual boundary knowledge translation. Ref-Net contains a Reference Segmentation Module (RSM) and a Boundary Knowledge Translation Module (BKTM). Inspired by the human recognition mechanism, RSM is devised only to segment the same category objects based on the features of the reference objects. BKTM, on the other hand, introduces two boundary discriminator branches to conduct inner and outer boundary segmentation of the target object in an adversarial manner, and translate the annotated boundary knowledge of open-source datasets into the segmentation network. Exhaustive experiments demonstrate that, with tens of finely-grained annotated samples as guidance, Ref-Net achieves results on par with fully supervised methods on six datasets. Our code can be found in the supplementary material. Lechao Cheng, Zunlei Feng, Xinchao Wang, Ya Jie Liu, Jie Lei 0002, Mingli Song |
IJCAI | 1 |
| 2019 | A Synthesis-by-Analysis Network with Applications in Image Super-Resolution
Lechao Cheng, Zhangye Wang |
CGI | 1 |
| 2018 | Intrinsic Image Transformation via Scale Space DecompositionabstractWe introduce a new network structure for decomposing an image into its intrinsic albedo and shading. We treat it as an image-to-image transformation problem and explore the scale space of the input and output. By expanding the output images (albedo and shading) into their Laplacian pyramid components, we develop a multi-channel architecture that learns the image-to-image transformation function in successive frequency bands in parallel, within each channel is a fully convolutional neural network. This network architecture is general and extensible, and has demonstrated excellent performance on the task of intrinsic image decomposition. We evaluate the network on two benchmark datasets: the MPI-Sintel dataset and the MIT Intrinsic Images dataset. Both quantitative and qualitative results show our model delivers a clear progression over state-of-the-art. Lechao Cheng, Zicheng Liao |
CVPR | 1 |
| 2015 | audeosynth: music-driven video montageabstractWe introduce music-driven video montage, a media format that offers a pleasant way to browse or summarize video clips collected from various occasions, including gatherings and adventures. In music-driven video montage, the music drives the composition of the video content. According to musical movement and beats, video clips are organized to form a montage that visually reflects the experiential properties of the music. Nonetheless, it takes enormous manual work and artistic expertise to create it. In this paper, we develop a framework for automatically generating music-driven video montages. The input is a set of video clips and a piece of background music. By analyzing the music and video content, our system extracts carefully designed temporal features from the input, and casts the synthesis problem as an optimization and solves the parameters through Markov Chain Monte Carlo sampling. The output is a video montage whose visual activities are cut and synchronized with the rhythm of the music, rendering a symphony of audio-visual resonance. Zicheng Liao, Yizhou Yu, Bingchen Gong, Lechao Cheng |
ACM Trans. Graph. | 4 |