EDBT 2026 Demo / reviewers in the wild / expert
Ke Cao 0001
dblp:89/2111-1
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0009-0000-4421-0817ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RelaCtrl: Relevance-Guided Efficient Control for Diffusion TransformersabstractThe Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource allocation due to their failure to account for the varying relevance of control information across different transformer layers. To address this, we propose the Relevance-Guided Efficient Controllable Generation framework, RelaCtrl, enabling efficient and resource-optimized integration of control signals into the Diffusion Transformer. First, we evaluate the relevance of each layer in the Diffusion Transformer to the control information by assessing the ControlNet Relevance Score, which measures the impact of skipping each control layer on both the quality of generation and the control effectiveness during inference. Based on the strength of the relevance, we then tailor the positioning, parameter scale, and modeling capacity of the control layers to reduce unnecessary parameters and redundant computations. Additionally, to further improve efficiency, we replace the self-attention and FFN in the commonly used copy block with the carefully designed Two-Dimensional Shuffle Mixer (TDSM), enabling efficient implementation of both the token mixer and channel mixer. Both qualitative and quantitative experimental results demonstrate that our approach achieves superior performance with only 15% of the parameters and computational complexity compared to PixArt-delta. Ke Cao 0001, Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Xuanhua He, Run Ling, Haozhe Wang 0002, Hongjuan Pei, Yihua Shao, Zhanjie Zhang, Jie Zhang 0033 |
AAAI | 1 |
| 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationabstractMulti-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality. Run Ling, Ke Cao 0001, Ao Ma 0005, Runze He, Changwei Wang 0001, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Guibing Guo, Jingjing Lv, Junjie Shen 0008, Ching Law, Xingwei Wang 0001 |
AAAI | 2 |
| 2026 | Self-supervised Multiplex Consensus Mamba for General Image FusionabstractImage fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks. Yingying Wang 0005, Rongjin Zhuang, Hui Zheng 0003, Xuanhua He, Ke Cao 0001, Xiaotong Tu, Xinghao Ding |
AAAI | 5 |
| 2026 | Shuffle Mamba: State Space Models With Random Shuffle for Multi-Modal Image FusionabstractMulti-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are proficient in long-range modeling with linear complexity, most Mamba-based approaches use fixed scanning strategies, which can introduce biased prior information. To mitigate this issue, we propose a novel Bayesian-inspired scanning strategy called Random Shuffle, supplemented by a theoretically feasible inverse shuffle to maintain information coordination invariance, aiming to eliminate biases associated with fixed sequence scanning. Based on this transformation pair, we customized the Shuffle Mamba Framework, penetrating modality-aware information representation and cross-modality information interaction across spatial and channel axes to ensure robust interaction and an unbiased global receptive field for multi-modal image fusion. Furthermore, we develop a testing methodology based on Monte-Carlo averaging to ensure the model’s output aligns more closely with expected results. Extensive experiments across multiple multi-modal image fusion tasks demonstrate the effectiveness of our proposed method, yielding excellent fusion quality compared to state-of-the-art alternatives. The code is available at https://github.com/caoke-963/Shuffle-Mamba. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Chengjun Xie, Man Zhou 0003, Jie Zhang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Distilling Textual Priors From LLM to Efficient Image FusionabstractMulti-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF. Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma 0005, Jiasong Feng, Ke Cao 0001, Jing Wang 0021, Yun Wang 0053, Quanwei Zhang, Zhanjie Zhang |
ICCV | 3 |
| 2025 | FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual GuidanceabstractSynthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations without frame-specific textual guidance. Thus, the model's capacity to comprehend the temporal logic conveyed in prompts and generate videos with coherent motion is restricted. To tackle this limitation, we introduce FancyVideo, an innovative video generator that improves the existing text-control mechanism with the well-designed Cross-frame Textual Guidance Module (CTGM). Specifically, CTGM incorporates the Temporal Information Injector (TII) and Temporal Affinity Refiner (TAR) at the beginning and end of cross-attention, respectively, to achieve frame-specific textual guidance. Firstly, TII injects frame-specific information from latent features into text conditions, thereby obtaining cross-frame textual conditions. Then, TAR refines the correlation matrix between cross-frame textual conditions and latent features along the time dimension. Extensive experiments comprising both quantitative and qualitative evaluations demonstrate the effectiveness of FancyVideo. Our approach achieves state-of-the-art T2V generation results on the EvalCrafter benchmark and facilitates the synthesis of dynamic and consistent videos. Note that the T2V process of FancyVideo essentially involves a text-to-image step followed by T+I2V. This means it also supports the generation of videos from user images, i.e., the image-to-video (I2V) task. A significant number of experiments have shown that its performance is also outstanding. Jiasong Feng, Ao Ma 0005, Jing Wang 0021, Ke Cao 0001, Zhanjie Zhang |
IJCAI | 4 |
| 2025 | WISA: World simulator assistant for physics-aware text-to-video generationabstractRecent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey physical laws. This limitation stems primarily from the lack of explicit physical guidance, caused by a significant gap between high-level physical concepts and the generative capabilities of current models. To address this challenge, we propose the **W**orld **S**imulator **A**ssistant (**WISA**), a novel framework designed to systematically decompose and integrate physical principles into T2V models. Specifically, WISA decomposes physical knowledge into three hierarchical levels: textual physical descriptions, qualitative physical categories, and quantitative physical properties. It then incorporates several carefully designed modules—such as Mixture-of-Physical-Experts Attention (MoPA) and a Physical Classifier—to effectively encode these attributes and enhance the model’s adherence to physical laws during generation. In addition, most existing video datasets feature only weak or implicit representations of physical phenomena, limiting their utility for learning explicit physical principles. To bridge this gap, we present **WISA-80K**, a new dataset comprising 80,000 human-curated videos that depict 17 fundamental physical laws across three core domains of physics: dynamics, thermodynamics, and optics. Experimental results show that WISA substantially improves the alignment of T2V models (such as CogVideoX and Wan2.1) with real-world physical laws, achieving notable gains on the VideoPhy benchmark. Our data, code, and models are available in the [Project Page](https://wisav1.github.io/WISA/). Jing Wang 0021, Ao Ma 0005, Ke Cao 0001, Jiasong Feng, Zhanjie Zhang, Wanyuan Pang, Xiaodan Liang |
NeurIPS | 3 |
| 2025 | Frequency Decoupled Domain-Irrelevant Feature Learning for Pan-SharpeningabstractPan-sharpening aims to generate high-detail multi-spectral images (HRMS) through the fusion of panchromatic (PAN) and multi-spectral (MS) images. However, existing pan-sharpening methods often suffer from significant performance degradation when dealing with out-of-distribution data, as they assume the training and test datasets are independent and identically distributed. To overcome this challenge, we propose a novel frequency domain-irrelevant feature learning framework that exhibits exceptional generalization capabilities. Our approach involves parallel extraction and processing of domain-irrelevant information from the amplitude and phase components of the input images. Specifically, we design a frequency information separation module to extract the amplitude and phase components of the paired images. The learnable high-pass filter is then employed to eliminate domain-specific information from the amplitude spectrums. After that, we devised two specialized sub-networks (AFL-Net and PFL-Net) to perform targeted learning of the frequency domain-irrelevant information. This allows our method to effectively capture the complementary domain-irrelevant information contained in the amplitude and phase spectra of the images. Finally, the information fusion and restoration module dynamically adjusts the feature channel weights, enabling the network to output high-quality HRMS images. Through this frequency domain-irrelevant feature learning framework, our method balances generalization capability and network performance on the distribution of training dataset. Extensive experiments conducted on various satellite datasets demonstrate the effectiveness of our method for generalized pan-sharpening. Our proposed network outperforms state-of-the-art methods in terms of both quantitative metrics and visual quality, showcasing its superior ability to handle diverse, out-of-distribution data. Jie Zhang 0033, Ke Cao 0001, Yunlong Lin, Xuanhua He, Yingying Wang 0005, Rui Li 0027, Chengjun Xie, Jun Zhang 0034, Man Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | PanDiT: A Few-Step Diffusion Transformer for High-Fidelity and Efficient PansharpeningabstractPansharpening plays a crucial role in remote sensing by fusing low-resolution multispectral (LRMS) images and high-resolution panchromatic (PAN) images to generate high-resolution multispectral (HRMS) images. While denoising diffusion models offer potential for high-fidelity image generation, their practical application in pansharpening has been severely hindered by huge computational costs from iterative sampling and naive conditioning strategies that struggle to fuse multi-modal information effectively. In this paper, we introduce PanDiT, a novel Diffusion Transformer framework designed to address these challenges. PanDiT is built on the core principle of a Decoupled Conditioning Mechanism, which explicitly disentangles and injects spatial and spectral guidance, and is engineered for practical, few-step inference. Our framework leverages a powerful Diffusion Transformer (DiT) backbone, where conditioning is achieved through two specialized injection blocks that capture spatial and time-frequency features. Crucially, by integrating an implicit sampling strategy, we accelerate the inference process to as few as two steps. Extensive experiments on multiple benchmark datasets demonstrate that PanDiT not only establishes a new state-of-the-art in fusion quality and achieves a good quality-efficiency trade-off. Code is available at https://github.com/para133/PanDIT. Jiabin Fang, Ke Cao 0001, Xuanhua He, Jie Zhang 0033, Man Zhou 0003, Liu Liu 0012 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Pan-Sharpening via Causal-Aware Feature Distribution CalibrationabstractIn this work, we reveal an interesting observation within the multi-spectral modality: high-frequency components exhibit a long-tailed distribution, in contrast to the Gaussian distribution of dominant low-frequency components. This dual-distribution characteristic presents a challenge for network optimization, leading to overfitting on low-frequency information while neglecting essential high-frequency details. Addressing this issue from a causal inference perspective, we identify optimizer momentum as a confounding factor that biases models towards focusing on the head part of the high-frequency distribution during training. To counteract this effect, we propose a novel optimization strategy and supplement the global-modeling network architecture to balance the frequencies learning. In the training stage, we employ the Recurrent Weighted Key-Value (RWKV) architecture, which features a global receptive field, to effectively learn the long-tailed distribution of high-frequency components and quantify the cumulative direction of feature bias. During the testing stage, we apply counterfactual reasoning to adjust feature distributions based on the quantified bias. To our knowledge, this is the first time to investigate the imbalance of frequency learning within pan-sharpening from the causal inference perspective. Extensive experiments on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, showcasing its effectiveness and robustness in pan-sharpening tasks. Xueheng Li, Tao Hu 0027, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Exploring Text-Guided Information Fusion Through Chain-of-Reasoning for PansharpeningabstractPan-sharpening aims to enhance the spatial resolution of low-resolution multispectral (LRMS) images by integrating high-frequency information from a corresponding texture-rich panchromatic (PAN) image, while maintaining the spectral integrity of the LRMS image. Although text-guided multi-modal learning has made considerable strides in the natural image domain, its potential to pan-sharpening remains underexplored, primarily due to the limited availability of multi-modal remote sensing datasets. To this end, we construct an entirely new pan-sharpening framework by making efforts from three key aspects: (1) text-equipped multi-modal data collection through chain-of-reasoning, (2) large model prior-driven multi-modal information fusion, and (3) visual information interaction through prompt engineering, leveraging textual information to guide the pan-sharpening process within a multi-modal fusion framework. We initially utilize the generic large language model priors to generate descriptive captions for MS images, forming a multi-modal pan-sharpening dataset. By integrating super-resolved imagery and segmentation maps generated by segment anything, we apply Chain-of-Thought (CoT) prompting to generate spatially focused captions across diverse satellite datasets. These captions enhance visual features and provide high-level contextual information, improving semantic understanding for pan-sharpening. Building on the aforementioned multi-modal data, we tailor two text-guided information fusion modules: Textual Enhancement Block (TEB) standing on large model prior and Textual Modulated Block (TMB) utilizing text information to effectively guide and refine the pan-sharpening fusion process. Extensive experiments on multiple satellite datasets demonstrate that our proposed framework outperforms state-of-the-art methods, highlighting its effectiveness and superior performance in pan-sharpening. Xueheng Li, Xuanhua He, Ke Cao 0001, Jie Zhang 0033, Chengjun Xie, Man Zhou 0003, Danfeng Hong, Bo Huang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Frequency Decomposition-Driven Network for JPEG Artifacts RemovalabstractJPEG compression, a widely adopted image format, often introduces visual artifacts and quality degradation in image quality. Removal of these JPEG artifacts, especially under high compression rates, proves challenging and typically results in overly smoothed images. This issue primarily arises due to the prevalence of low-frequency regions in natural images. This distribution leads models towards capturing low-frequency information and generating over-smoothing results. To address this issue, we propose the Frequency Decomposition-Driven Network (FDDNet) for JPEG artifact removal. FDDNet incorporates three core modules: the Decomposition Module (DM), inspired by wavelet lifting schemes, extracts both low-frequency and high-frequency components by considering feature channel relationships. The lightweight Low-Frequency Restoration Module and the High-Frequency Refinement Module, are each adept at handling distinct frequency components effectively. By emphasizing high-frequency components, our method surpasses existing approaches in terms of both quantitative metrics and visual quality across various datasets. Ke Cao 0001, Xuanhua He, Tao Hu 0027, Rui Li 0027, Chengjun Xie, Jie Zhang 0033 |
ICME | 1 |
| 2024 | Spatially-Adaptive Large-Kernel Network for Efficient Image Super-ResolutionabstractIn the realm of image super-resolution (SR), efficiency on low-power devices remains a significant challenge due to the high computational demands of current methods. In response to this issue, we propose a novel Spatially-Adaptive Large-Kernel Network(SALKN) tailored for efficient image super-resolution. Inspired by the spectrum convolution theorem, we introduce a spatially-adaptive large kernel convolution unit integrated into a vision transformer architecture. This approach implements dynamic large kernel convolution through an input-adaptive frequency domain multiplication and multi-head mechanism, significantly reducing computational overhead. Our implementation approach, which involves dynamically generating spatially-adaptive frequency filters, enables the dynamic large kernel convolution to be performed with reduced computational costs, facilitating the realization of a global receptive field and promoting scale diversity within the features. Extensive experiments validate that SALKN outperforms existing efficient super-resolution methods with reduced complexity, achieving state-of-the-art performance. Xuanhua He, Ke Cao 0001, Tao Hu 0027, Jie Zhang 0033, Rui Li 0027 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Pan-Sharpening With Wavelet-Enhanced High-Frequency InformationabstractPan-sharpening is essentially a panchromatic (PAN)-guided super-resolution process, primarily focused on enhancing multi-spectral image quality. This methodology intricately incorporates the high-frequency derived from texture-rich PAN images into the lower-resolution multi-spectral (LRMS) counterparts. However, current spatial domain techniques frequently face challenges in accurately restoring texture details, while frequency domain methods lack efficient interaction with spatial domains, thus restricting the overall model performance. In response to these challenges, we introduce a novel High-frequency Wavelet Network that capitalizes on the spatial-frequency interaction and frequency division capabilities inherent in wavelet transform. In particular, our approach consists of two fundamental modules: the Wavelet-Inspired Fusion Block and the High-Frequency Enhancement Block. The former is inspired by wavelet lifting schemes, enabling the fusion of frequencies and facilitating information exchange across various subbands. The latter harnesses wavelet’s frequency division attributes to enhance high-frequency information learning. Comprehensive experiments over multiple satellite datasets demonstrate that our approach outperforms state-of-the-art techniques in both quantitative and qualitative assessments. Moreover, our model showcases exceptional generalization capabilities in real-world scenarios. Code is available at https://github.com/alexhe101/WINet. Jie Zhang 0033, Xuanhua He, Ke Cao 0001, Rui Li 0027, Chengjun Xie, Man Zhou 0003, Danfeng Hong |
IEEE Trans. Geosci. Remote. Sens. | 4 |