Kecheng Chen

dblp:244/8268 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios
abstract
Large language models (LLMs) exhibit substantial variability in performance and computational cost across tasks and queries, motivating routing systems that select models to meet user-specific cost--performance trade-offs. However, existing routers generalize poorly in cold-start scenarios where in-domain training data is unavailable. We address this limitation with a multi-level task-profile--guided data synthesis framework that constructs a hierarchical task taxonomy and produces diverse question--answer pairs to approximate the test-time query distribution. Building on this, we introduce TRouter, a task-type--aware router approach that models query-conditioned cost and performance via latent task-type variables, with prior regularization derived from the synthesized task taxonomy. This design enhances TRouter's routing utility under both cold-start and in-domain settings. Across multiple benchmarks, we show that our synthesis framework alleviates cold-start issues and that TRouter delivers effective LLM routing. ©2026 Association for Computational Linguistics
Hui Liu 0036, Kecheng Chen, Jie Liu 0044, Wenya Wang 0001, Haoliang Li
ACL (1)3
2026 When Video Compression Meets Multimodal Large Language Models: A Unified Paradigm for Cross-Modality Video Compression
abstract
Traditional video compression methods perform well at high bitrates but struggle to preserve fine-grained semantic information at low bitrates. Recently, with the blossoming of Multimodal Large Language Models (MLLMs), Cross-modal compression techniques offer prospective solutions for improving video compression under low-bitrate conditions. In this paper, we propose a unified Cross-Modality Video Compression (CMVC) framework that integrates multimodal representations and video generative models. The encoder disentangles video into spatial and temporal components, which are mapped to compact cross modal representations using MLLMs. During decoding, different encoding-decoding modes are employed to acquire various video reconstruction qualities, including Text-Text-to-Video (TT2V) for semantic preservation and Image-Text-to-Video (IT2V) for perceptual consistency. Additionally, we elaborate on an efficient frame interpolation model using Low-Rank Adaptation (LoRA) to improve the perceptual quality. Experimental results demon strate that TT2V achieves effective semantic reconstruction, while IT2V ensures competitive perceptual consistency. These findings suggest the potential of leveraging multimodal priors to improve video compression, offering promising future research directions.
Jinlong Li 0003, Kecheng Chen, Meng Wang 0017, Long Xu 0001, Haoliang Li, Nicu Sebe, Sam Kwong, Shiqi Wang 0001
IEEE Signal Process. Lett.3
2025 Test-time Adaptation for Foundation Medical Segmentation Model without Parametric Updates
abstract
Foundation medical segmentation models, with MedSAM being the most popular, have achieved promising performance across organs and lesions. However, MedSAM still suffers from compromised performance on specific lesions with intricate structures and appearance, as well as bounding box prompt-induced perturbations. Although current test-time adaptation (TTA) methods for medical image segmentation may tackle this issue, partial (e.g., batch normalization) or whole parametric updates restrict their effectiveness due to limited update signals or catastrophic forgetting in large models. Meanwhile, these approaches ignore the computational complexity during adaptation, which is particularly significant for modern foundation models. To this end, our theoretical analyses reveal that directly refining image embeddings is feasible to approach the same goal as parametric updates under the MedSAM architecture, which enables us to realize high computational efficiency and segmentation performance without the risk of catastrophic forgetting. Under this framework, we propose to encourage maximizing factorized conditional probabilities of the posterior prediction probability using a proposed distribution-approximated latent conditional random field loss combined with an entropy minimization loss. Experiments show that we achieve about 3\% Dice score improvements across three datasets while reducing computational complexity by over 7 times.
Kecheng Chen, Xinyu Luo, Tiexin Qin, Jie Liu 0044, Hui Liu 0036, Victor Ho-fun Lee, Hong Yan 0001, Haoliang Li
ICCV1
2025 Test-time Adaptation for Image Compression with Distribution Regularization
abstract
Current test- or compression-time adaptation image compression (TTA-IC) approaches, which leverage both latent and decoder refinements as a two-step adaptation scheme, have potentially enhanced the rate-distortion (R-D) performance of learned image compression models on cross-domain compression tasks, \textit{e.g.,} from natural to screen content images. However, compared with the emergence of various decoder refinement variants, the latent refinement, as an inseparable ingredient, is barely tailored to cross-domain scenarios. To this end, we are interested in developing an advanced latent refinement method by extending the effective hybrid latent refinement (HLR) method, which is designed for \textit{in-domain} inference improvement but shows noticeable degradation of the rate cost in \textit{cross-domain} tasks. Specifically, we first provide theoretical analyses, in a cue of marginalization approximation from in- to cross-domain scenarios, to uncover that the vanilla HLR suffers from an underlying mismatch between refined Gaussian conditional and hyperprior distributions, leading to deteriorated joint probability approximation of marginal distribution with increased rate consumption. To remedy this issue, we introduce a simple Bayesian approximation-endowed \textit{distribution regularization} to encourage learning a better joint probability approximation in a plug-and-play manner. Extensive experiments on six in- and cross-domain datasets demonstrate that our proposed method not only improves the R-D performance compared with other latent refinement counterparts, but also can be flexibly integrated into existing TTA-IC methods with incremental benefits.
Kecheng Chen, Tiexin Qin, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li
ICLR1
2025 Lightweight Medical Image Restoration via Integrating Reliable Lesion-Semantic Driven Prior
abstract
Medical image restoration tasks aim to recover high-quality images from degraded observations, exhibiting emergent desires in many clinical scenarios, such as low-dose CT image denoising, MRI super-resolution, and MRI artifact removal. Despite the success achieved by existing deep learning-based restoration methods with sophisticated modules, they struggle with rendering computationally-efficient reconstruction results. Moreover, they usually ignore the reliability of the restoration results, which is much more urgent in medical systems. To alleviate these issues, we present LRformer, a Lightweight Transformer-based method via Reliability-guided learning in the frequency domain. Specifically, inspired by the uncertainty quantification in Bayesian neural networks (BNNs), we develop a Reliable Lesion-Semantic Prior Producer (RLPP). RLPP leverages Monte Carlo (MC) estimators with stochastic sampling operations to generate sufficiently-reliable priors by performing multiple inferences on the foundational medical image segmentation model, MedSAM. Additionally, instead of directly incorporating the priors in the spatial domain, we decompose the cross-attention (CA) mechanism into real symmetric and imaginary anti-symmetric parts via fast Fourier transform (FFT), resulting in the design of the Guided Frequency Cross-Attention (GFCA) solver. By leveraging the conjugated symmetric property of FFT, GFCA reduces the computational complexity of naive CA by nearly half. Extensive experimental results in various tasks demonstrate the superiority of the proposed LRformer in both effectiveness and efficiency.
Kecheng Chen, Jiaxin Huang 0006, Yazhou Ren 0001, Xiaorong Pu
ACM Multimedia2
2025 Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need
abstract
We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e.g.,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs.
Kecheng Chen, Hui Liu 0036, Jie Liu 0044, Yibing Liu, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li
NeurIPS1
2025 SPACE: SPike-Aware Consistency Enhancement for Test-Time Adaptation in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs), as a biologically plausible alternative to Artificial Neural Networks (ANNs), have demonstrated advantages in terms of energy efficiency, temporal processing, and biological plausibility. However, SNNs are highly sensitive to distribution shifts, which can significantly degrade their performance in real-world scenarios. Traditional test-time adaptation (TTA) methods designed for ANNs often fail to address the unique computational dynamics of SNNs, such as sparsity and temporal spiking behavior. To address these challenges, we propose SPike-Aware Consistency Enhancement (SPACE), the first source-free and single-instance TTA method specifically designed for SNNs. SPACE leverages the inherent spike dynamics of SNNs to maximize the consistency of spike-behavior-based local feature maps across augmented versions of a single test sample, enabling robust adaptation without requiring source data. We evaluate SPACE on multiple datasets. Furthermore, SPACE exhibits robust generalization across diverse network architectures, consistently enhancing the performance of SNNs on CNNs, Transformer, and ConvLSTM architectures. Experimental results show that SPACE outperforms state-of-the-art ANN methods while maintaining lower computational cost, highlighting its effectiveness and robustness for SNNs in real-world settings. The code will be available at https://github.com/ethanxyluo/SPACE.
Xinyu Luo, Kecheng Chen, Pao-Sheng Sun, Chris Xing Tian, Arindam Basu, Haoliang Li
NeurIPS2
2025 DP-TTA: Test-Time Adaptation for Transient Electromagnetic Signal Denoising via Dictionary-Driven Prior Regularization
abstract
Transient Electromagnetic (TEM) method is widely used in various geophysical applications, providing valuable insights into subsurface properties. However, time-domain TEM signals are often submerged in various types of noise. While recent deep learning-based denoising models have shown strong performance, these models are mostly trained on simulated or single real-world scenario data, overlooking the significant differences in noise characteristics from different geographical regions. Intuitively, models trained in one environment often struggle to perform well in new settings due to differences in geological conditions, equipment, and external interference, leading to reduced denoising performance. To this end, we propose the Dictionary-driven Prior Regularization Test-time Adaptation (DP-TTA). Our key insight is that TEM signals possess intrinsic physical characteristics, such as exponential decay and smoothness, which remain consistent across different regions regardless of external conditions. These intrinsic characteristics serve as ideal prior knowledge for guiding the TTA strategy, which helps the pre-trained model dynamically adjust parameters by utilizing self-supervised losses, improving denoising performance in new scenarios. To implement this, we customized a network, named DTEMDNet. Specifically, we first use dictionary learning to encode these intrinsic characteristics as a dictionary-driven prior, which is integrated into the model during training. At the testing stage, this prior guides the model to adapt dynamically to new environments by minimizing self-supervised losses derived from the dictionary-driven consistency and the signal one-order variation. Extensive experimental results demonstrate that the proposed method achieves much better performance than existing TEM denoising methods and TTA methods.
Meng Yang 0030, Kecheng Chen, Xianjie Chen, Yong Jia, Fanqiang Lin
IEEE Trans. Geosci. Remote. Sens.2
2025 Unsupervised Domain Adaptation for Low-Dose CT Reconstruction via Bayesian Uncertainty Alignment
abstract
Low-dose computed tomography (LDCT) image reconstruction techniques can reduce patient radiation exposure while maintaining acceptable imaging quality. Deep learning (DL) is widely used in this problem, but the performance of testing data (also known as target domain) is often degraded in clinical scenarios due to the variations that were not encountered in training data (also known as source domain). Unsupervised domain adaptation (UDA) of LDCT reconstruction has been proposed to solve this problem through distribution alignment. However, existing UDA methods fail to explore the usage of uncertainty quantification, which is crucial for reliable intelligent medical systems in clinical scenarios with unexpected variations. Moreover, existing direct alignment for different patients would lead to content mismatch issues. To address these issues, we propose to leverage a probabilistic reconstruction framework to conduct a joint discrepancy minimization between source and target domains in both the latent and image spaces. In the latent space, we devise a Bayesian uncertainty alignment to reduce the epistemic gap between the two domains. This approach reduces the uncertainty level of target domain data, making it more likely to render well-reconstructed results on target domains. In the image space, we propose a sharpness-aware distribution alignment (SDA) to achieve a match of second-order information, which can ensure that the reconstructed images from the target domain have similar sharpness to normal-dose CT (NDCT) images from the source domain. Experimental results on two simulated datasets and one clinical low-dose imaging dataset show that our proposed method outperforms other methods in quantitative and visualized performance.
Kecheng Chen, Jie Liu 0044, Renjie Wan, Victor Ho-fun Lee, Varut Vardhanabhuti, Hong Yan 0001, Haoliang Li
IEEE Trans. Neural Networks Learn. Syst.1
2024 Frequency-Constraint VQ-VAE for Adaptive MRI Segmentation
Kecheng Chen, Yazhou Ren 0001, Xiaorong Pu
ICONIP (9)5
2024 Domain Generalization with Small Data
abstract
Abstract In this work, we propose to tackle the problem of domain generalization in the context of insufficient samples. Instead of extracting latent feature embeddings based on deterministic models, we propose to learn a domain-invariant representation based on the probabilistic framework by mapping each data point into probabilistic embeddings. Specifically, we first extend empirical maximum mean discrepancy (MMD) to a novel probabilistic MMD that can measure the discrepancy between mixture distributions (i.e., source domains) consisting of a series of latent distributions rather than latent points. Moreover, instead of imposing the contrastive semantic alignment (CSA) loss based on pairs of latent points, a novel probabilistic CSA loss encourages positive probabilistic embedding pairs to be closer while pulling other negative ones apart. Benefiting from the learned representation captured by probabilistic models, our proposed method can marriage the measurement on the distribution over distributions (i.e., the global perspective alignment) and the distribution-based contrastive semantic alignment (i.e., the local perspective alignment). Extensive experimental results on three challenging medical datasets show the effectiveness of our proposed method in the context of insufficient data compared with state-of-the-art methods.
Kecheng Chen, Elena Gal, Hong Yan 0001, Haoliang Li
Int. J. Comput. Vis.1
2024 AutoAssign+: Automatic Shared Embedding Assignment in streaming recommendation
Ziru Liu, Kecheng Chen, Fengyi Song, Bo Chen 0023, Xiangyu Zhao 0001, Huifeng Guo, Ruiming Tang
Knowl. Inf. Syst.2
2024 Enhancing Transferability of Adversarial Examples Through Mixed-Frequency Inputs
abstract
Recent studies have shown that Deep Neural Networks (DNNs) are easily deceived by adversarial examples, revealing their serious vulnerability. Due to the transferability, adversarial examples can attack across multiple models with different architectures, called transfer-based black-box attacks. Input transformation is one of the most effective methods to improve adversarial transferability. In particular, the attacks fusing other categories of image information reveal the potential direction of adversarial attacks. However, the current techniques rely on input transformations in the spatial domain, which ignore the frequency information of the image and limit its transferability. To tackle this issue, we propose Mixed-Frequency Inputs (MFI) based on a frequency domain perspective. MFI alleviates the overfitting of adversarial examples to the source model by considering high-frequency components from various kinds of images in the process of calculating the gradient. By accumulating these high-frequency components, MFI acquires a more steady gradient direction in each iteration, leading to the discovery of better local maxima and enhancing transferability. Extensive experimental results on the ImageNet-compatible datasets demonstrate that MFI outperforms existing transform-based attacks with a clear margin on both Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), which proves MFI is more suitable for realistic black-box scenarios.
Yaguan Qian, Kecheng Chen, Bin Wang 0062, Zhaoquan Gu, Shouling Ji, Wei Wang 0012, Yanchun Zhang
IEEE Trans. Inf. Forensics Secur.2
2024 Learning Robust Shape Regularization for Generalizable Medical Image Segmentation
abstract
Generalizable medical image segmentation enables models to generalize to unseen target domains under domain shift issues. Recent progress demonstrates that the shape of the segmentation objective, with its high consistency and robustness across domains, can serve as a reliable regularization to aid the model for better cross-domain performance, where existing methods typically seek a shared framework to render segmentation maps and shape prior concurrently. However, due to the inherent texture and style preference of modern deep neural networks, the edge or silhouette of the extracted shape will inevitably be undermined by those domain-specific texture and style interferences of medical images under domain shifts. To address this limitation, we devise a novel framework with a separation between the shape regularization and the segmentation map. Specifically, we first customize a novel whitening transform-based probabilistic shape regularization extractor namely WT-PSE to suppress undesirable domain-specific texture and style interferences, leading to more robust and high-quality shape representations. Second, we deliver a Wasserstein distance-guided knowledge distillation scheme to help the WT-PSE to achieve more flexible shape extraction during the inference phase. Finally, by incorporating domain knowledge of medical images, we propose a novel instance-domain whitening transform method to facilitate a more stable training process with improved performance. Experiments demonstrate the performance of our proposed method on both multi-domain and single-domain generalization.
Kecheng Chen, Tiexin Qin, Victor Ho-fun Lee, Hong Yan 0001, Haoliang Li
IEEE Trans. Medical Imaging1
2024 Cross-Domain Low-Dose CT Image Denoising With Semantic Preservation and Noise Alignment
abstract
Deep learning (DL)-based Low-dose CT (LDCT) image denoising methods may face domain shift problem, where data from different domains (i.e., hospitals) may have similar anatomical regions but exhibit different intrinsic noise characteristics. Therefore, we propose a plug-and-play model called Lowand High-frequency Alignment (LHFA) to address this issue by leveraging semantic features and aligning noise distributions of different CT datasets, while maintaining diagnostic image quality and suppressing noise. Specifically, the LHFA model consists of a Low-frequency Alignment (LFA) module that preserves semantic features (i.e., low-frequency components) with fewer perturbations from both domains for reconstruction. Notably, a Highfrequency Alignment (HFA) module is proposed to quantify the discrepancy between noise representations (i.e., high-frequency components) in a latent space mapped by an auto-encoder. Experimental results demonstrate that the LHFA model effectively alleviates the domain shift problem and significantly improves the performance of DL-based methods on cross-domain LDCT image denoising task, outperforming other domain adaptationbased methods.
Jiaxin Huang 0006, Kecheng Chen, Yazhou Ren 0001, Xiaorong Pu, Ce Zhu
IEEE Trans. Multim.2
2023 Cross-Domain Object Classification Via Successive Subspace Alignment
abstract
Recently, successive subspace learning (SSL)-based methods have shown to be effective for the task of visual object classification with mild data desire and mathematically transparent interpretable capability. However, existing SSL-based methods rely heavily on the data-centric subspace representations, leading to potential performance degradation problem in case of the domain shift between the training (a.k.a., source domain) and testing (a.k.a., target domain) data. To address this limitation, we propose an effective successive subspace learning method based on existing SSL-based methods. Specifically, we introduce a novel linear transformation layer to align eigenvectors in SSL module between source and target domains, as such, the discrepancy between source and target domains will be reduced, resulting in better cross-domain performance. The effectiveness of our proposed method is demonstrated on the Office-Caltech-10 and Office-31 benchmark datasets by using features extracted from pre-trained deep neural networks as input.
Kecheng Chen, Haoliang Li, Hong Yan 0001
ICASSP1
2022 Cross Domain Low-Dose CT Image Denoising With Semantic Information Alignment
abstract
Recently, cross domain adaptation has been applied into quite a few image restoration tasks. While promising performance has been achieved, the domain shift problem between the training set (a.k.a., source domain) and the testing set (a.k.a., target domain) in Low-dose Computed Tomography (LDCT) image denoising tasks is typically ignored by most existing methods. This is prone to the degradation of the denoising performance due to large discrepancy of feature distribution in each dataset from various vendors. Therefore, a simple yet effective LDCT denoising approach has been proposed in this paper to alleviate the domain shift between source and target domains through a novel semantic information alignment. Specifically, we first propose an adaptive version of random frequency mask (RFM) to extract the shared semantic information of cross domains. Then, we incorporate the mask into the existing denoiser to construct a semantic-information-guided objective. Experiments on synthetic and real datasets show our proposed method achieves impressive performance.
Jiaxin Huang 0006, Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001
ICIP2
2022 ClusterUDA: Latent Space Clustering in Unsupervised Domain Adaption for Pulmonary Nodule Detection
Kecheng Chen, Xiaorong Pu, Chao Li 0034, Yazhou Ren 0001
ICONIP (6)4
2022 TEMDnet: A Novel Deep Denoising Network for Transient Electromagnetic Signal With Signal-to-Image Transformation
abstract
The considerable prospecting depth and accurate subsurface characteristics can be obtained by the transient electromagnetic method (TEM) in geophysics. Nevertheless, the time-domain TEM signal received by the coil is easily disturbed by environmental background noise, artificial noise, and electronic noise of the equipment. Recently, deep neural networks (DNNs) have been used to solve the TEM denoising problem and have achieved better performance than traditional methods. However, the existing denoising method with DNN adopts fully connected neural networks and is therefore not flexible enough to deal with various signal scales. To address these issues, a novel denoising framework with deep convolutional neural networks (CNNs) of transforming the TEM signal denoising task into an image denoising task (namely, TEMDnet) is proposed in this article. Specifically, a novel signal-to-image transformation method is developed first to preserve the structural features of TEM signals. Then, a novel deep CNN-based denoiser is proposed to further perform feature learning, in which the residual learning mechanism is adopted to model the noise estimation image for different signal features. Extensive experiments demonstrate that the proposed framework can achieve much better performance compared with other state-of-the-art approaches on both simulated signals and real-world signals from a landfill leachate treatment plant in Chengdu, Sichuan, China. Models and code are available at https://github.com/tonyckc/TEMDnet_demo.
Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001, Hang Qiu 0002, Fanqiang Lin, Saimin Zhang
IEEE Trans. Geosci. Remote. Sens.1
2021 Lesion-Inspired Denoising Network: Connecting Medical Image Denoising and Lesion Detection
abstract
Deep learning has achieved notable performance in the denoising task of low-quality medical images and the detection task of lesions, respectively. However, existing low-quality medical image denoising approaches are disconnected from the detection task of lesions. Intuitively, the quality of denoised images will influence the lesion detection accuracy that in turn can be used to affect the denoising performance. To this end, we propose a play-and-plug medical image denoising framework, namely Lesion-Inspired Denoising Network (LIDnet), to collaboratively improve both denoising performance and detection accuracy of denoised medical images. Specifically, we propose to insert the feedback of downstream detection task into existing denoising framework by jointly learning a multi-loss objective. Instead of using perceptual loss calculated on the entire feature map, a novel region-of-interest (ROI) perceptual loss induced by the lesion detection task is proposed to further connect these two tasks. To achieve better optimization for overall framework, we propose a customized collaborative training strategy for LIDnet. On consideration of clinical usability and imaging characteristics, three low-dose CT images datasets are used to evaluate the effectiveness of the proposed LIDnet. Experiments show that, by equipping with LIDnet, both of the denoising and lesion detection performance of baseline methods can be significantly improved.
Kecheng Chen, Kun Long, Yazhou Ren 0001, Xiaorong Pu
ACM Multimedia1
2020 Low-Dose CT Image Blind Denoising with Graph Convolutional Networks
Kecheng Chen, Xiaorong Pu, Yazhou Ren 0001, Hang Qiu 0002, Haoliang Li
ICONIP (1)1