Chenxin Li

dblp:00/9389 · DBLP profile ↗
← Back
50ranked-venue papers
12as first author
48since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 8 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 8 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online Communities
abstract
Counterspeech offers a non-repressive approach to moderate hate speech in online communities. Research has examined how counterspeech chatbots restrain hate speakers and support targets, but their impact on bystanders remains unclear. Therefore, we developed a counterspeech strategy framework and built Civilbot for a mixed-method within-subjects study. Bystanders generally viewed Civilbot as credible and normative, though its shallow reasoning limited persuasiveness. Its behavioural effects were subtle: when performing well, it could guide participation or act as a stand-in; when performing poorly, it could discourage bystanders or motivate them to step in. Strategy proved critical: cognitive strategies that appeal to reason, especially when paired with a positive tone, were relatively effective, while mismatch of contexts and strategies could weaken impact. Based on these findings, we offer design insights for mobilizing bystanders and shaping online discourse, highlighting when to intervene and how to do so through reasoning-driven and context-aware strategies.
Shuai Ma 0005, Peng Zhang 0060, Chenxin Li, Ning Gu 0001, Tun Lu
CHI5
2026 ST-MTLSM: A Device-Partitioned Multi-Tier LSM-Tree with Non-Blocking Snapshot Publication for Massive Spatio-Temporal IoT Data
Jianwen Yang, Qiuhong Zhang, Zhiming Ding, Xinguo Chen, Chenxin Li, Jian Miao, Xueyu Gao
DEXA (2)6
2026 Attribute knowledge inheritance and evolution for generalized zero-shot learning
Zhijie Rao, Jingcai Guo, Chenxin Li, Song Guo 0001
Inf. Sci.3
2026 Harnessing Lightweight Transformer With Contextual Synergic Enhancement for Efficient 3D Medical Image Segmentation
abstract
Transformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their applicability. To address these challenges, we consider two crucial aspects: model efficiency and data efficiency. Specifically, we propose Light-UNETR, a lightweight transformer designed to achieve model efficiency. Light-UNETR features a Lightweight Dimension Reductive Attention (LIDR) module, which reduces spatial and channel dimensions while capturing both global and local features via multi-branch attention. Additionally, we introduce a Compact Gated Linear Unit (CGLU) to selectively control channel interaction with minimal parameters. Furthermore, we introduce a Contextual Synergic Enhancement (CSE) learning strategy, which aims to boost the data efficiency of Transformers. It first leverages the extrinsic contextual information to support the learning of unlabeled data with Attention-Guided Replacement, then applies Spatial Masking Consistency that utilizes intrinsic contextual information to enhance the spatial context reasoning for unlabeled data. Extensive experiments on various benchmarks demonstrate the superiority of our approach in both performance and efficiency. For example, with only 10% labeled data on the Left Atrial Segmentation dataset, our method surpasses BCP by 1.43% Jaccard while drastically reducing the FLOPs by 90.8% and parameters by 85.8%.
Xinyu Liu 0001, Zhen Chen 0013, Wuyang Li, Chenxin Li, Yixuan Yuan
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 De-LightSAM: Modality-Decoupled Lightweight SAM for Generalizable Medical Segmentation
abstract
The universality of deep neural networks across different modalities and their generalization capabilities to unseen domains play an essential role in medical image segmentation. The recent segment anything model (SAM) has demonstrated strong adaptability across diverse natural scenarios. However, the huge computational costs, demand for manual annotations as prompts and conflict-prone decoding process of SAM degrade its generalization capabilities in medical scenarios. To address these limitations, we propose a modality-decoupled lightweight SAM for domain-generalized medical image segmentation, named De-LightSAM. Specifically, we first devise a lightweight domain-controllable image encoder (DC-Encoder) that produces discriminative visual features for diverse modalities. Further, we introduce the self-patch prompt generator (SP-Generator) to automatically generate high-quality dense prompt embeddings for guiding segmentation decoding. Finally, we design the query-decoupled modality decoder (QM-Decoder) that leverages a one-to-one strategy to provide an independent decoding channel for every modality, preventing mutual knowledge interference of different modalities. Moreover, we design a multi-modal decoupled knowledge distillation (MDKD) strategy to leverage robust common knowledge to complement domain-specific medical feature representations. Extensive experiments indicate that De-LightSAM outperforms state-of-the-arts in diverse medical imaging segmentation tasks, displaying superior modality universality and generalization capabilities. Especially, De-LightSAM uses only 2.0% parameters compared to SAM-H. The source code is available at https://github.com/xq141839/De-LightSAM.
Qing Xu 0014, Xiangjian He, Chenxin Li, Fiseha B. Tesema, Wenting Duan, Zhen Chen 0013, Rong Qu, Jonathan M. Garibaldi, Chang Wen Chen
IEEE Trans. Circuits Syst. Video Technol.4
2025 U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation
abstract
U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures.
Chenxin Li, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Hengyu Liu 0007, Yifan Liu 0010, Zhen Chen 0013, Yixuan Yuan
AAAI1
2025 Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
abstract
Video anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos—either by identifying anomalous frames or objects—they often neglect finer-grained analysis, such as anomalous pixels, which limits their ability to capture a broader range of anomalies. To address this challenge, we propose an innovative VAD framework called Track Any Anomalous Object (TAO), which introduces a Granular Video Anomaly Detection Framework that, for the first time, integrates the detection of multiple fine-grained anomalous objects into a unified framework. Unlike methods that assign anomaly scores to every pixel at each moment, our approach transforms the problem into pixel-level tracking of anomalous objects. By linking anomaly scores to subsequent tasks such as image segmentation and video tracking, our method eliminates the need for threshold selection and achieves more precise anomaly localization, even in long and challenging video sequences. Experiments on extensive datasets demonstrate that TAO achieves state-of-the-art performance, setting a new progress for VAD by providing a practical, granular, and holistic solution. For more information, visit the project page at: https://tao-25.github.io/
Yuzhi Huang, Chenxin Li, Zixu Lin, Yunlong Lin, Hengyu Liu 0007, Wuyang Li, Xinyu Liu 0001, Jiechao Gao, Yue Huang 0001, Xinghao Ding, Yixuan Yuan
CVPR2
2025 JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
abstract
Vision-centric perception systems struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world conditions, we propose JarvisIR, a VLM-powered agent that leverages the VLM as a controller to manage multiple expert restoration models. To further enhance system robustness, reduce hallucinations, and improve generalizability in real-world adverse weather, JarvisIR employs a novel two-stage framework consisting of supervised fine-tuning and human feedback alignment. Specifically, to address the lack of paired data in real-world scenarios, the human feedback alignment enables the VLM to be fine-tuned effectively on large-scale real-world data in an unsupervised manner. To support the training and evaluation of JarvisIR, we introduce CleanBench, a comprehensive dataset consisting of high-quality and large-scale instruction-responses pairs, including 150K synthetic entries and 80K real entries. Extensive experiments demonstrate that JarvisIR exhibits superior decision-making and restoration capabilities. Compared with existing methods, it achieves a 50% improvement in the average of all perception metrics on CleanBench-Real.
Yunlong Lin, Zixu Lin, Haoyu Chen 0003, Panwang Pan, Chenxin Li, Sixiang Chen, Kairun Wen, Yeying Jin, Wenbo Li 0002, Xinghao Ding
CVPR5
2025 MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
abstract
Recent advances in generalizable 3D Gaussian Splatting have demonstrated promising results in real-time high-fidelity rendering without per-scene optimization, yet existing approaches still struggle to handle unfamiliar visual content during inference on novel scenes due to limited generalizability. To address this challenge, we introduce MonoSplat, a novel framework that leverages rich visual priors from pre-trained monocular depth foundation models for robust Gaussian reconstruction. Our approach consists of two key components: a Mono-Multi Feature Adapter that transforms monocular features into multi-view representations, coupled with an Integrated Gaussian Prediction module that effectively fuses both feature types for precise Gaussian generation. Through the Adapter’s lightweight attention mechanism, features are seamlessly aligned and aggregated across views while preserving valuable monocular priors, enabling the Prediction module to generate Gaussian primitives with accurate geometry and appearance. Through extensive experiments on diverse real-world datasets, we convincingly demonstrate that MonoSplat achieves superior reconstruction quality and generalization capability compared to existing methods while maintaining computational efficiency with minimal trainable parameters. Codes are available at https://github.com/CUHK-AIM-Group/MonoSplat.
Yifan Liu 0010, Keyu Fan, Weihao Yu 0005, Chenxin Li, Hao Lu 0003, Yixuan Yuan
CVPR4
2025 FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
abstract
3D Gaussian splatting (3DGS) has enabled various applications in 3D scene representation and novel view synthesis due to its efficient rendering capabilities. However, 3DGS demands relatively significant GPU memory, limiting its use on devices with restricted computational resources. Previous approaches have focused on pruning less important Gaussians, effectively compressing 3DGS but often requiring a fine-tuning stage and lacking adaptability for the specific memory needs of different devices. In this work, we present an elastic inference method for 3DGS. Given an input for the desired model size, our method selects and transforms a subset of Gaussians, achieving substantial rendering performance without additional fine-tuning. We introduce a tiny learnable module that controls Gaussian selection based on the input percentage, along with a transformation module that adjusts the selected Gaussians to complement the performance of the reduced model. Comprehensive experiments on ZipNeRF, MipNeRF and Tanks&Temples scenes demonstrate the effectiveness of our approach. Code is available at https://flexgs.github.io/.
Hengyu Liu 0007, Yuehao Wang, Chenxin Li, Ruisi Cai, Wuyang Li, Pavlo Molchanov 0001, Peihao Wang, Zhangyang Wang
CVPR3
2025 ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
abstract
As 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. While steganography has advanced significantly in common 3D media like meshes and Neural Radiance Fields (NeRF), research into steganography for 3D- GS representations remains largely unexplored. To address this gap, we propose ConcealGS, a novel 3D steganography method that embeds implicit information into the explicit 3D representation of Gaussian Splatting. By introducing a consistency strategy for the decoder and a gradient optimization approach, ConcealGS overcomes limitations of NeRF-based models, enhancing both the robustness of implicit information and the quality of 3D reconstruction. Extensive evaluations across various potential application scenarios demonstrate that ConcealGS successfully recovers implicit information with negligible impact on rendering quality, offering a groundbreaking approach for embedding invisible yet recoverable information into 3D models. This work paves the way for advanced copyright protection and secure data transmission in the evolving landscape of 3D content creation and distribution. Code is available at https://github.com/zxk1212/ConcealGS.
Hengyu Liu 0007, Chenxin Li, Yining Sun, Wuyang Li, Yifan Liu 0010, Yiyang Lin, Yixuan Yuan, Nanyang Ye 0001
ICASSP3
2025 InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling
Chenxin Li, Yifan Liu 0010, Panwang Pan, Hengyu Liu 0007, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Weihao Yu 0004, Yiyang Lin, Yixuan Yuan
ICCV1
2025 Metascope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
Wuyang Li, Wentao Pan 0001, Zhendong Luo, Chenxin Li, Hengyu Liu 0007, Din Ping Tsai, Mu Ku Chen, Yixuan Yuan
ICCV5
2025 Dissecting Generalized Category Discovery: Multiplex Consensus under Self-Deconstruction
Luyao Tang, Kunze Huang, Chaoqi Chen, Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICCV5
2025 $\mathbf{X}^{\mathbf{2}}$-Gaussian: 4D Radiative Gaussian Splatting for Continuous-Time Tomographic Reconstruction
Weihao Yu 0005, Yuanhao Cai, Ruyi Zha, Zhiwen Fan, Chenxin Li, Yixuan Yuan
ICCV5
2025 InstantSplamp: Fast and Generalizable Stenography Framework for Generative Gaussian Splatting
abstract
With the rapid development of large generative models for 3D, especially the evolution from NeRF representations to more efficient Gaussian Splatting, the synthesis of 3D assets has become increasingly fast and efficient, enabling the large-scale publication and sharing of generated 3D objects. However, while existing methods can add watermarks or steganographic information to individual 3D assets, they often require time-consuming per-scene training and optimization, leading to watermarking overheads that can far exceed the time required for asset generation itself, making deployment impractical for generating large collections of 3D objects. To address this, we propose InstantSplamp a framework that seamlessly integrates the 3D steganography pipeline into large 3D generative models without introducing explicit additional time costs. Guided by visual foundation models,InstantSplamp subtly injects hidden information like copyright tags during asset generation, enabling effective embedding and recovery of watermarks within generated 3D assets while preserving original visual quality. Experiments across various potential deployment scenarios demonstrate that \model~strikes an optimal balance between rendering quality and hiding fidelity, as well as between hiding performance and speed. Compared to existing per-scene optimization techniques for 3D assets, InstantSplamp reduces their watermarking training overheads that are multiples of generation time to nearly zero, paving the way for real-world deployment at scale. Project page: https://gaussian-stego.github.io/.
Chenxin Li, Hengyu Liu 0007, Zhiwen Fan, Wuyang Li, Yifan Liu 0010, Panwang Pan, Yixuan Yuan
ICLR1
2025 Hide-in-Motion: Embedding Steganographic Copyright Information into 4D Gaussian Splatting Assets
abstract
As 4D extensions of 3D Gaussian Splatting (4D-GS) emerge as groundbreaking techniques for dynamic scene reconstruction and novel view synthesis in robotics and computer vision, ensuring the security and trustworthiness of these assets becomes crucial. While steganography has advanced significantly in 2D and 3D media, existing methods are inadequate for the complex, dynamic nature of 4D-GS representations. To address this gap, we propose Hide-in-Motion, a novel 4D steganography method for hiding information through deformation in Gaussian splatting. Our approach introduces a composite attribute and a Decouple Feature Field for coarse-to-fine deformation modeling and embedding implicit information, along with an Opacity-Guided Adaptive strategy. Hide-in-Motion overcomes the limitations of previous techniques, enhancing both the robustness of embedded information and the quality of 4D reconstruction. Extensive evaluations demonstrate that our method successfully embeds and recovers implicit information across various modalities while maintaining high rendering quality in dynamic scenes. This work not only advances copyright protection and secure data transmission for 4D assets but also paves the way for enhancing the security and integrity of 4D digital assets. Code is available at https://github.com/CUHK-AIM-Group/Hide-in-Motion.
Hengyu Liu 0007, Chenxin Li, Wentao Pan 0001, Zhiqin Yang, Yifan Liu 0010, Wuyang Li, Yixuan Yuan
ICRA2
2025 Pan-LUT: Efficient Pan-sharpening via Learnable Look-Up Tables
abstract
Recently, deep learning-based pan-sharpening algorithms have achieved notable advancements over traditional methods. However, deep learning-based methods incur substantial computational overhead during inference, especially with large images. This excessive computational demand limits the applicability of these methods in real-world scenarios, particularly in the absence of dedicated computing devices such as GPUs and TPUs. To address these challenges, we propose Pan-LUT, a novel learnable look-up table (LUT) framework for pan-sharpening that strikes a balance between performance and computational efficiency for large remote sensing images. Our method makes it possible to process 15K$\times$15K remote sensing images on a 24GB GPU. To finely control the spectral transformation, we devise the PAN-guided look-up table (PGLUT) for channel-wise spectral mapping. To effectively capture fine-grained spatial details, we introduce the spatial details look-up table (SDLUT). Furthermore, to adaptively aggregate channel information for generating high-resolution multispectral images, we design an adaptive output look-up table (AOLUT). Our model contains fewer than 700K parameters and processes a 9K$\times$9K image in under 1 ms using one RTX 2080 Ti GPU, demonstrating significantly faster performance compared to other methods. Experiments reveal that Pan-LUT efficiently processes large remote sensing images in a lightweight manner, bridging the gap to real-world applications. Furthermore, our model surpasses SOTA methods in full-resolution scenes under real-world conditions, highlighting its effectiveness and efficiency. We also extend our method to general image fusion tasks.
Zhongnan Cai, Yingying Wang 0005, Hui Zheng 0003, Panwang Pan, Zixu Lin, Ge Meng, Chenxin Li, Chunming He, Jiaxin Xie, Yunlong Lin, Junbin Lu, Yue Huang 0001, Xinghao Ding
NeurIPS7
2025 JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
abstract
Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based solutions provide automation but often suffer from limited adjustability and poor generalization, failing to meet diverse and personalized editing needs. To bridge this gap, we introduce JarvisArt, a multi-modal large language model (MLLM)-driven agent that understands user intent, mimics the reasoning process of professional artists, and intelligently coordinates over 200 retouching tools within Lightroom. JarvisArt undergoes a two-stage training process: an initial Chain-of-Thought supervised fine-tuning to establish basic reasoning and tool-use skills, followed by Group Relative Policy Optimization for Retouching (GRPO-R) to further enhance its decision-making and tool proficiency. We also propose the Agent-to-Lightroom Protocol to facilitate seamless integration with Lightroom. To evaluate performance, we develop MMArt-Bench, a novel benchmark constructed from real-world user edits. JarvisArt demonstrates user-friendly interaction, superior generalization, and fine-grained control over both global and local adjustments, paving a new avenue for intelligent photo retouching. Notably, it outperforms GPT-4o with a 60\% improvement in average pixel-level metrics on MMArt-Bench for content fidelity, while maintaining comparable instruction-following capabilities.
Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen 0003, Zhongdao Wang, Xinghao Ding, Wenbo Li 0002, Shuicheng Yan
NeurIPS6
2025 IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
abstract
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active creation rather than passive recognition. Grounded in the analysis-by-synthesis paradigm, IR3D-Bench tasks Vision-Language Agents (VLAs) with actively using programming and rendering tools to recreate the underlying 3D structure of an input image, achieving agentic inverse rendering through tool use. This ''understanding-by-creating'' approach probes the tool-using generative capacity of VLAs, moving beyond the descriptive or conversational capacity measured by traditional scene understanding benchmarks. We provide a comprehensive suite of metrics to evaluate geometric accuracy, spatial relations, appearance attributes, and overall plausibility. Initial experiments on agentic inverse rendering powered by various state-of-the-art VLMs highlight current limitations, particularly in visual precision rather than basic tool usage. IR3D-Bench, including data and evaluation protocols, is released to facilitate systematic study and development of tool-using VLAs towards genuine scene understanding by creating.
Hengyu Liu 0007, Chenxin Li, Yipeng Wu, Wuyang Li, Zhiqin Yang, Zhenyuan Zhang 0001, Yunlong Lin, Sirui Han, Brandon Yushan Feng
NeurIPS2
2025 HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
abstract
Recent advances in generative models have achieved high-fidelity in 3D human reconstruction, yet their utility for specific tasks (e.g., human 3D segmentation) remains constrained. We propose HumanCrafter, a unified framework that enables the joint modeling of appearance and human-part semantics from a single image in a feed-forward manner. Specifically, we integrate human geometric priors in the reconstruction stage and self-supervised semantic priors in the segmentation stage. To address labeled 3D human datasets scarcity, we further develop an interactive annotation procedure for generating high-quality data-label pairs. Our pixel-aligned aggregation enables cross-task synergy, while the multi-task objective simultaneously optimizes texture modeling fidelity and semantic consistency. Extensive experiments demonstrate that HumanCrafter surpasses existing state-of-the-art methods in both 3D human-part segmentation and 3D human reconstruction **from a single image**.
Panwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin, Kairun Wen, Yixuan Yuan
NeurIPS3
2025 DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
abstract
Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structure-from-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce **DynamicVerse**, a physical‑scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consists of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods.
Kairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 0003, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Junbin Lu, Chenguo Lin, Dilin Wang, Zhicheng Yan 0001, Hongyu Xu, Justin Theiss, Yue Huang 0001, Xinghao Ding, Zhiwen Fan
NeurIPS7
2025 FedGPS: Statistical Rectification Against Data Heterogeneity in Federated Learning
abstract
Federated Learning (FL) confronts a significant challenge known as data heterogeneity, which impairs model performance and convergence. Existing methods have made notable progress in addressing this issue. However, improving performance in certain heterogeneity scenarios remains an overlooked question: _How robust are these methods to deploy under diverse heterogeneity scenarios?_ To answer this, we conduct comprehensive evaluations across varied heterogeneity scenarios, showing that most existing methods exhibit limited robustness. Meanwhile, insights from these experiments highlight that sharing statistical information can mitigate heterogeneity by enabling clients to update with a global perspective. Motivated by this, we propose **FedGPS** (**Fed**erated **G**oal-**P**ath **S**ynergy), a novel framework that seamlessly integrates statistical distribution and gradient information from others. Specifically, FedGPS statically modifies each client’s learning objective to implicitly model the global data distribution using surrogate information, while dynamically adjusting local update directions with gradient information from other clients at each round. Extensive experiments show that FedGPS outperforms state-of-the-art methods across diverse heterogeneity scenarios, validating its effectiveness and robustness. The code is available at: <https://github.com/CUHK-AIM-Group/FedGPS>.
Zhiqin Yang, Yonggang Zhang 0003, Chenxin Li, Yiu-Ming Cheung, Bo Han 0003, Yixuan Yuan
NeurIPS3
2025 ADCNet: a unified framework for predicting the activity of antibody-drug conjugates
abstract
Antibody-drug conjugates (ADCs) have revolutionized the field of cancer treatment in the era of precision medicine due to their ability to precisely target cancer cells and release highly effective drugs. Nevertheless, the rational design and discovery of ADCs remain challenging because the relationship between their quintuple structures and activities is difficult to explore and understand. To address this issue, we first introduce a unified deep learning framework called ADCNet to explore such relationship and help design potential ADCs. The ADCNet highly integrates the protein representation learning language model ESM-2 and small-molecule representation learning language model functional group-based bidirectional encoder representations from transformers to achieve activity prediction through learning meaningful features from antigen and antibody protein sequences of ADC, SMILES strings of linker and payload, and drug-antibody ratio (DAR) value. Based on a carefully designed and manually tailored ADC data set, extensive evaluation results reveal that ADCNet performs best on the test set compared to baseline machine learning models across all evaluation metrics. For example, it achieves an average prediction accuracy of 87.12%, a balanced accuracy of 0.8689, and an area under receiver operating characteristic curve of 0.9293 on the test set. In addition, cross-validation, ablation experiments, and external independent testing results further prove the stability, advancement, and robustness of the ADCNet architecture. For the convenience of the community, we develop the first online platform (https://ADCNet.idruglab.cn) for the prediction of ADCs activity based on the optimal ADCNet model, and the source code is publicly available at https://github.com/idrugLab/ADCNet.
Liye Chen, Biaoshun Li, Mujie Lin, Chenxin Li, Ling Wang 0008
Briefings Bioinform.6
2025 Human-to-robot handovers based on multimodal perception
Chunfang Liu, Weifan Wang 0008, Chenxin Li, Ruitian Pang, Yan Shang, Yitong Gao
Neurocomputing3
2025 NuSegDG: Integration of heterogeneous space and Gaussian kernel for domain-generalized nuclei segmentation
Zhenye Lou, Qing Xu 0014, Zekun Jiang, Xiangjian He, Chenxin Li, Zhen Chen 0013, Yi Wang 0037, Maggie M. He, Wenting Duan
Knowl. Based Syst.5
2025 Foundation Model-Guided Gaussian Splatting for 4D Reconstruction of Deformable Tissues
abstract
Reconstructing deformable anatomical structures from endoscopic videos is a pivotal and promising research topic that can enable advanced surgical applications and improve patient outcomes. While existing surgical scene reconstruction methods have made notable progress, they often suffer from slow rendering speeds due to using neural radiance fields, limiting their practical viability in real-world applications. To overcome this bottleneck, we propose EndoGaussian, a framework that integrates the strengths of 3D Gaussian Splatting representations, allowing for high-fidelity tissue reconstruction, efficient training, and real-time rendering. Specifically, we dedicate a Foundation Model-driven Initialization (FMI) module, which distills 3D cues from multiple vision foundation models (VFMs) to swiftly construct the preliminary scene structure for Gaussian initialization. Then, a Spatio-temporal Gaussian Tracking (SGT) is designed, efficiently modeling scene dynamics using the multi-scale HexPlane with spatio-temporal priors. Furthermore, to improve the dynamics modeling ability for scenes with large deformation, EndoGaussian integrates Motion-aware Frame Synthesis (MFS) to adaptively synthesize new frames as extra training constraints. Experimental results on public datasets demonstrate EndoGaussian's efficacy against prior state-of-the-art methods, including superior rendering speed (168 FPS, real-time), enhanced rendering quality (38.555 PSNR), and reduced training overhead (within 2 min/scene). These results underscore EndoGaussian's potential to significantly advance intraoperative surgery applications, paving the way for more accurate and efficient real-time surgical guidance and decision-making in clinical scenarios. Code is available at: https://github.com/CUHK-AIM-Group/EndoGaussian.
Yifan Liu 0010, Chenxin Li, Hengyu Liu 0007, Chen Yang 0026, Yixuan Yuan
IEEE Trans. Medical Imaging2
2024 Learning to Estimate 6DoF Pose from Limited Data: A Few-Shot, Generalizable Approach using RGB Images
abstract
The accurate estimation of six degrees-of-freedom (6DoF) object poses is essential for many applications in robotics and augmented reality. However, existing methods for 6DoF pose estimation often depend on CAD templates or dense support views, restricting their usefulness in real-world situations. In this study, we present a new cascade framework named Cas6D for few-shot 6DoF pose estimation that is generalizable and uses only RGB images. To address the false positives of target object detection in the extreme few-shot setting, our framework utilizes a self-supervised pre-trained ViT to learn robust feature representations. Then, we initialize the nearest top-K pose candidates based on similarity score and refine the initial poses using feature pyramids to formulate and update the cascade warped feature volume, which encodes context at increasingly finer scales. By discretizing the pose search range using multiple pose bins and progressively narrowing the pose search range in each stage using predictions from the previous stage, Cas6D can overcome the large gap between pose candidates and ground truth poses, which is a common failure mode in sparse-view scenarios. Experimental results on the LINEMOD and GenMOP datasets demonstrate that Cas6D outperforms state-of-the-art methods by 9.2% and 3.8& accuracy (Proj-5) under the 32-shot setting compared to OnePose++ and Gen6D. Our framework also performs best under the full-shot setting with all support views. Code is avilable at https://github.com/github.com/paulpanwang/Cas6D.
Panwang Pan, Zhiwen Fan, Brandon Yushan Feng, Peihao Wang, Chenxin Li, Zhangyang Wang
3DV5
2024 PV-SSM: Exploring Pure Visual State Space Model for High-dimensional Medical Data Analysis
abstract
Despite previous endeavors to utilize Convolutional Neural Networks and Transformers as base networks for medical image analysis, their architectures still harbor inherent limitations: either an inability to model long-range dependencies or colossal computational consumption due to global self-attention. Recently, State Space Models (SSMs) have exhibited impressive capabilities in modeling long-term dependencies with satisfactory linear computational complexity. Nevertheless, extant medical visual SSMs are constrained by their limited capacity to capture inter-patch relationships and inefficient modeling due to the introduction of additional depth convolutions to handle high-dimensional data. In this paper, we propose a novel, Pure Visual State Space Model (PV-SSM) for high-dimensional medical data analysis. Different from prior medical visual SSMs, our proposed framework does not involve any convolutional or global attention operations while leverages a series of Pure-SSM blocks that employ a novel parallel-SSM mechanism to simultaneously extract feature data across different dimensions. Furthermore, we propose a learnable Parameterized Positional Encoding, which incorporates absolute positional information into patch features, effectively endowing inter-patch relationships with stronger inferential capabilities. We conducted extensive validation on various modalities of medical imaging data. Experimental results demonstrate superior performance and efficacy of our model against existing models. Our codes are available at https://github.com/chengwang96/PV-SSM
Cheng Wang 0043, Xinyu Liu 0001, Chenxin Li, Yifan Liu 0010, Yixuan Yuan
BIBM3
2024 Dual Adaptive Compression for Efficient Communication in Heterogeneous Federated Learning
abstract
In federated learning, multiple rounds of communication are involved between clients and the server to train a global model. The extensive model updates transmitted during the training lead to significant communication costs. Previous methods usually employ quantization or sparsification to compress model updates. However, the lossy compression leads to a decline in accuracy, it is challenging to strike a balance between communication efficiency and model accuracy. Meanwhile, due to the data heterogeneity, local updates among different clients are biased towards each other. Employing the same compression ratios for each local updates will further degrade the model accuracy. To achieve the trade-off between communication efficiency and model accuracy, we propose FedDAC, a Dual Adaptive Compression method in heterogeneous federated learning. In the local computation phase, the loss queue is adopted to detect the convergence trends within each client. FedDAC can then dynamically quantify model updates and allow for various compression ratios among heterogeneous clients. In the global aggregation phase, FedDAC can determine the fluctuations in training based on the similarity between clients and the server, thereby adjusting the sparsity ratio flexibly. To alleviate the reduction in model accuracy caused by lossy compression, we introduce residual updates in the local computation and global aggregation phases to maintain model accuracy. Experiment results show that compared with one-way compression methods NAGC and AdaQuantFL, FedDAC can maintain comparable accuracy while the accumulated communication volume is reduced by about 29.6 times, and 22.8 times, respectively. Moreover, the global model accuracy of FedDAC surpasses the two-way compression method T-FedAvg by about 2.4%, and the accumulated communication volume is about 2.5 times lower than T-FedAvg.
Yingchi Mao, Chenxin Li, Jiakai Zhang, Shufang Xu, Jie Wu 0001
CCGrid3
2024 GTP-4o: Modality-Prompted Heterogeneous Graph Learning for Omni-Modal Biomedical Representation
Chenxin Li, Xinyu Liu 0001, Cheng Wang 0043, Yifan Liu 0010, Weihao Yu 0005, Yixuan Yuan
ECCV (4)1
2024 Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification
abstract
Recent advances in fine-tuning Vision-Language Models (VLMs) have witnessed the success of prompt tuning and adapter tuning, while the classic model fine-tuning on inherent parameters seems to be overlooked.It is believed that fine-tuning the parameters of VLMs with few-shot samples corrupts the pre-trained knowledge since fine-tuning the CLIP model even degrades performance.In this paper, we revisit this viewpoint, and propose a new perspective: fine-tuning the specific parameters instead of all will uncover the power of classic model fine-tuning on VLMs.Through our meticulous study, we propose ClipFit, a simple yet effective method to fine-tune CLIP without introducing any overhead of extra parameters.We demonstrate that by only fine-tuning the specific bias terms and normalization layers, ClipFit can improve the performance of zero-shot CLIP by 7.27% average harmonic mean accuracy.Lastly, to understand how fine-tuning in CLIPFit affects the pre-trained models, we conducted extensive experimental analyses w.r.t.changes in internal parameters and representations.We found that low-level text bias layers and the first layer normalization layer change much more than other layers.The code is available at https://github.com/minglllli/CLIPFit.
Jike Zhong, Chenxin Li, Liuzhuozheng Li, Nie Lin, Masashi Sugiyama
EMNLP3
2024 Perceptual Learning in Lexical Tone: Phonetic Similarity vs. Phonological Categories
abstract
In speech comprehension, listeners recalibrate their interpretation of variable speech signals through exposure and disambiguating information. Recalibration is attested both segmentally and suprasegmentally, but little is known about what constrains it in lexical tone. This project investigated the effects of phonological categories and phonetic similarity on perceptual learning. We exposed Chinese listeners to pitch contours ambiguous between two tone categories (realised with a level or rising pitch contour) and lexically biased their perception to one interpretation. Crucially, the rising pitch contour could be from two different phonological tone categories. Perceptual learning was observed not only in the rising contours used for exposure but also across phonetically similar but phonologically different rising contours, suggesting that perceptual learning in tone is not constrained by phonological tone categories but is facilitated by the phonetic similarity of pitch contours.
Ariëlle Reitsema, Chenxin Li, Leanne van Lambalgen, Laura Preining, Saskia Galindo Jong, Xinyi Wen, Yiya Chen
INTERSPEECH2
2024 EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting
Chenxin Li, Brandon Yushan Feng, Yifan Liu 0010, Hengyu Liu 0007, Cheng Wang 0043, Weihao Yu 0005, Yixuan Yuan
MICCAI (6)1
2024 👦 Endora: Video Generation Models as Endoscopy Simulators
Chenxin Li, Hengyu Liu 0007, Yifan Liu 0010, Brandon Yushan Feng, Wuyang Li, Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan
MICCAI (6)1
2024 LGS: A Light-Weight 4D Gaussian Splatting for Efficient Surgical Scene Reconstruction
Hengyu Liu 0007, Yifan Liu 0010, Chenxin Li, Wuyang Li, Yixuan Yuan
MICCAI (3)3
2024 P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images
abstract
Generating diverse plausible outputs from a single input is crucial for addressing visual ambiguities, exemplified in medical imaging where experts may provide varying semantic segmentation annotations for the same image.Existing methods handles ambiguous segmentation relying on probabilistic modeling and extensive multi-output annotated data while often struggles with limited ambiguously labeled datasets common in real-world applications.To surmount the challenge, we propose P²SAM, a novel framework that leverages the Segment Anything Model (SAM)'s prior knowledge for ambiguous object segmentation. By transforming SAM's sensitivity to prompts into an advantage, we introduce a prior probabilistic space for prompts.Experimental results show that P²SAM significantly enhances medical segmentation precision and diversity using minimal ambiguously annotated samples. Benchmarking against state-of-the-art methods demonstrates superior performance with just 5.5% of the training data (+12% Dmax). This approach marks a significant advancement towards deploying probabilistic models in data-limited real-world scenarios.
Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu 0007, Haote Xu, Yifan Liu 0010, Yue Huang 0001, Xinghao Ding, Xiaotong Tu, Yixuan Yuan
ACM Multimedia2
2024 Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM
abstract
As the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https://a-sa-m.github.io/}.
Chenxin Li, Yuzhi Huang, Wuyang Li, Hengyu Liu 0007, Xinyu Liu 0001, Qing Xu 0014, Zhen Chen 0013, Yue Huang 0001, Yixuan Yuan
NeurIPS1
2023 Differential Privacy in Federated Dynamic Gradient Clipping Based on Gradient Norm
Yingchi Mao, Chenxin Li, Zijian Tu, Ping Ping
ICA3PP (4)2
2023 Hint-Dynamic Knowledge Distillation
abstract
Knowledge Distillation (KD) transfers the knowledge from a high-capacity teacher model to promote a smaller student model. Existing efforts guide the distillation by matching their prediction logits, feature embedding, etc., while leaving how to efficiently utilize them in junction less explored. In this paper, we propose Hint-dynamic Knowledge Distillation, dubbed HKD, which excavates the knowledge from the teacher’s hints in a dynamic scheme. The guidance effect from the knowledge hints usually varies in different instances and learning stages, which motivates us to customize a specific hint-learning manner for each instance adaptively. Specifically, a meta-weight network is introduced to generate the instance-wise weight coefficients about knowledge hints in the perception of the dynamical learning progress of the student model. We further present a weight ensembling strategy to eliminate the potential bias of coefficient estimation by exploiting the historical statics. Experiments on standard benchmarks of CIFAR-100 and Tiny-ImageNet manifest that the proposed HKD well boost the effect of knowledge distillation tasks.
Chenxin Li, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ICASSP2
2023 StegaNeRF: Embedding Invisible Information within Neural Radiance Fields
abstract
Recent advancements in neural rendering have paved the way for a future marked by the widespread distribution of visual data through the sharing of Neural Radiance Field (NeRF) model weights. However, while established techniques exist for embedding ownership or copyright information within conventional visual data such as images and videos, the challenges posed by the emerging NeRF format have remained unaddressed. In this paper, we introduce StegaNeRF, an innovative approach for steganographic information embedding within NeRF renderings. We have meticulously developed an optimization framework that enables precise retrieval of hidden information from images generated by NeRF, while ensuring the original visual quality of the rendered images to remain intact. Through rigorous experimentation, we assess the efficacy of our methodology across various potential deployment scenarios. Furthermore, we delve into the insights gleaned from our analysis. StegaNeRF represents an initial foray into the intriguing realm of infusing NeRF renderings with customizable, imperceptible, and recoverable information, all while minimizing any discernible impact on the rendered images. For more details, please visit our project page: https://xggnet.github.io/StegaNeRF/
Chenxin Li, Brandon Yushan Feng, Zhiwen Fan, Panwang Pan, Zhangyang Wang
ICCV1
2022 Knowledge Condensation Distillation
Chenxin Li, Mingbao Lin, Zhiyuan Ding, Nie Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Liujuan Cao
ECCV (11)1
2022 Unsupervised Anomaly Segmentation for Brain Lesions Using Dual Semantic-Manifold Reconstruction
Zhiyuan Ding, Haote Xu, Chenxin Li, Xinghao Ding, Yue Huang 0001
ICONIP (3)4
2022 Hierarchical deep network with uncertainty-aware semi-supervised learning for vessel segmentation
Chenxin Li, Wenao Ma, Liyan Sun, Xinghao Ding, Yue Huang 0001, Guisheng Wang, Yizhou Yu
Neural Comput. Appl.1
2021 Unsupervised Large-Scale Social Network Alignment via Cross Network Embedding
abstract
Nowadays, it is common for a person to possess different identities on multiple social platforms. Social network alignment aims to match the identities that from different networks. Recently, unsupervised network alignment methods have received significant attention since no identity anchor is required. However, to capture the relevance between identities, the existing unsupervised methods generally rely heavily on user profiles, which is unobtainable and unreliable in real-world scenarios. In this paper, we propose an unsupervised alignment framework named Large-Scale Network Alignment (LSNA) to integrate the network information and reduce the requirement on user profile. The embedding module of LSNA, named Cross Network Embedding Model (CNEM), aims to integrate the topology information and the network correlation to simultaneously guide the embedding process. Moreover, in order to adapt LSNA to large-scale networks, we propose a network disassembling strategy to divide the costly large-scale network alignment problem into multiple executable sub-problems. The proposed method is evaluated over multiple real-world social network datasets, and the results demonstrate that the proposed method outperforms the state-of-the-art methods.
Zhehan Liang, Yu Rong 0001, Chenxin Li, Yue Huang 0001, Tingyang Xu, Xinghao Ding, Junzhou Huang
CIKM3
2021 Consistent Posterior Distributions Under Vessel-Mixing: A Regularization For Cross-Domain Retinal Artery/Vein Classification
abstract
Retinal artery/vein (A/V) classification is a critical technique for diagnosing diabetes and cardiovascular diseases. Although deep learning based methods achieve impressive results in A/V classification, the performance usually degrades when directly apply the models that trained on one dataset to another set, due to the domain shift, e.g., caused by the variations in imaging protocols. In this paper, we propose a novel method to improve cross-domain generalization for pixel-wise retinal A/V classification. That is, vessel-mixing based consistency regularization, which regularizes the models to give consistent posterior distributions for vessel-mixing samples. The proposed method achieves the state-of-the-art performance on extensive experiments for cross-domain A/V classification, which is even close to the performance of fully supervised learning on target domain in some cases.
Chenxin Li, Zhehan Liang, Wenao Ma, Yue Huang 0001, Xinghao Ding
ICIP1
2021 Generator Versus Segmentor: Pseudo-healthy Synthesis
Chenxin Li, Liyan Sun, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu
MICCAI (6)2
2021 Diamond: a multi-modal DIA mass spectrometry data processing pipeline
abstract
SUMMARY: Currently, various software tools are used to support two mainstream workflows for data-independent acquisition (DIA) mass spectrometry (MS) data processing, namely, spectrum-centric scoring (SCS) and peptide-centric scoring (PCS). However, a fully automatic, easily reproducible and freely accessible pipeline that simultaneously integrates SCS and PCS strategies and supports both library-free and library-based modes is absent. We developed Diamond, a Nextflow-based, containerized, multi-modal DIA-MS data processing pipeline for peptide identification and quantification. Diamond integrated two mainstream workflows for DIA data analysis, namely, SCS and PCS, for use cases both with and without assay libraries. This multi-modal pipeline serves as a versatile, easy-to-use and easily extendable toolbox for large-scale DIA data processing. AVAILABILITY: Diamond is hosted on GitHub (https://github.com/xmuyulab/Diamond) and is released under the highly permissive MIT license to encourage further customization and modification. The Docker image for Diamond is freely accessible at https://hub.docker.com/r/zeroli/diamond.
Chenxin Li, Mingxuan Gao, Chuanqi Zhong, Rongshan Yu
Bioinform.1
2019 Enhanced Resource Selection Mechanism for LTE-V2X Sidelink Multi-Carrier Operation
abstract
Vehicle-to-everything (V2X) can improve the road safety, the traffic efficiency, and the availability of infotainment services. In order to meet the requirement of higher data rate for Long Term Evolution (LTE)-V2X, the carrier aggregation scheme was introduced in 3GPP Release 15. The enhanced resource selection mechanism for sidelink multi-carrier was studied, to support UEs with limited transmission capability performing multi- carrier operations and to alleviate the half- duplex impact. The system level simulation results show that the half-duplex impact can be suppressed and the proposed mechanisms in this paper can improve the reliability of LTE-V2X sidelink multi- carrier operation.
Jin-Ling Hu, Chenxin Li, Jia-Yi Fang, Yan Shi 0002
VTC Spring2
2018 The Performance Comparison of LTE-V2X and IEEE 802.11p
abstract
Vehicle-to-everything (V2X) can improve the road safety, the traffic efficiency, and the availability of infotainment services. Standardization of Long Term Evolution (LTE)-V2X has been actively conducted by the Third Generation Partnership Project (3GPP) to provide the solutions for V2X. In this paper, the challenges and detailed design issues in LTE V2X are discussed. The performance comparison of LTE-V2X and IEEE 802.11p is presented in typical urban and freeway scenarios.
Jia-Yi Fang, Jin-Ling Hu, Yan Shi 0002, Chenxin Li
VTC Spring7