Cheng Wang 0043

dblp:54/2062-43 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-6506-1221ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Fiber HGNN: Heterogeneous Graph Neural Network for Fiber Tract Segmentation
abstract
Fiber tract segmentation is crucial for clinical applications such as brain function interpretation and surgical planning. Existing methods typically adopt either a cortical-parcellation-based or fiber clustering approach, but fail to simultaneously integrate heterogeneous information (e.g., streamline shape, point position, anatomical priors). In this work, we propose Fiber HGNN, a novel heterogeneous graph neural network that explicitly models and integrates heterogeneous information of fibers for accurate fiber tract segmentation. We construct a heterogeneous graph comprising three types of nodes: streamline, fiber keypoint and anatomical region. Specifically, fiber keypoints are representative points sampled along each streamline to characterize local geometric features, while anatomical regions provide contextual priors derived from brain atlas. This design enables the network to jointly capture the complementary information of streamline shape, local geometry, and anatomical priors, thus facilitating the learning of more discriminative feature representations. To further leverage implicit anatomical connectivity, we design a Metapath-guided Heterogeneous Information Aggregation (MHIA) network. By analyzing the spatial relationships between streamline keypoints and anatomical regions, the heterogeneous graph is decomposed into anatomical subgraphs for each streamline. In each subgraph, heterogeneous information from metapath-linked nodes is aggregated to obtain the final fiber representation. We evaluate the effectiveness of our framework on the HCP105 and TractoInferno datasets. The experimental results demonstrate that our method significantly outperforms previous state-of-the-art methods. The source code is available at https://github.com/CUHK-AIM-Group/Fiber-HGNN.
Cheng Wang 0043, Wuyang Li, Xinyu Liu 0001, Yifan Liu 0010, Jian Cheng 0002, Yixuan Yuan
IEEE Trans. Medical Imaging1
2025 U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation
abstract
U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures.
Chenxin Li, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Hengyu Liu 0007, Yifan Liu 0010, Zhen Chen 0013, Yixuan Yuan
AAAI4
2025 InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling
Chenxin Li, Yifan Liu 0010, Panwang Pan, Hengyu Liu 0007, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Weihao Yu 0004, Yiyang Lin, Yixuan Yuan
ICCV7
2025 MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
abstract
High-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected low-resolution (LR) medical videos present unique challenges for video super-resolution (VSR) models, including camera shake, noise, and abrupt frame transitions, which result in significant optical flow errors and alignment difficulties. Additionally, tissues and organs exhibit continuous and nuanced structures, but current VSR models are prone to introducing artifacts and distorted features that can mislead doctors. To this end, we propose MedVSR, a tailored framework for medical VSR. It first employs Cross State-Space Propagation (CSSP) to address the imprecise alignment by projecting distant frames as control matrices within state-space models, enabling the selective propagation of consistent and informative features to neighboring frames for effective alignment. Moreover, we design an Inner State-Space Reconstruction (ISSR) module that enhances tissue structures and reduces artifacts with joint long-range spatial feature learning and large-kernel short-range information aggregation. Experiments across four datasets in diverse medical scenarios, including endoscopy and cataract surgeries, show that MedVSR significantly outperforms existing VSR models in reconstruction performance and efficiency. Code released at https://github.com/CUHK-AIM-Group/MedVSR.
Xinyu Liu 0001, Guolei Sun, Cheng Wang 0043, Yixuan Yuan, Ender Konukoglu
ICCV3
2025 EndoGen: Conditional Autoregressive Endoscopic Video Generation
Xinyu Liu 0001, Hengyu Liu 0007, Cheng Wang 0043, Tianming Liu 0001, Yixuan Yuan
MICCAI (10)3
2024 PV-SSM: Exploring Pure Visual State Space Model for High-dimensional Medical Data Analysis
abstract
Despite previous endeavors to utilize Convolutional Neural Networks and Transformers as base networks for medical image analysis, their architectures still harbor inherent limitations: either an inability to model long-range dependencies or colossal computational consumption due to global self-attention. Recently, State Space Models (SSMs) have exhibited impressive capabilities in modeling long-term dependencies with satisfactory linear computational complexity. Nevertheless, extant medical visual SSMs are constrained by their limited capacity to capture inter-patch relationships and inefficient modeling due to the introduction of additional depth convolutions to handle high-dimensional data. In this paper, we propose a novel, Pure Visual State Space Model (PV-SSM) for high-dimensional medical data analysis. Different from prior medical visual SSMs, our proposed framework does not involve any convolutional or global attention operations while leverages a series of Pure-SSM blocks that employ a novel parallel-SSM mechanism to simultaneously extract feature data across different dimensions. Furthermore, we propose a learnable Parameterized Positional Encoding, which incorporates absolute positional information into patch features, effectively endowing inter-patch relationships with stronger inferential capabilities. We conducted extensive validation on various modalities of medical imaging data. Experimental results demonstrate superior performance and efficacy of our model against existing models. Our codes are available at https://github.com/chengwang96/PV-SSM
Cheng Wang 0043, Xinyu Liu 0001, Chenxin Li, Yifan Liu 0010, Yixuan Yuan
BIBM1
2024 GTP-4o: Modality-Prompted Heterogeneous Graph Learning for Omni-Modal Biomedical Representation
Chenxin Li, Xinyu Liu 0001, Cheng Wang 0043, Yifan Liu 0010, Weihao Yu 0005, Yixuan Yuan
ECCV (4)3
2024 EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting
Chenxin Li, Brandon Yushan Feng, Yifan Liu 0010, Hengyu Liu 0007, Cheng Wang 0043, Weihao Yu 0005, Yixuan Yuan
MICCAI (6)5
2024 When 3D Partial Points Meets SAM: Tooth Point Cloud Segmentation with Sparse Labels
Yifan Liu 0010, Wuyang Li, Cheng Wang 0043, Hui Chen 0032, Yixuan Yuan
MICCAI (11)3
2023 Person Search by a Bi-Directional Task-Consistent Learning Model
abstract
Two-stage person search methods achieve the state-of-the-art performance by separate detection and re-ID stages, but neglect the consistency needs between these two stages. The re-ID stage needs more accurate query bounding boxes and fewer boxes of distractors; The detection stage needs the re-ID stage to have robustness against unavailable detection errors. In this paper, we introduce a novel Bi-directional Task-Consistent Learning (BTCL) person search framework, including a Target-Specific Detector (TSD) and a re-ID model with Dynamic Adaptive Learning Structure (DALS). For the former consistency need, we add a verification head for predicting the similarity scores between query and proposals in parallel with the existing heads for bounding box recognition. Thus, TSD generates accurate boxes for the query-like pedestrians, which are suitable for the re-ID stage. For the re-ID robustness need, DALS dynamically generates a large number of possible detection results in line with the real distribution. By training the re-ID model on data with different types of detection errors, DLAS improves the model robustness to detection inputs. Experimental results show our framework achieves state-of-the-art performance on two widely-used person search datasets.
Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Multim.1
2022 MonoEF: Extrinsic Parameter Free Monocular 3D Object Detection
abstract
Monocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing methods on open datasets neglect the camera pose information, which inevitably results in the detector being susceptible to camera extrinsic parameters. The perturbation of objects is very popular in most autonomous driving cases for industrial products. To this end, we propose a novel method to capture camera pose to formulate the detector free from extrinsic perturbation. Specifically, the proposed framework predicts camera extrinsic parameters by detecting vanishing point and horizon change. A converter is designed to rectify perturbative features in the latent space. By doing so, our 3D detector works independent of the extrinsic parameter variations and produces accurate results in realistic cases, e.g., potholed and uneven roads, where almost all existing monocular detectors fail to handle. Experiments demonstrate our method yields the best performance compared with the other state-of-the-arts by a large margin on both KITTI 3D and nuScenes datasets.
Yunsong Zhou, Hongzi Zhu, Cheng Wang 0043, Qinhong Jiang
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Monocular 3D Object Detection: An Extrinsic Parameter Free Approach
abstract
Monocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing methods on open datasets neglect the cam-era pose information, which inevitably results in the detector being susceptible to camera extrinsic parameters. The perturbation of objects is very popular in most autonomous driving cases for industrial products. To this end, we propose a novel method to capture camera pose to formulate the detector free from extrinsic perturbation. Specifically, the proposed framework predicts camera extrinsic parameters by detecting vanishing point and horizon change. A converter is designed to rectify perturbative features in the latent space. By doing so, our 3D detector works independent of the extrinsic parameter variations and produces accurate results in realistic cases, e.g., potholed and uneven roads, where almost all existing monocular detectors fail to handle. Experiments demonstrate our method yields the best performance compared with the other state-of-the-arts by a large margin on both KITTI 3D and nuScenes datasets.
Yunsong Zhou, Hongzi Zhu, Cheng Wang 0043, Qinhong Jiang
CVPR4
2020 TCTS: A Task-Consistent Two-Stage Framework for Person Search
abstract
The state of the art person search methods separate person search into detection and re-ID stages, but ignore the consistency between these two stages. The general person detector has no special attention on the query target; The re-ID model is trained on hand-drawn bounding boxes which are not available in person search. To address the consistency problem, we introduce a Task-Consist Two-Stage (TCTS) person search framework, includes an identity-guided query (IDGQ) detector and a Detection Results Adapted (DRA) re-ID model. In the detection stage, the IDGQ detector learns an auxiliary identity branch to compute query similarity scores for proposals. With consideration of the query similarity scores and foreground score, IDGQ produces query-like bounding boxes for the re-ID stage. In the re-ID stage, we predict identity labels of detected bounding boxes, and use these examples to construct a more practical mixed train set for the DRA model. Training on the mixed train set improves the robustness of the re-ID stage to inaccurate detection. We evaluate our method on two benchmark datasets, CUHK-SYSU and PRW. Our framework achieves 93.9% of mAP and 95.1% of rank1 accuracy on CUHK-SYSU, outperforming the previous state of the art methods.
Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
CVPR1