VLDB 2026 Research / reviewers in the wild / expert
Tom Weidong Cai
dblp:c/WeidongCai · also Weidong Cai 0001
· DBLP profile ↗
188ranked-venue papers
2as first author
91since 2021 · last 2026
0000-0003-3706-8896ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 117 · 1 first-author · 60 since 2021Artificial intelligence and machine learning · 78 · 45 since 2021Applied, interdisciplinary, general and emerging computing · 67 · 1 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding LearningabstractUniversal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and lack diversity in negative samples. Moreover, the embeddings exhibit limited discriminative ability in distinguishing false and hard negatives. In this paper, we leverage the advanced understanding capabilities of MLLMs to enhance representation learning, and present a novel Universal Multimodal Embedding(UniME-V2) model. Our approach first constructs a potential hard negative set through global retrieval. We then introduce the MLLM-as-a-Judge mechanism, which utilizes MLLMs to assess the semantic alignment of query-candidate pairs and generate soft semantic matching scores. These scores serve as a foundation for hard negative mining, mitigating the impact of false negatives and enabling the identification of diverse, high-quality hard negatives. Furthermore, the semantic matching scores are used as soft labels to mitigate the rigid one-to-one mapping constraint. By aligning the similarity matrix with the soft semantic matching score matrix, the model learns semantic distinctions among candidates, significantly enhancing its discriminative capacity. To further improve performance, we propose UniME-V2, a reranking model trained on our mined hard negatives through a joint pairwise and listwise optimization approach. We conduct comprehensive experiments on the MMEB benchmark and multiple retrieval tasks, demonstrating that our method achieves state-of-the-art performance across all tasks. Tiancheng Gu, Kaicheng Yang 0002, Kaichen Zhang, Xiang An, Ziyong Feng, Tom Weidong Cai, Jiankang Deng, Lidong Bing |
AAAI | 7 |
| 2026 | Gotta Hear Them All: Towards Sound Source Aware Audio GenerationabstractAudio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or texts. However, the immersiveness and expressiveness of the generation are limited. One possible problem is that existing methods solely rely on the global scene and overlook details of local sounding objects (i.e., sound sources). To address this issue, we propose a Sound Source-Aware Audio (SS2A) generator. SS2A is able to locally perceive multimodal sound sources from a scene with visual detection and cross-modality translation. It then contrastively learns a Cross-Modal Sound Source (CMSS) Manifold to semantically disambiguate each source. Finally, we attentively mix their CMSS semantics into a rich audio representation, from which a pretrained audio generator outputs the sound. To model the CMSS manifold, we curate a novel single-sound-source visual-audio dataset VGGS3 from VGGSound. We also design a Sound Source Matching Score to clearly measure localized audio relevance. With the effectiveness of explicit sound source modeling, SS2A achieves state-of-the-art performance in extensive image-to-audio tasks. We also qualitatively demonstrate SS2A's ability to achieve intuitive synthesis control by compositing vision, text, and audio conditions. Furthermore, we show that our sound source modeling can achieve competitive video-to-audio performance with a straightforward temporal aggregation mechanism. Heng Wang 0007, Tom Weidong Cai |
AAAI | 4 |
| 2026 | HiFusion: Hierarchical Intra-Spot Alignment and Regional Context Fusion for Spatial Gene Expression Prediction from HistopathologyabstractSpatial transcriptomics (ST) bridges gene expression and tissue morphology but faces clinical adoption barriers due to technical complexity and prohibitive costs. While computational methods predict gene expression from H&E-stained whole-slide images (WSIs), existing approaches often fail to capture the intricate biological heterogeneity within spots and are susceptible to morphological noise when integrating contextual information from surrounding tissue. To overcome these limitations, we propose HiFusion, a novel deep learning framework that integrates two complementary components. First, we introduce the Hierarchical Intra-Spot Modeling module that extracts fine-grained morphological representations through multi-resolution sub-patch decomposition, guided by a feature alignment loss to ensure semantic consistency across scales. Concurrently, we present the Context-aware Cross-scale Fusion module, which employs cross-attention to selectively incorporate biologically relevant regional context, thereby enhancing representational capacity. This architecture enables comprehensive modeling of both cellular-level features and tissue microenvironmental cues, which are essential for accurate gene expression prediction. Extensive experiments on two benchmark ST datasets demonstrate that HiFusion achieves state-of-the-art performance across both 2D slide-wise cross-validation and more challenging 3D sample-specific scenarios. These results underscore HiFusion’s potential as a robust, accurate, and scalable solution for ST inference from routine histopathology. Ziqiao Weng, Yaoyu Fang, Jiahe Qian, Xinkun Wang, Lee A. Cooper, Tom Weidong Cai, Bo Zhou 0009 |
AAAI | 6 |
| 2026 | Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM DecodingabstractExisting vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not grounded in the visual input. Efforts to address this issue without model finetuning primarily mitigate hallucination by contrastively reducing language biases or amplifying the weights of visual embedding during decoding. However, these approaches remain limited in their ability to capture fine-grained visual details. In this work, we propose the Perception Magnifier (PM), a novel visual decoding method that iteratively isolates relevant visual tokens based on attention and magnifies the corresponding regions, spurring the model to concentrate on fine-grained visual details during decoding. By magnifying critical regions while preserving the structural and contextual information at each decoding step, PM allows the VLM to enhance its scrutiny of the visual input, hence producing more accurate and faithful responses. Extensive experimental results demonstrate that PM not only achieves superior hallucination mitigation but also enhances language generation while preserving strong reasoning capabilities. Code can be found at https://github.com/ShunqiM/PM. Shunqi Mao, Chaoyi Zhang, Tom Weidong Cai |
ACL (1) | 3 |
| 2026 | ART-ASyn: Anatomy-aware Realistic Texture-based Anomaly Synthesis Framework for Chest X-RaysabstractUnsupervised anomaly detection aims to identify anomalies without pixel-level annotations. Synthetic anomaly-based methods exhibit a unique capacity to introduce controllable irregularities with known masks, enabling explicit supervision during training. However, existing methods often produce synthetic anomalies that are visually distinct from real pathological patterns and ignore anatomical structure. This paper presents a novel Anatomy-aware Realistic Texture-based Anomaly Synthesis framework (ART-ASyn) for chest X-rays that generates realistic and anatomically consistent lung opacity related anomalies using texture-based augmentation guided by our proposed Progressive Binary Thresholding Segmentation method (PBTSeg) for lung segmentation. The generated paired samples of synthetic anomalies and their corresponding precise pixel-level anomaly mask for each normal sample enable explicit segmentation supervision. In contrast to prior work limited to one-class classification, ART-ASyn is further evaluated for zero-shot anomaly segmentation, demonstrating generalizability on an unseen dataset without target-domain annotations. Code availability is available at https://github.com/angelacao-hub/ART-ASyn. Qinyi Cao, Jianan Fan, Tom Weidong Cai |
WACV | 3 |
| 2026 | Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology ImagesabstractAccurately predicting gene expression from histopathology images offers a scalable and non-invasive approach to molecular profiling, with significant implications for precision medicine and computational pathology. However, existing methods often underutilize the cross-modal representation alignment between histopathology images and gene expression profiles across multiple representational levels, thereby limiting their prediction performance. To address this, we propose Gene-DML, a unified framework that structures latent space through Dual-pathway Multi-Level discrimination to enhance correspondence between morphological and transcriptional modalities. The multi-scale instance-level discrimination pathway aligns hierarchical histopathology representations extracted at local, neighbor, and global levels with gene expression profiles, capturing scale-aware morphological-transcriptional relationships. In parallel, the cross-level instance-group discrimination pathway enforces structural consistency between individual (image/gene) instances and modality-crossed (gene/image, respectively) groups, strengthening the alignment across modalities. By jointly modeling fine-grained and structural-level discrimination, Gene-DML is able to learn robust cross-modal representations, enhancing both predictive accuracy and generalization across diverse biological contexts. Extensive experiments on public spatial transcriptomics datasets demonstrate that Gene-DML achieves state-of-the-art performance in gene expression prediction. The code and processed datasets are available at https://github.com/YXSong000/Gene-DML. Yaxuan Song, Jianan Fan, Hang Chang, Tom Weidong Cai |
WACV | 4 |
| 2026 | IUGC: A benchmark of landmark detection in end-to-end intrapartum ultrasound biometry
Jieyun Bai, Yitong Tang, Xiao Liu 0037, Jiale Hu, Yunda Li, Xufan Chen, Yunshu Li, Bowen Guo, Jing Jiao, Lifei Li, Yuzhang Ma, Xiaoxin Han, Haochen Shao, Qingchen Liu, Jingfan Kuang, Shanglin Song, Anirvan Krishna, Zaid Ahmed Khan, Zelan Li, Zhengyang Zhang, Hansen Zhang, Xuezhi Zhang, Lyuyang Tong, Bo Du 0004, Yu Chen 0099, Zilun Peng, Saeid Rezaei, Tom Weidong Cai, Fangyijie Wang, Kathleen M. Curran, Guénolé C. M. Silvestre, Isaac Khobo, Yaosheng Lu, Dong Ni 0001, Mohammad Yaqub, Jun Ma 0016, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 39 |
| 2026 | Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 9 |
| 2026 | VGM-UNet: A hybrid visual graph deformable mamba with fourier neural operator U-Net for medical image segmentationabstractDeep learning methods have demonstrated remarkable advancements in medical image segmentation. However, achieving high accuracy remains a prominent challenge. In this paper, we present a new architecture, named VGM-UNet, that improves U-shaped segmentation models in performance and expressiveness by introducing the Structured State Space Duality algorithm to combine Graph Neural Networks, sparse attention, and Mamba-2, into U-Net and yield the best of these designs. This is accomplished through three primary modifications: we first adopt a novel approach, constructing a 2D State Space Model and an eight-way multi-scanning module, thereby creating the Vision Mamba-2, which serves as the foundation for building a hierarchical visual backbone that can be directly applied to the graph structure of image patches. Then, based on the Fast Fourier transform, we construct a Feed-Forward Network module, as a complement to Mamba, to model channel contents and improve the accuracy of capturing small objects. Moreover, on the basis of the modular architecture, we build a simple yet powerful U-shaped hybrid network, which simplifies the model design and enhances the model's expressiveness. These changes alleviate the limitations of conventional U-shaped architectures in accuracy improvements and achieve impressive results. We validate VGM-UNet through extensive experiments, demonstrating that our model outperforms existing state-of-the-art models in terms of segmentation accuracy on the Synapse and ACDC benchmark datasets. The experimental results also indicate that the Visual Graph State Space module can be conveniently applied to various medical image segmentation tasks. Jianan Fan, Tom Weidong Cai |
Neural Networks | 3 |
| 2026 | NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes Without ReferencesabstractNeural View Synthesis (NVS), such as NeRF and 3D Gaussian Splatting, effectively creates photorealistic scenes from sparse viewpoints, typically evaluated by quality assessment methods like PSNR, SSIM, and LPIPS. However, these full-reference methods, which compare synthesized views to reference views, may not fully capture the perceptual quality of neurally synthesized scenes (NSS), particularly due to the limited availability of dense reference views. Furthermore, the challenges in acquiring human perceptual labels hinder the creation of extensive labeled datasets, risking model overfitting and reduced generalizability. To address these issues, we propose NVS-SQA, a NSS quality assessment method to learn no-reference quality representations through self-supervision without reliance on human labels. Traditional self-supervised learning predominantly relies on the "same instance, similar representation" assumption and extensive datasets. However, given that these conditions do not apply in NSS quality assessment, we employ heuristic cues and quality scores as learning objectives, along with a specialized contrastive pair preparation process to improve the effectiveness and efficiency of learning. The results show that NVS-SQA outperforms 17 no-reference methods by a large margin (i.e., on average 109.5% in SRCC, 98.6% in PLCC, and 91.5% in KRCC over the second best) and even exceeds 16 full-reference methods across all evaluation metrics (i.e., 22.9% in SRCC, 19.1% in PLCC, and 18.6% in KRCC over the second best). Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Cell as Point: One-stage framework for efficient cell trackingabstractConventional multi-stage cell tracking approaches rely heavily on detection or segmentation in each frame as a prerequisite, requiring substantial resources for high-quality segmentation masks and increasing the overall prediction time. To address these limitations, we propose CAP , a novel end-to-end one-stage framework that reimagines cell tracking by treating C ell a s P oint. Unlike traditional methods, CAP eliminates the need for explicit detection or segmentation, instead jointly tracking cells for sequences in one stage by leveraging the inherent correlations among their trajectories. This simplification reduces both labeling requirements and pipeline complexity. However, directly processing the entire sequence in one stage poses challenges related to data imbalance in capturing cell division events and long sequence inference. To solve these challenges, CAP introduces two key innovations: (1) adaptive event-guided (AEG) sampling, which prioritizes cell division events to mitigate the occurrence imbalance of cell events, and (2) the rolling-as-window (RAW) inference strategy, which ensures continuous and stable tracking of newly emerging cells over extended sequences. By removing the dependency on segmentation-based preprocessing while addressing the challenges of imbalanced occurrence of cell events and long-sequence tracking, CAP demonstrates promising cell tracking performance and is 8 to 32 times more efficient than existing methods. The code and model checkpoints are available at https://github.com/YXSong000/CAP . Yaxuan Song, Jianan Fan, Heng Huang 0001, Tom Weidong Cai |
Pattern Recognit. | 5 |
| 2026 | FUGC: Benchmarking Semi-Supervised Learning Methods for Cervical SegmentationabstractAccurate segmentation of cervical structures in transvaginal ultrasound (TVS) is critical for assessing the risk of spontaneous preterm birth (PTB), yet the scarcity of labeled data limits the performance of supervised learning approaches. This paper introduces the Fetal Ultrasound Grand Challenge (FUGC), the first benchmark for semi-supervised learning in cervical segmentation, hosted at ISBI 2025. FUGC provides a dataset of 890 TVS images, including 500 training images, 90 validation images, and 300 test images. Methods were evaluated using the Dice Similarity Coefficient (DSC), Hausdorff Distance (HD), and runtime (RT), with a weighted combination of 0.4/0.4/0.2. The challenge attracted 10 teams with 82 participants submitting innovative solutions. The best-performing methods for each individual metric achieved 90.26% mDSC, 38.88 mHD, and 32.85 ms RT, respectively. FUGC establishes a standardized benchmark for cervical segmentation, demonstrates the efficacy of semi-supervised methods with limited labeled data, and provides a foundation for AI-assisted clinical PTB risk assessment. Jieyun Bai, Yitong Tang, Mahdi Islam, Musarrat Tabassum, Enrique Almar-Munoz, Nianjiang Lv, Yu Chen 0099, Zilun Peng, Yusong Xiao, Li Xiao 0002, Nam-Khanh Tran, Dac-Phu Phan-Le, Hai-Dang Nguyen, Xiao Liu 0037, Jiale Hu, Mingxu Huang, Jitao Liang, Chaolu Feng, Xuezhi Zhang, Lyuyang Tong, Bo Du 0001, Ha-Hieu Pham, Thanh-Huy Nguyen, Min Xu 0009, Juntao Jiang, Jiangning Zhang, Yong Liu 0007, Md. Kamrul Hasan 0002, Zhuonan Liang, Tom Weidong Cai, Gongning Luo, Mohammad Yaqub, Karim Lekadir |
IEEE Trans. Medical Imaging | 35 |
| 2026 | MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and RetentionabstractHistopathology and transcriptomics are fundamental modalities in cancer diagnostics, encapsulating the morphological and molecular characteristics of the disease. Multi-modal self-supervised learning has demonstrated remarkable potential in learning pathological representations by integrating diverse data sources. Conventional multi-modal integration methods primarily emphasize modality alignment, while paying insufficient attention to retaining the modality-specific intrinsic structures. However, unlike conventional scenarios where multi-modal inputs often share highly overlapping features, histopathology and transcriptomics exhibit pronounced heterogeneity, offering orthogonal yet complementary insights. Histopathology data provides morphological and spatial context, elucidating tissue architecture and cellular topology, whereas transcriptomics data delineates molecular signatures through quantifying gene expression patterns. This inherent disparity introduces a major challenge in aligning these modalities while maintaining modality-specific fidelity. To address these challenges, we present MIRROR, a novel multi-modal representation learning framework designed to foster both modality alignment and retention. MIRROR employs dedicated encoders to extract comprehensive feature representations for each modality, which is further complemented by a modality alignment module to achieve seamless integration between phenotype patterns and molecular profiles. Furthermore, a modality retention module safeguards unique attributes from each modality, while a style clustering module mitigates redundancy and enhances disease-relevant information by modeling and aligning consistent pathological signatures within a clustering space. Extensive evaluations on The Cancer Genome Atlas (TCGA) cohorts for cancer subtyping and survival analysis highlight MIRROR's superior performance, demonstrating its effectiveness in constructing comprehensive oncological feature representations and benefiting the cancer diagnosis. Code is available at https://github.com/TianyiFranklinWang/MIRROR. Jianan Fan, Dingxin Zhang 0001, Dongnan Liu, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 7 |
| 2025 | CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationabstractContrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks. However, the effectiveness of CLIP heavily relies on a substantial corpus of pre-training data, resulting in notable consumption of computational resources. Although knowledge distillation has been widely applied in single modality models, how to efficiently expand knowledge distillation to vision-language foundation models with extensive data remains relatively unexplored. In this paper, we introduce CLIP-CID, a novel distillation mechanism that effectively transfers knowledge from a large vision-language foundation model to a smaller model. We initially propose a simple but efficient image semantic balance method to reduce transfer learning bias and improve distillation efficiency. This method filters out 43.7% of image-text pairs from the LAION400M while maintaining superior performance. After that, we leverage cluster-instance discrimination to facilitate knowledge transfer from the teacher model to the student model, thereby empowering the student model to acquire a holistic semantic comprehension of the pre-training data. Experimental results demonstrate that CLIP-CID achieves state-of-the-art performance on various downstream tasks including linear probe and zero-shot classification. Kaicheng Yang 0002, Tiancheng Gu, Xiang An, Haiqiang Jiang, Xiangzi Dai, Ziyong Feng, Tom Weidong Cai, Jiankang Deng |
AAAI | 7 |
| 2025 | Multi-Scale Visual Prompting for Robust Visual Question Answering in Medical ImagingabstractMedical imaging inherently exhibits multi-scale characteristics, encompassing both global anatomical structures and localized pathological details. However, most existing multimodal large language models (MLLMs) for medical visual question answering (VQA) rely mainly on global features, limiting fine-grained reasoning across spatial levels. To address this, we propose MSFormer (Multi-Scale Transformer), a vision-language architecture that dynamically integrates hierarchical image features across multiple scales. MSFormer extracts multi-resolution embeddings via a vision backbone and refines them using a Multi-Scale Positional Embedding (MSPE) module to maintain spatial alignment. Its core Multi-Scale Grouped Attention (MSGA) mechanism enables learnable queries to jointly attend to features from different scales, adaptively focusing on context relevant to each question. Through contrastive pretraining, instruction tuning, and fine-tuning, MSFormer effectively aligns multi-scale visual and textual representations, substantially improving both open- and closed-ended medical VQA. Extensive experiments show that MSFormer consistently surpasses prior state-of-the-art models, underscoring the value of scale-aware visual prompting for enhanced interpretability and clinical reasoning in multimodal medical AI. Dongang Wang, Michael Barnett 0006, Dingxuan Zhou, Tom Weidong Cai, Chenyu Wang 0001 |
BIBM | 6 |
| 2025 | ScSAM: Debiasing Morphology and Distributional Variability in Subcellular Semantic SegmentationabstractThe significant morphological and distributional variability among subcellular components poses a long-standing challenge for learning-based organelle segmentation models, significantly increasing the risk of biased feature learning. Existing methods often rely on single mapping relationships, overlooking feature diversity and thereby inducing biased training. Although the Segment Anything Model (SAM) provides rich feature representations, its application to subcellular scenarios is hindered by two key challenges: (1) The variability in subcellular morphology and distribution creates gaps in the label space, leading the model to learn spurious or biased features. (2) SAM focuses on global contextual understanding and often ignores fine-grained spatial details, making it challenging to capture subtle structural alterations and cope with skewed data distributions. To address these challenges, we introduce ScSAM, a method that enhances feature robustness by fusing pre-trained SAM with Masked Autoencoder (MAE)-guided cellular prior knowledge to alleviate training bias from data imbalance. Specifically, we design a feature alignment and fusion module to align pre-trained embeddings to the same feature space and efficiently combine different representations. Moreover, we present a cosine similarity matrix-based class prompt encoder to activate class-specific features to recognize subcellular categories. Extensive experiments on diverse subcellular image datasets demonstrate that ScSAM outperforms state-of-the-art methods. Jianan Fan, Dongnan Liu, Hang Chang, Gerald J. Shami, Filip Braet, Tom Weidong Cai |
ECAI | 7 |
| 2025 | VRM: Knowledge Distillation via Virtual Relation Matching
Tom Weidong Cai, Chao Ma 0004 |
ICCV | 3 |
| 2025 | HFBRI-MAE: Handcrafted Feature Based Rotation-Invariant Masked Autoencoder for 3D Point Cloud AnalysisabstractSelf-supervised learning (SSL) has demonstrated remarkable success in 3D point cloud analysis, particularly through masked autoencoders (MAEs). However, existing MAE-based methods lack rotation invariance, leading to significant performance degradation when processing arbitrarily rotated point clouds in real-world scenarios. To address this limitation, we introduce Handcrafted Feature-Based Rotation-Invariant Masked Autoencoder (HFBRI-MAE), a novel framework that refines the MAE design with rotation-invariant handcrafted features to ensure stable feature learning across different orientations. By leveraging both rotation-invariant local and global features for token embedding and position embedding, HFBRI-MAE effectively eliminates rotational dependencies while preserving rich geometric structures. Additionally, we redefine the reconstruction target to a canonically aligned version of the input, mitigating rotational ambiguities. Extensive experiments on ModelNet40, ScanObjectNN, and ShapeNetPart demonstrate that HFBRI-MAE consistently outperforms existing methods in object classification, segmentation, and few-shot learning, highlighting its robustness and strong generalization ability in real-world 3D applications. Xuanhua Yin, Dingxin Zhang 0001, Jianhui Yu, Tom Weidong Cai |
IJCNN | 4 |
| 2025 | CA-W3D: Leveraging Context-Aware Knowledge for Weakly Supervised Monocular 3D DetectionabstractWeakly supervised monocular 3D detection, while less annotation-intensive, often struggles to capture the global context required for reliable 3D reasoning. Conventional label-efficient methods focus on object-centric features, neglecting contextual semantic relationships that are critical in complex scenes. In this work, we propose a Context-Aware Weak Supervision for Monocular 3D object detection, namely CA-W3D, to address this limitation in a two-stage training paradigm. Specifically, we first introduce a pre-training stage employing Region-wise Object Contrastive Matching (ROCM), which aligns regional object embeddings derived from a trainable monocular 3D encoder and a frozen open-vocabulary 2D visual grounding model. This alignment encourages the monocular encoder to discriminate scene-specific attributes and acquire richer contextual knowledge. In the second stage, we incorporate a pseudo-label training process with a Dual-to-One Distillation (D2OD) mechanism, which effectively transfers contextual priors into the monocular encoder while preserving spatial fidelity and maintaining computational efficiency during inference. Extensive experiments conducted on the public KITTI benchmark demonstrate the effectiveness of our approach, surpassing the SoTA method over all metrics, highlighting the importance of contextual-aware knowledge in weakly-supervised monocular 3D detection. For implementation details: CAW3D Chupeng Liu, Runkai Zhao, Tom Weidong Cai |
IROS | 3 |
| 2025 | Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMsabstractThe Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key limitations: (1) text token truncation, (2) isolated image-text encoding, and (3) deficient compositionality due to bag-of-words behavior. While recent Multimodal Large Language Models (MLLMs) have demonstrated significant advances in generalized vision-language understanding, their potential for learning transferable multimodal representations remains underexplored. In this work, we present UniME (Universal Multimodal Embedding), a novel two-stage framework that leverages MLLMs to learn discriminative representations for diverse downstream tasks. In the first stage, we perform textual discriminative knowledge distillation from a powerful LLM-based teacher model to enhance the embedding capability of the MLLM's language component. In the second stage, we introduce hard negative enhanced instruction tuning to further advance discriminative representation learning. Specifically, we initially mitigate false negative contamination and then sample multiple hard negatives per instance within each batch, forcing the model to focus on challenging samples. This approach not only improves discriminative power but also enhances instruction-following ability in downstream tasks. We conduct extensive experiments on the MMEB benchmark and multiple retrieval tasks, including short & long caption retrieval and compositional retrieval. Results demonstrate that UniME achieves consistent performance improvement across all tasks, exhibiting superior discriminative and compositional capabilities. The code will be released in https://garygutc.github.io/UniME. Tiancheng Gu, Kaicheng Yang 0002, Ziyong Feng, Yanzhao Zhang, Dingkun Long, Yingda Chen, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 8 |
| 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation ParadigmabstractAfter pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn. Tiancheng Gu, Kaicheng Yang 0002, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 8 |
| 2025 | ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent MotionabstractModern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance videos that harmonize with both musical rhythm and user-defined choreography styles, limiting their applicability in real-world creative contexts. To address this gap, we introduce ChoreoMuse, a diffusion-based framework that uses SMPL format parameters and their variation version as intermediaries between music and video generation, thereby overcoming the usual constraints imposed by video resolution. Critically, ChoreoMuse supports style-controllable, high-fidelity dance video generation across diverse musical genres and individual dancer characteristics, including the flexibility to handle any reference individual at any resolution. Our method employs a novel music encoder MotionTune to capture motion cues from audio, ensuring that the generated choreography closely follows the beat and expressive qualities of the input music. To quantitatively evaluate how well the generated dances match both musical and choreographic styles, we introduce two new metrics that measure alignment with the intended stylistic cues. Extensive experiments confirm that ChoreoMuse achieves state-of-the-art performance across multiple dimensions, including video quality, beat alignment, dance diversity, and style adherence, demonstrating its potential as a robust solution for a wide range of creative applications. Video results can be found on our project page: https://choreomuse.github.io. Xuanchen Wang, Heng Wang 0007, Tom Weidong Cai |
ACM Multimedia | 3 |
| 2025 | KeyRegionPose: Region-Aware Feature Interaction and Multi-Scale Token Pruning for Efficient Human Pose EstimationabstractThe primary challenge in deploying Human Pose Estimation (HPE) methods in real-world applications lies in balancing computational speed, model compactness, and prediction accuracy. Existing methods achieve strong performance in one or two aspects, but usually at the expense of the remaining one. To overcome this trade-off, we propose KeyRegionPose, a novel framework that achieves high accuracy while reducing model size and computational cost. Central to our design is the Region Focus Mechanism, which enables the model to concentrate on keypoint-relevant regions rather than the entire image. During training, we generate intermediate keypoint proposals to estimate keypoint-specific areas, from which the model learns region-focused features and refines predictions. To ensure accurate keypoint localization and enhance final pose estimation performance, we introduce a Cross-Representation Consistency Loss (CRC Loss) that enforces alignment between the predicted heatmaps and the regressed keypoint coordinates. Additionally, we propose Progressive Multi-Scale Token Pruning (PMTP), a strategy that prunes irrelevant tokens across multiple feature scales to accelerate inference. KeyRegionPose achieves 76.0 AP on the COCO validation set and 75.4 AP on the test-dev set, with only 20.0 million parameters and 8.6 GFLOPs—representing a 27.3% reduction in parameter count, 21.8% decrease in GFLOPs, and a competitive result (+0.2%) over state-of-the-art lightweight HPE models. Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai |
MMAsia | 4 |
| 2025 | ORID: Organ-Regional Information Driven Framework for Radiology Report GenerationabstractThe objective of Radiology Report Generation (RRG) is to automatically generate coherent textual analyses of diseases based on radiological images, thereby alleviating the workload of radiologists. Current AI-based methods for RRG primarily focus on modifications to the encoder-decoder model architecture. To advance these approaches, this paper introduces an Organ-Regional Information Driven (ORID) framework which can effectively integrate multi-modal information and reduce the influence of noise from unrelated organs. Specifically, based on the LLaVA-Med, we first construct an RRG-related instruction dataset to improve organ-regional diagnosis description ability and get the LLaVA-Med-RRG. After that, we propose an organ-based cross-modal fusion module to effectively combine the information from the organ-regional diagnosis description and radiology image. To further reduce the influence of noise from unrelated organs on the radiology report generation, we introduce an organ importance coefficient analysis module, which leverages Graph Neural Network (GNN) to examine the interconnections of the cross-modal information of each organ region. Extensive experiments and comparisons with state-of-the-art methods across various evaluation metrics demonstrate the superior performance of our proposed method. Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai |
WACV | 6 |
| 2025 | AMNCutter: Affinity-Attention-Guided Multi-View Normalized Cutter for Unsupervised Surgical Instrument SegmentationabstractSurgical instrument segmentation (SIS) is pivotal for robotic-assisted minimally invasive surgery, assisting surgeons by identifying surgical instruments in endoscopic video frames. Recent unsupervised surgical instrument segmentation (USIS) methods primarily rely on pseudo-labels derived from low-level features such as color and optical flow, but these methods show limited effective-ness and generalizability in complex and unseen endo-scopic scenarios. In this work, we propose a label-free unsupervised model featuring a novel module named Multi-View Normalized Cutter (m-NCutter). Different from previous USIS works, our model is trained using a graph-cutting loss function that leverages patch affini-ties for supervision, eliminating the need for pseudo-labels. The framework adaptively determines which affini-ties from which levels should be prioritized. Therefore, the low- and high-level features and their affinities are effectively integrated to train a label-free unsupervised model, showing superior effectiveness and generalization abil-ity. We conduct comprehensive experiments across mul-tiple SIS datasets to validate our approach's state-of-the-art (SOTA) performance, robustness, and exceptional potential as a pre-trained model. Our code is released at https://github.com/MingyuShengSMYIAMNCutter. Mingyu Sheng, Jianan Fan, Dongnan Liu, Ron Kikinis, Tom Weidong Cai |
WACV | 5 |
| 2025 | Dance any Beat: Blending Beats with Visuals in Dance Video GenerationabstractGenerating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces their real-world applicability. These methods also require precise keypoint annotations, complicating data collection and limiting the use of self-collected video datasets. To overcome these challenges, we introduce a novel task: generating dance videos directly from images of individuals guided by music. This task enables the dance generation of specific individuals without requiring keypoint annotations, making it more versatile and applicable to various situations. Our solution, the Dance Any Beat Diffusion model (DabFusion), utilizes a reference image and a music piece to generate dance videos featuring various dance types and choreographies. The music is analyzed by our specially designed music encoder, which identifies essential features including dance style, movement, and rhythm. DabFusion excels in generating dance videos not only for individuals in the training dataset but also for any previously unseen person. This versatility stems from its approach of generating latent optical flow, which contains all necessary motion information to animate any person in the image. We evaluate DabFusion's performance using the AIST + + dataset, focusing on video quality, audio-video synchronization, and motion-music alignment. We propose a 2D Motion-Music Alignment Score (2D-MM Align), which builds on the Beat Alignment Score to more effectively evaluate motion-music alignment for this new task. Experiments show that our DabFusion establishes a solid baseline for this innovative task. Video results can be found on our project page: https://DabFusion.github.io. Xuanchen Wang, Heng Wang 0007, Dongnan Liu, Tom Weidong Cai |
WACV | 4 |
| 2025 | TractGraphFormer: Anatomically informed hybrid graph CNN-transformer network for interpretable sex and age prediction from diffusion MRI tractography
Yuqian Chen, Fan Zhang 0013, Leo R. Zekelman, Suheyla Cetin Karayumak, Tengfei Xue, Chaoyi Zhang, Yang Song 0001, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 12 |
| 2025 | On Structuring Hyperspherical Manifold for Probing Novel Biomedical EntitiesabstractThe insufficient high-throughput modeling capability for high-dimensional, multiscale, and nonlinear real-world observations and measurements stands as one of the major impediments for modern science advancements. In this regard, machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven disposition. With the ever-increasing stream of research data collection, it would be appealing to automate the exploration of patterns and insights from observational data for discovering novel classes of phenotypes and entities. However, in the discipline of biomedical investigation, the cumulative data is intrinsically subjected to non-i.i.d. distribution and severe biases amongst different clusters, inducing disorganization and ambiguity in the learned representation space. To contend with the inherent challenges, in this paper, we present a geometry- constrained probabilistic modeling treatment on hyperspherical manifolds. It firstly parameterizes the approximated posterior of instance-wise embedding as a marginal von MisesFisher distribution to account for the interference of distributional latent shift, and thereafter incorporates a suite of critical inductive biases to organically shape the layout of tailored embedding space. Together, these advancements offer a systematic solution to regularize the uncontrollable risk for unseen class learning and prospecting. Furthermore, we propose a spectral graph-theoretic method to efficiently estimate the number of potential novel classes and endow the prediction with adorable taxonomy adaptability. Through extensive experiments under various settings, we demonstrate the effectiveness and general applicability of the proposed methods in recognizing and structurally phenotyping novel visual concepts. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Contrastive Neuron Pruning for Backdoor DefenseabstractRecent studies have revealed that deep neural networks (DNNs) are susceptible to backdoor attacks, in which attackers insert a pre-defined backdoor into a DNN model by poisoning a few training samples. A small subset of neurons in DNN is responsible for activating this backdoor and pruning these backdoor-associated neurons has been shown to mitigate the impact of such attacks. Current neuron pruning techniques often face challenges in accurately identifying these critical neurons, and they typically depend on the availability of labeled clean data, which is not always feasible. To address these challenges, we propose a novel defense strategy called Contrastive Neuron Pruning (CNP). This approach is based on the observation that poisoned samples tend to cluster together and are distinguishable from benign samples in the feature space of a backdoored model. Given a backdoored model, we initially apply a reversed trigger to benign samples, generating multiple positive (benign-benign) and negative (benign-poisoned) feature pairs from the backdoored model. We then employ contrastive learning on these pairs to improve the separation between benign and poisoned features. Subsequently, we identify and prune neurons in the Batch Normalization layers that show significant response differences to the generated pairs. By removing these backdoor-associated neurons, CNP effectively defends against backdoor attacks while requiring the pruning of only about 1% of the total neurons. Comprehensive experiments conducted on various benchmarks validate the efficacy of CNP, demonstrating its robustness and effectiveness in mitigating backdoor attacks compared to existing methods. Benteng Ma, Dongnan Liu, Yanning Zhang 0001, Tom Weidong Cai, Yong Xia 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation ModelsabstractBuilding artificial intelligence (AI) systems on top of a set of foundation models (FMs) is becoming a new paradigm in AI research. Their representative and generative abilities learnt from vast amounts of data can be easily adapted and transferred to a wide range of downstream tasks without extra training from scratch. However, leveraging FMs in cross-modal generation remains under-researched when audio modality is involved. On the other hand, automatically generating semantically-relevant sound from visual input is an important problem in cross-modal generation studies. To solve this vision-to-audio (V2A) generation problem, existing methods tend to design and build complex systems from scratch using modestly sized datasets. In this paper, we propose a lightweight solution to this problem by leveraging foundation models, specifically CLIP, CLAP, and AudioLDM. We first investigate the domain gap between the latent space of the visual CLIP and the auditory CLAP models. Then we propose a simple yet effective mapper mechanism (V2A-Mapper) to bridge the domain gap by translating the visual input between CLIP and CLAP spaces. Conditioned on the translated CLAP embedding, pretrained audio generative FM AudioLDM is adopted to produce high-fidelity and visually-aligned sound. Compared to previous approaches, our method only requires a quick training of the V2A-Mapper. We further analyze and conduct extensive experiments on the choice of the V2A-Mapper and show that a generative mapper is better at fidelity and variability (FD) while a regression mapper is slightly better at relevance (CS). Both objective and subjective evaluation on two V2A datasets demonstrate the superiority of our proposed method compared to current state-of-the-art approaches - trained with 86% fewer parameters but achieving 53% and 19% improvement in FD and CS, respectively. Supplementary materials such as audio samples are provided at our demo website: https://v2a-mapper.github.io/. Heng Wang 0007, Santiago Pascual, Richard Cartwright, Tom Weidong Cai |
AAAI | 5 |
| 2024 | PaintHuman: Towards High-Fidelity Text-to-3D Human Texturing via Denoised Score DistillationabstractRecent advances in zero-shot text-to-3D human generation, which employ the human model prior (e.g., SMPL) or Score Distillation Sampling (SDS) with pre-trained text-to-image diffusion models, have been groundbreaking. However, SDS may provide inaccurate gradient directions under the weak diffusion guidance, as it tends to produce over-smoothed results and generate body textures that are inconsistent with the detailed mesh geometry. Therefore, directly leveraging existing strategies for high-fidelity text-to-3D human texturing is challenging. In this work, we propose a model called PaintHuman to addresses the challenges from two perspectives. We first propose a novel score function, Denoised Score Distillation (DSD), which directly modifies the SDS by introducing negative gradient components to iteratively correct the gradient direction and generate high-quality textures. In addition, we use the depth map as a geometric guide to ensure that the texture is semantically aligned to human mesh surfaces. To guarantee the quality of rendered results, we employ geometry-aware networks to predict surface materials and render realistic human textures. Extensive experiments, benchmarked against state-of-the-art (SoTA) methods, validate the efficacy of our approach.Project page: https://painthuman.github.io/. Jianhui Yu, Liming Jiang 0001, Chen Change Loy, Tom Weidong Cai, Wayne Wu |
AAAI | 5 |
| 2024 | Enhancing Robustness to Noise Corruption for Point Cloud Recognition via Spatial Sorting and Set-Mixing Aggregation Module
Dingxin Zhang 0001, Jianhui Yu, Tengfei Xue, Chaoyi Zhang, Dongnan Liu, Tom Weidong Cai |
ACCV (9) | 6 |
| 2024 | Unsupervised Domain Adaptation for Tubular Structure Segmentation Across Different Anatomical Sources
Yuxiang An, Dongnan Liu, Tom Weidong Cai |
BMVC | 3 |
| 2024 | Rethinking Domain Adaptive Optic Disc and Cup Segmentation in Fundus Image through Dynamic Diffusion Flow
Canran Li, Dongnan Liu, Tom Weidong Cai |
BMVC | 3 |
| 2024 | Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic ModelingabstractMachine learning holds tremendous promise for trans-forming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection, it would be appealing to autonomously explore patterns and insights from obser-vational data for discovering novel classes of phenotypes and concepts. However, in the biomedical domain, there are several challenges inherently presented in the cumu-lated data which hamper the progress of novel class dis-covery. The non-i.i.d. data distribution accompanied by the severe imbalance among different groups of classes es-sentially leads to ambiguous and biased semantic represen-tations. In this work, we present a geometry-constrained probabilistic modeling treatment to resolve the identified is-sues. First, we propose to parameterize the approximated posterior of instance embedding as a marginal von Mises-Fisher distribution to account for the interference of distri-butional latent bias. Then, we incorporate a suite of critical geometric properties to impose proper constraints on the layout of constructed embedding space, which in turn min-imizes the uncontrollable risk for unknown class learning and structuring. Furthermore, a spectral graph-theoretic method is devised to estimate the number of potential novel classes. It inherits two intriguing merits compared to exis-tent approaches, namely high computational efficiency and flexibility for taxonomy-adaptive estimation. Extensive ex-periments across various biomedical scenarios substantiate the effectiveness and general applicability of our method. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
CVPR | 6 |
| 2024 | Device-Wise Federated Network PruningabstractNeural network pruning, particularly channel pruning, is a widely used technique for compressing deep learning models to enable their deployment on edge devices with limited resources. Typically, redundant weights or structures are removed to achieve the target resource budget. Although data-driven pruning approaches have proven to be more effective, they cannot be directly applied to federated learning (FL), which has emerged as a popular technique in edge computing applications, because of distributed and confidential datasets. In response to this challenge, we design a new network pruning method for FL. We propose device-wise sub-networks for each device, assuming that the data distribution is similar within each device. These sub-networks are generated through sub-network embeddings and a hypernetwork. To further minimize memory usage and communication costs, we permanently prune the full model to remove weights that are not useful for all devices. During the FL process, we simultaneously train the device-wise sub-networks and the base sub-network to facilitate the pruning process. We then finetune the pruned model with device-wise sub-networks to regain performance. Moreover, we provided the theoretical guarantee of convergence for our method. Our method achieves better performance and resource trade-off than other well-established network pruning baselines, as demonstrated through extensive experiments on CIFAR-10, CIFAR-100, and TinyImageNet. Shangqian Gao, Junyi Li 0002, Yanfu Zhang, Tom Weidong Cai, Heng Huang 0001 |
CVPR | 5 |
| 2024 | Revisiting Adaptive Cellular Recognition Under Domain Shifts: A Contextual Correspondence View
Jianan Fan, Dongnan Liu, Canran Li, Hang Chang, Heng Huang 0001, Filip Braet, Tom Weidong Cai |
ECCV (73) | 8 |
| 2024 | Controllable Contextualized Image Captioning: Directing the Visual Narrative Through User-Defined Highlights
Shunqi Mao, Chaoyi Zhang, Hwanjun Song, Igor Shalyminov, Tom Weidong Cai |
ECCV (50) | 6 |
| 2024 | RWKV-CLIP: A Robust Vision-Language Representation LearnerabstractContrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web.This paper further explores CLIP from the perspectives of data and model architecture.To mitigate the impact of the noise data and enhance the quality of large-scale image-text data crawled from the internet, we introduce a diverse description generation framework that can leverage Large Language Models (LLMs) to combine and refine information from web-based image-text pairs, synthetic captions, and detection tags.Additionally, we propose RWKV-CLIP, the first RWKV-driven vision-language representation learning model that combines the effective parallel training of transformers with the efficient inference of RNNs.Extensive experiments across different model scales and pre-training datasets demonstrate that RWKV-CLIP is a robust vision-language representation learner and it achieves state-of-the-art performance across multiple downstream tasks, including linear probing, zero-shot classification, and zero-shot image-text retrieval.To facilitate future research, the code and pre-trained models are released at https: //github.com/deepglint/RWKV-CLIP. Tiancheng Gu, Kaicheng Yang 0002, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
EMNLP | 6 |
| 2024 | Enhancing Advanced Visual Reasoning Ability of Large Language ModelsabstractRecent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability. Traditional Vision-Language models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Conversely, Large Language Models (LLMs) demonstrate robust text reasoning capabilities; however, they lack visual acuity. To bridge this gap, we propose Complex Visual Reasoning Large Language Models (CVR-LLM), capitalizing on VLMs’ visual perception proficiency and LLMs’ extensive reasoning capability. Unlike recent multimodal large language models (MLLMs) that require a projection layer, our approach transforms images into detailed, context-aware descriptions using an iterative self-refinement loop and leverages LLMs’ text knowledge for accurate predictions without extra training. We also introduce a novel multi-modal in-context learning (ICL) methodology to enhance LLMs’ contextual understanding and reasoning. Additionally, we introduce Chain-of-Comparison (CoC), a step-by-step comparison technique enabling contrasting various aspects of predictions. Our CVR-LLM presents the first comprehensive study across a wide array of complex visual reasoning tasks and achieves SOTA performance among all. Dongnan Liu, Chaoyi Zhang, Heng Wang 0007, Tengfei Xue, Tom Weidong Cai |
EMNLP | 6 |
| 2024 | Enhancing Angular Resolution via Directionality Encoding and Geometric Constraints in Brain Diffusion Tensor Imaging
Zihao Tang 0002, Mariano Cabezas, Xinyi Wang 0015, Arkiev D'Souza, Michael Barnett 0006, Fernando Calamante, Tom Weidong Cai, Chenyu Wang 0001 |
ICONIP (4) | 8 |
| 2024 | Learning to Synthesize Graphics Programs for Geometric Artworks
Qi Bing, Chaoyi Zhang, Tom Weidong Cai |
ICPR (18) | 3 |
| 2024 | Advancements in 3D Lane Detection Using LiDAR Point Clouds: From Data Collection to Model DevelopmentabstractAdvanced Driver-Assistance Systems (ADAS) have successfully integrated learning-based techniques into vehicle perception and decision-making. However, their application in 3D lane detection for effective driving environment perception is hindered by the lack of comprehensive LiDAR datasets. The sparse nature of LiDAR point cloud data prevents an efficient manual annotation process. To solve this problem, we present LiSV-3DLane, a large-scale 3D lane dataset that comprises 20k frames of surround-view LiDAR point clouds with enriched semantic annotation. Unlike existing datasets confined to a frontal perspective, LiSV-3DLane provides a full 360-degree spatial panorama around the ego vehicle, capturing complex lane patterns in both urban and highway environments. We leverage the geometric traits of lane lines and the intrinsic spatial attributes of LiDAR data to design a simple yet effective automatic annotation pipeline for generating finer lane labels. To propel future research, we propose a novel LiDAR-based 3D lane detection model, LiLaDet, incorporating the spatial geometry learning of the LiDAR point cloud into Bird’s Eye View (BEV) based lane identification. Experimental results indicate that LiLaDet outperforms existing camera- and LiDAR-based approaches in the 3D lane detection task on the K-Lane dataset and our LiSV-3DLane. The project code will be available at https://github.com/RunkaiZhao/LiLaDet. Runkai Zhao, Yuwen Heng, Heng Wang 0007, Yuanda Gao, Shilei Liu, Changhao Yao, Tom Weidong Cai |
ICRA | 8 |
| 2024 | Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis
Dongang Wang, Lynette Masters, Michael Barnett 0006, Tom Weidong Cai, Chenyu Wang 0001 |
MICCAI (12) | 6 |
| 2024 | Cross-View Consistency Regularisation for Knowledge DistillationabstractKnowledge distillation (KD) is an established paradigm for transferring privileged knowledge from a cumbersome model to a more lightweight and efficient one. In recent years, logit-based KD methods are quickly catching up in performance with their feature-based counterparts. However, previous research has pointed out that logit-based methods are still fundamentally limited by two major issues in their training process, namely overconfident teacher and confirmation bias. Inspired by the success of cross-view learning in fields such as semi-supervised learning, in this work we introduce within-view and cross-view regularisations to standard logit-based distillation frameworks to combat the above cruxes. We also perform confidence-based soft label mining to improve the quality of distilling signals from the teacher, which further mitigates the confirmation bias problem. Despite its apparent simplicity, the proposed Consistency-Regularisation-based Logit Distillation (CRLD) significantly boosts student learning, setting new state-of-the-art results on the standard CIFAR-100, Tiny-ImageNet, and ImageNet datasets across a diversity of teacher and student architectures, whilst introducing no extra network parameters. Orthogonal to on-going logit-based distillation research, our method enjoys excellent generalisation properties and, without bells and whistles, boosts the performance of various existing approaches by considerable margins. Dongnan Liu, Tom Weidong Cai, Chao Ma 0004 |
ACM Multimedia | 3 |
| 2024 | LaneCMKT: Boosting Monocular 3D Lane Detection with Cross-Modal Knowledge TransferabstractDetecting 3D lane lines from monocular images is garnering increasing attention in the Autonomous Driving (AD) area due to its cost-effective edge. However, current monocular image models capture road scenes lacking 3D spatial awareness, which is error-prone to adverse circumstance changes. In this work, we design a novel cross-modal knowledge transfer scheme, namely LaneCMKT, to address this challenge by transferring 3D geometric cues learned from a pre-trained LiDAR model to the image model. Performing on the unified Bird's-Eye-View (BEV) grid, our monocular image model acts as a student network and benefits from the spatial guidance of the 3D LiDAR teacher model over the intermediate feature space. Since LiDAR points and image pixels are intrinsically two different modalities, to facilitate such heterogeneous feature transfer learning at matching levels, we propose a dual-path knowledge transfer mechanism. We divide the feature space into shallow and deep paths where the image student model is prompted to focus on lane-favored geometric cues from the LiDAR teacher model. We conduct extensive experiments and thorough analysis on the large-scale public benchmark OpenLane. Our model achieves notable improvements over the image baseline by 5.3% and the current BEV-driven SoTA method by 2.7% in the F1 score, without introducing any extra computational overhead. We also observe that the 3D abilities grabbed from the teacher model are critical for dealing with complex spatial lane properties from a 2D perspective. Runkai Zhao, Heng Wang 0007, Tom Weidong Cai |
ACM Multimedia | 3 |
| 2024 | Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
Dongnan Liu, Heng Wang 0007, Chaoyi Zhang, Tom Weidong Cai |
MMAsia | 5 |
| 2024 | Fibre Population-guided Pre-training for 3D Spatial Super-Resolution on Multimodal Brain Diffusion MR Imaging
Zihao Tang 0002, Xinyi Wang 0015, Mariano Cabezas, Arkiev D'Souza, Michael Barnett 0006, Fernando Calamante, Tom Weidong Cai, Chenyu Wang 0001 |
MMAsia | 7 |
| 2024 | Complex Organ Mask Guided Radiology Report GenerationabstractThe goal of automatic report generation is to generate a clinically accurate and coherent phrase from a single given X-ray image, which could alleviate the workload of traditional radiology reporting. However, in a real-world scenario, radiologists frequently face the challenge of producing extensive reports derived from numerous medical images, thereby medical report generation from multi-image perspective is needed. In this paper, we propose the Complex Organ Mask Guided (termed as COMG) report generation model, which incorporates masks from multiple organs (e.g., bones, lungs, heart, and mediastinum), to pro-vide more detailed information and guide the model’s attention to these crucial body regions. Specifically, we leverage prior knowledge of the disease corresponding to each organ in the fusion process to enhance the disease identification phase during the report generation process. Additionally, cosine similarity loss is introduced as target function to ensure the convergence of cross-modal consistency and facilitate model optimization. Experimental results on two public datasets show that COMG achieves a 11.4% and 9.7% improvement in terms of BLEU@4 scores over the SOTA model KiUT on IU-Xray and MIMIC, respectively. The code is publicly available at https://github.com/GaryGuTC/COMG_model. Tiancheng Gu, Dongnan Liu, Tom Weidong Cai |
WACV | 4 |
| 2024 | Alleviating Foreground Sparsity for Semi-Supervised Monocular 3D Object DetectionabstractMonocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image. In this paper, we strive to boost currently underperforming monocular 3D object detectors by leveraging an abundance of unlabelled data via semi-supervised learning. Our proposed ODM3D framework entails cross-modal knowledge distillation at various levels to inject LiDAR-domain knowledge into a monocular detector during training. By identifying foreground sparsity as a main culprit behind existing methods’ suboptimal training, we exploit the precise localisation information embedded in LiDAR points to enable more foreground-attentive and efficient distillation via the proposed BEV occupancy guidance mask, leading to notably improved knowledge transfer and M3OD performance. Besides, motivated by insights into why existing cross-modal GT-sampling techniques fail on our task at hand, we further design a novel cross-modal object-wise data augmentation strategy for effective RGB-LiDAR joint learning. Our method ranks 1stin both KITTI validation and test benchmarks, significantly surpassing all existing monocular methods, supervised or semi-supervised, on both BEV and 3D detection metrics. Code will be released at https://github.com/arcaninez/odm3d. Dongnan Liu, Chao Ma 0004, Tom Weidong Cai |
WACV | 4 |
| 2024 | Improving multiple sclerosis lesion segmentation across clinical sites: A federated learning approach with noise-resilient trainingabstractAccurately measuring the evolution of Multiple Sclerosis (MS) with magnetic resonance imaging (MRI) critically informs understanding of disease progression and helps to direct therapeutic strategy. Deep learning models have shown promise for automatically segmenting MS lesions, but the scarcity of accurately annotated data hinders progress in this area. Obtaining sufficient data from a single clinical site is challenging and does not address the heterogeneous need for model robustness. Conversely, the collection of data from multiple sites introduces data privacy concerns and potential label noise due to varying annotation standards. To address this dilemma, we explore the use of the federated learning framework while considering label noise. Our approach enables collaboration among multiple clinical sites without compromising data privacy under a federated learning paradigm that incorporates a noise-robust training strategy based on label correction. Specifically, we introduce a Decoupled Hard Label Correction (DHLC) strategy that considers the imbalanced distribution and fuzzy boundaries of MS lesions, enabling the correction of false annotations based on prediction confidence. We also introduce a Centrally Enhanced Label Correction (CELC) strategy, which leverages the aggregated central model as a correction teacher for all sites, enhancing the reliability of the correction process. Extensive experiments conducted on two multi-site datasets demonstrate the effectiveness and robustness of our proposed methods, indicating their potential for clinical applications in multi-site collaborations to train better deep learning models with lower cost in data collection and annotation. Lei Bai 0001, Dongang Wang, Hengrui Wang, Michael Barnett 0006, Mariano Cabezas, Tom Weidong Cai, Fernando Calamante, Kain Kyle, Dongnan Liu, Linda Ly, Aria Nguyen, Chun-Chien Shieh, Ryan Sullivan, Geng Zhan, Wanli Ouyang, Chenyu Wang 0001 |
Artif. Intell. Medicine | 6 |
| 2024 | Learning to Generalize over Subpartitions for Heterogeneity-Aware Domain Adaptive Nuclei SegmentationabstractAbstract Annotation scarcity and cross-modality/stain data distribution shifts are two major obstacles hindering the application of deep learning models for nuclei analysis, which holds a broad spectrum of potential applications in digital pathology. Recently, unsupervised domain adaptation (UDA) methods have been proposed to mitigate the distributional gap between different imaging modalities for unsupervised nuclei segmentation in histopathology images. However, existing UDA methods are built upon the assumption that data distributions within each domain should be uniform. Based on the over-simplified supposition, they propose to align the histopathology target domain with the source domain integrally, neglecting severe intra-domain discrepancy over subpartitions incurred by mixed cancer types and sampling organs. In this paper, for the first time, we propose to explicitly consider the heterogeneity within the histopathology domain and introduce open compound domain adaptation (OCDA) to resolve the crux. In specific, a two-stage disentanglement framework is proposed to acquire domain-invariant feature representations at both image and instance levels. The holistic design addresses the limitations of existing OCDA approaches which struggle to capture instance-wise variations. Two regularization strategies are specifically devised herein to leverage the rich subpartition-specific characteristics in histopathology images and facilitate subdomain decomposition. Moreover, we propose a dual-branch nucleus shape and structure preserving module to prevent nucleus over-generation and deformation in the synthesized images. Experimental results on both cross-modality and cross-stain scenarios over a broad range of diverse datasets demonstrate the superiority of our method compared with state-of-the-art UDA and OCDA methods. Graphical abstract Jianan Fan, Dongnan Liu, Hang Chang, Tom Weidong Cai |
Int. J. Comput. Vis. | 4 |
| 2024 | TractGeoNet: A geometric deep learning framework for pointwise analysis of tract microstructure to predict language assessment performance
Yuqian Chen, Leo R. Zekelman, Chaoyi Zhang, Tengfei Xue, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
Medical Image Anal. | 9 |
| 2024 | A review of image watermarking for identity protection and verificationabstractAbstract Identity protection is an indispensable feature of any information security system. An identity can exist in the form of digitally written signatures, biometric information, logos, etc. It serves the vital purpose of the owners’ verification and provides them with a safety net against their imposters, so its protection is essential. Numerous security mechanisms are being developed to achieve this goal, and information embedding is prominent among all. It consists of cryptography, steganography, and watermarking; collectively, they are known as data hiding (DH) techniques. In addition to providing insight into various DH techniques, this review prominently covers the image watermarking works that have positively influenced its relevant research area. To that end, one of the main aspects of this study is its inclusive nature in reviewing watermarking techniques, via which itaimsto provide a 360 $$^{\circ }$$ ∘ view of the watermarking technology. The main contributions of this study are summarised below. The proposed study covers more than 100 major watermarking works that have positively influenced the field and continue to do so. This approach makes the discussion effective as it allows us to pivot on the vital watermarking works that have positively influenced the research area instead of just highlighting as many existing methods as possible. Moreover, it also empowers us to provide the readers with an insight into the current research trends, the pros and cons of the state-of-the-art methods, and recommendations for future works. In addition to reviewing the state-of-the-art watermarking works, this study solves the issue of reverse-engineering the main existing watermarking methods. For instance, most recent surveys have focused primarily on reviewing as many watermarking works as possible without probing into the actual working of the techniques. This approach can leave the readership without a vital understanding of implementing or reverse-engineering a watermarking method. This issue is especially prevalent among newcomers to the watermarking field; hence, this study presents the breakdown of the well-known watermarking techniques. A new systematisation of classifying existing watermarking methods is proposed. It classifies watermarking techniques into two phases. The first phase divides watermarking methods into three categories based on the domain employed during watermark embedding. The methods are further classified based on other watermarking attributes in the following phase. Sunpreet Sharma, Ju Jia Zou, Gu Fang 0001, Pancham Shukla, Tom Weidong Cai |
Multim. Tools Appl. | 5 |
| 2024 | Exploiting Structural Consistency of Chest Anatomy for Unsupervised Anomaly Detection in Radiography ImagesabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. Exploiting this structured information could potentially ease the detection of anomalies from radiography images. To this end, we propose a Simple Space-Aware Memory Matrix for In-painting and Detecting anomalies from radiography images (abbreviated as SimSID). We formulate anomaly detection as an image reconstruction task, consisting of a space-aware memory matrix and an in-painting block in the feature space. During the training, SimSID can taxonomize the ingrained anatomical structures into recurrent visual patterns, and in the inference, it can identify anomalies (unseen/modified visual patterns) from the test image. Our SimSID surpasses the state of the arts in unsupervised anomaly detection by +8.0%, +5.0%, and +9.9% AUC scores on ZhangLab, COVIDx, and CheXpert benchmark datasets, respectively. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Rethinking Rotation Invariance with Point Cloud RegistrationabstractRecent investigations on rotation invariance for 3D point clouds have been devoted to devising rotation-invariant feature descriptors or learning canonical spaces where objects are semantically aligned. Examinations of learning frameworks for invariance have seldom been looked into. In this work, we review rotation invariance (RI) in terms of point cloud registration (PCR) and propose an effective framework for rotation invariance learning via three sequential stages, namely rotation-invariant shape encoding, aligned feature integration, and deep feature registration. We first encode shape descriptors constructed with respect to reference frames defined over different scales, e.g., local patches and global topology, to generate rotation-invariant latent shape codes. Within the integration stage, we propose an Aligned Integration Transformer (AIT) to produce a discriminative feature representation by integrating point-wise self- and cross-relations established within the shape codes. Meanwhile, we adopt rigid transformations between reference frames to align the shape codes for feature consistency across different scales. Finally, the deep integrated feature is registered to both rotation-invariant shape codes to maximize their feature similarities, such that rotation invariance of the integrated feature is preserved and shared semantic information is implicitly extracted from shape codes. Experimental results on 3D shape classification, part segmentation, and retrieval tasks prove the feasibility of our framework. Our project page is released at: https://rotation3d.github.io/. Jianhui Yu, Chaoyi Zhang, Tom Weidong Cai |
AAAI | 3 |
| 2023 | PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose RestorationabstractRecent interest in point cloud analysis has led rapid progress in designing deep learning methods for 3D models. However, state-of-the-art models are not robust to rotations, which remains an unknown prior to real applications and harms the model performance. In this work, we introduce a novel Patch-wise Rotation-invariant network (PaRot), which achieves rotation invariance via feature disentanglement and produces consistent predictions for samples with arbitrary rotations. Specifically, we design a siamese training module which disentangles rotation invariance and equivariance from patches defined over different scales, e.g., the local geometry and global shape, via a pair of rotations. However, our disentangled invariant feature loses the intrinsic pose information of each patch. To solve this problem, we propose a rotation-invariant geometric relation to restore the relative pose with equivariant information for patches defined over different scales. Utilising the pose information, we propose a hierarchical module which implements intra-scale and inter-scale feature aggregation for 3D shape learning. Moreover, we introduce a pose-aware feature propagation process with the rotation-invariant relative pose information embedded. Experiments show that our disentanglement module extracts high-quality rotation-robust features and the proposed lightweight model achieves competitive results in rotated 3D object classification and part segmentation tasks. Dingxin Zhang 0001, Jianhui Yu, Chaoyi Zhang, Tom Weidong Cai |
AAAI | 4 |
| 2023 | SQUID: Deep Feature In-Painting for Unsupervised Anomaly DetectionabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies from radiography images (abbreviated as SQUID). We show that SQUID can taxonomize the ingrained anatomical structures into recurrent patterns; and in the inference, it can identify anomalies (unseen/modified patterns) in the image. SQUID surpasses 13 state-of-the-art methods in unsupervised anomaly detection by at least 5 points on two chest X-ray benchmark datasets measured by the Area Under the Curve (AUC). Additionally, we have created a new dataset (DigitAnatomy), which synthesizes the spatial correlation and consistent shape in chest anatomy. We hope DigitAnatomy can prompt the development, evaluation, and interpretability of anomaly detection methods. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
CVPR | 6 |
| 2023 | CelebV-Text: A Large-Scale Facial Text-Video DatasetabstractText-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents Celeb V- Text, a large-scale, di-verse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to- video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts gen-erated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The supe-riority of CelebV- Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV- Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available11Project page: https://celebv-text.github.io. Jianhui Yu, Liming Jiang 0001, Chen Change Loy, Tom Weidong Cai, Wayne Wu |
CVPR | 5 |
| 2023 | Taxonomy Adaptive Cross-Domain Adaptation in Medical Imaging via Optimization Trajectory DistillationabstractThe success of automated medical image analysis depends on large-scale and expert-annotated training sets. Unsupervised domain adaptation (UDA) has been raised as a promising approach to alleviate the burden of labeled data collection. However, they generally operate under the closed-set adaptation setting assuming an identical label set between the source and target domains, which is over-restrictive in clinical practice where new classes commonly exist across datasets due to taxonomic inconsistency. While several methods have been presented to tackle both domain shifts and incoherent label sets, none of them take into account the common characteristics of the two issues and consider the learning dynamics along network training. In this work, we propose optimization trajectory distillation, a unified approach to address the two technical challenges from a new perspective. It exploits the low-rank nature of gradient space and devises a dual-stream distillation algorithm to regularize the learning dynamics of insufficiently annotated domain and classes with the external guidance obtained from reliable sources. Our approach resolves the issue of inadequate navigation along network optimization, which is the major obstacle in the taxonomy adaptive cross-domain adaptation scenario. We evaluate the proposed method extensively on several tasks towards various endpoints with clinical and open-world significance. The results demonstrate its effectiveness and improvements over previous methods. Code is available at https://github.com/camwew/TADA-MI. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
ICCV | 6 |
| 2023 | Unsupervised Domain Adaptation for Neuron Membrane Segmentation based on Structural FeaturesabstractAI-enhanced segmentation of neuronal boundaries in electron microscopy (EM) images is crucial for automatic and accurate neuroinformatics studies. To enhance the limited generalization ability of typical deep learning frameworks for medical image analysis, unsupervised domain adaptation (UDA) methods have been applied. In this work, we propose to improve the performance of UDA methods on cross-domain neuron membrane segmentation in EM images. First, we designed a feature weight module considering the structural features during adaptation. Second, we introduced a structural feature-based super-resolution approach to alleviating the domain gap by adjusting the cross-domain image resolutions. Third, we proposed an orthogonal decomposition module to facilitate the extraction of domain-invariant features. Extensive experiments on two domain adaptive membrane segmentation applications have indicated the effectiveness of our method. Yuxiang An, Dongnan Liu, Tom Weidong Cai |
ICME | 3 |
| 2023 | ASRCD: Adaptive Serial Relation-Based Model for Cognitive Diagnosis
Zhuonan Liang, Dongnan Liu, Caiyun Sun, Tom Weidong Cai, Peng Fu 0003 |
ICONIP (14) | 5 |
| 2023 | Topology Repairing of Disconnected Pulmonary Airways and Vessels: Baselines and a Dataset
Ziqiao Weng, Jiancheng Yang, Dongnan Liu, Tom Weidong Cai |
MICCAI (7) | 4 |
| 2023 | TractCloud: Registration-Free Tractography Parcellation with a Novel Local-Global Streamline Point Cloud Representation
Tengfei Xue, Yuqian Chen, Chaoyi Zhang, Alexandra J. Golby, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
MICCAI (8) | 7 |
| 2023 | PointNeuron: 3D Neuron Reconstruction via Geometry and Topology Learning of Point CloudsabstractDigital neuron reconstruction from 3D microscopy images is an essential technique for investigating brain connectomics and neuron morphology. Existing reconstruction frameworks use convolution-based segmentation networks to partition the neuron from noisy backgrounds before applying the tracing algorithm. The tracing results are sensitive to the raw image quality and segmentation accuracy. In this paper, we propose a novel framework for 3D neuron reconstruction. Our key idea is to use the geometric representation power of the point cloud to better explore the intrinsic structural information of neurons. Our proposed framework adopts one graph convolutional network to predict the neural skeleton points and another one to produce the connectivity of these points. We finally generate the target SWC file through the interpretation of the predicted point coordinates, radius, and connections. Evaluated on the Janelia-Fly dataset from the BigNeuron project, we show that our framework achieves competitive neuron reconstruction performance. Our geometry and topology learning of point clouds could further benefit 3D medical image analysis, such as cardiac surface reconstruction. Our code is available at https://github.com/RunkaiZhao/PointNeuron. Runkai Zhao, Heng Wang 0007, Chaoyi Zhang, Tom Weidong Cai |
WACV | 4 |
| 2023 | Superficial white matter analysis: An efficient point-cloud-based deep learning framework with supervised contrastive learning for consistent tractography parcellation across populations and dMRI acquisitions
Tengfei Xue, Fan Zhang 0013, Chaoyi Zhang, Yuqian Chen, Yang Song 0001, Alexandra J. Golby, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 9 |
| 2023 | Decompose to Adapt: Cross-Domain Object Detection Via Feature DisentanglementabstractRecent advances in unsupervised domain adaptation (UDA) techniques have witnessed great success in cross-domain computer vision tasks, enhancing the generalization ability of data-driven deep learning architectures by bridging the domain distribution gaps. For the UDA-based cross-domain object detection methods, the majority of them alleviate the domain bias by inducing the domain-invariant feature generation via adversarial learning strategy. However, their domain discriminators have limited classification ability due to the unstable adversarial training process. Therefore, the extracted features induced by them cannot be perfectly domain-invariant and still contain domain-private factors, bringing obstacles to further alleviate the cross-domain discrepancy. To tackle this issue, we design a Domain Disentanglement Faster-RCNN (DDF) to eliminate the source-specific information in the features for detection task learning. Our DDF method facilitates the feature disentanglement at the global and local stages, with a Global Triplet Disentanglement (GTD) module and an Instance Similarity Disentanglement (ISD) module, respectively. By outperforming state-of-the-art methods on four benchmark UDA object detection tasks, our DDF method is demonstrated to be effective with wide applicability. Dongnan Liu, Chaoyi Zhang, Yang Song 0001, Heng Huang 0001, Chenyu Wang 0001, Michael Barnett 0006, Tom Weidong Cai |
IEEE Trans. Multim. | 7 |
| 2023 | EV-LFV: Synthesizing Light Field Event Streams from an Event Camera and Multiple RGB CamerasabstractLight field videos captured in RGB frames (RGB-LFV) can provide users with a 6 degree-of-freedom immersive video experience by capturing dense multi-subview video. Despite its potential benefits, the processing of dense multi-subview video is extremely resource-intensive, which currently limits the frame rate of RGB-LFV (i.e., lower than 30 fps) and results in blurred frames when capturing fast motion. To address this issue, we propose leveraging event cameras, which provide high temporal resolution for capturing fast motion. However, the cost of current event camera models makes it prohibitive to use multiple event cameras for RGB-LFV platforms. Therefore, we propose EV-LFV, an event synthesis framework that generates full multi-subview event-based RGB-LFV with only one event camera and multiple traditional RGB cameras. EV-LFV utilizes spatial-angular convolution, ConvLSTM, and Transformer to model RGB-LFV's angular features, temporal features, and long-range dependency, respectively, to effectively synthesize event streams for RGB-LFV. To train EV-LFV, we construct the first event-to-LFV dataset consisting of 200 RGB-LFV sequences with ground-truth event streams. Experimental results demonstrate that EV-LFV outperforms state-of-the-art event synthesis methods for generating event-based RGB-LFV, effectively alleviating motion blur in the reconstructed RGB-LFV. Zhicheng Lu, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Yiran Shen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | LFACon: Introducing Anglewise Attention to No-Reference Quality Assessment in Light Field SpaceabstractLight field imaging can capture both the intensity information and the direction information of light rays. It naturally enables a six-degrees-of-freedom viewing experience and deep user engagement in virtual reality. Compared to 2D image assessment, light field image quality assessment (LFIQA) needs to consider not only the image quality in the spatial domain but also the quality consistency in the angular domain. However, there is a lack of metrics to effectively reflect the angular consistency and thus the angular quality of a light field image (LFI). Furthermore, the existing LFIQA metrics suffer from high computational costs due to the excessive data volume of LFIs. In this paper, we propose a novel concept of "anglewise attention" by introducing a multihead self-attention mechanism to the angular domain of an LFl. This mechanism better reflects the LFI quality. In particular, we propose three new attention kernels, including anglewise self-attention, anglewise grid attention, and anglewise central attention. These attention kernels can realize angular self-attention, extract multiangled features globally or selectively, and reduce the computational cost of feature extraction. By effectively incorporating the proposed kernels, we further propose our light field attentional convolutional neural network (LFACon) as an LFIQA metric. Our experimental results show that the proposed LFACon metric significantly outperforms the state-of-the-art LFIQA metrics. For the majority of distortion types, LFACon attains the best performance with lower complexity and less computational time. Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Unsupervised Domain Adaptive Fundus Image Segmentation with Few Labeled Source Data
Qianbi Yu, Dongnan Liu, Chaoyi Zhang, Xinwen Zhang, Tom Weidong Cai |
BMVC | 5 |
| 2022 | Channel-Position Self-Attention with Query Refinement Skeleton Graph Neural Network in Human Pose EstimationabstractHuman Pose Estimation (HPE) is a long-standing yet challenging task in computer vision. The nature of the problem requires comprehensive global contextual reasoning among joints in different locations. In this work, we explore how to incorporate two popular and effective concepts, self-attention and Graph Neural Network (GNN), to model long-range information in HPE. Three different ways to implement self-attention in 3D feature maps are studied, where the best result is achieved via the channel-position version. Accuracy is further improved by refining the queries via an efficient channel-wise parallel GNN that explicitly models the human joint graphical relationships. We are able to improve prediction accuracy on strong baseline models and achieve state-of-the-art results. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 4 |
| 2022 | Spatiality-guided Transformer for 3D Dense Captioning on Point CloudsabstractDense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance-level label of natural language description on visual appearance and spatial relations for each scene object of interest. To detect and describe objects in a scene, following the spirit of neural machine translation, we propose a transformer-based encoder-decoder architecture, namely SpaCap3D, to transform objects into descriptions, where we especially investigate the relative spatiality of objects in 3D scenes and design a spatiality-guided encoder via a token-to-token spatial relation learning objective and an object-centric decoder for precise and spatiality-enhanced object caption generation. Evaluated on two benchmark datasets, ScanRefer and ReferIt3D, our proposed SpaCap3D outperforms the baseline method Scan2Cap by 4.94% and 9.61% in [email protected], respectively. Our project page with source code and supplementary files is available at https://SpaCap3D.github.io/. Heng Wang 0007, Chaoyi Zhang, Jianhui Yu, Tom Weidong Cai |
IJCAI | 4 |
| 2022 | White Matter Tracts are Point Clouds: Neuropsychological Score Prediction and Critical Region Localization via Geometric Deep Learning
Yuqian Chen, Fan Zhang 0013, Chaoyi Zhang, Tengfei Xue, Leo R. Zekelman, Jianzhong He 0001, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Lauren O'Donnell |
MICCAI (1) | 11 |
| 2022 | Domain Adaptive Nuclei Instance Segmentation and Classification via Category-Aware Feature Alignment and Pseudo-Labelling
Canran Li, Dongnan Liu, Haoran Li 0024, Zheng Zhang 0006, Guangming Lu 0002, Xiaojun Chang, Tom Weidong Cai |
MICCAI (8) | 7 |
| 2022 | TractoFormer: A Novel Fiber-Level Whole Brain Tractography Analysis Framework Using Spectral Embedding and Vision Transformers
Fan Zhang 0013, Tengfei Xue, Tom Weidong Cai, Yogesh Rathi, Carl-Fredrik Westin, Lauren O'Donnell |
MICCAI (1) | 3 |
| 2022 | Deep-learning-based solution for data deficient satellite image segmentation
Henry Wing Fung Yeung, Vera Chung, Grant Moule, Wayne Thompson, Wanli Ouyang, Tom Weidong Cai, Mohammed Bennamoun |
Expert Syst. Appl. | 7 |
| 2022 | Towards bi-directional skip connections in encoder-decoder architectures and beyond
Tiange Xiang, Chaoyi Zhang, Xinyi Wang 0015, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
Medical Image Anal. | 7 |
| 2022 | Learning multi-scale synergic discriminative features for prostate image segmentation
Haozhe Jia, Tom Weidong Cai, Heng Huang 0001, Yong Xia 0001 |
Pattern Recognit. | 2 |
| 2022 | Multiple Sclerosis Lesion Analysis in Brain Magnetic Resonance Images: Techniques and Clinical ApplicationsabstractMultiple sclerosis (MS) is a chronic inflammatory and degenerative disease of the central nervous system, characterized by the appearance of focal lesions in the white and gray matter that topographically correlate with an individual patient's neurological symptoms and signs. Magnetic resonance imaging (MRI) provides detailed in-vivo structural information, permitting the quantification and categorization of MS lesions that critically inform disease management. Traditionally, MS lesions have been manually annotated on 2D MRI slices, a process that is inefficient and prone to inter-/intra-observer errors. Recently, automated statistical imaging analysis techniques have been proposed to detect and segment MS lesions based on MRI voxel intensity. However, their effectiveness is limited by the heterogeneity of both MRI data acquisition techniques and the appearance of MS lesions. By learning complex lesion representations directly from images, deep learning techniques have achieved remarkable breakthroughs in the MS lesion segmentation task. Here, we provide a comprehensive review of state-of-the-art automatic statistical and deep-learning MS segmentation methods and discuss current and future clinical applications. Further, we review technical strategies, such as domain adaptation, to enhance MS lesion segmentation in real-world clinical settings. Chaoyi Zhang, Mariano Cabezas, Yang Song 0001, Zihao Tang 0002, Dongnan Liu, Tom Weidong Cai, Michael Barnett 0006, Chenyu Wang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | DSNet: A Dual-Stream Framework for Weakly-Supervised Gigapixel Pathology Image AnalysisabstractWe present a novel weakly-supervised framework for classifying whole slide images (WSIs). WSIs, due to their gigapixel resolution, are commonly processed by patch-wise classification with patch-level labels. However, patch-level labels require precise annotations, which is expensive and usually unavailable on clinical data. With image-level labels only, patch-wise classification would be sub-optimal due to inconsistency between the patch appearance and image-level label. To address this issue, we posit that WSI analysis can be effectively conducted by integrating information at both high magnification (local) and low magnification (regional) levels. We auto-encode the visual signals in each patch into a latent embedding vector representing local information, and down-sample the raw WSI to hardware-acceptable thumbnails representing regional information. The WSI label is then predicted with a Dual-Stream Network (DSNet), which takes the transformed local patch embeddings and multi-scale thumbnail images as inputs and can be trained by the image-level label only. Experiments conducted on three large-scale public datasets demonstrate that our method outperforms all recent state-of-the-art weakly-supervised WSI classification methods. Tiange Xiang, Yang Song 0001, Chaoyi Zhang, Dongnan Liu, Fan Zhang 0013, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Network Pruning via Performance MaximizationabstractChannel pruning is a class of powerful methods for model compression. When pruning a neural network, it's ideal to obtain a sub-network with higher accuracy. However, a sub-network does not necessarily have high accuracy with low classification loss (loss-metric mismatch). In the paper, we first consider the loss-metric mismatch problem for pruning and propose a novel channel pruning method for Convolutional Neural Networks (CNNs) by directly maximizing the performance (i.e., accuracy) of sub-networks. Specifically, we train a stand-alone neural network to predict sub-networks' performance and then maximize the output of the network as a proxy of accuracy to guide pruning. Training such a performance prediction network efficiently is not an easy task, and it may potentially suffer from the problem of catastrophic forgetting and the imbalance distribution of sub-networks. To deal with this challenge, we introduce a corresponding episodic memory to update and collect sub-networks during the pruning process. In the experiment section, we further demonstrate that the gradients from the performance prediction network and the classification loss have different directions. Extensive experimental results show that the proposed method can achieve state-of-the-art performance with ResNet, MobileNetV2, and ShuffleNetV2+ on ImageNet and CIFAR-10. Shangqian Gao, Feihu Huang 0001, Tom Weidong Cai, Heng Huang 0001 |
CVPR | 3 |
| 2021 | Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph AnalysisabstractScene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene under-standing via three sequential stages, namely scene graph construction, reasoning, and inference. Within the reasoning stage, an EDGE-oriented Graph Convolutional Network (EdgeGCN) is created to exploit multi-dimensional edge features for explicit relationship modeling, together with the exploration of two associated twinning interaction mechanisms between nodes and edges for the independent evolution of scene graph representations. Overall, our integrated SGGpointframework is established to seek and infer scene structures of interest from both real-world and synthetic 3D point-based scenes. Our experimental results show promising edge-oriented reasoning effects on scene graph generation studies. We also demonstrate our method advantage on several traditional graph representation learning benchmark datasets, including the node-wise classification on citation networks and whole-graph recognition problems for molecular analysis. Chaoyi Zhang, Jianhui Yu, Yang Song 0001, Tom Weidong Cai |
CVPR | 4 |
| 2021 | Walk in the Cloud: Learning Curves for Point Clouds Shape AnalysisabstractDiscrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequently aggregated back to augment their pointwise features. We provide an effective implementation of the proposed aggregation strategy including a novel curve grouping operator followed by a curve aggregation operator. Our method was benchmarked on several point cloud analysis tasks where we achieved the state-of-the-art classification accuracy of 94.2% on the ModelNet40 classification task, instance IoU of 86.8% on the ShapeNetPart segmentation task and cosine error of 0.11 on the ModelNet40 normal estimation task. Our project page with source code is available at: https://curvenet.github.io/. Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Jianhui Yu, Tom Weidong Cai |
ICCV | 5 |
| 2021 | Iterative Subnetwork With Linear Hierarchical Ordering for Human Pose EstimationabstractHuman pose estimation is a long-standing and challenging problem in computer vision. Many recent advancements in the field have relied on complex structure refinement and specific human joint graphical relations. However, progress has been saturated in terms of accuracy. Each time, new state-of-the-art approaches only improve accuracy by less than 0.3% in the MPII test set despite using complicated model structures. Most recent developments can be summarized into two main ideas: 1) refinement subnetwork to improve predictions iteratively and 2) exploitation of human joint graphical relations. In this work, we present how efficient and simple iterative subnetworks with linear hierarchical ordering based on the aforementioned ideas can help to improve accuracy on strong backbone models. Different versions of iterative subnetwork are examined. Significant improvements on difficult body part predictions such as wrists and ankles using simple convolution subnetwork are observed. Further improvements can be made by using a large receptive field subnetwork such as axial-transformer [1]. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 4 |
| 2021 | ICE-GAN: Identity-Aware and Capsule-Enhanced GAN with Graph-Based Reasoning for Micro-Expression Recognition and SynthesisabstractMicro-expressions are reflections of people's true feelings and motives, which attract an increasing number of researchers into the study of automatic facial micro-expression recognition. The short detection window, the subtle facial muscle movements, and the limited training samples make micro-expression recognition challenging. To this end, we propose a novel Identity-aware and Capsule-Enhanced Generative Adversarial Network with graph-based reasoning (ICE-GAN), introducing micro-expression synthesis as an auxiliary task to assist recognition. The generator produces synthetic faces with controllable micro-expressions and identity-aware features, whose long-ranged dependencies are captured through the graph reasoning module (GRM), and the discriminator detects the image authenticity and expression classes. Our ICE-GAN was evaluated on Micro-Expression Grand Challenge 2019 (MEGC2019) with a significant improvement (12.9%) over the winner and surpassed other state-of-the-art methods. Jianhui Yu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
IJCNN | 4 |
| 2021 | Deep Fiber Clustering: Anatomically Informed Unsupervised Deep Learning for Fast and Effective White Matter Parcellation
Yuqian Chen, Chaoyi Zhang, Yang Song 0001, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
MICCAI (7) | 6 |
| 2021 | LG-Net: Lesion Gate Network for Multiple Sclerosis Lesion Inpainting
Zihao Tang 0002, Mariano Cabezas, Dongnan Liu, Michael Barnett 0006, Tom Weidong Cai, Chenyu Wang 0001 |
MICCAI (7) | 5 |
| 2021 | BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
Xinyi Wang 0015, Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 7 |
| 2021 | Development and Validation of an Unsupervised Feature Learning System for Leukocyte Characterization and Classification: A Multi-Hospital Study
Xuanyu Mao, Yongquan Xia, Chengbin Wang, Xuejing Xu, Xie Zhao, Guoye Liu, Zhiqiong Wang, Tom Weidong Cai, Hang Chang |
Int. J. Comput. Vis. | 18 |
| 2021 | Panoptic Feature Fusion Net: A Novel Instance Segmentation Paradigm for Biomedical and Biological ImagesabstractInstance segmentation is an important task for biomedical and biological image analysis. Due to the complicated background components, the high variability of object appearances, numerous overlapping objects, and ambiguous object boundaries, this task still remains challenging. Recently, deep learning based methods have been widely employed to solve these problems and can be categorized into proposal-free and proposal-based methods. However, both proposal-free and proposal-based methods suffer from information loss, as they focus on either global-level semantic or local-level instance features. To tackle this issue, we present a Panoptic Feature Fusion Net (PFFNet) that unifies the semantic and instance features in this work. Specifically, our proposed PFFNet contains a residual attention feature fusion mechanism to incorporate the instance prediction with the semantic features, in order to facilitate the semantic contextual information learning in the instance branch. Then, a mask quality sub-branch is designed to align the confidence score of each object with the quality of the mask prediction. Furthermore, a consistency regularization mechanism is designed between the semantic segmentation tasks in the semantic and instance branches, for the robust learning of both tasks. Extensive experiments demonstrate the effectiveness of our proposed PFFNet, which outperforms several state-of-the-art methods on various biomedical and biological datasets. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Image Process. | 5 |
| 2021 | PDAM: A Panoptic-Level Feature Alignment Framework for Unsupervised Domain Adaptive Instance Segmentation in Microscopy ImagesabstractIn this work, we present an unsupervised domain adaptation (UDA) method, named Panoptic Domain Adaptive Mask R-CNN (PDAM), for unsupervised instance segmentation in microscopy images. Since there currently lack methods particularly for UDA instance segmentation, we first design a Domain Adaptive Mask R-CNN (DAM) as the baseline, with cross-domain feature alignment at the image and instance levels. In addition to the image- and instance-level domain discrepancy, there also exists domain bias at the semantic level in the contextual information. Next, we, therefore, design a semantic segmentation branch with a domain discriminator to bridge the domain gap at the contextual level. By integrating the semantic- and instance-level feature adaptation, our method aligns the cross-domain features at the panoptic level. Third, we propose a task re-weighting mechanism to assign trade-off weights for the detection and segmentation loss functions. The task re-weighting mechanism solves the domain bias issue by alleviating the task learning for some iterations when the features contain source-specific factors. Furthermore, we design a feature similarity maximization mechanism to facilitate instance-level feature adaptation from the perspective of representational learning. Different from the typical feature alignment methods, our feature similarity maximization mechanism separates the domain-invariant and domain-specific features by enlarging their feature distribution dependency. Experimental results on three UDA instance segmentation scenarios with five datasets demonstrate the effectiveness of our proposed PDAM method, which outperforms state-of-the-art UDA methods by a large margin. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Shape-Oriented Convolution Neural Network for Point Cloud AnalysisabstractPoint cloud is a principal data structure adopted for 3D geometric information encoding. Unlike other conventional visual data, such as images and videos, these irregular points describe the complex shape features of 3D objects, which makes shape feature learning an essential component of point cloud analysis. To this end, a shape-oriented message passing scheme dubbed ShapeConv is proposed to focus on the representation learning of the underlying shape formed by each local neighboring point. Despite this intra-shape relationship learning, ShapeConv is also designed to incorporate the contextual effects from the inter-shape relationship through capturing the long-ranged dependencies between local underlying shapes. This shape-oriented operator is stacked into our hierarchical learning architecture, namely Shape-Oriented Convolutional Neural Network (SOCNN), developed for point cloud analysis. Extensive experiments have been performed to evaluate its significance in the tasks of point cloud classification and part segmentation. Chaoyi Zhang, Yang Song 0001, Lina Yao 0001, Tom Weidong Cai |
AAAI | 4 |
| 2020 | Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-WeightingabstractUnsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architecture for unsupervised nuclei segmentation in histopathology images, by learning from fluorescence microscopy images. More specifically, we first propose a nuclei inpainting mechanism to remove the auxiliary generated objects in the synthesized images. Secondly, a semantic branch with a domain discriminator is designed to achieve panoptic-level domain adaptation. Thirdly, in order to avoid the influence of the source-biased features, we propose a task re-weighting mechanism to dynamically add trade-off weights for the task-specific loss functions. Experimental results on three datasets indicate that our proposed method outperforms state-of-the-art UDA methods significantly, and demonstrates a similar performance as fully supervised methods. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai |
CVPR | 8 |
| 2020 | Automatic Dropout for Deep Neural Networks
Veena Dodballapur, Rajanish Calisa, Yang Song 0001, Tom Weidong Cai |
ICONIP (3) | 4 |
| 2020 | Predicting Potential Propensity of Adolescents to Drugs via New Semi-supervised Deep Ordinal Regression Model
Alireza Ganjdanesh, Kamran Ghasedi, Liang Zhan, Tom Weidong Cai, Heng Huang 0001 |
MICCAI (1) | 4 |
| 2020 | Learning High-Resolution and Efficient Non-local Features for Brain Glioma Segmentation in MR Images
Haozhe Jia, Yong Xia 0001, Tom Weidong Cai, Heng Huang 0001 |
MICCAI (4) | 3 |
| 2020 | BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture
Tiange Xiang, Chaoyi Zhang, Dongnan Liu, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 6 |
| 2020 | DeepAntigen: a novel method for neoantigen prioritization via 3D genome and deep sparse learningabstractMOTIVATION: The mutations of cancers can encode the seeds of their own destruction, in the form of T-cell recognizable immunogenic peptides, also known as neoantigens. It is computationally challenging, however, to accurately prioritize the potential neoantigen candidates according to their ability of activating the T-cell immunoresponse, especially when the somatic mutations are abundant. Although a few neoantigen prioritization methods have been proposed to address this issue, advanced machine learning model that is specifically designed to tackle this problem is still lacking. Moreover, none of the existing methods considers the original DNA loci of the neoantigens in the perspective of 3D genome which may provide key information for inferring neoantigens' immunogenicity. RESULTS: In this study, we discovered that DNA loci of the immunopositive and immunonegative MHC-I neoantigens have distinct spatial distribution patterns across the genome. We therefore used the 3D genome information along with an ensemble pMHC-I coding strategy, and developed a group feature selection-based deep sparse neural network model (DNN-GFS) that is optimized for neoantigen prioritization. DNN-GFS demonstrated increased neoantigen prioritization power comparing to existing sequence-based approaches. We also developed a webserver named deepAntigen (http://yishi.sjtu.edu.cn/deepAntigen) that implements the DNN-GFS as well as other machine learning methods. We believe that this work provides a new perspective toward more accurate neoantigen prediction which eventually contribute to personalized cancer immunotherapy. AVAILABILITY AND IMPLEMENTATION: Data and implementation are available on webserver: http://yishi.sjtu.edu.cn/deepAntigen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Shi 0007, Zehua Guo 0004, Xianbin Su, Luming Meng, Minhua Zheng, Xueyin Shang, Wangqiu Cheng, Yaoliang Yu, Yujia Cai, Chaoyi Zhang, Tom Weidong Cai, Guang He, Zeguang Han |
Bioinform. | 15 |
| 2020 | NFN+: A novel network followed network for retinal vessel segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
Neural Networks | 5 |
| 2020 | 3D APA-Net: 3D Adversarial Pyramid Anisotropic Convolutional Network for Prostate Segmentation in MR ImagesabstractAccurate and reliable segmentation of the prostate gland using magnetic resonance (MR) imaging has critical importance for the diagnosis and treatment of prostate diseases, especially prostate cancer. Although many automated segmentation approaches, including those based on deep learning have been proposed, the segmentation performance still has room for improvement due to the large variability in image appearance, imaging interference, and anisotropic spatial resolution. In this paper, we propose the 3D adversarial pyramid anisotropic convolutional deep neural network (3D APA-Net) for prostate segmentation in MR images. This model is composed of a generator (i.e., 3D PA-Net) that performs image segmentation and a discriminator (i.e., a six-layer convolutional neural network) that differentiates between a segmentation result and its corresponding ground truth. The 3D PA-Net has an encoder-decoder architecture, which consists of a 3D ResNet encoder, an anisotropic convolutional decoder, and multi-level pyramid convolutional skip connections. The anisotropic convolutional blocks can exploit the 3D context information of the MR images with anisotropic resolution, the pyramid convolutional blocks address both voxel classification and gland localization issues, and the adversarial training regularizes 3D PA-Net and thus enables it to generate spatially consistent and continuous segmentation results. We evaluated the proposed 3D APA-Net against several state-of-the-art deep learning-based segmentation approaches on two public databases and the hybrid of the two. Our results suggest that the proposed model outperforms the compared approaches on three databases and could be used in a routine clinical workflow. Haozhe Jia, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Heng Huang 0001, Yanning Zhang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Human Pose Estimation Using Deep Convolutional Densenet Hourglass Network with Intermediate Points VotingabstractHuman pose estimation is a long-standing and challenging problem in computer vision. The problem involves high freedom of articulation of body limbs, different occlusions such as self-occlusion or occlusion by other objects or persons, various clothing, various background in the natural image and foreshortening due to different capturing angle of the camera. In this work we present 1) how the DenseNet module can be used to improve the original ResNet hourglass model, 2) how intermediate points derived from ground truth joint segments can be used as output augmentation of a convolutional neural network (ConvNet) to improve the prediction accuracy. Further improvement has also been made via intermediate points voting by optimizing the joint probability distribution of human joints and the intermediate points. Experimental results on the effects of intermediate point and optimization scheme are presented. We are able to achieve competitive results to the state-of-the-art methods by the proposed method. Shek Wai Chu, Yang Song 0001, Ju Jia Zou, Tom Weidong Cai |
ICIP | 4 |
| 2019 | Nuclei Segmentation via a Deep Panoptic Model with Semantic Feature FusionabstractAutomated detection and segmentation of individual nuclei in histopathology images is important for cancer diagnosis and prognosis. Due to the high variability of nuclei appearances and numerous overlapping objects, this task still remains challenging. Deep learning based semantic and instance segmentation models have been proposed to address the challenges, but these methods tend to concentrate on either the global or local features and hence still suffer from information loss. In this work, we propose a panoptic segmentation model which incorporates an auxiliary semantic segmentation branch with the instance branch to integrate global and local features. Furthermore, we design a feature map fusion mechanism in the instance branch and a new mask generator to prevent information loss. Experimental results on three different histopathology datasets demonstrate that our method outperforms the state-of-the-art nuclei segmentation methods and popular semantic and instance segmentation models by a large margin. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Chaoyi Zhang, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai |
IJCAI | 7 |
| 2019 | HD-Net: Hybrid Discriminative Network for Prostate Segmentation in MR Images
Haozhe Jia, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai, Yong Xia 0001 |
MICCAI (2) | 4 |
| 2019 | Vessel-Net: Retinal Vessel Segmentation Under Multi-path Supervision
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Dongnan Liu, Chaoyi Zhang, Tom Weidong Cai |
MICCAI (1) | 7 |
| 2019 | Integrating Heterogeneous Brain Networks for Predicting Brain Disease Conditions
Yanfu Zhang, Liang Zhan, Tom Weidong Cai, Paul M. Thompson, Heng Huang 0001 |
MICCAI (4) | 3 |
| 2019 | Feature-Based Patch Matching for Moving Object DetectionabstractIn this paper, a new background subtraction framework is proposed to deal with possible scenarios occurring in natural scenes. In this method, a combination of two feature descriptors, namely color information in HSV color format and global texture descriptor T, are introduced to effectively identify background points under varying conditions. Using these features, an adaptive background model is constructed to automatically adapt to scene changes. The proposed framework is evaluated on common change detection datasets, showing improved performance compared to three well-known methods. Mosin Russell, Ju Jia Zou, Gu Fang 0001, Tom Weidong Cai |
VCIP | 4 |
| 2019 | Feature-Based Image Patch Classification for Moving Shadow DetectionabstractThe presence of shadows in images significantly affects the performance of many computer vision tasks and visual processing applications, such as object tracking, object classification, and behavior recognition. Most methods have been designed to detect shadows in specific situations, but they often fail to distinguish shadow points from the foreground object in many problematic situations, such as chromatic shadows, non-textured and dark surfaces, and foreground-background camouflage. In this paper, we propose a new feature-based image patch approximation and multi-independent sparse representation technique to tackle these environmental problems. In this method, two illumination-invariant features-binary patterns of local color constancy and light-based gradient matching-are introduced, along with the intensity-reduction histogram. These features are extracted from image patches and are used to construct two over-complete dictionaries for objects and shadows, respectively. Given a new image patch, its best approximation for a number of iterations is found from each dictionary. For each iteration, an independent class assignment is performed by finding its distances from the reference dictionaries. The patch is then assigned to a class based on its probability of occurrence. The proposed framework is evaluated on common shadow detection data sets, and it shows improved performance in terms of the shadow detection rate and discrimination rate compared with the state-of-the-art methods. Mosin Russell, Ju Jia Zou, Gu Fang 0001, Tom Weidong Cai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Knowledge-based Collaborative Deep Learning for Benign-Malignant Lung Nodule Classification on Chest CTabstractThe accurate identification of malignant lung nodules on chest CT is critical for the early detection of lung cancer, which also offers patients the best chance of cure. Deep learning methods have recently been successfully introduced to computer vision problems, although substantial challenges remain in the detection of malignant nodules due to the lack of large training data sets. In this paper, we propose a multi-view knowledge-based collaborative (MV-KBC) deep model to separate malignant from benign nodules using limited chest CT data. Our model learns 3-D lung nodule characteristics by decomposing a 3-D nodule into nine fixed views. For each view, we construct a knowledge-based collaborative (KBC) submodel, where three types of image patches are designed to fine-tune three pre-trained ResNet-50 networks that characterize the nodules' overall appearance, voxel, and shape heterogeneity, respectively. We jointly use the nine KBC submodels to classify lung nodules with an adaptive weighting scheme learned during the error back propagation, which enables the MV-KBC model to be trained in an end-to-end manner. The penalty loss function is used for better reduction of the false negative rate with a minimal effect on the overall performance of the MV-KBC model. We tested our method on the benchmark LIDC-IDRI data set and compared it to the five state-of-the-art classification approaches. Our results show that the MV-KBC model achieved an accuracy of 91.60% for lung nodule classification with an AUC of 95.70%. These results are markedly superior to the state-of-the-art approaches. Yutong Xie 0001, Yong Xia 0001, Yang Song 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 7 |
| 2018 | Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy ImagesabstractStructural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance. Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai |
ICIP | 8 |
| 2018 | Whole Slide Image Classification via Iterative Patch LabellingabstractBrain tumor can be a fatal disease in the world. With the aim of improving survival rates, many computerized algorithms have been proposed to assist the pathologists to make a diagnosis' using Whole Slide Pathology Images (WSI). Most methods focus on performing patch-level classification and aggregating the patch-level results to obtain the image classification. Since not all patches carry diagnostic information, it is thus important for our algorithm to recognize discriminative and non-discriminative patches. In this study, we propose an iterative patch labelling algorithm based on the Convolutional Neural Network (CNN), with a well-designed thresholding scheme, a training policy and a novel discriminative model architecture, to distinguish patches and use the discriminative ones to achieve WSI -classification. Our method is evaluated on the MICCAI 2015 Challenge Dataset, and shows a large improvement over the baseline approaches. Chaoyi Zhang, Yang Song 0001, Donghao Zhang 0004, Sidong Liu, Tom Weidong Cai |
ICIP | 6 |
| 2018 | 3D Large Kernel Anisotropic Network for Brain Tumor Segmentation
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai |
ICONIP (7) | 6 |
| 2018 | 3D Anisotropic Hybrid Network: Transferring Convolutional Features from 2D Images to 3D Anisotropic Volumes
Siqi Liu 0001, Daguang Xu, Shaohua Kevin Zhou, Olivier Pauly, Sasa Grbic, Thomas Mertelmeier, Julia Wicklein, Anna K. Jerebko, Tom Weidong Cai, Dorin Comaniciu |
MICCAI (2) | 9 |
| 2018 | Temporal Correlation Structure Learning for MCI Conversion Prediction
Xiaoqian Wang 0001, Tom Weidong Cai, Dinggang Shen, Heng Huang 0001 |
MICCAI (3) | 2 |
| 2018 | Multiscale Network Followed Network Model for Retinal Vessel Segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
MICCAI (2) | 5 |
| 2018 | Panoptic Segmentation with an End-to-End Cell R-CNN for Pathology Image Analysis
Donghao Zhang 0004, Yang Song 0001, Dongnan Liu, Haozhe Jia, Siqi Liu 0001, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (2) | 8 |
| 2018 | Atlas registration and ensemble deep convolutional neural network-based prostate segmentation using magnetic resonance imaging
Haozhe Jia, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 4 |
| 2018 | Merged region based image retrieval
Fanjie Meng, Dalong Shan, Ruixia Shi, Yang Song 0001, Baolong Guo 0001, Tom Weidong Cai |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Locality constrained encoding of frequency and spatial information for image classification
Yongsheng Pan, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai |
Multim. Tools Appl. | 4 |
| 2018 | Dense and Sparse Labeling With Multidimensional Features for Saliency DetectionabstractConventional low-level feature-based saliency detection methods tend to use nonrobust prior knowledge and do not perform well in complex or low-contrast images. In this paper, to address these issues in existing methods, we propose a novel deep neural network (DNN)-based dense and sparse labeling (DSL) framework for saliency detection. DSL consists of three major steps, namely, dense labeling (DL), sparse labeling (SL), and deep convolutional (DC) network. The DL and SL steps conduct initial saliency estimations with macro object contours and low-level image features, respectively, which effectively approximate the location of the salient object and generate accurate guidance channels for the DC step; the DC step, on the other hand, takes in the results of DL and SL, establishes a six-channeled input data structure (including local superpixel information), and conducts accurate final saliency classification. Our DSL framework exploits the saliency estimation guidance from both macro object contours and local low-level features, as well as utilizing the DNN for high-level saliency feature extraction. Extensive experiments are conducted on six well-recognized public data sets against 16 state-of-the-art saliency detection methods, including ten conventional feature-based methods and six learning-based methods. The results demonstrate the superior performance of DSL on various challenging cases in terms of both accuracy and robustness. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Reversion Correction and Regularized Random Walk Ranking for Saliency DetectionabstractIn recent saliency detection research, many graph-based algorithms have applied boundary priors as background queries, which may generate completely "reversed" saliency maps if the salient objects are on the image boundaries. Moreover, these algorithms usually depend heavily on pre-processed superpixel segmentation, which may lead to notable degradation in image detail features. In this paper, a novel saliency detection method is proposed to overcome the above issues. First, we propose a saliency reversion correction process, which locates and removes the boundary-adjacent foreground superpixels, and thereby increases the accuracy and robustness of the boundary prior-based saliency estimations. Second, we propose a regularized random walk ranking model, which introduces prior saliency estimation to every pixel in the image by taking both region and pixel image features into account, thus leading to pixel-detailed and superpixel-independent saliency maps. Experiments are conducted on four well-recognized data sets; the results indicate the superiority of our proposed method against 14 state-of-the-art methods, and demonstrate its general extensibility as a saliency optimization algorithm. We further evaluate our method on a new data set comprised of images that we define as boundary adjacent object saliency, on which our method performs better than the comparison methods. Yuchen Yuan, ChangYang Li, Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Image Process. | 4 |
| 2018 | Automated 3-D Neuron Tracing With Precise Branch Erasing and Confidence Controlled Back TrackingabstractThe automatic reconstruction of single neurons from microscopic images is essential to enable large-scale data-driven investigations in neuron morphology research. However, few previous methods were able to generate satisfactory results automatically from 3-D microscopic images without human intervention. In this paper, we developed a new algorithm for automatic 3-D neuron reconstruction. The main idea of the proposed algorithm is to iteratively track backward from the potential neuronal termini to the soma centre. An online confidence score is computed to decide if a tracing iteration should be stopped and discarded from the final reconstruction. The performance improvements comparing with the previous methods are mainly introduced by a more accurate estimation of the traced area and the confidence controlled back-tracking algorithm. The proposed algorithm supports large-scale batch-processing by requiring only one user specified parameter for background segmentation. We bench tested the proposed algorithm on the images obtained from both the DIADEM challenge and the BigNeuron challenge. Our proposed algorithm achieved the state-of-the-art results. Siqi Liu 0001, Donghao Zhang 0004, Yang Song 0001, Hanchuan Peng, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 5 |
| 2018 | Multi-Pass Fast Watershed for Accurate Segmentation of Overlapping Cervical CellsabstractThe task of segmenting cell nuclei and cytoplasm in pap smear images is one of the most challenging tasks in automated cervix cytological analysis due to specifically the presence of overlapping cells. This paper introduces a multi-pass fast watershed-based method (MPFW) to segment both nucleus and cytoplasm from large cell masses of overlapping cervical cells in three watershed passes. The first pass locates the nuclei with barrier-based watershed on the gradient-based edge map of a pre-processed image. The next pass segments the isolated, touching, and partially overlapping cells with a watershed transform adapted to the cell shape and location. The final pass introduces mutual iterative watersheds separately applied to each nucleus in the largely overlapping clusters to estimate the cell shape. In MPFW, the line-shaped contours of the watershed cells are deformed with ellipse fitting and contour adjustment to give a better representation of cell shapes. The performance of the proposed method has been evaluated using synthetic, real extended depth-of-field, and multi-layers cervical cytology images provided by the first and second overlapping cervical cytology image segmentation challenges in ISBI 2014 and ISBI 2015. The experimental results demonstrate superior performance of the proposed MPFW in terms of segmentation accuracy, detection rate, and time complexity, compared with recent peer methods. Afaf Tareef, Yang Song 0001, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 7 |
| 2017 | Video Recovery via Learning Variation and Consistency of ImagesabstractMatrix completion algorithms have been popularly used to recover images with missing entries, and they are proved to be very effective. Recent works utilized tensor completion models in video recovery assuming that all video frames are homogeneous and correlated. However, real videos are made up of different episodes or scenes, i.e. heterogeneous. Therefore, a video recovery model which utilizes both video spatiotemporal consistency and variation is necessary. To solve this problem, we propose a new video recovery method Sectional Trace Norm with Variation and Consistency Constraints (STN-VCC). In our model, capped L1-norm regularization is utilized to learn the spatial-temporal consistency and variation between consecutive frames in video clips. Meanwhile, we introduce a new low-rank model to capture the low-rank structure in video frames with a better approximation of rank minimization than traditional trace norm. An efficient optimization algorithm is proposed, and we also provide a proof of convergence in the paper. We evaluate the proposed method via several video recovery tasks and experiment results show that our new method consistently outperforms other related approaches. Zhouyuan Huo, Shangqian Gao, Tom Weidong Cai, Heng Huang 0001 |
AAAI | 3 |
| 2017 | Deep Clustering via Joint Convolutional Autoencoder Embedding and Relative Entropy MinimizationabstractIn this paper, we propose a new clustering model, called DEeP Embedded Regularized ClusTering (DEPICT), which efficiently maps data into a discriminative embedding subspace and precisely predicts cluster assignments. DEPICT generally consists of a multinomial logistic regression function stacked on top of a multi-layer convolutional autoencoder. We define a clustering objective function using relative entropy (KL divergence) minimization, regularized by a prior for the frequency of cluster assignments. An alternating strategy is then derived to optimize the objective by updating parameters and estimating cluster assignments. Furthermore, we employ the reconstruction loss functions in our autoencoder, as a data-dependent regularization term, to prevent the deep embedding function from overfitting. In order to benefit from end-to-end optimization and eliminate the necessity for layer-wise pre-training, we introduce a joint learning framework to minimize the unified clustering and reconstruction loss functions together and train all network layers simultaneously. Experimental results indicate the superiority and faster running time of DEPICT in real-world clustering tasks, where no labeled data is available for hyper-parameter tuning. Kamran Ghasedi Dizaji, Amirhossein Herandi, Cheng Deng 0002, Tom Weidong Cai, Heng Huang 0001 |
ICCV | 4 |
| 2017 | Locally-Transferred Fisher Vectors for Texture ClassificationabstractTexture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. However, by truncating the CNN model at the last convolutional layer, the CNN-based FV descriptors would not incorporate the full capability of neural networks in feature learning. In this study, we propose that we can further transform the CNN-based FV descriptors in a neural network model to obtain more discriminative feature representations. In particular, we design a locally-transferred Fisher vector (LFV) method, which involves a multi-layer neural network model containing locally connected layers to transform the input FV descriptors with filters of locally shared weights. The network is optimized based on the hinge loss of classification, and transferred FV descriptors are then used for image classification. Our results on three challenging texture image datasets show improved performance over the state-of-the-art approaches. Yang Song 0001, Fan Zhang 0013, Qing Li 0012, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai |
ICCV | 6 |
| 2017 | Supervised Intra-embedding of Fisher Vectors for Histopathology Image Classification
Yang Song 0001, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
MICCAI (3) | 4 |
| 2017 | Transferable Multi-model Ensemble for Benign-Malignant Lung Nodule Classification on Chest CT
Yutong Xie 0001, Yong Xia 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai |
MICCAI (3) | 6 |
| 2017 | Supra-Threshold Fiber Cluster Statistics for Data-Driven Whole Brain Tractography Analysis
Fan Zhang 0013, Weining Wu, Lipeng Ning, Gloria McAnulty, Deborah P. Waber, Borjan A. Gagoski, Kiera Sarill, Hesham M. Hamoda, Yang Song 0001, Tom Weidong Cai, Yogesh Rathi, Lauren O'Donnell |
MICCAI (1) | 10 |
| 2017 | Regularized Modal Regression with Applications in Cognitive Impairment PredictionabstractLinear regression models have been successfully used to function estimation and model selection in high-dimensional data analysis. However, most existing methods are built on least squares with the mean square error (MSE) criterion, which are sensitive to outliers and their performance may be degraded for heavy-tailed noise. In this paper, we go beyond this criterion by investigating the regularized modal regression from a statistical learning viewpoint. A new regularized modal regression model is proposed for estimation and variable selection, which is robust to outliers, heavy-tailed noise, and skewed noise. On the theoretical side, we establish the approximation estimate for learning the conditional mode function, the sparsity analysis for variable selection, and the robustness characterization. On the application side, we applied our model to successfully improve the cognitive impairment prediction using the Alzheimer’s Disease Neuroimaging Initiative (ADNI) cohort data. Xiaoqian Wang 0001, Hong Chen 0004, Tom Weidong Cai, Dinggang Shen, Heng Huang 0001 |
NIPS | 3 |
| 2017 | A spatially cohesive superpixel model for image noise level estimation
Peng Fu 0003, ChangYang Li, Tom Weidong Cai, Quan-Sen Sun |
Neurocomputing | 3 |
| 2017 | Automatic segmentation of overlapping cervical smear cells based on local distinctive features and guided shape deformation
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Hang Chang, Yue Joseph Wang, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 3 |
| 2017 | Optimizing the cervix cytological examination based on deep learning and dynamic shape modeling
Afaf Tareef, Yang Song 0001, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng, Tom Weidong Cai |
Neurocomputing | 7 |
| 2017 | Guest Editorial: Special issue on advances in computing techniques for big medical image data
Yuanjie Zheng, Shaoting Zhang 0001, Junzhou Huang, Tom Weidong Cai |
Neurocomputing | 4 |
| 2017 | Dual discriminative local coding for tissue aging analysis
Yang Song 0001, Qing Li 0012, Fan Zhang 0013, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
Medical Image Anal. | 8 |
| 2017 | Low Dimensional Representation of Fisher Vectors for Microscopy Image ClassificationabstractMicroscopy image classification is important in various biomedical applications, such as cancer subtype identification, and protein localization for high content screening. To achieve automated and effective microscopy image classification, the representative and discriminative capability of image feature descriptors is essential. To this end, in this paper, we propose a new feature representation algorithm to facilitate automated microscopy image classification. In particular, we incorporate Fisher vector (FV) encoding with multiple types of local features that are handcrafted or learned, and we design a separation-guided dimension reduction method to reduce the descriptor dimension while increasing its discriminative capability. Our method is evaluated on four publicly available microscopy image data sets of different imaging types and applications, including the UCSB breast cancer data set, MICCAI 2015 CBTC challenge data set, and IICBU malignant lymphoma, and RNAi data sets. Our experimental results demonstrate the advantage of the proposed low-dimensional FV representation, showing consistent performance improvement over the existing state of the art and the commonly used dimension reduction techniques. Yang Song 0001, Qing Li 0012, Heng Huang 0001, David Dagan Feng, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 6 |
| 2016 | Integrative Analysis of Cellular Morphometric Context Reveals Clinically Relevant Signatures in Lower Grade Glioma
Ju Han, Yunfu Wang, Tom Weidong Cai, Alexander Borowsky, Bahram Parvin, Hang Chang |
MICCAI (1) | 3 |
| 2016 | Error Analysis of Generalized Nyström Kernel RegressionabstractNystr\"{o}m method has been used successfully to improve the computational efficiency of kernel ridge regression (KRR). Recently, theoretical analysis of Nystr\"{o}m KRR, including generalization bound and convergence rate, has been established based on reproducing kernel Hilbert space (RKHS) associated with the symmetric positive semi-definite kernel. However, in real world applications, RKHS is not always optimal and kernel function is not necessary to be symmetric or positive semi-definite. In this paper, we consider the generalized Nystr\"{o}m kernel regression (GNKR) with $\ell_2$ coefficient regularization, where the kernel just requires the continuity and boundedness. Error analysis is provided to characterize its generalization performance and the column norm sampling is introduced to construct the refined hypothesis space. In particular, the fast learning rate with polynomial decay is reached for the GNKR. Experimental analysis demonstrates the satisfactory performance of GNKR with the column norm sampling. Hong Chen 0004, Haifeng Xia, Heng Huang 0001, Tom Weidong Cai |
NIPS | 4 |
| 2016 | Bioimage classification with subcategory discriminant transform of high dimensional visual descriptorsabstractBACKGROUND: Bioimage classification is a fundamental problem for many important biological studies that require accurate cell phenotype recognition, subcellular localization, and histopathological classification. In this paper, we present a new bioimage classification method that can be generally applicable to a wide variety of classification problems. We propose to use a high-dimensional multi-modal descriptor that combines multiple texture features. We also design a novel subcategory discriminant transform (SDT) algorithm to further enhance the discriminative power of descriptors by learning convolution kernels to reduce the within-class variation and increase the between-class difference. RESULTS: We evaluate our method on eight different bioimage classification tasks using the publicly available IICBU 2008 database. Each task comprises a separate dataset, and the collection represents typical subcellular, cellular, and tissue level classification problems. Our method demonstrates improved classification accuracy (0.9 to 9%) on six tasks when compared to state-of-the-art approaches. We also find that SDT outperforms the well-known dimension reduction techniques, with for example 0.2 to 13% improvement over linear discriminant analysis. CONCLUSIONS: We present a general bioimage classification method, which comprises a highly descriptive visual feature representation and a learning-based discriminative feature transformation algorithm. Our evaluation on the IICBU 2008 database demonstrates improved performance over the state-of-the-art for six different classification tasks. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang |
BMC Bioinform. | 2 |
| 2016 | DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutationsabstractBACKGROUND: With the developments of DNA sequencing technology, large amounts of sequencing data have become available in recent years and provide unprecedented opportunities for advanced association studies between somatic point mutations and cancer types/subtypes, which may contribute to more accurate somatic point mutation based cancer classification (SMCC). However in existing SMCC methods, issues like high data sparsity, small volume of sample size, and the application of simple linear classifiers, are major obstacles in improving the classification performance. RESULTS: To address the obstacles in existing SMCC studies, we propose DeepGene, an advanced deep neural network (DNN) based classifier, that consists of three steps: firstly, the clustered gene filtering (CGF) concentrates the gene data by mutation occurrence frequency, filtering out the majority of irrelevant genes; secondly, the indexed sparsity reduction (ISR) converts the gene data into indexes of its non-zero elements, thereby significantly suppressing the impact of data sparsity; finally, the data after CGF and ISR is fed into a DNN classifier, which extracts high-level features for accurate classification. Experimental results on our curated TCGA-DeepGene dataset, which is a reformulated subset of the TCGA dataset containing 12 selected types of cancer, show that CGF, ISR and DNN all contribute in improving the overall classification performance. We further compare DeepGene with three widely adopted classifiers and demonstrate that DeepGene has at least 24% performance improvement in terms of testing accuracy. CONCLUSIONS: Based on deep learning and somatic point mutation data, we devise DeepGene, an advanced cancer type classifier, which addresses the obstacles in existing SMCC studies. Experiments indicate that DeepGene outperforms three widely adopted existing classifiers, which is mainly attributed to its deep learning module that is able to extract the high level features between combinatorial somatic point mutations and cancer types. Yuchen Yuan, Yi Shi 0007, ChangYang Li, Jinman Kim, Tom Weidong Cai, Zeguang Han, David Dagan Feng |
BMC Bioinform. | 5 |
| 2016 | Texture image classification with discriminative neural networksabstractTexture provides an important cue for many computer vision applications, and texture image classification has been an active research area over the past years. Recently, deep learning techniques using convolutional neural networks (CNN) have emerged as the state-of-the-art: CNN-based features provide a significant performance improvement over previous handcrafted features. In this study, we demonstrate that we can further improve the discriminative power of CNN-based features and achieve more accurate classification of texture images. In particular, we have designed a discriminative neural network-based feature transformation (NFT) method, with which the CNN-based features are transformed to lower dimensionality descriptors based on an ensemble of neural networks optimized for the classification objective. For evaluation, we used three standard benchmark datasets (KTH-TIPS2, FMD, and DTD) for texture image classification. Our experimental results show enhanced classification performance over the state-of-the-art. Yang Song 0001, Qing Li 0012, David Dagan Feng, Ju Jia Zou, Tom Weidong Cai |
Comput. Vis. Media | 5 |
| 2016 | Dictionary pruning with visual word significance for medical image retrieval
Fan Zhang 0013, Yang Song 0001, Tom Weidong Cai, Alex Hauptmann 0001, Sidong Liu, Sonia Pujol, Ron Kikinis, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 3 |
| 2015 | Robust Capped Norm Nonnegative Matrix Factorization: Capped Norm NMFabstractAs an important matrix factorization model, Nonnegative Matrix Factorization (NMF) has been widely used in information retrieval and data mining research. Standard Nonnegative Matrix Factorization is known to use the Frobenius norm to calculate the residual, making it sensitive to noises and outliers. It is desirable to use robust NMF models for practical applications, in which usually there are many data outliers. It has been studied that the 2,1, or 1-norm can be used for robust NMF formulations to deal with data outliers. However, these alternatives still suffer from the extreme data outliers. In this paper, we present a novel robust capped norm orthogonal Nonnegative Matrix Factorization model, which utilizes the capped norm for the objective to handle these extreme outliers. Meanwhile, we derive a new efficient optimization algorithm to solve the proposed non-convex non-smooth objective. Extensive experiments on both synthetic and real datasets show our proposed new robust NMF method consistently outperforms related approaches. Hongchang Gao, Feiping Nie 0001, Tom Weidong Cai, Heng Huang 0001 |
CIKM | 3 |
| 2015 | Weakly Supervised Natural Language Processing Framework for Abstractive Multi-Document Summarization: Weakly Supervised Abstractive Multi-Document SummarizationabstractIn this paper, we propose a new weakly supervised abstractive news summarization framework using pattern based approaches. Our system first generates meaningful patterns from sentences. Then, in order to precisely cluster patterns, we propose a novel semisupervised pattern learning algorithm that leverages a hand-crafted list of topic-relevant keywords, which are the only weakly supervised information used by our framework to generate aspect-oriented summarization. After that, our system generates new patterns by fusing existing patterns and selecting top ranked new patterns via the recurrent neural network language model. Finally, we introduce a new pattern based surface realization algorithm to generate abstractive summaries. Automatic and manual evaluations demonstrate the effectiveness and advantages of our new methods. Code is available at: https://github.com/jerryli1981 Peng Li 0056, Tom Weidong Cai, Heng Huang 0001 |
CIKM | 2 |
| 2015 | Robust saliency detection via regularized random walks rankingabstractIn the field of saliency detection, many graph-based algorithms heavily depend on the accuracy of the pre-processed superpixel segmentation, which leads to significant sacrifice of detail information from the input image. In this paper, we propose a novel bottom-up saliency detection approach that takes advantage of both region-based features and image details. To provide more accurate saliency estimations, we first optimize the image boundary selection by the proposed erroneous boundary removal. By taking the image details and region-based estimations into account, we then propose the regularized random walks ranking to formulate pixel-wised saliency maps from the superpixel-based background and foreground saliency estimations. Experiment results on two public datasets indicate the significantly improved accuracy and robustness of the proposed algorithm in comparison with 12 state-of-the-art saliency detection approaches. ChangYang Li, Yuchen Yuan, Tom Weidong Cai, Yong Xia 0001, David Dagan Feng |
CVPR | 3 |
| 2015 | Fusing subcategory probabilities for texture classificationabstractTexture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these issues, in this paper, we propose a sub-categorization model for texture classification. By clustering each class into subcategories, classification probabilities at the subcategory-level are computed based on between-subcategory distinctiveness and within-subcategory representativeness. These subcategory probabilities are then fused based on their contribution levels and cluster qualities. This fused probability is added to the multiclass classification probability to obtain the final class label. Our method was applied to texture classification on three challenging datasets - KTH-TIPS2, FMD and DTD, and has shown excellent performance in comparison with the state-of-the-art approaches. Yang Song 0001, Tom Weidong Cai, Qing Li 0012, Fan Zhang 0013, David Dagan Feng, Heng Huang 0001 |
CVPR | 2 |
| 2015 | Subject-centered multi-view feature fusion for neuroimaging retrieval and classificationabstractMulti-View neuroimaging retrieval and classification play an important role in computer-aided-diagnosis of brain disorders, as multi-view features could provide more insights of the disease pathology and potentially lead to more accurate diagnosis than single-view features. The large inter-feature and inter-subject variations make the multi-view neuroimaging analysis a challenging task. Many multi-view or multi-modal feature fusion methods have been proposed to reduce the impact of inter-feature variations in neuroimaging data. However, there is not much in-depth work focusing on the inter-subject variations. In this study, we propose a subject-centered multi-view feature fusion method for neuroimaging retrieval and classification based on the propagation graph fusion (PGF) algorithm. Two main advantages of the proposed method are: 1) it evaluates the query online and adaptively reshapes the connections between subjects according to the query; 2) it measures the affinity of the query to the subjects using the subject-centered affinity matrices, which can be easily combined and efficiently solved. Evaluated using a public accessible neuroimaging database, our algorithm outperforms the state-of-the-art methods in retrieval and achieves comparable performance in classification. Sidong Liu, Tom Weidong Cai, Siqi Liu 0001, Sonia Pujol, Ron Kikinis, David Dagan Feng |
ICIP | 2 |
| 2015 | Beating cilia identification in fluorescence microscope images for accurate CBF measurementabstractCiliary beating frequency (CBF) is a regulated quantitative measurement to describe ciliary beating properties. It is widely used for diagnosis of defective mucociliary clearance diseases. Image-based methods can be effective for CBF estimation but also affected by the moving objects such as ciliated cells and debris. In this work, we propose a CBF estimation method by removing these unfavorable objects, which we refer to as foreground, so that we can focus on observing the beating cilia only. We firstly design a graph-based method to divide the cilia image into different regions. Next, the foreground regions are extracted and removed from the region division result. The beating cilia are then recognized from the background and used to compute the CBF. Our method conducts the CBF estimation by incorporating the cilia regions only and thus can provide a more accurate description of ciliary beating properties. Preliminary experimental results on cilia images showed the proposed method's potentials for accurate CBF measurement. Fan Zhang 0013, Tom Weidong Cai, Yang Song 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, David Dagan Feng |
ICIP | 2 |
| 2015 | Learning Shape-Driven Segmentation Based on Neural Network and Sparse Reconstruction Toward Automated Cell Analysis of Cervical Smears
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
ICONIP (1) | 3 |
| 2015 | Anatomical Annotations for Drosophila Gene Expression Patterns via Multi-Dimensional Visual Descriptors Integration: Multi-Dimensional Feature LearningabstractIn Drosophila gene expression pattern research, the in situ hybridization (ISH) image has become the standard technique to visualize and study the spatial distribution of RNA. To facilitate the search and comparison of Drosophila gene expression patterns during Drosophila embryogenesis, it is highly desirable to annotate the tissue-level anatomical ontology terms for ISH images. In ISH image annotations, the image content representation is crucial to achieve satisfactory results. However, existing methods mainly focus on improving the classification algorithms and only using simple visual descriptor. If we integrate the effective local and holistic visual descriptors via proper learning method, we can achieve more accurate image annotation results than using individual visual descriptor. Hongchang Gao, Lin Yan 0003, Tom Weidong Cai, Heng Huang 0001 |
KDD | 3 |
| 2015 | Motion Representation of Ciliated Cell Images with Contour-Alignment for Automated CBF Estimation
Fan Zhang 0013, Yang Song 0001, Siqi Liu 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, Sidong Liu, David Dagan Feng, Tom Weidong Cai |
MICCAI (3) | 11 |
| 2015 | Locality-constrained Subcluster Representation Ensemble for lung image classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, Yue Joseph Wang, David Dagan Feng |
Medical Image Anal. | 2 |
| 2015 | Large Margin Local Estimate With Applications to Medical Image ClassificationabstractMedical images usually exhibit large intra-class variation and inter-class ambiguity in the feature space, which could affect classification accuracy. To tackle this issue, we propose a new Large Margin Local Estimate (LMLE) classification model with sub-categorization based sparse representation. We first sub-categorize the reference sets of different classes into multiple clusters, to reduce feature variation within each subcategory compared to the entire reference set. Local estimates are generated for the test image using sparse representation with reference subcategories as the dictionaries. The similarity between the test image and each class is then computed by fusing the distances with the local estimates in a learning-based large margin aggregation construct to alleviate the problem of inter-class ambiguity. The derived similarities are finally used to determine the class label. We demonstrate that our LMLE model is generally applicable to different imaging modalities, and applied it to three tasks: interstitial lung disease (ILD) classification on high-resolution computed tomography (HRCT) images, phenotype binary classification and continuous regression on brain magnetic resonance (MR) imaging. Our experimental results show statistically significant performance improvements over existing popular classifiers. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng, Yue Joseph Wang, Michael J. Fulham |
IEEE Trans. Medical Imaging | 2 |
| 2014 | Medical image classification with convolutional neural networkabstractImage patch classification is an important task in many different medical imaging applications. In this work, we have designed a customized Convolutional Neural Networks (CNN) with shallow convolution layer to classify lung image patches with interstitial lung disease (ILD). While many feature descriptors have been proposed over the past years, they can be quite complicated and domain-specific. Our customized CNN framework can, on the other hand, automatically and efficiently learn the intrinsic image features from lung image patches that are most suitable for the classification purpose. The same architecture can be generalized to perform other medical image or texture classification tasks. Qing Li 0012, Tom Weidong Cai, Xiaogang Wang 0001, Yun Zhou 0006, David Dagan Feng |
ICARCV | 2 |
| 2014 | Propagation graph fusion for multi-modal medical content-based retrievalabstractMedical content-based retrieval (MCBR) plays an important role in computer aided diagnosis and clinical decision support. Multi-modal imaging data have been increasingly used in MCBR, as they could provide more insights of the diseases and complement the deficiencies of single-modal data. However, it is very challenging to fuse data in different modalities since they have different physical fundamentals and large value range variations. In this study, we propose a novel Propagation Graph Fusion (PGF) framework for multi-modal medical data retrieval. PGF models the subjects' relationships in single modalities using the directed propagation graphs, and then fuses the graphs into a single graph by summing up the edge weights. Our proposed PGF method could reduce the large inter-modality and inter-subject variations, and can be solved efficiently using the PageRank algorithm. We test the proposed method on a public medical database with 331 subjects using features extracted from two imaging modalities, PET and MRI. The preliminary results show that our PGF method could enhance multi-modal retrieval and modestly outperform the state-of-the-art single-modal and multi-modal retrieval methods. Sidong Liu, Siqi Liu 0001, Sonia Pujol, Ron Kikinis, David Dagan Feng, Tom Weidong Cai |
ICARCV | 6 |
| 2014 | Automated three-stage nucleus and cytoplasm segmentation of overlapping cellsabstractDeveloping segmentation techniques for overlapping cells has become a major hurdle for automated analysis of cervical cells. In this paper, an automated three-stage segmentation approach to segment the nucleus and cytoplasm of each overlapping cell is described. First, superpixel clustering is conducted to segment the image into small coherent clusters that are used to generate a refined superpixel map. The refined superpixel map is passed to an adaptive thresholding step to initially segment the image into cellular clumps and background. Second, a linear classifier with superpixel-based features is designed to finalize the separation between nuclei and cytoplasm. Finally, edge and region based cell segmentation are performed based on edge enhancement process, gradient thresholding, morphological operations, and region properties evaluation on all detected nuclei and cytoplasm pairs. The proposed framework has been evaluated using the ISBI 2014 challenge dataset. The dataset consists of 45 synthetic cell images, yielding 270 cells in total. Compared with the state-of-the-art approaches, our approach provides more accurate nuclei boundaries, as well as successfully segments most of overlapping cells. Afaf Tareef, Yang Song 0001, Tom Weidong Cai, David Dagan Feng |
ICARCV | 3 |
| 2014 | Image noise level estimation based on a new adaptive superpixel classificationabstractAccurate estimation of noise level in images plays an important role in different image processing applications. The current algorithms can precisely estimate noise with smooth images, but it is still the challenge to approximate noise level from richly textured images. In this paper, we proposed a new adaptive superpixel classification algorithm for noise estimation in complicated textured images. Firstly, our new superpixel algorithm adapts the finite Gaussian clustering approach, which can better approximate homogeneous patches in noisy images. Then noise information is obtained locally from each superpixel patch. Finally, the best estimation of noise level is calculated with a statistical approach. Experimental results with various kinds of images demonstrate that our method is more accurate and robust compared to the five existing common used algorithms. Peng Fu 0003, ChangYang Li, Quan-Sen Sun, Tom Weidong Cai, David Dagan Feng |
ICIP | 4 |
| 2014 | Large Margin Aggregation of Local Estimates for Medical Image Classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng |
MICCAI (2) | 2 |
| 2014 | Human Connectome Module Pattern Detection Using a New Multi-graph MinMax Cut Model
De Wang, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
MICCAI (3) | 5 |
| 2014 | Lesion Detection and Characterization With Context Driven Approximation in Thoracic FDG PET-CT Images of NSCLC StudiesabstractWe present a lesion detection and characterization method for (18)F-fluorodeoxyglucose positron emission tomography-computed tomography (FDG PET-CT) images of the thorax in the evaluation of patients with primary nonsmall cell lung cancer (NSCLC) with regional nodal disease. Lesion detection can be difficult due to low contrast between lesions and normal anatomical structures. Lesion characterization is also challenging due to similar spatial characteristics between the lung tumors and abnormal lymph nodes. To tackle these problems, we propose a context driven approximation (CDA) method. There are two main components of our method. First, a sparse representation technique with region-level contexts was designed for lesion detection. To discriminate low-contrast data with sparse representation, we propose a reference consistency constraint and a spatial consistent constraint. Second, a multi-atlas technique with image-level contexts was designed to represent the spatial characteristics for lesion characterization. To accommodate inter-subject variation in a multi-atlas model, we propose an appearance constraint and a similarity constraint. The CDA method is effective with a simple feature set, and does not require parametric modeling of feature space separation. The experiments on a clinical FDG PET-CT dataset show promising performance improvement over the state-of-the-art. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Yun Zhou 0006, Michael J. Fulham, David Dagan Feng |
IEEE Trans. Medical Imaging | 2 |
| 2013 | New Graph Structured Sparsity Model for Multi-label Image AnnotationsabstractIn multi-label image annotations, because each image is associated to multiple categories, the semantic terms (label classes) are not mutually exclusive. Previous research showed that such label correlations can largely boost the annotation accuracy. However, all existing methods only directly apply the label correlation matrix to enhance the label inference and assignment without further learning the structural information among classes. In this paper, we model the label correlations using the relational graph, and propose a novel graph structured sparse learning model to incorporate the topological constraints of relation graph in multi-label classifications. As a result, our new method will capture and utilize the hidden class structures in relational graph to improve the annotation results. In proposed objective, a large number of structured sparsity-inducing norms are utilized, thus the optimization becomes difficult. To solve this problem, we derive an efficient optimization algorithm with proved convergence. We perform extensive experiments on six multi-label image annotation benchmark data sets. In all empirical results, our new method shows better annotation results than the state-of-the-art approaches. Feiping Nie 0001, Tom Weidong Cai, Heng Huang 0001 |
ICCV | 3 |
| 2013 | Heterogeneous Image Features Integration via Multi-modal Semi-supervised Learning ModelabstractAutomatic image categorization has become increasingly important with the development of Internet and the growth in the size of image databases. Although the image categorization can be formulated as a typical multi-class classification problem, two major challenges have been raised by the real-world images. On one hand, though using more labeled training data may improve the prediction performance, obtaining the image labels is a time consuming as well as biased process. On the other hand, more and more visual descriptors have been proposed to describe objects and scenes appearing in images and different features describe different aspects of the visual characteristics. Therefore, how to integrate heterogeneous visual features to do the semi-supervised learning is crucial for categorizing large-scale image data. In this paper, we propose a novel approach to integrate heterogeneous features by performing multi-modal semi-supervised classification on unlabeled as well as unsegmented images. Considering each type of feature as one modality, taking advantage of the large amount of unlabeled data information, our new adaptive multi-modal semi-supervised classification (AMMSS) algorithm learns a commonly shared class indicator matrix and the weights for different modalities (image features) simultaneously. Feiping Nie 0001, Tom Weidong Cai, Heng Huang 0001 |
ICCV | 3 |
| 2013 | Semi-supervised Robust Dictionary Learning via Efficient l-Norms MinimizationabstractRepresenting the raw input of a data set by a set of relevant codes is crucial to many computer vision applications. Due to the intrinsic sparse property of real-world data, dictionary learning, in which the linear decomposition of a data point uses a set of learned dictionary bases, i.e., codes, has demonstrated state-of-the-art performance. However, traditional dictionary learning methods suffer from three weaknesses: sensitivity to noisy and outlier samples, difficulty to determine the optimal dictionary size, and incapability to incorporate supervision information. In this paper, we address these weaknesses by learning a Semi-Supervised Robust Dictionary (SSR-D). Specifically, we use the l2,0+-norm as the loss function to improve the robustness against outliers, and develop a new structured sparse regularization to incorporate the supervision information in dictionary learning, without incurring additional parameters. Moreover, the optimal dictionary size is automatically learned from the input data. Minimizing the derived objective function is challenging because it involves many non-smooth l2,0+-norm terms. We present an efficient algorithm to solve the problem with a rigorous proof of the convergence of the algorithm. Extensive experiments are presented to show the superior performance of the proposed method. Hua Wang 0007, Feiping Nie 0001, Tom Weidong Cai, Heng Huang 0001 |
ICCV | 3 |
| 2013 | A supervised multiview spectral embedding method for neuroimaging classificationabstractThe multi-view/multi-modal features are commonly used in neuroimaging classification because they could provide complementary information to each other and thus result in better classification performance than single-view features. However, it is very challenging to effectively integrate such rich features, since straightforward concatenation or singleview spectral embedding methods rarely leads to physically meaningful integration. In this paper, we present a supervised multi-view/multi-modal spectral embedding method (SMSE) for neuroimaging classification. This method embeds the high dimensional multi-view features derived from multi-modal neuroimaging data into a low dimensional feature space and preserves the optimal local embeddings among different views. The proposed SMSE algorithm, validated using three groups of neuroimaging data, is able to achieve significant classification improvement over the state-of-the-art multi-view spectral embedding methods. Sidong Liu, Lelin Zhang, Tom Weidong Cai, Yang Song 0001, Zhiyong Wang 0001, Lingfeng Wen, David Dagan Feng |
ICIP | 3 |
| 2013 | Graph cuts based relevance feedback in image retrievalabstractRelevance feedback (RF) allows users to be actively involved in the information retrieval process and has been widely used in various information retrieval tasks. While most existing RF methods in content-based image retrieval (CBIR) focus on visual features of individual images only, in this paper we formulate the relevance feedback process as an energy minimization problem. The energy function takes into account both the feature aspect of each image and the manifold structure among individual images. The solution of labelling images as relevant or irrelevant is obtained with the graph cuts method. As a result, our method enables flexibly partitioning the feature space and labelling of images and is capable of handling challenging scenarios (or queries). Experimental results demonstrate that our proposed method outperforms the popular RF methods. Lelin Zhang, Sidong Liu, Zhiyong Wang 0001, Tom Weidong Cai, Yang Song 0001, David Dagan Feng |
ICIP | 4 |
| 2013 | A New Sparse Simplex Model for Brain Anatomical and Genetic Network Analysis
Heng Huang 0001, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001 |
MICCAI (2) | 5 |
| 2013 | Multifold Bayesian Kernelization in Alzheimer's Diagnosis
Sidong Liu, Yang Song 0001, Tom Weidong Cai, Sonia Pujol, Ron Kikinis, Xiaogang Wang 0001, David Dagan Feng |
MICCAI (2) | 3 |
| 2013 | Discriminative Data Transform for Image Feature Extraction and Classification
Yang Song 0001, Tom Weidong Cai, Seungil Huh, Takeo Kanade, Yun Zhou 0006, David Dagan Feng |
MICCAI (2) | 2 |
| 2013 | Similarity Guided Feature Labeling for Lesion Detection
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
MICCAI (1) | 2 |
| 2013 | Region-based progressive localization of cell nuclei in microscopic images with data adaptive modelingabstractBACKGROUND: Segmenting cell nuclei in microscopic images has become one of the most important routines in modern biological applications. With the vast amount of data, automatic localization, i.e. detection and segmentation, of cell nuclei is highly desirable compared to time-consuming manual processes. However, automated segmentation is challenging due to large intensity inhomogeneities in the cell nuclei and the background. RESULTS: We present a new method for automated progressive localization of cell nuclei using data-adaptive models that can better handle the inhomogeneity problem. We perform localization in a three-stage approach: first identify all interest regions with contrast-enhanced salient region detection, then process the clusters to identify true cell nuclei with probability estimation via feature-distance profiles of reference regions, and finally refine the contours of detected regions with regional contrast-based graphical model. The proposed region-based progressive localization (RPL) method is evaluated on three different datasets, with the first two containing grayscale images, and the third one comprising of color images with cytoplasm in addition to cell nuclei. We demonstrate performance improvement over the state-of-the-art. For example, compared to the second best approach, on the first dataset, our method achieves 2.8 and 3.7 reduction in Hausdorff distance and false negatives; on the second dataset that has larger intensity inhomogeneity, our method achieves 5% increase in Dice coefficient and Rand index; on the third dataset, our method achieves 4% increase in object-level accuracy. CONCLUSIONS: To tackle the intensity inhomogeneities in cell nuclei and background, a region-based progressive localization method is proposed for cell nuclei localization in fluorescence microscopy images. The RPL method is demonstrated highly effective on three different public datasets, with on average 3.5% and 7% improvement of region- and contour-based segmentation performance over the state-of-the-art. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
BMC Bioinform. | 2 |
| 2013 | Feature-Based Image Patch Approximation for Lung Tissue ClassificationabstractIn this paper, we propose a new classification method for five categories of lung tissues in high-resolution computed tomography (HRCT) images, with feature-based image patch approximation. We design two new feature descriptors for higher feature descriptiveness, namely the rotation-invariant Gabor-local binary patterns (RGLBP) texture descriptor and multi-coordinate histogram of oriented gradients (MCHOG) gradient descriptor. Together with intensity features, each image patch is then labeled based on its feature approximation from reference image patches. And a new patch-adaptive sparse approximation (PASA) method is designed with the following main components: minimum discrepancy criteria for sparse-based classification, patch-specific adaptation for discriminative approximation, and feature-space weighting for distance computation. The patch-wise labelings are then accumulated as probabilistic estimations for region-level classification. The proposed method is evaluated on a publicly available ILD database, showing encouraging performance improvements over the state-of-the-arts. Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng |
IEEE Trans. Medical Imaging | 2 |
| 2012 | Multiscale and multiorientation feature extraction with degenerative patterns for 3D neuroimaging retrievalabstractAccurate neuroimaging feature extraction is essential for effective content-based management of the large neuroimaging databases, as well as achieving improved diagnosis. In this paper, we presented a multiscale and multi-orientation neuroimaging feature extraction algorithm with degenerative patterns for content-based 3D neuroimaging analysis and retrieval, based on the localized 3D Gabor wavelets. Our proposed approach was evaluated with 209 3D clinical neurological imaging studies and compared with the 3D discrete curvelet transform based method and the 3D spatial grey level co-occurrence matrices based method. The preliminary results suggested that our algorithm could support more reliable 3D neuroimaging retrieval. Sidong Liu, Tom Weidong Cai, Lingfeng Wen, David Dagan Feng |
ICIP | 2 |
| 2012 | Thoracic Abnormality Detection with Data Adaptive Structure Estimation
Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng |
MICCAI (1) | 2 |
| 2012 | A Multistage Discriminative Model for Tumor and Lymph Node Detection in Thoracic ImagesabstractAnalysis of primary lung tumors and disease in regional lymph nodes is important for lung cancer staging, and an automated system that can detect both types of abnormalities will be helpful for clinical routine. In this paper, we present a new method to automatically detect both tumors and abnormal lymph nodes simultaneously from positron emission tomography-computed tomography thoracic images. We perform the detection in a multistage approach, by first detecting all potential abnormalities, then differentiate between tumors and lymph nodes, and finally refine the detected tumors for false positive reduction. Each stage is designed with a discriminative model based on support vector machines and conditional random fields, exploiting intensity, spatial and contextual features. The method is designed to handle a wide and complex variety of abnormal patterns found in clinical datasets, consisting of different spatial contexts of tumors and abnormal lymph nodes. We evaluated the proposed method thoroughly on clinical datasets, and encouraging results were obtained. Yang Song 0001, Tom Weidong Cai, Jinman Kim, David Dagan Feng |
IEEE Trans. Medical Imaging | 2 |
| 2011 | Discriminative Pathological Context Detection in Thoracic Images Based on Multi-level Inference
Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
MICCAI (3) | 2 |
| 2010 | Localized multiscale texture based retrieval of neurological imageabstractThe volume and complexity of neurological images have significantly increased, which leads to challenges in efficient data management and retrieval. In this paper, we developed a new content-based image retrieval framework with the localized multiscale Discrete Curvelet Transform (DCvT) features extracted from parametric neurological images. We also compared the performance of three different irregular-to-regular shape padding methods. 142 patient data with neurodegenerative disorders were used in the evaluation. The preliminary results show that our proposed framework supports fast neuroimaging retrieval, and the orthographic projection method can reduce the computational complexity and has a great potential to improve the retrieval for indefinite cases. Sidong Liu, Tom Weidong Cai, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
CBMS | 3 |
| 2010 | A content-based image retrieval framework for multi-modality lung imagesabstractThis paper presents a framework for effective and fast content-based image retrieval for multi-modality PET-CT lung scans. PET-CT scans present significant advantages in tumor staging, but also place new challenges in computerized image analysis and retrieval. Our framework comprises 5 major components: lung field estimation, texture feature extraction, feature categorization, refinement using SVM, and similarity measure. Clinical data from lung cancer patients are used as case studies, and effective retrieval performance is demonstrated. Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
CBMS | 2 |
| 2010 | 3D neurological image retrieval with localized pathology-centric CMRGlc patternsabstractFunctional neuroimaging has an important role in non-invasive diagnosis of neurodegenerative disorders. There are now large volumes of imaging data generated by functional imaging technologies and so there is a need to efficiently manage and retrieve these data. In this paper, we propose a new scheme for efficient 3D content-based neurological image retrieval. 3D pathology-centric masks were adaptively designed and applied for extracting CMRGlc (cerebral metabolic rate of glucose consumption) texture features with volumetric co-occurrence matrices from neurological FDG PET images. Our results, using 93 clinical dementia studies, show that our approach offers a robust and efficient retrieval mechanism for relevant clinical cases and provides advantages in image data analysis and management. Tom Weidong Cai, Sidong Liu, Lingfeng Wen, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
ICIP | 1 |
| 2010 | Robust, accurate and efficient face recognition from a single training image: A uniform pursuit approach
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 4 |
| 2010 | Emulating biological strategies for uncontrolled face recognition
Weihong Deng, Jiani Hu, Jun Guo 0002, Tom Weidong Cai, David Dagan Feng |
Pattern Recognit. | 4 |
| 2008 | New Block-Based Motion Estimation for Sequences with Brightness Variation and Its Application to Static Sprite Generation for Video CompressionabstractIn this brief, a new local motion estimator is proposed which can accurately estimate motion activities under varying strong brightness conditions. The proposed estimator makes use of a new block division technique which manages practically to get rid of the adverse influence caused by brightness changes between frames. We also propose a new static sprite coding system using the proposed local motion estimator. The system is characterized not only with the features of accurate motion estimation under varying brightness conditions, but also possesses the capability of coding the brightness variability of the background scene using a single layered sprite image. Experimental results show that the resulting static sprite coding system improves the PSNR by 6.32 dB as compared with the conventional static sprite coding system when the background scenes of the video sequences involve strong brightness variations in the spatial and time domains. Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Tom Weidong Cai |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | Real-Time Volume Rendering Visualization of Dual-Modality PET/CT Images With Interactive Fuzzy Thresholding SegmentationabstractThree-dimensional (3-D) visualization has become an essential part for imaging applications, including image-guided surgery, radiotherapy planning, and computer-aided diagnosis. In the visualization of dual-modality positron emission tomography and computed tomography (PET/CT), 3-D volume rendering is often limited to rendering of a single image volume and by high computational demand. Furthermore, incorporation of segmentation in volume rendering is usually restricted to visualizing the presegmented volumes of interest. In this paper, we investigated the integration of interactive segmentation into real-time volume rendering of dual-modality PET/CT images. We present and validate a fuzzy thresholding segmentation technique based on fuzzy cluster analysis, which allows interactive and real-time optimization of the segmentation results. This technique is then incorporated into a real-time multi-volume rendering of PET/CT images. Our method allows a real-time fusion and interchangeability of segmentation volume with PET or CT volumes, as well as the usual fusion of PET/CT volumes. Volume manipulations such as window level adjustments and lookup table can be applied to individual volumes, which are then fused together in real time as adjustments are made. We demonstrate the benefit of our method in integrating segmentation with volume rendering in its application to PET/CT images. Responsive frame rates are achieved by utilizing a texture-based volume rendering algorithm and the rapid transfer capability of the high-memory bandwidth available in low-cost graphic hardware. Jinman Kim, Tom Weidong Cai, Stefan Eberl, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2007 | Fast and Reliable Estimation of Multiple Parametric Images Using an Integrated Method for Dynamic SPECTabstractDynamic single photon emission computed tomography (SPECT) has demonstrated the potential to quantitatively estimate physiological parameters in the brain and the heart. The generalized linear least square (GLLS) method is a well-established method for solving linear compartment models with fast computational speed. However, the high level of noise intrinsic in the SPECT data leads to reliability and instability problems of GLLS for generating parametric images. An integrated method is proposed to restrict the noise in both the temporal and spatial domains to estimate multiple parametric images for dynamic SPECT. This method comprises three steps which are optimum image sampling schedule in the projection space, cluster analysis applied postreconstruction and parametric image generation with GLLS. The simulation and experimental studies for the neuronal nicotine acetylcholine receptor tracer of 5-[123I]-iodo-A-85380 were employed to evaluate the performance of the proposed method. The results of influx rate of K1 and volume of distribution of Vd demonstrated that the integrated method was successful in generating low noise parametric images for high noise SPECT data without enhancing the partial volume effect. Furthermore, the integrated method is computationally efficient for potential clinical applications. Lingfeng Wen, Stefan Eberl, David Dagan Feng, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 4 |
| 2006 | Segmentation of VOI From Multidimensional Dynamic PET Images by Integrating Spatial and Temporal FeaturesabstractSegmentation of multidimensional dynamic positron emission tomography (PET) images into volumes of interest (VOIs) exhibiting similar temporal behavior and spatial features is a challenging task due to inherently poor signal-to-noise ratio and spatial resolution. In this study, we propose VOI segmentation of dynamic PET images by utilizing both the three-dimensional (3-D) spatial and temporal domain information in a hybrid technique that integrates two independent segmentation techniques of cluster analysis and region growing. The proposed technique starts with a cluster analysis that partitions the image based on temporal similarities. The resulting temporal partitions, together with the 3-D spatial information are utilized in the region growing segmentation. The technique was evaluated with dynamic 2-[18F] fluoro-2-deoxy-D-glucose PET simulations and clinical studies of the human brain and compared with the k-means and fuzzy c-means cluster analysis segmentation methods. The quantitative evaluation with simulated images demonstrated that the proposed technique can segment the dynamic PET images into VOIs of different kinetic structures and outperforms the cluster analysis approaches with notable improvements in the smoothness of the segmented VOIs with fewer disconnected or spurious segmentation clusters. In clinical studies, the hybrid technique was only superior to the other techniques in segmenting the white matter. In the gray matter segmentation, the other technique tended to perform slightly better than the hybrid technique, but the differences did not reach significance. The hybrid technique generally formed smoother VOIs with better separation of the background. Overall, the proposed technique demonstrated potential usefulness in the diagnosis and evaluation of dynamic PET neurological imaging studies. Jinman Kim, Tom Weidong Cai, David Dagan Feng, Stefan Eberl |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2006 | A New Way for Multidimensional Medical Data Management: Volume of Interest (VOI)-Based Retrieval of Medical Images With Visual and Functional FeaturesabstractThe advances in digital medical imaging and storage in integrated databases are resulting in growing demands for efficient image retrieval and management. Content-based image retrieval (CBIR) refers to the retrieval of images from a database, using the visual features derived from the information in the image, and has become an attractive approach to managing large medical image archives. In conventional CBIR systems for medical images, images are often segmented into regions which are used to derive two-dimensional visual features for region-based queries. Although such approach has the advantage of including only relevant regions in the formulation of a query, medical images that are inherently multidimensional can potentially benefit from the multidimensional feature extraction which could open up new opportunities in visual feature extraction and retrieval. In this study, we present a volume of interest (VOI) based content-based retrieval of four-dimensional (three spatial and one temporal) dynamic PET images. By segmenting the images into VOIs consisting of functionally similar voxels (e.g., a tumor structure), multidimensional visual and functional features were extracted and used as region-based query features. A prototype VOI-based functional image retrieval system (VOI-FIRS) has been designed to demonstrate the proposed multidimensional feature extraction and retrieval. Experimental results show that the proposed system allows for the retrieval of related images that constitute similar visual and functional VOI features, and can find potential applications in medical data management, such as to aid in education, diagnosis, and statistical analysis. Jinman Kim, Tom Weidong Cai, David Dagan Feng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2002 | Content access and distribution of multimedia medical data in E-healthabstractE-health is greatly impacting on information distribution and availability within the health services, hospitals and to the public. Previous research has addressed the development of system architectures with the aim of integrating the distributed and heterogeneous medical information systems. Easing the difficulties in the sharing and management of multimedia medical data and the timely accessibility to these data are critical needs for health care providers. We have proposed a client-server agent that integrates and allows a portal to every permitted information system of the hospital that consists of picture archiving and communication systems (PACS), radiology information system (RIS) and hospital information system (HIS) via the intranet and the Internet. Our proposed agent enables remote access into the usually closed information system of the hospital and a server that manages all the multimedia medical data and allows for in-depth and complex search queries for content access and automatic creation of patient reports for distribution. Jinman Kim, David Dagan Feng, Tom Weidong Cai, Stefan Eberl |
ICME (2) | 3 |
| 2002 | Dynamic image data compression in the spatial and temporal domains: clinical issues and assessmentabstractIn our previous work, we developed a novel approach to dynamic image data compression, and demonstrated that very high compression ratios can be achieved while preserving relevant kinetic information. However, the technique has not yet been assessed with clinical data. Many issues need to be addressed to tailor the method for clinical use. In this paper, we apply the compression technique to dynamic [18F] 2-fluoro-deoxy-glucose (FDG) brain positron emission tomography (PET) data, using a five-parameter model to include cerebral blood volume (CBV) and partial volume (PV) effects. Functional images generated from the compressed data are compared with those from the original uncompressed data. We show that the storage requirements for a typical clinical dynamic PET image data set can be reduced by more than 95%, without degradation of image quality. Furthermore, the technique greatly reduces the computational complexity of further clinical image postprocessing such as smoothing and generation of functional images. It is expected that the compression technique will be of benefit in image data management and telemedicine. David Dagan Feng, Tom Weidong Cai, Roger R. Fulton |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2000 | Visualization of Biomedical Processes: Local Quantitative Physiological Functions in Living Human BodyabstractFunctional imaging with dynamic positron emission tomography (PET) has been playing a crucial and expanding role in biomedical research and clinical diagnosis, providing image-wide quantitative and qualitative physiological functions in the human body, and supporting visualization of the distribution of these functions corresponding to anatomical structures. A number of parametric imaging algorithms have been developed. We give a brief study on some existing and our recently, developed techniques for generating parametric images. An integrated system for functional image data processing and visualization, and a Web-based application are presented. David Dagan Feng, Tom Weidong Cai |
Computer Graphics International | 2 |
| 2000 | Content-based retrieval of dynamic PET functional imagesabstractThe recent information explosion has led to massively increased demand for multimedia data storage in integrated database systems. Content-based retrieval is an important alternative and complement to traditional keyword-based searching for multimedia data and can greatly enhance information management. However, current content-based image retrieval techniques have some deficiencies when applied in the biomedical functional imaging domain. In this paper, we presented a prototype design for a content-based functional image retrieval database system for dynamic positron emission tomography. The system supports efficient content-based retrieval based on physiological kinetic features and reduces image storage requirements. This design makes it possible to maintain a large number of patient data sets online and to rapidly retrieve dynamic functional image sequences for interpretation and generation of physiological parametric images, and offers potential advantages in medical image data management and telemedicine, as well as providing possible opportunities in the statistical and comparative analysis of functional image data. Tom Weidong Cai, David Dagan Feng, Roger R. Fulton |
IEEE Trans. Inf. Technol. Biomed. | 1 |