EDBT 2026 Demo / reviewers in the wild / expert
Yang Song 0001
dblp:24/4470-1
· DBLP profile ↗
153ranked-venue papers
20as first author
100since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 78 · 9 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 66 · 17 first-author · 37 since 2021Artificial intelligence and machine learning · 64 · 3 first-author · 48 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Principles2Plan: LLM-Guided System for Operationalising Ethical Principles into PlansabstractEthical awareness is critical for robots operating in human environments, yet existing automated planning tools provide little support. Manually specifying ethical rules is labour-intensive and highly context-specific. We present Principles2Plan, an interactive research prototype demonstrating how a human and a Large Language Model (LLM) can collaborate to produce context-sensitive ethical rules and guide automated planning. A domain expert provides the planning domain, problem details, and relevant high-level principles such as beneficence and privacy. The system generates operationalisable ethical rules consistent with these principles, which the user can review, prioritise, and supply to a planner to produce ethically-informed plans. To our knowledge, no prior system supports users in generating principle-grounded rules for classical planning contexts. Principles2Plan showcases the potential of human-LLM collaboration for making ethical automated planning more practical and feasible. Tammy Zhong, Yang Song 0001, Maurice Pagnucco |
AAAI | 2 |
| 2026 | Slide-aware deep feature prompting for enhanced whole slide image classificationabstractThe advent of Whole Slide Imaging (WSI) has revolutionised digital pathology by enabling computational analysis of gigapixel-scale images. To handle their large size, most deep learning models divide WSIs into patches and apply Multiple Instance Learning (MIL) for slide-level classification. However, MIL models often depend on pre-trained feature extractors, resulting in domain gaps between natural and pathological images. Parameter-Efficient Fine-Tuning (PEFT) via visual prompting has emerged to bridge this gap with minimal overhead. Nevertheless, existing visual prompts are typically attached at the image level and tightly coupled with specific architectures such as CNNs or ViTs, limiting generalisability and scalability in WSI tasks. To overcome these limitations, we propose Slide-aware Deep Feature Prompt (S-DFP), a novel visual prompting method which derives task-specific information directly from feature embeddings and is initialised with slide-specific cues, thereby enhancing compatibility with diverse feature extractors and MIL frameworks. Experiments on four benchmark datasets, CAMELYON16, BRIGHT, TCGA-IDH, and UniToPath, demonstrate that S-DFP consistently boosts MIL model performance by 2–5% in AUC while introducing less than 0.02% additional parameters. Furthermore, when integrated with recent pathology foundation models, S-DFP yields additional performance gains. The code is publicly available at S-DFP . Cong Cong 0001, Yang Song 0001, Antonio Di Ieva, Qiangguo Jin, Lei Fan 0007, Angela Chou, Anthony J. Gill, Sidong Liu |
Expert Syst. Appl. | 2 |
| 2026 | Medical hierarchical image classification via dual-geometry image-text learningabstractHierarchical image classification is a fundamental challenge in medical image analysis, as tree-structured taxonomies inherently reflect biological and clinical relationships, spanning the general categorisation of disease entities and fine-grained cellular distinctions. Existing approaches primarily rely on multi-task learning and fine-grained detection, often requiring intricate model design and complex training strategies. In this paper, we aim to exploit the negative curvature property of hyperbolic space, which allows efficient representation of hierarchical structures. We propose a dual-geometry image-text framework, termed H 2 CL. Specifically, we introduce a lightweight classifier head on top of image backbones to extract both Euclidean and hyperbolic features, which are then combined to simultaneously preserve taxonomic consistency from an etiological perspective and enhance instance discrimination from a morphological perspective. Furthermore, a text branch is incorporated to integrate label semantics, where an entailment loss is employed to jointly model image–text alignment and inter-sample relationships. Extensive experiments on cervical cell, skin lesion, and gallbladder disease datasets demonstrate that our framework consistently outperforms advanced methods. Compared to the standard Swin Transformer, H 2 CL achieves an average accuracy improvement of 7% across all three datasets at the fine-grained level, with similarly consistent gains observed when integrated with other backbone models. The source code is publicly available at https://github.com/MCPathology/H2CL . Lei Fan 0007, Arcot Sowmya, Erik Meijering, ZongYuan Ge, Yang Song 0001 |
Medical Image Anal. | 6 |
| 2026 | M 3 Surv : Fusing Multi-slide and Multi-omics for Memory-augmented robust Survival predictionabstractMultimodal survival prediction is crucial for personalized oncology. However, existing methods typically integrate only Formalin-Fixed Paraffin-Embedded (FFPE) slides with a single omics type, such as genomics, overlooking Fresh Frozen (FF) slides that better preserve molecular information, as well as richer multi-omics data like proteomics and transcriptomics. More critically, the complete absence of certain modalities due to clinical constraints ( e.g. , time or cost) severely limits the applicability of conventional fusion models that rely on inter-modality correlations. To address these gaps, we propose M 3 Surv, a framework designed to integrate multi-pathology slides (both FF and FFPE) with multi-omics profiles. For multi-slide fusion, we design a divide-and-conquer hypergraph learning approach to capture both intra-slide higher-order cellular structures and inter-slide relationships, yielding a unified pathology representation. To enrich the biological context, we integrate multi-omics data and employ interactive cross-attention to fuse the pathological and omics modalities. To tackle the missing modality, we introduce a prototype-based memory bank. During training, this memory bank learns and stores representative pathology-omics feature prototypes. At inference, even if a modality is entirely missing, the model can query the bank with available features and robustly impute information from the most similar prototype. Extensive experiments on five TCGA cancer datasets and an in-house dataset demonstrate that M 3 Surv outperforms state-of-the-art methods, achieving an average 2.2% improvement in C-Index. The framework also shows strong stability across various missing modality scenarios, highlighting its clinical potential in real-world, data-incomplete scenarios. Mingcheng Qu, Donglin Di, Yue Gao 0002, Yang Song 0001, Lei Fan 0007 |
Medical Image Anal. | 5 |
| 2026 | STAG: Biologically guided spatial transcriptomics prediction via hypergraph learningabstractSpatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG. Mingcheng Qu, Yuchuan Zhao, Donglin Di, Xiu Su, Hongyan Xu 0002, Yang Song 0001, Lei Fan 0007 |
Medical Image Anal. | 7 |
| 2026 | Curvi-Tracker: Curvilinear structure segmentation refinement by iterative trackingabstract• Curvi-Tracker refines curvilinear structure segmentation using intelligent tracker agents. • Novel Direction-Net and Forward-Net improve connectivity and preserve topology. • Extensive experiments demonstrate effectiveness of the proposed Curvi-Tracker. Curvilinear structures are ubiquitous in various domains, such as blood vessels in medical images or roads in satellite images. The automation of curvilinear structure segmentation is highly beneficial because of the laborious and error-prone process of manual annotation. Existing methods produce segmentation results with decent pixel-level performance, but still with presence of incorrect connectivity. To overcome the challenge, this paper proposes Curvi-Tracker, a novel refinement framework that improves initial coarse segmentation results by deploying tracker agents on detected foreground pixels. The proposed framework has two main components: a Direction-Net and a Forward-Net, which jointly guide the movement of trackers in order to track the curvilinear object. A Direction-Aware Multi-Label loss and a Stepwise Masked loss are proposed for accurate tracking of curvilinear structures. Experiments on public datasets of various curvilinear objects including retinal vessels, roads and pavement cracks demonstrate that the proposed method consistently improves the topological correctness of coarse segmentation results coarse segmentation results, averaging overall 10 % of improvement in all three topological metrics. Zhan Heng, Maurice Pagnucco, Erik Meijering, Yang Song 0001 |
Pattern Recognit. | 4 |
| 2026 | Noise-aware cross attention for image manipulation localizationabstract• A Gated Noise Extractor that dynamically captures noise features from multiple strategies. • Dual-granularity contrastive learning for more discriminative noise extraction. • Noise-domain guided fusion module t • reduce interference from irrelevant in- formation. • An efficient model with low parameter count and computational complexity. Modern image manipulation techniques have achieved visual realism that often deceives the human eye and semantic-based detectors. However, manipulation operations typically disturb the intrinsic statistical properties of images. Unlike high-level semantic content, which remains visually consistent, such disturbances manifest as anomalies in noise characteristics, including inconsistencies in sensor pattern noise, distinct high-frequency residuals, and unnatural frequency-domain artifacts introduced by resampling or synthesis. These subtle forensic cues provide more reliable evidence for manipulation localization but are often suppressed by standard RGB-domain feature extractors. Existing IML methods often rely on a single noise feature extraction strategy or treat all tampering techniques uniformly, leading to two major limitations, incomplete noise characterization and insufficient tampering-type awareness . We propose a Noise-aware Contrastive localization Network (NC-Net), which introduces two key modules. Firstly, a Gated Noise Extractor that captures mixed noise-domain patterns using a gated network combining features derived from BayarConv and Discrete Wavelet Transform (DWT) operations. This extractor is further enhanced by a dual-granularity contrastive learning strategy, which models distributional discrepancies both within images (between manipulated and authentic regions) and across images (among different manipulation types). Secondly, a Multi-Scale Fusion Module that adaptively integrates noise-domain and RGB-domain semantic features via a cross-domain attention mechanism and a top-down feature pyramid. A lightweight decoder then produces the final localization map with high precision. NC-Net enables end-to-end joint optimization of the noise extraction and RGB branches, achieving state-of-the-art performance with competitive computational overhead. Extensive experiments demonstrate its superiority over existing methods. Source code is available at https://github.com/HIT-liar/NC-Net . Hongshi Zhang, Tonghua Su, Fuxiang Yang, Donglin Di, Yang Song 0001, Lei Fan 0007 |
Pattern Recognit. | 6 |
| 2026 | Unifying View-Specific Learning and Ambiguity Awareness With Emotions for Multimodal Misinformation DetectionabstractFake news detection is essential for safeguarding social media users, as the rapid spread of unverified information on these platforms facilitates the easy dissemination of fake news, creating significant risks to society. This underscores the urgent need for automated fake news detection systems that leverage advanced technologies. Current multimodal fake news detection approaches encounter challenges including inconsistent decision-making, prioritizing shared subspace information fusion over view-specific data extraction, and insufficient attention given to analyzing the sentimental elements of the news. To this end, we introduce sentiment-aware fake news detection with view-specific feature extraction (SentiView), which integrates multimodal and view-specific features and authors’ sentiments into a unified framework. The proposed framework comprises four modules: a) modal-specific encoder, which generates encodings from text and image modalities; b) multimodal consistency learning that captures the interrelationships to recognize the inherent ambiguity between the different modalities; c) view-aware feature extractor that captures information from different views within the image modality, integrating an orthogonal constraint within the shared subspace, thus facilitating the utilization of distinctive discriminative details; and d) sentiment extractor that captures the sentiments reflected to extract the underlying intention and potential bias of the news author. Extensive experiments on three publicly available datasets X, Weibo, and GossipCop show that SentiView outperforms state-of-the-art fake news detection approaches with an accuracy of 93.8%, 94.5%, and 87.3%, on the datasets respectively. Marium Malik, Yang Song 0001, Jiaojiao Jiang 0001, Sanjay K. Jha |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | GRPose: Learning Graph Relations for Human Image Generation with Pose PriorsabstractRecent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models. Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001 |
AAAI | 7 |
| 2025 | Structure based SAT dataset for analysing GNN generalisationabstractSatisfiability (SAT) solvers based on techniques such as conflict driven clause learning (CDCL) have produced excellent performance on both synthetic and real world industrial problems. While these CDCL solvers only operate on a per-problem basis, graph neural network (GNN) based solvers bring new benefits to the field by allowing practitioners to exploit knowledge gained from previously solved problems to expedite solving of new SAT problems. However, one specific area that is often studied in the context of CDCL solvers, but largely overlooked in GNN solvers, is the relationship between graph theoretic measure of structure in SAT problems and the generalisation ability of GNN solvers. To bridge the gap between structural graph properties (e.g., modularity, self-similarity) and the generalisability (or lack thereof) of GNN based SAT solvers, we present StructureSAT: a curated dataset, along with code to further generate novel examples, containing a diverse set of SAT problems from well known problem domains. Furthermore, we utilise a novel splitting method that focuses on deconstructing the families into more detailed hierarchies based on their structural properties. With the new dataset, we aim to help explain problematic generalisation in existing GNN SAT solvers by exploiting knowledge of structural graph properties. We conclude with multiple future directions that can help researchers in GNN based SAT solving develop more effective and generalisable SAT solvers. Anthony Tompkins, Yang Song 0001, Maurice Pagnucco |
AISTATS | 3 |
| 2025 | Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation LearningabstractUniversal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H&E enriches H&E-based features with diverse, biologically meaningful information. However, progress is limited by the scarcity of well-aligned multi-stain datasets. Inter-stain Misalignment shifts corresponding tissue across slides, hindering consistent patch-level features and degrading slide-level embeddings. To address this, we curated a slide-level aligned, five-stain dataset (H&E, HER2, KI67, ER, PGR) to enable paired H&E-IHC learning and robust cross-stain representation. Leveraging this dataset, we propose Cross-Stain Contrastive Learning (CSCL), a two-stage pretraining framework: a lightweight adapter trained with patch-wise contrastive alignment to improve the compatibility of H&E features with corresponding IHC-derived contextual cues; and slide-level representation learning with Multiple Instance Learning (MIL), which uses a cross-stain attention fusion module to integrate stain-specific patch features and a crossstain global alignment module to enforce consistency among slide-level embeddings across different stains. Experiments on cancer subtype classification, IHC biomarker status classification, and survival prediction, show consistent gains by yielding high-quality, transferable H&E slide-level representations. The code and data are available at: https://github.com/lily-zyz/CSCL. Yizhi Zhang, Lei Fan 0007, Zhulin Tao, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001 |
BIBM | 5 |
| 2025 | MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny ObjectsabstractWe present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distinct viewpoints to ensure comprehensive object coverage. The text component consists of two subsets: Declarative Knowledge, including 875 words that describe common anomalies across various domains and specific categories, with detailed explanations for ⟨what, why, how⟩, including causes and visual characteristics; and Constructivist Learning, providing 2K multiple-choice questions with varying levels of difficulty, each paired with images and corresponded answer explanations. We also propose a baseline for visual-text tasks and conduct extensive benchmarking experiments to evaluate advanced methods across different settings, highlighting the challenges and efficacy of our dataset. Lei Fan 0007, Dongdong Fan, Zhiguang Hu, Yiwen Ding, Donglin Di, Kai Yi, Maurice Pagnucco, Yang Song 0001 |
CVPR | 8 |
| 2025 | Prototype-Based Image Prompting for Weakly Supervised Histopathological Image SegmentationabstractWeakly supervised image segmentation with image-level labels has drawn attention due to the high cost of pixel-level annotations. Traditional methods using Class Activation Maps (CAMs) often highlight only the most discriminative regions, leading to incomplete masks. Recent approaches that introduce textual information struggle with histopathological images due to inter-class homogeneity and intra-class heterogeneity. In this paper, we propose a prototype-based image prompting framework for histopathological image segmentation. It constructs an image bank from the training set using clustering, extracting multiple prototype features per class to capture intra-class heterogeneity. By designing a matching loss between input features and class-specific prototypes using contrastive learning, our method addresses inter-class homogeneity and guides the model to generate more accurate CAMs. Experiments on four datasets (LUAD-HistoSeg, BCSS-WSSS, GCSS, and BCSS) show that our method outperforms existing weakly supervised segmentation approaches, setting new benchmarks in histopathological image segmentation.1 Qingchen Tang, Lei Fan 0007, Maurice Pagnucco, Yang Song 0001 |
CVPR | 4 |
| 2025 | Interpretable Image Classification via Non-parametric Part Prototype LearningabstractClassifying images with an interpretable decision-making process is a long-standing problem in computer vision. In recent years, Prototypical Part Networks has gained traction as an approach for self-explainable neural networks, due to their ability to mimic human visual reasoning by providing explanations based on prototypical object parts. However, the quality of the explanations generated by these methods leaves room for improvement, as the prototypes usually focus on repetitive and redundant concepts. Leveraging recent advances in prototype learning, we present a framework for part-based interpretable image classification that learns a set of semantically distinctive object parts for each class, and provides diverse and comprehensive explanations. The core of our method is to learn the partprototypes in a non-parametric fashion, through clustering deep features extracted from foundation vision models that encode robust semantic information. To quantitatively evaluate the quality of explanations provided by ProtoPNets, we introduce Distinctiveness Score and Comprehensiveness Score. Through evaluation on CUB-200-2011, Stanford Cars and Stanford Dogs datasets, we show that our framework compares favourably against existing ProtoPNets while achieving better interpretability. Code is available at: https://github.com/zijizhu/protonon-param. Zhijie Zhu, Lei Fan 0007, Maurice Pagnucco, Yang Song 0001 |
CVPR | 4 |
| 2025 | Enhancing Change Detection in Remote Sensing: Integrating Synthetic Data with Semi-Supervised LearningabstractChange detection (CD) in remote sensing is a crucial yet challenging task, particularly due to the labor-intensive nature of labeling bi-temporal images. We introduce a novel framework that leverages synthetic datasets, style transfer, and semi-supervised learning to enhance CD model performance while reducing the dependency on labeled data. Our approach begins with a GAN-based style transfer model that transforms synthetic images to align with real-world scenarios, narrowing the domain gap. These transformed images, combined with a small amount of labeled real data, are used for supervised training to build a robust initial model. We then apply a mean teacher model to integrate unlabeled real images, allowing for effective semi-supervised learning. Our method achieves state-of-the-art performance on the LEVIR and WHU-CD datasets, demonstrating its robustness and accuracy across diverse geographical and temporal conditions. Yafei Luo, Erik Meijering, Yang Song 0001 |
ICASSP | 3 |
| 2025 | Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection
Lei Fan 0007, Donglin Di, Anyang Su, Tianyou Song, Maurice Pagnucco, Yang Song 0001 |
ICCV | 7 |
| 2025 | RipGAN: A GAN-Based Rip Current Data Augmentation MethodabstractRip currents are a major hazard on beaches worldwide, and their strong, offshore-directed currents can place even experienced beachgoers at risk of drowning. While it is intuitive to consider developing an automated rip current detection system to assist lifeguards in protecting beachgoers, rip current detection is in its infancy due to the lack of high-quality large-scale annotated rip current datasets. Also, the collection and annotation of rip current images require expert knowledge, which makes it more difficult to build datasets. So, this paper proposes a GAN-based rip current data augmentation method, RipGAN, to improve the performance of rip current detectors by increasing representative training data. To create new training images, RipGAN, has two branches. One is a texture generator that enriches the pattern and texture details of waves, making the image more realistic. The other is a rip generator based on FFFM-Unet. FFFM (Fast Fourier Fusion Module) uses Fast Fourier convolution to fuse the features from the low and the high layers, so as to further optimise the generated image. Furthermore, we trained Yolov8, YOLOv10, DINO and RT-DETR as rip current detectors to prove the effectiveness of RipGAN. The detectors' rnAP50:95improved by 2.67% on the test set and AP50by 4.93% on real-scene videos, outperforming other data augmentation methods. Besides, abundant ablation studies have been conducted to further evaluate each component of RipGAN. Shenyang Qian, Mitchell Harley, Muhammad Imran Razzak, Yang Song 0001 |
ICRA | 4 |
| 2025 | SynerGuard: A Robust Framework for Point Cloud Classification via Local Geometry and Spatial TopologyabstractPoint cloud recognition models are known to be vulnerable to adversarial attacks. The state-of-the-art defense solutions either focus on partial features of the point cloud, limiting their effectiveness, or rely heavily on known adversarial examples, reducing their generalizability, while others, like point cloud reconstruction, will degrade the classifier's accuracy on clean examples. To address this, we introduce SynerGuard, a novel robust point cloud classification framework mitigating adversarial attacks by considering comprehensive geometric and topological attributes of the point cloud, without relying on known adversarial examples while attaining classification accuracies on clean examples. We comprehensively test SynerGuard against seven attack types from three leading adversarial attack approaches on two widely used datasets, ModelNet40 and ShapeNetPart. The results demonstrate SynERGUARD's superiority against existing defenses in mitigating adversarial attacks, as well as managing clean examples. Haonan Zhong, Maurice Pagnucco, Yang Song 0001 |
ICRA | 4 |
| 2025 | Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality RebalanceabstractMultimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. Furthermore, the disparity in data granularity and dimensionality between pathology and genomics leads to a significant modality imbalance. The high spatial resolution inherent in pathology data renders it a dominant role while overshadowing genomics in multimodal integration. In this paper, we propose a multimodal survival prediction framework that incorporates hypergraph learning to effectively capture both contextual and hierarchical details from pathology images. Moreover, it employs a modality rebalance mechanism and an interactive alignment fusion strategy to dynamically reweight the contributions of the two modalities, thereby mitigating the pathology-genomics imbalance. Quantitative and qualitative experiments are conducted on five TCGA datasets, demonstrating that our model outperforms advanced methods by over 3.4% in C-Index performance. Code: https://github.com/MCPathology/MRePath. Mingcheng Qu, Donglin Di, Tonghua Su, Yue Gao 0002, Yang Song 0001, Lei Fan 0007 |
IJCAI | 6 |
| 2025 | Source-free Few-shot Segmentation for Rarer Brain TumorsabstractSince the inception of BraTS challenge, a series of methods has been developed for brain tumor segmentation over the past years. Although these methods achieved promising results, they mostly focus on glioma segmentation, largely due to their relatively high incidence. These fully-supervised methods may not be applicable as they rely on abundant labeled data, which is intrinsically inaccessible for rarer types of brain tumors. Data-efficient transfer learning approaches like few-shot learning and domain adaptation assume full access to source data, which may not be feasible in real-life scenarios due to privacy and confidentiality concerns. In this work, we propose a new source-free few-shot learning framework for rarer brain tumor segmentation that adapts source model trained on gliomas to other less common brain tumors such as meningioma, metastasis and pediatric tumors with only a few labeled target data. The proposed framework follows a dual-branch prototypes learning structure that harmonize preservation of common knowledge from source class and learning new features from target. We show that our method gains a 6% increase in Dice score over representative source-free domain adaptation methods, and achieves comparable performance against its fully-supervised counterpart. Shenghui Yan, Sidong Liu, Antonio Di Ieva, Maurice Pagnucco, Yang Song 0001 |
IJCNN | 5 |
| 2025 | Location-Aware Parameter Fine-Tuning for Multimodal Image Segmentation
Sicong Gao, Maurice Pagnucco, Yang Song 0001 |
MICCAI (1) | 3 |
| 2025 | Spatially Gene Expression Prediction Using Dual-Scale Contrastive Learning
Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
MICCAI (15) | 6 |
| 2025 | Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
Mingcheng Qu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
MICCAI (10) | 6 |
| 2025 | Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
Junze Wang, Lei Fan 0007, Weipeng Jing 0001, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001 |
MICCAI (11) | 5 |
| 2025 | LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural PlanningabstractWhile large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To tackle these challenges, we introduce LLaPa, a vision-language model framework designed for multimodal procedural planning. LLaPa generates executable action sequences from textual task descriptions and visual environmental images using vision-language models (VLMs). Furthermore, we enhance LLaPa with two auxiliary modules to improve procedural planning. The first module, the Task-Environment Reranker (TER), leverages task-oriented segmentation to create a task-sensitive feature space, aligning textual descriptions with visual environments and emphasizing critical regions for procedural execution. The second module, the Counterfactual Activities Retriever (CAR), identifies and emphasizes potential counterfactual conditions, enhancing the model's reasoning capability in counterfactual scenarios. Extensive experiments on ActPlan-1K and ALFRED benchmarks demonstrate that LLaPa generates higher-quality plans with superior LCS and correctness, outperforming advanced models. The code and models are available https://github.com/sunshibo1234/LLaPa. Shibo Sun, Xue Li 0011, Donglin Di, Lanshun Nie, Weinan Zhang 0003, Dechen Zhan, Yang Song 0001, Lei Fan 0007 |
ACM Multimedia | 8 |
| 2025 | SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware AlignmentabstractWhile Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle with industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model, and dataset are available at https://github.com/amoreZgx1n/SAGE. Guoxin Zang, Xue Li 0011, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song 0001, Lei Fan 0007 |
ACM Multimedia | 6 |
| 2025 | Computational Machine Ethics: A SurveyabstractComputational Machine Ethics (CME) is an interdisciplinary field that integrates moral philosophy into an agent’s decision-making process, contributing to the broader domain of Artificial Intelligence Ethics. Technological advancements have transformed the world, where technology has become an integral part of society, progressively given more autonomy in making judgments within various domains in our lives. Inevitably, issues of ethics come into play in these judgments, making ethical decision-making in machines an increasingly critical problem to solve. This survey provides an overview of CME, highlighting the breadth of directions and the use of techniques within the field. We also provide some background on the ethical dimension before introducing our taxonomy used to categorise and detail the variety of existing approaches from a more technical perspective. Finally, we identify limitations in the research and suggest potential open challenges for future work. Tammy Zhong, Yang Song 0001, Raynaldio Limarga, Maurice Pagnucco |
J. Artif. Intell. Res. | 2 |
| 2025 | TractGraphFormer: Anatomically informed hybrid graph CNN-transformer network for interpretable sex and age prediction from diffusion MRI tractography
Yuqian Chen, Fan Zhang 0013, Leo R. Zekelman, Suheyla Cetin Karayumak, Tengfei Xue, Chaoyi Zhang, Yang Song 0001, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 8 |
| 2025 | Improving cross-domain generalizability of medical image segmentation using uncertainty and shape-aware continual test-time domain adaptationabstractContinual test-time adaptation (CTTA) aims to continuously adapt a source-trained model to a target domain with minimal performance loss while assuming no access to the source data. Typically, source models are trained with empirical risk minimization (ERM) and assumed to perform reasonably on the target domain to allow for further adaptation. However, ERM-trained models often fail to perform adequately on a severely drifted target domain, resulting in unsatisfactory adaptation results. To tackle this issue, we propose a generalizable CTTA framework. First, we incorporate domain-invariant shape modeling into the model and train it using domain-generalization (DG) techniques, promoting target-domain adaptability regardless of the severity of the domain shift. Then, an uncertainty and shape-aware mean teacher network performs adaptation with uncertainty-weighted pseudo-labels and shape information. As part of this process, a novel uncertainty-ranked cross-task regularization scheme is proposed to impose consistency between segmentation maps and their corresponding shape representations, both produced by the student model, at the patch and global levels to enhance performance further. Lastly, small portions of the model's weights are stochastically reset to the initial domain-generalized state at each adaptation step, preventing the model from 'diving too deep' into any specific test samples. The proposed method demonstrates strong continual adaptability and outperforms its peers on five cross-domain segmentation tasks, showcasing its effectiveness and generalizability. Bart Bolsterlee, Yang Song 0001, Erik Meijering |
Medical Image Anal. | 3 |
| 2025 | Multiscope topology learning with conditional updating for airway segmentationabstractAbstract Airway segmentation is essential in computer-assisted diagnosis and screening of bronchial diseases due to the inherent difficulty in obtaining a direct and clear visualization of airway trees from raw CT images. Although medical image segmentation technology is gradually maturing and beginning to be applied in clinics, challenges like breakages and leakages remain in airway segmentation. We propose a novel framework that enhances UNet3D with large-kernel attention for improved global and local feature extraction. A multitask prediction head across voxel, neighborhood, and surface scopes is introduced to better capture airway topology, supervised by customized loss functions. Additionally, a conditional updating strategy leverages a shared encoder and dual decoders to improve segmentation of thin branches by balancing over- and under-segmentation. Specifically, one decoder is optimized with hard examples to encourage over-segmentation, and the other refines results for accurate segmentation using all samples. Our model is quantitatively evaluated on the Binary Airway Segmentation dataset, achieving a Dice score of 0.904, precision of 0.954, tree detection rate of 0.950, and branch detection rate of 0.915, outperforming several recent methods in topological accuracy. In the Airway Tree Modeling Challenge 2022 validation set, our method ranks second overall by a mean position score across all metrics. Our future work aims to enhance prediction confidence and adaptability in ambiguous regions and improve generalizability and interpretability across diverse clinical datasets. Erik Meijering, Yang Song 0001 |
Pattern Anal. Appl. | 3 |
| 2025 | Multi-modal hypergraph contrastive learning for medical image segmentation
Weipeng Jing 0001, Junze Wang, Donglin Di, Yang Song 0001, Lei Fan 0007 |
Pattern Recognit. | 5 |
| 2025 | Learning Frequency-Domain Fusion for Multimodal Remote Sensing Semantic Segmentation
Guangsheng Chen, Fangyu Sun, Weipeng Jing 0001, Weitao Zou, Donglin Di, Yang Song 0001, Lei Fan 0007 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | GrainBrain: Multiview Identification and Stratification of Defective Grain KernelsabstractGrain appearance inspection is crucial for evaluating grain quality and determining seed stratification. Typically, trained inspectors manually examine each grain kernel to identify and remove defective ones, which is time-consuming and error-prone. In this article, we present GrainBrain, a robotic vision-based system comprising a hardware prototype (A100) and a deep learning model (GrainAD). A100 is equipped with five cameras to capture high-quality, multiview images of each kernel. The identification of defective kernels is treated as an unsupervised anomaly detection task. GrainAD trains a classifier to distinguish between healthy and pseudoanomaly samples generated at both image and feature levels, and a supervised contrastive learning loss is employed to obtain compact feature representations of healthy kernels. In addition, we release a large-scale dataset containing over 100K annotated images of four types of cereal grains. Extensive experiments were conducted to verify the superiority of our system, achieving an average AUROC of 94.4/90.4% at the image/pixel level. Our system excelled in both efficiency and consistency, as demonstrated by experiments comparing human experts to the system. Lei Fan 0007, Dongdong Fan, Yiwen Ding, Donglin Di, Maurice Pagnucco, Yang Song 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | Exploring Multi-Feature Relationship in Retinex Decomposition for Low-Light Image EnhancementabstractDespite the recent advancements in deep learning techniques, existing unsupervised low-light image enhancement methods fail to improve global brightness and restore colour due to the lack of high-quality training targets. Moreover, real-world low-light images inevitably contain noise, which significantly reduces image visibility and quality, further complicating the enhancement process. However, current unsupervised approaches tend to oversimplify or ignore the noise in low-light images. To address these issues, we first revise the traditional Retinex decomposition to better integrate with unsupervised deep learning frameworks. Then, we design a Local and Global Illumination-Guided Network for removing corruption from the reflectance component, which improves enhancement quality by not only investigating multi-feature similarity and attention mechanism based on the Retinex theory but also leveraging local details and long-range dependencies. Furthermore, by analysing the attributes of corruption within the reflectance component, we introduce a novel reflectance enhancement loss to effectively remove noise without using ground truth. Ruoyu Guo, Maurice Pagnucco, Yang Song 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Vision-Based Multi-Future Trajectory Prediction: A SurveyabstractVision-based trajectory prediction is an important task that supports safe and intelligent behaviors in autonomous systems. Many advanced approaches have been proposed over the years with improved spatial and temporal feature extraction. However, human behavior is naturally diverse and uncertain. Given the past trajectory and surrounding environment information, an agent can have multiple plausible trajectories in the future. To tackle this problem, an essential task named multi-future trajectory prediction (MTP) has recently been studied. This task aims to generate a diverse, acceptable, and explainable distribution of future predictions for each agent. In this article, we present the first survey for MTP with our unique taxonomies and a comprehensive analysis of frameworks, datasets, and evaluation metrics. We also compare models on existing MTP datasets and conduct experiments on the ForkingPath dataset. Finally, we discuss multiple future directions that can help researchers develop novel MTP systems and other diverse learning tasks similar to MTP. Renhao Huang, Hao Xue 0001, Maurice Pagnucco, Flora D. Salim, Yang Song 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Decoupled Optimisation for Long-Tailed Visual RecognitionabstractWhen training on a long-tailed dataset, conventional learning algorithms tend to exhibit a bias towards classes with a larger sample size. Our investigation has revealed that this biased learning tendency originates from the model parameters, which are trained to disproportionately contribute to the classes characterised by their sample size (e.g., many, medium, and few classes). To balance the overall parameter contribution across all classes, we investigate the importance of each model parameter to the learning of different class groups, and propose a multistage parameter Decouple and Optimisation (DO) framework that decouples parameters into different groups with each group learning a specific portion of classes. To optimise the parameter learning, we apply different training objectives with a collaborative optimisation step to learn complementary information about each class group. Extensive experiments on long-tailed datasets, including CIFAR100, Places-LT, ImageNet-LT, and iNaturaList 2018, show that our framework achieves competitive performance compared to the state-of-the-art. Cong Cong 0001, Shiyu Xuan, Sidong Liu, Shiliang Zhang, Maurice Pagnucco, Yang Song 0001 |
AAAI | 6 |
| 2024 | Boundary-Guided Learning for Gene Expression Prediction in Spatial TranscriptomicsabstractSpatial transcriptomics (ST) has emerged as an advanced technology that provides spatial context to gene expression. Recently, deep learning-based methods have shown the capability to predict gene expression from WSI data using ST data. Existing approaches typically extract features from images and the neighboring regions using pretrained models, and then develop methods to fuse this information to generate the final output. However, these methods often fail to account for the cellular structure similarity, cellular density and the interactions within the microenvironment.In this paper, we propose a framework named BG-TRIPLEX, which leverages boundary information extracted from pathological images as guiding features to enhance gene expression prediction from WSIs. Specifically, our model consists of three branches: the spot, in-context and global branches. In the spot and in-context branches, boundary information, including edge and nuclei characteristics, is extracted using pretrained models. These boundary features guide the learning of cellular morphology and the characteristics of microenvironment through Multi-Head Cross-Attention. Finally, these features are integrated with global features to predict the final output.Extensive experiments were conducted on three public ST datasets. The results demonstrate that our BG-TRIPLEX consistently outperforms existing methods in terms of Pearson Correlation Coefficient (PCC). This method highlights the crucial role of boundary features in understanding the complex interactions between WSI and gene expression, offering a promising direction for future research. Codes are available at: https://github.com/WcloudC0416/BG-TRIPLEX Mingcheng Qu, Yuncong Wu, Donglin Di, Anyang Su, Tonghua Su, Yang Song 0001, Lei Fan 0007 |
BIBM | 6 |
| 2024 | Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image GenerationabstractVanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limbs. Existing methods address this issue mostly by fine-tuning the model with extra images or adding additional control - human-centric priors such as pose or depth maps - during the image generation phase. This paper explores the integration of these human-centric priors directly into the model fine-tuning stage, essentially eliminating the need for extra conditions at the inference stage. We realize this idea by proposing a human-centric alignment loss to strengthen human-related information from the textual prompts within the cross-attention maps. To ensure semantic detail richness and human structural accuracy during fine-tuning, we introduce scale-aware and step-wise constraints within the diffusion process, according to an indepth analysis of the cross-attention layer. Extensive experiments show that our method largely improves over state-of-the-art text-to-image models to synthesize high-quality human images based on user-written prompts. Project page: https://hcplayercvpr2024.github.io. Junyan Wang 0001, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen, Hao Li 0030, Cheng Zhang 0014, Yang Song 0001 |
CVPR | 8 |
| 2024 | Domain Generalised Cell Nuclei Segmentation in Histopathology Images Using Domain-Aware Curriculum Learning and Colour-Perceived Meta LearningabstractCell nuclei segmentation in histopathology images is critical in computer-aided diagnosis and treatment planning. However, this task is challenging due to inherent heterogeneity in histopathology images especially when originating from different domains, caused by variations in imaging protocols, staining techniques, and tissue preparation methods. Such domain shifts can significantly affect segmentation performance when the segmentation model is trained and tested on different domains. In this work, we present a novel gradient-based meta-learning approach for domain generalisation in histopathology cell nuclei segmentation. Specifically, we propose a domain-aware regularisation to correct each pixel’s classification based on the specific domain. We also embed a novel network module to preserve the colour features in histopathology images via an enhanced feature extraction procedure. We demonstrate that our proposed framework can achieve consistent and accurate segmentation performance across domains through extensive experiments on multiple histopathology datasets from diverse sources. Our code is available at: https://github.com/winnie172026/DG. Kunzi Xie, Ruoyu Guo, Cong Cong 0001, Maurice Pagnucco, Yang Song 0001 |
ECAI | 5 |
| 2024 | MFVIEW: Multi-modal Fake News Detection with View-Specific Information Extraction
Marium Malik, Jiaojiao Jiang 0001, Yang Song 0001, Sanjay K. Jha |
ECIR (3) | 3 |
| 2024 | T-JEPA: A Joint-Embedding Predictive Architecture for Trajectory Similarity ComputationabstractTrajectory similarity computation is crucial for analyzing movement patterns in applications like traffic management and wildlife tracking. Recent self-supervised learning methods such as contrastive learning have made advancements in trajectory representation learning but rely on predefined data augmentation schemes, limiting generalized and robust high-level semantic understanding. We introduce T-JEPA, a self-supervised method using Joint-Embedding Predictive Architecture (JEPA) to enhance trajectory representation learning. By sampling and predicting in representation space, T-JEPA infers high-level trajectory semantics without manual intervention. Extensive experiments conducted on three urban and two Foursquare datasets verify the effectiveness of T-JEPA in trajectory similarity computation. Lihuan Li, Hao Xue 0001, Yang Song 0001, Flora D. Salim |
SIGSPATIAL/GIS | 3 |
| 2024 | Enriching Degradation Features for Fundus Image Enhancement via Multi-colour Dynamic Filter Network
Ruoyu Guo, Maurice Pagnucco, Yang Song 0001 |
ICONIP (8) | 3 |
| 2024 | Formalisation and Evaluation of Properties for Consequentialist Machine Ethics
Raynaldio Limarga, Yang Song 0001, Abhaya C. Nayak, David Rajaratnam, Maurice Pagnucco |
IJCAI | 2 |
| 2024 | GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-SpeechabstractThis paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the limitations of current zero-shot speaker adaptive Text-to-Speech (TTS) systems that exhibit poor generalizability in adapting to speakers with accents.Compared to commonly used English corpora, such as LibriTTS and VCTK, GLOBE is unique in its inclusion of utterances from 23,519 speakers and covers 164 accents worldwide, along with detailed metadata for these speakers.Compared to its original corpus, i.e., Common Voice, GLOBE significantly improves the quality of the speech data through rigorous filtering and enhancement processes, while also populating all missing speaker metadata.The final curated GLOBE corpus includes 535 hours of speech data at a 24 kHz sampling rate.Our benchmark results indicate that the speaker adaptive TTS model trained on the GLOBE corpus can synthesize speech with better speaker similarity and comparable naturalness than that trained on other popular corpora.We will release GLOBE publicly after acceptance.The GLOBE dataset is available at https://globecorpus.github.io/. Yang Song 0001, Sanjay K. Jha |
INTERSPEECH | 2 |
| 2024 | Refining Airway Segmentation Through Breakage Filling and Leakage Reduction Using Point CloudsabstractBronchoscopy reveals air passages and internal tissues for accurate diagnosis of various lung diseases. Robot-assisted bronchoscopy using an airway tree model can help path planning before surgery and navigation during surgery. In airway tree modeling, though volumetric deep learning methods have achieved good performance for airway segmentation, it remains a challenge due to the breakages and leakages. Some existing methods adopt post-processing using traditional methods like morphological and fuzzy connected algorithms. Also, some methods convert the volumetric data to point cloud format to refine segmentation. In this paper, we develop a new point cloud-based approach to refine volumetric segmentation. To address the breakage issue, we approach it as a regression problem of the branch extension direction and length. To tackle the leakage issue, we approach it as a segmentation task to eliminate leakages caused by breakage filling and from volumetric segmentation. Moreover, the direction information of branches is crucial for constructing the airway tree while point clouds do not naturally encode it. To introduce this information, we propose a directional feature aggregation, which first decomposes features of neighboring points based on their locations and aggregates decomposed features to aid the network in capturing the directional information effectively. Our proposed model has been evaluated on two public datasets, and the results show that our refinement can improve the volumetric segmentation. Erik Meijering, Yang Song 0001 |
IROS | 3 |
| 2024 | XTranPrune: eXplainability-Aware Transformer Pruning for Bias Mitigation in Dermatological Disease Classification
Ali Ghadiri, Maurice Pagnucco, Yang Song 0001 |
MICCAI (10) | 3 |
| 2024 | Fully Distributed, Flexible Compositional Visual Representations via Soft Tensor ProductsabstractSince the inception of the classicalist vs. connectionist debate, it has been argued that the ability to systematically combine symbol-like entities into compositional representations is crucial for human intelligence. In connectionist systems, the field of disentanglement has gained prominence for its ability to produce explicitly compositional representations; however, it relies on a fundamentally *symbolic, concatenative* representation of compositional structure that clashes with the *continuous, distributed* foundations of deep learning. To resolve this tension, we extend Smolensky's Tensor Product Representation (TPR) and introduce *Soft TPR*, a representational form that encodes compositional structure in an inherently *distributed, flexible* manner, along with *Soft TPR Autoencoder*, a theoretically-principled architecture designed specifically to learn Soft TPRs. Comprehensive evaluations in the visual representation learning domain demonstrate that the Soft TPR framework consistently outperforms conventional disentanglement alternatives -- achieving state-of-the-art disentanglement, boosting representation learner convergence, and delivering superior sample efficiency and low-sample regime performance in downstream tasks. These findings highlight the promise of a *distributed* and *flexible* approach to representing compositional structure by potentially enhancing alignment with the core principles of deep learning over the conventional symbolic approach. Bethia Sun, Maurice Pagnucco, Yang Song 0001 |
NeurIPS | 3 |
| 2024 | Large Language Models for Next Point-of-Interest RecommendationabstractThe next Point of Interest (POI) recommendation task is to predict users' immediate next POI visit given their historical data. Location-Based Social Network (LBSN) data, which is often used for the next POI recommendation task, comes with challenges. One frequently disregarded challenge is how to effectively use the abundant contextual information present in LBSN data. Previous methods are limited by their numerical nature and fail to address this challenge. In this paper, we propose a framework that uses pretrained Large Language Models (LLMs) to tackle this challenge. Our framework allows us to preserve heterogeneous LBSN data in its original format, hence avoiding the loss of contextual information. Furthermore, our framework is capable of comprehending the inherent meaning of contextual information due to the inclusion of commonsense knowledge. In experiments, we test our framework on three real-world LBSN datasets. Our results show that the proposed framework outperforms the state-of-the-art models in all three datasets. Our analysis demonstrates the effectiveness of the proposed framework in using contextual information as well as alleviating the commonly encountered cold-start and short trajectory problems. Peibo Li 0001, Maarten de Rijke, Hao Xue 0001, Shuang Ao, Yang Song 0001, Flora D. Salim |
SIGIR | 5 |
| 2024 | Adaptive unified contrastive learning with graph-based feature aggregator for imbalanced medical image classificationabstractMedical image datasets are often imbalanced due to biases in data collection and limitations in acquiring data for rare conditions. Addressing class imbalance is crucial for developing reliable deep-learning algorithms capable of effectively handling all classes. Recent class imbalanced methods have investigated the effectiveness of self-supervised learning (SSL) and demonstrated that such learned features offer increased resilience to class imbalance issues and obtain much improved performances over other types of class imbalanced methods. However, existing SSL methods either lack end-to-end capabilities or require substantial memory resources, potentially resulting in sub-optimal features and classifiers and limiting their practical usage. Moreover, the conventional pooling operations (e.g., max-pooling, or average-pooling) tend to generate less discriminative features when datasets pose high inter-class similarities. To alleviate the above issues, in this study, we present a novel end-to-end self-supervised learning framework tailored for imbalanced medical image datasets. Our framework constitutes an adaptive contrastive loss that can dynamically adjust the model’s learning focus between feature learning and classifier learning and a feature aggregation mechanism based on Graph Neural Networks to further enhance feature discriminability. We evaluate the effectiveness of our framework on four medical datasets, and the experimental results highlight its superior performance in imbalanced image classification tasks. Cong Cong 0001, Sidong Liu, Priyanka Rana, Maurice Pagnucco, Antonio Di Ieva, Shlomo Berkovsky, Yang Song 0001 |
Expert Syst. Appl. | 7 |
| 2024 | TractGeoNet: A geometric deep learning framework for pointwise analysis of tract microstructure to predict language assessment performance
Yuqian Chen, Leo R. Zekelman, Chaoyi Zhang, Tengfei Xue, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
Medical Image Anal. | 5 |
| 2024 | Multi-degradation-adaptation network for fundus image enhancement with degradation representation learningabstractFundus image quality serves a crucial asset for medical diagnosis and applications. However, such images often suffer degradation during image acquisition where multiple types of degradation can occur in each image. Although recent deep learning based methods have shown promising results in image enhancement, they tend to focus on restoring one aspect of degradation and lack generalisability to multiple modes of degradation. We propose an adaptive image enhancement network that can simultaneously handle a mixture of different degradations. The main contribution of this work is to introduce our Multi-Degradation-Adaptive module which dynamically generates filters for different types of degradation. Moreover, we explore degradation representation learning and propose the degradation representation network and Multi-Degradation-Adaptive discriminator for our accompanying image enhancement network. Experimental results demonstrate that our method outperforms several existing state-of-the-art methods in fundus image enhancement. Code will be available at https://github.com/RuoyuGuo/MDA-Net. Ruoyu Guo, Anthony Tompkins, Maurice Pagnucco, Yang Song 0001 |
Medical Image Anal. | 5 |
| 2024 | USAT: A Universal Speaker-Adaptive Text-to-Speech ApproachabstractConventional text-to-speech (TTS) research has predominantly focused on enhancing the quality of synthesized speech for speakers in the training dataset. The challenge of synthesizing lifelike speech for unseen, out-of-dataset speakers, especially those with limited reference data, remains a significant and unresolved problem. While zero-shot or few-shot speakeradaptive TTS approaches have been explored, they have many limitations. Zero-shot approaches tend to suffer from insufficient generalization performance to reproduce the voice of speakers with heavy accents. While few-shot methods can reproduce highly varying accents, they bring a significant storage burden and the risk of overfitting and catastrophic forgetting. In addition, prior approaches only provide either zero-shot or few-shot adaptation, constraining their utility across varied real-world scenarios with different demands. Besides, most current evaluations of speakeradaptive TTS are conducted only on datasets of native speakers, inadvertently neglecting a vast portion of non-native speakers with diverse accents. Our proposed framework unifies both zeroshot and few-shot speaker adaptation strategies, which we term as “instant” and “fine-grained” adaptations, respectively, based on their merits. To alleviate the insufficient generalization performance observed in zero-shot speaker adaptation, we designed two innovative discriminators and introduced a memory mechanism for the speech decoder. To prevent catastrophic forgetting and reduce storage implications for few-shot speaker adaptation, we designed two adapters and a unique adaptation procedure. Additionally, we introduce a new TTS dataset that encompasses 44,000 English utterances from 134 non-native speakers, capturing a wide array of non-native English accents. This dataset is intended to enhance holistic evaluations of adaptive TTS capabilities. Through comprehensive experiments on multiple datasets comprising both native and non-native speakers, our approach outperforms contemporary methodologies across various subjective and objective metrics. Yang Song 0001, Sanjay K. Jha |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Deconfounding Causal Inference for Zero-Shot Action RecognitionabstractZero-shot action recognition (ZSAR) aims to recognize unseen action categories in the test set without corresponding training examples. Most existing zero-shot methods follow the feature generation framework to transfer knowledge from seen action categories to model the feature distribution of unseen categories. However, due to the complexity and diversity of actions, it remains challenging to generate unseen feature distribution, especially for the cross-dataset scenario when there is a potentially larger domain shift. This article proposes aDeconfoundingCaUSAlGAN (DeCalGAN) for generating unseen action video features with the following technical contributions: 1) Our model unifies compositional ZSAR with traditional visual-semantic models to incorporate local object information with global semantic information for feature generation. 2) A GAN-based architecture is proposed for causal inference and unseen distribution discovery. 3) A deconfounding module is proposed to refine representations of local objects and global semantic information confounder in the training data. Action descriptions and random object features after causal inference are then used to discover unseen distributions of novel actions in different datasets. Our extensive experiments onCross-DatasetZero-ShotActionRecognition (CD-ZSAR) demonstrate substantial improvement over the UCF101 and HMDB51 standard benchmarks for this problem. Junyan Wang 0001, Yiqi Jiang, Yang Long 0001, Xiuyu Sun, Maurice Pagnucco, Yang Song 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Identifying the Defective: Detecting Damaged Grains for Cereal Appearance InspectionabstractCereal grain plays a crucial role in the human diet as a major source of essential nutrients. Grain Appearance Inspection (GAI) serves as an essential process to determine grain quality and facilitate grain circulation and processing. However, GAI is routinely performed manually by inspectors with cumbersome procedures, which poses a significant bottleneck in smart agriculture. In this paper, we endeavor to develop an automated GAI system: AI4GrainInsp. By analyzing the distinctive characteristics of grain kernels, we formulate GAI as a ubiquitous problem: Anomaly Detection (AD), in which healthy and edible kernels are considered normal samples while damaged grains or unknown objects are regarded as anomalies. We further propose an AD model, called AD-GAI, which is trained using only normal samples yet can identify anomalies during inference. Moreover, we customize a prototype device for data acquisition and create a large-scale dataset including 220K high-quality images of wheat and maize kernels. Through extensive experiments, AD-GAI achieves considerable performance in comparison with advanced AD methods, and AI4GrainInsp has highly consistent performance compared to human experts and excels at inspection efficiency over 20× speedup. The dataset, code and models will be released at https://github.com/hellodfan/AI4GrainInsp. Lei Fan 0007, Yiwen Ding, Dongdong Fan, Maurice Pagnucco, Yang Song 0001 |
ECAI | 6 |
| 2023 | Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition
Junyan Wang 0001, Zhenhong Sun, Yichen Qian, Dong Gong, Xiuyu Sun, Ming Lin 0002, Maurice Pagnucco, Yang Song 0001 |
ICLR | 8 |
| 2023 | Community-Aware Federated Video SummarizationabstractVideo summarization aims to extract representative frames to retain high-level information. Increasing concerns about privacy issues have been raised because conventional large-scale training requires users to upload video samples that may inevitably release sensitive information. In this paper, we thoroughly discuss the Federated Video Summarization problem, i.e., how to obtain a robust video summarization model when video data is distributed on private data islands. Our key contribution includes 1) We propose a fundamental Frame-Based aggregation method to video-related tasks, which differs from the sample-based aggregation in conventional FedAvg. 2) To mitigate the heterogeneous distribution due to community diversity, we propose the Community-Aware Clustering Federated Video Summarization Framework (CFed-VS) that clusters clients via a novel data-driven clustering approach. 3) We further tackle the challenging non-IID setting with a proposed Mixture Transformer, which manifests state-of-the-art performance via extensive quantitative and qualitative experiments on TVSum and SumMe datasets. Fan Wan, Junyan Wang 0001, Haoran Duan 0001, Yang Song 0001, Maurice Pagnucco, Yang Long 0001 |
IJCNN | 4 |
| 2023 | Generalizable Zero-Shot Speaker Adaptive Speech Synthesis with Disentangled Representations
Yang Song 0001, Sanjay K. Jha |
INTERSPEECH | 2 |
| 2023 | HyperTraj: Towards Simple and Fast Scene-Compliant Endpoint Conditioned Trajectory PredictionabstractAn important task in trajectory prediction is to model the uncertainty of agents' motions, which requires the system to propose multiple plausible future trajectories for agents based on their pastmovements. Recently, many approaches have been developed following an endpointconditioned deep learning framework by firstly predicting the distribution of endpoints, then sampling endpoints from it and finally completing their waypoints. However, this framework suffers a severe efficiency issue as it needs to repeatedly execute a separate decoder conditioned on multiple sampled endpoints. In this work, we propose a simple and fast endpoint conditioned fully convolutional trajectory prediction framework, called HyperTraj, by using dynamic convolutions to generate multiple trajectories, with the main benefits that (1) our prediction is conditioned on endpoint but takes almost constant time when the number of goals increases and (2) our model benefits from convolutional based predictions, such as the acceptance of various scene sizes and better modeling of agent-scene interactions. In our experiment, our model shows comparable or even better accuracy than our state-of-the-art baselines on SDD and VIRAT datasets with around 84% of acceleration and 90% model weight reduction for waypoint decoding. Renhao Huang, Maurice Pagnucco, Yang Song 0001 |
IROS | 3 |
| 2023 | Uncertainty and Shape-Aware Continual Test-Time Adaptation for Cross-Domain Segmentation of Medical Images
Bart Bolsterlee, Brian V. Y. Chow, Yang Song 0001, Erik Meijering |
MICCAI (3) | 4 |
| 2023 | Draw2Edit: Mask-Free Sketch-Guided Image ManipulationabstractSketch-based image modification is an interactive approach for image editing, where users indicate their intention of modifications in the images by drawing sketches on the input image and then the model generates the modified image based on the input sketch. Existing methods often necessitate specifying the region to be modified through a pixel-level mask, transforming the image modification process into a sketch-based inpainting task. Such approaches, however, present a limitation: the mask can cause loss of essential semantic information, compelling the model to perform restoration rather than editing the image. To address this challenge, we propose a novel mask-free image modification method, named Draw2Edit, which enables direct drawing of sketches and editing of images without pixel-level masks, simplifying the editing process. In addition, we employ the free-form deformation to generate structurally corresponding sketches and training images, effectively addressing the challenge of collecting paired sketches and images for training while enhancing the model's effectiveness for sketch-guided tasks. We evaluate our proposed method on commonly-used sketch-guided inpainting datasets, including CelebA-HQ and Places2, and demonstrate its state-of-the-art performance in both quantitative evaluation and user studies. Our code is available at https://github.com/YiwenXu/Draw2Edit. Ruoyu Guo, Maurice Pagnucco, Yang Song 0001 |
ACM Multimedia | 4 |
| 2023 | Speech-Gesture GAN: Gesture Generation for Robots and Embodied AgentsabstractEmbodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this capability is also required for embodied agents in order to enhance the quality and effectiveness of their interactions with humans. In this paper, we propose a novel framework that can generate sequences of joint angles from the speech text and speech audio utterances. Based on a conditional Generative Adversarial Network (GAN), our proposed neural network model learns the relationships between the co-speech gestures and both semantic and acoustic features from the speech input. In order to train our neural network model, we employ a public dataset containing co-speech gestures with corresponding speech audio utterances, which were captured from a single male native English speaker. The results from both objective and subjective evaluations demonstrate the efficacy of our gesture-generation framework for Robots and Embodied Agents. Carson Yu Liu, Gelareh Mohammadi, Yang Song 0001, Wafa Johal |
RO-MAN | 3 |
| 2023 | Imbalanced classification for protein subcellular localization with multilabel oversamplingabstractMOTIVATION: Subcellular localization of human proteins is essential to comprehend their functions and roles in physiological processes, which in turn helps in diagnostic and prognostic studies of pathological conditions and impacts clinical decision-making. Since proteins reside at multiple locations at the same time and few subcellular locations host far more proteins than other locations, the computational task for their subcellular localization is to train a multilabel classifier while handling data imbalance. In imbalanced data, minority classes are underrepresented, thus leading to a heavy bias towards the majority classes and the degradation of predictive capability for the minority classes. Furthermore, data imbalance in multilabel settings is an even more complex problem due to the coexistence of majority and minority classes. RESULTS: Our studies reveal that based on the extent of concurrence of majority and minority classes, oversampling of minority samples through appropriate data augmentation techniques holds promising scope for boosting the classification performance for the minority classes. We measured the magnitude of data imbalance per class and the concurrence of majority and minority classes in the dataset. Based on the obtained values, we identified minority and medium classes, and a new oversampling method is proposed that includes non-linear mixup, geometric and colour transformations for data augmentation and a sampling approach to prepare minibatches. Performance evaluation on the Human Protein Atlas Kaggle challenge dataset shows that the proposed method is capable of achieving better predictions for minority classes than existing methods. AVAILABILITY AND IMPLEMENTATION: Data used in this study are available at https://www.kaggle.com/competitions/human-protein-atlas-image-classification/data. Source code is available at https://github.com/priyarana/Protein-subcellular-localisation-method. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Priyanka Rana, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
Bioinform. | 4 |
| 2023 | An edge guided coarse-to-fine generative network for image outpainting
Maurice Pagnucco, Yang Song 0001 |
Neurocomputing | 3 |
| 2023 | SAC-Net: Learning with weak and noisy labels in histopathology image segmentation
Ruoyu Guo, Kunzi Xie, Maurice Pagnucco, Yang Song 0001 |
Medical Image Anal. | 4 |
| 2023 | Superficial white matter analysis: An efficient point-cloud-based deep learning framework with supervised contrastive learning for consistent tractography parcellation across populations and dMRI acquisitions
Tengfei Xue, Fan Zhang 0013, Chaoyi Zhang, Yuqian Chen, Yang Song 0001, Alexandra J. Golby, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 5 |
| 2023 | Multi-scale multi-reception attention network for bone age assessment in X-ray images
Zhichao Yang 0003, Cong Cong 0001, Maurice Pagnucco, Yang Song 0001 |
Neural Networks | 4 |
| 2023 | Cancer Survival Prediction From Whole Slide Images With Self-Supervised Learning and Slide ConsistencyabstractHistopathological Whole Slide Images (WSIs) at giga-pixel resolution are the gold standard for cancer analysis and prognosis. Due to the scarcity of pixel- or patch-level annotations of WSIs, many existing methods attempt to predict survival outcomes based on a three-stage strategy that includes patch selection, patch-level feature extraction and aggregation. However, the patch features are usually extracted by using truncated models (e.g. ResNet) pretrained on ImageNet without fine-tuning on WSI tasks, and the aggregation stage does not consider the many-to-one relationship between multiple WSIs and the patient. In this paper, we propose a novel survival prediction framework that consists of patch sampling, feature extraction and patient-level survival prediction. Specifically, we employ two kinds of self-supervised learning methods, i.e. colorization and cross-channel, as pretext tasks to train convnet-based models that are tailored for extracting features from WSIs. Then, at the patient-level survival prediction we explicitly aggregate features from multiple WSIs, using consistency and contrastive losses to normalize slide-level features at the patient level. We conduct extensive experiments on three large-scale datasets: TCGA-GBM, TCGA-LUSC and NLST. Experimental results demonstrate the effectiveness of our proposed framework, as it achieves state-of-the-art performance in comparison with previous studies, with concordance index of 0.670, 0.679 and 0.711 on TCGA-GBM, TCGA-LUSC and NLST, respectively. Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Decompose to Adapt: Cross-Domain Object Detection Via Feature DisentanglementabstractRecent advances in unsupervised domain adaptation (UDA) techniques have witnessed great success in cross-domain computer vision tasks, enhancing the generalization ability of data-driven deep learning architectures by bridging the domain distribution gaps. For the UDA-based cross-domain object detection methods, the majority of them alleviate the domain bias by inducing the domain-invariant feature generation via adversarial learning strategy. However, their domain discriminators have limited classification ability due to the unstable adversarial training process. Therefore, the extracted features induced by them cannot be perfectly domain-invariant and still contain domain-private factors, bringing obstacles to further alleviate the cross-domain discrepancy. To tackle this issue, we design a Domain Disentanglement Faster-RCNN (DDF) to eliminate the source-specific information in the features for detection task learning. Our DDF method facilitates the feature disentanglement at the global and local stages, with a Global Triplet Disentanglement (GTD) module and an Instance Similarity Disentanglement (ISD) module, respectively. By outperforming state-of-the-art methods on four benchmark UDA object detection tasks, our DDF method is demonstrated to be effective with wide applicability. Dongnan Liu, Chaoyi Zhang, Yang Song 0001, Heng Huang 0001, Chenyu Wang 0001, Michael Barnett 0006, Tom Weidong Cai |
IEEE Trans. Multim. | 3 |
| 2023 | InterREC: An Interpretable Method for Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) aims to locate the target object in the image according to a referring expression. This is a challenging task owing to the need for understanding both natural language and visual information and interpretable reasoning between them. Most existing implicit reasoning-based REC methods lack interpretability, while explicit reasoning-based REC methods have lower accuracy. To achieve competitive accuracy while providing adequate interpretability, in this work, we propose a novel explicit reasoning-based method named InterREC. First, in order to address the challenge of multi-modal understanding, we design two neural network modules based on text-image representation learning: a Text-Region Matching Module to align objects in the image and noun phrases in the expression, and a Text-Relation Matching Module to align relations between objects in the image and relational phrases in the expression. Additionally, we design a Reasoning Order Tree for handling complex expressions, which can reduce complex expressions to multiple object-relation-object triplets and therefore identify the inference order and reduce the difficulty of reasoning. At the same time, to achieve an interpretable reasoning step, we design a Bayesian Network-based explicit reasoning method. Based on the comparative evaluation on various datasets, our method achieves higher accuracy than existing explicit reasoning-based REC methods, and the visualization results demonstrate the method's high interpretability. Maurice Pagnucco, Chengpei Xu, Yang Song 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | DHG-GAN: Diverse Image Outpainting via Decoupled High Frequency Semantics
Maurice Pagnucco, Yang Song 0001 |
ACCV (7) | 3 |
| 2022 | Towards Unified Multi-Excitation for Unsupervised Video Prediction
Junyan Wang 0001, Likun Qin, Peng Zhang 0058, Yang Long 0001, Bingzhang Hu, Maurice Pagnucco, Shizheng Wang, Yang Song 0001 |
BMVC | 8 |
| 2022 | GrainSpace: A Large-scale Dataset for Fine-grained and Domain-adaptive Recognition of Cereal GrainsabstractCereal grains are a vital part of human diets and are important commodities for people's livelihood and international trade. Grain Appearance Inspection (GAI) serves as one of the crucial steps for the determination of grain quality and grain stratification for proper circulation, storage and food processing, etc. GAI is routinely performed manually by qualified inspectors with the aid of some hand tools. Automated GAI has the benefit of greatly assisting inspectors with their jobs but has been limited due to the lack of datasets and clear definitions of the tasks. In this paper we formulate GAI as three ubiquitous computer vision tasks: fine-grained recognition, domain adaptation and out-of-distribution recognition. We present a large-scale and publicly available cereal grains dataset called GrainSpace. Specifically, we construct three types of device prototypes for data acquisition, and a total of 5.25 million images determined by professional inspectors. The grain samples including wheat, maize and rice are collected from five countries and more than 30 regions. We also develop a comprehensive benchmark based on semi-supervised learning and self-supervised learning techniques. To the best of our knowledge, GrainSpace is the first publicly released dataset for cereal grain inspection, https://github.com/hellodfan/GrainSpace. Lei Fan 0007, Yiwen Ding, Dongdong Fan, Donglin Di, Maurice Pagnucco, Yang Song 0001 |
CVPR | 6 |
| 2022 | Graph-based Spatial Transformer with Memory Replay for Multi-future Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is an essential and challenging task for a variety of real-life applications such as autonomous driving and robotic motion planning. Besides generating a single future path, predicting multiple plausible future paths is becoming popular in some recent work on trajectory prediction. However, existing methods typically emphasize spatial interactions between pedestrians and surrounding areas but ignore the smoothness and temporal consistency of predictions. Our model aims to forecast multiple paths based on a historical trajectory by modeling multi-scale graph-based spatial transformers combined with a trajectory smoothing algorithm named “Memory Replay” utilizing a memory graph. Our method can comprehensively exploit the spatial information as well as correct the temporally inconsistent trajectories (e.g., sharp turns). We also propose a new evaluation metric named “Percentage of Trajectory Usage” to evaluate the comprehensiveness of diverse multi-future predictions. Our extensive experiments show that the proposed model achieves state-of-the-art performance on multi-future prediction and competitive results for single-future prediction. Code released at https://github.com/Jacobieee/ST-MR. Lihuan Li, Maurice Pagnucco, Yang Song 0001 |
CVPR | 3 |
| 2022 | Channel-Position Self-Attention with Query Refinement Skeleton Graph Neural Network in Human Pose EstimationabstractHuman Pose Estimation (HPE) is a long-standing yet challenging task in computer vision. The nature of the problem requires comprehensive global contextual reasoning among joints in different locations. In this work, we explore how to incorporate two popular and effective concepts, self-attention and Graph Neural Network (GNN), to model long-range information in HPE. Three different ways to implement self-attention in 3D feature maps are studied, where the best result is achieved via the channel-position version. Accuracy is further improved by refining the queries via an efficient channel-wise parallel GNN that explicitly models the human joint graphical relationships. We are able to improve prediction accuracy on strong baseline models and achieve state-of-the-art results. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 3 |
| 2022 | Autolv: Automatic Lecture Video GeneratorabstractWe propose an end-to-end lecture video generation system that can generate realistic and complete lecture videos directly from annotated slides, instructor’s reference voice and instructor’s reference portrait video. Our system is primarily composed of a speech synthesis module with few-shot speaker adaptation and an adversarial learning-based talking-head generation module. It is capable of not only reducing instructors’ workload but also changing the language and accent which can help the students follow the lecture more easily and enable a wider dissemination of lecture contents. Our experimental results show that the proposed model outperforms other current approaches in terms of authenticity, naturalness and accuracy. Here is a video demonstration of how our system works, and the outcomes of the evaluation and comparison: https://youtu.be/cY6TYkI0cog. Yang Song 0001, Sanjay K. Jha |
ICIP | 2 |
| 2022 | CoGNet: Cooperative Graph Neural NetworksabstractGraph representation learning has received increasing attention in recent years for many real-world applications. A major challenge in graph representation learning is the lack of labeled data. To address this challenge, Graph Neural Networks (GNNs) use message passing frameworks to combine information from unlabeled data with labeled data. However, the use of unlabeled data under the message passing framework is indirect in the training process where unlabeled data does not supervise the training process. To fully exploit the potential of unlabeled data, we propose a novel dual-view cooperative training framework for graph data where unlabeled data is involved in the training process for supervision. Specifically, we regard different views as the reasoning processes of two GNN models with which the models make predictions, integrating the understanding of different models on the underlying graph. To exchange information between models, we design a pseudo-label-based approach, where the two models mutually provide pseudo labels to each other iteratively. Moreover, to ensure the quality of pseudo labels, we propose an entropy-based pseudo-labels selection procedure and we adopt GNNExplainer to visualize different views in our framework. Our comprehensive experimental evaluation shows that our methods can boost the performance of state-of-the-art models. Peibo Li 0001, Yixing Yang, Maurice Pagnucco, Yang Song 0001 |
IJCNN | 4 |
| 2022 | White Matter Tracts are Point Clouds: Neuropsychological Score Prediction and Critical Region Localization via Geometric Deep Learning
Yuqian Chen, Fan Zhang 0013, Chaoyi Zhang, Tengfei Xue, Leo R. Zekelman, Jianzhong He 0001, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Lauren O'Donnell |
MICCAI (1) | 7 |
| 2022 | Fast FF-to-FFPE Whole Slide Image Translation via Laplacian Pyramid and Contrastive Learning
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
MICCAI (2) | 4 |
| 2022 | Electron Microscope Image Registration Using Laplacian Sharpening Transformer U-Net
Kunzi Xie, Yixing Yang, Maurice Pagnucco, Yang Song 0001 |
MICCAI (6) | 4 |
| 2022 | Colour adaptive generative networks for stain normalisation of histopathology images
Cong Cong 0001, Sidong Liu, Antonio Di Ieva, Maurice Pagnucco, Shlomo Berkovsky, Yang Song 0001 |
Medical Image Anal. | 6 |
| 2022 | Towards bi-directional skip connections in encoder-decoder architectures and beyond
Tiange Xiang, Chaoyi Zhang, Xinyi Wang 0015, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
Medical Image Anal. | 4 |
| 2022 | Multiple Sclerosis Lesion Analysis in Brain Magnetic Resonance Images: Techniques and Clinical ApplicationsabstractMultiple sclerosis (MS) is a chronic inflammatory and degenerative disease of the central nervous system, characterized by the appearance of focal lesions in the white and gray matter that topographically correlate with an individual patient's neurological symptoms and signs. Magnetic resonance imaging (MRI) provides detailed in-vivo structural information, permitting the quantification and categorization of MS lesions that critically inform disease management. Traditionally, MS lesions have been manually annotated on 2D MRI slices, a process that is inefficient and prone to inter-/intra-observer errors. Recently, automated statistical imaging analysis techniques have been proposed to detect and segment MS lesions based on MRI voxel intensity. However, their effectiveness is limited by the heterogeneity of both MRI data acquisition techniques and the appearance of MS lesions. By learning complex lesion representations directly from images, deep learning techniques have achieved remarkable breakthroughs in the MS lesion segmentation task. Here, we provide a comprehensive review of state-of-the-art automatic statistical and deep-learning MS segmentation methods and discuss current and future clinical applications. Further, we review technical strategies, such as domain adaptation, to enhance MS lesion segmentation in real-world clinical settings. Chaoyi Zhang, Mariano Cabezas, Yang Song 0001, Zihao Tang 0002, Dongnan Liu, Tom Weidong Cai, Michael Barnett 0006, Chenyu Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | DSNet: A Dual-Stream Framework for Weakly-Supervised Gigapixel Pathology Image AnalysisabstractWe present a novel weakly-supervised framework for classifying whole slide images (WSIs). WSIs, due to their gigapixel resolution, are commonly processed by patch-wise classification with patch-level labels. However, patch-level labels require precise annotations, which is expensive and usually unavailable on clinical data. With image-level labels only, patch-wise classification would be sub-optimal due to inconsistency between the patch appearance and image-level label. To address this issue, we posit that WSI analysis can be effectively conducted by integrating information at both high magnification (local) and low magnification (regional) levels. We auto-encode the visual signals in each patch into a latent embedding vector representing local information, and down-sample the raw WSI to hardware-acceptable thumbnails representing regional information. The WSI label is then predicted with a Dual-Stream Network (DSNet), which takes the transformed local patch embeddings and multi-scale thumbnail images as inputs and can be trained by the image-level label only. Experiments conducted on three large-scale public datasets demonstrate that our method outperforms all recent state-of-the-art weakly-supervised WSI classification methods. Tiange Xiang, Yang Song 0001, Chaoyi Zhang, Dongnan Liu, Fan Zhang 0013, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Epistemic Reasoning for Machine Ethics with Situation CalculusabstractWith the rapid development of autonomous machines such as selfdriving vehicles and social robots, there is increasing realisation that machine ethics is important for widespread acceptance of autonomous machines. Our objective is to encode ethical reasoning into autonomous machines following well-defined ethical principles and behavioural norms. We provide an approach to reasoning about actions that incorporates ethical considerations. It builds on Scherl and Levesque's [29, 30] approach to knowledge in the situation calculus. We show how reasoning about knowledge in a dynamic setting can be used to guide ethical and moral choices, aligned with consequentialist and deontological approaches to ethics. We apply our approach to autonomous driving and social robot scenarios, and provide an implementation framework. Maurice Pagnucco, David Rajaratnam, Raynaldio Limarga, Abhaya C. Nayak, Yang Song 0001 |
AIES | 5 |
| 2021 | Dynamic Graph Warping Transformer for Video Alignment
Junyan Wang 0001, Yang Long 0001, Maurice Pagnucco, Yang Song 0001 |
BMVC | 4 |
| 2021 | Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph AnalysisabstractScene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene under-standing via three sequential stages, namely scene graph construction, reasoning, and inference. Within the reasoning stage, an EDGE-oriented Graph Convolutional Network (EdgeGCN) is created to exploit multi-dimensional edge features for explicit relationship modeling, together with the exploration of two associated twinning interaction mechanisms between nodes and edges for the independent evolution of scene graph representations. Overall, our integrated SGGpointframework is established to seek and infer scene structures of interest from both real-world and synthetic 3D point-based scenes. Our experimental results show promising edge-oriented reasoning effects on scene graph generation studies. We also demonstrate our method advantage on several traditional graph representation learning benchmark datasets, including the node-wise classification on citation networks and whole-graph recognition problems for molecular analysis. Chaoyi Zhang, Jianhui Yu, Yang Song 0001, Tom Weidong Cai |
CVPR | 3 |
| 2021 | Speech-based Gesture Generation for Robots and Embodied Agents: A Scoping ReviewabstractHumans use gestures as a means of non-verbal communication. Often accompanying speech, these gestures have several purposes but in general, aim to convey an intended message to the receiver. Researchers have tried to develop systems to allow embodied agents to be better communicators when interacting with humans via using gestures. In this article, we present a scoping literature review of the methods and the metrics used to generate and evaluate co-speech gestures. After collecting a set of papers using a term search on the Scopus database, we analysed the content of these papers based on methodology (i.e., model, the dataset used), evaluation measures (i.e., objective and subjective) and limitations. The results indicate that data-driven approaches are used more frequently. In terms of evaluation measures, we found a trend of combining objective and subjective metrics, while no standards exist for either. This literature review provides an overview of the research in the area and, more specifically insights the trends and the challenges to be met in building a system to automatically generate gestures for embodied agents. Gelareh Mohammadi, Yang Song 0001, Wafa Johal |
HAI | 3 |
| 2021 | Walk in the Cloud: Learning Curves for Point Clouds Shape AnalysisabstractDiscrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequently aggregated back to augment their pointwise features. We provide an effective implementation of the proposed aggregation strategy including a novel curve grouping operator followed by a curve aggregation operator. Our method was benchmarked on several point cloud analysis tasks where we achieved the state-of-the-art classification accuracy of 94.2% on the ModelNet40 classification task, instance IoU of 86.8% on the ShapeNetPart segmentation task and cosine error of 0.11 on the ModelNet40 normal estimation task. Our project page with source code is available at: https://curvenet.github.io/. Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Jianhui Yu, Tom Weidong Cai |
ICCV | 3 |
| 2021 | Iterative Subnetwork With Linear Hierarchical Ordering for Human Pose EstimationabstractHuman pose estimation is a long-standing and challenging problem in computer vision. Many recent advancements in the field have relied on complex structure refinement and specific human joint graphical relations. However, progress has been saturated in terms of accuracy. Each time, new state-of-the-art approaches only improve accuracy by less than 0.3% in the MPII test set despite using complicated model structures. Most recent developments can be summarized into two main ideas: 1) refinement subnetwork to improve predictions iteratively and 2) exploitation of human joint graphical relations. In this work, we present how efficient and simple iterative subnetworks with linear hierarchical ordering based on the aforementioned ideas can help to improve accuracy on strong backbone models. Different versions of iterative subnetwork are examined. Significant improvements on difficult body part predictions such as wrists and ankles using simple convolution subnetwork are observed. Further improvements can be made by using a large receptive field subnetwork such as axial-transformer [1]. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 3 |
| 2021 | Deformable Convolution and Semi-supervised Learning in Point Clouds for Aneurysm Classification and Segmentation
Erik Meijering, Yong Xia 0001, Yang Song 0001 |
ICONIP (6) | 4 |
| 2021 | ICE-GAN: Identity-Aware and Capsule-Enhanced GAN with Graph-Based Reasoning for Micro-Expression Recognition and SynthesisabstractMicro-expressions are reflections of people's true feelings and motives, which attract an increasing number of researchers into the study of automatic facial micro-expression recognition. The short detection window, the subtle facial muscle movements, and the limited training samples make micro-expression recognition challenging. To this end, we propose a novel Identity-aware and Capsule-Enhanced Generative Adversarial Network with graph-based reasoning (ICE-GAN), introducing micro-expression synthesis as an auxiliary task to assist recognition. The generator produces synthetic faces with controllable micro-expressions and identity-aware features, whose long-ranged dependencies are captured through the graph reasoning module (GRM), and the discriminator detects the image authenticity and expression classes. Our ICE-GAN was evaluated on Micro-Expression Grand Challenge 2019 (MEGC2019) with a significant improvement (12.9%) over the winner and surpassed other state-of-the-art methods. Jianhui Yu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
IJCNN | 3 |
| 2021 | Deep Fiber Clustering: Anatomically Informed Unsupervised Deep Learning for Fast and Effective White Matter Parcellation
Yuqian Chen, Chaoyi Zhang, Yang Song 0001, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
MICCAI (7) | 3 |
| 2021 | Semi-supervised Adversarial Learning for Stain Normalisation in Histopathology Images
Cong Cong 0001, Sidong Liu, Antonio Di Ieva, Maurice Pagnucco, Shlomo Berkovsky, Yang Song 0001 |
MICCAI (8) | 6 |
| 2021 | Learning Visual Features by Colorization for Slide-Consistent Survival Prediction from Whole Slide Images
Lei Fan 0007, Arcot Sowmya, Erik Meijering, Yang Song 0001 |
MICCAI (8) | 4 |
| 2021 | Learning with Noise: Mask-Guided Attention Model for Weakly Supervised Nuclei Segmentation
Ruoyu Guo, Maurice Pagnucco, Yang Song 0001 |
MICCAI (2) | 3 |
| 2021 | BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
Xinyi Wang 0015, Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 4 |
| 2021 | Discriminative Latent Semantic Graph for Video CaptioningabstractVideo captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption. Yang Bai 0011, Junyan Wang 0001, Yang Long 0001, Bingzhang Hu, Yang Song 0001, Maurice Pagnucco, Yu Guan 0001 |
ACM Multimedia | 5 |
| 2021 | Panoptic Feature Fusion Net: A Novel Instance Segmentation Paradigm for Biomedical and Biological ImagesabstractInstance segmentation is an important task for biomedical and biological image analysis. Due to the complicated background components, the high variability of object appearances, numerous overlapping objects, and ambiguous object boundaries, this task still remains challenging. Recently, deep learning based methods have been widely employed to solve these problems and can be categorized into proposal-free and proposal-based methods. However, both proposal-free and proposal-based methods suffer from information loss, as they focus on either global-level semantic or local-level instance features. To tackle this issue, we present a Panoptic Feature Fusion Net (PFFNet) that unifies the semantic and instance features in this work. Specifically, our proposed PFFNet contains a residual attention feature fusion mechanism to incorporate the instance prediction with the semantic features, in order to facilitate the semantic contextual information learning in the instance branch. Then, a mask quality sub-branch is designed to align the confidence score of each object with the quality of the mask prediction. Furthermore, a consistency regularization mechanism is designed between the semantic segmentation tasks in the semantic and instance branches, for the robust learning of both tasks. Extensive experiments demonstrate the effectiveness of our proposed PFFNet, which outperforms several state-of-the-art methods on various biomedical and biological datasets. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Image Process. | 3 |
| 2021 | Learning to Recommend With Multiple Cascading BehaviorsabstractMost existing recommender systems leverage user behavior data of one type only, such as the purchase behavior in E-commerce that is directly related to the business Key Performance Indicator (KPI) of conversion rate. Besides the key behavioral data, we argue that other forms of user behaviors also provide valuable signal, such as views, clicks, adding a product to shopping carts and so on. They should be taken into account properly to provide quality recommendation for users. In this work, we contribute a new solution named short for Neural Multi-Task Recommendation (NMTR) for learning recommender systems from user multi-behavior data. We develop a neural network model to capture the complicated and multi-type interactions between users and items. In particular, our model accounts for the cascading relationship among different types of behaviors (e.g., a user must click on a product before purchasing it). To fully exploit the signal in the data of multiple types of behaviors, we perform a joint optimization based on the multi-task learning framework, where the optimization on a behavior is treated as a task. Extensive experiments on two real-world datasets demonstrate that NMTR significantly outperforms state-of-the-art recommender systems that are designed to learn from both single-behavior data and multi-behavior data. Further analysis shows that modeling multiple behaviors is particularly useful for providing recommendation for sparse users that have very few interactions. Chen Gao 0001, Xiangnan He 0001, Dahua Gan, Xiangning Chen, Fuli Feng, Yong Li 0008, Tat-Seng Chua, Lina Yao 0001, Yang Song 0001, Depeng Jin |
IEEE Trans. Knowl. Data Eng. | 9 |
| 2021 | PDAM: A Panoptic-Level Feature Alignment Framework for Unsupervised Domain Adaptive Instance Segmentation in Microscopy ImagesabstractIn this work, we present an unsupervised domain adaptation (UDA) method, named Panoptic Domain Adaptive Mask R-CNN (PDAM), for unsupervised instance segmentation in microscopy images. Since there currently lack methods particularly for UDA instance segmentation, we first design a Domain Adaptive Mask R-CNN (DAM) as the baseline, with cross-domain feature alignment at the image and instance levels. In addition to the image- and instance-level domain discrepancy, there also exists domain bias at the semantic level in the contextual information. Next, we, therefore, design a semantic segmentation branch with a domain discriminator to bridge the domain gap at the contextual level. By integrating the semantic- and instance-level feature adaptation, our method aligns the cross-domain features at the panoptic level. Third, we propose a task re-weighting mechanism to assign trade-off weights for the detection and segmentation loss functions. The task re-weighting mechanism solves the domain bias issue by alleviating the task learning for some iterations when the features contain source-specific factors. Furthermore, we design a feature similarity maximization mechanism to facilitate instance-level feature adaptation from the perspective of representational learning. Different from the typical feature alignment methods, our feature similarity maximization mechanism separates the domain-invariant and domain-specific features by enlarging their feature distribution dependency. Experimental results on three UDA instance segmentation scenarios with five datasets demonstrate the effectiveness of our proposed PDAM method, which outperforms state-of-the-art UDA methods by a large margin. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Shape-Oriented Convolution Neural Network for Point Cloud AnalysisabstractPoint cloud is a principal data structure adopted for 3D geometric information encoding. Unlike other conventional visual data, such as images and videos, these irregular points describe the complex shape features of 3D objects, which makes shape feature learning an essential component of point cloud analysis. To this end, a shape-oriented message passing scheme dubbed ShapeConv is proposed to focus on the representation learning of the underlying shape formed by each local neighboring point. Despite this intra-shape relationship learning, ShapeConv is also designed to incorporate the contextual effects from the inter-shape relationship through capturing the long-ranged dependencies between local underlying shapes. This shape-oriented operator is stacked into our hierarchical learning architecture, namely Shape-Oriented Convolutional Neural Network (SOCNN), developed for point cloud analysis. Extensive experiments have been performed to evaluate its significance in the tasks of point cloud classification and part segmentation. Chaoyi Zhang, Yang Song 0001, Lina Yao 0001, Tom Weidong Cai |
AAAI | 2 |
| 2020 | Unsupervised Instance Segmentation in Microscopy Images via Panoptic Domain Adaptation and Task Re-WeightingabstractUnsupervised domain adaptation (UDA) for nuclei instance segmentation is important for digital pathology, as it alleviates the burden of labor-intensive annotation and domain shift across datasets. In this work, we propose a Cycle Consistency Panoptic Domain Adaptive Mask R-CNN (CyC-PDAM) architecture for unsupervised nuclei segmentation in histopathology images, by learning from fluorescence microscopy images. More specifically, we first propose a nuclei inpainting mechanism to remove the auxiliary generated objects in the synthesized images. Secondly, a semantic branch with a domain discriminator is designed to achieve panoptic-level domain adaptation. Thirdly, in order to avoid the influence of the source-biased features, we propose a task re-weighting mechanism to dynamically add trade-off weights for the task-specific loss functions. Experimental results on three datasets indicate that our proposed method outperforms state-of-the-art UDA methods significantly, and demonstrates a similar performance as fully supervised methods. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Heng Huang 0001, Tom Weidong Cai |
CVPR | 3 |
| 2020 | Towards Enforcing Social Distancing Regulations with Occlusion-Aware Crowd DetectionabstractIn this paper, we present a video analysis method that automatically detects crowds violating social distancing regulations in public spaces, which is widely accepted to be essential to minimise the spreading of COVID-19. While various approaches have been published online to tackle this problem, our work presents a systematic study with comprehensive quantitative analysis of different deep learning models on multiple datasets. We experimented with two types of one-stage pedestrian detection models and further optimised their performance with a repulsion loss to address occlusions in crowds. We also propose a distance computation technique with locally adaptive threshold to approximate the actual spatial distance between pedestrians in the real world. In addition, since there is no existing dataset providing ground truth annotations of distances, we manually annotated three public datasets with such information to perform quantitative evaluation of our crowd detection method. Our comprehensive evaluation shows that our method achieves good detection performance with improvement provided by repulsion loss. Our code and ground truth annotations can be obtained from https://github.com/thomascong121/SocialDistance. Cong Cong 0001, Zhichao Yang 0003, Yang Song 0001, Maurice Pagnucco |
ICARCV | 3 |
| 2020 | Automatic Dropout for Deep Neural Networks
Veena Dodballapur, Rajanish Calisa, Yang Song 0001, Tom Weidong Cai |
ICONIP (3) | 3 |
| 2020 | BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture
Tiange Xiang, Chaoyi Zhang, Dongnan Liu, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 4 |
| 2020 | NFN+: A novel network followed network for retinal vessel segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
Neural Networks | 3 |
| 2020 | 3D APA-Net: 3D Adversarial Pyramid Anisotropic Convolutional Network for Prostate Segmentation in MR ImagesabstractAccurate and reliable segmentation of the prostate gland using magnetic resonance (MR) imaging has critical importance for the diagnosis and treatment of prostate diseases, especially prostate cancer. Although many automated segmentation approaches, including those based on deep learning have been proposed, the segmentation performance still has room for improvement due to the large variability in image appearance, imaging interference, and anisotropic spatial resolution. In this paper, we propose the 3D adversarial pyramid anisotropic convolutional deep neural network (3D APA-Net) for prostate segmentation in MR images. This model is composed of a generator (i.e., 3D PA-Net) that performs image segmentation and a discriminator (i.e., a six-layer convolutional neural network) that differentiates between a segmentation result and its corresponding ground truth. The 3D PA-Net has an encoder-decoder architecture, which consists of a 3D ResNet encoder, an anisotropic convolutional decoder, and multi-level pyramid convolutional skip connections. The anisotropic convolutional blocks can exploit the 3D context information of the MR images with anisotropic resolution, the pyramid convolutional blocks address both voxel classification and gland localization issues, and the adversarial training regularizes 3D PA-Net and thus enables it to generate spatially consistent and continuous segmentation results. We evaluated the proposed 3D APA-Net against several state-of-the-art deep learning-based segmentation approaches on two public databases and the hybrid of the two. Our results suggest that the proposed model outperforms the compared approaches on three databases and could be used in a routine clinical workflow. Haozhe Jia, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Heng Huang 0001, Yanning Zhang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Human Pose Estimation Using Deep Convolutional Densenet Hourglass Network with Intermediate Points VotingabstractHuman pose estimation is a long-standing and challenging problem in computer vision. The problem involves high freedom of articulation of body limbs, different occlusions such as self-occlusion or occlusion by other objects or persons, various clothing, various background in the natural image and foreshortening due to different capturing angle of the camera. In this work we present 1) how the DenseNet module can be used to improve the original ResNet hourglass model, 2) how intermediate points derived from ground truth joint segments can be used as output augmentation of a convolutional neural network (ConvNet) to improve the prediction accuracy. Further improvement has also been made via intermediate points voting by optimizing the joint probability distribution of human joints and the intermediate points. Experimental results on the effects of intermediate point and optimization scheme are presented. We are able to achieve competitive results to the state-of-the-art methods by the proposed method. Shek Wai Chu, Yang Song 0001, Ju Jia Zou, Tom Weidong Cai |
ICIP | 2 |
| 2019 | Nuclei Segmentation via a Deep Panoptic Model with Semantic Feature FusionabstractAutomated detection and segmentation of individual nuclei in histopathology images is important for cancer diagnosis and prognosis. Due to the high variability of nuclei appearances and numerous overlapping objects, this task still remains challenging. Deep learning based semantic and instance segmentation models have been proposed to address the challenges, but these methods tend to concentrate on either the global or local features and hence still suffer from information loss. In this work, we propose a panoptic segmentation model which incorporates an auxiliary semantic segmentation branch with the instance branch to integrate global and local features. Furthermore, we design a feature map fusion mechanism in the instance branch and a new mask generator to prevent information loss. Experimental results on three different histopathology datasets demonstrate that our method outperforms the state-of-the-art nuclei segmentation methods and popular semantic and instance segmentation models by a large margin. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Chaoyi Zhang, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai |
IJCAI | 3 |
| 2019 | HD-Net: Hybrid Discriminative Network for Prostate Segmentation in MR Images
Haozhe Jia, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai, Yong Xia 0001 |
MICCAI (2) | 2 |
| 2019 | Vessel-Net: Retinal Vessel Segmentation Under Multi-path Supervision
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Dongnan Liu, Chaoyi Zhang, Tom Weidong Cai |
MICCAI (1) | 3 |
| 2019 | Knowledge-based Collaborative Deep Learning for Benign-Malignant Lung Nodule Classification on Chest CTabstractThe accurate identification of malignant lung nodules on chest CT is critical for the early detection of lung cancer, which also offers patients the best chance of cure. Deep learning methods have recently been successfully introduced to computer vision problems, although substantial challenges remain in the detection of malignant nodules due to the lack of large training data sets. In this paper, we propose a multi-view knowledge-based collaborative (MV-KBC) deep model to separate malignant from benign nodules using limited chest CT data. Our model learns 3-D lung nodule characteristics by decomposing a 3-D nodule into nine fixed views. For each view, we construct a knowledge-based collaborative (KBC) submodel, where three types of image patches are designed to fine-tune three pre-trained ResNet-50 networks that characterize the nodules' overall appearance, voxel, and shape heterogeneity, respectively. We jointly use the nine KBC submodels to classify lung nodules with an adaptive weighting scheme learned during the error back propagation, which enables the MV-KBC model to be trained in an end-to-end manner. The penalty loss function is used for better reduction of the false negative rate with a minimal effect on the overall performance of the MV-KBC model. We tested our method on the benchmark LIDC-IDRI data set and compared it to the five state-of-the-art classification approaches. Our results show that the MV-KBC model achieved an accuracy of 91.60% for lung nodule classification with an AUC of 95.70%. These results are markedly superior to the state-of-the-art approaches. Yutong Xie 0001, Yong Xia 0001, Yang Song 0001, David Dagan Feng, Michael J. Fulham, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Densely Connected Large Kernel Convolutional Network for Semantic Membrane Segmentation in Microscopy ImagesabstractStructural analysis of neurons can provide valuable insights of brain function. Semantic segmentation of neurons thus becomes an important technique in bioinformatics. Deep learning approaches have shown promising performance in various semantic segmentation problems. However, segmentation of neurons in Electron Microscopy (EM) images has some differences compared with typical segmentation tasks due to the image noise and the disturbance of the intracellular structures. In our work, we propose a network with a ResNet encoder and densely connected decoder with large kernels, and then refinement with simple morphological post-possessing. Two main advantages of our method are: 1) the network can prevent the loss of high-resolution information and enlarge the reception field; 2) the post-processing method is simple and can be directly applied to the probability map from the network to enhance the unconfident area. Evaluated on the ISBI2012 EM membrane segmentation challenge, the proposed method achieves competitive performance. Dongnan Liu, Donghao Zhang 0004, Siqi Liu 0001, Yang Song 0001, Haozhe Jia, David Dagan Feng, Yong Xia 0001, Tom Weidong Cai |
ICIP | 4 |
| 2018 | Whole Slide Image Classification via Iterative Patch LabellingabstractBrain tumor can be a fatal disease in the world. With the aim of improving survival rates, many computerized algorithms have been proposed to assist the pathologists to make a diagnosis' using Whole Slide Pathology Images (WSI). Most methods focus on performing patch-level classification and aggregating the patch-level results to obtain the image classification. Since not all patches carry diagnostic information, it is thus important for our algorithm to recognize discriminative and non-discriminative patches. In this study, we propose an iterative patch labelling algorithm based on the Convolutional Neural Network (CNN), with a well-designed thresholding scheme, a training policy and a novel discriminative model architecture, to distinguish patches and use the discriminative ones to achieve WSI -classification. Our method is evaluated on the MICCAI 2015 Challenge Dataset, and shows a large improvement over the baseline approaches. Chaoyi Zhang, Yang Song 0001, Donghao Zhang 0004, Sidong Liu, Tom Weidong Cai |
ICIP | 2 |
| 2018 | 3D Large Kernel Anisotropic Network for Brain Tumor Segmentation
Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai |
ICONIP (7) | 3 |
| 2018 | Multiscale Network Followed Network Model for Retinal Vessel Segmentation
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Yanning Zhang 0001, Tom Weidong Cai |
MICCAI (2) | 3 |
| 2018 | Panoptic Segmentation with an End-to-End Cell R-CNN for Pathology Image Analysis
Donghao Zhang 0004, Yang Song 0001, Dongnan Liu, Haozhe Jia, Siqi Liu 0001, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (2) | 2 |
| 2018 | Atlas registration and ensemble deep convolutional neural network-based prostate segmentation using magnetic resonance imaging
Haozhe Jia, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 3 |
| 2018 | Merged region based image retrieval
Fanjie Meng, Dalong Shan, Ruixia Shi, Yang Song 0001, Baolong Guo 0001, Tom Weidong Cai |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Locality constrained encoding of frequency and spatial information for image classification
Yongsheng Pan, Yong Xia 0001, Yang Song 0001, Tom Weidong Cai |
Multim. Tools Appl. | 3 |
| 2018 | Automated 3-D Neuron Tracing With Precise Branch Erasing and Confidence Controlled Back TrackingabstractThe automatic reconstruction of single neurons from microscopic images is essential to enable large-scale data-driven investigations in neuron morphology research. However, few previous methods were able to generate satisfactory results automatically from 3-D microscopic images without human intervention. In this paper, we developed a new algorithm for automatic 3-D neuron reconstruction. The main idea of the proposed algorithm is to iteratively track backward from the potential neuronal termini to the soma centre. An online confidence score is computed to decide if a tracing iteration should be stopped and discarded from the final reconstruction. The performance improvements comparing with the previous methods are mainly introduced by a more accurate estimation of the traced area and the confidence controlled back-tracking algorithm. The proposed algorithm supports large-scale batch-processing by requiring only one user specified parameter for background segmentation. We bench tested the proposed algorithm on the images obtained from both the DIADEM challenge and the BigNeuron challenge. Our proposed algorithm achieved the state-of-the-art results. Siqi Liu 0001, Donghao Zhang 0004, Yang Song 0001, Hanchuan Peng, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Multi-Pass Fast Watershed for Accurate Segmentation of Overlapping Cervical CellsabstractThe task of segmenting cell nuclei and cytoplasm in pap smear images is one of the most challenging tasks in automated cervix cytological analysis due to specifically the presence of overlapping cells. This paper introduces a multi-pass fast watershed-based method (MPFW) to segment both nucleus and cytoplasm from large cell masses of overlapping cervical cells in three watershed passes. The first pass locates the nuclei with barrier-based watershed on the gradient-based edge map of a pre-processed image. The next pass segments the isolated, touching, and partially overlapping cells with a watershed transform adapted to the cell shape and location. The final pass introduces mutual iterative watersheds separately applied to each nucleus in the largely overlapping clusters to estimate the cell shape. In MPFW, the line-shaped contours of the watershed cells are deformed with ellipse fitting and contour adjustment to give a better representation of cell shapes. The performance of the proposed method has been evaluated using synthetic, real extended depth-of-field, and multi-layers cervical cytology images provided by the first and second overlapping cervical cytology image segmentation challenges in ISBI 2014 and ISBI 2015. The experimental results demonstrate superior performance of the proposed MPFW in terms of segmentation accuracy, detection rate, and time complexity, compared with recent peer methods. Afaf Tareef, Yang Song 0001, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Locally-Transferred Fisher Vectors for Texture ClassificationabstractTexture classification has been extensively studied in computer vision. Recent research shows that the combination of Fisher vector (FV) encoding and convolutional neural network (CNN) provides significant improvement in texture classification over the previous feature representation methods. However, by truncating the CNN model at the last convolutional layer, the CNN-based FV descriptors would not incorporate the full capability of neural networks in feature learning. In this study, we propose that we can further transform the CNN-based FV descriptors in a neural network model to obtain more discriminative feature representations. In particular, we design a locally-transferred Fisher vector (LFV) method, which involves a multi-layer neural network model containing locally connected layers to transform the input FV descriptors with filters of locally shared weights. The network is optimized based on the hinge loss of classification, and transferred FV descriptors are then used for image classification. Our results on three challenging texture image datasets show improved performance over the state-of-the-art approaches. Yang Song 0001, Fan Zhang 0013, Qing Li 0012, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai |
ICCV | 1 |
| 2017 | Supervised Intra-embedding of Fisher Vectors for Histopathology Image Classification
Yang Song 0001, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
MICCAI (3) | 1 |
| 2017 | Supra-Threshold Fiber Cluster Statistics for Data-Driven Whole Brain Tractography Analysis
Fan Zhang 0013, Weining Wu, Lipeng Ning, Gloria McAnulty, Deborah P. Waber, Borjan A. Gagoski, Kiera Sarill, Hesham M. Hamoda, Yang Song 0001, Tom Weidong Cai, Yogesh Rathi, Lauren O'Donnell |
MICCAI (1) | 9 |
| 2017 | Automatic segmentation of overlapping cervical smear cells based on local distinctive features and guided shape deformation
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Hang Chang, Yue Joseph Wang, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 2 |
| 2017 | Optimizing the cervix cytological examination based on deep learning and dynamic shape modeling
Afaf Tareef, Yang Song 0001, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng, Tom Weidong Cai |
Neurocomputing | 2 |
| 2017 | Dual discriminative local coding for tissue aging analysis
Yang Song 0001, Qing Li 0012, Fan Zhang 0013, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang, Tom Weidong Cai |
Medical Image Anal. | 1 |
| 2017 | Low Dimensional Representation of Fisher Vectors for Microscopy Image ClassificationabstractMicroscopy image classification is important in various biomedical applications, such as cancer subtype identification, and protein localization for high content screening. To achieve automated and effective microscopy image classification, the representative and discriminative capability of image feature descriptors is essential. To this end, in this paper, we propose a new feature representation algorithm to facilitate automated microscopy image classification. In particular, we incorporate Fisher vector (FV) encoding with multiple types of local features that are handcrafted or learned, and we design a separation-guided dimension reduction method to reduce the descriptor dimension while increasing its discriminative capability. Our method is evaluated on four publicly available microscopy image data sets of different imaging types and applications, including the UCSB breast cancer data set, MICCAI 2015 CBTC challenge data set, and IICBU malignant lymphoma, and RNAi data sets. Our experimental results demonstrate the advantage of the proposed low-dimensional FV representation, showing consistent performance improvement over the existing state of the art and the commonly used dimension reduction techniques. Yang Song 0001, Qing Li 0012, Heng Huang 0001, David Dagan Feng, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Bioimage classification with subcategory discriminant transform of high dimensional visual descriptorsabstractBACKGROUND: Bioimage classification is a fundamental problem for many important biological studies that require accurate cell phenotype recognition, subcellular localization, and histopathological classification. In this paper, we present a new bioimage classification method that can be generally applicable to a wide variety of classification problems. We propose to use a high-dimensional multi-modal descriptor that combines multiple texture features. We also design a novel subcategory discriminant transform (SDT) algorithm to further enhance the discriminative power of descriptors by learning convolution kernels to reduce the within-class variation and increase the between-class difference. RESULTS: We evaluate our method on eight different bioimage classification tasks using the publicly available IICBU 2008 database. Each task comprises a separate dataset, and the collection represents typical subcellular, cellular, and tissue level classification problems. Our method demonstrates improved classification accuracy (0.9 to 9%) on six tasks when compared to state-of-the-art approaches. We also find that SDT outperforms the well-known dimension reduction techniques, with for example 0.2 to 13% improvement over linear discriminant analysis. CONCLUSIONS: We present a general bioimage classification method, which comprises a highly descriptive visual feature representation and a learning-based discriminative feature transformation algorithm. Our evaluation on the IICBU 2008 database demonstrates improved performance over the state-of-the-art for six different classification tasks. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, David Dagan Feng, Yue Joseph Wang |
BMC Bioinform. | 1 |
| 2016 | Texture image classification with discriminative neural networksabstractTexture provides an important cue for many computer vision applications, and texture image classification has been an active research area over the past years. Recently, deep learning techniques using convolutional neural networks (CNN) have emerged as the state-of-the-art: CNN-based features provide a significant performance improvement over previous handcrafted features. In this study, we demonstrate that we can further improve the discriminative power of CNN-based features and achieve more accurate classification of texture images. In particular, we have designed a discriminative neural network-based feature transformation (NFT) method, with which the CNN-based features are transformed to lower dimensionality descriptors based on an ensemble of neural networks optimized for the classification objective. For evaluation, we used three standard benchmark datasets (KTH-TIPS2, FMD, and DTD) for texture image classification. Our experimental results show enhanced classification performance over the state-of-the-art. Yang Song 0001, Qing Li 0012, David Dagan Feng, Ju Jia Zou, Tom Weidong Cai |
Comput. Vis. Media | 1 |
| 2016 | Dictionary pruning with visual word significance for medical image retrieval
Fan Zhang 0013, Yang Song 0001, Tom Weidong Cai, Alex Hauptmann 0001, Sidong Liu, Sonia Pujol, Ron Kikinis, Michael J. Fulham, David Dagan Feng |
Neurocomputing | 2 |
| 2015 | Fusing subcategory probabilities for texture classificationabstractTexture, as a fundamental characteristic of objects, has attracted much attention in computer vision research. Performance of texture classification is however still lacking for some challenging cases, largely due to the high intra-class variation and low inter-class distinction. To tackle these issues, in this paper, we propose a sub-categorization model for texture classification. By clustering each class into subcategories, classification probabilities at the subcategory-level are computed based on between-subcategory distinctiveness and within-subcategory representativeness. These subcategory probabilities are then fused based on their contribution levels and cluster qualities. This fused probability is added to the multiclass classification probability to obtain the final class label. Our method was applied to texture classification on three challenging datasets - KTH-TIPS2, FMD and DTD, and has shown excellent performance in comparison with the state-of-the-art approaches. Yang Song 0001, Tom Weidong Cai, Qing Li 0012, Fan Zhang 0013, David Dagan Feng, Heng Huang 0001 |
CVPR | 1 |
| 2015 | Beating cilia identification in fluorescence microscope images for accurate CBF measurementabstractCiliary beating frequency (CBF) is a regulated quantitative measurement to describe ciliary beating properties. It is widely used for diagnosis of defective mucociliary clearance diseases. Image-based methods can be effective for CBF estimation but also affected by the moving objects such as ciliated cells and debris. In this work, we propose a CBF estimation method by removing these unfavorable objects, which we refer to as foreground, so that we can focus on observing the beating cilia only. We firstly design a graph-based method to divide the cilia image into different regions. Next, the foreground regions are extracted and removed from the region division result. The beating cilia are then recognized from the background and used to compute the CBF. Our method conducts the CBF estimation by incorporating the cilia regions only and thus can provide a more accurate description of ciliary beating properties. Preliminary experimental results on cilia images showed the proposed method's potentials for accurate CBF measurement. Fan Zhang 0013, Tom Weidong Cai, Yang Song 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, David Dagan Feng |
ICIP | 3 |
| 2015 | Learning Shape-Driven Segmentation Based on Neural Network and Sparse Reconstruction Toward Automated Cell Analysis of Cervical Smears
Afaf Tareef, Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
ICONIP (1) | 2 |
| 2015 | Motion Representation of Ciliated Cell Images with Contour-Alignment for Automated CBF Estimation
Fan Zhang 0013, Yang Song 0001, Siqi Liu 0001, Paul M. Young, Daniela Traini, Lucy Morgan, Hui-Xin Ong, Lachlan Buddle, Sidong Liu, David Dagan Feng, Tom Weidong Cai |
MICCAI (3) | 2 |
| 2015 | Locality-constrained Subcluster Representation Ensemble for lung image classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, Yue Joseph Wang, David Dagan Feng |
Medical Image Anal. | 1 |
| 2015 | Large Margin Local Estimate With Applications to Medical Image ClassificationabstractMedical images usually exhibit large intra-class variation and inter-class ambiguity in the feature space, which could affect classification accuracy. To tackle this issue, we propose a new Large Margin Local Estimate (LMLE) classification model with sub-categorization based sparse representation. We first sub-categorize the reference sets of different classes into multiple clusters, to reduce feature variation within each subcategory compared to the entire reference set. Local estimates are generated for the test image using sparse representation with reference subcategories as the dictionaries. The similarity between the test image and each class is then computed by fusing the distances with the local estimates in a learning-based large margin aggregation construct to alleviate the problem of inter-class ambiguity. The derived similarities are finally used to determine the class label. We demonstrate that our LMLE model is generally applicable to different imaging modalities, and applied it to three tasks: interstitial lung disease (ILD) classification on high-resolution computed tomography (HRCT) images, phenotype binary classification and continuous regression on brain magnetic resonance (MR) imaging. Our experimental results show statistically significant performance improvements over existing popular classifiers. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng, Yue Joseph Wang, Michael J. Fulham |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Automated three-stage nucleus and cytoplasm segmentation of overlapping cellsabstractDeveloping segmentation techniques for overlapping cells has become a major hurdle for automated analysis of cervical cells. In this paper, an automated three-stage segmentation approach to segment the nucleus and cytoplasm of each overlapping cell is described. First, superpixel clustering is conducted to segment the image into small coherent clusters that are used to generate a refined superpixel map. The refined superpixel map is passed to an adaptive thresholding step to initially segment the image into cellular clumps and background. Second, a linear classifier with superpixel-based features is designed to finalize the separation between nuclei and cytoplasm. Finally, edge and region based cell segmentation are performed based on edge enhancement process, gradient thresholding, morphological operations, and region properties evaluation on all detected nuclei and cytoplasm pairs. The proposed framework has been evaluated using the ISBI 2014 challenge dataset. The dataset consists of 45 synthetic cell images, yielding 270 cells in total. Compared with the state-of-the-art approaches, our approach provides more accurate nuclei boundaries, as well as successfully segments most of overlapping cells. Afaf Tareef, Yang Song 0001, Tom Weidong Cai, David Dagan Feng |
ICARCV | 2 |
| 2014 | Large Margin Aggregation of Local Estimates for Medical Image Classification
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yun Zhou 0006, David Dagan Feng |
MICCAI (2) | 1 |
| 2014 | Lesion Detection and Characterization With Context Driven Approximation in Thoracic FDG PET-CT Images of NSCLC StudiesabstractWe present a lesion detection and characterization method for (18)F-fluorodeoxyglucose positron emission tomography-computed tomography (FDG PET-CT) images of the thorax in the evaluation of patients with primary nonsmall cell lung cancer (NSCLC) with regional nodal disease. Lesion detection can be difficult due to low contrast between lesions and normal anatomical structures. Lesion characterization is also challenging due to similar spatial characteristics between the lung tumors and abnormal lymph nodes. To tackle these problems, we propose a context driven approximation (CDA) method. There are two main components of our method. First, a sparse representation technique with region-level contexts was designed for lesion detection. To discriminate low-contrast data with sparse representation, we propose a reference consistency constraint and a spatial consistent constraint. Second, a multi-atlas technique with image-level contexts was designed to represent the spatial characteristics for lesion characterization. To accommodate inter-subject variation in a multi-atlas model, we propose an appearance constraint and a similarity constraint. The CDA method is effective with a simple feature set, and does not require parametric modeling of feature space separation. The experiments on a clinical FDG PET-CT dataset show promising performance improvement over the state-of-the-art. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Yun Zhou 0006, Michael J. Fulham, David Dagan Feng |
IEEE Trans. Medical Imaging | 1 |
| 2013 | A supervised multiview spectral embedding method for neuroimaging classificationabstractThe multi-view/multi-modal features are commonly used in neuroimaging classification because they could provide complementary information to each other and thus result in better classification performance than single-view features. However, it is very challenging to effectively integrate such rich features, since straightforward concatenation or singleview spectral embedding methods rarely leads to physically meaningful integration. In this paper, we present a supervised multi-view/multi-modal spectral embedding method (SMSE) for neuroimaging classification. This method embeds the high dimensional multi-view features derived from multi-modal neuroimaging data into a low dimensional feature space and preserves the optimal local embeddings among different views. The proposed SMSE algorithm, validated using three groups of neuroimaging data, is able to achieve significant classification improvement over the state-of-the-art multi-view spectral embedding methods. Sidong Liu, Lelin Zhang, Tom Weidong Cai, Yang Song 0001, Zhiyong Wang 0001, Lingfeng Wen, David Dagan Feng |
ICIP | 4 |
| 2013 | Graph cuts based relevance feedback in image retrievalabstractRelevance feedback (RF) allows users to be actively involved in the information retrieval process and has been widely used in various information retrieval tasks. While most existing RF methods in content-based image retrieval (CBIR) focus on visual features of individual images only, in this paper we formulate the relevance feedback process as an energy minimization problem. The energy function takes into account both the feature aspect of each image and the manifold structure among individual images. The solution of labelling images as relevant or irrelevant is obtained with the graph cuts method. As a result, our method enables flexibly partitioning the feature space and labelling of images and is capable of handling challenging scenarios (or queries). Experimental results demonstrate that our proposed method outperforms the popular RF methods. Lelin Zhang, Sidong Liu, Zhiyong Wang 0001, Tom Weidong Cai, Yang Song 0001, David Dagan Feng |
ICIP | 5 |
| 2013 | Multifold Bayesian Kernelization in Alzheimer's Diagnosis
Sidong Liu, Yang Song 0001, Tom Weidong Cai, Sonia Pujol, Ron Kikinis, Xiaogang Wang 0001, David Dagan Feng |
MICCAI (2) | 2 |
| 2013 | Discriminative Data Transform for Image Feature Extraction and Classification
Yang Song 0001, Tom Weidong Cai, Seungil Huh, Takeo Kanade, Yun Zhou 0006, David Dagan Feng |
MICCAI (2) | 1 |
| 2013 | Similarity Guided Feature Labeling for Lesion Detection
Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Xiaogang Wang 0001, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
MICCAI (1) | 1 |
| 2013 | Region-based progressive localization of cell nuclei in microscopic images with data adaptive modelingabstractBACKGROUND: Segmenting cell nuclei in microscopic images has become one of the most important routines in modern biological applications. With the vast amount of data, automatic localization, i.e. detection and segmentation, of cell nuclei is highly desirable compared to time-consuming manual processes. However, automated segmentation is challenging due to large intensity inhomogeneities in the cell nuclei and the background. RESULTS: We present a new method for automated progressive localization of cell nuclei using data-adaptive models that can better handle the inhomogeneity problem. We perform localization in a three-stage approach: first identify all interest regions with contrast-enhanced salient region detection, then process the clusters to identify true cell nuclei with probability estimation via feature-distance profiles of reference regions, and finally refine the contours of detected regions with regional contrast-based graphical model. The proposed region-based progressive localization (RPL) method is evaluated on three different datasets, with the first two containing grayscale images, and the third one comprising of color images with cytoplasm in addition to cell nuclei. We demonstrate performance improvement over the state-of-the-art. For example, compared to the second best approach, on the first dataset, our method achieves 2.8 and 3.7 reduction in Hausdorff distance and false negatives; on the second dataset that has larger intensity inhomogeneity, our method achieves 5% increase in Dice coefficient and Rand index; on the third dataset, our method achieves 4% increase in object-level accuracy. CONCLUSIONS: To tackle the intensity inhomogeneities in cell nuclei and background, a region-based progressive localization method is proposed for cell nuclei localization in fluorescence microscopy images. The RPL method is demonstrated highly effective on three different public datasets, with on average 3.5% and 7% improvement of region- and contour-based segmentation performance over the state-of-the-art. Yang Song 0001, Tom Weidong Cai, Heng Huang 0001, Yue Joseph Wang, David Dagan Feng |
BMC Bioinform. | 1 |
| 2013 | Feature-Based Image Patch Approximation for Lung Tissue ClassificationabstractIn this paper, we propose a new classification method for five categories of lung tissues in high-resolution computed tomography (HRCT) images, with feature-based image patch approximation. We design two new feature descriptors for higher feature descriptiveness, namely the rotation-invariant Gabor-local binary patterns (RGLBP) texture descriptor and multi-coordinate histogram of oriented gradients (MCHOG) gradient descriptor. Together with intensity features, each image patch is then labeled based on its feature approximation from reference image patches. And a new patch-adaptive sparse approximation (PASA) method is designed with the following main components: minimum discrepancy criteria for sparse-based classification, patch-specific adaptation for discriminative approximation, and feature-space weighting for distance computation. The patch-wise labelings are then accumulated as probabilistic estimations for region-level classification. The proposed method is evaluated on a publicly available ILD database, showing encouraging performance improvements over the state-of-the-arts. Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng |
IEEE Trans. Medical Imaging | 1 |
| 2012 | Automatic Adaptation of Software Applications to Database Evolution by Graph Differencing and AOP-Based Dynamic PatchingabstractModern information systems, such as enterprise applications and e-commerce applications, often consist of databases surrounded by a large variety of software applications depending on the databases. During the evolution and deployment of such information systems, developers have to ensure the global consistency between database schemas and surrounding software applications. However, in such situations as Enterprise Application Integration (EAI), databases are shared by a number of software applications contributed by different independent parties, and the developers of those applications often have little or no control on when and how database schema evolves over time. As a result, databases and software applications may not always remain in sync. Such inconsistency may lead to data loss, program failures, or decreased performance. The fundamental challenge in evolving and deploying such database-centric information systems is the fact that databases and their surrounding software applications are subject to independent, asynchronous, and potentially conflicting evolution processes. In this paper, we present an approach to automatically adapting software applications to the evolution of their underlying databases by graph-based schema differencing and aspect-oriented dynamic patching. Our empirical study shows that our approach can automatically adapt software applications to a number of common types of database schema evolution, which accounts for over 87.5% of all schema evolution in the subject system. Our approach allows database schema maintainer to evolve database schema more freely without being afraid of breaking surrounding software applications; it also allows application developers to catch up database schema evolution more quickly without diverting too much from their main business concerns. Yang Song 0001, Xin Peng 0001, Zhenchang Xing, Wenyun Zhao |
COMPSAC | 1 |
| 2012 | Thoracic Abnormality Detection with Data Adaptive Structure Estimation
Yang Song 0001, Tom Weidong Cai, Yun Zhou 0006, David Dagan Feng |
MICCAI (1) | 1 |
| 2012 | A Multistage Discriminative Model for Tumor and Lymph Node Detection in Thoracic ImagesabstractAnalysis of primary lung tumors and disease in regional lymph nodes is important for lung cancer staging, and an automated system that can detect both types of abnormalities will be helpful for clinical routine. In this paper, we present a new method to automatically detect both tumors and abnormal lymph nodes simultaneously from positron emission tomography-computed tomography thoracic images. We perform the detection in a multistage approach, by first detecting all potential abnormalities, then differentiate between tumors and lymph nodes, and finally refine the detected tumors for false positive reduction. Each stage is designed with a discriminative model based on support vector machines and conditional random fields, exploiting intensity, spatial and contextual features. The method is designed to handle a wide and complex variety of abnormal patterns found in clinical datasets, consisting of different spatial contexts of tumors and abnormal lymph nodes. We evaluated the proposed method thoroughly on clinical datasets, and encouraging results were obtained. Yang Song 0001, Tom Weidong Cai, Jinman Kim, David Dagan Feng |
IEEE Trans. Medical Imaging | 1 |
| 2011 | Discriminative Pathological Context Detection in Thoracic Images Based on Multi-level Inference
Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
MICCAI (3) | 1 |
| 2010 | A content-based image retrieval framework for multi-modality lung imagesabstractThis paper presents a framework for effective and fast content-based image retrieval for multi-modality PET-CT lung scans. PET-CT scans present significant advantages in tumor staging, but also place new challenges in computerized image analysis and retrieval. Our framework comprises 5 major components: lung field estimation, texture feature extraction, feature categorization, refinement using SVM, and similarity measure. Clinical data from lung cancer patients are used as case studies, and effective retrieval performance is demonstrated. Yang Song 0001, Tom Weidong Cai, Stefan Eberl, Michael J. Fulham, David Dagan Feng |
CBMS | 1 |