VLDB 2026 Research / reviewers in the wild / expert
Pheng-Ann Heng
dblp:52/2889 · also PhengAnn Heng
· DBLP profile ↗
430ranked-venue papers
1as first author
193since 2021 · last 2026
0000-0003-3055-5034ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 220 · 86 since 2021Applied, interdisciplinary, general and emerging computing · 169 · 1 first-author · 82 since 2021Artificial intelligence and machine learning · 153 · 80 since 2021Systems, architecture and hardware · 22 · 17 since 2021Human-computer interaction and ubiquitous computing · 14Databases, data management, data science and information retrieval · 7 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurgPub-Video: A Comprehensive Surgical Video Framework for Enhanced Surgical Intelligence in Vision-Language ModelabstractVision-Language Models (VLMs) have shown significant potential in surgical scene analysis, yet existing models are limited by frame-level datasets and lack high-quality video data with procedural surgical knowledge. To address these challenges, we make the following contributions: (i) SurgPub-Video, a comprehensive dataset of over 3,000 surgical videos and 25 million annotated frames across 11 specialities, sourced from peer-reviewed clinical journals, (ii) SurgLLaVA-Video, a specialized VLM for surgical video understanding, built upon the TinyLLaVA-Video architecture that supports both video-level and frame-level inputs, and (iii) a video-level surgical Visual Question Answering (VQA) benchmark, covering diverse 11 surgical specialities, such as vascular, cardiology, and thoracic. Extensive experiments, conducted on the proposed benchmark and three additional surgical downstream tasks (action recognition, skill assessment, and triplet recognition), show that SurgLLaVA-Video significantly outperforms both general-purpose and surgical-specific VLMs with only three billion parameters. Yaoqian Li, Xikai Yang, Dunyuan Xu, Litao Zhao, Xiaowei Hu 0001, Jinpeng Li 0004, Pheng-Ann Heng |
AAAI | 8 |
| 2026 | IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
Guibao Shen, Quande Liu, Jialin Gao, Lan Du 0002, Cunjian Chen, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 11 |
| 2026 | MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingabstractMixture-of-experts (MoE) architectures enable scalable Large Language Models (LLMs) with reduced computational overhead, yet their deployment on memory-constrained edge devices is hindered by substantial memory demands. Traditional expert-offloading techniques mitigate memory constraints but often significantly increase inference latency. We introduce MoE-APEX, an Adaptive Precision EXpert offloading system that optimizes MoE inference for edge architectures by dynamically managing expert precision. Our core innovation is to replace less critical cache-miss experts with low-precision variants, reducing loading latency while maintaining accuracy. MoE-APEX introduces three innovative techniques that map the natural hierarchy of MoE computation: (1) a token-level dynamic expert loading mechanism, (2) a layer-level adaptive expert prefetching technique, and (3) a sequence-level cost-aware expert caching policy. These innovations enable MoE-APEX to leverage the benefits of mixed-precision expert inference fully. Implemented atop Llama.cpp, MoE-APEX achieves decoding speedups ranging from 1.34x to 9.75x compared to state-of-the-art MoE offloading systems across diverse edge devices, offering a robust solution for efficient MoE deployment in resource-constrained environments. Jiacheng Liu 0001, Xiaofeng Hou, Yi-Fei Pu, Jing Wang 0055, Pheng-Ann Heng, Chao Li 0009, Minyi Guo |
ASPLOS (2) | 6 |
| 2026 | Tackling missing modalities with memory-efficient modality-complementary prompt learning for robust brain tumor segmentation
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
Expert Syst. Appl. | 4 |
| 2026 | Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
Xiaowei Hu 0001, Zhenghao Xing, Tianyu Wang 0003, Chi-Wing Fu, Pheng-Ann Heng |
Int. J. Comput. Vis. | 5 |
| 2026 | Depth-induced prompt learning for laparoscopic liver landmark detectionabstract• A new liver landmark detection dataset, L3D-2K, comprising 2,000 keyframes sourced from surgical videos with professional annotations. • A novel deep learning framework D2GPLand+ that utilizes RGB-D information for laparoscopic liver landmark detections. • Proposing the DPE module, which incorporates learnable prompts with contrastive learning to discriminate the geometric features of different landmark categories from depth clues. • Introducing the CUMamba block that concurrently conducts cross-modal interactions on spatial dimension and feature reparameterization on channel dimension for effective RGB-D fusion. • Introducing the AFA scheme to highlight anatomical structures by implicit and explicit edge emphasis and controlling detail levels. Laparoscopic liver surgery presents a highly intricate intraoperative environment with significant liver deformation, posing challenges for surgeons in locating critical liver structures. Anatomical liver landmarks can greatly assist surgeons in spatial perception in laparoscopic scenarios and facilitate preoperative-to-intraoperative registration. To advance research in liver landmark detection, we develop a new dataset called L3D-2K , comprising 2,000 keyframes with expert landmark annotations from surgical videos of 47 patients. Accordingly, we propose a baseline, D 2 GPLand+, which effectively leverages depth modality to boost landmark detection performance. Concretely, we introduce a Depth-aware Prompt Embedding (DPE) scheme, which dynamically extracts class-related global geometric cues with the guidance of self-supervised prompts from the SAM encoder. Further, a Cross-dimension Unified Mamba (CUMamba) block is designed to comprehensively incorporate RGB and depth features with the concurrent spatial and channel scanning mechanism. Besides, we bring out an Anatomical Feature Augmentation (AFA) module that captures anatomical cues and emphasizes key structures by optimizing feature granularity. For benchmarking purposes, we evaluate our method and 17 mainstream detection models on L3D, L3D-2K, and P2ILF datasets. Experimental results demonstrate that D 2 GPLand+ obtains superior performance on all three datasets. Our approach provides surgeons with guiding clues that facilitate surgical operations and decision-making in complex laparoscopic surgery. Our code and dataset are available at https://github.com/cuiruize/D2GPLand-Plus . Ruize Cui, Weixin Si, Zhixi Li, Kai Wang 0092, Jialun Pei, Pheng-Ann Heng, Harry Qin |
Medical Image Anal. | 6 |
| 2026 | Advancing radiograph representation learning via cascading graph alignment for vision-language clinical concepts
Xilin Dang, Kang Li 0007, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2026 | HiVLR: Hierarchical Vision-Language Reasoning for interpretable zero-shot radiography image understanding
Xilin Dang, Kang Li 0007, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2026 | LungRes80: Towards tangled surgical workflow recognition in video-assisted thoracoscopic surgeryabstractVideo-Assisted Thoracoscopic Surgery (VATS) is a minimally invasive procedure developed to remove specific lung segments for the treatment of early-stage lung diseases. The surgical procedure involves intricate vascular and bronchial anatomy to preserve as much lung tissue as possible, minimizing impact on the pulmonary function. To assist in monitoring and early warning of this high-risk surgical workflow, we build a new dataset, LungRes80, including 269,806 video frames with phase annotations sampled from 80 VATS cases. LungRes80 presents unique challenges for hierarchical temporal modeling due to diverse short-term transitions between segmentectomy phases and latent long-term causal relations. To this end, we introduce an online baseline model termed LungReco. This framework employs Masked Causal Reasoning (MCR) to perform causal reasoning with semantic modeling from continuously updated memories along with pre-trained Large Language Models (LLMs), and combines it with Concurrent Spatial-Temporal encoding (CoST) for holistic bi-modal co-spatial-temporal aggregation across short- and long-term memories. Furthermore, a new metric, called the Attentional Distraction Coefficient (ADC), is proposed to quantify the costs of intraoperative distraction and postoperative corrections by wrong predictions. We establish a comprehensive benchmark for surgical workflow recognition by evaluating representative models on LungRes80, AutoLaparo, and Cholec80, where our method consistently achieves state-of-the-art performance. Code and data are available at LungRes80. Diandian Guo, Jialun Pei, Jiaao Li, Yanhui Wan, Hao Chen 0011, Pheng-Ann Heng |
Medical Image Anal. | 7 |
| 2026 | SAM-driven cross prompting with adaptive sampling consistency for semi-supervised medical image segmentationabstractSemi-supervised learning (SSL) has achieved notable progress in medical image segmentation. To achieve effective SSL, a model needs to be able to efficiently learn from limited labeled data and effectively exploit knowledge from abundant unlabeled data. Recent developments in visual foundation models, such as the Segment Anything Model (SAM), have demonstrated remarkable adaptability with improved sample efficiency. To seamlessly harness foundation models in SSL, we propose a SAM-driven cross prompting framework with adaptive sampling and prompt consistency for semi-supervised medical image segmentation, named CPAC-SAM. Our method employs SAM's unique prompt design and innovates a cross prompting strategy within a dual-branch framework to automatically generate prompts and supervision across two decoder branches, enabling effective learning from both scarce labeled and valuable unlabeled data. To ensure the quality of prompts for unlabeled data and provide meaningful supervision in the cross prompting scheme, we propose an innovative prototype-guided grid sampling strategy with adaptive intervals to simultaneously improve the reliability of the prompt selection area and ensure both adequate prompt density and complete target coverage. We further design a novel prompt consistency regularization to reduce SAM's prompt sensitivity and to enhance the output invariance under different prompts. We validate our method on five medical image segmentation tasks, encompassing both 2D and 3D scenarios. The extensive experiments with different labeled-data ratios and modalities demonstrate the superiority of our proposed method over the state-of-the-art SSL methods, with more than 4.1% and 3.8% Dice improvement on the breast cancer segmentation task and left atrium segmentation task, respectively. Our code is available at: https://github.com/JuzhengMiao/CPAC-SAM. Juzheng Miao, Cheng Chen 0013, Yuchen Yuan, Quanzheng Li, Pheng-Ann Heng |
Medical Image Anal. | 5 |
| 2026 | Automatic prediction of depth of invasion in oral tongue squamous cell carcinoma using a multimodal regression network fusing prior text and anatomical knowledge
Jiangchang Xu, Weiqing Tang, Pheng-Ann Heng, Xiaojun Chen 0003 |
Medical Image Anal. | 3 |
| 2026 | Dual Domain-Attribute Learning Framework With Asynchronous Adapters for Continual Test-Time AdaptationabstractContinual test-time domain adaptation (CTTA) aims to adapt a pre-trained source model to a stream of continually evolving unlabeled target domains, facilitating model deployment in dynamic and non-stationary environments. Contemporary works usually encode domain-specific (DS) style information in a domain-agnostic manner, synchronizing with the learning of domain-invariant (DI) semantic information. This scheme forces DS information to be optimized using the weights of the previous domain, corrupted by cross-domain discrepancies, and hence leads to error accumulation and catastrophic forgetting issues. Inspired by the Attribute Memory Model (AMM) in brain neuroscience, we propose a dual domain-attribute learning framework based on independent asynchronous updates, aiming to imitate how brain learns new knowledge without forgetting. Concretely, we explicitly decompose the continual adaptation process into two complementary systems: an event-based learning system (ELS) that captures DS style representations and a knowledge-based learning system (KLS) that concentrates on the DI structural characteristics. The ELS first detects differences in the distribution of data streams, and actively builds an adapter pool for new latent domains. The KLS adopts a cross-domain shared adapter emphasizing general knowledge, and cooperates with the adapter from ELS to jointly guide adaptation. To make DS and DI knowledge collaboratively working, we exploit a gradient conflict solver to ease the conflict between the past and current DI knowledge, realizing a win-win game (i.e., no interference adaptation) across evolving domains. Our framework have been extensively evaluated on four benchmarks and outperformed the state-of-the-art approaches on both segmentation and classification CTTA tasks. Yuntong Tian, Kang Li 0007, Tianyang He, Pheng-Ann Heng, Wei Feng 0005 |
IEEE Trans. Image Process. | 5 |
| 2026 | MT-SAM: A Mamba-Transformer Enhanced SAM With Prior-Guided Prompting for Multi-Modal Prostate Cancer DelineationabstractClinically, bi-parametric MRI (bp-MRI), including T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient map, offers essential prior localization of biopsy and focal therapy for suspicious clinically significant prostate cancer (csPCa), and accurate csPCa delineation from bp-MRI is crucial for better outcomes. However, due to the complexity and high variability in appearance, size, shape, and indistinct boundaries, delineating csPCa remains challenging, time-consuming, and heavily relies on the clinician's experience. To address these issues, we propose MT-SAM, a novel framework that enhances SAM with higher-quality feature extraction and a prior-guided automatic prompting strategy. Specifically, we introduce a mamba-transformer network to extract multi-stage multi-modal features from bp-MRI and fuse them into the SAM encoder via cross-mamba modules. Moreover, we propose a prior-guided pyramid-mamba prompting strategy to strengthen the model's attention on the targets. We extensively evaluate our method on both public and private datasets, and the experimental results show that our method achieves up to 5.6-34.1% higher Dice scores than state-of-the-art methods. Code is available at https://github.com/LuckLT/MT-SAM. Litao Zhao, Yuhan Zhang 0001, Libiao Ji, Caizi Li, Chi-Fai Ng, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle ManeuverabstractPringle maneuver (PM) in laparoscopic liver resection aims to reduce blood loss and provide a clear surgical view by intermittently blocking blood inflow of the liver, whereas prolonged PM may cause ischemic injury. To comprehensively monitor this surgical procedure and provide timely warnings of ineffective and prolonged blocking, we suggest two complementary AI-assisted surgical monitoring tasks: workflow recognition and blocking effectiveness detection in liver resections. The former presents challenges in real-time capturing of short-term PM, while the latter involves the intraoperative discrimination of long-term liver ischemia states. To address these challenges, we meticulously collect a novel dataset, called PmLR50, consisting of 25,037 video frames covering various surgical phases from 50 laparoscopic liver resection procedures. Additionally, we develop an online baseline for PmLR50, termed PmNet. This model embraces Masked Temporal Encoding (MTE) and Compressed Sequence Modeling (CSM) for efficient short-term and long-term temporal information modeling, and embeds Contrastive Prototype Separation (CPS) to enhance action discrimination between similar intraoperative operations. Experimental results demonstrate that PmNet outperforms existing state-of-the-art surgical workflow recognition methods on the PmLR50 benchmark. Our research offers potential clinical applications for the laparoscopic liver surgery community. Diandian Guo, Weixin Si, Zhixi Li, Jialun Pei, Pheng-Ann Heng |
AAAI | 5 |
| 2025 | MM-Mixing: Multi-Modal Mixing Alignment for 3D UnderstandingabstractWe introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversity and improving alignment across modalities. Our proposed two-stage training pipeline combines feature-level and input-level mixing to optimize the 3D encoder. The first stage employs feature-level mixing with contrastive learning to align 3D features with their corresponding modalities. The second stage incorporates both feature-level and input-level mixing, introducing mixed point cloud inputs to further refine 3D feature representations. MM-Mixing enhances intermodality relationships, promotes generalization, and ensures feature consistency while providing diverse and realistic training samples. We demonstrate that MM-Mixing significantly improves baseline performance across various learning scenarios, including zero-shot 3D classification, linear probing 3D classification, and cross-modal 3D shape retrieval. Notably, we improved the zero-shot classification accuracy on ScanObjectNN from 51.3% to 61.9%, and on Objaverse-LVIS from 46.8% to 51.4%. Our findings highlight the potential of multi-modal mixing-based alignment to significantly advance 3D object recognition and understanding while remaining straightforward to implement and integrate into existing frameworks. Renrui Zhang, Guangyong Chen, Anfeng Liu, Pheng-Ann Heng |
AAAI | 8 |
| 2025 | UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose EstimationabstractEstimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when applied to the other scenario. In this paper, we propose UniHOPE, a unified approach for general 3D hand-object pose estimation, flexibly adapting both scenarios. Technically, we design a grasp-aware feature fusion module to integrate hand-object features with an object switcher to dynamically control the hand-object pose estimation according to grasping status. Further, to uplift the robustness of hand pose estimation regardless of object presence, we generate realistic de-occluded image pairs to train the model to learn object-induced hand occlusions, and formulate multi-level feature enhancement techniques for learning occlusion-invariant features. Extensive experiments on three commonly-used benchmarks demonstrate UniHOPE’s SOTA performance in addressing hand-only and hand-object scenarios. Code will be released on https://github.com/JoyboyWang/UniHOPE_Pytorch. Yinqiao Wang, Hao Xu 0018, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 3 |
| 2025 | EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual InsightsabstractTraffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints to anomaly scenarios such as crashes and honking. Our contributions are twofold. First, we compile AV-TAU, the first large-scale audio-visual dataset for TAU, providing 29,865 traffic anomaly videos and 149,325 Q&A pairs, while supporting five essential TAU tasks. Second, we develop EchoTraffic, a multimodal LLM that integrates audio and visual data for TAU, through our audio-insight frame selector and dynamic connector to effectively extract crucial audio cues for anomaly understanding with a two-phase training framework. Experimental results on AV-TAU manifest that EchoTraffic sets a new SOTA performance in TAU, outperforming the existing multimodal LLMs. Our contributions, including AV-TAU and EchoTraffic, pave a new direction for multimodal TAU. Zhenghao Xing, Hao Chen 0193, Binzhu Xie, Xuemiao Xu, Jianye Hao, Chi-Wing Fu, Xiaowei Hu 0001, Pheng-Ann Heng |
CVPR | 10 |
| 2025 | Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian SplattingabstractLifting multi-view 2D instance segmentation to a radiance field has proven effective to enhance 3D understanding. Existing works rely on direct matching for end-to-end lifting, yielding inferior results, or employ a two-stage solution constrained by complex preor post-processing. In this work, we design Unified-Lift, a new end-to-end object-aware lifting approach that aims for high-quality 3D segmentation based on our object-aware 3D Gaussian representation. To start, we augment each Gaussian point with a Gaussian-level feature learned using a contrastive loss to encode instance information. Importantly, we introduce a learnable object-level codebook to account for individual objects in the scene for an explicit object-level understanding and associate the encoded object-level features with the Gaussian-level point features for segmentation predictions. While promising, achieving effective codebook learning is nontrivial and a naive solution leads to degraded performance. Hence, we formulate the association learning module and the noisy label filtering module for effective and robust codebook learning. We conduct experiments on three benchmarks LERF-Masked, Replica, and Messy Rooms. Both qualitative and quantitative results manifest that our Unified-Lift clearly outperforms existing methods in terms of segmentation quality and time efficiency. Runsong Zhu, Shi Qiu 0001, Zhengzhe Liu, Ka-Hei Hui, Qianyi Wu, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 6 |
| 2025 | DiTAC: Discrete Teamwork Abstraction for Ad Hoc CollaborationabstractTraining autonomous agents to collaborate with unknown teammates in cooperative multi-agent environments remains a fundamental challenge in ad hoc teamwork research. Conventional approaches rely heavily on online interactions with arbitrary teammates under the assumption of full observability. However, in real-world scenarios, teammate policies are often inaccessible, making historical trajectory rollouts a more practical alternative. We propose DiTAC, a method that learns discrete teamwork abstractions for ad hoc collaboration by automatically extracting latent cooperation patterns from short trajectory segments and adapting effectively to diverse teammate behaviors. To mitigate the out-of-distribution challenge, we constrain learned representations within a discrete code-book. Furthermore, we employ a masked bidirectional transformer architecture to infer teammate behaviors from local observations, thereby relaxing the full observability assumption. Empirical results demonstrate that DiTAC significantly outperforms existing baselines and its variants across widely-used ad hoc teamwork tasks. Jing Wang 0055, Pengjie Gu, Mengchen Zhao, Guangyong Chen, Furui Liu, Pheng-Ann Heng |
ECAI | 6 |
| 2025 | Circuit Synthesis based on Hierarchical Conditional Diffusion
Xinyi Zhou 0010, Xing Li 0023, Yingzhao Lian, Lei Chen 0031, Mingxuan Yuan, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
ACM Great Lakes Symposium on VLSI | 9 |
| 2025 | Fast Image Super-Resolution via Consistency Rectified Flow
Wenbo Li 0002, Haoze Sun, Zhixin Wang, Long Peng 0003, Xiaowei Hu 0001, Renjing Pei, Pheng-Ann Heng |
ICCV | 11 |
| 2025 | SCJD: Sparse Correlation and Joint Distillation for Efficient 3D Human Pose EstimationabstractExisting 3D Human Pose Estimation (HPE) methods achieve high accuracy but suffer from computational overhead and slow inference, while knowledge distillation methods fail to address spatial relationships between joints and temporal correlations in multi-frame inputs. In this paper, we propose Sparse Correlation and Joint Distillation (SCJD), a novel framework that balances efficiency and accuracy for 3D HPE. SCJD introduces Sparse Correlation Input Sequence Downsampling to reduce redundancy in student network inputs while preserving inter-frame correlations. For effective knowledge transfer, we propose Dynamic Joint Spatial Attention Distillation, which includes Dynamic Joint Embedding Distillation to enhance the student’s feature representation using the teacher’s multi-frame context feature, and Adjacent Joint Attention Distillation to improve the student network’s focus on adjacent joint relationships for better spatial understanding. Additionally, Temporal Consistency Distillation aligns the temporal correlations between teacher and student networks through upsampling and global supervision. Extensive experiments demonstrate that SCJD achieves state-of-the-art performance. Code is available at https://github.com/wileychan/SCJD. Xuemiao Xu, Haoxin Yang, Huaidong Zhang, Pheng-Ann Heng |
ICME | 8 |
| 2025 | Dual Ensembled Multiagent Q-Learning with Hypernet Regularizer
Yaodong Yang 0002, Guangyong Chen, Hongyao Tang, Furui Liu, Danruo Deng, Pheng-Ann Heng |
AAMAS | 6 |
| 2025 | MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion ModelsabstractText-to-image diffusion models can generate high-quality images but lack fine-grained control of visual concepts, limiting their creativity. Thus, we introduce component-controllable personalization, a new task that enables users to customize and reconfigure individual components within concepts. This task faces two challenges: semantic pollution, where undesired elements disrupt the target concept, and semantic imbalance, which causes disproportionate learning of the target concept and component. To address these, we design MagicTailor, a framework that uses Dynamic Masked Degradation to adaptively perturb unwanted visual semantics and Dual-Stream Balancing for more balanced learning of desired visual semantics. The experimental results show that MagicTailor achieves superior performance in this task and enables more personalized and creative image generation. Jiancheng Huang, Jinbin Bai, Hao Chen 0193, Guangyong Chen, Xiaowei Hu 0001, Pheng-Ann Heng |
IJCAI | 8 |
| 2025 | Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
Ruize Cui, Jiaan Zhang, Jialun Pei, Kai Wang 0092, Pheng-Ann Heng, Harry Qin |
MICCAI (10) | 5 |
| 2025 | ClipGS: Clippable Gaussian Splatting for Interactive Cinematic Visualization of Volumetric Medical Data
Chengkun Li, Yuqi Tong, Kai Chen 0028, Zhenya Yang, Shi Qiu 0001, Jason Ying-Kuen Chan, Pheng-Ann Heng, Qi Dou 0001 |
MICCAI (10) | 8 |
| 2025 | Sequence-Independent Continual Test-Time Adaptation with Mixture of Incremental Experts for Cross-Domain Segmentation
Dunyuan Xu, Yuchen Yuan, Xikai Yang, Jingyang Zhang, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (16) | 7 |
| 2025 | Medical Large Vision Language Models with Multi-image Visual Ability
Xikai Yang, Juzheng Miao, Yuchen Yuan, Qi Dou 0001, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (5) | 7 |
| 2025 | GLID$^2$E: A Gradient-Free Lightweight Fine-tune Approach for Discrete Biological Sequence DesignabstractThe design of biological sequences is essential for engineering functional biomolecules that contribute to advancements in human health and biotechnology. Recent advances in diffusion models, with their generative power and efficient conditional sampling, have made them a promising approach for sequence generation. To enhance model performance on limited data and enable multi-objective design and optimization, reinforcement learning (RL)-based fine-tuning has shown great potential. However, existing post-sampling and fine-tuning methods either lack stability in discrete optimization when avoiding gradients or incur high computational costs when employing gradient-based approaches, creating significant challenges for achieving both control and stability in the tuning process.
To address these limitations, we propose GLID$^2$E, a gradient-free RL-based tuning approach for discrete diffusion models. Our method introduces a clipped likelihood constraint to regulate the exploration space and implements reward shaping to better align the generative process with design objectives, ensuring a more stable and efficient tuning process.
By integrating these techniques, GLID$^2$E mitigates training instabilities commonly encountered in RL and diffusion-based frameworks, enabling robust optimization even in challenging biological design tasks. In the DNA sequence and protein sequence design systems, GLID$^2$E achieves competitive performance in function-based design while maintaining computational efficiency and a flexible tuning mechanism. Hanqun Cao, Haosen Shi 0003, Sinno Jialin Pan, Pheng-Ann Heng |
NeurIPS | 5 |
| 2025 | Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningabstractGeneralized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency. To mitigate this dilemma, dual-system approaches have been proposed to leverage a VLM-based System 2 module for handling high-level decision-making, and a separate System 1 action module for ensuring real-time control. However, existing designs maintain both systems as separate models, limiting System 1 from fully leveraging the rich pretrained knowledge from the VLM-based System 2. In this work, we propose Fast-in-Slow (FiS), a unified dual-system vision-language-action (VLA) model that embeds the System 1 execution module within the VLM-based System 2 by partially sharing parameters. This innovative paradigm not only enables high-frequency execution in System 1, but also facilitates coordination between multimodal reasoning and execution components within a single foundation model of System 2. Given their fundamentally distinct roles within FiS-VLA, we design the two systems to incorporate heterogeneous modality inputs alongside asynchronous operating frequencies, enabling both fast and precise manipulation. To enable coordination between the two systems, a dual-aware co-training strategy is proposed that equips System 1 with action generation capabilities while preserving System 2’s contextual understanding to provide stable latent conditions for System 1. For evaluation, FiS-VLA outperforms previous state-of-the-art methods by 8% in simulation and 11% in real-world tasks in terms of average success rate, while achieving a 117.7 Hz control frequency with action chunk set to eight. Project web page: https://fast-in-slow.github.io. Hao Chen 0193, Jiaming Liu 0003, Chenyang Gu, Zhuoyang Liu, Renrui Zhang, Xiaoqi Li 0020, Yandong Guo, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng |
NeurIPS | 11 |
| 2025 | T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTabstractRecent advancements in large language models have demonstrated how chain-of-thought (CoT) and reinforcement learning (RL) can improve performance. However, applying such reasoning strategies to the visual generation domain remains largely unexplored. In this paper, we present **T2I-R1**, a novel reasoning-enhanced text-to-image generation model, powered by RL with a bi-level CoT reasoning process. Specifically, we identify two levels of CoT that can be utilized to enhance different stages of generation: (1) the semantic-level CoT for high-level planning of the prompt and (2) the token-level CoT for low-level pixel processing during patch-by-patch generation. To better coordinate these two levels of CoT, we introduce **BiCoT-GRPO** with an ensemble of generation rewards, which seamlessly optimizes both generated CoTs within the same training step. By applying our reasoning strategies to the baseline model, Janus-Pro, we achieve superior performance with 13% improvement on T2I-CompBench and 19% improvement on the WISE benchmark, even surpassing the state-of-the-art model FLUX.1. All the training code is in the supplementary material and will be made public. Dongzhi Jiang, Renrui Zhang, Zhuofan Zong, Hao Li 0069, Le Zhuo, Shilin Yan, Pheng-Ann Heng, Hongsheng Li 0001 |
NeurIPS | 8 |
| 2025 | SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyabstractRecent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial role of scenes in storytelling, which restricts their creativity in practice. This paper introduces scene-oriented story generation, addressing two key challenges: (i) scene planning, where current methods fail to ensure scene-level narrative coherence by relying solely on text descriptions, and (ii) scene consistency, which remains largely unexplored in terms of maintaining scene consistency across multiple stories. We propose SceneDecorator, a training-free framework that employs VLM-Guided Scene Planning to ensure narrative coherence across different scenes in a ``global-to-local'' manner, and Long-Term Scene-Sharing Attention to maintain long-term scene consistency and subject diversity across generated stories. Extensive experiments demonstrate the superior performance of SceneDecorator, highlighting its potential to unleash creativity in the fields of arts, films, and games. Quanjian Song, Fei Shen 0004, Xiaowei Hu 0001, Cunjian Chen, Pheng-Ann Heng |
NeurIPS | 8 |
| 2025 | Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPOabstractRecent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are central to these developments, showcasing different pros and cons. Autoregressive image generation, also interpretable as a sequential CoT reasoning process, presents unique challenges distinct from LLM-based CoT reasoning. These encompass ensuring text-image consistency, improving image aesthetic quality, and designing sophisticated reward models, rather than relying on simpler rule-based rewards. While recent efforts have extended RL to this domain, these explorations typically lack an in-depth analysis of the domain-specific challenges and the characteristics of different RL strategies. To bridge this gap, we provide the first comprehensive investigation of the GRPO and DPO algorithms in autoregressive image generation, evaluating their ***in-domain*** performance and ***out-of-domain*** generalization, while scrutinizing the impact of ***different reward models*** on their respective capabilities. Our findings reveal that GRPO and DPO exhibit distinct advantages, and crucially, that reward models possessing stronger intrinsic generalization capabilities potentially enhance the generalization potential of the applied RL algorithms. Furthermore, we systematically explore ***three prevalent scaling strategies*** to enhance both their in-domain and out-of-domain proficiency, deriving unique insights into efficiently scaling performance for each paradigm. We hope our study paves a new path for inspiring future work on developing more effective RL algorithms to achieve robust CoT reasoning in the realm of autoregressive image generation. Chengzhuo Tong, Renrui Zhang, Wenyu Shan, Zhenghao Xing, Hongsheng Li 0001, Pheng-Ann Heng |
NeurIPS | 8 |
| 2025 | What We Miss Matters: Learning from the Overlooked in Point Cloud TransformersabstractPoint Cloud Transformers have become a cornerstone in 3D representation for their ability to model long-range dependencies via self-attention. However, these models tend to overemphasize salient regions while neglecting other informative regions, which limits feature diversity and compromises robustness.
To address this challenge, we introduce BlindFormer, a novel contrastive attention learning framework that redefines saliency by explicitly incorporating features typically neglected by the model. The proposed Attentional Blindspot Mining (ABM) suppresses highly attended regions during training, thereby guiding the model to explore its own blind spots. This redirection of attention expands the model’s perceptual field and uncovers richer geometric cues.
To consolidate these overlooked features, BlindFormer employs Blindspot-Aware Joint Optimization (BJO), a joint learning objective that integrates blindspot feature alignment with the original pretext task. BJO enhances feature discrimination while preserving performance on the primary task, leading to more robust and generalizable representations.
We validate BlindFormer on several challenging benchmarks and demonstrate consistent performance gains across multiple Transformer backbones. Notably, it improves Point-MAE by +13.4\% and PointGPT-S by +6.3\% on OBJ-BG under Gaussian noise. These results highlight the importance of mitigating attentional biases in 3D representation learning, revealing BlindFormer’s superior ability to handle perturbations and improve feature discrimination. Renrui Zhang, Guangyong Chen, Anfeng Liu, Pheng-Ann Heng |
NeurIPS | 8 |
| 2025 | Protein Inverse Folding From Structure FeedbackabstractThe inverse folding problem, aiming to design amino acid sequences that fold into desired three-dimensional structures, is pivotal for various biotechnological applications.
Here, we introduce a novel approach leveraging Direct Preference Optimization (DPO) to fine-tune an inverse folding model using feedback from a protein folding model.
Given a target protein structure, we begin by sampling candidate sequences from the inverse‐folding model, then predict the three‐dimensional structure of each sequence with the folding model to generate pairwise structural‐preference labels.
These labels are used to fine‐tune the inverse‐folding model under the DPO objective.
Our results on the CATH 4.2 test set demonstrate that DPO fine-tuning not only improves sequence recovery of baseline models but also leads to a significant improvement in average TM-Score from 0.77 to 0.81, indicating enhanced structure similarity.
Furthermore, iterative application of our DPO-based method on challenging protein structures yields substantial gains, with an average TM-Score increase of 79.5\% with regard to the baseline model.
This work establishes a promising direction for enhancing protein sequence design ability from structure feedback by effectively utilizing preference optimization. Junde Xu, Zijun Gao, Xinyi Zhou 0010, Xingyi Cheng, Guangyong Chen, Pheng-Ann Heng, Jiezhong Qiu |
NeurIPS | 8 |
| 2025 | CellVerse: Do Large Language Models Really Understand Cell Biology?abstractRecent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging 160M $\rightarrow$ 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis. Project Page: https://cellverse-cuhk.github.io Fan Zhang 0111, Tianyu Liu 0001, Zhihong Zhu 0001, Hao Wu 0094, Haixin Wang 0003, Yefeng Zheng 0001, Kun Wang 0056, Xian Wu 0001, Pheng-Ann Heng |
NeurIPS | 10 |
| 2025 | COS3D: Collaborative Open-Vocabulary 3D SegmentationabstractOpen-vocabulary 3D segmentation is a fundamental yet challenging task, requiring a mutual understanding of both segmentation and language. However, existing Gaussian-splatting-based methods rely either on a single 3D language field, leading to inferior segmentation, or on pre-computed class-agnostic segmentations, suffering from error accumulation. To address these limitations, we present COS3D, a new collaborative prompt-segmentation framework that contributes to effectively integrating complementary language and segmentation cues throughout its entire pipeline. We first introduce the new concept of collaborative field, comprising an instance field and a language field, as the cornerstone for collaboration. During training, to effectively construct the collaborative field, our key idea is to capture the intrinsic relationship between the instance field and language field, through a novel instance-to-language feature mapping and designing an efficient two-stage training strategy. During inference, to bridge distinct characteristics of the two fields, we further design an adaptive language-to-instance prompt refinement, promoting high-quality prompt-segmentation inference. Extensive experiments not only demonstrate COS3D's leading performance over existing methods on two widely-used benchmarks but also show its high potential to various applications,~\ie, novel image-based 3D segmentation, hierarchical segmentation, and robotics. Runsong Zhu, Ka-Hei Hui, Zhengzhe Liu, Qianyi Wu, Weiliang Tang, Shi Qiu 0001, Pheng-Ann Heng, Chi-Wing Fu |
NeurIPS | 7 |
| 2025 | Gated-GPS: enhancing protein-protein interaction site prediction with scalable learning and imbalance-aware optimizationabstractIn protein-protein interaction site (PPIS) prediction, existing machine learning models struggle with small datasets, limiting their predictive accuracy for unseen proteins. Additionally, class imbalance in protein complexes, where binding residues constitute a small fraction of all residues, hinders model performance. To address these challenges, we constructed a training dataset 9$\times $ larger than previous benchmarks by filtering the latest protein-protein complex data, improving diversity and generalization. We propose Gated-GPS, a Graph Transformer model with a novel gating mechanism designed to effectively leverage this expanded dataset. Additionally, we integrate cross-entropy loss with Tversky Loss to adjust sensitivity to positive and negative samples, mitigating class imbalance by emphasizing underrepresented binding residues. Experimental results show that Gated-GPS outperforms state-of-the-art (SOTA) models across four test sets. Notably, on the UBTest dataset, designed to evaluate generalization on unbounded proteins, our method improves MCC and AUPRC by 18.5% and 21.4%, respectively, over the previous SOTA. In a case study of snake venom toxin-protein interactions, our model accurately identified interaction sites, demonstrating its potential for therapeutic design and advancing the understanding of complex protein interactions. Xin Gao 0026, Hanqun Cao, Jinpeng Li 0004, Jiezhong Qiu, Guangyong Chen, Pheng-Ann Heng |
Briefings Bioinform. | 6 |
| 2025 | DivPro: diverse protein sequence design with direct structure recovery guidanceabstractMOTIVATION: Structure-based protein design is crucial for designing proteins with novel structures and functions, which aims to generate sequences that fold into desired structures. Current deep learning-based methods primarily focus on training and evaluating models using sequence recovery-based metrics. However, this approach overlooks the inherent ambiguity in the relationship between protein sequences and structures. Relying solely on sequence recovery as a training objective limits the models' ability to produce diverse sequences that maintain similar structures. These limitations become more pronounced when dealing with remote homologous proteins, which share functional and structural similarities despite low-sequence identity. RESULTS: Here, we present DivPro, a model that learns to design diverse sequences that can fold into similar structures. To improve sequence diversity, instead of learning a single fixed sequence representation for an input structure as in existing methods, DivPro learns a probabilistic sequence space from which diverse sequences could be sampled. We leverage the recent advancements in in silico protein structure prediction. By incorporating structure prediction results as training guidance, DivPro ensures that sequences sampled from this learned space reliably fold into the target structure. We conducted extensive experiments on three sequence design benchmarks and evaluated the structures of designed sequences using structure prediction models including AlphaFold2. Results show that DivPro can maintain high structure recovery while significantly improving the sequence diversity. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/veghen/DivPro. Xinyi Zhou 0010, Guibao Shen, Ying-Cong Chen, Guangyong Chen, Pheng-Ann Heng |
Bioinform. | 5 |
| 2025 | Unsupervised multi-source domain adaptation via contrastive learning for EEG classification
Chengjian Xu, Yonghao Song, Qingqing Zheng, Qiong Wang 0001, Pheng-Ann Heng |
Expert Syst. Appl. | 5 |
| 2025 | Enhancing source-free domain adaptation in Medical Image Segmentation via regulated model self-training
Kang Li 0007, Shi Gu, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2025 | Unambiguous granularity distillation for asymmetric image retrieval
Haoquan Zhang, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng-Ann Heng, Shengfeng He |
Neural Networks | 9 |
| 2025 | Unifying Physically-Informed Weather Priors in a Single Model for Image Restoration Across Multiple Adverse Weather ConditionsabstractImage restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather-related artifacts vary with the particle’s distance to the camera according to the established scene visibility analysis, where close and faraway regions are more affected by falling drops and fog effects, respectively. In challenging weather conditions, existing image restoration methods fall short by not accounting for the varying impact of adverse weather on different scene regions. We develop a novel unified imaging model combined with a weather-prior-based network that directly incorporates weather-specific physical imaging processes into the restoration process. This approach not only enhances visibility in both near and distant regions affected by drops but also outperforms current state-of-the-art methods by effectively mitigating artifacts such as fog. Our contributions include a comprehensive analysis of weather-related visual factors and the development of an innovative network architecture that leverages estimated occlusion and transmission to restore scene details. Experimental results on three synthetic benchmarks, including our Weather30K dataset, along with two all-weather datasets, and a real-world benchmark with challenging mixed weather conditions, show the superiority of our method against state-of-the-art methods. Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Multi-Organ Segmentation From Partially Labeled and Unaligned Multi-Modal MRI in Thyroid-Associated OrbitopathyabstractThyroid-associated orbitopathy (TAO) is a prevalent inflammatory autoimmune disorder, leading to orbital disfigurement and visual disability. Automatic comprehensive segmentation tailored for quantitative multi-modal MRI assessment of TAO holds enormous promise but is still lacking. In this paper, we propose a novel method, named cross-modal attentive self-training (CMAST), for the multi-organ segmentation in TAO using partially labeled and unaligned multi-modal MRI data. Our method first introduces a dedicatedly designed cross-modal pseudo label self-training scheme, which leverages self-training to refine the initial pseudo labels generated by cross-modal registration, so as to complete the label sets for comprehensive segmentation. With the obtained pseudo labels, we further devise a learnable attentive fusion module to aggregate multi-modal knowledge based on learned cross-modal feature attention, which relaxes the requirement of pixel-wise alignment across modalities. A prototypical contrastive learning loss is further incorporated to facilitate cross-modal feature alignment. We evaluate our method on a large clinical TAO cohort with 100 cases of multi-modal orbital MRI. The experimental results demonstrate the promising performance of our method in achieving comprehensive segmentation of TAO-affected organs on both T1 and T1c modalities, outperforming previous methods by a large margin. Our code is available at: https://github.com/cchen-cc/CMAST. Cheng Chen 0013, Yuan Zhong 0003, Jinyue Cai, Karen Kar Wun Chan, Qi Dou 0001, Kelvin Kam Lung Chong, Pheng-Ann Heng, Winnie Chiu-Wing Chu |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Boosting Few-Shot Semantic Segmentation of 3D Medical Images via Collaborative Slice AlignmentabstractFew-shot semantic segmentation (FSS) of 3D medical images requires finding a 2D slice from the labeled volume as support to 'query' slices of the unlabeled one. Accurately determining support slices is crucial for learning representative prototypical features, thereby enhancing segmentation accuracy. The existing methods typically resort to the true position of the query target to align the query with support slices or simply exploit one key support slice to segment all query slices, which inevitably results in poor practicality and mis-segmentation. In this regard, we seek a practical and efficient solution by proposing a novel Collaborative Slice Alignment (CSA) module, which densely assigns each query slice its own fittest support without knowing the target prior. Concretely, our CSA first estimates the confidence scores of slices from the sorting task to implicitly reflect their physical location in the human body. The estimated scores are considered as spatial references for aligning support slices and query slices so that each matching pair shares the most similar image contents. Moreover, the self-learnable ranking objective allows CSA to transfer internal knowledge into both support and query features to further boost the FSS performance. Additionally, we introduce an Information Reconciliation (InRe) module to mitigate the inconsistent feature distribution caused by the individual differences between support and query images. Experimental results demonstrate that the combination of CSA and InRe achieves an average Dice score improvement of at least 8.61% across three datasets, consistently outperforming other state-of-the-art methods. Jialun Pei, Zhiwei Wang 0002, Qiang Li 0018, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Multi-Scale Spatio-Temporal Transformer-Based Imbalanced Longitudinal Learning for Glaucoma Forecasting From Irregular Time Series ImagesabstractGlaucoma is one of the major eye diseases that leads to progressive optic nerve fiber damage and irreversible blindness, afflicting millions of individuals. Glaucoma forecast is a good solution to early screening and intervention of potential patients, which is helpful to prevent further deterioration of the disease. It leverages a series of historical fundus images of an eye and forecasts the likelihood of glaucoma occurrence in the future. However, the irregular sampling nature and the imbalanced class distribution are two challenges in the development of disease forecasting approaches. To this end, we introduce the Multi-scale Spatio-temporal Transformer Network (MST-former) based on the transformer architecture tailored for sequential image inputs, which can effectively learn representative semantic information from sequential images on both temporal and spatial dimensions. Specifically, we employ a multi-scale structure to extract features at various resolutions, which can largely exploit rich spatial information encoded in each image. Besides, we design a time distance matrix to scale time attention in a non-linear manner, which could effectively deal with the irregularly sampled data. Furthermore, we introduce a temperature-controlled Balanced Softmax Cross-entropy loss to address the class imbalance issue. Extensive experiments on the Sequential fundus Images for Glaucoma Forecast (SIGF) dataset demonstrate the superiority of the proposed MST-former method, achieving an AUC of 96.6% for glaucoma forecasting. Besides, our method shows excellent generalization capability on the Alzheimer's Disease Neuroimaging Initiative (ADNI) MRI dataset, with an accuracy of 88.2% for mild cognitive impairment and Alzheimer's disease prediction, outperforming the compared method by a large margin. A series of ablation studies further verify the contribution of our proposed components in addressing the irregular sampled and class imbalanced problems. Xikai Yang, Xi Wang 0013, Yuchen Yuan, Jinpeng Li 0004, Guangyong Chen, Ning Li Wang, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | BAT: A Versatile Bipartite Attention-Based Approach for Comprehensive Truth Inference in Mobile CrowdsourcingabstractThe proliferation of smart mobile devices has catalyzed the growth of Mobile CrowdSourcing (MCS) as a distributed problem-solving paradigm. MCS platforms heavily rely on advanced truth inference techniques to extract reliable information from diverse and potentially noisy crowd-contributed data. Existing truth inference models often made simplified assumptions about workers or tasks, employing complex Bayesian models or stringent data aggregation methods. These approaches tend to be task-specific, primarily limited to categorical labeling, making adaptations to other mobile computing scenarios labor-intensive. To address these limitations, we introduce the Bipartite Attention-driven Truth (BAT), a versatile approach tailored for mobile computing environments. BAT utilizes an Attributed Bipartite Graph (ABG) to holistically model the MCS process, with workers and tasks as nodes connected by edges representing answer-specific attributes. The approach employs a bipartite graph neural network with an innovative attention mechanism to assess the importance of different answers. BAT extends beyond categorical tasks to support numerical ones by incorporating novel feature representations and model extensions. Theoretical analyses clarify the link between answer similarity and worker expertise. Extensive experiments using diverse real-world datasets demonstrate BAT's superior performance compared to state-of-the-art categorical and numerical truth inference models, highlighting its effectiveness in mobile computing scenarios. Jiacheng Liu 0001, Feilong Tang 0001, Hao Liu 0085, Long Chen 0025, Yichuan Yu, Yanmin Zhu 0006, Jiadi Yu, Xiaofeng Hou, Pheng-Ann Heng |
IEEE Trans. Mob. Comput. | 9 |
| 2025 | LMT++: Adaptively Collaborating LLMs With Multi-Specialized Teachers for Continual VQA in Robotic Surgical VideosabstractVisual question answering (VQA) plays a vital role in advancing surgical education. However, due to the privacy concern of patient data, training VQA model with previously used data becomes restricted, making it necessary to use the exemplar-free continual learning (CL) approach. Previous CL studies in the surgical field neglected two critical issues: i) significant domain shifts caused by the wide range of surgical procedures collected from various sources, and ii) the data imbalance problem caused by the unequal occurrence of medical instruments or surgical procedures. This paper addresses these challenges with a multimodal large language model (LLM) and an adaptive weight assignment strategy. First, we developed a novel LLM-assisted multi-teacher CL framework (named LMT++), which could harness the strength of a multimodal LLM as a supplementary teacher. The LLM's strong generalization ability, as well as its good understanding of the surgical domain, help to address the knowledge gap arising from domain shifts and data imbalances. To incorporate the LLM in our CL framework, we further proposed an innovative approach to process the training data, which involves the conversion of complex LLM embeddings into logits value used within our CL training framework. Moreover, we design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of conventional VQA models obtained in previous model training processes within the CL framework. Finally, we created a new surgical VQA dataset for model evaluation. Comprehensive experimental findings on these datasets show that our approach surpasses state-of-the-art CL methods. Yuyang Du 0001, Kexin Chen 0003, Yue Zhan, Chang Han Low, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 9 |
| 2025 | S²Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in ORabstractScene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic scene graphs depend on intermediate processes with pose estimation and object detection. This pipeline may potentially compromise the flexibility of learning multimodal representations, consequently constraining the overall effectiveness. In this study, we introduce a novel single-stage bi-modal transformer framework for SGG in the OR, termed S2Former-OR, aimed to complementally leverage multi-view 2D scenes and 3D point clouds for SGG in an end-to-end manner. Concretely, our model embraces a View-Sync Transfusion scheme to encourage multi-view visual information interaction. Concurrently, a Geometry-Visual Cohesion operation is designed to integrate the synergic 2D semantic features into 3D point cloud features. Moreover, based on the augmented feature, we propose a novel relation-sensitive transformer decoder that embeds dynamic entity-pair queries and relational trait priors, which enables the direct prediction of entity-pair relations for graph generation without intermediate steps. Extensive experiments have validated the superior SGG performance and lower computational cost of S2Former-OR on 4D-OR benchmark, compared with current OR-SGG methods, e.g., 3 percentage points Precision increase and 24.2M reduction in model parameters. We further compared our method with generic single-stage SGG methods with broader metrics for a comprehensive evaluation, with consistently better performance achieved. Our source code can be made available at: https://github.com/PJLallen/S2Former-OR. Jialun Pei, Diandian Guo, Jingyang Zhang, Manxi Lin, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Instrument-Tissue-Guided Surgical Action Triplet Detection via Textual-Temporal Trail ExplorationabstractSurgical action triplet detection offers intuitive intraoperative scene analysis for dynamically perceiving laparoscopic surgical workflows and analyzing the interaction between instruments and tissues. The current challenge of this task lies in simultaneously localizing surgical instruments while performing more accurate surgical triplet recognition to enhance a comprehensive understanding of intraoperative surgical scenes. To fully leverage the spatial localization of surgical instruments for associating with triplet detection, we propose an Instrument-Tissue-Guided Triplet detector, termed ITG-Trip, which navigates the confluence of surgical action cues through instrument and tissue pseudo-localization labeling to optimize action triplet detection. For exploiting textual and temporal trails, our framework embraces a Visual-Linguistic Association (VLA) module that exploits a pre-trained text encoder to distill textual prior knowledge, enhancing semantic information in global visual features and compensating rare interaction class perception. Besides, we introduce a Mamba-enhanced Spatial-temporal Perception (MSP) decoder, which weaves Mamba and Transformer blocks to explore subject- and object-aware spatial and temporal information to improve the accuracy of action triplet detection in long-time sequence surgical videos. Experimental results on the CholecT50 benchmark indicate that our method significantly outperforms existing state-of-the-art methods in both instrument localization and action triplet detection. The code is available at: github.com/PJLallen/ITG-Trip. Jialun Pei, Jiaan Zhang, Guanyi Qin, Kai Wang 0092, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2025 | PAL: Boosting Skin Lesion Segmentation via Probabilistic Attribute LearningabstractSkin lesion segmentation is vital for the early detection, diagnosis, and treatment of melanoma, yet it remains challenging due to significant variations in lesion attributes (e.g., color, size, shape), ambiguous boundaries, and noise interference. Recent advancements have focused on capturing contextual information and incorporating boundary priors to handle challenging lesions. However, there has been limited exploration on the explicit analysis of the inherent patterns of skin lesions, a crucial aspect of the knowledge-driven decision-making process used by clinical experts. In this work, we introduce a novel approach called Probabilistic Attribute Learning (PAL), which leverages knowledge of lesion patterns to achieve enhanced performance on challenging lesions. Recognizing that the lesion patterns exhibited in each image can be properly depicted by disentangled attributes, we begin by explicitly estimating the distributions of these attributes as distinct Gaussian distributions, with mean and variance indicating the most likely pattern of that attribute and its variation. Using Monte Carlo Sampling, we iteratively draw multiple samples from these distributions to capture various potential patterns for each attribute. These samples are then merged through an effective attribute fusion technique, resulting in diverse representations that comprehensively depict the lesion class. By performing pixel-class proximity matching between each pixel-wise representation and the diverse class-wise representations, we significantly enhance the model's robustness. Extensive experiments on two public skin lesion datasets and one unified polyp lesion dataset demonstrate the effectiveness and strong generalization ability of our method. Codes are available at https://github.com/IsYuchenYuan/PAL. Yuchen Yuan, Xi Wang 0013, Jinpeng Li 0004, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Effective Semi-Supervised Medical Image Segmentation With Probabilistic Representations and Prototype LearningabstractLabel scarcity, class imbalance and data uncertainty are three primary challenges that are commonly encountered in the semi-supervised medical image segmentation. In this work, we focus on the data uncertainty issue that is overlooked by previous literature. To address this issue, we propose a probabilistic prototype-based classifier that introduces uncertainty estimation into the entire pixel classification process, including probabilistic representation formulation, probabilistic pixel-prototype proximity matching, and distribution prototype update, leveraging principles from probability theory. By explicitly modeling data uncertainty at the pixel level, model robustness of our proposed framework to tricky pixels, such as ambiguous boundaries and noises, is greatly enhanced when compared to its deterministic counterpart and other uncertainty-aware strategy. Empirical evaluations on three publicly available datasets that exhibit severe boundary ambiguity show the superiority of our method over several competitors. Moreover, our method also demonstrates a stronger model robustness to simulated noisy data. Code is available at https://github.com/IsYuchenYuan/PPC. Yuchen Yuan, Xi Wang 0013, Xikai Yang, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2025 | DC²T: Disentanglement-Guided Consolidation and Consistency Training for Semi-Supervised Cross-Site Continual SegmentationabstractContinual Learning (CL) is recognized to be a storage-efficient and privacy-protecting approach for learning from sequentially-arriving medical sites. However, most existing CL methods assume that each site is fully labeled, which is impractical due to budget and expertise constraint. This paper studies the Semi-Supervised Continual Learning (SSCL) that adopts partially-labeled sites arriving over time, with each site delivering only limited labeled data while the majority remains unlabeled. In this regard, it is challenging to effectively utilize unlabeled data under dynamic cross-site domain gaps, leading to intractable model forgetting on such unlabeled data. To address this problem, we introduce a novel Disentanglement-guided Consolidation and Consistency Training (DC2T) framework, which roots in an Online Semi-Supervised representation Disentanglement (OSSD) perspective to excavate content representations of partially labeled data from sites arriving over time. Moreover, these content representations are required to be consolidated for site-invariance and calibrated for style-robustness, in order to alleviate forgetting even in the absence of ground truth. Specifically, for the invariance on previous sites, we retain historical content representations when learning on a new site, via a Content-inspired Parameter Consolidation (CPC) method that prevents altering the model parameters crucial for content preservation. For the robustness against style variation, we develop a Style-induced Consistency Training (SCT) scheme that enforces segmentation consistency over style-related perturbations to recalibrate content encoding. We extensively evaluate our method on fundus and cardiac image segmentation, indicating the advantage over existing SSCL methods for alleviating forgetting on unlabeled data. Jingyang Zhang, Jialun Pei, Dunyuan Xu, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Landmark-Free Preoperative-to-Intraoperative Registration in Laparoscopic Liver ResectionabstractLiver registration by overlaying preoperative 3D models onto intraoperative 2D frames can assist surgeons in perceiving the spatial anatomy of the liver clearly for a higher surgical success rate. Existing registration methods rely heavily on anatomical landmark-based workflows, which encounter two major limitations: 1) ambiguous landmark definitions fail to provide efficient markers for registration; 2) insufficient integration of intraoperative liver visual information in shape deformation modeling. To address these challenges, in this paper, we propose a landmark-free preoperative-to-intraoperative registration framework utilizing effective self-supervised learning, termed Self-P2IR. This framework transforms the conventional 3D-2D workflow into a 3D-3D registration pipeline, which is then decoupled into rigid and non-rigid registration subtasks. Self-P2IR first introduces a feature-disentangled transformer to learn robust correspondences for recovering rigid transformations. Further, a structure-regularized deformation network is designed to adjust the preoperative model to align with the intraoperative liver surface. This network captures structural correlations through geometry similarity modeling in a low-rank transformer network. To facilitate the validation of the registration performance, we also construct an in-vivo registration dataset containing liver resection videos of 21 patients, called P2I-LReg, which contains 346 keyframes that provide a global view of the liver together with liver mask annotations and calibrated camera intrinsic parameters. Extensive experiments and user studies on both synthetic and in-vivo datasets demonstrate the superiority and potential clinical applicability of our method. The code and dataset are available at https://github.com/junzastar/Self-P2IR. Jun Zhou 0029, Bingchen Gao, Kai Wang 0092, Jialun Pei, Pheng-Ann Heng, Harry Qin |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Delving Into Quaternion Wavelet Transformer for Facial Expression Recognition in the WildabstractThe Facial Expression Recognition (FER) technique has increasingly matured over time. However, recognizing facial expressions in wild environments poses great challenges in achieving promising performance. The main obstacles arise from various factors, such as illumination changes, head pose variations, and occlusions. To overcome interferences from external environments and improve recognition accuracy, we propose a novel Quaternion Wavelet TRansformer (QWTR) model for FER in the wild. Specifically, we present a Quaternion Value Transformer (QVT) network that combines quaternion multi-head attention with quaternion CNN to capture emotional cues from global and local perception. To preserve the color structure while enhancing image contrast and brightness, we introduce a Quaternion Histogram Equalization (QHE) representation to transform color images into quaternion matrices representation. After that, to alleviate the impact of head pose and occlusion together with feature redundancy, a Quaternion Wavelet Feature Selection (QWFS) scheme is designed to decompose quaternion features and select the most correlated signals. Extensive experiments have been conducted on four in-the-wild FER datasets and several specific FER benchmarks under various conditions. The qualitative and quantitative results demonstrate thatQWTRoutperforms other state-of-the-art methods in FER benchmarks, e.g., 68.37% vs. 66.31% accuracy on the AffectNet dataset. Yu Zhou 0049, Jialun Pei, Weixin Si, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Multim. | 5 |
| 2025 | Scale-Aware Super-Resolution Network With Dual Affinity Learning for Lesion Segmentation From Medical ImagesabstractConvolutional neural networks (CNNs) have shown remarkable progress in medical image segmentation. However, the lesion segmentation remains a challenge to state-of-the-art CNN-based algorithms due to the variance in scales and shapes. On the one hand, tiny lesions are hard to delineate precisely from the medical images which are often of low resolutions. On the other hand, segmenting large-size lesions requires large receptive fields, which exacerbates the first challenge. In this article, we present a scale-aware super-resolution (SR) network to adaptively segment lesions of various sizes from low-resolution (LR) medical images. Our proposed network contains dual branches to simultaneously conduct lesion mask SR (LMSR) and lesion image SR (LISR). Meanwhile, we introduce scale-aware dilated convolution (SDC) blocks into the multitask decoders to adaptively adjust the receptive fields of the convolutional kernels according to the lesion sizes. To guide the segmentation branch to learn from richer high-resolution (HR) features, we propose a feature affinity (FA) module and a scale affinity (SA) module to enhance the multitask learning of the dual branches. On multiple challenging lesion segmentation datasets, our proposed network achieved consistent improvements compared with other state-of-the-art methods. Code will be available at: https://github.com/poiuohke/SASR_Net. Luyang Luo, Yanwen Li, Zhizhong Chai, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Hand-Shadow PoserabstractHand shadow art is a captivating art form, creatively using hand shadows to reproduce expressive shapes on the wall. In this work, we study an inverse problem: given a target shape, find the poses of left and right hands that together best produce a shadow resembling the input. This problem is nontrivial, since the design space of 3D hand poses is huge while being restrictive due to anatomical constraints. Also, we need to attend to the input's shape and crucial features, though the input is colorless and textureless. To meet these challenges, we design Hand-Shadow Poser, a three-stage pipeline, to decouple the anatomical constraints (by hand) and semantic constraints (by shadow shape): (i) a generative hand assignment module to explore diverse but reasonable left/right-hand shape hypotheses; (ii) a generalized hand-shadow alignment module to infer coarse hand poses with a similarity-driven strategy for selecting hypotheses; and (iii) a shadow-feature-aware refinement module to optimize the hand poses for physical plausibility and shadow feature preservation. Further, we design our pipeline to be trainable on generic public hand data, thus avoiding the need for any specialized training dataset. For method validation, we build a benchmark of 210 diverse shadow shapes of varying complexity and a comprehensive set of metrics, including a novel DINOv2-based evaluation metric. Through extensive comparisons with multiple baselines and user studies, our approach is demonstrated to effectively generate bimanual hand poses for a large variety of hand shapes for over 85% of the benchmark cases. Hao Xu 0018, Yinqiao Wang, Niloy J. Mitra, Shuaicheng Liu, Pheng-Ann Heng, Chi-Wing Fu |
ACM Trans. Graph. | 5 |
| 2025 | An Adaptive and Interpretable Congestion Control Service Based on Multi-Objective Reinforcement LearningabstractThe need for an adaptive congestion control (CC) service is crucial due to the heterogeneity of systems and the diversity of applications. Traditional CC methods often fail to adaptively balance throughput and delay, struggling to meet the varied demands of different network applications. In this work, we introduceAuto, a novel CC service that employs Multi-Objective Reinforcement Learning (MORL) to transcend these limitations. Unlike conventional approaches,Autooptimizes policies within a single model to cater to all potential preferences for balancing throughput and delay, making it ideal for diverse and heterogeneous network environments. To enhance operational transparency, we developed an interpretation algorithm that translates MORL into a human- readable decision tree, essential for service computing where clarity and interpretability are crucial. Furthermore,Autoallows users to explicitly set flow priorities and target sending rates, meeting varied application demands. Our extensive evaluations show thatAutonot only consistently outperforms existing CC methods in diverse network conditions but also exhibits robustness to stochastic packet loss and rapid network changes. These capabilities establishAutoas a pioneering solution for next-generation congestion control in networking services. Jiacheng Liu 0001, Xu Li 0012, Feilong Tang 0001, Peng Li 0017, Long Chen 0025, Jiadi Yu, Yanmin Zhu 0006, Pheng-Ann Heng, Laurence T. Yang |
IEEE Trans. Serv. Comput. | 8 |
| 2025 | HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001 |
Vis. Comput. | 24 |
| 2024 | DR-Label: Label Deconstruction and Reconstruction of GNN Models for Catalysis SystemsabstractAttaining the equilibrium geometry of a catalyst-adsorbate system is key to fundamentally assessing its effective properties, such as adsorption energy. While machine learning methods with advanced representation or supervision strategies have been applied to boost and guide the relaxation processes of catalysis systems, existing methods that produce linearly aggregated geometry predictions are susceptible to edge representations ambiguity, and are therefore vulnerable to graph variations. In this paper, we present a novel graph neural network (GNN) supervision and prediction strategy DR-Label. Our approach mitigates the multiplicity of solutions in edge representation and encourages model predictions that are independent of graph structural variations. DR-Label first Deconstructs finer-grained equilibrium state information to the model by projecting the node-level supervision signal to each edge. Reversely, the model Reconstructs a more robust equilibrium state prediction by converting edge-level predictions to node-level via a sphere-fitting algorithm. When applied to three fundamentally different models, DR-Label consistently enhanced performance. Leveraging the graph structure invariance of the DR-Label strategy, we further propose DRFormer, which applied explicit intermediate positional update and achieves a new state-of-the-art performance on the Open Catalyst 2020 (OC20) dataset and the Cu-based single-atom alloys CO adsorption (SAA) dataset. We expect our work to highlight vital principles for advancing geometric GNN models for catalysis systems and beyond. Our code is available at https://github.com/bowenwang77/DR-Label Bowen Wang 0017, Jiezhong Qiu, Furui Liu, Shaogang Hao, Dong Li 0016, Guangyong Chen, Xiaolong Zou, Pheng-Ann Heng |
AAAI | 10 |
| 2024 | PointPatchMix: Point Cloud Mixing with Patch ScoringabstractData augmentation is an effective regularization strategy for mitigating overfitting in deep neural networks, and it plays a crucial role in 3D vision tasks, where the point cloud data is relatively limited. While mixing-based augmentation has shown promise for point clouds, previous methods mix point clouds either on block level or point level, which has constrained their ability to strike a balance between generating diverse training samples and preserving the local characteristics of point clouds. The significance of each part component of the point clouds has not been fully considered, as not all parts contribute equally to the classification task, and some parts may contain unimportant or redundant information. To overcome these challenges, we propose PointPatchMix, a novel approach that mixes point clouds at the patch level and integrates a patch scoring module to generate content-based targets for mixed point clouds. Our approach preserves local features at the patch level, while the patch scoring module assigns targets based on the content-based significance score from a pre-trained teacher model. We evaluate PointPatchMix on two benchmark datasets including ModelNet40 and ScanObjectNN, and demonstrate significant improvements over various baselines in both synthetic and real-world datasets, as well as few-shot settings. With Point-MAE as our baseline, our model surpasses previous methods by a significant margin. Furthermore, our approach shows strong generalization across various point cloud methods and enhances the robustness of the baseline model. Code is available at https://jiazewang.com/projects/pointpatchmix.html. Jinpeng Li 0004, Guangyong Chen, Anfeng Liu, Pheng-Ann Heng |
AAAI | 7 |
| 2024 | SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationabstractEstimating 3D hand mesh from RGB images is a longstanding track, in which occlusion is one of the most challenging problems. Existing attempts towards this task often fail when the occlusion dominates the image space. In this paper, we propose SiMA-Hand, aiming to boost the mesh reconstruction performance by Single-to-Multi-view Adaptation. First, we design a multi-view hand reconstructor to fuse information across multiple views by holistically adopting feature fusion at image, joint, and vertex levels. Then, we introduce a single-view hand reconstructor equipped with SiMA. Though taking only one view as input at inference, the shape and orientation features in the single-view reconstructor can be enriched by learning non-occluded knowledge from the extra views at training, enhancing the reconstruction precision on the occluded regions. We conduct experiments on the Dex-YCB and HanCo benchmarks with challenging object- and self-caused occlusion cases, manifesting that SiMA-Hand consistently achieves superior performance over the state of the arts. Code will be released on https://github.com/JoyboyWang/SiMA-Hand Pytorch. Yinqiao Wang, Hao Xu 0018, Pheng-Ann Heng, Chi-Wing Fu |
AAAI | 3 |
| 2024 | ANEDL: Adaptive Negative Evidential Deep Learning for Open-Set Semi-supervised LearningabstractSemi-supervised learning (SSL) methods assume that labeled data, unlabeled data and test data are from the same distribution. Open-set semi-supervised learning (Open-set SSL) con- siders a more practical scenario, where unlabeled data and test data contain new categories (outliers) not observed in labeled data (inliers). Most previous works focused on out- lier detection via binary classifiers, which suffer from insufficient scalability and inability to distinguish different types of uncertainty. In this paper, we propose a novel framework, Adaptive Negative Evidential Deep Learning (ANEDL) to tackle these limitations. Concretely, we first introduce evidential deep learning (EDL) as an outlier detector to quantify different types of uncertainty, and design different uncertainty metrics for self-training and inference. Furthermore, we propose a novel adaptive negative optimization strategy, making EDL more tailored to the unlabeled dataset containing both inliers and outliers. As demonstrated empirically, our proposed method outperforms existing state-of-the-art methods across four datasets. Yang Yu 0070, Danruo Deng, Furui Liu, Qi Dou 0001, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
AAAI | 7 |
| 2024 | Memory-Efficient Prompt Tuning for Incremental Histopathology ClassificationabstractRecent studies have made remarkable progress in histopathology classification. Based on current successes, contemporary works proposed to further upgrade the model towards a more generalizable and robust direction through incrementally learning from the sequentially delivered domains. Unlike previous parameter isolation based approaches that usually demand massive computation resources during model updating, we present a memory-efficient prompt tuning framework to cultivate model generalization potential in economical memory cost. For each incoming domain, we reuse the existing parameters of the initial classification model and attach lightweight trainable prompts into it for customized tuning. Considering the domain heterogeneity, we perform decoupled prompt tuning, where we adopt a domain-specific prompt for each domain to independently investigate its distinctive characteristics, and one domain-invariant prompt shared across all domains to continually explore the common content embedding throughout time. All domain-specific prompts will be appended to the prompt bank and isolated from further changes to prevent forgetting the distinctive features of early-seen domains. While the domain-invariant prompt will be passed on and iteratively evolve by style-augmented prompt refining to improve model generalization capability over time. In specific, we construct a graph with existing prompts and build a style-augmented graph attention network to guide the domain-invariant prompt exploring the overlapped latent embedding among all delivered domains for more domain-generic representations. We have extensively evaluated our framework with two histopathology tasks, i.e., breast cancer metastasis classification and epithelium-stroma tissue classification, where our approach yielded superior performance and memory efficiency over the competing methods. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
AAAI | 4 |
| 2024 | Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor SegmentationabstractMulti-modal brain tumor segmentation typically involves four magnetic resonance imaging (MRI) modalities, while incomplete modalities significantly degrade performance. Existing solutions employ explicit or implicit modality adaptation, aligning features across modalities or learning a fused feature robust to modality incompleteness. They share a common goal of encouraging each modality to express both itself and the others. However, the two expression abilities are entangled as a whole in a seamless feature space, resulting in prohibitive learning burdens. In this paper, we propose DeMoSeg to enhance the modality adaptation by Decoupling the task of representing the ego and other Modalities for robust incomplete multi-modal Segmentation. The decoupling is super lightweight by simply using two convolutions to map each modality onto four feature sub-spaces. The first sub-space expresses itself (Self-feature), while the remaining sub-spaces substitute for other modalities (Mutual-features). The Self- and Mutual-features interactively guide each other through a carefully-designed Channel-wised Sparse Self-Attention (CSSA). After that, a Radiologist-mimic Cross-modality expression Relationships (RCR) is introduced to have available modalities provide Self-feature and also ‘lend’ their Mutual-features to compensate for the absent ones by exploiting the clinical prior knowledge. The benchmark results on BraTS2020, BraTS2018 and BraTS2015 verify the DeMoSeg’s superiority thanks to the alleviated modality adaptation difficulty. Concretely, for BraTS2020, DeMoSeg increases Dice by at least 0.92%, 2.95% and 4.95% on whole tumor, tumor core and enhanced tumor regions, respectively, compared to other state-of-the-arts. Codes are at https://github.com/kk42yy/DeMoSeg. Kaixiang Yang 0004, Wenqi Shan, Xikai Yang, Xi Wang 0013, Pheng-Ann Heng, Qiang Li 0018, Zhiwei Wang 0002 |
BIBM | 7 |
| 2024 | SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
Hao Chen 0193, Jinpeng Li 0004, Chenyong Guan, Guangyong Chen, Pheng-Ann Heng |
BMVC | 9 |
| 2024 | Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
Jinpeng Li 0004, Jiancheng Huang, Qiang Nie, Yong Liu 0032, Bin-Bin Gao, Qiong Wang 0001, Pheng-Ann Heng, Guangyong Chen |
BMVC | 9 |
| 2024 | Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
Mengyang Wu, Xiaohu You 0001, Chi-Wing Fu, Qi Dou 0001, Pheng-Ann Heng |
ECCV (18) | 6 |
| 2024 | PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion
Runsong Zhu, Shi Qiu 0001, Qianyi Wu, Ka-Hei Hui, Pheng-Ann Heng, Chi-Wing Fu |
ECCV (2) | 5 |
| 2024 | Sample-Efficient Multiagent Reinforcement Learning with Reset ReplayabstractThe popularity of multiagent reinforcement learning (MARL) is growing rapidly with the demand for real-world tasks that require swarm intelligence. However, a noticeable drawback of MARL is its low sample efficiency, which leads to a huge amount of interactions with the environment. Surprisingly, few MARL works focus on this practical problem especially in the parallel environment setting, which greatly hampers the application of MARL into the real world. In response to this gap, in this paper, we propose Multiagent Reinforcement Learning with Reset Replay (MARR) to greatly improve the sample efficiency of MARL by enabling MARL training at a high replay ratio in the parallel environment setting for the first time. To achieve this, first, a reset strategy is introduced for maintaining the network plasticity to ensure that MARL continually learns with a high replay ratio. Second, MARR incorporates a data augmentation technique to boost the sample efficiency further. Extensive experiments in SMAC and MPE show that MARR significantly improves the performance of various MARL approaches with much fewer environment interactions. Yaodong Yang 0002, Guangyong Chen, Jianye Hao, Pheng-Ann Heng |
ICML | 4 |
| 2024 | LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic SurgeryabstractVisual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQA system by a sequential data stream from multiple resources is demanded in robotic surgery to address new tasks. In surgical scenarios, the privacy issue of patient data often restricts the availability of old data when updating the model, necessitating an exemplar-free continual learning (CL) setup. However, prior studies overlooked two vital problems of the surgical domain: i) large domain shifts from diverse surgical operations collected from multiple departments or clinical centers, and ii) severe data imbalance arising from the uneven presence of surgical instruments or activities during surgical procedures. This paper proposes to address these two problems with a multimodal large language model (LLM) and an adaptive weight assignment methodology. We first develop a new multi-teacher CL framework that leverages a multimodal LLM as the additional teacher. The strong generalization ability of the LLM can bridge the knowledge gap when domain shifts and data imbalances occur. We then put forth a novel data processing method that transforms complex LLM embeddings into logits compatible with our CL framework. We also design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of the old CL model. Finally, we construct a new dataset for surgical VQA tasks. Extensive experimental results demonstrate the superiority of our method to other advanced CL models. Kexin Chen 0003, Yuyang Du 0001, Tao You, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
ICRA | 8 |
| 2024 | Adaptive Federated Learning for EEG Emotion RecognitionabstractEmotion classification based on electroencephalogram (EEG) signals has drawn huge attention in affective brain computer interface (BCI). Recently, plenty of deep learning approaches have been proposed to improve the performance of EEG emotion recognition, especially the application of domain adaptation methods to tackle the challenge of large individual differences of EEG signals from subject to subject. However, these conventional transfer learning methods would result in information leakage during the sharing of domain data to enhance the accuracy of the target tasks. Therefore, in this paper, we proposed a distributed deep learning method, named adaptive federated learning (AdaFL) for EEG emotion recognition. In AdaFL, a server collaboratively learns a global model by adaptively aggregating the local models according to their importance in several communication rounds. In particular, an importance function is developed to evaluate each client, which would determine to select a subset of optimal clients for subsequent global model aggregation. The function is a simple transformation of the training loss and the sample size of the local models. Then, the resulting importance scores of selected clients are further converted into aggregation coefficients to measure the weights of the local models for global model aggregation. The distinct advantage of AdaFL is that the cross-subject information could be well utilized and the information leakage risk could be significantly reduced. To validate the efficacy of the proposed AdaFL, we conduct extensive experiments on two real EEG emotion datasets, i.e., SEED and DEAP. The experimental results show that the proposed AdaFL has achieved 94.95 ± 0.40% and 91.06 ± 0.29% average classification accuracy on the SEED and DEAP datasets, respectively, which reflects the superiority of our method over the state-of-the-art approaches. Calvin Chan, Qingqing Zheng, Chengjian Xu, Qiong Wang 0001, Pheng-Ann Heng |
IJCNN | 5 |
| 2024 | Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
Diandian Guo, Manxi Lin, Jialun Pei, He Tang 0002, Yueming Jin, Pheng-Ann Heng |
MICCAI (6) | 6 |
| 2024 | Noise Level Adaptive Diffusion Model for Robust Reconstruction of Accelerated MRI
Shoujin Huang, Guanxiong Luo, Xi Wang 0013, Ziran Chen, Yuwan Wang, Huaishui Yang, Pheng-Ann Heng, Mengye Lyu |
MICCAI (7) | 7 |
| 2024 | Epicardium Prompt-Guided Real-Time Cardiac Ultrasound Frame-to-Volume Registration
Long Lei, Jun Zhou 0007, Jialun Pei, Baoliang Zhao, Yueming Jin, Jeremy Yuen-Chun Teoh, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 8 |
| 2024 | Comprehensive Generative Replay for Task-Incremental Segmentation with Concurrent Appearance and Semantic Forgetting
Jingyang Zhang, Pheng-Ann Heng, Lixu Gu |
MICCAI (8) | 3 |
| 2024 | Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation
Juzheng Miao, Cheng Chen 0013, Keli Zhang, Jie Chuai, Quanzheng Li, Pheng-Ann Heng |
MICCAI (11) | 6 |
| 2024 | FM-OSD: Foundation Model-Enabled One-Shot Detection of Anatomical Landmarks
Juzheng Miao, Cheng Chen 0013, Keli Zhang, Jie Chuai, Quanzheng Li, Pheng-Ann Heng |
MICCAI (11) | 6 |
| 2024 | Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
Jialun Pei, Ruize Cui, Yaoqian Li, Weixin Si, Harry Qin, Pheng-Ann Heng |
MICCAI (6) | 6 |
| 2024 | Coarse-to-Fine Latent Diffusion Model for Glaucoma Forecast on Sequential Fundus Images
Yuhan Zhang 0001, Xikai Yang, Xiao Ma 0011, Ningli Wang, Xi Wang 0013, Pheng-Ann Heng |
MICCAI (5) | 8 |
| 2024 | Weakly-Supervised Medical Image Segmentation with Gaze Annotations
Yuan Zhong 0003, Chenhui Tang, Ruoxi Qi, Yuqi Gong, Pheng-Ann Heng, Janet Hui-wen Hsiao, Qi Dou 0001 |
MICCAI (3) | 7 |
| 2024 | Unveiling the Generalization Power of Fine-Tuned Large Language ModelsabstractHaoran Yang, Yumeng Zhang, Jiaqi Xu, Hongyuan Lu, Pheng-Ann Heng, Wai Lam. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hongyuan Lu, Pheng-Ann Heng, Wai Lam |
NAACL-HLT | 5 |
| 2024 | Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningabstractAs a marriage between offline RL and meta-RL, the advent of offline meta-reinforcement learning (OMRL) has shown great promise in enabling RL agents to multi-task and quickly adapt while acquiring knowledge safely. Among which, context-based OMRL (COMRL) as a popular paradigm, aims to learn a universal policy conditioned on effective task representations. In this work, by examining several key milestones in the field of COMRL, we propose to integrate these seemingly independent methodologies into a unified framework. Most importantly, we show that the pre-existing COMRL algorithms are essentially optimizing the same mutual information objective between the task variable $M$ and its latent representation $Z$ by implementing various approximate bounds. Such theoretical insight offers ample design freedom for novel algorithms. As demonstrations, we propose a supervised and a self-supervised implementation of $I(Z; M)$, and empirically show that the corresponding optimization algorithms exhibit remarkable generalization across a broad spectrum of RL benchmarks, context shift scenarios, data qualities and deep learning architectures. This work lays the information theoretic foundation for COMRL methods, leading to a better understanding of task representation learning in the context of reinforcement learning. Given its
generality, we envision our framework as a promising offline pre-training paradigm of foundation models for decision making. Lanqing Li, Shatong Zhu, Junqiao Zhao, Pheng-Ann Heng |
NeurIPS | 7 |
| 2024 | Neural P3M: A Long-Range Interaction Modeling Enhancer for Geometric GNNsabstractGeometric graph neural networks (GNNs) have emerged as powerful tools for modeling molecular geometry. However, they encounter limitations in effectively capturing long-range interactions in large molecular systems. To address this challenge, we introduce **Neural P$^3$M**, a versatile enhancer of geometric GNNs to expand the scope of their capabilities by incorporating mesh points alongside atoms and reimaging traditional mathematical operations in a trainable manner. Neural P$^3$M exhibits flexibility across a wide range of molecular systems and demonstrates remarkable accuracy in predicting energies and forces, outperforming on benchmarks such as the MD22 dataset.
It also achieves an average improvement of 22% on the OE62 dataset while integrating with various architectures. Codes are available at https://github.com/OnlyLoveKFC/Neural_P3M. Chaoran Cheng, Shaoning Li, Yuxuan Ren, Bin Shao 0002, Pheng-Ann Heng, Nanning Zheng 0001 |
NeurIPS | 7 |
| 2024 | DPPMask: Masked Image Modeling with Determinantal Point ProcessesabstractMasked Image Modeling (MIM) has achieved impressive representative performance with the aim of reconstructing randomly masked images. Despite the empirical success, most previous works have neglected the important fact that it is unreasonable to force the model to reconstruct something beyond recovery, such as those masked objects. In this work, we show that uniformly random masking widely used in previous works unavoidably loses some key objects and changes original semantic information, resulting in a misalignment problem and hurting the representative learning eventually. To address this issue, we augment MIM with a new masking strategy namely the DPPMask by substituting the random process with Determinantal Point Process (DPPs) to reduce the semantic change of the image after masking. Our method is simple yet effective and requires no extra learnable parameters when implemented within various frameworks. In particular, we evaluate our method on two representative MIM frameworks, MAE and iBOT. We show that DPPMask surpassed random sampling under both lower and higher masking ratios, indicating that DPP-Mask makes the reconstruction task more reasonable. We further test our method on the background challenge and multi-class classification tasks, showing that our method is more robust at various tasks. Junde Xu, Zikai Lin, Yaodong Yang 0002, Xiangyun Liao, Qiong Wang 0001, Guangyong Chen, Pheng-Ann Heng |
WACV | 9 |
| 2024 | SSP: Semi-signed prioritized neural fitting for surface reconstruction from unoriented point cloudsabstractReconstructing 3D geometry from unoriented point clouds can benefit many downstream tasks. Recent shape modeling methods mostly adopt implicit neural representation to fit a signed distance field (SDF) and optimize the network by unsigned supervision. However, these methods occasionally have difficulty in finding the coarse shape for complicated objects, especially suffering from the "ghost" surfaces (i.e., fake surfaces that should not exist). To guide the network quickly fit the coarse shape, we propose to utilize the signed supervision in regions that are obviously outside the object and can be easily determined, resulting in our semi-signed supervision. To better recover high-fidelity details, a novel loss-based region sampling strategy and a progressive positional encoding (PE) method are applied to prioritize the optimization towards underfitting and complicated regions. Specifically, we voxelize and partition the object space into sign-known and sign-uncertain regions, in which different supervisions are applied. Besides, we adaptively adjust the sampling rate of each voxel according to the tracked reconstruction loss, so that the network can focus more on the complicated under-fitting regions. We conduct extensive experiments to demonstrate that our method achieves state-of-the-art performance compared to the existing fitting-based methods and comparable performance to learning-based methods on multiple datasets. The code is publicly available at https://github.com/Runsong123/SSP. Runsong Zhu, Ka-Hei Hui, Shi Qiu 0001, Linchao Bao, Pheng-Ann Heng, Chi-Wing Fu |
WACV | 8 |
| 2024 | siRNADiscovery: a graph neural network for siRNA efficacy prediction via deep RNA sequence analysisabstractThe clinical adoption of small interfering RNAs (siRNAs) has prompted the development of various computational strategies for siRNA design, from traditional data analysis to advanced machine learning techniques. However, previous studies have inadequately considered the full complexity of the siRNA silencing mechanism, neglecting critical elements such as siRNA positioning on mRNA, RNA base-pairing probabilities, and RNA-AGO2 interactions, thereby limiting the insight and accuracy of existing models. Here, we introduce siRNADiscovery, a Graph Neural Network (GNN) framework that leverages both non-empirical and empirical rule-based features of siRNA and mRNA to effectively capture the complex dynamics of gene silencing. On multiple internal datasets, siRNADiscovery achieves state-of-the-art performance. Significantly, siRNADiscovery also outperforms existing methodologies in in vitro studies and on an externally validated dataset. Additionally, we develop a new data-splitting methodology that addresses the data leakage issue, a frequently overlooked problem in previous studies, ensuring the robustness and stability of our model under various experimental settings. Through rigorous testing, siRNADiscovery has demonstrated remarkable predictive accuracy and robustness, making significant contributions to the field of gene silencing. Furthermore, our approach to redefining data-splitting standards aims to set new benchmarks for future research in the domain of predictive biological modeling for siRNA. Rongzhuo Long, Da Han, Boxiang Liu, Xudong Yuan, Guangyong Chen, Pheng-Ann Heng |
Briefings Bioinform. | 7 |
| 2024 | MA-SAM: Modality-agnostic SAM adaptation for 3D medical image segmentation
Cheng Chen 0013, Juzheng Miao, Dufan Wu, Aoxiao Zhong, Zhiling Yan, Sekeun Kim, Zhengliang Liu, Lichao Sun 0001, Xiang Li 0001, Tianming Liu 0001, Pheng-Ann Heng, Quanzheng Li |
Medical Image Anal. | 12 |
| 2024 | 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
Shizhan Gong, Yuan Zhong 0003, Wenao Ma, Jinpeng Li 0004, Zhao Wang 0006, Jingyang Zhang, Pheng-Ann Heng, Qi Dou 0001 |
Medical Image Anal. | 7 |
| 2024 | Robotic Needle Insertion With 2D Ultrasound-3D CT Fusion GuidanceabstractPuncture robots pave a new way for stable, accurate and safe percutaneous liver tumor puncture operation. However, affected by respiratory motion, intraoperative accurate location of the tumor and its surrounding anatomical structures remains a difficult problem in existing robot-assisted puncture operations. In this paper, a dual-arm robotic needle insertion system with guidance of intraoperative 2D ultrasound (US) and preoperative 3D computed tomography (CT) fusion is proposed, addressing the shortcomings of existing puncture robots. To deal with the challenge of cross-modal and cross-dimensional registration between 2D US and 3D CT, a decoupled two-stage registration approach combining initial vessel structure-based 3D US – 3D CT registration with intraoperative intensity-based 2D US -3D US registration is proposed. To achieve fast and robust ultrasound probe calibration, a method based on an improved N-wire phantom is proposed. Twenty puncture experiments are performed in different breath-holding positions on a respiratory motion simulation platform, and experimental results show that the mean puncture error is 2.48 mm, which can meet the requirements in a wide of clinical scenariosNote to Practitioners—In clinical percutaneous liver tumor puncture operation, due to the lack of real-time and clear image guidance, it is difficult to locate the tumor and its surrounding vital anatomical structures. In addition, the stability and accuracy of manual operation are poor. The development of a puncture robot is an effective solution for these problems. However, existing CT and magnetic resonance imaging (MRI) guided robots do not consider the tumor localization errors caused by inconsistent breath-holding positions between preoperative scan period and intraoperative puncture period, and US guided robots are limited by the poor image quality and the narrow field of vision. In this paper, a dual-arm robotic needle insertion system with guidance of intraoperative 2D US and preoperative 3D CT fusion is proposed. This system can take advantage of the real-time ultrasound and clear CT images at the same time, and can provide real-time, clear and all-round guidance for percutaneous liver tumor puncture operation, which has obvious advantages over the existing puncture robots. Phantom experiments have been completed and animal experiments will be carried out in the future. Long Lei, Baoliang Zhao, Xiaozhi Qi, Rui Mi, Hai Ye, Peng Zhang 0012, Qiong Wang 0001, Pheng-Ann Heng, Ying Hu 0001 |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2024 | G²Face: High-Fidelity Reversible Face Anonymization via Generative and Geometric PriorsabstractReversible face anonymization, unlike traditional face pixelization, seeks to replace sensitive identity information in facial images with synthesized alternatives, preserving privacy without sacrificing image clarity. Traditional methods, such as encoder-decoder networks, often result in significant loss of facial details due to their limited learning capacity. Additionally, relying on latent manipulation in pre-trained GANs can lead to changes in ID-irrelevant attributes, adversely affecting data utility due to GAN inversion inaccuracies. This paper introduces G2Face, which leverages both generative and geometric priors to enhance identity manipulation, achieving high-quality reversible face anonymization without compromising data utility. We utilize a 3D face model to extract geometric information from the input face, integrating it with a pre-trained GAN-based decoder. This synergy of generative and geometric priors allows the decoder to produce realistic anonymized faces with consistent geometry. Moreover, multi-scale facial features are extracted from the original face and combined with the decoder using our novel identity-aware feature fusion blocks (IFF). This integration enables precise blending of the generated facial patterns with the original ID-irrelevant features, resulting in accurate identity manipulation. Extensive experiments demonstrate that our method outperforms existing state-of-the-art techniques in face anonymization and recovery, while preserving high data utility. Code is available athttps://github.com/Harxis/G2Face. Haoxin Yang, Xuemiao Xu, Huaidong Zhang, Harry Qin, Yi Wang 0017, Pheng-Ann Heng, Shengfeng He |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | CalibNet: Dual-Branch Cross-Modal Calibration for RGB-D Salient Instance SegmentationabstractIn this study, we propose a novel approach for RGB-D salient instance segmentation using a dual-branch cross-modal feature calibration architecture called CalibNet. Our method simultaneously calibrates depth and RGB features in the kernel and mask branches to generate instance-aware kernels and mask features. CalibNet consists of three simple modules, a dynamic interactive kernel (DIK) and a weight-sharing fusion (WSF), which work together to generate effective instance-aware kernels and integrate cross-modal features. To improve the quality of depth features, we incorporate a depth similarity assessment (DSA) module prior to DIK and WSF. In addition, we further contribute a new DSIS dataset, which contains 1,940 images with elaborate instance-level annotations. Extensive experiments on three challenging benchmarks show that CalibNet yields a promising result, i.e., 58.0% AP with 320×480 input size on the COME15K-E test set, which significantly surpasses the alternative frameworks. Our code and dataset will be publicly available at: https://github.com/PJLallen/CalibNet. Jialun Pei, Tao Jiang 0002, He Tang 0002, Nian Liu 0002, Yueming Jin, Deng-Ping Fan, Pheng-Ann Heng |
IEEE Trans. Image Process. | 7 |
| 2024 | Video Instance Shadow Detection Under the Sun and SkyabstractInstance shadow detection, crucial for applications such as photo editing and light direction estimation, has undergone significant advancements in predicting shadow instances, object instances, and their associations. The extension of this task to videos presents challenges in annotating diverse video data and addressing complexities arising from occlusion and temporary disappearances within associations. In response to these challenges, we introduce ViShadow, a semi-supervised video instance shadow detection framework that leverages both labeled image data and unlabeled video data for training. ViShadow features a two-stage training pipeline: the first stage, utilizing labeled image data, identifies shadow and object instances through contrastive learning for cross-frame pairing. The second stage employs unlabeled videos, incorporating an associated cycle consistency loss to enhance tracking ability. A retrieval mechanism is introduced to manage temporary disappearances, ensuring tracking continuity. The SOBA-VID dataset, comprising unlabeled training videos and labeled testing videos, along with the SOAP-VID metric, is introduced for the quantitative evaluation of VISD solutions. The effectiveness of ViShadow is further demonstrated through various video-level applications such as video inpainting, instance cloning, shadow editing, and text-instructed shadow-object manipulation. Zhenghao Xing, Tianyu Wang 0003, Xiaowei Hu 0001, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Image Process. | 6 |
| 2024 | A Survey on Generative Diffusion ModelsabstractDeep generative models have unlocked another profound realm of human creativity. By capturing and generalizing patterns within data, we have entered the epoch of all-encompassing Artificial Intelligence for General Creativity (AIGC). Notably, diffusion models, recognized as one of the paramount generative models, materialize human ideation into tangible instances across diverse domains, encompassing imagery, text, speech, biology, and healthcare. To provide advanced and comprehensive insights into diffusion, this survey comprehensively elucidates its developmental trajectory and future directions from three distinct angles: the fundamental formulation of diffusion, algorithmic enhancements, and the manifold applications of diffusion. Each layer is meticulously explored to offer a profound comprehension of its evolution. Structured and summarized approaches are presented here. Hanqun Cao, Cheng Tan 0012, Zhangyang Gao, Guangyong Chen, Pheng-Ann Heng, Stan Z. Li |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | Deep Omni-Supervised Learning for Rib Fracture Detection From Chest Radiology ImagesabstractDeep learning (DL)-based rib fracture detection has shown promise of playing an important role in preventing mortality and improving patient outcome. Normally, developing DL-based object detection models requires a huge amount of bounding box annotation. However, annotating medical data is time-consuming and expertise-demanding, making obtaining a large amount of fine-grained annotations extremely infeasible. This poses a pressing need for developing label-efficient detection models to alleviate radiologists' labeling burden. To tackle this challenge, the literature on object detection has witnessed an increase of weakly-supervised and semi-supervised approaches, yet still lacks a unified framework that leverages various forms of fully-labeled, weakly-labeled, and unlabeled data. In this paper, we present a novel omni-supervised object detection network, ORF-Netv2, to leverage as much available supervision as possible. Specifically, a multi-branch omni-supervised detection head is introduced with each branch trained with a specific type of supervision. A co-training-based dynamic label assignment strategy is then proposed to enable flexible and robust learning from the weakly-labeled and unlabeled data. Extensive evaluation was conducted for the proposed framework with three rib fracture datasets on both chest CT and X-ray. By leveraging all forms of supervision, ORF-Netv2 achieves mAPs of 34.7, 44.7, and 19.4 on the three datasets, respectively, surpassing the baseline detector which uses only box annotations by mAP gains of 3.8, 4.8, and 5.0, respectively. Furthermore, ORF-Netv2 consistently outperforms other competitive label-efficient methods over various scenarios, showing a promising framework for label-efficient fracture detection. The code is available at: https://github.com/zhizhongchai/ORF-Net. Zhizhong Chai, Luyang Luo, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | A Dual Enrichment Synergistic Strategy to Handle Data Heterogeneity for Domain Incremental Cardiac SegmentationabstractUpon remarkable progress in cardiac image segmentation, contemporary studies dedicate to further upgrading model functionality toward perfection, through progressively exploring the sequentially delivered datasets over time by domain incremental learning. Existing works mainly concentrated on addressing the heterogeneous style variations, but overlooked the critical shape variations across domains hidden behind the sub-disease composition discrepancy. In case the updated model catastrophically forgets the sub-diseases that were learned in past domains but are no longer present in the subsequent domains, we proposed a dual enrichment synergistic strategy to incrementally broaden model competence for a growing number of sub-diseases. The data-enriched scheme aims to diversify the shape composition of current training data via displacement-aware shape encoding and decoding, to gradually build up the robustness against cross-domain shape variations. Meanwhile, the model-enriched scheme intends to strengthen model capabilities by progressively appending and consolidating the latest expertise into a dynamically-expanded multi-expert network, to gradually cultivate the generalization ability over style-variated domains. The above two schemes work in synergy to collaboratively upgrade model capabilities in two-pronged manners. We have extensively evaluated our network with the ACDC and M&Ms datasets in single-domain and compound-domain incremental learning settings. Our approach outperformed other competing methods and achieved comparable results to the upper bound. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Semantic-Oriented Visual Prompt Learning for Diabetic Retinopathy Grading on Fundus ImagesabstractDiabetic retinopathy (DR) is a serious ocular condition that requires effective monitoring and treatment by ophthalmologists. However, constructing a reliable DR grading model remains a challenging and costly task, heavily reliant on high-quality training sets and adequate hardware resources. In this paper, we investigate the knowledge transferability of large-scale pre-trained models (LPMs) to fundus images based on prompt learning to construct a DR grading model efficiently. Unlike full-tuning which fine-tunes all parameters of LPMs, prompt learning only involves a minimal number of additional learnable parameters while achieving a competitive effect as full-tuning. Inspired by visual prompt tuning, we propose Semantic-oriented Visual Prompt Learning (SVPL) to enhance the semantic perception ability for better extracting task-specific knowledge from LPMs, without any additional annotations. Specifically, SVPL assigns a group of learnable prompts for each DR level to fit the complex pathological manifestations and then aligns each prompt group to task-specific semantic space via a contrastive group alignment (CGA) module. We also propose a plug-and-play adapter module, Hierarchical Semantic Delivery (HSD), which allows the semantic transition of prompt groups from shallow to deep layers to facilitate efficient knowledge mining and model convergence. Our extensive experiments on three public DR grading datasets demonstrate that SVPL achieves superior results compared to other transfer tuning and DR grading methods. Further analysis suggests that the generalized knowledge from LPMs is advantageous for constructing the DR grading model on fundus images. Yuhan Zhang 0001, Xiao Ma 0011, Mingchao Li 0002, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Kine-Appendage: Enhancing Freehand VR Interaction Through Transformations of Virtual AppendagesabstractKinesthetic feedback, the feeling of restriction or resistance when hands contact objects, is essential for natural freehand interaction in VR. However, inducing kinesthetic feedback using mechanical hardware can be cumbersome and hard to control in commodity VR systems. We propose the kine-appendage concept to compensate for the loss of kinesthetic feedback in virtual environments, i.e., a virtual appendage is added to the user's avatar hand; when the appendage contacts a virtual object, it exhibits transformations (rotation and deformation); when it disengages from the contact, it recovers its original appearance. A proof-of-concept kine-appendage technique, BrittleStylus, was designed to enhance isomorphic typing. Our empirical evaluations demonstrated that (i) BrittleStylus significantly reduced the uncorrected error rate of naive isomorphic typing from 6.53% to 1.92% without compromising the typing speed; (ii) BrittleStylus could induce the sense of kinesthetic feedback, the degree of which was parity with that induced by pseudo-haptic (+ visual cue) methods; and (iii) participants preferred BrittleStylus over pseudo-haptic (+ visual cue) methods because of not only good performance but also fluent hand movements. Yang Tian 0008, Hualong Bai, Shengdong Zhao 0001, Chi-Wing Fu, Chun Yu, Haozhao Qin, Qiong Wang 0001, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2023 | On the Pitfall of Mixup for Uncertainty CalibrationabstractBy simply taking convex combinations between pairs of samples and their labels, mixup training has been shown to easily improve predictive accuracy. It has been recently found that models trained with mixup also perform well on uncertainty calibration. However, in this study, we found that mixup training usually makes models less calibratable than vanilla empirical risk minimization, which means that it would harm uncertainty estimation when post-hoc calibration is considered. By decomposing the mixup process into data transformation and random perturbation, we suggest that the confidence penalty nature of the data transformation is the reason of calibration degradation. To mitigate this problem, we first investigate the mixup inference strategy and found that despite it improves calibration on mixup, this ensemble-like strategy does not necessarily out-perform simple ensemble. Then, we propose a general strategy named mixup inference in training, which adopts a simple decoupling principle for recovering the outputs of raw samples at the end of forward network pass. By embedding the mixup inference, models can be learned from the original one-hot labels and hence avoid the negative impact of confidence penalty. Our experiments show this strategy properly solves mixup's calibration issue without sacrificing the predictive performance, while even improves accuracy than vanilla mixup. Dengbao Wang, Lanqing Li, Peilin Zhao, Pheng-Ann Heng, Min-Ling Zhang |
CVPR | 4 |
| 2023 | Video Dehazing via a Multi-Range Temporal Alignment Network with Physical PriorabstractVideo dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggregate temporal information. Specifically, we design a memory-based physical prior guidance module to encode the prior-related features into long-range memory. Besides, we formulate a multi-range scene radiance recovery module to capture space-time dependencies in multiple space-time ranges, which helps to effectively aggregate temporal information from adjacent frames. Moreover, we construct the first large-scale outdoor video dehazing benchmark dataset, which contains videos in various real-world scenarios. Experimental results on both synthetic and real conditions show the superiority of our proposed method. Xiaowei Hu 0001, Lei Zhu 0003, Qi Dou 0001, Jifeng Dai, Yu Qiao 0001, Pheng-Ann Heng |
CVPR | 7 |
| 2023 | RepMode: Learning to Re-Parameterize Diverse Experts for Subcellular Structure PredictionabstractIn biological research, fluorescence staining is a key technique to reveal the locations and morphology of subcellular structures. However, it is slow, expensive, and harmful to cells. In this paper, we model it as a deep learning task termed subcellular structure prediction (SSP), aiming to predict the 3D fluorescent images of multiple subcellular structures from a 3D transmitted-light image. Unfortunately, due to the limitations of current biotechnology, each image is partially labeled in SSP. Besides, naturally, subcellular structures vary considerably in size, which causes the multi-scale issue of SSP. To overcome these challenges, we propose Re-parameterizing Mixture-of-Diverse-Experts (RepMode), a network that dynamically organizes its parameters with task-aware priors to handle specified single-label prediction tasks. In RepMode, the Mixture-of-Diverse-Experts (MoDE) block is designed to learn the generalized parameters for all tasks, and gating re-parameterization (GatRep) is performed to generate the specialized parameters for each task, by which RepMode can maintain a compact practical topology exactly like a plain network, and meanwhile achieves a powerful theoretical topology. Comprehensive experiments show that RepMode can achieve state-of-the-art overall performance in SSP. Chunbin Gu, Junde Xu, Furui Liu, Qiong Wang 0001, Guangyong Chen, Pheng-Ann Heng |
CVPR | 7 |
| 2023 | Class-Conditional Sharpness-Aware Minimization for Deep Long-Tailed RecognitionabstractIt's widely acknowledged that deep learning models with flatter minima in its loss landscape tend to generalize better. However, such property is under-explored in deep long-tailed recognition (DLTR), a practical problem where the model is required to generalize equally well across all classes when trained on highly imbalanced label distribution. In this paper, through empirical observations, we argue that sharp minima are in fact prevalent in deep long-tailed models, whereas naï ve integration of existing flattening operations into long-tailed learning algorithms brings little improvement. Instead, we propose an effective two-stage sharpness-aware optimization approach based on the decoupling paradigm in DLTR. In the first stage, both the feature extractor and classifier are trained under parameter perturbations at a class-conditioned scale, which is theoretically motivated by the characteristic radius of flat minima under the PAC-Bayesian framework. In the second stage, we generate adversarial features with class-balanced sampling to further robustify the classifier with the backbone frozen. Extensive experiments on multiple long-tailed visual recognition benchmarks show that, our proposed Class-Conditional Sharpness-Aware Minimization (CC-SAM), achieves competitive performance compared to the state-of-the-arts. Code is available at https://github.com/zzpustc/CC-SAM. Lanqing Li, Peilin Zhao, Pheng-Ann Heng, Wei Gong 0001 |
CVPR | 4 |
| 2023 | Traj-MAE: Masked Autoencoders for Trajectory PredictionabstractTrajectory prediction has been a crucial task in building a reliable autonomous driving system by anticipating possible dangers. One key issue is to generate consistent trajectory predictions without colliding. To overcome the challenge, we propose an efficient masked autoencoder for trajectory prediction (Traj-MAE) that better represents the complicated behaviors of agents in the driving environment. Specifically, our Traj-MAE employs diverse masking strategies to pre-train the trajectory encoder and map encoder, allowing for the capture of social and temporal information among agents while leveraging the effect of environment from multiple granularities. To address the catastrophic forgetting problem that arises when pre-training the network with multiple masking strategies, we introduce a continual pre-training framework, which can help Traj-MAE learn valuable and diverse information from various strategies efficiently. Our experimental results in both multi-agent and single-agent settings demonstrate that Traj-MAE achieves competitive results with state-of-the-art methods and significantly outperforms our baseline model. Project page: https://jiazewang.com/projects/trajmae.html. Hao Chen 0193, Kun Shao, Furui Liu, Jianye Hao, Chenyong Guan, Guangyong Chen, Pheng-Ann Heng |
ICCV | 8 |
| 2023 | CauSSL: Causality-inspired Semi-supervised Learning for Medical Image SegmentationabstractSemi-supervised learning (SSL) has recently demonstrated great success in medical image segmentation, significantly enhancing data efficiency with limited annotations. However, despite its empirical benefits, there are still concerns in the literature about the theoretical foundation and explanation of semi-supervised segmentation. To explore this problem, this study first proposes a novel causal diagram to provide a theoretical foundation for the mainstream semi-supervised segmentation methods. Our causal diagram takes two additional intermediate variables into account, which are neglected in previous work. Drawing from this proposed causal diagram, we then introduce a causality-inspired SSL approach on top of co-training frameworks called CauSSL, to improve SSL for medical image segmentation. Specifically, we first point out the importance of algorithmic independence between two networks or branches in SSL, which is often overlooked in the literature. We then propose a novel statistical quantification of the uncomputable algorithmic independence and further enhance the independence via a min-max optimization process. Our method can be flexibly incorporated into different existing SSL methods to improve their performance. Our method has been evaluated on three challenging medical image segmentation tasks using both 2D and 3D network architectures and has shown consistent improvements over state-of-the-art methods. Our code is publicly available at: https://github.com/JuzhengMiao/CauSSL. Juzheng Miao, Cheng Chen 0013, Furui Liu, Pheng-Ann Heng |
ICCV | 5 |
| 2023 | Uncertainty Estimation by Fisher Information-based Evidential Deep LearningabstractUncertainty estimation is a key factor that makes deep learning reliable in practical applications. Recently proposed evidential neural networks explicitly account for different uncertainties by treating the network's outputs as evidence to parameterize the Dirichlet distribution, and achieve impressive performance in uncertainty estimation. However, for high data uncertainty samples but annotated with the one-hot label, the evidence-learning process for those mislabeled classes is over-penalized and remains hindered. To address this problem, we propose a novel method, Fisher Information-based Evidential Deep Learning ($\mathcal{I}$-EDL). In particular, we introduce Fisher Information Matrix (FIM) to measure the informativeness of evidence carried by each sample, according to which we can dynamically reweight the objective loss terms to make the network more focus on the representation learning of uncertain classes. The generalization ability of our network is further improved by optimizing the PAC-Bayesian bound. As demonstrated empirically, our proposed method consistently outperforms traditional EDL-related algorithms in multiple uncertainty estimation tasks, especially in the more challenging few-shot classification settings. Danruo Deng, Guangyong Chen, Furui Liu, Pheng-Ann Heng |
ICML | 5 |
| 2023 | On Improving Boundary Quality of Instance Segmentation in Cluttered and Chaotic ScenariosabstractInstance segmentation is a long-standing task for supporting robotic bin picking. However, objects of diverse classes can be closely packed with occlusions in cluttered and chaotic scenes, hence, even recent methods could have difficulty in locating clear and precise boundaries to distinguish nearby objects. In this work, we aim to improve the boundary quality of the instance masks for robust and precise instance segmentation in these challenging scenarios. Technical-wise, we first formulate an IoU-based Boundary-aware Mask head (IBM head) for predicting the instance-level mask, boundary, and their corresponding IoU scores. With this core module, we then follow the coarse-to-fine strategy and design our pipeline with two stages: an 1IoUNet to learn localization-based objectness cue and a hierarchical mask refiner to produce sharper and cleaner boundaries. We deploy the IBM head throughout the framework. Extensive experimental results on three grasping benchmarks manifest that our method attains the best instance segmentation performance, compared with the state-of-the-art approaches. Practically, we conduct real-world picking tests to show that with the objectness and boundary IoU scores as guidance, we are able to filter invalid (occluded) instances and select high-fidelity (exposed) instances for grasping. Biqi Yang, Xianzhi Li 0001, Yun-Hui Liu 0001, Chi-Wing Fu, Pheng-Ann Heng |
ICRA | 6 |
| 2023 | Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-trainingabstractMasked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric correlation between 2D and 3D. In this paper, we explore how the 2D modality can benefit 3D masked autoencoding, and propose Joint-MAE, a 2D-3D joint MAE framework for self-supervised 3D point cloud pre-training. Joint-MAE randomly masks an input 3D point cloud and its projected 2D images, and then reconstructs the masked information of the two modalities. For better cross-modal interaction, we construct our JointMAE by two hierarchical 2D-3D embedding modules, a joint encoder, and a joint decoder with modal-shared and model-specific decoders. On top of this, we further introduce two cross-modal strategies to boost the 3D representation learning, which are local-aligned attention mechanisms for 2D-3D semantic cues, and a cross-reconstruction loss for 2D-3D geometric constraints. By our pre-training paradigm, Joint-MAE achieves superior performance on multiple downstream tasks, e.g., 92.4% accuracy for linear SVM on ModelNet40 and 86.07% accuracy on the hardest split of ScanObjectNN. Renrui Zhang, Longtian Qiu, Xianzhi Li 0001, Pheng-Ann Heng |
IJCAI | 5 |
| 2023 | SDF-Pack: Towards Compact Bin Packing with Signed-Distance-Field MinimizationabstractRobotic bin packing is very challenging, especially when considering practical needs such as object variety and packing compactness. This paper presents SDF-Pack, a new approach based on signed distance field (SDF) to model the geometric condition of objects in a container and compute the object placement locations and packing orders for achieving a more compact bin packing. Our method adopts a truncated SDF representation to localize the computation, and based on it, we formulate the SDF -minimization heuristic to find optimized placements to compactly pack objects with the existing ones. To further improve space utilization, if the packing sequence is controllable, our method can suggest which object to be packed next. Experimental results on a large variety of everyday objects show that our method can consistently achieve higher packing compactness over 1,000 packing cases, enabling us to pack more objects into the container, compared with the existing heuristics under various packing settings. The code is publicly available at: https://github.com/kwpoon/SDF-Pack. Jia-Hui Pan, Ka-Hei Hui, Shize Zhu, Yun-Hui Liu 0001, Pheng-Ann Heng, Chi-Wing Fu |
IROS | 6 |
| 2023 | Fast Non-Markovian Diffusion Model for Weakly Supervised Anomaly Detection in Brain MR Images
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (5) | 7 |
| 2023 | Learning Robust Classifier for Imbalanced Medical Image Dataset with Noisy Labels by Minimizing Invariant Risk
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (6) | 7 |
| 2023 | Semi-supervised Class Imbalanced Deep Learning for Cardiac MRI Segmentation
Yuchen Yuan, Xi Wang 0013, Xikai Yang, Ruijiang Li, Pheng-Ann Heng |
MICCAI (4) | 5 |
| 2023 | SATTA: Semantic-Aware Test-Time Adaptation for Cross-Domain Medical Image Segmentation
Yuhan Zhang 0001, Cheng Chen 0013, Qiang Chen 0004, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2023 | Unite-Divide-Unite: Joint Boosting Trunk and Structure for High-accuracy Dichotomous Image SegmentationabstractHigh-accuracy Dichotomous Image Segmentation (DIS) aims to pinpoint category-agnostic foreground objects from natural scenes. The main challenge for DIS involves identifying the highly accurate dominant area while rendering detailed object structure. However, directly using a general encoder-decoder architecture may result in an oversupply of high-level features and neglect the shallow spatial information necessary for partitioning meticulous structures. To fill this gap, we introduce a novel Unite-Divide-Unite Network (UDUN) that restructures and bipartitely arranges complementary features to simultaneously boost the effectiveness of trunk and structure identification. The proposed UDUN proceeds from several strengths. First, a dual-size input feeds into the shared backbone to produce more holistic and detailed features while keeping the model lightweight. Second, a simple Divide-and-Conquer Module (DCM) is proposed to decouple multiscale low- and high-level features into our structure decoder and trunk decoder to obtain structure and trunk information respectively. Moreover, we design a Trunk-Structure Aggregation module (TSA) in our union decoder that performs cascade integration for uniform high-accuracy segmentation. As a result, UDUN performs favorably against state-of-the-art competitors in all six evaluation metrics on overall DIS-TE, i.e., achieving 0.772 weighted F-measure and 977 HCE. Using 1024X1024 input, our model enables real-time inference at 65.3 fps with ResNet-18. The source code is available at https://github.com/PJLallen/UDUN. Jialun Pei, Zhangjun Zhou, Yueming Jin, He Tang 0002, Pheng-Ann Heng |
ACM Multimedia | 5 |
| 2023 | Prototypical Variational Autoencoder for 3D Few-shot Object DetectionabstractFew-Shot 3D Point Cloud Object Detection (FS3D) is a challenging task, aiming to detect 3D objects of novel classes using only limited annotated samples for training. Considering that the detection performance highly relies on the quality of the latent features, we design a VAE-based prototype learning scheme, named prototypical VAE (P-VAE), to learn a probabilistic latent space for enhancing the diversity and distinctiveness of the sampled features. The network encodes a multi-center GMM-like posterior, in which each distribution centers at a prototype. For regularization, P-VAE incorporates a reconstruction task to preserve geometric information. To adopt P-VAE for the detection framework, we formulate Geometric-informative Prototypical VAE (GP-VAE) to handle varying geometric components and Class-specific Prototypical VAE (CP-VAE) to handle varying object categories. In the first stage, we harness GP-VAE to aid feature extraction from the input scene. In the second stage, we cluster the geometric-informative features into per-instance features and use CP-VAE to refine each instance feature with category-level guidance. Experimental results show the top performance of our approach over the state of the arts on two FS3D benchmarks. Quantitative ablations and qualitative prototype analysis further demonstrate that our probabilistic modeling can significantly boost prototype learning for FS3D. Weiliang Tang, Biqi Yang, Xianzhi Li 0001, Yun-Hui Liu 0001, Pheng-Ann Heng, Chi-Wing Fu |
NeurIPS | 5 |
| 2023 | RLMixer: A Reinforcement Learning Approach for Integrated Ranking with Contrastive User Preference Modeling
Jing Wang 0055, Mengchen Zhao, Wei Xia 0001, Zhenhua Dong, Ruiming Tang, Rui Zhang 0003, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
PAKDD (3) | 9 |
| 2023 | Semi-Supervised Intracranial Aneurysm Segmentation from CTA Images via Weight-Perceptual Self-Ensembling Model
Caizi Li, Ruiqiang Liu, Huan-Xin Zhong, Jun-Ming Fan, Weixin Si, Pheng-Ann Heng |
J. Comput. Sci. Technol. | 7 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 50 |
| 2023 | Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmarkabstractPURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery. Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt |
Medical Image Anal. | 29 |
| 2023 | Deep semi-supervised multiple instance learning with self-correction for DME classification from OCT images
Xi Wang 0013, Fangyao Tang, Hao Chen 0011, Carol Y. Cheung, Pheng-Ann Heng |
Medical Image Anal. | 5 |
| 2023 | Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification
Yuhan Zhang 0001, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2023 | Instance Shadow Detection With a Single-Stage DetectorabstractThis article formulates a new problem, instance shadow detection, which aims to detect shadow instance and the associated object instance that cast each shadow in the input image. To approach this task, we first compile a new dataset with the masks for shadow instances, object instances, and shadow-object associations. We then design an evaluation metric for quantitative evaluation of the performance of instance shadow detection. Further, we design a single-stage detector to perform instance shadow detection in an end-to-end manner, where the bidirectional relation learning module and the deformable maskIoU head are proposed in the detector to directly learn the relation between shadow instances and object instances and to improve the accuracy of the predicted masks. Finally, we quantitatively and qualitatively evaluate our method on the benchmark dataset of instance shadow detection and show the applicability of our method on light direction estimation and photo editing. Tianyu Wang 0003, Xiaowei Hu 0001, Pheng-Ann Heng, Chi-Wing Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Deep Texture-Aware Features for Camouflaged Object DetectionabstractCamouflaged object detection is a challenging task that aims to identify objects having similar texture to the surroundings. This paper presents to amplify the subtle texture difference between camouflaged objects and the background for camouflaged object detection by formulating multiple texture-aware refinement modules to learn the texture-aware features in a deep convolutional neural network. The texture-aware refinement module computes the biased co-variance matrices of feature responses to extract the texture information, adopts an affinity loss to learn a set of parameter maps that help to separate the texture between camouflaged objects and the background, and leverages a boundary-consistency loss to explore the structures of object details. We evaluate our network on the benchmark datasets for camouflaged object detection both qualitatively and quantitatively. Experimental results show that our approach outperforms various state-of-the-art methods by a large margin. Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Yangyang Xu 0003, Weiming Wang 0002, Zijun Deng, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | Representative Feature Alignment for Adaptive Object DetectionabstractUnsupervised domain adaptation for object detection aims to generalize the object detector trained on the label-rich source domain to the unlabeled target domain. Recently, existing works adopt the instance-level alignment or pixel-level alignment to perform domain transfer, which can effectively avoid the negative transfer due to the diverse background between domains. However, we find that they treat all the regions of an instance feature equally without suppressing background area. They do not segment the specific texture and discriminative regions of objects, which are transferable during adaptation. We call the features that combine the local structure feature and semantic discriminant features as representative features. We propose a novel Representative Feature Alignment (RFA) model to align the features extracted from representative patterns of objects, i.e. representative features, for domain adaptation. Specifically, the representative features are extracted by the Representative Feature Extraction (RFE) submodules. The RFE submodules take the features extracted from different intermediate layers of the detector as input, and filter out the representative features layer-by-layer via integrating class weighting generator, category selection and class activation mapping. Then the representative features from multi-layers are further adaptively aggregated to obtain the final representative features, which are utilized to conduct feature alignment in a class-aware manner. Our representative features are free of untransferable regions and background areas, which leads to better feature alignment. Extensive experimental results show that the proposed model outperforms state-of-the-art methods on a few benchmark datasets. Shan Xu 0006, Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Yangyang Xu 0003, Liangui Dai, Kup-Sze Choi, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | Dual Multiscale Mean Teacher Network for Semi-Supervised Infection Segmentation in Chest CT Volume for COVID-19abstractAutomated detecting lung infections from computed tomography (CT) data plays an important role for combating coronavirus 2019 (COVID-19). However, there are still some challenges for developing AI system: 1) most current COVID-19 infection segmentation methods mainly relied on 2-D CT images, which lack 3-D sequential constraint; 2) existing 3-D CT segmentation methods focus on single-scale representations, which do not achieve the multiple level receptive field sizes on 3-D volume; and 3) the emergent breaking out of COVID-19 makes it hard to annotate sufficient CT volumes for training deep model. To address these issues, we first build a multiple dimensional-attention convolutional neural network (MDA-CNN) to aggregate multiscale information along different dimension of input feature maps and impose supervision on multiple predictions from different convolutional neural networks (CNNs) layers. Second, we assign this MDA-CNN as a basic network into a novel dual multiscale mean teacher network (DM [Formula: see text]-Net) for semi-supervised COVID-19 lung infection segmentation on CT volumes by leveraging unlabeled data and exploring the multiscale information. Our DM [Formula: see text]-Net encourages multiple predictions at different CNN layers from the student and teacher networks to be consistent for computing a multiscale consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from multiple predictions of MDA-CNN. Third, we collect two COVID-19 segmentation datasets to evaluate our method. The experimental results show that our network consistently outperforms the compared state-of-the-art methods. Liansheng Wang 0002, Jiacheng Wang 0002, Lei Zhu 0003, Huazhu Fu, Ping Li 0016, Gary Cheng 0001, Shuo Li 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 9 |
| 2023 | Contrastive-ACE: Domain Generalization Through Alignment of Causal MechanismsabstractDomain generalization aims to learn knowledge invariant across different distributions while semantically meaningful for downstream tasks from multiple source domains, to improve the model's generalization ability on unseen target domains. The fundamental objective is to understand the underlying "invariance" behind these observational distributions and such invariance has been shown to have a close connection to causality. While many existing approaches make use of the property that causal features are invariant across domains, we consider the invariance of the average causal effect of the features to the labels. This invariance regularizes our training approach in which interventions are performed on features to enforce stability of the causal prediction by the classifier across domains. Our work thus sheds some light on the domain generalization problem by introducing invariance of the mechanisms into the learning process. Experiments on several benchmark datasets demonstrate the performance of the proposed method against SOTAs. The codes are available at: https://github.com/lithostark/Contrastive-ACE. Furui Liu, Zhitang Chen, Yik-Chung Wu, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Image Process. | 7 |
| 2023 | Domain-Incremental Cardiac Image Segmentation With Style-Oriented Replay and Domain-Sensitive Feature WhiteningabstractContemporary methods have shown promising results on cardiac image segmentation, but merely in static learning, i.e., optimizing the network once for all, ignoring potential needs for model updating. In real-world scenarios, new data continues to be gathered from multiple institutions over time and new demands keep growing to pursue more satisfying performance. The desired model should incrementally learn from each incoming dataset and progressively update with improved functionality as time goes by. As the datasets sequentially delivered from multiple sites are normally heterogenous with domain discrepancy, each updated model should not catastrophically forget previously learned domains while well generalizing to currently arrived domains or even unseen domains. In medical scenarios, this is particularly challenging as accessing or storing past data is commonly not allowed due to data privacy. To this end, we propose a novel domain-incremental learning framework to recover past domain inputs first and then regularly replay them during model optimization. Particularly, we first present a style-oriented replay module to enable structure-realistic and memory-efficient reproduction of past data, and then incorporate the replayed past data to jointly optimize the model with current data to alleviate catastrophic forgetting. During optimization, we additionally perform domain-sensitive feature whitening to suppress model's dependency on features that are sensitive to domain changes (e.g., domain-distinctive style features) to assist domain-invariant feature exploration and gradually improve the generalization performance of the network. We have extensively evaluated our approach with the M&Ms Dataset in single-domain and compound-domain incremental learning settings. Our approach outperforms other comparison methods with less forgetting on past domains and better generalization on current domains and unseen domains. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Single-Domain Generalization in Medical Image Segmentation via Test-Time Adaptation from Shape DictionaryabstractDomain generalization typically requires data from multiple source domains for model learning. However, such strong assumption may not always hold in practice, especially in medical field where the data sharing is highly concerned and sometimes prohibitive due to privacy issue. This paper studies the important yet challenging single domain generalization problem, in which a model is learned under the worst-case scenario with only one source domain to directly generalize to different unseen target domains. We present a novel approach to address this problem in medical image segmentation, which extracts and integrates the semantic shape prior information of segmentation that are invariant across domains and can be well-captured even from single domain data to facilitate segmentation under distribution shifts. Besides, a test-time adaptation strategy with dual-consistency regularization is further devised to promote dynamic incorporation of these shape priors under each unseen domain to improve model generalizability. Extensive experiments on two medical image segmentation tasks demonstrate the consistent improvements of our method across various unseen domains, as well as its superiority over state-of-the-art approaches in addressing domain generalization under the worst-case scenario. Quande Liu, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
AAAI | 4 |
| 2022 | LHNN: lattice hypergraph neural network for VLSI congestion predictionabstractPrecise congestion prediction from a placement solution plays a crucial role in circuit placement. This work proposes the lattice hypergraph (LH-graph), a novel graph formulation for circuits, which preserves netlist data during the whole learning process, and enables the congestion information propagated geometrically and topologically. Based on the formulation, we further developed a heterogeneous graph neural network architecture LHNN, jointing the routing demand regression to support the congestion spot classification. LHNN constantly achieves more than 35% improvements compared with U-nets and Pix2Pix on the F1 score. We expect our work shall highlight essential procedures using machine learning for congestion prediction. Bowen Wang 0017, Guibao Shen, Dong Li 0016, Jianye Hao, Wulong Liu, Yu Huang 0005, Hongzhong Wu, Yibo Lin, Guangyong Chen, Pheng-Ann Heng |
DAC | 10 |
| 2022 | Acknowledging the Unknown for Multi-label Learning with Single Positive Labels
Pengfei Chen 0003, Qiong Wang 0001, Guangyong Chen, Pheng-Ann Heng |
ECCV (24) | 5 |
| 2022 | Heterogeneous Graph Neural Network-Based Imitation Learning for Gate Sizing AccelerationabstractGate Sizing is an important step in logic synthesis, where the cells are resized to optimize metrics such as area, timing, power, leakage, etc. In this work, we consider the gate sizing problem for leakage power optimization with timing constraints. Lagrangian Relaxation is a widely employed optimization method for gate sizing problems. We accelerate Lagrangian Relaxation-based algorithms by narrowing down the range of cells to resize. In particular, we formulate a heterogeneous directed graph to represent the timing graph, propose a heterogeneous graph neural network as the encoder, and train in the way of imitation learning to mimic the selection behavior of each iteration in Lagrangian Relaxation. This network is used to predict the set of cells that need to be changed during the optimization process of Lagrangian Relaxation. Experiments show that our accelerated gate sizer could achieve comparable performance to the baseline with an average of 22.5% runtime reduction. Xinyi Zhou 0010, Junjie Ye 0002, Chak-Wa Pui, Kun Shao, Guangliang Zhang, Bin Wang 0034, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
ICCAD | 9 |
| 2022 | Towards Robust Part-aware Instance Segmentation for Industrial Bin PickingabstractIndustrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion. To address these challenges, we formulate a novel part-aware instance segmentation pipeline. The key idea is to decompose industrial objects into correlated approximate convex parts and enhance the object-level segmentation with part-level segmentation. We design a part-aware network to predict part masks and part-to-part offsets, followed by a part aggregation module to assemble the recognized parts into instances. To guide the network learning, we also propose an automatic label decoupling scheme to generate ground-truth part-level labels from instance-level labels. Finally, we contribute the first instance segmentation dataset, which contains a variety of industrial objects that are thin and have non-trivial shapes. Extensive experimental results on various industrial objects demonstrate that our method can achieve the best segmentation results compared with the state-of-the-art approaches. Yidan Feng, Biqi Yang, Xianzhi Li 0001, Chi-Wing Fu, Kai Chen 0028, Qi Dou 0001, Mingqiang Wei, Yun-Hui Liu 0001, Pheng-Ann Heng |
ICRA | 10 |
| 2022 | 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic SurgeryabstractAutomatic laparoscope motion control is fundamentally important for surgeons to efficiently perform operations. However, its traditional control methods based on tool tracking without considering information hidden in surgical scenes are not intelligent enough, while the latest supervised imitation learning (IL)-based methods require expensive sensor data and suffer from distribution mismatch issues caused by limited demonstrations. In this paper, we propose a novel Imitation Learning framework for Laparoscope Control (ILLC) with reinforcement learning (RL), which can efficiently learn the control policy from limited surgical video clips. Specially, we first extract surgical laparoscope trajectories from unlabeled videos as the demonstrations and reconstruct the corresponding surgical scenes. To fully learn from limited motion trajectory demonstrations, we propose Shape Preserving Trajectory Augmentation (SPTA) to augment these data, and build a simulation environment that supports parallel RGB-D rendering to reinforce the RL policy for interacting with the environment efficiently. With adversarial training for IL, we obtain the laparoscope control policy based on the generated rollouts and surgical demonstrations. Extensive experiments are conducted in unseen reconstructed surgical scenes, and our method outperforms the previous IL methods, which proves the feasibility of our unified learning-based framework for laparoscope control. Bin Li 0082, Ruofeng Wei, Bo Lu 0001, Chi Hang Yee, Chi-Fai Ng, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 7 |
| 2022 | TraSeTR: Track-to-Segment Transformer with Contrastive Query for Instance-level Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation - in general a pixel classification task - is fundamentally crucial for promoting cognitive intelligence in robot-assisted surgery (RAS). However, previous methods are struggling with discriminating instrument types and instances. To address above issues, we explore a mask classification paradigm that produces per-segment predictions. We propose TraSeTR, a novel Track-to-Segment Transformer that wisely exploits tracking cues to assist surgical instrument segmentation. TraSeTR jointly reasons about the instrument type, location, and identity with instance-level predictions i.e., a set of class-bbox-mask pairs, by decoding query embeddings. Specifically, we introduce the prior query that encoded with previous temporal knowledge, to transfer tracking signals to current instances via identity matching. A contrastive query learning strategy is further applied to reshape the query feature space, which greatly alleviates the tracking difficulty caused by large temporal variations. The effectiveness of our method is demonstrated with state-of-the-art instrument type segmentation results on three public datasets, including two RAS benchmarks from EndoVis Challenges and one cataract surgery dataset CaDIs. Yueming Jin, Pheng-Ann Heng |
ICRA | 3 |
| 2022 | RePFormer: Refinement Pyramid Transformer for Robust Facial Landmark DetectionabstractThis paper presents a Refinement Pyramid Transformer (RePFormer) for robust facial landmark detection. Most facial landmark detectors focus on learning representative image features. However, these CNN-based feature representations are not robust enough to handle complex real-world scenarios due to ignoring the internal structure of landmarks, as well as the relations between landmarks and context. In this work, we formulate the facial landmark detection task as refining landmark queries along pyramid memories. Specifically, a pyramid transformer head (PTH) is introduced to build both homologous relations among landmarks and heterologous relations between landmarks and cross-scale contexts. Besides, a dynamic landmark refinement (DLR) module is designed to decompose the landmark regression into an end-to-end refinement procedure, where the dynamically aggregated queries are transformed to residual coordinates predictions. Extensive experimental results on four facial landmark detection benchmarks and their various subsets demonstrate the superior performance and high robustness of our framework. Jinpeng Li 0004, Haibo Jin, Shengcai Liao, Ling Shao 0001, Pheng-Ann Heng |
IJCAI | 5 |
| 2022 | SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin PickingabstractInstance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any human annotations on real data, our proposed self-ensembling sim-to-real network, namely SESR, is able to generate precise instance masks for a wide variety of supermarket goods. We design our SESR with a teacher model and a student model trained with a self-ensembling strategy. We adopt different levels of consistency to bridge the sim-to-real gap and boost the model generalization ability. Also, we compile an auto-store bin-picking dataset covering various goods. Extensive experiments on both unseen scenarios and unseen objects validate the effectiveness and superiority of our method over others, and the robot arm demonstrations further show that our segmentation results can support real-time auto-store bin picking. Biqi Yang, Kai Chen 0028, Yidan Feng, Xianzhi Li 0001, Qi Dou 0001, Chi-Wing Fu, Yun-Hui Liu 0001, Pheng-Ann Heng |
IROS | 10 |
| 2022 | Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited AnnotationsabstractSurgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene segmentation from robotic surgical video, which is practically essential yet rarely explored before. We consider a clinically suitable annotation situation under the equidistant sampling. We then propose PGV-CL, a novel pseudo-label guided cross-video contrast learning method to boost scene segmentation. It effectively leverages unlabeled data for a trusty and global model regularization that produces more discriminative feature representation. Concretely, for trusty representation learning, we propose to incorporate pseudo labels to instruct the pair selection, obtaining more reliable representation pairs for pixel contrast. Moreover, we expand the representation learning space from previous image-level to cross-video, which can capture the global semantics to benefit the learning process. We extensively evaluate our method on a public robotic surgery dataset EndoVis18 and a public cataract dataset CaDIS. Experimental results demonstrate the effectiveness of our method, consistently outperforming the state-of-the-art semi-supervised methods under different labeling ratios, and even surpassing fully supervised training on EndoVis18 with 10.1% labeling. Our code is available at https://github.com/yangyu-cuhk/PGV-CL. Yang Yu 0070, Yueming Jin, Guangyong Chen, Qi Dou 0001, Pheng-Ann Heng |
IROS | 6 |
| 2022 | ORF-Net: Deep Omni-Supervised Rib Fracture Detection from Chest CT Scans
Zhizhong Chai, Huangjing Lin, Luyang Luo, Pheng-Ann Heng, Hao Chen 0011 |
MICCAI (3) | 4 |
| 2022 | Dynamic Bank Learning for Semi-supervised Federated Image Diagnosis with Class Imbalance
Meirui Jiang, Hongzheng Yang, Quande Liu, Pheng-Ann Heng, Qi Dou 0001 |
MICCAI (3) | 5 |
| 2022 | Flat-Aware Cross-Stage Distilled Framework for Imbalanced Medical Image Classification
Jinpeng Li 0004, Guangyong Chen, Hangyu Mao, Danruo Deng, Dong Li 0016, Jianye Hao, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 8 |
| 2022 | Pseudo Bias-Balanced Learning for Debiased Chest X-Ray Classification
Luyang Luo, Dunyuan Xu, Hao Chen 0011, Tien-Tsin Wong, Pheng-Ann Heng |
MICCAI (8) | 5 |
| 2022 | Transformer-based Working Memory for Multiagent Reinforcement Learning with Action ParsingabstractLearning in real-world multiagent tasks is challenging due to the usual partial observability of each agent. Previous efforts alleviate the partial observability by historical hidden states with Recurrent Neural Networks, however, they do not consider the multiagent characters that either the multiagent observation consists of a number of object entities or the action space shows clear entity interactions. To tackle these issues, we propose the Agent Transformer Memory (ATM) network with a transformer-based memory. First, ATM utilizes the transformer to enable the unified processing of the factored environmental entities and memory. Inspired by the human’s working memory process where a limited capacity of information temporarily held in mind can effectively guide the decision-making, ATM updates its fixed-capacity memory with the working memory updating schema. Second, as agents' each action has its particular interaction entities in the environment, ATM parses the action space to introduce this action’s semantic inductive bias by binding each action with its specified involving entity to predict the state-action value or logit. Extensive experiments on the challenging SMAC and Level-Based Foraging environments validate that ATM could boost existing multiagent RL algorithms with impressive learning acceleration and performance improvement. Yaodong Yang 0002, Guangyong Chen, Weixun Wang, Xiaotian Hao, Jianye Hao, Pheng-Ann Heng |
NeurIPS | 6 |
| 2022 | Amplitude-frequency-aware deep fusion network for optimal contact selection on STN-DBS electrodes
Linxia Xiao, Caizi Li, Yanjiang Wang 0001, Weixin Si, Doudou Zhang, Xiaodong Cai, Pheng-Ann Heng |
Sci. China Inf. Sci. | 8 |
| 2022 | Towards reliable cardiac image segmentation: Assessing image-level and pixel-level segmentation quality via self-reflective references
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2022 | Real-time landmark detection for precise endoscopic submucosal dissection via shape-aware relation network
Jiacheng Wang 0002, Yueming Jin, Shuntian Cai, Hongzhi Xu, Pheng-Ann Heng, Harry Qin, Liansheng Wang 0002 |
Medical Image Anal. | 5 |
| 2022 | Unsupervised feature disentanglement for video retrieval in minimally invasive surgery
Ziyi Wang 0006, Bo Lu 0001, Yueming Jin, Zerui Wang, Tak Hong Cheung, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
Medical Image Anal. | 7 |
| 2022 | Toward Image-Guided Automated Suture Grasping Under Complex Environments: A Learning-Enabled and Optimization-Based Holistic FrameworkabstractTo realize a higher-level autonomy of surgical knot tying in minimally invasive surgery (MIS), automated suture grasping, which bridges the suture stitching and looping procedures, is an important yet challenging task needs to be achieved. This paper presents a holistic framework with image-guided and automation techniques to robotize this operation even under complex environments. The whole task is initialized by suture segmentation, in which we propose a novel semi-supervised learning architecture featured with a suture-aware loss to pertinently learn its slender information using both annotated and unannotated data. With successful segmentation in stereo-camera, we develop a Sampling-based Sliding Pairing (SSP) algorithm to online optimize the suture’s 3D shape. By jointly studying the robotic configuration and the suture’s spatial characteristics, a target function is introduced to find the optimal grasping pose of the surgical tool with Remote Center of Motion (RCM) constraints. To compensate for inherent errors and practical uncertainties, a unified grasping strategy with a novel vision-based mechanism is introduced to autonomously accomplish this grasping task. Our framework is extensively evaluated from learning-based segmentation, 3D reconstruction, and image-guided grasping on the da Vinci Research Kit (dVRK) platform, where we achieve high performances and successful rates in perceptions and robotic manipulations. These results prove the feasibility of our approach in automating the suture grasping task, and this work fills the gap between automated surgical stitching and looping, stepping towards a higher-level of task autonomy in surgical knot tying. Note to Practitioners—This paper aims to automate the suture grasping task in surgical knot tying by leveraging stereo visual guidance. To effectively robotize this procedure, it requires multidisciplinary knowledge to achieve suture segmentation, 3D shape reconstruction, and reliable automated grasping, while there are no existing works tackling this procedure especially using robots with RCM kinematics constraints and under complex environments. In this article, we propose a learning-driven method along with a 3D shape optimizer, which can conduct the suture segmentation and output its accurate spatial coordinates, serving as guidance for automated grasping operation. Apart from this, we introduce a unified function to optimize the grasping pose, and a vision-based grasping strategy is also proposed to intelligently complete this task. The experiments extensively validate the feasibility of our framework for automated suture grasp, and its successful completion can serve as a basis for the following looping manipulation, hence filling a step gap in robot-assisted knot tying. This framework can be also encapsulated into the medical robotic system, and by simply indicating (e.g. mouse click) the rough position of the suture’s tip in one camera frame, the overall framework can be initialized and further accomplish the suture grasping task, which further prompts a full autonomy of surgical knot tying in the near future. Bo Lu 0001, Bin Li 0082, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2022 | Learning With Privileged Multimodal Knowledge for Unimodal SegmentationabstractMultimodal learning usually requires a complete set of modalities during inference to maintain performance. Although training data can be well-prepared with high-quality multiple modalities, in many cases of clinical practice, only one modality can be acquired and important clinical evaluations have to be made based on the limited single modality information. In this work, we propose a privileged knowledge learning framework with the 'Teacher-Student' architecture, in which the complete multimodal knowledge that is only available in the training data (called privileged information) is transferred from a multimodal teacher network to a unimodal student network, via both a pixel-level and an image-level distillation scheme. Specifically, for the pixel-level distillation, we introduce a regularized knowledge distillation loss which encourages the student to mimic the teacher's softened outputs in a pixel-wise manner and incorporates a regularization factor to reduce the effect of incorrect predictions from the teacher. For the image-level distillation, we propose a contrastive knowledge distillation loss which encodes image-level structured information to enrich the knowledge encoding in combination with the pixel-level distillation. We extensively evaluate our method on two different multi-class segmentation tasks, i.e., cardiac substructure segmentation and brain tumor segmentation. Experimental results on both tasks demonstrate that our privileged knowledge learning is effective in improving unimodal segmentation and outperforms previous methods. Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Quande Liu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene SegmentationabstractAutomatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL. Yueming Jin, Yang Yu 0070, Cheng Chen 0013, Pheng-Ann Heng, Danail Stoyanov |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Robust Medical Image Classification From Noisy Labeled Data With Global and Local Representation Guided Co-TrainingabstractDeep neural networks have achieved remarkable success in a wide variety of natural image and medical image computing tasks. However, these achievements indispensably rely on accurately annotated training data. If encountering some noisy-labeled images, the network training procedure would suffer from difficulties, leading to a sub-optimal classifier. This problem is even more severe in the medical image analysis field, as the annotation quality of medical images heavily relies on the expertise and experience of annotators. In this paper, we propose a novel collaborative training paradigm with global and local representation learning for robust medical image classification from noisy-labeled data to combat the lack of high quality annotated medical data. Specifically, we employ the self-ensemble model with a noisy label filter to efficiently select the clean and noisy samples. Then, the clean samples are trained by a collaborative training strategy to eliminate the disturbance from imperfect labeled samples. Notably, we further design a novel global and local representation learning scheme to implicitly regularize the networks to utilize noisy samples in a self-supervised manner. We evaluated our proposed robust learning strategy on four public medical image classification datasets with three types of label noise, i.e., random noise, computer-generated label noise, and inter-observer variability noise. Our method outperforms other learning from noisy label methods and we also conducted extensive experiments to analyze each component of our method. Cheng Xue 0003, Lequan Yu, Pengfei Chen 0003, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2022 | DLTTA: Dynamic Learning Rate for Test-Time Adaptation on Cross-Domain Medical ImagesabstractTest-time adaptation (TTA) has increasingly been an important topic to efficiently tackle the cross-domain distribution shift at test time for medical images from different institutions. Previous TTA methods have a common limitation of using a fixed learning rate for all the test samples. Such a practice would be sub-optimal for TTA, because test data may arrive sequentially therefore the scale of distribution shift would change frequently. To address this problem, we propose a novel dynamic learning rate adjustment method for test-time adaptation, called DLTTA, which dynamically modulates the amount of weights update for each test image to account for the differences in their distribution shift. Specifically, our DLTTA is equipped with a memory bank based estimation scheme to effectively measure the discrepancy of a given test sample. Based on this estimated discrepancy, a dynamic learning rate adjustment strategy is then developed to achieve a suitable degree of adaptation for each test sample. The effectiveness and general applicability of our DLTTA is extensively demonstrated on three tasks including retinal optical coherence tomography (OCT) segmentation, histopathological image classification, and prostate 3D MRI segmentation. Our method achieves effective and fast test-time adaptation with consistent performance improvement over current state-of-the-art test-time adaptation methods. Code is available at https://github.com/med-air/DLTTA. Hongzheng Yang, Cheng Chen 0013, Meirui Jiang, Quande Liu, Jianfeng Cao, Pheng-Ann Heng, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Item Relationship Graph Neural Networks for E-CommerceabstractIn a modern e-commerce recommender system, it is important to understand the relationships among products. Recognizing product relationships-such as complements or substitutes-accurately is an essential task for generating better recommendation results, as well as improving explainability in recommendation. Products and their associated relationships naturally form a product graph, yet existing efforts do not fully exploit the product graph's topological structure. They usually only consider the information from directly connected products. In fact, the connectivity of products a few hops away also contains rich semantics and could be utilized for improved relationship prediction. In this work, we formulate the problem as a multilabel link prediction task and propose a novel graph neural network-based framework, item relationship graph neural network (IRGNN), for discovering multiple complex relationships simultaneously. We incorporate multihop relationships of products by recursively updating node embeddings using the messages from their neighbors. An edge relational network is designed to effectively capture relational information between products. Extensive experiments are conducted on real-world product data, validating the effectiveness of IRGNN, especially on large and sparse product graphs. Weiwen Liu, Yin Zhang 0011, Jianling Wang, Yun He 0001, James Caverlee, Patrick P. K. Chan, Daniel S. Yeung, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2022 | A Rotation-Invariant Framework for Deep Point Cloud AnalysisabstractRecently, many deep neural networks were designed to process 3D point clouds, but a common drawback is that rotation invariance is not ensured, leading to poor generalization to arbitrary orientations. In this article, we introduce a new low-level purely rotation-invariant representation to replace common 3D Cartesian coordinates as the network inputs. Also, we present a network architecture to embed these representations into features, encoding local relations between points and their neighbors, and the global shape structure. To alleviate inevitable global information loss caused by the rotation-invariant representations, we further introduce a region relation convolution to encode local and non-local information. We evaluate our method on multiple point cloud analysis tasks, including (i) shape classification, (ii) part segmentation, and (iii) shape retrieval. Extensive experimental results show that our method achieves consistent, and also the best performance, on inputs at arbitrary orientations, compared with all the state-of-the-art methods. Xianzhi Li 0001, Ruihui Li, Guangyong Chen, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Beyond Class-Conditional Assumption: A Primary Attempt to Combat Instance-Dependent Label NoiseabstractSupervised learning under label noise has seen numerous advances recently, while existing theoretical findings and empirical results broadly build up on the class-conditional noise (CCN) assumption that the noise is independent of input features given the true label. In this work, we present a theoretical hypothesis testing and prove that noise in real-world dataset is unlikely to be CCN, which confirms that label noise should depend on the instance and justifies the urgent need to go beyond the CCN assumption.The theoretical results motivate us to study the more general and practical-relevant instance-dependent noise (IDN). To stimulate the development of theory and methodology on IDN, we formalize an algorithm to generate controllable IDN and present both theoretical and empirical evidence to show that IDN is semantically meaningful and challenging. As a primary attempt to combat IDN, we present a tiny algorithm termed self-evolution average label (SEAL), which not only stands out under IDN with various noise fractions, but also improves the generalization on real-world noise benchmark Clothing1M. Our code is released. Notably, our theoretical analysis in Section 2 provides rigorous motivations for studying IDN, which is an important topic that deserves more research attention in future. Pengfei Chen 0003, Junjie Ye 0002, Guangyong Chen, Pheng-Ann Heng |
AAAI | 5 |
| 2021 | Robustness of Accuracy Metric and its Inspirations in Learning with Noisy LabelsabstractFor multi-class classification under class-conditional label noise, we prove that the accuracy metric itself can be robust. We concretize this finding's inspiration in two essential aspects: training and validation, with which we address critical issues in learning with noisy labels. For training, we show that maximizing training accuracy on sufficiently many noisy samples yields an approximately optimal classifier. For validation, we prove that a noisy validation set is reliable, addressing the critical demand of model selection in scenarios like hyperparameter-tuning and early stopping. Previously, model selection using noisy validation samples has not been theoretically justified. We verify our theoretical results and additional claims with extensive experiments. We show characterizations of models trained with noisy labels, motivated by our theoretical results, and verify the utility of a noisy validation set by showing the impressive performance of a framework termed noisy best teacher and student (NTS). Our code is released. Pengfei Chen 0003, Junjie Ye 0002, Guangyong Chen, Pheng-Ann Heng |
AAAI | 5 |
| 2021 | Learning Semantic Context from Normal Samples for Unsupervised Anomaly DetectionabstractUnsupervised anomaly detection aims to identify data samples that have low probability density from a set of input samples, and only the normal samples are provided for model training. The inference of abnormal regions on the input image requires an understanding of the surrounding semantic context. This work presents a Semantic Context based Anomaly Detection Network, SCADN, for unsupervised anomaly detection by learning the semantic context from the normal samples. To achieve this, we first generate multi-scale striped masks to remove a part of regions from the normal samples, and then train a generative adversarial network to reconstruct the unseen regions. Note that the masks are designed in multiple scales and stripe directions, and various training examples are generated to obtain the rich semantic context . In testing, we obtain an error map by computing the difference between the reconstructed image and the input image for all samples, and infer the abnormal samples based on the error maps. Finally, we perform various experiments on three public benchmark datasets and a new dataset LaceAD collected by us, and show that our method clearly outperforms the current state-of-the-art methods. Huaidong Zhang, Xuemiao Xu, Xiaowei Hu 0001, Pheng-Ann Heng |
AAAI | 5 |
| 2021 | Single-Stage Instance Shadow Detection With Bidirectional Relation LearningabstractInstance shadow detection aims to find shadow instances paired with the objects that cast the shadows. The previous work adopts a two-stage framework to first predict shadow instances, object instances, and shadow-object associations from the region proposals, then leverage a post-processing to match the predictions to form the final shadow-object pairs. In this paper, we present a new single-stage fully-convolutional network architecture with a bidirectional relation learning module to directly learn the relations of shadow and object instances in an end-to-end manner. Compared with the prior work, our method actively explores the internal relationship between shadows and objects to learn a better pairing between them, thus improving the overall performance for instance shadow detection. We evaluate our method on the benchmark dataset for instance shadow detection, both quantitatively and visually. The experimental results demonstrate that our method clearly outperforms the state-of-the-art method. Tianyu Wang 0003, Xiaowei Hu 0001, Chi-Wing Fu, Pheng-Ann Heng |
CVPR | 4 |
| 2021 | Point Cloud Upsampling via Disentangled RefinementabstractPoint clouds produced by 3D scanning are often sparse, non-uniform, and noisy. Recent upsampling approaches aim to generate a dense point set, while achieving both distribution uniformity and proximity-to-surface, and possibly amending small holes, all in a single network. After revisiting the task, we propose to disentangle the task based on its multi-objective nature and formulate two cascaded sub-networks, a dense generator and a spatial refiner. The dense generator infers a coarse but dense out-put that roughly describes the underlying surface, while the spatial refiner further fine-tunes the coarse output by adjusting the location of each point. Specifically, we design a pair of local and global refinement units in the spatial refiner to evolve a coarse feature map. Also, in the spatial refiner, we regress a per-point offset vector to further adjust the coarse outputs in fine scale. Extensive qualitative and quantitative results on both synthetic and real-scanned datasets demonstrate the superiority of our method over the state-of-the-arts. The code is publicly available at https://github.com/liruihui/Dis-PU. Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 3 |
| 2021 | FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency SpaceabstractFederated learning allows distributed medical institutions to collaboratively learn a shared prediction model with privacy protection. While at clinical deployment, the models trained in federated learning can still suffer from performance drop when applied to completely unseen hospitals outside the federation. In this paper, we point out and solve a novel problem setting of federated domain generalization (FedDG), which aims to learn a federated model from multiple distributed source domains such that it can directly generalize to unseen target domains. We present a novel approach, named as Episodic Learning in Continuous Frequency Space (ELCFS), for this problem by enabling each client to exploit multi-source data distributions under the challenging constraint of data decentralization. Our approach transmits the distribution information across clients in a privacy-protecting way through an effective continuous frequency space interpolation mechanism. With the transferred multi-source distributions, we further carefully design a boundary-oriented episodic learning paradigm to expose the local learning to domain distribution shifts and particularly meet the challenges of model generalization in medical image segmentation scenario. The effectiveness of our method is demonstrated with superior performance over state-of-the-arts and in-depth ablation experiments on two medical image segmentation tasks. The code is available at https://github.com/liuquande/FedDG-ELCFS. Quande Liu, Cheng Chen 0013, Harry Qin, Qi Dou 0001, Pheng-Ann Heng |
CVPR | 5 |
| 2021 | C3-SemiSeg: Contrastive Semi-supervised Segmentation via Cross-set Learning and Dynamic Class-balancingabstractThe semi-supervised semantic segmentation methods utilize the unlabeled data to increase the feature discriminative ability to alleviate the burden of the annotated data. However, the dominant consistency learning diagram is limited by a) the misalignment between features from labeled and unlabeled data; b) treating each image and region separately without considering crucial semantic dependencies among classes. In this work, we introduce a novel C3-SemiSeg to improve consistency-based semi-supervised learning by exploiting better feature alignment under perturbations and enhancing the capability of discriminative feature cross images. Specifically, we first introduce a cross-set region-level data augmentation strategy to reduce the feature discrepancy between labeled data and unlabeled data. Cross-set pixel-wise contrastive learning is further integrated into the pipeline to facilitate feature representation ability. To stabilize training from the noisy label, we propose a dynamic confidence region selection strategy to focus on the high confidence region for loss calculation. We validate the proposed approach on Cityscapes and BDD100K dataset, which significantly outperforms other state-of-the-art semi-supervised semantic segmentation methods. Yanning Zhou 0001, Hang Xu 0004, Wei Zhang 0196, Pheng-Ann Heng |
ICCV | 5 |
| 2021 | Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningabstractThis paper presents a self-supervised method for learning reliable visual correspondence from unlabeled videos. We formulate the correspondence as finding paths in a joint space-time graph, where nodes are grid patches sampled from frames, and are linked by two type of edges: (i) neighbor relations that determine the aggregation strength from intra-frame neighbors in space, and (ii) similarity relations that indicate the transition probability of inter-frame paths across time. Leveraging the cycle-consistency in videos, our contrastive learning objective discriminates dynamic objects from both their neighboring views and temporal views. Compared with prior works, our approach actively explores the neighbor relations of central instances to learn a latent association between center-neighbor pairs (e.g., "hand – arm") across time, thus improving the instance discrimination. Without fine-tuning, our learned representation outperforms the state-of-the-art self-supervised methods on a variety of visual tasks including video object propagation, part propagation, and pose keypoint tracking. Our self-supervised method also surpasses some fully supervised algorithms designed for the specific tasks. Yueming Jin, Pheng-Ann Heng |
ICCV | 3 |
| 2021 | Noise against noise: stochastic label noise helps combat inherent label noise
Pengfei Chen 0003, Guangyong Chen, Junjie Ye 0002, Pheng-Ann Heng |
ICLR | 5 |
| 2021 | Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic SurgeryabstractAutomatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html. Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001 |
ICRA | 7 |
| 2021 | One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical VideoabstractSurgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is impractical and expensive to collect and annotate sufficient data from every new domain. To greatly increase the label efficiency, we explore a new problem, i.e., adaptive instrument segmentation, which is to effectively adapt one source model to new robotic surgical videos from multiple target domains, only given the annotated instruments in the first frame. We propose MDAL, a meta-learning based dynamic online adaptive learning scheme with a two-stage framework to fast adapt the model parameters on the first frame and partial subsequent frames while predicting the results. MDAL learns the general knowledge of instruments and the fast adaptation ability through the video-specific meta-learning paradigm. The added gradient gate excludes the noisy supervision from pseudo masks for dynamic online adaptation on target videos. We demonstrate empirically that MDAL outperforms other state-of-the-art methods on two datasets (including a real-world RAS dataset). The promising performance on ex-vivo scenes also benefits the downstream tasks such as robot-assisted suturing and camera control. Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Qi Dou 0001, Yun-Hui Liu 0001, Pheng-Ann Heng |
ICRA | 7 |
| 2021 | Accurate Grid Keypoint Learning for Efficient Video PredictionabstractVideo prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only as continuous coordinates, so noise from semantically insignificant deviations in videos easily disrupt learning stability, leading to inaccurate keypoint modeling. In this paper, we design a new grid keypoint learning framework, aiming at a robust and explainable intermediate keypoint representation for long-term efficient video prediction. We have two major technical contributions. First, we detect keypoints by jumping among candidate locations in our raised grid space and formulate a condensation loss to encourage meaningful keypoints with strong representative capability. Second, we introduce a 2D binary map to represent the detected grid keypoints and then suggest propagating keypoint locations with stochasticity by selecting entries in the discrete grid space, thus preserving the spatial structure of keypoints in the long-term horizon for better future frame generation. Extensive experiments verify that our method outperforms the state-of-the-art stochastic video prediction methods while saves more than 98% of computing resources. We also demonstrate our method on a robotic-assisted surgery dataset with promising results. Our code is available at https://github.com/xjgaocs/Grid-Keypoint-Learning. Yueming Jin, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IROS | 5 |
| 2021 | Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data AlignmentabstractAutomated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance degradation when a domain gap exists, such as the simulator and real robot. In this paper, we propose a novel unsupervised domain adaptation framework which can simultaneously transfer multi-modality knowledge, i.e., both kinematic and visual data, from simulator to real robot. It remedies the domain gap with enhanced transferable features by using temporal cues in videos, and inherent correlations in multi-modal towards recognizing gesture. Specifically, we first propose a Motion Direction Oriented Kinematics feature alignment (MDO-K) to align kinematics, which exploits temporal continuity to transfer motion directions with smaller gap rather than position values, relieving the adaptation burden. Moreover, we propose a Kinematic and Visual Relation Attention (KV-Relation-ATT) to transfer the co-occurrence signals of kinematics and vision. Such features attended by correlation similarity are more informative for enhancing domain-irreverent of the model. Two feature alignment strategies benefit the model mutually during the end-to-end learning process. We extensively evaluate our method for gesture recognition using DESK dataset with peg transfer procedure. Results show that our approach recovers the performance with great improvement gains, up to 12.91% in Accuracy and 20.16% in F1score without using any annotations in real robot. Xueying Shi, Yueming Jin, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IROS | 5 |
| 2021 | SurRoL: An Open-source Reinforcement Learning Centered and dVRK Compatible Platform for Surgical Robot LearningabstractAutonomous surgical execution relieves tedious routines and surgeon’s fatigue. Recent learning-based methods, especially reinforcement learning (RL) based methods, achieve promising performance for dexterous manipulation, which usually requires the simulation to collect data efficiently and reduce the hardware cost. The existing learning-based simulation platforms for medical robots suffer from limited scenarios and simplified physical interactions, which degrades the real-world performance of learned policies. In this work, we designed SurRoL, an RL-centered simulation platform for surgical robot learning compatible with the da Vinci Research Kit (dVRK). The designed SurRoL integrates a user-friendly RL library for algorithm development and a real-time physics engine, which is able to support more PSM/ECM scenarios and more realistic physical interactions. Ten learning-based surgical tasks are built in the platform, which are common in the real autonomous surgical execution. We evaluate SurRoL using RL algorithms in simulation, provide in-depth analysis, deploy the trained policies on the real dVRK, and show that our SurRoL achieves better transferability in the real world. Bin Li 0082, Bo Lu 0001, Yun-Hui Liu 0001, Qi Dou 0001, Pheng-Ann Heng |
IROS | 6 |
| 2021 | Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling
Cheng Chen 0013, Quande Liu, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 5 |
| 2021 | Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
Yueming Jin, Yonghao Long 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (4) | 5 |
| 2021 | Dual-Consistency Semi-supervised Learning with Uncertainty Quantification for COVID-19 Lesion Segmentation from CT Images
Yanwen Li, Luyang Luo, Huangjing Lin, Hao Chen 0011, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2021 | Federated Semi-supervised Medical Image Classification via Inter-client Relation Matching
Quande Liu, Hongzheng Yang, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 4 |
| 2021 | OXnet: Deep Omni-Supervised Thoracic Disease Detection from Chest X-Rays
Luyang Luo, Hao Chen 0011, Yanning Zhou 0001, Huangjing Lin, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2021 | Efficient Global-Local Memory for Real-Time Instrument Segmentation of Robotic Surgical Video
Jiacheng Wang 0002, Yueming Jin, Liansheng Wang 0002, Shuntian Cai, Pheng-Ann Heng, Harry Qin |
MICCAI (4) | 5 |
| 2021 | Learning Regularizer for Monocular Depth Estimation with Adversarial GuidanceabstractMonocular Depth Estimation (MDE) is a fundamental task in computer vision and multimedia. With the wide applications of deep Convolutional Neural Networks (CNNs), learning-based methods have achieved superior performance on MDE tasks in recent years. Because loss functions are important to train an accurate CNN with good generalization performance, nearly all previous efforts contribute to proposing powerful loss functions with careful hand-crafted regularizers(e.g., gradient loss and normal loss) added to the basic depth L1-Loss. However, the hand-crafted regularizers require rich domain knowledge, while their performance can still not be guaranteed. In this paper, we learn a new regularizer, approximated by a tiny CNN Regularizrer-Net(RN), and train it in an adversarial way. As demonstrated experimentally, our learned regularizer can notably outperform the current state-of-the-art methods by both quantitative evaluation and qualitative visualization on the benchmark NYU-Depth-v2 dataset, and well generalize to the new ScanNet dataset without any further training. Our code will be released soon. Guibao Shen, Yingkui Zhang, Mingqiang Wei, Qiong Wang 0001, Guangyong Chen, Pheng-Ann Heng |
ACM Multimedia | 7 |
| 2021 | Flattening Sharpness for Dynamic Gradient Projection Memory Benefits Continual LearningabstractThe backpropagation networks are notably susceptible to catastrophic forgetting, where networks tend to forget previously learned skills upon learning new ones. To address such the 'sensitivity-stability' dilemma, most previous efforts have been contributed to minimizing the empirical risk with different parameter regularization terms and episodic memory, but rarely exploring the usages of the weight loss landscape. In this paper, we investigate the relationship between the weight loss landscape and sensitivity-stability in the continual learning scenario, based on which, we propose a novel method, Flattening Sharpness for Dynamic Gradient Projection Memory (FS-DGPM). In particular, we introduce a soft weight to represent the importance of each basis representing past tasks in GPM, which can be adaptively learned during the learning process, so that less important bases can be dynamically released to improve the sensitivity of new skill learning. We further introduce Flattening Sharpness (FS) to reduce the generalization gap by explicitly regulating the flatness of the weight loss landscape of all seen tasks. As demonstrated empirically, our proposed method consistently outperforms baselines with the superior ability to learn new skills while alleviating forgetting effectively. Danruo Deng, Guangyong Chen, Jianye Hao, Qiong Wang 0001, Pheng-Ann Heng |
NeurIPS | 5 |
| 2021 | Deep virtual adversarial self-training with consistency regularization for semi-supervised medical image classification
Xi Wang 0013, Hao Chen 0011, Huiling Xiang, Huangjing Lin, Pheng-Ann Heng |
Medical Image Anal. | 6 |
| 2021 | Dual-path network with synergistic grouping loss and evidence driven risk stratification for whole slide cervical image analysis
Huangjing Lin, Hao Chen 0011, Xi Wang 0013, Qiong Wang 0001, Liansheng Wang 0002, Pheng-Ann Heng |
Medical Image Anal. | 6 |
| 2021 | Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challengeabstractIntraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts). Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein |
Medical Image Anal. | 21 |
| 2021 | Semi-supervised learning with progressive unlabeled data excavation for label-efficient surgical workflow recognition
Xueying Shi, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2021 | A global benchmark of algorithms for segmenting the left atrium from late gadolinium-enhanced cardiac magnetic resonance imaging
Zhaohan Xiong, Qing Xia 0002, Cheng Bian, Yefeng Zheng 0001, Sulaiman Vesal, Nishant Ravikumar, Andreas K. Maier, Xin Yang 0009, Pheng-Ann Heng, Dong Ni 0001, Caizi Li, Qianqian Tong 0001, Weixin Si, Élodie Puybareau, Younes Khoudli, Thierry Géraud, Jichao Zhao |
Medical Image Anal. | 11 |
| 2021 | Global guidance network for breast lesion segmentation in ultrasound images
Cheng Xue 0003, Lei Zhu 0003, Huazhu Fu, Xiaowei Hu 0001, Xiaomeng Li 0001, Pheng-Ann Heng |
Medical Image Anal. | 7 |
| 2021 | Anchor-guided online meta adaptation for fast one-Shot instrument segmentation from robotic surgical videos
Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Yun-Hui Liu 0001, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 8 |
| 2021 | SAC-Net: Spatial Attenuation Context for Salient Object DetectionabstractThis paper presents a new deep neural network design for salient object detection by maximizing the integration of local and global image context within, around, and beyond the salient objects. Our key idea is to adaptively propagate and aggregate the image context features with variable attenuation over the entire feature maps. To achieve this, we design the spatial attenuation context (SAC) module to recurrently translate and aggregate the context features independently with different attenuation factors and then to attentively learn the weights to adaptively integrate the aggregated context features. By further embedding the module to process individual layers in a deep network, namely SAC-Net, we can train the network end-To-end and optimize the context features for detecting salient objects. Compared with 29 state-of-The-Art methods, experimental results show that our method performs favorably over all the others on six common benchmark data, both quantitatively and visually. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Tianyu Wang 0003, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Learning Gated Non-Local Residual for Single-Image Rain Streak RemovalabstractThis work presents a gated non-local deep residual learning framework for image deraining. It can avoid the over-deraining or under-deraining caused by the global residual learning in existing deraining networks, since the learned soft gate in our method adaptively adjusts the amount of global residual to be passed for generating the final derained result. To generate feature maps for global residual prediction, we develop a non-local guided attention module (NLAM), which first obtains non-local features by exploiting spatial inter-dependencies among all the feature positions of local features produced by convolutional neural network (CNN), and then leverages the attention mechanism to merge the local and non-local features based on their complementary relation. Moreover, we develop a channel-wise gated prediction module to learn a soft gate on the global residual by explicitly modelling channel inter-dependencies of the feature maps obtained from NLAM. Experiments on four deraining benchmark datasets and real-world rainy images show that our network has a quantitative and qualitative improvement over state-of-the-arts. Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Haoran Xie 0001, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Multitask Feature Learning Meets Robust Tensor Decomposition for EEG ClassificationabstractIn this article, we study a tensor-based multitask learning (MTL) method for classification. Taking into account the fact that in many real-world applications, the given training samples are limited and can be inherently arranged into multidimensional arrays (tensors), we are motivated by the advantages of MTL, where the shared structural information among related tasks can be leveraged to produce better generalization performance. We propose a regularized tensor-based MTL method for joint feature selection and classification. For feature selection, we employ the Fisher discriminant criterion to both select discriminative features and control the within-class nonstationarity. For classification, we take both shared and task-specific structural information into consideration. We decompose the regression tensor for each task into a linear combination of a shared tensor and a task-specific tensor and propose a composite tensor norm. Specifically, we use the scaled latent trace norm for regularizing the shared tensor and the$\ell _{1}$-norm for task-specific tensor. Further, we give a computationally efficient optimization algorithm based on the alternating direction method of multipliers (ADMMs) to tackle the joint learning of discriminative features and multitask classification. The experimental results on real electroencephalography (EEG) datasets demonstrate the superiority of our method over the state-of-the-art techniques. Qingqing Zheng, Yi Wang 0031, Pheng-Ann Heng |
IEEE Trans. Cybern. | 3 |
| 2021 | Revisiting Shadow Detection: A New Benchmark Dataset for Complex WorldabstractShadow detection in general photos is a nontrivial problem, due to the complexity of the real world. Though recent shadow detectors have already achieved remarkable performance on various benchmark data, their performance is still limited for general real-world situations. In this work, we collected shadow images for multiple scenarios and compiled a new dataset of 10,500 shadow images, each with labeled ground-truth mask, for supporting shadow detection in the complex world. Our dataset covers a rich variety of scene categories, with diverse shadow sizes, locations, contrasts, and types. Further, we comprehensively analyze the complexity of the dataset, present a fast shadow detection network with a detail enhancement module to harvest shadow details, and demonstrate the effectiveness of our method to detect shadows in general situations. Xiaowei Hu 0001, Tianyu Wang 0003, Chi-Wing Fu, Yitong Jiang, Qiong Wang 0001, Pheng-Ann Heng |
IEEE Trans. Image Process. | 6 |
| 2021 | Single-Image Real-Time Rain Removal Based on Depth-Guided Non-Local FeaturesabstractRain is a common weather phenomenon that affects environmental monitoring and surveillance systems. According to an established rain model (Garg and Nayar, 2007), the scene visibility in the rain varies with the depth from the camera, where objects faraway are visually blocked more by the fog than by the rain streaks. However, existing datasets and methods for rain removal ignore these physical properties, thus limiting the rain removal efficiency on real photos. In this work, we analyze the visual effects of rain subject to scene depth and formulate a rain imaging model that collectively considers rain streaks and fog. Also, we prepare a dataset called RainCityscapes on real outdoor photos. Furthermore, we design a novel real-time end-to-end deep neural network, for which we train to learn the depth-guided non-local features and to regress a residual map to produce a rain-free output image. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to show its superiority over others. Xiaowei Hu 0001, Lei Zhu 0003, Tianyu Wang 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Image Process. | 5 |
| 2021 | Self-Ensembling Co-Training Framework for Semi-Supervised COVID-19 CT SegmentationabstractThe coronavirus disease 2019 (COVID-19) has become a severe worldwide health emergency and is spreading at a rapid rate. Segmentation of COVID lesions from computed tomography (CT) scans is of great importance for supervising disease progression and further clinical treatment. As labeling COVID-19 CT scans is labor-intensive and time-consuming, it is essential to develop a segmentation method based on limited labeled data to conduct this task. In this paper, we propose a self-ensembled co-training framework, which is trained by limited labeled data and large-scale unlabeled data, to automatically extract COVID lesions from CT scans. Specifically, to enrich the diversity of unsupervised information, we build a co-training framework consisting of two collaborative models, in which the two models teach each other during training by using their respective predicted pseudo-labels of unlabeled data. Moreover, to alleviate the adverse impacts of noisy pseudo-labels for each model, we propose a self-ensembling strategy to perform consistency regularization for the up-to-date predictions of unlabeled data, in which the predictions of unlabeled data are gradually ensembled via moving average at the end of every training epoch. We evaluate our framework on a COVID-19 dataset containing 103 CT scans. Experimental results show that our proposed method achieves better performance in the case of only 4 labeled CT scans compared to the state-of-the-art semi-supervised segmentation networks. Caizi Li, Qi Dou 0001, Fan Lin, Kebao Zhang, Zuxin Feng, Weixin Si, Xuesong Deng, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 10 |
| 2021 | SALMNet: A Structure-Aware Lane Marking Detection NetworkabstractLane marking detection is a fundamental task, which serves as an important prerequisite for automatic driving or driver-assistance systems. However, the complex and uncontrollable driving road environment as well as the discontinuous lane marking appearance make this task challenging. In this work, a novel deep neural network architecture is presented to detect lane markings in a complex environment by analyzing their structure information. There are two contributions to the network design. Firstly, a semantic-guided channel attention (SGCA) module is developed to select the low-level features of a deep convolutional neural network by taking the high-level features as the guidance. Secondly, a pyramid deformable convolution (PDC) module is formulated to enlarge the receptive fields and to capture the complex structures of lane markings by applying deformable convolutions on multiple feature maps with different scales. Hence, our network can better reduce false detection and enhance lane marking structures simultaneously. The experimental results on three benchmark datasets for lane marking detection show that our method outperforms other methods on all the benchmark datasets. Xuemiao Xu, Tianfei Yu, Xiaowei Hu 0001, Wing W. Y. Ng, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Temporal Memory Relation Network for Workflow Recognition From Surgical VideoabstractAutomatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset). Yueming Jin, Yonghao Long 0001, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Rotation-Oriented Collaborative Self-Supervised Learning for Retinal Disease DiagnosisabstractThe automatic diagnosis of various conventional ophthalmic diseases from fundus images is important in clinical practice. However, developing such automatic solutions is challenging due to the requirement of a large amount of training data and the expensive annotations for medical images. This paper presents a novel self-supervised learning framework for retinal disease diagnosis to reduce the annotation efforts by learning the visual features from the unlabeled images. To achieve this, we present a rotation-oriented collaborative method that explores rotation-related and rotation-invariant features, which capture discriminative structures from fundus images and also explore the invariant property used for retinal disease classification. We evaluate the proposed method on two public benchmark datasets for retinal disease classification. The experimental results demonstrate that our method outperforms other self-supervised feature learning methods (around 4.2% area under the curve (AUC)). With a large amount of unlabeled data available, our method can surpass the supervised baseline for pathologic myopia (PM) and is very close to the supervised baseline for age-related macular degeneration (AMD), showing the potential benefit of our method in clinical practice. Xiaomeng Li 0001, Xiaowei Hu 0001, Xiaojuan Qi 0001, Lequan Yu, Wei Zhao 0029, Pheng-Ann Heng, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Dual-Teacher++: Exploiting Intra-Domain and Inter-Domain Knowledge With Reliable Transfer for Cardiac SegmentationabstractAnnotation scarcity is a long-standing problem in medical image analysis area. To efficiently leverage limited annotations, abundant unlabeled data are additionally exploited in semi-supervised learning, while well-established cross-modality data are investigated in domain adaptation. In this paper, we aim to explore the feasibility of concurrently leveraging both unlabeled data and cross-modality data for annotation-efficient cardiac segmentation. To this end, we propose a cutting-edge semi-supervised domain adaptation framework, namely Dual-Teacher++. Besides directly learning from limited labeled target domain data (e.g., CT) via a student model adopted by previous literature, we design novel dual teacher models, including an inter-domain teacher model to explore cross-modality priors from source domain (e.g., MR) and an intra-domain teacher model to investigate the knowledge beneath unlabeled target domain. In this way, the dual teacher models would transfer acquired inter- and intra-domain knowledge to the student model for further integration and exploitation. Moreover, to encourag reliable dual-domain knowledge transfer, we enhance the inter-domain knowledge transfer on the samples with higher similarity to target domain after appearance alignment, and also strengthen intra-domain knowledge transfer of unlabeled target data with higher prediction confidence. In this way, the student model can obtain reliable dual-domain knowledge and yield improved performance on target domain data. We extensively evaluated the feasibility of our method on the MM-WHS 2017 challenge dataset. The experiments have demonstrated the superiority of our framework over other semi-supervised learning and domain adaptation methods. Moreover, our performance gains could be yielded in bidirections, i.e., adapting from MR to CT, and from CT to MR. Our code will be available at https://github.com/kli-lalala/Dual-Teacher-. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Transformation-Consistent Self-Ensembling Model for Semisupervised Medical Image SegmentationabstractA common shortfall of supervised deep learning for medical imaging is the lack of labeled data, which is often expensive and time consuming to collect. This article presents a new semisupervised method for medical image segmentation, where the network is optimized by a weighted combination of a common supervised loss only for the labeled inputs and a regularization loss for both the labeled and unlabeled data. To utilize the unlabeled data, our method encourages consistent predictions of the network-in-training for the same input under different perturbations. With the semisupervised segmentation tasks, we introduce a transformation-consistent strategy in the self-ensembling model to enhance the regularization effect for pixel-level predictions. To further improve the regularization effects, we extend the transformation in a more generalized form including scaling and optimize the consistency loss with a teacher model, which is an averaging of the student model weights. We extensively validated the proposed semisupervised method on three typical yet challenging medical image segmentation tasks: 1) skin lesion segmentation from dermoscopy images in the International Skin Imaging Collaboration (ISIC) 2017 data set; 2) optic disk (OD) segmentation from fundus images in the Retinal Fundus Glaucoma Challenge (REFUGE) data set; and 3) liver segmentation from volumetric CT scans in the Liver Tumor Segmentation Challenge (LiTS) data set. Compared with state-of-the-art, our method shows superior performance on the challenging 2-D/3-D medical images, demonstrating the effectiveness of our semisupervised method for medical image segmentation. Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | DNF-Net: A Deep Normal Filtering Network for Mesh DenoisingabstractThis article presents a deep normal filtering network, called DNF-Net, for mesh denoising. To better capture local geometry, our network processes the mesh in terms of local patches extracted from the mesh. Overall, DNF-Net is an end-to-end network that takes patches of facet normals as inputs and directly outputs the corresponding denoised facet normals of the patches. In this way, we can reconstruct the geometry from the denoised normals with feature preservation. Besides the overall network architecture, our contributions include a novel multi-scale feature embedding unit, a residual learning strategy to remove noise, and a deeply-supervised joint loss function. Compared with the recent data-driven works on mesh denoising, DNF-Net does not require manual input to extract features and better utilizes the training data to enhance its denoising performance. Finally, we present comprehensive experiments to evaluate our method and demonstrate its superiority over the state of the art on both synthetic and real-scanned meshes. Xianzhi Li 0001, Ruihui Li, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Towards Cross-Modality Medical Image Segmentation with Online Mutual Knowledge DistillationabstractThe success of deep convolutional neural networks is partially attributed to the massive amount of annotated training data. However, in practice, medical data annotations are usually expensive and time-consuming to be obtained. Considering multi-modality data with the same anatomic structures are widely available in clinic routine, in this paper, we aim to exploit the prior knowledge (e.g., shape priors) learned from one modality (aka., assistant modality) to improve the segmentation performance on another modality (aka., target modality) to make up annotation scarcity. To alleviate the learning difficulties caused by modality-specific appearance discrepancy, we first present an Image Alignment Module (IAM) to narrow the appearance gap between assistant and target modality data. We then propose a novel Mutual Knowledge Distillation (MKD) scheme to thoroughly exploit the modality-shared knowledge to facilitate the target-modality segmentation. To be specific, we formulate our framework as an integration of two individual segmentors. Each segmentor not only explicitly extracts one modality knowledge from corresponding annotations, but also implicitly explores another modality knowledge from its counterpart in mutual-guided manner. The ensemble of two segmentors would further integrate the knowledge from both modalities and generate reliable segmentation results on target modality. Experimental results on the public multi-class cardiac segmentation data, i.e., MM-WHS 2017, show that our method achieves large improvements on CT segmentation by utilizing additional MRI data and outperforms other state-of-the-art multi-modality learning methods. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
AAAI | 4 |
| 2020 | Virtually-Extended Proprioception: Providing Spatial Reference in VR through an Appended Virtual LimbabstractSelecting targets directly in the virtual world is difficult due to the lack of haptic feedback and inaccurate estimation of egocentric distances. Proprioception, the sense of self-movement and body position, can be utilized to improve virtual target selection by placing targets on or around one's body. However, its effective scope is limited closely around one's body. We explore the concept of virtually-extended proprioception by appending virtual body parts mimicking real body parts to users' avatars, to provide spatial reference to virtual targets. Our studies suggest that our approach facilitates more efficient target selection in VR as compared to no reference or using an everyday object as reference. Besides, by cultivating users' sense of ownership on the appended virtual body part, we can further enhance target selection performance. The effects of transparency and granularity of the virtual body part on target selection performance are also discussed. Yang Tian 0008, Yuming Bai, Shengdong Zhao 0001, Chi-Wing Fu, Tianpei Yang, Pheng-Ann Heng |
CHI | 6 |
| 2020 | Personalized Re-ranking with Item Relationships for E-commerceabstractRe-ranking is a critical task for large-scale commercial recommender systems. Given the initial ranked lists, top candidates are re-ranked to improve the accuracy of the ranking results. However, existing re-ranking strategies are sub-optimal due to (i) most prior works do not consider explicit item relationships, like being substitutable or complementary, which may mutually influence the user satisfaction on other items in the lists, and (ii) they usually apply an identical re-ranking strategy for all users, with personalized user preferences and intents ignored. To resolve the problem, we construct a heterogeneous graph to fuse the initial scoring information and item relationships information. We develop a graph neural network based framework, IRGPR, to explicitly model transitive item relationships by recursively aggregating relational information from multi-hop neighborhoods. We also incorporate a novel intent embedding network to embed personalized user intents into the propagation. We conduct extensive experiments on real-world datasets, demonstrating the effectiveness of IRGPR in re-ranking. Further analysis reveals that modeling the item relationships and personalized intents are particularly useful for improving the performance of re-ranking. Weiwen Liu, Qing Liu 0020, Ruiming Tang, Junyang Chen 0001, Xiuqiang He 0001, Pheng-Ann Heng |
CIKM | 6 |
| 2020 | A Multi-Task Mean Teacher for Semi-Supervised Shadow DetectionabstractExisting shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow detection by leveraging unlabeled data and exploring the learning of multiple information of shadows simultaneously. To be specific, we first build a multi-task baseline model to simultaneously detect shadow regions, shadow edges, and shadow count by leveraging their complementary information and assign this baseline model to the student and teacher network. After that, we encourage the predictions of the three tasks from the student and teacher networks to be consistent for computing a consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from the predictions of the multi-task baseline model. Experimental results on three widely-used benchmark datasets show that our method consistently outperforms all the compared state-of- the-art methods, which verifies that the proposed network can effectively leverage additional unlabeled data to boost the shadow detection performance. Zhihao Chen 0004, Lei Zhu 0003, Song Wang 0002, Wei Feng 0005, Pheng-Ann Heng |
CVPR | 6 |
| 2020 | PointAugment: An Auto-Augmentation Framework for Point Cloud ClassificationabstractWe present PointAugment, a new auto-augmentation framework that automatically optimizes and augments point cloud samples to enrich the data diversity when we train a classification network. Different from existing auto-augmentation methods for 2D images, PointAugment is sample-aware and takes an adversarial learning strategy to jointly optimize an augmentor network and a classifier network, such that the augmentor can learn to produce augmented samples that best fit the classifier. Moreover, we formulate a learnable point augmentation function with a shape-wise transformation and a point-wise displacement, and carefully design loss functions to adopt the augmented samples based on the learning progress of the classifier. Extensive experiments also confirm PointAugment's effectiveness and robustness to improve the performance of various networks on shape classification and retrival. Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 3 |
| 2020 | Instance Shadow DetectionabstractInstance shadow detection is a brand new problem, aiming to find shadow instances paired with object instances. To approach it, we first prepare a new dataset called SOBA, named after Shadow-OBject Association, with 3,623 pairs of shadow and object instances in 1,000 photos, each with individual labeled masks. Second, we design LISA, named after Light-guided Instance Shadow-object Association, an end-to-end framework to automatically predict the shadow and object instances, together with the shadow-object associations and light direction. Then, we pair up the predicted shadow and object instances, and match them with the predicted shadow-object associations to generate the final results. In our evaluations, we formulate a new metric named the shadow-object average precision to measure the performance of our results. Further, we conducted various experiments and demonstrate our method's applicability on light direction estimation and photo editing. Tianyu Wang 0003, Xiaowei Hu 0001, Qiong Wang 0001, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 4 |
| 2020 | Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization
Lequan Yu, Caizi Li, Chi-Wing Fu, Pheng-Ann Heng |
ECCV (9) | 5 |
| 2020 | Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree SearchabstractAutomatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame individually and produce the outcomes without effective consideration on future information. In this paper, we propose a framework based on reinforcement learning and tree search for joint surgical gesture segmentation and classification. An agent is trained to segment and classify the surgical video in a human-like manner whose direct decisions are re-considered by tree search appropriately. Our proposed tree search algorithm unites the outputs from two designed neural networks, i.e., policy and value network. With the integration of complementary information from distinct models, our framework is able to achieve the better performance than baseline methods using either of the neural networks. For an overall evaluation, our developed approach consistently outperforms the existing methods on the suturing task of JIGSAWS dataset in terms of accuracy, edit score and F1 score. Our study highlights the utilization of tree search to refine actions in reinforcement learning framework for surgical robotic applications. Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
ICRA | 4 |
| 2020 | A Learning-Driven Framework with Spatial Optimization For Surgical Suture Thread Reconstruction and Autonomous Grasping Under Multiple Topologies and Environmental NoisesabstractSurgical knot tying is one of the most fundamental and important procedures in surgery, and a high-quality knot can significantly benefit the postoperative recovery of the patient. However, a longtime operation may easily cause fatigue to surgeons, especially during the tedious wound closure task. In this paper, we present a vision-based method to automate the suture thread grasping, which is a sub-task in surgical knot tying and an intermediate step between the stitching and looping manipulations. To achieve this goal, the acquisition of a suture's three-dimensional (3D) information is critical. Towards this objective, we adopt a transfer-learning strategy first to fine-tune a pre-trained model by learning the information from large legacy surgical data and images obtained by the onsite equipment. Thus, a robust suture segmentation can be achieved regardless of inherent environment noises. We further leverage a searching strategy with termination policies for a suture's sequence inference based on the analysis of multiple topologies. Exact results of the pixel-level sequence along a suture can be obtained, and they can be further applied for a 3D shape reconstruction using our optimized shortest path approach. The grasping point considering the suturing criterion can be ultimately acquired. Experiments regarding the suture 2D segmentation and ordering sequence inference under environmental noises were extensively evaluated. Results related to the automated grasping operation were demonstrated by simulations in V-REP and by robot experiments using Universal Robot (UR) together with the da Vinci Research Kit (dVRK) adopting our learning-driven framework. Bo Lu 0001, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Henry K. Chu, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IROS | 7 |
| 2020 | Dual-Teacher: Integrating Intra-domain and Inter-domain Teachers for Annotation-Efficient Cardiac Segmentation
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
MICCAI (1) | 4 |
| 2020 | Difficulty-Aware Meta-learning for Rare Disease Diagnosis
Xiaomeng Li 0001, Lequan Yu, Yueming Jin, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
MICCAI (1) | 6 |
| 2020 | Shape-Aware Meta-learning for Generalizing Prostate MRI Segmentation to Unseen Domains
Quande Liu, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (2) | 3 |
| 2020 | Cascaded Robust Learning at Imperfect Labels for Chest X-ray Segmentation
Cheng Xue 0003, Xiaomeng Li 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (6) | 5 |
| 2020 | Learning Motion Flows for Semi-supervised Instrument Segmentation from Robotic Surgical Video
Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 5 |
| 2020 | Deep Semi-supervised Knowledge Distillation for Overlapping Cervical Cell Instance Segmentation
Yanning Zhou 0001, Hao Chen 0011, Huangjing Lin, Pheng-Ann Heng |
MICCAI (1) | 4 |
| 2020 | A Second-Order Subregion Pooling Network for Breast Lesion Segmentation in Ultrasound
Lei Zhu 0003, Rongzhen Chen, Huazhu Fu, Liansheng Wang 0002, Pheng-Ann Heng |
MICCAI (6) | 7 |
| 2020 | Balancing Between Accuracy and Fairness for Interactive Recommendation with Reinforcement Learning
Weiwen Liu, Feng Liu 0034, Ruiming Tang, Ben Liao, Guangyong Chen, Pheng-Ann Heng |
PAKDD (1) | 6 |
| 2020 | NormalF-Net: Normal Filtering Neural Network for Feature-preserving Mesh Denoising
Zhiqi Li 0002, Yingkui Zhang, Yidan Feng, Xingyu Xie, Qiong Wang 0001, Mingqiang Wei, Pheng-Ann Heng |
Comput. Aided Des. | 7 |
| 2020 | Revisiting metric learning for few-shot image classification
Xiaomeng Li 0001, Lequan Yu, Chi-Wing Fu, Pheng-Ann Heng |
Neurocomputing | 5 |
| 2020 | Multi-task recurrent convolutional network with correlation loss for surgical video analysis
Yueming Jin, Huaxia Li, Qi Dou 0001, Hao Chen 0011, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
Medical Image Anal. | 7 |
| 2020 | REFUGE Challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographsabstractGlaucoma is one of the leading causes of irreversible but preventable blindness in working age populations. Color fundus photography (CFP) is the most cost-effective imaging modality to screen for retinal disorders. However, its application to glaucoma has been limited to the computation of a few related biomarkers such as the vertical cup-to-disc ratio. Deep learning approaches, although widely applied for medical image analysis, have not been extensively used for glaucoma assessment due to the limited size of the available data sets. Furthermore, the lack of a standardize benchmark strategy makes difficult to compare existing methods in a uniform way. In order to overcome these issues we set up the Retinal Fundus Glaucoma Challenge, REFUGE (https://refuge.grand-challenge.org), held in conjunction with MICCAI 2018. The challenge consisted of two primary tasks, namely optic disc/cup segmentation and glaucoma classification. As part of REFUGE, we have publicly released a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one. We have also built an evaluation framework to ease and ensure fairness in the comparison of different models, encouraging the development of novel techniques in the field. 12 teams qualified and participated in the online challenge. This paper summarizes their methods and analyzes their corresponding results. In particular, we observed that two of the top-ranked teams outperformed two human experts in the glaucoma classification task. Furthermore, the segmentation results were in general consistent with the ground truth annotations, with complementary outcomes that can be further exploited by ensembling the results. José Ignacio Orlando, Huazhu Fu, João Barbosa Breda, Karel van Keer, Deepti R. Bathula, Andres Diaz-Pinto, Ruogu Fang, Pheng-Ann Heng, Jeyoung Kim, Joonseok Lee, Peng Liu 0049, Shuai Lu 0003, Balamurali Murugesan, Valery Naranjo, Sai Samarth R. Phaye, Sharath M. Shankaranarayana, Hrvoje Bogunovic |
Medical Image Anal. | 8 |
| 2020 | Towards multi-center glaucoma OCT image screening with semi-supervised joint structure and function multi-task learning
Xi Wang 0013, Hao Chen 0011, An-ran Ran, Luyang Luo, Poemen P. Chan, Clement C. Tham, Robert T. Chang, Suria S. Mannil, Carol Y. Cheung, Pheng-Ann Heng |
Medical Image Anal. | 10 |
| 2020 | Direction-Aware Spatial Context Features for Shadow Detection and RemovalabstractShadow detection and shadow removal are fundamental and challenging tasks, requiring an understanding of the global image semantics. This paper presents a novel deep neural network design for shadow detection and removal by analyzing the spatial image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting and removing shadows. This design is developed into the DSC module and embedded in a convolutional neural network (CNN) to learn the DSC features at different levels. Moreover, we design a weighted cross entropy loss to make effective the training for shadow detection and further adopt the network for shadow removal by using a euclidean loss function and formulating a color transfer function to address the color and luminosity inconsistencies in the training pairs. We employed two shadow detection benchmark datasets and two shadow removal benchmark datasets, and performed various experiments to evaluate our method. Experimental results show that our method performs favorably against the state-of-the-art methods for both shadow detection and shadow removal. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Aggregating Attentional Dilated Features for Salient Object DetectionabstractThis paper presents a novel deep learning model to aggregate the attentional dilated features for salient object detection by exploring the complementary information between the global and local context in a convolutional neural network. There are two technical contributions to our network design. First, we develop an attentional dense atrous (dilated) spatial pyramid pooling (AD-ASPP) module to selectively use the local saliency cues captured by dilated convolutions with a small rate and the global saliency cues captured by dilated convolutions with a large rate. Second, taking the feature pyramid network as the backbone, we develop an aggregation network to integrate the refined features by formulating two consecutive chains of residual learning based modules: one chain from deep to shallow layers while another chain from shallow to deep layers. We evaluate our network on seven widely-used saliency detection benchmarks by comparing it against 21 state-of-the-art methods. Experimental results show that our network outperforms others on all the seven benchmark datasets. Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | Weakly Supervised Deep Learning for Whole Slide Lung Cancer Image AnalysisabstractHistopathology image analysis serves as the gold standard for cancer diagnosis. Efficient and precise diagnosis is quite critical for the subsequent therapeutic treatment of patients. So far, computer-aided diagnosis has not been widely applied in pathological field yet as currently well-addressed tasks are only the tip of the iceberg. Whole slide image (WSI) classification is a quite challenging problem. First, the scarcity of annotations heavily impedes the pace of developing effective approaches. Pixelwise delineated annotations on WSIs are time consuming and tedious, which poses difficulties in building a large-scale training dataset. In addition, a variety of heterogeneous patterns of tumor existing in high magnification field are actually the major obstacle. Furthermore, a gigapixel scale WSI cannot be directly analyzed due to the immeasurable computational cost. How to design the weakly supervised learning methods to maximize the use of available WSI-level labels that can be readily obtained in clinical practice is quite appealing. To overcome these challenges, we present a weakly supervised approach in this article for fast and effective classification on the whole slide lung cancer images. Our method first takes advantage of a patch-based fully convolutional network (FCN) to retrieve discriminative blocks and provides representative deep features with high efficiency. Then, different context-aware block selection and feature aggregation strategies are explored to generate globally holistic WSI descriptor which is ultimately fed into a random forest (RF) classifier for the image-level prediction. To the best of our knowledge, this is the first study to exploit the potential of image-level labels along with some coarse annotations for weakly supervised learning. A large-scale lung cancer WSI dataset is constructed in this article for evaluation, which validates the effectiveness and feasibility of the proposed method. Extensive experiments demonstrate the superior performance of our method that surpasses the state-of-the-art approaches by a significant margin with an accuracy of 97.3%. In addition, our method also achieves the best performance on the public lung cancer WSIs dataset from The Cancer Genome Atlas (TCGA). We highlight that a small number of coarse annotations can contribute to further accuracy improvement. We believe that weakly supervised learning methods have great potential to assist pathologists in histology image diagnosis in the near future. Xi Wang 0013, Hao Chen 0011, Caixia Gan, Huangjing Lin, Qi Dou 0001, Efstratios Tsougenis, Qitao Huang, Muyan Cai, Pheng-Ann Heng |
IEEE Trans. Cybern. | 9 |
| 2020 | UD-MIL: Uncertainty-Driven Deep Multiple Instance Learning for OCT Image ClassificationabstractDeep learning has achieved remarkable success in the optical coherence tomography (OCT) image classification task with substantial labelled B-scan images available. However, obtaining such fine-grained expert annotations is usually quite difficult and expensive. How to leverage the volume-level labels to develop a robust classifier is very appealing. In this paper, we propose a weakly supervised deep learning framework with uncertainty estimation to address the macula-related disease classification problem from OCT images with the only volume-level label being available. First, a convolutional neural network (CNN) based instance-level classifier is iteratively refined by using the proposed uncertainty-driven deep multiple instance learning scheme. To our best knowledge, we are the first to incorporate the uncertainty evaluation mechanism into multiple instance learning (MIL) for training a robust instance classifier. The classifier is able to detect suspicious abnormal instances and abstract the corresponding deep embedding with high representation capability simultaneously. Second, a recurrent neural network (RNN) takes instance features from the same bag as input and generates the final bag-level prediction by considering the individually local instance information and globally aggregated bag-level representation. For more comprehensive validation, we built two large diabetic macular edema (DME) OCT datasets from different devices and imaging protocols to evaluate the efficacy of our method, which are composed of 30,151 B-scans in 1,396 volumes from 274 patients (Heidelberg-DME dataset) and 38,976 B-scans in 3,248 volumes from 490 patients (Triton-DME dataset), respectively. We compare the proposed method with the state-of-the-art approaches, and experimentally demonstrate that our method is superior to alternative methods, achieving volume-level accuracy, F1-score and area under the receiver operating characteristic curve (AUC) of 95.1%, 0.939 and 0.990 on Heidelberg-DME and those of 95.1%, 0.935 and 0.986 on Triton-DME, respectively. Furthermore, the proposed method also yields competitive results on another public age-related macular degeneration OCT dataset, indicating the high potential as an effective screening tool in the clinical practice. Xi Wang 0013, Fangyao Tang, Hao Chen 0011, Luyang Luo, Ziqi Tang, An-ran Ran, Carol Y. Cheung, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 8 |
| 2020 | Unsupervised Bidirectional Cross-Modality Adaptation via Deeply Synergistic Image and Feature Alignment for Medical Image SegmentationabstractUnsupervised domain adaptation has increasingly gained interest in medical image computing, aiming to tackle the performance degradation of deep neural networks when being deployed to unseen data with heterogeneous characteristics. In this work, we present a novel unsupervised domain adaptation framework, named as Synergistic Image and Feature Alignment (SIFA), to effectively adapt a segmentation network to an unlabeled target domain. Our proposed SIFA conducts synergistic alignment of domains from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features by leveraging adversarial learning in multiple aspects and with a deeply supervised mechanism. The feature encoder is shared between both adaptive perspectives to leverage their mutual benefits via end-to-end learning. We have extensively evaluated our method with cardiac substructure segmentation and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our SIFA method is effective in improving segmentation performance on unlabeled target images, and outperforms the state-of-the-art domain adaptation approaches by a large margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Unpaired Multi-Modal Segmentation via Knowledge DistillationabstractMulti-modal learning is typically performed with network architectures containing modality-specific layers and shared layers, utilizing co-registered images of different modalities. We propose a novel learning scheme for unpaired cross-modality image segmentation, with a highly compact architecture achieving superior segmentation accuracy. In our method, we heavily reuse network parameters, by sharing all convolutional kernels across CT and MRI, and only employ modality-specific internal normalization layers which compute respective statistics. To effectively train such a highly compact model, we introduce a novel loss term inspired by knowledge distillation, by explicitly constraining the KL-divergence of our derived prediction distributions between modalities. We have extensively validated our approach on two multi-class segmentation problems: i) cardiac structure segmentation, and ii) abdominal organ segmentation. Different network settings, i.e., 2D dilated network and 3D U-net, are utilized to investigate our method's general efficacy. Experimental results on both tasks demonstrate that our novel multi-modal learning scheme consistently outperforms single-modal training and previous multi-modal approaches. Qi Dou 0001, Quande Liu, Pheng-Ann Heng, Ben Glocker |
IEEE Trans. Medical Imaging | 3 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 8 |
| 2020 | CANet: Cross-Disease Attention Network for Joint Diabetic Retinopathy and Diabetic Macular Edema GradingabstractDiabetic retinopathy (DR) and diabetic macular edema (DME) are the leading causes of permanent blindness in the working-age population. Automatic grading of DR and DME helps ophthalmologists design tailored treatments to patients, thus is of vital importance in the clinical practice. However, prior works either grade DR or DME, and ignore the correlation between DR and its complication, i.e., DME. Moreover, the location information, e.g., macula and soft hard exhaust annotations, are widely used as a prior for grading. Such annotations are costly to obtain, hence it is desirable to develop automatic grading methods with only image-level supervision. In this article, we present a novel cross-disease attention network (CANet) to jointly grade DR and DME by exploring the internal relationship between the diseases with only image-level supervision. Our key contributions include the disease-specific attention module to selectively learn useful features for individual diseases, and the disease-dependent attention module to further capture the internal relationship between the two diseases. We integrate these two attention modules in a deep network to produce disease-specific and disease-dependent features, and to maximize the overall performance jointly for grading DR and DME. We evaluate our network on two public benchmark datasets, i.e., ISBI 2018 IDRiD challenge dataset and Messidor dataset. Our method achieves the best result on the ISBI 2018 IDRiD challenge dataset and outperforms other methods on the Messidor dataset. Our code is publicly available at https://github.com/xmengli999/CANet. Xiaomeng Li 0001, Xiaowei Hu 0001, Lequan Yu, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Multi-Task Deep Model With Margin Ranking Loss for Lung Nodule AnalysisabstractLung cancer is the leading cause of cancer deaths worldwide and early diagnosis of lung nodule is of great importance for therapeutic treatment and saving lives. Automated lung nodule analysis requires both accurate lung nodule benign-malignant classification and attribute score regression. However, this is quite challenging due to the considerable difficulty of lung nodule heterogeneity modeling and the limited discrimination capability on ambiguous cases. To solve these challenges, we propose a Multi-Task deep model with Margin Ranking loss (referred as MTMR-Net) for automated lung nodule analysis. Compared to existing methods which consider these two tasks separately, the relatedness between lung nodule classification and attribute score regression is explicitly explored in a cause-and-effect manner within our multi-task deep model, which can contribute to the performance gains of both tasks. The results of different tasks can be yielded simultaneously for assisting the radiologists in diagnosis interpretation. Furthermore, a Siamese network with a margin ranking loss is elaborately designed to enhance the discrimination capability on ambiguous nodule cases. To further explore the internal relationship between two tasks and validate the effectiveness of the proposed model, we use the recursive feature elimination method to iteratively rank the most malignancy-related features. We validate the efficacy of our method MTMR-Net on the public benchmark LIDC-IDRI dataset. Extensive experiments show that the diagnosis results with internal relationship explicitly explored in our model has met some similar patterns in clinical usage and also demonstrate that our approach can achieve competitive classification performance and more accurate scoring on attributes over the state-of-the-arts. Codes are publicly available at: https://github.com/CaptainWilliam/MTMR-NET. Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2020 | MS-Net: Multi-Site Network for Improving Prostate Segmentation With Heterogeneous MRI DataabstractAutomated prostate segmentation in MRI is highly demanded for computer-assisted diagnosis. Recently, a variety of deep learning methods have achieved remarkable progress in this task, usually relying on large amounts of training data. Due to the nature of scarcity for medical images, it is important to effectively aggregate data from multiple sites for robust model training, to alleviate the insufficiency of single-site samples. However, the prostate MRIs from different sites present heterogeneity due to the differences in scanners and imaging protocols, raising challenges for effective ways of aggregating multi-site data for network training. In this paper, we propose a novel multi-site network (MS-Net) for improving prostate segmentation by learning robust representations, leveraging multiple sources of data. To compensate for the inter-site heterogeneity of different MRI datasets, we develop Domain-Specific Batch Normalization layers in the network backbone, enabling the network to estimate statistics and perform feature normalization for each site separately. Considering the difficulty of capturing the shared knowledge from multiple datasets, a novel learning paradigm, i.e., Multi-site-guided Knowledge Transfer, is proposed to enhance the kernels to extract more generic representations from multi-site data. Extensive experiments on three heterogeneous prostate MRI datasets demonstrate that our MS-Net improves the performance across all datasets consistently, and outperforms state-of-the-art methods for multi-site learning. Quande Liu, Qi Dou 0001, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2020 | ψ-Net: Stacking Densely Convolutional LSTMs for Sub-Cortical Brain Structure SegmentationabstractSub-cortical brain structure segmentation is of great importance for diagnosing neuropsychiatric disorders. However, developing an automatic approach to segmenting sub-cortical brain structures remains very challenging due to the ambiguous boundaries, complex anatomical structures, and large variance of shapes. This paper presents a novel deep network architecture, namely Ψ -Net, for sub-cortical brain structure segmentation, aiming at selectively aggregating features and boosting the information propagation in a deep convolutional neural network (CNN). To achieve this, we first formulate a densely convolutional LSTM module (DC-LSTM) to selectively aggregate the convolutional features with the same spatial resolution at the same stage of a CNN. This helps to promote the discriminativeness of features at each CNN stage. Second, we stack multiple DC-LSTMs from the deepest stage to the shallowest stage to progressively enrich low-level feature maps with high-level context. We employ two benchmark datasets on sub-cortical brain structure segmentation, and perform various experiments to evaluate the proposed Ψ -Net. The experimental results show that our network performs favorably against the state-of-the-art methods on both benchmark datasets. Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Semi-Supervised Medical Image Classification With Relation-Driven Self-Ensembling ModelabstractTraining deep neural networks usually requires a large amount of labeled data to obtain good performance. However, in medical image analysis, obtaining high-quality labels for the data is laborious and expensive, as accurately annotating medical images demands expertise knowledge of the clinicians. In this paper, we present a novel relation-driven semi-supervised framework for medical image classification. It is a consistency-based method which exploits the unlabeled data by encouraging the prediction consistency of given input under perturbations, and leverages a self-ensembling model to produce high-quality consistency targets for the unlabeled data. Considering that human diagnosis often refers to previous analogous cases to make reliable decisions, we introduce a novel sample relation consistency (SRC) paradigm to effectively exploit unlabeled data by modeling the relationship information among different samples. Superior to existing consistency-based methods which simply enforce consistency of individual predictions, our framework explicitly enforces the consistency of semantic relation among different samples under perturbations, encouraging the model to explore extra semantic information from unlabeled data. We have conducted extensive experiments to evaluate our method on two public benchmark medical image classification datasets, i.e., skin lesion diagnosis with ISIC 2018 challenge and thorax disease classification with ChestX-ray14. Our method outperforms many state-of-the-art semi-supervised learning methods on both single-label and multi-label image classification scenarios. Quande Liu, Lequan Yu, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Deep Mining External Imperfect Data for Chest X-Ray Disease ScreeningabstractDeep learning approaches have demonstrated remarkable progress in automatic Chest X-ray analysis. The data-driven feature of deep models requires training data to cover a large distribution. Therefore, it is substantial to integrate knowledge from multiple datasets, especially for medical images. However, learning a disease classification model with extra Chest X-ray (CXR) data is yet challenging. Recent researches have demonstrated that performance bottleneck exists in joint training on different CXR datasets, and few made efforts to address the obstacle. In this paper, we argue that incorporating an external CXR dataset leads to imperfect training data, which raises the challenges. Specifically, the imperfect data is in two folds: domain discrepancy, as the image appearances vary across datasets; and label discrepancy, as different datasets are partially labeled. To this end, we formulate the multi-label thoracic disease classification problem as weighted independent binary tasks according to the categories. For common categories shared across domains, we adopt task-specific adversarial training to alleviate the feature differences. For categories existing in a single dataset, we present uncertainty-aware temporal ensembling of model predictions to mine the information from the missing labels further. In this way, our framework simultaneously models and tackles the domain and label discrepancies, enabling superior knowledge mining ability. We conduct extensive experiments on three datasets with more than 360,000 Chest X-ray images. Our method outperforms other competing models and sets state-of-the-art performance on the official NIH test set with 0.8349 AUC, demonstrating its effectiveness of utilizing the external dataset to improve the internal classification. Luyang Luo, Lequan Yu, Hao Chen 0011, Quande Liu, Xi Wang 0013, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 7 |
| 2020 | DoFE: Domain-Oriented Feature Embedding for Generalizable Fundus Image Segmentation on Unseen DatasetsabstractDeep convolutional neural networks have significantly boosted the performance of fundus image segmentation when test datasets have the same distribution as the training datasets. However, in clinical practice, medical images often exhibit variations in appearance for various reasons, e.g., different scanner vendors and image quality. These distribution discrepancies could lead the deep networks to over-fit on the training datasets and lack generalization ability on the unseen test datasets. To alleviate this issue, we present a novel Domain-oriented Feature Embedding (DoFE) framework to improve the generalization ability of CNNs on unseen target domains by exploring the knowledge from multiple source domains. Our DoFE framework dynamically enriches the image features with additional domain prior knowledge learned from multi-source domains to make the semantic features more discriminative. Specifically, we introduce a Domain Knowledge Pool to learn and memorize the prior information extracted from multi-source domains. Then the original image features are augmented with domain-oriented aggregated features, which are induced from the knowledge pool based on the similarity between the input image and multi-source domain images. We further design a novel domain code prediction branch to infer this similarity and employ an attention-guided mechanism to dynamically combine the aggregated features with the semantic features. We comprehensively evaluate our DoFE framework on two fundus image segmentation tasks, including the optic cup and disc segmentation and vessel segmentation. Our DoFE framework generates satisfying segmentation results on unseen datasets and surpasses other domain generalization and network regularization methods. Lequan Yu, Kang Li 0007, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Unsupervised Detection of Distinctive Regions on 3D ShapesabstractThis article presents a novel approach to learn and detect distinctive regions on 3D shapes. Unlike previous works, which require labeled data, our method is unsupervised. We conduct the analysis on point sets sampled from 3D shapes, then formulate and train a deep neural network for an unsupervised shape clustering task to learn local and global features for distinguishing shapes with respect to a given shape set. To drive the network to learn in an unsupervised manner, we design a clustering-based nonparametric softmax classifier with an iterative re-clustering of shapes, and an adapted contrastive loss for enhancing the feature embedding quality and stabilizing the learning process. By then, we encourage the network to learn the point distinctiveness on the input shapes. We extensively evaluate various aspects of our approach and present its applications for distinctiveness-guided shape retrieval, sampling, and view selection in 3D scenes. Xianzhi Li 0001, Lequan Yu, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ACM Trans. Graph. | 5 |
| 2020 | Saliency-Aware Texture SmoothingabstractTexture smoothing aims to smooth out textures in images, while retaining the prominent structures. This paper presents a saliency-aware approach to the problem with two key contributions. First, we design a deep saliency network with guided non-local blocks (GNLBs) for learning long-range pixel dependencies by taking the predicted saliency map at former layer as the guidance image to help suppress the non-saliency regions in the shallow layer. The GNLB computes the saliency response at a position by a weighted sum of features at all positions, and enables us to produce results that outperform existing deep saliency models. Second, we formulate a joint optimization framework to take saliency information when iteratively separating textures from structures: on the texture layer, we smooth out structures with the help of the saliency information and migrate structures from the texture to structure layer, while on the structure layer, we adopt another deep model to detect edges and simultaneous sparse coding to push textures back to the texture layer. We tested our method on a rich variety of images and compared it with several state-of-the-art methods. Both visual and quantitative comparison results show that our method better preserves structures while removing the texture components. Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image SegmentationabstractThis paper presents a novel unsupervised domain adaptation framework, called Synergistic Image and Feature Adaptation (SIFA), to effectively tackle the problem of domain shift. Domain adaptation has become an important and hot topic in recent studies on deep learning, aiming to recover performance degradation when applying the neural networks to new testing domains. Our proposed SIFA is an elegant learning diagram which presents synergistic fusion of adaptations from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features towards the segmentation task. The feature encoder layers are shared by both perspectives to grasp their mutual benefits during the end-to-end learning procedure. Without using any annotation from the target domain, the learning of our unified model is guided by adversarial losses, with multiple discriminators employed from various aspects. We have extensively validated our method with a challenging application of crossmodality medical image segmentation of cardiac structures. Experimental results demonstrate that our SIFA model recovers the degraded performance from 17.2% to 73.0%, and outperforms the state-of-the-art methods by a significant margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 5 |
| 2019 | Depth-Attentional Features for Single-Image Rain RemovalabstractRain is a common weather phenomenon, where object visibility varies with depth from the camera and objects faraway are visually blocked more by fog than by rain streaks. Existing methods and datasets for rain removal, however, ignore these physical properties, thereby limiting the rain removal efficiency on real photos. In this work, we first analyze the visual effects of rain subject to scene depth and formulate a rain imaging model collectively with rain streaks and fog; by then, we prepare a new dataset called RainCityscapes with rain streaks and fog on real outdoor photos. Furthermore, we design an end-to-end deep neural network, where we train it to learn depth-attentional features via a depth-guided attention mechanism, and regress a residual map to produce the rain-free image output. We performed various experiments to visually and quantitatively compare our method with several state-of-the-art methods to demonstrate its superiority over the others. Xiaowei Hu 0001, Chi-Wing Fu, Lei Zhu 0003, Pheng-Ann Heng |
CVPR | 4 |
| 2019 | Deep Multi-Model Fusion for Single-Image DehazingabstractThis paper presents a deep multi-model fusion network to attentively integrate multiple models to separate layers and boost the performance in single-image dehazing. To do so, we first formulate the attentional feature integration module to maximize the integration of the convolutional neural network (CNN) features at different CNN layers and generate the attentional multi-level integrated features (AMLIF). Then, from the AMLIF, we further predict a haze-free result for an atmospheric scattering model, as well as for four haze-layer separation models, and then fuse the results together to produce the final haze-free image. To evaluate the effectiveness of our method, we compare our network with several state-of-the-art methods on two widely-used dehazing benchmark datasets, as well as on two sets of real-world hazy images. Experimental results demonstrate clear quantitative and qualitative improvements of our method over the state-of-the-arts. Zijun Deng, Lei Zhu 0003, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Qing Zhang 0006, Harry Qin, Pheng-Ann Heng |
ICCV | 8 |
| 2019 | Mask-ShadowGAN: Learning to Remove Shadows From Unpaired DataabstractThis paper presents a new method for shadow removal using unpaired data, enabling us to avoid tedious annotations and obtain more diverse training samples. However, directly employing adversarial learning and cycle-consistency constraints is insufficient to learn the underlying relationship between the shadow and shadow-free domains, since the mapping between shadow and shadow-free images is not simply one-to-one. To address the problem, we formulate Mask-ShadowGAN, a new deep framework that automatically learns to produce a shadow mask from the input shadow image and then takes the mask to guide the shadow generation via re-formulated cycle-consistency constraints. Particularly, the framework simultaneously learns to produce shadow masks and learns to remove shadows, to maximize the overall performance. Also, we prepared an unpaired dataset for shadow removal and demonstrated the effectiveness of Mask-ShadowGAN on various experiments, even it was trained on unpaired data. Xiaowei Hu 0001, Yitong Jiang, Chi-Wing Fu, Pheng-Ann Heng |
ICCV | 4 |
| 2019 | PU-GAN: A Point Cloud Upsampling Adversarial NetworkabstractPoint clouds acquired from range scans are often sparse, noisy, and non-uniform. This paper presents a new point cloud upsampling network called PU-GAN1, which is formulated based on a generative adversarial network (GAN), to learn a rich variety of point distributions from the latent space and upsample points over patches on object surfaces. To realize a working GAN network, we construct an up-down-up expansion unit in the generator for upsampling point features with error feedback and self-correction, and formulate a self-attention unit to enhance the feature integration. Further, we design a compound loss with adversarial, uniform and reconstruction terms, to encourage the discriminator to learn more latent patterns and enhance the output point distribution uniformity. Qualitative and quantitative evaluations demonstrate the quality of our results over the state-of-the-arts in terms of distribution uniformity, proximity-to-surface, and 3D reconstruction quality. Ruihui Li, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ICCV | 5 |
| 2019 | FetusMap: Fetal Pose Estimation in 3D Ultrasound
Xin Yang 0009, Wenlong Shi, Haoran Dou, Jikuan Qian, Yi Wang 0031, Wufeng Xue, Shengli Li 0001, Dong Ni 0001, Pheng-Ann Heng |
MICCAI (5) | 9 |
| 2019 | Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion
Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 6 |
| 2019 | Agent with Warm Start and Active Termination for Plane Localization in 3D Ultrasound
Haoran Dou, Xin Yang 0009, Jikuan Qian, Wufeng Xue, Xu Wang 0017, Lequan Yu, Yi Xiong 0001, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (5) | 10 |
| 2019 | Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video
Yueming Jin, Keyun Cheng, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 4 |
| 2019 | Probabilistic Multilayer Regularization Network for Unsupervised 3D Brain Image Registration
Xiaowei Hu 0001, Lei Zhu 0003, Pheng-Ann Heng |
MICCAI (2) | 4 |
| 2019 | Deep Angular Embedding and Feature Correlation Attention for Breast MRI Cancer Analysis
Luyang Luo, Hao Chen 0011, Xi Wang 0013, Qi Dou 0001, Huangjing Lin, Gongjie Li, Pheng-Ann Heng |
MICCAI (4) | 8 |
| 2019 | Unifying Structure Analysis and Surrogate-Driven Function Regression for Glaucoma OCT Image Screening
Xi Wang 0013, Hao Chen 0011, Luyang Luo, An-ran Ran, Poemen P. Chan, Clement C. Tham, Carol Y. Cheung, Pheng-Ann Heng |
MICCAI (1) | 8 |
| 2019 | Boundary and Entropy-Driven Adversarial Learning for Fundus Image Segmentation
Lequan Yu, Kang Li 0007, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
MICCAI (1) | 6 |
| 2019 | Uncertainty-Aware Self-ensembling Model for Semi-supervised 3D Left Atrium Segmentation
Lequan Yu, Xiaomeng Li 0001, Chi-Wing Fu, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2019 | PFA-ScanNet: Pyramidal Feature Aggregation with Synergistic Learning for Breast Cancer Metastasis Analysis
Huangjing Lin, Hao Chen 0011, Pheng-Ann Heng |
MICCAI (1) | 4 |
| 2019 | IRNet: Instance Relation Network for Overlapping Cervical Cell Segmentation
Yanning Zhou 0001, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (1) | 5 |
| 2019 | Versatile numerical fractures removal for SPH-based free surface liquids
Weixin Si, Xiangyun Liao, Yinling Qian, Qiong Wang 0001, Pheng-Ann Heng |
Comput. Graph. | 5 |
| 2019 | Layered leaf texturing using structure-guided model
Yinling Qian, Hanqiu Sun, Lei Ma 0008, Yanyun Chen, Qiong Wang 0001, Pheng-Ann Heng |
Graph. Model. | 7 |
| 2019 | Mixed reality based respiratory liver tumor puncture navigationabstractThis paper presents a novel mixed reality based navigation system for accurate respiratory liver tumor punctures in radiofrequency ablation (RFA). Our system contains an optical see-through head-mounted display device (OST-HMD), Microsoft HoloLens for perfectly overlaying the virtual information on the patient, and a optical tracking system NDI Polaris for calibrating the surgical utilities in the surgical scene. Compared with traditional navigation method with CT, our system aligns the virtual guidance information and real patient and real-timely updates the view of virtual guidance via a position tracking system. In addition, to alleviate the difficulty during needle placement induced by respiratory motion, we reconstruct the patient-specific respiratory liver motion through statistical motion model to assist doctors precisely puncture liver tumors. The proposed system has been experimentally validated on vivo pigs with an accurate real-time registration approximately 5-mm mean FRE and TRE, which has the potential to be applied in clinical RFA guidance. Ruotong Li, Weixin Si, Xiangyun Liao, Qiong Wang 0001, Reinhard Klein, Pheng-Ann Heng |
Comput. Vis. Media | 6 |
| 2019 | Webthetics: Quantifying webpage aesthetics with deep learning
Qi Dou 0001, Xianjun Sam Zheng, Tongfang Sun, Pheng-Ann Heng |
Int. J. Hum. Comput. Stud. | 4 |
| 2019 | MILD-Net: Minimal information loss dilated network for gland instance segmentation in colon histology images
Simon Graham, Hao Chen 0011, Jevgenij Gamper, Qi Dou 0001, Pheng-Ann Heng, David R. J. Snead, Yee-Wah Tsang, Nasir M. Rajpoot |
Medical Image Anal. | 5 |
| 2019 | RMDL: Recalibrated multi-instance deep learning for whole slide gastric image classification
Yaxi Zhu, Lequan Yu, Hao Chen 0011, Huangjing Lin, Xiangbo Wan, Xinjuan Fan, Pheng-Ann Heng |
Medical Image Anal. | 8 |
| 2019 | Evaluation of algorithms for Multi-Modality Whole Heart Segmentation: An open-access grand challengeabstractKnowledge of whole heart anatomy is a prerequisite for many clinical applications. Whole heart segmentation (WHS), which delineates substructures of the heart, can be very valuable for modeling and analysis of the anatomy and functions of the heart. However, automating this segmentation can be challenging due to the large variation of the heart shape, and different image qualities of the clinical data. To achieve this goal, an initial set of training data is generally needed for constructing priors or for training. Furthermore, it is difficult to perform comparisons between different methods, largely due to differences in the datasets and evaluation metrics used. This manuscript presents the methodologies and evaluation results for the WHS algorithms selected from the submissions to the Multi-Modality Whole Heart Segmentation (MM-WHS) challenge, in conjunction with MICCAI 2017. The challenge provided 120 three-dimensional cardiac images covering the whole heart, including 60 CT and 60 MRI volumes, all acquired in clinical environments with manual delineation. Ten algorithms for CT data and eleven algorithms for MRI data, submitted from twelve groups, have been evaluated. The results showed that the performance of CT WHS was generally better than that of MRI WHS. The segmentation of the substructures for different categories of patients could present different levels of challenge due to the difference in imaging and variations of heart shapes. The deep learning (DL)-based methods demonstrated great potential, though several of them reported poor results in the blinded evaluation. Their performance could vary greatly across different network structures and training strategies. The conventional algorithms, mainly based on multi-atlas segmentation, demonstrated good performance, though the accuracy and computational efficiency could be limited. The challenge, including provision of the annotated training data and the blinded evaluation for submitted algorithms on the test data, continues as an ongoing benchmarking resource via its homepage (www.sdspeople.fudan.edu.cn/zhuangxiahai/0/mmwhs/). Xiahai Zhuang, Lei Li 0020, Christian Payer, Darko Stern, Martin Urschler, Mattias P. Heinrich, Julien Oster, Chunliang Wang, Örjan Smedby, Cheng Bian, Xin Yang 0009, Pheng-Ann Heng, Aliasghar Mortazi, Ulas Bagci, Guanyu Yang 0001, Chenchen Sun, Gaetan Galisot, Jean-Yves Ramel, Guang Yang 0006 |
Medical Image Anal. | 12 |
| 2019 | Online Subspace Learning from Gradient Orientations for Robust Image AlignmentabstractRobust and efficient image alignment remains a challenging task, due to the massiveness of images, great illumination variations between images, partial occlusion, and corruption. To address these challenges, we propose an online image alignment method via subspace learning from image gradient orientations (IGOs). The proposed method integrates the subspace learning, transformed the IGO reconstruction and image alignment into a unified online framework, which is robust for aligning images with severe intensity distortions. Our method is motivated by a principal component analysis (PCA) from gradient orientations that provides more reliable low-dimensional subspace than that from pixel intensities. Instead of processing in the intensity-domain-like conventional methods, we seek alignment in the IGO domain, such that the aligned IGO of the newly arrived image can be decomposed as the sum of a sparse error and a linear composition of the IGO-PCA basis learned from previously well-aligned ones. The optimization problem is tackled by an iterative linearization that minimizes the ℓ1-norm of the sparse error. Furthermore, the IGO-PCA basis is adaptively updated based on incremental thin singular value decomposition, which takes the shift of IGO mean into consideration. The efficacy of the proposed method is validated on the extensive challenging datasets through image alignment, medical atlas construction, and face recognition. The experimental results demonstrate that our algorithm provides more illumination- and occlusion-robust image alignment than the state-of-the-art methods. Qingqing Zheng, Yi Wang 0031, Pheng-Ann Heng |
IEEE Trans. Image Process. | 3 |
| 2019 | SINet: A Scale-Insensitive Convolutional Neural Network for Fast Vehicle DetectionabstractVision-based vehicle detection approaches achieve incredible success in recent years with the development of deep convolutional neural network (CNN). However, existing CNN-based algorithms suffer from the problem that the convolutional features are scale-sensitive in object detection task but it is common that traffic images and videos contain vehicles with a large variance of scales. In this paper, we delve into the source of scale sensitivity, and reveal two key issues: 1) existing RoI pooling destroys the structure of small scale objects and 2) the large intra-class distance for a large variance of scales exceeds the representation capability of a single network. Based on these findings, we present a scale-insensitive convolutional neural network (SINet) for fast detecting vehicles with a large variance of scales. First, we present a context-aware RoI pooling to maintain the contextual information and original structure of small scale objects. Second, we present a multi-branch decision network to minimize the intra-class distance of features. These lightweight techniques bring zero extra time complexity but prominent detection accuracy improvement. The proposed techniques can be equipped with any deep network architectures and keep them trained end-to-end. Our SINet achieves state-of-the-art performance in terms of accuracy and speed (up to 37 FPS) on the KITTI benchmark and a new highway dataset, which contains a large variance of scales and extremely small objects. Xiaowei Hu 0001, Xuemiao Xu, Yongjie Xiao, Hao Chen 0011, Shengfeng He, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2019 | Fast ScanNet: Fast and Dense Analysis of Multi-Gigapixel Whole-Slide Images for Cancer Metastasis DetectionabstractLymph node metastasis is one of the most important indicators in breast cancer diagnosis, that is traditionally observed under the microscope by pathologists. In recent years, with the dramatic advance of high-throughput scanning and deep learning technology, automatic analysis of histology from whole-slide images has received a wealth of interest in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, the automatic detection of lymph node metastases from whole-slide images remains a key challenge because such images are typically very large, where they can often be multiple gigabytes in size. Also, the presence of hard mimics may result in a large number of false positives. In this paper, we propose a novel method with anchor layers for model conversion, which not only leverages the efficiency of fully convolutional architectures to meet the speed requirement in clinical practice but also densely scans the whole-slide image to achieve accurate predictions on both micro- and macro-metastases. Incorporating the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. The efficacy of our method is corroborated on the benchmark dataset of 2016 Camelyon Grand Challenge. Our method achieved significant improvements in comparison with the state-of-the-art methods on tumor localization accuracy with a much faster speed and even surpassed human performance on both challenge tasks. Huangjing Lin, Hao Chen 0011, Simon Graham, Qi Dou 0001, Nasir M. Rajpoot, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2019 | Deep Attentive Features for Prostate Segmentation in 3D Transrectal UltrasoundabstractAutomatic prostate segmentation in transrectal ultrasound (TRUS) images is of essential importance for image-guided prostate interventions and treatment planning. However, developing such automatic solutions remains very challenging due to the missing/ambiguous boundary and inhomogeneous intensity distribution of the prostate in TRUS, as well as the large variability in prostate shapes. This paper develops a novel 3D deep neural network equipped with attention modules for better prostate segmentation in TRUS by fully exploiting the complementary information encoded in different layers of the convolutional neural network (CNN). Our attention module utilizes the attention mechanism to selectively leverage the multi-level features integrated from different layers to refine the features at each individual layer, suppressing the non-prostate noise at shallow layers of the CNN and increasing more prostate details into features at deep layers. Experimental results on challenging 3D TRUS volumes show that our method attains satisfactory segmentation performance. The proposed attention mechanism is a general strategy to aggregate multi-level deep features and has the potential to be used for other medical image segmentation tasks. The code is publicly available at https://github.com/wulalago/DAF3D. Yi Wang 0031, Dong Ni 0001, Haoran Dou, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Harry Qin, Pheng-Ann Heng, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2019 | Patch-Based Output Space Adversarial Learning for Joint Optic Disc and Cup SegmentationabstractGlaucoma is a leading cause of irreversible blindness. Accurate segmentation of the optic disc (OD) and optic cup (OC) from fundus images is beneficial to glaucoma screening and diagnosis. Recently, convolutional neural networks demonstrate promising progress in the joint OD and OC segmentation. However, affected by the domain shift among different datasets, deep networks are severely hindered in generalizing across different scanners and institutions. In this paper, we present a novel patch-based output space adversarial learning framework ( p OSAL) to jointly and robustly segment the OD and OC from different fundus image datasets. We first devise a lightweight and efficient segmentation network as a backbone. Considering the specific morphology of OD and OC, a novel morphology-aware segmentation loss is proposed to guide the network to generate accurate and smooth segmentation. Our p OSAL framework then exploits unsupervised domain adaptation to address the domain shift challenge by encouraging the segmentation in the target domain to be similar to the source ones. Since the whole-segmentation-based adversarial loss is insufficient to drive the network to capture segmentation details, we further design the p OSAL in a patch-based fashion to enable fine-grained discrimination on local segmentation details. We extensively evaluate our p OSAL framework and demonstrate its effectiveness in improving the segmentation performance on three public retinal fundus image datasets, i.e., Drishti-GS, RIM-ONE-r3, and REFUGE. Furthermore, our p OSAL framework achieved the first place in the OD and OC segmentation tasks in the MICCAI 2018 Retinal Fundus Glaucoma Challenge. Lequan Yu, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Towards Automated Semantic Segmentation in Prenatal Volumetric UltrasoundabstractVolumetric ultrasound is rapidly emerging as a viable imaging modality for routine prenatal examinations. Biometrics obtained from the volumetric segmentation shed light on the reformation of precise maternal and fetal health monitoring. However, the poor image quality, low contrast, boundary ambiguity, and complex anatomy shapes conspire toward a great lack of efficient tools for the segmentation. It makes 3-D ultrasound difficult to interpret and hinders the widespread of 3-D ultrasound in obstetrics. In this paper, we are looking at the problem of semantic segmentation in prenatal ultrasound volumes. Our contribution is threefold: 1) we propose the first and fully automatic framework to simultaneously segment multiple anatomical structures with intensive clinical interest, including fetus, gestational sac, and placenta, which remains a rarely studied and arduous challenge; 2) we propose a composite architecture for dense labeling, in which a customized 3-D fully convolutional network explores spatial intensity concurrency for initial labeling, while a multi-directional recurrent neural network (RNN) encodes spatial sequentiality to combat boundary ambiguity for significant refinement; and 3) we introduce a hierarchical deep supervision mechanism to boost the information flow within RNN and fit the latent sequence hierarchy in fine scales, and further improve the segmentation results. Extensively verified on in-house large data sets, our method illustrates a superior segmentation performance, decent agreements with expert measurements and high reproducibilities against scanning variations, and thus is promising in advancing the prenatal ultrasound examinations. Xin Yang 0009, Lequan Yu, Shengli Li 0001, Huaxuan Wen, Dandan Luo, Cheng Bian, Harry Qin, Dong Ni 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 9 |
| 2019 | Bas-Relief Modeling from Normal LayersabstractBas-relief is characterized by its unique presentation of intrinsic shape properties and/or detailed appearance using materials raised up in different degrees above a background. However, many bas-relief modeling methods could not manipulate scene details well. We propose a simple and effective solution for two kinds of bas-relief modeling (i.e., structure-preserving and detail-preserving) which is different from the prior tone mapping alike methods. Our idea originates from an observation on typical 3D models, which are decomposed into a piecewise smooth base layer and a detail layer in normal field. Proper manipulation of the two layers contributes to both structure-preserving and detail-preserving bas-relief modeling. We solve the modeling problem in a discrete geometry processing setup that uses normal-based mesh processing as a theoretical foundation. Specifically, using the two-step mesh smoothing mechanism as a bridge, we transfer the bas-relief modeling problem into a discrete space, and solve it in a least-squares manner. Experiments and comparisons to other methods show that (i) geometry details are better preserved in the scenario with high compression ratios, and (ii) structures are clearly preserved without shape distortion and interference from details. Mingqiang Wei, Yang Tian 0008, Wai-Man Pang, Charlie C. L. Wang, Mingyong Pang, Jun Wang 0039, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2018 | Recurrently Aggregating Deep Features for Salient Object DetectionabstractSalient object detection is a fundamental yet challenging problem in computer vision, aiming to highlight the most visually distinctive objects or regions in an image. Recent works benefit from the development of fully convolutional neural networks (FCNs) and achieve great success by integrating features from multiple layers of FCNs. However, the integrated features tend to include non-salient regions (due to low level features of the FCN) or lost details of salient objects (due to high level features of the FCN) when producing the saliency maps. In this paper, we develop a novel deep saliency network equipped with recurrently aggregated deep features (RADF) to more accurately detect salient objects from an image by fully exploiting the complementary saliency information captured in different layers. The RADF utilizes the multi-level features integrated from different layers of a FCN to recurrently refine the features at each layer, suppressing the non-salient noise at low-level of the FCN and increasing more salient details into features at high layers. We perform experiments to evaluate the effectiveness of the proposed network on 5 famous saliency detection benchmarks and compare it with 15 state-of-the-art methods. Our method ranks first in 4 of the 5 datasets and second in the left dataset. Xiaowei Hu 0001, Lei Zhu 0003, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
AAAI | 5 |
| 2018 | SFCN-OPI: Detection and Fine-Grained Classification of Nuclei Using Sibling FCN With Objectness Prior InteractionabstractCell nuclei detection and fine-grained classification have been fundamental yet challenging problems in histopathology image analysis. Due to the nuclei tiny size, significant inter-/intra-class variances, as well as the inferior image quality, previous automated methods would easily suffer from limited accuracy and robustness. In the meanwhile, existing approaches usually deal with these two tasks independently, which would neglect the close relatedness of them. In this paper, we present a novel method of sibling fully convolutional network with prior objectness interaction (called SFCN-OPI) to tackle the two tasks simultaneously and interactively using a unified end-to-end framework. Specifically, the sibling FCN branches share features in earlier layers while holding respective higher layers for specific tasks. More importantly, the detection branch outputs the objectness prior which dynamically interacts with the fine-grained classification sibling branch during the training and testing processes. With this mechanism, the fine-grained classification successfully focuses on regions with high confidence of nuclei existence and outputs the conditional probability, which in turn benefits the detection through back propagation. Extensive experiments on colon cancer histology images have validated the effectiveness of our proposed SFCN-OPI and our method has outperformed the state-of-the-art methods by a large margin. Yanning Zhou 0001, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 5 |
| 2018 | Semi-supervised Skin Lesion Segmentation via Transformation Consistent Self-ensembling Model
Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Pheng-Ann Heng |
BMVC | 5 |
| 2018 | Direction-Aware Spatial Context Features for Shadow DetectionabstractShadow detection is a fundamental and challenging task, since it requires an understanding of global image semantics and there are various backgrounds around shadows. This paper presents a novel network for shadow detection by analyzing image context in a direction-aware manner. To achieve this, we first formulate the direction-aware attention mechanism in a spatial recurrent neural network (RNN) by introducing attention weights when aggregating spatial context features in the RNN. By learning these weights through training, we can recover direction-aware spatial context (DSC) for detecting shadows. This design is developed into the DSC module and embedded in a CNN to learn DSC features at different levels. Moreover, a weighted cross entropy loss is designed to make the training more effective. We employ two common shadow detection benchmark datasets and perform various experiments to evaluate our network. Experimental results show that our network outperforms state-of-the-art methods and achieves 97% accuracy and 38% reduction on balance error rate. Xiaowei Hu 0001, Lei Zhu 0003, Chi-Wing Fu, Harry Qin, Pheng-Ann Heng |
CVPR | 5 |
| 2018 | PU-Net: Point Cloud Upsampling NetworkabstractLearning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch convolution unit implicitly in feature space. The expanded feature is then split to a multitude of features, which are then reconstructed to an upsampled point set. Our network is applied at a patch-level, with a joint loss function that encourages the upsampled points to remain on the underlying surface with a uniform distribution. We conduct various experiments using synthesis and scan data to evaluate our method and demonstrate its superiority over some baseline methods and an optimization-based method. Results show that our upsampled points have better uniformity and are located closer to the underlying surfaces. Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
CVPR | 5 |
| 2018 | EC-Net: An Edge-Aware Point Set Consolidation Network
Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ECCV (7) | 5 |
| 2018 | Bidirectional Feature Pyramid Network with Recurrent Attention Residual Modules for Shadow Detection
Lei Zhu 0003, Zijun Deng, Xiaowei Hu 0001, Chi-Wing Fu, Xuemiao Xu, Harry Qin, Pheng-Ann Heng |
ECCV (6) | 7 |
| 2018 | R³Net: Recurrent Residual Refinement Network for Saliency DetectionabstractSaliency detection is a fundamental yet challenging task in computer vision, aiming at highlighting the most visually distinctive objects in an image. We propose a novel recurrent residual refinement network (R^3Net) equipped with residual refinement blocks (RRBs) to more accurately detect salient regions of an input image. Our RRBs learn the residual between the intermediate saliency prediction and the ground truth by alternatively leveraging the low-level integrated features and the high-level integrated features of a fully convolutional network (FCN). While the low-level integrated features are capable of capturing more saliency details, the high-level integrated features can reduce non-salient regions in the intermediate prediction. Furthermore, the RRBs can obtain complementary saliency information of the intermediate prediction, and add the residual into the intermediate prediction to refine the saliency maps. We evaluate the proposed R^3Net on five widely-used saliency detection benchmarks by comparing it with 16 state-of-the-art saliency detectors. Experimental results show that our network outperforms our competitors in all the benchmark datasets. Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xuemiao Xu, Harry Qin, Guoqiang Han 0002, Pheng-Ann Heng |
IJCAI | 7 |
| 2018 | Unsupervised Cross-Modality Domain Adaptation of ConvNets for Biomedical Image Segmentations with Adversarial LossabstractConvolutional networks (ConvNets) have achieved great successes in various challenging vision tasks. However, the performance of ConvNets would degrade when encountering the domain shift. The domain adaptation is more significant while challenging in the field of biomedical image analysis, where cross-modality data have largely different distributions. Given that annotating the medical data is especially expensive, the supervised transfer learning approaches are not quite optimal. In this paper, we propose an unsupervised domain adaptation framework with adversarial learning for cross-modality biomedical image segmentations. Specifically, our model is based on a dilated fully convolutional network for pixel-wise prediction. Moreover, we build a plug-and-play domain adaptation module (DAM) to map the target input to features which are aligned with source domain feature space. A domain critic module (DCM) is set up for discriminating the feature space of both domains. We optimize the DAM and DCM via an adversarial loss without using any target domain label. Our proposed method is validated by adapting a ConvNet trained with MRI images to unpaired CT data for cardiac structures segmentations, and achieved very promising results. Qi Dou 0001, Cheng Ouyang, Cheng Chen 0013, Hao Chen 0011, Pheng-Ann Heng |
IJCAI | 5 |
| 2018 | Deep Attentional Features for Prostate Segmentation in Ultrasound
Yi Wang 0031, Zijun Deng, Xiaowei Hu 0001, Lei Zhu 0003, Xin Yang 0009, Xuemiao Xu, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (4) | 7 |
| 2018 | Generalizing Deep Models for Ultrasound Image Segmentation
Xin Yang 0009, Haoran Dou, Xu Wang 0017, Cheng Bian, Shengli Li 0001, Dong Ni 0001, Pheng-Ann Heng |
MICCAI (4) | 8 |
| 2018 | Augmented Reality-Based Personalized Virtual Operative Anatomy for Neurosurgical Guidance and TrainingabstractThis paper presents a novel augmented reality (AR) interactive environment for neurosurgical training. Comparing with traditional virtual reality based neurosurgical simulator, our system provides a more natural and intuitive fashion for surgeons. To achieve holographic visualization of virtual brain on 3D-printed skull (workspace), the first step is to reconstruct the personalized anatomy structure from segmented MR imaging. Then, tailored to the computational power of HoloLens, we employ the mass-spring method to model the mechanical response of brain. After that, a precise registration method is employed to map the virtual-real spatial information, which can overlay the virtual operative brain on workspace. In addition, bimanual haptic interface is also integrated into our simulator, which is more similar with real neurosurgery. In experiments, we conduct accuracy validation on our registration method, as well as the validity test on the developed simulators. The results demonstrate that our simulator can provide high-accuracy augmented visualization effects and deep immersion for novice surgeons. Weixin Si, Xianavun Liao, Qiong Wang 0001, Pheng-Ann Heng |
VR | 4 |
| 2018 | ScanNet: A Fast and Dense Scanning Framework for Metastastic Breast Cancer Detection from Whole-Slide ImageabstractLymph node metastasis is one of the most significant diagnostic indicators in breast cancer, which is traditionally observed under the microscope by pathologists. In recent years, computerized histology diagnosis has become one of the most rapidly expanding directions in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, automatic detection of lymph node metastases from whole slide images remains a challenging problem, due to the large-scale data with enormous resolutions and existence of hard mimics resulting in a large number of false positives. In this paper, we propose a novel framework by leveraging fully convolutional networks for efficient inference to meet the speed requirement for clinical practice, while reconstructing dense predictions under different offsets for ensuring accurate detection on both microand macro-metastases. Incorporating with the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. Extensive experiments on the benchmark dataset of 2016 Camelyon Grand Challenge corroborated the efficacy of our method. Compared with the state-of-the-art methods, our method achieved superior performance with a faster speed on the tumor localization task and even surpassed human performance on the WSI classification task. Huangjing Lin, Hao Chen 0011, Qi Dou 0001, Liansheng Wang 0002, Harry Qin, Pheng-Ann Heng |
WACV | 6 |
| 2018 | Non-Local Low-Rank Normal Filtering for Mesh DenoisingabstractAbstract This paper presents a non‐local low‐rank normal filtering method for mesh denoising. By exploring the geometric similarity between local surface patches on 3D meshes in the form of normal fields, we devise a low‐rank recovery model that filters normal vectors by means of patch groups. In summary, our method has the following key contributions. First, we present the guided normal patch covariance descriptor to analyze the similarity between patches. Second, we pack normal vectors on similar patches into the normal‐field patch‐group (NPG) matrix for rank analysis. Third, we formulate mesh denoising as a low‐rank matrix recovery problem based on the prior that the rank of the NPG matrix is high for raw meshes with noise, but can be significantly reduced for denoised meshes, whose normal vectors across similar patches should be more strongly correlated. Furthermore, we devise an objective function based on an improved truncated γ norm, and derive an op tim ization procedure using the alternative direction method of multipliers and iteratively re‐weighted least squares techniques. We conducted several experiments to evaluate our method using various 3D models, and compared our results against several state‐of‐the‐art methods. Experimental results show that our method consistently outperforms other methods and better preserves the fine details. Xianzhi Li 0001, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
Comput. Graph. Forum | 4 |
| 2018 | Multiclass support matrix machine for single trial EEG classification
Qingqing Zheng, Harry Qin, Pheng-Ann Heng |
Neurocomputing | 4 |
| 2018 | Feature-preserving ultrasound speckle reduction via L0 minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Chi-Wing Fu, Pheng-Ann Heng |
Neurocomputing | 9 |
| 2018 | 3D multi-scale FCN with random modality voxel dropout learning for Intervertebral Disc Localization and Segmentation from Multi-modality MR Images
Xiaomeng Li 0001, Qi Dou 0001, Hao Chen 0011, Chi-Wing Fu, Xiaojuan Qi 0001, Daniel L. Belavy, Gabriele Armbrecht, Dieter Felsenberg, Guoyan Zheng, Pheng-Ann Heng |
Medical Image Anal. | 10 |
| 2018 | Tracking topology structure adaptively with deep neural networks
Xueying Shi, Guangyong Chen, Pheng-Ann Heng, Zhang Yi 0001 |
Neural Comput. Appl. | 3 |
| 2018 | Sparse Support Matrix Machine
Qingqing Zheng, Harry Qin, Badong Chen, Pheng-Ann Heng |
Pattern Recognit. | 5 |
| 2018 | Large-Scale Bayesian Probabilistic Matrix Factorization with Memo-Free Distributed Variational InferenceabstractBayesian Probabilistic Matrix Factorization (BPMF) is a powerful model in many dyadic data prediction problems, especially the applications of Recommender system. However, its poor scalability has limited its wide applications on massive data. Based on the conditional independence property of observed entries in BPMF model, we propose a novel distributed memo-free variational inference method for large-scale matrix factorization problems. Compared with the state-of-the-art methods, the proposed method is favored for several attractive properties. Specifically, it does not require tuning of learning rate carefully, shuffling the training set at each iteration, or storing massive redundant variables, and can introduce new agents into the computations on the fly. We conduct extensive experiments on both synthetic and real-world datasets. The experimental results show that our method can converge significantly faster with better prediction performance than alternative algorithms. Guangyong Chen, Pheng-Ann Heng |
ACM Trans. Knowl. Discov. Data | 3 |
| 2018 | Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved?abstractDelineation of the left ventricular cavity, myocardium, and right ventricle from cardiac magnetic resonance images (multi-slice 2-D cine MRI) is a common clinical task to establish diagnosis. The automation of the corresponding tasks has thus been the subject of intense research over the past decades. In this paper, we introduce the "Automatic Cardiac Diagnosis Challenge" dataset (ACDC), the largest publicly available and fully annotated dataset for the purpose of cardiac MRI (CMR) assessment. The dataset contains data from 150 multi-equipments CMRI recordings with reference measurements and classification from two medical experts. The overarching objective of this paper is to measure how far state-of-the-art deep learning methods can go at assessing CMRI, i.e., segmenting the myocardium and the two ventricles as well as classifying pathologies. In the wake of the 2017 MICCAI-ACDC challenge, we report results from deep learning methods provided by nine research groups for the segmentation task and four groups for the classification task. Results show that the best methods faithfully reproduce the expert analysis, leading to a mean value of 0.97 correlation score for the automatic extraction of clinical indices and an accuracy of 0.96 for automatic diagnosis. These results clearly open the door to highly accurate and fully automatic analysis of cardiac CMRI. We also identify scenarios for which deep learning methods are still failing. Both the dataset and detailed results are publicly available online, while the platform will remain open for new submissions. Olivier Bernard 0001, Alain Lalande, Clément Zotti, Frederic Cervenansky, Xin Yang 0009, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara 0001, Miguel Ángel González Ballester, Gerard Sanroma, Sandy Napel, Steffen E. Petersen, Georgios Tziritas, Ilias Grinias, Mahendra Khened, Alex Varghese, Ganapathy Krishnamurthi, Marc-Michel Rohé, Xavier Pennec, Maxime Sermesant, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein, Peter M. Full, Ivo Wolf, Sandy Engelhardt, Christian F. Baumgartner, Lisa M. Koch, Jelmer M. Wolterink, Ivana Isgum, Yeonggul Jang, Yoonmi Hong, Jay Patravali, Shubham Jain 0006, Olivier Humbert, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 6 |
| 2018 | SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional NetworkabstractWe propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to form a novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin. Yueming Jin, Qi Dou 0001, Hao Chen 0011, Lequan Yu, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 7 |
| 2018 | H-DenseUNet: Hybrid Densely Connected UNet for Liver and Tumor Segmentation From CT VolumesabstractLiver cancer is one of the leading causes of cancer death. To assist doctors in hepatocellular carcinoma diagnosis and treatment planning, an accurate and automatic liver and tumor segmentation method is highly demanded in the clinical practice. Recently, fully convolutional neural networks (FCNs), including 2-D and 3-D FCNs, serve as the backbone in many volumetric image segmentation. However, 2-D convolutions cannot fully leverage the spatial information along the third dimension while 3-D convolutions suffer from high computational cost and GPU memory consumption. To address these issues, we propose a novel hybrid densely connected UNet (H-DenseUNet), which consists of a 2-D DenseUNet for efficiently extracting intra-slice features and a 3-D counterpart for hierarchically aggregating volumetric contexts under the spirit of the auto-context algorithm for liver and tumor segmentation. We formulate the learning process of the H-DenseUNet in an end-to-end manner, where the intra-slice representations and inter-slice features can be jointly optimized through a hybrid feature fusion layer. We extensively evaluated our method on the data set of the MICCAI 2017 Liver Tumor Segmentation Challenge and 3DIRCADb data set. Our method outperformed other state-of-the-arts on the segmentation results of tumors and achieved very competitive performance for liver segmentation even with a single model. Xiaomeng Li 0001, Hao Chen 0011, Xiaojuan Qi 0001, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Online Robust Projective Dictionary Learning: Shape Modeling for MR-TRUS RegistrationabstractRobust and effective shape prior modeling from a set of training data remains a challenging task, since the shape variation is complicated, and shape models should preserve local details as well as handle shape noises. To address these challenges, a novel robust projective dictionary learning (RPDL) scheme is proposed in this paper. Specifically, the RPDL method integrates the dimension reduction and dictionary learning into a unified framework for shape prior modeling, which can not only learn a robust and representative dictionary with the energy preservation of the training data, but also reduce the dimensionality and computational cost via the subspace learning. In addition, the proposed RPDL algorithm is regularized by using the norm to handle the outliers and noises, and is embedded in an online framework so that of memory and time efficiency. The proposed method is employed to model prostate shape prior for the application of magnetic resonance transrectal ultrasound registration. The experimental results demonstrate that our method provides more accurate and robust shape modeling than the state-of-the-art methods do. The proposed RPDL method is applicable for modeling other organs, and hence, a general solution for the problem of shape prior modeling. Yi Wang 0031, Qingqing Zheng, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Thin-Feature-Aware Transport-Velocity Formulation for SPH-Based Liquid AnimationabstractRealistic liquid animations with thin sheets or streams are crucial for creating fluid effects in digital media. However, it is challenging to simulate these appealing thin sheets or streams in the framework of smoothed particle hydrodynamics (SPH). The underlying reason for this challenge mainly lies in the inherent numerical instability of SPH due to inconsistent kernel interpolation, which is caused by the incomplete kernel support on the free surface and the particles' disorder dispersion within the simulation domain. To address this challenge, we propose a novel and effective approach to ensure the consistency of kernel interpolation at both internal flow and the free surface during the simulation such that these thin features can always be well maintained. First, we introduce a transport-velocity formulation to alleviate the disorder dispersion in the liquid domain. However, this formulation can only work in the internal flow, and it fails at the free surface because it cannot accurately estimate the density of particles there. To this end, we propose adaptively correcting the underestimated density caused by the incomplete kernel support of free-surface particles, which are identified by a geometry-aware anisotropic kernel, to counteract the inconsistent interpolation on the free surface. Then, we propose a novel scheme to further filter the background pressure to enhance the interactions between the internal flow and the free surface, as well as liquid and solid, such that the thin features generated from such interactions can be realistically simulated. The proposed approach can also achieve anticlumping and regularization effects in the entire simulation domain and, hence, further enhance the thin features in liquids. We evaluate our method on a variety of benchmark examples, and the results demonstrate that our method can achieve more appealing visual effects than state-of-the-art methods by realistically simulating more vivid thin features. Weixin Si, Harry Qin, Zhuchao Chen, Xiangyun Liao, Qiong Wang 0001, Pheng-Ann Heng |
IEEE Trans. Multim. | 6 |
| 2018 | Animating Wall-Bounded Turbulent Smoke via Filament-Mesh Particle-Particle MethodabstractTurbulent vortices in smoke flows are crucial for a visually interesting appearance. Unfortunately, it is challenging to efficiently simulate these appealing effects in the framework of vortex filament methods. The vortex filaments in grids scheme allows to efficiently generate turbulent smoke with macroscopic vortical structures, but suffers from the projection-related dissipation, and thus the small-scale vortical structures under grid resolution are hard to capture. In addition, this scheme cannot be applied in wall-bounded turbulent smoke simulation, which requires efficiently handling smoke-obstacle interaction and creating vorticity at the obstacle boundary. To tackle above issues, we propose an effective filament-mesh particle-particle (FMPP) method for fast wall-bounded turbulent smoke simulation with ample details. The Filament-Mesh component approximates the smooth long-range interactions by splatting vortex filaments on grid, solving the Poisson problem with a fast solver, and then interpolating back to smoke particles. The Particle-Particle component introduces smoothed particle hydrodynamics (SPH) turbulence model for particles in the same grid, where interactions between particles cannot be properly captured under grid resolution. Then, we sample the surface of obstacles with boundary particles, allowing the interaction between smoke and obstacle being treated as pressure forces in SPH. Besides, the vortex formation region is defined at the back of obstacles, providing smoke particles flowing by the separation particles with a vorticity force to simulate the subsequent vortex shedding phenomenon. The proposed approach can synthesize the lost small-scale vortical structures and also achieve the smoke-obstacle interaction with vortex shedding at obstacle boundaries in a lightweight manner. The experimental results demonstrate that our FMPP method can achieve more appealing visual effects than vortex filaments in grids scheme by efficiently simulating more vivid thin turbulent features. Xiangyun Liao, Weixin Si, Hanqiu Sun, Harry Qin, Qiong Wang 0001, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2017 | Fine-Grained Recurrent Neural Networks for Automatic Prostate Segmentation in Ultrasound ImagesabstractBoundary incompleteness raises great challenges to automatic prostate segmentation in ultrasound images. Shape prior can provide strong guidance in estimating the missing boundary, but traditional shape models often suffer from hand-crafted descriptors and local information loss in the fitting procedure. In this paper, we attempt to address those issues with a novel framework. The proposed framework can seamlessly integrate feature extraction and shape prior exploring, and estimate the complete boundary with a sequential manner. Our framework is composed of three key modules. Firstly, we serialize the static 2D prostate ultrasound images into dynamic sequences and then predict prostate shapes by sequentially exploring shape priors. Intuitively, we propose to learn the shape prior with the biologically plausible Recurrent Neural Networks (RNNs). This module is corroborated to be effective in dealing with the boundary incompleteness. Secondly, to alleviate the bias caused by different serialization manners, we propose a multi-view fusion strategy to merge shape predictions obtained from different perspectives. Thirdly, we further implant the RNN core into a multiscale Auto-Context scheme to successively refine the details of the shape prediction map. With extensive validation on challenging prostate ultrasound images, our framework bridges severe boundary incompleteness and achieves the best performance in prostate boundary delineation when compared with several advanced methods. Additionally, our approach is general and can be extended to other medical image segmentation tasks, where boundary incompleteness is one of the main challenges. Xin Yang 0009, Lequan Yu, Lingyun Wu, Yi Wang 0031, Dong Ni 0001, Harry Qin, Pheng-Ann Heng |
AAAI | 7 |
| 2017 | Volumetric ConvNets with Mixed Residual Connections for Automated Prostate Segmentation from 3D MR ImagesabstractAutomated prostate segmentation from 3D MR images is very challenging due to large variations of prostate shape and indistinct prostate boundaries. We propose a novel volumetric convolutional neural network (ConvNet) with mixed residual connections to cope with this challenging problem. Compared with previous methods, our volumetric ConvNet has two compelling advantages. First, it is implemented in a 3D manner and can fully exploit the 3D spatial contextual information of input data to perform efficient, precise and volume-to-volume prediction. Second and more important, the novel combination of residual connections (i.e., long and short) can greatly improve the training efficiency and discriminative capability of our network by enhancing the information propagation within the ConvNet both locally and globally. While the forward propagation of location information can improve the segmentation accuracy, the smooth backward propagation of gradient flow can accelerate the convergence speed and enhance the discrimination capability. Extensive experiments on the open MICCAI PROMISE12 challenge dataset corroborated the effectiveness of the proposed volumetric ConvNet with mixed residual connections. Our method ranked the first in the challenge, outperforming other competitors by a large margin with respect to most of evaluation metrics. The proposed volumetric ConvNet is general enough and can be easily extended to other medical image analysis tasks, especially ones with limited training data. Lequan Yu, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 5 |
| 2017 | A Non-local Low-Rank Framework for Ultrasound Speckle ReductionabstractSpeckle refers to the granular patterns that occur in ultrasound images due to wave interference. Speckle removal can greatly improve the visibility of the underlying structures in an ultrasound image and enhance subsequent post processing. We present a novel framework for speckle removal based on low-rank non-local filtering. Our approach works by first computing a guidance image that assists in the selection of candidate patches for non-local filtering in the face of significant speckles. The candidate patches are further refined using a low-rank minimization estimated using a truncated weighted nuclear norm (TWNN) and structured sparsity. We show that the proposed filtering framework produces results that outperform state-of-the-art methods both qualitatively and quantitatively. This framework also provides better segmentation results when used for pre-processing ultrasound images. Lei Zhu 0003, Chi-Wing Fu, Michael S. Brown, Pheng-Ann Heng |
CVPR | 4 |
| 2017 | Cascaded Feature Network for Semantic Segmentation of RGB-D ImagesabstractFully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional information to improve the segmentation performance. In this paper, we present a neural network with multiple branches for segmenting RGB-D images. Our approach is to use the available depth to split the image into layers with common visual characteristic of objects/scenes, or common “scene-resolution”. We introduce context-aware receptive field (CaRF) which provides a better control on the relevant contextual information of the learned features. Equipped with CaRF, each branch of the network semantically segments relevant similar scene-resolution, leading to a more focused domain which is easier to learn. Furthermore, our network is cascaded with features from one branch augmenting the features of adjacent branch. We show that such cascading of features enriches the contextual information of each branch and enhances the overall performance. The accuracy that our network achieves outperforms the state-of-the-art methods on two public datasets. Di Lin 0002, Guangyong Chen, Daniel Cohen-Or, Pheng-Ann Heng, Hui Huang 0004 |
ICCV | 4 |
| 2017 | Online Robust Image Alignment via Subspace Learning from Gradient OrientationsabstractRobust and efficient image alignment remains a challenging task, due to the massiveness of images, great illumination variations between images, partial occlusion and corruption. To address these challenges, we propose an online image alignment method via subspace learning from image gradient orientations (IGO). The proposed method integrates the subspace learning, transformed IGO reconstruction and image alignment into a unified online framework, which is robust for aligning images with severe intensity distortions. Our method is motivated by principal component analysis (PCA) from gradient orientations provides more reliable low-dimensional subspace than that from pixel intensities. Instead of processing in the intensity domain like conventional methods, we seek alignment in the IGO domain such that the aligned IGO of the newly arrived image can be decomposed as the sum of a sparse error and a linear composition of the IGO-PCA basis learned from previously well-aligned ones. The optimization problem is accomplished by an iterative linearization that minimizes the L1-norm of the sparse error. Furthermore, the IGO-PCA basis is adaptively updated based on incremental thin singular value decomposition (SVD) which takes the shift of IGO mean into consideration. The efficacy of the proposed method is validated on extensive challenging datasets through image alignment and face recognition. Experimental results demonstrate that our algorithm provides more illumination- and occlusion-robust image alignment than state-of-the-art methods do. Qingqing Zheng, Yi Wang 0003, Pheng-Ann Heng |
ICCV | 3 |
| 2017 | Joint Bi-layer Optimization for Single-Image Rain Streak RemovalabstractWe present a novel method for removing rain streaks from a single input image by decomposing it into a rain-free background layer B and a rain-streak layer R. A joint optimization process is used that alternates between removing rain-streak details from B and removing non-streak details from R. The process is assisted by three novel image priors. Observing that rain streaks typically span a narrow range of directions, we first analyze the local gradient statistics in the rain image to identify image regions that are dominated by rain streaks. From these regions, we estimate the dominant rain streak direction and extract a collection of rain-dominated patches. Next, we define two priors on the background layer B, one based on a centralized sparse representation and another based on the estimated rain direction. A third prior is defined on the rain-streak layer R, based on similarity of patches to the extracted rain patches. Both visual and quantitative comparisons demonstrate that our method outperforms the state-of-the-art. Lei Zhu 0003, Chi-Wing Fu, Dani Lischinski, Pheng-Ann Heng |
ICCV | 4 |
| 2017 | Learning to Aggregate Ordinal Labels by Maximizing Separating WidthabstractWhile crowdsourcing has been a cost and time efficient method to label massive samples, one critical issue is quality control, for which the key challenge is to infer the ground truth from noisy or even adversarial data by various users. A large class of crowdsourcing problems, such as those involving age, grade, level, or stage, have an ordinal structure in their labels. Based on a technique of sampling estimated label from the posterior distribution, we define a novel separating width among the labeled observations to characterize the quality of sampled labels, and develop an efficient algorithm to optimize it through solving multiple linear decision boundaries and adjusting prior distributions. Our algorithm is empirically evaluated on several real world datasets, and demonstrates its supremacy over state-of-the-art methods. Guangyong Chen, Shengyu Zhang 0002, Di Lin 0002, Hui Huang 0004, Pheng-Ann Heng |
ICML | 5 |
| 2017 | Automated Pulmonary Nodule Detection via 3D ConvNets with Online Sample Filtering and Hybrid-Loss Residual Learning
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Huangjing Lin, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 6 |
| 2017 | Towards Automatic Semantic Segmentation in Volumetric Ultrasound
Xin Yang 0009, Lequan Yu, Shengli Li 0001, Xu Wang 0017, Harry Qin, Dong Ni 0001, Pheng-Ann Heng |
MICCAI (1) | 8 |
| 2017 | Automatic 3D Cardiovascular MR Segmentation with Densely-Connected Volumetric ConvNets
Lequan Yu, Jie-Zhi Cheng, Qi Dou 0001, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 7 |
| 2017 | Patch green coordinates based interactive embedded deformable modelabstractVirtual surgery is a serious game which provides an opportunity to acquire cognitive and technical surgical skills via virtual surgical training and planning. However, interactively and realistically manipulating the human organ and simulating its motion under interaction is still a challenging task in this field. The underlying reason for this issue is the conflict requirements for physical constraints with high fidelity and real-time performance. To achieve realistic simulation of human organ motion with volume conservation, smooth interpolation under large deformation and precise frictional contact mechanics of global behavior in surgical scenario. This paper presents a novel and effective patch Green coordinates based interpolation for embedded deformable model to achieve the volume-preserving and smooth interpolation effects. Besides, we resolve the frictional contact mechanics for embedded deformable model, and further provide the precise boundary conditions for mechanical solver. In addition, our embedded deformable model is based on the total lagrangian explicit dynamics (TLED) finite element method (FEM) solver, which can well handle the large biological tissue deformation with both nonlinear geometric and material properties. In real compression experiments, our method can achieve liver deformation with average accuracy of 3.02 mm. Besides, the experimental results demonstrate that our method can also achieve smoother interpolation and volume-preserving effects than original embedded deformable model, and allows complex and accurate organ motion with mechanical interactions in virtual surgery. Weixin Si, Xiangyun Liao, Qiong Wang 0001, Harry Qin, Pheng-Ann Heng |
MIG | 6 |
| 2017 | Filament-based realistic turbulent wake synthesisabstractAbstract Turbulent wake is crucial for the visually appealing effects of liquid. Unfortunately, it is challenging to realistically simulate this phenomenon with ring‐shaped vortical structures. To tackle this issue, we propose a filament‐based turbulent wake synthesis method for realistically simulating the turbulent wake with ring‐shaped vortical structures. The filaments are sampled at the separation points on the obstacle surface and emitted into the liquid flow to generate structured turbulent wake. Besides, the surface tension model is incorporated to generate natural turbulent wake diffusion visual effects in liquid by the anticurvature effects. The proposed approach can realistically and effectively synthesize the turbulent wake with ring‐shaped vortical structures and make it diffuse naturally. The experimental results demonstrate that our method outperforms than the vortex particle‐based method in synthesizing appealing turbulent wake. Xiangyun Liao, Weixin Si, Qiong Wang 0001, Pheng-Ann Heng |
Comput. Animat. Virtual Worlds | 6 |
| 2017 | DCAN: Deep contour-aware networks for object instance segmentation from histology images
Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 6 |
| 2017 | 3D deeply supervised network for automated segmentation of volumetric medical images
Qi Dou 0001, Lequan Yu, Hao Chen 0011, Yueming Jin, Xin Yang 0009, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 7 |
| 2017 | Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: The LUNA16 challenge
Arnaud A. A. Setio, Alberto Traverso, Thomas de Bel, Moira S. N. Berens, Cas van den Bogaard, Piergiorgio Cerello, Hao Chen 0011, Qi Dou 0001, Maria Evelina Fantacci, Bram Geurts, Robbert van der Gugten, Pheng-Ann Heng, Bart Jansen 0001, Michael M. J. de Kaste, Valentin Kotov, Jack Yu-Hung Lin, Jeroen T. M. C. Manders, Alexander Sóñora-Mengana, Juan Carlos García-Naranjo, Evgenia Papavasileiou, Mathias Prokop |
Medical Image Anal. | 12 |
| 2017 | Gland segmentation in colon histology images: The glas challenge contest
Korsuk Sirinukunwattana, Josien P. W. Pluim, Hao Chen 0011, Xiaojuan Qi 0001, Pheng-Ann Heng, Li Yang Wang, Bogdan J. Matuszewski, Elia Bruni, Urko Sanchez, Anton Böhm, Olaf Ronneberger, Bassem Ben Cheikh, Daniel Racoceanu, Philipp Kainz, Michael Pfeiffer 0001, Martin Urschler, David R. J. Snead, Nasir M. Rajpoot |
Medical Image Anal. | 5 |
| 2017 | Evaluation and comparison of 3D intervertebral disc localization and segmentation methods for 3D T2 MR data: A grand challenge
Guoyan Zheng, Chengwen Chu, Daniel L. Belavy, Bulat Ibragimov, Robert Korez, Tomaz Vrtovec, Hugo Hutt, Richard M. Everson, Judith Meakin, Isabel Lopez Andrade, Ben Glocker, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng, Chunliang Wang, Daniel Forsberg, Ales Neubert, Jurgen Fripp, Martin Urschler, Darko Stern, Maria Wimmer 0002 |
Medical Image Anal. | 14 |
| 2017 | Blind Image Denoising via Dependent Dirichlet Process TreeabstractMost existing image denoising approaches assumed the noise to be homogeneous white Gaussian distributed with known intensity. However, in real noisy images, the noise models are usually unknown beforehand and can be much more complex. This paper addresses this problem and proposes a novel blind image denoising algorithm to recover the clean image from noisy one with the unknown noise model. To model the empirical noise of an image, our method introduces the mixture of Gaussian distribution, which is flexible enough to approximate different continuous distributions. The problem of blind image denoising is reformulated as a learning problem. The procedure is to first build a two-layer structural model for noisy patches and consider the clean ones as latent variable. To control the complexity of the noisy patch model, this work proposes a novel Bayesian nonparametric prior called "Dependent Dirichlet Process Tree" to build the model. Then, this study derives a variational inference algorithm to estimate model parameters and recover clean patches. We apply our method on synthesis and real noisy images with different noise models. Comparing with previous approaches, ours achieves better performance. The experimental results indicate the efficiency of the proposed algorithm to cope with practical image denoising tasks. Guangyong Chen, Jianye Hao, Pheng-Ann Heng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Fast feature-preserving speckle reduction for ultrasound images via phase congruency
Lei Zhu 0003, Weiming Wang 0002, Harry Qin, Kin Hong Wong, Kup-Sze Choi, Pheng-Ann Heng |
Signal Process. | 6 |
| 2017 | Ultrasound Standard Plane Detection Using a Composite Neural Network FrameworkabstractUltrasound (US) imaging is a widely used screening tool for obstetric examination and diagnosis. Accurate acquisition of fetal standard planes with key anatomical structures is very crucial for substantial biometric measurement and diagnosis. However, the standard plane acquisition is a labor-intensive task and requires operator equipped with a thorough knowledge of fetal anatomy. Therefore, automatic approaches are highly demanded in clinical practice to alleviate the workload and boost the examination efficiency. The automatic detection of standard planes from US videos remains a challenging problem due to the high intraclass and low interclass variations of standard planes, and the relatively low image quality. Unlike previous studies which were specifically designed for individual anatomical standard planes, respectively, we present a general framework for the automatic identification of different standard planes from US videos. Distinct from conventional way that devises hand-crafted visual features for detection, our framework explores in- and between-plane feature learning with a novel composite framework of the convolutional and recurrent neural networks. To further address the issue of limited training data, a multitask learning framework is implemented to exploit common knowledge across detection tasks of distinctive standard planes for the augmentation of feature learning. Extensive experiments have been conducted on hundreds of US fetus videos to corroborate the better efficacy of the proposed framework on the difficult standard plane detection problem. Hao Chen 0011, Lingyun Wu, Qi Dou 0001, Harry Qin, Shengli Li 0001, Jie-Zhi Cheng, Dong Ni 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 8 |
| 2017 | Integrating Online and Offline Three-Dimensional Deep Learning for Automated Polyp Detection in Colonoscopy VideosabstractAutomated polyp detection in colonoscopy videos has been demonstrated to be a promising way for colorectal cancer prevention and diagnosis. Traditional manual screening is time consuming, operator dependent, and error prone; hence, automated detection approach is highly demanded in clinical practice. However, automated polyp detection is very challenging due to high intraclass variations in polyp size, color, shape, and texture, and low interclass variations between polyps and hard mimics. In this paper, we propose a novel offline and online three-dimensional (3-D) deep learning integration framework by leveraging the 3-D fully convolutional network (3D-FCN) to tackle this challenging problem. Compared with the previous methods employing hand-crafted features or 2-D convolutional neural network, the 3D-FCN is capable of learning more representative spatio-temporal features from colonoscopy videos, and hence has more powerful discrimination capability. More importantly, we propose a novel online learning scheme to deal with the problem of limited training data by harnessing the specific information of an input video in the learning process. We integrate offline and online learning to effectively reduce the number of false positives generated by the offline network and further improve the detection performance. Extensive experiments on the dataset of MICCAI 2015 Challenge on Polyp Detection demonstrated the better performance of our method when compared with other competitors. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 5 |
| 2017 | Automated Melanoma Recognition in Dermoscopy Images via Very Deep Residual NetworksabstractAutomated melanoma recognition in dermoscopy images is a very challenging task due to the low contrast of skin lesions, the huge intraclass variation of melanomas, the high degree of visual similarity between melanoma and non-melanoma lesions, and the existence of many artifacts in the image. In order to meet these challenges, we propose a novel method for melanoma recognition by leveraging very deep convolutional neural networks (CNNs). Compared with existing methods employing either low-level hand-crafted features or CNNs with shallower architectures, our substantially deeper networks (more than 50 layers) can acquire richer and more discriminative features for more accurate recognition. To take full advantage of very deep networks, we propose a set of schemes to ensure effective training and learning under limited training data. First, we apply the residual learning to cope with the degradation and overfitting problems when a network goes deeper. This technique can ensure that our networks benefit from the performance gains achieved by increasing network depth. Then, we construct a fully convolutional residual network (FCRN) for accurate skin lesion segmentation, and further enhance its capability by incorporating a multi-scale contextual information integration scheme. Finally, we seamlessly integrate the proposed FCRN (for segmentation) and other very deep residual networks (for classification) to form a two-stage framework. This framework enables the classification network to extract more representative and specific features based on segmented results instead of the whole dermoscopy images, further alleviating the insufficiency of training data. The proposed framework is extensively evaluated on ISBI 2016 Skin Lesion Analysis Towards Melanoma Detection Challenge dataset. Experimental results demonstrate the significant performance gains of the proposed framework, ranking the first in classification and the second in segmentation among 25 teams and 28 teams, respectively. This study corroborates that very deep CNNs with effective training mechanisms can be employed to solve complicated medical image analysis tasks, even with limited training data. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Reconfigurable interlocking furnitureabstractReconfigurable assemblies consist of a common set of parts that can be assembled into different forms for use in different situations. Designing these assemblies is a complex problem, since it requires a compatible decomposition of shapes with correspondence across forms, and a planning of well-matched joints to connect parts in each form. This paper presents computational methods as tools to assist the design and construction of reconfigurable assemblies, typically for furniture. There are three key contributions in this work. First, we present the compatible decomposition as a weakly-constrained dissection problem, and derive its solution based on a dynamic bipartite graph to construct parts across multiple forms; particularly, we optimize the parts reuse and preserve the geometric semantics. Second, we develop a joint connection graph to model the solution space of reconfigurable assemblies with part and joint compatibility across different forms. Third, we formulate the backward interlocking and multi-key interlocking models, with which we iteratively plan the joints consistently over multiple forms. We show the applicability of our approach by constructing reconfigurable furniture of various complexities, extend it with recursive connections to generate extensible and hierarchical structures, and fabricate a number of results using 3D printing, 2D laser cutting, and woodworking. Peng Song 0001, Chi-Wing Fu, Yueming Jin, Hongfei Xu, Ligang Liu 0001, Pheng-Ann Heng, Daniel Cohen-Or |
ACM Trans. Graph. | 6 |
| 2016 | Mitosis Detection in Breast Cancer Histology Images via Deep Cascaded NetworksabstractThe number of mitoses per tissue area gives an important aggressiveness indication of the invasive breast carcinoma.However, automatic mitosis detection in histology images remains a challenging problem. Traditional methods either employ hand-crafted features to discriminate mitoses from other cells or construct a pixel-wise classifier to label every pixel in a sliding window way. While the former suffers from the large shape variation of mitoses and the existence of many mimics with similar appearance, the slow speed of the later prohibits its use in clinical practice.In order to overcome these shortcomings, we propose a fast and accurate method to detect mitosis by designing a novel deep cascaded convolutional neural network, which is composed of two components. First, by leveraging the fully convolutional neural network, we propose a coarse retrieval model to identify and locate the candidates of mitosis while preserving a high sensitivity.Based on these candidates, a fine discrimination model utilizing knowledge transferred from cross-domain is developed to further single out mitoses from hard mimics.Our approach outperformed other methods by a large margin in 2014 ICPR MITOS-ATYPIA challenge in terms of detection accuracy. When compared with the state-of-the-art methods on the 2012 ICPR MITOSIS data (a smaller and less challenging dataset), our method achieved comparable or better results with a roughly 60 times faster speed. Hao Chen 0011, Qi Dou 0001, Xi Wang 0013, Harry Qin, Pheng-Ann Heng |
AAAI | 5 |
| 2016 | Deep Contextual Networks for Neuronal Structure SegmentationabstractThe goal of connectomics is to manifest the interconnections of neural system with the Electron Microscopy (EM) images. However, the formidable size of EM image data renders human annotation impractical, as it may take decades to fulfill the whole job. An alternative way to reconstruct the connectome can be attained with the computerized scheme that can automatically segment the neuronal structures. The segmentation of EM images is very challenging as the depicted structures can be very diverse.To address this difficult problem, a deep contextual network is proposed here by leveraging multi-level contextual information from the deep hierarchical structure to achieve better segmentation performance.To further improve the robustness against the vanishing gradients and strengthen the capability of the back-propagation of gradient flow, auxiliary classifiers are incorporated in the architecture of our deep neural network. It will be shown that our method can effectively parse the semantic meaning from the images with the underlying neural network and accurately delineate the structural boundaries with the reference of low-level contextual cues. Experimental results on the benchmark dataset of 2012 ISBI segmentation challenge of neuronal structures suggest that the proposed method can outperform the state-of-the-art methods by a large margin with respect to different evaluation measurements. Our method can potentially facilitate the automatic connectome analysis from EM images with less human intervention effort. Hao Chen 0011, Xiaojuan Qi 0001, Jie-Zhi Cheng, Pheng-Ann Heng |
AAAI | 4 |
| 2016 | Ultrasound Speckle Reduction via L_0 Minimization
Lei Zhu 0003, Weiming Wang 0002, Xiaomeng Li 0001, Qiong Wang 0001, Harry Qin, Kin Hong Wong, Pheng-Ann Heng |
ACCV (3) | 7 |
| 2016 | DCAN: Deep Contour-Aware Networks for Accurate Gland SegmentationabstractThe morphology of glands has been used routinely by pathologists to assess the malignancy degree of adenocarcinomas. Accurate segmentation of glands from histology images is a crucial step to obtain reliable morphological statistics for quantitative diagnosis. In this paper, we proposed an efficient deep contour-aware network (DCAN) to solve this challenging problem under a unified multi-task learning framework. In the proposed network, multi-level contextual features from the hierarchical architecture are explored with auxiliary supervision for accurate gland segmentation. When incorporated with multi-task regularization during the training, the discriminative capability of intermediate features can be further improved. Moreover, our network can not only output accurate probability maps of glands, but also depict clear contours simultaneously for separating clustered objects, which further boosts the gland segmentation performance. This unified framework can be efficient when applied to large-scale histopathological data without resorting to additional steps to generate contours based on low-level cues for post-separating. Our method won the 2015 MICCAI Gland Segmentation Challenge out of 13 competitive teams, surpassing all the other methods by a significant margin. Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Pheng-Ann Heng |
CVPR | 4 |
| 2016 | From Noise Modeling to Blind Image DenoisingabstractTraditional image denoising algorithms always assume the noise to be homogeneous white Gaussian distributed. However, the noise on real images can be much more complex empirically. This paper addresses this problem and proposes a novel blind image denoising algorithm which can cope with real-world noisy images even when the noise model is not provided. It is realized by modeling image noise with mixture of Gaussian distribution (MoG) which can approximate large varieties of continuous distributions. As the number of components for MoG is unknown practically, this work adopts Bayesian nonparametric technique and proposes a novel Low-rank MoG filter (LR-MoG) to recover clean signals (patches) from noisy ones contaminated by MoG noise. Based on LR-MoG, a novel blind image denoising approach is developed. To test the proposed method, this study conducts extensive experiments on synthesis and real images. Our method achieves the state-of the-art performance consistently. Guangyong Chen, Pheng-Ann Heng |
CVPR | 3 |
| 2016 | A Bayesian Nonparametric Approach to Dynamic Dyadic Data PredictionabstractAn important issue of using matrix factorization for recommender systems is to capture the dynamics of user preference over time for more accurate prediction. We find that considering the existence of clusters among users with respect to evolution behavior of their preference can improve performance effectively. This is especially important to commercial recommender systems, where the evolution of preference for different users is heterogeneous, and historical ratings are not enough to estimate the preference of each user individually. Based on this, we propose a novel Bayesian nonparametric method based on the Dirichlet process, to detect users sharing the same evolution behavior of their preference. For each community, we use vector autoregressive model (VAR) to capture the evolution to explore higher-order dependency on historical user preference, and incorporate this feature with a novel adaptive prior strategy. We also derive variational inference approach to infer our method. Finally, we conduct extensive empirical experiments to show the advantage of our method over state-of-the-art algorithms. Guangyong Chen, Pheng-Ann Heng |
ICDM | 3 |
| 2016 | Iterative Multi-domain Regularized Deep Learning for Anatomical Structure Detection and Segmentation from Ultrasound Images
Hao Chen 0011, Yefeng Zheng 0001, Jin Hyeong Park, Pheng-Ann Heng, Shaohua Kevin Zhou |
MICCAI (2) | 4 |
| 2016 | 3D Deeply Supervised Network for Automatic Liver Segmentation from CT Volumes
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Lequan Yu, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 6 |
| 2016 | Non-Local Sparse and Low-Rank Regularization for Structure-Preserving Image SmoothingabstractAbstract This paper presents a new image smoothing method that better preserves prominent structures. Our method is inspired by the recent non‐local image processing techniques on the patch grouping and filtering. Overall, it has three major contributions over previous works. First, we employ the diffusion map as the guidance image to improve the accuracy of patch similarity estimation using the region covariance descriptor. Second, we model structure‐preserving image smoothing as a low‐rank matrix recovery problem, aiming at effectively filtering the texture information in similar patches. Lastly, we devise an objective function, namely the weighted robust principle component analysis (WRPCA), by regularizing the low rank with the weighted nuclear norm and sparsity pursuit with L1norm, and solve this non‐convex WRPCA optimization problem by adopting the alternative direction method of multipliers (ADMM) technique. We experiment our method with a wide variety of images and compare it against several state‐of‐the‐art methods. The results show that our method achieves better structure preservation and texture suppression as compared to other methods. We also show the applicability of our method on several image processing tasks such as edge detection, texture enhancement and seam carving. Lei Zhu 0003, Chi-Wing Fu, Yueming Jin, Mingqiang Wei, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 6 |
| 2016 | Automatic Detection of Cerebral Microbleeds From MR Images via 3D Convolutional Neural NetworksabstractCerebral microbleeds (CMBs) are small haemorrhages nearby blood vessels. They have been recognized as important diagnostic biomarkers for many cerebrovascular diseases and cognitive dysfunctions. In current clinical routine, CMBs are manually labelled by radiologists but this procedure is laborious, time-consuming, and error prone. In this paper, we propose a novel automatic method to detect CMBs from magnetic resonance (MR) images by exploiting the 3D convolutional neural network (CNN). Compared with previous methods that employed either low-level hand-crafted descriptors or 2D CNNs, our method can take full advantage of spatial contextual information in MR volumes to extract more representative high-level features for CMBs, and hence achieve a much better detection accuracy. To further improve the detection performance while reducing the computational cost, we propose a cascaded framework under 3D CNNs for the task of CMB detection. We first exploit a 3D fully convolutional network (FCN) strategy to retrieve the candidates with high probabilities of being CMBs, and then apply a well-trained 3D CNN discrimination model to distinguish CMBs from hard mimics. Compared with traditional sliding window strategy, the proposed 3D FCN strategy can remove massive redundant computations and dramatically speed up the detection process. We constructed a large dataset with 320 volumetric MR scans and performed extensive experiments to validate the proposed method, which achieved a high sensitivity of 93.16% with an average number of 2.74 false positives per subject, outperforming previous methods using low-level descriptors or 2D CNNs by a significant margin. The proposed method, in principle, can be adapted to other biomarker detection tasks from volumetric medical data. Qi Dou 0001, Hao Chen 0011, Lequan Yu, Lei Zhao 0003, Harry Qin, Defeng Wang, Vincent C. T. Mok, Lin Shi 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 9 |
| 2016 | Towards Personalized Statistical Deformable Model and Hybrid Point Matching for Robust MR-TRUS RegistrationabstractRegistration and fusion of magnetic resonance (MR) and 3D transrectal ultrasound (TRUS) images of the prostate gland can provide high-quality guidance for prostate interventions. However, accurate MR-TRUS registration remains a challenging task, due to the great intensity variation between two modalities, the lack of intrinsic fiducials within the prostate, the large gland deformation caused by the TRUS probe insertion, and distinctive biomechanical properties in patients and prostate zones. To address these challenges, a personalized model-to-surface registration approach is proposed in this study. The main contributions of this paper can be threefold. First, a new personalized statistical deformable model (PSDM) is proposed with the finite element analysis and the patient-specific tissue parameters measured from the ultrasound elastography. Second, a hybrid point matching method is developed by introducing the modality independent neighborhood descriptor (MIND) to weight the Euclidean distance between points to establish reliable surface point correspondence. Third, the hybrid point matching is further guided by the PSDM for more physically plausible deformation estimation. Eighteen sets of patient data are included to test the efficacy of the proposed method. The experimental results demonstrate that our approach provides more accurate and robust MR-TRUS registration than state-of-the-art methods do. The averaged target registration error is 1.44 mm, which meets the clinical requirement of 1.9 mm for the accurate tumor volume detection. It can be concluded that the presented method can effectively fuse the heterogeneous image information in the elastography, MR, and TRUS to attain satisfactory image alignment performance. Yi Wang 0031, Jie-Zhi Cheng, Dong Ni 0001, Muqing Lin, Harry Qin, Xióngbiao Luó, Xiaoyan Xie, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 9 |
| 2016 | Globally optimal toon trackingabstractThe ability to identify objects or region correspondences between consecutive frames of a given hand-drawn animation sequence is an indispensable tool for automating animation modification tasks such as sequence-wide recoloring or shape-editing of a specific animated character. Existing correspondence identification methods heavily rely on appearance features, but these features alone are insufficient to reliably identify region correspondences when there exist occlusions or when two or more objects share similar appearances. To resolve the above problems, manual assistance is often required. In this paper, we propose a new correspondence identification method which considers both appearance features and motions of regions in a global manner. We formulate correspondence likelihoods between temporal region pairs as a network flow graph problem which can be solved by a well-established optimization algorithm. We have evaluated our method with various animation sequences and results show that our method consistently outperforms the state-of-the-art methods without any user guidance. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 4 |
| 2015 | The Dynamic Chinese Restaurant Process via Birth and Death ProcessesabstractWe develop the Dynamic Chinese Restaurant Process (DCRP) which incorporates time-evolutionary feature in dependent Dirichlet Process mixture models. This model can capture the dynamic change of mixture components, allowing clusters to emerge, vanish and vary over time. All these macroscopic changes are controlled by tracing the birth and death of every single element. We investigate the properties of dependent Dirichlet Process mixture model based on DCRP and develop corresponding Gibbs Sampler for posterior inference. We also conduct simulation and empirical studies to compare this model with traditional CRP and related models. The results show that this model can provide better results for sequential data, especially for data with heterogeneous lifetime distribution. Pheng-Ann Heng |
AAAI | 3 |
| 2015 | An Efficient Statistical Method for Image Noise Level EstimationabstractIn this paper, we address the problem of estimating noise level from a single image contaminated by additive zero-mean Gaussian noise. We first provide rigorous analysis on the statistical relationship between the noise variance and the eigenvalues of the covariance matrix of patches within an image, which shows that many state-of-the-art noise estimation methods underestimate the noise level of an image. To this end, we derive a new nonparametric algorithm for efficient noise level estimation based on the observation that patches decomposed from a clean image often lie around a low-dimensional subspace. The performance of our method has been guaranteed both theoretically and empirically. Specifically, our method outperforms existing state-of-the-art algorithms on estimating noise level with the least executing time in our experiments. We further demonstrate that the denoising algorithm BM3D algorithm achieves optimal performance using noise variance estimated by our algorithm. Guangyong Chen, Pheng-Ann Heng |
ICCV | 3 |
| 2015 | Automatic Fetal Ultrasound Standard Plane Detection Using Knowledge Transferred Recurrent Neural Networks
Hao Chen 0011, Qi Dou 0001, Dong Ni 0001, Jie-Zhi Cheng, Harry Qin, Shengli Li 0001, Pheng-Ann Heng |
MICCAI (1) | 7 |
| 2015 | Automatic Localization and Identification of Vertebrae in Spine CT via a Joint Learning Model with Deep Neural Networks
Hao Chen 0011, Chiyao Shen, Harry Qin, Dong Ni 0001, Lin Shi 0001, Jack Chun-Yiu Cheng, Pheng-Ann Heng |
MICCAI (1) | 7 |
| 2015 | Morphology-preserving smoothing on polygonized isosurfaces of inhomogeneous binary volumes
Mingqiang Wei, Lei Zhu 0003, Jinze Yu 0001, Jun Wang 0039, Wai-Man Pang, Jianhuang Wu, Harry Qin, Pheng-Ann Heng |
Comput. Aided Des. | 8 |
| 2015 | Standard Plane Localization in Fetal Ultrasound via Domain Transferred Deep Neural NetworksabstractAutomatic localization of the standard plane containing complicated anatomical structures in ultrasound (US) videos remains a challenging problem. In this paper, we present a learning-based approach to locate the fetal abdominal standard plane (FASP) in US videos by constructing a domain transferred deep convolutional neural network (CNN). Compared with previous works based on low-level features, our approach is able to represent the complicated appearance of the FASP and hence achieve better classification performance. More importantly, in order to reduce the overfitting problem caused by the small amount of training samples, we propose a transfer learning strategy, which transfers the knowledge in the low layers of a base CNN trained from a large database of natural images to our task-specific CNN. Extensive experiments demonstrate that our approach outperforms the state-of-the-art method for the FASP localization as well as the CNN only trained on the limited US training samples. The proposed approach can be easily extended to other similar medical image computing problems, which often suffer from the insufficient training samples when exploiting the deep CNN to represent high-level features. Hao Chen 0011, Dong Ni 0001, Harry Qin, Shengli Li 0001, Xin Yang 0009, Tianfu Wang 0001, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 7 |
| 2015 | Closure-aware sketch simplificationabstractIn this paper, we propose a novel approach to simplify sketch drawings. The core problem is how to group sketchy strokes meaningfully, and this depends on how humans understand the sketches. The existing methods mainly rely on thresholding low-level geometric properties among the strokes, such as proximity, continuity and parallelism. However, it is not uncommon to have strokes with equal geometric properties but different semantics. The lack of semantic analysis will lead to the inability in differentiating the above semantically different scenarios. In this paper, we point out that, due to the gestalt phenomenon of closure , the grouping of strokes is actually highly influenced by the interpretation of regions. On the other hand, the interpretation of regions is also influenced by the interpretation of strokes since regions are formed and depicted by strokes. This is actually a chicken-or-the-egg dilemma and we solve it by an iterative cyclic refinement approach. Once the formed stroke groups are stabilized, we can simplify the sketchy strokes by replacing each stroke group with a smooth curve. We evaluate our method on a wide range of different sketch styles and semantically meaningful simplification results can be obtained in all test cases. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2015 | Bi-Normal Filtering for Mesh DenoisingabstractMost mesh denoising techniques utilize only either the facet normal field or the vertex normal field of a mesh surface. The two normal fields, though contain some redundant geometry information of the same model, can provide additional information that the other field lacks. Thus, considering only one normal field is likely to overlook some geometric features. In this paper, we take advantage of the piecewise consistent property of the two normal fields and propose an effective framework in which they are filtered and integrated using a novel method to guide the denoising process. Our key observation is that, decomposing the inconsistent field at challenging regions into multiple piecewise consistent fields makes the two fields complementary to each other and produces better results. Our approach consists of three steps: vertex classification, bi-normal filtering, and vertex position update. The classification step allows us to filter the two fields on a piecewise smooth surface rather than a surface that is smooth everywhere. Based on the piecewise consistence of the two normal fields, we filtered them using a piecewise smooth region clustering strategy. To benefit from the bi-normal filtering, we design a quadratic optimization algorithm for vertex position update. Experimental results on synthetic and real data show that our algorithm achieves higher quality results than current approaches on surfaces with multifarious geometric features and irregular surface sampling. Mingqiang Wei, Jinze Yu 0001, Wai-Man Pang, Jun Wang 0039, Harry Qin, Ligang Liu 0001, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2014 | Visualization of Needle Access Pathway and a Five-DoF EvaluationabstractIt is a common practice nowadays to plan needle access pathways to the volumetric organs before performing surgeries. An enormous amount of needle access planning systems has been proposed in recent years. Recent works mainly focus on the system usability or target accessibility. Visualization of the planned access pathways has drawn little attention and its effect on insertion quality is left unattended. We aim to address this problem by introducing an all-round evaluation framework that links up with human motions and computer graphics. Our evaluation framework provides an objective and quantitative analysis of the illustrativeness of the needle access pathway visualization techniques to an extent of five degrees of freedom. Our experimental results show that the visualization method adopted greatly influences insertion accuracy. Based on this finding, we propose a new visualization technique that intuitively conveys placement and orientation information. We also show that our method better conveys pathway orientation and thus enables a higher quality of insertion. Wing-Yin Chan, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | Coarse-to-Fine Normal Filtering for Feature-Preserving Mesh Denoising Based on Isotropic SubneighborhoodsabstractAbstract State‐of‐theart normal filters usually denoise each face normal using its entire anisotropic neighborhood. However, enforcing these filters indiscriminately on the anisotropic neighborhood will lead to feature blurring, especially in challenging regions with shallow features. We develop a novel mesh denoising framework which can effectively preserve features with various sizes. Our idea is inspired by the observation that the underlying surface of a noisy mesh is piecewise smooth. In this regard, it is more desirable that we denoise each face normal within its piecewise smooth region (we call such a region as an isotropic subneighborhood) instead of using the anisotropic neighborhood. To achieve this, we first classify mesh faces into several types using a face normal tensor voting and then perform a normal filter to obtain a denoised coarse normal field. Based on the results of normal classification and the denoised coarse normal field, we segment the anisotropic neighborhood of every feature face into a number of isotropic subneighborhoods via local spectral clustering. Thus face normal filtering can be performed again on the isotropic subneighborhoods and produce a more accurate normal field. Extensive tests on various models demonstrate that our method can achieve better performance than state‐of‐theart normal filters, especially in challenging regions with features. Lei Zhu 0003, Mingqiang Wei, Jinze Yu 0001, Weiming Wang 0002, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 6 |
| 2013 | Parallel structure-aware halftoning
Huisi Wu, Tien-Tsin Wong, Pheng-Ann Heng |
Multim. Tools Appl. | 3 |
| 2012 | A virtual surgical simulator for mandibular angle reduction based on patient specific dataabstractIn our work, a virtual reality-based surgical simulator for the mandibular angle reduction was designed and implemented on CUDA-based platform. High-fidelity visual and haptic feedbacks between the surgical instruments and the bone material are provided to enhance the perception in a realistic virtual surgical environment. Impulse-based dynamics haptic model was employed to simulate the contact forces generated on the high-speed instruments, including the reciprocating saw and the round burr. The validity of the simulated contact forces was verified by comparing against the actual force data measured through the constructed mechanical platform. An empirical study based on the patient specified data was conducted to evaluate the ability of the proposed system in training surgeons with various experiences. The results confirm the validity of our simulator. Qiong Wang 0001, Hui Chen 0020, Wen Wu 0001, Hai-yang Jin, Pheng-Ann Heng |
VR | 5 |
| 2012 | A Serious Game for Learning Ultrasound-Guided Needle Placement SkillsabstractUltrasound-guided needle placement is a key step in a lot of radiological intervention procedures such as biopsy, local anesthesia and fluid drainage. To help training future intervention radiologists, we develop a serious game to teach the skills involved. We introduce novel techniques for realistic simulation and integrate game elements for active and effective learning. This game is designed in the context of needle placement training based on the some essential characteristics of serious games. Training scenarios are interactively generated via a block-based construction scheme. A novel example-based texture synthesis technique is proposed to simulate corresponding ultrasound images. Game levels are defined based on the difficulties of the generated scenarios. Interactive recommendation of desirable insertion paths is provided during the training as an adaptation mechanism. We also develop a fast physics-based approach to reproduce the shadowing effect of needles in ultrasound images. Game elements such as time-attack tasks, hints and performance evaluation tools are also integrated in our system. Extensive experiments are performed to validate its feasibility for training. Wing-Yin Chan, Harry Qin, Yim-Pan Chui, Pheng-Ann Heng |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2012 | Real-Time Mandibular Angle Reduction Surgical Simulation With Haptic RenderingabstractMandibular angle reduction is a popular and efficient procedure widely used to alter the facial contour. The primary surgical instruments, the reciprocating saw and the round burr, employed in the surgery have a common feature: operating at a high-speed. Generally, inexperienced surgeons need a long-time practice to learn how to minimize the risks caused by the uncontrolled contacts and cutting motions in manipulation of instruments with high-speed reciprocation or rotation. A virtual reality-based surgical simulator for the mandibular angle reduction was designed and implemented on a CUDA-based platform in this paper. High-fidelity visual and haptic feedbacks are provided to enhance the perception in a realistic virtual surgical environment. The impulse-based haptic models were employed to simulate the contact forces and torques on the instruments. It provides convincing haptic sensation for surgeons to control the instruments under different reciprocation or rotation velocities. The real-time methods for bone removal and reconstruction during surgical procedures have been proposed to support realistic visual feedbacks. The simulated contact forces were verified by comparing against the actual force data measured through the constructed mechanical platform. An empirical study based on the patient-specific data was conducted to evaluate the ability of the proposed system in training surgeons with various experiences. The results confirm the validity of our simulator. Qiong Wang 0001, Hui Chen 0020, Wen Wu 0001, Hai-yang Jin, Pheng-Ann Heng |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2012 | Binocular tone mappingabstractBy extending from monocular displays to binocular displays, one additional image domain is introduced. Existing binocular display systems only utilize this additional image domain for stereopsis. Our human vision is not only able to fuse two displaced images, but also two images with difference in detail, contrast and luminance, up to a certain limit. This phenomenon is known as binocular single vision . Humans can perceive more visual content via binocular fusion than just a linear blending of two views. In this paper, we make a first attempt in computer graphics to utilize this human vision phenomenon, and propose a binocular tone mapping framework. The proposed framework generates a binocular low-dynamic range (LDR) image pair that preserves more human-perceivable visual content than a single LDR image using the additional image domain. Given a tone-mapped LDR image (left, without loss of generality), our framework optimally synthesizes its counterpart (right) in the image pair from the same source HDR image. The two LDR images are different, so that they can aggregately present more human-perceivable visual richness than a single arbitrary LDR image, without triggering visual discomfort . To achieve this goal, a novel binocular viewing comfort predictor (BVCP) is also proposed to prevent such visual discomfort. The design of BVCP is based on the findings in vision science. Through our user studies, we demonstrate the increase of human-perceivable visual richness and the effectiveness of the proposed BVCP in conservatively predicting the visual discomfort threshold of human observers. Xuan S. Yang, Linling Zhang, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 4 |
| 2011 | WYSIWYF: exploring and annotating volume data with a tangible handheld deviceabstractVisual exploration of volume data often requires the user to manipulate the orientation and position of a slicing plane in order to observe, annotate or measure its internal structures. Such operations, with its many degrees of freedom in 3D space, map poorly into interaction modalities afforded by mouse-keyboard interfaces or flat multi-touch displays alone. We addressed this problem using a what-you-see-is-what-you-feel (WYSIWYF) approach, which integrates the natural user interface of a multi-touch wall display with the untethered physical dexterity provided by a handheld device with multi-touch and 3D-tilt sensing capabilities. A slicing plane can be directly and intuitively manipulated at any desired position within the displayed volume data using a commonly available mobile device such as the iPod touch. 2D image slices can be transferred wirelessly to this small touch screen device, where a novel fast fat finger annotation technique (F3AT) is proposed to perform accurate and speedy contour drawings. Our user studies support the efficacy of our proposed visual exploration and annotation interaction designs. Peng Song 0001, Wooi-Boon Goh, Chi-Wing Fu, Pheng-Ann Heng |
CHI | 5 |
| 2010 | Structure-preserving multiscale vessel enhancing diffusion filterabstractEnhancement of vessels in medical images is still an unsolved problem. Multiscale approaches were proposed to improve the vessel enhancement effect based on the structure size and image resolution. Vessel enhancing diffusion (VED) filter is one of the multiscale approaches, which was based on the scale space theory. VED performs well on enhancing vessel structures but cannot preserve complex structures such as the vessel junctions. In this paper, a structure-preserving diffusion tensor is defined in the diffusion equation, which brings a structure-preserving vessel enhancing diffusion filter. Through the multiscale framework, the proposed method enhances the vessel structures especially the complex structure such as junctions. Experimental evaluation performed on various vessel data sets demonstrated the effectiveness of the proposed method. Yiping Chen 0002, Liansheng Wang 0002, Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Xiang Li 0014 |
ICIP | 5 |
| 2010 | Adaptive total variation denoising based on difference curvature
Qiang Chen 0004, Philippe Montesinos, Quan-Sen Sun, Pheng-Ann Heng, De-Shen Xia |
Image Vis. Comput. | 4 |
| 2010 | Two-Stage Object Tracking Method Based on Kernel and Active ContourabstractThis letter presents a two-stage object tracking method by combining a region-based method and a contour-based method. First, a kernel-based method is adopted to locate the object region. Then the diffusion snake is used to evolve the object contour in order to improve the tracking precision. In the first object localization stage, the initial target position is predicted and evaluated by the Kalman filter and the Bhattacharyya coefficient, respectively. In the contour evolution stage, the active contour is evolved on the basis of an object feature image generated with the color information in the initial object region. In the process of the evolution, similarities of the target region are compared to ensure that the object contour evolves in the right way. The comparison between our method and the kernel-based method demonstrates that our method can effectively cope with the severe deformation of object contour, so the tracking precision of our method is higher. Qiang Chen 0004, Quan-Sen Sun, Pheng-Ann Heng, De-Shen Xia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Fast and accurate 3-D registration of HR-pQCT imagesabstractHigh-resolution peripheral quantitative computed tomography (HR-pQCT) is a new noninvasive bone imaging technology that generates high-resolution 3-D images for quantitatively analysis of the bone microarchitecture in human. To enable quantitative evaluation of bone changes, either bone gain or loss, accurate alignment between the baseline and follow-up scans of the same individual is necessary. The major difficulties in achieving efficient and automatic registration of the HR-pQCT data are the large data size, deformations in the nonskeletal structures, and the complexity of the trabecular bone geometry. In this paper, we propose an automatic surface-based approach for fast and accurate registration of the HR-pQCT data, where the rigid registration is applied on the surfaces of the bony structures extracted from the grayscale HR-pQCT. The bony structure segmentation is performed via an automatic method that can adaptively determine the thresholds for separating the bony structure from the background and nonskeletal tissues. Experimental results performed on ten pairs of baseline and follow-up wrist scans of five adolescents and five elderly patients with osteoporosis showed the advantage of the proposed method in the high degree of automation, while the resultant parameters describing bone mineral density and trabecular architecture after registration were comparable with the outputs of the scanner's software. This automatic and accurate matching procedure may contribute to the clinical application and research of HR-pQCT. Lin Shi 0001, Defeng Wang, Vivian W. Y. Hung, Benson H. Y. Yeung, James F. Griffith, Winnie Chiu-Wing Chu, Pheng-Ann Heng, Jack Chun-Yiu Cheng |
IEEE Trans. Inf. Technol. Biomed. | 7 |
| 2010 | Resizing by symmetry-summarizationabstractImage resizing can be achieved more effectively if we have a better understanding of the image semantics. In this paper, we analyze the translational symmetry , which exists in many real-world images. By detecting the symmetric lattice in an image, we can summarize , instead of only distorting or cropping, the image content. This opens a new space for image resizing that allows us to manipulate, not only image pixels, but also the semantic cells in the lattice. As a general image contains both symmetry & non-symmetry regions and their natures are different, we propose to resize symmetry regions by summarization and non-symmetry region by warping. The difference in resizing strategy induces discontinuity at their shared boundary. We demonstrate how to reduce the artifact. To achieve practical resizing applications for general images, we developed a fast symmetry detection method that can detect multiple disjoint symmetry regions, even when the lattices are curved and perspectively viewed. Comparisons to state-of-the-art resizing techniques and a user study were conducted to validate the proposed method. Convincing visual results are shown to demonstrate its effectiveness. Huisi Wu, Yu-Shuen Wang, Kun-Chuan Feng, Tien-Tsin Wong, Tong-Yee Lee, Pheng-Ann Heng |
ACM Trans. Graph. | 6 |
| 2010 | Combined X-ray and facial videos for phoneme-level articulator dynamics
Hui Chen 0020, Wenxi Liu, Pheng-Ann Heng |
Vis. Comput. | 4 |
| 2009 | A Fast and Flexible Sorting Algorithm with CUDA
Shifu Chen, Harry Qin, Yongming Xie, Junping Zhao, Pheng-Ann Heng |
ICA3PP | 5 |
| 2009 | Solving the CLM Problem by Discrete-Time Linear Threshold Recurrent Neural Networks
Lei Zhang 0005, Pheng-Ann Heng, Zhang Yi 0001 |
ICANN (1) | 2 |
| 2009 | A Physically-Based Modeling and Simulation Framework for Facial AnimationabstractRealistic facial animation is important in many graphics applications, like animated feature films and computer games, to enrich human computer interaction. In this paper, we propose a physically-based facial animation approach employing knowledge from the anatomy and biomechanics of human facial muscles. First, the 3D face mesh is generated automatically by commercial software and the facial skin is represented by a nonlinear mass-spring system which simulates realistic elastic dynamics of human dermis. Then, a structure skull is attached to fit the face mesh. A set of anatomically consistent facial muscles are incorporated to model the forces deforming the face mesh. Finally, we extend Waters' muscle model to improve the combination of multiple muscle actions and to generate realistic expression. Experiments show that our method is superior to the traditional geometric model and can achieve comparative results with the commercial software FaceGen. Weiming Wang 0002, Xiaoqi Yan, Yongming Xie, Harry Qin, Wai-Man Pang, Pheng-Ann Heng |
ICIG | 6 |
| 2009 | Realistic grass withering simulation using time-varying texelsabstractGrass is one of the crucial elements in representing real nature scenes. Various methods have been proposed for grass simulation, for example, Boulanger and his colleagues render realistic grass in real-time with dynamic lighting, shadows and animations. Till now, few approaches have investigated the realistic simulation of withering grassland which is still a big challenge in computer graphics, due to the great complexity of both geometry and withering mechanism. In our work, a novel concept called Time-Varying Texels (TVT) is proposed to extend the static texel[Kajiya and Kay 1989] and make it time-serialized. In a grass TVT structure, time-dependent texels are arranged hierarchically in a Time-Space Partitioning (TSP) tree, so that the grass withering process can be efficiently approximated, with acceptable spatial and temporal errors. In this way we facilitate LOD rendering. Traditional texel structures could hardly undertake physical based calculation in each grass blade, so we introduce a point based structure (PBS) to increase the flexibility. Dynamic processes such as geometric deformation and material transformation can be achieved on PBS during the whole withering procedure. Additionally, by clustering the pre-computed TVT samples in low densities using a mingling algorithm, we obtain high-density grass TVT so as to significantly improve the efficiency of TVT generation. Shaohui Jiao, Pheng-Ann Heng, Enhua Wu |
SIGGRAPH ASIA Sketches | 2 |
| 2009 | Automatic detection of breast cancers in mammograms using structured support vector machines
Defeng Wang, Lin Shi 0001, Pheng-Ann Heng |
Neurocomputing | 3 |
| 2009 | Some multistability properties of bidirectional associative memory recurrent neural networks with unsaturating piecewise linear transfer functions
Lei Zhang 0005, Zhang Yi 0001, Pheng-Ann Heng |
Neurocomputing | 4 |
| 2009 | GL4D: A GPU-based Architecture for Interactive 4D VisualizationabstractThis paper describes GL4D, an interactive system for visualizing 2-manifolds and 3-manifolds embedded in four Euclidean dimensions and illuminated by 4D light sources. It is a tetrahedron-based rendering pipeline that projects geometry into volume images, an exact parallel to the conventional triangle-based rendering pipeline for 3D graphics. Novel features include GPU-based algorithms for real-time 4D occlusion handling and transparency compositing; we thus enable a previously impossible level of quality and interactivity for exploring lit 4D objects. The 4D tetrahedrons are stored in GPU memory as vertex buffer objects, and the vertex shader is used to perform per-vertex 4D modelview transformations and 4D-to-3D projection. The geometry shader extension is utilized to slice the projected tetrahedrons and rasterize the slices into individual 2D layers of voxel fragments. Finally, the fragment shader performs per-voxel operations such as lighting and alpha blending with previously computed layers. We account for 4D voxel occlusion along the 4D-to-3D projection ray by supporting a multi-pass back-to-front fragment composition along the projection ray; to accomplish this, we exploit a new adaptation of the dual depth peeling technique to produce correct volume image data and to simultaneously render the resulting volume data using 3D transfer functions into the final 2D image. Previous CPU implementations of the rendering of 4D-embedded 3-manifolds could not perform either the 4D depth-buffered projection or manipulation of the volume-rendered image in real-time; in particular, the dual depth peeling algorithm is a novel GPU-based solution to the real-time 4D depth-buffering problem. GL4D is implemented as an integrated OpenGL-style API library, so that the underlying shader operations are as transparent as possible to the user. Alan Chu, Chi-Wing Fu, Andrew J. Hanson, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | Generating massive high-quality random numbers using GPUabstractPseudo-random number generators (PRNG) have been intensively used in many stochastic algorithms in artificial intelligence, computer graphics and other scientific computing. However, the current commodity GPU design does not facilitate the efficient implementation of high-quality PRNGs that require high-precision integer arithmetics and bitwise operations. In this paper, we propose a framework to generate a high-quality PRNG shader for all kinds of GPUs. We adopt the cellular automata (CA) PRNG to facilitate high speed and parallel random number generation. The configuration of the CA PRNG is completed automatically by optimizing an objective function that accounts for quality of generated random sequences. To visually evaluate the result, we apply the best PRNG shader to photon mapping. Timing statistics show that our GPU parallelized PRNG is much faster than a pure CPU implementation. Wai-Man Pang, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Volumetric Ultrasound Panorama Based on 3D SIFT
Dong Ni 0001, Yingge Qu, Xuan S. Yang, Yim-Pan Chui, Tien-Tsin Wong, Simon S. M. Ho, Pheng-Ann Heng |
MICCAI (2) | 7 |
| 2008 | A double-threshold image binarization method based on edge detector
Qiang Chen 0004, Quan-Sen Sun, Pheng-Ann Heng, De-Shen Xia |
Pattern Recognit. | 3 |
| 2008 | Discriminative analysis of skull morphology in adolescent idiopathic scoliosis patients: Comparative study with normal controls
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
Pattern Recognit. | 3 |
| 2008 | Erratum to: "Discriminative analysis of skull morphology in adolescent idiopathic scoliosis patients: Comparative study with normal controls" [Pattern Recognition 41 (9) 2800-2811]
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
Pattern Recognit. | 3 |
| 2008 | Shape matching and modeling using skeletal context
Jun Xie 0001, Pheng-Ann Heng, Mubarak Shah |
Pattern Recognit. | 2 |
| 2008 | Parametric active contours for object tracking based on matching degree image of object contour points
Qiang Chen 0004, Quan-Sen Sun, Pheng-Ann Heng, De-Shen Xia |
Pattern Recognit. Lett. | 3 |
| 2008 | Image Diffusion Using Saliency Bilateral FilterabstractImage diffusion can smooth away noise and small-scale structures while retaining important features, thus improving the performances for many image processing algorithms. In this paper, we present a novel diffusion algorithm for which the filtering kernels vary according to the perceptual saliency of boundaries. The effectiveness of the proposed approach is validated by experiments on various medical images. Jun Xie 0001, Pheng-Ann Heng, Mubarak Shah |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2008 | Intrinsic colorizationabstractIn this paper, we present an example-based colorization technique robust to illumination differences between grayscale target and color reference images. To achieve this goal, our method performs color transfer in an illumination-independent domain that is relatively free of shadows and highlights. It first recovers an illumination-independent intrinsic reflectance image of the target scene from multiple color references obtained by web search. The reference images from the web search may be taken from different vantage points, under different illumination conditions, and with different cameras. Grayscale versions of these reference images are then used in decomposing the grayscale target image into its intrinsic reflectance and illumination components. We transfer color from the color reflectance image to the grayscale reflectance image, and obtain the final result by relighting with the illumination component of the target image. We demonstrate via several examples that our method generates results with excellent color consistency. Xiaopei Liu, Yingge Qu, Tien-Tsin Wong, Stephen Lin 0001, Andrew Chi-Sing Leung, Pheng-Ann Heng |
ACM Trans. Graph. | 7 |
| 2008 | Structure-aware halftoningabstractThis paper presents an optimization-based halftoning technique that preserves the structure and tone similarities between the original and the halftone images. By optimizing an objective function consisting of both the structure and the tone metrics, the generated halftone images preserve visually sensitive texture details as well as the local tone. It possesses the blue-noise property and does not introduce annoying patterns. Unlike the existing edge-enhancement halftoning, the proposed method does not suffer from the deficiencies of edge detector. Our method is tested on various types of images. In multiple experiments and the user study, our method consistently obtains the best scores among all tested methods. Wai-Man Pang, Yingge Qu, Tien-Tsin Wong, Daniel Cohen-Or, Pheng-Ann Heng |
ACM Trans. Graph. | 5 |
| 2008 | Richness-preserving manga screeningabstractDue to the tediousness and labor intensive cost, some manga artists have already employed computer-assisted methods for converting color photographs to manga backgrounds. However, existing bitonal image generation methods usually produce unsatisfactory uniform screening results that are not consistent with traditional mangas, in which the artist employs a rich set of screens. In this paper, we propose a novel method for generating bitonal manga backgrounds from color photographs. Our goal is to preserve the visual richness in the original photograph by utilizing not only screen density, but also the variety of screen patterns. To achieve the goal, we select screens for different regions in order to preserve the tone similarity, texture similarity, and chromaticity distinguishability. The multi-dimensional scaling technique is employed in such a color-to-pattern matching for maintaining pattern dissimilarity of the screens. Users can control the mapping by a few parameters and interactively fine-tune the result. Several results are presented to demonstrate the effectiveness and convenience of the proposed method. Yingge Qu, Wai-Man Pang, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 4 |
| 2008 | Fuzzified Choquet Integral With a Fuzzy-Valued Integrand and Its Application on Temperature PredictionabstractIn this paper, the original Choquet integral is generalized as a Fuzzified Choquet Integral with a Fuzzy-valued Integrand (FCIFI), which supports a fuzzy-valued integrand and an integration result. The calculation of the FCIFI is established on the Choquet integral with an interval-valued integrand (CIII). The definitions, properties, and calculation algorithms of the CIII and the FCIFI are discussed and proposed in this paper. As a specific application scheme, we designed a CIII regression model for the regression problems involving interval-valued data. This CIII regression model has a self-learning ability through a double genetic algorithm. Finally, a daily temperature predictor based on the CIII regression model is discussed, where a series of experiments is implemented to validate the performance of the predictor by real weather records from the Hong Kong Observatory. Rong Yang 0006, Zhenyuan Wang, Pheng-Ann Heng, Kwong-Sak Leung |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2007 | Robust Metric Reconstruction from Challenging Video SequencesabstractAlthough camera self-calibration and metric reconstruction have been extensively studied during the past decades, automatic metric reconstruction from long video sequences with varying focal length is still very challenging. Several critical issues in practical implementations are not adequately addressed. For example, how to select the initial frames for initializing the projective reconstruction? What criteria should be used? How to handle the large zooming problem? How to choose an appropriate moment for upgrading the projective reconstruction to a metric one? This paper gives a careful investigation of all these issues. Practical and effective approaches are proposed. In particular, we show that existing image-based distance is not an adequate measurement for selecting the initial frames. We propose a novel measurement to take into account the zoom degree, the self-calibration quality, as well as image-based distance. We then introduce a new strategy to decide when to upgrade the projective reconstruction to a metric one. Finally, to alleviate the heavy computational cost in the bundle adjustment, a local on-demand approach is proposed. Our method is also extensively compared with the state-of-the-art commercial software to evidence its robustness and stability. Guofeng Zhang 0001, Xueying Qin, Wei Hua 0002, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao |
CVPR | 5 |
| 2007 | Moving Object Extraction with a Hand-held CameraabstractThis paper presents a new method to detect and accurately extract the moving object from a video sequence taken by a hand-held camera. In order to extract the high quality moving foreground, previous approaches usually assume that the background is static or through only planar-perspective transformation. In our method, based on the robust motion estimation, we are capable of handling challenging videos where the background contains complex depth and the camera undergoes unknown motions. We propose the appearance and structure consistency constraint in 3D warping to robustly model the background, which greatly improves the foreground separation even on the object boundary. The estimated dense motion field and the bi- layer segmentation result are iteratively refined where continuous and discrete optimizations are alternatively used. Experimental results of high quality moving object extraction from challenging videos demonstrate the effectiveness of our method. Guofeng Zhang 0001, Jiaya Jia, Tien-Tsin Wong, Pheng-Ann Heng, Hujun Bao |
ICCV | 5 |
| 2007 | Orthopedics Surgery Trainer with PPU-Accelerated Blood and Tissue Simulation
Wai-Man Pang, Harry Qin, Yim-Pan Chui, Tien-Tsin Wong, Kwok-Sui Leung, Pheng-Ann Heng |
MICCAI (2) | 6 |
| 2007 | Landmark Correspondence Optimization for Coupled Surfaces
Lin Shi 0001, Defeng Wang, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
MICCAI (2) | 3 |
| 2007 | Dynamic touch-enabled virtual palpationabstractAbstract Palpation is an important method of feeling with hands during a physical examination, in which the doctor presses on the surface of the patient body to feel the organs or tissues underneath. In current surgical simulation systems, the lack of an effective sense of touch is still a major problem. In this paper, a dynamic touch‐enabled virtual palpation model is proposed. The palpation force sensing between the index finger and virtual tissues is simulated through a body‐based haptic interaction model. Both contact and frictional forces are evaluated based on Hertz's contact theory, and the press distribution within the contact area is also specified. The non‐linear viscoelastic behavior of typical tissues is mimicked via a volumetric tetrahedral mass‐spring system. Reaction during the palpation is restricted to a local area to highly reduce the order of the dynamic equation of the entire system to guarantee a fast working rate. Mechanical tests have been performed to evaluate the palpation force perception and the realistic behavior of typical human tissues. Copyright © 2007 John Wiley & Sons, Ltd. Hui Chen 0011, Wen Wu 0001, Hanqiu Sun, Pheng-Ann Heng |
Comput. Animat. Virtual Worlds | 4 |
| 2007 | Fast and active texture segmentation based on orientation and local variance
Qiang Chen 0004, Pheng-Ann Heng, De-Shen Xia |
J. Vis. Commun. Image Represent. | 3 |
| 2007 | Ellipsoidal support vector clustering for functional MRI analysis
Defeng Wang, Lin Shi 0001, Daniel S. Yeung, Eric C. C. Tsang, Pheng-Ann Heng |
Pattern Recognit. | 5 |
| 2007 | Classification of Heterogeneous Fuzzy Data by Choquet Integral With Fuzzy-Valued IntegrandabstractAs a fuzzification of the Choquet integral, the defuzzified choquet integral with fuzzy-valued integrand (DCIFI) takes a fuzzy-valued integrand and gives a crisp-valued integration result. In this paper, the DCIFI acts as a projection to project high-dimensional heterogeneous fuzzy data to one-dimensional crisp data to handle the classification problems involving different data forms, such as crisp data, interval values, fuzzy numbers, and linguistic variables, simultaneously. The nonadditivity of the signed fuzzy measure applied in the DCIFI can represent the interaction among the measurements of features towards the discrimination of classes. Values of the signed fuzzy measure in the DCIFI are considered to be unknown parameters which should be learned before the classifier is used to classify new data. We have implemented a genetic algorithm (GA)-based adaptive classifier-learning algorithm to optimally learn the signed fuzzy measure values and the classified boundaries simultaneously. The performance of our algorithm has been tested both on synthetic and real data. The experimental results are satisfactory and outperform those of existing methods, such as the fuzzy decision trees and the fuzzy-neuro networks. Rong Yang 0006, Zhenyuan Wang, Pheng-Ann Heng, Kwong-Sak Leung |
IEEE Trans. Fuzzy Syst. | 3 |
| 2007 | A Computational Framework for Approximating Boundary Surfaces in 3-D Biomedical ImagesabstractWe propose a new method for detecting and approximating the boundary surfaces in three-dimensional (3-D) biomedical images. Using this method, each boundary surface in the original 3-D image is normalized as a zero-value isosurface of a new 3-D image transformed from the original 3-D image. A novel computational framework is proposed to perform such an image transformation. According to this framework, we first detect boundary surfaces from the original 3-D image and compute discrete samplings of the boundary surfaces. Based on these discrete samplings, a new 3-D image is constructed for each boundary surface such that the boundary surface can be well approximated by a zero-value isosurface in the new 3-D image. In this way, the complex problem of reconstructing boundary surfaces in the original 3-D image is converted into a task to extract a zero-value isosurface from the new 3-D image. The proposed technique is not only capable of adequately reconstructing complex boundary surfaces in 3-D biomedical images, but it also overcomes vital limitations encountered by the isosurface-extracting method when the method is used to reconstruct boundary surfaces from 3-D images. The performances and advantages of the proposed computational framework are illustrated by many examples from different 3-D biomedical images. Lisheng Wang, Jing Bai 0001, Pheng-Ann Heng, Xuan S. Yang |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2007 | Discrete Wavelet Transform on Consumer-Level Graphics HardwareabstractDiscrete wavelet transform (DWT) has been heavily studied and developed in various scientific and engineering fields. Its multiresolution and locality nature facilitates applications requiring progressiveness and capturing high-frequency details. However, when dealing with enormous data volume, its performance may drastically reduce. On the other hand, with the recent advances in consumer-level graphics hardware, personal computers nowadays usually equip with a graphics processing unit (GPU) based graphics accelerator which offers SIMD-based parallel processing power. This paper presents a SIMD algorithm that performs the convolution-based DWT completely on a GPU, which brings us significant performance gain on a normal PC without extra cost. Although the forward and inverse wavelet transforms are mathematically different, the proposed algorithm unifies them to an almost identical process that can be efficiently implemented on GPU. Different wavelet kernels and boundary extension schemes can be easily incorporated by simply modifying input parameters. To demonstrate its applicability and performance, we apply it to wavelet-based geometric design, stylized image processing, texture-illuminance decoupling, and JPEG2000 image encoding Tien-Tsin Wong, Andrew Chi-Sing Leung, Pheng-Ann Heng, Jianqing Wang |
IEEE Trans. Multim. | 3 |
| 2007 | Tileable BTFabstractThis paper presents a modular framework to efficiently apply the bidirectional texture functions (BTF) onto object surfaces. The basic building blocks are the BTF tiles. By constructing one set of BTF tiles, a wide variety of objects can be textured seamlessly without re-synthesizing the BTF. The proposed framework nicely decouples the surface appearance from the geometry. With this appearance-geometry decoupling, one can build a library of BTF tile sets to instantaneously dress and render various objects under variable lighting and viewing conditions. The core of our framework is a novel method for synthesizing seamless high-dimensional BTF tiles, that are difficult for existing synthesis techniques. Its key is to shorten the cutting paths and broaden the choices of samples so as to increase the chance of synthesizing seamless BTF tiles. To tackle the enormous data, the tile synthesis process is performed in compressed domain. This not just allows the handling of large BTF data during the synthesis, but also facilitates compact storage of the BTF in GPU memory during the rendering. Man-Kang Leung, Wai-Man Pang, Chi-Wing Fu, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2006 | An Indirect and Efficient Approach for Solving Uncorrelated Optimal Discriminant Vectors
Quan-Sen Sun, Zhong Jin, Pheng-Ann Heng, De-Shen Xia |
ICIC (2) | 3 |
| 2006 | Morphometric Analysis for Pathological Abnormality Detection in the Skull Vaults of Adolescent Idiopathic Scoliosis Girls
Lin Shi 0001, Pheng-Ann Heng, Tien-Tsin Wong, Winnie Chiu-Wing Chu, Benson H. Y. Yeung, Jack Chun-Yiu Cheng |
MICCAI (1) | 2 |
| 2006 | Image Diffusion Using Saliency Bilateral Filter
Jun Xie 0001, Pheng-Ann Heng, Simon S. M. Ho, Mubarak Shah |
MICCAI (2) | 2 |
| 2006 | Real-valued Choquet integrals with fuzzy-valued integrand
Zhenyuan Wang, Rong Yang 0006, Pheng-Ann Heng, Kwong-Sak Leung |
Fuzzy Sets Syst. | 3 |
| 2006 | GPU-friendly warped display for scope-maintained video surveillance
Tien-Tsin Wong, Pheng-Ann Heng |
Multim. Syst. | 3 |
| 2006 | Shape Statistics Variational Approach for the Outer Contour Segmentation of Left Ventricle MR ImagesabstractSegmentation of left ventricles is one of the important research topics in cardiac magnetic resonance (MR) imaging. The segmentation precision influences the authenticity of ventricular motion reconstruction. In left ventricle MR images, the weak and broken boundary increases the difficulty of segmenting the outer contour precisely. In this paper, we present an improved shape statistics variational approach for the outer contour segmentation of left ventricle MR images. We use the Mumford-Shah model in an object feature space and incorporate the shape statistics and an edge image to the variational framework. The introduction of shape statistics can improve the segmentation with broken boundaries. The edge image can enhance the weak boundary and thus improve the segmentation precision. The generation of the object feature image, which has homogenous "intensities" in the left ventricle, facilitates the application of the Mumford-Shah model. A comparison of mean absolute distance analysis between different contours generated with our algorithm and that generated by hand demonstrated that our method can achieve a higher segmentation precision and a better stability than various approaches. It is a semiautomatic way for the segmentation of the outer contour of the left ventricle in clinical applications. Qiang Chen 0004, Ze Ming Zhou, Pheng-Ann Heng, De-Shen Xia |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2006 | Intelligent Inferencing and Haptic Simulation for Chinese Acupuncture Learning and TrainingabstractThis paper presents an intelligent virtual environment for Chinese acupuncture learning and training using state-of-the-art virtual reality technology. It is the first step toward developing a comprehensive virtual human model for studying Chinese medicine. Students can learn and practice acupuncture in the proposed 3-D interactive virtual environment that supports a force feedback interface for needle insertion. Thus, students not only "see" but also "touch" the virtual patient. With high performance computers, highly informative and flexible visualization of acupuncture points of various related meridian and collateral can be highlighted to guide the students during training. A computer-based expert system using our newly proposed intelligent fuzzy petri net is designed and implemented to train the students to treat different diseases using acupuncture. Such an intelligent virtual reality system can provide an interesting and effective learning environment for Chinese acupuncture. Pheng-Ann Heng, Tien-Tsin Wong, Rong Yang 0006, Yim-Pan Chui, Yongming Xie, Kwong-Sak Leung, P.-C. Leung |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | Manga colorizationabstractThis paper proposes a novel colorization technique that propagates color over regions exhibiting pattern-continuity as well as intensity-continuity. The proposed method works effectively on colorizing black-and-white manga which contains intensive amount of strokes, hatching, halftoning and screening. Such fine details and discontinuities in intensity introduce many difficulties to intensity-based colorization methods. Once the user scribbles on the drawing, a local, statistical based pattern feature obtained with Gabor wavelet filters is applied to measure the pattern-continuity. The boundary is then propagated by the level set method that monitors the pattern-continuity. Regions with open boundaries or multiple disjointed regions with similar patterns can be sensibly segmented by a single scribble. With the segmented regions, various colorization techniques can be applied to replace colors, colorize with stroke preservation, or even convert pattern to shading. Several results are shown to demonstrate the effectiveness and convenience of the proposed method. Yingge Qu, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2006 | Deringing cartoons by image analogiesabstractIn this article, we propose a novel method to reduce ringing artifacts in BDCT-encoded cartoon images using image analogies. The quantization procedure of BDCT compression (such as JPEG and MPEG) introduces annoying visual artifacts. Our main focus is on the removal of ringing artifacts that is seldom addressed by existing methods. In the proposed method, the contaminated image is modeled as a Markov random field (MRF). We “learn” the behavior of contamination by extracting massive numbers of artifact patterns from a training set, and organizing them using tree-structured vector quantization (TSVQ). Instead of postfiltering the input contaminated image, we synthesize an artifact-reduced image. Our method is noniterative and hence, can remove artifacts within a very short period of time. We show that substantial improvement is achieved using the proposed method in terms of visual quality and statistics. Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 3 |
| 2005 | Face Recognition Based on Generalized Canonical Correlation Analysis
Quan-Sen Sun, Pheng-Ann Heng, Zhong Jin, De-Shen Xia |
ICIC (2) | 2 |
| 2005 | Support Vector Clustering for Brain Activation Detection
Defeng Wang, Lin Shi 0001, Daniel S. Yeung, Pheng-Ann Heng, Tien-Tsin Wong, Eric C. C. Tsang |
MICCAI | 4 |
| 2005 | Shape Modeling Using Automatic Landmarking
Jun Xie 0001, Pheng-Ann Heng |
MICCAI (2) | 2 |
| 2005 | Fuzzy numbers and fuzzification of the Choquet integral
Rong Yang 0006, Zhenyuan Wang, Pheng-Ann Heng, Kwong-Sak Leung |
Fuzzy Sets Syst. | 3 |
| 2005 | A theorem on the generalized canonical projective vectors
Quan-Sen Sun, Zhengdong Liu, Pheng-Ann Heng, De-Shen Xia |
Pattern Recognit. | 3 |
| 2005 | A new method of feature fusion and its application in image recognition
Quan-Sen Sun, Sheng-Gen Zeng, Pheng-Ann Heng, De-Shen Xia |
Pattern Recognit. | 4 |
| 2005 | LV shape and motion: B-spline-based deformable model and sequential motion decompositionabstractIn this paper, we extend a previous work by J. Park and propose a uniform framework to reconstruct left ventricle (LV) geometry/motion from tagged MR images. In our work, the LV is modeled as a generalized prolate spheroid, and its motion is decomposed into four components-global translation, polar radial/z-axis compression, twisting, and bending. By formulating model parameters as tensor products of B-splines, we develop efficient algorithms to quickly reconstruct LV geometry/motion from extracted boundary contours and tracked planar tags. Experiments on both synthesized and in vivo data are also reported. Guo Luo, Pheng-Ann Heng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2005 | A system for real-time panorama generation and display in tele-immersive applicationsabstractWide field-of-view (FOV) is necessary for many industrial applications, such as air traffic control, large vehicle driving and navigation. Unfortunately, the supporting structure/frame in most systems usually blocks part of the view, results in "blind spot" and raises the risk to the pilot. In this paper, we introduce a video-based tele-immersive system, called the immersive cockpit. It captures live videos from the working site and recreates an immersive environment at the remote site where the pilot situates. It immerses the pilot at the remote site with a panoramic view of the environment, and hence improves interactivity and safety. The design goals of our system are real-time, live, low-cost, and scalable. We stitch multiple video streams captured from ordinary charged couple device cameras to generate a panoramic video. To avoid being blocked by the supporting frame, we allow a flexible placement of cameras. This approach trades the accuracy of the generated panoramic image for a larger FOV. To reduce the computation, parameters for stitching are determined once during the system initialization. The panoramic video is presented on an immersive display which covers the FOV of the viewer. We discuss how to correctly present the panoramic video on this nonplanar immersive display screen by sweet spot relocation. We also present the result and the performance evaluation of the system. Wai-Kwan Tang, Tien-Tsin Wong, Pheng-Ann Heng |
IEEE Trans. Multim. | 3 |
| 2005 | An improved scheme of an interactive finite element model for 3D soft-tissue cutting and deformation
Wen Wu 0001, Pheng-Ann Heng |
Vis. Comput. | 2 |
| 2004 | Feature fusion method based on canonical correlation analysis and handwritten character recognitionabstractA new feature extraction method, based on feature fusion, according to the idea of canonical correlation analysis (CCA), is proposed in this paper. A framework of CCA used in pattern recognition is described. The overall process comprises: extracting two groups of feature vectors with the same pattern; establishing the correlation criterion function between the two groups of feature vectors, and extract their canonical correlation features in order to form effective discriminant vectors for recognition. The inherent essence of this method used in recognition is theoretically analyzed. This method uses correlation features between two groups of feature vectors as effective discriminant information, so it not only is suitable for information fusion, but also eliminates redundant information within features, a new way for classification is proposed. Experimental results of our method applying on Concordia University CENPARMI handwritten numeral database has shown that our recognition rate is higher than that of the algorithm adopting single feature or the existing fusion algorithm. Quan-Sen Sun, Sheng-Gen Zeng, Pheng-Ann Heng, De-Shen Xia |
ICARCV | 3 |
| 2004 | A real-time Cantonese text-to-audiovisual speech synthesizerabstractThis paper describes the design and development of a Cantonese TTVS synthesizer, which can generate highly natural synthetic speech that is precisely time-synchronized with a real-time 3D face rendering. Our Cantonese TTVS synthesizer utilizes a homegrown Cantonese syllable-based concatenative text-to-speech system named CU VOCAL. This paper describes the extension of CU VOCAL to output syllable labels and durations that correspond to the output acoustic wave file. The syllables are decomposed and their initials/finals are mapped to the nearest IPA symbols that correspond to static viseme models. We have authored sixteen static viseme models together with two emotion-based face models. In order to achieve 3D face rendering, we have designed and implemented a blending technique that computes the linear combinations of the static face models to effect smooth transitions in between models. We demonstrate that this design and implementation of a TTVS synthesizer can achieve real-time performance in generation. Jianqing Wang, Ka-Ho Wong, Pheng-Ann Heng, Helen M. Meng, Tien-Tsin Wong |
ICASSP (1) | 3 |
| 2004 | Segmentation of Left Ventricle via Level Set Method Based on Enriched Speed Term
Yingge Qu, Qiang Chen 0004, Pheng-Ann Heng, Tien-Tsin Wong |
MICCAI (1) | 3 |
| 2004 | Digital photo similarity analysis in frequency domain and photo album compressionabstractWith the increasing popularity of digital camera, organizing and managing the large collection of digital photos effectively are therefore required. In this paper, we study the techniques of photo album sorting, clustering and compression in DCT frequency domain without having to decompress JPEG photos into spatial domain firstly. We utilize the first several non-zero DCT coefficients to build our feature set and calculate the energy histograms in frequency domain directly. We then calculate the similarity distances of every two photos, and perform photo album sorting and adaptive clustering algorithms to group the most similar photos together. We further compress those clustered photos by a MPEG-like algorithm with variable IBP frames and adaptive search windows. Our methods provide a compact and reasonable format for people to store and transmit their large number of digital photos. Experiments prove that our algorithm is efficient and effective for digital photo processing. Tien-Tsin Wong, Pheng-Ann Heng |
MUM | 3 |
| 2004 | An efficient and scalable deformable model for virtual reality-based medical applications
Kup-Sze Choi, Hanqiu Sun, Pheng-Ann Heng |
Artif. Intell. Medicine | 3 |
| 2004 | Deformable simulation using force propagation model with finite element optimization
Kup-Sze Choi, Hanqiu Sun, Pheng-Ann Heng |
Comput. Graph. | 3 |
| 2004 | Attitude dead reckoning in a collaborative virtual environment using cumulative polynomial extrapolation of quaternionsabstractAbstract In this paper, we propose a new attitude dead‐reckoning paradigm in a collaborative virtual environment (CVE). We derive a general polynomial construction scheme of attitude trajectory that uses a number of previous packets in order to extrapolate the future trajectory of objects by quaternion representation. The scheme allows consecutive attitudes received from the network to propagate in a smooth manner. This cumulative trajectory construction scheme helps in developing our adaptive prediction and convergence mechanism of the overall estimation paradigm. Aiming to facilitate the remote rendering of objects, our proposed cumulative polynomial extrapolation technique provides a robust management of rotational states of objects and at the same time reduces the bandwidth consumption compared with the traditional method. By devising a quaternion‐based attitude estimation paradigm, the complete predictive management for shared states is built; this permits us to estimate rotational trajectory of objects as well as camera views. The proposed algorithm can perform remote rendering accurately in the presence of network latency. Experiments are carried out to illustrate the effectiveness of our proposed algorithm. Copyright © 2004 John Wiley & Sons, Ltd. Yim-Pan Chui, Pheng-Ann Heng |
Concurr. Pract. Exp. | 2 |
| 2004 | A hybrid condensed finite element model with GPU acceleration for interactive 3D soft tissue cuttingabstractAbstract To meet the requirement of computer‐aided medical operations, apart from the real‐time deformation, it is also necessary in the design to simulate the tissue cutting and suturing in a surgery simulation. In this paper, we present a model on topology change and deformation of soft tissue, referred to as the hybrid condensed finite element model, based on the volumetric finite element method. The most important advantage of our model is its ability to achieve an interactive frame rate for the topology change in surgical simulation on standard PC platform. This is achieved through two innovations. One is to apply the condensation technique, by fully calculating the volumetric deformation in the operation part while only calculating the surface nodes in the non‐operation part. Secondly, the major calculation work in the Conjugate Gradient solver for cutting and deformation is migrated from the CPU to the contemporary GPU to promote the calculation. Test examples have been given to show the feasibility and efficiency of the model. Copyright © 2004 John Wiley & Sons, Ltd. Wen Wu 0001, Pheng-Ann Heng |
Comput. Animat. Virtual Worlds | 2 |