EDBT 2026 Demo / reviewers in the wild / expert
Terrence Chen
dblp:51/4242
· DBLP profile ↗
82ranked-venue papers
9as first author
42since 2021 · last 2026
0009-0001-6697-7098ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 42 · 5 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 2 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Memorization-Utilization Gap: Near-Lossless Context Compression via Reinforcement LearningabstractDespite recent progress in context compression, we identify a fundamental memorizationutilization gap where models can compress context with near-perfect fidelity yet fail to effectively utilize these compressed representations for downstream tasks.We address this with a holistic training paradigm spanning pretraining, instruction tuning, and reinforcement learning, built upon an average pooling compression.Our key innovation uses outcomebased RL to enable implicit expansion: the model learns to adaptively unfold task-relevant details during generation, interleaving reconstruction with reasoning.We achieve nearlossless 16× context compression (≈5.3× decoder sequence-length reduction in our current implementation) across 7B and 32B models, recovering over 98% of full-context QA performance and outperforming prior methods by 11 points.Our 32B model demonstrates strong out-of-distribution and length generalization, robustly scaling to 120k-token contexts despite training on no more than 4k tokens, matching full-context performance on NIAH, Long-Bench v2, and multi-hop reasoning.We verify the implicit expansion behavior in experiments. Yujan Ting, Terrence Chen, Weijing Huang |
ACL (1) | 3 |
| 2026 | Leveraging Diffusion Model and Image Foundation Model for Improved Correspondence Matching in Coronary AngiographyabstractAccurate correspondence matching in coronary angiography images is crucial for reconstructing 3D coronary artery structures, which is essential for precise diagnosis and treatment planning of coronary artery disease (CAD). Traditional matching methods for natural images often fail to generalize to X-ray images due to inherent differences such as lack of texture, lower contrast, and overlapping structures, compounded by insufficient training data. To address these challenges, we propose a novel pipeline that generates realistic paired coronary angiography images using a diffusion model conditioned on 2D projections of 3D reconstructed meshes from Coronary Computed Tomography Angiography (CCTA), providing high-quality synthetic data for training. Additionally, we employ large-scale image foundation models to guide feature aggregation, enhancing correspondence matching accuracy by focusing on semantically relevant regions and keypoints. Our approach demonstrates superior matching performance on synthetic datasets and effectively generalizes to real-world datasets, offering a practical solution for this task. Furthermore, our work investigates the efficacy of different foundation models in correspondence matching, providing novel insights into leveraging advanced image foundation models for medical imaging applications. Lin Zhao 0004, Yikang Liu 0001, Xiao Chen 0013, Eric Z. Chen, Terrence Chen, Shanhui Sun |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac FluoroscopyabstractThe accurate segmentation of guidewires in interventional cardiac fluoroscopy videos is crucial for computer-aided navigation tasks. Although deep learning methods have demonstrated high accuracy and robustness in wire segmentation, they require substantial annotated datasets for generalizability, underscoring the need for extensive labeled data to enhance model performance. To address this challenge, we propose the Segmentation-guided Frame-consistency Video Diffusion Model (SF-VD) to generate large collections of labeled fluoroscopy videos, augmenting the training data for wire segmentation networks. SF-VD leverages videos with limited annotations by independently modeling scene distribution and motion distribution. It first samples the scene distribution by generating 2D fluoroscopy images with wires positioned according to a specified input mask, and then samples the motion distribution by progressively generating subsequent frames, ensuring frame-to-frame coherence through a frame-consistency strategy. A segmentation-guided mechanism further refines the process by adjusting wire contrast, ensuring a diverse range of visibility in the synthesized image. Evaluation on a fluoroscopy dataset confirms the superior quality of the generated videos and shows significant improvements in guidewire segmentation. Shaoyan Pan, Yikang Liu 0001, Lin Zhao 0004, Eric Z. Chen, Xiao Chen 0013, Terrence Chen, Shanhui Sun |
AAAI | 6 |
| 2025 | Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingabstractTemporal awareness is essential for video large language models (LLMs) to understand and reason about events within long videos, enabling applications like dense video captioning and temporal video grounding in a unified system. However, the scarcity of long videos with detailed captions and precise temporal annotations limits their temporal awareness. In this paper, we propose Seq2Time, a data-oriented training paradigm that leverages sequences of images and short video clips to enhance temporal awareness in long videos. By converting sequence positions into temporal annotations, we transform large-scale image and clip captioning datasets into sequences that mimic the temporal structure of long videos, enabling self-supervised training with abundant time-sensitive data. To enable sequence-to-time knowledge transfer, we introduce a novel time representation that unifies positional information across image sequences, clip sequences, and long videos. Experiments demonstrate the effectiveness of our method, achieving a 27.6% improvement in F1 score and 44.8% in CIDEr on the YouCook2 benchmark and a 14.7% increase in recall on the Charades-STA benchmark compared to the baseline. Project available at: https://seq2time.github.io/ Andong Deng, Zhongpai Gao, Anwesa Choudhuri, Benjamin Planche, Meng Zheng 0002, Bin Wang 0068, Terrence Chen, Chen Chen 0001, Ziyan Wu 0001 |
CVPR | 7 |
| 2025 | CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single ImageabstractReconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free environment. Thus, when encountering in-the-wild occluded images, these algorithms produce multiview inconsistent and fragmented reconstructions. Additionally, most algorithms for monocular 3D human reconstruction leverage geometric priors such as SMPL annotations for training and inference, which are extremely challenging to acquire in real-world applications. To address these limitations, we propose CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-ConsistEncy from a Single Image, a novel pipeline designed to reconstruct occlusion-resilient 3D humans with multiview consistency from a single occluded image, without requiring either ground-truth geometric prior annotations or 3D supervision. Specifically, CHROME leverages a multiview diffusion model to first synthesize occlusion-free human images from the occluded input, compatible with off-the-shelf pose control to explicitly enforce cross-view consistency during synthesis. A 3D reconstruction model is then trained to predict a set of 3D Gaussians conditioned on both the occluded input and synthesized views, aligning cross-view details to produce a cohesive and accurate 3D representation. CHROME achieves significant improvements in terms of both novel view synthesis (upto 3 db PSNR) and geometric reconstruction under challenging conditions. Arindam Dutta, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Anwesa Choudhuri, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 6 |
| 2025 | 7DGS: Unified Spatial-Temporal-Angular Gaussian SplattingabstractReal-time rendering of dynamic scenes with view-dependent effects remains a fundamental challenge in computer graphics. While recent advances in Gaussian Splatting have shown promising results separately handling dynamic scenes (4DGS) and view-dependent effects (6DGS), no existing method unifies these capabilities while maintaining real-time performance. We present 7D Gaussian Splatting (7DGS), a unified framework representing scene elements as seven-dimensional Gaussians spanning position (3D), time (1D), and viewing direction (3D). Our key contribution is an efficient conditional slicing mechanism that transforms 7D Gaussians into view- and time-conditioned 3D Gaussians, maintaining compatibility with existing 3D Gaussian Splatting pipelines while enabling joint optimization. Experiments demonstrate that 7DGS outperforms prior methods by up to 7.36 dB in PSNR while achieving real-time rendering (401 FPS) on challenging dynamic scenes with complex view-dependent effects. The project page is: https://gaozhongpai.github.io/7dgs/. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Ziyan Wu 0001 |
ICCV | 5 |
| 2025 | Order-aware Interactive SegmentationabstractInteractive segmentation aims to accurately segment target objects with minimal user interactions. However, current methods often fail to accurately separate target objects from the background, due to a limited understanding of order, the relative depth between objects in a scene. To address this issue, we propose OIS: order-aware interactive segmentation, where we explicitly encode the relative depth between objects into order maps. We introduce a novel order-aware attention, where the order maps seamlessly guide the user interactions (in the form of clicks) to attend to the image features. We further present an object-aware attention module to incorporate a strong object-level understanding to better differentiate objects with similar order. Our approach allows both dense and sparse integration of user clicks, enhancing both accuracy and efficiency as compared to prior works. Experimental results demonstrate that OIS achieves state-of-the-art performance, improving mIoU after one click by 7.61 on the HQSeg44K dataset and 1.32 on the DAVIS dataset as compared to the previous state-of-the-art SegNext, while also doubling inference speed compared to current leading methods. Bin Wang 0068, Anwesa Choudhuri, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Andong Deng, Qin Liu 0008, Terrence Chen, Ulas Bagci, Ziyan Wu 0001 |
ICLR | 8 |
| 2025 | 6DGS: Enhanced Direction-Aware Gaussian Splatting for Volumetric RenderingabstractNovel view synthesis has advanced significantly with the development of neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS). However, achieving high quality without compromising real-time rendering remains challenging, particularly for physically-based rendering using ray/path tracing with view-dependent effects. Recently, N-dimensional Gaussians (N-DG) introduced a 6D spatial-angular representation to better incorporate view-dependent effects, but the Gaussian representation and control scheme are sub-optimal. In this paper, we revisit 6D Gaussians and introduce 6D Gaussian Splatting (6DGS), which enhances color and opacity representations and leverages the additional directional information in the 6D space for optimized Gaussian control. Our approach is fully compatible with the 3DGS framework and significantly improves real-time radiance field rendering by better modeling view-dependent effects and fine details. Experiments demonstrate that 6DGS significantly outperforms 3DGS and N-DG, achieving up to a 15.73 dB improvement in PSNR with a reduction of 66.5\% Gaussian points compared to 3DGS. The project page is: https://gaozhongpai.github.io/6dgs/. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Ziyan Wu 0001 |
ICLR | 5 |
| 2025 | 3D Vision-Language Gaussian SplattingabstractRecent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D vision-language Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin. Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng 0002, Anwesa Choudhuri, Terrence Chen, Chen Chen 0001, Ziyan Wu 0001 |
ICLR | 6 |
| 2025 | PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis
Anwesa Choudhuri, Zhongpai Gao, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
MICCAI (11) | 5 |
| 2025 | Adapting Vision Foundation Models for Real-Time Ultrasound Image Segmentation
Eric Z. Chen, Lin Zhao 0004, Xiao Chen 0013, Yikang Liu 0001, Boris Maihe, James S. Duncan, Terrence Chen, Shanhui Sun |
MICCAI (5) | 8 |
| 2025 | Automated Patient Positioning with Learned 3D Hand Gestures
Zhongpai Gao, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
WACV | 5 |
| 2025 | Retrieval-Augmented Few-Shot Medical Image Segmentation With Foundation ModelsabstractMedical image segmentation is crucial for clinical decision-making, but the scarcity of annotated data presents significant challenges. Few-shot segmentation (FSS) methods show promise but often require training on the target domain and struggle to generalize across different modalities. Similarly, adapting foundation models such as the segment anything model (SAM) for medical imaging has limitations, including the need for fine-tuning and domain-specific adaptation. To address these issues, we propose a novel method that adapts DINOv2 and SAM 2 for retrieval-augmented few-shot medical image segmentation. Our approach uses DINOv2's feature as query to retrieve similar samples from limited annotated data, which are then encoded as memories and stored in memory bank. With the memory attention mechanism of SAM 2, the model leverages these memories as conditions to generate accurate segmentation of the target image. We evaluated our framework on three medical image segmentation tasks, demonstrating superior performance and generalizability across various modalities without the need for any retraining or fine-tuning. Overall, this method offers a practical and effective solution for few-shot medical image segmentation and holds significant potential as a valuable annotation tool in clinical applications. Lin Zhao 0004, Xiao Chen 0013, Eric Z. Chen, Yikang Liu 0001, Terrence Chen, Shanhui Sun |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Disguise without Disruption: Utility-Preserving Face De-identificationabstractWith the rise of cameras and smart sensors, humanity generates an exponential amount of data. This valuable information, including underrepresented cases like AI in medical settings, can fuel new deep-learning tools. However, data scientists must prioritize ensuring privacy for individuals in these untapped datasets, especially for images or videos with faces, which are prime targets for identification methods. Proposed solutions to de-identify such images often compromise non-identifying facial attributes relevant to downstream tasks. In this paper, we introduce Disguise, a novel algorithm that seamlessly de-identifies facial images while ensuring the usability of the modified data. Unlike previous approaches, our solution is firmly grounded in the domains of differential privacy and ensemble-learning research. Our method involves extracting and substituting depicted identities with synthetic ones, generated using variational mechanisms to maximize obfuscation and non-invertibility. Additionally, we leverage supervision from a mixture-of-experts to disentangle and preserve other utility attributes. We extensively evaluate our method using multiple datasets, demonstrating a higher de-identification rate and superior consistency compared to prior approaches in various downstream tasks. Zikui Cai, Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Terrence Chen, Muhammad Salman Asif, Ziyan Wu 0001 |
AAAI | 5 |
| 2024 | Implicit Modeling of Non-rigid Objects with Cross-Category SignalsabstractDeep implicit functions (DIFs) have emerged as a potent and articulate means of representing 3D shapes. However, methods modeling object categories or non-rigid entities have mainly focused on single-object scenarios. In this work, we propose MODIF, a multi-object deep implicit function that jointly learns the deformation fields and instance-specific latent codes for multiple objects at once. Our emphasis is on non-rigid, non-interpenetrating entities such as organs. To effectively capture the interrelation between these entities and ensure precise, collision-free representations, our approach facilitates signaling between category-specific fields to adequately rectify shapes. We also introduce novel inter-object supervision: an attraction-repulsion loss is formulated to refine contact regions between objects. Our approach is demonstrated on various medical benchmarks, involving modeling different groups of intricate anatomical entities. Experimental results illustrate that our model can proficiently learn the shape representation of each organ and their relations to others, to the point that shapes missing from unseen instances can be consistently recovered by our method. Finally, MODIF can also propagate semantic information throughout the population via accurate point correspondences. Yuchun Liu, Benjamin Planche, Meng Zheng 0002, Zhongpai Gao, Pierre Sibut-Bourde, Fan Yang 0035, Terrence Chen, Ziyan Wu 0001 |
AAAI | 7 |
| 2024 | DaReNeRF: Direction-aware Representation for Dynamic ScenesabstractAddressing the intricate challenge of modeling and re-rendering dynamic scenes, most recent approaches have sought to simplify these complexities using plane-based explicit representations, overcoming the slow training time issues associated with methods like Neural Radiance Fields (NeRF) and implicit representations. However, the straight-forward decomposition of 4D dynamic scenes into multiple 2D plane-based representations proves insufficient for re-rendering high-fidelity scenes with complex motions. In response, we present a novel direction-aware representation (DaRe) approach that captures scene dynamics from six different directions. This learned representation under-goes an inverse dual-tree complex wavelet transformation (DTCWT) to recover plane-based information. DaReNeRF computes features for each space-time point by fusing vectors from these recovered planes. Combining DaReNeRF with a tiny MLP for color regression and leveraging volume rendering in training yield state-of-the-art performance in novel view synthesis for complex dynamic scenes. Notably, to address redundancy introduced by the six real and six imag-inary direction-aware wavelet coefficients, we introduce a trainable masking approach, mitigating storage issues without significant performance decline. Moreover, DaReNeRF maintains a 2 × reduction in training time compared to prior art while delivering superior performance. Ange Lou, Benjamin Planche, Zhongpai Gao, Tianyu Luan, Hao Ding 0021, Terrence Chen, Jack H. Noble, Ziyan Wu 0001 |
CVPR | 7 |
| 2024 | Divide and Fuse: Body Part Mesh Recovery from Partially Visible Human Images
Tianyu Luan, Zhongpai Gao, Luyuan Xie, Hao Ding 0021, Benjamin Planche, Meng Zheng 0002, Ange Lou, Terrence Chen, Junsong Yuan 0001, Ziyan Wu 0001 |
ECCV (24) | 9 |
| 2024 | PBADet: A One-Stage Anchor-Free Approach for Part-Body AssociationabstractThe detection of human parts (e.g., hands, face) and their correct association with individuals is an essential task, e.g., for ubiquitous human-machine interfaces and action recognition. Traditional methods often employ multi-stage processes, rely on cumbersome anchor-based systems, or do not scale well to larger part sets. This paper presents PBADet, a novel one-stage, anchor-free approach for part-body association detection. Building upon the anchor-free object representation across multi-scale feature maps, we introduce a singular part-to-body center offset that effectively encapsulates the relationship between parts and their parent bodies. Our design is inherently versatile and capable of managing multiple parts-to-body associations without compromising on detection accuracy or robustness. Comprehensive experiments on various datasets underscore the efficacy of our approach, which not only outperforms existing state-of-the-art techniques but also offers a more streamlined and efficient solution to the part-body association challenge. Zhongpai Gao, Huayi Zhou 0001, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Ziyan Wu 0001 |
ICLR | 6 |
| 2024 | Auxiliary Input in Training: Incorporating Catheter Features into Deep Learning Models for ECG-Free Dynamic Coronary Roadmapping
Yikang Liu 0001, Lin Zhao 0004, Eric Z. Chen, Xiao Chen 0013, Terrence Chen, Shanhui Sun |
MICCAI (6) | 5 |
| 2024 | Few-Shot 3D Volumetric Segmentation with Multi-surrogate Fusion
Meng Zheng 0002, Benjamin Planche, Zhongpai Gao, Terrence Chen, Richard J. Radke, Ziyan Wu 0001 |
MICCAI (9) | 4 |
| 2024 | DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume RenderingabstractDigitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extremely computationally intensity. Analytical DRR renderers are much more efficient, but at the price of ignoring anisotropic X-ray image formation phenomena such as Compton scattering. We propose a novel approach that balances realistic physics-inspired X-ray simulation with efficient, differentiable DRR generation using 3D Gaussian splatting (3DGS). Our direction-disentangled 3DGS (DDGS) method decomposes the radiosity contribution into isotropic and direction-dependent components, able to approximate complex anisotropic interactions without complex runtime simulations. Additionally, we adapt the 3DGS initialization to account for tomography data properties, enhancing accuracy and efficiency. Our method outperforms state-of-the-art techniques in image accuracy and inference speed, demonstrating its potential for intraoperative applications and inverse problems like pose registration. Zhongpai Gao, Benjamin Planche, Meng Zheng 0002, Terrence Chen, Ziyan Wu 0001 |
NeurIPS | 5 |
| 2023 | Progressive Multi-View Human Mesh Recovery with Self-SupervisionabstractTo date, little attention has been given to multi-view 3D human mesh estimation, despite real-life applicability (e.g., motion capture, sport analysis) and robustness to single-view ambiguities. Existing solutions typically suffer from poor generalization performance to new settings, largely due to the limited diversity of image/3D-mesh pairs in multi-view training data. To address this shortcoming, people have explored the use of synthetic images. But besides the usual impact of visual gap between rendered and target data, synthetic-data-driven multi-view estimators also suffer from overfitting to the camera viewpoint distribution sampled during training which usually differs from real-world distributions. Tackling both challenges, we propose a novel simulation-based training pipeline for multi-view human mesh recovery, which (a) relies on intermediate 2D representations which are more robust to synthetic-to-real domain gap; (b) leverages learnable calibration and triangulation to adapt to more diversified camera setups; and (c) progressively aggregates multi-view information in a canonical 3D space to remove ambiguities in 2D representations. Through extensive benchmarking, we demonstrate the superiority of the proposed solution especially for unseen in-the-wild scenarios. Liangchen Song, Meng Zheng 0002, Benjamin Planche, Terrence Chen, Junsong Yuan 0001, David S. Doermann, Ziyan Wu 0001 |
AAAI | 5 |
| 2023 | Computationally Efficient 3D MRI Reconstruction with Adaptive MLP
Eric Z. Chen, Xiao Chen 0013, Yikang Liu 0001, Terrence Chen, Shanhui Sun |
MICCAI (10) | 5 |
| 2023 | Federated Learning With Privacy-Preserving Ensemble Attention DistillationabstractFederated Learning (FL) is a machine learning paradigm where many local nodes collaboratively train a central model while keeping the training data decentralized. This is particularly relevant for clinical applications since patient data are usually not allowed to be transferred out of medical facilities, leading to the need for FL. Existing FL methods typically share model parameters or employ co-distillation to address the issue of unbalanced data distribution. However, they also require numerous rounds of synchronized communication and, more importantly, suffer from a privacy leakage risk. We propose a privacy-preserving FL framework leveraging unlabeled public data for one-way offline knowledge distillation in this work. The central model is learned from local knowledge via ensemble attention distillation. Our technique uses decentralized and heterogeneous local data like existing FL approaches, but more importantly, it significantly reduces the risk of privacy leakage. We demonstrate that our method achieves very competitive performance with more robust privacy preservation based on extensive experiments on image classification, segmentation, and reconstruction tasks. Liangchen Song, Rishi Vedula, Meng Zheng 0002, Benjamin Planche, Arun Innanje, Terrence Chen, Junsong Yuan 0001, David S. Doermann, Ziyan Wu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2022 | Preserving Privacy in Federated Learning with Ensemble Cross-Domain Knowledge DistillationabstractFederated Learning (FL) is a machine learning paradigm where local nodes collaboratively train a central model while the training data remains decentralized. Existing FL methods typically share model parameters or employ co-distillation to address the issue of unbalanced data distribution. However, they suffer from communication bottlenecks. More importantly, they risk privacy leakage risk. In this work, we develop a privacy preserving and communication efficient method in a FL framework with one-shot offline knowledge distillation using unlabeled, cross-domain, non-sensitive public data. We propose a quantized and noisy ensemble of local predictions from completely trained local models for stronger privacy guarantees without sacrificing accuracy. Based on extensive experiments on image classification and text classification tasks, we show that our method outperforms baseline FL algorithms with superior performance in both accuracy and data privacy preservation. Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, David S. Doermann, Arun Innanje |
AAAI | 5 |
| 2022 | SMPL-A: Modeling Person-Specific Deformable AnatomyabstractA variety of diagnostic and therapeutic protocols rely on locating in vivo target anatomical structures, which can be obtained from medical scans. However, organs move and deform as the patient changes his/her pose. In order to obtain accurate target location information, clinicians have to either conduct frequent intraoperative scans, resulting in higher exposition of patients to radiations, or adopt proxy procedures (e.g., creating and using custom molds to keep patients in the exact same pose during both preoperative organ scanning and subsequent treatment. Such custom proxy methods are typically sub-optimal, constraining the clinicians and costing precious time and money to the patients. To the best of our knowledge, this work is the first to present a learning-based approach to estimate the patient's internal organ deformation for arbitrary human poses in order to assist with radiotherapy and similar medical protocols. The underlying method first leverages medical scans to learn a patient-specific representation that potentially encodes the organ's shape and elastic properties. During inference, given the patient's current body pose information and the organ's representation extracted from previous medical scans, our method can estimate their current organ deformation to offer guidance to clinicians. We conduct experiments on a well-sized dataset which is augmented through real clinical data using finite element modeling. Our results suggest that pose-dependent organ deformation can be learned through a point cloud autoencoder conditioned on the parametric pose input. We hope that this work can be a starting point for future research towards closing the loop between human mesh recovery and anatomical reconstruction, with applications beyond the medical domain. Hengtao Guo, Benjamin Planche, Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Ziyan Wu 0001 |
CVPR | 5 |
| 2022 | Self-supervised Human Mesh Recovery with Cross-Representation Alignment
Meng Zheng 0002, Benjamin Planche, Srikrishna Karanam, Terrence Chen, David S. Doermann, Ziyan Wu 0001 |
ECCV (1) | 5 |
| 2022 | PseudoClick: Interactive Image Segmentation with Click Imitation
Qin Liu 0008, Meng Zheng 0002, Benjamin Planche, Srikrishna Karanam, Terrence Chen, Marc Niethammer, Ziyan Wu 0001 |
ECCV (6) | 5 |
| 2022 | PREF: Predictability Regularized Neural Motion Fields
Liangchen Song, Benjamin Planche, Meng Zheng 0002, David S. Doermann, Junsong Yuan 0001, Terrence Chen, Ziyan Wu 0001 |
ECCV (22) | 7 |
| 2022 | Visual Similarity AttentionabstractWhile there has been substantial progress in learning suitable distance metrics, these techniques in general lack transparency and decision reasoning, i.e., explaining why the input set of images is similar or dissimilar. In this work, we solve this key problem by proposing the first method to generate generic visual similarity explanations with gradient-based attention. We demonstrate that our technique is agnostic to the specific similarity model type, e.g., we show applicability to Siamese, triplet, and quadruplet models. Furthermore, we make our proposed similarity attention a principled part of the learning process, resulting in a new paradigm for learning similarity functions. We demonstrate that our learning mechanism results in more generalizable, as well as explainable, similarity models. Finally, we demonstrate the generality of our framework by means of experiments on a variety of tasks, including image retrieval, person re-identification, and low-shot semantic segmentation. Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Richard J. Radke, Ziyan Wu 0001 |
IJCAI | 3 |
| 2022 | Invertible Sharpening Network for MRI Reconstruction Enhancement
Siyuan Dong, Eric Z. Chen, Lin Zhao 0004, Xiao Chen 0013, Yikang Liu 0001, Terrence Chen, Shanhui Sun |
MICCAI (6) | 6 |
| 2022 | Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
Meng Zheng 0002, Benjamin Planche, Fan Yang 0035, Terrence Chen, Ziyan Wu 0001 |
MICCAI (8) | 5 |
| 2022 | Forecasting Human Trajectory from Scene HistoryabstractPredicting the future trajectory of a person remains a challenging problem, due to randomness and subjectivity. However, the moving patterns of human in constrained scenario typically conform to a limited number of regularities to a certain extent, because of the scenario restrictions (\eg, floor plan, roads and obstacles) and person-person or person-object interactivity. Thus, an individual person in this scenario should follow one of the regularities as well. In other words, a person's subsequent trajectory has likely been traveled by others. Based on this hypothesis, we propose to forecast a person's future trajectory by learning from the implicit scene regularities. We call the regularities, inherently derived from the past dynamics of the people and the environment in the scene, \emph{scene history}. We categorize scene history information into two types: historical group trajectories and individual-surroundings interaction. To exploit these information for trajectory prediction, we propose a novel framework Scene History Excavating Network (SHENet), where the scene history is leveraged in a simple yet effective approach. In particular, we design two components, the group trajectory bank module to extract representative group trajectories as the candidate for future path, and the cross-modal interaction module to model the interaction between individual past trajectory and its surroundings for trajectory refinement, respectively. In addition, to mitigate the uncertainty in the evaluation, caused by the aforementioned randomness and subjectivity, we propose to include smoothness into evaluation metrics. We conduct extensive evaluations to validate the efficacy of proposed framework on ETH, UCY, as well as a new, challenging benchmark dataset PAV, demonstrating superior performance compared to state-of-the-art methods. Mancheng Meng, Ziyan Wu 0001, Terrence Chen, Xiran Cai, Xiang Sean Zhou, Fan Yang 0054, Dinggang Shen |
NeurIPS | 3 |
| 2022 | Multi-motion and Appearance Self-Supervised Moving Object DetectionabstractIn this work, we consider the problem of self-supervised Moving Object Detection (MOD) in video, where no ground truth is involved in both training and inference phases. Recently, an adversarial learning framework is proposed [32] to leverage inherent temporal information for MOD. While showing great promising results, it uses single scale temporal information and may meet problems when dealing with a deformable object under multi-scale motion in different parts. Additional challenges can arise from the moving camera, which results in the failure of the motion independence hypothesis and locally independent background motion. To deal with these problems, we propose a Multimotion and Appearance Self-supervised Network (MASNet) to introduce multi-scale motion information and appearance information of scene for MOD. In particular, a moving object, especially the deformable, usually consists of moving regions at various temporal scales. Introducing multiscale motion can aggregate these regions to form a more complete detection. Appearance information can serve as another cue for MOD when the motion independence is not reliable and for removing false detection in background caused by locally independent background motion. To encode multi-scale motion and appearance, in MASNet we respectively design a multi-branch flow encoding module and an image inpainter module. The proposed modules and MASNet are extensively evaluated on the DAVIS dataset to demonstrate the effectiveness and superiority to state-of-the-art self-supervised methods. Fan Yang 0035, Srikrishna Karanam, Meng Zheng 0002, Terrence Chen, Haibin Ling, Ziyan Wu 0001 |
WACV | 4 |
| 2022 | Pyramid Convolutional RNN for MRI Image ReconstructionabstractFast and accurate MRI image reconstruction from undersampled data is crucial in clinical practice. Deep learning based reconstruction methods have shown promising advances in recent years. However, recovering fine details from undersampled data is still challenging. In this paper, we introduce a novel deep learning based method, Pyramid Convolutional RNN (PC-RNN), to reconstruct images from multiple scales. Based on the formulation of MRI reconstruction as an inverse problem, we design the PC-RNN model with three convolutional RNN (ConvRNN) modules to iteratively learn the features in multiple scales. Each ConvRNN module reconstructs images at different scales and the reconstructed images are combined by a final CNN module in a pyramid fashion. The multi-scale ConvRNN modules learn a coarse-to-fine image reconstruction. Unlike other common reconstruction methods for parallel imaging, PC-RNN does not employ coil sensitive maps for multi-coil data and directly model the multiple coils as multi-channel inputs. The coil compression technique is applied to standardize data with various coil numbers, leading to more efficient training. We evaluate our model on the fastMRI knee and brain datasets and the results show that the proposed model outperforms other methods and can recover more details. The proposed method is one of the winner solutions in the 2019 fastMRI competition. Eric Z. Chen, Puyang Wang, Xiao Chen 0013, Terrence Chen, Shanhui Sun |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Learning Local Recurrent Models for Human Mesh RecoveryabstractWe consider the problem of estimating frame-level full human body meshes given a video of a person with natural motion dynamics. While much progress in this field has been in single image-based mesh estimation, there has been a recent uptick in efforts to infer mesh dynamics from video given its role in alleviating issues such as depth ambiguity and occlusions. However, a key limitation of existing work is the assumption that all the observed motion dynamics can be modeled using one dynamical/recurrent model. While this may work well in cases with relatively simplistic dynamics, inference with in-the-wild videos presents many challenges. In particular, it is typically the case that different body parts of a person undergo different dynamics in the video, e.g., legs may move in a way that may be dynamically different from hands (e.g., a person dancing). To address these issues, we present a new method for video mesh recovery that divides the human mesh into several local parts following the standard skeletal model. We then model the dynamics of each local part with separate recurrent models, with each model conditioned appropriately based on the known kinematic structure of the human body. This results in a structure-informed local recurrent learning architecture that can be trained in an end-to-end fashion with available annotations. We conduct a variety of experiments on standard video mesh recovery benchmark datasets such as Human3.6M, MPI-INF-3DHP, and 3DPW, demonstrating the efficacy of our design of modeling local dynamics as well as establishing state-of-the-art results based on standard evaluation metrics. Runze Li 0003, Srikrishna Karanam, Terrence Chen, Bir Bhanu, Ziyan Wu 0001 |
3DV | 4 |
| 2021 | Everybody Is Unique: Towards Unbiased Human Mesh Recovery
Srikrishna Karanam, Meng Zheng 0002, Terrence Chen, Ziyan Wu 0001 |
BMVC | 4 |
| 2021 | A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsabstractDespite substantial progress in applying neural networks (NN) to a wide variety of areas, they still largely suffer from a lack of transparency and interpretability. While recent developments in explainable artificial intelligence attempt to bridge this gap (e.g., by visualizing the correlation between input pixels and final outputs), these approaches are limited to explaining low-level relationships, and crucially, do not provide insights on error correction. In this work, we propose a framework (VRX) to interpret classification NNs with intuitive structural visual concepts. Given a trained classification model, the proposed VRX extracts relevant class-specific visual concepts and organizes them using structural concept graphs (SCG) based on pairwise concept relationships. By means of knowledge distillation, we show VRX can take a step towards mimicking the reasoning process of NNs and provide logical, concept-level explanations for final model decisions. With extensive experiments, we empirically show VRX can meaningfully answer "why" and "why not" questions about the prediction, providing easy-to-understand insights about the reasoning process. We also show that these insights can potentially provide guidance on improving NN’s performance. Yunhao Ge, Zhi Xu 0013, Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Laurent Itti, Ziyan Wu 0001 |
CVPR | 6 |
| 2021 | Spatio-Temporal Representation Factorization for Video-based Person Re-IdentificationabstractDespite much recent progress in video-based person re-identification (re-ID), the current state-of-the-art still suffers from common real-world challenges such as appearance similarity among various people, occlusions, and frame misalignment. To alleviate these problems, we propose Spatio-Temporal Representation Factorization (STRF), a flexible new computational unit that can be used in conjunction with most existing 3D convolutional neural network architectures for re-ID. The key innovations of STRF over prior work include explicit pathways for learning discriminative temporal and spatial features, with each component further factorized to capture complementary person-specific appearance and motion information. Specifically, temporal factorization comprises two branches, one each for static features (e.g., the color of clothes) that do not change much over time, and dynamic features (e.g., walking patterns) that change over time. Further, spatial factorization also comprises two branches to learn both global (coarse segments) as well as local (finer segments) appearance features, with the local features particularly useful in cases of occlusion or spatial misalignment. These two factorization operations taken together result in a modular architecture for our parameter-wise light STRF unit that can be plugged in between any two 3D convolutional layers, resulting in an end-to-end learning framework. We empirically show that STRF improves performance of various existing baseline architectures while demonstrating new state-of-the-art results using standard person re-ID evaluation protocols on three benchmarks. Abhishek Aich, Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Amit K. Roy-Chowdhury, Ziyan Wu 0001 |
ICCV | 4 |
| 2021 | Ensemble Attention Distillation for Privacy-Preserving Federated LearningabstractWe consider the problem of Federated Learning (FL) where numerous decentralized computational nodes collaborate with each other to train a centralized machine learning model without explicitly sharing their local data samples. Such decentralized training naturally leads to issues of imbalanced or differing data distributions among the local models and challenges in fusing them into a central model. Existing FL methods deal with these issues by either sharing local parameters or fusing models via online distillation. However, such a design leads to multiple rounds of inter-node communication resulting in substantial band-width consumption, while also increasing the risk of data leakage and consequent privacy issues. To address these problems, we propose a new distillation-based FL frame-work that can preserve privacy by design, while also consuming substantially less network communication resources when compared to the current methods. Our framework engages in inter-node communication using only publicly available and approved datasets, thereby giving explicit privacy control to the user. To distill knowledge among the various local models, our framework involves a novel ensemble distillation algorithm that uses both final prediction as well as model attention. This algorithm explicitly considers the diversity among various local nodes while also seeking consensus among them. This results in a comprehensive technique to distill knowledge from various decentralized nodes. We demonstrate the various aspects and the associated benefits of our FL framework through extensive experiments that produce state-of-the-art results on both classification and segmentation tasks on natural and medical images. Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, David S. Doermann, Arun Innanje |
ICCV | 5 |
| 2021 | Multi-scale Neural ODEs for 3D Medical Image Registration
Junshen Xu, Eric Z. Chen, Xiao Chen 0013, Terrence Chen, Shanhui Sun |
MICCAI (4) | 4 |
| 2021 | Learning Hierarchical Attention for Weakly-Supervised Chest X-Ray Abnormality Localization and DiagnosisabstractWe consider the problem of abnormality localization for clinical applications. While deep learning has driven much recent progress in medical imaging, many clinical challenges are not fully addressed, limiting its broader usage. While recent methods report high diagnostic accuracies, physicians have concerns trusting these algorithm results for diagnostic decision-making purposes because of a general lack of algorithm decision reasoning and interpretability. One potential way to address this problem is to further train these models to localize abnormalities in addition to just classifying them. However, doing this accurately will require a large amount of disease localization annotations by clinical experts, a task that is prohibitively expensive to accomplish for most applications. In this work, we take a step towards addressing these issues by means of a new attention-driven weakly supervised algorithm comprising a hierarchical attention mining framework that unifies activation- and gradient-based visual attention in a holistic manner. Our key algorithmic innovations include the design of explicit ordinal attention constraints, enabling principled model training in a weakly-supervised fashion, while also facilitating the generation of visual-attention-driven model explanations by means of localization cues. On two large-scale chest X-ray datasets (NIH ChestX-ray14 and CheXpert), we demonstrate significant localization performance improvements over the current state of the art while also achieving competitive classification performance. Xi Ouyang, Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, Jiayu Huo, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng |
IEEE Trans. Medical Imaging | 4 |
| 2020 | FOAL: Fast Online Adaptive Learning for Cardiac Motion EstimationabstractMotion estimation of cardiac MRI videos is crucial for the evaluation of human heart anatomy and function. Recent researches show promising results with deep learning-based methods. In clinical deployment, however, they suffer dramatic performance drops due to mismatched distributions between training and testing datasets, commonly encountered in the clinical environment. On the other hand, it is arguably impossible to collect all representative datasets and to train a universal tracker before deployment. In this context, we proposed a novel fast online adaptive learning (FOAL) framework: an online gradient descent based optimizer that is optimized by a meta-learner. The meta-learner enables the online optimizer to perform a fast and robust adaptation. We evaluated our method through extensive experiments on two public clinical datasets. The results showed the superior performance of FOAL in accuracy compared to the offline-trained tracking method. On average, the FOAL took only 0.4 second per video for online optimization. Hanchao Yu, Shanhui Sun, Haichao Yu, Xiao Chen 0013, Humphrey Shi, Thomas S. Huang, Terrence Chen |
CVPR | 7 |
| 2020 | Hierarchical Kinematic Human Mesh Recovery
Georgios Georgakis, Srikrishna Karanam, Terrence Chen, Jana Kosecka, Ziyan Wu 0001 |
ECCV (17) | 4 |
| 2020 | MRI Image Reconstruction via Learning Optimization Using Neural ODEs
Eric Z. Chen, Terrence Chen, Shanhui Sun |
MICCAI (2) | 2 |
| 2020 | Robust Multi-modal 3D Patient Body Modeling
Fan Yang 0035, Georgios Georgakis, Srikrishna Karanam, Terrence Chen, Haibin Ling, Ziyan Wu 0001 |
MICCAI (3) | 5 |
| 2020 | Motion Pyramid Networks for Accurate and Efficient Cardiac Motion Estimation
Hanchao Yu, Xiao Chen 0013, Humphrey Shi, Terrence Chen, Thomas S. Huang, Shanhui Sun |
MICCAI (6) | 4 |
| 2020 | Towards Contactless Patient PositioningabstractThe ongoing COVID-19 pandemic, caused by the highly contagious SARS-CoV-2 virus, has overwhelmed healthcare systems worldwide, putting medical professionals at a high risk of getting infected themselves due to a global shortage of personal protective equipment. This has in-turn led to understaffed hospitals unable to handle new patient influx. To help alleviate these problems, we design and develop a contactless patient positioning system that can enable scanning patients in a completely remote and contactless fashion. Our key design objective is to reduce the physical contact time with a patient as much as possible, which we achieve with our contactless workflow. Our system comprises automated calibration, positioning, and multi-view synthesis components that enable patient scan without physical proximity. Our calibration routine ensures system calibration at all times and can be executed without any manual intervention. Our patient positioning routine comprises a novel robust dynamic fusion (RDF) algorithm for accurate 3D patient body modeling. With its multi-modal inference capability, RDF can be trained once and used across different applications (without re-training) having various sensor choices, a key feature to enable system deployment at scale. Our multi-view synthesizer ensures multi-view positioning visualization for the technician to verify positioning accuracy prior to initiating the patient scan. We conduct extensive experiments with publicly available and proprietary datasets to demonstrate efficacy. Our system has already been used, and had a positive impact on, hospitals and technicians on the front lines of the COVID-19 pandemic, and we expect to see its use increase substantially globally. Srikrishna Karanam, Fan Yang 0035, Terrence Chen, Ziyan Wu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Cascade Attention Machine for Occluded Landmark Detection in 2D X-Ray AngiographyabstractIn cardiac interventions, localization of guiding catheter tip in 2D fluoroscopic images is important to specify vessel branches and calibrate vessels with stenosis. While detection of guiding catheter tip is not trivial in contrast-free images due to low dose radiation as well as occlusion by other devices, it is even more challenging in contrast-filled images. As contrast-filled vessels become visible in X-ray imaging, the landmark of guiding catheter tip can often be completely occluded by the contrast medium. It is difficult even for human eyes to precisely localize the catheter tip from a single angiography image. Physicians have to rely on information before the inject of contrast medium to localize the guiding catheter tip occluded by contrast medium. Automatic landmark detection when occlusion happens is important and can significantly simplify the intervention workflow. To address this problem, we propose a novel Cascade Attention Machine (CAM) model. It borrows the idea of how human experts localize the catheter tip by first per-forming landmark detection when occlusion does not happen, then leveraging this information as prior knowledge to assist the occluded detection. Attention maps are computed from non-occluded detection to further refine the heatmaps for occluded detection to guide the inference focusing on related regions. Experiments on X-ray angiography demonstrate the promising performance compared with the state-of-the-art baselines. It shows that the CAM can capture the relation between situations with and without occlusion to achieve precise detection of occluded landmark. Liheng Zhang, Vivek K. Singh 0002, Guo-Jun Qi, Terrence Chen |
WACV | 4 |
| 2018 | Structure-Aware Shape SynthesisabstractWe propose a new procedure to guide training of a data-driven shape generative model using a structure-aware loss function. Complex 3D shapes often can be summarized using a coarsely defined structure which is consistent and robust across variety of observations. However, existing synthesis techniques do not account for structure during training, and thus often generate implausible and structurally unrealistic shapes. During training, we enforce structural constraints in order to enforce consistency and structure across the entire manifold. We propose a novel methodology for training 3D generative models that incorporates structural information into an end-to-end training pipeline. Elena Sizikova, Vivek K. Singh 0002, Jiangping Wang, Brian Teixeira, Terrence Chen, Thomas A. Funkhouser |
3DV | 5 |
| 2018 | Generating Synthetic X-Ray Images of a Person From the Surface Geometry
Brian Teixeira, Vivek K. Singh 0002, Terrence Chen, Birgi Tamersoy, Elena Sizikova, Dorin Comaniciu |
CVPR | 3 |
| 2018 | Towards Generating Personalized Volumetric Phantom from Patient's Surface Geometry
Vivek K. Singh 0002, Brian Teixeira, Birgi Tamersoy, Andreas Krauss, Terrence Chen |
MICCAI (1) | 7 |
| 2017 | DepthSynth: Real-Time Realistic Synthetic Data Generation from CAD Models for 2.5D RecognitionabstractRecent progress in computer vision has been dominated by deep neural networks trained over larges amount of labeled data. Collecting such datasets is however a tedious, often impossible task; hence a surge in approaches relying solely on synthetic data for their training. For depth images however, discrepancies with real scans still noticeably affect the end performance. We thus propose an end-to-end framework which simulates the whole mechanism of these devices, generating realistic depth data from 3D models by comprehensively modeling vital factors e.g. sensor noise, material reflectance, surface geometry. Not only does our solution cover a wider range of sensors and achieve more realistic results than previous methods, assessed through extended evaluation, but we go further by measuring the impact on the training of neural networks for various recognition tasks; demonstrating how our pipeline seamlessly integrates such architectures and consistently enhances their performance. Benjamin Planche, Ziyan Wu 0001, Shanhui Sun, Stefan Kluckner, Oliver Lehmann, Terrence Chen, Andreas Hutter, Sergey Zakharov, Harald Kosch, Jan Ernst |
3DV | 7 |
| 2017 | Multimodal Image Registration with Deep Context Reinforcement Learning
Jiangping Wang, Birgi Tamersoy, Yao-Jen Chang, Andreas Wimmer, Terrence Chen |
MICCAI (1) | 7 |
| 2017 | DARWIN: Deformable Patient Avatar Representation With Deep Image Network
Birgi Tamersoy, Yao-Jen Chang, Andreas Wimmer, Thomas O'Donnell, Terrence Chen |
MICCAI (2) | 7 |
| 2016 | Deep Decision Network for Multi-class Image ClassificationabstractIn this paper, we present a novel Deep Decision Network (DDN) that provides an alternative approach towards building an efficient deep learning network. During the learning phase, starting from the root network node, DDN automatically builds a network that splits the data into disjoint clusters of classes which would be handled by the subsequent expert networks. This results in a tree-like structured network driven by the data. The proposed method provides an insight into the data by identifying the group of classes that are hard to classify and require more attention when compared to others. DDN also has the ability to make early decisions thus making it suitable for timesensitive applications. We validate DDN on two publicly available benchmark datasets: CIFAR-10 and CIFAR-100 and it yields state-of-the-art classification performance on both the datasets. The proposed algorithm has no limitations to be applied to any generic classification problems. Venkatesh N. Murthy, Vivek K. Singh 0002, Terrence Chen, R. Manmatha, Dorin Comaniciu |
CVPR | 3 |
| 2016 | Towards Automated Ultrasound Transesophageal Echocardiography and X-Ray Fluoroscopy Fusion Using an Image-Based Co-registration Method
Shanhui Sun, Shun Miao, Tobias Heimann, Terrence Chen, Markus Kaiser 0003, Matthias John 0001, Erin Girard, Rui Liao |
MICCAI (1) | 4 |
| 2015 | BodyPrint: Pose Invariant 3D Shape Matching of Human Bodiesabstract3D human body shape matching has large potential on many real world applications, especially with the recent advances in the 3D range sensing technology. We address this problem by proposing a novel holistic human body shape descriptor called BodyPrint. To compute the bodyprint for a given body scan, we fit a deformable human body mesh and project the mesh parameters to a low-dimensional subspace which improves discriminability across different persons. Experiments are carried out on three real-world human body datasets to demonstrate that BodyPrint is robust to pose variation as well as missing information and sensor noise. It improves the matching accuracy significantly compared to conventional 3D shape matching techniques using local features. To facilitate practical applications where the shape database may grow over time, we also extend our learning framework to handle online updates. Jiangping Wang, Thomas S. Huang, Terrence Chen |
ICCV | 5 |
| 2015 | Towards an Efficient Computational Framework for Guiding Surgical Resection through Intra-operative Endo-microscopic Pathology
Shaohua Wan 0005, Shanhui Sun, Subhabrata Bhattacharya, Stefan Kluckner, Alexander Gigler, Elfriede Simon, Maximilian Fleischer, Patra Charalampaki, Terrence Chen, Ali Kamen |
MICCAI (1) | 9 |
| 2015 | Robust object tracking using semi-supervised appearance dictionary learning
Lei Zhang 0036, Wen Wu 0004, Terrence Chen, Norbert Strobel, Dorin Comaniciu |
Pattern Recognit. Lett. | 3 |
| 2014 | Estimating a Patient Surface Model for Optimizing the Medical Scanning Workflow
Yao-Jen Chang, Michael Wels, Grzegorz Soza, Terrence Chen |
MICCAI (1) | 6 |
| 2014 | Model-Guided Extraction of Coronary Vessel Structures in 2D X-Ray Angiograms
Shih-Yu Sun, Peng Wang 0005, Shanhui Sun, Terrence Chen |
MICCAI (2) | 4 |
| 2013 | Image-based Co-Registration of Angiography and Intravascular Ultrasound ImagesabstractIn image-guided cardiac interventions, X-ray imaging and intravascular ultrasound (IVUS) imaging are two often used modalities. Interventional X-ray images, including angiography and fluoroscopy, are used to assess the lumen of the coronary arteries and to monitor devices in real time. IVUS provides rich intravascular information, such as vessel wall composition, plaque, and stent expansions, but lacks spatial orientations. Since the two imaging modalities are complementary to each other, it is highly desirable to co-register the two modalities to provide a comprehensive picture of the coronaries for interventional cardiologists. In this paper, we present a solution for co-registering 2-D angiography and IVUS through image-based device tracking. The presented framework includes learning-based vessel detection and device detections, model-based tracking, and geodesic distance-based registration. The system first interactively detects the coronary branch under investigation in a reference angiography image. During the pullback of the IVUS transducers, the system acquires both ECG-triggered fluoroscopy and IVUS images, and automatically tracks the position of the medical devices in fluoroscopy. The localization of tracked IVUS transducers and guiding catheter tips is used to associate an IVUS imaging plane to a corresponding location on the vessel branch under investigation. The presented image-based solution can be conveniently integrated into existing cardiology workflow. The system is validated with a set of clinical cases, and achieves good accuracy and robustness. Peng Wang 0005, Olivier Ecabert, Terrence Chen, Michael Wels, Johannes Rieber, Martin Ostermeier, Dorin Comaniciu |
IEEE Trans. Medical Imaging | 3 |
| 2012 | Automatic Localization of Balloon Markers and Guidewire in Rotational Fluoroscopy with Application to 3D Stent Reconstruction
Yu Wang 0032, Terrence Chen, Peng Wang 0005, Christopher Rohkohl, Dorin Comaniciu |
ECCV (6) | 2 |
| 2012 | Real Time Assistance for Stent Positioning and Assessment by Self-initialized Tracking
Terrence Chen, Yu Wang 0032, Peter Durlak, Dorin Comaniciu |
MICCAI (1) | 1 |
| 2012 | Ultrasound and Fluoroscopic Images Fusion by Autonomous Ultrasound Probe Detection
Peter Mountney, Razvan Ioan Ionasec, Markus Kaiser 0003, Sina Mamaghani, Wen Wu 0004, Terrence Chen, Matthias John 0001, Jan M. Boese, Dorin Comaniciu |
MICCAI (2) | 6 |
| 2011 | Robust discriminative wire structure modeling with application to stent enhancement in fluoroscopyabstractLearning-based methods have been widely used in detecting landmarks or anatomical structures in various medical imaging applications. The performance of discriminative learning techniques has been demonstrated superior to traditional low-level filtering in robustness and scalability. Nevertheless, some structures and patterns are more difficult to be defined by such methods and complicated and ad-hoc methods still need to be used, e.g. a non-rigid and highly deformable wire structure. In this paper, we propose a novel scheme to train classifiers to detect the markers and guide wire segment anchored by markers. The classifier utilizes the markers as the end point and parameterizes the wire in-between them. The probabilities of the markers and the wire are integrated in a Bayesian framework. As a result, both the marker and the wire detection are improved by such a unified approach. Promising results are demonstrated by quantitative evaluation on 263 fluoroscopic sequences with 12495 frames. Our training scheme can further be generalized to localize longer guidewire with higher degrees of parameterization. Xiaoguang Lu, Terrence Chen, Dorin Comaniciu |
CVPR | 2 |
| 2011 | Learning-based hypothesis fusion for robust catheter tracking in 2D X-ray fluoroscopyabstractCatheter tracking has become more and more important in recent interventional applications. It provides real time navigation for the physicians and can be used to control a motion compensated fluoro overlay reference image for other means of guidance, e.g. involving a 3D anatomical model. Tracking the coronary sinus (CS) catheter is effective to compensate respiratory and cardiac motion for 3D overlay navigation to assist positioning the ablation catheter in Atrial Fibrillation (Afib) treatments. During interventions, the CS catheter performs rapid motion and non-rigid deformation due to the beating heart and respiration. In this paper, we model the CS catheter as a set of electrodes. Novelly designed hypotheses generated by a number of learning-based detectors are fused. Robust hypothesis matching through a Bayesian framework is then used to select the best hypothesis for each frame. As a result, our tracking method achieves very high robustness against challenging scenarios such as low SNR, occlusion, foreshortening, non-rigid deformation, as well as the catheter moving in and out of ROI. Quantitative evaluation has been conducted on a database of 13221 frames from 1073 sequences. Our approach obtains 0.50mm median error and 0.76mm mean error. 97.8% of evaluated data have errors less than 2.00mm. The speed of our tracking algorithm reaches 5 frames-per-second on most data sets. Our approach is not limited to the catheters inside the CS but can be extended to track other types of catheters, such as ablation catheters or circumferential mapping catheters. Wen Wu 0004, Terrence Chen, Peng Wang 0005, Shaohua Kevin Zhou, Dorin Comaniciu, Adrian Barbu, Norbert Strobel |
CVPR | 2 |
| 2011 | Combined Cardiac and Respiratory Motion Compensation for Atrial Fibrillation Ablation Procedures
Alexander Brost, Wen Wu 0004, Martin Koch 0002, Andreas Wimmer, Terrence Chen, Rui Liao, Joachim Hornegger, Norbert Strobel |
MICCAI (1) | 5 |
| 2011 | Robust and Fast Contrast Inflow Detection for 2D X-ray Fluoroscopy
Terrence Chen, Gareth Funka-Lea, Dorin Comaniciu |
MICCAI (1) | 1 |
| 2011 | Image-Based Device Tracking for the Co-registration of Angiography and Intravascular Ultrasound Images
Peng Wang 0005, Terrence Chen, Olivier Ecabert, Simone Prummer, Martin Ostermeier, Dorin Comaniciu |
MICCAI (1) | 2 |
| 2010 | Graph Based Interactive Detection of Curve Structures in 2D Fluoroscopy
Peng Wang 0005, Wei-shing Liao, Terrence Chen, Shaohua Kevin Zhou, Dorin Comaniciu |
MICCAI (3) | 3 |
| 2009 | Robust guidewire tracking in fluoroscopyabstractA guidewire is a medical device inserted into vessels during image guided interventions for balloon inflation. During interventions, the guidewire undergoes non-rigid deformation due to patients' breathing and cardiac motions, and such 3D motions are complicated when being projected onto the 2D fluoroscopy. Furthermore, in fluoroscopy there exist severe image artifacts and other wire-like structures. All these make robust guidewire tracking challenging. To address these challenges, this paper presents a probabilistic framework for robust guidewire tracking. We first introduce a semantic guidewire model that contains three parts, including a catheter tip, a guidewire tip and a guidewire body. Measurements of different parts are integrated into a Bayesian framework as measurements of a whole guidewire for robust guidewire tracking. Moreover, for each part, two types of measurements, one from learning-based detectors and the other from online appearance models, are applied and combined. A hierarchical and multi-resolution tracking scheme is then developed based on kernel-based measurement smoothing to track guidewires effectively and efficiently in a coarse-to-fine manner. The presented framework has been validated on a test set of 47 sequences, and achieves a mean tracking error of less than 2 pixels. This demonstrates the great potential of our method for clinical applications. Peng Wang 0005, Terrence Chen, Ying Zhu 0006, Wei Zhang 0018, Shaohua Kevin Zhou, Dorin Comaniciu |
CVPR | 2 |
| 2009 | Automatic ovarian follicle quantification from 3D ultrasound data using global/local context with database guided segmentationabstractIn this paper, we present a novel probabilistic framework for automatic follicle quantification in 3D ultrasound data. The proposed framework robustly estimates size and location of each individual ovarian follicle by fusing the information from both global and local context. Follicle candidates at detected locations are then segmented by a novel database guided segmentation method. To efficiently search hypothesis in a high dimensional space for multiple object detection, a clustered marginal space learning approach is introduced. Extensive evaluations conducted on 501 volumes containing 8108 follicles showed that our method is able to detect and segment ovarian follicles with high robustness and accuracy. It is also much faster than the current ultrasound manual workflow. The proposed method is able to streamline the clinical workflow and improve the accuracy of existing follicular measurements. Terrence Chen, Wei Zhang 0018, Sara Good, Shaohua Kevin Zhou, Dorin Comaniciu |
ICCV | 1 |
| 2009 | Dynamic Layer Separation for Coronary DSA and Enhancement in Fluoroscopic Sequences
Ying Zhu 0006, Simone Prummer, Peng Wang 0005, Terrence Chen, Dorin Comaniciu, Martin Ostermeier |
MICCAI (1) | 4 |
| 2007 | Optimizing Image Registration by Mutually Exclusive Scale ComponentsabstractLocal optimum has been one of the most difficult problems in image registration notwithstanding the extensive research effort that has been put into solving it. Local optimums occur when a portion of patterns in the floating image coincide with a portion of patterns in the reference image even though the two images are not entirely matched. Existing hierarchical or multi-scale methods suffer from this problem mainly because some redundant information that causes local optimums appears in multiple scales. We propose to avoid it by decomposing an image into several mutually exclusive scale components so that minimal redundant information is present. Our method is evaluated and compared with existing methods using high resolution satellite imagery where thousands of local optimum traps are hidden. We show that our method has significant improvement over existing solutions in both robustness and efficiency. Terrence Chen, Thomas S. Huang |
ICCV | 1 |
| 2006 | Scale-Driven Iterative Optimization for Brain Extraction and RegistrationabstractWe present a novel framework to automatically separate brain region from other non-brain regions in head images. The idea of the proposed method is to estimate larger scale patterns in an image and then correct the boundaries iteratively. The scale estimation is based on the recently proposed total variation (TV) regularized L^1 functional. An iterative optimization method is used to refine non-convex and acute angle boundaries. The final algorithm is able to extract large scale patterns with arbitrary shapes, which is particularly suitable for brain extraction. In order to reduce the computation overhead in 3D data, a multi-level technique is proposed to exponentially improve the speed of the brain extraction process. Based on accurate results of brain extraction, a non-rigid brain registration algorithm is proposed to improve accuracy and consistency of existing registration methods. Experimental results on real 3D brain MR images demonstrate that the proposed methods outperform existing solutions. In addition, results are provided to show that the proposed algorithm can also be used to segment large scale patterns in general images. Terrence Chen, Thomas S. Huang |
CVPR (2) | 1 |
| 2006 | Total Variation Models for Variable Lighting Face RecognitionabstractIn this paper, we present the logarithmic total variation (LTV) model for face recognition under varying illumination, including natural lighting conditions, where we rarely know the strength, direction, or number of light sources. The proposed LTV model has the ability to factorize a single face image and obtain the illumination invariant facial structure, which is then used for face recognition. Our model is inspired by the SQI model but has better edge-preserving ability and simpler parameter selection. The merit of this model is that neither does it require any lighting assumption nor does it need any training. The LTV model reaches very high recognition rates in the tests using both Yale and CMU PIE face databases as well as a face database containing 765 subjects under outdoor lighting conditions. Terrence Chen, Wotao Yin, Xiang Sean Zhou, Dorin Comaniciu, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Illumination Normalization for Face Recognition and Uneven Background Correction Using Total Variation Based Image ModelsabstractWe present a new algorithm for illumination normalization and uneven background correction in images, utilizing the recently proposed TV+L/sup 1/ model: minimizing the total variation of the output cartoon while subject to an L/sup 1/-norm fidelity term. We give intuitive proofs of its main advantages, including the well-known edge preserving capability, minimal signal distortion, and scale-dependent but intensity-independent foreground extraction. We then propose a novel TV-based quotient image model (TVQI) for illumination normalization, an important preprocessing for face recognition under different lighting conditions. Using this model, we achieve 100% face recognition rate on Yale face database B if the reference images are under good lighting condition and 99.45% if not. These results, compared to the average 65% recognition rate of the quotient image model and the average 95% recognition rate of the more recent self quotient image model, show a clear improvement. In addition, this model requires no training data, no assumption on the light source, and no alignment between different images for illumination normalization. We also present the results of the related applications - uneven background correction for cDNA mic roar ray films and digital microscope images. We believe the proposed works can serve important roles in the related fields. Terrence Chen, Wotao Yin, Xiang Sean Zhou, Dorin Comaniciu, Thomas S. Huang |
CVPR (2) | 1 |
| 2005 | Background correction for cDNA microarray images using the TV+L1 modelabstractMOTIVATION: Background correction is an important preprocess in cDNA microarray data analysis. A variety of methods have been used for this purpose. However, many kinds of backgrounds, especially inhomogeneous ones, cannot be estimated correctly using any of the existing methods. In this paper, we propose the use of the TV+L1 model, which minimizes the total variation (TV) of the image subject to an L1-fidelity term, to correct background bias. We demonstrate its advantages over the existing methods by both analytically discussing its properties and numerically comparing it with morphological opening. RESULTS: Experimental results on both synthetic data and real microarray images demonstrate that the TV+L1 model gives the restored intensity that is closer to the true data than morphological opening. As a result, this method can serve an important role in the preprocessing of cDNA microarray data. Wotao Yin, Terrence Chen, Xiang Sean Zhou, Amit Chakraborty |
Bioinform. | 2 |
| 2003 | A New Tracking Technique: Object Tracking and Identification from Motion
Terrence Chen, Yihong Gong, Thomas S. Huang |
CAIP | 1 |
| 2002 | Speeding up the similarity search in multimedia databaseabstractIn recent years, applications in multimedia databases have been more and more important. One new capability is to search by similarity in low-level image features (such as color, texture, shape, and motion). The performance of the similarity search highly relies on efficient index structures. However, current high-dimensional indexing techniques have limitations and a simple sequential scan algorithm can outperform them in many cases. We notice that few researchers have really evaluated the utilization of the actual memory size and the trade-off between I/O access time and computation time. Some methods tried to reduce the I/O access time but caused computation overhead. We propose a novel indexing technique, the RA-Blocks (Region Approximated Blocks), to overcome these limitations and improve the similarity search in multimedia databases. We also demonstrate the better performance of our work in experiments. Terrence Chen, Munehiro Nakazato, Thomas S. Huang |
ICME (2) | 1 |