EDBT 2026 Demo / reviewers in the wild / expert
Ying Sun 0001
dblp:10/5415-1
· DBLP profile ↗
41ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-7224-6726ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong BaselineabstractMetalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processes are often necessary to recover high-quality RGB images from metalens-captured inputs. While recent deep learning-based restoration methods show promise, they typically (1) blur or distort peripheral regions, or (2) fail entirely under unseen illumination conditions. To advance metalens image restoration, we introduce IlluMeta---the first and largest real-world, illumination-aware metalens image dataset—captured across diverse lighting environments. In addition, we propose a novel end-to-end restoration framework that directs attention to challenging regions and adaptively adjusts to varying illuminations via reinforcement learning. Experiments show that our method can be applied in a plug-and-play manner to enhance existing models, significantly improving image restoration quality, especially under unseen lighting conditions, paving the way for broader real-world deployment of metalens technologies. Fen Fang, Xinan Liang, Muli Yang, Jinghong Zheng 0001, Tobias Wilhelm W. Mass, Ying Sun 0001, Xulei Yang, Xuewu Xu, Zhengguo Li |
AAAI | 6 |
| 2026 | Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If CalibratedabstractDespite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world. Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002 |
AAAI | 8 |
| 2026 | Knowledge-guided multi-modality transformer for multi-label genetic mutation prediction
Gexin Huang, Chenfei Wu, Mingjie Li 0006, Xiaojun Chang, Ying Sun 0001, Lei Xing 0001, Xiaodan Liang, Liang Lin 0004 |
Pattern Recognit. | 5 |
| 2025 | Unveiling the Tapestry: The Interplay of Generalization and Forgetting in Continual LearningabstractIn artificial intelligence (AI), generalization refers to a model's ability to perform well on out-of-distribution data related to the given task, beyond the data it was trained on. For an AI agent to excel, it must also possess the continual learning capability, whereby an agent incrementally learns to perform a sequence of tasks without forgetting the previously acquired knowledge to solve the old tasks. Intuitively, generalization within a task allows the model to learn underlying features that can readily be applied to novel tasks, facilitating quicker learning and enhanced performance in subsequent tasks within a continual learning framework. Conversely, continual learning methods often include mechanisms to mitigate catastrophic forgetting, ensuring that knowledge from earlier tasks is retained. This preservation of knowledge over tasks plays a role in enhancing generalization for the ongoing task at hand. Despite the intuitive appeal of the interplay of both abilities, existing literature on continual learning and generalization has proceeded separately. In the preliminary effort to promote studies that bridge both fields, we first present empirical evidence showing that each of these fields has a mutually positive effect on the other. Next, building upon this finding, we introduce a simple and effective technique known as shape-texture consistency regularization (STCR), which caters to continual learning. STCR learns both shape and texture representations for each task, consequently enhancing generalization and thereby mitigating forgetting. Remarkably, extensive experiments validate that our STCR, can be seamlessly integrated with existing continual learning methods, including replay-free approaches. Its performance surpasses these continual learning methods in isolation or when combined with established generalization techniques by a large margin. Our data and source code are available at https://github.com/ZhangLab-DeepNeuroCogLab/distillation-style-cnn. Zenglin Shi, Jie Jing 0001, Ying Sun 0001, Joo-Hwee Lim, Mengmi Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Training-free Object Counting with PromptsabstractThis paper tackles the problem of object counting in images. Existing approaches rely on extensive training data with point annotations for each object, making data collection labor-intensive and time-consuming. To overcome this, we propose a training-free object counter that treats the counting task as a segmentation problem. Our approach leverages the Segment Anything Model (SAM), known for its high-quality masks and zero-shot segmentation capability. However, the vanilla mask generation method of SAM lacks class-specific information in the masks, resulting in inferior counting accuracy. To overcome this limitation, we introduce a prior-guided mask generation method that incorporates three types of priors into the segmentation process, enhancing efficiency and accuracy. Additionally, we tackle the issue of counting objects specified through text by proposing a two-stage approach that combines reference object selection and prior-guided mask generation. Extensive experiments on standard datasets demonstrate the competitive performance of our training-free counter compared to learning-based approaches. This paper presents a promising solution for counting objects in various scenarios without the need for extensive data collection and counting-specific training. Code is available at https://github.com/shizenglin/training-free-object-counter. Zenglin Shi, Ying Sun 0001, Mengmi Zhang |
WACV | 2 |
| 2024 | Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question AnsweringabstractThe main challenge in video question answering (VideoQA) is to capture and understand the complex spatial and temporal relations between objects based on given questions. Existing graph-based methods for VideoQA usually ignore keywords in questions and employ a simple graph to aggregate features without considering relative relations between objects, which may lead to inferior performance. In this paper, we propose a Keyword-aware Relative Spatio-Temporal (KRST) graph network for VideoQA. First, to make question features aware of keywords, we employ an attention mechanism to assign high weights to keywords during question encoding. The keyword-aware question features are then used to guide video graph construction. Second, because relations are relative, we integrate the relative relation modeling to better capture the spatio-temporal dynamics among object nodes. Moreover, we disentangle the spatio-temporal reasoning into an object-level spatial graph and a frame-level temporal graph, which reduces the impact of spatial and temporal relation reasoning on each other. Extensive experiments on the TGIF-QA, MSVD-QA and MSRVTT-QA datasets demonstrate the superiority of our KRST over multiple state-of-the-art methods. Hehe Fan, Dongyun Lin, Ying Sun 0001, Mohan Kankanhalli, Joo-Hwee Lim |
IEEE Trans. Multim. | 4 |
| 2024 | Controllable Video Generation With Text-Based InstructionsabstractMost of the existing studies on controllable video generation either transfer disentangled motion to an appearance without detailed control over motion or generate videos of simple actions such as the movement of arbitrary objects conditioned on a control signal from users. In this study, we introduce Controllable Video Generation with text-based Instructions (CVGI) framework that allows text-based control over action performed on a video. CVGI generates videos where hands interact with objects to perform the desired action by generating hand motions with detailed control through text-based instruction from users. By incorporating the motion estimation layer, we divide the task into two sub-tasks: (1) control signal estimation and (2) action generation. In control signal estimation, an encoder models actions as a set of simple motions by estimating low-level control signals for text-based instructions with given initial frames. In action generation, generative adversarial networks (GANs) generate realistic hand-based action videos as a combination of hand motions conditioned on the estimated low control level signal. Evaluations on several datasets (EPIC-Kitchens-55, BAIR robot pushing, and Atari Breakout) show the effectiveness of CVGI in generating realistic videos and in the control over actions. Ali Koksal, Kenan E. Ak, Ying Sun 0001, Deepu Rajan, Joo-Hwee Lim |
IEEE Trans. Multim. | 3 |
| 2023 | Counterfactual Dynamics Forecasting - a New Setting of Quantitative ReasoningabstractRethinking and introspection are important elements of human intelligence. To mimic these capabilities, counterfactual reasoning has attracted attention of AI researchers recently, which aims to forecast the alternative outcomes for hypothetical scenarios (“what-if”). However, most existing approaches focused on qualitative reasoning (e.g., casual-effect relationship). It lacks a well-defined description of the differences between counterfactuals and facts, as well as how these differences evolve over time. This paper defines a new problem formulation - counterfactual dynamics forecasting - which is described in middle-level abstraction under the structural causal models (SCM) framework and derived as ordinary differential equations (ODEs) as low-level quantitative computation. Based on it, we propose a method to infer counterfactual dynamics considering the factual dynamics as demonstration. Moreover, the evolution of differences between facts and counterfactuals are modelled by an explicit temporal component. The experimental results on two dynamical systems demonstrate the effectiveness of the proposed method. Yanzhu Liu, Ying Sun 0001, Joo-Hwee Lim |
AAAI | 2 |
| 2023 | Meta Compositional Referring Expression SegmentationabstractReferring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling this task may not be able to fully capture semantics and visual representations of individual concepts, which limits their generalization capability, especially when handling novel compositions of learned concepts. In this work, through the lens of meta learning, we propose a Meta Compositional Referring Expression Segmentation (MCRES) framework to enhance model compositional generalization performance. Specifically, to handle various levels of novel compositions, our framework first uses training data to construct a virtual training set and multiple virtual testing sets, where data samples in each virtual testing set contain a level of novel compositions w.r.t. the virtual training set. Then, following a novel meta optimization scheme to optimize the model to obtain good testing performance on the virtual testing sets after training on the virtual training set, our framework can effectively drive the model to better capture semantics and visual representations of individual concepts, and thus obtain robust generalization performance even when handling novel compositions. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our framework. Mark He Huang, Xindi Shang, Zehuan Yuan, Ying Sun 0001, Jun Liu 0036 |
CVPR | 5 |
| 2023 | Learning to Learn: How to Continuously Teach Humans and MachinesabstractCurriculum design is a fundamental component of education. For example, when we learn mathematics at school, we build upon our knowledge of addition to learn multiplication. These and other concepts must be mastered before our first algebra lesson, which also reinforces our addition and multiplication skills. Designing a curriculum for teaching either a human or a machine shares the underlying goal of maximizing knowledge transfer from earlier to later tasks, while also minimizing forgetting of learned tasks. Prior research on curriculum design for image classification focuses on the ordering of training examples during a single offline task. Here, we investigate the effect of the order in which multiple distinct tasks are learned in a sequence. We focus on the online class-incremental continual learning setting, where algorithms or humans must learn image classes one at a time during a single pass through a dataset. We find that curriculum consistently influences learning outcomes for humans and for multiple continual machine learning algorithms across several benchmark datasets. We introduce a novel-object recognition dataset for human curriculum learning experiments and observe that curricula that are effective for humans are highly correlated with those that are effective for machines. As an initial step towards automated curriculum design for online class-incremental learning, we propose a novel algorithm, dubbed Curriculum Designer (CD), that designs and ranks curricula based on inter-class feature similarities. We find significant overlap between curricula that are empirically highly effective and those that are highly ranked by our CD. Our study establishes a framework for further research on teaching humans and machines to learn continuously using optimized curricula. Our code and data are available through this link. Parantak Singh, Ankur Sikarwar, Weixian Lei, Difei Gao, Morgan B. Talbot, Ying Sun 0001, Zheng Shou 0001, Gabriel Kreiman, Mengmi Zhang |
ICCV | 7 |
| 2023 | Learning by Imagination: A Joint Framework for Text-Based Image Manipulation and Change CaptioningabstractImage and text are dual modalities of our semantic interpretation. Changing images based on text descriptions allows us to imagine and visualize the world (a.k.a. text-based image manipulation (TIM)). In this paper, we introduce a framework that combines TIM with change captioning (CC) and utilizes the benefits of co-training. CC aims to describe what has changed in a scene and can be regarded as the inverse version of TIM where both tasks rely on generative networks. These generative networks can be regarded as data producers of each other and unlike previous methods, we discover that integrating their learning procedures can benefit both. Since the CC module describes differences between two images as text, the CC module can be used as evaluation criteria and provide feedback. Furthermore, we utilize a shared attention mechanism in TIM and CC modules to localize towards prominent regions as well as enabling a change-aware discriminator. In the opposite direction, the output image synthesized by the TIM module can be assessed with the CC module, by checking whether the ground truth text description can be redescribed. Following this insight, not only do we boost the training of the TIM module, but we also utilize the TIM module as additional supervision for the CC training. Experimental results show that our framework outperforms existing TIM methods on several datasets substantially and we achieve marginal improvements in the CC module. To our best knowledge, this is the first study dedicated to the joint training of TIM and CC tasks. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Multim. | 2 |
| 2022 | Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation with Reliable Voted Pseudo LabelsabstractIn this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unla-beled data by scaling up or down point clouds and then predicting the scales. Second, to capture the local structure in a self-supervised manner, we propose to project a 3D local area onto a 2D plane and then learn to reconstruct the squeezed region. Moreover, to effectively transfer the knowledge from source domain, we propose to vote pseudo labels for target samples based on the labels of their nearest source neighbors in the shared feature space. To avoid the noise caused by incorrect pseudo labels, we only select re-liable target samples, whose voting consistencies are high enough, for enhancing adaptation. The voting method is able to adaptively select more and more target samples during training, which in return facilitates adaptation because the amount of labeled target data increases. Experiments on PointDA (ModelNet-10, ShapeNet-10 and ScanNet-10) and Sim-to-Real (ModelNet-11, ScanObjectNN-11, ShapeNet-9 and ScanObjectNN-9) demonstrate the effectiveness of our method. Hehe Fan, Xiaojun Chang, Wanyue Zhang, Ying Sun 0001, Mohan Kankanhalli |
CVPR | 5 |
| 2022 | Entropy guided attention network for weakly-supervised action localization
Ying Sun 0001, Hehe Fan, Tao Zhuo, Joo-Hwee Lim, Mohan Kankanhalli |
Pattern Recognit. | 2 |
| 2022 | Image Understanding With Reinforcement Learning: Auto-Tuning Image Attributes and Model Parameters for Object Detection and SegmentationabstractModels for image semantics understanding, such as deep learning (DL) models and mathematical models, are often trained on specific dataset or configured with specific parameters. When deploying such models on new tasks in a different test environment, it requires considerable effort to re-train the model or extensive expertise to tune the parameters. In this paper, we propose a smart reinforcement learning (RL) agent that could learn to tune parameters automatically to enhance model performance. The learning process is formulated as a generic control task for parameter adjustment, and applied to two use scenarios: (1) image attributes tuning to improve object detection performance on fixed DL model, and (2) parameter tuning of the mathematical model (Level Set) for image segmentation. We design a novel dynamic threshold mechanism in a multi-branch RL agent to effectively tune parameters of image qualities (for object detection) and Level Set models (for object segmentation). We conduct experiments on Pascal-VOC testing set, MS COCO validation set and a proprietary dataset of industrial components, where we achieve substantial improvement on object detection accuracy. We also perform experiments on the automatic parameter tuning of Level Set models. Results show that our method facilitates considerable performance improvement on public datasets compared with baseline method. Fen Fang, Qianli Xu, Ying Sun 0001, Joo-Hwee Lim |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | TAILOR: Teaching with Active and Incremental Learning for Object RegistrationabstractWhen deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and labor- intensive. We present TAILOR - a method and system for ob- ject registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informa- tive images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox as- sembly task through natural interactions. Qianli Xu, Nicolas Gauthier, Wenyu Liang, Fen Fang, Hui Li Tan, Ying Sun 0001, Yan Wu 0002, Liyuan Li, Joo-Hwee Lim |
AAAI | 6 |
| 2021 | Robust Multi-Frame Future Prediction By Leveraging View SynthesisabstractIn this paper, we focus on the problem of video prediction, i.e., future frame prediction. Most state-of-the-art techniques focus on synthesizing a single future frame at each step. However, this leads to utilizing the model’s own predicted frames when synthesizing multi-step prediction, resulting in gradual performance degradation due to accumulating errors in pixels. To alleviate this issue, we propose a model that can handle multi-step prediction. Additionally, we employ techniques to leverage from view synthesis for future frame prediction, where both problems are treated independently in the literature. Our proposed method employs multiview camera pose prediction and depth-prediction networks to project the last available frame to desired future frames via differentiable point cloud renderer. For the synthesis of moving objects, we utilize an additional refinement stage. In experiments, we show that the proposed framework outperforms state-of-theart methods in both KITTI and Cityscapes datasets. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 2 |
| 2021 | Action Relational Graph for Weakly-Supervised Temporal Action LocalizationabstractThe task of weakly-supervised temporal action localization (WTAL) is to recognize plentiful unstructured actions in untrimmed videos with only video-level class labels. As various actions may occur in an untrimmed video, it is desirable to capture the correlation among different actions to effectively identify the target actions. In this paper, we propose a novel Action Relational Graph Network (ARG-Net) to model the correlation between action labels. Specifically, we build a co-occurrence graph using Graph Convolutional Network (GCN), where the graph nodes and edges are represented by word embedding of action labels and relations between two labels, respectively. Then we apply the GCNs to project the action label embeddings into a set of correlated action classifiers which are multiplied with the learned video representations for video-level classification. To facilitate discriminative video representation learning, we employ the attention mechanism to model the probability of a frame containing action instances. A new Action Normalization Loss (ANL) is proposed to further alleviate the confusion from irrelevant background frames (i.e., frames containing no actions). Experimental results on THUMOS14 and ActivityNet1.2 datasets demonstrate that our ARG-Net outperforms the state-of-the-art methods. Ying Sun 0001, Dongyun Lin, Joo-Hwee Lim |
ICIP | 2 |
| 2021 | A Diagnostic Study Of Visual Question Answering With Analogical ReasoningabstractThe deep learning community has made rapid progress in low-level visual perception tasks such as object localization, detection and segmentation. However, for tasks such as Visual Question Answering (VQA) and visual language grounding that require high-level reasoning abilities, huge gaps still exist between artificial systems and human intelligence. In this work, we perform a diagnostic study on recent popular VQA in terms of analogical reasoning. We term it as Analogical VQA, where a system needs to reason on a group of images to find analogical relations among them in order to correctly answer a natural language question. To study the task in depth, we propose an initial diagnostic synthetic dataset CLEVR-Analogy, which tests a range of analogical reasoning abilities (e.g. reasoning on object attributes, spatial relationships, existence, and arithmetic analogies). We benchmark various recent state-of-the-art methods on our dataset and compare the results against human performance, and discover that existing systems fall shorts when facing analogical reasoning involving spatial relationships. The dataset and code will be publicly available to facilitate future research. Hongyuan Zhu 0002, Ying Sun 0001, Dongkyu Choi, Cheston Tan, Joo-Hwee Lim |
ICIP | 3 |
| 2020 | Learning Cross-Modal Representations for Language-Based Image ManipulationabstractIn this paper, we propose a generative architecture for manipulating images/scenes with natural language descriptions. This is a challenging task as the generative network is expected to perform the given text instruction without changing the non-affiliating contents of the input image. Two main drawbacks of the existing methods are their limitation of performing changes that would affect only a limited region and the inability of handling complex instructions. The proposed approach, designed to address these limitations initially uses two sets of networks to extract the image and text features respectively. Rather than a simple combination of these two modalities during the image manipulation process, we use an improved technique to compose image and text features. Additionally, the generative network utilizes similarity learning to improve text manipulation which also enforces only the text-relevant changes on the input image. Our experiments on CSS and Fashion Synthesis datasets show that the proposed approach performs remarkably well and outperforms the baseline frameworks in terms of R-precision and FID. Kenan E. Ak, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 2 |
| 2020 | Task-Oriented Multi-Modal Question Answering For Collaborative ApplicationsabstractCobots that can work in human workspaces and adapt to human need to understand and respond to human’s inquiry and instruction. In this paper, we propose new question answering (QA) task and dataset for human-robot collaboration on task-oriented operation, i.e., task-oriented collaborative QA (TCQA). Differing from conventional video QA for answering questions about what happened in video clips constrained by scripts and subtitles, TC-QA aims to share common ground for task-oriented operation through question answering. We propose an open-end (OE) format of answer with text reply, image with annotated related objects, and video with operation duration to guide operation execution. Designed for grounding, the TC-QA dataset comprises query videos and questions to seek acknowledgement, correction, attention to task-related objects, and information on objects or operation. Due to the flexibility of real-world task with limited training sample, we propose and evaluate a baseline method based on a hybrid approach. The hybrid approach employs deep learning methods for object detection, hand detection and gesture recognition, and symbolic reasoning to ground question on observation for providing the answer. Our experiments show that the hybrid method is effective for the TC-QA task. Hui Li Tan, Mei Chee Leong, Qianli Xu, Liyuan Li, Fen Fang, Nicolas Gauthier, Ying Sun 0001, Joo-Hwee Lim |
ICIP | 8 |
| 2020 | 6D Pose Estimation with Correlation Fusionabstract6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illumination, so it is important to complement them with depth information. However, existing methods using RGB-D data cannot adequately exploit consistent and complementary information between RGB and depth modalities. In this paper, we present a novel method to effectively consider the correlation within and across both modalities with attention mechanism to learn discriminative and compact multi-modal features. Then, effective fusion strategies for intra- and inter-correlation modules are explored to ensure efficient information flow between RGB and depth. To our best knowledge, this is the first work to explore effective intra- and inter-modality fusion in 6D pose estimation. The experimental results show that our method can achieve the state-of-the-art performance on LineMOD and YCB-Video dataset. We also demonstrate that the proposed method can benefit a real-world robot grasping task by providing accurate object pose estimation. Hongyuan Zhu 0002, Ying Sun 0001, Cihan Acar, Yan Wu 0002, Liyuan Li, Cheston Tan, Joo-Hwee Lim |
ICPR | 3 |
| 2018 | Dual-Resolution U-Net: Building Extraction from Aerial ImagesabstractDeep learning has been applied to segment buildings from high-resolution images with promising results. However, there still exist the problems stemming from training on split patches and class imbalances. To overcome these problems, we propose a dual-resolution U-Net that uses pairs of images as inputs to capture both high and low resolution features. We also employ a soft Jaccard loss to place more emphasis on the sparse and low accuracy samples. The images from different regions are further balanced according to their building densities. With our architecture, we achieved state-of-the-art results on the Inria aerial image labeling dataset without any post-processing. Kangkang Lu 0001, Ying Sun 0001, Sim Heng Ong |
ICPR | 2 |
| 2015 | Fast Reconstruction of Accelerated Dynamic MRI Using Manifold Kernel Regression
Kanwal K. Bhatia, Jose Caballero, Anthony N. Price, Ying Sun 0001, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (3) | 4 |
| 2013 | Three-dimensional segmentation of the left ventricle in late gadolinium enhanced MR images of chronic infarction combining long- and short-axis information
Dong Wei 0004, Ying Sun 0001, Sim Heng Ong, Ping Chai, Lynette L. Teo, Adrian F. Low |
Medical Image Anal. | 2 |
| 2013 | A Review of Recent Advances in Registration Techniques Applied to Minimally Invasive TherapyabstractMinimally invasive and less invasive procedure is becoming more and more common in medical therapy. Image guidance is an indispensable component in minimally invasive procedures by providing critical information about the position of the target sites and the optimal manipulation of the devices, while the field of view is limited to naked eyes due to the small incision. Registration is one of the enabling technologies for computer-aided image guidance, which brings high-resolution pre-operative data into the operating room to provide more realistic information about the patient's anatomy. In this paper, we survey the recent advances in registration techniques applied to minimally and/or less invasive therapy, including a wide variety of therapies in surgery, endoscopy, interventional cardiology, interventional radiology, and hybrid procedures. The registration approaches are categorized into several groups, including projection-to-volume, slice-to-volume, video-to-volume, and volume-to-volume registration. The focus is on recent advances in registration techniques that are specifically developed for minimally and/or less invasive procedures in the following medical specialties: neuroradiology and neurosurgery, cardiac applications, and thoracic-abdominal interventions. Rui Liao, Li Zhang 0024, Ying Sun 0001, Shun Miao, Christophe Chefd'Hotel |
IEEE Trans. Multim. | 3 |
| 2012 | Registration of Pre-Operative CT and Non-Contrast-Enhanced C-Arm CT: An Application to Trans-Catheter Aortic Valve Implantation (TAVI)
Yongning Lu, Ying Sun 0001, Rui Liao, Sim Heng Ong |
ACCV (2) | 2 |
| 2012 | Learning-based deformable registration using weighted mutual information
Yongning Lu, Rui Liao, Li Zhang 0024, Ying Sun 0001, Christophe Chefd'Hotel, Sim Heng Ong |
ICPR | 4 |
| 2012 | Integrating Segmentation Information for Improved MRF-Based Elastic Image RegistrationabstractIn this paper, we propose a method to exploit segmentation information for elastic image registration using a Markov-random-field (MRF)-based objective function. MRFs are suitable for discrete labeling problems, and the labels are defined as the joint occurrence of displacement fields (for registration) and segmentation class probability. The data penalty is a combination of the image intensity (or gradient information) and the mutual dependence of registration and segmentation information. The smoothness is a function of the interaction between the defined labels. Since both terms are a function of registration and segmentation labels, the overall objective function captures their mutual dependence. A multiscale graph-cut approach is used to achieve subpixel registration and reduce the computation time. The user defines the object to be registered in the floating image, which is rigidly registered before applying our method. We test our method on synthetic image data sets with known levels of added noise and simulated deformations, and also on natural and medical images. Compared with other registration methods not using segmentation information, our proposed method exhibits greater robustness to noise and improved registration accuracy. Dwarikanath Mahapatra, Ying Sun 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Orientation Histograms as Shape Priors for Left Ventricle Segmentation Using Graph Cuts
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (3) | 2 |
| 2011 | Myocardial Segmentation of Late Gadolinium Enhanced MR Images by Propagation of Contours from Cine MR Images
Dong Wei 0004, Ying Sun 0001, Ping Chai, Adrian F. Low, Sim Heng Ong |
MICCAI (3) | 2 |
| 2011 | Pseudo ground truth based nonrigid registration of myocardial perfusion MRI
Ying Sun 0001, Ping Chai |
Medical Image Anal. | 2 |
| 2010 | A patch-based spatiotemporal phase unwrapping method for phase contrast MRI using graph cutsabstractPhase unwrapping is an important and challenging problem in phase contrast magnetic resonance imaging (PC-MRI). In this paper, we propose a new algorithm for phase unwrapping based on graph cuts. Our algorithm takes a patch-based approach which has the advantages of simplicity and robustness to phase noise. To make use of temporal information from the neighboring frames as well as spatial information from the current frame, the energy function is designed to combine both spatial and temporal constraints. The proposed method has been tested with real PC-MRI data. Experimental results demonstrate that our algorithm is capable of unwrapping images with severe phase wrapping, and that it outperforms other existing popular phase unwrapping algorithms both qualitatively and quantitatively. Wenyu Xie, Ying Sun 0001, Sim Heng Ong |
ICARCV | 2 |
| 2010 | Active image: A shape and topology preserving segmentation method using B-spline free form deformationsabstractUnlike most conventional active contour models that drive the initial contours to match the object boundaries, we propose an “active image” segmentation method that deforms the image to match the initial contours, and hence can simultaneously segment multiple objects. The deformation field is modeled by B-spline free form deformations (FFD). Penalizing the bending energy of B-spline FFD enables us to preserve both the local shape and the topology of objects of interest. Preliminary results on both synthetic and real world images show that the proposed method is able to overcome low contrast, occlusion, and other defects by using simple segmentation criteria. Ying Sun 0001 |
ICIP | 2 |
| 2010 | An MRF framework for joint registration and segmentation of natural and perfusion imagesabstractRegistration and segmentation provide complementary information about each other. In this paper we propose a method for the joint registration and segmentation (JRS) of images using Markov random fields (MRFs). The use of MRFs allows us to formulate the problem as one of labeling and apply fast discrete optimization techniques like graph cuts. Graph cuts is able to overcome the limitations of previously used active contour frameworks namely, large number of iterations, risk of being trapped in local minima, and sensitivity to initialization. The labels in the MRF formulation indicate joint occurrence of displacement vectors and segmentation class and the energy formulation is able to capture their mutual dependency. Experiments on real patient perfusion data and natural images show that JRS gives better performance than conventional registration and segmentation methods. Dwarikanath Mahapatra, Ying Sun 0001 |
ICIP | 2 |
| 2010 | Joint Registration and Segmentation of Dynamic Cardiac Perfusion Images Using MRFs
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (1) | 2 |
| 2009 | Nonrigid Registration of Myocardial Perfusion MRI Using Pseudo Ground Truth
Ying Sun 0001 |
MICCAI (1) | 2 |
| 2008 | Automatic opacity detection in retro-illumination images for cortical cataract diagnosisabstractComputer aided analysis of medical images, a unique type of non-text media, can facilitate clinical diagnosis. As an example, an automatic opacity detection approach is proposed in this paper to grade cortical cataract more objectively. The automatic pupil detection is performed by detecting the strongest edges on the convex hull and ellipse fitting using nonlinear least square method. The cortical opacity is detected by radial edge detection and post-processing. The automatic grades are assigned following Wisconsin cataract grading protocol. The accuracy of pupil detection is 98.2%. The mean error of opacity area detection is 7 percent compared with the result of human grader. And 86.3% accurate grades of cortical cataract are achieved. This is the first time that the spoke-like feature is utilized in the automatic detection of cortical cataract to separate from other opacity types. The encouraging results show that it is probable to apply the proposed approach to clinical diagnosis later. Huiqi Li, Liling Ko, Joo-Hwee Lim, Jiang Liu 0001, Damon Wing Kee Wong, Tien Yin Wong, Ying Sun 0001 |
ICME | 7 |
| 2008 | Illumination invariant tracking in office environments using neurobiology-saliency based particle filterabstractBackground subtraction is a commonly employed approach for tracking in scenarios where the ambience is more or less constant in terms of illumination and number of objects. However in office environments, where the illumination can very easily change by switching off or on lights, the background subtraction method can lead to erroneous tracking. In this paper we propose a neurobiology-saliency based particle filter approach that uses low-level features like color, luminance and edge information along with motion cues to track a single person. We have tested our method on clips showing a single person carrying out a range of activities expected in an office environment. Our method performs better than a background subtraction method using a Kalman filter, in terms of the number of frames showing correct tracking and change detection for automatic initialization of tracks. Dwarikanath Mahapatra, Mukesh Saini, Ying Sun 0001 |
ICME | 3 |
| 2008 | Nonrigid Registration of Dynamic Renal MR Images Using a Saliency Based MRF Model
Dwarikanath Mahapatra, Ying Sun 0001 |
MICCAI (1) | 2 |
| 2004 | Integrated registration of dynamic renal perfusion MR images
Ying Sun 0001, Marie-Pierre Jolly, José M. F. Moura |
ICIP | 1 |
| 2004 | Contrast-Invariant Registration of Cardiac and Renal MR Perfusion Images
Ying Sun 0001, Marie-Pierre Jolly, José M. F. Moura |
MICCAI (1) | 1 |