Si Yong Yeo

dblp:38/9164 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-6403-6019ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 MAGF: Multi-scale attention and gated fusion for multi-modal glaucoma grading
Haixi Cheng, Bo Zhang 0002, Huihui Fang, Yanwu Xu 0001, Si Yong Yeo
Expert Syst. Appl.6
2025 MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations
abstract
Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model’s ability to synthesize coherent and novel visual representations from textual prompts, thereby reducing the effectiveness of multi-modal learning. In this work, we propose MedUnifier, a unified VLP framework tailored for medical data. MedUnifier seamlessly integrates text-grounded image generation capabilities with multi-modal learning strategies, including image-text contrastive alignment, image-text matching and image-grounded text generation. Unlike traditional methods that reply on continuous visual representations, our approach employs visual vector quantization, which not only facilitates a more cohesive learning strategy for cross-modal understanding but also enhances multi-modal generation quality by effectively leveraging discrete representations. Our framework’s effectiveness is evidenced by the experiments on established benchmarks, including uni-modal tasks, cross-modal tasks, and multi-modal tasks, where it achieves state-of-the-art performance across various tasks. MedUnifier also offers a highly adaptable tool for a wide range of language and vision tasks in healthcare, marking advancement toward the development of a generalizable AI model for medical applications.
Yang Yu 0079, Yucheng Chen 0002, Xulei Yang, Si Yong Yeo
CVPR5
2025 Future-Aware Interaction Network for Motion Forecasting
Shijie Li 0006, Xun Xu 0002, Si Yong Yeo, Xulei Yang
ICCV4
2025 PVChat: Personalized Video Chat with One-Shot Learning
abstract
Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthcare and smart home environments. To address this limitation, we propose a one-shot learning framework PVChat, the first personalized ViLLM that enables subject-aware question answering (QA) from a single video for each subject. Our approach optimizes a Mixture-of-Heads (MoH) enhanced ViLLM on a synthetically augmented video-QA dataset, leveraging a progressive image-to-video learning strategy. Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. To enhance subject-specific learning, we propose a ReLU Routing MoH attention mechanism, alongside two novel objectives: (1) Smooth Proximity Regularization for progressive learning through exponential distance scaling and (2) Head Activation Enhancement for balanced attention routing. Finally, we adopt a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations. We evaluate PVChat on diverse datasets covering medical scenarios, TV series, anime, and real-world footage, demonstrating its superiority in personalized feature understanding after learning from a single video, compared to state-of-the-art ViLLMs.
Weilong Yan, Zhenxi Li, Si Yong Yeo
ICCV9
2025 Prior-Guided Prototype Aggregation Learning for Alzheimer's Disease Diagnosis
Yueqin Diao, Huihui Fang, Hanyi Yu, Yaling Tao, Ziyan Huang, Si Yong Yeo, Yanwu Xu 0001
MICCAI (15)7
2025 EFFDNet: A Scribble-Supervised Medical Image Segmentation Method with Enhanced Foreground Feature Discrimination
Jinhua Liu 0003, Shu Yun Tan, Xulei Yang, Yanwu Xu 0004, Si Yong Yeo
MICCAI (16)5
2025 ColonNeRF: High-fidelity neural reconstruction of long colonoscopy
Yufei Shi 0003, Beijia Lu, Jia-Wei Liu, Ming Li 0073, Si Yong Yeo, Zheng Shou 0001
Neurocomputing5
2024 Edge-Guided and Cross-Scale Feature Fusion Network for Efficient Multi-contrast MRI Super-Resolution
Bo Zhang 0002, Si Yong Yeo
ICPR (27)4
2023 Chaotic World: A Large and Challenging Benchmark for Human Behavior Understanding in Chaotic Events
abstract
Understanding and analyzing human behaviors (actions and interactions of people), voices, and sounds in chaotic events is crucial in many applications, e.g., crowd management, emergency response services. Different from human behaviors in daily life, human behaviors in chaotic events are generally different in how they behave and influence others, and hence are often much more complex. However, currently there is lack of a large video dataset for analyzing human behaviors in chaotic situations. To this end, we create the first large and challenging multi-modal dataset, Chaotic World, that simultaneously provides different levels of fine-grained and dense spatio-temporal annotations of sounds, individual actions and group interaction graphs, and even text descriptions for each scene in each video, thereby enabling a thorough analysis of complicated behaviors in crowds and chaos. Our dataset consists of a total of 299,923 annotated instances for detecting human behaviors for Spatiotemporal Action Localization in chaotic events, 224,275 instances for identifying interactions between people for Behavior Graph Analysis in chaotic events, 336,390 instances for localizing relevant scenes of interest in long videos for Spatiotemporal Event Grounding, and 378,093 instances for triangulating the source of sound for Event Sound Source Localization. Given the practical complexity and challenges in chaotic events (e.g., large crowds, serious occlusions, complicated interaction patterns), our dataset shall be able to facilitate the community to develop, adapt, and evaluate various types of advanced models for analyzing human behaviors in chaotic events. We also design a simple yet effective IntelliCare model with a Dynamic Knowledge Pathfinder module that intelligently learns from multiple tasks and can analyze various aspects of a chaotic scene in a unified architecture. This method achieves promising results in experiments. Dataset and code can be found at https://github.com/sutdcv/Chaotic-World.
Kian Eng Ong, Xun Long Ng, Wenjie Ai, Kuangyi Zhao, Si Yong Yeo, Jun Liu 0036
ICCV6
2022 Animal Kingdom: A Large and Diverse Dataset for Animal Behavior Understanding
abstract
Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks, and also limited variations in environmental conditions and viewpoints. To address these limitations, we create a large and diverse dataset, Animal Kingdom, that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors. The wild animal footages used in our dataset record different times of the day in extensive range of environments containing variations in backgrounds, viewpoints, illumination and weather conditions. More specifically, our dataset contains 50 hours of annotated videos to localize relevant animal behavior segments in long videos for the video grounding task, 30K video sequences for the fine-grained multi-label action recognition task, and 33K frames for the pose estimation task, which correspond to a diverse range of animals with 850 species across 6 major animal classes. Such a challenging and comprehensive dataset shall be able to facilitate the community to develop, adapt, and evaluate various types of advanced methods for animal behavior analysis. Moreover, we propose a Collaborative Action Recognition (CARe) model that learns general and specific features for action recognition with unseen new animals. This method achieves promising performance in our experiments. Our dataset can be found at https://sutdcv.github.io/Animal-Kingdom.
Xun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni, Si Yong Yeo, Jun Liu 0036
CVPR5
2020 Automatic detection of anatomical landmarks in brain MR scanning using multi-task deep neural networks
Xulei Yang, Wai Teng Tang, Gabriel Tjio, Si Yong Yeo, Yi Su 0001
Neurocomputing4
2016 Cardiac image segmentation by random walks with dynamic shape constraint
abstract
The quantitative analysis of the left ventricle (LV) contractile function is one of the key steps in the assessment of cardiovascular disease. Such analysis greatly depends on the accurate delineation of LV boundary from cardiac sequences. However, segmentation of the LV still remains a challenging problem due to its subtle boundary, occlusion, and image inhomogeneity. To overcome such difficulties, the authors propose a novel segmentation method by incorporating a dynamic shape constraint into the weighting function of the random walks segmentation algorithm. This approach involves iterative updates on the intermediate result to achieve the desired solution. The inclusion of a shape constraint restricts the solution space of the segmentation result to handle misleading information that may come from noise, weak boundaries and clutter, leading to increased robustness of the algorithm. The authors describe the details of the proposed method and demonstrate its effectiveness in segmenting the LV from real cardiac magnetic resonance (CMR) image sets. The experimental results demonstrate that the proposed method obtains better segmentation performance than the standard method.
Xulei Yang, Yi Su 0001, Rubing Duan, Haijin Fan, Si Yong Yeo, Calvin Chi-Wan Lim, Liang Zhong 0001, Ru-San Tan
IET Comput. Vis.5
2013 Right Ventricle Segmentation by Temporal Information Constrained Gradient Vector Flow
abstract
Evaluation of right ventricular (RV) structure and function is of importance in the management of most cardiac disorders. But the segmentation of RV has always been considered challenging due to low contrast of the myocardium with surrounding and high shape variability of the RV. In this paper, we present a 2D + T active contour model for segmentation and tracking of RV endocardium on cardiac magnetic resonance (MR) images. To take into account the temporal information between adjacent frames, we propose to integrate the time-dependent constraints into the energy functional of the classical gradient vector flow (GVF). As a result, the prior motion knowledge of RV is introduced in the deformation process through the time-dependent constraints in the proposed GVF-T model. A weighting parameter is introduced to adjust the weight of the temporal information against the image data itself. The additional external edge forces retrieved from the temporal constraints may be useful for the RV segmentation, such that lead to a better segmentation performance. The effectiveness of the proposed approach is supported by experimental results on synthetic and cardiac MR images.
Xulei Yang, Si Yong Yeo, Yi Su 0001, Calvin Chi-Wan Lim, Liang Zhong 0001, Ru-San Tan
SMC2
2012 Segmentation of Vessel Geometries from Medical Images using GPF Deformable Model
Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu
ICPRAM (1)1
2012 Implicit active contours for N-dimensional biomedical image segmentation
abstract
The segmentation of shapes from biomedical images has a wide range of uses such as image based modelling and bioimage analysis. In this paper, an active contour model is proposed for the segmentation of N-dimensional biomedical images. The proposed model uses a curvature smoothing flow and an image attraction force derived from the interactions between the geometries of the active contour model and the image objects. The active contour model is formulated using the level set method so as to handle topological changes automatically. The magnitude and orientation of the image attraction force is based on the relative geometric configurations between the active contour model and the image object boundaries. The vector force field is therefore dynamic, and the active contour model can propagate through narrow structures to segment complex shapes efficiently. The proposed model utilizes pixel interactions across the image domain, which gives a coherent representation of the image object shapes. This allows the active contour model to be robust to image noise and weak object edges. The proposed model is compared against widely used active contour models in the segmentation of anatomical shapes from biomedical images. It is shown that the proposed model has several advantages over existing techniques and can be used for the segmentation of biomedical images efficiently.
Si Yong Yeo
SMC1
2011 Level set segmentation with robust image gradient energy and statistical shape prior
abstract
We propose a new level set segmentation method with statistical shape prior using a variational approach. The image energy is derived from a robust image gradient feature. This gives the active contour a global representation of the geometric configuration, making it more robust to image noise, weak edges and initial configurations. Statistical shape information is incorporated using nonparametric shape density distribution, which allows the model to handle relatively large shape variations. Comparative examples using both synthetic and real images show the robustness and efficiency of the proposed method.
Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu
ICIP1
2011 Geometrically Induced Force Interaction for Three-Dimensional Deformable Models
abstract
In this paper, we propose a novel 3-D deformable model that is based upon a geometrically induced external force field which can be conveniently generalized to arbitrary dimensions. This external force field is based upon hypothesized interactions between the relative geometries of the deformable model and the object boundary characterized by image gradient. The evolution of the deformable model is solved using the level set method so that topological changes are handled automatically. The relative geometrical configurations between the deformable model and the object boundaries contribute to a dynamic vector force field that changes accordingly as the deformable model evolves. The geometrically induced dynamic interaction force has been shown to greatly improve the deformable model performance in acquiring complex geometries and highly concave boundaries, and it gives the deformable model a high invariancy in initialization configurations. The voxel interactions across the whole image domain provide a global view of the object boundary representation, giving the external force a long attraction range. The bidirectionality of the external force field allows the new deformable model to deal with arbitrary cross-boundary initializations, and facilitates the handling of weak edges and broken boundaries. In addition, we show that by enhancing the geometrical interaction field with a nonlocal edge-preserving algorithm, the new deformable model can effectively overcome image noise. We provide a comparative study on the segmentation of various geometries with different topologies from both synthetic and real images, and show that the proposed method achieves significant improvements against existing image gradient techniques.
Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu
IEEE Trans. Image Process.1
2009 Geometric Potential Force for the Deformable Model
abstract
We propose a new external force field for deformable models which can be conve-niently generalized to high dimensions. The external force field is based on hypothesized interactions between the relative geometries of the deformable model and image gradi-ents. The evolution of the deformable model is solved using the level set method. The dynamic interaction forces between the geometries can greatly improve the deformable model performance in acquiring complex geometries and highly concave boundaries, and in dealing with weak image edges. The new deformable model can handle arbi-trary cross-boundary initializations. Here, we show that the proposed method achieve significant improvements when compared against existing state-of-the-art techniques. 1
Si Yong Yeo, Xianghua Xie, Igor Sazonov, Perumal Nithiarasu
BMVC1