VLDB 2026 Research / reviewers in the wild / expert
Yong Liu 0026
dblp:29/4867-26
· DBLP profile ↗
65ranked-venue papers
2as first author
50since 2021 · last 2026
0000-0002-1590-2029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 28 since 2021Artificial intelligence and machine learning · 29 · 25 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 22 since 2021Computer networks · 6 · 2 first-authorSystems, architecture and hardware · 4Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical NotesabstractEffective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refinement. To address this gap, we propose Note2Chat, a note-driven framework that trains LLMs to conduct structured history taking and diagnosis by learning from widely available medical notes. Instead of relying on scarce and sensitive dialogue data, we convert real-world medical notes into high-quality doctor-patient dialogues using a decision tree-guided generation and refinement pipeline. We then propose a three-stage fine-tuning strategy combining supervised learning, simulated data augmentation, and preference learning. Furthermore, we propose a novel single-turn reasoning paradigm that reframes history taking as a sequence of single-turn reasoning problems. This design enhances interpretability and enables local supervision, dynamic adaptation, and greater sample efficiency. Experimental results show that our method substantially improves clinical reasoning, achieving gains of +16.9 F1 and +21.0 Top-1 diagnostic accuracy over GPT-4o. Yang Zhou 0017, Zhenting Sheng, Mingrui Tan, Yuting Song, Jun Zhou 0014, Yu Heng Kwan, Lian Leng Low, Yang Bai 0011, Yong Liu 0026 |
AAAI | 9 |
| 2026 | Self -adaptive neural networks for domain generalization in medical image segmentation
Yan Wang 0015, Zizhou Wang, Yangqin Feng, Lei Zhang 0005, Rick Siow Mong Goh, Yong Liu 0026, Liangli Zhen |
Expert Syst. Appl. | 6 |
| 2026 | Annotation-efficient medical image segmentation via cross-latent graphs and vector-quantized memory
Yanyu Xu 0001, Menghan Zhou, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Li-Zhen Cui 0001 |
Medical Image Anal. | 6 |
| 2026 | Improving Learning of New Diseases Through Knowledge-Enhanced Initialization for Federated Adapter TuningabstractIn healthcare, federated learning (FL) is a widely adopted framework that enables privacy-preserving collaboration among medical institutions. With large foundation models (FMs) demonstrating impressive capabilities, using FMs in FL through cost-efficient adapter tuning has become a popular approach. Given the rapidly evolving healthcare environment, it is crucial for individual clients to quickly adapt to new tasks or diseases by tuning adapters while drawing upon past experiences. In this work, we introduce Federated Knowledge-Enhanced Initialization (FedKEI), a novel framework that leverages cross-client and cross-task transfer from past knowledge to generate informed initializations for learning new tasks with adapters. FedKEI begins with a global clustering process at the server to generalize knowledge across tasks, followed by the optimization of aggregation weights across clusters (inter-cluster weights) and within each cluster (intra-cluster weights) to personalize knowledge transfer for each new task. To facilitate more effective learning of the inter- and intra-cluster weights, we adopt a bi-level optimization scheme that collaboratively learns the global intra-cluster weights across clients and optimizes the local inter-cluster weights toward each client's task objective. Extensive experiments on three benchmark datasets of different modalities, including dermatology, chest X-rays, and retinal OCT, demonstrate FedKEI's advantage in adapting to new diseases compared to state-of-the-art methods. Danni Peng, Yuan Wang 0008, Kangning Cai, Peiyan Ning, Jiming Xu, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei, Huazhu Fu |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Single-Domain Generalization via Path Flatness-Aware Optimization of Loss LandscapesabstractDomain generalization (DG) methods traditionally rely on multiple source domains to achieve the robust performance across unseen target domains. However, single-DG (SDG) presents a more practical paradigm by learning from a single source domain, addressing scenarios where access to multiple domains is limited. While existing SDG approaches primarily focus on data augmentation and style transfer techniques to enhance the model robustness, these methods often incur substantial computational overhead and may inadequately capture the complexity of real-world domain shifts. In this article, we propose path flatness-aware optimization (PFO), an optimization framework that addresses the fundamental challenges of SDG. Unlike conventional approaches that rely on the synthetic data generation, PFO identifies and exploits regions of flat minima within the optimization landscape of deep neural networks. The framework employs an iterative optimization strategy to construct a path through the parameter space along which an ensemble of candidate models achieves the minimal empirical risk. The initialization of this optimization path is achieved through the strategic interconnection of model instances, each originating from carefully selected anchor points that are computationally determined through the systematic analysis of classification decision manifolds. This optimization path serves as a mechanism for implicit distribution alignment between source and target domains within the loss landscape, consequently enhancing the model's capacity for cross-DG. Empirical evaluation on multiple benchmark datasets demonstrates significant performance improvements in cross-DG, validating the efficacy of our approach. Zizhou Wang, Yan Wang 0015, Yangqin Feng, Jiawei Du 0002, Joey Tianyi Zhou, Rick Siow Mong Goh, Yong Liu 0026, Liangli Zhen |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2026 | $AiRacleX$: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt GenerationabstractDecentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose$AiRacleX$, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models (LLMs). Our approach begins with domain-specific knowledge extraction, where an LLM model synthesizes precise insights about price oracle vulnerabilities, eliminating the need for profound expertise from developers or auditors. This knowledge forms the foundation for a second LLM model to generate structured, context-aware Chain-of-Thought prompts, which guide a third LLM model in accurately identifying manipulation patterns in smart contracts. We evaluate$AiRacleX$on 60 known vulnerabilities from 44 real-world DeFi exploits and Code4rena projects spanning 2021-2023. The results show that$AiRacleX$achieves a 2.58 times improvement in recall over the state-of-the-art GPTScan, with comparable precision. Our framework also demonstrates strong extensibility and efficiency, and supports deployment with open-source LLMs to enhance security and reduce operational cost. Yuan Wang 0008, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh, David Lo 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | VQA4CIR: Boosting Composed Image Retrieval with Visual Question AnsweringabstractAlbeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performance of CIR. The resulting VQA4CIR is a post-processing approach and can be directly plugged into existing CIR methods. Given the top-C retrieved images by a CIR method, VQA4CIR aims to decrease the adverse effect of the failure retrieval results being inconsistent with the relative caption. To find the retrieved images inconsistent with the relative caption, we resort to the "QA generation → VQA" self-verification pipeline. For QA generation, we suggest fine-tuning LLM (e.g., LLaMA) to generate several pairs of questions and answers from each relative caption. We then fine-tune LVLM (e.g., LLaVA) to obtain the VQA model. By feeding the retrieved image and question to the VQA model, one can find the images inconsistent with relative caption when the answer by VQA is inconsistent with the answer in the QA pair. Consequently, the CIR performance can be boosted by modifying the ranks of inconsistently retrieved images. Experimental results show that our proposed method outperforms state-of-the-art CIR methods on the CIRR and Fashion-IQ datasets. Chun-Mei Feng 0001, Yang Bai 0011, Tao Luo 0014, Zhen Li 0026, Salman Khan 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026 |
AAAI | 8 |
| 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter TuningabstractPersonalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignoring their previous updates, which may result in suboptimal personalized model learning. To bridge this gap, we propose a novel framework termed pFedSeq, designed for personalizing adapters to fine-tune a foundation model in FL. In pFedSeq, the server maintains and trains a sequential learner, which processes a sequence of past adapter updates from clients and generates calibrations for personalized adapters. To effectively capture the cross-client and cross-step relations hidden in previous updates and generate high-performing personalized adapters, pFedSeq adopts the powerful selective state space model (SSM) as the architecture of sequential learner. Through extensive experiments on four public benchmark datasets, we demonstrate the superiority of pFedSeq over state-of-the-art PFL methods. Danni Peng, Yuan Wang 0008, Huazhu Fu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
AAAI | 5 |
| 2025 | History-Aware and Dynamic Client Contribution in Federated LearningabstractFederated Learning (FL) is a collaborative machine learning (ML) approach, where multiple clients participate in training an ML model without exposing their private data. Fair and accurate assessment of client contributions facilitates incentive allocation in FL and encourages diverse clients to participate in a unified model training. Existing methods for contribution assessment adopts a co-operative game-theoretic concept, called Shapley value, but under restricted assumptions, e.g., all clients’ participating in all epochs or at least in one epoch of FL. We propose a history-aware client contribution assessment framework, called FLContrib, where client-participation is dynamic, i.e., a subset of clients participates in each epoch. The theoretical underpinning of FLContrib is based on the Markovian training process of FL. Under this setting, we directly apply the linearity property of Shapley value and compute a historical timeline of client contributions. Considering the possibility of a limited computational budget, we propose a two-sided fairness criteria to schedule Shapley value computation in a subset of epochs. Empirically, FLContrib is efficient and consistently accurate in estimating contribution across multiple utility functions. As a practical application, we apply FLContrib to detect dishonest clients in FL based on historical Shaplee values. Bishwamittra Ghosh, Debabrota Basu, Huazhu Fu, Yuan Wang 0008, Renuga Kanagavelu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
ECAI | 7 |
| 2025 | On the Importance of Language-driven Representation Learning for Heterogeneous Federated LearningabstractNon-Independent and Identically Distributed (Non-IID) training data significantly challenge federated learning (FL), impairing the performance of the global model in distributed frameworks. Inspired by the superior performance and generalizability of language-driven representation learning in centralized settings, we explore its potential to enhance FL for handling non-IID data. In specific, this paper introduces FedGLCL, a novel language-driven FL framework for image-text learning that uniquely integrates global language and local image features through contrastive learning, offering a new approach to tackle non-IID data in FL. FedGLCL redefines FL by avoiding separate local training models for each client. Instead, it uses contrastive learning to harmonize local image features with global textual data, enabling uniform feature learning across different local models. The utilization of a pre-trained text encoder in FedGLCL serves a dual purpose: it not only reduces the variance in local feature representations within FL by providing a stable and rich language context but also aids in mitigating overfitting, particularly to majority classes, by leveraging broad linguistic knowledge. Extensive experiments show that FedGLCL significantly outperforms state-of-the-art FL algorithms across different non-IID scenarios. Yunlu Yan, Chun-Mei Feng 0001, Wangmeng Zuo, Salman Khan 0001, Yong Liu 0026, Lei Zhu 0003 |
ICLR | 5 |
| 2025 | Federated Residual Low-Rank Adaptation of Large Language ModelsabstractLow-Rank Adaptation (LoRA) presents an effective solution for federated fine-tuning of Large Language Models (LLMs), as it substantially reduces communication overhead. However, a straightforward combination of FedAvg and LoRA results in suboptimal performance, especially under data heterogeneity. We noted this stems from both intrinsic (i.e., constrained parameter space) and extrinsic (i.e., client drift) limitations, which hinder it effectively learn global knowledge. In this work, we proposed a novel Federated Residual Low-Rank Adaption method, namely FRLoRA, to tackle above two limitations. It directly sums the weight of the global model parameters with a residual low-rank matrix product (\ie, weight change) during the global update step, and synchronizes this update for all local models. By this, FRLoRA performs global updates in a higher-rank parameter space, enabling a better representation of complex knowledge structure. Furthermore, FRLoRA reinitializes the local low-rank matrices with the principal singular values and vectors of the pre-trained weights in each round, to calibrate their inconsistent convergence, thereby mitigating client drift. Our extensive experiments demonstrate that FRLoRA consistently outperforms various state-of-the-art FL methods across nine different benchmarks in natural language understanding and generation under different FL scenarios. Yunlu Yan, Chun-Mei Feng 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026, Lei Zhu 0003 |
ICLR | 5 |
| 2025 | AdvMIM: Adversarial Masked Image Modeling for Semi-supervised Medical Image Segmentation
Lei Zhu 0003, Jun Zhou 0014, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (16) | 4 |
| 2025 | Ad2Mix: Adversarial and Adaptive Mixup for Unsupervised Domain AdaptationabstractTransformer has recently gained tremendous popularity in unsupervised domain adaptation tasks due to its superior generalization ability. State-of-the-art methods leverage mixup to build an intermediate domain to reduce domain gap. However, such strategy becomes less effective when the domain gap becomes large, as the domain gap between intermediate domain and source domain is not minimized and the constructed intermediate domain is non informative. How to address the adaptation problem when domain gap becomes large is an important research problem in domain adaptation. In this paper, we propose an adversarial and adaptive mixup (Ad2mix) framework which gradually aligns the intermediate domain towards source domain to fully unleash the potential of both the transformer architecture and mixup to address the large domain gap problem. Specifically, we formulate a general framework for intermediate domain learning with mixup. We propose adversarial mixup with a specially designed mixup alike adversarial adaptation operation to reduce the domain gap between the intermediate domain and source domain. To construct an informative intermediate domain, unlike existing methods which utilize a Beta distribution to generate mixup coefficients to interpolate source and target data, we adaptively assign mixup coefficient for each target data instance based on their transferability and discriminativity information. Our framework creates a natural curriculum of intermediate domains from near source domain to near target domain for gradual adaptation. Extensive experimental studies and evaluations on three public domain adaptation benchmark datasets and one medical domain adaptation task demonstrate the superiority of our framework. Lei Zhu 0003, Yanyu Xu 0001, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
WACV | 3 |
| 2025 | Diffusion-Enhanced Test-Time Adaptation with Text and Image Augmentation
Chun-Mei Feng 0001, Yuanyang He, Jian Zou 0005, Salman Khan 0001, Huan Xiong, Zhen Li 0026, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026 |
Int. J. Comput. Vis. | 9 |
| 2025 | MDDIP: Efficient single-layer pixel-based metasurface denoising via deep image prior
Manna Dai, Feng Yang 0011, Joyjit Chattoraj, Yingzhi Xia, Xinxing Xu, Weijiang Zhao, My Ha Dao, Yong Liu 0026 |
Knowl. Based Syst. | 9 |
| 2025 | Text to Image for Multi-Label Image Recognition With Joint Prompt-Adapter LearningabstractBenefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be similar to the corresponding text features, the modality gap remains a nontrivial issue and limits image recognition performance of TaI. Using multi-label image recognition (MLR) as an example, we present a novel method, called T2I-PAL to tackle the modality gap issue when using only text captions for PEFT. The core design of T2I-PAL is to leverage pre-trained text-to-image generation models to generate photo-realistic and diverse images from text captions, thereby reducing the modality gap. To further enhance MLR, T2I-PAL incorporates a class-wise heatmap and learnable prototypes. This aggregates local similarities, making the representation of local visual features more robust and informative for multi-label recognition. For better PEFT, we further combine both prompt tuning and adapter learning to enhance classification performance. T2I-PAL offers significant advantages: it eliminates the need for fully semantically annotated training images, thereby reducing the manual annotation workload, and it preserves the intrinsic mode of the CLIP model, allowing for seamless integration with any existing CLIP framework. Extensive experiments on multiple benchmarks, including MS-COCO, VOC2007, and NUS-WIDE, show that our T2I-PAL can boost recognition performance by 3.47% in average above the top-ranked state-of-the-art methods. Chun-Mei Feng 0001, Kai Yu 0009, Xinxing Xu, Salman Khan 0001, Rick Siow Mong Goh, Wangmeng Zuo, Yong Liu 0026 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 9 |
| 2025 | Toward Reliable Medical Image Segmentation by Modeling Evidential Calibrated UncertaintyabstractMedical image segmentation is critical for disease diagnosis and treatment assessment. However, concerns regarding the reliability of segmentation regions persist among clinicians, mainly attributed to the absence of confidence assessment, robustness, and calibration to accuracy. To address this, we introduce deep evidential segmentation model (DEviS), an easily implementable foundational model that seamlessly integrates into various medical image segmentation networks. DEviS not only enhances the calibration and robustness of baseline segmentation accuracy but also provides high-efficiency uncertainty estimation for reliable predictions. By leveraging subjective logic theory, we explicitly model probability and uncertainty for medical image segmentation. Here, the Dirichlet distribution parameterizes the distribution of probabilities for different classes of the segmentation results. To generate calibrated predictions and uncertainty, we develop a trainable calibrated uncertainty penalty. Furthermore, DEviS incorporates an uncertainty-aware filtering (UAF) module, which designs the metric of uncertainty-calibrated error to filter out-of-distribution (OOD) data. We conducted validation studies on publicly available datasets, including ISIC2018, KiTS2021, LiTS2017, and BraTS2019, to assess the accuracy and robustness of different backbone segmentation models enhanced by DEviS, as well as the efficiency and reliability of uncertainty estimation. Additionally, two potential clinical trials were conducted using the UAF module. The clinical application conducted on the Johns Hopkins OCT and Duke OCT-DME datasets demonstrated the effectiveness of the model in filtering OOD data. The second trial evaluated its efficacy in filtering high-quality data on the FIVES datasets. At last, the proposed DEviS method was extended to semi-supervised medical image segmentation, where it exhibited strong robustness under noisy conditions. Our code has been released in https://github.com/Cocofeat/DEviS. Ke Zou, Ling Huang 0003, Xuedong Yuan, Xiaojing Shen, Meng Wang 0038, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Cybern. | 9 |
| 2025 | Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound DevicesabstractHandheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA. Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Medical Imaging | 11 |
| 2025 | Continuous Disentangled Joint Space Learning for Domain GeneralizationabstractDomain generalization (DG) aims to learn a model on one or multiple observed source domains that can generalize to unseen target test domains. Previous approaches have focused on extracting domain-invariant information from multiple source domains, but domain-specific information is also closely tied to semantics in individual domains and is not well-suited for generalization to the target domain. In this article, we propose a novel DG method called continuous disentangled joint space learning (CJSL), which leverages both domain-invariant and domain-specific information for more effective DG. The key idea behind CJSL is to formulate and learn a continuous joint space (CJS) for domain-specific representations from source domains through iterative feature disentanglement. This learned CJS can then be used to simulate domain-specific representations for test samples from a mixture of multiple domains via Monte Carlo sampling during the inference stage. Unlike existing approaches, which exploit domain-invariant feature vectors only or aim to learn a universal domain-specific feature extractor, we simulate domain-specific representations via sampling the latent vectors in the learned CJS for the test sample to fully use the power of multiple domain-specific classifiers for robust prediction. Empirical results demonstrate that CJSL outperforms 19 state-of-the-art (SOTA) methods on seven benchmarks, indicating the effectiveness of our proposed method. Zizhou Wang, Yan Wang 0015, Yangqin Feng, Jiawei Du 0002, Yong Liu 0026, Rick Siow Mong Goh, Liangli Zhen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | RLPeri: Accelerating Visual Perimetry Test with Reinforcement Learning and Convolutional Feature ExtractionabstractVisual perimetry is an important eye examination that helps detect vision problems caused by ocular or neurological conditions. During the test, a patient's gaze is fixed at a specific location while light stimuli of varying intensities are presented in central and peripheral vision. Based on the patient's responses to the stimuli, the visual field mapping and sensitivity are determined. However, maintaining high levels of concentration throughout the test can be challenging for patients, leading to increased examination times and decreased accuracy. In this work, we present RLPeri, a reinforcement learning-based approach to optimize visual perimetry testing. By determining the optimal sequence of locations and initial stimulus values, we aim to reduce the examination time without compromising accuracy. Additionally, we incorporate reward shaping techniques to further improve the testing performance. To monitor the patient's responses over time during testing, we represent the test's state as a pair of 3D matrices. We apply two different convolutional kernels to extract spatial features across locations as well as features across different stimulus values for each location. Through experiments, we demonstrate that our approach results in a 10-20% reduction in examination time while maintaining the accuracy as compared to state-of-the-art methods. With the presented approach, we aim to make visual perimetry testing more efficient and patient-friendly, while still providing accurate results. Tanvi Verma, Linh Le Dinh, Nicholas Tan, Xinxing Xu, Ching Yu Cheng, Yong Liu 0026 |
AAAI | 6 |
| 2024 | An Aggregation-Free Federated Learning for Tackling Data HeterogeneityabstractThe performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework, where clients update local models based on a global model aggregated by the server from the previous training round. This process can cause client drift, especially with significant cross-client data heterogeneity, impacting model performance and convergence of the FL algorithm. To address these challenges, we introduce FedAF, a novel aggregation-free FL algorithm. In this framework, clients collaboratively learn condensed data by leveraging peer knowledge, the server subsequently trains the global model using the condensed data and soft labels received from the clients. FedAF inherently avoids the issue of client drift, enhances the quality of condensed data amid notable data heterogeneity, and improves the global model performance. Extensive numerical studies on several popular benchmark datasets show FedAF surpasses various state-of-the-art FL algorithms in handling label-skew and feature-skew data heterogeneity, leading to superior global model accuracy and faster convergence. Yuan Wang 0008, Huazhu Fu, Renuga Kanagavelu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
CVPR | 5 |
| 2024 | Sentence-level Prompts Benefit Composed Image RetrievalabstractComposed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested to generate a pseudo-word token from the reference image, which is further integrated into the relative caption for CIR. However, these pseudo-word-based prompting methods have limitations when target image encompasses complex changes on reference image, e.g., object removal and attribute modification. In this work, we demonstrate that learning an appropriate sentence-level prompt for the relative caption (SPRC) is sufficient for achieving effective composed image retrieval. Instead of relying on pseudo- word-based prompts, we propose to leverage pretrained V-L models, e.g., BLIP-2, to generate sentence-level prompts. By concatenating the learned sentence-level prompt with the relative caption, one can readily use existing text-based image retrieval models to enhance CIR performance. Furthermore, we introduce both image-text contrastive loss and text prompt alignment loss to enforce the learning of suitable sentence-level prompts. Experiments show that our proposed method performs favorably against the state-of-the-art CIR methods on the Fashion-IQ and CIRR datasets. Yang Bai 0011, Xinxing Xu, Yong Liu 0026, Salman Khan 0001, Fahad Shahbaz Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng 0001 |
ICLR | 3 |
| 2024 | Memory-Efficient High-Resolution OCT Volume Synthesis with Cascaded Amortized Latent Diffusion Models
Xiao Ma 0011, Yuhan Zhang 0001, Songtao Yuan, Yong Liu 0026, Qiang Chen 0004, Huazhu Fu |
MICCAI (7) | 6 |
| 2024 | Towards a Benchmark for Colorectal Cancer Segmentation in Endorectal Ultrasound Videos: Dataset and Model Development
Yuncheng Jiang 0002, Yiwen Hu 0001, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Xuemei Tang, Yong Liu 0026, Shuguang Cui, Zhen Li 0026 |
MICCAI (8) | 8 |
| 2024 | MedSynth: Leveraging Generative Model for Healthcare Data Sharing
Renuga Kanagavelu, Madhav Walia, Yuan Wang 0008, Huazhu Fu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
MICCAI (12) | 6 |
| 2024 | Multi-Scale Region-Aware Implicit Neural Network for Medical Images Matting
Yanyu Xu 0001, Yingzhi Xia, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Xinxing Xu |
MICCAI (9) | 5 |
| 2024 | A New Perspective to Boost Performance Fairness For Medical Federated Learning
Yunlu Yan, Lei Zhu 0003, Yuexiang Li, Xinxing Xu, Rick Siow Mong Goh, Yong Liu 0026, Salman Khan 0001, Chun-Mei Feng 0001 |
MICCAI (10) | 6 |
| 2024 | UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
Kai Yu 0009, Yang Zhou 0017, Yang Bai 0011, Zhi Da Soh, Xinxing Xu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026 |
MICCAI (12) | 8 |
| 2024 | MedMLP: An Efficient MLP-Like Network for Zero-Shot Retinal Image Classification
Menghan Zhou, Yanyu Xu 0001, Zhi Da Soh, Huazhu Fu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026, Liangli Zhen |
MICCAI (3) | 7 |
| 2024 | Class Balance Matters to Active Class-Incremental LearningabstractFew-Shot Class-Incremental Learning has shown remarkable efficacy in efficient learning new concepts with limited annotations. Nevertheless, the heuristic few-shot annotations may not always cover the most informative samples, which largely restricts the capability of incremental learner. We aim to start from a pool of large-scale unlabeled data and then annotate the most informative samples for incremental learning. Based on this premise, Based on this purpose, this paper introduces the Active Class-Incremental Learning (ACIL). The objective of ACIL is to select the most informative samples from the unlabeled pool to effectively train an incremental learner, aiming to maximize the performance of the resulting model. Note that vanilla active learning algorithms suffer from class-imbalanced distribution among annotated samples, which restricts the ability of incremental learning. To achieve both class balance and informativeness in chosen samples, we propose Class-Balanced Selection (CBS) strategy. Specifically, we first cluster the features of all unlabeled images into multiple groups. Then for each cluster, we employ greedy selection strategy to ensure that the Gaussian distribution of the sampled features closely matches the Gaussian distribution of all unlabeled features within the cluster.Our CBS can be plugged and played into those CIL methods which are based on pretrained models with prompts tunning technique.Extensive experiments under ACIL protocol across five diverse datasets demonstrate that CBS outperforms both random selection and other SOTA active learning approaches. Zitong Huang, Yuanze Li, Bowen Dong 0001, Erjin Zhou, Yong Liu 0026, Rick Siow Mong Goh, Chun-Mei Feng 0001, Wangmeng Zuo |
ACM Multimedia | 6 |
| 2024 | BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-RaysabstractMedical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX. Yang Zhou 0017, Tan Li Hui Faith, Yanyu Xu 0001, Sicong Leng, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh |
NeurIPS | 6 |
| 2024 | A surrogate-assisted extended generative adversarial network for parameter optimization in free-form metasurface design
Manna Dai, Feng Yang 0011, Joyjit Chattoraj, Yingzhi Xia, Xinxing Xu, Weijiang Zhao, My Ha Dao, Yong Liu 0026 |
Neural Networks | 9 |
| 2024 | MedNAS: Multiscale Training-Free Neural Architecture Search for Medical Image AnalysisabstractDeep neural networks have demonstrated impressive results in medical image analysis, but designing suitable architectures for each specific task is expertise-dependent and time-consuming. Neural architecture search (NAS) offers an effective means of discovering architectures. It has been highly successful in numerous applications, particularly in natural image classification. Yet, medical images possess unique characteristics, such as small regions and a wide variety of lesion sizes, that differentiate them from natural images. Furthermore, most current NAS methods struggle with high computational costs, especially when dealing with high-resolution image datasets. In this paper, we present a novel evolutionary neural architecture search method called Multi-Scale Training-Free Neural Architecture Search to address these challenges. Specifically, to accommodate the broad range of lesion region sizes in disease diagnosis, we develop a new reduction cell search space that enables the search algorithm to explicitly identify the optimal scale combination for multi-scale feature extraction. To overcome the issue of high computational costs, we utilize training-free indicators as performance measures for candidate architectures, which allows us to search for the optimal architecture more efficiently. More specifically, by considering the capability and simplicity of various networks, we formulate a multi-objective optimization problem that involves two training-free indicators and model complexity for candidate architectures. Extensive experiments on a large medical image benchmark and a publicly available breast cancer detection dataset are conducted. The empirical results demonstrate that our MSTF-NAS outperforms both human-designed architectures and current state-of-the-art NAS algorithms on both datasets, indicating the effectiveness of our proposed method. Yan Wang 0015, Liangli Zhen, Jianwei Zhang 0016, Miqing Li, Lei Zhang 0005, Zizhou Wang, Yangqin Feng, Yu Xue 0003, Xiao Wang 0004, Zheng Chen 0012, Tao Luo 0014, Rick Siow Mong Goh, Yong Liu 0026 |
IEEE Trans. Evol. Comput. | 13 |
| 2024 | Geometric Correspondence-Based Multimodal Learning for Ophthalmic Image AnalysisabstractColor fundus photography (CFP) and Optical coherence tomography (OCT) images are two of the most widely used modalities in the clinical diagnosis and management of retinal diseases. Despite the widespread use of multimodal imaging in clinical practice, few methods for automated diagnosis of eye diseases utilize correlated and complementary information from multiple modalities effectively. This paper explores how to leverage the information from CFP and OCT images to improve the automated diagnosis of retinal diseases. We propose a novel multimodal learning method, named geometric correspondence-based multimodal learning network (GeCoM-Net), to achieve the fusion of CFP and OCT images. Specifically, inspired by clinical observations, we consider the geometric correspondence between the OCT slice and the CFP region to learn the correlated features of the two modalities for robust fusion. Furthermore, we design a new feature selection strategy to extract discriminative OCT representations by automatically selecting the important feature maps from OCT slices. Unlike the existing multimodal learning methods, GeCoM-Net is the first method that formulates the geometric relationships between the OCT slice and the corresponding region of the CFP image explicitly for CFP and OCT fusion. Experiments have been conducted on a large-scale private dataset and a publicly available dataset to evaluate the effectiveness of GeCoM-Net for diagnosing diabetic macular edema (DME), impaired visual acuity (VA) and glaucoma. The empirical results show that our method outperforms the current state-of-the-art multimodal learning methods by improving the AUROC score 0.4%, 1.9% and 2.9% for DME, VA and glaucoma detection, respectively. Yan Wang 0015, Liangli Zhen, Tien-En Tan, Huazhu Fu, Yangqin Feng, Zizhou Wang, Xinxing Xu, Rick Siow Mong Goh, Yipin Ng, Claire Calhoun, Gavin Siew Wei Tan, Jennifer K. Sun, Yong Liu 0026, Daniel S. W. Ting |
IEEE Trans. Medical Imaging | 13 |
| 2023 | Learning Federated Visual Prompt in Null Space for MRI ReconstructionabstractFederated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI protocols, insufficient local training data, and limited communication bandwidth inevitably impair global model convergence and updating. In this paper, we propose a new algorithm, FedPR, to learn federated visual prompts in the null space of global prompt for MRI reconstruction. FedPR is a new federated paradigm that adopts a powerful pre-trained model while only learning and communicating the prompts with few learnable parameters, thereby significantly reducing communication costs and achieving competitive performance on limited local data. Moreover, to deal with catastrophic forgetting caused by data heterogeneity, FedPR also updates efficient federated visual prompts that project the local prompts into an approximate null space of the global prompt, thereby suppressing the interference of gradients on the server performance. Extensive experiments on federated MRI show that FedPR significantly outperforms state-of-the-art FL algorithms with < 6% of communication costs when given the limited amount of local training data. Chun-Mei Feng 0001, Bangjun Li, Xinxing Xu, Yong Liu 0026, Huazhu Fu, Wangmeng Zuo |
CVPR | 4 |
| 2023 | Diverse Data Augmentation with Diffusions for Effective Test-time Prompt TuningabstractBenefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper, we focus on a particular setting of learning adaptive prompts on the fly for each test sample from an unseen new domain, which is known as test-time prompt tuning (TPT). Existing TPT methods typically rely on data augmentation and confidence selection. However, conventional data augmentation techniques, e.g., random resized crops, suffers from the lack of data diversity, while entropy-based confidence selection alone is not sufficient to guarantee prediction fidelity. To address these issues, we propose a novel TPT method, named DiffTPT, which leverages pre-trained diffusion models to generate diverse and informative new data. Specifically, we incorporate augmented data by both conventional method and pre-trained stable diffusion to exploit their respective merits, improving the model’s ability to adapt to unknown new test data. Moreover, to ensure the prediction fidelity of generated data, we introduce a cosine similarity-based filtration technique to select the generated data with higher similarity to the single test sample. Our experiments on test datasets with distribution shifts and unseen categories demonstrate that DiffTPT improves the zero-shot accuracy by an average of 5.13% compared to the state-of-the-art TPT method. Chun-Mei Feng 0001, Kai Yu 0009, Yong Liu 0026, Salman Khan 0001, Wangmeng Zuo |
ICCV | 3 |
| 2023 | Generative Gradient Inversion via Over-Parameterized Networks in Federated LearningabstractFederated learning has gained recognitions as a secure approach for safeguarding local private data in collaborative learning. But the advent of gradient inversion research has posed significant challenges to this premise by enabling a third-party to recover groundtruth images via gradients. While prior research has predominantly focused on low-resolution images and small batch sizes, this study highlights the feasibility of reconstructing complex images with high resolutions and large batch sizes. The success of the proposed method is contingent on constructing an over-parameterized convolutional network, so that images are generated before fitting to the gradient matching requirement. Practical experiments demonstrate that the proposed algorithm achieves high-fidelity image recovery, surpassing state-of-the-art competitors that commonly fail in more intricate scenarios. Consequently, our study shows that local participants in a federated learning system are vulnerable to potential data leakage issues. Source code is available at https://github.com/czhang024/CI-Net. Chi Zhang 0123, Xiaoman Zhang, Ekanut Sotthiwat, Yanyu Xu 0001, Ping Liu 0004, Liangli Zhen, Yong Liu 0026 |
ICCV | 7 |
| 2023 | Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment
Zhihao Chen 0004, Yang Zhou 0017, Junting Zhao, Gideon Ooi, Lionel Tim-Ee Cheng, Choon Hua Thng, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
MICCAI (7) | 10 |
| 2023 | Category-Independent Visual Explanation for Medical Deep Network Understanding
Yiming Qian, Liangzhi Li 0004, Huazhu Fu, Meng Wang 0001, Qingsheng Peng, Ching Yu Cheng, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
MICCAI (2) | 8 |
| 2023 | Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
MICCAI (2) | 7 |
| 2023 | Minimal-Supervised Medical Image Segmentation via Vector Quantization Memory
Yanyu Xu 0001, Menghan Zhou, Yangqin Feng, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (3) | 7 |
| 2023 | Contrastive domain adaptation with consistency match for automated pneumonia diagnosis
Yangqin Feng, Zizhou Wang, Xinxing Xu, Yan Wang 0015, Huazhu Fu, Shaohua Li 0003, Liangli Zhen, Xiaofeng Lei, Yingnan Cui, Jordan Zheng Ting Sim, Yonghan Ting, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Cher Heng Tan |
Medical Image Anal. | 13 |
| 2022 | CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowabstractOptical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft. Xiuchao Sui, Shaohua Li 0003, Xue Geng, Yan Wu 0002, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh, Hongyuan Zhu 0002 |
CVPR | 6 |
| 2022 | Airfoil Inverse Design using Conditional Generative Adversarial NetworksabstractCreating an aerodynamic shape, like an airfoil wing, requires many factors to be considered, especially aerodynamic properties such as its lift-to-drag ratio (L/D). Currently, generating feasible airfoil shapes usually requires computationally expensive tools, such as Computational Fluid Dynamics (CFD). In recent years, increasing work has been directed to utilizing machine learning algorithms to synthesize accurate airfoil shapes while reducing the required computational cost. Generative Adversarial Network (GAN) is one of many algorithms to see success in airfoil shape optimization and is shown to generate good airfoils given a small set of training examples. This paper focuses on implementing a conditional GAN (cGAN) based framework with various filters for airfoil inverse design problem. By labelling the training dataset with aerodynamic characteristics separated by pre-defined thresholds to lift-to-drag ratio (L/D) and shape area, the class labels will be able to guide the network to generate different classes of airfoils influenced by these characteristics. Together with layers of Savitzky-Golay (SG) filter and B-Spline Interpolation, the developed model was shown to achieve good performance in generating new airfoils. In addition, we explored the viability of adding Wasserstein loss from Wasserstein GAN into the network architecture, forming a cWGAN-GP. Testing results showed that cWGAN-GP was able to achieve better performance for a specific airfoil class. Xavier Tan, Manna Dai, Joyjit Chattoraj, Yong Liu 0026, Xinxing Xu, My Ha Dao, Feng Yang 0011 |
ICARCV | 4 |
| 2022 | Adversarial multimodal fusion with attention mechanism for skin lesion classification using clinical and dermoscopic images
Yan Wang 0015, Yangqin Feng, Lei Zhang 0005, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Liangli Zhen |
Medical Image Anal. | 5 |
| 2022 | Deep Supervised Domain Adaptation for Pneumonia Diagnosis From Chest X-Ray ImagesabstractPneumonia is one of the most common treatable causes of death, and early diagnosis allows for early intervention. Automated diagnosis of pneumonia can therefore improve outcomes. However, it is challenging to develop high-performance deep learning models due to the lack of well-annotated data for training. This paper proposes a novel method, called Deep Supervised Domain Adaptation (DSDA), to automatically diagnose pneumonia from chest X-ray images. Specifically, we propose to transfer the knowledge from a publicly available large-scale source dataset (ChestX-ray14) to a well-annotated but small-scale target dataset (the TTSH dataset). DSDA aligns the distributions of the source domain and the target domain according to the underlying semantics of the training samples. It includes two task-specific sub-networks for the source domain and the target domain, respectively. These two sub-networks share the feature extraction layers and are trained in an end-to-end manner. Unlike most existing domain adaptation approaches that perform the same tasks in the source domain and the target domain, we attempt to transfer the knowledge from a multi-label classification task in the source domain to a binary classification task in the target domain. To evaluate the effectiveness of our method, we compare it with several existing peer methods. The experimental results show that our method can achieve promising performance for automated pneumonia diagnosis. Yangqin Feng, Xinxing Xu, Yan Wang 0015, Xiaofeng Lei, Soo Kng Teo, Jordan Zheng Ting Sim, Yonghan Ting, Liangli Zhen, Joey Tianyi Zhou, Yong Liu 0026, Cher Heng Tan |
IEEE J. Biomed. Health Informatics | 10 |
| 2021 | Medical Image Segmentation using Squeeze-and-Expansion TransformersabstractMedical image segmentation is important for computer-aided diagnosis. Good segmentation demands the model to see the big picture and fine details simultaneously, i.e., to learn image features that incorporate large context while keep high spatial resolutions. To approach this goal, the most widely used methods -- U-Net and variants, extract and fuse multi-scale features. However, the fused features still have small "effective receptive fields" with a focus on local image cues, limiting their performance. In this work, we propose Segtran, an alternative segmentation framework based on transformers, which have unlimited "effective receptive fields" even at high feature resolutions. The core of Segtran is a novel Squeeze-and-Expansion transformer: a squeezed attention block regularizes the self attention of transformers, and an expansion block learns diversified representations. Additionally, we propose a new positional encoding scheme for transformers, imposing a continuity inductive bias for images. Experiments were performed on 2D and 3D medical image segmentation tasks: optic disc/cup segmentation in fundus images (REFUGE'20 challenge), polyp segmentation in colonoscopy images, and brain tumor segmentation in MRI scans (BraTS'19 challenge). Compared with representative existing methods, Segtran consistently achieved the highest segmentation accuracy, and exhibited good cross-domain generalization capabilities. Shaohua Li 0003, Xiuchao Sui, Xiangde Luo, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh |
IJCAI | 5 |
| 2021 | Few-Shot Domain Adaptation with Polymorphic Transformers
Shaohua Li 0003, Xiuchao Sui, Huazhu Fu, Xiangde Luo, Yangqin Feng, Xinxing Xu, Yong Liu 0026, Daniel S. W. Ting, Rick Siow Mong Goh |
MICCAI (2) | 8 |
| 2021 | Partially-Supervised Learning for Vessel Segmentation in Ocular Images
Yanyu Xu 0001, Xinxing Xu, Shenghua Gao, Rick Siow Mong Goh, Daniel S. W. Ting, Yong Liu 0026 |
MICCAI (1) | 7 |
| 2020 | RoboCoDraw: Robotic Avatar Drawing with GAN-Based Style Transfer and Time-Efficient Path OptimizationabstractRobotic drawing has become increasingly popular as an entertainment and interactive tool. In this paper we present RoboCoDraw, a real-time collaborative robot-based drawing system that draws stylized human face sketches interactively in front of human users, by using the Generative Adversarial Network (GAN)-based style transfer and a Random-Key Genetic Algorithm (RKGA)-based path optimization. The proposed RoboCoDraw system takes a real human face image as input, converts it to a stylized avatar, then draws it with a robotic arm. A core component in this system is the AvatarGAN proposed by us, which generates a cartoon avatar face image from a real human face. AvatarGAN is trained with unpaired face and avatar images only and can generate avatar images of much better likeness with human face images in comparison with the vanilla CycleGAN. After the avatar image is generated, it is fed to a line extraction algorithm and converted to sketches. An RKGA-based path optimization algorithm is applied to find a time-efficient robotic drawing path to be executed by the robotic arm. We demonstrate the capability of RoboCoDraw on various face images using a lightweight, safe collaborative robot UR5. Tianying Wang, Wei Qi Toh, Hao Zhang 0048, Xiuchao Sui, Shaohua Li 0003, Yong Liu 0026 |
AAAI | 6 |
| 2019 | Coverage Path Planning using Path Primitive Sampling and Primitive Coverage Graph for Visual InspectionabstractPlanning the path to gather the surface information of the target objects is crucial to improve the efficiency of and reduce the overall cost, for visual inspection applications with Unmanned Aerial Vehicles (UAVs). Coverage Path Planning (CPP) problem is often formulated for these inspection applications because of the coverage requirement. Traditionally, researchers usually plan and optimize the viewpoints to capture the surface information first, and then optimize the path to visit the selected viewpoints. In this paper, we propose a novel planning method to directly sample and plan the inspection path for a camera-equipped UAV to acquire visual and geometric information of the target structures as a video stream setting in complex 3D environment. The proposed planning method first generates via-points and path primitives around the target object by using sampling methods based on voxel dilation and subtraction. A novel Primitive Coverage Graph (PCG) is then proposed to encode the topological information, flying distances, and visibility information, with the sampled via-points and path primitives. Finally graph search is performed to find the resultant path in the PCG to complete the inspection task with the coverage requirements. The effectiveness of the proposed method is demonstrated through simulation and field tests in this paper. Di Deng, Yong Liu 0026, Kenji Shimada |
IROS | 4 |
| 2019 | Multi-Instance Multi-Scale CNN for Medical Image Classification
Shaohua Li 0003, Yong Liu 0026, Xiuchao Sui, Cheng Chen 0008, Gabriel Tjio, Daniel S. W. Ting, Rick Siow Mong Goh |
MICCAI (4) | 2 |
| 2019 | AnomalyNet: An Anomaly Detection Network for Video SurveillanceabstractSparse coding-based anomaly detection has shown promising performance, of which the keys are feature learning, sparse representation, and dictionary learning. In this paper, we propose a new neural network for anomaly detection (termed AnomalyNet) by deeply achieving feature learning, sparse representation, and dictionary learning in three joint neural processing blocks. Specifically, to learn better features, we design a motion fusion block accompanied by a feature transfer block to enjoy the advantages of eliminating noisy background, capturing motion, and alleviating data deficiency. Furthermore, to address some disadvantages (e.g., nonadaptive updating) of the existing sparse coding optimizers and embrace the merits of neural network (e.g., parallel computing), we design a novel recurrent neural network to learn sparse representation and dictionary by proposing an adaptive iterative hard-thresholding algorithm (adaptive ISTA) and reformulating the adaptive ISTA as a new long short-term memory (LSTM). To the best of our knowledge, this could be one of the first works to bridge the$\ell _{1}$ -solver and LSTM and may provide novel insight into understanding LSTM and model-based optimization (or named differentiable programming), as well as sparse coding-based anomaly detection. Extensive experiments show the state-of-the-art performance of our method in the abnormal events detection task. Joey Tianyi Zhou, Jiawei Du 0002, Hongyuan Zhu 0002, Xi Peng 0001, Yong Liu 0026, Rick Siow Mong Goh |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2018 | SC2Net: Sparse LSTMs for Sparse CodingabstractThe iterative hard-thresholding algorithm (ISTA) is one of the most popular optimization solvers to achieve sparse codes. However, ISTA suffers from following problems: 1) ISTA employs non-adaptive updating strategy to learn the parameters on each dimension with a fixed learning rate. Such a strategy may lead to inferior performance due to the scarcity of diversity; 2) ISTA does not incorporate the historical information into the updating rules, and the historical information has been proven helpful to speed up the convergence. To address these challenging issues, we propose a novel formulation of ISTA (named as adaptive ISTA) by introducing a novel \textit{adaptive momentum vector}. To efficiently solve the proposed adaptive ISTA, we recast it as a recurrent neural network unit and show its connection with the well-known long short term memory (LSTM) model. With a new proposed unit, we present a neural network (termed SC2Net) to achieve sparse codes in an end-to-end manner. To the best of our knowledge, this is one of the first works to bridge the $\ell_1$-solver and LSTM, and may provide novel insights in understanding model-based optimization and LSTM. Extensive experiments show the effectiveness of our method on both unsupervised and supervised tasks. Joey Tianyi Zhou, Kai Di, Jiawei Du 0002, Xi Peng 0001, Hao Yang 0033, Sinno Jialin Pan, Ivor W. Tsang, Yong Liu 0026, Zheng Qin 0004, Rick Siow Mong Goh |
AAAI | 8 |
| 2018 | A Learning-based Approach for Error Compensation of Industrial Manipulator with Hybrid ModelabstractThe industrial robot usually has high repeatability but relatively lower accuracy. Therefore, error compensation plays a pivotal role in many industrial robotic applications with high accuracy requirement. In this paper, we present a novel computational method that utilizes a hybrid model that consists of Local Product-Of-Exponential (POE) and Gaussian Process Regression (GPR) to compensate the positioning errors of the industrial robotic manipulator for high accuracy industrial robotic applications. Specifically in the proposed method, the Local POE calibration method is first applied to calibrate the robot forward kinematic model to reduce the geometric error. Then the GPR is applied to learn the inverse kinematic model to further compensate the residual error in task space. We also demonstrate the robustness and effectiveness of our proposed method by showing the reduction of norm pose error by up to 37.2%, compared to the existing methods with multiple datasets. Joey Tianyi Zhou, Yong Liu 0026, Pey Yuen Tao, Guilin Yang |
ICARCV | 4 |
| 2018 | EMRShare: A Cross-Organizational Medical Data Sharing and Management Framework Using Permissioned BlockchainabstractWith the development of information and storage technologies, electronic document recording has become an unalterable trend, which transforms the way that people store, access and operate the data generated in various applications. Healthcare is the leader application domain that pioneers in usage of electronic medical records (EMRs). Cross-organizational EMRs' sharing has many constructive effects in motivating the domain innovation, introducing better domain understanding and overall the domain intelligence. However, the privacy concern, trust issue as well as the sophisticated legal regulation of the sensitive EMRs' use leads to inefficiency in the data sharing process. In this paper, we propose a cross-organizational medical data sharing framework based on permissioned blockchain technology, named “EMRShare“, to resolve the trust concern existing in EMRs sharing practice among different participants like patients, clinicians and researchers, and other relevant parties such as the insurance agent and government, to make medical data sharing and access secure, efficient, transparent, immutable, traceable and auditable. A working prototype system is implemented to demonstrate the key features for the cross-organizational medical data sharing and access management. The objective of this work targets at explaining the essential design considerations along with the working principle and operation logics using the blockchain technology to facilitate the medical data sharing in a highly-cooperative healthcare ecosystem. Zengxiang Li, Yong Liu 0026, Thanarit Lertwuthikarn, Rick Siow Mong Goh |
ICPADS | 3 |
| 2018 | Cost-Efficient and Latency-Aware Workflow Scheduling Policy for Container-Based SystemsabstractContainer technology is being adopted to simplify workflow execution. In this paper, we investigate a workflow scheduling policy for container-based systems. A workflow, representing an application, consists of a set of tasks. Each task can be executed in a container within a virtual machine (VM), where the container packaging the function for the task should be loaded into the VM before task execution. To reduce the workflow execution time and the network bandwidth consumption, we propose a cost-efficient and latency-aware workflow scheduling algorithm that strategically loads the containers into VMs and executes the tasks on the VMs. The algorithm is based on “Stretch Out and Compact”, which can stretch out the tasks along the resources by critical path analysis and then find the inefficient slots within the computing resources and eventually compact the tasks into those slots. We introduce a concept of “virtual task” into the algorithm, where container loading is regarded as a virtual task that should be executed before the real task. The introduction of the virtual task can be more effective in finding the inefficient slots for the compaction, thus resulting in a more efficient workflow scheduling policy. Simulation results show that compared to the algorithms that fully or selectively load the dockers, the proposed algorithm can achieve less execution time while saving network bandwidth consumption for loading dockers. Yong Liu 0026, Long Wang 0005, Zengxiang Li, Rick Siow Mong Goh |
ICPADS | 2 |
| 2017 | Performance Modelling and Cost Effective Execution for Distributed Graph Processing on Configurable VMsabstractGraph Processing has been widely used to capture complex data dependency and uncover relationship insights. Due to the ever-growing graph scale and algorithm complexity, distributed graph processing has become more and more popular. In this paper, we investigate how to balance performance and cost for large scale graph processing on configurable virtual machines (VMs). We analyze the system architecture and implementation details of a Pregel-like distributed graph processing framework and develop a system-aware model to predict the execution time. Consequently, cost effective execution scenarios are recommended by selecting a certain number of VMs with specified capability subject to the predefined resource price and user preference. Experiments using synthetic and real world graphs have verified that system-aware model can achieve much higher prediction accuracy than popular machine-learning models which treat graph processing framework as a black box. As a result, the recommended execution scenarios have comparable cost efficiency to the optimal scenarios. Zengxiang Li, Shen Ren, Yong Liu 0026, Zheng Qin 0004, Rick Siow Mong Goh, Gurusamy Mohan |
CCGrid | 4 |
| 2011 | Fast spanning tree reconnection mechanism for resilient Metro Ethernet networks
Yong Liu 0026, Gurusamy Mohan, Kee Chaing Chua |
Comput. Networks | 2 |
| 2011 | Local restoration with multiple spanning trees in metro ethernet networksabstractEthernet is becoming a preferred technology to be extended to metropolitan area networks (MANs) due to its low cost, simplicity, and ubiquity. However, current Ethernet lacks a fast failure recovery mechanism as it reconstructs the spanning tree after the failure is detected, which commonly requires tens of seconds. Some fast failure-handling methods based on multiple spanning trees have been proposed in the literature, but these approaches are either centralized or require periodic message broadcasting over the entire network. In this paper, we propose a local restoration mechanism for metro Ethernet using multiple spanning trees, which is distributed and fast and does not need failure notification. Upon failure of a single link, the upstream switch locally restores traffic to preconfigured backup spanning trees. We propose two restoration approaches, connection-based and destination-based, to select backup trees. We formulate the tree preconfiguration problem that includes working spanning tree assignment and backup spanning tree configuration. We prove that the preconfiguration problem is NP-complete and develop an integer linear programming model. We also develop heuristic algorithms for each restoration approach to reduce the computation complexity. To evaluate the effectiveness of our heuristic algorithms, we carry out the simulation on grid and random networks. The simulation results show that our heuristic algorithms have comparable performance close to the optimal solutions, and both restoration approaches can efficiently utilize the network bandwidth to handle single link failures. Gurusamy Mohan, Kee Chaing Chua, Yong Liu 0026 |
IEEE/ACM Trans. Netw. | 4 |
| 2010 | Achieving High Performance Burst Transmission for Bursty Traffic using Optical Burst Chain Switching in WDM NetworksabstractIn Optical Burst Switching (OBS) network architecture, edge nodes assemble bursts and send them to the core network arbitrarily, which will lead to inevitable collisions and low network utilization in the core network. Our objective is to achieve high performance burst transmission for bursty traffic while avoiding the use of wavelength converters and excessive optical buffers. To achieve this objective, we propose a novel optical burst chain switching (OBCS) mechanism, which combines the merits of optical circuit switching and optical burst switching with affordable signaling overhead. In the proposed mechanism, switching unit is a burst chain which consists of multiple non-periodic and non-consecutive bursts in one wavelength. Theoretical analysis for the throughput and queuing delay of the proposed scheme is carried out and is also verified by simulation results. We present extensive simulation results to demonstrate its superior performance over OBS networks with/without wavelength converters. Yong Liu 0026, Kee Chaing Chua, Gurusamy Mohan |
IEEE Trans. Commun. | 1 |
| 2009 | Handling Double-Link Failures in Metro Ethernet Networks Using Fast Spanning Tree ReconnectionabstractEthernet is becoming a preferred technology to be deployed in metro domain due to its low cost, simplicity and ubiquity. However, traditional spanning tree based Ethernet protocol does not meet the requirement for metro area networks in terms of network resilience. In the work of Qiu et al. (2009), we proposed a fast spanning tree reconnection (FSTR) mechanism for metro Ethernet networks to handle single link failures. Upon failure of a link on a spanning tree, FSTR mechanism activates a reconnect-link to reconnect the broken spanning tree. FSTR mechanism has the features of fast recovery, simplicity, and guaranteed protection. However, when more than one link fail in the network, the FSTR mechanism would generate unexpected loops and cannot function properly. In this paper, we propose a fast spanning tree reconnection mechanism to handle double-link failures with protection grade guarantees. The mechanism is distributed and can alleviate the problem in previous FSTR mechanism. We formulate the reconnect-link pre-configuration problem for double-link failures as an integer linear programming problem. Through numerical results we demonstrate that the proposed mechanism can satisfy the protection grade required for each connection by efficiently utilizing the network capacity. Gurusamy Mohan, Kee Chaing Chua, Yong Liu 0026 |
GLOBECOM | 4 |
| 2009 | Fast Spanning Tree Reconnection for Resilient Metro Ethernet NetworksabstractEthernet is becoming a preferred technology to be deployed in metro domain due to its low cost, simplicity and ubiquity. However, spanning tree based Ethernet protocol does not meet the requirement for Metro Area Networks in terms of network resilience, despite the advancement of Ethernet standardization and commercialization. In this paper, we propose a fast spanning tree reconnection (FSTR) mechanism for Metro Ethernet networks to handle any single link failure, which has features of fast recovery, backup capacity guarantees and ease of implementation. Upon failure of a link on a spanning tree, a distributed failure recovery protocol is activated to reconnect the broken spanning tree using a reconnect-link not on the spanning tree. We present the details of the protocol, including failure notification and forwarding table reconfiguration procedures. The pre-configuration of the reconnect-links to reconnect each spanning tree is formulated as an integer linear programming (ILP) problem. The optimization results show that with lower implementation cost, fast spanning tree reconnection mechanism can achieve comparative or considerably better performance than other resilient mechanisms for Metro Ethernet networks. Yong Liu 0026, Gurusamy Mohan, Kee Chaing Chua |
ICC | 2 |
| 2009 | Multipath traffic engineering in WDM optical burst switching networksabstractIn this paper, we investigate the problem of multipath traffic engineering in optical burst switching (OBS) networks. The main goal of this work is to minimize burst loss rate in the network by adaptively balancing the burst traffic among multiple paths based on the measurement and analysis of path congestion. We develop three distributed traffic splitting algorithms by considering the special features of OBS networks. Among these three algorithms, one algorithm considers burst level traffic distribution, and the other two algorithms consider flow level traffic distribution to avoid packet reordering problem. The proposed multipath traffic engineering algorithms achieve significant performance improvement in reducing burst loss probability in a distributed manner in OBS networks. Further, through theoretical analysis and simulations, we show that the proposed algorithms converge quickly on a sudden traffic increase in the network. Yong Liu 0026, Gurusamy Mohan, Kee Chaing Chua |
IEEE Trans. Commun. | 1 |