Rick Siow Mong Goh

dblp:24/2670 · also Rich Siow Mong Goh · DBLP profile ↗
← Back
126ranked-venue papers
0as first author
70since 2021 · last 2026
0000-0001-9116-1595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 45 · 12 since 2021Artificial intelligence and machine learning · 41 · 30 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 20 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Self -adaptive neural networks for domain generalization in medical image segmentation
Yan Wang 0015, Zizhou Wang, Yangqin Feng, Lei Zhang 0005, Rick Siow Mong Goh, Yong Liu 0026, Liangli Zhen
Expert Syst. Appl.5
2026 Is quantum optimization ready? An effort towards neural network compression using adiabatic quantum computing
Zhehui Wang, Benjamin Chen Ming Choong, Tian Huang, Daniel Gerlinghoff, Rick Siow Mong Goh, Cheng Liu 0008, Tao Luo 0014
Future Gener. Comput. Syst.5
2026 Annotation-efficient medical image segmentation via cross-latent graphs and vector-quantized memory
Yanyu Xu 0001, Menghan Zhou, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Li-Zhen Cui 0001
Medical Image Anal.5
2026 Addressing Client Drift in Federated Learning via Class-Prototype Similarity Distillation and Adaptive Mask
abstract
Federated learning (FL) enables multiple clients to learn collaboratively in a distributed way, allowing for privacy protection. However, the real-world nonindependent and identically distributed (non-IID) data will lead to client drift, which degrades the performance of FL. Interestingly, we find that the logit difference between the local and global models increases as the model is continuously updated, which is the primary factor behind performance degradation. This is mainly due to catastrophic forgetting caused by non-IID data between clients. To alleviate this problem, we propose a new algorithm, named FedCSD, a class-prototype similarity distillation in a federated framework to align the logits of local and global models. FedCSD does not simply transfer global knowledge to local clients, as an insufficiently trained global model cannot provide reliable knowledge, i.e., class similarity information, and its wrong soft labels will mislead the optimization of local models. Concretely, FedCSD leverages the similarity between local logits and the global prototype to refine the global logits, thereby enhancing its class similarity information. Furthermore, FedCSD adopts an adaptive mask to filter out the terrible soft labels of the global models, thereby preventing them from misleading local optimization. Extensive experiments demonstrate the superiority of our method over the state-of-the-art FL approaches in various non-IID settings. Code is publicly available at https://github.com/IAMJackYan/FedCSD.
Yunlu Yan, Chun-Mei Feng 0001, Mang Ye, Wangmeng Zuo, Ping Li 0016, Rick Siow Mong Goh, Lei Zhu 0003, C. L. Philip Chen
IEEE Trans. Cybern.6
2026 BriDe Arbitrager: Enhancing Arbitrage in Ethereum 2.0 via Bribery-Enabled Delayed Block Production
abstract
The advent of Ethereum 2.0 has introduced significant changes, particularly the shift to Proof-of-Stake consensus. This change presents new opportunities and challenges for arbitrage. Amidst these changes, we introduce BriDe Arbitrager, a novel tool designed for Ethereum 2.0 that leveragesBribery-driven attacks toDelay block production and increase arbitrage gains. The main idea is to allow malicious proposers to delay block production by bribing validators/proposers, thereby gaining more time to identify arbitrage opportunities. Through analysing the bribery process, we design an adaptive bribery strategy. Additionally, we propose a Delayed Transaction Ordering Algorithm to leverage the delayed time to amplify arbitrage profits for malicious proposers. To ensure fairness and automate the bribery process, we design and implement a bribery smart contract and a bribery client. As a result, BriDe Arbitrager enables adversaries controlling a limited ($\lt 1/4$) fraction of the voting powers to delay block production via bribery and arbitrage more profit. Extensive experimental results based on Ethereum historical transactions demonstrate that BriDe Arbitrager yields an average of 8.78 ETH (16,687.88 USD) daily profits. Furthermore, our approach does not trigger any slashing mechanisms and remains effective even under Proposer Builder Separation and other potential mechanisms will be adopted by Ethereum.
Hulin Yang, Jin Zhang 0001, Alia Asheralieva, Qingsong Wei, Rick Siow Mong Goh
IEEE Trans. Dependable Secur. Comput.6
2026 Improving Learning of New Diseases Through Knowledge-Enhanced Initialization for Federated Adapter Tuning
abstract
In healthcare, federated learning (FL) is a widely adopted framework that enables privacy-preserving collaboration among medical institutions. With large foundation models (FMs) demonstrating impressive capabilities, using FMs in FL through cost-efficient adapter tuning has become a popular approach. Given the rapidly evolving healthcare environment, it is crucial for individual clients to quickly adapt to new tasks or diseases by tuning adapters while drawing upon past experiences. In this work, we introduce Federated Knowledge-Enhanced Initialization (FedKEI), a novel framework that leverages cross-client and cross-task transfer from past knowledge to generate informed initializations for learning new tasks with adapters. FedKEI begins with a global clustering process at the server to generalize knowledge across tasks, followed by the optimization of aggregation weights across clusters (inter-cluster weights) and within each cluster (intra-cluster weights) to personalize knowledge transfer for each new task. To facilitate more effective learning of the inter- and intra-cluster weights, we adopt a bi-level optimization scheme that collaboratively learns the global intra-cluster weights across clients and optimizes the local inter-cluster weights toward each client's task objective. Extensive experiments on three benchmark datasets of different modalities, including dermatology, chest X-rays, and retinal OCT, demonstrate FedKEI's advantage in adapting to new diseases compared to state-of-the-art methods.
Danni Peng, Yuan Wang 0008, Kangning Cai, Peiyan Ning, Jiming Xu, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei, Huazhu Fu
IEEE Trans. Medical Imaging7
2026 Single-Domain Generalization via Path Flatness-Aware Optimization of Loss Landscapes
abstract
Domain generalization (DG) methods traditionally rely on multiple source domains to achieve the robust performance across unseen target domains. However, single-DG (SDG) presents a more practical paradigm by learning from a single source domain, addressing scenarios where access to multiple domains is limited. While existing SDG approaches primarily focus on data augmentation and style transfer techniques to enhance the model robustness, these methods often incur substantial computational overhead and may inadequately capture the complexity of real-world domain shifts. In this article, we propose path flatness-aware optimization (PFO), an optimization framework that addresses the fundamental challenges of SDG. Unlike conventional approaches that rely on the synthetic data generation, PFO identifies and exploits regions of flat minima within the optimization landscape of deep neural networks. The framework employs an iterative optimization strategy to construct a path through the parameter space along which an ensemble of candidate models achieves the minimal empirical risk. The initialization of this optimization path is achieved through the strategic interconnection of model instances, each originating from carefully selected anchor points that are computationally determined through the systematic analysis of classification decision manifolds. This optimization path serves as a mechanism for implicit distribution alignment between source and target domains within the loss landscape, consequently enhancing the model's capacity for cross-DG. Empirical evaluation on multiple benchmark datasets demonstrates significant performance improvements in cross-DG, validating the efficacy of our approach.
Zizhou Wang, Yan Wang 0015, Yangqin Feng, Jiawei Du 0002, Joey Tianyi Zhou, Rick Siow Mong Goh, Yong Liu 0026, Liangli Zhen
IEEE Trans. Neural Networks Learn. Syst.6
2026 $AiRacleX$: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
abstract
Decentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose$AiRacleX$, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models (LLMs). Our approach begins with domain-specific knowledge extraction, where an LLM model synthesizes precise insights about price oracle vulnerabilities, eliminating the need for profound expertise from developers or auditors. This knowledge forms the foundation for a second LLM model to generate structured, context-aware Chain-of-Thought prompts, which guide a third LLM model in accurately identifying manipulation patterns in smart contracts. We evaluate$AiRacleX$on 60 known vulnerabilities from 44 real-world DeFi exploits and Code4rena projects spanning 2021-2023. The results show that$AiRacleX$achieves a 2.58 times improvement in recall over the state-of-the-art GPTScan, with comparable precision. Our framework also demonstrates strong extensibility and efficiency, and supports deployment with open-source LLMs to enhance security and reduce operational cost.
Yuan Wang 0008, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh, David Lo 0001
IEEE Trans. Serv. Comput.5
2025 VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering
abstract
Albeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performance of CIR. The resulting VQA4CIR is a post-processing approach and can be directly plugged into existing CIR methods. Given the top-C retrieved images by a CIR method, VQA4CIR aims to decrease the adverse effect of the failure retrieval results being inconsistent with the relative caption. To find the retrieved images inconsistent with the relative caption, we resort to the "QA generation → VQA" self-verification pipeline. For QA generation, we suggest fine-tuning LLM (e.g., LLaMA) to generate several pairs of questions and answers from each relative caption. We then fine-tune LVLM (e.g., LLaVA) to obtain the VQA model. By feeding the retrieved image and question to the VQA model, one can find the images inconsistent with relative caption when the answer by VQA is inconsistent with the answer in the QA pair. Consequently, the CIR performance can be boosted by modifying the ranks of inconsistently retrieved images. Experimental results show that our proposed method outperforms state-of-the-art CIR methods on the CIRR and Fashion-IQ datasets.
Chun-Mei Feng 0001, Yang Bai 0011, Tao Luo 0014, Zhen Li 0026, Salman Khan 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026
AAAI7
2025 Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning
abstract
Personalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignoring their previous updates, which may result in suboptimal personalized model learning. To bridge this gap, we propose a novel framework termed pFedSeq, designed for personalizing adapters to fine-tune a foundation model in FL. In pFedSeq, the server maintains and trains a sequential learner, which processes a sequence of past adapter updates from clients and generates calibrations for personalized adapters. To effectively capture the cross-client and cross-step relations hidden in previous updates and generate high-performing personalized adapters, pFedSeq adopts the powerful selective state space model (SSM) as the architecture of sequential learner. Through extensive experiments on four public benchmark datasets, we demonstrate the superiority of pFedSeq over state-of-the-art PFL methods.
Danni Peng, Yuan Wang 0008, Huazhu Fu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei
AAAI6
2025 History-Aware and Dynamic Client Contribution in Federated Learning
abstract
Federated Learning (FL) is a collaborative machine learning (ML) approach, where multiple clients participate in training an ML model without exposing their private data. Fair and accurate assessment of client contributions facilitates incentive allocation in FL and encourages diverse clients to participate in a unified model training. Existing methods for contribution assessment adopts a co-operative game-theoretic concept, called Shapley value, but under restricted assumptions, e.g., all clients’ participating in all epochs or at least in one epoch of FL. We propose a history-aware client contribution assessment framework, called FLContrib, where client-participation is dynamic, i.e., a subset of clients participates in each epoch. The theoretical underpinning of FLContrib is based on the Markovian training process of FL. Under this setting, we directly apply the linearity property of Shapley value and compute a historical timeline of client contributions. Considering the possibility of a limited computational budget, we propose a two-sided fairness criteria to schedule Shapley value computation in a subset of epochs. Empirically, FLContrib is efficient and consistently accurate in estimating contribution across multiple utility functions. As a practical application, we apply FLContrib to detect dishonest clients in FL based on historical Shaplee values.
Bishwamittra Ghosh, Debabrota Basu, Huazhu Fu, Yuan Wang 0008, Renuga Kanagavelu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei
ECAI8
2025 Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
abstract
Learned activation functions in models like Kolmogorov-Arnold Networks (KANs) outperform fixed-activation architectures in terms of accuracy and interpretability; however, their computational complexity poses critical challenges for energy-constrained edge AI deployments. Conventional CPUs/GPUs incur prohibitive latency and power costs when evaluating higher order activations, limiting deployability under ultra-tight energy budgets. We address this via a reconfigurable lookup architecture with edge FPGAs. By coupling fine-grained quantization with adaptive lookup tables, our design minimizes energy-intensive arithmetic operations while preserving activation fidelity. FPGA reconfigurability enables dynamic hardware specialization for learned functions, a key advantage for edge systems that require post-deployment adaptability. Evaluations using KANs - where unique activation functions play a critical role—demonstrate that our FPGA-based design achieves superior computational speed and over 104times higher energy efficiency compared to edge CPUs and GPUs, while maintaining matching accuracy and minimal footprint overhead. This breakthrough positions our approach as a practical enabler for energy-critical edge AI, where computational intensity and power constraints traditionally preclude the use of adaptive activation networks.
Mengyuan Yin, Benjamin Chen Ming Choong, Chuping Qu, Rick Siow Mong Goh, Weng-Fai Wong, Tao Luo 0014
ICCAD4
2025 Federated Residual Low-Rank Adaptation of Large Language Models
abstract
Low-Rank Adaptation (LoRA) presents an effective solution for federated fine-tuning of Large Language Models (LLMs), as it substantially reduces communication overhead. However, a straightforward combination of FedAvg and LoRA results in suboptimal performance, especially under data heterogeneity. We noted this stems from both intrinsic (i.e., constrained parameter space) and extrinsic (i.e., client drift) limitations, which hinder it effectively learn global knowledge. In this work, we proposed a novel Federated Residual Low-Rank Adaption method, namely FRLoRA, to tackle above two limitations. It directly sums the weight of the global model parameters with a residual low-rank matrix product (\ie, weight change) during the global update step, and synchronizes this update for all local models. By this, FRLoRA performs global updates in a higher-rank parameter space, enabling a better representation of complex knowledge structure. Furthermore, FRLoRA reinitializes the local low-rank matrices with the principal singular values and vectors of the pre-trained weights in each round, to calibrate their inconsistent convergence, thereby mitigating client drift. Our extensive experiments demonstrate that FRLoRA consistently outperforms various state-of-the-art FL methods across nine different benchmarks in natural language understanding and generation under different FL scenarios.
Yunlu Yan, Chun-Mei Feng 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026, Lei Zhu 0003
ICLR4
2025 AdvMIM: Adversarial Masked Image Modeling for Semi-supervised Medical Image Segmentation
Lei Zhu 0003, Jun Zhou 0014, Rick Siow Mong Goh, Yong Liu 0026
MICCAI (16)3
2025 Ad2Mix: Adversarial and Adaptive Mixup for Unsupervised Domain Adaptation
abstract
Transformer has recently gained tremendous popularity in unsupervised domain adaptation tasks due to its superior generalization ability. State-of-the-art methods leverage mixup to build an intermediate domain to reduce domain gap. However, such strategy becomes less effective when the domain gap becomes large, as the domain gap between intermediate domain and source domain is not minimized and the constructed intermediate domain is non informative. How to address the adaptation problem when domain gap becomes large is an important research problem in domain adaptation. In this paper, we propose an adversarial and adaptive mixup (Ad2mix) framework which gradually aligns the intermediate domain towards source domain to fully unleash the potential of both the transformer architecture and mixup to address the large domain gap problem. Specifically, we formulate a general framework for intermediate domain learning with mixup. We propose adversarial mixup with a specially designed mixup alike adversarial adaptation operation to reduce the domain gap between the intermediate domain and source domain. To construct an informative intermediate domain, unlike existing methods which utilize a Beta distribution to generate mixup coefficients to interpolate source and target data, we adaptively assign mixup coefficient for each target data instance based on their transferability and discriminativity information. Our framework creates a natural curriculum of intermediate domains from near source domain to near target domain for gradual adaptation. Extensive experimental studies and evaluations on three public domain adaptation benchmark datasets and one medical domain adaptation task demonstrate the superiority of our framework.
Lei Zhu 0003, Yanyu Xu 0001, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu
WACV4
2025 Diffusion-Enhanced Test-Time Adaptation with Text and Image Augmentation
Chun-Mei Feng 0001, Yuanyang He, Jian Zou 0005, Salman Khan 0001, Huan Xiong, Zhen Li 0026, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026
Int. J. Comput. Vis.8
2025 Self-distillation with model averaging
Xiaozhe Gu, Zixun Zhang, Rick Siow Mong Goh, Tao Luo 0014
Inf. Sci.4
2025 Text to Image for Multi-Label Image Recognition With Joint Prompt-Adapter Learning
abstract
Benefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be similar to the corresponding text features, the modality gap remains a nontrivial issue and limits image recognition performance of TaI. Using multi-label image recognition (MLR) as an example, we present a novel method, called T2I-PAL to tackle the modality gap issue when using only text captions for PEFT. The core design of T2I-PAL is to leverage pre-trained text-to-image generation models to generate photo-realistic and diverse images from text captions, thereby reducing the modality gap. To further enhance MLR, T2I-PAL incorporates a class-wise heatmap and learnable prototypes. This aggregates local similarities, making the representation of local visual features more robust and informative for multi-label recognition. For better PEFT, we further combine both prompt tuning and adapter learning to enhance classification performance. T2I-PAL offers significant advantages: it eliminates the need for fully semantically annotated training images, thereby reducing the manual annotation workload, and it preserves the intrinsic mode of the CLIP model, allowing for seamless integration with any existing CLIP framework. Extensive experiments on multiple benchmarks, including MS-COCO, VOC2007, and NUS-WIDE, show that our T2I-PAL can boost recognition performance by 3.47% in average above the top-ranked state-of-the-art methods.
Chun-Mei Feng 0001, Kai Yu 0009, Xinxing Xu, Salman Khan 0001, Rick Siow Mong Goh, Wangmeng Zuo, Yong Liu 0026
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small
abstract
Large language models (LLMs) have garnered substantial attention due to their promising applications in diverse domains. Nevertheless, the increasing size of LLMs comes with a significant surge in the computational requirements for training and deployment. Memristor crossbars have emerged as a promising solution, which demonstrated a small footprint and remarkably high energy efficiency in computer vision (CV) models. Memristors possess higher density compared to conventional memory technologies, making them highly suitable for effectively managing the extreme model size associated with LLMs. However, deploying LLMs on memristor crossbars faces three major challenges. First, the size of LLMs increases rapidly, already surpassing the capabilities of state-of-the-art memristor chips. Second, LLMs often incorporate multi-head attention blocks, which involve non-weight stationary multiplications that traditional memristor crossbars cannot support. Third, while memristor crossbars excel at performing linear operations, they are not capable of executing complex nonlinear operations in LLM such as softmax and layer normalization. To address these challenges, we present a novel architecture for the memristor crossbar that enables the deployment of state-of-the-art LLM on a single chip or package, eliminating the energy and time inefficiencies associated with off-chip communication. Our testing on BERT showed negligible accuracy loss. Compared to traditional memristor crossbars, our architecture achieves enhancements of up to in area overhead and in energy consumption. Compared to modern TPU/GPU systems, our architecture demonstrates at least a reduction in the area-delay product and a significant 69% energy consumption reduction.
Zhehui Wang, Tao Luo 0014, Cheng Liu 0008, Weichen Liu 0001, Rick Siow Mong Goh, Weng-Fai Wong
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Reliable Federated Disentangling Network for Non-IID Domain Feature
abstract
Federated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability.
Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu
IEEE Trans. Big Data7
2025 Toward Reliable Medical Image Segmentation by Modeling Evidential Calibrated Uncertainty
abstract
Medical image segmentation is critical for disease diagnosis and treatment assessment. However, concerns regarding the reliability of segmentation regions persist among clinicians, mainly attributed to the absence of confidence assessment, robustness, and calibration to accuracy. To address this, we introduce deep evidential segmentation model (DEviS), an easily implementable foundational model that seamlessly integrates into various medical image segmentation networks. DEviS not only enhances the calibration and robustness of baseline segmentation accuracy but also provides high-efficiency uncertainty estimation for reliable predictions. By leveraging subjective logic theory, we explicitly model probability and uncertainty for medical image segmentation. Here, the Dirichlet distribution parameterizes the distribution of probabilities for different classes of the segmentation results. To generate calibrated predictions and uncertainty, we develop a trainable calibrated uncertainty penalty. Furthermore, DEviS incorporates an uncertainty-aware filtering (UAF) module, which designs the metric of uncertainty-calibrated error to filter out-of-distribution (OOD) data. We conducted validation studies on publicly available datasets, including ISIC2018, KiTS2021, LiTS2017, and BraTS2019, to assess the accuracy and robustness of different backbone segmentation models enhanced by DEviS, as well as the efficiency and reliability of uncertainty estimation. Additionally, two potential clinical trials were conducted using the UAF module. The clinical application conducted on the Johns Hopkins OCT and Duke OCT-DME datasets demonstrated the effectiveness of the model in filtering OOD data. The second trial evaluated its efficacy in filtering high-quality data on the FIVES datasets. At last, the proposed DEviS method was extended to semi-supervised medical image segmentation, where it exhibited strong robustness under noisy conditions. Our code has been released in https://github.com/Cocofeat/DEviS.
Ke Zou, Ling Huang 0003, Xuedong Yuan, Xiaojing Shen, Meng Wang 0038, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu
IEEE Trans. Cybern.8
2025 Federated Pseudo Modality Generation for Incomplete Multi-Modal MRI Reconstruction
abstract
While multi-modal learning has been widely used for MRI reconstruction, it relies on paired multi-modal data, which is difficult to acquire in real clinical scenarios. Especially in the federated setting, there is a common issue that several medical institutions suffer from missing modalities or even only have single-modal data. Therefore, it is infeasible to deploy a standard federated learning framework in such conditions. In this paper, we propose a novel communication-efficient federated learning framework (namely Fed-PMG) to address the missing modality challenge in federated multi-modal MRI reconstruction. Specifically, we utilize a pseudo modality generation mechanism to recover the missing modality for each single-modal client by sharing the distribution information of the amplitude spectrum in frequency space. However, the step of sharing the original amplitude spectrum leads to heavy communication costs. To reduce the communication cost, we introduce a clustering scheme to project the set of amplitude spectrum into a finite number of cluster centroids and share them among the clients. With such an elaborate design, our approach can effectively complete the missing modality within an acceptable communication cost. Extensive experimental results demonstrate that our proposed method can outperform state-of-the-art methods and reach a performance similar to the ideal scenario (i.e., all clients have the full set of modalities).
Yunlu Yan, Chun-Mei Feng 0001, Yuexiang Li, Ping Li 0016, Rick Siow Mong Goh, Bai Ying Lei, Weiming Wang 0002, David Dagan Feng, Lei Zhu 0003
IEEE J. Biomed. Health Informatics5
2025 Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound Devices
abstract
Handheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA.
Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu
IEEE Trans. Medical Imaging10
2025 Generative Image Reconstruction From Gradients
abstract
In this article, we propose a method, generative image reconstruction from gradients (GIRG), for recovering training images from gradients in a federated learning (FL) setting, where privacy is preserved by sharing model weights and gradients rather than raw training data. Previous studies have shown the potential for revealing clients' private information or even pixel-level recovery of training images from shared gradients. However, existing methods are limited to low-resolution images and small batch sizes (BSs) or require prior knowledge about the client data. GIRG utilizes a conditional generative model to reconstruct training images and their corresponding labels from the shared gradients. Unlike previous generative model-based methods, GIRG does not require prior knowledge of the training data. Furthermore, GIRG optimizes the weights of the conditional generative model to generate highly accurate "dummy" images instead of optimizing the input vectors of the generative model. Comprehensive empirical results show that GIRG is able to recover high-resolution images with large BSs and can even recover images from the aggregation of gradients from multiple participants. These results reveal the vulnerability of current FL practices and call for immediate efforts to prevent inversion attacks in gradient-sharing-based collaborative training.
Ekanut Sotthiwat, Liangli Zhen, Chi Zhang 0123, Zengxiang Li, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.5
2025 Continuous Disentangled Joint Space Learning for Domain Generalization
abstract
Domain generalization (DG) aims to learn a model on one or multiple observed source domains that can generalize to unseen target test domains. Previous approaches have focused on extracting domain-invariant information from multiple source domains, but domain-specific information is also closely tied to semantics in individual domains and is not well-suited for generalization to the target domain. In this article, we propose a novel DG method called continuous disentangled joint space learning (CJSL), which leverages both domain-invariant and domain-specific information for more effective DG. The key idea behind CJSL is to formulate and learn a continuous joint space (CJS) for domain-specific representations from source domains through iterative feature disentanglement. This learned CJS can then be used to simulate domain-specific representations for test samples from a mixture of multiple domains via Monte Carlo sampling during the inference stage. Unlike existing approaches, which exploit domain-invariant feature vectors only or aim to learn a universal domain-specific feature extractor, we simulate domain-specific representations via sampling the latent vectors in the learned CJS for the test sample to fully use the power of multiple domain-specific classifiers for robust prediction. Empirical results demonstrate that CJSL outperforms 19 state-of-the-art (SOTA) methods on seven benchmarks, indicating the effectiveness of our proposed method.
Zizhou Wang, Yan Wang 0015, Yangqin Feng, Jiawei Du 0002, Yong Liu 0026, Rick Siow Mong Goh, Liangli Zhen
IEEE Trans. Neural Networks Learn. Syst.6
2025 Atomic Smart Contract Interoperability With High Efficiency via Cross-Chain Integrated Execution
abstract
With the development of Ethereum, numerous blockchains compatible with Ethereum's execution environment (i.e., Ethereum Virtual Machine, EVM) have emerged. Developers can leverage smart contracts to run various complex decentralized applications on top of blockchains. However, the increasing number of EVM-compatible blockchains has introduced significant challenges in cross-chain interoperability, particularly in ensuring efficiency and atomicity for the whole cross-chain application. Existing solutions areeither limited in guaranteeing overall atomicity for the cross-chain application, or inefficient due to the need for multiple rounds of cross-chain smart contract execution.To address this gap, we proposeIntegrateX, an efficient cross-chain interoperability system that ensures the overall atomicity of cross-chain smart contract invocations. The core idea is todeploy the logic required for cross-chain execution onto a single blockchain, where it can be executed in an integrated manner.This allows cross-chain applications to perform all cross-chain logic efficiently within the same blockchain.IntegrateXconsists of across-chain smart contract deployment protocoland across-chain smart contract integrated execution protocol.The former achieves efficient and secure cross-chain deployment by decoupling smart contract logic from state, and employing an off-chain cross-chain deployment mechanism combined with on-chain cross-chain verification. The latter ensures atomicity of cross-chain invocations through a 2PC-based mechanism, and enhances performance through transaction aggregation and fine-grained state lock. We implement a prototype ofIntegrateX. Extensive experiments demonstrate that it reduces up to 61.2% latency compared to the state-of-the-art baseline while maintaining low gas consumption.
Chaoyue Yin, Jin Zhang 0001, You Lin, Qingsong Wei, Rick Siow Mong Goh
IEEE Trans. Parallel Distributed Syst.6
2024 An Aggregation-Free Federated Learning for Tackling Data Heterogeneity
abstract
The performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework, where clients update local models based on a global model aggregated by the server from the previous training round. This process can cause client drift, especially with significant cross-client data heterogeneity, impacting model performance and convergence of the FL algorithm. To address these challenges, we introduce FedAF, a novel aggregation-free FL algorithm. In this framework, clients collaboratively learn condensed data by leveraging peer knowledge, the server subsequently trains the global model using the condensed data and soft labels received from the clients. FedAF inherently avoids the issue of client drift, enhances the quality of condensed data amid notable data heterogeneity, and improves the global model performance. Extensive numerical studies on several popular benchmark datasets show FedAF surpasses various state-of-the-art FL algorithms in handling label-skew and feature-skew data heterogeneity, leading to superior global model accuracy and faster convergence.
Yuan Wang 0008, Huazhu Fu, Renuga Kanagavelu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh
CVPR6
2024 Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
abstract
Recent advancements in neural network quantisation have yielded remarkable outcomes, with three-bit networks reaching state-of-the-art full-precision accuracy in complex tasks. These achievements present valuable opportunities for accelerating neural networks by computing in reduced precision. Implementing it on FPGAs can take advantage of bit-level reconfigurability, which is not available on conventional CPUs and GPUs. Simultaneously, the high data intensity of neural network processing has inspired computing-in-memory paradigms, including on FPGA platforms. By programming the effects of trained model weights as lookup operations in soft logic, the transfer of weight data from memory units can be avoided, alleviating the memory bottleneck. However, previous methods face poor scalability - the high logic utilisation limiting them to small networks/sub-networks of binary models with low accuracy. In this paper, we introduce Table Lookup Multiply-Accumulate (TLMAC) as a framework to compile and optimise quantised neural networks for scalable lookup-based processing. TLMAC clusters and maps unique groups of weights to lookup-based processing elements, enabling highly parallel computation while taking advantage of parameter redundancy. Further place and route algorithms are proposed to reduce LUT utilisation and routing congestion. We demonstrate that TLMAC significantly improves the scalability of previous related works. Our efficient logic mapping and high degree of reuse enables entire ImageNet-scale quantised models with full-precision accuracy to be implemented using lookup-based computing on one commercially available FPGA.
Daniel Gerlinghoff, Benjamin Chen Ming Choong, Rick Siow Mong Goh, Weng-Fai Wong, Tao Luo 0014
FPGA3
2024 Sentence-level Prompts Benefit Composed Image Retrieval
abstract
Composed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested to generate a pseudo-word token from the reference image, which is further integrated into the relative caption for CIR. However, these pseudo-word-based prompting methods have limitations when target image encompasses complex changes on reference image, e.g., object removal and attribute modification. In this work, we demonstrate that learning an appropriate sentence-level prompt for the relative caption (SPRC) is sufficient for achieving effective composed image retrieval. Instead of relying on pseudo- word-based prompts, we propose to leverage pretrained V-L models, e.g., BLIP-2, to generate sentence-level prompts. By concatenating the learned sentence-level prompt with the relative caption, one can readily use existing text-based image retrieval models to enhance CIR performance. Furthermore, we introduce both image-text contrastive loss and text prompt alignment loss to enforce the learning of suitable sentence-level prompts. Experiments show that our proposed method performs favorably against the state-of-the-art CIR methods on the Fashion-IQ and CIRR datasets.
Yang Bai 0011, Xinxing Xu, Yong Liu 0026, Salman Khan 0001, Fahad Shahbaz Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng 0001
ICLR7
2024 MedSynth: Leveraging Generative Model for Healthcare Data Sharing
Renuga Kanagavelu, Madhav Walia, Yuan Wang 0008, Huazhu Fu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh
MICCAI (12)7
2024 Multi-Scale Region-Aware Implicit Neural Network for Medical Images Matting
Yanyu Xu 0001, Yingzhi Xia, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Xinxing Xu
MICCAI (9)4
2024 A New Perspective to Boost Performance Fairness For Medical Federated Learning
Yunlu Yan, Lei Zhu 0003, Yuexiang Li, Xinxing Xu, Rick Siow Mong Goh, Yong Liu 0026, Salman Khan 0001, Chun-Mei Feng 0001
MICCAI (10)5
2024 UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
Kai Yu 0009, Yang Zhou 0017, Yang Bai 0011, Zhi Da Soh, Xinxing Xu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026
MICCAI (12)6
2024 MedMLP: An Efficient MLP-Like Network for Zero-Shot Retinal Image Classification
Menghan Zhou, Yanyu Xu 0001, Zhi Da Soh, Huazhu Fu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026, Liangli Zhen
MICCAI (3)5
2024 Class Balance Matters to Active Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning has shown remarkable efficacy in efficient learning new concepts with limited annotations. Nevertheless, the heuristic few-shot annotations may not always cover the most informative samples, which largely restricts the capability of incremental learner. We aim to start from a pool of large-scale unlabeled data and then annotate the most informative samples for incremental learning. Based on this premise, Based on this purpose, this paper introduces the Active Class-Incremental Learning (ACIL). The objective of ACIL is to select the most informative samples from the unlabeled pool to effectively train an incremental learner, aiming to maximize the performance of the resulting model. Note that vanilla active learning algorithms suffer from class-imbalanced distribution among annotated samples, which restricts the ability of incremental learning. To achieve both class balance and informativeness in chosen samples, we propose Class-Balanced Selection (CBS) strategy. Specifically, we first cluster the features of all unlabeled images into multiple groups. Then for each cluster, we employ greedy selection strategy to ensure that the Gaussian distribution of the sampled features closely matches the Gaussian distribution of all unlabeled features within the cluster.Our CBS can be plugged and played into those CIL methods which are based on pretrained models with prompts tunning technique.Extensive experiments under ACIL protocol across five diverse datasets demonstrate that CBS outperforms both random selection and other SOTA active learning approaches.
Zitong Huang, Yuanze Li, Bowen Dong 0001, Erjin Zhou, Yong Liu 0026, Rick Siow Mong Goh, Chun-Mei Feng 0001, Wangmeng Zuo
ACM Multimedia7
2024 BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays
abstract
Medical Vision-Language Pretraining (MedVLP) shows promise in learning generalizable and transferable visual representations from paired and unpaired medical images and reports. MedVLP can provide useful features to downstream tasks and facilitate adapting task-specific models to new setups using fewer examples. However, existing MedVLP methods often differ in terms of datasets, preprocessing, and finetuning implementations. This pose great challenges in evaluating how well a MedVLP method generalizes to various clinically-relevant tasks due to the lack of unified, standardized, and comprehensive benchmark. To fill this gap, we propose BenchX, a unified benchmark framework that enables head-to-head comparison and systematical analysis between MedVLP methods using public chest X-ray datasets. Specifically, BenchX is composed of three components: 1) Comprehensive datasets covering nine datasets and four medical tasks; 2) Benchmark suites to standardize data preprocessing, train-test splits, and parameter selection; 3) Unified finetuning protocols that accommodate heterogeneous MedVLP methods for consistent task adaptation in classification, segmentation, and report generation, respectively. Utilizing BenchX, we establish baselines for nine state-of-the-art MedVLP methods and found that the performance of some early MedVLP methods can be enhanced to surpass more recent ones, prompting a revisiting of the developments and conclusions from prior works in MedVLP. Our code are available at https://github.com/yangzhou12/BenchX.
Yang Zhou 0017, Tan Li Hui Faith, Yanyu Xu 0001, Sicong Leng, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh
NeurIPS7
2024 MedNAS: Multiscale Training-Free Neural Architecture Search for Medical Image Analysis
abstract
Deep neural networks have demonstrated impressive results in medical image analysis, but designing suitable architectures for each specific task is expertise-dependent and time-consuming. Neural architecture search (NAS) offers an effective means of discovering architectures. It has been highly successful in numerous applications, particularly in natural image classification. Yet, medical images possess unique characteristics, such as small regions and a wide variety of lesion sizes, that differentiate them from natural images. Furthermore, most current NAS methods struggle with high computational costs, especially when dealing with high-resolution image datasets. In this paper, we present a novel evolutionary neural architecture search method called Multi-Scale Training-Free Neural Architecture Search to address these challenges. Specifically, to accommodate the broad range of lesion region sizes in disease diagnosis, we develop a new reduction cell search space that enables the search algorithm to explicitly identify the optimal scale combination for multi-scale feature extraction. To overcome the issue of high computational costs, we utilize training-free indicators as performance measures for candidate architectures, which allows us to search for the optimal architecture more efficiently. More specifically, by considering the capability and simplicity of various networks, we formulate a multi-objective optimization problem that involves two training-free indicators and model complexity for candidate architectures. Extensive experiments on a large medical image benchmark and a publicly available breast cancer detection dataset are conducted. The empirical results demonstrate that our MSTF-NAS outperforms both human-designed architectures and current state-of-the-art NAS algorithms on both datasets, indicating the effectiveness of our proposed method.
Yan Wang 0015, Liangli Zhen, Jianwei Zhang 0016, Miqing Li, Lei Zhang 0005, Zizhou Wang, Yangqin Feng, Yu Xue 0003, Xiao Wang 0004, Zheng Chen 0012, Tao Luo 0014, Rick Siow Mong Goh, Yong Liu 0026
IEEE Trans. Evol. Comput.12
2024 Geometric Correspondence-Based Multimodal Learning for Ophthalmic Image Analysis
abstract
Color fundus photography (CFP) and Optical coherence tomography (OCT) images are two of the most widely used modalities in the clinical diagnosis and management of retinal diseases. Despite the widespread use of multimodal imaging in clinical practice, few methods for automated diagnosis of eye diseases utilize correlated and complementary information from multiple modalities effectively. This paper explores how to leverage the information from CFP and OCT images to improve the automated diagnosis of retinal diseases. We propose a novel multimodal learning method, named geometric correspondence-based multimodal learning network (GeCoM-Net), to achieve the fusion of CFP and OCT images. Specifically, inspired by clinical observations, we consider the geometric correspondence between the OCT slice and the CFP region to learn the correlated features of the two modalities for robust fusion. Furthermore, we design a new feature selection strategy to extract discriminative OCT representations by automatically selecting the important feature maps from OCT slices. Unlike the existing multimodal learning methods, GeCoM-Net is the first method that formulates the geometric relationships between the OCT slice and the corresponding region of the CFP image explicitly for CFP and OCT fusion. Experiments have been conducted on a large-scale private dataset and a publicly available dataset to evaluate the effectiveness of GeCoM-Net for diagnosing diabetic macular edema (DME), impaired visual acuity (VA) and glaucoma. The empirical results show that our method outperforms the current state-of-the-art multimodal learning methods by improving the AUROC score 0.4%, 1.9% and 2.9% for DME, VA and glaucoma detection, respectively.
Yan Wang 0015, Liangli Zhen, Tien-En Tan, Huazhu Fu, Yangqin Feng, Zizhou Wang, Xinxing Xu, Rick Siow Mong Goh, Yipin Ng, Claire Calhoun, Gavin Siew Wei Tan, Jennifer K. Sun, Yong Liu 0026, Daniel S. W. Ting
IEEE Trans. Medical Imaging8
2024 RCT: Resource Constrained Training for Edge AI
abstract
Efficient neural network training is essential for in situ training of edge artificial intelligence (AI) and carbon footprint reduction in general. Train neural network on the edge is challenging because there is a large gap between limited resources on edge and the resource requirement of current training methods. Existing training methods are based on the assumption that the underlying computing infrastructure has sufficient memory and energy supplies. These methods involve two copies of the model parameters, which is usually beyond the capacity of on-chip memory in processors. The data movement between off-chip and on-chip memory incurs large amounts of energy. We propose resource constrained training (RCT) to realize resource-efficient training for edge devices and servers. RCT only keeps a quantized model throughout the training so that the memory requirement for model parameters in training is reduced. It adjusts per-layer bitwidth dynamically to save energy when a model can learn effectively with lower precision. We carry out experiments with representative models and tasks in image classification, natural language processing, and crowd counting applications. Experiments show that on average, 8-15-bit weight update is sufficient for achieving SOTA performance in these applications. RCT saves 63.5%-80% memory for model parameters and saves more energy for communications. Through experiments, we observe that the common practice on the first/last layer in model compression does not apply to efficient training. Also, interestingly, the more challenging a dataset is, the lower bitwidth is required for efficient training.
Tian Huang, Tao Luo 0014, Ming Yan 0007, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.5
2024 Efficient Spiking Neural Networks With Radix Encoding
abstract
Spiking neural networks (SNNs) have advantages in latency and energy efficiency over traditional artificial neural networks (ANNs) due to their event-driven computation mechanism and the replacement of energy-consuming weight multiplication with addition. However, to achieve high accuracy, it usually requires long spike trains to ensure accuracy, usually more than 1000 time steps. This offsets the computation efficiency brought by SNNs because a longer spike train means a larger number of operations and larger latency. In this article, we propose a radix-encoded SNN, which has ultrashort spike trains. Specifically, it is able to use less than six time steps to achieve even higher accuracy than its traditional counterpart. We also develop a method to fit our radix encoding technique into the ANN-to-SNN conversion approach so that we can train radix-encoded SNNs more efficiently on mature platforms and hardware. Experiments show that our radix encoding can achieve 25× improvement in latency and 1.7% improvement in accuracy compared to the state-of-the-art method using the VGG-16 network on the CIFAR-10 dataset.
Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Joey Tianyi Zhou, Tao Luo 0014
IEEE Trans. Neural Networks Learn. Syst.3
2024 EDCompress: Energy-Aware Model Compression for Dataflows
abstract
Edge devices demand low energy consumption, cost, and small form factor. To efficiently deploy convolutional neural network (CNN) models on the edge device, energy-aware model compression becomes extremely important. However, existing work did not study this problem well because of the lack of considering the diversity of dataflow types in hardware architectures. In this article, we propose EDCompress (EDC), an energy-aware model compression method for various dataflows. It can effectively reduce the energy consumption of various edge devices, with different dataflow types. Considering the very nature of model compression procedures, we recast the optimization process to a multistep problem and solve it by reinforcement learning algorithms. We also propose a multidimensional multistep (MDMS) optimization method, which shows higher compressing capability than the traditional multistep method. Experiments show that EDC could improve 20x, 17x, and 26x energy efficiency in VGG-16, MobileNet, and LeNet-5 networks, respectively, with negligible loss of accuracy. EDC could also indicate the optimal dataflow type for specific neural networks in terms of energy consumption, which can guide the deployment of CNN on hardware.
Zhehui Wang, Tao Luo 0014, Rick Siow Mong Goh, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.3
2024 Optimizing for In-Memory Deep Learning With Emerging Memory Technology
abstract
In-memory deep learning executes neural network models where they are stored, thus avoiding long-distance communication between memory and computation units, resulting in considerable savings in energy and time. In-memory deep learning has already demonstrated orders of magnitude higher performance density and energy efficiency. The use of emerging memory technology (EMT) promises to increase density, energy, and performance even further. However, EMT is intrinsically unstable, resulting in random data read fluctuations. This can translate to nonnegligible accuracy loss, potentially nullifying the gains. In this article, we propose three optimization techniques that can mathematically overcome the instability problem of EMT. They can improve the accuracy of the in-memory deep learning model while maximizing its energy efficiency. Experiments show that our solution can fully recover most models' state-of-the-art (SOTA) accuracy and achieves at least an order of magnitude higher energy efficiency than the SOTA.
Zhehui Wang, Tao Luo 0014, Rick Siow Mong Goh, Wei Zhang 0012, Weng-Fai Wong
IEEE Trans. Neural Networks Learn. Syst.3
2023 MA-BERT: Towards Matrix Arithmetic-only BERT Inference by Eliminating Complex Non-Linear Functions
Neo Wei Ming, Zhehui Wang, Cheng Liu 0008, Rick Siow Mong Goh, Tao Luo 0014
ICLR4
2023 Category-Independent Visual Explanation for Medical Deep Network Understanding
Yiming Qian, Liangzhi Li 0004, Huazhu Fu, Meng Wang 0001, Qingsheng Peng, Ching Yu Cheng, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu
MICCAI (2)9
2023 Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu
MICCAI (2)6
2023 Minimal-Supervised Medical Image Segmentation via Vector Quantization Memory
Yanyu Xu 0001, Menghan Zhou, Yangqin Feng, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026
MICCAI (3)6
2023 Desire backpropagation: A lightweight training algorithm for multi-layer spiking neural networks based on spike-timing-dependent plasticity
Daniel Gerlinghoff, Tao Luo 0014, Rick Siow Mong Goh, Weng-Fai Wong
Neurocomputing3
2023 Contrastive domain adaptation with consistency match for automated pneumonia diagnosis
Yangqin Feng, Zizhou Wang, Xinxing Xu, Yan Wang 0015, Huazhu Fu, Shaohua Li 0003, Liangli Zhen, Xiaofeng Lei, Yingnan Cui, Jordan Zheng Ting Sim, Yonghan Ting, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Cher Heng Tan
Medical Image Anal.14
2023 DeepFire2: A Convolutional Spiking Neural Network Accelerator on FPGAs
abstract
Brain-inspired spiking neural networks (SNNs) replace the multiply-accumulate operations of traditional neural networks by integrate-and-fire neurons, with the goal of achieving greater energy efficiency. Specialized hardware implementations of those neurons clearly have advantages over general-purpose devices in terms of power and performance, but exhibit poor scalability when it comes to accelerating large neural networks. DeepFire2 introduces a hardware architecture which can map large network layers efficiently across multiple super logic regions in a multi-die FPGA. That gives more control over resource allocation and parallelism, benefiting both throughput and energy consumption. Avoiding the use of lookup tables to implement theANDoperations of an SNN, prevents the layer size to be limited by logic resources. A deep pipeline does not only lead to an increased clock speed of up to 600 MHz. We double the throughput and power efficiency compared to our previous version of DeepFire, which equates to an almost 10-fold improvement over other previous implementations. Importantly, we are able to deploy a large ImageNet model, while maintaining a throughput of over 1500 frames per second.
Myat Thu Linn Aung, Daniel Gerlinghoff, Chuping Qu, Tian Huang, Rick Siow Mong Goh, Tao Luo 0014, Weng-Fai Wong
IEEE Trans. Computers6
2023 Benchmarking Quantum(-Inspired) Annealing Hardware on Practical Use Cases
abstract
Quantum(-inspired) annealers show promise in solving combinatorial optimisation problems in practice. There has been extensive researches demonstrating the utility of D-Wave quantum annealer and quantum-inspired annealer, i.e., Fujitsu Digital Annealer on various applications, but few works are comparing these platforms. In this paper, we benchmark quantum(-inspired) annealers with three combinatorial optimisation problems ranging from generic scientific problems to complex problems in practical use. In the case where the problem size goes beyond the capacity of a quantum(-inspired) computer, we evaluate them in the context of decomposition. Experiments suggest that both annealers are effective on problems with small size and simple settings, but lose their utility when facing problems in practical size and settings. Decomposition methods extend the scalability of annealers, but they are still far away from practical use. Based on the experiments and comparison, we discuss the advantages and limitations of quantum(-inspired) annealers, as well as the research directions that may improve the utility and scalability of the these emerging computing technologies.
Tian Huang, Tao Luo 0014, Xiaozhe Gu, Rick Siow Mong Goh, Weng-Fai Wong
IEEE Trans. Computers5
2023 Simeuro: A Hybrid CPU-GPU Parallel Simulator for Neuromorphic Computing Chips
abstract
With the success of deep learning, there have been numerous efforts to build hardware for it. One approach that is gaining momentum is neuromorphic computing with spiking neural networks (SNNs), which are multiplication-free and open the possibility of using analog computing via novel technologies. However, to design effective and efficient hardware for such architectures, a fast and accurate software simulator is key. This article presents Simeuro, a fast and scalable system-level simulator for SNN models used in neuromorphic accelerators. The simulator uses spike-level details and configurable architectural constraints that are independent of the underlying hardware implementation. Simeuro supports a wide range of features including analog computing, novel memory (currently, RRAM is supported), and a full network-on-chip. The simulator can provide detailed simulation results such as routing statistics, energy consumption, delay, and accuracy of arbitrarily defined SNN architectures. Our simulator leverages a CPU-GPU hybrid environment to expedite the simulation by scaling out to multi-nodes equipped with multi-GPUs. We are able to conduct core simulations for a system-scale SNN chip of 20,000 neuromorphic cores on up to 512 A100 GPUs in a few minutes.
Huaipeng Zhang, Nhut-Minh Ho, Dogukan Yigit Polat, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen, Jintao Meng 0001, Rick Siow Mong Goh, Satoshi Matsuoka, Tao Luo 0014, Weng-Fai Wong
IEEE Trans. Parallel Distributed Syst.8
2022 CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
abstract
Optical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft.
Xiuchao Sui, Shaohua Li 0003, Xue Geng, Yan Wu 0002, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh, Hongyuan Zhu 0002
CVPR7
2022 A Resource-efficient Spiking Neural Network Accelerator Supporting Emerging Neural Encoding
abstract
Spiking neural networks (SNNs) recently gained momentum due to their low-power multiplication-free computing and the closer resemblance of biological processes in the nervous system of humans. However, SNNs require very long spike trains (up to 1000) to reach an accuracy similar to their artificial neural network (ANN) counterparts for large models, which offsets efficiency and inhibits its application to low-power systems for real-world use cases. To alleviate this problem, emerging neural encoding schemes are proposed to shorten the spike train while maintaining the high accuracy. However, current accelerators for SNN cannot well support the emerging encoding schemes. In this work, we present a novel hardware architecture that can efficiently support SNN with emerging neural encoding. Our implementation features energy and area efficient processing units with increased parallelism and reduced memory accesses. We verified the accelerator on FPGA and achieve 25% and 90% improvement over previous work in power consumption and latency, respectively. At the same time, high area efficiency allows us to scale for large neural network models. To the best of our knowledge, this is the first work to deploy the large neural network model VGG on physical FPGA-based neuromorphic hardware.
Daniel Gerlinghoff, Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Tao Luo 0014
DATE4
2022 Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
Jiawei Du 0002, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, Vincent Y. F. Tan
ICLR6
2022 Adversarial Semantic Hallucination for Domain Generalized Semantic Segmentation
abstract
Convolutional neural networks typically perform poorly when the test (target domain) and training (source domain) data have significantly different distributions. While this problem can be mitigated by using the target domain data to align the source and target domain feature representations, the target domain data may be unavailable due to privacy concerns. Consequently, there is a need for methods that generalize well despite restricted access to target domain data during training. In this work, we propose an adversarial semantic hallucination approach (ASH), which combines a class-conditioned hallucination module and a semantic segmentation module. Since the segmentation performance varies across different classes, we design a semantic-conditioned style hallucination module to generate affine transformation parameters from semantic information in the segmentation probability maps of the source domain image. Unlike previous adaptation approaches, which treat all classes equally, ASH considers the class-wise differences. The segmentation module and the hallucination module compete adversarially, with the hallucination module generating increasingly "difficult" stylized images to challenge the segmentation module. In response, the segmentation module improves as it is trained with generated samples at an appropriate class-wise difficulty level. Our results on the Cityscapes and Mapillary benchmark datasets show that our method is competitive with state of the art work. Code is made available at https://github.com/gabriel-tjio/ASH.
Gabriel Tjio, Ping Liu 0004, Joey Tianyi Zhou, Rick Siow Mong Goh
WACV4
2022 Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong
Neurocomputing11
2022 Corrigendum to "Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency" [Neurocomputing (2022) 128-140]
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong
Neurocomputing11
2022 Adversarial multimodal fusion with attention mechanism for skin lesion classification using clinical and dermoscopic images
Yan Wang 0015, Yangqin Feng, Lei Zhang 0005, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Liangli Zhen
Medical Image Anal.6
2022 Natural Language Video Localization: A Revisit in Span-Based Question Answering Framework
abstract
Natural Language Video Localization (NLVL) aims to locate a target moment from an untrimmed video that semantically corresponds to a text query. Existing approaches mainly solve the NLVL problem from the perspective of computer vision by formulating it as ranking, anchor, or regression tasks. These methods suffer from large performance degradation when localizing on long videos. In this work, we address the NLVL from a new perspective, i.e., span-based question answering (QA), by treating the input video as a text passage. We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework (named VSLBase), to address NLVL. VSLNet tackles the differences between NLVL and span-based QA through a simple yet effective query-guided highlighting (QGH) strategy. QGH guides VSLNet to search for the matching video span within a highlighted region. To address the performance degradation on long videos, we further extend VSLNet to VSLNet-L by applying a multi-scale split-and-concatenation strategy. VSLNet-L first splits the untrimmed video into short clip segments; then, it predicts which clip segment contains the target moment and suppresses the importance of other segments. Finally, the clip segments are concatenated, with different confidences, to locate the target moment accurately. Extensive experiments on three benchmark datasets show that the proposed VSLNet and VSLNet-L outperform the state-of-the-art methods; VSLNet-L addresses the issue of performance degradation on long videos. Our study suggests that the span-based QA framework is an effective strategy to solve the NLVL problem.
Hao Zhang 0048, Aixin Sun, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 NC-Net: Efficient Neuromorphic Computing Using Aggregated Subnets on a Crossbar-Based Architecture With Nonvolatile Memory
abstract
Neuromorphic computing chips consisting of crossbar arrays of emergent nonvolatile memory (NVM) have the potential of achieving both high energy efficiency and throughput as the low-power implementation of convolutional neural network (CNN) inference engines. However, such hardware has design constraints, such as its limited fan-in/fan-out and resource-inefficient mapping, that make the design and deployment of CNN on them challenging. As a result, the user has to design the CNN model with intricate knowledge of the hardware architecture and even cannot fit the models in the hardware for CNN with high resolution image input. In this article, we propose the use ofaggregated subnets, NC-net, which is a constrained form of the traditional layer structure, to solve these issues. With our method, we put forward an energy-efficient buffer- and analogue-to-digital converter and digital-to-analogue converter (ADC/DAC)-free architecture and a scalable end-to-end solution that automatically satisfies the hardware constraints of crossbar architectures, while optimizing the resource usage. In our solution, the exploration and deployment of a CNN for a neuromorphic crossbar hardware start with a design front end based onTensorFlow. Our automated design flow maps the NC-net network fromTensorFlowto the crossbar architecture. We tested our designs on both a simulator and a field-programmable gate array (FPGA) emulator with various benchmarks. In addition to general benchmarks, including MNIST, SVHN, CIFAR-10, and CIFAR-100, we tested our system on a real-world application, human detection with high resolution (224$\times $224) images as the input. Our system achieves the state-of-the-art accuracy for these benchmarks on the crossbar-based neuromorphic hardware, with an accuracy of more than 90% for the latter. It also yielded up to$4.25\times $improvement in the efficiency of spiking core usage compared to TrueNorth.
Tao Luo 0014, Huaipeng Zhang, Chuping Qu, Yingnan Cui, Weng-Fai Wong, Rick Siow Mong Goh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2022 Big Data Driven Vessel Trajectory and Navigating State Prediction With Adaptive Learning, Motion Modeling and Particle Filtering Techniques
abstract
The predictive vessel surveillance is one of the indispensable functional components in intelligent maritime traffic system. Vessel trajectory prediction serves as a prerequisite for collision detection and risk assessment. Perceiving the forthcoming traffic situation in advance helps decide the succeeding actions to mitigate the potential risk. The availability of maritime big data brings great potential to extract vessel movement patterns to support trajectory forecasting. In this paper, a novel vessel trajectory and navigating state prediction methodology is proposed based on AIS data, which synergizes properly designed learning, motion modelling and knowledge base assisted particle filtering processes. The primary contributions of this work also comprise several critical research findings to handle the key challenges in vessel trajectory and navigating state prediction problem, such as the adaptive training window determination for the learning process and effective knowledge storage and searching algorithm intended to reduce the query time of waterway pattern retrieval. The studies for these challenges are still missing in the reported literatures but they are essentially important for improving the prediction accuracy, efficiency and practicality. With the maritime traffic data collected for Singapore water, a thorough evaluation of the prediction performance has been conducted for different navigating scenarios. It is also observed that better prediction outperforms on account of allowing earlier alert in risk detection.
Xiuju Fu, Wanbing Zhang, Ryan Wen Liu, Rick Siow Mong Goh
IEEE Trans. Intell. Transp. Syst.7
2022 Deep Multimodal Transfer Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval (CMR) enables flexible retrieval experience across different modalities (e.g., texts versus images), which maximally benefits us from the abundance of multimedia data. Existing deep CMR approaches commonly require a large amount of labeled data for training to achieve high performance. However, it is time-consuming and expensive to annotate the multimedia data manually. Thus, how to transfer valuable knowledge from existing annotated data to new data, especially from the known categories to new categories, becomes attractive for real-world applications. To achieve this end, we propose a deep multimodal transfer learning (DMTL) approach to transfer the knowledge from the previously labeled categories (source domain) to improve the retrieval performance on the unlabeled new categories (target domain). Specifically, we employ a joint learning paradigm to transfer knowledge by assigning a pseudolabel to each target sample. During training, the pseudolabel is iteratively updated and passed through our model in a self-supervised manner. At the same time, to reduce the domain discrepancy of different modalities, we construct multiple modality-specific neural networks to learn a shared semantic space for different modalities by enforcing the compactness of homoinstance samples and the scatters of heteroinstance samples. Our method is remarkably different from most of the existing transfer learning approaches. To be specific, previous works usually assume that the source domain and the target domain have the same label set. In contrast, our method considers a more challenging multimodal learning situation where the label sets of the two domains are different or even disjoint. Experimental studies on four widely used benchmarks validate the effectiveness of the proposed method in multimodal transfer learning and demonstrate its superior performance in CMR compared with 11 state-of-the-art methods.
Liangli Zhen, Peng Hu 0002, Xi Peng 0001, Rick Siow Mong Goh, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.4
2022 E3NE: An End-to-End Framework for Accelerating Spiking Neural Networks With Emerging Neural Encoding on FPGAs
abstract
Compiler frameworks are crucial for the widespread use of FPGA-based deep learning accelerators. They allow researchers and developers, who are not familiar with hardware engineering, to harness the performance attained by domain-specific logic. There exists a variety of frameworks for conventional artificial neural networks. However, not much research effort has been put into the creation of frameworks optimized for spiking neural networks (SNNs). This new generation of neural networks becomes increasingly interesting for the deployment of AI on edge devices, which have tight power and resource constraints. Our end-to-end framework E3NE automates the generation of efficient SNN inference logic for FPGAs. Based on a PyTorch model and user parameters, it applies various optimizations and assesses trade-offs inherent to spike-based accelerators. Multiple levels of parallelism and the use of an emerging neural encoding scheme result in an efficiency superior to previous SNN hardware implementations. For a similar model, E3NE uses less than 50% of hardware resources and 20% less power, while reducing the latency by an order of magnitude. Furthermore, scalability and generality allowed the deployment of the large-scale SNN models AlexNet and VGG.
Daniel Gerlinghoff, Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Tao Luo 0014
IEEE Trans. Parallel Distributed Syst.4
2021 DeepFire: Acceleration of Convolutional Spiking Neural Network on Modern Field Programmable Gate Arrays
abstract
Spiking neural networks (SNN) with their ‘integrate and fire’ (I&F) neurons replace the hardware-intensive multiply-accumulate (MAC) operations in convolutional neural networks (CNN) with accumulate operations — not only making it easy to implement on FPGAs but also opening up the opportunities for energy-efficient hardware acceleration. In this paper, we propose DeepFire — the high-performance RTL IP — for accelerating convolutional SNN inference. The IP exploits various resources available on modern FPGAs, and it outperforms existing SNN implementations by more than 10× in terms of both frame per second (FPS) and performance per watt (FPS/Watt). Our design achieves up to 40.1kFPS and 28.3kFPS on MNIST and CIFAR-10/SVHN datasets with 99.14% and 81.8%/93.1% accuracies respectively. IP was evaluated with 7-series and Ultrascale+ FPGAs from Xilinx achieving Fmax of 375MHz and 500MHz respectively.
Myat Thu Linn Aung, Chuping Qu, Tao Luo 0014, Rick Siow Mong Goh, Weng-Fai Wong
FPL5
2021 Medical Image Segmentation using Squeeze-and-Expansion Transformers
abstract
Medical image segmentation is important for computer-aided diagnosis. Good segmentation demands the model to see the big picture and fine details simultaneously, i.e., to learn image features that incorporate large context while keep high spatial resolutions. To approach this goal, the most widely used methods -- U-Net and variants, extract and fuse multi-scale features. However, the fused features still have small "effective receptive fields" with a focus on local image cues, limiting their performance. In this work, we propose Segtran, an alternative segmentation framework based on transformers, which have unlimited "effective receptive fields" even at high feature resolutions. The core of Segtran is a novel Squeeze-and-Expansion transformer: a squeezed attention block regularizes the self attention of transformers, and an expansion block learns diversified representations. Additionally, we propose a new positional encoding scheme for transformers, imposing a continuity inductive bias for images. Experiments were performed on 2D and 3D medical image segmentation tasks: optic disc/cup segmentation in fundus images (REFUGE'20 challenge), polyp segmentation in colonoscopy images, and brain tumor segmentation in MRI scans (BraTS'19 challenge). Compared with representative existing methods, Segtran consistently achieved the highest segmentation accuracy, and exhibited good cross-domain generalization capabilities.
Shaohua Li 0003, Xiuchao Sui, Xiangde Luo, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh
IJCAI6
2021 Few-Shot Domain Adaptation with Polymorphic Transformers
Shaohua Li 0003, Xiuchao Sui, Huazhu Fu, Xiangde Luo, Yangqin Feng, Xinxing Xu, Yong Liu 0026, Daniel S. W. Ting, Rick Siow Mong Goh
MICCAI (2)10
2021 Partially-Supervised Learning for Vessel Segmentation in Ocular Images
Yanyu Xu 0001, Xinxing Xu, Shenghua Gao, Rick Siow Mong Goh, Daniel S. W. Ting, Yong Liu 0026
MICCAI (1)5
2021 Video Corpus Moment Retrieval with Contrastive Learning
abstract
Given a collection of untrimmed and unsegmented videos, video corpus moment retrieval (VCMR) is to retrieve a temporal moment (i.e., a fraction of a video) that semantically corresponds to a given text query. As video and text are from two distinct feature spaces, there are two general approaches to address VCMR: (i) to separately encode each modality representations, then align the two modality representations for query processing, and (ii) to adopt fine-grained cross-modal interaction to learn multi-modal representations for query processing. While the second approach often leads to better retrieval accuracy, the first approach is far more efficient. In this paper, we propose a Retrieval and Localization Network with Contrastive Learning (ReLoCLNet) for VCMR. We adopt the first approach and introduce two contrastive learning objectives to refine video encoder and text encoder to learn video and text representations separately but with better alignment for VCMR. The video contrastive learning (VideoCL) is to maximize mutual information between query and candidate video at video-level. The frame contrastive learning (FrameCL) aims to highlight the moment region corresponds to the query at frame-level, within a video. Experimental results show that, although ReLoCLNet encodes text and video separately for efficiency, its retrieval accuracy is comparable with baselines adopting cross-modal interaction learning.
Hao Zhang 0048, Aixin Sun, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh
SIGIR7
2021 Low Latency Big Data Processing without Prior Information
abstract
Job scheduling plays an important role in improving the overall system performance in big data processing frameworks. Simple job scheduling policies, such as Fair and FIFO scheduling, do not consider job sizes and may degrade the performance when jobs of varying sizes arrive. More elaborate job scheduling policies make the convenient assumption that jobs are recurring, and complete information about their sizes is available from their prior runs. In this paper, we design and implement an efficient and practical job scheduler for big data processing systems to achieve better performance even without prior information about job sizes. The superior performance of our job scheduler originates from the design of multiple level priority queues, where jobs are demoted to lower priority queues if the amount of service consumed so far reaches a certain threshold. In this case, jobs in need of a small amount of service can finish in the topmost several levels of queues, while jobs that need a large amount of service to complete are moved to lower priority queues to avoid head-of-line blocking. Our new job scheduler can effectively mimic the shortest job first scheduling policy without knowing the job sizes in advance. To demonstrate its performance, we have implemented our new job scheduler in YARN, a popular resource manager used by Hadoop/Spark, and validated its performance with experiments in both real testbeds including Amazon EC2 and large-scale trace-driven simulations. Our experimental and simulation results have strongly confirmed the effectiveness of our design: our new job scheduler can reduce the average job response time of the Fair scheduler by up to 45 percent and achieve better fairness at the same time.
Zhiming Hu 0001, Baochun Li, Zheng Qin 0004, Rick Siow Mong Goh
IEEE Trans. Cloud Comput.4
2021 Concurrent Processing Cluster Design to Empower Simultaneous Prediction for Hundreds of Vessels' Trajectories in Near Real-Time
abstract
The automatic identification system (AIS) plays a vital role in maritime traffic surveillance. AIS is designed for remotely tracking vessels but nowadays also becomes a useful data source to enable vessels' trajectory prediction so as to facilitate early alert of potential collision risks. Recent studies focus on improving prediction accuracy through the machine learning and knowledge-based technologies but the computational cost also greatly increases due to the model complexity of these methodologies. It becomes a practical challenge to realize near real-time (NRT) trajectory prediction of a large volume of ships. For risk alert application, forecasting timeliness is one of the key design considerations. In this paper, we propose a concurrent processing cluster solution to empower advanced trajectory forecasting for hundreds of vessels in NRT. The proposed solution relies on the properly determined frameworks by customizing the desired features for an integrated solution. Meanwhile, a novel task-based load balancing strategy with newly defined metrics are proposed, which aims to reduce the makespan of jobs and outperforms the existing load balancing algorithms. A practicable cluster system has been successfully implemented, serving as a step toward unlocking the power of advanced maritime traffic forecasting technologies and enabling the benefit from the latest progress on the methodological innovation.
Xiuju Fu, Wanbing Zhang, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Syst. Man Cybern. Syst.6
2020 Two-Phase Multi-Party Computation Enabled Privacy-Preserving Federated Learning
abstract
Countries across the globe have been pushing strict regulations on the protection of personal or private data collected. The traditional centralized machine learning method, where data is collected from end-users or IoT devices, so that it can discover insights behind real-world data, may not be feasible for many data-driven industry applications in light of such regulations. A new machine learning method, coined by Google as Federated Learning (FL) enables multiple participants to train a machine learning model collectively without directly exchanging data. However, recent studies have shown that there is still a possibility to exploit the shared models to extract personal or confidential data. In this paper, we propose to adopt Multi-Party Computation (MPC) to achieve privacy-preserving model aggregation for FL. The MPC-enabled model aggregation in a peer-to-peer manner incurs high communication overhead with low scalability. To address this problem, the authors proposed to develop a two-phase mechanism by 1) electing a small committee and 2) providing MPC-enabled model aggregation service to a larger number of participants through the committee. The MPC-enabled FL framework has been integrated in an IoT platform for smart manufacturing. It enables a set of companies to train high quality models collectively by leveraging their complementary data-sets on their own premises, without compromising privacy, model accuracy vis-a`-vis traditional machine learning methods and execution efficiency in terms of communication cost and execution time.
Renuga Kanagavelu, Zengxiang Li, Juniarto Samsudin, Yechao Yang, Feng Yang 0011, Rick Siow Mong Goh, Merivyn Cheah, Praewpiraya Wiwatphonthana, Khajonpong Akkarajitsakul, Shangguang Wang
CCGRID6
2020 An FPGA-Based Hardware Emulator for Neuromorphic Chip With RRAM
abstract
Neuromorphic chip with RRAM devices has been demonstrated as a promising computing platform for neural network-based applications. By directly mapping the weight matrices of neural networks onto RRAM-based crossbar arrays, high energy, and area efficiency can be achieved. However, the design of an RRAM-based neuromorphic chip faces many constraints due to the variability and limitations of RRAM. Simulation and emulation can help in the design of a neuromorphic chip prior to fabrication. However, software-based chip simulation on CPU is slow, especially for large-scale network-on-chip (NoC)-based chip design. In this paper, we present a hardware emulator on field-programmable gate array (FPGA) for an RRAM-based neuromorphic chip. Our emulator supports the emulation of static and dynamic variation of the RRAM-based crossbars used in the neural cores of a neuromorphic chip. Furthermore, an NoC is also implemented on FPGA to emulate the communication between the neural cores. Using the emulator, we show that effects, such as RRAM write and read noise and stuck-at faults affect the accuracy of an application on a neuromorphic chip. We also demonstrate the utility of the emulator in investigating NoC topologies, routing buffer depths, and neural core mappings.
Tao Luo 0014, Chuping Qu, Matthew Kay Fei Lee, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2020 Blockchain and IoT for Insurance: A Case Study and Cyberinfrastructure Solution on Fine-Grained Transportation Insurance
abstract
In this study, we initiate a cyberinfrastructure solution by synergizing both the blockchain and Internet of Things (IoT) technologies for transportation insurance. The insurance premium related services are encapsulated in “on-chain” chaincodes to perform over the facts on vehicle's trip and driver's behavior, which are deduced through “off-chain” analytic services using the sensing data collected from vehicles' on-board sensors. A hybrid scheme coordinating both the permissioned (Hyperledger) and public (Ethereum) blockchains is proposed to exploit their respective capabilities in terms of high transaction throughput and built-in cryptocurrency. A working prototype platform is implemented with a basic premium calculation model. The prototype system is deployed across Amazon Web Services (AWSs) cloud in a real-world Internet environment. A comprehensive performance study from the aspects of throughput, latency, and resource usage under different configurations is presented to show the solution's feasibility. The design practice and research findings are concluded in consort with the experience gained for further enhancing the proposed solution and extending the functional features such as a more realistic insurance policy to be applied in generic vehicle insurance applications.
Zengxiang Li, Yechao Yang, Piao Chen, Ryan Wen Liu, Yauheni Pyrloh, Ekanut Sotthiwat, Rick Siow Mong Goh
IEEE Trans. Comput. Soc. Syst.9
2020 Traffic Pattern Mining and Forecasting Technologies in Maritime Traffic Service Networks: A Comprehensive Survey
abstract
Maritime traffic service networks and information systems play a vital role in maritime traffic safety management. The data collected from the maritime traffic networks are essential for the perception of traffic dynamics and predictive traffic regulation. This paper is devoted to surveying the key processing components in maritime traffic networks. Specifically, the latest progress on maritime traffic data mining technologies for maritime traffic pattern extraction and the recent effort on vessels' motion forecasting for better situation awareness are reviewed. Through the review, we highlight that the traffic pattern knowledge presents valued insights for wide-spectrum domain application purposes, and serves as a prerequisite for the knowledge based forecasting techniques that are growing in popularity. The development of maritime traffic research in pattern mining and traffic forecasting reviewed in this paper affirms the importance of advanced maritime traffic studies and the great potential in maritime traffic safety and intelligence enhancement to accommodate the implementation of the Internet of Things, artificial intelligence technologies, and knowledge engineering and big data computing solution.
Xiuju Fu, Rick Siow Mong Goh
IEEE Trans. Intell. Transp. Syst.4
2020 Fast Recovery MapReduce (FAR-MR) to accelerate failure recovery in big data applications
Juniarto Samsudin, Renuga Kanagavelu, Long Wang 0005, Theint Theint Aye, Rick Siow Mong Goh
J. Supercomput.7
2019 Dual Adversarial Neural Transfer for Low-Resource Named Entity Recognition
abstract
We propose a new neural transfer method termed Dual Adversarial Transfer Network (DATNet) for addressing low-resource Named Entity Recognition (NER).Specifically, two variants of DATNet, i.e., DATNet-F and DATNet-P, are investigated to explore effective feature fusion between high and low resource.To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD).Additionally, adversarial training is adopted to boost model generalization.In experiments, we examine the effects of different components in DATNet across domains and languages, and show that significant improvement can be obtained especially for lowresource data, without augmenting any additional hand-crafted features and pre-trained language model.
Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Hongyuan Zhu 0002, Rick Siow Mong Goh, Kenneth Kwok
ACL (1)6
2019 Efficient Multi-party Computation Algorithm Design for Real-World Applications
abstract
Secure Multi-Party Computation (MPC) is a promising privacy-preserving technology to enable multiple trustless parties to compute a function jointly without revealing private inputs to each other. With the fast development of MPC protocols, software implementation, and underlying computation infrastructure, MPC has developed from purely theoretical interest to tangible platform implementations for real-world applications. In this paper, we investigate multiple mechanisms to design efficient MPC algorithms by avoiding costly MPC operations and leveraging parallel operations. In order to speed up database table searching, a machine learning-based approach is proposed to completely avoid equality-check operations, playing the trade-off between efficiency and accuracy. According to our experimental results, a well-designed MPC algorithm could improve performance and scalability significantly, and thus make MPC technology practicable.
Zengxiang Li, Chutima Kitcharoenpaisan, Phond Phunchongharn, Yechao Yang, Rick Siow Mong Goh, Yusen Li
ICPADS5
2019 Multi-Instance Multi-Scale CNN for Medical Image Classification
Shaohua Li 0003, Yong Liu 0026, Xiuchao Sui, Cheng Chen 0008, Gabriel Tjio, Daniel S. W. Ting, Rick Siow Mong Goh
MICCAI (4)7
2019 A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips
abstract
Advances in non-volatile resistive switching random access memory (RRAM) have made it a promising memory technology with potential applications in low-power and embedded in-memory computing devices owing to a number of advantages such as low-energy consumption, low area cost and good scaling. There have been proposals to employ RRAM in architecting chips for neuromorphic computing and artificial neural networks where matrix-vector multiplication can be computed in the analog domain in a single timestep. However, it is challenging to employ RRAM devices in neuromorphic chips owing to the non-ideal behavior of RRAM. In this article, we propose a cycle-accurate and scalable system-level simulator that can be used to study the effects of using RRAM devices in neuromorphic computing chips. The simulator models a spatial neuromorphic chip architecture containing many neural cores with RRAM crossbars connected via a Network-on-Chip (NoC). We focus on system-level simulation and demonstrate the effectiveness of our simulator in understanding how non-linear RRAM effects such as stuck-at-faults (SAFs), write variability, and random telegraph noise (RTN) can impact an application’s behavior. By using our simulator, we show that RTN and write variability can have adverse effects on an application. Nevertheless, we show that these effects can be mitigated through proper design choices and the implementation of a write-verify scheme.
Matthew Kay Fei Lee, Yingnan Cui, Thannirmalai Somu, Tao Luo 0014, Jun Zhou 0014, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh
ACM Trans. Archit. Code Optim.8
2019 AnomalyNet: An Anomaly Detection Network for Video Surveillance
abstract
Sparse coding-based anomaly detection has shown promising performance, of which the keys are feature learning, sparse representation, and dictionary learning. In this paper, we propose a new neural network for anomaly detection (termed AnomalyNet) by deeply achieving feature learning, sparse representation, and dictionary learning in three joint neural processing blocks. Specifically, to learn better features, we design a motion fusion block accompanied by a feature transfer block to enjoy the advantages of eliminating noisy background, capturing motion, and alleviating data deficiency. Furthermore, to address some disadvantages (e.g., nonadaptive updating) of the existing sparse coding optimizers and embrace the merits of neural network (e.g., parallel computing), we design a novel recurrent neural network to learn sparse representation and dictionary by proposing an adaptive iterative hard-thresholding algorithm (adaptive ISTA) and reformulating the adaptive ISTA as a new long short-term memory (LSTM). To the best of our knowledge, this could be one of the first works to bridge the$\ell _{1}$ -solver and LSTM and may provide novel insight into understanding LSTM and model-based optimization (or named differentiable programming), as well as sparse coding-based anomaly detection. Extensive experiments show the state-of-the-art performance of our method in the abnormal events detection task.
Joey Tianyi Zhou, Jiawei Du 0002, Hongyuan Zhu 0002, Xi Peng 0001, Yong Liu 0026, Rick Siow Mong Goh
IEEE Trans. Inf. Forensics Secur.6
2019 Learning With Annotation of Various Degrees
abstract
In this paper, we study a new problem in the scenario of sequences labeling. To be exact, we consider that the training data are with annotation of various degrees, namely, fully labeled, unlabeled, and partially labeled sequences. The learning with fully un/labeled sequence refers to the standard setting in traditional un/supervised learning, and the proposed partially labeling specifies the subject that the element does not belong to. The partially labeled data are cheaper to obtain compared with the fully labeled data though it is less informative, especially when the tasks require a lot of domain knowledge. To solve such a practical challenge, we propose a novel deep conditional random field (CRF) model which utilizes an end-to-end learning manner to smoothly handle fully/un/partially labeled sequences within a unified framework. To the best of our knowledge, this could be one of the first works to utilize the partially labeled instance for sequence labeling, and the proposed algorithm unifies the deep learning and CRF in an end-to-end framework. Extensive experiments show that our method achieves state-of-the-art performance in two sequence labeling tasks on some popular data sets.
Joey Tianyi Zhou, Hao Zhang 0048, Chen Gong 0002, Xi Peng 0001, Zhiguo Cao 0001, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.7
2018 SC2Net: Sparse LSTMs for Sparse Coding
abstract
The iterative hard-thresholding algorithm (ISTA) is one of the most popular optimization solvers to achieve sparse codes. However, ISTA suffers from following problems: 1) ISTA employs non-adaptive updating strategy to learn the parameters on each dimension with a fixed learning rate. Such a strategy may lead to inferior performance due to the scarcity of diversity; 2) ISTA does not incorporate the historical information into the updating rules, and the historical information has been proven helpful to speed up the convergence. To address these challenging issues, we propose a novel formulation of ISTA (named as adaptive ISTA) by introducing a novel \textit{adaptive momentum vector}. To efficiently solve the proposed adaptive ISTA, we recast it as a recurrent neural network unit and show its connection with the well-known long short term memory (LSTM) model. With a new proposed unit, we present a neural network (termed SC2Net) to achieve sparse codes in an end-to-end manner. To the best of our knowledge, this is one of the first works to bridge the $\ell_1$-solver and LSTM, and may provide novel insights in understanding model-based optimization and LSTM. Extensive experiments show the effectiveness of our method on both unsupervised and supervised tasks.
Joey Tianyi Zhou, Kai Di, Jiawei Du 0002, Xi Peng 0001, Hao Yang 0033, Sinno Jialin Pan, Ivor W. Tsang, Yong Liu 0026, Zheng Qin 0004, Rick Siow Mong Goh
AAAI10
2018 Concurrent Hybrid Breadth-First-Search on Distributed PowerGraph for Skewed Graphs
abstract
Large-scale graph-structured computation is becoming increasingly important for various data analytics applications. However, most distributed graph processing frameworks do not directly support efficient implementation of sophisticated algorithms requiring massive graph traversals. In this paper, we propose concurrent hybrid breadth-first-search (BFS) algorithm on a popular distributed graph processing framework (Power-Graph and its optimized version PowerLyra). It leverages the small-world property of skewed graphs which apply to most realworld data sets. Hybrid BFS algorithm changes graph traversal direction to save computational workload and message transmissions. Concurrent BFS algorithm enables running multiple BFS simultaneously while sharing vertex explorations efficiently with bit operations. Extensive experiments are conducted to evaluate performance and scalability on a multi-core computer and a cluster of AWS instances. Concurrent hybrid BFS algorithm dramatically increases graph traversal efficiency with respect to the increasing number of concurrent BFS traversals. Compared with PowerGraph, PowerLyra uses fewer resources, while achieving significantly higher performance and better scalability.
Zengxiang Li, Shen Ren, Sifei Lu, Jiachun Guo, Wentong Cai 0001, Qin Zheng 0002, Rick Siow Mong Goh
ICPADS7
2018 Blockchain and IoT Data Analytics for Fine-Grained Transportation Insurance
abstract
Innovations such as the Cloud, Internet of Things (IoT) and data analytics have already dramatically altered the customer experience in many, if not all, industries. Blockchain, as another emerging technology, is expected to be the next generation infrastructure to established trusted multiparty collaborations. In this paper, we investigated the convergence of aforementioned technologies, by presenting a prototype of fine-grained transportation insurance. Insurance premium were assessed based on vehicles usage and driver's behavior, which were deduced from streaming IoT data collected from mobile sensors. This incentive mechanism promotes fairness among drivers and encourages safer driving style. The prototype takes advantage of both private blockchain (e.g., high transaction rate in Hyperledger) and public blockchain (e.g., inbuilt cryptocurrency). Besides system architecture and implementation details, preliminary performance evaluations are presented and discussed.
Zengxiang Li, Quanqing Xu, Ekanut Sotthiwat, Rick Siow Mong Goh, Xueping Liang
ICPADS5
2018 Optimal Fee Structure for Efficient Lightning Networks
abstract
Off-chain transaction handling like the Lightning Network (LN) is among the most promising solutions to solve the scaling challenges in blockchain technology like Bitcoin. At the same time, the LN faces its own challenges like transaction path lengths, centralization of channels (hubs), channel imbalances (depletion), etc. Here we study the effects of payment channel fees on these various factors. To get realistic insights, we based our study on empirical Bitcoin transaction patterns and existing LN structure, and apply a simple form of fee structure with only one tunable parameter, α, that is most influential on the transaction routing paths. We assume the transactions through the LN take the path of the lowest aggregate fees, and found as a consequence that one cannot have short average path lengths and low overall channel imbalances at the same time. A good compromise is to have fees proportional to the square root of the channel capacity, such that reasonably short path lengths and overall balanced channel capacities can be achieved that makes the operation of the LN more sustainable.
Alvin Heng Jun Ren, Siew Ann Cheong, Rick Siow Mong Goh
ICPADS4
2018 EMRShare: A Cross-Organizational Medical Data Sharing and Management Framework Using Permissioned Blockchain
abstract
With the development of information and storage technologies, electronic document recording has become an unalterable trend, which transforms the way that people store, access and operate the data generated in various applications. Healthcare is the leader application domain that pioneers in usage of electronic medical records (EMRs). Cross-organizational EMRs' sharing has many constructive effects in motivating the domain innovation, introducing better domain understanding and overall the domain intelligence. However, the privacy concern, trust issue as well as the sophisticated legal regulation of the sensitive EMRs' use leads to inefficiency in the data sharing process. In this paper, we propose a cross-organizational medical data sharing framework based on permissioned blockchain technology, named “EMRShare“, to resolve the trust concern existing in EMRs sharing practice among different participants like patients, clinicians and researchers, and other relevant parties such as the insurance agent and government, to make medical data sharing and access secure, efficient, transparent, immutable, traceable and auditable. A working prototype system is implemented to demonstrate the key features for the cross-organizational medical data sharing and access management. The objective of this work targets at explaining the essential design considerations along with the working principle and operation logics using the blockchain technology to facilitate the medical data sharing in a highly-cooperative healthcare ecosystem.
Zengxiang Li, Yong Liu 0026, Thanarit Lertwuthikarn, Rick Siow Mong Goh
ICPADS7
2018 Building an Ethereum and IPFS-Based Decentralized Social Network System
abstract
Evolvement of blockchain technology has greatly changed the network and it makes many applications to be distributed, decentralized without loss of security. Ethereum is an open-source blockchain platform that provides a runtime environment for running smart contracts, which is called Ethereum Virtual Machine (EVM). Ethereum-based applications are usually referred to as Decentralized Applications (DApps), since they are based on the decentralized EVM, and its smart contracts. Meanwhile, distributed data store also evolves fast with the blockchain technology. Distributed storage develops to reduce the cost of the server side hardware and increase data availability. InterPlanetary File System (IPFS) is a protocol for distributed storage. IPFS stores immutable data, remove duplication, and obtain address information for storage nodes to search for files in the network. Many DApps have been created with the use of these technologies and one example is to use this design for a decentralized Twitter-like system that is resistant to censorship and single point of failure. This paper involves researching the blockchain technology and implementing a decentralised social network application on the Ethereum private blockchain with the use of smart contract and IPFS. In addition, it examines how blockchain and distributed storage can enable more functionalities of traditional social network systems.
Quanqing Xu, Zhiwen Song, Rick Siow Mong Goh, Yongjun Li 0006
ICPADS3
2018 Cost-Efficient and Latency-Aware Workflow Scheduling Policy for Container-Based Systems
abstract
Container technology is being adopted to simplify workflow execution. In this paper, we investigate a workflow scheduling policy for container-based systems. A workflow, representing an application, consists of a set of tasks. Each task can be executed in a container within a virtual machine (VM), where the container packaging the function for the task should be loaded into the VM before task execution. To reduce the workflow execution time and the network bandwidth consumption, we propose a cost-efficient and latency-aware workflow scheduling algorithm that strategically loads the containers into VMs and executes the tasks on the VMs. The algorithm is based on “Stretch Out and Compact”, which can stretch out the tasks along the resources by critical path analysis and then find the inefficient slots within the computing resources and eventually compact the tasks into those slots. We introduce a concept of “virtual task” into the algorithm, where container loading is regarded as a virtual task that should be executed before the real task. The introduction of the virtual task can be more effective in finding the inefficient slots for the compaction, thus resulting in a more efficient workflow scheduling policy. Simulation results show that compared to the algorithms that fully or selectively load the dockers, the proposed algorithm can achieve less execution time while saving network bandwidth consumption for loading dockers.
Yong Liu 0026, Long Wang 0005, Zengxiang Li, Rick Siow Mong Goh
ICPADS5
2018 Exploiting Sparsity to Accelerate Fully Connected Layers of CNN-Based Applications on Mobile SoCs
abstract
Convolutional neural networks (CNNs) are widely employed in many image recognition applications. With the proliferation of embedded and mobile devices, such applications are becoming commonplace on mobile devices. Network pruning is a commonly used strategy to reduce the memory and storage footprints of CNNs on mobile devices. In this article, we propose customized versions of the sparse matrix multiplication algorithm to speed up inference on mobile devices and make it more energy efficient. Specifically, we propose a Block Compressed Sparse Column algorithm and a bit-representation-based algorithm (BitsGEMM) that exploit sparsity to accelerate the fully connected layers of a network on the NVIDIA Jetson TK1 platform. We evaluate the proposed algorithms using real-world object classification and object detection applications. Experiments show that performance speedups can be achieved over the original baseline implementation using cuBLAS. On object detection CNNs, an average speedup of 1.82× is obtained over baseline cuBLAS in the fully connected layer of the VGG model, whereas on classification CNNs, an average speedup of 1.51× is achieved for the fully connected layer of the pruned-VGG model. Energy consumption reduction of 43--46% is also observed due to decreased computational and memory bandwidth demands.
Xinfeng Xie, Dayou Du, Qian Li 0027, Yun Liang 0001, Wai Teng Tang, Zhongliang Ong, Mian Lu, Huynh Phung Huynh, Rick Siow Mong Goh
ACM Trans. Embed. Comput. Syst.9
2018 Data Privacy-Preserving Automation Architecture for Industrial Data Exchange in Smart Cities
abstract
Information exchange across different entities, aiming at bridging various information “islands” that enclose specific domain understandings, has become an important means to achieve advanced intelligence toward smart cities. However, the concern of data privacy hinders the progress to establish a highly cooperative information sharing ecosystem. Existing data sharing platforms work as separate systems without incorporating the data privacy processing feature. Data privacy processing often needs to be manually handled offline using detached toolkits or systems before publishing-lack of automation, which makes it difficult to meet the evergrowing data exchange demand in both volume and frequency. In this paper, driven by real-world needs, a novel backend computing architecture, named data privacy-preserving automation architecture (DPA), is proposed to facilitate online privacy-protection processing automation and secure data privacy, which is able to seamlessly integrate with companies' principal application system in an interruption-free manner, allowing for adaption to flexible models and quality of service (QoS) guarantee for cross-entity data exchange. A novel QoS management approach, based on actor mode concurrency, is proposed for privacy processing task prioritization in application layer. A prototype system has been implemented based on real-world mobility application to demonstrate the main features of the DPA architecture. The DPA architecture can be flexibly adapted for various domain applications of smart city development.
Xiuju Fu, Rick Siow Mong Goh
IEEE Trans. Ind. Informatics3
2018 QLDS: A Novel Design Scheme for Trajectory Privacy Protection with Utility Guarantee in Participatory Sensing
abstract
Participatory sensing, leveraging on the ubiquity of cheap sensors in mobile devices, enables various promising applications of great social benefit. However, its ubiquitous sampling and openness results in serious privacy concerns. People's activity trajectories may reveal their private information such as home and work place and thus requires proper protection. In this paper, we propose a novel design scheme, called query logics detached storage (QLDS), for trajectory privacy protection. The core idea of QLDS is to extract the query logics for personal trajectory retrieval and make actual trajectory tuples not clustered to any route-identity or user-identity at server end, which introduces fine-granularity anonymity. The QLDS design scheme stores the extracted query logics at users own client devices and the un-clustered location tuples at the backend server, guaranteeing trajectory reconstruction at client-end and privacy preservation at server-end. Besides, integration of QLDS with other privacy protection approaches can further enhance the protection strength and bring flexible configurations to meet individuals' variable privacy concerns. The theoretical analytics is provided for the privacy and utility evaluation. The data retrieval performance of QLDS is experimentally evaluated in real-world internet environment.
Ming Huang 0003, Loganathan Ponnambalam, Xiuju Fu, Rick Siow Mong Goh
IEEE Trans. Mob. Comput.6
2018 Transfer Hashing: From Shallow to Deep
abstract
One major assumption used in most existing hashing approaches is that the domain of interest (i.e., the target domain) could provide sufficient training data, either labeled or unlabeled. However, this assumption may be violated in practice. To address this so-called data sparsity issue in hashing, a new framework termed transfer hashing with privileged information (THPI) is proposed, which marriages hashing and transfer learning (TL). To show the efficacy of THPI, we propose three variants of the well-known iterative quantization (ITQ) as a showcase. The proposed methods, ITQ+, LapITQ+, and deep transfer hashing (DTH), solve the aforementioned data sparsity issue from different aspects. Specifically, ITQ+ is a shallow model, which makes ITQ achieve hashing in a TL manner. ITQ+ learns a new slack function from the source domain to approximate the quantization error on the target domain given by ITQ. To further improve the performance of ITQ+, LapITQ+ is proposed by embedding the geometric relationship of the source domain into the target domain. Moreover, DTH is proposed to show the generality of our framework by utilizing the powerful representative capacity of deep learning. To the best of our knowledge, this could be one of the first DTH works. Extensive experiments on several popular data sets demonstrate the effectiveness of our shallow and DTH approaches comparing with several state-of-the-art hashing approaches.
Joey Tianyi Zhou, Heng Zhao 0004, Xi Peng 0001, Zheng Qin 0004, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.6
2017 Performance Modelling and Cost Effective Execution for Distributed Graph Processing on Configurable VMs
abstract
Graph Processing has been widely used to capture complex data dependency and uncover relationship insights. Due to the ever-growing graph scale and algorithm complexity, distributed graph processing has become more and more popular. In this paper, we investigate how to balance performance and cost for large scale graph processing on configurable virtual machines (VMs). We analyze the system architecture and implementation details of a Pregel-like distributed graph processing framework and develop a system-aware model to predict the execution time. Consequently, cost effective execution scenarios are recommended by selecting a certain number of VMs with specified capability subject to the predefined resource price and user preference. Experiments using synthetic and real world graphs have verified that system-aware model can achieve much higher prediction accuracy than popular machine-learning models which treat graph processing framework as a black box. As a result, the recommended execution scenarios have comparable cost efficiency to the optimal scenarios.
Zengxiang Li, Shen Ren, Yong Liu 0026, Zheng Qin 0004, Rick Siow Mong Goh, Gurusamy Mohan
CCGrid6
2017 Job Scheduling without Prior Information in Big Data Processing Systems
abstract
Job scheduling plays an important role in improving the overall system performance in big data processing frameworks. Simple job scheduling policies, such as Fair and FIFO scheduling, do not consider job sizes and may degrade the performance when jobs of varying sizes arrive. More elaborate job scheduling policies make the convenient assumption that jobs are recurring, and complete information about their sizes is available from their prior runs. In this paper, we design and implement an efficient and practical job scheduler for big data processing systems to achieve better performance even without prior information about job sizes. The superior performance of our job scheduler originates from the design of multiple level priority queues, where jobs are demoted to lower priority queues if the amount of service consumed so far reaches a certain threshold. In this case, jobs in need of a small amount of service can finish in the topmost several levels of queues, while jobs that need a large amount of service to complete are moved to lower priority queues to avoid head-of-line blocking. Our new job scheduler can effectively mimic the shortest job first scheduling policy without knowing the job sizes in advance. To demonstrate its performance, we have implemented our new job scheduler in YARN, a popular resource manager used by Hadoop/Spark, and validated its performance with both experiments on real datasets and large-scale trace-driven simulations. Our experimental and simulation results have strongly confirmed the effectiveness of our design: our new job scheduler can reduce the average job response time of the Fair scheduler by up to 45%.
Zhiming Hu 0001, Baochun Li, Zheng Qin 0004, Rick Siow Mong Goh
ICDCS4
2017 Scale-Free Sparse Matrix-Vector Multiplication on Many-Core Architectures
abstract
Sparse matrix-vector multiplication (SpMV) is one of the most important kernels for many applications. In this paper, we study the implementation of SpMV for scale-free matrices on many-core architectures including graphic processing units and Xeon Phi coprocessors. We first propose a hardware oblivious implementation for heterogeneous many-core processors using OpenCL. Our OpenCL implementation uses a novel SpMV format called hybrid COO+CSR (HCC), which employs 2-D jagged partitioning to balance the workload among a large number of cores and improve the data locality. Moreover, the OpenCL implementation is designed to be parametric, which allows systematic performance tuning. We conduct experiments to evaluate the efficiency of our hardware oblivious implementation. Experiments show that it achieves comparable performance to the Intel MKL and state-of-the-art OpenCL-based ViennaCL library implementation. Although the OpenCL implementation provides functional portability for heterogeneous systems, it fails to take advantage of the low-level architectural features. To further improve the performance, we propose a hardware conscious implementation using the native parallel programming language. We use the Xeon Phi platform as a case study. In our hardware conscious implementation, we ensure that the HCC format efficiently utilizes the vector process units on Xeon Phi by employing low-level intrinsics, and improve the overall performance through locality-aware block mapping, and intrablock tiling. Experiments using a wide range of representative scale-free matrices demonstrate that compared with the OpenCL-based hardware oblivious implementation, the hardware conscious implementation achieves 2.2× speedup on average. Compared with MKL, the hardware conscious implementation achieves 3.1× speedup on Xeon Phi.
Yun Liang 0001, Wai Teng Tang, Ruizhe Zhao, Mian Lu, Huynh Phung Huynh, Rick Siow Mong Goh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2016 Market fluctuation risk evaluation model for planning in manufacturing
abstract
Many manufacturing companies are experiencing more and more wider variation in market growth and customer demand. Studies have been conducted over the years to improve the accuracy of the manufacturing forecast. However, the theoretical forecast models so far haven't achieved much success and the gaps between forecast and actual demand are still wide that a buffer has to be built up especially for semiconductor manufacturing company to cope with the market fluctuation. An effective approach to assess the vulnerability of the demand forecast to market fluctuation for manufacturing is critical. In this work, a cost-based model for forecast vulnerability evaluation through simulation and visualization of the impact of various demand fluctuation is proposed. The model is able to evaluate the robustness of the forecast of semiconductor manufacturers and enable companies to prepare and respond to unexpected market fluctuation risk. It can also assist manufacturing production and resource planning proactively with a robust forecast. The case study shows that the model is able to identify at-risk product groups, facilities and distribution centers. Company can then decide how its forecast shall be adjusted by simulation and visualization of the forecast vulnerability under various market trends.
Xiaofeng Yin, Xiuju Fu, Loganathan Ponnambalam, Haiyan Xu 0002, Rick Siow Mong Goh
ICARCV5
2016 Transfer Hashing with Privileged Information
Joey Tianyi Zhou, Xinxing Xu, Sinno Jialin Pan, Ivor W. Tsang, Zheng Qin 0004, Rick Siow Mong Goh
IJCAI6
2016 Efficient Query Processing on Many-core Architectures: A Case Study with Intel Xeon Phi Processor
abstract
Recently, Intel Xeon Phi is emerging as a many-core processor with up to 61 x86 cores. In this demonstration, we present PhiDB, an OLAP query processor with simultaneous multi-threading (SMT) capabilities on Xeon Phi as a case study for parallel database performance on future many-core processors. With the trend towards many-core architectures, query operator optimizations, and efficient query scheduling on such many-core architectures remain as challenging issues. This motivates us to redesign and evaluate query processors. In PhiDB, we apply Xeon Phi aware optimizations on query operators to exploit hardware features of Xeon Phi, and design a heuristic algorithm to schedule the concurrent execution of query operators for better performance, to demonstrate the performance impact of Xeon Phi aware optimizations. We have also developed a user interface for users to explore the underlying performance impacts of hardware-conscious optimizations and scheduling plans.
Xuntao Cheng, Bingsheng He, Mian Lu, Chiew Tong Lau, Huynh Phung Huynh, Rick Siow Mong Goh
SIGMOD Conference6
2015 Optimizing and auto-tuning scale-free sparse matrix-vector multiplication on Intel Xeon Phi
abstract
Recently, the Intel Xeon Phi coprocessor has received increasing attention in high performance computing due to its simple programming model and highly parallel architecture. In this paper, we implement sparse matrix vector multiplication (SpMV) for scale-free matrices on the Xeon Phi architecture and optimize its performance. Scale-free sparse matrices are widely used in various application domains, such as in the study of social networks, gene networks and web graphs. We propose a novel SpMV format called vectorized hybrid COO+CSR (VHCC). Our SpMV implementation employs 2D jagged partitioning, tiling and vectorized prefix sum computations to improve hardware resource utilization, and thus overall performance. As the achieved performance depends on the number of vertical panels, we also develop a performance tuning method to guide its selection. Experimental results demonstrate that our SpMV implementation achieves an average 3× speedup over Intel MKL for a wide range of scale-free matrices.
Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang 0001, Huynh Phung Huyng, Xibai Li, Rick Siow Mong Goh
CGO7
2015 Integrated QoS-aware Resource Provisioning for Parallel and Distributed Applications
abstract
With more parallel and distributed applications moving to Cloud and data centers, it is challenging to provide predictable and controllable resources to multiple tenants, and thus guarantee application performance. In this paper, we propose an integrated QoS-aware resource provisioning platform based on virtualization technology for computing, storage and network resources. Coarse-grained CPU mapping and fine-grained CPU scheduling mechanisms are proposed to enable adjustable computing power. A hierarchical distributed scheduling mechanism is implemented on a scalable storage system to guarantee I/O throughput for individual tenants and applications. A network manager has also been developed to guarantee the data transmission rate. Web-based interface enables users to monitor real time resource utilization and to adjust resource QoS levels on the fly. According to our experimental results, the resource cost can be saved up to 45% without degrading the performance of a distributed data processing benchmark, and the performance of a parallel agent-based simulation can be improved by 91% using the same amount of resources.
Zengxiang Li, Long Wang 0005, Yu Zhang 0028, Tram Truong Huu, En Sheng Lim, Purnima Murali Mohan, Shibin Cheng, Shu Qin Ren, Gurusamy Mohan, Zheng Qin 0004, Rick Siow Mong Goh
DS-RT11
2015 Efficient GPU Spatial-Temporal Multitasking
abstract
Heterogeneous computing nodes are now pervasive throughout computing, and GPUs have emerged as a leading computing device for application acceleration. GPUs have tremendous computing potential for data-parallel applications, and the emergence of GPUs has led to proliferation of GPU-accelerated applications. This proliferation has also led to systems in which many applications are competing for access to GPU resources, and efficient utilization of the GPU resources is critical to system performance. Prior techniques of temporal multitasking can be employed with GPU resources as well, but not all GPU kernels make full use of the GPU resources. There is, therefore, an unmet need for spatial multitasking in GPUs. Resources used inefficiently by one kernel can be instead assigned to another kernel that can more effectively use the resources. In this paper we propose a software-hardware solution for efficient spatial-temporal multitasking and a software based emulation framework for our system. We pair an efficient heuristic in software with hardware leaky-bucket based thread-block interleaving to implement spatial-temporal multitasking. We demonstrate our techniques on various GPU architecture using nine representative benchmarks from CUDA SDK. Our experiments on Fermi GTX480 demonstrate performance improvement by up to 46% (average 26%) over sequential GPU task execution and 37% (average 18%) over default concurrent multitasking. Compared with the state-of-the-art Kepler K20 using Hyper-Q technology, our technique achieves up to 40% (average 17%) performance improvement over default concurrent multitasking.
Yun Liang 0001, Huynh Phung Huynh, Kyle Rupnow, Rick Siow Mong Goh, Deming Chen
IEEE Trans. Parallel Distributed Syst.4
2015 MrPhi: An Optimized MapReduce Framework on Intel Xeon Phi Coprocessors
abstract
In this work, we develop MrPhi, an optimized MapReduce framework on a heterogeneous computing platform, particularly equipped with multiple Intel Xeon Phi coprocessors. To the best of our knowledge, this is the first work to optimize the MapReduce framework on the Xeon Phi. We first focus on employing advanced features of the Xeon Phi to achieve high performance on a single coprocessor. We propose a vectorization friendly technique and SIMD hash computation algorithms to utilize the SIMD vectors. Then we pipeline the map and reduce phases to improve the resource utilization. Furthermore, we eliminate multiple local arrays but use low cost atomic operations on the global array to improve the thread scalability. For a given application, our framework is able to automatically detect suitable techniques to apply. Moreover, we extend our framework to a heterogeneous platform to utilize all hardware resource effectively. We adopt non-blocking data transfer to hide the communication overhead. We also adopt aligned memory transfer in order to fully utilize the PCIe bandwidth between the host and coprocessor. We conduct comprehensive experiments to benchmark the Xeon Phi and compare our optimized MapReduce framework with a state-of-the-art multi-core based MapReduce framework (Phoenix++). By evaluating six real-world applications, the experimental results show that our optimized framework is 1.2 to 38× faster than Phoenix++ for various applications on a single Xeon Phi. Additionally, the performance of four applications is able to achieve linear scalability on a platform equipped with up to four Xeon Phi coprocessors.
Mian Lu, Yun Liang 0001, Huynh Phung Huynh, Zhongliang Ong, Bingsheng He, Rick Siow Mong Goh
IEEE Trans. Parallel Distributed Syst.6
2015 A Code Generation Framework for Targeting Optimized Library Calls for Multiple Platforms
abstract
Directive-based programming approaches such as OpenMP and OpenACC have gained popularity due to their ease of programming. These programming models typically involve adding compiler directives to code sections such as loops in order to parallelize them for execution on multicore CPUs or GPUs. However, one problem with this approach is that existing compilers generate code directly from the annotated sections and do not make use of hardware-specific architectural features. As a result, the generated code is unable to fully exploit the capabilities of the underlying hardware. Alternatively, we propose a code generation framework in which linear algebraic operations in the annotated codes are recognized, extracted and mapped to optimized vendor-provided platform-specific library calls. We demonstrate that such an approach can result in better performance in the generated code compared to those which are generated by existing compilers. This is substantiated by experimental results on multicore CPUs and GPUs.
Wen Jun Tan, Wai Teng Tang, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong
IEEE Trans. Parallel Distributed Syst.3
2015 A Family of Bit-Representation-Optimized Formats for Fast Sparse Matrix-Vector Multiplication on the GPU
abstract
Sparse matrix-vector multiplication (SpMV) is an important kernel that is used in many iterative algorithms for solving scientific and engineering problems. One of the main challenges of SpMV is its memory-boundedness due to the low arithmetic intensity of the kernel. Although compression has been proposed previously to improve SpMV performance on CPUs, its use has not been demonstrated on the GPU because of the serial nature of many compression and decompression schemes. In this paper, we introduce a family of bit-representation-optimized (BRO) compression formats for representing sparse matrices on GPUs. The proposed formats - BRO-CSR, BRO-ELL and BRO-HYB, perform compression on index data and help to speed up SpMV on GPUs through the reduction of memory traffic. We also propose two other hybrid BRO formats which can potentially perform better than both HYB and BRO-HYB formats. Experimental results demonstrate that compared to uncompressed CSR and ELLPACK formats, our proposed compressed BRO-CSR and BRO-ELL formats are able to achieve average speedups of 2× and 1.4× respectively. Furthermore, we demonstrate that by using BRO-ELL, the preconditioned conjugate gradient method is able to achieve an average speedup of 1.3× over ELLPACK.
Wai Teng Tang, Wen Jun Tan, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong
IEEE Trans. Parallel Distributed Syst.3
2014 Towards building and evaluating a personalized location-based recommender system
abstract
Personalized location-based service recommendation is an important trend in the development of online ecommerce applications. In this work, we integrate the application of location-based service with recommendation technologies to present a hybrid recommendation model and a prototype system (HiPerData) to evaluate and measure the validity based on the Yelp dataset. In order to solve the four recommendation problems, we improve a predictive feature-based regression model, and combine the results of a set of collaborative filtering algorithms, which includes: SVD (Singular value decomposition), SVR (Support vector regression), SGD (Stochastic gradient descent), etc. Unlike previous approaches, we apply multiple methods to pre- and post-process the dataset and predict ratings, for example, a weighted pairwise preference regression for cold start problems, etc. We enhance the neighborhood-based approach leading to a substantial improvement of prediction accuracy. Our method gave the best overall results with a root mean square error of 1.22724.
Rubing Duan, Rick Siow Mong Goh, Feng Yang 0011, Yong Kiam Tan, Jesus F. B. Valenzuela
IEEE BigData2
2014 An Efficient Co-processing Framework for Large-Scale Scientific Applications
abstract
As scientific applications like Computational Fluid Dynamics (CFD) simulations generate more and more data, co-processing becomes the most cost effective way to process the vast amount of data generated by these simulation. In a co-processing environment, analysis and/or visualization of intermediate results occur concurrently to the simulation itself. Improved efficiency and early insight into the simulation process and results are potential advantages in comparison to postprocessing, where analysis and/or visualization are performed after the completion of the simulation. To enable co-processing, however, intermediate data needs to be shared between simulation and data analysis, and some degree of coordination may be required to maintain the correctness of both simulation and data analysis. The overhead incurred to facilitate data sharing and coordination may well offset benefits gained, particularly where distributed, large-scale systems are involved as workload sharing, processor affinity and data locality introduce significant effects to the overall performance. In this paper, we propose a co-processing framework to address these issues. The empirical benchmarking results suggest that co-processing overhead tasks scale well with the system size, the overall gain of about 20% in turnaround time compared to post-processing and that the coprocessing framework allows simulation and data analysis task to scale up to their individual limits.
Rubing Duan, Rick Siow Mong Goh, Lily Rachmawati, Long Wang 0005, Henry Novianus Palit, Xiaorong Li, Chi Keong Goh, Partha Sarathi Dutta, Leigh Lapworth, David Knott
CloudCom2
2014 Hierarchical Parallelization and Runtime Scheduling for Pregel-Like Graph Processing Systems
abstract
Graph processing has become popular for various big data analytic applications. Google's Pregel framework enables vertex-centric graph processing in distributed environment based on Bulk Synchronous Parallel (BSP) model. However, the BSP model is inefficient for many complex graph algorithms requiring graph traversals, as only a small number of vertices really update states in each super step. In this paper, we propose an hierarchical parallelization mechanism, taking the advantages of both synchronous (warp-level) and asynchronous (task-level) parallelization approaches. In addition, a runtime task scheduling mechanism is proposed, relying on real-time monitoring or prediction of resource utilization. Experiments have verified that the hierarchical parallelization mechanism can expose greater parallelism, and thus, increase resource utilization significantly. Moreover, the runtime scheduling mechanism can avoid aggressive resource competition, and thus, further enhance the performance of the parallelized graph processing.
Zengxiang Li, Rubing Duan, Long Wang 0005, Sifei Lu, Zheng Qin 0004, Rick Siow Mong Goh
CloudCom6
2014 An intelligent analysis and prediction model for on-demand cloud computing systems
abstract
In this paper, an intelligent model for analyzing and predicting cloud computing resource utilization is proposed to enhance on-demand services in cloud computing systems. The model is with the capability to discover active users and mine the system storage utilization patterns. This model is also with learning capabilities to adapt the dynamics in the cloud computing platform by capturing changing patterns of system storage utilization, and it employs data mining means for computing the practical model to be used for prediction and providing inputs for intelligent management in the on-demand cloud computing system. We have evaluated the proposed analysis and prediction model in a cloud computing platform. High prediction accuracies of 95% and 86% have been achieved in 1-day ahead and 7-day ahead system utilization prediction, respectively.
Xiuju Fu, Xiaorong Li, Lipo Wang 0001, Rick Siow Mong Goh
IJCNN5
2014 Special Issue: Recent Advances in Parallel and Distributed Systems, ICPADS 2012 Selected Papers
Xueyan Tang, Wentong Cai 0001, Rick Siow Mong Goh
Future Gener. Comput. Syst.3
2014 Mapping Streaming Applications onto GPU Systems
abstract
Graphics processing units leverage on a large array of parallel processing cores to boost the performance of the streaming computation patterns frequently found in graphics applications. Unfortunately, while many other general purpose applications also exhibit streaming behavior, they possess unfavorable data layout and poor computation-to-communication ratios that may penalize any straight-forward GPU implementation. In this paper we describe a performance-driven code generation framework that maps general purpose streaming applications onto GPU systems. This automated framework takes into account the idiosyncrasies of the GPU pipeline and the unique memory hierarchy. The framework has been implemented as a back-end to the StreamIt programming language compiler. Several key features in this framework ensure maximized performance and scalability. First, the generated code increases the effectiveness of the on-chip memory hierarchy by employing a heterogeneous mix of compute and memory access threads. Our scheme goes against the conventional wisdom of GPU programming which is to use a large number of homogeneous threads. Second, we utilise an efficient stream graph partitioning algorithm to handle larger applications and achieve the best performance under the given on-chip memory constraints. Lastly, the framework maps complex applications onto multiple GPUs using a highly effective parallel execution scheme. Our comprehensive experiments show its scalability and significant speedup compared to a previous state-of-the-art solution.
Huynh Phung Huynh, Andrei Hagiescu, Zhongliang Ong, Weng-Fai Wong, Rick Siow Mong Goh
IEEE Trans. Parallel Distributed Syst.5
2013 Optimizing the MapReduce framework on Intel Xeon Phi coprocessor
abstract
MapReduce has become one of the most popular framework for building big-data applications. It was originally designed for distributed-computing, and has been extended to various hardware architectures, e.g., multi-core CPUs, GPUs and FPGAs. In this work, we develop the first MapReduce framework on the recently released Intel Xeon Phi coprocessor. We utilize advanced features of the Xeon Phi to achieve high performance. In order to take advantage of the SIMD vector processing units, we propose a vectorization friendly technique to assist the auto-vectorization as well as develop SIMD hash computation algorithms. Furthermore, we utilize MIMD hyper-threading to pipeline the map and reduce phases to improve the resource utilization. We also eliminate multiple local arrays but use low cost atomic operations on the global array for some applications, which can improve the thread scalability and data locality. We conduct comprehensive experiments to compare our optimized MapReduce framework with a state-of-the-art multi-core based MapReduce framework (Phoenix++). By evaluating six real-world applications, the experimental results show that our optimized framework is 1.2X to 38X faster than Phoenix++ for various applications on the Xeon Phi.
Mian Lu, Lei Zhang 0005, Huynh Phung Huynh, Zhongliang Ong, Yun Liang 0001, Bingsheng He, Rick Siow Mong Goh, Richard Huynh
IEEE BigData7
2013 Hierarchical Parallel Algorithm for Modularity-Based Community Detection Using GPUs
Chun Yew Cheong, Huynh Phung Huynh, David Lo 0001, Rick Siow Mong Goh
Euro-Par4
2013 Optimizing and Auto-Tuning Iterative Stencil Loops for GPUs with the In-Plane Method
abstract
Stencils represent an important class of computations that are used in many scientific disciplines. Increasingly, many of the stencil computations in scientific applications are being offloaded to GPUs to improve running times. Since a large part of the simulation time is spent inside the stencil kernels, optimizing the kernel is therefore important in the context of achieving greater computation efficiencies and reducing simulation time. In this work, we proposed a novel in-plane method for stencil computations on GPUs and compared its performance with the conventional method implemented in the Nvidia SDK. We also implemented an auto-tuning framework for our method to select the optimal parameters for different GPU architectures. A performance model was developed for our proposed method, and is used to speed up the auto-tuning process. Our results show that a speedup of nearly 2× can be achieved compared to Nvidia's implementation.
Wai Teng Tang, Wen Jun Tan, Ratna Krishnamoorthy, Yi Wen Wong, Shyh-Hao Kuo, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong
IPDPS6
2013 Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemes
abstract
The sparse matrix-vector (SpMV) multiplication routine is an important building block used in many iterative algorithms for solving scientific and engineering problems. One of the main challenges of SpMV is its memory-boundedness. Although compression has been proposed previously to improve SpMV performance on CPUs, its use has not been demonstrated on the GPU because of the serial nature of many compression and decompression schemes. In this paper, we introduce a family of bit-representation-optimized (BRO) compression schemes for representing sparse matrices on GPUs. The proposed schemes, BRO-ELL, BRO-COO, and BRO-HYB, perform compression on index data and help to speed up SpMV on GPUs through reduction of memory traffic. Furthermore, we formulate a BRO-aware matrix reordering scheme as a data clustering problem and use it to increase compression ratios. With the proposed schemes, experiments show that average speedups of 1.5x compared to ELLPACK and HYB can be achieved for SpMV on GPUs.
Wai Teng Tang, Wen Jun Tan, Rajarshi Ray 0001, Yi Wen Wong, Weiguang Chen, Shyh-Hao Kuo, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong
SC7
2012 QoS-Aware Revenue-Cost Optimization for Latency-Sensitive Services in IaaS Clouds
abstract
Recently, application service providers have been employing Infrastructure-as-a-Service (IaaS) clouds such as Amazon EC2 to scale their computing resources on-demand to adapt to dynamic workloads. Existing research has been focusing more on cloud resource scaling in batch processing, non latency-sensitive applications. In this paper, we consider the problem of revenue-cost optimization in cloud-based application service providers with stringent QoS requirements, e.g., online gaming services. We propose an integrated approach which combines resource provisioning algorithms and request scheduling disciplines. The main goal is to maximize the service provider's revenue via satisfying pre-defined QoS requirements, and at the same time, to minimize cloud resource cost. We have implemented the proposed resource provisioning algorithms and scheduling disciplines into a cloud scaling framework developed in our previous work. Extensive experiments have been conducted with a fully functional implementation and realistic workloads modeled after real traces of popular online game servers. The results demonstrated the effectiveness of our proposed approach.
Ta Nguyen Binh Duong, Xiaorong Li, Rick Siow Mong Goh, Xueyan Tang, Wentong Cai 0001
DS-RT3
2012 Tulipse: A Visualization Framework for User-Guided Parallelization
Yi Wen Wong, Tomasz Dubrownik, Wai Teng Tang, Wen Jun Tan, Rubing Duan, Rick Siow Mong Goh, Shyh-Hao Kuo, Stephen John Turner, Weng-Fai Wong
Euro-Par6
2012 GPGPU for Real-Time Data Analytics
abstract
The demand for real-time data analytics (RTDA) has been on the rise in the past decades and is ever-growing with the proliferation of different data collection devices.GPGPU (General-Purpose computation on Graphics Processing Units) is an emerging research area in HPC (high performance computing). With the massive computation power and high memory bandwidth, GPUs have become a sharp weapon to address the performance requirement of RTDA. Designed as co-processors, GPUs pose a number of technical challenges for RTDA in terms of efficiency and programmability. On the one hand, while new generation GPUs can have over an order of magnitude higher memory bandwidth and higher computation power (in terms of GFLOPS) than CPUs, novel GPGPU algorithmic design and implementation are a must to unleash the hardware power. On the other hand, writing a correct and efficient GPU program is still challenging in general, and even more difficult for RTDA with streaming updates and real-time multi-tasking.
Bingsheng He, Huynh Phung Huynh, Rick Siow Mong Goh
ICPADS3
2012 Automatic Refactoring of Legacy Fortran Code to the Array Slicing Notation
abstract
There are many legacy Fortran programs still in use today, especially scientific codes which were written decades ago. Many of these codes use explicit DO-loops in programs that tend to clutter the code and make it harder to understand and maintain. Modern features of the Fortran language, such as the array slicing notation and introduction of commonly used intrinsic functions, go a long way in helping programmers write code that is easier to read and maintain. We introduce a refactoring tool that can help to transform code to make use of the array slicing notation and related intrinsic functions.
Chandrasehar Rajaseharan, Wen Jun Tan, Wai Teng Tang, Stephen John Turner, Shyh-Hao Kuo, Rick Siow Mong Goh, Weng-Fai Wong
ICPADS6
2012 Visualization for Anomaly Detection and Data Management by Leveraging Network, Sensor and GIS Techniques
abstract
This paper studies the importance of visualization for discerning and interpreting patterns of data and its application for solving real problems, such as anomaly detection and data management. There are various ways to realize visualization to cater to the needs of numerous real life applications. Depending on needs, a combination of some of these ways may be required for presenting an effective visualization. The authors present visualization schemes for anomaly detection/condition monitoring and data management by leveraging network techniques and combining them with modern techniques such as sensor, database, mobile communication, GPS and GIS techniques. Two case studies are presented and analyzed. By stepping through the design and implementation processes of these projects, this paper aims to serve as a guide for other designers or researchers to create visual analysis tools or implement projects requiring such visualization.
Zhaoxia Wang 0001, Chee Seng Chong, Rick Siow Mong Goh, Wanqing Zhou, Dan Peng, Hoong Chor Chin
ICPADS3
2012 Scalable framework for mapping streaming applications onto multi-GPU systems
abstract
Graphics processing units leverage on a large array of parallel processing cores to boost the performance of a specific streaming computation pattern frequently found in graphics applications. Unfortunately, while many other general purpose applications do exhibit the required streaming behavior, they also possess unfavorable data layout and poor computation-to-communication ratios that penalize any straight-forward execution on the GPU. In this paper we describe an efficient and scalable code generation framework that can map general purpose streaming applications onto a multi-GPU system. This framework spans the entire core and memory hierarchy exposed by the multi-GPU system. Several key features in our framework ensure the scalability required by complex streaming applications. First, we propose an efficient stream graph partitioning algorithm that partitions the complex application to achieve the best performance under a given shared memory constraint. Next, the resulting partitions are mapped to multiple GPUs using an efficient architecture-driven strategy. The mapping balances the workload while considering the communication overhead. Finally, a highly effective pipeline execution is employed for the execution of the partitions on the multi-GPU system. The framework has been implemented as a back-end of the StreamIt programming language compiler. Our comprehensive experiments show its scalability and significant performance speedup compared with a previous state-of-the-art solution.
Huynh Phung Huynh, Andrei Hagiescu, Weng-Fai Wong, Rick Siow Mong Goh
PPoPP4
2011 A Framework for Dynamic Resource Provisioning and Adaptation in IaaS Clouds
abstract
Infrastructure-as-a-Service (IaaS) cloud computing provides the ability to dynamically acquire extra or release existing computing resources on-demand to adapt to dynamic application workloads. In this paper, we propose an extensible framework for on-demand cloud resource provisioning and adaptation. The core of the framework is a set of resource adaptation algorithms that are capable of making informed provisioning decisions to adapt to workload fluctuations. The framework is designed to manage multiple sets of resources acquired from different cloud providers, and to interact with different local resource managers. We have developed a fully functional web-service based prototype of this framework, and used it for performance evaluation of various resource adaptation algorithms under different realistic settings, e.g. when input data such as jobs' wall times are inaccurate. Extensive experiments have been conducted with both synthetic and real workload traces obtained from the Grid Workload Archives, more specifically the traces from the Large Hadron Collider Computing Grid. The results demonstrate the effectiveness and robustness of our proposed algorithms.
Ta Nguyen Binh Duong, Xiaorong Li, Rick Siow Mong Goh
CloudCom3
2011 Automated Architecture-Aware Mapping of Streaming Applications Onto GPUs
abstract
Graphic Processing Units (GPUs) are made up of many streaming multiprocessors, each consisting of processing cores that interleave the execution of a large number of threads. Groups of threads - called warps and wave fronts, respectively, in nVidia and AMD literature - are selected by the hardware scheduler and executed in lockstep on the available cores. If threads in such a group access the slow off-chip global memory, the entire group has to be stalled, and another group is scheduled instead. The utilization of a given multiprocessor will remain high if there is a sufficient number of alternative thread groups to select from. Many parallel general purpose applications have been efficiently mapped to GPUs. Unfortunately, many stream processing applications exhibit unfavorable data movement patterns and low computation-to-communication ratio that may lead to poor performance. In this paper, we describe an automated compilation flow that maps most stream processing applications onto GPUs by taking into consideration two important architectural features of nVidia GPUs, namely interleaved execution as well as the small amount of shared memory available in each streaming multiprocessors. In particular, we show that using a small number of compute threads such that the memory footprint is reduced, we can achieve high utilization of the GPU cores. Our scheme goes against the conventional wisdom of GPU programming which is to use a large number of homogeneous threads. Instead, it uses a mix of compute and memory access threads, together with a carefully crafted schedule that exploits parallelism in the streaming application, while maximizing the effectiveness of the unique memory hierarchy. We have implemented our scheme in the compiler of the Stream It programming language, and our results show a significant speedup compared to the state-of-the-art solutions.
Andrei Hagiescu, Huynh Phung Huynh, Weng-Fai Wong, Rick Siow Mong Goh
IPDPS4
2010 Multi-objective optimization of large scale berth allocation and quay crane assignment problems
abstract
This paper describes the use of multi-objective optimization on a berth allocation and quay crane assignment problem (BACAP). The BACAP involves the simultaneous optimization of two highly-coupled container terminal operations, namely berth allocation and quay crane assignment, which have been traditionally solved as individual problems. The developed multi-objective evolutionary algorithm (MOEA) is validated on a large scale BACAP dataset, consisting of 23 berths and 87 quay cranes, generated based on the port conditions at the Pasir Panjang container terminal, which is the largest container terminal in Singapore. Optimization results show that the multi-objective optimization approach offers the port manager flexibility in selecting a desirable solution for implementation.
Chun Yew Cheong, Mohamed Salahuddin Habibullah, Rick Siow Mong Goh, Xiuju Fu
SMC3
2009 Data mining analysis to validate performance tuning practices for HPL
abstract
Applications performance is a criterion for system evaluation, and hence performance tuning for these applications is of great interest. One such benchmark application is High Performance Linpack (HPL). Although guidelines exist for HPL tuning, validating these guidelines on various systems is a challenging task as a large number of configurations need to be tested. In this work, we use data mining analysis to reduce the number of configurations to be tested in validating the HPL tuning guidelines on the Ranger System. We validate that NB, P and Q are the three most important parameters to tune HPL, and that PMAP does not have a significant impact on HPL performance. We also validate the practice of tuning HPL at small N using data mining analysis. We find that the value of N selected for tuning should not be significantly smaller than the largest N that can fit into the system memory. Our results indicate that data mining could be further applied to application performance tuning.
Tuan Zea Tan, Rick Siow Mong Goh, Verdi March, Simon See
CLUSTER2
2009 A Tabu Search for the Heterogeneous DAG Scheduling Problem
abstract
Scheduling parallel applications on heterogeneous processors/architectures with different computational speed is a difficult problem. Here, a tabu search metaheuristic is developed to improve the schedule generated by list scheduling. Three neighbourhoods variants are proposed and examined, including a novel neighbourhood that takes the shape of the task graph into account. The effectiveness is evaluated based on a set of modified random benchmark graphs, including task graphs of real-world applications. Factors affecting algorithm performance are also examined. We have found that the variants proposed were able to reduce the schedule length produced by HEFT up to an average of 30% and up to on average 20% for the standard graphs. The results also show that using information about the shape of the task graph is a viable strategy.
Yi Wen Wong, Rick Siow Mong Goh, Shyh-Hao Kuo, Malcolm Y. H. Low
ICPADS2
2009 A GA-SVM Feature Selection Model Based on High Performance Computing Techniques
abstract
Supervised learning is well-known and widely applied in many domains including bioinformatics, cheminformatics and financial forecasting. However, the interference from irrelevant features may lead to the poor accuracy of classifiers. As a popular feature selection model, GA-SVM is desirable in many of those cases to filter out irrelevant features and improve the learning performance subsequently. However, the high computational cost strongly discourages the application of GA-SVM in large-scale datasets. In this paper, an HPC-enabled GA-SVM (HGA-SVM) is proposed by integrating data parallelization, multithreading and heuristic techniques with the ultimate goal of robustness and low computational cost. Our proposed model is comprised of four improvement strategies: 1) GA parallelization, 2) SVM parallelization, 3) neighbor search and 4) evaluation caching. All the four strategies improve various aspects of the feature selection model and contribute collectively towards higher computational throughput.
Tianyou Zhang, Xiuju Fu, Rick Siow Mong Goh, Chee Keong Kwoh 0001, Gary Kee Khoon Lee
SMC3