Simon See

dblp:62/6547 · DBLP profile ↗
← Back
97ranked-venue papers
1as first author
65since 2021 · last 2026
0000-0002-4958-9237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 28 since 2021Systems, architecture and hardware · 19 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 XToM: Exploring the Multilingual Theory of Mind for Large Language Models
abstract
Theory of Mind (ToM), the ability to infer mental states in others, is pivotal for human social cognition. Existing evaluations of ToM in LLMs are largely limited to English, neglecting the linguistic diversity that shapes human cognition. This limitation raises a critical question: can LLMs exhibit Multilingual Theory of Mind, which is the capacity to reason about mental states across diverse linguistic contexts? To address this gap, we present XToM, a rigorously validated multilingual benchmark that evaluates ToM across five languages and incorporates diverse, contextually rich task scenarios. Using XToM, we systematically evaluate LLMs (e.g., DeepSeek R1), revealing a pronounced dissonance: while models excel in multilingual language understanding, their ToM performance varies across languages. Our findings expose limitations in LLMs' ability to replicate human-like mentalizing across linguistic contexts.
Chunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou, Xinyuan Cheng, Zhifan Sun, Zheye Deng, Kawai Chung, Yuzhuo Ao, Yixiang Fan, Cheng Jiayang, Ercong Nie, Ginny Y. Wong, Helmut Schmid, Hinrich Schütze, Simon See, Yangqiu Song
ACL (1)16
2026 IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning
abstract
Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani, Erik Cambria, Zhengchen Zhang, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Navya Gupta, Rishitej Reddy Vyalla, Avinash Anand, Chhavi Kirtani, Erik Cambria, Zhengchen Zhang, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah
ACL (1)10
2026 Cross-domain Few-shot Classification via Invariant-content Feature Reconstruction
abstract
Abstract In cross-domain few-shot classification (CFC), mainstream studies aim to train a simple module (e.g. a linear transformation head) to select or transform features (a.k.a., the high-level semantic features) for previously unseen domains with a few labeled training data available on top of a powerful pre-trained model. These studies usually assume that high-level semantic features are shared across these domains, and just simple feature selection or transformations are enough to adapt features to previously unseen domains. However, in this paper, we find that the simply transformed features are too general to fully cover the key content features regarding each class. Thus, we propose an effective method, invariant-content feature reconstruction (IFR), to train a simple module that simultaneously considers both high-level and fine-grained invariant-content features for the previously unseen domains. Specifically, the fine-grained invariant-content features are considered as a set of informative and discriminative features learned from a few labeled training data of tasks sampled from unseen domains and are extracted by retrieving features that are invariant to style modifications from a set of content-preserving augmented data in pixel level with an attention module. Extensive experiments on the Meta-Dataset benchmark show that IFR achieves good generalization performance on unseen domains, which demonstrates the effectiveness of the fusion of the high-level features and the fine-grained invariant-content features. Specifically, IFR improves the average accuracy on unseen domains by 1.6% and 6.5% respectively under two different cross-domain few-shot classification settings.
Hongduan Tian, Feng Liu 0003, Ka Chun Cheung, Zhen Fang 0001, Simon See, Tongliang Liu, Bo Han 0003
Int. J. Comput. Vis.5
2026 Multi-process thermodynamic graph learning for 2D fluid simulation
abstract
Solving partial differential equations (PDEs) for fluid simulation is computationally expensive, especially when dealing with complex geometries and high-resolution meshes. Recent advances in physics-informed graph neural networks (PIGNNs) have demonstrated potential in approximating such simulations more efficiently. Particularly, thermodynamic informed graph neural networks (TIGNNs) offer a promising data-driven alternative to traditional PDE solvers for fluid simulations. However, existing TIGNN implementations suffer from significant training inefficiencies, requiring prolonged runtimes and high memory consumption due to the need to maintain large parameter matrices in GPU memory. Inspired by multi-processor strategies in deformable solid simulations, we propose a novel multi-processor thermodynamic-informed graph neural network (MP-TIGNN) architecture to significantly accelerate training without compromising accuracy. Our approach enables faster convergence and reduces memory usage with mixed-precision training by leveraging fully sharded data parallelism (FSDP) across multiple GPUs. Experimental results show that our approach reduces training time by approximately 70% compared to the original setup while maintaining similar prediction accuracy.
Aik Beng Ng, Simon See, Zhengkui Wang, Frank Guan
Virtual Real. Intell. Hardw.3
2025 M3Net: Multimodal Multi-task Learning for 3D Detection, Segmentation, and Occupancy Prediction in Autonomous Driving
abstract
The perception system for autonomous driving generally requires to handle multiple diverse sub-tasks. However, current algorithms typically tackle individual sub-tasks separately, which leads to low efficiency when aiming at obtaining full-perception results. Some multi-task learning methods try to unify multiple tasks with one model, but do not solve the conflicts in multi-task learning. In this paper, we introduce M3Net, a novel multimodal and multi-task network that simultaneously tackles detection, segmentation, and 3D occupancy prediction for autonomous driving and achieves superior performance than single task model. M3Net takes multimodal data as input and multiple tasks via query-token interactions. To enhance the integration of multi-modal features for multi-task learning, we first propose the Modality-Adaptive Feature Integration (MAFI) module, which enables single-modality features to predict channel-wise attention weights for their high-performing tasks, respectively. Based on integrated features, we then develop task-specific query initialization strategies to accommodate the needs of detection/segmentation and 3D occupancy prediction. Leveraging the properly initialized queries, a shared decoder transforms queries and BEV features layer-wise, facilitating multi-task learning. Furthermore, we propose a Task-oriented Channel Scaling (TCS) module in the decoder to mitigate conflicts between optimizing for different tasks. Additionally, our proposed multi-task querying and TCS module support both Transformer-based decoder and Mamba-based decoder, demonstrating its flexibility to different architectures. M3Net achieves state-of-the-art multi-task learning performance on the nuScenes benchmarks.
Xuesong Chen 0001, Shaoshuai Shi, Tao Ma 0002, Jingqiu Zhou, Simon See, Ka Chun Cheung, Hongsheng Li 0001
AAAI5
2025 Test-Time Adaptation on Noisy Data via Model-Pruning-Based Filtering and Flatness-Aware Entropy Minimization
abstract
Test-time adaptation (TTA) deals with domain shifts during inference by training models based on only unlabeled test samples. Test samples may include noisy samples, which degrade domain adaptation. Existing methods rely on the model's output prediction to detect and filter noisy samples, and further search for flat regions during optimization, which makes the optimization more robust on noisy samples. However, there are two issues: (1) the output prediction tends to be inaccurate due to domain shifts, weakening noisy-sample detection; (2) current approaches for searching flat regions focus on optimization to enhance the worst case, which ignores achieving flatness by avoiding the quick changing of losses. To address these challenges, we propose a model pruning-based test-time adaptation model for noisy data streams, named MoTTA, which leverages a new proposed filtering, output difference under pruning (ODP)-based filtering, and a flatness-aware entropy minimization (FlatEM). Specifically, to reduce the impact of inaccurate output predictions, ODP-based filtering measures the output difference of a sample before and after model pruning, which works even under inaccurate output. To improve the search for flat loss surfaces, FlatEM integrates zeroth-order flatness and first-order flatness (minimize the maximal gradient normalization with a weight perturbation constrained in a small Euclidean ball) on entropy minimization. To solve these hard maximum problems, we leverage Taylor expansion to obtain approximated results for optimization. FlatEM also adopts a parameter regularization to mitigate incorrect updates from noisy samples. The experiments show our advantages in dealing with noisy data streams at TTA comparable to existing baselines.
Xingzhi Zhou 0002, Zhiliang Tian, Ka Chun Cheung, Simon See, Nevin Lianwen Zhang
AAAI6
2025 Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
abstract
Hateful meme detection aims to prevent the proliferation of hateful memes on various social media platforms. Considering its impact on social environments, this paper introduces a previously ignored but significant threat to hateful meme detection: backdoor attacks. By injecting specific triggers into meme samples, backdoor attackers can manipulate the detector to output their desired outcomes. To explore this, we propose the Meme Trojan framework to initiate backdoor attacks on hateful meme detection. Meme Trojan involves creating a novel Cross-Modal Trigger (CMT) and a learnable trigger augmentor to enhance the trigger pattern according to each input sample. Due to the cross-modal property, the proposed CMT can effectively initiate backdoor attacks on hateful meme detectors under an automatic application scenario. Additionally, the injection position and size of our triggers are adaptive to the texts contained in the meme, which ensures that the trigger is seamlessly integrated with the meme content. Our approach outperforms the state-of-the-art backdoor attack methods, showing significant improvements in effectiveness and stealthiness. We believe that this paper will draw more attention to the potential threat posed by backdoor attacks on hateful meme detection.
Ruofei Wang, Hongzhan Lin 0001, Ziyuan Luo, Ka Chun Cheung, Simon See, Jing Ma 0004, Renjie Wan
AAAI5
2025 A Lightweight Method for Generating Precise Number of 3D Object Instances from Text Prompts
abstract
Generating the precise number of 3D object instances from a user prompt is critical for building virtual worlds and augmenting 3D datasets. However, current generative models often struggle to honor explicit quantity instructions, even in straightforward prompts involving a single object type (e.g. “six chairs”), mainly because they process content generation holistically, prioritizing visual realism and global coherence, which often leads to the fuzziness in quantity fidelity of the results. We propose BatchGen, a new framework that enhances the precision of generating the exact number of individual 3D object instances from text, consistently delivering accurate results. There are two stages to our design: 1) A joint intent-slot model processes and interprets the user's utterance, and 2) the slot outputs serve as the text prompts to a tailored Text-to-3D model, which generates the correct number of specified 3D object instances. Experiments demonstrate that compared to existing methods, BatchGen excels at this task. Conceptually, our approach introduces a semantic enforcement mechanism towards ensuring that the generated output aligns exactly with the user's intent. This paves the way for exact, semantically accurate generation. In this paper, we focus specifically on ensuring the correct quantity of objects is generated.
Jin Qi Yeo, Aik Beng Ng, Simon See, Frank Guan
CW4
2025 LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
abstract
Tianshi Zheng, Cheng Jiayang, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tianshi Zheng, Cheng Jiayang, Zihao Wang 0001, Jiaxin Bai, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP9
2025 Resilient Test-Time Adaptation by Mitigating Batch-Normalization Overfitting
abstract
Test-time domain adaptation adjusts a source domain model to accommodate previously unseen domain shifts in a target domain during inference. In real-world scenarios, domain shifts continually evolve, and test data are often non-independent and identically distributed (non-i.i.d.). Existing methods update batch normalization (BN) statistics (mean and variance) based on test batch statistics to mitigate domain shifts and use a memory bank to provide approximate i.i.d. sampling by selectively storing samples. However, excessive updates to BN statistics lead to overfitting to specific domain shifts. To address this issue, we propose a resilient practical test-time adaptation method (ResiTTA), employing soft constraints on the BN statistics and a low-entropy sampling strategy, which reduces overfitting on domain shifts and enables rapid adaptation. Specifically, we develop a resilient batch normalization (BN) with estimated statistics and soft constraints between the source and the estimated statistics. The soft constraints regularize the estimated statistics to mitigate overfitting caused by the excessive updates. To avoid overfitting, we design a low-entropy memory bank that accounts for sample uncertainty and class balance. We adapt the source domain model via a teacher-student self-training adaptation on the samples from the memory, incorporating the soft constraints’ updates to BN. Our ResiTTA obtains state-of-the-art results on various benchmarks. We release our code1.
Xingzhi Zhou 0002, Zhiliang Tian, Xin Niu 0002, Ka Chun Cheung, Simon See, Nevin Lianwen Zhang
ICASSP7
2025 Masked Sensory-Temporal Attention for Sensor Generalization in Quadruped Locomotion
abstract
With the rising focus on quadrupeds, a generalized policy capable of handling different robot models and sensor inputs becomes highly beneficial. Although several methods have been proposed to address different morphologies, it remains a challenge for learning-based policies to manage various combinations of proprioceptive information. This paper presents Masked Sensory-Temporal Attention (MSTA), a novel transformer-based mechanism with masking for quadruped locomotion. It employs direct sensor-level attention to enhance the sensory-temporal understanding and handle different combinations of sensor data, serving as a foundation for incorporating unseen information. MSTA can effectively understand its states even with a large portion of missing information, and is flexible enough to be deployed on physical systems despite the long input sequence.
Dikai Liu, Tianwei Zhang 0004, Jianxiong Yin, Simon See
ICRA4
2025 Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection
abstract
Human-object interaction (HOI) detection is essential for accurately localizing and characterizing interactions between humans and objects, providing a comprehensive understanding of complex visual scenes across various domains. However, existing HOI detectors often struggle to deliver reliable predictions efficiently, relying on resource-intensive training methods and inefficient architectures. To address these challenges, we conceptualize a wavelet attention-like backbone and a novel ray-based encoder architecture tailored for HOI detection. Our wavelet backbone addresses the limitations of expressing middle-order interactions by aggregating discriminative features from the low- and high-order interactions extracted from diverse convolutional filters. Concurrently, the ray-based encoder facilitates multi-scale attention by optimizing the focus of the decoder on relevant regions of interest and mitigating computational overhead. As a result of harnessing the attenuated intensity of learnable ray origins, our decoder aligns query embeddings with emphasized regions of interest for accurate predictions. Experimental results on benchmark datasets, including ImageNet and HICO-DET, showcase the potential of our proposed architecture. The code is publicly available at [https://github.com/henrypay/RayEncoder].
Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See
IJCNN5
2025 SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
abstract
The resurgence of convolutional neural networks (CNNs) in visual recognition tasks, exemplified by ConvNeXt, has demonstrated their capability to rival transformer-based architectures through advanced training methodologies and ViTinspired design principles. However, both CNNs and transformers exhibit a simplicity bias, favoring straightforward features over complex structural representations. Furthermore, modern CNNs often integrate MLP-like blocks akin to those in transformers, but these blocks suffer from significant information redundancies, necessitating high expansion ratios to sustain competitive performance. To address these limitations, we propose SpaRTAN, a lightweight architectural design that enhances spatial and channel-wise information processing. SpaRTAN employs kernels with varying receptive fields, controlled by kernel size and dilation factor, to capture discriminative multi-order spatial features effectively. A wave-based channel aggregation module further modulates and reinforces pixel interactions, mitigating channel-wise redundancies. Combining the two modules, the proposed network can efficiently gather and dynamically contextualize discriminative features. Experimental results in ImageNet and COCO demonstrate that SpaRTAN achieves remarkable parameter efficiency while maintaining competitive performance. In particular, on the ImageNet-1k benchmark, SpaRTAN achieves 77. 7% accuracy with only 3.8M parameters and approximately 1.0 GFLOPs, demonstrating its ability to deliver strong performance through an efficient design. On the COCO benchmark, it achieves 50.0% AP, surpassing the previous benchmark by 1.2% with only 21.5M parameters. The code is publicly available at [https://github.com/henry-pay/SpaRTAN].
Quan Bi Pay, Vishnu Monn Baskaran, Junn Yong Loo, Koksheik Wong, Simon See
IJCNN5
2025 Unified Locomotion Transformer with Simultaneous Sim-to-Real Transfer for Quadrupeds
abstract
Quadrupeds have gained rapid advancement in their capability of traversing across complex terrains. The adoption of deep Reinforcement Learning (RL), transformers and various knowledge transfer techniques can greatly reduce the sim-to-real gap. However, the classical teacher-student framework commonly used in existing locomotion policies requires a pre-trained teacher and leverages the privilege information to guide the student policy. With the implementation of large-scale models in robotics controllers, especially transformers-based ones, this knowledge distillation technique starts to show its weakness in efficiency, due to the requirement of multiple supervised stages. In this paper, we propose Unified Locomotion Transformer (ULT), a new transformer-based framework to unify the processes of knowledge transfer and policy optimization in a single network while still taking advantage of privilege information. The policies are optimized with reinforcement learning, next state-action prediction, and action imitation, all in just one training stage, to achieve zero-shot deployment. Evaluation results demonstrate that with ULT, optimal teacher and student policies can be obtained at the same time, greatly easing the difficulty in knowledge transfer, even with complex transformer-based models.
Dikai Liu, Tianwei Zhang 0004, Jianxiong Yin, Simon See
IROS4
2025 Align 3D Representation and Text Embedding for 3D Content Personalization
abstract
Recent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D personalization approaches predominantly rely on knowledge distillation-based methods, which require computationally expensive retraining procedures. To address this challenge, we propose Invert3D, a novel framework for convenient 3D content personalization. Nowadays, vision-language models such as CLIP enable direct image personalization through aligned vision-text embedding spaces. However, the inherent structural differences between 3D content and 2D images preclude direct application of these techniques to 3D personalization. Our approach bridges this gap by establishing alignment between 3D representations and text embedding spaces. Specifically, we develop a camera-conditioned 3D-to-text inverse mechanism that projects 3D contents into a 3D embedding aligned with text embeddings. This alignment enables efficient manipulation and personalization of 3D content through natural language prompts, eliminating the need for computationally retraining procedures. Extensive experiments demonstrate that Invert3D achieves effective personalization of 3D content.
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
ACM Multimedia4
2025 Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction
abstract
Generalizable 3D Gaussian Splatting reconstruction showcases advanced Image-to-3D content creation but requires substantial computational resources and large datasets, posing challenges to training models from scratch. Current methods usually entangle the prediction of 3D Gaussian geometry and appearance, which rely heavily on data-driven priors and result in slow regression speeds. To address this, we propose Stereo-GS, a disentangled framework for efficient 3D Gaussian prediction. Our method extracts features from local image pairs using a stereo vision backbone and fuses them via global attention blocks. Dedicated point and Gaussian prediction heads generate multi-view point-maps for geometry and Gaussian features for appearance, combined as GS-maps to represent the 3DGS object. A refinement network enhances these GSmaps for high-quality reconstruction. Unlike existing methods that depend on camera parameters, our approach achieves pose-free 3D reconstruction, improving robustness and practicality. By reducing resource demands while maintaining high-quality outputs, Stereo- GS provides an efficient, scalable solution for real-world 3D content generation. Project page: https://kevinhuangxf.github.io/stereo-gs.
Xiufeng Huang, Ka Chun Cheung, Runmin Cong, Simon See, Renjie Wan
ACM Multimedia4
2025 ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
abstract
The widespread adoption of Retrieval-Augmented Image Generation (RAIG) has raised significant concerns about the unauthorized use of private image datasets. While these systems have shown remarkable capabilities in enhancing generation quality through reference images, protecting visual datasets from unauthorized use in such systems remains a challenging problem. Traditional digital watermarking approaches face limitations in RAIG systems, as the complex feature extraction and recombination processes fail to preserve watermark signals during generation. To address these challenges, we propose ImageSentinel, a novel framework for protecting visual datasets in RAIG. Our framework synthesizes sentinel images that maintain visual consistency with the original dataset. These sentinels enable protection verification through randomly generated character sequences that serve as retrieval keys. To ensure seamless integration, we leverage vision-language models to generate the sentinel images. Experimental results demonstrate that ImageSentinel effectively detects unauthorized dataset usage while preserving generation quality for authorized applications.
Ziyuan Luo, Yangyi Zhao, Ka Chun Cheung, Simon See, Renjie Wan
NeurIPS4
2025 MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
abstract
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Zhaowei Wang 0003, Wenhao Yu 0002, Xiyu Ren, Yu Zhao 0043, Rohit Saxena, Ginny Y. Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman
NeurIPS9
2025 Guest Editorial: Edge-Intelligence for Real-Time Computer Vision in 6G
Guodong Zhao 0001, Changyang She, Hao Su 0001, Dusit Niyato, Simon See, Dimitrios P. Pezaros
IEEE J. Sel. Areas Commun.5
2025 Replay Master: Automatic Sample Selection and Effective Memory Utilization for Continual Semantic Segmentation
abstract
Continual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in this task, replay methods can be adopted, constructing a memory buffer that stores a small number of samples from previous classes for future replay. However, existing replay approaches in CSS often lack a thorough exploration of two critical issues: how to find the most suitable memory samples and how to utilize them for replay more effectively. Common strategies either randomly select samples or rely on hand-crafted, single-factor-driven methods that are hard to be optimal, and often employ conventional training techniques for replay that do not account for class imbalance problem resulting from limited memory capacity. In this work, we tackle these challenges by introducing a novel memory sample selection method that leverages a reinforcement learning framework with innovative state representations and a dual-stage action scheme to automatically learn a selection policy. Additionally, we propose an expert mechanism and a dual-phase training method to address the class imbalance issue, thereby enhancing the effectiveness of replay training by making better use of memory samples. Incorporating the proposed automatic sample selection and effective memory utilization methods, we develop a novel and effective replay-based pipeline for CSS. Our extensive experiments on Pascal VOC 2012 and ADE20 K datasets demonstrate the effectiveness of our approach, which achieves state-of-the-art (SOTA) performance and outperforms previous advanced methods significantly.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, De Wen Soh, Jun Liu 0036
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Nutrition Estimation for Dietary Management: A Transformer Approach With Depth Sensing
abstract
Nutrition estimation is crucial for effective dietary management and overall health and well-being. Existing methods often struggle with sub-optimal accuracy and can be time-consuming. In this paper, we propose NuNet, a transformer-basednetwork designed fornutrition estimation that utilizes both RGB and depth information from food images. We have designed and implemented a multi-scale encoder and decoder, along with two types of feature fusion modules, specialized for estimating five nutritional factors. These modules effectively balance the efficiency and effectiveness of feature extraction with flexible usage of our customized attention mechanisms and fusion strategies. Our experimental study shows that NuNet significantly outperforms its variants and existing solutions for nutrition estimation. It achieves an error rate of 15.65%, the lowest known to us, largely due to our multi-scale architecture and fusion modules. This research holds practical values for dietary management with huge potential for transnational research and deployment and could inspire other applications involving multiple data types with varying degrees of importance.
Zhengyi Kwan, Wei Zhang 0082, Zhengkui Wang, Aik Beng Ng, Simon See
IEEE Trans. Multim.5
2024 AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
abstract
Zhaowei Wang, Wei Fan, Qing Zong, Hongming Zhang, Sehyun Choi, Tianqing Fang, Xin Liu, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaowei Wang 0003, Wei Fan 0001, Qing Zong, Hongming Zhang 0009, Sehyun Choi, Tianqing Fang, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See
ACL (1)10
2024 TCM-FTP: Fine-Tuning Large Language Models for Herbal Prescription Prediction
abstract
Traditional Chinese medicine (TCM) has relied on specific combinations of herbs in prescriptions to treat various symptoms and signs for thousands of years. Predicting TCM prescriptions poses a fascinating technical challenge with significant practical implications. However, this task faces limitations due to the scarcity of high-quality clinical datasets and the complex relationship between symptoms and herbs. To address these issues, we introduce DigestDS, a novel dataset comprising practical medical records from experienced experts in digestive system diseases. We also propose a method, TCM-FTP (TCM Fine-Tuning Pre-trained), to leverage pre-trained large language models (LLMs) via supervised fine-tuning on DigestDS. Additionally, we enhance computational efficiency using a low-rank adaptation technique. Moreover, TCM-FTP incorporates data augmentation by permuting herbs within prescriptions, exploiting their order-agnostic nature. Impressively, TCM-FTP achieves an F1-score of 0.8031, significantly outperforming previous methods. Furthermore, it demonstrates remarkable accuracy in dosage prediction, achieving a normalized mean square error of 0.0604. In contrast, LLMs without fine-tuning exhibit poor performance. Although LLMs have demonstrated wide-ranging capabilities, our work underscores the necessity of fine-tuning for TCM prescription prediction and presents an effective way to accomplish this.
Xingzhi Zhou 0002, Xin Dong 0017, Chunhao Li, Yuning Bai, Ka Chun Cheung, Simon See, Xinpeng Song, Runshun Zhang, Xuezhong Zhou, Nevin Lianwen Zhang
BIBM7
2024 AHIVE: Anatomy-Aware Hierarchical Vision Encoding for Interactive Radiology Report Retrieval
abstract
Automatic radiology report generation using deep learning models has been recently explored and found promising. Neural decoders are commonly used for the report generation, where irrelevant and unfaithful contents are unavoidable. The retrieval-based approach alleviates the limitation by identifying reports which are relevant to the input to assist the generation. To achieve clinically accurate report retrieval, we make reference to clinicians' diagnostic steps of examining a radiology image where anatomical and diagnostic details are typically focused, and propose a novel hierarchical visual concept representation called anatomy-aware hierarchical vision encoding (AHIVE). To learn AHIVE, we first derive a methodology to extract hierarchical diagnostic descriptions from radiology reports and develop a CLIP-based framework for the model training. Also, the hierarchical architecture of AHIVE is designed to support interactive report retrieval so that report revision made at one layer can be propagated to the subsequent ones to trigger other necessary revisions. We conduct extensive experiments and show that AHIVE can outperform the SOTA vision-language retrieval methods in terms of clinical accuracy by a large margin. We provide also a case study to illustrate how it enables interactive report retrieval.
Sixing Yan, William Kwok-Wai Cheung, Ivor W. Tsang, Wan Hang Keith Chiu, Terence M. Tong, Ka Chun Cheung, Simon See
CVPR7
2024 Addressing Background Context Bias in Few-Shot Segmentation Through Iterative Modulation
abstract
Existing few-shot segmentation methods usually extract foreground prototypes from support images to guide query image segmentation. However, different background contexts of support and query images can cause their foreground features to be misaligned. This phenomenon, known as background context bias, can hinder the effectiveness of support prototypes in guiding query image segmentation. In this work, we propose a novel framework with an it-erative structure to address this problem. In each iteration of the framework, we first generate a query prediction based on a support foreground feature. Next, we extract background context from the query image to modulate the support foreground feature, thus eliminating the foreground feature misalignment caused by the different backgrounds. After that, we design a confidence-biased attention to eliminate noise and cleanse information. By integrating these components through an iterative structure, we create a novel network that can leverage the synergies between different modules to improve their performance in a mutually reinforcing manner. Through these carefully designed components and structures, our network can effectively elimi-nate background context bias in few-shot segmentation, thus achieving outstanding performance. We conduct extensive experiments on the PASCAL-5iand COCO-20idatasets and achieve state-of-the-art (SOTA) results, which demonstrate the effectiveness of our approach.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
CVPR4
2024 Audience Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
abstract
Debate is the process of exchanging viewpoints or convincing others on a particular issue. Recent research has provided empirical evidence that the persuasiveness of an argument is determined not only by language usage but also by communicator characteristics. Researchers have paid much attention to aspects of languages, such as linguistic features and discourse structures, but combining argument persuasiveness and impact with the social personae of the audience has not been explored due to the difficulty and complexity. We have observed the impressive simulation and personification capability of ChatGPT, indicating a giant pre-trained language model may function as an individual to provide personae and exert unique influences based on diverse background knowledge. Therefore, we propose a persona knowledge-aligned framework for argument quality assessment tasks from the audience side. This is the first work that leverages the emergence of ChatGPT and injects such audience personae knowledge into smaller language models via prompt tuning. The performance of our pipeline demonstrates significant and consistent improvement compared to competitive architectures.
Chunkit Chan, Cheng Jiayang, Xin Liu 0039, Yauwai Yim, Zheye Deng, Haoran Li 0003, Yangqiu Song, Ginny Y. Wong, Simon See
ECAI10
2024 GeometrySticker: Enabling Ownership Claim of Recolorized Neural Radiance Fields
Xiufeng Huang, Ka Chun Cheung, Simon See, Renjie Wan
ECCV (9)3
2024 Protecting NeRFs' Copyright via Plug-And-Play Watermarking Base Model
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
ECCV (11)4
2024 Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
abstract
Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets.
Meng Shen 0002, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu 0001, Simon See
MMAsia6
2024 Geometry Cloak: Preventing TGS-based 3D Reconstruction from Copyrighted Images
abstract
Single-view 3D reconstruction methods like Triplane Gaussian Splatting (TGS) have enabled high-quality 3D model generation from just a single image input within seconds. However, this capability raises concerns about potential misuse, where malicious users could exploit TGS to create unauthorized 3D models from copyrighted images. To prevent such infringement, we propose a novel image protection approach that embeds invisible geometry perturbations, termed ``geometry cloaks'', into images before supplying them to TGS. These carefully crafted perturbations encode a customized message that is revealed when TGS attempts 3D reconstructions of the cloaked image. Unlike conventional adversarial attacks that simply degrade output quality, our method forces TGS to fail the 3D reconstruction in a specific way - by generating an identifiable customized pattern that acts as a watermark. This watermark allows copyright holders to assert ownership over any attempted 3D reconstructions made from their protected images. Extensive experiments have verified the effectiveness of our geometry cloak.
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan
NeurIPS4
2024 Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow
abstract
Existing Maximum-Entropy (MaxEnt) Reinforcement Learning (RL) methods for continuous action spaces are typically formulated based on actor-critic frameworks and optimized through alternating steps of policy evaluation and policy improvement. In the policy evaluation steps, the critic is updated to capture the soft Q-function. In the policy improvement steps, the actor is adjusted in accordance with the updated soft Q-function. In this paper, we introduce a new MaxEnt RL framework modeled using Energy-Based Normalizing Flows (EBFlow). This framework integrates the policy evaluation steps and the policy improvement steps, resulting in a single objective training process. Our method enables the calculation of the soft value function used in the policy evaluation target without Monte Carlo approximation. Moreover, this design supports the modeling of multi-modal action distributions while facilitating efficient action sampling. To evaluate the performance of our method, we conducted experiments on the MuJoCo benchmark suite and a number of high-dimensional robotic tasks simulated by Omniverse Isaac Gym. The evaluation results demonstrate that our method achieves superior performance compared to widely-adopted representative baselines.
Chen-Hao Chao, Chien Feng, Wei-Fang Sun, Cheng-Kuang Lee, Simon See, Chun-Yi Lee
NeurIPS5
2024 GaussianMarker: Uncertainty-Aware Copyright Protection of 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has become a crucial method for acquiring 3D assets. To protect the copyright of these assets, digital watermarking techniques can be applied to embed ownership information discreetly within 3DGS mod- els. However, existing watermarking methods for meshes, point clouds, and implicit radiance fields cannot be directly applied to 3DGS models, as 3DGS models use explicit 3D Gaussians with distinct structures and do not rely on neural networks. Naively embedding the watermark on a pre-trained 3DGS can cause obvious distortion in rendered images. In our work, we propose an uncertainty- based method that constrains the perturbation of model parameters to achieve invisible watermarking for 3DGS. At the message decoding stage, the copyright messages can be reliably extracted from both 3D Gaussians and 2D rendered im- ages even under various forms of 3D and 2D distortions. We conduct extensive experiments on the Blender, LLFF, and MipNeRF-360 datasets to validate the effectiveness of our proposed method, demonstrating state-of-the-art performance on both message decoding accuracy and view synthesis quality.
Xiufeng Huang, Yiu-Ming Cheung, Ka Chun Cheung, Simon See, Renjie Wan
NeurIPS5
2024 Natural generative noise diffusion model imputation
abstract
Imputation is a critical method for enhancing dataset quality, essential for ensuring accurate analysis and insights. This research proposes an advanced imputation algorithm utilizing a Diffusion Model enhanced with Perlin noise generation. We introduce Perlin noise at each step of the diffusion process and incorporate a cosine scheduler to optimize performance. Our approach demonstrates improvements in imputing non-normal data, validated through tests on ten real datasets. Due to the substantial slope of distribution properties produced by perlin noise, noisy data gradually contaminates nonnormally distributed data, making it more like to a Gaussian distribution. The Perlin noise distribution increases the normality of the noisy data provided noise when it enters the deep neural process of diffusion imputation. We assess our proposed approach by simulating the missing data rate using three scenarios: Missing Completely at Random (MCAR), Missing Not At Random (MNAR), and Missing at Random (MAR). Every case is handled similarly, with 20 % to 80 % missing data. Compared to other deep learning imputation methods, our proposed methods and improvements contribute to lowering the RMSE value up to 10 % on non-normal distributed data imputation.
Ari Wibisono, Denny, Petrus Mursanto, Simon See
Knowl. Based Syst.4
2024 Self-Supervised Video Representation Learning by Video Incoherence Detection
abstract
This article introduces a novel self-supervised method that leverages incoherence detection for video representation learning. It stems from the observation that the visual system of human beings can easily identify video incoherence based on their comprehensive understanding of videos. Specifically, we construct the incoherent clip by multiple subclips hierarchically sampled from the same raw video with various lengths of incoherence. The network is trained to learn the high-level representation by predicting the location and length of incoherence given the incoherent clip as input. Additionally, we introduce intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video. We evaluate our proposed method through extensive experiments on action recognition and video retrieval using various backbone networks. Experiments show that our proposed method achieves remarkable performance across different backbone networks and different datasets compared to previous coherence-based methods.
Haozhi Cao, Yuecong Xu, Kezhi Mao, Lihua Xie 0001, Jianxiong Yin, Simon See, Qianwen Xu 0001, Jianfei Yang 0001
IEEE Trans. Cybern.6
2024 CAMANet: Class Activation Map Guided Attention Network for Radiology Report Generation
abstract
Radiology report generation (RRG) has gained increasing research attention because of its huge potential to mitigate medical resource shortages and aid the process of disease decision making by radiologists. Recent advancements in Radiology Report Generation (RRG) are largely driven by improving a model's capabilities in encoding single-modal feature representations, while few studies explicitly explore the cross-modal alignment between image regions and words. Radiologists typically focus first on abnormal image regions before composing the corresponding text descriptions, thus cross-modal alignment is of great importance to learn a RRG model which is aware of abnormalities in the image. Motivated by this, we propose a Class Activation Map guided Attention Network (CAMANet) which explicitly promotes cross-modal alignment by employing aggregated class activation maps to supervise cross-modal attention learning, and simultaneously enrich the discriminative information. Experimental results demonstrate that CAMANet outperforms previous SOTA methods on two commonly used RRG benchmarks.
Jun Wang 0121, Abhir Bhalerao, Terry Yin, Simon See, Yulan He 0001
IEEE J. Biomed. Health Informatics4
2023 NAS-LID: Efficient Neural Architecture Search with Local Intrinsic Dimension
abstract
One-shot neural architecture search (NAS) substantially improves the search efficiency by training one supernet to estimate the performance of every possible child architecture (i.e., subnet). However, the inconsistency of characteristics among subnets incurs serious interference in the optimization, resulting in poor performance ranking correlation of subnets. Subsequent explorations decompose supernet weights via a particular criterion, e.g., gradient matching, to reduce the interference; yet they suffer from huge computational cost and low space separability. In this work, we propose a lightweight and effective local intrinsic dimension (LID)-based method NAS-LID. NAS-LID evaluates the geometrical properties of architectures by calculating the low-cost LID features layer-by-layer, and the similarity characterized by LID enjoys better separability compared with gradients, which thus effectively reduces the interference among subnets. Extensive experiments on NASBench-201 indicate that NAS-LID achieves superior performance with better efficiency. Specifically, compared to the gradient-driven method, NAS-LID can save up to 86% of GPU memory overhead when searching on NASBench-201. We also demonstrate the effectiveness of NAS-LID on ProxylessNAS and OFA spaces. Source code:https://github.com/marsggbo/NAS-LID.
Xin He 0019, Jiangchao Yao, Yuxin Wang 0003, Zhenheng Tang, Ka Chun Cheung, Simon See, Bo Han 0003, Xiaowen Chu 0001
AAAI6
2023 COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference Perspective
abstract
Zhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhaowei Wang 0003, Quyet V. Do, Hongming Zhang 0009, Jiayao Zhang 0001, Weiqi Wang 0001, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See
ACL (1)9
2023 A Simple Baseline for Video Restoration with Grouped Spatial-Temporal Shift
abstract
Video restoration, which aims to restore clear frames from degraded videos, has numerous important applications. The key to video restoration depends on utilizing inter-frame information. However, existing deep learning methods often rely on complicated network architectures, such as optical flow estimation, deformable convo-lution, and cross-frame self-attention layers, resulting in high computational costs. In this study, we propose a sim-ple yet effective framework for video restoration. Our approach is based on grouped spatial-temporal shift, which is a lightweight and straightforward technique that can implicitly capture inter-frame correspondences for multi-frame aggregation. By introducing grouped spatial shift, we attain expansive effective receptive fields. Combined with basic 2D convolution, this simple framework can effectively aggregate inter-frame information. Extensive experiments demonstrate that our framework outperforms the previous state-of-the-art method, while using less than a quarter of its computational cost, on both video deblurring and video denoising tasks. These results indicate the potential for our approach to significantly reduce computational overhead while maintaining high-quality results. Code is avaliable at https://github.com/dasonglil/Shift-Net.
Dasong Li, Xiaoyu Shi 0002, Yi Zhang 0108, Ka Chun Cheung, Simon See, Xiaogang Wang 0001, Hongwei Qin, Hongsheng Li 0001
CVPR5
2023 FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation
abstract
FlowFormer [24] introduces a transformer architecture into optical flow estimation and achieves state-of-the-art performance. The core component of FlowFormer is the transformer-based cost-volume encoder. Inspired by the recent success of masked autoencoding (MAE) pretraining in unleashing transformers' capacity of encoding visual representation, we propose Masked Cost Volume Autoencoding (MCVA) to enhance FlowFormer by pretraining the cost-volume encoder with a novel MAE scheme. Firstly, we introduce a block-sharing masking strategy to prevent masked information leakage, as the cost maps of neighboring source pixels are highly correlated. Secondly, we propose a novel pre-text reconstruction task, which encourages the cost-volume encoder to aggregate long-range information and ensures pretraining-finetuning consistency. We also show how to modify the FlowFormer architecture to accommodate masks during pretraining. Pretrained with MCVA, FlowFormer++ ranks 1st among published methods on both Sintel and KITTI-2015 benchmarks. Specifically, FlowFormer++ achieves 1.07 and 1.94 average end-point error (AEPE) on the clean and final pass of Sintel benchmark, leading to 7.76% and 7.18% error reductions from FlowFormer. FlowFormer++ obtains 4.52 F1-all on the KITTI-2015 test set, improving FlowFormer by 0.16.
Xiaoyu Shi 0002, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001
CVPR6
2023 Continual Semantic Segmentation with Automatic Memory Sample Selection
abstract
Continual Semantic Segmentation (CSS) extends static semantic segmentation by incrementally introducing new classes for training. To alleviate the catastrophic forgetting issue in CSS, a memory buffer that stores a small number of samples from the previous classes is constructed for replay. However, existing methods select the memory samples either randomly or based on a single-factor-driven handcrafted strategy, which has no guarantee to be optimal. In this work, we propose a novel memory sample selection mechanism that selects informative samples for effective replay in a fully automatic way by considering comprehensive factors including sample diversity and class performance. Our mechanism regards the selection operation as a decision-making process and learns an optimal selection policy that directly maximizes the validation performance on a reward set. To facilitate the selection decision, we design a novel state representation and a dual-stage action space. Our extensive experiments on Pascal-VOC 2012 and ADE 20K datasets demonstrate the effectiveness of our approach with state-of-the-art (SOTA) performance achieved, outperforming the second-place one by 12.54% for the 6-stage setting on Pascal-VOC 2012.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
CVPR4
2023 Deep Transfer Learning Application for Intelligent Marine Debris Detection
Kai Yuan Chia, Cheng Siong Chin, Simon See
EANN3
2023 TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses
abstract
3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MOT has made important progress in recent years. However, these methods only use the detection boxes of the current frame to obtain trajectory-box association results, which makes it impossible for the tracker to recover objects missed by the detector. In this paper, we present TrajectoryFormer, a novel point-cloud-based 3D MOT framework. To recover the missed object by detector, we generates multiple trajectory hypotheses with hybrid candidate boxes, including temporally predicted boxes and current-frame detection boxes, for trajectory-box association. The predicted boxes can propagate object’s history trajectory information to the current frame and thus the network can tolerate short-term miss detection of the tracked objects. We combine long-term object motion feature and short-term object appearance feature to create per-hypothesis feature embedding, which reduces the computational overhead for spatial-temporal encoding. Additionally, we introduce a Global-Local Interaction Module to conduct information interaction among all hypotheses and models their spatial relations, leading to accurate estimation of hypotheses. Our TrajectoryFormer achieves state-of-the-art performance on the Waymo 3D MOT benchmarks. Code is available at https://github.com/poodarchu/EFG.
Xuesong Chen 0001, Shaoshuai Shi, Benjin Zhu, Qiang Wang 0023, Ka Chun Cheung, Simon See, Hongsheng Li 0001
ICCV7
2023 CopyRNeRF: Protecting the CopyRight of Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have the potential to be a major representation of media. Since training a NeRF has never been an easy task, the protection of its model copyright should be a priority. In this paper, by analyzing the pros and cons of possible copyright protection solutions, we propose to protect the copyright of NeRF models by replacing the original color representation in NeRF with a watermarked color representation. Then, a distortion-resistant rendering scheme is designed to guarantee robust message extraction in 2D renderings of NeRF. Our proposed method can directly protect the copyright of NeRF models while maintaining high rendering quality and bit accuracy when compared among optional solutions. Project page: https://luo-ziyuan.github.io/copyrnerf.
Ziyuan Luo, Qing Guo 0005, Ka Chun Cheung, Simon See, Renjie Wan
ICCV4
2023 VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow Estimation
abstract
We introduce VideoFlow, a novel optical flow estimation framework for videos. In contrast to previous methods that learn to estimate optical flow from two frames, VideoFlow concurrently estimates bi-directional optical flows for multiple frames that are available in videos by sufficiently exploiting temporal cues.We first propose a TRi-frame Optical Flow (TROF) module that estimates bi-directional optical flows for the center frame in a three-frame manner. The information of the frame triplet is iteratively fused onto the center frame. To extend TROF for handling more frames, we further propose a MOtion Propagation (MOP) module that bridges multiple TROFs and propagates motion features between adjacent TROFs. With the iterative flow estimation refinement, the information fused in individual TROFs can be propagated into the whole sequence via MOP. By effectively exploiting video information, VideoFlow presents extraordinary performance, ranking 1st on all public benchmarks. On the Sintel benchmark, VideoFlow achieves 1.649 and 0.991 average end-point-error (AEPE) on the final and clean passes, a 15.1% and 7.6% error reduction from the best published results (1.943 and 1.073 from FlowFormer++). On the KITTI-2015 benchmark, VideoFlow achieves an F1-all error of 3.65%, a 19.2% error reduction from the best published result (4.52% from FlowFormer++). Code is released at https://github.com/XiaoyuShi97/VideoFlow.
Xiaoyu Shi 0002, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001
ICCV7
2023 Learning Gabor Texture Features for Fine-Grained Recognition
abstract
Extracting and using class-discriminative features is critical for fine-grained recognition. Existing works have demonstrated the possibility of applying deep CNNs to exploit features that distinguish similar classes. However, CNNs suffer from problems including frequency bias and loss of detailed local information, which restricts the performance of recognizing fine-grained categories. To address the challenge, we propose a novel texture branch as complimentary to the CNN branch for feature extraction. We innovatively utilize Gabor filters as a powerful extractor to exploit texture features, motivated by the capability of Gabor filters in effectively capturing multi-frequency features and detailed local information. We implement several designs to enhance the effectiveness of Gabor filters, including imposing constraints on parameter values and developing a learning method to determine the optimal parameters. Moreover, we introduce a statistical feature extractor to utilize informative statistical information from the signals captured by Gabor filters, and a gate selection mechanism to enable efficient computation by only considering qualified regions as input for texture extraction. Through the integration of features from the Gabor-filter-based texture branch and CNN-based semantic branch, we achieve comprehensive information extraction. We demonstrate the efficacy of our method on multiple datasets, including CUB-200-2011, NA-bird, Stanford Dogs, and GTOS-mobile. State-of-the-art performance is achieved using our approach.
Lanyun Zhu, Tianrun Chen, Jianxiong Yin, Simon See, Jun Liu 0036
ICCV4
2023 Logical Message Passing Networks with One-hop Inference on Atomic Formulas
Zihao Wang 0001, Yangqiu Song, Ginny Y. Wong, Simon See
ICLR4
2023 Self-Consistent Narrative Prompts on Abductive Natural Language Inference
abstract
Chunkit Chan, Xin Liu, Tsz Ho Chan, Jiayang Cheng, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chunkit Chan, Xin Liu 0039, Tsz Ho Chan, Cheng Jiayang, Yangqiu Song, Ginny Y. Wong, Simon See
IJCNLP (1)7
2023 Towards Balanced Active Learning for Multimodal Classification
abstract
Training multimodal networks requires a vast amount of data due to their larger parameter space compared to unimodal networks. Active learning is a widely used technique for reducing data annotation costs by selecting only those samples that could contribute to improving model performance. However, current active learning strategies are mostly designed for unimodal tasks, and when applied to multimodal data, they often result in biased sample selection from the dominant modality. This unfairness hinders balanced multimodal learning, which is crucial for achieving optimal performance. To address this issue, we propose three guidelines for designing a more balanced multimodal active learning strategy. Following these guidelines, a novel approach is proposed to achieve more fair data selection by modulating the gradient embedding with the dominance degree among modalities. Our studies demonstrate that the proposed method achieves more balanced multimodal learning by avoiding greedy sample selection from the dominant modality. Our approach outperforms existing active learning strategies on a variety of multimodal classification tasks. Overall, our work highlights the importance of balancing sample selection in multimodal active learning and provides a practical solution for achieving more balanced active learning for multimodal classification.
Meng Shen 0002, Yizheng Huang 0001, Jianxiong Yin, Heqing Zou, Deepu Rajan, Simon See
ACM Multimedia6
2023 IRCasTRF: Inverse Rendering by Optimizing Cascaded Tensorial Radiance Fields, Lighting, and Materials From Multi-view Images
abstract
We propose an inverse rendering pipeline that simultaneously reconstructs scene geometry, lighting, and spatially-varying material from a set of multi-view images. Specifically, the proposed pipeline involves volume and physics-based rendering, which are performed separately in two steps: exploration and exploitation. During the exploration step, our method utilizes the compactness of neural radiance fields and a flexible differentiable volume rendering technique to learn an initial volumetric field. Here, we introduce a novel cascaded tensorial radiance field method on top of the Canonical Polyadic (CP) decomposition to boost model compactness beyond conventional methods. In the exploitation step, a shading pass that incorporates a differentiable physics-based shading method is applied to jointly optimize the scene's geometry, spatially-varying materials, and lighting, using image reconstruction loss. Experimental results demonstrate that our proposed inverse rendering pipeline, IRCasTRF, outperforms prior works in inverse rendering quality. The final output is highly compatible with downstream applications like scene editing and advanced simulations. Further details are available on the project page: https://ircasrf.github.io/.
Wenpeng Xing, Jie Chen 0026, Ka Chun Cheung, Simon See
ACM Multimedia4
2023 Reducing the Carbon Footprint of Ensemble Weather Forecasting with GPUs
abstract
Climate and Weather Modelling is a highly complex and computationally intensive task which consumes substantial amounts of energy. A desire to improve forecast skill demands further advances in these forecasts, such as increased model fidelity and more comprehensive physical representations of the underlying processes. Another driver towards better forecasts is the goal of uncertainty quantification, with ensembles of forecasts a popular technique. But ensembles place much higher demands on the computational workload as N ensemble members require N times the compute cycles. This leads to an even higher energy demand. This study examines one approach to reducing the energy demands of ensembles by taking advantage of a hardware feature in modern NVIDIA GPUs known as Multi Instance GPU (MIG). This feature allows us to run multiple ensemble members on hardware-isolated GPU slices to maximize efficient use of the GPU resources and subsequently reduce the Carbon Footprint for an Ensemble forecast. We examine both small and large test cases across a range of setups to determine the optimal runtime configuration. Our study shows a 2.S-2.8x reduction in CO2 emissions across all cases which translates into a savings of between 141–171 tonnes of carbon emissions annually per GPU.
Jeff Adie, Terry Yin, Stan Posey, Simon See
TENCON4
2023 A Unified Framework for Factorizing Distributional Value Functions for Multi-Agent Reinforcement Learning
abstract
In fully cooperative multi-agent reinforcement learning (MARL) settings, environments are highly stochastic due to the partial observability of each agent and the continuously changing policies of other agents. To address the above issues, we proposed a unified framework, called DFAC, for integrating distributional RL with value function factorization methods. This framework generalizes expected value function factorization methods to enable the factorization of return distributions. To validate DFAC, we first demonstrate its ability to factorize the value functions of a simple matrix game with stochastic rewards. Then, we perform experiments on all Super Hard maps of the StarCraft Multi-Agent Challenge and six self-designed Ultra Hard maps, showing that DFAC is able to outperform a number of baselines.
Wei-Fang Sun, Cheng-Kuang Lee, Simon See, Chun-Yi Lee
J. Mach. Learn. Res.3
2023 Transfer Learning With Singular Value Decomposition of Multichannel Convolution Matrices
abstract
The task of transfer learning using pretrained convolutional neural networks is considered. We propose a convolution-SVD layer to analyze the convolution operators with a singular value decomposition computed in the Fourier domain. Singular vectors extracted from the source domain are transferred to the target domain, whereas the singular values are fine-tuned with a target data set. In this way, dimension reduction is achieved to avoid overfitting, while some flexibility to fine-tune the convolution kernels is maintained. We extend an existing convolution kernel reconstruction algorithm to allow for a reconstruction from an arbitrary set of learned singular values. A generalization bound for a single convolution-SVD layer is devised to show the consistency between training and testing errors. We further introduce a notion of transfer learning gap. We prove that the testing error for a single convolution-SVD layer is bounded in terms of the gap, which motivates us to develop a regularization model with the gap as the regularizer. Numerical experiments are conducted to demonstrate the superiority of the proposed model in solving classification problems and the influence of various parameters. In particular, the regularization is shown to yield a significantly higher prediction accuracy.
Tak Shing Au Yeung, Ka Chun Cheung, Michael Kwok-Po Ng, Simon See, Andy M. Yip
Neural Comput.4
2023 Attributed Abnormality Graph Embedding for Clinically Accurate X-Ray Report Generation
abstract
Despite the recent success of deep learning models for text generation, generating clinically accurate reports remains challenging. More precisely modeling the relationships of the abnormalities revealed in an X-ray image has been found promising to enhance the clinical accuracy. In this paper, we first introduce a novel knowledge graph structure called an attributed abnormality graph (ATAG). It consists of interconnected abnormality nodes and attribute nodes for better capturing more fine-grained abnormality details. In contrast to the existing methods where the abnormality graph are constructed manually, we propose a methodology to automatically construct the fine-grained graph structure based on annotated X-ray reports and the RadLex radiology lexicon. We then learn the ATAG embeddings as part of a deep model with an encoder-decoder architecture for the report generation. In particular, graph attention networks are explored to encode the relationships among the abnormalities and their attributes. A hierarchical attention attention and a gating mechanism are specifically designed to further enhance the generation quality. We carry out extensive experiments based on the benchmark datasets, and show that the proposed ATAG-based deep model outperforms the SOTA methods by a large margin in ensuring the clinical accuracy of the generated reports.
Sixing Yan, William Kwok-Wai Cheung, Wan Hang Keith Chiu, Terence M. Tong, Ka Chun Cheung, Simon See
IEEE Trans. Medical Imaging6
2022 Visual Marine Debris Detection using Yolov5s for Autonomous Underwater Vehicle
abstract
The trash in the ocean is causing harm to the marine environment. The current most used removal technique is the use of trawlers. It is a highly laborious job and requires high costs as well. With the help of Autonomous Underwater Vehicles (AUVs), removing marine debris could be one of the best and cheapest solutions available. This paper evaluates the use of the You Only Live Once Version 5 Small (YOLOv5s) to compare with the other networks used to identify marine debris. Without fine-tuning the YOLOv5s model in this study, it can achieve a Mean Average Precision (mAP) of 0.681 and an inference speed of 153.9 frames per second. It shows an improvement in mAP compared to YOLOv2, Tiny-YOLO, and Single Shot Multibox Detector.
Cheng Siong Chin, Aloysius Bo Hui Neo, Simon See
ICIS3
2022 Max Fusing Gated Recurrent Units and Ensemble Classifier for Intelligent Acoustic Classification
abstract
The paper presents a wavelet scattering feature extraction using an averaged data augmentation to include unseen devices in training. The multiple classifiers are applied to the extracted features. The outputs from Gated Recurrent Units (GRUs), and the ensemble classifiers are maximally fused to classify sound coming from different devices and scenes. The average device-wise accuracy has more than 5.4% improvement with the proposed GRUs network than its counterpart Long short-term memory (LSTM) in device-wise classification. The device-wise classification accuracy for the proposed max-fusion exhibits approximately 19.1% better than the baseline results in DCASE2020-Task1A. The proposed max-fusion also demonstrates around 22.9% higher classification accuracy in scene-wise classification. Lastly, the proposed max-fusion method produces comparative results with the Snapshot ensembles that won the DCASE2020 Challenges-Task 1A.
Cheng Siong Chin, Simon See
ICIS2
2022 SubeventWriter: Iterative Sub-event Sequence Generation with Coherence Controller
abstract
In this paper, we propose a new task of subevent generation for an unseen process to evaluate the understanding of the coherence of subevent actions and objects.To solve the problem, we design SubeventWriter, a sub-event sequence generation framework with a coherence controller.Given an unseen process, the framework can iteratively construct the subevent sequence by generating one sub-event at each iteration.We also design a very effective coherence controller to decode more coherent sub-events.As our extensive experiments and analysis indicate, SubeventWriter 1 can generate more reliable and meaningful sub-event sequences for unseen processes.
Zhaowei Wang 0003, Hongming Zhang 0009, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP6
2022 Complex Hyperbolic Knowledge Graph Embeddings with Fast Fourier Transform
abstract
The choice of geometric space for knowledge graph (KG) embeddings can have significant effects on the performance of KG completion tasks.The hyperbolic geometry has been shown to capture the hierarchical patterns due to its tree-like metrics, which addressed the limitations of the Euclidean embedding models.Recent explorations of the complex hyperbolic geometry further improved the hyperbolic embeddings for capturing a variety of hierarchical structures.However, the performance of the hyperbolic KG embedding models for nontransitive relations is still unpromising, while the complex hyperbolic embeddings do not deal with multi-relations.This paper aims to utilize the representation capacity of the complex hyperbolic geometry in multi-relational KG embeddings.To apply the geometric transformations which account for different relations and the attention mechanism in the complex hyperbolic space, we propose to use the fast Fourier transform (FFT) as the conversion between the real and complex hyperbolic space.Constructing the attention-based transformations in the complex space is very challenging, while the proposed Fourier transform-based complex hyperbolic approaches provide a simple and effective solution.Experimental results show that our methods outperform the baselines, including the Euclidean and the real hyperbolic embedding models.
Huiru Xiao, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP5
2022 Unified Recurrence Modeling for Video Action Anticipation
abstract
Forecasting future events based on evidence of current conditions is an innate skill of human beings, and key for predicting the outcome of any decision making. In artificial vision for example, we would like to predict the next human action before it happens, without observing the future video frames associated to it. Computer vision models for action anticipation are expected to collect the subtle evidence in the preamble of the target actions. In prior studies recurrence modeling often leads to better performance, the strong temporal inference is assumed to be a key element for reasonable prediction. To this end, we propose a unified recurrence modeling for video action anticipation via message passing framework. The information flow in space-time can be described by the interaction between vertices and edges, and the changes of vertices for each incoming frame reflects the underlying dynamics. Our model leverages self-attention as the building blocks for each of the message passing functions. In addition, we introduce different edge learning strategies that can be end-to-end optimized to gain better flexibility for the connectivity between vertices. Our experimental results demonstrate that our proposed method outperforms previous works on the large-scale EPIC-Kitchen dataset.
Tsung-Ming Tai, Giuseppe Fiameni, Cheng-Kuang Lee, Simon See, Oswald Lanz
ICPR4
2021 ACT: an Attentive Convolutional Transformer for Efficient Text Classification
abstract
Recently, Transformer has been demonstrating promising performance in many NLP tasks and showing a trend of replacing Recurrent Neural Network (RNN). Meanwhile, less attention is drawn to Convolutional Neural Network (CNN) due to its weak ability in capturing sequential and long-distance dependencies, although it has excellent local feature extraction capability. In this paper, we introduce an Attentive Convolutional Transformer (ACT) that takes the advantages of both Transformer and CNN for efficient text classification. Specifically, we propose a novel attentive convolution mechanism that utilizes the semantic meaning of convolutional filters attentively to transform text from complex word space to a more informative convolutional filter space where important n-grams are captured. ACT is able to capture both local and global dependencies effectively while preserving sequential information. Experiments on various text classification tasks and detailed analyses show that ACT is a lightweight, fast, and effective universal text classifier, outperforming CNNs, RNNs, and attentive models including Transformer.
Peixiang Zhong, Kezhi Mao, Dongzhe Wang, Xuefeng Yang, Jianxiong Yin, Simon See
AAAI8
2021 Retrospective Class Incremental Learning
abstract
Existing works study the Class Incremental learning (CIL) problem with the assumption that the data for previous classes are absent, or only a small subset of samples (known as exemplars) are accessible. Differently, we propose a new and practical setting called retrospective CIL, where all the previous data are accessible, but with bounded training budgets for old data replay. Since only a small subset of old samples can be replayed, it brings a new research problem, i.e., dynamically sampling old data along the incremental training process. As incremental learning particularly suffers from catastrophic forgetting, we propose to use the forgettability of the old samples as the sampling priorities to favour the forgotten samples during the dynamic sampling process. To achieve this, we introduce a forgetting rate metric with graph- based propagation to estimate the sample forgettability. The proposed method brings improvements on two benchmark datasets.
Qingyi Tao, Chen Change Loy, Jianfei Cai 0001, ZongYuan Ge, Simon See
ICME5
2021 Recent advance in machine learning for partial differential equation
Ka Chun Cheung, Simon See
CCF Trans. High Perform. Comput.2
2021 Challenges and opportunities for a hybrid modelling approach to earth system science
Simon See, Jeff Adie
CCF Trans. High Perform. Comput.1
2021 Exploiting inter-frame regional correlation for efficient action recognition
Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Expert Syst. Appl.5
2021 PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition
Yuecong Xu, Haozhi Cao, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Neurocomputing6
2021 Effective action recognition with embedded key point shifts
Haozhi Cao, Yuecong Xu, Jianfei Yang 0001, Kezhi Mao, Jianxiong Yin, Simon See
Pattern Recognit.6
2019 DeepHunter: a coverage-guided fuzz testing framework for deep neural networks
abstract
The past decade has seen the great potential of applying deep neural network (DNN) based software to safety-critical scenarios, such as autonomous driving. Similar to traditional software, DNNs could exhibit incorrect behaviors, caused by hidden defects, leading to severe accidents and losses. In this paper, we propose DeepHunter, a coverage-guided fuzz testing framework for detecting potential defects of general-purpose DNNs. To this end, we first propose a metamorphic mutation strategy to generate new semantically preserved tests, and leverage multiple extensible coverage criteria as feedback to guide the test generation. We further propose a seed selection strategy that combines both diversity-based and recency-based seed selection. We implement and incorporate 5 existing testing criteria and 4 seed selection strategies in DeepHunter. Large-scale experiments demonstrate that (1) our metamorphic mutation strategy is useful to generate new valid tests with the same semantics as the original seed, by up to a 98% validity ratio; (2) the diversity-based seed selection generally weighs more than recency-based seed selection in boosting the coverage and in detecting defects; (3) DeepHunter outperforms the state of the arts by coverage as well as the quantity and diversity of defects identified; (4) guided by corner-region based criteria, DeepHunter is useful to capture defects during the DNN quantization for platform migration.
Xiaofei Xie, Lei Ma 0003, Felix Juefei-Xu, Minhui Xue 0001, Hongxu Chen 0001, Yang Liu 0003, Jianjun Zhao 0001, Bo Li 0026, Jianxiong Yin, Simon See
ISSTA10
2019 Improving Deep Lesion Detection Using 3D Contextual and Spatial Attention
Qingyi Tao, ZongYuan Ge, Jianfei Cai 0001, Jianxiong Yin, Simon See
MICCAI (6)5
2018 Deep Learning with Evolutionary and Genomic Profiles for Identifying Cancer Subtypes
abstract
Cancer subtype identification is an unmet need in precision diagnosis. Recently, evolutionary conservation has been indicated containing understandable signatures for functional significance in cancers. However, the importance of evolutionary conservation in distinguishing cancer subtypes remains unclear. Here, we identified the evolutionarily conserved genes (i.e., core gene) and observed that they are mainly involved in the pathways relevant to cell growth and metabolisms. By using these core genes, we integrated their evolutionary and genomic profiles with deep learning to develop a feature-based strategy (FES) and an image-based strategy (IMS). In comparison with FES using the random set and the strategy using the PAM50 classifier, core gene set-based FES has higher accuracy for identifying breast cancer subtypes. Moreover, the IMS with data augmentation yields better performance than the other strategies. Comprehensive analysis of eight TCGA cancer data demonstrates that our evolutionary conservation-based models provide a valid and helpful approach to identify cancer subtypes and the core gene set offers distinguishable clues of cancer subtypes.
Chun-Yu Lin 0003, Peiying Ruan, Ruiming Li, Jinn-Moon Yang, Simon See, Tatsuya Akutsu
BIBE5
2018 Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
abstract
It is desirable to train convolutional networks (CNNs) to run more efficiently during inference. In many cases however, the computational budget that the system has for inference cannot be known beforehand during training, or the inference budget is dependent on the changing real-time resource availability. Thus, it is inadequate to train just inference-efficient CNNs, whose inference costs are not adjustable and cannot adapt to varied inference budgets. We propose a novel approach for cost-adjustable inference in CNNs - Stochastic Downsampling Point (SDPoint). During training, SDPoint applies feature map downsampling to a random point in the layer hierarchy, with a random downsampling ratio. The different stochastic downsampling configurations known as SDPoint instances (of the same model) have computational costs different from each other, while being trained to minimize the same prediction loss. Sharing network parameters across different instances provides significant regularization boost. During inference, one may handpick a SDPoint instance that best fits the inference budget. The effectiveness of SDPoint, as both a cost-adjustable inference approach and a regularizer, is validated through extensive experiments on image classification.
Jason Kuen, Xiangfei Kong, Zhe Lin 0001, Gang Wang 0012, Jianxiong Yin, Simon See, Yap-Peng Tan
CVPR6
2018 Optimizing Deep Learning Frameworks Incrementally to Get Linear Speedup: A Comparison Between IPoIB and RDMA Verbs
abstract
Deeper models and larger datasets are two major ingredients for applying deep learning (DL) on real-world problems, which inevitably shifts model training from on a single GPU card to on a GPU clusters due to limited GPU memory and time-to-solution requirements. High-speed low-latency RDMA-capable network fabrics like Infiniband and RoCE play an important role on coping with enoumous data exchanged during training. DL frameworks are built upon these fabrics with various APIs including IPoIB, MPI and RDMA Verbs. Tradeoffs are made between performance and usability when adapting DL frameworks onto RDMA-capable networks, which may result in high-performance yet hard-to-maintain and hard-to-merge code if improper design choices are made. This paper presents our approach to adapt MXNet, a modular versatile DL framework onto RDMA-capable networks. Dividing the training process on MXN et into P2P communication and A11Reduce commnunication, we add incremental optimizations on its message passing code. Experiments show that our approach exhibits near-linear speedups, whose parallel efficiency reaches 96% compared to 53% of the original IPoIB version when scaling to 100 GPU cards. In contrast to other MPI-based porting approach, our modifications are limited within MXNet's Parameter Server module, which is transparent for upper-layer operations, thus making no sacrifice on features like auto recovery and user-controlled consistency view.
Jianwen Wei, Yichao Wang 0001, Minhua Wen, Simon See, James Lin 0001
ICPADS5
2018 Fast MPEG-CDVS Encoder With GPU-CPU Hybrid Computing
abstract
The compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search.
Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
IEEE Trans. Image Process.7
2017 Real-time GPU-accelerated social media sentiment processing and visualization
abstract
Data visualization is an important aspect of data analytics in an age where decisions are all based on information. Approaches in data visualization, particularly those that have the capability of processing large-scale textual datasets and visualize them as structured information in real-time can be useful for monitoring trends in social media. In this article, we present our GPU accelerated project, which uses CUDA to distribute and parallelize the processing and analysis of textual data in order to visualize information in real-time, or close to real-time as a foundational system for the future of real-time applications which monitors trends in social media, applicable to political elections, social media analytics, and other needs in computational social sciences which are time-critical.
Eugene Ch'ng, Simon See
DS-RT3
2017 CRNN: A Joint Neural Network for Redundancy Detection
abstract
This paper proposes a novel framework for detecting redundancy in supervised sentence categorisation. Unlike traditional singleton neural network, our model incorporates character- aware convolutional neural network (Char-CNN) with character-aware recurrent neural network (Char-RNN) to form a convolutional recurrent neural network (CRNN). Our model benefits from Char-CNN in that only salient features are selected and fed into the integrated Char-RNN. Char-RNN effectively learns long sequence semantics via sophisticated update mechanism. We compare our framework against the state-of-the- art text classification algorithms on four popular benchmarking corpus. For instance, our model achieves competing precision rate, recall ratio, and F1 score on the Google-news data-set. For twenty- news-groups data stream, our algorithm obtains the optimum on precision rate, recall ratio, and F1 score. For Brown Corpus, our framework obtains the best F1 score and almost equivalent precision rate and recall ratio over the top competitor. For the question classification collection, CRNN produces the optimal recall rate and F1 score and comparable precision rate. We also analyse three different RNN hidden recurrent cells' impact on performance and their runtime efficiency. We observe that MGU achieves the optimal runtime and comparable performance against GRU and LSTM. For TFIDF based algorithms, we experiment with word2vec, GloVe, and sent2vec embeddings and report their performance differences.
Xinyu Fu 0001, Eugene Ch'ng, Uwe Aickelin, Simon See
SMARTCOMP4
2016 Learning Common and Specific Features for RGB-D Semantic Segmentation with Deconvolutional Networks
Zhenhua Wang 0002, Dacheng Tao, Simon See, Gang Wang 0012
ECCV (5)4
2016 Performance and Portability Studies with OpenACC Accelerated Version of GTC-P
abstract
Accelerator-based heterogeneous computing is of paramount importance to High Performance Computing. The increasing complexity of the cluster architectures requires more generic, high-level programming models. OpenACC is a directive-based parallel programming model, which provides performance on and portability across a wide variety of platforms, including GPU, multicore CPU, and many-core processors. GTC-P is a discovery-science-capable real-world application code based on the Particle-In-Cell (PIC) algorithm that is well-established in the HPC area. Several native versions of GTC-P have been developed for supercomputers on TOP500 with different architectures, including Titan, Mira, etc. Motivated by the state-of-art portability, we implemented the first OpenACC version of GTC-P and evaluated its performance portability across NVIDIA GPUs, Intel x86 and OpenPOWER CPUs. In this paper, we also proposed two key optimization methods for OpenACC implementation of PIC algorithm on multicore CPU and GPU including removing atomic operation and taking advantage of shared memory. OpenACC shows both impressive productivity and performance in a perspective of portability and scalability. The OpenACC version achieves more than 90% performance compared with the native versions with only about 300 LOC.
Yueming Wei, Yichao Wang 0001, Linjin Cai, William Tang 0002, Bei Wang 0002, Stéphane Ethier, Simon See, James Lin 0001
PDCAT7
2015 An Evaluation of Unified Memory Technology on NVIDIA GPUs
abstract
Unified Memory is an emerging technology which is supported by CUDA 6.X. Before CUDA 6.X, the existing CUDA programming model relies on programmers to explicitly manage data between CPU and GPU and hence increases programming complexity. CUDA 6.X provides a new technology which is called as Unified Memory to provide a new programming model that defines CPU and GPU memory space as a single coherent memory (imaging as a same common address space). The system manages data access between CPU and GPU without explicit memory copy functions. This paper is to evaluate the Unified Memory technology through different applications on different GPUs to show the users how to use the Unified Memory technology of CUDA 6.X efficiently. The applications include Diffusion3D Benchmark, Parboil Benchmark Suite, and Matrix Multiplication from the CUDA SDK Samples. We changed those applications to corresponding Unified Memory versions and compare those with the original ones. We selected the NVIDIA Keller K40 and the Jetson TK1, which can represent the latest GPUs with Keller architecture and the first mobile platform of NVIDIA series with Keller GPU. This paper shows that Unified Memory versions cause 10% performance loss on average. Furthermore, we used the NVIDIA Visual Profiler to dig the reason of the performance loss by the Unified Memory technology.
Guanghao Jin, Xuewen Cui, Simon See
CCGRID4
2014 Gravitational Search Algorithm Using CUDA
abstract
Many scientific and technical problems with massive computation requirements could benefit from the Graphics Processing Units (GPUs) using Compute Unified Device Architecture (CUDA) for high speed processing. Gravitational Search Algorithm (GSA) is a population-based metaheuristic algorithm that can be effectively implemented on GPU to reduce the execution time. In this paper we discuss possible approaches to parallelize GSA on graphics hardware using CUDA. An in-depth study of the computation efficiency of parallel algorithms and capability to effectively exploit the architecture of GPU is performed. Additionally, a comparative study of parallel and sequential GSA was carried out on a set of standard benchmark optimization functions. The results show a significant speedup that re-emphasizes the utility of CUDA based implementation for complex and computationally intensive parallel applications.
Amirreza Zarrabi, Ettikan Kandasamy Karuppiah, Keh Kok Yong, Ngo Chuan Hai, Simon See
PDCAT5
2013 GPU-accelerated adaptive compression framework for genomics data
abstract
Genomics data is being produced at an unprecedented rate, especially in the context of clinical applications and grand challenge questions. There are various types of data in genomics research, most of which are stored as plain text tables. A data compression framework tailored to this file type is introduced in this paper, featuring a combination of generic compression algorithms, GPU acceleration, and column-major storage. This approach is the first to achieve both compression and decompression rates of around 100MB/s on commodity hardware without compromising compression ratio. By selecting appropriate compression schemes for each column of data, this framework efficiently exploits data redundancy while remaining applicable to a wide range of formats. The GPU-accelerated implementation also properly exploits the parallelism of compression algorithms. Finally, this paper presents a novel first-order Markov model based transformation, with evidence that it is at least as effective as Burrows-Wheeler and Move-To-Front in some contexts.
GuiXin Guo, Zhiqiang Ye, Bingqiang Wang, Mian Lu, Simon See, Rui Mao 0001
IEEE BigData7
2011 Understanding Off-Chip Memory Contention of Parallel Programs in Multicore Systems
abstract
Memory contention is an important performance issue in current multicore architectures. In this paper, we focus on understanding how off-chip memory contention affects the performance of parallel applications. Using measurements conducted on state-of-the-art multicore systems, we observed that off-chip memory traffic is not always bursty, as it was previously reported in literature. Burstiness depends on the problem size. Small problem sizes lead to bursty memory traffic, and generate small off-chip contention. In contrast, when large program sizes cause memory contention, the memory traffic is non-bursty. Based on these observations, we propose an analytical model that relates the growth of memory contention to the number of active cores and to the problem size, for both uniform (UMA) and non-uniform memory access (NUMA) systems. Our model differs from measurements on average by less than 14\%. Contention for off-chip memory grows exponentially with the number of active cores, but adding additional memory controllers reduces the memory contention. For programs such as the penta diagonal solver SP from NPB benchmark, with a large matrix of $162^3$ elements (input size C), our analysis shows that memory contention increases the total number of processor cycles to execute the program by more than ten times on a machine with 24 cores.
Bogdan Marius Tudor, Yong Meng Teo, Simon See
ICPP3
2009 Data mining analysis to validate performance tuning practices for HPL
abstract
Applications performance is a criterion for system evaluation, and hence performance tuning for these applications is of great interest. One such benchmark application is High Performance Linpack (HPL). Although guidelines exist for HPL tuning, validating these guidelines on various systems is a challenging task as a large number of configurations need to be tested. In this work, we use data mining analysis to reduce the number of configurations to be tested in validating the HPL tuning guidelines on the Ranger System. We validate that NB, P and Q are the three most important parameters to tune HPL, and that PMAP does not have a significant impact on HPL performance. We also validate the practice of tuning HPL at small N using data mining analysis. We find that the value of N selected for tuning should not be significantly smaller than the largest N that can fit into the system memory. Our results indicate that data mining could be further applied to application performance tuning.
Tuan Zea Tan, Rick Siow Mong Goh, Verdi March, Simon See
CLUSTER4
2009 Towards Predictive Modeling of Message-Passing Communication
abstract
Communication has been shown to be a performance bottleneck and a limiting factor of many large parallel applications. As such, predicting the application scalability necessitates a communication performance model. This paper investigates the LogGP communication performance model for predicting message-passing communications when the system configuration (i.e., number of nodes) is varied. The cost functions for the message-passing operations are based on MVAPICH2 1.0, and the experiments are conducted on the Ranger system using up to 256 nodes connected with an InfiniBand net-work. For point-to-point communications, we observe that the LogGP model accurately predicts the communication performance. However, the results for three collective operations, i.e.,MPI_Barrier, MPI_Alltoall, andMPI_Bcast, are varying. ForMPI_Bcast, the LogGP model is able to predict its scalability up to 256 nodes, and the prediction error is at most a factor of two on 256 nodes. For the remaining collectives, the scalability— bar that ofMPI_Alltoallon small messages (m= 2 bytes) — is predicted by LogGP, but the prediction error for 256 nodes is 3.5-12 times of the measured performance.
Verdi March, Vijayaraghavan Murali, Yong Meng Teo, Simon See, James T. Himer
HPCC4
2009 An approach for matching communication patterns in parallel applications
abstract
Interprocessor communication is an important factor in determining the performance scalability of parallel systems. The communication requirements of a parallel application can be quantified to understand its communication pattern and communication pattern similarities among applications can be determined. This is essential for the efficient mapping of applications on parallel systems and leads to better interprocessor communication implementation among others. This paper proposes a methodology to compare the communication pattern of distributed-memory programs. Communication correlation coefficient quantifies the degree of similarity between two applications based on the communication metrics selected to characterize the applications. To capture the network topology requirements, we extract the communication graph of each applications and quantities this similarity. We apply this methodology to four applications in the NAS parallel benchmark suite and evaluate the communication patterns by studying the effects of varying problem size and the number of logical processes (LPs).
Chao Ma 0008, Yong Meng Teo, Verdi March, Naixue Xiong, I. R. Pop, Yanxiang He, Simon See
IPDPS7
2008 Grid Discovery Zone: Virtualization for Exploiting Easy-Management and High-Utilization in Grid
abstract
In this paper, we present a new virtualization strategy for grid called grid discovery zone (GDZ). The virtualization ability of the GDZ is used to exploit the resource utilization and scheduling, easy management and isolation for users in cluster or grid environment. We describe the infrastructure and architecture of GDZ in this paper, which uses Solaris Zone, Sun Grid Engine, Sun Studio 12, JES portal server, ZFS, and other Sun related technologies and product. Using the infrastructure of GDZ, many applications in different industries can be discovered in a single multi-core machine or grid/cluster consisting of multi-core nodes. The load can also be optimal balanced in the GDZ using the powerful resource control and scheduling ability of Solaris Zones and Sun Grid Engine.
Ruihua Zhang, Simon See, Che Cheung, Chi-Hung Chan
HPCC3
2008 Survey on Parallel Programming Model
Henry Kasim, Verdi March, Rita Zhang, Simon See
NPC4
2007 A multi-dimensional scheduling scheme in a Grid computing environment
Benjamin Khoo Boon Tat, Bharadwaj Veeravalli, Terence Hung, Simon See
J. Parallel Distributed Comput.4
2006 Service Registry Discovery using GridSearch P2P Framework
Melvin Koh, Jie Song 0005, Simon See
CCGRID4
2006 A Co-ordinate Based Resource Allocation Strategy for Grid Environments
abstract
In this paper, we propose a novel resource scheduling strategy, referred to as the Multi-Resource Scheduling (MRS) algorithm, which is capable of handling several resources to be used among jobs that arrive at a Grid Computing Environment. We propose a model in which the job and resource characteristics are captured together and are used in the scheduling strategy. To do so, we introduce the concept of virtual map and resource potential. Based on the proposed model, simulations with realistic workload traces were conducted to quantify the performance. We compare our strategy with some of the commonly used algorithms, and show that MRS renders a higher performance in all cases. Our experimental results clearly show that MRS outperforms other strategies and we highlight the impact and importance of our strategy.
Benjamin Khoo Boon Tat, Bharadwaj Veeravalli, Terence Hung, Simon See
CCGRID4
2005 Performance Monitoring for Distributed Service Oriented Grid Architecture
Melvin Koh, Jie Song 0005, Simon See
ICA3PP4
2005 Sensor Grid: Integration ofWireless Sensor Networks and the Grid
abstract
Wireless sensor networks have emerged as an exciting technology for a wide range of important applications that acquire and process information from the physical world. Grid computing has evolved as a standards-based approach for coordinated resource sharing. Sensor grids combine these two promising technologies by extending the grid computing paradigm to the sharing of sensor resources in wireless sensor networks. There are several issues and challenges in the design of sensor grids. In this paper, we propose a sensor grid architecture, called the scalable proxy-based architecture for sensor grid (SPRING), to address these design issues. We also developed a sensor grid testbed to study the design issues of sensor grids and to improve our sensor grid architecture design
Hock-Beng Lim, Yong Meng Teo, Protik Mukherjee, Vinh The Lam, Weng-Fai Wong, Simon See
LCN6
2004 Application Partitionability in Computational Grids
abstract
Summary form only given. Computation grids provide large volume of computing resources and have become an attractive alternative for scientific computing. It is desired that the applications are developed to utilize the globally distributed computing resources. Partitioning is one important way to achieve this goal. However, whether partitioning an application for computational grids is profitable or not is a basic problem and it is not fully addressed by the existing work. We call it partitionability problem. In our work, we try to quantify this problem and define the concept of computation density and partitionability based on the criteria of response time. We theoretically analyze the relationship between partitionability and application attributes such as I/O and internal communication data size. We show that with given workloads, those applications with higher computation density result in higher partitionability. We also propose a global resource registration mechanism so that the up-to-date resource information is available in partitioning. Our experiments with the simulated map image matching application shows that the proposed concept and framework improve the response time of the application by almost 40%.
Simon See
IPDPS2
2004 Benchmark Performance on Cluster Grid with NGB
abstract
Summary form only given. Along with the development of computational grids, it is more and more desirable to evaluate performance of the grids by using standard benchmarks. However, few benchmark suites have been developed and widely used so far, which results in an obstacle to the wide acceptance of grid in current e-science age. Meanwhile, the dynamic and heterogeneous nature of grid environment makes performance evaluation and its generic applicability challenging. NGB (NAS Grid Benchmarks) is a recently proposed benchmark suite for computational grids. In this paper, we use NGB to evaluate the performance of cluster grid with Sun Grid Engine and Globus, which are basic components of the campus grid in our university. The performance metrics we used include turnaround time and CPU utilization. Our preliminary experimental results show that the Sun Grid Engine's overhead is less than that of Globus because Globus has more procedures in job submission, both of them become negligible when the problem size increases. Moreover, in NGB less loosely coupled grid applications may not necessary result in higher resource utilization, due to various characteristics of applications.
Simon See, Jie Song 0005, Appie Stoelwinder, Hoon Kang Neo
IPDPS2
2004 Agent-Mediated Genetic Super-Scheduling in Grid Environments
Gang Chen 0002, Simon See, Jie Song 0005
PDCAT3
2004 Scheduler Oriented Grid Performance Evaluation
Simon See
PDCAT2
2004 Investigating Super Scheduling Algorithms for Grid Computing: A Simulation Approach
Jie Song 0005, Simon See
PDCAT3
2004 A prototype of distributed molecular visualization on computational grids
Huabing Zhu, Tony Kai Yun Chan, Lizhe Wang 0001, Wentong Cai 0001, Simon See
Future Gener. Comput. Syst.5
2003 DPBP: A Sort-First Parallel Rendering Algorithm for Distributed Rendering Environments
abstract
In some visualization systems, the data and computational resources are distributed globally and users need to interact with these resources easily and efficiently. Real-time rendering for massive datasets is a computation intensive task. one solution is to distributes the rendering tasks over a set of computation units to achieve high rendering performance. This paper presents a recursive sort-first partitioning algorithm named Dynamic Pixel Bucket Partition (DPBP) for parallel rendering alone with their implementation and performance in a distributed rendering environment. This algorithm distributes rendering work loads evenly to individual rendering units to achieve fast, high quality rendering of massive data. Test results in a multi-cluster environment demonstrate the practicality of this rendering algorithm.
Huabing Zhu, Kai-Yun Chan, Lizhe Wang 0001, Wentong Cai 0001, Simon See
CW5
2003 A Distributed Rendering Environment for Massive Data on Computational Grids
abstract
Scientific visualization, especially for massive data sets, has emerged in different disciplines recently. Generally, distributed scientific visualization applications require multiple resources, e.g., high-end computing resources to process data, high speed network for data transfer and large size database for data storage. Furthermore, these applications will meet research challenges, e.g., heterogeneous resources, geographically distributed environment and considerable communication delay. We study an application of distributed massive data rendering. We present infrastructure of the distributed rendering environment and explain how grid technologies are used in this application. Dynamic pixel bucket partition (DPBP) algorithm is a new algorithm proposed for task allocation of distributed rendering application in computational grids. Experiments in real test ted shows the performance of DPBP algorithm and the framework.
Huabing Zhu, Lizhe Wang 0001, Kai-Yun Chan, Wentong Cai 0001, Simon See
Peer-to-Peer Computing5