Alvin Chan

dblp:163/6518 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
abstract
Multimodal learning often relies on aligning representations across modalities to enable effective information integration—an approach traditionally assumed to be universally beneficial. However, prior research has primarily taken an observational approach, examining naturally occurring alignment in multimodal data and exploring its correlation with model performance, without systematically studying the direct effects of explicitly enforced alignment between representations of different modalities. In this work, we investigate how explicit alignment influences both model performance and representation alignment under different modality-specific information structures. Specifically, we introduce a controllable contrastive learning module that enables precise manipulation of alignment strength during training, allowing us to explore when explicit alignment improves or hinders performance. Our results on synthetic and real datasets under different data characteristics show that the impact of explicit alignment on the performance of unimodal models is related to the characteristics of the data: the optimal level of alignment depends on the amount of redundancy between the different modalities. We can find an optimal alignment strength that balances modality-specific signals and shared redundancy in the mixed information distributions. This work can help practitioners on when and how to enforce alignment for optimal unimodal encoder performance.
Wanlong Fang, Alvin Chan
AAAI3
2025 How to Make Large Language Models Generate 100% Valid Molecules?
abstract
Molecule generation is key to drug discovery and materials science, enabling the design of novel compounds with specific properties.Large language models (LLMs) can learn to perform a wide range of tasks from just a few examples.However, generating valid molecules using representations like SMILES is challenging for LLMs in few-shot settings.In this work, we explore how LLMs can generate 100% valid molecules.We evaluate whether LLMs can use SELFIES, a representation where every string corresponds to a valid molecule, for valid molecule generation but find that LLMs perform worse with SELFIES than with SMILES.We then examine LLMs' ability to correct invalid SMILES and find their capacity limited.Finally, we introduce SmiSelf, a cross-chemical language framework for invalid SMILES correction.SmiSelf converts invalid SMILES to SELFIES using grammatical rules, leveraging SELFIES' mechanisms to correct the invalid SMILES.Experiments show that SmiSelf ensures 100% validity while preserving molecular characteristics and maintaining or even enhancing performance on other metrics.SmiSelf helps expand LLMs' practical applications in biomedicine and is compatible with all SMILES-based generative models.Code is available at https: //github.com/wentao228/SmiSelf.
Wen Tao, Jing Tang 0004, Alvin Chan, Bryan Hooi, Baolong Bi, Nanyun Peng 0001, Yuansheng Liu, Yiwei Wang 0001
EMNLP3
2025 Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
abstract
The remarkable performance of Large Language Models (LLMs) can be enhanced with test-time computation, which relies on external tools and even other deep learning models. However, existing approaches for integrating non-text modality representations into LLMs typically require additional costly supervised training, restricting on-the-fly adaptation to new domains and modalities. In this work, we explore the feasibility of integrating representations from non-text foundational models (FMs) into text-based LLMs in a training-free manner. We propose In-Context Representation Learning (ICRL) as a proof-of-concept to allow LLMs to adaptively utilize non-text modality representations with few-shot learning. Unlike traditional in-context learning, which incorporates text-label pairs, ICRL replaces text inputs with FM representations, enabling the LLM to perform multi-modal inference without fine-tuning. We evaluate ICRL on a suite of tasks in the molecular domain, investigating three core research questions: (i) how to map FM representations into LLMs in a training-free manner, (ii) what factors influence ICRL performance, and (iii) what mechanisms underlie the effectiveness of ICRL. To the best of our knowledge, ICRL is the first training-free framework for integrating non-text modality representations into text-based LLMs, presenting a promising direction for adaptable, multi-modal generalization.
Wanlong Fang, Jonathan Woo, Paridhi Latawa, Deepak A. Subramanian, Alvin Chan
NeurIPS6
2022 How Does Frequency Bias Affect the Robustness of Neural Image Classifiers against Common Corruption and Adversarial Perturbations?
abstract
Model robustness is vital for the reliable deployment of machine learning models in real-world applications. Recent studies have shown that data augmentation can result in model over-relying on features in the low-frequency domain, sacrificing performance against low-frequency corruptions, highlighting a connection between frequency and robustness. Here, we take one step further to more directly study the frequency bias of a model through the lens of its Jacobians and its implication to model robustness. To achieve this, we propose Jacobian frequency regularization for models' Jacobians to have a larger ratio of low-frequency components. Through experiments on four image datasets, we show that biasing classifiers towards low (high)-frequency components can bring performance gain against high (low)-frequency corruption and adversarial perturbation, albeit with a tradeoff in performance for low (high)-frequency corruption. Our approach elucidates a more direct connection between the frequency bias and robustness of deep learning models.
Alvin Chan, Yew-Soon Ong, Clement Tan
IJCAI1
2022 Anti-Forensic Deepfake Personas and How To Spot Them
abstract
Forensic systems have recently been studied to detect and prevent deepfakes abuses such as fake personas, frauds, misinformation, or harassment. At the same time, anti-forensic deepfakes are being investigated to understand the gaps in these detection systems and pave the way for improvement. In this paper, we investigate the threat of anti-forensic fake personas, where a fraudster creates a fake personal profile from multiple anti-forensic deepfake images portraying a single identity. To comprehensively study this threat model, we consider three approaches that an attacker may use to conduct such attacks, encompassing both white- and black-box scenarios. A range of defense strategies is then proposed with the aim to improve the robustness of current forensic systems against such threats. Experimental result shows that while the attacks can bypass current detection, our proposed defense approaches that consider the multi-image nature of a fake persona can effectively mitigate this threat by lowering the attack success rate.
Nguyen Hong Ngoc, Alvin Chan, Huynh Thi Thanh Binh, Yew-Soon Ong
IJCNN2
2022 Breaking Neural Reasoning Architectures With Metamorphic Relation-Based Adversarial Examples
abstract
The ability to read, reason, and infer lies at the heart of neural reasoning architectures. After all, the ability to perform logical reasoning over language remains a coveted goal of Artificial Intelligence. To this end, models such as the Turing-complete differentiable neural computer (DNC) boast of real logical reasoning capabilities, along with the ability to reason beyond simple surface-level matching. In this brief, we propose the first probe into DNC's logical reasoning capabilities with a focus on text-based question answering (QA). More concretely, we propose a conceptually simple but effective adversarial attack based on metamorphic relations. Our proposed adversarial attack reduces DNCs' state-of-the-art accuracy from 100% to 1.5% in the worst case, exposing weaknesses and susceptibilities in modern neural reasoning architectures. We further empirically explore possibilities to defend against such attacks and demonstrate the utility of our adversarial framework as a simple scalable method to improve model adversarial robustness.
Alvin Chan, Lei Ma 0003, Felix Juefei-Xu, Yew-Soon Ong, Xiaofei Xie, Minhui Xue 0001, Yang Liu 0003
IEEE Trans. Neural Networks Learn. Syst.1
2021 CoCon: A Self-Supervised Approach for Controlled Text Generation
Alvin Chan, Yew-Soon Ong, Bill Pung, Aston Zhang
ICLR1
2021 Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters
Aston Zhang, Yi Tay, Shuai Zhang 0007, Alvin Chan, Anh Tuan Luu, Siu Cheung Hui, Jie Fu 0001
ICLR4
2021 Deep Extrapolation for Attribute-Enhanced Generation
abstract
Attribute extrapolation in sample generation is challenging for deep neural networks operating beyond the training distribution. We formulate a new task for extrapolation in sequence generation, focusing on natural language and proteins, and propose GENhance, a generative framework that enhances attributes through a learned latent space. Trained on movie reviews and a computed protein stability dataset, GENhance can generate strongly-positive text reviews and highly stable protein sequences without being exposed to similar data during training. We release our benchmark tasks and models to contribute to the study of generative modeling extrapolation and data-driven design in biology and chemistry.
Alvin Chan, Ali Madani, Ben Krause, Nikhil Naik 0002
NeurIPS1
2021 Self-Instantiated Recurrent Units with Dynamic Soft Recursion
abstract
While standard recurrent neural networks explicitly impose a chain structure on different forms of data, they do not have an explicit bias towards recursive self-instantiation where the extent of recursion is dynamic. Given diverse and even growing data modalities (e.g., logic, algorithmic input and output, music, code, images, and language) that can be expressed in sequences and may benefit from more architectural flexibility, we propose the self-instantiated recurrent unit (Self-IRU) with a novel inductive bias towards dynamic soft recursion. On one hand, theSelf-IRU is characterized by recursive self-instantiation via its gating functions, i.e., gating mechanisms of the Self-IRU are controlled by instances of the Self-IRU itself, which are repeatedly invoked in a recursive fashion. On the other hand, the extent of the Self-IRU recursion is controlled by gates whose values are between 0 and 1 and may vary across the temporal dimension of sequences, enabling dynamic soft recursion depth at each time step. The architectural flexibility and effectiveness of our proposed approach are demonstrated across multiple data modalities. For example, the Self-IRU achieves state-of-the-art performance on the logical inference dataset [Bowman et al., 2014] even when comparing with competitive models that have access to ground-truth syntactic information.
Aston Zhang, Yi Tay, Yikang Shen, Alvin Chan, Shuai Zhang 0007
NeurIPS4
2021 Player Identification in Hockey Broadcast Videos
Alvin Chan, Martin D. Levine, Mehrsan Javan Roshtkhari
Expert Syst. Appl.1
2020 Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social Preferences
abstract
Understanding human preferences, along with cultural and social nuances, lives at the heart of natural language understanding.Concretely, we present a new task and corpus for learning alignments between machine and human preferences.Our newly introduced problem is concerned with predicting the preferable options from two sentences describing scenarios that may involve social and cultural situations.Our problem is framed as a natural language inference task with crowd-sourced preference votes by human players, obtained from a gamified voting platform.We benchmark several state-of-the-art neural models, along with BERT and friends on this task.Our experimental results show that current state-ofthe-art NLP models still leave much room for improvement.
Yi Tay, Donovan Ong, Jie Fu 0001, Alvin Chan, Nancy F. Chen, Anh Tuan Luu, Christopher Joseph Pal
ACL4
2020 What It Thinks Is Important Is Important: Robustness Transfers Through Input Gradients
abstract
Adversarial perturbations are imperceptible changes to input pixels that can change the prediction of deep learning models. Learned weights of models robust to such perturbations are previously found to be transferable across different tasks but this applies only if the model architecture for the source and target tasks is the same. Input gradients characterize how small changes at each input pixel affect the model output. Using only natural images, we show here that training a student model's input gradients to match those of a robust teacher model can gain robustness close to a strong baseline that is robustly trained from scratch. Through experiments in MNIST, CIFAR-10, CIFAR-100 and Tiny-ImageNet, we show that our proposed method, input gradient adversarial matching, can transfer robustness across different tasks and even across different model architectures. This demonstrates that directly targeting the semantics of input gradients is a feasible way towards adversarial robustness.
Alvin Chan, Yi Tay, Yew-Soon Ong
CVPR1
2020 Jacobian Adversarially Regularized Networks for Robustness
Alvin Chan, Yi Tay, Yew-Soon Ong
ICLR1
2015 STEP: A Time-Efficient Tag Searching Protocol in Large RFID Systems
abstract
The radio frequency identification (RFID) technology is greatly revolutionizing applications such as warehouse management and inventory control in retail industry. In large RFID systems, an important and practical issue is tag searching: Given a particular set of tags called wanted tags, tag searching aims to determine which of them are currently present in the system and which are not. As an RFID system usually contains a large number of tags, the intuitive solution that collects IDs of all the tags in the system and compares them with the wanted tag IDs to obtain the result is highly time inefficient. In this paper, we design a novel technique called testing slot, with which a reader can quickly figure out which wanted tags are absent from its interrogation region without tag ID transmissions. The testing slot technique thus greatly reduces transmission overhead during the searching process. Based on this technique, we propose two protocols to perform time-efficient tag searching in practical large RFID systems containing multiple readers. In our protocols, each reader first employs the testing slot technique to obtain its local searching result by iteratively eliminating wanted tags that are absent from its interrogation region. The local searching results of readers are then combined to form the final searching result. The proposed protocols outperform existing solutions in both time efficiency and searching precision. Simulation results show that, compared with the state-of-the-art solution, our best protocol reduces execution time by up to 60 percent, meanwhile promotes the searching precision by nearly an order of magnitude.
Xuan Liu 0001, Bin Xiao 0001, Shigeng Zhang, Kai Bu, Alvin Chan
IEEE Trans. Computers5