Ginny Y. Wong

dblp:78/8568 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0001-7432-8496ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 XToM: Exploring the Multilingual Theory of Mind for Large Language Models
abstract
Theory of Mind (ToM), the ability to infer mental states in others, is pivotal for human social cognition. Existing evaluations of ToM in LLMs are largely limited to English, neglecting the linguistic diversity that shapes human cognition. This limitation raises a critical question: can LLMs exhibit Multilingual Theory of Mind, which is the capacity to reason about mental states across diverse linguistic contexts? To address this gap, we present XToM, a rigorously validated multilingual benchmark that evaluates ToM across five languages and incorporates diverse, contextually rich task scenarios. Using XToM, we systematically evaluate LLMs (e.g., DeepSeek R1), revealing a pronounced dissonance: while models excel in multilingual language understanding, their ToM performance varies across languages. Our findings expose limitations in LLMs' ability to replicate human-like mentalizing across linguistic contexts.
Chunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou, Xinyuan Cheng, Zhifan Sun, Zheye Deng, Kawai Chung, Yuzhuo Ao, Yixiang Fan, Cheng Jiayang, Ercong Nie, Ginny Y. Wong, Helmut Schmid, Hinrich Schütze, Simon See, Yangqiu Song
ACL (1)13
2025 LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
abstract
Tianshi Zheng, Cheng Jiayang, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tianshi Zheng, Cheng Jiayang, Zihao Wang 0001, Jiaxin Bai, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP8
2025 MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
abstract
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
Zhaowei Wang 0003, Wenhao Yu 0002, Xiyu Ren, Yu Zhao 0043, Rohit Saxena, Ginny Y. Wong, Simon See, Pasquale Minervini, Yangqiu Song, Mark Steedman
NeurIPS8
2024 AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
abstract
Zhaowei Wang, Wei Fan, Qing Zong, Hongming Zhang, Sehyun Choi, Tianqing Fang, Xin Liu, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaowei Wang 0003, Wei Fan 0001, Qing Zong, Hongming Zhang 0009, Sehyun Choi, Tianqing Fang, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See
ACL (1)9
2024 Audience Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
abstract
Debate is the process of exchanging viewpoints or convincing others on a particular issue. Recent research has provided empirical evidence that the persuasiveness of an argument is determined not only by language usage but also by communicator characteristics. Researchers have paid much attention to aspects of languages, such as linguistic features and discourse structures, but combining argument persuasiveness and impact with the social personae of the audience has not been explored due to the difficulty and complexity. We have observed the impressive simulation and personification capability of ChatGPT, indicating a giant pre-trained language model may function as an individual to provide personae and exert unique influences based on diverse background knowledge. Therefore, we propose a persona knowledge-aligned framework for argument quality assessment tasks from the audience side. This is the first work that leverages the emergence of ChatGPT and injects such audience personae knowledge into smaller language models via prompt tuning. The performance of our pipeline demonstrates significant and consistent improvement compared to competitive architectures.
Chunkit Chan, Cheng Jiayang, Xin Liu 0039, Yauwai Yim, Zheye Deng, Haoran Li 0003, Yangqiu Song, Ginny Y. Wong, Simon See
ECAI9
2023 COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference Perspective
abstract
Zhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhaowei Wang 0003, Quyet V. Do, Hongming Zhang 0009, Jiayao Zhang 0001, Weiqi Wang 0001, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See
ACL (1)8
2023 Logical Message Passing Networks with One-hop Inference on Atomic Formulas
Zihao Wang 0001, Yangqiu Song, Ginny Y. Wong, Simon See
ICLR3
2023 Self-Consistent Narrative Prompts on Abductive Natural Language Inference
abstract
Chunkit Chan, Xin Liu, Tsz Ho Chan, Jiayang Cheng, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chunkit Chan, Xin Liu 0039, Tsz Ho Chan, Cheng Jiayang, Yangqiu Song, Ginny Y. Wong, Simon See
IJCNLP (1)6
2022 SubeventWriter: Iterative Sub-event Sequence Generation with Coherence Controller
abstract
In this paper, we propose a new task of subevent generation for an unseen process to evaluate the understanding of the coherence of subevent actions and objects.To solve the problem, we design SubeventWriter, a sub-event sequence generation framework with a coherence controller.Given an unseen process, the framework can iteratively construct the subevent sequence by generating one sub-event at each iteration.We also design a very effective coherence controller to decode more coherent sub-events.As our extensive experiments and analysis indicate, SubeventWriter 1 can generate more reliable and meaningful sub-event sequences for unseen processes.
Zhaowei Wang 0003, Hongming Zhang 0009, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP5
2022 Complex Hyperbolic Knowledge Graph Embeddings with Fast Fourier Transform
abstract
The choice of geometric space for knowledge graph (KG) embeddings can have significant effects on the performance of KG completion tasks.The hyperbolic geometry has been shown to capture the hierarchical patterns due to its tree-like metrics, which addressed the limitations of the Euclidean embedding models.Recent explorations of the complex hyperbolic geometry further improved the hyperbolic embeddings for capturing a variety of hierarchical structures.However, the performance of the hyperbolic KG embedding models for nontransitive relations is still unpromising, while the complex hyperbolic embeddings do not deal with multi-relations.This paper aims to utilize the representation capacity of the complex hyperbolic geometry in multi-relational KG embeddings.To apply the geometric transformations which account for different relations and the attention mechanism in the complex hyperbolic space, we propose to use the fast Fourier transform (FFT) as the conversion between the real and complex hyperbolic space.Constructing the attention-based transformations in the complex space is very challenging, while the proposed Fourier transform-based complex hyperbolic approaches provide a simple and effective solution.Experimental results show that our methods outperform the baselines, including the Euclidean and the real hyperbolic embedding models.
Huiru Xiao, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See
EMNLP4
2018 A hybrid evolutionary preprocessing method for imbalanced datasets
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
Inf. Sci.1
2016 Identification of protein-ligand binding site using multi-clustering and Support Vector Machine
abstract
Multi-clustering has been widely used. It acts as a pre-training process for identifying protein-ligand binding in structure-based drug design. Then, the Support Vector Machine (SVM) is employed to classify the sites most likely for binding ligands. Three types of attributes are used, namely geometry-based, energy-based, and sequence conservation. Comparison is made on 198 drug-target protein complexes with LIGSITECSC, SURFNET, Fpocket, Q-SiteFinder, ConCavity, and MetaPocket. The results show an improved success rate of up to 86%.
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
IECON1
2014 An under-sampling method based on fuzzy logic for large imbalanced dataset
abstract
Large imbalanced datasets have introduced difficulties to classification problems. They cause a high error rate of the minority class samples and a long training time of the classification model. Therefore, re-sampling and data size reduction have become important steps to pre-process the data. In this paper, a sampling strategy over a large imbalanced dataset is proposed, in which the samples of the larger class are selected based on fuzzy logic. To further reduce the data size, the evolutionary computational method of CHC is employed. The evaluation is done by applying a Support Vector Machine (SVM) to train a classification model from the re-sampled training sets. From experimental results, it can be seen that our proposed method improves both the F-measure and AUC. The complexity of the classification model is also compared. It is found that our proposed method is superior to all other compared methods.
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
FUZZ-IEEE1
2013 A novel evolutionary preprocessing method based on over-sampling and under-sampling for imbalanced datasets
abstract
Imbalanced datasets are commonly encountered in real-world classification problems. However, many machine learning algorithms are originally designed for well-balanced datasets. Re-sampling has become an important step to preprocess imbalanced dataset. It aims at balancing the datasets by increasing the sample size of the smaller class or decreasing the sample size of the larger class, which are known as over-sampling and under-sampling respectively. In this paper, a novel sampling strategy based on both over-sampling and under-sampling is proposed, in which the new samples of the smaller class are created by the Synthetic Minority Over-sampling Technique (SMOTE). The improvement of the datasets is done by the evolutionary computational method of CHC that works on both the minority class and majority class samples. The result is a hybrid data preprocessing method that combines both over-sampling and under-sampling techniques to re-sample datasets. The evaluation is done by applying the learning algorithm C4.5 to obtain a classification model from the re-sampled datasets. Experimental results reported that the proposed approach can decrease the over-sampling rate about 50% with only around 3% discrepancy on the accuracy.
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
IECON1
2013 Predicting Protein-Ligand Binding Site Using Support Vector Machine with Protein Properties
abstract
Identification of protein-ligand binding site is an important task in structure-based drug design and docking algorithms. In the past two decades, different approaches have been developed to predict the binding site, such as the geometric, energetic, and sequence-based methods. When scores are calculated from these methods, the algorithm for doing classification becomes very important and can affect the prediction results greatly. In this paper, the support vector machine (SVM) is used to cluster the pockets that are most likely to bind ligands with the attributes of geometric characteristics, interaction potential, offset from protein, conservation score, and properties surrounding the pockets. Our approach is compared to LIGSITE, LIGSITE(CSC), SURFNET, Fpocket, PocketFinder, Q-SiteFinder, ConCavity, and MetaPocket on the data set LigASite and 198 drug-target protein complexes. The results show that our approach improves the success rate from 60 to 80 percent at AUC measure and from 61 to 66 percent at top 1 prediction. Our method also provides more comprehensive results than the others.
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
IEEE ACM Trans. Comput. Biol. Bioinform.1
2012 Predicting protein-ligand binding site with differential evolution and support vector machine
abstract
Identification of protein-ligand binding site is an important task in structure-based drug design and docking algorithms. In these two decades, many different approaches have been developed to predict the binding site, such as geometric, energetic and sequence-based methods. When the scores are calculated from these methods, the method of classification is very important and can affect the prediction results greatly. A developed support vector machine (SVM) is used to classify the pockets, which are most likely to bind ligands with the attributes of grid value, interaction potential, offset from protein, conservation score and the information around the pockets. Since SVM is sensitive to the input parameters and the positive samples are more relevant than negative samples, differential evolution (DE) is applied to find out the suitable parameters for SVM. We compare our algorithm to four other approaches: LIGSITE, SURFNET, PocketFinder and Concavity. Our algorithm is found to provide the highest success rate.
Ginny Y. Wong, Frank H. F. Leung, Sai-Ho Ling
IJCNN1
2010 Predicting protein-ligand binding site with support vector machine
abstract
Identification of protein-ligand binding site is an important task in structure-based drug design and docking algorithms. In these two decades, many different approaches have been developed to predict the binding site, such as geometric, energetic and sequence-based methods. We present the binding site prediction algorithm that takes advantage of both sequence conservation and geometric methods for pocket finding (LIGSITE and SURFNET). SVM is used to cluster the pockets, which are most likely to bind ligands with the attributes of grid value, interaction potential and offset from protein. We compare our algorithm to four other approaches: LIGSITE, SURFNET, PocketFinder and Concavity. Our algorithm is found to provide the highest success rate.
Ginny Y. Wong, Frank H. F. Leung
IEEE Congress on Evolutionary Computation1
2009 Relaxed stability conditions for discrete-time fuzzy-model-based control systems
abstract
This paper presents the stability conditions for discrete-time fuzzy-model-based control systems subject to parameter uncertainties. The nonlinear plant subject to parameter uncertainties is represented by a Takagi-Sugeno fuzzy model with uncertain grades of memberships. Relaxed stability conditions for this class of fuzzy control systems will be derived to guarantee the system stability. A numerical example will be presented to show the effectiveness of the proposed approach.
Ginny Y. Wong, Frank H. F. Leung, Hak-Keung Lam
FUZZ-IEEE1