EDBT 2026 Demo / reviewers in the wild / expert
Billy Chiu
dblp:144/1140
· DBLP profile ↗
16ranked-venue papers
3as first author
10since 2021 · last 2027
0000-0001-6683-3249ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Localize-then-summarize: Enhancing scientific multimodal summarization with facet-aware cross-modal memory
Zusheng Tan, Jing-Yu Ji, Ngai Fung Ng, Jeff K. T. Tang, Ken Fong, Jing Li 0034, Sam Kwong, Billy Chiu |
Inf. Process. Manag. | 10 |
| 2026 | MFG-SciSum: A multimodal faceted graph framework for scientific summarization
Zusheng Tan, Jing Li 0034, Shen Gao, Wai Lam, Sam Kwong, Billy Chiu |
Inf. Sci. | 8 |
| 2025 | Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective OptimizationabstractThe rapid development of LLMs has raised concerns about their potential misuse, leading to various watermarking schemes that typically offer high detectability.
However, existing watermarking techniques often face trade-off between watermark detectability and generated text quality.
In this paper, we introduce Learning to Watermark (LTW), a novel selective watermarking framework that leverages multi-objective optimization to effectively balance these competing goals.
LTW features a lightweight network that adaptively decides when to apply the watermark by analyzing sentence embeddings, token entropy, and current watermarking ratio.
Training of the network involves two specifically constructed loss functions that guide the model toward Pareto-optimal solutions, thereby harmonizing watermark detectability and text quality.
By integrating LTW with two baseline watermarking methods, our experimental evaluations demonstrate that LTW significantly enhances text quality without compromising detectability.
Our selective watermarking approach offers a new perspective for designing watermarks for LLMs and a way to preserve high text quality for watermarks. The code is publicly available at: https://github.com/fattyray/learning-to-watermark Chenrui Wang, Junyi Shu, Billy Chiu, Yu Li 0007, Saleh Alharbi, Min Zhang 0005, Jing Li 0034 |
NeurIPS | 3 |
| 2025 | SMSMO: Learning to generate multimodal summary for scientific papers
Xinyi Zhong, Zusheng Tan, Shen Gao, Jing Li 0034, Jiaxing Shen, Jing-Yu Ji, Jeff K. T. Tang, Billy Chiu |
Knowl. Based Syst. | 8 |
| 2025 | Scientific poster generation: A new dataset and approach
Xinyi Zhong, Zusheng Tan, Jing Li 0034, Shen Gao, Jing Ma 0004, Shanshan Feng 0001, Billy Chiu |
Pattern Recognit. | 7 |
| 2024 | An Improved Gradient-Based Repair Method for Constrained Numerical OptimizationabstractRecently, gradient-based repair methods have been commonly introduced into constraint-handling techniques to handle linear and non-linear constraints. These gradient-based repair methods are only designed to reduce constraint violations, and they do not act on the objective function in the repairing process. Nevertheless, the gradient descent optimization method is originally proposed to optimize the objective function without constraints. Motivated by this consideration, this study develops an improved gradient-based repair method that incorporates the objective function to handle the constraints and optimize the objective function simultaneously. The proposed repair method is integrated into a multiobjective differential evolution framework to investigate its effectiveness. Experiments have been conducted on 57 real-world constrained benchmark test functions. The empirical result shows that, compared to the selected state-of-the-art algorithms, our proposed gradient-based repair method can assist the adopted constrained optimization approach to obtain high-quality feasible solutions. Jing-Yu Ji, Kwan-Yeung Lee, Billy Chiu, Man Leung Wong, Sam Kwong |
CEC | 4 |
| 2024 | Few-Shot Relation Extraction With Dual Graph Neural Network InteractionabstractRecent advances in relation extraction with deep neural architectures have achieved excellent performance. However, current models still suffer from two main drawbacks: 1) they require enormous volumes of training data to avoid model overfitting and 2) there is a sharp decrease in performance when the data distribution during training and testing shift from one domain to the other. It is thus vital to reduce the data requirement in training and explicitly model the distribution difference when transferring knowledge from one domain to another. In this work, we concentrate on few-shot relation extraction under domain adaptation settings. Specifically, we propose DUAL GRAPH, a novel graph neural network (GNN) based approach for few-shot relation extraction. DUAL GRAPH leverages an edge-labeling dual graph (i.e., an instance graph and a distribution graph) to explicitly model the intraclass similarity and interclass dissimilarity in each individual graph, as well as the instance-level and distribution-level relations across graphs. A dual graph interaction mechanism is proposed to adequately fuse the information between the two graphs in a cyclic flow manner. We extensively evaluate DUAL GRAPH on FewRel1.0 and FewRel2.0 benchmarks under four few-shot configurations. The experimental results demonstrate that DUAL GRAPH can match or outperform previously published approaches. We also perform experiments to further investigate the parameter settings and architectural choices, and we offer a qualitative analysis. Jing Li 0034, Shanshan Feng 0001, Billy Chiu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Few-Shot Named Entity Recognition via Meta-Learning (Extended Abstract)abstractNamed entity recognition (NER) is typically framed as a sequence labeling problem where the entity classes are inherently entangled together because the entity number and classes in a sentence are not known in advance, leaving the N-way K-shot NER problem so far unexplored. In our TKDE paper, we first formally define a more suitable N-way K-shot setting for NER. Then we propose FewNER, a novel meta-learning approach for few-shot NER. FewNER separates the entire network into a task-independent part and a task-specific part. During training in FewNER, the task-independent part is meta-learned across multiple tasks and the task-specific part is learned for each individual task in a low-dimensional space. At test time, FewNER keeps the task-independent part fixed and adapts to a new task via gradient descent by updating only the task-specific part, resulting in it being less prone to overfitting and more computationally efficient. Compared with pre-trained language models (e.g., BERT and ELMo) which obtain the transferability in an implicit manner (i.e., relying on large-scale corpora), FewNER explicitly optimizes the capability of "learning to adapt quickly" through meta-learning. The results demonstrate that FewNER achieves state-of-the-art performance against nine baseline methods by significant margins on three adaptation experiments (i.e., intra-domain cross-type, cross-domain intra-type and cross-domain cross-type). Jing Li 0034, Billy Chiu, Shanshan Feng 0001, Hao Wang 0013 |
ICDE | 2 |
| 2022 | Few-Shot Named Entity Recognition via Meta-LearningabstractFew-shot learning under the$N$-way$K$-shot setting (i.e.,$K$annotated samples for each of$N$classes) has been widely studied in relation extraction (e.g., FewRel) and image classification (e.g., Mini-ImageNet). Named entity recognition (NER) is typically framed as a sequence labeling problem where the entity classes are inherently entangled together because the entity number and classes in a sentence are not known in advance, leaving the$N$-way$K$-shot NER problem so far unexplored. In this paper, we first formally define a more suitable$N$-way$K$-shot setting for NER. Then we proposeFewNER, a novel meta-learning approach for few-shot NER.FewNERseparates the entire network into a task-independent part and a task-specific part. During training inFewNER, the task-independent part is meta-learned across multiple tasks and the task-specific part is learned for each individual task in a low-dimensional space. At test time,FewNERkeeps the task-independent part fixed and adapts to a new task via gradient descent by updating only the task-specific part, resulting in it being less prone to overfitting and more computationally efficient. Compared with pre-trained language models (e.g., BERT and ELMo) which obtain the transferability in an implicit manner (i.e., relying on large-scale corpora),FewNERexplicitly optimizes the capability of “learning to adapt quickly” through meta-learning. The results demonstrate thatFewNERachieves state-of-the-art performance against nine baseline methods by significant margins on three adaptation experiments (i.e., intra-domain cross-type, cross-domain intra-type and cross-domain cross-type). Jing Li 0034, Billy Chiu, Shanshan Feng 0001, Hao Wang 0013 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Neural Text Segmentation and its Application to Sentiment AnalysisabstractText segmentation is a fundamental task in natural language processing. Depending on the levels of granularity, the task can be defined as segmenting a document into topical segments, or segmenting a sentence into elementary discourse units (EDUs). Traditional solutions to the two tasks heavily rely on carefully designed features. The recently proposed neural models do not need manual feature engineering, but they either suffer from sparse boundary tags or cannot efficiently handle the issue of variable size output vocabulary. In light of such limitations, we propose a generic end-to-end segmentation model, namely${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$, which first uses a bidirectional recurrent neural network to encode an input text sequence.${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$then uses another recurrent neural networks, together with a pointer network, to select text boundaries in the input sequence. In this way,${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$does not require any hand-crafted features. More importantly,${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$inherently handles the issue of variable size output vocabulary and the issue of sparse boundary tags. In our experiments,${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$outperforms state-of-the-art models on two tasks: document-level topic segmentation and sentence-level EDU segmentation. As a downstream application, we further propose a hierarchical attention model for sentence-level sentiment analysis based on the outcomes of${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$. The hierarchical model can make full use of both word-level and EDU-level information simultaneously for sentence-level sentiment analysis. In particular, it can effectively exploit EDU-level information, such as the inner properties of EDUs, which cannot be fully encoded in word-level features. Experimental results show that our hierarchical model achieves new state-of-the-art results on the Movie Review and Stanford Sentiment Treebank benchmarks. Jing Li 0034, Billy Chiu, Shuo Shang, Ling Shao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Autoencoding Keyword Correlation Graph for Document ClusteringabstractDocument clustering requires a deep understanding of the complex structure of longtext; in particular, the intra-sentential (local) and inter-sentential features (global).Existing representation learning models do not fully capture these features.To address this, we present a novel graph-based representation for document clustering that builds a graph autoencoder (GAE) on a Keyword Correlation Graph.The graph is constructed with topical keywords as nodes and multiple local and global features as edges.A GAE is employed to aggregate the two sets of features by learning a latent representation which can jointly reconstruct them.Clustering is then performed on the learned representations, using vector dimensions as features for inducing document classes.Extensive experiments on two datasets show that the features learned by our approach can achieve better clustering performance than other existing features, including term frequency-inverse document frequency and average embedding. Billy Chiu, Sunil Kumar Sahu, Derek Thomas, Neha Sengupta, Mohammady Mahdy |
ACL | 1 |
| 2020 | Relation Extraction with Self-determined Graph Convolutional NetworkabstractRelation Extraction is a way of obtaining the semantic relationship between entities in text. The state-of-the-art methods use linguistic tools to build a graph for the text in which the entities appear and then a Graph Convolutional Network (GCN) is employed to encode the pre-built graphs. Although their performance is promising, the reliance on linguistic tools results in a non end-to-end process. In this work, we propose a novel model, the Self-determined Graph Convolutional Network (SGCN), which determines a weighted graph using a self-attention mechanism, rather using any linguistic tool. Then, the self-determined graph is encoded using a GCN. We test our model on the TACRED dataset and achieve the state-of-the-art result. Our experiments show that SGCN outperforms the traditional GCN, which uses dependency parsing tools to build the graph. Sunil Kumar Sahu, Derek Thomas, Billy Chiu, Neha Sengupta, Mohammady Mahdy |
CIKM | 3 |
| 2020 | Attending to Inter-sentential Features in Neural Text ClassificationabstractText classification requires a deep understanding of the linguistic features in text; in particular, the intra-sentential (local) and inter-sentential features (global). Models that operate on word sequences have been successfully used to capture the local features, yet they are not effective in capturing the global features in long-text. We investigate graph-level extensions to such models and propose a novel architecture for combining alternative text features. It uses an attention mechanism to dynamically decide how much information to use from a sequence- or graph-level component. We evaluated different architectures on a range of text classification datasets, and graph-level extensions were found to improve performance on most benchmarks. In addition, the attention-based architecture, as adaptively-learned from the data, outperforms the generic and fixed-value concatenation ones. Billy Chiu, Sunil Kumar Sahu, Neha Sengupta, Derek Thomas, Mohammady Mahdy |
SIGIR | 1 |
| 2018 | Bio-SimVerb and Bio-SimLex: wide-coverage evaluation sets of word similarity in biomedicineabstractBACKGROUND: Word representations support a variety of Natural Language Processing (NLP) tasks. The quality of these representations is typically assessed by comparing the distances in the induced vector spaces against human similarity judgements. Whereas comprehensive evaluation resources have recently been developed for the general domain, similar resources for biomedicine currently suffer from the lack of coverage, both in terms of word types included and with respect to the semantic distinctions. Notably, verbs have been excluded, although they are essential for the interpretation of biomedical language. Further, current resources do not discern between semantic similarity and semantic relatedness, although this has been proven as an important predictor of the usefulness of word representations and their performance in downstream applications. RESULTS: We present two novel comprehensive resources targeting the evaluation of word representations in biomedicine. These resources, Bio-SimVerb and Bio-SimLex, address the previously mentioned problems, and can be used for evaluations of verb and noun representations respectively. In our experiments, we have computed the Pearson's correlation between performances on intrinsic and extrinsic tasks using twelve popular state-of-the-art representation models (e.g. word2vec models). The intrinsic-extrinsic correlations using our datasets are notably higher than with previous intrinsic evaluation benchmarks such as UMNSRS and MayoSRS. In addition, when evaluating representation models for their abilities to capture verb and noun semantics individually, we show a considerable variation between performances across all models. CONCLUSION: Bio-SimVerb and Bio-SimLex enable intrinsic evaluation of word representations. This evaluation can serve as a predictor of performance on various downstream tasks in the biomedical domain. The results on Bio-SimVerb and Bio-SimLex using standard word representation models highlight the importance of developing dedicated evaluation resources for NLP in biomedicine for particular word classes (e.g. verbs). These are needed to identify the most accurate methods for learning class-specific representations. Bio-SimVerb and Bio-SimLex are publicly available. Billy Chiu, Sampo Pyysalo, Ivan Vulic, Anna Korhonen |
BMC Bioinform. | 1 |
| 2017 | A neural network multi-task learning approach to biomedical named entity recognitionabstractBACKGROUND: Named Entity Recognition (NER) is a key task in biomedical text mining. Accurate NER systems require task-specific, manually-annotated datasets, which are expensive to develop and thus limited in size. Since such datasets contain related but different information, an interesting question is whether it might be possible to use them together to improve NER performance. To investigate this, we develop supervised, multi-task, convolutional neural network models and apply them to a large number of varied existing biomedical named entity datasets. Additionally, we investigated the effect of dataset size on performance in both single- and multi-task settings. RESULTS: We present a single-task model for NER, a Multi-output multi-task model and a Dependent multi-task model. We apply the three models to 15 biomedical datasets containing multiple named entities including Anatomy, Chemical, Disease, Gene/Protein and Species. Each dataset represent a task. The results from the single-task model and the multi-task models are then compared for evidence of benefits from Multi-task Learning. With the Multi-output multi-task model we observed an average F-score improvement of 0.8% when compared to the single-task model from an average baseline of 78.4%. Although there was a significant drop in performance on one dataset, performance improves significantly for five datasets by up to 6.3%. For the Dependent multi-task model we observed an average improvement of 0.4% when compared to the single-task model. There were no significant drops in performance on any dataset, and performance improves significantly for six datasets by up to 1.1%. The dataset size experiments found that as dataset size decreased, the multi-output model's performance increased compared to the single-task model's. Using 50, 25 and 10% of the training data resulted in an average drop of approximately 3.4, 8 and 16.7% respectively for the single-task model but approximately 0.2, 3.0 and 9.8% for the multi-task model. CONCLUSIONS: Our results show that, on average, the multi-task models produced better NER results than the single-task models trained on a single NER dataset. We also found that Multi-task Learning is beneficial for small datasets. Across the various settings the improvements are significant, demonstrating the benefit of Multi-task Learning for this task. Gamal K. O. Crichton, Sampo Pyysalo, Billy Chiu, Anna Korhonen |
BMC Bioinform. | 3 |
| 2016 | Syllable based DNN-HMM Cantonese Speech to Text System
Timothy Wong, Claire Li, Sam Lam, Billy Chiu, Qin Lu 0001, Minglei Li 0001, Dan Xiong, Roy Shing Yu, Vincent T. Y. Ng |
LREC | 4 |