Annie En-Shiun Lee

dblp:10/8510 · also En-Shiun Annie Lee · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-4592-3522ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 first-authorHuman-computer interaction and ubiquitous computing · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
abstract
Despite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic resources. In this paper, we introduce \dsname, a novel fine-grained dataset that builds on existing parallel corpora to provide error span, error type, and error severity annotations in machine-translated examples from English to Mandarin, Cantonese, and Wu Chinese, along with a Mandarin-Hokkien component derived from a non-parallel source. Our dataset serves as a resource for the MT community to fine-tune models with error detection capabilities, supporting research on translation quality estimation, error-aware generation, and low-resource language evaluation. We also establish baseline results using language models to benchmark translation error detection performance. Specifically, we evaluate multiple open source and closed source LLMs using span-level and correlation-based MQM metrics, revealing their limited precision, underscoring the need for our dataset. Finally, we report our rigorous annotation process by native speakers, with analyses on pilot studies, iterative feedback, insights, and patterns in error type and severity.
Hannah Liu, Junghyun Min, Annie En-Shiun Lee, Ethan Yue Heng Cheung, Shou-Yi Hung, Elsie Chan, Shiyao Qian, Runtong Liang, Kimlan Huynh, Wing Yu Yip, York Hay Ng, Tsz Fung Yau, Ka Ieng Charlotte Lo, You-Wei Wu, Richard Tzong-Han Tsai
LREC3
2026 OasisSimp: An Open-source Asian-English Sentence Simplification Dataset
Hannah Liu, Murphy Tian, Iqra Ali, Haonan Gao, Qiaoyiwen Wu, Blair Yang, Uthayasanker Thayasivam, Annie En-Shiun Lee, Pakawat Nakwijit, Surangika Ranathunga, Ravi Shekhar
LREC8
2026 Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+
abstract
The URIEL+ linguistic knowledge base supports multilingual research by encoding languages through geographic, genetic, and typological vectors. However, data sparsity (e.g. missing feature types, incomplete language entries, and limited genealogical coverage) remains prevalent. This limits the usefulness of URIEL+ in cross-lingual transfer, particularly for supporting low-resource languages. To address this sparsity, we extend URIEL+ by introducing script vectors to represent writing system properties for 7,488 languages, integrating Glottolog to add 18,710 additional languages, and expanding lineage imputation for 26,449 languages by propagating typological and script features across genealogies. These improvements reduce feature sparsity by 14% for script vectors, increase language coverage by up to 19,015 languages (1,007%), and boost imputation quality metrics by up to 35%. Our benchmark on cross-lingual transfer tasks (oriented around low-resource languages) shows occasionally divergent performance compared to URIEL+, with performance gains up to 6% in certain setups.
Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang, A. Seza Dogruöz, Annie En-Shiun Lee
LREC7
2026 Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-Resource Language Translation
abstract
Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language’s representation in the model are limited. This restricts the capabilities of domain-specific NMT systems for low-resource languages (LRLs). As a solution, parallel data from auxiliary domains can be used either to fine-tune or to further pre-train the msLM. We present an evaluation of the effectiveness of these two techniques in the context of domain-specific LRL-NMT. We also explore the impact of domain divergence on NMT model performance. We recommend several strategies for utilizing auxiliary parallel data in building domain-specific NMT models for LRLs.
Surangika Ranathunga, Shravan Nayak, Annie En-Shiun Lee, Shih-Ting Cindy Huang, Yuchen Zeng 0001, Yanke Mao, Yun-Hsiang Ray Chan, Songchen Yuan, Anthony Rinaldi
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2025 INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African Languages
abstract
Slot-filling and intent detection are well-established tasks in Conversational AI. However, current large-scale benchmarks for these tasks often exclude evaluations of low-resource languages and rely on translations from English benchmarks, thereby predominantly reflecting Western-centric concepts. In this paper, we introduce “INJONGO” - a multicultural, open-source benchmark dataset for 16 African languages with utterances generated by native speakers across diverse domains, including banking, travel, home, and dining. Through extensive experiments, we benchmark fine-tuning multilingual transformer models and prompting large language models (LLMs), and show the advantage of leveraging African-cultural utterances over Western-centric utterances for improving cross-lingual transfer from the English language. Experimental results reveal that current LLMs struggle with the slot-filling task, with GPT-4o achieving an average performance of 26 F1. In contrast, intent detection performance is notably better, with an average accuracy of 70.6%, though it still falls short of fine-tuning baselines. When compared to the English language, GPT-4o and fine-tuning baselines perform similarly on intent detection, achieving an accuracy of approximately 81%. Our findings suggest that LLMs performance is still behind for many low-resource African languages, and more work is needed to further improve their downstream performance.
Jesujoba O. Alabi, Andiswa Bukula, Jian Yun Zhuang, Annie En-Shiun Lee, Tadesse Kebede Guge, Israel Abebe Azime, Happy Buzaaba, Blessing K. Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Shamsuddeen Hassan Muhammad, Salomey Osei, Sokhar Samb, Dietrich Klakow, David Ifeoluwa Adelani
ACL (1)5
2025 URIEL+: Enhancing Linguistic Inclusion and Usability in a Typological and Multilingual Knowledge Base
abstract
URIEL is a knowledge base offering geographical, phylogenetic, and typological vector representations for 7970 languages. It includes distance measures between these vectors for 4005 languages, which are accessible via the lang2vec tool. Despite being frequently cited, URIEL is limited in terms of linguistic inclusion and overall usability. To tackle these challenges, we introduce URIEL+, an enhanced version of URIEL and lang2vec that addresses these limitations. In addition to expanding typological feature coverage for 2898 languages, URIEL+ improves the user experience with robust, customizable distance calculations to better suit the needs of users. These upgrades also offer competitive performance on downstream tasks and provide distances that better align with linguistic distance studies.
Aditya Armaan Khan, Mason Shipton, David Anugraha, Kaiyao Duan, Phuong Hanh Hoang, Eric Khiu, A. Seza Dogruöz, Annie En-Shiun Lee
COLING8
2025 Less is More: The Effectiveness of Compact Typological Language Representations
abstract
Linguistic feature datasets such as URIEL+ are valuable for modelling cross-lingual relationships, but their high dimensionality and sparsity, especially for low-resource languages, limit the effectiveness of distance metrics.We propose a pipeline to optimize the URIEL+ typological feature space by combining feature selection and imputation, producing compact yet interpretable typological representations.We evaluate these feature subsets on linguistic distance alignment and downstream tasks, demonstrating that reduced-size representations of language typology can yield more informative distance metrics and improve performance in multilingual NLP applications.
York Hay Ng, Phuong Hanh Hoang, Annie En-Shiun Lee
EMNLP3
2025 IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
abstract
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba Oluwadara Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing Kudzaishe Sibanda, Godson Koffi Kalipe, Jonathan Mukiibi, Salomon Kabongo Kabenamualu, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Salomey Osei, Shamsuddeen Hassan Muhammad, Sokhar Samb, Tadesse Kebede Guge, Tombekai Vangoni Sherman, Pontus Stenetorp. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O. Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, Annie En-Shiun Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing K. Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuya, Salomey Osei, Shamsuddeen Hassan Muhammad, Sokhar Samb, Tadesse Kebede Guge, Tombekai Vangoni Sherman, Pontus Stenetorp
NAACL (Long Papers)10
2025 WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
abstract
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Stephanie Yulia Salim, Yi Zhou 0019, Yinxuan Gui, David Ifeoluwa Adelani, Annie En-Shiun Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Wijaya, Alice Oh, Chong-Wah Ngo
NAACL (Long Papers)44
2025 Creating a Joint-Faculty Artificial Intelligence Concentration within a Graduate Program
abstract
The global demand for Artificial Intelligence (AI) talent is growing at a rapid pace, leading to significant AI skills shortages. This experience report describes the creation of a new and uniquely joint-faculty AI concentration within our existing Master's program, characterized by its applied research internship in industry. We describe the experience of creating a multi-disciplinary AI program through broad consultation across two of the largest faculties in our institution. The new concentration has been well-received by the collaborating faculties and with industry partners; its popularity with applicants has resulted in admitting exceptional candidates from around the world and becoming the most popular concentration in the program. We offer this experience report in the hope that it may serve as a model for other practitioners who are considering navigating the creation of a joint-faculty concentration within a graduate program, especially in the popular field of AI.
Annie En-Shiun Lee, Arvind Gupta, Amane Takeuchi, Stacey A. Koornneef
SIGCSE (2)1
2025 Crafting for Career Agility: An Outcome-Based Redesign of a Machine Learning Curriculum within a Program Bundle
abstract
In this paper, we focus on the certificate redesign of our machine learning program within the context of a three-program bundle; we used an outcome-based approach with the student's targeted career outcome in mind. The purpose of this curriculum redesign is threefold: (1) To align the course content with the industry's rapidly changing needs and demands of the job market; (2) To identify the proper sequence of laddering structure for students taking multiple programs; and (3) To create a natural and streamline flow of the learning experience (i.e., removing overlapping content). We show that the resulting redesign is a holistic set of certificate programs with industry-relevant content tailored for skills-based experiential learning.
Annie En-Shiun Lee, Sean Woodhead, Karthik Kuber, Hashmat Rohian, Stacey A. Koornneef
SIGCSE (2)1
2024 Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing Systems
abstract
Machine translation focuses mainly on high-resource languages (HRLs), while low-resource languages (LRLs) like Taiwanese Hokkien are relatively under-explored. The study aims to address this gap by developing a dual translation model between Taiwanese Hokkien and both Traditional Mandarin Chinese and English. We employ a pre-trained LLaMA 2-7B model specialized in Traditional Mandarin Chinese to leverage the orthographic similarities between Taiwanese Hokkien Han and Traditional Mandarin Chinese. Our comprehensive experiments involve translation tasks across various writing systems of Taiwanese Hokkien as well as between Taiwanese Hokkien and other HRLs. We find that the use of a limited monolingual corpus still further improves the model’s Taiwanese Hokkien capabilities. We then utilize our translation model to standardize all Taiwanese Hokkien writing systems into Hokkien Han, resulting in further performance improvements. Additionally, we introduce an evaluation method incorporating back-translation and GPT-4 to ensure reliable translation quality assessment even for LRLs. The study contributes to narrowing the resource gap for Taiwanese Hokkien and empirically investigates the advantages and limitations of pre-training and fine-tuning based on LLaMA 2.
Bo-Han Lu, Annie En-Shiun Lee, Richard Tzong-Han Tsai
LREC/COLING3
2024 SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects
abstract
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, En-Shiun Annie Lee. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen 0001, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, Annie En-Shiun Lee
EACL (1)8
2024 Exploring Student Motivation in Integration of Soft Skills Training within Three Levels of Computer Science Programs
abstract
In computer science education, cultivating soft skills alongside technical competencies is increasingly recognized as crucial for successful careers in industry and research. However, integrating soft skills training into curricula often remains a secondary consideration, separate from the primary program delivery or managed by a separate unit within the university. In this paper, we present three comprehensive curricula that intertwine soft skills within academic training for three distinct tiers of computer science education: undergraduate, master's, and professional levels. We identify common intrinsic motivations that we found most impactful for learner success within these programs, and present a detailed exploration of each program's curriculum and goals. By highlighting effective strategies and potential pitfalls, we offer valuable insights into harnessing these motivational drivers to enhance student engagement and learning. Furthermore, we outline emerging opportunities and challenges within the integrated curricula, inviting discussion on the broader implications for computer science education.
Annie En-Shiun Lee, Luki Danukarjanto, Sadia Sharmin, Shou-Yi Hung, Sicong Huang 0001
SIGCSE (1)1
2021 Pillars of Program Design and Delivery: A Case Study using Self-Directed, Problem-Based, and Supportive Learning
abstract
As machine learning (ML) becomes prevalent in industries and businesses, the need to use these algorithms to solve real-world problems grows rapidly. However, there is a serious deficit of qualified talent in this field and thus a corresponding shortage of educational programs. To address the shortage of ML specialists in the field, universities are offering continuing education programs that fast-track the development of technical and transverse skills needed for success in the field. The award-winning machine learning program described in this paper is carefully designed with industry and community partners while focusing on practical skills and participation in the local industry network. This program can be summarized in three learning principles: 1) learners are encouraged to build their knowledge and skills in a self-directed manner; 2) group projects in both simulated and workplace settings are incorporated to support problem-based learning; and 3) supportive learning environment is established to encourage open and safe learning. This paper reports on the instructors' experiences in teaching the four courses based on these principles, which has resulted in high satisfaction from students, successfully placing students in industry, and winning a national award. We offer this experiential report in the hope that it may serve as a point of reference for other instructors and programs for mature technical learners in machine learning.
Annie En-Shiun Lee, Karthik Kuber, Hashmat Rohian, Sean Woodhead
SIGCSE1
2017 Discovering Protein-DNA Binding Cores by Aligned Pattern Clustering
abstract
Understanding binding cores is of fundamental importance in deciphering Protein-DNA (TF-TFBS) binding and gene regulation. Limited by expensive experiments, it is promising to discover them with variations directly from sequence data. Although existing computational methods have produced satisfactory results, they are one-to-one mappings with no site-specific information on residue/nucleotide variations, where these variations in binding cores may impact binding specificity. This study presents a new representation for modeling binding cores by incorporating variations and an algorithm to discover them from only sequence data. Our algorithm takes protein and DNA sequences from TRANSFAC (a Protein-DNA Binding Database) as input; discovers from both sets of sequences conserved regions in Aligned Pattern Clusters (APCs); associates them as Protein-DNA Co-Occurring APCs; ranks the Protein-DNA Co-Occurring APCs according to their co-occurrence, and among the top ones, finds three-dimensional structures to support each binding core candidate. If successful, candidates are verified as binding cores. Otherwise, homology modeling is applied to their close matches in PDB to attain new chemically feasible binding cores. Our algorithm obtains binding cores with higher precision and much faster runtime ( ≥ 1,600x) than that of its contemporaries, discovering candidates that do not co-occur as one-to-one associated patterns in the raw data. AVAILABILITY: http://www.pami.uwaterloo.ca/~ealee/files/tcbbPnDna2015/Release.zip.
Annie En-Shiun Lee, Ho-Yin Sze-To, Man Hon Wong 0001, Kwong-Sak Leung, Terrence Chi-Kong Lau, Andrew K. C. Wong
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Partitioning and correlating subgroup characteristics from Aligned Pattern Clusters
abstract
MOTIVATION: Evolutionarily conserved amino acids within proteins characterize functional or structural regions. Conversely, less conserved amino acids within these regions are generally areas of evolutionary divergence. A priori knowledge of biological function and species can help interpret the amino acid differences between sequences. However, this information is often erroneous or unavailable, hampering discovery with supervised algorithms. Also, most of the current unsupervised methods depend on full sequence similarity, which become inaccurate when proteins diverge (e.g. inversions, deletions, insertions). Due to these and other shortcomings, we developed a novel unsupervised algorithm which discovers highly conserved regions and uses two types of information measures: (i) data measures computed from input sequences; and (ii) class measures computed using a priori class groupings in order to reveal subgroups (i.e. classes) or functional characteristics. RESULTS: Using known and putative sequences of two proteins belonging to a relatively uncharacterized protein family we were able to group evolutionarily related sequences and identify conserved regions, which are strong homologous association patterns called Aligned Pattern Clusters, within individual proteins and across the members of this family. An initial synthetic demonstration and in silico results reveal that (i) the data measures are unbiased and (ii) our class measures can accurately rank the quality of the evolutionarily relevant groupings. Furthermore, combining our data and class measures allowed us to interpret the results by inferring regions of biological importance within the binding domain of these proteins. Compared to popular supervised methods, our algorithm has a superior runtime and comparable accuracy. AVAILABILITY AND IMPLEMENTATION: The dataset and results are available at www.pami.uwaterloo.ca/∼ealee/files/classification2015 CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Annie En-Shiun Lee, Fiona J. Whelan, Dawn M. E. Bowdish, Andrew K. C. Wong
Bioinform.1
2015 Predicting Protein-protein interaction using co-occurring Aligned Pattern Clusters
abstract
Understanding Protein-protein interaction (PPI) is of fundamental importance in deciphering cellular processes. Predicting PPIs is thus critical in making new discoveries in the biological domains. Traditionally, new PPIs are identified through biochemical experiments but such methods are labor-intensive, expensive, time-consuming and technically ineffective due to high false positive rates. Computational docking is an alternative but requires the three-dimensional structures of the target proteins which are not always accessible. Sequence-based prediction is the most readily applicable and cost-effective method. It exploits known PPI Databases to construct classifiers for predicting unknown PPIs based only on sequence data. However, existing methods, adopting features that fix the pattern length and use exact patterns, are biologically unrealistic. Also, those based on SVM and String Kernel are hardly biologically interpretable since they do not compute the features. Recently, we have developed a new method for predicting PPI known as WeMine-P2P based on our WeMine Aligned Pattern Clustering algorithm which discovers and identifies the localized and co-occurring conserved patterns and regions allowing variable length and pattern variations. As our first attempt, under 40 independent experiments, we showed that (1) WeMine-P2P outperforms the well-known algorithm, PIPE2 which also utilizes co-occurring amino acid sequence segments but does not allow variable lengths and pattern variations; (2) Unlike SVM-based methods, WeMine-P2P renders interpretable biological features; (3) WeMine-P2P achieves satisfactory PPI prediction performance, comparable to the SVM-based methods particularly in unseen protein sequences, with a potential reduction of feature dimension of 1280x. WeMine-P2P is extendable to other biosequence interactions such as predicting Protein-DNA interactions.
Ho-Yin Sze-To, Sanderz Fung, Annie En-Shiun Lee, Andrew K. C. Wong
BIBM3
2014 Discovering protein-DNA binding cores by aligned pattern clustering
abstract
Understanding binding cores is of fundamental importance in deciphering Protein-DNA (TF-TFBS) binding and gene regulation. Variations (or mutations) in binding cores are ubiquitous and have different levels of effects on the binding specificity. To alleviate expensive experiments, we have developed a new method to discover directly from sequence data binding cores and study the effect due to variations. Although existing computational methods have produced satisfactory TF-TFBS binding cores, they are only one-to-one mappings with no site-specific information on residue/nucleotide variations; and also are largely overlapped. In this study, we propose a new representation for modeling TF-TFBS binding with variants known as TF-TFBS Co-Supportive Aligned Pattern Clusters (APCs), which are more compact, with more details for site-specific variants, and biologically more intuitive for analysis. To achieve this task, we have also developed an algorithm to discover TF-TFBS Co-Supportive APCs to capture binding cores at a higher precision with much faster runtime (≥1600X) comparing to other methods. The variants in TF-TFBS Co-Supportive APCs are also statistically analyzed and demonstrated that they can assist homology modeling to synthesize new biological knowledge.
Annie En-Shiun Lee, Kwong-Sak Leung, Ho-Yin Sze-To, Terrence Chi-Kong Lau, Man Hon Wong 0001, Andrew K. C. Wong
BIBM1
2014 Discovering co-occurring patterns and their biological significance in protein families
abstract
BACKGROUND: The large influx of biological sequences poses the importance of identifying and correlating conserved regions in homologous sequences to acquire valuable biological knowledge. These conserved regions contain statistically significant residue associations as sequence patterns. Thus, patterns from two conserved regions co-occurring frequently on the same sequences are inferred to have joint functionality. A method for finding conserved regions in protein families with frequent co-occurrence patterns is proposed. The biological significance of the discovered clusters of conserved regions with co-occurrences patterns can be validated by their three-dimensional closeness of amino acids and the biological functionality found in those regions as supported by published work. METHODS: Using existing algorithms, we discovered statistically significant amino acid associations as sequence patterns. We then aligned and clustered them into Aligned Pattern Clusters (APCs) corresponding to conserved regions with amino acid conservation and variation. When one APC frequently co-occurred with another APC, the two APCs have high co-occurrence. We then clustered APCs with high co-occurrence into what we refer to as Co-occurrence APC Clusters (Co-occurrence Clusters). RESULTS: Our results show that for Co-occurrence Clusters, the three-dimensional distance between their amino acids is closer than average amino acid distances. For the Co-occurrence Clusters of the ubiquitin and the cytochrome c families, we observed biological significance among the residing amino acids of the APCs within the same cluster. In ubiquitin, the residues are responsible for ubiquitination as well as conventional and unconventional ubiquitin-bindings. In cytochrome c, amino acids in the first co-occurrence cluster contribute to binding of other proteins in the electron transport chain, and amino acids in the second co-occurrence cluster contribute to the stability of the axial heme ligand. CONCLUSIONS: Thus, our co-occurrence clustering algorithm can efficiently find and rank conserved regions that contain patterns that frequently co-occurring on the same proteins. Co-occurring patterns are biologically significant due to their three-dimensional closeness and other evidences reported in literature. These results play an important role in drug discovery as biologists can quickly identify the target for drugs to conduct detailed preclinical studies.
Annie En-Shiun Lee, Sanderz Fung, Ho-Yin Sze-To, Andrew K. C. Wong
BMC Bioinform.1
2014 Aligning and Clustering Patterns to Reveal the Protein Functionality of Sequences
abstract
Discovering sequence patterns with variations unveils significant functions of a protein family. Existing combinatorial methods of discovering patterns with variations are computationally expensive, and probabilistic methods require more elaborate probabilistic representation of the amino acid associations. To overcome these shortcomings, this paper presents a new computationally efficient method for representing patterns with variations in a compact representation called Aligned Pattern Cluster (AP Cluster). To tackle the runtime, our method discovers a shortened list of non-redundant statistically significant sequence associations based on our previous work. To address the representation of protein functional regions, our pattern alignment and clustering step, presented in this paper captures the conservations and variations of the aligned patterns. We further refine our solution to allow more coverage of sequences via extending the AP Clusters containing only statistically significant patterns to Weak and Conserved AP Clusters. When applied to the cytochrome c, the ubiquitin, and the triosephosphate isomerase protein families, our algorithm identifies the binding segments as well as the binding residues. When compared to other methods, ours discovers all binding sites in the AP Clusters with superior entropy and coverage. The identification of patterns with variations help biologists to avoid time-consuming simulations and experimentations. (Software available upon request).
Andrew K. C. Wong, Annie En-Shiun Lee
IEEE ACM Trans. Comput. Biol. Bioinform.2
2013 Confirming biological significance of co-occurrence clusters of aligned pattern clusters
abstract
Advances in bioinformatics have provided researchers with a large influx of novel sequences, thus making the analysis of the sequences for inherent biological knowledge crucial. By using pattern discovery and pattern synthesis on protein family sequences, conserved protein segments can be represented by Aligned Pattern Clusters (APC), which is more knowledge-rich in statistical association comparing to probabilistic models. Such representation enabled us to exploit their co-occurrence on the same protein sequence to identify functional regions. In this paper, we developed an efficient algorithm to identify the frequently co-occurring patterns using only homologous protein sequences as input. We applied our algorithm to triosephosphate isomerase and ubiquitin for a detailed study. We found that the discovered co-occurring patterns are close in spatial distance in most cases, by comparing to corresponding 3D structures. We also found that the co-occurrence of patterns are biologically significant. Residues which play important and co-operative roles in the glycolytic pathway of triosephosphate isomerase and residues which are responsible for ubiquitination and ubiquitin-binding of ubiquitin are all covered in our co-occurring APCs. These results demonstrate the power of our algorithm to reveal the concurrent distant functional and structural relation of proteins sequences based on co-occurrence clusters of APCs.
Annie En-Shiun Lee, Sanderz Fung, Ho-Yin Sze-To, Andrew K. C. Wong
BIBM1
2013 Regrouping of pattern clusters to reveal characteristics of distinct classes and related classes
abstract
Discovering protein patterns for amino acids and their biochemical properties is important for revealing the underlying biophysical models. From this, pattern clustering was introduced in order to relate the discovered protein patterns to taxonomic classes in a localized region of a protein. This paper proposes an algorithm to synthesize and re-group pattern clusters, maximizing their separability in order to reveal class characteristics of the localized region of the protein based on our previous work. To evaluate the pattern clustering and regrouping pattern clusters results, we introduce three evaluation measures: F-measure, class entropy measure, and attribute entropy measure. To validate our proposed algorithm, experiments are run on synthetic data, protein family for amino acid attributes, and chemical property attributes. The experimental results show that: a) the result for regrouping pattern clusters is more accurate in class separation than only using pattern clustering; b) The clusters after regrouping are more distinctly separable with each other than only using pattern clustering; c) two types of pattern clusters are found, with one pertaining to distinct classes and the other associating with two or more related classes; and d) class characteristics are clearly revealed in the data subspace containing the patterns in the pattern clusters. The datasets with chemical properties show that unsupervised techniques can reveal common chemical attributes in the inherent classes as more of the common properties shared by different amino acids are taken into account
Pei-Yuan Zhou, Annie En-Shiun Lee, Andrew K. C. Wong
BIBM2
2012 Identifying protein binding functionality of protein family sequences by Aligned Pattern clusters
abstract
A basic task in protein analysis is to discover a set of sequence patterns that reflect the function of a protein family. This set of sequence patterns contains non-exact significant residue associations. Currently, the existing combinatorial methods are computationally expensive and probabilistic methods require richer representation of the amino acid associations. To undertake this task, we create a synthesized pattern representation called an Aligned Pattern (AP) Cluster that identifies the residue associations in the binding segment and the site variations in the aligned residues. In this paper, our algorithm identifies the binding segments for two protein families: the Cytochrome Complex and the Ubiquitin protein families. For each of the experiments, the AP Clusters obtained correspond to protein binding segments including a few beyond those identified by the other protein databases, PROSITE and pFam. Furthermore, the columns of aligned sites that exist only as a single value in the AP Clusters also corresponds to the binding residues. Additional information retained by the AP Clusters can reveal the amino acid residues of interest, thus averting time-consuming simulations and experimentation.
Annie En-Shiun Lee, Andrew K. C. Wong
BIBM1
2012 Discovery of Delta Closed Patterns and Noninduced Patterns from Sequences
abstract
Discovering patterns from sequence data has significant impact in many aspects of science and society, especially in genomics and proteomics. Here we consider multiple strings as input sequence data and substrings as patterns. In the real world, usually a large set of patterns could be discovered yet many of them are redundant, thus degrading the output quality. This paper improves the output quality by removing two types of redundant patterns. First, the notion of delta tolerance closed itemset is employed to remove redundant patterns that are not delta closed. Second, the concept of statistically induced patterns is proposed to capture redundant patterns which seem to be statistically significant yet their significance is induced by their strong significant subpatterns. It is computationally intense to mine these nonredundant patterns (delta closed patterns and noninduced patterns). To efficiently discover these patterns in very large sequence data, two efficient algorithms have been developed through innovative use of suffix tree. Three sets of experiments were conducted to evaluate their performance. They render excellent results when applying to genomics. The experiments confirm that the proposed algorithms are efficient and that they produce a relatively small set of patterns which reveal interesting information in the sequences.
Andrew K. C. Wong, Dennis Zhuang, Gary C. L. Li, Annie En-Shiun Lee
IEEE Trans. Knowl. Data Eng.4