Randy Goebel

dblp:g/RandyGoebel · also Randy G. Goebel · DBLP profile ↗
← Back
103ranked-venue papers
13as first author
18since 2021 · last 2026
0000-0002-0739-2946ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 3 since 2021Theory of computation · 20 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 17 · 3 since 2021Human-computer interaction and ubiquitous computing · 14 · 4 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 Feature-Level Interaction Explanations in Multimodal Transformers
Yeji Kim, Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
ICPR (14)4
2026 Reason2Decide: Rationale-Driven Multi-Task Learning
abstract
Despite the wide adoption of Large Language Models (LLM)s, clinical decision support systems face a critical challenge: achieving high predictive accuracy while generating explanations aligned with the predictions. Current approaches suffer from exposure bias leading to misaligned explanations. We propose Reason2Decide, a two-stage training framework that addresses key challenges in self-rationalization, including exposure bias and task separation. In Stage-1, our model is trained on rationale generation, while in Stage-2, we jointly train on label prediction and rationale generation, applying scheduled sampling to gradually transition from conditioning on gold labels to model predictions. We evaluate Reason2Decide on three medical datasets, including a proprietary triage dataset and public biomedical QA datasets. Across model sizes, Reason2Decide outperforms other fine-tuning baselines and some zero-shot LLMs in prediction (F1) and rationale fidelity (BERTScore, BLEU, LLM-as-a-Judge). In triage, Reason2Decide is rationale source-robust across LLM-generated, nurse-authored, and nurse-post-processed rationales. In our experiments, while using only LLM-generated rationales in Stage-1, Reason2Decide outperforms other fine-tuning variants. This indicates that LLM-generated rationales are suitable for pretraining models, reducing reliance on human annotations. Remarkably, Reason2Decide achieves these gains with models 40x smaller than contemporary foundation models, making clinical reasoning more accessible for resource-constrained deployments while still providing explainable decision support.
H. M. Quamran Hasan, Housam Khalifa Bashier Babiker, Jiayi Dai, Mi-Young Kim, Randy Goebel
LREC5
2026 Approximation schemes for multiprocessor scheduling within budget
abstract
Scheduling within a limited budget is closely related to several much studied scheduling with compression and rescheduling problems, and is also a variant of the more recent model of scheduling with testing. In this problem, one seeks to minimize the makespan for a set of jobs that are to be processed on a number of parallel identical machines. Each job J j is given an upper bound u j on its actual processing time p j , and a testing fee c j . In the offline case, the processing time p j is known to the scheduler; while in the oblivious case, p j is revealed to the scheduler only if the testing fee c j is paid. So the scheduler can choose to execute J j on any one of the machines non-preemptively for u j time or pay for the testing and then execute the job for p j time. The scheduler is given a budget B to pay for testing and seeks to minimize the makespan within that budget. The offline problem is denoted as P ∣ u j , p j , c j , B ∣ C max , and the oblivious problem is denoted as P ∣ u j , − , c j , B ∣ C max , where P stands for multiple parallel identical machines with the number of machines being part of the input. We contribute a polynomial-time approximation scheme (PTAS) for the offline problem P ∣ u j , p j , c j , B ∣ C max , which leads to an almost tight ( 2 + ϵ ) -competitive algorithm for the oblivious problem P ∣ u j , − , c j , B ∣ C max .
Mingyang Gong, Randy Goebel, Guohui Lin, Bing Su 0002
Theor. Comput. Sci.2
2025 MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
abstract
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP26
2025 An Overview of the COLIEE 2025 Competition: Legal Case Law and Statute Law Information Retrieval and Entailment
abstract
We summarize the 12th Competition on Legal Information Extraction and Entailment. In this edition, the competition included four tasks on case law and statute law, plus a new pilot task on Tort law. The case law component includes an information retrieval task (Task 1), and the confirmation of an entailment relation between an existing case and an unseen case (Task 2). The statute law component includes an information retrieval task (Task 3), and an entailment/question-answering task based on retrieved civil code statutes (Task 4). The new pilot task is tort prediction (TP) and its rationale extraction (RE).
Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Calum Kwan, Ken Satoh, Hiroaki Yamada 0002, Masaharu Yoshioka
ICAIL1
2025 Getting SMARTER for Motion Planning in Autonomous Driving Systems
abstract
Motion planning is a fundamental problem in autonomous driving and perhaps the most challenging to comprehensively evaluate because of the associated risks and expenses of real-world deployment. Therefore, simulations play an important role in efficient development of planning algorithms. To be effective, simulations must be accurate and realistic, both in terms of dynamics and behavior modeling, and also highly customizable in order to accommodate a broad spectrum of research frameworks. In this paper, we introduce SMARTS 2.0, the second generation of our motion planning simulator which, in addition to being highly optimized for large-scale simulation, provides many new features, such as realistic map integration, vehicle-to-vehicle (V2V) communication, traffic and pedestrian simulation, and a broad variety of sensor models. Moreover, we present a novel benchmark suite for evaluating planning algorithms in various highly challenging scenarios, including interactive driving, such as turning at intersections, and adaptive driving, in which the task is to closely follow a lead vehicle without any explicit knowledge of its intention. Each scenario is characterized by a variety of traffic patterns and road structures. We further propose a series of common and task-specific metrics to effectively evaluate the performance of the planning algorithms. At the end, we evaluate common motion planning algorithms using the proposed benchmark and highlight the challenges the proposed scenarios impose. The new SMARTS 2.0 features and the benchmark are publicly available at github.com/huawei-noah/SMARTS.
Montgomery Alban, Ehsan Ahmadi, Randy Goebel, Amir Rasouli
IV3
2025 Safety Implications of Explainable Artificial Intelligence in End-to-End Autonomous Driving
abstract
The end-to-end learning pipeline is gradually creating a paradigm shift in the ongoing development of highly autonomous vehicles (AVs), largely due to advances in deep learning, the availability of large-scale training datasets, and improvements in integrated sensor devices. However, a lack of explainability in real-time decisions with contemporary learning methods impedes user trust and attenuates the widespread deployment and commercialization of such vehicles. Moreover, the issue is exacerbated when these vehicles are involved in or cause traffic accidents. Consequently, explainability in end-to-end autonomous driving is essential to build trust in vehicular automation. With that said, automotive researchers have not yet rigorously explored safety benefits and consequences of explanations in end-to-end autonomous driving. This paper aims to bridge the gaps between these topics and seeks to answer the following research question: What are safety implications of explanations in end-to-end autonomous driving? In this regard, we first revisit established safety and explainability concepts in end-to-end driving. Furthermore, we present critical case studies and show the pivotal role of explanations in enhancing driving safety. Finally, we describe insights from empirical studies and reveal potential value, limitations, and caveats of practical explainable AI methods with respect to their potential impacts on safety of end-to-end driving.
Shahin Atakishiyev, Mohammad Salameh, Randy Goebel
IEEE Trans. Intell. Transp. Syst.3
2024 Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language Models
abstract
Developing the capacity to effectively search for requisite datasets is an urgent requirement to assist data users in identifying relevant datasets considering the very limited available metadata. For this challenge, the utilization of third-party data is emerging as a valuable source for improvement. Our research introduces a new architecture for data exploration which employs a form of Retrieval-Augmented Generation (RAG) to enhance metadata-based data discovery. The system integrates large language models (LLMs) with external vector databases to identify semantic relationships among diverse types of datasets. The proposed framework offers a new method for evaluating semantic similarity among heterogeneous data sources and for improving data exploration. Our study includes experimental results on four critical tasks: 1) recommending similar datasets, 2) suggesting combinable datasets, 3) estimating tags, and 4) predicting variables. Our results demonstrate that RAG can enhance the selection of relevant datasets, particularly from different categories, when compared to conventional metadata approaches. However, performance varied across tasks and models, which confirms the significance of selecting appropriate techniques based on specific use cases. The findings suggest that this approach holds promise for addressing challenges in data exploration and discovery, although further refinement is necessary for estimation tasks.
Teruaki Hayashi, Hiroki Sakaji, Jiayi Dai, Randy Goebel
IEEE Big Data4
2024 Juris-Informatics: Law for AI and Law of AI
abstract
This paper presents an outline of our research project developed at our research center for "Juris-Informatics". "Juris-Informatics" is a research field based on two main topics; "Law by AI" and "Law of AI". "Law by Ai" is a research field where we investigate a support tool by AI for legal activities such as legal reasoning and legal document processing. "Law of AI" is a research field where we conduct research on legal control of AI such as considering the legal responsibility of AI and legal compliance of AI.
Ken Satoh, Hideaki Takeda 0001, Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Juliano Rabelo 0001, Masaharu Yoshioka
IEEE Big Data3
2024 Incorporating Explanations into Human-Machine Interfaces for Trust and Situation Awareness in Autonomous Vehicles
abstract
Autonomous vehicles often make complex decisions via machine learning-based predictive models applied to collected sensor data. While this combination of methods provides a foundation for real-time actions, self-driving behavior primarily remains opaque to end users. In this sense, explainability of real-time decisions is a crucial and natural requirement for building trust in autonomous vehicles. Moreover, as autonomous vehicles still cause serious traffic accidents for various reasons, timely conveyance of upcoming hazards to road users can help improve scene understanding and prevent potential risks. Hence, there is also a need to supply autonomous vehicles with user-friendly interfaces for effective human-machine teaming. Motivated by this problem, we study the role of explainable AI and human-machine interface jointly in building trust in vehicle autonomy. We first present a broad context of the explanatory human-machine systems with the "3W1H" (what, whom, when, how) approach. Based on these findings, we present a situation awareness framework for calibrating users’ trust in self-driving behavior. Finally, we perform an experiment on our framework, conduct a user study on it, and validate the empirical findings with hypothesis testing.
Shahin Atakishiyev, Mohammad Salameh, Randy Goebel
IV3
2023 The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy Measure
abstract
The popular Proximal Policy Optimization (PPO) algorithm approximates the solution in a clipped policy space. Does there exist better policies outside of this space? By using a novel surrogate objective that employs the sigmoid function (which provides an interesting way of exploration), we found that the answer is "YES", and the better policies are in fact located very far from the clipped space. We show that PPO is insufficient in "off-policyness", according to an off-policy metric called DEON. Our algorithm explores in a much larger policy space than PPO, and it maximizes the Conservative Policy Iteration (CPI) objective better than PPO during training. To the best of our knowledge, all current PPO methods have the clipping operation and optimize in the clipped policy space. Our method is the first of this kind, which advances the understanding of CPI optimization and policy gradient methods. Code is available at https://github.com/raincchio/P3O.
Xing Chen 0022, Dongcui Diao, Hechang Chen, Hengshuai Yao, Haiyin Piao, Zhixiao Sun, Zhiwei Yang 0005, Randy Goebel, Bei Jiang, Yi Chang 0001
AAAI8
2023 From Intermediate Representations to Explanations: Exploring Hierarchical Structures in NLP
abstract
Interpretation methods for learned models used in natural language processing (NLP) applications usually provide support for local (specific) explanations, such as quantifying the contribution of each word to the predicted class. But they typically ignore the potential interaction amongst those word tokens. Unlike currently popular methods, we propose a deep model which uses feature attribution and identification of dependencies to support the learning of interpretable representations that will support creation of hierarchical explanations. In addition, hierarchical explanations provide a basis for visualizing how words and phrases are combined at different levels of abstraction, which enables end-users to better understand the prediction process of a deep network. Our study uses multiple well-known datasets to demonstrate the effectiveness of our approach, and provides both automatic and human evaluation.
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
ECAI3
2023 Summary of the Competition on Legal Information, Extraction/Entailment (COLIEE) 2023
abstract
We summarize the 10th Competition on Legal Information Extraction and Entailment. In this edition, the competition included four tasks on case law and statute law. The case law component includes an information retrieval task (Task 1), and the confirmation of an entailment relation between an existing case and an unseen case (Task 2). The statute law component includes an information retrieval task (Task 3), and an entailment/question answering task based on retrieved civil code statutes (Task 4). Participation was open to any group based on any approach. Ten different teams participated in the case law competition tasks, most of them in more than one task. We received results from 8 teams for Task 1 (22 runs) and seven teams for Task 2 (18 runs). On the statute law task, there were 9 different teams participating, most in more than one task. 6 teams submitted a total of 16 runs for Task 3, and 9 teams submitted a total of 26 runs for Task 4. We describe the variety of approaches, our official evaluation, and analysis of our data and submission results.
Randy Goebel, Yoshinobu Kano, Mi-Young Kim, Juliano Rabelo 0001, Ken Satoh, Masaharu Yoshioka
ICAIL1
2023 LawGiBa - Combining GPT, Knowledge Bases, and Logic Programming in a Legal Assistance System
abstract
We present LawGiBa, a proof-of-concept demonstration system for legal assistance that combines GPT, legal knowledge bases, and Prolog’s logic programming structure to provide explanations for legal queries. This novel combination effectively and feasibly addresses the hallucination issue of large language models (LLMs) in critical domains, such as law. Through this system, we demonstrate how incorporating a legal knowledge base and logical reasoning can enhance the accuracy and reliability of legal advice provided by AI models like GPT. Though our work is primarily a demonstration, it provides a framework to explore how knowledge bases and logic programming structures can be further integrated with generative AI systems, to achieve improved results across various natural languages and legal systems.
Ha-Thanh Nguyen, Randy Goebel, Francesca Toni, Kostas Stathis, Ken Satoh
JURIX2
2022 Locally Distributed Activation Vectors for Guided Feature Attribution
abstract
Explaining the predictions of a deep neural network (DNN) is a challenging problem. Many attempts at interpreting those predictions have focused on attribution-based methods, which assess the contributions of individual features to each model prediction. However, attribution-based explanations do not always provide faithful explanations to the target model, e.g., noisy gradients can result in unfaithful feature attribution for back-propagation methods. We present a method to learn explanations-specific representations while constructing deep network models for text classification. These representations can be used to faithfully interpret black-box predictions, i.e., highlighting the most important input features and their role in any particular prediction. We show that learning specific representations improves model interpretability across various tasks, for both qualitative and quantitative evaluations, while preserving predictive performance.
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
COLING3
2022 Neural Networks with Feature Attribution and Contrastive Explanations
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
ECML/PKDD (1)3
2021 DISK-CSV: Distilling Interpretable Semantic Knowledge with a Class Semantic Vector
abstract
Neural networks (NN) applied to natural language processing (NLP) are becoming deeper and more complex, making them increasingly difficult to understand and interpret.Even in applications of limited scope on fixed data, the creation of these complex "black-boxes" creates substantial challenges for debugging, understanding, and generalization.But rapid development in this field has now lead to building more straightforward and interpretable models.We propose a new technique (DISK-CSV) to distill knowledge concurrently from any neural network architecture for text classification, captured as a lightweight interpretable/explainable classifier.Across multiple datasets, our approach achieves better performance than the target black-box.In addition, our approach provides better explanations than existing techniques.
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
EACL3
2021 Explainable Zero-Shot Modelling of Clinical Depression Symptoms from Text
abstract
We focus on exploring various approaches of Zero-Shot Learning (ZSL) and their explainability for a challenging yet important supervised learning task, notorious for training data scarcity, i.e. Depression Symptoms Detection (DSD) from text. We start with a comprehensive synthesis of different components of our ZSL modelling and analysis of our ground truth samples and Depression symptom clues curation process with the help of a practicing Clinician. We next analyze the accuracy of various state-of-the-art ZSL models and their potential enhancements for our task. Further, we sketch a framework for the use of ZSL for hierarchical text-based explanation mechanism, which we call, Syntax Tree-Guided Semantic Explanation (STEP). Finally, we summarize experiments from which we conclude that we can use ZSL models and achieve reasonable accuracy and explainability, measured by a proposed Explainability Index (EI). This work is, to our knowledge, the first work to exhaustively explore the efficacy of ZSL models for DSD task, both in terms of accuracy and explainability.
Nawshad Farruque, Randy Goebel, Osmar R. Zaïane, Sudhakar Sivapalan
ICMLA2
2020 Explainable Artificial Intelligence: Concepts, Applications, Research Challenges and Visions
Luca Longo, Randy Goebel, Freddy Lécué, Peter Kieseberg, Andreas Holzinger
CD-MAKE2
2020 RANCC: Rationalizing Neural Networks via Concept Clustering
abstract
We propose a new self-explainable model for Natural Language Processing (NLP) text classification tasks.Our approach constructs explanations concurrently with the formulation of classification predictions.To do so, we extract a rationale from the text, then use it to predict a concept of interest as the final prediction.We provide three types of explanations: 1) rationale extraction, 2) a measure of feature importance, and 3) clustering of concepts.In addition, we show how our model can be compressed without applying complicated compression techniques.We experimentally demonstrate our explainability approach on a number of well-known text classification datasets.
Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
COLING3
2020 Open-shop scheduling for unit jobs under precedence constraints
Yong Chen 0002, Randy Goebel, Guohui Lin, Bing Su 0002, An Zhang 0001
Theor. Comput. Sci.2
2020 Approximation algorithms for the three-machine proportionate mixed shop scheduling
Longcheng Liu, Yong Chen 0002, Randy Goebel, Guohui Lin, Guanqun Ni, Bing Su 0002, An Zhang 0001
Theor. Comput. Sci.4
2019 A Multi-Task Learning Framework for Abstractive Text Summarization
abstract
We propose a Multi-task learning approach for Abstractive Text Summarization (MATS), motivated by the fact that humans have no difficulty performing such task because they have the capabilities of multiple domains. Specifically, MATS consists of three components: (i) a text categorization model that learns rich category-specific text representations using a bi-LSTM encoder; (ii) a syntax labeling model that learns to improve the syntax-aware LSTM decoder; and (iii) an abstractive text summarization model that shares its encoder and decoder with the text categorization and the syntax labeling tasks, respectively. In particular, the abstractive text summarization model enjoys significant benefit from the additional text categorization and syntax knowledge. Our experimental results show that MATS outperforms the competitors.1
Linqing Liu, Zhile Jiang, Min Yang 0007, Randy Goebel
AAAI5
2019 Basic and Depression Specific Emotions Identification in Tweets: Multi-label Classification Experiments
Nawshad Farruque, Chenyang Huang 0001, Osmar R. Zaïane, Randy Goebel
CICLing (2)4
2019 Statute Law Information Retrieval and Entailment
abstract
Our Yes/No statute law question answering system combines components for both statute law information retrieval and confirmation of textual entailment between statues and legal questions. We describe a statute law question answering system that exploits TF-IDF and a language model for information retrieval, and inter-paragraph entailment. We have evaluated our system using the data from the competition on legal information extraction/entailment (COLIEE-2019). The competition consists of four tasks: Tasks 1 and 2 are for the case law information extraction/entailment, and Tasks 3 and 4 are for the statute law information extraction/entailment. Here we explain our methods and evaluation results for Tasks 3 and 4. Task 3 requires the identification of civil law articles relevant to Japan legal bar exam query. For this task, we used TF-IDF and language model-based information retrieval approaches. Task 4 requires a decision on yes/no answer for previously unseen queries given relevant civil law articles. Our approach compares the approximate meanings of queries with relevant articles. Because many statute law and queries consist of more than one paragraph, we need an inter-paragraph entailment method. Our inter-paragraph entailment process exploits an analysis of statute law structure, and negation patterns to predict entailments. Using our heuristic selection of attributes, we perform two experiments which provide the basis for making a decision on the yes/no questions. One experiment uses an SVM model, and the other uses a general heuristic rule. Our experimental evaluation demonstrates the value of our method, and the results show that our method was ranked No. 1 in both of the Tasks 3 and 4 in COLIEE 2019.
Mi-Young Kim, Juliano Rabelo 0001, Randy Goebel
ICAIL3
2019 Combining Similarity and Transformer Methods for Case Law Entailment
abstract
We tackle the complex problem of determining entailment relationships between case law documents, one of the tasks in the Competition on Legal Information Extraction and Entailment (COLIEE). With input of an entailed fragment from a case coupled with a candidate entailing paragraph from a noticed case, our approach relies on four main components: (1) extraction of similarity measures between the two pieces of text; (2) application of a transformer-based technique on the input text; (3) applying a threshold-based classifier; and (4) post-processing the results considering the a priori probability determined by the data distribution on the training samples and combining the results of (1) and (2). Our experiments achieved an F-score of 0.70 on the official COLIEE test dataset, ranking first among all competitors for that task in the 2019 competition.
Juliano Rabelo 0001, Mi-Young Kim, Randy Goebel
ICAIL3
2019 A 21/16-Approximation for the Minimum 3-Path Partition Problem
abstract
The minimum k-path partition (Min-k-PP for short) problem targets to partition an input graph into the smallest number of paths, each of which has order at most k. We focus on the special case when k=3. Existing literature mainly concentrates on the exact algorithms for special graphs, such as trees. Because of the challenge of NP-hardness on general graphs, the approximability of the Min-3-PP problem attracts researchers' attention. The first approximation algorithm dates back about 10 years and achieves an approximation ratio of 3/2, which was recently improved to 13/9 and further to 4/3. We investigate the 3/2-approximation algorithm for the Min-3-PP problem and discover several interesting structural properties. Instead of studying the unweighted Min-3-PP problem directly, we design a novel weight schema for l-paths, l in {1, 2, 3}, and investigate the weighted version. A greedy local search algorithm is proposed to generate a heavy path partition. We show the achieved path partition has the least 1-paths, which is also the key ingredient for the algorithms with ratios 13/9 and 4/3. When switching back to the unweighted objective function, we prove the approximation ratio 21/16 via amortized analysis.
Yong Chen 0002, Randy Goebel, Bing Su 0002, Weitian Tong, An Zhang 0001
ISAAC2
2019 Augmenting Semantic Representation of Depressive Language: From Forums to Microblogs
Nawshad Farruque, Osmar R. Zaïane, Randy Goebel
ECML/PKDD (3)3
2018 Approximation Algorithms and a Hardness Result for the Three-Machine Proportionate Mixed Shop
Longcheng Liu, Guanqun Ni, Yong Chen 0002, Randy Goebel, An Zhang 0001, Guohui Lin
AAIM4
2018 Explainable AI: The New 42?
Randy Goebel, Ajay Chander, Katharina Holzinger, Freddy Lécué, Zeynep Akata, Simone Stumpf, Peter Kieseberg, Andreas Holzinger
CD-MAKE1
2018 Open-Shop Scheduling for Unit Jobs Under Precedence Constraints
An Zhang 0001, Yong Chen 0002, Randy Goebel, Guohui Lin
COCOA3
2018 Approximation Algorithms for Two-Machine Flow-Shop Scheduling with a Conflict Graph
Yinhui Cai, Guangting Chen, Yong Chen 0002, Randy Goebel, Guohui Lin, Longcheng Liu, An Zhang 0001
COCOON4
2018 Algorithms for Communication Scheduling in Data Gathering Network with Data Compression
Wenchang Luo, Boyuan Gu, Weitian Tong, Randy Goebel, Guohui Lin
Algorithmica5
2018 An approximation scheme for minimizing the makespan of the parallel identical multi-stage flow-shops
Weitian Tong, Eiji Miyano, Randy Goebel, Guohui Lin
Theor. Comput. Sci.3
2017 Two-step cascaded textual entailment for legal bar exam question answering
abstract
Our legal question answering system combines legal information retrieval and textual entailment, and exploits semantic information using a logic-based representation. We have evaluated our system using the data from the competition on legal information extraction/entailment (COLIEE)-2017. The competition focuses on the legal information processing required to answer yes/no questions from Japanese legal bar exams, and it consists of two phases: ad hoc legal information retrieval (Phase 1), and textual entailment (Phase 2). Phase 1 requires the identification of Japan civil law articles relevant to a legal bar exam query. For this phase, we have used an information retrieval approach using TF-IDF combined with a simple language model. Phase 2 requires a yes/no decision for previously unseen queries, which we approach by comparing the approximate meanings of queries with relevant statutes. Our meaning extraction process uses a selection of features based on a kind of paraphrase, coupled with a condition/conclusion/exception analysis of articles and queries. We also extract and exploit negation patterns from the articles. We construct a logic-based representation as a semantic analysis result, and then classify questions into easy and difficult types by analyzing the logic representation. If a question is in our easy category, we simply obtain the entailment answer from the logic representation; otherwise we use an unsupervised learning method to obtain the entailment answer. Experimental evaluation shows that our result ranked highest in the Phase 2 amongst all COLIEE-2017 competitors.
Mi-Young Kim, Randy Goebel
ICAIL2
2017 Facial expression recognition using SVM classification on mic-macro patterns
abstract
The identification of facial expressions is a fundamental topic in the area of human computer interaction and pattern recognition. The research has gained significant attention in recent years. However many challenges still exist. This is because an individual might display different expressions at different times for the same mood. Expressions can also be influenced by health. Our proposed framework aims to capture unique information related to expressions from salient patches. We extract representative feature patterns at both micro and macro levels within a pixel-patch, and use a support vector machine (SVM) classifier to label expressions. Our experimental results using the Japanese facial expression (JAFEE) and Cohn-Kanade (CK) datasets achieve high recognition rate and efficient computation time, outperforming existing work.
Housam Khalifa Bashier Babiker, Randy Goebel, Irene Cheng 0001
ICIP2
2016 Why Visualization is an AI-complete Problem (and Why That Matters)
abstract
Artificial Intelligence (AI) has infiltrated almost every scientific and social endeavour, including everything from medical research to the sociology of crowd control. But the foundation of AI continues to be based on digital representations of knowledge, and computational reasoning therewith. Because so much of modern knowledge infrastructure and social behaviour is connected to AI, understanding the role of AI in each such endeavour not only helps accelerate progress in those fields in which it applies, but also creates the challenges to extend the foundation for modern AI methods. The simple hypothesis herein is that so-called AI-complete problems have a role in helping to articulate the appropriate integration of AI within other disciplines. With the current growth of interest in "big data" and visualization, we argue that relatively simple formal structures provide a basis for the claim that visualization is an AI-complete problem. The value of confirming this claim is largely to encourage stronger formalizations of the visualization process in terms of the AI foundations of representation and reasoning. This connection will help ensure that relevant components of AI are appropriately applied and integrated, to provide value for a basis of a theory of visualization. The sketch of this claim here is based on the simple idea that visualization is an abstraction process, and that abstractions from partial information, however voluminous, directly confronts the non monotonic reasoning challenge, thus the need for caution in engineering visualization systems without carefully considering the consequences of visual abstraction. This is particularly important with interactive visualization, which has recently formed the basis for such fields as visual analytics.
Randy Goebel
IV1
2016 Smoothed heights of tries and patricia tries
Weitian Tong, Randy Goebel, Guohui Lin
Theor. Comput. Sci.2
2015 A Visualization-Analytics-Interaction Workflow Framework for Exploratory and Explanatory Search on Geo-located Search Data Using the Meme Media Digital Dashboard
abstract
Modern geo-position system (GPS) enabled smart phones are generating an increasing volume of information about their users, including geo-located search, movement, and transaction data. While this kind of data is increasingly rich and offers many grand opportunities to identify patterns and predict behaviour of groups and individuals, it is not immediately obvious how to develop a framework for extracting plausible inferences from these data. In our case, we have access to a large volume (more than half a billion individual records) of real user data from the Point smart phone application, and we have developed a generic and layered system architecture to incrementally find aggregate items of interest within that data. "Interest" is based on the semantics of the data, so include time and space correlations, e.g., Are people searching for dinner and a movie, distributions of usage patterns and platforms, e.g., Geographic distribution of Android, Apple, and Black-Berry users, and clustering to identify interesting and relatively complex search and movement patterns, e.g., Consumer trajectories from key word searches. Our integration of visualization tools is thus guided top-down, by semantic concepts in the application domain, rather than by bottom-up tool development. Our presentation here is preliminary in that we provide sketches of case-studies that demonstrate an application specific integration of the three major components of modern visual analytics: visualization, analytics, and interaction (VAI). Our case-study sketches show how an interactive system for visual data exploration can be used to alternate between exploratory search -- looking for ideas and new hypothesis in data -- and explanatory search -- looking for evidence to support a hypothesis. While we have not yet formulated experiments to directly measure the cognitive efficacy of our experimental system, we believe that our semantically-driven VAI workflows and the integration of visual methods and interaction provides some useful ideas about how to extend current frameworks for visual analytics systems.
Jonas Sjöbergh, Xingkai Li, Randy Goebel, Yuzuru Tanaka
IV3
2015 Recognition of Patient-Related Named Entities in Noisy Tele-Health Texts
abstract
We explore methods for effectively extracting information from clinical narratives that are captured in a public health consulting phone service called HealthLink. Our research investigates the application of state-of-the-art natural language processing and machine learning to clinical narratives to extract information of interest. The currently available data consist of dialogues constructed by nurses while consulting patients by phone. Since the data are interviews transcribed by nurses during phone conversations, they include a significant volume and variety of noise. When we extract the patient-related information from the noisy data, we have to remove or correct at least two kinds of noise: explicit noise , which includes spelling errors, unfinished sentences, omission of sentence delimiters, and variants of terms, and implicit noise , which includes non-patient information and patient's untrustworthy information. To filter explicit noise, we propose our own biomedical term detection/normalization method: it resolves misspelling, term variations, and arbitrary abbreviation of terms by nurses. In detecting temporal terms, temperature, and other types of named entities (which show patients’ personal information such as age and sex), we propose a bootstrapping-based pattern learning process to detect a variety of arbitrary variations of named entities. To address implicit noise, we propose a dependency path-based filtering method. The result of our denoising is the extraction of normalized patient information, and we visualize the named entities by constructing a graph that shows the relations between named entities. The objective of this knowledge discovery task is to identify associations between biomedical terms and to clearly expose the trends of patients’ symptoms and concern; the experimental results show that we achieve reasonable performance with our noise reduction methods.
Mi-Young Kim, Ying Xu 0003, Osmar R. Zaïane, Randy Goebel
ACM Trans. Intell. Syst. Technol.4
2014 On the Smoothed Heights of Trie and Patricia Index Trees
Weitian Tong, Randy Goebel, Guohui Lin
COCOON2
2014 Model Selection for Semi-Supervised Clustering
abstract
Although there is a large and growing literature that tackles the semi-supervised clustering problem (i.e., using some labeled objects or cluster-guiding constraints like \\must-link" or \\cannot-link"), the evaluation of semi-supervised clustering approaches has rarely been discussed. The application of cross-validation techniques, for example, is far from straightforward in the semi-supervised setting, yet the problems associated with evaluation have yet to be addressed. Here we \nsummarize these problems and provide a solution. \nFurthermore, in order to demonstrate practical applicability of semi-supervised clustering methods, we provide a method for model selection in semi-supervised clustering based on this sound evaluation procedure. Our method allows the user to select, based on the available information \n(labels or constraints), the most appropriate clustering model (e.g., number of clusters, density-parameters) for a given problem.
Mojgan Pourrajabi, Davoud Moulavi, Ricardo J. G. B. Campello, Arthur Zimek, Jörg Sander 0001, Randy Goebel
EDBT6
2014 The Challenge of Semantic Symmetry in Visualization
abstract
We present a fundamental problem which arises within an emerging theory of visualization, and provide examples that illustrate the challenge of what we call semantic symmetry. This theory of visualization distinguishes data domains (e.g., Numbers and symbols) from picture domains (e.g., Shapes, shading, colour), and provides a framework for specifying a variety of mappings between data and picture domains. Visualization is about enabling inferences about data within the human visual system, so crucially depends on the management of mappings from the data to the picture domain. But there are many possible choices for these mappings, and only in the last decade has there emerged any serious assessment of how one might measure the quality of a visualization. The situation is further complicated by what is now called visual analytics, where data to picture mappings allow manipulation of that picture to further understand or reveal the underlying data relationships. This kind of picture manipulation is exactly the departure point for our presentation of the problem of semantic symmetry. Semantic symmetry considers the problem of how to couple data and picture so that changes in one are accurately reflected in the other. We illustrate the foundational nature of the problems arising from the desire for semantic symmetry, and explain the kinds of constraints and framework that are necessary in order to be able to support a more complete theory of visualization.
Randy Goebel, Yuzuru Tanaka
IV1
2014 Towards the Identification of Consumer Trajectories in Geo-Located Search Data
abstract
Modern geo-positioning system (GPS) enabled smart phones are generating an increasing volume of information about their users, including geo-located search, movement, and transaction data. While this kind of data is increasingly rich and offers many grand opportunities to identify patterns and predict behaviour of groups and individuals, it is not immediately obvious how to develop a framework for extracting plausible inferences from these data. In our case, we have access to a large volume of real user data from the Point smart phone application, and we have developed a generic and layered system architecture to incrementally find aggregate items of interest within that data. This includes time and space correlations, e.g., are people searching for dinner and a movie, distributions of usage patterns and platforms, e.g., geographic distribution of Android, Apple, and BlackBerry users, and clustering to identify relatively complex search and movement patterns we call "consumer trajectories." Our pursuit of these kinds of patterns has helped guide our development of information extraction, machine learning, and visualization methods that provide systematic tools for investigating the geo-located data, and for the development of both conceptual tools and visualization tools in aid of finding both interesting and useful patterns in that data. Included in our system architecture is the ability to consider the difference between exploratory and explanatory hypotheses on data patterns, as well as the deployment of multiple visualization methods that can provide alternatives to help expose interesting patterns. In our introduction to our framework here, we provide examples of formulating hypotheses on geo-located behaviour, and how a variety of methods including those from machine learning and visualization, can help confirm or deny the value of such hypotheses as they emerge. In this particular case, we provide an initial basis for identifying semantically motivated data artifacts we call geo-located consumer trajectories. We investigate their plausibility with a variety of time and space series clustering and visualization models.
Xingkai Li, Randy Goebel, Jonas Sjöbergh
IV2
2014 On the approximability of the exemplar adjacency number problem for genomes with gene repetitions
Zhixiang Chen 0001, Randy Goebel, Guohui Lin, Weitian Tong, Jinhui Xu 0001, Boting Yang, Binhai Zhu
Theor. Comput. Sci.3
2014 Approximating the minimum independent dominating set in perturbed graphs
Weitian Tong, Randy Goebel, Guohui Lin
Theor. Comput. Sci.2
2014 Approximating the maximum multiple RNA interaction problem
Weitian Tong, Randy Goebel, Tian Liu 0001, Guohui Lin
Theor. Comput. Sci.2
2013 Patient information extraction in noisy tele-health texts
abstract
We explore methods for effectively extracting information from clinical narratives, which are captured in a public health consulting phone service called HealthLink. The currently available data consists of dialogues constructed by nurses while consulting patients on the phone. Since the data are interviews transcribed by nurses during phone conversations, they include a significant volume and variety of noise: First is explicit noise, which includes spelling errors, unfinished sentences, omission of sentence delimiters, variants of terms, etc. Second is implicit noise, which includes non-patient's information and negation of patient's information. To filter explicit noise, we propose our biomedical term detection/normalization method: it resolves misspelling, term variations, and arbitrary abbreviation of terms by nurses. In detecting temporal terms and other types of named entities (which show patients' personal information such as age, and sex), we propose a bootstrapping-based pattern learning to detect all kinds of arbitrary variations of the named entities. To address implicit noise, we propose a dependency path-based filtering method. The result of our denoising is the extraction of normalized patient information. The experimental results show that we achieve reasonable performance with our noise reduction methods.
Mi-Young Kim, Ying Xu 0003, Osmar R. Zaïane, Randy Goebel
BIBM4
2013 Approximation Algorithms for the Maximum Multiple RNA Interaction Problem
Weitian Tong, Randy Goebel, Tian Liu 0001, Guohui Lin
COCOA2
2013 Approximating the Minimum Independent Dominating Set in Perturbed Graphs
Weitian Tong, Randy Goebel, Guohui Lin
COCOON2
2013 The Role of Direct Manipulation of Visualizations in the Development and Use of Multi-level Knowledge Models
abstract
The proliferation of touch sensitive display screens has created a new generation of human-computer interaction styles which are so natural and common that even the youngest of users now perceive ordinary static media like a glossy magazine as a broken iPad. The volume of users who expect to be able to pinch, grab, twist and manipulate images on screen is rapidly growing; they drive a renewed interest in developing, assessing, and delivering new direct manipulation systems. Our premise is that one can exploit new technologies to develop new repertoires of direct manipulation, but with increasing pressure to provide semantically-coupled direct manipulation methods to experiment with computational information models. We develop this premise by noting highlights in the evolution of direct manipulation interfaces, and suggest that their selection and deployment can be tailored as visual experiments to debug and extend more complex computational models of information systems and processes. These systems and processes include those of natural systems such as arise in systems biology (e.g., modelling multiple levels of protein structure), but also in "unnatural" systems such as in the identification of hubs and authorities in artificial systems like the World Wide Web (WWW). The immediate consequence of our premise suggests that the design of direct manipulation tools should proceed with the semantics of the modelled systems in mind, so that each users' manipulations provide a new perspective on the concept of "data mining" of large data sets. This will allow users to not just expose implicit relationships, but to incrementally combine explanatory and exploratory investigation by direct manipulation, to adjust and improve the computational knowledge models that emerge from the underlying data.
Randy Goebel, Yuzuru Tanaka
IV1
2013 Open Information Extraction with Tree Kernels
Ying Xu 0003, Mi-Young Kim, Kevin Quinn 0002, Randy Goebel, Denilson Barbosa 0001
HLT-NAACL4
2012 Visualizing Community Centric Network Layouts
abstract
We present our COMmunity Boundary (COMB) and COMmunity Circles (COMC) network layout algorithms that focus on revealing the structure of discovered communities and the relationships between these communities. We believe this information is vital when developing new community mining algorithms as it allows the viewer to more quickly assess the quality of a mining result without appealing to large tables of statistics. To implement our algorithms we have introduced numerous modifications to the existing Fruchterman-Reingold layout, including support for multi-sized vertices, removal of the bounding frame, introduction of circular bounding boxes, and a novel slotting system. Our evaluation argues that both COMB and COMC outperform existing alternatives in their ability to reveal community structure and emphasize inter-community relations.
Justin Fagnan, Osmar R. Zaïane, Randy Goebel
IV3
2012 An improved approximation algorithm for the complementary maximal strip recovery problem
Guohui Lin, Randy Goebel, Lusheng Wang 0001
J. Comput. Syst. Sci.2
2012 Adaptive-capacity and robust natural language watermarking for agglutinative languages
abstract
ABSTRACT We present a robust and adaptive‐capacity watermarking algorithm for agglutinative languages. All processes, including the selection of sentences to be watermarked, watermark embedding, and watermark extraction, are based on syntactic dependency trees. We show that it is more robust to use syntactic dependency trees than the surface forms of sentences in text watermarking. For the agglutinative languages, we embed watermark using the two main characteristics of the languages. First, because a word consists of several morphemes, we can watermark sentences using morphological division/combination without deep linguistic analysis. Second, they permit relatively free word order, so we can move a syntactic constituent within its clause. Finally, to increase the information‐hiding capacity, we adaptively compute the number of watermark bits to be embedded for each sentence. We perform three kinds of evaluation: perceptibility, robustness, and capacity of our method. High capacity is achieved by dynamically determining possibly embedded watermark bits for each sentence. The secret rank based on a syntactic dependency tree strengthens robustness of our method. Finally, we show that the displacement of syntactic constituents and morphological division/combination does not affect the style and naturalness of the text. Copyright © 2011 John Wiley & Sons, Ltd.
Mi-Young Kim, Randy Goebel
Secur. Commun. Networks2
2011 What is Knowledge Visualization? Perspectives on an Emerging Discipline
abstract
This paper collates eight expert opinions about Knowledge Visualization, what it is and what it should be. An average of 581 words long, topics span from representation, storytelling and criticizing the lack of theory, to communication, analytics for the masses and reasoning, to trendy Visual Thinking and creativity beyond PowerPoint. These individual views provide a picture of the present and the future of a discipline that could not be more timely, aiming for a common understanding of the visualization of knowledge.
Stefan Bertschi, Sabrina Bresciani, Tom Crawford, Randy Goebel, Wolfgang Kienreich, Martin Lindner, Vedran Sabol, Andrew Vande Moere
IV4
2011 Strong Equivalence of Logic Programs with Abstract Constraint Atoms
Randy Goebel, Tomi Janhunen, Ilkka Niemelä, Jia-Huai You
LPNMR2
2011 The expansion continues: Stitching together the breadth of disciplines impinging on Artificial Intelligence
Randy Goebel, Mary-Anne Williams
Artif. Intell.1
2011 Size-constrained tree partitioning: Approximating the multicast k-tree routing problem
Zhipeng Cai 0001, Randy Goebel, Guohui Lin
Theor. Comput. Sci.2
2010 William J. Raynor Jr., International Dictionary of Artificial Intelligence (2nd edition), Global Professional Publishing (2009) ISBN 978-0-85297-657-9 242 pp
Randy Goebel
Artif. Intell.1
2010 The expanding breadth of artificial intelligence research
Randy Goebel, Mary-Anne Williams
Artif. Intell.1
2009 Local Community Identification in Social Networks
abstract
There has been much recent research on identifying global community structure in networks. However, most existing approaches require complete information of the graph in question, which is impractical for some networks, e.g. the World Wide Web (WWW). Algorithms for local community detection have been proposed but their results usually contain many outliers. In this paper, we propose a new measure of local community structure, coupled with a two-phase algorithm that extracts all possible candidates first, and then optimizes the community hierarchy. We compare our results with previous methods on real world networks such as the co-purchase network from Amazon. Experimental results verify the feasibility and effectiveness of our approach.
Jiyang Chen, Osmar R. Zaïane, Randy Goebel
ASONAM3
2009 A Visual Data Mining Approach to Find Overlapping Communities in Networks
abstract
Communities in social networks may overlap, with some hub nodes belonging to multiple communities. They may also have outliers, which are nodes that belong to no community. The criterion to locate hubs or outliers is network dependent. Previous methods usually require this information as input parameters, e.g., an expected number of communities, with no intuition or assistance. Here we present a visual data mining approach, which first helps the user to make appropriate parameter selections by observing initial data visualizations, and then finds and extracts overlapping community structures from the network. Experimental results verify the scalability and accuracy of our approach on real network data and show its advantages over previous methods.
Jiyang Chen, Osmar R. Zaïane, Randy Goebel
ASONAM3
2009 Size-Constrained Tree Partitioning: A Story on Approximation Algorithm Design for the Multicast k-Tree Routing Problem
Zhipeng Cai 0001, Randy Goebel, Guohui Lin
COCOA2
2009 Glen, Glenda or Glendale: Unsupervised and Semi-supervised Learning of English Noun Gender
Shane Bergsma, Dekang Lin, Randy Goebel
CoNLL3
2009 Web-Scale N-gram Models for Lexical Disambiguation
Shane Bergsma, Dekang Lin, Randy Goebel
IJCAI3
2009 A Continuum-Based Approach for Tightness Analysis of Chinese Semantic Units
Ying Xu 0003, Christoph Ringlstetter, Randy Goebel
PACLIC3
2009 Detecting Communities in Social Networks Using Max-Min Modularity
abstract
Many datasets can be described in the form of graphs or networks where nodes in the graph represent entities and edges represent relationships between pairs of entities. A common property of these networks is their community structure, considered as clusters of densely connected groups of vertices, with only sparser connections between groups. The identification of such communities relies on some notion of clustering or density measure. which defines the communities that can be found. However, previous community detection methods usually apply the same structural measure on all kinds of networks, despite their distinct dissimilar features. In this paper, we present a new community mining measure, Max-Min Modularity, which considers both connected pairs and criteria defined by domain experts in finding communities, and then specify a hierarchical clustering algorithm to detect communities in networks. When applied to real world networks for which the community structures are already known, our method shows improvement over previous algorithms. In addition, when applied to randomly generated networks for which we only have approximate information about communities, it gives promising results which shows the algorithm's robustness against noise.
Jiyang Chen, Osmar R. Zaïane, Randy Goebel
SDM3
2009 Most parsimonious haplotype allele sharing determination
abstract
BACKGROUND: The "common disease--common variant" hypothesis and genome-wide association studies have achieved numerous successes in the last three years, particularly in genetic mapping in human diseases. Nevertheless, the power of the association study methods are still low, in particular on quantitative traits, and the description of the full allelic spectrum is deemed still far from reach. Given increasing density of single nucleotide polymorphisms available and suggested by the block-like structure of the human genome, a popular and prosperous strategy is to use haplotypes to try to capture the correlation structure of SNPs in regions of little recombination. The key to the success of this strategy is thus the ability to unambiguously determine the haplotype allele sharing status among the members. The association studies based on haplotype sharing status would have significantly reduced degrees of freedom and be able to capture the combined effects of tightly linked causal variants. RESULTS: For pedigree genotype datasets of medium density of SNPs, we present two methods for haplotype allele sharing status determination among the pedigree members. Extensive simulation study showed that both methods performed nearly perfectly on breakpoint discovery, mutation haplotype allele discovery, and shared chromosomal region discovery. CONCLUSION: For pedigree genotype datasets, the haplotype allele sharing status among the members can be deterministically, efficiently, and accurately determined, even for very small pedigrees. Given their excellent performance, the presented haplotype allele sharing status determination programs can be useful in many downstream applications including haplotype based association studies.
Zhipeng Cai 0001, Hadi Sabaa, Randy Goebel, Jiaofen Xu, Paul Stothard, Guohui Lin
BMC Bioinform.4
2008 Distributional Identification of Non-Referential Pronouns
Shane Bergsma, Dekang Lin, Randy Goebel
ACL3
2008 Discriminative Learning of Selectional Preference from Unlabeled Text
Shane Bergsma, Dekang Lin, Randy Goebel
EMNLP3
2008 Self-Tutoring, Teaching and Testing: An Intelligent Process Analyzer
abstract
Mastery of basic concepts and logical units is a prerequisite for solving complex problems. The divide-and-conquer strategy has been used successfully in a variety of problem solving situations, including algorithm design and software engineering. In this paper, we adopt a similar approach and propose a process analyzer in education. The goal is to help students improve their problem solving skills, as well as to assist teachers to monitor student response so that guidance can be provided as required. Complex and tedious processes are often encountered in curricula such as physics, chemistry, and mathematics, where the final answer is built upon the results of numerous smaller processes. Therefore the process analyzer defines the top-level process as a hierarchy composed of smaller integral parts so that students who are not able to solve the problem on their own are able to look at lower-level simpler processes and follow the hints leading to the correct answer. Question designers can define hints, as well as the score in each step. Responses are recorded for modeling student performance. The students interact with the Process Analyzer through a graphical interface, which provides an engaging and motivating environment. The positive feedback on our process analyzer shows the feasibility of our approach. We describe the design and implementation of the process analyzer, and present our future plan.
Irene Cheng 0001, Nathaniel Rossol, Randy Goebel
ICALT3
2008 Targeting Chinese Nominal Compounds in Corpora
Weiruo Qu, Christoph Ringlstetter, Randy Goebel
LREC3
2008 An Unsupervised Approach to Cluster Web Search Results Based on Word Sense Communities
abstract
Effectively organizing web search results into clusters is important to facilitate quick user navigation to relevant documents. Previous methods may rely on a training process and do not provide a measure for whether page clustering is actually required. In this paper, we reformalize the clustering problem as a word sense discovery problem. Given a query and a list of result pages, our unsupervised method detects word sense communities in the extracted keyword network. The documents are assigned to several refined word sense communities to form clusters. We use the modularity score of the discovered keyword community structure to measure page clustering necessity. Experimental results verify our method's feasibility and effectiveness.
Jiyang Chen, Osmar R. Zaïane, Randy Goebel
Web Intelligence3
2008 Identifying a few foot-and-mouth disease virus signature nucleotide strings for computational genotyping
abstract
BACKGROUND: Serotypes of the Foot-and-Mouth disease viruses (FMDVs) were generally determined by biological experiments. The computational genotyping is not well studied even with the availability of whole viral genomes, due to uneven evolution among genes as well as frequent genetic recombination. Naively using sequence comparison for genotyping is only able to achieve a limited extent of success. RESULTS: We used 129 FMDV strains with known serotype as training strains to select as many as 140 most serotype-specific nucleotide strings. We then constructed a linear-kernel Support Vector Machine classifier using these 140 strings. Under the leave-one-out cross validation scheme, this classifier was able to assign correct serotype to 127 of these 129 strains, achieving 98.45% accuracy. It also assigned serotype correctly to an independent test set of 83 other FMDV strains downloaded separately from NCBI GenBank. CONCLUSION: Computational genotyping is much faster and much cheaper than the wet-lab based biological experiments, upon the availability of the detailed molecular sequences. The high accuracy of our proposed method suggests the potential of utilizing a few signature nucleotide strings instead of whole genomes to determine the serotypes of novel FMDV strains.
Guohui Lin, Zhipeng Cai 0001, Xiu-Feng Wan, Lizhe Xu, Randy Goebel
BMC Bioinform.6
2007 Selecting Genes with Dissimilar Discrimination Strength for Sample Class Prediction
Zhipeng Cai 0001, Randy Goebel, Mohammad R. Salavatipour, Yi Shi 0005, Lizhe Xu, Guohui Lin
APBC2
2007 Iterated Belief Contraction from First Principles
Abhaya C. Nayak, Randy Goebel, Mehmet A. Orgun
IJCAI2
2007 Visualizing Web Navigation Data with Polygon Graphs
abstract
As the volume of digitally accessible information grows, there is increasing pressure on the development of data visualization methods to enable humans to interpret that data. We provide a description of our WebViz system, as a tool to visualize both the structure and usage of web sites. We illustrate the use of our visualization paradigm by introducing polygonal graphs layered on top of our adaptation of radial disk trees. In our system, the structure of a web segment is rendered as a radial tree, and usage data can be extracted and layered as polygonal graphs. By interactively creating and adjusting these layers, a user can develop real time insight into the data. We present the system, show the idea of interactive visual operators, and provide some examples that help show the value of the specific visualization techniques, as well as the interactive use of those techniques.
Jiyang Chen, Tong Zheng 0006, William Thorne, Daniel Huntley, Osmar R. Zaïane, Randy Goebel
IV6
2007 Visual Data Mining of Web Navigational Data
abstract
Discovering web navigational trends and understanding data mining results is undeniably advantageous to web designers and web-based application builders. It is also desirable to interactively investigate web access data and patterns, to allows ad-hoc discovery and examination of patterns that are not apriori known. Visualizing the usage data in the context of the web site structure is of major importance, as it puts web access requests and their connectivity in perspective. Various visualization tools have been developed for this task, but often fail to provide visual data mining functionalities to generate new patterns. Here we present our visual data mining system, WebViz, which allows interactive investigation of web usage data within their structure context, as well as ad-hoc knowledge pattern discovery on web navigational behaviour.
Jiyang Chen, Tong Zheng 0006, William Thorne, Osmar R. Zaïane, Randy Goebel
IV5
2007 Nucleotide composition string selection in HIV-1 subtyping using whole genomes
abstract
MOTIVATION: The availability of the whole genomic sequences of HIV-1 viruses provides an excellent resource for studying the HIV-1 phylogenies using all the genetic materials. However, such huge volumes of data create computational challenges in both memory consumption and CPU usage. RESULTS: We propose the complete composition vector representation for an HIV-1 strain, and a string scoring method to extract the nucleotide composition strings that contain the richest evolutionary information for phylogenetic analysis. In this way, a large-scale whole genome phylogenetic analysis for thousands of strains can be done both efficiently and effectively. By using 42 carefully curated strains as references, we apply our method to subtype 1156 HIV-1 strains (10.5 million nucleotides in total), which include 825 pure subtype strains and 331 recombinants. Our results show that our nucleotide composition string selection scheme is computationally efficient, and is able to define both pure subtypes and recombinant forms for HIV-1 strains using the 5000 top ranked nucleotide strings. AVAILABILITY: The Java executable and the HIV-1 datasets are accessible through 'http://www.cs.ualberta.ca/~ghlin/src/WebTools/hiv.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaomeng Wu, Zhipeng Cai 0001, Xiu-Feng Wan, Tin Hoang, Randy Goebel, Guohui Lin
Bioinform.5
2007 Selecting dissimilar genes for multi-class classification, an application in cancer subtyping
abstract
BACKGROUND: Gene expression microarray is a powerful technology for genetic profiling diseases and their associated treatments. Such a process involves a key step of biomarker identification, which are expected to be closely related to the disease. A most important task of these identified genes is that they can be used to construct a classifier which can effectively diagnose disease and even recognize the disease subtypes. Binary classification, for example, diseased or healthy, in microarray data analysis has been successful, while multi-class classification, such as cancer subtyping, remains challenging. RESULTS: We target on the challenging multi-class classification in microarray data analysis, especially on the cancer subtyping using gene expression microarray. We present a novel class discrimination strength vector to represent individual genes and introduce a new measurement to quantify the class discrimination strength difference between two genes. Such a new distance measure is employed in gene clustering, and subsequently the gene cluster information is exploited to select a set of genes which can be used to construct a sample classifier. We tested our method on four real cancer microarray datasets each contains multiple subtypes of cancer patients. The experimental results show that the constructed classifiers all achieved a higher classification accuracy than the previously best classification results obtained on these four datasets. Additional tests show that the selected genes by our method are less correlated and they all contribute statistically significantly to the more accurate cancer subtyping. CONCLUSION: The proposed novel class discrimination strength vector is a better representation than the gene expression vector, in the sense that it can be used to effectively eliminate highly correlated but redundant genes for classifier construction. Such a method can build a classifier to achieve a higher classification accuracy, which is demonstrated via cancer subtyping.
Zhipeng Cai 0001, Randy Goebel, Mohammad R. Salavatipour, Guohui Lin
BMC Bioinform.2
2006 Using Gene Clustering to Identify Discriminatory Genes with Higher Classification Accuracy
abstract
A single DNA microarray measures thousands to tens of thousands of gene expression levels, but experimental datasets normally consist of much fewer such arrays, typically in tens to hundreds, taken over a selection of tissue samples. The biological interpretation of these data relies on identifying subsets of induced or repressed genes that can be used to discriminate various categories of tissue, to provide experimental evidence for connections between a subset of genes and the tissue pathology. A variety of methods can be used to identify discriminatory gene subsets, which can be ranked by classification accuracy. But the high dimensionality of the gene expression space, coupled with relatively fewer tissue samples, creates the dimensionality problem: gene subsets that are too large to provide convincing evidence for any plausible causal connection between that gene subset and the tissue pathology. We propose a new gene selection method, clustered gene selection (CGS) which, when coupled with existing methods, can identify gene subsets that overcome the dimensionality problem and improve classification accuracy. Experiments on eight real datasets showed that CGS can identify many more cancer related genes and clearly improve classification accuracy, compared with three other non-CGS based gene selection methods
Zhipeng Cai 0001, Lizhe Xu, Yi Shi 0005, Mohammad R. Salavatipour, Randy Goebel, Guohui Lin
BIBE5
2006 A Model-Free Greedy Gene Selection for Microarray Sample Class Prediction
abstract
Microarray data analysis is notoriously challenging as it involves a huge number of genes compared to only a limited number of samples. Gene selection, to detect the most significantly differentially expressed genes under different categories of conditions, is both computationally and biologically interesting, and has become a central research focus in all studies that use gene expression microarray technology. Despite many existing efforts, better gene selection methods that can effectively identify biologically significant biomarkers, yet computationally efficient, are still in need. In this paper, a model-free greedy (MFG) gene selection method is proposed, which implements several intuitive heuristics but doesn't assume any statistical distribution on the expression data. The experimental results on three real microarray datasets showed that the MFG method combined with a support vector machine (SVM) classifier or a k-nearest neighbor (KNN) classifier is efficient and robust in identifying discriminatory genes
Yi Shi 0005, Zhipeng Cai 0001, Lizhe Xu, Randy Goebel, Guohui Lin
CIBCB5
2006 Perceptually Enhanced Multimedia Processing, Visualization and Transmission
abstract
Data reduction has long been a method for adaptation to limited computational and network resources. But one major concern is the tradeoff between preserving visual quality and reducing data size. Furthermore, the presence of multi-modal data, e.g. visual and aural, is common and therefore distributing competing resources among multi-modal data to achieve optimal visual quality becomes a major. Since humans are typically the penultimate viewer of multimedia data, it is reasonable to take human perception into consideration during the data reduction and resource distribution process, in order to estimate and control the resulting visual quality. Psychophysical experiments reported in the literature have shown that better performance can be achieved in multimedia processing, visualization and transmission, by incorporating perceptual factors. This paper gives an overview of how human perception plays a role in the development of multimedia applications, so as to inspire and inform future research in this direction
Irene Cheng 0001, Randy Goebel
ISM2
2006 Taking Levi Identity Seriously: A Plea for Iterated Belief Contraction
Abhaya C. Nayak, Randy Goebel, Mehmet A. Orgun, Tam Pham
KSEM2
2004 Visualizing and Discovering Web Navigational Patterns
abstract
Web site structures are complex to analyze. Cross-referencing the web structure with navigational behaviour adds to the complexity of the analysis. However, this convoluted analysis is necessary to discover useful patterns and understand the navigational behaviour of web site visitors, whether to improve web site structures, provide intelligent on-line tools or offer support to human decision makers. Moreover, interactive investigation of web access logs is often desired since it allows ad hoc discovery and examination of patterns not a priori known. Various visualization tools have been provided for this task but they often lack the functionality to conveniently generate new patterns. In this paper we propose a visualization tool to visualize web graphs, representations of web structure overlaid with information and pattern tiers. We also propose a web graph algebra to manipulate and combine web graphs and their layers in order to discover new patterns in an ad hoc manner.
Jiyang Chen, Lisheng Sun, Osmar R. Zaïane, Randy Goebel
WebDB4
2004 Iterated Belief Change
abstract
Most existing formalizations treat belief change as a single‐step process, and ignore several problems that become important when a theory, or belief state, is revised over several steps. This paper identifies these problems, and argues for the need to retain all of the multiple possible outcomes of a belief change step, and for a framework in which the effects of a belief change step persist as long as is consistently possible. To demonstrate that such a formalization is indeed possible, we develop a framework, which uses the language of PJ‐default logic (Delgrande and Jackson 1991) to represent a belief state, and which enables the effects of a belief change step to persist by propagating belief constraints. Belief change in this framework maps one belief state to another, where each belief state is a collection of theories given by the set of extensions of the PJ‐default theory representing that belief state. Belief constraints do not need to be separately recorded; they are encoded as clearly identifiable components of a PJ‐default theory. The framework meets the requirements for iterated belief change that we identify and satisfies most of the AGM postulates (Alchourrón, Gärdenfors, and Makinson 1985) as well.
Aditya Ghose, Pablo O. Hadjinian, Abdul Sattar 0001, Jia-Huai You, Randy Goebel
Comput. Intell.5
2003 WebKIV: Visualizing Structure and Navigation forWeb Mining Applications
abstract
A significant part of the Web mining problem is simply in understanding the value of any mining method. For example, the value of Web mining to improve user navigation is even more challenging if one can't visualize the differences over a large collection of Web pages or a significant structure within the existing Web. We present WebKIV, a tool we've developed to help us visualize our own results in Web mining. WebKIV combines strategies from several other Web visualization tools, to provide a single method of visualizing Web structure, and the results of Web mining on that structure. We summarize the value of Web visualization tools along the dimensions of scale (can one visualize small and large structures), navigation dynamics (can one visualize navigation dynamically or statically), and cumulative usage (can one distinguish individual and aggregate Web usage). We then show how WebKIV provides a way of visualizing the results of Web mining in a way that distinguishes properties along all three of these dimensions.
Yonghe Niu, Tong Zheng 0006, Jiyang Chen, Randy Goebel
Web Intelligence4
2002 WebFrame: In Pursuit of Computationally and Cognitively Efficient Web Mining
Tong Zheng 0006, Yonghe Niu, Randy Goebel
PAKDD3
2001 Towards a Novel OLAP Interface for Distributed Data Warehouses
Ayman Ammoura, Osmar R. Zaïane, Randy Goebel
DaWaK3
2000 Knowledge Representation, Belief Revision, and the Challenge of Optimality
Randy Goebel
PRICAI1
1999 Connections Between Default Reasoning and Partial Constraint Satisfaction
Aditya Ghose, Grigoris Antoniou, Randy Goebel, Abdul Sattar 0001
Inf. Sci.3
1998 Belief States as Default Theories: Studies in Non-Prioritized Belief Change
Aditya Ghose, Randy Goebel
ECAI2
1997 An Abductive Semantics for Disjunctive Logic Programs and Its Proof Procedure
Jia-Huai You, Li-Yan Yuan, Randy Goebel
FSTTCS3
1996 Anytime Default Inference
Aditya Ghose, Randy Goebel
PRICAI2
1991 Meta-reasoning: An Incremental Compilation Approach
abstract
An incremental compilation approach to meta-reasoning is presented together with a method to update dynamically changing knowledge bases. The compilation process translates meta-level specification of facts and hypotheses into sentences of clausal logic. It then incrementally computes inconsistent sets of instances of hypotheses and records potential crucial literals. The extra information computed during compilation enables the theorem prover to avoid redundant computations and to efficiently update the compiled knowledge. Whenever a new fact is learned the effects of the fact are computed incrementally, without recompiling. A relationship between potential crucial literals and Reiter and de Kleer's prime implicants shows that this approach may be useful in incrementally computing and maintaining the prime implicants, as well.>
Abdul Sattar 0001, Randy Goebel
ICDE2
1991 A Message Passing Algorithm for Plan Recognition
Dekang Lin, Randy Goebel
IJCAI2
1991 Using crucial literals to select better theories
abstract
When Horn clause theories are combined with integrity constraints to produce potentially refutable theories, Seki and Takeuchi have shown how crucial literals can be used to discriminate two mutually incompatible theories. A literal is crucial with respect to two theories if only one of the two theories supports the derivation of that literal. In other words, actually determining the truth value of the crucial literal will refute one of the two incompatible theories. This paper presents an integration of the idea of crucial literal with Theorist, a logic‐based system for hypothetical reasoning. Theorist is a goal‐directed nonmonotonic reasoning system that classifies logical formulas as possible hypotheses, facts, and observations. As Theorist uses full clausal logic, it does not require Seki and Takeuchi's notion of integrity constraint to define refutable theories. In attempting to deduce observation sentences, Theorist identifies instances of possible hypotheses as nomological explanations: consistent sets of hypothesis instances required to deduce observations. As multiple and mutually incompatible explanations are possible, the notion of crucial literal provides the basis for proposing experiments that distinguish competing explanations. We attempt to make three contributions. First, we adapt Seki and Takeuchi's method for Theorist. To do so, we incrementally use crucial literals as experiments, whose results are used to reduce the total number of explanations generated for a given set of observations. Next, we specify an extension which incrementally constructs a table of all possible crucial literals for any pair of theories. This extension is more efficient and provides the user with greater opportunity to conduct experiments to eliminate falsifiable theories. A prototype is implemented in CProlog, and several examples of diagnosis are considered to show its empirical efficiency. Finally, we point out that assumption‐based truth maintenance systems (ATMS), as used in the multiple fault diagnosis system of de Kleer and Williams, are interesting special cases of this more general method of distinguishing explanatory theories.
Abdul Sattar 0001, Randy Goebel
Comput. Intell.2
1990 Integrating probabilistic, taxonomic and causal knowledge in abductive diagnosis
Dekang Lin, Randy Goebel
UAI2
1988 Exhuming the criticism of the logicist
abstract
McDermott has recently explained his fundamental philosophical shift on the methodology of artificial intelligence (AI) and has further suggested that the shift is both necessary and inevitable. The shift results from a perception that a trend towards overformalisation has detached the real problems from the research results. McDermott's criticism is an enlightened exhumation of the criticisms of the seventies and explains new ways in which the logical methodology can be abused. I argue that McDermott's criticism should not discourage the use of logic, but force a timely reexamination of its fundamental role in AI.
Randy Goebel
Comput. Intell.1
1986 Using Definite Clauses and Integrity Constraints as the Basis for a Theory Formation Approach to Diagnostic Reasoning
Randy Goebel, Koichi Furukawa, David Poole 0001
ICLP1
1986 Gracefully adding negation and disjunction to Prolog
David Poole 0001, Randy Goebel
ICLP2
1985 Interpreting Descriptions in a Prolog-based Knowledge Representation System
Randy Goebel
IJCAI1