VLDB 2026 Research / reviewers in the wild / expert
Xumin Liu
dblp:61/5010
· DBLP profile ↗
52ranked-venue papers
24as first author
9since 2021 · last 2026
0000-0002-6109-4851ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 11 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 12 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Teaching Data Science without the Programming Barrier: Design and Evaluation of an Integrated Learning Platform
Xumin Liu, Erik Golen |
ITiCSE (1) | 1 |
| 2025 | Lung Cancer Detection Using Fine-Tuned Data-Efficient Models
Jared Nobles, Xumin Liu |
IEEE Big Data | 2 |
| 2024 | Balancing Feature Similarity and Label Variability for Optimal Size-Aware One-shot Subset SelectionabstractSubset or core-set selection offers a data-efficient way for training deep learning models. One-shot subset selection poses additional challenges as subset selection is only performed once and full set data become unavailable after the selection. However, most existing methods tend to choose either diverse or difficult data samples, which fail to faithfully represent the joint data distribution that is comprised of both feature and label information. The selection is also performed independently from the subset size, which plays an essential role in choosing what types of samples. To address this critical gap, we propose to conduct Feature similarity and Label variability Balanced One-shot Subset Selection (BOSS), aiming to construct an optimal size-aware subset for data-efficient deep learning. We show that a novel balanced core-set loss bound theoretically justifies the need to simultaneously consider both diversity and difficulty to form an optimal subset. It also reveals how the subset size influences the bound. We further connect the inaccessible bound to a practical surrogate target which is tailored to subset sizes and varying levels of overall difficulty. We design a novel Beta-scoring importance function to delicately control the optimal balance of diversity and difficulty. Comprehensive experiments conducted on both synthetic and real data justify the important theoretical properties and demonstrate the superior performance of BOSS as compared with the competitive baselines. Abhinab Acharya, Dayou Yu, Qi Yu 0001, Xumin Liu |
ICML | 4 |
| 2023 | A Web-Based Learning Platform for Teaching Data Science to Non-Computer MajorsabstractA web-based learning platform is useful as it allows students with limited or no programming background to conduct in-depth hands-on practice in data science. Background: The need for data science coursework for non-computing majors has grown in recent years, given the demand in various disciplines. However, a substantial number of current data science courses are inappropriate for non-computing majors as they typically require a long chain of prerequisite courses in computer science and mathematics. Moreover, courses designed for computing majors do not match the preparation and interests of students majoring in other disciplines. Outcomes: This paper presents a platform for Learning Data Science (DSLP), a web-based platform, which assists in the teaching and learning of data science topics by students with limited or no coding experience, including those that have completed a high school AP Computer Science Principles (CSP) class or an equivalent CSP course increasingly offered in many colleges. Application Design: The platform helps students understand fundamental data science concepts and techniques, as well as provides them with an in-depth hands-on experience that goes beyond their coding capabilities. The platform offers various data visualization supports to help students understand data and analysis results. Students can use the platform to work on in-house datasets or their own data. This allows students to focus more on how to solve data science problems in various domains than how to write code. The platform also has several unique features that make it particularly helpful for teaching and learning data science topics such as code exemplification and sandbox, informative instructions, and progress monitoring. Findings: The platform has been used multiple times in data science courses for non-computing majors offered at the authors' institution. Preliminary student feedback indicated that the platform is effective in terms of improving student understanding and interest in the topics. Xumin Liu, Erik Golen, Rajendra K. Raj, Kimberly Fluet |
FIE | 1 |
| 2022 | DSLP: A Web-based Data Science Learning Platform to Support DS Education for Non-Computing MajorsabstractThe presenters will demo a web-based Data Science Learning Platform (DSLP) that makes data science education accessible to students with limited or no programming background. The DSLP platform offers students with several benefits such as: (1) learn a web-based user interface to perform data science tasks without requiring coding, (2) explore popular Python data science libraries (e.g., Pandas, Matplotlib, Numpy, or Scikit-Learn) through real-time code exemplification to prepare them for advanced data science topics, (3) become familiar with the on-site user guide and helpful tips to make the platform easy to use, (4) write their own code within a sandbox, and (5) monitor their own progress by tracking their platform usage. The demo will walk through the steps of using the DSLP to perform various data science tasks and the participants will be able to try out the features mentioned above. The demo will also cover the design of course materials, including hands-on practices and lab assignments using the DSLP platform. The typical participants include instructors who are interested in teaching introductory-level data science to high school students or non-computing college majors with little or no programming background. Participants need to have a laptop with access to the Internet to attend the hands-on exercises workshop. The laptop should have a current web browser (e.g., Safari or Chrome) installed to access the web-based learning platform. This demo describes work supported by the National Science Foundation under Award 2021287. Xumin Liu, Erik Golen, Rajendra K. Raj |
SIGCSE (2) | 1 |
| 2022 | Introducing Data Science Topics to Non-Computing MajorsabstractData science knowledge and skills have become indispensable to STEM and non-STEM disciplines alike. As a result, it has become crucial for students in non-computing majors to learn data science techniques, particularly in the context of their own disciplines. A majority of current university data science coursework, however, requires sufficient depth in programming and statistical skills related to managing, manipulating, and analyzing data, which reduces their usefulness for entry-level non-computing majors. This workshop presents a set of hands-on exercises to introduce data science to entry-level non-computing majors. The exercises cover the data science lifecycle, including data acquisition, preparation, model development and deployment, visualization, and storytelling. A freely-available web-based Data Science Learning Platform (DSLP) will be presented to show how to perform hands-on data science exercises with little or no coding background. The presenters will also share their experiences in using the DSLP tool in an entry-level data science course to non-computing majors at RIT. Both the tool and course materials will be shared with workshop participants. The typical workshop participant is a high school teacher or a college instructor interested in teaching data science at the introductory level. No prior programming or data science experience is needed, thus making the workshop materials usable by a wide audience. Participants need to have a laptop with access to the Internet to attend the hands-on exercises workshop. The laptop should have a current web browser (e.g., Safari or Chrome) installed to access the web-based learning platform. This work was supported by the National Science Foundation under Award 2021287. Xumin Liu, Erik Golen, Rajendra K. Raj |
SIGCSE (2) | 1 |
| 2022 | Hierarchical Bayesian multi-kernel learning for integrated classification and summarization of app reviewsabstractApp stores enable users to share their experiences directly with the developers in the form of app reviews. Recent studies have shown that the feedback received from users is a valuable source of information for requirements extraction, which encourages app developers to leverage the reviews for app update and maintenance purposes. Follow-up studies proposed automated techniques to help developers filter the large volume of daily and noisy reviews and/or summarize their content. However, all previous studies approached the app reviews classification and summarization as separate tasks, which complicated the process and introduced unnecessary overhead. Moreover, none of those approaches explored the potential of utilizing the hierarchical relationships that exist between the labels of app reviews for the purpose of building a more accurate model. In this work, we propose Hierarchical Multi-Kernel Relevance Vector Machines (HMK-RVM), a Bayesian multi-kernel technique that integrates app review classification and summarization using a unified model. Moreover, it can provide insights into the learned patterns and underlying data for easier model interpretation. We evaluated our proposed approach on two real-world datasets and showed that in addition to the gained insights, the model produces equal or better results than the state of the art. Moayad Alshangiti, Weishi Shi, Eduardo Lima, Xumin Liu, Qi Yu 0001 |
ESEC/SIGSOFT FSE | 4 |
| 2021 | Deep Reinforced Attention Regression for Partial Sketch Based Image RetrievalabstractFine-Grained Sketch-Based Image Retrieval (FG-SBIR) aims at finding a specific image from a large gallery given a query sketch. Despite the widespread applicability of FG-SBIR in many critical domains (e.g., crime activity tracking), existing approaches still suffer from a low accuracy while being sensitive to external noises such as unnecessary strokes in the sketch. The retrieval performance will further deteriorate under a more practical on-the-fly setting, where only a partially complete sketch with only a few (noisy) strokes are available to retrieve corresponding images. We propose a novel framework that leverages a uniquely designed deep reinforcement learning model that performs a dual-level exploration to deal with partial sketch training and attention region selection. By enforcing the model’s attention on the important regions of the original sketches, it remains robust to unnecessary stroke noises and improve the retrieval accuracy by a large margin. To sufficiently explore partial sketches and locate the important regions to attend, the model performs bootstrapped policy gradient for global exploration while adjusting a standard deviation term that governs a locator network for local exploration. The training process is guided by a hybrid loss that integrates a reinforcement loss and a supervised loss. A dynamic ranking reward is developed to fit the on-the-fly image retrieval process using partial sketches. The extensive experimentation performed on three public datasets shows that our proposed approach achieves the state-of-the-art performance on partial sketch based image retrieval. Dingrong Wang, Hitesh Sapkota, Xumin Liu, Qi Yu 0001 |
ICDM | 3 |
| 2021 | A Structure Alignment Deep Graph Model for Mashup Recommendation
Eduardo Lima, Xumin Liu |
ICSOC | 2 |
| 2020 | A Bayesian learning model for design-phase service mashup popularity prediction
Moayad Alshangiti, Weishi Shi, Xumin Liu, Qi Yu 0001 |
Expert Syst. Appl. | 3 |
| 2019 | Why is Developing Machine Learning Applications Challenging? A Study on Stack Overflow PostsabstractBackground: As smart and automated applications pervade our lives, an increasing number of software developers are required to incorporate machine learning (ML) techniques into application development. However, acquiring the ML skill set can be nontrivial for software developers owing to both the breadth and depth of the ML domain. Aims: We seek to understand the challenges developers face in the process of ML application development and offer insights to simplify the process. Despite its importance, there has been little research on this topic. A few existing studies on development challenges with ML are outdated, small scale, or they do no involve a representative set of developers. Method: We conduct an empirical study of ML-related developer posts on Stack Overflow. We perform in-depth quantitative and qualitative analyses focusing on a series of research questions related to the challenges of developing ML applications and the directions to address them. Results: Our findings include: (1) ML questions suffer from a much higher percentage of unanswered questions on Stack Overflow than other domains; (2) there is a lack of ML experts in the Stack Overflow QA community; (3) the data preprocessing and model deployment phases are where most of the challenges lay; and (4) addressing most of these challenges require more ML implementation knowledge than ML conceptual knowledge. Conclusions: Our findings suggest that most challenges are under the data preparation and model deployment phases, i.e., early and late stages. Also, the implementation aspect of ML shows much higher difficulty level among developers than the conceptual aspect. Moayad Alshangiti, Hitesh Sapkota, Pradeep K. Murukannaiah, Xumin Liu, Qi Yu 0001 |
ESEM | 4 |
| 2019 | Integrating Multi-level Tag Recommendation with External Knowledge Bases for Automatic Question AnsweringabstractWe focus on using natural language unstructured textual Knowledge Bases (KBs) to answer questions from community-based Question-and-Answer (Q8A) websites. We propose a novel framework that integrates multi-level tag recommendation with external KBs to retrieve the most relevant KB articles to answer user posted questions. Different from many existing efforts that primarily rely on the Q8A sites’ own historical data (e.g., user answers), retrieving answers from authoritative external KBs (e.g., online programming documentation repositories) has the potential to provide rich information to help users better understand the problem, acquire the knowledge, and hence avoid asking similar questions in future. The proposed multi-level tag recommendation best leverages the rich tag information by first categorizing them into different semantic levels based on their usage frequencies. A post-tag co-clustering model, augmented by a two-step tag recommender, is used to predict tags at different levels for a given user posted question. A KB article retrieval component leverages the recommended multi-level tags to select the appropriate KBs and search/rank the matching articles thereof. We conduct extensive experiments using real-world data from a Q8A site and multiple external KBs to demonstrate the effectiveness of the proposed question-answering framework. Eduardo Lima, Weishi Shi, Xumin Liu, Qi Yu 0001 |
ACM Trans. Internet Techn. | 3 |
| 2018 | Log sequence clustering for workflow mining in multi-workflow systems
Xumin Liu, Moayad Alshangiti, Chen Ding 0004, Qi Yu 0001 |
Data Knowl. Eng. | 1 |
| 2018 | A Web service search engine for large-scale Web service discovery based on the probabilistic topic modeling and clustering
Afnan Bukhari, Xumin Liu |
Serv. Oriented Comput. Appl. | 2 |
| 2017 | Recommending Services for New Mashups through Service Factors and Top-K NeighborsabstractOne of the most interesting research directions in service computing is to leverage current recommendation system solutions to suggest web services for a mashup application. Existing approaches are mainly based on collaborative filtering techniques, which can suffer from the heavy rely on human input, data sparsity and cold start issues, resulting in low accuracy. In this paper, we leverage advanced probabilistic model based approaches to tackle these issues. Our solution is to make service recommendation based on the service features and historical usage. We use the Hierarchical Dirichlet Process (HDP), a nonparametric Bayesian approach to intelligently discover the functionally relevant services based on their specifications. We leverage Probabilistic Matrix Factorization (PMF) to recommend services based on historical usage and tackle the cold start issues for new mashups through their top-K neighbors. We integrate the suggesting results from these two approaches through the Bayesian theorem and take the indicator of quality of service into account to make the final suggestion. We compared our approach with some existing approaches using a real world data set and the result indicates that our approach performs the best. Priyanka Samanta, Xumin Liu |
ICWS | 2 |
| 2017 | Correlation-Aware Multi-Label Active Learning for Web Service Tag RecommendationabstractTag recommendation has gained significant popularity for annotating various web-based resources including web services. Compared with other approaches, tag recommendation based on supervised learning models usually lead to good accuracy. However, a high-quality training data set is needed, which demands manual tagging efforts from domain experts. While we could leverage the tags of existing web services assigned by their developers, the quality of these tags may not be good enough to build accurate classifiers for tag recommendation. In this paper, a novel multi-label active learning approach is proposed for web service tag recommendation. The proposed approach is able to identify a small number of most informative web services to be tagged by domain experts. We further minimize the domain expert efforts by learning and leveraging the correlations among tags to improve the active learning process. We conduct a comprehensive experimental study on a real-world data set and results demonstrate the effectiveness of our approach. Weishi Shi, Xumin Liu, Qi Yu 0001 |
ICWS | 2 |
| 2017 | Statistical Learning of Domain-Specific Quality-of-Service Features from User ReviewsabstractWith the fast increase of online services of all kinds, users start to care more about the Quality of Service (QoS) that a service provider can offer besides the functionalities of the services. As a result, QoS-based service selection and recommendation have received significant attention since the mid-2000s. However, existing approaches primarily consider a small number of standard QoS parameters, most of which relate to the response time, fee, availability of services, and so on. As online services start to diversify significantly over different domains, these small set of QoS parameters will not be able to capture the different quality aspects that users truly care about over different domains. Most existing approaches for QoS data collection depend on the information from service providers, which are sensitive to the trustworthiness of the providers. Some service monitoring mechanisms collect QoS data through actual service invocations but may be affected by actual hardware/software configurations. In either case, domain-specific QoS data that capture what users truly care about have not been successfully collected or analyzed by existing works in service computing. To address this demanding issue, we develop a statistical learning approach to extract domain-specific QoS features from user-provided service reviews. In particular, we aim to classify user reviews based on their sentiment orientations into either a positive or negative category. Meanwhile, statistical feature selection is performed to identify statistically nontrivial terms from review text, which can serve as candidate QoS features. We also develop a topic models-based approach that automatically groups relevant terms and returns the term groups to users, where each term group corresponds to one high-level quality aspect of services. We have conducted extensive experiments on three real-world datasets to demonstrates the effectiveness of our approach. Xumin Liu, Weishi Shi, Arpeet Kale, Chen Ding 0004, Qi Yu 0001 |
ACM Trans. Internet Techn. | 1 |
| 2016 | A Testbed for Collecting QoS Data of Cloud-Based Analytic ServicesabstractQoS-based service selection has been an active research area in both Service Computing and Cloud Computing communities in recent years. One of the obstacles for many researchers is lack of benchmark QoS datasets, especially for cloud service selection. In this paper, we propose a testbed system which can be used to collect the QoS data for software services hosted in the cloud. We mainly focus on analytic services, whose QoS values could be dependent on the data they are used to process and analyze. Using the testbed, we can publish services (software or infrastructure services), we can invoke software services hosted on certain infrastructure services, and system can monitor and record the QoS values of all the invocations. We have implemented a proof-of-concept prototype system and used it to collect a sample QoS dataset. Md Shahinur Rahman, Chen Ding 0004, Xumin Liu, Chihung Chi |
CLOUD | 3 |
| 2016 | Data-Dependent QoS-Based Service Selection
Navati Jain, Chen Ding 0004, Xumin Liu |
ICSOC | 3 |
| 2016 | An LDA-SVM Active Learning Framework for Web Service ClassificationabstractClassifying Web services and labeling them based on their functional features have played a major role in several fundamental service management tasks, such as service discovery, selection, ranking, and recommendation. Existing approaches leverage text mining techniques and follow a supervised learning process, which involves building a classifier from a training set of services and applying the classifier to other services. This process requires intensive human effort on labeling services in the training set. In this paper, we propose to leverage the idea of pool-based active learning to realize a scalable service classification approach. Instead of manually labeling a large number of services to construct a complete training set, the approach starts with a base classifier with a small set of training set and iteratively asks for the labels of the most informative services outside of the initial training set. By doing this, the classifier can achieve comparable accuracy compared to traditional classification method with much smaller size of training set. We use SVM as the base classifier due to its effectiveness in text classification. We also incorporate probabilistic topic models to address the issues caused by sparse term vectors generated from service descriptions and reduce the dimensions to improve the efficiency. We conducted a comprehensive experimental study on real-world service data to demonstrate the effectiveness of the proposed approach. Xumin Liu, Shaleen Agarwal, Chen Ding 0004, Qi Yu 0001 |
ICWS | 1 |
| 2015 | Aggregating Functionality, Use History, and Popularity of APIs to Recommend Mashup Creation
Xumin Liu, Qi Yu 0001 |
ICSOC | 2 |
| 2015 | Incorporating User, Topic, and Service Related Latent Factors into Web Service RecommendationabstractDue to the large and increasing number of web services, it is very helpful to provide a proactive feed on what is available to users, i.e., Recommending web services. As collaborative filtering (CF) is an effective recommendation method by capturing latent factors, it has been used for service recommendation as well. However, the majority of current CF-based service recommendation approaches predict users' interests through the historical usage data, but not the service description. This makes them suitable for making QoS-based recommendation, but not for functionality-based recommendation. In this paper, we propose to use machine learning approaches to recommend web services to users from both historical usage data and service descriptions. Considering the great popularity of Restful services, our approach is applicable to both structured and unstructured service description, i.e., Free text descriptions. We exploit the idea of collaborative topic regression, which combines both probabilistic matrix factorization and probabilistic topic modeling, to form user-related, service-related, and topic related latent factor models and use them to predict user interests. We extracted public web service data and developer invocation history from Programmable Web and conducted a comprehensive experiment study. The result indicates that this approach is effective and outperforms other representative recommendation methods. Xumin Liu, Isankumar Fulia |
ICWS | 1 |
| 2015 | Extracting, Ranking, and Evaluating Quality Features of Web Services through User Review Sentiment AnalysisabstractQuality of Service (QoS) has become a standard way of evaluating web services and selecting the one that suites user interests the best. Traditional methods adopt a fixed set of QoS parameters and typical ones include response time, fee, and availability. There currently lacks an effective way of identifying quality features that users are actually interested in when choosing a service. Meanwhile, the traditional way of collecting QoS values relies on either public information released by service providers or test results from repeatedly invoking a service. Therefore, the values can be heavily affected by authenticity of the provider offered information or the quality/configuration of the test code/environment. As a result, existing QoS evaluation methods are not applicable to subject features, such as usability and affordability, where the values depend on user personal judgement. In this paper, we propose a novel approach to extracting domain-related QoS features, ranking those features based on their interestingness, evaluating the value of these features through sentiment analysis on user reviews. More specifically, we leverage natural language processing techniques and machine learning approaches to identify top QoS features that users are interested in and simultaneously learn their sentiment orientation towards those features. We model the problem as sentiment classification, where relevant terms in a review are modeled as features that determine whether a review is positive or negative. Logistic regression is used so that the impact of these terms are learned simultaneously when the classifier is learned through a supervised learning process. The nontrivial terms are selected as the candidate QoS featured. A comprehensive experiment has been conducted on a real-world dataset and the result demonstrates the effectiveness of our approach. Xumin Liu, Arpeet Kale, Javed Wasani, Chen Ding 0004, Qi Yu 0001 |
ICWS | 1 |
| 2015 | Efficient agglomerative hierarchical clustering
Athman Bouguettaya, Qi Yu 0001, Xumin Liu, Xiangmin Zhou, Andy Song |
Expert Syst. Appl. | 3 |
| 2015 | Personalized Decision-Strategy based Web Service Selection using a Learning-to-Rank AlgorithmabstractIn order to choose from a list of functionally similar services, users often need to make their decisions based on multiple QoS criteria they require on the target service. In this process, different users may follow different decision making strategies, some are compensatory in which only an overall value on all the criteria is evaluated, some evaluate one criterion at a time in the order of their importance levels, while others count on the number of winning criteria. Most of the current QoS-based service selection systems do not consider these decision strategies in the ranking process, which we believe are crucial for generating accurate ranking results for individual users. In this paper, we propose a decision strategy based service ranking model. Furthermore, considering that different users follow different strategies in different contexts at different times, we apply a machine learning algorithm to learn a personalized ranking model for individual users based on how they select services in the past. We have implemented and tested the proposed approach, and our experiment results show the effectiveness of the approach. Muhammad Suleman Saleem, Chen Ding 0004, Xumin Liu, Chihung Chi |
IEEE Trans. Serv. Comput. | 3 |
| 2014 | Unraveling and Learning Workflow Models from Interleaved Event LogsabstractBusiness process mining is to extract process knowledge from a system's log in order to reconstruct workflow models. Existing approaches treat a log record as an instance of one workflow model. They do not deal with interleaved logs, where each log record is a mixture of multiple workflow traces. However, such an interleaved log is typical for many systems especially web-based ones where all the user-system interaction traces are recorded and maintained by a web server. Dealing with interleaved logs is challenging due to the lack of prior knowledge of workflow models and noises contained in the log data. In this paper, we propose a two-phase workflow learning process. During the first phase, we use a probabilistic approach to learn the links between operations and the hidden workflow models. We consider a workflow model as a probabilistic distributions over operations and derive it through likelihood maximization. This allows us to identify the membership of an operation to a workflow model, which can be used to unravel a log record and generate a set of workflow instances from it. During the second phase, the sequential patterns between operations within each workflow model are derived from all its instances. We have conducted a comprehensive experimental study, which indicates the effectiveness of the proposed solution. Xumin Liu |
ICWS | 1 |
| 2014 | Personalized Decision Making for QoS-Based Service SelectionabstractIn order to choose from a list of functionally similar services, users often need to make their decisions based on multiple QoS criteria they require on the target service. In this process, different users may follow different decision making strategies, some are compensatory in which only an overall value on all the criteria is evaluated, some evaluate one criterion at a time in the order of their importance levels, while others count on the number of winning criteria. Most of the current QoS-based service selection systems do not consider these decision strategies in the ranking process, which we believe are crucial for generating accurate ranking results for individual users. In this paper, we propose a decision strategy based service ranking model. Furthermore, considering that different users follow different strategies in different contexts at different times, we apply a machine learning algorithm to learn a personalized ranking model for individual users based on how they select services in the past. Our experiment result shows the effectiveness of the proposed approach. Muhammad Suleman Saleem, Chen Ding 0004, Xumin Liu, Chihung Chi |
ICWS | 3 |
| 2014 | Teaching service-oriented programming to CS and SE undergraduate students (abstract only)abstractNo abstract available. Xumin Liu, Rajendra K. Raj, Thomas Reichlmayr, Alex Pantaleev |
SIGCSE | 1 |
| 2013 | Teaching Service-Oriented Programming to CS and SE undergraduate studentsabstractService-Oriented Programming (SOP) is a relatively new programming paradigm that supports the development of new software applications using existing services as building blocks. SOP has gained significant popularity in industry as it increases software reuse and productivity. As the SOP paradigm can improve modern software development, the presenters have created a course-module based approach for incorporating SOP into Computer Science (CS) and Software Engineering (SE) curricula; a course module is a distinct curricular unit such as a lab or teaching component that an instructor may incorporate into an existing course typically without requiring formal curricular approval. SOP course modules have been developed for inclusion in standard courses in many CS and SE programs; for example, an introductory SOP course module in a CS2 course while advanced modules for courses such as Programming Language Concepts, Software Engineering, or Web Services. This workshop will present basic concepts and techniques of SOP and describe how the course-module approach toward SOP can be adapted for the participants' own teaching. The typical participant would be a faculty member with some background in programming, and is interested in learning more about SOP but does need not to have prior web service programming experience. Xumin Liu, Rajendra K. Raj, Thomas Reichlmayr, Alex Pantaleev |
FIE | 1 |
| 2013 | Incorporating Service-Oriented Programming techniques into undergraduate CS and SE curriculaabstractService-Oriented Programming (SOP) has emerged as a new programming paradigm that allows the wrapping of existing software as web services, thus permitting the development of new software applications by using existing web services as building blocks. SOP has attracted great attention from industry as it dramatically increases software reuse. Despite the growing demand for an SOP-trained workforce, SOP has not been adequately covered in coursework for undergraduate students in Computer Science (CS) and Software Engineering (SE). This project addresses this curricular shortcoming via the design and creation of SOP materials for undergraduate CS and SE. The concept of course modules-self-contained units of instruction that can be incorporated into several existing courses-is used to make these materials accessible at multiple educational institutions. This paper describes an exemplification and visualization framework that supports the teaching of SOP, along with three course modules that can be folded into typical courses currently offered to CS or SE undergraduates. Xumin Liu, Rajendra K. Raj, Thomas Reichlmayr, Alex Pantaleev |
FIE | 1 |
| 2013 | Teaching business analyticsabstractIt is essential to prepare students with knowledge and skills in area of business analytics (BA) which will help business to process data, find patterns and relations, develop insights from past transactions, and make prediction. We develop hands-on labs to teach business analytics to students in Computer Science, Information Technology, and Software Engineering disciplines. Our hands-on labs can be adopted in courses such as database systems, data warehousing, data mining, etc. We use enterprise BA tools including MS SQL Server Business Intelligence and Cognos 10 platforms, which are essential to increase student interests, improve student learning, and enhance student confidence. Our hands-on labs contain three parts with one is built upon another: 1) Data integration; 2) Data Warehouse; and 3) Business analytics. Li Yang 0001, Xumin Liu |
FIE | 2 |
| 2013 | Incorporating User Behavior Patterns to Discover Workflow Models from Event LogsabstractWe propose a novel approach to discover workflow models from event logs. The proposed approach addresses two major limitations of current process mining approaches. First, they assume either a single workflow model for the entire event log or the availability of workflow ids that can be used to group logs associated with the same workflow model together. Nonetheless, these assumptions are oversimplified as a complex system typically runs multiple workflow models, all of which share the same log system. Second, existing process mining approaches do not consider the usage patterns of workflow users. Most systems support multi-users and each user is typically associated with (or use) certain number of operation sequences, which may all follow one or several workflow models. Hence, we propose to leverage User Behavior Patterns (or UBPs) to improve the outcome of process mining. In particular, we exploit machine learning techniques to incorporate UBPs into sequence clustering for workflow model discovery. We model a UBP as a probabilistic distribution on sequences, which allows to compute the distance between a UBP and any sequence. We apply three-way matrix factorization onto a UBP-sequence distance matrix to co-cluster users and sequences. In this way, users that share similar UBPs are grouped together while the clustering of similar sequences will lead to the discovery of workflow models. An comprehensive experimental study is conducted to demonstrate the effectiveness and efficiency of the proposed approach. Xumin Liu, Hua Liu 0001, Chen Ding 0004 |
ICWS | 1 |
| 2013 | Ev-LCS: A System for the Evolution of Long-Term Composed ServicesabstractWe propose a system, called EVolution of Long-term Composed Services (Ev-LCS), to address the change management issues in long-term composed services (LCSs). An LCS is a dynamic collaboration between autonomous web services that collectively provide a value-added service. It has a long-term commitment to its users. We first present a formal model, which provides the grounding semantics to support the automation of change management. We present a set of change operators that allow to specify a change in a precise and formal manner. We then propose a change enactment strategy that actually implements the changes. We develop a prototype system for the proposed Ev-LCS to demonstrate its effectiveness. We also conduct an experimental study to assess the performance of the change management approach. Xumin Liu, Athman Bouguettaya, Jemma Wu |
IEEE Trans. Serv. Comput. | 1 |
| 2012 | Collaboration visualization on large dataset for protein-protein interaction networkabstractThe accumulation of huge amount of biology data and their heterogeneity has become a bottleneck in the analysis of protein-protein interaction (PPI) networks, especially in the visualization of PPI networks. Because the format of the data generated from different experimental groups is diverse, and the databases for the storage and management of the data are different, network visualization of the heterogeneous data by integrating different derived data is challenging and is a key to comprehensively understanding the mechanism of biology system. To visualize the interactions of proteins, we first utilize the robot crawl technique to dynamically integrate the information of protein-protein interactions from all the related public databases such as Protein Interaction Database (PID), Human Protein Reference Database (HPRD) and Reactome. Second, we use a graph algorithm to partition the complex network into different sub-networks to discover the `Hub' proteins in the visualization protein-protein networks. Finally, we develop a protocol for the collaboration of different researchers based on the visualization of the PPI networks. Hui Li 0032, Xumin Liu |
CIBCB | 3 |
| 2012 | An asynchronous based approach to improve concurrency control in mobile web serversabstractThe recent boom in mobile computing has created a juncture where it is now possible to host web services on a mobile web server. However, there remain challenges due to limitations in available mobile hardware and software capable of managing resource intensive applications. This paper addresses the Sudipan Mishra, Siddharth Sarasvati, Xumin Liu |
CollaborateCom | 4 |
| 2012 | Automatic Abstract Service Generation from Web Service CommunitiesabstractThe concept of abstract services has been widely adopted in service computing to specify the functionality of certain types of Web services. It significantly benefits key service management tasks, such as service discovery and composition, as these tasks can be first applied to a small number of abstract services and then mapped to the large scale actual services. However, how to generate abstract services is non-trivial. Current approaches either assume the existence of abstract services or adopt a manual process that demands intensive human intervention. We propose a novel approach to fully automate the generation of abstract services from a service community that consists of a set of functionally similar services. A set of candidate outputs are first discovered based on predefined support ratio, which determines the minimum number of services that produce the outputs. Then, the matching inputs are identified to form the abstract services. We propose a set of heuristics to effectively prune a large number of candidate abstract services. An comprehensive experimental study on real world web service data is conducted to demonstrate the effectiveness and efficiency of the proposed approach. Xumin Liu, Hua Liu 0001 |
ICWS | 1 |
| 2012 | Automating Reusable Workflow Development from Design to InstantiationabstractThe proliferation of web services in both number and variety implies the co-existence of a wide number of service options, input/output data types, and encapsulations. Consequently, composing services into usable workflows has become increasingly development intensive. In order to leverage the design of a workflow and facilitate its reusability and maintenance, many research efforts have advocated to compose services at the type level instead of the instance level while using customized glue code to map service types to service instances. Another challenge then appears: Service types, either manually defined by domain experts or automatically generated from an ontology, cannot be automatically instantiated into concrete services due to the coarse granularity of service types and the complexity of input/output parameter mapping. This paper proposes a platform to automatically extract instantiable abstract operations from registered services and that enables the automatic generation of glue code that links concrete services to service types in order to produce reusable executable workflows. Yasmine Charif, Hua Liu 0001, Andres Quiroz, Xumin Liu |
SERVICES | 4 |
| 2011 | Bootstrapping operation-level web service ontology: A bottom-up approachabstractOntology is the key ingredient of semantic Web service technologies, which support systematic management of Web services, such as automatic service discovery, service composition, and change management. It is crucial and challenging to reduce the human efforts for developing ontologies. We propose Xumin Liu, Hua Liu 0001 |
CollaborateCom | 1 |
| 2011 | Constructing Operation-Level Ontologies for Web ServicesabstractWe propose an integrated framework that extracts semantics from WSDL descriptions and constructs operation-level ontologies for Web services. The semantics mainly focus on the functional features of Web services, which facilitates the efficient usage of Web services, such as service discovery and service composition. We use service operations as the first class objects to define service functionalities. We first create service ontologies by measuring the relevance between service operations and clustering operations into functionally relevant groups. We then construct the structure of the service ontologies through a hierarchical clustering algorithm. Xumin Liu, Hua Liu 0001 |
ICWS | 1 |
| 2011 | Evolutionary spectral co-clusteringabstractCo-clustering is the problem of deriving sub-matrices from the larger data matrix by simultaneously clustering rows and columns of the data matrix. Traditional co-clustering techniques are inapplicable to problems where the relationship between the instances (rows) and features (columns) evolve over time. Not only is it important for the clustering algorithm to adapt to the recent changes in the evolving data, but it also needs to take the historical relationship between the instances and features into consideration. We present ESCC, a general framework for evolutionary spectral co-clustering. We are able to efficiently co-cluster evolving data by incorporation of historical clustering results. Under the proposed framework, we present two approaches, Respect To the Current (RTC), and Respect To Historical (RTH). The two approaches differ in the way the historical cost is computed. In RTC, the present clustering quality is of most importance and historical cost is calculated with only one previous time-step. RTH, on the other hand, attempts to keep instances and features tied to the same clusters between time-steps. Extensive experiments performed on synthetic and real world data, demonstrate the effectiveness of the approach. Nathan Green, Manjeet Rege, Xumin Liu, Reynold J. Bailey |
IJCNN | 3 |
| 2011 | Efficient change management in long-term composed services
Xumin Liu, Athman Bouguettaya, Qi Yu 0001, Zaki Malik |
Serv. Oriented Comput. Appl. | 1 |
| 2011 | Service-Centric Framework for a Digital Government ApplicationabstractThis paper presents a service-oriented digital government infrastructure focused on efficiently providing customized services to senior citizens. We designed and developed a Web Service Management System (WSMS), called WebSenior, which provides a service-centric framework to deliver government services to senior citizens. The proposed WSMS manages the entire life cycle of third-party web services. These act as proxies for real government services. Due to the specific requirements of our digital government application, we focus on the following key components of WebSenior: service composition, service optimization, and service privacy preservation. These components form the nucleus that achieves seamless cooperation among government agencies to provide prompt and customized services to senior citizens. Athman Bouguettaya, Qi Yu 0001, Xumin Liu, Zaki Malik |
IEEE Trans. Serv. Comput. | 3 |
| 2010 | Peptide Sequence Tag-Based Blind Identification-based SVM ModelabstractIdentifying the ion types for a mass spectrum is essential for interpreting the spectrum and deriving its peptide sequence. In this paper, we proposed a novel method for identifying ion types and deriving matched peptide sequences for tandem mass spectra. We first divided our dataset into a training set and a testing set and then preprocessed the data using a Support Vector Machine and a 5-fold cross validation based dual denoting model. Then we constructed a syntax tree and generated a rule set to match the mass values from experimental mass spectra with the mass spectral values from corresponding theoretical mass spectra. Finally we applied the proposed algorithm to a tandem mass spectral dataset consisting of 2656 spectra from yeast. Compared with other methods, the experimental results showed that the proposed method can effectively filter noise and successfully derive peptide sequences. Hui Li 0032, Xumin Liu, Macire Diakite, Legand L. Burge III, Abdul-Aziz Yakubu, William M. Southerland |
ICMLA | 3 |
| 2010 | Semantic Support for Adaptive Long Term Composed ServicesabstractWe propose an integrated framework that manages changes in long term composed services. The main procedure of change reaction is presented. One of the most challenging research issues of change management is how to automate the process of change reaction. To address this issue, we propose a semantic support, which centers around a tree-structured Web service ontology. The ontology is expected to provide sufficient semantic for change reaction. We propose a set of algorithms for efficiently querying semantics from the ontology. We conduct a set of experiments to evaluate the performance of the proposed algorithms. Xumin Liu, Manjeet Rege, Athman Bouguettaya |
ICWS | 1 |
| 2010 | VDictionary: Automatically Generate Visual Dictionary via Wikimedias
Yanling Wu, Guangda Li, Zhiping Luo, Tat-Seng Chua, Xumin Liu |
MMM | 6 |
| 2010 | Hyperbolic polynomial uniform B-spline curves and surfaces with shape parameter
Xumin Liu |
Graph. Model. | 1 |
| 2010 | End-to-End Service Support for MashupsabstractWe propose a service-oriented approach to generate and manage mashups. The proposed approach is realized using the Mashup Services System (MSS), a novel platform to support users to create, use, and manage mashups with little or no programming effort. The proposed approach relieves users from programming-intensive, error-prone, and largely nonreusable output process for creating and maintaining mashups. We describe the overall design of MSS and discuss and evaluate its main enabling technologies. Athman Bouguettaya, Surya Nepal, Wanita Sherchan, Xuan Zhou 0001, Jemma Wu, Shiping Chen 0001, Dongxi Liu, Lily Li 0002, Xumin Liu |
IEEE Trans. Serv. Comput. | 10 |
| 2009 | Efficient Access to Composite M-servicesabstractWireless Web services, also called mobile services or M-services, provide access to Web services through wireless networks. In this paper, we propose novel access methods and multi-channel organization for mobile users to effectively access composite M-services in wireless broadcast networks. We define a few semantics for accessing broadcast based M-services and study their impact on access efficiency. Xu Yang 0008, Athman Bouguettaya, Xumin Liu |
ICWS | 3 |
| 2008 | Ontology Support for Managing Top-Down Changes in Composite Services
Xumin Liu, Athman Bouguettaya |
CollaborateCom | 1 |
| 2008 | Deploying and managing Web services: issues, solutions, and directions
Qi Yu 0001, Xumin Liu, Athman Bouguettaya, Brahim Medjahed |
VLDB J. | 2 |
| 2007 | Reacting to functional changes in service-oriented enterprisesabstractIn this paper, we focus on the changes that trigger the modification of a service-oriented enterprisepsilas functionality. We present a framework that helps an SOE automatically modify its functional schema based on a change specification. The central component of this framework is a change reaction manager. It provides mechanisms to interpret a change specification, modify the member services in an SOE, and rearrange their cooperation correspondingly. A domain knowledge provider offers the semantics that facilitates the change reaction process. We also use a logging mechanism to keep track of the change reaction process. Xumin Liu, Athman Bouguettaya |
CollaborateCom | 1 |
| 2007 | Managing Top-down Changes in Service-Oriented EnterprisesabstractA Service Oriented Enterprise (SOE) provides an efficient and flexible platform where multiple Web services can cooperate together to provide a value-added service. Change management is one of the fundamental issues in enabling SOEs. In this paper, we propose a framework that facilitates in automatically managing top-down changes in SOEs. We start with formalizing a SOE's schema since it is a central concept for specifying and managing top-down changes in SOEs. We then propose a change model as a guide to react to changes. Algorithms are proposed to implement changes by refining a SOE's behavior. Xumin Liu, Athman Bouguettaya |
ICWS | 1 |