San-Yih Hwang

dblp:22/852 · DBLP profile ↗
← Back
45ranked-venue papers
22as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 10 · 5 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 LITA: An Efficient LLM-Assisted Iterative Topic Augmentation Framework
Chia-Hsuan Chang, Jui-Tse Tsai, Yi-Hang Tsai, San-Yih Hwang
PAKDD (1)4
2022 An Integration of TextGCN and Autoencoder into Aspect-Based Sentiment Analysis
Yi-Hang Tsai, Kun-Hsiang Chen, San-Yih Hwang
DaWaK4
2021 A word embedding-based approach to cross-lingual topic modeling
Chia-Hsuan Chang, San-Yih Hwang
Knowl. Inf. Syst.2
2015 A Smart Physical World Based on Service Technologies, Big Data, and Game-Based Crowd Sourcing
abstract
The physical world is becoming smarter and smarter due to the advances in smart devices and CPS/IoT technologies. In this paper, we investigate the roles various cutting-edge technologies, such as service computing, big data analytics, crowd sourcing, gaming technologies, etc., can play to significantly enhance the intelligence of our physical environment and subsequently benefit the society. We consider a smart physical world (SPW) consisting of physical entities, cyber entities, and human. Service technologies can be used to model the entities in SPW and specify their capabilities. Service discovery and composition techniques can be used to, based on the modeling, compose these capabilities to solve real-world problems. To enable higher level intelligence, we further discuss how various technologies, such as big data analytics and artificial intelligence, can be used in the smart world to reason from sensor inputs to derive the situation facts, and from the situation facts to derive the reactive actions. From the derived actions (tasks), the service technologies can be used to compose the capabilities of the entities to realize the task. However, current technologies and artificial intelligence may fall short in many situations. We further propose a gaming-based crowd sourcing platform to make use of human intelligence to enable successful completion of some challenging reasoning and control tasks.
I-Ling Yen, Guang Zhou, Wei Zhu 0002, Farokh B. Bastani, San-Yih Hwang
ICWS5
2015 Personalized Hotel Recommendation Using Text Mining and Mobile Browsing Tracking
abstract
With the prevalence of mobile devices such as smartphones and tablets, the ways people access to the Internet have changed enormously. In addition to the information that can be recorded by traditional Web-based e-commerce like frequent online shopping stores and browsing histories, mobile devices are capable of tracking sophisticated browsing behavior. The aim of this study is to utilize users' browsing behavior of reading hotel reviews on mobile devices and subsequently apply text-mining techniques to construct user interest profiles to make personalized hotel recommendations. Specifically, we design and implement an app where the user can search hotels and browse hotel reviews, and every gesture the user has performed on the touch screen when reading the hotel reviews is recorded. We then identify the paragraphs of hotel reviews that a user has shown interests based on the gestures the user has performed. Text mining techniques are applied to construct the interest profile of the user according to the review content the user has seriously read. We collect more than 5,000 reviews of hotels in Taipei, the largest metropolitan area of Taiwan, and recruit 18 users to participate in the experiment. Experimental results demonstrate that the recommendations made by our system better match the user's hotel selections than previous approaches.
Keng-Pei Lin, Chia-Yu Lai, Po-Cheng Chen, San-Yih Hwang
SMC4
2015 Service Selection for Web Services with Probabilistic QoS
abstract
Web services can be specified from two perspectives, namely functional and non-functional properties. Multiple services may possess the same function while vary in their non-functional properties, or called quality-of-service (QoS). QoS values are important criteria for service selection or recommendation. Most of the former works in web service selection and recommendation treat the QoS values as constants. However, QoS values of a service as perceived by a given user are intrinsically random variables because QoS value prediction can never be precise and there are always some unobserved random effects. In this work, we address the service selection problem by representing services' QoS values as discrete random variables with probability mass functions. The goal is to select a set of atomic services for composing a composite service such that the probability of satisfying constraints imposed on the composite service is high and the execution time is reasonable. Our proposed method starts with an initial web service assignment and incrementally adjusts it using simulated annealing. We conduct several experiments and the results show that our approach generally performs better than previous works, such as the integer programming method and the cost-driven method.
San-Yih Hwang, Chien-Ching Hsu, Chien-Hsiang Lee
IEEE Trans. Serv. Comput.1
2013 Toward Ontology and Service Paradigm for Enhanced Carbon Footprint Management and Labeling
abstract
As the green house gas emission becomes a serious problem, a lot of researches now focus on how to monitor and manage carbon footprint (CF) of a production process or transportation, especially in the supply chain. Usually, most of the carbon footprint management systems are based on databases. But database is not sufficient in describing the production and transportation processes and the facilities used in these processes. In this paper, we develop an SOA based model for the carbon footprint management and labeling (CFML), using ontology and OWL-S techniques. We use OWL-S to describe the processes and workflows for production and transportation and extend it to specify the methods for deriving CF of them. We use the existing energy conversion formula to derive the CF when the energy data can be collected separately. We also derive an approach to separate the CFs when data of different processes have to be collected together. In a supply chain, a production company may have different choices of suppliers to provide certain components. To balance the tradeoff between carbon dioxide emission and cost of the overall production process, we design a supplier selection algorithm to derive the optimal solution.
Wei Zhu 0002, Guang Zhou, I-Ling Yen, San-Yih Hwang
ICWS4
2013 Reliable Web service selection in choreographed environments
San-Yih Hwang, Chien-Hsiang Lee
Decis. Support Syst.1
2013 iTravel: A recommender system in mobile peer-to-peer environment
Wan-Shiou Yang, San-Yih Hwang
J. Syst. Softw.2
2012 A Service Pattern Model for Flexible Service Composition
abstract
Although reuse is the main goal of SOA, composing existing services to realize different user requirements is still a difficult and time-consuming task. Research on workflow templates and design patterns can facilitate reuse and assist with the service composition task. Nevertheless, workflow templates are too specific, whereas design patterns may be too abstract; their effectiveness in assisting the composition process may be limited. In this paper, we propose a comprehensive service pattern model that is more flexible than workflow/service templates while allowing systematic instantiation into the concrete workflows. It can help ease the composition process and enable flexible pattern-based reuse.
Chien-Hsiang Lee, San-Yih Hwang, I-Ling Yen
ICWS2
2010 Web Services Selection in Support of Reliable Web Service Choreography
abstract
There are two approaches to specifying the composition of Web services: orchestration and choreography. Previous works in Web services selection are mostly based on the orchestration model which focuses on the interactions with a single party. However, in many application scenarios, business goals are achieved by a number of pair-wise interactions among a set of Web services, and there does not exist a single entity that is in charge of selecting Web services for all tasks. Each Web service will autonomously perform Web services selection. In such a choreographic environment, we study the kind of information that each Web service should provide to its partner Web services and how each Web service should perform Web service selection so as to maximize the chance of successfully accomplishing a business goal. The proposed approach is evaluated by simulation, and the experimental results show that our proposed method is close to centralized method and better than the other two distributed Web services selection methods.
San-Yih Hwang, Wen-Po Liao, Chien-Hsiang Lee
ICWS1
2010 Automatic index construction for multimedia digital libraries
San-Yih Hwang, Wan-Shiou Yang, Kang-Di Ting
Inf. Process. Manag.1
2009 Using trust for collaborative filtering in eCommerce
abstract
Personalization and customization have been shown to be an indispensable function in today's eCommerce businesses and highly applauded by their customers. Collaborative filtering is one of the two major techniques commonly employed by today's recommender systems and has found its way into the recommendation of many diversified types of products. However, collaborative filtering technique suffers from sparsity and cold start problems. The recent emergence of Web 2.0 offers an opportunity to remedy these problems by incorporating the trust relationships explicitly expressed by the users, as evident by some recent research. Previous work in using trust for making recommendation mainly focuses on inferring trust weights for unspecified trust relations. In this paper, we model the problem of using trust for recommendation as a linear program. We then describe two heuristics that leverages trust to estimate the ratings of unseen products by a given user. Finally we develop various strategies of giving continuous trust weights by considering the contextual information pertaining to trust statements and examine their impact on recommendation accuracy using the empirical data collected from Epinion.com. The experimental results show that assigning continuous trust weights using some of the proposed strategies yields higher recommendation accuracy when compared to the baseline approach that gives Boolean trust values.
San-Yih Hwang, Lung-Shian Chen
ICEC1
2008 A Probability-Based Framework for Dynamic Resource Scheduling in Grid Environment
San-Yih Hwang, Jian Tang 0001, Hong-Yang Lin
GPC1
2008 Discovering Generalized Profile-Association Rules for the Targeted Advertising of New Products
abstract
We propose a data-mining approach for the targeted marketing of new products that have never been rated or purchased by customers. This approach uncovers associations between customer types and product genres that frequently occurred in previous transaction records. Customer types are defined in terms of demographic attribute values that can be aggregated through concept hierarchies; product types can be generalized through product taxonomies. We use generalized profile-association rules (GP association rules) to identify the advertising targets for a given new product. In addition, we propose two algorithms—GP-Apriori and Merge-prune—to mine GP association rules and develop a value-based targeted advertising algorithm to select prospective customers of a new product on the basis of the discovered rules. We evaluate the proposed approach using both synthetic data and library-circulation data.
San-Yih Hwang, Wan-Shiou Yang
INFORMS J. Comput.1
2008 Dynamic Web Service Selection for Reliable Web Service Composition
abstract
This paper studies the dynamic web service selection problem in a failure-prone environment, which aims to determine a subset of Web services to be invoked at run-time so as to successfully orchestrate a composite web service. We observe that both the composite and constituent web services often constrain the sequences of invoking their operations and therefore propose to use finite state machine to model the permitted invocation sequences of Web service operations. We assign each state of execution an aggregated reliability to measure the probability that the given state will lead to successful execution in the context where each web service may fail with some probability. We show that the computation of aggregated reliabilities is equivalent to eigenvector computation and adopt the power method to efficiently derive aggregated reliabilities. In orchestrating a composite Web service, we propose two strategies to select Web services that are likely to successfully complete the execution of a given sequence of operations. A prototype that implements the proposed approach using BPEL for specifying the invocation order of a web service is developed and served as a testbed for comparing our proposed strategies and other baseline Web service selection strategies.
San-Yih Hwang, Ee-Peng Lim, Chien-Hsiang Lee, Cheng-Hung Chen
IEEE Trans. Serv. Comput.1
2008 Efficient algorithms for mining maximal valid groups
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
VLDB J.3
2007 On Composing a Reliable Composite Web Service: A Study of Dynamic Web Service Selection
abstract
Dynamic Web service selection refers to determining a subset of component Web services to be invoked so as to orchestrate a composite Web service. Previous work in Web service selection usually assumes the invocations of Web service operations to be independent of one another. This assumption however does not hold in practice as both the composite and component Web services often impose some orderings on the invocation of their operations. Such orderings constrain the selection of component Web services to orchestrate the composite Web service. We therefore propose to use finite state machine (FSM) to model the invocation order of Web service operations. We define a measure, called aggregated reliability, to measure the probability that a given state in the composite Web service will lead to successful execution in the context where each component Web service may fail with some probability. We show that the computation of aggregated reliabilities is equivalent to eigenvector computation. The power method is hence adopted to efficiently derive aggregated reliabilities. In orchestrating a composite Web service, we propose two strategies to select component Web services that are likely to successfully complete the execution of a given sequence of operations. Our experiments on a synthetically generated set of Web service operation execution sequences show that our proposed strategies perform better than the baseline random selection strategy.
San-Yih Hwang, Ee-Peng Lim, Chien-Hsiang Lee, Cheng-Hung Chen
ICWS1
2007 A probabilistic approach to modeling and estimating the QoS of web-services-based workflows
San-Yih Hwang, Haojun Wang, Jian Tang 0001, Jaideep Srivastava
Inf. Sci.1
2006 Efficient mining of group patterns from user movement data
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
Data Knowl. Eng.3
2006 A process-mining framework for the detection of healthcare fraud and abuse
Wan-Shiou Yang, San-Yih Hwang
Expert Syst. Appl.2
2005 Estimating missed actual positives using independent classifiers
abstract
Data mining is increasingly being applied in environments having very high rate of data generation like network intrusion detection [7], where routers generate about 300,000 -- 500,000 connections every minute. In such rare class data domains, the cost of missing a rare-class instance is much higher than that of other classes. However, the high cost for manual labeling of instances, the high rate at which data is collected as well as real-time response constraints do not always allow one to determine the actual classes for the collected unlabeled datasets. In our previous work [9], this problem of missed false negatives was explained in context of two different domains -- "network intrusion detection" and "business opportunity classification". In such cases, an estimate for the number of such missed high-cost, rare instances will aid in the evaluation of the performance of the modeling technique (e.g. classification) used. A capture-recapture method was used for estimating false negatives, using two or more learning methods (i.e. classifiers). This paper focuses on the dependence between the class labels assigned by such learners. We define the conditional independence for classifiers given a class label and show its relation to the conditional independence of the features sets (used by the classifiers) given a class label. The later is a computationally expensive problem and hence, a heuristic algorithm is proposed for obtaining conditionally independent (or less dependent) feature sets for the classifiers. Initial results of this algorithm on synthetic datasets are promising and further research is being pursued.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang
KDD3
2005 Mining Mobile Group Patterns: A Trajectory-Based Approach
San-Yih Hwang, Ying-Han Liu, Jeng-Kuen Chiu, Ee-Peng Lim
PAKDD1
2004 Efficient Group Pattern Mining Using Data Summarization
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
DASFAA3
2004 A Probabilistic QoS Model and Computation Framework for Web Services-Based Workflows
San-Yih Hwang, Haojun Wang, Jaideep Srivastava, Raymond A. Paul
ER1
2004 Estimation of False Negatives in Classification
abstract
In many classification problems such as spam detection and network intrusion, a large number of unlabeled test instances are predicted negative by the classifier However, the high costs as well as time constraints on an expert's time prevent further analysis of the "predicted false" class instances in order to segregate the false negatives from the true negatives. A systematic method is thus required to obtain an estimate of the number of false negatives. A capture-recapture based method can be used to obtain an ML-estimate of false negatives when two or more independent classifiers are available. In the case for which independence does not hold, we can apply log-linear models to obtain an estimate of false negatives. However, as shown in this paper, lesser the dependencies among the classifiers, better is the estimate obtained for false negatives. Thus, ideally independent classifiers should be used to estimate the false negatives in an unlabeled dataset. Experimental results on the spam dataset from the UCI machine learning repository are presented.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang, Jamshid A. Vayghan
ICDM3
2004 Consulting past exceptions to facilitate workflow exception handling
San-Yih Hwang, Jian Tang 0001
Decis. Support Syst.1
2003 On Mining Group Patterns of Mobile Users
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
DEXA3
2003 Personal Workflows: Modeling and Management
San-Yih Hwang, Ya-Fan Chen
Mobile Data Management1
2002 A Case for Analytical Customer Relationship Management
Jaideep Srivastava, Jau-Hwang Wang, Ee-Peng Lim, San-Yih Hwang
PAKDD4
2002 On the discovery of process models from their instances
San-Yih Hwang, Wan-Shiou Yang
Decis. Support Syst.1
2001 Personal Workflow Management in Support of Pervasive Computing
San-Yih Hwang, Jeng-Kuen Chiu, Wan-Shiou Yang
Mobile Data Management1
1999 Mining Exception Instances to Facilitate Workflow Exception Handling
abstract
The importance of exception handling within the context of workflow management has been widely recognized. While some exceptions are expected at design time and thus can be incorporated into the workflow design via some flexible mechanism, others are totally unexpected. Previous work in handling unexpected workflow exceptions focuses on the run-time support to, for example, allow the rollback of some already completed activities, validate the correctness of dynamic workflow change, and deploy a solution to handle exceptions. Authorized persons are responsible for deriving solutions to handle exceptions. We propose a novel approach to facilitating users in proposing solutions for resolving a given exception. Specifically, our approach scans through the previous records in handling exceptions, looking for those that are close to the current exception. The ways in which those exceptions were handled serve as useful information in determining how to handle the current one. Several algorithms are proposed and evaluated through both theoretical analysis and a synthetic data set.
San-Yih Hwang, Sun-Fa Ho, Jian Tang 0001
DASFAA1
1998 Component and Data Distribution in a Distributed Workflow Management System
abstract
Designing a distributed workflow management system (WFMS) has become a topic with tremendous research interests in recent years. Unlike a traditional client-server WFMS that executes only at a single site, a distributed WFMS allows the execution of a workflow instance to be governed by more than one servers so as to achieve availability and efficiency. A consensus has generally been reached that the component-based approach is a promising one to designing a distributed WFMS. The component-based approach divides the functionality of a WFMS into a number of components, each of which follows some well-accepted distributed object standard, e.g. CORBA or DCOM. These components can be distributed transparently to different sites for execution, and cooperatively they achieve the goal of the underlying business process. Following the component-based approach, this paper proposes and discusses several alternatives on how to distribute the components and the data of a distributed WFMS. Specifically, the component distribution problem can be transformed into two combinatorial problems, namely the integer programming problem and the weighted matching problem of a bipartite graph under different conditions. Regarding the data distribution, we recommend the full replication of workflow definition data and suggest different strategies in placing different types of instance data.
San-Yih Hwang, Chi-Ten Yang
APSEC1
1998 A Scheme to Specify and Implement Ad-Hoc Recovery in Workflow Systems
Jian Tang 0001, San-Yih Hwang
EDBT2
1997 Parallel Array Object I/O Support on Distributed Environments
Jenq Kuen Lee, Ing-Kuen Tsaur, San-Yih Hwang
J. Parallel Distributed Comput.3
1996 Handling Uncertainties in Workflow Applications
abstract
Workflow techniques are currently seen as the key techniques to improve the effectiveness and productivity of business processes.Most workflow models take a static view toward workflow applications.The static view requires that all the components are known in advance and the structural information alone can uniquely determine an execution path.Some applications may not have either or both of these properties, In these applications, uncertainty exists.In this paper, we study the issues related to uncertainties.We characterize uncertainties into three categories, domain uncertainty, structural uncertainty and implementation uncertainty, and discuss the impact they have on workflow specification and implementation.We then propose approaches to coping with each of them.
Jian Tang 0001, San-Yih Hwang
CIKM2
1996 Data Replication in a Distributed System: A Performance Study
San-Yih Hwang, Keith K. S. Lee
DEXA1
1996 Coping with Mismatched Semantics of Dependencies in Workflow Applications
Jian Tang 0001, San-Yih Hwang
DEXA2
1995 The Design and Implementation of a Full-Fledged Multiple DBMS
abstract
We have described our design of the multiple DBMS (MDBMS). This MDBMS enables users to access data controlled by different DBMSs as if data were managed by a single DBMS. It supports facilities for SQL queries and transactions, and considers security functions. In addition, an ODBC driver at the client site has been realized to ease the development of MDBMS applications. Several popular commercial DBMSs, including Oracle, Informix and Sybase, have been successfully integrated. The MDBMS is in operation now. However, we found the performance to be unsatisfactory. It took about several seconds to process an SQL query with single join on two relations of hundreds of tuples. We have identified the performance bottleneck to be on the retrieval of meta data. The current MDBMS Server employs a commercial DBMS to store meta data, which is necessary for processing a global query. The processing of a query is slow because it needs to retrieve the schema information via an external DBMS several times. We are currently designing a core storage manager and an access manager specifically for maintaining the meta data and the intermediate results of a global query. We expect this design to significantly improve the performance.>
Shu-Chin Su Chen, Chih-Shing Yu, Yen-Yao Yao, San-Yih Hwang, B. Paul Lin
ICDE4
1995 An Algebraic Transformation Framework for Multidatabase Queries
Ee-Peng Lim, Jaideep Srivastava, San-Yih Hwang
Distributed Parallel Databases3
1995 Myriad: Design and Implementation of a Federated Database Prototype
abstract
Abstract A key problem in providing ‘enterprise‐wide’ information is the integration of databases that have been independently developed. An important requirement is to accommodate heterogeneity and maintain the autonomy of component databases. Myriad is a federated database prototype developed at the University of Minnesota, to provide a testbed for investigating alternatives in architecture and algorithms for database integration, query processing and optimization, and concurrency control and recovery. The system incorporates our group's research results in these areas. This paper describes our experiences in the design and implementation of Myriad, and in the project management. Special emphasis is given to discussing design alternatives and their impact on Myriad. This paper also presents the software engineering principles and the project management techniques we used in developing Myriad and the lessons we learned. We believe these lessons would be useful for practitioners who wish to develop a similar system. Handling heterogeneity and autonomy were prime objectives throughout the prototyping effort. We are convinced that a prototype federated database is an important infrastructural requirement for the overall goal of ‘enterprise‐integration’, and believe Myriad to be a significant contribution towards this.
Ee-Peng Lim, San-Yih Hwang, Jaideep Srivastava, Dave Clements, M. Ganesh 0001
Softw. Pract. Exp.2
1994 The MYRIAD Federated Database Prototype
abstract
No abstract available.
San-Yih Hwang, Ee-Peng Lim, H.-R. Yang, S. Musukula, K. Mediratta, M. Ganesh 0001, Dave Clements, J. Stenoien, Jaideep Srivastava
SIGMOD Conference1
1994 Transaction Recovery in Federated Autonomous Databases
San-Yih Hwang, Jaideep Srivastava, Jianzhong Li 0001
Distributed Parallel Databases1
1993 Concurrency Control in Federated Databases: A Dynamic Approach
abstract
Article Free Access Share on Concurrency control in federated databases: a dynamic approach Authors: San-Yih Hwang Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile , Jiandong Huang Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MN Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MNView Profile , Jaideep Srivastava Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile Authors Info & Claims CIKM '93: Proceedings of the second international conference on Information and knowledge managementDecember 1993Pages 694–703https://doi.org/10.1145/170088.170458Published:01 December 1993Publication History 3citation370DownloadsMetricsTotal Citations3Total Downloads370Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
San-Yih Hwang, Jiandong Huang, Jaideep Srivastava
CIKM1