San-Yih Hwang

dblp:22/852 · DBLP profile ↗
← Back
26ranked-venue papers in the field
11as first author
3since 2021 · last 2025
—ORCID · none

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 14 (6 first)Data Mining & Knowledge Discovery · 7 (1 first)Information Retrieval & Web Search · 3 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Business Process & Enterprise Data · 1 (1 first)
YearPublicationVenuePosition
2025 LITA: An Efficient LLM-Assisted Iterative Topic Augmentation Framework
Chia-Hsuan Chang, Jui-Tse Tsai, Yi-Hang Tsai, San-Yih Hwang
PAKDD (1)4
2022 An Integration of TextGCN and Autoencoder into Aspect-Based Sentiment Analysis
Yi-Hang Tsai, Kun-Hsiang Chen, San-Yih Hwang
DaWaK4
2021 A word embedding-based approach to cross-lingual topic modeling
Chia-Hsuan Chang, San-Yih Hwang
Knowl. Inf. Syst.2
2010 Automatic index construction for multimedia digital libraries
San-Yih Hwang, Wan-Shiou Yang, Kang-Di Ting
Inf. Process. Manag.1
2008 Efficient algorithms for mining maximal valid groups
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
VLDB J.3
2007 A probabilistic approach to modeling and estimating the QoS of web-services-based workflows
San-Yih Hwang, Haojun Wang, Jian Tang 0001, Jaideep Srivastava
Inf. Sci.1
2006 Efficient mining of group patterns from user movement data
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
Data Knowl. Eng.3
2005 Estimating missed actual positives using independent classifiers
abstract
Data mining is increasingly being applied in environments having very high rate of data generation like network intrusion detection [7], where routers generate about 300,000 -- 500,000 connections every minute. In such rare class data domains, the cost of missing a rare-class instance is much higher than that of other classes. However, the high cost for manual labeling of instances, the high rate at which data is collected as well as real-time response constraints do not always allow one to determine the actual classes for the collected unlabeled datasets. In our previous work [9], this problem of missed false negatives was explained in context of two different domains -- "network intrusion detection" and "business opportunity classification". In such cases, an estimate for the number of such missed high-cost, rare instances will aid in the evaluation of the performance of the modeling technique (e.g. classification) used. A capture-recapture method was used for estimating false negatives, using two or more learning methods (i.e. classifiers). This paper focuses on the dependence between the class labels assigned by such learners. We define the conditional independence for classifiers given a class label and show its relation to the conditional independence of the features sets (used by the classifiers) given a class label. The later is a computationally expensive problem and hence, a heuristic algorithm is proposed for obtaining conditionally independent (or less dependent) feature sets for the classifiers. Initial results of this algorithm on synthetic datasets are promising and further research is being pursued.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang
KDD3
2005 Mining Mobile Group Patterns: A Trajectory-Based Approach
San-Yih Hwang, Ying-Han Liu, Jeng-Kuen Chiu, Ee-Peng Lim
PAKDD1
2004 Efficient Group Pattern Mining Using Data Summarization
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
DASFAA3
2004 A Probabilistic QoS Model and Computation Framework for Web Services-Based Workflows
San-Yih Hwang, Haojun Wang, Jaideep Srivastava, Raymond A. Paul
ER1
2004 Estimation of False Negatives in Classification
abstract
In many classification problems such as spam detection and network intrusion, a large number of unlabeled test instances are predicted negative by the classifier However, the high costs as well as time constraints on an expert's time prevent further analysis of the "predicted false" class instances in order to segregate the false negatives from the true negatives. A systematic method is thus required to obtain an estimate of the number of false negatives. A capture-recapture based method can be used to obtain an ML-estimate of false negatives when two or more independent classifiers are available. In the case for which independence does not hold, we can apply log-linear models to obtain an estimate of false negatives. However, as shown in this paper, lesser the dependencies among the classifiers, better is the estimate obtained for false negatives. Thus, ideally independent classifiers should be used to estimate the false negatives in an unlabeled dataset. Experimental results on the spam dataset from the UCI machine learning repository are presented.
Sandeep Mane, Jaideep Srivastava, San-Yih Hwang, Jamshid A. Vayghan
ICDM3
2003 On Mining Group Patterns of Mobile Users
Yida Wang 0002, Ee-Peng Lim, San-Yih Hwang
DEXA3
2003 Personal Workflows: Modeling and Management
San-Yih Hwang, Ya-Fan Chen
Mobile Data Management1
2002 A Case for Analytical Customer Relationship Management
Jaideep Srivastava, Jau-Hwang Wang, Ee-Peng Lim, San-Yih Hwang
PAKDD4
2001 Personal Workflow Management in Support of Pervasive Computing
San-Yih Hwang, Jeng-Kuen Chiu, Wan-Shiou Yang
Mobile Data Management1
1999 Mining Exception Instances to Facilitate Workflow Exception Handling
abstract
The importance of exception handling within the context of workflow management has been widely recognized. While some exceptions are expected at design time and thus can be incorporated into the workflow design via some flexible mechanism, others are totally unexpected. Previous work in handling unexpected workflow exceptions focuses on the run-time support to, for example, allow the rollback of some already completed activities, validate the correctness of dynamic workflow change, and deploy a solution to handle exceptions. Authorized persons are responsible for deriving solutions to handle exceptions. We propose a novel approach to facilitating users in proposing solutions for resolving a given exception. Specifically, our approach scans through the previous records in handling exceptions, looking for those that are close to the current exception. The ways in which those exceptions were handled serve as useful information in determining how to handle the current one. Several algorithms are proposed and evaluated through both theoretical analysis and a synthetic data set.
San-Yih Hwang, Sun-Fa Ho, Jian Tang 0001
DASFAA1
1998 A Scheme to Specify and Implement Ad-Hoc Recovery in Workflow Systems
Jian Tang 0001, San-Yih Hwang
EDBT2
1996 Handling Uncertainties in Workflow Applications
abstract
Workflow techniques are currently seen as the key techniques to improve the effectiveness and productivity of business processes.Most workflow models take a static view toward workflow applications.The static view requires that all the components are known in advance and the structural information alone can uniquely determine an execution path.Some applications may not have either or both of these properties, In these applications, uncertainty exists.In this paper, we study the issues related to uncertainties.We characterize uncertainties into three categories, domain uncertainty, structural uncertainty and implementation uncertainty, and discuss the impact they have on workflow specification and implementation.We then propose approaches to coping with each of them.
Jian Tang 0001, San-Yih Hwang
CIKM2
1996 Data Replication in a Distributed System: A Performance Study
San-Yih Hwang, Keith K. S. Lee
DEXA1
1996 Coping with Mismatched Semantics of Dependencies in Workflow Applications
Jian Tang 0001, San-Yih Hwang
DEXA2
1995 The Design and Implementation of a Full-Fledged Multiple DBMS
abstract
We have described our design of the multiple DBMS (MDBMS). This MDBMS enables users to access data controlled by different DBMSs as if data were managed by a single DBMS. It supports facilities for SQL queries and transactions, and considers security functions. In addition, an ODBC driver at the client site has been realized to ease the development of MDBMS applications. Several popular commercial DBMSs, including Oracle, Informix and Sybase, have been successfully integrated. The MDBMS is in operation now. However, we found the performance to be unsatisfactory. It took about several seconds to process an SQL query with single join on two relations of hundreds of tuples. We have identified the performance bottleneck to be on the retrieval of meta data. The current MDBMS Server employs a commercial DBMS to store meta data, which is necessary for processing a global query. The processing of a query is slow because it needs to retrieve the schema information via an external DBMS several times. We are currently designing a core storage manager and an access manager specifically for maintaining the meta data and the intermediate results of a global query. We expect this design to significantly improve the performance.>
Shu-Chin Su Chen, Chih-Shing Yu, Yen-Yao Yao, San-Yih Hwang, B. Paul Lin
ICDE4
1995 An Algebraic Transformation Framework for Multidatabase Queries
Ee-Peng Lim, Jaideep Srivastava, San-Yih Hwang
Distributed Parallel Databases3
1994 The MYRIAD Federated Database Prototype
abstract
No abstract available.
San-Yih Hwang, Ee-Peng Lim, H.-R. Yang, S. Musukula, K. Mediratta, M. Ganesh 0001, Dave Clements, J. Stenoien, Jaideep Srivastava
SIGMOD Conference1
1994 Transaction Recovery in Federated Autonomous Databases
San-Yih Hwang, Jaideep Srivastava, Jianzhong Li 0001
Distributed Parallel Databases1
1993 Concurrency Control in Federated Databases: A Dynamic Approach
abstract
Article Free Access Share on Concurrency control in federated databases: a dynamic approach Authors: San-Yih Hwang Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile , Jiandong Huang Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MN Sensor and System Development Center, Honeywell, 3660 Technology Drive, Minneapolis, MNView Profile , Jaideep Srivastava Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MN Computer Science Department, University of Minnesota, 200 Union St. SE, Minneapolis, MNView Profile Authors Info & Claims CIKM '93: Proceedings of the second international conference on Information and knowledge managementDecember 1993Pages 694–703https://doi.org/10.1145/170088.170458Published:01 December 1993Publication History 3citation370DownloadsMetricsTotal Citations3Total Downloads370Last 12 Months8Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
San-Yih Hwang, Jiandong Huang, Jaideep Srivastava
CIKM1