Michael Chau

dblp:25/1069 · DBLP profile ↗
← Back
36ranked-venue papers
15as first author
4since 2021 · last 2026
0000-0003-4579-4329ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSecurity and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Impact of copyright sharing on the success of non-fungible token collections
Cheng Tao 0004, Michael Chau, Daning Hu
Inf. Manag.3
2025 Watermarking Large Language Models: An Unbiased and Low-risk Method
abstract
Recent advancements in large language models (LLMs) have highlighted the risk of misusing them, raising the need for accurate detection of LLM-generated content.In response, a viable solution is to inject imperceptible identifiers into LLMs, known as watermarks.Our research extends the existing watermarking methods by proposing the novel Sampling One Then Accepting (STA-1) method.STA-1 is an unbiased watermark that preserves the original token distribution in expectation and has a lower risk of producing unsatisfactory outputs in low-entropy scenarios compared to existing unbiased watermarks.In watermark detection, STA-1 does not require prompts or a white-box LLM, provides statistical guarantees, demonstrates high efficiency in detection time, and remains robust against various watermarking attacks.Experimental results on low-entropy and high-entropy datasets demonstrate that STA-1 achieves the above properties simultaneously, making it a desirable solution for watermarking LLMs.Implementation codes for this study are available online. 1
Minjia Mao, Dongjun Wei, Michael Chau
ACL (1)5
2024 Are real-time volunteer apps really helping visually impaired people? A social justice perspective
Huilin Gao, Evelyn Ng, Bingjie Deng, Michael Chau
Inf. Manag.4
2021 The Role of Attitude toward Challenge in Serious Game Design
abstract
This paper proposes that challenge, an orthogonal game attribute, can be used to improve game effectiveness. The results of our study suggest that attitude toward challenge should be considered when we tune challenge in the serious game for better learning results. A significant moderating role of attitude toward challenge is found in the relationship between challenge and self-efficacy. Investigating from the perspective of game attributes together with players’ attitude toward the attributes is a good approach to improving game design. Our approach of game improvement is a clear, explicit and one-to-one approach to relate game attributes and attitudes toward them.
Philip T. Y. Lee, Michael Chau, Richard W. C. Lui
J. Comput. Inf. Syst.2
2016 Identifying features for detecting fraudulent loan requests on P2P platforms
abstract
This exploratory study is intended to address the problem of fraudulent loan requests on peer-to-peer (P2P) platforms. We propose a set of features that capture the behavioral characteristics (e.g., learning, past performance, social networking, and herding manipulation) of malevolent borrowers, who intentionally create loan requests to acquire funds from lenders but default later on. We found that using the widely adopted classification methods such as Random Forest and Support Vector Machines, the proposed feature set outperform the baseline feature set in helping detect fraudulent loan requests. Although the performance (e.g., Recall or Sensitivity) is still not up to its optimum, this study demonstrates that by analyzing the transaction records of confirmed malevolent borrowers, it is possible to capture some useful behavioral patterns for fraud detection. Such features and methods would possibly help lenders identify loan request frauds and avoid financial losses.
Jennifer Jie Xu 0001, Michael Chau
ISI3
2013 Using 3D virtual environments to facilitate students in constructivist learning
Michael Chau, Ada Wong, Minhong Wang 0001, Songnia Lai, Kristal W. Y. Chan, Tim M. H. Li, Debbie Chu, Ian K. W. Chan, Wai-Ki Sung
Decis. Support Syst.1
2011 Enterprise risk and security management: Data, text and Web mining
Hsinchun Chen, Michael Chau, Shu-Hsing Li
Decis. Support Syst.2
2011 Metric and trigonometric pruning for clustering of uncertain data in 2D geometric space
Wang Kay Ngai, Ben Kao, Reynold Cheng, Michael Chau, Sau Dan Lee, David Wai-Lok Cheung, Kevin Y. Yip
Inf. Syst.4
2010 PutMode: prediction of uncertain trajectories in moving objects databases
Shaojie Qiao, Changjie Tang, Huidong Jin 0001, Shucheng Dai, Yungchang Ku, Michael Chau
Appl. Intell.7
2010 Designing the user interface and functions of a search engine development tool
Michael Chau, Cho Hung Wong
Decis. Support Syst.1
2010 Evaluating the use of search engine development tools in IT education
abstract
Abstract It is important for education in computer science and information systems to keep up to date with the latest development in technology. With the rapid development of the Internet and the Web, many schools have included Internet‐related technologies, such as Web search engines and e‐commerce, as part of their curricula. Previous research has shown that it is effective to use search engine development tools to facilitate students' learning. However, the effectiveness of these tools in the classroom has not been evaluated. In this article, we review the design of three search engine development tools, SpidersRUs, Greenstone, and Alkaline, followed by an evaluation study that compared the three tools in the classroom. In the study, 33 students were divided into 13 groups and each group used the three tools to develop three independent search engines in a class project. Our evaluation results showed that SpidersRUs performed better than the two other tools in overall satisfaction and the level of knowledge gained in their learning experience when using the tools for a class project on Internet applications development.
Michael Chau, Cho Hung Wong, Yilu Zhou, Jialun Qin, Hsinchun Chen
J. Assoc. Inf. Sci. Technol.1
2010 Guest Editors' Introduction: Special Section on Mining Large Uncertain and Probabilistic Databases
abstract
The four papers in this special section were selected from 23 submissions and represent recent advances in the mining of uncertain databases. The works present new techniques for mining patterns, clustering, and ranking on uncertain data.
Reynold Cheng, Michael Chau, Minos N. Garofalakis, Jeffrey Xu Yu
IEEE Trans. Knowl. Data Eng.2
2009 Characteristics of character usage in Chinese Web searching
Michael Chau, Christopher C. Yang
Inf. Process. Manag.1
2008 A machine learning approach to web page filtering using content and structure analysis
Michael Chau, Hsinchun Chen
Decis. Support Syst.1
2008 SpidersRUs: Creating specialized search engines in multiple languages
Michael Chau, Jialun Qin, Yilu Zhou, Chunju Tseng, Hsinchun Chen
Decis. Support Syst.1
2007 Mining communities and their relationships in blogs: A study of online hate groups
Michael Chau, Jennifer Jie Xu 0001
Int. J. Hum. Comput. Stud.1
2007 Web searching in Chinese: A study of a search engine in Hong Kong
abstract
Abstract The number of non‐English resources has been increasing rapidly on the Web. Although many studies have been conducted on the query logs in search engines that are primarily English‐based (e.g., Excite and AltaVista), only a few of them have studied the information‐seeking behavior on the Web in non‐English languages. In this article, we report the analysis of the search‐query logs of a search engine that focused on Chinese. Three months of search‐query logs of Timway, a search engine based in Hong Kong, were collected and analyzed. Metrics on sessions, queries, search topics, and character usage are reported. N‐gram analysis also has been applied to perform character‐based analysis. Our analysis suggests that some characteristics identified in the search log, such as search topics and the mean number of queries per sessions, are similar to those in English search engines; however, other characteristics, such as the use of operators in query formulation, are significantly different. The analysis also shows that only a very small number of unique Chinese characters are used in search queries. We believe the findings from this study have provided some insights into further research in non‐English Web searching.
Michael Chau, Christopher C. Yang
J. Assoc. Inf. Sci. Technol.1
2007 Redips: Backlink search and analysis on the Web for business intelligence analysis
abstract
Abstract The World Wide Web presents significant opportunities for business intelligence analysis as it can provide information about a company's external environment and its stakeholders. Traditional business intelligence analysis on the Web has focused on simple keyword searching. Recently, it has been suggested that the incoming links, or backlinks, of a company's Web site (i.e., other Web pages that have a hyperlink pointing to the company of interest) can provide important insights about the company's “online communities.” Although analysis of these communities can provide useful signals for a company and information about its stakeholder groups, the manual analysis process can be very time‐consuming for business analysts and consultants. In this article, we present a tool called Redips that automatically integrates backlink meta‐searching and text‐mining techniques to facilitate users in performing such business intelligence analysis on the Web. The architectural design and implementation of the tool are presented in the article. To evaluate the effectiveness, efficiency, and user satisfaction of Redips, an experiment was conducted to compare the tool with two popular business intelligence analysis methods—using backlink search engines and manual browsing. The experiment results showed that Redips was statistically more effective than both benchmark methods (in terms of Recall and F‐measure) but required more time in search tasks. In terms of user satisfaction, Redips scored statistically higher than backlink search engines in all five measures used, and also statistically higher than manual browsing in three measures.
Michael Chau, Boby Shiu, Ivy Chan, Hsinchun Chen
J. Assoc. Inf. Sci. Technol.1
2007 Automated criminal link analysis based on domain knowledge
abstract
Abstract Link (association) analysis has been used in the criminal justice domain to search large datasets for associations between crime entities in order to facilitate crime investigations. However, link analysis still faces many challenging problems, such as information overload, high search complexity, and heavy reliance on domain knowledge. To address these challenges, this article proposes several techniques for automated, effective, and efficient link analysis. These techniques include the co‐occurrence analysis, the shortest path algorithm, and a heuristic approach to identifying associations and determining their importance. We developed a prototype system called CrimeLink Explorer based on the proposed techniques. Results of a user study with 10 crime investigators from the Tucson Police Department showed that our system could help subjects conduct link analysis more efficiently than traditional single‐level link analysis tools. Moreover, subjects believed that association paths found based on the heuristic approach were more accurate than those found based solely on the co‐occurrence analysis and that the automated link analysis system would be of great help in crime investigations.
Jennifer Schroeder, Jennifer Jie Xu 0001, Hsinchun Chen, Michael Chau
J. Assoc. Inf. Sci. Technol.4
2007 ServiceFinder: A method towards enhancing service portals
abstract
The rapid advancement of Internet technologies enables more and more educational institutes, companies, and government agencies to provide services, namely online services, through web portals. With hundreds of online services provided through a web portal, it is critical to design web portals, namely service portals, through which online services can be easily accessed by their consumers. This article addresses this critical issue from the perspective of service selection, that is, how to select a small number of service-links (i.e., hyperlinks pointing to online services) to be featured in the homepage of a service portal such that users can be directed to find the online services they seek most effectively. We propose a mathematically formulated metric to measure the effectiveness of the selected service-links in directing users to locate their desired online services and formally define the service selection problem. A solution method, ServiceFinder, is then proposed. Using real-world data obtained from the Utah State Government service portal, we show that ServiceFinder outperforms both the current practice of service selection and previous algorithms for adaptive website design. We also show that the performance of ServiceFinder is close to that of the optimal solution resulting from exhaustive search.
Olivia R. Liu Sheng, Michael Chau
ACM Trans. Inf. Syst.3
2007 Incorporating Web Analysis Into Neural Networks: An Example in Hopfield Net Searching
abstract
Neural networks have been used in various applications on the World Wide Web, but most of them only rely on the available input-output examples without incorporating Web-specific knowledge, such as Web link analysis, into the network design. In this paper, we propose a new approach in which the Web is modeled as an asymmetric Hopfield Net. Each neuron in the network represents a Web page, and the connections between neurons represent the hyperlinks between Web pages. Web content analysis and Web link analysis are also incorporated into the model by adding a page content score function and a link score function into the weights of the neurons and the synapses, respectively. A simulation study was conducted to compare the proposed model with traditional Web search algorithms, namely, a breadth-first search and a best-first search using PageRank as the heuristic. The results showed that the proposed model performed more efficiently and effectively in searching for domain-specific Web pages. We believe that the model can also be useful in other Web applications such as Web page clustering and search result ranking
Michael Chau, Hsinchun Chen
IEEE Trans. Syst. Man Cybern. Part C1
2006 Efficient Clustering of Uncertain Data
abstract
We study the problem of clustering data objects whose locations are uncertain. A data object is represented by an uncertainty region over which a probability density function (pdf) is defined. One method to cluster uncertain objects of this sort is to apply the UK-means algorithm, which is based on the traditional K-means algorithm. In UK-means, an object is assigned to the cluster whose representative has the smallest expected distance to the object. For arbitrary pdf, calculating the expected distance between an object and a cluster representative requires expensive integration computation. We study various pruning methods to avoid such expensive expected distance calculation.
Wang Kay Ngai, Ben Kao, Chun Kit Chui, Reynold Cheng, Michael Chau, Kevin Y. Yip
ICDM5
2006 Uncertain Data Mining: An Example in Clustering Location Data
Michael Chau, Reynold Cheng, Ben Kao, Jackey Ng
PAKDD1
2006 Building a scientific knowledge web portal: The NanoPort experience
Michael Chau, Zan Huang, Jialun Qin, Yilu Zhou, Hsinchun Chen
Decis. Support Syst.1
2006 Process-driven collaboration support for intra-agency crime analysis
J. Leon Zhao, Henry H. Bi, Hsinchun Chen, Daniel Dajun Zeng, Chienting Lin, Michael Chau
Decis. Support Syst.6
2006 Multilingual Web retrieval: An experiment in English-Chinese business intelligence
abstract
Abstract As increasing numbers of non‐English resources have become available on the Web, the interesting and important issue of how Web users can retrieve documents in different languages has arisen. Cross‐language information retrieval (CLIR), the study of retrieving information in one language by queries expressed in another language, is a promising approach to the problem. Cross‐language information retrieval has attracted much attention in recent years. Most research systems have achieved satisfactory performance on standard Text REtrieval Conference (TREC) collections such as news articles, but CLIR techniques have not been widely studied and evaluated for applications such as Web portals. In this article, the authors present their research in developing and evaluating a multilingual English–Chinese Web portal that incorporates various CLIR techniques for use in the business domain. A dictionary‐based approach was adopted and combines phrasal translation, co‐occurrence analysis, and pre‐ and posttranslation query expansion. The portal was evaluated by domain experts, using a set of queries in both English and Chinese. The experimental results showed that co‐occurrence‐based phrasal translation achieved a 74.6% improvement in precision over simple word‐by‐word translation. When used together, pre‐ and posttranslation query expansion improved the performance slightly, achieving a 78.0% improvement over the baseline word‐by‐word translation approach. In general, applying CLIR techniques in Web applications shows promise.
Jialun Qin, Yilu Zhou, Michael Chau, Hsinchun Chen
J. Assoc. Inf. Sci. Technol.3
2005 Visualizing criminal relationships: comparison of a hyperbolic tree and a hierarchical list
Michael Chau, Homa Atabakhsh, Hsinchun Chen
Decis. Support Syst.2
2005 Analysis of the query logs of a Web site search engine
abstract
Abstract A large number of studies have investigated the transaction log of general‐purpose search engines such as Excite and AltaVista, but few studies have reported on the analysis of search logs for search engines that are limited to particular Web sites, namely, Web site search engines. In this article, we report our research on analyzing the search logs of the search engine of the Utah state government Web site. Our results show that some statistics, such as the number of search terms per query, of Web users are the same for general‐purpose search engines and Web site search engines, but others, such as the search topics and the terms used, are considerably different. Possible reasons for the differences include the focused domain of Web site search engines and users' different information needs. The findings are useful for Web site developers to improve the performance of their services provided on the Web and for researchers to conduct further research in this area. The analysis also can be applied in e‐government research by investigating how information should be delivered to users in government Web sites.
Michael Chau, Olivia R. Liu Sheng
J. Assoc. Inf. Sci. Technol.1
2004 3D Geographic Visualization: The Marine GIS
Christopher M. Gold, Michael Chau, Marcin Dzieszko, Rafel Goralski
SDH2
2003 COPLINK Agent: An Architecture for Information Monitoring and Sharing in Law Enforcement
Daniel Dajun Zeng, Hsinchun Chen, Damien Daspit, Fu Shan, Suresh Nandiraju, Michael Chau, Chienting Lin
ISI6
2003 Design and evaluation of a multi-agent collaborative Web mining system
Michael Chau, Daniel Dajun Zeng, Hsinchun Chen, David Hendriawan
Decis. Support Syst.1
2003 Testing a Cancer Meta Spider
Hsinchun Chen, Haiyan Fan, Michael Chau, Daniel Dajun Zeng
Int. J. Hum. Comput. Stud.3
2003 elpfulMed: Intelligent searching for medical information over the internet
abstract
Abstract Medical professionals and researchers need information from reputable sources to accomplish their work. Unfortunately, the Web has a large number of documents that are irrelevant to their work, even those documents that purport to be “medically‐related.” This paper describes an architecture designed to integrate advanced searching and indexing algorithms, an automatic thesaurus, or “concept space,” and Kohonen‐based Self‐Organizing Map (SOM) technologies to provide searchers with fine‐grained results. Initial results indicate that these systems provide complementary retrieval functionalities. HelpfulMed not only allows users to search Web pages and other online databases, but also allows them to build searches through the use of an automatic thesaurus and browse a graphical display of medical‐related topics. Evaluation results for each of the different components are included. Our spidering algorithm outperformed both breadth‐first search and PageRank spiders on a test collection of 100,000 Web pages. The automatically generated thesaurus performed as well as both MeSH and UMLS—systems which require human mediation for currency. Lastly, a variant of the Kohonen SOM was comparable to MeSH terms in perceived cluster precision and significantly better at perceived cluster recall.
Hsinchun Chen, Ann M. Lally, Bin Zhu 0001, Michael Chau
J. Assoc. Inf. Sci. Technol.4
2003 Teaching key topics in computer science and information systems through a web search engine project
abstract
Advances in computer and Internet technologies have made it more and more important for information technology professionals to acquire experience in a variety of aspects, including new technologies, system integration, database administration, and project management. To provide students with a chance to acquire such skills, we designed a project called "Build Your Search Engine in 90 Days," in which students were required to build a domain-specific Web search engine in a semester. In this paper we review the tools and resources available to students and report our experiences in having students to work on this project in a course at the University of Arizona. We also review two tools, called AI Spider and AI Indexer, we developed for students in this project. We highlight a few search engines that were created by the students and suggest some future directions in improving the tools and expanding the project.
Michael Chau, Zan Huang, Hsinchun Chen
ACM J. Educ. Resour. Comput.1
2002 CI Spider: a tool for competitive intelligence on the Web
Hsinchun Chen, Michael Chau, Daniel Dajun Zeng
Decis. Support Syst.2
2001 MetaSpider: Meta-searching and categorization on the Web
abstract
Abstract It has become increasingly difficult to locate relevant information on the Web, even with the help of Web search engines. Two approaches to addressing the low precision and poor presentation of search results of current search tools are studied: meta‐search and document categorization. Meta‐search engines improve precision by selecting and integrating search results from generic or domain‐specific Web search engines or other resources. Document categorization promises better organization and presentation of retrieved results. This article introduces MetaSpider, a meta‐search engine that has real‐time indexing and categorizing functions. We report in this paper the major components of MetaSpider and discuss related technical approaches. Initial results of a user evaluation study comparing MetaSpider, NorthernLight, and MetaCrawler in terms of clustering performance and of time and effort expended show that MetaSpider performed best in precision rate, but disclose no statistically significant differences in recall rate and time requirements. Our experimental study also reveals that MetaSpider exhibited a higher level of automation than the other two systems and facilitated efficient searching by providing the user with an organized, comprehensive view of the retrieved documents.
Hsinchun Chen, Haiyan Fan, Michael Chau, Daniel Dajun Zeng
J. Assoc. Inf. Sci. Technol.3