Bi-Ru Dai

dblp:16/1480 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
4since 2021 · last 2026
0000-0003-4791-9914ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 24 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-authorComputer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Overlapping Community-aware Social Recommendation via Graph Attention Network Method
abstract
Social influence plays a crucial role in shaping user preferences and behaviors, making social recommendation an effective approach for alleviating the cold-start problem. However, most existing social recommendation methods either model social influence at the individual level or assume non-overlapping community structures, which fails to reflect the fact that users typically belong to multiple communities simultaneously. As a result, the influence from different communities and their cross-community interactions are not explicitly modeled. In this article, we formally study the problem of overlapping community-aware social recommendation, where a user’s preference is jointly influenced by personal behavior, social neighbors, and multiple overlapping communities, each contributing differently depending on the target item. To address this problem, we propose GANOC, a unified framework that decomposes user preference into three complementary domains: personal, social, and community. We employ graph attention networks to model influence propagation in both social and community graphs and design an item-aware attention mechanism to selectively aggregate cross-community influences. Furthermore, a domain attention network is introduced to adaptively integrate preference representations from different domains for rating prediction. Extensive experiments on three real-world benchmark datasets demonstrate that GANOC outperforms state-of-the-art social recommendation methods, particularly for cold-start users, validating the effectiveness of explicitly modeling overlapping community influence.
Bi-Ru Dai, Pao-Yun Ma
ACM Trans. Knowl. Discov. Data1
2024 STL-ConvTransformer: Series Decomposition and Convolution-Infused Transformer Architecture in Multivariate Time Series Anomaly Detection
Yu-Xiang Wu, Bi-Ru Dai
PAKDD (1)2
2022 A Greedy Algorithm for Budgeted Multiple-Product Profit Maximization in Social Network
abstract
Due to the development of technology in recent years, the propagation of information has become very conve-nient. People can easily disseminate information, such as product advertisements on the Internet, and this way of attracting users on social networks is called viral marketing. Because companies usually launch a variety of similar products for consumers to choose from, finding the most potential influencer for products with different prices and a limited seed budget is a very hot topic. However, the purchasing ability of a consumer is generally assumed to be unrestricted in some related studies. This assumption does not conform to the business model in reality since the demand for a product is usually limited for a customer. Considering the above situation, we study a novel problem, called the Budgeted Multiple-Product Profit Maximization (BMPM) problem. We will find suitable seeds for multiple products to maximize the overall profit when the overall budget is limited. Therefore, we propose a greedy algorithm named BG and a new graph structure PWDAG to effectively solve the BMPM problem. Our experiments on real world datasets show that BG can maximize the overall profit with high efficiency and effectiveness.
Chun-Cheng Fang, Chia-Chun Ho, Bi-Ru Dai
MDM3
2021 Semantic Analysis and Preference Capturing on Attentive Networks for Rating Prediction
abstract
Nowadays, people receive an enormous amount of information from day to day. However, they are only interested in information which matches their preferences. Thus, retrieving such information becomes an significant task, in our case, the reviews composed by users. Matrix Factorization (MF) based methods achieve fairly good performances on recommendation tasks. However, there exist several crucial issues with MF - based methods such as cold-start problems and data sparseness. In order to address the above issues, numerous recommendation models are proposed which obtained stellar performances. Nonetheless, we figured that there is not a more comprehensive framework that enhances its performance through retrieving user preference and item trend. Hence, we propose a novel approach to tackle the aforementioned issues. A hierarchical construction with user preference and item trend capturing is employed in this proposed framework. The performance excels in comparison to state-of-the-art models by testing on several real-world datasets. Experimental results verified that our framework can extract useful features even under sparse data.
Cheng-Han Chou, Bi-Ru Dai
GLOBECOM2
2020 Multi-Laplacian GAN with Edge Enhancement for Face Super Resolution
abstract
Face image super-resolution has become a research hotspot in the field of image processing. Nowadays, more and more researches add additional information, such as landmark, identity, to reconstruct high resolution images from low resolution ones, and have a good performance in quantitative terms and perceptual quality. However, these additional information is hard to obtain in many cases. In this work, we focus on reconstructing face images by extracting useful information from face images directly rather than using additional information. By observing edge information in each scale of face images, we propose a method to reconstruct high resolution face images with enhanced edge information. In additional, with the proposed training procedure, our method reconstructs photo-realistic images in upscaling factor 8× and outperforms state-of-the-art methods both in quantitative terms and perceptual quality.
Shanlei Ko, Bi-Ru Dai
ICPR2
2020 Live Stream Highlight Detection Using Chat Messages
abstract
In recent years, live-streaming services have been booming and are still continuing to grow on the Internet. Differing from TV shows and movies, live-streaming can have variable and longer lengths with no specific content restrictions. Traditional methods of video highlight detection, which are based on visual features, will suffer the difficulties of data scale and inconsistency. To address these issues, we alternatively extract information from the audience discussion in a chat room for high-light detection. In this paper, an attention-based model, LSTA, is proposed to integrate the long term and short term information in a chat room to determine which fragments should be identified as highlights. Our results demonstrate the improvement over both state-of-the-art visual and textual content-based approaches.
Chieh-Ming Liaw, Bi-Ru Dai
MDM2
2018 A Multi-Label Threshold Learning Framework for Propagation Algorithms on a Non-Feature Network
abstract
Recently, with the exponential growth of network data, collecting whole features correctly is time- consuming and expensive. For a classification problem on networks, traditional propagation algorithms, which rely on the feature information to build transition matrix to propagate label information on networks, generally do not perform well when the feature information is not available. Our observation shows that the problem of minority ignorance occurs on the propagation process of traditional algorithms. In this paper, we propose a LPBC framework to allow a propagation algorithm to deal with multi-label classification problem on networks. With a novel threshold training process, LPBC reduces the minority ignorance when the label information is propagated. Experimental results demonstrated the effectiveness and the performance improvement of the proposed framework.
Ye-Yan Zeng, Bi-Ru Dai
GLOBECOM2
2017 Reweighting Forest for Extreme Multi-label Classification
Zhun-Zheng Lin, Bi-Ru Dai
DaWaK2
2017 Accelerating K-Means by Grouping Points Automatically
Qiao Yu 0003, Bi-Ru Dai
DaWaK2
2016 A Framework of the Semi-supervised Multi-label Classification with Non-uniformly Distributed Incomplete Labels
Chih-Heng Chung, Bi-Ru Dai
DaWaK2
2016 Power of Bosom Friends, POI Recommendation by Learning Preference of Close Friends and Similar Users
Mu-Yao Fang, Bi-Ru Dai
DaWaK2
2016 A G-Means Update Ensemble Learning Approach for the Imbalanced Data Stream with Concept Drifts
Sin-Kai Wang, Bi-Ru Dai
DaWaK2
2015 A Model of Relevant Common Author and Citation Authority Propagation for Citation Recommendation
abstract
In academia, as more and more papers are published, it is difficult to search needed papers quickly. Therefore, scholarly search engines and many assessment methods have begun to appear. We propose Related Paper with Common Author method which can effectively filter search results to provide potential recommendation papers without the full text, and a citation network method Citation Authority Propagation method which is a recommending method that uses authority author propagation. The experiment results demonstrate the effectiveness of the proposed system.
Bo-Yu Hsiao, Chih-Heng Chung, Bi-Ru Dai
MDM (2)3
2015 A Weighted Distance Similarity Model to Improve the Accuracy of Collaborative Recommender System
abstract
Collaborative filtering is one of the most widely used methods to provide product recommendation in online stores. The key component of the method is to find similar users or items by using user-item matrix so that products can be recommended based on the similarities. However, traditional collaborative filtering approaches compute the similarity between a target user and the other user without considering a target item. More specifically, they give an equal weight to each of the items which are rated by both users. However, we think that the similarity between the target item and each of the co-rated items is a very important factor when we calculate the similarity between two users. Therefore, in this paper we propose a new similarity function that takes similarities between a target item and each of the co-rated items and the proportion of common ratings into account. Experimental results from Movie Lens dataset show that the method improves accuracy of recommender system significantly.
Bing-Hao Huang, Bi-Ru Dai
MDM (2)2
2014 A fragment-based iterative consensus clustering algorithm with a robust similarity
Chih-Heng Chung, Bi-Ru Dai
Knowl. Inf. Syst.2
2013 Opinion Mining on Social Media Data
abstract
Microblogging (Twitter or Facebook) has become a very popular communication tool among Internet users in recent years. Information is generated and managed through either computer or mobile devices by one person and is consumed by many other persons, with most of this user-generated content being textual information. As there are a lot of raw data of people posting real time messages about their opinions on a variety of topics in daily life, it is a worthwhile research endeavor to collect and analyze these data, which may be useful for users or managers to make informed decisions, for example. However this problem is challenging because a micro-blog post is usually very short and colloquial, and traditional opinion mining algorithms do not work well in such type of text. Therefore, in this paper, we propose a new system architecture that can automatically analyze the sentiments of these messages. We combine this system with manually annotated data from Twitter, one of the most popular microblogging platforms, for the task of sentiment analysis. In this system, machines can learn how to automatically extract the set of messages which contain opinions, filter out nonopinion messages and determine their sentiment directions (i.e. positive, negative). Experimental results verify the effectiveness of our system on sentiment analysis in real microblogging applications.
Po-Wei Liang, Bi-Ru Dai
MDM (2)2
2012 Efficient Map/Reduce-Based DBSCAN Algorithm with Optimized Data Partition
abstract
DBSCAN is a well-known algorithm for density-based clustering because it can identify the groups of arbitrary shapes and deal with noisy datasets. However, with the increasing amount of data, DBSCAN algorithm running on a single machine has to face the scalability problem. In this paper, we propose a Map/Reduce-based DBSCAN algorithm called DBSCAN-MR to solve the scalability problem. In DBSCAN-MR, the input dataset is partitioned into smaller parts and then parallel processed on the Hadoop platform. However, choosing different partition mechanisms will affect the execution efficiency and load balance of each node. Therefore, we propose a method, partition with reduce boundary points (PRBP), to select partition boundaries based on the distribution of data points. Our experimental results show that DBSCAN-MR with the design of PRBP has higher efficiency and scalability than competitors.
Bi-Ru Dai, I-Chang Lin
IEEE CLOUD1
2012 LF-CARS: A Loose Fragment-Based Consensus Clustering Algorithm with a Robust Similarity
Bi-Ru Dai, Chih-Heng Chung
Discovery Science1
2011 A Framework of Recommendation System Based on Both Network Structure and Messages
abstract
The evolving of Internet technology allows people to communicate even they are far away from each other. More and more people share information and exchange their thoughts via the communities on the websites and become friends. A larger community usually attracts more users, therefore, how to enhance the development of a social network on the website is an important issue for the survival of a website. In this paper, we combine the social network features into the recommendation system. In addition to messages between nodes, the features of network structure are taken into consideration. Experimental results show that the recommendation accuracy of our method is higher than the existing method which is based on the message ratio.
Bi-Ru Dai, Chang-Yi Lee, Chih-Heng Chung
ASONAM1
2011 An Instance Selection Algorithm Based on Reverse Nearest Neighbor
Bi-Ru Dai, Shu-Ming Hsu
PAKDD (1)1
2011 Improved inverse halftoning using vector and texture-lookup table-based learning approach
Yong-Huai Huang, Kuo-Liang Chung, Bi-Ru Dai
Expert Syst. Appl.3
2009 iTM: An Efficient Algorithm for Frequent Pattern Mining in the Incremental Database without Rescanning
Bi-Ru Dai, Pai-Yu Lin
IEA/AIE1
2009 A Decision Tree Based Quasi-Identifier Perturbation Technique for Preserving Privacy in Data Mining
abstract
Classification is an important issue in data mining, and decision tree is one of the most popular techniques for classification analysis. Some data sources contain private personal information that people are unwilling to reveal. The disclosure of person-specific data is possible to endanger thousands of people, and therefore the dataset should be protected before it is released for mining. However, techniques to hide private information usually modify the original dataset without considering influences on the prediction accuracy of a classification model. In this paper, we propose an algorithm to protect personal privacy for classification model based on decision tree. Our goal is to hide all person-specific information with minimized data perturbation. Furthermore, the prediction capability of the decision tree classifier can be maintained. As demonstrated in the experiments, the proposed algorithm can successfully hide private information with fewer disturbances of the classifier.
Bi-Ru Dai, Yang-Tze Lin
RCIS1
2008 Hiding Frequent Patterns under Multiple Sensitive Thresholds
Ya-Ping Kuo, Pai-Yu Lin, Bi-Ru Dai
DEXA3
2007 Incremental Clustering in Geography and Optimization Spaces
Chih-Hua Tai, Bi-Ru Dai, Ming-Syan Chen
PAKDD2
2007 Twain: Two-end association miner with precise frequent exhibition periods
abstract
We investigate the general model of mining associations in a temporal database, where the exhibition periods of items are allowed to be different from one to another. The database is divided into partitions according to the time granularity imposed. Such temporal association rules allow us to observe short-term but interesting patterns that are absent when the whole range of the database is evaluated altogether. Prior work may omit some temporal association rules and thus have limited practicability. To remedy this and to give more precise frequent exhibition periods of frequent temporal itemsets, we devise an efficient algorithm Twain (standing for TWo end AssocIation miNer .) Twain not only generates frequent patterns with more precise frequent exhibition periods, but also discovers more interesting frequent patterns. Twain employs Start time and End time of each item to provide precise frequent exhibition period while progressively handling itemsets from one partition to another. Along with one scan of the database, Twain can generate frequent 2-itemsets directly according to the cumulative filtering threshold. Then, Twain adopts the scan reduction technique to generate all frequent k -itemsets ( k > 2) from the generated frequent 2-itemsets. Theoretical properties of Twain are derived as well in this article. The experimental results show that Twain outperforms the prior works in the quality of frequent patterns, execution time, I/O cost, CPU overhead and scalability.
Jen-Wei Huang, Bi-Ru Dai, Ming-Syan Chen
ACM Trans. Knowl. Discov. Data2
2007 Clustering over Multiple Evolving Streams by Events and Correlations
abstract
In applications of multiple data streams such as stock market trading and sensor network data analysis, the clusters of streams change at different time because of the data evolution. The information of evolving cluster is valuable to support corresponding online decisions. In this paper, we present a framework for Clustering Over Multiple Evolving sTreams by CORrelations and Events, which, abbreviated as COMETCORE, monitors the distribution of clusters over multiple data streams based on their correlation. Instead of directly clustering the multiple data streams periodically, COMET-CORE applies efficient cluster split and merge processes only when significant cluster evolution happens. Accordingly, we devise an event detection mechanism to signal the cluster adjustments. The coming streams are smoothed as sequences of end points by employing piecewise linear approximation. At the time when end points are generated, weighted correlations between streams are updated. End points are good indicators of significant change in streams, and this is a main cause of cluster evolution event. When an event occurs, through split and merge operations we can report the latest clustering results. As shown in our experimental studies, COMET-CORE can be performed effectively with good clustering quality.
Mi-Yen Yeh, Bi-Ru Dai, Ming-Syan Chen
IEEE Trans. Knowl. Data Eng.2
2007 Constrained data clustering by depth control and progressive constraint relaxation
Bi-Ru Dai, Cheng-Ru Lin, Ming-Syan Chen
VLDB J.1
2006 COMET: Event-Driven Clustering over Multiple Evolving Streams
Mi-Yen Yeh, Bi-Ru Dai, Ming-Syan Chen
PAKDD2
2006 Adaptive Clustering for Multiple Evolving Streams
abstract
In the data stream environment, the patterns generated at different time instances are different due to data evolution. As time progresses, the behavior and members of clusters usually change. Hence, clustering continuous data streams allows us to observe the changes of group behavior. In order to support flexible clustering requirements, we devise in this paper a Clustering on Demand framework, abbreviated as COD framework, to dynamically cluster multiple data streams. While providing a general framework of clustering on multiple data streams, the COD framework has two advantageous features, namely, one data scan for online statistics collection and compact multiresolution approximations, which are designed to address, respectively, the time and the space constraints in a data stream environment. The COD framework consists of two phases, i.e., the online maintenance phase and the offline clustering phase. The online maintenance phase provides an efficient mechanism to maintain summary hierarchies of data streams with multiple resolutions in time linear in both the number of streams and the number of data points in each stream. On the other hand, an adaptive clustering algorithm is devised for the offline phase to retrieve approximations of desired substreams from summary hierarchies according to clustering queries. We propose two summarization techniques, based on wavelet and regression analyses, to construct the summary hierarchies. The regression-based summary hierarchy approximates the data stream more precisely and provides better clustering results, at the cost of slightly longer time than and twice the storage space as the wavelet-based one. An adaptive version of COD framework is designed to make a selection between a wavelet-based model and a regression-based model for building the summary hierarchy. By the adaptive COD, we can obtain clustering results with almost the same quality as the regression-based COD while using much less storage space for the summary hierarchy. As shown in the complexity analyses and also validated by our empirical studies, the COD framework performs very efficiently in the data stream environment while producing clustering results of very high quality.
Bi-Ru Dai, Jen-Wei Huang, Mi-Yen Yeh, Ming-Syan Chen
IEEE Trans. Knowl. Data Eng.1
2004 Clustering on Demand for Multiple Data Streams
abstract
In the data stream environment, the patterns generated by the mining techniques are usually distinct at different time because of the evolution of data. In order to deal with various types of multiple data streams and to support flexible mining requirements, we devise in this paper a clustering on demand framework, abbreviated as COD framework, to dynamically cluster multiple data streams. While providing a general framework of clustering on multiple data streams, the COD framework has two major features, namely one data scan for online statistics collection and compact multiresolution approximations, which are designed to address, respectively, the time and the space constraints in a data stream environment. Furthermore, with the multiresolution approximations of data streams, flexible clustering demands can be supported.
Bi-Ru Dai, Jen-Wei Huang, Mi-Yen Yeh, Ming-Syan Chen
ICDM1
2003 On the Techniques for Data Clustering with Numerical Constraints
abstract
In this paper, the attributes employed to model the constraints are called constraint attributes and those attributes involved in the objective function to be optimized are called cost-optimal attributes. The constrained clustering considered is conducted in such a way that the objective function of cost-optimal attributes is optimized subject to the condition that the imposed constraint is satisfied. Explicitly, we address the problem of constrained clustering with numerical constraints, in which the constraint attribute values of any two data items in the same cluster are required to be within the corresponding constraint range. We devise an effective and efficient algorithm with complete-link to solve this clustering problem. It is noted that due to the intrinsic nature of the numerical constrained clustering, there is an order dependency on the process of attaining the clustering, which in many cases degrades the clustering results. In view of this, we devise a progressive constraint relaxation technique to remedy this drawback and improve the overall performance of clustering results. Explicitly, by using a smaller (tighter) constraint range in earlier iterations of merge, we will have more room to relax the constraint and seek for better solutions in subsequent iterations. It is empirically shown that the progressive constraint relaxation technique is able to improve not only the execution efficiency but also the clustering quality.
Bi-Ru Dai, Cheng-Ru Lin, Ming-Syan Chen
SDM1