EDBT 2026 Demo / reviewers in the wild / expert
Korris Fu-Lai Chung
dblp:c/KFLChung · also Fu-Lai Chung, Fulai Chung
· DBLP profile ↗
41ranked-venue papers in the field
1as first author
10since 2021 · last 2025
0000-0001-5294-8168ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 14Knowledge Engineering, Semantic Web & Information Systems · 11Data Mining & Knowledge Discovery · 10 (1 first)Database Systems & Data Management · 3Other / Interdisciplinary · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A calibrated fully interpretable fuzzy classifier via Vapnik-Chervonenkis-dimension minimization learning
Korris Fu-Lai Chung, Yusuke Nojima, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2025 | A fully interpretable stacking fuzzy classifier with stochastic configuration-based learning for high-dimensional data
Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2023 | Timestamps as Prompts for Geography-Aware Location RecommendationabstractLocation recommendation plays a vital role in improving users' travel experience. The timestamp of the POI to be predicted is of great significance, since a user will go to different places at different times. However, most existing methods either do not use this kind of temporal information, or just implicitly fuse it with other contextual information. In this paper, we revisit the problem of location recommendation and point out that explicitly modeling temporal information is a great help when the model needs to predict not only the next location but also further locations. In addition, state-of-the-art methods do not make effective use of geographic information and suffer from the hard boundary problem when encoding geographic information by gridding. To this end, a Temporal Prompt-based and Geography-aware (TPG) framework is proposed. The temporal prompt is firstly designed to incorporate temporal information of any further check-in. A shifted window mechanism is then devised to augment geographic data for addressing the hard boundary problem. Via extensive comparisons with existing methods and ablation studies on five real-world datasets, we demonstrate the effectiveness and superiority of the proposed method under various settings. Most importantly, our proposed model has the superior ability of interval prediction. In particular, the model can predict the location that a user wants to go to at a certain time while the most recent check-in behavioral data is masked, or it can predict specific future check-in (not just the next one) at a given timestamp. Haoyi Duan, Ye Liu 0002, Korris Fu-Lai Chung |
CIKM | 4 |
| 2023 | Fuzzy style flat-based clusteringabstractThe recently developed fuzzy style k -plane clustering (S-KPC) algorithm displays promising clustering quality by leveraging both similarities and distinguishable styles between samples on stylistic data. However, S-KPC becomes vulnerable to similar styles that are not easily distinguishable. In this study, a novel f uzzy s tyle f lat-based c lustering (FSFC) algorithm is proposed to overcome this vulnerability. In FSFC, a style flat matrix (SFM) is designed to project samples onto appropriate flats while maintaining the styles of different clusters in a reasonable manner. Based on SFM, the core of FSFC is to learn the potentially intersecting manifold structures of clusters in the projected flat space to make samples with the same style close to the cluster center and simultaneously far away from the other cluster centers. Furthermore, the objective function of FSFC can provide scale flexibility for each flat in the projected flat space. In particular, the optimization problem of FSFC can be decomposed into a series of sub-problems about the flat parameters, which can be locally optimized using the concave-convex procedure (CCCP). Extensive experiments on both synthetic and real-world datasets demonstrate the competitive clustering performance of FSFC. Moreover, FSFC outperforms some state-of-the-art manifold clustering algorithms on six case studies about stylistic data. Suhang Gu, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2023 | MetaGC-MC: A graph-based meta-learning approach to cold-start recommendation with/without auxiliary information
Honglin Shu, Korris Fu-Lai Chung |
Inf. Sci. | 2 |
| 2023 | Improving Generalizability of Graph Anomaly Detection Models via Data AugmentationabstractGraph anomaly detection (GAD) has wide applications in real-world networked systems. In many scenarios, people need to identify anomalies on new (sub)graphs, but they may lack labels to train an effective detection model. Since recent semi-supervised GAD methods, which can leverage the available labels as prior knowledge, have achieved superior performance than unsupervised methods, one natural idea is to directly adopt a trained semi-supervised GAD model to the new (sub)graphs for testing. However, we find that existing semi-supervised GAD methods suffer from poor generalization issues, i.e., well-trained models could not perform well on an unseen area (i.e., not accessible in training) of the graph. Motivated by this, we formally define the problem of generalized graph anomaly detection that aims to effectively identify anomalies on both the training-domain graph(s) and the unseen test graph(s). Nevertheless, it is a challenging task since only limited labels are available, and the normal data distribution may differ between training and testing data. Accordingly, we propose a data augmentation method namedAugAN(Augmentation forAnomaly andNormal distributions) to enrich training data and adopt a customized episodic training strategy for learning with the augmented data. Extensive experiments verify the effectiveness ofAugANin improving model generalizability. Shuang Zhou 0012, Xiao Huang 0001, Ninghao Liu 0001, Huachi Zhou, Korris Fu-Lai Chung, Long-Kai Huang |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Urban Region Profiling via Multi-Graph Representation LearningabstractProfiling urban regions is essential for urban analytics and planning. Although existing studies have made great efforts to learn urban region representation from multi-source urban data, there are still limitations on modelling local-level signals, developing an effective yet integrated fusion framework, and performing well in regions with high variance socioeconomic attributes. Thus, we propose a multi-graph representation learning framework, called Region2Vec, for urban region profiling. Specifically, except that human mobility is encoded for inter-region relations, geographic neighborhood is introduced for capturing geographical contextual information while POI side information is adopted for representing intra-region information. Then, graphs are used to capture accessibility, vicinity, and functionality correlations among regions. An encoder-decoder multi-graph fusion module is further proposed to jointly learn comprehensive representations. Experiments on real-world datasets show that Region2Vec can be employed in three applications and outperforms all state-of-the-art baselines. Particularly, Region2Vec has better performance than previous studies in regions with high variance socioeconomic attributes. Korris Fu-Lai Chung, Kai Chen 0005 |
CIKM | 2 |
| 2022 | Unseen Anomaly Detection on Networks via Multi-Hypersphere LearningabstractNetwork anomaly detection is a crucial task since a few anomalies can cause huge losses. Semi-supervised anomaly detection methods can effectively leverage a small number of labels as prior knowledge to enhance detection accuracy. But in real-world scenarios, novel types of anomalies (i.e., unseen anomalies) usually exist on networks which may present different characteristics with the seen anomalies and are hard to be identified by prior semi-supervised anomaly detection methods. In this paper, we propose the novel problem of unseen network anomaly detection that aims to identify both seen and unseen anomalies to eliminate potential dangers. Accordingly, we propose a method called Multi-hypersphere Graph Learning (MHGL) to effectively leverage existing labels by learning fine-grained normal patterns to discriminate anomalies. Experiments demonstrate that MHGL outperforms state-of-the-art methods significantly. Shuang Zhou 0012, Xiao Huang 0001, Ninghao Liu 0001, Qiaoyu Tan, Korris Fu-Lai Chung |
SDM | 5 |
| 2022 | Multi-view clustering by virtually passing mutually supervised smooth messages
Suhang Gu, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2021 | Subtractive Aggregation for Attributed Network Anomaly DetectionabstractAttributed network anomaly detection is essential in various networked systems. It aims to detect nodes that significantly deviate from their corresponding background. In conventional anomaly detection, the background is defined as the vast majority. But in networks, anomalies can be local and look normal when compared with the majority. While several efforts have explored to consider communities as the background, it remains challenging to learn suitable communities for effective anomaly detection. Also, the patterns of anomalies are unknown and it is nontrivial to define criteria of anomalies. To bridge the gap, in this paper, we argue that, by using appropriate models, it is sufficient to simply consider neighbor nodes as the background to detect anomalies. Correspondingly, we propose a novel abnormality-aware graph neural network (AAGNN). It utilizes subtractive aggregation to represent each node as the deviation from its neighbors (the background). Normal nodes with high confidence are employed as labels to learn a tailored hypersphere as the criterion of anomalies. Experiments demonstrate that AAGNN surpasses state-of-the-art methods significantly. Shuang Zhou 0012, Qiaoyu Tan, Xiao Huang 0001, Korris Fu-Lai Chung |
CIKM | 5 |
| 2020 | Towards Self-Tuning Parameter ServersabstractRecent years, many applications have been driven advances by the use of Machine Learning (ML). Nowadays, it is common to see industrial-strength machine learning jobs that involve millions of model parameters, terabytes of training data, and weeks of training. Good efficiency, i.e., fast completion time of running a specific ML training job, therefore, is a key feature of a successful ML system. While the completion time of a long-running ML job is determined by the time required to reach model convergence, that is also largely influenced by the values of various system settings. In this paper, we contribute techniques towards building self-tuning parameter servers. Parameter Server (PS) is a popular system architecture for large-scale machine learning systems; and by self-tuning we mean while a long-running ML job is iteratively training the expert-suggested model, the system is also iteratively learning which system setting is more efficient for that job and applies it online. Our techniques are general enough to various PS-style ML systems. Experiments on TensorFlow show that our techniques can reduce the completion times of a variety of long-running TensorFlow jobs from 1.4× to 18×. Chris Liu, Bo Tang 0016, Hang Shen 0001, Ziliang Lai, Eric Lo 0001, Korris Fu-Lai Chung |
IEEE BigData | 7 |
| 2019 | Leveraging Label-Specific Discriminant Mapping Features for Multi-Label LearningabstractAs an important machine learning task, multi-label learning deals with the problem where each sample instance (feature vector) is associated with multiple labels simultaneously. Most existing approaches focus on manipulating the label space, such as exploiting correlations between labels and reducing label space dimension, with identical feature space in the process of classification. One potential drawback of this traditional strategy is that each label might have its own specific characteristics and using identical features for all label cannot lead to optimized performance. In this article, we propose an effective algorithm named LSDM, i.e., leveraging label-specific discriminant mapping features for multi-label learning , to overcome the drawback. LSDM sets diverse ratio parameter values to conduct cluster analysis on the positive and negative instances of identical label. It reconstructs label-specific feature space which includes distance information and spatial topology information. Our experimental results show that combining these two parts of information in the new feature representation can better exploit the clustering results in the learning process. Due to the problem of diverse combinations for identical label, we employ simplified linear discriminant analysis to efficiently excavate optimal one for each label and perform classification by querying the corresponding results. Comparison with the state-of-the-art algorithms on a total of 20 benchmark datasets clearly manifests the competitiveness of LSDM. Yumeng Guo, Korris Fu-Lai Chung, Guo-Zheng Li 0001, Jiancong Wang, James C. Gee |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Stacked Robust Adaptively Regularized Auto-Regressions for Domain AdaptationabstractDomain adaptation is the situation for supervised learning in which the training data are sampled from the source domain while the test data are sampled from the target domain that follows a different distribution. The key to solving such a problem is to reduce effects of the discrepancy between the training data and test data. Recently, deep learning methods that employ stacked denoising auto-encoders (SDAs) to learn new representations for both domains have been successfully applied in domain adaptation. And, remarkable performance on multi-domain sentiment analysis datasets has been reported, making deep learning a promising approach to domain adaptation problems. In this paper, a deep learning method called Stacked Robust Adaptively Regularized Auto-regressions (SRARAs) is proposed to learn useful representations for domain adaptation problems. Each layer of SRARAs contains two steps: a linear transformation step, which is based on robust adaptively regularized auto-regression, and a non-linear squashing transformation step. The first step aims at reducing the discrepancy between the training data and test data, and the second step is to introduce non-linearity and control the range of the elements in the outputs. The experimental results on text and image datasets demonstrate that the proposed method is very effective. Hongchang Gao, Wei Lu 0006, Wei Liu 0005, Korris Fu-Lai Chung, Heng Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | A Deep Bayesian Tensor-Based System for Video RecommendationabstractWith the availability of abundant online multi-relational video information, recommender systems that can effectively exploit these sorts of data and suggest creatively interesting items will become increasingly important. Recent research illustrates that tensor models offer effective approaches for complex multi-relational data learning and missing element completion. So far, most tensor-based user clustering models have focused on the accuracy of recommendation. Given the dynamic nature of online media, recommendation in this setting is more challenging as it is difficult to capture the users’ dynamic topic distributions in sparse data settings as well as to identify unseen items as candidates of recommendation. Targeting at constructing a recommender system that can encourage more creativity, a deep Bayesian probabilistic tensor framework for tag and item recommendation is proposed. During the score ranking processes, a metric called Bayesian surprise is incorporated to increase the creativity of the recommended candidates. The new algorithm, called Deep Canonical PARAFAC Factorization (DCPF), is evaluated on both synthetic and large-scale real-world problems. An empirical study for video recommendation demonstrates the superiority of the proposed model and indicates that it can better capture the latent patterns of interactions and generates interesting recommendations based on creative tag combinations. Wei Lu 0006, Korris Fu-Lai Chung, Martin Ester, Wei Liu 0005 |
ACM Trans. Inf. Syst. | 2 |
| 2018 | Deep Domain Adaptation Based on Multi-layer Joint Kernelized DistanceabstractDomain adaptation refers to the learning scenario where a model learned from the source data is applied on the target data which have the same categories but different distributions. In information retrieval, there exist application scenarios like cross domain recommendation characterized similarly. In this paper, by utilizing deep features extracted from the deep networks, we proposed to compute the multi-layer joint kernelized mean distance between the k th target data predicted as the i th category and all the source data of the j th category $d_ij ^k$. Then, target data $T_m$ that are most likely to belong to the i th category can be found by calculating the relative distance $d_ii ^k/\sum_j d_ij ^k$. By iteratively adding $T_m$ to the training data, the finetuned deep model can adapt on the target data progressively. Our results demonstrate that the proposed method can achieve a better performance compared to a number of state-of-the-art methods. Sitong Mao, Xiao Shen 0001, Korris Fu-Lai Chung |
SIGIR | 3 |
| 2017 | Deep Network Embedding with Aggregated Proximity PreservingabstractNetwork embedding is an effective method to learn a low-dimensional feature vector representation for each node of a given network. In this paper, we propose a deep network embedding model with aggregated proximity preserving (DNE-APP). Firstly, an overall network proximity matrix is generated to capture both local and global network structural information, by aggregating different k-th order network proximities between different nodes. Then, a semi-supervised stacked auto-encoder is employed to learn the hidden representations which can best preserve the aggregated proximity in the original network, and also map the node pairs with higher proximity closer to each other in the embedding space. With the hidden representations learned by DNE-APP, we apply vector-based machine learning techniques to conduct node classification and link label prediction tasks on the real-world datasets. Experimental results demonstrate the superiority of our proposed DNE-APP model over the state-of-the-art network embedding algorithms. Xiao Shen 0001, Korris Fu-Lai Chung |
ASONAM | 2 |
| 2017 | Leveraging Cross-Network Information for Graph Sparsification in Influence MaximizationabstractWhen tackling large-scale influence maximization (IM) problem, one effective strategy is to employ graph sparsification as a pre-processing step, by removing a fraction of edges to make original networks become more concise and tractable for the task. In this work, a Cross-Network Graph Sparsification (CNGS) model is proposed to leverage the influence backbone knowledge pre-detected in a source network to predict and remove the edges least likely to contribute to the influence propagation in the target networks. Experimental results demonstrate that conducting graph sparsification by the proposed CNGS model can obtain a good trade-off between efficiency and effectiveness of IM, i.e., existing IM greedy algorithms can run more efficiently, while the loss of influence spread can be made as small as possible in the sparse target networks. Xiao Shen 0001, Korris Fu-Lai Chung, Sitong Mao |
SIGIR | 2 |
| 2016 | Scarce Feature Topic Mining for Video RecommendationabstractRecommendation for user generated content sites has gained significant attention nowadays. To satisfy the niche tastes of users, product recommendation poses more challenges due to the data sparsity issue. This work is motivated by a real world online video recommendation problem, where the click records database suffers from sparseness of video inventory and video tags. Targeting the long tail phenomena of user behavior and sparsity of item features, we propose a personalized compound recommendation framework for online video recommendation called Dirichlet mixture probit model for information scarcity (DPIS). Assuming that each record is generated from a representation of user preferences, DPIS is a probit classifier utilizing record topical clustering on the user part for recommendation. As demonstrated by the real-world application, the proposed DPIS achieves better performance than traditional methods. Wei Lu 0006, Korris Fu-Lai Chung, Kunfeng Lai |
CIKM | 2 |
| 2016 | Computational Creativity Based Video RecommendationabstractComputational creativity, as an emerging domain of application, emphasizes the use of big data to automatically design new knowledge. Based on the availability of complex multi-relational data, one aspect of computational creativity is to infer unexplored regions of feature space and novel learning paradigm, which is particularly useful for online recommendation. Tensor models offer effective approaches for complex multi-relational data learning and missing element completion. Targeting at constructing a recommender system that can compromise between accuracy and creativity for users, a deep Bayesian probabilistic tensor framework for tag and item recommending is adopted. Empirical results demonstrate the superiority of the proposed method and indicate that it can better capture latent patterns of interaction relationships and generate interesting recommendations based on creative tag combinations. Wei Lu 0006, Korris Fu-Lai Chung |
SIGIR | 2 |
| 2016 | Semi-supervised classification method through oversampling and common hidden space
Aimei Dong, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2016 | Transfer affinity propagation-based clustering
Wenlong Hang, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2016 | A novel multi-task TSK fuzzy classifier and its enhanced version for labeling-risk-aware multi-task classification
Yizhang Jiang, Zhaohong Deng, Kup-Sze Choi, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 4 |
| 2015 | Multi-task TSK fuzzy system modeling using inter-task correlation information
Yizhang Jiang, Zhaohong Deng, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 3 |
| 2015 | Support vector machine with manifold regularization and partially labeling privacy protection
Tongguang Ni, Korris Fu-Lai Chung, Shitong Wang 0001 |
Inf. Sci. | 2 |
| 2014 | Persuasion driven influence propagation in social networksabstractWith the explosive growth of online social network services such as Facebook and Twitter, people now are providing ample opportunities to share information, ideas, and innovations among others. Studying social influence and information diffusion in online social networks can be extremely useful for various real-life applications, notably influencer marketing and viral marketing. Influence maximization problem, motivated by viral marketing applications, has received extensive attentions from social network analysis community in recent years. The goal of the problem is to find a small set of k seed nodes in a social network that maximize the influence spread under certain influence propagation models. Based on the widely adopted Independent Cascade model, a Persuasiveness Aware Cascade (PAC) model which considers social persuasion in influence propagation is proposed. In this model, the user-to-user influence probability is estimated by three types of social persuasion, namely, tie strength, peer conformity, and authority. Experiments conducted over real-world social networks suggest that the proposed model with the new social persuasion measures is more effective in describing real-world influence propagation than the well-studied propagation models for influence maximization. Terrence Leung, Korris Fu-Lai Chung |
ASONAM | 2 |
| 2014 | Scaling Up Synchronization-Inspired Partitioning ClusteringabstractBased on the extensive Kuramoto model, synchronization-inspired partitioning clustering algorithm was recently proposed and is attracting more and more attentions, due to the fact that it simulates the synchronization phenomena in clustering where each data object is regarded as a phase oscillator and the dynamic behavior of the objects is simulated over time. In order to circumvent the serious difficulty that its existing version can only be effectively carried out on considerably small/medium datasets, a novel scalable synchronization-inspired partitioning clustering algorithm termed LSSPC, based on the center-constrained minimal enclosing ball and the reduced set density estimator, is proposed for large dataset applications. LSSPC first condenses a large scale dataset into its reduced dataset by using a fast minimal-enclosing-ball based approximation for the reduced set density estimator, thus achieving an asymptotic time complexity that is linear in the size of dataset and a space complexity that is independent of this size. Then it carries out clustering adaptively on the obtained reduced dataset by using Sync with the Davies-Bouldin clustering criterion and a new order parameter which can help us observe the degree of local synchronization. Finally, it finishes clustering by using the proposed algorithm CRD on the remaining objects in the large dataset, which can capture the outliers and isolated clusters effectively. The effectiveness of the proposed clustering algorithm LSSPC for large datasets is theoretically analyzed and experimentally verified by running on artificial and real datasets. Wenhao Ying, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Fuzzy partition based soft subspace clustering and its applications in high dimensional data
Jun Wang 0024, Shitong Wang 0001, Korris Fu-Lai Chung, Zhaohong Deng |
Inf. Sci. | 3 |
| 2012 | Semiconducting bilinear deep learning for incomplete image recognitionabstractImage recognition with incomplete data is a well-known hard problem in multimedia content analysis. This paper proposes a novel deep learning technique called semiconducting bilinear deep belief networks (SBDBN) by referencing human's visual cortex and intelligent perception. Inheriting from deep models, SBDBN simulates the laminar structure of human's cerebral cortex and the neural loop in human's visual areas. To address the special difficulties of image recognition with incomplete data, we design a novel second-order deep architecture with semiconducting restricted boltzmann machines. Moreover, two peaks activation of human's perception is implemented by three learning stages of semiconducting bilinear discriminant initialization, greedy layer-wise reconstruction, and global fine-tuning. Owing to exploiting the embedding information according to the reliable features rather than any completion of missing features, the proposed SBDBN has demonstrated outstanding recognition ability on two standard datasets and one constructed dataset, comparing with both incomplete image recognition techniques and existing deep learning models. Shenghua Zhong, Yan Liu 0004, Korris Fu-Lai Chung, Gangshan Wu |
ICMR | 3 |
| 2012 | Transfer Spectral Clustering
Korris Fu-Lai Chung |
ECML/PKDD (2) | 2 |
| 2011 | Water reflection recognition via minimizing reflection cost based on motion blur invariant momentsabstractWater reflection, a kind of typical imperfect reflection symmetry problem, plays an important role in image content analysis. However, existing techniques of symmetry recognition cannot recognize water reflection images correctly because of the complex and various distortions caused by water wave. To address this difficulty, we construct a novel feature space which is composed of motion blur invariant moments. Moreover, we propose an efficient detection algorithm to determine the reflection axis in images with water reflection. By experimenting on real image dataset with different tasks, the proposed techniques demonstrate impressive results in the water reflection image classification, the reflection axis detection, and the retrieval of the images with water reflection. Shenghua Zhong, Yan Liu 0004, Ling Shao 0001, Korris Fu-Lai Chung |
ICMR | 4 |
| 2010 | A new context-dependent term weight computed by boost and discount using relevance informationabstractAbstract We studied the effectiveness of a new class of context‐dependent term weights for information retrieval. Unlike the traditional term frequency–inverse document frequency (TF–IDF), the new weighting of a term t in a document d depends not only on the occurrence statistics of t alone but also on the terms found within a text window (or “document‐context”) centered on t. We introduce a Boost and Discount (B&D) procedure which utilizes partial relevance information to compute the context‐dependent term weights of query terms according to a logistic regression model. We investigate the effectiveness of the new term weights compared with the context‐independent BM25 weights in the setting of relevance feedback. We performed experiments with title queries of the TREC‐6, ‐7, ‐8, and 2005 collections, comparing the residual Mean Average Precision (MAP) measures obtained using B&D term weights and those obtained by a baseline using BM25 weights. Given either 10 or 20 relevance judgments of the top retrieved documents, using the new term weights yields improvement over the baseline for all collections tested. The MAP obtained with the new weights has relative improvement over the baseline by 3.3 to 15.2%, with statistical significance at the 95% confidence level across all four collections. Edward K. F. Dang, Robert Wing Pong Luk, James Allan 0001, Edward Kei Shiu Ho, Stephen Chi-fai Chan, Korris Fu-Lai Chung, Dik Lun Lee |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2008 | Discovering the Correlation between Stock Time Series and Financial NewsabstractIt is always expected that a correlation exists between the movement of stock prices (technical analysis) and news sentiment (fundamental analysis). If we can determine such a correlation, further interesting research directions will certainly be generated. In this paper, a system prototype is proposed for investigating the correlation between stock prices and news sentiment. Our primary target market is Hong Kong and the system is customized for Chinese language. Different methods and the impacts of various design parameters are tested in the experiments. Tak-Chung Fu, Ka-ki Lee, Donahue C. M. Sze, Korris Fu-Lai Chung, Chak-man Ng |
Web Intelligence | 4 |
| 2007 | Improving weak ad-hoc queries using wikipedia as external corpusabstractIn an ad-hoc retrieval task, the query is usually short and the user expects to find the relevant documents in the first several result pages. We explored the possibilities of using Wikipedia's articles as an external corpus to expand ad-hoc queries. Results show promising improvements over measures that emphasize on weak queries. Robert Wing Pong Luk, Edward Kei Shiu Ho, Korris Fu-Lai Chung |
SIGIR | 4 |
| 2006 | A collaborative filtering framework based on fuzzy association rules and multiple-level similarity
Cane Wing-ki Leung, Stephen Chi-fai Chan, Korris Fu-Lai Chung |
Knowl. Inf. Syst. | 3 |
| 2005 | Incremental stock time series data delivery and visualizationabstractSB-Tree is a binary tree data structure proposed to represent time series according to the importance of data points. Its use in stock data management is distinguished by preserving the critical data points' attribute values, retrieving time series data according to the importance of data points and facilitating multi-resolution time series retrieval. As new stock data are available continuously, an effective updating mechanism for SB-Tree is needed. In this paper, a study of different updating approaches is reported. Three families of updating methods are proposed. They are periodic rebuild, batch update and point-by-point update. Their efficiency, effectiveness and characteristics are compared and reported. Tak-Chung Fu, Korris Fu-Lai Chung, Pui-ying Tang, Robert Wing Pong Luk, Chak-man Ng |
CIKM | 2 |
| 2005 | Text Categorization Based on Subtopic Clusters
Francis C. Y. Chik, Robert Wing Pong Luk, Korris Fu-Lai Chung |
NLDB | 3 |
| 2005 | On the use of hierarchical information in sequential mining-based XML document similarity computation
Ho-pong Leung, Korris Fu-Lai Chung, Stephen Chi-fai Chan |
Knowl. Inf. Syst. | 2 |
| 2003 | A New Sequential Mining Approach to XML Document Similarity Computation
Ho-pong Leung, Korris Fu-Lai Chung, Stephen Chi-fai Chan |
PAKDD | 2 |
| 2002 | Evolutionary Time Series Segmentation for Stock Data MiningabstractStock data in the form of multiple time series are difficult to process, analyze and mine. However, when they can be transformed into meaningful symbols like technical patterns, it becomes easier. Most recent work on time series queries concentrates only on how to identify a given pattern from a time series. Researchers do not consider the problem of identifying a suitable set of time points for segmenting the time series in accordance with a given set of pattern templates (e.g., a set of technical patterns for stock analysis). On the other hand, using fixed length segmentation is a primitive approach to this problem; hence, a dynamic approach (with high controllability) is preferred so that the time series can be segmented flexibly and effectively according to the needs of users and applications. In view of the fact that such a segmentation problem is an optimization problem and evolutionary computation is an appropriate tool to solve it, we propose an evolutionary time series segmentation algorithm. This approach allows a sizeable set of stock patterns to be generated for mining or query. In addition, defining the similarity between time series (or time series segments) is of fundamental importance in fitness computation. By identifying perceptually important points directly from the time domain, time series segments and templates of different lengths can be compared and intuitive pattern matching can be carried out in an effective and efficient manner. Encouraging experimental results are reported from tests that segment the time series of selected Hong Kong stocks. Korris Fu-Lai Chung, Tak-Chung Fu, Robert Wing Pong Luk, Vincent T. Y. Ng |
ICDM | 1 |
| 2000 | Discovery of Generalized Association Rules with Multiple Minimum Supports
Chung-Leung Lui, Korris Fu-Lai Chung |
PKDD | 2 |
| 1997 | Offline Handwritten Chinese Character Recognition viaRadical Extraction and RecognitionabstractDespite the fact that Chinese characters are composed of radicals and that Chinese people usually formulate their knowledge of Chinese characters as a combination of radicals, very few studies have focused on a character decomposition approach to recognition, i.e., recognizing a character by first extracting and recognizing its radicals. Such an approach is adopted and the problem of how to extract radical sub-images from character images is particularly addressed. A radical extraction algorithm based on deformable templates (DTs) has been developed. The advantage of the character decomposition approach is demonstrated by feeding the extracted radical images to an adopted structural based Chinese character recognizer whose outputs are then combined to produce the class label of the input character. Simulation results show that the performance of the adopted Chinese character recognition system can be improved significantly when the character decomposition approach is used. Wilson W. S. Ip, Korris Fu-Lai Chung, Daniel S. Yeung |
ICDAR | 2 |