Vincent T. Y. Ng

dblp:n/VTYNg · also To-Yee Ng, Vincent Ng 0002, Vincent To-Yee Ng · DBLP profile ↗
← Back
54ranked-venue papers
10as first author
0since 2021 · last 2017
0000-0002-5114-3883ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 22 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 19 · 5 first-authorArtificial intelligence and machine learning · 9 · 2 first-authorDatabases, data management, data science and information retrieval · 9 · 1 first-authorComputer networks · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorSecurity and privacy · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data mining · 100%
Human-computer interaction and pervasive computing
1 paper
Learning and educational technologies · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
pattern mining
0.132002
Evolutionary Time Series Segmentation for Stock Data Mining · ICDM 2002
Efficient Mining of Association Rules in Distributed Databases · IEEE Trans. Knowl. Data Eng. 1996
Maintenance of Discovered Knowledge: A Case in Multi-Level Association Rules · KDD 1996
Data mining › pattern mining
association rule mining
0.031996
Efficient Mining of Association Rules in Distributed Databases · IEEE Trans. Knowl. Data Eng. 1996
Maintenance of Discovered Knowledge: A Case in Multi-Level Association Rules · KDD 1996
Maintenance of Discovered Association Rules in Large Databases: An Incremental Updating Technique · ICDE 1996
Data mining › temporal data mining
time series mining
0.012002
Evolutionary Time Series Segmentation for Stock Data Mining · ICDM 2002
Data mining › time series analysis
time series segmentation
0.012002
Evolutionary Time Series Segmentation for Stock Data Mining · ICDM 2002
Data mining › pattern mining › association rule mining
distributed association rule mining
0.011996
Efficient Mining of Association Rules in Distributed Databases · IEEE Trans. Knowl. Data Eng. 1996
Data mining › incremental mining
incremental updating
0.011996
Maintenance of Discovered Association Rules in Large Databases: An Incremental Updating Technique · ICDE 1996
Distributed systems › distributed data processing
distributed data mining
0.011996
Efficient Mining of Association Rules in Distributed Databases · IEEE Trans. Knowl. Data Eng. 1996
Mathematical optimization
evolutionary computation
0.012002
Evolutionary Time Series Segmentation for Stock Data Mining · ICDM 2002

Methods — techniques the papers use, named apart from their topics

perceptually important points · 0.1evolutionary computation · 0.1candidate set pruning · 0.0incremental updating technique · 0.0association rule mining · 0.0
YearPublicationVenuePosition
2017 Predicting new and unusual mobility patterns
abstract
Traditional location-based service profiles user's traits by looking for patterns in historical mobility behaviors. Yet, from time to time, people are adventurous and would often like to go to unvisited places, or follow new transition paths. At that time, their next movements will be inconsistent with any previous patterns, making location-based recommendations inaccurate and irrelevant to user's real need. Under such circumstance, an alternative strategy is to figure out user's destination and intention before recommendation, where the ability to predict new and unusual mobility patterns plays a critical role. In this paper, we define the next location that breaks the earliest on-going patterns as a Point of Change (POC). To predict POCs, we introduce a mobility model, called ST-Pattern Network, to learn the occurrences of POCs under the regularity of spatial-temporal patterns. By computing the similarities of matched patterns, our model can online predict future POCs as well as fit recent trajectory via a pattern network. Experiments show 10% accuracy improvement on POC prediction can be achieved over traditional Markov models. Furthermore, we are able to categorize the POC into more refined scenarios so that different recommendations can be suggested under different circumstances.
Victor C. Liang, Vincent T. Y. Ng
CSCWD2
2017 Anomaly Detection with Attribute Conflict Identification in Bank Customer Data
abstract
In commercial banks, data centers often integrates different data sources, which represent complex and independent business systems. Due to the inherent data variability and measurement or execution errors, there may exist some abnormal customer records (data). Existing automatic abnormal customer detection methods are outlier detection which focuses on the differences between customers, and it ignores the other possible abnormal customers caused by the inner features confliction of each customer. In this paper, we designed a method to identify abnormal customer information whose inner attributes are conflicting (confliction detection). We integrate the outlier detection and the confliction identification techniques together, as the final abnormality detection. This can provide a complete and accurate support of customer data for commercial bank's decision making. Finally, we have performed experiments on a dataset from a Chinese commercial bank to demonstrate the effectiveness of our method.
Vincent T. Y. Ng
SMARTCOMP2
2016 Syllable based DNN-HMM Cantonese Speech to Text System
Timothy Wong, Claire Li, Sam Lam, Billy Chiu, Qin Lu 0001, Minglei Li 0001, Dan Xiong, Roy Shing Yu, Vincent T. Y. Ng
LREC9
2016 Perception-oriented video saliency detection via spatio-temporal attention analysis
Shenghua Zhong, Yan Liu 0004, Vincent T. Y. Ng, Yang Liu 0007
Neurocomputing3
2016 A Hierarchical Ensemble of ECOC for cancer classification based on multi-class microarray data
Kunhong Liu 0001, Vincent T. Y. Ng
Inf. Sci.3
2015 Associating sentimental orientation of Chinese neologism in social media data
abstract
Sentiment analysis has always found its practical use in collecting people's preferences towards any subject in the context of social media. Unlike normal words available in dictionaries, neologisms are not easy to be labeled with a sentimental orientation while they have been widely used in conveying people's feelings and opinions. In order to conduct a reliable sentiment analysis for neologisms, a neologism discovery method is first required. Next, a sentimental analysis based on the discovery results can be performed. This paper proposes a 2-step novel solution by having a Chinese neologism discovery method and then a sentimental orientation determination algorithm based on varied TF-IDF. For neologism discovery, statistical data include frequency, duration of appearance and the number of users using a neologism. For sentimental orientation determination, we consider keyword term frequency, and document frequency together and use a varied TF-IDF algorithm. The preliminary experimental results show good precision rate and recall rates for a collection of social media data in both neologism discovery and sentimental analysis.
Lifeng Huang, Vincent T. Y. Ng
CSCWD3
2014 Analyzing sentimental influence of posts on social networks
abstract
Lots of effort has been conducted to analyze information of social networks, such as sentiment trend analysis of social network users. Our aim is to analyze the sentimental influence of posts and compare the result on various topics and different social media platforms. Large amounts of posts are generated on social networks every day. People are curious in finding the influence among them. Most researchers measured the influence of a post through the number of replies it received. However, we are not sure if the influence is made positively or negatively on other posts if their sentimental information is not considered. In this paper, three research questions are raised and methodologies are proposed for the measure of sentimental influence of posts. Finally, a preliminary experiment is designed and carried out with some interesting results found.
Beiming Sun, Vincent T. Y. Ng
CSCWD2
2014 Collaborative discovery of Chinese neologisms in social media
abstract
The emergence of neologism in social media has bought the researchers' attention. Traditional ways of text mining are not sufficient to handle the unique properties of messages in the new media. New methods have been developed to extract neologisms in order to help researchers to understand about community behavior in different media. In this paper, we propose a collaborative framework to detect neologisms from various social media. There are 4 different types of agents working collaboratively. Among them, the summarizing agent is using the life span parameter to confirm if an unknown character pattern is a neologism. Preliminary experiments have been performed to investigate the possible popularity patterns of some known neologisms.
Shek Lung Lai, Vincent T. Y. Ng
SMC2
2013 Predicting short interval tracking polls with online social media
abstract
The of behavioral patterns in online social media are often reflecting the happenings in our society. These patterns, which can be considered as opinions, are often correlated with public opinion polling. However, many correlation analyses done previously were for subsequent discoveries and not being able to handle short interval polling opinions. For opinions obtained from tracking polling with short opinion collection interval, like rolling polling, it cannot perform well in tracing the latest trends. This paper describes an extended correlation model for such kind of polling in examining the correlation between opinion in online social media and the public opinion from tracking poll. It has been tested with a recent rolling polling and it outperformed the previous correlation models.
Li Ho Leung, Vincent T. Y. Ng, Simon C. K. Shiu
CSCWD2
2013 Collaborative Discovering Influences of News in Social Network Sites
abstract
In many cases, news articles are shared in Online Social Network (OSN) sites by their members. Disparate kinds of information are often propagated from the news media to the online social networks. Little efforts were devoted to measure or estimate the social influences of news media on the online social networks when compared to the identifications of their connectivity. In this paper, we took references on measurements of social influences among individuals in online social network and derived absolute, relative and combined measurements in order to deduce the social influences of single news article. Those measurements can be extended to a set of news articles and a social medium hence they can be used to evaluate and compare the social influences of the numerous news media. Experiments to compare the performances of the proposed measurements will also be demonstrated and discussed.
Li Ho Leung, Vincent T. Y. Ng, Simon C. K. Shiu
SMC2
2013 Designing i*CATch: A multipurpose, education-friendly construction kit for physical and wearable computing
abstract
This article presents the design and development of i*CATch, a construction kit for physical and wearable computing that was designed to be scalable, plug-and-play, and to provide support for iterative and exploratory learning. It consists of a standardized construction interface that can be adapted for a wide range of soft textiles or electronic boards, a set of functional components, and an easy-to-use hybrid text-graphical integrated development environment. The objective was to design an easily usable, manufacturable and extensible construction kit that can be used in a wide range of teaching tasks for a wide variety of student demographic profiles. We present detailed specifications of our construction kit and explain some of the major design decisions. Experiences in using the kit in multiple teaching environments, ranging from elementary school to postgraduate, demonstrate that the design objectives have been achieved.
Grace Ngai, Stephen Chi-fai Chan, Hong Va Leong, Vincent T. Y. Ng
ACM Trans. Comput. Educ.4
2013 Discovering associations between news and contents in social network sites with the D-Miner service framework
Li Ho Leung, Vincent T. Y. Ng
J. Netw. Comput. Appl.2
2012 Analyzing social networks with D-miner Cloud
abstract
The Online Social Network (OSN) sites have been getting more and more popular in recent years and there are interests of having a tool to automatically retrieve and analyze the information in order to understand their related social behavior. This paper presents a framework of D-miner Cloud (DMC) based on a cloud architecture which can provide collection of the information from heterogeneous OSN sites, managing and analyzing the collected information afterwards. It has 5 components consisting of frontend devices, service gateway, service unit, central repository and OSN sites. We present the approach to implement the framework, integrate DMC and the Application Program Interfaces (API) of OSN sites, organize the communications in DMC framework and implement the mobile frontend design.
Li Ho Leung, Vincent T. Y. Ng
CSCWD2
2012 Improving database performance with a mixed fragmentation design
Narasimhaiah Gorla, Vincent T. Y. Ng, Dik Man Law
J. Intell. Inf. Syst.2
2011 Multi-agent system for shipper's truck freight collaboration
abstract
This paper discusses how to reduce the cost of transportation of shipper and increases the utilization of the truck loading of the truck company by using the collaborative multi-agent system. We have designed an Intelligent Agent Cargo Booking Framework (IACBF) to facilitate the shipper grouping, partnership recognition and cost allocation with a multi-agent system solution. A prototype has been developed to illustrate the operations of our framework.
Pearl C. C. Shum, Vincent T. Y. Ng
CSCWD2
2011 Lifespan and popularity measurement of online content on social networks
abstract
With rapid development and increased popularity of social networks, more interests have been made in obtaining information from such social networking websites. Analysis on the popularity of online contents is one of the hottest interests which has triggered intensive research. Our research focuses on measuring the popularity of online posts within specified topics, such as “drug abuse”, as it can be used to detect the crime and discover potential drug abusers. We measure the lifespan and the popularity of drug related posts in order to know the level of influence they made. Identifying popular posts online can help us reveal the trend, and also detect the latent danger and may prevent future crime. In this paper, the Comment Arrival Model is proposed to identify the lifespan and the comment frequency pattern of posts, which are considered the main factors to define post popularity. And we present 4 general models to measure the popularity of posts which can be applied in different social network platforms. We also did experiments and evaluated the performance of models in two popular social networks in Hong Kong, HK Discussion and Twitter.
Beiming Sun, Vincent T. Y. Ng
ISI2
2011 Metasample-Based Sparse Representation for Tumor Classification
abstract
A reliable and accurate identification of the type of tumors is crucial to the proper treatment of cancers. In recent years, it has been shown that sparse representation (SR) by l1-norm minimization is robust to noise, outliers and even incomplete measurements, and SR has been successfully used for classification. This paper presents a new SR-based method for tumor classification using gene expression data. A set of metasamples are extracted from the training samples, and then an input testing sample is represented as the linear combination of these metasamples by l1-regularized least square method. Classification is achieved by using a discriminating function defined on the representation coefficients. Since l1-norm minimization leads to a sparse solution, the proposed method is called metasample-based SR classification (MSRC). Extensive experiments on publicly available gene expression data sets show that MSRC is efficient for tumor classification, achieving higher accuracy than many existing representative schemes.
Chun-Hou Zheng 0001, Lei Zhang 0006, Vincent T. Y. Ng, Simon C. K. Shiu, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2011 Molecular Pattern Discovery Based on Penalized Matrix Decomposition
abstract
A reliable and precise identification of the type of tumors is crucial to the effective treatment of cancer. With the rapid development of microarray technologies, tumor clustering based on gene expression data is becoming a powerful approach to cancer class discovery. In this paper, we apply the penalized matrix decomposition (PMD) to gene expression data to extract metasamples for clustering. The extracted metasamples capture the inherent structures of samples belong to the same class. At the same time, the PMD factors of a sample over the metasamples can be used as its class indicator in return. Compared with the conventional methods such as hierarchical clustering (HC), self-organizing maps (SOM), affinity propagation (AP) and nonnegative matrix factorization (NMF), the proposed method can identify the samples with complex classes. Moreover, the factor of PMD can be used as an index to determine the cluster number. The proposed method provides a reasonable explanation of the inconsistent classifications made by the conventional methods. In addition, it is able to discover the modules in gene expression data of conterminous developmental stages. Experiments on two representative problems show that the proposed PMD-based method is very promising to discover biological phenotypes.
Chun-Hou Zheng 0001, Lei Zhang 0006, Vincent T. Y. Ng, Simon C. K. Shiu, De-Shuang Huang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2010 i*CATch: a scalable plug-n-play wearable computing framework for novices and children
abstract
There has been much recent work in wearable computing that is directed at democratization of the field, to make it more accessible to the general public and more easily used by the hobbyist user. As the field becomes more diversified, there has also been a shift away from the highly specialized functionality of earlier applications towards aesthetics, creativity, design and self-expression, as well as a push towards using wearable computing as an outreach tool to broaden interest and exposure in engineering and computing.
Grace Ngai, Stephen Chi-fai Chan, Vincent T. Y. Ng, Joey C. Y. Cheung, Sam S. S. Choy, Winnie W. Y. Lau, Jason T. P. Tse
CHI3
2010 Multi-site collaboration design for air cargo planning
abstract
The paper is about the design of an agent-based framework of multi-site cargo planning system for the air freight forwarding industry. Currently, most of the air freight forwarders still generate loading plans individually. We propose to streamline effectiveness of cargo handling with a multi-site collaborative solution. Our approach is based on the intelligent agent technology, bin packing heuristic and the inter-site workflow protocol. The major problems being tackled are (1) how to make the best estimation of the equipments required; (2) how to load the cargoes under the minimum cost in the standalone site case and (3) optimization of space and cargo loading cost under the multiple sites case. A prototype has been developed to illustrate the operation s of our framework.
Pearl C. C. Shum, Vincent T. Y. Ng
CSCWD2
2010 Inferring the Transcriptional Modules Using Penalized Matrix Decomposition
Chun-Hou Zheng 0001, Lei Zhang 0006, Vincent T. Y. Ng, Simon C. K. Shiu, Shu-Lin Wang
ICIC (2)3
2010 HPC - privacy model for collaborative skyline processing
abstract
In general, skyline query is defined as finding a set of interesting database objects, which are not dominated to one another objects. A typical example is to find the hotel that is cheap and close to the beach. Since the introduction of skyline operator by Borzsonyi et al into database community, there has been a number of research works evolving and related publications related in last decade. However, there is only a few of them working on distributed skyline processing in collaborative computing environments. None of them considered the issue of privacy enforcement. The problem is that server has to disclose the sub-skylines (the actual skyline points) without privacy protection. In this paper, we propose the Hierarchical Piecewise Curve (HPC) model to enforce privacy during collaborative skyline processing and the private information can be released in a hierarchically controllable manner.
Boris Y. L. Chan, Vincent T. Y. Ng
ISI2
2010 Discovering potential drug abuse with fuzzy sets
abstract
According to the recent CRDA and drug statistics in Hong Kong, the age range for students abusing psychotropic substance is getting younger emphatically. In 2009, we have collaborated with the RPCO KE of the Hong Kong Police for the development of a web miner to detect information related to potential drug abusers on the web. This paper describes the D-MinerB algorithm for detecting possible patterns posted on blogs. Fuzzy set and some language parsing techniques have been adopted. For the preliminary test of the collected set, we are able to identify 21 blog posts from 472 blog posts while over 84% of the identified ones are relevant.
Li Ho Leung, Vincent T. Y. Ng
SMC2
2009 Alternative Feature Mapping for Heterogeneous Gene Data Classification
abstract
In order to overcome the limitation on small sizes of gene datasets, many meta-classification methods which ensemble classifiers with different datasets have been developed. However, due to discrepancies of the characteristics within heterogeneous or cross-platform datasets, the number of common and significant genes is usually small. Instead of matching common genes between heterogeneous datasets, we propose a novel solution, alternative feature mapping approach (AFM), to utilize related and discriminative gene expressions while not necessarily having exact matches. Genes in the training dataset are clustered and mapped to the test dataset as gene groups. Through analyzing the correlation within gene groups between training and test datasets, related significant genes can be applied for classification. We conducted experiments consisting of 8 heterogeneous datasets with different cancer types and platforms to test the effectiveness of AFM. Our experiments show that classification performance is greatly improved using suitable significant genes selected by AFM.
Victor C. Liang, Vincent T. Y. Ng
BIBE2
2009 Multi-site collaborative airfreight cargo packing
abstract
Airfreight forwarding is becoming more important with the increase of international trade. For an international company, sometimes, it would have several branches associating with different airports in the same region. When a branch is overloaded, nearby branches of different airports can offer help and share the workload. A framework of agents has been proposed to support collaborative cargo packing for multi-sites in the same region. Besides defining a set of XML message exchange standards in the framework, a core component is the packing module which is supported by RFID tags with information encoded. This paper discusses how to apply First Fit Decreasing and Best Fit Decreasing for packing airfreight cargoes and their preliminary performance.
Vincent T. Y. Ng, Suo Na, Stephen Chi-fai Chan
CSCWD1
2009 Dynamic collaborative robotic platform - A brief introduction
abstract
This paper presents a design for a platform of collaborative robots and electronic devices. The platform design is based on the similarity between the type, functionality and characteristics of the robot or device. We categorize the participating devices by functionality rather than by architecture, therefore making it easy to support new robots or devices with similar functions but different architectures. This approach also allows users to develop, implement and port applications quickly and easily. We demonstrate the efficacy and correctness of our platform through a variety of robotic applications ranging from research to teaching.
Jason T. P. Tse, Stephen Chi-fai Chan, Grace Ngai, Joey C. Y. Cheung, Vincent T. Y. Ng
CSCWD5
2009 Hierarchical Pareto Curve (HPC) Model for Privacy Skyline
abstract
Privacy is an essential issue in database publishing. Since the introduction of skyline operator in database community, there was a few researches working on the privacy skyline and related the privacy theory, framework and model in last few years. For those algorithms (e.g. Skyline Check and Privacy Diagnostics), centralized database is assumed and the consideration of concurrency and parallelism is in lack. In this paper, we propose the hierarchical Pareto curve (HPC) model for private skyline processing. In HPC, answers to the skyline query are interpolated by spline function and represented by a set of polynomial Pareto curves. Hence, skyline querying requests can be satisfied without disclosing the actual data points. Moreover, the accuracy of a skyline query can be controlled by setting the order of the polynomial expression and total number of Pareto curves. The HPC model can be extended for distributed and cooperative computing environments. With privacy embedded in piecewise Pareto curves and merging operators developed, distributed skyline processing becomes practical. From our preliminary experiments, the results show supportive indications towards the HPC model.
Boris Y. L. Chan, Vincent T. Y. Ng, Jacob Sun
SMC2
2009 Collaborative Group Assignment with Competency Trees
abstract
In collaborative commerce, task groups are often formed from experts of several companies to perform collaborative projects. The proper position assignments to different staff members from multi-organizations have become an important task for the success of projects. This paper proposes a collaborative framework with two components for companies to assign professionals to different roles in forming a group to complete a task. The first component is further divided into 2 parts. The first part uses a tree competency model to define competencies of staff (agents) as well as role requirements. The second part contains a tree similarity algorithm and bonus point algorithm to support the role assignment. The adoption of clustering techniques and negotiation protocol between team requester and multi-agent providers are the second component to support the group role assignments. Preliminary experiments have been performed with some interesting results.
Wan Hok Man, Vincent T. Y. Ng
SMC2
2008 Dynamic Pattern Analysis Framework for cooperative crime prevention
abstract
Spatial analysis plays a key role in crime prevention. Traditional approaches such as clustering can find static patterns but do not consider the change of spatial patterns over time. In this paper, we introduce a new analysis framework, Dynamic Pattern Analysis Framework (DPA Framework) focusing on two types of related dynamic patterns: the displacement or diffusion of spatial patterns over time and the similarity between spatial patterns of different periods. The new framework aims to support cooperative crime prevention in a district of Hong Kong.
Kelvin Leong, Junco Li, Stephen Chi-fai Chan, Vincent T. Y. Ng
CSCWD4
2008 Quality service assignments for role-based web services
abstract
Many collaborative systems involve business processes and workflows performed on the web. Users would need to issue web service requests to complete their business transactions. The requests are of different types and priorities according to their roles in their respective business processes. One immediate problem is to ensure the timely and cost effective completion of the requests. In this paper, we have proposed to utilize service brokers for adequately assigning online continuous incoming service requests so as to achieve some pre-defined optimization goals. Here, the aim is to minimize the number of service providers required to fulfill the requests from the users. With this goal, the service broker can reduce the costs of service providers with the Semi-online Pseudo-open Multi-dimensional Vector-bin-packing (SoPomVa) algorithm. Preliminary experiments have been conducted to investigate the effectiveness of the SoPomVa proposed in comparison with some traditional bin packing algorithms and good performance results are obtained.
Vincent T. Y. Ng, Boris Y. L. Chan, Louis L. Y. Shun, Ringo Tsang
SMC1
2008 RRSi: indexing XML data for proximity twig queries
Patrick K. L. Ng, Vincent T. Y. Ng
Knowl. Inf. Syst.2
2007 A Novel Structural Similarity Measure on XML Data for Integrated Document Management
Patrick K. L. Ng, Vincent T. Y. Ng
J. Comput. Inf. Syst.2
2006 Automatic Template Detection for Structured Web Pages
abstract
Similar Web pages of Web sites on the World Wide Web are usually encoded from an underlying structured source, and generated dynamically from a pre-defined template, such as books' information pages in Amazon.com. By giving a set of Web pages from a common Website, it is possible to extract the template by analyzing common patterns between the Web pages. In our work, we developed the CF-EXALG (collaborative finer-EXALG), based on EXALG, to decompose Web pages and finding their common structures. In our system, templates that are used to generate Web pages can be discovered automatically and stored in XML format. Hence, data encoded in Web pages can be easily extracted and the template can be stored for future manipulation. In our preliminary experiments, CF-EXALG has shown to be more accurate and efficient when compared with other similar systems
Lawrence Lo, Vincent T. Y. Ng, Patrick Ng, Stephen Chi-fai Chan
CSCWD2
2006 Co-assembler - Supporting Collaborative Product Configuration for Customer-oriented e-Commerce
abstract
The ability to interactively and collaboratively configure and re-configure 3D products on the Internet is an important facilitator for e-commerce. We describe the design of such a system, Co-assembler, that enables casual customers to perform limited configurations of 3D products over the Internet. We discuss the technical challenges and outline our solution. We also describe the application of artificial intelligence to assist the customer in configuring the desired product, The system aims to capture customers' needs and convert them into technical specifications, helping to facilitate e-commerce for manufacturing enterprises.
Sophia M. K. Soo, Stephen Chi-fai Chan, Vincent T. Y. Ng
CSCWD3
2005 Quality guarantee for WSDL-based services
abstract
In recent years, Web services are getting popular. In order to have more reliable services on the Internet, QoS (quality of service) plays an important role. We extend the WSDL with an optional attribute in order to control the level of QoS. The extension enables Web services providers to support partial request satisfaction by a priority attribute controlling the priorities of WSDL requests and the corresponding sub-requests. We have performed a set of experiments to demonstrate the significance of the partial request satisfaction by with a range of values of the priority attribute. The results are interesting and demonstrated the usefulness of the new parameter.
Boris Y. L. Chan, Vincent T. Y. Ng, Stephen Chi-fai Chan
CSCWD (1)2
2005 Adaptive Learning in a Mobile Environment
Tat Wai Ho, Vincent T. Y. Ng, Stephen Chi-fai Chan, King Kang Tsoi
ICCE2
2005 Verification of Prerequisite Relationship among Learning Objects
King Kang Tsoi, Vincent T. Y. Ng
ICCE2
2004 Mapping XML Schema to Relations Using Genetic Algorithm
Vincent T. Y. Ng, Chi-Kong Chan, Stephen Chi-fai Chan
KES1
2004 Supporting metasearch with XSL
Robert Wing Pong Luk, Tharam S. Dillon, Vincent T. Y. Ng
J. Syst. Softw.3
2004 An evolutionary approach to pattern-based time series segmentation
abstract
Time series data, due to their numerical and continuous nature, are difficult to process, analyze, and mine. However, these tasks become easier when the data can be transformed into meaningful symbols. Most recent works on time series only address how to identify a given pattern from a time series and do not consider the problem of identifying a suitable set of time points for segmenting the time series in accordance with a given set of pattern templates (e.g., a set of technical patterns for stock analysis). However, the use of fixed-length segmentation is an oversimplified approach to this problem; hence, a dynamic approach (with high controllability) is preferable so that the time series can be segmented flexibly and effectively according to the needs of the users and the applications. In view of the fact that this segmentation problem is an optimization problem and evolutionary computation is an appropriate tool to solve it, we propose an evolutionary time series segmentation algorithm. This approach allows a sizeable set of pattern templates to be generated for mining or query. In addition, defining similarity between time series (or time series segments) is of fundamental importance in fitness computation. By identifying the perceptually important points directly from the time domain, time series segments and templates of different lengths can be compared and intuitive pattern matching can be carried out in an effective and efficient manner. Encouraging experimental results are reported from tests that segment both artificial time series generated from the combinations of pattern templates and the time series of selected Hong Kong stocks.
Korris Fu-Lai Chung, Tak-Chung Fu, Vincent T. Y. Ng, Robert Wing Pong Luk
IEEE Trans. Evol. Comput.3
2003 Determining the Asymmetries of Skin Lesions with Fuzzy Borders
abstract
Malignant melanoma is a popular cancer among youth; it is desirable to have a fast and convenience way to determine this disease in its early stage. One of the clinical features in diagnosis is related to the shape of lesions. In previous studies, circularity is commonly used as the asymmetric measurement of skin lesions. However, this measurement depends very much on the accuracy of the segmentation result. In this paper, we present an artificial neural network model to improve the measurements of the asymmetries of lesions that may have fuzzy borders. The main idea is enhancing the symmetric distant (eSD) with a number of variations. Results from experiments, which use the digitized images front the Lesion Clinic in Vancouver, Canada have shown the good discriminating power of the neural network model.
Vincent T. Y. Ng, Tim K. Lee, Benny Y. M. Fung
BIBE1
2003 Customer Loyalty on Recurring Loans
Vincent T. Y. Ng, Ida Ng
IDEAL1
2002 Evolutionary Time Series Segmentation for Stock Data Mining
abstract
Stock data in the form of multiple time series are difficult to process, analyze and mine. However, when they can be transformed into meaningful symbols like technical patterns, it becomes easier. Most recent work on time series queries concentrates only on how to identify a given pattern from a time series. Researchers do not consider the problem of identifying a suitable set of time points for segmenting the time series in accordance with a given set of pattern templates (e.g., a set of technical patterns for stock analysis). On the other hand, using fixed length segmentation is a primitive approach to this problem; hence, a dynamic approach (with high controllability) is preferred so that the time series can be segmented flexibly and effectively according to the needs of users and applications. In view of the fact that such a segmentation problem is an optimization problem and evolutionary computation is an appropriate tool to solve it, we propose an evolutionary time series segmentation algorithm. This approach allows a sizeable set of stock patterns to be generated for mining or query. In addition, defining the similarity between time series (or time series segments) is of fundamental importance in fitness computation. By identifying perceptually important points directly from the time domain, time series segments and templates of different lengths can be compared and intuitive pattern matching can be carried out in an effective and efficient manner. Encouraging experimental results are reported from tests that segment the time series of selected Hong Kong stocks.
Korris Fu-Lai Chung, Tak-Chung Fu, Robert Wing Pong Luk, Vincent T. Y. Ng
ICDM4
2001 Evolutionary segmentation of financial time series into subsequences
abstract
Time series data are difficult to manipulate. When they can be transformed into meaningful symbols, it becomes an easy task to query and understand them. While most recent works in time series query only concentrate on how to identify a given pattern from a time series, they do not consider the problem of identifying a suitable set of time points based upon which the time series can be segmented in accordance with a given set of pattern templates, e.g., a set of technical analysis patterns for stock analysis. On the other hand, using fixed length segmentation is only a primitive approach to such kind of problem and hence a dynamic approach is preferred so that the time series can be segmented flexibly and effectively. In view of the fact that such a segmentation problem is actually an optimization problem and evolutionary computation is an appropriate tool to solve it, we propose an evolutionary segmentation algorithm in this paper. Encouraging experimental results in segmenting the Hong Kong Hang Seng Index using 22 technical analysis patterns are reported.
Tak-Chung Fu, Korris Fu-Lai Chung, Vincent T. Y. Ng, Robert Wing Pong Luk
CEC3
2001 Web Agents for Spatial Mining on Air Pollution Meteorology
abstract
This paper describes an agent framework to support spatial data mining of air pollution data. A Web-based solution, APMF (Air Pollution Mining Framework), is developed to investigate the effect of meteorological and air pollutant elements on air pollution. It makes use of three types of agents: collect agents, coordinate agents and query/mining agents. Query/mining agents interact with the user, receive user queries and mining requests, and deliver the results. Coordinate agents help to prepare the data resources for the querying tasks. Collect agents provide access to a heterogeneous collection of Web sources. In order to standardize data exchanges and query/mining process descriptions, an XML-based language, DQML, is developed. A prototype of the APMF is implemented with a Java servlet and JavaLite.
Vincent T. Y. Ng, Stephen Chi-fai Chan, Sandra Au
CSCWD1
1998 Quantitative association rules over incomplete data
abstract
This paper explores the use of principle component analysis (PCA) to estimate missing values during the mining of quantitative association rules. An example of such association may be "15% of customers spend $100-$300 every month will have two cable outlets at home". In our algorithm, instead of imputing missing values before the mining process, we propose to integrate the imputation step within the process. The idea is to reduce the unnecessary imputation effort and to improve the overall performance. First, only attributes with enough support counts and with missing values are required to perform imputations. Thus, effort will not be wasted on unimportant attributes. Further, rather than estimating the actual value of a missing data, the possible range of the value is guessed. This will not affect the resultant quantitative association rules much but will cut down the guessing effort.
Vincent T. Y. Ng
SMC1
1998 A solid modeling library for the World Wide Web
Stephen Chi-fai Chan, Vincent T. Y. Ng, Albert S. F. Au
Comput. Networks2
1998 Touring Hong Kong via the WWW
Vincent T. Y. Ng, Stephen Chi-fai Chan, Chak Man Ng, Fergus Tang
Comput. Networks1
1997 A content-based search engine on medical images for telemedicine
abstract
Retrieving images by content and forming visual queries are important functionality of an image database system. Using textual descriptions to specify queries on image content is another important component of content-based search. The authors describe a medical image database system MIQS which supports visual queries such as query by example and query by sketch. In addition, it supports textual queries on spatial relationships between the objects of an image. MIQS is designed as a client-server application in which the client accesses the database and its images via the WWW.
David Wai-Lok Cheung, Chihung Lee, Vincent T. Y. Ng
COMPSAC3
1997 Concurrent access to point data
abstract
The B/sup +/-tree, R-tree and K-D-B tree have been proposed as the index structures for point data in a d-dimensional space. We apply the lock coupling technique to them to allow concurrent accesses. We discuss the two common user operations, search and insert, and present concurrency control algorithms for them. We have implemented these concurrency control algorithms in the SR distributed programming language. We report on our preliminary results comparing their relative performance.
Vincent T. Y. Ng, Tiko Kameda
COMPSAC1
1996 Maintenance of Discovered Association Rules in Large Databases: An Incremental Updating Technique
abstract
An incremental updating technique is developed for maintenance of the association rules discovered by database mining. There have been many studies on efficient discovery of association rules in large databases. However, it is nontrivial to maintain such discovered rules in large databases because a database may allow frequent or occasional updates and such updates may not only invalidate some existing strong association rules but also turn some weak rules into strong ones. An incremental updating technique is proposed for efficient maintenance of discovered association rules when new transaction data are added to a transaction database.
David Wai-Lok Cheung, Jiawei Han 0001, Vincent T. Y. Ng, C. Y. Wong
ICDE3
1996 Maintenance of Discovered Knowledge: A Case in Multi-Level Association Rules
David Wai-Lok Cheung, Vincent T. Y. Ng, Benjamin W. Tam
KDD2
1996 Efficient Mining of Association Rules in Distributed Databases
abstract
Many sequential algorithms have been proposed for the mining of association rules. However, very little work has been done in mining association rules in distributed databases. A direct application of sequential algorithms to distributed databases is not effective, because it requires a large amount of communication overhead. In this study, an efficient algorithm called DMA (Distributed Mining of Association rules), is proposed. It generates a small number of candidate sets and requires only O(n) messages for support-count exchange for each candidate set, where n is the number of sites in a distributed database. The algorithm has been implemented on an experimental testbed, and its performance is studied. The results show that DMA has superior performance, when compared with the direct application of a popular sequential algorithm, in distributed databases.
David Wai-Lok Cheung, Vincent T. Y. Ng, Ada Wai-Chee Fu, Yongjian Fu 0001
IEEE Trans. Knowl. Data Eng.2
1994 The R-Link Tree: A Recoverable Index Structure for Spatial Data
Vincent T. Y. Ng, Tiko Kameda
DEXA1