Yao-Chung Fan

dblp:85/6913 · DBLP profile ↗
← Back
33ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0002-6894-015XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Systems, architecture and hardware · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CGPT: Cluster-Guided Partial Tables with LLM-Generated Supervision for Table Retrieval
abstract
General-purpose embedding models have demonstrated strong performance in text retrieval but remain suboptimal for table retrieval, where highly structured content leads to semantic compression and query–table mismatch. Recent LLM-based retrieval augmentation methods mitigate this issue by generating synthetic queries, yet they often rely on heuristic partial-table selection and seldom leverage these synthetic queries as supervision to improve the embedding model. We introduce CGPT, a training framework that enhances table retrieval through LLM-generated supervision. CGPT constructs semantically diverse partial tables by clustering table instances using K-means and sampling across clusters to broaden semantic coverage. An LLM then generates synthetic queries for these partial tables, which are used in hard-negative contrastive fine-tuning to refine the embedding model. Experiments across four public benchmarks (MimoTable, OTTQA, FetaQA, and E2E-WTQ) show that CGPT consistently outperforms the existing baselines with an average R@1 improvement of 16.54%. Under cross-domain evaluation, CGPT further demonstrates strong cross-domain generalization and remains effective even when using smaller LLMs for synthetic query generation. These results indicate that semantically guided partial-table construction, combined with contrastive training from LLM-generated supervision, provides an effective and scalable paradigm for large-scale table retrieval. Our code is available at https://github.com/yumeow0122/CGPT.
Tsung-Hsiang Chou, Chen-Jui Yu, Shui-Hsiang Hsu, Yao-Chung Fan
WWW4
2026 STAR: Semantic Table Representation with Header-Aware Clustering and Adaptive Weighted Fusion
abstract
Table retrieval is the task of retrieving the most relevant tables from large-scale corpora given natural language queries. However, structural and semantic discrepancies between unstructured text and structured tables make embedding alignment particularly challenging. Recent methods such as QGpT attempt to enrich table semantics by generating synthetic queries, yet they still rely on coarse partial-table sampling and simple fusion strategies, which limit semantic diversity and hinder effective query–table alignment. We propose STAR (Semantic Table Representation), a lightweight framework that improves semantic table representation through semantic clustering and weighted fusion. STAR first applies header-aware K-means clustering to group semantically similar rows and selects representative centroid instances to construct a diverse partial table. It then generates cluster-specific synthetic queries to comprehensively cover the table's semantic space. Finally, STAR employs weighted fusion strategies to integrate table and query embeddings, enabling fine-grained semantic alignment. This design enables STAR to capture complementary information from structured and textual sources, improving the expressiveness of table representations. Experiments on five benchmarks show that STAR achieves consistently higher Recall than QGpT on all datasets, demonstrating the effectiveness of semantic clustering and weighted fusion for robust table representation. Our code is available at https://github.com/adsl135789/STAR.
Shui-Hsiang Hsu, Tsung-Hsiang Chou, Chen-Jui Yu, Yao-Chung Fan
WWW4
2024 JH-Ranker: Enhancing Keigo Recognition in Japanese Sentences through Multi-Task Learning
abstract
In the realm of Japanese language processing, the correct use and understanding of Keigo (honorific language) presents unique challenges due to its complex nature and context-dependent usage. This study introduces JH-Ranker, a novel multitask learning framework designed to enhance the recognition and understanding of Keigo in Japanese sentences. We are employing a multi-task learning approach by integrating Keigo Level, Contextual Field, and the Classification of Keigo Forms. This integration aims to foster a comprehensive understanding of Keigo usage from small data, leveraging the synergy between multiple tasks to achieve a deeper understanding with limited resources. Utilizing the BERT-base-japanese pre-trained model, JH-Ranker takes into account various Keigo Forms and contextual scenarios, going beyond mere syntactic analysis. Our methodology also emphasizes the importance of considering both the form and context of Keigo, addressing challenges such as label imbalance and overfitting in multi-label classification tasks. The experimental results demonstrate a significant improvement in Keigo recognition capabilities over existing baseline models, particularly in discerning the subtle nuances of Keigo Levels and Forms. This study contributes to the realm of natural language processing by presenting a refined and contextually sensitive method for Japanese Keigo. It demonstrates the language model’s adaptability and understanding in dealing with intricate honorific language scenarios.
Shih-Wei Guo, Ya-Chun Wen, Jui-Ling Yu, Yao-Chung Fan
IJCNN4
2024 Personalized Cloze Test Generation with Large Language Models: Streamlining MCQ Development and Enhancing Adaptive Learning
abstract
Cloze multiple-choice questions (MCQs) are essential for assessing comprehension in educational settings, but manually designing effective distractors is time-consuming.Addressing this, recent research has automated distractor generation, yet such methods often neglect to adjust the difficulty level to the learner's abilities, resulting in non-personalized assessments.This study introduces the Personalized Cloze Test Generation (PCGL) Framework, utilizing Large Language Models (LLMs) to generate cloze tests tailored to individual proficiency levels.Our PCGL Framework simplifies test creation by generating question stems and distractors from a single input word and adjusting the difficulty to match the learners proficiency.The framework significantly reduces the effort in creating tests and enhances personalized learning by dynamically adapting to the needs of each learner.
Chih-Hsuan Shen, Yi-Li Kuo, Yao-Chung Fan
INLG3
2024 Automating True-False Multiple-Choice Question Generation and Evaluation with Retrieval-based Accuracy Differential
abstract
Creating high-quality True-False (TF) multiple-choice questions (MCQs), with accurate distractors, is a challenging and time-consuming task in education.This paper introduces True-False Distractor Generation (TFDG), a pipeline that leverages pre-trained language models and sentence retrieval techniques to automate the generation of TF-type MCQ distractors.Furthermore, the evaluation of generated TF questions presents a challenge.Traditional metrics like BLEU and ROUGE are unsuitable for this task.To address this, we propose a new evaluation metric called Retrieval-based Accuracy Differential (RAD).RAD assesses the discriminative power of TF questions by comparing model accuracy with and without access to reference texts.It quantitatively evaluates how well questions differentiate between students with varying knowledge levels.This research benefits educators and assessment developers, facilitating the efficient automatic generation of high-quality TF-type MCQs and their reliable evaluation.
Chen-Jui Yu, Wen Hung Lee, Lin Tse Ke, Shih-Wei Guo, Yao-Chung Fan
INLG5
2024 X-Phishing-Writer: A Framework for Cross-lingual Phishing E-mail Generation
abstract
Cybercrime is projected to cause annual business losses of $10.5 trillion by 2025, a significant concern given that a majority of security breaches are due to human errors, especially through phishing attacks. The rapid increase in daily identified phishing sites over the past decade underscores the pressing need to enhance defenses against such attacks. Social Engineering Drills (SEDs) are essential in raising awareness about phishing yet face challenges in creating effective and diverse phishing e-mail content. These challenges are exacerbated by the limited availability of public datasets and concerns over using external language models like ChatGPT for phishing e-mail generation. To address these issues, this article introduces X-Phishing-Writer, a novel cross-lingual Few-shot phishing e-mail generation framework. X-Phishing-Writer allows for the generation of e-mails based on minimal user input, leverages single-language datasets for multilingual e-mail generation, and is designed for internal deployment using a lightweight, open-source language model. Incorporating Adapters into an Encoder–Decoder architecture, X-Phishing-Writer marks a significant advancement in the field, demonstrating superior performance in generating phishing e-mails across 25 languages when compared to baseline models. Experimental results and real-world drills involving 1,682 users showcase a 17.67% e-mail open rate and a 13.33% hyperlink click-through rate, affirming the framework’s effectiveness and practicality in enhancing phishing awareness and defense.
Shih-Wei Guo, Yao-Chung Fan
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2024 Handover QG: Question Generation by Decoder Fusion and Reinforcement Learning
abstract
In recent years, Question Generation (QG) has gained significant attention as a research topic, particularly in the context of its potential to support automatic reading comprehension assessment preparation. However, current QG models are mostly trained on factoid-type datasets, which tend to produce questions that are too simple for assessing advanced abilities. One promising alternative is to train QG models on exam-type datasets, which contain questions that require content reasoning. Unfortunately, there is a shortage of such training data compared to factoid-type questions. To address this issue and improve the quality of QG for generating advanced questions, we propose theHandover QGframework. This framework involves the joint training of exam-type QG and factoid-type QG, and controls the question generation process by interleavingly using the exam-type QG decoder and the factoid-type QG decoder. Furthermore, we employ reinforcement learning to enhance QG performance. Our experimental evaluation shows that our model significantly outperforms the compared baselines, with a BLEU-4 score increase from 5.31 to 6.48. Human evaluation also confirms that the questions generated by our model are answerable and appropriately difficult. Overall, theHandover QGframework offers a promising solution for improving QG performance in generating advanced questions for reading comprehension assessment.
Ho-Lam Chung, Ying-Hong Chan, Yao-Chung Fan
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 How Is the Stroke? Inferring Shot Influence in Badminton Matches via Long Short-term Dependencies
abstract
Identifying significant shots in a rally is important for evaluating players’ performance in badminton matches. While there are several studies that have quantified player performance in other sports, analyzing badminton data has remained untouched. In this article, we introduce a badminton language to fully describe the process of the shot, and we propose a deep-learning model composed of a novel short-term extractor and a long-term encoder for capturing a shot-by-shot sequence in a badminton rally by framing the problem as predicting a rally result. Our model incorporates an attention mechanism to enable the transparency between the action sequence and the rally result, which is essential for badminton experts to gain interpretable predictions. Experimental evaluation based on a real-world dataset demonstrates that our proposed model outperforms the strong baselines. We also conducted case studies to show the ability to enhance players’ decision-making confidence and to provide advanced insights for coaching, which benefits the badminton analysis community and bridges the gap between the field of badminton and computer science.
Wei-Yao Wang, Teng-Fong Chan, Wen-Chih Peng, Hui-Kuo Yang, Chih-Chuan Wang, Yao-Chung Fan
ACM Trans. Intell. Syst. Technol.6
2022 Extracting Crime Prosecution Elements based on Neural Machine Reading Comprehension Model
abstract
In this paper, we explore the task of extracting prosecution elements (text description about prosecution elements) in an indictment. We approach the prosecution element extraction problem by formulating it as a reading comprehension task. Specifically, our idea is to train a reading comprehension model to extract a text span to indicate the statement of a crime element according to an asked question. By such a reformulation, we leverage the power of neural machine reading models to prosecution element extraction task. Experimental evaluation demonstrates the feasibility of the machine reading reformulation. We also make our code and data available on https://github.com/NCHU-NLP-Lab/Legal-Document-Question-Answering
Jui-Ching Tsou, Kai-Yu Hsieh, Chen-Hua Huang, Yu-An Shih, Han-Cheng Yu, Yao-Chung Fan
IEEE Big Data6
2022 Hierarchical Cache Transformer: Dynamic Early Exit for Language Translation
abstract
The transformer model significantly improves the performance of natural language processing tasks. However, the downside of employing the transformer-based model is its heavy inference cost, which raises the concern of putting transformer-based models into industrial operations. Thus, studies for im-proving inference performance were reported. However, the major studies mainly consider classification tasks but not for natural language generation (NLG). In this paper, we propose Hierarchical Cache (HC) Transformer model tailored to NLG tasks. Our experiments show the feasible results on German-English translation dataset. The experiment result demonstrates that HC-Transformer can speed up inference by 32% with a 3% loss in performance.
Chih-Shuo Tsai, Ying-Hong Chan, Yao-Chung Fan
IJCNN3
2022 Misleading Inference Generation via Proximal Policy Optimization
Hsien-Yung Peng, Ho-Lam Chung, Ying-Hong Chan, Yao-Chung Fan
PAKDD (1)4
2022 Word Embedding Quantization for Personalized Recommendation on Storage-Constrained Edge Devices in a Smart Store
Yao-Chung Fan, Si-Ying Huang, Yung-Yu Chen, Lun-Chi Chen, Fang-Yie Leu
Mob. Networks Appl.1
2022 Emotion-cause pair extraction based on machine reading comprehension model
Ting-Wei Chang, Yao-Chung Fan, Arbee L. P. Chen
Multim. Tools Appl.2
2021 Exploring the Long Short-Term Dependencies to Infer Shot Influence in Badminton Matches
abstract
Identifying significant shots in a rally is important for evaluating players’ performance in badminton matches. While there are several studies that have quantified player performance in other sports, analyzing badminton data is remained untouched. In this paper, we introduce a badminton language to fully describe the process of the shot and propose a deep learning model composed of a novel short-term extractor and a long-term encoder for capturing a shot-by-shot sequence in a badminton rally by framing the problem as predicting a rally result. Our model incorporates an attention mechanism to enable the transparency of the action sequence to the rally result, which is essential for badminton experts to gain interpretable predictions. Experimental evaluation based on a real-world dataset demonstrates that our proposed model outperforms the strong baselines. The source code is publicly available at https://github.com/wywyWang/Shot-Influence.
Wei-Yao Wang, Teng-Fong Chan, Hui-Kuo Yang, Chih-Chuan Wang, Yao-Chung Fan, Wen-Chih Peng
ICDM5
2020 Product Quality Prediction with Convolutional Encoder-Decoder Architecture and Transfer Learning
abstract
Mining data collected from industrial manufacturing process plays an important role for intelligent manufacturing in Industry 4.0. In this paper, we propose a deep convolutional model for predicting wafer fabrication quality in an intelligent integrated-circuit manufacturing application. The wafer fabrication quality prediction is motivated by the need for improving product line efficiency and reducing manufacturing cost by detecting potential defective work-in-process (WIP) wafers. This work considers the following two crucial data characteristics for wafer fabrication. First, our model is designed to learn spatial correlation between quality measurements on WIP wafers and fabrication results through an encoder-decoder neural network. Second, we leverage the fact that different products share the same raw manufacturing process to enable the knowledge transferring between prediction models of different products. Performance evaluation on real data sets is conducted to validate the strengths of our model on quality prediction, model interpretability, and feasibility of transferring knowledge.
Hao-Yi Chih, Yao-Chung Fan, Wen-Chih Peng, Hai-Yuan Kuo
CIKM2
2020 Extractive Summarization by Rouge Score Regression Based on BERT
Bing-Hong Tsai, Yao-Chung Fan, Fang-Yie Leu
CISIS2
2019 Interpretable Multi-task Learning for Product Quality Prediction with Attention Mechanism
abstract
In this paper, we investigate the problem of mining multivariate time series data generated from sensors mounted on manufacturing stations for early product quality prediction. In addition to accurate quality prediction, another crucial requirement for industrial production scenarios is model interpretability, i.e., to understand the significance of an individual time series with respect to the final quality. Aiming at the goals, this paper proposes a multi-task learning model with an encoder-decoder architecture augmented by the matrix factorization technique and the attention mechanism. Our model design brings two major advantages. First, by jointly considering the input multivariate time series reconstruction task and the quality prediction in a multi-task learning model, the performance of the quality prediction task is boosted. Second, by incorporating the matrix factorization technique, we enable the proposed model to pay/learn attentions on the component of the multivariate time series rather than on the time axis. With the attention on components, the correlation between a sensor reading and a final quality measure can be quantized to improve the model interpretability. Comprehensive performance evaluation on real data sets is conducted. The experimental results validate that strengths of the proposed model on quality prediction and model interpretability.
Cheng-Han Yeh, Yao-Chung Fan, Wen-Chih Peng
ICDE2
2019 BERT for Question Generation
abstract
In this study, we investigate the employment of the pre-trained BERT language model to tackle question generation tasks.We introduce two neural architectures built on top of BERT for question generation tasks.The first one is a straightforward BERT employment, which reveals the defects of directly using BERT for text generation.And, the second one remedies the first one by restructuring the BERT employment into a sequential manner for taking information from previous decoded results.Our models are trained and evaluated on the question-answering dataset SQuAD.Experiment results show that our best model yields state-of-the-art performance which advances the BLEU 4 score of existing best models from 16.85 to 21.04.
Ying-Hong Chan, Yao-Chung Fan
INLG2
2018 Personalized Item-of-Interest Recommendation on Storage Constrained Smartphone Based on Word Embedding Quantization
Si-Ying Huang, Yung-Yu Chen, Hung-Yuan Chen, Lun-Chi Chen, Yao-Chung Fan
PAKDD (3)5
2018 On the semantic annotation of Wi-Fi SSID logs in mobile applications
Yao-Chung Fan, Chih-Wei Chang, Wang-Chien Lee, Kuo-Chen Wu, Arbee L. P. Chen
Pervasive Mob. Comput.1
2017 Enabling in-network aggregation by diffusion units for urban scale M2M networks
Yao-Chung Fan, Huan Chen 0002, Fang-Yie Leu, Ilsun You
J. Netw. Comput. Appl.1
2016 A framework for enabling user preference profiling through Wi-Fi logs
abstract
Understanding users is a key for many business applications. In this paper, we propose to pursue user preference understanding by their Wi-Fi logs collected from their mobile devices. As shown, Wi-Fi data are essentially of various information types and with noises. The challenges lie in how to refine relevant information from noisy Wi-Fi data. Aiming at the challenges, this paper proposes a data cleaning and information enrichment framework for enabling user preference understanding through Wi-Fi logs, and introduces a series of filters for cleaning, correcting, and refining Wi-Fi logs. A comprehensive experiment with real data collected from users is made to verify the effectiveness of the proposed techniques for cleaning noisy Wi-Fi data for user preference profiling. To the best of our knowledge, this work is the first attempt to study user behavior understanding by mining Wi-Fi logs.
Yao-Chung Fan, Kuan-Chieh Tung, Kuo-Chen Wu, Arbee L. P. Chen
ICDE1
2016 A Framework for Enabling User Preference Profiling through Wi-Fi Logs
abstract
Nowadays, mobile devices have become a ubiquitous medium supporting various forms of functionality and are widely accepted for commons. In this study, we investigate using Wi-Fi logs from a mobile device to discover user preferences. The core ideas are two folds. First, every Wi-Fi access point is with a network name, normally a human-readable string, called SSID (Service Set Identifier). Since SSIDs are often with semantics, from which we can infer the place where the user stayed. Second, a Wi-Fi log is produced when the user is near a Wi-Fi access point. A high frequency of a consecutively observed SSID implies a long stay duration at a place. To the best of our knowledge, our work is the first attempting to understand users from the collected Wi-Fi logs from mobile devices. However, Wi-Fi logs are essentially of various information types and with noises. How to assess the information types, eliminate irrelevant information, and clean up the noises within partial-informative SSIDs are therefore keys for profiling user preferences over Wi-Fi logs. In this paper, we propose a data cleaning and information enrichment framework for enabling the user preference understanding through collected Wi-Fi logs, and introduce a data clean framework for cleaning, correcting, and refining Wi-Fi logs. In addition, a comprehensive experiment with data collected from users is made to verify the effectiveness of the proposed techniques for cleaning noisy Wi-Fi data for user preferences profiling. The experiment results demonstrate the effectiveness of the proposed framework for profiling user preferences through Wi-Fi logs.
Yao-Chung Fan, Kuan-Chieh Tung, Kuo-Chen Wu, Arbee L. P. Chen
IEEE Trans. Knowl. Data Eng.1
2013 GreenSensing: A Fine Grained Power Monitoring System for a Network of Computers
abstract
Recent studies have shown the necessity of fine grained power usage visibility to encourage user behavior energy conservation. Existing works toward this direction are mainly focused on residential monitoring scenarios. This study considers another important scenario of providing fine grained power usage information over a set of networked computers. Such scenario could be applied to a wide range of workspaces, including research labs in colleges and business offices in companies. To our best knowledge, no existing systems provide cost-effective solutions for such scenario. To this end, we present GreenSensing which provides real time individual power usage estimation by leveraging the fact that the amount of power usage is highly correlated to CPU usage rate of a computer. As CPU usage can be obtained by a software program, we can have the power information without real metering instruments, and therefore have the benefit of zero infrastructure cost. However, due to the indirectness, before the estimation works, a proper calibration for computers is required, which needs lots of human intervention. To this problem, we propose a novel framework that empowers GreenSensing to automatically and simultaneously calibrate multiple computers on the fly. We show through experiments GreenSensing provides only 4.1% to 6.2% estimation error on individual power usage.
Yao-Chung Fan, Huan Chen 0002
DCOSS1
2013 TeleEye: Enabling Real-time Geospatial Query Answering with Mobile Crowd
abstract
In this paper, we present TeleEye, a real time geographic information inquiring system based on the mobile crowd sourcing. The TeleEye system supports novel information inquiring services beyond traditional geospatial information system by providing a real time information, such as if a tennis court is now occupied. This paper summarizes the design, the architecture, and the prototype of the TeleEye system.
Yao-Chung Fan, Cheng Teng Iam, Gia Hao Syu, Wei Hong Lee
DCOSS1
2013 LocalSense: An Infrastructure-Mediated Sensing Method for Locating Appliance Usage Events in Homes
abstract
In this paper, we introduce a novel technique called Local Sense for detecting appliance usage events in a household by leveraging existing home power infrastructure. Specifically, the Local Sense technique works by purely analyzing the records taken from the main power meter. The main power meter is installed in a household to measure the overall real time aggregated power consumption in terms of wattage and voltage for the whole load of appliances in a household. The Local Sense technique requires no additional apparatuses being installed in existing household power infrastructure, which significantly reduces the cost for enabling the location-aware power usage information over existing techniques. The idea behind the Local Sense system is to reflect the fact that different outlets are with different power line impedances, which causes different power dissipation even for the same appliance usage. By identifying the additional power dissipation, detecting the appliance usage events at outlet levels is therefore possible. The proposed technique is validated in a research lab and preliminary results demonstrate the effectiveness of the proposed system.
Hung-Yuan Chen, Chien-Liang Lai, Huan Chen 0002, Lun-Chia Kuo, Hsi-Chuan Chen, Jyh-Shyan Lin, Yao-Chung Fan
ICPADS7
2013 Indoor Place Name Annotations with Mobile Crowd
abstract
With the popularity of mobile devices, numerous mobile applications have been and will continue to be developed for various interesting usage scenarios. Riding this trend, recent research community envisions a novel information retrieving and information-sharing platform, which views the users with mobile devices and being willing to accept crowd sourcing tasks as crowd sensors. With the neat idea, a set of crowd sensors applications have emerged. Among the applications, the geospatial information systems based on crowd sensors show significant potentials beyond traditional ones by providing real time geospatial information. In the applications, user positioning is of great importance. However, existing positioning techniques have their own disadvantages. In this paper, we study using pervasive Wi-Fi access point as a position indicator. The major challenge for using Wi-Fi access point is that there is no mechanism for mapping observed Wi-Fi signals to human-defined places. To this end, our idea is to employ crowd sourcing model to perform place name annotations by mobile participants to bridge the gap between signals and human-defined places. In this paper, we propose schemes for effectively enabling based-based place name annotation, and conduct real trials with recruited participants to study the effectiveness of the proposed schemes. The experiment results demonstrate the effectiveness of the proposed schemes over existing solutions.
Yao-Chung Fan, Wei Hong Lee, Cheng Teng Iam, Gia Hao Syu
ICPADS1
2012 Energy Efficient Schemes for Accuracy-Guaranteed Sensor Data Aggregation Using Scalable Counting
abstract
Sensor networks have received considerable attention in recent years, and are employed in many applications. In these applications, statistical aggregates such as Sum over the readings of a group of sensor nodes are often needed. One challenge for computing sensor data aggregates comes from the communication failures, which are common in sensor networks. To enhance the robustness of the aggregate computation, multipath-based aggregation is often used. However, the multipath-based aggregation suffers from the problem of overcounting sensor readings. The approaches using the multipath-based aggregation therefore need to incorporate techniques that avoid overcounting sensor readings. In this paper, we present a novel technique named scalable counting for efficiently avoiding the overcounting problem. We focus on having an (ε, δ) accuracy guarantee for computing an aggregate, which ensures that the error in computing the aggregate is within a factor of ε with probability (1 - δ). Our schemes using the scalable counting technique efficiently compute the aggregates under a given accuracy guarantee. We provide theoretical analyses that show the advantages of the scalable counting technique over previously proposed techniques. Furthermore, extensive experiments are made to validate the theoretical results and manifest the advantages of using the scalable counting technique for sensor data aggregation.
Yao-Chung Fan, Arbee L. P. Chen
IEEE Trans. Knowl. Data Eng.1
2010 Efficient and Robust Schemes for Sensor Data Aggregation Based on Linear Counting
abstract
Sensor networks have received considerable attention in recent years, and are often employed in the applications where data are difficult or expensive to collect. In these applications, in addition to individual sensor readings, statistical aggregates such as Min and Count over the readings of a group of sensor nodes are often needed. To conserve resources for sensor nodes, in-network strategies are adopted to process the aggregates. One primitive in-network aggregation strategy is the tree-based aggregation, where the aggregates are computed from leaves to the root of a spanning tree over a sensor network. However, a shortcoming with the tree-based aggregation is that it is not robust against communication failures, which are common in sensor networks. One of the solutions to overcome this shortcoming is to enable multipath routing, by which each node broadcasts its reading or a partial aggregate to multiple neighbors. However, multipath routing-based aggregation typically suffers from the problem of overcounting sensor readings. In this study, we propose two schemes based on the linear counting technique to deal with the overcounting problem. These two schemes process aggregates by statically and dynamically, respectively, allocating space for the use of the linear counting technique. Both schemes provide the same accuracy guarantee but involve different communication costs. Through extensive experiments with real-world and synthetic data, we demonstrate the efficiency and effectiveness of using these two schemes as solutions for processing aggregates in a sensor network. The experiments also show that the scheme that dynamically allocates the space often outperforms the other one in terms of energy conservation since it requires less space to satisfy an accuracy constraint.
Yao-Chung Fan, Arbee L. P. Chen
IEEE Trans. Parallel Distributed Syst.1
2009 An Approximation Algorithm for Optimizing Multiple Path Tracking Queries over Sensor Data Streams
Yao-Chung Fan, Arbee L. P. Chen
DEXA1
2009 Energy-Efficient Sensor Data Acquisition Based on Periodic Patterns
abstract
Wireless sensor networks have received considerable attention in recent years and played an important role in data collection applications. Sensor nodes usually have limited supply of energy. Therefore, a major consideration for developing sensor network applications is to conserve the energy for sensor nodes. In this paper, we propose a novel energy-efficient data acquisition algorithm based on the periodic patterns derived from past sensor readings. Our key observation is that sensor readings often exhibit periodic patterns, e.g., the daily cycle of temperature readings, and the patterns provide opportunities for reducing energy consumption for sensor data acquisition. We exploit the patterns and use the patterns to build a statistic model for predicting sensor readings. In our approach, sensor data acquisition is needed only when acquired readings are unpredictable. Therefore the energy for sensor data acquisition and the associated radio communications can be conserved. The experiments performed with real data validate the effectiveness and efficiency of our approach.
Guan-Rong Lin, Yao-Chung Fan, En Tzu Wang, Arbee L. P. Chen
ICPADS2
2008 Efficient and robust sensor data aggregation using linear counting sketches
abstract
Sensor networks have received considerable attention in recent years, and are often employed in the applications where data are difficult or expensive to collect. In these applications, in addition to individual sensor readings, statistical aggregates such as Min and Count over the readings of a group of sensor nodes are often needed. To conserve resources for sensor nodes, in-network strategies are adopted to process the aggregates. One primitive in-network aggregation strategy is the tree-based aggregation, where the aggregates are computed along a spanning tree over a sensor network. However, a shortcoming with the tree-based aggregation is that it is not robust against communication failures, which are common in sensor networks. One of the solutions to overcome this shortcoming is to enable multi-path routing, by which each node broadcasts its reading or a partial aggregate to multiple neighbors. However, multi-path routing based aggregation typically suffers from the problem of overcounting sensor readings. In this study, we propose using the linear counting sketches for multi-path routing based in-network aggregation. We claim that the use of the linear counting sketches makes our approach considerably more accurate than previous approaches using the same sketch space. Our approach also enjoys low variances in term of the aggregate accuracy, and low overheads either in computations or sketch space. Through extensive experiments with real-world and synthetic data, we demonstrate the efficiency and effectiveness of using the linear counting sketches as a solution for the in-network aggregation.
Yao-Chung Fan, Arbee L. P. Chen
IPDPS1
2004 Compressing a Directed Massive Graph using Small World Model
abstract
In this article we propose a method that can compress a small-world-like massive digraph (directed graph) into at most half size of the original representation represented by using adjacency list. This method also provides a fast decompression algorithm that works as quickly as adjacency list does. In this paper we deal with the problem of finding a compact representation of a graph from which the vertices adjacent to any specified vertex can be easily determined.
Fang-Yie Leu, Yao-Chung Fan
Data Compression Conference2