VLDB 2026 Research / reviewers in the wild / expert
Chihung Chi
dblp:31/10001 · also Chi-Hung Chi
· DBLP profile ↗
132ranked-venue papers
29as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 35 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 32 · 6 first-author · 5 since 2021Systems, architecture and hardware · 23 · 11 first-author · 5 since 2021Software engineering, systems software and programming languages · 21 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 3 since 2021Computer networks · 8 · 1 first-author · 5 since 2021Security and privacy · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sentences Based Adversarial Attack on AI-Generated Text DetectorsabstractThe widespread use of AI-generated text has introduced significant security concerns, driving the need for reliable detection systems. However, recent studies reveal that neural network-based detectors are vulnerable to adversarial examples. To improve the robustness of such classifiers, a number of adversarial attack strategies have been developed, particularly in the context of text sentiment classification. Most existing adversarial attack methods focus on the semantics of individual words or sentences, often neglecting the broader contextual semantics of the entire text-particularly in the case of long AI-generated text. This limitation frequently results in adversarial examples that lack fluency and coherence. In this paper, we propose a novel method calledSentence-based Adversarial attack on AI-Generated Text detectors (SAGT), which generates linguistically fluent adversarial examples by inserting model-generated sentences into the original text. To ensure contextual semantic consistency, we extract important keywords from the original text-selected based on changes in the detector's confidence score-and incorporate them into the generated sentences. Extensive experimental results demonstrate that adversarial examples crafted bySAGTcan effectively evade AI-generated text detectors. Rongxin Tu, Xiangui Kang, Chee-Wei Tan 0001, Chihung Chi, Kwok-Yan Lam |
IEEE Trans. Big Data | 4 |
| 2026 | TTACO: Trusted Time-Aware Computing Offloading in Air-Ground Integrated NetworksabstractAs efficiency and security requirements emerge in computing offloading fields, trusted computing offloading has grabbed tremendous sights, especially in Air-Ground Integrated Networks (AGINs). Although traditional computing offloading studies have attempted to employ trust to identify malicious devices without mobility, the integration of trust management and computing offloading has not been explored in a high-hazardous environment. Then, with the increasing requirement of adaptation and security in new-generation network architectures (i.e., AGINs), trust management plays a pivotal role in expanding the implementation of trusted computing offloading. Therefore, we propose a trusted time-aware computing offloading mechanism in AGINs based on communication, computing offloading, and trust modeling. Specifically, a maximum reward optimization formulation is designed to generate an optimal offloading strategy, considering service utility trust, opinion service trust, computing efficiency trust, time constraints, rewards, and task volumes. Based on problem analysis, a solution paradigm is condensed to improve the probability of seeking an efficient solution. Moreover, a greedy search algorithm is designed to find the possible efficient solution, including the default solution setting stage, the boundary search stage, and the greedy reallocation stage. Due to the trust threshold affecting identification results, a dichotomy-based dynamic trust threshold method is employed to empower trusted computing offloading mechanism with the adaptive capability in AGINs. Extensive experiments show that our mechanism outperforms other baselines in terms of accuracy, precision, recall, F-measure, task success rate, and average response time. Yue Cao 0002, Zhenning Wang, Chihung Chi, Wei Ren 0002, Wei Wang 0050 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection SystemsabstractAudio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval. Kwok Chin Yuen, Jia Qi Yip, Chihung Chi, Kwok-Yan Lam |
INTERSPEECH | 4 |
| 2025 | Reliability Assessment of Multiprocessor System Based on Exchanged Crossed Cube NetworksabstractABSTRACT With the increasingly widespread application of multiprocessor systems, some processors in multiprocessor systems are inevitably prone to malfunctions. The reliability and effectiveness of the system are key issues. As a standard for measuring system fault tolerance, connectivity, and edge connectivity have many drawbacks. Therefore, Haray proposed conditional connectivity by restricting the connected components in disconnected subgraphs to satisfy certain properties, where and represent the interconnection network and its set of faulty vertices, respectively. Restricted connectivity is a special type of conditional connectivity. Exchanged crossed cube, as a deformation of hypercube, has more favorable properties, such as smaller diameter, smaller link size, and lower cost. We prove that the 2‐restricted connectivity of the exchanged crossed cubes is for . Xuanli Liu, Weibei Fan, Jing He 0004, Zhijie Han 0001, Chihung Chi |
Concurr. Comput. Pract. Exp. | 5 |
| 2025 | Heterogeneous Parallel Key-Insulated Multi-Receiver Signcryption Scheme for IoVabstractThe rapid growth of electric vehicle and autonomous vehicle populations has led to explosive expansion of IoV data being transmitted in the wireless communication infrastructure. Advances in IoV technologies also resulted in more complex and dynamic communication protocols/patterns, which are hard for the underlying wireless network to satisfy. Besides, security considerations of IoV communications require that key management must be stringently prohibit global failure mode of key management, meaning that, if a single IoV node compromises its private key, it will not lead to total security failure of the entire IoV network. To address these issues, in this paper, we propose a heterogeneous parallel key-insulated multi-receiver signcryption scheme for IoV (HPKI-MRSC). Firstly, the proposed scheme can realize one-to-many heterogeneous transmission, in which RSUs are deployed on certificateless cryptography (CLC) system, while vehicles are allocated in identity-based cryptography (IBC) system. In this manner, we observe that message transmission efficiency is improved greatly. Secondly, the parallel key-insulated mechanism can employ two helper keys to update private key periodically, and then solve key disclosure problem. Finally, when the number of receiver n is greater than or equal to 3, the proposed scheme has a lower signcryption overhead than other comparative schemes, and thus it is more suitable for IoV. Yingzhe Hou, Yue Cao 0002, Hu Xiong, Debiao He, Chihung Chi, Kwok-Yan Lam |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | All Points Guided Adversarial Generator for Targeted Attack Against Deep Hashing RetrievalabstractDeep hashing has been widely used in image retrieval tasks, while deep hashing networks are vulnerable to adversarial example attacks. To improve the deep hashing networks’ robustness, it is essential to investigate adversarial attacks on the networks, especially targeted attacks. Among the existing targeted attacks for hashing, the generation-based targeted attack methods have attracted increasing attention due to their efficiency in generating adversarial examples. However, these methods supervise the generation of adversarial examples solely with the hash codes of positive samples, without employing the hash codes of all points in the training set to directly participate in supervisory training, thereby making the attack less effective. Since the hash codes of the training set samples are generated by a well-trained hashing model, these hash codes retain rich semantic information of their corresponding samples, highlighting the necessity of sufficiently utilizing them. Therefore, in this paper, we propose a targeted attack method that utilizes all points’ hash codes in the training set to guide the generation of adversarial attack examples directly. Specifically, we first decode the target label to obtain the corresponding feature map. Then, we concatenate the feature map with the query image and feed them into an encoder-decoder network that employs a skip-connection strategy to obtain a perturbed example. Furthermore, to guide adversarial example generation, we introduce a loss function that exploits the similarities between the perturbed example’s hash code and all points’ hash codes in the training set, thereby making sufficient utilization of the rich semantic information in these hash codes. Experimental results illustrate that our method outperforms the state-of-the-art targeted attack methods in targeted attack effectiveness and transferability. The code is available athttps://github.com/rongxintu3/APGA. Rongxin Tu, Xiangui Kang, Chee-Wei Tan 0001, Chihung Chi, Kwok-Yan Lam |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | A Hierarchical Encrypted Compression Scheme for Intra-Vehicle NetworkabstractThe CAN bus is the most widely used bus for intra-vehicle communication due to its high transmission stability, excellent real-time communication capability, and relatively low cost. As the number of ECUs grows, the CAN bus load increases and thus raises the possibility of data transmission delays and errors. Message compression based on the differential algorithm has been proposed to reduce the CAN bus load. However, current works do not consider the security problems of the CAN bus. Attackers can manage to acquire the original messages before compression and disturb the message statistics to decrease compression rate by injecting malicious frames. In this paper, we propose a secure compression mechanism for the intra-vehicle network, including an improved compression algorithm, a stream key distribution scheme, and a hierarchical encryption scheme. Formal verification results show that the proposed scheme can achieve mutual authentication, message confidentiality and integrity, resist replay attacks, and support secure compression. Evaluations using real vehicle data on 16 MHz boards show the average communication overhead can be reduced by 46.38% compared to the original messages. Performance analysis results show our scheme can reduce computational overhead on compression by 31.82% and 19.43% on decompression compared to related schemes. Jin Cao 0001, Zejian Li, Ben Niu 0001, Kwok-Yan Lam, Chihung Chi, Hui Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | An Accountable GAKA Protocol With Changeable Thresholds and Verifiable Shares in UAVs-Assisted IoVs for Emergency RescueabstractUnmanned aerial vehicles (UAVs) equipped with high line-of-sight communications have been explored as a complement to emergency communication vehicles (ECVs), particularly when ground Internet of Vehicles (IoVs) is dysfunctional resulting from natural disasters. To protect the communication security and integrity of this air-ground integrated networks created in an untrusted and open wireless environment, authentication and key agreement (AKA) is an essential mechanism to establish a secure communication channel by negotiating a session key. Nonetheless, general AKA protocols are typically not ideal for time-sensitive and computing-intensive UAVs-assisted IoVs emergency rescue, in that they are unable to simultaneously satisfy efficiency and key security requirements including accountability, resilience, thresholds changeability, shares verifiability, group adaptability, and key update. To address this challenge, this paper proposes an accountable group authentication and key agreement (GAKA) protocol with changeable thresholds and verifiable shares supporting adaptive group memberships and updatable keys (i.e. SecER). To safeguard accurate rescue decisions for ground ECVs enabled by reliable collaboration among multiple UAVs, we propose a pseudonym mechanism that aims to provide accountability for UAVs. To achieve resilience and thresholds changeability, our SecER utilizes secret sharing coupled with random parameters to seamlessly transform our GAKA protocol into one running with a new threshold in one round (i.e. round-optimal), rather than two rounds. Finally, verifiable parameters and updatable keys are respectively applied to counter deception attacks caused by maliciously distributed shares and to support adaptive network topology (i.e. UAVs joining and leaving). Extensive simulations show that compared to the state-of-the-art approaches, our SecER is superior in balancing security and efficiency. Di Wang 0025, Yue Cao 0002, Kwok-Yan Lam, Chihung Chi, Kim-Kwang Raymond Choo |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Multi-Layer Transfer Learning for Cross-Domain Recommendation Based on Graph Node Representation Enhancement
Jie Nie, Niantai Jing, Jianliang Xu, Xiaodong Wang 0006, Xuesong Gao, Mingxing Jiang, Chihung Chi, Zhiqiang Wei 0002 |
IEEE Trans. Multim. | 8 |
| 2025 | Copy-Move Forgery Image Detection Based on Cross-Scale Modeling and Alternating RefinementabstractThe detection of image tampering, specifically copy detection, is an important problem in many domains such as military, media, and public opinion outlets. Effective means to detect such tampering is crucial in controlling the dissemination of false information. However, a major challenge in achieving high detection accuracy lies in the variability of the scale of the copied targets. To tackle this problem, we introduce an all-encompassing methodology called Cross-Scale Modeling and Alternating Refinement (CANet) to detect the genuine source and tampered region at the pixel level. CANet consists of three modules: the Cross-Scale Similar Region Detection (CS) module, the Edge-Supervised Tamper Region Detection (ET) module, and the Alternating Refinement (AR) module. The CS module extracts coarse similar region features by cross-scale correlation modeling, which can alleviate the scale gap between the source and tampered region. The obtained coarse similar region feature is refined by the AR module, in which we introduce the source and the tampered region as the auxiliary information and employ a two-stage process that sequentially models their global feature representations. The tampered region used in the AR module is obtained from the ET module using edge supervision with a salient edge selection scheme, and the source region is generated by the implicit modeling. We conducted experiments on the USC-ISI, CASIA v2.0, CoMoFoD, and MICC-F220 datasets separately. Results show that our method outperforms the state-of-the-art. Jingyu Wang 0005, Jie Nie, Niantai Jing, Xiaodong Wang 0006, Chihung Chi, Zhiqiang Wei 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | Trustworthy and Responsible AI for Information and Knowledge Management SystemabstractThe way research and business manage and utilize knowledge is undergoing a significant transformation, driven by Artificial Intelligence (AI). Deep learning and machine learning are emerging as powerful tools for optimizing knowledge management systems, leading to more informed and productive development. AI offers unique solutions for organizations struggling with information overload and inefficient knowledge transfer. These AI models can significantly improve data management and utilization. Imagine an AI-powered system that streamlines onboarding processes, provides precise answers to various queries, and even captures the valuable tacit knowledge (implicit skills and expertise) often residing within individuals. AI bridges the gap between explicit knowledge (easily documented information) and tacit knowledge, fostering a more comprehensive and accessible knowledge base. However, such AI systems solicit trustworthy and responsible approaches to mitigate potential misuse and malfunction. In this workshop, we aim to gather researchers and engineers from academia and industry to discuss the latest advances in trustworthy and responsible AI solutions for information and knowledge management systems. Huaming Chen, Jun Zhuang 0004, Yu Yao 0005, Wei Jin 0009, Haohan Wang, Yong Xie 0002, Chihung Chi, Kim-Kwang Raymond Choo |
CIKM | 7 |
| 2024 | A Message-based Lightweight Session Key Distribution Scheme for Intra-Vehicle NetworkabstractModern vehicles are equipped with ECU nodes and intra-vehicle buses. Among these, the CAN bus stands out as the most widely utilized intra-vehicle bus due to its affordability and straightforward deployment. However, the CAN bus suffers from significant security vulnerabilities, such as the absence of access control, identity authentication, message encryption, and authentication. In this paper, we propose a lightweight message-based key distribution scheme aimed at addressing these vulnerabilities. Our scheme facilitates mutual authentication during key distribution and assigns a unique key to each class of message. Formal verification using the Scyther tool demonstrates that our protocol achieves mutual authentication and effectively mitigates several protocol attacks, including replay, tampering, and manin-the-middle attacks. We evaluate our scheme using Arduino UNO boards. Performance analysis indicates that our scheme exhibits superior security capabilities compared to other related schemes and outperforms them in terms of communication and computation overheads. Jin Cao 0001, Zejian Li, Yurong Luo, Kwok-Yan Lam, Chihung Chi, Hui Li 0006 |
GLOBECOM | 6 |
| 2024 | Early Discovery of Key Innovative Publications by Analyzing Emerging Topic Trends
Junfeng Wu 0010, Xiangmin Zhou, Guangyan Huang, Borui Cai, Guang-Li Huang, Hui Zheng 0001, Chihung Chi, Jing He 0004 |
WISE (1) | 7 |
| 2024 | SE-shapelets: Semi-supervised Clustering of Time Series Using Representative ShapeletsabstractShapelets that discriminate time series using local features (subsequences) are promising for time series clustering. Existing time series clustering methods may fail to capture representative shapelets because they discover shapelets from a large pool of uninformative subsequences, and thus result in low clustering accuracy. This paper proposes a Semi-supervised Clustering of Time Series Using Representative Shapelets (SE-Shapelets) method, which utilizes a small number of labeled and propagated pseudo-labeled time series to help discover representative shapelets, thereby improving the clustering accuracy. In SE-Shapelets, we propose two techniques to discover representative shapelets for the effective clustering of time series. (1) A salient subsequence chain (SSC) that can extract salient subsequences (as candidate shapelets) of a labeled/pseudo-labeled time series, which helps remove massive uninformative subsequences from the pool. (2) A linear discriminant selection (LDS) algorithm to identify shapelets that can capture representative local features of time series in different classes, for convenient clustering. Experiments on UCR time series datasets demonstrate that SE-shapelets discovers representative shapelets and achieves higher clustering accuracy than counterpart semi-supervised time series clustering methods. Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi |
Expert Syst. Appl. | 5 |
| 2024 | Secure and privacy-preserving sharing of personal health records with multi-party pre-authorization verification
Kheng Leong Tan, Chihung Chi, Kwok-Yan Lam |
Wirel. Networks | 2 |
| 2023 | Interdependence analysis on heterogeneous data via behavior interior dimensionsabstractInterdependent dimensions including categorical and continuous variables can be seen commonly as heterogeneous behavioral data in the real world. Mixed-type objects are more or less associated in terms of certain coupling relationships. The usual representation of such behavioral data is an information table with explicit behavior exterior dimensions (i.e. the original attributes to describe data heterogeneity), assuming the independence of dimensions and the independence of objects. However, both variables and objects are actually very often interdependent on one another either explicitly or implicitly in functional and semantic manners. Limited research has been done in analyzing such interactions among dimensions and those relationships among objects, leading to the learning results to be more local than global. This paper proposes the interdependence analysis to capture the functional multifarious relationships among attributes and among objects in heterogeneous data by addressing the coupling context and coupling weights in unsupervised learning. Such global couplings consider the interactions within discrete dimensions, within numerical attributes and across them, as well as the relationships within an individual object and between multiple objects, to form the attribute-based and object-based coupled data representation schemes based on feature conversion and neighborhood calculation. In addition, we interpret both the representation models via implicit behavior interior dimensions (i.e. the newly defined attributes to model data interdependence) to explain the intrinsic rationales for the superiority of our proposed methods. This work explicitly models the coupling of multiple attributes and the coupling of multiple objects for heterogeneous data sets, demonstrated by various data mining and machine learning applications, such as cluster structure analysis, data clustering evaluation, and data density comparison. Moreover, the sensitivity study is carried out to tune the neighborhood parameter and weight parameter, and the scalability analysis is explored to test the robustness of both models. Extensive experiments on a series of synthetic data sets and multiple UCI data sets show that our proposed framework can effectively capture the global couplings of both heterogeneous variables and mixed-type objects, and is superior to the traditional way as well as the state-of-the-art approaches, which is also verified by statistical analysis. Can Wang 0004, Chihung Chi, Lina Yao 0001, Alan Wee-Chung Liew, Hong Shen 0001 |
Knowl. Based Syst. | 2 |
| 2023 | Entity alignment via graph neural networks: a component-level studyabstractAbstract Entity alignment plays an essential role in the integration of knowledge graphs (KGs) as it seeks to identify entities that refer to the same real-world objects across different KGs. Recent research has primarily centred on embedding-based approaches. Among these approaches, there is a growing interest in graph neural networks (GNNs) due to their ability to capture complex relationships and incorporate node attributes within KGs. Despite the presence of several surveys in this area, they often lack comprehensive investigations specifically targeting GNN-based approaches. Moreover, they tend to evaluate overall performance without analysing the impact of individual components and methods. To bridge these gaps, this paper presents a framework for GNN-based entity alignment that captures the key characteristics of these approaches. We conduct a fine-grained analysis of individual components and assess their influences on alignment results. Our findings highlight specific module options that significantly affect the alignment outcomes. By carefully selecting suitable methods for combination, even basic GNN networks can achieve competitive alignment results. Yanfeng Shu, Guangyan Huang, Chihung Chi, Jing He 0004 |
World Wide Web (WWW) | 4 |
| 2022 | Emerging Scientific Topic Discovery by Finding Infrequent Synonymous Biterms
Junfeng Wu 0010, Guangyan Huang, Roozbeh Zarei, Jianxin Li 0001, Guang-Li Huang, Hui Zheng 0001, Jing He 0004, Chihung Chi |
PAKDD (1) | 8 |
| 2022 | A polynomial-time algorithm for simple undirected graph isomorphismabstractIn the author list, "Ferry Sansoto" should be Ferry Susanto.• To reflect more accurately the contribution of the article, the title should be changed to "A permutation and equinumerosity based polynomial-time algorithm for simple undirected graph isomorphism."• In the abstract, the "Pythagorean Triples Theorem" should be removed.• In the abstract, "squared sums of elements" should be "nth power sums."• In Section 2.2, "and the sum of the individual squared elements.By checking two sums," should be ", the sum of the individual squared elements and until the sum of the nth power of the nth element in the array.By checking these sums,"• In Section 2.2, "For both vertex and edge arrays of row/column sum based on the vertex and edge adjacency matrices, if and only if one array is a permutation of another one, the corresponding two graphs are isomorphic."should be "For both the vertex and edge arrays of row/column sum based on the vertex and edge adjacency matrices, if and only if one array is a permutation of another one and the corresponding edge and vertex's adjacent relationship has been preserved, the corresponding two graphs are isomorphic." Jing He 0004, Guangyan Huang, Jie Cao 0001, Zhiwang Zhang, Hui Zheng 0001, Peng Zhang 0063, Roozbeh Zarei, Ferry Susanto, Ruchuan Wang 0001, Yimu Ji 0001, Weibei Fan, Zhijun Xie, Xiancheng Wang, Mengjiao Guo, Chihung Chi, Jiekui Zhang, Youtao Li, Xiaojun Chen 0001, Yong Shi 0001, André Van Zundert |
Concurr. Comput. Pract. Exp. | 15 |
| 2022 | A multiple feature fusion framework for video emotion recognition in the wildabstractSummary Human emotions can be recognized from facial expressions captured in videos. It is a growing research area in which many have attempted to improve video emotion detection in both lab‐controlled and unconstrained environments. While existing methods show a decent recognition accuracy on lab‐controlled datasets, they deliver much lower accuracy in a real‐world uncontrolled environment, where a variety of challenges need to be addressed such as variations in illumination, head pose, and individual appearance. Moreover, automatically identifying the key frames consisting of the expression from real‐world videos is another challenge. In this article, to overcome these challenges, we provide a video emotion recognition via multiple feature fusion method. First, a uniform local binary pattern (LBP) and the scale‐invariant feature transform features are extracted from each frame in the video sequences. By applying a random forest classifier, all of the static frames are then labelled by the related emotion class. In this way, the key frames can be automatically identified, including neutral and other expressions. Furthermore, from the key frames, a new geometric feature vector and the LBP from three orthogonal planes are extracted. To further improve robustness, audio features are extracted from the video sequences as an additional dimension to augmenting visual facial expression analysis. The audio and visual features are fused through a kernel multimodal sparse representation. Finally, the corresponding emotion labels to the video sequences can be assigned when a multimodal quality measure specifies the quality of each modality and its role in the decision. The results on both acted facial expressions in the Wild and MMI datasets demonstrate that the proposed method outperforms several counterpart video emotion recognition methods. Najmeh Samadiani, Guangyan Huang, Wei Luo 0001, Chihung Chi, Yanfeng Shu, Rui Wang 0008, Tuba Kocaturk |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | An adaptive simulated annealing and artificial fish swarm algorithm for the optimization of multi-depot express delivery vehicle routingabstractIn this paper, the Capacitated Vehicle Routing Problem (CVRP) of multi-depot express delivery is investigated based on the actual express delivery business in Beijing and driving intention-based road network. An Adaptive Simulated Annealing and Artificial Fish Swarm Algorithm (A-SAAFSA) is proposed to solve the CVRP. The basic ideas are use a “certainty” probability to accept the worst solution through the Metropolis criterion in the search process, and a strategy of adjusting the swimming direction to avoid falling into the local optimal solution. Moreover, an adaptive visual strategy, which adjusts the visual range adaptively in real time according to the current solution quality, is used to ensure the efficient searching and accuracy of the algorithm. Experimental results show that the A-SAAFSA algorithm outperforms four well-known algorithms, namely simulated annealing and artificial fish swarm algorithm, artificial fish swarm algorithm, simulated annealing algorithm, and genetic algorithm. Mengfei Yuan, Xiu Kan, Chihung Chi, Le Cao, Huisheng Shu, Yixuan Fan |
Intell. Data Anal. | 3 |
| 2022 | sAuth: a hierarchical implicit authentication mechanism for service robotsabstractAbstract With advances in networks, Artificial Intelligence (AI), and the Internet of Things, humanoid robots are rising in many areas, including elderly care, companion, education, and services in public sectors. Given their sensing and communication functionality, information leakage and unauthorized access will be of big concern. Very often, authentication techniques for service robots, especially those related to behavioral identification, have been developed, in which behavior models are created using raw data from sensors. However, behavioral-based authentication and re-authentication is still an open area for research, including cold start problems, accuracy, and uncertainty. This paper proposes a hierarchical implicit authentication system by joint built-in sensors and trust evaluation, coined sAuth, which exploits sensor data-based sliding window trust model to identify the service robot and its expected users. In order to mitigate the fluctuations of identification results in the real world environment, the trust evaluation is computed via combining the weighted intermediate identification probability of various small sliding windows. The performance of sAuth is evaluated under different scenarios where we show that (i) approximately 5–7% higher accuracy and 2–18% lower equal error rate can be achieved by our method compared to other works; and (ii) the hierarchical scheme with joint sensors and trust sliding windows improves the authentication accuracy significantly by comparing it with only sensor-based authentication. Pengming Zhang, Chihung Chi |
J. Supercomput. | 5 |
| 2022 | MUSH: Multi-Stimuli Hawkes Process Based Sybil Attacker Detector for User-Review Social NetworksabstractUser-Review Social Networks (URSNs) now become the targets of Sybil attacks, where fake reviews are posted by attackers to raise the reputation of listed services or products. Unlike previous fake accounts on Twitter or Weibo, Sybil attackers on URSNs are also genuine users in many cases, which presents a challenge for the existing fake review detection system. Hence, the main purpose of detecting Sybil attackers is to profile abnormal behaviors of users, and we propose a novel Sybil detection model named MUSH to extract long-term features of user behaviors on URSNs. First, aiming to measure the uncertainty of user behaviors, four entropy-based preference models are designed to quantify user preferences, including catering preferences, price preferences, word-of-mouth preferences, and rating preferences. Second, in order to extract temporal logic features, a novel multi-stimuli Hawkes process is introduced by combining external incentives and internal incentives, which can detect abnormal event sequences of posted reviews. This approach is quite different from previous solutions which mostly use direct/indirect graph models or user-related features for detection. Finally, by integrating entropy-based preferences with temporal logic features, a smart Sybil detection model is proposed based on a binary classification approach. Extensive experimental results indicate that MUSH can effectively detect Sybil attackers with a high detection accuracy. By comparing with other approaches, MUSH is quite suitable for adversarial network environments, which could provide better security services to social networks. Chen Lyu 0002, Chihung Chi |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Predictable Model for Detecting Sybil Attacks in Mobile Social NetworksabstractMobile Social Networks have become one of the most convenient services for users to share information everywhere. This crowdsourced information is often meaningful and recommended to users, e.g., reviews on Yelp or high marks on Dianping, which poses the threat of Sybil attacks. To address the problem of Sybil attacks, previous solutions mostly use indirect/direct graph model or clickstream model to detect fake accounts. However, they are either dependent on strong connections or solely preserved by servers of social networks. In this paper, we propose a novel predictable approach by exploiting users' custom patterns to distinguish Sybil attackers from normal users for the application of recommendation in mobile social networks. First, we introduce the entropy of spatial-temporal features to profile the mobility traces of normal users, which is quite different from Sybil attackers. Second, we develop discriminative entropy-based features, i.e., users' preference features, to measure the uncertainty of users' behaviors. Third, we design a smart Sybil detection model based on a binary classification approach by combining our entropy-based features with traditional behavior-based features. Finally, we examine our model and carry out extensive experiments on a real-world dataset from Dianping. Our results have demonstrated that the model can significantly improve the detection accuracy of Sybil attacks. Chen Lyu 0002, Qingyao Jia, Chihung Chi, Yang Xu 0012 |
WCNC | 6 |
| 2021 | Systematic evaluation of machine learning methods for identifying human-pathogen protein-protein interactionsabstractIn recent years, high-throughput experimental techniques have significantly enhanced the accuracy and coverage of protein-protein interaction identification, including human-pathogen protein-protein interactions (HP-PPIs). Despite this progress, experimental methods are, in general, expensive in terms of both time and labour costs, especially considering that there are enormous amounts of potential protein-interacting partners. Developing computational methods to predict interactions between human and bacteria pathogen has thus become critical and meaningful, in both facilitating the detection of interactions and mining incomplete interaction maps. In this paper, we present a systematic evaluation of machine learning-based computational methods for human-bacterium protein-protein interactions (HB-PPIs). We first reviewed a vast number of publicly available databases of HP-PPIs and then critically evaluate the availability of these databases. Benefitting from its well-structured nature, we subsequently preprocess the data and identified six bacterium pathogens that could be used to study bacterium subjects in which a human was the host. Additionally, we thoroughly reviewed the literature on 'host-pathogen interactions' whereby existing models were summarized that we used to jointly study the impact of different feature representation algorithms and evaluate the performance of existing machine learning computational models. Owing to the abundance of sequence information and the limited scale of other protein-related information, we adopted the primary protocol from the literature and dedicated our analysis to a comprehensive assessment of sequence information and machine learning models. A systematic evaluation of machine learning models and a wide range of feature representation algorithms based on sequence information are presented as a comparison survey towards the prediction performance evaluation of HB-PPIs. Huaming Chen, Fuyi Li, Lei Wang 0001, Yaochu Jin, Chihung Chi, Lukasz A. Kurgan, Jiangning Song, Jun Shen 0001 |
Briefings Bioinform. | 5 |
| 2021 | A polynomial-time algorithm for simple undirected graph isomorphismabstractSummary The graph isomorphism problem is to determine two finite graphs that are isomorphic which is not known with a polynomial‐time solution. This paper solves the simple undirected graph isomorphism problem with an algorithmic approach as NP=P and proposes a polynomial‐time solution to check if two simple undirected graphs are isomorphic or not. Three new representation methods of a graph as vertex/edge adjacency matrix and triple tuple are proposed. A duality of edge and vertex and a reflexivity between vertex adjacency matrix and edge adjacency matrix were first introduced to present the core idea. Beyond this, the mathematical approval is based on an equivalence between permutation and bijection. Because only addition and multiplication operations satisfy the commutative law, we propose a permutation theorem to check fast whether one of two sets of arrays is a permutation of another or not. The permutation theorem was mathematically approved by Integer Factorization Theory, Pythagorean Triples Theorem, and Fundamental Theorem of Arithmetic. For each of two n ‐ary arrays, the linear and squared sums of elements were respectively calculated to produce the results. Jing He 0004, Jinjun Chen, Guangyan Huang, Jie Cao 0001, Zhiwang Zhang, Hui Zheng 0001, Peng Zhang 0063, Roozbeh Zarei, Ferry Sansoto, Ruchuan Wang 0001, Yimu Ji 0001, Weibei Fan, Zhijun Xie, Xiancheng Wang, Mengjiao Guo, Chihung Chi, Paulo A. de Souza, Jiekui Zhang, Youtao Li, Xiaojun Chen 0001, Yong Shi 0001, David G. Green, Taraporewalla Kersi, André Van Zundert |
Concurr. Comput. Pract. Exp. | 16 |
| 2021 | An Accurate Negative Survey Using Answer Confidence LevelabstractNegative survey is an effective method to protect the privacy of survey participants and has many applications. Different from normal survey, it hides personal privacy by requiring an individual participant to provide a negative option as the answer to the survey question. Answers of participants can further be reconstructed into the distribution of positive options. However, the current negative survey that assumes all participants are 100 percent confident in their answers introduces some precision loss to the reconstruction process. In this paper, we provide an accurate negative survey by allowing participants to annotate their confidence levels to the answers. In particular, we develop a novel Negative Survey to Positive Survey with Answer Confidence Level (NStoPS-CL) algorithm to reconstruct the negative survey with answer confidence level and further increase the accuracy of reconstruction. Experiments demonstrate that NStoPS-CL reliably improves the reconstruction accuracy by testing answer confidence level under different conditions (i.e., dataset size, number of question options and dataset distributions), while balancing reconstruction accuracy and privacy well. Borui Cai, Guangyan Huang, Chihung Chi, Yanfeng Shu |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | Absorbing Diagonal Algorithm: An Eigensolver of $O\left(n^{2.584963}\log \frac{1}{\varepsilon }\right)$On2.584963log1ɛ Complexity at Accuracy $\varepsilon$ɛabstractEigenvalue decomposition is widely used in dimensionality reduction for knowledge engineering, in particular principal component analysis and other similar spectral methods. Traditional eigenvalue decomposition algorithms for decomposing a matrix of size n ×nn×n are usually of complexity O(n3)O(n3), due to a bottleneck in using Householder/Givens transforms to convert a general matrix to a tri-diagonal one. It is proposed in this article a new algorithm that takes only O(n2.584963log 1/ε) computational complexity to achieve accuracy ε of eigenvalue decomposition for any ε > 0ε>0. The basic idea of our algorithm is to convert a matrix into a diagonal form in multi-scale divide and conquer scheme, and the conversion is to iteratively and recursively apply two phases of operations called diagonal attractions and diagonal absorptions respectively. In a diagonal attraction, it attracts the off-diagonal entries to make the entries nearer to the diagonal larger in magnitude than those farther away from the diagonal. In a diagonal absorption, it absorbs the near-to-diagonal nonzero entries into the diagonal. In such a scheme, no Householder or Givens transforms are involved. Moreover, diagonal attractions and diagonal absorptions can be implemented with fast matrix multiplications. The scheme's divide and conquer pattern also allows our algorithm to be easily mapped to modern computer hardware. Our algorithm also complements well the family of randomized eigenvalue/SVD algorithms using sampling techinques, which are of complexity O(nαpolylog(1/ε)) with small α but very large overheads in the polylog. Their strength in the small exponent αα of nn in complexity was easily cancelled by the exploding overheads in the polylog. Now, with their low-accuracy estimate refined by our algorithm for high accuracy, their strength can be boosted significantly. Junfeng Wu 0010, Jing He 0004, Chihung Chi, Guangyan Huang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | A Fuzzy Theory Based Topological Distance Measurement for Undirected MultigraphsabstractThe topological distance is to measure the structural difference between two graphs in a metric space. Graphs are ubiquitous, and topological measurements over graphs arise in diverse areas, including, e.g. COVID-19 structural analysis, DNA/RNA alignment, discovering the Isomers, checking the code plagiarism. Unfortunately, popular distance scores used in these applications, that scale over large graphs, are not metrics, and the computation usually becomes NP-hard. While, fuzzy measurement is an uncertain representation to apply for a polynomial-time solution for undirected multigraph isomorphism. But the graph isomorphism problem is to determine two finite graphs that are isomorphic, which is not known with a polynomial-time solution. This paper solves the undirected multigraph isomorphism problem with an algorithmic approach as NP=P and proposes a polynomial-time solution to check if two undirected multigraphs are isomorphic or not. Based on the solution, we define a new fuzzy measurement based on graph isomorphism for topological distance/structural similarity between two graphs. Thus, this paper proposed a fuzzy measure of the topological distance between two undirected multigraphs. If two graphs are isomorphic, the topological distance is 0; if not, we will calculate the Euclidean distance among eight extracted features and provide the fuzzy distance. The fuzzy measurement executes more efficiently and accurately than the current methods. Jing He 0004, Jinjun Chen, Guangyan Huang, Mengjiao Guo, Zhiwang Zhang, Hui Zheng 0001, Yunyao Li 0002, Ruchuan Wang 0001, Weibei Fan, Chihung Chi, Weiping Ding 0001, Paulo A. de Souza, Run-Wei Li, André Van Zundert |
FUZZ-IEEE | 10 |
| 2020 | HIME: Mining and Ensembling Heterogeneous Information for Protein-Protein Interactions PredictionabstractResearch on protein-protein interactions (PPIs) data paves the way towards understanding the mechanisms of infectious diseases, however improving the prediction performance of PPIs of inter-species remains a challenge. Since one single type of sequence data such as amino acid composition may be deficient for high-quality prediction of protein interactions, we have investigated a broader range of heterogeneous information of sequences data. This paper proposes a novel framework for PPIs prediction based on Heterogeneous Information Mining and Ensembling (HIME) process to effectively learn from the interaction data. In particular, the proposed approach introduces an ensemble process together with substantial features that generate better performance of PPIs prediction task. The performance of the proposed framework is validated on real protein interaction datasets. The extensive experiments show that HIME achieves higher performance over all existing methods reported in literature so far. Huaming Chen, Yaochu Jin, Lei Wang 0001, Chihung Chi, Jun Shen 0001 |
IJCNN | 4 |
| 2020 | Clustering Hashtags Using Temporal Patterns
Borui Cai, Guangyan Huang, Shuiqiao Yang, Yong Xiang 0001, Chihung Chi |
WISE (1) | 5 |
| 2020 | APEX2S: A two-layer machine learning model for discovery of host-pathogen protein-protein interactions on cloud-based multiomics dataabstractSummary Presented by the avalanche of biological interactions data, computational biology is now facing greater challenges on big data analysis and solicits more studies to mine and integrate cloud‐based multiomics data, especially when the data are related to infectious diseases. Meanwhile, machine learning techniques have recently succeeded in different computational biology tasks. In this article, we have calibrated the focus for host‐pathogen protein‐protein interactions study, aiming to apply the machine learning techniques for learning the interactions data and making predictions. A comprehensive and practical workflow to harness different cloud‐based multiomics data is discussed. In particular, a novel two‐layer machine learning model, namely APEX2S, is proposed for discovery of the protein‐protein interactions data. The results show that our model can better learn and predict from the accumulated host‐pathogen protein‐protein interactions. Huaming Chen, Jun Shen 0001, Lei Wang 0001, Chihung Chi |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Top-N Hashtag Prediction via Coupling Social Influence and Homophily
Can Wang 0004, Yunwei Zhao, Chihung Chi, Willem-Jan van den Heuvel, Kwok-Yan Lam, Bela Stantic |
ADMA | 4 |
| 2019 | TOM: A Threat Operating Model for Early Warning of Cyber Security Threats
Can Wang 0004, Yunwei Zhao, Kwok-Yan Lam, Chihung Chi, Hui Tian 0001 |
ADMA | 6 |
| 2019 | A Methodology for Resolving Heterogeneity and Interdependence in Data Analytics
Yunwei Zhao, Can Wang 0004, Min Shu, Chihung Chi, Yonghong Yu |
ADMA | 6 |
| 2019 | Unfolding the Mixed and Intertwined: A Multilevel View of Topic Evolution on Twitter
Yunwei Zhao, Can Wang 0004, Willem-Jan van den Heuvel, Chihung Chi, Weimin Li 0001 |
ADMA | 5 |
| 2019 | Hyperparameter Estimation in SVM with GPU Acceleration for Prediction of Protein-Protein InteractionsabstractFor classification tasks, such as protein-protein interactions (PPI), support vector machines (SVMs) have been continually utilised as a standard machine learning model. However, most practices in PPIs classifications are limited to common circumstances with small datasets and low feature dimensions, due to the big computation burden of kernel functions and quadratic optimization of SVM. Alternatively, these practical experiences might tend to employ a linear model once the dataset becomes larger, which may have exclusively lost the kernel function's potential. Since there are different defined kernels and various groups of hyperparameter, the time costs in estimating a best set of hyperparameter by traditional grid search are subsequently tremendous for PPI classification. To address this challenge, in this paper, we present a more efficient solution of hyperparameter estimation by gaining acceleration with GPU, which trains SVM efficiently and accurately with kernel functions calculation accelerated on various PPI datasets. The experiments are firstly conducted on PPI classification task, and we have exclusively evaluated the effectiveness on five public classification datasets. Our solution demonstrates a faster and more accurate performance comparing with the state-of-the-art. Huaming Chen, Lei Wang 0001, Yaochu Jin, Chihung Chi, Fucun Li, Huaiyuan Chu, Jun Shen 0001 |
IEEE BigData | 4 |
| 2019 | Beyond the Power of Mere Repetition: Forms of Social Communication on Twitter through the Lens of Information Flows and Its Effect on Topic EvolutionabstractUnderstanding how people interact and exchange messages on social networks is significant for managing online contents and making predictions of future behaviors. Most existing research on the communication characteristics simply focuses on the user involvement. The current work largely neglects the content changes that imply how wide and deep the discussion in a topic goes, and to what degree people set forth their own views with the additional information supplemented. We are highly motivated to propose a theoretical framework to target those issues. In this paper, we define the communication modality constructs, and classify topics based on three dimensions: user involvement, information flow depth, and topic inter-relations, which substantially extend the traditional focus in user interaction analysis. The communication modality constructs comprise of (i) topic dialogicity, (ii) discussion intensiveness, and (iii) discussion extensibility. We introduce a quantitative model based on the topology of information flow graph, and use the information addition as well as the emotion attachment along the path to measure the pattern divergence between topic groups. Our model is empirically validated by using 78 million tweets, and experiments on Twitter demonstrate our contributions. Yunwei Zhao, Can Wang 0004, Chihung Chi, Willem-Jan van den Heuvel, Kwok-Yan Lam, Min Shu |
IJCNN | 3 |
| 2019 | A lightweight authentication scheme with privacy protection for smart grid communications
Liping Zhang 0003, Lanchao Zhao, Shuijun Yin, Chihung Chi |
Future Gener. Comput. Syst. | 4 |
| 2018 | Towards Biological Sequence Data Service with InsightsabstractTestable prediction outcomes generated by computational models based on available databases are the primary sources helping to design biological experiments. Although numerous databases have been designed by collecting data either only from literature manually or together with prediction outcomes from computational models, there is currently not a comprehensive data service framework delivering better insights for these results. In this paper, we introduce a biological sequence data service towards delivering deeper insights and helping better biological experiments design. The service includes following major components: a comprehensive database for storing biological data, data analytics tools for analysing biological data, and computational models for delivering testable prediction outcomes. Specifically, we present this service in a framework for studies on host-pathogen interactions. The design of this framework aims to improve the understanding of host-pathogen interactions. The relationships of hierarchical databases and their working mechanism, specifically between PPIs and DDIs, are also presented in this framework. Finally, the preliminary and practical experiences of building computational model for prediction is discussed. Huaming Chen, Jun Shen 0001, Lei Wang 0001, Chihung Chi |
IEEE BigData | 4 |
| 2018 | A Comparative Study of Transactional and Semantic Approaches for Predicting Cascades on TwitterabstractThe availability of massive social media data has enabled the prediction of people’s future behavioral trends at an unprecedented large scale. Information cascades study on Twitter has been an integral part of behavior analysis. A number of methods based on the transactional features (such as keyword frequency) and the semantic features (such as sentiment) have been proposed to predict the future cascading trends. However, an in-depth understanding of the pros and cons of semantic and transactional models is lacking. This paper conducts a comparative study of both approaches in predicting information diffusion with three mechanisms: retweet cascade, url cascade, and hashtag cascade. Experiments on Twitter data show that the semantic model outperforms the transactional model, if the exterior pattern is less directly observable (i.e. hashtag cascade). When it becomes more directly observable (i.e. retweet and url cascades), the semantic method yet delivers approximate accuracy (i.e. url cascade) or even worse accuracy (i.e. retweet cascade). Further, we demonstrate that the transactional and semantic models are not independent, and the performance gets greatly enhanced when combining both. Yunwei Zhao, Can Wang 0004, Chihung Chi, Kwok-Yan Lam, Sen Wang 0001 |
IJCAI | 3 |
| 2018 | A Multi-protocol Authentication Shibboleth Framework and Implementation for Identity Federation
Mengyi Li, Chihung Chi, Chen Ding 0004, Raymond K. Wong 0001, Zhong She |
SecureComm (2) | 2 |
| 2018 | Coupled Clustering Ensemble by Exploring Data InterdependenceabstractClustering ensembles combine multiple partitions of data into a single clustering solution. It is an effective technique for improving the quality of clustering results. Current clustering ensemble algorithms are usually built on the pairwise agreements between clusterings that focus on the similarity via consensus functions, between data objects that induce similarity measures from partitions and re-cluster objects, and between clusters that collapse groups of clusters into meta-clusters. In most of those models, there is a strong assumption on IIDness (i.e., independent and identical distribution), which states that base clusterings perform independently of one another and all objects are also independent. In the real world, however, objects are generally likely related to each other through features that are either explicit or even implicit. There is also latent but definite relationship among intermediate base clusterings because they are derived from the same set of data. All these demand a further investigation of clustering ensembles that explores the interdependence characteristics of data. To solve this problem, a new coupled clustering ensemble (CCE) framework that works on the interdependence nature of objects and intermediate base clusterings is proposed in this article. The main idea is to model the coupling relationship between objects by aggregating the similarity of base clusterings, and the interactive relationship among objects by addressing their neighborhood domains. Once these interdependence relationships are discovered, they will act as critical supplements to clustering ensembles. We verified our proposed framework by using three types of consensus function: clustering-based, object-based, and cluster-based. Substantial experiments on multiple synthetic and real-life benchmark datasets indicate thatCCEcan effectively capture the implicit interdependence relationships among base clusterings and among objects with higher clustering accuracy, stability, and robustness compared to 14 state-of-the-art techniques, supported by statistical analysis. In addition, we show that the final clustering quality is dependent on the data characteristics (e.g., quality and consistency) of base clusterings in terms of sensitivity analysis. Finally, the applications in document clustering, as well as on the datasets with much larger size and dimensionality, further demonstrate the effectiveness, efficiency, and scalability of our proposed models. Can Wang 0004, Chihung Chi, Zhong She, Longbing Cao, Bela Stantic |
ACM Trans. Knowl. Discov. Data | 2 |
| 2017 | Emerging Service Orchestration Discovery and MonitoringabstractDue to the popularity of web services on the Internet, it is important to have a clear view of their utilization behaviors. Despite asynchronous service invocations and distributed executions can provide better user experience, our views are blurred by out-of-order and fragmented service logs. Researchers have been trying various methods to reveal emerging service orchestration patterns, but nearly all of them have taken deterministic approaches. Hence, they do not natively cater for incomplete data and noises. In this paper, we propose to address these problems by using topic models aiming to reveal service orchestration patterns from sparse service logs. Probabilistic approaches do not only tolerate data defects, but their associated approximation methods also overcome combinatorial explosion. We first investigate the implications of sparsity on topic models. Secondly, we propose an extended time-series form of susceptible-infectious-recovered model to monitor the dynamics of emerging service orchestrations. We quantify their emerging-potential by estimated effective-reproduction-number, which is obtained incrementally by Bayesian parameter estimations. Guided by our proposed emerging-potential measure, one can profile and categorize emerging service orchestration patterns, and generate automated alerts on upcoming consumption peaks. In practice, our model enables service providers to better allocate their resources to meet demands dynamically. While our findings affirm that biterm topic model can be applied to service logs with short and sparse log entries, the effectiveness of our proposed monitoring solutions is also shown by experiments. Victor W. Chu, Raymond K. Wong 0001, Simon Fong 0001, Chihung Chi |
IEEE Trans. Serv. Comput. | 4 |
| 2017 | Efficient Role Mining for Context-Aware Service Recommendation Using a High-Performance ClusterabstractService recommendation systems have been trying to utilize context-aware information to recommend services that better meet the needs of the service consumers. However, current context-aware service recommendation techniques are mainly based on individual intelligence or the local knowledge of users, and do not take into consideration the common knowledge among different users. To address this, recent research has attempted to use role-based approaches to recommend services to other members within the same context group. However, these proposed algorithms are inefficient and may not scale to cope with the large amount of mobile traffic in the real-world. This paper proposes novel algorithms with better runtime complexity, and further extends them to a MapReduce style to take advantage of popular distributed computing platforms. Experiments running on a medium-sized high performance computing cluster demonstrate that our proposed algorithms outperform previous work in runtime complexity and scalability. Raymond K. Wong 0001, Chihung Chi |
IEEE Trans. Serv. Comput. | 3 |
| 2016 | A Testbed for Collecting QoS Data of Cloud-Based Analytic ServicesabstractQoS-based service selection has been an active research area in both Service Computing and Cloud Computing communities in recent years. One of the obstacles for many researchers is lack of benchmark QoS datasets, especially for cloud service selection. In this paper, we propose a testbed system which can be used to collect the QoS data for software services hosted in the cloud. We mainly focus on analytic services, whose QoS values could be dependent on the data they are used to process and analyze. Using the testbed, we can publish services (software or infrastructure services), we can invoke software services hosted on certain infrastructure services, and system can monitor and record the QoS values of all the invocations. We have implemented a proof-of-concept prototype system and used it to collect a sample QoS dataset. Md Shahinur Rahman, Chen Ding 0004, Xumin Liu, Chihung Chi |
CLOUD | 4 |
| 2016 | Interrelationships of Service Orchestrations
Victor W. Chu, Raymond K. Wong 0001, Fang Chen 0001, Chihung Chi |
ADMA | 4 |
| 2016 | Identity in the Internet-of-Things (IoT): New Challenges and Opportunities
Kwok-Yan Lam, Chihung Chi |
ICICS | 2 |
| 2016 | Special Issue on Trajectory-based Behaviour Analytics
Chihung Chi, Can Wang 0004 |
J. Comput. Syst. Sci. | 1 |
| 2015 | Coupled Interdependent Attribute Analysis on Mixed DataabstractIn the real-world applications, heterogeneous interdependent attributes that consist of both discrete and numerical variables can be observed ubiquitously. The usual representation of these data sets is an information table, assuming the independence of attributes. However, very often, they are actually interdependent on one another, either explicitly or implicitly. Limited research has been conducted in analyzing such attribute interactions, which causes the analysis results to be more local than global. This paper proposes the coupled heterogeneous attribute analysis to capture the interdependence among mixed data by addressing coupling context and coupling weights in unsupervised learning. Such global couplings integrate the interactions within discrete attributes, within numerical attributes and across them to form the coupled representation for mixed type objects based on dimension conversion and feature selection. This work makes one step forward towards explicitly modeling the interdependence of heterogeneous attributes among mixed data, verified by the applications in data structure analysis, data clustering evaluation, and density comparison. Substantial experiments on 12 UCI data sets show that our approach can effectively capture the global couplings of heterogeneous attributes and outperforms the state-of-the-art methods, supported by statistical analysis. Can Wang 0004, Chihung Chi, Wei Zhou 0011, Raymond K. Wong 0001 |
AAAI | 2 |
| 2015 | Imperfect referees: Reducing the impact of multiple biases in peer reviewabstractBias in peer review entails systematic prejudice that prevents accurate and objective assessment of scientific studies. The disparity between referees' opinions on the same paper typically makes it difficult to judge the paper's quality. This article presents a comprehensive study of peer review biases with regard to 2 aspects of referees: the static profiles (factual authority and self‐reported confidence) and the dynamic behavioral context (the temporal ordering of reviews by a single reviewer), exploiting anonymized, real‐world review reports of 2 different international conferences in information systems / computer science. Our work extends conventional bias research by considering multiple biases occurring simultaneously. Our findings show that the referees' static profiles are more dominant in peer review bias when compared to their dynamic behavioral context. Of the static profiles, self‐reported confidence improved both conference fitness and impact‐based bias reductions, while factual authority could only contribute to conference fitness‐based bias reduction. Our results also clearly show that the reliability of referees' judgments varies along their static profiles and is contingent on the temporal interval between 2 consecutive reviews. Yun Wei Zhao, Chihung Chi, Willem-Jan van den Heuvel |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2015 | Coupled Attribute Similarity Learning on Categorical DataabstractAttribute independence has been taken as a major assumption in the limited research that has been conducted on similarity analysis for categorical data, especially unsupervised learning. However, in real-world data sources, attributes are more or less associated with each other in terms of certain coupling relationships. Accordingly, recent works on attribute dependency aggregation have introduced the co-occurrence of attribute values to explore attribute coupling, but they only present a local picture in analyzing categorical data similarity. This is inadequate for deep analysis, and the computational complexity grows exponentially when the data scale increases. This paper proposes an efficient data-driven similarity learning approach that generates a coupled attribute similarity measure for nominal objects with attribute couplings to capture a global picture of attribute similarity. It involves the frequency-based intra-coupled similarity within an attribute and the inter-coupled similarity upon value co-occurrences between attributes, as well as their integration on the object level. In particular, four measures are designed for the inter-coupled similarity to calculate the similarity between two categorical values by considering their relationships with other attributes in terms of power set, universal set, joint set, and intersection set. The theoretical analysis reveals the equivalent accuracy and superior efficiency of the measure based on the intersection set, particularly for large-scale data sets. Intensive experiments of data structure and clustering algorithms incorporating the coupled dissimilarity metric achieve a significant performance improvement on state-of-the-art measures and algorithms on 13 UCI data sets, which is confirmed by the statistical analysis. The experiment results show that the proposed coupled attribute similarity is generic, and can effectively and efficiently capture the intrinsic and global interactions within and between attributes for especially large-scale categorical data sets. In addition, two new coupled categorical clustering algorithms, i.e., CROCK and CLIMBO are proposed, and they both outperform the original ones in terms of clustering quality on UCI data sets and bibliographic data. Can Wang 0004, Xiangjun Dong 0001, Fei Zhou 0001, Longbing Cao, Chihung Chi |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Personalized Decision-Strategy based Web Service Selection using a Learning-to-Rank AlgorithmabstractIn order to choose from a list of functionally similar services, users often need to make their decisions based on multiple QoS criteria they require on the target service. In this process, different users may follow different decision making strategies, some are compensatory in which only an overall value on all the criteria is evaluated, some evaluate one criterion at a time in the order of their importance levels, while others count on the number of winning criteria. Most of the current QoS-based service selection systems do not consider these decision strategies in the ranking process, which we believe are crucial for generating accurate ranking results for individual users. In this paper, we propose a decision strategy based service ranking model. Furthermore, considering that different users follow different strategies in different contexts at different times, we apply a machine learning algorithm to learn a personalized ranking model for individual users based on how they select services in the past. We have implemented and tested the proposed approach, and our experiment results show the effectiveness of the approach. Muhammad Suleman Saleem, Chen Ding 0004, Xumin Liu, Chihung Chi |
IEEE Trans. Serv. Comput. | 4 |
| 2015 | Formalization and Verification of Group Behavior InteractionsabstractGroup behavior interactions, such as multirobot teamwork and group communications in social networks, are widely seen in both natural, social, and artificial behavior-related applications. Behavior interactions in a group are often associated with varying coupling relationships, for instance, conjunction or disjunction. Such coupling relationships challenge existing behavior representation methods, because they involve multiple behaviors from different actors, constraints on the interactions, and behavior evolution. In addition, the quality of behavior interactions are not checked through verification techniques. In this paper, we propose an ontology-based behavior modeling and checking system (OntoB for short) to explicitly represent and verify complex behavior relationships, aggregations, and constraints. The OntoB system provides both a visual behavior model and an abstract behavior tuple to capture behavioral elements, as well as building blocks. It formalizes various intra-coupled interactions (behaviors conducted by the same actor) via transition systems (TSs), and inter-coupled behavior aggregations (behaviors conducted by different actors) from temporal, inferential, and party-based perspectives. OntoB converts a behavior-oriented application into a TS and temporal logic formulas for further verification and refinement. We demonstrate and evaluate the effectiveness of the OntoB in modeling multirobot behaviors and their interactions in the Robocup soccer competition game. We show, that the OntoB system can effectively model complex behavior interactions, verify and refine the modeling of complex group behavior interactions in a sound manner. Can Wang 0004, Longbing Cao, Chihung Chi |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2014 | Microblog Topic Contagiousness Measurement and Emerging Outbreak MonitoringabstractA recent study on collective attention in Twitter shows that an epidemic spreading of hashtags is predominantly driven by external factors. We extend a time-series form of susceptible-infectious-recovered (SIR) model to monitor microblog emerging outbreaks by considering both endogenous and exogenous drivers. In addition, we adopt partially labeled Dirichlet allocation (PLDA) model to generate both background latent topics and hashtag topics. It overcomes the problem of small available samples in hashtag analysis by including related but unlabeled tweets through inference. We standardize hashtag topic contagiousness measure as the estimated effective-reproduction-number R derived from epidemiology. It is obtained by Bayesian parameter estimation. Guided by R, one can profile and categorize emerging topics, and generate alerts on potential outbreaks. Experiment results confirm the effectiveness of this approach. Victor W. Chu, Raymond K. Wong 0001, Fang Chen 0001, Chihung Chi |
CIKM | 4 |
| 2014 | Web Service Orchestration Topic MiningabstractDue to the popularity of using web services to deliver services on the Web, a clear view of how they are being consumed is becoming critical. Researchers have been trying multiple methods to reveal actual service orchestration patterns from service logs. However, most of the discovery methods have taken deterministic approaches, and hence, they do not provide enough allowance to cater for incomplete data and noises. On the other hand, most investigations do not take combinatorial explosion into consideration leading to scalability problem. Moreover, asynchronous web service invocations and distributed executions also make it difficult to identify service patterns due to the randomness in log record generation. In this paper, probabilistic topic mining class of solutions are applied to reveal web service orchestration patterns from service logs, in which robust approximation methods are available to provide scalability. Data sparsity problem in service log is also investigated by using biterm topic model (BTM) and comparing its results with traditional latent Dirichlet allocation (LDA) model. In addition, a topic matching method is introduced based on the Hungarian method on Jensen-Shannon divergence matrix, whilst notions of aggJSD and autoJSD are also introduced to measure topic diversity between matched topic sets and within a single topic set respectively. Experiment results confirm that BTM can be used for service logs with short log entries and with sparsity larger than 90% approximately. Victor W. Chu, Raymond K. Wong 0001, Chihung Chi, Patrick C. K. Hung |
ICWS | 3 |
| 2014 | Personalized Decision Making for QoS-Based Service SelectionabstractIn order to choose from a list of functionally similar services, users often need to make their decisions based on multiple QoS criteria they require on the target service. In this process, different users may follow different decision making strategies, some are compensatory in which only an overall value on all the criteria is evaluated, some evaluate one criterion at a time in the order of their importance levels, while others count on the number of winning criteria. Most of the current QoS-based service selection systems do not consider these decision strategies in the ranking process, which we believe are crucial for generating accurate ranking results for individual users. In this paper, we propose a decision strategy based service ranking model. Furthermore, considering that different users follow different strategies in different contexts at different times, we apply a machine learning algorithm to learn a personalized ranking model for individual users based on how they select services in the past. Our experiment result shows the effectiveness of the proposed approach. Muhammad Suleman Saleem, Chen Ding 0004, Xumin Liu, Chihung Chi |
ICWS | 4 |
| 2013 | Enhancing tag-based collaborative filtering via integrated social networking informationabstractRecently, researchers have taken tremendous strides in attempting to synthesize conventional social judgments and automated filtering within recommender systems. In this study, we aim to enhance recommendation efficiency via integrating social networking information with traditional recommendation algorithms. To achieve this objective, we first propose a new user similarity metric that not only considers tagging activities of users, but also incorporates their social relationships, such as friendship and membership, in measuring the closeness of two users. Subsequently, we define a new item prediction method which makes use of both user-to-user similarity and item-to-item similarity. Experimental outcomes on Last.fm show some positive results that attest the efficiency of our proposed approach. Sogol Naseri, Arash Bahrehmand, Chen Ding 0004, Chihung Chi |
ASONAM | 4 |
| 2013 | Scalable context-aware role mining with MapReduceabstractCloud computing platforms facilitate efficiently processing complicated computing problems of which the time cost used to be unacceptable. Recent research has attempted to use role-based approaches for context-aware service recommendation, yet role mining problem has been proven to be difficult to compute. Currently proposed role-mining algorithms are inefficient and may not scale to cope with the huge amount of data in the real-world. This paper proposes a novel algorithm with much better runtime complexity, and in MapReduce style to take advantage of popular distributed computing platforms. Experiments running on a medium-sized high performance computing cluster demonstrate that our proposed algorithm works well with both running time complexity and scalability. Raymond K. Wong 0001, Chihung Chi |
IEEE BigData | 3 |
| 2013 | Online Role Mining without Over-Fitting for Service RecommendationabstractDue to the popularity of smartphones, finding and recommending suitable services on mobile devices are increasingly important. Recent research has attempted to use role-based approaches to recommend mobile services to other members among the same group in a context dependent manner. However, the traditional role mining approaches originated from the domain of security control tend to be rigid and may not be able to capture human behaviors adequately. In particular, during the course of role mining process, these approaches easily result in over-fitting, i.e., too many roles with slightly different service consumption patterns are found. As a result, they fail to reveal the true common preferences within the user community. This paper proposes an online role mining algorithm with a residual term that automatically group users according to their interests and habits without losing sight of their individual preferences. Moreover, to resolve the over-fitting problem, we relax the role mining mechanism by introducing quasi-roles based on the concept of quasi-bicliques. Most importantly, the new concept allows us to propose a monitoring framework to detect and correct over-fitting in online role mining such that recommendations can be made based on the latest and genuine common preferences. To the best of our knowledge, this is a new area in service recommendation that is yet to be fully explored. Victor W. Chu, Raymond K. Wong 0001, Chihung Chi |
ICWS | 3 |
| 2013 | Non-functional Requirement Analysis and Recommendation for Software ServicesabstractNon-Functional (NF) requirement is very important for the success of a software service. Considering that there could be multiple services implementing a same function, it is crucial for software providers to understand the real NF demands from consumers so that they can meet these demands and attract users. It is also crucial for consumers to know what is being offered so that they can pose realistic NF requests. We address both issues here by proposing a NF requirement analysis and recommendation system which works for both providers and consumers. NF requirements from various sources are first collected, and then we apply the factor analysis technique to identify those independent latent factors which contribute to those observable NF values. Finally we use cluster analysis to summarize the popular NF demands. Our experiment result shows the effectiveness of this approach. Chihung Chi, Chen Ding 0004, Raymond K. Wong 0001 |
ICWS | 2 |
| 2012 | Scalable Join Queries in Cloud Data StoresabstractCloud data stores provide scalability and high availability properties for Web applications, but do not support complex queries such as joins. Web application developers must therefore design their programs according to the peculiarities of No SQL data stores rather than established software engineering practice. This results in complex and error-prone code, especially with respect to subtle issues such as data consistency under concurrent read/write queries. We present join query support in Cloud TPS, a middleware layer which stands between a Web application and its data store. The system enforces strong data consistency and scales linearly under a demanding workload composed of join queries and read-write transactions. In large-scale deployments, Cloud TPS outperforms replicated Postgre SQL up to three times. Guillaume Pierre, Chihung Chi |
CCGRID | 3 |
| 2012 | Efficient QoS-based Service Selection with Consideration of User RequirementsabstractMany QoS-based service selection algorithms are complex and time-consuming, and sometimes require a lot of manual efforts from users. As a result, users' real-time searching experiences may not be good when there are many candidate services, because it could be very slow to calculate the ranking scores of services. In this paper, we want to improve the efficiency of the selection process by using a simple vector space model. We also consider the actual QoS requirements from users in the ranking process, which is missing in many current systems. Our experiment results show a big improvement on the system efficiency without losing much on the accuracy when compared with a well-known algorithm -- skyline computation. Anita Mohebi, Chen Ding 0004, Chihung Chi |
EDOC | 3 |
| 2012 | A State Synchronization Mechanism for Orchestrated ProcessesabstractTwo orchestrated processes interacting with each other have to maintain their own states. Messages are used to synchronize states between orchestrated processes. Server crash and network failure may result in loss of messages and therefore result in a state change performed by only one party. Thus, the states of the parties are no longer synchronized, resulting in state inconsistencies and in worst case deadlocks. In this paper, we propose a mechanism for guaranteed state synchronization of orchestrated processes with system and network failures. Our mechanism is based on interaction patterns and process transformations. The basic idea is to redesign the original processes into their state synchronization-enabled counterparts via process transformations that can be automated. The transformation mechanism is formalized based on Colored Petri Nets. We present the formal proof of the correctness of our mechanism and give the overhead analysis to illustrate its practicability. Lei Wang 0111, Andreas Wombacher, Luís Ferreira Pires, Marten van Sinderen, Chihung Chi |
EDOC | 5 |
| 2012 | Modeling User's Non-functional Preferences for Personalized Service Ranking
Rozita Mirmotalebi, Chen Ding 0004, Chihung Chi |
ICSOC | 3 |
| 2012 | CloudTPS: Scalable Transactions for Web Applications in the CloudabstractNoSQL cloud data stores provide scalability and high availability properties for web applications, but at the same time they sacrifice data consistency. However, many applications cannot afford any data inconsistency. CloudTPS is a scalable transaction manager which guarantees full ACID properties for multi-item transactions issued by web applications, even in the presence of server failures and network partitions. We implement this approach on top of the two main families of scalable data layers: Bigtable and SimpleDB. Performance evaluation on top of HBase (an open-source version of Bigtable) in our local cluster and Amazon SimpleDB in the Amazon cloud shows that our system scales linearly at least up to 40 nodes in our local cluster and 80 nodes in the Amazon cloud. Guillaume Pierre, Chihung Chi |
IEEE Trans. Serv. Comput. | 3 |
| 2011 | A mathematical value system model for agent in online social networkabstractAgent-based modeling is a powerful method for studying user's behaviors in online social networks. Agent uses its value system to estimate the value of its behavior. A good value system model for agent can make the agent more intelligence and hence facilitate the development of intelligence system. However, the issue of modeling agent's value system is not properly studied yet. In this paper, we propose a mathematical value system model for estimating the value of agent's behavior in online social network. We use three evaluation dimensions as the universal equivalents for behaviors, and propose a weight function mechanism to estimate each dimension's weight. We show the potential usage of our model by applying it to the case of online word-of-mouth advertising. Donghong Huang, Chihung Chi |
HIS | 2 |
| 2011 | Collaborative Filtering Based Service Ranking Using Invocation HistoriesabstractCollaborative filtering based recommender systems are very successful on dealing with the information overload problem and providing personalized recommendations to users. When more and more web services are published online, this technique can also help recommend and select services which satisfy users' particular Quality of Service (QoS) requirements and preferences. In this paper, we propose a novel collaborative filtering based service ranking mechanism, in which the invocation and query histories are used to infer the user behavior, and user similarity is calculated based on similar invocations and queries. To overcome some of the inherent problems with the collaborative filtering systems such as the cold start and data sparsity problem, the final ranking score is a combination of the QoS-based matching score and the collaborative filtering based score. The experiment using a simulated dataset proves the effectiveness of the algorithm. Chen Ding 0004, Chihung Chi |
ICWS | 3 |
| 2011 | Ontology-Based Business Process Customization for Composite Web ServicesabstractA key goal of the Semantic Web is to shift social interaction patterns from a producer-centric paradigm to a consumer-centric one. Treating customers as the most valuable assets and making the business models work better for them are at the core of building successful consumer-centric business models. It follows that customizing business processes constitutes a major concern in the realm of a knowledge-pull-based human semantic Web. This paper conceptualizes the customization of service-based business processes leveraging the existing knowledge of Web services and business processes. We represent this conceptualization as a new Extensible Markup Language (XML) markup language Web Ontology Language-Business Process Customization (OWL-BPC), based on the de facto semantic markup language for Web-based information [Web Ontology Language (OWL)]. Furthermore, we report a framework, built on OWL-BPC, for customizing service-based business processes, which supports customization detection and enactment. Customization detection is enabled by a business-goal analysis, and customization enactment is enabled via event-condition-action rule inference. Our solution and framework have the following capabilities in dealing with inconsistencies and misalignments in business process interactions: 1) resolve semantic mismatch of process parameters; 2) handle behavioral mismatches which may or may not be compatible; and 3) process misaligned rendezvous requirements. Such capabilities are applicable to business processes with heterogeneous domain ontology. We present an architectural description of the implementation and a walk-through of an example of solving a customization problem as a validation of the proposed approach. Qianhui Althea Liang, Xindong Wu 0001, E. K. Park, Taghi M. Khoshgoftaar, Chihung Chi |
IEEE Trans. Syst. Man Cybern. Part A | 5 |
| 2010 | Non-I.I.D. Multi-Instance Dimensionality Reduction by Learning a Maximum Bag Margin SubspaceabstractMulti-instance learning, as other machine learning tasks, also suffers from the curse of dimensionality. Although dimensionality reduction methods have been investigated for many years, multi-instance dimensionality reduction methods remain untouched. On the other hand, most algorithms in multi- instance framework treat instances in each bag as independently and identically distributed samples, which fails to utilize the structure information conveyed by instances in a bag. In this paper, we propose a multi-instance dimensionality reduction method, which treats instances in each bag as non-i.i.d. samples. We regard every bag as a whole entity and define a bag margin objective function. By maximizing the margin of positive and negative bags, we learn a subspace to obtain more salient representation of original data. Experiments demonstrate the effectiveness of the proposed method. Wei Ping, Kexin Ren, Chihung Chi, Furao Shen |
AAAI | 4 |
| 2010 | Secure Data Aggregation through Proactive DefenseabstractGossip based aggregation protocols are a promising approach to monitoring large-scale decentralized IT infrastructures. Compared to traditional approaches they exhibit good properties of scalability, tolerance of churn, and communication overhead. Gossip-based protocols can compute statistical aggregates such as the average, sum or statistical distribution of an attribute across a large system. However, such protocols are extremely vulnerable to malicious attacks, and even a small number of attackers in the system can largely undermine aggregation results. This paper presents a secure protocol for computing attribute averages. In this system, each node autonomously judges whether its neighbors are malicious, and may subsequently stop any interaction with them. A node appearing malicious to its neighbors quickly gets excluded from the system. Instead of defining malicious behavior (and excluding nodes that follow the definition of maliciousness), our system defines correct behavior (and excludes any node that behaves differently). This allows in principle our system to address arbitrary types of attacks. Simulations based on real-world attribute data demonstrate that our system offers good resistance against four different types of attacks. Guillaume Pierre, Chihung Chi |
ICCCN | 3 |
| 2010 | High Availability Data Model for P2P Storage Network
Bangyu Wu, Chihung Chi, Zhiheng Xie, Chen Ding 0004 |
WISE | 2 |
| 2010 | Autonomous resource provisioning for multi-service web applicationsabstractDynamic resource provisioning aims at maintaining the end-to-end response time of a web application within a pre-defined SLA. Although the topic has been well studied for monolithic applications, provisioning resources for applications composed of multiple services remains a challenge. When the SLA is violated, one must decide which service(s) should be reprovisioned for optimal effect. We propose to assign an SLA only to the front-end service. Other services are not given any particular response time objectives. Services are autonomously responsible for their own provisioning operations and collaboratively negotiate performance objectives with each other to decide the provisioning service(s). We demonstrate through extensive experiments that our system can add/remove/shift both servers and caches within an entire multi-service application under varying workloads to meet the SLA target and improve resource utilization. Dejun Jiang 0003, Guillaume Pierre, Chihung Chi |
WWW | 3 |
| 2009 | Compatibility-driven and adaptable service compositionabstractServices participating in the composition are usually co-ordinated according to a workflow, composed by several activities, each of which carried out by a service. The binding of services to workflow activities may be affected by several parameters (e.g., QoS, price, reputation, etc.). In this paper, we propose a service binding driven by a further important requirement, that is, the incompatibilities among services participating into the composition. To achieve a compatibility-driven composition we propose a solution where services assignment is configured directly by the engine coordinating the composite service. Moreover, the composition is generated such to implement a failure recovery strategy, that is, in case of some service failure the engine dynamically replaces the unavailable service. Barbara Carminati, Chihung Chi, Elena Ferrari 0001, Lianghuan Yu |
APSCC | 2 |
| 2009 | Scalable Transactions for Web Applications in the Cloud
Guillaume Pierre, Chihung Chi |
Euro-Par | 3 |
| 2009 | Selection Strategy of Rescue Servers Under Hot-Spot CongestionabstractIn this paper, we investigate selection strategies for rescue servers in a fully collaborative in which heterogeneous systems share spare resource to address hotspot problems experienced by individual members. Different from content replication, delays due to application replication and startup time of the resulting rescue servers will be taken into consideration. Details about our replication strategy, such as the maximum set size margin on the total number of discarded requests, overall latency for enabling rescue servers, are given. Simulation result shows that our selection strategies perform much better than existing mechanisms in which replication and startup delay are ignored. Chihung Chi, Chen Ding 0004 |
ICIW | 2 |
| 2009 | Workflow-based resource allocation to optimize overall performance of composite services
Bangyu Wu, Chihung Chi, Ming Gu 0001, Jia-Guang Sun 0001 |
Future Gener. Comput. Syst. | 2 |
| 2009 | QoS Requirement Generation and Algorithm Selection for Composite Service Based on Reference Vector
Bangyu Wu, Chihung Chi, Ming Gu 0001, Jia-Guang Sun 0001 |
J. Comput. Sci. Technol. | 2 |
| 2008 | Quantifying Trust Based on Service Level Agreement for Software as a ServiceabstractIn this paper, we aim to construct quantifying trust model based on service level agreements (SLAs) as well as feedbacks from service consumers. The proposed model can aggregate the feedbacks weighted by both trust factor and similarity factor to give quantitative prediction of the target services' trustworthiness. By comparing with traditional prediction algorithm, we show that our methodology gives better performance in trust estimation. Chihung Chi, Jianming Deng |
COMPSAC | 2 |
| 2008 | A Novel Ownership Scheme to Maintain Web Content Consistency
Chihung Chi, Choon-Keng Chua, Weihong Song |
GPC | 1 |
| 2008 | Adaptive Quality Recommendation Mechanism for Software Service ProvisioningabstractIn this paper, we propose an adaptive quality recommendation mechanism to help software service providers understand the dynamism of quality demand from the majority of requesters accurately. The unique feature of our approach is that based on the intra-cluster proximity index (called the icp-index) that we propose, the granularity of service clustering can be adjusted dynamically to meet the wide variation of service providers who want to target their services to different groups of clients. Experiments show that our approach is more accurate and flexible to identify the need of service quality of requesters than existing solutions such as simple averaging, minimum-maximum-mean, or traditional interval-range data clustering. SiMing Li, Chen Ding 0004, Chihung Chi, Jianming Deng |
ICWS | 3 |
| 2008 | A Clustering-Based Approach for Assisting Semantic Web Service RetrievalabstractThe discovery of suitable web services for a given task from a brokerage system is the core part of current Service Oriented Architecture (SOA). With the number of registered web services growing, organization of search results is critical for improving the utility of any service search engine. A clustering view of search results is more effective than traditional ranked-list style in helping users to navigate into relevant services quickly and accurately. In this paper, we will propose an efficient clustering algorithm for organizing returned services. Our experiments validated the efficiency of the proposed method. Dong Shou, Chihung Chi |
ICWS | 2 |
| 2008 | Service-oriented data denormalization for scalable web applicationsabstractMany techniques have been proposed to scale web applications. However, the data interdependencies between the database queries and transactions issued by the applications limit their efficiency. We claim that major scalability improvements can be gained by restructuring the web application data into multiple independent data services with exclusive access to their private data store. While this restructuring does not provide performance gains by itself, the implied simplification of each database workload allows a much more efficient use of classical techniques. We illustrate the data denormalization process on three benchmark applications: TPC-W, RUBiS and RUBBoS. We deploy the resulting service-oriented implementation of TPC-W across an 85-node cluster and show that restructuring its data can provide at least an order of magnitude improvement in the maximum sustainable throughput compared to master-slave database replication, while preserving strong consistency and transactional properties. Dejun Jiang 0003, Guillaume Pierre, Chihung Chi, Maarten van Steen |
WWW | 4 |
| 2007 | On-Demand Capacity Framework
Chihung Chi |
ICA3PP | 1 |
| 2006 | A Risk Assessment Model for Enterprise Network Security
Fu-Hong Yang, Chihung Chi, Lin Liu 0001 |
ATC | 2 |
| 2006 | Quantitative Analysis on the Cacheability Factors of Web ObjectsabstractIn this paper, we would like to quantify factors that contribute to the cacheability decision of web objects. Our study is based on a comprehensive study of object attributes that are related to the HTTP protocol transfer, proxy preference, content availability, freshness revalidation as well as the inter-relationship and dependency among them. Furthermore, the relative importance of these factors towards the final cacheability decision of web objects are also investigated. This study is important because it provides hints on how to optimize web caching performance. Chihung Chi, Lin Liu 0001, LuWei Zhang |
COMPSAC (1) | 1 |
| 2006 | Proxy-Based Pervasive Multimedia Content DeliveryabstractAs the Web client devices become more and more heterogeneous, the Web content providers should cater for multi-presentation services. Transcoding is seen as an alternative to deliver such services. When deployed offline, transcoding is space-consuming, inflexible, and difficult to maintain. On-demand transcoding, on the other hand, is CPU-consuming and time-delaying. Inspired by new multimedia standards - like JPEG 2000 and MPEG-4 -, a new adaptation called modulation is devised. Unlike transcoding, modulation does not involve complex computations and can offer high data reuse, which is beneficial to Internet data access. The specifications of modulation are presented, and so is its comparison with transcoding Henry Novianus Palit, Chihung Chi, Lin Liu 0001 |
COMPSAC (1) | 2 |
| 2006 | Data Integrity Related Markup Language and HTTP Protocol Support for Web Intermediaries
Chihung Chi, Lin Liu 0001, Xiaoyin Yu |
EUC | 1 |
| 2006 | Task Delegation in Active Web Intermediary NetworkabstractIn this paper, we formalize the concept of a delegation network for Web intermediaries and present formal semantics for the responsibility of these actors. Key properties of the network are proven and method to judge an actor's responsibility is given. This work is important because it determines the accuracy of task execution and the feasibility of content reuse in the network Chihung Chi, Lin Liu 0001, Xiaoyin Yu |
ICWS | 1 |
| 2006 | Quantitative Trust Based on ActionsabstractIn this paper, we propose an automatic, self-adaptive trust model for component-based service-oriented architecture that is based on Bayesian model. The focus of this model is the "quantitative trust on action". Both the computation model and its updating mechanism are given. Through our implementation of this model using Monte-Carlo algorithm, we show the feasibility and practicability of our proposal as well as its ability to rectify any unfair advertisement and to track the capability of highly dynamic service Chihung Chi |
ICWS | 2 |
| 2006 | Truthful application-layer multicast in mesh-based selfish overlaysabstractBuilding a multicast tree in an overlay network is the major problem in application layer multicast. In a selfish overlay network, the end hosts may not be cooperative. They are willing to maximize their own benefits instead of being loyal to the protocol or altruistically contributing to the benefits of the whole overlay network. In this paper, we design a strategyproof mechanism in building a truthful minimum cost multicast tree in the mesh-tree way, which could simplify the overlay construction and maintenance. We model the network in both simple and general cost model. We conduct simulation to study the relationship between utility of cheating node and other nodes Wei Zhou 0011, Chihung Chi |
IPCCC | 4 |
| 2006 | AN.P2P - an Active Peer-to-peer System
Chihung Chi, Mu Su, Lin Liu 0001, Hongguang Wang |
SEKE | 1 |
| 2006 | Web Object Cacheability How Much Do We Know?
Chihung Chi, Jun-Li Yuan, Lin Liu 0001 |
SEKE | 1 |
| 2006 | Object Placement and Caching Strategies on AN.P2P
Mu Su, Chihung Chi, Lin Liu 0001, Hongguang Wang |
WAIM | 2 |
| 2005 | Architecture and Performance of Application Networking in Pervasive Content DeliveryabstractThis paper proposes Application Networking (App.Net) architecture that enables Web server to deploy intermediate response with associated service logic to edge proxies in form of a workflow. The recipient proxy is allowed to instantiate the workflow using local services or by downloading mobile applications from remote sites. The final response is the output from the workflow execution fed with the intermediate result. In this paper, we defined a workflow modulation method to represent service logic and manipulate it in uniform operations. An App.Net caching method is designed to cache intermediate results as well as final response presentations. According to cost model measuring the bandwidth usage, we designed workflow placement algorithms to deploy intermediate response objects to the proxy in optimal or greedy ways. Finally, we developed an App.Net prototype and implemented a wide range of applications existed in the Web. Our simulation results show that using workflow end dynamic application deployment, App.Net can achieve better performance than conventional methods. Mu Su, Chihung Chi |
ICDE | 2 |
| 2005 | Modes of Real-Time Content Transformation for Web Intermediaries in Active NetworkabstractActive content transformation in web is a hot topic in Internet content delivery research. In this paper, we proposed 3 different modes content transformation, which are whole-file buffering, byte-streaming, and chunk buffering. Based on the chunk dependence graph, we studied the performance for different modes and argued that the chunk streaming is the most appropriate transform model for Web intermediaries based real-time content transformation Haiguang Xie, Chihung Chi |
MMSP | 4 |
| 2005 | Quantitative Modeling for Web Objects' Cacheability
Chen Ding 0004, Chihung Chi, Lin Liu 0001, LuWei Zhang, Hongguang Wang |
WAIM | 2 |
| 2005 | A Framework of HTML Content Selection and Adaptive Delivery
Chen Ding 0004, Chihung Chi |
WAIM | 3 |
| 2005 | Modulation for Scalable Multimedia Content Delivery
Henry Novianus Palit, Chihung Chi |
WAIM | 2 |
| 2005 | Automatic Keyword Extraction by Server Log Analysis
Chen Ding 0004, Chihung Chi |
WISE | 3 |
| 2004 | Evaluation of Data Streaming Effect on Real-Time Content Transformation in Active Proxies
Chihung Chi, Hongguang Wang |
EUC | 1 |
| 2004 | Survey on the Technological Aspects of Digital Rights Management
William Ku, Chihung Chi |
ISC | 2 |
| 2004 | A Method to Obtain Signatures from Honeypots Data
Chihung Chi, Ming Li 0002, Dongxi Liu |
NPC | 1 |
| 2003 | Secure Route Structures for the Fast Dispatch of Large-Scale Mobile Agents
Yan Wang 0002, Chihung Chi, Tieyan Li |
ICICS | 2 |
| 2003 | Efficient Key Assignment Scheme for Mobile Agent SystemsabstractAgent security is always a key concern to many mobile agent systems such as distributed intrusion detection systems (DIDS), where an attacker might turn to the mobile agent directory server if he fails to locate the critical IDS hosts/agents. Previous work on mobile agent security (e.g. Volker and Mehrdad's scheme, 1998) often suffers from large code size of the agent and the high computational cost. In this paper, we would like to address these two issues with a new RSA-based key assignment scheme. Both the details of the scheme and its implementation are given. Our analysis shows that not only the above two issues are improved significantly, but the scheme also provides agent's adaptability through dynamic key management and access control strategy. Wai-Meng Chew, Chihung Chi, Tieyan Li |
ICTAI | 2 |
| 2003 | A Generalized Site Ranking Model for Web IRabstractNormally, the unit for a ranking model in a Web IR system is a Web page, which is, sometimes, just an information fragment. A larger unit considering the linkage information may be desired to reduce the cognitive overload for users to identify the complete information from the interconnected Web. We propose a ranking model to measure the relevance of the whole Web site. We take some illustrations to show the idea and provide evidences to indicate its effectiveness. Chen Ding 0004, Chihung Chi |
Web Intelligence | 2 |
| 2002 | Runtime Association of Software Prefetch Control to Memory Access Instructions (Research Note)
Chihung Chi, Jun-Li Yuan |
Euro-Par | 1 |
| 2002 | Scalable multimedia content delivery on InternetabstractWe propose a new multimedia data model, called the scalable data model (SDM), to describe how scalable pervasive Internet access with efficient data reuse can be achieved. This new data model not only provides a theoretical foundation for maximum possible reuse of transcoded data for scalable pervasive Internet access, but its supporting architecture also ensures the feasibility and practicability because of its matching to the emerging trends in HTTP and multimedia data format. Chihung Chi, Tiejian Luo |
ICME (1) | 1 |
| 2002 | Context Query in Information RetrievalabstractThere is an important query requirement missing for search engines. With the wide variation of domain knowledge and user interest, a user would like to retrieve documents in which one query term is discussed in the context of another. Based on existing query mechanisms, what can be specified at most is the co-occurrence of multiple terms in a query. This is insufficient because the co-occurrence of two terms does not necessarily mean that one is discussed in the context of the other. In this paper we propose the context query for Web searching. A new query operator, called the 'in' operator, is used to specify context inclusion between two terms. Heuristic rules to identify context inclusion are suggested and implementation of the 'in' operator in search engines is proposed. Results show that both the precision and ranking relevance of Web searching are improved significantly. Chihung Chi, Chen Ding 0004, Kwok-Yan Lam |
ICTAI | 1 |
| 2002 | Study for Fusion of Different Sources to Determine RelevanceabstractThe relevance of a Web document could be measured not only by its text content, but also by some other factors such as the link connectivity and usage patterns. In previous data fusion researches, the text is the only source to determine the relevance, and only the different runs (e.g. by different retrieval models, different query or document representations) on this same source are combined. It is the purpose of this paper to investigate whether the different sources can be combined to determine the relevance with a better accuracy than any single source. We conducted a preliminary experiment to test its feasibility and effectiveness and a positive result was obtained. Chihung Chi, Chen Ding 0004, Kwok-Yan Lam |
ICTAI | 1 |
| 2002 | Agent Warehouse: A New Paradigm for Mobile Agent DeploymentabstractThis paper describes a novel concept of agent warehouse. Non-mobile Web agents typically operate from their users' computer and make request for data possibly from a very far location. In addition, much of this data will be irrelevant to the user, thus aggravating the bandwidth scarcity problem of the Internet. With current mobile agent paradigm, many of these problems such as bandwidth reduction and off-line autonomous negotiation are solved. However, this paradigm does have some significant limitations in its common deployment scenarios; system resource consumption, server collaboration, and accumulated agent size along the travelling path, etc. are some typical ones. These limitations are becoming more important when multiple visits to the same server host are required: updating time of host information is nondeterministic and the decision of negotiation is also not simultaneous. In this paper, the intermediate "proxy-like" agent warehouse is proposed to address these issues. The agent warehouse locates near the data sources and supports agent execution, thus allowing agents to operate much closer to these data sources and minimising the effect of discarded search results. In addition, it is able to provide more resources than a normal Web server host does as it is dedicated to cater for agents. More importantly, even if the remote site does not support agent execution, the agent will still be able to complete its task through the warehouse. This changes the typical approach of how agents can be deployed by providing a more generic, flexible system environment for agents to execute. Chihung Chi, John Sim, Kwok-Yan Lam |
ICTAI | 1 |
| 2002 | Progressive proxy-based multimedia transcoding system with maximum data reuseabstractIn this paper, we propose a new proxy-based transcoding system based on a novel, scalable, and progressive transfer model for multimedia data. This system features on-the-fly transcoding on streaming web data with negligible overhead, maximum data reuse among different transcoded versions of a web object, and the optimal bandwidth usage between the proxy and the server. The importance of the system reflects not only in its feasibility in implementation but also in its matching with the new HTTP protocol support for partial object retrieval and the emerging new multimedia data formats such as JPEG2000 and MPEG4. Chihung Chi |
ACM Multimedia | 1 |
| 2002 | An Improved Usage-Based Ranking
Chen Ding 0004, Chihung Chi, Tiejian Luo |
WAIM | 2 |
| 2002 | A matching-based algorithm for page access sequencing in join processing
Andrew Lim 0001, Wee-Chong Oon, Chihung Chi |
J. Syst. Softw. | 3 |
| 2001 | Automatic proxy-based watermarking for WWW
Chihung Chi, Yan-Hong Lin, Tat-Seng Chua |
Comput. Commun. | 1 |
| 2001 | Load-balancing data prefetching techniques
Chihung Chi, Jun-Li Yuan |
Future Gener. Comput. Syst. | 1 |
| 2000 | Reverse mapping of referral links from storage hierarchy for Web documentsabstractDue to the lack of back-pointers, information about the referral parent(s) of a given Web page is not usually available to Web surfers. This significantly reduces the effectiveness of Web surfing and information discovery. We propose a mechanism to locate the referral parent(s) for a given fragment of a Web document (in terms of an URL address) from the client side. The basic idea is to explore the hierarchical hyperlink structure of a Web document fragment from its storage path directory information that is embedded in its URL. Two mechanisms were proposed: exhaustibility first search (EFS) and generalization first search (GFS). Results showed that through direct mapping of the storage hierarchy path to the URL address and then performing reverse tracing along the hyperlink structure, the chance of discovering a referral parent ranged from 41% to 51%. Compared to EFS, GFS was found to reduce the average search space from 26 pages to 4 pages. Chen Ding 0004, Chihung Chi, Vincent W. L. Tam |
ICTAI | 2 |
| 2000 | Towards an adaptive and task-specific ranking mechanism in Web searchingabstractNo abstract available. Chen Ding 0004, Chihung Chi |
SIGIR | 2 |
| 2000 | Beyond the traditional query operatorsabstractNo abstract available. Chen Ding 0004, Chihung Chi |
SIGIR | 2 |
| 1999 | Word Segmentation and Recognition for Web Document FrameworkabstractIt is observed that a better approach to Web information understanding is to base on its document framework, which is mainly consisted of (i) the title and the URL name of the page, (ii) the titles and the URL names of the Web pages that it points to, (iii) the alternative information source for the embedded Web objects, and (iv) its linkage to other Web pages of the same document. Investigation reveals that a high percentage of words inside the document framework are “compound words” which cannot be understood by ordinary dictionaries. They might be abbreviations or acronyms, or concatenations of several (partial) words. To recover the content hierarchy of Web documents, we propose a new word segmentation and recognition mechanism to understand the information derived from the Web document framework. A maximal bi-directional matching algorithm with heuristic rules is used to resolve ambiguous segmentation and meaning in compound words. An adaptive training process is further employed to build a dictionary of recognisable abbreviations and acronyms. Empirical results show that over 75% of the compound words found in the Web document framework can be understood by our mechanism. With the training process, the success rate of recognising compound words can be increased to about 90%. Chihung Chi, Chen Ding 0004, Andrew Lim 0001 |
CIKM | 1 |
| 1999 | Design Consideration for Multi-lingual Cascading Text CompressorsabstractSummary form only given. We study the cascading of LZ variants to Huffman coding for multilingual documents. Two models are proposed: the static model and the adaptive (dynamic) model. The static model makes use of the dictionary generated by the LZW algorithm in Chinese dictionary-based Huffman compression to achieve better performance. The dynamic model is an extension of the static cascading model. During the insertion of phrases into the dictionary the frequency count of the phrases is updated so that a dynamic Huffman tree with variable length output tokens is obtained. We propose a new method to capture the "LZW dictionary" "by picking up the dictionary entries during decompression. The general idea is the adding of delimiters during the decompression process so that the decompressed files are segmented into phrases that reflect how the LZW compressor makes use of its dictionary phrases to encode the source. The idea of the adaptive cascading model can be thought as an extension of the Chinese LZW compression. Since the size of the header is one important performance bottleneck in the static cascading model we propose the adaptive cascading model to address this issue. The LZW compressor is now outputting not a fixed length token, but a variable length Huffman code from the Huffman tree. It is expected that such a compressor can achieve very good compression performance. In our adaptive cascading model we choose LZW instead of LZSS because the LZW algorithm preserves more information than the LZSS algorithm does. This characteristic is found to be very useful in helping Chinese compressors to attain better performance. Chihung Chi, Yan Zhang IV |
Data Compression Conference | 1 |
| 1999 | Design Considerations of High Performance Data Cache with Prefetching
Chihung Chi, Jun-Li Yuan |
Euro-Par | 1 |
| 1999 | Load-Balancing Branch Target Cache and Prefetch BufferabstractSophisticated branch prediction and compiler optimization technologies result in a higher predictability of instruction references, thus making the branch target cache and prefetch buffer (BTC+PB) design appealing. However, it is surprising to find that this BTC+PB design actually performs worse than the non-partitioned instruction cache. Further investigation shows that this degradation is mainly due to the limited bus bandwidth available for prefetching. To make up for this situation, we propose two load-balancing mechanisms for the BTC+PB design: multi-blocks target (MBT) and dynamic prefetched instruction placement (DIP) techniques. The basic ideas of these two techniques are to tradeoff cache space for bus bandwidth once the bus is found to be overloaded by prefetching. The resulting cache, called the LB+PB design, is found to have superior performance over current non-partitioned instruction cache designs do. Based on the SPEC95, the memory latency due to instruction references can be reduced by an average of 5% to 15%, with some benchmarks whose improvement can go up to over 50%. Chihung Chi, Jun-Li Yuan |
ICCD | 1 |
| 1999 | Cyclic dependence based data reference predictionabstractCurrent work in data prediction and prefetching is mainly focused on one individual predictor per each class of data references.Despite the increasing complexity of these hybrid predictors, their prefetch coverage is still very limited.To reduce the complexity of the predictors and to expand the coverage of data prediction, we propose a novel data predictor, called the Cyclic Dependence based data Predictor (CDP), in this paper.Based on the runtime analysis for value dependence, registers used in the address calculation of a memory access instruction in a loop are classified into cyclic dependent registers and acyclic dependent registers.The complexity and predictability of data references will be determined by the path length of the cycle that contains the index pointer registers.To further illustrate its importance, we generalize our previously proposed Reference Value Prediction Cache [5] with this CDP predictor.Simulation shows that significant reduction in memory latency can be obtained, especially for those using complex data pointer structures. Chihung Chi, Jun-Li Yuan, Chin-Ming Cheung |
International Conference on Supercomputing | 1 |
| 1998 | Study on Mult-lingual LZ77 and LZ78 Text CompressionabstractSummary form only given. We studied the effectiveness of this multi-lingual character sampling on Lempel-Ziv (LZ) compression algorithms. LZSS and LZW algorithms were chosen to represent LZ77 and LZ78 compression respectively in the study. They were modified to adapt the characteristics of non-English information such as Chinese. It is interesting to see that the Chinese LZW compression outperforms the original one by a larger percentage than the Chinese LZSS compression does (14.5% vs. 3.7% on average). CLZW also performs better than CLZSS. This can be explained by two factors: the overall dictionary size and the constraints in each of the algorithms. The dictionary size in the LZ78 algorithms or the sliding window size in the LZ77 algorithms determines how much previous content that the compressor can make use of in order to find repeated phrases. Our result shows that the Chinese LZ78 compressor can make use of a larger dictionary much more effectively than the sliding window in LZ77 family does without introducing any bad side-effects. This also illustrates that previous content is particularly helpful in compressing Chinese text. In terms of the linguistic structure of the Chinese language, the occurrence of repeated phrases in Chinese text does not occur as often as that in English. In other words, within a small, fixed amount of text, it is easier to find repeated phrases in English text than that in Chinese text. Since the LZW preserves large volume of previous content, the Chinese implementation can make good use of it. The difference constraints in the two algorithms also contribute to their performance difference. From the analysis, we can conclude that the LZ88 algorithm (and thus the LZW) is a more suitable Lempel-Ziv family to extend for multi-lingual text compression than the LZ77 does. Chihung Chi |
Data Compression Conference | 1 |
| 1998 | Data prefetching with co-operative cachingabstractRecent research in data cache prefetching is found to be selective in nature: achieving high prediction accuracy over a set of selected references such as array access with constant strides. As a result, for applications where the memory latency is mainly due to data accesses in the set of non selected references of a program, they lose their effectiveness. In fact, their performance might be worse than that of the traditional, less accurate prefetch-on-miss scheme. To overcome this situation, we propose three cooperative cache techniques to assist data prefetching. They are: [1] default prefetching to increase the overall prefetch coverage; [2] block concept to perform variable distance lookahead prefetching; and [3] a spatial data buffer with load balancing to reduce the interference between spatial data and temporal data. To illustrate the potentials of these techniques, they were implemented on top of our previously proposed Instruction Opcode-Based Prefetching (IOBP) scheme (T.F. Chen, 1993). Trace driven simulation on SPEC92 showed that a 8 Kbytes data cache with a 512 bytes spatial buffer can achieve similar performance as a 32 Kbytes data cache through these techniques. Chihung Chi, Siu-Chung Lau |
HiPC | 1 |
| 1998 | Hardware-driven Prefetching for Pointer Data ReferencesabstractProceedings of the International Conference on Supercomputing Chihung Chi, Chin-Ming Cheung |
International Conference on Supercomputing | 1 |
| 1995 | Reducing data access penalty using intelligent opcode-driven cache prefetchingabstractIn the latest processor architectures such as IBM PowerPC and HP Precision Architecture (PA), it is found that certain important compound opcodes such as LOAD-UPDATE and LOAD-MODIFY contain accurate information about how data will be referenced in the near future. Furthermore, these opcodes have been fully utilized by the compiler in the program code generation. With the migration of data cache onto the processor chip, it is now possible for the on-chip cache controller to perform intelligent data prefetching based on the information from the instruction decode unit. In this paper, a novel hardware-driven data prefetching scheme, called the Instruction Opcode-Based Prefetching (IOBP), is proposed. Our simulation shows that this IOBP scheme is very effective in reducing processor stall time due to memory accesses, especially for array or pointer references with constant strides. Chihung Chi, Siu-Chung Lau |
ICCD | 1 |
| 1994 | Compiler Optimization Technique for Data Cache Prefetching Using a Small CAM ArrayabstractWith advances in compiler optimization and program flow analysis, software assisted cache prefetching schemes using PREFETCH instructions are now possible. Although data can be prefetched accurately into the cache, the runtime overhead associated with these schemes often limits their practical use. In this paper, we propose a new scheme, called the Stride_CAM Data Prefetching (SCP), to prefetch array references with constant strides accurately. Compared to current software assisted data prefetching schemes, the SCP scheme has much lower runtime overhead without sacrificing prefetching accuracy. Our result showed that the SCP scheme is particularly suitable for computing intensive scientific applications where cache misses are mainly due to array references with constant strides and they can be prefetched very accurately by this SCP scheme. Chihung Chi |
ICPP (1) | 1 |
| 1990 | Improving instruction cache behavior by reducing cache pollutionabstractCompiler techniques for improving instruction cache performance are described. Through repositioning of the code in the main memory, leaving memory locations unused, code duplication, and code propagation, the effectiveness of the cache can be improved due to reduced cache pollution and fewer cache misses. Results of experiments indicate that significant reduction in bus traffic results from the use of these techniques. Since memory bandwidth is a critical resource in shared memory multiprocessors, such systems can benefit from the techniques described. The notion of control dependence is used to decide when instructions belonging to different basic blocks can be allowed to share the same cache line without increasing cache pollution.> Rajiv Gupta 0001, Chihung Chi |
SC | 2 |
| 1989 | Unified Management of Registers and Cache Using Liveness and Cache BypassabstractIn current computer memory system hierarchy, registers and cache are both used to bridge the reference delay gap between the fast processor(s) and the slow main memory. While registers are managed by the compiler using program flow analysis, cache is mainly controlled by hardware without any program understanding. Due to the lack of coordination in managing these two memory structures, significant loss of system performance results because: Chihung Chi, Henry G. Dietz |
PLDI | 1 |
| 1988 | CRegs: a new kind of memory for referencing arrays and pointersabstractPointer and subscripted array references often touch memory locations for which there are several possible aliases; hence these references cannot be made from registers. Although conventional caches can increase performance somewhat, they do not provide many of the benefits of registers, and do not permit the compiler to perform many optimizations associated with register references. The CReg (pronounced 'C-Reg') mechanism combines the hardware structures of cache and registers to create a novel kind of memory structure, which can be used either as processor registers or as a replacement for conventional cache memory. By permitting aliased names to be grouped together. CRegs resolve ambiguous alias problems in hardware, resulting in more efficient execution that even the combination of conventional registers and cache can provide. The authors discuss both the conceptual CReg hardware structure and the compiler analysis and optimization techniques to manage that structure.> Henry G. Dietz, Chihung Chi |
SC | 2 |