Dan Ye 0004

dblp:37/3880-4 · DBLP profile ↗
← Back
33ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0001-5409-6464ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 19 · 11 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ALERT: Adversarial Learning Enhanced Stability-aware Routing Transformer for Adaptive Depression Detection
Liangyi Kang, Jie Liu 0008, Dan Ye 0004
AAAI5
2025 Root Cause Analysis of RISC-V Build Failures via LLM and MCTS Reasoning
abstract
Build failures are a major obstacle in RISC-V software migration, often involving complex interactions across logs, configurations, and environments. Traditional diagnostic tools struggle with the unstructured, multi-phase nature of build logs and lack semantic reasoning.We propose a two-stage framework for automated root cause analysis. RV-LAD compresses logs using template-based filtering and applies phase-aware anomaly detection via few-shot LLM prompting. MCTS-RCA integrates a domain-specific knowledge base with Monte Carlo Tree Search to perform LLM-guided multi-source reasoning under classification constraints.To support evaluation, we construct a curated dataset of 117 real-world RISC-V build failures, each annotated with logs, spec files, and repair records. Experiments show our approach achieves 75.2% diagnosis accuracy, surpassing previous LLM-based and rule-based methods. It also offers interpretable reasoning traces, enabling practical and transparent diagnosis. This work provides an effective and extensible solution for RCA in emerging software ecosystems like RISC-V, bridging large language models with domain-aware inference.
Weipeng Shuai, Jie Liu 0008, Zhirou Ma, Liangyi Kang, Dan Ye 0004, Wei Wang 0049
ASE7
2025 SCodeGen: A Real-Time Trustworthy Constrained Decoding Framework for Secure Code Generation with LLMs
abstract
Large language models (LLMs) are increasingly integrated into software development workflows to accelerate code generation, but often produce insecure and uncontrollable code due to vulnerable training data and unconstrained decoding strategies. This poses severe risks in security-critical systems, where post-generation vulnerability detection and manual remediation incur significant overhead. While constrained decoding offers a practical mitigation strategy, existing methods suffer from degraded trustworthiness, constraint conflicts, and high latency—especially when enforcing multiple concurrent security constraints.We propose SCodeGen, a real-time constrained decoding framework designed to enforce fine-grained security controls during LLM code generation. To improve trustworthiness and controllability, SCodeGen introduces (1) a matching-length-aware logit modulation strategy that enhances trustworthiness and controllability without semantic disruption, and (2) a two-stage low-latency decoding architecture, which compiles constraint phrases into a runtime-enforceable constraint automaton (RCA) with precomputed logit bias vectors for efficient online decoding. Extensive evaluations on CodeGuard+ show that SCodeGen significantly improves secure pass rates under both single and multi-constraint settings, while maintaining latency comparable to unconstrained decoding. This work demonstrates a practical and scalable solution toward trustworthy LLM-assisted software development under security constraints.
Muzi Qu, Jie Liu 0008, Liangyi Kang, Shuyi Ling, Dan Ye 0004, Tao Huang 0001
TrustCom6
2025 Characterizing and detecting Python version incompatibilities caused by inconsistent version specifications
Haocheng Gao, Wei Chen 0018, Yi Li 0008, Haoxiang Tian 0001, Dan Ye 0004
J. Syst. Softw.7
2024 Chorus: More Efficient Machine Learning on Serverless Platform
Jie Liu 0008, Muzi Qu, Dan Ye 0004, Hua Zhong 0007
DEXA (1)5
2024 Context-Aware Dual Attention Network for Multimodal Sarcasm Detection
abstract
Multimodal sarcasm is often used to express strong emotions online through the discrepancy of the literal-figurative scene across multi-modalities. Current researches retrofit transform-based pretrained language models to integrate text and image to detect sarcasm. However, these methods struggle to distinguish subtle semantic and emotional differences between image and text within the same instance. To address this issue, this paper proposes a new context-aware dual attention network that collaboratively performs textual and visual attentions using a shared memory module. This approach enables us to reason about the interconnected portions involving sarcasm in both text and image. Additionally, we use implicit context derived from multimodal commonsense graph to establish a holistic perspective that encompasses semantics and emotions across modalities. Finally, multi-view cross-modal matching technique is employed to effectively identify contradictions. We evaluate our method on the widely used HFM dataset and achieve 1.01% improvements on the F1-score. Extensive experiments demonstrate the effectiveness of the proposed method.
Liangyi Kang, Jie Liu 0008, Dan Ye 0004
ICASSP3
2024 Dynamic Scoring Code Token Tree: A Novel Decoding Strategy for Generating High-Performance Code
abstract
Within the realms of scientific computing, large-scale data processing, and artificial intelligence-powered computation, disparities in performance, which originate from differing code implementations, directly influence the practicality of the code. Although existing works tried to utilize code knowledge to enhance the execution performance of codes generated by large language models, they neglect code evaluation outcomes which directly refer to the code execution details, resulting in inefficient computation. To address this issue, we propose DSCT-Decode, an innovative adaptive decoding strategy for large language models, that employs a data structure named 'Code Token Tree' (CTT), which guides token selection based on code evaluation outcomes. DSCT-Decode assesses generated code across three dimensions---correctness, performance, and similarity---and utilizes a dynamic penalty-based boundary intersection method to compute multi-objective scores, which are then used to adjust the scores of nodes in the CTT during backpropagation. By maintaining a balance between exploration, through token selection probabilities, and exploitation, through multi-objective scoring, DSCT-Decode effectively navigates the code space to swiftly identify high-performance code solutions. To substantiate our framework, we developed a new benchmark, big-DS-1000, which is an extension of DS-1000. This benchmark is the first of its kind to specifically evaluate code generation methods based on execution performance. Comparative evaluations with leading large language models, such as CodeLlama and GPT-4, show that our framework achieves an average performance enhancement of nearly 30%. Furthermore, 30% of the codes exhibited a performance improvement of more than 20%, underscoring the effectiveness and potential of our framework for practical applications.
Muzi Qu, Jie Liu 0008, Liangyi Kang, Dan Ye 0004, Tao Huang 0001
ASE5
2023 Fixing Robust Out-of-distribution Detection for Deep Neural Networks
abstract
Deep Neural Network (DNN) classifiers easily yield high confidence for Out-of-Distribution (OOD) examples beyond the training distribution, i.e., In-Distribution (ID), leading to classification errors. Detecting and rejecting various OOD examples is crucial for the reliability of DNNs. More challenging, well-built detections can also suffer from being re-bypassed by adversarial attacks perturbing unseen OOD examples. Some existing works introduce adversarial training on the auxiliary outliers to improve the robustness of OOD detection. However, in this work, we find that applying adversarial training on the auxiliary outliers is insufficient to make the detection robust to strong adaptive attacks. To fix this bug of OOD detection, we propose a semi-supervised adversarial training approach, RobDet, which mines adversarially perturbed ID examples from within the neighborhood of clean ID ones as auxiliary outliers and uses multiple "other" classes to train them together with other auxiliary clean and adversarially perturbed outliers to enhance the robustness of OOD detection without significantly sacrificing the performance on clean OOD examples. Experiments show that RobDet has a significant advantage in detecting malicious OOD examples generated by strong adaptive attacks while maintaining advanced performance in detecting clean OOD examples.
Jie Liu 0008, Wensheng Dou, Liangyi Kang, Muzi Qu, Dan Ye 0004
ISSRE7
2023 EasyPip: Detect and Fix Dependency Problems in Python Dependency Declaration Files
abstract
Environment configuration is the basis for software reuse, enabling developers to reuse specific functions.However, the lack of uniform practice in dependency declaration specifications of Python projects can cause problems for developers trying to install third-party libraries.Existing package management tools are often inadequate to help fix these problems.Fixing these errors requires expensive hours and domain knowledge for developers.To help address related problems, some studies focus on well-maintained and popular Python projects about dependency conflict problems caused by PIP's installation rules.However, many projects in the wild are outside of this scope.We carefully investigate 110 issues in 110 projects in the wild.Based on the comprehensive study, we design and implement EasyPip to automatically detect and fix problems in Python dependency declaration files.Dif-
Jie Liu 0008, Haoxiang Tian 0001, Wei Chen 0018, Liangyi Kang, Dan Ye 0004
SEKE7
2022 Differentially Testing Database Transactions for Fun and Profit
abstract
Database Management Systems (DBMSs) utilize transactions to ensure the consistency and integrity of data. Incorrect transaction implementations in DBMSs can lead to severe consequences, e.g., incorrect database states and query results. Therefore, it is critical to ensure the reliability of transaction implementations.
Ziyu Cui, Wensheng Dou, Qianwang Dai, Jiansen Song, Wei Wang 0049, Jun Wei 0001, Dan Ye 0004
ASE7
2022 Generating Critical Test Scenarios for Autonomous Driving Systems via Influential Behavior Patterns
abstract
Autonomous Driving Systems (ADSs) are safety-critical, and must be fully tested before being deployed on real-world roads. To comprehensively evaluate the performance of ADSs, it is essential to generate various safety-critical scenarios. Most of existing studies assess ADSs either by searching high-dimensional input space, or using simple and pre-defined test scenarios, which are not efficient or not adequate. To better test ADSs, this paper proposes to automatically generate safety-critical test scenarios for ADSs by influential behavior patterns, which are mined from real traffic trajectories. Based on influential behavior patterns, a novel scenario generation technique, CRISCO, is presented to generate safety-critical scenarios for ADSs testing. CRISCO assigns participants to perform influential behaviors to challenge the ADS. It generates different test scenarios by solving trajectory constraints, and improves the challenge of those non-critical scenarios by adding participants’ behavior from influential behavior patterns incrementally. We demonstrate CRISCO on an industrial-grade ADS platform, Baidu Apollo. The experiment results show that our approach can effectively and efficiently generate critical scenarios to crash ADS, and it exposes 13 distinct types of safety violations in 12 hours. It also outperforms two state-of-art ADS testing techniques by exposing more 5 distinct types of safety violations on the same roads.
Haoxiang Tian 0001, Guoquan Wu, Jiren Yan, Jun Wei 0001, Wei Chen 0018, Dan Ye 0004
ASE8
2022 MOSAT: finding safety violations of autonomous driving systems using multi-objective genetic algorithm
abstract
Autonomous Driving Systems (ADSs) are safety-critical systems, and safety violations of Autonomous Vehicles (AVs) in real traffic will cause huge losses. Therefore, ADSs must be fully tested before deployed on real world roads. Simulation testing is essential to find safety violations of ADS. This paper proposes MOSAT, a multi-objective search-based testing framework, which constructs diverse and adversarial driving environment to expose safety violations of ADSs. Specifically, based on atomic driving maneuvers, MOSAT introduces motif pattern, which describes a sequence of maneuvers that can challenge ADS effectively. MOSAT constructs test scenarios by atomic maneuvers and motif patterns, and uses multi-objective genetic algorithm to search for adversarial and diverse test scenarios. Moreover, in order to test the performance of ADS comprehensively during long-mile driving, we design a novel continuous simulation testing technique, which runs the scenarios generated by multiple parallel search processes alternately in the simulator and can continuously create different perturbations to ADS. We demonstrate MOSAT on an industrial-grade platform, Baidu Apollo, and the experimental results show that MOSAT can effectively generate safety-critical scenarios to crash ADSs and it exposes 11 distinct types of safety violations in a short period of time. It also outperforms state-of-the-art techniques by finding more 6 distinct safety violations on the same road.
Haoxiang Tian 0001, Guoquan Wu, Jiren Yan, Jun Wei 0001, Wei Chen 0018, Dan Ye 0004
ESEC/SIGSOFT FSE8
2021 FaasRS: Remote Sensing Image Processing System on Serverless Platform
abstract
Big data processing is now the primary mission in remote sensing processing, fortunately, cloud computing provides a feasible approach to perform it efficiently. But the work of resource provisioning, scheduling, and scaling is still inevitable in most cloud computing solutions, it poses a considerable challenge to data analyst. The emerging serverless architecture presents a new paradigm to provide a cloud service, the user only needs to upload function codes and leaves all the other server management jobs to the service provider. It reveals a new possibility of remote sensing processing. This paper presents FaasRS, a framework to process remote sensing images upon serverless platform. FaasRS is built on AWS Lambda, it exposes only simple APIs to operate images, and builds DAG for user’s algorithm. FaasRS splits task by splitting the image into small tiles based on geospatial region, and uses each Lambda worker to perform the computation for one tile. To reduce the redundant operations, we also make optimizations based on the algorithm DAG. FaasRS shows favorable performance and scalability in our evaluation. In the comparison with Spark and Ray, FaasRS shows a significant performance improvement in different type of RS processing jobs.
Jie Liu 0008, Muzi Qu, Dan Ye 0004, Hua Zhong 0007
COMPSAC5
2021 Label Definitions Augmented Interaction Model for Legal Charge Prediction
Liangyi Kang, Jie Liu 0008, Lingqiao Liu, Dan Ye 0004
ECIR (1)4
2021 Meta-graph Embedding in Heterogeneous Information Network for Top-N Recommendation
abstract
Heterogeneous Information Network (HIN) is a graph that contains variety of nodes and their relationships. It can provide abundant auxiliary information for the feature engineering of the recommendation model and thus help to improve its recommendation performance. Most work applying the auxiliary information is to calculate node similarities over meta-paths or meta-graphs of HIN and then recommend based on those similarities through matrix factorization or other analogous recommendation algorithms. In this paper, we propose a novel meta-graph embedding based deep learning recommendation model, MGRec. Types of meta-graphs of HIN are embedded as input features through multiple same structured Attention-enhanced CNNs, which help to learn the weight of each node and get a more accurate vector representation of the meta-graph. Besides, a Wide&Multi-Deep structured recommendation framework is designed to learn both the shallow and deep interactions among features, in which multiple independent deep modules are used to learn the distinguishable correlation degree of each type of meta-graph to the target user and item to highlight the distinguishable contribution of each meta-graph to the recommendation. Experiments on two real-world datasets show that, compared with other popular recommendation models, our MGRec model achieves the best performance in multiple evaluation metrics.
Chengye Cai, Jie Liu 0008, Dan Ye 0004
IJCNN4
2021 Identity-linked Group Channel Pruning for Deep Neural Networks
abstract
Channel pruning is a commonly used model compression in convolutional neural network. The structured pruning using sparse constraints can automatically learn the importance of parameters during the training process by imposing sparse constraints on parameters. However, existing pruning methods based on sparse constraints cannot process the final convolutional layer of the residual module with complex connections. Due to the existence of residual connection, if the final convolutional layer of the residual module is pruned, the sparse channel of the feature map from residual connection does not correspond to the feature map from module output, which will cause the parameters to be unable to be pruned. This paper studies this problem and proposes an identity association group pruning algorithm, which we call IGP. IGP groups the parameters and channels that generate the corresponding feature maps, uses Group Lasso to sparse the same group of parameters as a whole, and forces the sparseness of the parameters with sparse correlation to be consistent with each other. Experiments show that when IGP compresses ResNet56 60% parameters, the model performance only drops 0.36 %, which is better than the existing pruning method based on sparse constraints. In the case of high compression ratio, IGP can compresses ResNet-50 compressesed with 87% parameters and the performance drops only 0.76%, which is 5.17 % higher than the existing methods.
Chenxin Zhang, Keqin Xu, Jie Liu 0008, Liangyi Kang, Dan Ye 0004
IJCNN6
2021 Semantic table structure identification in spreadsheets
abstract
Spreadsheets are widely used in various business tasks, and contain amounts of valuable data. However, spreadsheet tables are usually organized in a semi-structured way, and contain complicated semantic structures, e.g., header types and relations among headers. Lack of documented semantic table structures, existing data analysis and error detection tools can hardly understand spreadsheet tables. Therefore, identifying semantic table structures in spreadsheet tables is of great importance, and can greatly promote various analysis tasks on spreadsheets.
Haoyu Dong 0001, Wensheng Dou, Shi Han, Dongmei Zhang 0001, Jun Wei 0001, Dan Ye 0004
ISSTA8
2021 DeepCon: Contribution Coverage Testing for Deep Learning Systems
abstract
Deep learning (DL) has been widely adopted in many safety-critical scenarios. Deep neural networks (DNNs) usually play the core part in these DL systems. Existing studies have shown that DNNs can suffer from various vulnerabilities, and cause severe consequences. To improve the testing adequacy of DNNs, researchers have proposed several coverage criteria, e.g., neuron coverage in DeepXplore. The prediction result of a DNN is jointly determined by the outputs of neurons and the connection weights that they connect into next-level neurons. However, existing coverage criteria use only the output of a neuron to determine the activation state of the neuron and ignore the connection weights it emits.In this paper, we propose DeepCon, a novel contribution coverage. In DeepCon, we define a term contribution as the combination of the output of a neuron and the connection weight it emits, and use the contribution coverage to gauge the testing adequacy of DNNs. DeepCon can thoroughly cover both neurons and the connection weights they emit and can scale well to large DNNs. We further propose a contribution coverage guided test generation approach, DeepCon-Gen, which can automatically generate tests and activate inactivated contributions of DNNs. We evaluate DeepCon and DeepCon-Gen on five different DNNs over two popular datasets. The experimental results show that DeepCon can well present the testing adequacy of these DNNs. DeepCon-Gen can effectively activate the inactivated contributions, and 62.6% of the generated tests can lead to mispredictions.
Wensheng Dou, Jie Liu 0008, Chenxin Zhang, Jun Wei 0001, Dan Ye 0004
SANER6
2021 Semi-supervised emotion recognition in textual conversation via a context-augmented auxiliary training task
Liangyi Kang, Jie Liu 0008, Lingqiao Liu, Dan Ye 0004
Inf. Process. Manag.5
2020 Learning to detect table clones in spreadsheets
abstract
In order to speed up spreadsheet development productivity, end users can create a spreadsheet table by copying and modifying an existing one. These two tables share the similar computational semantics, and form a table clone. End users may modify the tables in a table clone, e.g., adding new rows and deleting columns, thus introducing structure changes into the table clone. Our empirical study on real-world spreadsheets shows that about 58.5% of table clones involve structure changes. However, existing table clone detection approaches in spreadsheets can only detect table clones with the same structures. Therefore, many table clones with structure changes cannot be detected.
Wensheng Dou, Jun Wei 0001, Dan Ye 0004
ISSTA7
2017 Fine-grained Patient Similarity Measuring using Deep Metric Learning
abstract
Patient similarity measuring plays a significant role in many healthcare applications, such as cohort study and treatment comparative effectiveness research. Existing methods mainly rely on supervised metric learning method to study patient similarity from Electronic Health Records (EHRs), facing the challenge of differentiating patients with a large number of fine-grained disease categories. Deep metric learning has gained noticeable success in fine-grained image categorization problem, however, it cannot be directly applied to classification of patients with hierarchical disease labels. In this paper, we present a novel three layer patient similarity deep metric learning framework (PSDML) by optimizing quadruple loss improved from triplet loss, to learn an embedding distance for disease classification among the patients. The context semantic relation of multi diagnosis labels encoding by ICD-10 is taken into account to compute the supervised distance of patients. To solve the diagnosis class imbalance, patient tuples that violate deep metric learning framework loss constraints are chosen prior as samples to accelerate the convergence of the neural network. We conducted KNN multi label classification experiment using the learned similarity metric on the real EHRs about stroke disease collected by Chinese Stroke Data Center. The results demonstrate substantial improvement over the baselines.
Jiazhi Ni, Jie Liu 0008, Chenxin Zhang, Dan Ye 0004, Zhirou Ma
CIKM4
2017 Fast and Precise recovery in Stream processing based on Distributed Cache
abstract
Stream processing system (SPS) faces the problem of node failure when running over a long period of time. In addition, "exactly once" precise semantic guarantee is more and more important for SPS in some scenarios. In general, the approaches to achieve precise semantic is by using global snapshot, which should store state and records to external reliable storage or rely on transactions. However, these approaches suffer from high recovery latency, because of large I/O disk overhead. In order to reduce excessive latency in failure recovery, we save the intermediate results which are produced during the stream processing, and propose an algorithm DCAS which asynchronously snapshots state to implements precise recovery. In addition, we use in-memory distributed cache to provide the storage of intermediate results and snapshots to reduce recovery latency. We evaluate our failure recovery approach in recovery latency and runtime overhead. The experimental results show that our approach is 2 to 6 times faster than other conventional failure recovery approaches, and induces a 6% runtime overhead.
Yingying Zheng, Wei Wang 0049, Lijie Xu, Zhongshan Ren, Jun Wei 0001, Dan Ye 0004
Internetware7
2016 Hug the Elephant: Migrating a Legacy Data Analytics Application to Hadoop Ecosystem
abstract
Big data applications that rely on relational databases gradually expose limitations on scalability and performance. In recent years, Hadoop ecosystem has been widely adopted as an evolving solution. This paper presents the migration of a legacy data analytics application in a provincial data center. The target platform follows "no one size fits all" method. Considering different workloads, data storage is hybrid with distributed file system (HDFS) and distributed NoSQL database. Beyond the architecture re-design, we focus on the problem of data model transformation from relational database to NoSQL database. We propose a query-aware approach to free developers from tedious manual work. The approach generates query-specific views (NoView) for NoSQL and re-structures the views to align with NoSQL's data model. Our results show that the migrated application achieves high scalability and high performance. We believe that our practice provides valuable insights (such as NoSQL data modeling methodology), and the techniques can be easily applied to other similar migrations.
Jie Liu 0008, Sa Wang, Lijie Xu, Jixin Ren, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001
ICSME7
2016 Parallel Materialization of Datalog Programs with Spark for Scalable Reasoning
Haijiang Wu, Jie Liu 0008, Tao Wang 0030, Dan Ye 0004, Jun Wei 0001, Hua Zhong 0007
WISE (1)4
2015 A Lightweight Evaluation Framework for Table Layouts in MapReduce Based Query Systems
Jie Liu 0008, Lijie Xu, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001
APWeb4
2014 Scalable Horn-Like Rule Inference of Semantic Data Using MapReduce
Haijiang Wu, Jie Liu 0008, Dan Ye 0004, Jun Wei 0001, Hua Zhong 0007
KSEM3
2013 Consistent Query Answering Based on Repairing Inconsistent Attributes with Nulls
Jie Liu 0008, Dan Ye 0004, Jun Wei 0001, Hua Zhong 0007
DASFAA (1)2
2013 A Distributed Cache Framework for Metadata Service of Distributed File Systems
abstract
Most recent distributed file systems have adopted architecture with an independent metadata server cluster. However, potential multiple hotspots and flash crowds access patterns often cause a metadata service that violates performance Service Level Objectives. To maximize the throughput of the metadata service, an adaptive request load balancing framework is critical. We present a distributed cache framework above the distributed metadata management schemes to manage hotspots rather than managing all metadata to achieve request load balancing. This benefits the metadata hierarchical locality and the system scalability. Compared with data, metadata has its own distinct characteristics, such as small size and large quantity. The cost of useless metadata prefetching is much less than data prefetching. In light of this, we devise a time period-based prefetching strategy and a perfecting-based adaptive replacement cache algorithm to improve the performance of the distributed caching layer to adapt constantly changing workloads. Finally, we evaluate our approach with a hadoop distributed file system cluster.
Jie Liu 0008, Dan Ye 0004, Hua Zhong 0007
ICPADS3
2013 A distributed rule execution mechanism based on MapReduce in sematic web reasoning
abstract
Rule execution is the core step of rule-based semantic web reasoning. However, most existing approaches are centralized, which cannot scale out to reason big semantic web datasets. In this paper, we described a kind of semantic web rule execution mechanism using MapReduce programming model, which not only can handle RDFS and OWL ter Horst semantic rules, but also can be used in SWRL reasoning. Theoretical analysis is present on the scalability of this rule execution mechanism. Result shows that it can scale well as Mapreduce framework.
Haijiang Wu, Jie Liu 0008, Dan Ye 0004, Hua Zhong 0007, Jun Wei 0001
Internetware3
2013 Mining user daily behavior patterns from access logs of massive software and websites
abstract
Everyone has a characteristic pattern of daily activities. This study applies cluster analysis to identify a computer user's daily behavior patterns based on 1000 China users' 4-weeks software and web usage. Clustering models are built for 4 different behavior definition methods with different time period divisions and feature measurement selections. With these patterns, we build classification models to predict new users' daily behavior pattern with their half day activity logs. For example, if we know one user use computer for entertainment in the morning, we can predict his behavior in the afternoon and evening. The prediction model can be used to recommend suitable items to users according to their current behavior status. Our method can get 92.5% prediction correctness for the best.
Jie Liu 0008, Dan Ye 0004, Jun Wei 0001
Internetware3
2010 A new approach to performance optimization of mashups via data flow refactoring
abstract
Mashup tools allow end users graphically build complex mashups using pipes to connect web data sources into a data flow. Because end users are of poor technical expertise, the designed data flows may be inefficient. This paper targets on enhancing the performance of mashups via automatically refactoring the structure of its data flows. First a set of operational semantics features are selected for annotating the operators in data flows and refactoring rules are defined to generate all candidate semantics equivalent data flows. Then a heuristic algorithm is described for accurately searching the data flow of minimal execution time by constructing a partially ordered set of data flows based on their cost estimation. This approach is applicable to general mashup data flows without knowing complete operational semantics of their operators and the efficiency improvement is demonstrated by experiments.
Jie Liu 0008, Jun Wei 0001, Dan Ye 0004, Tao Huang 0001
Internetware3
2009 ETL Workflow Analysis and Verification Using Backwards Constraint Propagation
Jie Liu 0008, Senlin Liang, Dan Ye 0004, Jun Wei 0001, Tao Huang 0001
CAiSE3
2004 POP beyond SODA, Reaching the New Horizon of Service Cooperation
abstract
As we have been gaining more experiences in services provision, online services are becoming increasingly complex. They have moved from simple service provision and invocation to very sophisticated service interaction and cooperation. As presented in this paper, service cooperation will be a promising computation model to achieve overall goals beyond individual capabilities. Based on the supreme wide spread of service-oriented development (SODA), the process-oriented platform (POP), with favourable flexibility derived from late binding, will be the optimum approach to this end. The PI production developed by us is such a system implementation, in which architecture, components, and functionalities are also introduced in detail. We believe that the service cooperation paradigm is a hopeful solution to future service evolution, and the PI system will be an instructive explorer to reach this new horizon
Shaohua Liu 0002, Dan Ye 0004, Jun Wei 0001, Yonglin Xia
COMPSAC2