Junfeng Zhao 0001

dblp:72/3918-1 · DBLP profile ↗
← Back
54ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-1268-5006ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 2 first-author · 23 since 2021Software engineering, systems software and programming languages · 19 · 2 first-authorDatabases, data management, data science and information retrieval · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
abstract
Improving large language models (LLMs) for electronic health record (EHR) reasoning is essential for enabling accurate and generalizable clinical predictions. While LLMs excel at medical text understanding, they underperform on EHR-based prediction tasks due to challenges in modeling temporally structured, high-dimensional data. Existing approaches often rely on hybrid paradigms, where LLMs serve merely as frozen prior retrievers while downstream deep learning (DL) models handle prediction, failing to improve the LLM’s intrinsic reasoning capacity and inheriting the generalization limitations of DL models. To this end, we propose EAG-RL, a novel two-stage training framework designed to intrinsically enhance LLMs’ EHR reasoning ability through expert attention guidance, where expert EHR models refer to task-specific DL models trained on EHR data. Concretely, EAG-RL first constructs high-quality, stepwise reasoning trajectories using expert-guided Monte Carlo Tree Search to effectively initialize the LLM’s policy. Then, EAG-RL further optimizes the policy via reinforcement learning by aligning the LLM’s attention with clinically salient features identified by expert EHR models. Extensive experiments on two real-world EHR datasets show that EAG-RL improves the intrinsic EHR reasoning ability of LLMs by an average of 14.62%, while also enhancing robustness to feature perturbations and generalization to unseen clinical domains. These results demonstrate the practical potential of EAG-RL for real-world deployment in clinical prediction tasks.
Jiaran Gao, Hongxin Ding, Xinke Jiang, Weibin Liao, Yongxin Xu, Yinghao Zhu, Zhibang Yang, Liantao Ma, Junfeng Zhao 0001, Yasha Wang
AAAI11
2026 ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
abstract
Hongxin Ding, Baixiang Huang, Yue Fang, Weibin Liao, Xinke Jiang, Jinyang Zhang, Yinghao Zhu, Zheng Li, Liantao Ma, Junfeng Zhao, Yasha Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hongxin Ding, Baixiang Huang, Weibin Liao, Xinke Jiang, Yinghao Zhu, Liantao Ma, Junfeng Zhao 0001, Yasha Wang
ACL (1)10
2026 DFAMS: Dynamic-flow guided Federated Alignment based Multi-prototype Search
abstract
Zhibang Yang, Xinke Jiang, Rihong Qiu, Ruiqing Li, Yihang Zhang, Yue Fang, Yongxin Xu, Hongxin Ding, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhibang Yang, Xinke Jiang, Rihong Qiu, Yongxin Xu, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
ACL (1)10
2026 Beyond Imputation: A Semantic Unification Framework for Data and its Missingness in Multimodal Healthcare Analytics
Chaohe Zhang, Liantao Ma, Shiwei Lyu, Junfeng Zhao 0001, Yasha Wang
ICDE5
2025 KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language Models
abstract
By integrating external knowledge, Retrieval-Augmented Generation (RAG) has become an effective strategy for mitigating the hallucination problems that large language models (LLMs) encounter when dealing with knowledge-intensive tasks. However, in the process of integrating external non-parametric supporting evidence with internal parametric knowledge, inevitable knowledge conflicts may arise, leading to confusion in the model's responses. To enhance the knowledge selection of LLMs in various contexts, some research has focused on refining their behavior patterns through instruction-tuning. Nonetheless, due to the absence of explicit negative signals and comparative objectives, models fine-tuned in this manner may still exhibit undesirable behaviors such as contextual ignorance and contextual overinclusion. To this end, we propose a Knowledge-aware Preference Optimization strategy, dubbed KnowPO, aimed at achieving adaptive knowledge selection based on contextual relevance in real retrieval scenarios. Concretely, we proposed a general paradigm for constructing knowledge conflict datasets, which comprehensively cover various error types and learn how to avoid these negative signals through preference optimization methods. Simultaneously, we proposed a rewriting strategy and data ratio optimization strategy to address preference imbalances. Experimental results show that KnowPO outperforms previous methods for handling knowledge conflicts by over 37%, while also exhibiting robust generalization across various out-of-distribution datasets.
Ruizhe Zhang 0013, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu, Xinke Jiang, Junfeng Zhao 0001, Yasha Wang
AAAI7
2025 DearLLM: Enhancing Personalized Healthcare via Large Language Models-Deduced Feature Correlations
abstract
Exploring the correlations between medical features is essential for extracting patient health patterns from electronic health records (EHR) data, and strengthening medical predictions and decision-making. To constrain the hypothesis space of pure data-driven deep learning in the context of limited annotated data, a common trend is to incorporate external knowledge, especially knowledge priors related to personalized health contexts, to optimize model training. However, most existing methods lack flexibility and are constrained by the uncertainties brought about by fixed feature correlation priors. In addition, in utilizing knowledge, these methods overlook the knowledge informative for personalized healthcare. To this end, we propose DearLLM, a novel and effective framework that leverages feature correlations deduced by large language models (LLMs) to enhance personalized healthcare. Concretely, DearLLM captures and learns quantitative correlations between medical features by calculating the conditional perplexity of LLMs’ deduction based on personalized patient backgrounds. Then, DearLLM enhances healthcare predictions by emphasizing knowledge that carries unique patient information through a feature-frequency-aware graph pooling method. Extensive experiments on two real-world benchmark datasets show significant performance gains brought by DearLLM. Furthermore, the discovered findings align well with medical literature, offering meaningful clinical interpretations.
Yongxin Xu, Xinke Jiang, Rihong Qiu, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
AAAI7
2025 HyKGE: A Hypothesis Knowledge Graph Enhanced RAG Framework for Accurate and Reliable Medical LLMs Responses
abstract
Xinke Jiang, Ruizhe Zhang, Yongxin Xu, Rihong Qiu, Yue Fang, Zhiyuan Wang, Jinyi Tang, Hongxin Ding, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinke Jiang, Ruizhe Zhang 0013, Yongxin Xu, Rihong Qiu, Jinyi Tang, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
ACL (1)10
2025 TC-RAG: Turing-Complete RAG's Case study on Medical LLM Systems
abstract
Xinke Jiang, Yue Fang, Rihong Qiu, Haoyu Zhang, Yongxin Xu, Hao Chen, Wentao Zhang, Ruizhe Zhang, Yuchen Fang, Xinyu Ma, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinke Jiang, Rihong Qiu, Yongxin Xu, Hao Chen 0103, Wentao Zhang 0008, Ruizhe Zhang 0013, Yuchen Fang 0001, Junfeng Zhao 0001, Yasha Wang
ACL (1)12
2025 Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning
abstract
Yongxin Xu, Ruizhe Zhang, Xinke Jiang, Yujie Feng, Yuzhen Xiao, Xinyu Ma, Runchuan Zhu, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yongxin Xu, Ruizhe Zhang 0013, Xinke Jiang, Yuzhen Xiao, Runchuan Zhu, Junfeng Zhao 0001, Yasha Wang
ACL (1)9
2025 3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
abstract
Hongxin Ding, Yue Fang, Runchuan Zhu, Xinke Jiang, Jinyang Zhang, Yongxin Xu, Weibin Liao, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hongxin Ding, Runchuan Zhu, Xinke Jiang, Yongxin Xu, Weibin Liao, Junfeng Zhao 0001, Yasha Wang
EMNLP9
2025 DRESSing Up LLM: Efficient Stylized Question-Answering via Style Subspace Editing
abstract
We introduce DRESS, a novel approach for generating stylized large language model (LLM) responses through representation editing. Existing methods like prompting and fine-tuning are either insufficient for complex style adaptation or computationally expensive, particularly in tasks like NPC creation or character role-playing. Our approach leverages the over-parameterized nature of LLMs to disentangle a style-relevant subspace within the model's representation space to conduct representation editing, ensuring a minimal impact on the original semantics. By applying adaptive editing strengths, we dynamically adjust the steering vectors in the style subspace to maintain both stylistic fidelity and semantic integrity. We develop two stylized QA benchmark datasets to validate the effectiveness of DRESS, and the results demonstrate significant improvements compared to baseline methods such as prompting and ITI. In short, DRESS is a lightweight, train-free solution for enhancing LLMs with flexible and effective style control, making it particularly useful for developing stylized conversational agents. Codes and benchmark datasets are available at https://github.com/ArthurLeoM/DRESS-LLM.
Tianlong Wang, Junfeng Zhao 0001, Yasha Wang
ICLR7
2025 Efficient Graph Continual Learning via Lightweight Graph Neural Tangent Kernels-based Dataset Distillation
abstract
Graph Neural Networks (GNNs) have emerged as a fundamental tool for modeling complex graph structures across diverse applications. However, directly applying pretrained GNNs to varied downstream tasks without fine-tuning-based continual learning remains challenging, as this approach incurs high computational costs and hinders the development of Large Graph Models (LGMs). In this paper, we investigate an efficient and generalizable dataset distillation framework for Graph Continual Learning (GCL) across multiple downstream tasks, implemented through a novel Lightweight Graph Neural Tangent Kernel (LIGHTGNTK). Specifically, LIGHTGNTK employs a low-rank approximation of the Laplacian matrix via Bernoulli sampling and linear association within the GNTK. This design enables efficient capture of both structural and feature relationships while supporting gradient-based dataset distillation. Additionally, LIGHTGNTK incorporates a unified subgraph anchoring strategy, allowing it to handle graph-level, node-level, and edge-level tasks under diverse input structures. Comprehensive experiments on several datasets show that LIGHTGNTK achieves state-of-the-art performance in GCL scenarios, promoting the development of adaptive and scalable LGMs.
Rihong Qiu, Xinke Jiang, Yuchen Fang 0001, Hongbin Lai, Hao Miao 0001, Junfeng Zhao 0001, Yasha Wang
ICML7
2025 MODEL SHAPLEY: Find Your Ideal Parameter Player via One Gradient Backpropagation
abstract
Measuring parameter importance is crucial for understanding and optimizing large language models (LLMs). Existing work predominantly focuses on pruning or probing at neuron/feature levels without fully considering the cooperative behaviors of model parameters. In this paper, we introduce a novel approach--Model Shapley to quantify parameter importance based on the Shapley value, a principled method from cooperative game theory that captures both individual and synergistic contributions among parameters, via only one gradient backpropagation. We derive a scalable second-order approximation to compute Shapley values at the parameter level, leveraging blockwise Fisher information for tractability in large-scale settings. Our method enables fine-grained differentiation of parameter importance, facilitating targeted knowledge injection and model compression. Through mini-batch Monte Carlo updates and efficient approximation of the Hessian structure, we achieve robust Shapley-based attribution with only modest computational overhead. Experimental results indicate that this cooperative game perspective enhances interpretability, guides more effective parameter-specific fine-tuning and model compressing, and paves the way for continuous model improvement in various downstream tasks.
Xinke Jiang, Rihong Qiu, Jiaran Gao, Junfeng Zhao 0001
NeurIPS5
2025 A three-tiered semi supervised MTL mechanism and its application in dating apps
abstract
Abstract A thorough understanding of the purpose of dating applications is crucial for service providers in order to optimize the design and user experience of the application. Despite the fact that many APPs prompt users to provide their usage purpose, many do not reveal this attribute. In this study, a three-module framework with semi-supervised and multitask learning mechanisms is proposed (T-SSMTL). Using the T-SSMTL mechanism, the purpose of the dating APP usage can be automatically inferred from the publicly available heterogeneous data of the user. The heterogeneous feature extraction module employs a number of techniques to extract semantic representations, maximizing the use of heterogeneous dating APP data. The multi-task module extracts task-specific knowledge for learning and solves the classification problem involving multiple labels. To alleviate the problem of label insufficiency, the semi-supervised module utilizes a large quantity of unlabeled data generated by users who do not report their usage purpose. A large-scale dataset containing 34,364 active dating APP users with their self-reported usage purpose, portrait image, profile, and posts was collected to evaluate the T-SSMTL framework. In the context of this dataset, simulation experiments have confirmed the efficacy of all three modules of the T-SSMTL framework, demonstrating its substantial theoretical significance as well as its excellent application value.
Junyi Ma, Yasha Wang, Xuanliang Wang, Jiangtao Wang 0001, Junfeng Zhao 0001
Neural Comput. Appl.5
2024 Parameter Efficient Quasi-Orthogonal Fine-Tuning via Givens Rotation
abstract
With the increasingly powerful performances and enormous scales of pretrained models, promoting parameter efficiency in fine-tuning has become a crucial need for effective and efficient adaptation to various downstream tasks. One representative line of fine-tuning methods is Orthogonal Fine-tuning (OFT), which rigorously preserves the angular distances within the parameter space to preserve the pretrained knowledge. Despite the empirical effectiveness, OFT still suffers low parameter efficiency at $\mathcal{O}(d^2)$ and limited capability of downstream adaptation. Inspired by Givens rotation, in this paper, we proposed quasi-Givens Orthogonal Fine-Tuning (qGOFT) to address the problems. We first use $\mathcal{O}(d)$ Givens rotations to accomplish arbitrary orthogonal transformation in $SO(d)$ with provable equivalence, reducing parameter complexity from $\mathcal{O}(d^2)$ to $\mathcal{O}(d)$. Then we introduce flexible norm and relative angular adjustments under soft orthogonality regularization to enhance the adaptation capability of downstream semantic deviations. Extensive experiments on various tasks and pretrained models validate the effectiveness of our methods.
Zhibang Yang, Junfeng Zhao 0001
ICML6
2024 ProtoMix: Augmenting Health Status Representation Learning via Prototype-based Mixup
abstract
With the widespread adoption of electronic health records (EHR) data, deep learning techniques have been broadly utilized for various health prediction tasks. Nevertheless, the labeled data scarcity issue restricts the prediction power of these deep models. To enhance the generalization capability of deep learning models when faced with such situations, a common trend is to train generative adversarial networks (GANs) or diffusion models for data augmentation. However, due to limitations in sample size and potential label imbalance issues, these methods are prone to mode collapse problems. This results in the generation of new samples that fail to preserve the subtype structure within EHR data, thereby limiting their practicality in health prediction tasks that generally require detailed patient phenotyping. Aiming at the above problems, we propose a Prototype-based Mixup method, dubbed ProtoMix, which combines prior knowledge of intrinsic data features from subtype centroids (i.e., prototypes) to guide the synthesis of new samples. Specifically, ProtoMix employs a prototype-guided mixup training task to shift the decision boundary away from the subtypes. Then, ProtoMix optimizes the sampling weights in different areas of the data manifold via a prototype-guided mixup sampling strategy. Throughout the training process, ProtoMix dynamically expands the training distribution using an adaptive mixing coefficient computation method. Experimental evaluations on three real-world datasets demonstrate the efficacy of ProtoMix.
Yongxin Xu, Xinke Jiang, Yuzhen Xiao, Chaohe Zhang, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
KDD7
2024 RAGraph: A General Retrieval-Augmented Graph Learning Framework
abstract
Graph Neural Networks (GNNs) have become essential in interpreting relational data across various domains, yet, they often struggle to generalize to unseen graph data that differs markedly from training instances. In this paper, we introduce a novel framework called General Retrieval-Augmented Graph Learning (RAGraph), which brings external graph data into the general graph foundation model to improve model generalization on unseen scenarios. On the top of our framework is a toy graph vector library that we established, which captures key attributes, such as features and task-specific label information. During inference, the RAGraph adeptly retrieves similar toy graphs based on key similarities in downstream tasks, integrating the retrieved data to enrich the learning context via the message-passing prompting mechanism. Our extensive experimental evaluations demonstrate that RAGraph significantly outperforms state-of-the-art graph learning methods in multiple tasks such as node classification, link prediction, and graph classification across both dynamic and static datasets. Furthermore, extensive testing confirms that RAGraph consistently maintains high performance without the need for task-specific fine-tuning, highlighting its adaptability, robustness, and broad applicability.
Xinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang 0008, Ruizhe Zhang 0013, Yuchen Fang 0001, Junfeng Zhao 0001, Yasha Wang
NeurIPS9
2024 SMART: Towards Pre-trained Missing-Aware Model for Patient Health Status Prediction
abstract
Electronic health record (EHR) data has emerged as a valuable resource for analyzing patient health status. However, the prevalence of missing data in EHR poses significant challenges to existing methods, leading to spurious correlations and suboptimal predictions. While various imputation techniques have been developed to address this issue, they often obsess difficult-to-interpolate details and may introduce additional noise when making clinical predictions. To tackle this problem, we propose SMART, a Self-Supervised Missing-Aware RepresenTation Learning approach for patient health status prediction, which encodes missing information via missing-aware temporal and variable attentions and learns to impute missing values through a novel self-supervised pre-training approach which reconstructs missing data representations in the latent space rather than in input space as usual. By adopting elaborated attentions and focusing on learning higher-order representations, SMART promotes better generalization and robustness to missing data. We validate the effectiveness of SMART through extensive experiments on six EHR tasks, demonstrating its superiority over state-of-the-art methods.
Zhihao Yu, Yujie Jin, Yasha Wang, Junfeng Zhao 0001
NeurIPS5
2023 KerPrint: Local-Global Knowledge Graph Enhanced Diagnosis Prediction for Retrospective and Prospective Interpretations
abstract
While recent developments of deep learning models have led to record-breaking achievements in many areas, the lack of sufficient interpretation remains a problem for many specific applications, such as the diagnosis prediction task in healthcare. The previous knowledge graph(KG) enhanced approaches mainly focus on learning clinically meaningful representations, the importance of medical concepts, and even the knowledge paths from inputs to labels. However, it is infeasible to interpret the diagnosis prediction, which needs to consider different medical concepts, various medical relationships, and the time-effectiveness of knowledge triples in different patient contexts. More importantly, the retrospective and prospective interpretations of disease processes are valuable to clinicians for the patients' confounding diseases. We propose KerPrint, a novel KG enhanced approach for retrospective and prospective interpretations to tackle these problems. Specifically, we propose a time-aware KG attention method to solve the problem of knowledge decay over time for trustworthy retrospective interpretation. We also propose a novel element-wise attention method to select candidate global knowledge using comprehensive representations from the local KG for prospective interpretation. We validate the effectiveness of our KerPrint through an extensive experimental study on a real-world dataset and a public dataset. The results show that our proposed approach not only achieves significant improvement over knowledge-enhanced methods but also gives the interpretability of diagnosis prediction in both retrospective and prospective views.
Kai Yang 0053, Yongxin Xu, Peinie Zou, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
AAAI5
2023 VecoCare: Visit Sequences-Clinical Notes Joint Learning for Diagnosis Prediction in Healthcare Data
abstract
Due to the insufficiency of electronic health records (EHR) data utilized in practical diagnosis prediction scenarios, most works are devoted to learning powerful patient representations either from structured EHR data (e.g., temporal medical events, lab test results, etc.) or unstructured data (e.g., clinical notes, etc.). However, synthesizing rich information from both of them still needs to be explored. Firstly, the heterogeneous semantic biases across them heavily hinder the synthesis of representation spaces, which is critical for diagnosis prediction. Secondly, the intermingled quality of partial clinical notes leads to inadequate representations of to-be-predicted patients. Thirdly, typical attention mechanisms mainly focus on aggregating information from similar patients, ignoring important auxiliary information from others. To tackle these challenges, we propose a novel visit sequences-clinical notes joint learning approach, dubbed VecoCare. It performs a Gromov-Wasserstein Distance (GWD)-based contrastive learning task and an adaptive masked language model task in a sequential pre-training manner to reduce heterogeneous semantic biases. After pre-training, VecoCare further aggregates information from both similar and dissimilar patients through a dual-channel retrieval mechanism. We conduct diagnosis prediction experiments on two real-world datasets, which indicates that VecoCare outperforms state-of-the-art approaches. Moreover, the findings discovered by VecoCare are consistent with the medical researches.
Yongxin Xu, Kai Yang 0053, Chaohe Zhang, Peinie Zou, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
IJCAI7
2023 Fused Gromov-Wasserstein Graph Mixup for Graph-level Classifications
abstract
Graph data augmentation has shown superiority in enhancing generalizability and robustness of GNNs in graph-level classifications. However, existing methods primarily focus on the augmentation in the graph signal space and the graph structure space independently, neglecting the joint interaction between them. In this paper, we address this limitation by formulating the problem as an optimal transport problem that aims to find an optimal inter-graph node matching strategy considering the interactions between graph structures and signals. To solve this problem, we propose a novel graph mixup algorithm called FGWMixup, which seeks a "midpoint" of source graphs in the Fused Gromov-Wasserstein (FGW) metric space. To enhance the scalability of our method, we introduce a relaxed FGW solver that accelerates FGWMixup by improving the convergence rate from $\mathcal{O}(t^{-1})$ to $\mathcal{O}(t^{-2})$. Extensive experiments conducted on five datasets using both classic (MPNNs) and advanced (Graphormers) GNN backbones demonstrate that \mname\xspace effectively improves the generalizability and robustness of GNNs. Codes are available at https://github.com/ArthurLeoM/FGWMixup.
Yasha Wang, Junfeng Zhao 0001, Liantao Ma, Wenwu Zhu 0001
NeurIPS5
2023 SeqCare: Sequential Training with External Medical Knowledge Graph for Diagnosis Prediction in Healthcare Data
abstract
Deep learning techniques are capable of capturing complex input-output relationships, and have been widely applied to the diagnosis prediction task based on web-based patient electronic health records (EHR) data. To improve the prediction and interpretability of pure data-driven deep learning with only a limited amount of labeled data, a pervasive trend is to assist the model training with knowledge priors from online medical knowledge graphs. However, they marginally investigated the label imbalance and the task-irrelevant noise in the external knowledge graph. The imbalanced label distribution would bias the learning and knowledge extraction towards the majority categories. The task-irrelevant noise introduces extra uncertainty to the model performance. To this end, aiming at by-passing the bias-variance trade-off dilemma, we introduce a new sequential learning framework, dubbed SeqCare, for diagnosis prediction with online medical knowledge graphs. Concretely, in the first step, SeqCare learns a bias-reduced space through a self-supervised graph contrastive learning task. Secondly, SeqCare reduces the learning uncertainty by refining the supervision signal and the graph structure of the knowledge graph simultaneously. Lastly, SeqCare trains the model in the bias-variance reduced space with a self-distillation to further filter out irrelevant information in the data. Experimental evaluations on two real-world datasets show that SeqCare outperforms state-of-the-art approaches. Case studies exemplify the interpretability of SeqCare. Moreover, the medical findings discovered by SeqCare are consistent with experts and medical literature.
Yongxin Xu, Kai Yang 0053, Peinie Zou, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang
WWW7
2023 An Adaptive Fusion Risk-Zone Detection Network and its Application
abstract
COVID-19 has caused a pandemic and adverse effects in many fields on a global scale. The city scale quarantine has demonstrated its effectiveness in controlling the epidemic. Conversely, it is costly and risky in inducing economic and social challenges. A compromised solution is to place quarantine measures at high-risk zones on a local scale. Therefore, it is important to investigate risk zones for conducting cost insensitive precautionary measures. The urban data depict the characteristics of different city zones, which offers an opportunity for detecting the high-risk zones. Yet, the high noise-to-signal ratio requires an efficient procedure to rule out irrelevant information in the informative raw urban data and adapt to the risk detection task. In this paper, we propose an Adaptive Fusion Risk-zone Detection Network (AFRDN), which fuses the static and dynamic multi-sourced urban data in an adaptive manner. Specifically, AFRDN first extracts diverse information-rich features from raw urban data with various encoders in the embedding learning module. Then, the AFRDN takes a hierarchical late fusion strategy by fusing the static embedding and the attentive hidden state of dynamic features in the deep latent space. To capture the most relevant information for risk-zone detection, the AFRDN adapts each dimension in the fused embedding with multi-head self-attention blocks. We have collected a real-world dataset including six Chinese cities and conducted extensive experiments to evaluate our framework. Simulation experiments and comparative analysis results show that the AFRDN is effective and feasible for early detection of infectious diseases high-risk zones.
Junyi Ma, Xuanliang Wang, Yasha Wang, Junfeng Zhao 0001
Int. J. Pattern Recognit. Artif. Intell.5
2023 Patient Health Representation Learning via Correlational Sparse Prior of Medical Features
abstract
Exploiting the correlations between medical features is essential to the success of healthcare data analysis. However, most existing methods are either suffering large estimation variance for data insufficiency or inflexible in terms of demanding task-specific medical knowledge. In this paper, we propose a novel patient health representation learning framework dubbedSAFARI.SAFARIlearns a compact representation by imposing a clinical-fact-inspired task-agnostic correlational sparsity prior to the correlations of medical feature pairs. Specifically, we learn the compact representation by solving the bi-level optimization problem, which involves solving the high-level inter-group correlations and the nested lower-level intra-group correlations. We leverage the Laplacian kernel as a robust metric for feature grouping and graph neural networks for solving the bi-level optimization problem following the optimal value reformulation paradigm. Experiments on five datasets of various inputs and tasks demonstrate the efficacy ofSAFARI. The discovered findings are also consistent with our insights and medical literature, which can provide valuable clinical explanations.
Yasha Wang, Liantao Ma, Wen Tang 0001, Junfeng Zhao 0001, Ye Yuan 0001, Guoren Wang
IEEE Trans. Knowl. Data Eng.6
2022 Enhancing Robust Text Classification via Category Description
abstract
Despite the success of deep neural networks on text classification, their large capacity also leads to capturing task-irrelevant patterns such as label noise. Label noise is usually introduced into the data during label collection and causes nontrivial declines in performance due to the memorization effect. Though effort has been devoted to combating the label noise in other systems such as image classification, high-quality input features are necessary for discovering task-relevant patterns before memorizing the label noise. However, such a high-quality input feature requirement is hard to be satisfied for text classification due to the nature of natural language. To combat the label noise with low-quality input features in the text classification, we propose a novel framework that exploits external category descriptions to construct prototypes that can be used to denoise the input representation and alleviate the over-fitting. However, there still remains a challenge that the external category descriptions from other corpora could be semantically discrepant with the underlying task-specific classes in the training corpus. To align their semantics, we propose two regularizers that penalize sample-wise semantic-based deviations at the local level and class-wise structure-based deviations at the global level, respectively. Our extensive experiments across two open datasets and one real-world case study demonstrate that our method is superior to state-of-the-art baselines under various settings of label noise.
Zhengye Zhu, Yasha Wang, Wenjie Ruan, Junfeng Zhao 0001
ICDM6
2022 M3Care: Learning with Missing Modalities in Multimodal Healthcare Data
abstract
Multimodal electronic health record (EHR) data are widely used in clinical applications. Conventional methods usually assume that each sample (patient) is associated with the unified observed modalities, and all modalities are available for each sample. However, missing modality caused by various clinical and social reasons is a common issue in real-world clinical scenarios. Existing methods mostly rely on solving a generative model that learns a mapping from the latent space to the original input space, which is an unstable ill-posed inverse problem. To relieve the underdetermined system, we propose a model solving a direct problem, dubbed learning with Missing Modalities in Multimodal healthcare data (M3Care). M3Care is an end-to-end model compensating the missing information of the patients with missing modalities to perform clinical analysis. Instead of generating raw missing data, M3Care imputes the task-related information of the missing modalities in the latent space by the auxiliary information from each patient's similar neighbors, measured by a task-guided modality-adaptive similarity metric, and thence conducts the clinical tasks. The task-guided modality-adaptive similarity metric utilizes the uncensored modalities of the patient and the other patients who also have the same uncensored modalities to find similar patients. Experiments on real-world datasets show that M3Care outperforms the state-of-the-art baselines. Moreover, the findings discovered by M3Care are consistent with experts and medical knowledge, demonstrating the capability and the potential of providing useful insights and explanations.
Chaohe Zhang, Liantao Ma, Yinghao Zhu, Yasha Wang, Jiangtao Wang 0001, Junfeng Zhao 0001
KDD7
2022 LDA-Reg: Knowledge Driven Regularization Using External Corpora
abstract
While recent developments of neural network (NN) models have led to a series of record-breaking achievements in many applications, the lack of sufficiently good datasets remains a problem for some applications. For such a problem, we can however exploit a large number of unstructured text corpora as an external knowledge to complement the training data, and most prevailing neural network solutions employ word embedding methods for such purposes. In this paper, we propose LDA-Reg, a novel knowledge driven regularization framework based on Latent Dirichlet Allocation (LDA) as an alternative to the word embedding methods to adaptively utilize abundant external knowledge and to interpret the NN model. For the joint learning of the parameters, we propose EM-SGD, an effective update method which incorporates Expectation Maximization (EM) and Stochastic Gradient Descent (SGD) to update parameters iteratively. Moreover, we also devise a lazy update and sparse update method for the high-dimensional inputs and sparse inputs respectively. We validate the effectiveness of our regularization framework through an extensive experimental study over real world and standard benchmark datasets. The results show that our proposed framework not only achieves significant improvement over state-of-the-art word embedding methods but also learns interpretable and significant topics for various tasks.
Kai Yang 0053, Zhaojing Luo, Jinyang Gao, Junfeng Zhao 0001, Beng Chin Ooi
IEEE Trans. Knowl. Data Eng.4
2020 COTSAE: CO-Training of Structure and Attribute Embeddings for Entity Alignment
abstract
Entity alignment is a fundamental and vital task in Knowledge Graph (KG) construction and fusion. Previous works mainly focus on capturing the structural semantics of entities by learning the entity embeddings on the relational triples and pre-aligned "seed entities". Some works also seek to incorporate the attribute information to assist refining the entity embeddings. However, there are still many problems not considered, which dramatically limits the utilization of attribute information in the entity alignment. Different KGs may have lots of different attribute types, and even the same attribute may have diverse data structures and value granularities. Most importantly, attributes may have various "contributions" to the entity alignment. To solve these problems, we propose COTSAE that combines the structure and attribute information of entities by co-training two embedding learning components, respectively. We also propose a joint attention method in our model to learn the attentions of attribute types and values cooperatively. We verified our COTSAE on several datasets from real-world KGs, and the results showed that it is significantly better than the latest entity alignment methods. The structure and attribute information can complement each other and both contribute to performance improvement.
Kai Yang 0053, Shaoqin Liu, Junfeng Zhao 0001, Yasha Wang
AAAI3
2019 Trip2Vec: a deep embedding approach for clustering and profiling taxi trip purposes
Chao Chen 0004, Chengwu Liao, Xuefeng Xie, Yasha Wang, Junfeng Zhao 0001
Pers. Ubiquitous Comput.5
2018 Toward accurate link between code and software documentation
Yingkui Cao, Yanzhen Zou, Yuxiang Luo, Junfeng Zhao 0001
Sci. China Inf. Sci.5
2017 TaGiTeD: Predictive Task Guided Tensor Decomposition for Representation Learning from Electronic Health Records
abstract
With the better availability of healthcare data, such as Electronic Health Records (EHR), more and more data analytics methodologies are developed aiming at digging insights from them to improve the quality of care delivery. There are many challenges on analyzing EHR, such as high dimensionality and event sparsity. Moreover, different from other application domains, the EHR analysis algorithms need to be highly interpretable to make them clinically useful. This makes representation learning from EHRs of key importance. In this paper, we propose an algorithm called Predictive Task Guided Tensor Decomposition (TaGiTeD), to analyze EHRs. Specifically, TaGiTeD learns event interaction patterns that are highly predictive for certain tasks from EHRs with supervised tensor decomposition. Compared with unsupervised methods, TaGiTeD can learn effective EHR representations in a more focused way. This is crucial because most of the medical problems have very limited patient samples, which are not enough for unsupervised algorithms to learn meaningful representations form. We apply TaGiTeD on real world EHR data warehouse and demonstrate that TaGiTeD can learn representations that are both interpretable and predictive.
Kai Yang 0053, Xiang Li 0013, Haifeng Liu 0005, Jing Mei, Guo Tong Xie, Junfeng Zhao 0001, Fei Wang 0001
AAAI6
2017 Refining Traceability Links between Code and Software Documents
abstract
Recovering traceability links between source code and software document can be very helpful for Software Maintenance and Software Reuse. Existing work has already achieved good results in extracting code elements (classes, methods, etc.) from software documents. However, it will lead to a lot of noise links if we link a document to all the code elements existing in it. In this paper, we propose an approach to identify the contextual code elements and the salient code elements in a software document, then we can weight the traceability links between source code and software document so that those noise traceability links can be filtered effectively. We measure the saliency of each code element in a document with four kinds of document-related features and three kinds of code-related features, and we adopt TransR-based code embedding technology to evaluate the distance between code elements. In the experiments, we get a precision of 70.7% in recognizing salient code elements of StackOverflow answer documents, which is more than 12% improvement compared with Rigby's work. At the same time, we can filter about 56.5%~69.3% noise traceability links compared with the RecoDoc approach. It will improve the quality of traceability links between source code and related software documents.
Yingkui Cao, Yanzhen Zou, Yuxiang Luo, Junfeng Zhao 0001
Internetware5
2017 Document Distance Estimation via Code Graph Embedding
abstract
Accurately representing the distance between two documents (i.e. pieces of textual information extracted from various software artifacts) has far-reaching applications in many automated software engineering approaches, such as concept location, bug location and traceability link recovery. This is a challenging task, since documents containing different words may have similar semantic meanings. In this paper, we propose a novel document distance estimation approach. This approach captures latent semantic associations between documents through analyzing structural information in software source code: first, we embed code elements as points in a shared representation space according to structural dependencies between them; then, we represent documents as weighted point clouds of code elements in the representation space and reduce the distance between two documents to an earth mover's distance transportation problem. We define a document classification task in StackOverflow dataset to evaluate the effectiveness of our approach. The empirical evaluation results show that our approach outperforms several state-of-the-art approaches.
Zeqi Lin, Junfeng Zhao 0001, Yanzhen Zou
Internetware2
2017 Automatically Generating Task-Oriented API Learning Guide
abstract
Learning and reusing open source API libraries remain a time consuming process due to the documentation quality and the knowledge gap between API providers and users. Some researchers and API providers have found that the development tasks would narrow the knowledge gap and meet the needs of busy developers. To our knowledge, there is no existing work to generating task oriented API documents. In this paper, we propose an automatic approach to generating task oriented API learning guide. The guide is organized by a hierarchical task list. We integrate the natural language processing techniques with an evidence-based filtering pipeline in our approach. We also employ a graph-based clustering procedure to generate a three-layer task list. Furthermore, we define the normal form of the task phrases as the metadata in our approach. The approach has been implemented as a tool, APITasks. We used it to generate the API documents for four libraries. In an empirical study, we evaluate the accuracy and completeness of our approach with the manually created benchmarks. The results affirm the capability of our approach.
Zixiao Zhu, Chenyan Hua, Yanzhen Zou, Junfeng Zhao 0001
Internetware5
2017 Improving software text retrieval using conceptual knowledge in source code
abstract
A large software project usually has lots of various textual learning resources about its API, such as tutorials, mailing lists, user forums, etc. Text retrieval technology allows developers to search these API learning resources for related documents using free-text queries, but it suffers from the lexical gap between search queries and documents. In this paper, we propose a novel approach for improving the retrieval of API learning resources through leveraging software-specific conceptual knowledge in software source code. The basic idea behind this approach is that the semantic relatedness between queries and documents could be measured according to software-specific concepts involved in them, and software source code contains a large amount of software-specific conceptual knowledge. In detail, firstly we extract an API graph from software source code and use it as software-specific conceptual knowledge. Then we discover API entities involved in queries and documents, and infer semantic document relatedness through analyzing structural relationships between these API entities. We evaluate our approach in three popular open source software projects. Comparing to the state-of-the-art text retrieval approaches, our approach lead to at least 13.77% improvement with respect to mean average precision (MAP).
Zeqi Lin, Yanzhen Zou, Junfeng Zhao 0001
ASE3
2017 Intelligent Development Environment and Software Knowledge Graph
Zeqi Lin, Yanzhen Zou, Junfeng Zhao 0001, Xuandong Li, Jun Wei 0001, Hailong Sun 0001, Gang Yin
J. Comput. Sci. Technol.4
2016 Probabilistic-Mismatch Anomaly Detection: Do One's Medications Match with the Diagnoses
abstract
Anomaly detection in healthcare data like patient records is no trivial task. The anomalies in these datasets are often caused by mismatches between different types of feature, e.g., medications that do not match with the diagnoses. Existing anomaly detection methods do not perform well when detecting "mismatches" between multiple types of feature, especially when the feature space is high-dimensional and sparse. This paper introduces a novel anomaly detection paradigm: Probabilistic-Mismatch Anomaly Detection (PMAD), which detects mismatches between features by modeling a normal instance with a common latent probability distribution that governs the generation of all types of feature. Under this paradigm, the target of anomaly detection is to find instances with dissimilar latent distributions. We further propose Topical PMAD based on an extended Latent Dirichlet Allocation (LDA) model, which is able to capture the latent relationship between features in a high-dimensional space. Experiments on both synthetic data and real-world patient records show that Topical PMAD can effectively detect anomalies with mismatched features, and is highly robust against high-dimensional data as well as inaccurate model selection. The real-world anomalies detected on a patient record dataset show a promising application prospect.
Lingxiao Zhang, Xiang Li 0013, Haifeng Liu 0005, Jing Mei, Gang Hu 0001, Junfeng Zhao 0001, Yanzhen Zou, Guo Tong Xie
ICDM6
2015 A Situation-Aware and Interactive System for Assisting People Fill out Paper Forms
abstract
People often need help when filling out paper forms, because they do not fully understand the meaning of form fields. Commonly adopted solutions are referring to the form filling instructions or consulting other people. However, they are either inefficient or inconvenient. In this paper, we propose a situation-aware and interactive system, named Interact Form, to help people fill out paper forms. First, it provides an easy-to-use tool to build the knowledge offline about any given form. The knowledge includes the instruction and examples of each form field, and the constraints between fields. Then, when people are filling out a paper form, with the equipped video camera, the system can determine the field-grained position of the pen and provide text or audio instructions based on the user's form filling situations. When a pre-defined constraint is possibly going to be violated, the provided pen will send out an alert by vibration, and an explanation of the alert will also be given at the same time. We evaluate Interact Form with 60 paper form filling activities on 10 real-world paper forms, and the results show the usability of the system.
Jiangtao Wang 0001, Yasha Wang, Junfeng Zhao 0001
COMPSAC4
2015 Helping Campaign Initiators Create Mobile Crowd Sensing Apps: A Supporting Framework
abstract
Mobile Crowd Sensing (MCS) refers to the sensing paradigm in which mobile users with sensing and computing devices are tasked to collect and contribute data in order to enable various applications. The initiators of mobile crowd sensing campaigns are suitable to be ones for creating MCS applications (MCSA). However, the relatively high requirements in software development skills become the barrier for initiators to create MCSA. In this paper, we propose a supporting framework for the initiators who lack of software development skills to build MCSA in a quick and simple way. It enables initiators to build MCSA by just doing some simple settings, which totally eliminates the requirement of programming skills. Finally, we evaluate the effectiveness of the framework as far as the functionality and efficiency are concerned.
Jiangtao Wang 0001, Yasha Wang, Junfeng Zhao 0001
COMPSAC3
2015 Data-Driven Composition for Service-Oriented Situational Web Applications
abstract
The convergence of Services Computing and Web 2.0 gains a large space of opportunities to compose “situational” web applications from web-delivered services. However, the large number of services and the complexity of composition constraints make manual composition difficult to application developers, who might be non-professional programmers or even end-users. This paper presents a systematic data-driven approach to assisting situational application development. We first propose a technique to extract useful information from multiple sources to abstract service capabilities with a set tags. This supports intuitive expression of user's desired composition goals by simple queries, without having to know underlying technical details. A planning technique then exploits composition solutions which can constitute the desired goals, even with some potential new interesting composition opportunities. A browser-based tool facilitates visual and iterative refinement of composition solutions, to finally come up with the satisfying outputs. A series of experiments demonstrate the efficiency and effectiveness of our approach.
Xuanzhe Liu, Yun Ma 0002, Gang Huang 0001, Junfeng Zhao 0001, Hong Mei 0001, Yunxin Liu 0001
IEEE Trans. Serv. Comput.4
2014 A graph database based crowdsourcing infrastructure for modelling and searching code structure
abstract
Software reuse offers a solution to eliminate repeated work and improve efficiency and quality in the software development. In order to reuse existing software resources, software developers usually need to understand code structure of them. However, code structure is usually too complex to figure out. Therefore, it is helpful to demonstrate software developers the code structure they want to know. This paper presents a graph database based crowdsourcing infrastructure for modelling and searching code structure. In this paper, a graph based modelling paradigm of code structure is provided, which solves the problem that how code structure should be demonstrated. Software developers' search purposes are analyzed by natural language processing technique. A crowdsourcing mechanism is provided to integrate different code structure analysis algorithms for these different search purposes. Our work improves the efficiency of software reuse, and it is validated through an industrial case study.
Zeqi Lin, Junfeng Zhao 0001
Internetware2
2013 Mining Cohesive Domain Topics from Source Code
Junfeng Zhao 0001, Yanzhen Zou
ICSR4
2013 STERS: A System for Service Trustworthiness Evaluation and Recommendation based on the Trust Network (S)
Yasha Wang, Jiangtao Wang 0001, Yuxing Teng, Junfeng Zhao 0001
SEKE4
2012 ARIMA Model-Based Web Services Trustworthiness Evaluation and Prediction
Zhebang Hua, Junfeng Zhao 0001, Yanzhen Zou
ICSOC3
2012 Towards Automatic Tagging for Web Services
abstract
Tagging technique is widely used to annotate objects in Web 2.0 applications. Tags can support web service understanding, categorizing and discovering, which are important tasks in a service-oriented software system. However, most of existing web services' tags are annotated manually. Manual tagging is time-consuming. In this paper, we propose a novel approach to tag web services automatically. Our approach consists of two tagging strategies, tag enriching and tag extraction. In the first strategy, we cluster web services using WSDL documents, and then we enrich tags for a service with the tags of other services in the same cluster. Considering our approach may not generate enough tags by tag enriching, we also extract tags from WSDL documents and related descriptions in the second step. To validate the effectiveness of our approach, a series of experiments are carried out based on web-scale web services. The experimental results show that our tagging method is effective, ensuring the number and quality of generated tags. We also show how to use tagging results to improve the performance of a web service search engine, which can prove that our work in this paper is useful and meaningful.
Junfeng Zhao 0001, Yanzhen Zou, Lingshuang Shao
ICWS4
2012 Propositional Logic-Based and Evidence-Rich Trustworthiness Evaluation for Web Services
abstract
With more and more Web services available on the Internet, users reuse these services in their own applications. Because Web services are delivered by third parties and hosted on remote servers, it becomes a big problem to determine the trustworthiness of the Web services. Many Web services trustworthiness evaluation approaches have been proposed, however, the trustworthy evidences used in these approaches are limited and the methods proposed lack customizability and extensibility, which makes them difficult to apply. In this paper, we propose a lightweight Propositional Logic-based Web services trustworthiness evaluation method, which is customizable, extensible, and easy to apply in reality. First we collect comprehensive trustworthy evidences including both objective trustworthy evidences (e.g. QoS) and subjective evidences (e.g. reputation) from the Internet. Then we propose a Propositional Logic-based Web services trustworthiness evaluation model, which is customizable and extensible, to capture users' trustworthiness requirements. Finally, the trustworthiness of all Web services are evaluated and returned to the users via a Web services search engine. To validate the effectiveness of our approach, two experiments are conducted on a large-scale real-world dataset. The experimental results show that our method is easy to use and can effectively evaluate Web services trustworthiness, which helps users to reuse Web services.
Junfeng Zhao 0001
SERVICES3
2011 CoWS: An Internet-Enriched and Quality-Aware Web Services Search Engine
abstract
With more and more Web services available on the Internet, many approaches have been proposed to help users discover and select desired services. However, existing approaches heavily rely on the information in UDDI repositories or WSDL files, which is quite limited in fact. The limitation of information weakens the effectiveness of existing approaches. In this paper, we present a novel Web services search engine named CoWS, which enriches Web services information using the information captured from the Internet to provide quality-aware Web services search. The information captured can be classified into two groups: functional descriptions and subjective feedbacks. We use the functional descriptions to enrich descriptions of Web services and the subjective feedbacks to calculate Web services' reputation. CoWS first ranks the services according to their functional similarities to a user's query, which are calculated using both descriptions in WSDL files and the enriched descriptions, and then refines and re-ranks the services with both objective quality constraints (QoS) and subjective quality constraints (reputation). The experiments on a large-scale dataset (including 31,129 Web services) show that CoWS can improve the effectiveness of both Web services discovery and selection comparing with existing approaches.
Junfeng Zhao 0001, Sibo Cai
ICWS2
2010 TSRR: A Software Resource Repository for Trustworthiness Resource Management and Reuse
Junfeng Zhao 0001, Yasha Wang
SEKE1
2009 User-Perceived Service Availability: A Metric and an Estimation Approach
abstract
Web-service-related techniques have become popular to improve system integration and interaction. In distributed and dynamic environment, Web services' availability has been regarded as one of the key properties for (critical) service-oriented applications. Quality of Service (QoS), including availability, has been regarded by IEEE as a user-perceived property. However, based on our investigation of monitoring invocation records of real Web services, existing availability metrics, which were proposed in traditional domains, have not addressed the "user-perceived'' characteristics. Based on analyzing the limitations of the existing availability metrics, we propose a status-based user-perceived service availability metric and a corresponding estimation approach. Experiments on monitoring and analyzing the invocation records of real services demonstrate that the new metric and the corresponding estimation approach could lead to a feasible estimation on Web services' availability from the user side.
Lingshuang Shao, Junfeng Zhao 0001, Tao Xie 0001, Lu Zhang 0023, Hong Mei 0001
ICWS2
2009 User Perceived Response-time Optimization Method for Composite Web Services
Junfeng Zhao 0001, Yasha Wang
SEKE1
2008 Dynamic Availability Estimation for Service Selection Based on Status Identification
abstract
With the popularity of service-oriented computing, how to construct highly available service-oriented applications is becoming a hot topic in both the research and industry communities. As a fundamental problem in dynamic service selection, availability estimation is challenging because of the dynamic nature of Web services. To grasp the dynamic nature of Web services, we set up an experimental environment for collecting runtime information of Web services. Based on the collected runtime information, we identify several characteristics of service failures and successes, and further define three typical service runtime statuses. Based on these statuses, we propose a novel approach to dynamic availability estimation, which is called status identification based availability estimation for service selection (SIBE). To evaluate our approach, we compare SIBE with other approaches in an experiment of dynamic service selection on the Internet. Experimental results show that SIBE can efficiently improve the success rate of selecting available services.
Lingshuang Shao, Lu Zhang 0023, Tao Xie 0001, Junfeng Zhao 0001, Hong Mei 0001
ICWS4
2007 Personalized QoS Prediction forWeb Services via Collaborative Filtering
abstract
Many researchers propose that, not only functional but also non-functional properties, also known as quality of service (QoS), should be taken into consideration when consumers select services. Consumers need to make prediction on quality of unused web services before selecting. Usually, this prediction is based on other consumers' experiences. Being aware of different QoS experiences of consumers, this paper proposes a collaborative filtering based approach to making similarity mining and prediction from consumers' experiences. Experimental results demonstrate that this approach can make significant improvement on the effectiveness of QoS prediction for web services.
Lingshuang Shao, Jing Zhang 0005, Junfeng Zhao 0001, Hong Mei 0001
ICWS4
2007 Using Scenario Oriented Response-time Management for Composite Web Services
abstract
In order to make a composite web service meet user's response-time requirement, the proper instance of each member web service should be selected and bound. In the literature, all the member web services of a composite web service are treated equally during the response-time management, but without considering their different capabilities of effecting the user's satisfaction. However, in a certain using scenario, some member services are more sensitive than others, that is to say, in a given composite web service, if these sensitive services delayed, the decline of users' satisfaction is remarkably greater than that when other services delayed. In this article, a using scenario oriented response-time management method is proposed to reduce the delaying risk of the time sensitive web services, and thus to improve the user's satisfaction of a composite web service. Our experiments validated the efficiency of the proposed method.
Yasha Wang, Junfeng Zhao 0001
ICWS2
2004 Towards an Optimization-Based Method for Consolidating Domain Variabilities in Domain-Specific Web Services Composition
Junfeng Zhao 0001, Lu Zhang 0023, Yasha Wang
ICTAC1