Jian Wu 0001

dblp:96/2744-1 · DBLP profile ↗
← Back
185ranked-venue papers
7as first author
105since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 56 since 2021Applied, interdisciplinary, general and emerging computing · 55 · 2 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 35 since 2021Software engineering, systems software and programming languages · 27 · 2 since 2021Databases, data management, data science and information retrieval · 27 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorComputer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
abstract
Qiyuan Chen, Hongsen Huang, Jiahe Chen, Qian Shao, Jintai Chen, Hongxia Xu, Renjie Hua, Ren Chuan, Jian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiyuan Chen 0003, Hongsen Huang, Qian Shao, Jintai Chen, Renjie Hua, Ren Chuan, Jian Wu 0001
ACL (1)9
2026 MT³: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation
abstract
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Zhijie Zhou, Wenxuan Huang, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaopeng Feng, Yupu Liang, Shaosheng Cao, Jiayuan Su, Jiahan Ren, Wenxuan Huang 0001, Jian Wu 0001, Zuozhu Liu
ACL (1)8
2026 Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering
abstract
Songtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, ZhiBo Yang, Jimeng Sun, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Songtao Jiang, Ruizhe Chen, Yan Zhang 0004, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, Jimeng Sun 0001, Jian Wu 0001, Zuozhu Liu
ACL (1)11
2026 Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation
abstract
Large Language Models enhanced with Retrieval Augmented Generation show strong potential in knowledge intensive tasks.However, they often encounter knowledge conflicts, where retrieved information contradicts the model's internal knowledge or exhibits internal inconsistencies.Existing methods force models into a binary choice between context and memory, leading to unreliable predictions.We argue that a more principled approach is to embrace contradictions as opportunities for deeper reasoning.To this end, we introduce Debate-of-Thoughts (DoT), a framework that transforms conflict resolution into an active deliberation process.DoT guides a single model through three phases: 1) hypothesis generation, which forms competing perspectives; 2) internal debate, where the model acts as both a proponent and a critic to stress test each view; and 3) adjudication, where the model acts as a judge to evaluate arguments based on evidence and logical consistency.We implement DoT via two complementary strategies: inference time prompt chaining and supervised fine tuning.Experiments across multiple conflict benchmarks show that DoT consistently outperforms state-of-the-art methods, while generating transparent debate transcripts that explain its decisions.By improving both accuracy and interpretability under knowledge conflicts, DoT establishes a more reliable paradigm for retrieval augmented generation systems. 1
Guocong Li, Qirui Hu, Guofeng Zhang 0001, Jian Wu 0001
ACL (1)5
2026 Trend-aware structure relearning framework for water quality prediction with coupled noise governance
Shuo Tong, Yuyang Xu, Jianming Sun, Fuzhen Zhuang, Jian Wu 0001, Guangdi Chen, Haochao Ying
Expert Syst. Appl.5
2026 LLM-SDaT: A knowledge-informed LLM framework for syndrome differentiation in TCM
Bingtao Guan, Shangde Gao, Dawei Zheng, Haoxiang Xia, Jian Wu 0001
Neural Networks6
2026 Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal Representation
abstract
Despite the rapid advancements of electrocardiogram (ECG) signal diagnosis and analysis methods through deep learning, two major hurdles still limit their clinical adoption: the lack of versatility in processing ECG signals with diverse configurations, and the inadequate detection of risk signals due to sample imbalances. Addressing these challenges, we introduceVersAtile andRisk-Sensitive cardiac diagnosis (VARS), an innovative approach that employs a graph-based representation to uniformly model heterogeneous ECG signals. VARS stands out by transforming ECG signals into versatile graph structures that capture critical diagnostic features, irrespective of signal diversity in the lead count, sampling frequency, and duration. This graph-centric formulation also enhances diagnostic sensitivity, enabling precise localization and identification of abnormal ECG patterns that often elude standard analysis methods. To facilitate representation transformation, our approach integrates denoising reconstruction with contrastive learning to preserve raw ECG information while highlighting pathognomonic patterns. We rigorously evaluate the efficacy of VARS on three distinct ECG datasets, encompassing a range of structural variations. The results demonstrate that VARS not only consistently surpasses existing state-of-the-art models across all these datasets but also exhibits substantial improvement in identifying risk signals. Additionally, VARS offers interpretability by pinpointing the exact waveforms that lead to specific model outputs, thereby assisting clinicians in making informed decisions. These findings suggest that our VARS will likely emerge as an invaluable tool for comprehensive cardiac health assessment.
Yuyang Xu, Renjun Hu, Fanqi Shen, Hanyun Jiang, Jun Wang 0072, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001, Haochao Ying
IEEE Trans. Big Data9
2026 MambaCapsule: Toward Transparent Cardiac Disease Diagnosis With Electrocardiography Using Mamba Capsule Network
abstract
Cardiac arrhythmia, a condition characterized by irregular heartbeats, often serves as an early indication of various heart ailments. With the advent of deep learning, numerous innovative models have been introduced for diagnosing arrhythmias using electrocardiogram (ECG) signals. However, recent studies solely focus on the performance of models, neglecting the interpretation of their results. This leads to a considerable lack of transparency, posing a significant risk in the actual diagnostic process. To solve this problem, this article introduces MambaCapsule, a deep neural network for ECG arrhythmias classification, which increases the explainability of the model while enhancing the accuracy. Our model utilizes Mamba for feature extraction and Capsule networks for prediction, providing not only a confidence score but also signal features. Akin to the processing mechanism of human brain, the model learns signal features and their relationship between them by reconstructing ECG signals in the predicted selection. The model evaluation was conducted on Massachusetts Institute of Technology - Beth Israel Hospital (MIT-BIH) and Physikalisch-Technische Bundesanstalt Diagnostic ECG Database (PTB) datasets, following the AAMI standard. MambaCapsule has achieved a total accuracy of 99.54% and 99.59% on the test sets, respectively. These results demonstrate the promising performance under the standard test protocol.
Yinlong Xu 0002, Zitai Kong, Yingzhou Lu, Jian Wu 0001
IEEE Trans. Comput. Soc. Syst.7
2026 Decouple, Reorganize, and Fuse: A Multimodal Framework for Cancer Survival Prediction
abstract
Cancer survival analysis commonly integrates information across diverse medical modalities to make survival-time predictions. Existing methods primarily focus on extracting different decoupled features of modalities and performing fusion operations such as concatenation, attention, and Mixture-of-Experts (MoE)-based fusion. However, these methods still face two key challenges: 1) fixed fusion schemes (concatenation and attention) can lead to model over-reliance on predefined feature combinations, limiting the dynamic fusion of decoupled features; and 2) in MoE-based fusion methods, each expert network handles separate decoupled features, which limits information interaction among the decoupled features. To address these challenges, we propose a novel Decoupling-Reorganization-Fusion framework (DeReF), which devises a random feature reorganization strategy between modalities decoupling and dynamic MoE fusion modules. Its advantages are: 1) it increases the diversity of feature combinations and granularity, enhancing the generalization ability of the subsequent expert networks; and 2) it overcomes the problem of information closure and helps expert networks better capture information among decoupled features. Additionally, we incorporate a regional cross-attention network within the modality decoupling module to improve the representation quality of decoupled features. Extensive experimental results on our in-house Liver Cancer (LC) and three widely used public datasets from The Cancer Genome Atlas (TCGA) confirm the effectiveness of our proposed method. Codes are available at https://github.com/ZJUMAI/DeReF.
Haochao Ying, Yuyang Xu, Qibo Qiu, Danny Ziyi Chen, Ying Sun 0015, Jian Wu 0001
IEEE Trans. Medical Imaging8
2026 MoE2: Optimizing Collaborative Inference for Edge Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for diverse emerging applications, as it enables greater cost-effectiveness and reduced latency. In this work, we introduceMixture-of-Edge-Experts (MoE2), a novel collaborative inference framework for edge LLMs. We formulate a joint gating and expert selection problem to optimize inference performance under energy and latency constraints. Unlike conventional MoE problems, LLM expert selection becomes significantly more challenging due to the combinatorial nature and the heterogeneity of edge LLMs across various attributes. To this end, we propose a two-level expert selection mechanism through which we uncover an optimality-preserving property of gating parameters across expert selections. This property enables the decomposition of the training and selection processes, significantly reducing complexity. Furthermore, we leverage the objective’s monotonicity and design a discrete monotonic optimization algorithm for optimal expert selection. We implement edge servers with NVIDIA Jetson AGX Orins and NVIDIA RTX 4090 GPUs, and perform extensive experiments. Our results validate the performance improvements for various LLM models and show that our MoE2 method can achieve optimal trade-offs among different delay and energy budgets, and outperforms baselines under various system resource constraints. We further demonstrate its strong robustness in dynamic, non-stationary environments and its effectiveness in achieving load balancing.
Lyudong Jin, Shurong Wang, Howard H. Yang, Jian Wu 0001, Meng Zhang 0013
IEEE Trans. Netw.6
2025 Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer
abstract
Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery.
Mingze Yin, Hanjing Zhou, Yiheng Zhu 0002, Jialu Wu, Wei Wu 0045, Kun Fu 0002, Zheng Wang 0027, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
AAAI11
2025 ProtCLIP: Function-Informed Protein Multi-Modal Learning
abstract
Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.
Hanjing Zhou, Mingze Yin, Wei Wu 0045, Kun Fu 0002, Jintai Chen, Jian Wu 0001, Zheng Wang 0027
AAAI7
2025 LLMs Can Simulate Standardized Patients via Agent Coevolution
abstract
Zhuoyun Du, LujieZheng LujieZheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, Haochao Ying. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhuoyun Du, Lujie Zheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun 0015, Wei Chen 0001, Jian Wu 0001, Haolei Cai, Haochao Ying
ACL (1)8
2025 M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation
abstract
Recent advancements in large language models (LLMs) have given rise to the LLM-as-a-judge paradigm, showcasing their potential to deliver human-like judgments. However, in the field of machine translation (MT) evaluation, current LLM-as-a-judge methods fall short of learned automatic metrics. In this paper, we propose Multidimensional Multi-Agent Debate (M-MAD), a systematic LLM-based multi-agent framework for advanced LLM-as-a-judge MT evaluation. Our findings demonstrate that M-MAD achieves significant advancements by (1) decoupling heuristic MQM criteria into distinct evaluation dimensions for fine-grained assessments; (2) employing multi-agent debates to harness the collaborative reasoning capabilities of LLMs; (3) synthesizing dimension-specific results into a final evaluation judgment to ensure robust and reliable outcomes. Comprehensive experiments show that M-MAD not only outperforms all existing LLM-as-a-judge methods but also competes with state-of-the-art reference-based automatic metrics, even when powered by a suboptimal model like GPT-4o mini. Detailed ablations and analysis highlight the superiority of our framework design, offering a fresh perspective for LLM-as-a-judge paradigm. Our code and data are publicly available at https://github.com/SU-JIAYUAN/M-MAD.
Zhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren, Yan Zhang 0004, Jian Wu 0001, Hongwei Wang 0001, Zuozhu Liu
ACL (1)6
2025 HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
abstract
Songtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang, Yangyang Wu, Yang Feng, Jian Wu, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Songtao Jiang, Yan Zhang 0004, Yeying Jin, Zhihang Tang, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
ACL (1)7
2025 Scalable Autoregressive Monocular Depth Estimation
abstract
This paper proposes a new autoregressive model as an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise causal mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-of-the-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities. Project page: https://depth-ar.github.io/.
Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001
CVPR8
2025 Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
abstract
Qiyuan Chen, Hongsen Huang, Qian Shao, Jiahe Chen, Jintai Chen, Hongxia Xu, Renjie Hua, Ren Chuan, Jian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Qiyuan Chen 0003, Hongsen Huang, Qian Shao, Jintai Chen, Renjie Hua, Ren Chuan, Jian Wu 0001
EMNLP9
2025 OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
Shuo Tong, Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001
ICCV10
2025 Small Models are LLM Knowledge Triggers for Medical Tabular Prediction
abstract
Recent development in large language models (LLMs) has demonstrated impressive domain proficiency on unstructured textual or multi-modal tasks. However, despite with intrinsic world knowledge, their application on structured tabular data prediction still lags behind, primarily due to the numerical insensitivity and modality discrepancy that brings a gap between LLM reasoning and statistical tabular learning. Unlike textual or vision data (e.g., electronic clinical notes or medical imaging data), tabular data is often presented in heterogeneous numerical values (e.g., CBC reports). This ubiquitous data format requires intensive expert annotation, and its numerical nature limits LLMs' capability to effectively transfer untapped domain expertise. In this paper, we propose SERSAL, a general self-prompting method by synergy learning with small models to enhance LLM tabular prediction in an unsupervised manner. Specifically, SERSAL utilizes the LLM's prior outcomes as original soft noisy annotations, which are dynamically leveraged to teach a better small student model. Reversely, the outcomes from the trained small model are used to teach the LLM to further refine its real capability. This process can be repeatedly applied to gradually distill refined knowledge for continuous progress. Comprehensive experiments on widely used medical domain tabular datasets show that, without access to gold labels, applying SERSAL to OpenAI GPT reasoning process attains substantial improvement compared to linguistic prompting methods, which serves as an orthogonal direction for tabular LLM, and increasing prompting bonus is observed as more powerful LLMs appear. Codes are available at https://github.com/jyansir/sersal.
Jiahuan Yan, Jintai Chen, Chaowen Hu, Bo Zheng 0011, Yaojun Hu, Jimeng Sun 0001, Jian Wu 0001
ICLR7
2025 Group-On: Boosting One-Shot Segmentation with Supportive Query
abstract
One-shot semantic segmentation aims to segment query images given only ONE annotated support image of the same class. This task is challenging because target objects in the support and query images can be largely different in appearance and pose (i.e., intra-class variation). Prior works suggested that incorporating more annotated support images in few-shot settings boosts performances but increases costs due to additional manual labeling. In this paper, we propose a novel and effective approach for ONE-shot semantic segmentation, called Group-On, which packs multiple query images in batches for the benefit of mutual knowledge support within the same category. Specifically, after coarse segmentation masks of the batch of queries are predicted, query-mask pairs act as pseudo support data to enhance mask predictions mutually. To effectively steer such process, we construct an innovative MoME module, where a flexible number of mask experts are guided by a scene-driven router and work together to make comprehensive decisions, fully promoting mutual benefits of queries. Comprehensive experiments on three standard benchmarks show that, in the ONE-shot setting, Group-On significantly outperforms previous works by considerable margins. With only one annotated support image, Group-On can be even competitive with the counterparts using 5 annotated images.
Hanjing Zhou, Mingze Yin, Danny Ziyi Chen, Jian Wu 0001, Jintai Chen
ICME4
2025 Dual-level Fuzzy Learning with Patch Guidance for Image Ordinal Regression
abstract
Ordinal regression bridges regression and classification by assigning objects to ordered classes. While human experts rely on discriminative patch-level features for decisions, current approaches are limited by the availability of only image-level ordinal labels, overlooking fine-grained patch-level characteristics. In this paper, we propose a Dual-level Fuzzy Learning with Patch Guidance framework, named DFPG that learns precise feature-based grading boundaries from ambiguous ordinal labels, with patch-level supervision. Specifically, we propose patch-labeling and filtering strategies to enable the model to focus on patch-level features exclusively with only image-level ordinal labels available. We further design a dual-level fuzzy learning module, which leverages fuzzy logic to quantitatively capture and handle label ambiguity from both patch-wise and channel-wise perspectives. Extensive experiments on various image ordinal regression datasets demonstrate the superiority of our proposed method, further confirming its ability in distinguishing samples from difficult-to-classify categories. The code is available at https://github.com/ZJUMAI/DFPG-ord.
Chunlai Dong, Haochao Ying, Qibo Qiu, Danny Ziyi Chen, Jian Wu 0001
IJCAI6
2025 Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
abstract
Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models.
Songtao Jiang, Yan Zhang 0004, Ruizhe Chen, Tianxiang Hu, Yeying Jin, Qinglin He, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
IJCAI8
2025 Matrix Factorization with Dynamic Multi-view Clustering for Recommender System
abstract
Matrix factorization (MF), a cornerstone of recommender systems, decomposes user-item interaction matrices into latent representations. Traditional MF approaches, however, employ a two-stage, non-end-to-end paradigm, sequentially performing recommendation and clustering, resulting in prohibitive computational costs for large-scale applications like e-commerce and IoT, where billions of users interact with trillions of items. To address this, we propose Matrix Factorization with Dynamic Multi-view Clustering (MFDMC), a unified framework that balances efficient end-to-end training with comprehensive utilization of web-scale data and enhances interpretability. MFDMC leverages dynamic multi-view clustering to learn user and item representations, adaptively pruning poorly formed clusters. Each entity's representation is modeled as a weighted projection of robust clusters, capturing its diverse roles across views. This design maximizes representation space utilization, improves interpretability, and ensures resilience for downstream tasks. Extensive experiments demonstrate MFDMC's superior performance in recommender systems and other representation learning domains, such as computer vision, highlighting its scalability and versatility.
Shangde Gao, Ke Liu 0012, Yichao Fu, Jian Wu 0001
IJCNN5
2025 Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
Songtao Jiang, Sibo Song, Yan Zhang 0004, Yeying Jin, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
MICCAI (11)7
2025 Uncertainty-Aware Multi-expert Knowledge Distillation for Imbalanced Disease Grading
Shuo Tong, Shangde Gao, Ke Liu 0012, Haochao Ying, Jian Wu 0001
MICCAI (13)7
2025 V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Shujian Gao, Zhihang Tang, Xiaotang Gai, Jian Wu 0001, Zuozhu Liu
MICCAI (5)7
2025 Fair-MoE: Medical Fairness-Oriented Mixture of Experts in Vision-Language Models
Peiran Wang, Linjie Tong, Jian Wu 0001, Zuozhu Liu
MICCAI (5)3
2025 Poster: MoE2: Optimizing Collaborative Inference for Edge Large Language Models
abstract
Large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks. Exploiting the heterogeneous capabilities of edge LLMs is crucial for emerging applications, enabling greater cost-effectiveness and reduced latency. In this work, we introduce Mixture-of-Edge-Experts (MoE2), a collaborative inference framework for edge LLMs. We formulate the joint gating and expert selection problem to optimize inference under energy and latency constraints. Unlike conventional MoE problems, expert selection here is more challenging due to the combinatorial nature and heterogeneity of edge LLMs. To address this, we propose a two-level expert selection mechanism and uncover an optimality-preserving property of gating parameters that decouples training and selection, reducing complexity. We further leverage the objective's monotonicity and design a discrete monotonic optimization algorithm. Implemented on Jetson Orin and RTX 4090 platforms, MoE2 achieves optimal trade-offs across delay and energy budgets, outperforming baselines under various resource constraints.
Lyudong Jin, Shurong Wang, Howard H. Yang, Jian Wu 0001, Meng Zhang 0013
MobiCom6
2025 3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks
abstract
Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset designed to advance 3D Med-VQA using radiology CT scans. The 3D-RAD dataset encompasses six diverse VQA tasks: anomaly detection, image observation, medical computation, existence detection, static temporal diagnosis, and longitudinal temporal diagnosis. It supports both open- and closed-ended questions while introducing complex reasoning challenges, including computational tasks and multi-stage temporal analysis, to enable comprehensive benchmarking. Extensive evaluations demonstrate that existing vision-language models (VLMs), especially medical VLMs exhibit limited generalization, particularly in multi-temporal tasks, underscoring the challenges of real-world 3D diagnostic reasoning. To drive future advancements, we release a high-quality training set 3D-RAD-T of 136,195 expert-aligned samples, showing that fine-tuning on this dataset could significantly enhance model performance. Our dataset and code, aiming to catalyze multimodal medical AI research and establish a robust foundation for 3D medical visual understanding, are publicly available.
Xiaotang Gai, Zijie Meng, Jian Wu 0001, Zuozhu Liu
NeurIPS5
2025 Advancing interpretable cardiac disease diagnosis via a transformer-convolutional hybrid network on electrocardiograms
abstract
Manual heart disease diagnosis with the electrocardiogram (ECG) is intractable due to the intertwined signal features and lengthy diagnosis procedure, especially for the 24-hour dynamic ECG signals. Consequently, even experienced cardiologists may face difficulty in producing all accurate ECG reports. In recent years, Artificial Intelligence (AI), particularly neural network-based automatic ECG diagnosis methods have exhibited promising performance, suggesting a potential alternative to the labor-intensive examination conducted by cardiologists. However, many existing approaches failed to adequately consider the temporal and channel dimensions when assembling features and ignored interpretability. And clinical theory underscores the necessity of prolonged signal observations for diagnosing certain ECG conditions such as tachycardia. Moreover, specific heart diseases manifest primarily through distinct ECG leads represented as channels. In response to these challenges, this paper introduces a novel neural network architecture for ECG classification (diagnosis). The proposed model incorporates Lead Fusing blocks, transformer-XL (meaning extra long) encoder-based Encoder modules, and hierarchical temporal attentions. Importantly, this classifier operates directly on raw ECG time-series signals rather than cardiac cycles. Signal integration begins with the Lead Fusing blocks, followed by the Encoder modules and hierarchical temporal attentions, enabling the extraction of long-dependent features. Furthermore, existing convolution-based methods have been argued to compromise interpretability, whereas the proposed neural network provides improved clarity in this regard. Experimental evaluations on a comprehensive public dataset confirm the superiority of the proposed classifier over state-of-the-art methods. Moreover, a visualization method was employed to generate a location map that demonstrates the areas of the signal emphasized by the model, thereby enhancing interpretability. • Our model extracts long-dependent features of ECG signals based on the Transformer-XL encoder. • The proposed network offers the improved interpretability. • Our classifier achieves superior performance over other state-of-the-art methods.
Yinlong Xu 0002, Siyu Long, Yisen Huang, Yingzhou Lu, Yingxuan Huang, Jian Wu 0001, Honghao Gao
Eng. Appl. Artif. Intell.10
2025 Cross-center Model Adaptive Tooth segmentation
Ruizhe Chen, Jianfei Yang 0001, Huimin Xiong, Ruiling Xu, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
Medical Image Anal.6
2025 Hulk: A Universal Knowledge Translator for Human-Centric Tasks
abstract
Human-centric perception tasks, e.g., pedestrian detection, skeleton-based action recognition, and pose estimation, have wide industrial applications, such as metaverse and sports analysis. There is a recent surge to develop human-centric foundation models that can benefit a broad range of human-centric perception tasks. While many human-centric foundation models have achieved success, they did not explore 3D and vision-language tasks for human-centric and required task-specific finetuning. These limitations restrict their application to more downstream tasks and situations. To tackle these problems, we present Hulk, the first multimodal human-centric generalist model, capable of addressing 2D vision, 3D vision, skeleton-based, and vision-language tasks without task-specific finetuning. The key to achieving this is condensing various task-specific heads into two general heads, one for discrete representations, e.g., languages, and the other for continuous representations, e.g., location coordinates. The outputs of two heads can be further stacked into four distinct input and output modalities. This uniform representation enables Hulk to treat diverse human-centric tasks as modality translation, integrating knowledge across a wide range of tasks. Comprehensive evaluations of Hulk on 12 benchmarks covering 8 human-centric tasks demonstrate the superiority of our proposed method, achieving state-of-the-art performance in 11 benchmarks.
Yizhou Wang 0007, Weizhen He, Xun Guo 0001, Feng Zhu 0006, Lei Bai 0001, Rui Zhao 0001, Jian Wu 0001, Tong He 0001, Wanli Ouyang, Shixiang Tang
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 A Progressively-Passing-Then-Disentangling Approach to Recipe Recommendation
abstract
The increasing popularity of online food blogs and food ordering services has made personalized recipe recommendation a vital aspect of our emotional well-being. However, existing solutions, mainly based on graph neural networks, still face significant challenges, such as (a) focusing on exploiting the user-recipe interactions while neglecting other crucial pairwise and high-order relationships, and (b) failing to explicitly distinguish the distinct factors, e.g., hedonic and healthy, that influence recipe selection. To address these issues, we propose a progressively-passing-then-disentangling approach named P2D. Our approach utilizes a three-stage progressive message-passing mechanism for better representation learning. Specifically, we incorporate the extra pairwise relationships between recipes and nutrients, ingredients, and visual contents to create fine-grained and multimodal recipe representations. We next refine these representations via message passing between high-order recipe relationships to learn people's shared food preferences. Based on them, we could derive comprehensive user representations, which are subsequently transformed into disentangled forms that correspond to various decision factors through contrastive and mutual information regularization. Experimental results demonstrate both the superiority and the rationality of our method: (a) P2D outperforms the state-of-the-art recipe recommendation methods by a large margin under various metrics, (b) ablation studies confirm the positive impact of each of its components, and (c) our visualization analysis empirically supports the advantage of explicitly disentangling decision factors.
Chunlai Dong, Haochao Ying, Renjun Hu, Yuyang Xu, Jintai Chen, Fuzhen Zhuang, Jian Wu 0001
IEEE Trans. Multim.7
2024 Arithmetic Feature Interaction Is Necessary for Deep Tabular Learning
abstract
Until recently, the question of the effective inductive bias of deep models on tabular data has remained unanswered. This paper investigates the hypothesis that arithmetic feature interaction is necessary for deep tabular learning. To test this point, we create a synthetic tabular dataset with a mild feature interaction assumption and examine a modified transformer architecture enabling arithmetical feature interactions, referred to as AMFormer. Results show that AMFormer outperforms strong counterparts in fine-grained tabular data modeling, data efficiency in training, and generalization. This is attributed to its parallel additive and multiplicative attention operators and prompt-based optimization, which facilitate the separation of tabular samples in an extended space with arithmetically-engineered features. Our extensive experiments on real-world data also validate the consistent effectiveness, efficiency, and rationale of AMFormer, suggesting it has established a strong inductive bias for deep learning on tabular data. Code is available at https://github.com/aigc-apps/AMFormer.
Renjun Hu, Haochao Ying, Jian Wu 0001, Wei Lin 0016
AAAI5
2024 Multi-rater Prompting for Ambiguous Medical Image Segmentation
abstract
Multi-rater annotations commonly occur when medical images are independently annotated by multiple experts (raters). In this paper, we tackle two challenges arisen in multi-rater annotations for medical image segmentation (called ambiguous medical image segmentation): (1) How to train a deep learning model when a group of raters produces a set of diverse but plausible annotations, and (2) how to fine-tune the model efficiently when computation resources are not available for retraining the entire model on a different dataset domain. We propose a multi-rater prompt-based approach to address these two challenges altogether. Specifically, we introduce a series of rater-aware prompts that can be plugged into the U-Net model for uncertainty estimation to handle multi-annotation cases. During the prompt-based fine-tuning process, only 0.3% of learnable parameters are required to be updated comparing to training the entire model. Further, in order to integrate expert consensus and disagreement, we explore different multi-rater incorporation strategies and design a mix-training strategy for comprehensive insight learning. Extensive experiments verify the effectiveness of our new approach for ambiguous medical image segmentation on two public datasets while alleviating the heavy burden of model re-training. Code will be made available.
Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
BIBM6
2024 MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual Invariant
abstract
Medical generative models, acknowledged for their high-quality sample generation ability, have accelerated the fast growth of medical applications. However, recent works concentrate on separate medical generation models for dis-tinct medical tasks and are restricted to inadequate medi-cal multimodal knowledge, constraining medical compre-hensive diagnosis. In this paper, we propose MedM2G, a Medical Multi-Modal Generative framework, with the key innovation to align, extract, and generate medical multimodal within a unified model. Extending beyond single or two medical modalities, we efficiently align medical multimodal through the central alignment approach in the unified space. Significantly, our framework extracts valuable clini-cal knowledge by preserving the medical visual invariant of each imaging modal, thereby enhancing specific medical information for multimodal generation. By conditioning the adaptive cross-guided parameters into the multi-flow diffusion framework, our model promotes flexible interactions among medical multimodalfor generation. MedM2G is the first medical generative model that unifies medical generation tasks of text-to-image, image-to-text, and unified generation of medical modalities (CT, MRI, X-ray). It performs 5 medical generation tasks across 10 datasets, consistently outperforming various state-of-the-art works.
Chenlu Zhan, Gaoang Wang, Hongwei Wang 0001, Jian Wu 0001
CVPR5
2024 DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
Yizhou Wang 0007, Shixiang Tang, Tong He 0001, Wanli Ouyang, Philip Torr 0001, Jian Wu 0001
ECCV (32)8
2024 Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications
abstract
Recently, large language models (LLMs) have achieved tremendous breakthroughs in the field of NLP, but still lack understanding of their internal neuron activities when processing different languages. We designed a method to convert dense LLMs into fine-grained MoE architectures, and then visually studied the multilingual activation patterns of LLMs through expert activation frequency heatmaps. Through comprehensive experiments on different model families, different model sizes, and different variants, we analyzed the similarities and differences in the internal neuron activation patterns of LLMs when processing different languages. Specifically, we investigated the distribution of high-frequency activated experts, multilingual shared experts, whether multilingual activation patterns are related to language families, and the impact of instruction tuning on activation patterns. We further explored leveraging the discovered differences in expert activation frequencies to guide sparse activation and pruning. Experimental results demonstrated that our method significantly outperformed random expert pruning and even exceeded the performance of unpruned models in some languages. Additionally, we found that configuring different pruning rates for different layers based on activation level differences could achieve better results. Our findings reveal the multilingual processing mechanisms within LLMs and utilize these insights to offer new perspectives for applications such as sparse activation and model pruning.
Weize Liu, Yinlong Xu 0002, Jintai Chen, Xuming Hu, Jian Wu 0001
EMNLP6
2024 FedLoGe: Joint Local and Generic Federated Learning under Long-tailed Data
abstract
Federated Long-Tailed Learning (Fed-LT), a paradigm wherein data collected from decentralized local clients manifests a globally prevalent long-tailed distribution, has garnered considerable attention in recent times. In the context of Fed-LT, existing works have predominantly centered on addressing the data imbalance issue to enhance the efficacy of the generic global model while neglecting the performance at the local level. In contrast, conventional Personalized Federated Learning (pFL) techniques are primarily devised to optimize personalized local models under the presumption of a balanced global data distribution. This paper introduces an approach termed Federated Local and Generic Model Training in Fed-LT (FedLoGe), which enhances both local and generic model performance through the integration of representation learning and classifier alignment within a neural collapse framework. Our investigation reveals the feasibility of employing a shared backbone as a foundational framework for capturing overarching global trends, while concurrently employing individualized classifiers to encapsulate distinct refinements stemming from each client’s local features. Building upon this discovery, we establish the Static Sparse Equiangular Tight Frame Classifier (SSE-C), inspired by neural collapse principles that naturally prune extraneous noisy features and foster the acquisition of potent data representations. Furthermore, leveraging insights from imbalance neural collapse's classifier norm patterns, we develop Global and Local Adaptive Feature Realignment (GLA-FR) via an auxiliary global classifier and personalized Euclidean norm transfer to align global features with client preferences. Extensive experimental results on CIFAR-10/100-LT, ImageNet, and iNaturalist demonstrate the advantage of our method over state-of-the-art pFL and Fed-LT approaches.
Zikai Xiao, Zihan Chen 0001, Liyinglan Liu, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Wanlu Liu, Howard H. Yang, Zuozhu Liu
ICLR6
2024 Making Pre-trained Language Models Great on Tabular Prediction
abstract
The transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime.
Jiahuan Yan, Bo Zheng 0011, Yiheng Zhu 0002, Danny Ziyi Chen, Jimeng Sun 0001, Jian Wu 0001, Jintai Chen
ICLR7
2024 Personalized Heart Disease Detection via ECG Digital Twin Generation
Yaojun Hu, Jintai Chen, Lianting Hu, Dantong Li, Jiahuan Yan, Haochao Ying, Huiying Liang, Jian Wu 0001
IJCAI8
2024 AI-Enhanced Virtual Reality in Medicine: A Comprehensive Survey
Kaiyuan Hu, Danny Ziyi Chen, Jian Wu 0001
IJCAI4
2024 Can a Deep Learning Model be a Sure Bet for Tabular Prediction?
abstract
Data organized in tabular format is ubiquitous in real-world applications, and users often craft tables with biased feature definitions and flexibly set prediction targets of their interests. Thus, a rapid development of a robust, effective, dataset-versatile, user-friendly tabular prediction approach is highly desired. While Gradient Boosting Decision Trees (GBDTs) and existing deep neural networks (DNNs) have been extensively utilized by professional users, they present several challenges for casual users, particularly: (i) the dilemma of model selection due to their different dataset preferences, and (ii) the need for heavy hyperparameter searching, failing which their performances are deemed inadequate. In this paper, we delve into this question: Can we develop a deep learning model that serves as a sure bet solution for a wide range of tabular prediction tasks, while also being user-friendly for casual users? We delve into three key drawbacks of deep tabular models, encompassing: (P1) lack of rotational variance property, (P2) large data demand, and (P3) over-smooth solution. We propose ExcelFormer, addressing these challenges through a semi-permeable attention module that effectively constrains the influence of less informative features to break the DNNs' rotational invariance property (for P1), data augmentation approaches tailored for tabular data (for P2), and attentive feedforward network to boost the model fitting capability (for P3). These designs collectively make ExcelFormer a sure bet solution for diverse tabular datasets. Extensive and stratified experiments conducted on real-world datasets demonstrate that our model outperforms previous approaches across diverse tabular data prediction tasks, and this framework can be friendly to casual users, offering ease of use without the heavy hyperparameter tuning. The codes are available at https://github.com/whatashot/excelformer.
Jintai Chen, Jiahuan Yan, Qiyuan Chen 0003, Danny Ziyi Chen, Jian Wu 0001, Jimeng Sun 0001
KDD5
2024 Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPs
abstract
Tabular datasets play a crucial role in various applications.Thus, developing efficient, effective, and widely compatible prediction algorithms for tabular data is important.Currently, two prominent model types, Gradient Boosted Decision Trees (GBDTs) and Deep Neural Networks (DNNs), have demonstrated performance advantages on distinct tabular prediction tasks.However, selecting an effective model for a specific tabular dataset is challenging, often demanding time-consuming hyperparameter tuning.To address this model selection dilemma, this paper proposes a new framework that amalgamates the advantages of both GBDTs and DNNs, resulting in a DNN algorithm that is as efficient as GBDTs and is competitively effective regardless of dataset preferences for GBDTs or DNNs.Our idea is rooted in an observation that deep learning (DL) offers a larger parameter space that can represent a well-performing GBDT model, yet the current back-propagation optimizer struggles to efficiently discover such optimal functionality.On the other hand, during GBDT development, hard tree pruning, entropy-driven feature gate, and model ensemble have proved to be more adaptable to tabular data.By combining these key components, we present a Tree-hybrid simple MLP (T-MLP).In our framework, a tensorized, rapidly trained GBDT feature gate, a DNN architecture pruning approach, as well as a vanilla back-propagation optimizer collaboratively train a randomly initialized MLP model.Comprehensive experiments show that T-MLP is competitive with extensively tuned DNNs and GBDTs in their dominating tabular benchmarks (88 datasets) respectively, all achieved with compact model storage and significantly reduced training duration.The codes and full experiment results are available at https://github.com/jyansir/tmlp.
Jiahuan Yan, Jintai Chen, Qianxing Wang, Danny Ziyi Chen, Jian Wu 0001
KDD5
2024 PX2Tooth: Reconstructing the 3D Point Cloud Teeth from a Single Panoramic X-Ray
Huikai Wu, Zikai Xiao, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
MICCAI (3)5
2024 🐍 LKM-UNet: Large Kernel Vision Mamba UNet for Medical Image Segmentation
Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
MICCAI (8)4
2024 TeleOR: Real-Time Telemedicine System for Full-Scene Operating Room
Kaiyuan Hu, Qian Shao, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
MICCAI (6)6
2024 Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
abstract
Weize Liu, Guocong Li, Kai Zhang, Bang Du, Qiyuan Chen, Xuming Hu, Hongxia Xu, Jintai Chen, Jian Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Weize Liu, Guocong Li, Kai Zhang 0053, Bang Du, Qiyuan Chen 0003, Xuming Hu, Jintai Chen, Jian Wu 0001
NAACL-HLT9
2024 Enhancing Semi-Supervised Learning via Representative and Diverse Sample Selection
abstract
Semi-Supervised Learning (SSL) has become a preferred paradigm in many deep learning tasks, which reduces the need for human labor. Previous studies primarily focus on effectively utilising the labelled and unlabeled data to improve performance. However, we observe that how to select samples for labelling also significantly impacts performance, particularly under extremely low-budget settings. The sample selection task in SSL has been under-explored for a long time. To fill in this gap, we propose a Representative and Diverse Sample Selection approach (RDSS). By adopting a modified Frank-Wolfe algorithm to minimise a novel criterion $\alpha$-Maximum Mean Discrepancy ($\alpha$-MMD), RDSS samples a representative and diverse subset for annotation from the unlabeled data. We demonstrate that minimizing $\alpha$-MMD enhances the generalization ability of low-budget learning. Experimental results show that RDSS consistently improves the performance of several popular SSL frameworks and outperforms the state-of-the-art sample selection approaches used in Active Learning (AL) and Semi-Supervised Active Learning (SSAL), even with constrained annotation budgets. Our code is available at [RDSS](https://github.com/YanhuiAILab/RDSS).
Qian Shao, Jiangrui Kang, Qiyuan Chen 0003, Zepeng Li 0002, Yiwen Cao, Jiajuan Liang, Jian Wu 0001
NeurIPS8
2024 SynergyX: a multi-modality mutual attention network for interpretable drug synergy prediction
abstract
Discovering effective anti-tumor drug combinations is crucial for advancing cancer therapy. Taking full account of intricate biological interactions is highly important in accurately predicting drug synergy. However, the extremely limited prior knowledge poses great challenges in developing current computational methods. To address this, we introduce SynergyX, a multi-modality mutual attention network to improve anti-tumor drug synergy prediction. It dynamically captures cross-modal interactions, allowing for the modeling of complex biological networks and drug interactions. A convolution-augmented attention structure is adopted to integrate multi-omic data in this framework effectively. Compared with other state-of-the-art models, SynergyX demonstrates superior predictive accuracy in both the General Test and Blind Test and cross-dataset validation. By exhaustively screening combinations of approved drugs, SynergyX reveals its ability to identify promising drug combination candidates for potential lung cancer treatment. Another notable advantage lies in its multidimensional interpretability. Taking Sorafenib and Vorinostat as an example, SynergyX serves as a powerful tool for uncovering drug-gene interactions and deciphering cell selectivity mechanisms. In summary, SynergyX provides an illuminating and interpretable framework, poised to catalyze the expedition of drug synergy discovery and deepen our comprehension of rational combination therapy.
Yue Guo 0008, Haitao Hu, Jian Wu 0001, Chang-Yu Hsieh, Qiaojun He
Briefings Bioinform.5
2024 DPML: Prior-guided multitask learning for dental object recognition on limited panoramic radiograph dataset
Zheng Cao 0005, Chengyu Feng, Yefeng Shen, Guanchen Ye, Jian Wu 0001, Zhendong Wu, Honghao Gao, Haihua Zhu 0002
Expert Syst. Appl.6
2024 Collaborative knowledge amalgamation: Preserving discriminability and transferability in unsupervised learning
Shangde Gao, Yichao Fu, Ke Liu 0012, Wei Gao 0001, Jian Wu 0001, Yuqiang Han
Inf. Sci.6
2024 A Corresponding Region Fusion Framework for Multi-Modal Cervical Lesion Detection
abstract
Cervical lesion detection (CLD) using colposcopic images of multi-modality (acetic and iodine) is critical to computer-aided diagnosis (CAD) systems for accurate, objective, and comprehensive cervical cancer screening. To robustly capture lesion features and conform with clinical diagnosis practice, we propose a novel corresponding region fusion network (CRFNet) for multi-modal CLD. CRFNet first extracts feature maps and generates proposals for each modality, then performs proposal shifting to obtain corresponding regions under large position shifts between modalities, and finally fuses those region features with a new corresponding channel attention to detect lesion regions on both modalities. To evaluate CRFNet, we build a large multi-modal colposcopic image dataset collected from our collaborative hospital. We show that our proposed CRFNet surpasses known single-modal and multi-modal CLD methods and achieves state-of-the-art performance, especially in terms of Average Precision.
Tingting Chen 0002, Heping Hu, Chunhua Luo, Jintai Chen, Chunnv Yuan, Weiguo Lu, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE Trans. Comput. Biol. Bioinform.10
2024 A Protein-Context Enhanced Master Slave Framework for Zero-Shot Drug Target Interaction Prediction
abstract
Drug Target Interaction (DTI) prediction plays a crucial role in in-silico drug discovery, especially for deep learning (DL) models. Along this line, existing methods usually first extract features from drugs and target proteins, and use drug-target pairs to train DL models. However, these DL-based methods essentially rely on similar structures and patterns defined by the homologous proteins from a large amount of data. When few drug-target interactions are known for a newly discovered protein and its homologous proteins, prediction performance can suffer notable reduction. In this paper, we propose a novel Protein-Context enhanced Master/Slave Framework (PCMS), for zero-shot DTI prediction. This framework facilitates the efficient discovery of ligands for newly discovered target proteins, addressing the challenge of predicting interactions without prior data. Specifically, the PCMS framework consists of two main components: a Master Learner and a Slave Learner. The Master Learner first learns the target protein context information, and then adaptively generates the corresponding parameters for the Slave Learner. The Slave Learner then perform zero-shot DTI prediction in different protein contexts. Extensive experiments verify the effectiveness of our PCMS compared to state-of-the-art methods in various metrics on two public datasets.
Yuyang Xu, Jingbo Zhou 0003, Haochao Ying, Jintai Chen, Wei Chen 0001, Danny Ziyi Chen, Jian Wu 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2024 Polygonal Approximation Learning for Convex Object Segmentation in Biomedical Images With Bounding Box Supervision
abstract
As a common and critical medical image analysis task, deep learning based biomedical image segmentation is hindered by the dependence on costly fine-grained annotations. To alleviate this data dependence, in this article, a novel approach, called Polygonal Approximation Learning (PAL), is proposed for convex object instance segmentation with only bounding-box supervision. The key idea behind PAL is that the detection model for convex objects already contains the necessary information for segmenting them since their convex hulls, which can be generated approximately by the intersection of bounding boxes, are equivalent to the masks representing the objects. To extract the essential information from the detection model, a repeated detection approach is employed on biomedical images where various rotation angles are applied and a dice loss with the projection of the rotated detection results is utilized as a supervised signal in training our segmentation model. In biomedical imaging tasks involving convex objects, such as nuclei instance segmentation, PAL outperforms the known models (e.g., BoxInst) that rely solely on box supervision. Furthermore, PAL achieves comparable performance with mask-supervised models including Mask R-CNN and Cascade Mask R-CNN. Interestingly, PAL also demonstrates remarkable performance on non-convex object instance segmentation tasks, for example, surgical instrument and organ instance segmentation.
Jintai Chen, Kai Zhang 0053, Jiahuan Yan, Bang Du, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE J. Biomed. Health Informatics10
2024 A Transformer-Based Knowledge Distillation Network for Cortical Cataract Grading
abstract
Cortical cataract, a common type of cataract, is particularly difficult to be diagnosed automatically due to the complex features of the lesions. Recently, many methods based on edge detection or deep learning were proposed for automatic cataract grading. However, these methods suffer a large performance drop in cortical cataract grading due to the more complex cortical opacities and uncertain data. In this paper, we propose a novel Transformer-based Knowledge Distillation Network, called TKD-Net, for cortical cataract grading. To tackle the complex opacity problem, we first devise a zone decomposition strategy to extract more refined features and introduce special sub-scores to consider critical factors of clinical cortical opacity assessment (location, area, density) for comprehensive quantification. Next, we develop a multi-modal mix-attention Transformer to efficiently fuse sub-scores and image modality for complex feature learning. However, obtaining the sub-score modality is a challenge in the clinic, which could cause the modality missing problem instead. To simultaneously alleviate the issues of modality missing and uncertain data, we further design a Transformer-based knowledge distillation method, which uses a teacher model with perfect data to guide a student model with modality-missing and uncertain data. We conduct extensive experiments on a dataset of commonly-used slit-lamp images annotated by the LOCS III grading system to demonstrate that our TKD-Net outperforms state-of-the-art methods, as well as the effectiveness of its key components. Codes are available at https://github.com/wjh892521292/Cataract_TKD-Net.
Haochao Ying, Tingting Chen 0002, Zuozhu Liu, Danny Ziyi Chen, Ke Yao, Jian Wu 0001
IEEE Trans. Medical Imaging9
2023 T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature Interaction
abstract
Recent development of deep neural networks (DNNs) for tabular learning has largely benefited from the capability of DNNs for automatic feature interaction. However, the heterogeneity nature of tabular features makes such features relatively independent, and developing effective methods to promote tabular feature interaction still remains an open problem. In this paper, we propose a novel Graph Estimator, which automatically estimates the relations among tabular features and builds graphs by assigning edges between related features. Such relation graphs organize independent tabular features into a kind of graph data such that interaction of nodes (tabular features) can be conducted in an orderly fashion. Based on our proposed Graph Estimator, we present a bespoke Transformer network tailored for tabular learning, called T2G-Former, which processes tabular data by performing tabular feature interaction guided by the relation graphs. A specific Cross-level Readout collects salient features predicted by the layers in T2G-Former across different levels, and attains global semantics for final prediction. Comprehensive experiments show that our T2G-Former achieves superior performance among DNNs and is competitive with non-deep Gradient Boosted Decision Tree models. The code and detailed results are available at https://github.com/jyansir/t2g-former.
Jiahuan Yan, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
AAAI5
2023 Ord2Seq: Regarding Ordinal Regression as Label Sequence Prediction
abstract
Ordinal regression refers to classifying object instances into ordinal categories. It has been widely studied in many scenarios, such as medical disease grading and movie rating. Known methods focused only on learning inter-class ordinal relationships, but still incur limitations in distinguishing adjacent categories thus far. In this paper, we propose a simple sequence prediction framework for ordinal regression called Ord2Seq, which, for the first time, transforms each ordinal category label into a special label sequence and thus regards an ordinal regression task as a sequence prediction process. In this way, we decompose an ordinal regression task into a series of recursive binary classification steps, so as to subtly distinguish adjacent categories. Comprehensive experiments show the effectiveness of distinguishing adjacent categories for performance improvement and our new approach exceeds state-of-the-art performances in four different scenarios. Codes are available at https://github.com/wjh892521292/Ord2Seq.
Jintai Chen, Tingting Chen 0002, Danny Ziyi Chen, Jian Wu 0001
ICCV6
2023 TabCaps: A Capsule Neural Network for Tabular Data Classification with BoW Routing
Jintai Chen, Kuanlun Liao, Yanwen Fang, Danny Ziyi Chen, Jian Wu 0001
ICLR5
2023 Robust Image Ordinal Regression with Controllable Image Generation
abstract
Image ordinal regression has been mainly studied along the line of exploiting the order of categories. However, the issues of class imbalance and category overlap that are very common in ordinal regression were largely overlooked. As a result, the performance on minority categories is often unsatisfactory. In this paper, we propose a novel framework called CIG based on controllable image generation to directly tackle these two issues. Our main idea is to generate extra training samples with specific labels near category boundaries, and the sample generation is biased toward the less-represented categories. To achieve controllable image generation, we seek to separate structural and categorical information of images based on structural similarity, categorical similarity, and reconstruction constraints. We evaluate the effectiveness of our new CIG approach in three different image ordinal regression scenarios. The results demonstrate that CIG can be flexibly integrated with off-the-shelf image encoders or ordinal regression models to achieve improvement, and further, the improvement is more significant for minority categories.
Haochao Ying, Renjun Hu, Xiao Zhang 0015, Danny Ziyi Chen, Jian Wu 0001
IJCAI8
2023 MolHF: A Hierarchical Normalizing Flow for Molecular Graph Generation
abstract
Molecular de novo design is a critical yet challenging task in scientific fields, aiming to design novel molecular structures with desired property profiles. Significant progress has been made by resorting to generative models for graphs. However, limited attention is paid to hierarchical generative models, which can exploit the inherent hierarchical structure (with rich semantic information) of the molecular graphs and generate complex molecules of larger size that we shall demonstrate to be difficult for most existing models. The primary challenge to hierarchical generation is the non-differentiable issue caused by the generation of intermediate discrete coarsened graph structures. To sidestep this issue, we cast the tricky hierarchical generation problem over discrete spaces as the reverse process of hierarchical representation learning and propose MolHF, a new hierarchical flow-based model that generates molecular graphs in a coarse-to-fine manner. Specifically, MolHF first generates bonds through a multi-scale architecture, then generates atoms based on the coarsened graph structure at each scale. We demonstrate that MolHF achieves state-of-the-art performance in random generation and property optimization, implying its high capacity to model data distribution. Furthermore, MolHF is the first flow-based model that can be applied to model larger molecules (polymer) with more than 100 heavy atoms. The code and models are available at https://github.com/violet-sto/MolHF.
Yiheng Zhu 0002, Zhenqiu Ouyang, Ben Liao, Jialu Wu, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
IJCAI8
2023 TSegFormer: 3D Tooth Segmentation in Intraoral Scans with Geometry Guided Transformer
Huimin Xiong, Kunle Li, Kaiyuan Tan, Yang Feng 0011, Joey Tianyi Zhou, Jin Hao, Haochao Ying, Jian Wu 0001, Zuozhu Liu
MICCAI (6)8
2023 GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels
abstract
Since annotating medical images for segmentation tasks commonly incurs expensive costs, it is highly desirable to design an annotation-efficient method to alleviate the annotation burden. Recently, contrastive learning has exhibited a great potential in learning robust representations to boost downstream tasks with limited labels. In medical imaging scenarios, ready-made meta labels (i.e., specific attribute information of medical images) inherently reveal semantic relationships among images, which have been used to define positive pairs in previous work. However, the multi-perspective semantics revealed by various meta labels are usually incompatible and can incur intractable "semantic contradiction" when combining different meta labels. In this paper, we tackle the issue of "semantic contradiction" in a gradient-guided manner using our proposed Gradient Mitigator method, which systematically unifies multi-perspective meta labels to enable a pre-trained model to attain a better high-level semantic recognition ability. Moreover, we emphasize that the fine-grained discrimination ability is vital for segmentation-oriented pre-training, and develop a novel method called Gradient Filter to dynamically screen pixel pairs with the most discriminating power based on the magnitude of gradients. Comprehensive experiments on four medical image segmentation datasets verify that our new method GCL: (1) learns informative image representations and considerably boosts segmentation performance with limited labels, and (2) shows promising generalizability on out-of-distribution datasets.
Jintai Chen, Jiahuan Yan, Yiheng Zhu 0002, Danny Ziyi Chen, Jian Wu 0001
ACM Multimedia6
2023 Towards Distribution-Agnostic Generalized Category Discovery
abstract
Data imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on classifying close-set samples and detecting open-set samples during testing, it's still essential to be able to classify unknown subjects as human beings. In this paper, we formally define a more realistic task as distribution-agnostic generalized category discovery (DA-GCD): generating fine-grained predictions for both close- and open-set classes in a long-tailed open-world setting. To tackle the challenging problem, we propose a Self-**Ba**lanced **Co**-Advice co**n**trastive framework (BaCon), which consists of a contrastive-learning branch and a pseudo-labeling branch, working collaboratively to provide interactive supervision to resolve the DA-GCD task. In particular, the contrastive-learning branch provides reliable distribution estimation to regularize the predictions of the pseudo-labeling branch, which in turn guides contrastive learning through self-balanced knowledge transfer and a proposed novel contrastive loss. We compare BaCon with state-of-the-art methods from two closely related fields: imbalanced semi-supervised learning and generalized category discovery. The effectiveness of BaCon is demonstrated with superior performance over all baselines and comprehensive analysis across various datasets. Our code is publicly available.
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li 0001, Joey Tianyi Zhou, Yang Feng 0011, Jian Wu 0001, Haoji Hu
NeurIPS9
2023 Fast Model DeBias with Machine Unlearning
abstract
Recent discoveries have revealed that deep neural networks might behave in a biased manner in many real-world scenarios. For instance, deep networks trained on a large-scale face recognition dataset CelebA tend to predict blonde hair for females and black hair for males. Such biases not only jeopardize the robustness of models but also perpetuate and amplify social biases, which is especially concerning for automated decision-making processes in healthcare, recruitment, etc., as they could exacerbate unfair economic and social inequalities among different groups. Existing debiasing methods suffer from high costs in bias labeling or model re-training, while also exhibiting a deficiency in terms of elucidating the origins of biases within the model. To this respect, we propose a fast model debiasing method (FMD) which offers an efficient approach to identify, evaluate and remove biases inherent in trained models. The FMD identifies biased attributes through an explicit counterfactual concept and quantifies the influence of data samples with influence functions. Moreover, we design a machine unlearning-based strategy to efficiently and effectively remove the bias in a trained model with a small counterfactual dataset. Experiments on the Colored MNIST, CelebA, and Adult Income datasets demonstrate that our method achieves superior or competing classification accuracies compared with state-of-the-art retraining-based methods while attaining significantly fewer biases and requiring much less debiasing cost. Notably, our method requires only a small external dataset and updating a minimal amount of model parameters, without the requirement of access to training data that may be too large or unavailable in practice.
Ruizhe Chen, Huimin Xiong, Jianhong Bai, Tianxiang Hu, Jin Hao, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Zuozhu Liu
NeurIPS9
2023 Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient Balancer
abstract
Data privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist.
Zikai Xiao, Zihan Chen 0001, Songshang Liu, Hualiang Wang, Yang Feng 0011, Jin Hao, Joey Tianyi Zhou, Jian Wu 0001, Howard H. Yang, Zuozhu Liu
NeurIPS8
2023 Sample-efficient Multi-objective Molecular Optimization with GFlowNets
abstract
Many crucial scientific problems involve designing novel molecules with desired properties, which can be formulated as a black-box optimization problem over the *discrete* chemical space. In practice, multiple conflicting objectives and costly evaluations (e.g., wet-lab experiments) make the *diversity* of candidates paramount. Computational methods have achieved initial success but still struggle with considering diversity in both objective and search space. To fill this gap, we propose a multi-objective Bayesian optimization (MOBO) algorithm leveraging the hypernetwork-based GFlowNets (HN-GFN) as an acquisition function optimizer, with the purpose of sampling a diverse batch of candidate molecular graphs from an approximate Pareto front. Using a single preference-conditioned hypernetwork, HN-GFN learns to explore various trade-offs between objectives. We further propose a hindsight-like off-policy strategy to share high-performing molecules among different preferences in order to speed up learning for HN-GFN. We empirically illustrate that HN-GFN has adequate capacity to generalize over preferences. Moreover, experiments in various real-world MOBO settings demonstrate that our framework predominantly outperforms existing methods in terms of candidate quality and sample efficiency. The code is available at https://github.com/violet-sto/HN-GFN.
Yiheng Zhu 0002, Jialu Wu, Chaowen Hu, Jiahuan Yan, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001
NeurIPS7
2023 Robust Training of Graph Neural Networks via Noise Governance
abstract
Graph Neural Networks (GNNs) have become widely-used models for semi-supervised learning. However, the robustness of GNNs in the presence of label noise remains a largely under-explored problem. In this paper, we consider an important yet challenging scenario where labels on nodes of graphs are not only noisy but also scarce. In this scenario, the performance of GNNs is prone to degrade due to label noise propagation and insufficient learning. To address these issues, we propose a novel RTGNN (Robust Training of Graph Neural Networks via Noise Governance) framework that achieves better robustness by learning to explicitly govern label noise. More specifically, we introduce self-reinforcement and consistency regularization as supplemental supervision. The self-reinforcement supervision is inspired by the memorization effects of deep neural networks and aims to correct noisy labels. Further, the consistency regularization prevents GNNs from overfitting to noisy labels via mimicry loss in both the inter-view and intra-view perspectives. To leverage such supervisions, we divide labels into clean and noisy types, rectify inaccurate labels, and further generate pseudo-labels on unlabeled nodes. Supervision for nodes with different types of labels is then chosen adaptively. This enables sufficient learning from clean labels while limiting the impact of noisy ones. We conduct extensive experiments to evaluate the effectiveness of our RTGNN framework, and the results validate its consistent superior performance over state-of-the-art methods with two types of label noises and various noise rates.
Siyi Qian, Haochao Ying, Renjun Hu, Jingbo Zhou 0003, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
WSDM7
2023 TransFoxMol: predicting molecular property with focused attention
abstract
Predicting the biological properties of molecules is crucial in computer-aided drug development, yet it's often impeded by data scarcity and imbalance in many practical applications. Existing approaches are based on self-supervised learning or 3D data and using an increasing number of parameters to improve performance. These approaches may not take full advantage of established chemical knowledge and could inadvertently introduce noise into the respective model. In this study, we introduce a more elegant transformer-based framework with focused attention for molecular representation (TransFoxMol) to improve the understanding of artificial intelligence (AI) of molecular structure property relationships. TransFoxMol incorporates a multi-scale 2D molecular environment into a graph neural network + Transformer module and uses prior chemical maps to obtain a more focused attention landscape compared to that obtained using existing approaches. Experimental results show that TransFoxMol achieves state-of-the-art performance on MoleculeNet benchmarks and surpasses the performance of baselines that use self-supervised learning or geometry-enhanced strategies on small-scale datasets. Subsequent analyses indicate that TransFoxMol's predictions are highly interpretable and the clever use of chemical knowledge enables AI to perceive molecules in a simple but rational way, enhancing performance.
Zheyuan Shen, Sikang Chen, Qingyu Bian, Yue Guo 0008, Liteng Shen, Jian Wu 0001, Binbin Zhou 0005, Tingjun Hou, Qiaojun He, Jinxin Che, Xiaowu Dong
Briefings Bioinform.10
2023 D-former: a U-shaped Dilated Transformer for 3D medical image segmentation
Kuanlun Liao, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
Neural Comput. Appl.7
2023 CariesNet: a deep learning approach for segmentation of multi-stage caries lesion from oral panoramic X-ray image
Haihua Zhu 0002, Zheng Cao 0005, Luya Lian, Guanchen Ye, Honghao Gao, Jian Wu 0001
Neural Comput. Appl.6
2023 Identifying Electrocardiogram Abnormalities Using a Handcrafted-Rule-Enhanced Neural Network
abstract
A large number of people suffer from life-threatening cardiac abnormalities, and electrocardiogram (ECG) analysis is beneficial to determining whether an individual is at risk of such abnormalities. Automatic ECG classification methods, especially the deep learning based ones, have been proposed to detect cardiac abnormalities using ECG records, showing good potential to improve clinical diagnosis and help early prevention of cardiovascular diseases. However, the predictions of the known neural networks still do not satisfactorily meet the needs of clinicians, and this phenomenon suggests that some information used in clinical diagnosis may not be well captured and utilized by these methods. In this paper, we introduce some rules into convolutional neural networks, which help present clinical knowledge to deep learning based ECG analysis, in order to improve automated ECG diagnosis performance. Specifically, we propose a Handcrafted-Rule-enhanced Neural Network (called HRNN) for ECG classification with standard 12-lead ECG input, which consists of a rule inference module and a deep learning module. Experiments on two large-scale public ECG datasets show that our new approach considerably outperforms existing state-of-the-art methods. Further, our proposed approach not only can improve the diagnosis performance, but also can assist in detecting mislabelled ECG samples.
Yuexin Bian, Jintai Chen, Xiaoxian Yang, Danny Ziyi Chen, Jian Wu 0001
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 Time-Aware Context-Gated Graph Attention Network for Clinical Risk Prediction
abstract
Clinical risk prediction based on Electronic Health Records (EHR) can assist doctors in better judgment and can make sense of early diagnosis. However, the prediction performance heavily relies on effective representations from multi-dimensional time-series EHR data. Existing solutions usually focus on temporal features or inherent relations between clinical event variables or extract both information in two separate phases. This usually leads to insufficient patient feature information and results in poor prediction performance. Moreover, existing methods based on Heterogeneous Graph Neural Network usually require manual selection of proper Meta-Paths. To solve these problems, we propose the Time-aware Context-Gated Graph Attention Network (T-ContextGGAN). Specifically, we design a GNN based module with Time-aware Meta-Paths and self-attention mechanism to extract both temporal semantic information and inherent relations of EHR data simultaneously and perform automatic Meta-Path selection. To evaluate the proposed model, we extract the first 48 hour EHR data in the first Intensive Care Unit (ICU) admission of three different tasks from two open-source datasets and model various clinical variables on the proposed EHRGraph. Extensive experimental results show the proposed model can effectively extract informative features, and outperform existing state-of-art models in terms of various prediction measures. Our code is available in https://github.com/OwlCitizen/TContext-GGAN.
Yuyang Xu, Haochao Ying, Siyi Qian, Fuzhen Zhuang, Xiao Zhang 0015, Deqing Wang 0001, Jian Wu 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.7
2023 A Robust Shape-Aware Rib Fracture Detection and Segmentation Framework With Contrastive Learning
abstract
The rib fracture is a common type of thoracic skeletal trauma, and its inspections using computed tomography (CT) scans are critical for clinical evaluation and treatment planning. However, it is often challenging for radiologists to quickly and accurately detect rib fractures due to tiny objects and blurriness in large 3D CT images. Previous diagnoses for automatic rib fracture mostly relied on deep learning (DL)-based object detection, which highly depends on label quality and quantity. Moreover, general object detection methods did not take into consideration the typically elongated and oblique shapes of ribs in 3D volumes. To address these issues, we propose a shape-aware method based on DL called SA-FracNet for rib fracture detection and segmentation. First, we design a pixel-level pretext task founded on contrastive learning on massive unlabeled CT images. Second, we train the fine-tuned rib fracture detection model based on the pre-trained weights. Third, we develop a fracture shape-aware multi-task segmentation network to delineate the fracture based on the detection result. Experiments demonstrate that our proposed SA-FracNet achieves state-of-the-art rib fracture detection and segmentation performance on the public RibFrac dataset, with a detection sensitivity of 0.926 and segmentation Dice of 0.754. Test on a private dataset also validates the robustness and generalization of our SA-FracNet.
Zheng Cao 0005, Liming Xu, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE Trans. Multim.5
2023 FraudAuditor: A Visual Analytics Approach for Collusive Fraud in Health Insurance
abstract
Collusive fraud, in which multiple fraudsters collude to defraud health insurance funds, threatens the operation of the healthcare system. However, existing statistical and machine learning-based methods have limited ability to detect fraud in the scenario of health insurance due to the high similarity of fraudulent behaviors to normal medical visits and the lack of labeled data. To ensure the accuracy of the detection results, expert knowledge needs to be integrated with the fraud detection process. By working closely with health insurance audit experts, we propose FraudAuditor, a three-stage visual analytics approach to collusive fraud detection in health insurance. Specifically, we first allow users to interactively construct a co-visit network to holistically model the visit relationships of different patients. Second, an improved community detection algorithm that considers the strength of fraud likelihood is designed to detect suspicious fraudulent groups. Finally, through our visual interface, users can compare, investigate, and verify suspicious patient behavior with tailored visualizations that support different time scales. We conducted case studies in a real-world healthcare scenario, i.e., to help locate the actual fraud group and exclude the false positive group. The results and expert feedback proved the effectiveness and usability of the approach.
Jiehui Zhou, Xumeng Wang, Huanliang Wang, Zihan Zhou 0009, Dongming Han, Haochao Ying, Jian Wu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.9
2022 DANets: Deep Abstract Networks for Tabular Data Classification and Regression
abstract
Tabular data are ubiquitous in real world applications. Although many commonly-used neural components (e.g., convolution) and extensible neural networks (e.g., ResNet) have been developed by the machine learning community, few of them were effective for tabular data and few designs were adequately tailored for tabular data structures. In this paper, we propose a novel and flexible neural component for tabular data, called Abstract Layer (AbstLay), which learns to explicitly group correlative input features and generate higher-level features for semantics abstraction. Also, we design a structure re-parameterization method to compress the trained AbstLay, thus reducing the computational complexity by a clear margin in the reference phase. A special basic block is built using AbstLays, and we construct a family of Deep Abstract Networks (DANets) for tabular data classification and regression by stacking such blocks. In DANets, a special shortcut path is introduced to fetch information from raw tabular features, assisting feature interactions across different levels. Comprehensive experiments on seven real-world tabular datasets show that our AbstLay and DANets are effective for tabular data classification and regression, and the computational complexity is superior to competitive methods. Besides, we evaluate the performance gains of DANet as it goes deep, verifying the extendibility of our method. Our code is available at https://github.com/WhatAShot/DANet.
Jintai Chen, Kuanlun Liao, Yao Wan 0001, Danny Ziyi Chen, Jian Wu 0001
AAAI5
2022 CTT-Net: A Multi-view Cross-token Transformer for Cataract Postoperative Visual Acuity Prediction
abstract
Surgery is the only viable treatment for cataract patients with visual acuity (VA) impairment. Clinically, to assess the necessity of cataract surgery, accurately predicting postoperative VA before surgery by analyzing multi-view optical coherence tomography (OCT) images is crucially needed. Unfortunately, due to complicated fundus conditions, determining postoperative VA remains difficult for medical experts. Deep learning methods for this problem were developed in recent years. Although effective, these methods still face several issues, such as not efficiently exploring potential relations between multi-view OCT images, neglecting the key role of clinical prior knowledge (e.g., preoperative VA value), and using only regression-based metrics which are lacking reference. In this paper, we propose a novel Cross-token Transformer Network (CTT-Net) for postoperative VA prediction by analyzing both the multi-view OCT images and preoperative VA. To effectively fuse multi-view features of OCT images, we develop cross-token attention that could restrict redundant/unnecessary attention flow. Further, we utilize the preoperative VA value to provide more information for postoperative VA prediction and facilitate fusion between views. Moreover, we design an auxiliary classification loss to improve model performance and assess VA recovery more sufficiently, avoiding the limitation by only using the regression metrics. To evaluate CTT-Net, we build a multi-view OCT image dataset collected from our collaborative hospital. A set of extensive experiments validate the effectiveness of our model compared to existing methods in various metrics. Code is available at: https://github.con wjh892521292/Cataract-OCT.
Tingting Chen 0002, Xingdi Wu, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001
BIBM10
2022 DialMed: A Dataset for Dialogue-based Medication Recommendation
abstract
Medication recommendation is a crucial task for intelligent healthcare systems. Previous studies mainly recommend medications with electronic health records (EHRs). However, some details of interactions between doctors and patients may be ignored or omitted in EHRs, which are essential for automatic medication recommendation. Therefore, we make the first attempt to recommend medications with the conversations between doctors and patients. In this work, we construct DIALMED, the first high-quality dataset for medical dialogue-based medication recommendation task. It contains 11, 996 medical dialogues related to 16 common diseases from 3 departments and 70 corresponding common medications. Furthermore, we propose a Dialogue structure and Disease knowledge aware Network (DDN), where a QA Dialogue Graph mechanism is designed to model the dialogue structure and the knowledge graph is used to introduce external disease knowledge. The extensive experimental results demonstrate that the proposed method is a promising solution to recommend medications with medical dialogues. The dataset and code are available at https://github.com/f-window/DialMed.
Zhenfeng He, Yuqiang Han, Zhenqiu Ouyang, Wei Gao 0001, Hongxu Chen 0002, Guandong Xu, Jian Wu 0001
COLING7
2022 ME-GAN: Learning Panoptic Electrocardio Representations for Multi-view ECG Synthesis Conditioned on Heart Diseases
abstract
Electrocardiogram (ECG) is a widely used non-invasive diagnostic tool for heart diseases. Many studies have devised ECG analysis models (e.g., classifiers) to assist diagnosis. As an upstream task, researches have built generative models to synthesize ECG data, which are beneficial to providing training samples, privacy protection, and annotation reduction. However, previous generative methods for ECG often neither synthesized multi-view data, nor dealt with heart disease conditions. In this paper, we propose a novel disease-aware generative adversarial network for multi-view ECG synthesis called ME-GAN, which attains panoptic electrocardio representations conditioned on heart diseases and projects the representations onto multiple standard views to yield ECG signals. Since ECG manifestations of heart diseases are often localized in specific waveforms, we propose a new "mixup normalization" to inject disease information precisely into suitable locations. In addition, we propose a "view discriminator" to revert disordered ECG views into a pre-determined order, supervising the generator to obtain ECG representing correct view characteristics. Besides, a new metric, rFID, is presented to assess the quality of the synthesized ECG signals. Comprehensive experiments verify that our ME-GAN performs well on multi-view ECG signal synthesis with trusty morbid manifestations.
Jintai Chen, Kuanlun Liao, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001
ICML6
2022 Automating Blastocyst Formation and Quality Prediction in Time-Lapse Imaging with Adaptive Key Frame Selection
Tingting Chen 0002, Zhaoxia Yang, Danny Ziyi Chen, Jian Wu 0001
MICCAI (4)7
2022 Self-learning and One-Shot Learning Based Single-Slice Annotation for 3D Medical Image Segmentation
Bo Zheng 0011, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001
MICCAI (8)5
2022 Out-of-the-box deep learning prediction of quantum-mechanical partial charges by graph representation and transfer learning
abstract
Accurate prediction of atomic partial charges with high-level quantum mechanics (QM) methods suffers from high computational cost. Numerous feature-engineered machine learning (ML)-based predictors with favorable computability and reliability have been developed as alternatives. However, extensive expertise effort was needed for feature engineering of atom chemical environment, which may consequently introduce domain bias. In this study, SuperAtomicCharge, a data-driven deep graph learning framework, was proposed to predict three important types of partial charges (i.e. RESP, DDEC4 and DDEC78) derived from high-level QM calculations based on the structures of molecules. SuperAtomicCharge was designed to simultaneously exploit the 2D and 3D structural information of molecules, which was proved to be an effective way to improve the prediction accuracy of the model. Moreover, a simple transfer learning strategy and a multitask learning strategy based on self-supervised descriptors were also employed to further improve the prediction accuracy of the proposed model. Compared with the latest baselines, including one GNN-based predictor and two ML-based predictors, SuperAtomicCharge showed better performance on all the three external test sets and had better usability and portability. Furthermore, the QM partial charges of new molecules predicted by SuperAtomicCharge can be efficiently used in drug design applications such as structure-based virtual screening, where the predicted RESP and DDEC4 charges of new molecules showed more robust scoring and screening power than the commonly used partial charges. Finally, two tools including an online server (http://cadd.zju.edu.cn/deepchargepredictor) and the source code command lines (https://github.com/zjujdj/SuperAtomicCharge) were developed for the easy access of the SuperAtomicCharge services.
Dejun Jiang 0002, Huiyong Sun, Jike Wang, Chang-Yu Hsieh, Zhenxing Wu, Dong-Sheng Cao 0001, Jian Wu 0001, Tingjun Hou
Briefings Bioinform.8
2022 MODIG: integrating multi-omics and multi-dimensional gene network for cancer driver gene identification based on graph attention network model
abstract
MOTIVATION: Identifying genes that play a causal role in cancer evolution remains one of the biggest challenges in cancer biology. With the accumulation of high-throughput multi-omics data over decades, it becomes a great challenge to effectively integrate these data into the identification of cancer driver genes. RESULTS: Here, we propose MODIG, a graph attention network (GAT)-based framework to identify cancer driver genes by combining multi-omics pan-cancer data (mutations, copy number variants, gene expression and methylation levels) with multi-dimensional gene networks. First, we established diverse types of gene relationship maps based on protein-protein interactions, gene sequence similarity, KEGG pathway co-occurrence, gene co-expression patterns and gene ontology. Then, we constructed a multi-dimensional gene network consisting of approximately 20 000 genes as nodes and five types of gene associations as multiplex edges. We applied a GAT to model within-dimension interactions to generate a gene representation for each dimension based on this graph. Moreover, we introduced a joint learning module to fuse multiple dimension-specific representations to generate general gene representations. Finally, we used the obtained gene representation to perform a semi-supervised driver gene identification task. The experiment results show that MODIG outperforms the baseline models in terms of area under precision-recall curves and area under the receiver operating characteristic curves. AVAILABILITY AND IMPLEMENTATION: The MODIG program is available at https://github.com/zjupgx/modig. The code and data underlying this article are also available on Zenodo, at https://doi.org/10.5281/zenodo.7057241. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenyi Zhao, Xun Gu 0002, Jian Wu 0001, Zhan Zhou
Bioinform.4
2022 TGSA: protein-protein association-based twin graph neural networks for drug response prediction with similarity augmentation
abstract
MOTIVATION: Drug response prediction (DRP) plays an important role in precision medicine (e.g. for cancer analysis and treatment). Recent advances in deep learning algorithms make it possible to predict drug responses accurately based on genetic profiles. However, existing methods ignore the potential relationships among genes. In addition, similarity among cell lines/drugs was rarely considered explicitly. RESULTS: We propose a novel DRP framework, called TGSA, to make better use of prior domain knowledge. TGSA consists of Twin Graph neural networks for Drug Response Prediction (TGDRP) and a Similarity Augmentation (SA) module to fuse fine-grained and coarse-grained information. Specifically, TGDRP abstracts cell lines as graphs based on STRING protein-protein association networks and uses Graph Neural Networks (GNNs) for representation learning. SA views DRP as an edge regression problem on a heterogeneous graph and utilizes GNNs to smooth the representations of similar cell lines/drugs. Besides, we introduce an auxiliary pre-training strategy to remedy the identified limitations of scarce data and poor out-of-distribution generalization. Extensive experiments on the GDSC2 dataset demonstrate that our TGSA consistently outperforms all the state-of-the-art baselines under various experimental settings. We further evaluate the effectiveness and contributions of each component of TGSA via ablation experiments. The promising performance of TGSA shows enormous potential for clinical applications in precision medicine. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/violet-sto/TGSA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yiheng Zhu 0002, Zhenqiu Ouyang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001
Bioinform.7
2022 Detecting Duplicate Questions in Stack Overflow via Source Code Modeling
abstract
Stack Overflow is one of the most popular Question-Answering sites for programmers. However, it faces the problem of question duplication, where newly created questions are identical to previous questions. Existing works on duplicate question detection in Stack Overflow extract a set of textual features on the question pairs and use supervised learning approaches to classify duplicate question pairs. However, they do not consider the source code information in the questions. While in some cases, the intention of a question is mainly represented by the source code. In this paper, we aim to learn the semantics of a question by combining both text features and source code features. We use word embedding and convolutional neural networks to extract textual features from questions to overcome the lexical gap issue. We use tree-based convolutional neural networks to extract structural and semantic features from source code. In addition, we perform multi-task learning by combining the duplication question detection task with a question tag prediction side task. We conduct extensive experiments on the Stack Overflow dataset and show that our approach can detect duplicate questions with higher recall and MRR compared with baseline approaches on Python and Java programming languages.
Wei Gao 0001, Jian Wu 0001, Guandong Xu
Int. J. Softw. Eng. Knowl. Eng.2
2022 ChroNet: A multi-task learning based approach for prediction of multiple chronic diseases
Ruiwei Feng, Xuechen Liu 0004, Tingting Chen 0002, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
Multim. Tools Appl.8
2022 Discriminative Cervical Lesion Detection in Colposcopic Images With Global Class Activation and Local Bin Excitation
abstract
Accurate cervical lesion detection (CLD) methods using colposcopic images are highly demanded in computer-aided diagnosis (CAD) for automatic diagnosis of High-grade Squamous Intraepithelial Lesions (HSIL). However, compared to natural scene images, the specific characteristics of colposcopic images, such as low contrast, visual similarity, and ambiguous lesion boundaries, pose difficulties to accurately locating HSIL regions and also significantly impede the performance improvement of existing CLD approaches. To tackle these difficulties and better capture cervical lesions, we develop novel feature enhancing mechanisms from both global and local perspectives, and propose a new discriminative CLD framework, called CervixNet, with a Global Class Activation (GCA) module and a Local Bin Excitation (LBE) module. Specifically, the GCA module learns discriminative features by introducing an auxiliary classifier, and guides our model to focus on HSIL regions while ignoring noisy regions. It globally facilitates the feature extraction process and helps boost feature discriminability. Further, our LBE module excites lesion features in a local manner, and allows the lesion regions to be more fine-grained enhanced by explicitly modelling the inter-dependencies among bins of proposal feature. Extensive experiments on a number of 9888 clinical colposcopic images verify the superiority of our method (AP$_{.75}$= 20.45) over state-of-the-art models on four widely used metrics.
Tingting Chen 0002, Xuechen Liu 0004, Ruiwei Feng, Wenzhe Wang, Chunnv Yuan, Weiguo Lu, Haizhen He, Honghao Gao, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001
IEEE J. Biomed. Health Informatics11
2022 A Task Decomposing and Cell Comparing Method for Cervical Lesion Cell Detection
abstract
Automatic detection of cervical lesion cells or cell clumps using cervical cytology images is critical to computer-aided diagnosis (CAD) for accurate, objective, and efficient cervical cancer screening. Recently, many methods based on modern object detectors were proposed and showed great potential for automatic cervical lesion detection. Although effective, several issues still hinder further performance improvement of such known methods, such as large appearance variances between single-cell and multi-cell lesion regions, neglecting normal cells, and visual similarity among abnormal cells. To tackle these issues, we propose a new task decomposing and cell comparing network, called TDCC-Net, for cervical lesion cell detection. Specifically, our task decomposing scheme decomposes the original detection task into two subtasks and models them separately, which aims to learn more efficient and useful feature representations for specific cell structures and then improve the detection performance of the original task. Our cell comparing scheme imitates clinical diagnosis of experts and performs cell comparison with a dynamic comparing module (normal-abnormal cells comparing) and an instance contrastive loss (abnormal-abnormal cells comparing). Comprehensive experiments on a large cervical cytology image dataset confirm the superiority of our method over state-of-the-art methods.
Tingting Chen 0002, Haochao Ying, Xiangyu Tan, Danny Ziyi Chen, Jian Wu 0001
IEEE Trans. Medical Imaging8
2022 Reinforcement-Learning-Guided Source Code Summarization Using Hierarchical Attention
abstract
Code summarization (aka comment generation) provides a high-level natural language description of the function performed by code, which can benefit the software maintenance, code categorization and retrieval. To the best of our knowledge, the state-of-the-art approaches follow an encoder-decoder framework which encodes source code into a hidden space and later decodes it into a natural language space. Such approaches suffer from the following drawbacks: (a) they are mainly input by representing code as a sequence of tokens while ignoring code hierarchy; (b) most of the encoders only input simple features (e.g., tokens) while ignoring the features that can help capture the correlations between comments and code; (c) the decoders are typically trained to predict subsequent words by maximizing the likelihood of subsequent ground truth words, while in real world, they are excepted to generate the entire word sequence from scratch. As a result, such drawbacks lead to inferior and inconsistent comment generation accuracy. To address the above limitations, this paper presents a new code summarization approach using hierarchical attention network by incorporating multiple code features, including type-augmented abstract syntax trees and program control flows. Such features, along with plain code sequences, are injected into a deep reinforcement learning (DRL) framework (e.g., actor-critic network) for comment generation. Our approach assigns weights (pays “attention”) to tokens and statements when constructing the code representation to reflect the hierarchical code structure under different contexts regarding code features (e.g., control flows and abstract syntax trees). Our reinforcement learning mechanism further strengthens the prediction results through the actor network and the critic network, where the actor network provides the confidence of predicting subsequent words based on the current state, and the critic network computes the reward values of all the possible extensions of the current state to provide global guidance for explorations. Eventually, we employ an advantage reward to train both networks and conduct a set of experiments on a real-world dataset. The experimental results demonstrate that our approach outperforms the baselines by around 22 to 45 percent in BLEU-1 and outperforms the state-of-the-art approaches by around 5 to 60 percent in terms of S-BLEU and C-BLEU.
Yuqun Zhang, Yulei Sui, Yao Wan 0001, Zhou Zhao 0001, Jian Wu 0001, Philip S. Yu, Guandong Xu
IEEE Trans. Software Eng.6
2021 To Choose or to Fuse? Scale Selection for Crowd Counting
abstract
In this paper, we address the large scale variation problem in crowd counting by taking full advantage of the multi-scale feature representations in a multi-level network. We implement such an idea by keeping the counting error of a patch as small as possible with a proper feature level selection strategy, since a specific feature level tends to perform better for a certain range of scales. However, without scale annotations, it is sub-optimal and error-prone to manually assign the predictions for heads of different scales to specific feature levels. Therefore, we propose a Scale-Adaptive Selection Network (SASNet), which automatically learns the internal correspondence between the scales and the feature levels. Instead of directly using the predictions from the most appropriate feature level as the final estimation, our SASNet also considers the predictions from other feature levels via weighted average, which helps to mitigate the gap between discrete feature levels and continuous scale variation. Since the heads in a local patch share roughly a same scale, we conduct the adaptive selection strategy in a patch-wise style. However, pixels within a patch contribute different counting errors due to the various difficulty degrees of learning. Thus, we further propose a Pyramid Region Awareness Loss (PRA Loss) to recursively select the most hard sub-regions within a patch until reaching the pixel level. With awareness of whether the parent patch is over-estimated or under-estimated, the fine-grained optimization with the PRA Loss for these region-aware hard pixels helps to alleviate the inconsistency problem between training target and evaluation metric. The state-of-the-art results on four datasets demonstrate the superiority of our approach. The code will be available at: https://github.com/TencentYoutuResearch/CrowdCounting-SASNet.
Qingyu Song 0001, Changan Wang, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Jian Wu 0001, Jiayi Ma 0001
AAAI7
2021 AGMI: Attention-Guided Multi-omics Integration for Drug Response Prediction with Graph Neural Networks
abstract
Accurate drug response prediction (DRP) is a crucial yet challenging task in precision medicine. This paper presents a novel Attention-Guided Multi-omics Integration (AGMI) approach for DRP, which first constructs a Multiedge Graph (MeG) for each cell line, and then aggregates multi-omics features to predict drug response using a novel structure, called Graph edge-aware Network (GeNet). For the first time, our AGMI approach explores gene constraint based multi-omics integration for DRP with the whole-genome using GNNs. Empirical experiments on the CCLE and GDSC datasets show that our AGMI largely outperforms state-of-the-art DRP methods by 8.3%-34.2% on four metrics. Our data and code are available at https://github.com/yivan-WYYGDSG/AGMI.
Ruiwei Feng, Minshan Lai, Danny Ziyi Chen, Jian Wu 0001
BIBM6
2021 A Receptor Skeleton for Capsule Neural Networks
abstract
In previous Capsule Neural Networks (CapsNets), routing algorithms often performed clustering processes to assemble the child capsules’ representations into parent capsules. Such routing algorithms were typically implemented with iterative processes and incurred high computing complexity. This paper presents a new capsule structure, which contains a set of optimizable receptors and a transmitter is devised on the capsule’s representation. Specifically, child capsules’ representations are sent to the parent capsules whose receptors match well the transmitters of the child capsules’ representations, avoiding applying computationally complex routing algorithms. To ensure the receptors in a CapsNet work cooperatively, we build a skeleton to organize the receptors in different capsule layers in a CapsNet. The receptor skeleton assigns a share-out objective for each receptor, making the CapsNet perform as a hierarchical agglomerative clustering process. Comprehensive experiments verify that our approach facilitates efficient clustering processes, and CapsNets with our approach significantly outperform CapsNets with previous routing algorithms on image classification, affine transformation generalization, overlapped object recognition, and representation semantic decoupling.
Jintai Chen, Hongyun Yu, Chengde Qian, Danny Ziyi Chen, Jian Wu 0001
ICML5
2021 Electrocardio Panorama: Synthesizing New ECG views with Self-supervision
abstract
Multi-lead electrocardiogram (ECG) provides clinical information of heartbeats from several fixed viewpoints determined by the lead positioning. However, it is often not satisfactory to visualize ECG signals in these fixed and limited views, as some clinically useful information is represented only from a few specific ECG viewpoints. For the first time, we propose a new concept, Electrocardio Panorama, which allows visualizing ECG signals from any queried viewpoints. To build Electrocardio Panorama, we assume that an underlying electrocardio field exists, representing locations, magnitudes, and directions of ECG signals. We present a Neural electrocardio field Network (Nef-Net), which first predicts the electrocardio field representation by using a sparse set of one or few input ECG views and then synthesizes Electrocardio Panorama based on the predicted representations. Specially, to better disentangle electrocardio field information from viewpoint biases, a new Angular Encoding is proposed to process viewpoint angles. Also, we propose a self-supervised learning approach called Standin Learning, which helps model the electrocardio field without direct supervision. Further, with very few modifications, Nef-Net can synthesize ECG signals from scratch. Experiments verify that our Nef-Net performs well on Electrocardio Panorama synthesis, and outperforms the previous work on the auxiliary tasks (ECG view transformation and ECG synthesis from scratch). The codes and the division labels of cardiac cycles and ECG deflections on Tianchi ECG and PTB datasets are available at https://github.com/WhatAShot/Electrocardio-Panorama.
Jintai Chen, Xiangshang Zheng, Hongyun Yu, Danny Ziyi Chen, Jian Wu 0001
IJCAI5
2021 CanDriS: posterior profiling of cancer-driving sites based on two-component evolutionary model
abstract
Current cancer genomics databases have accumulated millions of somatic mutations that remain to be further explored. Due to the over-excess mutations unrelated to cancer, the great challenge is to identify somatic mutations that are cancer-driven. Under the notion that carcinogenesis is a form of somatic-cell evolution, we developed a two-component mixture model: while the ground component corresponds to passenger mutations, the rapidly evolving component corresponds to driver mutations. Then, we implemented an empirical Bayesian procedure to calculate the posterior probability of a site being cancer-driven. Based on these, we developed a software CanDriS (Cancer Driver Sites) to profile the potential cancer-driving sites for thousands of tumor samples from the Cancer Genome Atlas and International Cancer Genome Consortium across tumor types and pan-cancer level. As a result, we identified that approximately 1% of the sites have posterior probabilities larger than 0.90 and listed potential cancer-wide and cancer-specific driver mutations. By comprehensively profiling all potential cancer-driving sites, CanDriS greatly enhances our ability to refine our knowledge of the genetic basis of cancer and might guide clinical medication in the upcoming era of precision medicine. The results were displayed in a database CandrisDB (http://biopharm.zju.edu.cn/candrisdb/).
Wenyi Zhao, Jingcheng Wu, Guoxing Cai, Jeffrey Haltom, Weijia Su, Michael J. Dong, Jian Wu 0001, Zhan Zhou, Xun Gu 0002
Briefings Bioinform.10
2021 Cascaded SE-ResUnet for segmentation of thoracic organs at risk
Zheng Cao 0005, Bohan Yu, Biwen Lei, Haochao Ying, Xiao Zhang 0015, Danny Ziyi Chen, Jian Wu 0001
Neurocomputing7
2021 Enhancing session-based social recommendation through item graph embedding and contextual friendship modeling
Pan Gu, Yuqiang Han, Wei Gao 0001, Guandong Xu, Jian Wu 0001
Neurocomputing5
2021 A semi-supervised deep convolutional framework for signet ring cell detection
Haochao Ying, Qingyu Song 0004, Jintai Chen, Tingting Liang, Jingjing Gu, Fuzhen Zhuang, Danny Ziyi Chen, Jian Wu 0001
Neurocomputing8
2021 Radiographs and texts fusion learning based deep networks for skeletal bone age assessment
Pengyi Hao, Taotao Ye, Xuhang Xie, Fuli Wu, Wuheng Zuo, Wei Chen 0001, Jian Wu 0001
Multim. Tools Appl.8
2021 Multi-modality fusion learning for the automatic diagnosis of optic neuropathy
Zheng Cao 0005, Chuanbin Sun, Wenzhe Wang, Xiangshang Zheng, Jian Wu 0001, Honghao Gao
Pattern Recognit. Lett.5
2021 A Transfer Learning Based Super-Resolution Microscopy for Biopsy Slice Images: The Joint Methods Perspective
abstract
Higher-resolution biopsy slice images reveal many details, which are widely used in medical practice. However, taking high-resolution slice images is more costly than taking low-resolution ones. In this paper, we propose a joint framework containing a novel transfer learning strategy and a deep super-resolution framework to generate high-resolution slice images from low-resolution ones. The super-resolution framework called SRFBN+ is proposed by modifying a state-of-the-art framework SRFBN. Specifically, the structure of the feedback block of SRFBN was modified to be more flexible. Besides, it is challenging to use typical transfer learning strategies directly for the tasks on slice images, as the patterns on different types of biopsy slice images are varying. To this end, we propose a novel transfer learning strategy, called Channel Fusion Transfer Learning (CF-Trans). CF-Trans builds a middle domain by fusing the data manifolds of the source domain and the target domain, serving as a springboard for knowledge transfer. Thus, in the transfer learning setting, SRFBN+ can be trained on the source domain and then the middle domain and finally the target domain. Experiments on biopsy slice images validate SRFBN+ works well in generating super-resolution slice images, and CF-Trans is an efficient transfer learning strategy.
Jintai Chen, Haochao Ying, Xuechen Liu 0004, Jingjing Gu, Ruiwei Feng, Tingting Chen 0002, Honghao Gao, Jian Wu 0001
IEEE ACM Trans. Comput. Biol. Bioinform.8
2021 A Deep Learning Approach for Colonoscopy Pathology WSI Analysis: Accurate Segmentation and Classification
abstract
Colorectal cancer (CRC) is one of the most life-threatening malignancies. Colonoscopy pathology examination can identify cells of early-stage colon tumors in small tissue image slices. But, such examination is time-consuming and exhausting on high resolution images. In this paper, we present a new framework for colonoscopy pathology whole slide image (WSI) analysis, including lesion segmentation and tissue diagnosis. Our framework contains an improved U-shape network with a VGG net as backbone, and two schemes for training and inference, respectively (the training scheme and inference scheme). Based on the characteristics of colonoscopy pathology WSI, we introduce a specific sampling strategy for sample selection and a transfer learning strategy for model training in our training scheme. Besides, we propose a specific loss function, class-wise DSC loss, to train the segmentation network. In our inference scheme, we apply a sliding-window based sampling strategy for patch generation and diploid ensemble (data ensemble and model ensemble) for the final prediction. We use the predicted segmentation mask to generate the classification probability for the likelihood of WSI being malignant. To our best knowledge, DigestPath 2019 is the first challenge and the first public dataset available on colonoscopy tissue screening and segmentation, and our proposed framework yields good performance on this dataset. Our new framework achieved a DSC of 0.7789 and AUC of 1 on the online test dataset, and we won the [Formula: see text] place in the DigestPath 2019 Challenge (task 2). Our code is available at https://github.com/bhfs9999/colonoscopy_tissue_screen_and_segmentation.
Ruiwei Feng, Xuechen Liu 0004, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE J. Biomed. Health Informatics6
2021 KerNet: A Novel Deep Learning Approach for Keratoconus and Sub-Clinical Keratoconus Detection Based on Raw Data of the Pentacam HR System
abstract
Keratoconus is one of the most severe corneal diseases, which is difficult to detect at the early stage (i.e., sub-clinical keratoconus) and possibly results in vision loss. In this paper, we propose a novel end-to-end deep learning approach, called KerNet, which processes the raw data of the Pentacam HR system (consisting of five numerical matrices) to detect keratoconus and sub-clinical keratoconus. Specifically, we propose a novel convolutional neural network, called KerNet, containing five branches as the backbone with a multi-level fusion architecture. The five branches receive five matrices separately and capture effectively the features of different matrices by several cascaded residual blocks. The multi-level fusion architecture (i.e., low-level fusion and high-level fusion) moderately takes into account the correlation among five slices and fuses the extracted features for better prediction. Experimental results show that: (1) our novel approach outperforms state-of-the-art methods on an in-house dataset, by ~1% for keratoconus detection accuracy and ~4 for sub-clinical keratoconus detection accuracy; (2) the attention maps visualized by Grad-CAM show that our KerNet places more attention on the inferior temporal part for sub-clinical keratoconus, which has been proved as the identifying regions for ophthalmologists to detect sub-clinical keratoconus in previous clinical studies. To our best knowledge, we are the first to propose an end-to-end deep learning approach utilizing raw data obtained by the Pentacam HR system for keratoconus and subclinical keratoconus detection. Further, the prediction performance and the clinical significance of our KerNet are well evaluated and proved by two clinical experts. Our code is available at https://github.com/upzheng/Keratoconus.
Ruiwei Feng, Xiangshang Zheng, Heping Hu, Xiuming Jin, Danny Ziyi Chen, Ke Yao, Jian Wu 0001
IEEE J. Biomed. Health Informatics8
2021 Interactive Few-Shot Learning: Limited Supervision, Better Medical Image Segmentation
abstract
Many known supervised deep learning methods for medical image segmentation suffer an expensive burden of data annotation for model training. Recently, few-shot segmentation methods were proposed to alleviate this burden, but such methods often showed poor adaptability to the target tasks. By prudently introducing interactive learning into the few-shot learning strategy, we develop a novel few-shot segmentation approach called Interactive Few-shot Learning (IFSL), which not only addresses the annotation burden of medical image segmentation models but also tackles the common issues of the known few-shot segmentation methods. First, we design a new few-shot segmentation structure, called Medical Prior-based Few-shot Learning Network (MPrNet), which uses only a few annotated samples (e.g., 10 samples) as support images to guide the segmentation of query images without any pre-training. Then, we propose an Interactive Learning-based Test Time Optimization Algorithm (IL-TTOA) to strengthen our MPrNet on the fly for the target task in an interactive fashion. To our best knowledge, our IFSL approach is the first to allow few-shot segmentation models to be optimized and strengthened on the target tasks in an interactive and controllable manner. Experiments on four few-shot segmentation tasks show that our IFSL approach outperforms the state-of-the-art methods by more than 20% in the DSC metric. Specifically, the interactive optimization algorithm (IL-TTOA) further contributes ~10% DSC improvement for the few-shot segmentation models.
Ruiwei Feng, Xiangshang Zheng, Tianxiang Gao, Jintai Chen, Wenzhe Wang, Danny Ziyi Chen, Jian Wu 0001
IEEE Trans. Medical Imaging7
2021 Aspect-level sentiment capsule network for micro-video click-through rate prediction
Yuqiang Han, Pan Gu, Wei Gao 0001, Guandong Xu, Jian Wu 0001
World Wide Web5
2021 Multiple interleaving interests modeling of sequential user behaviors in e-commerce platform
Yuqiang Han, Qian Li 0016, Hucheng Zhou, Zhenglu Yang, Jian Wu 0001
World Wide Web6
2020 Flow-Mixup: Classifying Multi-labeled Medical Images with Corrupted Labels
abstract
In clinical practice, medical image interpretation often involves multi-labeled classification, since the affected parts of a patient tend to present multiple symptoms or comorbidities. Recently, deep learning based frameworks have attained expertlevel performance on medical image interpretation, which can be attributed partially to large amounts of accurate annotations. However, manually annotating massive amounts of medical images is impractical, while automatic annotation is fast but imprecise (possibly introducing corrupted labels). In this work, we propose a new regularization approach, called Flow-Mixup, for multi-labeled medical image classification with corrupted labels. Flow-Mixup guides the models to capture robust features for each abnormality, thus helping handle corrupted labels effectively and making it possible to apply automatic annotation. Specifically, Flow-Mixup decouples the extracted features by adding constraints to the hidden states of the models. Also, FlowMixup is more stable and effective comparing to other known regularization methods, as shown by theoretical and empirical analyses. Experiments on two electrocardiogram datasets and a chest X-ray dataset containing corrupted labels verify that FlowMixup is effective and insensitive to corrupted labels.
Jintai Chen, Hongyun Yu, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001
BIBM5
2020 A Hierarchical Graph Network for 3D Object Detection on Point Clouds
abstract
3D object detection on point clouds finds many applications. However, most known point cloud object detection methods did not adequately accommodate the characteristics (e.g., sparsity) of point clouds, and thus some key semantic information (e.g., shape information) is not well captured. In this paper, we propose a new graph convolution (GConv) based hierarchical graph network (HGNet) for 3D object detection, which processes raw point clouds directly to predict 3D bounding boxes. HGNet effectively captures the relationship of the points and utilizes the multi-level semantics for object detection. Specially, we propose a novel shape-attentive GConv (SA-GConv) to capture the local shape features, by modelling the relative geometric positions of points to describe object shapes. An SA-GConv based U-shape network captures the multi-level features, which are mapped into an identical feature space by an improved voting module and then further utilized to generate proposals. Next, a new GConv based Proposal Reasoning Module reasons on the proposals considering the global scene semantics, and the bounding boxes are then predicted. Consequently, our new framework outperforms state-of-the-art methods on two large-scale point cloud datasets, by ~4% mean average precision (mAP) on SUN RGB-D and by ~3% mAP on ScanNet-V2.
Jintai Chen, Biwen Lei, Qingyu Song 0004, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001
CVPR6
2020 UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation
abstract
Recently, a growing interest has been seen in deep learning-based semantic segmentation. UNet, which is one of deep learning networks with an encoder-decoder architecture, is widely used in medical image segmentation. Combining multi-scale features is one of important factors for accurate segmentation. UNet++ was developed as a modified Unet by designing an architecture with nested and dense skip connections. However, it does not explore sufficient information from full scales and there is still a large room for improvement. In this paper, we propose a novel UNet 3+, which takes advantage of full-scale skip connections and deep supervisions. The full-scale skip connections incorporate low-level details with high-level semantics from feature maps in different scales; while the deep supervision learns hierarchical representations from the full-scale aggregated feature maps. The proposed method is especially benefiting for organs that appear at varying scales. In addition to accuracy improvements, the proposed UNet 3+ can reduce the network parameters to improve the computation efficiency. We further propose a hybrid loss function and devise a classification-guided module to enhance the organ boundary and reduce the over-segmentation in a non-organ image, yielding more accurate segmentation results. The effectiveness of the proposed method is demonstrated on two datasets. The code is available at: github.com/ZJUGiveLab/UNet-Version.
Huimin Huang 0002, Lanfen Lin, Ruofeng Tong 0001, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Jian Wu 0001
ICASSP9
2020 Doctor Imitator: A Graph-Based Bone Age Assessment Framework Using Hand Radiographs
Jintai Chen, Bohan Yu, Biwen Lei, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001
MICCAI (6)6
2020 Dual-Level Selective Transfer Learning for Intrahepatic Cholangiocarcinoma Segmentation in Non-enhanced Abdominal CT
Wenzhe Wang, Qingyu Song 0004, Jiarong Zhou, Ruiwei Feng, Tingting Chen 0002, Wenhao Ge, Danny Ziyi Chen, Shaohua Kevin Zhou, Jian Wu 0001
MICCAI (1)10
2020 Texture branch network for chronic kidney disease screening based on ultrasound images
abstract
Chronic kidney disease (CKD) is a widespread renal disease throughout the world. Once it develops to the advanced stage, serious complications and high risk of death will follow. Hence, early screening is crucial for the treatment of CKD. Since ultrasonography has no side effects and enables radiologists to dynamically observe the morphology and pathological features of the kidney, it is commonly used for kidney examination. In this study, we propose a novel convolutional neural network (CNN) framework named the texture branch network to screen CKD based on ultrasound images. This introduces a texture branch into a typical CNN to extract and optimize texture features. The model can automatically generate texture features and deep features from input images, and use the fused information as the basis of classification. Furthermore, we train the base part of the network by means of transfer learning, and conduct experiments on a dataset with 226 ultrasound images. Experimental results demonstrate the effectiveness of the proposed approach, achieving an accuracy of 96.01% and a sensitivity of 99.44%.
Pengyi Hao, Shu-yuan Tian, Fuli Wu, Wei Chen 0001, Jian Wu 0001
Frontiers Inf. Technol. Electron. Eng.6
2020 CAMAR: a broad learning based context-aware recommender for mobile applications
Tingting Liang, Lifang He 0001, Chun-Ta Lu, Liang Chen 0001, Haochao Ying, Philip S. Yu, Jian Wu 0001
Knowl. Inf. Syst.7
2020 Multi-view factorization machines for mobile app recommendation based on hierarchical attention
Tingting Liang, Lei Zheng 0001, Liang Chen 0001, Yao Wan 0001, Philip S. Yu, Jian Wu 0001
Knowl. Based Syst.6
2020 Fast and Efficient Facial Expression Recognition Using a Gabor Convolutional Network
abstract
Automatic facial expression recognition (FER) is a fundamental topic in computer vision. Many studies have indicated that facial emotion changes are strongly related to certain regions of interest (ROIs), such as the mouth, eyes, eyebrows, and nose; therefore, the features of these facial ROIs are very important for identifying expressions. Since Gabor filters are very efficient in extracting visual content, Gabor orientation filters (GoFs) modulated by Gabor kernels and traditional convolutional filters can capture such ROI information better than conventional convolutional filters. Consequently, this letter presents a light Gabor convolutional network (GCN) consisting of only four Gabor convolutional layers and two linear layers for FER tasks. Extensive experiments on the FER2013, FERPlus and Real-world Affective Faces (RAF) databases demonstrate that the proposed method achieves good recognition accuracy and requires very low computational costs. The source code can be found at https://github.com/general515/Facial_Expression_Recognition_Using _GCN.
Ping Jiang 0004, Bo Wan 0002, Quan Wang 0006, Jian Wu 0001
IEEE Signal Process. Lett.4
2019 A Dual-Attention Dilated Residual Network for Liver Lesion Classification and Localization on CT Images
abstract
Automatic liver lesion classification on computed tomography images is of great importance to early cancer diagnosis and remains a challenging task. State-of-the-art liver lesion classification algorithms are currently based on manually selected regions of interest (ROIs) or automatically detected ROIs. However, liver lesions usually vary in size and shape, which makes the ROI selection process labor-intensive and also poses an obstacle to automatic lesion detection. In this paper, we propose a dual-attention dilated residual network (DADRN) as a potential solution to lesion classification task without manual ROI selection or automatic lesion detection. We incorporated a novel dual-attention module in order to capture the non-local feature dependencies and help the deep neural network focus on the lesion area by enlarging the difference between the lesion area and nonlesion area. To the best of our knowledge, we are the first to employ the self-attention mechanism to address liver lesion classification task. In addition, the well-trained DADRN can be used for weakly-supervised lesion localization without any architectural change or retraining. Experiment results show that DADRN could achieve a lesion classification accuracy comparable to that of the state-of-the-art ROI-based method and outperformed state-of-the-art attention-based approaches in both liver lesion classification and localization tasks.
Xiao Chen 0016, Jian Wu 0001, Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001
ICIP2
2019 Multi-Stream Scale-Insensitive Convolutional and Recurrent Neural Networks for Liver Tumor Detection in Dynamic Ct Images
abstract
Convolutional neural networks (CNNs) have achieved great success in numerous challenging vision tasks, and have great potential for object detection in natural images. Compared with the natural images, medical images exhibit some unique characteristics. Therefore, substantial challenges still remain in this field. The first challenge is to develop a method for effectively distilling enhancement patterns from the dynamic CT images. Moreover, since tumor sizes vary greatly and small lesions are important for early liver tumor detection, lesion detection with a widely variable scale is another challenge. In this paper, we propose a multi-stream scale-insensitive convolutional and recurrent neural network (MSCR) for liver tumor detection. Specifically, we propose the use of grouped convolutional long short-term memory (GCLSTM) to extract enhancement patterns, which is developed as a plug-and-play module. Experiments show that the MSCR framework exhibits superior performance over state-of-the-art approaches, achieving an average precision of 77.06% for detection of focal liver lesions. We have released the code of MSCR in1.
Ruofeng Tong 0001, Jian Wu 0001, Lanfen Lin, Xiao Chen 0016, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001
ICIP3
2019 Multi-modal Attention Network Learning for Semantic Source Code Retrieval
abstract
Code retrieval techniques and tools have been playing a key role in facilitating software developers to retrieve existing code fragments from available open-source repositories given a user query (e.g., a short natural language text describing the functionality for retrieving a particular code snippet). Despite the existing efforts in improving the effectiveness of code retrieval, there are still two main issues hindering them from being used to accurately retrieve satisfiable code fragments from large-scale repositories when answering complicated queries. First, the existing approaches only consider shallow features of source code such as method names and code tokens, but ignoring structured features such as abstract syntax trees (ASTs) and control-flow graphs (CFGs) of source code, which contains rich and well-defined semantics of source code. Second, although the deep learning-based approach performs well on the representation of source code, it lacks the explainability, making it hard to interpret the retrieval results and almost impossible to understand which features of source code contribute more to the final results. To tackle the two aforementioned issues, this paper proposes MMAN, a novel Multi-Modal Attention Network for semantic source code retrieval. A comprehensive multi-modal representation is developed for representing unstructured and structured features of source code, with one LSTM for the sequential tokens of code, a Tree-LSTM for the AST of code and a GGNN (Gated Graph Neural Network) for the CFG of code. Furthermore, a multi-modal attention fusion layer is applied to assign weights to different parts of each modality of source code and then integrate them into a single hybrid representation. Comprehensive experiments and analysis on a large-scale real-world dataset show that our proposed model can accurately retrieve code snippets and outperforms the state-of-the-art methods.
Yao Wan 0001, Jingdong Shu, Yulei Sui, Guandong Xu, Zhou Zhao 0001, Jian Wu 0001, Philip S. Yu
ASE6
2019 Multi-view Learning with Feature Level Fusion for Cervical Dysplasia Diagnosis
Tingting Chen 0002, Xinjun Ma, Xuechen Liu 0004, Wenzhe Wang, Ruiwei Feng, Jintai Chen, Chunnv Yuan, Weiguo Lu, Danny Ziyi Chen, Jian Wu 0001
MICCAI (1)10
2019 LSRC: A Long-Short Range Context-Fusing Framework for Automatic 3D Vertebra Localization
Jintai Chen, Ruoqian Guo, Bohan Yu, Tingting Chen 0002, Wenzhe Wang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001
MICCAI (6)9
2019 Semi-supervised Segmentation of Liver Using Adversarial Learning with Deep Atlas Prior
Lanfen Lin, Hongjie Hu, Qiaowei Zhang, Qingqing Chen 0001, Yutaro Iwamoto, Xianhua Han, Yen-Wei Chen 0001, Ruofeng Tong 0001, Jian Wu 0001
MICCAI (6)10
2019 SMS: A Framework for Service Discovery by Incorporating Social Media Information
abstract
With the explosive growth of services, including Web services, cloud services, APIs and mashups, discovering the appropriate services for consumers is becoming an imperative issue. The traditional service discovery approaches mainly face two challenges: 1) the single source of description documents limits the effectiveness of discovery due to the insufficiency of semantic information; 2) more factors should be considered with the generally increasing functional and nonfunctional requirements of consumers. In this paper, we propose a novel framework, called SMS, for effectively discovering the appropriate services by incorporating social media information. Specifically, we present different methods to measure four social factors (semantic similarity, popularity, activity, decay factor) collected from Twitter. Latent Semantic Indexing (LSI) model is applied to mine semantic information of services from meta-data of Twitter Lists that contains them. In addition, we assume the target query-service matching function as a linear combination of multiple social factors and design a weight learning algorithm to learn an optimal combination of the measured social factors. Comprehensive experiments based on a real-world dataset crawled from Twitter demonstrate the effectiveness of the proposed framework SMS, through some compared approaches.
Tingting Liang, Liang Chen 0001, Jian Wu 0001, Guandong Xu, Zhaohui Wu 0001
IEEE Trans. Serv. Comput.3
2019 Time-aware metric embedding with asymmetric projection for successive POI recommendation
Haochao Ying, Jian Wu 0001, Guandong Xu, Yanchi Liu, Tingting Liang, Xiao Zhang 0015, Hui Xiong 0001
World Wide Web2
2018 Improved Dynamic Memory Network for Dialogue Act Classification with Adversarial Training
abstract
Dialogue Act (DA) classification is a challenging problem in dialogue interpretation, which aims to attach semantic labels to utterances and characterize the speaker's intention. Currently, many existing approaches formulate the DA classification problem ranging from multi-classification to structured prediction, which suffer from two limitations: a) these methods are either handcrafted feature-based or have limited memories. b) adversarial examples can't be correctly classified by traditional training methods. To address these issues, in this paper we first cast the problem into a question and answering problem and proposed an improved dynamic memory networks with hierarchical pyramidal utterance encoder. Moreover, we apply adversarial training to train our proposed model. We evaluate our model on two public datasets, i.e., Switchboard dialogue act corpus and the MapTask corpus. Extensive experiments show that our proposed model is not only robust, but also achieves better performance when compared with some state-of-the-art baselines.
Yao Wan 0001, Wenqiang Yan, Jianwei Gao, Zhou Zhao 0001, Jian Wu 0001, Philip S. Yu
IEEE BigData5
2018 Sequential Recommender System based on Hierarchical Attention Networks
abstract
With a large amount of user activity data accumulated, it is crucial to exploit user sequential behavior for sequential recommendations. Conventionally, user general taste and recent demand are combined to promote recommendation performances. However, existing methods often neglect that user long-term preference keep evolving over time, and building a static representation for user general taste may not adequately reflect the dynamic characters. Moreover, they integrate user-item or item-item interactions through a linear way which limits the capability of model. To this end, in this paper, we propose a novel two-layer hierarchical attention network, which takes the above properties into account, to recommend the next item user might be interested. Specifically, the first attention layer learns user long-term preferences based on the historical purchased item representation, while the second one outputs final user representation through coupling user long-term and short-term preferences. The experimental study demonstrates the superiority of our method compared with other state-of-the-art ones.
Haochao Ying, Fuzhen Zhuang, Yanchi Liu, Guandong Xu, Xing Xie 0001, Hui Xiong 0001, Jian Wu 0001
IJCAI8
2018 Improving automatic source code summarization via deep reinforcement learning
abstract
Code summarization provides a high level natural language description of the function performed by code, as it can benefit the software maintenance, code categorization and retrieval. To the best of our knowledge, most state-of-the-art approaches follow an encoder-decoder framework which encodes the code into a hidden space and then decode it into natural language space, suffering from two major drawbacks: a) Their encoders only consider the sequential content of code, ignoring the tree structure which is also critical for the task of code summarization; b) Their decoders are typically trained to predict the next word by maximizing the likelihood of next ground-truth word with previous ground-truth word given. However, it is expected to generate the entire sequence from scratch at test time. This discrepancy can cause an exposure bias issue, making the learnt decoder suboptimal. In this paper, we incorporate an abstract syntax tree structure as well as sequential content of code snippets into a deep reinforcement learning framework (i.e., actor-critic network). The actor network provides the confidence of predicting the next word according to current state. On the other hand, the critic network evaluates the reward value of all possible extensions of the current state and can provide global guidance for explorations. We employ an advantage reward composed of BLEU metric to train both networks. Comprehensive experiments on a real-world dataset show the effectiveness of our proposed model when compared with some state-of-the-art methods.
Yao Wan 0001, Zhou Zhao 0001, Min Yang 0007, Guandong Xu, Haochao Ying, Jian Wu 0001, Philip S. Yu
ASE6
2018 A Framework for Identifying Diabetic Retinopathy Based on Anti-noise Detection and Attention-Based Fusion
Zhiwen Lin, Ruoqian Guo, Tingting Chen 0002, Wenzhe Wang, Danny Ziyi Chen, Jian Wu 0001
MICCAI (2)8
2018 Deep Active Self-paced Learning for Accurate Pulmonary Nodule Segmentation
Wenzhe Wang, Tingting Chen 0002, Danny Ziyi Chen, Jian Wu 0001
MICCAI (2)6
2018 Exploiting cross-source knowledge for warming up community question answering services
Yao Wan 0001, Guandong Xu, Liang Chen 0001, Zhou Zhao 0001, Jian Wu 0001
Neurocomputing5
2018 SCSMiner: mining social coding sites for software developer recommendation with relevance propagation
Yao Wan 0001, Liang Chen 0001, Guandong Xu, Zhou Zhao 0001, Jie Tang 0001, Jian Wu 0001
World Wide Web6
2017 A Broad Learning Approach for Context-Aware Mobile Application Recommendation
abstract
With the rapid development of mobile apps, the availability of a large number of mobile apps in application stores brings challenges to locate appropriate apps for users. Providing accurate mobile app recommendation for users becomes an imperative task. Conventional approaches mainly focus on learning users' preferences and app features to predict the user-app ratings. However, most of them did not consider the interactions among the context information of apps. To address this issue, we propose a broad learning approach for Context-Aware app recommendation with Tensor Analysis (CATA). Specifically, we utilize a tensor-based framework to effectively integrate app category information and multi-view features on users and apps, respectively, to facilitate the performance of rating prediction. The multidimensional structure is employed to capture the hidden relationships among the app categories and the multiview features. We develop an efficient factorization method which applies Tucker decomposition to learn the full-order interactions among the app categories and features. Furthermore, we employ a group ℓ1-norm regularization to learn the group-wise feature importance of each view with respect to each app category. Experiments on a real-world mobile app dataset demonstrate the effectiveness of the proposed method.
Tingting Liang, Lifang He 0001, Chun-Ta Lu, Liang Chen 0001, Philip S. Yu, Jian Wu 0001
ICDM6
2017 A Novel Framework for Service Set Recommendation in Mashup Creation
abstract
With an overwhelming number of web services online, recommending services for automatic mashup creation greatly facilitates the composition process of developers. Various approaches have been proposed for the task. However, these approaches concentrate on improving the recommending accuracy of an individual service, which give rise to two problems: (1) Top-ranked services may be highly redundant with the same functionality, and (2) The cooperation relations among services are ignored. Therefore, we argue that services should be recommended not individually, but collectively. In this paper, we focus on the problem of recommending service sets instead of services. A service set contains a list of functionally distinct services that collectively match different aspects of functional requirements and are more inclined to compose together following mashup composition patterns. To this end, we propose a novel recommendation framework consisting of two stages: Service Set Generation Stage and Service Set Ranking Stage. We also perform an experimental evaluation on ProgrammableWeb dataset to demonstrate the effectiveness of our framework.
Wei Gao 0001, Jian Wu 0001
ICWS2
2017 Mobile Application Rating Prediction via Feature-Oriented Matrix Factorization
abstract
With the proliferation of mobile application (app) markets (e.g., Google Play, Apple App Store), predicting user preferences on apps becomes a challenging problem. Different from previous work, we assume that a user likes an app because he/she likes certain features of the app (e.g., permission, genre, topic). Based on this assumption, we propose a feature-oriented approach to predict user preferences on apps. Specifically, we transform the original app rating matrix to feature rating data and predict the unknown ratings on the features through a latent factor model, instead of directly predicting ratings on apps. The predicted user ratings on features can be used to generate the ratings on apps. Two integration methods are presented to give different significance for feature preferences. The approach has some obvious advantages: as it integrates feature information to analyze the details of user preference, it can generalize better as the feature rating data is denser, and improve the interpretation of the prediction of app ratings. Experimental results on a real-world dataset demonstrate the effectiveness of the proposed approach.
Tingting Liang, Liang Chen 0001, Xingde Ying, Philip S. Yu, Jian Wu 0001, Zibin Zheng
ICWS5
2017 Exploiting Geographical Location for Team Formation in Social Coding Sites
Yuqiang Han, Yao Wan 0001, Liang Chen 0001, Guandong Xu, Jian Wu 0001
PAKDD (1)5
2016 Personalized API Recommendation via Implicit Preference Modeling
Wei Gao 0001, Liang Chen 0001, Jian Wu 0001, Hai Dong 0001, Athman Bouguettaya
ICSOC3
2016 Meta-Path Based Service Recommendation in Heterogeneous Information Networks
Tingting Liang, Liang Chen 0001, Jian Wu 0001, Hai Dong 0001, Athman Bouguettaya
ICSOC3
2016 Joint Modeling Users, Services, Mashups, and Topics for Service Recommendation
abstract
With an increasing number of Web services on-line, service recommendation becomes an important approach to help user discover suitable services. However, current service recommendation approaches only consider user's historical preferences and functional or non-functional properties of services. The compositional side of services to form mashups is overlooked. However, the service composition information is a valuable information for recommendation since it gives an implicit measure of relatedness of services. In this paper, we recognize the importance of mashups on service recommendation and propose to jointly model historical preference of users on mashups and services, compositional information of services as mashups, and functional information of services and mashups in a single framework. Specifically, we extract topics from functional descriptions of services and model the relations between user, mashup and service and topics as a quadripartite graph. We use a personalized ranking on graph algorithm to learn the proposed model and simultaneously recommend services and mashups for user. We conduct a comprehensive experimental study using real world data from ProgrammbleWeb and results show that our recommendation method outperforms other representative recommendation approaches.
Wei Gao 0001, Liang Chen 0001, Jian Wu 0001, Athman Bouguettaya
ICWS3
2016 Exploiting Heterogeneous Information for Tag Recommendation in API Management
abstract
As web-enabled software becomes the standard for business processes, the ways organizations, partners and customers interface with it have become a critical differentiator in the market place, i.e., API Economy. With the rapid proliferation of APIs, it is increasingly important for users to effectively manage objective APIs in kinds of API markets, e.g., ProgramableWeb (PW), Mashape, etc. In this paper, to facilitate the process of API management, we propose a graphbased recommendation approach called ATRec to automatically assign tags to unlabeled APIs by exploiting both graph structure information and semantic similarity. Specifically, ATRec first leverages the multi-type relations (i.e., among APIs, mashups, and mashup assigned tags) to construct a heterogeneous network, in which a Random Walk with Restart (RWR) model is applied to alleviate the total cold start problem where no API has ever been tagged. Furthermore, we apply the recommended API tags in two API management scenarios (API search, API recommendation). Comprehensive experiments based on a real dataset crawled from PW demonstrate the effectiveness of the proposed approach.
Tingting Liang, Liang Chen 0001, Jian Wu 0001, Athman Bouguettaya
ICWS3
2016 Incorporating Heterogeneous Information for Mashup Discovery with Consistent Regularization
Yao Wan 0001, Liang Chen 0001, Qi Yu 0001, Tingting Liang, Jian Wu 0001
PAKDD (1)5
2016 Collaborative Deep Ranking: A Hybrid Pair-Wise Recommendation Algorithm with Implicit Feedback
Haochao Ying, Liang Chen 0001, Yuwen Xiong, Jian Wu 0001
PAKDD (2)4
2016 Temporal Pattern Based QoS Prediction
Liang Chen 0001, Haochao Ying, Qibo Qiu, Jian Wu 0001, Hai Dong 0001, Athman Bouguettaya
WISE (2)4
2016 Modern Service Industry and Crossover Services: Development and Trends in China
abstract
Modern service industry (MSI) is an information and knowledge intensive service industry, relying on information technology and modern management philosophy. The development of MSI is of high significance for promoting rapid global economy, accelerating social progress, and building an innovation-oriented society and harmonious realm. This paper first presents the history and trends of MIS in China in terms of the worldwide MSI development status. It then proposes the concept of Crossover Services based upon a survey of 62 MSI-related listed firms in China, whose business models, products, and services evidently exploit the concept. After elaborating on the crossover, convergence, and complex characteristics of crossover services, it proposes a technical framework that facilities addressing the scientific issues and technical challenges on crossover service realization. Finally, it illustrates how a new cloud-based middleware platform, named JTang++, supports the realization of crossover services.
Zhaohui Wu 0001, Jianwei Yin, Shuiguang Deng, Jian Wu 0001, Ying Li 0001, Liang Chen 0001
IEEE Trans. Serv. Comput.4
2015 CASE: A Platform for Crowdsourcing Based API Search
Tingting Liang, Liang Chen 0001, Zhining Xie, Jian Wu 0001
ICSOC5
2015 WS-HFS: A Heterogeneous Feature Selection Framework for Web Services Mining
abstract
With the development of Service Computing and Big Data research, more and more heterogeneous data generated in the process of Service Computing attracts our attention. Combining correlated data sources may help improve the performance of a given task. For example, in service recommendation, one can combine (1) user profile data (e.g. Genders, age, etc.), (2) user log data (e.g., Click through data, service invocation records, etc.), (3) QoS data (e.g. Response time, cost, etc.), (4) service functional description (e.g., Service name, WSDL document, etc.) and (5) service tagging data (i.e., Tags annotated by users) to build a recommendation model. All these data sources provide informative but heterogeneous features. For instance, user profile and QoS data usually have nominal features reflecting users' background and services' qualities, log data provides term-based features about users' historical behaviors, and service functional description and tagging data have term-based features reflecting services' functionalities and users' collective opinions. Given multiple heterogeneous data sources, one important challenge is to find a unified feature subspace to capture the knowledge from all data sources. To handle this problem, in this paper, we propose a Heterogeneous Feature Selection framework, named as WS-HFS, in which the consensus and the weight of different sources are both considered. Moreover, we apply the proposed framework to Web service clustering as a case study, and compare it with the state of the art approaches. The comprehensive experiments based on real data demonstrate the effectiveness of WS-HFS.
Liang Chen 0001, Qi Yu 0001, Philip S. Yu, Jian Wu 0001
ICWS4
2015 Manifold-Learning Based API Recommendation for Mashup Creation
abstract
With the wide adoption of Service-Oriented Architecture (SOA), the number of web accessible services and their compositions is increasing rapidly. Among huge number of services, how to recommend appropriate ones for automatic composition satisfying users' need is challenging. We investigate services and their compositions in Programmable Web which characterize services as APIs and their compositions as mashups. We study the problem of recommending suitable APIs satisfying users' need for mash up creation. To this end, we propose a manifold ranking framework for API recommendation. First, we categorize existing mashups into functionally similar clusters. Then we recommend APIs for each mash up cluster using manifold ranking algorithm which incorporate the relationships between mashups, between APIs and between mashups and APIs. Intuitively, we take three factors into consideration: (1) We recommend APIs that are in functionally similar mashups. (2) We recommend APIs that are popular in the mashups. (3) We recommend APIs that are similar to each other. Finally, we map a user's requirement for mash up creation to a mash up cluster and recommend APIs generated by the algorithm to user. Experiments based on real dataset crawled from Programmble Web demonstrate the effectiveness of the proposed approach in terms of precision, recall, and NDCG.
Wei Gao 0001, Liang Chen 0001, Jian Wu 0001, Honghao Gao
ICWS3
2015 Time-Aware API Popularity Prediction via Heterogeneous Features
abstract
Application Programming Interfaces (APIs), which are emerging web services in general, are increasing with a rapid speed in recent years. With so many APIs, many management platforms have been developed and deployed, leading to the boom of API markets, that are similar to the mobile App markets. Meanwhile, it has become more and more difficult to select and manage APIs. In reality, most existing management platforms typically recommend currently popular APIs to developers. However, the fact that popularity of API varies over time is ignored in those platforms, leading to the difficulty of recommending APIs that are just released but may be popular in the near future. To tackle this challenge, an approach of predicting the popularity of APIs is proposed in this paper. Predicting the popularity of API can not only be used for API ranking, recommendation and selection, but also make it more convenient for API providers and consumers to manage or select API respectively. In this paper, we propose a time-aware linear model to predict the API popularity, using time series feature of APIs and API's self-features such as its' provider ranking and description features, which are called heterogeneous features in our paper. Comprehensive experiments have been conducted on a real-world Programmable Web dataset with 613 real APIs. The experimental results show that our model has a better performance, when compared with some other state-of-the-art prediction models.
Yao Wan 0001, Liang Chen 0001, Jian Wu 0001, Qi Yu 0001
ICWS3
2015 Trust-aware media recommendation in heterogeneous social networks
Jian Wu 0001, Liang Chen 0001, Qi Yu 0001, Panpan Han, Zhaohui Wu 0001
World Wide Web1
2014 Data Augmented Maximum Margin Matrix Factorization for Flickr Group Recommendation
Liang Chen 0001, Yilun Wang 0001, Tingting Liang, Lichuan Ji, Jian Wu 0001
PAKDD (1)5
2014 SLQ: a user-friendly graph querying system
abstract
Querying complex graph databases such as knowledge graphs is a challenging task for non-professional users. In this demo, we present SLQ, a user-friendly graph querying system enabling schemales and structures graph querying, where a user need not describe queries precisely as required by most databases. SLQ system combines searching and ranking: it leverages a set of transformation functions, including abbreviation, ontology, synonym, etc., that map keywords and linkages from a query to their matches in a data graph, based on an automatically learned ranking model. To help users better understand search results at different levels of granularity, it supports effective result summarization with "drill-down" and "roll-up" operations. Better still, the architecture of SLQ is elastic for new transformation functions, query logs and user feedback, to iteratively refine the ranking model. SLQ significantly improves the usability of graph querying. This demonstration highlights (1) SLQ can automatically learn an effective ranking model, without assuming manually labeled training examples, (2) it can efficiently return top ranked matches over noisy, large data graphs, (3) it can summarize the query matches to help users easily access, explore and understand query results, and (4) its GUI can interact with users to help them construct queries, explore data graphs and inspect matches in a user-friendly manner.
Shengqi Yang, Yanan Xie, Yinghui Wu 0001, Huan Sun 0001, Jian Wu 0001, Xifeng Yan
SIGMOD Conference6
2014 Trust-Based Personalized Service Recommendation: A Network Perspective
Shuiguang Deng, Longtao Huang, Jian Wu 0001, Zhaohui Wu 0001
J. Comput. Sci. Technol.3
2014 Modeling and exploiting tag relevance for Web service mining
Liang Chen 0001, Jian Wu 0001, Zibin Zheng, Michael R. Lyu, Zhaohui Wu 0001
Knowl. Inf. Syst.2
2014 Clustering Web services to facilitate service discovery
Jian Wu 0001, Liang Chen 0001, Zibin Zheng, Michael R. Lyu, Zhaohui Wu 0001
Knowl. Inf. Syst.1
2014 Instant Recommendation for Web Services Composition
abstract
Web service composition helps users integrate services to create new large-granularity and value-added composite services. Most recent studies have focused on automatic AI-Planning-based static or dynamic composition at functional- or process-level. However in industry, most business applications are still composed manually or semi-automatically with abundant domain expertise. Consequently, to build a good and reliable composite service is really a time-consuming and professional task. Inspired by the Instant Search of Google, we propose an Instant recommendation approach to provide optimal suggestions while a composition process incrementally proceeds. In our model, we fully utilize the execution log of composite services, and intend to identify appropriate services which have been proved to be more reliable and robust, therefore those services have higher probability to fulfill users' demands. To find the top-k possible composite services in real-time, we adopt the A* search algorithm with various pruning heuristics to dynamically expand the search space efficiently. Experiments on a real-world dataset with 15,959 real Web services crawled from the Internet demonstrate the effectiveness and efficiency of the proposed approach.
Liang Chen 0001, Jian Wu 0001, Hengyi Jian, Hongbo Deng, Zhaohui Wu 0001
IEEE Trans. Serv. Comput.2
2013 iNewsBox: modeling and exploiting implicit feedback for building personalized news radio
abstract
Online news reading has become the major method to know about the world as web provide more information than other media like TV and radio. However, traditional online news reading interface is inconvenient for many types of people, especially for those who are disabled or taking a bus. This paper presents a mobile application iNewsBox enabling users to listen to news collected from the Internet. In order to simplify necessary interactions of getting valuable news, we also propose a framework for using implicit feedback to recommend news in this paper. Experiment shows our algorithms in iNewsBox are effective.
Yanan Xie, Liang Chen 0001, Kunyang Jia, Lichuan Ji, Jian Wu 0001
CIKM5
2013 gIceberg: Towards iceberg analysis in large graphs
abstract
Traditional multi-dimensional data analysis techniques such as iceberg cube cannot be directly applied to graphs for finding interesting or anomalous vertices due to the lack of dimensionality in graphs. In this paper, we introduce the concept of graph icebergs that refer to vertices for which the concentration (aggregation) of an attribute in their vicinities is abnormally high. Intuitively, these vertices shall be “close” to the attribute of interest in the graph space. Based on this intuition, we propose a novel framework, called gIceberg, which performs aggregation using random walks, rather than traditional SUM and AVG aggregate functions. This proposed framework scores vertices by their different levels of interestingness and finds important vertices that meet a user-specified threshold. To improve scalability, two aggregation strategies, forward and backward aggregation, are proposed with corresponding optimization techniques and bounds. Experiments on both real-world and synthetic large graphs demonstrate that gIceberg is effective and scalable.
Ziyu Guan, Lijie Ren, Jian Wu 0001, Jiawei Han 0001, Xifeng Yan
ICDE4
2013 WT-LDA: User Tagging Augmented LDA for Web Service Clustering
Liang Chen 0001, Yilun Wang 0001, Qi Yu 0001, Zibin Zheng, Jian Wu 0001
ICSOC5
2013 Static and Dynamic Structural Correlations in Graphs
abstract
Real-life graphs not only contain nodes and edges, but also have events taking place, e.g., product sales in social networks. Among different events, some exhibit strong correlations with the network structure, while others do not. Such structural correlations will shed light on viral influence existing in the corresponding network. Unfortunately, the traditional association mining concept is not applicable in graphs because it only works on homogeneous data sets like transactions and baskets. We propose a novel measure for assessing such structural correlations in heterogeneous graph data sets with events. The measure applies hitting time to aggregate the proximity among nodes that have the same event. To calculate the correlation scores for many events in a large network, we develop a scalable framework, called gScore, using sampling and approximation. By comparing to the situation where events are randomly distributed in the same network, our method is able to discover events that are highly correlated with the graph structure. We test gScore's effectiveness by synthetic events on the DBLP coauthor network and report interesting correlation results in a social network extracted from TaoBao.com, the largest online shopping network in China. Scalability of gScore is tested on the Twitter network. Since an event is essentially a temporal phenomenon, we also propose a dynamic measure, which reveals structural correlations at specific time steps and can be used for discovering detailed evolutionary patterns.
Jian Wu 0001, Ziyu Guan, Ambuj K. Singh, Xifeng Yan
IEEE Trans. Knowl. Data Eng.1
2013 Predicting Quality of Service for Selection by Neighborhood-Based Collaborative Filtering
abstract
Quality-of-service-based (QoS) service selection is an important issue of service-oriented computing. A common premise of previous research is that the QoS values of services to target users are supposed to be all known. However, many of QoS values are unknown in reality. This paper presents a neighborhood-based collaborative filtering approach to predict such unknown values for QoS-based selection. Compared with existing methods, the proposed method has three new features: 1) the adjusted-cosine-based similarity calculation to remove the impact of different QoS scale; 2) a data smoothing process to improve prediction accuracy; and 3) a similarity fusion approach to handle the data sparsity problem. In addition, a two-phase neighbor selection strategy is proposed to improve its scalability. An extensive performance study based on a public data set demonstrates its effectiveness.
Jian Wu 0001, Liang Chen 0001, Yipeng Feng, Zibin Zheng, MengChu Zhou, Zhaohui Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2012 WSTRank: Ranking Tags to Facilitate Web Service Mining
Liang Chen 0001, Zibin Zheng, Yipeng Feng, Jian Wu 0001, Michael R. Lyu
ICSOC4
2012 WSCRec: Utilizing Historical Information to Facilitate Web Service Composition
abstract
In this paper, we propose a novel framework WSCRec, in which historical information (i.e., execution logs) are utilized to facilitate the process of Web service composition. According to the execution logs of composite services, appropriate services which have been proved to be more reliable and robust and have higher probability to fulfill users' demands are located.
Liang Chen 0001, Hengyi Jian, Jian Wu 0001
ICWS3
2011 WTCluster: Utilizing Tags for Web Services Clustering
Liang Chen 0001, Liukai Hu, Zibin Zheng, Jian Wu 0001, Jianwei Yin, Ying Li 0001, Shuiguang Deng
ICSOC4
2011 Data-Dependency Aware Trust Evaluation for Service Choreography
abstract
This paper proposes a novel trust evaluation method for service choreography. Compared with current work towards this problem, it considers not only the trust for individual partner services and the explicit trust relation among partner services that have logical dependencies for each other, but also the implicit trust relation implied in data-dependencies among services. A serial of experiments, using the simulation tool Net Logo, are carried out to compare the evaluation results between the proposed method and the method without data-dependency consideration. The result shows that taking consideration of the data-dependency trust improves the accuracy of trust evaluation to a great extent.
Longtao Huang, Shuiguang Deng, Ying Li 0001, Jian Wu 0001, Jianwei Yin
ICWS4
2011 AWSP: An Automatic Web Service Planner Based on Heuristic State Space Search
abstract
With the number of available Web services is rapidly increasing, how to compose multiple Web services automatically to fulfill a given request has attracted much attention. This paper proposes a dedicated planner named AWSP (Automatic Web Service Planner) toward this problem. Compared with other AI planners for automatic Web service composition, AWSP is characterized by its two different heuristic functions to reduce the search space greatly. A series of experiments based on test sets generated by WSBen show that 1) AWSP performs well even when the scale of the test set expands significantly. 2) AWSP has a smaller search space and performs better when using the backward search strategy than using the forward search strategy, 3) AWSP with the A* heuristic function can get the solution with the shortest invocation path.
Shuiguang Deng, Ying Li 0001, Jian Wu 0001, Jianwei Yin
ICWS4
2011 Assessing and ranking structural correlations in graphs
abstract
Real-life graphs not only have nodes and edges, but also have events taking place, e.g., product sales in social networks and virus infection in communication networks. Among different events, some exhibit strong correlation with the network structure, while others do not. Such structural correlation will shed light on viral influence existing in the corresponding network. Unfortunately, the traditional association mining concept is not applicable in graphs since it only works on homogeneous datasets like transactions and baskets.
Ziyu Guan, Jian Wu 0001, Ambuj K. Singh, Xifeng Yan
SIGMOD Conference2
2010 Recommendation on Uncertain Services
abstract
In this paper, we propose a time-sensitive probability skyline (TPS) approach to recommend services with uncertainty. We project services to n-dimensional data space and recommend services in TPS. Experimental evaluation on real data shows the great performance of TPS in service recommendation by comparing the experiment result with results of other approaches.
Liang Chen 0001, Jian Wu 0001, Ru Jia, Shuiguang Deng, Ying Li 0001
ICWS2
2010 Analyzing Behavioral Substitution of Web Services Based on Pi-calculus
abstract
The behavioral analysis for Web services provides a priori detection of errors to ensure successful interactions in services invocation and composition, and the behavioral substitution of Web services is one of the most important issues in such analysis. In this paper, we propose to formalize the behavior of a Web service by π-calculus. Based on the formalization, we introduce two notions of behavioral substitution of Web services namely strong and weak simulation. Furthermore, we propose a derivative approach to analyzing the behavioral substitution of services according to the given notions, which is implemented based on an existing tool of π-calculus. The proposed approach takes advantage of formalization and theory of π-calculus, so that the formalized services can be naturally analyzed and the behavioral substitution of them can be easily determined.
Li Kuang, Yingjie Xia, Shuiguang Deng, Jian Wu 0001
ICWS4
2010 QoS-Driven Dynamic Reconfiguration of the SOA Based Software
abstract
SOA based software is typically based on dynamic reconfiguration, since it is the composition of services. But few works focus on the non-functional reconfiguration of the SOA-based software. This paper presents an approach for QoS driven dynamic reconfiguration of the SOA based Software. The approach can reconfigure a SOA based software to comply with a new QoS constrains by replacing its individual or multiple component services. The individual component services are replaced according to the descending order relative to the critical factors. While if the attempts fail, multiple component services will be replaced together. In our case study, an example is given to show the approach is efficient to reconfigure a SOA based software to meet a new QoS constraints.
Ying Li 0001, Yuyu Yin, Jian Wu 0001
ICSS4
2010 Automatic Composition of Semantic Web Services An Enhanced State Space Search Approach
abstract
This paper presents a novel approach for semantic web service composition based on traditional state space search approach. We regard automatic web service composition problem as an AI problem-solving problem and propose an enhanced state space search approach toward web service composition domain. This approach can not only be used for automatic service composition, but also for general problem-solving domain. In addition, in order to validate the feasibility of our approach, a prototype system is implemented.
Jian Wu 0001, Shuiguang Deng, Ying Li 0001, Jianwei Yin
ICSS2
2010 A C_net-based Verification of Web Service Compositions
abstract
Composition of Web Services has emerged as a new method to support business-to-business application integration. And the industrial world has already proposed several xml-based business protocol specification languages. In order to address the correct integration of web services, we specify a new method based on the c_net formal model for verifying the composition of web services. This new method can use the semantic information included in the c_net model to verify the flow composition at the semantic level. The mapping between c_net and BPEL4WS is also discussed.
Jian Wu 0001, Ying Li 0001
ICSS2
2010 Service Recommendation: Similarity-Based Representative Skyline
abstract
Skyline attracts more and more attention from academic circle and industrial circle because of its application in multi-criterion decision support, preference answering and data analysis. However, it seems unnecessary to recommend all services in skyline while the number of skyline points is large. The number of services in skyline is always large for the reason that comparability decreases with the increase of data dimensionality. Users always want to get only 2 or 3 recommendations instead of all services in skyline. Motivated by this, we propose to compute the representative skyline which contains some points that best describe the contour of the full skyline. In this paper, we propose a new definition which we call “similarity-based representative skyline”. We provide an algorithm SBRSA, which is based on a traversal approach to compute the value of similarity. In particular, we propose an algorithm to maintain the result of SBRSA in dynamic data environment. An extensive performance study using real and synthetic service data is reported to verify its great performance in representation and computing cost.
Liang Chen 0001, Jian Wu 0001, Shuiguang Deng, Ying Li 0001
SERVICES2
2009 Improving Scalability of Software Cloud for Composite Web Services
abstract
Most of the work on cloud scalability has focused on the granularity of applications and systems deployed on the cloud and on how to adjust their resource assignment according to the scale or volume of the requests. In this paper, we present a scheme for improving the scalability of service-based applications in a cloud from the granularity of the constituent services and their individual placement in the cloud. The approach we take is to analyze the communication patterns among the service operations of service-based applications and to analyze the assignment of the involved services to the available servers. We define a notion of scalability for service-based applications in a cloud and a framework to measure the scalability. We then propose an optimized assignment strategy to improve the scalability of composite Web services in terms of the productivity of such services. We report preliminary simulation experimental results that show the effectiveness of our scheme.
Jian Wu 0001, Qianhui Althea Liang, Elisa Bertino
IEEE CLOUD1
2009 Bayesian network based services recommendation
abstract
Nowadays, how to efficiently compose web services has become a hotspot. In this paper, we introduce a method of recommending an optimal service sequence based on the original service sequence for a composite service. This method uses a Bayesian-based approach and selects the service sequence that has the largest probability as the best choice. Compared with existing methods, this method has two advantages: firstly, service sequences recommended by this method are robust; secondly, this method produces a composite service with a high quality and it does this efficiently. We have conducted experiments to illustrate how our work helps facilitate web service composition.
Jian Wu 0001, Qianhui Althea Liang, Hengyi Jian
APSCC1
2009 Towards Adaptation of Service Interface Semantics
abstract
Interoperability promised by Web service makes it a most promising technology for the development of next generation distributed heterogeneous software systems. Services should be compliant at signature, behavioral and semantic level to make the interoperation successful and correct. Service adaptation provides an effective approach to bridge the incompatibility of services to make them interoperate as well as possible. In this paper, we aim to contribute to the definition of a methodology to develop adaptors that are capable of making two incompatible services interoperate not only successfully but also correctly at semantic level. To achieve this goal, we proposed service specifications for both atomic and composite services with semantic dependency between outputs and inputs specified; then we proposed adaptor specification consisting of three parts, which are message mapping, action mapping and treatment for non-mapping messages. Based on service and adaptor specifications, an incremental derivation approach of a concrete adaptor is given.
Li Kuang, Shuiguang Deng, Jian Wu 0001, Ying Li 0001
ICWS3
2009 Simulation Study of Public Goods Experiment
Zhi Tang 0009, Jian Wu 0001, Hanyi Xu
KES-AMSTA2
2009 Computing compatibility in dynamic service composition
Zhaohui Wu 0001, Shuiguang Deng, Ying Li 0001, Jian Wu 0001
Knowl. Inf. Syst.4
2008 Service Behavioral Adaptation Based on Dependency Graph
abstract
Service adaptation is one of the most important issues in SOC (Service Oriented Computing). This paper focuses on the issue of service behavioral adaptation and proposes an adapting method based on dependency graph. It can be divided into three sequential sub-problems: (1) service description-the foundation of service adaptation. We propose a formal approach to describing service behavior protocols; (2) mismatch definition-the identification of service mismatches. We define several kinds of behavior mismatches based on dependency graph; (3) service adaptation-the adaptor construction process. We detect all possible behavior mismatches and then generate different adaptors correspondingly.
Shuiguang Deng, Jian Wu 0001, Ying Li 0001, Li Kuang, Jianwei Yin
APSCC3
2008 An efficient two-phase service discovery mechanism
abstract
We bring forward a two-phase semantic service discovery mechanism which supports both the operation matchmaking and operation-composition matchmaking. A serial of experiments on a service management framework show that the mechanism gains better performance on both discovery recall rate and precision than a traditional matchmaker.
Shuiguang Deng, Zhaohui Wu 0001, Jian Wu 0001, Ying Li 0001
WWW3
2007 Inverted Indexing for Composition-Oriented Service Discovery
abstract
Service discovery becomes a key to hastening the evolution of web services as the number of services is expected to increase dramatically. In this paper, we propose to index all the ontology-annotated outputs in registered services. For each ontology-annotated output, there is a service list which records all the services in the registry that deliver the output. Based on the indexing, we propose a composition-oriented service discovery algorithm, which greatly accelerates the filtering of irrelevant atomic services by making use of the inverted indexing, and increases the likelihood of finding a possible candidate by exploring service composition. Experimental results show that the proposed algorithm provides a better performance on response time than the sequential matchmaking, and a better recall rate than the algorithms without the exploration of composition.
Li Kuang, Ying Li 0001, Jian Wu 0001, Shuiguang Deng, Zhaohui Wu 0001
ICWS3
2007 Using Improved FOAF to Enhance BPEL-extracted RBAC Capability
abstract
BPEL can automate orchestrations for cross-organizational Web services; however, it meets a serious challenge from modeling human-intensive business activities, especially from addressing access control for human coordination considering complex interpersonal relationship in modern business. This paper analyzes the importance of human-intensive processes and introduces several additional types of BPEL constructs, then discusses RBAC Model extracted from BPEL process, finally uses improved FOAF to enhance RBAC Model in BPEL. The goal of our work is to enhance human coordination capability in BPEL-based business processes by using RBAC model and improved FOAF.
Jian Wu 0001, Ying Li 0001, Zhaohui Wu 0001
Web Intelligence2
2006 Describing and Verifying Web Service Using Type Theory
abstract
A Web service is a basic software component that can be accessed by standard Internet protocols. It provides a new approach to cooperative and federated computing among different organizational units. There are many specifications which can describe the elements of Web services and make the end-users interact with each other. However, they are remaining at the descriptive level, without supporting any kind of mechanisms or tools for the verifying the specified attributes of the Web services. In the paper, we provide a mathematical scheme, type theory, to describe the basic elements of Web services and the specified attributes. We also present the mechanism to deduce the automated programs in other languages (for example ML) from the type theory. Thus we can verify the behavioral properties of a Web service as well as analyzing and verifying Web services composition
Jian Wu 0001, Shuiguang Deng, Ying Li 0001, Zhaohui Wu 0001
CSCWD2
2006 Service Classification Using Adaptive Back-Propagation Neural Network and Semantic Similarity
abstract
With the growing population of Web services, the discovery of services is a key to the development of Web services. While extensive researches focus mainly on service matchmaking algorithms, service classification that is also a meaningful approach to accelerating service discovery only receives little attention. In this paper, we propose to use adaptive back-propagation neural network model (BPM) to perform service recognition. During the training process, the feature vectors of training services and their categories are learned by the BPM. The element in the feature vector is the semantic similarity between the feature word in the system dictionary and the occurrence in the feature set for a service. During the recognition process, the characteristics of the test service are analyzed by the BPM and the output shows the category of the test service. Furthermore, the BPM is adapted with correctly recognized test services, which results in better modeling over time. Based on extensive experiments, we show that using the adaptive BPM is a promising way to realize automatic service classification, and the adaptive BPM using semantic similarity as the element in feature vector provides better performance than the general BPM using word frequency
Li Kuang, Jian Wu 0001, Shuiguang Deng, Ying Li 0001, Zhaohui Wu 0001
CSCWD2
2006 Modeling Service Compatibility with Pi-calculus for Choreography
Shuiguang Deng, Zhaohui Wu 0001, MengChu Zhou, Ying Li 0001, Jian Wu 0001
ER5
2006 Service Matchmaking Based on Semantics and Interface Dependencies
Shuiguang Deng, Jian Wu 0001, Ying Li 0001, Zhaohui Wu 0001
WAIM2
2006 Expressing Service and Query Behavior Using pi-Calculus for Matchmaking
abstract
Service discovery becomes a key to accelerating the evolution of Web services as the number of services is expected to increase dramatically. Foregoing work on service discovery is primarily based on the interfaces of services through the use of ontology. Ongoing work targets at service behavior, with not only individual message exchanges being captured, but also constraints between these message exchanges. In this paper, we propose a formal approach to expressing the service and query behavior using pi-calculus for service matchmaking. The resulting pi-calculus expressions of services and queries are precise in defining single operations involving message exchanges as well as execution sequence between operations. Based on the formalizations, service matchmaking between a service query and a service description is reasoned through the capability of pi-calculus. Expressing service behavior using pi-calculus is expected to be a promising way to realize intelligent service discovery
Li Kuang, Ying Li 0001, Shuiguang Deng, Jian Wu 0001, Zhaohui Wu 0001
Web Intelligence4
2006 Intelligent Transportation Information Sharing and Service Integration in Semantic Grid Environment
abstract
ITSGrid is an undergoing joint engineering project designed and developed by advanced computing and system (CCNT) lab in Zhejiang University and Hangzhou Enjoyor Electronics Co. Ltd (Enjoyor). The new features of ITSGrid are originated from two important research projects - DartGrid and DartFlow, and one key engineering project - JTang application server, in CCNT lab. Its goal is to build an integrated intelligent transportation information and service platform (ITISP), to integrate traffic data resources collected by Enjoyor and cooperate existing ITS subsystems and services deployed by Enjoyor, finally serve for transportation construction in China. During building this project, we utilize systematically the grid technology, the semantic Web technology, the Web service technology, the messaging oriented middleware technology
Jian Wu 0001, Ying Li 0001, Li Kuang
Web Intelligence2
2005 DART-Man: a management platform for Web services based on semantic Web technologies
abstract
Today's industries exhibit a growing trend towards benefiting from Web services. However, management complexity also grows with the number and type of services deployed. How to help users discover an appropriate service from tremendous ones, how to provide security access to target service, how to supply reliable service, and how to integrate commerce into services executing process become significant challenges. Therefore, companies soon require facilities that support and automate their management efforts. Furthermore, existing Web services management facilities, which lack structure and computer understandable metadata, prohibit users from managing Web services effectively. To address above issues, this paper elaborates a semantic Web technologies based service-management solution, called Dart-Man, towards making Web service management more flexible and automated.
Jian Wu 0001, Zhaohui Wu 0001
CSCWD (2)1