Fanghua Ye 0001

dblp:203/0957 · DBLP profile ↗
← Back
38ranked-venue papers
12as first author
23since 2021 · last 2025
0000-0002-1063-6534ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 8 first-author · 18 since 2021Databases, data management, data science and information retrieval · 13 · 5 first-author · 5 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Understanding Large Language Model Vulnerabilities to Social Bias Attacks
abstract
Large Language Models (LLMs) have become foundational in human-computer interaction, demonstrating remarkable linguistic capabilities across various tasks. However, there is a growing concern about their potential to perpetuate social biases present in their training data. In this paper, we comprehensively investigate the vulnerabilities of contemporary LLMs to various social bias attacks, including prefix injection, refusal suppression, and learned attack prompts. We evaluate popular models such as LLaMA-2, GPT-3.5, and GPT-4 across gender, racial, and religious bias types. Our findings reveal that models are generally more susceptible to gender bias attacks compared to racial or religious biases. We also explore novel aspects such as cross-bias and multiple-bias attacks, finding varying degrees of transferability across bias types. Additionally, our results show that larger models and pretrained base models often exhibit higher susceptibility to bias attacks. These insights contribute to the development of more inclusive and ethically responsible LLMs, emphasizing the importance of understanding and mitigating potential bias vulnerabilities. We offer recommendations for model developers and users to enhance the robustness of LLMs against social bias attacks.
Jiaxu Zhao 0002, Fanghua Ye 0001, Joey Tianyi Zhou, Mykola Pechenizkiy
ACL (1)3
2025 UNComp: Can Matrix Entropy Uncover Sparsity? - A Compressor Design from an Uncertainty-Aware Perspective
abstract
Jing Xiong, Jianghan Shen, Fanghua Ye, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Xun Wu, Chuanyang Zheng, Zhijiang Guo, Min Yang, Lingpeng Kong, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jianghan Shen, Fanghua Ye 0001, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Chuanyang Zheng, Zhijiang Guo, Min Yang 0007, Lingpeng Kong, Ngai Wong 0001
EMNLP3
2025 ParallelComp: Parallel Long-Context Compressor for Length Extrapolation
abstract
Extrapolating ultra-long contexts (text length $>$128K) remains a major challenge for large language models (LLMs), as most training-free extrapolation methods are not only severely limited by memory bottlenecks, but also suffer from the attention sink, which restricts their scalability and effectiveness in practice. In this work, we propose ParallelComp, a parallel long-context compression method that effectively overcomes the memory bottleneck, enabling 8B-parameter LLMs to extrapolate from 8K to 128K tokens on a single A100 80GB GPU in a training-free setting. ParallelComp splits the input into chunks, dynamically evicting redundant chunks and irrelevant tokens, supported by a parallel KV cache eviction mechanism. Importantly, we present a systematic theoretical and empirical analysis of attention biases in parallel attention—including the attention sink, recency bias, and middle bias—and reveal that these biases exhibit distinctive patterns under ultra-long context settings. We further design a KV cache eviction technique to mitigate this phenomenon. Experimental results show that ParallelComp enables an 8B model (trained on 8K context) to achieve 91.17% of GPT-4’s performance under ultra-long contexts, outperforming closed-source models such as Claude-2 and Kimi-Chat. We achieve a 1.76x improvement in chunk throughput, thereby achieving a 23.50x acceleration in the prefill stage with negligible performance loss and pave the way for scalable and robust ultra-long contexts extrapolation in LLMs. We release the code at https://github.com/menik1126/ParallelComp.
Jianghan Shen, Chuanyang Zheng, Zhongwei Wan, Chiwun Yang, Fanghua Ye 0001, Hongxia Yang, Lingpeng Kong, Ngai Wong 0001
ICML7
2025 SkipGPT: Each Token is One of a Kind
abstract
Large language models (LLMs) achieve remarkable performance across tasks but incur substantial computational costs due to their deep, multi-layered architectures. Layer pruning has emerged as a strategy to alleviate these inefficiencies, but conventional static pruning methods overlook two critical dynamics inherent to LLM inference: (1) *horizontal dynamics*, where token-level heterogeneity demands context-aware pruning decisions, and (2) *vertical dynamics*, where the distinct functional roles of MLP and self-attention layers necessitate component-specific pruning policies. We introduce **SkipGPT**, a dynamic layer pruning framework designed to optimize computational resource allocation through two core innovations: (1) global token-aware routing to prioritize critical tokens and (2) decoupled pruning policies for MLP and self-attention components. To mitigate training instability, we propose a two-stage optimization paradigm: first, a disentangled training phase that learns routing strategies via soft parameterization to avoid premature pruning decisions, followed by parameter-efficient LoRA fine-tuning to restore performance impacted by layer removal. Extensive experiments demonstrate that SkipGPT reduces over 40% model parameters while matching or exceeding the performance of the original dense model across benchmarks. By harmonizing dynamic efficiency with preserved expressivity, SkipGPT advances the practical deployment of scalable, resource-aware LLMs. Our code is publicly available at: https://github.com/EIT-NLP/SkipGPT.
Anhao Zhao, Fanghua Ye 0001, Yingqi Fan, Junlong Tong, Zhiwei Fei, Hui Su, Xiaoyu Shen 0001
ICML2
2025 Soft-consensual Federated Learning for Data Heterogeneity via Multiple Paths
abstract
Federated learning enables collaborative training while preserving the privacy of all participants. However, the heterogeneity in data distribution across multiple training nodes poses significant challenges to the construction of federated models. Prior studies were dedicated to mitigating the effects of data heterogeneity by using global information as a blueprint and restricting the local update of the model for reaching a "hard consensus". But this practice makes it difficult to balance local and global information, and it neglects to negotiate amicably between local and global models to reach mutually agreeable results, called ``soft consensus". In this paper, a multiple-path solving method is proposed to balance global and local features and combine these two feature preference paths to reach a soft consensus. Rather than relying on global information as the sole criterion, a negotiation process is employed to address the same objective by accommodating diverse feature preferences, thereby facilitating the discovery of a more plausible solution through multiple distinct pathways. Considering the overwhelming power of local features during local training, a swapping strategy is applied to weaken them to balance the solution paths. Moreover, to minimize the additional communication cost caused by the introduction of multiple paths, the solution of the task network is converted into data adaptation to reduce the amount of parameter transmission. Extensive experiments are conducted to demonstrate the advantages of the proposed method.
Lele Fu, Fanghua Ye 0001, Tianchi Liao, Bowen Deng 0002, Chuanfu Zhang, Chuan Chen 0001
NeurIPS3
2025 Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
abstract
Abstract The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017) that have acted as benchmarks for progress in this field. This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation of long sentences, attention model as word alignment, and sub-optimal beam search. Our empirical findings show that LLMs effectively reduce reliance on parallel data for major languages during pretraining and significantly improve translation of long sentences containing approximately 80 words, even translating documents up to 512 words. Despite these improvements, challenges in domain mismatch and rare word prediction persist. While NMT-specific challenges like word alignment and beam search may not apply to LLMs, we identify three new challenges in LLM-based translation: inference efficiency, translation of low-resource languages during pretraining, and human-aligned evaluation.
Jianhui Pang, Fanghua Ye 0001, Derek F. Wong, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu, Longyue Wang
Trans. Assoc. Comput. Linguistics2
2024 Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism
abstract
Large language models (LLMs) exhibit remarkable in-context learning (ICL) capabilities.However, the underlying working mechanism of ICL remains poorly understood.Recent research presents two conflicting views on ICL: One emphasizes the impact of similar examples in the demonstrations, stressing the need for label correctness and more shots.The other attributes it to LLMs' inherent ability of task recognition, deeming label correctness and shot numbers of demonstrations as not crucial.In this work, we provide a Two-Dimensional Coordinate System that unifies both views into a systematic framework.The framework explains the behavior of ICL through two orthogonal variables: whether similar examples are presented in the demonstrations (perception) and whether LLMs can recognize the task (cognition).We propose the peak inverse rank metric to detect the task recognition ability of LLMs and study LLMs' reactions to different definitions of similarity.Based on these, we conduct extensive experiments to elucidate how ICL functions across each quadrant on multiple representative classification tasks.Finally, we extend our analyses to generation tasks, showing that our coordinate system can also be used to interpret ICL for generation tasks effectively.
Anhao Zhao, Fanghua Ye 0001, Jinlan Fu, Xiaoyu Shen 0001
EMNLP2
2024 Benchmarking LLMs via Uncertainty Quantification
abstract
The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods. However, current evaluation platforms, such as the widely recognized HuggingFace open LLM leaderboard, neglect a crucial aspect -- uncertainty, which is vital for thoroughly assessing LLMs. To bridge this gap, we introduce a new benchmarking approach for LLMs that integrates uncertainty quantification. Our examination involves nine LLMs (LLM series) spanning five representative natural language processing tasks. Our findings reveal that: I) LLMs with higher accuracy may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty compared to their smaller counterparts; and III) Instruction-finetuning tends to increase the uncertainty of LLMs. These results underscore the significance of incorporating uncertainty in the evaluation of LLMs. Our implementation is available at https://github.com/smartyfh/LLM-Uncertainty-Bench.
Fanghua Ye 0001, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi 0001, Zhaopeng Tu
NeurIPS1
2024 UTMGAT: a unified transformer with memory encoder and graph attention networks for multidomain dialogue state tracking
Bhuyan Kaibalya Prasad, Guilin Qi, Fanghua Ye 0001, Zafar Ali, Irfan Ullah 0001, Pavlos Kefalas
Appl. Intell.5
2024 Semi-supervised anomaly detection with contamination-resilience and incremental training
Liheng Yuan, Fanghua Ye 0001, Heng Li 0008, Cuiying Gao, Chengqing Yu, Wei Yuan 0001, Xinge You
Eng. Appl. Artif. Intell.2
2024 Concept drift adaptation with scarce labels: A novel approach based on diffusion and adversarial learning
Liheng Yuan, Fanghua Ye 0001, Wei Yuan 0001, Xinge You
Eng. Appl. Artif. Intell.2
2023 Modeling User Satisfaction Dynamics in Dialogue via Hawkes Process
abstract
Dialogue systems have received increasing attention while automatically evaluating their performance remains challenging.User satisfaction estimation (USE) has been proposed as an alternative.It assumes that the performance of a dialogue system can be measured by user satisfaction and uses an estimator to simulate users.The effectiveness of USE depends heavily on the estimator.Existing estimators independently predict user satisfaction at each turn and ignore satisfaction dynamics across turns within a dialogue.In order to fully simulate users, it is crucial to take satisfaction dynamics into account.To fill this gap, we propose a new estimator ASAP (sAtisfaction eStimation via HAwkes Process) that treats user satisfaction across turns as an event sequence and employs a Hawkes process to effectively model the dynamics in this sequence.Experimental results on four benchmark dialogue datasets demonstrate that ASAP can substantially outperform state-of-the-art baseline estimators.
Fanghua Ye 0001, Emine Yilmaz
ACL (1)1
2023 Turn-Level Active Learning for Dialogue State Tracking
abstract
Dialogue state tracking (DST) plays an important role in task-oriented dialogue systems.However, collecting a large amount of turnby-turn annotated dialogue data is costly and inefficient.In this paper, we propose a novel turn-level active learning framework for DST to actively select turns in dialogues to annotate.Given the limited labelling budget, experimental results demonstrate the effectiveness of selective annotation of dialogue turns.Additionally, our approach can effectively achieve comparable DST performance to traditional training approaches with significantly less annotated data, which provides a more efficient way to annotate new dialogue data 1 . Turn User SystemCan you tell me some info on the Avalon hotel?The Avalon is a 4 star moderately priced guesthouse in the north with free internet.Would you like to book there?Yes. Can you book it for 5 people on Saturday?We need rooms for 4 nights. Dialogue StateYour taxi has been booked to take you from Avalon to Frankie and Bennys at 17:45.Your taxi will be a black Tesla and the contact number is 07715682347.That sounds great.Thank you very much.…..
Fanghua Ye 0001, Ling Chen 0006, Mohammad-Reza Namazi-Rad
EMNLP3
2023 Lending Interaction Wings to Recommender Systems with Conversational Agents
abstract
An intelligent conversational agent (a.k.a., chat-bot) could embrace conversational technologies to obtain user preferences online, to overcome inherent limitations of recommender systems trained over the offline historical user behaviors. In this paper, we propose CORE, a new offline-training and online-checking framework to plug a COnversational agent into REcommender systems. Unlike most prior conversational recommendation approaches that systemically combine conversational and recommender parts through a reinforcement learning framework, CORE bridges the conversational agent and recommender system through a unified uncertainty minimization framework, which can be easily applied to any existing recommendation approach. Concretely, CORE treats a recommender system as an offline estimator to produce an estimated relevance score for each item, while CORE regards a conversational agent as an online checker that checks these estimated scores in each online session. We define uncertainty as the sum of unchecked relevance scores. In this regard, the conversational agent acts to minimize uncertainty via querying either attributes or items. Towards uncertainty minimization, we derive the certainty gain of querying each attribute and item, and develop a novel online decision tree algorithm to decide what to query at each turn. Our theoretical analysis reveals the bound of the expected number of turns of CORE in a cold-start setting. Experimental results demonstrate that CORE can be seamlessly employed on a variety of recommendation approaches, and can consistently bring significant improvements in both hot-start and cold-start settings.
Jiarui Jin, Fanghua Ye 0001, Mengyue Yang, Yue Feng 0002, Weinan Zhang 0001, Yong Yu 0001, Jun Wang 0012
NeurIPS3
2022 Dynamic Schema Graph Fusion Network for Multi-Domain Dialogue State Tracking
abstract
Dialogue State Tracking (DST) aims to keep track of users' intentions during the course of a conversation.In DST, modelling the relations among domains and slots is still an under-studied problem.Existing approaches that have considered such relations generally fall short in: (1) fusing prior slot-domain membership relations and dialogue-aware dynamic slot relations explicitly, and (2) generalizing to unseen domains.To address these issues, we propose a novel Dynamic Schema Graph Fusion Network (DSGFNet), which generates a dynamic schema graph to explicitly fuse the prior slot-domain membership relations and dialogue-aware dynamic slot relations.It also uses the schemata to facilitate knowledge transfer to new domains.DSGFNet consists of a dialogue utterance encoder, a schema graph encoder, a dialogue-aware schema graph evolving network, and a schema graph enhanced dialogue state decoder.Empirical results on benchmark datasets (i.e., SGD, MultiWOZ2.1,and MultiWOZ2.2),show that DSGFNet outperforms existing methods.
Yue Feng 0002, Aldo Lipani, Fanghua Ye 0001, Qiang Zhang 0026, Emine Yilmaz
ACL (1)3
2022 MetaASSIST: Robust Dialogue State Tracking with Meta Learning
abstract
Existing dialogue datasets contain lots of noise in their state annotations.Such noise can hurt model training and ultimately lead to poor generalization performance.A general framework named ASSIST has recently been proposed to train robust dialogue state tracking (DST) models.It introduces an auxiliary model to generate pseudo labels for the noisy training set.These pseudo labels are combined with vanilla labels by a common fixed weighting parameter to train the primary DST model.Notwithstanding the improvements of ASSIST on DST, tuning the weighting parameter is challenging.Moreover, a single parameter shared by all slots and all instances may be suboptimal.To overcome these limitations, we propose a meta learning-based framework MetaASSIST to adaptively learn the weighting parameter.Specifically, we propose three schemes with varying degrees of flexibility, ranging from slot-wise to both slot-wise and instance-wise, to convert the weighting parameter into learnable functions.These functions are trained in a meta-learning manner by taking the validation set as meta data.Experimental results demonstrate that all three schemes can achieve competitive performance.Most impressively, we achieve a state-of-the-art joint goal accuracy of 80.10% on MultiWOZ 2.4.
Fanghua Ye 0001, Xi Wang 0012, Jie Huang 0009, Shenghui Li, Samuel Stern, Emine Yilmaz
EMNLP1
2022 MultiWOZ 2.4: A Multi-Domain Task-Oriented Dialogue Dataset with Essential Annotation Corrections to Improve State Tracking Evaluation
abstract
The MultiWOZ 2.0 dataset has greatly stimulated the research of task-oriented dialogue systems.However, its state annotations contain substantial noise, which hinders a proper evaluation of model performance.To address this issue, massive efforts were devoted to correcting the annotations.Three improved versions (i.e., MultiWOZ 2.1-2.3)have then been released.Nonetheless, there are still plenty of incorrect and inconsistent annotations.This work introduces MultiWOZ 2.4, which refines the annotations in the validation set and test set of MultiWOZ 2.1.The annotations in the training set remain unchanged (same as MultiWOZ 2.1) to elicit robust and noise-resilient model training.We benchmark eight state-of-the-art dialogue state tracking models on MultiWOZ 2.4.All of them demonstrate much higher performance than on MultiWOZ 2.1 1 .Error Type Conversation Example MultiWOZ 2.1 MultiWOZ 2.4 (I) Context Mismatch Usr: Hello, I would like to book a taxi from restaurant 2 two to the museum of classical archaeology.taxi-destination=museum of archaelogy and anthropology taxi-destination=museum of classical archaeology Usr: I am looking for a restaurant that serves Portuguese food.rest.-food=Portugeserest.-food=Portuguese
Fanghua Ye 0001, Jarana Manotumruksa, Emine Yilmaz
SIGDIAL1
2022 Auto-weighted Robust Federated Learning with Corrupted Data Sources
abstract
Federated learning provides a communication-efficient and privacy-preserving training process by enabling learning statistical models with massive participants without accessing their local data. Standard federated learning techniques that naively minimize an average loss function are vulnerable to data corruptions from outliers, systematic mislabeling, or even adversaries. In this article, we address this challenge by proposing Auto-weighted Robust Federated Learning ( ARFL ), a novel approach that jointly learns the global model and the weights of local updates to provide robustness against corrupted data sources. We prove a learning bound on the expected loss with respect to the predictor and the weights of clients, which guides the definition of the objective for robust federated learning. We present an objective that minimizes the weighted sum of empirical risk of clients with a regularization term, where the weights can be allocated by comparing the empirical risk of each client with the average empirical risk of the best \( p \) clients. This method can downweight the clients with significantly higher losses, thereby lowering their contributions to the global model. We show that this approach achieves robustness when the data of corrupted clients is distributed differently from the benign ones. To optimize the objective function, we propose a communication-efficient algorithm based on the blockwise minimization paradigm. We conduct extensive experiments on multiple benchmark datasets, including CIFAR-10, FEMNIST, and Shakespeare, considering different neural network models. The results show that our solution is robust against different scenarios, including label shuffling, label flipping, and noisy features, and outperforms the state-of-the-art methods in most scenarios.
Shenghui Li, Edith C. H. Ngai, Fanghua Ye 0001, Thiemo Voigt
ACM Trans. Intell. Syst. Technol.3
2021 Outlier-Resilient Web Service QoS Prediction
abstract
The proliferation of Web services makes it difficult for users to select the most appropriate one among numerous functionally identical or similar service candidates. Quality-of-Service (QoS) describes the non-functional characteristics of Web services, and it has become the key differentiator for service selection. However, users cannot invoke all Web services to obtain the corresponding QoS values due to high time cost and huge resource overhead. Thus, it is essential to predict unknown QoS values. Although various QoS prediction methods have been proposed, few of them have taken outliers into consideration, which may dramatically degrade the prediction performance. To overcome this limitation, we propose an outlier-resilient QoS prediction method in this paper. Our method utilizes Cauchy loss to measure the discrepancy between the observed QoS values and the predicted ones. Owing to the robustness of Cauchy loss, our method is resilient to outliers. We further extend our method to provide time-aware QoS prediction results by taking the temporal information into consideration. Finally, we conduct extensive experiments on both static and dynamic datasets. The results demonstrate that our method is able to achieve better performance than state-of-the-art baseline methods.
Fanghua Ye 0001, Chuan Chen 0001, Zibin Zheng, Hong Huang 0001
WWW1
2021 Slot Self-Attentive Dialogue State Tracking
abstract
An indispensable component in task-oriented dialogue systems is the dialogue state tracker, which keeps track of users’ intentions in the course of conversation. The typical approach towards this goal is to fill in multiple pre-defined slots that are essential to complete the task. Although various dialogue state tracking methods have been proposed in recent years, most of them predict the value of each slot separately and fail to consider the correlations among slots. In this paper, we propose a slot self-attention mechanism that can learn the slot correlations automatically. Specifically, a slot-token attention is first utilized to obtain slot-specific features from the dialogue context. Then a stacked slot self-attention is applied on these features to learn the correlations among slots. We conduct comprehensive experiments on two multi-domain task-oriented dialogue datasets, including MultiWOZ 2.0 and MultiWOZ 2.1. The experimental results demonstrate that our approach achieves state-of-the-art performance on both datasets, verifying the necessity and effectiveness of taking slot correlations into consideration.
Fanghua Ye 0001, Jarana Manotumruksa, Qiang Zhang 0026, Shenghui Li, Emine Yilmaz
WWW1
2021 Learning deep discriminative representations with pseudo supervision for image clustering
Weibo Hu, Chuan Chen 0001, Fanghua Ye 0001, Zibin Zheng, Yunfei Du 0001
Inf. Sci.3
2021 Enhancing Graph Neural Networks via auxiliary training for semi-supervised node classification
Yu Song 0005, Hong Huang 0001, Fanghua Ye 0001, Xing Xie 0001, Hai Jin 0001
Knowl. Based Syst.4
2021 Multi-Stage Network Embedding for Exploring Heterogeneous Edges
abstract
The relationships between objects in a network are typically diverse and complex, leading to the heterogeneous edges with different semantic information. In this article, we focus on exploring the heterogeneous edges for network representation learning. By considering each relationship as a view that depicts a specific type of proximity between nodes, we propose a multi-stage non-negative matrix factorization (MNMF) model, committed to utilizing abundant information in multiple views to learn robust network representations. In fact, most existing network embedding methods are closely related to implicitly factorizing the complex proximity matrix. However, the approximation error is usually quite large, since a single low-rank matrix is insufficient to capture the original information. Through a multi-stage matrix factorization process motivated by gradient boosting, our MNMF model achieves lower approximation error. Meanwhile, the multi-stage structure of MNMF gives the feasibility of designing two kinds of non-negative matrix factorization (NMF) manners to preserve network information better. The united NMF aims to preserve the consensus information between different views, and the independent NMF aims to preserve unique information of each view. Concrete experimental results on realistic datasets indicate that our model outperforms three types of baselines in practical applications.
Hong Huang 0001, Yu Song 0005, Fanghua Ye 0001, Xing Xie 0001, Xuanhua Shi, Hai Jin 0001
ACM Trans. Knowl. Discov. Data3
2020 Modeling Heterogeneous Edges to Represent Networks with Graph Auto-Encoder
Lu Wang 0002, Yu Song 0005, Hong Huang 0001, Fanghua Ye 0001, Xuanhua Shi, Hai Jin 0001
DASFAA (2)4
2020 Homophily Preserving Community Detection
abstract
As a fundamental problem in social network analysis, community detection has recently attracted wide attention, accompanied by the output of numerous community detection methods. However, most existing methods are developed by only exploiting link topology, without taking node homophily (i.e., node similarity) into consideration. Thus, much useful information that can be utilized to improve the quality of detected communities is ignored. To overcome this limitation, we propose a new community detection approach based on nonnegative matrix factorization (NMF), namely, homophily preserving NMF (HPNMF), which models not only link topology but also node homophily of networks. As such, HPNMF is able to better reflect the inherent properties of community structure. In order to capture node homophily from scratch, we provide three similarity measurements that naturally reveal the association relationships between nodes. We further present an efficient learning algorithm with convergence guarantee to solve the proposed model. Finally, extensive experiments are conducted, and the results demonstrate that HPNMF has strong ability to outperform the state-of-the-art baseline methods.
Fanghua Ye 0001, Chuan Chen 0001, Zibin Zheng, Wuhui Chen
IEEE Trans. Neural Networks Learn. Syst.1
2020 Nonuniform Hyper-Network Embedding with Dual Mechanism
abstract
Network embedding which aims to learn the low-dimensional representations for vertices in networks has been extensively studied in recent years. Although there are various models designed for networks with different properties and different structures for different tasks, most of them are only applied to normal networks which only contain pairwise relationships between vertices. In many realistic cases, relationships among objects are not pairwise and such relationships can be better modeled by a hyper-network in which each edge can connect an uncertain number of vertices. In this article, we focus on two properties of hyper-networks: nonuniform and dual property. In order to make full use of these two properties, we firstly propose a flexible model called Hyper2vec to learn the embeddings of hyper-networks by applying a biased second order random walk strategy to hyper-networks in the framework of Skip-gram. Then, we combine the features of hyperedges by considering the dual hyper-networks to build a further model called NHNE based on 1D convolutional neural networks, and train a tuplewise similarity function for the nonuniform relationships in hyper-networks. Extensive experiments demonstrate the significant effectiveness of our methods for hyper-network embedding.
Jie Huang 0009, Chuan Chen 0001, Fanghua Ye 0001, Weibo Hu, Zibin Zheng
ACM Trans. Inf. Syst.3
2020 Finding skyline communities in multi-valued networks
Rong-Hua Li 0001, Lu Qin 0001, Fanghua Ye 0001, Guoren Wang, Jeffrey Xu Yu, Xiaokui Xiao, Nong Xiao 0001, Zibin Zheng
VLDB J.3
2019 Discrete Overlapping Community Detection with Pseudo Supervision
abstract
Community detection is of significant importance in understanding the structures and functions of networks. Recently, overlapping community detection has drawn much attention due to the ubiquity of overlapping community structures in real-world networks. Nonnegative matrix factorization (NMF), as an emerging standard framework, has been widely employed for overlapping community detection, which obtains nodes' soft community memberships by factorizing the adjacency matrix into low-rank factor matrices. However, in order to determine the ultimate community memberships, we have to post-process the real-valued factor matrix by manually specifying a threshold on it, which is undoubtedly a difficult task. Even worse, a unified threshold may not be suitable for all nodes. To circumvent the cumbersome post-processing step, we propose a novel discrete overlapping community detection approach, i.e., Discrete Nonnegative Matrix Factorization (DNMF), which seeks for a discrete (binary) community membership matrix directly. Thus DNMF is able to assign explicit community memberships to nodes without post-processing. Moreover, DNMF incorporates a pseudo supervision module into it to exploit the discriminative information in an unsupervised manner, which further enhances its robustness. We thoroughly evaluate DNMF using both synthetic and real-world networks. Experiments show that DNMF has the ability to outperform state-of-the-art baseline approaches.
Fanghua Ye 0001, Chuan Chen 0001, Zibin Zheng, Rong-Hua Li 0001, Jeffrey Xu Yu
ICDM1
2019 Digging into it: Community detection via hidden attributes analysis
Li Rui Jie, Fanghua Ye 0001, Shaoan Xie, Chuan Chen 0001, Zibin Zheng
Neurocomputing2
2019 STD: An Automatic Evaluation Metric for Machine Translation Based on Word Embeddings
abstract
Lexical-based metrics such as BLEU, NIST, and WER have been widely used in machine translation (MT) evaluation. However, these metrics badly represent semantic relationships and impose strict identity matching, leading to moderate correlation with human judgments. In this paper, we propose a novel MT automatic evaluation metric Semantic Travel Distance (STD) based on word embeddings. STD incorporates both semantic and lexical features (word embeddings and n-gram and word order) into one metric. It measures the semantic distance between the hypothesis and reference by calculating the minimum cumulative cost that the embedded n-grams of the hypothesis need to “travel” to reach the embedded n-grams of the reference. Experiment results show that STD has a better and more robust performance than a range of state-of-the-art metrics for both the segment-level and system-level evaluation.
Pairui Li, Chuan Chen 0001, Wujie Zheng, Yuetang Deng, Fanghua Ye 0001, Zibin Zheng
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Deep Autoencoder-like Nonnegative Matrix Factorization for Community Detection
abstract
Community structure is ubiquitous in real-world complex networks. The task of community detection over these networks is of paramount importance in a variety of applications. Recently, nonnegative matrix factorization (NMF) has been widely adopted for community detection due to its great interpretability and its natural fitness for capturing the community membership of nodes. However, the existing NMF-based community detection approaches are shallow methods. They learn the community assignment by mapping the original network to the community membership space directly. Considering the complicated and diversified topology structures of real-world networks, it is highly possible that the mapping between the original network and the community membership space contains rather complex hierarchical information, which cannot be interpreted by classic shallow NMF-based approaches. Inspired by the unique feature representation learning capability of deep autoencoder, we propose a novel model, named Deep Autoencoder-like NMF (DANMF), for community detection. Similar to deep autoencoder, DANMF consists of an encoder component and a decoder component. This architecture empowers DANMF to learn the hierarchical mappings between the original network and the final community assignment with implicit low-to-high level hidden attributes of the original network learnt in the intermediate layers. Thus, DANMF should be better suited to the community detection task. Extensive experiments on benchmark datasets demonstrate that DANMF can achieve better performance than the state-of-the-art NMF-based community detection approaches.
Fanghua Ye 0001, Chuan Chen 0001, Zibin Zheng
CIKM1
2018 Adaptive Affinity Learning for Accurate Community Detection
abstract
The task of community detection has become a fundamental research problem in complex network analysis. Intuitively, similar nodes are more likely to be contained in the same community. However, most existing community detection methods cannot extract the intrinsic similarity between nodes. Thus, they may fail to identify the real community structures. In this paper, we propose to learn an affinity matrix adaptively, which can capture the intrinsic similarity between nodes accurately, and therefore benefit the community detection results. Specifically, the proposed model first embeds each node into a low-dimensional space through a transformation matrix with the community structures being preserved. Then, our model learns the affinity matrix in this low-dimensional space. The affinity matrix is further utilized to guide the learning of the community membership matrix via manifold regularization. The above three matrices are learned simultaneously and updated iteratively under the framework of Alternating Direction Method of Multipliers (ADMM). Extensive experiments show that our model can outperform the state-of-the-art approaches.
Fanghua Ye 0001, Shenghui Li, Chuan Chen 0001, Zibin Zheng
ICDM1
2018 An Adaptive Semi-local Algorithm for Node Ranking in Large Complex Networks
Fanghua Ye 0001, Chuan Chen 0001, Jiajing Wu, Zibin Zheng
ICSOC1
2018 Identifying Influential Nodes in Complex Networks via Semi-Local Centrality
abstract
Node influence refers to the ability of a node to disseminate information. The faster and wider the node spreads, the more influential it is. With its great theoretical and practical significance, identifying influential nodes in complex networks becomes one of the most attractive research topics in recent years. Consequently, a variety of different methods, such as betweenness, closeness, and pagerank, have been proposed to identify influential nodes. However, most of the existing methods cannot lead to a tradeoff between the ranking accuracy and time complexity, which limits their application on many real-world complex networks. In this paper, we propose a novel and efficient ranking method named semi-local centrality, to evaluate the influence of nodes more accurately. The method performs random walk to collect the influential surround nodes which are used to evaluate the nodes' influence. We use susceptible-infected-recovered (SIR) model to evaluate the performance of our method. The experimental results on four real-world networks show that the proposed method can identify influential nodes more effectively compared with five state-of-the-art methods.
Jiali Dong, Fanghua Ye 0001, Wuhui Chen, Jiajing Wu
ISCAS2
2018 Measuring Cohesion of Software Systems Using Weighted Directed Complex Networks
abstract
Network theory has been demonstrated as an effective approach for better understanding and analysis of software systems from a systematic perspective. In this paper, we develop a directed and weighted software dependency network model to analyse software systems from a complex network perspective. In particular, to measure the "High Cohesion and Low Coupling" nature of object-oriented software systems, we propose to use a directed and weighted modularity index, which can better reflect the cohesion of software systems. Experiments are conducted on a series of open source object-oriented software systems with various scales, and the results show that the directed and weighted modularity can better characterize software systems with different cohesion.
Jiajing Wu, Yongxiang Xia, Fanghua Ye 0001
ISCAS4
2018 Skyline Community Search in Multi-valued Networks
abstract
Given a scientific collaboration network, how can we find a group of collaborators with high research indicator (e.g., h-index) and diverse research interests? Given a social network, how can we identify the communities that have high influence (e.g., PageRank) and also have similar interests to a specified user? In such settings, the network can be modeled as a multi-valued network where each node has d ($d \ge 1$) numerical attributes (i.e., h-index, diversity, PageRank, similarity score, etc.). In the multi-valued network, we want to find communities that are not dominated by the other communities in terms of d numerical attributes. Most existing community search algorithms either completely ignore the numerical attributes or only consider one numerical attribute of the nodes. To capture d numerical attributes, we propose a novel community model, called skyline community, based on the concepts of k-core and skyline. A skyline community is a maximal connected k-core that cannot be dominated by the other connected k-cores in the d-dimensional attribute space. We develop an elegant space-partition algorithm to efficiently compute the skyline communities. Two striking advantages of our algorithm are that (1) its time complexity relies mainly on the size of the answer s (i.e., the number of skyline communities), thus it is very efficient if s is small; and (2) it can progressively output the skyline communities, which is very useful for applications that only require part of the skyline communities. Extensive experiments on both synthetic and real-world networks demonstrate the efficiency, scalability, and effectiveness of the proposed algorithm.
Rong-Hua Li 0001, Lu Qin 0001, Fanghua Ye 0001, Jeffrey Xu Yu, Xiaokui Xiao, Nong Xiao 0001, Zibin Zheng
SIGMOD Conference3
2017 Efficient Influential Individuals Discovery on Service-Oriented Social Networks: A Community-Based Approach
Fanghua Ye 0001, Chuan Chen 0001, Guohui Ling, Zibin Zheng
ICSOC1
2017 Finding weighted k-truss communities in large networks
Zibin Zheng, Fanghua Ye 0001, Rong-Hua Li 0001, Guohui Ling, Tan Jin
Inf. Sci.2