Yanhui Zhu

dblp:12/6541 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HSE-GTA: Generative Augmentation Based Multimodal Entity Alignment
Fuwen Tang, Yanhui Zhu, Xujian Ying, Zhipeng Gan
ICIC (24)2
2024 Fairness in Monotone k-submodular Maximization: Algorithms and Applications
abstract
Submodular optimization has become increasingly prominent in machine learning, and fairness has drawn much attention. In this paper, we propose to study the fair k-submodular maximization problem and develop a 1/3-approximation greedy algorithm with a running time of O(knB). Our theoretical guarantee matches the best-known k-submodular maximization results without fairness constraints. In addition, we have developed a faster threshold-based algorithm that achieves a (1/3 ϵ) approximation with ${\mathcal{O}}\left({\frac{{kn}}{\varepsilon }\log \frac{B}{\varepsilon }}\right)$ evaluations of the function−f. Furthermore, for both algorithms, we provide approximation guarantees when the k-submodular function is not accessible but only can be approximately accessed. We have extensively validated our theoretical findings through empirical study and examined the practical implications of fairness. The experimental results show that the fairness constraints do not significantly undermine the quality of solutions.
Yanhui Zhu, Samik Basu 0001, Aduri Pavan
IEEE Big Data1
2024 Regularized Unconstrained Weakly Submodular Maximization
abstract
Submodular optimization finds applications in machine learning and data mining. In this paper, we study the problem of maximizing functions of the form h = f-c, where f is a monotone, non-negative, weakly submodular set function and c is a modular function. We design a deterministic approximation algorithm that runs with O(n/ε log n/(γ ε) ) oracle calls to function h, and outputs a set S such that h(S) ≥ γ(1-ε)f(OPT)-c(OPT)-c(OPT)/γ(1-ε) log f(OPT)/c(OPT), where γ is the submodularity ratio of f. Existing algorithms for this problem either admit a worse approximation ratio or have quadratic runtime. We also present an approximation ratio of our algorithm for this problem with an approximate oracle of f. We validate our theoretical results through extensive empirical evaluations on real-world applications, including vertex cover and influence diffusion problems for submodular utility function f, and Bayesian A-Optimal design for weakly submodular f. Our experimental results demonstrate that our algorithms efficiently achieve high-quality solutions.
Yanhui Zhu, Samik Basu 0001, Aduri Pavan
CIKM1
2024 Submodular Optimization: Variants, Theory and Applications
abstract
Submodular function optimization is a fundamental tool in modeling complex interactions in machine learning and graph mining problems. We propose to study constrained submodular optimization to improve the current state of the art. Our goals are to design evolutionary algorithms with stronger approximation guarantees, the study of submodular maximization under submodular constraints, fairness in submodular optimization, and k-submodular optimization. We begin by exploring a basic submodular maximization problem in the context of viral marketing/information diffusion in social networks. From there, we broaden our scope to include a broader range of scenarios and formulate them into more optimization problems. Looking ahead, we plan to tackle scalable submodular optimization problems in fairness, dynamic constraints, dynamic streams, and distributed fashion. Lastly, we discuss a series of real-world applications that can be formulated as submodular optimization problems. In the future, we aim to apply algorithmic ideas to solve more real-world problems.
Yanhui Zhu
CIKM1
2024 Improved Evolutionary Algorithms for Submodular Maximization with Cost Constraints
Yanhui Zhu, Samik Basu 0001, Aduri Pavan
IJCAI1
2024 A Knowledge Graph Completion Method Based on Gated Adaptive Fusion and Conditional Generative Adversarial Networks
abstract
The multimodal knowledge graph completion (MMKGC) task aims to acquire accurate entity representations by learning different modal information about entities, which can be used to predict missing entities or relations in knowledge graphs (KGs). The lack of multimodal information accumulation and the noise in the various modal information collected leads to different modal information contributing differently to knowledge graph completion (KGC). Therefore, it becomes crucial to effectively fusion and fully utilize the multimodal information. To solve the above problems, we propose a knowledge graph completion method based on gated adaptive fusion and conditional generative adversarial networks (GAF-CGAN). GAF-CGAN is based on a gating mechanism for deeply extracting critical information in different modal features. After this, the model fuses the features by adapting the weights between different modal features. Ultimately, we utilize the relations in the triples as conditions to guide the generator in generating fake samples with specific meanings for adversarial training, enhancing the model’s robustness. We conduct link prediction experiments on two publicly available datasets, DB15K and FB15K-237, and the results show that our method significantly outperforms existing benchmark methods and effectively improves the model’s capacity for KGC.
Yanhui Zhu, Yuezhong Wu, Fangteng Man, Xujian Ying
TrustCom2
2024 Remote sensing scene classification using multi-domain sematic high-order network
Yanhui Zhu
Image Vis. Comput.2
2024 SCOREH+: A High-Order Node Proximity Spectral Clustering on Ratios-of-Eigenvectors Algorithm for Community Detection
abstract
The research on complex networks has achieved significant progress in revealing the mesoscopic features of networks. Community detection is an important aspect of understanding real-world complex systems. We present in this paper a High-order node proximity Spectral Clustering on Ratios-of-Eigenvectors (SCOREH+) algorithm for locating communities in complex networks. The algorithm improves SCORE and SCORE+ and preserves high-order transitivity information of the network affinity matrix. We optimize the high-order proximity matrix from the initial affinity matrix using the Radial Basis Functions (RBFs) and Katz index. In addition to the optimization of the Laplacian matrix, we implement a procedure that joins an additional eigenvector (the (k + 1)th leading eigenvector) to the spectrum domain for clustering if the network is considered to be a “weak signal” graph. The algorithm has been successfully applied to both real-world and synthetic data sets. The proposed algorithm is compared with state-of-art algorithms, such as ASE, Louvain, Fast-Greedy, Spectral Clustering (SC), SCORE, and SCORE+. To demonstrate the high efficacy of the proposed method, we conducted comparison experiments on eleven real-world networks and a number of synthetic networks with noise. The experimental results in most of these networks demonstrate that SCOREH+ outperforms the baseline methods. Moreover, by tuning the RBFs and their shaping parameters, we may generate state-of-the-art community structures on all real-world networks and even on noisy synthetic networks.
Yanhui Zhu, Fang Hu 0001, Lei Hsin Kuo, Jia Liu 0071
IEEE Trans. Big Data1
2023 A Framework for Adapting Offline Algorithms to Solve Combinatorial Multi-Armed Bandit Problems with Bandit Feedback
abstract
We investigate the problem of stochastic, combinatorial multi-armed bandits where the learner only has access to bandit feedback and the reward function can be non-linear. We provide a general framework for adapting discrete offline approximation algorithms into sublinear $\alpha$-regret methods that only require bandit feedback, achieving $\mathcal{O}\left(T^\frac{2}{3}\log(T)^\frac{1}{3}\right)$ expected cumulative $\alpha$-regret dependence on the horizon $T$. The framework only requires the offline algorithms to be robust to small errors in function evaluation. The adaptation procedure does not even require explicit knowledge of the offline approximation algorithm — the offline algorithm can be used as black box subroutine. To demonstrate the utility of the proposed framework, the proposed framework is applied to multiple problems in submodular maximization, adapting approximation algorithms for cardinality and for knapsack constraints. The new CMAB algorithms for knapsack constraints outperform a full-bandit method developed for the adversarial setting in experiments with real-world data.
Guanyu Nie, Yididiya Y. Nadew, Yanhui Zhu, Vaneet Aggarwal, Christopher J. Quinn
ICML3
2023 Size-constrained k-submodular maximization in near-linear time
abstract
We investigate the problems of maximizing k-submodular functions over total size constraints and over individual size constraints. k-submodularity is a generalization of submodularity beyond just picking items of a ground set, instead associating one of k types to chosen items. For sensor selection problems, for instance, this enables modeling of which type of sensor to put at a location, not simply whether to put a sensor or not. We propose and analyze threshold-greedy algorithms for both types of constraints. We prove that our proposed algorithms achieve the best known approximation ratios for both constraint types, up to a user-chosen parameter that balances computational complexity and the approximation ratio, while only using a number of function evaluations that depends linearly (up to poly-logarithmic terms) on the number of elements n, the number of types k, and the inverse of the user chosen parameter. Other algorithms that achieve the best-known deterministic approximation ratios require a number of function evaluations that depends linearly on the budget B, while our methods do not. We empirically demonstrate our algorithms’ performance in applications of sensor placement with k types and influence maximization with k topics.
Guanyu Nie, Yanhui Zhu, Yididiya Y. Nadew, Samik Basu 0001, Aduri Pavan, Christopher J. Quinn
UAI2
2023 Maximizing submodular functions under submodular constraints
abstract
We consider the problem of maximizing submodular functions under submodular constraints by formulating the problem in two ways: SCSKC and DiffC. Given two submodular functions f and g where f is monotone, the objective of SCSKC problem is to find a set S of size at most k that maximizes f(S) under the constraint that g(S) < theta, for a given value of theta. The problem of DiffC focuses on finding a set S of size at most k such that h(S) = f(S)-g(S) is maximized. It is known that these problems are highly inapproximable and do not admit any constant factor multiplicative approximation algorithms unless NP is easy. Known approximation algorithms involve data-dependent approximation factors that are not efficiently computable. We initiate a study of the design of approximation algorithms where the approximation factors are efficiently computable. For the problem of SCSKC, we prove that the greedy algorithm produces a solution whose value is at least (1-1/e)f(OPT) - A, where A is the data-dependent additive error. For the DiffC problem, we design an algorithm that uses the SCSKC greedy algorithm as a subroutine. This algorithm produces a solution whose value is at least (1-1/e)h(OPT)-B, where B is also a data-dependent additive error. A salient feature of our approach is that the additive error terms can be computed efficiently, thus enabling us to ascertain the quality of the solutions produced.
Madhavan R. Padmanabhan, Yanhui Zhu, Samik Basu 0001, Aduri Pavan
UAI2
2023 Prediction of Depression Severity Based on the Prosodic and Semantic Features With Bidirectional LSTM and Time Distributed CNN
abstract
Depression is increasingly impacting individuals both physically and psychologically worldwide. It has become a global major public health problem and attracts attention from various research fields. Traditionally, the diagnosis of depression is formulated through semi-structured interviews and supplementary questionnaires, which makes the diagnosis heavily relying on physicians’ experience and is subject to bias. However, since the pathogenic mechanism of depression is still under investigation, it is difficult for physicians to diagnose and treat, especially in the early clinical stage. As smart devices and artificial intelligence advance rapidly, understanding how depression associates with daily behaviors can be beneficial for the early stage depression diagnosis, which reduces labor costs and the likelihood of clinical mistakes as well as physicians bias. Furthermore, mental health monitoring and cloud-based remote diagnosis can be implemented through an automated depression diagnosis system. In this article, we propose an attention-based multimodality speech and text representation for depression prediction. Our model is trained to estimate the depression severity of participants using the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WOZ) dataset. For the audio modality, we use the collaborative voice analysis repository (COVAREP) features provided by the dataset and employ a Bidirectional Long Short-Term Memory Network (Bi-LSTM) followed by a Time-distributed Convolutional Neural Network (T-CNN). For the text modality, we use global vectors for word representation (GloVe) to perform word embeddings and the embeddings are fed into the Bi-LSTM network. Results show that both audio and text models perform well on the depression severity estimation task, with best sequence level$F_{1}$score of 0.9870 and patient-level$F_{1}$score of 0.9074 for the audio model over five classes (healthy, mild, moderate, moderately severe, and severe), as well as sequence level$F_{1}$score of 0.9709 and patient-level$F_{1}$score of 0.9245 for the text model over five classes. Results are similar for the multimodality fused model, with the highest$F_{1}$score of 0.9580 on the patient-level depression detection task over five classes. Experiments show statistically significant improvements over previous works.
Kaining Mao, Deborah Baofeng Wang, Rongqi Jiao, Yanhui Zhu, Tiansheng Zheng, Lei Qian 0002, Wei Lyu, Minjie Ye, Jie Chen 0002
IEEE Trans. Affect. Comput.6
2019 Support Vector Machine Optimized Using the Improved Fish Swarm Optimization Algorithm and Its Application to Face Recognition
abstract
Support vector machine (SVM) is always used for face recognition. However, kernel function selection is a key problem for SVM. This paper tries to make some contributions to this problem with focus on optimizing the parameters in the selected kernel function to improve the accuracy of classification and recognition of SVM. Firstly, an improved artificial fish swarm optimization algorithm (IAFSA) is proposed to optimize the parameters in SVM. In the improved version of artificial fish swarm optimization algorithm, the visual distance and the step size of artificial fish are adjusted adaptively. In the early stage of convergence, artificial fish are widely distributed, and the visual distance and step size take larger values to accelerate the convergence of the algorithm. In the later stage of convergence, artificial fish gathered gradually, and the visual distance and the step size were given small values to prevent oscillation. Then the optimized SVM is used to recognize face images. Simultaneously, in order to improve the accuracy rate of face recognition, an improved local binary pattern (ILBP) is proposed to extract features of face images. Numerical results show the advantage of our new algorithm over a range of existing algorithms.
Wenqiu Zhu, Haixing Bao, Zhigao Zeng, Zhiqiang Wen, Yanhui Zhu, Huazheng Xiang
Int. J. Pattern Recognit. Artif. Intell.5
2012 Translation and localization of SNOMED CT in China: A pilot study
Yanhui Zhu, Huiting Pan, Ana Chen, Ulrich Andersen, Shuxiang Pan, Lixin Tian, Jianbo Lei
Artif. Intell. Medicine1