VLDB 2026 Research / reviewers in the wild / expert
Zhijie Xie
dblp:188/0735
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Password Guessing Based on Hidden Weak Password AnalysisabstractPassword has become the mainstream method of authentication today. To improve password security, researchers evaluate the strength of target password datasets through early brute-force attacks to current password guessing methods, aiming to help users reduce the use of weak passwords. With users becoming more aware of security, they make local variations on weak passwords to improve the password strength while being easy to remember. These transformations render passwords more complex and enhance the score in password strength meter. However, such variations do not genuinely enhance password security, as human habits tend to converge. This allows attackers to deduce the modification patterns and consequently crack these passwords. Motivated by this, this paper defines the hidden weak passwords, a local variant of explicit weak passwords, which appear to enhance password security yet remain vulnerable. We systematically analyze transformation behavior between explicit and hidden weak passwords. Then we design an automated rule generation algorithm to identify hidden weak passwords and generate transformation rules. Based on automatically mined rules, we generate a large number of password guesses and fuses them with existing methods to improve password guessing performance. Finally, we demonstrate the effectiveness of the proposed method through password guessing experiments on eight real-world datasets, where the cracking rate improves on all five state-of-the-art methods. Min Zhang 0054, Zhijie Xie, Shasha Guo 0001, Yuliang Lu, Fan Shi 0003, Yi Shen 0012 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Efficient Long Speech Sequence Modelling for Time-Domain Depression Level EstimationabstractDepression significantly affects emotions, thoughts, and daily activities. Recent research indicates that speech signals contain vital cues about depression, sparking interest in audiobased deep-learning methods for estimating its severity. However, most methods rely on time-frequency representations of speech which have recently been criticized for their limitations due to the loss of information when performing time-frequency projections, e.g. Fourier transform, and Mel-scale transformation. Furthermore, segmenting real-world speech into brief intervals risks losing critical interconnections between recordings. Additionally, such an approach may not adequately reflect real-world scenarios, as individuals with depression often pause and slow down in their conversations and interactions. Building on these observations, we present an efficient method for depression level estimation using long speech signals in the time domain. The proposed method leverages a state space model coupled with the dual-path structure-based long sequence modelling module and temporal external attention module to reconstruct and enhance the detection of depression-related cues hidden in the raw audio waveforms. Experimental results on the AVEC2013 and AVEC2014 datasets show promising results in capturing consequential long-sequence depression cues and demonstrate outstanding performance over the state-of-the-art. Shuanglin Li, Zhijie Xie, Syed M. Naqvi |
ICASSP | 2 |
| 2025 | Client Selection for Federated Policy Optimization with Environment HeterogeneityabstractThe development of Policy Iteration (PI) has inspired many recent algorithms for Reinforcement Learning (RL), including several policy gradient methods that gained both theoretical soundness and empirical success on a variety of tasks. The theory of PI is rich in the context of centralized learning, but its study under the federated setting is still in the infant stage. This paper investigates the federated version of Approximate PI (API) and derives its error bound, taking into account the approximation error introduced by environment heterogeneity. We theoretically prove that a proper client selection scheme can reduce this error bound. Based on the theoretical result, we propose a client selection algorithm to alleviate the additional approximation error caused by environment heterogeneity. Experiment results show that the proposed algorithm outperforms other biased and unbiased client selection methods on the federated mountain car problem, the MuJoCo Hopper problem, and the SUMO-based autonomous vehicle training problem by effectively selecting clients with a lower level of heterogeneity from the population distribution. Zhijie Xie, Shenghui Song 0001 |
J. Mach. Learn. Res. | 1 |
| 2025 | Understanding and Characterizing the Adoption of Internationalized Domain Names in PracticeabstractInternationalized Domain Names (IDNs) allow users to access the internet using domain names in their native languages. This technology provides significant convenience for non-English speaking users. However, despite the widespread acceptance and use of IDNs, the risks associated with using IDNs remain unclear in practice, such as the IDN homograph problem. To address this issue, we conduct a systematic analysis of the IDN homograph problem and explore the adoption characteristics of IDNs in practice. Specifically, we design and implement an effective IDN analysis framework, named as IDNMon. We perform a large-scale measurement study covering 863 top-level domain zone files and historical top lists based on IDNMon. Our findings indicate that the IDN registration and usage in Europe exceeds that in East Asia. Our results confirm that the IDN homograph problem is universal (12.32% of 2,623,161 IDNs face this problem), which raises serious challenges when designing protection strategies for browsers. Our work provides new insights into the adoption of IDNs in practice, contributes to a better understanding, and promotes the development of IDNs. Chengxi Xu, Fan Shi 0003, Min Zhang 0054, Yuwei Li 0002, Zhijie Xie |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | MFSFormer: A Novel End-to-End Rotating Machinery Fault Diagnosis Framework for Noise and Small SamplesabstractConvolutional neural networks and transformers have both achieved remarkable success in the field of fault diagnosis owing to their proficiency in extracting local and global features. However, in real industrial production, the diagnostic performance of many methods is often limited by factors such as environmental noise and sample quantity. To address these challenges, this article presents a novel fault diagnosis framework called MFSFormer. First, it employs embedded convolutional layers with kernels of various sizes to extract multiscale receptive field features from vibration signals. Second, a Fuse-Shuffle attention block is employed to capture dependencies between feature channels and windows, facilitating the implementation of subfeature flow. The case studies conducted on two public datasets demonstrate that the proposed method achieves an average diagnostic accuracy that surpasses the suboptimal method by 3.91% and 7.00% in noisy environments and small-sample conditions, respectively. The results indicate that the proposed method not only enhances the robustness of fault diagnosis in noisy environments, but also improves diagnostic accuracy under limited sample conditions, and the method has practical value. Xueyi Li 0004, Sixin Li, Feibin Zhang, Zhijie Xie, Fulei Chu |
IEEE Trans. Reliab. | 6 |
| 2024 | GuessFuse: Hybrid Password Guessing With Multi-ViewabstractPassword guessing is a primary method for password strength evaluation. Despite various password guessing models have been proposed, there is still a significant gap between their guessing effectiveness and the actual cracking capabilities of attackers. Integrating multiple models for password guessing, also known as hybrid password guessing, could better capture the cracking capabilities of real attackers. However, the reason why hybrid password guessing can enhance cracking capabilities, and how to effectively integrate multiple heterogeneous password guessing models, are still not well understood. To address these issues, this paper draws inspiration from the concept of multi-view learning. We regard the guess lists generated by various password guessing models as multiple views of the data. Through a comprehensive analysis of these guess lists, we have identified the key reason why hybrid password guessing can enhance the cracking capabilities: integrating more diverse views allows for the coverage of a wider range of heterogeneous password characteristics, and provides more detailed information on effective password distributions. Based on the these findings, we propose a new hybrid password guessing framework, namedGuessFuse.GuessFuseemploys the multi-view subset extraction module and segment splitting selection module to accurately extract and reorganize the effective password from multiple guess lists. Experimental results on six large-scale datasets demonstrate the effectiveness ofGuessFuse. By combining two (resp. five) guess lists,GuessFuseoutperforms its foremost counterparts by an average of 11.00% ~ 59.62% (resp. 4.70% ~ 17.66%) within 107guesses.GuessFusecan effectively improve the cracking success rate under a limited number of guesses, approaching the actual cracking capabilities of attackers. Zhijie Xie, Fan Shi 0003, Min Zhang 0054, Huimin Ma 0004, Huaixi Wang, Zhenhan Li |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | FedKL: Tackling Data Heterogeneity in Federated Reinforcement Learning by Penalizing KL DivergenceabstractOne of the fundamental issues for Federated Learning (FL) is data heterogeneity, which causes accuracy degradation, slow convergence, and the communication bottleneck issue. Although the impact of data heterogeneity on supervised FL has been widely studied, the related investigation for Federated Reinforcement Learning (FRL) is still in its infancy. In this paper, we first define the type and level of data heterogeneity for FRL systems. By inspecting the connection between the global and local objective functions, we prove that local training can benefit the global objective, if the local update is properly penalized by the total variation (TV) distance between the local and global policies. A necessary condition for the global policy to be learn-able from the local environments is also derived, which is directly related to the heterogeneity level. Based on the theoretical result, a Kullback-Leibler (KL) divergence based penalty is proposed to directly constrain the model outputs in the distribution space and the convergence proof of the proposed algorithm is also provided. By jointly penalizing the divergence of the local policy from the global policy with a global penalty and penalizing each iteration of the local training with a local penalty, the proposed method achieves a better trade-off between training speed (step size) and convergence. Experiment results on two popular Reinforcement Learning (RL) experiment platforms demonstrate the advantage of the proposed algorithm over existing methods in accelerating and stabilizing the training process with heterogeneous data. Zhijie Xie, Shenghui Song 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2020 | A New Targeted Password Guessing Model
Zhijie Xie, Anqi Yin, Zhenhan Li |
ACISP | 1 |
| 2020 | Modified Password Guessing Methods Based on TarGuess-IabstractTarGuess − I is a leading online targeted password guessing model using users’ personally identifiable information (PII) proposed at ACM CCS 2016 by Wang et al. It has attracted widespread attention in password security owing to its superior guessing performance. Yet, after analyzing the users’ vulnerable behaviors of using popular passwords and constructing passwords with users’ PII, we find that this model does not take into account popular passwords, keyboard patterns, and the special strings. The special strings are the strings related to users but do not appear in the users’ demographic information. Thus, we propose TarGuess − I + K P X , a modified password guessing model with three semantic methods, including (1) identifying popular passwords by generating top-300 lists from similar websites, (2) recognizing keyboard patterns by relative position, and (3) catching the special strings by extracting continuous characters from user-generated PII. We conduct a series of evaluations on six large-scale real-world leaked password datasets. The experimental results show that our modified model outperforms TarGuess − I by 2.62% within 100 guesses. Zhijie Xie, Min Zhang 0054, Yuqi Guo 0002, Zhenhan Li, Hongjun Wang 0010 |
Wirel. Commun. Mob. Comput. | 1 |
| 2017 | A periodic repair algorithm for dynamic scheduling in home health care using agent-based modelabstractThis paper presents a periodic repair algorithm for dynamic home visit scheduling with the objective of reducing the service cost in home healthcare. In our setting, the health care agency needs to assign practitioners to cover all home visit requests and, at the same time, respect practitioner's availability, eligibility and patient's visit time constraints. We consider a dynamic scheduling problem in which dynamic events occur along with the execution of existing schedules. The occurrence of dynamic events, such as newly added requests, visit cancellations or availability changes, will render the existing schedule infeasible. We propose a repair-based rescheduling algorithm which periodically revises the existing schedule to accommodate the collection of dynamic events during a certain time period. An agent-based simulation model is developed using AnyLogic to validate the efficiency of the proposed algorithm. Simulation results show that the service cost of the solutions generated by the proposed algorithm is on average 14% lower than that of the solutions generated by first-come-first-serve policy which is commonly used in home health care scheduling practice. Zhijie Xie |
CSCWD | 1 |