VLDB 2026 Research / reviewers in the wild / expert
Yaxin Zhao
dblp:246/8132
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HalluClean: A Unified Framework to Combat Hallucinations in LLMsabstractLarge language models (LLMs) have achieved impressive performance across a wide range of natural language processing tasks, yet they often produce hallucinated content that undermines factual reliability. To address this challenge, we introduce HalluClean, a lightweight and task-agnostic framework for detecting and correcting hallucinations in LLM-generated text. HalluClean adopts a reasoning-enhanced paradigm, explicitly decomposing the process into planning, execution, and revision stages to identify and refine unsupported claims. It employs minimal task-routing prompts to enable zero-shot generalization across diverse domains, without relying on external knowledge sources or supervised detectors. We conduct extensive evaluations on five representative tasks—question answering, dialogue, summarization, math word problems, and contradiction detection. Experimental results show that HalluClean significantly improves factual consistency and outperforms competitive baselines, demonstrating its potential to enhance the trustworthiness of LLM outputs in real-world applications. Yaxin Zhao |
AAAI | 1 |
| 2025 | Explanation-Based Anonymization Methods for Motion Privacy
Thomas Carr 0001, Yaxin Zhao, Depeng Xu 0001, Aidong Lu |
PAKDD (4) | 2 |
| 2025 | Heterogeneous Privacy-Preserving Blockchain-Enabled Federated Learning for Social FintechabstractSocial fintech integrates financial technology with social networking to enhance financial services’ accessibility and personalization by leveraging social interactions and user data. This approach raises privacy security concerns, particularly in application based on centralized artificial intelligence systems. To address these issues, blockchain-enabled federated learning (BEFL) offers a decentralized solution, improving robustness and privacy but facing challenges such as privacy attacks and heterogeneous crypto system. In response, a novel PKI and identity-based heterogeneous authenticated asymmetric group key agreement (PKI-IB-HAAGKA) protocol was proposed, which resolves crypto system heterogeneity issues. What's more, a PKI and identity-based heterogeneous batch multisignature (PKI-IB-HBMS) was proposed as a building block of PKI-IB-HAAGKA. This article presents the heterogeneous privacy-preserving blockchain-enabled federated learning (HPP-BEFL) system, designed to enhance privacy, security, and efficiency in social fintech applications. It effectively mitigates man-in-the-middle and inference attacks while improving overall system performance. Through security analysis and experiment results, it is demonstrated that the proposed PKI-IB-HAAGKA, PKI-IB-HBMS, and HPP-BEFL are provably secure and highly efficient, which can be applied to large-scale heterogeneous privacy-preserving model training scenarios. Hu Xiong, Yaxin Zhao, Abubaker Wahaballa, Kuo-Hui Yeh |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | DA-FL: Blockchain Empowered Secure and Private Federated Learning With Anonymous AuthenticationabstractFederated learning (FL) is a secure multiparty machine learning that addresses the issue of data silos by allowing nodes to train locally. Nonetheless, the lack of trusted environments, node supervision, and privacy protection measures in centralized FL limit its large-scale promotion. To address these issues, a blockchain-based decentralized FL framework is proposed, namely, decentralized federated learning with node anonymous authentication (DA-FL). Specifically, DA-FL introduces blockchain for local model storage and global model aggregation in the absence of centralized server, and uses differential privacy to reduce the risk of model privacy leakage. In addition, a consensus mechanism proof of accuracy is designed to effectively reduce the computational load of consensus and mitigate the impact of low-quality models on the aggregation results. To achieve node supervision, distributed key generation and revocable ring signature technologies are being integrated. This ensures the anonymous authentication of nodes while also allowing for the revocation of the anonymity of malicious nodes when necessary. Finally, the security and functionality of DA-FL are evaluated through simulation experiments conducted on real datasets. The numerical results show that the proposed FL scheme has significant performance advantages over other schemes. Hu Xiong, Yaxin Zhao, Kuo-Hui Yeh |
IEEE Trans. Reliab. | 2 |
| 2024 | How accessibility affects other quality attributes of software? A case study of GitHub
Yaxin Zhao, Lina Gong, Wenhua Yang 0001, Yu Zhou 0010 |
Sci. Comput. Program. | 1 |
| 2024 | A Conditional Privacy-Preserving Mutual Authentication Protocol With Fine-Grained Forward and Backward Security in IoVabstractWith the rise of intelligent transportation, various mobile value-added services can be provided by the service provider (SP) in the Internet of Vehicles (IoV). To guarantee the dependability of services, it is essential to implement a mutual authentication protocol between the vehicles and the SP. Existing mutual authentication protocols to secure the communication between the SP and the vehicle raise challenges such as providing fine-grained forward security for the SP and achieving backward security for the vehicle. To handle these challenges, this paper proposes a conditional privacy-preserving mutual authentication protocol featured with fine-grained forward security and backward security for IoV, which can be implemented via two building blocks we have constructed. Specifically, we present a new puncturable signature (PS) scheme without false-positive probability and the update of the public key as well as the first proxy re-signature scheme with parallel key-insulation (PKI-PRS). What’s more, both the proposed PKI-PRS and PS still have interest beyond this protocol. Then, an anonymous mutual authentication protocol with resistance to key leakage is constructed by incorporating the above signature schemes. The proposed protocol not only provides fine-grained forward security for the SP, but also ensures forward security as well as backward security for the vehicles. Besides, the approach to achieving anonymous authentication can efficiently provide conditional privacy-preserving for the vehicles. With the support of the random oracle model and experimental simulations, the formal security proof and the superiority of the proposed protocol is explicitly given. Hu Xiong, Ting Yao 0002, Yaxin Zhao, Lingxiao Gong, Kuo-Hui Yeh |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Memory-Augmented Contrastive Learning for Talking Head GenerationabstractGiven one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The same speech clip can generate multiple possible lip and head movements, that is, there is no one-to-one mapping relationship between them. To overcome this problem, we propose a Speech Feature Extractor (SFE) based on memory-augmented self-supervised contrastive learning, which introduces the memory module to store multiple different speech mapping results. In addition, we introduce the Mixed Density Networks (MDN) into the landmark regression task to generate multiple predicted facial landmarks. Extensive qualitative and quantitative experiments show that the quality of our facial animation is significantly superior to that of the state-of-the-art (SOTA). The code has been released at https://github.com/Yaxinzhao97/MACL.git. Jianrong Wang, Yaxin Zhao, Hongkai Fan, Li Liu 0036 |
ICASSP | 2 |
| 2023 | TSFCloNet: Clothing Classification Algorithm Based on Two-Stream Network StructureabstractIn the fashion field, with the increasing diversity of clothing types and styles, accurate clothing classification becomes very important. However, the complex background and diverse styles of clothing images bring challenges to feature extraction. Classification based on texture features alone may focus too much on details and ignore the overall shape information, thus reducing the accuracy and stability of classification. In order to achieve fast and accurate clothing classification, this paper proposes a two-stream network structure clothing classification algorithm based on shape texture features and multi-feature fusion (TSFCloNet). Its main core is as follows: 1) using the two-stream network structure to extract texture and shape features from the input data set respectively; 2) in the shape feature extraction stream, the clothing shape acquisition module is first used to process the input clothing data set, and the obtained clothing shape data set is input into the ShapeNet feature extraction module to obtain shape feature information; 3) the FFCE (Feature Fusion Channel Enhancement) module is used to fuse the features obtained by the two branches of the structure respectively, and the DSAConv module is used to enhance feature extraction, and the final features are sent to the trained classifier to obtain the clothing style classification results. A large number of experimental results show that the proposed TSFCloNet network achieves higher classification accuracy when dealing with diverse and changeable fashion styles, significantly improving the performance of fashion image classification. Minghua Jiang, Yaxin Zhao, Li Liu 0047, Feng Yu 0017 |
ICPADS | 3 |
| 2023 | Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksabstractGiven an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made some achievements in the past, most of them only focused on lip and audio synchronization and lack the ability to reproduce the facial expressions of the target person. To this end, we propose a talking head generation model consisting of a Memory-Sharing Emotion Feature extractor (MSEF) and an Attention-Augmented Translator based on U-net (AATU). Firstly, MSEF can extract implicit emotional auxiliary features from audio to estimate more accurate emotional face landmarks. Secondly, AATU acts as a translator between the estimated landmarks and the photo-realistic video frames. Extensive qualitative and quantitative experiments have shown the superiority of the proposed method to the previous works. Codes will be made publicly available. Jianrong Wang, Yaxin Zhao, Li Liu 0036 |
INTERSPEECH | 2 |
| 2022 | Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Captureabstract3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To address these problems, we propose a new 3D virtual try-on network via multi-scale characteristic capture (VTON-MC), which can produce an exact 3D model with the generated photo-realistic monocular image. The main processes are as follows: 1) predicting the human semantic-map and aligning the in-shop garment in the human pose using the appearance flow method, 2) synthesizing the human body and the warped garment to gain the image try-on result, and 3) estimating the human double-depth map of the image try-on result to reconstruct desired 3D try-on mesh by designed Depth Estimation Network (DEN). Extensive experiments on existing benchmark datasets demonstrate that VTON-MC outperforms state-of-the-art approaches efficiently. Chenghu Du, Feng Yu 0017, Minghua Jiang, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
ICASSP | 4 |
| 2022 | High fidelity virtual try-on network via semantic adaptation and distributed componentizationabstractImage-based virtual try-on systems have significant commercial value in online garment shopping. However, prior methods fail to appropriately handle details, so are defective in maintaining the original appearance of organizational items including arms, the neck, and in-shop garments. We propose a novel high fidelity virtual try-on network to generate realistic results. Specifically, a distributed pipeline is used for simultaneous generation of organizational items. First, the in-shop garment is warped using thin plate splines (TPS) to give a coarse shape reference, and then a corresponding target semantic map is generated, which can adaptively respond to the distribution of different items triggered by different garments. Second, organizational items are componentized separately using our novel semantic map-based image adjustment network (SMIAN) to avoid interference between body parts. Finally, all components are integrated to generate the overall result by SMIAN. A priori dual-modal information is incorporated in the tail layers of SMIAN to improve the convergence rate of the network. Experiments demonstrate that the proposed method can retain better details of condition information than current methods. Our method achieves convincing quantitative and qualitative results on existing benchmark datasets. Chenghu Du, Feng Yu 0017, Minghua Jiang, Ailing Hua, Yaxin Zhao, Tao Peng 0006, Xinrong Hu |
Comput. Vis. Media | 5 |
| 2020 | M-Sosanet: An Efficient Convolution Network Backbone For Embedding DevicesabstractIn this paper, we build a lightweight convolution neural network M-SOSAnet that combines efficiency and accuracy for edge devices. Just as DenseNet connects each layer to every other layer in the neural network. Although dense connection can effectively keep information between the middle layers, the increasing input channels by dense connection leads to resource consumption, which greatly increases the amount of parameters and is inefficient. MobileNet series use depthwise separable convolutions, which reduces the amount of calculation and parameters, but will reduce the accuracy. ESPnetv2 uses group pointwise and depthwise dilated separable convolution to learn representation from a large receptive field with fewer flops and parameters but will get a little accuracy decrease. In order to solve these problems, we propose two module: 1. Mobile SOSA Module 2. Downsampling Block with scoring mechanism, and also introduce the attention mechanism. Compare to some past methods, our model M-SOSAnet has lower flops, higher accuracy, maintains the similar amount of parameters. We evaluated model in the areas of image classification, and semantic segmentation to prove that our methods have better performance. Tangkun Zhang, Jichao Jiao, Chengkai Zhang, Yaxin Zhao, Xinping Chen |
ICIP | 4 |
| 2020 | MANet: Multimodal Attention Network based Point-View Fusion for 3D Shape Recognitionabstract3D shape recognition has attracted more and more attention as a task of 3D vision research. The proliferation of 3D data encourages various deep learning methods based on 3D data. Now there have been many deep learning models based on point-cloud data or multi-view data alone. However, in the era of big data, integrating data of two different modals to obtain a unified 3D shape descriptor is bound to improve the recognition accuracy. Therefore, this paper proposes a fusion network based on multimodal attention mechanism for 3D shape recognition. Considering the limitations of multi-view data, we introduce a soft attention scheme, which can use the global point-cloud features to filter the multi-view features, and then realize the effective fusion of the two features. More specifically, we obtain the enhanced multi-view features by mining the contribution of each multi-view image to the overall shape recognition, and then fuse the point-cloud features and the enhanced multi-view features to obtain a more discriminative 3D shape descriptor. We have performed relevant experiments on the ModelNet40 dataset, and experimental results verify the effectiveness of our method. Yaxin Zhao, Jichao Jiao, Ning Li 0015 |
ICPR | 1 |
| 2019 | Combining VSM and BTM to Improve Requirements Trace Links GenerationabstractTrace links between software artifacts provide available traceability information and in-depth insights for different stakeholders.Unfortunately, establishing trace links is a fallible, tedious, and labor-intensive task.To alleviate these problems, many Information Retrieval (IR) methods, such as Vector Space Model (VSM), Latent Semantic Indexing (LSI) and their variants, have been proposed to establish trace links automatically.In recent years, short-text artifacts (or even lack of documentation) become a new trend as more and more software systems are developed abiding by agile methodologies.It makes the effects of traditional IR-based trace links generation methods even worse.In this paper, Biterm Topic Model (BTM), which is good at dealing with short text, is introduced to solve the problem.A hybrid method combining VSM and BTM is proposed to generate requirements trace links.The empirical experiments conducted on three real and frequently-used datasets indicate that the hybrid method can achieve better performance, and the results can reach the "acceptable level" directly. Bangchao Wang, Rong Peng, Yaxin Zhao |
SEKE | 4 |