VLDB 2026 Research / reviewers in the wild / expert
Zhenhua Yang
dblp:42/4517
· DBLP profile ↗
25ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLMabstractInscriptions are invaluable cultural heritage, yet centuries of degradation (e.g., fractures, erosion, oxidation) have rendered many partially illegible. Existing Historical Inscription Restoration (HIR) methods rely on task-separated pipelines with irreversible error accumulation and patch-based generation that sacrifices page-level consistency. Therefore, we present UniHIR, the first unified MLLM for end-to-end historical inscription restoration. It integrates two novel designs, Draft-Guided Localization and Hierarchical Self-Refinement, to enable accurate damage localization and illegible-content prediction via iterative reasoning and self-correction. This unified approach enables true page-level restoration with consistent typography and style. To support training under high-resolution inputs and long sequences, we design UHIRFactory and construct HIRBench, enabling step-wise, memory-efficient instruction tuning with step-aware annotations for intermediate drafts and refinements. Experiments demonstrate that UniHIR achieves superior performance in both text restoration accuracy and appearance restoration quality, validating that HIR can be effectively tackled by a standalone model in a unified manner. The model and code are available at https://github.com/ZZXF11/UniHIR. Yuyi Zhang 0002, Junle Liu, Peirong Zhang 0001, Jianliang Liu, Zhenhua Yang |
ACL (1) | 5 |
| 2026 | Three-Level Pressure-Based Authentication on Touch ScreensabstractSince pressure is invisible in nature, it has been applied to enhance the authentication security. However, previous work focused on only two-level pressure detection, which makes it feasible for an attacker to guess a pressure level from a user's pressing behavior in the shoulder surfing attack. To mitigate the above risk, this article extends pressure-based authentication from two levels to three levels. It systematically evaluates its usability and security with 98 young adults in a mid-west university in two user studies. The first study with 67 participants showed that a three-level pressure-based password did not increase the difficulty of memorization, and it was more resistant to the shoulder surfing attack than a two-level pressure-based password (two-level = 79.39% successful attacks versus three-level = 44.71% successful attacks). However, three-level passwords have a higher false negative rate than two-level passwords, which implies that users lack a consistent pressing pattern on the medium level. To address the above issue, we designed an adaptive training process that assists users in forming a consistent pressing pattern. The training process features an automatic detection of consistent pressing patterns, which potentially reduces the training length. The second user study, with 31 participants, indicated that the false negative rate in three-level passwords is decreased significantly after the training from 59.90% to 2.15%, while maintaining a reasonable false acceptance rate of 20%. Zhenhua Yang, Juan Li 0004, Tongxin Shi |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2026 | Robust GNSS Positioning via Variational Bayesian Factor Graph Optimization With Dirichlet Process Mixture Models
Zhenhua Yang, Yongqing Wang 0002, Yuyao Shen |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Predicting the Original Appearance of Damaged Historical DocumentsabstractHistorical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily focus on binarization, enhancement, etc., neglecting the repair of these damages. To this end, we present a new task, termed Historical Document Repair (HDR), which aims to predict the original appearance of damaged historical documents. To fill the gap in this field, we propose a large-scale dataset HDR28K and a diffusion-based network DiffHDR for historical document repair. Specifically, HDR28K contains 28,552 damaged-repaired image pairs with character-level annotations and multi-style degradations. Moreover, DiffHDR augments the vanilla diffusion framework with semantic and spatial information and a meticulously designed character perceptual loss for contextual and visual coherence. Experimental results demonstrate that the proposed DiffHDR trained on HDR28K significantly surpasses existing approaches and exhibits remarkable performance in handling real scenarios. Notably, DiffHDR can also be extended to document editing and text block generation, showcasing its high flexibility and generalization capacity. We believe this study could pioneer a new direction of document processing and contribute to the inheritance of invaluable cultures and civilizations. Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 0002, Chongyu Liu |
AAAI | 1 |
| 2025 | Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationabstractYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuyi Zhang 0002, Peirong Zhang 0001, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo |
ACL (1) | 3 |
| 2025 | RAPID: Reliable and efficient Automatic generation of submission rePortIng checklists with large language moDelsabstractOBJECTIVE: To evaluate an automated reporting checklist generation tool using large language models and retrieval augmentation generation technology, called RAPID. MATERIALS AND METHODS: This study utilized large language models to develop a retrieval augmentation generation architecture. To assess its performance, a total of 91 published journal articles were collected and manually annotated in accordance with the CONSORT and CONSORT-AI medical reporting guidelines. These articles comprised 50 randomized controlled trials conducted without AI intervention and 41 randomized controlled trials that incorporated AI tools. RESULTS: Fifty RCT articles without the intervention of AI tools and 41 RCT articles with the intervention of AI tools were collected as CONSORT and CONSORT-AI datasets. All of the CONSORT reporting items (37) were included in the tool. RAPID achieved a high average accuracy rate of 92.11% and a content consistency score of 81.14% on the CONSORT dataset. Of the CONSORT-AI reporting items, 11 items related to the intervention of AI tools were included in the tool. RAPID achieved an average accuracy of 83.81% with a content consistency score of 72.51% on the CONSORT-AI dataset. DISCUSSION: RAPID may effectively save time and improve working efficiency for different user groups such as medical authors, researchers, editors, and reviewers. CONCLUSION: RAPID has strong scalability, which can be easily adapted to different medical reporting guidelines without transfer learning on a large dataset. RAPID got state-of-the-art performance on 2 datasets for 2 different checklists compared to other methods. Xufei Luo, Zhenhua Yang, Bingyi Wang, Long Ge, Zhaoxiang Bian, Yaolong Chen, Lu Zhang 0061, Dongrui Peng, Honghao Lai, Minjie Duan, Shilin Tang |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Multi-dimensional Topological Association Strengthening Clustering NetworkabstractDeep graph clustering, as a fundamental task in data mining, has attracted widespread attention. Recently, excellent performance has been achieved by integrating graph structure and node attributes to generate consensus latent embeddings. However, existing clustering methods are limited by redundant information and unreliable clustering distribution, which hinders the discriminative power of the latent embeddings. To address this issue, we propose a novel deep graph clustering framework called Multi-dimensional Topological Association Strengthening Clustering Network (MTASCN). Specifically, we design a Multi-dimensional Feature Association Mechanism (MFAM), which extracts the competitive or cooperative relationship between features to alleviate the interference of redundant features and enhance the dominant features. In addition, we develop a Structure-oriented Multi-order Loss Module (SMLM) that reinforces the generation of clustering distribution under reliable structure information guidance by calculating the multi-order similarity between the latent embeddings and the original graph structure. Extensive experiments on five benchmark datasets have demonstrated that MTASCN consistently outperforms other clustering methods. Mengzhe Sun, Renda Han, Moxuan Zeng, Zhenhua Yang, Jingxin Liu 0006, Wen Xin, Jingmei Feng |
Neural Process. Lett. | 6 |
| 2025 | MegaHan97K: A large-scale dataset for mega-category Chinese character recognition with over 97K categories
Yuyi Zhang 0002, Yongxin Shi, Peirong Zhang 0001, Yixin Zhao, Zhenhua Yang |
Pattern Recognit. | 5 |
| 2025 | HierCode: A lightweight hierarchical codebook for zero-shot Chinese text recognition
Yuyi Zhang 0002, Dezhi Peng, Peirong Zhang 0001, Zhenhua Yang, Zhibo Yang 0003, Cong Yao |
Pattern Recognit. | 5 |
| 2025 | Dual Feature Enhancement Graph Clustering Network
Renda Han, Mengzhe Sun, Zhenhua Yang, Jingxin Liu 0006 |
Pattern Recognit. Lett. | 6 |
| 2025 | SDRS: Sentiment-Aware Disentangled Representation Shifting for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) aims to leverage the complementary information from multiple modalities for affective understanding of user-generated videos. Existing methods mainly focused on designing sophisticated feature fusion strategies to integrate the separately extracted multimodal representations, ignoring the interference of the information irrelevant to sentiment. In this paper, we propose to disentangle the unimodal representations into sentiment-specific and sentiment-independent features, the former of which are fused for the MSA task. Specifically, we design a novel Sentiment-aware Disentangled Representation Shifting framework, termed SDRS, with two components.Interactive sentiment-aware representation disentanglementaims to extract sentiment-specific feature representations for each nonverbal modality by considering the contextual influence of other modalities with the newly developed cross-attention autoencoder.Attentive cross-modal representation shiftingtries to shift the textual representation in a latent token space using the nonverbal sentiment-specific representations after projection. The shifted representation is finally employed to fine-tune a pre-trained language model for multimodal sentiment analysis. Extensive experiments are conducted on three public benchmark datasets, i.e., CMU-MOSI, CMU-MOSEI, and CH-SIMS. The results demonstrate that the proposed SDRS framework not only obtains state-of-the-art results based solely on multimodal labels but also outperforms the methods that additionally require the labels of each modality. Sicheng Zhao, Zhenhua Yang, Henglin Shi, Lingpengkun Meng, Bing Qin 0001, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningabstractAutomatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations. To address these issues, we propose FontDiffuser, a diffusion-based image-to-image one-shot font generation method, which innovatively models the font imitation task as a noise-to-denoise paradigm. In our method, we introduce a Multi-scale Content Aggregation (MCA) block, which effectively combines global and local content cues across different scales, leading to enhanced preservation of intricate strokes of complex characters. Moreover, to better manage the large variations in style transfer, we propose a Style Contrastive Refinement (SCR) module, which is a novel structure for style representation learning. It utilizes a style extractor to disentangle styles from images, subsequently supervising the diffusion model via a meticulously designed style contrastive loss. Extensive experiments demonstrate FontDiffuser's state-of-the-art performance in generating diverse characters and styles. It consistently excels on complex characters and large style changes compared to previous methods. The code is available at https://github.com/yeungchenwa/FontDiffuser. Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 0002, Cong Yao |
AAAI | 1 |
| 2024 | Evolutionary Dynamic Optimization-Based Calibration Framework for Agent-Based Financial Market SimulatorsabstractThe agent-based financial market simulators serve as an important validation tool for trading strategies. For high-fidelity simulation, it is pivotal to calibrate the parameters of a simulator so that the generated simulation data resembles the observed real market data of interest. In traditional calibration methods, it is typical that the parameters of the simulator are set to be time-invariant. However, the dynamic nature of the real financial market introduces various variability into the behaviors of the market participants over different time intervals, posing in-herent limitations to the traditional methods. A more reasonable approach might involve employing a simulator with time-variant parameters. This suggests that the model parameters can be dynamically adjusted at different stages of the simulation to adapt to the evolving market. Consequently, the calibration problem of the financial market simulators can be treated as a dynamic optimization problem. To dynamically calibrate the simulators, we introduce an Evolutionary Dynamic Optimization (EDO) framework. By monitoring the changes of the best fitness, the whole simulation time interval is adaptively divided into multiple stages. Then the Negatively Correlated Search (NCS) algorithm is employed to effectively adjust the parameters at different simulation stages to better simulate the real financial market. Empirical results on both synthetic and real data verify that our dynamic calibration framework significantly outperforms traditional calibration methods that fixing a parameter for the whole simulation interval. The proposed strategy of detecting dynamic changes is also shown to be more reliable than the naive method of manually segmenting stages. In terms of calibration time, our proposed method significantly improves by nearly 93% compared to the fixed parameter setting, and approximately 61% compared to manual segmentation calibration. Zhenhua Yang, Muyao Zhong, Peng Yang 0008 |
CEC | 1 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 2 |
| 2024 | Improving Zero-Shot Coordination with Diversely Rewarded Partner AgentsabstractZero-shot coordination studies the training of well-generalizing human-AI coordination agents in the scenario where human data is unavailable. To obtain a coordination agent generalize to unseen humans, prevailing methods generate a population of partner agents as proxy models of human partners and then train a coordination agent with these partner agents. Constructed partner agents are expected to be as diverse as possible to cover a wide range of human behaviors, preventing a distribution shift between training and testing stages. Recent works concentrate on studying effective methods of creating a group of high-reward while diverse partner agents to model unseen human partners. However, the resulting high-reward partner agents do not accurately reflect real-world situations, considering that human decisions are not always optimal and may sometimes even hinder the progression of coordination. Therefore, these studies still struggle to capture the potential characteristics of human partners. In this work, reinforcement learning (RL) and supervised learning (SL) are integrated to train a reward-conditioned policy. By conditioned on different desired rewards, a reward-conditioned policy simulates both low-reward and high-reward partners. Additionally, a reward-bucketed replay buffer and curriculum learning are applied to enhance reward diversity and boost the training of coordination agents. Experiments demonstrate that the proposed reward-conditioned policy is capable of generating agents with different rewards. Moreover, the zero-shot coordination performance of agents trained with these partners surpasses previous methods in the majority of scenarios within the Overcooked human-AI coordination benchmark. Zhenhua Yang, Peng Yang 0008 |
IJCNN | 2 |
| 2024 | Cue-based two factor authentication
Zhenhua Yang |
Comput. Secur. | 1 |
| 2023 | Fine-Grained Tuple Transfer for Pipelined Query Execution on CPU-GPU Coprocessor
Zhenhua Yang, Qingfeng Pan |
DASFAA (1) | 1 |
| 2023 | Understanding the Continuance Participation of Enterprise Social Media Using the Self-Determined Theory: The Moderating Role of Communication VisibilityabstractEnterprise social media (ESM) technologies are increasingly adopted by organizations to improve organizational processes. Since ESM can afford high communication visibility, employees may be afraid of being supervised. Drawing from self-determination theory, this paper examined how communication visibility moderates the process from two types of motivations (i.e., autonomous motivation and controlled motivation) to employees' subsequent participation provision. Factor-based structural equation modeling was used to analyze a recall-based sample, including 358 employees who are using an ESM tool named DingTalk, within their organizations. The results revealed that autonomous/controlled motivation could indirectly increase/decrease continuance participation intention by positive affect. Also, communication visibility can undermine the positive effects of autonomous motivation and then continuance participation intention. These findings suggest the necessity of stimulating ESM participation with autonomous motivation and balancing the communication visibility afforded by ESM. Zhenhua Yang, Jihui Shi, Shengjun Wang, Daqing Zheng |
J. Glob. Inf. Manag. | 2 |
| 2023 | Neighbor Reweighted Local Centroid for Geometric Feature IdentificationabstractIdentifying geometric features from sampled surfaces is a significant and fundamental task. The existing curvature-based methods that can identify ridge and valley features are generally sensitive to noise. Without requiring high-order differential operators, most statistics-based methods sacrifice certain extents of the feature descriptive powers in exchange for robustness. However, neither of these types of methods can treat the surface boundary features simultaneously. In this paper, we propose a novel neighbor reweighted local centroid (NRLC) computational algorithm to identify geometric features for point cloud models. It constructs a feature descriptor for the considered point via decomposing each of its neighboring vectors into two orthogonal directions. A neighboring vector starts from the considered point and ends with the corresponding neighbor. The decomposed neighboring vectors are then accumulated with different weights to generate the NRLC. With the defined NRLC, we design a probability set for each candidate feature point so that the convex, concave and surface boundary points can be recognized concurrently. In addition, we introduce a pair of feature operators, including assimilation and dissimilation, to further strengthen the identified geometric features. Finally, we test NRLC on a large body of point cloud models derived from different data sources. Several groups of the comparison experiments are conducted, and the results verify the validity and efficiency of our NRLC method. Zhenhua Yang, Shaojun Hu, Zhiyi Zhang 0002, Chunxia Xiao, Xiaohu Guo, Long Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Exploiting Unblocking Checkpoint for Fault-Tolerance in Pregel-Like Systems
Zhenhua Yang |
WISE (1) | 2 |
| 2020 | Censoring-Aware Deep Ordinal Regression for Survival Prediction from Pathological Images
Lichao Xiao, Jin-Gang Yu, Jiarong Ou, Shule Deng, Zhenhua Yang, Yuanqing Li 0001 |
MICCAI (5) | 6 |
| 2017 | A clustering-based algorithm for automatic detection of automobile dashboardabstractThis paper presents an automatic detection system capable of detecting an automobile dashboard with high accuracy. Since the structure of an automobile dashboard is quite different from general instruments, commonly used algorithms for instrument detection can hardly meet the accuracy and robustness. In this paper, a novel approach is presented to detect an automobile dashboard. The contour retrieving algorithm is first performed to extract contour image of dashboard. Progressive Probabilistic Hough Transform (PPHT) is then applied to fit accurate pointer position. A K-means clustering based method is then proposed recognizing tick marks. Since the result of K-means clustering always includes noise items, Z-test is introduced to insure precise clustering result. Lagrange interpolation method is used to describe the precise linear relationship between angle and reading of the instrument. The experimental test shows robustness and precision of this detection algorithm. Zhenhua Yang, Fengyu Guo |
IECON | 2 |
| 2017 | Query Optimal k-Plex Based Community in GraphsabstractCommunity search problem, which is to find good communities given a set of query nodes in a graph, has attracted increasing research interest recently. Though various measurement models have been proposed to define and solve community search problem. Few of them could define a community concisely and have good quality of query results. They either involve additional constraints for modeling communities, such as size and diameter, or suffer from the free rider effect, i.e., include irrelevant subgraphs. In this paper, we propose a new k-plex based community model for community search. We show that our model not only is simple and clear, but also meets with basic requirements of defining a community search problem. We formulate the maximum k-plex community query (MCKPQ) problem, that is, given a set of query nodes Q , searching for optimal k-plex containing Q . We prove that MCKPQ is NP-hard, and it is hard to approximate in any constant factor. We first give exact solutions. Then, we propose an efficient branch-and-bound (B&B) method and design an effective upper bound function and a pruning strategy. Furthermore, we optimize the basic B&B by fast candidate generation. We also give a fast heuristic solution, which produces high-quality results in practice. The effectiveness of our model of community and the efficiency of our methods are verified by elaborate experiments. Yue Wang 0012, Xun Jian 0001, Zhenhua Yang |
Data Sci. Eng. | 3 |
| 2017 | Correction to: Query Optimal k-Plex Based Community in GraphsabstractIn the initial publication, first name and family name of the second author Xun Jian were switched around. The original article has been corrected. Yue Wang 0012, Xun Jian 0001, Zhenhua Yang |
Data Sci. Eng. | 3 |
| 2015 | A unified design of channel coding for LTE uplink control information
Wei Yang 0029, Linyuan Zhang, Zhenhua Yang, Changlong Xu, Young-Il Kim |
Wirel. Networks | 3 |