VLDB 2026 Research / reviewers in the wild / expert
Hongchang Zhang
dblp:36/9348
· DBLP profile ↗
17ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCNet: Differential computing-driven network for infrared and visible image fusion
Longtao Wang, Hongchang Zhang |
Image Vis. Comput. | 3 |
| 2025 | One-Step Multi-Frame Inpainting Framework for Real-Time Lip-Sync Digital Human GenerationabstractIn recent times, audio-driven lip-synching generation for digital humans has attracted considerable attention. However, the prevailing methodologies frequently encounter challenges pertaining to elevated computational complexity and deficient real-time performance. Although the MuseTalk framework has achieved notable progress in inference efficiency through its end-to-end, latent-space-based single-step generation algorithm, it still suffers from noticeable lip jitter and insufficient synchronization between audio and lip movements. To address these limitations, we propose an enhanced multi-frame inpainting framework that integrates Variational Autoencoders (VAE) and a multi-scale U-Net architecture. Specifically, our approach directly synthesizes the occluded lip region by leveraging multi-frame visual references combined with corresponding audio embeddings, thereby effectively improving lip synchronization and maintaining identity consistency. Furthermore, we introduce a landmark-guided multi-frame sampling strategy designed to enhance model attention towards lip dynamics. To facilitate deeper feature extraction and fusion, we propose a hierarchical latent-space feature fusion network (FusionNet), incorporating global and local residual connections and an enhanced Convolutional Block Attention Module. Additionally, a frame interpolation technique is employed during inference to further smooth lip movements and significantly mitigate lip jitter. The model has been trained on a large-scale Chinese dataset and comprehensively evaluated using both Chinese and English datasets. The experimental results demonstrate that the proposed framework achieves high visual accuracy, consistent lip synchronization, and efficient real-time inference, highlighting its strong cross-lingual generalization capability. Yijun Bei, Yunze Qi, Hengrui Lou, Erteng Liu, Hongchang Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2025 | MalRGBDet: Windows Malware Detection Method Based on RGB Image Representation and Heterogeneous Neural NetworkabstractAs the world’s most widely used operating system, Windows has long been a primary target for malware attacks, causing severe economic losses and threats to data security for users and enterprises. Existing detection methods often struggle with low accuracy when dealing with complex malware, suffering from high false-negative and false-positive rates. Additionally, malware detection in Windows faces challenges such as limited datasets, a lack of benign sample contrast and insufficient original feature information. To address these issues, we propose a malware detection method based on RGB image representation and heterogeneous neural network (MalRGBDet). First, we collected malware samples from the GitHub and VirusShare platforms, along with benign software from Windows systems, to build a dataset named MalDet. This data set contains unprocessed malicious and benign samples, providing original feature information and addressing the lack of benign samples in existing data sets. Next, we extracted three key features from the malware samples: code sections, data sections and API call sequences. These features closely relate to the behavior of malware and accurately describe its operations. We then transformed these features into uniformly sized RGB images, which helped reveal hidden patterns. Finally, we employ a heterogeneous neural network that integrates ResNet and AlexNet for classification. ResNet, with its deep architecture and residual learning mechanism, significantly enhances the model’s representation capability and classification performance, thereby improving detection accuracy. Meanwhile, AlexNet’s Dropout regularization strategy effectively boosts the model’s generalization ability. In our data set of 1952 Windows software samples, MalRGBDet achieved more than 95% in accuracy, precision, recall and F1-score, improving these metrics by up to 4% compared to the latest methods. Furthermore, false-negative and false-positive rates were kept below 5%. Rong Ren, Hongchang Zhang, Bing Zhang 0011, Haitao He, Guoyan Huang, Qian Wang 0009 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2025 | DFF-Net: Deep Feature Fusion Network for low-light image enhancement
Hongchang Zhang, Longtao Wang, Qizhan Zou, Juan Zeng |
Image Vis. Comput. | 1 |
| 2024 | Approach to Detect Windows Malware Based on Malicious Tendency Image and ResNet AlgorithmabstractTimely detection of self-replicating malware in the high market share Windows operating system can effectively prevent personal or corporate financial losses. The form and characteristics of malware are constantly evolving, leading to a concept drift issue that gradually decreases the effectiveness of traditional detection methods. Therefore, we propose WinMDet, a Windows malware detection method based on malicious tendency image and ResNet algorithm. First, to tackle the complexity and difficulty in accurately characterizing malware features, WinMDet retains detailed malware features and encodes them into malicious tendency images to better describe malware across different periods. Secondly, WinMDet utilizes previously generated malicious tendency images to train the initial detection model. Then, to alleviate the issue of malware concept drift, WinMDet employs Local Maximum Mean Discrepancy (LMMD) as the criterion for model transfer, enhancing the initial detection model’s ability to distinguish between malware and benign software. We conducted a comprehensive evaluation of WinMDet using common metrics such as accuracy, precision and recall. The results indicate that WinMDet performs remarkably well in terms of accuracy, exceeding 82%. Additionally, significant improvements were observed in precision and recall, surpassing 82.42% and 82.06%, respectively. After employing our LMMD-based transfer method, the initial detection model improved the detection accuracy of malware in 2021 and 2022 by approximately 4.22% to 8.06%. The false negative rate decreased by at most 4.34%, and the false positive rate decreased by at most 4.61%. Bing Zhang 0011, Hongchang Zhang, Rong Ren, Qian Wang 0009 |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2023 | DARL: Distance-Aware Uncertainty Estimation for Offline Reinforcement LearningabstractTo facilitate offline reinforcement learning, uncertainty estimation is commonly used to detect out-of-distribution data. By inspecting, we show that current explicit uncertainty estimators such as Monte Carlo Dropout and model ensemble are not competent to provide trustworthy uncertainty estimation in offline reinforcement learning. Accordingly, we propose a non-parametric distance-aware uncertainty estimator which is sensitive to the change in the input space for offline reinforcement learning. Based on our new estimator, adaptive truncated quantile critics are proposed to underestimate the out-of-distribution samples. We show that the proposed distance-aware uncertainty estimator is able to offer better uncertainty estimation compared to previous methods. Experimental results demonstrate that our proposed DARL method is competitive to the state-of-the-art methods in offline evaluation tasks. Hongchang Zhang, Jianzhun Shao, Shuncheng He, Yuhang Jiang 0001, Xiangyang Ji |
AAAI | 1 |
| 2023 | In-sample Actor Critic for Offline Reinforcement Learning
Hongchang Zhang, Yixiu Mao, Shuncheng He, Yi Xu 0008, Xiangyang Ji |
ICLR | 1 |
| 2023 | Supported Trust Region Optimization for Offline Reinforcement LearningabstractOffline reinforcement learning suffers from the out-of-distribution issue and extrapolation error. Most policy constraint methods regularize the density of the trained policy towards the behavior policy, which is too restrictive in most cases. We propose Supported Trust Region optimization (STR) which performs trust region policy optimization with the policy constrained within the support of the behavior policy, enjoying the less restrictive support constraint. We show that, when assuming no approximation and sampling error, STR guarantees strict policy improvement until convergence to the optimal support-constrained policy in the dataset. Further with both errors incorporated, STR still guarantees safe policy improvement for each step. Empirical results validate the theory of STR and demonstrate its state-of-the-art performance on MuJoCo locomotion domains and much more challenging AntMaze domains. Yixiu Mao, Hongchang Zhang, Yi Xu 0008, Xiangyang Ji |
ICML | 2 |
| 2023 | Complementary Attention for Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning, centralized training with decentralized execution (CTDE) shows great promise for a trade-off between independent Q-learning and joint action learning. However, vanilla CTDE methods assumed a fixed number of agents could hardly adapt to real-world scenarios where dynamic team compositions typically suffer from dramatically variant partial observability. Specifically, agents with extensive sight ranges are prone to be affected by trivial environmental substrates, dubbed the "distracted attention" issue; ones with limited observation can hardly sense their teammates, degrading the cooperation quality. In this paper, we propose Complementary Attention for Multi-Agent reinforcement learning (CAMA), which applies a divide-and-conquer strategy on input entities accompanied with the complementary attention of enhancement and replenishment. Concretely, to tackle the distracted attention issue, highly contributed entities' attention is enhanced by the execution-related representation extracted via action prediction with an inverse model. For better out-of-sight-range cooperation, the lowly contributed ones are compressed to brief messages with a conditional mutual information estimator. Our CAMA facilitates stable and sustainable teamwork, which is justified by the impressive results reported on the challenging StarCraftII, MPE, and Traffic Junction benchmarks. Jianzhun Shao, Hongchang Zhang, Yun Qu 0002, Chang Liu 0030, Shuncheng He, Yuhang Jiang 0001, Xiangyang Ji |
ICML | 2 |
| 2023 | Supported Value Regularization for Offline Reinforcement LearningabstractOffline reinforcement learning suffers from the extrapolation error and value overestimation caused by out-of-distribution (OOD) actions. To mitigate this issue, value regularization approaches aim to penalize the learned value functions to assign lower values to OOD actions. However, existing value regularization methods lack a proper distinction between the regularization effects on in-distribution (ID) and OOD actions, and fail to guarantee optimal convergence results of the policy. To this end, we propose Supported Value Regularization (SVR), which penalizes the Q-values for all OOD actions while maintaining standard Bellman updates for ID ones. Specifically, we utilize the bias of importance sampling to compute the summation of Q-values over the entire OOD region, which serves as the penalty for policy evaluation. This design automatically separates the regularization for ID and OOD actions without manually distinguishing between them. In tabular MDP, we show that the policy evaluation operator of SVR is a contraction, whose fixed point outputs unbiased Q-values for ID actions and underestimated Q-values for OOD actions. Furthermore, the policy iteration with SVR guarantees strict policy improvement until convergence to the optimal support-constrained policy in the dataset. Empirically, we validate the theoretical properties of SVR in a tabular maze environment and demonstrate its state-of-the-art performance on a range of continuous control tasks in the D4RL benchmark. Yixiu Mao, Hongchang Zhang, Yi Xu 0008, Xiangyang Ji |
NeurIPS | 2 |
| 2023 | Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningabstractOffline multi-agent reinforcement learning is challenging due to the coupling effect of both distribution shift issue common in offline setting and the high dimension issue common in multi-agent setting, making the action out-of-distribution (OOD) and value overestimation phenomenon excessively severe. To mitigate this problem, we propose a novel multi-agent offline RL algorithm, named CounterFactual Conservative Q-Learning (CFCQL) to conduct conservative value estimation. Rather than regarding all the agents as a high dimensional single one and directly applying single agent conservative methods to it, CFCQL calculates conservative regularization for each agent separately in a counterfactual way and then linearly combines them to realize an overall conservative value estimation. We prove that it still enjoys the underestimation property and the performance guarantee as those single agent conservative methods do, but the induced regularization and safe policy improvement bound are independent of the agent number, which is therefore theoretically superior to the direct treatment referred to above, especially when the agent number is large. We further conduct experiments on four environments including both discrete and continuous action settings on both existing and our man-made datasets, demonstrating that CFCQL outperforms existing methods on most datasets and even with a remarkable margin on some of them. Jianzhun Shao, Yun Qu 0002, Hongchang Zhang, Xiangyang Ji |
NeurIPS | 4 |
| 2022 | Wasserstein Unsupervised Reinforcement LearningabstractUnsupervised reinforcement learning aims to train agents to learn a handful of policies or skills in environments without external reward. These pre-trained policies can accelerate learning when endowed with external reward, and can also be used as primitive options in hierarchical reinforcement learning. Conventional approaches of unsupervised skill discovery feed a latent variable to the agent and shed its empowerment on agent’s behavior by mutual information (MI) maximization. However, the policies learned by MI-based methods cannot sufficiently explore the state space, despite they can be successfully identified from each other. Therefore we propose a new framework Wasserstein unsupervised reinforcement learning (WURL) where we directly maximize the distance of state distributions induced by different policies. Additionally, we overcome difficulties in simultaneously training N(N>2) policies, and amortizing the overall reward to each step. Experiments show policies learned by our approach outperform MI-based methods on the metric of Wasserstein distance while keeping high discriminability. Furthermore, the agents trained by WURL can sufficiently explore the state space in mazes and MuJoCo tasks and the pre-trained policies can be applied to downstream tasks by hierarchical learning. Shuncheng He, Yuhang Jiang 0001, Hongchang Zhang, Jianzhun Shao, Xiangyang Ji |
AAAI | 3 |
| 2022 | State Deviation Correction for Offline Reinforcement LearningabstractOffline reinforcement learning aims to maximize the expected cumulative rewards with a fixed collection of data. The basic principle of current offline reinforcement learning methods is to restrict the policy to the offline dataset action space. However, they ignore the case where the dataset's trajectories fail to cover the state space completely. Especially, when the dataset's size is limited, it is likely that the agent would encounter unseen states during test time. Prior policy-constrained methods are incapable of correcting the state deviation, and may lead the agent to its unexpected regions further. In this paper, we propose the state deviation correction (SDC) method to constrain the policy's induced state distribution by penalizing the out-of-distribution states which might appear during the test period. We first perturb the states sampled from the logged dataset, then simulate noisy next states on the basis of a dynamics model and the policy. We then train the policy to minimize the distances between the noisy next states and the offline dataset. In this manner, we allow the trained policy to guide the agent to its familiar regions. Experimental results demonstrate that our proposed method is competitive with the state-of-the-art methods in a GridWorld setup, offline Mujoco control suite, and a modified offline Mujoco dataset with a finite number of valuable samples. Hongchang Zhang, Jianzhun Shao, Yuhang Jiang 0001, Shuncheng He, Guanwen Zhang, Xiangyang Ji |
AAAI | 1 |
| 2022 | SPD: Synergy Pattern Diversifying Oriented Unsupervised Multi-agent Reinforcement LearningabstractReinforcement learning typically relies heavily on a well-designed reward signal, which gets more challenging in cooperative multi-agent reinforcement learning. Alternatively, unsupervised reinforcement learning (URL) has delivered on its promise in the recent past to learn useful skills and explore the environment without external supervised signals. These approaches mainly aimed for the single agent to reach distinguishable states, insufficient for multi-agent systems due to that each agent interacts with not only the environment, but also the other agents. We propose Synergy Pattern Diversifying Oriented Unsupervised Multi-agent Reinforcement Learning (SPD) to learn generic coordination policies for agents with no extrinsic reward. Specifically, we devise the Synergy Pattern Graph (SPG), a graph depicting the relationships of agents at each time step. Furthermore, we propose an episode-wise divergence measurement to approximate the discrepancy of synergy patterns. To overcome the challenge of sparse return, we decompose the discrepancy of synergy patterns to per-time-step pseudo-reward. Empirically, we show the capacity of SPD to acquire meaningful coordination policies, such as maintaining specific formations in Multi-Agent Particle Environment and pass-and-shoot in Google Research Football. Furthermore, we demonstrate that the same instructive pretrained policy's parameters can serve as a good initialization for a series of downstream tasks' policies, achieving higher data efficiency and outperforming state-of-the-art approaches in Google Research Football. Yuhang Jiang 0001, Jianzhun Shao, Shuncheng He, Hongchang Zhang, Xiangyang Ji |
NeurIPS | 4 |
| 2022 | Self-Organized Group for Cooperative Multi-agent Reinforcement LearningabstractCentralized training with decentralized execution (CTDE) has achieved great success in cooperative multi-agent reinforcement learning (MARL) in practical applications. However, CTDE-based methods typically suffer from poor zero-shot generalization ability with dynamic team composition and varying partial observability. To tackle these issues, we propose a spontaneously grouping mechanism, termed Self-Organized Group (SOG), which is featured with conductor election (CE) and message summary (MS). In CE, a certain number of conductors are elected every $T$ time-steps to temporally construct groups, each with conductor-follower consensus where the followers are constrained to only communicate with their conductor. In MS, each conductor summarize and distribute the received messages to all affiliate group members to hold a unified scheduling. SOG provides zero-shot generalization ability to the dynamic number of agents and the varying partial observability. Sufficient experiments on mainstream multi-agent benchmarks exhibit superiority of SOG. Jianzhun Shao, Zhiqiang Lou, Hongchang Zhang, Yuhang Jiang 0001, Shuncheng He, Xiangyang Ji |
NeurIPS | 3 |
| 2022 | Optimizing the dynamic treatment regime of in-hospital warfarin anticoagulation in patients after surgical valve replacement using reinforcement learningabstractOBJECTIVE: Warfarin anticoagulation management requires sequential decision-making to adjust dosages based on patients' evolving states continuously. We aimed to leverage reinforcement learning (RL) to optimize the dynamic in-hospital warfarin dosing in patients after surgical valve replacement (SVR). MATERIALS AND METHODS: 10 408 SVR cases with warfarin dosage-response data were retrospectively collected to develop and test an RL algorithm that can continuously recommend daily warfarin doses based on patients' evolving multidimensional states. The RL algorithm was compared with clinicians' actual practice and other machine learning and clinical decision rule-based algorithms. The primary outcome was the ratio of patients without in-hospital INRs >3.0 and the INR at discharge within the target range (1.8-2.5) (excellent responders). The secondary outcomes were the safety responder ratio (no INRs >3.0) and the target responder ratio (the discharge INR within 1.8-2.5). RESULTS: In the test set (n = 1260), the excellent responder ratio under clinicians' guidance was significantly lower than the RL algorithm: 41.6% versus 80.8% (relative risk [RR], 0.51; 95% confidence interval [CI], 0.48-0.55), also the safety responder ratio: 83.1% versus 99.5% (RR, 0.83; 95% CI, 0.81-0.86), and the target responder ratio: 49.7% versus 81.1% (RR, 0.61; 95% CI, 0.58-0.65). The RL algorithms performed significantly better than all the other algorithms. Compared with clinicians' actual practice, the RL-optimized INR trajectory reached and maintained within the target range significantly faster and longer. DISCUSSION: RL could offer interactive, practical clinical decision support for sequential decision-making tasks and is potentially adaptable for varied clinical scenarios. Prospective validation is needed. CONCLUSION: An RL algorithm significantly optimized the post-operation warfarin anticoagulation quality compared with clinicians' actual practice, suggesting its potential for challenging sequential decision-making tasks. Juntong Zeng, Jianzhun Shao, Hongchang Zhang, Xiaoting Su, Xiaocong Lian, Xiangyang Ji |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | An approach of class integration test order determination based on test levelsabstractIn recent years, many approaches have been developed to determine the order of tested classes in interclass integration test. However, existing approaches are inaccurate, as they ignore the influence of abstract classes and polymorphism. In this paper, we propose a test-level-based approach to deal with class-integration-test order, in which both abstract classes and polymorphism are taken into account. First, based on interclass dependence analysis, we develop an edge-removing algorithm to eliminate cycles caused by static and dynamic dependencies, taking abstract classes and polymorphism into account. Then, after eliminating cycles, we propose a class-integration-test order algorithm based on test levels, including static and dynamic test levels. In this algorithm, we take into account the fact of some test levels infeasible caused by the characteristic of abstract classes that they cannot be instantiated and offer corresponding adjustment strategy. Finally, we design and implement a test level order generator. The experimental results show that the proposed strategy needs less test stubs than the most typically graph-based approaches. Copyright © 2014 John Wiley & Sons, Ltd. Shujuan Jiang, Guan Yuan, Xiaolin Ju, Hongchang Zhang |
Softw. Pract. Exp. | 5 |