VLDB 2026 Research / reviewers in the wild / expert
Haorui Wang
dblp:272/9103
· DBLP profile ↗
17ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RLMR: Reinforcement Learning with Mixed Rewards for Creative WritingabstractLarge language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constraint following (e.g., format requirements and word limits). Existing reinforcement learning methods struggle to balance these two aspects: single reward strategies fail to improve both abilities simultaneously, while fixed-weight mixed-reward methods lack the ability to adapt to different writing scenarios. To address this problem, we propose Reinforcement Learning with Mixed Rewards (RLMR), utilizing a dynamically mixed reward system from a writing reward model evaluating subjective writing quality and a constraint verification model assessing objective constraint following. The constraint following reward weight is adjusted dynamically according to the writing quality within sampled groups, ensuring that samples violating constraints get negative advantage in GRPO and thus penalized during training, which is the key innovation of this proposed method. We conduct automated and manual evaluations across diverse model families from 8B to 72B parameters. Additionally, we construct a real-world writing benchmark named WriteEval for comprehensive evaluation. Results illustrate that our method achieves consistent improvements in both instruction following (IFEval from 83.36% to 86.65%) and writing quality (72.75% win rate in manual expert pairwise evaluations on WriteEval). To the best of our knowledge, RLMR is the first work to combine subjective preferences with objective verification in online RL training, providing an effective solution for multi-dimensional creative writing optimization. Jianxing Liao, Yusong Zhang, Haorui Wang, Bosi Wen, Ziying Wang, Runzhi Shi |
AAAI | 5 |
| 2026 | HyMed: An Event-Driven Multi-agent Framework for Smart Hospitals
Haorui Wang, Jingjing Pan, Chuanlei Zhang |
ICIC (29) | 1 |
| 2025 | Diffusion Models as Constrained Samplers for Optimization with Unknown ConstraintsabstractAddressing real-world optimization problems becomes particularly challenging when analytic objective functions or constraints are unavailable. While numerous studies have addressed the issue of unknown objectives, limited research has focused on scenarios where feasibility constraints are not given explicitly. Overlooking these constraints can lead to spurious solutions that are unrealistic in practice. To deal with such unknown constraints, we propose to perform optimization within the data manifold using diffusion models. To constrain the optimization process to the data manifold, we reformulate the original optimization problem as a sampling problem from the product of the Boltzmann distribution defined by the objective function and the data distribution learned by the diffusion model. Depending on the differentiability of the objective function, we propose two different sampling methods. For differentiable objectives, we propose a two-stage framework that begins with a guided diffusion process for warm-up, followed by a Langevin dynamics stage for further correction. For non-differentiable objectives, we propose an iterative importance sampling strategy using the diffusion model as the proposal distribution. Comprehensive experiments on a synthetic dataset, six real-world black-box optimization datasets, and a multi-objective molecule optimization dataset show that our method achieves better or comparable performance with previous state-of-the-art baselines. Yuanqi Du, Wenhao Mu, Kirill Neklyudov, Valentin De Bortoli, Dongxia Wu, Haorui Wang, Aaron M. Ferber, Yi-An Ma, Carla P. Gomes, Chao Zhang 0014 |
AISTATS | 7 |
| 2025 | Synergizing Multimodal Temporal Knowledge Graphs and Large Language Models for Social Relation RecognitionabstractRecent years have witnessed remarkable advances in Large Language Models (LLMs).However, in the task of social relation recognition, Large Language Models (LLMs) encounter significant challenges due to their reliance on sequential training data, which inherently restricts their capacity to effectively model complex graph-structured relationships.To address this limitation, we propose a novel low-coupling method synergizing multimodal temporal Knowledge Graphs and Large Language Models (mtKG-LLM) for social relation reasoning.Specifically, we extract multimodal information from the videos and model the social networks as spatial Knowledge Graphs (KGs) for each scene.Temporal KGs are constructed based on spatial KGs and updated along the timeline for long-term reasoning.Subsequently, we retrieve multi-scale information from the graph-structured knowledge for LLMs to recognize the underlying social relation.Extensive experiments demonstrate that our method has achieved state-ofthe-art performance in social relation recognition.Furthermore, our framework exhibits effectiveness in bridging the gap between KGs and LLMs.We release our code at https: //github.com/HarryWgCN/mtKG-LLM. Haorui Wang |
EMNLP | 1 |
| 2025 | Efficient Evolutionary Search Over Chemical Space with Large Language ModelsabstractMolecular discovery, when formulated as an optimization problem, presents significant computational challenges because optimization objectives can be non-differentiable. Evolutionary Algorithms (EAs), often used to optimize black-box objectives in molecular discovery, traverse chemical space by performing random mutations and crossovers, leading to a large number of expensive objective evaluations. In this work, we ameliorate this shortcoming by incorporating chemistry-aware Large Language Models (LLMs) into EAs. Namely, we redesign crossover and mutation operations in EAs using LLMs trained on large corpora of chemical information. We perform extensive empirical studies on both commercial and open-source models on multiple tasks involving property optimization, molecular rediscovery, and structure-based drug design, demonstrating that the joint usage of LLMs with EAs yields superior performance over all baseline models across single- and multi-objective settings. We demonstrate that our algorithm improves both the quality of the final solution and convergence speed, thereby reducing the number of required objective evaluations. Haorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao 0001, Felix Strieth-Kalthoff, Chenru Duan, Yuchen Zhuang, Yue Yu 0001, Yanqiao Zhu 0001, Yuanqi Du, Alán Aspuru-Guzik, Kirill Neklyudov, Chao Zhang 0014 |
ICLR | 1 |
| 2025 | LLM-Augmented Chemical Synthesis and Design Decision ProgramsabstractRetrosynthesis, the process of breaking down a target molecule into simpler precursors through a series of valid reactions, stands at the core of organic chemistry and drug development. Although recent machine learning (ML) research has advanced single-step retrosynthetic modeling and subsequent route searches, these solutions remain restricted by the extensive combinatorial space of possible pathways. Concurrently, large language models (LLMs) have exhibited remarkable chemical knowledge, hinting at their potential to tackle complex decision-making tasks in chemistry. In this work, we explore whether LLMs can successfully navigate the highly constrained, multi-step retrosynthesis planning problem. We introduce an efficient scheme for encoding reaction pathways and present a new route-level search strategy, moving beyond the conventional step-by-step reactant prediction. Through comprehensive evaluations, we show that our LLM-augmented approach excels at retrosynthesis planning and extends naturally to the broader challenge of synthesizable molecular design. Haorui Wang, Jeff Guo, Rampi Ramprasad, Philippe Schwaller, Yuanqi Du, Chao Zhang 0014 |
ICML | 1 |
| 2025 | Improving Graph Contrastive Learning with LLMs: A New Hardness-Aware Negative Sampling StrategyabstractGraph contrastive learning plays a crucial role in many fields, such as social network analysis and recommendation, and has become a hot research area in recent years. The key to improving the quality of contrastive learning is to obtain high-quality sample pairs that reflect the structural features of the data more effectively. Existing methods typically sample based on node similarity and can only select negative samples of a fixed difficulty level, leading to issues like false positives and false negatives. This study proposes a novel paradigm called Hardness-aware Negative Sampling Graph Contrastive Learning with LLMs. We outline the criteria that the adaptive hardness sampling method should follow and provide concrete instances. By adaptively selecting negative samples of appropriate hardness during the training process, this method effectively mitigates false negative and false positive problems. Additionally, we leverage large language models to integrate data with supplementary textual attributes. Extensive experiments and analyses demonstrate the superiority and effectiveness of our proposed method. Shuai Zhong, Haorui Wang, Xinming Chen, Yuanxing Xu, Bin Wu 0001 |
IJCNN | 3 |
| 2025 | Cause and Effect: Video Social Relationship Recognition from Causal PerspectiveabstractVideo social relation recognition is a fundamental task in video understanding, which is dedicated to the construction of multi-modal knowledge graphs. Previous work mainly focuses on multi-modal fusion and the construction of special character graphs. However, they often treat the global frame sequence equally, ignoring the influence of key frame sequence on relation recognition. Specifically, the key frame sequence that significantly reflect character relationships in a video tends to be sparse and short. At the same time, the key frames have not only temporal but also strong causal relationship. Therefore, we propose a novel Video Local Causal Frame (VLCF) model to explore the causal relationship between frames. Inspired by Granger causality theory, we estimate inter-frame causal relationships by comparing the predicted result frames with and without masking the premise frame. We then construct global connections between video frames. Multiple local causal frame sequences and global frame sequences are extracted to capture the key information and global information in the video. Extensive experiments conducted on the ViSR dataset and the MovieGraphs dataset demonstrate that the proposed model achieves state-of-the-art performance. Yangfu Zhu, Haorui Wang, Guangyao Su, Bin Wu 0001 |
ACM Multimedia | 5 |
| 2024 | Two Birds with One Stone: Enhancing Uncertainty Quantification and Interpretability with Graph Functional Neural ProcessabstractGraph neural networks (GNNs) are powerful tools on graph data. However, their predictions are mis-calibrated and lack interpretability, limiting their adoption in critical applications. To address this issue, we propose a new uncertainty-aware and interpretable graph classification model that combines graph functional neural process and graph generative model. The core of our method is to assume a set of latent rationales which can be mapped to a probabilistic embedding space; the predictive distribution of the classifier is conditioned on such rationale embeddings by learning a stochastic correlation matrix. The graph generator serves to decode the graph structure of the rationales from the embedding space for model interpretability. For efficient model training, we adopt an alternating optimization procedure which mimics the well known Expectation-Maximization (EM) algorithm. The proposed method is general and can be applied to any existing GNN architecture. Extensive experiments on five graph classification datasets demonstrate that our framework outperforms state-of-the-art methods in both uncertainty quantification and GNN interpretability. We also conduct case studies to show that the decoded rationale structure can provide meaningful explanations. Yuchen Zhuang, Haorui Wang, Wenhao Mu, Chao Zhang 0014 |
AISTATS | 4 |
| 2024 | A Dynamic pre-trained Model for Chinese Classical Poetry
Xuanning Liu, Haorui Wang, Bin Wu 0001 |
DASFAA (2) | 3 |
| 2024 | Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case StudyabstractYinghao Li, Haorui Wang, Chao Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Haorui Wang, Chao Zhang 0014 |
NAACL-HLT | 2 |
| 2024 | Aligning Large Language Models with Representation Editing: A Control PerspectiveabstractAligning large language models (LLMs) with human objectives is crucial for real-world applications. However, fine-tuning LLMs for alignment often suffers from unstable training and requires substantial computing resources. Test-time alignment techniques, such as prompting and guided decoding, do not modify the underlying model, and their performance remains dependent on the original model's capabilities. To address these challenges, we propose aligning LLMs through representation editing. The core of our method is to view a pre-trained autoregressive LLM as a discrete-time stochastic dynamical system. To achieve alignment for specific objectives, we introduce external control signals into the state space of this language dynamical system. We train a value function directly on the hidden states according to the Bellman equation, enabling gradient-based optimization to obtain the optimal control signals at test time. Our experiments demonstrate that our method outperforms existing test-time alignment techniques while requiring significantly fewer resources compared to fine-tuning methods. Our code is available at [https://github.com/Lingkai-Kong/RE-Control](https://github.com/Lingkai-Kong/RE-Control). Haorui Wang, Wenhao Mu, Yuanqi Du, Yuchen Zhuang, Rongzhi Zhang, Kai Wang 0036, Chao Zhang 0014 |
NeurIPS | 2 |
| 2024 | Camellia oleifera trunks detection and identification based on improved YOLOv7abstractSummary Camellia oleifera typically thrives in unstructured environments, making the identification of its trunks crucial for advancing agricultural robots towards modernization and sustainability. Traditional target detection algorithms, however, fall short in accurately identifying Camellia oleifera trunks, especially in scenarios characterized by small targets and poor lighting. This article introduces an enhanced trunk detection algorithm for Camellia oleifera based on an improved YOLOv7 model. This model incorporates dynamic snake convolution instead of standard convolutions to bolster its feature extraction capabilities. It integrates more contextual information, thus enhancing the model's generalization ability across various scenes. Additionally, coordinate attention is introduced to refine the model's spatial feature representation, amplifying the network's focus on essential target region features, which in turn boosts detection accuracy and robustness. This feature selectively strengthens response levels across different channels, prioritizing key attributes for classification and localization. Moreover, the original coordinate loss function of YOLOv7 is replaced with EIoU loss, further enhancing the model's robustness and convergence speed. Experimental results demonstrate a recall rate of 96%, a mean average precision (mAP) of 87.9%, an F1 score of 0.87, and a detection speed of 18 milliseconds per frame. When compared with other models like Faster‐RCNN, YOLOv3, ScaledYOLOv4, YOLOv5, and the original YOLOv7, our improved model shows mAP increases of 8.1%, 7.0%, 7.5%, and 6.6% respectively. Occupying only 70.8 MB, our model requires 9.8 MB less memory than the original YOLOv7. This model not only achieves high accuracy and detection efficiency but is also easily deployable on mobile devices, providing a robust foundation for future intelligent harvesting technologies. Haorui Wang, Yang Liu 0449, Yuanyin Luo |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Shifted GCN-GAT and Cumulative-Transformer based Social Relation Recognition for Long VideosabstractSocial Relation Recognition is an important part of Video Understanding, providing insights into the information that videos convey. Most previous works mainly focused on graph generation for characters, instead of edges which are more suitable for relation modelling. Furthermore, previous methods tend to recognize social relations for single frames or short video clips within their receptive fields, neglecting the importance of continuous reasoning throughout the entire video. To tackle these challenges, we propose a novel Shifted GCN-GAT and Cumulative-Transformer framework, named SGCAT-CT. The overall architecture consists of an SGCAT module for shifted graph operations on novel relation graphs and a CT module for temporal processing with memory. SGCAT-CT conducts continuous recognition of social relations and memorizes information from as early as the beginning of a long video. Experiments conducted on several video datasets demonstrate encouraging performance on long videos. Our code will be released at https://github.com/HarryWgCN/SGCAT-CT. Haorui Wang, Yibo Hu 0005, Yangfu Zhu, Jinsheng Qi, Bin Wu 0001 |
ACM Multimedia | 1 |
| 2023 | Social Relation Graph Generation on Untrimmed Video
Yibo Hu 0005, Chenghao Yan, Chenyu Cao, Haorui Wang, Bin Wu 0001 |
MMM (2) | 4 |
| 2022 | Equivariant and Stable Positional Encoding for More Powerful Graph Neural Networks
Haorui Wang, Haoteng Yin, Muhan Zhang, Pan Li 0005 |
ICLR | 1 |
| 2021 | Predicting Drug-miRNA Resistance with Layer Attention Graph Convolution Network and Multi Channel Feature ExtractionabstractMicroRNA (miRNA) has became an increasingly important class of attractive drug targets in recent studies. However, there are only few computational tools aiming to predict drugmi-RNA resistance associations. Hence, it is of great significance to develop effective and high accuracy methods for predicting drugmi-RNA resistance associations. In this work, we propose a novel method abbreviated as “DMR-GCN”, which enhances drugmi-RNA resistance interaction prediction by using layer attention graph convolution network and multi channel feature extraction. Specifically, DMR-GCN first constructs a heterogeneous network based on known drug-miRNA interactions, drug-drug similarities and miRNA-miRNA similarities. Secondly, layer attention graph convolution network is used to extract drug representations from the drug molecular graph and the heterogeneous network. We concatenate the extracted representations from molecular graph and heterogeneous network as the drug embedding vectors. Similarly, miRNA representations extracted from the heterogeneous network and the miRNA expression features embedded by MLP are concatenated as the miRNA embedding vectors. Further, we utilize Multi-Layer Perceptron (MLP), Generalized Tensor Factorization (GTF) and Compressed Tensor Network (CTN) to extract node-pair representations from different aspects. Finally, the predictive scores for unobserved drug-miRNA resistance associations are given by a fully connection layer with the integrated embeddings. In the evaluation experiments, DMR-GCN achieves an area under the precision-recall curve of 0.2920 and an area under the receiver-operating characteristic curve of 0.9433, which are better than the state-of-the-art prediction methods. The experimental results demonstrate that layer attention mechanism produces satisfying results for learning representations from graph, and integrating multi channel feature extraction can make further improvements. In conclusion, DMR-GCN is a promising method for predicting drug-miRNA resistance associations. Haorui Wang, Shahanavaj Khan, Shichao Liu 0002, Fang Zheng 0010, Wen Zhang 0008 |
BIBM | 1 |