VLDB 2026 Research / reviewers in the wild / expert
Qiaozhi He
dblp:233/6822
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 42% Trustworthy machine learning · 21% Language models and text generation · 20% | |
| Computer graphics and multimedia
2 papers |
Computer animation and physical simulation · 79% Visual content generation and editing · 21% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
2.0 | 3 | 2026 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
1.9 | 2 | 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing · IEEE Trans. Pattern Anal. Mach. Intell. 2026 StickMotion: Generating 3D Human Motions by Drawing a Stickman · CVPR 2025 |
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
reward model interpretability |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Visual content generation and editing
3d content creation |
1.0 | 1 | 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer animation and physical simulation
motion synthesis |
1.0 | 1 | 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 |
Computer vision › Vision and language › vision-language model
vision-language model alignment |
0.9 | 1 | 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation |
0.9 | 1 | 2025 | StickMotion: Generating 3D Human Motions by Drawing a Stickman · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.3 | 1 | 2025 | StickMotion: Generating 3D Human Motions by Drawing a Stickman · CVPR 2025 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | StickMotion: Generating 3D Human Motions by Drawing a Stickman · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.3 | 1 | 2025 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
multi-condition fusion · 1.7dynamic supervision · 1.7diffusion model · 1.7probing · 1.0multi-objective optimization · 1.0freehand drawing · 1.0optimal transport · 0.9label smoothing · 0.9generative model · 0.9direct preference optimization · 0.9best-of-n sampling · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsabstractPrevious methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models. Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001 |
AAAI | 5 |
| 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing
Tao Wang 0011, Lei Jin 0003, Qiaozhi He, Jiaming Chu, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Shuicheng Yan, Li Wang 0039 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Cross-layer Attention Sharing for Pre-trained Large Language ModelsabstractAbstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa. Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang 0002, Jiali Zeng, Qiaozhi He, Murun Yang, Fandong Meng, Jie Zhou 0016, Tong Xiao 0001 |
Trans. Assoc. Comput. Linguistics | 7 |
| 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataabstractLarge vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 6 |
| 2025 | StickMotion: Generating 3D Human Motions by Drawing a StickmanabstractText-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, which generates desired motions based on traditional text and our proposed stickman conditions for global and local control of these motions, respectively. We address the challenges introduced by the user-friendly stickman from three perspectives: 1) Data generation. We develop an algorithm to generate hand-drawn stickmen automatically across different dataset formats. 2) Multi-condition fusion. We propose a multi-condition module that integrates into the diffusion process and obtains outputs of all possible condition combinations, reducing computational complexity and enhancing StickMotion’s performance compared to conventional approaches with the self-attention module. 3) Dynamic supervision. We empower StickMotion to make minor adjustments to the stickman’s position within the output sequences, generating more natural movements through our proposed dynamic supervision strategy. Through quantitative experiments and user studies, sketching stickmen saves users about 51.5% of their time generating motions consistent with their imagination. Our codes, demos, and relevant data will be released in https:// github.com/InvertedForest/StickMotion. Tao Wang 0011, Qiaozhi He, Jiaming Chu, Ling Qian, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Lei Jin 0003 |
CVPR | 3 |
| 2025 | Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal ModelsabstractPrevious work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes providing a few demonstrations and optimizing image descriptions to be more detailed and logical. However, as demand for more complex and flexible image descriptions grows, enhancing comprehension of input text within the ICL paradigm remains a critical yet underexplored area. In this work, we extend this line of research by constructing parallel multilingual prompts aimed at harnessing the multilingual capabilities of LMMs. More specifically, we translate the input text into several languages and provide the models with both the original text and the translations. Experiments on two LMMs across 3 benchmarks show that our method, PMT2I, achieves superior performance in general, compositional, and fine-grained assessments, especially in human preference alignment Additionally, with its advantage of generating more diverse images, PMT2I significantly outperforms baseline prompts when incorporated with reranking methods. Our code and parallel multilingual data can be found at https://github.com/takagi97/PMT2I. Yongyu Mu, Junxin Wang, Xiaoxuan Zhou, Chenglong Wang 0002, Yingfeng Luo, Qiaozhi He, Tong Xiao 0001, Guocheng Chen |
ICASSP | 7 |
| 2025 | GRAM: A Generative Foundation Reward Model for Reward GeneralizationabstractIn aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
ICML | 5 |
| 2023 | Learning Reliable Neural Networks with Distributed Architecture RepresentationsabstractNeural architecture search (NAS) has shown the strong performance of learning neural models automatically in recent years. But most NAS systems are unreliable due to the architecture gap brought by discrete representations of atomic architectures. In this article, we improve the performance and robustness of NAS via narrowing the gap between architecture representations. More specifically, we apply a general contraction mapping to model neural networks with distributed representations (Neural Architecture Search with Distributed Architecture Representations (ArchDAR)). Moreover, for a better search result, we present a joint learning approach to integrating distributed representations with advanced architecture search methods. We implement our ArchDAR in a differentiable architecture search model and test learned architectures on the language modeling task. On the Penn Treebank data, it outperforms a strong baseline significantly by 1.8 perplexity scores. Also, the search process with distributed representations is more stable, which yields a faster structural convergence when it works with the differentiable architecture search model. Yinqiao Li, Runzhe Cao, Qiaozhi He, Tong Xiao 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |