EDBT 2026 Demo / reviewers in the wild / expert
Shiyu Huang 0001
dblp:198/1495
· DBLP profile ↗
20ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-0500-0141ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniDPO: A Preference Optimization Framework to Address Omni-Modal HallucinationabstractRecently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text modality tend to dominate, leading OLLMs to rely more heavily on textual cues while neglecting visual and audio information. In addition, fully multimodal scenarios introduce new challenges. Most existing models align visual or auditory modalities with text independently during training, while ignoring the intrinsic correlations between video and its corresponding audio. This oversight results in hallucinations when reasoning requires interpreting hidden audio cues embedded in video content. To address these challenges, we propose OmniDPO, a preference-alignment framework designed to mitigate hallucinations in OLLMs. Specifically, OmniDPO incorporates two strategies: (1) constructing text-preference sample pairs to enhance the model’s understanding of audio-video interactions; and (2) constructing multimodal-preference sample pairs to strengthen the model’s attention to visual and auditory information. By tackling both challenges, OmniDPO effectively improves multimodal grounding and reduces hallucination. Experiments conducted on two OLLMs demonstrate that OmniDPO not only effectively mitigates multimodal hallucinations but also significantly enhances the models' reasoning capabilities across modalities. Junzhe Chen 0001, Tianshu Zhang 0002, Shiyu Huang 0001, Yuwei Niu, Rongzhou Zhang, Guanyu Zhou, Lijie Wen 0001 |
AAAI | 3 |
| 2026 | VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video GenerationabstractVisual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore. Jiazheng Xu, Yuanming Yang, Wenbo Duan, Shen Yang 0001, Qunlin Jin, Shurun Li, Jiayan Teng, Zhuoyi Yang, Wendi Zheng, Xiao Liu 0036, Ming Ding 0004, Shiyu Huang 0001, Xiaotao Gu, Minlie Huang, Jie Tang 0001, Yuxiao Dong |
AAAI | 18 |
| 2026 | AutoSAT: Automatically Optimize SAT Solvers via Large Language ModelsabstractBackground: Conflict-Driven Clause Learning (CDCL) is a dominant framework for solving the Satisfiability problem (SAT). Modern CDCL solvers rely heavily on various heuristics, which significantly influence their performance. Established solvers such as MiniSat and Kissat typically incorporate multiple heuristics and therefore require substantial manual effort and domain expertise for fine-tuning in practice. Objectives: The emergence of Large Language Models (LLMs) offers a promising opportunity to automate SAT solver optimization. However, generating a complete CDCL solver from scratch using LLMs is impractical due to the complexity and large context volume of modern SAT solvers. To address this challenge, we propose AutoSAT, a framework that automatically optimizes heuristics within a CDCL solver, EasySAT. Our goal is to leverage LLMs to discover effective heuristic designs beyond conventional parameter tuning. Methods: Unlike traditional automated algorithm design approaches that mainly focus on hyperparameter tuning and operator selection, AutoSAT can generate new efficient heuristics for CDCL solvers. In this first attempt to leverage LLMs for SAT solver optimization, we integrate several search strategies, including the greedy hill climber and the (1 + 1) Evolutionary Algorithm, to guide LLMs in searching for better heuristics. Results: Experimental results demonstrate that LLMs can consistently enhance the performance of CDCL solvers. Empirically, AutoSAT outperforms MiniSat and its parameter-tuning variants on 12 of the 19 test datasets, and even surpasses the state-of-the-art hybrid solver Kissat and its parameter-tuning variants on 4 datasets. Conclusions: These results suggest that LLMs can serve as effective agents for automatically improving SAT solver heuristics. AutoSAT provides an initial step toward automated heuristic discovery for CDCL solvers, showing the potential of LLM-based methods to complement and reduce the manual expertise traditionally required in SAT solver design. Furong Ye, Xianyin Zhang, Shiyu Huang 0001, Binzhen Zhang, Shaowei Cai 0001 |
J. Artif. Intell. Res. | 4 |
| 2025 | Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?abstractLeyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Leyi Pan, Aiwei Liu, Shiyu Huang 0001, Yijian Lu, Xuming Hu, Lijie Wen 0001, Irwin King, Philip S. Yu |
ACL (1) | 3 |
| 2025 | ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsabstractDespite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical application in real-world scenarios that demand high levels of precision. Existing methods typically either fine-tune the LVLMs using additional data, which incurs extra costs in manual annotation and computational resources or perform comparisons at the decoding stage, which may eliminate useful language priors for reasoning while introducing inference time overhead. Therefore, we propose ICT, a lightweight, training-free method that calculates an intervention direction to shift the model’s focus towards different levels of visual information, enhancing its attention to high-level and fine-grained visual details. During the forward pass stage, the intervention is applied to the attention heads that encode the overall image information and the fine-grained object details, effectively mitigating the phenomenon of overly language priors, and thereby alleviating hallucinations. Extensive experiments demonstrate that ICT achieves strong performance with a small amount of data and generalizes well across different datasets and models. Our codes are publicly available at:https://github.com/THU-BPM/ICT/. Junzhe Chen 0001, Tianshu Zhang 0002, Shiyu Huang 0001, Yuwei Niu, Linfeng Zhang 0001, Lijie Wen 0001, Xuming Hu |
CVPR | 3 |
| 2025 | MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language ModelsabstractIn recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability — fine-grained motion comprehension — remains under-explored in current benchmarks. To address this gap, we propose MotionBench, a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. MotionBench evaluates models’ motion-level perception through six primary categories of motion-oriented question types and includes data collected from diverse sources, ensuring a broad representation of real-world video content. Experimental results reveal that existing VLMs perform poorly in understanding fine-grained motions. To enhance VLM’s ability to perceive fine-grained motion within a limited sequence length of LLM, we conduct extensive experiments reviewing VLM architectures optimized for video feature compression and propose a novel and efficient Through-Encoder (TE) Fusion method. Experiments show that higher frame rate inputs and TE Fusion yield improvements in motion understanding, yet there is still substantial room for enhancement. Our benchmark aims to guide and motivate the development of more capable video understanding models, emphasizing the importance of fine-grained motion comprehension. Project page: https://motion-bench.github.io. Wenyi Hong, Yean Cheng, Zhuoyi Yang, Lefan Wang, Xiaotao Gu, Shiyu Huang 0001, Yuxiao Dong, Jie Tang 0001 |
CVPR | 7 |
| 2025 | LVBench: An Extreme Long Video Understanding BenchmarkabstractRecent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io. Zehai He, Wenyi Hong, Yean Cheng, Ji Qi 0003, Ming Ding 0004, Xiaotao Gu, Shiyu Huang 0001, Bin Xu 0001, Yuxiao Dong, Jie Tang 0001 |
ICCV | 9 |
| 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerabstractWe present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels.
Previous video generation models often struggled with limited motion and short durations.
It is especially difficult to generate videos with coherent narratives based on text.
We propose several designs to address these issues.
First, we introduce a 3D Variational Autoencoder (VAE) to compress videos across spatial and temporal dimensions, enhancing both the compression rate and video fidelity.
Second, to improve text-video alignment, we propose an expert transformer with expert adaptive LayerNorm to facilitate the deep fusion between the two modalities.
Third, by employing progressive training and multi-resolution frame packing, CogVideoX excels at generating coherent, long-duration videos with diverse shapes and dynamic movements.
In addition, we develop an effective pipeline that includes various pre-processing strategies for text and video data.
Our innovative video captioning model significantly improves generation quality and semantic alignment.
Results show that CogVideoX achieves state-of-the-art performance in both automated benchmarks and human evaluation.
We publish the code and model checkpoints of CogVideoX along with our VAE model and video captioning model at https://github.com/THUDM/CogVideo. Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding 0004, Shiyu Huang 0001, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Guanyu Feng, Da Yin, Yean Cheng, Bin Xu 0001, Xiaotao Gu, Yuxiao Dong, Jie Tang 0001 |
ICLR | 5 |
| 2025 | Can Large Language Models Master Complex Card Games?abstractComplex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players in Go and Chess, garnering widespread societal attention towards artificial intelligence. Concurrently, large language models (LLMs) have exhibited remarkable capabilities across various tasks, raising the question of whether LLMs can achieve similar success in complex games. In this paper, we explore the potential of LLMs in mastering complex card games. We systematically assess the learning capabilities of LLMs across eight diverse card games, evaluating the impact of fine-tuning on high-quality gameplay data, and examining the models' ability to retain general capabilities while mastering these games. Our findings indicate that: (1) LLMs can approach the performance of strong game AIs through supervised fine-tuning on high-quality data, (2) LLMs can achieve a certain level of proficiency in multiple complex card games simultaneously, with performance augmentation for games with similar rules and conflicts for dissimilar ones, and (3) LLMs experience a decline in general capabilities when mastering complex games, but this decline can be mitigated by integrating a certain amount of general instruction data. The evaluation results demonstrate strong learning ability and versatility of LLMs. The code is available at
https://github.com/THUDM/LLM4CardGame Wei Wang 0074, Fuqing Bie, Junzhe Chen 0001, Shiyu Huang 0001, Evgeny Kharlamov, Jie Tang 0001 |
NeurIPS | 5 |
| 2024 | DGPO: Discovering Multiple Strategies with Diversity-Guided Policy OptimizationabstractMost reinforcement learning algorithms seek a single optimal strategy that solves a given task. However, it can often be valuable to learn a diverse set of solutions, for instance, to make an agent's interaction with users more engaging, or improve the robustness of a policy to an unexpected perturbance. We propose Diversity-Guided Policy Optimization (DGPO), an on-policy algorithm that discovers multiple strategies for solving a given task. Unlike prior work, it achieves this with a shared policy network trained over a single run. Specifically, we design an intrinsic reward based on an information-theoretic diversity objective. Our final objective alternately constraints on the diversity of the strategies and on the extrinsic reward. We solve the constrained optimization problem by casting it as a probabilistic inference task and use policy iteration to maximize the derived lower bound. Experimental results show that our method efficiently discovers diverse strategies in a wide variety of reinforcement learning tasks. Compared to baseline methods, DGPO achieves comparable rewards, while discovering more diverse strategies, and often with better sample efficiency. Wentse Chen, Shiyu Huang 0001, Yuan Chiang, Tim Pearce, Wei-Wei Tu, Ting Chen 0006, Jun Zhu 0001 |
AAAI | 2 |
| 2024 | LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent EnvironmentsabstractJunzhe Chen, Xuming Hu, Shuodi Liu, Shiyu Huang, Wei-Wei Tu, Zhaofeng He, Lijie Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junzhe Chen 0001, Xuming Hu, Shuodi Liu, Shiyu Huang 0001, Wei-Wei Tu, Zhaofeng He 0001, Lijie Wen 0001 |
ACL (1) | 4 |
| 2024 | MQE: Unleashing the Power of Interaction with Multi-agent Quadruped EnvironmentabstractThe advent of deep reinforcement learning (DRL) has significantly advanced the field of robotics, particularly in the control and coordination of quadruped robots. However, the complexity of real-world tasks often necessitates the deployment of multi-robot systems capable of sophisticated interaction and collaboration. To address this need, we introduce the Multi-agent Quadruped Environment (MQE), a novel platform designed to facilitate the development and evaluation of multi-agent reinforcement learning (MARL) algorithms in realistic and dynamic scenarios. MQE emphasizes complex interactions between robots and objects, hierarchical policy structures, and challenging evaluation scenarios that reflect real-world applications. We present a series of collaborative and competitive tasks within MQE, ranging from simple coordination to complex adversarial interactions, and benchmark state-of-the-art MARL algorithms. Our findings indicate that hierarchical reinforcement learning can simplify task learning, but also highlight the need for advanced algorithms capable of handling the intricate dynamics of multi-agent interactions. MQE serves as a stepping stone towards bridging the gap between simulation and practical deployment, offering a rich environment for future research in multi-agent systems and robot learning. For open-sourced code and more details of MQE, please refer to https://ziyanx02.github.io/multiagent-quadruped-environment/. Ziyan Xiong, Shiyu Huang 0001, Wei-Wei Tu, Zhaofeng He 0001 |
IROS | 3 |
| 2023 | SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksabstractWe introduce SwiftSage, a novel agent framework inspired by the dual-process theory of human cognition, designed to excel in action planning for complex interactive reasoning tasks. SwiftSage integrates the strengths of behavior cloning and prompting large language models (LLMs) to enhance task completion performance. The framework comprises two primary modules: the Swift module, representing fast and intuitive thinking, and the Sage module, emulating deliberate thought processes. The Swift module is a small encoder-decoder LM fine-tuned on the oracle agent's action trajectories, while the Sage module employs LLMs such as GPT-4 for subgoal planning and grounding. We develop a heuristic method to harmoniously integrate the two modules, resulting in a more efficient and robust problem-solving process. In 30 tasks from the ScienceWorld benchmark, SwiftSage significantly outperforms other methods such as SayCan, ReAct, and Reflexion, demonstrating its effectiveness in solving complex interactive tasks. Bill Y. Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang 0001, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi 0001, Xiang Ren 0001 |
NeurIPS | 5 |
| 2022 | VMAPD: Generate Diverse Solutions for Multi-Agent Games with Recurrent Trajectory DiscriminatorsabstractRecent algorithms designed for multi-agent tasks focus on finding a single optimal solution for all the agents. However, in many tasks (e.g., matrix games and transportation dispatching), there may exist more than one optimal solution, while previous algorithms can only converge to one of them. In many practical applications, it is important to develop reasonable agents with diverse behaviors. In this paper, we propose ”variational multi-agent policy diversification” (VMAPD), an on-policy framework for discovering diverse policies for coordination patterns of multiple agents. By taking advantage of latent variables and exploiting the connection between variational inference and multi-agent reinforcement learning, we derive a tractable evidence lower bound (ELBO) on the trajectories of all agents. Our algorithm uses policy iteration to maximize the derived lower bound and can be simply implemented by adding a pseudo reward during centralized learning. And the trained agents do not need to access the pseudo reward during decentralized execution. We demonstrate the effectiveness of our algorithm on several popular multi-agent testbeds. Experimental results show that VMAPD finds more solutions with similar sample complexity compared with other baselines. Shiyu Huang 0001, Chao Yu 0005, Bin Wang 0034, Dong Li 0016, Yu Wang 0002, Ting Chen 0006, Jun Zhu 0001 |
CoG | 1 |
| 2022 | Diverse Policies Converge in Reward-Free Markov Decision Processes
Fanqi Lin, Shiyu Huang 0001, Wei-Wei Tu |
PRICAI (1) | 2 |
| 2022 | Deep reinforcement learning with credit assignment for combinatorial optimization
Jiayi Weng, Shiyu Huang 0001, Chongxuan Li, Yichi Zhou, Hang Su 0006, Jun Zhu 0001 |
Pattern Recognit. | 3 |
| 2021 | Off-Policy Training for Truncated TD(λ) Boosted Soft Actor-Critic
Shiyu Huang 0001, Bin Wang 0034, Hang Su 0006, Dong Li 0016, Jianye Hao, Jun Zhu 0001, Ting Chen 0006 |
PRICAI (3) | 1 |
| 2020 | SVQN: Sequential Variational Soft Q-Learning Networks
Shiyu Huang 0001, Hang Su 0006, Jun Zhu 0001, Ting Chen 0006 |
ICLR | 1 |
| 2019 | Combo-Action: Training Agent For FPS Game with Auxiliary TasksabstractDeep reinforcement learning (DRL) has achieved surpassing human performance on Atari games, using raw pixels and rewards to learn everything. However, first-person-shooter (FPS) games in 3D environments contain higher levels of human concepts (enemy, weapon, spatial structure, etc.) and a large action space. In this paper, we explore a novel method which can plan on temporally-extended action sequences, which we refer as Combo-Action to compress the action space. We further train a deep recurrent Q-learning network model as a high-level controller, called supervisory network, to manage the Combo-Actions. Our method can be boosted with auxiliary tasks (enemy detection and depth prediction), which enable the agent to extract high-level concepts in the FPS games. Extensive experiments show that our method is efficient in training process and outperforms previous stateof-the-art approaches by a large margin. Ablation study experiments also indicate that our method can boost the performance of the FPS agent in a reasonable way. Shiyu Huang 0001, Hang Su 0006, Jun Zhu 0001, Ting Chen 0006 |
AAAI | 1 |
| 2017 | Expecting the Unexpected: Training Detectors for Unusual Pedestrians with Adversarial ImpostersabstractAs autonomous vehicles become an every-day reality, high-accuracy pedestrian detection is of paramount practical importance. Pedestrian detection is a highly researched topic with mature methods, but most datasets (for both training and evaluation) focus on common scenes of people engaged in typical walking poses on sidewalks. But performance is most crucial for dangerous scenarios that are rarely observed, such as children playing in the street and people using bicycles/skateboards in unexpected ways. Such in-the-tail data is notoriously hard to observe, making both training and testing difficult. To analyze this problem, we have collected a novel annotated dataset of dangerous scenarios called the Precarious Pedestrian dataset. Even given a dedicated collection effort, it is relatively small by contemporary standards (≈ 1000 images). To explore large-scale data-driven learning, we explore the use of synthetic data generated by a game engine. A significant challenge is selected the right priors or parameters for synthesis: we would like realistic data with realistic poses and object configurations. Inspired by Generative Adversarial Networks, we generate a massive amount of synthetic data and train a discriminative classifier to select a realistic subset (that fools the classifier), which we deem Synthetic Imposters. We demonstrate that this pipeline allows one to generate realistic (or adverserial) training data by making use of rendering/animation engines. Interestingly, we also demonstrate that such data can be used to rank algorithms, suggesting that Synthetic Imposters can also be used for in-the-tail validation at test-time, a notoriously difficult challenge for real-world deployment. Shiyu Huang 0001, Deva Ramanan |
CVPR | 1 |