EDBT 2026 Demo / reviewers in the wild / expert
Sirui Han
dblp:373/4263
· DBLP profile ↗
33ranked-venue papers
1as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryabstractData scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availability of precise and reasonable control instructions. Current methods primarily rely on 2D trajectories as instruction prompts, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3Dfor generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length, avoiding collisions and retiming. Next, we employ a latent editing technique to create video sequences from the initial image latent, text instruction and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method significantly reduces human intervention requirements by autonomously planing plausible 3D trajectories. Experimental results demonstrate its superior visual quality and precision. Ying Li 0128, Xiaobao Wei, Xiaowei Chi, Zhongyu Zhao, Hao Wang 0073, Ningning Ma, Ming Lu 0002, Sirui Han |
AAAI | 9 |
| 2026 | Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingabstractMixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods aim to achieve greater efficiency by consolidating several experts, they are fundamentally hindered by parameter conflicts arising from expert specialization. In this paper, we present Sub-MoE, a novel MoE compression framework via Subspace Expert Merging. Our key insight is to perform joint Singular Value Decomposition (SVD) on concatenated expert weights, reducing conflicting parameters by extracting shared U-matrices while enabling effective merging of the expert-specific V components. Specifically, Sub-MoE consists of two innovative stages: (1) Adaptive Expert Clustering, which groups functionally coherent experts via K-means clustering based on cosine similarity of expert outputs; and (2) Subspace Expert Merging, which first performs Experts Union Decomposition to derive the shared U-matrix across experts in the same group, then applies frequency-based merging for individual V-matrices, and completes expert reconstruction using the merged V-matrix. In this way, we align and fuse experts in a shared subspace. Additionally, the framework can be extended with intra-expert compression for further inference optimization. Extensive experiments on Mixtral, DeepSeek, and Qwen-1.5/3 MoE LLMs demonstrate that our Sub-MoE significantly outperforms existing expert pruning and merging methods. Notably, our Sub-MoE maintains 96%/86% of original performance with 25%/50% expert reduction on Mixtral-8×7B in zero-shot benchmarks. Lujun Li 0001, Qiyuan Zhu, Xiaoyu Qin 0001, Wei Li 0286, Hao Gu 0001, Sirui Han, Yike Guo |
AAAI | 7 |
| 2026 | What, Whether and How? Unveiling Process Reward Models for Thinking with Images ReasoningabstractThe rapid advancement of Large Vision Language Models (LVLMs) has demonstrated excellent abilities in various visual tasks. Building upon these developments, the thinking with images paradigm has emerged, enabling models to dynamically edit and re-encode visual information at each reasoning step, mirroring human visual processing. However, this paradigm introduces significant challenges as diverse errors may occur during reasoning processes. This necessitates Process Reward Models (PRMs) for distinguishing positive and negative reasoning steps, yet existing benchmarks for PRMs are predominantly text-centric and lack comprehensive assessment under this paradigm. To address these gaps, this work introduces the first comprehensive benchmark specifically designed for evaluating PRMs under the thinking with images paradigm. Our main contributions are: (1) Through extensive analysis of reasoning trajectories and guided search experiments with PRMs, we define 7 fine-grained error types and demonstrate both the necessity for specialized PRMs and the potential for improvement. (2) We construct a comprehensive benchmark comprising 1,206 manually annotated thinking with images reasoning trajectories spanning 4 categories and 16 subcategories for fine-grained evaluation of PRMs. (3) Our experimental analysis reveals that current LVLMs fall short as effective PRMs, exhibiting limited capabilities in visual reasoning process evaluation with significant performance disparities across error types, positive evaluation bias, and sensitivity to reasoning step positions. These findings demonstrate the effectiveness of our benchmark and establish crucial foundations for advancing PRMs in LVLMs. Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Chi-Min Chan, Sirui Han |
AAAI | 9 |
| 2026 | Outlier Matters: Efficient Long-to-Short Reasoning via Outlier-Guided Model MergingabstractLarge Reasoning Language Models (LRMs) have recently shown remarkable performance in complex reasoning tasks, but their extensive reasoning chains incur substantial computational overhead. To address this challenge, we propose Outlier-aware Reasoning Conciseness Adaptive Merge (ORCA), a novel plug-and-play model merging framework that leverages outlier activation patterns to fuse base models with reasoning models. Our ORCA introduces three key innovations: (1) adaptive alignment that reduces conflicts between disparate activation patterns during merging, (2) outlier-guided allocation that assigns merging coefficients proportional to each layer's reasoning importance as indicated by outlier concentrations, and (3) dynamic probe-based adjustment that adapts merging coefficients during inference based on input-specific activation characteristics. These strategies allow seamless integration into existing merging pipelines while creating unified models that maintain reasoning accuracy with significantly reduced response verbosity. Comprehensive evaluation across six benchmarks using Qwen and LLaMA models shows ORCA reduces average response length by 55% while improving accuracy by 2.4∼5.7% over existing methods. Qiyuan Zhu, Lujun Li 0001, Xiaoyu Qin 0001, Wei Li 0286, Hao Gu 0001, Sirui Han, Yike Guo |
AAAI | 8 |
| 2026 | Benchmarking Fine-Grained Error Detection in Multimodal ReasoningabstractChi-Min Chan, Han Zhu, Chunyang Jiang, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi-Min Chan, Jiaming Ji, Juntao Dai, Wei Xue 0002, Sirui Han, Yike Guo |
ACL (1) | 7 |
| 2026 | Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesabstractChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue 0002, Sirui Han, Yike Guo |
ACL (1) | 8 |
| 2026 | Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language ModelsabstractLarge Language Models (LLMs) have achieved exceptional performance in complex reasoning via Chain-of-Thought (CoT), yet the associated computational costs remain prohibitive. CoT reasoning contains significant untapped efficiency potential across two dimensions: temporal redundancy, where reasoning steps may be unnecessary, and spatial redundancy, where computations can be performed at reduced precision. While current optimization techniques often necessitate resource-intensive fine-tuning or data curation, we introduce ASTRO (Adaptive Spatial and Temporal Redundancy Optimization), a training-free framework that simultaneously addresses both dimensions. ASTRO leverages Dewey’s reflective thinking model to segment reasoning phases, applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. Empirical results across diverse reasoning benchmarks demonstrate that ASTRO achieves up to an 11.3 \times efficiency gain without compromising accuracy, highlighting the advantages of holistic multi-dimensional redundancy management over isolated optimization methods. Pengyu Cheng, Qiyuan Zhu, Hao Gu 0001, Ruijie Shen, Xiaofeng Hou, Sirui Han, Jiacheng Liu 0001 |
ACL (1) | 9 |
| 2026 | Perception, Understanding and Reasoning: A Multimodal Benchmark for Video Fake News DetectionabstractCui Yakun, Peng Qi, Fushuo Huo, Hang Du, Weijie Shi, Juntao Dai, Zhenghao Zhu, Sirui Han. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yakun Cui, Fushuo Huo, Juntao Dai, Zhenghao Zhu, Sirui Han |
ACL (1) | 8 |
| 2026 | BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary CodebookabstractHao Gu, Lujun Li, Hao Wang, Lei Wang, Zheyu Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Sirui Han, Yike Guo |
ACL (1) | 9 |
| 2026 | ContextLens: Modeling Imperfect Privacy and Safety Context for Legal ComplianceabstractHaoran Li, Yulin Chen, Huihao Jing, Wenbin Hu, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haoran Li 0003, Huihao Jing, Wenbin Hu 0001, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song |
ACL (1) | 8 |
| 2026 | Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning ModelsabstractHao Wang, Hao Gu, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue, Sirui Han, Yike Guo, Dapeng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Wang 0193, Hao Gu 0001, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue 0001, Sirui Han, Yike Guo, Dapeng Oliver Wu |
ACL (1) | 7 |
| 2026 | Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMsabstractBinxing Xu, Hao Gu, Lujun Li, Hao Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Xintong Yang, Chao Li, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Binxing Xu, Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Xintong Yang, Chao Li 0009, Sirui Han, Yike Guo |
ACL (1) | 10 |
| 2026 | SafeMT: Multi-turn Safety for Multimodal Language ModelsabstractHan Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Juntao Dai, Jiaming Ji, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen 0008, Yaodong Yang 0001, Sirui Han, Yike Guo |
ACL (1) | 10 |
| 2026 | Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI CollaborationabstractFact verification is a critical yet underexplored component of non-litigation legal practice. While existing research has examined automation in legal workflow and human-AI collaboration in high-stakes domains, little is known about how GenAI can support fact verification, a task that demands prudent judgment and strict accountability. To address this, we conducted semi-structured interviews with 18 lawyers to understand their current verification practices, attitudes toward GenAI adoption, and expectations for future systems. We found that while lawyers use GenAI for low-risk tasks like drafting and language optimization, concerns over accuracy, confidentiality, and liability are currently limiting its adoption for fact verification. These concerns translate into core design requirements for AI systems that are trustworthy and accountable. Based on these, we contribute design insights for human-AI collaboration in legal fact verification, emphasizing the development of auditable systems that balance efficiency with professional judgment and uphold ethical and legal accountability in high-stakes practice. Sirui Han, Yuyao Zhang 0006, Yidan Huang, Chengzhong Liu, Yike Guo |
CHI | 1 |
| 2026 | CareerCraft: Supporting New Graduates on Job Hunting with LLM-Assisted Self-Construction of Career ProfileabstractStarting the job hunt is often challenging for new graduates, who face barriers in translating experiences into actionable career profiles due to limited self-awareness and unclear skill mapping. Through formative study with new graduates and early-career professionals, we concluded specific challenges in experience extraction, skill organization, and expressive confidence. Drawing on these insights, we designed CareerCraft, an interactive system that scaffolds the construction of coherent career stories and supports tailored job searching via experience card extraction, guided profile building, and LLM-powered recommendations. In a within-subject evaluation (N=16), participants rated the efficacy of CareerCraft against the baseline condition without the tool in improving profile structuring, clarifying their self-awareness and competencies, and supporting informed job direction choices. Based on the findings, we concluded that CareerCraft offered a promising pathway to career readiness among new graduates to the workforce. We further summarized the design considerations for LLM products emphasizing on users’ self-exploration. Xinyue Qi, Chengzhong Liu, Xiangyu Long, Zhizhuo Kou, Sirui Han, Yike Guo |
CHI | 5 |
| 2026 | An image steganography algorithm using selective timestep embedding and diffusion model
Jiajun Han, Guodong Ye, Sirui Han |
Expert Syst. Appl. | 3 |
| 2025 | FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning EvaluationabstractJunyu Luo, Zhizhuo Kou, Liming Yang, Xiao Luo, Jinsheng Huang, Zhiping Xiao, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang, Yike Guo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junyu Luo 0002, Zhizhuo Kou, Xiao Luo 0001, Jinsheng Huang, Zhiping Xiao 0001, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang 0004, Yike Guo |
ACL (1) | 11 |
| 2025 | PrivaCI-Bench: Evaluating Privacy with Contextual Integrity and Legal ComplianceabstractHaoran Li, Wenbin Hu, Huihao Jing, Yulin Chen, Qi Hu, Sirui Han, Tianshu Chu, Peizhao Hu, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haoran Li 0003, Wenbin Hu 0001, Huihao Jing, Sirui Han, Peizhao Hu, Yangqiu Song |
ACL (1) | 6 |
| 2025 | PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceabstractJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001 |
ACL (1) | 11 |
| 2025 | LegalReasoner: Step-wised Verification-Correction for Legal Judgment ReasoningabstractWeijie Shi, Han Zhu, Jiaming Ji, Mengze Li, Jipeng Zhang, Ruiyuan Zhang, Jia Zhu, Jiajie Xu, Sirui Han, Yike Guo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Mengze Li 0001, Ruiyuan Zhang, Jia Zhu 0003, Jiajie Xu 0001, Sirui Han, Yike Guo |
ACL (1) | 9 |
| 2025 | Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem ProvingabstractChuxue Cao, Mengze Li, Juntao Dai, Jinluan Yang, Zijian Zhao, Shengyu Zhang, Weijie Shi, Chengzhong Liu, Sirui Han, Yike Guo. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chuxue Cao, Mengze Li 0001, Juntao Dai, Jinluan Yang, Zijian Zhao 0002, Shengyu Zhang 0001, Chengzhong Liu, Sirui Han, Yike Guo |
EMNLP | 9 |
| 2025 | Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement LearningabstractWenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Xu Heli, Tianshu Chu, Peizhao Hu, Yangqiu Song. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wenbin Hu 0001, Haoran Li 0003, Huihao Jing, Ziqian Zeng, Sirui Han, Heli Xu, Peizhao Hu, Yangqiu Song |
EMNLP | 6 |
| 2025 | DIDS: Domain Impact-aware Data Sampling for Large Language Model TrainingabstractWeijie Shi, Jipeng Zhang, Yaguang Wu, Jingzhi Fang, Shibo Zhang, Yao Zhao, Hao Chen, Ruiyuan Zhang, Yue Cui, Jia Zhu, Sirui Han, Jiajie Xu, Xiaofang Zhou. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yaguang Wu, Jingzhi Fang, Ruiyuan Zhang, Yue Cui 0001, Jia Zhu 0003, Sirui Han, Jiajie Xu 0001, Xiaofang Zhou 0001 |
EMNLP | 11 |
| 2025 | Efficient Fine-Tuning of Large Models Via Nested Low-Rank Adaptation
Lujun Li 0001, Cheng Lin 0001, You-Liang Huang, Wei Li 0286, Jie Zou 0001, Wei Xue 0002, Sirui Han, Yike Guo |
ICCV | 9 |
| 2025 | AIRA: Activation-Informed Low-Rank Adaptation for Large ModelsabstractLow-Rank Adaptation (LoRA) is a widely used method for efficiently fine-tuning large models by introducing lowrank matrices into weight updates. However, existing LoRA techniques fail to account for activation information, such as outliers, which significantly impact model performance. This omission leads to suboptimal adaptation and slower convergence. To address this limitation, we present Activation-Informed Low-Rank Adaptation (AIRA), a novel approach that integrates activation information into initialization, training, and rank assignment to enhance model performance. Specifically, AIRA introduces: (1) Outlierweighted SVD decomposition to reduce approximation errors in low-rank weight initialization, (2) Outlier-driven dynamic rank assignment using offline optimization for better layer-wise adaptation, and (3) Activation-informed training to amplify updates on significant weights. This cascaded activation-informed paradigm enables faster convergence and fewer fine-tuned parameters while maintaining high performance. Extensive experiments on multiple large models demonstrate that AIRA outperforms state-of-the-art LoRA variants, achieving superior performance-efficiency trade-offs in vision-language instruction tuning, few-shot learning, and image generation. Codes are available at https://github.com/lliai/LoRA-Zoo. Lujun Li 0001, Cheng Lin 0001, Wei Li 0286, Wei Xue 0002, Sirui Han, Yike Guo |
ICCV | 6 |
| 2025 | DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsabstractGenerating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multiturn editable dance dataset comprising the prompt featuring over 25.3 M dance frames and 84.5 K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multimodal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/. Xingqun Qi, Muyi Sun, Siye Wang, Man Zhang 0005, Sirui Han |
ICCV | 8 |
| 2025 | Consistent and Invariant Generalization Learning for Short-video Misinformation DetectionabstractShort-video misinformation detection has attracted wide attention in the multi-modal domain, aiming to accurately identify the misinformation in the video format accompanied by the corresponding audio. Despite significant advancements, current models in this field, trained on particular domains (source domains), often exhibit unsatisfactory performance on unseen domains (target domains) due to domain gaps. To effectively realize such domain generalization on the short-video misinformation detection task, we propose deep insights into the characteristics of different domains: (1) The detection on various domains may mainly rely on different modalities (i.e., mainly focusing on videos or audios). To enhance domain generalization, it is crucial to achieve optimal model performance on all modalities simultaneously. (2) For some domains focusing on cross-modal joint fraud, a comprehensive analysis relying on cross-modal fusion is necessary. However, domain biases located in each modality (especially in each frame of videos) will be accumulated in this fusion process, which may seriously damage the final identification of misinformation. To address these issues, we propose a new DOmain generalization model via ConsisTency and invariance learning for shORt-video misinformation detection (named DOCTOR), which contains two characteristic modules: (1) We involve the cross-modal feature interpolation to map multiple modalities into a shared space and the interpolation distillation to synchronize multi-modal learning; (2) We design the diffusion model to add noise to retain core features of multi modal and enhance domain invariant features through cross-modal guided denoising. Extensive experiments demonstrate the effectiveness of our proposed DOCTOR model. Our code is publicly available at https://github.com/ghh1125/DOCTOR. Hanghui Guo, Mengze Li 0001, Juncheng Li 0006, Yue Cui 0001, Jiajie Xu 0001, Jia Zhu 0003, Zhangze Chen, Sirui Han |
ACM Multimedia | 11 |
| 2025 | Outlier-Aware Model Merging for Efficient Multitask InferenceabstractModel merging techniques aim to consolidate multiple fine-tuned models into a single unified model, reducing both storage and computational overhead while retaining task-specific performance. However, existing methods face several limitations: monotonous compression techniques that fail to account for task-specific weight distribution characteristics, weight-magnitude-based compression that fails to consider functional importance revealed by activation patterns, and non-adaptive allocation strategies that ignores task-specific layer importance. To overcome these challenges, we propose OA-Merge, a novel Outlier-Aware Model Merging framework that leverages task activation outliers to enable adaptive compression and resource allocation across tasks. OA-Merge comprises three key components: (1) dynamic hybrid decomposition technique that formulates task vectors as tailored combinations of low-rank and sparse components adapted to task-specific statistical distributions, (2) activation-informed compression methodology that incorporates task-specific activation statistics to prioritize functionally important weights, and (3) task-related allocation that optimizes the distribution of compression resources according to layer-specific importance metrics derived from activation outlier analysis. These hybrid outlier-aware strategies adapt dynamically to each task's intrinsic characteristics, avoiding the pitfalls of one-size-fits-all ways. Extensive experiments on both vision models (e.g., ViT) and language models (e.g., RoBERTa, Qwen) demonstrate that OA-Merge outperforms state-of-the-art baselines, achieving average performance gains of 3.2% on vision tasks and 2.8% on language tasks. Qiyuan Zhu, Lujun Li 0001, Jiacheng Liu 0001, Pengyu Cheng, Sirui Han, Yike Guo |
ACM Multimedia | 7 |
| 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human FeedbackabstractAs multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Boyuan Chen 0008, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Kaile Wang, Juntao Dai, Xuyao Wang, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 13 |
| 2025 | Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackabstractMultimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: (I) BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); (II) Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model’s overall safety by an average of 40.9%. (II) Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%. Jiaming Ji, Donghai Hong, Boyuan Chen 0008, Kaile Wang, Juntao Dai, Chi-Min Chan, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 12 |
| 2025 | IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse RenderingabstractVision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active creation rather than passive recognition. Grounded in the analysis-by-synthesis paradigm, IR3D-Bench tasks Vision-Language Agents (VLAs) with actively using programming and rendering tools to recreate the underlying 3D structure of an input image, achieving agentic inverse rendering through tool use. This ''understanding-by-creating'' approach probes the tool-using generative capacity of VLAs, moving beyond the descriptive or conversational capacity measured by traditional scene understanding benchmarks. We provide a comprehensive suite of metrics to evaluate geometric accuracy, spatial relations, appearance attributes, and overall plausibility. Initial experiments on agentic inverse rendering powered by various state-of-the-art VLMs highlight current limitations, particularly in visual precision rather than basic tool usage. IR3D-Bench, including data and evaluation protocols, is released to facilitate systematic study and development of tool-using VLAs towards genuine scene understanding by creating. Hengyu Liu 0007, Chenxin Li, Yipeng Wu, Wuyang Li, Zhiqin Yang, Zhenyuan Zhang 0001, Yunlong Lin, Sirui Han, Brandon Yushan Feng |
NeurIPS | 9 |
| 2025 | Semantic-guided Diverse Decoding for Large Language ModelabstractDiverse decoding of large language models is crucial for applications requiring multiple semantically distinct responses, yet existing methods primarily achieve lexical rather than semantic diversity. This limitation significantly constrains Best-of-N strategies, group-based reinforcement learning, and data synthesis. While temperature sampling and diverse beam search modify token distributions or apply n-gram penalties, they fail to ensure meaningful semantic differentiation. We introduce Semantic-guided Diverse Decoding (SemDiD), operating directly in embedding space that balances quality with diversity through three complementary mechanisms: orthogonal directional guidance, dynamic inter-group repulsion, and position-debiased probability assessment. SemDiD harmonizes these competing objectives using adaptive gain functions and constraint optimization, ensuring both quality thresholds and maximal semantic differentiation. Experiments show SemDiD consistently outperforms existing methods, improving Best-of-N coverage by 1.4-5.2% across diverse tasks and accelerating RLHF training convergence by 15% while increasing accuracy by up to 2.1%. Yue Cui 0001, Yaguang Wu, Jingzhi Fang, Mengze Li 0001, Sirui Han, Jia Zhu 0003, Jiajie Xu 0001, Xiaofang Zhou 0001 |
NeurIPS | 7 |
| 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceabstractTraining multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses. Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 7 |