Bohan Wu

dblp:133/4453 · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 11 since 2021Systems, architecture and hardware · 5 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CogniTrust: Cognitive Memory-Driven Verifiable Supervision for Robust Hashing
abstract
In this paper, we study the problem of robust multi-label hashing, where label noise hinders the learning of a reliable semantic structure from data. Many existing methods rely on heuristic sample selection or consistency-based training, but lack a unified mechanism to validate and refine supervision across structural and semantic levels. Inspired by cognitive theories of human memory, we propose a novel framework called CogniTrust that unifies verifiable supervision with a triadic memory model: a) In episodic memory, feature activations are decomposed into spatial patterns that support the assessment of structural evidence and the estimation of label reliability; b) Semantic memory keeps track of class-level prototypes from structurally attentive regions to estimate the semantic plausibility of labels; c) Reconstructive memory simulates memory recall through interpolation between images using a diffusion-based mixup process, which enriches the training signals for semantically uncertain regions. These components work together, allowing supervision to be refined through the joint consideration of spatial structure and semantic information. Extensive experiments on noisy hashing benchmarks demonstrate that CogniTrust consistently outperforms a range of state-of-the-art baselines. Our results show that cognitive memory mechanisms offer a principled basis for more reliable label denoising and robust hashing.
Yiyang Gu, Bohan Wu, Yifang Qin, Jiaru Tang, Rongcheng Tu, Zhiping Xiao 0001, Taian Guo, Junyu Luo 0002, Wei Ju 0001, Xiao Luo 0001, Dacheng Tao, Ming Zhang 0004
AAAI2
2026 SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
abstract
Yiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan, Bin Feng, Yingce Xia, Shufang Xie, Kaili Liu, Bohan Wu, Qi Shi, Haoran Li, Beier Xiao, Zhiping Xiao, Xiao Luo, Weizhi Zhang, Philip S. Yu, Zequn Liu, Ming Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yiyang Gu, Junyu Luo 0002, Ye Yuan 0016, Yingce Xia, Shufang Xie 0003, Kaili Liu, Bohan Wu, Haoran Li 0003, Beier Xiao, Zhiping Xiao 0001, Xiao Luo 0001, Weizhi Zhang 0001, Philip S. Yu, Zequn Liu, Ming Zhang 0004
ACL (1)9
2026 EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language Models
abstract
Retrieval-augmented generation (RAG) is widely adopted for radiology report generation with medical vision-language models, leveraging external reports as linguistic references. However, existing RAG methods rely primarily on dense embedding similarity, which may retrieve reports that are semantically related yet clinically inconsistent with respect to presence or laterality constraints. Such inconsistencies are often propagated into generation, resulting in contradictory or unsupported findings. We propose an evidence-guided retrieval-augmented framework EviRAG that decomposes retrieval into structured and unstructured alignment levels. First, we induce structured clinical triplets from both query and database cases through targeted visual interrogation, projecting images into a shared evidence space. Triplet-level alignment enforces explicit agreement over presence and laterality variables, yielding a clinically admissible candidate set via structural ranking. Within this constrained space, we perform semantic alignment in a shared multimodal embedding space to capture nuanced descriptive correspondence. The top-ranked reports and query image are jointly fed into a medical vision-language model for report generation. Comprehensive experiments on radiology report generation benchmarks show that EviRAG substantially reduces clinical inconsistencies compared to strong medical vision-language baselines. The source code is available at https://github.com/liamgu06/EviRAG.
Yiyang Gu, Jiayue Fan, Kaili Liu, Bohan Wu, Binqi Chen, Zequn Liu, Zhiping Xiao 0001, Rongcheng Tu, Xiao Luo 0001, Ming Zhang 0004
SIGIR4
2026 Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
abstract
Variational inference (VI) has emerged as a popular method for approximate inference for high-dimensional Bayesian models. In this paper, we propose a novel VI method that extends the naive mean field via entropic regularization, referred to as $\Xi$-variational inference ($\Xi$-VI). $\Xi$-VI has a close connection to the entropic optimal transport problem and benefits from the computationally efficient Sinkhorn algorithm. We show that $\Xi$-variational posteriors effectively recover the true posterior dependency, where the likelihood function is downweighted by a regularization parameter. We analyze the role of dimensionality of the parameter space on the accuracy of $\Xi$-variational approximation and the computational complexity of computing the approximate distribution, providing a rough characterization of the statistical-computational trade-off in $\Xi$-VI, where higher statistical accuracy requires greater computational effort. We also investigate the frequentist properties of $\Xi$-VI and establish results on consistency, asymptotic normality, high-dimensional asymptotics, and algorithmic stability. We provide sufficient criteria for our algorithm to achieve polynomial-time convergence. Finally, we show the inferential benefits of using $\Xi$-VI over mean-field VI and other competing methods, such as normalizing flow, on simulated and real datasets.
Bohan Wu, David M. Blei
J. Mach. Learn. Res.1
2025 A Survey on Efficient Large Language Model Training: From Data-centric Perspectives
abstract
Junyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao, Yiqiao Jin, Rong-Cheng Tu, Nan Yin, Yifan Wang, Jingyang Yuan, Wei Ju, Ming Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Junyu Luo 0002, Bohan Wu, Xiao Luo 0001, Zhiping Xiao 0001, Yiqiao Jin, Rongcheng Tu, Yifan Wang 0014, Jingyang Yuan, Wei Ju 0001, Ming Zhang 0004
ACL (1)2
2025 GeT-USE: Learning Generalized Tool Usage for Bimanual Mobile Manipulation via Simulated Embodiment Extensions
abstract
The ability to use random objects as tools in a generalizable manner is a missing piece in robots’ intelligence today to boost their versatility and problem-solving capabilities. State-of-the-art robotic tool usage methods focused on procedurally generating or crowd-sourcing datasets of tools for a task to learn how to grasp and manipulate them for that task. However, these methods assume that only one object is provided and that it is possible, with the correct grasp, to perform the task; they are not capable of identifying, grasping, and using the best object for a task when many are available, especially when the optimal tool is absent. In this work, we propose GeT-USE, a two-step procedure that learns to perform real-robot generalized tool usage by learning first to extend the robot’s embodiment in simulation and then transferring the learned strategies to real-robot visuomotor policies. Our key insight is that by exploring a robot’s embodiment extensions (i.e., building new end-effectors) in simulation, the robot can identify the general tool geometries most beneficial for a task. This learned geometric knowledge can then be distilled to perform generalized tool usage tasks by selecting and using the best available real-world object as tool. On a real robot with 22 degrees of freedom (DOFs), GeT-USE outperforms state-of-the-art methods by 30-60% success rates across three vision-based bimanual mobile manipulation tool-usage tasks.
Bohan Wu, Paul de La Sayette, Li Fei-Fei 0001, Roberto Martin Martin
IROS1
2025 MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
abstract
Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jinsheng Huang, Liang Chen 0024, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan 0016, Haozhe Zhao, Zhihui Guo, Yichi Zhang 0010, Jingyang Yuan, Wei Ju 0001, Luchen Liu, Tianyu Liu 0001, Baobao Chang, Ming Zhang 0004
NAACL (Long Papers)6
2025 SEGA: Shaping Semantic Geometry for Robust Hashing under Noisy Supervision
abstract
This paper studies the problem of learning hash codes from noisy supervision, which is a practical yet challenging task. This problem is important in extensive real-world applications such as image retrieval and cross-modal retrieval. However, most of the existing methods focus on label denoising to address this problem, but ignore the geometric structure of the hash space, which is critical for learning stable hash codes. Towards this end, this paper proposes a novel framework named Semantic Geometry Shaping (SEGA) that explicitly refines the semantic geometry of hash space. Specifically, we first learn dynamic class prototypes as semantic anchors and cluster hash embeddings around these prototypes to keep structural stability. We then leverage both the energy of predicted distributions and structure-based divergence to estimate the uncertainty of instances and calibrate the supervision in a soft manner. Moreover, we introduce structure-aware interpolation to improve the class boundaries. To verify the effectiveness of our design, we give the theoretical analysis for the proposed framework. Experiments on a range of widely-used retrieval datasets justify the superiority of our SEGA over extensive strong baselines under noisy supervision.
Yiyang Gu, Bohan Wu, Qinghua Ran, Rongcheng Tu, Xiao Luo 0001, Zhiping Xiao 0001, Wei Ju 0001, Dacheng Tao, Ming Zhang 0004
NeurIPS2
2025 Frequentist Guarantees of Distributed (Non)-Bayesian Inference
abstract
We establish frequentist properties, i.e., posterior consistency, asymptotic normality, and posterior contraction rates, for the distributed (non-)Bayesian inference problem for a set of agents connected over a network. These results are motivated by the need to analyze large, decentralized datasets, where distributed (non)-Bayesian inference has become a critical research area across multiple fields, including statistics, machine learning, and economics. Our results show that, under appropriate assumptions on the communication graph, distributed (non)-Bayesian inference retains parametric efficiency while enhancing robustness in uncertainty quantification. We also explore the trade-off between statistical efficiency and communication efficiency by examining how the design and size of the communication graph impact the posterior contraction rate. Furthermore, we extend our analysis to time-varying graphs and apply our results to exponential family models, distributed logistic regression, and decentralized detection models.
Bohan Wu, César A. Uribe
J. Mach. Learn. Res.1
2024 Target conversation extraction: Source separation using turn-taking dynamics
abstract
Extracting the speech of participants in a conversation amidst interfering speakers and noise presents a challenging problem. In this paper, we introduce the novel task of target conversation extraction, where the goal is to extract the audio of a target conversation based on the speaker embedding of one of its participants. To accomplish this, we propose leveraging temporal patterns inherent in human conversations, particularly turn-taking dynamics, which uniquely characterize speakers engaged in conversation and distinguish them from interfering speakers and noise. Using neural networks, we show the feasibility of our approach on English and Mandarin conversation datasets. In the presence of interfering speakers, our results show an 8.19 dB improvement in signal-to-noise ratio for 2-speaker conversations and a 7.92 dB improvement for 2-4-speaker conversations. Code, dataset available at https://github.com/chentuochao/Target-Conversation-Extraction.
Tuochao Chen, Bohan Wu, Malek Itani, Sefik Emre Eskimez, Takuya Yoshioka, Shyamnath Gollakota
INTERSPEECH3
2024 Charging non-deterministic mobile nodes in a transfer learning approach
Bohan Wu, Shuqi Tang, Feng Lin 0010
Ad Hoc Networks1
2023 M-EMBER: Tackling Long-Horizon Mobile Manipulation via Factorized Domain Transfer
abstract
In this paper, we propose a novel method to create visuomotor mobile manipulation solutions to long-horizon activities. We propose to leverage the recent advances in robot simulation to train robust visual solutions in simulation that can transfer to the real world. While previous works have shown success applying this procedure to autonomous visual navigation and stationary manipulation, applying it to long-horizon visuomotor mobile manipulation is still an open challenge that demands both perceptual and compositional generalization of multiple skills. In this work, we develop Mobile-EMBER, or M-EMBER, a factorized method that decomposes a long-horizon mobile manipulation activity into a repertoire of primitive visual skills, reinforcement-learns each skill in simulation, and composes these skills to a long-horizon mobile manipulation activity. On a real mobile manipulation robot, we find that M-EMBER completes a long-horizon mobile manipulation activity, cleaning_kitchen, achieving over 50% success rate. This requires successfully planning and executing five factorized, learned visual skills, in sequences of up to 48 skills long.
Bohan Wu, Roberto Martin Martin, Li Fei-Fei 0001
ICRA1
2022 Inattentional Blindness in Augmented Reality Head-Up Display-Assisted Driving
abstract
Augmented reality head-up display (AR HUD) is a new technology in assisted driving, which can add extra information to the driving environment in real-time to help the driver better perceive road situation. AR HUD can enhance driving safety but may also encourage inattentional blindness. Hence, this study aims to examine whether AR HUD-induces inattentional blindness and determine whether workload intensifies their relationship. In experiment 1, 60 participants were randomly assigned to three groups and watched three types of augmented reality (AR)-augmented driving videos, respectively. They were instructed to respond to any critical events, but only their responses to road-crossing pedestrians were recorded. Results show that AR HUD reduces inattentional blindness when pedestrians are augmented but encourages inattentional blindness when pedestrians are not augmented. In experiment 2, 20 participants viewed AR-augmented driving videos of high and low workloads. Pedestrians were not augmented in all videos. Result reveals that a high workload induces more inattentional blindness than low workload. The finding confirms that AR HUD induces inattentional blindness, and a high workload will intensify this relationship. The future design of the AR HUD assisted-driving system should consider the risk of inattentional blindness and come up with corresponding countermeasures.
Yimin Wu, Bohan Wu, Duming Wang, Hongting Li, Zhen Yang 0033
Int. J. Hum. Comput. Interact.4
2021 Greedy Hierarchical Variational Autoencoders for Large-Scale Video Prediction
abstract
A video prediction model that generalizes to diverse scenes would enable intelligent agents such as robots to perform a variety of tasks via planning with the model. However, while existing video prediction models have produced promising results on small datasets, they suffer from severe underfitting when trained on large and diverse datasets. To address this underfitting challenge, we first observe that the ability to train larger video prediction models is often bottlenecked by the memory constraints of GPUs or TPUs. In parallel, deep hierarchical latent variable models can produce higher quality predictions by capturing the multi-level stochasticity of future observations, but end-to-end optimization of such models is notably difficult. Our key insight is that greedy and modular optimization of hierarchical autoencoders can simultaneously address both the memory constraints and the optimization challenges of large-scale video prediction. We introduce Greedy Hierarchical Variational Autoencoders (GHVAEs), a method that learns highfidelity video predictions by greedily training each level of a hierarchical autoencoder. In comparison to state- of-the-art models, GHVAEs provide 17-55% gains in prediction performance on four video datasets, a 35–40% higher success rate on real robot tasks, and can improve performance monotonically by simply adding more modules. Visualization and more details are at https://sites.google.com/view/ghvae.
Bohan Wu, Suraj Nair 0003, Roberto Martin Martin, Li Fei-Fei 0001, Chelsea Finn
CVPR1
2020 SQUIRL: Robust and Efficient Learning from Video Demonstration of Long-Horizon Robotic Manipulation Tasks
abstract
Recent advances in deep reinforcement learning (RL) have demonstrated its potential to learn complex robotic manipulation tasks. However, RL still requires the robot to collect a large amount of real-world experience. To address this problem, recent works have proposed learning from expert demonstrations (LfD), particularly via inverse reinforcement learning (IRL), given its ability to achieve robust performance with only a small number of expert demonstrations. Nevertheless, deploying IRL on real robots is still challenging due to the large number of robot experiences it requires. This paper aims to address this scalability challenge with a robust, sample-efficient, and general meta-IRL algorithm, SQUIRL, that performs a new but related long-horizon task robustly given only a single video demonstration. First, this algorithm bootstraps the learning of a task encoder and a task-conditioned policy using behavioral cloning (BC). It then collects real-robot experiences and bypasses reward learning by directly recovering a Q-function from the combined robot and expert trajectories. Next, this algorithm uses the learned Q-function to re-evaluate all cumulative experiences collected by the robot to improve the policy quickly. In the end, the policy performs more robustly (90%+ success) than BC on new tasks while requiring no experiences at test time. Finally, our real-robot and simulated experiments demonstrate our algorithm's generality across different state spaces, action spaces, and vision-based manipulation tasks, e.g., pick-pour-place and pick-carry-drop.
Bohan Wu, Zhanpeng He, Abhi Gupta, Peter K. Allen
IROS1
2020 Model primitives for hierarchical lifelong reinforcement learning
Bohan Wu, Jayesh K. Gupta, Mykel J. Kochenderfer
Auton. Agents Multi Agent Syst.1
2020 Effect of Warning Graphics Location on Driving Performance: An Eye Movement Study
abstract
With the development of cutting-edge technology in the area of driving performance, driver warning systems based on head up displays (HUD) are considered to have the potential to improve driving safety in the future. The location of HUD warning graphics is a vital component to ensure that drivers obtain information the first time and avoid cognitive tunneling when coming across hazards; however, few studies have critically examined this. The present study investigated the advantages of HUD in presenting warning graphics in comparison with traditional head down display (HDD) in vehicles, and further explored the effect of HUD location based on comprehensive indicators, including behavior performance, eye movement data, and subjective assessment. The results revealed that compared with HDD, presenting warning graphics to drivers on HUD could significantly improve driving performance and eye movement patterns, and HUD was the preferential mode for drivers. Results also demonstrated that presenting HUD warning graphics at a location of 8°below the sight line was associated with the worst results in driving performance, eye movement patterns and subjective assessment. Other locations of HUD presentation were not associated with any significant differences for most indicators. These findings have some reference implications for automobile designers as they construct and implement HUD warning systems.
Zhen Yang 0033, Jinlei Shi, Bohan Wu, Chunyan Kang, Wei Zhang 0348, Hongting Li, Changxu Wu
Int. J. Hum. Comput. Interact.3
2019 Pixel-Attentive Policy Gradient for Multi-Fingered Grasping in Cluttered Scenes
abstract
Recent advances in on-policy reinforcement learning (RL) methods enabled learning agents in virtual environments to master complex tasks with high-dimensional and continuous observation and action spaces. However, leveraging this family of algorithms in multi-fingered robotic grasping remains a challenge due to large sim-to-real fidelity gaps and the high sample complexity of on-policy RL algorithms. This work aims to bridge these gaps by first reinforcement-learning a multi-fingered robotic grasping policy in simulation that operates in the pixel space of the input: a single depth image. Using a mapping from pixel space to Cartesian space according to the depth map, this method transfers to the real world with high fidelity and introduces a novel attention mechanism that substantially improves grasp success rate in cluttered environments. Finally, the direct-generative nature of this method allows learning of multi-fingered grasps that have flexible end-effector positions, orientations and rotations, as well as all degrees of freedom of the hand.
Bohan Wu, Iretiayo Akinola, Peter K. Allen
IROS1
2013 A novel frequency search algorithm to achieve fast locking without phase tracking in ADPLL
abstract
A novel frequency search algorithm is proposed in this paper to achieve fast locking in all digital PLL (ADPLL) with no phase tracking being required. According to phase and frequency error, the normalized tuning word (NTW) is calculated so that the output frequency reaches the desired frequency immediately. As the non-idealities, such as DCO gain estimation error and TDC finite resolution, greatly affect the accuracy of the calculation, the output frequency is continuously measured and frequency error is averaged to minimize those impacts. With 0.13um CMOS process, the proposed ADPLL operates at 2.7 GHz and achieves 0.35 us locking time while consuming 7.47mW.
Bohan Wu, Weixin Gai, Te Han
ISCAS1