EDBT 2026 Demo / reviewers in the wild / expert
Bohan Yu
dblp:250/5778
· DBLP profile ↗
26ranked-venue papers
6as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Based Multi-Range Radiance Separation and 3D Reconstruction via Line-Scan Pseudo-Square IlluminationabstractDecomposing scene radiance into physically meaningful components, including direct reflection, interreflection, and scattering, enables a deeper understanding of scene appearance. In this paper, we propose the first method to perform multi-range radiance component separation using only events captured by an event camera, without requiring any additional frame-based measurements. Our approach scans the scene by swiping line-shaped illumination across it, while exploiting the event camera's high temporal resolution and wide dynamic range to recover both direct and multiple global components corresponding to different light propagation distances. To address the noise inherent in event-integration-based radiance recovery, we present a pixel-wise calibration strategy that leverages the reproducibility of per-pixel noise patterns. We demonstrate that this calibration is highly effective in suppressing noise, enabling stable recovery from subtle signals. Moreover, we show that by detecting the timing at which the scanning line passes each pixel, the same line-scan event data can be exploited for coarse 3D reconstruction. Experimental results on real scenes show that our event-based approach achieves faster and finer component separation, while also enabling coarse depth estimation without the exposure control required by frame-based cameras. Ryuji Hashimoto, Yuta Asano, Shin Ishihara, Bohan Yu, Chu Zhou, Boxin Shi, Imari Sato |
3DV | 4 |
| 2026 | TextShield-R1: Reinforced Reasoning for Tampered Text DetectionabstractThe growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal large language models (MLLMs) demonstrate strong potential in analyzing tampered images and generating interpretations. However, they still struggle with identifying micro-level artifacts, exhibit low accuracy in localizing tampered text regions, and heavily rely on expensive annotations for forgery interpretation. To this end, we introduce TextShield-R1, the first reinforcement learning based MLLM solution for tampered text detection and reasoning. Specifically, our approach introduces Forensic Continual pre-training, an easy-to-hard curriculum that well prepares the MLLM for tampered text detection by harnessing the large-scale cheap data from natural image forensic and OCR tasks. During fine-tuning, we perform Group Relative Policy Optimization with novel reward functions to reduce annotation dependency and improve reasoning capabilities. At inference time, we enhance localization accuracy via OCR Rectification, a method that leverages the MLLM’s strong text recognition abilities to refine its predictions. Furthermore, to support rigorous evaluation, we introduce Text Forensics Reasoning (TFR) benchmark, comprising over 45k real and tampered images across 16 languages, 10 tampering techniques, and diverse domains. Rich reasoning-style annotations are included, allowing for comprehensive assessment. Our TFR benchmark simultaneously addresses seven major limitations of existing benchmarks and enables robust evaluation under cross-style, cross-method, and cross-language conditions. Extensive experiments demonstrate that TextShield-R1 significantly advances the state of the art in interpretable tampered text detection. Chenfan Qu, Yiwu Zhong, Xuekang Zhu, Bohan Yu |
AAAI | 5 |
| 2026 | SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised AttentionabstractThis paper proposes SR-KI, a novel approach for integrating real-time and large-scale structured knowledge bases (KBs) into large language models (LLMs). SR-KI begins by encoding KBs into key-value pairs using a pretrained encoder, and injects them into LLMs' KV cache. Building on this representation, we employ a two-stage training paradigm: first locating a dedicated retrieval layer within the LLM, and then applying an attention-based loss at this layer to explicitly supervise attention toward relevant KB entries. Unlike traditional retrieval-augmented generation methods that rely heavily on the performance of external retrievers and multi-stage pipelines, SR-KI supports end-to-end inference by performing retrieval entirely within the model’s latent space. This design enables efficient compression of injected knowledge and facilitates dynamic knowledge updates. Comprehensive experiments demonstrate that SR-KI enables the integration of up to 40K KBs into a 7B LLM on a single A100 40GB GPU, and achieves strong retrieval performance—maintaining over 98% Recall@10 on the best-performing task and exceeding 88% on average across all tasks. Task performance on question answering and KB ID generation also demonstrates that SR-KI maintains strong performance while achieving up to 99.75% compression of the injected KBs. Bohan Yu |
AAAI | 1 |
| 2026 | Hetero-Designer: Automated Design of Multi-Agent Systems with Heterogeneous LLMsabstractLLM-based Multi-agent systems (MAS) have shown strong capabilities across a wide range of domains.Their success largely hinges on the collaboration topology design, which has emerged as a central research focus in the automated MAS design.However, existing approaches are fundamentally constrained by their reliance on homogeneous LLMs, which significantly limits overall system intelligence.In response to this limitation, we for the first time propose the concept of Automated Design of Heterogeneous-LLMs-based MAS (ADHM).ADHM sheds light on a promising avenue for advancing collective intelligence, which focuses on the automated design of costeffective MAS composed of diverse LLMs and roles to suit various queries.Toward this challenging goal, we propose Hetero-Designer, a novel pipeline that efficiently encodes intricate dependencies among queries, LLMs and roles through a novel Binary-Star Transformer and constructs Hetero-MAS in an autoregressive graph generation process.Extensive experiments demonstrate that Hetero-Designer is: (i) high-performing on various benchmarks, (ii) economical in reducing overhead, (iii) extensible to unseen LLMs and roles. Yuanzhe Zhang, Bohan Yu, Daojian Zeng |
ACL (1) | 3 |
| 2025 | CaDRL: Document-level Relation Extraction via Context-aware Differentiable Rule LearningabstractDocument-level Relation Extraction (DocRE) aims to extract relations from documents. Compared with sentence-level relation extraction, it is necessary to extract long-distance dependencies. Existing methods enhance the output of trained DocRE models either by learning logical rules or by extracting rules from annotated data and then injecting them into the model. However, these approaches can result in suboptimal performance due to incorrect rule set constraints. To mitigate this issue, we propose Context-aware differentiable rule learning or CaDRL for short, a novel differentiable rule-based framework that learns the doc-specific logical rule to avoid generating suboptimal constraints. Specifically, we utilize Transformer-based relation attention to encode document and relation information, thereby learning the contextual information of the relation. We employ a sequence-generated differentiable rule decoder to generate relational probabilistic logic rules at each reasoning step. We also introduce a parameter sharing training mechanism in CaDRL to reconcile the DocRE model and the rule learning module. Extensive experimental results on three DocRE datasets demonstrate that CaDRL outperforms existing rule-based frameworks, significantly improving DocRE performance and making predictions more interpretable and logical. Kunli Zhang, Bohan Yu, Kejun Wu, Aoze Zheng, Xiyang Huang, Chenkang Zhu, Hongying Zan |
COLING | 3 |
| 2025 | EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event CameraabstractSimultaneously acquisition of the surface normal and reflectance parameters is a crucial but challenging technique in the field of computer vision and graphics. It requires capturing multiple high dynamic range (HDR) images in existing methods using frame-based cameras. In this paper, we propose EventPSR, the first work to recover surface normal and reflectance parameters (e.g., metallic and roughness) simultaneously using an event camera. Compared with the existing methods based on photometric stereo or neural radiance fields, EventPSR is a robust and efficient approach that works consistently with different materials. Thanks to the extremely high temporal resolution and high dynamic range coverage of event cameras, EventPSR can recover accurate surface normal and reflectance of objects with various materials in 10 seconds. Extensive experiments on both synthetic data and real objects show that compared with existing methods using more than 100 HDR images, EventPSR recovers comparable surface normal and reflectance parameters with only about 30% of the data rate. Bohan Yu, Jin Han 0001, Boxin Shi, Imari Sato |
CVPR | 1 |
| 2025 | Active Hyperspectral Imaging Using an Event CameraabstractHyperspectral imaging plays a critical role in numerous scientific and industrial fields. Conventional hyperspectral imaging systems often struggle with the trade-off between capture speed, spectral resolution, and bandwidth, particularly in dynamic environments. In this work, we present a novel event-based active hyperspectral imaging system designed for real-time capture with low bandwidth in dynamic scenes. By combining an event camera with a dynamic illumination strategy, our system achieves unprecedented temporal resolution while maintaining high spectral fidelity, all at a fraction of the bandwidth requirements of traditional systems. Unlike basis-based methods that sacrifice spectral resolution for efficiency, our approach enables continuous spectral sampling through an innovative "sweeping rainbow" illumination pattern synchronized with a rotating mirror array. The key insight is leveraging the sparse, asynchronous nature of event cameras to encode spectral variations as temporal contrasts, effectively transforming the spectral reconstruction problem into a series of geometric constraints. Extensive evaluations of both synthetic and real data demonstrate that our system outperforms state-of-the-art methods in temporal resolution while maintaining competitive spectral reconstruction quality. Bohan Yu, Jinxiu Liang, Zhuofeng Wang, Bin Fan 0002, Art Subpa-Asa, Boxin Shi, Imari Sato |
CVPR | 1 |
| 2025 | TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question AnsweringabstractLLMs have shown impressive progress in natural language processing.However, they still face significant challenges in TableQA, where real-world complexities such as diverse table structures, multilingual data, and domain-specific reasoning are crucial.Existing TableQA benchmarks are often limited by their focus on simple flat tables and suffer from data leakage.Furthermore, most benchmarks are monolingual and fail to capture the crosslingual and cross-domain variability in practical applications.To address these limitations, we introduce TableEval, a new benchmark designed to evaluate LLMs on realistic TableQA tasks.Specifically, TableEval includes tables with various structures (such as concise, hierarchical, and nested tables) collected from four domains (including government, finance, academia, and industry reports).Additionally, TableEval features cross-lingual scenarios with tables in Simplified Chinese, Traditional Chinese, and English.To reduce potential data leakage, we curate data from recent real-world documents.Considering that existing TableQA metrics fail to capture semantic accuracy, we further propose SEAT, a new evaluation framework that assesses the alignment between model responses and reference answers at the sub-question level.Experimental results have shown that SEAT achieves high agreement with human judgment.Extensive experiments on TableEval reveal critical gaps in the ability of state-of-the-art LLMs to handle these complex, real-world TableQA tasks, offering insights for future improvements.We make our dataset available here: Junnan Zhu, Bohan Yu |
EMNLP | 3 |
| 2025 | EventUPS: Uncalibrated Photometric Stereo Using an Event Camera
Jinxiu Liang, Bohan Yu, Haotian Zhuang, Jieji Ren, Peiqi Duan 0002, Boxin Shi |
ICCV | 2 |
| 2025 | Development of an Efficient Stiffness Modulation Mechanism in Fish-like Robots for Enhanced Swimming PerformanceabstractDrawing inspiration from the ability of fish to maintain efficient swimming over a wide range of speeds by tuning the stiffness of their tails, researchers have explored stiffness adjustment mechanisms in fish-like robots. Typically, existing mechanisms require extra actuators or power sources only for tuning stiffness, resulting in additional energy consumption and more complex structures. To address this, our study introduces an innovative fishtail featuring an online stiffness modulation mechanism that does not require additional actuators or power sources solely for stiffness adjustment. Through model-based simulations and experimental testing, we evaluated the effectiveness of the proposed method. The results demonstrate that the designed mechanism enables efficient swimming across a broader frequency range (0–4 Hz) compared to most servo-actuated platforms with adjustable stiffness reported in existing studies. The robot achieves a maximum average speed of 1.4 BL/s and a minimum cost of transport of 9.5 J/(m•kg). Xu Chao, Bohan Yu, David Navarro-Alarcon, Xing Jian Jing |
IROS | 2 |
| 2025 | Knowledge-Enhanced and Event-Rule Guided Framework for Fine-Grained Argument Mining in Chinese Essays
Bohan Yu, Aoze Zheng, Kunli Zhang, Hongying Zan |
NLPCC (4) | 4 |
| 2025 | Comprehensive Argument Mining for Chinese Argumentative Essays Using Large Language Models
Bohan Yu, Aoze Zheng, Hongying Zan, Kunli Zhang |
NLPCC (4) | 2 |
| 2025 | Logical Rule-Constrained Large Language Models for Document-Level Relation Extraction
Kunli Zhang, Bohan Yu, Hongying Zan |
NLPCC (1) | 3 |
| 2024 | Latency Correction for Event-Guided Deblurring and Frame InterpolationabstractEvent cameras, with their high temporal resolution, dynamic range, and low power consumption, are particu-larly good at time-sensitive applications like deblurring and frame interpolation. However, their performance is hindered by latency variability, especially under low-light conditions and with fast-moving objects. This paper addresses the challenge of latency in event cameras - the temporal discrepancy between the actual occurrence of changes in the corresponding timestamp assigned by the sensor. Focusing on event-guided deblurring and frame interpolation tasks, we propose a latency correction method based on a parameterized latency model. To enable data-driven learning, we develop an event-based temporal fidelity to describe the sharpness of latent images reconstructed from events and the corresponding blurry images, and reformulate the event-based double integral model differentiable to latency. The proposed method is validated using synthetic and real-world datasets, demonstrating the benefits of latency correction for deblurring and interpolation across different lighting conditions. Yixin Yang 0008, Jinxiu Liang, Bohan Yu, Jimmy S. J. Ren, Boxin Shi |
CVPR | 3 |
| 2024 | EventPS: Real-Time Photometric Stereo Using an Event CameraabstractPhotometric stereo is a well-established technique to es-timate the surface normal of an object. However, the re-quirement of capturing multiple high dynamic range images under different illumination conditions limits the speed and real-time applications. This paper introduces EventPS, a novel approach to real-time photometric stereo using an event camera. Capitalizing on the exceptional temporal resolution, dynamic range, and low bandwidth character-istics of event cameras, EventPS estimates surface nor-mal only from the radiance changes, significantly enhancing data efficiency. EventPS seamlessly integrates with both optimization-based and deep-learning-based photo-metric stereo techniques to offer a robust solution for non-Lambertian surfaces. Extensive experiments validate the effectiveness and efficiency of EventPS compared to frame-based counterparts. Our algorithm runs at over 30 fps in real-world scenarios, unleashing the potential of EventPS in time-sensitive and high-speed downstream applications.11Code available: https://codeberg.org/ybh1998/EventPS Bohan Yu, Jieji Ren, Jin Han 0001, Feishi Wang, Jinxiu Liang, Boxin Shi |
CVPR | 1 |
| 2024 | Towards an Interpretable Representation of Speaker Identity via Perceptual Voice QualitiesabstractUnlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at high-level demographic information, such as gender or age. In this paper, we propose a possible interpretable representation of speaker identity based on perceptual voice qualities (PQs). By adding gendered PQs to the pathology-focused Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) protocol, our PQ-based approach provides a perceptual latent space of the character of adult voices that is an intermediary of abstraction between high-level demographics and low-level acoustic, physical, or learned representations. Contrary to prior belief, we demonstrate that these PQs are hearable by ensembles of non-experts, and further demonstrate that the information encoded in a PQ-based representation is predictable by various speech representations. Robert Netzorg, Bohan Yu, Andrea Guzman, Peter Wu, Luna McNulty, Gopala Krishna Anumanchipalli |
ICASSP | 2 |
| 2024 | Multimodal Segmentation for Vocal Tract Modeling
Bohan Yu, Peter Wu, Tejas S. Prabhune, Gopala Krishna Anumanchipalli |
INTERSPEECH | 2 |
| 2024 | Towards EMG-to-Speech with Necklace Form Factor
Peter Wu, Ryan Kaveh, Raghav Nautiyal, Christine Zhang, Albert Guo, Anvitha Kachinthaya, Tavish Mishra, Bohan Yu, Alan W. Black, Rikky Muller, Gopala Krishna Anumanchipalli |
INTERSPEECH | 8 |
| 2024 | Fast, High-Quality and Parameter-Efficient Articulatory Synthesis Using Differentiable DSPabstractArticulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal processing (DDSP) is a parameter-efficient framework for audio synthesis. Therefore, integrating low-dimensional EMA features with DDSP can significantly enhance the computational efficiency of speech synthesis. In this paper, we propose a fast, high-quality, and parameter-efficient DDSP articulatory vocoder that can synthesize speech from EMA, F0, and loudness. We incorporate several techniques to solve the harmonics / noise imbalance problem, and add a multiresolution adversarial loss for better synthesis quality. Our model achieves a transcription word error rate (WER) of 6.67% and a mean opinion score (MOS) of 3.74, with an improvement of 1.63% and 0.16 compared to the state-of-the-art (SOTA) baseline. Our DDSP vocoder is 4.9 x faster than the baseline on CPU during inference, and can generate speech of comparable quality with only 0.4 M parameters, in contrast to the 9 M parameters required by the SOTA. Yisi Liu, Bohan Yu, Drake Lin, Peter Wu, Cheol Jun Cho, Gopala Krishna Anumanchipalli |
SLT | 2 |
| 2023 | ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric StereoabstractIllumination planning in photometric stereo aims to find a balance between surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illumination, and the choice of photometric stereo backbones, which are too complex to be handled by existing methods based on handcrafted illumination planning rules. This paper proposes a learning-based illumination planning method that jointly considers these factors via integrating a neural network and a generalized image formation model. As it is impractical to supervise illumination planning due to the enormous search space for ground truth light configurations, we formulate illumination planning using reinforcement learning, which explores the light space in a photometric stereo-aware and reward-driven manner. Experiments on synthetic and real-world datasets demonstrate that photometric stereo under the 20-light configurations from our method is comparable to, or even surpasses that of using lights from all available directions. Jun Hoong Chan, Bohan Yu, Heng Guo 0003, Jieji Ren, Zongqing Lu 0002, Boxin Shi |
ICCV | 2 |
| 2023 | MILO: Multi-Bounce Inverse Rendering for Indoor Scene With Light-Emitting ObjectsabstractRecently, many advances in inverse rendering are achieved by high-dimensional lighting representations and differentiable rendering. However, multi-bounce lighting effects can hardly be handled correctly in scene editing using high-dimensional lighting representations, and light source model deviation and ambiguities exist in differentiable rendering methods. These problems limit the applications of inverse rendering. In this paper, we present a multi-bounce inverse rendering method based on Monte Carlo path tracing, to enable correct complex multi-bounce lighting effects rendering in scene editing. We propose a novel light source model that is more suitable for light source editing in indoor scenes, and design a specific neural network with corresponding disambiguation constraints to alleviate ambiguities during the inverse rendering. We evaluate our method on both synthetic and real indoor scenes through virtual object insertion, material editing, relighting tasks, and so on. The results demonstrate that our method achieves better photo-realistic quality. Bohan Yu, Xuanning Cui, Siyan Dong, Baoquan Chen, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Cascaded SE-ResUnet for segmentation of thoracic organs at risk
Zheng Cao 0005, Bohan Yu, Biwen Lei, Haochao Ying, Xiao Zhang 0015, Danny Ziyi Chen, Jian Wu 0001 |
Neurocomputing | 2 |
| 2021 | WiFi-Sleep: Sleep Stage Monitoring Using Commodity Wi-Fi DevicesabstractSleep monitoring is essential to people's health and wellbeing, which can also assist in the diagnosis and treatment of sleep disorder. Compared with contact-based solutions, contactless sleep monitoring does not attach any device to the human body; hence, it has attracted increasing attention in recent years. Inspired by the recent advances in Wi-Fi-based sensing, this article proposes a low-cost and nonintrusive sleep monitoring system using commodity Wi-Fi devices, namely, WiFi-Sleep. We leverage the fine-grained channel state information from multiple antennas and propose advanced fusion and signal processing methods to extract accurate respiration and body movement information. We introduce a deep learning method combined with clinical sleep medicine prior knowledge to achieve four-stage sleep monitoring with limited data sources (i.e., only respiration and body movement information). We benchmark the performance of WiFi-Sleep with polysomnography, the gold reference standard. Results show that WiFi-Sleep achieves an accuracy of 81.8%, which is comparable to the state-of-the-art sleep stage monitoring using expensive radar devices. Bohan Yu, Kai Niu 0003, Youwei Zeng, Tao Gu 0001, Leye Wang, Cuntai Guan, Daqing Zhang 0001 |
IEEE Internet Things J. | 1 |
| 2020 | Doctor Imitator: A Graph-Based Bone Age Assessment Framework Using Hand Radiographs
Jintai Chen, Bohan Yu, Biwen Lei, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 2 |
| 2020 | Teaching Platform for Network Communication and Protocols Using a Micro: bit Based Wheeled RobotabstractIn this study, we presented a lightweight inverted curriculum\cite1 for teaching the essential details of network communication and protocols to undergraduate students major in computer science. Students are instructed to construct a wireless communication and control system connecting a computer to a wheeled robot using the Micro:bit platform. This platform consists of a micro-controller loaded with a Python interpreter and an additional extension board integrated with motors and sensors. In this study, we describe how the students were instructed to build the system step by step, from establishing a wired connection to implementing a TCP server on the PC-side for wireless control. Students can learn these knowledge through practice, which improves classroom engagement as a consequence. Zizhang Luo, Bohan Yu |
SIGCSE | 3 |
| 2019 | LSRC: A Long-Short Range Context-Fusing Framework for Automatic 3D Vertebra Localization
Jintai Chen, Ruoqian Guo, Bohan Yu, Tingting Chen 0002, Wenzhe Wang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 4 |