EDBT 2026 Demo / reviewers in the wild / expert
Yanhao Li
dblp:221/8689
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelabstractLarge Language Models (LLMs) often produce unnecessarily lengthy reasoning traces, which significantly increase computational cost and latency.Existing approaches typically rely on fixed length penalties, but such penalties are hard to tune and fail to adapt to the evolving reasoning abilities of LLMs, leading to suboptimal trade-offs between accuracy and conciseness.To address this challenge, we propose LEASH (adaptive LEngth penAlty and reward SHaping), a reinforcement learning framework for efficient reasoning in LLMs.We formulate length control as a constrained optimization problem and employ a Lagrangian primal-dual method to dynamically adjust the penalty coefficient.When generations exceed the target length, the penalty is intensified; when they are shorter, it is relaxed.This adaptive mechanism guides models toward producing concise reasoning without sacrificing task performance.Experiments on Deepseek-R1-Distill-Qwen-1.5B and Qwen3-4B-Thinking-2507 show that LEASH reduces the average reasoning length by 60% across diverse tasks-including indistribution mathematical reasoning and outof-distribution domains such as coding and instruction following-while maintaining competitive performance.Our work thus presents a practical and effective paradigm for developing controllable and efficient LLMs that balance reasoning capabilities with computational budgets. Yanhao Li, Jiaran Zhang, Lexiang Tang, Guibo Luo |
ACL (1) | 1 |
| 2025 | FMT: Simulating Clinical Reasoning via Forward-Inference Framework in Medical Dialogue SystemsabstractLarge language models (LLMs) show great potential for supporting diagnostic decision-making in medical dialogue systems. However, most existing approaches mainly focus on reproducing surface-level responses from physicians while neglecting the underlying clinical reasoning process. Consequently, these systems often rely heavily on shallow data patterns-particularly in complex or ambiguous clinical scenarios-which limits their ability to accurately model real-world decision making. Recent efforts have attempted to enhance diagnostic performance by incorporating chain-of-thought (CoT) reasoning to enable knowledge distillation from LLMs. Nonetheless, these methods typically depend on reasoning paths constructed after the correct diagnosis is already known, which restricts both the depth and the authenticity of the reasoning process. To overcome these limitations, we propose FMT, a novel framework that simulates clinical reasoning by generating prior diagnostic chains through a forward-reasoning paradigm. Specifically, FMT employs an iterative Reason-Review-Reflect mechanism to autonomously construct high-quality reasoning chains that more closely mirror real-world diagnostic thought patterns. Experiments conducted on two publicly available datasets, MedDialog and KaMed, demonstrate that our framework substantially outperforms existing baselines across multiple evaluation metrics. We believe this work provides a highly promising direction for building more reliable, interpretable, and clinically grounded medical dialogue systems. Guandong Wang, Yanhao Li, Jie Zhai |
BIBM | 3 |
| 2025 | Unraveling the Mystery: Defending Against Jailbreak Attacks Via Unearthing Real IntentionabstractAs Large Language Models (LLMs) become more advanced, the security risks they pose also increase. Ensuring that LLM behavior aligns with human values, particularly in mitigating jailbreak attacks with elusive and implicit intentions, has become a significant challenge. To address this issue, we propose a jailbreak defense method called Real Intentions Defense (RID), which involves two phases: soft extraction and hard deletion. In the soft extraction phase, LLMs are leveraged to extract unbiased, genuine intentions, while in the hard deletion phase, a greedy gradient-based algorithm is used to remove the least important parts of a sentence, based on the insight that words with smaller gradients have less impact on its meaning. We conduct extensive experiments on Vicuna and Llama2 models using eight state-of-the-art jailbreak attacks and six benchmark datasets. Our results show a significant reduction in both Attack Success Rate (ASR) and Harmful Score of jailbreak attacks, while maintaining overall model performance. Further analysis sheds light on the underlying mechanisms of our approach. Yanhao Li, Hongshen Chen, Zhiwei Ge, Sulong Xu, Guibo Luo |
COLING | 1 |
| 2025 | RETAIN: Reliable Topology Augmentation for both Heterophilic and Homophilic GraphsabstractCurrent graph topology augmentation methods are mostly static and heavily rely on the assumption of homophily, where connected nodes are presumed to share the same labels by default. Due to the complexity of real-world graphs, the underlying assumption is often disrupted, thus performance declines, demonstrating their limited adaptability. Although learnable methods flexibly change augmentation strategies based on data, ignorance of balancing consistency and diversity leads to suboptimal performance. This gap highlights the need for universally applicable graph augmentation strategies that ensure these two aspects, thereby enhancing model robustness. To address these challenges, we propose the RETAIN framework as an adaptive data augmentation method for graphs with different homophily levels. This method can dynamically adjust the graph structure based on a learned conditional distribution with the aid of a graph explainer, thus balancing the consistency and diversity of the augmented data. By framing graph topology modification and model refinement as a joint optimization problem, RETAIN facilitates concurrent learning from augmented data and model optimization. Empirical evaluations across diverse benchmarks on node classification tasks reveal that RETAIN can be effectively combined with other methods in a plug-and-play manner and consistently yields performance improvement across a diverse set of benchmarks for both homophilic and heterophilic graphs. Ziyun Zou, Lian Shen, Yanhao Li, Xiangrong Liu |
ICASSP | 3 |
| 2025 | Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time ComputeabstractRecent advancements in software engineering agents have demonstrated promising capabilities in automating program improvements. However, their reliance on closed-source or resource-intensive models introduces significant deployment challenges in private environments, prompting a critical question: How can personally deployable open-source LLMs (e.g., 32B models running on a single GPU) achieve comparable code reasoning performanceƒ To this end, we propose a unified Test-Time Compute (TTC) scaling framework that leverages increased inference-time computation instead of larger models. Our framework incorporates two complementary strategies: internal TTC and external TTC. Internally, we introduce a development-contextualized trajectory synthesis method leveraging real-world software repositories to bootstrap multi-stage reasoning processes, such as fault localization and patch generation. We further enhance trajectory quality through rejection sampling, rigorously evaluating trajectories along accuracy and complexity. Externally, we propose a novel development-process-based search strategy guided by reward models and execution verification. This approach enables targeted computational allocation at critical development decision points, overcoming limitations of existing "end-point only" verification methods.Evaluations on SWE-bench Verified demonstrate our 32B model achieves a 46% issue resolution rate, surpassing significantly larger models such as DeepSeek R1 671B and OpenAI o1. Additionally, we provide the empirical validation of the test-time scaling phenomenon within SWE agents, revealing that models dynamically allocate more tokens to increasingly challenging problems, effectively enhancing reasoning capabilities. We publicly release all training data, models, and code to facilitate future research.1. In fact, our method has been deployed in Tongyi Lingma, an IDE-based coding assistant developed by Alibaba Cloud, where it helps developers solve real-world programming problems. Yingwei Ma, Yongbin Li 0001, Yihong Dong, Yanhao Li, Rongyu Cao, Jue Chen 0003, Fei Huang 0002, Binhua Li |
ASE | 5 |
| 2025 | PipeQS: Pipeline-Based Adaptive Quantization and Staleness-Aware Distributed GNN Training System
Donghang Wu, Lian Shen, Changzhi Jiang, Yanhao Li, Xiangrong Liu |
ECML/PKDD (2) | 4 |
| 2025 | Optimizing Data Acquisitions in Multi-Robot SystemsabstractWe present ROSfs, a novel user-level file system designed to address critical data query inefficiencies in multi-robot systems (MRS). ROSfs introduces an innovative file organization model where robot data is structured as labeled sub-files, coupled with a time-indexed architecture that enables efficient querying of actively modified data. This design enables real-time cross-robot data acquisition and collaboration capabilities previously unattainable in MRS deployments. Our implementation integrates seamlessly with the Robot Operating System (ROS) and has been extensively evaluated using both physical UAV/UGV platforms and data servers. Experimental results demonstrate that ROSfs achieves a 7x reduction in online data query latency under wireless network conditions compared to conventional ROS storage methods, while simultaneously improving data freshness (Age of Information) by up to 271x. These advancements position ROSfs as a transformative solution for high-performance robotic data management in distributed systems. Yanhao Li, Xuanjun Wen, Guancheng Li, Shu Yin 0001 |
SC | 1 |
| 2025 | An Image Emotion Classification Framework Based on Instruction-Guided Triplets With Descriptive CaptionsabstractABSTRACT Image emotion classification remains a challenging task due to the intrinsic subjectivity of emotional perception and the semantic ambiguity inherent in visual content. Although recent studies have applied language supervision to exploit semantic cues, but most methods rely on fixed templates that exhibit limited emotional relevance, while generating high‐quality affective textual descriptions typically involves substantial manual effort and cost. To overcome these limitations, this paper proposes an emotion classification framework that integrates language supervision with instruction tuning. The proposed approach significantly improves emotion classification performance through three key components: (1) instruction‐guided generation of descriptive emotion captions (DECs), (2) cross‐modal pseudo‐label construction, and (3) adaptive multimodal fusion. We design emotion‐centric instructional prompts to guide large pre‐trained vision‐language models in generating semantically rich DECs, thereby surpassing the constraints of conventional template‐based methods. Instruction tuning is further employed to generate structured emotion pseudo‐labels, forming image‐caption‐pseudo‐label triplets that strengthen cross‐modal alignment. Finally, an adaptive fusion mechanism combined with a multi‐branch loss function is introduced to optimize classification efficacy. Extensive experiments conducted across multiple domain‐specific datasets demonstrate that our method achieves state‐of‐the‐art accuracy. Ablation studies confirm the critical role of multimodal collaboration in enhancing model performance. Furthermore, detailed linguistic analysis shows that DECs achieve high levels of emotional expressiveness, as evidenced by their length, degree of abstraction, and affective distribution, closely approximating the quality of human‐authored descriptions. This work demonstrates fine‐grained emotion description generation without manual annotation and offers a potential solution for multi‐source image emotion recognition. Fuxiao Zhang, Yanhao Li |
IET Image Process. | 4 |
| 2024 | Optimal Design of Electric Bus Network Considering on-Board Photovoltaic Power SupplyabstractThe scale of electric buses is experiencing rapid growth in the progress of electrification for public transport. While in many cities traditional charging mode is still adapted, which is highly dependent on charging piles deployed at fixed charging stations. However, due to the limited charging station areas and insufficient power supply, vehicles would suffer from long detours and queuing time, which would burden the sustainable operation of electric buses. To address these issues, integrating solar panels as on-board photovoltaic equipment to provide en-route power supply for electric buses has become a potential feasible solution to ease such pressures. Considering the impact of time varying light intensity and demand density, a continuous approximation model was constructed on the grid network, considering passengers' time costs, agency costs, and charging facilities costs and a mixed heuristic was designed to solve the model. Take a region in Panjin city, China as an example, the results indicate that: (1) On-board photovoltaic power supply can significantly improve the service level of electric buses without increasing system costs, and alleviate the dependence of electric buses on charging stations; (2) Increasing light intensity and coverage of solar panels have positive implications for reducing the number of charging piles and queuing time at stations. When light intensity exceeds 6.4×1041x, the charging queue time will be reduced by 97.32% with 100% coverage compared to 20% coverage. Chengdong Zhang, Yanhao Li, Yanxi Zhang |
INDIN | 3 |
| 2023 | A Heterogeneous Ranking Contrastive Learning Method for Drug-Target Interaction PredictionabstractRecently, with the in-depth study of biological network structures, methods such as graph neural networks and graph contrastive learning have attracted significant attention and demonstrated notable advantages in DTI prediction. Nevertheless, it remains a challenging task that predicting new DTIs by a small number of known data in heterogeneous biological networks. Graph contrastive learning as an effective method to solute this issue has been proposed recently. Although contrastive learning have shown significant advantages, the objective function heavily relies on unbiased positive and negative samples. Inspired by this issue, a heterogeneous ranking contrastive learning method (HRCL-DTI) for DTI prediction is proposed. Specifically, multiple graph encoders are employed to capture the topological relationships in heterogeneous biological networks. After that, the prediction scores are calculated in the ranking module of HRCL-DTI to select reliable positive and negative samples for the heterogeneous contrastive learning module, aiming to enhance the consistency of DTI representations. Experimental results demonstrate that HRCL-DTI outperforms existing stateof-the-art baselines on multiple datasets and it possesses strong generalization ability and practical effectiveness. Desheng Kong, Maoqiang Xie, Yanhao Li, Yiran Wan, Yalou Huang |
BIBM | 5 |
| 2023 | A Contrario Detection of H.264 Video Double CompressionabstractVideo manipulation detection plays a vital role in modern multimedia forensics. In particular, double compression detection provides significant clues leading to the video edition history and hinting at potential malevolent manipulation. While such an analysis is well-understood on images, the research on this subject remains lacking in videos and existing methods are not yet able to reliably detect double-compressed videos. This work presents a novel method for identifying double compression in H.264 codec videos. Our technique exploits the periodicity of frame residuals caused by fixed Group of Pictures in the initial compression, and employs an a contrario framework to minimize and control false detections. The proposed method can reliably detect double compression in videos. It does not require threshold tuning, thus enabling automatic detection. The code is available at https://github.com/li-yanhao/gop_detection. Yanhao Li, Marina Gardella, Quentin Bammey, Tina Nikoukhah, Jean-Michel Morel, Miguel Colom, Rafael Grompone von Gioi |
ICIP | 1 |
| 2022 | The Impact of JPEG Compression on Prior Image NoiseabstractJPEG compression is widely used to store digital images and extensive studies analysed its impact on the image quality; in particular the quantization noise and artefacts created by JPEG. Nevertheless, there is little work on the impact of JPEG compression on the noise already present in the image. In this paper, we propose a model predicting how the noise power is affected by JPEG compression. This allows for a better understanding the noise traces on the image, which is crucial for image forensic analysis and image restoration. An interactive demo for this article is available at https://ipolcore.ipol.im/demo/clientApp/demo.html?id=77777000136 Marina Gardella, Tina Nikoukhah, Yanhao Li, Quentin Bammey |
ICASSP | 3 |
| 2022 | Video Signal-Dependent Noise Estimation via Inter-Frame PredictionabstractWe propose a block-based signal-dependent noise estimation method on videos, that leverages inter-frame redundancy to separate noise from signal. Block matching is applied to find block pairs between two consecutive frames with similar signal. Then Ponomarenko’s method is extended by sorting pairs by their low-frequency energy and estimating noise in the high frequencies. Experiments on three datasets show that this method improves on the state of the art. Yanhao Li, Marina Gardella, Quentin Bammey, Tina Nikoukhah, Rafael Grompone von Gioi, Miguel Colom, Jean-Michel Morel |
ICIP | 1 |
| 2021 | LiDAR-Based Initial Global Localization Using Two-Dimensional (2D) Submap Projection Image (SPI)abstractInitial global localization is important to mobile robotics in terms of navigation initialization (or re-initialization) and loop closure in SLAM. 3D LiDARs are commonly used for mobile robotics, yet LiDAR-based initial global localization (especially at large scale such as in outdoor environments) is still challenging due to lack of salient features in LiDAR range data. Inspired by visual SLAM oriented initial global localization methods, we propose a method of LiDAR-based initial global localization using 2D submap projection image (SPI). For this, global descriptors from SPIs are extracted for place recognition; pose estimation is realized by feature point matching between the queried SPI and SPIs from a global map database. The proposed initial global localization module runs at 2.4 Hz with precision of 1.2 m and for translation and 1.2° for rotation, which can serve as a suitable initial estimate for subsequent pose estimation refinement via existing mature point cloud registration methods. Yanhao Li, Hao Li 0024 |
ICRA | 1 |
| 2020 | A Collaborative Relative Localization Method for Vehicles using Vision and LiDAR SensorsabstractInter-vehicle relative localization is essential to various multi-vehicle collaborative tasks such as environment perception augmentation, long range path planning and communication-assisted driving applications. We propose a real-time collaborative relative localization framework to estimate inter-vehicle poses via vision and LiDAR sensors. The vision sensor is used for inter-vehicle pose initialization, whereas the LiDAR sensor is used for SLAM and pose refinement. A comparative study between our proposed method and a representative baseline method based on ORB-SLAM2-stereo is performed in the CARLA simulator, which demonstrates advantages of the proposed method in terms of accuracy and robustness in real-time operation. Yanhao Li, Hao Li 0024 |
ICARCV | 1 |