Bin Li 0083

dblp:89/6764-83 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-6508-5071ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation
abstract
Accurate segmentation of aortic vascular structures is critical for diagnosing and treating cardiovascular diseases. Traditional Transformer-based models have shown promise in this domain by capturing long-range dependencies between vascular features. However, their reliance on fixed-size rectangular patches often influences the integrity of complex vascular structures, leading to suboptimal segmentation accuracy. To address this challenge, we propose the adaptive Morph-Patch Transformer (MPT), a novel architecture specifically designed for aortic vascular segmentation. Specifically, MPT introduces an adaptive patch partitioning strategy that dynamically generates morphology-aware patches aligned with complex vascular structures. This strategy can preserve semantic integrity of complex vascular structures within individual patches. Moreover, a Semantic Clustering Attention (SCA) method is proposed to dynamically aggregate features from various patches with similar semantic characteristics. This method enhances the model's capability to segment vessels of varying sizes, preserving the integrity of vascular structures. Extensive experiments on three open-source datasets (AVT, AortaSeg24 and TBAD) demonstrate that MPT achieves state-of-the-art performance, with improvements in segmenting intricate vascular structures.
Fuchen Zheng, Adnan Iltaf, Yifei Han, Zhenyu Chen 0001, Yue Du, Bin Li 0083, Tianyong Liu, Shoujun Zhou
AAAI7
2026 Boosting Large Language Models for Mental Manipulation Detection via Data Augmentation and Distillation
abstract
Mental manipulation on social media poses a covert yet serious threat to individuals' psychological well-being and the integrity of online interactions. Detecting such behavior is challenging due to the difficult-to-annotate training data, its highly covert and multi-turn nature, and the lack of real-world datasets. To address these challenges, we propose MentalMAD, a framework that enhances large language models for mental manipulation detection. Our approach consists of three key components: EvoSA, an annotation-free data augmentation method that combines evolutionary operations with speech-act-aware prompting; teacher-model-generated complementary-task supervision; and Complementary-Convergent Distillation, a phase-wise strategy for transferring manipulation-specific knowledge to student models. We then constructed the ReaMent dataset, comprising 5,000 real-world-sourced dialogues. Extensive experiments show that MentalMAD improves accuracy by 14.0%, macro-F1 by 27.3%, and weighted F1 by 15.1% over the strongest baseline. The code and the dataset are publicly available at https://github.com/Yuansheng-Gao/MentalMAD.
Yuansheng Gao, Bin Li 0083, Jixiang Luo, Zonghui Wang, Wenzhi Chen
WWW4
2025 Learning to Unify Audio, Visual and Text for Audio-Enhanced Visual Answer Localization
abstract
The goal of Visual Answer Localization (VAL) is to locate a video segment that answers a given question. Existing methods either focus solely on visual modality or integrate visual and subtitle modalities. However, these methods neglect the audio modality in videos, consequently leading to incomplete input information and poor consistency in the VAL task. In this paper, we propose a unified Audio-Visual-Textual Span Localization (AVTSL) method that incorporates audio modality to augment both visual and textual representations for the VAL task. Specifically, we design a network structure with three interacting predictors, each corresponding to the audio, visual, and text modalities. Each predictor generates predictions based on its respective modality. To maintain consistency across the predicted results, we design an Audio-Visual-Textual Consistency module. This module utilizes a Dynamic Triangular Loss (DTL) function, allowing each modality’s predictor to dynamically learn from the others. This collaborative learning ensures that the model generates consistent and comprehensive answers. Extensive experiments show that our proposed method outperforms several state-of-the-art (SOTA) methods, which demonstrates the effectiveness of the audio modality. Code is available at https://github.com/binbin2xs/AVTSL.
Zhibin Wen, Bin Li 0083
ICME2
2025 Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
Bin Li 0083, Shenxi Liu, Yixuan Weng, Yue Du, Yuhang Tian 0002, Shoujun Zhou
NLPCC (4)1
2025 FocusMorph: A novel multi-scale fusion network for 3D brain MR image registration
Tianyong Liu, Guojia Fan, Chengwu Xu, Bin Li 0083, Shoujun Zhou
Pattern Recognit.6
2025 MambaVesselNet: A Novel Approach to Blood Vessel Segmentation Based on State-Space Models
abstract
Three-dimensional blood vessel segmentation is an important and challenging task that faces two main difficulties: (1) blood vessel structures are small, making them hard to capture by the network, and vessel edges are difficult to segment accurately; (2) false positives are prone to occur due to the presence of artifacts and noise. This paper proposes a novel blood vessel segmentation method called MambaVesselNet. This method is based on a state-space model and employs a selective state-space time series modeling strategy to achieve a larger receptive field. To better capture fine vascular structures and accurately segment edges, this paper introduces an edge enhancement module and a feature selection module. In terms of data preprocessing, nnUNet's preprocessing strategy is adopted to ensure spatial consistency of the input data. Evaluation on three standard vascular segmentation benchmarks shows that MambaVesselNet achieves state-of-the-art performance. Specifically, on cardiovascular and liver vessel datasets, the Dice coefficient is improved by 1.38% and 2.69%, respectively. The contributions of this paper include the proposal of a new module for enhancing blood vessel edge features, the development of a feature selection module with long sequence modeling capability, and the adoption of nnUNet's data preprocessing strategy, setting a new benchmark for blood vessel segmentation technology.
Tianyong Liu, Guojia Fan, Bin Li 0083, Shoujun Zhou, Chengwu Xu, Fuxia Yang
IEEE J. Biomed. Health Informatics4
2025 Small but mighty: enhancing time series forecasting with lightweight LLMs
Haoran Fan, Bin Li 0083, Yixuan Weng, Shoujun Zhou
J. Supercomput.2
2025 Enhancing video temporal grounding with large language model-based data augmentation
Bin Li 0083
J. Supercomput.4
2024 Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural Networks
abstract
Language models' (LMs) proficiency in handling deterministic symbolic reasoning and rule-based tasks remains limited due to their dependency implicit learning on textual data. To endow LMs with genuine rule comprehension abilities, we propose "Neural Comprehension" - a framework that synergistically integrates compiled neural networks (CoNNs) into the standard transformer architecture. CoNNs are neural modules designed to explicitly encode rules through artificially generated attention weights. By incorporating CoNN modules, the Neural Comprehension framework enables LMs to accurately and robustly execute rule-intensive symbolic tasks. Extensive experiments demonstrate the superiority of our approach over existing techniques in terms of length generalization, efficiency, and interpretability for symbolic operations. Furthermore, it can be applied to LMs across different model scales, outperforming tool-calling methods in arithmetic reasoning tasks while maintaining superior inference efficiency. Our work highlights the potential of seamlessly unifying explicit rule learning via CoNNs and implicit pattern learning in LMs, paving the way for true symbolic comprehension capabilities. The code is released at: \url{https://github.com/wengsyx/Neural-Comprehension}.
Yixuan Weng, Minjun Zhu, Bin Li 0083, Shizhu He, Kang Liu 0001, Jun Zhao 0001
ICLR4
2024 Overview of the NLPCC 2024 Shared Task 7: Multi-lingual Medical Instructional Video Question Answering
Bin Li 0083, Yixuan Weng, Qiya Song, Lianhui Liang, Xianwen Min, Shoujun Zhou
NLPCC (5)1
2024 Large Language Models With Holistically Thought Could Be Better Doctors
Yixuan Weng, Bin Li 0083, Minjun Zhu, Bin Sun 0001, Shizhu He, Shengping Liu, Kang Liu 0001, Shutao Li 0001, Jun Zhao 0001
NLPCC (2)2
2024 Distinct but correct: generating diversified and entity-revised medical response
Bin Li 0083, Bin Sun 0001, Shutao Li 0001, Encheng Chen, Hongru Liu, Yixuan Weng, Yongping Bai, Meiling Hu
Sci. China Inf. Sci.1
2024 Towards better Chinese-centric neural machine translation for low-resource languages
Bin Li 0083, Yixuan Weng, Hanjun Deng
Comput. Speech Lang.1
2024 Towards Visual-Prompt Temporal Answer Grounding in Instructional Video
abstract
Temporal answer grounding in instructional video (TAGV) is a new task naturally derived from temporal sentence grounding in general video (TSGV). Given an untrimmed instructional video and a text question, this task aims at locating the frame span from the video that can semantically answer the question, i.e., visual answer. Existing methods tend to solve the TAGV problem with a visual span-based predictor, taking visual information to predict the start and end frames in the video. However, due to the weak correlations between the semantic features of the textual question and visual answer, current methods using the visual span-based predictor do not work well in the TAGV task. In this paper, we propose a visual-prompt text span localization (VPTSL) method, which introduces the timestamped subtitles for a text span-based predictor. Specifically, the visual prompt is a learnable feature embedding, which brings visual knowledge to the pre-trained language model. Meanwhile, the text span-based predictor learns joint semantic representations from the input text question, video subtitles, and visual prompt feature with the pre-trained language model. Thus, the TAGV is reformulated as the task of the visual-prompt subtitle span localization for the visual answer. Extensive experiments on five instructional video datasets, namely MedVidQA, TutorialVQA, VehicleVQA, CrossTalk and Coin, show that the proposed method outperforms several state-of-the-art (SOTA) methods by a large margin in terms of mIoU score, which demonstrates the effectiveness of the proposed visual prompt and text span-based predictor.
Shutao Li 0001, Bin Li 0083, Bin Sun 0001, Yixuan Weng
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Learning To Locate Visual Answer In Video Corpus Using Question
abstract
We introduce a new task, named video corpus visual answer localization (VCVAL), which aims to locate the visual answer in a large collection of untrimmed instructional videos using a natural language question. This task requires a range of skills - the interaction between vision and language, video retrieval, passage comprehension, and visual answer localization. In this paper, we propose a cross-modal contrastive global-span (CCGS) method for the VCVAL, jointly training the video corpus retrieval and visual answer localization subtasks with the global-span matrix. We have reconstructed a dataset named MedVidCQA, on which the VCVAL task is benchmarked. Experimental results show that the proposed method outperforms other competitive methods both in the video corpus retrieval and visual answer localization sub-tasks. Most importantly, we perform detailed analyses on extensive experiments, paving a new path for understanding the instructional videos, which ushers in further research1.
Bin Li 0083, Yixuan Weng, Bin Sun 0001, Shutao Li 0001
ICASSP1
2023 Visual Answer Localization with Cross-Modal Mutual Knowledge Transfer
abstract
The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early methods are based on the interaction modelling between video and text to predict the visual answer by the visual predictor. Later, using the textual predictor with subtitles for the VAL proves to be more precise. However, these existing methods still have cross-modal knowledge deviations from visual frames or textual subtitles. In this paper, we propose a cross-modal mutual knowledge transfer span localization (MutualSL) method to reduce the knowledge deviation. MutualSL has both visual predictor and textual predictor, where we expect the prediction results of these both to be consistent, so as to promote semantic knowledge understanding between cross-modalities. On this basis, we design a one-way dynamic loss function to dynamically adjust the proportion of knowledge transfer. We have conducted extensive experiments on three public datasets for evaluation. The experimental results show that our method outperforms other competitive state-of-the-art (SOTA) methods, demonstrating its effectiveness1.
Yixuan Weng, Bin Li 0083
ICASSP2
2023 Overview of the NLPCC 2023 Shared Task: Chinese Medical Instructional Video Question Answering
Bin Li 0083, Yixuan Weng, Hu Guo, Bin Sun 0001, Shutao Li 0001, Mengyao Qi, Xufei Liu, Yuwei Han, Haiwen Liang, Shuting Gao
NLPCC (3)1
2023 Bilateral personalized dialogue generation with contrastive learning
Bin Li 0083, Hanjun Deng
Soft Comput.1
2022 Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation
Bin Li 0083, Yixuan Weng, Ziyu Ma, Bin Sun 0001, Shutao Li 0001
NLPCC (2)1