EDBT 2026 Demo / reviewers in the wild / expert
Baoyu Fan
dblp:218/2727
· DBLP profile ↗
20ranked-venue papers
3as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 12 |
| 2026 | T-MSA: Transformer-Driven Multi-Strategy Adaptive Microarchitecture Design Space ExplorationabstractThe design of modern processors ignores the topological relationships among all design parameters, leading to significant simulation costs wasted on invalid designs. Therefore, we propose the T-MSA to address this issue. It is a Transformer-driven multi-strategy adaptive design space exploration scheme. A customized lightweight Transformer (LiteFormer) is devised to model topological relationships among arbitrary design parameters, constructing an implicit interaction graph in the latent space. Secondly, we design a dynamic active learning (DynamicAL) strategy to extract sparse and high-quality initial points via sparse centroid initialization and hybrid sampling. Finally, a triple Pareto frontier acquisition function (TriPFAF) is devised to guide optimization direction based on gains from three types of Pareto frontiers, dynamically balancing exploration and exploitation. We conducted rigorous experiments on two BOOM evaluation platforms, demonstrating that T-MSA efficiently and comprehensively optimizes the performance-power-area (PPA) objective. The designs it identifies achieve significant improvements over state-of-the-art DSE algorithms on Pareto hypervolume (HV). When attaining the same HV value, T-MSA outperforms BOOM-Explorer by 188.24% and 133.33% on two platforms. Fan Yang 0032, Xiaochuan Li 0001, Cong Xu 0001, RenGang Li, Baoyu Fan |
DATE | 8 |
| 2026 | A knowledge-guided hierarchical multi-agent deep reinforcement learning approach for crowd evacuation
Hong Liu 0013, Baoyu Fan, Xiaochuan Li 0001, Wenhao Li 0006 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Xiaoli Gong, Jin Zhang 0003 |
Signal Process. | 1 |
| 2026 | A Fault-Aware Architecture for Reliable Sparse Matrix Multiplication
Yuxuan Qiao, Changxu Liu, Junjie Zuo, Baoyu Fan, Fan Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | Fine-Grained Audio-Visual Event LocalizationabstractAudio-visual event localization (AVEL) aims to recognize events in videos by associating audio-visual information. However, events involved in existing AVEL tasks are usually coarse-grained events. Actually, finer-grained events are sometimes necessary to be distinguished, especially in certain expert-level applications or rich-content-generation studies. However, this is challenging because they are more difficult to detect or distinguish compared with coarse-grained events. To better address this problem, we discuss a new setting of fine-grained AVEL from dataset to method. First, we constructed the first fine-grained audio-visual event dataset, which is called IT-AVE, relying on videos of playing musical instruments, containing 13k video clips and over 52k audio-visual events. All events are labeled from professional music practitioners, and the event categories are all derived from playing techniques, which are fine-grained with little interclass variation. Next, we designed a new fine-grained event localization method, spatial-temporal video event detector (SVED), which focuses on the challenges that fine-grained events are more imperceptible and prone to be disturbed. Finally, we conduct extensive experiments based on the proposed IT-AVE dataset versus fine-grained versions of two existing related datasets, including UnAV-22 derived from UnAV-100 and FineAction-AV derived from FineAction. Experimental results demonstrate the effectiveness of our method. We hope that this work will contribute to the exploration of an integrated understanding of audio-visual videos. Baoyu Fan, Lu Liu 0009, Xiaochuan Li 0001, Jin Zhang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan |
ICCV | 13 |
| 2025 | SSmokeDet: A novel network dedicated to small-scale smoke detection
Li Wang 0040, Xiaochuan Li 0001, Baoyu Fan |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Image Content Generation with Causal ReasoningabstractThe emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning is primarily limited to the domain of language generation, such as in models like GPT-3. In visual modality, there is currently no equivalent research. Considering causal reasoning in visual content generation is significant. This is because visual information contains infinite granularity. Particularly, images can provide more intuitive and specific demonstrations for certain reasoning tasks, especially when compared to coarse-grained text. Hence, we propose a new image generation task called visual question answering with image (VQAI) and establish a dataset of the same name based on the classic Tom and Jerry animated series. Additionally, we develop a new paradigm for image generation to tackle the challenges of this task. Finally, we perform extensive experiments and analyses, including visualizations of the generated content and discussions on the potentials and limitations. The code and data are publicly available under the license of CC BY-NC-SA 4.0 for academic and non-commercial usage at: https://github.com/IEIT-AGI/MIX-Shannon/blob/main/projects/VQAI/lgd_vqai.md. Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li |
AAAI | 2 |
| 2024 | A Distributed Method for Negative Content Spread Minimization on Social Networks
Ruidong Yan, Weili Wu 0001, Baoyu Fan |
AAIM (1) | 4 |
| 2024 | FDNet: Feature Decoupling Framework for Trajectory PredictionabstractTrajectory prediction plays a significant role in autonomous driving, with current challenges primarily focused on capturing complex interactions in traffic scenes. Previous methods usually directly encode non-interactive and interactive information together, and then decode them for trajectory prediction. However, given the complexity inherent property in the trajectory generation process (e.g., the generation of trajectory points are influenced by the interactions among multiple moving agents, as well as the interactions between agents and the static environment), previous approaches fail to precisely capture separate variations of the trajectory generation process. In this paper, we propose a general and plug-and-play feature decoupling framework for trajectory prediction called FDNet, which can learn the interactive and non-interactive factors in the latent space to capture separate variations of the trajectory generation process. At its core, FDNet is comprised of a Non-interactive Feature Extraction Module to extract non-interactive features, and an Interactive Feature Decoupling Module to decouple interactive features. Extensive experiments conducted on Argoverse and nuScenes demonstrate that FDNet significantly improves the performance of existing methods. Yuhang Li 0007, Baoyu Fan, Rongqing Li, Dongchun Ren, Ye Yuan 0001, Guoren Wang |
IROS | 3 |
| 2024 | SyncIntellects: Orchestrating LLM Inference with Progressive Prediction and QoS-Friendly ControlabstractLarge Language Models (LLMs) have shown impressive capabilities, especially in the realm of Human-Machine Chat Systems. Nevertheless, these models entail significant computational expenses, particularly when generating tokens. As a remedy to enhance system throughput and hardware utilization, batch scheduling is commonly adopted. This method involves initiating a batch of inference requests concurrently and then waiting for their completion. A significant challenge encountered with task-batching is the need to group requests with similar response lengths. However, accurately predicting response length proves to be a daunting task, and the inherent variability in response length leads to suboptimal resource utilization.In this paper, we introduce SyncIntellects, a framework designed to orchestrate Large Language Model (LLM) Inference with fine-grained response length prediction and Quality of Service (QoS)-Friendly length control. Specifically, SyncIntellects enhances response length prediction by leveraging embedding information during token generation through a transformer-based model. Subsequently, a dynamic response length controller based on Prompt Engineering techniques is employed to ensure alignment of response lengths without compromising the QoS of the responses. We have implemented SyncIntellects and seamlessly integrated it with a chatbot engine based on the llama2 7B model. We conduct comprehensive experiments on an NVIDIA A100-based testbed, and the results demonstrate a significant reduction in latency by 17.76% on average, along with an increase in throughput by 9.34%. Xue Lin 0006, Peining Yue, Haoran Li 0014, Jin Zhang 0003, Baoyu Fan, Huayou Su, Xiaoli Gong |
IWQoS | 6 |
| 2024 | Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and BaselineabstractExisting video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI. Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li |
NeurIPS | 2 |
| 2024 | MVIndEmo: a dataset for micro video public-induced emotion prediction on social mediaabstractAbstract Distinct from the realm of perceived emotion research, induced emotion pertains to the emotional responses engendered within content consumers. This facet has garnered considerable attention and finds extensive application in the analysis of public social media. However, the advent of micro videos presents unique challenges when attempting to discern the induced emotional patterns exhibited by content consumers, owing to their free-style representation and other factors. Consequently, we have put forth two novel tasks concerning the recognition of public-induced emotion on micro videos: emotion polarity and emotion classification. Additionally, we have introduced a accessible dataset specifically tailored for the analysis of public-induced emotion on micro videos. The data corpus has been meticulously collected from Tiktok, a burgeoning social media platform renowned for its trendsetting content. To construct the dataset, we have selected eight captivating topics that elicit vibrant social discussions. In devising our label generation strategy, we have employed an automated approach characterized by the fusion of multiple expert models. This strategy incorporates a confidence measure method that relies on three distinct models for effectively aggregating user comments. To accommodate adaptable benchmark configurations, we provide both binary classification labels and probability distribution labels. The dataset encompasses a vast collection of 7,153 labeled micro videos. We have undertaken an extensive statistical analysis of the dataset to provide a comprehensive overview composition. It is our earnest aspiration that this dataset will serve as a catalyst for pioneering research avenues in the analysis of emotional patterns and the understanding of multi-modal information. Zhenhua Guo 0003, Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Yaqian Zhao, RenGang Li |
Multim. Syst. | 3 |
| 2024 | Inexactly Matched Referring Expression Comprehension With RationaleabstractReferring Expression Comprehension (REC) is a multimodal comprehension task that aims to locate an object in an image, given a text description. Traditionally, during the existing REC tasks, there has been a basic assumption that the given text expression and the image are usually exactly matched to each other. However, in real-world scenarios, there is uncertainty in how well the image and text match each other exactly. Illegible objects in the image or ambiguous phrases in the text have the potential to significantly degrade the performance of conventional REC tasks. To overcome these limitations, we consider a more practical and comprehensive REC task, where the given image and its referring text expression can be inexactly matched. Our models aim to correct such inexact matching and supply corresponding interpretations. We refer to this task asFurther REC (FREC). This task is divided into three subtasks: 1) correcting the erroneous text expression using visual information, 2) generating the rationale for this input expression, and 3) localizing the proper object based on the corrected expression. We introduce three new datasets for FREC:Further-RefCOCOs,Further-CopsrefandFurther-Talk2Car. These datasets are based on the existing REC datasets, including RefCOCO and Talk2Car. We developed a novel pipeline architecture to execute the three subtasks simultaneously in an end-to-end fashion. Next, we developed an elastic masked language modeling (EMLM) training head to rectify text errors with uncertain lengths. Our experimental results demonstrate the validity of our proposed pipeline. We hope this work sparks more research focused on inexactly matched REC. Xiaochuan Li 0001, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li |
IEEE Trans. Multim. | 2 |
| 2022 | Towards Further Comprehension on Referring Expression with RationaleabstractReferring Expression Comprehension (REC) is one important research branch in visual grounding, where the goal of REC is to localize a relevant object in the image, given an expression in the form of text to exactly describe a specific object. However, existing REC tasks aim at text content filtering and image object locating, which are evaluated based on the precision of the detection boxes. This may lead models to skip the learning process of multimodal comprehension directly and achieve good performance. In this paper, we work on how to enable an artificial agent to understand RE further and propose a more comprehensive task, called Further Comprehension on Referring Expression (FREC). In this task, we mainly focus on three sub-tasks: 1) correcting the erroneous text expression based on visual information; 2) generating the rationale of this input expression; 3) localizing the proper object based on the corrected expression. Accordingly, we make a new dataset named Further-RefCOCOs based on the RefCOCO, RefCOCO+, RefCOCOg benchmark datasets for this new task and make it publicly available. After that, we design a novel end-to-end pipeline to achieve these sub-tasks simultaneously. The experimental results demonstrate the validity of the proposed pipeline. We believe this work will motivate more researchers to explore along with this direction, and promote the development of visual grounding. RenGang Li, Baoyu Fan, Xiaochuan Li 0001, Zhenhua Guo 0003, Yaqian Zhao, Weifeng Gong, Endong Wang |
ACM Multimedia | 2 |
| 2022 | AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityabstractVisual Question Answering (VQA) serves as a proxy for evaluating the scene understanding of an intelligent agent by answering questions about images. Most VQA benchmarks to date are focused on those questions that can be answered through understanding visual content in the scene, such as simple counting, visual attributes, and even a little challenging questions that require extra encyclopedic knowledge. However, humans have a remarkable capacity to reason dynamic interaction on the scene, which is beyond the literal content of an image and has not been investigated so far. In this paper, we propose Agent Interaction Visual Question Answering (AI-VQA), a task investigating deep scene understanding if the agent takes a certain action. For this task, a model not only needs to answer action-related questions but also to locate the objects in which the interaction occurs for guaranteeing it truly comprehends the action. Accordingly, we make a new dataset based on Visual Genome and ATOMIC knowledge graph, including more than 19,000 manually annotated questions, and will make it publicly available. Besides, we also provide an annotation of the reasoning path while developing the answer for each question. Based on the dataset, we further propose a novel method, called ARE, that can comprehend the interaction and explain the reason based on a given event knowledge base. Experimental results show that our proposed method outperforms the baseline by a clear margin. RenGang Li, Cong Xu 0001, Zhenhua Guo 0003, Baoyu Fan, Yaqian Zhao, Weifeng Gong, Endong Wang |
ACM Multimedia | 4 |
| 2021 | Knowledge-Supervised Learning: Knowledge Consensus Constraints for Person Re-IdentificationabstractThe consensus of multiple views on the same data will provide extra regularization, thereby improving accuracy. Based on this idea, we proposed a novel Knowledge-Supervised Learning (KSL) method for person re-identification (Re-ID), which can improve the performance without introducing extra inference cost. Firstly, we introduce isomorphic auxiliary training strategy to conduct basic multiple views that simultaneously train multiple classifier heads of the same network on the same training data. The consensus constraints aim to maximize the agreement among multiple views. To introduce this regular constraint, inspired by knowledge distillation that paired branches can be trained collaboratively through mutual imitation learning. Three novel constraints losses are proposed to distill the knowledge that needs to be transferred across different branches: similarity of predicted classification probability for cosine space constraints, distance of embedding features for euclidean space constraints, hard sample mutual mining for hard sample space constraints. From different perspectives, these losses complement each other. Experiments on four mainstream Re-ID datasets show that a standard model with KSL method trained from scratch outperforms its ImageNet pre-training results by a clear margin. With KSL method, a lightweight model without ImageNet pre-training outperforms most large models. We expect that these discoveries can attract some attention from the current de facto paradigm of "pre-training and fine-tuning" in Re-ID task to the knowledge discovery during model training. Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong, Endong Wang |
ACM Multimedia | 2 |
| 2020 | Dense-Scale Feature Learning in Person Re-identification
Li Wang 0040, Baoyu Fan, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong |
ACCV (6) | 2 |
| 2020 | Contextual Multi-Scale Feature Learning for Person Re-IdentificationabstractRepresenting features at multiple scales is significant for person re-identification (Re-ID). Most existing methods learn the multi-scale features by stacking streams and convolutions without considering the cooperation of multiple scales at a granular level. However, most scales are more discriminative only when they integrate other scales as contextual information. We termed that contextual multi-scale. In this paper, we proposed a novel architecture, namely contextual multi-scale network (CMSNet), for learning common and contextual multi-scale representations simultaneously. The building block of CMSNet obtains contextual multi-scale representations by bidirectionally hierarchical connection groups: the forward hierarchical connection group for stepwise inter-scale information fusion and the backward hierarchical connection group for leap-frogging inter-scale information fusion. Too rich scale features without a selection will confuse the discrimination. Additionally, we introduced a new channel-wise scale selection module to dynamically select scale features for corresponding input image. To the best of our knowledge, CMSNet is the most lightweight model for person Re-ID and it achieves state-of-the-art performance on four commonly used Re-ID datasets, surpassing most large-scale models. Baoyu Fan, Li Wang 0040, Zhenhua Guo 0003, Yaqian Zhao, RenGang Li, Weifeng Gong |
ACM Multimedia | 1 |