EDBT 2026 Demo / reviewers in the wild / expert
Baoqun Yin
dblp:96/3589
· DBLP profile ↗
44ranked-venue papers
3as first author
32since 2021 · last 2026
0000-0002-5225-4729ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud RegistrationabstractTypical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect correspondences. Moreover, without dedicated designs, it remains difficult to effectively select informative and correlated representations across modalities, thereby limiting the robustness and accuracy of registration. To address these challenges, we propose a novel cross-modal registration framework composed of two key modules: the Iterative Agents Selection (IAS) module and the Reliable Agents Interaction (RAI) module. IAS enhances structural feature awareness with phase maps and employs reinforcement learning principles to efficiently select reliable agents. RAI then leverages these selected agents to guide cross-modal interactions, effectively reducing mismatches and improving overall robustness. Extensive experiments on the RGB-D Scenes v2 and 7-Scenes benchmarks demonstrate that our method consistently achieves state-of-the-art performance. Zhixin Cheng, Xiaotian Yin, Jiacheng Deng 0002, Bohao Liao, Baoqun Yin, Tianzhu Zhang 0001 |
AAAI | 7 |
| 2026 | GCL: Group-shared continual learning fine-tuning for sparse LLMs
Baoqun Yin |
Neurocomputing | 2 |
| 2026 | Multi-scale pyramid fusion with overlap density attention module for crowd counting
Avinash Rohra, Baoqun Yin, Aakash Kumar, Ajeet Kumar Bhatia, Hazrat Bilal, Munawar Ali |
Neural Networks | 2 |
| 2026 | Multi-scale feature fusion with cross-view head re-identification module for crowd-counting
Avinash Rohra, Baoqun Yin, Aakash Kumar, Ajeet Kumar Bhatia, Izis Kanjarawy |
Pattern Anal. Appl. | 2 |
| 2026 | GLASS: Geometry-Aware Local Alignment and Structure Synchronization Network for 2D-3D RegistrationabstractImage-to-point cloud registration methods typically follow a coarse-to-fine pipeline, extracting patch-level correspondences and refining them into dense pixel-to-point matches. However, in scenes with repetitive patterns, images often lack sufficient 3D structural cues and alignment with point clouds, leading to incorrect matches. Moreover, prior methods usually overlook structural consistency, limiting the full exploitation of correspondences. To address these issues, we propose two novel modules: the Local Geometry Enhancement (LGE) module and the Graph Distribution Consistency (GDC) module. LGE enhances both image and point cloud features with normal vectors, injecting geometric structure into image features to reduce mismatches. GDC constructs a graph from matched points to update features and explicitly constrain similarity distributions. Extensive experiments and ablations on two benchmarks, RGB-D Scenes v2 and 7-Scenes, demonstrate that our approach achieves state-of-the-art performance in image-to-point cloud registration. Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Bohao Liao, Li Liu 0067, Xiaotian Yin, Baoqun Yin, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain AlignmentabstractThe method for image-to-point cloud registration typically determines the rigid transformation using a coarse-to-fine pipeline. However, directly and uniformly matching image patches with point cloud patches may lead to focusing on incorrect noise patches during matching while ignoring key ones. Moreover, due to the significant differences between image and point cloud modalities, it may be challenging to bridge the domain gap without specific improvements in design. To address the above issues, we innovatively propose the Uncertainty-aware Hierarchical Matching Module (UHMM) and the Adversarial Modal Alignment Module (AMAM). Within the UHMM, we model the uncertainty of critical information in image patches and facilitate multi-level fusion interactions between image and point cloud features. In the AMAM, we design an adversarial approach to reduce the domain gap between image and point cloud. Extensive experiments and ablation studies on RGB-D Scene V2 and 7-Scenes benchmarks demonstrate the superiority of our method, making it a state-of-the-art approach for image-to-point cloud registration tasks. Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Baoqun Yin, Tianzhu Zhang 0001 |
AAAI | 4 |
| 2025 | Towards Precise Scaling Laws for Video Diffusion TransformersabstractAchieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in language models to predict performance, their existence and accurate derivation in visual generation models remain underexplored. In this paper, we systematically analyze scaling laws for video diffusion transformers and confirm their presence. Moreover, we discover that, unlike language models, video diffusion models are more sensitive to learning rate and batch size—two hyperparameters often not precisely modeled. To address this, we propose a new scaling law that predicts optimal hyperparameters for any model size and compute budget. Under these optimal settings, we achieve comparable performance and reduce inference costs by 40.1% compared to conventional scaling methods, within a compute budget of 1e10 TFlops. Furthermore, we establish a more generalized and precise relationship among validation loss, any model size, and compute budget. This enables performance prediction for non-optimal model sizes, which may also be appealed under practical inference cost constraints, achieving a better trade-off. Yuanyang Yin, Mingwu Zheng, Jiarong Ou, Victor Shea-Jay Huang, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Baoqun Yin, Wentao Zhang 0001, Kun Gai |
CVPR | 12 |
| 2025 | Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual KnowledgeabstractDoes seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the large language model (LLM). Specifically, the VE’s representation of visual information may not fully align with LLM’s cognitive framework, leading to a mismatch where visual features exceed the language model’s interpretive range. To address this, we investigate how variations in VE representations influence LVLM comprehension, especially when the LLM faces VE-Unknown data—images whose ambiguous visual representations challenge the VE’s interpretive precision. Accordingly, we construct a multi-granularity landmark dataset and systematically examine the impact of VE-Known and VE-Unknown data on interpretive abilities. Our results show that VE-Unknown data limits LVLM’s capacity for accurate understanding, while VE-Known data, rich in distinctive features, helps reduce cognitive misalignment. Building on these insights, we propose Entity-Enhanced Cognitive Alignment (EECA), a method that employs multi-granularity supervision to generate visually enriched, well-aligned tokens that not only integrate within the embedding space but also align with the LLM’s cognitive framework. This alignment markedly enhances LVLM performance in landmark recognition. Our findings underscore the challenges posed by VE-Unknown data and highlight the essential role of cognitive alignment in advancing multimodal systems. Yuanyang Yin, Victor Shea-Jay Huang, Weipeng Chen, Baoqun Yin, Zenan Zhou |
CVPR | 8 |
| 2025 | CA-I2P: Channel-Adaptive Registration Network with Global Optimal SelectionabstractDetection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching results, ultimately impairing registration accuracy. Furthermore, similar structures in the scene could lead to redundant correspondences in cross-modal matching. To address these issues, we propose Channel Adaptive Adjustment Module (CAA) and Global Optimal Selection Module (GOS). CAA enhances intra-modal features and suppresses cross-modal sensitivity, while GOS replaces local selection with global optimization. Experiments on RGB-D Scenes V2 and 7-Scenes demonstrate the superiority of our method, achieving state-of-the-art performance in image-to-point cloud registration. Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Xiaotian Yin, Bohao Liao, Baoqun Yin, Wenfei Yang, Tianzhu Zhang 0001 |
ICCV | 6 |
| 2025 | CBQ: Cross-Block Quantization for Large Language ModelsabstractPost-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from significant performance degradation at ultra-low bits precision. To dissect this issue, we conducted an in-depth analysis of quantization errors specific to LLMs and surprisingly discovered that, unlike traditional sources of quantization errors, the growing number of model parameters, combined with the reduction in quantization bits, intensifies inter-layer and intra-layer dependencies, which severely impact quantization accuracy. This finding highlights a critical challenge in quantizing LLMs. To address this, we propose CBQ, a cross-block reconstruction-based PTQ method for LLMs. CBQ leverages a cross-block dependency to establish long-range dependencies across multiple blocks and integrates an adaptive LoRA-Rounding technique to manage intra-layer dependencies. To further enhance performance, CBQ incorporates a coarse-to-fine pre-processing mechanism for processing weights and activations. Extensive experiments show that CBQ achieves superior low-bit quantization (W4A4, W4A8, W2A16) and outperforms existing state-of-the-art methods across various LLMs and datasets. Notably, CBQ only takes 4.3 hours to quantize a weight-only quantization of a 4-bit LLAMA1-65B model, achieving a commendable trade off between performance and efficiency. Xiaoyu Liu 0006, Zhijun Tu, Wei Li 0002, Jie Hu 0021, Hanting Chen, Yehui Tang 0001, Zhiwei Xiong, Baoqun Yin, Yunhe Wang 0001 |
ICLR | 10 |
| 2025 | DCFT: Dependency-aware continual learning fine-tuning for sparse LLMs
Baoqun Yin |
Neurocomputing | 3 |
| 2025 | MSFFNet: multi-scale feature fusion network with semantic optimization for crowd counting
Avinash Rohra, Baoqun Yin, Hazrat Bilal, Aakash Kumar, Munawar Ali |
Pattern Anal. Appl. | 2 |
| 2025 | End-to-end model compression via pruning and knowledge distillation for lightweight image super resolution
Avinash Rohra, Baoqun Yin |
Pattern Anal. Appl. | 4 |
| 2024 | Cross-Level Feature Relocation: Mitigating Information Loss in Cross-Layer Feature Fusion for Crowd Counting
Yuanyang Yin, Baoqun Yin |
ACML | 2 |
| 2024 | PQ-SAM: Post-training Quantization for Segment Anything Model
Xiaoyu Liu 0006, Yuanyuan Xi, Wei Li 0002, Zhijun Tu, Jie Hu 0021, Hanting Chen, Baoqun Yin, Zhiwei Xiong |
ECCV (10) | 9 |
| 2024 | Online Fault Diagnosis of Industrial Robot Using IoRT and Hybrid Deep Learning Techniques: An Experimental ApproachabstractThe Internet of Robotic Things (IoRT) is growing rapidly with new applications. Co-operatory robotics enables the sharing of information, autonomy, and fail-safe interaction with environment, humans, and other robots. They can also self-maintain, self-aware, and self-heal. To provide reliable and robust online monitoring of the industrial manipulator joint status, this article proposes a new IoRT architecture based on transfer learning (TL) techniques to detect manipulator fault. Robotic manipulator joint status are detected with high accuracy using a hybrid 1-D multichannel convolutional neural network (1D-MCNN), including matrix kernels and recurrent neural network (MCNN-RNN) technique. Moreover, a timestamp mapping method addresses the challenges associated with inconsistencies in sensor data timestamps. Existing data-driven methods struggle with the diverse operating conditions of industrial robots, where load and speed constantly fluctuate. To address this limitation, we propose a novel TL-based MCNN-RNN approach for joint fault diagnosis under varying work conditions. This method leverages the adaptability of TL while incorporating the inherent relations between different failure modes, enhancing the TL process. To demonstrate the performance of the suggested IoRT topology, various experimental scenarios are performed with data acquisition on six degree-of-freedom (DOF) UR16e (universal robot) manipulator. Based on the results, the proposed IoRT architecture can effectively visualize the joint fault status of the manipulator. As a result, TL architecture combined with MCNN-RNN provides an excellent accuracy of 99.03%in detecting faults on manipulator joints, which is significantly higher than traditional convolutional neural network (CNN), deep belief network (DBN), domain adversarial neural network (DANN), and conditional domain-adversarial network (CDAN). Hazrat Bilal, Mohammad S. Obaidat, Muhammad Shamrooz Aslam, Jing Zhang 0015, Baoqun Yin, Khalid Mahmood 0002 |
IEEE Internet Things J. | 5 |
| 2024 | Advanced efficient strategy for detection of dark objects based on spiking network with multi-box detection
Munawar Ali, Baoqun Yin, Hazrat Bilal, Aakash Kumar, Ali Muhammad Shaikh, Avinash Rohra |
Multim. Tools Appl. | 2 |
| 2024 | A Semi-supervised crowd counting method based on patch crowds statistics
Sifan Peng, Baoqun Yin, Yinfeng Xia |
Pattern Anal. Appl. | 2 |
| 2023 | HTNet: A Hybrid Model Boosted by Triple Self-attention for Crowd Counting
Baoqun Yin |
PRCV (12) | 2 |
| 2023 | Consecutive layer collaborative filter similarity for differentiable neural network pruning
Xuan Zu, Baoqun Yin |
Neurocomputing | 3 |
| 2023 | Exploring density rectification and domain adaption method for crowd counting
Sifan Peng, Baoqun Yin |
Neural Comput. Appl. | 2 |
| 2022 | Jointly attention network for crowd counting
Yuqiang He, Yinfeng Xia, Baoqun Yin |
Neurocomputing | 4 |
| 2022 | Weight-Dependent Gates for Network PruningabstractIn this paper, a simple yet effective network pruning framework is proposed to simultaneously address the problems of pruning indicator, pruning ratio, and efficiency constraint. This paper argues that the pruning decision should depend on the convolutional weights, and thus proposes novel weight-dependent gates (W-Gates) to learn the information from filter weights and obtain binary gates to prune or keep the filters automatically. To prune the network under efficiency constraints, a switchable Efficiency Module is constructed to predict the hardware latency or FLOPs of candidate pruned networks. Combined with the proposed Efficiency Module, W-Gates can perform filter pruning in an efficiency-aware manner and achieve a compact network with a better accuracy-efficiency trade-off. We have demonstrated the effectiveness of the proposed method on ResNet34, ResNet50, and MobileNet V2, respectively achieving up to 1.33/1.28/1.1 higher Top-1 accuracy with lower hardware latency on ImageNet. Compared with state-of-the-art methods, W-Gates also achieves superior performance. Zechun Liu, Weiqun Wu, Xiangyu Zhang 0005, Chi Zhang 0026, Baoqun Yin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Scale-aware and Anti-interference Convolutional Network for Crowd CountingabstractCrowd counting is a challenging vision task which aims to estimate the count and distribution of the crowd in a single image accurately. To this end, we propose a novel end - to-end trainable architecture called Scale-aware and Anti-interference Convolutional Network (SACN) to learn a mapping from the input image to the corresponding crowd density map, which concentrates on dealing with the scale variation and background interference of input images for the crowd counting problem. In specific, aiming to cope with the scale variation, we propose Scale-aware Feature Extraction Module (SFEM) to extract multiscale feature maps and learn corresponding pixel-level weight maps for assigning appropriate weights to features of different scales to generate scale-aware features. Futhermore, Regression and Classification Double-head Module (RCDM) is placed at the end of the network to resist the background interference, where the regression head is designed to perform the density map regression and the classification head provides the density map regression task with attention masks related to the foreground and background. In addition, we supervise the intermediate information to help optimize the network and alleviate the gradient vanishing phenomenon. Extensive experiments on three challenging datasets including ShanghaiTech, UCF-QNRF and JHU-CROWD++ were conducted to demonstrate the superiority of the proposed approach compared with the state-of-the-art. Xiaoliang Hao, Yinfeng Xia, Sifan Peng, Baoqun Yin |
IJCNN | 5 |
| 2021 | Boosting Mobile CNN Inference through Semantic MemoryabstractHuman brains are known to be capable of speeding up visual recognition of repeatedly presented objects through faster memory encoding and accessing procedures on activated neurons. For the first time, we borrow and distill such a capability into a semantic memory design, namely SMTM, to improve on-device CNN inference. SMTM employs a hierarchical memory architecture to leverage the long-tail distribution of objects of interest, and further incorporates several novel techniques to put it into effects: (1) it encodes high-dimensional feature maps into low-dimensional, semantic vectors for low-cost yet accurate cache and lookup; (2) it uses a novel metric in determining the exit timing considering different layers' inherent characteristics; (3) it adaptively adjusts the cache size and semantic vectors to fit the scene dynamics. SMTM is prototyped on commodity CNN engine and runs on both mobile CPU and GPU. Extensive experiments on large-scale datasets and models show that SMTM can significantly speed up the model inference over standard approach (up to 2×) and prior cache designs (up to 1.5x), with acceptable accuracy loss. Chen Zhang 0001, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu 0001, Mengwei Xu 0001 |
ACM Multimedia | 5 |
| 2021 | ARNet: Accurate and Real-Time Network for Crowd Counting
Yinfeng Xia, Wenyue Wei, Baoqun Yin |
PRICAI (2) | 4 |
| 2021 | Pruning filters with L1-norm and capped L1-norm for CNN compression
Aakash Kumar, Ali Muhammad Shaikh, Hazrat Bilal, Baoqun Yin |
Appl. Intell. | 5 |
| 2021 | Multi-level feature fusion network for crowd countingabstractAbstract Crowd counting has become a noteworthy vision task due to the needs of numerous practical applications, but it remains challenging. State‐of‐the‐art methods generally estimate the density map of the crowd image with the high‐level semantic features of various deep convolutional networks. However, the absence of low‐level spatial information may result in counting errors in the local details of the density map. To this end, a novel framework named Multi‐level Feature Fusion Network (MFFN) for single image crowd counting is proposed. The proposed MFFN, which is constructed in an encoder–decoder fashion, incorporates semantic and spatial information for generating high‐resolution density maps of input crowd images. Skip connections are developed between the encoder and the decoder so that low‐level spatial information and high‐level semantic features can be combined by element‐wise addition. In addition, a dense dilated convolution block is placed behind the encoder, extracting multi‐scale context features to guide feature fusion by a channel attention mechanism. The model is trained by multi‐task learning; semantic segmentation supervision is introduced to enhance feature representation. Extensive experiments are conducted on three crowd counting datasets (ShanghaiTech, UCF_CC_50, UCF‐QNRF), and the results show that MFFN outperforms state‐of‐the‐art methods. In addition, sufficient ablation studies are performed to verify the effectiveness of each component in our proposed method. Sifan Peng, Baoqun Yin |
IET Comput. Vis. | 5 |
| 2021 | EDENet: Elaborate density estimation network for crowd counting
Yinfeng Xia, Yuqiang He, Sifan Peng, Xiaoliang Hao, Baoqun Yin |
Neurocomputing | 6 |
| 2021 | CFFNet: Coordinated feature fusion network for crowd counting
Yinfeng Xia, Yuqiang He, Sifan Peng, Baoqun Yin |
Image Vis. Comput. | 5 |
| 2021 | Adaptive weighted crowd receptive field network for crowd counting
Sifan Peng, Baoqun Yin, Yinfeng Xia, Xiaoliang Hao |
Pattern Anal. Appl. | 3 |
| 2021 | Depth and edge auxiliary learning for still image crowd density estimation
Sifan Peng, Baoqun Yin, Xiaoliang Hao, Aakash Kumar |
Pattern Anal. Appl. | 2 |
| 2019 | Using Feature Entropy to Guide Filter Pruning for Efficient Convolutional Networks
Sifan Peng, Aakash Kumar, Baoqun Yin |
ICANN (2) | 5 |
| 2019 | Removing background interference for crowd counting via de-background detail convolutional network
Baoqun Yin |
Neurocomputing | 2 |
| 2018 | A modified artificial bee colony approach for the 0-1 knapsack problem
Baoqun Yin, Xiaonong Lu, Yu Kang 0001 |
Appl. Intell. | 2 |
| 2018 | Skip-connection convolutional neural network for still image crowd counting
Baoqun Yin, Aixin Guo |
Appl. Intell. | 2 |
| 2018 | A SMDP-based forwarding scheme in named data networking
Jinfa Yao, Baoqun Yin, Xiaobin Tan |
Neurocomputing | 2 |
| 2018 | A novel POMDP-based server RAM caching algorithm for VoD systems
Baoqun Yin, Yu Kang 0001, Xiaonong Lu, Xiaofeng Jiang |
Multim. Tools Appl. | 1 |
| 2017 | A POMDP framework for forwarding mechanism in named data networking
Jinfa Yao, Baoqun Yin, Xiaobin Tan, Xiaofeng Jiang |
Comput. Networks | 2 |
| 2016 | Analysis of topology dynamics for unstructured P2P networks
Baoqun Yin, Xiaonong Lu, Yu Kang 0001 |
Comput. Commun. | 1 |
| 2008 | Event-related optimization for a class of resource location with admission controlabstractA class of resource location service for distributed VoD system, which combines one-hop k-random walk and global centralized indexing service, is studied. First, in order to minimizing the cost of communication and guaranteeing the response time performance, a Markov model is proposed to describe the queue phenomenon, admission control and the process of location. In this model, control is related with not only states but also events, which introduce more information as the control basis. Then, an optimization algorithm that combines policy gradient estimation and stochastic approximation is proposed. This algorithm can deal with constraints and depend on no system parameter. Finally, an illustrative simulation is performed to demonstrate the effectiveness of model and algorithm. Chenfeng Xu, Jian Yang 0014, Hongsheng Xi, Baoqun Yin |
IJCNN | 5 |
| 2008 | Partially Observable Markov Decision Processes and Performance Sensitivity AnalysisabstractThe sensitivity-based optimization of Markov systems has become an increasingly important area. From the perspective of performance sensitivity analysis, policy-iteration algorithms and gradient estimation methods can be directly obtained for Markov decision processes (MDPs). In this correspondence, the sensitivity-based optimization is extended to average reward partially observable MDPs (POMDPs). We derive the performance-difference and performance-derivative formulas of POMDPs. On the basis of the performance-derivative formula, we present a new method to estimate the performance gradients. From the performance-difference formula, we obtain a sufficient optimality condition without the discounted reward formulation. We also propose a policy-iteration algorithm to obtain a nearly optimal finite-state-controller policy. Baoqun Yin, Hongsheng Xi |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Sensitivity analysis and estimates of the performance for M/G/1 queueing systems
Baoqun Yin, Guiping Dai, Hongsheng Xi |
Perform. Evaluation | 1 |
| 2005 | Simulation-Based Optimization of Singularly Perturbed Markov Reward Processes with States Aggregation
Dali Zhang, Hongsheng Xi, Baoqun Yin |
ICIC (2) | 3 |