VLDB 2026 Research / reviewers in the wild / expert
Boan Chen
dblp:299/6727
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-4484-3416ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-modal attention fusion of RGB and skeleton for multimodal-driven video anomaly detection
Boan Chen, Weide Liu, Jinmei Liu, Baoquan Zhao, Yang Liu 0246 |
Pattern Recognit. | 1 |
| 2025 | SVR-YOLO11: Real-Time Animal Detection for Situational Awareness in SAR OperationsabstractReal-time animal detection in search-and-rescue (SAR) operations represents a critical challenge for situational awareness systems, particularly when deploying lightweight solutions on resource-constrained edge computing platforms. Current detection methods suffer from computational bottlenecks that compromise either accuracy or real-time performance, limiting their effectiveness in time-critical rescue scenarios. This paper presents SVR-YOLO11, an enhanced detection network specifically optimized for real-time animal identification in SAR applications. The proposed architecture introduces three key in-novations: (1) a Slim-Neck design paradigm that preserves inter-channel connections while reducing computational complexity, (2) integration of Large Separable Kernel Attention (LSKA) modules with dynamic upsampling to improve detection accuracy for morphologically similar animal species across diverse natural backgrounds, and (3) direct deployment capability on edge devices without requiring complex model conversion processes. Experimental validation on the Animals. v4i dataset demonstrates that SVR-YOLO11 achieves superior performance with 68.3% mAP50-95 while maintaining only 2.843M parameters and 6.5 GFLOPs computational cost. Real-time testing on multiple edge computing platforms, including Raspberry Pi and NVIDIA Jetson devices, confirms the network's practical applicability for field deployment. The system's effectiveness is further validated through a comprehensive head-mounted display demonstration using Unity Engine simulations that replicate dynamic SAR environments. These results establish SVR-YOLO11 as a ro-bust solution for real-time animal detection in distributed edge computing scenarios, directly supporting enhanced situational awareness in critical rescue operations. Zhongyi He, Yang Liu 0246, Hao Yang 0055, Peng Sun 0007, Boan Chen |
DS-RT | 5 |
| 2025 | Melody Structure Transfer Network: Generating Music with Separable Self-AttentionabstractMost existing symbolic music generation methods focus on generating short pieces, typically less than 8 bars and occasionally up to 32 bars. Generating long music sequences requires effective representation of coherent musical structures. Vanilla self-attention face challenges in capturing subtle long-term musical structures. We propose an approach to transfer the structural characteristics of training samples for generating music. We introduce a separable self-attention-based model that facilitates the learning and transfer of structural embeddings. It can generate music sequences of up to 100 bars, producing compositions with interpretable structures that closely resemble the structural and compositional techniques of the training set. Experiments show the model’s ability to generate music with targeted structures while maintaining good diversity. Ning Zhang 0040, Boan Chen, Huanxi Liu, Junchi Yan |
ICASSP | 4 |
| 2025 | LaTeXNet: A Specialized Model for Converting Visual Tables and Equations to LaTeX CodeabstractLaTeX provides precise representation of complex elements (i.e., tables and equations) in scientific documents. However, the automated transcription of visual representations into LaTeX code is challenging and prone to errors. This paper introduces LaTeXNet, a specialized model designed to automate the conversion of visual tables and equations into LaTeX code. First, we develop an automated annotation tool that extracts image-LaTeX pairs for tables, equations, and text paragraphs with inline equations from arXiv platform, creating the MM-LaTeX dataset with over 2.5M pairs. Moreover, we design the LaTeXNet model, trained on MM-LaTeX, which unifies the conversion of Tables, Equations, and TextEqs. Our experimental results indicate that LaTeXNet surpasses both open-source and commercial, closed-source models in Table-to-LaTeX, Equation-to-LaTeX and TextEq-to-LaTeX tasks. Renqiu Xia, Hongbin Zhou, Ziming Feng, Huanxi Liu, Boan Chen, Junchi Yan |
ICASSP | 5 |
| 2025 | Towards Green VAE: A Light Pixel-weighting Technique to Enhance Variational AutoEncoderabstractVariational autoencoders (VAEs) has been a popular generative model for its effectiveness, mathematical foundation, and its impact to other approaches in deep generative learning. For its relatively light-weights and easiness for training, compared with Generative Adversarial Networks (GANs) or other large-scale model e.g. Diffusion, VAEs become a viable tool in cost-efficient generative applications especially considering the so-called green AI. However, in comparison to GANs, the performance of VAEs in generating realistic images is still inferior to the state-of-the-art generative adversarial network (GAN). In this paper, we argue that this problem is at least partly due to the irrational reconstruction in VAE that all pixels are equally weighted, which is harmful to the generating ability (or density estimate). Motivated by this, we propose to compute the weights of pixels. First, we formulate the problem of finding the appropriate weights into an optimal problem, and then give an analytical solution. Moreover, we propose a method to apply the computed weights of pixels into the training pipeline with almost no computation overhead which fits with the spirit of green AI for more energy-saving AI especially for deep learning models. Experiments on MNIST, Fashion-MNIST and CIFAR-10 show that our method can significantly improves VAE in terms of FID. Ziming Feng, Boan Chen, Junchi Yan |
ICASSP | 4 |
| 2025 | SAM2-ProMem: Enhancing Zero-Shot 3D Segmentation with Stochastic Propagation and Memory Search
Juntao Huang, Dazhu Liang, Fangzhou Liao, Boan Chen |
MICCAI (6) | 6 |
| 2025 | M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content becomes increasingly complex with diverse behavioral patterns and contextual scenarios, traditional VAD approaches struggle to provide robust assessment for modern surveillance systems. Existing methods either lack comprehensive spatial-temporal modeling or require excessive computational resources for real-time applications. In this regard, we present a Mamba-based multi-scale spatial-temporal learning (M2S2L) framework in this paper. The proposed method employs hierarchical spatial encoders operating at multiple granularities and multi-temporal encoders capturing motion dynamics across different time scales. We also introduce a feature decomposition mechanism to enable task-specific optimization for appearance and motion reconstruction, facilitating more nuanced behavioral modeling and quality-aware anomaly assessment. Experiments on three benchmark datasets demonstrate that M2S2L framework achieves 98.5%, 92.1%, and 77.9% frame-level AUCs on UCSD Ped2, CUHK Avenue, and ShanghaiTech respectively, while maintaining efficiency with 20.1G FLOPs and 45 FPS inference speed, making it suitable for practical surveillance deployment. Yang Liu 0246, Boan Chen, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Wei Zhou 0013 |
VCIP | 2 |
| 2024 | View Crafting For Instance-Level Representation from Scene ImagesabstractExisting image-level self-supervised learning (SSL) methods pre-trained on natural scene data can have difficulty in adating to dense prediction tasks. However, scene images contain multiple varied instances. We devise two techniques to craft high-quality scene and instance views for instance-level SSL. Firstly, we leverage the prior from the image-level pre-trained model to discover the salient instances and the correspondence between cross-image instance pairs. Then, we augment multiple instances to compose the scene views, where both scene similarity and variance are promoted. Meanwhile, we encourage the scene and inner instances to align in feature space to model the scene-instance correlation. Experiments show its SOTA performance in dense prediction tasks. Bin Liu 0054, Shaofeng Zhang, Zehuan Yuan, Changdong Xu, Boan Chen, Junchi Yan |
ICASSP | 6 |
| 2024 | Improved Message Mechanism-Based Cross-Domain Security Control Model in Mobile TerminalsabstractDual-domain terminal with two built-in independent operating systems - Life Domain and Work Domain, provides convenience for daily use and mobile office. However, the security isolation between the two domains also causes that message reminders cannot be delivered and viewed across domains, which restricts the improvement of work efficiency and the expansion of mobile services. This paper conducts an in-depth study on this pain point and proposes the concept and implementation method of a cross-domain instant messaging reminder service system for mobile office, focusing on solving the problems of: cross-domain isolated boundary exchange of message reminders, timeliness and delivery rate guarantee of message reminders, and security check filtering of message contents. Technically, on the side of mobile office platform, based on AMQP technical framework and protocol, the cross-domain isolated border message queue push and synchronization services are built, which are real-time, reliable and high-throughput. Zhijie Fan, Boan Chen, Zidong Cheng, Shijun Xu |
Int. J. Inf. Secur. Priv. | 3 |
| 2024 | Hierarchical GNN Framework for Earth's Surface Anomaly Detection in Single Satellite ImageryabstractSudden-onset Earth’s surface anomalies, such as natural disasters and man-made incidents, pose severe threats to human life and property security, emphasizing the crucial role of accurate detection and rapid response in Humanitarian Assistance and Disaster Response (HADR). In this work, we propose a hierarchical graph neural network (GNN) based framework for Earth’s surface anomaly detection, called L2S-Net, to integrate from local to semantic (L2S) information for rapid and accurate detection of multi-class anomalies. Specifically, L2S-Net only utilizes a single very high-resolution (VHR) image as input to expedite processing speed, while employing a hierarchical graph representation for better image understanding. Meanwhile, drawing from brain-inspired research and graph theory, we design a local-to-semantic fusion network, called L2S-GNN, to explicitly learn relationships between nodes at different levels facilitating accurate detection of Earth’s surface anomaly. L2S-Net significantly reduces data requirements while capturing valuable higher-order information from images, achieving a superior balance between accuracy and efficiency. Furthermore, due to the lack of a public dataset for Earth’s surface anomaly detection, we create a novel and large-scale benchmark dataset ESADv2. Extensive experiments on the ESADv2 dataset and two real-world cases demonstrate that the proposed L2S-Net outperforms many state-of-the-art methods in both model size and performance while exhibiting exceptional generalizability and robustness. Boan Chen, Zhi Gao 0005, Ziyao Li, Aohan Hu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Semi-Supervised Few-Shot Classification With Multitask Learning and Iterative Label CorrectionabstractFew-shot learning enables rapid generalization from extremely limited training examples. While previous efforts have utilized meta-learning or data augmentation methods to mitigate the problem of data scarcity, such approaches may struggle to maintain robustness and generalize effectively due to overfitting and noise sensitivity. In this paper, we propose a novel approach, the Semi-Supervised Label Correction method for Few-Shot Learning (SSLC-FSL), which leverages the data distribution of readily available and easily obtainable unlabeled data. SSLC-FSL iteratively corrects the labels of testing samples with alternating steps of pseudo-labeling and sample selection. The objective of pseudo-labeling is to repurpose graph-based semi-supervised learning for joint prediction of the entire testing set. We then introduce a Modulation Selection Network (MSN) to rank testing samples by learning with noisy labels. The training set is expanded by selecting confident pseudo-labeled samples. In the MSN, a Modulation Aggregation Layer is designed to encode support class information into each testing sample, thereby highlighting target category features and mitigating the negative impact of incorrect labels. The iterative label correction process is repeated until all testing samples are recalled to the expanded support set. To boost the SSLC-FSL algorithm, we pre-train a feature extractor to produce general-purpose representations. Particularly, we investigate two types of auxiliary tasks and their collaborative learning to acquire transferable visual information via an end-to-end multi-task learning model. Our SSLC-FSL outperforms current state-of-the-art methods in any shot and all data settings, with up to +27.74% on standard remote sensing benchmarks and +5.70% on standard natural scene benchmarks. Zhi Gao 0005, Ziyao Li, Boan Chen, Yanzhang Li, Zhicheng Shi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Adaptive Embedding and Distribution Re-margin for Long-Tail Recognition
Yulin Su, Boan Chen, Ziming Feng, Junchi Yan |
ICANN (8) | 2 |
| 2023 | Image Deblurring With Image BlurringabstractDeep learning (DL) based methods for motion deblurring, taking advantage of large-scale datasets and sophisticated network structures, have reported promising results. However, two challenges still remain: existing methods usually perform well on synthetic datasets but cannot deal with complex real-world blur, and in addition, over- and under-estimation of the blur will result in restored images that remain blurred and even introduce unwanted distortion. We propose a motion deblurring framework that includes a Blur Space Disentangled Network (BSDNet) and a Hierarchical Scale-recurrent Deblurring Network (HSDNet) to address these issues. Specifically, we train an image blurring model to facilitate learning a better image deblurring model. Firstly, BSDNet learns how to separate the blur features from blurry images, which is adaptable for blur transferring, dataset augmentation, and ultimately directing the deblurring model. Secondly, to gradually recover sharp information in a coarse-to-fine manner, HSDNet makes full use of the blur features acquired by BSDNet as a priori and breaks down the non-uniform deblurring task into various subtasks. Moreover, the motion blur dataset created by BSDNet also bridges the gap between training images and actual blur. Extensive experiments on real-world blur datasets demonstrate that our method works effectively on complex scenarios, resulting in the best performance that significantly outperforms many state-of-the-art approaches. Ziyao Li, Zhi Gao 0005, Han Yi, Boan Chen |
IEEE Trans. Image Process. | 5 |