Chengpei Xu

dblp:246/5905 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-8860-7701ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Learning 3D Occupancy from Beam Overlap in 2D Rotating mmWave Radar
abstract
Robust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation information, they typically suffer from limited coverage and range. In this work, we exploit a key observation about mechanically rotating 2D mmWave Radars: in each sweep, an overlap exists between adjacent azimuth beam coverage due to the width of the main lobe, which makes the reflected intensity difference imply object materials and geometric shapes, including elevation. With this observation, we propose a method that learns 3D occupancy by disentangling bird’s-eye view (BEV) layout and elevation estimation from one frame Radar scan. Specifically, we partition one sweep into two interleaved subsets, corresponding to overlapping beam directions, and utilize them to infer coarse geometric structure through spatial differences and intensity patterns. Extensive quantitative and qualitative evaluations on two real-world datasets demonstrate that our proposed method outperforms existing baselines. The codes will be publicly available.
Ruifeng Nie, Long Ma 0002, Chengpei Xu, Yu Liu 0012, Weimin Wang 0007
AAAI4
2026 HiEn: Hierarchical ensemble learning for semi-supervised medical image segmentation
Long Ma 0002, Xinwei Xue, Chengpei Xu, Weimin Wang 0007, Yi Wang 0037
Neurocomputing5
2026 Visible-infrared joint image deraining for harsh rain conditions with cross-modal semantic consistency
Xin Li 0005, Chengpei Xu, Zhenyu Wang 0008, Weisheng Dong
Pattern Recognit.4
2026 SWG-Fusion: Soft weather-guided multimodal fusion with VLM-assistance for BEV object detection under harsh weather
Weimin Wang 0007, Ruifeng Nie, Yingchi Liu, Long Ma 0002, Chengpei Xu, Qi Jia 0001, Yu Liu 0012, Na Lei
Pattern Recognit.5
2026 Enhancing Underwater Images via Resonant Fusion
abstract
Recent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main.
Xinwei Xue, Zimeng Xu, Jincheng Yuan, Jingchun Zhou, Chengpei Xu, Xiaoke Shang, Long Ma 0002, Weimin Wang 0007
IEEE Trans. Image Process.6
2026 LKAFormer: A Lightweight Kolmogorov-Arnold Transformer Model for Image Semantic Segmentation
abstract
Transformer-based semantic segmentation methods have demonstrated outstanding performance by leveraging global self-attention to effectively capture long-range dependence. However, there still exist two issues in existing works: (1) Most of them utilize the full-rank weight matrix to support the self-attention mechanism and feed-forward network in modelling long-range dependence between patches/pixels, resulting in a high computational cost during both training and inference. (2) Most of them ignore information interactions between high-level semantics and low-level structures during the image resolution recovery, which leads to the performance degradation in segmenting objects with complex boundaries. To tackle these challenges, a lightweight Kolmogorov-Arnold Transformer model (LKAFormer) is proposed for the image semantic segmentation, containing a two-stream lightweight Transformer encoder and a graph feature pyramid aggregation KAN-decoder. The former constructs a hierarchical feature cross-scale fusion pipeline to obtain sufficient semantics containing comprehensive multi-scale information via setting coarse-grained and fine-grained streams with different-size patches of images. In that pipeline, feature lightweight focusing modules model complex and long-range dependence across patches/pixels to refine image semantics with less computational costs by lightweight multi-head self-attention and lightweight feed-forward network designs. The latter leverages the learnable nonlinear transformation mechanism of the Kolmogorov-Arnold Transformer architecture to adaptively capture spatial structure dependence of distinct sub-regions of images. And then, it jointly performs the intra-scale graph fusion and cross-scale graph fusion during the image resolution recovery to enhance information interactions between high-level semantics and low-level structures, which achieves the robust boundary localization and texture refinement of segmentation objects. Finally, plentiful experiments are conducted on three challenging datasets, and the results show LKAFormer sets a new baseline in the image segmentation task in comparison with 11 methods.
Shoulin Yin, Liguo Wang 0001, Tao Chen 0002, Huafei Huang 0001, Jing Gao 0007, Jianing Zhang 0001, Meng Liu 0025, Peng Li 0027, Chengpei Xu
ACM Trans. Intell. Syst. Technol.9
2025 FairGP: A Scalable and Fair Graph Transformer Using Graph Partitioning
abstract
Recent studies have highlighted significant fairness issues in Graph Transformer (GT) models, particularly against subgroups defined by sensitive features. Additionally, GTs are computationally intensive and memory-demanding, limiting their application to large-scale graphs. Our experiments demonstrate that graph partitioning can enhance the fairness of GT models while reducing computational complexity. To understand this improvement, we conducted a theoretical investigation into the root causes of fairness issues in GT models. We found that the sensitive features of higher-order nodes disproportionately influence lower-order nodes, resulting in sensitive feature bias. We propose Fairness-aware scalable GT based on Graph Partitioning (FairGP), which partitions the graph to minimize the negative impact of higher-order nodes. By optimizing attention mechanisms, FairGP mitigates the bias introduced by global attention, thereby enhancing fairness. Extensive empirical evaluations on six real-world datasets validate the superior performance of FairGP in achieving fairness compared to state-of-the-art methods.
Renqiang Luo, Huafei Huang 0001, Ivan Lee 0001, Chengpei Xu, Jianzhong Qi 0001, Feng Xia 0001
AAAI4
2025 CoA: Towards Real Image Dehazing via Compression-and-Adaptation
abstract
Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both efficiency and adaptability to address real image dehazing effectively. This work proposes a Compression-and-Adaptation (CoA) computational flow to tackle these challenges from a divide-and-conquer perspective. First, model compression is performed in the synthetic domain to develop a compact dehazing parameter space, satisfying efficiency demands. Then, a bilevel adaptation in the real domain is introduced to be fearless in unknown real environments by aggregating the synthetic dehazing capabilities during the learning process. Leveraging a succinct design free from additional constraints, our CoA exhibits domain-irrelevant stability and model-agnostic flexibility, effectively bridging the model chasm between synthetic and real domains to further improve its practical utility. Extensive evaluations and analyses underscore the approach's superiority and effectiveness. The code is publicly available at https://github.com/fyxnl/COA.
Long Ma 0002, Yan Zhang 0002, Jinyuan Liu 0001, Weimin Wang 0007, Guang-Yong Chen, Chengpei Xu, Zhuo Su 0001
CVPR7
2025 Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond
abstract
Recently, enhancing image quality in the original RAW domain has garnered significant attention, with denoising and reconstruction emerging as fundamental tasks. Although some works attempt to couple these tasks, they primarily focus on cascade learning while neglecting task associativity within a broader parameter space, leading to suboptimal performance. This work introduces a novel approach by rethinking denoising and reconstruction from a "backbone-head" perspective, leveraging the stronger shared parameter space offered by the backbone, compared to the encoder used in existing works. We derive task-specific heads with fewer parameters to mitigate learning pressure. By incorporating chromaticity-and-noise perception module into the backbone and introducing task-specific supervision during training, we enable simultaneous high-quality results for reconstruction and denoising. Additionally, we design a dual-head interaction module to capture the latent correspondence between the two tasks, significantly enhancing multi-task accuracy. Extensive experiments validate the superiority of the proposed method. Code is available at: https://github.com/csmty/CANS.
Tengyu Ma 0004, Long Ma 0002, Ziye Li, Yuetong Wang, Jinyuan Liu 0001, Chengpei Xu, Risheng Liu
CVPR6
2025 Bright to Dark: Stage-wise Bilevel Knowledge Transfer for Seeing Text in the Dark
abstract
Localizing text under low-light conditions has gained attention, with typical approaches relying on two stage cascading modules that combine low-light enhancement and text localization. However, these often require additional enhancement modules and cause inefficiency in joint optimization. In this work, we address the challenge by adopting a novel approach: tailoring the detector for low light conditions through knowledge distillation from normal light conditions, without relying on any enhancement module. First, we design a Graph Topological Aggregation (GTA) model that utilizes the message passing mechanism of graph neural networks to structurally represent text topology and facilitate structured feature expression in knowledge transfer. We then introduce two specially designed knowledge transfer constraints aimed at enhancing the learning of text's multi-scale features and topological knowledge. Finally,we propose a Stage-wise Bilevel Knowledge Transfer learning strategy that designates the low-light learning process as the upper-level task, while treating normal light learning as the lower-level task, effectively addressing the coupling issues and sequential dependencies prevalent during the distillation process. Extensive experiments underscore the approach's superiority.
Chengpei Xu, Long Ma 0002, Weimin Wang 0007, Feng Xia 0001, Binghao Li, Wenjie Zhang 0001
ACM Multimedia1
2025 SEHG: Bridging Interpretability and Prediction in Self-Explainable Heterogeneous Graph Neural Networks
abstract
Heterogeneous Graph Neural Networks (HGNNs) are extensively applied in modeling web-based applications that involve heterogeneous graph structures. Explanation models for HGNNs aim to address their ''black box'' nature. Enhancing the interpretability of HGNNs leads to a better understanding and can potentially improve predictive performance. However, existing post-hoc HGNN explanation methods cannot impact the HGNN's predictions. Self-explainable homogeneous models also perform poorly on heterogeneous graphs. To address these challenges, we present a Self-Explainable Heterogeneous Graph Neural Network (SEHG), a novel architecture that integrates explanation generation into the learning process of HGNN through two alternative stages. The first stage focuses on producing high-quality explanations while providing predictions alongside. The second stage enhances prediction accuracy by a contrastive learning strategy. Unlike the current methods that rely on manually defined metapaths for structural explanations, SEHG generates important structure and feature explanations by learnable heterogeneous masks. To ensure high-quality and sparsity explanation, these masks are regulated by a uniquely designed range-based penalty during training. Moreover, we introduce HetBA, a collection of synthetic heterogeneous datasets designed to quantify and visualize explanations or heterogeneous graphs. Extensive experiments demonstrate the effectiveness of SEHG, which surpasses strong baselines in real-world node classification tasks by notable margins of up to 3.91%. SEHG also achieves state-of-the-art performance on synthetic datasets with improvement of up to 9.44%, and records the highest fidelity scores in explanation tasks, improving by up to 46.57%. To our knowledge, SEHG is a pioneering self-explainable HGNN framework that achieves state-of-the-art performance on both heterogeneous graph explanation and prediction tasks.
Zhenhua Huang 0002, Xiuyang Wu, Chengpei Xu, Junfeng Fang, Linyuan Lu, Feng Xia 0001
WWW5
2025 Deep dive into clarity: Leveraging signal-to-noise ratio awareness and knowledge distillation for underwater image enhancement
Jingchun Zhou, Chengpei Xu
Expert Syst. Appl.3
2025 Underwater variable zoom: Depth-guided perception network for underwater image enhancement
Zhixiong Huang, Xinying Wang 0005, Chengpei Xu, Jinjiang Li 0001, Lin Feng 0001
Expert Syst. Appl.3
2025 FSCMF: A Dual-Branch Frequency-Spatial Joint Perception Cross-Modality Network for visible and infrared image fusion
Chengpei Xu, Zhen Hua, Jinjiang Li 0001, Jingchun Zhou
Neurocomputing2
2025 Learning With Self-Calibrator for Fast and Robust Low-Light Image Enhancement
abstract
Convolutional Neural Networks (CNNs) have shown significant success in the low-light image enhancement task. However, most of existing works encounter challenges in balancing quality and efficiency simultaneously. This limitation hinders practical applicability in real-world scenarios and downstream vision tasks. To overcome these obstacles, we propose a Self-Calibrated Illumination (SCI) learning scheme, introducing a new perspective to boost the model's capability. Based on a weight-sharing illumination estimation process, we construct an embedded self-calibrator to accelerate stage-level convergence, yielding gains that utilize only a single basic block for inference, which drastically diminishes computation cost. Additionally, by introducing the additivity condition on the basic block, we acquire a reinforced version dubbed SCI++, which disentangles the relationship between the self-calibrator and illumination estimator, providing a more interpretable and effective learning paradigm with faster convergence and better stability. We assess the proposed enhancers on standard benchmarks and in-the-wild datasets, confirming that they can restore clean images from diverse scenes with higher quality and efficiency. The verification on different levels of low-light vision tasks shows our applicability against other methods.
Long Ma 0002, Tengyu Ma 0004, Chengpei Xu, Jinyuan Liu 0001, Xin Fan 0001, Zhongxuan Luo, Risheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Towards Multimodal Metaphor Understanding: A Chinese Dataset and Model for Metaphor Mapping Identification
abstract
Metaphors play a crucial role in human communication, yet their comprehension remains a significant challenge for natural language processing (NLP) due to the cognitive complexity involved. According to Conceptual Metaphor Theory (CMT), metaphors map a target domain onto a source domain, and understanding this mapping is essential for grasping the nature of metaphors. Existing NLP research has focused on tasks like metaphor detection and sentiment analysis. However, there has been limited attention to identifying mappings between source and target domains. Moreover, non-English multimodal metaphor resources remain largely neglected in the literature, hindering a deeper understanding of the key elements involved in metaphor interpretation. To address this gap, we developed a Chinese multimodal metaphor advertisement dataset (namely CM3D) that includes annotations of specific target and source domains. This dataset aims at fostering further research into metaphor comprehension, particularly in non-English languages. Furthermore, we propose a Chain-of-Thought (CoT) Prompting-based Metaphor Mapping Identification Model (CPMMIM), which simulates the human cognitive process for identifying these mappings. Drawing inspiration from CoT reasoning and Bi-Level Optimization (BLO), we treat the task as a hierarchical identification problem, enabling more accurate and interpretable metaphor mapping. Our experimental results demonstrate the effectiveness of CPMMIM, highlighting its potential for advancing metaphor comprehension in NLP. Our dataset and code are both publicly available to encourage further advancements in this field.
Dongyu Zhang 0001, Shengcheng Yin, Jingwei Yu, Zhiyao Wu, Zhen Li 0014, Chengpei Xu, Feng Xia 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2025 SPRMamba: A Mamba-Based Saliency Proportion Reconciliatory Network With Squeezed Windows for Remote Sensing Change Detection
abstract
Remote sensing (RS) change detection (CD) faces challenges in effectively identifying non-salient change regions, such as subtle architectural modifications or changes that closely resemble the background. The primary difficulties stem from the weak feature representation of non-salient changes, which results in insufficient model response, and the high similarity between background and change regions, leading to misdetection or omission. To address this issue, we propose SPRMamba, a Mamba-based saliency proportion reconciliatory network with squeezed windows for remote sensing change detection. It introduces state space models with windowing operations and cross-window interaction mechanisms to improve the response to weak signals. To dynamically balance the representation of salient and non-salient features, we design the saliency proportion reconciler (SPR) to optimize the discrimination between background and change regions. In addition, we introduce a sparse saliency loss function, which imposes sparsity constraints on salient regions to enhance the feature representation of non-salient change regions. Experimental results show that SPRMamba significantly outperforms existing methods on several public datasets. Our code will be available at https://github.com/boomstarzzn/SPRMamba.
Shengning Zhou, Chengpei Xu, Jinjiang Li 0001, Zhen Hua, Jingchun Zhou
IEEE Trans. Geosci. Remote. Sens.2
2024 Seeing Text in the Dark: Algorithm and Benchmark
abstract
Localizing text in low-light environments is challenging due to visual degradations. Although a straightforward solution involves a two-stage pipeline with low-light image enhancement (LLE) as the initial step followed by detection, LLE is primarily designed for human vision rather than machine vision and can accumulate errors. In this work, we propose an efficient and effective single-stage approach for localizing text in the dark that circumvents the need for LLE. We introduce a constrained learning module as an auxiliary mechanism during the training stage of the text detector. This module is designed to guide the text detector in preserving textual spatial features amidst feature map resizing, thus minimizing the loss of spatial information in texts under low-light visual degradations. Specifically, we incorporate spatial reconstruction and spatial semantic constraints within this module to ensure the text detector acquires essential positional and contextual range knowledge. Our approach enhances the original text detector's ability to identify text's local topological features using a dynamic snake feature pyramid network and adopts a bottom-up contour shaping strategy with a novel rectangular accumulation technique for accurate delineation of streamlined text features. In addition, we present a comprehensive low-light dataset for arbitrary-shaped text, encompassing diverse scenes and languages. Notably, our method achieves state-of-the-art results on this low-light dataset and exhibits comparable performance on standard normal light datasets. The code and dataset will be released.
Chengpei Xu, Hao Fu 0004, Long Ma 0002, Wenjing Jia, Chengqi Zhang, Feng Xia 0001, Xiaoyu Ai, Binghao Li, Wenjie Zhang 0001
ACM Multimedia1
2024 FeatSync: 3D point cloud multiview registration with attention feature-based refinement
abstract
Current studies often decompose multiview registration into several individual tasks, ignoring the correlation between each stage and making certain assumptions on noise distribution without knowledge from previous stages. These issues bring difficulties in generalization to real cases. In this paper, we propose an end-to-end feature-based multiview registration model that takes a set of raw 3D point cloud fragments as input and outputs the global transformation. Unlike previous works, our method allows the exchange of information between stages. We firstly estimate pairwise registration by a attention-based model to assist feature learning. In the next stage, we utilize iteratively reweighted least squares (IRLS) algorithm to refine and obtain the global transformation. In each iteration, instead of making assumptions on noises, we directly construct a model to infer the outliers from pairwise registration so that such an inference can help synchronization produce more reliable results. To follow the process in IRLS algorithm, we propose a simple yet effective refinement module to boost feature-based pairwise estimations in an iterative manner, which can be seamlessly integrated into the IRLS procedure. Extensive experiments conducted on benchmark datasets show that the results of our proposed method outperformed existing methods.
Yiheng Hu, Binghao Li, Chengpei Xu, Sarp Saydam, Wenjie Zhang 0001
Neurocomputing3
2023 InterREC: An Interpretable Method for Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC) aims to locate the target object in the image according to a referring expression. This is a challenging task owing to the need for understanding both natural language and visual information and interpretable reasoning between them. Most existing implicit reasoning-based REC methods lack interpretability, while explicit reasoning-based REC methods have lower accuracy. To achieve competitive accuracy while providing adequate interpretability, in this work, we propose a novel explicit reasoning-based method named InterREC. First, in order to address the challenge of multi-modal understanding, we design two neural network modules based on text-image representation learning: a Text-Region Matching Module to align objects in the image and noun phrases in the expression, and a Text-Relation Matching Module to align relations between objects in the image and relational phrases in the expression. Additionally, we design a Reasoning Order Tree for handling complex expressions, which can reduce complex expressions to multiple object-relation-object triplets and therefore identify the inference order and reduce the difficulty of reasoning. At the same time, to achieve an interpretable reasoning step, we design a Bayesian Network-based explicit reasoning method. Based on the comparative evaluation on various datasets, our method achieves higher accuracy than existing explicit reasoning-based REC methods, and the visualization results demonstrate the method's high interpretability.
Maurice Pagnucco, Chengpei Xu, Yang Song 0001
IEEE Trans. Multim.3
2023 Arbitrary-Shape Scene Text Detection via Visual-Relational Rectification and Contour Approximation
abstract
One trend in the latest bottom-up approaches for arbitrary-shape scene text detection is to determine the links between text segments using Graph Convolutional Networks (GCNs). However, the performance of these bottom-up methods is still inferior to that of state-of-the-art top-down methods even with the help of GCNs. We argue that a cause of this is that bottom-up methods fail to make proper use of visual-relational features, which results in accumulated false detection, as well as the error-prone route-finding used for grouping text segments. In this paper, we improve classic bottom-up text detection frameworks by fusing the visual-relational features of text with two effective false positive/negative suppression (FPNS) mechanisms and developing a new shape-approximation strategy. First, dense overlapping text segments depicting the “characterness” and “streamline” properties of text are constructed and used in weakly supervised node classification to filter the falsely detected text segments. Then, relational features and visual features of text segments are fused with a novel Location-Aware Transfer (LAT) module and Fuse Decoding (FD) module to jointly rectify the detected text segments. Finally, a novel multiple-text-map-aware contour-approximation strategy is developed based on the rectified text segments, instead of the error-prone route-finding process, to generate the final contour of the detected text. Experiments conducted on five benchmark datasets demonstrate that our method outperforms the state-of-the-art performance when embedded in a classic text detection framework, which revitalizes the strengths of bottom-up methods.
Chengpei Xu, Wenjing Jia, Tingcheng Cui, Ruomei Wang 0001, Yuan-fang Zhang, Xiangjian He
IEEE Trans. Multim.1
2023 MorphText: Deep Morphology Regularized Accurate Arbitrary-Shape Scene Text Detection
abstract
Bottom-up text detection methods play an important role in arbitrary-shape scene text detection but there are two restrictions preventing them from achieving their great potential, i.e., 1) the accumulation of false text segment detections, which affects subsequent processing, and 2) the difficulty of building reliable connections between text segments. Targeting these two problems, we propose a novel approach, named ``MorphText", to capture the regularity of texts by embedding deep morphology for arbitrary-shape text detection. Towards this end, two deep morphological modules are designed to regularize text segments and determine the linkage between them. First, a Deep Morphological Opening (DMOP) module is constructed to remove false text segment detections generated in the feature extraction process. Then, a Deep Morphological Closing (DMCL) module is proposed to allow text instances of various shapes to stretch their morphology along their most significant orientation while deriving their connections.Extensive experiments conducted on four challenging benchmark datasets (CTW1500, Total-Text, MSRA-TD500 and ICDAR2017) demonstrate that our proposed MorphText outperforms both top-down and bottom-up state-of-the-art arbitrary-shape scene text detection approaches.
Chengpei Xu, Wenjing Jia, Ruomei Wang 0001, Xiangjian He
IEEE Trans. Multim.1
2021 Rethinking feature aggregation for deep RGB-D salient object detection
Yuanfang Zhang, Jiangbin Zheng 0001, Long Li 0008, Nian Liu 0002, Wenjing Jia, Xiaochen Fan, Chengpei Xu, Xiangjian He
Neurocomputing7
2019 Lecture2Note: Automatic Generation of Lecture Notes from Slide-Based Educational Videos
abstract
Given rapid development witnessed by open educational resources (OER) in the past few decades, a considerable number of online educational videos emerge on various MOOC platforms such as Coursera and YouTube. Nevertheless, most educational videos on the internet are lengthy and lack of elaborate annotations, which poses a challenge for learners to explore and locate content of interest efficiently. To address this, we present an automatic note-generating method to establish correspondences between visual entities in the slide-based lecture video and their descriptive speech texts by evaluating the semantic relationship. Firstly, the visual entities are extracted and recognised from the presentation slides. Then, each of visual entities is associated with its corresponding descriptive speech text. Finally, a placement optimisation scheme is put forward to pack the visual entities and speech texts into a note-like layout in a compact fashion, which can help learners to improve their learning efficiency. The experimental results show that the efficient performances about visual entity extraction and correspondence matching are efficient. The user study is also designed to investigate the performance of Lecture2Note in facilitating learning. Compared with peer methods, the auto-generated note created by our method achieves a higher user satisfaction level regarding a properly structured layout as well as efficient content navigation and exploration.
Chengpei Xu, Ruomei Wang 0001, Shujin Lin, Baoquan Zhao, Lijie Shao, Mengqiu Hu
ICME1