VLDB 2026 Research / reviewers in the wild / expert
Lele Xie
dblp:214/0068
· DBLP profile ↗
12ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Image recognition and object detection · 51% Knowledge representation and reasoning · 28% Information extraction and text analysis · 10% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
scene text detection |
0.9 | 3 | 2019 | Omnidirectional Scene Text Detection with Sequential-free Box Discretization · IJCAI 2019 Tightness-Aware Evaluation Protocol for Scene Text Detection · CVPR 2019 DeRPN: Taking a Further Step toward More General Object Detection · AAAI 2019 |
Computer vision › Image recognition and object detection › object detection
detector training |
0.8 | 1 | 2024 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer · Int. J. Comput. Vis. 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
first-order logic constraints |
0.8 | 1 | 2024 | LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints · ICLR 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
markov logic networks |
0.8 | 1 | 2024 | LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints · ICLR 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.8 | 1 | 2024 | LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic Constraints · ICLR 2024 |
Computer vision › Image recognition and object detection
object detection |
0.8 | 2 | 2019 | Omnidirectional Scene Text Detection with Sequential-free Box Discretization · IJCAI 2019 DeRPN: Taking a Further Step toward More General Object Detection · AAAI 2019 |
Computer vision › Image recognition and object detection › object detection
training from scratch |
0.8 | 1 | 2024 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer · Int. J. Comput. Vis. 2024 |
Computer vision › Image recognition and object detection
scene text recognition |
0.7 | 1 | 2023 | Fine-grained Pseudo Labels for Scene Text Recognition · ACM Multimedia 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction › multimodal information extraction
visual information extraction |
0.5 | 1 | 2021 | MatchVIE: Exploiting Match Relevancy between Entities for Visual Information Extraction · IJCAI 2021 |
Computer vision › Image recognition and object detection › object detection
anchor-free object detection |
0.4 | 1 | 2019 | DeRPN: Taking a Further Step toward More General Object Detection · AAAI 2019 |
Machine learning › Deep learning architectures and training
loss function design |
0.4 | 1 | 2019 | Aggregation Cross-Entropy for Sequence Recognition · CVPR 2019 |
Computer vision › Image recognition and object detection › object detection
object proposal generation |
0.4 | 1 | 2019 | DeRPN: Taking a Further Step toward More General Object Detection · AAAI 2019 |
Computer vision › Image recognition and object detection › object detection
oriented object detection |
0.4 | 1 | 2019 | Omnidirectional Scene Text Detection with Sequential-free Box Discretization · IJCAI 2019 |
Natural language and speech › Speech recognition and synthesis
sequence recognition |
0.4 | 1 | 2019 | Aggregation Cross-Entropy for Sequence Recognition · CVPR 2019 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.2 | 1 | 2024 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer · Int. J. Comput. Vis. 2024 |
Computer vision › Image recognition and object detection › object detection › category-specific object detection
text detection |
0.1 | 1 | 2019 | Omnidirectional Scene Text Detection with Sequential-free Box Discretization · IJCAI 2019 |
Methods — techniques the papers use, named apart from their topics
mean-field variational inference · 0.8teacher-student framework · 0.7pseudo-labeling · 0.7curriculum learning · 0.7adaptive distribution regularization · 0.7num2vec · 0.5graph neural network · 0.5scale-sensitive loss · 0.4intersection over union · 0.4dimension decomposition · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Dynamic Multi-Objective Optimization Time Series Ensemble Prediction Framework Based on Correlation Type DetectionabstractDynamic multi-objective optimization problems (DMOPs) exhibit variations in constraints, decision parameter quantities, and the number of objectives over time, leading to changes in their optimal solutions. Prediction-based methods currently stand as the prevailing paradigm for handling DMOPs. Implicit correlations among solutions ob-tained at consecutive time steps while solving DMOPs con-tinuously can guide predictions for future time increments in the time series. Most existing methods rely on a single linear or nonlinear model to capture the correlations of historical optimal solutions. However, a linear model struggles with the nonlinear scenario, and using a nonlinear model to fit the linear relationship of the historical optimal solutions increases the training time because the nonlinear model has more parameters. This paper proposes a novel framework for time series ensemble prediction based on correlation type detection TSEPFCTD. The framework leverages correlation type detection strategy to identify the autocorrelation types in the time series of historical solutions. Subsequently, distinct prediction models are applied for different types, enhancing prediction accuracy. We implement this prediction framework on MOEA/D-DE, constructing a novel algorithm for solving DMOPs. Experimental results on a series of test suites demonstrate the effectiveness of our algorithm in providing robust solution outcomes. Lele Xie |
CEC | 1 |
| 2024 | LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic ConstraintsabstractIntegrating first-order logic constraints (FOLCs) with neural networks is a crucial but challenging problem since it involves modeling intricate correlations to satisfy the constraints. This paper proposes a novel neural layer, LogicMP, which performs mean-field variational inference over a Markov Logic Network (MLN). It can be plugged into any off-the-shelf neural network to encode FOLCs while retaining modularity and efficiency. By exploiting the structure and symmetries in MLNs, we theoretically demonstrate that our well-designed, efficient mean-field iterations greatly mitigate the difficulty of MLN inference, reducing the inference from sequential calculation to a series of parallel tensor operations. Empirical results in three kinds of tasks over images, graphs, and text show that LogicMP outperforms advanced competitors in both performance and efficiency. Weidi Xu, Lele Xie, Jianshan He, Hongting Zhou, Taifeng Wang, Xiaopei Wan, Jingdong Chen, Chao Qu |
ICLR | 3 |
| 2024 | Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerabstractAbstract Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performance of self-attention mechanism in the language field, transformers tailored for visual data have drawn significant attention and triumphed over CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train and fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier works in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to a vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights to help other researchers and practitioners, and inspire more interesting research in other fields, such as remote sensing, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on the MS COCO dataset demonstrate that vision transformer based detectors trained from scratch can also achieve similar performance to their counterparts with ImageNet pre-training. Weixiang Hong 0001, Wang Ren, Jiangwei Lao, Lele Xie, Liheng Zhong, Jian Wang 0108, Jingdong Chen, Honghai Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | A Novel Mode Selection-Based Fast Intra Prediction Algorithm for Spatial SHVCabstractDue to multi-layer encoding and Inter-layer prediction, Spatial Scalable High-Efficiency Video Coding (SSHVC) has extremely high coding complexity. It is very crucial to improve its coding speed so as to promote widespread and cost-effective SSHVC applications. In this paper, we have proposed a novel Mode Selection-Based Fast Intra Prediction algorithm for SSHVC. We reveal the RD costs of Inter-layer Reference (ILR) mode and Intra mode have a significant difference, and the RD costs of these two modes follow Gaussian distribution. Based on this observation, we propose to apply the classic Gaussian Mixture Model and Expectation Maximization in machine learning to determine whether ILR is the best mode so as to skip the Intra mode. Experimental results demonstrate that the proposed algorithm can significantly improve the coding speed with negligible coding efficiency loss. Yu Sun 0003, Weisheng Li 0001, Lele Xie, Xin Lu 0001, Frédéric Dufaux, Ce Zhu |
ICASSP | 4 |
| 2023 | Fine-grained Pseudo Labels for Scene Text RecognitionabstractPseudo-Labeling based semi-supervised learning has shown promising advantages in Scene Text Recognition (STR). Most of them usually use a pre-trained model to generate sequence-level pseudo labels for text images and then re-train the model. Recently, conducting Pseudo-Labeling in a teacher-student framework (a student model is supervised by the pseudo labels from a teacher model) has become increasingly popular, which trains in an end-to-end manner and yields outstanding performance in semi-supervised learning. However, applying this framework directly to Pseudo-Labeling STR exhibits unstable convergence, as generating pseudo labels at the coarse-grained sequence-level leads to inefficient utilization of unlabelled data. Furthermore, the inherent domain shift between labeled and unlabeled data results in low quality of derived pseudo labels. To mitigate the above issues, we propose a novel Cross-domain Pseudo-Labeling (CPL) approach for scene text recognition, which makes better utilization of unlabeled data at the character-level and provides more accurate pseudo labels. Specifically, our proposed Pseudo-Labeled Curriculum Learning dynamically adjusts the thresholds for different character classes according to the model's learning status. Moreover, an Adaptive Distribution Regularizer is employed to bridge the domain gap and improve the quality of pseudo labels. Extensive experiments show that CPL boosts those representative STR models to achieve state-of-the-art results on six challenging STR benchmarks. Besides, it can be effectively generalized to handwritten text. Xiaoxue Chen, Zuming Huang, Lele Xie, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 4 |
| 2021 | MatchVIE: Exploiting Match Relevancy between Entities for Visual Information ExtractionabstractVisual Information Extraction (VIE) task aims to extract key information from multifarious document images (e.g., invoices and purchase receipts). Most previous methods treat the VIE task simply as a sequence labeling problem or classification problem, which requires models to carefully identify each kind of semantics by introducing multimodal features, such as font, color, layout. But simply introducing multimodal features can't work well when faced with numeric semantic categories or some ambiguous texts. To address this issue, in this paper we propose a novel key-value matching model based on a graph neural network for VIE (MatchVIE). Through key-value matching based on relevancy evaluation, the proposed MatchVIE can bypass the recognitions to various semantics, and simply focuses on the strong relevancy between entities. Besides, we introduce a simple but effective operation, Num2Vec, to tackle the instability of encoded values, which helps model converge more smoothly. Comprehensive experiments demonstrate that the proposed MatchVIE can significantly outperform previous methods. Notably, to the best of our knowledge, MatchVIE may be the first attempt to tackle the VIE task by modeling the relevancy between keys and values and it is a good complement to the existing methods. Guozhi Tang, Lele Xie, Jingdong Chen, Qianying Wang 0002, Yaqiang Wu |
IJCAI | 2 |
| 2019 | DeRPN: Taking a Further Step toward More General Object DetectionabstractMost current detection methods have adopted anchor boxes as regression references. However, the detection performance is sensitive to the setting of the anchor boxes. A proper setting of anchor boxes may vary significantly across different datasets, which severely limits the universality of the detectors. To improve the adaptivity of the detectors, in this paper, we present a novel dimension-decomposition region proposal network (DeRPN) that can perfectly displace the traditional Region Proposal Network (RPN). DeRPN utilizes an anchor string mechanism to independently match object widths and heights, which is conducive to treating variant object shapes. In addition, a novel scale-sensitive loss is designed to address the imbalanced loss computations of different scaled objects, which can avoid the small objects being overwhelmed by larger ones. Comprehensive experiments conducted on both general object detection datasets (Pascal VOC 2007, 2012 and MS COCO) and scene text detection datasets (ICDAR 2013 and COCO-Text) all prove that our DeRPN can significantly outperform RPN. It is worth mentioning that the proposed DeRPN can be employed directly on different models, tasks, and datasets without any modifications of hyperparameters or specialized optimization, which further demonstrates its adaptivity. The code has been released at https://github.com/HCIILAB/DeRPN. Lele Xie, Zecheng Xie |
AAAI | 1 |
| 2019 | Tightness-Aware Evaluation Protocol for Scene Text DetectionabstractEvaluation protocols play key role in the developmental progress of text detection methods. There are strict requirements to ensure that the evaluation methods are fair, objective and reasonable. However, existing metrics exhibit some obvious drawbacks: 1) They are not goal-oriented; 2) they cannot recognize the tightness of detection methods; 3) existing one-to-many and many-to-one solutions involve inherent loopholes and deficiencies. Therefore, this paper proposes a novel evaluation protocol called Tightness-aware Intersect-over-Union (TIoU) metric that could quantify completeness of ground truth, compactness of detection, and tightness of matching degree. Specifically, instead of merely using the IoU value, two common detection behaviors are properly considered; meanwhile, directly using the score of TIoU to recognize the tightness. In addition, we further propose a straightforward method to address the annotation granularity issue, which can fairly evaluate word and text-line detections simultaneously. By adopting the detection results from published methods and general object detection frameworks, comprehensive experiments on ICDAR 2013 and ICDAR 2015 datasets are conducted to compare recent metrics and the proposed TIoU metric. The comparison demonstrated some promising new prospects, e.g., determining the methods and frameworks for which the detection is tighter and more beneficial to recognize. Our method is extremely simple; however, the novelty is none other than the proposed metric can utilize simplest but reasonable improvements to lead to many interesting and insightful prospects and solving most the issues of the previous metrics. The code is publicly available at https://github.com/Yuliang-Liu/TIoU-metric. Zecheng Xie, Canjie Luo, Shuaitao Zhang, Lele Xie |
CVPR | 6 |
| 2019 | Aggregation Cross-Entropy for Sequence RecognitionabstractIn this paper, we propose a novel method, aggregation cross-entropy (ACE), for sequence recognition from a brand new perspective. The ACE loss function exhibits competitive performance to CTC and the attention mechanism, with much quicker implementation (as it involves only four fundamental formulas), faster inference\back-propagation (approximately O(1) in parallel), less storage requirement (no parameter and negligible runtime memory), and convenient employment (by replacing CTC with ACE). Furthermore, the proposed ACE loss function exhibits two noteworthy properties: (1) it can be directly applied for 2D prediction by flattening the 2D prediction into 1D prediction as the input and (2) it requires only characters and their numbers in the sequence annotation for supervision, which allows it to advance beyond sequence recognition, e.g., counting problem. The code is publicly available at https://github.com/summerlvsong/Aggregation-Cross-Entropy. Zecheng Xie, Yaoxiong Huang, Lele Xie |
CVPR | 6 |
| 2019 | Omnidirectional Scene Text Detection with Sequential-free Box DiscretizationabstractScene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection methods. However, recent researches reveal that introducing quadrilateral bounding box for scene text detection will bring a label confusion issue which is easily overlooked, and this issue may significantly undermine the detection performance. To address this issue, in this paper, we propose a novel method called Sequential-free Box Discretization (SBD) by discretizing the bounding box into key edges (KE) which can further derive more effective methods to improve detection performance. Experiments showed that the proposed method can outperform state-of-the-art methods in many popular scene text benchmarks, including ICDAR 2015, MLT, and MSRA-TD500. Ablation study also showed that simply integrating the SBD into Mask R-CNN framework, the detection performance can be substantially improved. Furthermore, an experiment on the general object dataset HRSC2016 (multi-oriented ships) showed that our method can outperform recent state-of-the-art methods by a large margin, demonstrating its powerful generalization ability. Sheng Zhang 0024, Lele Xie, Yaqiang Wu, Zhepeng Wang 0002 |
IJCAI | 4 |
| 2018 | Detecting Heads using Feature Refine Net and Cascaded Multi-scale ArchitectureabstractThis paper presents a method that can accurately detect heads especially small heads under the indoor scene. To achieve this, we propose a novel method, Feature Refine Net (FRN), and a cascaded multi-scale architecture. FRN exploits the multi-scale hierarchical features created by deep convolutional neural networks. The proposed channel weighting method enables FRN to make use of features alternatively and effectively. To improve the performance of small head detection, we propose a cascaded multi-scale architecture which has two detectors. One called global detector is responsible for detecting large objects and acquiring the global distribution information. The other called local detector is designed for small objects detection and makes use of the information provided by global detector. Due to the lack of head detection datasets, we have collected and labeled a new large dataset named SCUT-HEAD which includes 4405 images with 111251 heads annotated. Experiments show that our method has achieved state-of-the-art performance on SCUT-HEAD. Dezhi Peng, Zikai Sun, Zirong Chen, Zirui Cai, Lele Xie |
ICPR | 5 |
| 2018 | A New CNN-Based Method for Multi-Directional Car License Plate DetectionabstractThis paper presents a novel convolutional neural network (CNN) -based method for high-accuracy real-time car license plate detection. Many contemporary methods for car license plate detection are reasonably effective under the specific conditions or strong assumptions only. However, they exhibit poor performance when the assessed car license plate images have a degree of rotation, as a result of manual capture by traffic police or deviation of the camera. Therefore, we propose the a CNN-based MD-YOLO framework for multi-directional car license plate detection. Using accurate rotation angle prediction and a fast intersection-over-union evaluation strategy, our proposed method can elegantly manage rotational problems in real-time scenarios. A series of experiments have been carried out to establish that the proposed method outperforms over other existing state-of-the-art methods in terms of better accuracy and lower computational cost. Lele Xie, Tasweer Ahmad, Sheng Zhang 0024 |
IEEE Trans. Intell. Transp. Syst. | 1 |