VLDB 2026 Research / reviewers in the wild / expert
Qingyong Li
dblp:29/4219
· DBLP profile ↗
67ranked-venue papers
6as first author
38since 2021 · last 2026
0000-0002-3860-4809ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASTAFN: Bridging the Gap Between Weather Foundation Models and Accurate Station-Level ForecastingabstractWeather Foundation Models (WFMs) have recently attracted significant attention for their exceptional performance and inference efficiency in global-scale weather forecasting. However, their coarse spatial resolution and inherent biases constrain their utility for station-level forecasting, which is crucial for applications such as renewable energy management and aviation safety. To address these limitations, we propose the Adaptive Spatiotemporal Alignment Fusion Network (ASTAFN), a novel framework designed for accurate station-level weather forecasting through the synergistic integration of WFMs and station observations. ASTAFN incorporates two complementary data sources: (1) recent station observations, which offer fine-grained local trend information, and (2) WFM-generated forecasts, which provide broad-scale weather patterns. The core innovation of ASTAFN lies in its proxy station learning mechanism, which aligns the spatial structure and corrects the biases of WFMs relative to actual station data, facilitating the extraction of homogeneous spatiotemporal features from both sources. These features are dynamically fused at each forecasting step using an adaptive strategy, effectively compensating for WFM biases and enhancing predictive accuracy. Experimental evaluations on three real-world datasets demonstrate that ASTAFN reduces mean absolute error by 20%–35% compared to baseline WFMs for station-level wind speed forecasting. ASTAFN has been deployed on the regional station-level weather forecasting and analysis platform of the Chinese Academy of Meteorological Sciences, currently serving the Guangdong and Yunnan provinces in southern China. Bihe Xu, Qingyong Li, Zhiqing Guo |
KDD (1) | 3 |
| 2026 | Progressive cross-scale semantic alignment for language-guided medical image segmentation
Hengzhi Xue, Yin Dai, Qingyong Li, Yu-Dong Yao, Yueyang Teng |
Knowl. Based Syst. | 3 |
| 2026 | GSPNet: Graph Spectral Projection Network Using Learnable Spectral TransformationabstractSpectral graph convolutional networks (SGCNs) are one of the leading tools to handle learning tasks with graph structure. SGCNs leverage graph structure to define the graph spectral transformation (i.e., graph Fourier transformation) and pursue a good spectral filter in the spectral domain. For efficiency reasons, the pursued spectral filter is usually approximated by a polynomial of the normalized adjacency matrix in the spatial domain, referred to as the polynomial propagation of SGCNs. In this paper, we present a theoretical analysis on the propagation matrix defined by SGCNs and show that it actually almost always incurs a non-trivial propagation error. We also derive an explicit bound for this error. To mitigate this issue, we propose the Graph Spectral Projection Network (GSPNet) by learning a better spectral transformation, which endows GSPNet with potential to eliminate the error encountered by the polynomial propagation of SGCNs. Experimental results on commonly-used graph datasets suggest that GSPNet can learn a propagation matrix beyond the theoretically optimal polynomial propagation and generate promising results over existing SGCN models. Yuxiao Dong, Wenzheng Feng, Qingyong Li, Jie Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Towards plastic and stable incremental learning: A dual-learner framework with cumulative parameter averaging
Wenju Sun, Qingyong Li, Wen Wang 0019 |
Pattern Recognit. | 2 |
| 2026 | DIDLM: A SLAM Dataset for Difficult Scenarios Featuring Infrared, Depth Cameras, LiDAR, 4D Radar, and Others Under Adverse Weather, Low Light Conditions, and Rough RoadsabstractAdverse weather conditions, low-light environments, and bumpy road surfaces pose significant challenges to SLAM in robotic navigation and autonomous driving. Existing datasets in this field predominantly rely on single sensors or combinations of LiDAR, cameras, and IMUs. However, 4D millimeter-wave radar demonstrates robustness in adverse weather, infrared cameras excel in capturing details under low-light conditions, and depth images provide richer spatial information. Multi-sensor fusion methods also show potential for better adaptation to bumpy roads. Despite some SLAM studies incorporating these sensors and conditions, there remains a lack of comprehensive datasets addressing low-light environments and bumpy road conditions, or featuring a sufficiently diverse range of sensor data. In this study, we introduce a multi-sensor dataset covering challenging scenarios such as snowy weather, rainy weather, nighttime conditions, speed bumps, and rough terrains. The dataset includes rarely utilized sensors for extreme conditions, such as 4D millimeter-wave radar, infrared cameras, and depth cameras, alongside 3D LiDAR, RGB cameras, GPS, and IMU. It supports both autonomous driving and ground robot applications and provides reliable GPS/INS ground truth data, covering structured and semi-structured terrains. We evaluated various SLAM algorithms using this dataset, including RGB images, infrared images, depth images, LiDAR, and 4D millimeter-wave radar. The dataset spans a total of 18.5 km, 69 minutes, and approximately 660 GB, offering a valuable resource for advancing SLAM research under complex and extreme conditions. Our dataset is available athttps://gongweisheng.github.io/DIDLM.github.io/ Weisheng Gong, Chen He 0002, Kaijie Su, Qingyong Li, Z. Jane Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingabstractMulti-task model merging offers a promising paradigm for integrating multiple expert models into a unified system without additional training. Existing state-of-the-art techniques, such as Task Arithmetic and its variants, merge models by accumulating task vectors—defined as the parameter differences between pre-trained and fine-tuned models. However, task vector accumulation is often hindered by knowledge conflicts, where conflicting components across different task vectors can lead to performance degradation during the merging process. To address this challenge, we propose Conflict-Aware Task Merging (CAT Merging), a novel training-free framework that selectively trims conflict-prone components from the task vectors. CAT Merging introduces several parameter-specific strategies, including projection for linear weights and masking for scaling and shifting parameters in normalization layers. Extensive experiments on vision and vision-language tasks demonstrate that CAT Merging effectively suppresses knowledge conflicts, achieving average accuracy improvements of up to 4.7% (ViT-B/32) and 2.0% (ViT-L/14) over state-of-the-art methods. Wenju Sun, Qingyong Li, Boyang Li 0001 |
ICML | 2 |
| 2025 | Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsabstractMulti-task model merging offers an efficient solution for integrating knowledge from multiple fine-tuned models, mitigating the significant computational and storage demands associated with multi-task training. As a key technique in this field, Task Arithmetic (TA) defines task vectors by subtracting the pre-trained model (0 pre) from the fine-tuned task models in parameter space, then adjusting the weight between these task vectors and 0 pre to balance task-generalized and task-specific knowledge. Despite the promising performance of TA, conflicts can arise among the task vectors, particularly when different tasks require distinct model adaptations. In this paper, we formally define this issue as knowledge conflicts, characterized by the performance degradation of one task after merging with a model fine-tuned for another task. Through in-depth analysis, we show that these conflicts stem primarily from the components of task vectors that align with the gradient of task-specific losses at 0 pre. To address this, we propose Task Arithmetic in Trust Region (TATR), which defines the trust region as dimensions in the model parameter space that cause only small changes (corresponding to the task vector components with gradient orthogonal direction) in the task-specific losses. Restricting parameter merging within this trust region, TATR can effectively alleviate knowledge conflicts. Moreover, TATR serves as a plug-and-play module compatible with a wide range of TA-based methods. Extensive empirical evaluations on visual and visual-language tasks robustly demonstrate that TATR improves the multi-task performance of several TA-based model merging methods. Wenju Sun, Qingyong Li, Wen Wang 0019, Boyang Li 0001 |
ACM Multimedia | 2 |
| 2025 | Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationabstractMulti-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing methods primarily approach this by minimizing differences between task-specific experts and the unified model, either from a parameter-level or a task-loss perspective. However, parameter-level methods exhibit a significant performance gap compared to the upper bound, while task-loss approaches entail costly secondary training procedures. In contrast, we observe that performance degradation closely correlates with feature drift, i.e., differences in feature representations of the same sample caused by model merging. Motivated by this observation, we propose Layer-wise Optimal Task Vector Merging (LOT Merging), a technique that explicitly minimizes feature drift between task-specific experts and the unified model in a layer-by-layer manner. LOT Merging can be formulated as a convex quadratic optimization problem, enabling us to analytically derive closed-form solutions for the parameters of linear and normalization layers. Consequently, LOT Merging achieves efficient model consolidation through basic matrix operations. Extensive experiments across vision and vision-language benchmarks demonstrate that LOT Merging significantly outperforms baseline methods, achieving improvements of up to 4.4% (ViT-B/32) over state-of-the-art approaches. The source code is available at https://github.com/SunWenJu123/model-merging. Wenju Sun, Qingyong Li, Wen Wang 0019, Yang Liu 0352, Boyang Li 0001 |
NeurIPS | 2 |
| 2025 | Decoupled likelihood modeling: A scalable approach for incremental generalized category discovery
Wenju Sun, Qingyong Li, Wen Wang 0019 |
Neurocomputing | 4 |
| 2025 | Hard-Label Black-Box Adversarial Attacks for Implicit Scene Interactions
Muxue Liang, Chuan Wang 0002, Siyuan Liang 0004, Aishan Liu, Yanan Cao 0006, Qingyong Li, Zeming Liu, Liang Yang 0002, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | Incremental Learning via Robust Parameter Posterior FusionabstractThe posterior estimation of parameters based on Bayesian theory is a crucial technique in Incremental Learning (IL). The estimated posterior is typically utilized to impose loss regularization, which aligns the current training model parameters with the previously learned posterior to mitigate catastrophic forgetting, a major challenge in IL. However, this additional loss regularization can also impose detriment to the model learning, preventing it from reaching the true global optimum. To overcome this limitation, this paper introduces a novel Bayesian IL framework, Robust Parameter Posterior Fusion (RP2F). Unlike traditional methods, RP2F directly estimates the parameter posterior for new data without introducing extra loss regularization, which allows the model to accommodate new knowledge more sufficiently. It then fuses this new posterior with the existing ones based on the Maximum A Posteriori (MAP) principle, ensuring effective knowledge sharing across tasks. Furthermore, RP2F incorporates a common parameter-robustness priori to facilitate a seamless integration during posterior fusion. Comprehensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets show that RP2F not only effectively mitigates catastrophic forgetting but also achieves backward knowledge transfer. Wenju Sun, Qingyong Li, Wen Wang 0019 |
ACM Multimedia | 2 |
| 2024 | Disentangling clusters from non-Euclidean data via graph frequency reorganization
Chong-Yung Chi, Wenju Sun, Jing Zhang 0058, Qingyong Li |
Inf. Sci. | 5 |
| 2024 | Spatial-temporal graph-guided global attention network for video-based person re-identification
Xiaobao Li, Wen Wang 0019, Qingyong Li |
Mach. Vis. Appl. | 3 |
| 2024 | CLDiff: Weakly Supervised Cloud Detection With Denoising Diffusion Probabilistic ModelsabstractCloud detection is an essential step in remote sensing (RS) image processing, contributing to various applications. However, existing fully supervised cloud detection methods rely on massive pixel-wise annotations, which are expensive and time-consuming. To alleviate the annotation burden, weakly supervised cloud detection (WSCD) has received extensive attention recently. One standard approach performs cloud detection within a classification paradigm, which inevitably faces category ambiguity when detecting semitransparent clouds. To tackle this problem, we propose a novel WSCD framework based on the diffusion model, termed CLDiff. Specifically, a multiscale feature rectification (MFR) module is introduced to extract multiscale semantic features in the encoder, enabling a definite identification of clouds and mitigating interference from bright objects in the background. Considering that clouds exhibit varying optical thicknesses, a diffusion decoder is developed to model the intraclass variations of clouds in a generative strategy, improving thin cloud detection. Initially, it devises a Gaussian modulation function to recalibrate ambiguous cloud activations and emphasize semitransparent clouds. Subsequently, these modulated activations serve as semantic guidance to optimize the diffusion process. This approach enables CLDiff to activate cloud contours under definite semantic conditions and avoids the additional branches for semantic learning as found in previous methods. Experimental results demonstrate that CLDiff achieves state-of-the-art performance in WSCD. A public reference implementation of this work in PyTorch is available athttps://github.com/YLiu-creator/CLDiff. Yang Liu 0352, Qingyong Li, Zhigang Yao, Tony Z. Qiu, Wen Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Causal Discovery on Discrete Data via Weighted Normalized Wasserstein DistanceabstractThe task of causal discovery from observational data (X,Y) is defined as the task of deciding whether X causes Y , or Y causes X or if there is no causal relationship between X and Y . Causal discovery from observational data is an important problem in many areas of science. In this study, we propose a method to address this problem when the cause-and-effect relationship is represented by a discrete additive noise model (ANM). First, assuming that X causes Y , we estimate the conditional distributions of the noise given X using regression. Similarly, assuming that Y causes X , we also estimate the conditional distributions of noise given Y . Based on the structural characteristics of the discrete ANM, we find that the dissimilarity of the conditional distributions of noise in the causal direction is smaller than that in the anticausal direction. Then, we propose a weighted normalized Wasserstein distance to measure the dissimilarity of the conditional distributions of noise. Finally, we propose a decision rule for casual discovery by comparing two computed weighted normalized Wasserstein distances. An empirical investigation demonstrates that our method performs well on synthetic data and outperforms state-of-the-art methods on real data. Li-Hui Lin, Dengming Zhu, Qingyong Li |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Decoupling Learning and Remembering: a Bilevel Memory Framework with Knowledge Projection for Task-Incremental LearningabstractThe dilemma between plasticity and stability arises as a common challenge for incremental learning. In contrast, the human memory system is able to remedy this dilemma owing to its multilevel memory structure, which motivates us to propose a Bilevel Memory system with Knowledge Projection (BMKP) for incremental learning. BMKP decouples the functions of learning and remembering via a bilevel-memory design: a working memory responsible for adaptively model learning, to ensure plasticity; a long-term memory in charge of enduringly storing the knowledge incorporated within the learned model, to guarantee stability. However, an emerging issue is how to extract the learned knowledge from the working memory and assimilate it into the long-term memory. To approach this issue, we reveal that the parameters learned by the working memory are actually residing in a redundant high-dimensional space, and the knowledge incorporated in the model can have a quite compact representation under a group of pattern basis shared by all incremental learning tasks. Therefore, we propose a knowledge projection process to adaptively maintain the shared basis, with which the loosely organized model knowledge of working memory is projected into the compact representation to be remembered in the long-term memory. We evaluate BMKP on CIFAR-10, CIFAR-100, and Tiny-ImageNet. The experimental results show that BMKP achieves state-of-the-art performance with lower memory usage11The code is available at https://github.com/SunWenJu123/BMKP. Wenju Sun, Qingyong Li, Jing Zhang 0058, Wen Wang 0019 |
CVPR | 2 |
| 2023 | Confidence-adapted meta-interaction for unsupervised person re-identification
Xiaobao Li, Qingyong Li, Wenyuan Xue, Yang Liu 0352, Fengjiao Liang, Wen Wang 0019 |
Appl. Intell. | 2 |
| 2023 | Multi-granularity Pseudo-label Collaboration for unsupervised person re-identification
Xiaobao Li, Qingyong Li, Fengjiao Liang, Wen Wang 0019 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Class Incremental Learning based on Identically Distributed Parallel One-Class Classifiers
Wenju Sun, Qingyong Li, Jing Zhang 0058, Wen Wang 0019 |
Neurocomputing | 2 |
| 2023 | A General Dual-Branch Framework for Land Cover Mapping Models With Multispectral DataabstractLand cover mapping based on multispectral images can, in principle, be considered an application of semantic segmentation, but land cover mapping inputs include near-infrared (NIR) data in addition to RGB data. It has been experimentally found that LULC mapping performance based on RGB data alone is better than that based on RGB and NIR data when using some established single-branch encoder–decoder models. To address this issue, we propose a dual-branch encoder–decoder (DBED) framework that can be applied to existing encoder–decoder models for semantic segmentation. First, the multispectral input data is divided into two parts: RGB and NIR and fed to the respective branch for encoding. The dual-branch structure facilitates cross-modal complementary information encoding without deteriorating the original RGB modality-specific feature extraction. Second, an attention module named the multispectral attention module (MSAM) is proposed to mine the contextual correlation between the multispectral feature maps, leading to further performance boosting. We apply this framework to three mainstream semantic segmentation models and validate it on the Gaofen Image Dataset (GID). Experimental results show that this structure brings performance improvements. The source code of DBED is publicly available athttps://github.com/MalignusCN/DBED. Qingyong Li, Yang Liu 0352, Wen Wang 0019 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Exemplar-free class incremental learning via discriminative and comparable parallel one-class classifiers
Wenju Sun, Qingyong Li, Jing Zhang 0058, Danyu Wang, Wen Wang 0019 |
Pattern Recognit. | 2 |
| 2023 | Leveraging Physical Rules for Weakly Supervised Cloud Detection in Remote Sensing ImagesabstractCloud detection plays a significant role in remote sensing image applications. Existing deep learning-based cloud detection methods rely on massive precise pixel-wise annotations, which are time-consuming and expensive. To alleviate this problem, we propose a weakly supervised cloud detection framework that leverages physical rules to generate weak supervision for cloud detection in remote sensing images. Specifically, a rule-based adaptive pseudo labeling (RAPL) algorithm is devised to adaptively annotate potential cloud pixels based on cloud spectral properties without manual intervention. Unlike existing physical annotations using fixed thresholds, RAPL employs the bidirectional threshold segmentation and adaptive gating mechanism to annotate cloud and boundary masks with more explicit semantic categories and spatial structures separately. Subsequently, these pseudo masks are treated as weak supervision to optimize the heuristic cloud detection network for pixel-wise segmentation. Considering that clouds appear as complex geometric structures and nonuniform spectral reflectance, a deformable boundary refining module is designed to enhance the modeling ability of spatial transformation and activate sharp boundaries from translucent cloud regions. Moreover, a harmonic loss is employed to recognize clouds with nonuniform spectral reflectance and suppress the interference of bright backgrounds. Extensive experiments on the GF-1, L8 Biome, and WDCD datasets demonstrate that the proposed method achieves state-of-the-art results. A public reference implementation of this work in PyTorch is available at https://github.com/NiAn-creator/HeuristicCloudDetection. Yang Liu 0352, Qingyong Li, Xiaobao Li, Shuyi He, Fengjiao Liang, Zhigang Yao, Wen Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CGF: A Category Guidance Based PM$_{2.5}$ Sequence Forecasting Training FrameworkabstractPM$_{2.5}$concentration forecasting is important yet challenging. First, complicated local fluctuations in PM$_{2.5}$concentrations disturb modeling global trends. Second, forecasting errors are often accumulated through an autoregressive process. To contend with the two challenges, we propose aCategoryGuidance based PM${_{2.5}}$sequenceForecasting training framework (CGF) to enhance the performance of existing PM${_{2.5}}$concentration forecasting models. CGF contains a Category based Representation Learning (CRL) module and a Category based Self-paced Learning (CSL) module, both of which utilize PM${_{2.5}}$category information that is easily obtained and publicly available. First, CRL employs category information to guide forecasting models to produce more robust hidden representations that are insensitive to local fluctuations, thus alleviating the negative impact of local fluctuations. Second, CSL adaptively selects real PM${_{2.5}}$concentration values versus autoregressive PM${_{2.5}}$forecast values when training forecasting models, helping alleviate error accumulations. The CGF framework is applied to existing PM${_{2.5}}$forecasting models, and the experimental results on two real-world datasets demonstrate that CGF is able to consistently improve the accuracy of existing forecasting models. Furthermore, to validate the generality of CGF, we conduct extensional experiments in two other time-series prediction tasks, including exchange rate forecasting and electricity forecasting. The experimental results also verify the effectiveness of CGF. Haomin Yu, Jilin Hu, Xinyuan Zhou, Chenjuan Guo, Bin Yang 0002, Qingyong Li |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | LightNet+: A dual-source lightning forecasting network with bi-direction spatiotemporal transformation
Xinyuan Zhou, Haomin Yu, Qingyong Li, Liangtao Xu, Yijun Zhang 0002 |
Appl. Intell. | 4 |
| 2022 | Distribution and gradient constrained embedding model for zero-shot learning with fewer seen samples
Jing Zhang 0058, Wen Wang 0019, Wenju Sun, Zhirong Yang, Qingyong Li |
Knowl. Based Syst. | 6 |
| 2022 | Semantic Segmentation of Remote Sensing Images With Self-Supervised Semantic-Aware InpaintingabstractSemantic segmentation of remote sensing imageries plays a crucial role in resource exploration, urban planning, weather forecasting, etc. For this task, deep learning-based methods have shown significant achievement, typically trained with large-scale labeled data. However, these methods often suffer the performance deterioration facing limited labeled data in real-world applications. To address this problem, a novel self-supervised semantic segmentation framework is proposed for remote sensing imageries with limited labeled data. Specifically, image inpainting is acted as pixel-level pretext task for learning dense feature representations suitable for semantic segmentation. Further, rather than trivially leveraging the conventional random inpainting strategy, a novel adversarial training scheme is proposed to drive the pretext task to adaptively mask and restore salient local regions. The adversarial training scheme consists of instructor network and inpainting network, the instructor network increasingly predicts meaningful salient regions as erased regions, and meanwhile the inpainting network seeks for restoring the corrupted image as pretext task to learn its intrinsic representation. Moreover, the structural similarity (SSIM) is applied as a patch-level loss function for semantic segmentation considering that remote sensing images are highly structured. The experimental results on the ISPRS Potsdam dataset demonstrate that our method outperforms state-of-the-art self-supervised methods and the ImageNet pre-training methods. The source code is available at https://github.com/JasmineBJTU/self-supervised_RSSS. Shuyi He, Qingyong Li, Yang Liu 0352, Wen Wang 0019 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | DCNet: A Deformable Convolutional Cloud Detection Network for Remote Sensing ImageryabstractRecently, deep convolutional neural networks (CNNs) have made important progress in cloud detection with powerful representation learning capability and yield significant performance. However, most existing CNN-based cloud detection methods still face serious challenges because of the variable geometry of clouds and the complexity of underlying surfaces. It is attributed that they only use the fixed grid to extract contextual information, which lacks internal mechanisms to handle the geometric transformations of clouds. To tackle this problem, we propose a deformable convolutional cloud detection network with an encoder-decoder architecture, named DCNet, which can enhance the adaptability of a model to cloud variations. Specifically, we introduce deformable convolution blocks at the encoder to capture saliency spatial contexts adaptively based on the morphological characteristics of clouds and generate high-level semantic representations. After this, we incorporate skip-connection mechanisms into the decoder that integrate low-level spatial contexts as guidance to recover high-level semantic pixel localization and export precise cloud-detection results. Extensive experiments on the GF-1 wide field-of-view (WFV) Satellite Imagery demonstrate that DCNet outperforms several state-of-the-art methods. A public reference implementation of our proposed model in PyTorch is available athttps://github.com/NiAn-creator/deformableCloudDetection.git. Yang Liu 0352, Wen Wang 0019, Qingyong Li, Min Min, Zhigang Yao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A zero-shot learning framework via cluster-prototype matching
Jing Zhang 0058, Qingyong Li, Wen Wang 0019, Wenju Sun, Chuan Shi 0001, Zhengming Ding |
Pattern Recognit. | 2 |
| 2022 | An Unsupervised Multi-Shot Person Re-Identification Method via Mutual Normalized Sparse Representation and Stepwise LearningabstractDue to abundant prior information and widespread applications, multi-shot based person re-identification has drawn increasing attention in recent years. In this paper, the high labeling cost and huge unlabeled data motivate us to focus on the unsupervised scenario and a unified coarse-to-fine framework is proposed, named by Mutual Normalized Sparse Representation (MNSR). Our method is an iteration procedure and each iteration involves two key steps: label estimation and metric model learning. In the former, we present a MNSR model to infer the pairwise labels of cross-camera by endowing sparse representation coefficient with the probability property. MNSR explicitly takes the mutually correlation between cameras into consideration and thus produces more accurate results. Meanwhile, we propose a probability-guided positive pairwise label prediction method to mine hard positive samples. For the latter, we learn a metric model with the estimated pairwise labels as supervision. In this procedure, we select some reliable labels for training by configuring with a stepwise learning method, rather than use all the estimated pair samples. This procedure helps to prevent the noise samples damaging the learning of discriminative metric model, especially for the initial iterations. Extensive experiments are conducted on four publicly available datasets, including PRID 2011, iLIDS-VID, SAIVT-SoftBio and MARS, and the results demonstrate the superior performance of the MNSR method in comparison with state-of-the-art unsupervised multi-shot person re-identification methods. Xiaobao Li, Qingyong Li, Wen Wang 0019, Lijun Guo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | DDGNet: A Dual-Stage Dynamic Spatio-Temporal Graph Network for PM2.5 ForecastingabstractAs air pollution problems become increasingly serious, PM2.5forecasting based on spatio-temporal observation data has received widespread attention. This forecasting task is full of challenges given the complicated producing factors and fickle transmission process of PM2.5. However, most existing forecasting methods only exploit the spatial dependency by graph networks with fixed adjacency matrices, ignoring the dynamic spatio-temporal correlation of PM2.5concentrations. In this paper, we propose a dual-stage dynamic spatio-temporal graph network (DDGNet) to model dynamic correlations for PM2.5prediction of different cities. Specifically, DDGNet consists of two major stages: (1) dynamic graph construction to identify potentially informative neighbors for each node (a city) in every forecasting period; (2) graph attention networks to dynamically determine linking weights for each vertex to its neighbors. We evaluate DDGNet on three real-world datasets and compare it with several baselines. The experimental results demonstrate that our method achieves the state-of-the-art performance. Haomin Yu, Xiaobao Li, Qingyong Li |
IEEE BigData | 5 |
| 2021 | TGRNet: A Table Graph Reconstruction Network for Table Structure RecognitionabstractA table arranging data in rows and columns is a very effective data structure, which has been widely used in business and scientific research. Considering large-scale tabular data in online and offline documents, automatic table recognition has attracted increasing attention from the document analysis community. Though human can easily understand the structure of tables, it remains a challenge for machines to understand that, especially due to a variety of different table layouts and styles. Existing methods usually model a table as either the markup sequence or the adjacency matrix between different table cells, failing to address the importance of the logical location of table cells, e.g., a cell is located in the first row and the second column of the table. In this paper, we reformulate the problem of table structure recognition as the table graph reconstruction, and propose an end-to-end trainable table graph reconstruction network (TGRNet) for table structure recognition. Specifically, the proposed method has two main branches, a cell detection branch and a cell logical location branch, to jointly predict the spatial location and the logical location of different cells. Experimental results on three popular table recognition datasets and a new dataset with table graph annotations (TableGraph-350K) demonstrate the effectiveness of the proposed TGRNet for table structure recognition. Code and annotations will be made publicly available at https://github.com/xuewenyuan/TGRNet. Wenyuan Xue, Baosheng Yu, Wen Wang 0019, Dacheng Tao, Qingyong Li |
ICCV | 5 |
| 2021 | HCADecoder: A Hybrid CTC-Attention Decoder for Chinese Text Recognition
Wenyuan Xue, Qingyong Li, Peng Zhao 0014 |
ICDAR (3) | 3 |
| 2021 | A Two-Stage Autoencoder For Visual Anomaly Detection
Yezhou Zhu, Jianzhu Wang, Jing Zhang 0058, Qingyong Li |
ICIP | 4 |
| 2021 | A Question Answering System for Unstructured Table ImagesabstractQuestion answering over tables is a very popular semantic parsing task in natural language processing (NLP). However, few existing methods focus on table images, even though there are usually large-scale unstructured tables in practice (e.g., table images). Table parsing from images is nontrivial since it is closely related to not only NLP but also computer vision (CV) to parse the tabular structure from an image. In this demo, we present a question answering system for unstructured table images. The proposed system mainly consists of 1) a table recognizer to recognize the tabular structure from an image and 2) a table parser to generate the answer to a natural language question over the table. In addition, to train the model, we further provide table images and structure annotations for two widely used semantic parsing datasets. Specifically, the test set is used for this demo, from where the users can either choose from default questions or enter a new custom question. Wenyuan Xue, Wen Wang 0019, Qingyong Li, Baosheng Yu, Yibing Zhan, Dacheng Tao |
ACM Multimedia | 4 |
| 2021 | Unsupervised Multi-shot Person Re-identification via Dynamic Bi-directional Normalized Sparse Representation
Xiaobao Li, Wen Wang 0019, Qingyong Li, Lijun Guo |
MMM (1) | 3 |
| 2021 | From Digital Model to Reality Application: A Domain Adaptation Method for Rail Defect Detection
Wenkai Cui, Jianzhu Wang, Haomin Yu, Wenjuan Peng, Qingyong Li |
PRCV (2) | 8 |
| 2021 | LRGAN: Visual anomaly detection using GAN with locality-preferred recoding
Jianzhu Wang, Qingyong Li |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Fine-Grained Image-Text Retrieval via Discriminative Latent Space LearningabstractFine-grained image-text retrieval aims at searching relevant images among fine-grained classes given a text query or in a reverse way. The challenges are not only bridging the gap between two heterogeneous modalities but also dealing with large inter-class similarity and intra-class variance existed in fine-grained data. To deal with the above challenges, we propose a Discriminative Latent Space Learning (DLSL) method for fine-grained image-text retrieval. Concretely, image and text features are extracted for capturing the subtle difference in fine-grained data. Subsequently, based on the extracted features, we perform couple dictionary learning to align the heterogeneous data in a uniform latent space. To make such alignment discriminative enough for the fine-grained task, the learned latent space is endowed with discriminative property via learning a discriminative map. Comprehensive experiments on fine-grained datasets demonstrate the effectiveness of our approach. Wen Wang 0019, Qingyong Li |
IEEE Signal Process. Lett. | 3 |
| 2020 | Multi-Component Graph Convolutional Collaborative FilteringabstractThe interactions of users and items in recommender system could be naturally modeled as a user-item bipartite graph. In recent years, we have witnessed an emerging research effort in exploring user-item graph for collaborative filtering methods. Nevertheless, the formation of user-item interactions typically arises from highly complex latent purchasing motivations, such as high cost performance or eye-catching appearance, which are indistinguishably represented by the edges. The existing approaches still remain the differences between various purchasing motivations unexplored, rendering the inability to capture fine-grained user preference. Therefore, in this paper we propose a novel Multi-Component graph convolutional Collaborative Filtering (MCCF) approach to distinguish the latent purchasing motivations underneath the observed explicit user-item interactions. Specifically, there are two elaborately designed modules, decomposer and combiner, inside MCCF. The former first decomposes the edges in user-item graph to identify the latent components that may cause the purchasing relationship; the latter then recombines these latent components automatically to obtain unified embeddings for prediction. Furthermore, the sparse regularizer and weighted random sample strategy are utilized to alleviate the overfitting problem and accelerate the optimization. Empirical results on three real datasets and a synthetic dataset not only show the significant performance gains of MCCF, but also well demonstrate the necessity of considering multiple components. Xiao Wang 0017, Chuan Shi 0001, Guojie Song, Qingyong Li |
AAAI | 5 |
| 2020 | AirNet: A Calibration Model for Low-Cost Air Monitoring Sensors Using Dual Sequence Encoder NetworksabstractAir pollution monitoring has attracted much attention in recent years. However, accurate and high-resolution monitoring of atmospheric pollution remains challenging. There are two types of devices for air pollution monitoring, i.e., static stations and mobile stations. Static stations can provide accurate pollution measurements but their spatial distribution is sparse because of their high expense. In contrast, mobile stations offer an effective solution for dense placement by utilizing low-cost air monitoring sensors, whereas their measurements are less accurate. In this work, we propose a data-driven model based on deep neural networks, referred to as AirNet, for calibrating low-cost air monitoring sensors. Unlike traditional methods, which treat the calibration task as a point-to-point regression problem, we model it as a sequence-to-point mapping problem by introducing historical data sequences from both a mobile station (to be calibrated) and the referred static station. Specifically, AirNet first extracts an observation trend feature of the mobile station and a reference trend feature of the static station via dual encoder neural networks. Then, a social-based guidance mechanism is designed to select periodic and adjacent features. Finally, the features are fused and fed into a decoder to obtain a calibrated measurement. We evaluate the proposed method on two real-world datasets and compare it with six baselines. The experimental results demonstrate that our method yields the best performance. Haomin Yu, Qingyong Li, Zhi Wei 0001 |
AAAI | 2 |
| 2020 | Query by Strings and Return Ranking Word Regions with Only One Look
Peng Zhao 0014, Wenyuan Xue, Qingyong Li |
ACCV (6) | 3 |
| 2020 | EvaNet: An Extreme Value Attention Network for Long-Term Air Quality PredictionabstractAir quality affects social activities and human health. Air quality prediction, especially for extreme events such as severe haze pollution, plays an essential guiding role in government decision-making and outdoor activity scheduling. Established prediction models face the challenges of forecasting extreme values and long-term tendency. In this paper, we propose an extreme value attention network (EvaNet) based on encoder and decoder framework to achieve long-term air quality prediction. This model designs an extreme value attention mechanism to alleviate the impact of sudden changes on prediction. In addition, to capture long-term dependence relationships, EvaNet introduces a temporal attention mechanism. Integrating the dual attention mechanisms, the extracted features are fed into a decoder to yield the final prediction. The experiments evaluated on two real-world air quality datasets show the superiority of our method against other state-of-the-art baselines. Zechuan Chen, Haomin Yu, Qingyong Li |
IEEE BigData | 4 |
| 2020 | More Than One: A Cluster-Prototype Matching Framework for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize unseen categories whose data is unavailable during the training stage. Most existing ZSL algorithms focus on learning an embedding space and determine the classes of test samples according to sample-prototype similarities in this space. However, we observe that, in contrast to the single sample-prototype relationship, an ensemble criterion usually benefits the final classification, just as the saying "more than one". Inspired by this, we introduce a novel cluster-prototype matching (CPM) strategy and propose a ZSL framework based on CPM. Firstly, we learn a mapping between the visual space and the semantic space utilizing a well-established ZSL algorithm. Via the learned mapping, all test samples are projected into the embedding space and clustered in this space. Secondly, two CPM methods, soft-CPM and hard-CPM, are proposed to match clusters and class prototypes, along with cluster-prototype similarities calculated. Finally, the label of each sample is determined by the combination of the sample-prototype similarity and the cluster-prototype similarity. We apply our framework to five basic ZSL methods and compare them with several advanced baselines of ZSL. The experimental results demonstrate that the proposed framework can significantly improve the performance of the basic ZSL models and help them achieve or beyond the state-of-the-art. Jing Zhang 0058, Qingyong Li, Chuan Shi 0001 |
CIKM | 3 |
| 2020 | A Heterogeneous Spatiotemporal Network for Lightning PredictionabstractLightning prediction is a complicated and challenging task requiring meteorologists to integrate information from multiple data sources to make decisions. Although some data-driven models have been proposed to make prediction automatically, most of them are based on a single data source or several basically-homogeneous data sources, making them hard to adapt to complex and diverse data in practice. In this work, we propose a heterogeneous spatiotemporal network (HSTN) for lightning prediction, aiming at mining knowledge from several heterogeneous spatiotemporal (ST) data sources. Specifically, HSTN comprises three modules: Gaussian diffusion module, ST encoder and ST decoder. Noting that most of meteorological data can be formatted into either a dense ST tensor or a sparse ST tensor, the ST encoder, with the help of the Gaussian diffusion module, is designed to extract information from both two types of tensors. On the other hand, ST decoder is responsible for merging all information from the other modules and generate the final prediction. By organically combining the three modules, HSTN can handle complex input with heterogeneity in both space and time domains. We conduct experimental evaluations on a real-world lightning dataset. The results demonstrate that HSTN achieves state-of-the-art performance compared with several established baselines. Qingyong Li, Tianyang Lin, Jing Zhang 0058, Liangtao Xu, Weitao Lyu, Heng Huang 0001 |
ICDM | 2 |
| 2020 | Surface Defect Detection via Entity Sparsity Pursuit With Intrinsic PriorsabstractComputer vision based methods have been widely used in surface defect inspection. However, most of these approaches are task specific, and it is hard to transfer them to similar detection scenarios. This paper proposes an entity sparsity pursuit (ESP) method to identify surface defects. Based on the observation that surface image textures usually form a low-rank structure and the structure can be violated by the presence of rare defects, we formulate the detection task as a low-rank and ESP problem. To alleviate the feature shortage issue existed in industrial gray-scale images, we customize a kind of intuitive features for surface defect inspection. Different from previous work utilizing complicated regularization terms, we resort to mine intrinsic priors of defect images, which can be neatly incorporated into the designed architecture. The proposed model is compact and able to detect surface defects in an unsupervised manner. To fully evaluate the presented method, we conduct a series of experiments using three real-world and one synthetic defect datasets. Experimental results demonstrate that ESP outperforms state-of-the-art methods. Jianzhu Wang, Qingyong Li, Jinrui Gan, Haomin Yu |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | Local-Density Subspace Distributed Clustering for High-Dimensional DataabstractDistributed clustering is emerging along with the advent of the era of big data. However, most existing established distributed clustering methods focus on problems caused by a large amount of data rather than caused by the large dimension of data. Consequently, they suffer the “curse” of dimensionality (e.g., poor performance and heavy network overhead) when high-dimensional (HD) data are clustered. In this article, we propose a distributed algorithm, referred to as Local Density Subspace Distributed Clustering (LDSDC) algorithm, to cluster large-scale HD data, motivated by the idea that a local dense region of a HD dataset is usually distributed in a low-dimensional (LD) subspace. LDSDC follows a local-global-local processing structure, including grouping of local dense regions (atom clusters) followed by subspace Gaussian model (SGM) fitting (flexible and scalable to data dimension) at each sub-site, merging of atom clusters at every sub-site according to the merging result broadcast from the global site. Moreover, we propose a fast method to estimate the parameters of SGM for HD data, together with its convergence proof. We evaluate LDSDC on both synthetic and real datasets and compare it with four state-of-the-art methods. The experimental results demonstrate that the proposed LDSDC yields best overall performance. Qingyong Li, Mingfei Liang, Chong-Yung Chi, Juan Tan, Heng Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Online Rail Surface Inspection Utilizing Spatial Consistency and ContinuityabstractRail surface inspection using visual inspection system is an important part of railway maintenance. However, accurate and efficient identification of possible defects remains challenging. This paper proposes a background-oriented defect inspector (BODI) to improve defect detection by considering specified characteristics of the track during inspection. Reformulating the inspection task in this manner offers a new way to model rail surface images. More specifically, BODI features a random sampling stage to obtain a compact background representation without any prior information. A sufficient number of random selections generates adequate and diverse background statistics, and defect-determination and a fusion of procedures then determine whether current pixel belongs to the background. Finally, a background update mechanism and parallelism ensure real-time applicability. The proposed BODI is evaluated on a working railway line. The experimental results demonstrate that it outperforms state-of-the-art methods. Jinrui Gan, Jianzhu Wang, Haomin Yu, Qingyong Li, Zhi-Ping Shi 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | ReS2TIM: Reconstruct Syntactic Structures from Table ImagesabstractTables often represent densely packed but structured data. Understanding table semantics is vital for effective information retrieval and data mining. Unlike web tables, whose semantics are readable directly from markup language and contents, the full analysis of tables published as images requires the conversion of discrete data into structured information. This paper presents a novel framework to convert a table image into its syntactic representation through the relationships between its cells. In order to reconstruct the syntactic structures of a table, we build a cell relationship network to predict the neighbors of each cell in four directions. During the training stage, a distance-based sample weight is proposed to handle the class imbalance problem. According to the detected relationships, the table is represented by a weighted graph that is then employed to infer the basic syntactic table structure. Experimental evaluation of the proposed framework using two datasets demonstrates the effectiveness of our model for cell relationship detection and table structure inference. Wenyuan Xue, Qingyong Li, Dacheng Tao |
ICDAR | 2 |
| 2019 | An End-to-End Abnormal Fastener Detection Method Based on Data SynthesisabstractFasteners that hold rails in a fixed position are essentially important infrastructure components, and abnormal fasteners may cause a train to derail. The periodical inspection of fasteners is a significant guarantee for the safety of railway operation, and machine vision systems are popularly applied for fastener inspection. In this paper, we present an end-to-end abnormal fastener detection method, which identifies abnormal fasteners from a track image that contains a rail, fasteners, sleepers, etc. The proposed method is inspired from the well-established Faster R-CNN, but customizes with two aspects according to the application of fastener inspection. On the one hand, a light-weight backbone network is employed instead of complicated network to quicken detecting speed, and a threshold pruning algorithm is designed to reduce false positive rate. On the other hand, a hybrid loss function, combining weighted softmax loss function with center loss function, is devised to handle both the problems of class imbalance and small inter-class differences. Furthermore, two data synthesis methods are brought forward to solve the small sample problem. The proposed method is verified on a real-world image set of railway inspection, and the experimental results demonstrate that our method outperforms four established baselines. Bangyi Dong, Qingyong Li, Jianzhu Wang |
ICTAI | 2 |
| 2019 | LightNet: A Dual Spatiotemporal Encoder Network Model for Lightning PredictionabstractLightning as a natural phenomenon poses serious threats to human life, aviation and electrical infrastructures. Lightning prediction plays a vital role in lightning disaster reduction. Existing prediction methods, usually based on numerical weather models, rely on lightning parameterization schemes for forecasting. These methods, however, have two drawbacks. Firstly, simulations of the numerical weather models usually have deviations in space and time domains, which introduces irreparable biases to subsequent parameterization processes. Secondly, the lightning parameterization schemes are designed manually by experts in meteorology, which means these schemes can hardly benefit from abundant historical data. In this work, we propose a data-driven model based on neural networks, referred to as LightNet, for lightning prediction. Unlike the conventional prediction methods which are fully based on numerical weather models, LightNet introduces recent lightning observations in an attempt to calibrate the simulations and assist the prediction. LightNet first extracts spatiotemporal features of the simulations and observations via dual encoders. These features are then combined by a fusion module. Finally, the fused features are fed into a spatiotemporal decoder to make forecasts. We conduct experimental evaluations on a real-world North China lightning dataset, which shows that LightNet achieves a threefold improvement in equitable threat score for six-hour prediction compared with three established forecast methods. Qingyong Li, Tianyang Lin, Liangtao Xu, Weitao Lyu, Yijun Zhang 0002 |
KDD | 2 |
| 2018 | A cyber-enabled visual inspection system for rail corrugation
Qingyong Li, Zhi-Ping Shi 0002, Huayan Zhang, Yunqiang Tan, Shengwei Ren |
Future Gener. Comput. Syst. | 1 |
| 2018 | RECOME: A new density-based clustering algorithm using relative KNN kernel density
Qingyong Li, Rong Zheng 0001, Fuzhen Zhuang, Ruisi He, Naixue Xiong |
Inf. Sci. | 2 |
| 2017 | Fabric defect detection based on improved low-rank and sparse matrix decompositionabstractIn this paper, we propose an effective approach to detect defects in fabrics. Based on the observation that fabric textures usually form a low-rank structure and the structure can be violated by the presence of defects, we formulate the task as a low-rank and sparse matrix decomposition problem. Moreover, the prior that defects tend to be continuous regions is considered in our model and the estimation of defect levels is properly solved by introducing an integration mechanism. Experimental results demonstrate that our proposed method can not only detect defects accurately but also have greater ability to preserve defect details than traditional approaches. Jianzhu Wang, Qingyong Li, Jinrui Gan, Haomin Yu |
ICIP | 2 |
| 2017 | REMOLD: An Efficient Model-Based Clustering Algorithm for Large Datasets with SparkabstractDensity-based clustering algorithms have the distinctive advantage of discovering arbitrarily shaped clusters, but they usually require a procedure to compute the distance between every pair of data points, and this procedure is prohibitive for large datasets since it has quadratic computation complexity. In this paper, we propose a new distributed clustering algorithm, named REstore MOdel with Local Density estimation (REMOLD). Firstly, REMODL applies a balanced partitioning method to evenly divide an large dataset based on Local Sensitive Hashing (LSH). Then, it locally clusters each partition of the dataset, and uses a Gaussian model to represent each local cluster based on the observation that the density distribution of each local cluster shares similar shape with Gaussian distribution. Finally, these models are aggregated on a server where REMOLD restores global clusters based on these local Gaussian models. More specifically, model connection, which measures the density connectivity between two models, are defined to merge local models with an optimized procedure. In this aggregation, REMOLD requires low cost of network transmission for local Gaussian models, since the number of Gaussian models is often less than that of core objects for each partition. We evaluate REMOLD on three synthetic datasets and three real-world datasets on Spark, and the experiment results demonstrate that REMOLD is efficient and effective to find out clusters with complex shapes and it outperforms the established methods. Mingfei Liang, Qingyong Li, Jianzhu Wang, Zhi Wei 0001 |
ICPADS | 2 |
| 2017 | An Automatic Clustering Algorithm for Multipath Components Based on Kernel-Power-DensityabstractIn the real-world environments, multipath components (MPCs) of wireless channels are generally distributed as groups, i.e., clusters. Modeling the clustered MPCs is important and necessary for channel modeling and an automatic clustering algorithm is thus required. This paper proposes a novel Kernel-power-density (KPD) based algorithm for MPC clustering. It uses the Kernel density to incorporate the modeled behavior of MPCs and takes into account the power of the MPCs. The proposed algorithm only considers the K nearest MPCs in the density estimation to better identify the local density variations of MPCs. Simulations validate the KPD algorithm and almost no performance degradation is found even with a large number of clusters and large cluster angular spread. The KPD algorithm enables applications with no prior knowledge about the clusters such as number and initial locations. It can be used for the cluster based channel modeling for 4G#x002F;5G communications. Ruisi He, Qingyong Li, Bo Ai 0001, Andreas F. Molisch, Vinod Kristem, Zhangdui Zhong, Jian Yu 0001 |
WCNC | 2 |
| 2017 | A Kernel-Power-Density-Based Algorithm for Channel Multipath Components ClusteringabstractCluster-based channel modeling has been an important trend in the development of channel model, as it maintains accuracy while reducing complexity. Whereas a large number of channel measurements have shown that multipath components (MPCs) are distributed as groups, i.e., clusters, existing clustering algorithms have various drawbacks with respect to complexity, threshold choices, and/or assumptions about prior knowledge. In this paper, a kernel-power-density (KPD)-based algorithm is proposed for MPC clustering. It uses the kernel density of MPCs to incorporate the modeled behavior of MPCs and takes into account the power of the MPCs. Furthermore, the KPD algorithm only considers the K nearest MPCs in the density estimation to better identify the local density variations of MPCs. A heuristic approach of cluster merging is used to improve the performance. Both simulation and channel measurements validate the KPD algorithm, and almost no performance degradation is found even with a large number of clusters and large cluster angular spread, which outperforming other algorithms. The KPD algorithm enables applications in multipleinput-multiple-output channels with no prior knowledge about the clusters, such as number and initial locations. It also has a fairly low computational complexity and can be used for clusterbased channel modeling. Ruisi He, Qingyong Li, Bo Ai 0001, Andreas F. Molisch, Vinod Kristem, Zhangdui Zhong, Jian Yu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | Two-Level Feature Representation for Aerial Scene ClassificationabstractEffective scene representation is a fundamental part of high-resolution scene classification systems. In this letter, we present a holistic scene representation method, i.e., the two-level feature representation (TLFR) model. The TLFR is composed of low-level and high-level features. Low-level features are obtained by computing the residual error between a local descriptor and its corresponding visual word, and the high-level features are obtained using a proposed selection-constrained sparse coding method. In addition, low-level features in a cluster are integrated by summation pooling, whereas high-level features are fused by maximization pooling. The holistic scene representation is finally generated by incorporating these two levels of features into the bag-of-visual-words framework. Experimental results show that the TLFR model is robust to translation and rotation variations and demonstrates promising performance with the Land Use and Land Cover Database data set and a newly released Singapore data set. Jinrui Gan, Qingyong Li, Jianzhu Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | An implicit relevance feedback method for CBIR with real-time eye tracking
Qingyong Li, Jinrui Sun |
Multim. Tools Appl. | 1 |
| 2012 | Learning sparse tag patterns for social image classificationabstractUser-generated tags associated with images from social media (e.g., Flickr) provide valuable textual resources for image classification. However, the noisy and huge tag vocabulary heavily degrades the effectiveness and efficiency of state-of-the-art image classification methods that exploited auxiliary web data. To alleviate the problem, we introduce a Sparse Tag Patterns (STP) model to discover sparsity constrained co-occurrence tag patterns from large scale user contributed tags among social data. To fulfill the compactness and discriminability, we formulate STP as a problem of minimizing a quadratic loss function regularized by the bi-layer l1norm. We treat the learned STP as alternative intermediate semantic image feature and verify its superiority within a search-based image classification framework. Experiments on 240K social images associated with millions of tags have demonstrated encouraging performance of the proposed method compared to the state-of-the-art. Jie Lin 0001, Ling-Yu Duan, Junsong Yuan 0001, Qingyong Li, Siwei Luo |
ICIP | 4 |
| 2012 | Thin Cloud Detection of All-Sky Images Using Markov Random FieldsabstractThin cloud detection for all-sky images is a challenge in ground-based sky-imaging systems because of low contrast and vague boundaries between cloud and sky regions. We treat cloud detection as a labeling problem based on the Markov random field model. In this model, each pixel is represented by a combined-feature vector that aims at improving the disparity between thin cloud and sky. The distribution of each label in the feature space is defined as a Gaussian model. Spatial information is coded by a generalized Potts model. During the estimation, thin cloud is detected by minimizing the posterior energy with an iterative procedure. Both subjective and objective evaluation results demonstrate higher accuracy of the algorithm compared with some other algorithms. Qingyong Li, Weitao Lu, James Z. Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | Extracting discriminative features for CBIR
Zhi-Ping Shi 0002, Xi Liu 0009, Qingyong Li, Qing He 0003, Zhongzhi Shi |
Multim. Tools Appl. | 3 |
| 2012 | A Visual Detection System for Rail Surface DefectsabstractDiscrete surface defects are the most common anomalies of rails and they should be carefully inspected. However, it is a challenge to detect such defects in a vision system because of illumination inequality and the variation of reflection property of rail surfaces. This paper presents an intelligent vision detection system (VDS) for discrete surface defects and focuses on two key issues of VDS: image enhancement and automatic thresholding. We propose the local Michelson-like contrast (MLC) measure to enhance rail images. MLC-based method is nonlinear and illumination independent; therefore, it notably improves the distinction between defects and background. In addition, we put forward the new automatic thresholding method-proportion emphasized maximum entropy (PEME) thresholding algorithm. PEME selects a threshold that maximizes the object entropy and meanwhile keeps the defect proportion in a low level. Our experimental results demonstrate that VDS detects the Type-II defects with a recall of 91.61% and Type-I defects with a recall of 88.53%, and the proposed MLC-based image enhancement method and PEME thresholding algorithm outperform the related well-established approaches. Qingyong Li, Shengwei Ren |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2010 | An associative sparse coding neural network and applications
Siwei Luo, Qingyong Li |
Neurocomputing | 3 |
| 2009 | Transfer Learning with Data Edit
Qingyong Li |
ADMA | 2 |
| 2007 | Image Retrieval Based on Fuzzy Color SemanticsabstractIn order to improve the performance of content-based image retrieval (CBIR) systems, the 'semantic gap' between the low-level visual features and the high-level semantic features attracts more and more research interest. We propose an approach to describe and to extract the fuzzy color semantics. According to human color perception model, we utilize the linguistic variable to describe the image color semantics, so it becomes possible to depict the image in linguistic expression such as mostly red. Furthermore, we apply the feedforward neural network to model the vagueness of human color perception and to extract the fuzzy semantic feature vector. Our experiments show that the color semantic features have good accordance with the human perception, and also have good retrieval performance. In some extent, our approach shows the potential to reduce the semantic gap in CBIR. Qingyong Li, Zhi-Ping Shi 0002, Siwei Luo |
FUZZ-IEEE | 1 |
| 2007 | A Neural Network Approach for Bridging the Semantic Gap in Texture Image RetrievalabstractOne of the big challenges faced by content-based image retrieval (CBIR) is the 'semantic gap' between the visual features and the richness of human semantics for image content. We put forward a neural network approach to extract the image fuzzy semantics ground on linguistic expression based image description framework (LEBID). We utilize the linguistic variable to depict the texture semantics according to Tamura texture model, so we can describe the image in linguistic expression such as coarse, very line-like. Moreover, we use feedforward neural network (NN) to model the vagueness of human visual perception and to extract the fuzzy semantic feature. Our experiments demonstrate that NN outperforms other method such as genetic algorithm on the complexity of model, and it also achieves good retrieval performance. Qingyong Li, Zhi-Ping Shi 0002, Siwei Luo |
IJCNN | 1 |
| 2006 | An Improved Multiobjective Evolutionary Algorithm Based on Dominating Tree
Chuan Shi 0001, Qingyong Li, Zhongzhi Shi |
PRICAI | 2 |