EDBT 2026 Demo / reviewers in the wild / expert
Yong Feng 0002
dblp:64/4429-2
· DBLP profile ↗
41ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-8820-8388ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Denoising diffusion bridge for cross-modal alignment in three-dimensional scene understanding
Yong Feng 0002, Wuyang Luan, Pu Xiao, Yanying Chen, Guofan Duan, Jintao Tan, Mingliang Zhou 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Enhancing surface defect detection in industrial products through few-shot learning via meta-learning
Quanyou Zhang, Yong Feng 0002, Yanying Chen, Baohua Qiang, Zhangli Lan |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Unsupervised deep metric learning based on context and attention weighting
Shiya Li, Yong Feng 0002, Pu Xiao, Shi-hao Yuan, Guofan Duan, Yanying Chen, Baohua Qiang, Xiancai Xiong, Mingliang Zhou 0001 |
Multim. Syst. | 2 |
| 2026 | Neighborhood Attention-based Feature Reconstruction for Image Anomaly Detection and LocalizationabstractWith the advancement of machine vision technology, automated vision inspection systems are needed in broad quality control scenarios. This article proposes a neighborhood attention-based feature reconstruction method for image anomaly detection and localization (NAFRAD). To address the challenges of data scarcity, low visibility, and irregular defect shapes in unsupervised anomaly detection, we introduce a feature reconstruction framework that preserves high-level abstract features rather than focusing on pixel-level reconstruction. This approach enhances model robustness and generalizability by leveraging neighborhood attention (NA) mechanisms, which simultaneously capture local details and the global context through a sliding window strategy. The NA-based autoencoder reconstructs normal features by aggregating local inductive biases with translational equivariance, enabling precise anomaly localization. Extensive experiments on the MVTec Anomaly Detection (MVTec AD) dataset—comprising 15 categories with 5,354 images—demonstrate the superiority of NAFRAD. It achieves state-of-the-art performance with AUROC \({}_{I}\) = 99.02, AUROC \({}_{P}\) = 98.99, and AP = 79.40, outperforming existing methods by 3.6% in AP and 0.89% in AUROC \({}_{P}\) . The framework’s effectiveness is validated through ablation studies, visualization of feature reconstruction, and comparisons with eight leading unsupervised methods. The code is made public at https://github.com/Math-Computer/NAFRAD . Weizhi Xian, Yichi Chen 0002, Bin Chen 0022, Leong Hou U, Shiyou Liu, Yong Feng 0002, Mingliang Zhou 0001, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual InferenceabstractExisting full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse scenarios. In this paper, we propose an FR-IQA method based on abductive counterfactual inference to investigate the causal relationships between deep network features and perceptual distortions. First, we explore the causal effects of deep features on perception and integrate causal reasoning with feature comparison, constructing a model that effectively handles complex distortion types across different IQA scenarios. Second, the analysis of the perceptual causal correlations of our proposed method is independent of the backbone architecture and thus can be applied to a variety of deep networks. Through abductive counterfactual experiments, we validate the proposed causal relationships, confirming the model’s superior perceptual relevance and interpretability of quality scores. The experimental results demonstrate the robustness and effectiveness of the method, providing competitive quality predictions across multiple benchmarks. The source code is available at https://anonymous.4open.science/r/DeepCausalQuality-25BC. Wenhao Shen, Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Huayan Pu, Weijia Jia 0001 |
CVPR | 5 |
| 2025 | Manifold and patch-based unsupervised deep metric learning for fine-grained image retrieval
Shi-hao Yuan, Yong Feng 0002, Agen Qiu, Guofan Duan, Mingliang Zhou 0001, Baohua Qiang, Yong-heng Wang |
Appl. Intell. | 2 |
| 2025 | Continuous reinforcement learning via advantage value difference reward shaping: A proximal policy optimization perspective
Xuekai Wei, Weizhi Xian, Jielu Yan, Leong Hou U, Yong Feng 0002, Zhaowei Shang, Mingliang Zhou 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Variational adversarial negative sampling for multimodal knowledge graph completion
Haohui Miao, Yong Feng 0002, Yanying Chen, Guofan Duan, Baohua Qiang, Mingliang Zhou 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | ViLNM: Visual-Language Noise Modeling for Text-to-Image Person RetrievalabstractText-to-image person retrieval (TPR) focuses on finding a specific person based on the textual description, and most methods implicitly assume the training image-text pairs are correctly aligned. In practice, the image-text pairs exist under-correlated or false-correlated due to the low quality of the images and annotation errors. Meanwhile, remarkable similarities between different person identities may lead to a mismatch between text and image. To tackle the two issues, we present a Visual-Language Noise Modeling (ViLNM) method that successfully captures robust cross-modal associations even with noise. Specifically, we design a Noise Token Aware (NTA) module that eliminates the words in the textual description that do not match the image, utilizing the matched words to establish a more reliable association. Besides, to enhance the recognition ability of the model for different person identities, we propose a Joint Inter and Intra-Modal Contrastive Loss (JII) and Local Aggregation (LA) module to increase the feature differences between different person identities. We conduct comprehensive experiments on three public benchmarks, and ViLNM performs best. Guolin Xu, Yong Feng 0002, Yanying Chen, Guofan Duan, Mingliang Zhou 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Boundary-Aware Feature Fusion With Dual-Stream Attention for Remote Sensing Small Object DetectionabstractDetecting small objects in remote sensing images poses significant challenges to the field of computer vision, primarily stemming from the complexity of backgrounds, limitations in pixel resolution, and information loss during the feature fusion process. While general object detection has significantly advanced in recent years, remote sensing small object detection remains an unsolved problem, with existing frameworks struggling to achieve high performance at small scales. In this article, we propose a novel framework called the boundary-aware feature fusion network (BAFNet), which significantly enhances the model’s ability to represent and locate small objects precisely within complex remote sensing scenarios. First, a dual-stream attention fusion module captures complementary foreground and background cues through bidirectional context modeling. Jointly attending to objects and their surroundings enhances discriminative power for distinguishing small objects. Additionally, we incorporate a boundary-aware branch to better preserve crucial detailed information vital for small-scale objects. This auxiliary component supervises the fusion of contextual semantics and spatial information, aiding in retaining critical boundary details that are prone to loss during cross-layer feature fusion. We conducted experiments on the challenging AI-TOD, VisDrone, DIOR, and LEVIR-Ship datasets. The results demonstrate the superiority of our approach over other state-of-the-art (SOTA) object detection methods, particularly in terms of precisely identifying small objects within remote sensing images. The code is available athttps://github.com/ooo1128/BAFNet. Jingnan Song, Mingliang Zhou 0001, Jun Luo 0006, Huayan Pu, Yong Feng 0002, Xuekai Wei, Weijia Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Sparse Reduced-Rank Fully Connected Layers with Its Applications in Detection and ClassificationabstractFully connected (FC) layers play a significant role in deep neural networks (DNNs) models. Owing to the complexity of its parameters, an FC layer has sufficient capacity to manage high-dimensional tasks, so a large amount of memory and powerful computing capabilities become essential requirements. However, the large number of parameters in an FC layer greatly limits the practical application of this model. To address this problem, we apply matrix optimization to an FC layer. First, an added penalty term properly maintains the sparsity of the imposed weights. Second, a rank constraint is applied to the two components of the factorized weight matrix. Our compression algorithm can effectively reduce the number of required network parameters, which not only reduces the computational complexity of the network but also results in better generalizability on a test dataset. Finally, the effectiveness of the proposed method is verified in two different computer vision task domains. Experiments show that our sparse reduced-rank method achieves a better compression ratio with a lower accuracy loss relative to the competing approaches. The code is available at https://github.com/cheer79/Compress_FC . Mingliang Zhou 0001, Xuekai Wei, Yong Feng 0002, Tao Xiang 0001, Bin Fang 0001, Zhaowei Shang, Fan Jia 0005, Xu Zhuang, Huayan Pu, Jun Luo 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Deep Multiscale Fine-Grained Hashing for Remote Sensing Cross-Modal RetrievalabstractHashing retrieval is a widely used technique in high spatial resolution remote sensing (RS) images due to its efficient retrieval speed and low memory overhead. However, existing hashing retrieval methods primarily focus on matching multilabel RS images, neglecting the extensive fine-grained semantic information in cross-modal RS data. Moreover, RS images exhibit notable object size differences and contain redundant features that lack effective multiscale feature extraction methods. To address these issues, we propose a novel deep multiscale fine-grained hashing (DMFH) method for cross-modal hashing retrieval of RS data. The DMFH method comprises two modules: the feature extraction module and hashing retrieval module. In the feature extraction module, we introduce a multiscale feature representation method to extract both low-level and high-level features from RS images while using a redundant optimizer to remove duplicate features. In addition, we used embedding vectors to extract fine-grained semantic information from description texts. The hashing retrieval module uses contrastive loss and triplet loss to guide the hash function toward learning and generating hash codes from extracted features. Our proposed DMFH method achieves state-of-the-art performance in two public RS image–text datasets (RSICD and RSITMD) through extensive experiments and ablation studies. Jiaxiang Huang, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | DMA-YOLO: multi-scale object detection method with attention mechanism for aerial images
Ya-ling Li, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Vis. Comput. | 2 |
| 2023 | Joint Robust Representation And Generalization Enhancement For Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification (cm-ReID) aims to match pedestrian images from visible and infrared cameras. Most existing methods ignore data bias due to different cameras and views and overlook the strong dependence between feature maps that hinders modal alignment. In this paper, we propose a unified method named Joint Robust Representation and Generalization Enhancement (RRGE) to alleviate the above issues. First, we propose a robust representation module (RRM), which can improve the model’s robustness for the global context, camera, and view change perturbations. Second, we propose a generalization enhancement module (GEM), which uses channel-level dropout to alleviate the dependencies between feature maps to improve the model’s generalization. Moreover, we balance the number of different modalities in each batch. Our method outperforms other state-of-the-art methods in terms of cross-modality person re-identification tasks. Heqing Cheng, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
ICASSP | 2 |
| 2023 | TIAR: Text-Image-Audio Retrieval with weighted multimodal re-ranking
Peide Chi, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Appl. Intell. | 2 |
| 2023 | Multi-modal transformer using two-level visual features for fake news detection
Yong Feng 0002, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Appl. Intell. | 2 |
| 2023 | Improve Session-Based Recommendation with Triplet Mining and Dynamic Perturbations Graph Neural NetworksabstractSession-based recommendation (SBR) emphasizes mining user interests to predict the next click based on recent interactions within sessions. Most current SBR methods suffer from insufficient interactive information problems and fail to distinguish session representations with high similarities, which can neglect the inherent features within sessions. To fill the gap, we propose a triplet mining enhanced graph neural networks (TME-GNN) approach to enhance the recommendation systems by mining structural and inherent information. Technically, we first generate anchor, positive and negative embeddings based on the given session and set a triplet mining task to improve the recommendation task with subtle features by pushing positive pairs close and pulling negative pairs away. Second, to robust the model, we employ a self-supervised auxiliary task by adding dynamic perturbations to the embedding space. We conduct extensive experiments to demonstrate the superiority of our method against other state-of-the-art algorithms. Our implementations are available on the following site https://github.com/Info4Rec/TME-GNN . Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang, Qin Mao, Bin Fang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2023 | Multi-level network based on transformer encoder for fine-grained image-text matching
Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Multim. Syst. | 2 |
| 2023 | Multilevel Similarity-Aware Deep Metric Learning for Fine-Grained Image RetrievalabstractFast and accurate image retrieval is an important and challenging task in massive image data scenarios. As the core technology of image retrieval tasks, deep metric learning aims at learning effective embedding representations that possess two properties among data points: positive concentrated and negative separated. In this work, we propose a multilevel similarity-aware method based on deep local descriptors for deep metric learning. We take the rich interclass similarity relationship based on the deep local invariant descriptors from the data into account to optimize sampling strategies for mining informative samples. The method dynamically adjusts the margin between data points to better match the true similarity relationship between classes. Specifically, for images in a batch, we first obtain deep local descriptors and calculate the similarity matrix of the channel, pixel, and spatial levels. Then, depending on the calculated comprehensive similarity matrix, we propose a multilevel similarity-aware loss function through the deviation between pairwise distance and violate margin to make full use of informative samples. The experimental results demonstrate that our proposed method outperforms other state-of-the-art methods in terms of fine-grained image retrieval and clustering tasks. Congcong Duan, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang, Weijia Jia 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Deep Adversarial Quantization Network for Cross-Modal RetrievalabstractIn this paper, we propose a seamless multimodal binary learning method for cross-modal retrieval. First, we utilize adversarial learning to learn modality-independent representations of different modalities. Second, we formulate loss function through the Bayesian approach, which aims to jointly maximize correlations of modality-independent representations and learn the common quantizer codebooks for both modalities. Based on the common quantizer codebooks, our method performs efficient and effective cross-modal retrieval with fast distance table lookup. Extensive experiments on three cross-modal datasets demonstrate that our method outperforms state-of-the-art methods. The source code is available at https://github.com/zhouyu1996/DAQN. Yu Zhou 0053, Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang, Leong Hou U, Jiajie Zhu 0001 |
ICASSP | 2 |
| 2021 | Distribution-Aware Hierarchical Weighting Method for Deep Metric LearningabstractIn this paper, we propose distribution-aware hierarchical weighting (DHW) method for deep metric learning. First, we formulate the distributions of different classes according to the form of gaussian curves, and update distributions as the training process. Second, depending on the learnable distribution, we propose a loss function named distribution-aware loss with dynamic mining margins and hierarchical degrees of weights to make full use of samples. The experimental results show that our algorithm outperforms other state-of-the-art methods in terms of retrieval and clustering tasks. Code is available at https://github.com/zhuyinong1/DHW-master. Yinong Zhu, Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang, Leong Hou U, Jiajie Zhu 0001 |
ICASSP | 2 |
| 2021 | Semantic ranking structure preserving for cross-modal retrieval
Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang |
Appl. Intell. | 2 |
| 2021 | Prototype-Based Discriminative Feature Representation for Class-incremental Cross-modal RetrievalabstractCross-modal retrieval aims to retrieve the related items from various modalities with respect to a query from any type. The key challenge of cross-modal retrieval is to learn more discriminative representations between different category, as well as expand to an unseen class retrieval in the open world retrieval task. To tackle the above problem, in this paper, we propose a prototype learning-based discriminative feature learning (PLDFL) to learn more discriminative representations in a common space. First, we utilize a prototype learning algorithm to cluster these samples labeled with the same semantic class, by jointly taking into consideration the intra-class compactness and inter-class sparsity without discriminative treatments. Second, we use the weight-sharing strategy to model the correlations of cross-modal samples to narrow down the modality gap. Finally, we apply the prototype to achieve class-incremental learning to prove the robustness of our proposed approach. According to our experimental results, significant retrieval performance in terms of mAP can be achieved on average compared to several state-of-the-art approaches. Shaoquan Zhu, Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang, Bin Fang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2020 | Relationship-Aware Hard Negative Generation in Deep Metric Learning
Yong Feng 0002, Mingliang Zhou 0001, Baohua Qiang |
KSEM (2) | 2 |
| 2020 | A Deep Sequence-to-Sequence Method for Aircraft Landing Speed Prediction Based on QAR Data
Zongwei Kang, Jiaxing Shang, Yong Feng 0002, Linjiang Zheng, Dajiang Liu, Baohua Qiang |
WISE (2) | 3 |
| 2020 | FRWCAE: joint faster-RCNN and Wasserstein convolutional auto-encoder for instance retrieval
Yong Feng 0002, Dajiang Liu, Jiaxing Shang, Baohua Qiang |
Appl. Intell. | 2 |
| 2020 | DSRPH: Deep semantic-aware ranking preserving hashing for efficient multi-label image retrieval
Yong Feng 0002, Bin Fang 0001, Mingliang Zhou 0001, Sam Kwong, Baohua Qiang |
Inf. Sci. | 2 |
| 2019 | Data-Flow Graph Mapping Optimization for CGRA With Deep Reinforcement LearningabstractCoarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their flexibility and energy efficiency. Data flow graphs (DFGs) are often mapped onto CGRAs for acceleration. The problem of DFG mapping is challenging due to the diverse structures from DFGs and constrained hardware from CGRAs. Consequently, it is difficult to find a valid and high quality solution simultaneously. Inspired from the great progress in deep reinforcement learning (RL) for AI problems, we consider building methods that learn to map DFGs onto spatially programmed CGRAs directly from experiences. We propose RLMap, a solution that formulates DFG mapping on CGRA as an agent in RL, which unifies placement, routing and processing element insertion by interchange actions of the agent. Experimental results show that RLMap performs comparably to state-of-the-art heuristics in mapping quality, adapts to different architecture, and converges quickly. Dajiang Liu, Shouyi Yin, Guojie Luo, Jiaxing Shang, Leibo Liu, Shaojun Wei, Yong Feng 0002, Shangbo Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2018 | SCE-MSPFS: A Novel Deep Convolutional Feature Selection Method for Image Retrieval
Dong-dong Niu, Yong Feng 0002, Jiaxing Shang, Baohua Qiang |
ICONIP (6) | 2 |
| 2018 | Wasserstein Generative Recurrent Adversarial Networks for Image GeneratingabstractMost generative models are generating images at a time, but in fact, painting is usually done iteratively and repeatedly. Generative Adversarial Networks (GAN) are well known for generating images, however, it is hard to train stably. To tackle this problem, we propose a framework named the Wasserstein generative recurrent adversarial networks (WGRAN), which merges Wasserstein distance with recurrent neural networks to iteratively generate realistic looking images and trains our model in an adversarial way. Therefore, our generative model is gradually generates images using the feedback of discriminate model. And our approach allows us to control the number of iterations of generation. We train our model on various image datasets and compare our model with the recurrent generative adversarial networks (GRAN) and other state-of-the-art generative models using Generative Adversarial Metric. From these experiments, we show evidence that our model has the ability to generate high quantity images. Chunping Zhang, Yong Feng 0002, Baohua Qiang, Jiaxing Shang |
ICPR | 2 |
| 2018 | Density core-based clustering algorithm with dynamic scanning radius
Jiang Xie 0002, Zhongyang Xiong, Yu-Fang Zhang, Yong Feng 0002, Jie Ma 0001 |
Knowl. Based Syst. | 4 |
| 2017 | App Uninstalls Prediction: A Machine Learning and Time Series Mining Approach
Jiaxing Shang, Hongchun Wu, Shangbo Zhou, Yong Feng 0002 |
ICONIP (5) | 6 |
| 2017 | A Linear Time Algorithm for Influence Maximization in Large-Scale Social Networks
Hongchun Wu, Jiaxing Shang, Shangbo Zhou, Yong Feng 0002 |
ICONIP (5) | 4 |
| 2016 | Improving Temporal Recommendation Accuracy and Diversity via Long and Short-Term Preference Transfer and Fusion Models
Yong Feng 0002 |
APWeb (2) | 2 |
| 2010 | Deep Web Sources Classifier Based on DSOM-EACO Clustering Model
Yong Feng 0002, Xianyong Chen |
ADMA (1) | 1 |
| 2010 | An enhanced swarm intelligence clustering-based RBFNN classifier and its application in deep Web sources classification
Yong Feng 0002, Zhongfu Wu, Chunxiao Ye, Kaigui Wu |
Frontiers Comput. Sci. China | 1 |
| 2009 | An Enhanced Swarm Intelligence Clustering-Based RBF Neural Network Web Text Classifier
Yong Feng 0002, Zhongfu Wu, Chunxiao Ye, Kaigui Wu |
ISNN (2) | 1 |
| 2008 | An Enhanced Swarm Intelligence Clustering-Based RBF Neural Network Detection Classifier
Yong Feng 0002, Zhongfu Wu, Chunxiao Ye, Kaigui Wu |
ICIC (2) | 1 |
| 2007 | Network Anomaly Detection Based on DSOM and ACO Clustering
Yong Feng 0002, Zhongyang Xiong, Chunxiao Ye, Kaigui Wu |
ISNN (2) | 1 |
| 2006 | Neuron Selection for RBF Neural Network Classifier Based on Multiple Granularities Immune Network
Chunxiao Ye, Yong Feng 0002, Zhongfu Wu |
ISNN (1) | 3 |
| 2005 | Intrusion Detection Based on Dynamic Self-organizing Map Neural Network Clustering
Yong Feng 0002, Kaigui Wu, Zhongfu Wu, Zhongyang Xiong |
ISNN (3) | 1 |