VLDB 2026 Research / reviewers in the wild / expert
Yijie Zeng
dblp:216/8272
· DBLP profile ↗
20ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-8843-0755ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional chain-of-thought for zero-shot object navigation
Haonan Luo 0002, Yijie Zeng, Zihang Wang 0002, Botao Jiang, Xiruo Jiang |
Frontiers Comput. Sci. | 3 |
| 2025 | MoPE: Mixture of Policy Experts and Verification with Multimodal Information for Instance ImageGoal NavigationabstractInstance ImageGoal Navigation (IIN) entails an agent autonomously seeking out a specific object instance depicted by a goal image in an unknown environment. While Large Language Models (LLMs) have shown promise in navigation tasks similar to IIN, their application to IIN remains unexplored. Furthermore, existing LLM-based exploration faces challenges such as inaccurate reasoning due to informative environmental information available to the agent, especially in the early episode stages, and the inability of reinforcement learning(RL) exploration to fully leverage gathered information. Moreover, previous IIN methods did not productively verify potentially distant goal objects discovered during exploration. This work proposes MoPE–Mixture of Policy Experts for exploration and potential goal verification with multimodal information when exploring. Specifically, the hybrid exploration policy comprises an LLM and an RL-based Policy Network (RLPN) to generate an exploration goal to explore efficiently. Our MoPE model surpasses prior approaches on the HM3D datasets significantly. Yijie Zeng, Kexun Chen, Zhixuan Shen, Haonan Luo 0002, Tianrui Li 0001 |
ICME | 1 |
| 2025 | A Triplet Optimization and Difference Detail Perception Network with Adaptive Feature Enhancement for Radiology Report GenerationabstractThe generation of radiation reports is an essential task in the field of medical artificial intelligence which aims to automatically generate text descriptions of radiology images. However, there are still several problems in this task: 1) existing methods lack global feature interaction when extracting image features, and their ability to represent images is limited; 2) previous models need to retrieve similar triplets input models from pre-constructed knowledge graphs, and the triplets lack entity relationship refinement; 3) existing approaches lack a correlation mechanism across samples that makes it difficult to effectively capture small abnormal regions; 4) the self-attention mechanism of the Transformer decoder is good at capturing global dependencies but ignores relationships between local contexts. To address these issues, we propose a triplet optimization and difference detail perception network with adaptive feature enhancement. In our model, we design an adaptive image feature enhancement module to dynamically capture global image features. Furthermore, we propose a multi-modal triplet optimization module that boosts capability for detecting abnormal regions by incorporating context-aware entity relationship refinement into the initial triplet. Moreover, we design a difference comparison weighting module to obtain fine-grained features between different samples and improve cross-sample correlation so that the model pays more attention to small details and anomalies that are easy to ignore. Finally, we design a detail-aware enhancement decoder to make the decoder pay more attention to the relationship between local contexts. We experimented and evaluated our model on the IU-Xray and MIMIC-CXR datasets to compare with other baseline models. Yijie Zeng, Wenfeng Jiang, Song Liu 0008 |
SMC | 1 |
| 2024 | A Novel 3D Medical Image Segmentation Model Using Improved SAMabstract3D medical image segmentation is an essential task in the medical image field, which aims to segment organs or tumours into different labels. A number of issues exist with the current 3D medical image segmentation task: existing models cannot simultaneously obtain the space correlation and depth correlation of 3D slices; previous models suffer from local detail loss of positional embedding in 3D images; previous approaches often have blurring of boundaries in segmenting 3D images. To solve these shortcomings, we propose a 3D medical image segmentation model named TPM-SAM. In our model, we design a twinchannel image encoder to simultaneously capture the space correlation and depth correlation of 3D slices through a multi-head attention mechanism and improved adapters. Furthermore, we design a prompt encoding generator, which divides the volumetric image into small blocks and better captures the local detail information. In addition, we introduce a multi-layer aggregation decoder by employing U-Net with multi-level skip connection to solve the blurring of boundaries in processing 3D images. Finally, we experimented and evaluated our model on KiTS21 and LiTS17 datasets to compare with other baseline models. Yuansen Kuang, Xitong Ma, Guangchen Wang, Yijie Zeng, Song Liu 0008 |
SMC | 5 |
| 2024 | A Cross-Modal Interactive Memory Network Based on Fine-Grained Medical Feature Extraction for Radiology Report GenerationabstractRadiology report generation is an essential task in the medical field, which aims to automate the generation of medical terminology descriptions of radiology images. However, this task currently suffers from several problems: 1) existing methods need to manually build knowledge graphs or templates (consuming time and effort) to introduce medical or prior knowledge to assist in report generation; 2) previous models cannot handle the problem of data bias well (anomaly reports and anomaly descriptions make up only a tiny portion of the dataset), causing the models to ignore the learning of anomaly descriptions easily; 3) existing approaches cannot robustly supervise the model, resulting in incomplete and inconsistent reports being generated. To address these issues, we propose a cross-modal interactive memory network based on fine-grained medical feature extraction. In our model, we design a cross-modal interactive memory network to automatically store and remember the required medical text knowledge and use this medical knowledge to help generate reports. Furthermore, we design an abnormal medical knowledge enhancement module to enhance the learning of abnormal fine-grained knowledge through the interaction of disease topics and their states to interact with text features. In addition, we design a cross-modal joint semantic loss unit to reduce semantic differences between different features and improve the visual representation ability of the model. We experimented and evaluated our model on MIMIC-CXR and IU-Xray datasets to compare with other baseline models. Xitong Ma, Yuansen Kuang, Yijie Zeng, Song Liu 0008 |
SMC | 5 |
| 2024 | VLAI: Exploration and Exploitation based on Visual-Language Aligned Information for Robotic Object Goal Navigation
Haonan Luo 0002, Yijie Zeng, Kexun Chen, Zhixuan Shen, Fengmao Lv |
Image Vis. Comput. | 2 |
| 2023 | E$^{3}$3Outlier: a Self-Supervised Framework for Unsupervised Deep Outlier DetectionabstractExisting unsupervised outlier detection (OD) solutions face a grave challenge with surging visual data like images. Although deep neural networks (DNNs) prove successful for visual data, deep OD remains difficult due to OD’s unsupervised nature. This paper proposes a novel framework namedE$^{3}$Outlierthat can performeffective andend-to-end deep outlier removal. Its core idea is to introduceself-supervisioninto deep OD. Specifically, our major solution is to adopt a discriminative learning paradigm that creates multiple pseudo classes from given unlabeled data by various data operations, which enables us to apply prevalent discriminative DNNs (e.g. ResNet) to the unsupervised OD problem. Then, with theoretical and empirical demonstration, we argue that inlier priority, a property that encourages DNN to prioritize inliers during self-supervised learning, makes it possible to perform end-to-end OD. Meanwhile, unlike frequently-used outlierness measures (e.g. density, proximity) in previous OD methods, we explore network uncertainty and validate it as a highly effective outlierness measure, while two practical score refinement strategies are also designed to improve OD performance. Finally, in addition to the discriminative learning paradigm above, we also explore the solutions that exploit other learning paradigms (i.e. generative learning and contrastive learning) to introduce self-supervision forE$^{3}$Outlier. Such extendibility not only brings further performance gain on relatively difficult datasets, but also enablesE$^{3}$Outlierto be applied to other OD applications like video abnormal event detection. Extensive experiments demonstrate thatE$^{3}$Outliercan considerably outperform state-of-the-art counterparts by 10%-30% AUROC. Demo codes are available athttps://github.com/demonzyj56/E3Outlier. Siqi Wang 0001, Yijie Zeng, Zhen Cheng 0004, Xinwang Liu 0002, Sihang Zhou 0001, En Zhu, Marius Kloft, Jianping Yin, Qing Liao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | End-to-end novel visual categories learning via auxiliary self-supervision
Yuanyuan Qing, Yijie Zeng, Qi Cao 0002, Guang-Bin Huang |
Neural Networks | 2 |
| 2021 | Label propagation via local geometry preserving for deep semi-supervised image recognition
Yuanyuan Qing, Yijie Zeng, Guang-Bin Huang |
Neural Networks | 2 |
| 2021 | Slice-Based Online Convolutional Dictionary LearningabstractConvolutional dictionary learning (CDL) aims to learn a structured and shift-invariant dictionary to decompose signals into sparse representations. While yielding superior results compared to traditional sparse coding methods on various signal and image processing tasks, most CDL methods have difficulties handling large data, because they have to process all images in the dataset in a single pass. Therefore, recent research has focused on online CDL (OCDL) which updates the dictionary with sequentially incoming signals. In this article, a novel OCDL algorithm is proposed based on a local, slice-based representation of sparse codes. Such representation has been found useful in batch CDL problems, where the convolutional sparse coding and dictionary learning problem could be handled in a local way similar to traditional sparse coding problems, but it has never been explored under online scenarios before. We show, in this article, that the proposed algorithm is a natural extension of the traditional patch-based online dictionary learning algorithm, and the dictionary is updated in a similar memory efficient way too. On the other hand, it can be viewed as an improvement of existing second-order OCDL algorithms. Theoretical analysis shows that our algorithm converges and has lower time complexity than existing counterpart that yields exactly the same output. Extensive experiments are performed on various benchmarking datasets, which show that our algorithm outperforms state-of-the-art batch and OCDL algorithms in terms of reconstruction objectives. Yijie Zeng, Jichao Chen, Guang-Bin Huang |
IEEE Trans. Cybern. | 1 |
| 2020 | Unsupervised feature selection based extreme learning machine for clustering
Jichao Chen, Yijie Zeng, Yue Li 0024, Guang-Bin Huang |
Neurocomputing | 2 |
| 2020 | Learning local discriminative representations via extreme learning machine for machine fault diagnosis
Yue Li 0024, Yijie Zeng, Yuanyuan Qing, Guang-Bin Huang |
Neurocomputing | 2 |
| 2020 | Deep and wide feature based extreme learning machine for image classification
Yuanyuan Qing, Yijie Zeng, Yue Li 0024, Guang-Bin Huang |
Neurocomputing | 2 |
| 2020 | Clustering via Adaptive and Locality-constrained Graph Learning and Unsupervised ELM
Yijie Zeng, Jichao Chen, Yue Li 0024, Yuanyuan Qing, Guang-Bin Huang |
Neurocomputing | 1 |
| 2020 | Simultaneously learning affinity matrix and data representations for machine fault diagnosis
Yue Li 0024, Yijie Zeng, Tianchi Liu 0001, Xiaofan Jia, Guang-Bin Huang |
Neural Networks | 2 |
| 2020 | ELM embedded discriminative dictionary learning for image classification
Yijie Zeng, Yue Li 0024, Jichao Chen, Xiaofan Jia, Guang-Bin Huang |
Neural Networks | 1 |
| 2019 | Effective End-to-end Unsupervised Outlier Detection via Inlier Priority of Discriminative NetworkabstractDespite the wide success of deep neural networks (DNN), little progress has been made on end-to-end unsupervised outlier detection (UOD) from high dimensional data like raw images. In this paper, we propose a framework named E^3Outlier, which can perform UOD in a both effective and end-to-end manner: First, instead of the commonly-used autoencoders in previous end-to-end UOD methods, E^3Outlier for the first time leverages a discriminative DNN for better representation learning, by using surrogate supervision to create multiple pseudo classes from original unlabelled data. Next, unlike classic UOD that utilizes data characteristics like density or proximity, we exploit a novel property named inlier priority to enable end-to-end UOD by discriminative DNN. We demonstrate theoretically and empirically that the intrinsic class imbalance of inliers/outliers will make the network prioritize minimizing inliers' loss when inliers/outliers are indiscriminately fed into the network for training, which enables us to differentiate outliers directly from DNN's outputs. Finally, based on inlier priority, we propose the negative entropy based score as a simple and effective outlierness measure. Extensive evaluations show that E^3Outlier significantly advances UOD performance by up to 30% AUROC against state-of-the-art counterparts, especially on relatively difficult benchmarks. Siqi Wang 0001, Yijie Zeng, Xinwang Liu 0002, En Zhu, Jianping Yin, Chuanfu Xu, Marius Kloft |
NeurIPS | 2 |
| 2018 | Data Driven Convolutional Sparse Coding for Visual RecognitionabstractConvolutional sparse coding (CSC) has become an important method in image processing and computer vision. In this paper we focus on visual recognition problems and apply CSC as a feature learning method. We propose a task-specific approach to treat the dictionary of CSC as parameters for a larger learning framework. These parameters are differentiable under mild conditions, and could be updated end-to-end using back-propagation when the errors from the task objectives are provided. We perform several experiments to show that such method provides a more discriminate representation compared with previous CSC methods, and this data driven approach is effective for visual recognition problems. Yijie Zeng, Jichao Chen, Guang-Bin Huang |
ICASSP | 1 |
| 2018 | Octree-based Convolutional Autoencoder Extreme Learning Machine for 3D Shape ClassificationabstractWe introduce Octree-based Convolutional Autoencoder Extreme Learning Machine (OCA-ELM) for 3D shape classification. This approach combines Convolutional Autoencoder Extreme Learning Machine (CAE-ELM) with octreebased con- volution to generate feature maps from several types of geometric data, and extract discriminative features with Extreme Learning Machine Autoencoder (ELM-AE). The extracted features can then be used for various computer graphics applications, such as 3D shape classification. Compared with other 3D classification methods, the proposed OCA-ELM has superior classification performance. Experiments on ModelNet40 show that OCA-ELM outperforms state-of-the-art CNN-based methods and surpasses CAE-ELM in classification accuracy by 3.69%, demonstrating the effectiveness of our method. Jichao Chen, Yijie Zeng, Siqi Wang 0001, Soh Ling Min, Guang-Bin Huang |
IJCNN | 2 |
| 2018 | Detecting Abnormality without Knowing Normality: A Two-stage Approach for Unsupervised Video Abnormal Event DetectionabstractAbnormal event detection in video surveillance is a valuable but challenging problem. Most methods adopt a supervised setting that requires collecting videos with only normal events for training. However, very few attempts are made under unsupervised setting that detects abnormality without priorly knowing normal events. Existing unsupervised methods detect drastic local changes as abnormality, which overlooks the global spatio-temporal context. This paper proposes a novel unsupervised approach, which not only avoids manually specifying normality for training as supervised methods do, but also takes the whole spatio-temporal context into consideration. Our approach consists of two stages: First, normality estimation stage trains an autoencoder and estimates the normal events globally from the entire unlabeled videos by a self-adaptive reconstruction loss thresholding scheme. Second, normality modeling stage feeds the estimated normal events from the previous stage into one-class support vector machine to build a refined normality model, which can further exclude abnormal events and enhance abnormality detection performance. Experiments on various benchmark datasets reveal that our method is not only able to outperform existing unsupervised methods by a large margin (up to 14.2% AUC gain), but also favorably yields comparable or even superior performance to state-of-the-art supervised methods. Siqi Wang 0001, Yijie Zeng, Qiang Liu 0004, Chengzhang Zhu, En Zhu, Jianping Yin |
ACM Multimedia | 2 |