EDBT 2026 Demo / reviewers in the wild / expert
Yingjie Cai
dblp:84/9538
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TextMSA: Multi-head Attention Model for Darknet Content Classification
Haihui Gao, Shaojie Guo, Yushan Xie, Yingjie Cai, Tianbo Lu |
ICIC (11) | 5 |
| 2026 | LCC-AKA: Lightweight certificateless cross-domain authentication key agreement protocol for IoT devices
Yingjie Cai, Tianbo Lu, Jiaze Shang, Qitai Gong, Hanrui Chen |
Comput. Networks | 1 |
| 2026 | IGANEEG: an IoMT-enabled generative framework for small-sample schizophrenia EEG augmentationabstractIn the Internet of Medical Things (IoMT), reliable electroencephalogram (EEG) analysis is crucial for early diagnosis of schizophrenia and other neurological disorders. However, limited schizophrenia EEG data restrict the generalization ability of deep learning models, hindering clinical deployment. Existing augmentation methods often produce low-quality synthetic samples or suffer from overfitting and vanishing gradients. To address these issues, we propose IGANEEG, an improved generative adversarial network for small-sample schizophrenia EEG augmentation. IGANEEG incorporates joint time–frequency–spatial feature groups from real EEG as conditional inputs to generate realistic synthetic samples retaining critical data characteristics. The generator is replaced by an autoencoder to reduce reconstruction loss, and a dynamic pause mechanism is proposed to enhance training stability while the learning capacity of the hidden layers is optimized. Experiments on a public dataset and a private dataset show that IGANEEG achieves recognition accuracies of 96.8% and 98.2%, respectively, substantially outperforming state-of-the-art methods. These results validate the framework's effectiveness in improving model generalization and offer valuable data support for intelligent healthcare systems in the IoMT. Heyan Huang, Yingjie Cai |
Connect. Sci. | 3 |
| 2026 | Centerless semi-supervised clustering via sparse distance optimization and unified K-means/Spectral framework
Jianyong Zhu, Jianyang Shen, Kaijun Jia, Hui Yang 0005, Yingjie Cai, Feiping Nie 0001 |
Inf. Sci. | 5 |
| 2026 | Multi-subspace graph clustering joint dimensionality reduction and feature selection
Yingjie Cai, Hui Yang 0005, Jianyong Zhu, Feiping Nie 0001 |
Pattern Recognit. | 1 |
| 2025 | VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous DrivingabstractThis paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to reconstruct multi-view representations using only images as supervision. Specifically, we introduce a self-supervised method for voxel velocity estimation. By warping voxels to adjacent frames and supervising the rendered outputs, the model effectively learns motion cues in the sequential data. Furthermore, we adopt a multi-frame photometric consistency approach to enhance geometric perception. It projects adjacent frames to the current frame based on rendered depths and relative poses, boosting the 3D geometric representation through pure image supervision. Extensive experiments on autonomous driving datasets demonstrate that VisionPAD significantly improves performance in 3D object detection, occupancy prediction and map segmentation, surpassing state-of-the-art pre-training strategies by a considerable margin. Haiming Zhang 0001, Wending Zhou, Yiyao Zhu, Xu Yan 0005, Jiantao Gao, Dongfeng Bai, Yingjie Cai, Shuguang Cui, Zhen Li 0026 |
CVPR | 7 |
| 2025 | DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image GenerationabstractIn the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting the subject-essential attributes from the visual prompt. This leads to subject-irrelevant attributes infiltrating the generation process, ultimately compromising the personalization quality in both editability and ID preservation. In this paper, we present $\textbf{DisEnvisioner}$, a novel approach for effectively extracting and enriching the subject-essential features while filtering out -irrelevant information, enabling exceptional customization performance, in a $\textbf{tuning-free}$ manner and using only $\textbf{a single image}$. Specifically, the feature of the subject and other irrelevant components are effectively separated into distinctive visual tokens, enabling a much more accurate customization. Aiming to further improving the ID consistency, we enrich the disentangled features, sculpting them into a more granular representation. Experiments demonstrate the superiority of our approach over existing methods in instruction response (editability), ID consistency, inference speed, and the overall image quality, highlighting the effectiveness and efficiency of DisEnvisioner. Yongzhe Hu, Guibao Shen, Yingjie Cai, Weichao Qiu, Ying-Cong Chen |
ICLR | 5 |
| 2025 | Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language ModelsabstractLarge Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode occupancy as input for the LLM and address the category imbalances associated with occupancy, we propose Motion Separation Variational Autoencoder (MS-VAE). This innovative approach utilizes prior knowledge to distinguish dynamic objects from static scenes before inputting them into a tailored Variational Autoencoder (VAE). This separation enhances the model's capacity to concentrate on dynamic trajectories while effectively reconstructing static scenes. The efficacy of Occ-LLM has been validated across key tasks, including 4D occupancy forecasting, self-ego planning, and occupancybased scene question answering. Comprehensive evaluations demonstrate that Occ-LLM significantly surpasses existing state-of-the-art methodologies, achieving gains of about 6% in Intersection over Union (IoU) and 4% in mean Intersection over Union (mIoU) for the task of 4D occupancy forecasting. These findings highlight the transformative potential of Occ-LLM in reshaping current paradigms within robotic and autonomous driving. Tianshuo Xu, Hao Lu 0009, Xu Yan 0005, Yingjie Cai, Ying-Cong Chen |
ICRA | 4 |
| 2025 | SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous DrivingabstractSparse Perception Models (SPMs) adopt a query-driven paradigm that forgoes explicit dense BEV or volumetric construction, enabling highly efficient computation and accelerated inference. In this paper, we introduce SQS, a novel query-based splatting pre-training specifically designed to advance SPMs in autonomous driving. SQS introduces a plug-in module that predicts 3D Gaussian representations from sparse queries during pre-training, leveraging self-supervised splatting to learn fine-grained contextual features through the reconstruction of multi-view images and depth maps. During fine-tuning, the pre-trained Gaussian queries are seamlessly integrated into downstream networks via query interaction mechanisms that explicitly connect pre-trained queries with task-specific queries, effectively accommodating the diverse requirements of occupancy prediction and 3D object detection. Extensive experiments on autonomous driving benchmarks demonstrate that SQS delivers considerable performance gains across multiple query-based 3D perception tasks, notably in occupancy prediction and 3D object detection, outperforming prior state-of-the-art pre-training approaches by a significant margin (i.e., +1.3 mIoU on occupancy prediction and +1.0 NDS on 3D detection). Haiming Zhang 0001, Yiyao Zhu, Wending Zhou, Xu Yan 0005, Yingjie Cai, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 5 |
| 2025 | DRaft: A double-layer structure for Raft consensus mechanism
Jiaze Shang, Tianbo Lu, Yingjie Cai |
J. Netw. Comput. Appl. | 3 |
| 2024 | DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and PerceptionabstractCurrent perceptive models heavily depend on resource-intensive datasets, prompting the need for innovative solutions. Leveraging recent advances in diffusion models, synthetic data, by constructing image inputs from various annotations, proves beneficial for downstream tasks. While prior methods have separately addressed generative and perceptive models, DetDiffusion, for the first time, harmonizes both, tackling the challenges in generating effective data for perceptive models. To enhance image generation with perceptive models, we introduce perception-aware loss (P.A. loss) through segmentation, improving both quality and controllability. To boost the performance of specific perceptive models, our method customizes data augmentation by extracting and utilizing perception-aware attribute (P.A. Attr) during generation. Experimental results from the object detection task highlight DetDiffusion's superior performance, establishing a new state-of-the-art in layout-guided generation. Furthermore, image syntheses from DetDiffusion can effectively augment training data, significantly enhancing downstream detection performance. Yibo Wang 0039, Ruiyuan Gao 0001, Kai Chen 0023, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu 0001, Kai Zhang 0008 |
CVPR | 5 |
| 2024 | FeatAug-DETR: Enriching One-to-Many Matching for DETRs With Feature AugmentationabstractOne-to-one matching is a crucial design in DETR- like object detection frameworks. It enables the DETR to perform end-to-end detection. However, it also faces challenges of lacking positive sample supervision and slow convergence speed. Several recent works proposed the one-to-many matching mechanism to accelerate training and boost detection performance. We revisit these methods and model them in a unified format of augmenting the object queries. In this paper, we propose two methods that realize one-to-many matching from a different perspective of augmenting images or image features. The first method is One-to-many Matching via Data Augmentation (denoted asDataAug-DETR). It spatially transforms the images and includes multiple augmented versions of each image in the same training batch. Such a simple augmentation strategy already achieves one-to-many matching and surprisingly improves DETR's performance. The second method is One-to-many matching via Feature Augmentation (denoted asFeatAug-DETR). UnlikeDataAug-DETR, it augments the image features instead of the original images and includes multiple augmented features in the same batch to realize one-to-many matching.FeatAug-DETRsignificantly accelerates DETR training and boosts detection performance while keeping the inference speed unchanged. We conduct extensive experiments to evaluate the effectiveness of the proposed approach on DETR variants, including DAB-DETR, Deformable-DETR, and$\mathcal {H}$-Deformable-DETR. Without extra training data,FeatAug-DETRshortens the training convergence periods of Deformable-DETR [1] to 24 epochs and achieves 58.3 AP on COCOval2017set with Swin-L as the backbone. Rongyao Fang, Peng Gao 0007, Aojun Zhou, Yingjie Cai, Si Liu 0001, Jifeng Dai, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceabstractMonocular 3D Semantic Scene Completion (SSC) has garnered significant attention in recent years due to its potential to predict complex semantics and geometry shapes from a single image, requiring no 3D inputs. In this paper, we identify several critical issues in current state-of-the-art methods, including the Feature Ambiguity of projected 2D features in the ray to the 3D space, the Pose Ambiguity of the 3D convolution, and the Computation Imbalance in the 3D convolution across different depth levels. To address these problems, we devise a novel Normalized Device Coordinates scene completion network (NDC-Scene) that directly extends the 2D feature map to a Normalized Device Coordinates (NDC) space, rather than to the world space directly, through progressive restoration of the dimension of depth with deconvolution operations. Experiment results demonstrate that transferring the majority of computation from the target 3D space to the proposed normalized device coordinates space benefits monocular SSC tasks. Additionally, we design a Depth-Adaptive Dual Decoder to simultaneously upsample and fuse the 2D and 3D feature maps, further improving overall performance. Our extensive experiments confirm that the proposed method consistently outperforms state-of-the-art methods on both outdoor SemanticKITTI and indoor NYUv2 datasets. Our code are available at https://github.com/Jiawei-Yao0812/NDCScene. Jiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai, Hao Li 0069, Wanli Ouyang, Hongsheng Li 0001 |
ICCV | 4 |
| 2023 | Quantitative feature classification for breast ultrasound images using improved naive bayesabstractAbstract The quantitative feature classification of breast ultrasound images can explore the intrinsic connection between image features and lesions, and become an important basis for the diagnosis of tumor nature in clinical medicine. However, the tumor features extracted from breast ultrasound images are generally based on the overall situation, which has an impact on the classification of features in ultrasound images. Therefore, a quantitative feature classification algorithm for breast ultrasound images using improved Naive Bayes (NB) is proposed. First, histogram features, grayscale co‐generation matrix and other texture features are extracted from the ultrasound image. Second, NB is combined with decision tree calculation to improve the traditional NB algorithm. Finally, feature classification is optimized using improved NB algorithm to achieve high accuracy in the classification of ultrasound images. The results show that the contour of the lesion recognition results obtained by the proposed algorithm is the most complete, and the other regions can maximize the retention of the texture features of the original image. The classification accuracy is high, and the classification time of quantitative features of breast ultrasound images is short. It has application value in the field of breast ultrasound diagnosis. Xiaofeng Li 0012, Yupeng Sang, Xianmin Ma, Yingjie Cai |
IET Image Process. | 4 |
| 2022 | Learning a Structured Latent Space for Unsupervised Point Cloud CompletionabstractUnsupervised point cloud completion aims at estimating the corresponding complete point cloud of a partial point cloud in an unpaired manner. It is a crucial but challenging problem since there is no paired partial-complete supervision that can be exploited directly. In this work, we pro-pose a novel framework, which learns a unified and structured latent space that encoding both partial and complete point clouds. Specifically, we map a series of related par-tial point clouds into multiple complete shape and occlusion code pairs and fuse the codes to obtain their repre-sentations in the unified latent space. To enforce the learning of such a structured latent space, the proposed method adopts a series of constraints including structured ranking regularization, latent code swapping constraint, and distribution supervision on the related partial point clouds. By establishing such a unified and structured latent space, better partial-complete geometry consistency and shape completion accuracy can be achieved. Extensive experi-ments show that our proposed method consistently outper-forms state-of-the-art unsupervised methods on both syn-thetic ShapeNet and real-world KITTI, ScanNet, and Mat- terport3D datasets. Yingjie Cai, Kwan-Yee Lin, Qiang Wang 0023, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 1 |
| 2022 | Potential escalator-related injury identification and prevention based on multi-module integrated system for public health
Zeyu Jiao, Huan Lei, Hengshan Zong, Yingjie Cai, Zhenyu Zhong |
Mach. Vis. Appl. | 4 |
| 2021 | Semantic Scene Completion via Integrating Instances and Scene In-the-LoopabstractSemantic Scene Completion aims at reconstructing a complete 3D scene with precise voxel-wise semantics from a single-view depth or RGBD image. It is a crucial but challenging problem for indoor scene understanding. In this work, we present a novel framework named Scene-Instance-Scene Network (SISNet), which takes advantages of both in-stance and scene level semantic information. Our method is capable of inferring fine-grained shape details as well as nearby objects whose semantic categories are easily mixed-up. The key insight is that we decouple the instances from a coarsely completed semantic scene instead of a raw input image to guide the reconstruction of instances and the over-all scene. SISNet conducts iterative scene-to-instance (SI) and instance-to-scene (IS) semantic completion. Specifically, the SI is able to encode objects’ surrounding context for effectively decoupling instances from the scene and each instance could be voxelized into higher resolution to capture finer details. With IS, fine-grained instance information can be integrated back into the 3D scene and thus leads to more accurate semantic scene completion. Utilizing such an iterative mechanism, the scene and instance completion benefits each other to achieve higher completion accuracy. Extensively experiments show that our proposed method consistently outperforms state-of-the-art methods on both real NYU, NYUCAD and synthetic SUNCG-RGBD datasets. The code and the supplementary material will be available at https://github.com/yjcaimeow/SISNet. Yingjie Cai, Xuesong Chen 0001, Kwan-Yee Lin, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 1 |
| 2021 | Dimension reduction of multimodal data by auto-weighted local discriminant analysis
Rongxiu Lu, Yingjie Cai, Jianyong Zhu, Feiping Nie 0001, Hui Yang 0005 |
Neurocomputing | 2 |
| 2021 | Automatic Annotation Algorithm of Medical Radiological Images using Convolutional Neural Network
Xiaofeng Li 0012, Yingjie Cai |
Pattern Recognit. Lett. | 3 |
| 2020 | Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationabstractMonocular 3D object detection task aims to predict the 3D bounding boxes of objects based on monocular RGB images. Since the location recovery in 3D space is quite difficult on account of absence of depth information, this paper proposes a novel unified framework which decomposes the detection problem into a structured polygon prediction task and a depth recovery task. Different from the widely studied 2D bounding boxes, the proposed novel structured polygon in the 2D image consists of several projected surfaces of the target object. Compared to the widely-used 3D bounding box proposals, it is shown to be a better representation for 3D detection. In order to inversely project the predicted 2D structured polygon to a cuboid in the 3D physical world, the following depth recovery task uses the object height prior to complete the inverse projection transformation with the given camera projection matrix. Moreover, a fine-grained 3D box refinement scheme is proposed to further rectify the 3D detection results. Experiments are conducted on the challenging KITTI benchmark, in which our method achieves state-of-the-art detection accuracy. Yingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 0001, Xingyu Zeng, Xiaogang Wang 0001 |
AAAI | 1 |