EDBT 2026 Demo / reviewers in the wild / expert
Jichao Jiao
dblp:160/0898
· DBLP profile ↗
17ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0002-1200-5525ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Sampling-Gaussian for Stereo MatchingabstractThe soft-argmax operation is widely adopted in neural network-based stereo matching methods to enable differentiable regression of disparity. However, networks trained with soft-argmax tend to predict multimodal probability distributions due to the absence of explicit constraints on the shape of the distribution. Previous methods leveraged Laplacian distributions and cross-entropy for training but failed to effectively improve accuracy and even increased the network’s processing time. In this paper, we propose a novel method called Sampling-Gaussian as a substitute for soft-argmax. It improves accuracy without increasing inference time. We innovatively interpret the training process as minimizing the distance in vector space and propose a combined loss of L1 loss and cosine similarity loss. We leveraged the normalized discrete Gaussian distribution for supervision. Moreover, we identified two issues in previous methods and proposed extending the disparity range and employing bilinear interpolation as solutions. We have conducted comprehensive experiments to demonstrate the superior performance of our Sampling-Gaussian method. The experimental results prove that we have achieved better accuracy on five baseline methods across four datasets. Moreover, we have achieved significant improvements on small datasets and models with weaker generalization capabilities. Our method is easy to implement, and the code is available online. Baiyu Pan, Jichao Jiao, Jianxin Pang, Jun Cheng 0002 |
IROS | 3 |
| 2025 | EVMDet: EfficientViM for Small Object Detection
Jichao Jiao, Ning Li 0015, Yuqing Peng, Yingchao Zeng, Ziyi Bao, Zimo Guo |
PRCV (17) | 2 |
| 2024 | ForceGNN: A Force-Based Hypergraph Neural Network for Multi-agent Pedestrian Trajectory Forecasting
Jiaqian Zhou, Jichao Jiao, Ning Li 0015 |
ICPR (14) | 2 |
| 2024 | Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge DevicesabstractIn recent years, numerous real-time stereo matching methods have been introduced, but they often lack accuracy. These methods attempt to improve accuracy by introducing new modules or integrating traditional methods. However, the improvements are only modest. In this paper, we propose a novel strategy by incorporating knowledge distillation and model pruning to overcome the inherent trade-off between speed and accuracy. As a result, we obtained a model that maintains real-time performance while delivering high accuracy on edge devices. Our proposed method involves three key steps. Firstly, we review state-of-the-art methods and design our lightweight model by removing redundant modules from those efficient models through a comparison of their contributions. Next, we leverage the efficient model as the teacher to distill knowledge into the lightweight model. Finally, we systematically prune the lightweight model to obtain the final model. Through extensive experiments conducted on two widely-used benchmarks, Sceneflow and KITTI, we perform ablation studies to analyze the effectiveness of each module and present our state-of-the-art results. Baiyu Pan, Jichao Jiao, Jianxing Pang, Jun Cheng 0002 |
ICRA | 2 |
| 2024 | ISFP: Iterative Simultaneous Optical Flow Estimation and Interest Points Extraction NetworkabstractIn this paper, we propose ISFP, an end-to-end deep learning method for robust feature point extraction and optical flow estimation simultaneously. We observe that the structure of the deep network for feature point extraction bears similarities with the feature extraction network used in optical flow. Leveraging this observation, we merge these two tasks into a unified approach, enabling the extraction of feature points and their corresponding optical flow displacements in an end-to-end manner. ISFP serves as a comprehensive deep learning solution for feature point tracking in monocular visual SLAM systems, seamlessly replacing the front-end tasks of feature point extraction and inter-frame tracking. Our network utilizes a pre-trained interest point generation network to automatically generate interest point labels on an annotated optical flow dataset, thus creating a new dataset with ground truth interest points for training. Extensive experiments demonstrate the effectiveness and practical applicability of our approach in real-world scenarios. Compared with RAFT, our method is trained on a single FlyingChairs data set. It improves by 0.3 on the Sintel final data set, maintains comparable accuracy in clean, and the inference speed is 1.5 times that of RAFT. Weiguang Chen, Jichao Jiao, Ben Ding |
IJCNN | 2 |
| 2024 | LIOP: Tightly Coupled LiDAR-Inertial Odometry and Prior Information System for Long-term LocalizationabstractLocalization is a crucial component of robotic systems. LiDAR odometry provides robust and real-time localization but suffers from significant cumulative errors in the absence of loop closures. On the other hand, local localization based on prior maps is not affected by cumulative errors but can fail when faced with dynamic changes in the map during long-term localization scenarios. To address these issues, we propose a tightly coupled LiDAR-inertial odometry(LIO) and prior information localization system that simultaneously solves the problems of cumulative error and localization failure in dynamic scenes. Our system employs local localization to quickly track the prior map, thereby reducing odometry’s cumulative error and ensuring robust performance even in areas with significant scene changes. Additionally, we introduce a global localization strategy to rapidly correct cumulative errors when odometry tracking of the prior map fails. We evaluated our approach on the MulRan dataset, and the results demonstrate that our method effectively resolves localization failures and cumulative error issues during long-term localization. Furthermore, we validated the feasibility of our method in a real-world logistics factory environment. The results indicate that our method performs well in complex and dynamic logistics factory scenarios. Zeyuan Zhao, Jichao Jiao, Ning Li 0015, Min Pang |
IPIN | 2 |
| 2024 | Task-decoupled interactive embedding network for object detection
Mai Liu, Jichao Jiao, Ning Li 0015, Min Pang |
Mach. Learn. | 2 |
| 2024 | Multimodal remote sensing image registration based on adaptive multi-scale PIIFD
Ning Li 0015, Jichao Jiao |
Multim. Tools Appl. | 3 |
| 2024 | Hyperspectral image super-resolution via double-flow pretreatment network
Ning Li 0015, Rubin Ma, Jichao Jiao, Wangjing Qi |
Multim. Tools Appl. | 3 |
| 2024 | SPCC: A superpixel and color clustering based camouflage assessment
Ning Li 0015, Wangjing Qi, Jichao Jiao, Liqun Li |
Multim. Tools Appl. | 3 |
| 2022 | Conmw Transformer: A General Vision Transformer Backbone With Merged-Window AttentionabstractRecently, the application of Transformer in computer vision has shown us the potential of this new paradigm. However, standard multi-head Attention (MSA) faces an explosion of computational cost as the input changes from a sequence of text to an image, and MSA is computationally redundant for images. In this paper, we propose a new backbone network combining window-based attention and convolutional neural networks named ConMW Transformer, introducing convolution into the Transformer to help it converge quickly and improve accuracy. ConMW Transformer use a hierarchical architecture, an inductive bias is incorporated during tokenization and feature projection. We reduce the computational cost by performing the attention operation within windows after partition the feature map, while allowing connections between multiple heads for a more appropriate joint representation. We also use large kernel convolution after the window-based attention to merge features between windows, which help maintaining the superiority of attention in global context modelling. With only ImageNet-1K pre-training using 224 × 224 resolution, our base model can achieve 83.7% top-1 accuracy on ImageNet-1K and 49.9 mIoU for semantic segmentation on ADE20K. Jichao Jiao, Ning Li 0015, Wangjing Qi, Min Pang |
ICIP | 2 |
| 2022 | Research status and development trend of image camouflage effect evaluation
Ning Li 0015, Liqun Li, Jichao Jiao, Wangjing Qi, Xiaohu Yan |
Multim. Tools Appl. | 3 |
| 2021 | Dyn-arcFace: dynamic additive angular margin loss for deep face recognition
Jichao Jiao, Weilun Liu, Yaokai Mo, Xinping Chen |
Multim. Tools Appl. | 1 |
| 2020 | M-Sosanet: An Efficient Convolution Network Backbone For Embedding DevicesabstractIn this paper, we build a lightweight convolution neural network M-SOSAnet that combines efficiency and accuracy for edge devices. Just as DenseNet connects each layer to every other layer in the neural network. Although dense connection can effectively keep information between the middle layers, the increasing input channels by dense connection leads to resource consumption, which greatly increases the amount of parameters and is inefficient. MobileNet series use depthwise separable convolutions, which reduces the amount of calculation and parameters, but will reduce the accuracy. ESPnetv2 uses group pointwise and depthwise dilated separable convolution to learn representation from a large receptive field with fewer flops and parameters but will get a little accuracy decrease. In order to solve these problems, we propose two module: 1. Mobile SOSA Module 2. Downsampling Block with scoring mechanism, and also introduce the attention mechanism. Compare to some past methods, our model M-SOSAnet has lower flops, higher accuracy, maintains the similar amount of parameters. We evaluated model in the areas of image classification, and semantic segmentation to prove that our methods have better performance. Tangkun Zhang, Jichao Jiao, Chengkai Zhang, Yaxin Zhao, Xinping Chen |
ICIP | 2 |
| 2020 | MANet: Multimodal Attention Network based Point-View Fusion for 3D Shape Recognitionabstract3D shape recognition has attracted more and more attention as a task of 3D vision research. The proliferation of 3D data encourages various deep learning methods based on 3D data. Now there have been many deep learning models based on point-cloud data or multi-view data alone. However, in the era of big data, integrating data of two different modals to obtain a unified 3D shape descriptor is bound to improve the recognition accuracy. Therefore, this paper proposes a fusion network based on multimodal attention mechanism for 3D shape recognition. Considering the limitations of multi-view data, we introduce a soft attention scheme, which can use the global point-cloud features to filter the multi-view features, and then realize the effective fusion of the two features. More specifically, we obtain the enhanced multi-view features by mining the contribution of each multi-view image to the overall shape recognition, and then fuse the point-cloud features and the enhanced multi-view features to obtain a more discriminative 3D shape descriptor. We have performed relevant experiments on the ModelNet40 dataset, and experimental results verify the effectiveness of our method. Yaxin Zhao, Jichao Jiao, Ning Li 0015 |
ICPR | 2 |
| 2019 | MaaFace: Multiplicative and Additive Angular Margin Loss for Deep Face Recognition
Weilun Liu, Jichao Jiao, Yaokai Mo |
ICIG (3) | 2 |
| 2016 | Smallest enclosing circle-based fingerprint clustering and modified-WKNN matching algorithm for indoor positioningabstractMethods to cluster fingerprints based on Smallest-Enclosing-Circle (SEC) and to modify Weighted-K-Nearest-Neighbor (WKNN) matching algorithm for indoor fingerprint positioning system are proposed. Based on the approach to computing the smallest k-enclosing circle, the method proposed clusters fingerprints in database by introducing reference points' coordinates, instead of their received signal strength (RSS). This approach performs higher accuracy of positioning areas compared to conventional clustering algorithms, which are based on RSS. Meanwhile, this paper analyses the transmission characteristics of wireless signals in dense cluttered environments, and derives a novel path-loss-model-based weight computational method for WKNN matching algorithm. A modified-WKNN (M-WKNN) matching algorithm for indoor fingerprint positioning system is proposed and experiments are implemented in China National Grand Theatre. Results show that the location area accuracy using the proposed clustering algorithm is improved by 30% compared to that using K-means algorithm, and the positioning accuracy of M-WKNN is 11.9% and 29.1% higher than that of WKNN and KNN, respectively. Wen Liu 0002, Xiao Fu 0002, Lianming Xu, Jichao Jiao |
IPIN | 5 |