EDBT 2026 Demo / reviewers in the wild / expert
Jiaxin Zhang 0014
dblp:287/9552
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0002-7926-4919ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAMAv2: A Vision-Centric Approach for Static Map Element AnnotationabstractThe recent advancement of Bird’s Eye View (BEV) perception algorithms requires extensive, high-quality annotated map data for effective training and deployment in real-world scenarios. However, existing HD map-based auto-labeling methods, like those found in public datasets, face significant issues related to efficiency and accuracy. For instance, the nuScenes dataset reveals considerable misalignment and inconsistency between images and their annotations, with an average reprojection error of approximately 8.03 pixels. Additionally, the dependence on HD maps limits the applicability of these methods for large-scale, real-world auto-labeling. To tackle these challenges, we introduce CAMAv2: a vision-centric approach for Consistent and Accurate Map Annotation. This pipeline primarily utilizes camera inputs to generate precise 3D annotations of static map elements, achieving high reprojection accuracy across all surrounding cameras and maintaining spatiotemporal consistency throughout the entire sequence. Importantly, CAMAv2 annotations show lower reprojection errors compared to the original nuScenes map elements, with errors of 4.96 pixels versus 8.03 pixels. Comprehensive evaluations across various public datasets confirm the feasibility and generalizability of our pipeline in diverse environments, as well as under different weather and lighting conditions. Shiyuan Chen, Jiaxin Zhang 0014, Ruohong Mei, Yingfeng Cai, Wei Sui |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | A Vision-Centric Approach for Static Map Element AnnotationabstractThe recent development of online static map element (a.k.a. HD Map) construction algorithms has raised a vast demand for data with ground truth annotations. However, available public datasets currently cannot provide high-quality training data regarding consistency and accuracy. To this end, we present CAMA: a vision-centric approach for Consistent and Accurate Map Annotation. Without LiDAR inputs, our proposed framework can still generate high-quality 3D annotations of static map elements. Specifically, the annotation can achieve high reprojection accuracy across all surrounding cameras and is spatial-temporal consistent across the whole sequence. We apply our proposed framework to the popular nuScenes dataset to provide efficient and highly accurate annotations. Compared with the original nuScenes static map element, models trained with annotations from CAMA achieve lower reprojection errors (e.g., 4.73 vs. 8.03 pixels). Jiaxin Zhang 0014, Shiyuan Chen, Ruohong Mei, Qian Zhang 0009, Wei Sui |
ICRA | 1 |
| 2024 | VRSO: Visual-Centric Reconstruction for Static Object AnnotationabstractAs a part of the perception results of intelligent driving systems, static object detection (SOD) in 3D space provides crucial cues for driving environment understanding. With the rapid deployment of deep neural networks for SOD tasks, the demand for high-quality training samples soars. The traditional, also reliable, way is manual labelling over the dense LiDAR point clouds and reference images. Though most public driving datasets adopt this strategy to provide SOD ground truth (GT), it is still expensive and time-consuming in practice. This paper introduces VRSO, a visual-centric approach for static object annotation. Experiments on the Waymo Open Dataset show that the mean reprojection error from VRSO annotation is only 2.6 pixels, around four times lower than the Waymo Open Dataset labels (10.6 pixels). VRSO is distinguished in low cost, high efficiency, and high quality: (1) It recovers static objects in 3D space with only camera images as input, and (2) manual annotation is barely involved since GT for SOD tasks is generated based on an automatic reconstruction and annotation pipeline. Chenyao Yu, Yingfeng Cai, Jiaxin Zhang 0014, Hui Kong 0001, Wei Sui |
IROS | 3 |
| 2021 | Deep Online Correction for Monocular Visual OdometryabstractIn this work, we propose a novel deep online correction (DOC) framework for monocular visual odometry. The whole pipeline has two stages: First, depth maps and initial poses are obtained from convolutional neural networks (CNNs) trained in self-supervised manners. Second, the poses predicted by CNNs are further improved by minimizing photometric errors via gradient updates of poses during inference phases. The benefits of our proposed method are twofold: 1) Different from online-learning methods, DOC does not need to calculate gradient propagation for parameters of CNNs. Thus, it saves more computation resources during inference phases. 2) Unlike hybrid methods that combine CNNs with traditional methods, DOC fully relies on deep learning (DL) frameworks. Though without complex back-end optimization modules, our method achieves outstanding performance with relative transform error (RTE) = 2.0% on KITTI Odometry benchmark for Seq. 09, which outperforms traditional monocular VO frameworks and is comparable to hybrid methods. Jiaxin Zhang 0014, Wei Sui, Xinggang Wang, Wenming Meng, Hongmei Zhu, Qian Zhang 0009 |
ICRA | 1 |
| 2020 | Binarizing Weights Wisely for Edge Intelligence: Guide for Partial Binarization of Deconvolution-Based GeneratorsabstractThis article explores the weight binarization of the deconvolution-based generator in a generative adversarial network (GAN) for memory saving and speedup of image construction on the edge. This article suggests that different from convolutional neural networks (including the discriminator) where all layers can be binarized, only some of the layers in the generator can be binarized without significant performance loss. Supported by theoretical analysis and verified by experiments, a direct metric based on the dimension of deconvolution operations is established, which can be used to quickly decide which layers in a generator can be binarized. Our results also indicate that both the generator and the discriminator should be binarized simultaneously for balanced competition and better performance during training. The experimental results on CelebA dataset with DCGAN and original loss functions suggest that directly applying state-of-the-art binarization techniques to all the layers of the generator will lead to 2.83× performance loss measured by sliced Wasserstein distance compared with the original generator, while applying them to selected layers only can yield up to 25.81× saving in memory consumption, and 1.96× and 1.32× speedup in inference and training, respectively, with little performance loss. Similar conclusions can also be drawn on other loss functions for different GANs. Jinglan Liu, Jiaxin Zhang 0014, Yukun Ding, Xiaowei Xu 0004, Meng Jiang 0001, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |