EDBT 2026 Demo / reviewers in the wild / expert
Zhi Chen 0026
dblp:05/1539-26
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Shape Prior Guided Attack: Sparser Perturbations on 3D Point CloudsabstractDeep neural networks are extremely vulnerable to malicious input data. As 3D data is increasingly used in vision tasks such as robots, autonomous driving and drones, the internal robustness of the classification models for 3D point cloud has received widespread attention. In this paper, we propose a novel method named SPGA (Shape Prior Guided Attack) to generate adversarial point cloud examples. We use shape prior information to make perturbations sparser and thus achieve imperceptible attacks. In particular, we propose a Spatially Logical Block (SLB) to apply adversarial points through sliding in the oriented bounding box. Moreover, we design an algorithm called FOFA for this type of task, which further refines the adversarial attack in the process of breaking down complicated problems into sub-problems. Compared with the methods of global perturbation, our attack method consumes significantly fewer computations, making it more efficient. Most importantly of all, SPGA can generate examples with a higher attack success rate (even in a defensive situation), less perturbation budget and stronger transferability. Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Zhidong Yu, Liusheng Huang |
AAAI | 2 |
| 2022 | AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image SynthesisabstractImage generation from text is a challenging and ill-posed task. Images generated from previous methods usually have low semantic consistency with texts and the achieved resolution is limited. To generate semantically consistent high-resolution images, we propose a novel method named AtHom, in which two attention modules are developed to extract the relationships from both independent modality and unified modality. The first is a novel Independent Modality Attention Module (IAM), which is presented to find out semantically important areas in generated images and to extract the informative context in texts. The second is a new module named Unified Semantic Space Attention Module (UAM), which is utilized to find out the relationships between extracted text context and essential areas in generated images. In particular, to bring the semantic features of texts and images closer in a unified semantic space, AtHom incorporates a homomorphic training mode by exploiting an extra discriminator to distinguish between two different modalities. Extensive experiments show that our AtHom surpasses previous methods by large margins. Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Liusheng Huang |
ACM Multimedia | 2 |
| 2022 | Public Curb Parking Demand Estimation With POI DistributionabstractWith the increasing quantity of private cars, curb parking has evolved into an important approach to mitigate parking pressure in urban cities. While some efforts have been made for the demand analysis of point-of-interest (POI) and pattern analysis of human mobility, which may indirectly reflect the parking situation in urban area, there is a lack of comprehensive models for the parking demand, so as to make a prediction for the road sections without parking lots. In this paper, by focusing on curb parking and designing a systemic framework, namedCurb Parking Demand Estimation(CPDE), we model the public parking demand in urban area, w.r.t. parking durations and regional characteristics. Specifically, we use taxi destinations and the distribution of POIs to quantitatively analyze the regional characteristics, designing corresponding features, and propose aK-means-basedLeast Square (KLS) method to relate parking characteristics, namely, the temporal parking durations and the corresponding demands, with these features. In this way, we effectively avoid the geographical sparsity of road parking sections and can finely estimate parking durations and demands for newly developed districts without parking data. Moreover, we give a strategy, namedParking Types Estimation(PTE), which projects estimated parking durations and demands onto Gaussian Mixture Model (GMM) to accurately measure the distribution of demands over different parking durations for a road section. At last, we conduct experiments on a real-world curb parking dataset in Hefei, a provincial city in China. This dataset contains parking orders of 2016 over the urban area of Hefei. The experimental results validate the effectiveness of our methods, and show that our framework outperforms the state-of-the-art baseline schemes. Yiwen Nie, Wei Yang 0011, Zhi Chen 0026, Nanxue Lu, Liusheng Huang, Huan Huang 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant KeypointsabstractIn this paper, we propose VK-Net, a neural network that learns to discover a set of category-specific keypoints from a single point cloud in an unsupervised manner. VK-Net is able to generate semantically consistent and rotation invariant keypoints across objects of the same category and different views. Particularly, we find that utilizing learned keypoints for the task of point cloud registration outperforms other traditional and learning-based approaches. Given the paired source and target point clouds, we can construct keypoint correspondences from learned keypoints using VK-Net. These keypoint correspondences are then employed to calculate a good pose initialization, after which an ICP is utilized to refine the registration. Extensive experiments on the ShapeNet dataset demonstrate that our model outperforms the state-of-the-art methods by a large margin. Zhi Chen 0026, Wei Yang 0011, Zhenbo Xu, Zhenbo Shi, Liusheng Huang |
ICASSP | 1 |
| 2021 | Mask4D: 4D Convolution Network for Light Field Occlusion RemovalabstractCurrent light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D, which is a 4D convolution-based encoder-decoder network. We propose to keep the spatial layout of SAIs and construct all SAIs as a 5D input tensor to fully exploit the spatial connection information between SAIs. In particular, except for center view reconstruction, we jointly predict the occlusion mask to disentangle the occlusion mask from the occluded content. Extensive evaluations demonstrate that our Mask4D surpasses the state-of-the-art approaches across different datasets. Moreover, visualizations show that Mask4D predicts the occlusion mask precisely and the reconstructed center view looks more realistic than other approaches. Our code will be publicly available. Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Zhenbo Shi, Liusheng Huang |
ICASSP | 4 |
| 2021 | Adversarial Attacks on Object Detectors with Limited PerturbationsabstractDeep convolutional neural networks are widely witnessed vulnerable to adversarial attacks. Recently, great progress has been achieved in attacking object detectors. However, current attacks neglect the practical utility and rely on global perturbations on the target image with a large number of patches or pixels. In this paper, we present a novel attack framework named DTTACK to fool both one-stage and two-stage object detectors with limited perturbations. A novel divergent patch shape consisting of four intersecting lines is proposed to effectively affect deep convolutional feature extraction with limited pixels. In particular, we introduce an instance-aware heat map as a self-attention module to help DTTACK focus on salient object areas, which further improves the attacking performance. Extensive experiments on PASCAL-VOC, MS-COCO, as well as an online detection system demonstrate that DTTACK surpasses the state-of-the-art methods by large margins. Zhenbo Shi, Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Liusheng Huang |
ICASSP | 4 |
| 2021 | Pointer Networks for Arbitrary-Shaped Text SpottingabstractCurrent text spotting methods perform text detection and text recognition separately. However, in complex scenes where bounding boxes of texts with various shapes are often overlapped, text detection becomes error-prone. By contrast, character detection is more non-ambiguous and easier to learn. In this paper, we present a highly efficient one-stage method named PointerNet for arbitrary-shaped text spotting. Unlike previous methods, PointerNet does not rely on text detection and opens a novel spotting-by-character-detection paradigm. In particular, to connect characters to texts, we propose a simple yet highly effective strategy named pointer that learns the 2D offset from the center of the current character to the center of the subsequent character. Evaluations demonstrate that our PointerNet achieves state-of-the-art performance and is more efficient than current methods (75ms vs. 133ms compared with FOTS). Our code will be publicly available. Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Liusheng Huang |
ICASSP | 5 |
| 2021 | Revealing the Reciprocal Relations between Self-Supervised Stereo and Monocular Depth EstimationabstractCurrent self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complementary knowledge of the two tasks. Our approach consists of three stages. In the first stage, the proposed stereo matching network termed StereoNet is trained on image pairs in a self-supervised manner. Second, we introduce an occlusion-aware distillation (OA Distillation) module, which leverages the predicted depths from StereoNet in non-occluded regions to train our monocular depth estimation network named SingleNet. At last, we design an occlusion-aware fusion module (OA Fusion), which generates more reliable depths by fusing estimated depths from StereoNet and SingleNet given the occlusion map. Furthermore, we also take the fused depths as pseudo labels to supervise StereoNet in turn, which brings StereoNet’s performance to a new height. Extensive experiments on KITTI dataset demonstrate the effectiveness of our proposed framework. We achieve new SOTA performance on both stereo and monocular depth estimation tasks. Zhi Chen 0026, Xiaoqing Ye, Wei Yang 0011, Zhenbo Xu, Xiao Tan 0001, Zhikang Zou, Errui Ding, Xinming Zhang 0001, Liusheng Huang |
ICCV | 1 |
| 2021 | Continuous Copy-Paste for One-stage Multi-object Tracking and SegmentationabstractCurrent one-step multi-object tracking and segmentation (MOTS) methods lag behind recent two-step methods. By separating the instance segmentation stage from the tracking stage, two-step methods can exploit non-video datasets as extra data for training instance segmentation. Moreover, instances belonging to different IDs on different frames, rather than limited numbers of instances in raw consecutive frames, can be gathered to allow more effective hard example mining in the training of trackers. In this paper, we bridge this gap by presenting a novel data augmentation strategy named continuous copy-paste (CCP). Our intuition behind CCP is to fully exploit the pixel-wise annotations provided by MOTS to actively increase the number of instances as well as unique instance IDs in training. Without any modifications to frameworks, current MOTS methods achieve significant performance gains when trained with CCP. Based on CCP, we propose the first effective one-stage online MOTS method named CCPNet, which generates instance masks as well as the tracking results in one shot. Our CCPNet surpasses all state-of-the-art methods by large margins (3.8% higher sMOTSA and 4.1% higher MOTSA for pedestrians on the KITTI MOTS Validation) and ranks 1st on the KITTI MOTS leaderboard. Evaluations across three datasets also demonstrate the effectiveness of both CCP and CCPNet. Our codes are publicly available at: https://github.com/detectRecog/CCP. Zhenbo Xu, Ajin Meng, Zhenbo Shi, Wei Yang 0011, Zhi Chen 0026, Liusheng Huang |
ICCV | 5 |
| 2021 | AggNet for Self-supervised Monocular Depth Estimation: Go An Aggressive Step FurtheabstractWithout appealing to exhaustive labeled data, self-supervised monocular depth estimation (MDE) plays a fundamental role in computer vision. Previous methods usually adopt a one-stage MDE network, which is insufficient to achieve high performance. In this paper, we dig deep into this task to propose an aggressive framework termed AggNet. The framework is based on a training-only progressive two-stage module to perform pseudo counter-surveillance as well as a simple yet effective dual-warp loss function between image pairs. In particular, we first propose a residual module, which follows the MDE network to learn a refined depth. The residual module takes both the initial depth generated from MDE and the initial color image as input to generate refined depth with residual depth learning. Then, the refined depth is leveraged to supervise the initial depth simultaneously during the training period. For inference, only the MDE network is retained to regress depth from a single image, which gains better performance without introducing extra computation. In addition to self-distillation loss, a simple yet effective dual-warp consistency loss is introduced to encourage the MDE network to keep depth consistency between stereo image pairs. Extensive experiments show that our AggNet achieves state-of-the-art performance on the KITTI and Make3D datasets. Zhi Chen 0026, Xiaoqing Ye, Liang Du 0004, Wei Yang 0011, Liusheng Huang, Xiao Tan 0001, Zhenbo Shi, Fumin Shen, Errui Ding |
ACM Multimedia | 1 |
| 2020 | DCNet: Dense Correspondence Neural Network for 6DoF Object Pose Estimation in Occluded Scenesabstract6DoF object pose estimation is essential for many real-world applications. Although great progress has been made, challenges still remain in estimating 6D pose for occluded objects. Current RGB-D approaches predict 6DoF pose directly, which is sensitive to occlusion in cluttered scenes. In this work, we propose DCNet, an end-to-end framework for estimating 6DoF object poses. DCNet first converts pixels in the image plane to point clouds in the camera coordinate system and then establishes dense correspondences between the camera coordinate system and the object coordinate system. Based on these two systems, we fuse 2D appearance and 3D geometric features by pixel-wise concatenation to construct dense correspondences, from which the pose is calculated through the least-squares fitting algorithm. Dense correspondences guarantee enough point pairs for a robust 6DoF pose estimation, even if the occlusion is heavy. Experimental results demonstrate that DCNet outperforms the state-of-the-art methods on LINEMOD, Occlusion LINEMOD and YCB-Video datasets, especially in terms of the robustness to occlusion scenes. Zhi Chen 0026, Wei Yang 0011, Zhenbo Xu, Xike Xie, Liusheng Huang |
ACM Multimedia | 1 |