EDBT 2026 Demo / reviewers in the wild / expert
Zhenbo Xu
dblp:57/10208
· DBLP profile ↗
35ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorComputer networks · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace AdaptationabstractDianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He 0001 |
ACL (1) | 5 |
| 2026 | SCVQ: Sparse-Compensated Vector Quantization for Large Language ModelsabstractZixuan Zhou, Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yujun Diao, Zicheng Kong, Dehua Ma, Zhenbo Xu, Pei Pei Li, Zhaofeng He 0001 |
ACL (1) | 5 |
| 2025 | Psyche-Wave: Fusing Vector-Quantized Morphology and LLM-Inferred Semantics from Millimeter-Wave SCG for Psychological State DecodingabstractThis paper introduces Psyche-Wave, a novel paradigm for non-contact psychological state assessment, addressing the challenge that existing methods struggle to reconcile signal representation robustness with deep physiological semantic understanding. The proposed framework is built upon high-fidelity Seismocardiogram (SCG) and respiratory signals, captured by a proprietary high-sampling-rate millimeter-wave (mmWave) radar system. Psyche-Wave features a parallel dual-branch architecture for complementary feature extraction. The first, a Data-Driven Morphological Branch, employs Vector Quantization (VQ) to encode the Mel spectrogram of the SCG signal into a codebook-based representation, yielding a noise-resilient morphological embedding. The second, a Knowledge-Driven Semantic Branch, leverages a Large Language Model (LLM) to infer deep contextual relationships from medically significant physiological parameters—including heart rate variability, cardiac time intervals, and cardiopulmonary coupling—outputting a rich semantic embedding. These complementary embeddings are then integrated through a dedicated fusion module and passed to a downstream classifier for precise emotion and personality trait evaluation. Comprehensive evaluations on a newly collected high-fidelity dataset, referred to as mmHeart-Pro, and the public ReMAP dataset demonstrate state-of-the-art performance. This work pioneers a new path that fuses data-driven morphological analysis with knowledge-driven semantic reasoning, significantly advancing the accuracy and interpretability of non-contact psychological sensing. Yiwei Ru, Zhenbo Xu, Yanlin Xu, Huijia Wu, Zhaofeng He 0001, Zhenan Sun |
BIBM | 2 |
| 2025 | FruitMMBench: A Multi-modal Benchmark for Fruit Quality AssessmentabstractThe rapid advancement of Large Vision-Language Models (LVLMs) has brought notable improvements in tasks like visual recognition and multi-modal understanding, demonstrating significant potential in real-world applications. However, their performances on issues related to daily life such as fruit quality assessment have never been explored due to the lack of a benchmark for fruit quality assessment. To bridge this gap, we introduce FruitMMBench, a comprehensive multi-modal benchmark designed to assess the ability of fruit quality assessment. A comprehensive metric is carefully designed by jointly considering many aspects. The resulting large-scale FruitMMBench contains 5,465 collected fruit images and dedicated quality labels including the following aspects: fruit classification, quantity recognition, maturity, surface condition, quality status, and edibility recommendation. By extensive evaluations on FruitMMBench, we find popular LVLMs struggle to provide reliable results in fruit quality assessment and their performances vary greatly. The test results reveal that the ability of current LVLMs to analyze fruit quality in real-world scenarios is still weak and needs to be paid attention to and enhanced in the future. All evaluation codes and dataset will be publicly accessible shortly. Gong Huang, Zhenbo Xu, Qinghong Yang |
ICASSP | 4 |
| 2025 | Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMsabstractFood recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the food recognition task still lags behind traditional vision models. To bridge this gap, this paper proposes two methods to improve the food recognition capabilities of LVLMs: Retrieval-Augmented Recognition (RAR) and Domain-Adaptive Recognition (DAR). On the one hand, the training-free RAR utilizes a vision model to retrieve relevant image-category pairs from an image-category memory pre-built from the training set, thus incorporating the categorical information into the input of LVLMs to enhance food recognition performance. On the other hand, DAR employs a two-stage training process by first pre-training LVLMs on diverse food analysis tasks and then fine-tuning LVLMs using food recognition data. Extensive evaluations on two large-scale food recognition datasets demonstrate that both RAR and DAR improve the food recognition performance of LVLMs and, compred to RAR, DAR achieves a higher precision that outperforms traditional vision models. Dehua Ma, Zhenbo Xu, Tianshun Xing, Huijia Wu, Zhaofeng He 0001 |
ICASSP | 2 |
| 2025 | FoodWeight1.4M: A Large-scale Multi-modal Dataset for Weight EstimationabstractLarge vision language models (VLMs) excel in visual tasks but struggle with weight estimation, hindering 3D perception and embodied intelligence. To address the lack of large-scale weight datasets, we present FoodWeight1.4M, derived from real-world supermarket scenarios. It contains 1.4 million high-quality images across 1,550 food categories, with weights precisely measured and rigorously filtered, making it the first large-scale weight estimation dataset. The weight estimation performance of current VLMs were tested and found to be unsatisfactory, which can be significantly improved by instruction tuning using Food-Weight1.4M. Moreover, we propose two strategies, Category-Guided and Reference Calibration, to enhance weight estimation without fine-tuning. Experiments confirm their effectiveness in improving multi-modal weight perception. Furthermore, experimental results show that pre-training on FoodWeight1.4M can benefit other food analysis tasks. Our dataset will be publicly available soon. Zhenbo Xu, Dehua Ma, Liuyu Xiang, Huijia Wu, Zhaofeng He 0001 |
ICME | 2 |
| 2025 | RecipeRAG: Advancing Recipe Generation with Reinforced Retrieval Augmented GenerationabstractGenerating accurate recipes from dish images is a challenging task that requires a deep understanding of food categories, ingredient combinations, cooking methods, and context. Current works mainly rely on the two-stage training method or supervised fine-tuning of vision-language models (VLMs). Two-stage models typically first predict ingredients from images and then generate recipes based on both ingredients and images. However, accumulated errors in ingredient prediction often lead to inaccurate recipes. Fine-tuning VLMs only fit the statistical patterns of the training data, lacking deep reasoning capabilities, which leads to severe hallucinations in the generated recipes. In this paper, we introduce a novel reinforced retrieval-augmented generation framework named RecipeRAG for recipe generation, and compare the supervised fine-tuning (SFT) paradigm and the reinforcement fine-tuning (RFT) paradigm. To effectively retrieve recipes relevant to the query image, we improve CLIP to obtain IR-CLIP as both our retriever and re-ranker by integrating metric learning and contrastive learning. The retrieved recipes are then used to enhance the generated results, improving accuracy and reducing hallucinations. However, the SFT VLM often fails to judge the quality of the retrieved recipe information and perform the complex recipe generation. Therefore, we furthermore investigate the two-phase RFT training framework. Firstly, the cold-start phase uses generated Chain-of-Thought (CoT) data for SFT to activate the reasoning capabilities of VLMs. Then, the reinforcement learning phase utilizes Group Relative Policy Optimization (GRPO) to generate multiple reasoning-answer pairs, further enhancing the generalization ability of VLMs in recipe generation tasks. Extensive evaluations on the large-scale Recipe1M dataset demonstrate that RecipeRAG outperforms all previous methods in recipe generation and exhibits strong generalization ability under the RL paradigm. Zhenbo Xu, Dehua Ma, Fei Liu 0008, Gong Huang, Zhaofeng He 0001 |
ACM Multimedia | 2 |
| 2024 | Instance-Based Continual Learning: A Real-World Dataset and Baseline for Fresh RecognitionabstractReal-time learning on real-world data streams with temporal relations is essential for intelligent agents. However, current online Continual Learning (CL) benchmarks adopt the mini-batch setting and are composed of temporally unrelated and disjoint tasks as well as pre-set class boundaries. In this paper, we delve into a real-world CL scenario for fresh recognition where algorithms are required to recognize a huge variety of products to facilitate the checkout speed. Products mainly consists of packaged cereals, seasonal fruits, and vegetables from local farms or shipped from overseas. Since algorithms process instance streams consisting of sequential images, we name this real-world CL problem as Instance-Based Continual Learning (IBCL) . Different from the current online CL setting, algorithms are required to perform instant testing and learning upon each incoming instance. Moreover, IBCL has no task boundaries or class boundaries and allows the evolution and the forgetting of old samples within each class. To promote the researches on real CL challenges, we propose the first real-world CL dataset coined the Continual Fresh Recognition (CFR) dataset, which consists of fresh recognition data streams (766 K labelled images in total) collected from 30 supermarkets. Based on the CFR dataset, we extensively evaluate the performance of current online CL methods under various settings and find that current prominent online CL methods operate at high latency and demand significant memory consumption to cache old samples for replaying. Therefore, we make the first attempt to design an efficient and effective Instant Training-Free Learning (ITFL) framework for IBCL. ITFL consists of feature extractors trained in the metric learning manner and reformulates CL as a temporal classification problem among several most similar classes. Unlike current online CL methods that cache image samples (150 KB per image) and rely on training to learn new knowledge, our framework only caches features (2 KB per image) and is free of training in deployment. Extensive evaluations across three datasets demonstrate that our method achieves comparable recognition accuracy to current methods with lower latency and less resource consumption. Our codes and datasets will be publicly available at https://github.com/detectRecog/IBCL . Zhenbo Xu, Hai-Miao Hu, Wenming Tan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | One-Shot Neural Band Selection for Spectral RecoveryabstractBand selection has a great impact on the spectral recovery quality. To solve this ill-posed inverse problem, most band selection methods adopt hand-crafted priors or exploit clustering or sparse regularization constraints to find most prominent bands. These methods are either very slow due to the computational cost of repeatedly training with respect to different selection frequencies or different band combinations. Many traditional methods rely on the scene prior and thus are not applicable to other scenarios. In this paper, we present a novel one-shot Neural Band Selection (NBS) framework for spectral recovery. Unlike conventional searching approaches with a discrete search space and a non-differentiable search strategy, our NBS is based on the continuous relaxation of the band selection process, thus allowing efficient band search using gradient descent. To enable the compatibility for selecting any number of bands in one-shot, we further exploit the band-wise correlation matrices to progressively suppress similar adjacent bands. Extensive evaluations on the NTIRE 2022 Spectral Reconstruction Challenge demonstrate that our NBS achieves consistent performance gains over competitive baselines when examined with four different spectral recovery methods. Our code will be publicly available. Hai-Miao Hu, Zhenbo Xu, Wenshuai Xu, You Song, YiTao Zhang, Zhilin Han, Ajin Meng |
ICASSP | 2 |
| 2023 | A Solution to Co-occurence Bias: Attributes Disentanglement via Mutual Information Minimization for Pedestrian Attribute RecognitionabstractRecent studies on pedestrian attribute recognition progress with either explicit or implicit modeling of the co-occurence among attributes. Considering that this known a prior is highly variable and unforeseeable regarding the specific scenarios, we show that current methods can actually suffer in generalizing such fitted attributes interdependencies onto scenes or identities off the dataset distribution, resulting in the underlined bias of attributes co-occurence. To render models robust in realistic scenes, we propose the attributes-disentangled feature learning to ensure the recognition of an attribute not inferring on the existence of others, and which is sequentially formulated as a problem of mutual information minimization. Rooting from it, practical strategies are devised to efficiently decouple attributes, which substantially improve the baseline and establish state-of-the-art performance on realistic datasets like PETAzs and RAPzs. Hai-Miao Hu, Jinzuo Yu, Zhenbo Xu, Weiqing Lu, Yuran Cao |
IJCAI | 4 |
| 2023 | Reinforcement Learning-based Adversarial Attacks on Object Detectors using Reward ShapingabstractIn the field of object detector attacks, previous methods primarily rely on fixed gradient optimization or patch-based cover techniques, often leading to suboptimal attack performance and excessive distortions. To address these limitations, we propose a novel attack method, Interactive Reinforcement-based Sparse Attack (IRSA), which employs Reinforcement Learning (RL) to discover the vulnerabilities of object detectors and systematically generate erroneous results. Specifically, we formulate the process of seeking optimal margins for adversarial examples as a Markov Decision Process (MDP). We tackle the RL convergence difficulty through innovative reward functions and a composite optimization method for effective and efficient policy training. Moreover, the perturbations generated by IRSA are more subtle and difficult to detect while requiring less computational effort. Our method also demonstrates strong generalization capabilities against various object detectors. In summary, IRSA is a refined, efficient, and scalable interactive, iterative, end-to-end algorithm. Zhenbo Shi, Wei Yang 0011, Zhenbo Xu, Zhidong Yu, Liusheng Huang |
ACM Multimedia | 3 |
| 2022 | Shape Prior Guided Attack: Sparser Perturbations on 3D Point CloudsabstractDeep neural networks are extremely vulnerable to malicious input data. As 3D data is increasingly used in vision tasks such as robots, autonomous driving and drones, the internal robustness of the classification models for 3D point cloud has received widespread attention. In this paper, we propose a novel method named SPGA (Shape Prior Guided Attack) to generate adversarial point cloud examples. We use shape prior information to make perturbations sparser and thus achieve imperceptible attacks. In particular, we propose a Spatially Logical Block (SLB) to apply adversarial points through sliding in the oriented bounding box. Moreover, we design an algorithm called FOFA for this type of task, which further refines the adversarial attack in the process of breaking down complicated problems into sub-problems. Compared with the methods of global perturbation, our attack method consumes significantly fewer computations, making it more efficient. Most importantly of all, SPGA can generate examples with a higher attack success rate (even in a defensive situation), less perturbation budget and stronger transferability. Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Zhidong Yu, Liusheng Huang |
AAAI | 3 |
| 2022 | AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image SynthesisabstractImage generation from text is a challenging and ill-posed task. Images generated from previous methods usually have low semantic consistency with texts and the achieved resolution is limited. To generate semantically consistent high-resolution images, we propose a novel method named AtHom, in which two attention modules are developed to extract the relationships from both independent modality and unified modality. The first is a novel Independent Modality Attention Module (IAM), which is presented to find out semantically important areas in generated images and to extract the informative context in texts. The second is a new module named Unified Semantic Space Attention Module (UAM), which is utilized to find out the relationships between extracted text context and essential areas in generated images. In particular, to bring the semantic features of texts and images closer in a unified semantic space, AtHom incorporates a homomorphic training mode by exploiting an extra discriminator to distinguish between two different modalities. Extensive experiments show that our AtHom surpasses previous methods by large margins. Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Liusheng Huang |
ACM Multimedia | 3 |
| 2022 | Revealing the real-world applicable setting of online continual learningabstractThe motivation of online continual learning (CL) is training agents to learn from an infinite stream of data and quickly accommodate changes in the data distribution. However, current online CL datasets are synthesized by common classification datasets by splitting all classes into disjoint tasks where disjoint task streams have little temporal relations, resulting in a CL setting far from realistic. In this paper, we ask two questions: (i) What are the characteristics of real-world CL scenarios? (ii) How existing methods perform on real-world CL scenarios? To answer the first question, we propose the first realistic CL setting coined instance-based continual learning (IBCL). IBCL has no task or class boundaries and requires algorithms to predict and learn from instance streams simultaneously. The life cycles of classes under IBCL are dynamic and instances belonging to the same class might evolve over time. For each sequentially arrival instance, algorithms are required to give the recognition result and then perform changes based on its label. No additional training resource are available except for the instance stream in evaluation. To answer the second question, on CORe50 and mini-ImageNet, we compare current online CL methods under the IBCL setting with both the traditional ResNet18 backbone as well as the recent transformer-based backbone ViT on the IBCL setting. Three aspects including the recognition performance, the latency, and the memory usage of current methods are analyzed. Experiment results show that current online CL methods perform poorly in the real CL scenarios, and methods using the transformer-based backbone perform better than the CNN-based counterparts. Zhenbo Xu, Hai-Miao Hu |
MMSP | 1 |
| 2022 | Segment as Points for Efficient and Effective Online Multi-Object Tracking and SegmentationabstractCurrent multi-object tracking and segmentation (MOTS) methods follow the tracking-by-detection paradigm and adopt 2D or 3D convolutions to extract instance embeddings for instance association. However, due to the large receptive field of deep convolutional neural networks, the foreground areas of the current instance and the surrounding areas containing the nearby instances or environments are usually mixed up in the learned instance embeddings, resulting in ambiguities in tracking. In this paper, we propose a highly effective method for learning instance embeddings based on segments by converting the compact image representation to un-ordered 2D point cloud representation. In this way, the non-overlapping nature of instance segments can be fully exploited by strictly separating the foreground point cloud and the background point cloud. Moreover, multiple informative data modalities are formulated as point-wise representations to enrich point-wise features. For each instance, the embedding is learned on the foreground 2D point cloud, the environment 2D point cloud, and the smallest circumscribed bounding box. Then, similarities between instance embeddings are measured for the inter-frame association. In addition, to enable the practical utility of MOTS, we modify the one-stage instance segmentation method SpatialEmbedding for instance segmentation. The resulting efficient and effective framework, named PointTrackV2, outperforms all the state-of-the-art methods including 3D tracking methods by large margins (4.8 percent higher sMOTSA for pedestrians over MOTSFusion) with the near real-time speed (20 FPS evaluated on a single 2080Ti). Extensive evaluations on three datasets demonstrate both the effectiveness and efficiency of our method. Furthermore, as crowded scenes for cars are insufficient in current MOTS datasets, we provide a more challenging dataset named APOLLO MOTS with a much higher instance density. Zhenbo Xu, Wei Yang 0011, Wei Zhang 0197, Xiao Tan 0001, Huan Huang 0004, Liusheng Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant KeypointsabstractIn this paper, we propose VK-Net, a neural network that learns to discover a set of category-specific keypoints from a single point cloud in an unsupervised manner. VK-Net is able to generate semantically consistent and rotation invariant keypoints across objects of the same category and different views. Particularly, we find that utilizing learned keypoints for the task of point cloud registration outperforms other traditional and learning-based approaches. Given the paired source and target point clouds, we can construct keypoint correspondences from learned keypoints using VK-Net. These keypoint correspondences are then employed to calculate a good pose initialization, after which an ICP is utilized to refine the registration. Extensive experiments on the ShapeNet dataset demonstrate that our model outperforms the state-of-the-art methods by a large margin. Zhi Chen 0026, Wei Yang 0011, Zhenbo Xu, Zhenbo Shi, Liusheng Huang |
ICASSP | 3 |
| 2021 | Mask4D: 4D Convolution Network for Light Field Occlusion RemovalabstractCurrent light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D, which is a 4D convolution-based encoder-decoder network. We propose to keep the spatial layout of SAIs and construct all SAIs as a 5D input tensor to fully exploit the spatial connection information between SAIs. In particular, except for center view reconstruction, we jointly predict the occlusion mask to disentangle the occlusion mask from the occluded content. Extensive evaluations demonstrate that our Mask4D surpasses the state-of-the-art approaches across different datasets. Moreover, visualizations show that Mask4D predicts the occlusion mask precisely and the reconstructed center view looks more realistic than other approaches. Our code will be publicly available. Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Zhenbo Shi, Liusheng Huang |
ICASSP | 3 |
| 2021 | Adversarial Attacks on Object Detectors with Limited PerturbationsabstractDeep convolutional neural networks are widely witnessed vulnerable to adversarial attacks. Recently, great progress has been achieved in attacking object detectors. However, current attacks neglect the practical utility and rely on global perturbations on the target image with a large number of patches or pixels. In this paper, we present a novel attack framework named DTTACK to fool both one-stage and two-stage object detectors with limited perturbations. A novel divergent patch shape consisting of four intersecting lines is proposed to effectively affect deep convolutional feature extraction with limited pixels. In particular, we introduce an instance-aware heat map as a self-attention module to help DTTACK focus on salient object areas, which further improves the attacking performance. Extensive experiments on PASCAL-VOC, MS-COCO, as well as an online detection system demonstrate that DTTACK surpasses the state-of-the-art methods by large margins. Zhenbo Shi, Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Liusheng Huang |
ICASSP | 3 |
| 2021 | Pointer Networks for Arbitrary-Shaped Text SpottingabstractCurrent text spotting methods perform text detection and text recognition separately. However, in complex scenes where bounding boxes of texts with various shapes are often overlapped, text detection becomes error-prone. By contrast, character detection is more non-ambiguous and easier to learn. In this paper, we present a highly efficient one-stage method named PointerNet for arbitrary-shaped text spotting. Unlike previous methods, PointerNet does not rely on text detection and opens a novel spotting-by-character-detection paradigm. In particular, to connect characters to texts, we propose a simple yet highly effective strategy named pointer that learns the 2D offset from the center of the current character to the center of the subsequent character. Evaluations demonstrate that our PointerNet achieves state-of-the-art performance and is more efficient than current methods (75ms vs. 133ms compared with FOTS). Our code will be publicly available. Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Liusheng Huang |
ICASSP | 3 |
| 2021 | Revealing the Reciprocal Relations between Self-Supervised Stereo and Monocular Depth EstimationabstractCurrent self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complementary knowledge of the two tasks. Our approach consists of three stages. In the first stage, the proposed stereo matching network termed StereoNet is trained on image pairs in a self-supervised manner. Second, we introduce an occlusion-aware distillation (OA Distillation) module, which leverages the predicted depths from StereoNet in non-occluded regions to train our monocular depth estimation network named SingleNet. At last, we design an occlusion-aware fusion module (OA Fusion), which generates more reliable depths by fusing estimated depths from StereoNet and SingleNet given the occlusion map. Furthermore, we also take the fused depths as pseudo labels to supervise StereoNet in turn, which brings StereoNet’s performance to a new height. Extensive experiments on KITTI dataset demonstrate the effectiveness of our proposed framework. We achieve new SOTA performance on both stereo and monocular depth estimation tasks. Zhi Chen 0026, Xiaoqing Ye, Wei Yang 0011, Zhenbo Xu, Xiao Tan 0001, Zhikang Zou, Errui Ding, Xinming Zhang 0001, Liusheng Huang |
ICCV | 4 |
| 2021 | Continuous Copy-Paste for One-stage Multi-object Tracking and SegmentationabstractCurrent one-step multi-object tracking and segmentation (MOTS) methods lag behind recent two-step methods. By separating the instance segmentation stage from the tracking stage, two-step methods can exploit non-video datasets as extra data for training instance segmentation. Moreover, instances belonging to different IDs on different frames, rather than limited numbers of instances in raw consecutive frames, can be gathered to allow more effective hard example mining in the training of trackers. In this paper, we bridge this gap by presenting a novel data augmentation strategy named continuous copy-paste (CCP). Our intuition behind CCP is to fully exploit the pixel-wise annotations provided by MOTS to actively increase the number of instances as well as unique instance IDs in training. Without any modifications to frameworks, current MOTS methods achieve significant performance gains when trained with CCP. Based on CCP, we propose the first effective one-stage online MOTS method named CCPNet, which generates instance masks as well as the tracking results in one shot. Our CCPNet surpasses all state-of-the-art methods by large margins (3.8% higher sMOTSA and 4.1% higher MOTSA for pedestrians on the KITTI MOTS Validation) and ranks 1st on the KITTI MOTS leaderboard. Evaluations across three datasets also demonstrate the effectiveness of both CCP and CCPNet. Our codes are publicly available at: https://github.com/detectRecog/CCP. Zhenbo Xu, Ajin Meng, Zhenbo Shi, Wei Yang 0011, Zhi Chen 0026, Liusheng Huang |
ICCV | 1 |
| 2021 | MDANet: Multi-Modal Deep Aggregation Network for Depth CompletionabstractDepth completion aims to recover the dense depth map from sparse depth data and RGB image respectively. However, due to the huge difference between the multi-modal signal input, vanilla convolutional neural network and simple fusion strategy cannot extract features from sparse data and aggregate multi-modal information effectively. To tackle this problem, we design a novel network architecture that takes full advantage of multi-modal features for depth completion. An effective Pre-completion algorithm is first put forward to increase the density of the input depth map and to provide distribution priors. Moreover, to effectively fuse the image features and the depth features, we propose a multi-modal deep aggregation block that consists of multiple connection and aggregation pathways for deeper fusion. Furthermore, based on the intuition that semantic image features are beneficial for accurate contour, we introduce the deformable guided fusion layer to guide the generation of the dense depth map. The resulting architecture, called MDANet, outperforms all the stateof-the-art methods on the popular KITTI Depth Completion Benchmark, meanwhile with fewer parameters than recent methods. The code of this work will be available at https://github.com/USTC-Keyanjie/MDANet_ICRA2021. Yanjie Ke, Wei Yang 0011, Zhenbo Xu, Dayang Hao, Liusheng Huang |
ICRA | 4 |
| 2021 | ABPNet: Adaptive Background Modeling for Generalized Few Shot SegmentationabstractExisting Few Shot Segmentation (FS-Seg) methods mostly study a restricted setting where only foreground and background are required to be discriminated and fall short at discriminating multiple classes. In this paper, we focus on a challenging but more practical variant: Generalized Few Shot Segmentation (GFS-Seg), where all SEEN and UNSEEN classes are segmented simultaneously. Previous methods treat the background as a regular class, leading to difficulty in differentiating UNSEEN classes from it at the test stage. To address this issue, we propose Adaptive Background Modeling and Prototype Query Network (ABPNet), in which the background is formulated as the complement of the set of interested classes. With the help of the attention mechanism and a novel meta-training strategy, it learns an effective set difference function that predicts task-specific background adaptively. Furthermore, we design a Prototype Querying (PQ) module that effectively transfers the learned knowledge to UNSEEN classes with a neural dictionary. Experimental results demonstrate that ABPNet significantly outperforms the state-of-the-art method CAPL on PASCAL-5i and COCO-20i, especially on UNSEEN classes. Also, without retraining, ABPNet can generalize well to FS-Seg. Kaiqi Dong, Wei Yang 0011, Zhenbo Xu, Liusheng Huang, Zhidong Yu |
ACM Multimedia | 3 |
| 2020 | ZoomNet: Part-Aware Adaptive Zooming Neural Network for 3D Object Detectionabstract3D object detection is an essential task in autonomous driving and robotics. Though great progress has been made, challenges remain in estimating 3D pose for distant and occluded objects. In this paper, we present a novel framework named ZoomNet for stereo imagery-based 3D detection. The pipeline of ZoomNet begins with an ordinary 2D object detection model which is used to obtain pairs of left-right bounding boxes. To further exploit the abundant texture cues in rgb images for more accurate disparity estimation, we introduce a conceptually straight-forward module – adaptive zooming, which simultaneously resizes 2D instance bounding boxes to a unified resolution and adjusts the camera intrinsic parameters accordingly. In this way, we are able to estimate higher-quality disparity maps from the resized box images then construct dense point clouds for both nearby and distant objects. Moreover, we introduce to learn part locations as complementary features to improve the resistance against occlusion and put forward the 3D fitting score to better estimate the 3D detection quality. Extensive experiments on the popular KITTI 3D detection dataset indicate ZoomNet surpasses all previous state-of-the-art methods by large margins (improved by 9.4% on APbv (IoU=0.7) over pseudo-LiDAR). Ablation study also demonstrates that our adaptive zooming strategy brings an improvement of over 10% on AP3d (IoU=0.7). In addition, since the official KITTI benchmark lacks fine-grained annotations like pixel-wise part locations, we also present our KFG dataset by augmenting KITTI with detailed instance-wise annotations including pixel-wise part location, pixel-wise disparity, etc.. Both the KFG dataset and our codes will be publicly available at https://github.com/detectRecog/ZoomNet. Zhenbo Xu, Wei Zhang 0197, Xiaoqing Ye, Xiao Tan 0001, Wei Yang 0011, Shilei Wen, Errui Ding, Ajin Meng, Liusheng Huang |
AAAI | 1 |
| 2020 | Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object DetectionabstractObject detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point cloud data. Designing robust feature representation against such appearance changes is hence the key issue in a 3D object detection method. In this paper, we innovatively propose a domain adaptation like approach to enhance the robustness of the feature representation. More specifically, we bridge the gap between the perceptual domain where the feature comes from a real scene and the conceptual domain where the feature is extracted from an augmented scene consisting of non-occlusion point cloud rich of detailed information. This domain adaptation approach mimics the functionality of the human brain when proceeding object perception. Extensive experiments demonstrate that our simple yet effective approach fundamentally boosts the performance of 3D point cloud object detection and achieves the state-of-the-art results. Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Jianfeng Feng, Zhenbo Xu, Errui Ding, Shilei Wen |
CVPR | 5 |
| 2020 | Segment as Points for Efficient Online Multi-Object Tracking and Segmentation
Zhenbo Xu, Wei Zhang 0197, Xiao Tan 0001, Wei Yang 0011, Huan Huang 0004, Shilei Wen, Errui Ding, Liusheng Huang |
ECCV (1) | 1 |
| 2020 | DCNet: Dense Correspondence Neural Network for 6DoF Object Pose Estimation in Occluded Scenesabstract6DoF object pose estimation is essential for many real-world applications. Although great progress has been made, challenges still remain in estimating 6D pose for occluded objects. Current RGB-D approaches predict 6DoF pose directly, which is sensitive to occlusion in cluttered scenes. In this work, we propose DCNet, an end-to-end framework for estimating 6DoF object poses. DCNet first converts pixels in the image plane to point clouds in the camera coordinate system and then establishes dense correspondences between the camera coordinate system and the object coordinate system. Based on these two systems, we fuse 2D appearance and 3D geometric features by pixel-wise concatenation to construct dense correspondences, from which the pose is calculated through the least-squares fitting algorithm. Dense correspondences guarantee enough point pairs for a robust 6DoF pose estimation, even if the occlusion is heavy. Experimental results demonstrate that DCNet outperforms the state-of-the-art methods on LINEMOD, Occlusion LINEMOD and YCB-Video datasets, especially in terms of the robustness to occlusion scenes. Zhi Chen 0026, Wei Yang 0011, Zhenbo Xu, Xike Xie, Liusheng Huang |
ACM Multimedia | 3 |
| 2018 | Towards End-to-End License Plate Detection and Recognition: A Large Dataset and Baseline
Zhenbo Xu, Wei Yang 0011, Ajin Meng, Nanxue Lu, Huan Huang 0004, Changchun Ying, Liusheng Huang |
ECCV (13) | 1 |
| 2018 | A robust and efficient method for license plate recognitionabstractLicense plate recognition is an essential step in automatic license plate recognition since it is a key technology to recognize detected license plates. Though there are extensive researches on license plate recognition, it is still challenging to recognize license plates under conditions like great tilt angles, uneven illuminations, and distortions. Based on the observation that an accurate shape correction can significantly improve the recognition accuracy on these images, this paper proposes a robust methodology named LCR for license plate recognition free of conventional image analysis operations. This approach is based on three neural networks for three different purposes: (i) predicting the locations of four vertices; (ii) predicting cutting locations; (iii) character classification. To the best of our knowledge, LCR is the first to address shape correction by designing neural networks to accurately predict the coordinates of license plates vertices. Experiments on over 250,000 unique images show that LCR significantly outperforms several state-of-the-art license plate recognition approaches. Moreover, in evaluations, the application of shape correction significantly improve the recognition accuracy. Ajin Meng, Wei Yang 0011, Zhenbo Xu, Huan Huang 0004, Liusheng Huang, Changchun Ying |
ICPR | 3 |
| 2018 | Link Us if You Can: Enabling Unlinkable Communication on the InternetabstractFor online conversations with top privacy, we often need to erase the existing contact behavior. Thus we want communications in which adversaries can not link you to the person you contact, namely communications with unlinkability. However, most current communication systems including variations of Mix networks fail to maintain unlinkability against global active adversaries (GAA) who can monitor global traffic and easily compromise clients and infrastructures. Therefore, designing an unlinkable communication system against GAA is challenging. By analyzing limitations of current communication systems, we propose two other features to assure unlinkability: covertness and deniability. In this paper, we design HTor, a novel and practical communication system with unlinkability, via a single web server. HTor interpolates the server to cut off the direct connection between two people in one communication and exploits covert channels (CCs) to hide communications between clients and the server. Considering servers might be corrupted, HTor utilizes a group mechanism to protect the receiver for each message. By extensive large-scale evaluations, we show that communications over HTor are robust and difficult to detect. Besides, HTor is easily implemented and, with multiple servers, it can provide enough bandwidth and relatively low latency for chatting. Zhenbo Xu, Wei Yang 0011, Yang Xu 0020, Ajin Meng, Qijian He, Liusheng Huang |
SECON | 1 |
| 2016 | The Floating-Point Extension of Symbolic Execution Engine for Bug DetectionabstractMany existing symbolic execution engines for bug detection often ignore floating-point types and operations. That will result in imprecise reasoning about the feasibility of program paths, which in turn leads to false positives and negatives. Recently, there are quite some progress in satisfiability modulo theories (SMT) solving, and some tools are able to support floating-point arithmetic. Nevertheless, naturally extending a symbolic execution engine and directly replacing the back-end with the new SMT solver will not make a good static analyzer for floating-point programs.In this paper, we extend an existing symbolic execution engine for C program bug finding, so that it can deal with floating-point arithmetic and mathematical functions. For the mathematical functions, we employ an abstract model to keep a balance between overhead and precision. We also introduce a strategy, Lazy-verification, to reduce the number of SMT solver calls. We implemented our approach as a tool called Canalyze-fp. Experiments with self-developed benchmarks and non-trivial open source programs show that the proposed approach can effectively avoid the false positives and negatives, without introducing too much overhead. Xingming Wu, Zhenbo Xu, Tianyong Wu, Jun Yan 0009, Jian Zhang 0001 |
APSEC | 2 |
| 2016 | Light-Weight, Inter-Procedural and Callback-Aware Resource Leak Detection for Android AppsabstractAndroid devices include many embedded resources such as Camera, Media Player and Sensors. These resources require programmers to explicitly request and release them. Missing release operations might cause serious problems such as performance degradation or system crash. This kind of defects is called resource leak. Despite a large body of existing works on testing and analyzing Android apps, there still remain several challenging problems. In this work, we present Relda2, a light-weight and precise static resource leak detection tool. We first systematically collected a resource table, which includes the resources that the Android reference requires developers release manually. Based on this table, we designed a general approach to automatically detect resource leaks. To make a more precise inter-procedural analysis, we construct a Function Call Graph for each Android application, which handles function calls of user-defined methods and the callbacks invoked by the Android framework at the same time. To evaluate Relda2's effectiveness and practical applicability, we downloaded 103 apps from popular app stores and an open source community, and found 67 real resource leaks, which we have confirmed manually. Tianyong Wu, Jierui Liu, Zhenbo Xu, Chaorong Guo, Jun Yan 0009, Jian Zhang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2015 | Melton: a practical and precise memory leak detection tool for C programs
Zhenbo Xu, Zhongxing Xu |
Frontiers Comput. Sci. | 1 |
| 2014 | Canalyze: a static bug-finding tool for C programsabstractSymbolic analysis is a commonly used approach for static bug finding. It usually performs a precise path-by-path symbolic simulation from program inputs. A major challenge is its scalability and precision on interprocedural analysis. The former limits the application to large programs. The latter may lead to many false alarms. Zhenbo Xu, Zhongxing Xu, Jiteng Wang |
ISSTA | 1 |
| 2011 | Memory Leak Detection Based on Memory State Transition GraphabstractMemory leak is a common type of defect that is hard to detect manually. Existing memory leak detection tools suffer from lack of precise interprocedural alias and path conditions. To address this problem, we present a static interprocedural analysis algorithm, which captures memory actions and path conditions precisely, to detect memory leak in C programs. Our algorithm uses path-sensitive symbolic execution to track the memory actions in different program paths guarded by path conditions. A novel analysis model called Memory State Transition Graph (MSTG) is proposed to describe the tracking process and its results. An MSTG is generated from a procedure. Nodes in an MSTG contain states of memory objects which record the function behaviors precisely. Edges in anMSTG are annotated with path conditions collected by symbolic execution. The path conditions are checked for satisfiability to reduce the number of false alarms and the path explosion. In order to do interprocedural analysis, our algorithm generates a summary for each procedure from the MSTG and applies the summary at the procedure's call sites. Our implemented tool has found several memory leak bugs in some open source programs and detected more bugs than other tools in some programs from the SPEC2000 benchmarks. In some cases, our tool produces many false positives, but most of them are caused by the same code patterns which are easy to check. Zhenbo Xu, Zhongxing Xu |
APSEC | 1 |