VLDB 2026 Research / reviewers in the wild / expert
Aoxue Li
dblp:152/6095
· DBLP profile ↗
35ranked-venue papers
11as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Optimal Acceleration Evaluation Model for Automated Driving Based on Importance Sampling MethodabstractIn the field of automated driving, scenario-based safety testing is essential for ensuring vehicle safety and promoting technological development. However, traditional mileage-based real-world testing methods are limited by the massive testing requirements and extremely long cycles, hardly meeting efficient validation demands. To address this, this paper proposes a Hybrid Optimal Acceleration Evaluation (HOAE) model based on virtual testing. First, a hybrid segmentation model based on fitting error is constructed to accurately characterize the distribution features of naturalistic driving data, and an acceleration model is designed with the number of tests as the optimization objective to solve for the optimal number of segments. Second, the Importance Sampling (IS) method is introduced to increase the occurrence probability of risk scenarios, and the optimal IS function is solved based on the above acceleration model. Finally, a simulation test platform is established, and comprehensive comparative experiments are conducted under various acceleration evaluation methods, indicators, and parameter distributions. The results demonstrate that the proposed HOAE method offers better universality and higher testing efficiency, and it can increase the test mileage to 104times of the Monte Carlo (MC) method, thereby advancing safety testing of automated vehicles toward greater efficiency and broader applicability. Haobin Jiang, Aoxue Li, Shidian Ma |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | CustomVideo: Customizing Text-to-Video Generation With Multiple SubjectsabstractCustomized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more challenging and practical scenario. In this work, our aim is to promote multi-subject guided text-to-video customization. We propose CustomVideo, a novel framework that can generate identity-preserving videos with the guidance of multiple subjects. To be specific, firstly, we encourage the co-occurrence of multiple subjects via composing them in a single image. Further, upon a basic text-to-video diffusion model, we design a simple yet effective attention control strategy to disentangle different subjects in the latent space of diffusion model. Moreover, to help the model focus on the specific area of the object, we segment the object from given reference images and provide a corresponding object mask for attention learning. Also, we collect a multi-subject text-to-video generation dataset as a comprehensive benchmark. Extensive qualitative, quantitative, and user study results demonstrate the superiority of our method compared to previous state-of-the-art approaches. Zhao Wang 0006, Aoxue Li, Lingting Zhu, Qi Dou 0001, Zhenguo Li |
IEEE Trans. Multim. | 2 |
| 2025 | Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
Wenbo Li 0002, Zhongdao Wang, Haoze Sun, Bangzhen Liu, Haoyu Chen 0003, Aoxue Li, Lei Zhu 0003 |
ICCV | 8 |
| 2025 | CARTS: Advancing Neural Theorem Proving with Diversified Tactic Calibration and Bias-Resistant Tree SearchabstractRecent advancements in neural theorem proving integrate large language models with tree search algorithms like Monte Carlo Tree Search (MCTS), where the language model suggests tactics and the tree search finds the complete proof path. However, many tactics proposed by the language model converge to semantically or strategically similar, reducing diversity and increasing search costs by expanding redundant proof paths. This issue exacerbates as computation scales and more tactics are explored per state. Furthermore, the trained value function suffers from false negatives, label imbalance, and domain gaps due to biased data construction. To address these challenges, we propose CARTS (diversified tactic CAlibration and bias-Resistant Tree Search), which balances tactic diversity and importance while calibrating model confidence. CARTS also introduce preference modeling and an adjustment term related to the ratio of valid tactics to improve the bias-resistance of the value function. Experimental results demonstrate that CARTS consistently outperforms previous methods achieving a pass@l rate of 49.6\% on the miniF2F-test benchmark. Further analysis confirms that CARTS improves tactic diversity and leads to a more balanced tree search. Xiao-Wen Yang, Aoxue Li, Wen-Da Wei, Zhenguo Li |
ICLR | 4 |
| 2025 | FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal VelocitiesabstractThe rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framework. However, most existing MLLMs rely on autoregressive (AR) architectures, which impose inherent limitations on future development, such as the raster-scan order in image generation and restricted reasoning abilities in causal context modeling. In this work, we challenge the dominance of AR-based approaches by introducing FUDOKI, a unified multimodal model purely based on discrete flow matching, as an alternative to conventional AR paradigms. By leveraging metric-induced probability paths with kinetic optimal velocities, our framework goes beyond the previous masking-based corruption process, enabling iterative refinement with self-correction capability and richer bidirectional context integration during generation. To mitigate the high cost of training from scratch, we initialize FUDOKI from pre-trained AR-based MLLMs and adaptively transition to the discrete flow matching paradigm. Experimental results show that FUDOKI achieves performance comparable to state-of-the-art AR-based MLLMs across both visual understanding and image generation tasks, highlighting its potential as a foundation for next-generation unified multimodal models. Furthermore, we show that applying test-time scaling techniques to FUDOKI yields significant performance gains, further underscoring its promise for future enhancement through reinforcement learning. Yao Lai, Aoxue Li, Ning Kang 0001, Chengyue Wu, Zhenguo Li, Ping Luo 0002 |
NeurIPS | 3 |
| 2025 | ZooKT: Task-adaptive knowledge transfer of Model Zoo for few-shot learning
Baoquan Zhang, Bingqi Shan, Aoxue Li, Chuyao Luo, Yunming Ye, Zhenguo Li |
Pattern Recognit. | 3 |
| 2025 | Generating G2 Continuity Reference Paths for Autonomous Vehicles at RoundaboutsabstractPlanning paths for Frenet-based autonomous vehicles (AVs) at roundabouts is difficult without complete and smooth reference paths. In such situations, the interpolating curve planner is often used to create segmented reference paths from simplified geometric roundabout data. While this method ensures curvature continuity within each curve segment, the continuity at the junctions of these segments is poor. Additionally, the determination of merging and diverging point positions at roundabouts has not been thoroughly explored. This paper introduces a novel approach using 5th-order Bézier curves to plan piecewise reference paths for AVs at roundabouts. The proposed method enhances endpoint curvature continuity of the Bézier curves and improves adaptability to non-standard roundabouts. A well-designed objective function is created to optimize both the geometric continuity parameters of the Bézier curves and the positions of merging and diverging points in the circulatory roadway. This function takes into account key factors, including path length and smoothness. Case studies validate the feasibility of maintaining curvature continuity at the endpoints and the method’s ability to generalize across various scenarios, proving its effectiveness for different roundabout structures. The results also confirm the method’s efficacy in generating paths from original geometric roundabout data. Lastly, the acceptable transverse deviations between real-world trajectories and reference paths demonstrate the rationality and practical applicability of this method. Qingyuan Shen, Haobin Jiang, Aoxue Li, Marco Cecotti, Chenhui Yin, You Gong |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | XRNeRF: View-guided Neural Radiance Fields for Occlusion RemovalabstractThe Neural Radiance Fields (NeRF) technique has gained significant popularity as a means of generating new views by representing 3D scenes using an implicit space. This work builds a controlled viewpoint model of a scene by using a multi-view stereo geometry approach, which advances the development of multi-camera group collaboration techniques. Despite NeRF being a spatial query function based on spatial coordinates and observation directions, it cannot efficiently acquire points in the region behind the occlusion when there is occlusion interference in the dataset. For obtaining images that are free from occlusion, we propose the utilization of a neural point cloud to impose constraints on the scene. Our method involves filtering out occlusion by fitting the light distribution and restoring the obscured region by leveraging the isotropic characteristics of the target points obtained from the multiview point feature. The experimental results on existing datasets validate the effectiveness of our method. Aoxue Li, Xinhua Zeng, Yunlong Du, Chengxin Pang |
CSCWD | 1 |
| 2024 | GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingabstractDespite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the generated images unreliable. Meanwhile, a single model tends to specialize in particular tasks and possess the corresponding capabilities, making it inadequate for fulfilling all user requirements. We propose GenArtist, a unified image generation and editing system, coordinated by a multimodal large language model (MLLM) agent. We integrate a comprehensive range of existing models into the tool library and utilize the agent for tool selection and execution. For a complex problem, the MLLM agent decomposes it into simpler sub-problems and constructs a tree structure to systematically plan the procedure of generation, editing, and self-correction with step-by-step verification. By automatically generating missing position-related inputs and incorporating position information, the appropriate tool can be effectively employed to address each sub-problem. Experiments demonstrate that GenArtist can perform various generation and editing tasks, achieving state-of-the-art performance and surpassing existing models such as SDXL and DALL-E 3, as can be seen in Fig. 1. We will open-source the code for future research and applications. Aoxue Li, Zhenguo Li, Xihui Liu |
NeurIPS | 2 |
| 2024 | V-PETL Bench: A Unified Visual Parameter-Efficient Transfer Learning BenchmarkabstractParameter-efficient transfer learning (PETL) methods show promise in adapting a pre-trained model to various downstream tasks while training only a few parameters. In the computer vision (CV) domain, numerous PETL algorithms have been proposed, but their direct employment or comparison remains inconvenient. To address this challenge, we construct a Unified Visual PETL Benchmark (V-PETL Bench) for the CV domain by selecting 30 diverse, challenging, and comprehensive datasets from image recognition, video action recognition, and dense prediction tasks. On these datasets, we systematically evaluate 25 dominant PETL algorithms and open-source a modular and extensible codebase for fair evaluation of these algorithms. V-PETL Bench runs on NVIDIA A800 GPUs and requires approximately 310 GPU days. We release all the benchmark, making it more efficient and friendly to researchers. Additionally, V-PETL Bench will be continuously updated for new PETL algorithms and CV tasks. Yi Xin 0003, Xuyang Liu 0002, Yuntao Du 0001, Haodi Zhou, Christina E. Lee, Junlong Du, Haozhe Wang 0002, Mingcai Chen, Ting Liu 0018, Guimin Hu, Zhongwei Wan, Rongchao Zhang, Aoxue Li, Mingyang Yi, Xiaohong Liu 0001 |
NeurIPS | 15 |
| 2024 | Towards Understanding the Working Mechanism of Text-to-Image Diffusion ModelabstractRecently, the strong latent Diffusion Probabilistic Model (DPM) has been applied to high-quality Text-to-Image (T2I) generation (e.g., Stable Diffusion), by injecting the encoded target text prompt into the gradually denoised diffusion image generator. Despite the success of DPM in practice, the mechanism behind it remains to be explored. To fill this blank, we begin by examining the intermediate statuses during the gradual denoising generation process in DPM. The empirical observations indicate, the shape of image is reconstructed after the first few denoising steps, and then the image is filled with details (e.g., texture). The phenomenon is because the low-frequency signal (shape relevant) of the noisy image is not corrupted until the final stage in the forward process (initial stage of generation) of adding noise in DPM. Inspired by the observations, we proceed to explore the influence of each token in the text prompt during the two stages. After a series of experiments of T2I generations conditioned on a set of text prompts. We conclude that in the earlier generation stage, the image is mostly decided by the special token [\texttt{EOS}] in the text prompt, and the information in the text prompt is already conveyed in this stage. After that, the diffusion model completes the details of generated images by information from themselves. Finally, we propose to apply this observation to accelerate the process of T2I generation by properly removing text guidance, which finally accelerates the sampling up to 25\%+. Mingyang Yi, Aoxue Li, Zhenguo Li |
NeurIPS | 2 |
| 2024 | Efficient Transferability Assessment for Selection of Pre-trained DetectorsabstractLarge-scale pre-training followed by downstream finetuning is an effective solution for transferring deeplearning-based models. Since finetuning all possible pretrained models is computational costly, we aim to predict the transferability performance of these pre-trained models in a computational efficient manner. Different from previous work that seek out suitable models for downstream classification and segmentation tasks, this paper studies the efficient transferability assessment of pre-trained object detectors. To this end, we build up a detector transferability benchmark which contains a large and diverse zoo of pre-trained detectors with various architectures, source datasets and training schemes. Given this zoo, we adopt 7 target datasets from 5 diverse domains as the downstream target tasks for evaluation. Further, we propose to assess classification and regression sub-tasks simultaneously in a unified framework. Additionally, we design a complementary metric for evaluating tasks with varying objects. Experimental results demonstrate that our method outperforms other state-of-the-art approaches in assessing transferability under different target domains while efficiently reducing wall-clock time 32× and requires a mere 5.2% memory footprint compared to brute-force fine-tuning of all pretrained detectors. Our assessment code and benchmark will be publicly available. Zhao Wang 0006, Aoxue Li, Zhenguo Li, Qi Dou 0001 |
WACV | 2 |
| 2024 | A Two-Dimensional Lane-Changing Dynamics Model Based on ForceabstractThe lane-changing behavior exerts a profound influence on the dynamic attributes of traffic flow and road safety. Accurate analysis of lane-changing behavior not only facilitates the comprehension of traffic phenomena and the prevention of traffic accidents but also contributes to constructing a dynamic and realistic background traffic flow for autonomous driving tests. In this paper, we propose a two-dimensional lane-changing dynamics model based on force by abstracting the motion of the vehicle as the variation in acceleration under the influence of forces. Specifically, based on the similarity analysis of factors affecting acceleration and constitutive relationships, the three types of forces to which the vehicle is subjected are represented by a combination of basic visco-elastic elements, while the driver’s attention to the current and target lanes decreases and grows during the lane changing process is analyzed. This model not only comprehensively delineates the lane-changing process but also represents the influences of driver characteristics, neighboring vehicles, and road conditions. In order to verify the performance of the proposed model, we identified the parameters using the least squares method (LSM) based on the highD naturalistic driving dataset. Comparative results illustrate that the lane-changing dynamics model effectively capture complex lane-changing behaviors and their interactions between ego vehicle and surrounding vehicles through a simple and unified way. Aoxue Li, Haobin Jiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
Zhao Wang 0006, Aoxue Li, Fengwei Zhou, Zhenguo Li, Qi Dou 0001 |
BMVC | 2 |
| 2023 | ContraNeRF: Generalizable Neural Radiance Fields for Synthetic-to-real Novel View Synthesis via Contrastive LearningabstractAlthough many recent works have investigated generalizable NeRF-based novel view synthesis for unseen scenes, they seldom consider the synthetic-to-real generalization, which is desired in many practical applications. In this work, we first investigate the effects of synthetic data in synthetic-to-real novel view synthesis and surprisingly observe that models trained with synthetic data tend to produce sharper but less accurate volume densities. For pixels where the volume densities are correct, fine-grained details will be obtained. Otherwise, severe artifacts will be produced. To maintain the advantages of using synthetic data while avoiding its negative effects, we propose to introduce geometry-aware contrastive learning to learn multi-view consistent features with geometric constraints. Meanwhile, we adopt cross-view attention to further enhance the geometry perception of features by querying features across input views. Experiments demonstrate that under the synthetic-to-real setting, our method can render images with higher quality and better fine-grained details, outperforming existing generalizable novel view synthesis methods in terms of PSNR, SSIM, and LPIPS. When trained on real data, our method also achieves state-of-the-art results. https://haoy945.github.io/contranerf/ Hao Yang 0044, Lanqing Hong, Aoxue Li, Tianyang Hu 0001, Zhenguo Li, Gim Hee Lee, Liwei Wang 0001 |
CVPR | 3 |
| 2023 | UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationabstractJointly processing information from multiple sensors is crucial to achieving accurate and robust perception for reliable autonomous driving systems. However, current 3D perception research follows a modality-specific paradigm, leading to additional computation overheads and inefficient collaboration between different sensor data. In this paper, we present an efficient multi-modal backbone for outdoor 3D perception named UniTR, which processes a variety of modalities with unified modeling and shared parameters. Unlike previous works, UniTR introduces a modality-agnostic transformer encoder to handle these view-discrepant sensor data for parallel modal-wise representation learning and automatic cross-modal interaction without additional fusion steps. More importantly, to make full use of these complementary sensor types, we present a novel multi-modal integration strategy by both considering semantic-abundant 2D perspective and geometry-aware 3D sparse neighborhood relations. UniTR is also a fundamentally task-agnostic backbone that naturally supports different 3D perception tasks. It sets a new state-of-the-art performance on the nuScenes benchmark, achieving +1.1 NDS higher for 3D object detection and +12.0 higher mIoU for BEV map segmentation with lower inference latency. Code will be available at https://github.com/Haiyang-W/UniTR. Hao Tang 0005, Shaoshuai Shi, Aoxue Li, Zhenguo Li, Bernt Schiele, Liwei Wang 0001 |
ICCV | 4 |
| 2022 | Semi-Supervised Object Detection via Multi-instance Alignment with Global Class PrototypesabstractSemi-Supervised object detection (SSOD) aims to improve the generalization ability of object detectors with large-scale unlabeled images. Current pseudo-labeling-based SSOD methods individually learn from labeled data and unlabeled data, without considering the relation be-tween them. To make full use of labeled data, we pro-pose a Multi-instance Alignment model which enhances the prediction consistency based on Global Class Proto-types (MA-GCP). Specifically, we impose the consistency between pseudo ground-truths and their high-IoU candi-dates by minimizing the cross-entropy loss of their class distributions computed based on global class prototypes. These global class prototypes are estimated with the whole labeled dataset via the exponential moving average algorithm. To evaluate the proposed MA-GCP model, we inte-grate it into the state-of-the-art SSOD framework and ex-periments on two benchmark datasets demonstrate the ef-fectiveness of our MA-GCP approach. Aoxue Li, Zhenguo Li |
CVPR | 1 |
| 2022 | CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point CloudsabstractWe present a novel two-stage fully sparse convolutional 3D object detection framework, named CAGroup3D. Our proposed method first generates some high-quality 3D proposals by leveraging the class-aware local group strategy on the object surface voxels with the same semantic predictions, which considers semantic consistency and diverse locality abandoned in previous bottom-up approaches. Then, to recover the features of missed voxels due to incorrect voxel-wise segmentation, we build a fully sparse convolutional RoI pooling module to directly aggregate fine-grained spatial information from backbone for further proposal refinement. It is memory-and-computation efficient and can better encode the geometry-specific features of each 3D proposal. Our model achieves state-of-the-art 3D detection performance with remarkable gains of +3.6% on ScanNet V2 and +2.6% on SUN RGB-D in term of [email protected]. Code will be available at https://github.com/Haiyang-W/CAGroup3D. Lihe Ding, Shaocong Dong, Shaoshuai Shi, Aoxue Li, Jianan Li 0001, Zhenguo Li, Liwei Wang 0001 |
NeurIPS | 5 |
| 2022 | STC-IDS: Spatial-temporal correlation feature analyzing based intrusion detection system for intelligent connected vehiclesabstractIntrusion detection is an important defensive measure for automotive communications security. Accurate frame detection models assist vehicles to avoid malicious attacks. Uncertainty and diversity regarding attack methods make this task challenging. However, the existing works have the limitation of only considering local features or the weak feature mapping of multifeatures. To address these limitations, we present a novel model for automotive intrusion detection by spatial–temporal correlation (STC) features of in-vehicle communication traffic (intrusion detection system [IDS]). Specifically, the proposed model exploits an encoding-detection architecture. In the encoder part, spatial and temporal relations are encoded simultaneously. To strengthen the relationship between features, the attention-based convolutional network still captures spatial and channel features to increase the receptive field, while attention-long short-term memory builds meaningful relationships from previous time series or crucial bytes. The encoded information is then passed to detector for generating forceful spatial–temporal attention features and enabling anomaly classification. In particular, single-frame and multiframe models are constructed to present different advantages, respectively. Under automatic hyperparameter selection based on Bayesian optimization, the model is trained to attain the best performance. Extensive empirical studies based on a real-world vehicle attack data set demonstrate that STC-IDS has outperformed baseline methods and obtains fewer false-alarm rates while maintaining efficiency. Pengzhou Cheng, Mu Han, Aoxue Li, Fengwei Zhang |
Int. J. Intell. Syst. | 3 |
| 2022 | Federated learning-based trajectory prediction model with privacy preserving for intelligent vehicleabstractThe existing trajectory prediction is mainly for specific road sections, which is poor to adapt to complex and changing traffic scenarios. Meanwhile, decentralized trajectory data is hard to be fully utilized in the data silo environment. To solve data silos in the intelligent vehicle industry, introducing federated learning methods in vehicular edge computing has attracted extensive attention. But traditional federated learning still has the potential to suffer from the mining of training data in the case of model privacy leakage. In this paper, a vehicle trajectory prediction method based on federated learning and homomorphic encryption has been presented, which adopts a three-layer architecture with a vehicle cluster, edge computing server, and cloud core network. Compared with traditional centralized deep learning methods, this approach enables joint modeling of multiple parties to improve the model's generalization performance. We used a proxy re-encryption algorithm to implement key distribution and also designed an encrypted federated network algorithm FAHEFL, which uses FAHE1 homomorphic encryption to protect the privacy of the model in parameter transmission. Each local model includes a driving behavior recognition module and trajectory output module. The driving behavior recognition module uses a 1D convolutional neural network to recognize the driving behavior, then input recognition results and historical trajectories to the trajectory output module, which uses LSTM neural network to output predicted trajectories. The experiment results show that the model built with FAHEFL has no more 1% error in the driving behavior recognition module than concentrated learning, while the minimum mean square error of the trajectory prediction module increased by only 2.5%. This article also discusses the performance between FAHEFL and well-known cryptographic federation network algorithms. Mu Han, Shidian Ma, Aoxue Li, Haobin Jiang |
Int. J. Intell. Syst. | 4 |
| 2021 | Dense Relation Distillation With Context-Aware Aggregation for Few-Shot Object DetectionabstractConventional deep learning based methods for object detection require a large amount of bounding box annotations for training, which is expensive to obtain such high quality annotated data. Few-shot object detection, which learns to adapt to novel classes with only a few annotated examples, is very challenging since the fine-grained feature of novel object can be easily overlooked with only a few data available. In this work, aiming to fully exploit features of annotated novel object and capture fine-grained features of query object, we propose Dense Relation Distillation with Context-aware Aggregation (DCNet) to tackle the few-shot detection problem. Built on the meta-learning based framework, Dense Relation Distillation module targets at fully exploiting support features, where support features and query feature are densely matched, covering all spatial locations in a feed-forward fashion. The abundant usage of the guidance information endows model the capability to handle common challenges such as appearance changes and occlusions. Moreover, to better capture scale-aware features, Context-aware Aggregation module adaptively harnesses features from different scales for a more comprehensive feature representation. Extensive experiments illustrate that our proposed approach achieves state-of-the-art results on PASCAL VOC and MS COCO datasets. Code will be made available at https://github.com/hzhupku/DCNet. Hanzhe Hu, Shuai Bai, Aoxue Li, Jinshi Cui, Liwei Wang 0001 |
CVPR | 3 |
| 2021 | Transformation Invariant Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to learn detectors that can be generalized to novel classes with only a few instances. Unlike previous attempts that exploit meta-learning techniques to facilitate FSOD, this work tackles the problem from the perspective of sample expansion. To this end, we propose a simple yet effective Transformation Invariant Principle (TIP) that can be flexibly applied to various meta-learning models for boosting the detection performance on novel class objects. Specifically, by introducing consistency regularization on predictions from various transformed images, we augment vanilla FSOD models with the generalization ability to objects perturbed by various transformation, such as occlusion and noise. Importantly, our approach can extend supervised FSOD models to naturally cope with unlabeled data, thus addressing a more practical and challenging semi-supervised FSOD problem. Extensive experiments on PASCAL VOC and MSCOCO datasets demonstrate the effectiveness of our TIP under both of the two FSOD settings. Aoxue Li, Zhenguo Li |
CVPR | 1 |
| 2021 | Zero and Few Shot Learning With Semantic Feature Synthesis and Competitive LearningabstractZero-shot learning (ZSL) is made possible by learning a projection function between a feature space and a semantic space (e.g., an attribute space). Key to ZSL is thus to learn a projection that is robust against the often large domain gap between the seen and unseen class domains. In this work, this is achieved by unseen class data synthesis and robust projection function learning. Specifically, a novel semantic data synthesis strategy is proposed, by which semantic class prototypes (e.g., attribute vectors) are used to simply perturb seen class data for generating unseen class ones. As in any data synthesis/hallucination approach, there are ambiguities and uncertainties on how well the synthesised data can capture the targeted unseen class data distribution. To cope with this, the second contribution of this work is a novel projection learning model termed competitive bidirectional projection learning (BPL) designed to best utilise the ambiguous synthesised data. Specifically, we assume that each synthesised data point can belong to any unseen class; and the most likely two class candidates are exploited to learn a robust projection function in a competitive fashion. As a third contribution, we show that the proposed ZSL model can be easily extended to few-shot learning (FSL) by again exploiting semantic (class prototype guided) feature synthesis and competitive BPL. Extensive experiments show that our model achieves the state-of-the-art results on both problems. Jiechao Guan, Zhiwu Lu 0001, Tao Xiang 0002, Aoxue Li, Ji-Rong Wen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Boosting Few-Shot Learning With Adaptive Margin LossabstractFew-shot learning (FSL) has attracted increasing attention in recent years but remains challenging, due to the intrinsic difficulty in learning to generalize from a few examples. This paper proposes an adaptive margin principle to improve the generalization ability of metric-based meta-learning approaches for few-shot learning problems. Specifically, we first develop a class-relevant additive margin loss, where semantic similarity between each pair of classes is considered to separate samples in the feature embedding space from similar classes. Further, we incorporate the semantic context among all classes in a sampled training task and develop a task-relevant additive margin loss to better distinguish samples from different classes. Our adaptive margin method can be easily extended to a more realistic generalized FSL setting. Extensive experiments demonstrate that the proposed method can boost the performance of current metric-based meta-learning approaches, under both the standard FSL and generalized FSL settings. Aoxue Li, Weiran Huang 0001, Xu Lan, Jiashi Feng, Zhenguo Li, Liwei Wang 0001 |
CVPR | 1 |
| 2020 | Transferrable Feature and Projection Learning with Class Hierarchy for Zero-Shot Learning
Aoxue Li, Zhiwu Lu 0001, Jiechao Guan, Tao Xiang 0002, Liwei Wang 0001, Ji-Rong Wen |
Int. J. Comput. Vis. | 1 |
| 2020 | Human-Like Trajectory Planning on Curved Road: Learning From Human DriversabstractThe ultimate goal of self-driving technologies is to offer a safe and human-like driving experience. As one of the most important enabling functionalities, trajectory planning has been extensively studied from the perspective of safety. However, human-like trajectory planning on curved roads has rarely been studied. In this paper, we characterize and model human driving using extensive experimental driving collected on an urban curved road with 30 participants (10 experienced and 20 novice drivers) and five vehicles of different types. Differential global positioning system (GPS) is used to measure vehicle positions in high precision. We study factors that affect the driving trajectory, including vehicle speed, road curvature, and sight distance. We find that the human drivers typically do not follow lane centerline and the human-driven trajectories are very different from planners like rapidly exploring random tree (RRT). To generate human-like driving trajectory, we develop a data-driven trajectory model using general regression neural network (GRNN). The model was validated in various cases with promising performance. Aoxue Li, Haobin Jiang, Zhaojian Li 0001, Jie Zhou 0019, Xinchen Zhou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | A New Microscopic Traffic Model Using a Spring-Mass-Damper-Clutch SystemabstractMicroscopic traffic models describe how cars interact with their neighbors in an uninterrupted traffic flow and are frequently used for reference in advanced vehicle control design. In this paper, we propose a novel mechanical system-inspired microscopic traffic model using a mass-spring-damper-clutch system. This model naturally captures the ego vehicle's resistance to large relative speed and deviation from a (driver- and speed-dependent) desired relative distance when following the lead vehicle. Compared with the existing car-following (CF) models, this model offers physically interpretable insights into the underlying CF dynamics and is able to characterize the impact of the ego vehicle on the lead vehicle, which is neglected in the existing CF models. Thanks to the nonlinear wave propagation analysis techniques for mechanical systems, the proposed model, therefore, has great scalability so that multiple mass-spring-damper-clutch systems can be chained to study the macroscopic traffic flow. We investigate the stability of the proposed model on the system parameters and the time delay using the spectral element method. We also develop a parallel recursive least square with inverse QR decomposition (PRLS-IQR) algorithm to identify the model parameters online. These real-time estimated parameters can be used to predict the driving trajectory that can be incorporated into advanced vehicle longitudinal control systems for improved safety and fuel efficiency. The PRLS-IQR is computationally efficient and numerically stable, and therefore, it is suitable for online implementation. The traffic model and the parameter identification algorithm are validated on both the simulations and naturalistic driving data from multiple drivers. Promising performance is demonstrated. Zhaojian Li 0001, Firas A. Khasawneh, Xiang Yin 0003, Aoxue Li, Ziyou Song |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | Large-Scale Few-Shot Learning: Knowledge Transfer With Class HierarchyabstractRecently, large-scale few-shot learning (FSL) becomes topical. It is discovered that, for a large-scale FSL problem with 1,000 classes in the source domain, a strong baseline emerges, that is, simply training a deep feature embedding model using the aggregated source classes and performing nearest neighbor (NN) search using the learned features on the target classes. The state-of-the-art large-scale FSL methods struggle to beat this baseline, indicating intrinsic limitations on scalability. To overcome the challenge, we propose a novel large-scale FSL model by learning transferable visual features with the class hierarchy which encodes the semantic relations between source and target classes. Extensive experiments show that the proposed model significantly outperforms not only the NN baseline but also the state-of-the-art alternatives. Furthermore, we show that the proposed model can be easily extended to the large-scale zero-shot learning (ZSL) problem and also achieves the state-of-the-art results. Aoxue Li, Tiange Luo, Zhiwu Lu 0001, Tao Xiang 0002, Liwei Wang 0001 |
CVPR | 1 |
| 2019 | Few-Shot Learning With Global Class RepresentationsabstractIn this paper, we propose to tackle the challenging few-shot learning (FSL) problem by learning global class representations using both base and novel class training samples. In each training episode, an episodic class mean computed from a support set is registered with the global representation via a registration module. This produces a registered global class representation for computing the classification loss using a query set. Though following a similar episodic training pipeline as existing meta learning based approaches, our method differs significantly in that novel class training samples are involved in the training from the beginning. To compensate for the lack of novel class training samples, an effective sample synthesis strategy is developed to avoid overfitting. Importantly, by joint base-novel class training, our approach can be easily extended to a more practical yet challenging FSL setting, i.e., generalized FSL, where the label space of test data is extended to both base and novel classes. Extensive experiments show that our approach is effective for both of the two FSL settings. Aoxue Li, Tiange Luo, Tao Xiang 0002, Weiran Huang 0001, Liwei Wang 0001 |
ICCV | 1 |
| 2018 | Large-Scale Sparse Learning From Noisy Tags for Semantic SegmentationabstractIn this paper, we present a large-scale sparse learning (LSSL) approach to solve the challenging task of semantic segmentation of images with noisy tags. Different from the traditional strongly supervised methods that exploit pixel-level labels for semantic segmentation, we make use of much weaker supervision (i.e., noisy tags of images) and then formulate the task of semantic segmentation as a weakly supervised learning (WSL) problem from the view point of noise reduction of superpixel labels. By learning the data manifolds, we transform the WSL problem into an LSSL problem. Based on nonlinear approximation and dimension reduction techniques, a linear-time-complexity algorithm is developed to solve the LSSL problem efficiently. We further extend the LSSL approach to visual feature refinement for semantic segmentation. The experiments demonstrate that the proposed LSSL approach can achieve promising results in semantic segmentation of images with noisy tags. Aoxue Li, Zhiwu Lu 0001, Liwei Wang 0001, Peng Han 0005, Ji-Rong Wen |
IEEE Trans. Cybern. | 1 |
| 2017 | Accurate Pulmonary Nodule Detection in Computed Tomography Images Using Deep Convolutional Neural Networks
Jia Ding, Aoxue Li, Liwei Wang 0001 |
MICCAI (3) | 2 |
| 2017 | Zero-Shot Scene Classification for High Spatial Resolution Remote Sensing ImagesabstractDue to the rapid technological development of various sensors, a huge volume of high spatial resolution (HSR) image data can now be acquired. How to efficiently recognize the scenes from such HSR image data has become a critical task. Conventional approaches to remote sensing scene classification only utilize information from HSR images. Therefore, they always need a large amount of labeled data and cannot recognize the images from an unseen scene class without any visual sample in the labeled data. To overcome this drawback, we propose a novel approach for recognizing images from unseen scene classes, i.e., zero-shot scene classification (ZSSC). In this approach, we first use the well-known natural language process model, word2vec, to map names of seen/unseen scene classes to semantic vectors. A semantic-directed graph is then constructed over the semantic vectors for describing the relationships between unseen classes and seen classes. To transfer knowledge from the images in seen classes to those in unseen classes, we make an initial label prediction on test images by an unsupervised domain adaptation model. With the semantic-directed graph and initial prediction, a label-propagation algorithm is then developed for ZSSC. By leveraging the visual similarity among images from the same scene class, a label refinement approach based on sparse learning is used to suppress the noise in the zero-shot classification results. Experimental results show that the proposed approach significantly outperforms the state-of-the-art approaches in ZSSC. Aoxue Li, Zhiwu Lu 0001, Liwei Wang 0001, Tao Xiang 0002, Ji-Rong Wen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Remote sensing image segmentation based on Wilcoxon rank sum test and mean absolute deviationabstractIn this paper, a novel threshold segmentation method for remote sensing images is proposed. The proposed method is based on Wilcoxon rank sum test and mean absolute deviation (MAD) model with color feature and can segment roads and residential areas from vegetation more accurately. Three steps are used to realize the new method. First, we use blue and green color components as paired sample on Wilcoxon rank sum test to partition the vegetation. Second, a road-residential area map is constructed by mean absolute deviation on an improved two dimensional histogram to get the optimal threshold for segmentation. Finally, we fuse vegetation and residential map to get the final segmentation result. Compared with several existing algorithms, the proposed method presents a more accurate segmentation. Libao Zhang, Qiaoyue Sun, Aoxue Li |
IGARSS | 4 |
| 2016 | Global and Local Saliency Analysis for the Extraction of Residential Areas in High-Spatial-Resolution Remote Sensing ImageabstractExtraction of residential areas plays an important role in remote sensing image processing. Extracted results can be applied to various scenarios, including disaster assessment, urban expansion, and environmental change research. Quality residential areas extracted from a remote sensing image must meet three requirements: well-defined boundaries, uniformly highlighted residential area, and no background redundancy in residential areas. Driven by these requirements, this study proposes a global and local saliency analysis model (GLSA) for the extraction of residential areas in high-spatial-resolution remote sensing images. In the proposed model, a global saliency map based on quaternion Fourier transform (QFT) and a global saliency map based on adaptive directional enhancement lifting wavelet transform (ADE-LWT) are generated along with a local saliency map, all of which are fused into a main saliency map based on complementarities. In order to analyze the correlation among spectrums in the remote sensing image, the phase spectrum information of QFT is used on the multispectral images for producing a global saliency map. To acquire the texture and edge features of different scales and orientations, the coefficients acquired by ADE-LWT are used to construct another global saliency map. To discard redundant backgrounds, the amplitude spectrum of the Fourier transform and the spatial relations among patches are introduced into the panchromatic image to generate the local saliency map. Experimental results indicate that the GLSA model can better define the boundaries of residential areas and achieve complete residential areas than current methods. Furthermore, the GLSA model can prevent redundant backgrounds in residential areas and thus acquire more accurate residential areas. Libao Zhang, Aoxue Li, Kaina Yang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Remote Sensing Image Segmentation Based on an Improved 2-D Gradient Histogram and MMAD ModelabstractA novel remote sensing image segmentation algorithm based on an improved 2-D gradient histogram and minimum mean absolute deviation (MMAD) model is proposed in this letter. We extract the global features as a 1-D histogram from an improved 2-D gradient histogram by diagonal projection and subsequently use the MMAD model on the 1-D histogram to implement the optimal threshold. Experiments on remote sensing images indicate that the new algorithm provides accurate segmentation results, particularly for images characterized by Laplace distribution histograms. Furthermore, the new algorithm has low time consumption. Libao Zhang, Aoxue Li, Shuaijing Xu, Xuye Yang |
IEEE Geosci. Remote. Sens. Lett. | 2 |