EDBT 2026 Demo / reviewers in the wild / expert
Hui Zhang 0091
dblp:181/2846-91
· DBLP profile ↗
20ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0001-7267-5179ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Discriminative to Generative: A Diffusion-Based Paradigm for Multi-Agent Collaborative PerceptionabstractCollaborative perception leveraging intermediate feature fusion has emerged as a leading paradigm to significantly enhance the environmental perception capabilities of autonomous driving systems. However, existing methods typically rely on discriminative supervision guided by downstream tasks. This paradigm compels models to learn minimal, task-specific representations, which conflicts with the goal of cooperative perception to capture comprehensive information, thereby limiting generalization. To address this issue, we propose DiGS-CP, a novel two-stage generative supervised collaborative perception framework. Specifically, we introduce a diffusion-based generative task that conditions on fused object-level features to generate representations of object-level point clouds. The proposed generative supervision provides fine-grained, task-agnostic signals that encourages the fusion module to learn comprehensive representations beyond task-specific requirements. By preserving and integrating complementary information from collaborative agents, our approach overcomes the limitations of task-specific learning and enhances the generalizability of the learned features. Furthermore, our two-stage architecture requires agents to transmit only object-level features, significantly reducing communication overhead. Extensive experiments on three benchmark datasets demonstrate that DiGS-CP achieves state-of-the-art performance in 3D object detection, while maintaining low bandwidth requirements and exhibiting excellent generalization ability. Kexin Gong, Puyi Yao, Guiyang Luo, Quan Yuan 0004, Tiange Fu, Hui Zhang 0091 |
AAAI | 6 |
| 2026 | MambaOcc: Visual state space models for BEV-based occupancy prediction with local adaptive reordering
Yonglin Tian, Songlin Bai, Zhiyao Luo, Yutong Wang 0001, Hui Zhang 0091, Baoqing Guo, Fei-Yue Wang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | CoDS: Enhancing Collaborative Perception in Heterogeneous Scenarios via Domain Separation
Yushan Han, Hui Zhang 0091, Honglei Zhang 0002, Chuntao Ding, Yuanzhouhan Cao, Yidong Li |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | CoDTS: Enhancing Sparsely Supervised Collaborative Perception with a Dual Teacher-Student FrameworkabstractCurrent collaborative perception methods often rely on fully annotated datasets, which can be expensive to obtain in practical situations. To reduce annotation costs, some works adopt sparsely supervised learning techniques and generate pseudo labels for the missing instances. However, these methods fail to achieve an optimal confidence threshold that harmonizes the quality and quantity of pseudo labels. To address this issue, we propose an end-to-end Collaborative perception Dual Teacher-Student framework (CoDTS), which employs adaptive complementary learning to produce both high-quality and high-quantity pseudo labels. Specifically, the Main Foreground Mining (MFM) module generates high-quality pseudo labels based on the prediction of the static teacher. Subsequently, the Supplement Foreground Mining (SFM) module ensures a balance between the quality and quantity of pseudo labels by adaptively identifying missing instances based on the prediction of the dynamic teacher. Additionally, the Neighbor Anchor Sampling (NAS) module is incorporated to enhance the representation of pseudo labels. To promote the adaptive complementary learning, we implement a staged training strategy that trains the student and dynamic teacher in a mutually beneficial manner. Extensive experiments demonstrate that the CoDTS effectively ensures an optimal balance of pseudo labels in both quality and quantity, establishing a new state-of-the-art in sparsely supervised collaborative perception. Yushan Han, Hui Zhang 0091, Honglei Zhang 0002, Yidong Li |
AAAI | 2 |
| 2025 | Selective Shift: Towards Personalized Domain Adaptation in Multi-Agent Collaborative PerceptionabstractGiven the scarcity of real data and the time-intensive nature of labeling, current multi-agent perception models often rely on simulated sensor data for training and validation. However, perception performance deteriorates significantly due to domain gap between simulated and real data. Existing adaptation methods focus on domain-generalized feature extraction while neglecting multi-agent shift uncertainty and relational semantic loss. To address this issue, we propose a Selective Shift Domain Adaptation method in multi-agent collaborative perception, called SSDA. SSDA incorporates two essential components: the frequency-decoupled feature shift adjustment (FSA) and the entropy-driven staged adaptive alignment (SAA). To mitigate the relational semantic loss, the FSA is proposed to simplify the representation of correlation features and remove redundant information from the source domain, thereby mitigating interference for domain adversarial scenarios. To tackle the shift uncertainty, the SAA is designed to achieve adaptive alignment from global to local guided by information entropy, which dynamically adjusts weights for samples according to their level of uncertainty. The results demonstrate that the SSDA is significantly superior to the SOTA, achieving up to 7.35% improvements on [email protected]. Hui Zhang 0091, Yiteng Xu, Yonglin Tian, Yidong Li, Tiago H. Falk, Fei-Yue Wang 0001 |
ACM Multimedia | 1 |
| 2025 | 24-h Lane Line Detection via Parallel Scene Information CollaborationabstractLane detection is a critical technology for autonomous driving, but current deep learning-based methods face significant challenges due to the lack of diverse datasets, especially for nighttime conditions. Most datasets are predominantly composed of daytime images, making it difficult to develop models that perform reliably around the clock. Inspired by the parallel system theory, we explore a novel approach to generate comprehensive 24-hour datasets from daytime images alone. In this paper, we propose the Parallel Scene Information Collaboration (PSIC) framework, designed to enhance 24-hour lane detection using only daytime data. The PSIC framework consists of three key components: artificial scene generation, information collaboration, and lane line detection. First, we address the limitations of existing datasets by proposing two generators—one that transforms daytime images into realistic nighttime scenes, and another that refines nighttime images by adding daytime characteristics. Next, to mitigate noise in the generated scenes, we propose a Multi-Spatial Feature Fusion (MSFF) module that effectively integrates features from both real and artificial scenes through spatial collaboration. Finally, the combined information is used by an anchor-based detection head to accurately identify lane positions. Our experiments on the TuSimple, Night TuSimple, and CULane datasets demonstrate that our method achieves state-of-the-art performance in 24-hour lane line detection, significantly improving reliability and robustness across varying conditions. Shaohua Duan, Chunjie Zhang 0001, Xiaolong Zheng 0001, Yutong Wang 0001, Hui Zhang 0091, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SASAN: Shape-Adaptive Set Abstraction Network for Point-Voxel 3D Object DetectionabstractPoint-voxel 3D object detectors have achieved impressive performance in complex traffic scenes. However, they utilize the 3D sparse convolution (spconv) layers with fixed receptive fields, such as voxel-based detectors, and inherit the fixed sphere radius from point-based methods for generating the features of keypoints, which make them weak in adaptively modeling various geometrical deformations and sizes of real objects. To tackle this issue, we propose a shape-adaptive set abstraction network (SASAN) for point-voxel 3D object detection. First, the proposal and offset generation module is adopted to learn the coordinates and confidences of 3D proposals and shape-adaptive offsets of the certain number of offset points for each voxel. Meanwhile, an extra offset supervision task is employed to guide the learning of shifting values of offset points, aiming at motivating the predicted offsets to preferably adapt to the various shapes of objects. Then, the shape-adaptive set abstraction module is proposed to extract multiscale keypoints features by grouping the neighboring offset points' features, as well as features learned from adjacent raw points and the 2-D bird-view map. Finally, the region of interest (RoI)-grid proposal refinement module is used to aggregate the keypoints features for further proposal refinement and confidence prediction. Extensive experiments on the competitive KITTI 3D detection benchmark demonstrate that the proposed SASAN gains superior performance as compared with state-of-the-art methods. Hui Zhang 0091, Guiyang Luo, Xiao Wang 0002, Yidong Li, Weiping Ding 0001, Fei-Yue Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | EdgeCooper: Network-Aware Cooperative LiDAR Perception for Enhanced Vehicular AwarenessabstractAutonomous driving vehicle (ADV) that is ready to transform our society and economy, is in desperate need of precise positioning over itself as well as surrounding environments. However, it is still a challenging issue for ADV to retrieve real-time positioning knowledge over road participants and dynamic surrounding environments, due to unsatisfied perception accuracy caused by sparse observations and limited perception range. Cooperative perception, which advocates cooperatively disseminating perception data among vehicles, has the potential to overcome the above limitations. To this end, this article proposes a novel edge-assisted multi-vehicle perception system to enhance vehicles’ awareness over surrounding environments, which is termed as EdgeCooper. EdgeCooper first schedules vehicles to share complementarity-enhanced and redundancy-minimized raw sensor data with an edge server, using multi-hop cooperative 5G V2X communications. Then, EdgeCooper merges vehicles’ individual views to form a holistic view with a higher resolution, thus enhancing perception robustness and enlarging perception range. We formulate multi-vehicle multi-hop cooperative data sharing as a minimum cost flow problem with conflict, and further prove that there exists no polynomial-time approximation algorithm with a constant performance ratio unless P = NP. Furthermore, a two-dimension graph coloring algorithm with guaranteed performance is proposed to eliminate conflict. We evaluate EdgeCooper by building a comprehensive simulation platform through a joint manipulation of SUMO, CARLA, NS3, and PyTorch. The experiment results show that, compared to a single vehicle’s perception, EdgeCooper performs effective and efficient in enhancing vehicular awareness, e.g., extending up to 3.6 times detection range and improving perception accuracy by 20%. Guiyang Luo, Chongzhang Shao, Nan Cheng 0001, Hui Zhang 0091, Quan Yuan 0004 |
IEEE J. Sel. Areas Commun. | 5 |
| 2024 | Scale-Disentangled and Uncertainty-Guided Alignment for Domain-Adaptive Object DetectionabstractUnsupervised domain adaptive object detection methods aim to transfer knowledge from the label-sufficient domain to the unlabeled domain. Most existing works minimize domain disparity by concentrating on different levels through adversarial learning. However, adversarial learning do not consider the different influences on under-aligned and well-aligned samples as they merely match distinct distributions with consistent weight. To address this issue, we design a novel scale-disentangled and uncertainty-guided alignment for domain-adaptive object detection (SDUGA), consisting of three main components: (1) Disentangled scale coarse module, which decouples scale information from global image features and performs individual alignment across domains for the corresponding scale by training domain classifiers in an adversarial learning manner; (2) Disentangled scale fine module, which generalizes the disentangled scale alignment to instance-level adaptation, reinforcing the distribution alignment across domains from multi-scale local instance level; (3) Uncertainty-guided coarse-to-fine attention alignment, which adjusts weights for various samples adaptively by generating the uncertainty-guided attention map, thus enforcing the detector to converge more on alignment for under-aligned samples and avoid misaligning well-aligned ones. Extensive experiments over three challenging domain-shift object detection scenarios demonstrate that SDUGA gains superior performance compared to state-of-the-art methods. Hui Zhang 0091, Guiyang Luo, Yuanzhouhan Cao, Xiao Wang 0002, Yidong Li, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | One Size Fits All: A Unified Traffic Predictor for Capturing the Essential Spatial-Temporal DependencyabstractTraffic prediction is a keystone for building smart cities in the new era and has found wide applications in traffic scheduling and management, environment policy making, public safety, and so on. Instead of creating a traffic predictor for each city, this article focuses on designing a unified network model that could be directly applied for traffic prediction in any city, by learning the essential spatial-temporal dependencies, i.e., the mutual relationship between traffic and the corresponding fine-grained road network. To achieve this goal, this article proposes a joint knowledge- and data-driven mechanism that novelly divides dependencies into three kinds of correlations, i.e., road segment, intra-intersection, and inter-intersection correlation, which capture the microcosmic, middle, and macroscopic dependencies between traffic and the road network, respectively. Specifically, we first construct traffic datasets that could cover all road segments from real-world trajectory datasets, which makes it possible to model the whole road network as a graph, with the help of fine-grained road topology. Then, we propose meta road segment learner, connection-aware spatial-temporal graph convolutional network (GCN), and multiscale residual networks for capturing the microcosmic, middle, and macroscopic dependencies, respectively. Our experiments on three real-world datasets demonstrate that our proposed method could: 1) achieve better prediction accuracy compared with several approaches and 2) capture the mutual relationship between traffic and the fine-grained road network since our model trained only using data from the source city achieves good performance when it is directly applied for traffic prediction in the target city, without any fine-tuning. The codes will be made publicly available. Guiyang Luo, Hui Zhang 0091, Quan Yuan 0004, Wendong Wang 0003, Fei-Yue Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | GLaLT: Global-Local Attention-Augmented Light Transformer for Scene Text RecognitionabstractRecent years have witnessed the growing popularity of connectionist temporal classification (CTC) and attention mechanism in scene text recognition (STR). CTC-based methods consume less time with few computational burdens, while they are not as effective as attention-based methods. To retain computational efficiency and effectiveness, we propose the global-local attention-augmented light Transformer (GLaLT), which adopts a Transformer-based encoder-decoder structure to orchestrate CTC and attention mechanism. The encoder integrates the self-attention module with the convolution module to augment the attention, where the self-attention module pays more attention to capturing long-term global dependencies and the convolution module focuses on local context modeling. The decoder consists of two parallel modules: one is the Transformer-decoder-based attention module and the other is the CTC module. The first one is removed in the testing phase and can guide the second one to extract robust features in the training phase. Extensive experiments on standard benchmarks demonstrate that GLaLT achieves state-of-the-art performance for both regular and irregular STR. In terms of tradeoffs, the proposed GLaLT is at or near the frontiers for maximizing speed, accuracy, and computational efficiency at the same time. Hui Zhang 0091, Guiyang Luo, Xiao Wang 0002, Fei-Yue Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multimodal Perception and Decision-Making Systems for Complex Roads Based on Foundation ModelsabstractSince the inception of Industry 5.0 in 2021, a growing number of researchers have begun to pay their attention to the revolutionary shift it brings. The principles of Industry 5.0, including human-centric, sustainability, and emphasis on ecological and social values, will become the new paradigm for future industrial development. In this transformative landscape, artificial intelligence (AI) plays a pivotal role, and foundation models based on ChatGPT are set to reshape the organizational structure of industries. In this article, we introduce a multimodal perception and decision-making system built upon a foundational model. This system integrates image and point cloud data to enhance perception accuracy and provide ample information for decision making. It is designed to achieve a deep integration of AI and human-centric autonomous driving within the context of Industry 5.0. We introduce a cross-domain learning approach in the system architecture, along with a model training method from foundation models to handle complex road conditions. The proposed method enables road drivable area segmentation on complex unstructured roads. To address the issue of increased variance caused by the residual structure employed in previous works, this article introduces a distribution correction module, which effectively mitigates this problem. Furthermore, to achieve high-performance perception systems in intricate road scenarios, we put forth a multimodal perception fusion method in this study. The experiments demonstrate the superiority of this approach over single-sensor perception. This work contributes to the ongoing discourse on the convergence of AI, human-centric values, and advanced driving systems within the framework of Industry 5.0. Lili Fan, Yutong Wang 0001, Hui Zhang 0091, Changxian Zeng, Yunjie Li, Chao Gou, Hui Yu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | AlphaRoute: Large-Scale Coordinated Route Planning via Monte Carlo Tree SearchabstractThis paper proposes AlphaRoute, an AlphaGo inspired algorithm for coordinating large-scale routes, built upon graph attention reinforcement learning and Monte Carlo Tree Search (MCTS). We first partition the road network into regions and model large-scale coordinated route planning as a Markov game, where each partitioned region is treated as a player instead of each driver. Then, AlphaRoute applies a bilevel optimization framework, consisting of several region planners and a global planner, where the region planner coordinates the route choices for vehicles located in the region and generates several strategies, and the global planner evaluates the combination of strategies. AlphaRoute is built on graph attention network for evaluating each state and MCTS algorithm for dynamically visiting and simulating the future state for narrowing down the search space. AlphaRoute is capable of 1) bridging user fairness and system efficiency, 2) achieving higher search efficiency by alleviating the curse of dimensionality problems, and 3) making an effective and informed route planning by simulating over the future to capture traffic dynamics. Comprehensive experiments are conducted on two real-world road networks as compared with several baselines to evaluate the performance, and results show that AlphaRoute achieves the lowest travel time, and is efficient and effective for coordinating large-scale routes and alleviating the traffic congestion problem. The code will be publicly available. Guiyang Luo, Yantao Wang, Hui Zhang 0091, Quan Yuan 0004 |
AAAI | 3 |
| 2023 | SSC3OD: Sparsely Supervised Collaborative 3D Object Detection from LiDAR Point CloudsabstractCollaborative 3D object detection, with its improved interaction advantage among multiple agents, has been widely explored in autonomous driving. However, existing collaborative 3D object detectors in a fully supervised paradigm heavily rely on large-scale annotated 3D bounding boxes, which is labor-intensive and time-consuming. To tackle this issue, we propose a sparsely supervised collaborative 3D object detection framework SSC3OD, which only requires each agent to randomly label one object in the scene. Specifically, this model consists of two novel components, i.e., the pillar-based masked autoencoder (Pillar-MAE) and the instance mining module. The Pillar-MAE module aims to reason over high-level semantics in a self-supervised manner, and the instance mining module generates high-quality pseudo labels for collaborative detectors online. By introducing these simple yet effective mechanisms, the proposed SSC3OD can alleviate the adverse impacts of incomplete annotations. We generate sparse labels based on collaborative perception datasets to evaluate our method. Extensive experiments on three large-scale datasets reveal that our proposed SSC3OD can effectively improve the performance of sparsely supervised collaborative 3D object detectors. Yushan Han, Hui Zhang 0091, Honglei Zhang 0002, Yidong Li |
SMC | 2 |
| 2023 | ClusterST: Clustering Spatial-Temporal Network for Traffic ForecastingabstractTraffic forecasting aims to capture complex spatial-temporal dependencies and non-linear dynamics, which plays an indispensable role in intelligent transportation systems and other domains like neuroscience, climate, etc. Most recent works rely on graph convolutional networks (GCN) to model the dependencies and the dynamics. However, the over-smoothing issue of GCN would produce indistinguishable features among nodes, leading to poor expressivity and weak capability of modeling complex dependencies and dynamics. To address this issue, we present a novel clustering spatial-temporal (ClusterST) unit, which incorporates unsupervised learning into GCN for extracting discriminative features. Specifically, we first exploit a neural network to learn a dynamic clustering, i.e., learning to partition the neighbors of each node into clusters at each time step. Two probabilistic losses are proposed to improve the separability of clusters. Then, the extracted features of different clusters can be distinguished. Based on the dynamically formed clusters, a vanilla GCN is applied to aggregate features within each cluster. By purely exploiting such a ClusterST unit, large improvements over the state-of-the-art are achieved. Furthermore, ClusterST units with a different number of clusters can be regarded as basic components to construct an inception-like ClusterST network for going deeper. We evaluate the framework on two real-world large-scale traffic datasets and observe an average improvement of 18.19% and 7.62% over state-of-the-art baselines, respectively. The code and models will be publicly available. Guiyang Luo, Hui Zhang 0091, Quan Yuan 0004, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Parallel Vision for Intelligent Transportation Systems in Metaverse: Challenges, Solutions, and Potential ApplicationsabstractMetaverse and intelligent transportation system (ITS) are disruptive technologies that have the potential to transform the current transportation system by decreasing traffic accidents and improving driving safety. The integration of Metaverse and transportation technology, called metaverse transportation system (MTS), can greatly improve the intelligence of real transportation system. The digital models built in MTS help to simulate the full life cycle of physical entities, which equip the virtual space with controllability and flexibility. In this article, we concentrate on the field of environment perception, which is the basic function of intelligent vehicles in MTS. To overcome the poor scalability of traditional environment perception methods, we develop the framework of parallel vision for ITS in metaverse (PVITS), consisting of construction of virtual transportation space, model learning based on computational experiments, and feedback optimization based on parallel execution. This article highlights opportunities brought by PVITS in terms of model precision and generalization improvement. Then, the challenges of PVITS are discussed, i.e., distribution difference between virtual and real transportation space, structure design and theoretical interpretation of vision models, and data security and privacy in virtual transportation space. After that, we present several solutions to tackle the application challenges and fully exploit the superior characteristics of PVITS while attenuating their negative side effects. Some potential applications are also given to represent the effectiveness and reliability of PVITS. Hui Zhang 0091, Guiyang Luo, Yidong Li, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Complementarity-Enhanced and Redundancy-Minimized Collaboration Network for Multi-agent PerceptionabstractMulti-agent collaborative perception depends on sharing sensory information to improve perception accuracy and robustness, as well as to extend coverage. The cooperative shared information between agents should achieve an equilibrium between redundancy and complementarity, thus creating a concise and composite representation. To this end, this paper presents a complementarity-enhanced and redundancy-minimized collaboration network (CRCNet), for efficiently guiding and supervising the fusion among shared features. Our key novelties lie in two aspects. First, each fused feature is forced to bring about a marginal gain by exploiting a contrastive loss, which can supervise our model to select complementary features. Second, mutual information is applied to measure the dependence between fused feature pairs and the upper bound of mutual information is minimized to encourage independence, thus guiding our model to select irredundant features. Furthermore, the above modules are incorporated into a feature fusion network CRCNet. Our quantitative and qualitative experiments in collaborative object detection show that CRCNet performs better than the state-of-the-art methods. Guiyang Luo, Hui Zhang 0091, Quan Yuan 0004 |
ACM Multimedia | 2 |
| 2022 | CMAN: Leaning Global Structure Correlation for Monocular 3D Object DetectionabstractThe key to 3D object detection is proper utilization of depth data. Compared with LiDAR based approaches, 3D object detection from a single image remains a challenging task due to the lack of structure information. Recent methods leverage monocular depth estimation as a way to produce 2D depth maps, and adopt the depth maps as additional source of input to explore structure information. However, these methods either encode local structure correlations, or encode long range structure correlations by iteratively passing local messages. In this work, we propose a cross modal attention network (CMAN) for monocular 3D object detection. It is built upon the self-attention module which learns attention map from single modal data. Our CMAN is able to encode structure correlations from depth data, and embed the structure correlations with appearance information which is learned from RGB data. Thanks to the attention learning mechanism, our CMAN learns global structure correlations without iteration. In order to reduce the computational burden, our CMAN adopts a novel node sampler to eliminate redundant nodes during the attention map calculation. Experiment results on benchmark KITTI3D dataset show that our proposed CMAN outperforms the state-of-the-art methods. Yuanzhouhan Cao, Hui Zhang 0091, Yidong Li, Chao Ren 0002, Congyan Lang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | ESTNet: Embedded Spatial-Temporal Network for Modeling Traffic Flow DynamicsabstractAccurate spatial-temporal prediction is a fundamental building block of many real-world applications such as traffic scheduling and management, environment policy making, and public safety. This problem is still challenging due to nonlinear, complicated, and dynamic spatial-temporal dependencies. To address these challenges, we propose a novel embedded spatial-temporal network (ESTNet), which extracts efficient features to model the dynamic correlations and then exploits three-dimension convolution to synchronously model the spatial-temporal dependencies. Specifically, we propose multi-range graph convolution networks for extracting multi-scale static features from the fine-grained road network. Meanwhile, dynamic features are extracted from real-time traffic using a gated recurrent unit network. These features can be applied to identify the dynamic and flexible correlations among sensors and make it possible to exploit a three-dimension convolution unit (3DCon) to simultaneously model the spatial-temporal dependencies. Furthermore, we propose a residual network by stacking multiple 3DCon to capture the nonlinear and complicated dependencies. The effectiveness and superiority of ESTNet are verified on two real-world datasets, and experiments show ESTNet outperforms the state-of-the-art with a significant margin. The code and models will be publicly available. Guiyang Luo, Hui Zhang 0091, Quan Yuan 0004, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | C2FDA: Coarse-to-Fine Domain Adaptation for Traffic Object DetectionabstractObject detection in traffic scenes has attracted considerable attention from both academia and industry recently. Modern detectors achieve excellent performance under a simple constrained environment while performing poorly under the actual complex and open traffic environment. Therefore, the capability of adapting to new and unseen domains is a key factor for the large-scale application and proliferation of detectors in autonomous driving. To this end, this paper proposes a novel category-induced coarse-to-fine domain adaptation approach (C2FDA) for cross-domain object detection, which consists of three pivotal components: (1) Attention-induced coarse-grained alignment module (ACGA), which strengthens the distribution alignment across disparate domains within the foreground features in category-agnostic way by the minimax optimization between the domain classifier and the backbone feature extractor; (2) Attention-induced feature selection module, which assists the model to emphasize the crucial foreground features and enables the ACGA to focus on the relevant and discriminative foreground features, without being affected by the distribution of inconsequential background features; (3) Category-induced fine-grained alignment module (CFGA), which reduces the domain shift in category-aware way by minimizing the distance of centroids with the same category from different domains and maximizing that of centroids with disparate categories. We evaluate the performance of our approach in various source/target domain pairs and comprehensive results demonstrate that C2FDA significantly outperforms the state-of-the-art on multiple domain adaptation scenarios, i.e., the synthetic-to-real adaptation, the weather adaptation, and the cross camera adaptation. Hui Zhang 0091, Guiyang Luo, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |