VLDB 2026 Research / reviewers in the wild / expert
Hongkai Yu
dblp:150/6670
· DBLP profile ↗
64ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0001-5383-8913ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 20 since 2021Systems, architecture and hardware · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ADVersa: Abductive Driving Accident Video UnderstandingabstractUnderstanding traffic accident scenes is a long-standing research for vision-based safe driving. It seeks to answer why accidents occur, how near-crash scenes develop, and what the key elements of an accident are. This research is challenging due to the scarcity and fragmentation of accident data, as well as the complex accident environments. To study this, we present a framework of Abductive Driving accident Video understanding (ADVersa), which infers a plausible visual and textual explanation for the absent near-crash scenes. ADVersa underscores three groups of tasks: 1) visual past recovery of near-crash scenes, 2) visual prediction of near-crash scenes, and 3) accident cause involved video synthesis. To support the study, we first contribute MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild driving accident videos with temporally aligned text descriptions, 2.23 million well-annotated object boxes, and 58,650 pairs of video-based accident cause texts. We then propose an Abductive CLIP model and a Contrastive Graph Video Pre-training (CGVP) model, which exploit relation-aware cross-modal semantic learning to drive spatially abductive and temporally abductive accident video diffusion. Extensive experiments verify the superiority of ADVersa to the state-of-the-art approaches on different tasks, i.e., historical near-crash video frame recovering, crashing video frame prediction, textual accident cause and category reasoning, normal-to-accident video synthesis, and accident video editing. With these efforts, we hope this research can advance the progress on multimodal accident video understanding. Lei-Lei Li, Jianwu Fang, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | FedSTA: Spatio-Temporal Alternation for Efficient Federated Multi-Task LearningabstractFederated multi-task learning (FMTL) faces significant challenges due to resource constraints and negative transfer among tasks. Existing methods, such as MAS, rely on task grouping and require multiple backbone models, resulting in increased complexity. To address this issue, we propose FedSTA, a novel FMTL framework that leverages affinity-based task partitions while maintaining a single shared backbone model. We design two task alternation strategies: spatial alternation, which assigns different task subsets to distinct clients within the same round, and temporal alternation, which cycles through task subsets across different rounds. Both strategies effectively exploit task synergies to mitigate negative transfer without the need to split the backbone. Additionally, we propose Global Proximal Min-max Optimization, a novel task weighting mechanism specifically designed for FMTL, capable of capturing global task difficulty distributions and adaptively modulating task optimization priorities to enhance training balance and robustness. Extensive experiments on multiple datasets demonstrate that FedSTA consistently outperforms existing multi-backbone approaches in overall performance while maintaining comparable computational and communication overhead using only a single backbone. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | An Efficient and Accurate Dynamic Sparse Training Framework Based on Parameter-FreezingabstractFederated learning is a decentralized machine learning approach that consists of servers and clients. It protects data privacy during model training by keeping the training data locally in each client. However, the requirement for the server and clients to frequently synchronize the parameters of the model brings a heavy burden to the communication links, especially when the model size has grown drastically in recent years. Several methods have been proposed to compress the model size by sparsification to reduce the communication overhead, albeit with significant accuracy degradation. In this work, we propose methods to better trade-off between model accuracy and training efficiency in federated learning. Our first proposed method is a novel sparse mask readjustment rule on the server and the second is a parameter-freezing method during training on the clients. Experimental results show that the model accuracy has significantly improved when combining our proposed methods. For example, compared with the previous state-of-the-art methods with the same total amount of communication cost and computation FLOPs, the accuracy increases on average by 4% and 6% in our methods for CIFAR-10 and CIFAR-100 datasets on ResNet-18, respectively. On the other hand, when targeting the same accuracy, the proposed method can reduce the communication cost by 4-8 times for different datasets with different sparsity levels. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
AAAI | 4 |
| 2025 | Robust Multi-task Adversarial Attacks Using Min-max OptimizationabstractDeep neural networks have achieved exceptional performance across a wide range of applications but remain susceptible to adversarial attacks. While most prior research has focused on single-task scenarios, increasing attention is being directed toward adversarial attacks targeting multiple tasks simultaneously. However, existing methods often fail to balance attack performance across tasks in a multi-task model. These approaches typically aim to maximize the model’s overall loss, neglecting task-specific attack difficulties, which results in imbalanced attack performance among tasks. To address this challenge, we propose a novel multi-task adversarial attack method that ensures robust and balanced attack performance across multiple tasks. Our approach dynamically updates task-specific weighting factors through a min-max optimization during the attack, optimizing the worst-case attack performance across all tasks. Experimental results demonstrate that our method significantly enhances the worst-case attack performance across diverse datasets and attack strategies compared to existing approaches. By dynamically adjusting the attack intensity on the least vulnerable tasks, the min-max optimization significantly improves overall attack effectiveness as well as the worst-case performance by balancing the task weights. Jiacheng Guo, Lei Li 0066, Haochen Yang 0002, Baocheng Geng, Hongkai Yu, Minghai Qin, Tianyun Zhang |
ICASSP | 5 |
| 2025 | Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
ICCV | 5 |
| 2025 | V2X-DGW: Domain Generalization for Multi-Agent Perception Under Adverse Weather ConditionsabstractCurrent LiDAR-based Vehicle-to-Everything (V2X) multi-agent perception systems have shown the significant success on 3D object detection. While these models perform well in the trained clean weather, they struggle in unseen adverse weather conditions with the domain gap. In this paper, we propose a Domain Generalization based approach, named V2X-DGW, for LiDAR-based 3D object detection on multi-agent perception system under adverse weather conditions. Our research aims to not only maintain favorable multi-agent performance in the clean weather but also promote the performance in the unseen adverse weather conditions by learning only on the clean weather data. To realize the Domain Generalization, we first introduce the Adaptive Weather Augmentation (AWA) to mimic the unseen adverse weather conditions, and then propose two alignments for generalizable representation learning: Trust-region Weatherinvariant Alignment (TWA) and Agent-aware Contrastive Alignment (ACA). To evaluate this research, we add Fog, Rain, Snow conditions on two publicized multi-agent datasets based on physics-based models, resulting in two new datasets: OPV2V-w and V2XSet-w. Extensive experiments demonstrate that our V2X-DGW achieved significant improvements in the unseen adverse weathers. The code is available at https://github.com/Baolu1998/V2X-DGW. Xinyu Liu 0009, Runsheng Xu, Zhengzhong Tu, Jiacheng Guo, Qin Zou 0001, Xiaopeng Li 0020, Hongkai Yu |
ICRA | 9 |
| 2025 | V2X-DG: Domain Generalization for Vehicle-to-Everything Cooperative PerceptionabstractLiDAR-based Vehicle-to-Everything (V2X) cooperative perception has demonstrated its impact on the safety and effectiveness of autonomous driving. Since current cooperative perception algorithms are trained and tested on the same dataset, the generalization ability of cooperative perception systems remains underexplored. This paper is the first work to study the Domain Generalization problem of LiDAR-based V2X cooperative perception (V2X-DG) for 3D detection based on four widely-used open source datasets: OPV2V, V2XSet, V2V4Real and DAIR-V2X. Our research seeks to sustain high performance not only within the source domain but also across other unseen domains, achieved solely through training on source domain. To this end, we propose Cooperative Mixup Augmentation based Generalization (CMAG) to improve the model generalization capability by simulating the unseen cooperation, which is designed compactly for the domain gaps in cooperative perception. Furthermore, we propose a constraint for the regularization of the robust generalized feature representation learning: Cooperation Feature Consistency (CFC), which aligns the intermediately fused features of the generalized cooperation by CMAG and the early fused features of the original cooperation in source domain. Extensive experiments demonstrate that our approach achieves significant performance gains when generalizing to other unseen datasets while it also maintains strong performance on the source dataset. Zongzhe Xu, Xinyu Liu 0009, Jianwu Fang, Xiaopeng Li 0020, Hongkai Yu |
ICRA | 7 |
| 2025 | CoMamba: Real-time Cooperative Perception Unlocked with State-Space ModelsabstractCooperative perception systems play a vital role in enhancing the safety and efficiency of vehicular autonomy. Although recent studies have highlighted the efficacy of vehicle-to-everything (V2X) communication techniques in autonomous driving, a significant challenge persists: how to efficiently integrate multiple high-bandwidth features across an expanding network of connected agents such as vehicles and infrastructure. In this paper, we introduce CoMamba, a novel cooperative 3D detection framework designed to leverage state-space models for real-time onboard vehicle perception. Compared to prior state-of-the-art transformer-based models, CoMamba enjoys being a more scalable 3D model using bidirectional state space models, bypassing the quadratic complexity pain-point of attention mechanisms. Through extensive experimentation on V2X/V2V datasets, CoMamba achieves superior performance compared to existing methods while maintaining real-time processing capabilities. The proposed framework not only enhances object detection accuracy but also significantly reduces processing time, making it a promising solution for next-generation cooperative perception systems in intelligent transportation networks. Xinyu Liu 0009, Runsheng Xu, Jiachen Li 0001, Hongkai Yu, Zhengzhong Tu |
IROS | 6 |
| 2025 | DA3D: Domain-Aware Dynamic Adaptation for All-Weather Multimodal 3D DetectionabstractLiDAR-Radar fusion has been widely regarded as an effective strategy for enhancing sensor-level robustness in 3D perception under adverse weather. However, it remains fundamentally insufficient to address feature-level domain shifts induced by diverse weather conditions - a critical yet often overlooked bottleneck in multimodal 3D object detection. In this work, we advocate a new perspective: all-weather 3D detection should be formulated as a lightweight capacity allocation problem, rather than simply enlarging or duplicating models for each weather domain. To this end, we propose DA3D, a Domain-Aware Dynamic Adaptation framework that leverages LoRA as a domain-adaptive capacity controller for efficient and scalable feature modulation. In addition, we introduce a domain-aware rank adaptation strategy that dynamically reallocates LoRA capacity based on domain difficulty, allowing the model to focus its representational power where it matters most. Extensive experiments on the K-Radar benchmark show that DA3D consistently improves 3D detection across both radar-only and LiDAR-Radar fusion backbones, achieving +4.9% AP3D on RTNH, +3.8% on 3D-LRF, and +8.1% on L4DR at IoU=0.5. Notably, DA3D outperforms existing multi-weather modeling methods under the same parameter budget, offering a practical and scalable solution for robust all-weather 3D perception. The code is available at https://github.com/Dawns14/DA3D. Haochen Yang 0002, Lei Li 0066, Jiacheng Guo, Minghai Qin, Hongkai Yu, Tianyun Zhang |
ACM Multimedia | 6 |
| 2025 | Task-Aware Federated Multi-Task LearningabstractFederated Multi-Task Learning (FMTL) enables collaborative training of multiple tasks across decentralized clients, but faces two key challenges in practice: negative transfer among tasks and scalability under resource constraints. Task differences can cause gradient conflicts that degrade overall performance, while limited computation and storage on edge devices make it difficult to maintain accuracy with low overhead. Existing methods address these issues either by adopting multi-backbone architectures, which split tasks to reduce interference but incur substantial parameter and computation costs, or by performing naive global averaging, which ignores inter-task differences and fails to effectively mitigate negative transfer. To overcome these limitations, we propose Task-Aware Federated Multi-Task Learning (TA-FMTL), a single-backbone framework that balances accuracy and efficiency. TA-FMTL integrates two lightweight components: a min–max task-difficulty weighting strategy that dynamically allocates more updates to harder tasks for balanced optimization, and a variance-aware reputation aggregation that down-weights clients with high overall loss or unstable cross-task performance. This design enables robust coordination across heterogeneous tasks without task splitting. Experiments on the Taskonomy benchmark show that TA-FMTL consistently achieves better or comparable accuracy to state-of-the-art MAS variants while reducing parameters by up to 77.2% and FLOPs by 37.3% in challenging 5-task and 9-task settings, demonstrating its scalability and practicality for real-world FMTL under heterogeneous and resource-limited conditions. Lei Li 0066, Haochen Yang 0002, Jiacheng Guo, Hongkai Yu, Minghai Qin, Tianyun Zhang |
MMAsia | 4 |
| 2025 | A min-max optimization framework for sparse multi-task deep neural network
Jiacheng Guo, Huiming Sun, Minghai Qin, Hongkai Yu, Tianyun Zhang |
Neurocomputing | 5 |
| 2025 | BEVFix: Deep feature enhancement for robust 3D object detection
Jian Zhou 0011, Chi Chen 0002, Hongkai Yu, Bo Du 0001, Qin Zou 0001 |
Neural Networks | 4 |
| 2025 | EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video SynthesisabstractTraffic Accident Anticipation (TAA) in traffic scenes is a challenging problem for achieving zero fatalities in the future. Current approaches typically treat TAA as a supervised learning task needing the laborious annotation of accident occurrence duration. However, the inherent long-tailed, uncertain, and fast-evolving nature of traffic scenes has the problem that real causal parts of accidents are difficult to identify and are easily dominated by data bias, resulting in a background confounding issue. Thus, we propose an Attentive Video Diffusion (AVD) model that synthesizes additional accident video clips by generating the causal part in dashcam videos, i.e., from normal clips to accident clips. AVD aims to generate causal video frames based on accident or accident-free text prompts while preserving the style and content of frames for TAA after video generation. This approach can be trained using datasets collected from various driving scenes without any extra annotations. Additionally, AVD facilitates an Equivariant TAA (EQ-TAA) with an equivariant triple loss for an anchor accident-free video clip, along with the generated pair of contrastivepseudo-normalandpseudo-accidentclips. Extensive experiments have been conducted to evaluate the performance of AVD and EQ-TAA, and competitive performance compared to state-of-the-art methods has been obtained. Jianwu Fang, Lei-Lei Li, Zhedong Zheng, Hongkai Yu, Jianru Xue, Zhengguo Li, Tat-Seng Chua |
IEEE Trans. Multim. | 4 |
| 2024 | SQLdepth: Generalizable Self-Supervised Fine-Structured Monocular Depth EstimationabstractRecently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and struggle to recover fine-grained scene details. In this paper, we introduce SQLdepth, a novel approach that can effectively learn fine-grained scene structure priors from ego-motion. In SQLdepth, we propose a novel Self Query Layer (SQL) to build a self-cost volume and infer depth from it, rather than inferring depth from feature maps. We show that, the self-cost volume is an effective inductive bias for geometry learning, which implicitly models the single-frame scene geometry, with each slice of it indicating a relative distance map between points and objects in a latent space. Experimental results on KITTI and Cityscapes show that our method attains remarkable state-of-the-art performance, and showcases computational efficiency, reduced training complexity, and the ability to recover fine-grained scene details. Moreover, the self-matching-oriented relative distance querying in SQL improves the robustness and zero-shot generalization capability of SQLdepth. Code is available at https://github.com/hisfog/SfMNeXt-Impl. Youhong Wang, Yunji Liang, Shaohui Jiao, Hongkai Yu |
AAAI | 5 |
| 2024 | Abductive Ego-View Accident Video Understanding for Safe Driving PerceptionabstractWe present MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild ego-view accident videos, each with temporally aligned text descriptions. We annotate over 2.23 mil-lion object boxes and 58,650 pairs of video-based accident reasons, covering 58 accident categories. MM-AU supports various accident understanding tasks, particularly multimodal video diffusion to understand accident cause-effect chains for safe driving. With MM-AU, we present an Abductive accident Video unders tanding framework for Safe Driving perception (AdVersa-SD). AdVersa-SD performs video diffusion via an Object-Centric Video Diffusion (OAVD) method which is driven by an abductive CLIP model. This model involves a contrastive interaction loss to learn the pair co-occurrence of normal, near-accident, accident frames with the corresponding text descriptions, such as accident reasons, prevention advice, and accident categories. OAVD enforces the object region learning while fixing the content of the original frame background in video generation, to find the dominant objects for certain accidents. Extensive experiments verify the abductive ability of AdVersa-SD and the superiority of OAVD against the state-of-the-art diffusion models. Additionally, we provide care-ful benchmark evaluations for object detection and accident reason answering since AdVersa-SD relies on precise object and accident reason information. Jianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv 0001, Jianru Xue, Tat-Seng Chua |
CVPR | 5 |
| 2024 | Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous DrivingabstractVision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability, especially compared to LiDAR-based systems. However, these systems often struggle in low-light conditions, potentially compromising their performance and safety. To address this, our paper introduces LightDiff, a domain-tailored framework designed to enhance the low-light image quality for autonomous driving applications. Specifically, we employ a multi-condition controlled diffusion model. LightDiff works without any human-collected paired data, leveraging a dynamic data degradation process instead. It incorporates a novel multi-condition adapter that adaptively controls the input weights from different modalities, including depth maps, RGB images, and text captions, to effectively illuminate dark scenes while maintaining context consistency. Furthermore, to align the enhanced images with the detection model's knowledge, LightDiff employs perception-specific scores as rewards to guide the diffusion training process through reinforcement learning. Extensive experiments on the nuScenes datasets demonstrate that LightDiff can significantly improve the performance of several state-of-the-art 3D detectors in night-time conditions while achieving high visual quality scores, highlighting its potential to safeguard autonomous driving. Zhengzhong Tu, Xinyu Liu 0009, Qing Guo 0005, Felix Juefei-Xu, Runsheng Xu, Hongkai Yu |
CVPR | 8 |
| 2024 | AdvGPS: Adversarial GPS for Multi-Agent Perception AttackabstractThe multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can be influenced by a range of factors, including wireless transmission and obstructions like buildings. Given the pivotal role of GPS signals in perception fusion and the potential for various interference, it becomes imperative to investigate whether specific GPS signals can easily mislead the multi-agent perception system. To address this concern, we frame the task as an adversarial attack challenge and introduce ADVGPS, a method capable of generating adversarial GPS signals which are also stealthy for individual agents within the system, significantly reducing object detection accuracy. To enhance the success rates of these attacks in a black-box scenario, we introduce three types of statistically sensitive natural discrepancies: appearance-based discrepancy, distribution-based discrepancy, and task-aware discrepancy. Our extensive experiments on the OPV2V dataset demonstrate that these attacks substantially undermine the performance of state-of-the-art methods, showcasing remarkable transferability across different point cloud based 3D detection systems. This alarming revelation underscores the pressing need to address security implications within multi-agent perception systems, thereby underscoring a critical area of research. The code is available at https://github.com/jinlong17/AdvGPS. Xinyu Liu 0009, Jianwu Fang, Felix Juefei-Xu, Qing Guo 0005, Hongkai Yu |
ICRA | 7 |
| 2024 | Breaking Data Silos: Cross-Domain Learning for Multi-Agent Perception from Independent Private SourcesabstractThe diverse agents in multi-agent perception systems may be from different companies. Each company might use the identical classic neural network architecture based encoder for feature extraction. However, the data source to train the various agents is independent and private in each company, leading to the Distribution Gap of different private data for training distinct agents in multi-agent perception system. The data silos by the above Distribution Gap could result in a significant performance decline in multi-agent perception. In this paper, we thoroughly examine the impact of the distribution gap on existing multi-agent perception systems. To break the data silos, we introduce the Feature Distribution-aware Aggregation (FDA) framework for cross-domain learning to mitigate the above Distribution Gap in multi-agent perception. FDA comprises two key components: Learnable Feature Compensation Module and Distribution-aware Statistical Consistency Module, both aimed at enhancing intermediate features to minimize the distribution gap among multi-agent features. Intensive experiments on the public OPV2V and V2XSet datasets underscore FDA’s effectiveness in point cloud-based 3D object detection, presenting it as an invaluable augmentation to existing multi-agent perception systems. The code is available at https://github.com/jinlong17/BDS-V2V. Xinyu Liu 0009, Runsheng Xu, Jiaqi Ma 0003, Hongkai Yu |
ICRA | 6 |
| 2024 | S2R-ViT for Multi-Agent Cooperative Perception: Bridging the Gap from Simulation to RealityabstractDue to the lack of enough real multi-agent data and time-consuming of labeling, existing multi-agent cooperative perception algorithms usually select the simulated sensor data for training and validating. However, the perception performance is degraded when these simulation-trained models are deployed to the real world, due to the significant domain gap between the simulated and real data. In this paper, we propose the first Simulation-to-Reality transfer learning framework for multi-agent cooperative perception using a novel Vision Transformer, named as S2R-ViT, which considers both the Deployment Gap and Feature Gap between simulated and real data. We investigate the effects of these two types of domain gaps and propose a novel uncertainty-aware vision transformer to effectively relief the Deployment Gap and an agent-based feature adaptation module with inter-agent and ego-agent discriminators to reduce the Feature Gap. Our intensive experiments on the public multi-agent cooperative perception datasets OPV2V and V2V4Real demonstrate that the proposed S2R-ViT can effectively bridge the gap from simulation to reality and outperform other methods significantly for point cloud-based 3D object detection. Runsheng Xu, Xinyu Liu 0009, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu |
ICRA | 7 |
| 2024 | Vehicle Behavior Prediction by Episodic-Memory Implanted NDTabstractIn autonomous driving, predicting the behavior (turning left, stopping, etc.) of target vehicles is crucial for the self-driving vehicle to make safe decisions and avoid accidents. Existing deep learning-based methods have shown excellent and accurate performance, but the black-box nature makes it untrustworthy to apply them in practical use. In this work, we explore the interpretability of behavior prediction of target vehicles by an Episodic Memory implanted Neural Decision Tree (abbrev. eMem-NDT). The structure of eMem-NDT is constructed by hierarchically clustering the text embedding of vehicle behavior descriptions. eMem-NDT is a neural-backed part of a pre-trained deep learning model by changing the soft-max layer of the deep model to eMem-NDT, for grouping and aligning the memory prototypes of the historical vehicle behavior features in training data on a neural decision tree. Each leaf node of eMem-NDT is modeled by a neural network for aligning the behavior memory prototypes. By eMem-NDT, we infer each instance in behavior prediction of vehicles by bottom-up Memory Prototype Matching (MPM) (searching the appropriate leaf node and the links to the root node) and top-down Leaf Link Aggregation (LLA) (obtaining the probability of future behaviors of vehicles for certain instances). We validate eMem-NDT on BLVD and LOKI datasets, and the results show that our model can obtain a superior performance to other methods with clear explainability. The code is available in https://github.com/JWFangit/eMem-NDT. Peining Shen, Jianwu Fang, Hongkai Yu, Jianru Xue |
ICRA | 3 |
| 2024 | A Min-Max Optimization Framework for Multi-task Deep Neural Network CompressionabstractMulti-task learning is a subfield of machine learning in which the data is trained with a shared model to solve different tasks simultaneously. Instead of training multiple models corresponding to different tasks, we only need to train a single model with shared parameters by using multi-task learning. Multi-task learning highly reduces the number of parameters in the machine learning models and thus reduces the computational and storage requirements. When we apply multi-task learning on deep neural networks (DNNs), we need to further compress the model since the model size of a single DNN is still a critical challenge to many computation systems, especially for edge platforms. However, when model compression is applied to multi-task learning, it is challenging to maintain the performance of all the different tasks. To deal with this challenge, we propose a min-max optimization framework for the training of highly compressed multi-task DNN models. Our proposed framework can automatically adjust the learnable weighting factors corresponding to different tasks to guarantee that the task with worst-case performance across all the different tasks will be optimized. Jiacheng Guo, Huiming Sun, Minghai Qin, Hongkai Yu, Tianyun Zhang |
ISCAS | 4 |
| 2024 | VehicleGAN: Pair-flexible Pose Guided Image Synthesis for Vehicle Re-identificationabstractVehicle Re-identification (Re-ID) has been broadly studied in the last decade; however, the different camera view angles leading to confused discrimination in the feature subspace for the vehicles of various poses, is still challenging for the Vehicle Re-ID models in the real world. To promote the Vehicle Re-ID models, this paper proposes to synthesize a large number of vehicle images in the target pose, whose idea is to project the vehicles of diverse poses into the unified target pose so as to enhance feature discrimination. Considering that the paired data of the same vehicles in different traffic surveillance cameras might be not available in the real world, we propose the first Pair-flexible Pose Guided Image Synthesis method for Vehicle Re-ID, named as VehicleGAN in this paper, which works for both supervised and unsupervised settings without the knowledge of geometric 3D models. Because of the feature distribution difference between real and synthetic data, simply training a traditional metric learning based Re-ID model with data-level fusion (i.e., data augmentation) is not satisfactory, therefore we propose a new Joint Metric Learning (JML) via effective feature-level fusion from both real and synthetic data. Intensive experimental results on the public VeRi-776 and VehicleID datasets prove the accuracy and effectiveness of our proposed VehicleGAN and JML. Ping Liu 0004, Lan Fu, Jianwu Fang, Zhigang Xu 0001, Hongkai Yu |
IV | 7 |
| 2024 | V2X-DSI: A Density-Sensitive Infrastructure LiDAR Benchmark for Economic Vehicle-to-Everything Cooperative PerceptionabstractRecent research has demonstrated that the Vehicle-to-Everything (V2X) communication techniques can fundamentally improve the perception system for autonomous driving by collaborating between vehicle and infrastructure sensors. LiDAR is the commonly-used sensor for V2X autonomous driving due to its robustness in challenging scenarios. However, the LiDAR sensor is expensive, so the cost of equipping LiDAR sensors to a large number of infrastructures on the large-scale roadway network is extremely high, which has limited the wide deployment of the V2X cooperative perception system. How to discover an economic V2X cooperative perception system is never been well studied before. Inspired by the cost difference of the various point cloud densities of LiDAR, we propose the first Density-Sensitive Infrastructure LiDAR benchmark for economic V2X cooperative perception, named V2X-DSI, in this paper. Using the proposed V2X-DSI benchmark, we analyze the effect of cooperative perception performance under different beam infrastructure LiDAR. We specifically assess three state-of-the-art methods, i.e., OPV2V, V2X-ViT, and CoBEVT, using our V2X-DSI dataset. The results indicate that varying beam infrastructure LiDAR sensors play a crucial role in influencing cooperative perception performance. Xinyu Liu 0009, Runsheng Xu, Jiaqi Ma 0003, Xiaopeng Li 0020, Hongkai Yu |
IV | 7 |
| 2024 | EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAVabstractVehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed for the vehicle detection and tracking in UAV images. Recent studies show that adding an adversarial patch on objects can fool the well-trained deep neural networks based object detectors, posing security concerns to the downstream tasks. However, the current public UAV datasets might ignore the diverse altitudes, vehicle attributes, fine-grained instance-level annotation in mostly side view with blurred vehicle roof, so none of them is good to study the adversarial patch based vehicle detection attack problem. In this paper, we propose a new dataset named EVD4UAV as an altitude-sensitive benchmark to evade vehicle detection in UAV with 6,284 images and 90,886 fine-grained annotated vehicles. The EVD4UAV dataset has diverse altitudes (50m, 70m, 90m), vehicle attributes (color, type), fine-grained annotation (horizontal and rotated bounding boxes, instance-level mask) in top view with clear vehicle roof. One white-box and two black-box patch based attack methods are implemented to attack three classic deep neural networks based object detectors on EVD4UAV. The experimental results show that these representative attack methods could not achieve the robust altitude-insensitive attack performance. Huiming Sun, Jiacheng Guo, Zibo Meng, Tianyun Zhang, Jianwu Fang, Yuewei Lin, Hongkai Yu |
IV | 7 |
| 2024 | Defense against Adversarial Cloud Attack on Remote Sensing Salient Object DetectionabstractDetecting the salient objects in a remote sensing image has wide applications. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images with remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original image, could result in a collapse for the well-trained deep learning model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learnable pre-processing to the adversarial cloudy images to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing dataset (EORSSD) show the promising defense against adversarial cloud attacks. Huiming Sun, Lan Fu, Qing Guo 0005, Zibo Meng, Tianyun Zhang, Yuewei Lin, Hongkai Yu |
WACV | 8 |
| 2024 | Adversarial Relighting Against Face RecognitionabstractDeep face recognition (FR) has achieved significantly high accuracy on several challenging datasets and fosters successful real-world applications, even showing high robustness to the illumination variation that is usually regarded as a main threat to the FR system. However, in the real world, illumination variation caused by diverse lighting conditions cannot be fully covered by the limited face dataset. In this paper, we study the threat of lighting against FR from a new angle,i.e.,adversarial attack, and identify a new task,i.e.,adversarial relighting. Given a face image, adversarial relighting aims to produce a naturally relighted counterpart while fooling the state-of-the-art deep FR methods. To this end, we first propose the physical model-based adversarial relighting attack (ARA) denoted asalbedo-quotient-based adversarial relighting attack (AQ-ARA). It generates natural adversarial lighting under the guidance of FR systems and synthesizes adversarially relighted face images. Moreover, we propose theauto-predictive adversarial relighting attack (AP-ARA)by training an adversarial relighting network (ARNet) to automatically predict the adversarial lighting in a one-step manner according to different input faces, allowing efficiency-sensitive applications. More importantly, we propose to transfer the above digital attacks tophysical ARA (Phy-ARA)through a precise relighting device, making the estimated adversarial lighting condition reproducible in the real world. We validate our methods on several state-of-the-art deep FR methods on two public datasets. The extensive and insightful results demonstrate our work can generate realistic adversarial relighted face images fooling face recognition tasks easily, revealing the threat of specific light directions and strengths. Qian Zhang 0051, Qing Guo 0005, Ruijun Gao, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Learning Cross-modality Interaction for Robust Depth Perception of Autonomous DrivingabstractAs one of the fundamental tasks of autonomous driving, depth perception aims to perceive physical objects in three dimensions and to judge their distances away from the ego vehicle. Although great efforts have been made for depth perception, LiDAR-based and camera-based solutions have limitations with low accuracy and poor robustness for noise input. With the integration of monocular cameras and LiDAR sensors in autonomous vehicles, in this article, we introduce a two-stream architecture to learn the modality interaction representation under the guidance of an image reconstruction task to compensate for the deficiencies of each modality in a parallel manner. Specifically, in the two-stream architecture, the multi-scale cross-modality interactions are preserved via a cascading interaction network under the guidance of the reconstruction task. Next, the shared representation of modality interaction is integrated to infer the dense depth map due to the complementarity and heterogeneity of the two modalities. We evaluated the proposed solution on the KITTI dataset and CALAR synthetic dataset. Our experimental results show that learning the coupled interaction of modalities under the guidance of an auxiliary task can lead to significant performance improvements. Furthermore, our approach is competitive against the state-of-the-art models and robust against the noisy input. The source code is available at https://github.com/tonyFengye/Code/tree/master . Yunji Liang, Nengzhen Chen, Zhiwen Yu 0001, Lei Tang 0002, Hongkai Yu, Bin Guo 0001, Daniel Dajun Zeng |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | V2V4Real: A Real-World Large-Scale Dataset for Vehicle-to-Vehicle Cooperative PerceptionabstractModern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has demonstrated that the Vehicle-to-Vehicle (V2V) cooperative perception system has great potential to revolutionize the autonomous driving industry. However, the lack of a real-world dataset hinders the progress of this field. To facilitate the development of cooperative perception, we present V2V4Real, the first large-scale real-world multi-modal dataset for V2V perception. The data is collected by two vehicles equipped with multi-modal sensors driving together through diverse scenarios. Our V2V4Real dataset covers a driving area of 410 km, comprising 20K LiDAR frames, 40K RGB frames, 240K annotated 3D bounding boxes for 5 classes, and HDMaps that cover all the driving routes. V2V4Real introduces three perception tasks, including cooperative 3D object detection, cooperative 3D object tracking, and Sim2Real domain adaptation for cooperative perception. We provide comprehensive benchmarks of recent cooperative perception algorithms on three tasks. The V2V4Real dataset can be found at research.seas.ucla.edu/mobility-lab/v2v4real/. Runsheng Xu, Xin Xia 0007, Hanzhao Li, Zhengzhong Tu, Zonglin Meng, Hao Xiang 0001, Rui Song 0007, Hongkai Yu, Bolei Zhou, Jiaqi Ma 0003 |
CVPR | 11 |
| 2023 | Bridging the Domain Gap for Multi-Agent PerceptionabstractExisting multi-agent perception algorithms usually select to share deep neural features extracted from raw sensing data between agents, achieving a trade-off between accuracy and communication bandwidth limit. However, these methods assume all agents have identical neural networks, which might not be practical in the real world. The transmitted features can have a large domain gap when the models differ, leading to a dramatic performance drop in multi-agent perception. In this paper, we propose the first lightweight framework to bridge such domain gaps for multi-agent perception, which can be a plug-in module for most of the existing systems while maintaining confidentiality. Our framework consists of a learnable feature resizer to align features in multiple dimensions and a sparse cross-domain transformer for domain adaption. Extensive experiments on the public multi-agent perception dataset V2XSet have demonstrated that our method can effectively bridge the gap for features from different domains and outperform other baseline methods significantly by at least 8% for point-cloud-based 3D object detection. Runsheng Xu, Hongkai Yu, Jiaqi Ma 0003 |
ICRA | 4 |
| 2023 | Domain Adaptive Object Detection for Autonomous Driving under Foggy WeatherabstractMost object detection methods for autonomous driving usually assume a consistent feature distribution between training and testing data, which is not always the case when weathers differ significantly. The object detection model trained under clear weather might be not effective enough on the foggy weather because of the domain gap. This paper proposes a novel domain adaptive object detection framework for autonomous driving under foggy weather. Our method leverages both image-level and object-level adaptation to diminish the domain discrepancy in image style and object appearance. To further enhance the model’s capabilities under challenging samples, we also come up with a new adversarial gradient reversal layer to perform adversarial mining for the hard examples together with domain adaptation. Moreover, we propose to generate an auxiliary domain by data augmentation to enforce a new domain-level metric regularization. Experimental results on public benchmarks show the effectiveness and accuracy of the proposed method. The code is available at https://github.com/jinlong17/DA-Detect. Runsheng Xu, Jin Ma 0005, Qin Zou 0001, Jiaqi Ma 0003, Hongkai Yu |
WACV | 6 |
| 2023 | Pik-Fix: Restoring and Colorizing Old PhotosabstractRestoring and inpainting the visual memories that are present, but often impaired, in old photos remains an intriguing but unsolved research topic. Decades-old photos often suffer from severe and commingled degradation such as cracks, defocus, and color-fading, which are difficult to treat individually and harder to repair when they interact. Deep learning presents a plausible avenue, but the lack of large-scale datasets of old photos makes addressing this restoration task very challenging. Here we present a novel reference-based end-to-end learning framework that is able to both repair and colorize old, degraded pictures. Our proposed framework consists of three modules: a restoration sub-network that conducts restoration from degradations, a similarity network that performs color histogram matching and color transfer, and a colorization subnet that learns to predict the chroma elements of images conditioned on chromatic reference signals. The overall system makes uses of color histogram priors from reference images, which greatly reduces the need for large-scale training data. We have also created a first-of-a-kind public dataset of real old photos that are paired with ground truth "pristine" photos that have been manually restored by PhotoShop experts. We conducted extensive experiments on this dataset and synthetic datasets, and found that our method significantly outperforms previous state-of-the-art models using both qualitative comparisons and quantitative measurements. The code is available at https://github.com/DerrickXuNu/Pik-Fix. Runsheng Xu, Zhengzhong Tu, Yuanqi Du, Zibo Meng, Jiaqi Ma 0003, Alan C. Bovik, Hongkai Yu |
WACV | 9 |
| 2023 | An end-to-end network for co-saliency detection in one single image
Yuanhao Yue, Qin Zou 0001, Hongkai Yu, Qian Wang 0002, Zhongyuan Wang 0001, Song Wang 0002 |
Sci. China Inf. Sci. | 3 |
| 2023 | Deep learning for image inpainting: A survey
Hanyu Xiang, Qin Zou 0001, Muhammad Ali Nawaz, Fan Zhang 0006, Hongkai Yu |
Pattern Recognit. | 6 |
| 2023 | Coarse-to-Fine Task-Driven Inpainting for Geoscience ImagesabstractThe processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that all the images are clear. However, in many real-world cases, the geoscience images might contain occlusions during the image acquisition. This problem actually implies the image inpainting problem in computer vision and multimedia. As far as we know, all the existing image inpainting algorithms learn to repair the occluded regions for a better visualization quality, they are excellent for natural images but not good enough for geoscience images, and they never consider the following gescience task when developing inpainting methods. This paper aims to repair the occluded regions for a better geoscience task performance and advanced visualization quality simultaneously, without changing the current deployed deep learning based geoscience models. Because of the complex context of geoscience images, we propose a coarse-to-fine encoder-decoder network with the help of designed coarse-to-fine adversarial context discriminators to reconstruct the occluded image regions. Due to the limited data of geoscience images, we propose a MaskMix based data augmentation method, which augments inpainting masks instead of augmenting original images, to exploit the limited geoscience image data. The experimental results on three public geoscience datasets for remote sensing scene recognition, cross-view geolocation and semantic segmentation tasks respectively show the effectiveness and accuracy of the proposed method. The code is available at:https://github.com/HMS97/Task-driven-Inpainting. Huiming Sun, Jin Ma 0005, Qing Guo 0005, Qin Zou 0001, Shaoyue Song, Yuewei Lin, Hongkai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Heterogeneous Trajectory Forecasting via Risk and Scene Graph LearningabstractHeterogeneous trajectory forecasting is critical for intelligent transportation systems, but it is challenging because of the difficulty of modeling the complex interaction relations among the heterogeneous road agents as well as their agent-environment constraints. In this work, we propose a risk and scene graph learning method for trajectory forecasting of heterogeneous road agents, which consists of a Heterogeneous Risk Graph (HRG) and a Hierarchical Scene Graph (HSG) from the aspects of agent category and their movable semantic regions. HRG groups each kind of road agent and calculates their interaction adjacency matrix based on an effective collision risk metric. HSG of the driving scene is modeled by inferring the relationship between road agents and road semantic layout aligned by the road scene grammar. Based on this formulation, we can obtain effective trajectory forecasting in driving situations, and comparable performance to other state-of-the-art approaches is presented by extensive experiments on the nuScenes, ApolloScape, and Argoverse datasets. Jianwu Fang, Pu Zhang 0001, Hongkai Yu, Jianru Xue |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods. Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 4 |
| 2022 | AdvTraffic: Obfuscating Encrypted Traffic with Adversarial ExamplesabstractWebsite fingerprinting can reveal which sensitive website a user visits over encrypted network traffic. Obfuscating encrypted traffic, e.g., adding dummy packets, is considered as a primary approach to defend against website fingerprinting. How-ever, existing defenses relying on traffic obfuscation are either ineffective or introduce significant overheads. As recent website fingerprinting attacks heavily rely on deep neural networks to achieve high accuracy, producing adversarial examples could be utilized as a new way to obfuscate encrypted traffic. Unfortunately, existing adversarial example algorithms are designed for images and do not consider unique challenges for network traffic.In this paper, we design a new method, named AdvTraffic, which can customize perturbations produced by any existing adversarial example algorithm on images and derive adversarial examples over encrypted traffic. Our experimental results show that the integration of AdvTraffic, particularly with Generative Adversarial Networks, can effectively mitigate the accuracy of website fingerprinting from 95.0% to 10.2%, even if an attacker retrains a classifier with defended traffic. Compared to other defenses, our method outperforms most of them in mitigating attack accuracy and offers the lowest bandwidth overhead. Jimmy Dani, Hongkai Yu, Wenhai Sun, Boyang Wang 0001 |
IWQoS | 3 |
| 2022 | Let There Be Light: Improved Traffic Surveillance via Detail Preserving Night-to-Day TransferabstractIn recent years, image and video surveillance have made considerable progresses to the Intelligent Transportation Systems (ITS) with the help of deep Convolutional Neural Networks (CNNs). As one of the state-of-the-art perception approaches, detecting the interested objects in each frame of video surveillance is widely desired by ITS. Currently, object detection shows remarkable efficiency and reliability in standard scenarios such as daytime scenes with favorable illumination conditions. However, in face of adverse conditions such as the nighttime, object detection loses its accuracy significantly. One of the main causes of the problem is the lack of sufficient annotated detection datasets of nighttime scenes. In this paper, we propose a framework to alleviate the accuracy decline when object detection is taken to adverse conditions by using image translation method. We propose to utilize style translation based StyleMix method to acquire pairs of day time image and nighttime image as training data for following nighttime to daytime image translation. To alleviate the detail corruptions caused by Generative Adversarial Networks (GANs), we propose to utilize Kernel Prediction Network (KPN) based method to refine the nighttime to daytime image translation. The KPN network is trained with object detection task together to adapt the trained daytime model to nighttime vehicle detection directly. Experiments on vehicle detection verified the accuracy and effectiveness of the proposed approach. Lan Fu, Hongkai Yu, Felix Juefei-Xu, Qing Guo 0005, Song Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Traffic Accident Detection via Self-Supervised Consistency Learning in Driving ScenariosabstractWith the rapid progress of autonomous driving and advanced driver assistance systems, there are growing efforts to promote their safety in natural driving scenarios, especially for the detection of the traffic accidents. However, because of the dynamic camera motion and complex scene in driving situations, traffic accident detection is still challenging. In this work, we aim to give the ability of Traffic Accident Detection for driving systems by proposing a Self-Supervised Consistency learning framework, termed as SSC-TAD, that involves the appearance, motion, and context consistency learning. The key formulation is to find the inconsistency of video frames, object locations and the spatial relation structure of scene temporally between different frames captured by the dashcam videos. Within this field, different from the previous works which concentrate on predicting the future object locations or frames, we further focus on predicting the visual scene context in driving scenarios and detecting the traffic accident by considering the temporal frame consistency, temporal object location consistency, and the spatial-temporal relation consistency of road participants. In this work, this formulation is fulfilled by a collaborative multi-task consistency learning network and the visual scene context feature is represented by a graph convolution network. The superiority to the state-of-the-art is verified by exhaustive evaluations on two large scale datasets, i.e., the AnAn Accident Detection (A3D) dataset and DADA-2000 dataset collected recently. Jianwu Fang, Jiahuan Qiao, Hongkai Yu, Jianru Xue |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | DADA: Driver Attention Prediction in Driving Accident ScenariosabstractDriver attention prediction is becoming an essential research problem in human-like driving systems. This work makes an attempt to predict thedriverattention indrivingaccident scenarios (DADA). However, challenges tread on the heels of that because of the dynamic traffic scene, intricate and imbalanced accident categories. In this work, we design a semantic context induced attentive fusion network (SCAFNet). We first segment the RGB video frames into the images with different semantic regions (i.e., semantic images), where each region denotes one semantic category of the scene (e.g., road, trees, etc.), and learn the spatio-temporal features of RGB frames and semantic images in two parallel paths simultaneously. Then, the learned features are fused by an attentive fusion network to find the semantic-induced scene variation in driver attention prediction. The contributions are three folds. 1) With the semantic images, we introduce their semantic context features and verify the manifest promotion effect for helping the driver attention prediction, where the semantic context features are modeled by a graph convolution network (GCN) on semantic images; 2) We fuse the semantic context features of semantic images and the features of RGB frames in an attentive strategy, and the fused details are transferred over frames by a convolutional LSTM module to obtain the attention map of each video frame with the consideration of historical scene variation in driving situations; 3) The superiority of the proposed method is evaluated on our previously collected dataset (named as DADA-2000) and two other challenging datasets with state-of-the-art methods. Jianwu Fang, Dingxin Yan, Jiahuan Qiao, Jianru Xue, Hongkai Yu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Deep Domain Adaptation Based Multi-Spectral Salient Object DetectionabstractSalient Object Detection (SOD) plays an important role in many image-related multimedia applications. Although there are many existing research works about the salient object detection in traditional RGB (visible-light spectrum) images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we mainly model this research problem as a deep learning based domain adaptation from the traditional RGB image data (source domain) to the multi-spectral data (target domain), and an adversarial deep domain adaptation model is proposed. We first collect and will publicize a large multi-spectral dataset, RGBN-SOD dataset, including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. Intensive experimental results show the effectiveness and accuracy of the proposed deep domain adaptation for the multi-spectral SOD. Besides, due to the absence of research on the field of multi-spectral co-saliency detection, we also collect 200 synchronized RGB and NIR image pairs in addition to explore the multi-spectral co-saliency detection. Shaoyue Song, Zhenjiang Miao, Hongkai Yu, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
IEEE Trans. Multim. | 3 |
| 2021 | Auto-Exposure Fusion for Single-Image Shadow RemovalabstractShadow removal is still a challenging task due to its inherent background-dependent1and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this task by formulating it as an exposure fusion problem to address the challenges. Intuitively, we first estimate multiple over-exposure images w.r.t. the input image to let the shadow regions in these images have the same color with shadow-free areas in the input image. Then, we fuse the original input with the over-exposure images to generate the final shadow-free counterpart. Nevertheless, the spatial-variant property of the shadow requires the fusion to be sufficiently ‘smart’, that is, it should automatically select proper over-exposure pixels from different images to make the final output natural. To address this challenge, we propose the shadow-aware FusionNet that takes the shadow image as input to generate fusion weight maps across all the over-exposure images. Moreover, we propose the boundary-aware RefineNet to eliminate the remaining shadow trace further. We conduct extensive experiments on the ISTD, ISTD+, and SRD datasets to validate our method’s effectiveness and show better performance in shadow regions and comparable performance in non-shadow regions over the state-of-the-art methods. We release the code in https://github.com/tsingqguo/exposure-fusion-shadow-removal. Lan Fu, Changqing Zhou, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 5 |
| 2021 | JPGNet: Joint Predictive Filtering and Generative Network for Image InpaintingabstractImage inpainting aims to restore the missing regions of corrupted images and make the recovery result identical to the originally complete image, which is different from the common generative task emphasizing the naturalness or realism of generated images. Nevertheless, existing works usually regard it as a pure generation problem and employ cutting-edge deep generative techniques to address it. The generative networks can fill the main missing parts with realistic contents but usually distort the local structures or introduce obvious artifacts. In this paper, for the first time, we formulate image inpainting as a mix of two problems, i.e., predictive filtering and deep generation. Predictive filtering is good at preserving local structures and removing artifacts but falls short to complete the large missing regions. The deep generative network can fill the numerous missing pixels based on the understanding of the whole scene but hardly restores the details identical to the original ones. To make use of their respective advantages, we propose the joint predictive filtering and generative network (JPGNet) that contains three branches: predictive filtering & uncertainty network (PFUNet), deep generative network, and uncertainty-aware fusion network (UAFNet). The PFUNet can adaptively predict pixel-wise kernels for filtering-based inpainting according to the input image and output an uncertainty map. This map indicates the pixels should be processed by filtering or generative networks, which is further fed to the UAFNet for a smart combination between filtering and generative results. Note that, our method as a novel framework for the image inpainting problem can benefit any existing generation-based methods. We validate our method on three public datasets, i.e., Dunhuang, Places2, and CelebA, and demonstrate that our method can enhance three state-of-the-art generative methods (i.e., StructFlow, EdgeConnect, and RFRNet) significantly with slightly extra time costs. We have released the code at https://github.com/tsingqguo/jpgnet. Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Yang Liu 0003, Song Wang 0002 |
ACM Multimedia | 4 |
| 2021 | Guest editorial: Graph learning for computer visionabstractMany fields in the real world involve a lot of structured data, such as social networks, transportation networks, communication networks etc., and their structures carry important information about the characteristics of the data. However, how to use its structural information to analyse and process the data efficiently has caused continuous research in the field. A graph provides an important means for dealing with structured data. It can describe the geometric structure of data intuitively and flexibly, especially in the representation of spatial irregular data. Graph learning refers to machine learning on graphs, which mainly utilises machine learning algorithms to extract the relevant features of graphs. In recent years, combined with specific applications, researchers have conducted in-depth research on graph learning and proposed various approaches. This Special Issue aims to introduce the latest studies in graph learning for computer vision and proposes new theories and approaches to solve the existing problems. It received a number of submissions from researchers in the field, which all went through a rigorous review process. After several rounds of review, six papers were accepted. These papers cover a variety of fields, such as medicine, remote sensing, and data mining. Specific tasks include image segmentation, knowledge graph reasoning, clustering, and image classification. These accepted papers are mainly divided into two categories. The first category covers the graph learning method guided by optimisation, which obtains the graph structure by establishing a clear model and solving the corresponding optimisation problem. The second category is the deep learning-oriented graph learning method, which combines convolutional neural network and graph neural network to construct the model. In the first paper of the Special Issue by Wang et al. entitled An Enhanced 3D U-Net with Graph-based Refining for Segmentation of Gastrointestinal Stromal Tumours, the authors propose a segmentation algorithm using an improved 3D U-Net to segment gastrointestinal stromal tumours. To enhance information transmission, multiple skip connections are attached into same size feature maps between an encoder and a decoder. Due to difficulties in tumour labelling and other reasons, the author transforms the small intestinal segmentation model into a gastrointestinal stromal tumour segmentation model. Since fully convolutional networks typically suffer from inaccuracies around the boundaries of small structures, the graph neural network is introduced to refine segmentation results. Experiments demonstrate that the proposed method presents superior performance over traditional U-Net. The second paper by Ma et al. entitled Hybrid Attention Mechanism for Few-Shot Relational Learning of Knowledge Graphs, develops a few-shot relationship learning framework. The authors first design an entity-enhanced encoder with weak attention networks and self-attention mechanisms to explore the influence of different levels for source entities. The local graph structure is then utilised to enhance the embedding of the source entity by combining explicit and implicit features. Finally, the model parameters are optimised to infer real entities in the candidate set of similar entities obtained by a loop-processing matching processor. The authors provide extensive experiments and confirm the excellent accuracy of the proposed model. The third paper by Zhao et al. entitled Incremental Multi-View Correlated Feature Learning Based on Non-Negative Matrix Factorization, studies multi-view data. The authors present an incremental multi-view correlated feature learning approach based on non-negative matrix factorization to analyse the uncorrelated items in each view. The algorithm separates uncorrelateditems across views and constructs incremental joint learning with uncorrelated and correlated features to study the common features for multi-view data. Subsequently, the authors design an incremental objective function and derive an effective updating scheme. The proposed method is proved to converge effectively, and its complexity is discussed. They evaluate the proposed solution on real-world datasets and report excellent performance in comparison with the existing state-of-the-art solutions. The fourth paper by Hu et al. entitled Complete/Incomplete Multi-view Subspace Clustering via Soft Block-Diagonal-Induced Regularizer, concentrates on complete and incomplete multi-view clustering problems. The proposed method adopts the self-representation model to individually construct the similarity graphs for each view. To fuse a shared affinity matrix for all views, the authors design the soft block-diagonal-induced regulariser to encourage the generation of a matrix with K diagonal blocks. Considering the incomplete multi-view data, the proposed method effectively utilises some indicator matrices to accurately mark the missing instances in each view. The authors analyse the complexity and convergence of the proposed method on four public datasets and demonstrate that it is better than the most advanced complete/incomplete clustering methods. The fifth paper by Gong et al. entitled Few-shot Learning with Relation Propagation and Constraint, aims to extract valuable information of pair-wise correlation between sparse training samples. The authors state that transductive relation propagation simply propagates the pair-wise relation without relation constraints. Thus, the paper develops a constrained relation–propagation network to capture the accurate relation so as to generate discriminative relational representations. To constrain the pair-wise relation, the proposed method introduces a relation constraint module to regularise the distilled relations between samples, which helps to calibrate the propagated correlation information. Extensive experiments conducted on several benchmark datasets indicate that the proposed method achieves remarkable performance compared to few-shot learning methods. The last paper by Guo et al. entitled CNN-Combined Graph Residual Network with Multilevel Feature Fusion for Hyperspectral Image Classification introduces graph convolutional networks to obtain more superpixel-level features with a topological structure. This paper develops an effective CNN-combined graph residual network with a multilevel feature fusion strategy. The main idea is to learn superpixeltopological information by using the graph residual network and pixel information by using the convolutional neural network. The strategy can adequately leverage the superpixel level and pixel-level features and capture the class boundary features, which further enhances the generalisation performance. Experiments report highly competitive performance in comparison to existing hyperspectral image classification approaches. The papers selected in this Special Issue highlight the extensive study of graph learning in computer vision. We hope that these papers can promote the theoretical study of graph learning as well as provide new ideas for more researchers who are committed to graph learning. Qi Wang 0009, Hongkai Yu, Song Wang 0002, Jianzhe Lin |
IET Comput. Vis. | 2 |
| 2020 | Multi-Spectral Salient Object Detection by Adversarial Domain AdaptationabstractAlthough there are many existing research works about the salient object detection (SOD) in RGB images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we first collect and will publicize a large multi-spectral dataset including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. We model this research problem as an adversarial domain adaptation from the existing RGB image dataset (source domain) to the collected multi-spectral dataset (target domain). Experimental results show the effectiveness and accuracy of the proposed adversarial domain adaptation for the multi-spectral SOD. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Jianwu Fang, Cong Ma 0004, Song Wang 0002 |
AAAI | 2 |
| 2020 | Weakly supervised easy-to-hard learning for object detection in image sequences
Hongkai Yu, Dazhou Guo, Zhipeng Yan, Lan Fu, Jeff P. Simmons, Craig Przybyla, Song Wang 0002 |
Neurocomputing | 1 |
| 2020 | Vehicle re-identification in tunnel scenes via synergistically cascade forests
Rixing Zhu, Jianwu Fang, Qi Wang 0009, Hongke Xu, Jianru Xue, Hongkai Yu |
Neurocomputing | 7 |
| 2020 | Degraded Image Semantic Segmentation With Dense-Gram NetworksabstractDegraded image semantic segmentation is of great importance in autonomous driving, highway navigation systems, and many other safety-related applications and it was not systematically studied before. In general, image degradations increase the difficulty of semantic segmentation, usually leading to decreased semantic segmentation accuracy. Therefore, performance on the underlying clean images can be treated as an upper bound of degraded image semantic segmentation. While the use of supervised deep learning has substantially improved the state of the art of semantic image segmentation, the gap between the feature distribution learned using the clean images and the feature distribution learned using the degraded images poses a major obstacle in improving the degraded image semantic segmentation performance. The conventional strategies for reducing the gap include: 1) Adding image-restoration based pre-processing modules; 2) Using both clean and the degraded images for training; 3) Fine-tuning the network pre-trained on the clean image. In this paper, we propose a novel Dense-Gram Network to more effectively reduce the gap than the conventional strategies and segment degraded images. Extensive experiments demonstrate that the proposed Dense-Gram Network yields stateof-the-art semantic segmentation performance on degraded images synthesized using PASCAL VOC 2012, SUNRGBD, CamVid, and CityScapes datasets. Dazhou Guo, Yanting Pei, Hongkai Yu, Song Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2020 | A New Method and Benchmark for Detecting Co-Saliency Within a Single ImageabstractRecently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision and multimedia communities. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new benchmark dataset of 664 color images (two subsets) for within-image co-saliency detection. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms. The experimental results also show that the proposed method can be applied to detect the repetitive patterns in a single image and detect the co-saliency in multiple images. Hongkai Yu, Jianwu Fang, Hao Guo 0002, Song Wang 0002 |
IEEE Trans. Multim. | 1 |
| 2019 | Visual Attention Consistency Under Image Transforms for Multi-Label Image ClassificationabstractHuman visual perception shows good consistency for many multi-label image classification tasks under certain spatial transforms, such as scaling, rotation, flipping and translation. This has motivated the data augmentation strategy widely used in CNN classifier training -- transformed images are included for training by assuming the same class labels as their original images. In this paper, we further propose the assumption of perceptual consistency of visual attention regions for classification under such transforms, i.e., the attention region for a classification follows the same transform if the input image is spatially transformed. While the attention regions of CNN classifiers can be derived as an attention heatmap in middle layers of the network, we find that their consistency under many transforms are not preserved. To address this problem, we propose a two-branch network with an original image and its transformed image as inputs and introduce a new attention consistency loss that measures the attention heatmap consistency between two branches. This new loss is then combined with multi-label image classification loss for network training. Experiments on three datasets verify the superiority of the proposed network by achieving new state-of-the-art classification performance. Hao Guo 0002, Xiaochuan Fan, Hongkai Yu, Song Wang 0002 |
CVPR | 4 |
| 2019 | RNN-based default logic for route planning in urban environments
Jiangtao Kong, Jian Huang 0010, Hongkai Yu, Hanqiang Deng, Jianxing Gong, Hao Chen 0099 |
Neurocomputing | 3 |
| 2019 | An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002 |
Neurocomputing | 2 |
| 2019 | Domain Adaptation for Convolutional Neural Networks-Based Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification plays an important role in the field of earth observation. With the rapid development of the RS techniques, a large number of RS scene images are available. As manually labeling large-scale RS scene images is both labor and time consuming, when a new unlabeled data set is obtained, how to use the existing labeled data sets to classify the new unlabeled images is an important research direction. Different RS scene image data sets may be taken from different type of sensors, and the images may vary from imaging modalities, spatial resolutions, and image scales, so the distribution discrepancy exists among different image data sets. As a result, simply applying convolutional neural networks (CNN) trained on source domain cannot accurately classify the images on target domain. Domain adaptation (DA) can be helpful to solve this problem. In this letter, we design a subspace alignment (SA) and CNN-based framework to solve the DA problem in RS scene image classification. A new SA layer is proposed and added into CNN models for DA, which could align the source and target domains in some feature subspace. Fine-tuning the modified CNN model with the added SA layer makes the CNN model adapt to the aligned feature subspace and helps to relieve the domain distribution discrepancy. The experiments conducted on two public data sets show that adding the SA layer into CNN improves the scene classification on the target domain. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Qiang Zhang 0030, Yuewei Lin, Song Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Small Object Sensitive Segmentation of Urban Street Scene With Spatial Adjacency Between Object ClassesabstractRecent advancements in deep learning have shown exciting promise in the urban street scene segmentation. However, many objects, such as poles and sign symbols, are relatively small and they usually cannot be accurately segmented since the larger objects usually contribute more to the segmentation loss. In this paper, we propose a new boundary-based metric that measures the level of spatial adjacency between each pair of object classes and find that this metric is robust against object size induced biases. We develop a new method to enforce this metric into the segmentation loss. We propose a network, which starts with a segmentation network, followed by a new encoder to compute the proposed boundary-based metric, and then trains this network in an end-to-end fashion. In deployment, we only use the trained segmentation network, without the encoder, to segment new unseen images. Experimentally, we evaluate the proposed method using CamVid and CityScapes datasets and achieve a favorable overall performance improvement and a substantial improvement in segmenting small objects. Dazhou Guo, Ligeng Zhu, Hongkai Yu, Song Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2018 | Co-Saliency Detection Within a Single ImageabstractRecently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision community. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the original image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new dataset of 364 color images with within-image cosaliency. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms. Hongkai Yu, Jianwu Fang, Hao Guo 0002, Wei Feng 0005, Song Wang 0002 |
AAAI | 1 |
| 2018 | Multiple human tracking in wearable camera videos with informationless intervals
Hongkai Yu, Haozhou Yu, Hao Guo 0002, Jeff P. Simmons, Qin Zou 0001, Wei Feng 0005, Song Wang 0002 |
Pattern Recognit. Lett. | 1 |
| 2017 | Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable CamerasabstractIn this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by optical flow. However, optical flow is view-variant - the same person's optical flow in different videos can be very different due to view angle change. In this paper, we attempt to utilize 3D human-skeleton sequences to learn a model that can extract view-invariant motion features from optical flows in different views. For this purpose, we use 3D Mocap database to build a synthetic optical flow dataset and train a Triplet Network (TN) consisting of three sub-networks: two for optical flow sequences from different views and one for the underlying 3D Mocap skeleton sequence. Finally, sub-networks for optical flows are used to extract view-invariant features for CVPI. Experimental results show that, using only the motion information, the proposed method can achieve comparable performance with the state-of-the-art methods. Further combination of the proposed method with an appearance-based method achieves new state-of-the-art performance. Xiaochuan Fan, Yuewei Lin, Hao Guo 0002, Hongkai Yu, Dazhou Guo, Song Wang 0002 |
ICCV | 5 |
| 2017 | Loosecut: Interactive image segmentation with loosely bounded boxesabstractOne popular approach to interactively segment an object of interest from an image is to annotate a bounding box that covers the object, followed by a binary labeling. However, the existing algorithms for such interactive image segmentation prefer a bounding box that tightly encloses the object. This increases the annotation burden, and prevents these algorithms from utilizing automatically detected bounding boxes. In this paper, we develop a new LooseCut algorithm that can handle cases where the bounding box only loosely covers the object. We propose a new Markov Random Fields (MRF) model for segmentation with loosely bounded boxes, including an additional energy term to encourage consistent labeling of similar-appearance pixels and a global similarity constraint to better distinguish the foreground and background. This MRF model is then solved by an iterated max-flow algorithm. We evaluate LooseCut in three public image datasets, and show its better performance against several state-of-the-art methods when increasing the bounding-box size. Hongkai Yu, Youjie Zhou, Hui Qian 0001, Min Xian, Song Wang 0002 |
ICIP | 1 |
| 2017 | Feature sampling strategies for action recognitionabstractAlthough dense local spatial-temporal features with bag-of-features representation achieve state-of-the-art performance for action recognition, the huge feature number and feature size prevent current methods from scaling up to real size problems. In this work, we investigate different types of feature sampling strategies for action recognition, namely dense sampling, uniformly random sampling and selective sampling. We propose two effective selective sampling methods using object proposal techniques. Experiments conducted on a large video dataset show that we are able to achieve better average recognition accuracy using 25% less features, through one of the proposed selective sampling methods, and even maintain comparable accuracy while discarding 70% features. Youjie Zhou, Hongkai Yu, Song Wang 0002 |
ICIP | 2 |
| 2016 | Groupwise Tracking of Crowded Similar-Appearance Targets from Low-Continuity Image SequencesabstractAutomatic tracking of large-scale crowded targets are of particular importance in many applications, such as crowded people/vehicle tracking in video surveillance, fiber tracking in materials science, and cell tracking in biomedical imaging. This problem becomes very challenging when the targets show similar appearance and the interslice/ inter-frame continuity is low due to sparse sampling, camera motion and target occlusion. The main challenge comes from the step of association which aims at matching the predictions and the observations of the multiple targets. In this paper we propose a new groupwise method to explore the target group information and employ the within-group correlations for association and tracking. In particular, the within-group association is modeled by a nonrigid 2D Thin-Plate transform and a sequence of group shrinking, group growing and group merging operations are then developed to refine the composition of each group. We apply the proposed method to track large-scale fibers from microscopy material images and compare its performance against several other multi-target tracking methods. We also apply the proposed method to track crowded people from videos with poor inter-frame continuity. Hongkai Yu, Youjie Zhou, Jeff P. Simmons, Craig Przybyla, Yuewei Lin, Xiaochuan Fan, Yang Mi, Song Wang 0002 |
CVPR | 1 |
| 2016 | Large-Scale Fiber Tracking Through Sparsely Sampled Image Sequences of Composite MaterialsabstractFast and accurate characterization of fiber micro-structures plays a central role for material scientists to analyze physical properties of continuous fiber reinforced composite materials. In materials science, this is usually achieved by continuously cross-sectioning a 3D material sample for a sequence of 2D microscopic images, followed by a fiber detection/tracking algorithm through the obtained image sequence. To speed up this process and be able to handle larger size material samples, this paper proposes sparse sampling with larger inter-slice distance in cross sectioning and develops a new algorithm that can robustly track large-scale fibers from such a sparsely sampled image sequence. In particular, the problem is formulated as multi-target tracking, and the Kalman filters are applied to track each fiber along the image sequence. One main challenge in this tracking process is to correctly associate each fiber to its observation given that: fiber observations are of large scale, crowded, and show very similar appearances in a 2D slice and there may be a large gap between the predicted location of a fiber and its observation in the sparse sampling. To address this challenge, a novel group-wise association algorithm is developed by leveraging the fact that fibers are implanted in bundles and the fibers in the same bundle are highly correlated through the image sequence. In experiments, the proposed algorithm is tested on three tiles of 100-slice S200 material samples and the tracking performance is evaluated using 1136 human annotated ground-truth fiber tracks. Both quantitative and qualitative results show that the proposed algorithm clearly outperforms the state-of-the-art multiple-target tracking algorithms on sparsely sampled image sequences. Youjie Zhou, Hongkai Yu, Jeff P. Simmons, Craig Przybyla, Song Wang 0002 |
IEEE Trans. Image Process. | 2 |
| 2015 | Co-Interest Person Detection from Multiple Wearable Camera VideosabstractWearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos taken by multiple wearable cameras. Our basic idea is to exploit the motion patterns of people and use them to correlate the persons across different videos, instead of performing appearance-based matching as in traditional video co-segmentation/localization. This way, we can identify CIP even if a group of people with similar appearance are present in the view. More specifically, we detect a set of persons on each frame as the candidates of the CIP and then build a Conditional Random Field (CRF) model to select the one with consistent motion patterns in different videos and high spacial-temporal consistency in each video. We collect three sets of wearable-camera videos for testing the proposed algorithm. All the involved people have similar appearances in the collected videos and the experiments demonstrate the effectiveness of the proposed algorithm. Yuewei Lin, Kareem Abdelfatah, Youjie Zhou, Xiaochuan Fan, Hongkai Yu, Hui Qian 0001, Song Wang 0002 |
ICCV | 5 |
| 2014 | Unsupervised co-segmentation based on a new global GMM constraint in MRFabstractThis paper proposes a new Markov Random Fields (MRF) optimization model for co-segmentation. The co-saliency model is incorporated into our model to make it fully unsupervised and work well for images with similar backgrounds. The Gaussian Mixture Model (GMM) based dissimilarity between foregrounds in each image and the common objects in the set is involved as a new global constraint (i.e., energy term) in our model. Finally, we introduce an alternative approximation to represent the energy function, which could be minimized by Graph Cuts iteratively. The experimental results on two datasets show that our algorithm achieves better or comparable accuracy when comparing with state-of-the-art algorithms. Hongkai Yu, Min Xian, Xiaojun Qi 0001 |
ICIP | 1 |
| 2014 | Unsupervised cosegmentation based on superpixel matching and FastgrabcutabstractThis paper proposes a novel unsupervised cosegmentation method which automatically segments the common objects in multiple images. It designs a simple superpixel matching algorithm to explore the inter-image similarity. It then constructs the object mask for each image using the matched superpixels. This object mask is a convex hull potentially containing the common objects and some backgrounds. Finally, it applies a new FastGrabCut algorithm, an improved GrabCut algorithm, on the object mask to simultaneously improve the segmentation efficiency and maintain the segmentation accuracy. This FastGrabcut algorithm introduces preliminary classification to accelerate convergence. It uses Expectation Maximization (EM) algorithm to estimate optimal Gaussian Mixture Model(GMM) parameters of the object and background and then applies Graph Cuts to minimize the energy function for each image. Experimental results on the iCoseg dataset demonstrate the accuracy and robustness of our cosegmentation method. Hongkai Yu, Xiaojun Qi 0001 |
ICME | 1 |