VLDB 2026 Research / reviewers in the wild / expert
Gaoang Wang
dblp:176/7523
· DBLP profile ↗
76ranked-venue papers
8as first author
68since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 8 first-author · 50 since 2021Artificial intelligence and machine learning · 34 · 4 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IGFuse: Interactive 3D Gaussian Scene Reconstruction via Multi-Scans FusionabstractReconstructing complete and interactive 3D scenes remains a fundamental challenge in computer vision and robotics, particularly due to persistent object occlusions and limited sensor coverage. Even multi-view observations from a single scene scan often fail to capture the full structural details. Existing approaches typically rely on multi-stage pipelines—such as segmentation, background completion, and inpainting—or require per-object dense scanning, both of which are error-prone, and not easily scalable. We propose IGFuse, a novel framework that reconstructs interactive Gaussian scene by fusing observations from multiple scans, where natural object rearrangement between captures reveal previously occluded regions. Our method constructs segmentation-aware Gaussian fields and enforces bi-directional photometric and semantic consistency across scans. To handle spatial misalignments, we introduce a pseudo-intermediate scene state for symmetric alignment, alongside collaborative co-pruning strategies to refine geometry. IGFuse enables high-fidelity rendering and object-level scene manipulation without dense observations or complex pipelines. Extensive experiments validate the framework’s strong generalization to novel scene configurations, demonstrating its effectiveness for real-world 3D reconstruction and real-to-simulation transfer. Zesheng Li, Haonan Zhou, Xuexiang Wen, Zhizhong Su, Gaoang Wang |
AAAI | 8 |
| 2026 | Understanding Dynamic Scenes in Ego Centric 4D Point CloudsabstractUnderstanding dynamic 4D scenes from an egocentric perspective—modeling changes in 3D spatial structure over time—is crucial for human–machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets contain dynamic scenes, they lack unified 4D annotations and task-driven evaluation protocols for fine-grained spatio-temporal reasoning, especially on motion of objects and human, together with their interactions. To address this gap, we introduce EgoDynamic4D, a novel QA benchmark on highly dynamic scenes, comprising RGB-D video, camera poses, globally unique instance masks, and 4D bounding boxes. We construct 927K QA pairs accompanied by explicit Chain-of-Thought (CoT), enabling verifiable, step-by-step spatio-temporal reasoning. We design 12 dynamic QA tasks covering agent motion, human–object interaction, trajectory prediction, relation understanding, and temporal–causal reasoning, with fine-grained, multidimensional metrics. To tackle these tasks, we propose an end-to-end spatio-temporal reasoning framework that unifies dynamic and static scene information, using instance-aware feature encoding, time and camera encoding, and spatially adaptive down-sampling to compress large 4D scenes into token sequences manageable by LLMs. Experiments on EgoDynamic4D show that our method consistently outperforms baselines, validating the effectiveness of multimodal temporal modeling for egocentric dynamic scene understanding. Shengyu Hao, Bocheng Hu, Hongwei Wang 0001, Gaoang Wang |
AAAI | 5 |
| 2026 | X-MoGen: Unified Motion Generation Across Humans and AnimalsabstractText-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representation and improved generalization. However, morphological differences across species remain a key challenge, often compromising motion plausibility. To address this, we propose X-MoGen, the first unified framework for cross-species text-driven motion generation covering both humans and animals. X-MoGen adopts a two-stage architecture. First, a conditional graph variational autoencoder learns canonical T-pose priors, while an autoencoder encodes motion into a shared latent space regularized by morphological loss. In the second stage, we perform masked motion modeling to generate motion embeddings conditioned on textual descriptions. During training, a morphological consistency module is employed to promote skeletal plausibility across species. To support unified modeling, we construct UniMo4D, a large-scale dataset of 115 species and 119k motion sequences, which integrates human and animal motions under a shared skeletal topology for joint training. Extensive experiments on UniMo4D demonstrate that X-MoGen outperforms state-of-the-art methods on both seen and unseen species. Kai Ruan, Liyang Qian, Guo Zhi Zhi, Gaoang Wang |
AAAI | 6 |
| 2026 | Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field
Wenhao Hu 0002, Wenhao Chai, Shengyu Hao, Xiaotong Cui, Xuexiang Wen, Jenq-Neng Hwang, Gaoang Wang |
Int. J. Comput. Vis. | 7 |
| 2026 | RDG-GS: Relative Depth Guidance with Gaussian Splatting for Real-time Sparse-View 3D Rendering
Chenlu Zhan, Yufei Zhang 0015, Gaoang Wang, Hongwei Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | OR-DARE: A deliberative framework for solving operations research problems via collaborative debate and iterative refinement
Jiawu Zhang, Der-Horng Lee, Gaoang Wang |
Inf. Sci. | 5 |
| 2026 | MovieChat+: Question-Aware Sparse Memory for Long Video Question AnsweringabstractRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific vision tasks. Yet, existing methods either employ complex spatial-temporal modules or rely heavily on additional perception models to extract temporal features for video understanding, performing well only on short videos. For long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges. Leveraging the hierarchical memory structure of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination, we propose MovieChat within a training-free memory consolidation mechanism to overcome these challenges, which transfers dense frames from short-term memory into sparse tokens in long-term memory by temporally merging adjacent frames. We lift pre-trained large multi-modal models for understanding long videos without additional trainable modules, employing a zero-shot approach. Additionally, in our new version, MovieChat+, we design an enhanced training-free vision-question matching-based memory consolidation mechanism to better anchor predictions to relevant visual content. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1 K benchmark with 1 K long video, 2 K temporal grounding labels, and 14 K manual annotations. Enxin Song, Wenhao Chai, Tian Ye 0001, Jenq-Neng Hwang, Xi Li 0001, Gaoang Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | CWPS: Efficient Channel-Wise Parameter Sharing for Knowledge TransferabstractKnowledge transfer aims to apply existing knowledge to different tasks or new data, and it has extensive applications in multi-domain and Multi-Task Learning. The key to this task is quickly identifying a fine-grained object for knowledge sharing and efficiently transferring knowledge. Current methods, such as fine-tuning, layer-wise parameter sharing, and task-specific adapters, only offer coarse-grained sharing solutions and struggle to effectively search for shared parameters, thus hindering the performance and efficiency of knowledge transfer. To address these issues, we propose Channel-Wise Parameter Sharing (CWPS), a novel fine-grained parameter-sharing method for knowledge transfer, which is efficient for parameter sharing, comprehensive, and plug-and-play. For the coarse-grained problem, we first achieve fine-grained parameter sharing by refining the granularity of shared parameters from the level of layers to the level of neurons. The knowledge learned from previous tasks can be utilized through the explicit composition of the model neurons. Besides, we promote an effective search strategy to minimize computational costs, simplifying the selection of shared weights. In addition, our CWPS has strong composability and generalization ability, which theoretically can be applied to any network consisting of linear and convolution layers. We introduce several datasets in both Incremental Learning and Multi-Task Learning scenarios. Our method has achieved state-of-the-art precision-to-parameter ratio performance with various backbones, demonstrating its efficiency and versatility. Mingxuan Cui, Xuewei Li 0003, Cunzheng Wang, Gaoang Wang, Chenyi Zhuang, Jinjie Gu, Xiubo Liang, Xi Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Hi-LSplat: Hierarchical 3D Language Gaussian SplattingabstractModeling 3D language fields with Gaussian Splatting for open-ended language queries has recently garnered increasing attention. However, recent 3DGS-based models leverage view-dependent 2D foundation models to refine 3D semantics but lack a unified 3D representation, leading to view inconsistencies. Additionally, inherent open-vocabulary challenges cause inconsistencies in object and relational descriptions, impeding hierarchical semantic understanding. In this paper, we propose Hi-LSplat, a view-consistent Hierarchical Language Gaussian Splatting work for 3D open-vocabulary querying. To achieve view-consistent 3D hierarchical semantics, we first lift 2D features to 3D features by constructing a 3D hierarchical semantic tree with layered instance clustering, which addresses the view inconsistency issue caused by 2D semantic features. Besides, we introduce instance-wise and part-wise contrastive losses to capture all-sided hierarchical semantic representations. Notably, we construct two hierarchical semantic datasets to better assess the model's ability to distinguish different semantic levels. Extensive experiments highlight our method's superiority in 3D open-vocabulary segmentation and localization. Its strong performance on hierarchical semantic datasets underscores its ability to capture complex hierarchical semantics within 3D scenes. Chenlu Zhan, Yufei Zhang 0015, Gaoang Wang, Hongwei Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action RecognitionabstractIn this paper, we propose a novel Temporal Sequence-Aware-Model (TSAM) for few-shot action recognition (FSAR), which incorporates a sequential perceiver adapter into the pre-training framework, to integrate both the spatial information and the sequential temporal dynamics into the feature embeddings. Different from the existing fine-tuning approaches that capture temporal information by exploring the relationships among all the frames, our perceiver-based adapter recurrently captures the sequential dynamics alongside the timeline, which could perceive the frame order change. To obtain the discriminative representations for each class, we extend a textual corpus for each class derived from the large language models (LLMs) and enrich the visual prototypes by integrating the contextual semantic information. Besides, We introduce an unbalanced optimal transport strategy for feature matching that mitigates the impact of class-unrelated features, thereby facilitating more effective decision-making. Experimental results on five FSAR datasets demonstrate that our method establishes a new benchmark, outperforming the second-best competitors. Bozheng Li, Mushui Liu, Gaoang Wang |
AAAI | 3 |
| 2025 | AniMo: Species-Aware Model for Text-Driven Animal Motion GenerationabstractText-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal ecology, and biomechanics. Animal motion modeling presents unique challenges due to species diversity, varied morphological structures, and different behavioral patterns in response to similar textual descriptions. To address these challenges, we propose AniMo for text-driven animal motion generation. AniMo consists of two stages: motion tokenization and text-to-motion generation. In the motion tokenization stage, we encode motions using a joint-aware spatiotemporal encoder with species-aware feature modulation, enabling the model to adapt to diverse skeletal structures across species. In the text-to-motion generation stage, we employ masked modeling to jointly learn the mapping from textual descriptions to motion tokens. Additionally, we introduce AniMo4D, a large-scale dataset containing 78,149 motion sequences and 185,435 textual descriptions across 114 animal species. Experimental results show that AniMo achieves superior performance on both the AniMo4D and AnimalML3D datasets, effectively capturing diverse morphological structures and behavioral patterns across animal species. Kai Ruan, Gaoang Wang |
CVPR | 4 |
| 2025 | Adaptive Graph Pruning for Multi-Agent CommunicationabstractLarge Language Model (LLM) based multi-agent systems have shown impressive performance across various fields of tasks, further enhanced through collaborative debate and communication using carefully designed communication topologies. However, existing methods typically employ a fixed number of agents or static communication structures, requiring manual pre-definition, and thus struggle to dynamically adapt the number of agents and topology simultaneously to varying task complexities. In this paper, we propose Adaptive Graph Pruning (AGP), a novel task-adaptive multi-agent collaboration framework that jointly optimizes agent quantity (hard-pruning) and communication topology (soft-pruning). Specifically, our method employs a two-stage training strategy: firstly, independently training soft-pruning networks for different agent quantities to determine optimal agent-quantity-specific complete graphs and positional masks across specific tasks; and then jointly optimizing hard-pruning and soft-pruning within a maximum complete graph to dynamically configure the number of agents and their communication topologies per task. Extensive experiments demonstrate that our approach is: (1) High-performing, achieving state-of-the-art results across six benchmarks and consistently generalizes across multiple mainstream LLM architectures, with a increase in performance of 2.58% ∼ 9.84%; (2) Task-adaptive, dynamically constructing optimized communication topologies tailored to specific tasks, with an extremely high performance in all three task categories (general reasoning, mathematical reasoning, and code generation); (3) Token-economical, having fewer training steps and token consumption at the same time, with a decrease in token consumption of 90%+; and (4) Training-efficient, achieving high performance with very few training steps compared with other methods. The performance will surpass the existing baselines after about ten steps of training under six benchmarks. Our code and demos are publicly available at https://resurgamm.github.io/AGP/. Boyi Li 0002, Zhonghan Zhao, Der-Horng Lee, Gaoang Wang |
ECAI | 4 |
| 2025 | RAPID: Recognition of Any-Possible DrIver Distraction via Multi-view Pose Generation ModelsabstractDriver distraction remains a pressing traffic safety issue. Drivers are often careless with their distraction behaviours, which may cause serious traffic accidents. However, current Driver Monitoring Systems (DMS) cannot be put into practical application well, which tend to have high latency, lack precision, and are unable to cover all distraction behaviours. In this paper, we assume driver distraction to be a One-Class Classification (OCC) problem and build an unsupervised learning baseline based on denoising diffusion probabilistic models (DDPM) called RAPID which aggregates future patterns generated by the diffusion process to detect distraction, considering the diversity of normal and abnormal situations. Besides, we propose a skeleton-based synchronized multi-view dataset with diverse distraction behaviours called sktDD (skeleton-based Driver Distraction dataset) to improve on existing datasets. RAPID facilitates a frame-level (0.03 second) and undefined prediction with AUC score beyond State-of-the-Art (SOTA) methods, surpassing currently typical DMS that rely on post-processing procedures and predefined actions. RAPID has the potential to bring significant advancements in the field of traffic safety, which can also be applied in future self-driving scenarios to determine whether the remote-driving operator’s current state is suitable to take over. Our dataset and code are available at https://github.com/jingyulei/rapid. Jingyu Lei, Shengyu Hao, Gaoang Wang, Der-Horng Lee |
ICASSP | 3 |
| 2025 | SCI-Gaussian: Optimizing 3D Gaussian Radiance Fields from a Snapshot Compressive ImageabstractSnapshot compressive imaging (SCI) is a compressed sensing (CS)-based high-speed imaging modality. Recent efforts have explored the underlying 3D representation from only an SCI image using neural radiance fields (NeRF), yet the training time, rendering computation cost, and reconstruction quality limitations are general issues that have limited wider adoption. This paper introduces SCI-Gaussian, the first 3D-aware SCI reconstruction based on 3D Gaussian splatting (3D-GS). This method utilizes an explicit 3D representation to achieve efficient and high-quality scene reconstruction. The motivation stems from the highly efficient representation and surprising quality of 3D-GS, despite when applied to SCI system, it encounters difficulties in generating point initialization for explicit Gaussians and accurate pose recovery from a single SCI measured image. Specifically, we effectively initialize these Gaussians through sampling a coarsely trained NeRF at various hash structures, then model the physical formation of the SCI measurement and jointly optimize Gaussians and camera trajectories with a bundle adjustment formulation during exposure time. Extensive experiments on synthetic and real-world datasets demonstrate that SCI-Gaussian outperforms the state-of-the-art (SOTA) methods, achieving comparable or better results with significantly 10× faster training and 1000× faster rendering speed than the most recent NeRF-based method. Xiaodong Wang 0026, Xin Yuan 0002, Mark D. Butala, Gaoang Wang |
ICASSP | 6 |
| 2025 | Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai, Xuexiang Wen, Tian Ye 0001, Gaoang Wang |
ICCV | 6 |
| 2025 | Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep DeploymentabstractSpiking Neural Networks (SNNs) are emerging as a brain-inspired alternative to traditional Artificial Neural Networks (ANNs), prized for their potential energy efficiency on neuromorphic hardware. Despite this, SNNs often suffer from accuracy degradation compared to ANNs and face deployment challenges due to fixed inference timesteps, which require retraining for adjustments, limiting operational flexibility. To address these issues, our work considers the spatio-temporal property inherent in SNNs, and proposes a novel distillation framework for deep SNNs that optimizes performance across full-range timesteps without specific retraining, enhancing both efficacy and deployment adaptability. We provide both theoretical analysis and empirical validations to illustrate that training guarantees the convergence of all implicit models across full-range timesteps. Experimental results on CIFAR-10, CIFAR-100, CIFAR10-DVS, and ImageNet demonstrate state-of-the-art performance among distillation-based SNNs training methods. Our code is available at https://github.com/Intelli-Chip-Lab/snn_temporal_decoupling_distillation. Chengting Yu, Xiaochen Zhao, Gaoang Wang, Erping Li 0001, Aili Wang 0002 |
ICML | 5 |
| 2025 | Distant supervised relation extraction with label entailment and collaborative denoising
Tingyu Xie, Qi Li 0042, Gaoang Wang, Hongwei Wang 0001 |
J. Intell. Inf. Syst. | 3 |
| 2025 | UnICLAM: Contrastive representation learning with adversarial masking for unified and interpretable Medical Vision Question Answering
Chenlu Zhan, Peng Peng 0006, Hongwei Wang 0001, Gaoang Wang, Hongsen Wang |
Medical Image Anal. | 4 |
| 2025 | Pose-Guided Transformer for Fine-Grained Action Quality AssessmentabstractAction Quality Assessment (AQA) is a task aimed at automatically and fairly evaluating the level of movement execution, which holds significant importance for action understanding. Previous methods, while adept at extracting video features, often neglect human regions. This leads to a limited capability to discern subtle action differences and results in a lack of interpretative depth. In this work, we propose a Pose-Guided Transformer framework, termed PGT, for assessing action quality more accurately. Essentially, this framework incorporates pose information to augment human region features during video feature extraction. The PGT framework incorporates two critical modules: a pose-guided attention layer and a global-local feature extractor. The former is designed to isolate body-specific features, effectively minimizing background noise, while the latter further delineates fine-grained features by utilizing decomposed information from various human body parts. The proposed PGT achieves significant results on various challenging AQA benchmarks. Notably, on MTL-AQA dataset, with a Spearman’s rank correlation of 0.9630. Additionally, on the AQA-7 dataset, our approach achieves an average Spearman’s rank correlation of 0.8673, further validating the effectiveness of our method. These findings demonstrate that our framework excels in the task of action quality assessment, providing a viable solution for accurate and fair evaluation of movement execution. Yanting Zhang 0001, Wenhao Chai, Cairong Yan, Wenhai Wang, Gaoang Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | S4Fusion: Saliency-Aware Selective State Space Model for Infrared and Visible Image FusionabstractThe preservation and the enhancement of complementary features between modalities are crucial for multi-modal image fusion and downstream vision tasks. However, existing methods are limited to local receptive fields (CNNs) or lack comprehensive utilization of spatial information from both modalities during interaction (transformers), which results in the inability to effectively retain useful information from both modalities in a comparative manner. Consequently, the fused images may exhibit a bias towards one modality, failing to adaptively preserve salient targets from all sources. Thus, a novel fusion framework (S4Fusion) based on the Saliency-aware Selective State Space is proposed. S4Fusion introduces the Cross-Modal Spatial Awareness Module (CMSA), which is designed to simultaneously capture global spatial information from all input modalities and promote effective cross-modal interaction. This enables a more comprehensive representation of complementary features. Furthermore, to guide the model in adaptively preserving salient objects, we propose a novel perception-enhanced loss function. This loss aims to enhance the retention of salient features by minimizing ambiguity or uncertainty, as measured at a pre-trained model's decision layer, within the fused images. The code is available at https://github.com/zipper112/S4Fusion. Haolong Ma, Hui Li 0037, Chunyang Cheng, Gaoang Wang, Xiaoning Song, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Efficient Transfer From Image-Based Large Multimodal Models to Video TasksabstractExtending image-based Large Multimodal Models (LMMs) to video-based LMMs always requires temporal modeling in the pre-training. However, training the temporal modules gradually erases the knowledge of visual features learned from various image-text-based scenarios, leading to degradation in some downstream tasks. % Adapting pre-trained video-based large language models (LLMs) to downstream fine-grained video understanding tasks always requires modeling on temporal modules. However, training the temporal modules during video pretraining gradually erases the knowledge of visual features learned from various image-text-based scenarios, leading to degradation in some downstream tasks. % Instead of tuning video-based LLMs to downstream tasks, To address this issue, in this paper, we introduce a novel, efficient transfer approach termed MTransLLAMA, which employs transfer learning from pre-trained image LMMs for fine-grained video tasks with only small-scale training sets. Our method enablesfewer trainable parametersand achievesfaster adaptationandhigher accuracythan pre-training video-based LMM models. Specifically, our method adopts early fusion between textual and visual features to capture fine-grained information, reuses spatial attention weights in temporal attentions for cyclical spatial-temporal reasoning, and introduces dynamic attention routing to capture both global and local information in spatial-temporal attentions. Experiments demonstrate that across multiple datasets and tasks, without relying on video pre-training, our model achieves state-of-the-art performance, enabling lightweight and efficient transfer from image-based LMMs to fine-grained video tasks. Shidong Cao, Zhonghan Zhao, Shengyu Hao, Wenhao Chai, Jenq-Neng Hwang, Hongwei Wang 0001, Gaoang Wang |
IEEE Trans. Multim. | 7 |
| 2025 | A Survey of Deep Learning in Sports Applications: Perception, Comprehension, and DecisionabstractDeep learning has the potential to revolutionize sports performance, with applications ranging from perception and comprehension to decision. This article presents a comprehensive survey of deep learning in sports performance, focusing on three main aspects: algorithms, datasets and virtual environments, and challenges. First, we discuss the hierarchical structure of deep learning algorithms in sports performance which includes perception, comprehension and decision while comparing their strengths and weaknesses. Second, we list widely used existing datasets in sports and highlight their characteristics and limitations. Finally, we summarize current challenges and point out future trends of deep learning in sports. Our survey provides valuable reference material for researchers interested in deep learning in sports applications. Zhonghan Zhao, Wenhao Chai, Shengyu Hao, Wenhao Hu 0002, Guanhong Wang, Shidong Cao, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2024 | Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsabstractDenoising Diffusion Probabilistic Models (DDPMs) have achieved significant success in generation tasks. Nevertheless, the exposure bias issue, i.e., the natural discrepancy between the training (the output of each step is calculated individually by a given input) and inference (the output of each step is calculated based on the input iteratively obtained based on the model), harms the performance of DDPMs. To our knowledge, few works have tried to tackle this issue by modifying the training process for DDPMs, but they still perform unsatisfactorily due to 1) partially modeling the discrepancy and 2) ignoring the prediction error accumulation. To address the above issues, in this paper, we propose a multi-step denoising scheduled sampling (MDSS) strategy to alleviate the exposure bias for DDPMs. Analyzing the formulations of the training and inference of DDPMs, MDSS 1) comprehensively considers the discrepancy influence of prediction errors on the output of the model (the Gaussian noise) and the output of the step (the calculated input signal of the next step), and 2) efficiently models the prediction error accumulation by using multiple iterations of a mathematical formulation initialized from one-step prediction error obtained from the model. The experimental results, compared with previous works, demonstrate that our approach is more effective in mitigating exposure bias in DDPM, DDIM, and DPM-solver. In particular, MDSS achieves an FID score of 3.86 in 100 sample steps of DDIM on the CIFAR-10 dataset, whereas the second best obtains 4.78. The code will be available on GitHub. Zhiyao Ren, Yibing Zhan, Liang Ding 0006, Gaoang Wang, Zhongyi Fan, Dacheng Tao |
AAAI | 4 |
| 2024 | UniAP: Towards Universal Animal Perception in Vision via Few-Shot LearningabstractAnimal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception model that can freely adapt to different animals across various perception tasks, due to the varying poses of a large diversity of animals, lacking data on rare species, and the semantic inconsistency of different tasks. We introduce UniAP, a novel Universal Animal Perception model that leverages few-shot learning to enable cross-species perception among various visual tasks. Our proposed model takes support images and labels as prompt guidance for a query image. Images and labels are processed through a Transformer-based encoder and a lightweight label encoder, respectively. Then a matching module is designed for aggregating information between prompt guidance and the query image, followed by a multi-head label decoder to generate outputs for various tasks. By capitalizing on the shared visual characteristics among different animals and tasks, UniAP enables the transfer of knowledge from well-studied species to those with limited labeled data or even unseen species. We demonstrate the effectiveness of UniAP through comprehensive experiments in pose estimation, segmentation, and classification tasks on diverse animal species, showcasing its ability to generalize and adapt to new classes with minimal labeled examples. Meiqi Sun, Zhonghan Zhao, Wenhao Chai, Hanjun Luo, Shidong Cao, Yanting Zhang 0001, Jenq-Neng Hwang, Gaoang Wang |
AAAI | 8 |
| 2024 | MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingabstractRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat. Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang |
CVPR | 13 |
| 2024 | BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionabstractVision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread. Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001 |
CVPR | 8 |
| 2024 | MedM2G: Unifying Medical Multi-Modal Generation via Cross-Guided Diffusion with Visual InvariantabstractMedical generative models, acknowledged for their high-quality sample generation ability, have accelerated the fast growth of medical applications. However, recent works concentrate on separate medical generation models for dis-tinct medical tasks and are restricted to inadequate medi-cal multimodal knowledge, constraining medical compre-hensive diagnosis. In this paper, we propose MedM2G, a Medical Multi-Modal Generative framework, with the key innovation to align, extract, and generate medical multimodal within a unified model. Extending beyond single or two medical modalities, we efficiently align medical multimodal through the central alignment approach in the unified space. Significantly, our framework extracts valuable clini-cal knowledge by preserving the medical visual invariant of each imaging modal, thereby enhancing specific medical information for multimodal generation. By conditioning the adaptive cross-guided parameters into the multi-flow diffusion framework, our model promotes flexible interactions among medical multimodalfor generation. MedM2G is the first medical generative model that unifies medical generation tasks of text-to-image, image-to-text, and unified generation of medical modalities (CT, MRI, X-ray). It performs 5 medical generation tasks across 10 datasets, consistently outperforming various state-of-the-art works. Chenlu Zhan, Gaoang Wang, Hongwei Wang 0001, Jian Wu 0001 |
CVPR | 3 |
| 2024 | See and Think: Embodied Agent in Virtual Environment
Zhonghan Zhao, Wenhao Chai, Boyi Li 0002, Shengyu Hao, Shidong Cao, Tian Ye 0001, Gaoang Wang |
ECCV (8) | 8 |
| 2024 | Blind Inpainting with Object-Aware Discrimination for Artificial Marker RemovalabstractMedical images often incorporate doctor-added markers that can hinder AI-based diagnosis. This issue highlights the need of inpainting techniques to restore the corrupted visual contents. However, existing methods require manual mask annotation as input, limiting the application scenarios. In this paper, we propose a novel blind inpainting method that automatically reconstructs visual contents within the corrupted regions without mask input as guidance. Our model includes a blind reconstruction network and an object-aware discriminator for adversarial training. The reconstruction network contains two branches that predict corrupted regions in images and simultaneously restore the missing visual contents. Leveraging the potent recognition capability of a dense object detector, the object-aware discriminator ensures markers undetectable after inpainting. Thus, the restored images closely resemble the clean ones. We evaluate our method on three datasets of various medical imaging modalities, confirming better performance over other state-of-the-art methods. Xuechen Guo, Wenhao Hu 0002, Chiming Ni, Wenhao Chai, Shiyan Li, Gaoang Wang |
ICASSP | 6 |
| 2024 | Sam-Guided Enhanced Fine-Grained Encoding with Mixed Semantic Learning for Medical Image CaptioningabstractWith the development of multimodality and large language models, the deep learning-based technique for medical image captioning holds the potential to offer valuable diagnostic recommendations. However, current generic text and image pre-trained models do not yield satisfactory results when it comes to describing intricate details within medical images. In this paper, we present a novel medical image captioning method guided by the segment anything model (SAM) to enable enhanced encoding with both general and detailed feature extraction. In addition, our approach employs a distinctive pre-training strategy with mixed semantic learning to simultaneously capture both the overall information and finer details within medical images. We demonstrate the effectiveness of this approach, as it outperforms the pre-trained BLIP2 model on various evaluation metrics for generating descriptions of medical images. Benlu Wang, Weijie Liang, Xuechen Guo, Guanhong Wang, Shiyan Li, Gaoang Wang |
ICASSP | 8 |
| 2024 | Vision meets mmWave Radar: 3D Object Perception Benchmark for Autonomous DrivingabstractSensor fusion is crucial for an accurate and robust perception system on autonomous vehicles. Most existing datasets and perception solutions focus on fusing cameras and LiDAR. However, the collaboration between camera and radar is significantly under-exploited. Incorporating rich semantic information from the camera and reliable 3D information from the radar can achieve an efficient, cheap, and portable solution for 3D perception tasks. It can also be robust to different lighting or all-weather driving scenarios due to the capability of mmWave radars. In this paper, we introduce the CRUW3D dataset, including 66K synchronized and well-calibrated camera, radar, and LiDAR frames in various driving scenarios. Unlike other large-scale autonomous driving datasets, our radar data is in the format of radio frequency (RF) tensors that contain not only 3D location information but also spatio-temporal semantic information. This kind of radar format can enable machine learning models to generate more reliable object perception results after interacting and fusing the information or features between the camera and radar. We run several camera- and radar-based baseline methods for 3D object detection and multi-object tracking on our dataset. We hope the CRUW3D dataset will foster radar and multi-modal 3D perception research. CRUW3D is available at https://huggingface.co/datasets/uwipl/CRUW3D Yizhou Wang 0005, Jen-Hao Cheng, Jui-Te Huang, Sheng-Yao Kuan, Qiqian Fu, Chiming Ni, Shengyu Hao, Gaoang Wang, Guanbin Xing, Hui Liu 0011, Jenq-Neng Hwang |
IV | 8 |
| 2024 | Enhanced Multimodal Trajectory Prediction for Autonomous Vehicles Using Advanced Diffusion Model TechniquesabstractVehicle trajectory prediction is crucial for ensuring the safety and reliability of autonomous driving systems. Due to the highly stochastic nature of road participants’ behaviors, it is vital that prediction models accommodate a wide range of possible scenarios to mitigate safety risks. To address this challenge, we propose a novel trajectory prediction model called DiffusionTrajPred, an innovative trajectory prediction model based on the diffusion model. This model uniquely combines forward and reverse processes, manipulating noise levels in trajectory data to forecast future paths. Through the application of a mask-based reverse process, the model can make full use of historical trajectory information and predict trajectories that combine accuracy and multiple possibilities. The model utilizes a Transformer architecture for learning the noise, which enables the model to extract richer temporal information from trajectory data, resulting in improved semantic comprehension. Furthermore, we have effectively encoded high-definition (HD) semantic map information and vehicle interaction dynamics as crucial input features, improving the model ’s predictive power. Extensive experiments on the widely recognized open-source dataset ’Argoverse’ reveal that our method outperformed the most existing state-of-the-art methods in terms of accuracy and multimodality, demonstrating the diffusion model’s unique advantage in addressing the stochastic nature of road scenarios in autonomous driving. Song Lian, Simon Hu 0001, Jianghan Hu, Gaoang Wang, José Escribano, Xiaoxiang Na, Sheng Jin 0001 |
IV | 5 |
| 2024 | LLaVA-Ultra: Large Chinese Language and Vision Assistant for UltrasoundabstractMultimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performing multimodal tasks. This boom begins to significantly impact medical field. However, general visual language model (VLM) lacks sophisticated comprehension for medical visual question answering (Med-VQA). Even models specifically tailored for medical domain tend to produce vague answers with weak visual relevance. In this paper, we propose a fine-grained adaptive VLM architecture for Chinese medical visual conversations through parameter-efficient tuning. Specifically, we devise a fusion module with fine-grained vision encoders to achieve enhancement for subtle medical visual semantics. Then we note data redundancy common to medical scenes is ignored in most prior works. In cases of a single text paired with multiple figures, we utilize weighted scoring with knowledge distillation to adaptively screen valid images mirroring text descriptions. For execution, we leverage a large-scale multimodal Chinese ultrasound dataset obtained from the hospital. We create instruction-following data based on text from professional doctors, which ensures effective tuning. With enhanced model and quality data, our Large Chinese Language and Vision Assistant for Ultra sound (LLaVA-Ultra) shows strong capability and robustness to medical scenarios. On three Med-VQA datasets, LLaVA-Ultra surpasses previous state-of-the-art models on various metrics. Xuechen Guo, Wenhao Chai, Shiyan Li, Gaoang Wang |
ACM Multimedia | 4 |
| 2024 | Ego3DT: Tracking Every 3D Object in Ego-centric VideosabstractThe growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localization and tracking of objects in ego-centric videos, primarily due to the substantial variability in viewing angles. Addressing this issue, this paper introduces a novel zero-shot approach for the 3D reconstruction and tracking of all objects from the ego-centric video. We present Ego3DT, a novel framework that initially identifies and extracts detection and segmentation information of objects within the ego environment. Utilizing information from adjacent video frames, Ego3DT dynamically constructs a 3D scene of the ego view using a pre-trained 3D scene reconstruction model. Additionally, we have innovated a dynamic hierarchical association mechanism for creating stable 3D tracking trajectories of objects in ego-centric videos. Moreover, the efficacy of our approach is corroborated by extensive experiments on two newly compiled datasets, with 1.04 × - 2.90× in HOTA, showcasing the robustness and accuracy of our method in diverse ego-centric scenarios. Shengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun, Wendi Hu, Jieyang Zhou, Yixian Zhao, Yizhou Wang 0005, Gaoang Wang |
ACM Multimedia | 11 |
| 2024 | An Efficient Multi-prior Hybrid Approach for Consistent 3D Generation from Single Images
Yichen Ouyang, Jiayi Ye, Wenhao Chai, Dapeng Tao, Yibing Zhan, Gaoang Wang |
MMAsia | 6 |
| 2024 | Advancing Training Efficiency of Deep Spiking Neural Networks through Rate-based BackpropagationabstractRecent insights have revealed that rate-coding is a primary form of information representation captured by surrogate-gradient-based Backpropagation Through Time (BPTT) in training deep Spiking Neural Networks (SNNs). Motivated by these findings, we propose rate-based backpropagation, a training strategy specifically designed to exploit rate-based representations to reduce the complexity of BPTT. Our method minimizes reliance on detailed temporal derivatives by focusing on averaged dynamics, streamlining the computational graph to reduce memory and computational demands of SNNs training. We substantiate the rationality of the gradient approximation between BPTT and the proposed method through both theoretical analysis and empirical observations. Comprehensive experiments on CIFAR-10, CIFAR-100, ImageNet, and CIFAR10-DVS validate that our method achieves comparable performance to BPTT counterparts, and surpasses state-of-the-art efficient training techniques. By leveraging the inherent benefits of rate-coding, this work sets the stage for more scalable and efficient SNNs training within resource-constrained environments. Chengting Yu, Gaoang Wang, Erping Li 0001, Aili Wang 0002 |
NeurIPS | 3 |
| 2024 | MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
Zhenyu Zhang 0030, Wenhao Chai, Zhongyu Jiang, Tian Ye 0001, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
PRCV (11) | 7 |
| 2024 | DIVOTrack: A Novel Dataset and Baseline Method for Cross-View Multi-Object Tracking in DIVerse Open Scenes
Shengyu Hao, Peiyuan Liu, Yibing Zhan, Kaixun Jin, Zuozhu Liu, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
Int. J. Comput. Vis. | 8 |
| 2024 | Knowledge-guided pre-training and fine-tuning: Video representation learning for action recognition
Guanhong Wang, Zhanhao He, Keyu Lu, Yang Feng 0011, Zuozhu Liu, Gaoang Wang |
Neurocomputing | 7 |
| 2024 | Self-Paced Multi-Grained Cross-Modal Interaction Modeling for Referring Expression ComprehensionabstractAs an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning. In addition, due to the diversity of visual scenes and the variation of linguistic expressions, some hard examples have much more abundant multi-grained information than others. How to aggregate multi-grained information from different modalities and extract abundant knowledge from hard examples is crucial in the REC task. To address aforementioned challenges, in this paper, we propose a Self-paced Multi-grained Cross-modal Interaction Modeling framework, which improves the language-to-vision localization ability through innovations in network structure and learning mechanism. Concretely, we design a transformer-based multi-grained cross-modal attention, which effectively utilizes the inherent multi-grained information in visual and linguistic encoders. Furthermore, considering the large variance of samples, we propose a self-paced sample informativeness learning to adaptively enhance the network learning for samples containing abundant multi-grained information. The proposed framework significantly outperforms state-of-the-art methods on widely used datasets, such as RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame datasets, demonstrating the effectiveness of our method. Peihan Miao 0002, Wei Su 0009, Gaoang Wang, Xuewei Li 0003, Xi Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | DiffFashion: Reference-Based Fashion Design With Structure-Aware Transfer by Diffusion ModelsabstractImage-based fashion design with AI techniques has attracted increasing attention in recent years. We focus on the reference-based fashion design task, where we aim to combine a reference appearance image and a clothing image to generate a new fashion clothing image. Although existing diffusion-based image translation methods have enabled flexible style transfer, it is often difficult to transfer the appearance of the image realistically during reverse diffusion. When the referenced appearance domain greatly differs from the source domain, it often leads to the collapse in the translation. To tackle this issue, we present a novel diffusion model-based unsupervised structure-aware transfer method, namelyDiffFashion. Our method is free of model tuning and structure-preserving and has high flexibility in transferring from images with large domain gaps. Specifically, based on the optimal transport properties, we keep a shared latent across the clothing image and reference appearance image to bridge the gap between the two domains in the denoising process, and the latent of the reference image is gradually adapted to the clothing domain. Simultaneously, the structure is transferred from the source clothing to the output fashion image with mixed guidance, including pre-trained Vision Transformer (ViT) guidance and a foreground mask guidance, to further preserve the structure and appearance semantics from source and reference images. Our experimental results show that the proposed method outperforms state-of-the-art baseline models, generating more realistic images in the fashion design task. Shidong Cao, Wenhao Chai, Shengyu Hao, Yanting Zhang 0001, Hangyue Chen, Gaoang Wang |
IEEE Trans. Multim. | 6 |
| 2024 | UniDCP: Unifying Multiple Medical Vision-Language Tasks via Dynamic Cross-Modal Learnable PromptsabstractMedical vision-language pre-training (Med-VLP) models have recently accelerated the fast-growing medical diagnostics application. However, most Med-VLP models learn task-specific representations independently from scratch, thereby leading to great inflexibility when they work across multiple fine-tuning tasks. In this work, we proposeUniDCP, aUnified medical vision-language model withDynamicCross-modal learnablePrompts, which can be plastically applied to multiple medical vision-language tasks within a unified model. Specifically, we explicitly construct a unified framework to harmonize diverse inputs from multiple pre-training tasks by leveraging cross-modal prompts for unification, which accordingly can accommodate heterogeneous medical fine-tuning tasks within a same model. Furthermore, we conceive a dynamic cross-modal prompt optimizing strategy that optimizes the prompts within the shareable space for implicitly processing the shareable clinic knowledge. UniDCP is the first Med-VLP model capable of performing all 8 medical uni-modal and cross-modal tasks over 14 corresponding datasets, consistently yielding superior results over diverse state-of-the-art methods. Chenlu Zhan, Yufei Zhang 0015, Gaoang Wang, Hongwei Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Language Adaptive Weight Generation for Multi-Task Visual GroundingabstractAlthough the impressive performance in visual grounding, the prevailing approaches usually exploit the visual backbone in a passive way, i.e., the visual backbone extracts features with fixed weights without expression-related hints. The passive perception may lead to mismatches (e.g., redundant and missing), limiting further performance improvement. Ideally, the visual backbone should actively extract visual features since the expressions already provide the blueprint of desired visual features. The active perception can take expressions as priors to extract relevant visual features, which can effectively alleviate the mismatches. Inspired by this, we propose an active perception Visual Grounding framework based on Language Adaptive Weights, called VG-LAW. The visual backbone serves as an expression-specific feature extractor through dynamic weights generated for various expressions. Benefiting from the specific and relevant visual features extracted from the language-aware visual backbone, VG-LAW does not require additional modules for cross-modal interaction. Along with a neat multi-task head, VG-LAW can be competent in referring expression comprehension and segmentation jointly. Extensive experiments on four representative datasets, i.e., RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame, validate the effectiveness of the proposed framework and demonstrate state-of-the-art performance. Wei Su 0009, Peihan Miao 0002, Huanzhang Dou, Gaoang Wang, Liang Qiao 0001, Zheyang Li, Xi Li 0001 |
CVPR | 4 |
| 2023 | TransLink: Transformer-Based Embedding for Tracklets' Global LinkabstractMulti-object tracking (MOT) is essential to many tasks related to the smart transportation. Detecting and tracking humans on the road can give a vital feedback for either the moving vehicle or traffic control to ensure better driving safety and traffic flow. However, most trackers face a common problem of identity (ID) switch, resulting in an incomplete human trajectory prediction. In this paper, we propose a Transformer-based tracklet linking method called TransLink to mitigate the association failures. Specifically, the self-attention mechanism is well exploited to get the feature representation for tracklets, followed by a multilayer perceptron to predict the association likelihood, which can be further used in determining the tracklet association. Experiments on the MOT dataset demonstrate the effectiveness of the proposed module in lifting the tracking performances. Yanting Zhang 0001, Shuanghong Wang, Yuxuan Fan, Gaoang Wang, Cairong Yan |
ICASSP | 4 |
| 2023 | StableVideo: Text-driven Consistency-aware Diffusion Video EditingabstractDiffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents diffusion models from being applied to natural video editing in practical scenarios. In this paper, we tackle this problem by introducing temporal dependency to existing text-driven diffusion models, which allows them to generate consistent appearance for the edited objects. Specifically, we develop a novel inter-frame propagation mechanism for diffusion video editing, which leverages the concept of layered representations to propagate the appearance information from one frame to the next. We then build up a text-driven video editing framework based on this mechanism, namely StableVideo, which can achieve consistency-aware video editing. Extensive experiments demonstrate the strong editing capability of our approach. Compared with state-of-the-art video editing methods, our approach shows superior qualitative and quantitative results. Our code is available at this https URL. Wenhao Chai, Xun Guo 0002, Gaoang Wang, Yan Lu 0001 |
ICCV | 3 |
| 2023 | Global Adaptation meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose EstimationabstractWhen applying a pre-trained 2D-to-3D human pose lifting model to a target unseen dataset, large performance degradation is commonly encountered due to domain shift issues. We observe that the degradation is caused by two factors: 1) the large distribution gap over global positions of poses between the source and target datasets due to variant camera parameters and settings, and 2) the deficient diversity of local structures of poses in training. To this end, we combine global adaptation and local generalization in PoseDA, a simple yet effective framework of unsupervised domain adaptation for 3D human pose estimation. Specifically, global adaptation aims to align global positions of poses from the source domain to the target domain with a proposed global position alignment (GPA) module. And local generalization is designed to enhance the diversity of 2D-3D pose mapping with a local pose augmentation (LPA) module. These modules bring significant performance improvement without introducing additional learnable parameters. In addition, we propose local pose augmentation (LPA) to enhance the diversity of 3D poses following an adversarial training scheme consisting of 1) a augmentation generator that generates the parameters of pre-defined pose transformations and 2) an anchor discriminator to ensure the reality and quality of the augmented data. Our approach can be applicable to almost all 2D-3D lifting models. PoseDA achieves 61.3 mm of MPJPE on MPI-INF-3DHP under a cross-dataset evaluation setup, improving upon the previous state-of-the-art method by 10.2%. Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang, Gaoang Wang |
ICCV | 4 |
| 2023 | Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionabstractKnowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to the optimal student classification scores for dense object detectors. This cross-task protocol inconsistency is critical, especially for dense object detectors, since the foreground categories are extremely imbalanced. To address the issue of protocol differences between distillation and classification, we propose a novel distillation method with cross-task consistent protocols, tailored for the dense object detection. For classification distillation, we address the cross-task protocol inconsistency problem by formulating the classification logit maps in both teacher and student models as multiple binary-classification maps and applying a binary-classification distillation loss to each map. For localization distillation, we design an IoU-based Localization Distillation Loss that is free from specific network structures and can be compared with existing localization distillation losses. Our proposed method is simple but effective, and experimental results demonstrate its superiority over existing methods. Code is available at https://github.com/TinyTigerPan/BCKD. Longrong Yang, Xianpan Zhou, Xuewei Li 0003, Liang Qiao 0001, Zheyang Li, Ziwei Yang 0004, Gaoang Wang, Xi Li 0001 |
ICCV | 7 |
| 2023 | Learning Discrimination from Contaminated Data: Multi-Instance Learning for Unsupervised Anomaly DetectionabstractAnomaly detection aims at identifying deviant samples from the normal data distribution. Much progress has been made in recent years for anomaly detection with self-supervised representation learning. However, most existing approaches assume the training set contains either only clean normal samples or some labeled abnormal samples. With a contaminated unlabeled training set, the performance is degraded with unclear discrimination between normal and abnormal samples. To address the above challenge, in this paper, we propose a novel unsupervised representation learning framework that takes advantage of the extra information provided by the anomalies in the unlabeled contaminated data for anomaly detection. Specifically, anomaly discrimination learning with pseudo normality score generation and multi-instance contrastive learning is proposed for distinguishing abnormal samples from the unlabeled set. Meanwhile, we combine the discrimination learning with mean-shifted contrastive learning and distribution-shifting transformation classification into a multi-task representation learning framework to improve the stability and robustness of the training. Our proposed method achieves state-of-the-art performance on CIFAR-10, ImageNet-10, MNIST and F-MNIST datasets. Extensive ablation studies further demonstrate the effectiveness of different components of the proposed framework. Wenhao Hu 0002, Xuanyu Chen, Gaoang Wang |
ICME | 5 |
| 2023 | SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic SegmentationabstractAs an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D properties of original 360 degree data. Therefore, their performance will drop a lot when inputting panoramic images with the 3D disturbance. To be more robust to 3D disturbance, we propose our Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation (SGAT4PASS), considering 3D spherical geometry knowledge. Specifically, a spherical geometry-aware framework is proposed for PASS. It includes three modules, i.e., spherical geometry-aware image projection, spherical deformable patch embedding, and a panorama-aware loss, which takes input images with 3D disturbance into account, adds a spherical geometry-aware constraint on the existing deformable patch embedding, and indicates the pixel density of original 360 degree data, respectively. Experimental results on Stanford2D3D Panoramic datasets show that SGAT4PASS significantly improves performance and robustness, with approximately a 2% increase in mIoU, and when small 3D disturbances occur in the data, the stability of our performance is improved by an order of magnitude. Our code and supplementary material are available at https://github.com/TencentARC/SGAT4PASS. Xuewei Li 0003, Zhongang Qi, Gaoang Wang, Ying Shan, Xi Li 0001 |
IJCAI | 4 |
| 2023 | Temporal Constrained Feasible Subspace Learning for Human Pose Forecasting
Gaoang Wang, Mingli Song |
IJCAI | 1 |
| 2023 | Debiasing Medical Visual Question Answering via Counterfactual Training
Chenlu Zhan, Peng Peng 0006, Hanrong Zhang, Haiyue Sun, Chunnan Shang, Hongsen Wang, Gaoang Wang, Hongwei Wang 0001 |
MICCAI (2) | 8 |
| 2023 | PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose EstimationabstractThe current 3D human pose estimators face challenges in adapting to new datasets due to the scarcity of 2D-3D pose pairs in target domain training sets. We present the Multi-Hypothesis Pose Synthesis Domain Adaptation (PoSynDA) framework to overcome this issue without extensive target domain annotation. Utilizing a diffusion-centric structure, PoSynDA simulates the 3D pose distribution in the target domain, filling the data diversity gap. By incorporating a multi-hypothesis network, it creates diverse pose hypotheses and aligns them with the target domain. Target-specific source augmentation obtains the target domain distribution data from the source domain by decoupling the scale and position parameters. The teacher-student paradigm and low-rank adaptation further refine the process. PoSynDA demonstrates competitive performance on benchmarks, such as Human3.6M, MPI-INF-3DHP, and 3DPW, even comparable with the target-trained MixSTE model. This work paves the way for the practical application of 3D human pose estimation1. The source code is available at https://github.com/hbing-l/PoSynDA. Jun-Yan He, Zhi-Qi Cheng, Wangmeng Xiang, Qize Yang, Wenhao Chai, Gaoang Wang, Xu Bao 0003, Bin Luo 0008, Yifeng Geng, Xuansong Xie |
ACM Multimedia | 7 |
| 2023 | User-Aware Prefix-Tuning Is a Good Learner for Personalized Image Captioning
Guanhong Wang, Wenhao Chai, Gaoang Wang |
PRCV (7) | 5 |
| 2023 | Handwritten Chinese signature detection with simple Copy-Paste augmentation on power plants technical documents
Jian Zhang 0083, Kaihong Yan, Hongwei Wang 0001, Gaoang Wang |
Serv. Oriented Comput. Appl. | 5 |
| 2023 | Hierarchical Self-Supervised Learning for 3D Tooth Segmentation in Intra-Oral Mesh ScansabstractAccurately delineating individual teeth and the gingiva in the three-dimension (3D) intraoral scanned (IOS) mesh data plays a pivotal role in many digital dental applications, e.g., orthodontics. Recent research shows that deep learning based methods can achieve promising results for 3D tooth segmentation, however, most of them rely on high-quality labeled dataset which is usually of small scales as annotating IOS meshes requires intensive human efforts. In this paper, we propose a novel self-supervised learning framework, named STSNet, to boost the performance of 3D tooth segmentation leveraging on large-scale unlabeled IOS data. The framework follows two-stage training, i.e., pre-training and fine-tuning. In pre-training, three hierarchical-level, i.e., point-level, region-level, cross-level, contrastive losses are proposed for unsupervised representation learning on a set of predefined matched points from different augmented views. The pretrained segmentation backbone is further fine-tuned in a supervised manner with a small number of labeled IOS meshes. With the same amount of annotated samples, our method can achieve an mIoU of 89.88%, significantly outperforming the supervised counterparts. The performance gain becomes more remarkable when only a small amount of labeled samples are available. Furthermore, STSNet can achieve better performance with only 40% of the annotated samples as compared to the fully supervised baselines. To the best of our knowledge, we present the first attempt of unsupervised pre-training for 3D tooth segmentation, demonstrating its strong potential in reducing human efforts for annotation and verification. Zuozhu Liu, Xiaoxuan He, Hualiang Wang, Huimin Xiong, Yan Zhang 0004, Gaoang Wang, Jin Hao, Yang Feng 0011, Fudong Zhu, Haoji Hu |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Split and Connect: A Universal Tracklet Booster for Multi-Object TrackingabstractMulti-object tracking (MOT) is an essential task in the computer vision field. With the fast development of deep learning technology in recent years, MOT has achieved great improvement. However, some challenges still remain, such as sensitiveness to occlusion, instability under different lighting conditions, and non-robustness to deformable objects, causing incorrect temporal associations. To address such common challenges in most of the existing trackers, in this paper, a tracklet booster (TBooster) algorithm is proposed to correct the association errors resulting from existing trackers. The correction of the association error from TBooster has two folds: split tracklets on potential ID-change positions and then connect multiple tracklets into one if they are from the same object. To achieve this goal, the TBooster consists of two components,i.e., Splitter and Connector. In Splitter, an architecture with stacked temporal dilated convolution blocks is employed for the splitting position prediction via label smoothing strategy with adaptive Gaussian kernels. In Connector, a multi-head self-attention-based encoder is exploited for the tracklet embedding, which is further used to connect tracklets into full tracks. We conduct sufficient experiments on MOT17 and MOT20 benchmark datasets and achieve promising results. Combined with the proposed tracklet booster, existing trackers can achieve large improvements on the IDF1 score, which shows the effectiveness of the proposed TBooster. Gaoang Wang, Yizhou Wang 0005, Renshu Gu, Weijie Hu, Jenq-Neng Hwang |
IEEE Trans. Multim. | 1 |
| 2022 | Hierarchical Semi-supervised Contrastive Learning for Contamination-Resistant Anomaly Detection
Gaoang Wang, Yibing Zhan, Xinchao Wang, Mingli Song, Klara Nahrstedt |
ECCV (25) | 1 |
| 2022 | ActiveMatch: End-To-End Semi-Supervised Active Representation LearningabstractSemi-supervised learning (SSL) is an efficient framework that can train models with both labeled and unlabeled data, but may generate ambiguous and non-distinguishable representations when lacking adequate labeled samples. With human-in-the-loop, active learning can iteratively select informative unlabeled samples for labeling and training to improve the performance in the SSL framework. However, most existing active learning approaches rely on pre-trained features, which is not suitable for end-to-end learning. To deal with the drawbacks of SSL, in this paper, we propose a novel end-to-end representation learning method, namely ActiveMatch, which combines SSL with contrastive learning and active learning to fully leverage the limited labels. Starting from a small amount of labeled data with unsupervised contrastive learning as a warm-up, ActiveMatch then combines SSL and supervised contrastive learning, and actively selects the most representative samples for labeling during the training, resulting in better representations towards the classification. Compared with MixMatch and FixMatch with the same amount of labeled data, we show that ActiveMatch achieves the state-of-the-art performance, with 89.24% accuracy on CIFAR-10 with 100 collected labels, and 92.20% accuracy with 200 collected labels. Xinkai Yuan, Zilinghan Li, Gaoang Wang |
ICIP | 3 |
| 2022 | Human-Centered Prior-Guided and Task-Dependent Multi-Task Representation Learning for Action Recognition Pre-TrainingabstractRecently, much progress has been made for self-supervised action recognition. Most existing approaches emphasize the contrastive relations among videos, including appearance and motion consistency. However, two main issues remain for existing pre-training methods: 1) the learned representation is neutral and not informative for a specific task; 2) multi-task learning-based pre-training sometimes leads to sub-optimal solutions due to inconsistent domains of different tasks. To address the above issues, we propose a novel action recognition pre-training framework, which exploits human-centered prior knowledge that generates more informative representation’ and avoids the conflict between multiple tasks by using task-dependent representations. Specifically, we distill knowledge from a human parsing model to enrich the semantic capability of representation. In addition, we combine knowledge distillation with contrastive learning to constitute a task-dependent multi-task framework. We achieve state-of-the-art performance on two popular benchmarks for action recognition task, i.e., UCF101 and HMDB51, verifying the effectiveness of our method. Guanhong Wang, Keyu Lu, Zhanhao He, Gaoang Wang |
ICME | 5 |
| 2022 | When Few-Shot Learning Meets Video Object DetectionabstractDifferent from static images, videos contain additional temporal and spatial information for better object detection. However, it is costly to obtain a large number of videos with bounding box annotations that are required for supervised deep learning. Although humans can easily learn to recognize new objects by watching only a few video clips, deep learning usually suffers from overfitting. This leads to an important question: how to effectively learn a video object detector from only a few labeled video clips? In this paper, we study the new problem of few-shot learning for video object detection. We first define the few-shot setting and create a new benchmark dataset for few-shot video object detection derived from the widely used ImageNet VID dataset. We employ a transfer-learning framework to effectively train the video object detector on a large number of base-class objects and a few video clips of novel-class objects. By analyzing the results of two methods under this framework (Joint and Freeze) on our designed weak and strong base datasets, we reveal insufficiency and overfitting problems. A simple but effective method, called Thaw, is naturally developed to trade off the two problems and validate our analysis. Extensive experiments on our proposed benchmark datasets with different scenarios demonstrate the effectiveness of our novel analysis in this new few-shot video object detection problem. Zhongjie Yu 0003, Gaoang Wang, Lin Chen 0021, Sebastian Raschka, Jiebo Luo 0001 |
ICPR | 2 |
| 2022 | pcnaDeep: a fast and robust single-cell tracking method using deep-learning mediated cell cycle profilingabstractSUMMARY: Computational methods that track single cells and quantify fluorescent biosensors in time-lapse microscopy images have revolutionized our approach in studying the molecular control of cellular decisions. One barrier that limits the adoption of single-cell analysis in biomedical research is the lack of efficient methods to robustly track single cells over cell division events. Here, we developed an application that automatically tracks and assigns mother-daughter relationships of single cells. By incorporating cell cycle information from a well-established fluorescent cell cycle reporter, we associate mitosis relationships enabling high fidelity long-term single-cell tracking. This was achieved by integrating a deep-learning-based fluorescent proliferative cell nuclear antigen signal instance segmentation module with a cell tracking and cell cycle resolving pipeline. The application offers a user-friendly interface and extensible APIs for customized cell cycle analysis and manual correction for various imaging configurations. AVAILABILITY AND IMPLEMENTATION: pcnaDeep is an open-source Python application under the Apache 2.0 licence. The source code, documentation and tutorials are available at https://github.com/chan-labsite/PCNAdeep. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yifan Gui, Shuangshuang Xie, Renzhi Yao, Xukai Gao, Yutian Dong, Gaoang Wang, Kuan Yoow Chan |
Bioinform. | 8 |
| 2022 | Unsupervised universal hierarchical multi-person 3D pose estimation for natural scenes
Renshu Gu, Zhongyu Jiang, Gaoang Wang, Kevin McQuade, Jenq-Neng Hwang |
Multim. Tools Appl. | 3 |
| 2022 | Forgery-Domain-Supervised Deepfake Detection With Non-Negative ConstraintabstractFake faces produced by deepfake techniques have attracted public concerns in recent years. Deepfake detection is a binary classification task that distinguishes fake faces from real ones. As the training data for deepfake detection is usually generated from real faces via various face forgery methods, it is difficult for a single binary decision boundary to distinguish fake faces. Besides that, the learned features often involve irrelevant information for identifying fake faces that are generated from diverse forgery methods. To deal with such challenges, unlike existing approaches that regard fake detection as a binary classification, we re-model the task as a multiclass forgery-domain classification task, where each forgery method is treated as a distinct class. This simplifies the complex decision boundary brought by the diversity of forgery patterns and provides more forgery-relevant information for the learning process. In addition, we introduce a non-negative constrained learning framework composed of non-negative features and a non-negative constrained classifier (NCC) to block irrelevant features with zero weights and enhance forgery-relevant features with positive weights, leading to a sparse structure of the classifier. Furthermore, to capture subtle and discriminative forgery-relevant features, we propose an integration module over augmented faces based on cross-attention. We demonstrate that our approach achieves competitive performance and generalization ability on widely-used benchmarks through extensive experiments. Yike Yuan, Xinghe Fu, Gaoang Wang, Xi Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Track without Appearance: Learn Box and Tracklet Embedding with Local and Global Motion Patterns for Vehicle TrackingabstractVehicle tracking is an essential task in the multi-object tracking (MOT) field. A distinct characteristic in vehicle tracking is that the trajectories of vehicles are fairly smooth in both the world coordinate and the image coordinate. Hence, models that capture motion consistencies are of high necessity. However, tracking with the standalone motion-based trackers is quite challenging because targets could get lost easily due to limited information, detection error and occlusion. Leveraging appearance information to assist object re-identification could resolve this challenge to some extent. However, doing so requires extra computation while appearance information is sensitive to occlusion as well. In this paper, we try to explore the significance of motion patterns for vehicle tracking without appearance information. We propose a novel approach that tackles the association issue for long-term tracking with the exclusive fully-exploited motion information. We address the tracklet embedding issue with the proposed reconstruct-to-embed strategy based on deep graph convolutional neural networks (GCN). Comprehensive experiments on the KITTI-car tracking dataset and UA-Detrac dataset show that the proposed method, though without appearance information, could achieve competitive performance with the state-of-the-art (SOTA) trackers. The source code will be available at https://github.com/GaoangW/LGMTracker. Gaoang Wang, Renshu Gu, Zuozhu Liu, Weijie Hu, Mingli Song, Jenq-Neng Hwang |
ICCV | 1 |
| 2021 | ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving ApplicationsabstractThe Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community. Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu |
ICMR | 3 |
| 2021 | Beware of the generic machine learning-based scoring functions in structure-based virtual screeningabstractMachine learning-based scoring functions (MLSFs) have attracted extensive attention recently and are expected to be potential rescoring tools for structure-based virtual screening (SBVS). However, a major concern nowadays is whether MLSFs trained for generic uses rather than a given target can consistently be applicable for VS. In this study, a systematic assessment was carried out to re-evaluate the effectiveness of 14 reported MLSFs in VS. Overall, most of these MLSFs could hardly achieve satisfactory results for any dataset, and they could even not outperform the baseline of classical SFs such as Glide SP. An exception was observed for RFscore-VS trained on the Directory of Useful Decoys-Enhanced dataset, which showed its superiority for most targets. However, in most cases, it clearly illustrated rather limited performance on the targets that were dissimilar to the proteins in the corresponding training sets. We also used the top three docking poses rather than the top one for rescoring and retrained the models with the updated versions of the training set, but only minor improvements were observed. Taken together, generic MLSFs may have poor generalization capabilities to be applicable for the real VS campaigns. Therefore, it should be quite cautious to use this type of methods for VS. Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Jinping Pang, Gaoang Wang, Haiyang Zhong, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 6 |
| 2021 | Can machine learning consistently improve the scoring power of classical scoring functions? Insights into the role of machine learning in scoring functionsabstractHow to accurately estimate protein-ligand binding affinity remains a key challenge in computer-aided drug design (CADD). In many cases, it has been shown that the binding affinities predicted by classical scoring functions (SFs) cannot correlate well with experimentally measured biological activities. In the past few years, machine learning (ML)-based SFs have gradually emerged as potential alternatives and outperformed classical SFs in a series of studies. In this study, to better recognize the potential of classical SFs, we have conducted a comparative assessment of 25 commonly used SFs. Accordingly, the scoring power was systematically estimated by using the state-of-the-art ML methods that replaced the original multiple linear regression method to refit individual energy terms. The results show that the newly-developed ML-based SFs consistently performed better than classical ones. In particular, gradient boosting decision tree (GBDT) and random forest (RF) achieved the best predictions in most cases. The newly-developed ML-based SFs were also tested on another benchmark modified from PDBbind v2007, and the impacts of structural and sequence similarities were evaluated. The results indicated that the superiority of the ML-based SFs could be fully guaranteed when sufficient similar targets were contained in the training set. Moreover, the effect of the combinations of features from multiple SFs was explored, and the results indicated that combining NNscore2.0 with one to four other classical SFs could yield the best scoring power. However, it was not applicable to derive a generic target-specific SF or SF combination. Chao Shen 0008, Zhe Wang 0041, Xujun Zhang, Haiyang Zhong, Gaoang Wang, Lei Xu 0035, Dong-Sheng Cao 0001, Tingjun Hou |
Briefings Bioinform. | 6 |
| 2021 | Weakly supervised instance segmentation using multi-prior fusion
Shengyu Hao, Gaoang Wang, Renshu Gu |
Comput. Vis. Image Underst. | 2 |
| 2020 | Exploring Severe Occlusion: Multi-Person 3D Pose Estimation with Gated Convolutionabstract3D human pose estimation (HPE) is crucial in many fields, such as human behavior analysis, augmented reality/virtual reality (AR/VR) applications, and self-driving industry. Videos that contain multiple potentially occluded people captured from freely moving monocular cameras are very common in realworld scenarios, while 3D HPE for such scenarios is quite challenging, partially because there is a lack of such data with accurate 3D ground truth labels in existing datasets. In this paper, we propose a temporal regression network with a gated convolution module to transform 2D joints to 3D and recover the missing occluded joints in the meantime. A simple yet effective localization approach is further conducted to transform the normalized pose to the global trajectory. To verify the effectiveness of our approach, we also collect a new moving camera multi-human (MMHuman) dataset that includes multiple people with heavy occlusion captured by moving cameras. The 3D ground truth joints are provided by accurate motion capture (MoCap) system. From the experiments on static-camera based Human3.6M data and our own collected moving-camera based data, we show that our proposed method outperforms most state-of-the-art 2D-to-3D pose estimation methods, especially for the scenarios with heavy occlusions. Renshu Gu, Gaoang Wang, Jenq-Neng Hwang |
ICPR | 2 |
| 2020 | DAIL: Dataset-Aware and Invariant Learning for Face RecognitionabstractTo achieve good performance in face recognition, a large scale training dataset is usually required. A simple yet effective way to improve the recognition performance is to use a dataset as large as possible by combining multiple datasets in the training. However, it is problematic and troublesome to naively combine different datasets due to two major issues. First, the same person can possibly appear in different datasets, leading to an identity overlapping issue between different datasets. Naively treating the same person as different classes in different datasets during training will affect back-propagation and generate nonrepresentative embeddings. On the other hand, manually cleaning labels may take formidable human efforts, especially when there are millions of images and thousands of identities. Second, different datasets are collected in different situations and thus will lead to different domain distributions. Naively combining datasets will make it difficult to learn domain invariant embeddings across different datasets. In this paper, we propose DAIL: Dataset-Aware and Invariant Learning to resolve the above-mentioned issues. To solve the first issue of identity overlapping, we propose a dataset-aware loss for multi-dataset training by reducing the penalty when the same person appears in multiple datasets. This can be readily achieved with a modified softmax loss with a dataset-aware term. To solve the second issue, domain adaptation with gradient reversal layers is employed for dataset invariant learning. The proposed approach not only achieves the state-of-the-art results on several commonly used face recognition validation sets, including LFW, CFP-FP, and AgeDB-30, but also shows great benefit for practical use. Gaoang Wang, Lin Chen 0021, Tianqiang Liu, Mingwei He, Jiebo Luo 0001 |
ICPR | 1 |
| 2020 | Multi-Person Hierarchical 3D Pose Estimation in Natural VideosabstractDespite the increasing need of analyzing human poses on the street and in the wild, multi-person 3D pose estimation using monocular static or moving camera in real-world scenarios remains a challenge, either requiring large-scale training data or high computation complexity due to the high degrees of freedom in 3D human poses. We propose a novel scheme to effectively track and hierarchically estimate 3D human poses in natural videos in an efficient fashion. Without the need of using labelled 3D training data, we formulate torso estimation as a Perspective-N-Point (PNP) problem, and limb pose estimation as an optimization problem, and hierarchically structure the high dimensional poses to efficiently address the challenge. Experiments show good performance and high efficiency of multi-person 3D pose estimation on real-world videos, including street scenarios and various human daily activities from fixed and moving cameras, resulting in great new opportunities to understand and predict human behaviors. Renshu Gu, Gaoang Wang, Zhongyu Jiang, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Exploit the Connectivity: Multi-Object Tracking with TrackletNetabstractMulti-object tracking (MOT) is an important topic and critical task related to both static and moving camera applications, such as traffic flow analysis, autonomous driving and robotic vision. However, due to unreliable detection, occlusion and fast camera motion, tracked targets can be easily lost, which makes MOT very challenging. Most recent works exploit spatial and temporal information for MOT, but how to combine appearance and temporal features is still not well addressed. In this paper, we propose an innovative and effective tracking method called TrackletNet Tracker (TNT) that combines temporal and appearance information together as a unified framework. First, we define a graph model which treats each tracklet as a vertex. The tracklets are generated by associating detection results frame by frame with the help of the appearance similarity and the spatial consistency. To compensate camera movement, epipolar constraints are taken into consideration in the association. Then, for every pair of two tracklets, the similarity, called the connectivity in the paper, is measured by our designed multi-scale TrackletNet. Afterwards, the tracklets are clustered into groups and each group represents a unique object ID. Our proposed TNT has the ability to handle most of the challenges in MOT, and achieves promising results on MOT16 and MOT17 benchmark datasets compared with other state-of-the-art methods. Gaoang Wang, Yizhou Wang 0005, Haotian Zhang 0005, Renshu Gu, Jenq-Neng Hwang |
ACM Multimedia | 1 |
| 2019 | Eye in the Sky: Drone-Based Object Tracking and 3D LocalizationabstractDrones, or general UAVs, equipped with a single camera have been widely deployed to a broad range of applications, such as aerial photography, fast goods delivery and most importantly, surveillance. Despite the great progress achieved in computer vision algorithms, these algorithms are not usually optimized for dealing with images or video sequences acquired by drones, due to various challenges such as occlusion, fast camera motion and pose variation. In this paper, a drone-based multi-object tracking and 3D localization scheme is proposed based on the deep learning based object detection. We first combine a multi-object tracking method called TrackletNet Tracker (TNT) which utilizes temporal and appearance information to track detected objects located on the ground for UAV applications. Then, we are also able to localize the tracked ground objects based on the group plane estimated from the Multi-View Stereo technique. The system deployed on the drone can not only detect and track the objects in a scene, but can also localize their 3D coordinates in meters with respect to the drone camera. The experiments have proved our tracker can reliably handle most of the detected objects captured by drones and achieve favorable 3D localization performance when compared with the state-of-the-art methods. Haotian Zhang 0005, Gaoang Wang, Zhichao Lei, Jenq-Neng Hwang |
ACM Multimedia | 2 |
| 2019 | Uncertainty-Based Active Learning via Sparse Modeling for Image ClassificationabstractUncertainty sampling-based active learning has been well studied for selecting informative samples to improve the performance of a classifier. In batch-mode active learning, a batch of samples are selected for a query at the same time. The samples with top uncertainty are encouraged to be selected. However, this selection strategy ignores the relations among the samples, because the selected samples may have much redundant information with each other. This paper addresses this problem by proposing a novel method that combines uncertainty, diversity, and density via sparse modeling in the sample selection. We use sparse linear combination to represent the uncertainty of unlabeled pool data with Gaussian kernels, in which the diversity and density are well incorporated. The selective sampling method is proposed before optimization to reduce the representation error. To deal with ${l}_{0}$ norm constraint in the sparse problem, two approximated approaches are adopted for efficient optimization. Four image classification data sets are used for evaluation. Extensive experiments related to batch size, feature space, seed size, significant analysis, data transform, and time efficiency demonstrate the advantages of the proposed method. Gaoang Wang, Jenq-Neng Hwang, Craig S. Rose, Farron Wallace |
IEEE Trans. Image Process. | 1 |
| 2017 | Uncertainty sampling based active learning with diversity constraint by sparse selectionabstractUncertainty based active learning has been well studied for selecting informative samples to improve the performance of the classifier. One of the simplest strategy is that we always select samples with top largest uncertainties for a query. However, the selected samples may be very similar to each other, which results in little information added to update the classifier. In other words, we should avoid selecting similar samples for training the classifier. This paper addresses this problem by proposing a novel method using uncertainty based active learning algorithm with diversity constraint by sparse selection. First, uncertainty scores of unlabeled samples are obtained based on the previously trained support vector machine (SVM) classifiers. Then the sample selection is represented as a sparse modeling problem and optimal samples up to the pre-defined batch size are selected for a query. Besides that, two approximated approaches are proposed to solve the sparse problem via greedy search and quadratic programming (QP), respectively. After selection, the SVM classifiers are re-trained with new labeled data and the performance is tested on the testing dataset. We conduct several experiments on three image datasets for image classification task. The experimental results show the proposed method outperforms other four different methods and achieves promising performance. Gaoang Wang, Jenq-Neng Hwang, Craig S. Rose, Farron Wallace |
MMSP | 1 |
| 2017 | An Open-Source Platform for Underwater Image and Video AnalyticsabstractGlobal fisheries and the future of sustainable seafood are predicated on healthy populations of various species of fish and shellfish. Recent developments in the collection of large-volume optical data by autonomous underwater vehicles (AUVs), stationary camera arrays, and towed vehicles has made it possible for fishery scientists to generate species-specific, size-structured abundance estimates for different species of marine organisms via imagery. The immense volume of data collected by such devices quickly exceeds manual processing capacity and creates a strong need for automatic image analysis. This paper presents an open-source computer vision software platform designed to integrate common image and video analytics, such as stereo calibration, object detection and object classification, into a sequential data processing pipeline that is easy to program, multi-threaded, and generic. The system provides a cross-language common interface for each of these components, multiple implementations of each, as well as unified methods for evaluating and visualizing the results of different methods for accomplishing the same task. Matthew Dawkins, Linus Sherrill, Keith Fieldhouse, Anthony Hoogs, Benjamin L. Richards, Lakshman Prasad, Kresimir Williams, Nathan Lauffenburger, Gaoang Wang |
WACV | 10 |