EDBT 2026 Demo / reviewers in the wild / expert
Yi Wang 0068
dblp:17/221-68
· DBLP profile ↗
58ranked-venue papers
9as first author
53since 2021 · last 2026
0000-0001-8659-4724ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 21 since 2021Systems, architecture and hardware · 9 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern GenerationabstractTextile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated TPG, existing image-to-image models appear to be natural candidates for this task. However, when applied directly, these methods often produce unfaithful results, failing to preserve fine-grained details due to feature confusion between complex textile patterns and the inherent non-rigid texture distortions in clothing images. In this paper, we propose a novel method, SLDDM-TPG, for faithful and high-fidelity TPG. Our method consists of two stages: (1) a latent disentangled network (LDN) that resolves feature confusion in clothing representations and constructs a multi-dimensional, independent clothing feature space; and (2) a semi-supervised latent diffusion model (S-LDM), which receives guidance signals from LDN and generates faithful results through semi-supervised diffusion training, combined with our designed fine-grained alignment strategy. Extensive evaluations show that SLDDM-TPG reduces FID by 4.1 and improves SSIM by up to 0.116 on our CTP-HD dataset, and also demonstrate good generalization on the VITON-HD dataset. Chenggong Hu, Yi Wang 0068, Mengqi Xue, Haofei Zhang, Jie Song 0011 |
AAAI | 2 |
| 2026 | CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory AugmentationabstractWhile previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-image comprehension tasks. This limitation stems from their predominant reliance on text-based intermediate reasoning processes. While for human, when engaging in sophisticated multi-image analysis, they typically perform two complementary cognitive operations: (1) continuous cross-image visual comparison through region-of-interest matching, and (2) dynamic memorization of critical visual concepts throughout the reasoning chain. Motivated by these observations, we propose the Complex Multi-Modal Chain-of-Thought (CMMCoT) framework, a multi-step reasoning framework that mimics human-like "slow thinking" for multi-image understanding. Our approach incorporates two key innovations: (1) The construction of interleaved multimodal multi-step reasoning chains, which utilize critical visual region tokens, extracted from intermediate reasoning steps, as supervisory signals. This mechanism not only facilitates comprehensive cross-modal understanding but also enhances model interpretability. (2) The introduction of a test-time memory augmentation module that expands the model’s reasoning capacity during inference while preserving parameter efficiency. Furthermore, to facilitate research in this direction, we have curated a novel multi-image slow-thinking dataset. Extensive experiments demonstrate the effectiveness of our model. Yan Xia 0006, Mushui Liu, Zhelun Yu, Haoyuan Li 0002, Wanggui He, Dong She, Yi Wang 0068, Hao Jiang 0014 |
AAAI | 9 |
| 2026 | Evolutionary Negative Module Pruning for Better LoRA MergingabstractAnda Cao, Zhuo Gou, Yi Wang, Kaixuan Chen, Yu Wang, Can Wang, Mingli Song, Jie Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Anda Cao, Zhuo Gou, Yi Wang 0068, Kai-Xuan Chen 0001, Yu Wang 0176, Mingli Song, Jie Song 0011 |
ACL (1) | 3 |
| 2026 | Egocentric human-object interaction detection: A new benchmark and method
Kunyuan Deng, Yi Wang 0068, Lap-Pui Chau |
Expert Syst. Appl. | 2 |
| 2026 | CaRe-Ego: Contact-aware relationship modeling for egocentric interactive hand-object segmentation
Yuejiao Su, Yi Wang 0068, Lap-Pui Chau |
Expert Syst. Appl. | 2 |
| 2026 | LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance SegmentationabstractQuery-based 3D scene instance segmentation from point clouds has attained notable performance. However, existing methods suffer from the query initialization dilemma due to the sparse nature of point clouds and rely on computationally intensive attention mechanisms in query decoders. We accordingly introduceLaSSM, prioritizing simplicity and efficiency while maintaining competitive performance. Specifically, we propose a hierarchical semantic-spatial query initializer to derive the query set from superpoints by considering both semantic cues and spatial distribution, achieving comprehensive scene coverage and accelerated convergence. We further present a coordinate-guided state space model (SSM) decoder that progressively refines queries. The novel decoder features a local aggregation scheme that restricts the model to focus on geometrically coherent regions and a spatial dual-path SSM block to capture underlying dependencies within the query set by integrating associated coordinates information. Our design enables efficient instance prediction, avoiding the incorporation of noisy information and reducing redundant computation. LaSSM ranksfirst placeon the latest ScanNet++ V2 leaderboard, outperforming the previous best method by 2.5% mAP with only 1/3 FLOPs, demonstrating its superiority in challenging large-scale scene instance segmentation. LaSSM also achieves competitive performance on ScanNet V2, ScanNet200, S3DIS and ScanNet++ V1 benchmarks with less computational cost. Extensive ablation studies and qualitative results validate the effectiveness of our design. The code and weights are available at https://github.com/RayYoh/LaSSM. Yi Wang 0068, Yawen Cui, Moyun Liu, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Fuzzy-Aware Loss for Source-Free Domain Adaptation in Visual Emotion RecognitionabstractSource-free domain adaptation in visual emotion recognition (SFDA-VER) is a highly challenging task that re quires adapting VER models to the target domain without relying on source data, which is of great significance for data privacy protection. However, due to the unignorable disparities between visual emotion data and traditional image classification data, existing SFDA methods perform poorly on this task. In this paper, we investigate the SFDA-VER task from a fuzzy perspective and identify two key issues: fuzzy emotion labels and fuzzy pseudo-labels. These issues arise from the inherent uncertainty of emotion annotations and the potential mispredictions in pseudo labels. To address these issues, we propose a novel fuzzy aware loss (FAL) to enable the VER model to better learn and adapt to new domains under fuzzy labels. Specifically, FAL modifies the standard cross entropy loss and focuses on adjusting the losses of non-predicted categories, which prevents a large number of uncertain or incorrect predictions from overwhelming the VER model during adaptation. In addition, we provide a theoretical analysis of FAL and prove its robustness in handling the noise in generated pseudo-labels. Extensive experiments on 26 domain adaptation sub-tasks across three benchmark datasets demonstrate the effectiveness of our method. Code is available at: https://github.com/zhengyinghit/FAL. Ying Zheng 0009, Yiyi Zhang 0001, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | EVA02-AT: Egocentric Video-Language Understanding With Spatial-Temporal Rotary Positional Embeddings and Symmetric OptimizationabstractEgocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2) Ineffective spatial-temporal encoding due to manually split 3D rotary positional embeddings that hinder feature interactions, and 3) Imprecise learning objectives in soft-label multi-instance retrieval, which neglect negative pair correlations. In this paper, we introduce EVA02-AT, a suite of EVA02-based video-language foundation models tailored to egocentric video understanding tasks. EVA02-AT first efficiently transfers an image-based CLIP model into a unified video encoder via a single-stage pretraining. Second, instead of applying rotary positional embeddings to isolated dimensions, we introduce spatial-temporal rotary positional embeddings along with joint attention, which can effectively encode both spatial and temporal information on the entire hidden dimension. This joint encoding of spatial-temporal features enables the model to learn cross-axis relationships, which are crucial for accurately modeling motion and interaction in videos. Third, focusing on multi-instance video-language retrieval tasks, we introduce the Symmetric Multi-Similarity (SMS) loss and a novel training framework that advances all soft labels for both positive and negative pairs, providing a more precise learning objective. Extensive experiments on Ego4D, EPIC-Kitchens-100, and Charades-Ego under zero-shot and fine-tuning settings demonstrate that EVA02-AT achieves state-of-the-art performance across diverse egocentric video-language tasks with fewer parameters. Models with our SMS loss also show significant performance gains on multi-instance retrieval benchmarks. Our code and models are publicly available at https://github.com/xqwang14/EVA02-AT. Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Image Process. | 2 |
| 2026 | PromptSR: Cascade Prompting for Lightweight Image Super-ResolutionabstractAlthough the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-based self-attention modeling. The quadratic computational complexity relative to window size restricts its ability to use a large window size for expanding the receptive field while maintaining low computational costs. To address this challenge, we propose PromptSR, a novel prompt-empowered lightweight image SR method. The core component is the proposed cascade prompting block (CPB), which enhances global information access and local refinement via three cascaded prompting layers: a global anchor prompting layer (GAPL) and two local prompting layers (LPLs). The GAPL leverages downscaled features as anchors to construct low-dimensional anchor prompts (APs) through cross-scale attention, significantly reducing computational costs. These APs, with enhanced global perception, are then used to provide global prompts, efficiently facilitating long-range token connections. The two LPLs subsequently combine category-based self-attention and window-based self-attention to refine the representation in a coarse-to-fine manner. They leverage attention maps from the GAPL as additional global prompts, enabling them to perceive features globally at different granularities for adaptive local refinement. In this way, the proposed CPB effectively combines global priors and local details, significantly enlarging the receptive field while maintaining the low computational costs of our PromptSR. The experimental results demonstrate the superiority of our method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations. Our code will be released at https://github.com/wenyang001/PromptSR. Wenyang Liu, Jianjun Gao 0005, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 5 |
| 2025 | MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image SynthesisabstractAuto-regressive models have made significant progress in the realm of text-to-image synthesis, yet devising an appropriate model architecture and training strategy to achieve a satisfactory level remains an important avenue of exploration. In this work, we introduce MARS, a novel framework for T2I generation that incorporates a specially designed Semantic Vision-Language Integration Expert (SemVIE). This innovative component integrates pre-trained LLMs by independently processing linguistic and visual information—freezing the textual component while fine-tuning the visual component. This methodology preserves the NLP capabilities of LLMs while imbuing them with exceptional visual understanding. Building upon the powerful base of the pre-trained Qwen-7B, MARS stands out with its bilingual generative capabilities corresponding to both English and Chinese language prompts and the capacity for joint image and text generation. The flexibility of this framework lends itself to migration towards any-to-any task adaptability. Furthermore, MARS employs a multi-stage training strategy that first establishes robust image-text alignment through complementary bidirectional tasks and subsequently concentrates on refining the T2I generation process, significantly augmenting text-image synchrony and the granularity of image details. Notably, MARS requires only 9% of the GPU days needed by SD1.5, yet it achieves remarkable results across a variety of benchmarks, illustrating the training efficiency and the potential for swift deployment in various applications. Wanggui He, Siming Fu, Mushui Liu, Xierui Wang, Wenyi Xiao, Fangxun Shu, Yi Wang 0068, Lei Zhang 0006, Zhelun Yu, Haoyuan Li 0002, Ziwei Huang 0005, Leilei Gan, Hao Jiang 0014 |
AAAI | 7 |
| 2025 | ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric InteractionabstractEgocentric interaction perception is one of the essential branches in investigating human-environment interaction, which lays the basis for developing next-generation intelligent systems. However, existing egocentric interaction understanding methods cannot yield coherent textual and pixel-level responses simultaneously according to user queries, which lacks flexibility for varying downstream application requirements. To comprehend egocentric interactions exhaustively, this paper presents a novel task named Egocentric Interaction Reasoning and pixel Grounding (Ego-IRG). Taking an egocentric image with the query as input, Ego-IRG is the first task that aims to resolve the interactions through three crucial steps: analyzing, answering, and pixel grounding, which results in fluent textual and fine-grained pixel-level responses. Another challenge is that existing datasets cannot meet the conditions for the Ego-IRG task. To address this limitation, this paper creates the Ego-IRGBench dataset based on extensive manual efforts, which includes over 20k egocentric images with 1.6 million queries and corresponding multimodal responses about interactions. Moreover, we design a unified ANNEXE model to generate text- and pixel-level outputs utilizing multimodal large language models, which enables a comprehensive interpretation of egocentric interactions. The experiments on the Ego-IRGBench exhibit the effectiveness of our ANNEXE model compared with other works. Yuejiao Su, Yi Wang 0068, Qiongyang Hu, Chuang Yang 0003, Lap-Pui Chau |
CVPR | 2 |
| 2025 | OccProphet: Pushing the Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with an Observer-Forecaster-Refiner FrameworkabstractPredicting variations in complex traffic environments is crucial for the safety of autonomous driving. Recent advancements in occupancy forecasting have enabled forecasting future 3D occupied status in driving environments by observing historical 2D images. However, high computational demands make occupancy forecasting less efficient during training and inference stages, hindering its feasibility for deployment on edge agents. In this paper, we propose a novel framework, \textit{i.e.}, OccProphet, to efficiently and effectively learn occupancy forecasting with significantly lower computational requirements while improving forecasting accuracy. OccProphet comprises three lightweight components: Observer, Forecaster, and Refiner. The Observer extracts spatio-temporal features from 3D multi-frame voxels using the proposed Efficient 4D Aggregation with Tripling-Attention Fusion, while the Forecaster and Refiner conditionally predict and refine future occupancy inferences. Experimental results on nuScenes, Lyft-Level5, and nuScenes-Occupancy datasets demonstrate that OccProphet is both training- and inference-friendly. OccProphet reduces 58\%$\sim$78\% of the computational cost with a 2.6$\times$ speedup compared with the state-of-the-art Cam4DOcc. Moreover, it achieves 4\%$\sim$18\% relatively higher forecasting accuracy. Code and models are publicly available at https://github.com/JLChen-C/OccProphet. Huaiyuan Xu, Yi Wang 0068, Lap-Pui Chau |
ICLR | 3 |
| 2025 | Restoration of Bitstream-Corrupted Images: A Mamba-based Thumbnail-guided NetworkabstractThis paper investigates the real-world JPEG image restoration problem with bit errors on the compressed bitstream. To mimic the effect of bit errors encountered in real images, we automatically inject various bit errors to generate damaged images, thereby simulating the bitstream-corrupted JPEG images in real situations. The image restoration problem is proposed to recover these images caused by bit errors that conventional decoders cannot perfectly decode. Typically, when a bit stream containing bit errors is decoded by the robust decoder, the resulting image exhibits two distinct characteristics: color casts and block shifts. To solve those problems, we propose a Mamba-based thumbnail-guided network to address the impact of color casts and block shifts on the image. The proposed framework is structurally divided into three blocks. Firstly, we use a feature aggregation (FA) block to reassemble the information from the corrupted image and the thumbnail image into a more acceptable input format for the subsequent networks, allowing for self-adjustment of the input format. Then, we design a point-to-point restoration (PPR) block as a decoder to parse the features of inputs and thumbnails to generate coarse images. Finally, with the guidance of coarse images, a pyramid fusion (PF) block is used to generate the refined images. Extensive experimental results demonstrate our model outperforms state-of-the-art methods. Ablation studies and comparisons with super-resolution methods illustrate the effectiveness of our approach. The code will be available at https://github.com/HU1qy/MambaThumbnail. Qiongyang Hu, Yi Wang 0068, Lap-Pui Chau |
ISCAS | 2 |
| 2025 | From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionabstractRecent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edge device deployment. Meanwhile, conventional GSR models often lack generalization ability, falling short in recognizing unseen and rare situations. In this paper, we exploit transferring knowledge from a teacher MLLM to a small GSR model to enhance its generalization and zero-shot abilities, thereby introducing the task of Open-vocabulary Grounded Situation Recognition (Ov-GSR). To achieve this, we propose Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations. Specifically, the MIPD framework first leverages the LLM-based Judgmental Rationales Generator (JRG) to construct positive and negative glimpse and gaze rationales enriched with contextual semantic information. The proposed scene-aware and instance-perception prompts are then introduced to align rationales with visual information from the MLLM teacher via the Negative-Guided Multimodal Prompting Alignment (NMPA) module, effectively capturing holistic and perceptual multimodal knowledge. Finally, the aligned multimodal knowledge is distilled into the student Ov-GSR model, providing a stronger foundation for generalization that enhances situation understanding, bridges the gap between seen and unseen scenarios, and mitigates prediction bias in rare cases. We evaluate MIPD on the refined Ov-SWiG dataset, achieving superior performance on seen, rare, and unseen situations, and further demonstrate improved unseen detection on the HICO-DET dataset. Jianjun Gao 0005, Wenyang Liu, Kejun Wu, Yi Wang 0068, Soo Chin Liew |
ACM Multimedia | 7 |
| 2025 | Probabilistic Mixture of Hyperbolic Mamba for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) grapples with the dual challenge of learning new classes from minimal labeled training data while alleviating catastrophic forgetting of previous learned classes. Compared with previous methods employing static adaptation on specific parameters, current works verify that dynamic weights and sequence modeling in Selective State Space Models (SSMs) can capture distinctive feature drifts in FSCIL. However, the flattening operation in SSMs fragments the latent semantic relationship, where the resulting task isolation and representation degeneration are detrimental to FSCIL. Toward this issue, this paper presents a novel framework named Probabilistic Mixture of Hyperbolic State Space Experts (PmH-SSE) for FSCIL. First, since SSMs rely on scanning as an alternative to self-attention, the Hyperbolic state space model with multi-scale hybrid scan is built to facilitate few-shot learning by providing an extra Hyperbolic geometry that encodes hierarchical relationships. Moreover, we propose the probabilistic mixture of Mamba to increase the model's flexibility in handling non-stationary data streams in FSCIL and enhance the stability of high-parameter models in few-shot conditions. Finally, under the same experimental conditions, the proposed PmH-SSE demonstrates superior performance in comprehensive experiments. The codes are available at https://github.com/yawencui/PmH-SSE. Yawen Cui, Wenbin Zou, Huiping Zhuang, Yi Wang 0068, Lap-Pui Chau |
ACM Multimedia | 4 |
| 2025 | Towards Blind Bitstream-corrupted Video Recovery: A Visual Foundation Model-driven FrameworkabstractVideo signals are vulnerable in multimedia communication and storage systems, as even slight bitstream-domain corruption can lead to significant pixel-domain degradation. To recover faithful spatio-temporal content from corrupted inputs, bitstream-corrupted video recovery has recently emerged as a challenging and understudied task. However, existing methods require time-consuming and labor-intensive annotation of corrupted regions for each corrupted video frame, resulting in a large workload in practice. In addition, high-quality recovery remains difficult as part of the local residual information in corrupted frames may mislead feature completion and successive content recovery. In this paper, we propose the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with recovery model, which is adapted to different types of corruption and bitstream-level prompts. Within the framework, the proposed Detect Any Corruption (DAC) model leverages the rich priors of the visual foundation model while incorporating bitstream and corruption knowledge to enhance corruption localization and blind recovery. Additionally, we introduce a novel Corruption-aware Feature Completion (CFC) module, which adaptively processes residual contributions based on high-level corruption understanding. With VFM-guided hierarchical feature augmentation and high-level coordination in a mixture-of-residual-experts (MoRE) structure, our method suppresses artifacts and enhances informative residuals. Comprehensive evaluations show that the proposed method achieves outstanding performance in bitstream-corrupted video recovery without requiring a manually labeled mask sequence. The demonstrated effectiveness will help to realize improved user experience, wider application scenarios, and more reliable multimedia communication and storage systems. Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
ACM Multimedia | 4 |
| 2025 | GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian SplattingabstractThe significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts demonstrating promising performance, the model collapse and structural information deficiency remain prevalent due to insufficient point discrimination difficulty, yielding unreliable expressions and suboptimal performance. In this paper, we present GaussianCross, a novel cross-modal self-supervised 3D representation learning architecture integrating feed-forward 3D Gaussian Splatting (3DGS) techniques to address current challenges. GaussianCross seamlessly converts scale-inconsistent 3D point clouds into a unified cuboid-normalized Gaussian representation without missing details, enabling stable and generalizable pre-training. Subsequently, a tri-attribute adaptive distillation splatting module is incorporated to construct a 3D feature field, facilitating synergetic feature capturing of appearance, geometry, and semantic cues to maintain cross-modal consistency. To validate GaussianCross, we perform extensive evaluations on various benchmarks, including ScanNet, ScanNet200, and S3DIS. In particular, GaussianCross shows a prominent parameter and data efficiency, achieving superior performance through linear probing (<0.1% parameters) and limited data training (1% of scenes) compared to state-of-the-art methods. Furthermore, GaussianCross demonstrates strong generalization capabilities, improving the full fine-tuning accuracy by 9.3% mIoU and 6.1% AP50 on ScanNet200 semantic and instance segmentation tasks, respectively, supporting the effectiveness of our approach. The code, weights, and visualizations are publicly available at https://rayyoh.github.io/GaussianCross/. Yi Wang 0068, Moyun Liu, Lap-Pui Chau |
ACM Multimedia | 2 |
| 2025 | Semantic Representation Attack against Aligned Large Language ModelsabstractLarge Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from limited convergence, unnatural prompts, and high computational costs. We introduce semantic representation attacks, a novel paradigm that fundamentally reconceptualizes adversarial objectives against aligned LLMs. Rather than targeting exact textual patterns, our approach exploits the semantic representation space that can elicit diverse responses that share equivalent harmful meanings. This innovation resolves the inherent trade-off between attack effectiveness and prompt naturalness that plagues existing methods. Our Semantic Representation Heuristic Search (SRHS) algorithm efficiently generates semantically coherent adversarial prompts by maintaining interpretability during incremental search. We establish rigorous theoretical guarantees for semantic convergence and demonstrate that SRHS achieves unprecedented attack success rates (89.4% averaged across 18 LLMs, including 100% on 11 models) while significantly reducing computational requirements. Extensive experiments show that our method consistently outperforms existing approaches. Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau |
NeurIPS | 4 |
| 2025 | REAL: Representation enhanced analytic learning for exemplar-free class-incremental learning
Run He, Di Fang 0004, Yizhu Chen, Kai Tong, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau, Huiping Zhuang |
Knowl. Based Syst. | 6 |
| 2025 | PADetBench: Towards benchmarking texture- and patch-based physical attacks against object detection
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau |
Knowl. Based Syst. | 4 |
| 2025 | Foundation model-assisted interpretable vehicle behavior decision making
Shiyu Meng, Yi Wang 0068, Yawen Cui, Lap-Pui Chau |
Knowl. Based Syst. | 2 |
| 2025 | Open World Object Detection: A SurveyabstractExploring new knowledge is a fundamental human ability that can be mirrored in the development of deep neural networks, especially in the field of object detection. Open world object detection (OWOD) is an emerging area of research that adapts this principle to explore new knowledge. It focuses on recognizing and learning from objects absent from initial training sets, thereby incrementally expanding its knowledge base when new class labels are introduced. This survey paper offers a thorough review of the OWOD domain, covering essential aspects, including problem definitions, benchmark datasets, source codes, evaluation metrics, and a comparative study of existing methods. Additionally, we investigate related areas like open set recognition (OSR) and incremental learning (IL), underlining their relevance to OWOD. Finally, the paper concludes by addressing the limitations and challenges faced by current OWOD algorithms and proposes directions for future research. To our knowledge, this is the first comprehensive survey of the emerging OWOD field with over one hundred references, marking a significant step forward for object detection technology. A comprehensive source code and benchmarks are archived and concluded athttps://github.com/ArminLee/OWOD_Review. Yi Wang 0068, Dan Lin 0008, Kim-Hui Yap |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | SGIFormer: Semantic-Guided and Geometric-Enhanced Interleaving Transformer for 3D Instance SegmentationabstractIn recent years, transformer-based models have exhibited considerable potential in point cloud instance segmentation. Despite the promising performance achieved by existing methods, they encounter challenges such as instance query initialization problems and excessive reliance on stacked layers, rendering them incompatible with large-scale 3D scenes. This paper introduces a novel method, named SGIFormer, for 3D instance segmentation, which is composed of the Semantic-guided Mix Query (SMQ) initialization and the Geometric-enhanced Interleaving Transformer (GIT) decoder. Specifically, the principle of our SMQ initialization scheme is to leverage the predicted voxel-wise semantic information to implicitly generate the scene-aware query, yielding adequate scene prior and compensating for the learnable query set. Subsequently, we feed the formed overall query into our GIT decoder to alternately refine instance query and global scene features for further capturing fine-grained information and reducing complex design intricacies simultaneously. To emphasize geometric property, we consider bias estimation as an auxiliary task and progressively integrate shifted point coordinates embedding to reinforce instance localization. SGIFormer attains state-of-the-art performance on ScanNet V2, ScanNet200, S3DIS datasets, and the challenging high-fidelity ScanNet ++ benchmark, striking a balance between accuracy and efficiency. The code, weights, and demo videos are publicly available athttps://rayyoh.github.io/SGIFormer/. Yi Wang 0068, Moyun Liu, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | CAMCFormer: Cross-Attention and Multicorrelation Aided Transformer for Few-Shot Object Detection in Optical Remote Sensing ImagesabstractFew-shot object detection (FSOD) enables the detection of novel-class objects in remote sensing images (RSIs) with limited labeled samples. Although convolutional neural networks (CNNs) are commonly used for this task, they suffer from two inherent constraints. First, their limited local receptive field fails to capture global context within a single image and the relational dependencies between query and support images. Second, an additional feature alignment mechanism is typically required to bridge the gap between query and support images. To address these challenges, this work introduces a novel cross-attention and multicorrelation aided transformer (CAMCFormer) FSOD framework tailored for global feature representation and multicorrelation modeling in complex and large-scale RSIs. Specifically, a long-distance cross-attention module (LDCAM) is devised to capture dependencies between distant elements across query and support images at each feature extraction layer. This module facilitates the exchange of contextual information between images, resulting in more comprehensive feature representations and eliminating the need for separate feature alignment and fusion modules. Multicorrelation aided heads (MAHs) are constructed to enhance detection performance further to model various relational aspects, i.e., channel-correlation detection head (CCDH), spatial-correlation detection head (SCDH), and cross-attention detection head (CADH). These aided heads contribute to more robust and accurate classification and localization. Comprehensive experiments have been conducted, demonstrating the superiority of the proposed framework compared to several state-of-the-art detectors, highlighting its potential as an effective solution for FSOD in remote sensing scenarios. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han, Yan Feng 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian TrackingabstractMultiple pedestrian tracking is crucial for enhancing safety and efficiency in intelligent transport and autonomous driving systems by predicting movements and enabling adaptive decision-making in dynamic environments. It optimizes traffic flow, facilitates human interaction, and ensures compliance with regulations. However, it faces the challenge of tracking pedestrians in the presence of occlusion. Existing methods overlook effects caused by abnormal detections during partial occlusion. Subsequently, these abnormal detections can lead to inaccurate motion estimation, unreliable appearance features, and unfair association. To address these issues, we propose an adaptive occlusion-aware multiple pedestrian tracker, OccluTrack, to mitigate the effects caused by partial occlusion. Specifically, we first introduce a plug-and-play abnormal motion suppression mechanism into the Kalman Filter to adaptively detect and suppress outlier motions caused by partial occlusion. Second, we develop a pose-guided re-identification (Re-ID) module to extract discriminative part features for partially occluded pedestrians. Last, we develop a new occlusion-aware association method towards fair Intersection over Union (IoU) and appearance embedding distance measurement for occluded pedestrians. Extensive evaluation results demonstrate that our method outperforms state-of-the-art methods on MOTChallenge and DanceTrack datasets. Particularly, the performance improvements on IDF1 and ID Switches, as well as visualized results, demonstrate the effectiveness of our method in multiple pedestrian tracking. Jianjun Gao 0005, Yi Wang 0068, Kim-Hui Yap, Kratika Garg, Boon Siew Han |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | SignEye: Traffic Sign Interpretation From Vehicle First-Person ViewabstractTraffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navigation instructions. However, current works are limited to basic sign understanding without considering the egocentric vehicle’s spatial position, which fails to support further regulation assessment and direction navigation. Following the above issues, we introduce a new task: traffic sign interpretation from the vehicle’s first-person view, referred to asTSI-FPV. Meanwhile, we develop a traffic guidance assistant (TGA) scenario application to re-explore the role of traffic signs in ADS as a complement to popular autonomous technologies (such as obstacle perception). Notably, TGA is not a replacement for electronic map navigation; rather, TGA can be an automatic tool for updating it and complementing it in situations such as offline conditions or temporary sign adjustments. Lastly, a spatial and semantic logic-aware stepwise reasoning pipeline (SignEye) is constructed to achieve the TSI-FPV and TGA, and an application-specific dataset (Traffic-CN) is built. Experiments show that TSI-FPV and TGA are achievable via our SignEye trained on Traffic-CN. The results also demonstrate that the TGA can provide complementary information to ADS beyond existing popular autonomous technologies. Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Yuejiao Su, Junyu Gao 0001, Hongyuan Zhang 0001, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | ByteNet: Rethinking Multimedia File Fragment Classification Through Visual PerspectivesabstractMultimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multimedia storage and communication. Existing MFFC methods typically treat fragments as 1D byte sequences and emphasize the relations between separate bytes (interbytes) for classification. However, the more informative relations inside bytes (intrabytes) are overlooked and seldom investigated. By looking inside bytes, the bit-level details of file fragments can be accessed, enabling a more accurate classification. Motivated by this, we first proposeByte2Image, a novel visual representation model that incorporates previously overlooked intrabyte information into file fragments and reinterprets these fragments as 2D grayscale images. This model involves a sliding byte window to reveal the intrabyte information and a rowwise stacking of intrabyte n-grams for embedding fragments into a 2D space. Thus, complex interbyte and intrabyte correlations can be mined simultaneously using powerful vision networks. Additionally, we propose an end-to-end dual-branch networkByteNetto enhance robust correlation mining and feature representation. ByteNet makes full use of the raw 1D byte sequence and the converted 2D image through a shallow byte branch feature extraction (BBFE) and a deep image branch feature extraction (IBFE) network. In particular, the BBFE, composed of a single fully-connected layer, adaptively recognizes the co-occurrence of several some specific bytes within the raw byte sequence, while the IBFE, built on a vision Transformer, effectively mines the complex interbyte and intrabyte correlations from the converted image. Experiments on the two representative benchmarks, including 14 cases, validate that our proposed method outperforms state-of-the-art approaches on different cases by up to 12.2%. Wenyang Liu, Kejun Wu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
IEEE Trans. Multim. | 4 |
| 2025 | WHANet:Wavelet-Based Hybrid Asymmetric Network for Spectral Super-Resolution From RGB InputsabstractThe reconstruction from three to dozens of spectral bands, known as spectral super resolution (SSR) has achieved remarkable progress with the continuous development of deep learning. However, the reconstructed hyperspectral images (HSIs) still suffer from the spatial degeneration due to the insufficient retention of high-frequency (HF) information during the SSR process. To remedy this issue, a novel Wavelet-based Hybrid Asymmetric Network (WHANet) is proposed to establish a RGB-to-HSI translation in wavelet domain, thus reserving and emphasizing the HF features in hyperspectral space. Basically, the backbone is designed in a hybrid asymmetric structure that learns the exact representations of decomposed wavelet coefficients in hyperspectral domain in a parallel way. Innovatively, a CNN-based HF reconstruction module (HFRM) and a transformer-based low frequency (LF) reconstruction module (LFRM) are delicately devised to perform the SSR process individually, which are able to process the discriminative wavelet coefficients contrapuntally. Furthermore, a hybrid loss function incorporated with the Fast Fourier loss (FFL) is proposed to directly regularize and emphasis the missing HF components. Eventually, experimental results over three benchmark datasets and one remote sensing dataset demonstrate that our WHANet is able to reach the state-of-the-art performance quantitatively and qualitatively. Nan Wang 0026, Shaohui Mei, Yi Wang 0068, Yifan Zhang 0006, Duo Zhan |
IEEE Trans. Multim. | 3 |
| 2025 | 3DGeoDet: General-Purpose Geometry-Aware Image-Based 3D Object DetectionabstractThis paper proposes 3DGeoDet, a novel geometry-aware 3D object detection approach that effectively handles single- and multi-view RGB images in indoor and outdoor environments, showcasing its general-purpose applicability. The key challenge for image-based 3D object detection tasks is the lack of 3D geometric cues, which leads to ambiguity in establishing correspondences between images and 3D representations. To tackle this problem, 3DGeoDet generates efficient 3D geometric representations in both explicit and implicit manners based on predicted depth information. Specifically, we utilize the predicted depth to learn voxel occupancy and optimize the voxelized 3D feature volume explicitly through the proposed voxel occupancy attention. To further enhance 3D awareness, the feature volume is integrated with an implicit 3D representation, the truncated signed distance function (TSDF). Without requiring supervision from 3D signals, we significantly improve the model's comprehension of 3D geometry by leveraging intermediate 3D representations and achieve end-to-end training. Our approach surpasses the performance of state-of-the-art image-based methods on both single- and multi-view benchmark datasets across diverse environments, achieving a 9.3 [email protected] improvement on the SUN RGB-D dataset, a 3.3 [email protected] improvement on the ScanNetV2 dataset, and a 0.19$\text{AP}_{\text{3D}}[email protected] improvement on the KITTI dataset. The project page is available at:https://cindy0725.github.io/3DGeoDet/ Yi Wang 0068, Yawen Cui, Lap-Pui Chau |
IEEE Trans. Multim. | 2 |
| 2024 | Weakly-Supervised Crowd Counting with Token Attention and Fusion: A Simple and Effective BaselineabstractConventional crowd counting methods exploit a large number of point annotations to train regression-based neural networks for density map estimation. However, laborious point annotations of human heads (strong supervision) are required in training. This paper presents a simple and effective crowd counting method with only image-level count annotations, i.e., the number of people in an image (weak supervision). Specifically, we first investigate three backbone networks and find the significance of the global information extracted by self-attention for weakly-supervised crowd counting. Then, we propose an effective network composed of a Transformer backbone and token channel attention module (T-CAM) in the counting head, where the attention in channels of tokens can compensate for the self-attention between tokens of the Transformer. Finally, a simple token fusion is proposed to obtain global information. Experimental results on two representative crowd counting benchmarks show the superiority of the proposed method, with an average 10% relative improvement compared with baselines. The code is publicly available at https://github.com/WangyiNTU/WSCC_TAF. Yi Wang 0068, Qiongyang Hu, Lap-Pui Chau |
ICASSP | 1 |
| 2024 | GemNet: Analysis and Prediction of Building Materials for Optimizing Indoor Wireless NetworksabstractThis paper investigates the correlation between building material properties and indoor network coverage, encompassing both indoor Wi-Fi and outdoor 5G technologies to provide customized network services tailored to users' needs in diverse areas. We first analyze the impact of building material characteristics, with a special focus on wall materials, on the distribution of wireless signal propagation. Then, a ray-tracing-based method is introduced to synthetically generate high-quality training data that covers fine-grained network scenarios with a wide range of wall materials, extending beyond traditional materials. This dataset serves as the foundation for our proposed Global Embedding Isomorphism Network (GemNet), a machine learning framework that facilitates the prediction of optimal material parameters for customized in-building coverage. This innovation enables architects and builders to design novel, network-friendly materials, ensuring ubiquitous and on-demand network services. Extensive evaluations consistently demonstrate a re-markable prediction accuracy of 90.52% on material parameters, underscoring the framework's ability to optimize indoor wireless network planning through the lens of material engineering. Zhijin Yang, Zhizhen Li, Yi Wang 0068, Jianqing Liu, Mingzhe Chen, Yuchen Liu 0001 |
ICC | 3 |
| 2024 | Temporal Sentence Grounding with Temporally Global Textual KnowledgeabstractTemporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent domain gap between different modalities. In this paper, we utilize pseudo-query features containing extensive temporally global textual knowledge sourced from the same video-query pair, to enhance the bridging of domain gaps and attain a heightened level of similarity between multi-modal features. Specifically, we propose a Pseudo-query Intermediary Network (PIN) to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning. Subsequently, we utilize learnable prompts to encapsulate the knowledge of pseudo-queries, propagating them into the textual encoder and multimodal fusion module, further enhancing the feature alignment between visual and language for better temporal grounding. Extensive experiments conducted on the Charades-STA and ActivityNet-Captions datasets demonstrate the effectiveness of our method. Runzhong Zhang, Jianjun Gao 0005, Kejun Wu, Kim-Hui Yap, Yi Wang 0068 |
ICME | 6 |
| 2024 | Multi-scale Attentive Fusion Network for Remote Sensing Image Change CaptioningabstractRemote-sensing Image Change Captioning (RSICC) aims to automatically generate sentences describing the difference of content in remote-sensing bitemporal images. Most of the methods often address shortcomings in model architecture to enhance previous work, overlooking the distinctive characteristics that set remote sensing images apart from natural images, such as recognizing the change of objects with various scales (e.g., small/large-scale objects). By considering the difference, we proposed a Multi-scale Attentive Fusion Network (MAF-Net) to adaptively capture and describe the object change with a wide range of scales. The MAF-Net first extracts multi-scale visual features of bitemporal images from different stages of the CNN backbone, then captures the changes in each pair of the features with the proposed Multi-scale Change Aware Encoders (MCAE). Specifically, the MCAE captures the change-aware discriminative information over the paired multi-scale bitemporal features by Transformer-based different and content cross-attention encoding. Furthermore, a Gated Attentive Fusion (GAF) module is introduced to adaptively aggregate the relevant change-aware features to enhance the change caption performance. We evaluate the effectiveness of our proposed method on two RSICC datasets (e.g., LEVIR-CC and LEVIRCCD), and experimental results demonstrate that our method achieves state-of-the-art performance. Yi Wang 0068, Kim-Hui Yap |
ISCAS | 2 |
| 2024 | Depth-powered Moving-obstacle Segmentation Under Bird-eye-view for Autonomous DrivingabstractSensing the moving obstacles accurately under birdeye view (BEV) is the foundation for reliable autonomous driving, providing straightforward information for the downstream tasks. However, accurately segmenting moving obstacles only through monocular camera views is extremely difficult due to the lack of depth information. It can easily generate the projected depth information from point clouds, but its sparsity provides incomplete depth information. Therefore, in this paper, we propose a dense depth-powered framework, dubbed DPMoSeg, to generate dense moving-obstacle segmentation observations under BEV space. To better represent the depth prediction, we design a sparse-dense attention module to fully combine the knowledge across non- homogeneous and homogeneous representations. The experimental results demonstrate the effectiveness and superiority of our proposed framework. Shiyu Meng, Yi Wang 0068, Lap-Pui Chau |
ISCAS | 2 |
| 2024 | Few-shot Class-agnostic Counting with Occlusion Augmentation and LocalizationabstractMost existing few-shot class-agnostic counting (FCAC) methods follow the extract-and-compare pipeline to count all instances of an arbitrary category in the query image given a few exemplars. However, these methods generate the density map rather than the exact instance location for counting, which is less intuitive and accurate than the latter. Besides, how to alleviate the problem of occlusion is ignored in most existing work. To solve the above problems, this paper proposes an Occlusion-Augmented Localization Network (OALNet), which extracts multiple occluded features of exemplars for comparison and utilizes the precise position of instances for more accurate and confident counting results. Specifically, the OALNet is in an extract-and-attention manner. It includes an Occluded Feature Generation module to deal with the occlusion problem in query images. Besides, the OALNet adopts the Feature Attention module to improve the extracted feature by self-attention and model the relationship between the exemplar features and query features by cross-attention. Compared with other FCAC methods, experimental results demonstrate that the proposed OALNet achieves superior performance. Yuejiao Su, Yi Wang 0068, Lap-Pui Chau |
ISCAS | 2 |
| 2024 | F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental LearningabstractOnline Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach—Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL’s robust performance in OCIL scenarios. Code is available at: https://github.com/liuyuchen-cz/F-OAL Huiping Zhuang, Yuchen Liu 0001, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau |
NeurIPS | 7 |
| 2024 | Top-down framework for weakly-supervised grounded image captioning
Suchen Wang, Kim-Hui Yap, Yi Wang 0068 |
Knowl. Based Syst. | 4 |
| 2024 | Intra- and inter-sector contextual information fusion with joint self-attention for file fragment classificationabstractFile fragment classification (FFC) aims to identify the file type of file fragments in memory sectors, which is of great importance in memory forensics and information security. Existing works focused on processing the bytes within sectors separately and ignoring contextual information between adjacent sectors. In this paper, we introduce a joint self-attention network (JSANet) for FFC to learn intra-sector local features and inter-sector contextual features. Specifically, we propose an end-to-end network with the byte, channel, and sector self-attention modules. Byte self-attention adaptively recognizes the intra-sector significant bytes, and channel self-attention re-calibrates the features between channels. Based on the insight that adjacent memory sectors are most likely to store a file fragment, sector self-attention leverages contextual information in neighboring sectors to enhance inter-sector feature representation. Extensive experiments on seven FFC benchmarks show the superiority of our method compared with state-of-the-art methods. Moreover, we construct VFF-16, a variable-length file fragment dataset to reflect file fragmentation. Integrated with sector self-attention, our method improves accuracy by more than 16.3% against the baseline on VFF-16, and the runtime achieves 5.1 s/GB with GPU acceleration. In addition, we extend our model to malware detection and show its applicability. Yi Wang 0068, Wenyang Liu, Kejun Wu, Kim-Hui Yap, Lap-Pui Chau |
Knowl. Based Syst. | 1 |
| 2024 | A three-stream fusion and self-differential attention network for multi-modal crowd counting
Haihan Tang, Yi Wang 0068, Zhiping Lin 0001, Lap-Pui Chau, Huiping Zhuang |
Pattern Recognit. Lett. | 2 |
| 2024 | Few-Shot Object Detection With Multilevel Information Interaction for Optical Remote Sensing ImagesabstractMetalearning has been widely applied to solve the few-shot object detection (FSOD) problem in natural scenes, which performs similarity measurement and information aggregation of the support set and the query set. However, regarding remote sensing images (RSIs), many difficulties caused by their disparities need to be further addressed, such as inconsistencies in imaging scale, direction, and background between support and query images. These result in feature misalignment and attention bias, interfering with model performance. In this article, a multilevel information interaction (MLII) strategy is proposed for FSOD to alleviate feature misalignment and attention bias. Information interactions are conducted within multiple scales of features and highlight similar regions of query and support features. A semantic enhancement module (SEM) is proposed to assist MLII in extracting key information and achieving more discriminative feature representation. Moreover, a feature cross-aggregation module (FCM) with separate classification losses is designed to train the detector to identify objects that coexist in query and support images. Extensive experiments demonstrate that the proposed method outperforms several state-of-the-art few-shot object detectors over commonly used benchmark datasets, i.e., DIOR and NWPU-10. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Bitstream-Corrupted JPEG Images are Restorable: Two-stage Compensation and Alignment Framework for Image RestorationabstractIn this paper, we study a real-world JPEG image restoration problem with bit errors on the encrypted bitstream. The bit errors bring unpredictable color casts and block shifts on decoded image contents, which cannot be resolved by existing image restoration methods mainly relying on pre-defined degradation models in the pixel domain. To address these challenges, we propose a robust JPEG decoder, followed by a two-stage compensation and alignment framework to restore bitstream-corrupted JPEC images. Specifically, the robust JPEC decoder adopts an error-resilient mechanism to decode the corrupted JPEG bitstream. The two-stage framework is composed of the self-compensation and alignment (SCA) stage and the guided-compensation and alignment (GCA) stage. The SCA adaptively performs block-wise image color compensation and alignment based on the estimated color and block offsets via image content similarity. The GCA leverages the extracted low-resolution thumbnail from the JPEG header to guide full-resolution pixel-wise image restoration in a coarse-to-fine manner. It is achieved by a coarse-guided pix2pix network and a refine-guided bi-directional Laplacian pyramid fusion network. We conduct experiments on three benchmarks with varying degrees of bit error rates. Experimental results and ablation studies demonstrate the superiority of our proposed method. The code will be released at https://github.com/wenyang001/Two-ACIR. Wenyang Liu, Yi Wang 0068, Kim-Hui Yap, Lap-Pui Chau |
CVPR | 2 |
| 2023 | A Spatial-Focal Error Concealment Scheme for Corrupted Focal Stack VideoabstractFocal stack image sequences can be regarded as successive frames of videos, which are densely captured by focusing on a stack of focal planes. This type of data is able to provide focus cues for display technologies. Before the displays on the user side, focal stack video is possibly corrupted during compression, storage and transmission chains, generating error frames on the decoder side. The error regions are difficult to be recovered due to the focal changes among frames. Conventional error concealment methods result in sharpness inconsistency between recovered regions and their spatial adjacent regions. Motivated by this, in this paper, we propose a spatial-focal error concealment scheme specialized for focal stack videos. The spatial adjacent regions around an error region are employed to reveal the prediction relations between error frame and focal adjacent frames. Gaussian blur filtering and Lucy-Richardson deblur filtering are applied to simulate the video focal changes. In this way, the error regions can be well recovered by exploiting the spatial-focal information. Experiment results show that the proposed scheme can achieve the highest objective quality in terms of PSNR and SSIM. It can also obtain the best subjective quality with sharpness consistency in recovered regions and without block effect. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
DCC | 2 |
| 2023 | Image Representation and Deep Inception-Attention for File-type and Malware ClassificationabstractFile-type classification aims to recognize the file types of files/fragments without file-system metadata, which is essential for memory forensics and data recovery. In this paper, we introduce an image representation and deep inception-attention manner for file-type classification. Specifically, we consider file-type classification as an image classification problem. Raw data sequences in the memory block are converted to 2D binary images, enriching the representation ability and visualization while retaining the completeness of the bitstream. With binary images as inputs, we propose a deep inception-attention network to extract discriminate horizontal features and re-calibrate the weights of feature maps, and finally, predict file types. Experiments on a large-scale benchmark show the superiority of the proposed model. Moreover, our method can be extended to a similar application, like malware classification, and achieve outstanding performance. Yi Wang 0068, Kejun Wu, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
ISCAS | 1 |
| 2023 | METFormer: A Motion Enhanced Transformer for Multiple Object TrackingabstractMultiple object tracking (MOT) is an important task in computer vision, especially video analytics. Transformer-based methods are emerging approaches using both tracking and detection queries. However, motion modeling in existing transformer-based methods lacks effective association capability. Thus, this paper introduces a new METFormer model, a Motion Enhanced TransFormer-based tracker with a novel global-local motion context learning technique to mitigate the lack of motion information in existing transformer-based methods. The global-local motion context learning technique first centers on difference-guided global motion learning to obtain temporal information from adjacent frames. Based on global motion, we leverage context-aware local object motion modelling to study motion patterns and enhance the feature representation for individual objects. Experimental results on the benchmark MOT17 dataset show that our proposed method can surpass the state-of-the-art Trackformer [21] by 1.8% on IDF1 and 21.7% on ID Switches under public detection settings. Jianjun Gao 0005, Kim-Hui Yap, Yi Wang 0068, Kratika Garg, Boon Siew Han |
ISCAS | 3 |
| 2023 | Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and MethodabstractThe past decade has witnessed great strides in video recovery by specialist technologies, like video inpainting, completion, and error concealment. However, they typically simulate the missing content by manual-designed error masks, thus failing to fill in the realistic video loss in video communication (e.g., telepresence, live streaming, and internet video) and multimedia forensics. To address this, we introduce the bitstream-corrupted video (BSCV) benchmark, the first benchmark dataset with more than 28,000 video clips, which can be used for bitstream-corrupted video recovery in the real world. The BSCV is a collection of 1) a proposed three-parameter corruption model for video bitstream, 2) a large-scale dataset containing rich error patterns, multiple corruption levels, and flexible dataset branches, and 3) a new video recovery framework that serves as a benchmark. We evaluate state-of-the-art video inpainting methods on the BSCV dataset, demonstrating existing approaches' limitations and our framework's advantages in solving the bitstream-corrupted video recovery problem. The benchmark and dataset are released at https://github.com/LIUTIGHE/BSCV-Dataset. Kejun Wu, Yi Wang 0068, Wenyang Liu, Kim-Hui Yap, Lap-Pui Chau |
NeurIPS | 3 |
| 2023 | SSN: Stockwell Scattering Network for SAR Image Change DetectionabstractRecently, synthetic aperture radar (SAR) image change detection has become an interesting yet challenging direction due to the presence of speckle noise. Although both traditional and modern learning-driven methods attempted to overcome this challenge, deep convolutional neural networks (DCNNs)-based methods are still hindered by the lack of interpretability and the requirement of large computation power. To overcome this drawback, wavelet scattering network (WSN) and Fourier scattering network (FSN) are proposed. Combining respective merits of WSN and FSN, we propose Stockwell scattering network (SSN) based on Stockwell transform (ST), which is widely applied against noisy signals and shows advantageous characteristics in speckle reduction. The proposed SSN provides noise-resilient feature representation and obtains state-of-the-art performance in SAR image change detection as well as high computational efficiency. Experimental results on three real SAR image datasets demonstrate the effectiveness of the proposed method. Gong Chen 0003, Yanan Zhao 0003, Yi Wang 0068, Kim-Hui Yap |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Moving Towards Centers: Re-Ranking With Attention and Memory for Re-IdentificationabstractRe-ranking utilizes contextual information to optimize the initial ranking list of person or vehicle re-identification (re-ID), which boosts the retrieval performance at post-processing steps. This paper proposes a re-ranking network to predict the correlations between the probe and top-ranked neighbor samples. Specifically, all the feature embeddings of query and gallery images are expanded and enhanced by a linear combination of their neighbors, with the correlation prediction serving as discriminative combination weights. The combination process is equivalent to moving independent embeddings toward the identity centers, improving cluster compactness. For correlation prediction, we first aggregate the contextual information for probe's$k$-nearest neighbors via the Transformer encoder. Then, we distill and refine the probe-related features into the Contextual Memory cell via attention mechanism. Like humans that retrieve images by not only considering probe images but also memorizing the retrieved ones, the Contextual Memory produces multi-view descriptions for each instance. Finally, the neighbors are reconstructed with features fetched from the Contextual Memory, and a binary classifier predicts their correlations with the probe. Experiments on six widely-used person and vehicle re-ID benchmarks demonstrate the effectiveness of the proposed method. Especially, our method surpasses the state-of-the-art re-ranking approaches on large-scale datasets by a significant margin, i.e., with an average 4.83% CMC@1 and 14.83% mAP improvements on VERI-Wild, MSMT17, and VehicleID datasets. Yunhao Zhou, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Multim. | 2 |
| 2022 | TAFNet: A Three-Stream Adaptive Fusion Network for RGB-T Crowd CountingabstractIn this paper, we propose a three-stream adaptive fusion network named TAFNet, which uses paired RGB and thermal images for crowd counting. Specifically, TAFNet is divided into one main stream and two auxiliary streams. We combine a pair of RGB and thermal images to constitute the input of main stream. Two auxiliary streams respectively exploit RGB image and thermal image to extract modality-specific features. Besides, we propose an Information Improvement Module (IIM) to fuse the modality-specific features into the main stream adaptively. Experiment results on RGBT-CC dataset show that our method achieves more than 20% improvement on mean average error and root mean squared error compared with state-of-the-art method. The source code will be publicly available at https://github.com/TANGHAIHAN/TAFNet. Haihan Tang, Yi Wang 0068, Lap-Pui Chau |
ISCAS | 2 |
| 2022 | Weakly-Supervised Part-Attention and Mentored Networks for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) aims to retrieve images with the same vehicle ID across different cameras. Current part-level feature learning methods typically detect vehicle parts via uniform division, outside tools, or attention modeling. However, such part features often require expensive additional annotations and cause sub-optimal performance in case of unreliable part mask predictions. In this paper, we propose a weakly-supervised Part-Attention Network (PANet) and Part-Mentored Network (PMNet) for Vehicle Re-ID. Firstly, PANet localizes vehicle parts via part-relevant channel recalibration and cluster-based mask generation without vehicle part supervisory information. Secondly, PMNet leverages teacher-student guided learning to distill vehicle part-specific features from PANet and performs multi-scale global-part feature extraction. During inference, PMNet can adaptively extract discriminative part features without part localization by PANet, preventing unstable part mask predictions. We address this Re-ID issue as a multi-task problem and adopt Homoscedastic Uncertainty to learn the optimal weighing of ID losses. Experiments are conducted on two public benchmarks, showing that our approach outperforms recent methods, which require no extra annotations by an average increase of 3.0% in CMC@5 on VehicleID and over 1.4% in mAP on VeRi776. Moreover, our method can extend to the occluded vehicle Re-ID task and exhibits good generalization ability. Lisha Tang, Yi Wang 0068, Lap-Pui Chau |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Rethinking and Designing a High-Performing Automatic License Plate Recognition ApproachabstractIn this paper, we propose a real-time and accurate automatic license plate recognition (ALPR) approach. Our study illustrates the outstanding design of ALPR with four insights: (1) the resampling-based cascaded framework is beneficial to both speed and accuracy; (2) the highly efficient license plate recognition should abundant additional character segmentation and recurrent neural network (RNN), but adopt a plain convolutional neural network (CNN); (3) in the case of CNN, taking advantage of vertex information on license plates improves the recognition performance; and (4) the weight-sharing character classifier addresses the lack of training images in small-scale datasets. Based on these insights, we propose a novel ALPR approach, termed VSNet. Specifically, VSNet includes two CNNs, i.e., VertexNet for license plate detection and SCR-Net for license plate recognition, integrated in a resampling-based cascaded manner. In VertexNet, we propose an efficient integration block to extract the spatial features of license plates. With vertex supervisory information, we propose a vertex-estimation branch in VertexNet such that license plates can be rectified as the input images of SCR-Net. In SCR-Net, we introduce a horizontal encoding technique for left-to-right feature extraction and propose a weight-sharing classifier for character recognition. Experimental results show that the proposed VSNet outperforms state-of-the-art methods by more than 50% relative improvement on error rate, achieving >99% recognition accuracy on CCPD and AOLP datasets with 149 FPS inference speed. Moreover, our method illustrates an outstanding generalization capability when evaluated on the unseen PKUData and CLPD datasets. Yi Wang 0068, Zhen-Peng Bian, Yunhao Zhou, Lap-Pui Chau |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Fully Decoupled Neural Network Learning Using Delayed GradientsabstractTraining neural networks with backpropagation (BP) requires a sequential passing of activations and gradients. This has been recognized as the lockings (i.e., the forward, backward, and update lockings) among modules (each module contains a stack of layers) inherited from the BP. In this brief, we propose a fully decoupled training scheme using delayed gradients (FDG) to break all these lockings. The FDG splits a neural network into multiple modules and trains them independently and asynchronously using different workers (e.g., GPUs). We also introduce a gradient shrinking process to reduce the stale gradient effect caused by the delayed gradients. Our theoretical proofs show that the FDG can converge to critical points under certain conditions. Experiments are conducted by training deep convolutional neural networks to perform classification tasks on several benchmark data sets. These experiments show comparable or better results of our approach compared with the state-of-the-art methods in terms of generalization and acceleration. We also show that the FDG is able to train various networks, including extremely deep ones (e.g., ResNet-1202), in a decoupled fashion. Huiping Zhuang, Yi Wang 0068, Qinglai Liu, Zhiping Lin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | A Self-Training Approach for Point-Supervised Object Detection and Counting in CrowdsabstractIn this article, we propose a novel self-training approach named Crowd-SDNet that enables a typical object detector trained only with point-level annotations (i.e., objects are labeled with points) to estimate both the center points and sizes of crowded objects. Specifically, during training, we utilize the available point annotations to supervise the estimation of the center points of objects directly. Based on a locally-uniform distribution assumption, we initialize pseudo object sizes from the point-level supervisory information, which are then leveraged to guide the regression of object sizes via a crowdedness-aware loss. Meanwhile, we propose a confidence and order-aware refinement scheme to continuously refine the initial pseudo object sizes such that the ability of the detector is increasingly boosted to detect and count objects in crowds simultaneously. Moreover, to address extremely crowded scenes, we propose an effective decoding method to improve the detector's representation ability. Experimental results on the WiderFace benchmark show that our approach significantly outperforms state-of-the-art point-supervised methods under both detection and counting tasks, i.e., our method improves the average precision by more than 10% and reduces the counting error by 31.2%. Besides, our method obtains the best results on the crowd counting and localization datasets (i.e., ShanghaiTech and NWPU-Crowd) and vehicle counting datasets (i.e., CARPK and PUCPR+) compared with state-of-the-art counting-by-detection methods. The code will be publicly available at https://github.com/WangyiNTU/Point-supervised-crowd-detection. Yi Wang 0068, Junhui Hou, Xinyu Hou, Lap-Pui Chau |
IEEE Trans. Image Process. | 1 |
| 2021 | Convolutional Neural Networks With Dynamic RegularizationabstractRegularization is commonly used for alleviating overfitting in machine learning. For convolutional neural networks (CNNs), regularization methods, such as DropBlock and Shake-Shake, have illustrated the improvement in the generalization performance. However, these methods lack a self-adaptive ability throughout training. That is, the regularization strength is fixed to a predefined schedule, and manual adjustments are required to adapt to various network architectures. In this article, we propose a dynamic regularization method for CNNs. Specifically, we model the regularization strength as a function of the training loss. According to the change of the training loss, our method can dynamically adjust the regularization strength in the training procedure, thereby balancing the underfitting and overfitting of CNNs. With dynamic regularization, a large-scale model is automatically regularized by the strong perturbation, and vice versa. Experimental results show that the proposed method can improve the generalization capability on off-the-shelf network architectures and outperform state-of-the-art regularization methods. Yi Wang 0068, Zhen-Peng Bian, Junhui Hou, Lap-Pui Chau |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Vehicle Tracking Using Deep SORT with Low Confidence Track FilteringabstractMulti-object tracking (MOT) becomes an attractive topic due to its wide range of usability in video surveillance and traffic monitoring. Recent improvements on MOT has focused on tracking-by-detection manner. However, as a relatively complicated and integrated computer vision mission, state-of-the-art tracking-by-detection techniques are still suffering from issues such as a large number of false-positive tracks. To reduce the effect of unreliable detections on vehicle tracking, in this paper, we propose to incorporate a low confidence track filtering into the Simple Online and Realtime Tracking with a Deep association metric (Deep SORT) algorithm. We present a self-generated UA-DETRAC vehicle re-identification dataset which can be used to train the convolutional neural network of Deep SORT for data association. We evaluate our proposed tracker on UA-DETRAC test dataset. Experimental results show that the proposed method can improve the original Deep SORT algorithm with a significant margin. Our tracker outperforms the state-of-the-art online trackers and is comparable with batch-mode trackers. Xinyu Hou, Yi Wang 0068, Lap-Pui Chau |
AVSS | 2 |
| 2019 | Object Counting in Video Surveillance Using Multi-scale Density Map RegressionabstractIn this paper, we present an effective convolutional neural network (CNN) for object counting in video surveillance, namely multi-scale density map regressor (MSDMR). In contrast to existing CNN-based methods that achieve high accuracy by means of empirically increasing the model capacity with more complex structures/layers, we focus on a compact CNN. Specifically, the MSDMR is mainly designed with the supervision of multi-scale outputs, in which two CNN stacks estimate coarse- and fine-scale density maps, respectively. The integral of the fine density map provides the count of objects. The two stacks are connected in a cascaded manner and jointly trained such that the overall model can learn discriminative and complementary features to produce expressive performance. Experimental results show that the proposed MSDMR can achieve higher accuracy compared with state-of-the-art methods on the surveillance datasets. Yi Wang 0068, Junhui Hou, Lap-Pui Chau |
ICASSP | 1 |
| 2019 | Airtight Estimation Based on Distant Region SegmentationabstractNatural images suffer from bad weather conditions, such as haze or fog, which decreases the contrast and degrades the color of observed images. Haze removal aims to recover haze-free images by the image degradation model. The global atmospheric light (airlight) estimation is an essential step for haze removal. With an assumption that the airlight exists in the infinite distance, we propose a novel learning-based framework for airlight estimation. Our framework is mainly composed of two steps: i) the airlight is initially determined by distant region segmentation based on U-Net; ii) the final airlight can be obtained by the weighted sum of the pixel values inside the distant region. Owing to lack of ground-truth airlight, we present a method to synthesize outdoor training examples. The proposed framework not only perform well on synthetic images but also has a good generalization ability for natural images. Experimental results demonstrate that our proposed approach can achieve more accurate estimate of airlight than state-of-the-art methods on both synthetic and natural images. Yi Wang 0068, Lap-Pui Chau, Xiaoxi Ma |
ISCAS | 1 |
| 2017 | Single underwater image restoration using attenuation-curve priorabstractUnderwater images suffer from low contrast and color distortion due to the existence of dust-like particles and light attenuation. Some previous works using the patch-based priors, e.g. adaptations of the dark channel prior, cannot achieve satisfactory results in both contrast enhancement and color restoration in the underwater environment. In this paper, we propose a novel underwater image restoration method based on a non-local prior, termed an attenuation-curve prior. This prior relies on the observation that colors of a clear image can be well approximated by several hundred distinct color clusters and the pixels in the same color cluster will form a power function curved line in RGB space after their colors are attenuated by water. Our work mainly contains two steps. Firstly, we estimate the waterlight based on its smoothness properties and the different attenuation coefficient of light. Secondly, we estimate the transmission map using the attenuation-curve prior. Once the waterlight and transmission are obtained, the clear underwater image can be restored. Experimental results demonstrate that our proposed method can achieve better results when comparing with state-of-the-art approaches. Yi Wang 0068, Hui Liu 0032, Lap-Pui Chau |
ISCAS | 1 |
| 2015 | Onboard image selection for small-satellite based remote sensing missionabstractThe contradiction that imaging system can acquire huge amount of image data while communication system can deliver only a very small part of them has become a bottleneck for the small-satellite based earth observation. In this paper, a novel onboard image selection strategy is designed to select most informative images that are acquired by the imaging system for transmission. Specifically, the image that is worst reconstructed by previously transmitted images, instead of the instantly acquired image, is selected since it possesses most distinguishing information to previously transmitted images. Experiment on simulated image sequence has demonstrated the effectiveness of the proposed onboard image selection algorithm. Yihang Wang 0001, Shaohui Mei, Shuai Wan, Yi Wang 0068 |
IGARSS | 4 |