EDBT 2026 Demo / reviewers in the wild / expert
Xiao Wu 0001
dblp:73/6038-1
· DBLP profile ↗
117ranked-venue papers
16as first author
59since 2021 · last 2026
0000-0002-8322-8558ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 9 first-author · 45 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 13 since 2021Computer networks · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry RefinementabstractAccurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi-scale feature aggregation and unflexible prior modeling. To overcome these limitations, MonoVPR is proposed, a novel framework integrating dynamic context adaptation and progressive geometry refinement. Specifically, a Hierarchical Dual-Context Attention (HDCA) module is introduced to resolve scale-dependent degradation through gated cross-attention across multi-resolution feature maps, dynamically fusing object-centric geometric cues with scene-centric semantics. For shape refinement, the Bounded Iterative Mesh Refiner (BIMR) progressively optimizes template-guided deformations via multi-head attention and a tanh-bounded correction loop, ensuring physically plausible reconstructions.Extensive experiments on the ApolloCar3D benchmark demonstrate MonoVPR achieves state-of-the-art performance, showing exceptional capability in reconstructing geometrically consistent shapes and precise poses for challenging long-range scenarios. Wei Li 0110, Long Ji, Xiao Wu 0001, Zhaoquan Yuan, Penglin Dai |
AAAI | 4 |
| 2026 | InfoCom: Kilobyte-Scale Communication-Efficient Collaborative Perception with Information BottleneckabstractPrecise environmental perception is critical for the reliability of autonomous driving systems. While collaborative perception mitigates the limitations of single-agent perception through information sharing, it encounters a fundamental communication-performance trade-off. Existing communication-efficient approaches typically assume MB-level data transmission per collaboration, which may fail due to practical network constraints. To address these issues, we propose InfoCom, an information-aware framework establishing the pioneering theoretical foundation for communication-efficient collaborative perception via extended Information Bottleneck principles. Departing from mainstream feature manipulation, InfoCom introduces a novel information purification paradigm that theoretically optimizes the extraction of minimal sufficient task-critical information under Information Bottleneck constraints. Its core innovations include: i) An Information-Aware Encoding condensing features into minimal messages while preserving perception-relevant information; ii) A Sparse Mask Generation identifying spatial cues with negligible communication cost; and iii) A Multi-Scale Decoding that progressively recovers perceptual information through mask-guided mechanisms rather than simple feature reconstruction. Comprehensive experiments across multiple datasets demonstrate that InfoCom achieves near-lossless perception while reducing communication overhead from megabyte to kilobyte-scale, representing 440-fold and 90-fold reductions per agent compared to Where2comm and ERMVP, respectively. Quanmin Wei, Penglin Dai, Wei Li 0110, Bingyi Liu, Xiao Wu 0001 |
AAAI | 5 |
| 2026 | I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-DisentangleabstractCompositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks. Zhaoquan Yuan, Yuankang Pan, Ao Luo, Wei Li 0110, Xiao Wu 0001, Changsheng Xu |
AAAI | 6 |
| 2026 | Rethinking Crowd Localization Evaluation via Optimal Transportation Cost
Jun-Xiu Li, Hong Liu 0009, Xiao Wu 0001, Yu-Pei Song, Zhenhua Zeng, Shin'ichi Satoh 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Learning Unknowns Without Forgetting Knowns: Compositional and Bidirectional Low-Rank Adaptive Open-World Detection TransformerabstractOpen-World Object Detection (OWOD) aims to detect unseen objects as “unknown” while incrementally learning them without catastrophic forgetting. This problem presents two major challenges: (1) the lack of annotations for unknown objects during training, and (2) the risk of catastrophic forgetting during model updates. To address these issues, we propose the COmpositional and Bidirectional low-Rank Adaptive open-world detection transformer (COBRA)-a novel framework built upon a pre-trained Deformable DETR model. Specifically, COBRA first employs an attentional filtering mechanism that prunes previously known (P-Known) and currently known (C-Known) objects, yielding a purified set of candidateunknowns. To system-atically pseudo-label theseunknowns, we introduce a Primitive Composition Recognition (PCR) module, which evaluates set-level similarity between candidate objects and learned primitives, enabling accurate labeling ofpseudo-unknowns. To mitigate catastrophic forgetting during incremental updates, COBRA leverages Bidirectional Low-Rank Adaptation (Bi-LoRA)-a parameter-efficient mechanism that supports forward knowledge transfer and stable backward integration. Together, these components form a synergistic pipeline for continual object discovery and knowledge consolidation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our rehearsal-free COBRA framework outperforms SAM-powered methods in unknown recall while achieving lower forgetting compared to rehearsal-based competitors. Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Wei Li 0110, Ao Luo, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Class-Specific Knowledge-Guided Multimodal Prompt Tuning for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) requires a model to learn the knowledge of new categories incrementally, using only a few samples, after being trained on a base session with ample categories and sample sizes. This task presents two major challenges: catastrophic forgetting and overfitting. Current approaches primarily enhance the model’s ability to extract knowledge during the base stage to improve adaptability to new tasks. Large-scale pre-trained models, known for their high robustness and zero-shot transfer capabilities, have demonstrated promising performance in FSCIL. The key to solving FSCIL lies in effectively fine-tuning such large models to balance the learning of new knowledge and the retention of old knowledge. Inspired by human-like knowledge retrieval mechanisms, we propose Class-specific Knowledge-Guided Prompt Tuning (CKGPT), which leverages class-specific prompts to guide the model in learning targeted knowledge reuse and integration effectively. When faced with novel tasks, the model selectively activates previously learned knowledge that is the most relevant, improving performance on new tasks while minimizing updates to irrelevant knowledge to reduce forgetting. By incorporating mechanisms that balance knowledge retention and transfer, CKGPT ensures a more robust adaptation to sequential tasks. Extensive experiments on multiple benchmarks validate the effectiveness of our method in achieving superior performance. Fangying Xiong, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Cooperative Perception of Multi-Agents Under the Spatio-Temporal Drift IssueabstractCooperative perception has significant potential to enhance perception performance compared to single-agent systems by integrating information from multiple agents through vehicle-to-everything (V2X) communication. However, several challenges hinder the attainment of high performance in cooperative perception, particularly positional errors arising from sensor data collection and time delays during data transmission. Existing research often addresses only one of these issues, making it unsuitable for scenarios where spatial-temporal errors coexist. In this paper, we focus on resolving the spatio-temporal drift issue caused by the interplay of spatial and temporal variations. To address this, we propose a novel end-to-end cooperativeperception framework called Multi-frame Grouping Multi-agent Perception (MGMP), which effectively fuses spatio-temporal perception features from multiple agents, including vehicles and road infrastructure. Our approach extracts the effective semantic information of the temporal context of multiple agents, leverage the cross-learning of window information through multi-scale window attention, and group and aggregate multiple agents to simultaneously address the spatio-temporal drift problem caused by positional errors and time delays. We validate the effectiveness of our method on the V2XSet, OPV2V and Dair-V2X datasets. Experimental results indicate that, compared to the state-of-the-art (SOTA) work, our method achieves improvements of 2.7%, 1.7%, and 1.2% on [email protected], respectively. Penglin Dai, Quanmin Wei, Xiao Wu 0001, Zhanbo Sun, Zhaofei Yu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | THMM-CLIP: Task-Guided Hierarchical Multi-Modal Alignment for Rehearsal-Free Class Incremental LearningabstractClass incremental learning (CIL) requires models to acquire knowledge from sequential tasks containing non-overlapping classes while avoiding catastrophic forgetting. While vision-language foundation models like CLIP demonstrate remarkable potential for CIL through their pre-trained cross-modal alignment capabilities, existing CLIP-based approaches critically overlook the progressive degradation of visual representations in incremental scenarios . Through feature space analysis, we identify a crucial dichotomy : textual embeddings maintain stable discriminative power across sequential tasks, whereas visual features exhibit progressive deterioration manifested by intra-task confusion (ambiguous decision boundaries between co-occurring classes) and inter-task interference (semantic collision between historical and novel categories). To address these dual challenges, we propose task-guided hierarchical multi-modal alignment (THMM-CLIP), a framework that establishes persistent visual-textual coherence through hierarchical multi-modal alignment (HMA) and robust prompt selection (RPS). HMA adapts lightweight task-specific prompt vectors to dynamically recalibrate the CLIP image encoder, thereby achieving: (i) intra-task alignment, (ii) inter-task discriminability alignment, and (iii) global structural alignment with textual features. RPS incorporates a dual-level task identifier that integrates class-level and task-level representative features to ensure precise prompt retrieval during inference. Ablation studies validate all components’ contributions, while t-SNE visualizations, confusion matrices, and Grad-CAM analyses confirm strengthened cross-modal alignment. Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | PSGNet: Pure Smoke Image Generation With Gradient and Style LearningabstractThe realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches. Jian-Jun Qiao, Xiao Wu 0001, Zhi-Qi Cheng, Wei Li 0110, Zhaoquan Yuan |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | CoPEFT: Fast Adaptation Framework for Multi-Agent Collaborative Perception with Parameter-Efficient Fine-TuningabstractMulti-agent collaborative perception is expected to significantly improve perception performance by overcoming the limitations of single-agent perception through exchanging complementary information. However, training a robust collaborative perception model requires collecting sufficient training data that covers all possible collaboration scenarios, which is impractical due to intolerable deployment costs. Hence, the trained model is not robust against new traffic scenarios with inconsistent data distribution and fundamentally restricts its real-world applicability. Further, existing methods, such as domain adaptation, have mitigated this issue by exposing the deployment data during the training stage but incur a high training cost, which is infeasible for resource-constrained agents. In this paper, we propose a Parameter-Efficient Fine-Tuning-based lightweight framework, CoPEFT, for fast adapting a trained collaborative perception model to new deployment environments under low-cost conditions. CoPEFT develops a Collaboration Adapter and Agent Prompt to perform macro-level and micro-level adaptations separately. Specifically, the Collaboration Adapter utilizes the inherent knowledge from training data and limited deployment data to adapt the feature map to new data distribution. The Agent Prompt further enhances the Collaboration Adapter by inserting fine-grained contextual information about the environment. Extensive experiments demonstrate that our CoPEFT surpasses existing methods with less than 1\% trainable parameters, proving the effectiveness and efficiency of our proposed method. Quanmin Wei, Penglin Dai, Wei Li 0110, Bingyi Liu, Xiao Wu 0001 |
AAAI | 5 |
| 2025 | POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position SearchabstractAchieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs three key contributions: (1) Pseudo-range multilateration is utilized to correct heatmap errors, improving landmark localization accuracy. By integrating multiple anchor points, it reduces the impact of individual heatmap inaccuracies, leading to robust overall positioning. (2) To enhance the pseudo-range accuracy of selected anchor points, a new loss function, named multilateration anchor loss, is proposed. This loss function enhances the accuracy of the distance map, mitigates the risk of local optima, and ensures optimal solutions. (3) A single-step parallel computation algorithm is introduced, boosting computational efficiency and reducing processing time. Extensive evaluations across five benchmark datasets demonstrate that POPoS consistently outperforms existing methods, particularly excelling in low-resolution heatmaps scenarios with minimal computational overhead. These advantages make POPoS as a highly efficient and accurate tool for FLD, with broad applicability in real-world scenarios. Chong-Yang Xiang, Jun-Yan He, Zhi-Qi Cheng, Xiao Wu 0001, Xian-Sheng Hua 0001 |
AAAI | 4 |
| 2025 | Dual-Rate Dynamic Teacher for Source-Free Domain Adaptive Object Detection
Qi He 0007, Xiao Wu 0001, Jun-Yan He, Shuai Li 0014 |
ICCV | 2 |
| 2025 | Contrastive Invariant Risk Minimization for Grounded Situation RecognitionabstractGrounded situation recognition (GSR) is a comprehensive structured scene understanding task that predicts the salient activity (verb), entities (nouns) involved in the activity with their roles, as well as the corresponding bounding-box groundings of the entities from the given image. Existing I.I.D.-based methods for GSR are limited in their ability to recognize novel verb-noun combinations. To address this problem, in this paper, we novelly consider GSR as a Non-I.I.D. task and focus on learning verb-invariant and role-specific representations for verb and noun predictions. Based on the causality, a novel Contrastive Invariant Risk Minimization (CIRM) model for GSR is proposed. In the proposed CIRM, invariant risk minimization is integrated into a transformer architecture to learn invariant representations for verb prediction. To enhance the intra-verb compactness and the inter-verb separability, contrastive learning is utilized to learn discriminative features. As far as we know, this is the first work that regards the task of GSR as a problem of out-of-distribution generalization. Extensive experiments on the benchmark SWiG dataset demonstrate the effectiveness of our proposed CIRM over other state-of-the-art methods in all evaluation metrics. Zhaoquan Yuan, Chengbin Zhao, Yuting Tang, Lishu Guo, Xiao Wu 0001, Changsheng Xu |
ICME | 5 |
| 2025 | Latent Interactiveness Field for Non-Contact Human Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection serves a broad spectrum of applications. Despite significant progress, current approaches encounter difficulties in effectively handling Non-Contact Human-Object Interaction (NCHOI) scenarios, where humans and objects remain physically apart. To address these challenges, this paper proposes a novel approach, named Latent Interactiveness Field Modeling (LIFM), which enhances HOI detection by capturing long-range contextual dependencies. Specifically, the Latent Interactiveness Field (LIF) is introduced to define potential interactive relationships between humans and objects. To complement this, the LIF Fusion Encoder is designed to adaptively fuse visual features with LIF, resulting in more informative and discriminative feature representations. The Mobile Scanning HOI Dataset (MSHD) is introduced as a comprehensive benchmark to systematically assess the robustness of existing methods on both common HOI and NCHOI in real-world applications. Extensive experimentation indicates that the proposed approach outperforms existing state-of-the-art techniques. It offers substantial improvements, particularly in NCHOI scenarios, which highlight its effectiveness in resolving issues related to long-range interactions. Xiang Huang 0004, Ao Luo, Xiao Wu 0001, Zhaoquan Yuan |
ACM Multimedia | 3 |
| 2025 | HOPNet: Learning Hand-Object-Person Interaction Network for Hand Contact State DetectionabstractThe detection of hand contact states, which involves identifying interactions between hands and objects or other entities, is essential for the development of human-computer interaction systems and the comprehension of social dynamics. Previous approaches have made progress in modeling hand-object interactions. Nonetheless, they neglect critical cues between their hands and bodies, as well as those of others, thus constraining their ability to accurately detect interpersonal contact. The task remains challenging due to frequent occlusions, especially in crowded multi-person scenarios with complex contexts. In this paper, a novel hand-object-person interaction network, called HOPNet, is proposed to model contextual information between hands and objects, as well as between hands and bodies. Specifically, HOPNet consists of two components: (i) the Hand-Object Relation (HOR) module analyzes interaction patterns between hands and objects, capturing spatial and semantic relationships; (ii) the Contrastive Spatial Refinement (CSR) module learns hand-body interactions through contrastive geometric embedding and relative spatial enhancement, improving interpersonal contact recognition in crowded scenarios. Experiments on ContactHands and 100DOH datasets demonstrate that HOPNet outperforms state-of-the-art methods. Wei Li 0110, Yizhao Wan, Xiao Wu 0001, Jianshuai Wang, Penglin Dai, Zhaoquan Yuan |
ACM Multimedia | 3 |
| 2025 | DualEnhance: External Multimodal Foundation Models Guidance and Internal Fast-Slow Teacher RegulationabstractSource-Free Domain Adaptive Object Detection addresses cross-domain detection on an unlabeled target domain without accessing source data. Existing methods implement self-training with Mean Teacher but are bottlenecked by error accumulation from noisy pseudo-labels generated via recursive teacher-student updates. This issue is handled through the proposed dual enhancements: (1) External Guidance via Multimodal Foundation Models (FMs); (2) Internal Regulation through Fast-Slow Teacher. First, despite FMs' multimodal comprehension, their semantic misalignment with a specific task introduces noise during adaptation. Bidirectional Distillation mitigates this by calibrating the FM using task-specific knowledge transferred from the source detector. The aligned cross-modal knowledge then propagates through high-quality pseudo-label generation. Second, the conventional Mean Teacher suffers from plasticity-stability dilemma, where rapid adaptation corrupts historical knowledge. Fast-Slow Teacher introduces dual-velocity knowledge consolidation: The Fast Teacher dynamically captures emerging domain features, while the Slow Teacher preserves stable historical knowledge and periodically resets the Fast Teacher, establishing an error-correcting dynamic equilibrium. Experiments show our method achieves significant improvements over SOTA. Qi He 0007, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Zhaoquan Yuan |
ACM Multimedia | 2 |
| 2025 | Pragmatic Heterogeneous Collaborative Perception via Generative Communication MechanismabstractMulti-agent collaboration enhances the perception capabilities of individual agents through information sharing. However, in real-world applications, differences in sensors and models across heterogeneous agents inevitably lead to domain gaps during collaboration. Existing approaches based on adaptation and reconstruction fail to support *pragmatic heterogeneous collaboration* due to two key limitations: (1) Intrusive retraining of the encoder or core modules disrupts the established semantic consistency among agents; and (2) accommodating new agents incurs high computational costs, limiting scalability. To address these challenges, we present a novel **Gen**erative **Comm**unication mechanism (GenComm) that facilitates seamless perception across heterogeneous multi-agent systems through feature generation, without altering the original network, and employs lightweight numerical alignment of spatial information to efficiently integrate new agents at minimal cost. Specifically, a tailored Deformable Message Extractor is designed to extract spatial message for each collaborator, which is then transmitted in place of intermediate features. The Spatial-Aware Feature Generator, utilizing a conditional diffusion model, generates features aligned with the ego agent's semantic space while preserving the spatial information of the collaborators. These generated features are further refined by a Channel Enhancer before fusion. Experiments conducted on the OPV2V-H, DAIR-V2X and V2X-Real datasets demonstrate that GenComm outperforms existing state-of-the-art methods, achieving an 81\% reduction in both computational cost and parameter count when incorporating new agents. Our code is available at https://github.com/jeffreychou777/GenComm. Junfei Zhou, Penglin Dai, Quanmin Wei, Bingyi Liu, Xiao Wu 0001 |
NeurIPS | 5 |
| 2025 | HighlightNet: Learning Highlight-Guided Attention Network for Nighttime Vehicle DetectionabstractVehicle detection at night is a crucial task in Intelligent Transportation Systems. Due to the complex lighting environment, vehicle detection at night remains a challenging task. Headlights and taillights are essential cues to identify vehicles at night. However, existing methods struggle to effectively utilize the light information of the vehicle. This paper proposes a novel highlight-guided framework to identify vehicles, named HighlightNet, by utilizing both the illumination data from the vehicle lights and the reflective properties of vehicles. The framework combines vehicle detection and highlight area recognition via dual-branch joint learning. To ensure that both branches focus on the highlighted regions, Feature Similarity Awareness Attention (FSAA) is introduced to capture the common attention regions of different branches. Highlight Region Perception (HRP) is proposed to exclude streetlights and other reflective illuminations from the FSAA output, which generates a mask map capable of differentiating the foreground from the background of highlighted areas. It improves the allocation of feature weights and adaptively modifies the distribution within the dual-branch configuration. Furthermore, to address the severe pixel imbalance between the highlighted area and the background, Adaptive Spatial Balance (ASB) loss is introduced to allocate the attention towards prospective vehicle regions while diminishing the emphasis on background regions. Extensive experiments conducted on the BDD100K-Night dataset and a newly acquired dataset specifically designed for nighttime surveillance, called the NightVehicle dataset, demonstrate that HighlightNet outperforms the state-of-the-art methods for nighttime vehicle detection. Yu-Pei Song, Xiao Wu 0001, Wei Li 0110, Tingquan He, Dongfeng Hu, Qiang Peng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Disentanglement-Based Equivariant Learning for Compositional VQAabstractCompositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of previously learned concepts. The current methods often overlook the disentanglement of underlying concepts and are restricted in terms of their ability to effectively capture the compositional variation mechanism. Moreover, the state-of-the-art techniques depend on additional clues for training, which is not feasible in real-world VQA scenarios. To address these issues, in this paper, we introduce a novelDisentanglement-basedEquivAriantLearning (DEAL) framework for compositional VQA, which is guided exclusively by ground-truth answers. In DEAL, we employ causality-inspired interventions to disentangle concepts derived from visual and textual inputs within a re-encoding framework. Based on the principle of equivariance, we subsequently perform a compositional transformation on the inference input and impose the equivariant constraint on the output to augment the compositional reasoning capacity of the model. Comprehensive experiments conducted on the benchmark CLEVR-CoGenT and GQA-SGL datasets validate the superiority of our proposed DEAL approach over the existing state-of-the-art methods for compositional VQA tasks in both visual and linguistic generalization settings. Zhou Du, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2025 | Active Cross-Modal Domain AdaptationabstractMost cross-modal methods assume that training and testing data come from the same domain, which is often not the case in real-world scenarios due to cross-modal domain shifts and potential unknown concepts. Moreover, cross-modal shifts hinder the capture of unknown concepts, and the presence of unknown concepts can in turn exacerbate the cross-modal shifts. To address these challenges, this paper proposes a new paradigm called Active Cross-Modal Domain Adaptation (ACM-DA), wherein only cross-modal data from the source domain and uni-modal data from the target domain are utilized. To concurrently mitigate the adverse effects of both cross-modal domain shifts and unknown concepts, we propose a Curiosity-Driven Active Adaptation Network (CD-A2N), selectively annotating samples to maximize performance gain. First, we present Curiosity Arousal within Cross-modal Domain Adaptation (CA-CDA) to explore the complexity and novelty characteristics of target samples, while reducing cross-modal discrepancy and aligning source and target domains. Second, Curiosity-driven Active Learning (CAL) is devised to strategically select a subset of target samples for annotation, aiming to achieve more valuable data selection at a small labeling cost. Finally, we jointly train CA-CDA and CAL with the newly labeled target domain sub-dataset to alleviate the above issues. Extensive experiments demonstrate that CD-A2N provides an effective solution for achieving ACM-DA. Code will be available athttps://github.com/Feliciaxyao/ACM-DA. Xuan Yao 0001, Junyu Gao 0002, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Multim. | 5 |
| 2025 | Person in Uniforms Re-IdentificationabstractPerson in Uniforms Re-identification (PU-ReID) is an emerging computer vision task for various intelligent video surveillance applications. PU-ReID is much understudied due to the absence of large-scale annotated datasets, also this task is extremely challenging because many individuals captured in surveillance videos wear same clothing, introducing significant interference for retrieval tasks owing to the high visual similarity of outfits and subtle differences among individuals. This research initiates the exploration of person in uniforms re-identification, a novel and challenging task tailored for real industrial scenarios. To address these issues, a novel framework is proposed for PU-ReID, which aims to reduce the visual impact of similar uniforms and learn the unique cues derived from human parts and detailed visual features. Specifically, several novel techniques are built in this study: first, a uniform feature separation method with orthogonal constraints is proposed to extract non-uniform features. Second, multi-view subspace feature alignment is introduced to integrate soft-biometrics including optics-related visual features, contextual information of human parts, and cloth-invariant biometric features. In addition, to close the gap between academic research and real-world settings, a new person in uniforms ReID dataset named PU-151 is constructed, which consists of 151 gas station employees in uniforms from 1,488 videos. At last, extensive experiments conducted on five datasets demonstrate that the proposed approach significantly outperforms the state-of-the-art methods. This advancement can drive further developments in re-identification and person search technologies. Chong-Yang Xiang, Xiao Wu 0001, Jun-Yan He, Zhaoquan Yuan, Tingquan He |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Multi-Agent Reinforcement Learning for Freshness-Aware Data Sensing Model in Vehicular Crowdsensing SystemsabstractVehicular Crowdsensing (VCS) is a promising paradigm for supporting urban sensing services, where Service Providers (SPs) engage Mobile Vehicles (MVs) to perform data sensing tasks with specific objectives. However, existing studies have predominantly focused on data sensing quality in terms of data collection completeness and geographic fairness, while largely neglecting the important aspect of data freshness. Moreover, effective mechanisms for optimizing data freshness through coordination of the behaviors of both SPs and MVs are still lacking. Accordingly, this paper proposes a Freshness-Aware Data Sensing (FDS) model by considering heterogeneous data freshness, varying sensing capabilities of MVs, and limited budgets of SPs. The FDS is formulated as a two-stage game model, where SPs and MVs iteratively determine their pricing and sensing strategies in a self-interested manner to maximize their individual gains. Further, we develop a multi-agent reinforcement learning-based approach to learn the pricing strategies based on historical observations, which allows SPs to make pricing decisions without global knowledge. Additionally, given the pricing strategies of SPs, the optimal solution for each MV is derived. Finally, we build the simulation model based on realistic vehicular traces, where the simulation results demonstrate the superiority of the proposed algorithm in various scenarios. Penglin Dai, Xin Wang 0190, Yue Xiang, Xiao Wu 0001, Kai Liu 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | PostureHMR: Posture Transformation for 3D Human Mesh RecoveryabstractHuman Mesh Recovery (HMR) aims to estimate the 3D human body from 2D images, which is a challenging task due to inherent ambiguities in translating 2D observations to 3D space. A novel approach called PostureHMR is pro-posed to leverage a multi-step diffusion-style process, which converts this task into a posture transformation from an SMPL T-pose mesh to the target mesh. To inject the learning process of posture transformation with the physical structure of the human body model, a kinematics-based forward process is proposed to interpolate the intermediate state with pose and shape decomposition. Moreover, a mesh-to-posture (M2P) decoder is designed, by combining the in-put of 3D and 2D mesh constraints estimated from the im-age to model the posture changes in the reverse process. It mitigates the difficulties of posture change learning directly from RGB pixels. To overcome the limitation of pixel-level misalignment of modeling results with the input image, a new trimap-based rendering loss is designed to highlight the areas with poor recognition. Experiments conducted on three widely used datasets demonstrate that the proposed approach outperforms the state-of-the-art methods. Yu-Pei Song, Xiao Wu 0001, Zhaoquan Yuanl, Jian-Jun Qiao, Qiang Peng |
CVPR | 2 |
| 2024 | Rethinking the Effect of Uninformative Class Name in Prompt LearningabstractLarge pre-trained vision-language models like CLIP have shown amazing zero-shot recognition performance. To adapt pre-trained vision-language models to downstream tasks, recent studies have focused on the learnable context + class name paradigm, which learns continuous prompt contexts on downstream datasets. In practice, the learned prompt context tends to overfit the base categories and cannot generalize well to novel categories out of the training data. Recent works have also noticed this problem and have proposed several improvements. In this work, we draw a new insight based on empirical analysis, that is, uninformative class names lead to degraded base-to-novel generalization performance in prompt learning, which is usually overlooked by existing works. Under this motivation, we advocate to improve the base-to-novel generalization performance of prompt learning by enhancing the semantic richness of class names. We coin our approach as the Information Disengagement based Associative Prompt Learning (IDAPL) mechanism which considers the associative, meanwhile, decoupled learning of prompt context and class name embedding. IDAPL can effectively alleviate the phenomenon of learnable context overfitting to base classes, meanwhile, learning more informative semantic representation of base classes by fine-tuning the class name embedding, leading to improved performance on both base and novel classes. Experimental results on eleven widely used few-shot learning benchmarks clearly validate the effectiveness of our proposed approach. Code is available at https://github.com/tiggers23/IDAPL Fengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang, Guosheng Lin, Xiao Wu 0001, Tianrui Li 0001 |
ACM Multimedia | 6 |
| 2024 | CAPNet: Cartoon Animal Parsing with Spatial Learning and Structural ModelingabstractCartoon animal parsing aims to segment the body parts such as heads, arms, legs and tails of cartoon animals. Different from previous parsing tasks, cartoon animal parsing faces new challenges, including irregular body structures, abstract drawing styles and diverse animal categories. Existing methods have difficulties when addressing these challenges caused by the spatial and structural properties of cartoon animals. To address these challenges, a novel spatial learning and structural modeling network, named CAPNet, is proposed for cartoon animal parsing. It aims to address the critical problems of spatial perception, structure modeling and spatial-structural consistency learning. A spatial-aware learning module integrates deformable convolutions to learn spatial features of diverse cartoon animals. The multi-task edge and center point prediction mechanism is incorporated to capture the intricate spatial patterns. A structural modeling method is proposed to model the complex structural representations of cartoon animals, which integrates a graph neural network with a shape-aware relation learning module. To mitigate the significant differences among animals, a spatial and structural consistency learning strategy is proposed to capture and learn feature correlations across different animal species. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods. Jian-Jun Qiao, Meng-Yu Duan, Xiao Wu 0001, Wei Li 0110 |
ACM Multimedia | 3 |
| 2024 | CartoonNet: Cartoon Parsing with Semantic Consistency and Structure CorrelationabstractCartoon parsing is an important task for cartoon-centric applications, which segments the body parts of cartoon images. Due to the complex appearances, abstract drawing styles, and irregular structures of cartoon characters, cartoon parsing remains a challenging task. In this paper, a novel approach, named CartoonNet, is proposed for cartoon parsing, in which semantic consistency and structure correlation are integrated to address the visual diversity and structural complexity for cartoon parsing. A memory-based semantic consistency module is designed to learn the diverse appearances exhibited by cartoon characters. The memory bank stores features of diverse samples and retrieves the samples related to new samples for consistency, which aims to improve the semantic reasoning capability of the network. A self-attention mechanism is employed to conduct consistency learning among diverse body parts belong to the retrieved samples and new samples. To capture the intricate structural information of cartoon images, a structure correlation module is proposed. Leveraging graph attention networks and a main body-aware mechanism, the proposed approach enables structural correlation, allowing it to parse cartoon images with complex structures. Experiments conducted on cartoon parsing and human parsing datasets demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art approaches for cartoon parsing and achieves competitive performance on human parsing. Jian-Jun Qiao, Meng-Yu Duan, Xiao Wu 0001, Yu-Pei Song |
ACM Multimedia | 3 |
| 2024 | MagicCartoon: 3D Pose and Shape Estimation for Bipedal Cartoon CharactersabstractThe 3D model can be estimated by regressing the pose and shape parameters from the image data of the digital model. The reconstruction of 3D cartoon characters poses a challenging task due to diverse visual representations and postural variations. This paper proposes a dual-branch structure named MagicCartoon for 3D bipedal cartoon character estimation, which models pose and shape independently through feature decoupling. Considering the correlation between category difference and shape parameters, a hybrid feature fusion technique is introduced, which integrates the global features of the original image with the corresponding local features expressed by the puzzle image, reducing the abstractness of understanding shape parameter differences. To semantically align image and geometric between feature space, a geometric-guided feedback loop is proposed in an iterative way, so that the pose of modeling results can be expressed consistently with the image. Moreover, a feature consistency loss is designed to augment the training data by incorporating the same character with different postures and the same posture of different characters. It enhances the correlation between the features extracted by the backbone network and the specific task. Experiments conducted on the 3DBiCar dataset demonstrate that MagicCartoon outperforms the state-of-the-art methods. Yu-Pei Song, Yuantong Liu, Xiao Wu 0001, Qi He 0007, Zhaoquan Yuan, Ao Luo |
ACM Multimedia | 3 |
| 2024 | TMM-CLIP: Task-guided Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu |
MMAsia | 3 |
| 2024 | Context-Aware Offloading for Edge-Assisted On-Device Video Analytics Through Online Learning ApproachabstractEdge computing has emerged as a powerful technology for enhancing the performance of on-device video analytics, which is critical to support real-time applications. Nevertheless, there still lack of effective metrics to guide the offloading decision of video analytics tasks between device and edge server. Additionally, these existing optimization mechanisms either presume prior knowledge of the ground-truth of previous inferences or involve high training overheads, thereby rendering them unsuitable for real-time situations. To address these challenges, this paper presents a system model of edge-assisted online video analytics, where a lightweight object tracking module and a complex DNN-based model are deployed at the device and edge server, respectively. We formulate the resolution and deviation-based offloading (RDO) problem by considering heterogeneous computation resources and dynamic network bandwidth, aiming at maximizing inference accuracy and processing rate concurrently. We propose a context-aware offloading (CO) algorithm based on Bayesian optimization, which learns the optimal parameter settings by evaluating reward based on Gaussian process. Notably, the CO is proved to offer near-optimal solution with sublinear regret. Finally, we build a testbed and test algorithm performance on three realistic video datasets. The simulation results illustrate that the proposed CO outperforms other existing solutions in various service scenarios. Penglin Dai, Yangyang Chao, Xiao Wu 0001, Kai Liu 0001, Songtao Guo |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Meta Reinforcement Learning for Multi-Task Offloading in Vehicular Edge ComputingabstractMobile edge computing has been a promising solution to enable real-time service in vehicular networks. However, due to high dynamics of mobile environment and heterogeneous features of vehicular services, traditional expert-based or learning-based strategies has to update handcrafted parameters or retrain learning model, which leads to intolerant overhead. Therefore, this paper investigates the problem of multi-task offloading (MTO), where there exist multiple offloading scenarios with varying parameters, such as task topology, resource requirement and transmission/computation capability. The objective is to design a unified solution to minimize task execution time under different MTO scenarios. Accordingly, we develop a Seq2seq-based Meta Reinforcement Learning algorithm for MTO (SMRL-MTO). Specifically, a bidirectional gated recurrent units integrated with attention mechanism is designed to determine offloading action by encoding sequential offloading actions and showing different preferences to different parts of input sequence. Particularly, a meta reinforcement learning framework is designed based on model-agnostic meta learning, which trains a meta policy offline and fast adapts to new MTO scenario within a few training steps. Finally, we conduct performance evaluation based on task generator DAGGEN and realistic vehicular traces, which shows that the SMRL-MTO reduces task execution time by 11.36% on average compared with greedy algorithm. Penglin Dai, Yaorong Huang, Kaiwen Hu, Xiao Wu 0001, Huanlai Xing, Zhaofei Yu |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Distributed Convex Relaxation for Heterogeneous Task Replication in Mobile Edge ComputingabstractMobile edge computing (MEC) is expected to support real-time services at wireless networks, where task replication is applied to guarantee job completion within a strict deadline through replicating multiple copies to different edge servers. Most of previous works focused on guaranteeing the reliability of individual task in MEC-based networks with the assumption of homogeneous task execution distribution. Further, these algorithms cannot suit dynamic network scales, due to overhigh communication or retraining overhead. Therefore, this paper formulates the problem of heterogeneous task replication in a finer level by modeling outage probability of individual replication, where the decisions of all tasks are jointly optimized within the constraints of both mobile users and MEC servers for minimizing job outage probability. To adapt to varying network scales, we develop centralized and distributed algorithms, respectively. The centralized algorithm is developed based on Interior Point Method, which obtains the optimal solution of relaxed model and then approximates to the solution of original problem. Further, the distributed algorithm decomposes the HTR into multiple subproblems and parallelly compute each local solution based on Distributed ADMM. Finally, we build a simulation model and conduct comprehensive results, which demonstrates that the proposed algorithms can achieve high-accuracy solution with fast convergence. Penglin Dai, Biao Han 0001, Xiao Wu 0001, Huanlai Xing, Bingyi Liu, Kai Liu 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Joint Optimization for Quality Selection and Resource Allocation of Live Video Streaming in Internet of VehiclesabstractLive Video Streaming (LVS) services are critical in supporting real-time applications in Internet of Vehicles (IoV) by transmitting real-time generated video content from streaming server to vehicles. Due to restricted spectrum resources and high vehicle mobility, LVS suffers from notable performance degradation. Moreover, existing strategies such as buffer size control and edge caching, are designed for video-on-demand service, which is ineffective for LVS in IoV. Accordingly, we investigate the problem of LVS-IoV by synthesizing multicasting and Scalable Video Coding-based encoding with the goal of maximizing Quality of Experience (QoE), which is defined as the weighted sum of video quality, rebuffering time, and quality variation. The LVS-IoV is decoupled into three sub-problems: vehicle grouping, quality selection, and resource allocation. Firstly, we propose a K-means-based vehicle grouping method that considers geographical distribution, velocity, and dynamic channels. Secondly, we determine the quality selection of each group based on the Value Decomposition Network for maximizing overall video quality. This network utilizes global value function decomposition and centralized training to achieve fast convergence, followed by distributed execution. Lastly, we propose a sub-gradient algorithm to achieve optimal resource allocation. We build simulation model and perform extensive evaluation, which demonstrates its superiority compared to other competitive methods. Penglin Dai, Meiting Wu, Ke Li 0020, Xiao Wu 0001, Yan Ding 0002 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | CPNet: Cartoon Parsing with Pixel and Part CorrelationabstractCartoon parsing, the task of segmenting constituent parts such as heads, arms, and legs of cartoon characters, holds substantial significance for applications in the animation industry and emerging metaverse. Nonetheless, this domain presents considerable challenges stemming from complex visual appearances, irregular structures, abstract drawing styles, among other factors. In this paper, a novel Cartoon Parsing Network (CPNet) is introduced to address these challenges. CPNet skillfully leverages the spatial and semantic correlations of pixels to discern intricate and visually akin appearances. Furthermore, it employs both local and global correlations of constituent parts to differentiate irregular and abstract body sections. Specifically, the pixels of the cartoon image are interconnected by capitalizing on the spatial and semantic correlations. To this end, a center point predictor, working in tandem with a pixel-aware attention, facilitates the exploration of pixel-level correlation learning. Additionally, the various constituent parts are meticulously organized to resonate with the intrinsic physiological structure of a cartoon character. The character's graph structure is assembled and analyzed by an edge-aware graph neural network, thereby linking adjacent parts and assimilating local correlations. A part-guided non-local attention mechanism is fashioned to correlate individual parts with the entire body, thereby modeling global connections. In addition, a new dataset named CartoonSet is curated and annotated explicitly for cartoon parsing. Experiments carried out on both cartoon parsing and human parsing datasets yield compelling results, thereby attesting to the efficacy and innovativeness of the proposed method. Jian-Jun Qiao, Jie Zhang 0179, Xiao Wu 0001, Yu-Pei Song, Wei Li 0110 |
ACM Multimedia | 3 |
| 2023 | Debunking Free Fusion Myth: Online Multi-view Anomaly Detection with Disentangled Product-of-Experts ModelingabstractMulti-view or even multi-modal data is appealing yet challenging for real-world applications. Detecting anomalies in multi-view data is a prominent recent research topic. However, most of the existing methods 1) are only suitable for two views or type-specific anomalies, 2) suffer from the issue of fusion disentanglement, and 3) do not support online detection after model deployment. To address these challenges, our main ideas in this paper are three-fold: multi-view learning, disentangled representation learning, and generative model. To this end, we propose dPoE, a novel multi-view variational autoencoder model that involves (1) a Product-of-Experts (PoE) layer in tackling multi-view data, (2) a Total Correction (TC) discriminator in disentangling view-common and view-specific representations, and (3) a joint loss function in wrapping up all components. In addition, we devise theoretical information bounds to control both view-common and view-specific representations. Extensive experiments on six real-world datasets demonstrate that the proposed dPoE outperforms baselines markedly. Hao Wang 0068, Zhi-Qi Cheng, Jingdong Sun, Xin Yang 0012, Xiao Wu 0001, Hongyang Chen 0001, Yan Yang 0001 |
ACM Multimedia | 5 |
| 2023 | Human-Object-Object Interaction: Towards Human-Centric Complex Interaction DetectionabstractLocalizing and recognizing interactive actions in videos is a pivotal yet intricate task that paves the way towards profound video comprehension. Recent advancements in Human-Object Interaction (HOI) detection, which involve detecting and localizing the interactions between human and object pairs, have undeniably marked significant progress. However, the realm of human-object-object interaction, an essential aspect of real-world industrial applications, remains largely uncharted. In this paper, we introduce a novel task referred to as Human-Object-Object Interaction (HOOI) detection and present a cutting-edge method named the Human-Object-Object Interaction Network (H2O-Net). The proposed H2O-Net is comprised of two principal modules: sequential motion feature extraction and HOOI modeling. The former module delves into the gradually evolving visual characteristics of entities throughout the HOOI process, harnessing spatial-temporal features across multiple fine-grained partitions. Conversely, the latter module aspires to encapsulate HOOI actions through intricate interactions between entities. It commences by capturing and amalgamating two sub-interaction features to extract comprehensive HOOI features, subsequently refining them using the interaction cues embedded within the long-term global context. Furthermore, we contribute to the research community by constructing a new video dataset, dubbed the HOOI dataset. The actions encompassed within this dataset pertain to pivotal operational behaviors in industrial manufacturing, imbuing it with substantial application potential and serving as a valuable addition to the existing repertoire of interaction action detection datasets. Experimental evaluations conducted on the proposed HOOI and widely-used AVA datasets demonstrate that our method outperforms existing state-of-the-art techniques by margins of 6.16 mAP and 1.9 mAP, respectively, thus substantiating its effectiveness. Mingxuan Zhang 0001, Xiao Wu 0001, Zhaoquan Yuan, Qi He 0007, Xiang Huang 0004 |
ACM Multimedia | 2 |
| 2023 | Improving Anomaly Segmentation with Multi-Granularity Cross-Domain AlignmentabstractAnomaly segmentation plays a crucial role in identifying anomalous objects within images, which facilitates the detection of road anomalies for autonomous driving. Although existing methods have shown impressive results in anomaly segmentation using synthetic training data, the domain discrepancies between synthetic training data and real test data are often neglected. To address this issue, Multi-Granularity Cross-Domain Alignment (MGCDA) framework is proposed for anomaly segmentation in complex driving environments. It uniquely combines a new Multi-source Domain Adversarial Training (MDAT) module and a novel Cross-domain Anomaly-aware Contrastive Learning (CACL) method to boost the generality of the model, seamlessly integrating multi-domain data at both scene and sample levels. Multi-source domain adversarial loss and a dynamic label smoothing strategy are integrated into MDAT module to facilitate the acquisition of domain-invariant features at the scene level, through adversarial training across multiple stages. CACL aligns sample-level representations with contrastive loss on cross-domain data, which utilizes an anomaly-aware sampling strategy to efficiently sample hard samples and anchors. The proposed framework has decent properties of parameter-free during the inference stage and is compatible with other anomaly segmentation networks. Experimental conducted on Fishyscapes and RoadAnomaly datasets demonstrate that the proposed framework achieves the state-of-the-art performance. Ji Zhang 0027, Xiao Wu 0001, Zhi-Qi Cheng, Qi He 0007, Wei Li 0110 |
ACM Multimedia | 2 |
| 2023 | Learning Surface-awareness Network for X-Ray Prohibited Item DetectionabstractX-ray image security detection is a crucial method used to identify various types of prohibited items in luggage. However, the unique characteristics of X-ray imaging can result in the loss of intricate surface details, leading to subpar detection of prohibited items within X-ray images. In this paper, a Surface-aware Prohibited Item X-ray Detection Network (SPIXDet) is proposed to address this issue, which incorporates two key components: the Boundary Aggregation Module (BAM) and the Global Cross-Feature Downsampling layer (GCFD). The BAM module effectively mines image edge information while minimizing the number of parameters involved. Meanwhile, the GCFD module is introduced to mitigate chaotic interference caused by undifferentiated boundary boosting. The surface-aware capability of the model can be enhanced through the BAM and GCFD module. Furthermore, the Focal-SIoU loss function is introduced to increase positioning accuracy and optimize the model training process. To validate the effectiveness of our model, extensive experiments are conducted on the SIXray100 dataset, and the results demonstrate the advantages of SPIXDet compared to other X-ray prohibited item detection methods. Wei Li 0110, Zhaoquan Yuan, Xiao Wu 0001 |
MMAsia | 4 |
| 2023 | Stacked denoising autoencoder for missing traffic data reconstruction via mobile edge computing
Penglin Dai, Jingtao Luo, Kangli Zhao, Huanlai Xing, Xiao Wu 0001 |
Neural Comput. Appl. | 5 |
| 2022 | Rethinking Spatial Invariance of Convolutional Networks for Object CountingabstractPrevious work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too strict pixel-level spatial invariance would cause overfit noise in the density map generation. In this paper, we try to use locally connected Gaussian kernels to replace the original convolution filter to estimate the spatial position in the density map. The purpose of this is to allow the feature extraction process to potentially stimulate the density map generation process to overcome the annotation noise. Inspired by previous work, we propose a low-rank approximation accompanied with translation invariance to favorably implement the approximation of massive Gaussian convolution. Our work points a new direction for follow-up research, which should investigate how to properly relax the overly strict pixel-level spatial invariance for object counting. We evaluate our methods on 4 mainstream object counting networks (i.e., MCNN, CSRNet, SANet, and ResNet-50). Extensive experiments were conducted on 7 popular benchmarks for 3 applications (i.e., crowd, vehicle, and plant counting). Experimental results show that our methods significantly outperform other state-of-the-art methods and achieve promising learning of the spatial position of objects11Code is at https://github.com/zhiqic/Rethinking-Counting. Zhi-Qi Cheng, Qi Dai 0001, Jingkuan Song, Xiao Wu 0001, Alex Hauptmann 0001 |
CVPR | 5 |
| 2022 | Learning Selective Assignment Network for Scene-Aware Vehicle DetectionabstractDeep learning has shown remarkable success in data-driven vehicle detection, relying on collected training samples from known scenes. A challenging problem arises when these detectors handle agnostic scenes, while keeping the performance of previous ones. To address this issue, a feasible remedy is to learn a set of domain-adaptive detectors by aligning the features from one scene to another. However, the improvement obtained in this way is inflexible despite the progress in object detection. An important reason is that the memory sizes grow massively with deliberately saving all scenes-independent detectors, while ignoring the relationship among different scenes. In this paper, we aim to bridge the gap between scene diversification and object consistency for scene-aware vehicle detection. Specifically, a novel structured network is proposed to integrate selective assignment of scene-specific parameters into the vehicle detection framework. Extensive experiments conducted on different scenes including BDD, Cityscapes-car, CARPK, etc, demonstrate that the proposed method achieves impressive performance, while keeping the performance of previous scenes as the scene changes. Zhenting Wang, Wei Li 0110, Xiao Wu 0001, Luhan Sheng |
ICIP | 3 |
| 2022 | Learning Graph-based Residual Aggregation Network for Group Activity RecognitionabstractGroup activity recognition aims to understand the overall behavior performed by a group of people. Recently, some graph-based methods have made progress by learning the relation graphs among multiple persons. However, the differences between an individual and others play an important role in identifying confusable group activities, which have not been elaborately explored by previous methods. In this paper, a novel Graph-based Residual AggregatIon Network (GRAIN) is proposed to model the differences among all persons of the whole group, which is end-to-end trainable. Specifically, a new local residual relation module is explicitly proposed to capture the local spatiotemporal differences of relevant persons, which is further combined with the multi-graph relation networks. Moreover, a weighted aggregation strategy is devised to adaptively select multi-level spatiotemporal features from the appearance-level information to high level relations. Finally, our model is capable of extracting a comprehensive representation and inferring the group activity in an end-to-end manner. The experimental results on two popular benchmarks for group activity recognition clearly demonstrate the superior performance of our method in comparison with the state-of-the-art methods. Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Zhaoquan Yuan |
IJCAI | 3 |
| 2022 | Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionabstractLearning spatial and temporal relations among people plays an important role in recognizing group activity. Recently, transformer-based methods have become popular solutions due to the proposal of self-attention mechanism. However, the person-level features are fed directly into the self-attention module without any refinement. Moreover, group activity in a clip often involves unbalanced spatio-temporal interactions, where only a few persons with special actions are critical to identifying different activities. It is difficult to learn the spatio-temporal interactions due to the lack of elaborately modeling the action dependencies among all people. In this paper, a novel Action-guided Spatio-Temporal transFormer (ASTFormer) is proposed to capture the interaction relations for group activity recognition by learning action-centric aggregation and modeling spatio-temporal action dependencies. Specifically, ASTFormer starts with assigning all persons in each frame to the latent actions, while an action-centric aggregation strategy is performed by weighting the sum of residuals for each latent action under the supervision of global action information. Then, a dual-branch transformer is proposed to refine the inter- and intra-frame action-level features, where two encoders with the self-attention mechanism are employed to select important tokens. Next, a semantic action graph is explicitly devised to model the dynamic action-wise dependencies. Finally, our model is capable of boosting group activity recognition by fusing these important cues, while only requiring video-level action labels. Extensive experiments on two popular benchmarks (Volleyball and Collective Activity) demonstrate the superior performance of our method in comparison with the state-of-the-art methods using only raw RGB frames as input. Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Xian-Jun Du, Jian-Jun Qiao |
ACM Multimedia | 3 |
| 2022 | Domain-Specific Conditional Jigsaw Adaptation for Enhancing transferability and DiscriminabilityabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a label-rich source domain to a target domain where the label is unavailable. Existing approaches tend to reduce the distribution discrepancy between the source and target domains or assign the pseudo target labels to implement a self-training strategy. However, the transferability or discriminability lackage of the traditional methods results in the limited ability to generalize on the target domain. To remedy this issue, a novel unsupervised domain adaptation framework called Domain-specific Conditional Jigsaw Adaptation Network (DCJAN) is proposed for UDA, which simultaneously encourages the network to extract transferable and discriminative features. To improve the discriminability, a conditional jigsaw module is presented to reconstruct class-aware features of the original images by reconstructing that of corresponding shuffled images. Moreover, in order to enhance the transferability, a domain-specific jigsaw adaptation is proposed to deal with the domain gaps, which utilizes the prior knowledge of jigsaw puzzles to reduce mismatching. It trains conditional jigsaw modules for each domain and updates the shared feature extractor to make the domain-specific conditional jigsaw modules could perform well not only on the corresponding domain but also on the other domain. A consistent conditioning strategy is proposed to ensure the safe training of conditional jigsaw. Experiments conducted on the widely-used Office-31, Office-Home, VisDA-2017, and DomainNet datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods. Qi He 0007, Zhaoquan Yuan, Xiao Wu 0001, Jun-Yan He |
ACM Multimedia | 3 |
| 2022 | Real-time Semantic Segmentation with Parallel Multiple Views Feature AugmentationabstractReal-time semantic segmentation is essential for many practical applications, which utilizes attention-based feature aggregation into lightweight structures to improve accuracy and efficiency. However, existing attention-based methods ignore 1) high-level and low-level feature augmentation guided by spatial information, and 2) low-level feature augmentation guided by semantic context, so that feature gaps between multi-level features and noise of low-level spatial details still exist. To address these problems, a new real-time semantic segmentation network, called MvFSeg, is proposed. In MvFSeg, parallel convolution with multiple depths is designed as a context head to generate and integrate multi-view features with larger receptive fields. Moreover, MvFSeg designs multiple views feature augmentation strategies that exploit spatial and semantic guidance for shallow and deep feature augmentation in an inter-layer and intra-layer manner. These strategies eliminate feature gaps between multi-level features, filter out the noise of spatial details, and provide spatial and semantic guidance for multi-level features. By combining multi-view features and augmented features from the lightweight networks with progressive dense aggregation structures, MvFSeg effectively captures invariance at various scales and generates high-quality segmentation results. Experiments conducted on Cityscapes and CamVid benchmark show that MvFSeg outperforms existing state-of-the-art methods. Jian-Jun Qiao, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Ji Zhang 0027 |
ACM Multimedia | 3 |
| 2022 | CrossNet: Boosting Crowd Counting with LocalizationabstractGenerating high-quality density maps is a crucial step in crowd counting. It is obvious that exploiting the head location of the people can naturally highlight the crowded area and eliminate the interference of background noise. However, existing crowd counting methods are still tricky to reasonably use location in density generation. In this paper, a novel location-guided framework named CrossNet is proposed for crowd counting, which integrates location supervision into density maps through dual-branch joint training. First, a new branching network is proposed to localize the potential positions of pedestrians. With the help of supervision induced from the localization branch, Location Enhancement (LE) module is designed to obtain high-quality density maps by positioning foreground regions. Second, Adaptive Density Awareness Attention (ADAA) module is engaged to enhance localization accuracy, which can efficiently use the density of the counting branch to adaptively capture the error-prone dense areas of the location maps. Finally, Density Awareness Localization (DAL) loss is offered to allocate attention to the crowd density levels, which delivers more focus on regions with high densities and less concentration on areas with low densities. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches both in crowd counting and crowd localization. Ji Zhang 0027, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Jian-Jun Qiao |
ACM Multimedia | 3 |
| 2022 | User-dependent interactive light field video streaming system
Qiang Peng, Eric Wang 0001, Wei Xiang 0001, Xiao Wu 0001 |
Multim. Tools Appl. | 5 |
| 2022 | Correction to: User‑dependent interactive light field video streaming system
Qiang Peng, Eric Wang 0001, Wei Xiang 0001, Xiao Wu 0001 |
Multim. Tools Appl. | 5 |
| 2022 | Learning-based high-efficiency compression framework for light field videos
Wei Xiang 0001, Eric Wang 0001, Qiang Peng, Pan Gao 0001, Xiao Wu 0001 |
Multim. Tools Appl. | 6 |
| 2022 | A Probabilistic Approach for Cooperative Computation Offloading in MEC-Assisted Vehicular NetworksabstractMobile edge computing (MEC) has been an effective paradigm for supporting computation-intensive applications by offloading resources at network edge. Especially in vehicular networks, the MEC server, is deployed as a small-scale computation server at the roadside and offloads computation-intensive task to its local server. However, due to the unique characteristics of vehicular networks, including high mobility of vehicles, dynamic distribution of vehicle densities and heterogeneous capacities of MEC servers, it is still challenging to implement efficient computation offloading mechanism in MEC-assisted vehicular networks. In this article, we investigate a novel scenario of computation offloading in MEC-assisted architecture, where task upload coordination between multiple vehicles, task migration between MEC/cloud servers and heterogeneous computation capabilities of MEC/cloud severs, are comprehensively investigated. On this basis, we formulate cooperative computation offloading (CCO) problem by modeling the procedure of task upload, migration and computation based on queuing theory, which aims at minimizing the delay of task completion. To tackle the CCO problem, we propose a probabilistic computation offloading (PCO) algorithm, which enables MEC server to independently make online scheduling based on the derived allocation probability. Specifically, the PCO transforms the objective function into augmented Lagrangian and achieves the optimal solution in an iterative way, based on a convex framework called Alternating Direction Method of Multipliers (ADMM). Last but not the least, we implement the simulation model. The comprehensive simulation results show the superiority of the proposed algorithm under a wide range of scenarios. Penglin Dai, Kaiwen Hu, Xiao Wu 0001, Huanlai Xing, Fei Teng 0001, Zhaofei Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | SWNet: A Deep Learning Based Approach for Splashed Water Detection on RoadabstractAdverse weather conditions seriously threaten the traffic safety, especially for rainy days with the ponding water on the road surface, which potentially result in vehicle crashes, person injuries and crash fatalities. Automatic splashed water detection based on surveillance videos is an attractive way to effectively prevent the traffic accidents. However, surveillance videos exhibit great variations with lighting changes, illumination conditions and complex backgrounds, which pose great difficulties in automatic recognition. In this paper, a novel deep learning based approach is proposed to detect the splashed water. To the best of our knowledge, this is the first work on this topic based on deep learning. An effective semantic segmentation network, called SWNet, is novelly proposed to extract the potential splashed water regions. An encoder-decoder structure is designed to capture the visual characteristics of splashed water. SWNet achieves high efficiency by reusing pooling indices and adopting the light-weight decoder. With the multi-scale feature fusion structure, SWNet integrates the coarse semantic information and detailed appearance information, which significantly boosts the accuracy and refines the edge segmentation. A weighted cross entropy loss for splashed water is adopted to cope with the unbalanced distribution between splashed water and backgrounds. Moreover, a splashed water attention module is designed to focus on the salient regions of moving vehicles and splashed water, by performing attention mechanism to integrate global contextual information in semantic segmentation. Experiments conducted on a newly collected splashed water dataset demonstrate the effectiveness and efficiency of the proposed approach, which outperforms the state-of-the-art methods. Jian-Jun Qiao, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Qiang Peng |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Meta-Learning Causal Feature Selection for Stable PredictionabstractConventional predictive models in machine learning are based on I.I.D. hypothesis between training and testing data. However, such a hypothesis is fragile in the real world, and the model minimizing empirical errors on training data does not perform well on testing data, which makes the prediction unstable. This instability can be found widely in domain generalization, active learning, and transfer learning, etc. In this paper, we propose a novel Meta-learning Causal Feature Selection (MCFS) model for general Non-I.I.D. image classification. In MCFS, we jointly optimize a convolutional network and a causal parameter for identifying causal variables on meta-training and meta-testing data which simulate the distribution shifts in Non-I.I.D. problems. Extensive experiments conducted on public VLCS and NICO datasets demonstrate the effectiveness of the proposed MCFS, which outperforms the state-of-the-art methods. Zhaoquan Yuan, Xiao Wu 0001, Bing-Kun Bao, Changsheng Xu |
ICME | 3 |
| 2021 | Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC-Empowered Vehicular NetworksabstractMobile edge computing (MEC) has been an effective paradigm to support real-time computation-intensive vehicular applications. However, due to highly dynamic vehicular topology, these existing centralized-based or distributed-based scheduling algorithms requiring high communication overhead, are not suitable for task offloading in vehicular networks. Therefore, we investigate a novel service scenario of MEC-based vehicular crowdsourcing, where each MEC server is an independent agent and responsible for making scheduling of processing traffic data sensed by crowdsourcing vehicles. On this basis, we formulate a data-driven task offloading problem by jointly optimizing offloading decision and bandwidth/computation resource allocation, and renting cost of heterogeneous servers, such as powerful vehicles, MEC servers and cloud, which is a mixed-integer programming problem and NP-hard. To reduce high time-complexity, we propose the solution in two stages. First, we design an asynchronous deep Q-learning to determine offloading decision, which achieves fast convergence by training the local DQN model at each agent in parallel and uploading for global model update asynchronously. Second, we decompose the remaining resource allocation problem into several independent subproblems and derive optimal analytic formula based on convex theory. Lastly, we build a simulation model and conduct comprehensive simulation, which demonstrates the superiority of the proposed algorithm. Penglin Dai, Kaiwen Hu, Xiao Wu 0001, Huanlai Xing, Zhaofei Yu |
INFOCOM | 3 |
| 2021 | Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal TransformerabstractDiagram question answering (DQA) is an effective way to evaluate the reasoning ability for diagram semantic understanding, which is a very challenging task and largely understudied compared with natural images. Existing separate two-stage methods for DQA are limited in ineffective feedback mechanisms. To address this problem, in this paper, we propose a novel structural parsing-integrated Hierarchical Multi-Task Learning (HMTL) model for diagram question answering based on a multi-modal transformer framework. In the proposed paradigm of multi-task learning, the two tasks of diagram structural parsing and question answering are in the different semantic levels and equipped with different transformer blocks, which constituents a hierarchical architecture. The structural parsing module encodes the information of constituents and their relationships in diagrams, while the diagram question answering module decodes the structural signals and combines question-answers to infer correct answers. Visual diagrams and textual question-answers are interplayed in the multi-modal transformer, which achieves cross-modal semantic comprehension and reasoning. Extensive experiments on the benchmark AI2D and FOODWEBS datasets demonstrate the effectiveness of our proposed HMTL over other state-of-the-art methods. Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu |
ACM Multimedia | 3 |
| 2021 | Vehicle Counting Network with Attention-based Mask Refinement and Spatial-awareness Block LossabstractVehicle counting aims to calculate the number of vehicles in congested traffic scenes. Although object detection and crowd counting have made tremendous progress with the development of deep learning, vehicle counting remains a challenging task, due to scale variations, viewpoint changes, inconsistent location distributions, diverse visual appearances and severe occlusions. In this paper, a well-designed Vehicle Counting Network (VCNet) is novelly proposed to alleviate the problem of scale variation and inconsistent spatial distribution in congested traffic scenes. Specifically, VCNet is composed of two major components: (i) To capture multi-scale vehicles across different types and camera viewpoints, an effective multi-scale density map estimation structure is designed by building an attention-based mask refinement module. The multi-branch structure with hybrid dilated convolution blocks is proposed to assign receptive fields to generate multi-scale density maps. To efficiently aggregate multi-scale density maps, the attention-based mask refinement is well-designed to highlight the vehicle regions, which enables each branch to suppress the scale interference from other branches. (ii) In order to capture the inconsistent spatial distributions, a spatial-awareness block loss (SBL) based on the region-weighted reward strategy is proposed to calculate the loss of different spatial regions including sparse, congested and occluded regions independently by dividing the density map into different regions. Extensive experiments conducted on three benchmark datasets, TRANCOS, VisDrone2019 Vehicle and CVCSet demonstrate that the proposed VCNet outperforms the state-of-the-art approaches in vehicle counting. Moreover, the proposed idea can be applicable for crowd counting, which produces competitive results on ShanghaiTech crowd counting dataset. Ji Zhang 0027, Jian-Jun Qiao, Xiao Wu 0001, Wei Li 0110 |
ACM Multimedia | 3 |
| 2021 | Contrastive Learning in Frequency Domain for Non-I.I.D. Image Classification
Huan Shao 0005, Zhaoquan Yuan, Xiao Wu 0001 |
MMM (1) | 4 |
| 2021 | A novel class restriction loss for unsupervised domain adaptation
Qi He 0007, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He |
Neurocomputing | 3 |
| 2021 | DB-LSTM: Densely-connected Bi-directional LSTM for human action recognition
Jun-Yan He, Xiao Wu 0001, Zhi-Qi Cheng, Zhaoquan Yuan, Yu-Gang Jiang 0001 |
Neurocomputing | 2 |
| 2021 | MGSeg: Multiple Granularity-Based Real-Time Semantic Segmentation NetworkabstractRecent works on semantic segmentation witness significant performance improvement by utilizing global contextual information. In this paper, an efficient multi-granularity based semantic segmentation network (MGSeg) is proposed for real-time semantic segmentation, by modeling the latent relevance between multi-scale geometric details and high-level semantics for fine granularity segmentation. In particular, a light-weight backbone ResNet-18 is first adopted to produce the hierarchical features. Hybrid Attention Feature Aggregation (HAFA) is designed to filter the noisy spatial details of features, acquire the scale-invariance representation, and alleviate the gradient vanishing problem of the early-stage feature learning. After aggregating the learned features, Fine Granularity Refinement (FGR) module is employed to explicitly model the relationship between the multi-level features and categories, generating proper weights for fusion. More importantly, to meet the real-time processing, a series of light-weight strategies and simplified structures are applied to accelerate the efficiency, including light-weight backbone, channel compression, narrow neck structure, and so on. Extensive experiments conducted on benchmark datasets Cityscapes and CamVid demonstrate that the proposed method achieves the state-of-the-art performance, 77.8%@50fps and 72.7%@127fps on Cityscapes and CamVid datasets, respectively, having the capability for real-time applications. Jun-Yan He, Shi-Hua Liang, Xiao Wu 0001, Bo Zhao 0032, Lei Zhang 0006 |
IEEE Trans. Image Process. | 3 |
| 2021 | Adversarial Multimodal Network for Movie Story Question AnsweringabstractVisual question answering by using information from multiple modalities has attracted more and more attention in recent years. However, it is a very challenging task, as the visual content and natural language have quite different statistical properties. In this work, we present a method called Adversarial Multimodal Network (AMN) to better understand video stories for question answering. In AMN, we propose to learn multimodal feature representations by finding a more coherent subspace for video clips and the corresponding texts (e.g., subtitles and questions) based on generative adversarial networks. Moreover, a self-attention mechanism is developed to enforce our newly introduced consistency constraint in order to preserve the self-correlation between the visual cues of the original video clips in the learned multimodal representations. Extensive experiments on the benchmark MovieQA and TVQA datasets show the effectiveness of our proposed AMN over other published state-of-the-art methods. Zhaoquan Yuan, Lixin Duan, Xiao Wu 0001, Changsheng Xu |
IEEE Trans. Multim. | 5 |
| 2020 | CODAN: Counting-driven Attention Network for Vehicle Detection in Congested ScenesabstractAlthough recent object detectors have shown excellent performance for vehicle detection, they are incompetent for scenarios with a relatively large number of vehicles. In this paper, we explore the dense vehicle detection given the number of vehicles. Existing crowd counting methods cannot directly applied for dense vehicle detection due to insufficient description of density map, and the lack of effective constraint for mining the spatial awareness of dense vehicles. Inspired by these observations, a conceptually simple yet efficient framework, called CODAN, is proposed for dense vehicle detection. The proposed approach is composed of three major components: (i) an efficient strategy for generating multi-scale density maps (MDM) is designed to represent the vehicle counting, which can capture the global semantics and spatial information of dense vehicles, (ii) a multi-branch attention module (MAM) is proposed to bridging the gap between object counting and vehicle detection framework, (iii) with the well-designed density maps as explicit supervision, an effective counting-awareness loss (C-Loss) is employed to guide the attention learning by building the pixel-level constrain. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods. The impressive results indicate that vehicle detection and counting can be mutually supportive, which is an important and meaningful finding. Wei Li 0110, Zhenting Wang, Xiao Wu 0001, Ji Zhang 0027, Qiang Peng, Hongliang Li 0001 |
ACM Multimedia | 3 |
| 2020 | Learning fashion compatibility across categories with deep multimodal neural networks
Guang-Lu Sun, Jun-Yan He, Xiao Wu 0001, Bo Zhao 0032, Qiang Peng |
Neurocomputing | 3 |
| 2019 | Joint Resource Optimization for Adaptive Multimedia Services in MEC-Based Vehicular NetworksabstractMobile edge computing (MEC) has been an emerging paradigm to support low-latency applications in vehicular networks by offloading resources at network edge. However, it is still challenging to apply MEC- based architecture to implement multimedia services due to varying wireless communication, high vehicle mobility and heterogeneous resource integration. In this paper, we investigate adaptive-bitrate (ABR)-based multimedia services (MS) in MEC-based vehicular networks, where each multimedia file is divided into multiple chunks and can be requested at different bitrate levels. Further, MEC servers can satisfy local vehicular requests by integrating heterogeneous cache and communication resources. Based on the above observation, we formulate joint resource optimization (JSO) problem by synthesizing cache placement, wireless bandwidth allocation and chunk quality adaptation. On this basis, we propose a reinforcement- learning-based cache placement (RLCP) algorithm, which determines the optimal offloaded chunks by learning the global knowledge of cache reward in an iterative way. Further, we design an adaptive-quality- based chunk selection (AQCS) algorithm, which can be adaptive to time-varying wireless channel by dynamically adjusting bandwidth allocation and quality level based on real-time service workload. Lastly, we build the simulation model and conduct an extensive performance evaluation, which demonstrates the superiority of proposed algorithms. Penglin Dai, Kai Liu 0001, Xiao Wu 0001, Huanlai Xing, Victor C. S. Lee |
GLOBECOM | 3 |
| 2019 | Generative Adversarial Networks Based Error Concealment for Low Resolution VideoabstractIn this paper, a novel deep generative model-based approach for video error concealment is proposed. Our method is comprised of completion network and two critics. The frame completion network is trained to fool the both the local and global critics, which requires completion network to conceal frame distortions with regard to overall consistency as well as in details. Specifically, mask attention convolution layer is proposed, which utilize not only the temporal information of the previous frame, but also the intact pixels of the current distorted frame to mask and re-normalize convolution features. Then, both qualitative and quantitative experiments validate the effectiveness and generality of our approach in advancing the error concealment on low resolution video. Chongyang Xiang, Chuan Yan, Qiang Peng, Xiao Wu 0001 |
ICASSP | 5 |
| 2019 | A Learning Algorithm for Real-Time Service in Vehicular Networks with Mobile-Edge ComputingabstractMobile edge computing (MEC) is an emerging paradigm to offload the server-side resources closer to the mobile terminals compared with cloud-based computing. However, due to highly vehicular mobility and limited wireless coverage, it is challenging to apply off-the-shelf MEC-based architecture to support the real-time services in vehicular networks, especially when the vehicle density changes dynamically. Hence, this paper investigates a novel service scenario in an MEC-based architecture, where the local MEC server has to complete the real-time services of mobile vehicles in its service range. On this basis, we formulate a novel problem of distributed real-time service scheduling (DRSS) by comprehensively considering the delay requirements of real-time services, the heterogeneous computing capabilities of MEC servers and the mobility features of vehicles, which targets at maximizing the service ratio. To resolve such an issue, we propose a multi-agent reinforcement learning algorithm called Utility-based Learning (UL), in which each local MEC server selects the optimal solution by learning the global knowledge online. Specifically, a utility table is established to determine the optimal solution by estimating the pending delay of service request at each MEC server and it will be updated periodically based on the feedback signal from the assigned MEC server. Lastly, we build the simulation model and conduct an extensive performance evaluation, which demonstrates the superiority of the proposed algorithm. Penglin Dai, Kai Liu 0001, Xiao Wu 0001, Huanlai Xing, Zhaofei Yu, Victor C. S. Lee |
ICC | 3 |
| 2019 | Learning Spatial Awareness to Improve Crowd CountingabstractThe aim of crowd counting is to estimate the number of people in images by leveraging the annotation of center positions for pedestrians' heads. Promising progresses have been made with the prevalence of deep Convolutional Neural Networks. Existing methods widely employ the Euclidean distance (i.e., L2loss) to optimize the model, which, however, has two main drawbacks: (1) the loss has difficulty in learning the spatial awareness (i.e., the position of head) since it struggles to retain the high-frequency variation in the density map, and (2) the loss is highly sensitive to various noises in crowd counting, such as the zeromean noise, head size changes, and occlusions. Although the Maximum Excess over SubArrays (MESA) loss has been previously proposed by [16] to address the above issues by finding the rectangular subregion whose predicted density map has the maximum difference from the ground truth, it cannot be solved by gradient descent, thus can hardly be integrated into the deep learning framework. In this paper, we present a novel architecture called SPatial Awareness Network (SPANet) to incorporate spatial context for crowd counting. The Maximum Excess over Pixels (MEP) loss is proposed to achieve this by finding the pixel-level subregion with high discrepancy to the ground truth. To this end, we devise a weakly supervised learning scheme to generate such region with a multi-branch architecture. The proposed framework can be integrated into existing deep crowd counting methods and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that our method can significantly improve the performance of baselines. More remarkably, our approach outperforms the state-of-the-art methods on all benchmark datasets. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Alex Hauptmann 0001 |
ICCV | 4 |
| 2019 | Improving the Learning of Multi-column Convolutional Neural Network for Crowd CountingabstractTremendous variation in the scale of people/head size is a critical problem for crowd counting. To improve the scale invariance of feature representation, recent works extensively employ Convolutional Neural Networks with multi-column structures to handle different scales and resolutions. However, due to the substantial redundant parameters in columns, existing multi-column networks invariably exhibit almost the same scale features in different columns, which severely affects counting accuracy and leads to overfitting. In this paper, we attack this problem by proposing a novel Multicolumn Mutual Learning (McML) strategy. It has two main innovations: 1) A statistical network is incorporated into the multi-column framework to estimate the mutual information between columns, which can approximately indicate the scale correlation between features from different columns. By minimizing the mutual information, each column is guided to learn features with different image scales. 2) We devise a mutual learning scheme that can alternately optimize each column while keeping the other columns fixed on each mini-batch training data. With such asynchronous parameter update process, each column is inclined to learn different feature representation from others, which can efficiently reduce the parameter redundancy and improve generalization ability. More remarkably, McML can be applied to all existing multi-column networks and is end-to-end trainable. Extensive experiments on four challenging benchmarks show that McML can significantly improve the original multi-column networks and outperform the other state-of-the-art approaches. Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2019 | CRA-Net: Composed Relation Attention Network for Visual Question AnsweringabstractThe task of Visual Question Answering (VQA) is to answer a natural language question tied to the content of a visual image. Most existing VQA models either apply attention mechanism to locate the relevant object regions and/or utilize the off-the-shelf methods of the relation reasoning to detect object relations. However, they 1) mostly encode the simple relations which cannot sufficiently provide sophisticated knowledge for answering complicated visual questions; 2) seldom leverage the harmony cooperation of the object appearance feature and relation feature. To address these problems, we propose a novel end-to-end VQA model, termed Composed Relation Attention Network (CRA-Net ). In specific, we devise two question-adaptive relation attention modules that can extract not only the fine-grained and precise binary relations but also the more sophisticated trinary relations. Both kinds of question-related relations can reveal deeper semantics, thereby enhancing the reasoning ability in question answering. Furthermore, our CRA-Net also combines the object appearance feature with the relation feature under the guidance of the corresponding question, which can reconcile the two types of features effectively. Extensive experiments on two large benchmark datasets, VQA-1.0 and VQA-2.0, demonstrate that our proposed model outperforms state-of-the-art approaches. Yang Yang 0002, Zheng Wang 0044, Xiao Wu 0001, Zi Huang |
ACM Multimedia | 4 |
| 2019 | Multi-objective Optimization for Network Resource Management in Heterogeneous Vehicular NetworksabstractHeterogeneous network integration is a promising technique to support efficient data services in vehicular networks. However, due to highly dynamics of vehicular mobility and heterogeneous performance of wireless interfaces, it is still challenging to design an efficient scheduling policy for information services in vehicular networks. In this paper, we propose a centralized service architecture for managing heterogeneous network resources. Particularly, we comprehensively investigate the heterogeneity of networks, as well as the diversity of service requests. On this basis, we formulate the heterogeneous network resource management (HNRM) problem as a multiple-objective problem, which aims at minimizing both the service delay and the network access cost simultaneously. Then, we propose a packet-encoding based multi-objective algorithm (PEMA), which consists of two components: packet encoding for data broadcast and multiobjective algorithm for network interface selection. Specifically, for improving bandwidth efficiency, we develop a multiple-packet encoding (MPE) technique to serve more requests simultaneously. For network selection, we propose a multi-objective evolutionary mechanism to further minimize both the service delay and the network access cost via population evolution. Finally, we give a comprehensive performance evaluation to demonstrate the superiority of PEMA under a wide range of scenarios. Penglin Dai, Kai Liu 0001, Xiao Wu 0001, Huanlai Xing, Victor C. S. Lee |
WCNC | 3 |
| 2019 | Cooperative Temporal Data Dissemination in SDN-Based Heterogeneous Vehicular NetworksabstractHeterogeneous network resources are expected to cooperate with each other to support temporal data services in vehicular networks. However, it is challenging to implement an efficient data scheduling strategy due to the following factors: first, there are different time constraints on services, which are imposed by the application requirements of both temporal data quality and transmission delay; second, the heterogeneity of wireless interfaces further complicates the transmission task assignment in dynamic vehicular environments. Therefore, this paper proposes an software-defined network-based architecture to enable unified management on heterogeneous network resources. Then, we formulate the cooperative temporal data dissemination (CTDD) problem by considering the property of temporal data, the heterogeneity of wireless interfaces, and the delay constraints on service requests. Further, we prove the NP-hardness of the CTDD by constructing a polynomial-time reduction from a well know NP-hard problem, classical knapsack problem. On this basis, we design a heuristic algorithm called priority-based task assignment (PTA), which synthesizes dynamic task assignment, broadcast efficiency, and service deadline into priority design. Accordingly, PTA is able to adaptively distribute broadcast tasks of each request among multiple interfaces, so as to improve overall system performance. Last but not least, we build the simulation model and implement the proposed algorithm. The comprehensive simulation results show the superiority of the proposed algorithm under a wide range of scenarios. Penglin Dai, Kai Liu 0001, Xiao Wu 0001, Zhaofei Yu, Huanlai Xing, Victor C. S. Lee |
IEEE Internet Things J. | 3 |
| 2019 | Temporal Information Services in Large-Scale Vehicular Networks Through Evolutionary Multi-Objective OptimizationabstractTemporal information services are critical in implementing emerging intelligent transportation systems. Nevertheless, it is challenging to realize timely temporal data update and dissemination due to an intermittent wireless connection and a limited communication bandwidth in dynamic vehicular networks. Some previous studies have considered the temporal data dissemination in vehicular networks, but they are limited to the service region, which is inside the coverage of roadside units. To enhance system scalability, it is imperative to exploit the synergic effect of vehicle-to-infrastructure (V2I) and vehicle-to-vehicle (V2V) communications for providing efficient temporal information services in such an environment. With the above motivations, we propose a novel system architecture to enable efficient data scheduling in hybrid V2I/V2V communications by having the global knowledge of network resources of the system. On this basis, we formulate a temporal data upload and dissemination (TDUD) problem, aiming at optimizing two conflict objectives simultaneously, which are enhancing the data quality and improving the delivery ratio. Furthermore, we propose an evolutionary multi-objective algorithm calledMO-TDUD, which consists of a decomposition scheme for handling multiple objectives, a scalable chromosome representation forTDUDsolution encoding, and an evolutionary operator designed forTDUDsolution reproduction. The proposedMO-TDUDcan be adaptive to different requirements on data quality and delivery ratio by selecting the best solution from the derived Pareto solutions. Last but not least, we build the simulation model and implementMO-TDUDfor performance evaluation. The comprehensive simulation results demonstrate the superiority of the proposed solution. Penglin Dai, Kai Liu 0001, Liang Feng 0001, Haijun Zhang 0002, Victor C. S. Lee, Sang Hyuk Son, Xiao Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2019 | BranchGAN: Unsupervised Mutual Image-to-Image Transfer With A Single Encoder and Dual DecodersabstractImage-to-image translation is a fundamental task for a wide range of applications, such as image style transfer, video effect generation, cross-domain retrieval, etc. Due to the limited number of labeled data, complex scenes, abstract semantics and various involved domains, image translation remains a challenging task. Compared to the supervised approaches for image translation that need a large collection of paired images for training, the unsupervised methods can significantly reduce the training cost. In this paper, an unsupervised end-to-end generative adversarial network is proposed, namedBranchGAN, for mutual image-to-image transfer between two domains. A structure with one single encoder and dual decoders is novelly proposed to capture the cross-domain distributions and generate the images in both domains. Three factors, that is, pixel-level overall style, region semantics, and domain distinguishability are comprehensively considered to constrain the training process of the proposed model, corresponding toreconstruction loss,encoding loss, andadversarial loss, respectively. Experiments conducted on three benchmark datasets demonstrate the effectiveness of the proposed method that outperforms the unsupervised state-of-the-art approaches and has the competitive performance as the supervised method. Yi-Fan Zhou, Runhao Jiang, Xiao Wu 0001, Jun-Yan He, Shuang Weng, Qiang Peng |
IEEE Trans. Multim. | 3 |
| 2018 | A Novel Weighted Boundary Matching Error Concealment Schema for HEVCabstractIn this paper, a novel weighted boundary matching error concealment schema for HEVC is proposed, which is based on the CU depths and PU partitions in reference frame. Firstly, the information of CU depths in reference frames is used for lost slices. For each LCU in a lost slice, the LCUs surrounding to the co-located LCU are used to calculate summed CU-depth weight, which is used to determine the conceal order of each CU. Then, the co-located partition decision from the reference frame is adopted for PUs in each lost CU. The sequence of PUs to conceal is sorted based on the texture randomness index weight and the PU with the largest weight will be concealed next. Finally, the best estimated motion vector for the lost PU is selected for concealment. The experimental results show that our method achieves higher PSNR gains and has a better visual quality than the state-of-the-art methods. Chuan Yan, Qiang Peng, Xiao Wu 0001 |
ICIP | 5 |
| 2018 | Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model SearchabstractAs attribute leaning brings mid-level semantic properties for objects, it can benefit many traditional learning problems in multimedia and computer vision communities. When facing the huge number of attributes, it is extremely challenging to automatically design a generalizable neural network for other attribute learning tasks. Even for a specific attribute domain, the exploration of the neural network architecture is always optimized by a combination of heuristics and grid search, from which there is a large space of possible choices to be searched. In this paper, Generalizable Attribute Learning Model (GALM) is proposed to automatically design the neural networks for generalizable attribute learning. The main novelty of GALM is that it fully exploits the Multi-Task Learning and Reinforcement Learning to speed up the search procedure. With the help of parameter sharing, GALM is able to transfer the pre-searched architecture to different attribute domains. In experiments, we comprehensively evaluate GALM on 251 attributes from three domains: animals, objects, and scenes. Extensive experimental results demonstrate that GALM significantly outperforms the state-of-the-art attribute learning approaches and previous neural architecture search methods on two generalizable attribute learning scenarios. Zhi-Qi Cheng, Xiao Wu 0001, Siyu Huang, Jun-Xiu Li, Alex Hauptmann 0001, Qiang Peng |
ACM Multimedia | 2 |
| 2018 | Multi-View Image Generation from a Single-ViewabstractHow to generate multi-view images with realistic-looking appearance from only a single view input is a challenging problem. In this paper, we attack this problem by proposing a novel image generation model termed VariGANs, which combines the merits of the variational inference and the Generative Adversarial Networks (GANs). It generates the target image in a coarse-to-fine manner instead of a single pass which suffers from severe artifacts. It first performs variational inference to model global appearance of the object (e.g., shape and color) and produces coarse images of different views. Conditioned on the generated coarse images, it then performs adversarial learning to fill details consistent with the input and generate the fine images. Extensive experiments conducted on two clothing datasets, MVC and DeepFashion, have demonstrated that the generated images with the proposed VariGANs are more plausible than those generated by existing approaches, which provide more consistent global appearance as well as richer and sharper details. Bo Zhao 0032, Xiao Wu 0001, Zhi-Qi Cheng, Hao Liu 0003, Zequn Jie, Jiashi Feng |
ACM Multimedia | 2 |
| 2018 | Personalized clothing recommendation combining user social circle and fashion style consistency
Guang-Lu Sun, Zhi-Qi Cheng, Xiao Wu 0001, Qiang Peng |
Multim. Tools Appl. | 3 |
| 2018 | Hookworm Detection in Wireless Capsule Endoscopy Images With Deep LearningabstractAs one of the most common human helminths, hookworm is a leading cause of maternal and child morbidity, which seriously threatens human health. Recently, wireless capsule endoscopy (WCE) has been applied to automatic hookworm detection. Unfortunately, it remains a challenging task. In recent years, deep convolutional neural network (CNN) has demonstrated impressive performance in various image and video analysis tasks. In this paper, a novel deep hookworm detection framework is proposed for WCE images, which simultaneously models visual appearances and tubular patterns of hookworms. This is the first deep learning framework specifically designed for hookworm detection in WCE images. Two CNN networks, namely edge extraction network and hookworm classification network, are seamlessly integrated in the proposed framework, which avoid the edge feature caching and speed up the classification. Two edge pooling layers are introduced to integrate the tubular regions induced from edge extraction network and the feature maps from hookworm classification network, leading to enhanced feature maps emphasizing the tubular regions. Experiments have been conducted on one of the largest WCE datasets with WCE images, which demonstrate the effectiveness of the proposed hookworm detection framework. It significantly outperforms the state-of-the-art approaches. The high sensitivity and accuracy of the proposed method in detecting hookworms shows its potential for clinical application. Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Qiang Peng, Ramesh Jain 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images
Zhi-Qi Cheng, Xiao Wu 0001, Yang Liu 0155, Xian-Sheng Hua 0001 |
CVPR | 2 |
| 2017 | Memory-Augmented Attribute Manipulation Networks for Interactive Fashion SearchabstractWe introduce a new fashion search protocol where attribute manipulation is allowed within the interaction between users and search engines, e.g. manipulating the color attribute of the clothing from red to blue. It is particularly useful for image-based search when the query image cannot perfectly match users expectation of the desired product. To build such a search engine, we propose a novel memory-augmented Attribute Manipulation Network (AMNet) which can manipulate image representation at the attribute level. Given a query image and some attributes that need to modify, AMNet can manipulate the intermediate representation encoding the unwanted attributes and change them to the desired ones through following four novel components: (1) a dual-path CNN architecture for discriminative deep attribute representation learning, (2) a memory block with an internal memory and a neural controller for prototype attribute representation learning and hosting, (3) an attribute manipulation network to modify the representation of the query image with the prototype feature retrieved from the memory block, (4) a loss layer which jointly optimizes the attribute classification loss and a triplet ranking loss over triplet images for facilitating precise attribute manipulation and image retrieving. Extensive experiments conducted on two large-scale fashion search datasets, i.e. DARN and DeepFashion, have demonstrated that AMNet is able to achieve remarkably good performance compared with well-designed baselines in terms of effectiveness of attribute manipulation and search accuracy. Bo Zhao 0032, Jiashi Feng, Xiao Wu 0001, Shuicheng Yan |
CVPR | 3 |
| 2017 | On the Selection of Anchors and Targets for Video HyperlinkingabstractA problem not well understood in video hyperlinking is what qualifies a fragment as an anchor or target. Ideally, anchors provide good starting points for navigation, and targets supplement anchors with additional details while not distracting users with irrelevant, false and redundant information. The problem is not trivial for intertwining relationship between data characteristics and user expectation. Imagine that in a large dataset, there are clusters of fragments spreading over the feature space. The nature of each cluster can be described by its size (implying popularity) and structure (implying complexity). A principle way of hyperlinking can be carried out by picking centers of clusters as anchors and from there reach out to targets within or outside of clusters with consideration of neighborhood complexity. The question is which fragments should be selected either as anchors or targets, in one way to reflect the rich content of a dataset, and meanwhile to minimize the risk of frustrating user experience. This paper provides some insights to this question from the perspective of hubness and local intrinsic dimensionality, which are two statistical properties in assessing the popularity and complexity of data space. Based these properties, two novel algorithms are proposed for low-risk automatic selection of anchors and targets. Zhi-Qi Cheng, Hao Zhang 0047, Xiao Wu 0001, Chong-Wah Ngo |
ICMR | 3 |
| 2017 | Sketch Recognition with Deep Visual-Sequential Fusion ModelabstractIn this paper, a deep end-to-end network for sketch recognition, named Deep Visual-Sequential Fusion model (DVSF) is proposed to model the visual and sequential patterns of the strokes. To capture the intermediate states of sketches, a three-way representation learner is first utilized to extract the visual features. These deep features are simultaneously fed into the visual and sequential networks to capture spatial and temporal properties, respectively. More specifically, visual networks are novelly proposed to learn the stroke patterns by stacking the Residual Fully-Connected (R-FC) layers, which integrate ReLU and Tanh activation functions to achieve the sparsity and generalization ability. To learn the patterns of stroke order, sequential networks are constructed by Residual Long Short-Term Memory (R-LSTM) units, which optimize the network architecture by skip connection. Finally, the visual and sequential representations of the sketches are seamlessly integrated with a fusion layer to obtain the final results. Experiments conducted on the benchmark sketch dataset TU-Berlin demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art approaches. Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Bo Zhao 0032, Qiang Peng |
ACM Multimedia | 2 |
| 2017 | Automatic content understanding with cascaded spatial-temporal deep framework for capsule endoscopy videos
Honghan Chen, Xiao Wu 0001, Tao Gan, Qiang Peng |
Neurocomputing | 2 |
| 2017 | Perception-based adaptive quantization for transform-domain Wyner-Ziv video coding
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Video eCommerce++: Toward Large Scale Online Video AdvertisingabstractThe prevalence of online videos provides an opportunity for e-commerce companies to recommend their products in videos. In this paper, we propose an online video advertising system named Video eCommerce ++, to exhibit appropriate product ads to particular users at proper time stamps of videos, which takes into account video semantics, user shopping preference, and viewing behavior feedback. First, an incremental co-relation regression (ICRR) model is novelly proposed to construct the semantic association between videos and products. To meet the requirement of online advertising, ICRR is implemented in an incremental way to reduce the time complexity. User preference diffusion (UPD) is induced under the framework of heterogeneous information network to construct user-product association from two different e-commerce platforms, Tmall and MagicBox, which alleviates the problems of data sparsity and cold start. A video scene importance model (VSIM) is proposed to model the scene importance by utilizing the user viewing behavior, so that ads can be embedded at the most attractive positions in the video stream. To combine the outputs of ICRR, UPD, and VSIM, a unified distributed heterogeneous relation matrix factorization (D-HRMF) is applied for online video advertising, which is efficiently conducted in parallel to address the real-time update problem, so that the whole system can be performed in real time. Extensive experiments conducted on a variety of online videos from Tmall MagicBox demonstrate that Video eCommerce++ significantly outperforms the state-of-the-art advertising methods, and can handle large-scale data in real time. Zhi-Qi Cheng, Xiao Wu 0001, Yang Liu 0155, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Integration of Diverse Data Sources for Spatial PM2.5 Data InterpolationabstractHeterogeneous data fusion from disparate geospatial sensors has drawn increasing attention in multimedia. Unfortunately, environmental sensors are usually sparsely and preferentially located, which restricts situation recognition of geographical regions and results in uncertainty in derived inferences. Spatial interpolation is an effective way to solve the problem of data sparsity, which demands the availability of related data sources. However, these data sources are usually in different resolutions, distributions, scales, and densities, which poses a major challenge in data integration. To address this problem, we present a novel spatial interpolation framework to incorporate diverse data sources and model the spatial processes explicitly at multiple resolutions. Spectral analysis is deployed to generate features at multiple spatial resolutions and to improve the interpolation accuracy at unobserved locations. A statistical operator based on the spatial Gaussian process is implemented and integrated into a geospatial situation recognition system, which can analyze heterogeneous spatio-temporal data streams derived from sensors. To verify the effectiveness and efficiency of the proposed framework, this framework is applied to the PM2.5 air pollution application. Experiments conducted in California, USA, demonstrate that the proposed method outperforms state-of-the-art approaches. Mengfan Tang, Xiao Wu 0001, Pranav Agrawal 0002, Siripen Pongpaichet, Ramesh Jain 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Diversified Visual Attention Networks for Fine-Grained Object ClassificationabstractFine-grained object classification attracts increasing attention in multimedia applications. However, it is a quite challenging problem due to the subtle interclass difference and large intraclass variation. Recently, visual attention models have been applied to automatically localize the discriminative regions of an image for better capturing critical difference, which have demonstrated promising performance. Unfortunately, without consideration of the diversity in attention process, most of existing attention models perform poorly in classifying fine-grained objects. In this paper, we propose a diversified visual attention network (DVAN) to address the problem of fine-grained object classification, which substantially relieves the dependency on strongly supervised information for learning to localize discriminative regions compared with attention-less models. More importantly, DVAN explicitly pursues the diversity of attention and is able to gather discriminative information to the maximal extent. Multiple attention canvases are generated to extract convolutional features for attention. An LSTM recurrent unit is employed to learn the attentiveness and discrimination of attention canvases. The proposed DVAN has the ability to attend the object from coarse to fine granularity, and a dynamic internal representation for classification is built up by incrementally combining the information from different locations and scales of the image. Extensive experiments conducted on CUB-2011, Stanford Dogs, and Stanford Cars datasets have demonstrated that the pro-posed DVAN achieves competitive performance compared to the state-of-the-art approaches, without using any prior knowledge, user interaction, or external resource in training and testing. Bo Zhao 0032, Xiao Wu 0001, Jiashi Feng, Qiang Peng, Shuicheng Yan |
IEEE Trans. Multim. | 2 |
| 2016 | Video eCommerce: Towards Online Video AdvertisingabstractThe prevalence of online videos provides an opportunity for e-commerce companies to exhibit their product ads in videos by recommendation. In this paper, we propose an advertising system named Video eCommerce to exhibit appropriate product ads to particular users at proper time stamps of videos, which takes into account video semantics, user shopping preference and viewing behavior feedback by a two-level strategy. At the first level, Co-Relation Regression (CRR) model is novelly proposed to construct the semantic association between keyframes and products. Heterogeneous information network (HIN) is adopted to build the user shopping preference from two different e-commerce platforms, Tmall and MagicBox, which alleviates the problems of data sparsity and cold start. In addition, Video Scene Importance Model (VSIM) utilizes the viewing behavior of users to embed ads at the most attractive position within the video stream. At the second level, taking the results of CRR, HIN and VSIM as the input, Heterogeneous Relation Matrix Factorization (HRMF) is applied for product advertising. Extensive evaluation on a variety of online videos from Tmall MagicBox demonstrates that Video eCommerce achieves promising performance, which significantly outperforms the state-of-the-art advertising methods. Zhi-Qi Cheng, Yang Liu 0155, Xiao Wu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |
| 2016 | Web video categorization using category-predictive classifiers and category-specific concept classifiers
Mehtab Afzal, Xiao Wu 0001, Honghan Chen, Yu-Gang Jiang 0001, Qiang Peng |
Neurocomputing | 2 |
| 2016 | Part-based clothing image annotation by visual neighbor retrieval
Guang-Lu Sun, Xiao Wu 0001, Qiang Peng |
Neurocomputing | 2 |
| 2016 | Detection of bird nests in overhead catenary system images for high-speed rail
Xiao Wu 0001, Ping Yuan, Qiang Peng, Chong-Wah Ngo, Jun-Yan He |
Pattern Recognit. | 1 |
| 2016 | Near-Duplicate Segments based news web video event mining
Chengde Zhang, Dianting Liu, Xiao Wu 0001, Guiru Zhao, Mei-Ling Shyu, Qiang Peng |
Signal Process. | 3 |
| 2016 | Integration of Visual Temporal Information and Textual Distribution Information for News Web Video Event MiningabstractNews web videos exhibit several characteristics, including a limited number of features, noisy text information, and error in near-duplicate keyframes (NDK) detection. Such characteristics have made the mining of the events from news web videos a challenging task. In this paper, a novel framework is proposed to better group the associated web videos to events. First, the data preprocessing stage performs feature selection and tag relevance learning. Next, multiple correspondence analysis is applied to explore the correlations between terms and events with the assistance of visual information. Cooccurrence and visual near-duplicate feature trajectory induced from NDKs are combined to calculate the similarity between NDKs and events. Finally, a probabilistic model is proposed for news web video event mining, where both visual temporal information and textual distribution information are integrated. Experiments on the news web videos from YouTube demonstrate that the integration of visual temporal information and textual distribution information outperforms the existing methods in the news web video event mining. Chengde Zhang, Xiao Wu 0001, Mei-Ling Shyu, Qiang Peng |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2016 | Automatic Hookworm Detection in Wireless Capsule Endoscopy ImagesabstractWireless capsule endoscopy (WCE) has become a widely used diagnostic technique to examine inflammatory bowel diseases and disorders. As one of the most common human helminths, hookworm is a kind of small tubular structure with grayish white or pinkish semi-transparent body, which is with a number of 600 million people infection around the world. Automatic hookworm detection is a challenging task due to poor quality of images, presence of extraneous matters, complex structure of gastrointestinal, and diverse appearances in terms of color and texture. This is the first few works to comprehensively explore the automatic hookworm detection for WCE images. To capture the properties of hookworms, the multi scale dual matched filter is first applied to detect the location of tubular structure. Piecewise parallel region detection method is then proposed to identify the potential regions having hookworm bodies. To discriminate the unique visual features for different components of gastrointestinal, the histogram of average intensity is proposed to represent their properties. In order to deal with the problem of imbalance data, Rusboost is deployed to classify WCE images. Experiments on a diverse and large scale dataset with 440 K WCE images demonstrate that the proposed approach achieves a promising performance and outperforms the state-of-the-art methods. Moreover, the high sensitivity in detecting hookworms indicates the potential of our approach for future clinical application. Xiao Wu 0001, Honghan Chen, Tao Gan, Junzhou Chen 0001, Chong-Wah Ngo, Qiang Peng |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Clothing Cosegmentation for Shopping Images With Cluttered BackgroundabstractIn this paper, we address an important and practical problem of clothing cosegmentation (CCS): given multiple fashion model photos with natural backgrounds on e-commerce websites, to automatically and simultaneously segment all images and extract the clothing regions. However, cluttered backgrounds, variations in colors and styles, and inconsistent human poses all make it a challenging task. In this paper, a novel CCS algorithm is proposed to improve the accuracy of clothing extraction by exploiting the properties of multiple clothing images with the same apparel. First, the co-salient objects are computed by detecting the upper bodies of fashion models and transferring their locations within multiple images. Based on the coarse clothing regions determined by the upper body localization and co-salient object detection, the foreground (clothing) and background Gaussian mixture models are estimated, respectively. Finally, the clothing region in each image is extracted through energy minimization based on graph cuts iteratively. The proposed cosegmentation algorithm is mainly designed for multiple clothing images. As a byproduct, it can also be applied to single image segmentation without any modification. The experiments demonstrate that the proposed approach outperforms the state-of-the-art cosegmentation methods as well as traditional single image segmentation solution for shopping images. Bo Zhao 0032, Xiao Wu 0001, Qiang Peng, Shuicheng Yan |
IEEE Trans. Multim. | 2 |
| 2014 | Clothing Extraction Using Region-Based Segmentation and Pixel-Level RefinementabstractIn this paper, we demonstrate an effective method for automatic extracting clothing object from fashion photographs, an extremely challenging problem due to the non-uniform natural backgrounds, various types of apparel and different poses of human models. This method consists of three phases: (1) coarse clothing area localization by pose estimation and super pixel segmentation, (2) region-level image segmentation, (3) pixel-level refinement using spatial information and Grab cut. Experiments on a dataset with 1000 images crawled from Taobao demonstrate that the proposed method outperforms other methods, which can extract clothing from images with complex background. Zhao-Rui Liu, Xiao Wu 0001, Bo Zhao 0032, Qiang Peng |
ISM | 2 |
| 2013 | An Error Resilient Depth Map Coding Scheme Using Adaptive Wyner-Ziv Frame
Xiangkai Liu, Qiang Peng, Xiao Wu 0001, Lei Zhang 0006, Ling-Yu Duan |
MMM (2) | 3 |
| 2013 | Clothing Extraction by Coarse Region Localization and Fine Foreground/Background Estimation
Xiao Wu 0001, Bo Zhao 0032, Ling-Ling Liang, Qiang Peng |
MMM (2) | 1 |
| 2013 | SSIM-Based End-to-End Distortion Model for Error Resilient Video Coding over Packet-Switched Networks
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001 |
MMM (1) | 3 |
| 2013 | A Novel Web Video Event Mining Framework with the Integration of Correlation and Co-Occurrence Information
Chengde Zhang, Xiao Wu 0001, Mei-Ling Shyu, Qiang Peng |
J. Comput. Sci. Technol. | 2 |
| 2013 | Guest editorial: selected papers from ICIMCS 2011
Chong-Wah Ngo, Changsheng Xu, Xiao Wu 0001, Abdulmotaleb El Saddik |
Multim. Syst. | 3 |
| 2012 | Boosting web video categorization with contextual information from social web
Xiao Wu 0001, Chong-Wah Ngo, Yi-Ming Zhu, Qiang Peng |
World Wide Web | 1 |
| 2010 | On the Annotation of Web Videos by Efficient Near-Duplicate SearchabstractWith the proliferation of Web 2.0 applications, user-supplied social tags are commonly available in social media as a means to bridge the semantic gap. On the other hand, the explosive expansion of social web makes an overwhelming number of web videos available, among which there exists a large number of near-duplicate videos. In this paper, we investigate techniques which allow effective annotation of web videos from a data-driven perspective. A novel classifier-free video annotation framework is proposed by first retrieving visual duplicates and then suggesting representative tags. The significance of this paper lies in the addressing of two timely issues for annotating query videos. First, we provide a novel solution for fast near-duplicate video retrieval. Second, based on the outcome of near-duplicate search, we explore the potential that the data-driven annotation could be successful when huge volume of tagged web videos is freely accessible online. Experiments on cross sources (annotating Google videos and Yahoo! videos using YouTube videos) and cross time periods (annotating YouTube videos using historical data) show the effectiveness and efficiency of the proposed classifier-free approach for web video tag annotation. Wanlei Zhao, Xiao Wu 0001, Chong-Wah Ngo |
IEEE Trans. Multim. | 2 |
| 2009 | Towards google challenge: combining contextual and social information for web video categorizationabstractWeb video categorization is a fundamental task for web video search. In this paper, we explore the Google challenge from a new perspective by combing contextual and social information under the scenario of social web. The semantic meaning of text (title and tags), video relevance from related videos, and user interest induced from user videos, are integrated to robustly determine the video category. Experiments on YouTube videos demonstrate the effectiveness of the proposed solution. The performance reaches 60% improvement compared to the traditional text based classifiers. Xiao Wu 0001, Wanlei Zhao, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2009 | Real-Time Near-Duplicate Elimination for Web Video Search With Content and ContextabstractWith the exponential growth of social media, there exist huge numbers of near-duplicate web videos, ranging from simple formatting to complex mixture of different editing effects. In addition to the abundant video content, the social web provides rich sets of context information associated with web videos, such as thumbnail image, time duration and so on. At the same time, the popularity of Web 2.0 demands for timely response to user queries. To balance the speed and accuracy aspects, in this paper, we combine the contextual information from time duration, number of views, and thumbnail images with the content analysis derived from color and local points to achieve real-time near-duplicate elimination. The results of 24 popular queries retrieved from YouTube show that the proposed approach integrating content and context can reach real-time novelty re-ranking of web videos with extremely high efficiency, where the majority of duplicates can be rapidly detected and removed from the top rankings. The speedup of the proposed approach can reach 164 times faster than the effective hierarchical method proposed in, with just a slight loss of performance. Xiao Wu 0001, Chong-Wah Ngo, Alex Hauptmann 0001, Hung-Khoon Tan |
IEEE Trans. Multim. | 1 |
| 2008 | Modeling video hyperlinks with hypergraph for web video rerankingabstractIn this paper, we investigate a novel approach of exploiting visual-duplicates for web video reranking using hypergraph. Current graph-based reranking approaches consider mainly the pair-wise linking of keyframes and ignore reliability issues that are inherent in such representation. We exploit higher order relation to overcome the issues of missing links in visual-duplicate keyframes and in addition identify the latent relationships among keyframes. Based on hypergraph, we consider two groups of video threads: visual near-duplicate threads and story threads, to hyperlink web videos and describe the higher order information existing in video content. To facilitate reranking using random walk algorithm, the hypergraph is converted to a star-like graph using star expansion algorithm. Experiments on a dataset of 12,790 web videos show that hypergraph reranking can improve web video retrieval up to 45% over the initial ranked result by the video sharing websites and 8.3% over the pair-wise based graph reranking in mean average precision (MAP). Hung-Khoon Tan, Chong-Wah Ngo, Xiao Wu 0001 |
ACM Multimedia | 3 |
| 2008 | Accelerating near-duplicate video matching by combining visual similarity and alignment distortionabstractIn this paper, we investigate a novel approach to accelerate the matching of two video clips by exploiting the temporal coherence property inherent in the keyframe sequence of a video. Motivated by the fact that keyframe correspondences between near-duplicate videos typically follow certain spatial arrangements, such property could be employed to guide the alignment of two keyframe sequences. We set the alignment problem as an integer quadratic programming problem, where the cost function takes into account both the visual similarity of the corresponding keyframes as well as the alignment distortion among the set of correspondences. The set of keyframe-pairs found by our algorithm provides our proposal on the list of candidate keyframe-pairs for near-duplicate detection using local interest points. This eliminates the need for exhaustive keyframe-pair comparisons, which significantly accelerates the matching speed. Experiments on a dataset of 12,790 web videos demonstrate that the proposed method maintains a similar near-duplicate video retrieval performance as the hierarchical method proposed in [12] but with a significantly reduced number of keyframe-pair comparisons. Hung-Khoon Tan, Xiao Wu 0001, Chong-Wah Ngo, Wanlei Zhao |
ACM Multimedia | 2 |
| 2008 | Measuring novelty and redundancy with multiple modalities in cross-lingual broadcast news
Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
Comput. Vis. Image Underst. | 1 |
| 2008 | Multimodal News Story Clustering With Pairwise Visual Near-Duplicate ConstraintabstractStory clustering is a critical step for news retrieval, topic mining, and summarization. Nonetheless, the task remains highly challenging owing to the fact that news topics exhibit clusters of varying densities, shapes, and sizes. Traditional algorithms are found to be ineffective in mining these types of clusters. This paper offers a new perspective by exploring the pairwise visual cues deriving from near-duplicate keyframes (NDK) for constraint-based clustering. We propose a constraint-driven co-clustering algorithm (CCC), which utilizes the near-duplicate constraints built on top of text, to mine topic-related stories and the outliers. With CCC, the duality between stories and their underlying multimodal features is exploited to transform features in low-dimensional space with normalized cut. The visual constraints are added directly to this new space, while the traditional DBSCAN is revisited to capitalize on the availability of constraints and the reduced dimensional space. We modify DBSCAN with two new characteristics for story clustering: 1) constraint-based centroid selection and 2) adaptive radius. Experiments on TRECVID-2004 corpus demonstrate that CCC with visual constraints is more capable of mining news topics of varying densities, shapes and sizes, compared with traditionalk-means, DBSCAN, and spectral co-clustering algorithms. Xiao Wu 0001, Chong-Wah Ngo, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 1 |
| 2007 | Efficient Near-Duplicate Keyframe Retrieval with Visual Language ModelsabstractNear-duplicate keyframe retrieval is a critical task for video similarity measure, video threading and tracking. In this paper, instead of using expensive point-to-point matching on keypoints, we investigate the visual language models built on visual keywords to speed up the near-duplicate keyframe retrieval. The main idea is to estimate a visual language model on visual keywords for each keyframe and compare keyframes by the likelihood of their visual language models. Experiments on a subset of TRECVID-2004 video corpus show that visual language models built on visual keywords demonstrate promising performance for near-duplicate keyframe retrieval, which greatly speed up the retrieval speed although sacrifice a little performance compared to expensive point-to-point matching. Xiao Wu 0001, Wanlei Zhao, Chong-Wah Ngo |
ICME | 1 |
| 2007 | Novelty detection for cross-lingual news stories with visual duplicates and speech transcriptsabstractAn overwhelming volume of news videos from different channels and languages is available today, which demands automatic management of this abundant information. To effectively search, retrieve, browse and track cross-lingual news stories, a news story similarity measure plays a critical role in assessing the novelty and redundancy among them. In this paper, we explore the novelty and redundancy detection with visual duplicates and speech transcripts for cross-lingual news stories. News stories are represented by a sequence of keyframes in the visual track and a set of words extracted from speech transcript in the audio track. A major difference to pure text documents is that the number of keyframes in one story is relatively small compared to the number of words and there exist a large number of non-near duplicate keyframes. These features make the behavior of similarity measures different compared to traditional textual collections. Furthermore, the textual features and visual features complement each other for news stories. They can be further combined to boost the performance. Experiments on the TRECVID-2005 cross-lingual news video corpus show that approaches on textual features and visual features demonstrate different performance, and measures on visual features are quite effective. Overall, the cosine distance on keyframes is still a robust measure. Language models built on visual features demonstrate promising performance. The fusion of textual and visual features improves overall performance. Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2007 | Practical elimination of near-duplicates from web video searchabstractCurrent web video search results rely exclusively on text keywords or user-supplied tags. A search on typical popular video often returns many duplicate and near-duplicate videos in the top results. This paper outlines ways to cluster and filter out the near-duplicate video using a hierarchical approach. Initial triage is performed using fast signatures derived from color histograms. Only when a video cannot be clearly classified as novel or near-duplicate using global signatures, we apply a more expensive local feature based near-duplicate detection which provides very accurate duplicate analysis through more costly computation. The results of 24 queries in a data set of 12,790 videos retrieved from Google, Yahoo! and YouTube show that this hierarchical approach can dramatically reduce redundant video displayed to the user in the top result set, at relatively small computational cost. Xiao Wu 0001, Alex Hauptmann 0001, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2007 | Near-Duplicate Keyframe Identification With Interest Point Matching and Pattern LearningabstractThis paper proposes a new approach for near-duplicate keyframe (NDK) identification by matching, filtering and learning of local interest points (LIPs) with PCA-SIFT descriptors. The issues in matching reliability, filtering efficiency and learning flexibility are novelly exploited to delve into the potential of LIP-based retrieval and detection. In matching, we propose a one-to-one symmetric matching (OOS) algorithm which is found to be highly reliable for NDK identification, due to its capability in excluding false LIP matches compared with other matching strategies. For rapid filtering, we address two issues: speed efficiency and search effectiveness, to support OOS with a new index structure called LIP-IS. By exploring the properties of PCA-SIFT, the filtering capability and speed of LIP-IS are asymptotically estimated and compared to locality sensitive hashing (LSH). Owing to the robustness consideration, the matching of LIPs across keyframes forms vivid patterns that are utilized for discriminative learning and detection with support vector machines. Experimental results on TRECVID-2003 corpus show that our proposed approach outperforms other popular methods including the techniques with LSH in terms of retrieval and detection effectiveness. In addition, the proposed LIP-IS successfully speeds up OOS for more than ten times and possesses several avorable properties compared to LSH. Wanlei Zhao, Chong-Wah Ngo, Hung-Khoon Tan, Xiao Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2006 | Scheduling real-time requests in on-demand data broadcast environments
Victor C. S. Lee, Xiao Wu 0001, Joseph Kee-Yin Ng |
Real Time Syst. | 2 |
| 2005 | Co-Clustering of Time-Evolving News Story with Transcript and KeyframeabstractThis paper presents techniques in clustering the same topic news stories according to event themes. We model the relationship of stories with textual and visual concepts under the representation of bipartite graph. The textual and visual concepts are extracted respectively from speech transcripts and keyframes. Co-clustering algorithm is employed to exploit the duality of stories and textual-visual concepts based on spectral graph partitioning. Experimental results on TRECVID-2004 corpus show that the co-clustering of news stories with textual-visual concepts is significantly better than the co-clustering with either textual or visual concept alone. Xiao Wu 0001, Chong-Wah Ngo, Qing Li 0001 |
ICME | 1 |
| 2005 | Threading stories and generating topic structures in news videos across different sourcesabstractNews videos delivered from different sources constitute a huge volume of daily information. These videos, overall, form a huge collection of news stories that are intertwined with various novel and old topic themes. To date, it remains a challenging task on how to automatically extract a concise view of news stories according to topic themes. This doctoral thesis studies the issues in story dependency threading and topical auto-documentary in news stories. Initially, a co-clustering algorithm is proposed to perform the news story clustering by exploiting the duality between stories and multi-modal concepts. Then, the novelty and redundancy detection is performed to capture the relationship among stories of a topic. To facilitate the fast navigation of news topic, a novel topic structure is then proposed to chains the dependencies of stories. A main thread is extracted to highlight the important aspects of a theme. A news video editing optimization algorithm can be directly applied to automatically select suitable video and speech contents from the original video source to create an edited video documentary. Xiao Wu 0001 |
ACM Multimedia | 1 |
| 2005 | A Preemptive Scheduling Algorithm for Wireless Real-Time On-Demand Data BroadcastabstractOn-demand broadcast is an attractive data dissemination method for mobile and wireless computing. In this paper, we propose a new online preemptive scheduling algorithm, called PRDS that incorporates the urgency, the data size and the number of pending requests for real-time on-demand broadcast system. Furthermore, we use pyramid preemption to optimize performance and reduce overhead. We have done a series of simulation experiments to evaluate the performance of our algorithm as compared with other previously proposed methods under a range of scenarios. The experimental results show that our algorithm can substantially outperform other algorithms without jeopardizing other performance metrics, such as response time and stretch. Xiao Wu 0001, Victor C. S. Lee, Joseph Kee-Yin Ng |
RTCSA | 1 |
| 2005 | Wireless real-time on-demand data broadcast scheduling with dual deadlines
Xiao Wu 0001, Victor C. S. Lee |
J. Parallel Distributed Comput. | 1 |
| 2004 | Preemptive Maximum Stretch Optimization Scheduling for Wireless On-Demand Data Broadcast
Xiao Wu 0001, Victor C. S. Lee |
IDEAS | 1 |