VLDB 2026 Research / reviewers in the wild / expert
Hao Wang 0073
dblp:181/2812-73
· DBLP profile ↗
41ranked-venue papers
14as first author
39since 2021 · last 2026
0000-0002-6956-7342ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 22 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryabstractData scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availability of precise and reasonable control instructions. Current methods primarily rely on 2D trajectories as instruction prompts, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3Dfor generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length, avoiding collisions and retiming. Next, we employ a latent editing technique to create video sequences from the initial image latent, text instruction and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method significantly reduces human intervention requirements by autonomously planing plausible 3D trajectories. Experimental results demonstrate its superior visual quality and precision. Ying Li 0128, Xiaobao Wei, Xiaowei Chi, Zhongyu Zhao, Hao Wang 0073, Ningning Ma, Ming Lu 0002, Sirui Han |
AAAI | 6 |
| 2026 | RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow MatchingabstractRGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with detail inconsistency and color deviation, due to the ill-posed nature of inverse ISP and the inherent information loss in quantized RGB images. To address these limitations, we pioneer a generative perspective by reformulating RGB-to-RAW reconstruction as a deterministic latent transport problem and introduce a novel framework named RAW-Flow, which leverages flow matching to learn a deterministic vector field in latent space, to effectively bridge the gap between RGB and RAW representations and enable accurate reconstruction of structural details and color information. To further enhance latent transport, we introduce a cross-scale context guidance module that injects hierarchical RGB features into the flow estimation process. Moreover, we design a Dual-domain Latent Autoencoder (DLAE) with a feature alignment constraint to support the proposed latent transport framework, which jointly encodes RGB and RAW inputs while promoting stable training and high-fidelity reconstruction. Extensive experiments demonstrate that RAW-Flow outperforms state-of-the-art approaches both quantitatively and visually. Zhen Liu 0022, Diedong Feng, Hai Jiang 0006, Liaoyuan Zeng, Hao Wang 0073, Chaoyu Feng, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 5 |
| 2026 | Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction MetricabstractWith the rapid advancement of 3D visualization, 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time, high-fidelity rendering. While prior research has emphasized algorithmic performance and visual fidelity, the perceptual quality of 3DGS-rendered content, especially under varying reconstruction conditions, remains largely underexplored. In practice, factors such as viewpoint sparsity, limited training iterations, point downsampling, noise, and color distortions can significantly degrade visual quality, yet their perceptual impact has not been systematically studied. To bridge this gap, we present 3DGS-QA, the first subjective quality assessment dataset for 3DGS. It comprises 225 degraded reconstructions across 15 object types, enabling a controlled investigation of common distortion factors. Based on this dataset, we introduce a no-reference quality prediction model that directly operates on native 3D Gaussian primitives, without requiring rendered images or ground-truth references. Our model extracts spatial and photometric cues from the Gaussian representation to estimate perceived quality in a structure-aware manner. We further benchmark existing quality assessment methods, spanning both traditional and learning-based approaches. Experimental results show that our method consistently achieves superior performance, highlighting its robustness and effectiveness for 3DGS content evaluation. The dataset and code are made publicly available to facilitate future research in 3DGS quality assessment. Zhaolin Wan, Yining Diao, Jingqi Xu, Hao Wang 0073, Zhiyang Li 0001, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao |
AAAI | 4 |
| 2026 | SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene SimulationabstractWhile 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/. Qingpo Wuwu, Xiaobao Wei, Peng Chen 0046, Zhongyu Zhao, Hao Wang 0073, Ming Lu 0002, Ningning Ma, Shanghang Zhang |
ICMR | 6 |
| 2026 | Conditional diffusion models for X-ray security image synthesis: Addressing data scarcity in security inspection tasks
Da Cai, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003, Dongyue Chen 0001 |
Neurocomputing | 4 |
| 2026 | DCART: A dual contrastive alignment residual transformer model for visual grounding
Dongyue Chen 0001, Hao Wang 0073, Tong Jia 0001, Shizhuo Deng |
Pattern Recognit. | 3 |
| 2026 | GAA-TSO: Geometry-Aware-Assisted Depth Completion for Transparent and Specular ObjectsabstractTransparent and specular objects are frequently encountered in daily life, factories, and laboratories. However, due to the unique optical properties, the depth information on these objects is usually incomplete and inaccurate, which poses significant challenges for downstream robotics tasks. Therefore, it is crucial to accurately restore the depth information of transparent and specular objects. Previous depth completion methods for these objects usually generate structure-less or ambiguous depth predictions. To address these issues, we propose a geometry-aware assisted depth completion method for transparent and specular objects, which focuses on exploring the 3D structural cues of the scene. Specifically, besides extracting 2D features from RGB-D input, we back-project the input depth to a point cloud and build the 3D branch to extract hierarchical scene-level 3D structural features. To exploit 3D geometric information, we design several gated cross-modal fusion modules to effectively propagate multi-level 3D geometric features to the image branch. In addition, we propose an adaptive correlation aggregation strategy to appropriately assign 3D features to the corresponding 2D features. Extensive experiments on ClearGrasp, OOD, TransCG, and STD datasets show that our method outperforms other state-of-the-art methods. We further demonstrate that our method significantly enhances the performance of downstream robotic grasping tasks. The code will be available at: https://github.com/lyz3356/GAA-TSO. Yizhe Liu, Tong Jia 0001, Jiahui Wei, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Contextual Style Coherence Network for X-Ray Prohibited Item Image SynthesisabstractProhibited item detection in X-Ray baggage images plays a crucial role for preventing the social security and stability. Well annotated X-Ray prohibited item training samples show necessity in achieving high detection performance for X-Ray inspection system. While collection of massive samples is extremely laborious and costly, especially for those X-Ray images, which need professional inspection machine. Synthesizing X-Ray images through Threat Image Projection (TIP) is a promising solution to overcome the data insufficient limitation in prohibited item detection. However, TIP based methods rarely consider the contextual style coherence between the foreground prohibited items and background images, resulting in generating low realistic X-Ray security images. For improving image quality and diversity, we propose a Contextual Style Coherence Network for X-Ray Prohibited item Image Synthesis. Specifically, we first propose a style fusion module to guarantee the style coherence and consistency between the foreground prohibited items and background images. We transfer the threat image projection from image space to feature space, and an affine transformation matrix is applied to uniformly sample the location, ratio and scale of the prohibited items to improve the sample diversity. We further normalize the features of the foreground prohibited item by implementing the style transfer through Gram matrix. Then, a mask partial convolution is designed for inpainting the non-object regions of the foreground prohibited items to achieve a better style transition, especially for the boundary parts. The whole network follows the adversarial training pipeline in an unsupervised manner guided by the incorporation of adversarial loss and total variation regularization. We evaluate the synthetic images generated by our method from different evaluating metrics including image quality and object detection performance on various prohibited item detection datasets. The results verify that our method can effectively generate realistic X-Ray prohibited item images and improve the detection performance. Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng |
IEEE Trans. Image Process. | 1 |
| 2026 | Global and Local Visual-Textual Alignment for Open Vocabulary Object DetectionabstractRecently, with the development of the Vision-Language Model (VLM), adopting such VLM (e.g., CLIP) into object detection framework has gradually become a promising and attractive research direction, and the resulted open vocabulary object detection methods can effectively alleviate the limitations in those close-set ones, making the detectors perceive the unseen world. The core issue in open vocabulary object detection is to design an effective and efficient alignment between the visual (e.g., image) and textual (e.g., caption) features in the semantic space, so that the detectors can capture more information around the open-set scene. Current approaches deploy extra uncurated image-text pairs to pre-train a detector for obtaining a better visual-textual alignment in the feature space. Besides, knowledge distillation technology is also adopted to design an appropriate information transferring flow for aligning the visual-textual knowledge. However, large-scale image-text pairs are not always available to obtain, and the pre-training process will inevitable introduce much more computation overhead. While knowledge distillation methods focus on aligning between the local region visual feature in RoI and the textual features of VLM, neglecting the global information alignment between the image and text. For addressing the dilemmas in these alignment manners, we propose a Global and Local Visual-Textual Alignment for Open Vocabulary Object Detection in this paper. Specifically, our proposed method integrates global image-caption and local region-prompt alignments into a unified learning paradigm. The global alignment takes the whole image and caption as the visual and textual inputs, respectively, and matches the image and caption representations from the detector and the text encoder in CLIP by contrastive learning from the overall perspective. Different from global alignment, the local one concentrates on the accordance between regions and prompts from the aspect of portion description. It extracts and aligns the embeddings for the visual patch RoIs from the image encoder in CLIP and discriminating textual token prompts from the text encoder. Moreover, we also design a prompt tuning strategy, which contains global and local components corresponding to the alignment procedure, for better adapting CLIP to downstream task object detection in a parameter-efficient learning manner. By implementation on Faster R-CNN, we conduct experiments on open vocabulary benchmarks OV-COCO and OV-LVIS, respectively. The results verify that our proposed method can achieve clear improvement over counterparts on novel categories, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Shizhuo Deng, Dongyue Chen 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2025 | DUPL: Domain-agnostic Unknown-aware Prompt Learning for Threshold-free Open-set Domain GeneralizationabstractOpen-set domain generalization (OSDG) aims to recognize known categories in unseen target domains without fine-tuning, while rejecting unknown categories. Existing methods show limited practicality for the requirement of extra generated or collected unknown samples, or for the assumption that the known categories appear in all source domains. Additionally, they need to determine an optimal threshold for distinguishing between known and unknown samples during testing, which is impractical in OSDG, as the target domain is unavailable during training. Besides, they are usually studied with conventional CNNs and shows unsatisfactory generalizability. To address these issues, we harness the transferable property of the pre-trained vision-language model CLIP, and propose domain-agnostic unknown-aware prompt learning (DUPL) framework to achieve threshold-free OSDG. Specifically, we train unknown tokens (UT) to enable threshold-free unknown rejection through unknown-aware prompt learning (UAPL) with only known data, and then introduce a Fourier-based data augmentation (FDA) strategy to obtain domain-agnostic prompts via domain-agnostic semantic consistency (DASC) regularization. Extensive experiments show that our method achieves state-of-the-art OSDG performance. Code is available at https://github.com/X-funbean/DUPL. Fangbin Xu, Dongyue Chen 0001, Shizhuo Deng, Tong Jia 0001, Hao Wang 0073 |
ICME | 5 |
| 2025 | OmniArch: Building Foundation Model for Scientific ComputingabstractFoundation models have revolutionized language modeling, while whether this success is replicated in scientific computing remains unexplored. We present OmniArch, the first prototype aiming at solving multi-scale and multi-physics scientific computing problems with physical alignment. We addressed all three challenges with one unified architecture. Its pre-training stage contains a Fourier Encoder-decoder fading out the disharmony across separated dimensions and a Transformer backbone integrating quantities through temporal dynamics, and the novel PDE-Aligner performs physics-informed fine-tuning under flexible conditions. As far as we know, we first conduct 1D-2D-3D united pre-training on the PDEBench, and it sets not only new performance benchmarks for 1D, 2D, and 3D PDEs but also demonstrates exceptional adaptability to new physics via in-context and zero-shot learning approaches, which supports realistic engineering applications and foresight physics discovery. Tianyu Chen 0017, Haoyi Zhou, Ying Li 0128, Hao Wang 0073, Chonghan Gao, Rongye Shi, Shanghang Zhang, Jianxin Li 0002 |
ICML | 4 |
| 2025 | SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representationabstract3D semantic occupancy prediction is a crucial task in visual perception, as it requires the simultaneous comprehension of both scene geometry and semantics. It plays a crucial role in understanding 3D scenes and has great potential for various applications, such as robotic vision perception and autonomous driving. Many existing works utilize planar-based representations such as Bird's Eye View (BEV) and Tri-Perspective View (TPV). These representations aim to simplify the complexity of 3D scenes while preserving essential object information, thereby facilitating efficient scene representation. However, in dense indoor environments with prevalent occlusions, directly applying these planar-based methods often leads to difficulties in capturing global semantic occupancy, ultimately degrading model performance. In this paper, we present a new vertical slice representation that divides the scene along the vertical axis and projects spatial point features onto the nearest pair of parallel planes. To utilize these slice features, we propose SliceOcc, an RGB camera-based model specifically tailored for indoor 3D semantic occupancy prediction. SliceOcc utilizes pairs of slice queries and cross-attention mechanisms to extract planar features from input images. These local planar features are then fused to form a global scene representation, which is employed for indoor occupancy prediction. Experimental results on the EmbodiedScan dataset demonstrate that SliceOcc achieves a mIoU of 15.45 % across 81 indoor categories, setting a new state-of-the-art performance among RGB camera-based models for indoor 3D semantic occupancy prediction. Jianing Li 0001, Ming Lu 0002, Hao Wang 0073, Chenyang Gu, Wenzhao Zheng, Shanghang Zhang |
ICRA | 4 |
| 2025 | FreqMoE: Dynamic Frequency Enhancement for Neural PDE SolversabstractFourier Neural Operators (FNO) have emerged as promising solutions for efficiently solving partial differential equations (PDEs) by learning infinite-dimensional function mappings through frequency domain transformations. However, the sparsity of high-frequency signals limits computational efficiency for high-dimensional inputs, and fixed-pattern truncation often causes high-frequency signal loss, reducing performance in scenarios such as high-resolution inputs or long-term predictions. To address these challenges, we propose FreqMoE, an efficient and progressive training framework that exploits the dependency of high-frequency signals on low-frequency components. The model first learns low-frequency weights and then applies a sparse upward-cycling strategy to construct a mixture of experts (MoE) in the frequency domain, effectively extending the learned weights to high-frequency regions. Experiments on both regular and irregular grid PDEs demonstrate that FreqMoE achieves up to 16.6 percent accuracy improvement while using merely 2.1 percent parameters (47.32x reduction) compared to dense FNO. Furthermore, the approach demonstrates remarkable stability in long-term predictions and generalizes seamlessly to various FNO variants and grid structures, establishing a new Low frequency Pretraining, High frequency Fine-tuning'' paradigm for solving PDEs. Tianyu Chen 0017, Haoyi Zhou, Ying Li 0128, Hao Wang 0073, Zhenzhe Zhang, Tianchen Zhu, Shanghang Zhang, Jianxin Li 0002 |
IJCAI | 4 |
| 2025 | RePST: Language Model Empowered Spatio-Temporal Forecasting via Semantic-Oriented ReprogrammingabstractSpatio-temporal forecasting is pivotal in numerous real-world applications, including transportation planning, energy management, and climate monitoring. In this work, we aim to harness the reasoning and generalization abilities of Pre-trained Language Models (PLMs) for more effective spatio-temporal forecasting, particularly in data-scarce scenarios. However, recent studies uncover that PLMs, which are primarily trained on textual data, often falter when tasked with modeling the intricate correlations in numerical time series, thereby limiting their effectiveness in comprehending spatio-temporal data. To bridge the gap, we propose RePST, a semantic-oriented PLM reprogramming framework tailored for spatio-temporal forecasting. Specifically, we first propose a semantic-oriented decomposer that adaptively disentangles spatially correlated time series into interpretable sub-components, which facilitates PLM to understand sophisticated spatio-temporal dynamics via a divide-and-conquer strategy. Moreover, we propose a selective discrete reprogramming scheme, which introduces an expanded spatio-temporal vocabulary space to project spatio-temporal series into discrete representations. This scheme minimizes the information loss during reprogramming and enriches the representations derived by PLMs. Extensive experiments on real-world datasets show that the proposed RePST outperforms twelve state-of-the-art baseline methods, particularly in data-scarce scenarios, highlighting the effectiveness and superior generalization capabilities of PLMs for spatio-temporal forecasting. Codes and Appendix can be found at https://github.com/usail-hkust/REPST. Hao Wang 0073, Jindong Han, Wei Fan 0010, Leilei Sun, Hao Liu 0026 |
IJCAI | 1 |
| 2025 | EmbodiedOcc++: Boosting Embodied 3D Occupancy Prediction with Plane Regularization and Uncertainty SamplerabstractOnline 3D occupancy prediction provides a comprehensive spatial understanding of embodied environments. While the innovative EmbodiedOcc framework utilizes 3D semantic Gaussians for progressive indoor occupancy prediction, it overlooks the geometric characteristics of indoor environments, which are primarily characterized by planar structures. This paper introduces EmbodiedOcc++, enhancing the original framework with two key innovations: a Geometry-guided Refinement Module (GRM) that constrains Gaussian updates through plane regularization, along with a Semantic-aware Uncertainty Sampler (SUS) that enables more effective updates in overlapping regions between consecutive frames. GRM regularizes the position update to align with surface normals. It determines the adaptive regularization weight using curvature-based and depth-based constraints, allowing semantic Gaussians to align accurately with planar surfaces while adapting in complex regions. To effectively improve geometric consistency from different views, SUS adaptively selects proper Gaussians to update. Comprehensive experiments on the EmbodiedOcc-ScanNet benchmark demonstrate that EmbodiedOcc++ achieves state-of-the-art performance across different settings. Our method demonstrates improved edge accuracy and retains more geometric details while ensuring computational efficiency, which is essential for online embodied perception. The code will be released at: https://github.com/PKUHaoWang/EmbodiedOcc2. Hao Wang 0073, Xiaobao Wei, Xiaoan Zhang, Jianing Li 0001, Chengyu Bai, Ying Li 0128, Ming Lu 0002, Wenzhao Zheng, Shanghang Zhang |
ACM Multimedia | 1 |
| 2025 | CSPCL: Category Semantic Prior Contrastive Learning for Deformable DETR-Based Prohibited Item DetectorsabstractProhibited item detection based on X-ray images is one of the most effective security inspection methods. However, the foreground-background feature coupling caused by the overlapping phenomenon specific to X-ray images makes general detectors designed for natural images perform poorly. To address this issue, we propose a Category Semantic Prior Contrastive Learning (CSPCL) mechanism, which aligns the class prototypes perceived by the classifier with the content queries to correct and supplement the missing semantic information responsible for classification, thereby enhancing the model sensitivity to foreground features. To achieve this alignment, we design a specific contrastive loss, CSP loss, which comprises the Intra-Class Truncated Attraction (ITA) loss and the Inter-Class Adaptive Repulsion (IAR) loss, and outperforms classic contrastive losses. Specifically, the ITA loss leverages class prototypes to attract intra-class content queries and preserves essential intra-class diversity via a gradient truncation function. The IAR loss employs class prototypes to adaptively repel inter-class content queries, with the repulsion strength scaled by prototype-prototype similarity, thereby improving inter-class discriminability, especially among similar categories. CSPCL is general and can be easily integrated into Deformable DETR-based models. Extensive experiments on the PIXray, OPIXray, PIDray, and CLCXray datasets demonstrate that CSPCL significantly enhances the performance of various state-of-the-art models without increasing inference complexity. The code is publicly available at https://github.com/Limingyuan001/CSPCL. Mingyuan Li 0003, Tong Jia 0001, Hao Wang 0073, Shiyi Guo, Da Cai, Dongyue Chen 0001 |
NeurIPS | 3 |
| 2025 | Detection of novel prohibited item categories for real-world security inspection
Shuyang Lin, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003, Dongyue Chen 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Scalable Pre-Training of Compact Urban Spatio-Temporal Predictive Models on Large-Scale Multi-Domain DataabstractSpatio-Temporal Prediction (STP) is crucial for various smart city applications, such as traffic management and resource allocation. However, training samples can be scarce in data-constrained scenarios, which often degrades the predictive capability of existing deep STP models. Although recent STP foundation models excel in few-shot and zero-shot learning through extensive pre-training on large-scale, multi-domain spatio-temporal data, they often rely on large parameter scale to achieve enhanced performance, resulting in high computational demands that hinder practical deployment. In response, we develop CompactST, an efficient, compact, and versatile pre-trained model for STP in data-scarce settings. Recognizing the complexities posed by large-scale, heterogeneous pre-training datasets, CompactST integrates three specialized components: (1) a mixture-of-normalizers module to address domain and spatial heterogeneity, (2) a multi-scale spatio-temporal mixer that captures diverse patterns from datasets with varying spatio-temporal resolutions, and (3) an adaptive dataset-oriented tuning module that transfers the handling of dataset-specific parameters from pre-training to fine-tuning stage. These tailored designs enable CompactST to maximize generalizability across diverse datasets while maintaining a compact model size ( i.e. , only 300K parameters). To validate its effectiveness, we pre-train CompactST on a substantial corpus of public spatio-temporal datasets spanning over 10 domains and encompassing 300 million data points. Extensive experimental results on ten real-world datasets demonstrate CompactST's significantly improved prediction accuracy and efficiency in data-scarce scenarios. Jindong Han, Hao Wang 0073, Hui Xiong 0001, Hao Liu 0026 |
Proc. VLDB Endow. | 2 |
| 2025 | RASP: Robot Active Scene Perception With Joint Viewpoint Planning and Depth Completion in Cluttered EnvironmentsabstractCapturing dense visual perception in cluttered environments using depth sensors is crucial for downstream robotics tasks. However, occlusions between objects and unreliable depth data make it very challenging for robots to perceive comprehensive and accurate scene information. To address these issues, we propose a novel robot active scene perception method called RASP, which is composed of two parts. First, we introduce a temporal attention-based view planning algorithm, which actively plans the minimum feasible viewpoint sequence based on latent dependencies in all previous observations to maximize the information perception of the cluttered environments. Subsequently, for the unreliable depth data obtained by the depth sensor, especially caused by transparent and specular objects in the scene, we design a geometry-guided depth completion network that fully utilizes the 3D scene information during the progressive perception process. Specifically, multi-level scene geometric features are extracted and projected into the image space, combining with the image features to guide the depth completion step. These two parts are learned jointly to achieve consistent and accurate results. Extensive experiments demonstrate that our method outperforms state-of-the-art methods. Furthermore, we show that our method significantly improves the performance of downstream grasping tasks. Yizhe Liu, Tong Jia 0001, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Automatic Label Assignment for Object DetectionabstractLabel assignment, which aims to classify region proposals as positive or negative samples depending on the correlations between their classification and localization predictions with the corresponding ground truth, is recognized as an essential ingredient in object detection and strongly affects the detection performance. Recently, some dynamic label assignment methods have been proposed to overcome the limitations of the static methods and achieve promising performance improvement. Despite eliminating the restrictions of the human prior sampling knowledge in static methods, existing dynamic principles usually suffer from two weaknesses. First, most of them deploy mixture models or implicit branch in prediction head to coarsely estimate the spatial distribution of the positive samples for objects. They give little attention to the effect of appearance information of the objects. Furthermore, these methods still cannot perceive the quality distribution of the positive samples, and these low-quality samples lead to adverse effects on the detection performance. To address issues, this paper presents a novel automatic label assignment for object detection. Specifically, our method first introduces an instance property branch into object detection pipeline to distinguish the foreground from the background. Then, an objectness prediction module which is composed by the confidence and weight mechanisms is developed to generate the positive and negative weight maps for the objects. The instance property branch and objectness prediction module can provide a coarse-to-fine optimization framework to make our method realize the appearance of the objects. Finally, a positive sample selection strategy is proposed to explore the quality statistical distribution of the positive samples, which are trained by different designed label targets. We evaluate our method on the MS COCO dataset and we achieve 48.4%, 47.9%, 48.0% and 49.3% on ResNet-101, ResNeXt-101, DCN-ResNet-101 and DCN-ResNeXt-101 in terms of AP0.5:0.95, respectively. We evaluate the timing complexity of ALA by calculating the inference speed and the frame per second (FPS) for these four backbones are 11.9, 10.4, 9.9 and 8.0, respectively. The experiment results demonstrate that we can obtain clear improvement over the competing methods with favorable performance compared to the state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Open-Vocabulary Prohibited Item Detection for Real-World X-Ray Security InspectionabstractComputer-aided prohibited item detection is applied in X-ray security inspection to maintain public safety. However, existing prohibited item detectors are limited to a small set of categories in current X-ray datasets, posing potential risks to public security. Since constructing bigger datasets and annotating hundreds of categories is time-consuming and labor-intensive, scaling detectors to more categories with minimal supervision is of great importance. To this end, in this paper, we adopt an open-vocabulary object detection (OVOD) method to detect arbitrary unlabeled novel categories of prohibited item. OVOD methods typically rely on datasets with caption annotations, which are lacking in the domain of prohibited item detection. To support the research on OVOD in X-ray security inspection scenarios, we contribute PIXray Caption dataset, the first X-ray dataset with image-caption pair annotations, which could benchmark and facilitate researches in the community. Further, we propose a novel Open-Vocabulary Prohibited Item Detection (OVPID) network to leverage textual information from captions. OVPID contains two core modules, i.e., Interference Resistant Module (IRM) and Prediction Module (PM). Specifically, IRM includes two submodules, namely Edge Perception (EP) and Foreground Activation (FA), which are designed to address the dilemma of interference caused by overlapping problem and complex background in X-ray images. PM consists of two branches for classification and localization. In classification branch, PM generates more accurate prompts for X-ray dataset via large multimodal model (LMM). In localization branch, PM aligns the student embeddings with both teacher and caption embeddings. Extensive experiments on PIXray Caption dataset demonstrate that OVPID outperforms other OVOD methods by delivering a higher accuracy on novel categories. Shuyang Lin, Tong Jia 0001, Hao Wang 0073, Mingyuan Li 0003 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Meta-TIP: An Unsupervised End-to-End Fusion Network for Multi-Dataset Style-Adaptive Threat Image ProjectionabstractThreat Image Projection (TIP) is a convenient and effective means to expand X-ray baggage images, which is essential for training both security personnel and computer-aided screening systems. Existing methods are primarily divided into two categories: X-ray imaging principle-based methods and GAN-based generative methods. The former cast prohibited items acquisition and projection as two individual steps and rarely consider the style consistency between the source prohibited items and target X-ray images from different datasets, making them less flexible and reliable for practical applications. Although GAN-based methods can directly generate visually consistent prohibited items on target images, they suffer from unstable training and lack of interpretability, which significantly impact the quality of the generated items. To overcome these limitations, we present a conceptually simple, flexible and unsupervised end-to-end TIP framework, termed as Meta-TIP, which superimposes the prohibited item distilled from the source image onto the target image in a style-adaptive manner. Specifically, Meta-TIP mainly applies three innovations: 1) reconstruct a pure prohibited item from a cluttered source image with a novel foreground-background contrastive loss; 2) a material-aware style-adaptive projection module learns two modulation parameters pertinently based on the style of similar material objects in the target image to control the appearance of prohibited items; 3) a novel logarithmic form loss is well-designed based on the principle of TIP to optimize synthetic results in an unsupervised manner. We comprehensively verify the authenticity and training effect of the synthetic X-ray images on four public datasets, i.e., SIXray, OPIXray, PIXray, and PIDray dataset, and the results confirm that our framework can flexibly generate very realistic synthetic images without any limitations. Tong Jia 0001, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | WS-SAM: Generalizing SAM to Weakly Supervised Object Detection With Category LabelabstractBuilding an effective object detector usually depends on large well-annotated training samples. While annotating such dataset is extremely laborious and costly, where box-level supervision which contains both accurate classification category and localization coordinate is required. Compared to above box-level supervised annotation, those weakly supervised learning manners (e.g,, category, point and scribble) need relatively less laborious annotation cost, and provide a feasible way to mitigate the reliance on the dataset. Because of the lack of sufficient supervised information, current weakly supervised methods cannot achieve satisfactory detection performance. Recently, Segment Anything Model (SAM) has appeared as a task-agnostic foundation model and shown promising performance improvement in many related works due to its powerful generalization and data processing abilities. The properties of the SAM inspire us to adopt such basic benchmark to weakly supervised object detection field to compensate the deficiencies in supervised information. However, directly deploying SAM on weakly supervised object detection task meets with two issues. Firstly, SAM needs meticulously-designed prompts, and such expert-level prompts restrict their applicability and practicality. Besides, SAM is a category unawareness model, and it cannot assign the category labels to the generated predictions. To solve above issues, we propose WS-SAM, which generalizes Segment Anything Model (SAM) to weakly supervised object detection with category label. Specifically, we design an adaptive prompt generator to take full advantages of the spatial and semantic information from the prompt. It employs in a self-prompting manner by taking the output of SAM from the previous iteration as the prompt input to guide the next iteration, where the prompts can be adaptively generated based on the classification activation map. We also develop a segmentation mask refinement module and formulate the label assignment process as a shortest path optimization problem by considering the similarity between each location and prompts. Furthermore, a bidirectional adapter is also implemented to resolve the domain discrepancy by incorporating domain-specific information. We evaluate the effectiveness of our method on several detection datasets (e.g., PASCAL VOC and MS COCO), and the experiment results show that our proposed method can achieve clear improvement over state-of-the-art methods, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2025 | AO-DETR: Anti-Overlapping DETR for X-Ray Prohibited Items DetectionabstractProhibited item detection in X-ray images is one of the most essential and highly effective methods widely employed in various security inspection scenarios. Considering the significant overlapping phenomenon in X-ray prohibited item images, we propose an anti-overlapping detection transformer (AO-DETR) based on one of the state-of-the-art (SOTA) general object detectors, DETR with improved denoising anchor boxes (DINO). Specifically, to address the feature coupling issue caused by overlapping phenomena, we introduce the category-specific one-to-one assignment (CSA) strategy to constrain category-specific object queries in predicting prohibited items of fixed categories, which can enhance their ability to extract features specific to prohibited items of a particular category from the overlapping foreground-background features. To address the edge blurring problem caused by overlapping phenomena, we propose the look forward densely (LFD) scheme, which improves the localization accuracy of reference boxes in mid-to-high-level decoder layers and enhances the ability to locate blurry edges of the final layer. Similar to DINO, our AO-DETR provides two different versions with distinct backbones, tailored to meet diverse application requirements. Extensive experiments on the PIXray, OPIXray, and HIXray datasets demonstrate that the proposed method surpasses the SOTA object detectors, indicating its potential applications in the field of prohibited item detection. The source code will be available at: https://github.com/Limingyuan001/AO-DETR. Mingyuan Li 0003, Tong Jia 0001, Hao Wang 0073, Shuyang Lin, Da Cai, Dongyue Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Self-Paced Unified Representation Learning for Hierarchical Multi-Label ClassificationabstractHierarchical Multi-Label Classification (HMLC) is a well-established problem that aims at assigning data instances to multiple classes stored in a hierarchical structure. Despite its importance, existing approaches often face two key limitations: (i) They employ dense networks to solely explore the class hierarchy as hard criterion for maintaining taxonomic consistency among predicted classes, yet without leveraging rich semantic relationships between instances and classes; (ii) They struggle to generalize in settings with deep class levels, since the mini-batches uniformly sampled from different levels ignore the varying complexities of data and result in a non-smooth model adaptation to sparse data. To mitigate these issues, we present a Self-Paced Unified Representation (SPUR) learning framework, which focuses on the interplay between instance and classes to flexibly organize the training process of HMLC algorithms. Our framework consists of two lightweight encoders designed to capture the semantics of input features and the topological information of the class hierarchy. These encoders generate unified embeddings of instances and class hierarchy, which enable SPUR to exploit semantic dependencies between them and produce predictions in line with taxonomic constraints. Furthermore, we introduce a dynamic hardness measurement strategy that considers both class hierarchy and instance features to estimate the learning difficulty of each instance. This strategy is achieved by incorporating the propagation loss obtained at each hierarchical level, allowing for a more comprehensive assessment of learning complexity. Extensive experiments on several empirical benchmarks demonstrate the effectiveness and efficiency of SPUR compared to state-of-the-art methods, especially in scenarios with missing features. Zixuan Yuan, Hao Liu 0026, Haoyi Zhou, Xiao Zhang 0015, Hao Wang 0073, Hui Xiong 0001 |
AAAI | 6 |
| 2024 | Self-Supervised Federated Learning for Personalized Human Activity RecognitionabstractPersonalized Human Activity Recognition (PHAR) based on wearable sensors is crucial in the medical, sports, industrial and other fields. PHAR faces challenges of privacy leakage and a shortage of labeled data. Therefore, we propose a framework called self-supervised federated learning for personalized human activity recognition (SSF-HAR) to implement private PHAR. To protect user privacy, our framework integrates federated learning (FL) to achieve the transmission of only model parameters between the cloud and clients, rather than user data. Besides, we propose a strategy of weighted aggregation to update the cloud model with the client models. To overcome the lack of labeled data, our framework introduces self-supervised learning (SSL) tasks to pretrain a feature extractor in the cloud. The proxy task of SSL transforms data and provides pseudo-labels in three forms. We test the performance on the benchmark datasets MotionSense and WIDSM. The experiments show that SSF-HAR outperforms other FL frameworks for PHAR. Shizhuo Deng, Da Teng, Zhubao Guo, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
ICME | 7 |
| 2024 | APPN: An Attention-based Pseudo-label Propagation Network for few-shot learning with noisy labels
Shizhuo Deng, Da Teng, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
Neurocomputing | 6 |
| 2024 | Toward Dual-View X-Ray Baggage Inspection: A Large-Scale Benchmark and Adaptive Hierarchical Cross Refinement for Prohibited Item DiscoveryabstractDual-view baggage inspection has been widely applied in real-world scenarios, where orthogonal viewpoints are deployed to capture diverse and complementary information. Compared with single-view, it can effectively improve the identification performance when rotation and overlay hinder the viewability of the objects. However, this topic has not been rigorously explored due to the scarcity of datasets. To overcome this limitation, we contribute the first fully public large-scale Dual-view X-ray dataset. Our dataset, named DvXray, contains 16,000 pairs, 32,000 X-ray images, in which 15 common classes of 5,496 prohibited items are manually labeled. Besides, we propose an approach named Adaptive Hierarchical Cross Refinement (AHCR) to establish a strong baseline for prohibited item discovery in dual-view X-ray images. AHCR hypothesizes that each input pair is sampled from one mixture distribution, hence gathering the non-overlapping and position-aware cues along the shared axis and complementarily delivering to the other in a hierarchical structure to enrich the feature discriminability of the objects of interest from background overlaps. Upon this structure, we propose an adaptive control strategy and a confidence-weighted view fusion term to make it robust to difficult samples. Extensive experiments on DvXray show that AHCR not only brings significant classification gains over various backbones, such as recent Swin Transformer and ConvNeXt, but also exhibits an impressively better ability to localize objects. In addition, AHCR performs favorably against the counterparts and some recent multi-view learning approaches, moving a step closer towards potential application in practice. Dataset and code are available at https://github.com/Mbwslib/DvXray. Tong Jia 0001, Mingyuan Li 0003, Songsheng Wu, Hao Wang 0073, Dongyue Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Delving Into Cluttered Prohibited Item Detection for Security Inspection SystemabstractProhibited item detection in X-Ray baggage images can efficiently prevent the social security and stability. With the development of deep learning, applying specific methods in computer vision tasks to prohibited item detection has shown promising perspectives, and the resulted intelligent security inspection system can effectively address the limitations of human inspection. Despite several deep learning based methods have been proposed to flourish this researching field, there are still two issues which have not been fully explored. First, most of them suffer from the shortage of dataset and collection of massive and well-annotated samples is extremely laborious and costly. Second, the designation of backbone modules for current prohibited item detection methods shares similar idea with the ones in object detection. While little consideration has been paid to the properties of prohibited item, where most of them are over-lapped and cluttered. To overcome limitations, this paper proposes a cluttered prohibited item detection method for security inspection system. Specifically, our method first generates synthetic X-Ray images through cut-and-paste strategy from the training samples in each training mini-batch, where the strategy effectively and efficiently augments the dataset and quality of the synthetic samples can be guaranteed. Then, a high-order dilated convolution module is developed for enriching the representation ability of the feature, further promoting the localization ability for over-lapped and cluttered prohibited items. Experiments show that our proposed method can be well generalized to various datasets, and achieve clear improvement over state-of-the-art methods. Hao Wang 0073, Tong Jia 0001, Dongyue Chen 0001, Shizhuo Deng |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Relation Knowledge Distillation by Auxiliary Learning for Object DetectionabstractBalancing the trade-off between accuracy and speed for obtaining higher performance without sacrificing the inference time is a challenging topic for object detection task. Knowledge distillation, which serves as a kind of model compression techniques, provides a potential and feasible way to handle above efficiency and effectiveness issue through transferring the dark knowledge from the sophisticated teacher detector to the simple student one. Despite demonstrating promising solutions to make harmonies between accuracy and speed, current knowledge distillation for object detection methods still suffer from two limitations. Firstly, most of the methods are inherited or refereed from the frameworks in image classification task, and deploy an implicit manner by imitating or constraining the features from the intermediate layers or the output predictions between the teacher and student models. While little consideration has been raised to the intrinsic relevance of the classification and localization predictions in object detection task. Besides, these methods fail to investigate the relationship between detection and distillation tasks in knowledge distillation pipeline, and they train the whole network by simply integrating losses from these two different tasks through hand-crafted designation parameters. For addressing the aforementioned issues, we propose a novel Relation Knowledge Distillation by Auxiliary Learning for Object Detection (ReAL) method in this paper. Specifically, we first design a prediction relation distillation module which makes the student model directly mimic the output predictions from the teacher one, and conduct self and mutual relation distillation losses to excavate the relation information between teacher and student models. Moreover, for better devolving into the relationship between different tasks in distillation pipeline, we introduce the auxiliary learning into knowledge distillation for object detection and develop a dynamic weight adaptation strategy. Through regarding detection task as primary task and treating distillation task as auxiliary task in auxiliary learning framework, we dynamically adjust and regularize the corresponding weights of the losses for these tasks during the training process. Experiments on MS COCO dataset are conducted using various detector combinations of teacher and student models and the results show that our proposed ReAL can achieve obvious improvement on different distillation model configurations, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2024 | LHAR: Lightweight Human Activity Recognition on Knowledge DistillationabstractSensor-based Human Activity Recognition (HAR) is widely used in daily life and is the basic-level bridge to virtual healthcare in the metaverse. The current challenge is the low recognition accuracy for personalized users on smart wearable devices. The limited resource cannot support large deep learning models updated locally. Besides, integrating and transmitting sensor data to the cloud would reduce the efficiency. Considering the tradeoff between performance and complexity, we propose a Lightweight Human Activity Recognition (LHAR) framework. In LHAR, we combine the cross-people HAR task with the lightweight model task. LHAR framework is designed on the teacher-student architecture and the student network consists of multiple depthwise separable convolution layers to achieve fewer parameters. The dark knowledge distilled from the complex teacher model enhances the generalization ability of LHAR. To achieve effective knowledge distillation, we propose two optimization methods. Firstly, we train the teacher model by ensemble learning to promote teacher performance. Secondly, a multi-channel data augmentation method is proposed for the diversity of the dataset, which is a plug-in operation for the ensemble teacher model. In the experiments, we compare LHAR with state-of-art models in comparison evaluation, ablation study and the hyperparameter analysis, which proves the better performance of LHAR in efficiency and effectiveness. Shizhuo Deng, Da Teng, Chuangui Yang, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | CBDMoE: Consistent-but-Diverse Mixture of Experts for Domain GeneralizationabstractMachine learning models often suffer from severe performance degradation due to distributional shifts between testing and training data. To address this issue, researchers have focused on domain generalization (DG), which aims to generalize a model trained on source domains to arbitrary unseen target domains. Recently, ensemble learning has emerged as a popular strategy for addressing the DG problem, and domain-specific experts are typically involved. However, the existing methods do not sufficiently consider the generalizability of individual experts or leverage the consistency and diversity among them, thus limiting the generalizability of the constructed models. In this paper, we propose a consistent-but-diverse mixture of experts (CBDMoE) algorithm, which is an improved MoE framework that effectively harnesses ensemble learning for solving the DG problem. Specifically, we introduce individual expert learning (IEL), which incorporates a novel domain-class-balanced subset division (DCBSD)-based sampling strategy to facilitate a generalizable expert learning process. Additionally, we present consistent-but-diverse learning (CBDL), which employs two regularizing losses to encourage consistency and diversity in the predictions of the experts. Our proposed strategy significantly enhances the generalizability of the MoE framework. Extensive experiments conducted on three popular DG benchmark datasets demonstrate that our method outperforms the state-of-the-art approaches. Fangbin Xu, Dongyue Chen 0001, Tong Jia 0001, Shizhuo Deng, Hao Wang 0073 |
IEEE Trans. Multim. | 5 |
| 2023 | UUKG: Unified Urban Knowledge Graph Dataset for Urban Spatiotemporal PredictionabstractAccurate Urban SpatioTemporal Prediction (USTP) is of great importance to the development and operation of the smart city. As an emerging building block, multi-sourced urban data are usually integrated as urban knowledge graphs (UrbanKGs) to provide critical knowledge for urban spatiotemporal prediction models. However, existing UrbanKGs are often tailored for specific downstream prediction tasks and are not publicly available, which limits the potential advancement. This paper presents UUKG, the unified urban knowledge graph dataset for knowledge-enhanced urban spatiotemporal predictions. Specifically, we first construct UrbanKGs consisting of millions of triplets for two metropolises by connecting heterogeneous urban entities such as administrative boroughs, POIs, and road segments. Moreover, we conduct qualitative and quantitative analysis on constructed UrbanKGs and uncover diverse high-order structural patterns, such as hierarchies and cycles, that can be leveraged to benefit downstream USTP tasks. To validate and facilitate the use of UrbanKGs, we implement and evaluate 15 KG embedding methods on the KG completion task and integrate the learned KG embeddings into 9 spatiotemporal models for five different USTP tasks. The extensive experimental results not only provide benchmarks of knowledge-enhanced USTP models under different task settings but also highlight the potential of state-of-the-art high-order structure-aware UrbanKG embedding methods. We hope the proposed UUKG fosters research on urban knowledge graphs and broad smart city applications. The dataset and source code are available at https://github.com/usail-hkust/UUKG/. Yansong Ning, Hao Liu 0026, Hao Wang 0073, Zhenyu Zeng, Hui Xiong 0001 |
NeurIPS | 3 |
| 2023 | Self-relation attention networks for weakly supervised few-shot activity recognition
Shizhuo Deng, Zhubao Guo, Da Teng, Boqian Lin, Dongyue Chen 0001, Tong Jia 0001, Hao Wang 0073 |
Knowl. Based Syst. | 7 |
| 2023 | Fully Cascade Consistency Learning for One-Stage Object DetectionabstractObject detection is usually solved by deploying one single prediction head including classification and localization branches to obtain the final results. Recently proposed works utilize several prediction heads in a cascade learning manner to improve the detection performance. Despite achieving promising performance, existing cascade learning manner methods still meet with two inconsistency issues. Firstly, most of them refine the bounding boxes in different prediction heads only by depending on the localization accuracy (i.e., IoU), while ignoring the inconsistency between classification confidence and localization accuracy. Moreover, simply increasing the IoU threshold by experience to select positive samples makes the inconsistency issue even worse. Secondly, little consideration has been paid on the feature inconsistency between detection-specific features from different prediction heads and detection-generalized ones from backbone model. The extracted feature from backbone model contains the general representation for the whole images. While prediction heads need to be carefully designed to have specific ability which contains more discriminative expressions for the two sub-tasks classification and regression. The different contexture representations of the output features from these two parts lead to the feature inconsistency between backbone model and prediction head in cascade learning architecture. To solve these two inconsistency issues, this paper proposes a novel cascade consistency learning method for one-stage detector. Specifically, a feature adaptation module is firstly developed to calibrate features from different prediction heads and backbone model for solving the feature inconsistency. Then, we design an automatic positive sample threshold selection strategy for further solve the inconsistency between the classification and localization predictions. Moreover, the quality of bounding boxes in cascade learning manner are evaluated by taking both the classification confidence and localization accuracy into consideration. Experiments on MS COCO show that our proposed cascade consistency learning manner (dubbed$\text{C}^{2}\text{L}$) can achieve clear improvement over counterparts based on several different one-stage detectors, while performing favorably against state-of-the-arts. Hao Wang 0073, Tong Jia 0001, Qilong Wang 0001, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Semi-Supervised Wide-Angle Portraits Correction by Multi-Scale TransformerabstractWe propose a semi-supervised network for wide-angle portraits correction. Wide-angle images often suffer from skew and distortion affected by perspective distortion, especially noticeable at the face regions. Previous deep learning based approaches need the ground-truth correction flow maps for training guidance. However, such labels are expensive, which can only be obtained manually. In this work, we design a semi-supervised scheme and build a high-quality unlabeled dataset with rich scenarios, allowing us to simultaneously use labeled and unlabeled data to improve performance. Specifically, our semi-supervised scheme takes advantage of the consistency mechanism, with several novel components such as direction and range consistency (DRC) and regression consistency (RC). Furthermore, different from the existing methods, we propose the Multi-Scale Swin-Unet (MS-Unet) based on the multi-scale swin transformer block (MSTB), which can simultaneously learn short-distance and long-distance information to avoid artifacts. Extensive experiments demonstrate that the proposed method is superior to the state-of-the-art methods and other representative baselines. The source code and dataset are available at https://github.corn/megvii-research/PortraitsCorrection Fushun Zhu, Shan Zhao 0010, Hao Wang 0073, Shuaicheng Liu |
CVPR | 4 |
| 2022 | CrabNet: Fully Task-Specific Feature Learning for One-Stage Object DetectionabstractObject detection is usually solved by learning a deep architecture involving classification and localization tasks, where feature learning for these two tasks is shared using the same backbone model. Recent works have shown that suitable disentanglement of classification and localization tasks has the great potential to improve performance of object detection. Despite the promising performance, existing feature disentanglement methods usually suffer from two limitations. First, most of them only focus on the disentangled proposals or predication heads for classification and localization tasks after RPN. While little consideration has been given to that the features for these two different tasks actually are obtained by a shared backbone model before RPN. Second, they are suggested for two-stage objectors and are not applicable to one-stage methods. To overcome these limitations, this paper presents a novel fully task-specific feature learning method for one-stage object detection. Specifically, our method first learns disentangled features for classification and localization tasks using two separated backbone models, where auxiliary classification and localization heads are inserted at the end of the two backbone models for providing a fully task-specific features for classification and localization. Then, a feature interaction module is developed for aligning and fusing task-specific features, which are further used to produce the final detection result. Experiments on MS COCO show that our proposed method (dubbed CrabNet) can achieve clear improvement over counterparts with increasing limited inference time, while performing favorably against state-of-the-arts. Hao Wang 0073, Qilong Wang 0001, Qinghua Hu, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2021 | Multi-scale structural kernel representation for object detection
Hao Wang 0073, Qilong Wang 0001, Peihua Li, Wangmeng Zuo |
Pattern Recognit. | 1 |
| 2021 | Constrained Online Cut-Paste for Object DetectionabstractWell-annotated training samples show necessity in achieving high performance of object detection, but collection of massive samples is extremely laborious and costly. Recently, cut-paste based methods show the potential to augment the training samples by cutting the foreground instances and pasting them on some background regions. However, existing cut-paste based methods hardly guarantee the quality of synthetic images due to lack of mechanism to ensure rationality of the pasted instances (e.g., context, geometry and diversity), limiting the effectiveness of data augmentation. To overcome above issues, this paper proposes a novel Constrained Online Cut-Paste (COCP) method, making an attempt to effectively and efficiently augment training data for improving performance of object detection. Specifically, our COCP generates synthetic images by switching instances of same class from various image pairs in each training mini-batch, ensuring context coherence between the cut instances and the pasted backgrounds. Furthermore, two constraints based on geometric consistency and sample diversity are developed to eliminate counterproductive and meaningless switched instances those suffer from significant geometric discrepancy or lack variations, further improving quality of the synthetic images. The experiments are conducted on both MS COCO and PASCAL VOC datasets using various state-of-the-art detectors (e.g., Faster R-CNN, RetinaNet, FCOS and Mask R-CNN). The results show that our proposed COCP can be well generalized to various datasets and detectors with clear performance gains, while performing favorably against its counterparts. Hao Wang 0073, Qilong Wang 0001, Jian Yang 0003, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Multi-Scale Location-Aware Kernel Representation for Object DetectionabstractAlthough Faster R-CNN and its variants have shown promising performance in object detection, they only exploit simple first-order representation of object proposals for final classification and regression. Recent classification methods demonstrate that the integration of high-order statistics into deep convolutional neural networks can achieve impressive improvement, but their goal is to model whole images by discarding location information so that they cannot be directly adopted to object detection. In this paper, we make an attempt to exploit high-order statistics in object detection, aiming at generating more discriminative representations for proposals to enhance the performance of detectors. To this end, we propose a novel Multi-scale Location-aware Kernel Representation (MLKP) to capture high-order statistics of deep features in proposals. Our MLKP can be efficiently computed on a modified multi-scale feature map using a low-dimensional polynomial kernel approximation. Moreover, different from existing orderless global representations based on high-order statistics, our proposed MLKP is location retentive and sensitive so that it can be flexibly adopted to object detection. Through integrating into Faster R-CNN schema, the proposed MLKP achieves very competitive performance with state-of-the-art methods, and improves Faster R-CNN by 4.9% (mAP), 4.7% (mAP) and 5.0% (AP at IOU=[0.5:0.05:0.95]) on PASCAL VOC 2007, VOC 2012 and MS COCO benchmarks, respectively. Code is available at: https://github.com/Hwang64/MLKP. Hao Wang 0073, Qilong Wang 0001, Mingqi Gao 0006, Peihua Li, Wangmeng Zuo |
CVPR | 1 |
| 2018 | Masking Effects Based Rate Control Scheme for High Efficiency Video CodingabstractThis paper presents a masking effects based rate control scheme for high efficiency video coding (HEVC). Rate control is regarded as a very effective tool to improve the performance of video coding under the limited bandwidth. However, the state-of-the-art rate control algorithm based on R-X model ignores the characteristics of human visual system (HVS), which leads to poor performance in subjective quality. Moreover, some structural similarity (SSIM) or saliency based perceptual rate control algorithms only consider spatial characteristics. Since spatial and temporal visual masking effects can better reflect the characteristics of HVS, in this paper masking effects based perceptual factor for coding tree unit (CTU) is proposed, which takes both texture complexity and motion information into account. Then the proposed perceptual factor is utilized to guide bit allocation in CTU-level rate control. Experimental results show that the proposed scheme can effectively improve the coding performance compared with the R-λ algorithm. Hao Wang 0073, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001 |
ISCAS | 1 |