VLDB 2026 Research / reviewers in the wild / expert
Yandong Guo
dblp:28/4272
· DBLP profile ↗
80ranked-venue papers
3as first author
66since 2021 · last 2026
0000-0002-4594-8415ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 3 first-author · 46 since 2021Artificial intelligence and machine learning · 47 · 2 first-author · 40 since 2021Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compression-Oriented Video Super-ResolutionabstractCurrent compressed video super-resolution methods have achieved promising performance, but they often assume that an input video is compressed under low-delay configurations. However, under random access configurations, those methods might struggle to leverage the metadata effectively due to the large variations of metadata in different compression configurations. In this work, we propose a Compression-Oriented Video Super-Resolution (COVSR) method that can address video super-resolution for both low-delay and random-access configurations. Specifically, we first introduce an efficient compression-aware propagation (ECAP) module that dynamically adjusts propagation routes in accordance with the compression configurations. Since existing methods require reconstructing frames in a frame-by-frame manner, it is difficult to achieve efficient parallelization. However, we find that by slightly relaxing sequential dependencies, our ECAP can significantly improve inference speed. Furthermore, existing methods typically perform alignment between adjacent frames or adjacent features. However, since ECAP may propagate features along non-adjacent reference routes, it introduces new challenges for accurate cross-frame feature alignment. In response, we propose a metadata-driven alignment (MDA) module that refines cross-frame motion vectors into dense, feature-level flow offsets, enabling precise alignment across temporally distant features. Extensive experimental results demonstrate that our COVSR not only achieves efficient and superior super-resolution performance but also is generalizable to various compression configurations. Our code will be available at https://covsr.github.io. Yanbin Liu 0003, Ming Lu 0002, Zhuojie Wu, Senmao Tian, Yandong Guo, Xin Yu 0002 |
IEEE Trans. Image Process. | 6 |
| 2025 | LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR UnderstandingabstractRecently, Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have shown promise in instruction following and image understanding. While these models are powerful, they have not yet been developed to comprehend the more challenging 3D geometric and physical scenes, especially when it comes to the sparse outdoor LiDAR data. In this paper, we introduce LiDAR-LLM, which takes raw LiDAR data as input and harnesses the remarkable reasoning capabilities of LLMs to gain a comprehensive understanding of outdoor 3D scenes. The central insight of our LiDAR-LLM is the reformulation of 3D outdoor scene cognition as a language modeling problem, encompassing tasks such as 3D captioning, 3D grounding, 3D question answering, etc. Specifically, due to the scarcity of 3D LiDAR-text pairing data, we introduce a three-stage training strategy and generate relevant datasets, progressively aligning the 3D modality with the language embedding of LLM. Furthermore, we design a Position-Aware Transformer (PAT) to connect the 3D encoder with the LLM, which effectively bridges the modality gap and enhances the LLM's spatial orientation comprehension of visual features. Our experiments demonstrate that LiDAR-LLM effectively comprehends a wide range of instructions related to 3D scenes, achieving a 40.9 BLEU-1 score on the 3D captioning dataset, a Grounded Captioning accuracy of 63.1%, and a BEV mIoU of 14.3%. Senqiao Yang, Jiaming Liu 0003, Renrui Zhang, Mingjie Pan, Xiaoqi Li 0009, Peng Gao 0007, Hongsheng Li 0001, Yandong Guo, Shanghang Zhang |
AAAI | 10 |
| 2025 | 3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level AlignmentabstractThe 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Category-level ambiguity arises from representing objects of fine-grained categories in a highly sparse point cloud format, making category distinction challenging. Instance-level complexity stems from multiple instances of the same category coexisting in a scene, leading to distractions during grounding. To address these challenges, we propose a novel weaklysupervised grounding approach that explicitly differentiates between categories and instances. In the category-level branch, we utilize extensive category knowledge from a pre-trained external detector to align object proposal features with sentencelevel category features, thereby enhancing category awareness. In the instance-level branch, we utilize spatial relationship descriptions from language queries to refine object proposal features, ensuring clear differentiation among objects. These designs enable our model to accurately identify target-category objects while distinguishing instances within the same category. Compared to previous methods, our approach achieves state-of-the-art performance on three widely used benchmarks: Nr3D, Sr3D, and ScanRef. Xiaoqi Li 0020, Jiaming Liu 0003, Nuowei Han, Liang Heng, Yandong Guo, Hao Dong 0003, Yang Liu 0105 |
ICRA | 5 |
| 2025 | Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningabstractGeneralized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning capabilities of internet-scale pretrained vision-language models (VLMs), they often suffer from low execution frequency. To mitigate this dilemma, dual-system approaches have been proposed to leverage a VLM-based System 2 module for handling high-level decision-making, and a separate System 1 action module for ensuring real-time control. However, existing designs maintain both systems as separate models, limiting System 1 from fully leveraging the rich pretrained knowledge from the VLM-based System 2. In this work, we propose Fast-in-Slow (FiS), a unified dual-system vision-language-action (VLA) model that embeds the System 1 execution module within the VLM-based System 2 by partially sharing parameters. This innovative paradigm not only enables high-frequency execution in System 1, but also facilitates coordination between multimodal reasoning and execution components within a single foundation model of System 2. Given their fundamentally distinct roles within FiS-VLA, we design the two systems to incorporate heterogeneous modality inputs alongside asynchronous operating frequencies, enabling both fast and precise manipulation. To enable coordination between the two systems, a dual-aware co-training strategy is proposed that equips System 1 with action generation capabilities while preserving System 2’s contextual understanding to provide stable latent conditions for System 1. For evaluation, FiS-VLA outperforms previous state-of-the-art methods by 8% in simulation and 11% in real-world tasks in terms of average success rate, while achieving a 117.7 Hz control frequency with action chunk set to eight. Project web page: https://fast-in-slow.github.io. Hao Chen 0193, Jiaming Liu 0003, Chenyang Gu, Zhuoyang Liu, Renrui Zhang, Xiaoqi Li 0020, Yandong Guo, Chi-Wing Fu, Shanghang Zhang, Pheng-Ann Heng |
NeurIPS | 8 |
| 2025 | MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning SegmentationabstractReasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demonstrated impressive 2D image reasoning segmentation, adapting these capabilities to 3D scenes remains underexplored. In this paper, we introduce MLLM-For3D, a simple yet effective framework that transfers knowledge from 2D MLLMs to 3D scene understanding. Specifically, we utilize MLLMs to generate multi-view pseudo-segmentation masks and corresponding text embeddings, then unproject 2D masks into 3D space and align them with the text embeddings. The primary challenge lies in the absence of 3D context and spatial consistency across multiple views, causing the model to hallucinate objects that do not exist and fail to target objects consistently. Training the 3D model with such irrelevant objects leads to performance degradation. To address this, we first filter irrelevant views using token attention. With these reliable pseudo-labels, we develop a token-for-Query approach for multimodal semantic alignment, enabling consistent identification of the same object across different views. Moreover, we introduce a spatial consistency strategy to enforce that segmentation masks remain coherent in the 3D space, effectively capturing the geometry of the scene. Extensive evaluations of various challenging indoor scene benchmarks demonstrate that, even without labeled 3D training data, MLLM-For3D outperforms existing 3D reasoning segmentation methods, effectively interpreting user intent, understanding 3D scenes, and reasoning about spatial relationships. Runnan Chen, Zhengqing Gao, Yandong Guo, Mingming Gong, Tongliang Liu |
NeurIPS | 6 |
| 2025 | Surprise3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D ScenesabstractThe integration of language and 3D perception is critical for embodied AI and robotic systems to perceive, understand, and interact with the physical world. Spatial reasoning, a key capability for understanding spatial relationships between objects, remains underexplored in current 3D vision-language research. Existing datasets often mix semantic cues (e.g., object name) with spatial context, leading models to rely on superficial shortcuts rather than genuinely interpreting spatial relationships. To address this gap, we introduce Surprise3D, a novel dataset designed to evaluate language-guided spatial reasoning segmentation in complex 3D scenes. Surprise3D consists of more than 200k vision language pairs across 900+ detailed indoor scenes from ScanNet++ v2, including more than 2.8k unique object classes. The dataset contains 89k+ human-annotated spatial queries deliberately crafted without object name, thereby mitigating shortcut biases in spatial understanding. These queries comprehensively cover various spatial reasoning skills, such as relative position, narrative perspective, parametric perspective, and absolute distance reasoning. Initial benchmarks demonstrate significant challenges for current state-of-the-art expert 3D visual grounding methods and 3D-LLMs, underscoring the necessity of our dataset and the accompanying 3D Spatial Reasoning Segmentation (3D-SRS) benchmark suite. Surprise3D and 3D-SRS aim to facilitate advancements in spatially aware AI, paving the way for effective embodied interaction and robotic planning. Hanlue Zhang, Runnan Chen, Zhengqing Gao, Yandong Guo, Wenping Wang 0001, Tongliang Liu, Mingming Gong |
NeurIPS | 7 |
| 2025 | Towards Accurate Time Series Forecasting via Implicit DecodingabstractRecent booming time series models have demonstrated remarkable forecasting performance. However, these methods often place greater focus on more effectively modelling the historical series, largely neglecting the forecasting phase, which generates long-term forecasts by separately predicting multiple time points. Given that real-world time series typically consist of various long short-term dynamics, independent predictions over individual time points may fail to express complex underlying patterns and can lead to a lack of global views. To address these issues, this work explores new perspectives from the forecasting phase and proposes a novel Implicit Forecaster (IF) as an additional decoding module. Inspired by decomposition forecasting, IF adopts a more nuanced approach by implicitly predicting constituent waves represented by their frequency, amplitude, and phase, thereby accurately forming the time series. Extensive experimental results from multiple real-world datasets show that IF can consistently boost mainstream time series models, achieving state-of-the-art forecasting performance. Code is available at this repository: [https://github.com/rakuyorain/Implicit-Forecaster](https://github.com/rakuyorain/Implicit-Forecaster). Hao Wang 0014, Haoxuan Li 0001, Liuhua Peng, Feng Liu 0003, Yandong Guo, Kun Zhang 0001, Mingming Gong |
NeurIPS | 7 |
| 2024 | Debiased Novel Category Discovering and LocalizationabstractIn recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often identified as background or incorrectly classified as pre-defined categories by the detectors. In this paper, we focus on the challenging problem of Novel Class Discovery and Localization (NCDL), aiming to train detectors that can detect the categories present in the training data, while also actively discover, localize, and cluster new categories. We analyze existing NCDL methods and identify the core issue: object detectors tend to be biased towards seen objects, and this leads to the neglect of unseen targets. To address this issue, we first propose an Debiased Region Mining (DRM) approach that combines class-agnostic Region Proposal Network (RPN) and class-aware RPN in a complementary manner. Additionally, we suggest to improve the representation network through semi-supervised contrastive learning by leveraging unlabeled data. Finally, we adopt a simple and efficient mini-batch K-means clustering method for novel class discovery. We conduct extensive experiments on the NCDL benchmark, and the results demonstrate that the proposed DRM approach significantly outperforms previous methods, establishing a new state-of-the-art. Juexiao Feng, Yuhong Yang 0008, Yanchun Xie, Yandong Guo, Liuyu Xiang, Guiguang Ding |
AAAI | 5 |
| 2024 | Continual-MAE: Adaptive Distribution Masked Autoencoders for Continual Test-Time AdaptationabstractContinual Test-Time Adaptation (CTTA) is proposed to migrate a source pretrained model to continually changing target distributions, addressing real-world dynamism. Existing CTTA methods mainly rely on entropy minimization or teacher-student pseudo-labeling schemes for knowledge extraction in unlabeled target domains. However, dynamic data distributions cause miscalibrated predictions and noisy pseudo-labels in existing self-supervised learning methods, hindering the effective mitigation of error accumulation and catastrophic forgetting problems during the continual adaptation process. To tackle these issues, we propose a continual self-supervised method, Adaptive Distribution Masked Autoencoders (ADMA), which enhances the extraction of target domain knowledge while mitigating the accumulation of distribution shifts. Specifically, we propose a Distribution-aware Masking (DaM) mechanism to adaptively sample masked positions, followed by establishing consistency constraints between the masked target samples and the original target samples. Additionally, for masked tokens, we utilize an efficient decoder to reconstruct a handcrafted feature descriptor (e.g., Histograms of Oriented Gradients), leveraging its invariant properties to boost task-relevant representations. Through conducting extensive experiments on four widely recognized benchmarks, our proposed method attains state-of-the-art performance in both classification and segmentation CTTA tasks. Jiaming Liu 0003, Ran Xu 0013, Senqiao Yang, Renrui Zhang, Qizhe Zhang, Yandong Guo, Shanghang Zhang |
CVPR | 7 |
| 2024 | NTO3D: Neural Target Object 3D Reconstruction with Segment AnythingabstractNeural 3D reconstruction from multi-view images has recently attracted increasing attention from the community. Existing methods normally learn a neural field for the whole scene, while it is still under-explored how to reconstruct a target object indicated by users. Considering the Segment Anything Model (SAM) has shown effectiveness in segmenting any 2D images, in this paper, we propose NTO3D, a novel high-quality Neural Target Object 3D (NTO3D) reconstruction method, which leverages the benefits of both neural field and SAM. We first propose a novel strategy to lift the multi-view 2D segmentation masks of SAM into a unified 3D occupancy field. The 3D occupancy field is then projected into 2D space and generates the new prompts for SAM. This process is iterative until convergence to sepa-rate the target object from the scene. After this, we then lift the 2D features of the SAM encoder into a 3D feature field in order to improve the reconstruction quality of the target object. NTO3D lifts the 2D masks and features of SAM into the 3D neural field for high-quality neural target object 3D reconstruction. We conduct detailed experiments on several benchmark datasets to demonstrate the advantages of our method. The code will be available at: https://github.com/ucwxb/NTO3D. Xiaobao Wei, Renrui Zhang, Jiarui Wu, Jiaming Liu 0003, Ming Lu 0002, Yandong Guo, Shanghang Zhang |
CVPR | 6 |
| 2024 | Tag2Text: Guiding Vision-Language Model via Image TaggingabstractThis paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance. Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001 |
ICLR | 8 |
| 2024 | ViDA: Homeostatic Visual Domain Adapter for Continual Test Time AdaptationabstractSince real-world machine systems are running in non-stationary environments, Continual Test-Time Adaptation (CTTA) task is proposed to adapt the pre-trained model to continually changing target domains. Recently, existing methods mainly focus on model-based adaptation, which aims to leverage a self-training manner to extract the target domain knowledge. However, pseudo labels can be noisy and the updated model parameters are unreliable under dynamic data distributions, leading to error accumulation and catastrophic forgetting in the continual adaptation process. To tackle these challenges and maintain the model plasticity, we design a Visual Domain Adapter (ViDA) for CTTA, explicitly handling both domain-specific and domain-shared knowledge. Specifically, we first comprehensively explore the different domain representations of the adapters with trainable high-rank or low-rank embedding spaces. Then we inject ViDAs into the pre-trained model, which leverages high-rank and low-rank features to adapt the current domain distribution and maintain the continual domain-shared knowledge, respectively. To exploit the low-rank and high-rank ViDAs more effectively, we further propose a Homeostatic Knowledge Allotment (HKA) strategy, which adaptively combines different knowledge from each ViDA. Extensive experiments conducted on four widely used benchmarks demonstrate that our proposed method achieves state-of-the-art performance in both classification and segmentation CTTA tasks. Note that, our method can be regarded as a novel transfer paradigm for large-scale models, delivering promising results in adaptation to continually changing distributions. Jiaming Liu 0003, Senqiao Yang, Peidong Jia, Renrui Zhang, Ming Lu 0002, Yandong Guo, Wei Xue 0002, Shanghang Zhang |
ICLR | 6 |
| 2024 | BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object DetectionabstractVision-centric bird-eye-view (BEV) perception has shown promising potential in autonomous driving. Recent works mainly focus on improving efficiency or accuracy but neglect the challenges when facing environment changing, resulting in severe degradation of transfer performance. For BEV perception, we figure out the significant domain gaps existing in typical real-world cross-domain scenarios and comprehensively solve the Domain Adaption (DA) problem for multi-view 3D object detection. Since BEV perception approaches are complicated and contain several components, the domain shift accumulation on multiple geometric spaces (i.e., 2D, 3D Voxel, BEV) makes BEV DA even challenging. In this paper, we propose a Multi-space Alignment Teacher-Student (MATS) framework to ease the domain shift accumulation, which consists of a Depth-Aware Teacher (DAT) and a Geometric-space Aligned Student (GAS) model. DAT tactfully combines target lidar and reliable depth prediction to construct depth-aware information, extracting target domain-specific knowledge in Voxel and BEV feature spaces. It then transfers the sufficient domain knowledge of multiple spaces to the student model. In order to jointly alleviate the domain shift, GAS projects multi-geometric space features to a shared geometric embedding space and decreases data distribution distance between two domains. To verify the effectiveness of our method, we conduct BEV 3D object detection experiments on three cross-domain scenarios and achieve state-of-the-art performance. Code: https://github.com/liujiaming1996/BEVUDA. Jiaming Liu 0003, Rongyu Zhang, Xiaoqi Li 0020, Xiaowei Chi, Ming Lu 0002, Yandong Guo, Shanghang Zhang |
ICRA | 7 |
| 2024 | RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and ManipulationabstractA fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) insufficient reasoning ability to tackle complex tasks, and (2) high computational costs for VLA model fine-tuning and inference. The recently proposed state space model (SSM) known as Mamba demonstrates promising capabilities in non-trivial sequence modeling with linear inference complexity. Inspired by this, we introduce RoboMamba, an end-to-end robotic VLA model that leverages Mamba to deliver both robotic reasoning and action capabilities, while maintaining efficient fine-tuning and inference. Specifically, we first integrate the vision encoder with Mamba, aligning visual tokens with language embedding through co-training, empowering our model with visual common sense and robotic-related reasoning. To further equip RoboMamba with SE(3) pose prediction abilities, we explore an efficient fine-tuning strategy with a simple policy head. We find that once RoboMamba possesses sufficient reasoning capability, it can acquire manipulation skills with minimal fine-tuning parameters (0.1\% of the model) and time. In experiments, RoboMamba demonstrates outstanding reasoning capabilities on general and robotic evaluation benchmarks. Meanwhile, our model showcases impressive pose prediction results in both simulation and real-world experiments, achieving inference speeds 3 times faster than existing VLA models. Jiaming Liu 0003, Zhenyu Wang 0002, Pengju An, Xiaoqi Li 0020, Kaichen Zhou, Senqiao Yang, Renrui Zhang, Yandong Guo, Shanghang Zhang |
NeurIPS | 9 |
| 2024 | Rethinking the Person Localization for Single-Stage Multi-Person Pose EstimationabstractSingle-stage models for multi-person pose estimation have garnered significant attention due to their streamlined approach in generating person position localization and body structure perception in a single pass. These two parts, however, are processed individually by existing methods, leading to suboptimal results, e.g., candidates with high confidences for person localization while poor structure estimations. To this end, we propose a simple yet effective approach, namely Structure-guided Person Localization (SPL), jointly leveraging the advantages of the two aspects to solve the multi-person pose estimation problem, with two complementary novelties. First, we propose to incorporate body structure perception to guide person position localization, consequently, we introduce the Structure-guided Center Learning (SCL) to unify the quality of the body structure perception in the displacement map with the confidence of the person existence in the center map, thus achieving more accurate keypoint position localization results even with extreme poses. Second, to facilitate the end-to-end training of SPL, we propose the efficient Agency-based Scale-adaptive Learning (ASL). Specifically, we predict an agency map of the same size as the center map, which focuses on the foreground area and can adaptively adjust the scale size for each central area with the body structure perception confidence. Comprehensive experiments on challenging benchmarks including COCO and CrowdPose clearly verify the superiority of our framework, which achieves new state-of-the-art single-stage multi-person pose estimation results. Specifically, SPL obtains 72.1 AP scores and 69.5 AP scores in COCO test-dev2017 and CrowdPose test set, respectively. Lei Jin 0003, Xuecheng Nie, Wendong Wang 0003, Yandong Guo, Shuicheng Yan, Jian Zhao 0006 |
IEEE Trans. Multim. | 5 |
| 2023 | CloSET: Modeling Clothed Humans on Continuous Surface with Explicit Template DecompositionabstractCreating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing details or hinder end-to-end learning. In this paper, we revisit point-based solutions and propose to decompose explicit garment-related templates and then add pose-dependent wrinkles to them. In this way, the clothing deformations are disentangled such that the pose-dependent wrinkles can be better learned and applied to unseen poses. Additionally, to tackle the seam artifact issues in recent state-of-the-art point-based methods, we propose to learn point features on a body surface, which establishes a continuous and compact feature space to capture the fine-grained and pose-dependent clothing geometry. To facilitate the research in this field, we also introduce a high-quality scan dataset of humans in real-world clothing. Our approach is validated on two existing datasets and our newly introduced dataset, showing better clothing deformation results in unseen poses. The project page with code and dataset can be found at https://www.liuyebin.com/closet. Hongwen Zhang 0001, Siyou Lin, Ruizhi Shao, Yuxiang Zhang 0006, Zerong Zheng, Han Huang 0005, Yandong Guo, Yebin Liu |
CVPR | 7 |
| 2023 | PiMAE: Point Cloud and Image Interactive Masked Autoencoders for 3D Object DetectionabstractMasked Autoencoders learn strong visual representations and achieve state-of-the-art results in several independent modalities, yet very few works have addressed their capabilities in multi-modality settings. In this work, we focus on point cloud and RGB image data, two modalities that are often presented together in the real world, and explore their meaningful interactions. To improve upon the cross-modal synergy in existing works, we propose Pi-MAE, a self-supervised pre-training framework that promotes 3D and 2D interaction through three aspects. Specifically, we first notice the importance of masking strategies between the two sources and utilize a projection module to complementarily align the mask and visible tokens of the two modalities. Then, we utilize a well-crafted two-branch MAE pipeline with a novel shared decoder to promote cross-modality interaction in the mask tokens. Finally, we design a unique cross-modal reconstruction module to enhance representation learning for both modalities. Through extensive experiments performed on large-scale RGB-D scene understanding benchmarks (SUN RGB-D and ScannetV2), we discover it is nontrivial to interactively learn point-image features, where we greatly improve multiple 3D detectors, 2D detectors, and few-shot classifiers by 2.9%, 6.7%, and 2.4%, respectively. Code is available at https://github.com/BLVLab/PiMAE. Anthony Chen, Renrui Zhang, Zihan Wang 0011, Yuheng Lu, Yandong Guo, Shanghang Zhang |
CVPR | 6 |
| 2023 | BEV-SAN: Accurate BEV 3D Object Detection via Slice Attention NetworksabstractBird'View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature extraction, BEV feature construction, and task heads. Among the three components, BEV feature construction is BEV-specific compared with 2D tasks. Existing methods aggregate the multi-view camera features to the flattened grid in order to construct the BEV feature. However, flattening the BEV space along the height dimension fails to emphasize the informative features of different heights. For example, the barrier is located at a low height while the truck is located at a high height. In this paper, we propose a novel method named BEV Slice Attention Network (BEV-SAN) for exploiting the intrinsic characteristics of different heights. Instead of flattening the BEV space, we first sample along the height dimension to build the global and local BEV slices. Then, the features of BEV slices are aggregated from the camera features and merged by the attention mechanism. Finally, we fuse the merged local and global BEV features by a transformer to generate the final feature map for task heads. The purpose of local BEV slices is to emphasize informative heights. In order to find them, we further propose a LiDAR-guided sampling strategy to leverage the statistical distribution of LiDAR to determine the heights of local slices. Compared with uniform sampling, LiDAR-guided sampling can determine more informative heights. We conduct detailed experiments to demonstrate the effectiveness of BEV-SAN. Code will be released. Xiaowei Chi, Jiaming Liu 0003, Ming Lu 0002, Rongyu Zhang, Zhaoqing Wang, Yandong Guo, Shanghang Zhang |
CVPR | 6 |
| 2023 | Box-Level Active DetectionabstractActive learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images. Furthermore, existing methods still perform image-level annotation, but equally scoring all targets within the same image incurs waste of budget and redundant labels. Having revealed above problems and limitations, we introduce a box-level active detection framework that controls a box-based budget per cycle, prioritizes informative targets and avoids redundancy for fair comparison and efficient application. Under the proposed box-level setting, we devise a novel pipeline, namely Complementary Pseudo Active Strategy (ComPAS). It exploits both human annotations and the model intelligence in a complementary fashion: an efficient input-end committee queries labels for informative objects only; meantime well-learned targets are identified by the model and compensated with pseudo-labels. ComPAS consistently outperforms 10 competitors under 4 settings in a unified codebase. With supervision from labeled data only, it achieves 100% supervised performance of VOC0712 with merely 19% box annotations. On the COCO dataset, it yields up to 4.3% mAP improvement over the second-best method. ComPAS also supports training with the unlabeled pool, where it surpasses 90% COCO supervised performance with 85% label reduction. Our source code is publicly available at https://github.com/lyumengyao/blad. Mengyao Lyu, Jundong Zhou, Hui Chen 0013, Dongdong Yu, Yandong Guo, Liuyu Xiang, Guiguang Ding |
CVPR | 7 |
| 2023 | Learning Audio-Visual Source Localization via False Negative Aware Contrastive LearningabstractSelf-supervised audio-visual source localization aims to locate sound-source objects in video frames without extra annotations. Recent methods often approach this goal with the help of contrastive learning, which assumes only the audio and visual contents from the same video are positive samples for each other. However, this assumption would suffer from false negative samples in real-world training. For example, for an audio sample, treating the frames from the same audio class as negative samples may mislead the model and therefore harm the learned representations (e.g., the audio of a siren wailing may reasonably correspond to the ambulances in multiple images). Based on this observation, we propose a new learning strategy named False Negative Aware Contrastive (FNAC) to mitigate the problem of misleading the training with such false negative samples. Specifically, we utilize the intra-modal similarities to identify potentially similar samples and construct corresponding adjacency matrices to guide contrastive learning. Further, we propose to strengthen the role of true negative samples by explicitly leveraging the visual features of sound sources to facilitate the differentiation of authentic sounding source regions. FNAC achieves state-of-the-art performances on Flickr-SoundNet, VGG-Sound, and AVSBench, which demonstrates the effectiveness of our method in mitigating the false negative issue. The code is available at https://github.com/OpenNLPLab/FNAC_AVL. Weixuan Sun, Zheyuan Liu 0002, Yiran Zhong, Tianpeng Feng, Yandong Guo, Nick Barnes |
CVPR | 7 |
| 2023 | CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network with Large InputabstractWith the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and memory cost, current methods first split the large input into local patches and then merge the SR patches into the output. These methods adaptively allocate a subnet for each patch. Quantization is a very important technique for network acceleration and has been used to design the subnets. Current methods train an MLP bit selector to determine the propoer bit for each layer. However, they uniformly sample subnets for training, making simple subnets overfitted and complicated subnets underfitted. Therefore, the trained bit selector fails to determine the optimal bit. Apart from this, the introduced bit selector brings additional cost to each layer of the$SR$network. In this paper, we propose a novel method named Content-Aware Bit Mapping (CABM), which can remove the bit selector without any performance loss. CABM also learns a bit selector for each layer during training. After training, we analyze the relation between the edge information of an input patch and the bit of each layer. We observe that the edge information can be an effective metric for the selected bit. Therefore, we design a strategy to build an Edge-to-Bit lookup table that maps the edge score of a patch to the bit of each layer during inference. The bit configuration of SR network can be determined by the lookup tables of all layers. Our strategy can find better bit configuration, resulting in more efficient mixed precision networks. We conduct detailed experiments to demonstrate the generalization ability of our method. The code will be released. Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Yurong Chen 0001, Shunli Zhang 0005 |
CVPR | 4 |
| 2023 | Ultra Real-Time Portrait Matting via Parallel Semantic GuidanceabstractMost existing portrait matting models either require expensive auxiliary information or try to decompose the task into sub-tasks that are usually resource-hungry. These challenges limit its application on low-power computing devices. In this paper, we propose an ultra-light-weighted portrait matting network via parallel semantic guidance (PSGNet) for real-time portrait matting without any auxiliary inputs. PSGNet leverages parallel multi-level semantic information to efficiently guide the feature representations to replace traditional sequential semantic hints from objective decomposition. We also introduce an efficient fusion module to effectively combine parallel branches of PSGNet to minimize the representation redundancy. Comprehensive experiments demonstrate that our PSGNet can achieve remarkable performance on both synthetic and real-world images. Our PSGNet is capable to process at 100fps thanks to its ultra-small number of parameters, which makes it deployable on low-power computing devices without compromising on the performance of real-time portrait matting. Jiake Xie, Bo Xu 0031, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
ICASSP | 7 |
| 2023 | A Comprehensive Comparison of Projections in Omnidirectional Super-ResolutionabstractSuper-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resolution of omnidirectional frames. In these works, omnidirectional frames are projected from the 3D sphere to a 2D plane by Equi-Rectangular Projection (ERP). Although ERP has been widely used for projection, it has severe projection distortion near poles. Current DNN-based SR methods use 2D convolution modules, which is more suitable for the regular grid. In this paper, we find that different projection methods have great impact on the performance of DNNs. To study this problem, a comprehensive comparison of projections in omnidirectional super-resolution is conducted. We compare the SR results of different projection methods. Experimental results show that Equi-Angular cube map projection (EAC), which has minimal distortion, achieves the best result in terms of WS-PSNR compared with other projections. Code and data will be released. Huicheng Pi, Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Shunli Zhang 0005 |
ICASSP | 5 |
| 2023 | Mixed Sample Augmentation for Online DistillationabstractMixed Sample Regularization (MSR), such as MixUp or CutMix, is a powerful data augmentation strategy to generalize convolutional neural networks. Previous empirical analysis has illustrated an orthogonal performance gain between MSR and conventional offline Knowledge Distillation (KD). To be more specific, student networks can be enhanced with the involvement of MSR in the training stage of sequential distillation. Yet, the interplay between MSR and online knowledge distillation, where an ensemble of peer students learn mutually from each other, remains unexplored. To bridge the gap, we make the first attempt at incorporating CutMix into online distillation, where we empirically observe a significant improvement. Encouraged by this fact, we propose an even stronger MSR specifically for online distillation, named as CutnMix. Furthermore, a novel online distillation framework is designed upon CutnMix, to enhance the distillation with feature level mutual learning and a self-ensemble teacher. Comprehensive evaluations on CIFAR10 and CIFAR100 with six network architectures show that our approach can consistently outperform state-of-the-art distillation methods. Yiqing Shen 0003, Liwu Xu, Yuzhe Yang 0001, Yandong Guo |
ICASSP | 5 |
| 2023 | Neural Reconstruction of Relightable Human Model from Monocular VideoabstractCreating relightable and animatable human characters from monocular video at a low cost is a critical task for digital human modeling and virtual reality applications. This task is complex due to intricate articulation motion, a wide range of ambient lighting conditions, and pose-dependent clothing deformations. In this paper, we introduce a novel self-supervised framework that takes a monocular video of a moving human as input and generates a 3D neural representation capable of being rendered with novel poses under arbitrary lighting conditions. Our framework decomposes dynamic humans under varying illumination into neural fields in canonical space, taking into account geometry and spatially varying BRDF material properties. Additionally, we introduce pose-driven deformation fields, enabling bidirectional mapping between canonical space and observation. Leveraging the proposed appearance decomposition and deformation fields, our framework learns in a self-supervised manner. Ultimately, based on pose-driven deformation, recovered appearance, and physically-based rendering, the reconstructed human figure becomes relightable and can be explicitly driven by novel poses. We demonstrate significant performance improvements over previous works and provide compelling examples of relighting from monocular videos of moving humans in challenging, uncontrolled capture scenarios. Wenzhang Sun, Yunlong Che, Yandong Guo, Han Huang 0005 |
ICCV | 3 |
| 2023 | Mosaic Representation Learning for Self-supervised Visual Pre-training
Zhaoqing Wang, Yandong Guo, Jun Yu 0001, Mingming Gong, Tongliang Liu |
ICLR | 4 |
| 2023 | Attribute-assisted Multimodal Network for Image Aesthetics AssessmentabstractImage aesthetics assessment (IAA) is challenging due to its highly abstract nature. Nowadays, people tend to share images and comment them on social networks, which can provide rich information for judging image aesthetics. As a result, user comments of an image can be jointly utilized to learn better feature representations for IAA. Previous researches have shown that aesthetic attributes are crucial factors in determining image aesthetic quality and influencing people’s aesthetic perception. Accordingly, when commenting an image, people usually give descriptions from the perspective of aesthetic attributes. Inspired by this, this paper presents a new Attribute-Assisted Multimodal network (AAM-Net) for image aesthetics assessment. Specifically, we propose a cross-modal attribute interaction module to explore the related aesthetic attribute semantics shared by an image and the corresponding aesthetic comments. Then, a cross-modal gate unit is introduced to further refine significant attribute semantics interactively. Finally, informative aesthetic features can be obtained for predicting image aesthetic distributions. Experimental results on two public multimodal IAA databases demonstrate the superiority of the proposed model over the state-of-the-art methods. Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo |
ICME | 7 |
| 2023 | ContrastMotion: Self-supervised Scene Motion Learning for Large-Scale LiDAR Point CloudsabstractIn this paper, we propose a novel self-supervised motion estimator for LiDAR-based autonomous driving via BEV representation. Different from usually adopted self-supervised strategies for data-level structure consistency, we predict scene motion via feature-level consistency between pillars in consecutive frames, which can eliminate the effect caused by noise points and view-changing point clouds in dynamic scenes. Specifically, we propose Soft Discriminative Loss that provides the network with more pseudo-supervised signals to learn discriminative and robust features in a contrastive learning manner. We also propose Gated Multi-Frame Fusion block that learns valid compensation between point cloud frames automatically to enhance feature extraction. Finally, pillar association is proposed to predict pillar correspondence probabilities based on feature distance, and whereby further predicts scene motion. Extensive experiments show the effectiveness and superiority of our ContrastMotion on both scene flow and motion prediction tasks. Xiangze Jia, Hui Zhou 0005, Xinge Zhu, Yandong Guo, Yuexin Ma |
IJCAI | 4 |
| 2023 | Data-Driven Based Cascading Orientation and Translation Estimation for Inertial NavigationabstractRecently, data-driven approaches have brought both opportunities and challenges for Inertial Navigation Systems. In this paper, we propose a novel data-driven method which is composed of cascading orientation and translation estimation with IMU-only measurements. For robust orientation estimation, we combine a CNN-based neural network with an EKF to eliminate orientation errors caused by sensor noises. We additionally propose a hybrid CNN-Transformer-based neural network which exploits both spatial and long-term temporal information to regress accurate translations. Specifically, we conduct detailed evaluations on datasets acquired by iPhone and Android devices. The result demonstrates that our method outperforms state-of-the-art methods in both orientation and translation errors. Xiangyu Deng, Shenyue Wang, Chunxiang Shan, Jinjie Lu, Jijunnan Li, Yandong Guo |
IROS | 7 |
| 2023 | SELVO: A Semantic-Enhanced Lidar-Visual OdometryabstractIn the face of complex external environment, single sensor information can no longer meet the accuracy requirements of low-drift SLAM. In this paper, we focus on the fusion scheme of cameras and lidar, and explore the gain of semantic information to SLAM system. A Semantic-Enhanced Lidar-Visual Odometry (SELVO) is proposed to achieve pose estimation with high accuracy and robustness by applying semantics and utilizing strategies of initialization and sensor fusion. In loop closure detection thread, we propose a novel place recognition method based on semantic information to maintain the global consistency of the map. In the back-end, we design a joint optimization framework including visual odometry, lidar odometry and loop closure detection, and innovatively propose to recognize degraded scenes with semantic information. We have conducted a large number of experiments on KITTI [1] and KITTI-360 [2] dataset, and the results show that our system can achieve the high accuracy and competitive performance in comparison with state-of-the-art methods. Jijunnan Li, Yandong Guo, Shijie Liu 0002, Chunlai Li 0003, Jianyu Wang 0017 |
IROS | 5 |
| 2023 | RepCaM: Re-parameterization Content-aware Modulation for Neural Video DeliveryabstractRecently, content-aware methods have been utilized to reduce the bandwidth and improve the quality of Internet video delivery. Existing methods train corresponding content-aware super-resolution (SR) models for each video chunk on the server and stream low-resolution (LR) video chunks along with SR models to the client. Previous works introduce additional partial parameters to privatize the models of different video chunks. However, this still leads to the accumulation of parameters and even fails to modulate when the length of video increases, bringing extra delivery costs and performance degradation. In this paper, we introduce a novel Re-parameterization Content-aware Modulation (RepCaM) method to modulate all the video chunks with an end-to-end training strategy. Our method adopts extra parallel-cascade parameters during training to fit multiple chunks while removing the additional parameters through re-parameterization during inference. Therefore, RepCaM increases no extra model size compared with the original SR model. Moreover, in order to improve the training efficiency on servers, we propose an online Video Patch Sampling (VPS) method to speed up the training convergence. We conduct extensive experiments on VSD4K and newly collected dataset (VSD4K-2022), achieving state-of-the-art results in video restoration quality and delivery bandwidth compression. Code is available at: https://github.com/Neural-video-delivery/RepCaM-Pytorch-NOSSDAV2023. Rongyu Zhang, Lixuan Du, Jiaming Liu 0003, Congcong Song, Fangxin Wang 0001, Xiaoqi Li 0009, Ming Lu 0002, Yandong Guo, Shanghang Zhang |
NOSSDAV | 8 |
| 2023 | Video Object Matting via Hierarchical Space-Time Semantic GuidanceabstractDifferent from most existing approaches that require trimap generation for each frame, we reformulate video object matting (VOM) by introducing improved semantic guidance propagation. The proposed approach can achieve a higher degree of temporal coherence between frames with only a single coarse mask as a reference. In this paper, we adapt the hierarchical memory matching mechanism into the space-time baseline to build an efficient and robust framework for semantic guidance propagation and alpha prediction. To enhance the temporal smoothness, we also propose a cross-frame attention refinement (CFAR) module that can refine the feature representations across multiple adjacent frames (both historical and current frames) based on the spatio-temporal correlation among the cross- frame pixels. Extensive experiments demonstrate the effectiveness of hierarchical spatio-temporal semantic guidance and the cross-video-frame attention refinement module, and our model outperforms the state-of-the-art VOM methods. We also analyze the significance of different components in our model. Bo Xu 0031, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
WACV | 6 |
| 2023 | AdvFAS: A robust face anti-spoofing framework against adversarial examplesabstractEnsuring the reliability of face recognition systems against presentation attacks necessitates the deployment of face anti-spoofing techniques. Despite considerable advancements in this domain, the ability of even the most state-of-the-art methods to defend against adversarial examples remains elusive. While several adversarial defense strategies have been proposed, they typically suffer from constrained practicability due to inevitable trade-offs between universality, effectiveness, and efficiency. To overcome these challenges, we thoroughly delve into the coupled relationship between adversarial detection and face anti-spoofing. Based on this, we propose a robust face anti-spoofing framework, namely AdvFAS, that leverages two coupled scores to accurately distinguish between correctly detected and wrongly detected face images. Extensive experiments demonstrate the effectiveness of our framework in a variety of settings, including different attacks, datasets, and backbones, meanwhile enjoying high accuracy on clean examples. Moreover, we successfully apply the proposed method to detect real-world adversarial examples. Xiao Yang 0028, Mingzhi Ma, Bihui Chen, Jianteng Peng, Yandong Guo, Zhao-Xia Yin, Hang Su 0006 |
Comput. Vis. Image Underst. | 7 |
| 2023 | Improving the robustness of adversarial attacks using an affine-invariant gradient estimator
Wenzhao Xiang 0001, Hang Su 0006, Chang Liu 0077, Yandong Guo, Shibao Zheng |
Comput. Vis. Image Underst. | 4 |
| 2023 | Anchor-based knowledge embedding for image aesthetics assessment
Leida Li, Tianwu Zhi, Guangming Shi, Yuzhe Yang 0001, Liwu Xu, Yandong Guo |
Neurocomputing | 7 |
| 2023 | SuperFast: 200× Video Frame Interpolation via Event CameraabstractTraditional frame-based video frame interpolation (VFI) methods rely on the linear motion assumption and brightness invariance assumption, which may lead to fatal errors confronting the scenarios with high-speed motions. To tackle the above challenge, inspired by the advantages of event cameras on asynchronously recording brightness changes at each pixel, we propose a Fast-Slow joint synthesis framework for event-enhanced high-speed video frame interpolation, named SuperFast, in this paper, which can generate high frame rate (5000 FPS, 200× faster) video from the input low frame rate (25 FPS) video and the corresponding event stream. In our framework, the task is divided into two sub-tasks, i.e., video frame interpolation for the contents with and without high-speed motions, which are tackled by two corresponding branches, i.e., the fast synthesis pathway and the slow synthesis pathway. The fast synthesis pathway leverages a spiking neural network to encode the input event stream, and combines boundary frames to generate intermediate results through synthesis and refinement, targeting on contents with high-speed motions. The slow synthesis pathway stacks the two input boundary frames and the event stream to synthesize intermediate results, focusing on relatively slow-motion contents. Finally, a fusion module with a comparison loss is utilized to generate the final video frame interpolation results. We also build a hybrid visual acquisition system containing an event camera and a high frame rate camera, and collect the first 5000 FPS High-Speed Event-enhanced Video frame Interpolation (THU[Formula: see text]) dataset. To evaluate the performance of our proposed framework, we have conducted experiments on our THU[Formula: see text] dataset and the existing HS-ERGB dataset. Experimental results demonstrate that our proposed framework can achieve state-of-the-art 200× video frame interpolation performance under high-speed motion scenarios. Yue Gao 0002, Siqi Li 0001, Yandong Guo, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Theme-Aware Visual Attribute Reasoning for Image Aesthetics AssessmentabstractPeople usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes. Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Explainable and Generalizable Blind Image Quality Assessment via Semantic Attribute ReasoningabstractBlind image quality assessment (BIQA) that can directly evaluate image quality without perfect-quality reference has been a long-standing research topic. Although the existing BIQA models have achieved very encouraging performance, the lack of explainability and generalization ability limits their real-world applications to a great extent. People usually assess image quality according to semantic attributes, e.g., brightness, color, contrast, noise and sharpness. Furthermore, judgment on image quality is also impacted by the scene presented in the image. Therefore, the inherent relationship between semantic attributes and scenes is crucial for image quality assessment, which has rarely been explored yet. With this motivation, this paper presents a Semantic Attribute Reasoning based image QUality Evaluator (SARQUE). Specifically, we propose a two-stream network to predict semantic attributes and scene categories from distorted images. To investigate the inherent relationship between the semantic attributes and scene category, a semantic reasoning module is further proposed based on the graph convolution network (GCN), producing the final quality score. Extensive experiments conducted on five in-the-wild image quality databases demonstrate the superiority of the proposed SARQUE model over the state-of-the-arts. Furthermore, the proposed model features better explainability and generalization ability due to the use of semantic attributes. Yipo Huang, Leida Li, Yuzhe Yang 0001, Yandong Guo |
IEEE Trans. Multim. | 5 |
| 2023 | Grouping by Center: Predicting Centripetal Offsets for the Bottom-up Human Pose EstimationabstractWe introduce Grouping by Center, a novel grouping approach for the bottom-up human pose estimation, which detects human joint first and then does grouping. The grouping strategy is the critical factor for the bottom-up pose estimation. To increase the conciseness and accuracy, we propose to use the center of the body as a grouping clue. More concretely, we predict the offsets from the keypoints to the body centers. Keypoints with aligned shifted results will be grouped as one person. However, the multi-scale variance of people can affect the prediction of the grouping clue, which has been neglected in previous research. To resolve the scale variance of the offset, we put forward a Multi-scale Translation Layer and an iterative refinement. Furthermore, we scheme a greedy grouping strategy with a dynamic threshold due to the various scales of instances. Through a comprehensive comparison, our framework is validated to be effective and practical. We also lay out the state-of-the-art performance revolving the bottom-up multi-person pose estimation on the MS-COCO dataset and the CrowdPose dataset. Lei Jin 0003, Xuecheng Nie, Luoqi Liu, Yandong Guo, Jian Zhao 0006 |
IEEE Trans. Multim. | 5 |
| 2023 | Knowledge-Guided Blind Image Quality Assessment With Few Training SamplesabstractBlind image quality assessment (BIQA) for in-the-wild images has achieved great progress by training advanced deep neural networks. However, the current BIQA models are suffering the generalization challenge, meaning that a well-trained BIQA model is still very limited in evaluating images with different distributions. Deep BIQA models are data-intensive, but the annotation of image quality labels is extremely expensive. To design a generalizable BIQA model with few training samples is highly desired. Motivated by the above fact, this paper presents a knowledge-guided BIQA (KG-IQA) framework by integrating domain knowledge from the human visual system (HVS) and natural scene statistics (NSS). Specifically, the quality-aware HVS and NSS features are first extracted as prior knowledge. Then, we embed the two types of knowledge into the conventional deep neural network by learning to predict the HVS and NSS features, producing the knowledge-enhanced quality features, based on which the final image quality score is obtained. We conduct extensive experiments and comparisons on five authentically distorted IQA datasets. The experimental results demonstrate that the introduction of knowledge greatly reduces the requirement on the amount of training images, and the proposed KG-IQA model achieves superior performance in terms of both prediction accuracy and generalization ability. Tianshu Song, Leida Li, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi |
IEEE Trans. Multim. | 6 |
| 2023 | Learning Personalized Image Aesthetics From Subjective and Objective AttributesabstractDue to the widespread popularity of social media, researchers have developed a strong interest in learning the personalized image aesthetics of online users. Personalized image aesthetics assessment (PIAA) aims to study the aesthetic preferences of individual users for images, which should be affected by the properties of both users and images. Existing PIAA approaches usually use the generic aesthetics learned from images as a prior model and adapt it to PIAA models through a small number of data annotated by individual users. However, the prior model merely learns the objective attributes of images, which is agnostic to the subjective attributes of users, complicating efficient learning of the personalized image aesthetics of individual users. Therefore, we propose a personalized image aesthetics assessment method that integrates the subjective attributes of users and objective attributes of images simultaneously. To characterize these two attributes jointly, an attribute extraction module is introduced to learn users’ personality traits and image aesthetic attributes. Then, an aesthetic prior model is built from numerous individual users’ annotated data, which leverages the personality traits of users and the aesthetic attributes of rated images as prior knowledge to model both the image aesthetic distribution and users’ residual scores relative to generic aesthetics simultaneously. Finally, a PIAA model is obtained by fine-tuning the aesthetic prior model with an individual user’s annotated data. Experiments demonstrate that the proposed method is superior to existing PIAA methods in learning individual users’ personalized image aesthetics. Hancheng Zhu, Yong Zhou 0003, Leida Li, Yandong Guo |
IEEE Trans. Multim. | 5 |
| 2022 | On the Efficacy of Small Self-Supervised Contrastive Models without Distillation SignalsabstractIt is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be suitable for some resource-restricted scenarios due to the huge computational expenses of deploying a large model. In this paper, we study the issue of training self-supervised small models without distillation signals. We first evaluate the representation spaces of the small models and make two non-negligible observations: (i) the small models can complete the pretext task without overfitting despite their limited capacity and (ii) they universally suffer the problem of over clustering. Then we verify multiple assumptions that are considered to alleviate the over-clustering phenomenon. Finally, we combine the validated techniques and improve the baseline performances of five small architectures with considerable margins, which indicates that training small self-supervised contrastive models is feasible even without distillation signals. The code is available at https://github.com/WOWNICE/ssl-small. Haizhou Shi, Youcai Zhang, Siliang Tang, Wenjie Zhu 0003, Yandong Guo, Yueting Zhuang |
AAAI | 6 |
| 2022 | Self-Distillation from the Last Mini-Batch for Consistency RegularizationabstractKnowledge distillation (KD) shows a bright promise as a powerful regularization strategy to boost generalization ability by leveraging learned sample-level soft targets. Yet, employing a complex pre-trained teacher network or an ensemble of peer students in existing KD is both timeconsuming and computationally costly. Various self KD methods have been proposed to achieve higher distillation efficiency. However, they either require extra network architecture modification or are difficult to parallelize. To cope with these challenges, we propose an efficient and reliable self-distillation framework, named Self-Distillation from Last Mini-Batch (DLB). Specifically, we rearrange the sequential sampling by constraining half of each mini-batch coinciding with the previous iteration. Meanwhile, the rest half will coincide with the upcoming iteration. Afterwards, the former half mini-batch distills on-the-fly soft targets generated in the previous iteration. Our proposed mechanism guides the training stability and consistency, resulting in robustness to label noise. Moreover, our method is easy to implement, without taking up extra run-time memory or requiring model structure modification. Experimental results on three classification benchmarks illustrate that our approach can consistently outperform state-of-the-art self-distillation approaches with different network architectures. Additionally, our method shows strong compatibility with augmentation strategies by gaining additional performance improvement. The code is available at https://github.com/Meta-knowledge-Lab/DLB. Yiqing Shen 0003, Liwu Xu, Yuzhe Yang 0001, Yandong Guo |
CVPR | 5 |
| 2022 | Single-Stage is Enough: Multi-Person Absolute 3D Pose EstimationabstractThe existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both efficiency and performance. To this end, we present an efficient single-stage solution, Decoupled Regression Model (DRM), with three distinct novelties. First, DRM introduces a new decoupled representation for 3D pose, which expresses the 2D pose in image plane and depth information of each 3D human instance via 2D center point (center of visible keypoints) and root point (denoted as pelvis), respectively. Second, to learn better feature representation for the human depth regression, DRM introduces a 2D Pose-guided Depth Query Module (PDQM) to extract the features in 2D pose regression branch, enabling the depth regression branch to perceive the scale information of instances. Third, DRM leverages a Decoupled Absolute Pose Loss (DAPL) to facilitate the absolute root depth and root-relative depth estimation, thus improving the accuracy of absolute 3D pose. Comprehensive experiments on challenging benchmarks including MuPoTS-3D and Panoptic clearly verify the superiority of our framework, which outperforms the state-of-the-art bottom-up absolute 3D pose estimation methods. Lei Jin 0003, Yabo Xiao, Yandong Guo, Xuecheng Nie, Jian Zhao 0006 |
CVPR | 5 |
| 2022 | CRIS: CLIP-Driven Referring Image SegmentationabstractReferring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet separately transfer the language/vision knowledge from pretrained models, ignoring the multi-modal corresponding information. Inspired by the recent advance in Contrastive Language-Image Pretraining (CLIP), in this paper, we propose an end-to-end CLIP-Driven Referring Image Segmen-tation framework (CRIS). To transfer the multi-modal knowledge effectively, CRIS resorts to vision-language decoding and contrastive learning for achieving the text-to-pixel alignment. More specifically, we design a vision-language decoder to propagate fine-grained semantic information from textual representations to each pixel-level activation, which promotes consistency between the two modalities. In addition, we present text-to-pixel contrastive learning to explicitly enforce the text feature similar to the related pixel-level features and dissimilar to the irrelevances. The experimental results on three benchmark datasets demonstrate that our proposed framework significantly outperforms the state-of-the-art performance without any post-processing. Zhaoqing Wang, Qiang Li 0024, Xunqiang Tao, Yandong Guo, Mingming Gong, Tongliang Liu |
CVPR | 5 |
| 2022 | Personalized Image Aesthetics Assessment with Rich AttributesabstractPersonalized image aesthetics assessment (PIAA) is challenging due to its highly subjective nature. People's aesthetic tastes depend on diversified factors, including image characteristics and subject characters. The existing PIAA databases are limited in terms of annotation diversity, especially the subject aspect, which can no longer meet the increasing demands of PIAA research. To solve the dilemma, we conduct so far, the most comprehensive subjective study of personalized image aesthetics and introduce a new Personalized image Aesthetics database with Rich Attributes (PARA), which consists of 31,220 images with annotations by 438 subjects. PARA features wealthy annotations, including 9 image-oriented objective attributes and 4 human-oriented subjective attributes. In addition, desensitized subject information, such as personality traits, is also provided to support study of PIAA and user portraits. A comprehensive analysis of the annotation data is provided and statistic study indicates that the aesthetic preferences can be mirrored by proposed subjective attributes. We also propose a conditional PIAA model by utilizing subject information as conditional prior. Experimental results indicate that the conditional PIAA model can outperform the control group, which is also the first attempt to demonstrate how image aesthetics and subject characters interact to produce the intricate personalized tastes on image aesthetics. We believe the database and the associated analysis would be useful for conducting next-generation PIAA study. The project page of PARA can be found at: https://cv-datasets.institutecv.com/#/data-sets. Yuzhe Yang 0001, Liwu Xu, Leida Li, Nan Qie, Yandong Guo |
CVPR | 7 |
| 2022 | Structured Local Radiance Fields for Human Avatar ModelingabstractIt is extremely challenging to create an animatable clothed human avatar from RGB videos, especially for loose clothes due to the difficulties in motion modeling. To address this problem, we introduce a novel representation on the basis of recent neural scene rendering techniques. The core of our representation is a set of structured local radiance fields, which are anchored to the pre-defined nodes sampled on a statistical human body template. These local radiance fields not only leverage the flexibility of implicit representation in shape and appearance modeling, but also factorize cloth deformations into skeleton motions, node residual translations and the dynamic detail variations inside each individual radiance field. To learn our representation from RGB data and facilitate pose generalization, we propose to learn the node translations and the detail variations in a conditional generative latent space. Overall, our method enables automatic construction of animatable human avatars for various types of clothes without the need for scanning subject-specific templates, and can generate realistic images with dynamic details for novel poses. Experiment show that our method outperforms state-of-the-art methods both qualitatively and quantitatively. Zerong Zheng, Han Huang 0005, Tao Yu 0007, Hongwen Zhang 0001, Yandong Guo, Yebin Liu |
CVPR | 5 |
| 2022 | Efficient Meta-Tuning for Content-Aware Neural Video Delivery
Xiaoqi Li 0009, Jiaming Liu 0003, Shizun Wang, Ming Lu 0002, Yurong Chen 0001, Anbang Yao, Yandong Guo, Shanghang Zhang |
ECCV (18) | 8 |
| 2022 | Adaptive Patch Exiting for Scalable Single Image Super-Resolution
Shizun Wang, Jiaming Liu 0003, Kaixin Chen 0001, Xiaoqi Li 0009, Ming Lu 0002, Yandong Guo |
ECCV (18) | 6 |
| 2022 | SDETR: Attention-Guided Salient Object Detection with TransformerabstractMost existing CNN-based salient object detection methods can identify fine-grained segmentation details like hair and animal fur, but often mispredict the salient object due to lack of global contextual information caused by locality convolution layers. The limited training data of the current SOD task adds additional difficulty to capture the saliency information. In this paper, we propose a two-stage predict-refine SDETR model to leverage both benefits of transformer and CNN layers that can produce results with accurate saliency prediction and fine-grained local details. We also propose a novel pre-train dataset annotation COCO SOD to erase the overfitting problem caused by insufficient training data. Comprehensive experiments on five benchmark datasets demonstrate that the SDETR outperforms state-of-the-art approaches on four evaluation metrics, and our COCO SOD can largely improve the model performance on DUTS, ECSSD, DUT, PASCAL-S datasets. Guanze Liu, Bo Xu 0031, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
ICASSP | 5 |
| 2022 | DARTS-PD: Differentiable Architecture Search with Path-Wise Weight Sharing DerivationabstractWith the advent of Neural Architecture Search (NAS), auto-designing of artificial neural networks has been made possible. Among various NAS methods, Differentiable Architecture Search (DARTS) has achieved significant progress due to its high calculating efficiency. However, it suffers from poor stability and obvious performance drop because of bi-level optimization and hard pruning. Besides, it only generates one best architecture at once. To alleviate the problems above, we design a three-stage framework with a path-wise weight sharing derivation. We first prune the supernet with differentiable methods to keep top-k operations on each edge instead of one. Then the pruned supernet is trained with our path-wise weight sharing method. At the derivation stage, the best candidate operations are selected with Evolutionary Search based on the validation accuracy of paths. Our weight sharing derivation is proved effective in improving searching stability as well as alleviating the performance drop. Furthermore, it also allows us to search for a large number of architectures with different parameter sizes at one time. Comprehensive experiments on CIFAR-10 and ImageNet show that we manage to find a group of state-of-the-art architectures (97.61% on CIFAR-10 and 76.4% on ImageNet). He Cai, Zhaokai Zhang, Tianpeng Feng, Yandong Guo |
ICIP | 4 |
| 2022 | Psychology Inspired Model for Hierarchical Image Aesthetic Attribute PredictionabstractDeep neural network has proved its effectiveness in image aesthetic quality assessment (IAQA), but still lacks reasonable interpretability. Aesthetic attributes provide rich intermediate-level information for understanding the underlying principles of image aesthetics, but has not been fully investigated. Psychological studies have shown that aesthetic experience involves hierarchical stages, i.e., human process image aesthetics following a staged information processing mechanism. Motivated by this, this paper presents a Hierarchical Image Aesthetic Attribute (HIAA) prediction model, aiming to imitate the staged mechanism of human aesthetic experience. Image aesthetic attributes are first divided into several hierarchical groups. Then, hierarchical features are extracted from the cascaded layers of the deep neural network to predict the aesthetic attributes in a group-wise manner. The overall image aesthetic score is also predicted by aggregating the hierarchical features. Experimental results demonstrate that the proposed HIAA model outperforms the state-of-the-arts in terms of both aesthetic attribute prediction and aesthetic score regression. Leida Li, Jiachen Duan, Yuzhe Yang 0001, Liwu Xu, Yandong Guo |
ICME | 6 |
| 2022 | ONavi: Data-driven based Multi-sensor Fusion Positioning System in Indoor EnvironmentsabstractThis paper proposes a multi-sensor fusion system, named ONavi, that fuses WiFi and IMU to provide an accurate positioning service on smartphones in indoor environments. In this system, a hybrid CNN-Transformer-based neural network is proposed for our Pedestrian Dead Reckoning(PDR), which outperforms existing state-of-the-art methods. Additionally, in our data-driven WiFi positioning module, instead of pure RSSI based WiFi feature, a “Fusion BSSID-RSSI” feature is proposed, which significantly improves positioning accuracy. Eventually, we use a loosely coupled optimization-based framework to fuse the aforementioned positioning results. Quantitative evaluations demonstrate that ONavi is capable of achieving outstanding performance of positioning estimation in real indoor environment. Jinjie Lu, Chunxiang Shan, Xiangyu Deng, Shenyue Wang, Yuepeng Wu, Jijunnan Li, Yandong Guo |
IPIN | 8 |
| 2022 | Pose Refinement with Joint Optimization of Visual Points and LinesabstractHigh-precision camera re-localization technology in a pre-established 3D environment map is the basis for many tasks, such as Augmented Reality, Robotics and Autonomous Driving. The point-based visual re-localization approaches are well-developed in recent decades, but are insufficient in some feature-less cases. In this paper, we design a complete pipeline for camera pose refinement with points and lines, which contains the innovatively designed line extracting CNN named VLSE, the line matching and the pose optimization approaches. We adopt a novel line representation and customize a hybrid convolution block based on the Stacked Hourglass network [1], to detect accurate and stable line features on images. Then we apply a geometric-based strategy to obtain precise 2D-3D line correspondences using epipolar constraint and reprojection filtering. A following point-line joint cost function is constructed to optimize the camera pose with the initial coarse pose from the pure point-based localization. Sufficient experiments are conducted on open datasets, i.e, line extractor on Wireframe and YorkUrban, localization performance on InLoc ducl and duc2, to confirm the effectiveness of our point-line joint pose optimization method. Jixiang Wan, Yishan Ping, Shuzhou Dong, Haikuan Ning, Jijunnan Li, Yandong Guo |
IROS | 9 |
| 2022 | CrossHuman: Learning Cross-guidance from Multi-frame Images for Human ReconstructionabstractWe propose CrossHuman, a novel method that learns cross-guidance from parametric human model and multi-frame RGB images to achieve high-quality 3D human reconstruction. To recover geometry details and texture even in invisible regions, we design a reconstruction pipeline combined with tracking-based methods and tracking-free methods. Given a monocular RGB sequence, we track the parametric human model in the whole sequence, the points (voxels) corresponding to the target frame are warped to reference frames by the parametric body motion. Guided by the geometry priors of the parametric body and spatially aligned features from RGB sequence, the robust implicit surface is fused. Moreover, a multi-frame transformer (MFT) and a self-supervised warp refinement module are integrated to the framework to relax the requirements of parametric body and help to deal with very loose cloth. Compared with previous works, our CrossHuman enables high-fidelity geometry details and texture in both visible and invisible regions and improves the accuracy of the human reconstruction even under estimated inaccurate parametric human models. The experiments demonstrate that our method achieves state-of-the-art (SOTA) performance. Liliang Chen, Jiaqi Li 0026, Han Huang 0005, Yandong Guo |
ACM Multimedia | 4 |
| 2022 | IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-trainingabstractVision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost. Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo |
ACM Multimedia | 9 |
| 2022 | Transductive Aesthetic Preference Propagation for Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (PIAA) aims at capturing individual aesthetic preference. Fine-tuning on personalized data has been proven to be effective in PIAA task. However, a fixed fine-tuning strategy may cause under/over-fitting on limited personal data and it also brings additional training cost. To alleviate these issues, we employ a meta learning-based Transductive Aesthetic Preference Propagation (TAPP-PIAA) algorithm under regression manner to substitute the fine-tuning strategy. Specifically, each user's data is regarded as a meta-task and spilt into support and query set. Then, we extract deep aesthetic features with a pre-trained generic image aesthetic assessment (GIAA) model. Next, we treat image features as graph nodes and their similarities as edge weights to construct an undirected nearest neighbor graph for inference. Instead of fine-tuning on support set, TAPP-PIAA propagates aesthetic preference from support to query set with a predefined propagation formula. Finally, to learn a generalizable aesthetic representation for various users, we optimize our TAPP-PIAA across different users with meta-learning framework. Experimental results indicate that our TAPP-PIAA can surpass the state-of-the-art methods on benchmark databases. Yuzhe Yang 0001, Huaxiong Li, Haoxing Chen, Liwu Xu, Leida Li, Yandong Guo |
ACM Multimedia | 8 |
| 2022 | Situational Perception Guided Image MattingabstractMost automatic matting methods try to separate the salient foreground from the background. However, the insufficient quantity and subjective bias of the current existing matting datasets make it difficult to fully explore the semantic association between object-to-object and object-to-environment in a given image. In this paper, we propose a Situational Perception Guided Image Matting (SPG-IM) method that mitigates subjective bias of matting annotations and captures sufficient situational perception information for better global saliency distilled from the visual-to-textual task. SPG-IM can better associate inter-objects and object-to-environment saliency, and compensate the subjective nature of image matting and its expensive annotation. We also introduce a textual Semantic Transformation (TST) module that can effectively transform and integrate the semantic feature stream to guide the visual representations. In addition, an Adaptive Focal Transformation (AFT) Refinement Network is proposed to adaptively switch multi-scale receptive fields and focal points to enhance both global and local details. Extensive experiments demonstrate the effectiveness of situational perception guidance from the visual-to-textual tasks on image matting, and our model outperforms the state-of-the-art methods. We also analyze the significance of different components in our model. Bo Xu 0031, Jiake Xie, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
ACM Multimedia | 7 |
| 2022 | FloRen: Real-time High-quality Human Performance Rendering via Appearance Flow Using Sparse RGB CamerasabstractWe propose FloRen, a novel system for real-time, high-resolution free-view human synthesis. Our system runs at 15fps in 1K resolution with very sparse RGB cameras. In FloRen, a coarse-level implicit geometry is recovered at first as initialization, and then processed by a neural rendering framework based on appearance flow. Our appearance flow-based rendering framework consists of three steps, namely view-dependent depth refinement, appearance flow estimation and occlusion-aware color rendering. In this way, we resolve the view synthesis problem in the image plane, where 2D convolutional neural networks can be efficiently applied, contributing to high speed performance. For robust appearance flow estimation, we explicitly combine data-driven human prior knowledge with multiview geometric constraints. The accurate appearance flow enables precise color mapping from input view to novel view, which greatly facilitates high-resolution novel view generation. We demonstrate that our system achieves state-of-the-art performance and even outperforms many offline methods. Ruizhi Shao, Liliang Chen, Zerong Zheng, Hongwen Zhang 0001, Yuxiang Zhang 0006, Han Huang 0005, Yandong Guo, Yebin Liu |
SIGGRAPH Asia | 7 |
| 2022 | Deep Two-Stream Video Inference for Human Body Pose and Shape EstimationabstractSeveral video-based 3D pose and shape estimation algorithms have been proposed to resolve the temporal inconsistency of single-image-based methods. However it still remains challenging to have stable and accurate reconstruction. In this paper, we propose a new framework Deep Two-Stream Video Inference for Human Body Pose and Shape Estimation (DTS-VIBE), to generate 3D human pose and mesh from RGB videos. We reformulate the task as a multi-modality problem that fuses RGB and optical flow for more reliable estimation. In order to fully utilize both sensory modalities (RGB or optical flow), we train a two-stream temporal network based on transformer to predict SMPL parameters. The supplementary modality, optical flow, helps to maintain temporal consistency by leveraging motion knowledge between two consecutive frames. The proposed algorithm is extensively evaluated on the Human3.6 and 3DPW datasets. The experimental results show that it outperforms other state-of-the-art methods by a significant margin. Bo Xu 0031, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
WACV | 5 |
| 2022 | SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu |
Comput. Graph. | 4 |
| 2021 | View-Guided Point Cloud CompletionabstractThis paper presents a view-guided solution for the task of point cloud completion. Unlike most existing methods directly inferring the missing points using shape priors, we address this task by introducing ViPC (view-guided point cloud completion) that takes the missing crucial global structure information from an extra single-view image. By leveraging a framework that sequentially performs effective cross-modality and cross-level fusions, our method achieves significantly superior results over typical existing solutions on a new large-scale dataset we collect for the view-guided point cloud completion task. Xuancheng Zhang, Yutong Feng, Siqi Li 0001, Changqing Zou, Hai Wan, Xibin Zhao, Yandong Guo, Yue Gao 0002 |
CVPR | 7 |
| 2021 | Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object InteractionabstractMost existing human matting algorithms tried to separate pure human-only foreground from the background. In this paper, we propose a Virtual Multi-modality Foreground Matting (VMFM) method to learn human-object interactive foreground (human and objects interacted with him or her) from a raw RGB image. The VMFM method requires no additional inputs, e.g. trimap or known background. We reformulate foreground matting as a self-supervised multi-modality problem: factor each input image into estimated depth map, segmentation mask, and interaction heatmap using three auto-encoders. In order to fully utilize the characteristics of each modality, we first train a dual encoder-to-decoder network to estimate the same alpha matte. Then we introduce a self-supervised method: Complementary Learning(CL) to predict deviation probability map and exchange reliable gradients across modalities without label. We conducted extensive experiments to analyze the effectiveness of each modality and the significance of different components in complementary learning. We demonstrate that our model outperforms the state-of-the-art methods. Bo Xu 0031, Han Huang 0005, Cheng Lu 0006, Yandong Guo |
ICCV | 5 |
| 2021 | Distance Restricted Transformer Encoder for Multi-Label ClassificationabstractMulti-label image classification is a fundamental but challenging task in Multimedia community. It aims to predict a set of labels presented in an image. Great progress has been made by exploring convolutional neural network with binary cross-entropy loss recently. However, conventional approaches are limited to highlight the key visual contents associated with target labels and pay little attention to confining the distances between visual and positive/negative label representations. To target these aspects, we firstly introduce a variant transformer encoder model for acquiring the underlying and crucial visual information related to ground truth labels. Specifically, a novel primal feature guided net is designed to maintain the original visual features during encoding process. Secondly, we exploit a distance restricted learning strategy in a common semantic space to shrink the distances of images with positive labels while expand with the negative ones during training stage. Extensive experiments are executed on MSCOCO and WIDER Attribute datasets and outstanding performance is achieved compared with other state-of-the-art models. Yandong Guo, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICME | 4 |
| 2021 | Retrieval and Localization with Observation ConstraintsabstractAccurate visual re-localization is very critical to many artificial intelligence applications, such as augmented reality, virtual reality, robotics and autonomous driving. To accomplish this task, we propose an integrated visual re-localization method called RLOCS by combining image retrieval, semantic consistency and geometry verification to achieve accurate estimations. The localization pipeline is designed as a coarse-to-fine paradigm. In the retrieval part, we cascade the architecture of ResNet101-GeM-ArcFace and employ DBSCAN followed by spatial verification to obtain a better initial coarse pose. We design a module called observation constraints, which combines geometry information and semantic consistency for filtering outliers. Comprehensive experiments are conducted on open datasets, including retrieval on R-Oxford5k and R-Paris6k, semantic segmentation on Cityscapes, localization on Aachen Day-Night and InLoc. By creatively modifying separate modules in the total pipeline, our method achieves many performance improvements on the challenging localization benchmarks. Huanhuan Fan, Jijunnan Li, Yandong Guo |
ICRA | 7 |
| 2021 | To See in the Dark: N2DGAN for Background Modeling in Nighttime SceneabstractDue to the deteriorated conditions of illumination lack and uneven lighting, the performance of traditional background modeling methods is greatly limited for the surveillance of nighttime video. To make background modeling under nighttime scene performs as well as in daytime condition, we put forward a promising generation-based background modeling framework for foreground surveillance. With a pre-specified daytime reference image as background frame, the GAN based generation model, called N2DGAN, is trained to transfer each frame of nighttime video to a virtual daytime image with the same scene to the reference image except for the foreground part. Specifically, to balance the preservation of background scene and the foreground object(s) in generating the virtual daytime image, we presented a two-pathway generation model, in which the global and local sub-networks were well combined with spatial and temporal consistency constraints. For the sequence of generated virtual daytime images, a multi-scale Bayes model was further proposed to characterize pertinently the temporal variation of background. We manually labeled ground truth on the collected nightime video datasets for performance evaluation. The impressive results illustrated in both the main paper and supplementary show the effectiveness of our proposed approach. Zhenfeng Zhu, Yingying Meng, Deqiang Kong, Xingxing Zhang 0001, Yandong Guo, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Discriminative Multi-Modality Speech RecognitionabstractVision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded to the multi-modality speech recognition (MSR). In this paper, we propose a two-stage speech recognition model. In the first stage, the target voice is separated from background noises with help from the corresponding visual information of lip movements, making the model ‘listen' clearly. At the second stage, the audio modality combines visual modality again to better understand the speech by a MSR sub-network, further improving the recognition rate. There are some other key contributions: we introduce a pseudo-3D residual convolution (P3D)-based visual front-end to extract more discriminative features; we upgrade the temporal convolution block from 1D ResNet with the temporal convolutional network (TCN), which is more suitable for the temporal tasks; the MSR sub-network is built on the top of Element-wise-Attention Gated Recurrent Unit (EleAtt-GRU), which is more effective than Transformer in long sequences. We conducted extensive experiments on the LRS3-TED and the LRW datasets. Our two-stage model (audio enhanced multi-modality speech recognition, AE-MSR) consistently achieves the state-of-the-art performance by a significant margin, which demonstrates the necessity and effectiveness of AE-MSR. Bo Xu 0031, Cheng Lu 0006, Yandong Guo, Jacob Wang |
CVPR | 3 |
| 2020 | Learning to Detect Head Movement in Unconstrained Remote Gaze Estimation in the WildabstractUnconstrained remote gaze estimation remains challenging mostly due to its vulnerability to the large variability in head-pose. Prior solutions struggle to maintain reliable accuracy in unconstrained remote gaze tracking. Among them, appearance-based solutions demonstrate tremendous potential in improving gaze accuracy. However, existing works still suffer from head movement and are not robust enough to handle real-world scenarios. Especially most of them study gaze estimation under controlled scenarios where the collected datasets often cover limited ranges of both head-pose and gaze which introduces further bias. In this paper, we propose novel end-to-end appearance-based gaze estimation methods that could more robustly incorporate different levels of head-pose representations into gaze estimation. Our method could generalize to real-world scenarios with low image quality, different lightings and scenarios where direct head-pose information is not available. To better demonstrate the advantage of our methods, we further propose a new benchmark dataset with the most rich distribution of head-gaze combination reflecting real-world scenarios. Extensive evaluations on several public datasets and our own dataset demonstrate that our method consistently outperforms the state-of-the-art by a significant margin. Zhecan Wang, Jian Zhao 0006, Cheng Lu 0006, Han Huang 0005, Fan Yang 0035, Lianji Li, Yandong Guo |
WACV | 7 |
| 2020 | Watch to Listen Clearly: Visual Speech Enhancement Driven Multi-modality Speech RecognitionabstractMulti-modality (talking face video and audio) information helps improve speech recognition performance compared to the single modality. In noisy environments, the effect of audio modality is weakened, which further affects the performance of multi-modality speech recognition (MSR). Most of the MSR methods use noisy audio signal as input of the audio modality without any enhancement (filtering the noisy components in the audio signal). In this paper, we propose an audio-enhanced multi-modality speech recognition model. In particular, the proposed model consists of two sub-networks, one is the visual speech enhancement (VE) sub-network and the other is the multi-modality speech recognition (MSR) sub-network. The VE sub-network is able to separate a speaker’s voice from background noises when given the corresponding talking face to enhance audio modality. Then the audio modality together with video modality are fed into the MSR sub-network to produce characters. We introduce a pseudo-3D residual network (P3D)based visual front-end to extract more advantageous visual features. The MSR sub-network is built on top of the Element-wise-Attention Gated Recurrent Unit (EleAttGRU) architecture which is more effective than Transformer in long sequences. We demonstrate the effectiveness of audio enhancement for MSR by extensive experiments. The proposed method surpasses the state-of-the-art MSR models on the LRS3-TED dataset and the LRW dataset. Bo Xu 0031, Jacob Wang, Cheng Lu 0006, Yandong Guo |
WACV | 4 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 26 |
| 2019 | Large Scale Incremental LearningabstractModern machine learning suffers from \textit{catastrophic forgetting} when learning new classes incrementally. The performance dramatically degrades due to the missing data of old classes. Incremental learning methods have been proposed to retain the knowledge acquired from the old classes, by using knowledge distilling and keeping a few exemplars from the old classes. However, these methods struggle to \textbf{scale up to a large number of classes}. We believe this is because of the combination of two factors: (a) the data imbalance between the old and new classes, and (b) the increasing number of visually similar classes. Distinguishing between an increasing number of visually similar classes is particularly challenging, when the training data is unbalanced. We propose a simple and effective method to address this data imbalance issue. We found that the last fully connected layer has a strong bias towards the new classes, and this bias can be corrected by a linear model. With two bias parameters, our method performs remarkably well on two large datasets: ImageNet (1000 classes) and MS-Celeb-1M (10000 classes), outperforming the state-of-the-art algorithms by 11.1\% and 13.2\% respectively. Yue Wu 0008, Yinpeng Chen, Yuancheng Ye, Zicheng Liu 0001, Yandong Guo, Yun Fu 0001 |
CVPR | 6 |
| 2019 | Deep class-skewed learning for face recognition
Pingyu Wang, Zhicheng Zhao 0001, Yandong Guo, Yanyun Zhao, Bojin Zhuang |
Neurocomputing | 4 |
| 2019 | Edge Heuristic GAN for Non-Uniform Blind DeblurringabstractNon-uniform blur, mainly caused by camera shake and motions of multiple objects, is one of the most common causes of image quality degradation. However, the traditional blind deblurring methods based on blur kernel estimation do not perform well on complicated non-uniform motion blurs. However, recent studies show that GAN-based approaches achieve impressive performance on deblurring tasks. In this letter, to further improve the performance of GAN-based methods on deblurring tasks, we propose an edge heuristic multi-scale generative adversarial network (GAN), which uses the coarse-to-fine scheme to restore clear images in an end-to-end manner. In particular, an edge-generated network is designed to generate sharp edges as auxiliary information to guide the deblurring process. Furthermore, We propose a hierarchical content loss function for deblurring tasks. Extensive experiments on different datasets show that our method achieves state-of-the-art performance in dynamic scene deblurring. Shuai Zheng 0005, Zhenfeng Zhu, Jian Cheng 0001, Yandong Guo, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | One-Shot Face Recognition via Generative LearningabstractOne-shot face recognition measures the ability to recognize persons with only seeing them once, which is a hallmark of human visual intelligence. It is challenging for existing machine learning approaches to mimic this way, since limited data cannot well represent the data variance. To this end, we propose to build a large-scale face recognizer, which is capable to fight off the data imbalance difficulty. To seek a more effective general classifier, we develop a novel generative model attempting to synthesize meaningful data for one-shot classes by adapting the data variances from other normal classes. Specifically, we formulate conditional generative adversarial networks and the general Softmax classifier into a unified framework. Such a two-player minimax optimization can guide the generation of more effective data, which benefit the classifier learning for one-shot classes. The experimental results on a large-scale face benchmark with 21K persons verify the effectiveness of our proposed algorithm in one-shot classification, as our generative model significantly improves the recognition coverage rate from 25:65% to 94:84% at the precision of 99% for the one-shot classes, while still keeps an overall Top-1 accuracy at 99:80% for the normal classes. Zhengming Ding, Yandong Guo, Yun Fu 0001 |
FG | 2 |
| 2017 | Model-Based Iterative Restoration for Binary Document Image Compression with Dictionary LearningabstractThe inherent noise in the observed (e.g., scanned) binary document image degrades the image quality and harms the compression ratio through breaking the pattern repentance and adding entropy to the document images. In this paper, we design a cost function in Bayesian framework with dictionary learning. Minimizing our cost function produces a restored image which has better quality than that of the observed noisy image, and a dictionary for representing and encoding the image. After the restoration, we use this dictionary (from the same cost function) to encode the restored image following the symbol-dictionary framework by JBIG2 standard with the lossless mode. Experimental results with a variety of document images demonstrate that our method improves the image quality compared with the observed image, and simultaneously improves the compression ratio. For the test images with synthetic noise, our method reduces the number of flipped pixels by 48.2% and improves the compression ratio by 36.36% as compared with the best encoding methods. For the test images with real noise, our method visually improves the image quality, and outperforms the cutting-edge method by 28.27% in terms of the compression ratio. Yandong Guo, Cheng Lu 0006, Jan P. Allebach, Charles A. Bouman |
CVPR | 1 |
| 2016 | MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition
Yandong Guo, Lei Zhang 0001, Yuxiao Hu 0001, Xiaodong He 0001, Jianfeng Gao 0001 |
ECCV (3) | 1 |
| 2015 | Image quality evaluation using image quality ruler and graphical modelabstractQuantifying image quality through subjective evaluation is very critical to image quality evaluation. Using the image quality ruler method, an average score per stimulus can be easily obtained in the unit of Just Noticeable Differences (JNDs). However, it requires a large number of subjects, since pure averaging does not consider the different judging quality of different subjects. In this paper, we propose an image quality evaluation framework using the image quality ruler method with a statistical model. By incorporating this model, we consider the quality score, the expertise of the subjects, and the difficulty of image rating task as three hidden variables. Then we use expectation-maximization (EM) to estimate these hidden variables. From our experimental results, we show that our method provides reliable results without using a large number of subjects. Preliminary results also demonstrate that the estimates of the parameters can guide us to better distribute the valuable human resources used to conduct psychophysical experiments. Weibao Wang, Jan P. Allebach, Yandong Guo |
ICIP | 3 |
| 2013 | Dynamic hierarchical dictionary design for multi-page binary document image compressionabstractThe JBIG2 standard is widely used for binary document image compression primarily because it achieves much higher compression ratios than conventional facsimile encoding standards. In this paper, we propose a dynamic hierarchical dictionary design method (DH) for multi-page binary document image compression with JBIG2. Our DH method outperforms other methods for multi-page compression by utilizing the information redundancy among pages with the following technologies. First, we build a hierarchical dictionary to keep more information per page for future usage. Second, we dynamically update the dictionary in memory to keep as much information as possible subject to the memory constraint. Third, we incorporate our conditional entropy estimation algorithm to utilize the saved information more effectively. Our experimental results show that the compression ratio improvement by our DH method is about 15% compared to the best existing multi-page encoding method. Yandong Guo, Dejan Depalov, Peter Bauer, Brent M. Bradburn, Jan P. Allebach, Charles A. Bouman |
ICIP | 1 |
| 2013 | Message passing with l1 penalized KL minimizationabstractBayesian inference is often hampered by large computational expense. As a generalization of belief propagation (BP), expectation propagation (EP) approximates exact Bayesian computation with efficient message passing updates. However, when an approximation family used by EP is far from exact posterior distributions, message passing may lead to poor approximation quality and suffer from divergence. To address this issue, we propose an approximate inference method, relaxed expectation propagation(REP), based on a new divergence with a l1 penalty. Minimizing this penalized divergence adaptively relaxes EP’s moment matching requirement for message passing. We apply REP to Gaussian process classification and experimental results demonstrate significant improvement of REP over EP and alpha-divergence based power EP – in terms of algorithmic stability, estimation accuracy, and predictive performance. Furthermore, we develop relaxed belief propagation(RBP), a special case of REP, to conduct inference on discrete Markov random fields (MRFs). Our results show improved estimation accuracy of RBP over BP and fractional BP when interactions between MRF nodes are strong. Yandong Guo |
ICML (3) | 2 |
| 2010 | High dimensional regression using the sparse matrix transform (SMT)abstractRegression from high dimensional observation vectors is particularly difficult when training data is limited. More specifically, if the number of sample vectors n is less than dimension of the sample vectors p, then accurate regression is difficult to perform without prior knowledge of the data covariance. In this paper, we propose a novel approach to high dimensional regression for application when n ≪ p. The approach works by first decorrelating the high dimensional observation vector using the sparse matrix transform (SMT) estimate of the data covariance. Then the decorrelated observations are used in a regularized regression procedure such as Lasso or shrinkage. Numerical results demonstrate that the proposed regression approach can significantly improve the prediction accuracy, especially when n is small and the signal to be predicted lies in the subspace of the observations corresponding to the small eigenvalues. Guangzhi Cao, Yandong Guo, Charles A. Bouman |
ICASSP | 2 |