EDBT 2026 Demo / reviewers in the wild / expert
Yue Liao
dblp:210/0033
· DBLP profile ↗
40ranked-venue papers
9as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 2.5 GHz GaN multiple quantum well micro-photodetector for high-speed visible light communication
Yue Liao, Runze Lin, Xugao Cui |
Sci. China Inf. Sci. | 1 |
| 2026 | MC#: Mixture Compressor for Mixture-of-Experts Large ModelsabstractMixture-of-Experts (MoE) has emerged as an effective and efficient scaling mechanism for large language models (LLMs) and vision-language models (VLMs). By expanding a single feed-forward network into multiple expert branches, MoE increases model capacity while maintaining efficiency through sparse activation. However, despite this sparsity, the need to preload all experts into memory and activate multiple experts per input introduces significant computational and memory overhead. The expert module becomes the dominant contributor to model size and inference cost, posing a major challenge for deployment. To address this, we propose MC# (Mixture-Compressor-sharp), a unified framework that combines static quantization and dynamic expert pruning by leveraging the significance of both experts and tokens to achieve aggressive compression of MoE-LLMs/VLMs. To reduce storage and loading overhead, we introduce Pre-Loading Mixed-Precision Quantization (PMQ), which formulates adaptive bit allocation as a linear programming problem. The objective function jointly considers expert importance and quantization error, producing a Pareto-optimal trade-off between model size and performance. To reduce runtime computation, we further introduce Online Top-any Pruning (OTP), which models expert activation per token as a learnable distribution via Gumbel-Softmax sampling. During inference, OTP dynamically selects a subset of experts for each token, allowing fine-grained control over activation. By combining PMQ's static bit-width optimization with OTP's dynamic routing, MC# achieves extreme compression with minimal accuracy degradation. On DeepSeek-VL2, MC# achieves a 6.2 × weight reduction at an average of 2.57 bits, with only a 1.7% drop across five multimodal benchmarks compared to the 16-bit baseline. Moreover, OTP further reduces expert activation by 20% with less than 1% performance loss, demonstrating strong potential for efficient deployment of MoE-based models. Wei Huang 0042, Yue Liao, Yukang Chen, Haoru Tan, Si Liu 0001, Shuicheng Yan, Xiaojuan Qi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame SelectionabstractThe advancement of Large Vision Language Models (LVLMs) has significantly improved multimodal understanding, yet challenges remain in video reasoning tasks due to the scarcity of high-quality, large-scale datasets. Existing video question-answering (VideoQA) datasets often rely on costly manual annotations with insufficient granularity or automatic construction methods with redundant frame-by-frame analysis, limiting their scalability and effectiveness for complex reasoning. To address these challenges, we introduce VideoEspresso, a novel dataset that features VideoQA pairs preserving essential spatial details and temporal coherence, along with multimodal annotations of intermediate reasoning steps. Our construction pipeline employs a semantic-aware method to reduce redundancy, followed by generating QA pairs using GPT-4o. We further develop video Chain-of-Thought (CoT) annotations to enrich reasoning processes, guiding GPT-4o in extracting logical relationships from QA pairs and video content. To exploit the potential of high-quality VideoQA pairs, we propose a Hybrid LVLMs Collaboration framework, featuring a Frame Selector and a two-stage instruction fine-tuned reasoning LVLM. This framework adaptively selects core frames and performs CoT reasoning using multimodal evidence. Evaluated on our proposed benchmark with 14 tasks against 9 popular LVLMs, our method outperforms existing baselines on most tasks, demonstrating superior video reasoning capabilities. Our code and dataset have been released at: https://github.com/hshjerry/VideoEspresso Songhao Han, Wei Huang 0042, Hairong Shi, Le Zhuo, Xiu Su, Xiaojuan Qi 0001, Yue Liao, Si Liu 0001 |
CVPR | 9 |
| 2025 | Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMsabstractPreference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly target hallucination factors, they overlook the factors essential for multi-modal comprehension capabilities, often narrowing their improvements on hallucination mitigation. To bridge this gap, we propose Instruction-oriented Preference Alignment (IPA), a scalable framework designed to automatically construct alignment preferences grounded in instruction fulfillment efficacy. Our method involves an automated preference construction coupled with a dedicated verification process that identifies instruction-oriented factors, avoiding significant variability in response representations. Additionally, IPA incorporates a progressive preference collection pipeline, further recalling challenging samples through model self-evolution and reference-guided refinement. Experiments conducted on Qwen2VL-7B demonstrate IPA's effectiveness across multiple benchmarks, including hallucination evaluation, visual question answering, and text understanding tasks, highlighting its capability to enhance general comprehension. Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Si Liu 0001 |
ICCV | 2 |
| 2025 | From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection TuningabstractRecent text-to-image diffusion models achieve impressive visual quality through extensive scaling of training data and model parameters, yet they often struggle with complex scenes and fine-grained details. Inspired by the self-reflection capabilities emergent in large language models, we propose ReflectionFlow, an inference-time framework enabling diffusion models to iteratively reflect upon and refine their outputs. ReflectionFlow introduces three complementary inference-time scaling axes: (1) noise-level scaling to optimize latent initialization; (2) prompt-level scaling for precise semantic guidance; and most notably, (3) reflection-level scaling, which explicitly provides actionable reflections to iteratively assess and correct previous generations. To facilitate reflection-level scaling, we construct GenRef, a large-scale dataset comprising 1 million triplets, each containing a reflection, a flawed image, and an enhanced image. Leveraging this dataset, we efficiently perform reflection tuning on state-of-the-art diffusion transformer, FLUX.1-dev, by jointly modeling multimodal inputs within a unified framework. Experimental results show that ReflectionFlow significantly outperforms naive noise-level scaling methods, offering a scalable and compute-efficient solution toward higher-quality image synthesis on challenging tasks. Le Zhuo, Liangbing Zhao, Sayak Paul, Yue Liao, Renrui Zhang, Yi Xin 0003, Peng Gao 0007, Mohamed Elhoseiny 0001, Hongsheng Li 0001 |
ICCV | 4 |
| 2025 | Mixture Compressor for Mixture-of-Experts LLMs Gains MoreabstractMixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant, as many tokens may only require a single expert. Motivated by these issues, we investigate the MoE-LLMs and make two key observations: a) different experts exhibit varying behaviors on activation reconstruction error, routing scores, and activated frequencies, highlighting their differing importance, and b) not all tokens are equally important-- only a small subset is critical. Building on these insights, we propose MC, a training-free Mixture-Compressor for MoE-LLMs, which leverages the significance of both experts and tokens to achieve an extreme compression. First, to mitigate storage and loading overheads, we introduce Pre-Loading Mixed-Precision Quantization (PMQ), which formulates the adaptive bit-width allocation as a Linear Programming (LP) problem, where the objective function balances multi-factors reflecting the importance of each expert. Additionally, we develop Online Dynamic Pruning (ODP), which identifies important tokens to retain and dynamically select activated experts for other tokens during inference to optimize efficiency while maintaining performance. Our MC integrates static quantization and dynamic pruning to collaboratively achieve extreme compression for MoE-LLMs with less accuracy loss, ensuring an optimal trade-off between performance and efficiency Extensive experiments confirm the effectiveness of our approach. For instance, at 2.54 bits, MC compresses 76.6% of the model, with only a 3.8% average accuracy loss. During dynamic inference, we further reduce activated parameters by 15%, with a performance drop of less than 0.6%. Remarkably, MC even surpasses floating-point 13b dense LLMs with significantly smaller parameter sizes, suggesting that mixture compression in MoE-LLMs has the potential to outperform both comparable and larger dense LLMs. Our code is
available at https://github.com/Aaronhuang-778/MC-MoE Wei Huang 0042, Yue Liao, Ruifei He, Haoru Tan, Hongsheng Li 0001, Si Liu 0001, Xiaojuan Qi 0001 |
ICLR | 2 |
| 2025 | LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationabstractWe introduce LLaVA-MoD, a novel framework designed to enable the efficient training of small-scale Multimodal Language Models ($s$-MLLM) distilling knowledge from large-scale MLLM ($l$-MLLM). Our approach tackles two fundamental challenges in MLLM distillation. First, we optimize the network structure of $s$-MLLM by integrating a sparse Mixture of Experts (MoE) architecture into the language model, striking a balance between computational efficiency and model expressiveness. Second, we propose a progressive knowledge transfer strategy for comprehensive knowledge transfer. This strategy begins with mimic distillation, where we minimize the Kullback-Leibler (KL) divergence between output distributions to enable $s$-MLLM to emulate $s$-MLLM's understanding. Following this, we introduce preference distillation via Preference Optimization (PO), where the key lies in treating $l$-MLLM as the reference model. During this phase, the $s$-MLLM's ability to discriminate between superior and inferior examples is significantly enhanced beyond $l$-MLLM, leading to a better $s$-MLLM that surpasses $l$-MLLM, particularly in hallucination benchmarks.
Extensive experiments demonstrate that LLaVA-MoD surpasses existing works across various benchmarks while maintaining a minimal activated parameters and low computational costs. Remarkably, LLaVA-MoD-2B surpasses Qwen-VL-Chat-7B with an average gain of 8.8\%, using merely $0.3\%$ of the training data and 23\% trainable parameters. The results underscore LLaVA-MoD's ability to effectively distill comprehensive knowledge from its teacher model, paving the way for developing efficient MLLMs. Fangxun Shu, Yue Liao, Lei Zhang 0006, Le Zhuo, Chenning Xu, Long Chan, Zhelun Yu, Wanggui He, Siming Fu, Haoyuan Li 0002, Si Liu 0001, Hongsheng Li 0001, Hao Jiang 0062 |
ICLR | 2 |
| 2025 | Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and MethodologyabstractDeveloping agents capable of navigating to a target location based on language instructions and visual information, known as vision-language navigation (VLN), has attracted widespread interest. Most research has focused on ground-based agents, while UAV-based VLN remains relatively underexplored. Recent efforts in UAV vision-language navigation predominantly adopt ground-based VLN settings, relying on predefined discrete action spaces and neglecting the inherent disparities in agent movement dynamics and the complexity of navigation tasks between ground and aerial environments. To address these disparities and challenges, we propose solutions from three perspectives: platform, benchmark, and methodology. To enable realistic UAV trajectory simulation in VLN tasks, we propose the OpenUAV platform, which features diverse environments, realistic flight control, and extensive algorithmic support. We further construct a target-oriented VLN dataset consisting of approximately 12k trajectories on this platform, serving as the first dataset specifically designed for realistic UAV VLN tasks. To tackle the challenges posed by complex aerial environments, we propose an assistant-guided UAV object search benchmark called UAV-Need-Help, which provides varying levels of guidance information to help UAVs better accomplish realistic VLN tasks. We also propose a UAV navigation LLM that, given multi-view images, task descriptions, and assistant instructions, leverages the multimodal understanding capabilities of the MLLM to jointly process visual and textual information, and performs hierarchical trajectory generation. The evaluation results of our method significantly outperform the baseline models, while there remains a considerable gap between our results and those achieved by human operators, underscoring the challenge presented by the UAV-Need-Help task. Donglin Yang, Ziqin Wang, Hohin Kwan, Hongsheng Li 0001, Yue Liao, Si Liu 0001 |
ICLR | 8 |
| 2025 | Pask: Providing Answer before AsKing toward Proactive AI agentabstractWe present Pask, a proactive AI agent that provides real-time, context-aware guidance and knowledge support in audio-centric media environments. Unlike passive assistants that follow the ''you ask, I answer'' model, Pask shifts toward ''answering before asking'' by continuously monitoring live audio, anticipating user needs, and proactively offering conceptual explanations and semantic clarifications. It integrates three core components: a silent copilot for in-situ explanation, a structured knowledge base for factual grounding, and a private memory module for personalized adaptation. Pask enhances comprehension and communication in scenarios such as online learning, media consumption, and live meetings through sustained, intelligent guidance. A live demo is available at https://www.youtube.com/watch?v=ki_CKiV9Oyk. Zhifei Xie, Hu Zongzheng, Guibin Zhang, Yue Liao, Chunyan Miao, Shuicheng Yan |
ACM Multimedia | 5 |
| 2025 | RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation EvaluationabstractRecent advances in vision-language models (VLMs) have enabled instruction-conditioned robotic systems with improved generalization. However, most existing work focuses on reactive System 1 policies, underutilizing VLMs’ strengths in semantic reasoning and long-horizon planning. These System 2 capabilities—characterized by deliberative, goal-directed thinking—remain underexplored due to the limited temporal scale and structural complexity of current benchmarks. To address this gap, we introduce RoboCerebra, a benchmark for evaluating high-level reasoning in long-horizon robotic manipulation. RoboCerebra includes: (1) a large-scale simulation dataset with extended task horizons and diverse subtask sequences in household environments; (2) a hierarchical framework combining a high-level VLM planner with a low-level vision-language-action (VLA) controller; and (3) an evaluation protocol targeting planning, reflection, and memory through structured System 1–System 2 interaction. The dataset is constructed via a top-down pipeline, where GPT generates task instructions and decomposes them into subtask sequences. Human operators execute the subtasks in simulation, yielding high-quality trajectories with dynamic object variations. Compared to prior benchmarks, RoboCerebra features significantly longer action sequences and denser annotations. We further benchmark state-of-the-art VLMs as System 2 modules and analyze their performance across key cognitive dimensions, advancing the development of more capable and generalizable robotic planners. Songhao Han, Boxiang Qiu, Yue Liao, Siyuan Huang 0004, Chen Gao 0005, Shuicheng Yan, Si Liu 0001 |
NeurIPS | 3 |
| 2025 | EnerVerse: Envisioning Embodied Future Space for Robotics ManipulationabstractWe introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page. Siyuan Huang 0004, Liliang Chen, Shengcong Chen, Yue Liao, Zhengkai Jiang 0001, Peng Gao 0007, Hongsheng Li 0001, Maoqing Yao, Guanghui Ren |
NeurIPS | 5 |
| 2025 | Nyström-Accelerated Primal LS-SVMs: Breaking the $O(an^3)$ Complexity Bottleneck for Scalable ODEs LearningabstractA major problem of kernel-based methods (e.g., least squares support vector machines, LS-SVMs) for solving linear/nonlinear ordinary differential equations (ODEs) is the prohibitive $O(an^3)$ ($a=1$ for linear ODEs and 27 for nonlinear ODEs) part of their computational complexity with increasing temporal discretization points $n$. We propose a novel Nyström-accelerated LS-SVMs framework that breaks this bottleneck by reformulating ODEs as primal-space constraints. Specifically, we derive for the first time an explicit Nyström-based mapping and its derivatives from one-dimensional temporal discretization points to a higher $m$-dimensional feature space ($1< m\le n$), enabling the learning process to solve linear/nonlinear equation systems with $m$-dependent complexity. Numerical experiments on sixteen benchmark ODEs demonstrate: 1) $10-6000$ times faster computation than classical LS-SVMs and physics-informed neural networks (PINNs), 2) comparable accuracy to LS-SVMs ($<0.13\%$ relative MAE, RMSE, and $\left \| y-\hat{y} \right \| _{\infty } $difference) while maximum surpassing PINNs by 72\% in RMSE, and 3) scalability to $n=10^4$ time steps with $m=50$ features. This work establishes a new paradigm for efficient kernel-based ODEs learning without significantly sacrificing the accuracy of the solution. Weikuo Wang, Yue Liao |
NeurIPS | 2 |
| 2025 | UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation LearningabstractUnmanned Aerial Vehicles (UAVs) are evolving into language-interactive platforms, enabling more intuitive forms of human-drone interaction. While prior works have primarily focused on high-level planning and long-horizon navigation, we shift attention to language-guided fine-grained trajectory control, where UAVs execute short-range, reactive flight behaviors in response to language instructions. We formalize this problem as the Flying-on-a-Word (Flow) task and introduce UAV imitation learning as an effective approach. In this framework, UAVs learn fine-grained control policies by mimicking expert pilot trajectories paired with atomic language instructions. To support this paradigm, we present UAV-Flow, the first real-world benchmark for language-conditioned, fine-grained UAV control. It includes a task formulation, a large-scale dataset collected in diverse environments, a deployable control framework, and a simulation suite for systematic evaluation. Our design enables UAVs to closely imitate the precise, expert-level flight trajectories of human pilots and supports direct deployment without sim-to-real gap. We conduct extensive experiments on UAV-Flow, benchmarking VLN and VLA paradigms. Results show that VLA models are superior to VLN baselines and highlight the critical role of spatial grounding in the fine-grained Flow setting. Donglin Yang, Yue Liao, Hongsheng Li 0001, Si Liu 0001 |
NeurIPS | 3 |
| 2025 | Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor RegressionabstractIn this paper, we focus on the challenging task of monocular 3D lane detection. Previous methods typically adopt inverse perspective mapping (IPM) to transform the Front-Viewed (FV) images or features into the Bird-Eye-Viewed (BEV) space for lane detection. However, IPM's dependence on flat ground assumption and context information loss in BEV representations lead to inaccurate 3D information estimation. Though efforts have been made to bypass BEV and directly predict 3D lanes from FV representations, their performances still fall behind BEV-based methods due to a lack of structured modeling of 3D lanes. In this paper, we propose a novel BEV-free method named Anchor3DLane++ which defines 3D lane anchors as structural representations and makes predictions directly from FV features. We also design a Prototype-based Adaptive Anchor Generation (PAAG) module to generate sample-adaptive sparse 3D anchors dynamically. In addition, an Equal-Width (EW) loss is developed to leverage the parallel property of lanes for regularization. Furthermore, camera-LiDAR fusion is also explored based on Anchor3DLane++ to leverage complementary information. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane++ outperforms previous state-of-the-art methods. Code is available at: https://github.com/tusen-ai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Yue Liao, Jizhong Han, Naiyan Wang, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Toward Faithful Scene-Adaptive Knowledge for Semantic Segmentation of Remote Sensing ImagesabstractSemantic segmentation is a critical procedure in remote sensing image analysis that backs up various applications. High resolution remote sensing images contain a wealth of ground object features, which are organized into various describable scenes. The visual content offset towards each scene is an intuitive sense for understanding geospatial objects. However, existing semantic segmentation methods for remote sensing images generally neglect this intuition and lack the ability to adjust their perception preference on different images. To address this problem, we propose a paradigm for collecting scene information and dynamically adjusting the model inference process to be scene-aware. Specifically, our method leverages the class feature from the image to enhance the fixed class representation from the model. The interaction of these information is facilitated by a neighbor-friendly embedding space, making it more faithful to associate the image features and model parameters. For the model to better understand the complex scenes, a manifold mixup method is proposed to expand the effective embedding space on intra-class and inter-class regions, forcing the model to challenge the ambiguous instances in remote sensing images. Extensive experiments on four publicly available datasets demonstrated that our proposed improved the accuracy of semantic segmentation models on remote sensing images, overcoming the state-of-the-art methods. Yue Liao, Wei He 0003, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
Penghui Du, Yifan Sun 0003, Luting Wang 0001, Yue Liao, Errui Ding, Yan Wang 0059, Jingdong Wang 0001, Si Liu 0001 |
ECCV (23) | 5 |
| 2024 | Mask-Enhanced Segment Anything Model for Tumor Lesion Semantic Segmentation
Hairong Shi, Songhao Han, Shaofei Huang 0001, Yue Liao, Guanbin Li, Xiangxing Kong, Xiaomu Wang, Si Liu 0001 |
MICCAI (8) | 4 |
| 2024 | PPDM++: Parallel Point Detection and Matching for Fast and Accurate HOI DetectionabstractHuman-Object Interaction (HOI) detection aims to understand human activities by detecting interaction triplets. Previous HOI detection methods adopt a two-stage instance-driven paradigm. Unfortunately, many non-interactive human-object pairs generated by the first stage are the main obstacle impeding HOI detectors from high efficiency and promising performance. To remedy this, we propose a novel top-down interaction-driven paradigm, detecting interactions first and bridging interactive human-object pairs through interactions. We formulate HOI as a point triplet human point, interaction point, object point and design a Parallel Point Detection and Matching (PPDM) framework. We further take advantage of two-stage methods and propose a novel framework, PPDM++, that detects the interactive human-object pairs by PPDM, then extracts region features for each pair to predict actions. The core of PPDM/PPDM++ is to convert the instance-driven bottom-up paradigm to an interaction-driven top-down paradigm, thus avoiding additional computation costs from traversing a tremendous number of non-interactive pairs. Benefiting from the advanced paradigm, PPDM/PPDM++ has achieved significant performance gains with high efficiency. PPDM-DLA-34 has achieved 19.94 mAP with 42 FPS as the first real-time HOI detector, and PPDM++-SwinB achieves 30.1 mAP with 17 FPS on HICO-DET dataset. We also built an application-oriented database named HOI-A, a supplement to the existing datasets. Yue Liao, Si Liu 0001, Yulu Gao, Aixi Zhang, Fei Wang 0032, Bo Li 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | MAC: Masked Contrastive Pre-Training for Efficient Video-Text RetrievalabstractWe present a simple yet effective end-to-end Video-language Pre-training (VidLP) framework, Masked Contrastive Video-language Pre-training (MAC), for video-text retrieval tasks. Our MAC aims to reduce video representation's spatial and temporal redundancy in the VidLP model by a mask sampling mechanism to improve pre-training efficiency. Comparing conventional temporal sparse sampling, we propose to randomly mask a high ratio of spatial regions and only take visible regions into the encoder as sparse spatial sampling. Similarly, we adopt the mask sampling technique for text inputs for consistency. Instead of blindly applying the mask-then-prediction paradigm from MAE, we propose a masked-then-alignment paradigm for efficient video-text alignment. The motivation is that video-text retrieval tasks rely on high-level alignment rather than low-level reconstruction, and multimodal alignment with masked modeling encourages the model to learn a robust and general multimodal representation from incomplete and unstable inputs. Coupling these designs enables efficient end-to-end pre-training: 3× speed up, 60%+ computation reduction, and 4%+ performance improvement. Our MAC achieves state-of-the-art results on various video-text retrieval datasets including MSR-VTT, DiDeMo, and ActivityNet. Our approach is omnivorous to input modalities. With minimal modifications, we achieve competitive results on image-text retrieval tasks. Fangxun Shu, Biaolong Chen, Yue Liao, Jinqiao Wang, Si Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Object-Aware Distillation Pyramid for Open-Vocabulary Object DetectionabstractOpen-vocabulary object detection aims to provide object detectors trained on a fixed set of object categories with the generalizability to detect objects described by arbitrary text queries. Previous methods adopt knowledge distillation to extract knowledge from Pretrained Vision-and-Language Models (PVLMs) and transfer it to detectors. However, due to the non-adaptive proposal cropping and single-level feature mimicking processes, they suffer from information destruction during knowledge extraction and inefficient knowledge transfer. To remedy these limitations, we propose an Object-Aware Distillation Pyramid (OADP) framework, including an Object-Aware Knowledge Extraction (OAKE) module and a Distillation Pyramid (DP) mechanism. When extracting object knowledge from PVLMs, the former adaptively transforms object proposals and adopts object-aware mask attention to obtain precise and complete knowledge of objects. The latter introduces global and block distillation for more comprehensive knowledge transfer to compensate for the missing relation information in object distillation. Extensive experiments show that our method achieves significant improvement compared to current methods. Especially on the MS-COCO dataset, our OADP framework reaches 35.6 mAPN50, surpassing the current state-of-the-art method by 3.3 mAPN50. Code is released at https://github.com/LutingWang/OADP. Luting Wang 0001, Yi Liu 0070, Penghui Du, Yue Liao, Qiaosong Qi, Biaolong Chen, Si Liu 0001 |
CVPR | 5 |
| 2023 | Video Background Music Generation: Dataset, Method and EvaluationabstractMusic is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it requires music-video datasets, efficient architectures for video-to-music generation, and reasonable metrics, none of which currently exist. To close this gap, we introduce a complete recipe including dataset, benchmark model, and evaluation metric for video background music generation. We present SymMV, a video and symbolic music dataset with various musical annotations. To the best of our knowledge, it is the first video-music dataset with rich musical annotations. We also propose a benchmark video background music generation framework named V-MusProd, which utilizes music priors of chords, melody, and accompaniment along with video-music relations of semantic, color, and motion features. To address the lack of objective metrics for video-music correspondence, we design a retrieval-based metric VMCP built upon a powerful video-music representation learning model. Experiments show that with our dataset, V-MusProd outperforms the state-of-the-art method in both music quality and correspondence with videos. We believe our dataset, benchmark model, and evaluation metric will boost the development of video background music generation. Our dataset and code are available at https://github.com/zhuole1025/SymMV. Le Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao, Chenxi Bao, Stanley Peng, Songhao Han, Aixi Zhang, Fei Fang 0002, Si Liu 0001 |
ICCV | 4 |
| 2023 | DiffDance: Cascaded Human Motion Diffusion Model for Dance GenerationabstractWhen hearing music, it is natural for people to dance to its rhythm. Automatic dance generation, however, is a challenging task due to the physical constraints of human motion and rhythmic alignment with target music. Conventional autoregressive methods introduce compounding errors during sampling and struggle to capture the long-term structure of dance sequences. To address these limitations, we present a novel cascaded motion diffusion model, DiffDance, designed for high-resolution, long-form dance generation. This model comprises a music-to-dance diffusion model and a sequence super-resolution diffusion model. To bridge the gap between music and motion for conditional generation, DiffDance employs a pretrained audio representation learning model to extract music embeddings and further align its embedding space to motion via contrastive loss. During training our cascaded diffusion model, we also incorporate multiple geometric losses to constrain the model outputs to be physically plausible and add a dynamic loss weight that adaptively changes over diffusion timesteps to facilitate sample diversity. Through comprehensive experiments performed on the benchmark dataset AIST++, we demonstrate that DiffDance is capable of generating realistic dance sequences that align effectively with the input music. These results are comparable to those achieved by state-of-the-art autoregressive methods. Qiaosong Qi, Le Zhuo, Aixi Zhang, Yue Liao, Fei Fang 0002, Si Liu 0001, Shuicheng Yan |
ACM Multimedia | 4 |
| 2023 | Simultaneously Training and Compressing Vision-and-Language Pre-Training ModelabstractModel compression is an essential step for large-scale pre-training models toward practical application and deployment on the edge device. However, when conventional compression methods following ‘pre-training then compressing’ two-phase pipeline are applied to Vision-and-Language Pre-training (VLP) models, it will lead to a high calculation and memory overhead. In this work, we break the two-phase pipeline and propose an efficient and effective one-phase VLP model compression mechanism, namedREDUCER, which stands for ‘simultaneously training and compREssing’ VLP model via progressive moDUle replaCing and nEtworkRewiring. Specifically, REDUCER consists of three insightful designs. Firstly, we design a one-phase compression framework to train and compress the VLP model simultaneously to avoid the extra calculation and memory cost caused by an isolated model compression phase in the conventional two-phase pipeline. Secondly, we propose an adaptive progressive module replacing mechanism to compress the model depth free from explicit knowledge distillation losses, relieving the multi-task optimization problems. Thirdly, we integrate pruning techniques into VLP model compression to simultaneously compress the model in width and depth. Overall, we obtain a lightweight VLP model with only one pre-training phase, and it is the first one-phase compression method for VLP models. Extensive experiments have been conducted on representative VLP models,i.e., ClipBERT and VICTOR, and the experimental results show a superior trade-off between performance and efficiency. Qiaosong Qi, Aixi Zhang, Yue Liao, Wenyu Sun, Si Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionabstractThe task of Human-Object Interaction (HOI) detection could be divided into two core problems, i.e., human-object association and interaction understanding. In this paper, we reveal and address the disadvantages of the conventional query-driven HOI detectors from the two aspects. For the association, previous two-branch methods suffer from complex and costly post-matching, while single-branch methods ignore the features distinction in different tasks. We propose Guided-Embedding Network (GEN) to attain a two-branch pipeline without post-matching. In GEN, we design an instance decoder to detect humans and objects with two independent query sets and a position Guided Embedding (p-GE) to mark the human and object in the same position as a pair. Besides, we design an interaction decoder to classify interactions, where the interaction queries are made of instance Guided Embeddings (i-GE) generated from the outputs of each instance decoder layer. For the interaction understanding, previous methods suffer from long-tailed distribution and zero-shot discovery. This paper proposes Visual-Linguistic Knowledge Transfer (VLKT) training strategy to enhance interaction understanding by transferring knowledge from a visual-linguistic pre-trained model CLIP. In specific, we extract text embeddings for all labels with CLIP to initialize the classifier and adopt a mimic loss to minimize the visual feature distance between GEN and CLIP. As a result, GEN-VLKT outperforms the state of the art by large margins on multiple datasets, e.g., +5.05 mAP on HICO-Det. The source codes are available at https://github.com/YueLiao/gen-vlkt. Yue Liao, Aixi Zhang, Miao Lu, Si Liu 0001 |
CVPR | 1 |
| 2022 | HEAD: HEtero-Assists Distillation for Heterogeneous Object Detectors
Luting Wang 0001, Yue Liao, Zeren Jiang, Jianlong Wu, Fei Wang 0032, Chen Qian 0006, Si Liu 0001 |
ECCV (9) | 3 |
| 2022 | Human-Centric Relation Segmentation: Dataset and SolutionabstractVision and language understanding techniques have achieved remarkable progress, but currently it is still difficult to well handle problems involving very fine-grained details. For example, when the robot is told to "bring me the book in the girl's left hand", most existing methods would fail if the girl holds one book respectively in her left and right hand. In this work, we introduce a new task named human-centric relation segmentation (HRS), as a fine-grained case of HOI-det. HRS aims to predict the relations between the human and surrounding entities and identify the relation-correlated human parts, which are represented as pixel-level masks. For the above exemplar case, our HRS task produces results in the form of relation triplets 〈girl [left hand], hold, book 〉 and exacts segmentation masks of the book, with which the robot can easily accomplish the grabbing task. Correspondingly, we collect a new Person In Context (PIC) dataset for this new task, which contains 17,122 high-resolution images and densely annotated entity segmentation and relations, including 141 object categories, 23 relation categories and 25 semantic human parts. We also propose a Simultaneous Matching and Segmentation (SMS) framework as a solution to the HRS task. It contains three parallel branches for entity segmentation, subject object matching and human parsing respectively. Specifically, the entity segmentation branch obtains entity masks by dynamically-generated conditional convolutions; the subject object matching branch detects the existence of any relations, links the corresponding subjects and objects by displacement estimation and classifies the interacted human parts; and the human parsing branch generates the pixelwise human part labels. Outputs of the three branches are fused to produce the final HRS results. Extensive experiments on PIC and V-COCO datasets show that the proposed SMS method outperforms baselines with the 36 FPS inference speed. Notably, SMS outperforms the best performing baseline m-KERN with only 17.6 percent time cost. The dataset and code will be released at http://picdataset.com/challenge/index/. Si Liu 0001, Zitian Wang, Yulu Gao, Lejian Ren, Yue Liao, Guanghui Ren, Bo Li 0006, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Human-Centric Spatio-Temporal Video Grounding With Visual TransformersabstractIn this work, we introduce a novel task – Human-centric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatio-temporal tube of the target person from an untrimmed video based on a given textural description. This task is useful, especially for healthcare and security related applications, where the surveillance videos can be extremely long but only a specific person during a specific period is concerned. HC-STVG is a video grounding task that requires both spatial (where) and temporal (when) localization. Unfortunately, the existing grounding methods cannot handle this task well. We tackle this task by proposing an effective baseline method named Spatio-Temporal Grounding with Visual Transformers (STGVT), which utilizes Visual Transformers to extract cross-modal representations for video-sentence matching and temporal localization. To facilitate this task, we also contribute an HC-STVG datasetThe new dataset is available athttps://github.com/tzhhhh123/HC-STVG. consisting of 5,660 video-sentence pairs on complex multi-person scenes. Specifically, each video lasts for 20 seconds, pairing with a natural query sentence with an average of 17.25 words. Extensive experiments are conducted on this dataset, demonstrating that the newly-proposed method outperforms the existing baseline methods. Zongheng Tang, Yue Liao, Si Liu 0001, Guanbin Li, Xiaojie Jin 0004, Qian Yu 0002, Dong Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Progressive Language-Customized Visual Feature Learning for One-Stage Visual GroundingabstractVisual grounding is a task to localize an object described by a sentence in an image. Conventional visual grounding methods extract visual and linguistic features isolatedly and then perform cross-modal interaction in a post-fusion manner. We argue that this post-fusion mechanism does not fully utilize the information in two modalities. Instead, it is more desired to perform cross-modal interaction during the extraction process of the visual and linguistic feature. In this paper, we propose a language-customized visual feature learning mechanism where linguistic information guides the extraction of visual feature from the very beginning. We instantiate the mechanism as a one-stage framework named Progressive Language-customized Visual feature learning (PLV). Our proposed PLV consists of a Progressive Language-customized Visual Encoder (PLVE) and a grounding module. We customize the visual feature with linguistic guidance at each stage of the PLVE by Channel-wise Language-guided Interaction Modules (CLIM). Our proposed PLV outperforms conventional state-of-the-art methods with large margins across five visual grounding datasets without pre-training on object detection datasets, while achieving real-time speed. The source code is available in the supplementary material. Yue Liao, Aixi Zhang, Zhiyuan Chen 0008, Tianrui Hui, Si Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | A Local-Global Dual-Stream Network for Building Extraction From Very-High-Resolution Remote Sensing ImagesabstractBuildings constitute one of the most important landscapes in remote sensing (RS) images and have been broadly analyzed in a wide range of applications from urban planning to other socioeconomic studies. As very-high-resolution (VHR) RS imagery becomes more accessible, the current building extraction methods are confronted with the challenges of the diverse appearances, various scales, and complicated structures of buildings in complex scenes. With the development of context-aware deep learning methods, it has been proven by numerous works that capturing contextual information can offer spatial relation cues for robust recognition and detection of the objects. In this article, we propose a novel local-global dual-stream network (DS-Net) that adaptively captures local and long-range information for the accurate mapping of building rooftops in VHR RS images. The local branch and the global branch of DS-Net work in a complementary manner to each other with different fields of view on the input image. Through a well-defined dual-stream architecture, DS-Net learns hierarchical representations for both the local and global branches, and a deep feature sharing strategy is further developed to enforce more collaborative integration of the two branches. Extensive experiments were carried out to verify the effectiveness of our model on three widely used VHR RS data sets: the Massachusetts buildings data set, the Inria Aerial Image Labeling data set, and the DeepGlobe Building Detection Challenge data set. Empirically, the proposed DS-Net achieves competitive or superior performance compared with the current state-of-the-art methods in terms of quantitative measures and visual evaluations. Hongyan Zhang 0001, Yue Liao, Honghai Yang, Liangpei Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Reformulating HOI Detection As Adaptive Set PredictionabstractDetermining which image regions to concentrate is critical for Human-Object Interaction (HOI) detection. Conventional HOI detectors focus on either detected human and object pairs or pre-defined interaction locations, which limits learning of the effective features. In this paper, we reformulate HOI detection as an adaptive set prediction problem, with this novel formulation, we propose an Adaptive Set-based one-stage framework (AS-Net) with parallel instance and interaction branches. To attain this, we map a trainable interaction query set to an interaction prediction set with transformer. Each query adaptively aggregates the interaction-relevant features from global contexts through multi-head co-attention. Besides, the training process is supervised adaptively by matching each ground-truth with the interaction prediction. Furthermore, we design an effective instance-aware attention module to introduce instructive features from the instance branch into the interaction branch. Our method outperforms previous state-of-the-art methods without any extra human pose and language features on three challenging HOI detection datasets. Especially, we achieve over 31% relative improvement on a large scale HICO-DET dataset. Code is available at https://github.com/yoyomimi/AS-Net. Mingfei Chen, Yue Liao, Si Liu 0001, Zhiyuan Chen 0008, Fei Wang 0032, Chen Qian 0006 |
CVPR | 2 |
| 2021 | Mining the Benefits of Two-stage and One-stage HOI DetectionabstractTwo-stage methods have dominated Human-Object Interaction~(HOI) detection for several years. Recently, one-stage HOI detection methods have become popular. In this paper, we aim to explore the essential pros and cons of two-stage and one-stage methods. With this as the goal, we find that conventional two-stage methods mainly suffer from positioning positive interactive human-object pairs, while one-stage methods are challenging to make an appropriate trade-off on multi-task learning, \emph{i.e.}, object detection, and interaction classification. Therefore, a core problem is how to take the essence and discard the dregs from the conventional two types of methods. To this end, we propose a novel one-stage framework with disentangling human-object detection and interaction classification in a cascade manner. In detail, we first design a human-object pair generator based on a state-of-the-art one-stage HOI detector by removing the interaction classification module or head and then design a relatively isolated interaction classifier to classify each human-object pair. Two cascade decoders in our proposed framework can focus on one specific task, detection or interaction classification. In terms of the specific implementation, we adopt a transformer-based HOI detector as our base model. The newly introduced disentangling paradigm outperforms existing methods by a large margin, with a significant relative mAP gain of 9.32% on HICO-Det. The source codes are available at https://github.com/YueLiao/CDN. Aixi Zhang, Yue Liao, Si Liu 0001, Miao Lu, Chen Gao 0005 |
NeurIPS | 2 |
| 2021 | Scene Graph Generation With Hierarchical ContextabstractScene graph generation has received increasing attention in recent years. Enhancing the predicate representations is an important entry point to this task. There are various methods to fully investigate the context of representation enhancement. In this brief, we analyze the decisive factors that can significantly affect the relation detection results. Our analysis shows that spatial correlations between objects, focused regions of objects, and global hints related to the relations have strong influences in relation prediction and contradiction elimination. Based on our analysis, we propose a hierarchical context network (HCNet) to generate a scene graph. HCNet consists of three contexts, including interaction context, depression context, and global context, which integrates information from pair, object, and graph levels. The experiments show that our method outperforms the state-of-the-art methods on the Visual Genome (VG) data set. Guanghui Ren, Lejian Ren, Yue Liao, Si Liu 0001, Bo Li 0006, Jizhong Han, Shuicheng Yan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object DetectionabstractKeypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. CentripetalNet predicts the position and the centripetal shift of the corner points and matches corners whose shifted results are aligned. Combining position information, our approach matches corner points more accurately than the conventional embedding approaches do. Corner pooling extracts information inside the bounding boxes onto the border. To make this information more aware at the corners, we design a cross-star deformable convolution network to conduct feature adaption. Furthermore, we explore instance segmentation on anchor-free detectors by equipping our CentripetalNet with a mask prediction module. On COCO test-dev, our CentripetalNet not only outperforms all existing anchor-free detectors with an AP of 48.0% but also achieves comparable performance to the state-of-the-art instance segmentation approaches with a 40.2% Mask AP. Code is available at https: //github.com/KiveeDong/CentripetalNet. Zhiwei Dong, Guoxuan Li, Yue Liao, Fei Wang 0032, Pengju Ren, Chen Qian 0006 |
CVPR | 3 |
| 2020 | A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionabstractReferring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time inference without accuracy drop. The reason for the relatively slow inference speed is that these methods artificially split the referring expression comprehension into two sequential stages including proposal generation and proposal ranking. It does not exactly conform to the habit of human cognition. To this end, we propose a novel Realtime Cross-modality Correlation Filtering method (RCCF). RCCF reformulates the referring expression comprehension as a correlation filtering process. The expression is first mapped from the language domain to the visual domain and then treated as a template (kernel) to perform correlation filtering on the image feature map. The peak value in the correlation heatmap indicates the center points of the target box. In addition, RCCF also regresses a 2-D object size and 2-D offset. The center point coordinates, object size and center point offset together to form the target bounding box. Our method runs at 40 FPS while achieving leading performance in RefClef, RefCOCO, RefCOCO+ and RefCOCOg benchmarks. In the challenging RefClef dataset, our methods almost double the state-of-the-art performance (34.70% increased to 63.79%). We hope this work can arouse more attention and studies to the new cross-modality correlation filtering framework as well as the one-stage framework for referring expression comprehension. Yue Liao, Si Liu 0001, Guanbin Li, Fei Wang 0032, Chen Qian 0006, Bo Li 0006 |
CVPR | 1 |
| 2020 | PPDM: Parallel Point Detection and Matching for Real-Time Human-Object Interaction DetectionabstractWe propose a single-stage Human-Object Interaction (HOI) detection method that has outperformed all existing methods on HICO-DET dataset at 37 fps on a single Titan XP GPU. It is the first real-time HOI detection method. Conventional HOI detection methods are composed of two stages, i.e., human-object proposals generation, and proposals classification. Their effectiveness and efficiency are limited by the sequential and separate architecture. In this paper, we propose a Parallel Point Detection and Matching (PPDM) HOI detection framework. In PPDM, an HOI is defined as a point triplet. Human and object points are the center of the detection boxes, and the interaction point is the midpoint of the human and object points. PPDM contains two parallel branches, namely point detection branch and point matching branch. The point detection branch predicts three points. Simultaneously, the point matching branch predicts two displacements from the interaction point to its corresponding human and object points. The human point and the object point originated from the same interaction point are considered as matched pairs. In our novel parallel architecture, the interaction points implicitly provide context and regularization for human and object detection. The isolated detection boxes unlikely to form meaningful HOI triplets are suppressed, which increases the precision of HOI detection. Moreover, the matching between human and object detection boxes is only applied around limited numbers of filtered candidate interaction points, which saves much computational cost. Additionally, we build a new application-oriented database named as HOI-A, which serves as a good supplement to the existing datasets. Yue Liao, Si Liu 0001, Fei Wang 0032, Chen Qian 0006, Jiashi Feng |
CVPR | 1 |
| 2020 | Local Correlation Consistency for Knowledge Distillation
Jianlong Wu, Hongyu Fang, Yue Liao, Fei Wang 0032, Chen Qian 0006 |
ECCV (12) | 4 |
| 2020 | Learning Discriminative Global and Local Features for Building Extraction from Aerial ImagesabstractBuildings constitute one of the most important landscapes in remote sensing images. Automatic building extraction methods towards the very high resolution remote sensing imagery feature both the local refinement of segmentation results and the context-aware reasoning for segmentation. In this paper, we propose a novel dual-stream convolutional neural network (DS-Net) to collaboratively incorporate local and global features for accurately segmenting buildings in very high resolution aerial images. We develop a hierachical representation and a deep feature sharing strategy for both the local branch and global branch in DS-Net to effectively exploit the complementarity between the two branches. Through extensive experiments on the large-scale building detection datasets, we show that the proposed DS-Net can benefit from both the local and global features, which significantly improves the accuracy of building extraction over diversified remote sensing scenes. Yue Liao, Hongyan Zhang 0001, Liangpei Zhang 0001 |
IGARSS | 1 |
| 2020 | Land Cover Mapping Based On Multi-Branch Fusion Of Object-Based And Pixel-Based Segmentation With Filtered LabelsabstractIn this paper, a multi-branch fusion framework is proposed to address the land cover mapping issue with low-resolution labels. To obtain homogeneous target objects, a multi-resolution segmentation (MRS) algorithm is applied to yield unsupervised object-based segmentation maps. Through an index-based judgement mechanism, a label filtering principle was designed and employed to screen out samples with noisy labels while retaining samples with clean labels, thus acquiring more accurate training data. A patch-to-point classification network was established based on these filtered training patches, which fully extracts the contextual features and generates pixel-based prediction results. A post-processing step, consisting of fusion and voting operations, was developed to merge the pixel-based and object-based results, and produce a final segmentation map. Verified through the competition website, the proposed method achieved an average accuracy (AA) of 57.22%, ranking second in the first track of 2020 IEEE GRSS Data Fusion Contest. Yu Xia 0032, Yue Liao, Hongyan Zhang 0001 |
IGARSS | 2 |
| 2020 | Cross-Modal Omni Interaction Modeling for Phrase GroundingabstractPhrase grounding aims to localize the objects described by phrases in a natural language specification. Previous works model the interaction of inputs from text modality and visual modality only in the intra-modal global level and consequently lacks the ability to capture the precise and complete context information. In this paper, we propose a novel Cross-Modal Omni Interaction network (COI Net) composed of a neighboring interaction module, a global interaction module, a cross-modal interaction module and a multilevel alignment module. Our approach formulates the complex spatial and semantic relationship among image regions and phrases through multi-level multi-modal interaction. We capture the local relationship using the interaction among neighboring regions and then collect the global context through the interaction among all regions using a transformer encoder. We further use a co-attention module to apply the interaction between two modalities to gather the cross-modal context for all image regions and phrases. In addition to the omni interaction modeling, we also leverage a straightforward yet effective multilevel alignment regularization to formulate the dependencies among all grounding decisions. We extensively validate the effectiveness of our model. Experiments show that our approach outperforms existing state-of-the-art methods by large margins on two popular datasets in terms of accuracy: 6.15% on Flickr30K Entities (71.36% increased to 77.51%) and 21.25% on ReferItGame (44.91% increased to 66.16%). The code of our implementation is available at https://github.com/yiranyyu/Phrase-Grounding. Tianyu Yu 0002, Tianrui Hui, Zhihao Yu, Yue Liao, Sansi Yu, Faxi Zhang, Si Liu 0001 |
ACM Multimedia | 4 |
| 2019 | GPS: Group People Segmentation with Detailed Part InferenceabstractNoticeable progress has been witnessed in general object detection, semantic segmentation and instance segmentation, while parsing a group of people is still a challenging task for human-centric visual understanding due to severe occlusion and various poses. In this paper, we present a new large-scale dataset named “GPS (Group People Segmentation)” to boost academical study and technology development. GPS contains 14000 elaborately annotated images with 20 fine-grained semantic category labels related to human, divided into two sub-datasets corresponding to indoor and outdoor scenes involving various poses, occlusion and background. We further propose a novel GPSNet for group people segmentation. GPSNet consists of a new “Adjusted RoI Align” module to adjust position of detected person and align RoI features, such that the network does not need to fit various positions of each person. A fusion of global and local features is also employed to refine parsing results. Compared with baseline methods, GPSNet achieves the best performance on GPS Dataset. Yue Liao, Si Liu 0001, Tianrui Hui, Chen Gao 0005, Yao Sun 0004, Bo Li 0006 |
ICME | 1 |