EDBT 2026 Demo / reviewers in the wild / expert
Ruimao Zhang
dblp:54/10697
· DBLP profile ↗
81ranked-venue papers
5as first author
54since 2021 · last 2026
0000-0001-9511-7532ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 60 · 3 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Aware Embedding Network for Label-Free Radio Map Construction
Zheng Xing 0001, Weibing Zhao, Mengru Wu, Wenjie Liu 0017, Cheng Zeng 0002, Huijun Xing, Ruimao Zhang |
IEEE Trans. Wirel. Commun. | 9 |
| 2025 | DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion GenerationabstractIn autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from performance degradation, known as Out-of-Distribution (OOD) problems. To address this, controllable Text-to-Image (T2I) diffusion offers a potential solution for training data enhancement, which is required to generate diverse OOD scenarios with precise 3D object geometry. Nevertheless, existing controllable T2I approaches are restricted by the limited scale of training data or struggle to preserve all annotated 3D objects. In this paper, we present DriveGEN, a method designed to improve the robustness of 3D detectors in Driving via Training-Free Controllable Text-to-Image Diffusion Generation. Without extra diffusion model training, DriveGEN consistently preserves objects with precise 3D geometry across diverse OOD generations, consisting of 2 stages: 1) Self-Prototype Extraction: We empirically find that self-attention features are semantic-aware but require accurate region selection for 3D objects. Thus, we extract precise object features via layouts to capture 3D object geometry, termed self-prototypes. 2) Prototype-Guided Diffusion: To preserve objects across various OOD scenarios, we perform semantic-aware feature alignment and shallow feature alignment during denoising. Extensive experiments demonstrate our effectiveness in improving 3D detection. The code is available at github.com/Hongbin98/DriveGEN. Zilu Guo, Yifan Zhang 0004, Shuaicheng Niu, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
CVPR | 6 |
| 2025 | ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation ModelabstractThe scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of 1e18. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law. Project page: https://shunlinlu.github.io/ScaMo/ Shunlin Lu, Jingbo Wang 0003, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai 0002, Ruimao Zhang |
CVPR | 9 |
| 2025 | RoboFactory: Exploring Embodied Agent Collaboration with Compositional ConstraintsabstractDesigning effective embodied multi-agent systems is critical for solving complex real-world tasks across domains. Due to the complexity of multi-agent embodied systems, existing methods fail to automatically generate safe and efficient training data for such systems. To this end, we propose the concept of compositional constraints for embodied multi-agent systems, addressing the challenges arising from collaboration among embodied agents. We design various interfaces tailored to different types of constraints, enabling seamless interaction with the physical world. Leveraging compositional constraints and specifically designed interfaces, we develop an automated data collection framework for embodied multi-agent systems and introduce the first benchmark for embodied multi-agent manipulation, RoboFactory. Based on RoboFactory benchmark, we adapt and evaluate the method of imitation learning and analyzed its performance in different difficulty agent tasks. Furthermore, we explore the architectures and training strategies for multi-agent imitation learning, aiming to build safe and efficient embodied multi-agent systems. Yiran Qin, Xiufeng Song, Zhenfei Yin, Xiaohong Liu 0001, Xihui Liu, Ruimao Zhang, Lei Bai 0001 |
ICCV | 7 |
| 2025 | High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian RepresentationabstractWeather nowcasting is an essential task that involves predicting future radar echo sequences based on current observations, offering significant benefits for disaster management, transportation, and urban planning. Current prediction methods are limited by training and storage efficiency, mainly focusing on 2D spatial predictions at specific altitudes. Meanwhile, 3D volumetric predictions at each timestamp remain largely unexplored. To address such a challenge, we introduce a comprehensive framework for 3D radar sequence prediction in weather nowcasting, using the newly proposed SpatioTemporal Coherent Gaussian Splatting (STC-GS) for dynamic radar representation and GauMamba for efficient and accurate forecasting. Specifically, rather than relying on a 4D Gaussian for dynamic scene reconstruction, STC-GS optimizes 3D scenes at each frame by employing a group of Gaussians while effectively capturing their movements across consecutive frames. It ensures consistent tracking of each Gaussian over time, making it particularly effective for prediction tasks. With the temporally correlated Gaussian groups established, we utilize them to train GauMamba, which integrates a memory mechanism into the Mamba framework. This allows the model to learn the temporal evolution of Gaussian groups while efficiently handling a large volume of Gaussian tokens. As a result, it achieves both efficiency and accuracy in forecasting a wide range of dynamic meteorological radar signals. The experimental results demonstrate that our STC-GS can efficiently represent 3D radar sequences with over $16\times$ higher spatial resolution compared with the existing 3D representation methods, while GauMamba outperforms state-of-the-art methods in forecasting a broad spectrum of high-dynamic weather conditions. Yiran Qin, Ruimao Zhang |
ICLR | 4 |
| 2025 | WorldSimBench: Towards Video Generation Models as World SimulatorsabstractRecent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing benchmarks are unable to effectively evaluate higher-capability, highly embodied predictive models from an embodied perspective. In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench. WorldSimBench includes Explicit Perceptual Evaluation and Implicit Manipulative Evaluation, encompassing human preference assessments from the visual perspective and action-level evaluations in embodied tasks, covering three representative embodied scenarios: Open-Ended Embodied Environment, Autonomous, Driving, and Robot Manipulation. In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulater. In the Implicit Manipulative Evaluation, we assess the video-action consistency of World Simulators by evaluating whether the generated situation-aware video can be accurately translated into the correct control signals in dynamic environments. Our comprehensive evaluation offers key insights that can drive further innovation in video generation models, positioning World Simulators as a pivotal advancement toward embodied artificial intelligence. Yiran Qin, Zhelun Shi, Jiwen Yu, Enshen Zhou, Zhenfei Yin, Xihui Liu, Lu Sheng, Lei Bai 0001, Ruimao Zhang |
ICML | 12 |
| 2025 | NavigateDiff: Visual Predictors are Zero-Shot Navigation AssistantsabstractNavigating unfamiliar environments presents significant challenges for household robots, requiring the ability to recognize and reason about novel decoration and layout. Existing reinforcement learning methods cannot be directly transferred to new environments, as they typically rely on extensive mapping and exploration, leading to time-consuming and inefficient. To address these challenges, we try to transfer the logical knowledge and the generalization ability of pretrained foundation models to zero-shot navigation. By integrating a large vision-language model with a diffusion network, our approach named NavigateDiff constructs a visual predictor that continuously predicts the agent's potential observations in the next step which can assist robots generate robust actions. Furthermore, to adapt the temporal property of navigation, we introduce temporal historical information to ensure that the predicted image is aligned with the navigation scene. We then carefully designed an information fusion framework that embeds the predicted future frames as guidance into goalreaching policy to solve downstream image navigation tasks. This approach enhances navigation control and generalization across both simulated and real-world environments. Through extensive experimentation, we demonstrate the robustness and versatility of our method, showcasing its potential to improve the efficiency and effectiveness of robotic navigation in diverse settings. Project Page: https://21styouth.github.io/NavigateDiff/. Yiran Qin, Yuze Hong, Benyou Wang, Ruimao Zhang |
ICRA | 5 |
| 2025 | Semantic-Supervised Spatial-Temporal Fusion for LiDAR-Based 3D Object DetectionabstractLiDAR-based 3D object detection presents significant challenges due to the inherent sparsity of LiDAR points. A common solution involves long-term temporal LiDAR data to densify the inputs. However, efficiently leveraging spatial-temporal information remains an open problem. In this paper, we propose a novel Semantic-Supervised Spatial-Temporal Fusion (ST-Fusion) method, which introduces a novel fusion module to relieve the spatial misalignment caused by the object motion over time and a feature-level semantic supervision to sufficiently unlock the capacity of the proposed fusion module. Specifically, the ST- Fusion consists of a Spatial Aggregation (SA) module and a Temporal Merging (TM) module. The SA module employs a convolutional layer with progressively expanding receptive fields to aggregate the object features from the local regions to alleviate the spatial misalignment, the TM module dynamically extracts object features from the preceding frames based on the attention mechanism for a comprehensive sequential presentation. Besides, in the semantic supervision, we propose a Semantic Injection method to enrich the sparse LiDAR data via injecting the point-wise semantic labels, using it for training a teacher model and providing a reconstruction target at the feature level supervised by the proposed object-aware loss. Extensive experiments on various LiDAR-based detectors demonstrate the effectiveness and universality of our proposal, yielding an improvement of approximately +2.8% in NDS based on the nuScenes benchmark. Chaoqun Wang 0012, Xiaobin Hong 0002, Ruimao Zhang |
ICRA | 4 |
| 2025 | Unlock the Power of Unlabeled Data in Language Driving ModelabstractRecent Vision-based Large Language Models (VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quality annotated data, which is costly and labor-intensive. To address this issue, we propose unlocking the value of abundant yet unlabeled data to improve the language-driving model in a semi-supervised learning manner. Specifically, we first introduce a series of template-based prompts to extract scene information, generating questions that create pseudo-answers for the unlabeled data based on a model trained with limited labeled data. Next, we propose a Self-Consistency Refinement method to improve the quality of these pseudo-annotations, which are later used for further training. By utilizing a pre-trained VisionLLM (e.g., InternVL), we build a strong Language Driving Model (LDM) for driving scene question-answering, outperforming previous state-of-theart methods. Extensive experiments on the DriveLM benchmark show that our approach performs well with just 5% labeled data, achieving competitive performance against models trained with full datasets. In particular, our LDM achieves 44.85% performance with limited labeled data, increasing to 54.27 % when using unlabeled data, while models trained with full datasets reach 60.68% on the DriveLM benchmark. Chaoqun Wang 0012, Xiaobin Hong 0002, Ruimao Zhang |
ICRA | 4 |
| 2025 | Chain-of-Imagination for Reliable Instruction Following in Decision MakingabstractEnabling the embodied agent to imagine step-by-step the future states and sequentially approach these situation-aware states can enhance its capability to make reliable action decisions from textual instructions. In this work, we introduce a simple but effective mechanism called Chain-of-Imagination (CoI), which repeatedly employs a Multimodal Large Language Model (MLLM) equipped with diffusion model to facilitate imagining and acting upon the series of intermediate situation-aware visual sub-goals one by one, resulting in more reliable instruction-following capability. Based on the CoI mechanism, we propose an embodied agent DecisionDreamer as the low-level controller that can be adapted to different open-world scenarios. Extensive experiments demonstrate that Decision-Dreamer can achieve more reliable and accurate decision-making and significantly outperform the state-of-the-art generalist agents in the Minecraft and CALVIN sandbox simulators, regarding the instruction-following capability. For more demos, please see https://sites.google.com/view/decisiondreamer. Enshen Zhou, Yiran Qin, Zhenfei Yin, Zhelun Shi, Yuzhou Huang, Ruimao Zhang, Lu Sheng |
IROS | 6 |
| 2025 | GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion PoliciesabstractDespite significant advances in robotic policy generation, effective coordination in embodied multi-agent systems remains a fundamental challenge—particularly in scenarios where agents must balance individual perspectives with global environmental awareness.
Existing approaches often struggle to balance fine-grained local control with comprehensive scene understanding, resulting in limited scalability and compromised collaboration quality.
In this paper, we present GauDP, a novel Gaussian-image synergistic representation that facilitates scalable, perception-aware imitation learning in multi-agent collaborative systems.
Specifically, GauDP reconstructs a globally consistent 3D Gaussian field from local-view RGB images, allowing all agents to dynamically query task-relevant features from a shared scene representation.
This design facilitates both fine-grained control and globally coherent behavior without requiring additional sensing modalities.
We evaluate GauDP on the RoboFactory benchmark, which includes diverse multi-arm manipulation tasks.
Our method achieves superior performance over existing image-based methods and approaches the effectiveness of point-cloud-driven methods, while maintaining strong scalability as the number of agents increases.
Extensive ablations and visualizations further demonstrate the robustness and efficiency of our unified local-global perception framework for multi-agent embodied learning. Yiran Qin, Jiahua Ma, Zhanglin Peng, Lei Bai 0001, Ruimao Zhang |
NeurIPS | 7 |
| 2025 | ED-Pose++: Enhanced Explicit Box Detection for Conventional and Interactive Multi-Object Keypoint DetectionabstractDetecting keypoints on diverse objects is essential for fine-grained visual understanding and analysis. This paper introduces Enhanced Explicit Box Detection (ED-Pose++), an end-to-end framework that leverages cascade box regression to realize both conventional and interactive multi-object keypoint detection. Unlike traditional one-stage methods, ED-Pose++ innovatively redefines multi-object keypoint detection as a dual-phase explicit box detection, achieving a unified representation and regression optimization process. Specifically, an object detection decoder first extracts each object's position and global features, establishing a good initialization for subsequent keypoint detection. To bring in contextual information near keypoints, we also regard each keypoint as a small box to learn both positions and their related local contents. In practice, an object-to-keypoint detection decoder adopts a collaborative learning strategy between object and keypoint features, facilitating efficient information propagation between global and local perspectives. Rooted on the architecture, we further equip dual-phase box detection with an interactive mechanism that enables the model to refine its predictions based on limited user feedback. During training, we incorporate an error correction scheme to equip the model with an adept self-correction capability for use during inference. The comprehensive experiments demonstrate ED-Pose++'s superior performance in conventional multi-object keypoint detection tasks. For the first time, ED-Pose++ outperforms heatmap-based top-down approaches across various benchmarks, despite operating within a fully end-to-end architecture. The interactive variant also dramatically reduces more than 10 times the labeling effort of 2D keypoint annotation compared with manual-only annotation. Ailing Zeng, Tianhe Ren, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Boosting 3D Object Detection via Self-Distilling Introspective Dataabstract3D object detection is a fundamental yet critical task for autonomous driving. In this paper, we investigate a novel self-distilling paradigm by proposing Self-distilling Introspective Data (SID) to boost the accuracy of 3D object detection in both LiDAR-based and LiDAR-Camera-based scenarios. The proposed SID significantly improves the applicability of the distillation approach since it does not require extra training data or complex teacher network design. Specifically, we first employ an introspective data augmentation method to enrich object-aware information in sparse point clouds through geometric or semantic injection. We then utilize this enhanced data to train a robust teacher model. In contrast to traditional distillation that relies on larger models to enhance the representations of smaller ones, the teacher model in SID shares the same architecture as the student model but exhibits exceptionally high discriminative ability. This enables the effective transfer of rich feature representations to the student model. Rooted on such a scheme, when conducting LiDAR-based detectors, SID significantly enhances the semantic representation capabilities of sparse point clouds. Additionally, in the LiDAR-Camera-based setting, SID also effectively supervises the fusion of the two modalities at the feature level, ensuring more reasonable cross-modal learning. Extensive experiments show the proposed SID improves a variety of detectors. For the LiDAR-based detector, the SID gains 2.31% mAP improvements for the hard objects in KITTI, while 1.76% NDS improvements on nuScenes. For the LiDAR-Camera-based detectors, the SID boosts the detection accuracy significantly, with 1.5% mAP promotion on KITTI and 2.15% NDS improvements on the nuScenes benchmark. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Yukai Shi, Zhen Li 0026, Ruimao Zhang |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferabstractThe field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using single-modal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge, outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D. Linglin Jing, Ying Xue 0003, Xu Yan 0005, Chaoda Zheng, Dong Wang 0028, Ruimao Zhang, Zhigang Wang 0002, Hui Fang 0003, Bin Zhao 0001, Zhen Li 0026 |
AAAI | 6 |
| 2024 | SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language ModelsabstractCurrent instruction-based image editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach of instruction-based image editing that leverages Multimodal Large Language Models (MLLMs) to enhance its understanding and reasoning capabilities. However, direct integration of these elements still faces challenges in situations requiring complex reasoning. To mitigate this, we propose a Bidirectional Interaction Module (BIM) that enables comprehensive bidirectional information interactions between the input image and the MLLM output. During training, we initially incorporate perception data to boost the perception and understanding capabilities of diffusion models. Subsequently, we demonstrate that a small amount of complex instruction editing data can effectively stimulate SmartEdit’ s editing capabilities for more complex instructions. We further construct a new evaluation dataset, Reason-Edit, specifically tailored for complex instruction-based image editing. Both quantitative and qualitative results on this evaluation dataset indicate that our SmartEdit surpasses previous methods, paving the way for the practical application of complex instruction-based image editing. Yuzhou Huang, Liangbin Xie, Xintao Wang 0002, Ziyang Yuan, Xiaodong Cun, Yixiao Ge, Jiantao Zhou 0001, Chao Dong 0005, Ruimao Zhang, Ying Shan |
CVPR | 10 |
| 2024 | SEED-Bench: Benchmarking Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given in-terleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of single image-text inputs, failing to keep up with the strides made in MLLMs. A comprehensive benchmark is imperative for investigating the progress and uncovering the limitations of current MLLMs. In this work, we categorize the capabilities of MLLMs into hierarchical levels from L0to L4based on the modalities they can ac-cept and generate, and propose SEED-Bench, a comprehensive benchmark that evaluates the hierarchical capa-bilities of MLLMs. Specifically, SEED-Bench comprises 24K multiple-choice questions with accurate human annotations, which span 27 dimensions, including the evaluation of both text and image generation. Multiple-choice questions with ground truth options derived from human annotation enable an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 22 prominent open-source MLLMs and summarize valuable observations. By revealing the limitations of existing MLLMs through extensive evaluations, we aim for SEED-Bench to provide insights that will mo-tivate future research toward the goal of General Artificial Intelligence. Dataset and evaluation code are available at https://github.com/AILab-CVC/SEED-Bench. Bohao Li 0002, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang 0092, Ruimao Zhang, Ying Shan |
CVPR | 6 |
| 2024 | MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active PerceptionabstractIt is a long-lasting goal to design an embodied system that can solve long-horizon open-world tasks in human-like ways. However, existing approaches usually struggle with compound difficulties caused by the logic-aware decompo-sition and context-aware execution of these tasks. To this end, we introduce MP5, an open-ended multimodal em-bodied system built upon the challenging Minecraft sim-ulator, which can decompose feasible sub-objectives, de-sign sophisticated situation-aware plans, and perform em-bodied action control, with frequent communication with a goal-conditioned active perception scheme. Specifically, MP5 is developed on top of recent advances in Multimodal Large Language Models (MLLMs), and the system is mod-ulated into functional modules that can be scheduled and collaborated to ultimately solve pre-defined context- and process-dependent tasks. Extensive experiments prove that MP5 can achieve a 22% success rate on difficult process-dependent tasks and a 91 % success rate on tasks that heav-ily depend on the context. Moreover, MP5 exhibits a re-markable ability to address many open-ended tasks that are entirely novel. Please see the project page at https: //iranqin. github.io/MP5. github.io/. Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao 0001 |
CVPR | 6 |
| 2024 | FreeMan: Towards Benchmarking 3D Human Pose Estimation Under Real-World ConditionsabstractEstimating the 3D structure of the human body from nat-ural scenes is afundamental aspect of visual perception. 3D human pose estimation is a vital step in advancing fields like AIGC and human-robot interaction, serving as a crucial tech-nique for understanding and interacting with human actions in real-world settings. However, the current datasets, often collected under single laboratory conditions using complex motion capture equipment and unvarying backgrounds, are insufficient. The absence of datasets on variable conditions is stalling the progress of this crucial task. To facilitate the development of 3D pose estimation, we present FreeMan, the first large-scale, multi-view dataset collected under the real-world conditions. FreeMan was captured by synchronizing 8 smartphones across diverse scenarios. It comprises 11M frames from 8000 sequences, viewed from different perspec-tives. These sequences cover 40 subjects across 10 different scenarios, each with varying lighting conditions. We have also established an semi-automated pipeline containing er-ror detection to reduce the workload of manual check and ensure precise annotation. We provide comprehensive eval-uation baselines for a range of tasks, underlining the sig-nificant challenges posed by FreeMan. Further evaluations of standard indoor/outdoor human sensing datasets reveal that FreeMan offers robust representation transferability in real and complex scenes. FreeMan is publicly available at https://wangjiongw.github.io/freeman. Fengyu Yang 0005, Bingliang Li, Wenbo Gou, Danqi Yan 0001, Ailing Zeng, Yijun Gao, Junle Wang, Yanqing Jing, Ruimao Zhang |
CVPR | 10 |
| 2024 | Open-World Human-Object Interaction Detection via Multi-Modal PromptsabstractIn this paper, we develop MP-HOI, a powerful Multimodal Prompt-based HOI detector designed to leverage both textual descriptions for open-set generalization and visual exemplars for handling high ambiguity in descriptions, realizing HOI detection in the open world. Specifically, it integrates visual prompts into existing language-guided-only HOI detectors to handle situations where textual descriptions face difficulties in generalization and to address complex scenarios with high interaction ambiguity. To facilitate MP-HOI training, we build a large-scale HOI dataset named Magic-HOI, which gathers six existing datasets into a unified label space, forming over 186K images with 2.4K objects, 1.2K actions, and 20K HOI interactions. Furthermore, to tackle the long-tail issue within the Magic-HOI dataset, we introduce an automated pipeline for generating realistically annotated HOI images and present SynHOI, a high-quality synthetic HOI dataset containing 100K images. Leveraging these two datasets, MP-HOI optimizes the HOI task as a similarity learning process between multi-modal prompts and objects/interactions via a unified contrastive loss, to learn generalizable and transferable objects/interactions representations from large-scale data. MP-HOI could serve as a generalist HOI detector, surpassing the HOI vocabulary of existing expert models by more than 30 times. Concurrently, our results demonstrate that MP-HOI exhibits remarkable zero-shot capability in real- world scenarios and consistently achieves a new state-of-the-art performance across various benchmarks. Our project homepage is available at https://MP-HOI.github.iol. Bingliang Li, Ailing Zeng, Ruimao Zhang |
CVPR | 5 |
| 2024 | F-HOI: Toward Fine-Grained Semantic-Aligned 3D Human-Object Interactions
Xuesong Niu, Ruimao Zhang, Siyuan Huang 0001 |
ECCV (37) | 4 |
| 2024 | X-Pose: Detecting Any Keypoints
Ailing Zeng, Ruimao Zhang |
ECCV (46) | 3 |
| 2024 | Enhancing Human-AI Collaboration Through Logic-Guided ReasoningabstractWe present a systematic framework designed to enhance human-robot perception and collaboration through the integration of logical rules and Theory of Mind (ToM). Logical rules provide interpretable predictions and generalize well across diverse tasks, making them valuable for learning and decision-making. Leveraging the ToM for understanding others' mental states, our approach facilitates effective collaboration. In this paper, we employ logic rules derived from observational data to infer human goals and guide human-like agents. These rules are treated as latent variables, and a rule encoder is trained alongside a multi-agent system in the robot's mind. We assess the posterior distribution of latent rules using learned embeddings, representing entities and relations. Confidence scores for each rule indicate their consistency with observed data. Then, we employ a hierarchical reinforcement learning model with ToM to plan robot actions for assisting humans. Extensive experiments validate each component of our framework, and results on multiple benchmarks demonstrate that our model outperforms the majority of existing approaches. Chengzhi Cao, Yinghao Fu, Ruimao Zhang, Shuang Li 0002 |
ICLR | 4 |
| 2024 | HumanTOMATO: Text-aligned Whole-body Motion GenerationabstractThis work targets a novel text-driven **whole-body** motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation tasks mainly have two limitations: they ignore the key role of fine-grained hand and face controlling in vivid whole-body motion generation, and lack a good alignment between text and motion. To address such limitations, we propose a Text-aligned whOle-body Motion generATiOn framework, named HumanTOMATO, which is the first attempt to our knowledge towards applicable holistic motion generation in this research area. To tackle this challenging task, our solution includes two key designs: (1) a Holistic Hierarchical VQ-VAE (aka H${}^{2}$VQ) and a Hierarchical-GPT for fine-grained body and hand motion reconstruction and generation with two structured codebooks; and (2) a pre-trained text-motion-alignment model to help generated motion align with the input textual description explicitly. Comprehensive experiments verify that our model has significant advantages in both the quality of generated motions and their alignment with text. Shunlin Lu, Ailing Zeng, Ruimao Zhang, Lei Zhang 0001, Harry Shum |
ICML | 5 |
| 2024 | Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and CalibrationabstractRecent camera-based 3D object detection is limited by the precision of transforming from image to 3D feature spaces, as well as the accuracy of object localization within the 3D space. This paper aims to address such a fundamental problem of camera-based 3D object detection: How to effectively learn depth information for accurate feature lifting and object localization. Different from previous methods which directly predict depth distributions by using a supervised estimation model, we propose a cascade framework consisting of two depth-aware learning paradigms. First, a depth estimation (DE) scheme leverages relative depth information to realize the effective feature lifting from 2D to 3D spaces. Furthermore, a depth calibration (DC) scheme introduces depth reconstruction to further adjust the 3D object localization perturbation along the depth axis. In practice, the DE is explicitly realized by using both the absolute and relative depth optimization loss to promote the precision of depth prediction, while the capability of DC is implicitly embedded into the detection Transformer through a depth denoising mechanism in the training phase. The entire model training is accomplished through an end-to-end manner. We propose a baseline detector and evaluate the effectiveness of our proposal with +2.2%/+2.7% NDS/mAP improvements on NuScenes benchmark, and gain a comparable performance with 55.9%/45.7% NDS/mAP. Furthermore, we conduct extensive experiments to demonstrate its generality based on various detectors with about +2% NDS improvements. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Ruimao Zhang |
ICRA | 5 |
| 2024 | KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionabstractRecent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pixel-level semantic details, e.g., the keypoints of an object. To bridge this gap, we introduce the novel challenge of Semantic Keypoint Comprehension, which aims to comprehend keypoints across different task scenarios, including keypoint semantic understanding, visual prompt-based keypoint detection, and textual prompt-based keypoint detection. Moreover, we introduce KptLLM, a unified multimodal model that utilizes an identify-then-detect strategy to effectively address these challenges. KptLLM underscores the initial discernment of semantics in keypoints, followed by the precise determination of their positions through a chain-of-thought process. With several carefully designed modules, KptLLM adeptly handles various modality inputs, facilitating the interpretation of both semantic contents and keypoint locations. Our extensive experiments demonstrate KptLLM's superiority in various keypoint detection benchmarks and its unique semantic capabilities in interpreting keypoints. Sheng Jin 0007, Lumin Xu, Wentao Liu 0002, Chen Qian 0006, Ruimao Zhang |
NeurIPS | 7 |
| 2024 | Hierarchical Weight Averaging for Deep Neural NetworksabstractDespite simplicity, stochastic gradient descent (SGD)-like algorithms are successful in training deep neural networks (DNNs). Among various attempts to improve SGD, weight averaging (WA), which averages the weights of multiple models, has recently received much attention in the literature. Broadly, WA falls into two categories: 1) online WA, which averages the weights of multiple models trained in parallel, is designed for reducing the gradient communication overhead of parallel mini-batch SGD and 2) offline WA, which averages the weights of one model at different checkpoints, is typically used to improve the generalization ability of DNNs. Though online and offline WA are similar in form, they are seldom associated with each other. Besides, these methods typically perform either offline parameter averaging or online parameter averaging, but not both. In this work, we first attempt to incorporate online and offline WA into a general training framework termed hierarchical WA (HWA). By leveraging both the online and offline averaging manners, HWA is able to achieve both faster convergence speed and superior generalization performance without any fancy learning rate adjustment. Besides, we also analyze the issues faced by the existing WA methods, and how our HWA addresses them, empirically. Finally, extensive experiments verify that HWA outperforms the state-of-the-art methods significantly. Xiaozhe Gu, Zixun Zhang, Yuncheng Jiang 0002, Tao Luo 0014, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label DomainsabstractThis paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios are addressed, termed universal parsing and dedicated parsing, where the former aims to learn homogeneous human representations from multiple label domains and switch predictions by only using different segmentation heads, and the latter aims to learn a specific domain prediction while distilling the semantic knowledge from other domains. The proposed SST has the following appealing benefits: (1) it can capably serve as an effective training scheme to embed semantic associations of human body parts from multiple label domains into the human representation learning process; (2) it is an extensible semantic transfer framework without predetermining the overall relations of multiple label domains, which allows continuously adding human parsing datasets to promote the training. (3) the relevant modules are only used for auxiliary training and can be removed during inference, eliminating the extra reasoning cost. Experimental results demonstrate SST can effectively achieve promising universal human parsing performance as well as impressive improvements compared to its counterparts on three human parsing benchmarks (i.e., PASCAL-Person-Part, ATR, and CIHP). Code is available at https://github.com/yangjie-cv/SST. Chaoqun Wang 0012, Zhen Li 0026, Junle Wang, Ruimao Zhang |
CVPR | 5 |
| 2023 | SupFusion: Supervised LiDAR-Camera Fusion for 3D Object DetectionabstractLiDAR-Camera fusion-based 3D detection is a critical task for automatic driving. In recent years, many LiDAR-Camera fusion approaches sprung up and gained promising performances compared with single-modal detectors, but always lack carefully designed and effective supervision for the fusion process. In this paper, we propose a novel training strategy called SupFusion, which provides an auxiliary feature level supervision for effective LiDAR-Camera fusion and significantly boosts detection performance. Our strategy involves a data enhancement method named Polar Sampling, which densifies sparse objects and trains an assistant model to generate high-quality features as the supervision. These features are then used to train the LiDAR-Camera fusion model, where the fusion feature is optimized to simulate the generated high-quality features. Furthermore, we propose a simple yet effective deep fusion module, which contiguously gains superior performance compared with previous fusion methods with SupFusion strategy. In such a manner, our proposal shares the following advantages. Firstly, SupFusion introduces auxiliary feature-level supervision which could boost LiDAR-Camera detection performance without introducing extra inference costs. Secondly, the proposed deep fusion could continuously improve the detector’s abilities. Our proposed SupFusion and deep fusion module is plug-and-play, we make extensive experiments to demon-strate its effectiveness. Specifically, we gain around 2% 3D mAP improvements on KITTI benchmark based on multiple LiDAR-Camera 3D detectors. Our code is available at https://github.com/IranQin/SupFusion. Yiran Qin, Chaoqun Wang 0012, Zijian Kang, Ningning Ma, Zhen Li 0026, Ruimao Zhang |
ICCV | 6 |
| 2023 | Neural Interactive Keypoint DetectionabstractThis work proposes an end-to-end neural interactive keypoint detection framework named Click-Pose, which can significantly reduce more than 10 times labeling costs of 2D keypoint annotation compared with manual-only annotation. Click-Pose explores how user feedback can cooperate with a neural keypoint detector to correct the predicted keypoints in an interactive way for a faster and more effective annotation process. Specifically, we design the pose error modeling strategy that inputs the ground truth pose combined with four typical pose errors into the decoder and trains the model to reconstruct the correct poses, which enhances the self-correction ability of the model. Then, we attach an interactive human-feedback loop that allows receiving users’ clicks to correct one or several predicted keypoints and iteratively utilizes the decoder to update all other keypoints with a minimum number of clicks (NoC) for efficient annotation. We validate Click-Pose in in-domain, out-of-domain scenes, and a new task of keypoint adaptation. For annotation, Click-Pose only needs 1.97 and 6.45 NoC@95 (at precision 95%) on COCO and Human-Art, reducing 31.4% and 36.3% efforts than the SOTA model (ViTPose) with manual correction, respectively. Besides, without user clicks, Click-Pose surpasses the previous end-to-end model by 1.4 AP on COCO and 3.0 AP on Human-Art. Ailing Zeng, Feng Li 0040, Shilong Liu 0004, Ruimao Zhang, Lei Zhang 0001 |
ICCV | 5 |
| 2023 | Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation
Ailing Zeng, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001 |
ICLR | 5 |
| 2023 | YONA: You Only Need One Adjacent Reference-Frame for Accurate and Fast Video Polyp Detection
Yuncheng Jiang 0002, Zixun Zhang, Ruimao Zhang, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
MICCAI (5) | 3 |
| 2023 | Dance with You: The Diversity Controllable Dancer Generation via Diffusion ModelsabstractRecently, digital humans for interpersonal interaction in virtual environments have gained significant attention. In this paper, we introduce a novel multi-dancer synthesis task called partner dancer generation, which involves synthesizing virtual human dancers capable of performing dance with users. The task aims to control the pose diversity between the lead dancer and the partner dancer. The core of this task is to ensure the controllable diversity of the generated partner dancer while maintaining temporal coordination with the lead dancer. This scenario varies from earlier research in generating dance motions driven by music, as our emphasis is on automatically designing partner dancer postures according to pre-defined diversity, the pose of lead dancer, as well as the accompanying tunes. To achieve this objective, we propose a three-stage framework called Dance-with-You (DanY). Initially, we employ a 3D Pose Collection stage to collect a wide range of basic dance poses as references for motion generation. Then, we introduce a hyper-parameter that coordinates the similarity between dancers by masking poses to prevent the generation of sequences that are over-diverse or consistent. To avoid the rigidity of movements, we design a Dance Pre-generated stage to pre-generate these masked poses instead of filling them with zeros. After that, a Dance Motion Transfer stage is adopted with leader sequences and music, in which a multi-conditional sampling formula is rewritten to transfer the pre-generated poses into a sequence with a partner style. In practice, to address the lack of multi-person datasets, we introduce AIST-M, a new dataset for partner dancer generation, which is publicly availiable at https://github.com/JJessicaYao/AIST-M-Dataset. Comprehensive evaluations on our AIST-M dataset demonstrate that the proposed DanY can synthesize satisfactory partner dancer results with controllable diversity. Siyue Yao, Mingjie Sun, Bingliang Li, Fengyu Yang 0005, Junle Wang, Ruimao Zhang |
ACM Multimedia | 6 |
| 2023 | Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human ActionsabstractWe propose an interpretable model to uncover the behavioral patterns of human movements by analyzing their trajectories. Our approach is based on the belief that human actions are driven by intentions and are influenced by environmental factors such as spatial relationships with surrounding objects. To model this, we use a set of spatial-temporal logic rules that include intention variables as principles. These rules are automatically discovered and used to capture the dynamics of human actions. To learn the model parameters and rule content, we design an EM learning algorithm that treats the unknown rule content as a latent variable. In the E-step, we evaluate the posterior over the latent rule content, and in the M-step, we optimize the rule generator and model parameters by maximizing the expected log-likelihood. Our model has wide-ranging applications in areas such as sports analytics, robotics, and autonomous cars. We demonstrate the model's superior interpretability and prediction performance on both pedestrian and NBA basketball player datasets, achieving promising results. Chengzhi Cao, Ruimao Zhang, Shuang Li 0002 |
NeurIPS | 3 |
| 2023 | Motion-X: A Large-scale 3D Expressive Whole-body Human Motion DatasetabstractIn this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes. Besides, Motion-X provides 15.6M frame-level whole-body pose descriptions and 81.1K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery. Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, Lei Zhang 0001 |
NeurIPS | 5 |
| 2023 | Community Channel-Net: Efficient channel-wise interactions via community graph topology
Qi Liu 0042, Zhanglin Peng, Ruimao Zhang, Rosa H. M. Chan |
Pattern Recognit. | 4 |
| 2023 | Multi-Stage Spatio-Temporal Aggregation Transformer for Video Person Re-IdentificationabstractIn recent years, the Transformer architecture has shown its superiority in the video-based person re-identification task. Inspired by video representation learning, these methods mainly focus on designing modules to extract informative spatial and temporal features. However, they are still limited in extracting local attributes and global identity information, which are critical for the person re-identification task. In this paper, we propose a novel Multi-Stage Spatial-Temporal Aggregation Transformer (MSTAT) with two novel designed proxy embedding modules to address the above issue. Specifically, MSTAT consists of three stages to encode the attribute-associated, the identity-associated, and the attribute-identity-associated information from the video clips, respectively, achieving the holistic perception of the input person. We combine the outputs of all the stages for the final identification. In practice, to save the computational cost, the Spatial-Temporal Aggregation (STA) modules are first adopted in each stage to conduct the self-attention operations along the spatial and temporal dimensions separately. We further introduce the Attribute-Aware and Identity-Aware Proxy embedding modules (AAP and IAP) to extract the informative and discriminative feature representations at different stages. All of them are realized by employing newly designed self-attention operations with specific meanings. Moreover, temporal patch shuffling is also introduced to further improve the robustness of the model. Extensive experimental results demonstrate the effectiveness of the proposed modules in extracting the informative and discriminative information from the videos, and illustrate the MSTAT can achieve state-of-the-art accuracies on various standard benchmarks. Ruimao Zhang, Zhanglin Peng, Jinrui Chen, Liang Lin 0004 |
IEEE Trans. Multim. | 2 |
| 2022 | Active Domain Adaptation with Multi-level Contrastive Units for Semantic Segmentation
Hao Zhang 0097, Ruimao Zhang |
ACCV (7) | 2 |
| 2022 | Weakly Supervised Object Localization via Transformer with Implicit Spatial Calibration
Haotian Bai, Ruimao Zhang |
ECCV (9) | 2 |
| 2022 | 2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds
Xu Yan 0005, Jiantao Gao, Chaoda Zheng, Chao Zheng 0004, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
ECCV (28) | 5 |
| 2022 | Polygon-Free: Unconstrained Scene Text Detection with Box AnnotationsabstractUnlike existing works that employ fully-supervised training with polygon annotations, this study proposes an unconstrained text detection system termed Polygon-free (PF), in which most existing polygon-based text detectors (e.g., PSENet [1]) are trained with only upright bounding box annotations. Our core idea is to transfer knowledge from synthetic data to real data to enhance the supervision information of upright bounding boxes. This is made possible with a simple segmentation network, namely Skeleton Attention Segmentation Network (SASN), that includes three vital components (i.e., channel attention, spatial attention and skeleton attention map) and one soft cross-entropy loss.Experiments demonstrate that the proposed Polygon-free yields surprisingly high-quality pixel-level results with only upright bounding box annotations. For example, without using polygon annotations, PSENet achieves an 80.5% F-score on TotalText (vs. 80.9% of fully supervised counterpart), 31.1% better than training directly with upright bounding box annotations, and saves 80%+ labeling costs. Weijia Wu 0001, Enze Xie, Ruimao Zhang, Wenhai Wang, Ping Luo 0002 |
ICIP | 3 |
| 2022 | Toward Clinically Assisted Colorectal Polyp Recognition via Structured Cross-Modal Representation Consistency
Weijie Ma 0001, Ruimao Zhang, Yiwen Hu 0001, Zhen Li 0026 |
MICCAI (3) | 3 |
| 2022 | Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisabstractAlthough recent point cloud analysis achieves impressive progress, the paradigm of representation learning from single modality gradually meets its bottleneck. In this work, we take a step towards more discriminative 3D point cloud representation using 2D images, which inherently contain richer appearance information, e.g., texture, color, and shade. Specifically, this paper introduces a simple but effective point cloud cross-modality training (PointCMT) strategy, which utilizes view-images, i.e., rendered or projected 2D images of the 3D object, to boost point cloud classification. In practice, to effectively acquire auxiliary knowledge from view-images, we develop a teacher-student framework and formulate the cross-modal learning as a knowledge distillation problem. Through novel feature and classifier enhancement criteria, PointCMT eliminates the distribution discrepancy between different modalities and avoid potential negative transfer effectively. Note that PointCMT efficiently improves the point-only representation without any architecture modification. Sufficient experiments verify significant gains on various datasets based on several backbones, i.e., equipped with PointCMT, PointNet++ and PointMLP achieve state-of-the-art performance on two benchmarks, i.e., 94.4% and 86.7% accuracy on ModelNet40 and ScanObjectNN, respectively. Xu Yan 0005, Heshen Zhan, Chaoda Zheng, Jiantao Gao, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 5 |
| 2022 | AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image SegmentationabstractDespite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting and labeling 3D medical data, most of the deep learning models to date are driven by datasets with a limited number of organs of interest or samples, which still limits the power of modern deep models and makes it difficult to provide a fully comprehensive and fair estimate of various methods. To mitigate the limitations, we present AMOS, a large-scale, diverse, clinical dataset for abdominal organ segmentation. AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. Information can be found at https://amos22.grand-challenge.org. Yuanfeng Ji, Haotian Bai, Chongjian Ge, Ruimao Zhang, Zhen Li 0026, Wanling Ma, Ping Luo 0002 |
NeurIPS | 6 |
| 2022 | PolarMask++: Enhanced Polar Representation for Single-Shot Instance Segmentation and BeyondabstractReducing complexity of the pipeline of instance segmentation is crucial for real-world applications. This work addresses this problem by introducing an anchor-box free and single-shot instance segmentation framework, termed PolarMask++, which reformulates the instance segmentation problem as predicting the contours of objects in the polar coordinate, leading to several appealing benefits. (1) The polar representation unifies instance segmentation (masks) and object detection (bounding boxes) into a single framework, reducing the design and computational complexity. (2) We carefully design two modules (soft polar centerness and polar IoU loss) to sample high-quality center examples and optimize polar contour regression, making the performance of PolarMask++ does not depend on the bounding box prediction and thus more efficient in training. (3) PolarMask++ is fully convolutional and can be easily embedded into most off-the-shelf detectors. To further improve the accuracy of the framework, a Refined Feature Pyramid is introduced to improve the feature representation at different scales. Extensive experiments demonstrate the effectiveness of PolarMask++, which achieves competitive results on COCO dataset, and new state-of-the-art results on text detection and cell segmentation datasets. We hope polar representation can provide a new perspective for designing algorithms to solve single-shot instance segmentation. Code is released at: github.com/xieenze/PolarMask. Enze Xie, Wenhai Wang, Mingyu Ding, Ruimao Zhang, Ping Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | MetaCloth: Learning Unseen Tasks of Dense Fashion Landmark Detection From a Few SamplesabstractRecent advanced methods for fashion landmark detection are mainly driven by training convolutional neural networks on large-scale fashion datasets, which has a large number of annotated landmarks. However, such large-scale annotations are difficult and expensive to obtain in real-world applications, thus models that can generalize well from a small amount of labelled data are desired. We investigate this problem of few-shot fashion landmark detection, where only a few labelled samples are available for an unseen task. This work proposes a novel framework named MetaCloth via meta-learning, which is able to learn unseen tasks of dense fashion landmark detection with only a few annotated samples. Unlike previous meta-learning work that focus on solving " N -way K -shot" tasks, where each task predicts N number of classes by training with K annotated samples for each class ( N is fixed for all seen and unseen tasks), a task in MetaCloth detects N different landmarks for different clothing categories using K samples, where N varies across tasks, because different clothing categories usually have various number of landmarks. Therefore, numbers of parameters are various for different seen and unseen tasks in MetaCloth. MetaCloth is carefully designed to dynamically generate different numbers of parameters for different tasks, and learn a generalizable feature extraction network from a few annotated samples with a set of good initialization parameters. Extensive experiments show that MetaCloth outperforms its counterparts by a large margin. Yuying Ge, Ruimao Zhang, Ping Luo 0002 |
IEEE Trans. Image Process. | 2 |
| 2022 | Crowd Counting Via Perspective-Guided Fractional-Dilation ConvolutionabstractCrowd counting is critical for numerous video surveillance scenarios. One of the main issues in this task is how to handle the dramatic scale variations of pedestrians caused by the perspective effect. To address this issue, this paper proposes a novel convolution neural network-based crowd counting method, termed Perspective-guided Fractional-Dilation Network (PFDNet). By modeling the continuous scale variations, the proposed PFDNet is able to select the proper fractional-dilation kernels for adapting to different spatial locations. It significantly improves the flexibility of the state-of-the-arts that only consider the discrete representative scales. In addition, by avoiding the multi-scale or multi-column architecture that used in other methods, it is computationally more efficient. In practice, the proposed PFDNet is constructed by stacking multiple Perspective-guided Fractional-Dilation Convolutions (PFC) on a VGG16-BN backbone. By introducing a novel generalized dilation convolution operation, the PFC can handle fractional dilation ratios in the spatial domain under the guidance of perspective annotations, achieving continuous scales modeling of pedestrians. To deal with the problem of unavailable perspective information in some cases, we further introduce an effective perspective estimation branch to the proposed PFDNet, which can be trained in either supervised or weakly-supervised setting once the branch has been pre-trained. Extensive experiments show that the proposed PFDNet outperforms state-of-the-art methods on ShanghaiTech A, ShanghaiTech B, WorldExpo’10, UCF-QNRF, UCF_CC_50 and TRANCOS dataset, achieving MAE 53.8, 6.5, 6.8, 84.3205.8, and 3.06 respectively. Zhaoyi Yan, Ruimao Zhang, Qingfu Zhang 0001, Wangmeng Zuo |
IEEE Trans. Multim. | 2 |
| 2021 | Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionabstractLiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel sparse LiDAR point cloud semantic segmentation framework assisted by learned contextual shape priors. In practice, an initial semantic segmentation (SS) of a single sweep point cloud can be achieved by any appealing network and then flows into the semantic scene completion (SSC) module as the input. By merging multiple frames in the LiDAR sequence as supervision, the optimized SSC module has learned the contextual shape priors from sequential LiDAR data, completing the sparse single sweep point cloud to the dense one. Thus, it inherently improves SS optimization through fully end-to-end training. Besides, a Point-Voxel Interaction (PVI) module is proposed to further enhance the knowledge fusion between SS and SSC tasks, i.e., promoting the interaction of incomplete local geometry of point cloud and complete voxel-wise global structure. Furthermore, the auxiliary SSC and PVI modules can be discarded during inference without extra burden for SS. Extensive experiments confirm that our JS3C-Net achieves superior performance on both SemanticKITTI and SemanticPOSS benchmarks, i.e., 4% and 3% improvement correspondingly. Xu Yan 0005, Jiantao Gao, Jie Li 0002, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
AAAI | 4 |
| 2021 | Parser-Free Virtual Try-On via Distilling Appearance FlowsabstractImage virtual try-on aims to fit a garment image (target clothes) to a person image. Prior methods are heavily based on human parsing. However, slightly-wrong segmentation results would lead to unrealistic try-on images with large artifacts. A recent pioneering work employed knowledge distillation to reduce the dependency of human parsing, where the try-on images produced by a parser-based method are used as supervisions to train a "student" network without relying on segmentation, making the student mimic the try-on ability of the parser-based model. However, the image quality of the student is bounded by the parser-based model. To address this problem, we propose a novel approach, "teacher-tutor-student" knowledge distillation, which is able to produce highly photo-realistic images without human parsing, possessing several appealing advantages compared to prior arts. (1) Unlike existing work, our approach treats the fake images produced by the parser-based method as "tutor knowledge", where the artifacts can be corrected by real "teacher knowledge", which is extracted from the real person images in a self-supervised way. (2) Other than using real images as supervisions, we formulate knowledge distillation in the try-on problem as distilling the appearance flows between the person image and the garment image, enabling us to find accurate dense correspondences between them to produce high-quality results. (3) Extensive evaluations show large superiority of our method (see Fig. 1). Yuying Ge, Yibing Song, Ruimao Zhang, Chongjian Ge, Wei Liu 0005, Ping Luo 0002 |
CVPR | 3 |
| 2021 | End-to-End Dense Video Captioning with Parallel DecodingabstractDense video captioning aims to generate multiple associated captions with their temporal locations from the video. Previous methods follow a sophisticated "localizethen-describe" scheme, which heavily relies on numerous hand-crafted components. In this paper, we proposed a simple yet effective framework for end-to-end dense video captioning with parallel decoding (PDVC), by formulating the dense caption generation as a set prediction task. In practice, through stacking a newly proposed event counter on the top of a transformer decoder, the PDVC precisely segments the video into a number of event pieces under the holistic understanding of the video content, which effectively increases the coherence and readability of predicted captions. Compared with prior arts, the PDVC has several appealing advantages: (1) Without relying on heuristic non-maximum suppression or a recurrent event sequence selection network to remove redundancy, PDVC directly produces an event set with an appropriate size; (2) In contrast to adopting the two-stage scheme, we feed the enhanced representations of event queries into the localization head and caption head in parallel, making these two sub-tasks deeply interrelated and mutually promoted through the optimization; (3) Without bells and whistles, extensive experiments on ActivityNet Captions and YouCook2 show that PDVC is capable of producing high-quality captioning results, surpassing the state-of-the-art two-stage methods when its localization accuracy is on par with them. Code is available at https://github.com/ttengwang/PDVC. Teng Wang 0007, Ruimao Zhang, Zhichao Lu, Feng Zheng 0001, Ran Cheng 0004, Ping Luo 0002 |
ICCV | 2 |
| 2021 | InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringabstractCompared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer1, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our model first predicts the target category from the language descriptions using a simple language classification model. Then, based on the category, our model sifts out a small number of instance candidates (usually less than 20) from the panoptic segmentation on point clouds. Thus, the non-trivial 3D visual grounding task has been effectively re-formulated as a simplified instance-matching problem, considering that instance-level candidates are more rational than the redundant 3D object proposals. Subsequently, for each candidate, we perform the multi-level contextual inference, i.e., referring from instance attribute perception, instance-to-instance relation perception, and instance-to-background global localization perception, respectively. Eventually, the most relevant candidate is selected and localized by ranking confidence scores, which are obtained by the cooperative holistic visual-language feature matching. Experiments confirm that our method outperforms previous state-of-the-arts on ScanRefer online benchmark and Nr3D/Sr3D datasets. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Ruimao Zhang, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 4 |
| 2021 | PointLIE: Locally Invertible Embedding for Point Cloud Sampling and RecoveryabstractPoint Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while preserving the local topology of discarded points in a case-agnostic manner (i.e., without additional storage for point relationships)? We propose a novel Locally Invertible Embedding (PointLIE) framework to unify the point cloud sampling and upsampling into one single framework through bi-directional learning. Specifically, PointLIE decouples the local geometric relationships between discarded points from the sampled points by progressively encoding the neighboring offsets to a latent variable. Once the latent variable is forced to obey a pre-defined distribution in the forward sampling path, the recovery can be achieved effectively through inverse operations. Taking the recover-pleasing sampled points and a latent embedding randomly drawn from the specified distribution as inputs, PointLIE can theoretically guarantee the fidelity of reconstruction and outperform state-of-the-arts quantitatively and qualitatively. Weibing Zhao, Xu Yan 0005, Jiantao Gao, Ruimao Zhang, Jiayan Zhang, Zhen Li 0026, Shuguang Cui |
IJCAI | 4 |
| 2021 | Multi-compound Transformer for Accurate Biomedical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Huijie Wang, Zhen Li 0026, Lingyun Wu, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 2 |
| 2021 | Shallow Attention Network for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Ruimao Zhang, Zhen Li 0026, Shaohua Kevin Zhou, Shuguang Cui |
MICCAI (1) | 3 |
| 2021 | Switchable Normalization for Learning-to-Normalize Deep RepresentationabstractWe address a learning-to-normalize problem by proposing Switchable Normalization (SN), which learns to select different normalizers for different normalization layers of a deep neural network. SN employs three distinct scopes to compute statistics (means and variances) including a channel, a layer, and a minibatch. SN switches between them by learning their importance weights in an end-to-end manner. It has several good properties. First, it adapts to various network architectures and tasks (see Fig. 1). Second, it is robust to a wide range of batch sizes, maintaining high performance even when small minibatch is presented (e.g., 2 images/GPU). Third, SN does not have sensitive hyper-parameter, unlike group normalization that searches the number of groups as a hyper-parameter. Without bells and whistles, SN outperforms its counterparts on various challenging benchmarks, such as ImageNet, COCO, CityScapes, ADE20K, MegaFace and Kinetics. Analyses of SN are also presented to answer the following three questions: (a) Is it useful to allow each normalization layer to select its own normalizer? (b) What impacts the choices of normalizers? (c) Do different tasks and datasets prefer different normalizers? We hope SN will help ease the usage and understand the normalization techniques in deep learning. The code of SN has been released at https://github.com/switchablenorms. Ping Luo 0002, Ruimao Zhang, Jiamin Ren, Zhanglin Peng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Towards Photo-Realistic Virtual Try-On by Adaptively Generating↔Preserving Image ContentabstractImage visual try-on aims at transferring a target clothes image onto a reference person, and has become a hot topic in recent years. Prior arts usually focus on preserving the character of a clothes image (e.g. texture, logo, embroidery) when warping it to arbitrary human pose. However, it remains a big challenge to generate photo-realistic try-on images when large occlusions and human poses are presented in the reference person. To address this issue, we propose a novel visual try-on network, namely Adaptive Content Generating and Preserving Network (ACGPN). In particular, ACGPN first predicts semantic layout of the reference image that will be changed after try-on (e.g.long sleeve shirt→arm, arm→jacket), and then determines whether its image content needs to be generated or preserved according to the predicted semantic layout, leading to photo-realistic try-on and rich clothes details. ACGPN generally involves three major modules. First, a semantic layout generation module utilizes semantic segmentation of the reference image to progressively predict the desired semantic layout after try-on. Second, a clothes warping module warps clothes image according to the generated semantic layout, where a second-order difference constraint is introduced to stabilize the warping process during training.Third, an inpainting module for content fusion integrates all information (e.g. reference image, semantic layout, warped clothes) to adaptively produce each semantic part of human body. In comparison to the state-of-the-art methods, ACGPN can generate photo-realistic images with much better perceptual quality and richer fine-details. Ruimao Zhang, Xiaobao Guo, Wei Liu 0005, Wangmeng Zuo, Ping Luo 0002 |
CVPR | 2 |
| 2020 | Exemplar Normalization for Learning Deep RepresentationabstractNormalization techniques are important in different advanced neural networks and different tasks. This work investigates a novel dynamic learning-to-normalize (L2N) problem by proposing Exemplar Normalization (EN), which is able to learn different normalization methods for different convolutional layers and image samples of a deep network. EN significantly improves the flexibility of the recently proposed switchable normalization (SN), which solves a static L2N problem by linearly combining several normalizers in each normalization layer (the combination is the same for all samples). Instead of directly employing a multi-layer perceptron (MLP) to learn data-dependent parameters as conditional batch normalization (cBN) did, the internal architecture of EN is carefully designed to stabilize its optimization, leading to many appealing benefits. (1) EN enables different convolutional layers, image samples, categories, benchmarks, and tasks to use different normalization methods, shedding light on analyzing them in a holistic view. (2) EN is effective for various network architectures and tasks. (3) It could replace any normalization layers in a deep network and still produce stable model training. Extensive experiments demonstrate the effectiveness of EN in a wide spectrum of tasks including image recognition, noisy label learning, and semantic segmentation. For example, by replacing BN in the ordinary ResNet50, improvement produced by EN is 300% more than that of SN on both ImageNet and the noisy WebVision dataset. The codes and models will be released. Ruimao Zhang, Zhanglin Peng, Lingyun Wu, Zhen Li 0026, Ping Luo 0002 |
CVPR | 1 |
| 2020 | Towards Content-Independent Multi-Reference Super-Resolution: Adaptive Pattern Matching and Feature Aggregation
Xu Yan 0005, Weibing Zhao, Kun Yuan 0004, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
ECCV (25) | 4 |
| 2020 | UXNet: Searching Multi-level Feature Aggregation for 3D Medical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Zhen Li 0026, Jiamin Ren, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 2 |
| 2020 | SSN: Learning Sparse Switchable Normalization via SparsestMax
Wenqi Shao, Jiamin Ren, Ruimao Zhang, Xiaogang Wang 0001, Ping Luo 0002 |
Int. J. Comput. Vis. | 4 |
| 2020 | SCN: Switchable Context Network for Semantic Segmentation of RGB-D ImagesabstractContext representations have been widely used to profit semantic image segmentation. The emergence of depth data provides additional information to construct more discriminating context representations. Depth data preserves the geometric relationship of objects in a scene, which is generally hard to be inferred from RGB images. While deep convolutional neural networks (CNNs) have been successful in solving semantic segmentation, we encounter the problem of optimizing CNN training for the informative context using depth data to enhance the segmentation accuracy. In this paper, we present a novel switchable context network (SCN) to facilitate semantic segmentation of RGB-D images. Depth data is used to identify objects existing in multiple image regions. The network analyzes the information in the image regions to identify different characteristics, which are then used selectively through switching network branches. With the content extracted from the inherent image structure, we are able to generate effective context representations that are aware of both image structures and object relationships, leading to a more coherent learning of semantic segmentation network. We demonstrate that our SCN outperforms state-of-the-art methods on two public datasets. Di Lin 0002, Ruimao Zhang, Yuanfeng Ji, Ping Li 0016, Hui Huang 0004 |
IEEE Trans. Cybern. | 2 |
| 2019 | DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing ImagesabstractUnderstanding fashion images has been advanced by benchmarks with rich annotations such as DeepFashion, whose labels include clothing categories, landmarks, and consumer-commercial image pairs. However, DeepFashion has nonnegligible issues such as single clothing-item per image, sparse landmarks (4∼8 only), and no per-pixel masks, making it had significant gap from real-world scenarios. We fill in the gap by presenting DeepFashion2 to address these issues. It is a versatile benchmark of four tasks including clothes detection, pose estimation, segmentation, and retrieval. It has 801K clothing items where each item has rich annotations such as style, scale, view- point, occlusion, bounding box, dense landmarks (e.g. 39 for ‘long sleeve outwear’ and 15 for ‘vest’), and masks. There are also 873K Commercial-Consumer clothes pairs. The annotations of DeepFashion2 are much larger than its counterparts such as 8× of FashionAI Global Challenge. A strong baseline is proposed, called Match R- CNN, which builds upon Mask R-CNN to solve the above four tasks in an end-to-end manner. Extensive evaluations are conducted with different criterions in Deep- Fashion2. DeepFashion2 Dataset will be released at : https://github.com/switchablenorms/DeepFashion2 Yuying Ge, Ruimao Zhang, Xiaogang Wang 0001, Xiaoou Tang, Ping Luo 0002 |
CVPR | 2 |
| 2019 | SSN: Learning Sparse Switchable Normalization via SparsestMaxabstractNormalization methods improve both optimization and generalization of ConvNets. To further boost performance, the recently-proposed switchable normalization (SN) provides a new perspective for deep learning: it learns to select different normalizers for different convolution layers of a ConvNet. However, SN uses softmax function to learn importance ratios to combine normalizers, leading to redundant computations compared to a single normalizer. This work addresses this issue by presenting Sparse Switchable Normalization (SSN) where the importance ratios are constrained to be sparse. Unlike $\ell_1$ and $\ell_0$ constraints that impose difficulties in optimization, we turn this constrained optimization problem into feed-forward computation by proposing SparsestMax, which is a sparse version of softmax. SSN has several appealing properties. (1) It inherits all benefits from SN such as applicability in various tasks and robustness to a wide range of batch sizes. (2) It is guaranteed to select only one normalizer for each normalization layer, avoiding redundant computations. (3) SSN can be transferred to various tasks in an end-to-end manner. Extensive experiments show that SSN outperforms its counterparts on various challenging benchmarks such as ImageNet, Cityscapes, ADE20K, and Kinetics. Code is available at \url{https://github.com/switchablenorms/Sparse_SwitchNorm}. Wenqi Shao, Tianjian Meng, Ruimao Zhang, Yudian Li, Xiaogang Wang 0001, Ping Luo 0002 |
CVPR | 4 |
| 2019 | Once a MAN: Towards Multi-Target Attack via Learning Multi-Target Adversarial Network OnceabstractModern deep neural networks are often vulnerable to adversarial samples. Based on the first optimization-based attacking method, many following methods are proposed to improve the attacking performance and speed. Recently, generation-based methods have received much attention since they directly use feed-forward networks to generate the adversarial samples, which avoid the time-consuming iterative attacking procedure in optimization-based and gradient-based methods. However, current generation-based methods are only able to attack one specific target (category) within one model, thus making them not applicable to real classification systems that often have hundreds/thousands of categories. In this paper, we propose the first Multi-target Adversarial Network (MAN), which can generate multi-target adversarial samples with a single model. By incorporating the specified category information into the intermediate features, it can attack any category of the target classification model during runtime. Experiments show that the proposed MAN can produce stronger attack results and also have better transferability than previous state-of-the-art methods in both multi-target attack task and single-target attack task. We further use the adversarial samples generated by our MAN to improve the robustness of the classification model. It can also achieve better classification accuracy than other methods when attacked by various methods. Jiangfan Han, Xiaoyi Dong, Ruimao Zhang, Dongdong Chen 0001, Weiming Zhang 0001, Nenghai Yu, Ping Luo 0002, Xiaogang Wang 0001 |
ICCV | 3 |
| 2019 | Differentiable Learning-to-Group Channels via Groupable Convolutional Neural NetworksabstractGroup convolution, which divides the channels of ConvNets into groups, has achieved impressive improvement over the regular convolution operation. However, existing models, \eg ResNext, still suffers from the sub-optimal performance due to manually defining the number of groups as a constant over all of the layers. Toward addressing this issue, we present Groupable ConvNet (GroupNet) built by using a novel dynamic grouping convolution (DGConv) operation, which is able to learn the number of groups in an end-to-end manner. The proposed approach has several appealing benefits. (1) DGConv provides a unified convolution representation and covers many existing convolution operations such as regular dense convolution, group convolution, and depthwise convolution. (2) DGConv is a differentiable and flexible operation which learns to perform various convolutions from training data. (3) GroupNet trained with DGConv learns different number of groups for different convolution layers. Extensive experiments demonstrate that GroupNet outperforms its counterparts such as ResNet and ResNeXt in terms of accuracy and computational complexity. We also present introspection and reproducibility study, for the first time, showing the learning dynamics of training group numbers. Zhaoyang Zhang 0004, Wenqi Shao, Zhanglin Peng, Ruimao Zhang, Xiaogang Wang 0001, Ping Luo 0002 |
ICCV | 5 |
| 2019 | Differentiable Learning-to-Normalize via Switchable Normalization
Ping Luo 0002, Jiamin Ren, Zhanglin Peng, Ruimao Zhang |
ICLR (Poster) | 4 |
| 2019 | Differentiable Dynamic Normalization for Learning Deep RepresentationabstractThis work presents Dynamic Normalization (DN), which is able to learn arbitrary normalization operations for different convolutional layers in a deep ConvNet. Unlike existing normalization approaches that predefined computations of the statistics (mean and variance), DN learns to estimate them. DN has several appealing benefits. First, it adapts to various networks, tasks, and batch sizes. Second, it can be easily implemented and trained in a differentiable end-to-end manner with merely small number of parameters. Third, its matrix formulation represents a wide range of normalization methods, shedding light on analyzing them theoretically. Extensive studies show that DN outperforms its counterparts in CIFAR10 and ImageNet. Ping Luo 0002, Zhanglin Peng, Wenqi Shao, Ruimao Zhang, Jiamin Ren, Lingyun Wu |
ICML | 4 |
| 2019 | Hierarchical Scene Parsing by Weakly Supervised Learning with Image DescriptionsabstractThis paper investigates a fundamental problem of scene understanding: how to parse a scene image into a structured configuration (i.e., a semantic object hierarchy with object interaction relations). We propose a deep architecture consisting of two networks: i) a convolutional neural network (CNN) extracting the image representation for pixel-wise object labeling and ii) a recursive neural network (RsNN) discovering the hierarchical object structure and the inter-object relations. Rather than relying on elaborative annotations (e.g., manually labeled semantic maps and relations), we train our deep model in a weakly-supervised learning manner by leveraging the descriptive sentences of the training images. Specifically, we decompose each sentence into a semantic tree consisting of nouns and verb phrases, and apply these tree structures to discover the configurations of the training images. Once these scene configurations are determined, then the parameters of both the CNN and RsNN are updated accordingly by back propagation. The entire model training is accomplished through an Expectation-Maximization method. Extensive experiments show that our model is capable of producing meaningful scene configurations and achieving more favorable scene labeling results on two benchmarks (i.e., PASCAL VOC 2012 and SYSU-Scenes) compared with other state-of-the-art weakly-supervised deep learning methods. In particular, SYSU-Scenes contains more than 5,000 scene images with their semantic sentence descriptions, which is created by us for advancing research on scene parsing. Ruimao Zhang, Liang Lin 0004, Guangrun Wang, Meng Wang 0001, Wangmeng Zuo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Progressively diffused networks for semantic visual parsing
Ruimao Zhang, Wei Yang 0019, Zhanglin Peng, Pengxu Wei, Xiaogang Wang 0001, Liang Lin 0004 |
Pattern Recognit. | 1 |
| 2019 | SCAN: Self-and-Collaborative Attention Network for Video Person Re-IdentificationabstractVideo person re-identification has attracted much attention in recent years. It aims to match image sequences of pedestrians from different camera views. Previous approaches usually improve this task from three aspects, including: 1) selecting more discriminative frames; 2) generating more informative temporal representations; and 3) developing more effective distance metrics. To address the above issues, we present a novel and practical deep architecture for video person re-identification termed self-and-collaborative attention network (SCAN), which adopts the video pairs as the input and outputs their matching scores. SCAN has several appealing properties. First, SCAN adopts a non-parametric attention mechanism to refine the intra-sequence and inter-sequence feature representation of videos and outputs self-and-collaborative feature representation for each video, making the discriminative frames aligned between the probe and gallery sequences. Second, beyond the existing models, a generalized pairwise similarity measurement is proposed to generate the similarity feature representation of video pair by calculating the Hadamard product of their self-representation difference and collaborative-representation difference. Thus, the matching result can be predicted by the binary classifier. Third, a dense clip segmentation strategy is also introduced to generate rich probe-gallery pairs to optimize the model. In the test phase, the final matching score of two videos is determined by averaging the scores of top-ranked clip-pairs. Extensive experiments demonstrate the effectiveness of SCAN, which outperforms the top-1 accuracies of the best-performing baselines on iLIDS-VID, PRID2011, and MARS datasets, respectively. Ruimao Zhang, Yuying Ge, Ping Luo 0002, Xiaogang Wang 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2018 | Scheduling Large-scale Distributed Training via Reinforcement LearningabstractScheduling the training procedure of deep neural networks (DNNs) such as tuning the learning rates is crucial to the successes of deep learning. Previous strategies such as piecewise and exponential learning rate schedulers have different arguments (hyper-parameters) that need to be tuned manually. With the expanding of data scale and model computation, searching for these arguments spends lots of empirical efforts. To address this issue, this work proposes policy schedular that determines the arguments of learning rate (lr) by reinforcement learning, significantly reducing costs to tune them. The policy schedular possesses several appealing benefits. First, instead of manually defining the values of initial lr and ultimate lr, it autonomously determines these values in training. Second, rather than using predefined functions to update lr, it adaptively oscillates lr by monitoring learning curves without human intervention. Third, it is able to select lr for each block or layer of a DNN. Experiments show that the DNNs trained with policy schedular achieve superior performances, outperforming previous work on various tasks and benchmarks such as ImageNet, COCO, and learning-to-learn. Zhanglin Peng, Jiamin Ren, Ruimao Zhang, Lingyun Wu, Xinjiang Wang, Ping Luo 0002 |
IEEE BigData | 3 |
| 2018 | CUImage: A Neverending Learning Platform on a Convolutional Knowledge Graph of Billion Web ImagesabstractPretraining visual features by image classification on ImageNet is an indispensable step towards many advanced perception systems in the last decade. ImageNet is the most prevalent database for supervised pretraining of image features. Unlike ImageNet assuming that the visual concepts are static and independent with each other, this work presents a neverending learning platform, termed CUImage, which learns visual representation on a knowledge graph of billions of images, whose data scale is several orders of magnitude larger than ImageNet. A novel dynamic graph convolutional network (GCN) is proposed to learn visual concepts. Once the new data are presented, the GCN is updated dynamically where new concepts can be discovered or existing concepts can be merged. This is enabled by three main components in CUImage, including Data Dispersion (DD), Data Management and Mining (DMM), and Data Evaluation (DE). These three components are built on top of a computer cluster with thousands of GPU/CPU cores and a parallel storage of petabytes. So far, CUImage has processed and managed more than 2 million visual concepts of 2 billion images. To evaluate the learned representation, we transfer the pretrained features to several challenging benchmarks such as image recognition on ImageNet and object detection in MS-COCO. We achieve state-of-the-art results, significantly surpassing the systems that used ImageNet for pretraining. The codes, data, and models will be released. Zhanglin Peng, Lingyun Wu, Jiamin Ren, Ruimao Zhang, Ping Luo 0002 |
IEEE BigData | 4 |
| 2018 | Attentive Crowd Flow MachinesabstractTraffic flow prediction is crucial for urban traffic management and public safety. Its key challenges lie in how to adaptively integrate the various factors that affect the flow changes. In this paper, we propose a unified neural network module to address this problem, called Attentive Crowd Flow Machine~(ACFM), which is able to infer the evolution of the crowd flow by learning dynamic representations of temporally-varying data with an attention mechanism. Specifically, the ACFM is composed of two progressive ConvLSTM units connected with a convolutional layer for spatial weight prediction. The first LSTM takes the sequential flow density representation as input and generates a hidden state at each time-step for attention map inference, while the second LSTM aims at learning the effective spatial-temporal feature expression from attentionally weighted crowd flow features. Based on the ACFM, we further build a deep architecture with the application to citywide crowd flow prediction, which naturally incorporates the sequential and periodic data as well as other external influences. Extensive experiments on two standard benchmarks (i.e., crowd flow in Beijing and New York City) show that the proposed method achieves significant improvements over the state-of-the-art methods. Lingbo Liu, Ruimao Zhang, Jiefeng Peng, Guanbin Li, Bowen Du 0001, Liang Lin 0004 |
ACM Multimedia | 2 |
| 2018 | Learning deep representations for semantic image parsing: a comprehensive overview
Lili Huang 0004, Jiefeng Peng, Ruimao Zhang, Guanbin Li, Liang Lin 0004 |
Frontiers Comput. Sci. | 3 |
| 2018 | Image-to-Video Person Re-Identification With Temporally Memorized Similarity LearningabstractWith the development of video surveillance in public safety field, there is an increasing research on person re-identification (re-id). In this paper, we address the image-to-video person re-id, in which the probe is an image and the gallery is consists of videos captured by nonoverlapping cameras. Compared with image, video sequence contains more temporal information that can be explored to improve the performance of re-identification system. However, it is challenging to model temporal information in the matching process of image-to-video person re-id. In this paper, we proposed a novel temporally memorized similarity learning neural network for this problem. In specific, the proposed network mainly consisted of two parts, including feature representation sub-network and similarity sub-network. In the first part, we adopted a convolutional neural network (CNN) to extract features from the input image. Given a video sequence of a person, features were first extracted from each its frame by using CNN and further forward to a long shot term memory (LSTM) network to encode the temporal information of video sequence. The outputs of LSTM were concatenated together as the feature vector of video sequences. Finally, the feature vectors of probe image and the video sequence were further forward to the similarity sub-network for distance metric learning. In the proposed framework, the feature representation and the similarity metric learning can be learned and optimized simultaneously. We evaluated the proposed framework on three public person re-id data sets, and the experimental results showed that the proposed approach is effective for the image-to-video person re-id. Dongyu Zhang 0002, Wenxi Wu, Hui Cheng 0002, Ruimao Zhang, Zhenjiang Dong, Zhaoquan Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Cost-Effective Active Learning for Deep Image ClassificationabstractRecent successes in learning-based image classification, however, heavily rely on the large number of annotated training samples, which may require considerable human effort. In this paper, we propose a novel active learning (AL) framework, which is capable of building a competitive classifier with optimal feature representation via a limited amount of labeled training instances in an incremental learning manner. Our approach advances the existing AL methods in two aspects. First, we incorporate deep convolutional neural networks into AL. Through the properly designed framework, the feature representation and the classifier can be simultaneously updated with progressively annotated informative samples. Second, we present a cost-effective sample selection strategy to improve the classification performance with less manual annotations. Unlike traditional methods focusing on only the uncertain samples of low prediction confidence, we especially discover the large amount of high-confidence samples from the unlabeled set for feature learning. Specifically, these high-confidence samples are automatically selected and iteratively assigned pseudolabels. We thus call our framework cost-effective AL (CEAL) standing for the two advantages. Extensive experiments demonstrate that the proposed CEAL framework can achieve promising results on two challenging image classification data sets, i.e., face recognition on the cross-age celebrity face recognition data set database and object categorization on Caltech-256. Keze Wang, Dongyu Zhang 0002, Ruimao Zhang, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Deep Structured Scene Parsing by Learning with Image DescriptionsabstractThis paper addresses a fundamental problem of scene understanding: How to parse the scene image into a structured configuration (i.e., a semantic object hierarchy with object interaction relations) that finely accords with human perception. We propose a deep architecture consisting of two networks: i) a convolutional neural network (CNN) extracting the image representation for pixelwise object labeling and ii) a recursive neural network (RNN) discovering the hierarchical object structure and the inter-object relations. Rather than relying on elaborative user annotations (e.g., manually labeling semantic maps and relations), we train our deep model in a weakly-supervised manner by leveraging the descriptive sentences of the training images. Specifically, we decompose each sentence into a semantic tree consisting of nouns and verb phrases, and facilitate these trees discovering the configurations of the training images. Once these scene configurations are determined, then the parameters of both the CNN and RNN are updated accordingly by back propagation. The entire model training is accomplished through an Expectation-Maximization method. Extensive experiments suggest that our model is capable of producing meaningful and structured scene configurations and achieving more favorable scene labeling performance on PASCAL VOC 2012 over other state-of-the-art weakly-supervised methods. Liang Lin 0004, Guangrun Wang, Ruimao Zhang, Xiaodan Liang, Wangmeng Zuo |
CVPR | 4 |
| 2016 | Geometric Scene Parsing with Hierarchical LSTM
Zhanglin Peng, Ruimao Zhang, Xiaodan Liang, Xiaobai Liu, Liang Lin 0004 |
IJCAI | 2 |
| 2015 | Adaptive Scene Category Discovery With Generative Learning and Compositional SamplingabstractThis paper investigates a general framework to discover categories of unlabeled scene images according to their appearances (i.e., textures and structures). We jointly solve the two coupled tasks in an unsupervised manner: 1) classifying images without predetermining the number of categories and 2) pursuing generative model for each category. In our method, each image is represented by two types of image descriptors that are effective to capture image appearances from different aspects. By treating each image as a graph vertex, we build up a graph and pose the image categorization as a graph partition process. Specifically, a partitioned subgraph can be regarded as a category of scenes and we define the probabilistic model of graph partition by accumulating the generative models of all separated categories. For efficient inference with the graph, we employ a stochastic cluster sampling algorithm, which is designed based on the Metropolis-Hasting mechanism. During the iterations of inference, the model of each category is analytically updated by a generative learning algorithm. In the experiments, our approach is validated on several challenging databases, and it outperforms other popular state-of-the-art methods. The implementation details and empirical analysis are presented as well. Liang Lin 0004, Ruimao Zhang, Xiaohua Duan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Bit-Scalable Deep Hashing With Regularized Similarity Learning for Image Retrieval and Person Re-IdentificationabstractExtracting informative image features and learning effective approximate hashing functions are two crucial steps in image retrieval. Conventional methods often study these two steps separately, e.g., learning hash functions from a predefined hand-crafted feature space. Meanwhile, the bit lengths of output hashing codes are preset in the most previous methods, neglecting the significance level of different bits and restricting their practical flexibility. To address these issues, we propose a supervised learning framework to generate compact and bit-scalable hashing codes directly from raw images. We pose hashing learning as a problem of regularized similarity learning. In particular, we organize the training images into a batch of triplet samples, each sample containing two images with the same label and one with a different label. With these triplet samples, we maximize the margin between the matched pairs and the mismatched pairs in the Hamming space. In addition, a regularization term is introduced to enforce the adjacency consistency, i.e., images of similar appearances should have similar codes. The deep convolutional neural network is utilized to train the model in an end-to-end fashion, where discriminative image features and hash functions are simultaneously optimized. Furthermore, each bit of our hashing codes is unequally weighted, so that we can manipulate the code lengths by truncating the insignificant bits. Our framework outperforms state-of-the-arts on public benchmarks of similar image search and also achieves promising results in the application of person re-identification in surveillance. It is also shown that the generated bit-scalable hashing codes well preserve the discriminative powers with shorter code lengths. Ruimao Zhang, Liang Lin 0004, Wangmeng Zuo, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2014 | Deep boosting: Layered feature mining for general image classificationabstractConstructing effective representations is a critical but challenging problem in multimedia understanding. The traditional handcraft features often rely on domain knowledge, limiting the performances of exiting methods. This paper discusses a novel computational architecture for general image feature mining, which assembles the primitive filters (i.e. Gabor wavelets) into compositional features in a layer-wise manner. In each layer, we produce a number of base classifiers (i.e. regression stumps) associated with the generated features, and discover informative compositions by using the boosting algorithm. The output compositional features of each layer are treated as the base components to build up the next layer. Our framework is able to generate expressive image representations while inducing very discriminate functions for image classification. The experiments are conducted on several public datasets, and we demonstrate superior performances over state-of-the-art approaches. Zhanglin Peng, Liang Lin 0004, Ruimao Zhang |
ICME | 3 |
| 2011 | Color style transfer by constraint locally linear embeddingabstractThis paper presents a new semi-automatic method for color style transfer between images, which enhances the artistic expression of a image while preserving the content. Our method consists of three steps. (1) We first parse an input image into several semantic objects according to different material properties using interactive algorithms. (2) We search for a proper reference image from library using the semantic information. (3) The dominant colors of the input image and reference image are computed by clustering, and then we propose a constrainted locally linear embedding (CLLE) algorithm to perform color style transfer on the input image. In the experiments, we apply the proposed method to several photos to produce expressive results. Ruimao Zhang, Xiaodan Lan, Yan Pan 0002, Liang Lin 0004 |
ICIP | 2 |