VLDB 2026 Research / reviewers in the wild / expert
Zhaoxiang Zhang 0001
dblp:55/2285-1 · also Zhao-Xiang Zhang 0001
· DBLP profile ↗
292ranked-venue papers
20as first author
163since 2021 · last 2026
0000-0003-2648-3875ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 222 · 14 first-author · 140 since 2021Graphics, computer vision, multimedia, augmented reality and games · 179 · 11 first-author · 89 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Security and privacy · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned AdaptationabstractThe surface pressure field of transportation systems, including cars, trains, and aircraft, is critical for aerodynamic analysis and design. In recent years, deep neural networks have emerged as promising and efficient methods for modeling surface pressure field, being alternatives to computationally expensive CFD simulations. Currently, large-scale public datasets are available for domains such as automotive aerodynamics. However, in many specialized areas, such as high-speed trains, data scarcity remains a fundamental challenge in aerodynamic modeling, severely limiting the effectiveness of standard neural network approaches. To address this limitation, we propose the Adaptive Field Learning Framework (AdaField), which pre-trains the model on public large-scale datasets to improve generalization in sub-domains with limited data. AdaField comprises two key components. First, we design the Semantic Aggregation Point Transformer (SAPT) as a high-performance backbone that efficiently handles large-scale point clouds for surface pressure prediction. Second, regarding the substantial differences in flow conditions and geometric scales across different aerodynamic subdomains, we propose Flow-Conditioned Adapter (FCA) and Physics-Informed Data Augmentation (PIDA). FCA enables the model to flexibly adapt to different flow conditions with a small set of trainable parameters, while PIDA expands the training data distribution to better cover variations in object scale and velocity. Our experiments show that AdaField achieves SOTA performance on the DrivAerNet++ dataset and can be effectively transferred to train and aircraft scenarios with minimal fine-tuning. These results highlight AdaField’s potential as a generalizable and transferable solution for surface pressure field modeling, supporting efficient aerodynamic design across a wide range of transportation systems. Junhong Zou, Zhenxu Sun, Zhaoxiang Zhang 0001, Xiangyu Zhu 0001 |
AAAI | 5 |
| 2026 | CriticLean: Critic-Guided Reinforcement Learning for Mathematical FormalizationabstractZhongyuan Peng, Yifan Yao, Kaijing Ma, Shuyue Guo, Yizhe Li, Yichi Zhang, Chenchen Zhang, Yifan Zhang, Zhouliang Yu, Luming Li, Minghao Liu, Yihang Xia, Jiawei Shen, Yuchen Wu, Yixin Cao, Zhaoxiang Zhang, Wenhao Huang, Jiaheng Liu, Ge Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhongyuan Peng, Kaijing Ma, Shuyue Guo, Yichi Zhang 0010, Zhouliang Yu, Luming Li, Minghao Liu 0003, Yihang Xia, Yixin Cao 0002, Zhaoxiang Zhang 0001, Wenhao Huang 0001, Ge Zhang 0009 |
ACL (1) | 16 |
| 2026 | LLM4CGDS: Large language model-based agents for Chinese graded document simplificationabstractGraded reading tailors text difficulty to learners’ proficiency by producing multiple versions of the same content—an approach long embraced in language education but still dependent on labor-intensive, expert-driven adaptation. In this paper, we introduce the task of C hinese G raded D ocument S implification (CGDS) for non-native learners, which seeks to automate the creation of multi-level reading materials in accordance with established proficiency standards. Guided by the three stages of the Hanyu Shuiping Kaoshi (HSK) 3.0 framework (Levels 1–3 for Advanced, Levels 4–6 for Intermediate, and Levels 7–9 for Beginner learners), we propose Large Language Model for Chinese Graded Document Simplification (LLM4CGDS), a rule-guided, large language model (LLM)-based framework that integrates HSK-level readability constraints and external knowledge retrieval to control document-level simplification without requiring supervised fine-tuning. To foster further research, we construct two complementary datasets: J ourney to the W est D ocument S implification (JWDS) and M ulti- D omain D ocument S implification (MDDS) that covering diverse genres and difficulty levels. Experimental evaluation on two datasets demonstrates that LLM4CGDS substantially outperforms direct prompting of state-of-the-art LLMs in both readability control and meaning preservation. Dengzhao Fang, Jipeng Qiang, Wenjie Hou, Yi Zhu 0006, Jingtong Gao, Zhaoxiang Zhang 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Learning Pseudo 3D Representation for Ego-centric 2D Multiple Object Tracking
Jiawei He 0002, Lue Fan, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
Yuxi Wang 0001, Junran Peng, Genghao Zhang, Chuanchen Luo, Shibiao Xu, Man Zhang 0005, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | Secure Quantum Uplink MIMO System With Rydberg Atomic MIMO ReceiversabstractPhysical layer security (PLS) provides information-theoretic confidentiality for next-generation wireless systems, yet its practical deployment is fundamentally constrained by the sensitivity and noise characteristics of classical radio receivers. Recent advances in quantum sensing, particularly Rydberg atom-based receivers (RAQRs), present a promising approach to address these limitations. This paper introduces a Rydberg atomic quantum multiple-input multiple-output (RAQ-MIMO) architecture for secure multi-user uplink communications, where an array of Rydberg vapor cells functions as a high-sensitivity, reconfigurable receiving front-end. We develop a comprehensive end-to-end signal and channel model that incorporates quantum sensor response, wireless fading, and eavesdropper behavior under passive interception. To maximize the achievable secrecy rate, we propose a joint optimization framework that adaptively tunes key system degrees of freedom, including the Rabi frequency of the local oscillator, phase configuration, digital combining matrix, and user transmit power. Extensive simulations validate that RAQ-MIMO substantially outperforms classical MIMO in terms of secrecy rate across varying transmission distances, power budgets, and user loads—particularly under low-signal to noise ratio (SNR) and extended-range regimes. Our work establishes a quantum-sensing-assisted security foundation for future high-assurance wireless networks, bridging atomic physics with scalable communication system design with scalable communication system design under narrowband operation. Jianxin Dai, Zhaohui Yang 0001, Zhaoxiang Zhang 0001, Youguo Wang |
IEEE Internet Things J. | 4 |
| 2026 | Practical Continual Forgetting for Pre-Trained Vision ModelsabstractFor privacy and security concerns, the need to erase unwanted information from pre-trained vision models is becoming evident nowadays. In real-world scenarios, erasure requests originate at any time from both users and model owners, and these requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify three key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. (iii) In real-world scenarios, the training samples may be scarce or partially missing during the process of forgetting. To address them, we first propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we introduce Low-Rank Adaptation (LoRA) modules to fine-tune the Feed-Forward Network (FFN) layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. To further extend GS-LoRA to more practical scenarios, we incorporate prototype information as additional supervision and introduce a more practical approach, GS-LoRA++. For each forgotten class, we move the logits away from its original prototype. For the remaining classes, we pull the logits closer to their respective prototypes. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that our method manages to forget specific classes with minimal impact on other classes. Hongbo Zhao 0006, Fei Zhu 0004, Bolin Ni, Gaofeng Meng, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decompositionabstract3D Gaussian Splatting has emerged as an alternative 3D representation for novel view synthesis, benefiting from its high-quality rendering results and real-time rendering speed. However, the 3D Gaussians learned by 3D-GS have ambiguous structures without any geometry constraints. This inherent issue in 3D-GS leads to a rough boundary when segmenting individual objects. To remedy these problems, we propose SAGD, a conceptually simple yet effective boundary-enhanced segmentation pipeline for 3D-GS to improve segmentation accuracy while preserving segmentation speed. Specifically, we introduce a Gaussian Decomposition scheme, which ingeniously utilizes the special structure of 3D Gaussians, finds out, and then decomposes the boundary Gaussians. Moreover, to achieve fast interactive 3D segmentation, we introduce a novel training-free pipeline by lifting a 2D foundation model to 3D-GS. Extensive experiments demonstrate that our approach achieves high-quality 3D segmentation without rough boundary issues, which can be easily applied to other scene editing tasks. Our code is publicly available at https://github.com/XuHu0529/SAGS. Yuxi Wang 0001, Lue Fan, Chuanchen Luo, Junsong Fan, Zhen Lei 0001, Qing Li 0001, Junran Peng, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 9 |
| 2025 | FIRM: Flexible Interactive Reflection ReMovalabstractRemoving reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfactory performance, they often lack the flexibility to adapt user guidance in different modalities, and dense user interactions further limit their practicality. To alleviate these problems, this paper presents FIRM, a novel framework for Flexible Interactive image Reflection reMoval with various forms of guidance, where users can provide sparse visual guidance (e.g., points, boxes, or strokes) or text descriptions for better reflection removal. Firstly, we design a novel user guidance conversion module (UGC) to transform different forms of guidance into unified contrastive masks. The contrastive masks provide explicit cues for identifying reflection and transmission layers in blended images. Secondly, we devise a contrastive mask-guided reflection removal network that comprises a newly proposed contrastive guidance interaction block (CGIB). This block leverages a unique cross-attention mechanism that merges contrastive masks with image features, allowing for precise layer separation. The proposed framework requires only 10% of the guidance time needed by previous interactive methods, which makes a step-change in flexibility. Extensive results on public real-world reflection removal datasets validate that our method demonstrates state-of-the-art reflection removal performance. Xiao Chen 0016, Yunkang Tao, Zhen Lei 0001, Qing Li 0001, Chenyang Lei, Zhaoxiang Zhang 0001 |
AAAI | 7 |
| 2025 | SceneX: Procedural Controllable Large-Scale Scene GenerationabstractDeveloping comprehensive explicit world models is crucial for understanding and simulating real-world scenarios. Recently, Procedural Controllable Generation (PCG) has gained significant attention in large-scale scene generation by enabling the creation of scalable, high-quality assets. However, PCG faces challenges such as limited modular diversity, high expertise requirements, and challenges in managing the diverse elements and structures in complex scenes. In this paper, we introduce a large-scale scene generation framework, SceneX, which can automatically produce high-quality procedural models according to designers' textual descriptions. Specifically, the proposed method comprises two components, PCGHub and PCGPlanner. The former encompasses an extensive collection of accessible procedural assets and thousands of hand-craft API documents to perform as a standard protocol for PCG controller. The latter aims to generate executable actions for Blender to produce controllable and precise 3D assets guided by the user's instructions. Extensive experiments demonstrated the capability of our method in controllable large-scale scene generation, including nature scenes and unbounded cities, as well as scene editing such as asset placement and season translation. Mengqi Zhou, Yuxi Wang 0001, Shougao Zhang, Chuanchen Luo, Junran Peng, Zhaoxiang Zhang 0001 |
AAAI | 8 |
| 2025 | Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?abstractYancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Z.y. Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yancheng He, Weixun Wang, Xingyuan Bu, Ge Zhang 0009, Z. Y. Peng, Zhaoxiang Zhang 0001, Zhicheng Zheng, Wenbo Su, Bo Zheng 0007 |
ACL (1) | 8 |
| 2025 | OpenCoder: The Open Cookbook for Top-Tier Code Large Language ModelsabstractSiming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001 |
ACL (1) | 17 |
| 2025 | AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMsabstractUser interface understanding with vision-language models (VLMs) has received much attention due to its potential for enhancing software automation.However, existing datasets used to build UI-VLMs either only contain large-scale context-free element annotations or contextualized functional descriptions for elements at a small scale.In this work, we propose the AutoGUI pipeline for automatically annotating UI elements with detailed functionality descriptions at scale.Specifically, we leverage large language models (LLMs) to infer element functionality by comparing UI state changes before and after simulated interactions. To improve annotation quality, we propose LLM-aided rejection and verification, eliminating invalid annotations without human labor.We construct a high-quality AutoGUI-704k dataset using the proposed pipeline, featuring diverse and detailed functionality annotations that are hardly provided by previous datasets.Human evaluation shows that we achieve annotation correctness comparable to a trained human annotator. Extensive experiments show that our dataset remarkably enhances VLM’s UI grounding capabilities and exhibits significant scaling effects. We also show the interesting potential use of our dataset in UI agent tasks. Please view our project at https://autogui-project.github.io/. Jingfan Chen, Jingran Su, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
ACL (1) | 6 |
| 2025 | M2RC-EVAL: Massively Multilingual Repository-level Code Completion EvaluationabstractJiaheng Liu, Ken Deng, Congnan Liu, Jian Yang, Shukai Liu, He Zhu, Peng Zhao, Linzheng Chai, Yanan Wu, JinKe JinKe, Ge Zhang, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang, Wenbo Su, Bo Zheng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ken Deng, Congnan Liu, Jian Yang 0030, Linzheng Chai, Ge Zhang 0009, Zekun Moore Wang, Guoan Zhang, Yingshui Tan, Bangyu Xiang, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007 |
ACL (1) | 16 |
| 2025 | FreeSim: Toward Free-viewpoint Camera Simulation in Driving ScenesabstractWe propose FreeSim, a camera simulation method for autonomous driving via 3D Gaussian Splatting and diffusion-based image generation. FreeSim emphasizes high-quality rendering from viewpoints beyond the recorded ego trajectories. In such viewpoints, previous methods have unacceptable degradation because the training data of these viewpoints is unavailable. To address such data scarcity, we first propose a generative enhancement model with a matched data construction strategy. The resulting model can generate high-quality images in a viewpoint slightly deviated from the recorded trajectories, conditioned on the degraded rendering of this viewpoint. We then propose a progressive reconstruction strategy, which progressively adds generated images of unrecorded views into the reconstruction process, starting from slightly off-trajectory viewpoints and moving progressively farther away. With this progressive generation-reconstruction pipeline, FreeSim supports high-quality off-trajectory view synthesis under large deviations of more than 3 meters. Lue Fan, Qitai Wang, Hongsheng Li 0001, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2025 | FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and RenderingabstractDriving scene reconstruction and rendering have advanced significantly using the 3D Gaussian Splatting. However, most prior research has focused on the rendering quality along a pre-recorded vehicle path and struggles to generalize to out-of-path viewpoints, which is caused by the lack of high-quality supervision in those out-of-path views. To address this issue, we introduce an Inverse View Warping technique to create compact and high-quality images as supervision for the reconstruction of the out-of-path views, enabling high-quality rendering results for those views. For accurate and robust inverse view warping, a depth bootstrap strategy is proposed to obtain on-the-fly dense depth maps during the optimization process, overcoming the sparsity and incompleteness of LiDAR depth data. Our method achieves superior in-path and out-of-path reconstruction and rendering performance on the widely used Waymo Open dataset. In addition, a simulator-based benchmark is proposed to obtain the out-of-path ground truth and quantitatively evaluate the performance of out-of-path rendering, where our method outperforms previous methods by a significant margin. Our code is available at https://github.com/zhou745/FlexDrive.git. Jingqiu Zhou, Lue Fan, Linjiang Huang, Xiaoyu Shi 0002, Si Liu 0001, Zhaoxiang Zhang 0001, Hongsheng Li 0001 |
CVPR | 6 |
| 2025 | MIO: A Foundation Model on Multimodal TokensabstractZekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001 |
EMNLP | 12 |
| 2025 | SflLLM: Efficient Split Federated Learning for Large Language Model over Wireless NetworksabstractFine-tuning large language models (LLM) in a distributed manner over edge devices with limited communication and computational resources presents substantial challenges in wireless networks. To tackle these issues, this paper proposes a novel Split Federated Learning framework tailored for LLM (SflLLM), which integrates split federated learning with parameter-efficient fine-tuning techniques. By employing model partitioning and low-rank adaptation (LoRA), SflLLM significantly reduces the computational load on edge devices. Moreover, the introduction of the federated server not only facilitates parallel training but also enhances privacy preservation. To accommodate the heterogeneous communication conditions and diverse computational capacities of edge devices—while accounting for the influence of LoRA rank selection on model convergence and training overhead—we formulate a joint optimization problem. This problem simultaneously optimizes subchannel allocation, power control, model split point selection, and LoRA rank configuration, with the objective of minimizing the overall training latency. An alternating optimization algorithm is developed to efficiently solve the proposed problem and accelerate the training process. Simulation results demonstrate that, compared to conventional methods, the proposed resource allocation scheme and adaptive LoRA rank selection strategy significantly reduce training latency. Mingzhe Chen, Chongwen Huang, Zhaohui Yang 0001, Zhaoxiang Zhang 0001 |
GLOBECOM | 6 |
| 2025 | DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerabstractSpeech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to improved performance. We suspect this is due to a shortage of paired audio-4D data, which is crucial for the Transformer to effectively perform as a denoiser within the Diffusion framework. To tackle this issue, we present DiffSpeaker, a Transformer-based network equipped with novel biased conditional attention modules. These modules serve as substitutes for the traditional self/cross-attention in standard Transformers, incorporating thoughtfully designed biases that steer the attention mechanisms to concentrate on both the relevant task-specific and diffusion-related conditions. We also explore the trade-off between accurate lip synchronization and non-verbal facial expressions within the Diffusion paradigm. Experiments show our model achieves state-of-the-art performance on existing benchmarks, and fast inference speed owing to its ability to generate facial motions in parallel. Our code is avalable at https://github.com/theEricMa/DiffSpeaker. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Chen Qian 0006, Shukai Chen, Guo-Jun Qi, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCB | 7 |
| 2025 | DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersabstractWorld model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specialize in visual generation but lack the flexibility to incorporate other modalities like action. In contrast, autoregressive transformers have demonstrated exceptional capability in modeling multimodal data. Our work aims to unify both driving model simulation and trajectory planning into a single sequence modeling problem. We introduce a multimodal driving language based on interleaved image and action tokens, and develop DrivingGPT to learn joint world modeling and planning through standard next-token prediction. Our DrivingGPT demonstrates strong performance in both action-conditioned video generation and end-to-end planning, outperforming strong baselines on large-scale nuPlan and NAVSIM benchmarks. Yuntao Chen, Yuqi Wang 0001, Zhaoxiang Zhang 0001 |
ICCV | 3 |
| 2025 | DexVLG: Dexterous Vision-Language-Grasp Model at ScaleabstractAs large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large models for human-like dexterous hands. In this paper, we introduce DexVLG, a large Vision-Language-Grasp model for Dexterous grasp pose prediction aligned with language instructions using single-view RGBD input. To accomplish this, we generate a dataset of 170 million dexterous grasp poses mapped to semantic parts across 174,000 objects in simulation, paired with detailed part-level captions. This large-scale dataset, named DexGraspNet 3.0, is used to train a VLM and flow-matching-based pose head capable of producing instruction-aligned grasp poses for tabletop objects. To assess DexVLG's performance, we create benchmarks in physics-based simulations and conduct real-world experiments. Extensive testing demonstrates DexVLG's strong zero-shot generalization capabilities-achieving over 76% zero-shot execution success rate and state-of-the-art part-grasp accuracy in simulation-and successful part-aligned grasps on physical objects in real-world scenarios. Jiawei He 0002, Danshi Li, Xinqiang Yu, Zekun Qi, Jiayi Chen 0003, Zhaoxiang Zhang 0001, Zhizheng Zhang 0011, Li Yi 0001, He Wang 0010 |
ICCV | 7 |
| 2025 | UIPro: Unleashing Superior Interaction Capability for GUI AgentsabstractBuilding autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing methods have tried developing GUI agents based on the multi-modal comprehension ability of vision-language models (VLMs). However, the limited scenario, insufficient size, and heterogeneous action spaces hinder the progress of building generalist GUI agents. To resolve these issues, this paper proposes \textbf{UIPro}, a novel generalist GUI agent trained with extensive multi-platform and multi-task GUI interaction data, coupled with a unified action space. We first curate a comprehensive dataset encompassing 20.6 million GUI understanding tasks to pre-train UIPro, granting it a strong GUI grounding capability, which is key to downstream GUI agent tasks. Subsequently, we establish a unified action space to harmonize heterogeneous GUI agent task datasets and produce a merged dataset to foster the action prediction ability of UIPro via continued fine-tuning. Experimental results demonstrate UIPro's superior performance across multiple GUI task benchmarks on various platforms, highlighting the effectiveness of our approach. Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
ICCV | 7 |
| 2025 | End-to-End Driving with Online Trajectory Evaluation via BEV World Model
Yingyan Li, Yuqi Wang 0001, Yang Liu 0347, Jiawei He 0002, Lue Fan, Zhaoxiang Zhang 0001 |
ICCV | 6 |
| 2025 | MCOP: Multi-UAV Collaborative Occupancy PredictionabstractUnmanned Aerial Vehicle (UAV) swarm systems necessitate efficient collaborative perception mechanisms for diverse operational scenarios. Current Bird's Eye View (BEV)-based approaches exhibit two main limitations: bounding-box representations fail to capture complete semantic and geometric information of the scene, and their performance significantly degrades when encountering undefined or occluded objects. To address these limitations, we propose a novel multi-UAV collaborative occupancy prediction framework. Our framework effectively preserves 3D spatial structures and semantics through integrating a Spatial-Aware Feature Encoder and Cross-Agent Feature Integration. To enhance efficiency, we further introduce Altitude-Aware Feature Reduction to compactly represent scene information, along with a Dual-Mask Perceptual Guidance mechanism to adaptively select features and reduce communication overhead. Due to the absence of suitable benchmark datasets, we extend three datasets for evaluation: two virtual datasets (Air-to-Pred-Occ and UAV3D-Occ) and one real-world dataset (GauUScene-Occ). Experiments results demonstrate that our method achieves state-of-the-art accuracy, significantly outperforming existing collaborative methods while reducing communication overhead to only a fraction of previous approaches. Zefu Lin, Xiaojuan Jin, Yuran Yang, Lue Fan, Zhaoxiang Zhang 0001 |
ICCV | 8 |
| 2025 | Ross3d: Reconstructive Visual Instruction Tuning With 3D-Awareness
Tiancai Wang, Haoqiang Fan, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001 |
ICCV | 6 |
| 2025 | LayerAnimate: Layer-Level Control for Animation
Yuxue Yang, Lue Fan, Zuzeng Lin, Feng Wang 0015, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2025 | CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale ScenesabstractRecently, 3D Gaussian Splatting (3DGS) has revolutionized radiance field reconstruction, manifesting efficient and high-fidelity novel view synthesis. However, accurately representing surfaces, especially in large and complex scenarios, remains a significant challenge due to the unstructured nature of 3DGS. In this paper, we present CityGaussianV2, a novel approach for large-scale scene reconstruction that addresses critical challenges related to geometric accuracy and efficiency. Building on the favorable generalization capabilities of 2D Gaussian Splatting (2DGS), we address its convergence and scalability issues. Specifically, we implement a decomposed-gradient-based densification and depth regression technique to eliminate blurry artifacts and accelerate convergence. To scale up, we introduce an elongation filter that mitigates Gaussian count explosion caused by 2DGS degeneration. Furthermore, we optimize the CityGaussian pipeline for parallel training, achieving up to 10$\times$ compression, at least 25\% savings in training time, and a 50\% decrease in memory usage. We also established standard geometry benchmarks under large-scale scenes. Experimental results demonstrate that our method strikes a promising balance between visual quality, geometric accuracy, as well as storage and training costs. Yang Liu 0347, Chuanchen Luo, Zhongkai Mao, Junran Peng, Zhaoxiang Zhang 0001 |
ICLR | 5 |
| 2025 | McEval: Massively Multilingual Code EvaluationabstractCode large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval. Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001 |
ICLR | 18 |
| 2025 | Enhancing End-to-End Autonomous Driving with Latent World ModelabstractIn autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we develop better scene feature representations to fully leverage sensor data in end-to-end driving? Self-supervised learning methods show great success in learning rich feature representations in NLP and computer vision. Inspired by this, we propose a novel self-supervised learning approach using the LAtent World model (LAW) for end-to-end driving. LAW predicts future latent scene features based on current features and ego trajectories. This self-supervised task can be seamlessly integrated into perception-free and perception-based frameworks, improving scene feature learning while optimizing trajectory prediction. LAW achieves state-of-the-art performance across multiple benchmarks, including real-world open-loop benchmark nuScenes, NAVSIM, and simulator-based closed-loop benchmark CARLA. The code will be released. Yingyan Li, Lue Fan, Jiawei He 0002, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001, Tieniu Tan |
ICLR | 6 |
| 2025 | FreeVS: Generative View Synthesis on Free Driving TrajectoryabstractExisting reconstruction-based novel view synthesis methods for driving scenes focus on synthesizing camera views along the recorded trajectory of the ego vehicle.
Their image rendering performance will severely degrade on viewpoints falling out of the recorded trajectory, where camera rays are untrained.
We propose FreeVS, a novel fully generative approach that can synthesize camera views on free new trajectories in real driving scenes.
To control the generation results to be 3D consistent with the real scenes and accurate in viewpoint pose, we propose the pseudo-image representation of view priors to control the generation process.
Viewpoint translation simulation is applied on pseudo-images to simulate camera movement in each direction.
Once trained, FreeVS can be applied to any validation sequences without reconstruction process and synthesis views on novel trajectories.
Moreover, we propose two new challenging benchmarks tailored to driving scenes, which are novel camera synthesis and novel trajectory synthesis, emphasizing the freedom of viewpoints.
Given that no ground truth images are available on novel trajectories, we also propose to evaluate the consistency of images synthesized on novel trajectories with 3D perception models.
Experiments on the Waymo Open Dataset show that FreeVS has a strong image synthesis performance on both the recorded trajectories and novel trajectories.
The code is released. Project page: https://freevs24.github.io/. Qitai Wang, Lue Fan, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
ICLR | 5 |
| 2025 | MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsabstractLarge Language Models (LLMs) have displayed massive improvements in reason- ing and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing datasets have the following limitations: (1). Insufficient evaluation scenarios (e.g., only cover limited tool-use scenes). (2). Extensive evaluation costs (e.g., GPT API costs). To address these limitations, in this work, we propose a multi-granularity tool-use benchmark for large language models called MTU-Bench. For the "multi-granularity" property, our MTU-Bench covers five tool usage scenes (i.e., single-turn and single-tool, single-turn and multiple-tool, multiple-turn and single-tool, multiple-turn and multiple-tool, and out-of-distribution tasks). Besides, all evaluation metrics of our MTU-Bench are based on the prediction results and the ground truth without using any GPT or human evaluation metrics. Moreover, our MTU-Bench is collected by transforming existing high-quality datasets to simulate real-world tool usage scenarios, and we also propose an instruction dataset called MTU-Instruct data to enhance the tool-use abilities of existing LLMs. Comprehensive experimental results demonstrate the effectiveness of our MTU-Bench. Noah Wang, Xiaoshuai Song, Z. Y. Peng, Ken Deng, Jiakai Wang, Junran Peng, Ge Zhang 0009, Hangyu Guo, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007 |
ICLR | 13 |
| 2025 | Reconstructive Visual Instruction TuningabstractThis paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to conventional visual instruction tuning approaches that exclusively supervise text outputs, ROSS prompts LMMs to supervise visual outputs via reconstructing input images. By doing so, it capitalizes on the inherent richness and detail present within input images themselves, which are often lost in pure text supervision. However, producing meaningful feedback from natural images is challenging due to the heavy spatial redundancy of visual signals. To address this issue, ROSS employs a denoising objective to reconstruct latent representations of input images, avoiding directly regressing exact raw RGB values. This intrinsic activation design inherently encourages LMMs to maintain image detail, thereby enhancing their fine-grained comprehension capabilities and reducing hallucinations. Empirically, ROSS consistently brings significant improvements across different visual encoders and language models. In comparison with extrinsic assistance state-of-the-art alternatives that aggregate multiple visual experts, ROSS delivers competitive performance with a single SigLIP visual encoder, demonstrating the efficacy of our vision-centric supervision tailored for visual outputs. The code will be made publicly available upon acceptance. Anlin Zheng, Tiancai Wang, Zheng Ge, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001 |
ICLR | 7 |
| 2025 | Language-Conditioned Waypoint Predictor for Continuous Vision-and-Language NavigationabstractWaypoint prediction is a popular technique for Vision-and-Language Navigation in Continuous Environments (VLN-CE), which abstracts navigable locations as waypoints to ease the subsequent action prediction. Nevertheless, we found current waypoint predictors are not always accurate, limiting navigation’s overall performance. One possible reason may be the lack of language context, leading to the failure to generate corresponding waypoints for critical locations mentioned in the instructions. To that end, we propose a novel framework to enable the training of the language-conditioned waypoint predictor. First, as the VLN-CE agents ground instructions with the environment when navigating, we employ a pre-trained agent to encode language for the waypoint predictor. Second, the language-conditioned waypoint predictor is trained with the data collected using the same agent. Third, we train the new VLN-CE navigation agent with the proposed waypoint predictor. Fourth, the disparity between the language encoder agent and the navigation agent drives us to devise a cycle training scheme to alternately train the agent and the waypoint predictor, further enhancing the performance of both the waypoint predictor and navigation agent. Experimental results show that our waypoint predictor’s performance surpasses all existing ones. With better waypoints, the gap between waypoint-based methods and their upper bound narrows by about 60%. Yuankai Qi, Xu Yang 0004, Zhaoxiang Zhang 0001 |
ICME | 6 |
| 2025 | Top-Down Guidance for Learning Object-Centric RepresentationsabstractHumans' innate ability to decompose scenes into objects allows for efficient understanding, predicting, and planning. In light of this, Object-Centric Learning (OCL) attempts to endow networks with similar capabilities, learning to represent scenes with the composition of objects. However, existing OCL models only learn through reconstructing the input images, which does not assist the model in distinguishing objects, resulting in suboptimal object-centric representations. This flaw limits current object-centric models to relatively simple downstream tasks. To address this issue, we draw on humans’ top-down vision pathway and propose Top-Down Guided Network (TDGNet), which includes a top-down pathway to improve object-centric representations. During training, the top-down pathway constructs guidance with high-level object-centric representations to optimize low-level grid features output by the backbone. While during inference, it refines object-centric representations by detecting and solving conflicts between low- and high-level features. We show that TDGNet outperforms current object-centric models on multiple datasets of varying complexity. In addition, we expand the downstream task scope of object-centric representations by applying TDGNet to the field of robotics, validating its effectiveness in downstream tasks including video prediction and visual planning. Code will be available at https://github.com/zoujunhong/RHGNet. Junhong Zou, Xiangyu Zhu 0001, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCAI | 3 |
| 2025 | OmniBench: Towards The Future of Universal Omni-Language ModelsabstractRecent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/. Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002 |
NeurIPS | 19 |
| 2025 | TC-Light: Temporally Coherent Generative Rendering for Realistic World TransferabstractIllumination and texture rerendering are critical dimensions for world-to-world transfer, which is valuable for applications including sim2real and real2real visual data scaling up for embodied AI. Existing techniques generatively re-render the input video to realize the transfer, such as video relighting models and conditioned world generation models. Nevertheless, these models are predominantly limited to the domain of training data (e.g., portrait) or fall into the bottleneck of temporal consistency and computation efficiency, especially when the input video involves complex dynamics and long durations. In this paper, we propose **TC-Light**, a novel paradigm characterized
by the proposed two-stage post optimization mechanism. Starting from the video preliminarily relighted by an inflated video relighting model, it optimizes appearance embedding in the first stage to align global illumination. Then it optimizes the proposed canonical video representation, i.e., **Unique Video Tensor (UVT)**, to align fine-grained texture and lighting in the second stage. To comprehensively evaluate performance, we also establish a long and highly dynamic video benchmark. Extensive experiments show that our method enables physically plausible re-rendering results with superior temporal coherence and low computation cost. The code and video demos are available at our [Project Page](https://dekuliutesla.github.io/tclight/). Yang Liu 0347, Chuanchen Luo, Zimo Tang, Yingyan Li, Yuran Yang, Yuanyong Ning, Lue Fan, Junran Peng, Zhaoxiang Zhang 0001 |
NeurIPS | 9 |
| 2025 | MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsabstractThe advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research. Yuanxing Zhang, Noah Wang, Ge Zhang 0009, Jian Yang 0037, Yanghai Wang, Xintao Wang 0002, Houyi Li, Wei Ji 0011, Pengfei Wan 0001, Wenhao Huang 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 15 |
| 2025 | DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving has substantially progressed by directly predicting future trajectories from raw perception inputs, which bypasses traditional modular pipelines. However, mainstream methods trained via imitation learning suffer from critical safety limitations, as they fail to distinguish between trajectories that appear human-like but are potentially unsafe. Some recent approaches attempt to address this by regressing multiple rule-driven scores but decoupling supervision from policy optimization, resulting in suboptimal performance. To tackle these challenges, we propose DriveDPO, a Safety Direct Preference Optimization Policy Learning framework. First, we distill a unified policy distribution from human imitation similarity and rule-based safety scores for direct policy optimization. Further, we introduce an iterative Direct Preference Optimization stage formulated as trajectory-level preference alignment. Extensive experiments on the NAVSIM benchmark demonstrate that DriveDPO achieves a new state-of-the-art PDMS of 90.0. Furthermore, qualitative results across diverse challenging scenarios highlight DriveDPO’s ability to produce safer and more reliable driving behaviors. Shuyao Shang, Yuntao Chen, Yuqi Wang 0001, Yingyan Li, Zhaoxiang Zhang 0001 |
NeurIPS | 5 |
| 2025 | KORGym: A Dynamic Game Platform for LLM Reasoning EvaluationabstractRecent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments. Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009 |
NeurIPS | 26 |
| 2025 | KansformerEPI: a deep learning framework integrating KAN and transformer for predicting enhancer-promoter interactionsabstractEnhancer-promoter interaction (EPI) is a critical component of gene regulation. Accurately predicting EPIs across diverse cell types can advance our understanding of the molecular mechanisms behind transcriptional regulation and provide valuable insights into the onset and progression of related diseases. At present, large-scale genome-wide EPI predictions typically rely on computational approaches. However, most of these methods focus on predicting EPIs within a single cell line and lack a global perspective encompassing multiple cell lines. Furthermore, they often fail to fully account for the nonlinear relationships between features, leading to suboptimal prediction accuracy. In this study, we propose KansformerEPI, a global EPI prediction model designed for multiple cell lines. The model is built on Kansformer, an encoder that integrates KAN and Transformer, effectively capturing the nonlinear relationships among various epigenetic and sequence features. We utilized KansformerEPI to achieve cross-tissue prediction of EPIs across different cell types. This approach enhances the model's scalability, eliminating the complexity of designing separate prediction models for individual tissues. As a result, our model is applicable to various tissues, thereby reducing dependency on extensive datasets. Experimental results demonstrate that KansformerEPI surpasses existing methods such as TransEPI, TargetFinder, and SPEID in both accuracy and stability of EPI predictions across datasets including HMEC, IMR90, K562, and NHEK. Saihong Shao, Zhongqian Zhao, Xingjie Zhao, Zhaoxiang Zhang 0001, Guohua Wang 0001 |
Briefings Bioinform. | 6 |
| 2025 | ROLA: real-world object-centric learning with attention optimization
Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
Sci. China Inf. Sci. | 5 |
| 2025 | Pulling Target to Source: A New Perspective on Domain Adaptive Semantic Segmentation
Yujun Shen, Jingjing Fei, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | Using Unreliable Pseudo-Labels for Label-Efficient Semantic Segmentation
Yujun Shen, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | FSD V2: Improving Fully Sparse 3D Object Detection With Virtual VoxelsabstractLiDAR-based fully sparse architecture has gained increasing attention. FSDv1 stands out as a representative work, achieving impressive efficacy and efficiency, albeit with intricate structures and handcrafted designs. In this paper, we present FSDv2, an evolution that aims to simplify the previous FSDv1 and eliminate the ad-hoc heuristics in its handcrafted instance-level representation, thus promoting better universality. To this end, we introduce virtual voxels, taking over the clustering-based instance segmentation in FSDv1. Virtual voxels not only address the notorious issue of the Center Feature Missing in fully sparse detectors but also endow the framework with a more elegant and streamlined approach. Besides, we develop a suite of components to complement the virtual voxel mechanism, including a virtual voxel encoder, a virtual voxel mixer, and a virtual voxel assignment strategy. We conduct experiments on three large-scale datasets: Waymo Open Dataset, Argoverse 2 dataset, and nuScenes dataset. Our results showcase state-of-the-art performance on all three datasets, highlighting the superiority of FSDv2 in long-range scenarios and its universality in achieving competitive performance across diverse scenarios. Moreover, we provide comprehensive experimental analysis to understand the workings of FSDv2. To facilitate further research, we have open-sourced the full code at https://github.com/tusen-ai/SST. Lue Fan, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Uncertain Object Representation for Image-Based 3D Object PerceptionabstractDue to the ill-posed nature of locating 3D objects based on image inputs, objects detected by camera-based detectors tend to have considerable uncertainty in their localization. Previous works in camera-based 3D detection and tracking represent each detected object as a single certain 3D bounding box, ignoring their localization uncertainty. We propose the uncertain representation of 3D objects to meet the indeterminacy of localizing objects in images. We model the localization uncertainty of objects during the detection process and represent the location of objects as a probability distribution in 3D space. For camera-based 3D detection, we propose to gather and suppress redundant predictions about an object to form its uncertain representation. For camera-based 3D multiple object tracking, we generalize the cross-frame association metric under the uncertain representation of objects for better-tracking objects with uncertain and unstable localization. As a plug-in module for camera 3D detectors, our proposed method brings a +3.5%/+3.2%/+3.7% NDS boost to BEVDet4D/BEVDet4D-Depth/DD3D on nuScenes validation set and a +4.7% NDS boost to BEVDet4D-Depth on nuScenes test set. With enhanced cross-frame association, our tracking method achieves a 48.2% AMOTA performance and reduces the remaining identity-switch cases to only 300 on nuScenes test set. Qitai Wang, Yuntao Chen, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Bootstrap Masked Visual Modeling via Hard Patch MiningabstractMasked visual modeling has attracted much attention due to its promising potential in learning generalizable representations. Typical approaches urge models to predict specific contents of masked tokens, which can be intuitively considered as teaching a student (the model) to solve given problems (predicting masked contents). Under such settings, the performance is highly correlated with mask strategies (the difficulty of provided problems). We argue that it is equally important for the model to stand in the shoes of a teacher to produce challenging problems by itself. Intuitively, patches with high values of reconstruction loss can be regarded as hard samples, and masking those hard patches naturally becomes a demanding reconstruction task. To empower the model as a teacher, we propose Hard Patch Mining (HPM), predicting patch-wise losses and subsequently determining where to mask. Technically, we introduce an auxiliary loss predictor, which is trained with a relative objective to prevent overfitting to exact loss values. To gradually guide the training procedure, we propose an easy-to-hard mask strategy. Empirically, HPM brings significant improvements under both image and video benchmarks. Interestingly, solely incorporating the extra loss prediction objective leads to better representations, verifying the efficacy of determining where is hard to reconstruct. Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Tiancai Wang, Xiangyu Zhang 0005, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | VAG: A Uniform Model for Cross-Modal Visual-Audio Mutual GenerationabstractConsidering both audio and visual modalities is helpful for understanding a video. In the face of harsh environmental interference or signal packet loss, automatically compensating for audio and vision is a challenging task. We propose a dynamic cross-modal visual-audio mutual generation model (VAMG), which includes audio to visual conversion, visual to audio conversion, audio self-generation, and visual self-generation. VAMG jointly optimizes modal reconstruction and adversarial constraints, effectively solving the problems of structural alignment and signal compensation in incomplete videos. We conducted an instrument-oriented and pose-oriented cross-modal audio-visual mutual generation experiment on the sub-University of Rochester Musical Performance dataset to verify the effectiveness of the model. He Guan, Zhaoxiang Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Fully Data-Driven Pseudo Label Estimation for Pointly-Supervised Panoptic SegmentationabstractThe core of pointly-supervised panoptic segmentation is estimating accurate dense pseudo labels from sparse point labels to train the panoptic head. Previous works generate pseudo labels mainly based on hand-crafted rules, such as connecting multiple points into polygon masks, or assigning the label information of labeled pixels to unlabeled pixels based on the artificially defined traversing distance. The accuracy of pseudo labels is limited by the quality of the hand-crafted rules (polygon masks are rough at object contour regions, and the traversing distance error will result in wrong pseudo labels). To overcome the limitation of hand-crafted rules, we estimate pseudo labels with a fully data-driven pseudo label branch, which is optimized by point labels end-to-end and predicts more accurate pseudo labels than previous methods. We also train an auxiliary semantic branch with point labels, it assists the training of the pseudo label branch by transferring semantic segmentation knowledge through shared parameters. Experiments on Pascal VOC and MS COCO demonstrate that our approach is effective and shows state-of-the-art performance compared with related works. Codes are available at https://github.com/BraveGroup/FDD. Jing Li 0112, Junsong Fan, Yuran Yang, Shuqi Mei, Jun Xiao 0005, Zhaoxiang Zhang 0001 |
AAAI | 6 |
| 2024 | Compositional Inversion for Stable Diffusion ModelsabstractInversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. It stems from the fact that during inversion, the irrelevant semantics in the user images are also encoded, forcing the inverted concepts to occupy locations far from the core distribution in the embedding space. To address this issue, we propose a method that guides the inversion process towards the core distribution for compositional embeddings. Additionally, we introduce a spatial regularization approach to balance the attention on the concepts being composed. Our method is designed as a post-training approach and can be seamlessly integrated with other inversion methods. Experimental results demonstrate the effectiveness of our proposed approach in mitigating the overfitting problem and generating more diverse and balanced compositions of concepts in the synthesized images. The source code is available at https://github.com/zhangxulu1996/Compositional-Inversion. Xulu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001 |
AAAI | 5 |
| 2024 | Driving Into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingabstractIn autonomous driving, predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to better plan their actions, enhancing safety and efficiency on the road. To this end, we propose Drive-Wm, the first driving world model compatible with existing end-to-end planning models. Through a joint spatial-temporal modeling facilitated by view factorization, our model generates high-fidelity multiview videos in driving scenes. Building on its powerful generation ability, we showcase the potential of applying the world model for safe driving planning for the first time. Particularly, our Drive-Wm enables driving into multiple futures based on distinct driving maneuvers, and determines the optimal trajectory according to the image-based rewards. Evaluation on real-world driving datasets verifies that our method could generate high-quality, consistent, and controllable multiview videos, opening up possibilities for real-world simulations and safe planning. Yuqi Wang 0001, Jiawei He 0002, Lue Fan, Yuntao Chen, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2024 | Continual Forgetting for Pre-Trained Vision ModelsabstractFor privacy and security concerns, the need to erase un-wanted information from pre-trained vision models is becoming evident nowadays. In real-world scenar-ios, erasure requests originate at any time from both users and model owners. These requests usually form a sequence. Therefore, under such a setting, selective information is expected to be continuously removed from a pre-trained model while maintaining the rest. We define this problem as continual forgetting and identify two key challenges. (i) For unwanted knowledge, efficient and effective deleting is crucial. (ii) For remaining knowledge, the impact brought by the forgetting procedure should be minimal. To address them, we propose Group Sparse LoRA (GS-LoRA). Specifically, towards (i), we use LoRA modules to fine-tune the FFN layers in Transformer blocks for each forgetting task independently, and towards (ii), a simple group sparse regularization is adopted, enabling automatic selection of specific LoRA groups and zeroing out the others. GS-LoRA is effective, parameter-efficient, data-efficient, and easy to implement. We conduct extensive experiments on face recognition, object detection and image classification and demonstrate that GS-LoRA manages to forget specific classes with minimal impact on other classes. Codes will be released on https://github.com/bjzhb666/GS-LoRA. Hongbo Zhao 0006, Bolin Ni, Junsong Fan, Yuxi Wang 0001, Yuntao Chen, Gaofeng Meng, Zhaoxiang Zhang 0001 |
CVPR | 7 |
| 2024 | Robust Depth Enhancement via Polarization Prompt Fusion TuningabstractExisting depth sensors are imperfect and may provide inaccurate depth values in challenging scenarios, such as in the presence of transparent or reflective objects. In this work, we present a general framework that leverages polarization imaging to improve inaccurate depth measurements from various depth sensors. Previous polarization-based depth enhancement methods focus on utilizing pure physics-based formulas for a single sensor. In contrast, our method first adopts a learning-based strategy where a neural network is trained to estimate a dense and complete depth map from polarization data and a sensor depth map from different sensors. To further improve the performance, we propose a Polarization Prompt Fusion Tuning (PPFT) strategy to effectively utilize RGB-based models pre-trained on large-scale datasets, as the size of the polarization dataset is limited to train a strong model from scratch. We conducted extensive experiments on a public dataset, and the results demonstrate that the proposed method performs favorably compared to existing depth enhancement baselines. Code and demos are available at https://lastbasket.github.io/PPFT/. Kei Ikemura, Yiming Huang 0007, Felix Heide, Zhaoxiang Zhang 0001, Qifeng Chen 0001, Chenyang Lei |
CVPR | 4 |
| 2024 | MemoNav: Working Memory Model for Visual NavigationabstractImage-goal navigation is a challenging task that requires an agent to navigate to a goal indicated by an image in unfamiliar environments. Existing methods utilizing diverse scene memories suffer from inefficient exploration since they use all historical observations for decision-making without considering the goal-relevant fraction. To address this limitation, we present MemoNav, a novel memory model for image-goal navigation, which utilizes a working memory-inspired pipeline to improve navigation performance. Specifically, we employ three types of navigation memory. The node features on a map are stored in the short-term memory (STM), as these features are dynamically updated. A forgetting module then retains the informative STM fraction to increase efficiency. We also introduce long-term memory (LTM) to learn global scene representations by progressively aggregating STM features. Subsequently, a graph attention module encodes the retained STM and the LTM to generate working memory (WM) which contains the scene features essential for efficient navigation. The synergy among these three memory types boosts navigation performance by enabling the agent to learn and leverage goal-relevant scene features within a topological map. Our evaluation on multi-goal tasks demonstrates that MemoNav significantly outperforms previous methods across all difficulty levels in both Gibson and Matterport3D scenes. Qualitative results further illustrate that MemoNav plans more efficient routes. Xu Yang 0004, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2024 | HardMo: A Large-Scale Hardcase Dataset for Motion CaptureabstractRecent years have witnessed rapid progress in monoc-ular human mesh recovery. Despite their impressive performance on public benchmarks, existing methods are vulnerable to unusual poses, which prevents them from deploying to challenging scenarios such as dance and martial arts. This issue is mainly attributed to the domain gap induced by the data scarcity in relevant cases. Most existing datasets are captured in constrained scenarios and lack samples of such complex movements. For this reason, we propose a data collection pipeline comprising automatic crawling, precise annotation, and hardcase mining. Based on this pipeline, we establish a large dataset in a short time. The dataset, named HardMo, contains 7M images along with precise annotations covering 15 categories of dance and 14 categories of martial arts. Empirically, we find that the prediction failure in dance and martial arts is mainly characterized by the misalignment of hand-wrist and foot-ankle. To dig deeper into the two hardcases, we leverage the proposed automatic pipeline to filter collected data and construct two subsets named HardMo-Hand and HardMo-Foot. Extensive experiments demonstrate the effectiveness of the annotation pipeline and the data-driven solution to failure cases. Specifically, after being trained on HardMo, HMR, an early pioneering method, can even outperform the current state of the art, 4DHumans, on our benchmarks. Dataset will be publicly available at https://ljqnb.github.io/HardMo.github.io. Jiaqi Liao, Chuanchen Luo, Yinuo Du, Yuxi Wang 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng |
CVPR | 7 |
| 2024 | Enhancing Visual Continual Learning with Language-Guided SupervisionabstractContinual learning (CL) aims to empower models to learn new tasks without forgetting previously acquired knowledge. Most prior works concentrate on the techniques of architectures, replay data, regularization, etc. However, the category name of each class is largely neglected. Existing methods commonly utilize the one-hot labels and randomly initialize the classifier head. We argue that the scarce semantic information conveyed by the one-hot labels hampers the effective knowledge transfer across tasks. In this paper, we revisit the role of the classifier head within the CL paradigm and replace the classifier with semantic knowledge from pretrained language models (PLMs). Specifically, we use PLMs to generate semantic targets for each class, which are frozen and serve as supervision signals during training. Such targets fully consider the semantic correlation between all classes across tasks. Empirical studies show that our approach mitigates forgetting by alleviating representation drifting and facilitating knowledge transfer across tasks. The proposed method is simple to implement and can seamlessly be plugged into existing methods with negligible adjustments. Extensive experiments based on eleven mainstream baselines demonstrate the effectiveness and generalizability of our approach to various protocols. For example, under the class-incremental learning setting on ImageNet-100, our method significantly improves the Top-1 accuracy by 3.2% to 6.1% while reducing the forgetting rate by 2.6% to 13.1%. Bolin Ni, Hongbo Zhao 0006, Chenghao Zhang 0003, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang |
CVPR | 6 |
| 2024 | PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic SegmentationabstractComprehensive modeling of the surrounding 3D world is crucial for the success of autonomous driving. However, existing perception tasks like object detection, road structure segmentation, depth & elevation estimation, and open-set object localization each only focus on a small facet of the holistic 3D scene understanding task. This divide-and-conquer strategy simplifies the algorithm development process but comes at the cost of losing an end-to-end unified solution to the problem. In this work, we address this limitation by studying camera-based 3D panoptic segmentation, aiming to achieve a unified occupancy representation for camera-only 3D scene understanding. To achieve this, we introduce a novel method called PanoOcc, which utilizes voxel queries to aggregate spatiotemporal information from multi-frame and multi-view images in a coarse-to-fine scheme, integrating feature learning and scene representation into a unified occupancy representation. We have conducted extensive ablation studies to validate the effectiveness and efficiency of the proposed method. Our approach achieves new state-of-the-art results for camera-based semantic segmentation and panoptic segmentation on the nuScenes dataset. Furthermore, our method can be easily extended to dense occupancy prediction and has demonstrated promising performance on the Occ3D benchmark. The code will be made available at https://github.com/Robertwyq/PanoOcc. Yuqi Wang 0001, Yuntao Chen, Xingyu Liao, Lue Fan, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2024 | RCL: Reliable Continual Learning for Unified Failure DetectionabstractDeep neural networks are known to be overconfident for what they don't know in the wild, which is undesirable for decision-making in high-stakes applications. Despite quan-tities of existing works, most of them focus on detecting out-of-distribution (OOD) samples from unseen classes, while ignoring large parts of relevant failure sources like mis-classified samples from known classes. In particular, recent studies reveal that prevalent OOD detection methods are actually harmful for misclassification detection (MisD), indicating that there seems to be a tradeoff between those two tasks. In this paper, we study the critical yet under-explored problem of unified failure detection, which aims to detect both misclassified and OOD examples. Concretely, we identify the failure of simply integrating learning objectives of misclassification and OOD detection, and show the potential of sequence learning. Inspired by this, we propose a reliable continual learning paradigm, whose spirit is to equip the model with MisD ability first, and then improve the OOD detection ability without degrading the al-ready adequate MisD performance. Extensive experiments demonstrate that our method achieves strong unified failure detection performance. The code is available at https://github.com/Impression2805/RCL. Fei Zhu 0004, Zhen Cheng 0003, Xu-Yao Zhang, Cheng-Lin Liu 0001, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2024 | Expanding Scene Graph Boundaries: Fully Open-Vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
Zuyao Chen, Jinlin Wu, Zhen Lei 0001, Zhaoxiang Zhang 0001, Chang Wen Chen |
ECCV (66) | 4 |
| 2024 | Point-Supervised Panoptic Segmentation via Estimating Pseudo Labels from Learnable Distance
Jing Li 0112, Junsong Fan, Zhaoxiang Zhang 0001 |
ECCV (16) | 3 |
| 2024 | CityGaussian: Real-Time High-Quality Large-Scale Scene Rendering with Gaussians
Yang Liu 0347, Chuanchen Luo, Lue Fan, Naiyan Wang, Junran Peng, Zhaoxiang Zhang 0001 |
ECCV (16) | 6 |
| 2024 | OneTrack: Demystifying the Conflict Between Detection and Tracking in End-to-End 3D Trackers
Qitai Wang, Jiawei He 0002, Yuntao Chen, Zhaoxiang Zhang 0001 |
ECCV (7) | 4 |
| 2024 | Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
Pengfei Wang 0012, Yuxi Wang 0001, Shuai Li 0014, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (15) | 4 |
| 2024 | Monocular Occupancy Prediction for Scalable Indoor Scenes
Hongxiao Yu, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
ECCV (30) | 4 |
| 2024 | CSOT: Cross-scan Object Transfer for Semi-Supervised LiDAR Object Detection
Jinglin Zhan, Tiejun Liu, RenGang Li, Zhaoxiang Zhang 0001, Yuntao Chen |
ECCV (17) | 4 |
| 2024 | General Geometry-Aware Weakly Supervised 3D Object Detection
Guowen Zhang, Junsong Fan, Liyi Chen 0002, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (51) | 4 |
| 2024 | MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object DetectionabstractLabel-efficient LiDAR-based 3D object detection is currently dominated by weakly/semi-supervised methods. Instead of exclusively following one of them, we propose MixSup, a more practical paradigm simultaneously utilizing massive cheap coarse labels and a limited number of accurate labels for Mixed-grained Supervision. We start by observing that point clouds are usually textureless, making it hard to learn semantics. However, point clouds are geometrically rich and scale-invariant to the distances from sensors, making it relatively easy to learn the geometry of objects, such as poses and shapes. Thus, MixSup leverages massive coarse cluster-level labels to learn semantics and a few expensive box-level labels to learn accurate poses and shapes. We redesign the label assignment in mainstream detectors, which allows them seamlessly integrated into MixSup, enabling practicality and universality. We validate its effectiveness in nuScenes, Waymo Open Dataset, and KITTI, employing various detectors. MixSup achieves up to 97.31% of fully supervised performance, using cheap cluster annotations and only 10% box annotations. Furthermore, we propose PointSAM based on the Segment Anything Model for automated coarse labeling, further reducing the annotation burden. The code is available at https://github.com/BraveGroup/PointSAM-for-MixSup. Yuxue Yang, Lue Fan, Zhaoxiang Zhang 0001 |
ICLR | 3 |
| 2024 | StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation FrameworkabstractThanks to the powerful generative capacity of diffusion models, recent years have witnessed rapid progress in human motion generation. Existing diffusion-based methods employ disparate network architectures and training strategies. The effect of the design of each component is still unclear. In addition, the iterative denoising process consumes considerable computational overhead, which is prohibitive for real-time scenarios such as virtual characters and humanoid robots. For this reason, we first conduct a comprehensive investigation into network architectures, training strategies, and inference process. Based on the profound analysis, we tailor each component for efficient high-quality human motion generation. Despite the promising performance, the tailored model still suffers from foot skating which is an ubiquitous issue in diffusion-based solutions. To eliminate footskate, we identify foot-ground contact and correct foot motions along the denoising process. By organically combining these well-designed components together, we present StableMoFusion, a robust and efficient framework for human motion generation. Extensive experimental results show that our StableMoFusion performs favorably against current state-of-the-art methods. Chuanchen Luo, Yuxi Wang 0001, Shibiao Xu, Zhaoxiang Zhang 0001, Man Zhang 0005, Junran Peng |
ACM Multimedia | 6 |
| 2024 | MaterialSeg3D: Segmenting Dense Materials from 2D Priors for 3D AssetsabstractDriven by powerful image diffusion models, recent research has achieved the automatic creation of 3D objects from textual or visual guidance. By performing score distillation sampling (SDS) iteratively across different views, these methods succeed in lifting 2D generative prior to the 3D space. However, such a 2D generative image prior bakes the effect of illumination and shadow into the texture. As a result, material maps optimized by SDS inevitably involve spurious correlated components. The absence of precise material definition makes it infeasible to relight the generated assets reasonably in novel scenes, which limits their application in downstream scenarios. In contrast, humans can effortlessly circumvent this ambiguity by deducing the material of the object from its appearance and semantics. Motivated by this insight, we propose MaterialSeg3D, a 3D asset material generation framework to infer underlying material from the 2D semantic prior. Based on such a prior model, we devise a mechanism to parse material in 3D space. We maintain a UV stack, each map of which is unprojected from a specific viewpoint. After traversing all viewpoints, we fuse the stack through a weighted voting scheme and then employ region unification to ensure the coherence of the object parts. To fuel the learning of semantics prior, we collect a material dataset, named Materialized Individual Objects (MIO), which features abundant images, diverse categories, and accurate annotations. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method. Ruitong Gan, Chuanchen Luo, Yuxi Wang 0001, Qing Li 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng |
ACM Multimedia | 10 |
| 2024 | Generative Active Learning for Image Synthesis PersonalizationabstractThis paper presents a pilot study that explores the application of active learning, traditionally studied in the context of discriminative models, to generative models. We specifically focus on image synthesis personalization tasks. The primary challenge in conducting active learning on generative models lies in the open-ended nature of querying, which differs from the closed form of querying in discriminative models that typically target a single concept. We introduce the concept of anchor directions to transform the querying process into a semi-open problem. We propose a direction-based uncertainty sampling strategy to enable generative active learning and tackle the exploitation-exploration dilemma. Extensive experiments are conducted to validate the effectiveness of our approach, demonstrating that an open-source model can achieve superior performance compared to closed-source models developed by large companies, such as Google's StyleDrop. The source code is available at https://github.com/zhangxulu1996/GAL4Personalization. Xulu Zhang, Wengyu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001 |
ACM Multimedia | 5 |
| 2024 | STODINE: Decompose video to Object-centric Spatial-Temporal Slots for physical reasoning
Xiangyu Zhu 0001, Qu Tang, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
MMAsia | 4 |
| 2024 | OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map ConstructionabstractIn this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving. Hongbo Zhao 0006, Lue Fan, Yuntao Chen, Yuran Yang, Xiaojuan Jin, Gaofeng Meng, Zhaoxiang Zhang 0001 |
NeurIPS | 9 |
| 2024 | RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from ScriptsabstractBelievable agents can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication. Recently, generative agents have been proposed to simulate believable human behavior by using Large Language Models. However, the existing method heavily relies on human-annotated agent profiles (e.g., name, age, personality, relationships with others, and so on) for the initialization of each agent, which cannot be scaled up easily. In this paper, we propose a scalable RoleAgent framework to generate high-quality role-playing agents from raw scripts, which includes building and interacting stages. Specifically, in the building stage, we use a hierarchical memory system to extract and summarize the structure and high-level information of each agent for the raw script. In the interacting stage, we propose a novel innovative mechanism with four steps to achieve a high-quality interaction between agents. Finally, we introduce a systematic and comprehensive evaluation benchmark called RoleAgentBench to evaluate the effectiveness of our RoleAgent, which includes 100 and 28 roles for 20 English and 5 Chinese scripts, respectively. Extensive experimental results on RoleAgentBench demonstrate the effectiveness of RoleAgent. Zehao Ni, Haoran Que, Tao Sun 0016, Noah Wang, Jian Yang 0030, Jiakai Wang, Hongcheng Guo, Zhongyuan Peng, Ge Zhang 0009, Xingyuan Bu, Ke Xu 0001, Wenge Rong, Junran Peng, Zhaoxiang Zhang 0001 |
NeurIPS | 16 |
| 2024 | DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World ModelabstractDriving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving datasets. We introduce DrivingDojo, the first dataset tailor-made for training interactive world models with complex driving dynamics. Our dataset features video clips with a complete set of driving maneuvers, diverse multi-agent interplay, and rich open-world driving knowledge, laying a stepping stone for future world model development. We further define an action instruction following (AIF) benchmark for world models and demonstrate the superiority of the proposed dataset for generating action-controlled future predictions. Yuqi Wang 0001, Jiawei He 0002, Qitai Wang, Hengchen Dai, Yuntao Chen, Zhaoxiang Zhang 0001 |
NeurIPS | 8 |
| 2024 | Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object DetectionabstractSerialization-based methods, which serialize the 3D voxels and group them into multiple sequences before inputting to Transformers, have demonstrated their effectiveness in 3D object detection. However, serializing 3D voxels into 1D sequences will inevitably sacrifice the voxel spatial proximity. Such an issue is hard to be addressed by enlarging the group size with existing serialization-based methods due to the quadratic complexity of Transformers with feature sizes. Inspired by the recent advances of state space models (SSMs), we present a Voxel SSM, termed as Voxel Mamba, which employs a group-free strategy to serialize the whole space of voxels into a single sequence. The linear complexity of SSMs encourages our group-free design, alleviating the loss of spatial proximity of voxels. To further enhance the spatial proximity, we propose a Dual-scale SSM Block to establish a hierarchical structure, enabling a larger receptive field in the 1D serialization curve, as well as more complete local regions in 3D space. Moreover, we implicitly apply window partition under the group-free framework by positional encoding, which further enhances spatial proximity by encoding voxel positional information. Our experiments on Waymo Open Dataset and nuScenes dataset show that Voxel Mamba not only achieves higher accuracy than state-of-the-art methods, but also demonstrates significant advantages in computational efficiency. The source code is available at https://github.com/gwenzhang/Voxel-Mamba. Guowen Zhang, Lue Fan, Chenhang He, Zhen Lei 0001, Zhaoxiang Zhang 0001, Lei Zhang 0006 |
NeurIPS | 5 |
| 2024 | VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationabstractBird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{generating} the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. \emph{The question is how to align the PV features with the generative models to facilitate the map estimation}. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on \url{https://github.com/Z1zyw/VQ-Map}. Fudong Ge, Guan Luo, Bing Li 0001, Zhaoxiang Zhang 0001, Haibin Ling, Weiming Hu 0004 |
NeurIPS | 6 |
| 2024 | GRAMO: geometric resampling augmentation for monocular 3D object detectionabstractAbstract Data augmentation is widely recognized as an effective means of bolstering model robustness. However, when applied to monocular 3D object detection, non-geometric image augmentation neglects the critical link between the image and physical space, resulting in the semantic collapse of the extended scene. To address this issue, we propose two geometric-level data augmentation operators named Geometric-Copy-Paste (Geo-CP) and Geometric-Crop-Shrink (Geo-CS). Both operators introduce geometric consistency based on the principle of perspective projection, complementing the options available for data augmentation in monocular 3D. Specifically, Geo-CP replicates local patches by reordering object depths to mitigate perspective occlusion conflicts, and Geo-CS re-crops local patches for simultaneous scaling of distance and scale to unify appearance and annotation. These operations ameliorate the problem of class imbalance in the monocular paradigm by increasing the quantity and distribution of geometrically consistent samples. Experiments demonstrate that our geometric-level augmentation operators effectively improve robustness and performance in the KITTI and Waymo monocular 3D detection benchmarks. He Guan, Chunfeng Song, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 3 |
| 2024 | Learnable Graph Matching: A Practical Paradigm for Data AssociationabstractData association is at the core of many computer vision tasks, e.g., multiple object tracking, image matching, and point cloud registration. however, current data association solutions have some defects: they mostly ignore the intra-view context information; besides, they either train deep association models in an end-to-end way and hardly utilize the advantage of optimization-based assignment methods, or only use an off-the-shelf neural network to extract features. In this paper, we propose a general learnable graph matching method to address these issues. Especially, we model the intra-view relationships as an undirected graph. Then data association turns into a general graph matching problem between graphs. Furthermore, to make optimization end-to-end differentiable, we relax the original graph matching problem into continuous quadratic programming and then incorporate training into a deep graph neural network with KKT conditions and implicit function theorem. In MOT task, our method achieves state-of-the-art performance on several MOT datasets. For image matching, our method outperforms state-of-the-art methods on a popular indoor dataset, ScanNet. For point cloud registration, we also achieve competitive results. Jiawei He 0002, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Fully Sparse Fusion for 3D Object DetectionabstractCurrently prevalent multi-modal 3D detection methods rely on dense detectors that usually use dense Bird's-Eye-View (BEV) feature maps. However, the cost of such BEV feature maps is quadratic to the detection range, making it not scalable for long-range detection. Recently, LiDAR-only fully sparse architecture has been gaining attention for its high efficiency in long-range perception. In this paper, we study how to develop a multi-modal fully sparse detector. Specifically, our proposed detector integrates the well-studied 2D instance segmentation into the LiDAR side, which is parallel to the 3D instance segmentation part in the LiDAR-only baseline. The proposed instance-based fusion framework maintains full sparsity while overcoming the constraints associated with the LiDAR-only fully sparse detector. Our framework showcases state-of-the-art performance on the widely used nuScenes dataset, Waymo Open Dataset, and the long-range Argoverse 2 dataset. Notably, the inference speed of our proposed method under the long-range perception setting is 2.7× faster than that of other state-of-the-art multimodal 3D detection methods. Yingyan Li, Lue Fan, Yang Liu 0347, Zehao Huang, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Large-Scale Object Detection in the Wild With Imbalanced Data Distribution, and Multi-LabelsabstractTraining with more data has always been the most stable and effective way of improving performance in the deep learning era. The Open Images dataset, the largest object detection dataset, presents significant opportunities and challenges for general and sophisticated scenarios. However, its semi-automatic collection and labeling process, designed to manage the huge data scale, leads to label-related problems, including explicit or implicit multiple labels per object and highly imbalanced label distribution. In this work, we quantitatively analyze the major problems in large-scale object detection and provide a detailed yet comprehensive demonstration of our solutions. First, we design a concurrent softmax to handle the multi-label problems in object detection and propose a soft-balance sampling method with a hybrid training scheduler to address the label imbalance. This approach yields a notable improvement of 3.34 points, achieving the best single-model performance with a mAP of 60.90% on the public object detection test set of Open Images. Then, we introduce a well-designed ensemble mechanism that substantially enhances the performance of the single model, achieving an overall mAP of 67.17%, which is 4.29 points higher than the best result from the Open Images public test 2018. Cong Pan 0001, Junran Peng, Xingyuan Bu, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Enhancing Sound Source Localization via False Negative EliminationabstractSound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling negatives in prior arts can lead to the false negative issue, where the sounds semantically similar to visual instance are sampled as negatives and incorrectly pushed away from the visual anchor/query. As a result, this misalignment of audio and visual features could yield inferior performance. To address this issue, we propose a novel audio-visual learning framework which is instantiated with two individual learning schemes: self-supervised predictive learning (SSPL) and semantic-aware contrastive learning (SACL). SSPL explores image-audio positive pairs alone to discover semantically coherent similarities between audio and visual features, while a predictive coding module for feature alignment is introduced to facilitate the positive-only learning. In this regard SSPL acts as a negative-free method to eliminate false negatives. By contrast, SACL is designed to compact visual features and remove false negatives, providing reliable visual anchor and audio negatives for contrast. Different from SSPL, SACL releases the potential of audio-visual contrastive learning, offering an effective alternative to achieve the same goal. Comprehensive experiments demonstrate the superiority of our approach over the state-of-the-arts. Furthermore, we highlight the versatility of the learned representation by extending the approach to audio-visual event classification and object detection tasks. Zengjie Song, Jiangshe Zhang 0001, Yuxi Wang 0001, Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A Curriculum-Style Self-Training Approach for Source-Free Semantic SegmentationabstractSource-free domain adaptation has developed rapidly in recent years, where the well-trained source model is adapted to the target domain instead of the source data, offering the potential for privacy concerns and intellectual property protection. However, a number of feature alignment techniques in prior domain adaptation methods are not feasible in this challenging problem setting. Thereby, we resort to probing inherent domain-invariant feature learning and propose a curriculum-style self-training approach for source-free domain adaptive semantic segmentation. In particular, we introduce a curriculum-style entropy minimization method to explore the implicit knowledge from the source model, which fits the trained source model to the target data using certain information from easy-to-hard predictions. We then train the segmentation network by the proposed complementary curriculum-style self-training, which utilizes the negative and positive pseudo labels following the curriculum-learning manner. Although negative pseudo-labels with high uncertainty cannot be identified with the correct labels, they can definitely indicate absent classes. Moreover, we employ an information propagation scheme to further reduce the intra-domain discrepancy within the target domain, which could act as a standard post-processing method for the domain adaptation field. Furthermore, we extend the proposed method to a more challenging black-box source model scenario where only the source model's predictions are available. Extensive experiments validate that our method yields state-of-the-art performance on source-free semantic segmentation tasks for both synthetic-to-real and adverse conditions datasets. Yuxi Wang 0001, Jian Liang 0001, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Reusable Architecture Growth for Continual Stereo MatchingabstractThe remarkable performance of recent stereo depth estimation models benefits from the successful use of convolutional neural networks to regress dense disparity. Akin to most tasks, this needs gathering training data that covers a number of heterogeneous scenes at deployment time. However, training samples are typically acquired continuously in practical applications, making the capability to learn new scenes continually even more crucial. For this purpose, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at inference. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth to learn new scenes continually in both supervised and self-supervised manners. It can maintain high reusability during growth by reusing previous units while obtaining good performance. Additionally, we present a Scene Router module to adaptively select the scene-specific architecture path at inference. Comprehensive experiments on numerous datasets show that our framework performs impressively in various weather, road, and city circumstances and surpasses the state-of-the-art methods in more challenging cross-dataset settings. Further experiments also demonstrate the adaptability of our method to unseen scenes, which can facilitate end-to-end stereo architecture learning and practical deployment. Chenghao Zhang 0003, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Prototype learning for adversarial domain adaptation
Yuchun Fang, Chen Chen 0114, Wei Zhang 0021, Zhaoxiang Zhang 0001, Shaorong Xie |
Pattern Recognit. | 5 |
| 2024 | Large-scale continual learning for ancient Chinese character recognition
Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2024 | MA-ST3D: Motion Associated Self-Training for Unsupervised Domain Adaptation on 3D Object DetectionabstractRecently, unsupervised domain adaptation (UDA) for 3D object detectors has increasingly garnered attention as a method to eliminate the prohibitive costs associated with generating extensive 3D annotations, which are crucial for effective model training. Self-training (ST) has emerged as a simple and effective technique for UDA. The major issue involved in ST-UDA for 3D object detection is refining the imprecise predictions caused by domain shift and generating accurate pseudo labels as supervisory signals. This study presents a novel ST-UDA framework to generate high-quality pseudo labels by associating predictions of 3D point cloud sequences during ego-motion according to spatial and temporal consistency, named motion-associated self-training for 3D object detection (MA-ST3D). MA-ST3D maintains a global-local pathway (GLP) architecture to generate high-quality pseudo-labels by leveraging both intra-frame and inter-frame consistencies along the spatial dimension of the LiDAR's ego-motion. It also equips two memory modules for both global and local pathways, called global memory and local memory, to suppress the temporal fluctuation of pseudo-labels during self-training iterations. In addition, a motion-aware loss is introduced to impose discriminated regulations on pseudo labels with different motion statuses, which mitigates the harmful spread of false positive pseudo labels. Finally, our method is evaluated on three representative domain adaptation tasks on authoritative 3D benchmark datasets (i.e. Waymo, Kitti, and nuScenes). MA-ST3D achieved SOTA performance on all evaluated UDA settings and even surpassed the weakly supervised DA methods on the Kitti and NuScenes object detection benchmark. Chi Zhang 0060, Wei Wang 0353, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Visually Guided Sound Source Separation With Audio-Visual Predictive CodingabstractThe framework of visually guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An ongoing trend in this field has been to tailor involved visual feature extractor for informative visual guidance and separately devise module for feature fusion, while utilizing U-Net by default for sound analysis. However, such a divide-and-conquer paradigm is parameter-inefficient and, meanwhile, may obtain suboptimal performance as jointly optimizing and harmonizing various model components is challengeable. By contrast, this article presents a novel approach, dubbed audio-visual predictive coding (AVPC), to tackle this task in a parameter-efficient and more effective manner. The network of AVPC features a simple ResNet-based video analysis network for deriving semantic visual features, and a predictive coding (PC)-based sound separation network that can extract audio features, fuse multimodal information, and predict sound separation masks in the same architecture. By iteratively minimizing the prediction error between features, AVPC integrates audio and visual information recursively, leading to progressively improved performance. In addition, we develop a valid self-supervised learning strategy for AVPC via copredicting two audio-visual representations of the same sound source. Extensive evaluations demonstrate that AVPC outperforms several baselines in separating musical instrument sounds, while reducing the model size significantly. Code is available at: https://github.com/zjsong/Audio-Visual-Predictive-Coding. Zengjie Song, Zhaoxiang Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Robust Feature Rectification of Pretrained Vision Models for Object RecognitionabstractPretrained vision models for object recognition often suffer a dramatic performance drop with degradations unseen during training. In this work, we propose a RObust FEature Rectification module (ROFER) to improve the performance of pretrained models against degradations. Specifically, ROFER first estimates the type and intensity of the degradation that corrupts the image features. Then, it leverages a Fully Convolutional Network (FCN) to rectify the features from the degradation by pulling them back to clear features. ROFER is a general-purpose module that can address various degradations simultaneously, including blur, noise, and low contrast. Besides, it can be plugged into pretrained models seamlessly to rectify the degraded features without retraining the whole model. Furthermore, ROFER can be easily extended to address composite degradations by adopting a beam search algorithm to find the composition order. Evaluations on CIFAR-10 and Tiny-ImageNet demonstrate that the accuracy of ROFER is 5% higher than that of SOTA methods on different degradations. With respect to composite degradations, ROFER improves the accuracy of a pretrained CNN by 10% and 6% on CIFAR-10 and Tiny-ImageNet respectively. Shengchao Zhou, Gaofeng Meng, Zhaoxiang Zhang 0001, Shiming Xiang |
AAAI | 3 |
| 2023 | 3D Video Object Detection with Learnable Object-Centric Global OptimizationabstractWe explore long-term temporal visual correspondence-based optimization for 3D video object detection in this work. Visual correspondence refers to one-to-one mappings for pixels across multiple images. Correspondence-based optimization is the cornerstone for 3D scene reconstruction but is less studied in 3D video object detection, because moving objects violate multi-view geometry constraints and are treated as outliers during scene reconstruction. We address this issue by treating objects as first-class citizens during correspondence-based optimization. In this work, we propose BA-Det, an end-to-end optimizable object detector with object-centric temporal correspondence learning and featuremetric object bundle adjustment. Empirically, we verify the effectiveness and efficiency of BA-Det for multiple baseline 3D detectors under various setups. Our BA-Det achieves SOTA performance on the large-scale Waymo Open Dataset (WOD) with only marginal computation cost. Our code is available at https://github.com/jiaweihe1996/BA-Det. Jiawei He 0002, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2023 | Blind Video Deflickering by Neural Filtering with a Flawed AtlasabstractMany videos contain flickering artifacts; common causes of flicker include video processing algorithms, video generation algorithms, and capturing videos under specific situations. Prior work usually requires specific guidance such as the flickering frequency, manual annotations, or extra consistent videos to remove the flicker. In this work, we propose a general flicker removal framework that only receives a single flickering video as input without additional guidance. Since it is blind to a specific flickering type or guidance, we name this “blind deflickering.” The core of our approach is utilizing the neural atlas in cooperation with a neural filtering strategy. The neural atlas is a unified representation for all frames in a video that provides temporal consistency guidance but is flawed in many cases. To this end, a neural network is trained to mimic a filter to learn the consistent features (e.g., color, brightness) and avoid introducing the artifacts in the atlas. To validate our method, we construct a dataset that contains diverse real-world flickering videos. Extensive experiments show that our method achieves satisfying deflickering performance and even outperforms baselines that use extra guidance on a public benchmark. The source code is publicly available at https://chenyanglei.github.io/deflicker. Chenyang Lei, Xuanchi Ren, Zhaoxiang Zhang 0001, Qifeng Chen 0001 |
CVPR | 3 |
| 2023 | BAEFormer: Bi-Directional and Early Interaction Transformers for Bird's Eye View Semantic SegmentationabstractBird's Eye View (BEV) semantic segmentation is a critical task in autonomous driving. However, existing Transformer-based methods confront difficulties in transforming Perspective View (PV) to BEV due to their unidirectional and posterior interaction mechanisms. To address this issue, we propose a novel Bi-directional and Early Interaction Transformers framework named BAEFormer, consisting of (i) an early-interaction PV-BEV pipeline and (ii) a bi-directional cross-attention mechanism. Moreover, we find that the image feature maps' resolution in the cross-attention module has a limited effect on the final performance. Under this critical observation, we propose to enlarge the size of input images and downsample the multi-view image features for cross-interaction, further improving the accuracy while keeping the amount of computation controllable. Our proposed method for BEV semantic segmentation achieves state-of-the-art performance in real-time inference speed on the nuScenes dataset, i.e., 38.9 mIoU at 45 FPS on a single A100 GPU. Cong Pan 0001, Yonghao He, Junran Peng, Qian Zhang 0009, Wei Sui, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2023 | Intrinsic Physical Concepts Discovery with Object-Centric Predictive ModelsabstractThe ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in terms of objects and relations in an unsupervised manner. Recent approaches learn object-centric represen-tations and capture visually observable concepts of objects, e.g., shape, size, and location. In this paper, we take a step forward and try to discover and represent intrinsic physical concepts such as mass and charge. We introduce the PHYsi-cal Concepts Inference NEtwork (PHYCINE), a system that infers physical concepts in different abstract levels with-out supervision. The key insights underlining PHYCINE are two-fold, commonsense knowledge emerges with pre-diction, and physical concepts of different abstract levels should be reasoned in a bottom-up fashion. Empirical eval-uation demonstrates that variables inferred by our system work in accordance with the properties of the corresponding physical concepts. We also show that object representations containing the discovered physical concepts variables could help achieve better performance in causal reasoning tasks, i.e., ComPhy. Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2023 | FrustumFormer: Adaptive Instance-aware Resampling for Multi-view 3D DetectionabstractThe transformation of features from 2D perspective space to 3D space is essential to multi-view 3D object detection. Recent approaches mainly focus on the design of view transformation, either pixel-wisely lifting perspective view features into 3D space with estimated depth or grid-wisely constructing BEV features via 3D projection, treating all pixels or grids equally. However, choosing what to transform is also important but has rarely been discussed before. The pixels of a moving car are more informative than the pixels of the sky. To fully utilize the information contained in images, the view transformation should be able to adapt to different image regions according to their contents. In this paper, we propose a novel framework named FrustumFormer, which pays more attention to the features in instance regions via adaptive instance-aware resampling. Specifically, the model obtains instance frustums on the bird's eye view by leveraging image view object proposals. An adaptive occupancy mask within the instance frustum is learned to refine the instance location. Moreover, the temporal frustum intersection could further reduce the localization uncertainty of objects. Comprehensive experiments on the nuScenes dataset demonstrate the effectiveness of FrustumFormer, and we achieve a new state-of-the-art performance on the benchmark. Codes and models will be made available at https://github.com/Robertwyq/Frustum. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2023 | Hard Patches Mining for Masked Image ModelingabstractMasked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually focus on predicting specific contents of masked patches, and their performances are highly related to pre-defined mask strategies. Intuitively, this procedure can be considered as training a student (the model) on solving given problems (predict masked patches). However, we argue that the model should not only focus on solving given problems, but also stand in the shoes of a teacher to produce a more challenging problem by itself. To this end, we propose Hard Patches Mining (HPM), a brand-new framework for MIM pre-training. We observe that the reconstruction loss can naturally be the metric of the difficulty of the pretraining task. Therefore, we introduce an auxiliary loss predictor, predicting patch-wise losses first and deciding where to mask next. It adopts a relative relationship learning strategy to prevent overfitting to exact reconstruction loss values. Experiments under various settings demonstrate the effectiveness of HPM in constructing masked images. Furthermore, we empirically find that solely introducing the loss prediction objective leads to powerful representations, verifying the efficacy of the ability to be aware of where is hard to reconstruct.11Code: https://github.com/Haochen-wang409/HPM Kaiyou Song, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2023 | Sharpness-Aware Gradient Matching for Domain GeneralizationabstractThe goal of domain generalization (DG) is to enhance the generalization capability of the model learned from a source domain to other unseen domains. The recently developed Sharpness-Aware Minimization (SAM) method aims to achieve this goal by minimizing the sharpness measure of the loss landscape. Though SAM and its variants have demonstrated impressive DG performance, they may not always converge to the desired flat region with a small loss value. In this paper, we present two conditions to ensure that the model could converge to a flat minimum with a small loss, and present an algorithm, named Sharpness-Aware Gradient Matching (SAGM), to meet the two conditions for improving model generalization capability. Specifically, the optimization objective of SAGM will simultaneously minimize the empirical risk, the perturbed loss (i.e., the maximum loss within a neighborhood in the parameter space), and the gap between them. By implicitly aligning the gradient directions between the empirical risk and the perturbed loss, SAGM improves the generalization capability over SAM and its variants without increasing the computational cost. Extensive experimental results show that our proposed SAGM method consistently outperforms the state-of-the-art methods on five DG benchmarks, including PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet. Codes are available at https://github.com/Wang-pengfei/SAGM. Pengfei Wang 0009, Zhaoxiang Zhang 0001, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 2 |
| 2023 | BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective SupervisionabstractWe present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and bet-suits modern image backbones. Existing state-of-the-art BEV detectors are often tied to certain depth pretrained backbones like Vo Vn et, hindering the synergy between booming image backbones and BEV detectors. To address this limitation, we prioritize easing the optimization of BEV detectors by introducing perspective view supervision. To this end, we propose a two-stage BEV detector; where proposals from the perspective head are fed into the bird’ s-eye-view head for final predictions. To evaluate the effectiveness of our model, we conduct extensive ablation studies focusing on the form of supervision and the gener-ality of the proposed detector. The proposed method is ver-ified with a wide spectrum of traditional and modern image backbones and achieves new SoTA results on the large-scale nuScenes dataset. The code shall be released soon. Yuntao Chen, Hao Tian 0006, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang 0001, Gao Huang 0001, Hongyang Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai |
CVPR | 6 |
| 2023 | Graphics Capsule: Learning Hierarchical 3D Face Representations from 2D ImagesabstractThe function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural networks. However, their descriptions are restricted to the 2D space, limiting their capacities to imitate the intrinsic 3D perception ability of humans. In this paper, we propose an Inverse Graphics Capsule Network (IGC-Net) to learn the hierarchical 3D face representations from large-scale unlabeled images. The core of IGC-Net is a new type of capsule, named graphics capsule, which represents 3D primitives with interpretable parameters in computer graphics (CG), including depth, albedo, and 3D pose. Specifically, IGC-Net first decomposes the objects into a set of semantic-consistent part-level descriptions and then assembles them into object-level descriptions to build the hierarchy. The learned graphics capsules reveal how the neural networks, oriented at visual perception, understand faces as a hierarchy of 3D models. Besides, the discovered parts can be deployed to the unsupervised face segmentation task to evaluate the semantic consistency of our method. Moreover, the part-level descriptions with explicit physical meanings provide insight into the face analysis that originally runs in a black box, such as the importance of shape and texture for face recognition. Experiments on CelebA, BP4D, and Multi-PIE demonstrate the characteristics of our IGC-Net. Chang Yu 0001, Xiangyu Zhu 0001, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
CVPR | 4 |
| 2023 | FPR: False Positive Rectification for Weakly Supervised Semantic SegmentationabstractMany weakly supervised semantic segmentation (WSSS) methods employ the class activation map (CAM) to generate the initial segmentation results. However, CAM often fails to distinguish the foreground from its co-occurred background (e.g., train and railroad), resulting in inaccurate activation from the background. Previous endeavors address this co-occurrence issue by introducing external supervision and human priors. In this paper, we present a False Positive Rectification (FPR) approach to tackle the co-occurrence problem by leveraging the false positives of CAM. Based on the observation that the CAM-activated regions of absent classes contain class-specific co-occurred background cues, we collect these false positives and utilize them to guide the training of CAM network by proposing a region-level contrast loss and a pixel-level rectification loss. Without introducing any external supervision and human priors, the proposed FPR effectively suppresses wrong activations from the background objects. Extensive experiments on the PASCAL VOC 2012 and MS COCO 2014 demonstrate that FPR brings significant improvements for off-the-shelf methods and achieves state-of-the-art performance. Code is available at https://github.com/mt-cly/FPR. Liyi Chen 0002, Chenyang Lei, Ruihuang Li, Shuai Li 0014, Zhaoxiang Zhang 0001, Lei Zhang 0006 |
ICCV | 5 |
| 2023 | Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object DetectionabstractThis paper aims for high-performance offline LiDAR-based 3D object detection. We first observe that experienced human annotators annotate objects from a track-centric perspective. They first label objects in a track with clear shapes, and then leverage the temporal coherence to infer the annotations of obscure objects. Drawing inspiration from this, we propose a high-performance offline detector in a track-centric perspective instead of the conventional object-centric perspective. Our method features a bidirectional tracking module and a track-centric learning module. Such design allows our detector to infer and refine a complete track once the object is detected at a certain moment. We refer this characteristic to "onCe detecTed, neveR Lost" and name the proposed system CTRL. Extensive experiments demonstrate the remarkable performance of our method, surpassing the human-level annotating accuracy and previous state-of-the-art methods in the highly competitive Waymo Open Dataset leaderboard without model ensemble. The code is available at https://github.com/tusen-ai/SST. Lue Fan, Yuxue Yang, Yiming Mao 0008, Feng Wang 0015, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 7 |
| 2023 | DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other datasets to extract features, which are not suitable enough for WTAL. To address this problem, researchers design several modules for feature enhancement, which improve the performance of the localization module, especially modeling the temporal relationship between snippets. However, all of them omit that ambiguous snippets deliver contradictory information, which would reduce the discriminability of linked snippets. Considering this phenomenon, we propose Discriminability-Driven Graph Network (DDG-Net), which explicitly models ambiguous snippets and discriminative snippets with well-designed connections, preventing the transmission of ambiguous information and enhancing the discriminability of snippet-level representations. Additionally, we propose feature consistency loss to prevent the assimilation of features and drive the graph convolution network to generate more discriminative representations. Extensive experiments on THUMOS14 and ActivityNet1.2 benchmarks demonstrate the effectiveness of DDG-Net, establishing new state-of-the-art results on both datasets. Source code is available at https://github.com/XiaojunTang22/ICCV2023-DDGNet. Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang 0001, Man Zhang 0005, Zongyuan Yang |
ICCV | 4 |
| 2023 | Informative Data Mining for One-shot Cross-Domain Semantic SegmentationabstractContemporary domain adaptation offers a practical solution for achieving cross-domain transfer of semantic segmentation between labelled source data and unlabeled target data. These solutions have gained significant popularity; however, they require the model to be retrained when the test environment changes. This can result in unbearable costs in certain applications due to the time-consuming training process and concerns regarding data privacy. One-shot domain adaptation methods attempt to overcome these challenges by transferring the pre-trained source model to the target domain using only one target data. Despite this, the referring style transfer module still faces issues with computation cost and over-fitting problems. To address this problem, we propose a novel framework called Informative Data Mining (IDM) that enables efficient one-shot domain adaptation for semantic segmentation. Specifically, IDM provides an uncertainty-based selection criterion to identify the most informative samples, which facilitates quick adaptation and reduces redundant training. We then perform a model adaptation method using these selected samples, which includes patch-wise mixing and prototype-based information maximization to update the model. This approach effectively enhances adaptation and mitigates the overfitting problem. In general, we provide empirical evidence of the effectiveness and efficiency of IDM. Our approach outperforms existing methods and achieves a new state-of-the-art one-shot performance of 56.7%/55.4% on the GTA5/SYNTHIA to Cityscapes adaptation tasks, respectively. The code will be released at https://github.com/yxiwang/IDM. Yuxi Wang 0001, Jian Liang 0001, Jun Xiao 0005, Shuqi Mei, Yuran Yang, Zhaoxiang Zhang 0001 |
ICCV | 6 |
| 2023 | SSF: Accelerating Training of Spiking Neural Networks with Stabilized Spiking FlowabstractSurrogate gradient (SG) is one of the most effective approaches for training spiking neural networks (SNNs). While assisting SNNs to achieve classification performance comparable to artificial neural networks, SG suffers from the problem of time-consuming training, preventing it from efficient learning. In this paper, we formally analyze the backward process of classic SG and find that the membrane accumulation through time leads to exponential growth of training time. With this discovery, we propose Stabilized Spiking Flow (SSF), a simple yet effective approach to accelerate training of SG-based SNNs. For each spiking neuron, SSF averages its input and output activations over time to yield stabilized input and output, respectively. Then, instead of back propagating all errors that are related to current neuron and inherently entangled in time domain, the auxiliary gradient is directly propagated from the stabilized output to input through a devised relationship mapping. Additionally, SSF method is suitable to different neuron models. Extensive experiments on both static and neuromorphic datasets demonstrate that SNNs trained with SSF approach can achieve performance comparable to the original counterparts, while reducing the training time significantly. In particular, SSF speeds up the training process of state-of-the-art SNN models up to 10× when time steps equal to 80. Zengjie Song, Yuxi Wang 0001, Jun Xiao 0005, Yuran Yang, Shuqi Mei, Zhaoxiang Zhang 0001 |
ICCV | 7 |
| 2023 | LMR: A Large-Scale Multi-Reference Dataset for Reference-based Super-ResolutionabstractIt is widely agreed that reference-based super-resolution (RefSR) achieves superior results by referring to similar high quality images, compared to single image super-resolution (SISR). Intuitively, the more references, the better performance. However, previous RefSR methods have all focused on single-reference image training, while multiple reference images are often available in testing or practical applications. The root cause of such training-testing mismatch is the absence of publicly available multi-reference SR training datasets, which greatly hinders research efforts on multi-reference super-resolution. To this end, we construct a large-scale, multi-reference super-resolution dataset, named LMR. It contains 112, 142 groups of 300×300 training images, which is 10× of the existing largest RefSR dataset. The image size is also some times larger. More importantly, each group is equipped with 5 reference images with different similarity levels. Furthermore, we propose a new baseline method for multi-reference super-resolution: MRefSR, including a Multi-Reference Attention Module (MAM) for feature fusion of an arbitrary number of reference images, and a Spatial Aware Filtering Module (SAFM) for the fused feature selection. The proposed MRefSR achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations. Our code and data are available at: https://github.com/wdmwhh/MRefSR. Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Errui Ding, Zhaoxiang Zhang 0001 |
ICCV | 6 |
| 2023 | SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsabstractComputer end users have spent billions of hours completing daily tasks like tabular data processing and project timeline scheduling. Most of these tasks are repetitive and error-prone, yet most end users lack the skill to automate these burdensome works. With the advent of large language models (LLMs), directing software with natural language user requests become a reachable goal. In this work, we propose a SheetCopilot agent that takes natural language task and control spreadsheet to fulfill the requirements. We propose a set of atomic actions as an abstraction of spreadsheet software functionalities. We further design a state machine-based task planning framework for LLMs to robustly interact with spreadsheets. We curate a representative dataset containing 221 spreadsheet control tasks and establish a fully automated evaluation pipeline for rigorously benchmarking the ability of LLMs in software control tasks. Our SheetCopilot correctly completes 44.3\% of tasks for a single generation, outperforming the strong code generation baseline by a wide margin. Our project page: https://sheetcopilot.github.io/. Jingran Su, Yuntao Chen, Qing Li 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 5 |
| 2023 | Echoes Beyond Points: Unleashing the Power of Raw Radar Data in Multi-modality FusionabstractRadar is ubiquitous in autonomous driving systems due to its low cost and good adaptability to bad weather. Nevertheless, the radar detection performance is usually inferior because its point cloud is sparse and not accurate due to the poor azimuth and elevation resolution. Moreover, point cloud generation algorithms already drop weak signals to reduce the false targets which may be suboptimal for the use of deep fusion. In this paper, we propose a novel method named EchoFusion to skip the existing radar signal processing pipeline and then incorporate the radar raw data with other sensors. Specifically, we first generate the Bird's Eye View (BEV) queries and then take corresponding spectrum features from radar to fuse with other sensors. By this approach, our method could utilize both rich and lossless distance and speed clues from radar echoes and rich semantic clues from images, making our method surpass all existing methods on the RADIal dataset, and approach the performance of LiDAR. The code will be released on https://github.com/tusen-ai/EchoFusion. Yang Liu 0347, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
NeurIPS | 4 |
| 2023 | DropPos: Pre-Training Vision Transformers by Reconstructing Dropped PositionsabstractAs it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enhances the location awareness of ViTs is becoming evident. To address this, we present DropPos, a novel pretext task designed to reconstruct Dropped Positions. The formulation of DropPos is simple: we first drop a large random subset of positional embeddings and then the model classifies the actual position for each non-overlapping patch among all possible positions solely based on their visual appearance. To avoid trivial solutions, we increase the difficulty of this task by keeping only a subset of patches visible. Additionally, considering there may be different patches with similar visual appearances, we propose position smoothing and attentive reconstruction strategies to relax this classification problem, since it is not necessary to reconstruct their exact positions in these cases. Empirical evaluations of DropPos show strong capabilities. DropPos outperforms supervised pre-training and achieves competitive results compared with state-of-the-art self-supervised alternatives on a wide range of downstream benchmarks. This suggests that explicitly encouraging spatial reasoning abilities, as DropPos does, indeed contributes to the improved location awareness of ViTs. The code is publicly available at https://github.com/Haochen-Wang409/DropPos. Junsong Fan, Yuxi Wang 0001, Kaiyou Song, Zhaoxiang Zhang 0001 |
NeurIPS | 6 |
| 2023 | Fairly Adaptive Negative Sampling for RecommendationsabstractPairwise learning strategies are prevalent for optimizing recommendation models on implicit feedback data, which usually learns user preference by discriminating between positive (i.e., clicked by a user) and negative items (i.e., obtained by negative sampling). However, the size of different item groups (specified by item attribute) is usually unevenly distributed. We empirically find that the commonly used uniform negative sampling strategy for pairwise algorithms (e.g., BPR) can inherit such data bias and oversample the majority item group as negative instances, severely countering group fairness on the item side. In this paper, we propose a Fairly adaptive Negative sampling approach (FairNeg), which improves item group fairness via adaptively adjusting the group-level negative sampling distribution in the training process. In particular, it first perceives the model’s unfairness status at each step and then adjusts the group-wise sampling distribution with an adaptive momentum update strategy for better facilitating fairness optimization. Moreover, a negative sampling distribution Mixup mechanism is proposed, which gracefully incorporates existing importance-aware sampling techniques intended for mining informative negative samples, thus allowing for achieving multiple optimization purposes. Extensive experiments on four public datasets show our proposed method’s superiority in group fairness enhancement and fairness-utility tradeoff. Xiao Chen 0016, Wenqi Fan, Jingfan Chen, Zitao Liu 0001, Zhaoxiang Zhang 0001, Qing Li 0001 |
WWW | 6 |
| 2023 | Toward Practical Weakly Supervised Semantic Segmentation via Point-Level Supervision
Junsong Fan, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Super Sparse 3D Object DetectionabstractAs the perception range of LiDAR expands, LiDAR-based 3D object detection contributes ever-increasingly to the long-range perception in autonomous driving. Mainstream 3D object detectors often build dense feature maps, where the cost is quadratic to the perception range, making them hardly scale up to the long-range settings. To enable efficient long-range detection, we first propose a fully sparse object detector termed FSD. FSD is built upon the general sparse voxel encoder and a novel sparse instance recognition (SIR) module. SIR groups the points into instances and applies highly-efficient instance-wise feature extraction. The instance-wise grouping sidesteps the issue of the center feature missing, which hinders the design of the fully sparse architecture. To further enjoy the benefit of fully sparse characteristic, we leverage temporal information to remove data redundancy and propose a super sparse detector named FSD++. FSD++ first generates residual points, which indicate the point changes between consecutive frames. The residual points, along with a few previous foreground points, form the super sparse input data, greatly reducing data redundancy and computational overhead. We comprehensively analyze our method on the large-scale Waymo Open Dataset, and state-of-the-art performance is reported. To showcase the superiority of our method in long-range detection, we also conduct experiments on Argoverse 2 Dataset, where the perception range ([Formula: see text] m) is much larger than Waymo Open Dataset ([Formula: see text] m). Lue Fan, Yuxue Yang, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Memory-Based Cross-Image Contexts for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) trains segmentation models by only weak labels, aiming to save the burden of expensive pixel-level annotations. This paper tackles the WSSS problem of utilizing image-level labels as the weak supervision. Previous approaches address this problem by focusing on generating better pseudo-masks from weak labels to train the segmentation model. However, they generally only consider every single image and overlook the potential cross-image contexts. We emphasize that the cross-image contexts among a group of images can provide complementary information for each other to obtain better pseudo-masks. To effectively employ cross-image contexts, we develop an end-to-end cross-image context module containing a memory bank mechanism and a transformer-based cross-image attention module. The former extracts cross-image contexts online from the feature encodings of input images and stores them as the memory. The latter mines useful information from the memorized contexts to provide the original queries with additional information for better pseudo-mask generation. We conduct detailed experiments on the Pascal VOC 2012 and the COCO dataset to demonstrate the advantage of utilizing cross-image contexts. Besides, state-of-the-art performance is also achieved. Codes are available at https://github.com/js-fan/MCIC.git. Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learning to Adapt Across Dual Discrepancy for Cross-Domain Person Re-IdentificationabstractThanks to the advent of deep neural networks, recent years have witnessed rapid progress in person re-identification (re-ID). Deep-learning-based methods dominate the leadership of large-scale benchmarks, some of which even surpass the human-level performance. Despite their impressive performance under the single-domain setup, current fully-supervised re-ID models degrade significantly when transplanted to an unseen domain. According to the characteristics of the re-ID task, such degradation is mainly attributed to the dramatic variation within the target domain and the severe shift between the source and target domain, which we call dual discrepancy in this paper. To achieve a model that generalizes well to the target domain, it is desirable to take such dual discrepancy into account. In terms of the former issue, a prevailing solution is to enforce consistency between nearest-neighbors in the embedding space. However, we find that the search of neighbors is highly biased in our case due to the discrepancy across cameras. For this reason, we equip the vanilla neighborhood invariance approach with a camera-aware learning scheme. As for the latter issue, we propose a novel cross-domain mixup scheme. It works in conjunction with virtual prototypes which are employed to handle the disjoint label space between the two domains. In this way, we can realize the smooth transfer by introducing the interpolation between the two domains as a transition state. Extensive experiments on four public benchmarks demonstrate the superiority of our method. Without any auxiliary models and offline clustering procedure, it achieves competitive performance against existing state-of-the-art methods. The code is available at https://github.com/LuckyDC/generalizing-reid-improved. Chuanchen Luo, Chunfeng Song, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | GAIA-Universe: Everything is Super-NetifyabstractPre-training on large-scale datasets has played an increasingly significant role in computer vision and natural language processing recently. However, as there exist numerous application scenarios that have distinctive demands such as certain latency constraints and specialized data distributions, it is prohibitively expensive to take advantage of large-scale pre-training for per-task requirements. we focus on two fundamental perception tasks (object detection and semantic segmentation) and present a complete and flexible system named GAIA-Universe(GAIA), which could automatically and efficiently give birth to customized solutions according to heterogeneous downstream needs through data union and super-net training. GAIA is capable of providing powerful pre-trained weights and searching models that conform to downstream demands such as hardware constraints, computation constraints, specified data domains, and telling relevant data for practitioners who have very few datapoints on their tasks. With GAIA, we achieve promising results on COCO, Objects365, Open Images, BDD100 k, and UODB which is a collection of datasets including KITTI, VOC, WiderFace, DOTA, Clipart, Comic, and more. Taking COCO as an example, GAIA is able to efficiently produce models covering a wide range of latency from 16 ms to 53 ms, and yields AP from 38.2 to 46.5 without whistles and bells. GAIA is released at https://github.com/GAIA-vision. Junran Peng, Xingyuan Bu, Lingxi Xie, Xiaopeng Zhang 0008, Qi Tian 0001, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | Weakly Supervised Semantic Segmentation via Box-Driven Masking and Filling Rate ShiftingabstractSemantic segmentation has achieved huge progress via adopting deep Fully Convolutional Networks (FCN). However, the performance of FCN-based models severely rely on the amounts of pixel-level annotations which are expensive and time-consuming. Considering that bounding boxes also contain abundant semantic and objective information, an intuitive solution is to learn the segmentation with weak supervisions from the bounding boxes. How to make full use of the class-level and region-level supervisions from bounding boxes to estimate the uncertain regions is the critical challenge for the weakly supervised learning task. In this paper, we propose a mixture model to address this problem. First, we introduce a box-driven class-wise masking model (BCM) to remove irrelevant regions of each class. Moreover, based on the pixel-level segment proposal generated from the bounding box supervision, we calculate the mean filling rates of each class to serve as an important prior cue to guide the model ignoring the wrongly labeled pixels in proposals. To realize the more fine-grained supervision at instance-level, we further propose the anchor-based filling rate shifting module. Unlike previous methods that directly train models with the generated noisy proposals, our method can adjust the model learning dynamically with the adaptive segmentation loss. Thus it can help reduce the negative impacts from wrongly labeled proposals. Besides, based on the learned high-quality proposals with above pipeline, we explore to further boost the performance through two-stage learning. The proposed method is evaluated on the challenging PASCAL VOC 2012 benchmark and achieves 74.9 % and 76.4 % mean IoU accuracy under weakly and semi-supervised modes, respectively. Extensive experimental results show that the proposed method is effective and is on par with, or even better than current state-of-the-art methods. Chunfeng Song, Wanli Ouyang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Object Affinity Learning: Towards Annotation-Free Instance SegmentationabstractWe address the problem of annotation-free instance segmentation in the wild, aiming to relieve the expensive cost of manual mask annotations. Existing approaches utilize appearance cues, such as color, edge, and texture information, to generate pseudo masks for instance segmentation. However, due to the ambiguity of defining an object by visual appearance alone, these methods fail to distinguish objects from the background under complex scenes. Beyond visual cues, objects are one-piece in space and move together over time, which indicates that geometry cues, such as spatial continuity and motion consistency, are also exploitable for this problem. To directly utilize geometry cues, we propose an affinity-based paradigm for annotation-free instance segmentation. The new paradigm is called object affinity learning, a proxy task of annotation-free instance segmentation, which aims to tell whether two pixels come from the same object by learning feature representation from geometry cues. During inference, the learned object affinity could be further converted into instance segmentation masks by some graph partition algorithms. The proposed object affinity learning achieves much better instance segmentation performance than existing pseudo-mask-based methods on the large-scale Waymo Open Dataset and KITTI dataset. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | MMT: Cross Domain Few-Shot Learning via Meta-Memory TransferabstractFew-shot learning aims to recognize novel categories solely relying on a few labeled samples, with existing few-shot methods primarily focusing on the categories sampled from the same distribution. Nevertheless, this assumption cannot always be ensured, and the actual domain shift problem significantly reduces the performance of few-shot learning. To remedy this problem, we investigate an interesting and challenging cross-domain few-shot learning task, where the training and testing tasks employ different domains. Specifically, we propose a Meta-Memory scheme to bridge the domain gap between source and target domains, leveraging style-memory and content-memory components. The former stores intra-domain style information from source domain instances and provides a richer feature distribution. The latter stores semantic information through exploration of knowledge of different categories. Under the contrastive learning strategy, our model effectively alleviates the cross-domain problem in few-shot learning. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on cross-domain few-shot semantic segmentation tasks on the COCO-20$^{i}$, PASCAL-5$^{i}$, FSS-1000, and SUIM datasets and positively affects few-shot classification tasks on Meta-Dataset. Wenjian Wang 0002, Lijuan Duan, Yuxi Wang 0001, Junsong Fan, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Extracting Semantic Knowledge From GANs With Unsupervised LearningabstractRecently, unsupervised learning has made impressive progress on various tasks. Despite the dominance of discriminative models, increasing attention is drawn to representations learned by generative models and in particular, Generative Adversarial Networks (GANs). Previous works on the interpretation of GANs reveal that GANs encode semantics in feature maps in a linearly separable form. In this work, we further find that GAN's features can be well clustered with the linear separability assumption. We propose a novel clustering algorithm, named KLiSH, which leverages the linear separability to cluster GAN's features. KLiSH succeeds in extracting fine-grained semantics of GANs trained on datasets of various objects, e.g., car, portrait, animals, and so on. With KLiSH, we can sample images from GANs along with their segmentation masks and synthesize paired image-segmentation datasets. Using the synthesized datasets, we enable two downstream applications. First, we train semantic segmentation networks on these datasets and test them on real images, realizing unsupervised semantic segmentation. Second, we train image-to-image translation networks on the synthesized datasets, enabling semantic-conditional image synthesis without human annotations. Jianjin Xu, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Coarse Mask Guided Interactive Object SegmentationabstractInteractive object segmentation aims to produce object masks with user interactions, such as clicks, bounding boxes, and scribbles. Click point is the most popular interactive cue for its efficiency, and related deep learning methods have attracted lots of interest in recent years. Most works encode click points as gaussian maps and concatenate them with images as the model's input. However, the spatial and semantic information of gaussian maps would be noised through multiple convolution layers and won't be fully exploited by top layers for mask prediction. To pass click information to top layers exactly and efficiently, we propose a coarse mask guided model (CMG) which predicts coarse masks with a coarse module to guide the object mask prediction. Specifically, the coarse module encodes user clicks as query features and enriches their semantic information with backbone features through transformer layers, coarse masks are generated based on the enriched query feature and fed into CMG's decoder. Benefiting from the efficiency of transformer, CMG's coarse module and decoder module are lightweight and computationally efficient, making the interaction process more smooth. Experiments on several segmentation benchmarks demonstrate the effectiveness of our method, and we get new state-of-the-art results compared with previous works. Jing Li 0112, Junsong Fan, Yuxi Wang 0001, Yuran Yang, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Adversarial Learning Guided Task Relatedness Refinement for Multi-Task Deep LearningabstractIn machine learning, the relatedness across multiple tasks is usually complex and entangled. Due to dataset bias, the relatedness among tasks might be distorted and mislead the training of the models with solid learning ability, such as the multi-task neural networks. In this paper, we propose the idea of Relatedness Refinement Multi-Task Learning (RRMTDL) by introducing adversarial learning in the multi-task deep neural network to tackle the problem. The RRMTDL deep learning model restrains the misleading relatedness task by adversarial training and extracts information sharing across tasks with valuable relatedness. With RRMTDL, multi-task deep learning can enhance the task-specific representation for the major tasks by excluding the misleading relatedness. We design tests with various combinations of task-relatedness to validate the proposed model. Experimental results show that the RRMTDL model can effectively refine the task relatedness and prominently outperform other multi-task deep learning models in datasets with entangled task labels. Yuchun Fang, Sirui Cai, Yiting Cao, Zhengchen Li, Zhaoxiang Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual PredictionabstractDiscovering the underneath causal relations is the fundamental ability for reasoning about the surrounding environment and predicting the future states in the physical world. Counterfactual prediction from visual input, which requires simulating future states based on unrealized situations in the past, is a vital component in causal relation tasks. In this paper, we work on the confounders that have effect on the physical dynamics, including masses, friction coefficients, etc., to bridge relations between the intervened variable and the affected variable whose future state may be altered. We propose a neural network framework combining Global Causal Relation Attention (GCRA) and Confounder Transmission Structure (CTS). The GCRA looks for the latent causal relations between different variables and estimates the confounders by capturing both spatial and temporal information. The CTS integrates and transmits the learnt confounders in a residual way, so that the estimated confounders can be encoded into the network as a constraint for object positions when performing counterfactual prediction. Without any access to ground truth information about confounders, our model outperforms the state-of-the-art method on various benchmarks by fully utilizing the constraints of confounders. Extensive experiments demonstrate that our model can generalize to unseen environments and maintain good performance. Zongzhao Li, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
AAAI | 4 |
| 2022 | DATA: Domain-Aware and Task-Aware Self-supervised LearningabstractThe paradigm of training models on massive data without label through self-supervised learning (SSL) and fine-tuning on many downstream tasks has become a trend recently. However, due to the high training costs and the un-consciousness of downstream usages, most self-supervised learning methods lack the capability to correspond to the diversities of downstream scenarios, as there are various data domains, different vision tasks and latency constraints on models. Neural architecture search (NAS) is one universally acknowledged fashion to conquer the issues above, but applying NAS on SSL seems impossible as there is no label or metric provided for judging model selection. In this paper, we present DATA, a simple yet effective NAS approach specialized for SSL that provides Domain-Aware and Task-Aware pre-training. Specifically, we (i) train a supernet which could be deemed as a set of millions of networks covering a wide range of model scales without any label, (ii) propose a flexible searching mechanism compatible with SSL that enables finding networks of different computation costs, for various downstream vision tasks and data domains without explicit metric provided. Instantiated With MoCo v2, our method achieves promising results across a wide range of computation costs on down-stream tasks, including image classification, object detection and semantic segmentation. DATA is orthogonal to most existing SSL methods and endows them the ability of customization on downstream needs. Extensive experiments on other SSL methods demonstrate the generalizability of the proposed method. Code is released at https://github.com/GAIA-vision/GAIA-ssl. Junran Peng, Lingxi Xie, Qi Tian 0001, Zhaoxiang Zhang 0001 |
CVPR | 7 |
| 2022 | Sparse Instance Activation for Real-Time Instance SegmentationabstractIn this paper, we propose a conceptually novel, efficient, and fully convolutional framework for real-time instance segmentation. Previously, most instance segmentation methods heavily rely on object detection and perform mask prediction based on bounding boxes or dense centers. In contrast, we propose a sparse set of instance activation maps, as a new object representation, to high-light informative regions for each foreground object. Then instance-level features are obtained by aggregating features according to the highlighted regions for recognition and segmentation. Moreover, based on bipartite matching, the instance activation maps can predict objects in a one-to-one style, thus avoiding non-maximum suppression (NMS) in post-processing. Owing to the simple yet effective designs with instance activation maps, SparseInst has extremely fast inference speed and achieves 40 FPS and 37.9 AP on the COCO benchmark, which significantly out-performs the counterparts in terms of speed and accuracy. Code and models are available at https://github.com/hustvl/SparseInst. Tianheng Cheng, Xinggang Wang, Shaoyu Chen, Qian Zhang 0009, Chang Huang, Zhaoxiang Zhang 0001, Wenyu Liu 0001 |
CVPR | 7 |
| 2022 | Embracing Single Stride 3D Object Detector with Sparse TransformerabstractIn LiDAR-based 3D object detection for autonomous driving, the ratio of the object size to input scene size is significantly smaller compared to 2D detection cases. Over-looking this difference, many 3D detectors directly follow the common practice of 2D detectors, which downsample the feature maps even after quantizing the point clouds. In this paper, we start by rethinking how such multi-stride stereotype affects the LiDAR-based 3D object detectors. Our experiments point out that the downsampling operations bring few advantages, and lead to inevitable information loss. To remedy this issue, we propose Single-stride Sparse Transformer (SST) to maintain the original resolution from the beginning to the end of the network. Armed with transformers, our method addresses the problem of insufficient receptive field in single-stride architectures. It also cooperates well with the sparsity of point clouds and naturally avoids expensive computation. Eventually, our SST achieves state-of-the-art results on the large-scale Waymo Open Dataset. It is worth mentioning that our method can achieve exciting performance (83.8 LEVEL_1 AP on validation split) on small object (pedestrian) detection due to the characteristic of single stride. Our codes will be public soon. Lue Fan, Ziqi Pang, Tianyuan Zhang 0002, Yu-Xiong Wang, Hang Zhao 0021, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 8 |
| 2022 | Towards Noiseless Object Contours for Weakly Supervised Semantic SegmentationabstractImage-level label based weakly supervised semantic segmentation has attracted much attention since image labels are very easy to obtain. Existing methods usually generate pseudo labels from class activation map (CAM) and then train a segmentation model. CAM usually highlights partial objects and produce incomplete pseudo labels. Some methods explore object contour by training a contour model with CAM seed label supervision and then propagate CAM score from discriminative regions to nondiscriminative regions with contour guidance. The propagation process suffers from the noisy intra-object contours, and inadequate propagation results produce incomplete pseudo labels. This is because the coarse CAM seed label lacks sufficient precise semantic information to suppress contour noise. In this paper, we train a SANCE model which utilizes an auxiliary segmentation module to supplement high-level semantic information for contour training by backbone feature sharing and online label supervision. The auxiliary segmentation module also provides more accurate localization map than CAM for pseudo label generation. We evaluate our approach on Pascal VOC 2012 and MS COCO 2014 benchmarks and achieve stateof- the-art performance, demonstrating the effectiveness of our method. The source code can be found at https://github.com/BraveGroup/SANCE Jing Li 0112, Junsong Fan, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2022 | Remember the Difference: Cross-Domain Few-Shot Semantic Segmentation via Meta-Memory TransferabstractFew-shot semantic segmentation intends to predict pixel-level categories using only a few labeled samples. Existing few-shot methods focus primarily on the categories sampled from the same distribution. Nevertheless, this assumption cannot always be ensured. The actual domain shift problem significantly reduces the performance of few-shot learning. To remedy this problem, we propose an interesting and challenging cross-domain few-shot semantic segmentation task, where the training and test tasks perform on different domains. Specifically, we first propose a meta-memory bank to improve the generalization of the segmentation network by bridging the domain gap between source and target domains. The meta-memory stores the intra-domain style information from source domain instances and transfers it to target samples. Subsequently, we adopt a new contrastive learning strategy to explore the knowledge of different categories during the training stage. The negative and positive pairs are obtained from the proposed memory-based style augmentation. Comprehensive experiments demon-strate that our proposed method achieves promising results on cross-domain few-shot semantic segmentation tasks on COCO-20i, PASCAL-Si, FSS-1000, and SUIM datasets. Wenjian Wang 0002, Lijuan Duan, Yuxi Wang 0001, Qing En, Junsong Fan, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2022 | HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule NetworkabstractCapsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures, which are suitable for capsules to describe, have not been explored. In this paper, we propose a Hierarchical Parsing Capsule Network (HP-Capsule) for unsupervised face subpart-part discovery. When browsing large-scale face images without labels, the network first encodes the frequently observed patterns with a set of explainable subpart capsules. Then, the subpart capsules are assembled into part-level capsules through a Transformer-based Parsing Module (TPM) to learn the compositional relations between them. During training as the face hierarchy is progressively built and refined, the part capsules adaptively encode the face parts with semantic consistency. HP-Capsule extends the application of capsule networks from digits to human faces and takes a step forward to show how the neural networks understand homologous objects without human intervention. Besides, HP-Capsule gives unsupervised face segmentation results by the covered regions of part capsules, enabling qualitative and quantitative evaluation. Experiments on BP4D and Multi-PIE datasets show the effectiveness of our method. Chang Yu 0001, Xiangyu Zhu 0001, Zidu Wang, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
CVPR | 5 |
| 2022 | Implicit Sample Extension for Unsupervised Person Re-IdentificationabstractMost existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters substantially hampers the Re-ID accuracy. Due to the limited samples in each identity, we suppose there may lack some underlying information to well reveal the accurate clusters. To discover these information, we propose an Implicit Sample Extension (ISE) method to generate what we call support samples around the cluster boundaries. Specifically, we generate support samples from actual samples and their neighbouring clusters in the embedding space through a progressive linear interpolation (PLI) strategy. PLI controls the generation with two critical factors, i.e., 1) the direction from the actual sample towards its K-nearest clusters and 2) the degree for mixing up the context information from the K-nearest clusters. Meanwhile, given the support samples, ISE further uses a label-preserving loss to pull them towards their corresponding actual samples, so as to compact each cluster. Consequently, ISE reduces the “sub and mixed” clustering errors, thus improving the Re-ID performance. Extensive experiments demonstrate that the proposed method is effective and achieves state-of-the-art performance for unsupervised person Re-ID. Code is available at: https://github.com/PaddlePaddle/PaddleClas. Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Errui Ding, Qinfeng Shi, Zhaoxiang Zhang 0001, Jingdong Wang 0001 |
CVPR | 7 |
| 2022 | Continual Stereo Matching of Continuous Driving Scenes with Growing ArchitectureabstractThe deep stereo models have achieved state-of-the-art performance on driving scenes, but they suffer from severe performance degradation when tested on unseen scenes. Although recent work has narrowed this performance gap through continuous online adaptation, this setup requires continuous gradient updates at inference and can hardly deal with rapidly changing scenes. To address these challenges, we propose to perform continual stereo matching where a model is tasked to 1) continually learn new scenes, 2) overcome forgetting previously learned scenes, and 3) continuously predict disparities at deployment. We achieve this goal by introducing a Reusable Architecture Growth (RAG) framework. RAG leverages task-specific neural unit search and architecture growth for continual learning of new scenes. During growth, it can maintain high reusability by reusing previous neural units while achieving good performance. A module named Scene Router is further introduced to adaptively select the scene-specific architecture path at inference. Experimental results demonstrate that our method achieves compelling performance in various types of challenging driving scenes. Chenghao Zhang 0003, Bin Fan 0001, Gaofeng Meng, Zhaoxiang Zhang 0001, Chunhong Pan |
CVPR | 5 |
| 2022 | The Devil Is in the Details: Window-based Attention for Image CompressionabstractLearned image compression methods have exhibited superior rate-distortion performance than classical image compression standards. Most existing learned image compression models are based on Convolutional Neural Networks (CNNs). Despite great contributions, a main drawback of CNN based model is that its structure is not designed for capturing local redundancy, especially the nonrepetitive textures, which severely affects the reconstruction quality. Therefore, how to make full use of both global structure and local texture becomes the core problem for learning-based image compression. Inspired by recent progresses of Vision Transformer (ViT) and Swin Transformer, we found that combining the local-aware attention mechanism with the global-related feature learning could meet the expectation in image compression. In this paper, we first extensively study the effects of multiple kinds of attention mechanisms for local features learning, then introduce a more straightforward yet effective window-based local attention block. The proposed window-based attention is very flexible which could work as a plug-and-play component to enhance CNN and Transformer models. Moreover, we propose a novel Symmetrical TransFormer (STF) framework with absolute transformer blocks in the down-sampling encoder and up-sampling decoder. Extensive experimental evaluations have shown that the proposed method is effective and outperforms the state-of-the-art methods. The code is publicly available at https://github.com/Googolxx/STF. Renjie Zou, Chunfeng Song, Zhaoxiang Zhang 0001 |
CVPR | 3 |
| 2022 | Pointly-Supervised Panoptic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (30) | 2 |
| 2022 | Densely Constrained Depth Estimator for Monocular 3D Object Detection
Yingyan Li, Yuntao Chen, Jiawei He 0002, Zhaoxiang Zhang 0001 |
ECCV (9) | 4 |
| 2022 | RRSR: Reciprocal Reference-Based Image Super-Resolution with Progressive Feature Alignment and Selection
Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Yili Wang 0003, Zhaoxiang Zhang 0001 |
ECCV (19) | 6 |
| 2022 | Stereo Depth Estimation with Echoes
Chenghao Zhang 0003, Bolin Ni, Gaofeng Meng, Bin Fan 0001, Zhaoxiang Zhang 0001, Chunhong Pan |
ECCV (27) | 6 |
| 2022 | Object Dynamics Distillation for Scene Decomposition and Representation
Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
ICLR | 4 |
| 2022 | Self-Guided Hard Negative Generation for Unsupervised Person Re-IdentificationabstractRecent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative samples, playing important role in training reID models, are significantly reduced. To alleviate this problem, we propose a self-guided hard negative generation method for unsupervised person re-ID. Specifically, a joint framework is developed which incorporates a hard negative generation network (HNGN) and a re-ID network. To continuously generate harder negative samples to provide effective supervisions in the contrastive learning, the two networks are alternately trained in an adversarial manner to improve each other, where the reID network guides HNGN to generate challenging data and HNGN enforces the re-ID network to enhance discrimination ability. During inference, the performance of re-ID network is improved without introducing any extra parameters. Extensive experiments demonstrate that the proposed method significantly outperforms a strong baseline and also achieves better results than state-of-the-art methods. Zhigang Wang 0002, Jian Wang 0066, Xinyu Zhang 0015, Errui Ding, Jingdong Wang 0001, Zhaoxiang Zhang 0001 |
IJCAI | 7 |
| 2022 | Interact with Open Scenes: A Life-long Evolution Framework for Interactive Segmentation ModelsabstractExisting interactive segmentation methods mainly focus on optimizing user interacting strategies, as well as making better use of clicks provided by users. However, the intention of the interactive segmentation model is to obtain high-quality masks with limited user interactions, which are supposed to be applied to unlabeled new images. But most existing methods overlooked the generalization ability of their models when witnessing new target scenes. To overcome this problem, we propose a life-long evolution framework for interactive models in this paper, which provides a possible solution for dealing with dynamic target scenes with one single model. Given several target scenes and an initial model trained with labels on the limited closed dataset, our framework arranges sequentially evolution steps on each target set. Specifically, we propose an interactive-prototype module to generate and refine pseudo masks, and apply a feature alignment module in order to adapt the model to a new target scene and keep the performance on previous images at the same time. All evolution steps above do not require ground truth labels as supervision. We conduct thorough experiments on PASCAL VOC, Cityscapes, and COCO datasets, demonstrating the effectiveness of our framework in solving new target datasets and maintaining performance on previous scenes at the same time. Ruitong Gan, Junsong Fan, Yuxi Wang 0001, Zhaoxiang Zhang 0001 |
ACM Multimedia | 4 |
| 2022 | Fully Sparse 3D Object DetectionabstractAs the perception range of LiDAR increases, LiDAR-based 3D object detection becomes a dominant task in the long-range perception task of autonomous driving. The mainstream 3D object detectors usually build dense feature maps in the network backbone and prediction head. However, the computational and spatial costs on the dense feature map are quadratic to the perception range, which makes them hardly scale up to the long-range setting. To enable efficient long-range LiDAR-based object detection, we build a fully sparse 3D object detector (FSD). The computational and spatial cost of FSD is roughly linear to the number of points and independent of the perception range. FSD is built upon the general sparse voxel encoder and a novel sparse instance recognition (SIR) module. SIR first groups the points into instances and then applies instance-wise feature extraction and prediction. In this way, SIR resolves the issue of center feature missing, which hinders the design of the fully sparse architecture for all center-based or anchor-based detectors. Moreover, SIR avoids the time-consuming neighbor queries in previous point-based methods by grouping points into instances. We conduct extensive experiments on the large-scale Waymo Open Dataset to reveal the working mechanism of FSD, and state-of-the-art performance is reported. To demonstrate the superiority of FSD in long-range detection, we also conduct experiments on Argoverse 2 Dataset, which has a much larger perception range ($200m$) than Waymo Open Dataset ($75m$). On such a large perception range, FSD achieves state-of-the-art performance and is 2.4$\times$ faster than the dense counterpart. Codes will be released. Lue Fan, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
NeurIPS | 4 |
| 2022 | 4D Unsupervised Object DiscoveryabstractObject discovery is a core task in computer vision. While fast progresses have been made in supervised object detection, its unsupervised counterpart remains largely unexplored. With the growth of data volume, the expensive cost of annotations is the major limitation hindering further study. Therefore, discovering objects without annotations has great significance. However, this task seems impractical on still-image or point cloud alone due to the lack of discriminative information. Previous studies underlook the crucial temporal information and constraints naturally behind multi-modal inputs. In this paper, we propose 4D unsupervised object discovery, jointly discovering objects from 4D data -- 3D point clouds and 2D RGB images with temporal information. We present the first practical approach for this task by proposing a ClusterNet on 3D point clouds, which is jointly iteratively optimized with a 2D localization network. Extensive experiments on the large-scale Waymo Open Dataset suggest that the localization network and ClusterNet achieve competitive performance on both class-agnostic 2D object detection and 3D instance segmentation, bridging the gap between unsupervised methods and full supervised ones. Codes and models will be made available at https://github.com/Robertwyq/LSMOL. Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001 |
NeurIPS | 3 |
| 2022 | Toward few-shot domain adaptation with perturbation-invariant representation and transferable prototypes
Junsong Fan, Yuxi Wang 0001, He Guan, Chunfeng Song, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 5 |
| 2022 | Improving Image Segmentation with Boundary Patch Refinement
Xiaolin Hu 0001, Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 6 |
| 2022 | From Individual to Whole: Reducing Intra-class Variance by Feature Aggregation
Zhaoxiang Zhang 0001, Chuanchen Luo, Haiping Wu, Yuntao Chen, Naiyan Wang, Chunfeng Song |
Int. J. Comput. Vis. | 1 |
| 2022 | Delving into the Effectiveness of Receptive Fields: Learning Scale-Transferrable Architectures for Practical Object Detection
Zhaoxiang Zhang 0001, Cong Pan 0001, Junran Peng |
Int. J. Comput. Vis. | 1 |
| 2022 | Identifying the key frames: An attention-aware sampling method for action recognition
Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
Pattern Recognit. | 2 |
| 2022 | Enhanced task attention with adversarial learning for dynamic multi-task CNN
Yuchun Fang, Shiwei Xiao, Menglu Zhou, Sirui Cai, Zhaoxiang Zhang 0001 |
Pattern Recognit. | 5 |
| 2022 | MonoPoly: A practical monocular 3D object detector
He Guan, Chunfeng Song, Zhaoxiang Zhang 0001, Tieniu Tan |
Pattern Recognit. | 3 |
| 2022 | Context-aware co-supervision for accurate object detection
Junran Peng, Haoquan Wang, Shaolong Yue, Zhaoxiang Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | Multimodal channel-wise attention transformer inspired by multisensory integration mechanisms of the brain
Qianqian Shi 0001, Junsong Fan, Zuoren Wang, Zhaoxiang Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | Alleviating Modality Bias Training for Infrared-Visible Person Re-IdentificationabstractThe task of infrared-visible person re-identification (IV-reID) is to recognize people across two modalities (i.e., RGB and IR). Existing cutting-edge approaches normally use a pair of images that have the same IDs (i.e., ID-tied cross-modality image pairs) and input them into an ImageNet-trained ResNet50. The ResNet50 backbone model can learn shared features across modalities to tolerate modality discrepancies between RGB and IR. This work will unveil a Modality Bias Training (MBT) problem that is less discussed in IV-reID, which will demonstrate that MBT significantly compromises the performance of IV-reID. Due to MBT, IR information can be overwhelmed by RGB information during training when the ResNet50 model is pretrained based on a large amount of RGB images from ImageNet. Thus, the trained models are more inclined to RGB information. Accordingly, the cross-modality generalization ability of the model is also compromised. To tackle this issue, we present a Dual-level Learning Strategy (DLS) that 1) enforces the focus of the network on ID-exclusive (rather than ID-tied) labels of cross-modality image pairs to mitigate the problem of MBT and 2) introduces third modality data that contain both RGB and IR information to further prevent the information from the IR modality from being overwhelmed during training. Our third modality images are generated by a generative adversarial network. A dynamic ID-exclusive Smooth (dIDeS) label is proposed for the generated third modality data. In experiments, comprehensive experiments are carried out to demonstrate the success of DLS in tackling the MBT issue exposed in IV-reID. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | Group-Wise Semantic Mining for Weakly Supervised Semantic SegmentationabstractAcquiring sufficient ground-truth supervision to train deep vi- sual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmen- tation, which requires pixel-level annotations. This work ad- dresses weakly supervised semantic segmentation (WSSS), with the goal of bridging the gap between image-level anno- tations and pixel-level segmentation. We formulate WSSS as a novel group-wise learning task that explicitly models se- mantic dependencies in a group of images to estimate more reliable pseudo ground-truths, which can be used for training more accurate segmentation models. In particular, we devise a graph neural network (GNN) for group-wise semantic min- ing, wherein input images are represented as graph nodes, and the underlying relations between a pair of images are char- acterized by an efficient co-attention mechanism. Moreover, in order to prevent the model from paying excessive atten- tion to common semantics only, we further propose a graph dropout layer, encouraging the model to learn more accurate and complete object responses. The whole network is end-to- end trainable by iterative message passing, which propagates interaction cues over the images to progressively improve the performance. We conduct experiments on the popular PAS- CAL VOC 2012 and COCO benchmarks, and our model yields state-of-the-art performance. Our code is available at: https://github.com/Lixy1997/Group-WSSS. Xueyi Li 0006, Tianfei Zhou, Jianwu Li, Yi Zhou 0007, Zhaoxiang Zhang 0001 |
AAAI | 5 |
| 2021 | GAIA: A Transfer Learning System of Object Detection That Fits Your NeedsabstractTransfer learning with pre-training on large-scale datasets has played an increasingly significant role in computer vision and natural language processing recently. However, as there exist numerous application scenarios that have distinctive demands such as certain latency constraints and specialized data distributions, it is prohibitively expensive to take advantage of large-scale pre-training for per-task requirements. In this paper, we focus on the area of object detection and present a transfer learning system named GAIA, which could automatically and efficiently give birth to customized solutions according to heterogeneous downstream needs. GAIA is capable of providing powerful pre-trained weights, selecting models that conform to downstream demands such as latency constraints and specified data domains, and collecting relevant data for practitioners who have very few datapoints for their tasks. With GAIA, we achieve promising results on COCO, Objects365, Open Images, Caltech, CityPersons, and UODB which is a collection of datasets including KITTI, VOC, WiderFace, DOTA, Clipart, Comic, and more. Taking COCO as an ex-ample, GAIA is able to efficiently produce models covering a wide range of latency from 16ms to 53ms, and yields AP from 38.2 to 46.5 without whistles and bells. To benefit every practitioner in the community of object detection, GAIA is released at https://github.com/GAIA-vision. Xingyuan Bu, Junran Peng, Tieniu Tan, Zhaoxiang Zhang 0001 |
CVPR | 5 |
| 2021 | Bottom-Up Human Pose Estimation via Disentangled Keypoint RegressionabstractIn this paper, we are interested in the bottom-up paradigm of estimating human poses from an image. We study the dense keypoint regression framework that is previously inferior to the keypoint detection and grouping framework. Our motivation is that regressing keypoint positions accurately needs to learn representations that focus on the keypoint regions.We present a simple yet effective approach, named disentangled keypoint regression (DEKR). We adopt adaptive convolutions through pixel-wise spatial transformer to activate the pixels in the keypoint regions and accordingly learn representations from them. We use a multi-branch structure for separate regression: each branch learns a representation with dedicated adaptive convolutions and regresses one keypoint. The resulting disentangled representations are able to attend to the keypoint regions, respectively, and thus the keypoint regression is spatially more accurate. We empirically show that the proposed direct regression method outperforms keypoint detection and grouping methods and achieves superior bottom-up pose estimation results on two benchmark datasets, COCO and CrowdPose. The code and models are available at https://github.com/HRNet/DEKR. Zigang Geng, Ke Sun 0009, Bin Xiao 0004, Zhaoxiang Zhang 0001, Jingdong Wang 0001 |
CVPR | 4 |
| 2021 | Learnable Graph Matching: Incorporating Graph Partitioning With Deep Feature Learning for Multiple Object TrackingabstractData association across frames is at the core of Multiple Object Tracking (MOT) task. This problem is usually solved by a traditional graph-based optimization or directly learned via deep learning. Despite their popularity, we find some points worth studying in current paradigm: 1) Existing methods mostly ignore the context information among tracklets and intra-frame detections, which makes the tracker hard to survive in challenging cases like severe occlusion. 2) The end-to-end association methods solely rely on the data fitting power of deep neural networks, while they hardly utilize the advantage of optimization-based assignment methods. 3) The graph-based optimization methods mostly utilize a separate neural network to extract features, which brings the inconsistency between training and inference. Therefore, in this paper we propose a novel learnable graph matching method to address these issues. Briefly speaking, we model the relationships between tracklets and the intra-frame detections as a general undirected graph. Then the association problem turns into a general graph matching between tracklet graph and detection graph. Furthermore, to make the optimization end-to-end differentiable, we relax the original graph matching into continuous quadratic programming and then incorporate the training of it into a deep graph network with the help of the implicit function theorem. Lastly, our method GMTracker, achieves state-of-the-art performance on several standard MOT datasets. Our code is available at https://github.com/jiaweihe1996/GMTracker. Jiawei He 0002, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2021 | Look Closer To Segment Better: Boundary Patch Refinement for Instance SegmentationabstractTremendous efforts have been made on instance segmentation but the mask quality is still not satisfactory. The boundaries of predicted instance masks are usually imprecise due to the low spatial resolution of feature maps and the imbalance problem caused by the extremely low proportion of boundary pixels. To address these issues, we propose a conceptually simple yet effective post-processing refinement framework to improve the boundary quality based on the results of any instance segmentation model, termed BPR. Following the idea of looking closer to segment boundaries better, we extract and refine a series of small boundary patches along the predicted instance boundaries. The refinement is accomplished by a boundary patch refinement network at higher resolution. The proposed BPR framework yields significant improvements over the Mask R-CNN baseline on Cityscapes benchmark, especially on the boundary-aware metrics. Moreover, by applying the BPR framework to the "PolyTransform + SegFix" baseline, we reached 1stplace on the Cityscapes leaderboard. Code is available at https://github.com/tinyalpha/BPR. Chufeng Tang, Hang Chen 0004, Xiao Li 0028, Jianmin Li 0001, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
CVPR | 5 |
| 2021 | Unsupervised Object Detection With LIDAR CluesabstractDespite the importance of unsupervised object detection, to the best of our knowledge, there is no previous work addressing this problem. One main issue, widely known to the community, is that object boundaries derived only from 2D image appearance are ambiguous and unreliable. To address this, we exploit LiDAR clues to aid unsupervised object detection. By exploiting the 3D scene structure, the issue of localization can be considerably mitigated. We further identify another major issue, seldom noticed by the community, that the long-tailed and open-ended (sub-)category distribution should be accommodated. In this paper, we present the first practical method for unsupervised object detection with the aid of LiDAR clues. In our approach, candidate object segments based on 3D point clouds are firstly generated. Then, an iterative segment labeling process is conducted to assign segment labels and to train a segment labeling network, which is based on features from both 2D images and 3D point clouds. The labeling process is carefully designed so as to mitigate the issue of long-tailed and open-ended distribution. The final segment labels are set as pseudo annotations for object detection network training. Extensive experiments on the large-scale Waymo Open dataset suggest that the derived unsupervised object detection method achieves reasonable accuracy compared with that of strong supervision within the LiDAR visible range. Hao Tian 0006, Yuntao Chen, Jifeng Dai, Zhaoxiang Zhang 0001, Xizhou Zhu |
CVPR | 4 |
| 2021 | RefineMask: Towards High-Quality Instance Segmentation With Fine-Grained FeaturesabstractThe two-stage methods for instance segmentation, e.g. Mask R-CNN, have achieved excellent performance recently. However, the segmented masks are still very coarse due to the downsampling operations in both the feature pyramid and the instance-wise pooling process, especially for large objects. In this work, we propose a new method called RefineMask for high-quality instance segmentation of objects and scenes, which incorporates fine-grained features during the instance-wise segmenting process in a multi-stage manner. Through fusing more detailed information stage by stage, RefineMask is able to refine high-quality masks consistently. RefineMask succeeds in segmenting hard cases such as bent parts of objects that are oversmoothed by most previous methods and outputs accurate boundaries. Without bells and whistles, RefineMask yields significant gains of 2.6, 3.4, 3.8 AP over Mask R-CNN on COCO, LVIS, and Cityscapes benchmarks respectively at a small amount of additional computational cost. Furthermore, our single-model result outperforms the winner of the LVIS Challenge 2020 by 1.3 points on the LVIS test-dev set and establishes a new state-of-the-art. Code will be available at https://github.com/zhanggang001/RefineMask. Xin Lu 0002, Jingru Tan, Jianmin Li 0001, Zhaoxiang Zhang 0001, Quanquan Li, Xiaolin Hu 0001 |
CVPR | 5 |
| 2021 | Distractor-Aware Fast Tracking via Dynamic Convolutions and MOT PhilosophyabstractA practical long-term tracker typically contains three key properties, i.e. an efficient model design, an effective global re-detection strategy and a robust distractor awareness mechanism. However, most state-of-the-art long-term trackers (e.g., Pseudo and re-detecting based ones) do not take all three key properties into account and therefore may either be time-consuming or drift to distractors. To address the issues, we propose a two-task tracking framework (named DMTrack), which utilizes two core components (i.e., one-shot detection and re-identification (re-id) association) to achieve distractor-aware fast tracking via Dynamic convolutions (d-convs) and Multiple object tracking (MOT) philosophy. To achieve precise and fast global detection, we construct a lightweight one-shot detector using a novel dynamic convolutions generation method, which provides a unified and more flexible way for fusing target information into the search field. To distinguish the target from distractors, we resort to the philosophy of MOT to reason distractors explicitly by maintaining all potential similarities’ tracklets. Benefited from the strength of high recall detection and explicit object association, our tracker achieves state-of-the-art performance on the LaSOT, Ox-UvA, TLP, VOT2018LT and VOT2019LT benchmarks and runs in real-time (3x faster than comparisons)1. Zikai Zhang 0003, Bineng Zhong 0001, Shengping Zhang, Zhenjun Tang, Xin Liu 0011, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2021 | Clothing Status Awareness for Long-Term Person Re-IdentificationabstractLong-Term person re-identification (LT-reID) exposes extreme challenges because of the longer time gaps between two recording footages where a person is likely to change clothing. There are two types of approaches for LT-reID: biometrics-based approach and data adaptation based approach. The former one is to seek clothing irrelevant biometric features. However, seeking high quality biometric feature is the main concern. The latter one adopts fine-tuning strategy by using data with significant clothing change. However, the performance is compromised when it is applied to cases without clothing change. This work argues that these approaches in fact are not aware of clothing status (i.e., change or no-change) of a pedestrian. Instead, they blindly assume all footages of a pedestrian have different clothes. To tackle this issue, a Regularization via Clothing Status Awareness Network (RCSANet) is proposed to regularize descriptions of a pedestrian by embedding the clothing status awareness. Consequently, the description can be enhanced to maintain the best ID discriminative feature while improving its robustness to real-world LT-reID where both clothing-change case and no-clothing-change case exist. Experiments show that RCSANet performs reasonably well on three LT-reID datasets. Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2021 | RangeDet: In Defense of Range View for LiDAR-based 3D Object DetectionabstractIn this paper, we propose an anchor-free single-stage LiDAR-based 3D object detector – RangeDet. The most notable difference with previous works is that our method is purely based on the range view representation. Compared with the commonly used voxelized or Bird’s Eye View (BEV) representations, the range view representation is more compact and without quantization error. Although there are works adopting it for semantic segmentation, its performance in object detection is largely behind voxelized or BEV counterparts. We first analyze the existing range-view-based methods and find two issues overlooked by previous works: 1) the scale variation between nearby and far away objects; 2) the inconsistency between the 2D range image coordinates used in feature extraction and the 3D Cartesian coordinates used in output. Then we deliberately design three components to address these issues in our RangeDet. We test our RangeDet in the large-scale Waymo Open Dataset (WOD). Our best model achieves 72.9/75.9/65.8 3D AP on vehicle/pedestrian/cyclist. These results outperform other range-view-based methods by a large margin, and are overall comparable with the state-of-the-art multi-view-based methods. Codes will be released at https://github.com/TuSimple/RangeDet. Lue Fan, Xuan Xiong, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 5 |
| 2021 | Uncertainty-aware Pseudo Label Refinery for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation aims to assign the pixel-level labels for unlabeled target domain by transferring knowledge from the labeled source domain. A typical self-supervised learning approach generates pseudo labels from the source model and then re-trains the model to fit the target distribution. However, it suffers from noisy pseudo labels due to the existence of domain shift. Related works alleviate this problem by selecting high-confidence predictions, but uncertain classes with low confidence scores have rarely been considered. This informative uncertainty is essential to enhance feature representation and align source and target domains. In this paper, we propose a novel uncertainty-aware pseudo label refinery framework considering two crucial factors simultaneously. First, we progressively enhance the feature alignment model via the target-guided uncertainty rectifying framework. Second, we provide an uncertainty-aware pseudo label assignment strategy without any manually de-signed threshold to reduce the noisy labels. Extensive experiments demonstrate the effectiveness of our proposed approach and achieve state-of-the-art performance on two standard synthetic-2-real tasks. Yuxi Wang 0001, Junran Peng, Zhaoxiang Zhang 0001 |
ICCV | 3 |
| 2021 | Biologically inspired visual computing: the state of the art
Ian Max Andolina, Wei Wang 0061, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 4 |
| 2021 | Unsupervised Domain Adaptation with Background Shift Mitigating for Person Re-Identification
Yan Huang 0023, Qiang Wu 0001, Jingsong Xu, Yi Zhong 0002, Zhaoxiang Zhang 0001 |
Int. J. Comput. Vis. | 5 |
| 2021 | Joint Multisource Saliency and Exemplar Mechanism for Weakly Supervised Video Object SegmentationabstractWeakly supervised video object segmentation (WSVOS) is a vital yet challenging task in which the aim is to segment pixel-level masks with only category labels. Existing methods still have certain limitations, e.g., difficulty in comprehending appropriate spatiotemporal knowledge and an inability to explore common semantic information with category labels. To overcome these challenges, we formulate a novel framework by integrating multisource saliency and incorporating an exemplar mechanism for WSVOS. Specifically, we propose a multisource saliency module to comprehend spatiotemporal knowledge by integrating spatial and temporal saliency as bottom-up cues, which can effectively eliminate disruptions due to confusing regions and identify attractive regions. Moreover, to our knowledge, we make the first attempt to incorporate an exemplar mechanism into WSVOS by proposing an adaptive exemplar module to process top-down cues, which can provide reliable guidance for co-occurring objects in intraclass videos and identify attentive regions. Our framework, which comprises the two aforementioned modules, offers a new perspective on directly constructing the correspondence between bottom-up cues and top-down cues when ground-truth information for the reference frames is lacking. Comprehensive experiments demonstrate that the proposed framework achieves state-of-the-art performance. Qing En, Lijuan Duan, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-Domain Image-to-Image Translation via a Unified Circular FrameworkabstractThe image-to-image translation aims to learn the corresponding information between the source and target domains. Several state-of-the-art works have made significant progress based on generative adversarial networks (GANs). However, most existing one-to-one translation methods ignore the correlations among different domain pairs. We argue that there is common information among different domain pairs and it is vital to multiple domain pairs translation. In this paper, we propose a unified circular framework for multiple domain pairs translation, leveraging a shared knowledge module across numerous domains. One selected translation pair can benefit from the complementary information from other pairs, and the sharing knowledge is conducive to mutual learning between domains. Moreover, absolute consistency loss is proposed and applied in the corresponding feature maps to ensure intra-domain consistency. Furthermore, our model can be trained in an end-to-end manner. Extensive experiments demonstrate the effectiveness of our approach on several complex translation scenarios, such as Thermal IR switching, weather changing, and semantic transfer tasks. Yuxi Wang 0001, Zhaoxiang Zhang 0001, Chunfeng Song |
IEEE Trans. Image Process. | 2 |
| 2021 | Attention Guided Multiple Source and Target Domain AdaptationabstractDomain adaptation aims to alleviate the distribution discrepancy between source and target domains. Most conventional methods focus on one target domain setting adapted from one or multiple source domains while neglecting the multi-target domain setting. We argue that different target domains also have complementary information, which is very important for performance improvement. In this paper, we propose an Attention-guided Multiple source-and-target Domain Adaptation (AMDA) method to capture the context dependency information on transferable regions among multiple source and target domains. The innovation points of this paper are as follows: (1) We use numerous adversarial strategies to harvest sufficient information from multiple source and target domains, which extends the generalization and robustness of the feature pools. (2) We propose an intra-domain and inter-domain attention module to explore transferable context information. The proposed attention module can learn domain-invariant representations and reduce the negative transfer by focusing on transferable knowledge. Extensive experiments validate the effectiveness of our method with achieving state-of-the-art performance on several unsupervised domain adaptation datasets. Yuxi Wang 0001, Zhaoxiang Zhang 0001, Chunfeng Song |
IEEE Trans. Image Process. | 2 |
| 2021 | Image Inpainting by End-to-End Cascaded Refinement With Mask AwarenessabstractInpainting arbitrary missing regions is challenging because learning valid features for various masked regions is nontrivial. Though U-shaped encoder-decoder frameworks have been witnessed to be successful, most of them share a common drawback of mask unawareness in feature extraction because all convolution windows (or regions), including those with various shapes of missing pixels, are treated equally and filtered with fixed learned kernels. To this end, we propose our novel mask-aware inpainting solution. Firstly, a Mask-Aware Dynamic Filtering (MADF) module is designed to effectively learn multi-scale features for missing regions in the encoding phase. Specifically, filters for each convolution window are generated from features of the corresponding region of the mask. The second fold of mask awareness is achieved by adopting Point-wise Normalization (PN) in our decoding phase, considering that statistical natures of features at masked points differentiate from those of unmasked points. The proposed PN can tackle this issue by dynamically assigning point-wise scaling factor and bias. Lastly, our model is designed to be an end-to-end cascaded refinement one. Supervision information such as reconstruction loss, perceptual loss and total variation loss is incrementally leveraged to boost the inpainting results from coarse to fine. Effectiveness of the proposed framework is validated both quantitatively and qualitatively via extensive experiments on three public datasets including Places2, CelebA and Paris StreetView. Manyu Zhu, Dongliang He, Xin Li 0106, Chao Li 0034, Fu Li 0003, Xiao Liu 0022, Errui Ding, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 8 |
| 2020 | CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with only image-level labels saves large human effort to annotate pixel-level labels. Cutting-edge approaches rely on various innovative constraints and heuristic rules to generate the masks for every single image. Although great progress has been achieved by these methods, they treat each image independently and do not take account of the relationships across different images. In this paper, however, we argue that the cross-image relationship is vital for weakly supervised segmentation. Because it connects related regions across images, where supplementary representations can be propagated to obtain more consistent and integral regions. To leverage this information, we propose an end-to-end cross-image affinity module, which exploits pixel-level cross-image relationships with only image-level labels. By means of this, our approach achieves 64.3% and 65.3% mIoU on Pascal VOC 2012 validation and test set respectively, which is a new state-of-the-art result by only using image-level labels for weakly supervised semantic segmentation, demonstrating the superiority of our approach. Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan, Chunfeng Song, Jun Xiao 0005 |
AAAI | 2 |
| 2020 | Cascading Convolutional Color ConstancyabstractRegressing the illumination of a scene from the representations of object appearances is popularly adopted in computational color constancy. However, it's still challenging due to intrinsic appearance and label ambiguities caused by unknown illuminants, diverse reflection properties of materials and extrinsic imaging factors (such as different camera sensors). In this paper, we introduce a novel algorithm – Cascading Convolutional Color Constancy (in short, C4) to improve robustness of regression learning and achieve stable generalization capability across datasets (different cameras and scenes) in a unique framework. The proposed C4 method ensembles a series of dependent illumination hypotheses from each cascade stage via introducing a weighted multiply-accumulate loss function, which can inherently capture different modes of illuminations and explicitly enforce coarse-to-fine network optimization. Experimental results on the public Color Checker and NUS 8-Camera benchmarks demonstrate superior performance of the proposed algorithm in comparison with the state-of-the-art methods, especially for more difficult scenes. Huanglin Yu, Ke Chen 0004, Kaiqi Wang, Yanlin Qian, Zhaoxiang Zhang 0001, Kui Jia |
AAAI | 5 |
| 2020 | Instance Guided Proposal Network for Person SearchabstractPerson detection networks have been widely used in person search. These detectors discriminate persons from the background and generate proposals of all the persons from a gallery of scene images for each query. However, such a large number of proposals have a negative influence on the following identity matching process because many distractors are involved. In this paper, we propose a new detection network for person search, named Instance Guided Proposal Network (IGPN), which can learn the similarity between query persons and proposals. Thus, we can decrease proposals according to the similarity scores. To incorporate information of the query into the detection network, we introduce the Siamese region proposal network to Faster-RCNN and we propose improved cross-correlation layers to alleviate the imbalance of parameters distribution. Furthermore, we design a local relation block and a global relation branch to leverage the proposal-proposal relations and query-scene relations, respectively. Extensive experiments show that our method improves the person search performance through decreasing proposals and achieves competitive performance on two large person search benchmark datasets, CUHK-SYSU and PRW. Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 2 |
| 2020 | Bi-Directional Interaction Network for Person SearchabstractExisting works have designed end-to-end frameworks based on Faster-RCNN for person search. Due to the large receptive fields in deep networks, the feature maps of each proposal, cropped from the stem feature maps, involve redundant context information outside the bounding boxes. However, person search is a fine-grained task which needs accurate appearance information. Such context information can make the model fail to focus on persons, so the learned representations lack the capacity to discriminate various identities. To address this issue, we propose a Siamese network which owns an additional instance-aware branch, named Bi-directional Interaction Network (BINet). During the training phase, in addition to scene images, BINet also takes as inputs person patches which help the model discriminate identities based on human appearance. Moreover, two interaction losses are designed to achieve bi-directional interaction between branches at two levels. The interaction can help the model learn more discriminative features for persons in the scene. At the inference stage, only the major branch is applied, so BINet introduces no additional computation. Extensive experiments on two widely used person search benchmarks, CUHK-SYSU and PRW, have shown that our BINet achieves state-of-the-art results among end-to-end methods without loss of efficiency. Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 2 |
| 2020 | Learning Integral Objects With Intra-Class Discriminator for Weakly-Supervised Semantic SegmentationabstractImage-level weakly-supervised semantic segmentation (WSSS) aims at learning semantic segmentation by adopting only image class labels. Existing approaches generally rely on class activation maps (CAM) to generate pseudo-masks and then train segmentation models. The main difficulty is that the CAM estimate only covers partial foreground objects. In this paper, we argue that the critical factor preventing to obtain the full object mask is the classification boundary mismatch problem in applying the CAM to WSSS. Because the CAM is optimized by the classification task, it focuses on the discrimination across different image-level classes. However, the WSSS requires to distinguish pixels sharing the same image-level class to separate them into the foreground and the background. To alleviate this contradiction, we propose an efficient end-to-end Intra-Class Discriminator (ICD) framework, which learns intra-class boundaries to help separate the foreground and the background within each image-level class. Without bells and whistles, our approach achieves the state-of-the-art performance of image label based WSSS, with mIoU 68.0% on the VOC 2012 semantic segmentation benchmark, demonstrating the effectiveness of the proposed approach. Junsong Fan, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 2 |
| 2020 | Large-Scale Object Detection in the Wild From Imbalanced Multi-LabelsabstractTraining with more data has always been the most stable and effective way of improving performance in deep learn-ing era. As the largest object detection dataset so far, OpenImages brings great opportunities and challenges for object detection in general and sophisticated scenarios. However, owing to its semi-automatic collecting and labeling pipeline to deal with the huge data scale, Open Images dataset suffers from label-related problems that objects may explicitly or implicitly have multiple labels and the label distribution is extremely imbalanced. In this work, we quantitatively analyze these label problems and provide a simple but effective solution. We design a concurrent softmax to handle the multi-label problems in object detection and propose a soft-sampling methods with hybrid training scheduler to deal with the label imbalance. Overall, our method yields a dramatic improvement of 3.34 points, leading to the best single model with 60.90 mAP on the public object detection test set of Open Images. And our ensembling result achieves 67.17mAP, which is 4.29 points higher than the first place method last year. Junran Peng, Xingyuan Bu, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
CVPR | 4 |
| 2020 | Context-Aware Attention Network for Image-Text RetrievalabstractAs a typical cross-modal problem, image-text bi-directional retrieval relies heavily on the joint embedding learning and similarity measure for each image-text pair. It remains challenging because prior works seldom explore semantic correspondences between modalities and semantic correlations in a single modality at the same time. In this work, we propose a unified Context-Aware Attention Network (CAAN), which selectively focuses on critical local fragments (regions and words) by aggregating the global context. Specifically, it simultaneously utilizes global inter-modal alignments and intra-modal correlations to discover latent semantic relations. Considering the interactions between images and sentences in the retrieval process, intra-modal correlations are derived from the second-order attention of region-word alignments instead of intuitively comparing the distance between original features. Our method achieves fairly competitive results on two generic image-text retrieval datasets Flickr30K and MS-COCO. Zhen Lei 0001, Zhaoxiang Zhang 0001, Stan Z. Li |
CVPR | 3 |
| 2020 | Boosting Decision-Based Black-Box Adversarial Attacks with Random Sign Flip
Weilun Chen, Zhaoxiang Zhang 0001, Xiaolin Hu 0001, Baoyuan Wu |
ECCV (15) | 2 |
| 2020 | Employing Multi-estimations for Weakly-Supervised Semantic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (17) | 2 |
| 2020 | Generalizing Person Re-Identification by Camera-Aware Invariance Learning and Cross-Domain Mixup
Chuanchen Luo, Chunfeng Song, Zhaoxiang Zhang 0001 |
ECCV (15) | 3 |
| 2020 | Attentive Part-aware Networks for Partial Person Re- identificationabstractPartial person re-identification (re-ID) refers to re-identify a person through occluded images. It suffers from two major challenges, i.e., insufficient training data and incomplete probe image. In this paper, we introduce a part-aware learning method for partial person re-identification. On the one hand, we adopt data augmentation operation to enrich the training data and improve the robustness of the model. On the other hand, we intuitively find that the partial person images usually have fixed percentages of parts, therefore, in partial person re-ID task, the probe image could be cropped from the pictures and divided into several different partial types following fixed ratios. Based on the cropped images, we propose the Cropping Type Consistency (CTC) loss to classify the cropping types of partial images. Moreover, in order to help the network better fit the generated and cropped data, we incorporate the Block Attention Mechanism (BAM) into the framework for attentive learning. To enhance the retrieval performance in the inference stage, we implement cropping on gallery images according to the predicted types of probe partial images. Through calculating feature distances between the partial image and the cropped holistic gallery images, the model can recognize the right person from the gallery. To validate the effectiveness of our approach, we conduct extensive experiments on the partial re- ID benchmarks and achieve state-of-the-art performance. Lijuan Huo, Chunfeng Song, Zhengyi Liu, Zhaoxiang Zhang 0001 |
ICPR | 4 |
| 2020 | Manual-Label Free 3D Detection via An Open-Source SimulatorabstractLiDAR based 3D object detectors typically need a large amount of detailed-labeled point cloud data for training, but these detailed labels are commonly expensive to acquire. In this paper, we propose a manual-label free 3D detection algorithm that leverages the CARLA simulator to generate a large amount of self-labeled training samples and introduces a novel Domain Adaptive VoxelNet (DA-VoxelNet) that can cross the distribution gap from the synthetic data to the real scenario. The self-labeled training samples are generated by a set of high quality 3D models embedded in a CARLA simulator and a proposed LiDAR-guided sampling algorithm. Then a DA-VoxelNet that integrates both a sample-level DA module and an anchor-level DA module is proposed to enable the detector trained by the synthetic data to adapt to real scenario. Experimental results show that the proposed unsupervised DA 3D detector on KITTI evaluation set can achieve 76.66% and 56.64% mAP on BEV mode and 3D mode respectively. The results reveal a promising perspective of training a LIDAR-based 3D detector without any hand-tagged label. Zhen Yang 0009, Chi Zhang 0060, Huiming Guo, Zhaoxiang Zhang 0001 |
ICPR | 4 |
| 2020 | SARPNET: Shape attention regional proposal network for liDAR-based 3D object detection
Yangyang Ye, Houjin Chen, Chi Zhang 0060, Xiaoli Hao, Zhaoxiang Zhang 0001 |
Neurocomputing | 5 |
| 2020 | Beyond Scalar Neuron: Adopting Vector-Neuron Capsules for Long-Term Person Re-IdentificationabstractCurrent person re-identification (re-ID) works mainly focus on the short-term scenario where a person is less likely to change clothes. However, in the long-term re-ID scenario, a person has a great chance to change clothes. A sophisticated re-ID system should take such changes into account. To facilitate the study of long-term re-ID, this paper introduces a large-scale re-ID dataset called “Celeb-reID” to the community. Unlike previous datasets, the same person can change clothes in the proposed Celeb-reID dataset. Images of Celeb-reID are acquired from the Internet using street snap-shots of celebrities. There is a total of 1,052 IDs with 34,186 images making Celeb-reID being the largest long-term re-ID dataset so far. To tackle the challenge of cloth changes, we propose to use vector-neuron (VN) capsules instead of the traditional scalar neurons (SN) to design our network. Compared with SN, one extra-dimensional information in VN can perceive cloth changes of the same person. We introduce a well-designed ReIDCaps network and integrate capsules to deal with the person re-ID task. Soft Embedding Attention (SEA) and Feature Sparse Representation (FSR) mechanisms are adopted in our network for performance boosting. Experiments are conducted on the proposed long-term re-ID dataset and two common short-term re-ID datasets. Comprehensive analyses are given to demonstrate the challenge exposed in our datasets. Experimental results show that our ReIDCaps can outperform existing state-of-the-art methods by a large margin in the long-term scenario.The new dataset and code will be released to facilitate future researches. Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Yi Zhong 0002, Peng Zhang 0057, Zhaoxiang Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | Attention-Aware Sampling via Deep Reinforcement Learning for Action RecognitionabstractDeep learning based methods have achieved remarkable progress in action recognition. Existing works mainly focus on designing novel deep architectures to achieve video representations learning for action recognition. Most methods treat sampled frames equally and average all the frame-level predictions at the testing stage. However, within a video, discriminative actions may occur sparsely in a few frames and most other frames are irrelevant to the ground truth and may even lead to a wrong prediction. As a result, we think that the strategy of selecting relevant frames would be a further important key to enhance the existing deep learning based action recognition. In this paper, we propose an attentionaware sampling method for action recognition, which aims to discard the irrelevant and misleading frames and preserve the most discriminative frames. We formulate the process of mining key frames from videos as a Markov decision process and train the attention agent through deep reinforcement learning without extra labels. The agent takes features and predictions from the baseline model as input and generates importance scores for all frames. Moreover, our approach is extensible, which can be applied to different existing deep learning based action recognition models. We achieve very competitive action recognition performance on two widely used action recognition datasets. Wenkai Dong, Zhaoxiang Zhang 0001, Tieniu Tan |
AAAI | 2 |
| 2019 | Human-Like Delicate Region Erasing Strategy for Weakly Supervised Detection
Qing En, Lijuan Duan, Zhaoxiang Zhang 0001, Xiang Bai |
AAAI | 3 |
| 2019 | Scale-Aware Trident Networks for Object DetectionabstractScale variation is one of the key challenges in object detection. In this work, we first present a controlled experiment to investigate the effect of receptive fields for scale variation in object detection. Based on the findings from the exploration experiments, we propose a novel Trident Network (TridentNet) aiming to generate scale-specific feature maps with a uniform representational power. We construct a parallel multi-branch architecture in which each branch shares the same transformation parameters but with different receptive fields. Then, we adopt a scale-aware training scheme to specialize each branch by sampling object instances of proper scales for training. As a bonus, a fast approximation version of TridentNet could achieve significant improvements without any additional parameters and computational cost compared with the vanilla detector. On the COCO dataset, our TridentNet with ResNet-101 backbone achieves state-of-the-art single-model results of 48.4 mAP. Codes are available at https://git.io/fj5vR. Yanghao Li, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2019 | Spectral Feature Transformation for Person Re-IdentificationabstractWith the surge of deep learning techniques, the field of person re-identification has witnessed rapid progress in recent years. Deep learning based methods focus on learning a discriminative feature space where data points are clustered compactly according to their corresponding identities. Most existing methods process data points individually or only involves a fraction of samples while building a similarity structure. They ignore dense informative connections among samples more or less. The lack of holistic observation eventually leads to inferior performance. To relieve the issue, we propose to formulate the whole data batch as a similarity graph. Inspired by spectral clustering, a novel module termed Spectral Feature Transformation is developed to facilitate the optimization of group-wise similarities. It adds no burden to the inference and can be applied to various scenarios. As a natural extension, we further derive a lightweight re-ranking method named Local Blurring Re-ranking which makes the underlying clustering structure around the probe set more compact. Empirical studies on four public benchmarks show the superiority of the proposed method. Code is available at https://github.com/LuckyDC/SFT_REID. Chuanchen Luo, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2019 | POD: Practical Object Detection With Scale-Sensitive NetworkabstractScale-sensitive object detection remains a challenging task, where most of the existing methods not learn it explicitly and not robust to scale variance. In addition, the most existing methods are less efficient during training or slow during inference, which are not friendly to real-time application. In this paper, we propose a practical object detection with scale-sensitive network.Our method first predicts a global continuous scale ,which shared by all position, for each convolution filter of each network stage. To effectively learn the scale, we average the spatial features and distill the scale from channels. For fast-deployment, we propose a scale decomposition method that transfers the robust fractional scale into combinations of fixed integral scales for each convolution filter, which exploit the dilated convolution. We demonstrate it on one-stage and two-stage algorithm under almost different configure. For practical application, training of our method is of efficiency and simplicity which gets rid of complex data sampling or optimize strategy. During testing, the proposed method requires no extra operation and is very friendly to hardware acceleration like TensorRT and TVM.On the COCO test-dev, our model could achieve a 41.5mAP on one-stage detector and 42.1 mAP on two-stage detectors based on ResNet-101, outperforming baselines by 2.4 and 2.1 respectively without extra FLOPS. Junran Peng, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
ICCV | 3 |
| 2019 | Improving Pedestrian Attribute Recognition With Weakly-Supervised Multi-Scale Attribute-Specific LocalizationabstractPedestrian attribute recognition has been an emerging research topic in the area of video surveillance. To predict the existence of a particular attribute, it is demanded to localize the regions related to the attribute. However, in this task, the region annotations are not available. How to carve out these attribute-related regions remains challenging. Existing methods applied attribute-agnostic visual attention or heuristic body-part localization mechanisms to enhance the local feature representations, while neglecting to employ attributes to define local feature areas. We propose a flexible Attribute Localization Module (ALM) to adaptively discover the most discriminative regions and learns the regional features for each attribute at multiple levels. Moreover, a feature pyramid architecture is also introduced to enhance the attribute-specific localization at low-levels with high-level semantic guidance. The proposed framework does not require additional region annotations and can be trained end-to-end with multi-level deep supervision. Extensive experiments show that the proposed method achieves state-of-the-art results on three pedestrian attribute datasets, including PETA, RAP, and PA-100K. Chufeng Tang, Lu Sheng, Zhaoxiang Zhang 0001, Xiaolin Hu 0001 |
ICCV | 3 |
| 2019 | Sequence Level Semantics Aggregation for Video Object DetectionabstractVideo objection detection (VID) has been a rising research direction in recent years. A central issue of VID is the appearance degradation of video frames caused by fast motion. This problem is essentially ill-posed for a single frame. Therefore, aggregating features from other frames becomes a natural choice. Existing methods rely heavily on optical flow or recurrent neural networks for feature aggregation. However, these methods emphasize more on the temporally nearby frames. In this work, we argue that aggregating features in the full-sequence level will lead to more discriminative and robust features for video object detection. To achieve this goal, we devise a novel Sequence Level Semantics Aggregation (SELSA) module. We further demonstrate the close relationship between the proposed method and the classic spectral clustering method, providing a novel view for understanding the VID problem. We test the proposed method on the ImageNet VID and the EPIC KITCHENS dataset and achieve new state-of-the-art results. Our method does not need complicated postprocessing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. Haiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2019 | CASIA-AHCDB: A Large-Scale Chinese Ancient Handwritten Characters DatabaseabstractThis paper introduces a Chinese Ancient Handwritten Characters Database (CASIA-AHCDB) for character recognition research. The database was built by annotating 11,937 pages of Chinese ancient handwritten documents. It consists of more than 2.2 million annotated handwritten character samples of 10,350 categories. According to the source of these documents, the database is divided into two datasets of different styles: Complete Library in Four Sections (AHCDB-style1) and Ancient Buddhist Scriptures (AHCDB-style2). Each dataset can be divided into three parts based on its applications. The first part, called basic category set, contains samples of common categories in two datasets, and is suitable for basic character recognition task. The second part, called enhanced category set, is mainly used for open-set character recognition task based on the basic character recognition. The third part, called the reserved category set, can be used in many pattern recognition tasks in the future. Based on the large category set, the various writing styles and the imbalanced sample number per category, CASIA-AHCDB can also be used for various classification and learning tasks such as transfer learning, few-shot learning. We performed experiments of basic character recognition on the basic category set, and report the results for benchmark. More techniques can be evaluated on this challenging database in the future. Dahan Wang, Xu-Yao Zhang, Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
ICDAR | 5 |
| 2019 | Relational Network for Skeleton-Based Action RecognitionabstractWith the fast development of effective and low-cost human skeleton capture systems, skeleton-based action recognition has attracted much attention recently. Most existing methods use Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) to extract spatio-temporal information embedded in the skeleton sequences for action recognition. However, these approaches are limited in the ability of relational modeling in a single skeleton, due to the loss of important structural information when converting the raw skeleton data to adapt to the input format of CNN or RNN. In this paper, we propose an Attentional Recurrent Relational Network-LSTM (ARRN-LSTM) to simultaneously model spatial configurations and temporal dynamics in skeletons for action recognition. We introduce the Recurrent Relational Network to learn the spatial features in a single skeleton, followed by a multi-layer LSTM to learn the temporal features in the skeleton sequences. Between the two modules, we design an adaptive attentional module to focus attention on the most discriminative parts in the single skeleton. To exploit the complementarity from different geometries in the skeleton for sufficient relational modeling, we design a two-stream architecture to learn the structural features among joints and lines simultaneously. Extensive experiments are conducted on several popular skeleton datasets and the results show that the proposed approach achieves better results than most mainstream methods. Wu Zheng, Zhaoxiang Zhang 0001, Yan Huang 0008, Liang Wang 0001 |
ICME | 3 |
| 2019 | Efficient Neural Architecture Transformation Search in Channel-Level for Object DetectionabstractRecently, Neural Architecture Search has achieved great success in large-scale image classification. In contrast, there have been limited works focusing on architecture search for object detection, mainly because the costly ImageNet pretraining is always required for detectors. Training from scratch, as a substitute, demands more epochs to converge and brings no computation saving. To overcome this obstacle, we introduce a practical neural architecture transformation search(NATS) algorithm for object detection in this paper. Instead of searching and constructing an entire network, NATS explores the architecture space on the base of existing network and reusing its weights. We propose a novel neural architecture search strategy in channel-level instead of path-level and devise a search space specially targeting at object detection. With the combination of these two designs, an architecture transformation scheme could be discovered to adapt a network designed for image classification to task of object detection. Since our method is gradient-based and only searches for a transformation scheme, the weights of models pretrained in ImageNet could be utilized in both searching and retraining stage, which makes the whole process very efficient. The transformed network requires no extra parameters and FLOPs, and is friendly to hardware optimization, which is practical to use in real-time application. In experiments, we demonstrate the effectiveness of NATS on networks like {\em ResNet} and {\em ResNeXt}. Our transformed networks, combined with various detection frameworks, achieve significant improvements on the COCO dataset while keeping fast. Junran Peng, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
NeurIPS | 3 |
| 2019 | Uncertainty-optimized deep learning model for small-scale person re-identification
Cairong Zhao, Di Zang, Zhaoxiang Zhang 0001, Wangmeng Zuo, Duoqian Miao 0001 |
Sci. China Inf. Sci. | 4 |
| 2019 | SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance RecognitionabstractObject detection and instance recognition play a central role in many AI applications like autonomous driving, video surveillance and medical image analysis. However, training object detection models on large scale datasets remains computationally expensive and time consuming. This paper presents an efficient and open source object detection framework called SimpleDet which enables the training of state-of-the-art detection models on consumer grade hardware at large scale. SimpleDet covers a wide range of models including both high-performance and high-speed ones. SimpleDet is well-optimized for both low precision training and distributed training and achieves 70% higher throughput for the Mask R-CNN detector compared with existing frameworks. Codes, examples and documents of SimpleDet can be found at https://github.com/tusimple/simpledet. Yuntao Chen, Chenxia Han, Yanghao Li, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
J. Mach. Learn. Res. | 7 |
| 2019 | Spatiotemporal distilled dense-connectivity network for video action recognition
Zhaoxiang Zhang 0001 |
Pattern Recognit. | 2 |
| 2019 | Semi-supervised domain adaptation via Fredholm integral based kernel methods
Wei Wang 0061, Hao Wang 0005, Zhaoxiang Zhang 0001, Chen Zhang 0003, Yang Gao 0021 |
Pattern Recognit. | 3 |
| 2019 | Image Caption Generation with Part of Speech Guidance
Xinwei He 0001, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang 0001, Weisheng Dong |
Pattern Recognit. Lett. | 5 |
| 2019 | Generative adversarial dehaze mapping nets
Ce Li 0001, Zhaoxiang Zhang 0001, Shaoyi Du |
Pattern Recognit. Lett. | 3 |
| 2019 | Deep Learning for Pattern Recognition
Zhaoxiang Zhang 0001, Shiguang Shan, Yi Fang 0006, Ling Shao 0001 |
Pattern Recognit. Lett. | 1 |
| 2019 | Multi-Pseudo Regularized Label for Generated Data in Person Re-IdentificationabstractSufficient training data normally is required to train deeply learned models. However, due to the expensive manual process for labelling large number of images (i.e., annotation), the amount of available training data (i.e., real data) is always limited. To produce more data for training a deep network, Generative Adversarial Network (GAN) can be used to generate artificial sample data (i.e., generated data). However, the generated data usually does not have annotation labels. To solve this problem, in this paper, we propose a virtual label called Multi-pseudo Regularized Label (MpRL) and assign it to the generated data. With MpRL, the generated data will be used as the supplementary of real training data to train a deep neural network in a semi-supervised learning fashion. To build the corresponding relationship between the real data and generated data, MpRL assigns each generated data a proper virtual label which reflects the likelihood of the affiliation of the generated data to predefined training classes in the real data domain. Unlike the traditional label which usually is a single integral number, the virtual label proposed in this work is a set of weight-based values each individual of which is a number in (0,1] called multi-pseudo label and reflects the degree of relation between each generated data to every pre-defined class of real data. A comprehensive evaluation is carried out by adopting two state-of-the-art convolutional neural networks (CNNs) in our experiments to verify the effectiveness of MpRL. Experiments demonstrate that by assigning MpRL to generated data, we can further improve the person re-ID performance on five re-ID datasets, i.e., Market-1501, DukeMTMC-reID, CUHK03, VIPeR, and CUHK01. The proposed method obtains +6.29%, +6.30%, +5.58%, +5.84%, and +3.48% improvements in rank-1 accuracy over a strong CNN baseline on the five datasets respectively, and outperforms state-of-the-art methods. Yan Huang 0023, Jingsong Xu, Qiang Wu 0001, Zhedong Zheng, Zhaoxiang Zhang 0001, Jian Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2019 | Dynamic Collaborative TrackingabstractCorrelation filter has been demonstrated remarkable success for visual tracking recently. However, most existing methods often face model drift caused by several factors, such as unlimited boundary effect, heavy occlusion, fast motion, and distracter perturbation. To address the issue, this paper proposes a unified dynamic collaborative tracking framework that can perform more flexible and robust position prediction. Specifically, the framework learns the object appearance model by jointly training the objective function with three components: target regression submodule, distracter suppression submodule, and maximum margin relation submodule. The first submodule mainly takes advantage of the circulant structure of training samples to obtain the distinguishing ability between the target and its surrounding background. The second submodule optimizes the label response of the possible distracting region close to zero for reducing the peak value of the confidence map in the distracting region. Inspired by the structure output support vector machines, the third submodule is introduced to utilize the differences between target appearance representation and distracter appearance representation in the discriminative mapping space for alleviating the disturbance of the most possible hard negative samples. In addition, a CUR filter as an assistant detector is embedded to provide effective object candidates for alleviating the model drift problem. Comprehensive experimental results show that the proposed approach achieves the state-of-the-art performance in several public benchmark data sets. Guibo Zhu, Zhaoxiang Zhang 0001, Jinqiao Wang, Yi Wu 0001, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities TransferabstractWe have witnessed rapid evolution of deep neural network architecture design in the past years. These latest progresses greatly facilitate the developments in various areas such as computer vision and natural language processing. However, along with the extraordinary performance, these state-of-the-art models also bring in expensive computational cost. Directly deploying these models into applications with real-time requirement is still infeasible. Recently, Hinton et al. have shown that the dark knowledge within a powerful teacher model can significantly help the training of a smaller and faster student network. These knowledge are vastly beneficial to improve the generalization ability of the student model. Inspired by their work, we introduce a new type of knowledge---cross sample similarities for model compression and acceleration. This knowledge can be naturally derived from deep metric learning model. To transfer them, we bring the "learning to rank" technique into deep metric learning formulation. We test our proposed DarkRank method on various metric learning tasks including pedestrian re-identification, image retrieval and image clustering. The results are quite encouraging. Our method can improve over the baseline method by a large margin. Moreover, it is fully compatible with other existing methods. When combined, the performance can be further boosted. Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
AAAI | 3 |
| 2018 | CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual GenerationabstractVisual and audio modalities are two symbiotic modalities underlying videos, which contain both common and complementary information. If they can be mined and fused sufficiently, performances of related video tasks can be significantly enhanced. However, due to the environmental interference or sensor fault, sometimes, only one modality exists while the other is abandoned or missing. By recovering the missing modality from the existing one based on the common information shared between them and the prior information of the specific modality, great bonus will be gained for various vision tasks. In this paper, we propose a Cross-Modal Cycle Generative Adversarial Network (CMCGAN) to handle cross-modal visual-audio mutual generation. Specifically, CMCGAN is composed of four kinds of subnetworks: audio-to-visual, visual-to-audio, audio-to-audio and visual-to-visual subnetworks respectively, which are organized in a cycle architecture. CMCGAN has several remarkable advantages. Firstly, CMCGAN unifies visual-audio mutual generation into a common framework by a joint corresponding adversarial loss. Secondly, through introducing a latent vector with Gaussian distribution, CMCGAN can handle dimension and structure asymmetry over visual and audio modalities effectively. Thirdly, CMCGAN can be trained end-to-end to achieve better convenience. Benefiting from CMCGAN, we develop a dynamic multimodal classification network to handle the modality missing problem. Abundant experiments have been conducted and validate that CMCGAN obtains the state-of-the-art cross-modal visual-audio generation results. Furthermore, it is shown that the generated modality achieves comparable effects with those of original modality, which demonstrates the effectiveness and advantages of our proposed method. Zhaoxiang Zhang 0001, He Guan |
AAAI | 2 |
| 2018 | Integrating Both Visual and Audio Cues for Enhanced Video CaptionabstractVideo caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring the synchronized audio cues. We propose three multimodal deep fusion strategies to maximize the benefits of visual-audio resonance information. The first one explores the impact on cross-modalities feature fusion from low to high order. The second establishes the visual-audio short-term dependency by sharing weights of corresponding front-end networks. The third extends the temporal dependency to long-term through sharing multimodal memory across visual and audio modalities. Extensive experiments have validated the effectiveness of our three cross-modalities fusion strategies on two benchmark datasets, including Microsoft Research Video to Text (MSRVTT) and Microsoft Video Description (MSVD). It is worth mentioning that sharing weight can coordinate visual- audio feature fusion effectively and achieve the state-of-art performance on both BELU and METEOR metrics. Furthermore, we first propose a dynamic multimodal feature fusion framework to deal with the part modalities missing case. Experimental results demonstrate that even in the audio absence mode, we can still obtain comparable results with the aid of the additional audio modality inference module. Zhaoxiang Zhang 0001, He Guan |
AAAI | 2 |
| 2018 | Hard-Aware Point-to-Set Deep Metric for Person Re-identification
Rui Yu 0002, Zhiyong Dou, Song Bai 0001, Zhaoxiang Zhang 0001, Yongchao Xu, Xiang Bai |
ECCV (16) | 4 |
| 2018 | Diversified Dual Domain-Adversarial Neural NetworksabstractThe application cost machine learning methods often rely on the availability of large-scale data collection and annotation, especially in the cases of cross-domain learning. One way to circumvent this cost is constructing models to synthesize data and provide automatic annotation. Although these models are attractive, they often can not be generalized from synthetic images to real-world images. Therefore, domain adaptive algorithm is needed to improve these models, so that they can be applied successfully. In this paper, we propose a novel unsupervised domain adaptive framework codenamed D-DANN inspired by the theory of adversarial learning. We apply the discriminator to diverse the features extracted from dual branch CNN. We can obtain more sufficient shared representation across domains by the proposed dual feature extractors. The framework can be easily adapt to most popular CNN models to improve the representation power. We implement the D-DANN with several popular CNN models including LeNet, AlexNet and so on. Using these D-DANN enhanced neural networks, we conduct extensive experiments on several pairs of domain adaptive validation datasets. The results show that our approach can efficiently enhance domain adaptive capability of general CNN models for unlabeled data. Yuchun Fang, Qiulong Yuan, Wei Zhang 0021, Zhaoxiang Zhang 0001 |
ICPR | 4 |
| 2018 | Inception Donut Convolution for Top-down Semantic SegmentationabstractOne of recent trends in network architecture design confirms that the inception-block convolutional group is efficient, since it can aggregate spatial context information in lower dimensions without causing significant loss in representative capabilities. We believe that not only the strong correlation between adjacent cells, multi-scale feature extraction also plays a vital role in this novel module. In this paper, we extend the profits of the block to a top-down donut convolutional network for semantic segmentation task. Our network automatically learns rich convolution kernels to capture more structure prior. In the inception-block design, it overcomes the limitations in larger kernel size and adaptively captures different object-scales contexts without chain sampling. Our experiments demonstrate that the proposed inception-block donut convolutional network is orthogonal and can further improve the performance of most off-the-shelf bottom-up based methods. He Guan, Zhaoxiang Zhang 0001, Tieniu Tan |
ICPR | 2 |
| 2018 | Deep Temporal Feature Encoding for Action RecognitionabstractHuman action recognition is an important task in computer vision. Recently, deep learning methods for video action recognition have developed rapidly. A popular way to tackle this problem is known as two-stream methods which take both spatial and temporal modalities into consideration. These methods often treat sparsely-sampled frames as input and video labels as supervision. Because of such sampling strategy, they are typically limited to processing shorter sequences, which might cause the problems such as suffering from the confusion by partial observation. In this paper we propose a novel video feature representation method, called Deep Temporal Feature Encoding (DTE). It could aggregate frame-level features into a robust and global video-level representation. Firstly, we sample enough RGB frames and optical flow stacks across the whole video. Then we use a deep temporal feature encoding layer to construct a strong video feature. Lastly, end-to-end training is applied so that our video representation could be global and sequence-aware. Comprehensive experiments are conducted on two public datasets: HMDB51 and UCF101. Experimental results demonstrate that DTE achieves the competitive state-of-the-art performance on both datasets. Zhaoxiang Zhang 0001, Yan Huang 0008, Liang Wang 0001 |
ICPR | 2 |
| 2018 | Rethinking ReLU to Train Better CNNsabstractMost of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the most widely used. In this paper, we argue that the designed structure with the equal ratio between these two layers may not be the best choice since it could result in the poor generalization ability. Thus, we try to investigate a more suitable method on using ReL U to explore the better network architectures. Specifically, we propose a proportional module to keep the ratio between convolution and ReLU amount to be N:m (n>m). The proportional module can be applied in almost all networks with no extra computational cost to improve the performance. Comprehensive experimental results indicate that the proposed method achieves better performance on different benchmarks with different network architectures, thus verify the superiority of our work. Gangming Zhao, Zhaoxiang Zhang 0001, He Guan, Peng Tang 0005, Jingdong Wang 0001 |
ICPR | 2 |
| 2018 | Accelerating the Classification of Very Deep Convolutional Network by A Cascading ApproachabstractLarge convolutional networks have achieved impressive classification performances recently. To achieve better performance, convolutional network tends to develop into deeper. However, the increase of network depth causes the linear growth of computational complexity, but cannot bring equivalent increase to the classification accuracy. To alleviate this inconsistence, we propose a cascading approach to accelerate the classification of very deep convolutional neural network. By exploiting the entropy metric to analyze the statistic differences of basic networks between the correctly and mistakenly classified images, we can assign the easily distinguished images to the shallow networks for reducing the computational complexity, and leave the difficultly classified images to the deep networks for maintaining the overall performance. Besides, the proposed cascaded networks can take advantage of the complementarity between different networks, which may boost the classification accuracy compared to the deepest network. We perform the experiments using residual networks of different depths on cifar100 dataset, on the condition of obtaining the similar accuracy to the deepest network, the results show that our cascaded ResNet32-ResNet110 and cascaded ResNet32-ResNet164 can reduce the computation time by 48.6% and 44.3% compared to ResNet110 and ResNet164, respectively. And the cascaded ResNet32-ResNet110-ResNet164 can reduce the computation time by 85.4% compared to the very deep Resnet1001. Wu Zheng, Zhaoxiang Zhang 0001 |
ICPR | 2 |
| 2018 | Multi-task Layout Analysis for Historical Handwritten Documents Using Fully Convolutional NetworksabstractLayout analysis is a fundamental process in document image analysis and understanding. It consists of several sub-processes such as page segmentation, text line segmentation, baseline detection and so on. In this work, we propose a multi-task layout analysis method that use a single FCN model to solve the above three problems simultaneously. The FCN is trained to segment the document image into different regions and detect the center line of each text line by classifying pixels into different categories. By supervised learning on document images with pixel-wise labels, the FCN can extract discriminative features and perform pixel-wise classification accurately. After pixel-wise classification, post-processing steps are taken to reduce noises, correct wrong segmentations and find out overlapping regions. Experimental results on the public dataset DIVA-HisDB containing challenging medieval manuscripts demonstrate the effectiveness and superiority of the proposed method. Zhaoxiang Zhang 0001, Cheng-Lin Liu 0001 |
IJCAI | 3 |
| 2018 | Deep Convolutional Neural Networks with Merge-and-Run MappingsabstractA deep residual network, built by stacking a sequence of residual blocks, is easy to train, because identity mappings skip residual branches and thus improve information flow. To further reduce the training difficulty, we present a simple network architecture, deep merge-and-run neural networks. The novelty lies in a modularized building block, merge-and-run block, which assembles residual branches in parallel through a merge-and-run mapping: average the inputs of these residual branches (Merge), and add the average to the output of each residual branch as the input of the subsequent residual branch (Run), respectively. We show that the merge-and-run mapping is a linear idempotent function in which the transformation matrix is idempotent, and thus improves information flow, making training easy. In comparison with residual networks, our networks enjoy compelling advantages: they contain much shorter paths and the width, i.e., the number of channels, is increased, and the time complexity remains unchanged. We evaluate the performance on the standard recognition tasks. Our approach demonstrates consistent improvements over ResNets with the comparable setup, and achieves competitive results (e.g., 3.06% testing error on CIFAR-10, 17.55% on CIFAR-100, 1.51% on SVHN). Mingjie Li 0007, Depu Meng, Xi Li 0001, Zhaoxiang Zhang 0001, Yueting Zhuang, Zhuowen Tu, Jingdong Wang 0001 |
IJCAI | 5 |
| 2018 | Conditional Expression Synthesis with Face Parsing TransformationabstractFacial expression synthesis with various intensities is a challenging synthesis task due to large identity appearance variations and a paucity of efficient means for intensity measurement. This paper advances the expression synthesis domain by the introduction of a Couple-Agent Face Parsing based Generative Adversarial Network (CAFP-GAN) that unites the knowledge of facial semantic regions and controllable expression signals. Specially, we employ a face parsing map as a controllable condition to guide facial texture generation with a special expression, which can provide a semantic representation of every pixel of facial regions. Our method consists of two sub-networks: face parsing prediction network (FPPN) uses controllable labels (expression and intensity) to generate a face parsing map transformation that corresponds to the labels from the input neutral face, and facial expression synthesis network (FESN) makes the pretrained FPPN as a part of it to provide the face parsing map as a guidance for expression synthesis. To enhance the reality of results, couple-agent discriminators are served to distinguish fake-real pairs in both two sub-nets. Moreover, we only need the neutral face and the labels to synthesize the unknown expression with different intensities. Experimental results on three popular facial expression databases show that our method has the compelling ability on continuous expression synthesis. Zhihe Lu, Tanhao Hu, Lingxiao Song, Zhaoxiang Zhang 0001, Ran He 0001 |
ACM Multimedia | 4 |
| 2018 | View Decomposition and Adversarial for Semantic Segmentation
He Guan, Zhaoxiang Zhang 0001 |
PRICAI | 2 |
| 2018 | Weakly-Supervised Object Localization by Cutting Background with Deep Reinforcement Learning
Wu Zheng, Zhaoxiang Zhang 0001 |
PRICAI | 2 |
| 2018 | On the role of sparsity in feature selection and an innovative method LRMI
Yuchun Fang, Qiulong Yuan, Zhaoxiang Zhang 0001 |
Neurocomputing | 3 |
| 2018 | Improving context-sensitive similarity via smooth neighborhood for object retrieval
Song Bai 0001, Shaoyan Sun, Xiang Bai, Zhaoxiang Zhang 0001, Qi Tian 0001 |
Pattern Recognit. | 4 |
| 2018 | Efficient auto-refocusing for light field camera
Chi Zhang 0060, Guangqi Hou, Zhaoxiang Zhang 0001, Zhenan Sun, Tieniu Tan |
Pattern Recognit. | 3 |
| 2018 | GII Representation-Based Cross-View Gait Recognition by Discriminative Projection With List-Wise ConstraintsabstractRemote person identification by gait is one of the most important topics in the field of computer vision and pattern recognition. However, gait recognition suffers severely from the appearance variance caused by the view change. It is very common that gait recognition has a high performance when the view is fixed but the performance will have a sharp decrease when the view variance becomes significant. Existing approaches have tried all kinds of strategies like tensor analysis or view transform models to slow down the trend of performance decrease but still have potential for further improvement. In this paper, a discriminative projection with list-wise constraints (DPLC) is proposed to deal with view variance in cross-view gait recognition, which has been further refined by introducing a rectification term to automatically capture the principal discriminative information. The DPLC with rectification (DPLCR) embeds list-wise relative similarity measurement among intraclass and inner-class individuals, which can learn a more discriminative and robust projection. Based on the original DPLCR, we have introduced the kernel trick to exploit nonlinear cross-view correlations and extended DPLCR to deal with the problem of multiview gait recognition. Moreover, a simple yet efficient gait representation, namely gait individuality image (GII), based on gait energy image is proposed, which could better capture the discriminative information for cross view gait recognition. Experiments have been conducted in the CASIA-B database and the experimental results demonstrate the outstanding performance of both the DPLCR framework and the new GII representation. It is shown that the DPLCR-based cross-view gait recognition has outperformed the-state-of-the-art approaches in almost all cases under large view variance. The combination of the GII representation and the DPLCR has further enhanced the performance to be a new benchmark for cross-view gait recognition. Zhaoxiang Zhang 0001, Jiaxin Chen 0002, Qiang Wu 0001, Ling Shao 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Dynamic Multi-Task Learning with Convolutional Neural NetworkabstractMulti-task learning and deep convolutional neural network (CNN) have been successfully used in various fields. This paper considers the integration of CNN and multi-task learning in a novel way to further improve the performance of multiple related tasks. Existing multi-task CNN models usually empirically combine different tasks into a group which is then trained jointly with a strong assumption of model commonality. Furthermore, traditional approaches usually only consider small number of tasks with rigid structure, which is not suitable for large-scale applications. In light of this, we propose a dynamic multi-task CNN model to handle these problems. The proposed model directly learns the task relations from data instead of subjective task grouping. Due to its flexible structure, it supports task-wise incremental training, which is useful for efficient training of massive tasks. Specifically, we add a new task transfer connection (TTC) between the layers of each task. The learned TTC is able to reflect the correlation among different tasks guiding the model dynamically adjusting the multiplexing of the information among different tasks. With the help of TTC, multiple related tasks can further boost the whole performance for each other. Experiments demonstrate that the proposed dynamic multi-task CNN model outperforms traditional approaches. Yuchun Fang, Zhengyan Ma, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Xiang Bai |
IJCAI | 3 |
| 2017 | Random Shifting for CNN: a Solution to Reduce Information Loss in Down-Sampling LayersabstractDown-sampling is widely adopted in deep convolutional neural networks (DCNN) for reducing the number of network parameters while preserving the transformation invariance. However, it cannot utilize information effectively because it only adopts a fixed stride strategy, which may result in poor generalization ability and information loss. In this paper, we propose a novel random strategy to alleviate these problems by embedding random shifting in the down-sampling layers during the training process. Random shifting can be universally applied to diverse DCNN models to dynamically adjust receptive fields by shifting kernel centers on feature maps in different directions. Thus, it can generate more robust features in networks and further enhance the transformation invariance of down-sampling operators. In addition, random shifting cannot only be integrated in all down-sampling layers including strided convolutional layers and pooling layers, but also improve performance of DCNN with negligible additional computational cost. We evaluate our method in different tasks (e.g., image classification and segmentation) with various network architectures (i.e., AlexNet, FCN and DFN-MR). Experimental results demonstrate the effectiveness of our proposed method. Gangming Zhao, Jingdong Wang 0001, Zhaoxiang Zhang 0001 |
IJCAI | 3 |
| 2017 | Diverse Neuron Type Selection for Convolutional Neural NetworksabstractThe activation function for neurons is a prominent element in the deep learning architecture for obtaining high performance. Inspired by neuroscience findings, we introduce and define two types of neurons with different activation functions for artificial neural networks: excitatory and inhibitory neurons, which can be adaptively selected by self-learning. Based on the definition of neurons, in the paper we not only unify the mainstream activation functions, but also discuss the complementariness among these types of neurons. In addition, through the cooperation of excitatory and inhibitory neurons, we present a compositional activation function that leads to new state-of-the-art performance comparing to rectifier linear units. Finally, we hope that our framework not only gives a basic unified framework of the existing activation neurons to provide guidance for future design, but also contributes neurobiological explanations which can be treated as a window to bridge the gap between biology and computer science. Guibo Zhu, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
IJCAI | 2 |
| 2017 | Local structured representation for generic object detection
Junge Zhang, Kaiqi Huang, Tieniu Tan, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 4 |
| 2017 | Spectral attribute learning for visual regression
Ke Chen 0004, Kui Jia, Zhaoxiang Zhang 0001, Joni-Kristian Kämäräinen |
Pattern Recognit. | 3 |
| 2017 | Learning to Classify Fine-Grained Categories with Privileged Visual-Semantic MisalignmentabstractImage categorisation is an active yet challenging research topic in computer vision, which is to classify the images according to their semantic content. Recently, fine-grained object categorisation has attracted wide attention and remains difficult due to feature inconsistency caused by smaller inter-class and larger intra-class variation as well as large varying poses. Most of the existing frameworks focused on exploiting a more discriminative imagery representation or developing a more robust classification framework to mitigate the suffering. The concern has recently been paid to discovering the dependency across fine-grained class labels based on Convolutional Neural Networks. Encouraged by the success of semantic label embedding to discover the fine-grained class labels' correlation, this paper exploits the misalignment between visual feature space and semantic label embedding space and incorporates it as a privileged information into a cost-sensitive learning framework. Owing to capturing both the variation of imagery feature representation and also the label correlation in the semantic label embedding space, such a visual-semantic misalignment can be employed to reflect the importance of instances, which is more informative that conventional cost-sensitivities. Experiment results demonstrate the effectiveness of the proposed framework on public fine-grained benchmarks with achieving superior performance to state-of-the-arts. Ke Chen 0004, Zhaoxiang Zhang 0001 |
IEEE Trans. Big Data | 2 |
| 2017 | GIFT: Towards Scalable 3D Shape RetrievalabstractProjective analysis is an important solution in three-dimensional (3D) shape retrieval, since human visual perceptions of 3D shapes rely on various 2D observations from different viewpoints. Although multiple informative and discriminative views are utilized, most projection-based retrieval systems suffer from heavy computational cost, and thus cannot satisfy the basic requirement of scalability for search engines. In the past three years, shape retrieval contest (SHREC) pays much attention to the scalability of 3D shape retrieval algorithms, and organizes several large scale tracks accordingly [1]- [3]. However, the experimental results indicate that conventional algorithms cannot be directly applied to large datasets. In this paper, we present a real-time 3D shape search engine based on the projective images of 3D shapes. The real-time property of our search engine results from the following aspects: (1) efficient projection and view feature extraction using GPU acceleration; (2) the first inverted file, called F-IF, is utilized to speed up the procedure of multiview matching; and (3) the second inverted file, which captures a local distribution of 3D shapes in the feature manifold, is adopted for efficient context-based reranking. As a result, for each query the retrieval task can be finished within one second despite the necessary cost of IO overhead. We name the proposed 3D shape search engine, which combines GPU acceleration and inverted file (t wice), as GIFT. Besides its high efficiency, GIFT also outperforms state-of-the-art methods significantly in retrieval accuracy on various shape benchmarks (ModelNet40 dataset, ModelNet10 dataset, PSB dataset, McGill dataset) and competitions (SHREC14LSGTB, ShapeNet Core55, WM-SHREC07). Song Bai 0001, Xiang Bai, Zhaoxiang Zhang 0001, Qi Tian 0001, Longin Jan Latecki |
IEEE Trans. Multim. | 4 |
| 2017 | Pedestrian Counting With Back-Propagated Information and Target Drift RemedyabstractPedestrian density is one of the important factors in designing visual surveillance and intelligent transportation systems, but it is challenging to obtain accurate and robust estimates because of both inconsistent crowd patterns in the scenes and target drift caused by imbalanced data distribution. Most of existing global regression frameworks focus on the former challenge to improve the robustness of regression learning, but very few work concerns on mitigating the suffering from the latter one. This paper proposes a novel counting-by-regression framework to utilize the importance of training samples to improve the robustness against inconsistent feature-target relationship based on a recently-proposed learning paradigm-learning with privileged information. To this end, the concept of back-propagation is for the first time considered to select more informative samples contributed to robust fitting performance. Moreover, the direction of target drift along the continuously-changing target dimension is discovered by learning local classifiers under different situation of pedestrian density, which can thus be exploited in our algorithm to further boost the performance. Experimental evaluation on the public UCSD and shopping Mall benchmarks verifies that our approach significantly beats the state-of-the-art counting-by-regression frameworks. Ke Chen 0004, Zhaoxiang Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | GIFT: A Real-Time and Scalable 3D Shape Search EngineabstractProjective analysis is an important solution for 3D shape retrieval, since human visual perceptions of 3D shapes rely on various 2D observations from different view points. Although multiple informative and discriminative views are utilized, most projection-based retrieval systems suffer from heavy computational cost, thus cannot satisfy the basic requirement of scalability for search engines. In this paper, we present a real-time 3D shape search engine based on the projective images of 3D shapes. The real-time property of our search engine results from the following aspects: (1) efficient projection and view feature extraction using GPU acceleration, (2) the first inverted file, referred as F-IF, is utilized to speed up the procedure of multi-view matching, (3) the second inverted file (S-IF), which captures a local distribution of 3D shapes in the feature manifold, is adopted for efficient context-based reranking. As a result, for each query the retrieval task can be finished within one second despite the necessary cost of IO overhead. We name the proposed 3D shape search engine, which combines GPU acceleration and Inverted File (Twice), as GIFT. Besides its high efficiency, GIFT also outperforms the state-of-the-art methods significantly in retrieval accuracy on various shape benchmarks and competitions. Song Bai 0001, Xiang Bai, Zhaoxiang Zhang 0001, Longin Jan Latecki |
CVPR | 4 |
| 2016 | Smooth Neighborhood Structure Mining on Multiple Affinity Graphs with Applications to Context-Sensitive Similarity
Song Bai 0001, Shaoyan Sun, Xiang Bai, Zhaoxiang Zhang 0001, Qi Tian 0001 |
ECCV (2) | 4 |
| 2016 | Facial Age Estimation Using Robust Label DistributionabstractFacial age estimation, to predict the persons' exact ages given facial images, usually encounters the data sparsity problem due to the difficulties in data annotation. To mitigate the suffering from sparse data, a recent label distribution learning (LDL) algorithm attempts to embed label correlation into a classification based framework. However, the conventional label distribution learning framework only considers correlations across the neighbouring variables (ages), which omits the intrinsic complexity of age classes during different ageing periods (age groups). In the light of this, we introduce a novel concept of robust label distribution for scalar-valued labels, which is designed to encode the age scalars into label distribution matrices, i.e. two-dimensional Gaussian distributions along age classes and age groups respectively. Overcoming the limitations of conventional hard group boundaries in age grouping and capturing intrinsic inter-group dependency, our framework achieves robust and competitive performance over the conventional algorithms on two popular benchmarks for human age estimation. Ke Chen 0004, Joni-Kristian Kämäräinen, Zhaoxiang Zhang 0001 |
ACM Multimedia | 3 |
| 2016 | Corrections to "Relevance Metric Learning for Person Re-Identification by Exploiting Listwise Similarities"
Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Compressive Sequential Learning for Action Similarity LabelingabstractHuman action recognition in videos has been extensively studied in recent years due to its wide range of applications. Instead of classifying video sequences into a number of action categories, in this paper, we focus on a particular problem of action similarity labeling (ASLAN), which aims at verifying whether a pair of videos contain the same type of action or not. To address this challenge, a novel approach called compressive sequential learning (CSL) is proposed by leveraging the compressive sensing theory and sequential learning. We first project data points to a low-dimensional space by effectively exploring an important property in compressive sensing: the restricted isometry property. In particular, a very sparse measurement matrix is adopted to reduce the dimensionality efficiently. We then learn an ensemble classifier for measuring similarities between pairwise videos by iteratively minimizing its empirical risk with the AdaBoost strategy on the training set. Unlike conventional AdaBoost, the weak learner for each iteration is not explicitly defined and its parameters are learned through greedy optimization. Furthermore, an alternative of CSL named compressive sequential encoding is developed as an encoding technique and followed by a linear classifier to address the similarity-labeling problem. Our method has been systematically evaluated on four action data sets: ASLAN, KTH, HMDB51, and Hollywood2, and the results show the effectiveness and superiority of our method for ASLAN. Jie Qin 0004, Li Liu 0004, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Dither modulation of significant amplitude difference for wavelet based robust watermarking
Chunlei Li 0004, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Bin Ma 0004, Di Huang 0001 |
Neurocomputing | 2 |
| 2015 | Enhancing person re-identification by integrating gait biometric
Zheng Liu 0014, Zhaoxiang Zhang 0001, Qiang Wu 0001, Yunhong Wang 0001 |
Neurocomputing | 2 |
| 2015 | Crowd counting in public video surveillance by label distribution learning
Zhaoxiang Zhang 0001, Xin Geng 0001 |
Neurocomputing | 1 |
| 2015 | Relevance Metric Learning for Person Re-Identification by Exploiting Listwise SimilaritiesabstractPerson re-identification aims to match people across non-overlapping camera views, which is an important but challenging task in video surveillance. In order to obtain a robust metric for matching, metric learning has been introduced recently. Most existing works focus on seeking a Mahalanobis distance by employing sparse pairwise constraints, which utilize image pairs with the same person identity as positive samples, and select a small portion of those with different identities as negative samples. However, this training strategy has abandoned a large amount of discriminative information, and ignored the relative similarities. In this paper, we propose a novel relevance metric learning method with listwise constraints (RMLLCs) by adopting listwise similarities, which consist of the similarity list of each image with respect to all remaining images. By virtue of listwise similarities, RMLLC could capture all pairwise similarities, and consequently learn a more discriminative metric by enforcing the metric to conserve predefined similarity lists in a low-dimensional projection subspace. Despite the performance enhancement, RMLLC using predefined similarity lists fails to capture the relative relevance information, which is often unavailable in practice. To address this problem, we further introduce a rectification term to automatically exploit the relative similarities, and develop an efficient alternating iterative algorithm to jointly learn the optimal metric and the rectification term. Extensive experiments on four publicly available benchmarking data sets are carried out and demonstrate that the proposed method is significantly superior to the state-of-the-art approaches. The results also show that the introduction of the rectification term could further boost the performance of RMLLC. Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Learning visual categories through a sparse representation classifier based cross-category knowledge transferabstractTo solve the challenging task of learning effective visual categories with limited training samples, we propose a new sparse representation classifier based transfer learning method, namely SparseTL, which propagates the cross-category knowledge from multiple source categories to the target category. Specifically, we enhance the target classification task in learning a both generative and discriminative sparse representation based classifier using pairs of source categories most positively and most negatively correlated to the target category. We further improve the discriminative ability of the classifier by choosing the most discriminative bins in the feature vector with a feature selection process. The experimental results show that the proposed method achieves competitive performance on the NUS-WIDE Scene database compared to several state of the art transfer learning algorithms while keeping a very efficient runtime. Ying Lu 0007, Liming Chen 0002, Alexandre Saidi, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIP | 4 |
| 2014 | Relevance Metric Learning for Person Re-identification by Exploiting Global SimilaritiesabstractPerson re-identification aims to match people across non-overlapping camera views, which is an important and challenging task. In order to obtain a robust metric for measuring (dis)similarities of (un)matched image pairs, metric learning has been introduced recently. Most existing works focus on seeking a Mahalanobis distance by employing sparse pair wise (dis)similarity constraints. However, the pair wise constraints have ignored a large portion of useful similarity information, and could not provide global similarity information. This paper proposes a novel metric learning method that could effectively exploit the global similarities. Specifically, we predefine lists of similarity scores, and measure (dis)similarities by the relevance of feature vectors. Subsequently, we learn a relevance metric by using the predefined list wise constraints, where the learnt metric is enforced to conserve predefined list wise similarities. Our main contributions lie on three folds: (1) we propose a metric learning method, which could effectively encode the global similarity information by using list wise constraints, (2) we formulate the relevance metric learning into a convex optimization problem, which could be solved efficiently, (3) we further kernelize the proposed method to support nonlinear mappings. The proposed method is experimentally validated on benchmark datasets, and outperforms state-of-the-art metric learning methods. Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICPR | 2 |
| 2014 | Enhanced Human Parsing with Multiple Feature Fusion and Augmented Pose ModelabstractWe address the problem of human pose estimation, which is a very challenging problem due to view angle variance, noise and occlusions. In this paper, we propose a novel human parsing method which can estimate diverse human poses from real world images. We merge the parallel lines feature and uniform LBP feature, thereby the new feature contains both shape and texture information, which can be used by discriminative body part detectors. The standard tree model is augmented by using virtual nodes in order to describe the correlations between originally unconnected nodes, which enhances the robustness of the traditional kinematic tree model. We test our method in a sports image dataset, and the experimental results demonstrate the advantages of the merged feature as well as the augmented pose model in real applications. Zhaoxiang Zhang 0001, Jianliang Hao, Yunhong Wang 0001 |
ICPR | 1 |
| 2014 | Object Classification in Traffic Scene Surveillance Based on Online Semi-supervised Active LearningabstractObject Classification in traffic scene surveillance has gained popularity in recent years. Traditional methods tend to utilize a large number of labeled training samples to achieve a satisfactory classification performance. However, labels of samples are not always available and manual labeling work is both time and labor consuming. To address the problem, a large number of semi-supervised learning based methods have been proposed, but most of them only focus on the offline settings. Motivated by an active learning framework, a novel online learning strategy is proposed in this paper. Furthermore, an intuitive semi-supervised learning method, which incorporates the spirits of both the online and active learning, is proposed and utilized in the scenario of traffic scene surveillance. The proposed learning framework is evaluated on the BUAA-IRIP traffic database, and the observed superior performance proves the effectiveness of our approach. Zhaoxiang Zhang 0001, Jie Qin 0004, Yunhong Wang 0001, Meng Liang |
ICPR | 1 |
| 2014 | Pan-sharpening based on weighted red black waveletsabstractPan‐sharpening is a technique which provides an efficient and economical solution to generate multi‐spectral (MS) images with high‐spatial resolution by fusing spectral information in MS images and spatial information in panchromatic (PAN) image. In this study, the authors propose a new pan‐sharpening method based on weighted red‐black (WRB) wavelets and adaptive principal component analysis (PCA), where the usage of WRB wavelet decomposition is to extract the spatial details in PAN image and the adaptive PCA is used to select the adequate principal component for injecting spatial details. WRB wavelets are data‐dependent second generation wavelets. Multi‐resolution analysis (MRA) based on WRB wavelet transform shows a better de‐correlation of the data compared with common linear translation‐invariant MRA, which makes it suitable for applications requiring manipulating image details. A local processing strategy is introduced to reduce the artefact effects and spectral distortions in the pan‐sharpened images. The proposed method is evaluated on the datasets acquired by QuickBird, IKONOS and Landsat‐7 ETM + satellites and compared with existing methods. Experimental results demonstrate that the authors method can provide promising fused MS images with high‐spatial resolution. Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IET Image Process. | 3 |
| 2014 | Face synthesis from low-resolution near-infrared to high-resolution visual light spectrum based on tensor analysis
Zhaoxiang Zhang 0001, Yunhong Wang 0001, Zeda Zhang |
Neurocomputing | 1 |
| 2014 | Secure multimodal biometric authentication with wavelet quantization based fingerprint watermarking
Bin Ma 0004, Yunhong Wang 0001, Chunlei Li 0004, Zhaoxiang Zhang 0001, Di Huang 0001 |
Multim. Tools Appl. | 4 |
| 2014 | Incremental learning patch-based bag of facial words representation for face recognition in videos
Chao Wang 0062, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2014 | On-line signature verification based on spatio-temporal correlation
Yunhong Wang 0001, Zhaoxiang Zhang 0001, Kaiyue Wang, Bin Ma 0004 |
Multim. Tools Appl. | 2 |
| 2013 | Face Tracking and Recognition via Incremental Local Sparse RepresentationabstractThis paper addresses the problem of tracking and recognizing faces via incremental local sparse representation. We first develop a robust face tracking algorithm based on the local sparse appearance. This sparse representation model exploits both partial and spatial information of the face based on a covariance pooling method. Following in the face recognition stage, with the employment of a novel template update strategy, our recognition algorithm adapts the template to appearance change and reduces the influence of occlusion and illumination variation. In the experiments, we test the quality of face recognition in real-world noisy videos on YouTube database. Our proposed method produces a high face recognition results on over 93% of all videos. The tracking results on challenging videos demonstrate that the proposed tracking algorithm performs favorably against several state-of-the-art methods. On the challenging data set in which faces are undergo occlusion and illumination variation, our proposed method also consistently demonstrates a high recognition rate. Chao Wang 0062, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIG | 3 |
| 2013 | Enhancing Person Re-identification by Robust Structural Metric LearningabstractPerson re-identification has become an important but also challenging task for video surveillance systems as it aims to match people across non-overlapping camera views. So far, most successful methods either focus on robust feature representation or sophisticated learners. Recently, metric learning has been applied in this task which aims to find a suitable feature subspace for matching samples from different cameras. However, most metric learning approaches rely on either pair wise or triplet-based distance comparison, which can be easily over-fitting in large scale and high dimension learning situation. Meanwhile, the performance of these methods can significantly decrease when the extracted features contain noisy information. In this paper, we propose a robust structural metric learning model for person re-identification with two main advantages: 1) it applies loss functions at the level of rankings rather than pair wise distances, 2) the proposed model is also robust to noisy information of the extracted features. The approach is verified on two available public datasets, and experimental results show that our method can get state-of-the-art performance. Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIG | 2 |
| 2013 | Semi-supervised learning in traffic scene surveillance based on label-propagationabstractObject classification in traffic scene surveillance has attracted much attention recent years. Traditional classification methods need lots of labeled samples to build a satisfying classifier. However, the acquisition of the labeled samples may cost lots of time and human labor. In this paper, we propose an label-propagation based semi-supervised learning method which uses the information of both labeled and un-labeled samples. Experiment results show that our method outperforms the traditional methods both in accuracy and robustness. Meng Liang, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIP | 2 |
| 2013 | Cross-view action recognition via transductive transfer learningabstractHuman action recognition is a hot topic in computer vision field. Various applicable approaches have been proposed to recognize different types of actions. However, the recognition performance deteriorates rapidly when the viewpoint changes. Traditional approaches aim to address the problem by inductive transfer learning, in which target-view samples are manually labeled. In this paper, we present a novel approach for cross-view action recognition based on transductive transfer learning. We address the problem by transferring instances across views. In our settings, both labels of examples from the target view and the corresponding relation between examples from pairwise views are dispensable. Experimental results on the IXMAS multi-view data set demonstrate the effectiveness of our approach, and are comparable to the state of the art. Jie Qin 0004, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIP | 2 |
| 2013 | Pixel-wise skin colour detection based on flexible neural treeabstractSkin colour detection plays an important role in image processing and computer vision. Selection of a suitable colour space is one key issue. The question that which colour space is most appropriate for pixel‐wise skin colour detection is not yet concluded. In this study, a pixel‐wise skin colour detection method is proposed based on the flexible neural tree (FNT) without considering the problem of selecting a suitable colour space. A FNT‐based skin model is constructed by using large skin data sets which identifies the important components of colour spaces automatically. Experimental results show improved accuracy and false positive rates (FPRs). The structure and parameters of FNT are optimised via genetic programming and particle swarm optimisation algorithms, respectively. In the experiments, nine FNT skin models are constructed and evaluated on features extracted from RGB, YCbCr, HSV and CIE‐Lab colour spaces. The Compaq and ECU datasets are used for constructing FNT‐based skin model and evaluating its performance compared with other skin detection methods. Without extra processing steps, the authors method achieves state of the art performance in skin pixel classification and better performance in terms of accuracy and FPRs. Tao Xu 0021, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IET Image Process. | 3 |
| 2013 | View independent object classification by exploring scene consistency information for traffic scene surveillance
Zhaoxiang Zhang 0001, Kaiqi Huang, Yunhong Wang 0001, Min Li 0022 |
Neurocomputing | 1 |
| 2013 | Cross-View Gait Recognition with Short Probe Sequences: from View Transformation Model to View-Independent stance-Independent Identity VectorabstractConsidering it is difficult to guarantee that at least one continuous complete gait cycle is captured in real applications, we address the multi-view gait recognition problem with short probe sequences. With unified multi-view population hidden markov models (umvpHMMs), the gait pattern is represented as fixed-length multi-view stances. By incorporating the multi-stance dynamics, the well-known view transformation model (VTM) is extended into a multi-linear projection model in a four-order tensor space, so that a view-independent stance-independent identity vector (VSIV) can be extracted. The main advantage is that the proposed VSIV is stable for each subject regardless of the camera location or the sequence length. Experiments show that our algorithm achieves encouraging performance for cross-view gait recognition even with short probe sequences. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | Multi-block dependency based fragile watermarking scheme for fingerprint images protection
Chunlei Li 0004, Yunhong Wang 0001, Bin Ma 0004, Zhaoxiang Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2013 | Estimation of view angles for gait using a robust regression method
Yunhong Wang 0001, Zhaoxiang Zhang 0001, Maodi Hu |
Multim. Tools Appl. | 3 |
| 2013 | Practical Camera Calibration From Moving Objects for Traffic Scene SurveillanceabstractWe address the problem of camera calibration for traffic scene surveillance, which supplies a connection between 2-D image features and 3-D measurement. It is helpful to deal with appearance distortion related to view angles, establish multiview correspondences, and make use of 3-D object models as prior information to enhance surveillance performance. A convenient and practical camera calibration method is proposed in this paper. With the camera heightHmeasured as the only user input, we can recover both intrinsic and extrinsic parameters of the camera based on redundant information supplied by moving objects in monocular videos. All cases of traffic scene layouts are considered and corresponding solutions are given to make our method applicable to almost all kinds of traffic scenes in reality. Numerous experiments are conducted in different scenes, and experimental results demonstrate the accuracy and practicability of our approach. It is shown that our approach can be effectively adopted in all kinds of traffic scene surveillance applications. Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Incremental Learning for Video-Based Gait Recognition With LBP FlowabstractGait analysis provides a feasible approach for identification in intelligent video surveillance. However, the effectiveness of the dominant silhouette-based approaches is overly dependent upon background subtraction. In this paper, we propose a novel incremental framework based on optical flow, including dynamics learning, pattern retrieval, and recognition. It can greatly improve the usability of gait traits in video surveillance applications. Local binary pattern (LBP) is employed to describe the texture information of optical flow. This representation is called LBP flow, which performs well as a static representation of gait movement. Dynamics within and among gait stances becomes the key consideration for multiframe detection and tracking, which is quite different from existing approaches. To simulate the natural way of knowledge acquisition, an individual hidden Markov model (HMM) representing the gait dynamics of a single subject incrementally evolves from a population model that reflects the average motion process of human gait. It is beneficial for both tracking and recognition and makes the training process of the HMM more robust to noise. Extensive experiments on widely adopted databases have been carried out to show that our proposed approach achieves excellent performance. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little |
IEEE Trans. Cybern. | 3 |
| 2013 | View-Invariant Discriminative Projection for Multi-View Gait-Based Human IdentificationabstractExisting methods for multi-view gait-based identification mainly focus on transforming the features of one view to the features of another view, which is technically sound but has limited practical utility. In this paper, we propose a view-invariant discriminative projection (ViDP) method, to improve the discriminative ability of multi-view gait features by a unitary linear projection. It is implemented by iteratively learning the low dimensional geometry and finding the optimal projection according to the geometry. By virtue of ViDP, the multi-view gait features can be directly matched without knowing or estimating the viewing angles. The ViDP feature projected from gait energy image achieves promising performance in the experiments of multi-view gait-based identification. We suggest that it is possible to construct a gait-based identification system for arbitrary probe views, by incorporating the information of gallery data with sufficient viewing angles. In addition, ViDP performs even better than the state-of-the-art view transformation methods, which are trained for the combination of gallery and probe viewing angles in every evaluation. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little, Di Huang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2013 | Transferring Training Instances for Convenient Cross-View Object Classification in SurveillanceabstractAutomatic object classification is an important issue in traffic scene surveillance. Appearance variation due to perspective distortion is one of the most difficult problems for moving object detection, tracking, and recognition. We propose an active transfer learning approach to bridge the gap between appearance variations under two different scenes. Only a small number of training samples are required in the target scene, which can be combined with transferred samples of the source scene to achieve a reliable object classifier in the target scene, and active learning strategy makes the algorithm more efficient. Abundant experiments are conducted and experimental results demonstrate the effectiveness and convenience of our approach. Zhaoxiang Zhang 0001, Yunhong Wang 0001, Jianyun Liu, Zhenjun Yao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Learning the Spherical Harmonic Features for 3-D Face RecognitionabstractIn this paper, a competitive method for 3-D face recognition (FR) using spherical harmonic features (SHF) is proposed. With this solution, 3-D face models are characterized by the energies contained in spherical harmonics with different frequencies, thereby enabling the capture of both gross shape and fine surface details of a 3-D facial surface. This is in clear contrast to most 3-D FR techniques which are either holistic or feature based, using local features extracted from distinctive points. First, 3-D face models are represented in a canonical representation, namely, spherical depth map, by which SHF can be calculated. Then, considering the predictive contribution of each SHF feature, especially in the presence of facial expression and occlusion, feature selection methods are used to improve the predictive performance and provide faster and more cost-effective predictors. Experiments have been carried out on three public 3-D face datasets, SHREC2007, FRGC v2.0, and Bosphorus, with increasing difficulties in terms of facial expression, pose, and occlusion, and which demonstrate the effectiveness of the proposed method. Peijiang Liu, Yunhong Wang 0001, Di Huang 0001, Zhaoxiang Zhang 0001, Liming Chen 0002 |
IEEE Trans. Image Process. | 4 |
| 2012 | Efficient Human Parsing Based on Sketch Representation
Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ACCV (1) | 2 |
| 2012 | Model-Based Multi-view Face Construction and Recognition in Videos
Chao Wang 0062, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIC (2) | 3 |
| 2012 | Cross-view object classification in traffic scene surveillance based on transductive transfer learningabstractObject classification in traffic scene surveillance has been a hot topic in image processing field. A big challenge is that shooting view changes in different scenes, which leads to sharp accuracy decrease since training and test samples do not share the same distribution. Inductive transfer learning methods try to bridge this gap by making use of manually labeled target samples. However, it is in line with reality to conduct unsupervised transfer without manually labeling. In this paper, we propose an intuitive transductive transfer method by transferring instances across view. Experimental results indicate that our method outperforms traditional approaches such as inductive SVM and cluster method, and could even achieve a comparable performance compared with manually labeling approach. Yi Mo, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIP | 2 |
| 2012 | Recognizing Occluded 3D Faces Using an Efficient ICP VariantabstractThis paper proposes an efficient variant of the Iterative Closest Point (ICP) algorithm for 3D face recognition in the presence of occlusion. The new ICP variant improves the original one in two aspects: the computational efficiency and the robustness to occlusion changes. For the former one, a facial surface is firstly described as a Spherical Depth Map (SDM), based on which uniform down-sampling can be conveniently applied to remove redundant vertices, aiming to decrease the consumed time of ICP. For the latter one, since occlusions can be considered as face outliers, a rejection strategy is embedded into ICP to eliminate their impacts. The proposed method is validated in face verification and identification scenarios on the Bosphorus database, and the experimental results clearly demonstrate its effectiveness and efficiency. Peijiang Liu, Yunhong Wang 0001, Di Huang 0001, Zhaoxiang Zhang 0001 |
ICME | 4 |
| 2012 | A Hybrid Transfer Learning Mechanism for Object Classification across ViewabstractObject classification in traffic scene is of vital importance to intelligent traffic surveillance. In real applications, the shooting view changes frequently in different scenes, which leads to sharp accuracy decrease since source and target domain samples do not follow the same distribution anymore. On the other hand, manual labeling training samples is time and labor consuming. Transfer learning approaches are to utilize the knowledge learnt from source view for target object classification. In this paper, we propose a hybrid transfer learning mechanism combining two single transfer approaches to gap the divergence of different domain distributions. An instance-based transfer approach is implemented to label target samples that represent target domain distribution best. And a feature-based transfer framework is to learn a strong classifier for target domain with both labeled source and target domain samples. Experimental results indicate that our approach outperforms traditional machine learning and single transfer learning methods. Yi Mo, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICMLA (1) | 2 |
| 2012 | Moving Object Detection in Aerial VideoabstractWe address the problem of moving object detection in aerial video. Moving object detection in aerial video is still a challenging problem for the reason that when capturing the video the camera (or the platform) is moving all the time. As a result, the problem is detecting moving object from moving background which is much more difficult than the case that the background is constant. To this end, a novel approach is proposed in this paper. Moving object detection in stationary scene usually modeling the pixel value changes over time, but in aerial video the change does not have regular patterns. Therefore, we model the motion of the background rather than modeling the background directly. The optical flow between every two adjacent frames is computed first to get the motion information for each pixel. Based on this, we define a notion named ``pixel motion process" which means the motion changes (the optical flow value changes) of a particular pixel over time, and transfer the Gaussian mixture model framework used for modeling background in the stationary scene to model the background motion. The result is an accurate, adaptive and general background motion model which is used to detect foreground moving objects. Experimental results demonstrate the effectiveness of our approach. Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICMLA (2) | 2 |
| 2012 | Locally linear embedding based example learning for pan-sharpening
Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 4 |
| 2012 | Pan-sharpening using weighted red-black wavelet
Qingjie Liu 0001, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 3 |
| 2012 | Enhancing biometric security with wavelet quantization watermarking based two-stage multimodal authentication
Bin Ma 0004, Chunlei Li 0004, Yunhong Wang 0001, Zhaoxiang Zhang 0001, Di Huang 0001 |
ICPR | 4 |
| 2012 | Enhancing cross-view object classification by feature-based transfer learning
Yi Mo, Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICPR | 2 |
| 2012 | Robust mobile spamming detection via graph patterns
Zhaoxiang Zhang 0001, Yunhong Wang 0001, Jianyun Liu |
ICPR | 2 |
| 2012 | Automatic object classification using motion blob based local feature fusion for traffic scene surveillance
Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
Frontiers Comput. Sci. | 1 |
| 2012 | Representing 3D Face from Point Cloud to Face-Aligned spherical Depth MapabstractWe propose a novel representation of 3D face shape which is a key step for feature extraction and face recognition. The input of the proposed methods is unstructured point cloud, which determines the wide applicability of the proposed representation. Our contributions mainly include two parts: Spherical Depth Map (SDM) and face alignment based on SDM. SDM, which can be adopted to many applications, is a special kind of range image utilizing the prior anatomical knowledge of human face. Useful characteristics of SDM facilitate face alignment with higher efficiency and accuracy. Experiments conducted on three popular 3D face databases verify the high efficacy and superiority of the proposed method. The accuracy of face alignment is up to 100% with our strategy. The face verification rates based on the standard protocols are all higher than the baseline performance of FRGC2.0. Peijiang Liu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | Spam Short Messages Detection via Mining Social Networks
Jianyun Liu, Zhaoxiang Zhang 0001, Yunhong Wang 0001, Xue-Mei Yuan, Zhenjiang Dong |
J. Comput. Sci. Technol. | 3 |
| 2012 | Three-Dimensional Deformable-Model-Based Localization and Recognition of Road VehiclesabstractWe address the problem of model-based object recognition. Our aim is to localize and recognize road vehicles from monocular images or videos in calibrated traffic scenes. A 3-D deformable vehicle model with 12 shape parameters is set up as prior information, and its pose is determined by three parameters, which are its position on the ground plane and its orientation about the vertical axis under ground-plane constraints. An efficient local gradient-based method is proposed to evaluate the fitness between the projection of the vehicle model and image data, which is combined into a novel evolutionary computing framework to estimate the 12 shape parameters and three pose parameters by iterative evolution. The recovery of pose parameters achieves vehicle localization, whereas the shape parameters are used for vehicle recognition. Numerous experiments are conducted in this paper to demonstrate the performance of our approach. It is shown that the local gradient-based method can evaluate accurately and efficiently the fitness between the projection of the vehicle model and the image data. The evolutionary computing framework is effective for vehicles of different types and poses is robust to all kinds of occlusion. Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Combining Tensor Space Analysis and Active Appearance Models for Aging Effect Simulation on Face ImagesabstractApplications of the simulation of adult aging effects are widespread nowadays, whereas the difficulties in certain aspects restrict its development. In this paper, a method is proposed for simulating adult facial aging effects by means of super-resolution. Accounting for the nature of multimodalities in the face image set, multilinear algebra is introduced to represent and process the whole image set in tensor space. To ameliorate the aging simulation results generated by merely the super-resolution method, we further adopt active appearance models to reduce the blurring effects of the results through adding normalization of the faces and postprocessing to the algorithm. To evaluate our aging simulation method, the aged faces obtained are compared with the ground-truth face images of the same individuals and also assessed by several volunteers mainly from two perspectives: the aged faces' perceived age and their preservation effects of the original identities of subjects in the test images. Additionally, objective experiments based on an automatic age estimator and a face recognition method using eigenfaces are also conducted as another way of the evaluation. Yunhong Wang 0001, Zhaoxiang Zhang 0001, Weixin Li 0001, Fangyuan Jiang |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2011 | On-line signature verification using wavelet packetabstractIn this paper, we propose a novel approach for on-line signature verification using wavelet packet. Signatures are first normalized and resampled, thus they have the same number of sample points. Then, several types of local features are extracted, so that wavelet transform can be applied on them. After that, we conduct experiments to select the best local features, wavelet bases and wavelet packet settings. Also, experiments are carried out to verify the reliability and efficiency of our approach, which performs better than discrete wavelet transform and competes with the state-of-arts. Kaiyue Wang, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IJCB | 3 |
| 2011 | Face synthesis from near-infrared to visual light via sparse representationabstractThis paper presents a novel method for synthesizing artificial visual light (VIS) face images from near-infrared (NIR) inputs. Active NIR imaging is now widely employed because it is unobtrusive, invariant of environmental illuminations, and can penetrate glasses and sweats. Unfortunately, NIR imaging exhibits discrepant photic properties compared with VIS imaging. Based on recent results of research on compressive sensing, natural images can be compressed and recovered with an overcomplete dictionary by sparse representation coefficients. In our approach a pairwise dictionary is trained from randomly sampled coupled face patches, which contains sparse coded base functions to reconstruct representation coefficients via l1-minimization. We will demonstrate that this method is robust to moderate pose and expression variations, and is efficient in computing. Comparative experiments are conducted with state-of-the-art algorithms. Zeda Zhang, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IJCB | 3 |
| 2011 | On-line Signature Verification Using Segment-to-Segment Graph MatchingabstractThis paper proposes a novel approach of on-line signature verification. Firstly, on-line signatures are partitioned into a series of segments, which are then represented by graphs. Four segmentation methods are taken into account. Secondly, graph matching techniques are adopted to compute edit distance between corresponding graphs, which measures the similarity of them. Finally, having been able to compare two signatures, limited genuine signatures are used to train user dependent classifiers for each user. Experiments are conducted to validate the effectiveness of the proposed method and promising results are achieved. Kaiyue Wang, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICDAR | 3 |
| 2011 | Visual Saliency Based Aerial Video Summarization by Online Scene ClassificationabstractCompared with traditional video summarization approaches, aerial video summarization is a new and challenging issue for its particular characteristics. Aerial video data is a massive data stream, without pre-edit structures such as sports or news video data, lack of camera motion such as zoom and pan. On account of these characteristics, we proposed a novel approach for summarization. First, we extract GIST features for each frame as the holistic scene representation. Then, we divide aerial video into temporal segments representing a visual scene using on-line clustering method by examine GIST features of each frame only once. Finally, we select several key frames from each scene for summarization according to visual saliency index (VSI) of each frame computed from their visual saliency map. In the paper, we proposed new criterion for estimation of temporal segmentation of streaming video. Experimental observations show the success of our approach on aerial video summarization. Jiewei Wang, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIG | 3 |
| 2011 | On-line Signature Verification Using Graph RepresentationabstractThis paper proposes a novel approach of on-line signature verification. Firstly, on-line signatures are represented by a series of graphs, whose nodes and edges describe certain properties at sample points and relationship between points respectively. Then, graph matching techniques are introduced to compute edit distance between graphs, which measures the similarity of graphs. Finally, having been able to compare any two signatures through the last two steps, user-dependent classifiers are trained using limited genuine signatures. The proposed method is tested on SUSIG online signature database and shows promising performance. Kaiyue Wang, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIG | 3 |
| 2011 | Codebook Reconstruction with Word Correlation Feedback MechanismabstractBag of feature model has been shown to be one of the most successful methods in generic image categorization problems. However, creating codebook by clustering local feature vectors (e.g. Kmeans) may lose holistic information of images. This paper presents a novel process called Correlation Feedback for codebook construction. It introduces semantic similarities of words by measuring correlation between distributions of them within one image. Further more, we employ label propagation process to spread the affinities among all features. An enhanced codebook is constructed based on fusion of the new similarity matrix with spectral clustering. Experimental results on Caltech101 and the 15 nature scenes datasets shows promising performance of importing the novel similarity to dictionary construction. Zhaoxiang Zhang 0001, Yunhong Wang 0001 |
ICIG | 2 |
| 2011 | Multi-view multi-stance gait identificationabstractView transformation in gait analysis has attracted more and more attentions recently. However, most of the existing methods are based on the entire gait dynamics, such as Gait Energy Image (GEI). And the distinctive characteristics of different walking phases are neglected. This paper proposes a multi-view multi-stance gait identification method using unified multi-view population Hidden Markov Models (pHMM-s), in which all the models share the same transition probabilities. Hence, the gait dynamics in each view can be normalized into fixed-length stances by Viterbi decoding. To optimize the view-independent and stance-independent identity vector, a multi-linear projection model is learned from tensor decomposition. The advantage of using tensor is that different types of information are integrated in the final optimal solution. Extensive experiments show that our algorithm achieves promising performances of multi-view gait identification even with incomplete gait cycles. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIP | 3 |
| 2011 | Face synthesis from near-infrared to visual light spectrum using quotient image and kernel-based multifactor analysisabstractThis paper addresses the problem of synthesizing an artificial visual light (VIS) facial image from near-infrared (NIR) input. After extensively assessing photic characteristics of tissues at human skin surface, we propose a framework for this task. Firstly, we take the quotient images for training and reconstruction, so that information related to face structure can be preserved. Secondly, to handle heterogeneous blur resulted from multiple scattering within tissues, we introduce kernel based strategy as a powerful nonlinear analyzing instrument. Finally, as in our application the image ensembles involve multiple factors, a tensor structure is employed to transform heterogeneous face data into uniform subspaces. Comparative results show that our synthesized images are both suited for human vision and discriminative for machine recognition. Zeda Zhang, Yunhong Wang 0001, Zhaoxiang Zhang 0001, Guangpeng Zhang |
ICME | 3 |
| 2011 | Gait-Based Gender Classification Using Mixed Conditional Random FieldabstractThis paper proposes a supervised modeling approach for gait-based gender classification. Different from traditional temporal modeling methods, male and female gait traits are competitively learned by the addition of gender labels. Shape appearance and temporal dynamics of both genders are integrated into a sequential model called mixed conditional random field (CRF) (MCRF), which provides an open framework applicable to various spatiotemporal features. In this paper, for the spatial part, pyramids of fitting coefficients are used to generate the gait shape descriptors; for the temporal part, neighborhood-preserving embeddings are clustered to allocate the stance indexes over gait cycles. During these processes, we employ evaluation functions like the partition index and Xie and Beni's index to improve the feature sparseness. By fusion of shape descriptors and stance indexes, the MCRF is constructed in coordination with intra- and intergender temporary Markov properties. Analogous to the maximum likelihood decision used in hidden Markov models (HMMs), several classification strategies on the MCRF are discussed. We use CASIA (Data set B) and IRIP Gait Databases for the experiments. The results show the superior performance of the MCRF over HMMs and separately trained CRFs. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2010 | Automatic and robust 3D face registration using multiresolution Spherical Depth MapabstractFace registration is a necessary preprocessing step for 3D face recognition. An entirely automatic method for 3D face registration is proposed in this paper with high accuracy and good robustness to pose and facial expression variations. Our method consists of the following three stages. Firstly, the face shape is represented by Fitting Sphere Representation (FSR). Secondly, generate the Spherical Depth Map (SDM) of face shape which is normalized roughly to a similar pose. Finally, accurately localize the nose tip using multiresolution SDM by a coarse-to-fine method. Then, in conjunction with the face orientation figured out by linear fitting and the center of fitting sphere, face can be registered completely. Extensive experiments are conducted on five popular 3D face databases. The registration accuracy is near 100 percent. Experimental results demonstrate the high robustness of the proposed methods to pose and expression variations. Peijiang Liu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICIP | 3 |
| 2010 | Combining Spatial and Temporal Information for Gait Based Gender ClassificationabstractIn this paper, we address the problem of gait based gender classification. The Gabor feature which is a new attempt for gait analysis, not only improves the robustness to the segmental noise, but also provides a feasible way to purge the additional influence factors like clothing and carrying condition changes before supervised learning. Furthermore, through the agency of Maximization of Mutual Information (MMI), the low dimensional discriminative representation is obtained as the Gabor-MMI feature. After that, gender related Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) are constructed for classification work. In this case, supervised learning reduces the dimension of parameter space, and significantly increases the gap between likelihoods of the gender models. In order to assess the performance of our proposed approach, we compare it with other methods on the standard CASIA Gait Databases (Dataset B). Experimental results demonstrate that our approach achieves better Correct Classification Rate (CCR) than the state of the art methods. Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 3 |
| 2010 | Block Pyramid Based Adaptive Quantization Watermarking for Multimodal Biometric AuthenticationabstractThis paper proposes a novel robust watermarking scheme to embed fingerprint minutiae into face images for multimodal biometric authentication. First, a block pyramid is layered according to the block-wise face region distinctiveness estimated by Adaboost; upper level indicates informative spacial regions. Then, we adopt a first-order statics QIM method to perform watermark embedding in each pyramid level. Numeric watermark bits with higher priority are embedded into upper pyramid level with a larger embedding strength. By joint differentiation of host image regions and watermark bits priority, our scheme achieves a trade-offs among watermarking robustness, capacity and fidelity. Experimental results demonstrate that our approach guarantees the robustness of hidden biometric data, while preserving the distinctiveness of host biometric images. Bin Ma 0004, Chunlei Li 0004, Yunhong Wang 0001, Zhaoxiang Zhang 0001 |
ICPR | 4 |
| 2010 | 3D Model Based Vehicle Tracking Using Gradient Based Fitness Evaluation under Particle Filter FrameworkabstractWe address the problem of 3D model based vehicle tracking from monocular videos of calibrated traffic scenes. A 3D wire-frame model is set up as prior information and an efficient fitness evaluation method based on image gradients is introduced to estimate the fitness score between the projection of vehicle model and image data, which is then combined into a particle filter based framework for robust vehicle tracking. Numerous experiments are conducted and experimental results demonstrate the effectiveness of our approach for accurate vehicle tracking and robustness to noise and occlusions. Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan, Yunhong Wang 0001 |
ICPR | 1 |
| 2009 | Rapid and robust human detection and tracking based on omega-shape featuresabstractThis paper proposes a novel method for rapid and robust human detection and tracking based on the omega-shape features of people's head-shoulder parts. There are two modules in this method. In the first module, a Viola-Jones type classifier and a local HOG (Histograms of Oriented Gradients) feature based AdaBoost classifier are combined to detect head-shoulders rapidly and effectively. Then, in the second module, each detected head-shoulder is tracked by a particle filter tracker using local HOG features to model target's appearance, which shows great robustness in scenarios of crowding, background distractors and partial occlusions. Experimental results demonstrate the effectiveness and efficiency of the proposed approach. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICIP | 2 |
| 2009 | Robust visual tracking based on simplified biologically inspired featuresabstractWe address the problem of robust appearance-based visual tracking. First, a set of simplified biologically inspired features (SBIF) is proposed for object representation and the Bhattacharyya coefficient is used to measure the similarity between the target model and candidate targets. Then, the proposed appearance model is combined into a Bayesian state inference tracking framework utilizing the SIR (sampling importance resampling) particle filter to propagate sample distributions over time. Numerous experiments are conducted and experimental results demonstrate that our algorithm is robust to partial occlusions and variations of illumination and pose, resistent to nearby distractors, as well as possesses the state-of-the-art tracking accuracy. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICIP | 2 |
| 2008 | Practical camera auto-calibration based on object appearance and motion for traffic scene visual surveillanceabstractCamera calibration, as a fundamental issue in computer vision, is indispensable in many visual surveillance applications. Firstly, calibrated camera can help to deal with perspective distortion of object appearance on image plane. Secondly, calibrated camera makes it possible to recover metrics from images which are robust to scene or view an gle changes. In addition, with calibrated cameras, we can make use of prior information of 3D models to estimate 3D pose of objects and make object detection or tracking more robust to noise and occlusions. In this paper, we propose an automatic method to recover camera models from traffic scene surveillance videos. With only the camera height H measured, we can completely recover both intrinsic and extrinsic parameters of cameras based on appearance and motion of objects in videos. Experiments are conducted in different scenes and experimental results demonstrate the effectiveness and practicability of our approach, which can be adopted in many traffic scene surveillance applications. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
CVPR | 1 |
| 2008 | Robust automated ground plane rectification based on moving vehicles for traffic scene surveillanceabstractMost outdoor visual surveillance scenes involve objects of interest moving on the ground plane. However, perspective distortion introduces many difficulties to various applications like object classification and activity recognition. In this paper, we propose a robust automated method for both affine and metric rectification of the ground plane based on appearance and motion of vehicles in traffic scene surveillance videos. This rectification enables normalization of object properties like size, length and velocity. Various useful applications are presented and experimental results demonstrate the effectiveness and robustness of the proposed method. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICIP | 1 |
| 2008 | Estimating the number of people in crowded scenes by MID based foreground segmentation and head-shoulder detectionabstractThis paper proposes a novel method to address the problem of estimating the number of people in surveillance scenes with people gathering and waiting. The proposed method combines a MID (mosaic image difference) based foreground segmentation algorithm and a HOG (histograms of oriented gradients) based head-shoulder detection algorithm to provide an accurate estimation of people counts in the observed area. In our framework, the MID-based foreground segmentation module provides active areas for the head-shoulder detection module to detect heads and count the number of people. Numerous experiments are conducted and convincing results demonstrate the effectiveness of our method. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICPR | 2 |
| 2008 | Boosting local feature descriptors for automatic objects classification in traffic scene surveillanceabstractWe address the problem of automatic object classification for traffic scene surveillance, which is very challenging for the low resolution videos, large intra-class variations and real-time requirement. In this paper, we propose a new strategy for object classification by boosting different local feature descriptors in motion blobs. We not only evaluate the performance of each local feature descriptor, but also fuse these descriptors to achieve better performance. Numerous experiments are conducted and experimental results demonstrate the effectiveness and efficiency of our approach with robustness to noise and variance of view angles, lighting conditions and environments. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICPR | 1 |
| 2008 | 3D model based vehicle localization by optimizing local gradient based fitness evaluationabstractWe address the problem of 3D model based vehicle localization in calibrated traffic scenes. A wire-frame vehicle model is set up as prior information and an efficient local gradient based method is proposed to evaluate the fitness between the projection of 3D model and image data, which illustrates smooth optimization surface and more conspicuous peak with low computational cost. Gradient decent is then applied to optimize the evaluation score for localization. Experimental results demonstrate the accuracy, efficiency and robustness of the proposed method for model based vehicle localization. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICPR | 1 |
| 2007 | EDA Approach for Model Based Localization and Recognition of VehiclesabstractWe address the problem of model based recognition. Our aim is to localize and recognize road vehicles from monocular images in calibrated scenes. A deformable 3D geometric vehicle model with 12 parameters is set up as prior information and Bayesian Classification Error is adopted for evaluation of fitness between the model and images. Using a novel evolutionary computing method called EDA (Estimation of Distribution Algorithm), we can not only determine the 3D pose of the vehicle, but also obtain a 12 dimensional vector which corresponds to the 12 shape parameters of the model. By clustering obtained vectors in the parameter space, we can recognize different types of vehicles. Experimental results demonstrate the effectiveness of the approach to vehicles of different types and poses. Thanks to EDA, we can not only localize and recognize vehicles, but also show the whole evolution procedure of the deformable model which gradually fits the image better and better. Zhaoxiang Zhang 0001, Weishan Dong, Kaiqi Huang, Tieniu Tan |
CVPR | 1 |
| 2007 | Real-Time Moving Object Classification with Automatic Scene DivisionabstractWe address the problem of moving object classification. Our aim is to classify moving objects of traffic scene videos into pedestrians, bicycles and vehicles. Instead of supervised learning and manual labeling of large training samples, our classifiers are initialized and refined online automatically. With efficient features extracted and organized, the approach can be real-time and achieve high classification accuracy. Once the view or scene changes detected, the algorithm can automatically refine the classifiers and adapt them to new environments. Experimental results demonstrate the effectiveness and robustness of the proposed approach. Zhaoxiang Zhang 0001, Yinghao Cai, Kaiqi Huang, Tieniu Tan |
ICIP (5) | 1 |