VLDB 2026 Research / reviewers in the wild / expert
Jianke Zhu
dblp:10/4016
· DBLP profile ↗
124ranked-venue papers
15as first author
56since 2021 · last 2026
0000-0003-1831-0106ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 10 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 12 first-author · 35 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BCA-IML: Bidirectional cross-attention guided multi-scale feature fusion for image manipulation localization
Yulin Cheng, Wenyu Liu 0005, Puning Zhao, Jianke Zhu |
Expert Syst. Appl. | 6 |
| 2026 | ChatTracker: Enhancing Visual Tracking via LLM-Driven Iterative Description RefinementabstractVisual object tracking focuses on locating a target object within a video sequence based on an initial bounding box. Recently, Vision-Language (VL) trackers have been proposed to utilize additional natural language descriptions to enhance versatility in various applications. Despite this potential, VL trackers still underperform the State-of-the-Art (SoTA) visual trackers in terms of tracking accuracy. We find that this inferiority is primarily due to their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we identify, for the first time, that over 10% of textual annotations in existing VL tracking datasets suffer from inaccuracies through manual evaluation. To address this problem, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel Reflection-based Language Description Refinement Module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed, which can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that ChatTracker achieves comparable performance to existing SoTA tracking methods. In addition, language descriptions generated by ChatTracker enhance the performance of various VL trackers and exhibit better text-to-image alignment than annotations in the original dataset. Moreover, our proposed framework can improve the performance of various visual tasks, including Referring Expression Comprehension (REC), Referring Expression Segmentation (RES), and Referring Video Object Segmentation (R-VOS) tasks by providing more accurate language descriptions, which demonstrates the universality of ChatTracker. We release the manual evaluation results and the generated textual descriptions, aiming to drive advancements in VL tracking. Yiming Sun 0006, Mi Zhang 0001, Shaoxiang Chen 0001, Yang Li 0041, Changbo Wang, Jianke Zhu, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | XHand: Real-Time Expressive Hand AvatarabstractHand avatars play a pivotal role in a wide array of digital interfaces. Fine-detailed and realistic hand representations enhance user immersion and facilitating natural interaction within virtual environments. While previous studies have focused on photo-realistic hand rendering, little attention has been paid to reconstruct the hand geometry with fine details, which is essential to rendering quality. In the realms of extended reality and gaming, on-the-fly rendering becomes imperative. To this end, we introduce an expressive hand avatar, named XHand, that is designed to comprehensively generate hand shape, appearance, and deformations in real-time. To obtain fine-grained hand meshes, we make use of three feature embedding modules to predict hand deformation displacements, albedo, and linear blending skinning weights, respectively. To achieve photo-realistic hand rendering on fine-grained meshes, our method employs a mesh-based neural renderer by leveraging mesh topological consistency and latent codes from embedding modules. During training, a part-aware Laplace smoothing strategy is proposed by incorporating the distinct levels of regularization to effectively maintain the necessary details and eliminate the undesired artifacts. The experimental evaluations on InterHand2.6M and DeepHandMesh datasets demonstrate the efficacy of XHand, which is able to recover high-fidelity geometry and texture for hand animations across diverse poses in real-time. To reproduce our results, we will make the full implementation publicly available at https://github.com/agnJason/XHand. Qijun Gan, Jianke Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | ScaleOT: Privacy-utility-scalable Offsite-tuning with Dynamic LayerReplace and Selective Rank CompressionabstractOffsite-tuning is a privacy-preserving method for tuning large language models (LLMs) by sharing a lossy compressed emulator from the LLM owners with data owners for downstream task tuning. This approach protects the privacy of both the model and data owners. However, current offsite tuning methods often suffer from adaptation degradation, high computational costs, and limited protection strength due to uniformly dropping LLM layers or relying on expensive knowledge distillation. To address these issues, we propose ScaleOT, a novel privacy-utility-scalable offsite-tuning framework that effectively balances privacy and utility. ScaleOT introduces a novel layerwise lossy compression algorithm that uses reinforcement learning to obtain the importance of each layer. It employs lightweight networks, termed harmonizers, to replace the raw LLM layers. By combining important original LLM layers and harmonizers in different ratios, ScaleOT generates emulators tailored for optimal performance with various model scales for enhanced privacy protection. Additionally, we present a rank reduction method to further compress the original LLM layers, significantly enhancing privacy with negligible impact on utility. Comprehensive experiments show that ScaleOT can achieve nearly lossless offsite tuning performance compared with full fine-tuning while obtaining better model privacy. Zhaorui Tan, Tiandi Ye, Lichun Li, Yuan Zhao 0015, Wenyan Liu 0001, Wei Wang 0002, Jianke Zhu |
AAAI | 8 |
| 2025 | GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language ModelsabstractKai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, Wei Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao 0015, Yixin Ji, Jianke Zhu, Wei Wang 0002 |
ACL (1) | 9 |
| 2025 | Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionabstractReliable high-definition (HD) map construction is crucial for the driving safety of autonomous vehicles. Although recent studies demonstrate improved performance, their generalization capability across unfamiliar driving scenes remains unexplored. To tackle this issue, we propose UIGenMap, an uncertainty-instructed structure injection approach for generalizable HD map vectorization, which concerns the uncertainty resampling in statistical distribution and employs explicit instance features to reduce excessive reliance on training data. Specifically, we introduce the perspective-view (PV) detection branch to obtain explicit structural features, in which the uncertainty-aware decoder is designed to dynamically sample probability distributions considering the difference in scenes. With probabilistic embedding and selection, UI2DPrompt is proposed to construct PV-learnable prompts. These PV prompts are integrated into the map decoder by designed hybrid injection to compensate for neglected instance structures. To ensure real-time inference, a lightweight Mimic Query Distillation is designed to learn from PV prompts, which can serve as an efficient alternative to the flow of PV branches. Extensive experiments on challenging geographically disjoint (geo-based) data splits demonstrate that our UIGen-Map achieves superior performance, with +5.7 mAP improvement on the nuScenes dataset. Source code is available at https://github.com/xiaolul2/UIGenMap. Ruizi Yang, Song Wang 0019, Wentong Li 0001, Junbo Chen, Jianke Zhu |
CVPR | 6 |
| 2025 | PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud LearningabstractSelf-supervised representation learning for point cloud has demonstrated effectiveness in improving pre-trained model performance across diverse tasks. However, as pre-trained models grow in complexity, fully fine-tuning them for downstream applications demands substantial computational and storage resources. Parameter-efficient fine-tuning (PEFT) methods offer a promising solution to mitigate these resource requirements, yet most current approaches rely on complex adapter and prompt mechanisms that increase tunable parameters. In this paper, we propose PointLoRA, a simple yet effective method that combines low-rank adaptation (LoRA) with multi-scale token selection to efficiently fine-tune point cloud models. Our approach embeds LoRA layers within the most parameter-intensive components of point cloud transformers, reducing the need for tunable parameters while enhancing global feature capture. Additionally, multi-scale token selection extracts critical local information to serve as prompts for downstream fine-tuning, effectively complementing the global context captured by LoRA. The experimental results across various pre-trained models and three challenging public datasets demonstrate that our approach achieves competitive performance with only 3.43% of the trainable parameters, making it highly effective for resource-constrained applications. Source code is available at: https://github.com/songw-zju/PointLoRA. Song Wang 0019, Lingdong Kong, Jianyun Xu, Chunyong Hu, Gongfan Fang, Wentong Li 0001, Jianke Zhu, Xinchao Wang |
CVPR | 8 |
| 2025 | Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction TuningabstractDespite encouraging progress in 3D scene understanding, it remains challenging to develop an effective Large Multimodal Model (LMM) that is capable of understanding and reasoning in complex 3D environments. Most previous methods typically encode 3D point and 2D image features separately, neglecting interactions between 2D semantics and 3D object properties, as well as the spatial relationships within the 3D environment. This limitation not only hinders comprehensive representations of 3D scene, but also compromises training and inference efficiency. To address these challenges, we propose a unified Instance-aware 3DLarge Multi-modal Model (Inst3D-LMM) to deal with multiple 3D scene understanding tasks simultaneously. To obtain the fine-grained instance-level visual tokens, we first introduce a novel Multi-view Cross-Modal Fusion (MCMF) module to inject the multi-view 2D semantics into their corresponding 3D geometric features. For scene-level relation-aware tokens, we further present a 3D Instance Spatial Relation (3D-ISR) module to capture the intricate pairwise spatial relationships among objects. Additionally, we perform end-to-end multi-task instruction tuning simultaneously without the subsequent task-specific fine-tuning. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods across 3D scene understanding, reasoning and grounding tasks. Source code is available at: https://github.com/hanxunyu/Inst3D-LMM. Hanxun Yu, Wentong Li 0001, Song Wang 0019, Junbo Chen, Jianke Zhu |
CVPR | 5 |
| 2025 | VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMabstractVideo Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal details. Besides, the lack of high-quality object-level video instruction data and a comprehensive benchmark further hinders their advancements. To tackle these challenges, we introduce the VideoRefer Suite to empower Video LLM for finer-level spatial-temporal video understanding, i.e., enabling perception and reasoning on any objects throughout the video. Specially, we thoroughly develop VideoRefer Suite across three essential aspects: dataset, model, and benchmark. Firstly, we introduce a multi-agent data engine to meticulously curate a largescale, high-quality object-level video instruction dataset, termed VideoRefer-700K. Next, we present the VideoRefer model, which equips a versatile spatial-temporal object encoder to capture precise regional and sequential representations. Finally, we meticulously create a VideoRefer-Bench to comprehensively assess the spatial-temporal understanding capability of a Video LLM, evaluating it across various aspects. Extensive experiments and analyses demonstrate that our VideoRefer model not only achieves promising performance on video referring benchmarks but also facilitates general video understanding capabilities. Yuqian Yuan, Wentong Li 0001, Zesen Cheng, Boqiang Zhang, Xin Li 0056, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, Lidong Bing |
CVPR | 11 |
| 2025 | SAM4D: Segment Anything in Camera and LiDAR StreamsabstractWe present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared 3D space, enabling seamless cross-modal prompting and interaction. Additionally, we propose Motion-aware Cross-modal Memory Attention (MCMA), which leverages ego-motion compensation to enhance temporal consistency and long-horizon feature retrieval, ensuring robust segmentation across dynamically changing autonomous driving scenes. To avoid annotation bottlenecks, we develop a multi-modal automated data engine that synergizes VFM-driven video masklets, spatiotemporal 4D reconstruction, and cross-modal masklet fusion. This framework generates camera-LiDAR aligned pseudo-labels at a speed orders of magnitude faster than human annotation while preserving VFM-derived semantic fidelity in point cloud representations. We conduct extensive experiments on the constructed Waymo-4DSeg, which demonstrate the powerful cross-modal segmentation ability and great potential in data annotation of proposed SAM4D. Jianyun Xu, Song Wang 0019, Ziqian Ni, Chunyong Hu, Sheng Yang 0007, Jianke Zhu |
ICCV | 6 |
| 2025 | LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video DiffusionabstractVideo Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or temporal movement to driven VDMs with constrained data, achieving precise control over both camera trajectories and object motion remains challenging due to the unstable fusion and non-linear scalability. To address these issues, we propose LiON-LoRA, a novel framework that rethinks LoRA fusion through three core principles: Linear scalability, Orthogonality, and Norm consistency. First, we analyze the orthogonality of LoRA features in shallow VDM layers, enabling decoupled low-level controllability. Second, norm consistency is enforced across layers to stabilize fusion during complex camera motion combinations. Third, a controllable token is integrated into the diffusion transformer (DiT) to linearly adjust motion amplitudes for both cameras and objects with a modified self-attention mechanism to ensure decoupled control. Additionally, we extend LiON-LoRA to temporal generation by leveraging static-camera videos, unifying spatial and temporal controllability. Experiments demonstrate that LiON-LoRA outperforms state-of-the-art methods in trajectory control accuracy and motion strength adjustment, achieving superior generalization with minimal training data. Project Page: https://fuchengsu.github.io/lionlora.github.io/ Yisu Zhang, Chenjie Cao, Chaohui Yu, Jianke Zhu |
ICCV | 4 |
| 2025 | PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano PerformanceabstractRecently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes. The source code and dataset can be accessed at https://github.com/agnJason/PianoMotion10M. Qijun Gan, Song Wang 0019, Shengtao Wu, Jianke Zhu |
ICLR | 4 |
| 2025 | Reliable and Calibrated Semantic Occupancy Prediction by Hybrid Uncertainty LearningabstractVision-centric semantic occupancy prediction plays a crucial role in autonomous driving, which requires accurate and reliable predictions from low-cost sensors. Although having notably narrowed the accuracy gap with LiDAR, there is still few research effort to explore the reliability and calibration in predicting semantic occupancy from camera. In this paper, we conduct a comprehensive evaluation of existing semantic occupancy prediction models from a reliability perspective for the first time. Despite the gradual alignment of camera-based models with LiDAR in terms of accuracy, a significant reliability gap still persists. To address this concern, we propose ReliOcc, a method designed to enhance the reliability of camera-based occupancy networks. ReliOcc provides a plug-and-play scheme for existing models, which integrates hybrid uncertainty from individual voxels with sampling-based noise and relative voxels through mix-up learning. Besides, an uncertainty-aware calibration strategy is devised to further improve model reliability in offline mode. Extensive experiments under various settings demonstrate that ReliOcc significantly enhances the reliability of learned model while maintaining the accuracy for both geometric and semantic predictions. Notably, our proposed approach exhibits robustness to sensor failures and out of domain noises during inference. Song Wang 0019, Zhongdao Wang, Wentong Li 0001, Bailan Feng, Junbo Chen, Jianke Zhu |
IJCAI | 7 |
| 2025 | A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy GroundingabstractVisual grounding aims at identifying objects or regions in a scene based on natural language descriptions, which is essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically depend on bounding boxes that often fail to capture fine-grained details. Not all voxels within a bounding box are occupied, resulting in inaccurate object representations. To address this, we introduce a benchmark for 3D occupancy grounding in challenging outdoor scenes. Built on the nuScenes dataset, it fuses natural language with voxel-level occupancy annotations, offering more precise object perception compared to the traditional grounding task. Moreover, we propose GroundingOcc, an end-to-end model designed for 3D occupancy grounding through multimodal learning. It combines visual, textual, and point cloud features to predict object location and occupancy information from coarse to fine. Specifically, GroundingOcc comprises a multimodal encoder for feature extraction, an occupancy head for voxel-wise predictions, and a grounding head for refining localization. Additionally, a 2D grounding module and a depth estimation module enhance geometric understanding, thereby boosting model performance. Extensive experiments on the benchmark demonstrate that our method outperforms existing baselines on 3D occupancy grounding. The dataset is available at https://github.com/RONINGOD/GroundingOcc. Song Wang 0019, Junbo Chen, Jianke Zhu |
IROS | 4 |
| 2025 | MambaMap: Online Vectorized HD Map Construction using State Space ModelabstractHigh-definition (HD) maps are essential for autonomous driving, as they provide precise road information for downstream tasks. Recent advances highlight the potential of temporal modeling in addressing challenges like occlusions and extended perception range. However, existing methods either fail to fully exploit temporal information or incur substantial computational overhead in handling extended sequences. To tackle these challenges, we propose MambaMap, a novel framework that efficiently fuses long-range temporal features in the state space to construct online vectorized HD maps. Specifically, MambaMap incorporates a memory bank to store and utilize information from historical frames, dynamically updating BEV features and instance queries to improve robustness against noise and occlusions. Moreover, we introduce a gating mechanism in the state space, selectively integrating dependencies of map elements in high computational efficiency. In addition, we design innovative multi-directional and spatial-temporal scanning strategies to enhance feature extraction at both BEV and instance levels. These strategies significantly boost the prediction accuracy of our approach while ensuring robust temporal consistency. Extensive experiments on the nuScenes and Argoverse2 datasets demonstrate that our proposed MambaMap approach outperforms state-of-the-art methods across various splits and perception ranges. Source code will be available at https://github.com/ZiziAmy/MambaMap. Ruizi Yang, Junbo Chen, Jianke Zhu |
IROS | 4 |
| 2025 | DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language ModelsabstractJianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Pei Wang, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Hui Huang 0021, Chenchen Jing, Xingwei Qu, Jihao Gu, Yangguang Li 0001, Jianke Zhu |
NAACL (Long Papers) | 17 |
| 2025 | Semantic-preserved point-based human avatar
Lixiang Lin, Jianke Zhu |
Comput. Vis. Image Underst. | 2 |
| 2025 | Hexagonal mesh-based neural rendering for real-time rendering and fast reconstruction
Yisu Zhang, Jianke Zhu, Lixiang Lin |
Comput. Vis. Image Underst. | 2 |
| 2025 | TokenPacker: Efficient Visual Projector for Multimodal LLM
Wentong Li 0001, Yuqian Yuan, Jian Liu 0012, Dongqi Tang, Song Wang 0019, Jie Qin 0004, Jianke Zhu, Lei Zhang 0006 |
Int. J. Comput. Vis. | 7 |
| 2025 | Warped convolutional neural networks for large homography transformation with psl(3) algebra
Xinrui Zhan, Wenyu Liu 0001, Risheng Yu, Jianke Zhu, Yang Li 0041 |
Neurocomputing | 4 |
| 2025 | DGNR: Density-Guided Neural Point Rendering of Large Driving ScenesabstractDespite the recent success of Neural Radiance Field (NeRF), it is still challenging to render large-scale driving scenes with long trajectories, particularly when the rendering quality and efficiency are in high demand. Existing methods for such scenes usually involve with spatial warping, geometric supervision from zero-shot normal or depth estimation, or scene division strategies, where the synthesized views are often blurry or fail to meet the requirement of efficient rendering. To address the above challenges, this paper presents a novel framework that learns a density space from the scenes to guide the construction of a point-based renderer, dubbed as DGNR (Density-Guided Neural Rendering). In DGNR, geometric priors are no longer needed, which can be intrinsically learned from the density space through volumetric rendering. Specifically, we make use of a differentiable renderer to synthesize images from the neural density features obtained from the learned density space. A density-based fusion module and geometric regularization are proposed to optimize the density space. By conducting experiments on a widely used autonomous driving dataset, we have validated the effectiveness of DGNR in synthesizing photorealistic driving scenes and achieving real-time capable rendering. Our project page is available athttps://github.com/JOP-Lee/DGNR-Rendering. Note to Practitioners—While Neural Radiance Field (NeRF) has been gaining attraction, it is still challenging to create highly detailed, efficient renderings of large driving scenes. Current methods often resort to spatial warping, geometric guidance from tools like zero-shot normal or depth estimates, or dividing the scene into smaller parts. Unfortunately, these techniques can result in blurred images or fail to meet efficiency needs. To solve these challenges, we introduce a learned density space to build a point-based renderer, termed Density-Guided Neural Rendering (DGNR). With DGNR, we no longer need geometric priors because the density space can inherently learn them through volume rendering. Specifically, we use a flexible renderer to create images from the neural density features derived from the learned density space. We have also proposed a density-based fusion module and geometric regularization to optimize the density space. We evaluated DGNR on a popular autonomous driving dataset and found it to be effective in creating realistic driving scenes and capable of real-time rendering. Project page:https://github.com/JOP-Lee/DGNR-Rendering. Zhuopeng Li, Chenming Wu, Liangjun Zhang, Jianke Zhu |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | LPM: Efficient 3D Content Creation From Single Image by Large-Scale Partial 3D ModelingabstractSynthesizing 3D content from single image has great potential in many real-world applications. To deal with the inherent ambiguity of single image, existing methods usually leverage pre-trained 2D diffusion models for computational intensive per-instance optimization. Although having been able to create 3D assets in a feed-forward manner, the efficacy of recent advances in 3D foundation models is still limited due to neglecting geometric cues from images. To address this issue, we propose an efficient 3D foundation model named LPM to synthesize 3D content from an image. Like the masked modeling in the 2D image domain, the key of our approach is to learn 3D representations from incomplete visible shapes. By taking advantage of a synthesis-by-analysis paradigm, we establish an efficient pipeline to first estimate the visible portions and then generate the complete 3D representations. Based on the principle that an image is the projection of 3D model, we initially estimate partial 3D voxel features from single image, which are further projected onto orthogonal planes to form an incomplete yet efficient triplane representation. Subsequently, an autoencoder is employed to model a complete triplane representation based on the incomplete parts. We train our model on massive data with over 250 million parameters to enhance its generalization capability. The experimental results show that LPM can generate high-fidelity 3D objects from an image within 0.1 seconds, which is more effective than the existing feed-forward approaches. Our implementation and pre-trained models will be made publicly available. Yisu Zhang, Chaohui Yu, Fan Wang 0019, Jianke Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Offboard Occupancy Refinement With Hybrid Propagation for Autonomous DrivingabstractVision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic estimation, continuity across varying viewpoints, and single-view occlusion. Our paper introduces OccFiner, a novel offboard framework designed to enhance the accuracy of vision-based occupancy predictions. OccFiner operates in two hybrid phases: 1) a multi-to-multi local propagation network that implicitly aligns and processes multiple local frames for correcting onboard model errors and consistently enhancing occupancy accuracy across all distances. 2) the region-centric global propagation, focuses on refining labels using explicit multi-view geometry and integrating sensor bias, particularly for increasing the accuracy of distant occupied voxels. Extensive experiments demonstrate that OccFiner improves both geometric and semantic accuracy across various types of coarse occupancy, setting a new state-of-the-art performance on the SemanticKITTI dataset. Notably, OccFiner significantly boosts the performance of vision-based SSC models, achieving accuracy levels competitive with established LiDAR-based onboard SSC methods. Furthermore, OccFiner is the first to achieve automatic annotation of SSC in a purely vision-based approach. Quantitative experiments prove that OccFiner successfully facilitates occupancy data loop-closure in autonomous driving. Additionally, we quantitatively and qualitatively validate the superiority of the offboard approach on city-level SSC static maps. The source code will be made publicly available at https://github.com/MasterHow/OccFiner Hao Shi 0004, Song Wang 0019, Jiaming Zhang 0001, Xiaoting Yin, Guangming Wang 0001, Jianke Zhu, Kailun Yang 0001, Kaiwei Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | FastHuman: Reconstructing High-Quality Clothed Human in MinutesabstractWe propose an approach for optimizing high-quality clothed human body shapes in minutes, using multi-view posed images. While traditional neural rendering methods struggle to disentangle geometry and appearance using only rendering loss, and are computationally intensive, our method uses a mesh-based patch warping technique to ensure multi-view photometric consistency, and sphere harmonics (SH) illumination to refine geometric details efficiently. We employ oriented point clouds’ shape representation and SH shading, which significantly reduces optimization and rendering times compared to implicit methods. Our approach has demonstrated promising results on both synthetic and real-world datasets, making it an effective solution for rapidly generating high-quality human body shapes. Project page with source code is https://l1346792580123.github.io/nccsfs/. Lixiang Lin, Songyou Peng, Qijun Gan, Jianke Zhu |
3DV | 4 |
| 2024 | Fine-Grained Multi-View Hand Reconstruction Using Inverse RenderingabstractReconstructing high-fidelity hand models with intricate textures plays a crucial role in enhancing human-object interaction and advancing real-world applications. Despite the state-of-the-art methods excelling in texture generation and image rendering, they often face challenges in accurately capturing geometric details. Learning-based approaches usually offer better robustness and faster inference, which tend to produce smoother results and require substantial amounts of training data. To address these issues, we present a novel fine-grained multi-view hand mesh reconstruction method that leverages inverse rendering to restore hand poses and intricate details. Firstly, our approach predicts a parametric hand mesh model through Graph Convolutional Networks (GCN) based method from multi-view images. We further introduce a novel Hand Albedo and Mesh (HAM) optimization module to refine both the hand mesh and textures, which is capable of preserving the mesh topology. In addition, we suggest an effective mesh-based neural rendering scheme to simultaneously generate photo-realistic image and optimize mesh geometry by fusing the pre-trained rendering network with vertex features. We conduct the comprehensive experiments on InterHand2.6M, DeepHandMesh and dataset collected by ourself, whose promising results show that our proposed approach outperforms the state-of-the-art methods on both reconstruction accuracy and rendering quality. Code and dataset are publicly available at https://github.com/agnJason/FMHR. Qijun Gan, Wentong Li 0001, Jinwei Ren, Jianke Zhu |
AAAI | 4 |
| 2024 | MGMap: Mask-Guided Learning for Online Vectorized HD Map ConstructionabstractCurrently, high-definition (HD) map construction leans towards a lightweight online generation tendency, which aims to preserve timely and reliable road scene information. However, map elements contain strong shape priors. Subtle and sparse annotations make current detection-based frameworks ambiguous in locating relevant feature scopes and cause the loss of detailed structures in prediction. To alleviate these problems, we propose MGMap, a mask-guided approach that effectively highlights the informative regions and achieves precise map element localization by introducing the learned masks. Specifically, MGMap employs learned masks based on the enhanced multi-scale BEV features from two perspectives. At the instance level, we propose the Mask-activated instance (MAI) decoder, which incorporates global instance and structural information into instance queries by the activation of instance masks. At the point level, a novel position-guided mask patch refinement (PG-MPR) module is designed to refine point locations from a finer-grained perspective, enabling the extraction of point-specific patch information. Compared to the baselines, our proposed MGMap achieves a notable improvement of around 10 mAP for different input modalities. Extensive experiments also demonstrate that our approach showcases strong robustness and generalization capabilities. Our code can be found at https://github.com/xiaolul2/MGMap. Song Wang 0019, Wentong Li 0001, Ruizi Yang, Junbo Chen, Jianke Zhu |
CVPR | 6 |
| 2024 | Not All Voxels are Equal: Hardness-Aware Semantic Scene Completion with Self-DistillationabstractSemantic scene completion, also known as semantic oc-cupancy prediction, can provide dense geometric and semantic information for autonomous vehicles, which attracts the increasing attention of both academia and industry. Un-fortunately, existing methods usually formulate this task as a voxel-wise classification problem and treat each voxel equally in 3D space during training. As the hard voxels have not been paid enough attention, the performance in some challenging regions is limited. The 3D dense space typically contains a large number of empty voxels, which are easy to learn but require amounts of computation due to handling all the voxels uniformly for the existing models. Further-more, the voxels in the boundary region are more challenging to differentiate than those in the interior. In this paper, we propose HASSC approach to train the semantic scene completion model with hardness-aware design. The global hardness from the network optimization process is defined for dynamical hard voxel selection. Then, the local hard-ness with geometric anisotropy is adopted for voxel- wise refinement. Besides, self-distillation strategy is introduced to make training process stable and consistent. Extensive experiments show that our HASSC scheme can effectively promote the accuracy of the baseline model without incur-ring the extra inference cost. Source code is available at: https://github.com/songw-zju/HASSC. Song Wang 0019, Wentong Li 0001, Wenyu Liu 0005, Junbo Chen, Jianke Zhu |
CVPR | 7 |
| 2024 | Osprey: Pixel Understanding with Visual Instruction TuningabstractMultimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their ad-vancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incor-porating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimen-tal results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey. Yuqian Yuan, Wentong Li 0001, Jian Liu 0012, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang 0006, Jianke Zhu |
CVPR | 8 |
| 2024 | HO-Gaussian: Hybrid Optimization of 3D Gaussian Splatting for Urban Scenes
Zhuopeng Li, Chenming Wu, Jianke Zhu, Liangjun Zhang |
ECCV (60) | 4 |
| 2024 | HVOFusion: Incremental Mesh Reconstruction Using Hybrid Voxel Octree
Shaofan Liu, Junbo Chen, Jianke Zhu |
IJCAI | 3 |
| 2024 | Label-efficient Semantic Scene Completion with Scribble Annotations
Song Wang 0019, Wentong Li 0001, Hao Shi 0004, Kailun Yang 0001, Junbo Chen, Jianke Zhu |
IJCAI | 7 |
| 2024 | Box2Mask: Box-Supervised Instance Segmentation via Level-Set EvolutionabstractIn contrast to fully supervised methods using pixel-wise mask labels, box-supervised instance segmentation takes advantage of simple box annotations, which has recently attracted increasing research attention. This paper presents a novel single-shot instance segmentation approach, namely Box2Mask, which integrates the classical level-set evolution model into deep neural network learning to achieve accurate mask prediction with only bounding box supervision. Specifically, both the input image and its deep features are employed to evolve the level-set curves implicitly, and a local consistency module based on a pixel affinity kernel is used to mine the local context and spatial relations. Two types of single-stage frameworks, i.e., CNN-based and transformer-based frameworks, are developed to empower the level-set evolution for box-supervised instance segmentation, and each framework consists of three essential components: instance-aware decoder, box-level matching assignment and level-set evolution. By minimizing the level-set energy function, the mask map of each instance can be iteratively optimized within its bounding box annotation. The experimental results on five challenging testbeds, covering general scenes, remote sensing, medical and scene text images, demonstrate the outstanding performance of our proposed Box2Mask approach for box-supervised instance segmentation. In particular, with the Swin-Transformer large backbone, our Box2Mask obtains 42.4% mask AP on COCO, which is on par with the recently developed fully mask-supervised methods. Wentong Li 0001, Wenyu Liu 0005, Jianke Zhu, Miaomiao Cui, Risheng Yu, Xian-Sheng Hua 0001, Lei Zhang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Pyramid Deep Fusion Network for Two-Hand Reconstruction From RGB-D ImagesabstractAccurately recovering the dense 3D mesh of both hands from monocular images poses considerable challenges due to occlusions and projection ambiguity. Most of the existing methods extract features from color images to estimate the root-aligned hand meshes, which neglect the crucial depth and scale information in the real world. Given the noisy sensor measurements with limited resolution, depth-based methods predict 3D keypoints rather than a dense mesh. These limitations motivate us to take advantage of these two complementary inputs to acquire dense hand meshes on a real-world scale. In this work, we propose an end-to-end framework for recovering dense meshes for both hands, which employ single-view RGB-D image pairs as input. The primary challenge lies in effectively utilizing two different input modalities to mitigate the blurring effects in RGB images and noises in depth images. Instead of directly treating depth maps as additional channels for RGB images, we encode the depth information into the unordered point cloud to preserve more geometric details. Specifically, our framework employs ResNet50 and PointNet++ to derive features from RGB and point cloud, respectively. Additionally, we introduce a novel pyramid deep fusion network (PDFNet) to aggregate features at different scales, which demonstrates superior efficacy compared to previous fusion strategies. Furthermore, we employ a GCN-based decoder to process the fused features and recover the corresponding 3D pose and dense mesh. Through comprehensive ablation experiments, we have not only demonstrated the effectiveness of our proposed fusion algorithm but also outperformed the state-of-the-art approaches on publicly available datasets. To reproduce the results, we will make our source code and models publicly available at https://github.com/zijinxuxu/PDFNet. Jinwei Ren, Jianke Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Domain Adaptation Transformer for Unsupervised Driving-Scene Segmentation in Adverse ConditionsabstractSemantic segmentation in driving scenarios is important for modern autonomous driving technology. While the existing methods have shown promising results in segmenting normal-condition images, their performance in adverse scenes remains unsatisfactory due to limited visual field and lack of annotation. To address this issue, we propose an unsupervised domain adaptation semantic segmentation method with the transformer architecture, namely ACSegFormer, for driving-scene adverse conditions, aiming at mining image features in visually restricted scenes. Three effective training strategies are proposed in ACSegFormer to learn the latent image context relations and to reduce the gaps between different domains: an entropy-based pseudo label correction scheme that refines the target domain predictions with the normal reference predictions, an optimal transport-based inter-domain alignment module that performs domain alignment on the outputs of transformer encoder, and a masked context learning module that enhances the model’s ability to perceive the missing information of target domain image. Our ACSegFormer has no additional training parameters on top of the existing transformer segmentation framework, which can be easily used for self-training-based unsupervised domain adaptation approaches. The experimental results show that our ACSegFormer achieves state-of-the-art performance on driving-scene segmentation benchmarks in adverse conditions, including Dark Zurich and ACDC. Codes and models are available athttps://github.com/wenyyu/ACSegFormer. Wenyu Liu 0005, Song Wang 0019, Jianke Zhu, Xuansong Xie, Lei Zhang 0006 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | SLED: Structure Learning based Denoising for RecommendationabstractIn recommender systems, click behaviors play a fundamental role in mining users’ interests and training models (clicked items as positive samples). Such signals are implicit feedback and are arguably less representative of users’ inherent interests. Most existing works denoise implicit feedback by introducing external signals, such as gaze, dwell time, and “like” behaviors. However, such explicit feedback is not always routinely available, or might be problematic to collect on a large scale. In this paper, we identify that an interaction’s related structural patterns in its neighborhood graph are potentially correlated with some outcome of implicit feedback (i.e., users’ ratings after consuming items), analogous to findings in other domains such as social networks. Inspired by this finding, we propose a novel Structure LEarning based Denoising (SLED) framework for denoising recommendation without explicit signals, which consists of two phases: center-aware graph structure learning and denoised recommendation . Phase 1 pre-trains a structural encoder in a self-supervised manner and learns to capture an interaction’s related structural patterns in its neighborhood graph. Phase 2 transfers the structure encoder to downstream recommendation datasets, which helps to down-weight the effect of noisy interactions on user interest modeling and loss calculation. We collect a relatively noisy industrial dataset across several days during a period of product promotion festival. Extensive experiments on this dataset and multiple public datasets demonstrate that the proposed SLED framework can significantly improve the recommendation quality over various base recommendation models. Shengyu Zhang 0001, Tan Jiang, Kun Kuang 0001, Fuli Feng, Zhou Zhao 0001, Jianke Zhu, Hongxia Yang, Tat-Seng Chua, Fei Wu 0001 |
ACM Trans. Inf. Syst. | 8 |
| 2023 | READ: Large-Scale Neural Scene Rendering for Autonomous DrivingabstractWith the development of advanced driver assistance systems~(ADAS) and autonomous vehicles, conducting experiments in various scenarios becomes an urgent need. Although having been capable of synthesizing photo-realistic street scenes, conventional image-to-image translation methods cannot produce coherent scenes due to the lack of 3D information. In this paper, a large-scale neural rendering method is proposed to synthesize the autonomous driving scene~(READ), which makes it possible to generate large-scale driving scenes in real time on a PC through a variety of sampling schemes. In order to effectively represent driving scenarios, we propose an ω-net rendering network to learn neural descriptors from sparse point clouds. Our model can not only synthesize photo-realistic driving scenes but also stitch and edit them. The promising experimental results show that our model performs well in large-scale driving scenarios. Zhuopeng Li, Jianke Zhu |
AAAI | 3 |
| 2023 | LiDAR2Map: In Defense of LiDAR-Based Semantic Map Construction Using Online Camera DistillationabstractSemantic map construction under bird's-eye view (BEV) plays an essential role in autonomous driving. In contrast to camera image, LiDAR provides the accurate 3D observations to project the captured 3D features onto BEV space inherently. However, the vanilla LiDAR-based BEV feature often contains many indefinite noises, where the spatial features have little texture and semantic cues. In this paper, we propose an effective LiDAR-based method to build semantic map. Specifically, we introduce a BEV pyramid feature decoder that learns the robust multi-scale BEV features for semantic map construction, which greatly boosts the accuracy of the LiDAR-based method. To mitigate the defects caused by lacking semantic cues in LiDAR data, we present an online Camera-to-LiDAR distillation scheme to facilitate the semantic learning from image to point cloud. Our distillation scheme consists of feature-level and log it-level distillation to absorb the semantic information from camera in BEV. The experimental results on challenging nuScenes dataset demonstrate the efficacy of our proposed LiDAR2Map on semantic map construction, which significantly outperforms the previous LiDAR-based methods over 27.9% mIoU and even performs better than the state-of-the-art camera-based approaches. Source code is available at: https://github.com/songw-zjuILiDAR2Map. Song Wang 0019, Wentong Li 0001, Wenyu Liu 0005, Jianke Zhu |
CVPR | 5 |
| 2023 | Multi-View Stereo Representation Revist: Region-Aware MVSNetabstractDeep learning-based multi-view stereo has emerged as a powerful paradigm for reconstructing the complete geometrically-detailed objects from multi-views. Most of the existing approaches only estimate the pixel-wise depth value by minimizing the gap between the predicted point and the intersection of ray and surface, which usually ignore the surface topology. It is essential to the textureless regions and surface boundary that cannot be properly reconstructed. To address this issue, we suggest to take advantage of point-to-surface distance so that the model is able to perceive a wider range of surfaces. To this end, we predict the distance volume from cost volume to estimate the signed distance of points around the surface. Our proposed RA-MVSNet is patch-awared, since the perception range is enhanced by associating hypothetical planes with a patch of surface. Therefore, it could increase the completion of textureless regions and reduce the outliers at the boundary. Moreover, the mesh topologies with fine details can be generated by the introduced distance volume. Comparing to the conventional deep learning-based multi-view stereo methods, our proposed RA-MVSNet approach obtains more complete reconstruction results by taking advantage of signed distance supervision. The experiments on both the DTU and Tanks & Temples datasets demonstrate that our proposed approach achieves the state-of-the-art results. Yisu Zhang, Jianke Zhu, Lixiang Lin |
CVPR | 2 |
| 2023 | Point2Mask: Point-supervised Panoptic Segmentation via Optimal TransportabstractWeakly-supervised image segmentation has recently attracted increasing research attentions, aiming to avoid the expensive pixel-wise labeling. In this paper, we present an effective method, namely Point2Mask, to achieve high-quality panoptic prediction using only a single random point annotation per target for training. Specifically, we formulate the panoptic pseudo-mask generation as an Optimal Transport (OT) problem, where each ground-truth (gt) point label and pixel sample are defined as the label supplier and consumer, respectively. The transportation cost is calculated by the introduced task-oriented maps, which focus on the category-wise and instance-wise differences among the various thing and stuff targets. Furthermore, a centroid-based scheme is proposed to set the accurate unit number for each gt point supplier. Hence, the pseudo-mask generation is converted into finding the optimal transport plan at a globally minimal transportation cost, which can be solved via the Sinkhorn-Knopp Iteration. Experimental results on Pascal VOC and COCO demonstrate the promising performance of our proposed Point2Mask approach to point-supervised panoptic segmentation. Source code is available at: https://github.com/LiWentomng/Point2Mask. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Jianke Zhu, Jianshu Li, Jian Liu 0012, Lei Zhang 0006 |
ICCV | 4 |
| 2023 | Structure First Detail Next: Image Inpainting with Pyramid GeneratorabstractRecent deep generative models have achieved promising performance in image inpainting. However, it is still challenging for a neural network to generate realistic image details and textures due to its inherent spectral bias. We suggest adopting a ‘structure first detail next’ workflow for image inpainting by knowing how artists work. Thus, we propose to build a Pyramid Generator by stacking several sub-generators, where lower-layer sub-generators focus on restoring image structures. In contrast, the higher-layer sub-generators emphasize image details. Our model progressively restores the input through the entire pyramid in a bottom-up fashion. Notably, our approach has a learning scheme of progressively increasing hole size, which allows it to restore large-hole images. In addition, our method could fully exploit the benefits of learning with high-resolution images and hence is suitable for high-resolution image inpainting. Extensive experimental results on benchmark datasets have validated the effectiveness of our approach compared with state-of-the-arts. Shuyi Qu, Zhenxing Niu, Jianke Zhu, Bin Dong 0003, Kaizhu Huang |
ICME | 3 |
| 2023 | SDFMAP: Neural Signed Distance Fields for Mapping and Positioning in Real-TimeabstractNeural surface reconstruction has recently gained a bit attention due to the promising result on scene rendering. Nevertheless, most of existing approaches either treat the camera parameters as the prior during training or indirectly estimate them through structure-from-motion. To tap the potential of implicit neural networks, we present a novel end-to-end neural network, termed SDFMAP, without any prior knowledge of the scene, like pre-computed camera parameters and pretrained geometric priors. Specifically, our method adopts a single multilayer perceptron to achieve simultaneously pose estimation and indoor scene reconstruction in real-time through learning the truncated signed distance function. Comparing to the recent neural implicit vSLAM systems, our approach achieves higher tracking speed via a lightweight network. Experiments on several challenging benchmark datasets show that our SDFMAP method achieves the state-of-the-art results on camera tracking and scene reconstruction. Shaofan Liu, Jianke Zhu |
IROS | 2 |
| 2023 | ECTLO: Effective Continuous-Time Odometry Using Range Image for LiDAR with Small FoVabstractPrism-based LiDARs are more compact and cheaper than the conventional mechanical multi-line spinning LiDARs, which have become increasingly popular in robotics, recently. However, there are several challenges for these new LiDAR sensors, including small field of view, severe motion distortions, and irregular patterns. These difficulties hinder them from being widely used in LiDAR odometry, practically. To tackle these problems, we present an effective continuous-time LiDAR odometry (ECTLO) method for the Risley-prism-based LiDARs with non-repetitive scanning patterns. A single range image covering historical points in LiDAR's small FoV is adopted for efficient map representation. To account for the noisy data from occlusions after map updating, a filter-based point-to-plane Gaussian Mixture Model is used for robust registration. Moreover, a LiDAR-only continuous-time motion model is employed to relieve the inevitable distortions. Extensive experiments have been conducted on various testbeds using the prism-based LiDARs with different scanning patterns, whose promising results demonstrate the efficacy of our proposed approach. Xin Zheng 0009, Jianke Zhu |
IROS | 2 |
| 2023 | Moiré Backdoor Attack (MBA): A Novel Trigger for Pedestrian Detectors in the Physical WorldabstractA backdoor attack is executed by injecting a few poisoned samples into the training dataset of Deep Neural Networks (DNNs), enabling attackers to implant a hidden manipulation. This manipulation can be triggered during inference to exhibit controlled behavior, posing risks in real-world deployments. In this paper, we specifically focus on the safety-critical task of pedestrian detection and propose a novel backdoor trigger by exploiting the Moiré effect. The Moiré effect, a common physical phenomenon, disrupts camera-captured images by introducing Moiré patterns and unavoidable interference. Our method comprises three key steps. Firstly, we analyze the Moiré effect's cause and simulate its patterns on pedestrians' clothing. Next, we embed these Moiré patterns as a backdoor trigger into digital images and use this dataset to train a backdoored detector. Finally, we physically test the trained detector by wearing clothing that generates Moiré patterns. We demonstrate that individuals wearing such clothes can effectively evade detection by the backdoored model while wearing regular clothes does not trigger the attack, ensuring the attack remains covert. Extensive experiments in both digital and physical spaces thoroughly demonstrate the effectiveness and efficacy of our proposed Moiré Backdoor Attack. Hui Wei 0004, Hanxun Yu, Zhixiang Wang 0001, Jianke Zhu, Zheng Wang 0007 |
ACM Multimedia | 5 |
| 2023 | Label-efficient Segmentation via Affinity PropagationabstractWeakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i.e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach. Wentong Li 0001, Yuqian Yuan, Song Wang 0019, Wenyu Liu 0005, Dongqi Tang, Jian Liu 0012, Jianke Zhu, Lei Zhang 0006 |
NeurIPS | 7 |
| 2023 | End-to-end weakly-supervised single-stage multiple 3D hand mesh reconstruction from a single RGB image
Jinwei Ren, Jianke Zhu |
Comput. Vis. Image Underst. | 2 |
| 2023 | Multiview Textured Mesh Recovery by Differentiable RenderingabstractAlthough having achieved the promising results on shape and color recovery through self-supervision, the multi-layer perceptrons-based methods usually suffer from heavy computational cost on learning the deep implicit surface representation. Since rendering each pixel requires a forward network inference, it is very computationally intensive to synthesize a whole image. To tackle these challenges, we propose an effective coarse-to-fine approach to recover the textured mesh from multi-views in this paper. Specifically, a differentiable Poisson Solver is employed to represent the object’s shape, which is able to produce topology-agnostic and watertight surfaces. To account for depth information, we optimize the shape geometry by minimizing the differences between the rendered mesh and the predicted depth from multi-view stereo. In contrast to the implicit neural representation on shape and color, we introduce a physically-based inverse rendering scheme to jointly estimate the environment lighting and object’s reflectance, which is able to render the high resolution image at real-time. The texture of reconstructed mesh is interpolated from a learnable dense texture grid. We have conducted the extensive experiments on several multi-view stereo datasets, whose promising results demonstrate the efficacy of our proposed approach. The code is available athttps://github.com/l1346792580123/diff. Lixiang Lin, Jianke Zhu, Yisu Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Improving Nighttime Driving-Scene Segmentation via Dual Image-Adaptive Learnable FiltersabstractSemantic segmentation on driving-scene images is vital for autonomous driving. Although encouraging performance has been achieved on daytime images, the performance on nighttime images are less satisfactory due to the insufficient exposure and the lack of labeled data. To address these issues, we present an add-on module called dual image-adaptive learnable filters (DIAL-Filters) to improve the semantic segmentation in nighttime driving conditions, aiming at exploiting the intrinsic features of driving-scene images under different illuminations. DIAL-Filters consist of two parts, including an image-adaptive processing module (IAPM) and a learnable guided filter (LGF). With DIAL-Filters, we design both unsupervised and supervised frameworks for nighttime driving-scene segmentation, which can be trained in an end-to-end manner. Specifically, the IAPM module consists of a small convolutional neural network with a set of differentiable image filters, where each image can be adaptively enhanced for better segmentation with respect to the different illuminations. The LGF is employed to enhance the output of segmentation network to get the final segmentation result. The DIAL-Filters are light-weight and efficient and they can be readily applied for both daytime and nighttime images. Our experiments show that DAIL-Filters can significantly improve the supervised segmentation performance on ACDC_Night and NightCity datasets, while it demonstrates the state-of-the-art performance on unsupervised nighttime semantic segmentation on Dark Zurich and Nighttime Driving testbeds. Codes and models are available athttps://github.com/wenyyu/IA-Seg. Wenyu Liu 0005, Wentong Li 0001, Jianke Zhu, Miaomiao Cui, Xuansong Xie, Lei Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Topology-preserved human reconstruction with details
Lixiang Lin, Jianke Zhu |
Vis. Comput. | 2 |
| 2022 | Image-Adaptive YOLO for Object Detection in Adverse Weather ConditionsabstractThough deep learning-based object detection methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. The existing methods either have difficulties in balancing the tasks of image enhancement and object detection, or often ignore the latent information beneficial for detection. To alleviate this problem, we propose a novel Image-Adaptive YOLO (IA-YOLO) framework, where each image can be adaptively enhanced for better detection performance. Specifically, a differentiable image processing (DIP) module is presented to take into account the adverse weather conditions for YOLO detector, whose parameters are predicted by a small convolutional neural network (CNN-PP). We learn CNN-PP and YOLOv3 jointly in an end-to-end fashion, which ensures that CNN-PP can learn an appropriate DIP to enhance the image for detection in a weakly supervised manner. Our proposed IA-YOLO approach can adaptively process images in both normal and adverse weather conditions. The experimental results are very encouraging, demonstrating the effectiveness of our proposed IA-YOLO method in both foggy and low-light scenarios. The source code can be found at https://github.com/wenyyu/Image-Adaptive-YOLO. Wenyu Liu 0005, Gaofeng Ren, Runsheng Yu, Shi Guo, Jianke Zhu, Lei Zhang 0006 |
AAAI | 5 |
| 2022 | Multimodal Adversarially Learned Inference with Factorized DiscriminatorsabstractLearning from multimodal data is an important research topic in machine learning, which has the potential to obtain better representations. In this work, we propose a novel approach to generative modeling of multimodal data based on generative adversarial networks. To learn a coherent multimodal generative model, we show that it is necessary to align different encoder distributions with the joint decoder distribution simultaneously. To this end, we construct a specific form of the discriminator to enable our model to utilize data efficiently, which can be trained constrastively. By taking advantage of contrastive learning through factorizing the discriminator, we train our model on unimodal data. We have conducted experiments on the benchmark datasets, whose promising results show that our proposed approach outperforms the-state-ofthe-art methods on a variety of metrics. The source code is publicly available at https://github.com/6b5d/mmali. Wenxue Chen, Jianke Zhu |
AAAI | 2 |
| 2022 | Homography Decomposition Networks for Planar Object TrackingabstractPlanar object tracking plays an important role in AI applications, such as robotics, visual servoing, and visual SLAM. Although the previous planar trackers work well in most scenarios, it is still a challenging task due to the rapid motion and large transformation between two consecutive frames. The essential reason behind this problem is that the condition number of such a non-linear system changes unstably when the searching range of the homography parameter space becomes larger. To this end, we propose a novel Homography Decomposition Networks~(HDN) approach that drastically reduces and stabilizes the condition number by decomposing the homography transformation into two groups. Specifically, a similarity transformation estimator is designed to predict the first group robustly by a deep convolution equivariant network. By taking advantage of the scale and rotation estimation with high confidence, a residual transformation is estimated by a simple regression model. Furthermore, the proposed end-to-end network is trained in a semi-supervised fashion. Extensive experiments show that our proposed approach outperforms the state-of-the-art planar tracking methods at a large margin on the challenging POT, UCSB and POIC datasets. Codes and models are available at https://github.com/zhanxinrui/HDN. Xinrui Zhan, Yueran Liu, Jianke Zhu, Yang Li 0041 |
AAAI | 3 |
| 2022 | Oriented RepPoints for Aerial Object DetectionabstractIn contrast to the generic object, aerial targets are often non-axis aligned with arbitrary orientations having the cluttered surroundings. Unlike the mainstreamed approaches regressing the bounding box orientations, this paper proposes an effective adaptive points learning approach to aerial object detection by taking advantage of the adaptive points representation, which is able to capture the geometric information of the arbitrary-oriented instances. To this end, three oriented conversion functions are presented to facilitate the classification and localization with accurate orientation. Moreover, we propose an effective quality assessment and sample assignment scheme for adaptive points learning toward choosing the representative oriented reppoints samples during training, which is able to capture the non-axis aligned features from adjacent objects or background noises. A spatial constraint is introduced to penalize the outlier points for roust adaptive learning. Experimental results on four challenging aerial datasets including DOTA, HRSC2016, UCAS-AOD and DIOR-R, demonstrate the efficacy of our proposed approach. The source code is availabel at: https://github.com/LiWentomng/OrientedRepPoints. Wentong Li 0001, Kaixuan Hu, Jianke Zhu |
CVPR | 4 |
| 2022 | Box-Supervised Instance Segmentation with Level Set Evolution
Wentong Li 0001, Wenyu Liu 0001, Jianke Zhu, Miaomiao Cui, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ECCV (29) | 3 |
| 2022 | Horizontal-to-Vertical Video ConversionabstractAt this blooming age of social media and mobile platform, mass consumers are migrating from horizontal video to vertical contents delivered on hand-held devices. Accordingly, revitalizing the exposure of horizontal video becomes vital and urgent, which is hereby tackled by our automated horizontal-to-vertical (abbreviated as H2V) video conversion framework. Essentially, the {\it \textbf{H2V}} framework performs subject-preserving video cropping instantiated in the proposed Rank-SS module. Rank-SS incorporates object detection to discover candidate subjects, from which we select the primary subject-to-preserve leveraging location, appearance, and salient cues in a convolutional neural network. In addition to converting horizontal videos vertically by cropping around the selected subject, automatic shot detection and multi-object tracking are integrated into the {\it \textbf{H2V}} framework to accommodate long and complex videos. To develop {\it \textbf{H2V}} systems, we collect an {\it \textbf{H2V-142K}} dataset containing 125 videos (132K frames) and 9,500 cover images annotated with primary subject bounding boxes. On {\it \textbf{H2V-142K}} and public object detection datasets, our method demonstrates promising results on the subject selection comparing to the related solutions. Furthermore, our {\it \textbf{H2V}} framework is industrially deployed hosting millions of daily active users and exhibits favorable H2V conversion performance. By making this dataset as well as our approach publicly available, we wish to pave the way for more horizontal-to-vertical video conversion research. Our collected H2V-142K dataset is available atH2V-142K website. Tun Zhu, Daoxin Zhang, Yao Hu 0002, Tianran Wang, Jianke Zhu |
IEEE Trans. Multim. | 6 |
| 2021 | Manifold adversarial training for supervised and semi-supervised learning
Shufei Zhang, Kaizhu Huang, Jianke Zhu |
Neural Networks | 3 |
| 2021 | Attribute-Aware Pedestrian Detection in a CrowdabstractPedestrian detection is an initial step to perform outdoor scene analysis, which plays an essential role in many real-world applications. Although having enjoyed the merits of deep learning frameworks from the generic object detectors, pedestrian detection is still a very challenging task due to heavy occlusions, and highly crowded group. Generally, the conventional detectors are unable to differentiate individuals from each other effectively under such a dense environment. To tackle this critical problem, we propose an attribute-aware pedestrian detector to explicitly model people's semantic attributes in a high-level feature detection fashion. Besides the typical semantic features, center position, target's scale, and offset, we introduce a pedestrian-oriented attribute feature to encode the high-level semantic differences among the crowd. Moreover, a novel attribute-feature-based Non-Maximum Suppression (NMS) is proposed to distinguish the person from a highly overlapped group by adaptively rejecting the false-positive results in a very crowd settings. Furthermore, an enhanced ground truth target is designed to alleviate the difficulties caused by the attribute configuration, and to ease the class imbalance issue during training. Finally, we evaluate our proposed attribute-aware pedestrian detector on three benchmark datasets including CityPerson, CrowdHuman, and EuroCityPerson, and achieves the state-of-the-art results. Lixiang Lin, Jianke Zhu, Yang Li 0041, Yun-chen Chen, Yao Hu 0002, Steven C. H. Hoi |
IEEE Trans. Multim. | 3 |
| 2020 | Deep Time-Stream Framework for Click-through Rate Prediction by Tracking Interest EvolutionabstractClick-through rate (CTR) prediction is an essential task in industrial applications such as video recommendation. Recently, deep learning models have been proposed to learn the representation of users' overall interests, while ignoring the fact that interests may dynamically change over time. We argue that it is necessary to consider the continuous-time information in CTR models to track user interest trend from rich historical behaviors. In this paper, we propose a novel Deep Time-Stream framework (DTS) which introduces the time information by an ordinary differential equations (ODE). DTS continuously models the evolution of interests using a neural network, and thus is able to tackle the challenge of dynamically representing users' interests based on their historical behaviors. In addition, our framework can be seamlessly applied to any existing deep CTR models by leveraging the additional Time-Stream Module, while no changes are made to the original CTR models. Experiments on public dataset as well as real industry dataset with billions of samples demonstrate the effectiveness of proposed approaches, which achieve superior performance compared with existing methods. Wenhao Zheng 0001, Yao Hu 0002, Jianke Zhu, Ming Li 0005 |
AAAI | 6 |
| 2020 | DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsabstractIn this paper, we propose to investigate the problem of out-of-domain visio-linguistic pretraining, where the pretraining data distribution differs from that of downstream data on which the pretrained model will be fine-tuned. Existing methods for this problem are purely likelihood-based, leading to the spurious correlations and hurt the generalization ability when transferred to out-of-domain downstream tasks. By spurious correlation, we mean that the conditional probability of one token (object or word) given another one can be high (due to the dataset biases) without robust (causal) relationships between them. To mitigate such dataset biases, we propose a Deconfounded Visio-Linguistic Bert framework, abbreviated as DeVLBert, to perform intervention-based learning. We borrow the idea of the backdoor adjustment from the research field of causality and propose several neural-network based architectures for Bert-style out-of-domain pretraining. The quantitative results on three downstream tasks, Image Retrieval (IR), Zero-shot IR, and Visual Question Answering, show the effectiveness of DeVLBert by boosting generalization ability. Shengyu Zhang 0001, Tan Jiang, Kun Kuang 0001, Zhou Zhao 0001, Jianke Zhu, Hongxia Yang, Fei Wu 0001 |
ACM Multimedia | 6 |
| 2020 | Two birds with one stone: Transforming and generating facial images with iterative GAN
Dan Ma 0005, Bin Liu 0022, Zhao Kang 0001, Jianke Zhu, Zenglin Xu |
Neurocomputing | 5 |
| 2020 | Single-shot bidirectional pyramid networks for high-quality object detection
Xiongwei Wu, Doyen Sahoo, Daoxin Zhang, Jianke Zhu, Steven C. H. Hoi |
Neurocomputing | 4 |
| 2020 | Feature agglomeration networks for single stage face detection
Xiongwei Wu, Steven C. H. Hoi, Jianke Zhu |
Neurocomputing | 4 |
| 2020 | DeepFacade: A Deep Learning Approach to Facade Parsing With Symmetric LossabstractParsing building facades into procedural grammars plays an important role for 3D building model generation tasks, which have been long desired in computer vision. Deep learning is a promising approach to facade parsing, however, a straightforward solution by directly applying standard deep learning approaches cannot always yield the optimal results. This is primarily due to two reasons: 1) it is nontrivial to train existing semantic segmentation networks for facade parsing, e.g., Fully-Convolutional Neural Networks (FCN) which are usually weak at predicting fine-grained shapes (J. Long et al., 2015); and 2) building facades are man-made architectures with highly regularized shape priors, and the prior knowledge plays an important role in facade parsing, for which how to integrate the prior knowledge into deep neural networks remains an open problem. In this paper, we present a novel symmetric loss function that can be used in deep neural networks for end-to-end training. This novel loss is based on the assumption that most of windows and doors have a highly symmetric rectangle shape, and it penalizes all window predictions that are non-rectangles. This prior knowledge is smoothly integrated into the end-to-end training process. Quantitative evaluation demonstrates that our method has outperformed previous state-of-art methods significantly on five popular facade parsing datasets. Qualitative results have shown that our method effectively aids deep convolutional neural networks to predict more accurate, visually pleasing, and symmetric shapes. To the best of our knowledge, we are the first to incorporate symmetry constraint into end-to-end training in deep neural networks for facade parsing. Hantang Liu, Jianke Zhu, Yang Li 0041, Steven C. H. Hoi |
IEEE Trans. Multim. | 4 |
| 2019 | Robust Estimation of Similarity Transformation for Visual Object TrackingabstractMost of existing correlation filter-based tracking approaches only estimate simple axis-aligned bounding boxes, and very few of them is capable of recovering the underlying similarity transformation. To tackle this challenging problem, in this paper, we propose a new correlation filter-based tracker with a novel robust estimation of similarity transformation on the large displacements. In order to efficiently search in such a large 4-DoF space in real-time, we formulate the problem into two 2-DoF sub-problems and apply an efficient Block Coordinates Descent solver to optimize the estimation result. Specifically, we employ an efficient phase correlation scheme to deal with both scale and rotation changes simultaneously in log-polar coordinates. Moreover, a variant of correlation filter is used to predict the translational motion individually. Our experimental results demonstrate that the proposed tracker achieves very promising prediction performance compared with the state-of-the-art visual object tracking methods while still retaining the advantages of high efficiency and simplicity in conventional correlation filter-based tracking methods. Yang Li 0041, Jianke Zhu, Steven C. H. Hoi, Wenjie Song 0002, Hantang Liu |
AAAI | 2 |
| 2019 | AI Coach: Deep Human Pose Estimation and Analysis for Personalized Athletic Training AssistanceabstractRecent years have witnessed an unprecedented growing of sport videos, as different types of sports activities can be widely-observed (i.e., from professional athletics to personal fitness). Existing approaches by computer vision have predominantly focused on creating experiences of content browsing and searching by video tagging and summarization. These techniques have already enabled a wide-range of applications for sports enthusiasts, such as text-based video search, highlight generation, and so on. In this paper, we take one step further to create an AI coach system to provide personalized athletic training experiences. Especially for sports activities which the training quality largely depends on the correctness of human poses in a video sequence. As sports videos often involve grand challenges of fast movement (e.g., skiing, skating) and complex actions (e.g., gymnastics), we propose to design the system with several distinct features: (1) trajectory extraction for a single human instance by leveraging deep visual tracking, (2) human pose estimation by proposing a novel human joints relation model in spatial and temporal domains, (3) pose correction by abnormal detection and exemplar-based visual suggestions. We have collected sports training videos from 30 sports enthusiasts, namely Freestyle Skiing Aerials dataset (63 clips). We show that the proposed system can lead to a remarkably better user training experience by extensive user studies. Kai Qiu 0001, Houwen Peng, Jianlong Fu, Jianke Zhu |
ACM Multimedia | 5 |
| 2019 | AI Coach: Deep Human Pose Estimation and Analysis for Personalized Athletic Training AssistanceabstractAccurate pose analysis in sport videos is beneficial to users to improve skills. In this paper, we propose an AI coach system to provide personalized athletic training experiences for posture-wise sports activities, in which the training quality largely depends on the correctness of human poses in a video sequence. we propose to design the system with several distinct features: (1) trajectory extraction for a single human instance by leveraging deep visual tracking, (2) human pose estimation by proposing a novel human joints relation model in spatial and temporal domains,(3) pose correction by abnormal detection, performance rating and exemplar-based visual suggestions. We build an online service of this AI coach system for sports enthusiasts and collect extensive feedbacks. Comparisons with some latest popular sport apps demonstrate the effectiveness of this AI coach system to improve skills for users. Kai Qiu 0001, Houwen Peng, Jianlong Fu, Jianke Zhu |
ACM Multimedia | 5 |
| 2019 | Dynamic Saliency-Aware Regularization for Correlation Filter-Based Object TrackingabstractWith a good balance between tracking accuracy and speed, correlation filter (CF) has become one of the best object tracking frameworks, based on which many successful trackers have been developed. Recently, spatially regularized CF tracking (SRDCF) has been developed to remedy the annoying boundary effects of CF tracking, thus further boosting the tracking performance. However, SRDCF uses a fixed spatial regularization map constructed from a loose bounding box and its performance inevitably degrades when the target or background show significant variations, such as object deformation or occlusion. To address this problem, we propose a new dynamic saliency-aware regularized CF tracking (DSAR-CF) scheme. In DSAR-CF, a simple yet effective energy function, which reflects the object saliency and tracking reliability in the spatial-temporal domain, is defined to guide the online updating of the regularization weight map using an efficient level-set algorithm. Extensive experiments validate that the proposed DSAR-CF leads to better performance in terms of accuracy and speed than the original SRDCF. Wei Feng 0005, Rui-Ze Han, Qing Guo 0005, Jianke Zhu, Song Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2018 | Noise-aware co-segmentation with local and global priors
Qingqun Ning, Jianke Zhu, Mingli Song, Jiajun Bu, Chun Chen 0001 |
Neurocomputing | 3 |
| 2018 | Temporally-adjusted correlation filter-based tracking
Wenjie Song 0002, Yang Li 0041, Jianke Zhu, Chun Chen 0001 |
Neurocomputing | 3 |
| 2017 | CFNN: Correlation Filter Neural Network for Visual Object TrackingabstractAlbeit convolutional neural network (CNN) has shown promising capacity in many computer vision tasks, applying it to visual tracking is yet far from solved. Existing methods either employ a large external dataset to undertake exhaustive pre-training or suffer from less satisfactory results in terms of accuracy and robustness. To track single target in a wide range of videos, we present a novel Correlation Filter Neural Network architecture, as well as a complete visual tracking pipeline, The proposed approach is a special case of CNN, whose initialization does not need any pre-training on the external dataset. The initialization of network enjoys the merits of cyclic sampling to achieve the appealing discriminative capability, while the network updating scheme adopts advantages from back-propagation in order to capture new appearance variations. The tracking pipeline integrates both aspects well by making them complementary to each other. We validate our tracker on OTB-2013 benchmark. The proposed tracker obtains the promising results compared to most of existing representative trackers. Yang Li 0041, Jianke Zhu |
IJCAI | 3 |
| 2017 | DeepFacade: A Deep Learning Approach to Facade ParsingabstractThe parsing of building facades is a key component to the problem of 3D street scenes reconstruction, which is long desired in computer vision. In this paper, we propose a deep learning based method for segmenting a facade into semantic categories. Man-made structures often present the characteristic of symmetry. Based on this observation, we propose a symmetric regularizer for training the neural network. Our proposed method can make use of both the power of deep neural networks and the structure of man-made architectures. We also propose a method to refine the segmentation results using bounding boxes generated by the Region Proposal Network. We test our method by training a FCN-8s network with the novel loss function. Experimental results show that our method has outperformed previous state-of-the-art methods significantly on both the ECP dataset and the eTRIMS dataset. As far as we know, we are the first to employ end-to-end deep convolutional neural network on full image scale in the task of building facades parsing. Hantang Liu, Jianke Zhu, Steven C. H. Hoi |
IJCAI | 3 |
| 2017 | Image Gradient-based Joint Direct Visual Odometry for Stereo CameraabstractVisual odometry is an important research problem for computer vision and robotics. In general, the feature-based visual odometry methods heavily rely on the accurate correspondences between local salient points, while the direct approaches could make full use of whole image and perform dense 3D reconstruction simultaneously. However, the direct visual odometry usually suffers from the drawback of getting stuck at local optimum especially with large displacement, which may lead to the inferior results. To tackle this critical problem, we propose a novel scheme for stereo odometry in this paper, which is able to improve the convergence with more accurate pose. The key of our approach is a dual Jacobian optimization that is fused into a multi-scale pyramid scheme. Moreover, we introduce a gradient-based feature representation, which enjoys the merit of being robust to illumination changes. Furthermore, a joint direct odometry approach is proposed to incorporate the information from the last frame and previous keyframes. We have conducted the experimental evaluation on the challenging KITTI odometry benchmark, whose promising results show that the proposed algorithm is very effective for stereo visual odometry. Jianke Zhu |
IJCAI | 1 |
| 2017 | Sparse Online Learning of Image SimilarityabstractLearning image similarity plays a critical role in real-world multimedia information retrieval applications, especially in Content-Based Image Retrieval (CBIR) tasks, in which an accurate retrieval of visually similar objects largely relies on an effective image similarity function. Crafting a good similarity function is very challenging because visual contents of images are often represented as feature vectors in high-dimensional spaces, for example, via bag-of-words (BoW) representations, and traditional rigid similarity functions, for example, cosine similarity, are often suboptimal for CBIR tasks. In this article, we address this fundamental problem, that is, learning to optimize image similarity with sparse and high-dimensional representations from large-scale training data, and propose a novel scheme of Sparse Online Learning of Image Similarity (SOLIS). In contrast to many existing image-similarity learning algorithms that are designed to work with low-dimensional data, SOLIS is able to learn image similarity from large-scale image data in sparse and high-dimensional spaces. Our encouraging results showed that the proposed new technique achieves highly competitive accuracy as compared to the state-of-the-art approaches but enjoys significant advantages in computational efficiency, model sparsity, and retrieval scalability, making it more practical for real-world multimedia retrieval applications. Xingyu Gao 0001, Steven C. H. Hoi, Yongdong Zhang 0001, Jianshe Zhou, Ji Wan, Zhenyu Chen 0003, Jintao Li 0001, Jianke Zhu |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2017 | Scalable Image Retrieval by Sparse Product QuantizationabstractFast approximate nearest neighbor (ANN) search technique for high-dimensional feature indexing and retrieval is the crux of large-scale image retrieval. A recent promising technique is product quantization, which attempts to index high-dimensional image features by decomposing the feature space into a Cartesian product of low-dimensional subspaces and quantizing each of them separately. Despite the promising results reported, their quantization approach follows the typical hard assignment of traditional quantization methods, which may result in large quantization errors, and thus, inferior search performance. Unlike the existing approaches, in this paper, we propose a novel approach called sparse product quantization (SPQ) to encoding the high-dimensional feature vectors into sparse representation. We optimize the sparse representations of the feature vectors by minimizing their quantization errors, making the resulting representation is essentially close to the original data in practice. Experiments show that the proposed SPQ technique is not only able to compress data, but also an effective encoding technique. We obtain state-of-the-art results for ANN search on four public image datasets and the promising results of content-based image retrieval further validate the efficacy of our proposed method. Qingqun Ning, Jianke Zhu, Zhiyuan Zhong, Steven C. H. Hoi, Chun Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Realtime and robust object matching with a large number of templates
Jianke Zhu, Jun Yu 0002, Jun Cheng 0002 |
Multim. Tools Appl. | 2 |
| 2016 | Image Alignment by Online Robust PCA via Stochastic Gradient DescentabstractAligning a given set of images is usually conducted in batch mode manner, which not only requires large amount of memory but also adjusts all the previous transformations to register an input image. To address this issue, we propose a novel approach to image alignment by incorporating the geometric transformation into online robust principal component analysis (PCA). Instead of calculating the warp update using noisy input samples like the conventional methods, we suggest directly linearizing the object function by performing warp update on the recovered samples, which corresponds to an efficient inverse composition algorithm. Since the basis matrix is kept constant for a given sample, both the latent vector and warp update can be very efficiently computed. Moreover, we present two basis updating methods for robust PCA, including the closed-form solution and stochastic gradient descent scheme. We have conducted the extensive experiments on the real-world tasks of background subtraction with camera motion and visual tracking on the challenging video sequences, whose promising results demonstrate the efficacy of our presented approach. Wenjie Song 0002, Jianke Zhu, Yang Li 0041, Chun Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | A Spatial-Temporal QoS Prediction Approach for Time-aware Web Service RecommendationabstractDue to the popularity of service-oriented architectures for various distributed systems, an increasing number of Web services have been deployed all over the world. Recently, Web service recommendation became a hot research topic, one that aims to accurately predict the quality of functional satisfactory services for each end user. Generally, the performance of Web service changes over time due to variations of service status and network conditions. Instead of employing the conventional temporal models, we propose a novel spatial-temporal QoS prediction approach for time-aware Web service recommendation, where a sparse representation is employed to model QoS variations. Specifically, we make a zero-mean Laplace prior distribution assumption on the residuals of the QoS prediction, which corresponds to a Lasso regression problem. To effectively select the nearest neighbor for the sparse representation of temporal QoS values, the geo-location of web service is employed to reduce searching range while improving prediction accuracy. The extensive experimental results demonstrate that the proposed approach outperforms state-of-art methods with more than 10% improvement on the accuracy of temporal QoS prediction for time-aware Web service recommendation. Xinyu Wang 0001, Jianke Zhu, Zibin Zheng, Wenjie Song 0002, Yuanhong Shen, Michael R. Lyu |
ACM Trans. Web | 2 |
| 2015 | Reliable Patch Trackers: Robust visual tracking by exploiting reliable patchesabstractMost modern trackers typically employ a bounding box given in the first frame to track visual objects, where their tracking results are often sensitive to the initialization. In this paper, we propose a new tracking method, Reliable Patch Trackers (RPT), which attempts to identify and exploit the reliable patches that can be tracked effectively through the whole tracking process. Specifically, we present a tracking reliability metric to measure how reliably a patch can be tracked, where a probability model is proposed to estimate the distribution of reliable patches under a sequential Monte Carlo framework. As the reliable patches distributed over the image, we exploit the motion trajectories to distinguish them from the background. Therefore, the visual object can be defined as the clustering of homo-trajectory patches, where a Hough voting-like scheme is employed to estimate the target state. Encouraging experimental results on a large set of sequences showed that the proposed approach is very effective and in comparison to the state-of-the-art trackers. The full source code of our implementation will be publicly available. Yang Li 0041, Jianke Zhu, Steven C. H. Hoi |
CVPR | 2 |
| 2015 | A survey of human pose estimation: The body parts parsing based methods
Jianke Zhu, Jiajun Bu, Chun Chen 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Treelets Binary Feature Retrieval for Fast Keypoint RecognitionabstractFast keypoint recognition is essential to many vision tasks. In contrast to the classification-based approaches, we directly formulate the keypoint recognition as an image patch retrieval problem, which enjoys the merit of finding the matched keypoint and its pose simultaneously. To effectively extract the binary features from each patch surrounding the keypoint, we make use of treelets transform that can group the highly correlated data together and reduce the noise through the local analysis. Treelets is a multiresolution analysis tool, which provides an orthogonal basis to reflect the geometry of the noise-free data. To facilitate the real-world applications, we have proposed two novel approaches. One is the convolutional treelets that capture the image patch information locally and globally while reducing the computational cost. The other is the higher-order treelets that reflect the relationship between the rows and columns within image patch. An efficient sub-signature-based locality sensitive hashing scheme is employed for fast approximate nearest neighbor search in patch retrieval. Experimental evaluations on both synthetic data and the real-world Oxford dataset have shown that our proposed treelets binary feature retrieval methods outperform the state-of-the-art feature descriptors and classification-based approaches. Jianke Zhu, Chenxia Wu, Chun Chen 0001, Deng Cai 0001 |
IEEE Trans. Cybern. | 1 |
| 2015 | Fast Object Retrieval Using Direct Spatial MatchingabstractThe conventional bag-of-visual-words (BoW) model is popular for the large-scale object retrieval system but suffers from the critical drawback of ignoring spatial information . RANSAC-based methods attempt to remedy this drawback, but often require traversing all the feature matches for each hypothesis , leading to the heavy computational cost which limits the number of gallery images to be verified for each online query. We propose an efficient direct spatial matching (DSM) approach to directly estimate the scale variation using region sizes, in which all feature matches voted for estimating geometric transformation . DSM is much faster than RANSAC-based methods and exhaustive enumeration approaches. A logarithmic term frequency- inverse document frequency (log tf-idf) weighting scheme is introduced to boost the performance of the base system. We have conducted extensive experimental evaluations on four benchmark datasets for object retrieval. The proposed DSM method, together with a carefully-tailored reranking scheme, achieves the state-of-the-art results on the Oxford buildings and Paris datasets, which demonstrates the efficacy and scalability of our novel DSM technique for large scale object retrieval systems. Zhiyuan Zhong, Jianke Zhu, Steven C. H. Hoi |
IEEE Trans. Multim. | 2 |
| 2015 | Network-Aware QoS Prediction for Service Composition Using GeolocationabstractQoS-aware web service composition intends to maximize the global QoS of a composite service with local and global QoS constraints while selecting the independent candidate services from different providers. With the increasing number of candidate services emerging from the Internet, the network delays often greatly affect the performance of the composite service, which are usually difficult to be collected beforehand. One remedy is to predict them for the composition. However, there are some new issues in network delay predictions for the composition, including prediction accuracy, on-demand measures to new services and runtime overhead. In this paper, we try to tackle these critical challenges by taking advantage of the geolocations of candidate services. We first describe a network-aware service composition problem. Then, we present a novel geolocation-based NQoS prediction and reprediction approach for service composition. Furthermore, a geolocation-based service selection algorithm is presented to make use of our NQoS prediction approach for the composition. We have conducted extensive experiments on the real-world data set collected from PlanetLab. Comparative experimental results demonstrate that our approach improves the prediction accuracy and predictability of the NQoS and reduces the runtime overheads in predicting the composition. Xinyu Wang 0001, Jianke Zhu, Yuanhong Shen |
IEEE Trans. Serv. Comput. | 2 |
| 2014 | Deep Learning for Content-Based Image Retrieval: A Comprehensive StudyabstractLearning effective feature representations and similarity measures are crucial to the retrieval performance of a content-based image retrieval (CBIR) system. Despite extensive research efforts for decades, it remains one of the most challenging open problems that considerably hinders the successes of real-world CBIR systems. The key challenge has been attributed to the well-known ``semantic gap'' issue that exists between low-level image pixels captured by machines and high-level semantic concepts perceived by human. Among various techniques, machine learning has been actively investigated as a possible direction to bridge the semantic gap in the long term. Inspired by recent successes of deep learning techniques for computer vision and other applications, in this paper, we attempt to address an open problem: if deep learning is a hope for bridging the semantic gap in CBIR and how much improvements in CBIR tasks can be achieved by exploring the state-of-the-art deep learning techniques for learning feature representations and similarity measures. Specifically, we investigate a framework of deep learning with application to CBIR tasks with an extensive set of empirical studies by examining a state-of-the-art deep learning method (Convolutional Neural Networks) for CBIR tasks under varied settings. From our empirical studies, we find some encouraging results and summarize some important insights for future research. Ji Wan, Steven C. H. Hoi, Jianke Zhu, Yongdong Zhang 0001, Jintao Li 0001 |
ACM Multimedia | 5 |
| 2014 | Object cosegmentation by nonrigid mapping
Jianke Zhu, Jiajun Bu, Chun Chen 0001 |
Neurocomputing | 2 |
| 2014 | Retrieval-Based Face Annotation by Weak Label Regularized Local Coordinate CodingabstractAuto face annotation, which aims to detect human faces from a facial image and assign them proper human names, is a fundamental research problem and beneficial to many real-world applications. In this work, we address this problem by investigating a retrieval-based annotation scheme of mining massive web facial images that are freely available over the Internet. In particular, given a facial image, we first retrieve the top $(n)$ similar instances from a large-scale web facial image database using content-based image retrieval techniques, and then use their labels for auto annotation. Such a scheme has two major challenges: 1) how to retrieve the similar facial images that truly match the query, and 2) how to exploit the noisy labels of the top similar facial images, which may be incorrect or incomplete due to the nature of web images. In this paper, we propose an effective Weak Label Regularized Local Coordinate Coding (WLRLCC) technique, which exploits the principle of local coordinate coding by learning sparse features, and employs the idea of graph-based weak label regularization to enhance the weak labels of the similar facial images. An efficient optimization algorithm is proposed to solve the WLRLCC problem. Moreover, an effective sparse reconstruction scheme is developed to perform the face annotation task. We conduct extensive empirical studies on several web facial image databases to evaluate the proposed WLRLCC algorithm from different aspects. The experimental results validate its efficacy. We share the two constructed databases "WDB" (714,454 images of 6,025 people) and "ADB" (126,070 images of 1,200 people) with the public. To further improve the efficiency and scalability, we also propose an offline approximation scheme (AWLRLCC) which generally maintains comparable results but significantly reduces the annotation time. Steven C. H. Hoi, Ying He 0001, Jianke Zhu, Tao Mei 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Feature Correlation Hypergraph: Exploiting High-order Potentials for Multimodal RecognitionabstractIn computer vision and multimedia analysis, it is common to use multiple features (or multimodal features) to represent an object. For example, to well characterize a natural scene image, we typically extract a set of visual features to represent its color, texture, and shape. However, it is challenging to integrate multimodal features optimally. Since they are usually high-order correlated, e.g., the histogram of gradient (HOG), bag of scale invariant feature transform descriptors, and wavelets are closely related because they collaboratively reflect the image texture. Nevertheless, the existing algorithms fail to capture the high-order correlation among multimodal features. To solve this problem, we present a new multimodal feature integration framework. Particularly, we first define a new measure to capture the high-order correlation among the multimodal features, which can be deemed as a direct extension of the previous binary correlation. Therefore, we construct a feature correlation hypergraph (FCH) to model the high-order relations among multimodal features. Finally, a clustering algorithm is performed on FCH to group the original multimodal features into a set of partitions. Moreover, a multiclass boosting strategy is developed to obtain a strong classifier by combining the weak classifiers learned from each partition. The experimental results on seven popular datasets show the effectiveness of our approach. Yinfu Feng, Jianke Zhu, Deng Cai 0001 |
IEEE Trans. Cybern. | 5 |
| 2014 | Mining Weakly Labeled Web Facial Images for Search-Based Face AnnotationabstractThis paper investigates a framework of search-based face annotation (SBFA) by mining weakly labeled facial images that are freely available on the World Wide Web (WWW). One challenging problem for search-based face annotation scheme is how to effectively perform annotation by exploiting the list of most similar facial images and their weak labels that are often noisy and incomplete. To tackle this problem, we propose an effective unsupervised label refinement (ULR) approach for refining the labels of web facial images using machine learning techniques. We formulate the learning problem as a convex optimization and develop effective optimization algorithms to solve the large-scale learning task efficiently. To further speed up the proposed scheme, we also propose a clustering-based approximation algorithm which can improve the scalability considerably. We have conducted an extensive set of empirical studies on a large-scale web facial image testbed, in which encouraging results showed that the proposed ULR algorithms can significantly boost the performance of the promising SBFA scheme. Steven C. H. Hoi, Ying He 0001, Jianke Zhu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Sparse Poisson coding for high dimensional document clusteringabstractDocument clustering plays an important role in large scale textual data analysis, which generally faces with great challenge of the high dimensional textual data. One remedy is to learn the high-level sparse representation by the sparse coding techniques. In contrast to traditional Gaussian noise-based sparse coding methods, in this paper, we employ a Poisson distribution model to represent the word-count frequency feature of a text for sparse coding. Moreover, a novel sparse-constrained Poisson regression algorithm is proposed to solve the induced optimization problem. Different from previous Poisson regression with the family of ℓ1-regularization to enhance the sparse solution, we introduce a sparsity ratio measure which make use of both ℓ1-norm and ℓ2-norm on the learned weight. An important advantage of the sparsity ratio is that it bounded in the range of 0 and 1. This makes it easy to set for practical applications. To further make the algorithm trackable for the high dimensional textual data, a projected gradient descent algorithm is proposed to solve the regression problem. Extensive experiments have been conducted to show that our proposed approach can achieve effective representation for document clustering compared with state-of-the-art regression methods. Chenxia Wu, Haiqin Yang, Jianke Zhu, Jiemi Zhang, Irwin King, Michael R. Lyu |
IEEE BigData | 3 |
| 2013 | Geographic Location-Based Network-Aware QoS Prediction for Service CompositionabstractQoS-aware service composition intends to maximize the global QoS of a composite service while selecting candidate services from different providers with local and global QoS constraints. With more and more candidate services emerging from all over the world, the network delays often greatly impact the performance of the composite service, which are usually not easy to be collected before the composition. One remedy is to predict them for the composition. However, new issues occur in predicting network delay for the composition, including prediction accuracy and on-demand measures to new services, which affect the performance of network-aware composite services. To solve these critical challenges, in this paper, we take advantage of the geographic location information of candidate services. We propose a network-aware QoS (NQoS) model for the composite service. Based on that, we present a novel geographic location-based NQoS prediction approach before composition, and a NQoS re-prediction approach during the execution of the composite service. Extensive experiments are conducted on the real-world dataset collected from PlanetLab. Comparative experiment results reveal our approach facilitates to improve the prediction accuracy and predictability of the NQoS values, and increase global NQoS of the composite service while ensuring its reliability constraints. Yuanhong Shen, Jianke Zhu, Xinyu Wang 0001, Xiaohu Yang 0001, Bo Zhou 0010 |
ICWS | 2 |
| 2013 | Bilevel Visual Words Coding for Image Classification
Jiemi Zhang, Chenxia Wu, Deng Cai 0001, Jianke Zhu |
IJCAI | 4 |
| 2013 | Learning to name faces: a multimodal learning scheme for search-based face annotationabstractAutomated face annotation aims to automatically detect human faces from a photo and further name the faces with the corresponding human names. In this paper, we tackle this open problem by investigating a search-based face annotation (SBFA) paradigm for mining large amounts of web facial images freely available on the WWW. Given a query facial image for annotation, the idea of SBFA is to first search for top-n similar facial images from a web facial image database and then exploit these top-ranked similar facial images and their weak labels for naming the query facial image. To fully mine those information, this paper proposes a novel framework of Learning to Name Faces (L2NF) -- a unified multimodal learning approach for search-based face annotation, which consists of the following major components: (i) we enhance the weak labels of top-ranked similar images by exploiting the "label smoothness" assumption; (ii) we construct the multimodal representations of a facial image by extracting different types of features; (iii) we optimize the distance measure for each type of features using distance metric learning techniques; and finally (iv) we learn the optimal combination of multiple modalities for annotation through a learning to rank scheme. We conduct a set of extensive empirical studies on two real-world facial image databases, in which encouraging results show that the proposed algorithms significantly boost the naming accuracy of search-based face annotation task. Steven C. H. Hoi, Jianke Zhu, Ying He 0001, Chunyan Miao |
SIGIR | 4 |
| 2013 | Hypergraph-based multi-example ranking with sparse representation for transductive learning image retrieval
Jianke Zhu |
Neurocomputing | 2 |
| 2013 | Semi-Supervised Nonlinear Hashing Using Bootstrap Sequential Projection LearningabstractIn this paper, we study the effective semi-supervised hashing method under the framework of regularized learning-based hashing. A nonlinear hash function is introduced to capture the underlying relationship among data points. Thus, the dimensionality of the matrix for computation is not only independent from the dimensionality of the original data space but also much smaller than the one using linear hash function. To effectively deal with the error accumulated during converting the real-value embeddings into the binary code after relaxation, we propose a semi-supervised nonlinear hashing algorithm using bootstrap sequential projection learning which effectively corrects the errors by taking into account of all the previous learned bits holistically without incurring the extra computational overhead. Experimental results on the six benchmark data sets demonstrate that the presented method outperforms the state-of-the-art hashing algorithms at a large margin. Chenxia Wu, Jianke Zhu, Deng Cai 0001, Chun Chen 0001, Jiajun Bu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | A Convolutional Treelets Binary Feature Approach to Fast Keypoint Recognition
Chenxia Wu, Jianke Zhu, Jiemi Zhang, Chun Chen 0001, Deng Cai 0001 |
ECCV (5) | 2 |
| 2012 | Realtime object matching with robust dominant orientation templates
Jianke Zhu, Mingli Song, Yinting Wang |
ICPR | 2 |
| 2012 | Unsupervised face-name association via commute distanceabstractRecently, the task of unsupervised face-name association has received a considerable interests in multimedia and information retrieval communities. It is quite different with the generic facial image annotation problem because of its unsupervised and ambiguous assignment properties. Specifically, the task of face-name association should obey the following three constraints: (1) a face can only be assigned to a name appearing in its associated caption or to null; (2) a name can be assigned to at most one face; and (3) a face can be assigned to at most one name. Many conventional methods have been proposed to tackle this task while suffering from some common problems, eg, many of them are computational expensive and hard to make the null assignment decision. In this paper, we design a novel framework named face-name association via commute distance (FACD), which judges face-name and face-null assignments under a unified framework via commute distance (CD) algorithm. Then, to further speed up the on-line processing, we propose a novel anchor-based commute distance (ACD) algorithm whose main idea is using the anchor point representation structure to accelerate the eigen-decomposition of the adjacency matrix of a graph. Systematic experiment results on a large scale and real world image-caption database with a total of 194,046 detected faces and 244,725 names show that our proposed approach outperforms many state-of-the-art methods in performance. Our framework is appropriate for a large scale and real-time system. Jiajun Bu, Bin Xu 0005, Chenxia Wu, Chun Chen 0001, Jianke Zhu, Deng Cai 0001, Xiaofei He 0001 |
ACM Multimedia | 5 |
| 2012 | Locally discriminative topic modeling
Jiajun Bu, Chun Chen 0001, Jianke Zhu, Lijun Zhang 0005, Haifeng Liu 0001, Can Wang 0001, Deng Cai 0001 |
Pattern Recognit. | 4 |
| 2012 | Learning Bregman Distance Functions for Semi-Supervised ClusteringabstractLearning distance functions with side information plays a key role in many data mining applications. Conventional distance metric learning approaches often assume that the target distance function is represented in some form of Mahalanobis distance. These approaches usually work well when data are in low dimensionality, but often become computationally expensive or even infeasible when handling high-dimensional data. In this paper, we propose a novel scheme of learning nonlinear distance functions with side information. It aims to learn a Bregman distance function using a nonparametric approach that is similar to Support Vector Machines. We emphasize that the proposed scheme is more general than the conventional approach for distance metric learning, and is able to handle high-dimensional data efficiently. We verify the efficacy of the proposed distance learning method with extensive experiments on semi-supervised clustering. The comparison with state-of-the-art approaches for learning distance functions with side information reveals clear advantages of the proposed technique. Lei Wu 0017, Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Nenghai Yu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2011 | Social Recommendation Using Low-Rank Semidefinite ProgramabstractThe most critical challenge for the recommendation system is to achieve the high prediction quality on the large scale sparse data contributed by the users. In this paper, we present a novel approach to the social recommendation problem, which takes the advantage of the graph Laplacian regularization to capture the underlying social relationship among the users. Differently from the previous approaches, that are based on the conventional gradient descent optimization, we formulate the presented graph Laplacian regularized social recommendation problem into a low-rank semidefinite program, which is able to be efficiently solved by the quasi-Newton algorithm. We have conducted the empirical evaluation on a large scale dataset of high sparsity, the promising experimental results show that our method is very effective and efficient for the social recommendation task. Jianke Zhu, Chun Chen 0001, Jiajun Bu |
AAAI | 1 |
| 2011 | Retrieval-based face annotation by weak label regularized local coordinate codingabstractThis paper investigates a retrieval-based annotation paradigm of mining web facial images for automated face annotation. In general, there are two key challenges for such an annotation paradigm. The first challenge is how to efficiently retrieve a short list of most similar facial images from facial image databases, and the second challenge is how to effectively perform annotation by exploiting these similar facial images and their weak labels which are often noisy and incomplete. In this paper, we mainly focus on tackling the second challenge of the retrieval-based face annotation paradigm. In particular, we propose an effective Weak Label Regularized Local Coordinate Coding (WLRLCC) technique, which exploits the local coordinate coding principle in learning sparse features, and at the same time employs the graph-based weak label regularization principle to enhance the weak labels of the short list of similar facial images. We present an efficient optimization algorithm to solve the WLRLCC task, and develop an effective sparse reconstruction scheme to perform the final face name annotation. We conduct a set of extensive empirical studies on a large-scale facial image database withatotal of 6,000 persons and over 600,000 web facial images, in which encouraging results show that the proposed WLRLCC algorithm significantly boosts the performance of the regular retrieval-based face annotation approaches. Steven C. H. Hoi, Ying He 0001, Jianke Zhu |
ACM Multimedia | 4 |
| 2011 | Distance metric learning from uncertain side information for automated photo taggingabstractAutomated photo tagging is an important technique for many intelligent multimedia information systems, for example, smart photo management system and intelligent digital media library. To attack the challenge, several machine learning techniques have been developed and applied for automated photo tagging. For example, supervised learning techniques have been applied to automated photo tagging by training statistical classifiers from a collection of manually labeled examples. Although the existing approaches work well for small testbeds with relatively small number of annotation words, due to the long-standing challenge of object recognition, they often perform poorly in large-scale problems. Another limitation of the existing approaches is that they require a set of high-quality labeled data, which is not only expensive to collect but also time consuming. In this article, we investigate a social image based annotation scheme by exploiting implicit side information that is available for a large number of social photos from the social web sites. The key challenge of our intelligent annotation scheme is how to learn an effective distance metric based on implicit side information (visual or textual) of social photos. To this end, we present a novel “Probabilistic Distance Metric Learning” (PDML) framework, which can learn optimized metrics by effectively exploiting the implicit side information vastly available on the social web. We apply the proposed technique to photo annotation tasks based on a large social image testbed with over 1 million tagged photos crawled from a social photo sharing portal. Encouraging results show that the proposed technique is effective and promising for social photo based annotation tasks. Lei Wu 0017, Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Nenghai Yu |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | Near-duplicate keyframe retrieval by semi-supervised learning and nonrigid image matchingabstractNear-duplicate keyframe (NDK) retrieval techniques are critical to many real-world multimedia applications. Over the last few years, we have witnessed a surge of attention on studying near-duplicate image/keyframe retrieval in the multimedia community. To facilitate an effective approach to NDK retrieval on large-scale data, we suggest an effective Multi-Level Ranking (MLR) scheme that effectively retrieves NDKs in a coarse-to-fine manner. One key stage of the MLR ranking scheme is how to learn an effective ranking function with extremely small training examples in a near-duplicate detection task. To attack this challenge, we employ a semi-supervised learning method, semi-supervised support vector machines, which is able to significantly improve the retrieval performance by exploiting unlabeled data. Another key stage of the MLR scheme is to perform a fine matching among a subset of keyframe candidates retrieved from the previous coarse ranking stage. In contrast to previous approaches based on either simple heuristics or rigid matching models, we propose a novel Nonrigid Image Matching (NIM) approach to tackle near-duplicate keyframe retrieval from real-world video corpora in order to conduct an effective fine matching. Compared with the conventional methods, the proposed NIM approach can recover explicit mapping between two near-duplicate images with a few deformation parameters and find out the correct correspondences from noisy data simultaneously. To evaluate the effectiveness of our proposed approach, we performed extensive experiments on two benchmark testbeds extracted from the TRECVID2003 and TRECVID2004 corpora. The promising results indicate that our proposed method is more effective than other state-of-the-art approaches for near-duplicate keyframe retrieval. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu, Shuicheng Yan |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2010 | Bridging the Semantic Gap Between Image Contents and TagsabstractWith the exponential growth of Web 2.0 applications, tags have been used extensively to describe the image contents on the Web. Due to the noisy and sparse nature in the human generated tags, how to understand and utilize these tags for image retrieval tasks has become an emerging research direction. As the low-level visual features can provide fruitful information, they are employed to improve the image retrieval results. However, it is challenging to bridge the semantic gap between image contents and tags. To attack this critical problem, we propose a unified framework in this paper which stems from a two-level data fusions between the image contents and tags: 1) A unified graph is built to fuse the visual feature-based image similarity graph with the image-tag bipartite graph; 2) A novel random walk model is then proposed, which utilizes a fusion parameter to balance the influences between the image contents and tags. Furthermore, the presented framework not only can naturally incorporate the pseudo relevance feedback process, but also it can be directly applied to applications such as content-based image retrieval, text-based image retrieval, and image annotation. Experimental analysis on a large Flickr dataset shows the effectiveness and efficiency of our proposed framework. Hao Ma 0001, Jianke Zhu, Michael R. Lyu, Irwin King |
IEEE Trans. Multim. | 2 |
| 2009 | Nonrigid shape recovery by Gaussian process regressionabstractMost state-of-the-art nonrigid shape recovery methods usually use explicit deformable mesh models to regularize surface deformation and constrain the search space. These triangulated mesh models heavily relying on the quadratic regularization term are difficult to accurately capture large deformations, such as severe bending. In this paper, we propose a novel Gaussian process regression approach to the nonrigid shape recovery problem, which does not require to involve a predefined triangulated mesh model. By taking advantage of our novel Gaussian process regression formulation together with a robust coarse-to-fine optimization scheme, the proposed method is fully automatic and is able to handle large deformations and outliers. We conducted a set of extensive experiments for performance evaluation in various environments. Encouraging experimental results show that our proposed approach is both effective and robust to nonrigid shape recovery with large deformations. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu |
CVPR | 1 |
| 2009 | Unsupervised face alignment by robust nonrigid mappingabstractWe propose a novel approach to unsupervised facial image alignment. Differently from previous approaches, that are confined to affine transformations on either the entire face or separate patches, we extract a nonrigid mapping between facial images. Based on a regularized face model, we frame unsupervised face alignment into the Lucas-Kanade image registration approach. We propose a robust optimization scheme to handle appearance variations. The method is fully automatic and can cope with pose variations and expressions, all in an unsupervised manner. Experiments on a large set of images showed that the approach is effective. Jianke Zhu, Luc Van Gool, Steven C. H. Hoi |
ICCV | 1 |
| 2009 | Distance metric learning from uncertain side information with application to automated photo taggingabstractAutomated photo tagging is an important technique for many intelligent multimedia information systems, e.g. smart photo management system and intelligent digital media library. To attack the challenge, several machine learning techniques have been developed and applied for automated photo tagging. For example, supervised learning techniques have been applied to automated photo tagging by training statistical classifiers from a collection of manually labeled examples. Although the existing approaches work well for small testbeds with relatively small number of annotation words, due to the long-standing challenge of object recognition, they often perform poorly in large-scale problems. Another limitation of the existing approaches is that they require a set of high quality labeled data, which is not only expensive to collect but also time consuming. In this paper, we investigate a social image based annotation scheme by exploiting implicit side information that is available for a large number of social photos from the social web sites. The key challenge of our intelligent annotation scheme is how to learn an effective distance metric based on implicit side information (visual or textual) of social photos. To this end, we present a novel “Probabilistic Distance Metric Learning ” (PDML) framework, which can learn optimized metrics by effectively exploiting the implicit side information vastly available on the social web. We apply Lei Wu 0017, Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Nenghai Yu |
ACM Multimedia | 4 |
| 2009 | Learning Bregman Distance Functions and Its Application for Semi-Supervised ClusteringabstractLearning distance functions with side information plays a key role in many machine learning and data mining applications. Conventional approaches often assume a Mahalanobis distance function. These approaches are limited in two aspects: (i) they are computationally expensive (even infeasible) for high dimensional data because the size of the metric is in the square of dimensionality; (ii) they assume a fixed metric for the entire input space and therefore are unable to handle heterogeneous data. In this paper, we propose a novel scheme that learns nonlinear Bregman distance functions from side information using a non-parametric approach that is similar to support vector machines. The proposed scheme avoids the assumption of fixed metric because its local distance metric is implicitly derived from the Hessian matrix of a convex function that is used to generate the Bregman distance function. We present an efficient learning algorithm for the proposed scheme for distance function learning. The extensive experiments with semi-supervised clustering show the proposed technique (i) outperforms the state-of-the-art approaches for distance function learning, and (ii) is computationally efficient for high dimensional data. Lei Wu 0017, Rong Jin 0001, Steven C. H. Hoi, Jianke Zhu, Nenghai Yu |
NIPS | 4 |
| 2009 | Adaptive Regularization for Transductive Support Vector MachineabstractWe discuss the framework of Transductive Support Vector Machine (TSVM) from the perspective of the regularization strength induced by the unlabeled data. In this framework, SVM and TSVM can be regarded as a learning machine without regularization and one with full regularization from the unlabeled data, respectively. Therefore, to supplement this framework of the regularization strength, it is necessary to introduce data-dependant partial regularization. To this end, we reformulate TSVM into a form with controllable regularization strength, which includes SVM and TSVM as special cases. Furthermore, we introduce a method of adaptive regularization that is data dependant and is based on the smoothness assumption. Experiments on a set of benchmark data sets indicate the promising results of the proposed work compared with state-of-the-art TSVM algorithms. Zenglin Xu, Rong Jin 0001, Jianke Zhu, Irwin King, Michael R. Lyu, Zhirong Yang |
NIPS | 3 |
| 2009 | A novel kernel-based maximum a posteriori classification method
Zenglin Xu, Kaizhu Huang, Jianke Zhu, Irwin King, Michael R. Lyu |
Neural Networks | 3 |
| 2009 | A Fast 2D Shape Recovery Approach by Fusing Features and AppearanceabstractIn this paper, we present a fusion approach to solve the nonrigid shape recovery problem, which takes advantage of both the appearance information and the local features. We have two major contributions. First, we propose a novel progressive finite Newton optimization scheme for the feature-based nonrigid surface detection problem, which is reduced to only solving a set of linear equations. The key is to formulate the nonrigid surface detection as an unconstrained quadratic optimization problem that has a closed-form solution for a given set of observations. Second, we propose a deformable Lucas-Kanade algorithm that triangulates the template image into small patches and constrains the deformation through the second-order derivatives of the mesh vertices. We formulate it into a sparse regularized least squares problem, which is able to reduce the computational cost and the memory requirement. The inverse compositional algorithm is applied to efficiently solve the optimization problem. We have conducted extensive experiments for performance evaluation on various environments, whose promising results show that the proposed algorithm is both efficient and effective. Jianke Zhu, Michael R. Lyu, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Semisupervised SVM batch mode active learning with applications to image retrievalabstractSupport vector machine (SVM) active learning is one popular and successful technique for relevance feedback in content-based image retrieval (CBIR). Despite the success, conventional SVM active learning has two main drawbacks. First, the performance of SVM is usually limited by the number of labeled examples. It often suffers a poor performance for the small-sized labeled examples, which is the case in relevance feedback. Second, conventional approaches do not take into account the redundancy among examples, and could select multiple examples that are similar (or even identical). In this work, we propose a novel scheme for explicitly addressing the drawbacks. It first learns a kernel function from a mixture of labeled and unlabeled data, and therefore alleviates the problem of small-sized training data. The kernel will then be used for a batch mode active learning method to identify the most informative and diverse examples via a min-max framework. Two novel algorithms are proposed to solve the related combinatorial optimization: the first approach approximates the problem into a quadratic program, and the second solves the combinatorial optimization approximately by a greedy algorithm that exploits the merits of submodular functions. Extensive experiments with image retrieval using both natural photo images and medical images show that the proposed algorithms are significantly more effective than the state-of-the-art approaches. A demo is available at http://msm.cais.ntu.edu.sg/LSCBIR/. Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Michael R. Lyu |
ACM Trans. Inf. Syst. | 3 |
| 2008 | Semi-supervised SVM batch mode active learning for image retrievalabstractActive learning has been shown as a key technique for improving content-based image retrieval (CBIR) performance. Among various methods, support vector machine (SVM) active learning is popular for its application to relevance feedback in CBIR. However, the regular SVM active learning has two main drawbacks when used for relevance feedback. First, SVM often suffers from learning with a small number of labeled examples, which is the case in relevance feedback. Second, SVM active learning usually does not take into account the redundancy among examples, and therefore could select multiple examples in relevance feedback that are similar (or even identical) to each other. In this paper, we propose a novel scheme that exploits both semi-supervised kernel learning and batch mode active learning for relevance feedback in CBIR. In particular, a kernel function is first learned from a mixture of labeled and unlabeled examples. The kernel will then be used to effectively identify the informative and diverse examples for active learning via a min-max framework. An empirical study with relevance feedback of CBIR showed that the proposed scheme is significantly more effective than other state-of-the-art approaches. Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Michael R. Lyu |
CVPR | 3 |
| 2008 | An Effective Approach to 3D Deformable Surface Tracking
Jianke Zhu, Steven C. H. Hoi, Zenglin Xu, Michael R. Lyu |
ECCV (3) | 1 |
| 2008 | Near-duplicate keyframe retrieval by nonrigid image matchingabstractNear-duplicate image retrieval plays an important role in many real-world multimedia applications. Most previous approaches have some limitations. For example, conventional appearance-based methods may suffer from the illumination variations and occlusion issue, and local feature correspondence-based methods often do not consider local deformations and the spatial coherence between two point sets. In this paper, we propose a novel and effective Nonrigid Image Matching (NIM) approach to tackle the task of near-duplicate keyframe retrieval from real-world video corpora. In contrast to previous approaches, the NIM technique can recover an explicit mapping between two nearduplicate images with a few deformation parameters and find out the correct correspondences from noisy data effectively. To make our technique applicable to large-scale applications, we suggest an effective multi-level ranking scheme that filters out the irrelevant results in a coarse-to-fine manner. In our ranking scheme, to overcome the extremely small training size challenge, we employ a semi-supervised learning method for improving the performance using unlabeled data. To evaluate the effectiveness of our solution, we have conducted extensive experiments on two benchmark testbeds extracted from the TRECVID2003 and TRECVID2004 corpora. The promising results show that our proposed method is more effective than other state-of-the-art approaches for near-duplicate keyframe retrieval. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu, Shuicheng Yan |
ACM Multimedia | 1 |
| 2008 | Face Annotation Using Transductive Kernel Fisher DiscriminantabstractFace annotation in images and videos enjoys many potential applications in multimedia information retrieval. Face annotation usually requires many training data labeled by hand in order to build effective classifiers. This is particularly challenging when annotating faces on large-scale collections of media data, in which huge labeling efforts would be very expensive. As a result, traditional supervised face annotation methods often suffer from insufficient training data. To attack this challenge, in this paper, we propose a novel Transductive Kernel Fisher Discriminant (TKFD) scheme for face annotation, which outperforms traditional supervised annotation methods with few training data. The main idea of our approach is to solve the Fisher's discriminant using deformed kernels incorporating the information of both labeled and unlabeled data. To evaluate the effectiveness of our method, we have conducted extensive experiments on three types of multimedia testbeds: the FRGC benchmark face dataset, the Yahoo! web image collection, and the TRECVID video data collection. The experimental results show that our TKFD algorithm is more effective than traditional supervised approaches, especially when there are very few training data. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu |
IEEE Trans. Multim. | 1 |
| 2008 | Robust Regularized Kernel RegressionabstractRobust regression techniques are critical to fitting data with noise in real-world applications. Most previous work of robust kernel regression is usually formulated into a dual form, which is then solved by some quadratic program solver consequently. In this correspondence, we propose a new formulation for robust regularized kernel regression under the theoretical framework of regularization networks and then tackle the optimization problem directly in the primal. We show that the primal and dual approaches are equivalent to achieving similar regression performance, but the primal formulation is more efficient and easier to be implemented than the dual one. Different from previous work, our approach also optimizes the bias term. In addition, we show that the proposed solution can be easily extended to other noise-reliable loss function, including the Huber- epsilon insensitive loss function. Finally, we conduct a set of experiments on both artificial and real data sets, in which promising results show that the proposed method is effective and more efficient than traditional approaches. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | A Multi-Scale Tikhonov Regularization Scheme for Implicit Surface ModellingabstractKernel machines have recently been considered as a promising solution for implicit surface modelling. A key challenge of machine learning solutions is how to fit implicit shape models from large-scale sets of point cloud samples efficiently. In this paper, we propose a fast solution for approximating implicit surfaces based on a multi-scale Tikhonov regularization scheme. The optimization of our scheme is formulated into a sparse linear equation system, which can be efficiently solved by factorization methods. Different from traditional approaches, our scheme does not employ auxiliary off-surface points, which not only saves the computational cost but also avoids the problem of injected noise. To further speedup our solution, we present a multi-scale surface fitting algorithm of coarse to fine modelling. We conduct comprehensive experiments to evaluate the performance of our solution on a number of datasets of different scales. The promising results show that our suggested scheme is considerably more efficient than the state-of-the-art approach. Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu |
CVPR | 1 |
| 2007 | Progressive Finite Newton Approach To Real-time Nonrigid Surface DetectionabstractDetecting nonrigid surfaces is an interesting research problem for computer vision and image analysis. One important challenge of nonrigid surface detection is how to register a nonrigid surface mesh having a large number of free deformation parameters. This is particularly significant for detecting nonrigid surfaces from noisy observations. Nonrigid surface detection is usually regarded as a robust parameter estimation problem, which is typically solved iteratively from a good initialization in order to avoid local minima. In this paper, we propose a novel progressive finite Newton optimization scheme for the non-rigid surface detection problem, which is reduced to only solving a set of linear equations. The key of our approach is to formulate the nonrigid surface detection as an unconstrained quadratic optimization problem which has a closed-form solution for a given set of observations. Moreover, we employ a progressive active-set selection scheme, which takes advantage of the rank information of detected correspondences. We have conducted extensive experiments for performance evaluation on various environments, whose promising results show that the proposed algorithm is more efficient and effective than the existing iterative methods. Jianke Zhu, Michael R. Lyu |
CVPR | 1 |
| 2007 | Kernel Maximum a Posteriori Classification with Error Bound Analysis
Zenglin Xu, Kaizhu Huang, Jianke Zhu, Irwin King, Michael R. Lyu |
ICONIP (1) | 3 |
| 2007 | Two-stage Multi-class AdaBoost for Facial Expression RecognitionabstractAlthough AdaBoost has achieved great success, it still suffers from following problems: (1) the training process could be unmanageable when the number of features is extremely large; (2) the same weak classifier may be learned multiple times from a weak classifier pool, which does not provide additional information for updating the model; (3) there is an imbalance between the amount of the positive samples and that of the negative samples for multi-class classification problems. In this paper, we propose a two-stage AdaBoost learning framework to select and fuse the discriminative feature effectively. Moreover, an improved AdaBoost algorithm is developed to select weak classifiers. Instead of boosting in the original feature space, whose dimensionality is usually very high, multiple feature subspaces with lower dimensionality are generated. In the first stage, boosting is carried out in each subspace. Then the trained classifiers are further combined with simple fusion method in the second stage. Experimental results on facial expression recognition data demonstrate that our proposed algorithms not only reduce the computational cost for training, but also achieve comparable classification performance. Hongbo Deng, Jianke Zhu, Michael R. Lyu, Irwin King |
IJCNN | 2 |
| 2007 | Maximum Margin based Semi-supervised Spectral Kernel LearningabstractSemi-supervised kernel learning is attracting increasing research interests recently. It works by learning an embedding of data from the input space to a Hilbert space using both labeled data and unlabeled data, and then searching for relations among the embedded data points. One of the most well-known semi-supervised kernel learning approaches is the spectral kernel learning methodology which usually tunes the spectral empirically or through optimizing some generalized performance measures. However, the kernel designing process does not involve the bias of a kernel-based learning algorithm, the deduced kernel matrix cannot necessarily facilitate a specific learning algorithm. To supplement the spectral kernel learning methods, this paper proposes a novel approach, which not only learns a kernel matrix by maximizing another generalized performance measure, the margin between two classes of data, but also leads directly to a convex optimization method for learning the margin parameters in support vector machines. Moreover, experimental results demonstrate that our proposed spectral kernel learning method achieves promising results against other spectral kernel learning methods. Zenglin Xu, Jianke Zhu, Michael R. Lyu, Irwin King |
IJCNN | 2 |
| 2007 | Efficient Convex Relaxation for Transductive Support Vector MachineabstractWe consider the problem of Support Vector Machine transduction, which involves a combinatorial problem with exponential computational complexity in the number of unlabeled examples. Although several studies are devoted to Transductive SVM, they suffer either from the high computation complexity or from the solutions of local optimum. To address this problem, we propose solving Transductive SVM via a convex relaxation, which converts the NP-hard problem to a semi-definite programming. Compared with the other SDP relaxation for Transductive SVM, the proposed algorithm is computationally more efficient with the number of free parameters reduced from O(n2) to O(n) where n is the number of examples. Empirical study with several benchmark data sets shows the promising performance of the proposed algorithm in comparison with other state-of-the-art implementations of Transductive SVM. Zenglin Xu, Rong Jin 0001, Jianke Zhu, Irwin King, Michael R. Lyu |
NIPS | 3 |
| 2006 | Real-Time Non-rigid Shape Recovery Via Active Appearance Models for Augmented Reality
Jianke Zhu, Steven C. H. Hoi, Michael R. Lyu |
ECCV (1) | 1 |
| 2006 | Batch mode active learning and its application to medical image classificationabstractThe goal of active learning is to select the most informative examples for manual labeling. Most of the previous studies in active learning have focused on selecting a single unlabeled example in each iteration. This could be inefficient since the classification model has to be retrained for every labeled example. In this paper, we present a framework for "batch mode active learning" that applies the Fisher information matrix to select a number of informative examples simultaneously. The key computational challenge is how to efficiently identify the subset of unlabeled examples that can result in the largest reduction in the Fisher information. To resolve this challenge, we propose an efficient greedy algorithm that is based on the property of submodular functions. Our empirical studies with five UCI datasets and one real-world medical image classification show that the proposed batch mode active learning algorithm is more effective than the state-of-the-art algorithms for active learning. Steven C. H. Hoi, Rong Jin 0001, Jianke Zhu, Michael R. Lyu |
ICML | 3 |
| 2004 | Gabor wavelets transform and extended nearest feature space classifier for face recognitionabstractThis paper proposes new hybrid approaches for face recognition. Gabor wavelets representation of face images is an effective approach for both facial action recognition and face identification. The performance of dimensionality reduction and linear discriminant analysis on the downsampled Gabor waveletfaces can increase the discriminant ability. The nearest feature space is extended to various similarity measures. In our experiments, proposed Gabor waveletfaces combined with extended nearest feature space classifier shows very good performance, which can achieve 99 % maximum correct recognition rate on ORL data set without any preprocessing step. Jianke Zhu, Mang I Vai, Peng Un Mak |
ICIG | 1 |