VLDB 2026 Research / reviewers in the wild / expert
Jianzhuang Liu
dblp:l/JianzhuangLiu
· DBLP profile ↗
221ranked-venue papers
9as first author
84since 2021 · last 2026
0000-0002-7960-9382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 165 · 2 first-author · 59 since 2021Artificial intelligence and machine learning · 151 · 9 first-author · 70 since 2021Security and privacy · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsabstractPersonalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation (LoRA) enable efficient single-concept customization by injecting lightweight, concept-specific adapters into pre-trained diffusion models. However, combining multiple LoRA modules for multi-concept generation often leads to identity missing and visual feature leakage. In this work, we identify two key issues behind these failures: (1) token-wise interference among different LoRA modules, and (2) spatial misalignment between the attention map of a rare token and its corresponding concept-specific region. To address these issues, we propose Token-Aware LoRA (TARA), which introduces a token mask to explicitly constrain each module to focus on its associated rare token to avoid interference, and a training objective that encourages the spatial attention of a rare token to align with its concept region. Our method enables training-free multi-concept composition by directly injecting multiple independently trained TARA modules at inference time. Experimental results demonstrate that TARA enables efficient multi-concept inference and effectively preserving the visual identity of each concept by avoiding mutual interference between LoRA modules. Yuqi Peng, Lingtao Zheng, Yi Huang 0035, Mingfu Yan, Jianzhuang Liu, Shifeng Chen |
AAAI | 6 |
| 2026 | OutDreamer: Video Outpainting With a Diffusion TransformerabstractVideo outpainting is a challenging task that generates new video content by extending beyond the boundaries of an original input video, requiring both temporal and spatial consistency. Many existing methods utilize latent diffusion models with U-Net backbones but still struggle to achieve high quality and adaptability in generated content. Diffusion transformers (DiTs) have emerged as a promising alternative because of their superior performance. We introduce OutDreamer, a DiT-based video outpainting framework comprising two main components: a video control branch and a conditional outpainting branch. The video control branch effectively extracts masked video information, while the conditional outpainting branch generates missing content based on these extracted conditions. Additionally, we propose a mask-driven self-attention layer that dynamically integrates the given mask information, further enhancing the model's adaptability to outpainting tasks. Furthermore, we introduce a latent alignment loss to maintain overall consistency both within and between frames. For long video outpainting, we employ a cross-video-clip refiner to iteratively generate missing content, ensuring temporal consistency across video clips. Extensive evaluations demonstrate that our OutDreamer outperforms existing video outpainting methods on widely recognized benchmarks. Linhao Zhong 0001, Yi Huang 0035, Jianzhuang Liu, Renjing Pei, Fenglong Song |
IEEE Trans. Image Process. | 4 |
| 2026 | Context Modeling With Multimodal Prompts for Emotion Recognition in ConversationabstractEmotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. After the extensive exploration of the text modality, visual and audio information has attracted considerable attention. Most of the existing approaches adopt either the attention mechanism or graph neural networks to conduct multi-modal fusion by directly utilizing pre-extracted single modal features, but they ignore the inherent priors and characteristics (highly related to emotions) contained in different modalities. For example, for visual modality, facial expression can well reveal a person's emotional state, and the intonation, speech rate and volume of audio modality also reflect emotional fluctuation. Thus, in this work, we aim to promote context fusion among multi-modal features by exploring inherent modal priors. Firstly, we create textual descriptions for different emotion categories belonging to different modalities, denoted as multimodal prompts which can be generated either from common sense or using large language models (e.g., ChatGPT). Then, we propose an adaptive gating fusion module which dynamically learns the weights between unimodal features and prompts, and allows the multimodal prompts to participate in encoding and enriching modal information. Finally, to facilitate multimodal fusion, we design a multimodal progressive encoder to learn inter-modal interactions among conversational utterances. Experimental results show that our model outperforms state-of-the-art models in ERC on two popular benchmark datasets. Weihong Ren, Yu Gao 0010, Jianzhuang Liu, Honghai Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Decoupling Appearance Variations with 3D Consistent Features in Gaussian SplattingabstractGaussian Splatting has emerged as a prominent 3D representation in novel view synthesis, but it still suffers from appearance variations, which are caused by various factors, such as modern camera ISPs, different time of day, weather conditions, and local light changes. These variations can lead to floaters and color distortions in the rendered images/videos. Recent appearance modeling approaches in Gaussian Splatting are either tightly coupled with the rendering process, hindering real-time rendering, or they only account for mild global variations, performing poorly in scenes with local light changes. In this paper, we propose DAVIGS, a method that decouples appearance variations in a plug-and-play and efficient manner. By transforming the rendering results at the image level instead of the Gaussian level, our approach can model appearance variations with minimal optimization time and memory overhead. Furthermore, our method gathers appearance-related information in 3D space to transform the rendered images, thus building 3D consistency across views implicitly. We validate our method on several appearance-variant scenes, and demonstrate that it achieves state-of-the-art rendering quality with minimal training time and memory usage, without compromising rendering speeds. Additionally, it provides performance improvements for different Gaussian Splatting baselines in a plug-and-play manner. Zhihao Li 0002, Binxiao Huang, Jianzhuang Liu, Shiyong Liu, Fenglong Song, Wenming Yang |
AAAI | 5 |
| 2025 | DIVE: Taming DINO for Subject-Driven Video EditingabstractBuilding on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these issues, this paper proposes DINO-guided Video Editing (DIVE), a framework designed to facilitate subject-driven editing in source videos conditioned on either target text prompts or reference images with specific identities. The core of DIVE lies in leveraging the powerful semantic features extracted from a pretrained DINOv2 model as implicit correspondences to guide the editing process. Specifically, to ensure temporal motion consistency, DIVE employs DINO features to align with the motion trajectory of the source video. For precise subject editing, DIVE incorporates the DINO features of reference images into a pretrained text-to-image model to learn Low-Rank Adaptations (LoRAs), effectively registering the target subject's identity. Extensive experiments on diverse real-world videos demonstrate that our framework can achieve high-quality editing results with robust motion consistency, highlighting the potential of DINO to contribute to video editing. Project page: https://dino-video-editing.github.io Yi Huang 0035, Wei Xiong 0008, He Zhang 0004, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen |
ICCV | 5 |
| 2025 | OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and RenderingabstractIn large-scale scene reconstruction using 3D Gaussian splatting, it is common to partition the scene into multiple smaller regions and reconstruct them individually. However, existing division methods are occlusion-agnostic, meaning that each region may contain areas with severe occlusions. As a result, the cameras within those regions are less correlated, leading to a low average contribution to the overall reconstruction. In this paper, we propose an occlusion-aware scene division strategy that clusters training cameras based on their positions and co-visibilities to acquire multiple regions. Cameras in such regions exhibit stronger correlations and a higher average contribution, facilitating high-quality scene reconstruction. We further propose a region-based rendering technique to accelerate large scene rendering, which culls Gaussians invisible to the region where the viewpoint is located. Such a technique significantly speeds up the rendering without compromising quality. Extensive experiments on multiple large scenes show that our method achieves superior reconstruction results with faster rendering speed compared to existing state-of-the-art approaches. Project page: https://occlugaussian.github.io. Shiyong Liu, Zhihao Li 0002, Yingfan He, Chongjie Ye, Jianzhuang Liu, Binxiao Huang, Shunbo Zhou |
ICCV | 6 |
| 2025 | IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image PromptsabstractRecent advances in 3D generation have been remarkable, with methods such as DreamFusion leveraging large-scale text-to-image diffusion-based models to guide 3D object generation. These methods enable the synthesis of detailed and photorealistic textured objects. However, the appearance of 3D objects produced by such text-to-3D models is often unpredictable, and it is hard for single-image-to-3D methods to deal with images lacking a clear subject, complicating the generation of appearance-controllable 3D objects from complex images. To address these challenges, we present IPDreamer, a novel method that captures intricate appearance features from complex **I**mage **P**rompts and aligns the synthesized 3D object with these extracted features, enabling high-fidelity, appearance-controllable 3D object generation. Our experiments demonstrate that IPDreamer consistently generates high-quality 3D objects that align with both the textual and complex image prompts, highlighting its promising capability in appearance-controlled, complex 3D object generation. Bohan Zeng, Shanglin Li, Yutang Feng, Ling Yang 0006, Hong Li 0016, Conghui He, Wentao Zhang 0001, Jianzhuang Liu, Baochang Zhang 0001, Shuicheng Yan |
ICLR | 10 |
| 2025 | InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsabstractIn multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality could achieve alone. Existing methods may struggle to effectively capture the full spectrum of synergistic information, leading to suboptimal performance in tasks where such interactions are critical. This is particularly problematic because synergistic information constitutes the fundamental value proposition of multimodal representation. To address this challenge, we introduce InfMasking, a contrastive synergistic information extraction method designed to enhance synergistic information through an Infinite Masking strategy. InfMasking stochastically occludes most features from each modality during fusion, preserving only partial information to create representations with varied synergistic patterns. Unmasked fused representations are then aligned with masked ones through mutual information maximization to encode comprehensive synergistic information. This infinite masking strategy enables capturing richer interactions by exposing the model to diverse partial modality combinations during training. As computing mutual information estimates with infinite masking is computationally prohibitive, we derive an InfMasking loss to approximate this calculation. Through controlled experiments, we demonstrate that InfMasking effectively enhances synergistic information between modalities. In evaluations on large-scale real-world datasets, InfMasking achieves state-of-the-art performance across seven benchmarks. Code is released at https://github.com/brightest66/InfMasking. Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai 0001, Dongkai Wang, Zhao Kang 0001, Jun Wang 0089, Zenglin Xu, Jiang Duan |
NeurIPS | 3 |
| 2025 | Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image EditingabstractText-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity. Jiancheng Huang, Yi Huang 0035, Jianzhuang Liu, Yifan Liu 0001, Shifeng Chen |
WACV | 3 |
| 2025 | Implicit Diffusion Models for Continuous Super-Resolution
Xuhui Liu, Sicheng Gao, Bohan Zeng, Tian Wang 0002, Jianzhuang Liu, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | TraDiffusion: Trajectory-Based Training-Free Image Generation
Mingrui Wu, Oucheng Huang, Jiayi Ji, Jianzhuang Liu, Xiaoshuai Sun, Liujuan Cao, Rongrong Ji |
Int. J. Comput. Vis. | 4 |
| 2025 | Diffusion Model-Based Image Editing: A SurveyabstractDenoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Human Motion Video Generation: A SurveyabstractHuman motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Language-Driven Visual Consensus for Zero-Shot Semantic SegmentationabstractThe pre-trained vision-language model, exemplified by CLIP, advances zero-shot semantic segmentation by aligning visual features with class embeddings through a transformer decoder to generate semantic masks. Despite its effectiveness, prevailing methods within this paradigm encounter challenges, including overfitting on seen classes and small fragmentation in segmentation masks. To mitigate these issues, we propose a Language-Driven Visual Consensus (LDVC) approach, fostering improved alignment of linguistic and visual information. Specifically, we leverage class embeddings as anchors due to their discrete and abstract nature, steering visual features toward class embeddings. Moreover, to achieve a more compact visual space, we introduce route attention into the transformer decoder to find visual consensus, thereby enhancing semantic consistency within the same object. Equipped with a vision-language prompting strategy, our approach significantly boosts the generalization capacity of segmentation models for unseen classes. Experimental results underscore the effectiveness of our approach, showcasing mIoU gains of 4.5% on the PASCAL VOC 2012 and 3.6% on the COCO-Stuff 164K for unseen classes compared with the state-of-the-art methods. Wei Ke 0003, Yi Zhu 0004, Xiaodan Liang, Jianzhuang Liu, Qixiang Ye, Tong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Controllable Mind Visual Diffusion ModelabstractBrain signal visualization has emerged as an active research area, serving as a critical interface between the human visual system and computer vision models. Diffusion-based methods have recently shown promise in analyzing functional magnetic resonance imaging (fMRI) data, including the reconstruction of high-quality images consistent with original visual stimuli. Nonetheless, it remains a critical challenge to effectively harness the semantic and silhouette information extracted from brain signals. In this paper, we propose a novel approach, termed as Controllable Mind Visual Diffusion Model (CMVDM). Specifically, CMVDM first extracts semantic and silhouette information from fMRI data using attribute alignment and assistant networks. Then, a control model is introduced in conjunction with a residual block to fully exploit the extracted information for image synthesis, generating high-quality images that closely resemble the original visual stimuli in both semantic content and silhouette characteristics. Through extensive experimentation, we demonstrate that CMVDM outperforms existing state-of-the-art methods both qualitatively and quantitatively. Our code is available at https://github.com/zengbohan0217/CMVDM. Bohan Zeng, Shanglin Li, Xuhui Liu, Sicheng Gao, Xu Tang 0007, Yao Hu 0002, Jianzhuang Liu, Baochang Zhang 0001 |
AAAI | 8 |
| 2024 | UV-IDM: Identity-Conditioned Latent Diffusion Model for Face UV-Texture Generationabstract3D face reconstruction aims at generating high-fidelity 3D face shapes and textures from single-view or multi-view images. However, current prevailing facial texture generation methods generally suffer from low-quality texture, identity information loss, and inadequate handling of occlusions. To solve these problems, we introduce an Identity-Conditioned Latent Diffusion Model for face UV-texture generation (UV-IDM) to generate photo-realistic textures based on the Basel Face Model (BFM). UV-IDM leverages the powerful texture generation capacity of a latent diffusion model (LDM) to obtain detailed facial textures. To preserve the identity during the reconstruction procedure, we design an identity-conditioned module that can utilize any in-the-wild image as a robust condition for the LDM to guide texture generation. UV-IDM can be easily adapted to different BFM-based methods as a high-fidelity texture generator. Furthermore, in light of the limited accessibility of most existing UV-texture datasets, we build a large-scale and publicly available UV-texture dataset based on BFM, termed BFM-UV. Extensive experiments show that our UV-IDM can generate high-fidelity textures in 3D face reconstruction within seconds while maintaining image consistency, bringing new state-of-the-art performance in facial texture generation. Hong Li 0016, Yutang Feng, Xuhui Liu, Bohan Zeng, Shanglin Li, Jianzhuang Liu, Shumin Han, Baochang Zhang 0001 |
CVPR | 8 |
| 2024 | ZONE: Zero-Shot Instruction-Guided Local EditingabstractRecent advances in vision-language models like Stable Diffusion have shown remarkable power in creative image synthesis and editing. However, most existing text-to-image editing methods encounter two obstacles: First, the text prompt needs to be carefully crafted to achieve good results, which is not intuitive or user-friendly. Second, they are in-sensitive to local edits and can irreversibly affect non-edited regions, leaving obvious editing traces. To tackle these problems, we propose a Zero-shot instructiON-guided local image Editing approach, termed ZONE. We first convert the editing intent from the user-provided instruction (e.g., “make his tie blue”) into specific image editing regions through InstructPix2Pix. We then propose a Region-loll scheme for precise image layer extraction from an off-the-shelf segment model. We further develop an edge smoother based on FFT for seamless blending between the layer and the image. Our method allows for arbitrary manipulation of a specific region with a single instruction while preserving the rest. Extensive experiments demonstrate that our Z ONE achieves remarkable local editing results and user-friendliness, outperforming state-of-the-art methods. Code is available at https://github.com/ls1001006/ZONE. Shanglin Li, Bohan Zeng, Yutang Feng, Sicheng Gao, Xiuhui Liu, Xu Tang 0007, Yao Hu 0002, Jianzhuang Liu, Baochang Zhang 0001 |
CVPR | 10 |
| 2024 | VastGaussian: Vast 3D Gaussians for Large Scene ReconstructionabstractExisting NeRF-based methods for large scene reconstruction often have limitations in visual quality and rendering speed. While the recent 3D Gaussian Splatting works well on small-scale and object-centric scenes, scaling it up to large scenes poses challenges due to limited video memory, long optimization time, and noticeable appearance variations. To address these challenges, we present VastGaussian, the first method for high-quality reconstruction and real-time rendering on large scenes based on 3D Gaussian Splatting. We propose a progressive partitioning strategy to divide a large scene into multiple cells, where the training cameras and point cloud are properly distributed with an airspace-aware visibility criterion. These cells are merged into a complete scene after parallel optimization. We also introduce decoupled appearance modeling into the optimization process to reduce appearance variations in the rendered images. Our approach outperforms existing NeRF-based methods and achieves state-of-the-art results on multiple large scene datasets, enabling fast optimization and high-fidelity real-time rendering. Project page: https://vastgaussian.github.io. Zhihao Li 0002, Jianzhuang Liu, Shiyong Liu, Jiayue Liu, Yangdi Lu, Songcen Xu, Youliang Yan, Wenming Yang |
CVPR | 4 |
| 2024 | CoSeR: Bridging Image and Language for Cognitive Super-ResolutionabstractExisting super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the intro-duction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the ca-pacity to comprehend low-resolution images. We achieve this by marrying image appearance and language under-standing to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To fur-ther improve image fidelity, we propose a novel condition injection scheme called “Ali-in-Attention ”, consolidating all conditional information into a single module. Conse-quently, our method successfully restores semantically cor-rect and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/ Haoze Sun, Wenbo Li 0002, Jianzhuang Liu, Haoyu Chen 0003, Renjing Pei, Xueyi Zou, Youliang Yan, Yujiu Yang 0001 |
CVPR | 3 |
| 2024 | MagicEraser: Erasing Any Objects via Semantics-Aware Control
Zixiao Zhang, Yi Huang 0035, Jianzhuang Liu, Renjing Pei, Songcen Xu |
ECCV (28) | 4 |
| 2024 | MirrorGaussian: Reflecting 3D Gaussians for Reconstructing Mirror Reflections
Jiayue Liu, Freeman Cheng, Roy Yang, Zhihao Li 0002, Jianzhuang Liu, Yi Huang 0035, Shiyong Liu, Songcen Xu, Chun Yuan 0003 |
ECCV (72) | 6 |
| 2024 | UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
Songlin Tang, Guangming Lu 0002, Jianzhuang Liu, Wenjie Pei |
ECCV (71) | 4 |
| 2024 | Wavelet-based network for high dynamic range imagingabstractHigh dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance. Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan |
Comput. Vis. Image Underst. | 4 |
| 2024 | Correctable Landmark Discovery via Large Models for Vision-Language NavigationabstractVision-Language Navigation (VLN) requires the agent to follow language instructions to reach a target position. A key factor for successful navigation is to align the landmarks implied in the instruction with diverse visual observations. However, previous VLN agents fail to perform accurate modality alignment especially in unexplored scenes, since they learn from limited navigation data and lack sufficient open-world alignment knowledge. In this work, we propose a new VLN paradigm, called COrrectable LaNdmark DiScOvery via Large ModEls (CONSOLE). In CONSOLE, we cast VLN as an open-world sequential landmark discovery problem, by introducing a novel correctable landmark discovery scheme based on two large models ChatGPT and CLIP. Specifically, we use ChatGPT to provide rich open-world landmark cooccurrence commonsense, and conduct CLIP-driven landmark discovery based on these commonsense priors. To mitigate the noise in the priors due to the lack of visual constraints, we introduce a learnable cooccurrence scoring module, which corrects the importance of each cooccurrence according to actual observations for accurate landmark discovery. We further design an observation enhancement strategy for an elegant combination of our framework with different VLN agents, where we utilize the corrected landmark features to obtain enhanced observation features for action decision. Extensive experimental results on multiple popular VLN benchmarks (R2R, REVERIE, R4R, RxR) show the significant superiority of CONSOLE over strong baselines. Especially, our CONSOLE establishes the new state-of-the-art results on R2R and R4R in unseen scenarios. Bingqian Lin, Yunshuang Nie, Ziming Wei 0001, Yi Zhu 0004, Hang Xu 0004, Shikui Ma, Jianzhuang Liu, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | MVEB: Self-Supervised Learning With Multi-View Entropy BottleneckabstractSelf-supervised learning aims to learn representation that can be effectively generalized to downstream tasks. Many self-supervised approaches regard two views of an image as both the input and the self-supervised signals, assuming that either view contains the same task-relevant information and the shared information is (approximately) sufficient for predicting downstream tasks. Recent studies show that discarding superfluous information not shared between the views can improve generalization. Hence, the ideal representation is sufficient for downstream tasks and contains minimal superfluous information, termed minimal sufficient representation. One can learn this representation by maximizing the mutual information between the representation and the supervised view while eliminating superfluous information. Nevertheless, the computation of mutual information is notoriously intractable. In this work, we propose an objective termed multi-view entropy bottleneck (MVEB) to learn minimal sufficient representation effectively. MVEB simplifies the minimal sufficient learning to maximizing both the agreement between the embeddings of two views and the differential entropy of the embedding distribution. Our experiments confirm that MVEB significantly improves performance. For example, it achieves top-1 accuracy of 76.9% on ImageNet with a vanilla ResNet-50 backbone on linear evaluation. To the best of our knowledge, this is the new state-of-the-art result with ResNet-50. Liangjian Wen, Xiasi Wang, Jianzhuang Liu, Zenglin Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | IFAST: Weakly Supervised Interpretable Face Anti-Spoofing From Single-Shot Binocular NIR ImagesabstractSingle-shot face anti-spoofing (FAS) is a key technique for securing face recognition systems, relying solely on static images as input. However, single-shot FAS remains a challenging and under-explored problem due to two reasons: 1) On the data side, learning FAS from RGB images is largely context-dependent, and single-shot images without additional annotations contain limited semantic information. 2) On the model side, existing single-shot FAS models struggle to provide proper evidence for their decisions, and FAS methods based on depth estimation require expensive per-pixel annotations. To address these issues, we construct and release a large binocular NIR image dataset named BNI-FAS, which contains more than 300,000 real face and plane attack images, and propose an Interpretable FAS Transformer (IFAST) that requires only weak supervision to produce interpretable predictions. Our IFAST generates pixel-wise disparity maps using the proposed disparity estimation Transformer with Dynamic Matching Attention (DMA) blocks. Besides, we design a confidence map generator to work in tandem with a dual-teacher distillation module to obtain the final discriminant results. Comprehensive experiments show that our IFAST achieves state-of-the-art performance on BNI-FAS, verifying its effectiveness of single-shot FAS on binocular NIR images. The project page is available athttps://ifast-bni.github.io/. Jiancheng Huang, Jianzhuang Liu, Linxiao Shi, Shifeng Chen |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Video Frame Interpolation With Stereo Event and Intensity CamerasabstractThe stereo event-intensity camera setup is widely applied to leverage the advantages of both event cameras with low latency and intensity cameras that capture accurate brightness and texture information. However, such a setup commonly encounters cross-modality parallax that is difficult to be eliminated solely with stereo rectification especially for real-world scenes with complex motions and varying depths, posing artifacts and distortion for existing Event-based Video Frame Interpolation (E-VFI) approaches. To tackle this problem, we propose a novel Stereo Event-based VFI (SE-VFI) network (SEVFI-Net) to generate high-quality intermediate frames and corresponding disparities from misaligned inputs consisting of two consecutive keyframes and event streams emitted between them. Specifically, we propose a Feature Aggregation Module (FAM) to alleviate the parallax and achieve spatial alignment in the feature domain. We then exploit the fused features accomplishing accurate optical flow and disparity estimation, and achieving better interpolated results through flow-based and synthesis-based ways. We also build a stereo visual acquisition system composed of an event camera and an RGB-D camera to collect a new Stereo Event-Intensity Dataset (SEID) containing diverse scenes with complex motions and varying depths. Experiments on public real-world stereo datasets, i.e., DSEC and MVSEC, and our SEID dataset, demonstrate that our proposed SEVFI-Net outperforms state-of-the-art methods by a large margin. The code and dataset are available athttps://dingchao1214.github.io/web_sevfi/. Chao Ding 0003, Mingyuan Lin, Jianzhuang Liu, Lei Yu 0006 |
IEEE Trans. Multim. | 4 |
| 2024 | WaveDM: Wavelet-Based Diffusion Models for Image RestorationabstractLatest diffusion-based methods for many image restoration tasks outperform traditional models, but they encounter the long-time inference problem. To tackle it, this paper proposes a Wavelet-Based Diffusion Model (WaveDM). WaveDM learns the distribution of clean images in the wavelet domain conditioned on the wavelet spectrum of degraded images after wavelet transform, which is more time-saving in each step of sampling than modeling in the spatial domain. To ensure restoration performance, a unique training strategy is proposed where the low-frequency and high-frequency spectrums are learned using distinct modules. In addition, an Efficient Conditional Sampling (ECS) strategy is developed from experiments, which reduces the number of total sampling steps to around 5. Evaluations on twelve benchmark datasets including image raindrop removal, rain steaks removal, dehazing, defocus deblurring, demoiréing, and denoising demonstrate that WaveDM achieves state-of-the-art performance with the efficiency that is comparable to traditional one-pass methods and over 100× faster than existing image restoration methods using vanilla diffusion models. The code is available athttps://github.com/stayalive16/WaveDM Yi Huang 0035, Jiancheng Huang, Jianzhuang Liu, Mingfu Yan, Jiaxi Lv, Chaoqi Chen, Shifeng Chen |
IEEE Trans. Multim. | 3 |
| 2023 | Low-Light Video Enhancement with Synthetic Event GuidanceabstractLow-light video enhancement (LLVE) is an important yet challenging task with many applications such as photographing and autonomous driving. Unlike single image low-light enhancement, most LLVE methods utilize temporal information from adjacent frames to restore the color and remove the noise of the target frame. However, these algorithms, based on the framework of multi-frame alignment and enhancement, may produce multi-frame fusion artifacts when encountering extreme low light or fast motion. In this paper, inspired by the low latency and high dynamic range of events, we use synthetic events from multiple frames to guide the enhancement and restoration of low-light videos. Our method contains three stages: 1) event synthesis and enhancement, 2) event and image fusion, and 3) low-light enhancement. In this framework, we design two novel modules (event-image fusion transform and event-guided dual branch) for the second and third stages, respectively. Extensive experiments show that our method outperforms existing low-light video or single image enhancement approaches on both synthetic and real LLVE datasets. Our code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/LLVE-SEG. Lin Liu 0016, Junfeng An, Jianzhuang Liu, Shanxin Yuan, Xiangyu Chen 0006, Wengang Zhou 0001, Houqiang Li, Yanfeng Wang 0001, Qi Tian 0001 |
AAAI | 3 |
| 2023 | Actional Atomic-Concept Learning for Demystifying Vision-Language NavigationabstractVision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existing VLN agents directly learn to align the raw directional features and visual features trained using one-hot labels to linguistic instruction features. However, the big semantic gap among these multi-modal inputs makes the alignment difficult and therefore limits the navigation performance. In this paper, we propose Actional Atomic-Concept Learning (AACL), which maps visual observations to actional atomic concepts for facilitating the alignment. Specifically, an actional atomic concept is a natural language phrase containing an atomic action and an object, e.g., ``go up stairs''. These actional atomic concepts, which serve as the bridge between observations and instructions, can effectively mitigate the semantic gap and simplify the alignment. AACL contains three core components: 1) a concept mapping module to map the observations to the actional atomic concept representations through the VLN environment and the recently proposed Contrastive Language-Image Pretraining (CLIP) model, 2) a concept refining adapter to encourage more instruction-oriented object concept extraction by re-ranking the predicted object concepts by CLIP, and 3) an observation co-embedding module which utilizes concept representations to regularize the observation representations. Our AACL establishes new state-of-the-art results on both fine-grained (R2R) and high-level (REVERIE and R2R-Last) VLN benchmarks. Moreover, the visualization shows that AACL significantly improves the interpretability in action decision. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/VLN-AACL. Bingqian Lin, Yi Zhu 0004, Xiaodan Liang, Liang Lin 0004, Jianzhuang Liu |
AAAI | 5 |
| 2023 | Implicit Diffusion Models for Continuous Super-ResolutionabstractImage super-resolution (SR) has attracted increasing attention due to its widespread applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continuous image super-resolution. IDM integrates an implicit neural representation and a denoising diffusion model in a unified end-to-end framework, where the implicit neural representation is adopted in the decoding process to learn continuous-resolution representation. Furthermore, we design a scale-adaptive conditioning mechanism that consists of a low-resolution (LR) conditioning network and a scaling factor. The scaling factor regulates the resolution and accordingly modulates the proportion of the LR information and generated features in the final output, which enables the model to accommodate the continuous-resolution requirement. Extensive experiments validate the effectiveness of our IDM and demonstrate its superior performance over prior arts. The source code will be available at https://github.com/Ree1s/IDM. Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu 0007, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, Baochang Zhang 0001 |
CVPR | 7 |
| 2023 | CLIPPING: Distilling CLIP-Based Models with a Student Base for Video-Language RetrievalabstractPre-training a vision-language model and then fine-tuning it on downstream tasks have become a popular paradigm. However, pre-trained vision-language models with the Transformer architecture usually take long inference time. Knowledge distillation has been an efficient technique to transfer the capability of a large model to a small one while maintaining the accuracy, which has achieved remarkable success in natural language processing. However, it faces many problems when applying KD to the multi-modality applications. In this paper, we propose a novel knowledge distillation method, named CLIPPING11In this paper, CLIPPING means cutting something to make it smaller through distilling., where the plentiful knowledge of a large teacher model that has been fine-tuned for video-language tasks with the powerful pre-trained CLIP can be effectively transferred to a small student only at the fine-tuning stage. Especially, a new layer-wise alignment with the student as the base is proposed for knowledge distillation of the intermediate layers in CLIPPING, which enables the student's layers to be the bases of the teacher, and thus allows the student to fully absorb the knowledge of the teacher. CLIPPING with MobileViT-v2 as the vision encoder without any vision-language pre-training achieves 88.1%-95.3% of the performance of its teacher on three video-language retrieval benchmarks, with its vision encoder being 19.5x smaller. CLIPPING also significantly outperforms a state-of-the-art small baseline (ALL-in-one-B) on the MSR-VTT dataset, obtaining relatively 7.4% performance gain, with 29% fewer parameters and 86.9% fewer flops. Moreover, CLIPPING is comparable or even superior to many large pre-training models. Renjing Pei, Jianzhuang Liu, Weimian Li, Songcen Xu, Peng Dai 0002, Juwei Lu, Youliang Yan |
CVPR | 2 |
| 2023 | AttriCLIP: A Non-Incremental Learner for Incremental Knowledge LearningabstractContinual learning aims to enable a model to incrementally learn knowledge from sequentially arrived data. Previous works adopt the conventional classification architecture, which consists of a feature extractor and a classifier. The feature extractor is shared across sequentially arrived tasks or classes, but one specific group of weights of the classifier corresponding to one new class should be incrementally expanded. Consequently, the parameters of a continual learner gradually increase. Moreover, as the classifier contains all historical arrived classes, a certain size of the memory is usually required to store rehearsal data to mitigate classifier bias and catastrophic forgetting. In this paper, we propose a non-incremental learner, named AttriCLIP, to incrementally extract knowledge of new classes or tasks. Specifically, AttriCLIP is built upon the pre-trained visual-language model CLIP. Its image encoder and text encoder are fixed to extract features from both images and text. Text consists of a category name and a fixed number of learnable parameters which are selected from our designed attribute word bank and serve as attributes. As we compute the visual and textual similarity for classification, AttriCLIP is a non-incremental learner. The attribute prompts, which encode the common knowledge useful for classification, can effectively mitigate the catastrophic forgetting and avoid constructing a replay memory. We evaluate our AttriCLIP and compare it with CLIP-based and previous state-of-the-art continual learning methods in realistic settings with domain-shift and long-sequence learning. The results show that our method performs favorably against previous state-of-the-arts. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/AttriCLIP. Runqi Wang, Xiaoyue Duan, Guoliang Kang, Jianzhuang Liu, Shaohui Lin, Songcen Xu, Jinhu Lü 0001, Baochang Zhang 0001 |
CVPR | 4 |
| 2023 | SmartAssign: Learning A Smart Knowledge Assignment Strategy for Deraining and DesnowingabstractExisting methods mainly handle single weather types. However, the connections of different weather conditions at deep representation level are usually ignored. These connections, if used properly, can generate complementary representations for each other to make up insufficient training data, obtaining positive performance gains and better generalization. In this paper, we focus on the very correlated rain and snow to explore their connections at deep representation level. Because sub-optimal connections may cause negative effect, another issue is that if rain and snow are handled in a multi-task learning way, how to find an optimal connection strategy to simultaneously improve deraining and desnowing performance. To build desired connection, we propose a smart knowledge assignment strategy, called SmartAssign, to optimally assign the knowledge learned from both tasks to a specific one. In order to further enhance the accuracy of knowledge assignment, we propose a novel knowledge contrast mechanism, so that the knowledge assigned to different tasks preserves better uniqueness. The inherited inductive biases usually limit the modelling ability of CNNs, we introduce a novel transformer block to constitute the backbone of our network to effectively combine long-range context dependency and local image details. Extensive experiments on seven benchmark datasets verify that proposed SmartAssign explores effective connection between rain and snow, and improves the performances of both deraining and desnowing apparently. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/SmartAssign. Yinglong Wang 0002, Chao Ma 0004, Jianzhuang Liu |
CVPR | 3 |
| 2023 | Few-Shot Learning with Visual Distribution Calibration and Cross-Modal Distribution AlignmentabstractPre-trained vision-language models have inspired much research on few-shot learning. However, with only a few training images, there exist two crucial problems: (1) the visual feature distributions are easily distracted by class-irrelevant information in images, and (2) the alignment between the visual and language feature distributions is difficult. To deal with the distraction problem, we propose a Selective Attack module, which consists of trainable adapters that generate spatial attention maps of images to guide the attacks on class-irrelevant image areas. By messing up these areas, the critical features are captured and the visual distributions of image features are calibrated. To better align the visual and language feature distributions that describe the same object class, we propose a cross-modal distribution alignment module, in which we introduce a vision-language prototype for each class to align the distributions, and adopt the Earth Mover's Distance (EMD) to optimize the prototypes. For efficient computation, the upper bound of EMD is derived. In addition, we propose an augmentation strategy to increase the diversity of the images and the text prompts, which can reduce overfitting to the few-shot training images. Extensive experiments on 11 datasets demonstrate that our method consistently outperforms prior arts in few-shot learning. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/SADA. Runqi Wang, Xiaoyue Duan, Jianzhuang Liu, Yuning Lu, Tian Wang 0002, Songcen Xu, Baochang Zhang 0001 |
CVPR | 4 |
| 2023 | Low-Light Image Enhancement with Illumination-Aware Gamma Correction and Complete Image Modelling NetworkabstractThis paper presents a novel network structure with illumination-aware gamma correction and complete image modelling to solve the low-light image enhancement problem. Low-light environments usually lead to less informative large-scale dark areas, directly learning deep representations from low-light images is insensitive to recovering normal illumination. We propose to integrate the effectiveness of gamma correction with the strong modelling capacities of deep networks, which enables the correction factor gamma to be learned in a coarse to elaborate manner via adaptively perceiving the deviated illumination. Because exponential operation introduces high computational complexity, we propose to use Taylor Series to approximate gamma correction, accelerating the training and inference speed. Dark areas usually occupy large scales in low-light images, common local modelling structures, e.g., CNN, SwinIR, are thus insufficient to recover accurate illumination across whole low-light images. We propose a novel Transformer block to completely simulate the dependencies of all pixels across images via a local-to-global hierarchical attention mechanism, so that dark areas could be inferred by borrowing the information from far informative regions in a highly effective manner. Extensive experiments on several benchmark datasets demonstrate that our approach outperforms state-of-the-art methods. Yinglong Wang 0002, Zhen Liu 0022, Jianzhuang Liu, Songcen Xu, Shuaicheng Liu |
ICCV | 3 |
| 2023 | MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationabstractRecently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at the pixel level and predicting accurate object masks. To address this issue, we propose MixReorg, a novel and straightforward pre-training paradigm for semantic segmentation that enhances a model’s ability to reorganize patches mixed across images, exploring both local visual relevance and global semantic coherence. Our approach involves generating fine-grained patch-text pairs data by mixing image patches while preserving the correspondence between patches and text. The model is then trained to minimize the segmentation loss of the mixed images and the two contrastive losses of the original and restored features. With MixReorg as a mask learner, conventional text-supervised semantic segmentation models can achieve highly generalizable pixel-semantic alignment ability, which is crucial for open-world segmentation. After training with large-scale image-text data, MixReorg models can be applied directly to segment visual objects of arbitrary categories, without the need for further fine-tuning. Our proposed framework demonstrates strong performance on popular zero-shot semantic segmentation benchmarks, outperforming GroupViT by significant margins of 5.0%, 6.2%, 2.5%, and 3.4% mIoU on PASCAL VOC2012, PASCAL Context, MS COCO, and ADE20K, respectively. Kaixin Cai, Pengzhen Ren, Yi Zhu 0004, Hang Xu 0004, Jianzhuang Liu, Guangrun Wang, Xiaodan Liang |
ICCV | 5 |
| 2023 | PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video RetrievalabstractText-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video retrieval. However, due to the modality difference between videos and images, how to effectively adapt CLIP to the video domain is still underexplored. In this paper, we investigate this problem from two aspects. First, we enhance the transferred image encoder of CLIP for fine-grained video understanding in a seamless fashion. Second, we conduct fine-grained contrast between videos and texts from both model improvement and loss design. Particularly, we propose a fine-grained contrastive model equipped with parallel isomeric attention and dynamic routing, namely PIDRo, for text-video retrieval. The parallel isomeric attention module is used as the video encoder, which consists of two parallel branches modeling the spatial-temporal information of videos from both patch and frame levels. The dynamic routing module is constructed to enhance the text encoder of CLIP, generating informative word representations by distributing the fine-grained information to the related word tokens within a sentence. Such model design provides us with informative patch, frame and word representations. We then conduct token-wise interaction upon them. With the enhanced encoders and the token-wise loss, we are able to achieve finer-grained text-video alignment and more accurate retrieval. PIDRo obtains state-of-the-art performance over various text-video retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, DiDeMo and ActivityNet. Peiyan Guan, Renjing Pei, Jianzhuang Liu, Weimian Li, Jiaxi Gu, Hang Xu 0004, Songcen Xu, Youliang Yan, Edmund Y. Lam |
ICCV | 4 |
| 2023 | HiVLP: Hierarchical Interactive Video-Language Pre-TrainingabstractVideo-Language Pre-training (VLP) has become one of the most popular research topics in deep learning. However, compared to image-language pre-training, VLP has lagged far behind due to the lack of large amounts of video-text pairs. In this work, we train a VLP model with a hybrid of image-text and video-text pairs, which significantly outperforms pre-training with only the video-text pairs. Besides, existing methods usually model the cross-modal interaction using cross-attention between single-scale visual tokens and textual tokens. These visual features are either of low resolutions lacking fine-grained information, or of high resolutions without high-level semantics. To address the issue, we propose Hierarchical interactive Video-Language Pre-training (HiVLP) that efficiently uses a hierarchical visual feature group for multi-modal cross-attention during pre-training. In the hierarchical framework, low-resolution features are learned with focus on more global high-level semantic information, while high-resolution features carry fine-grained details. As a result, HiVLP has the ability to effectively learn both the global and fine-grained representations to achieve better alignment between video and text inputs. Furthermore, we design a hierarchical multi-scale vision contrastive loss for self-supervised learning to boost the interaction between them. Experimental results show that HiVLP establishes new state-of-the-art results in three downstream tasks, text-video retrieval, video-text retrieval, and video captioning. Jianzhuang Liu, Renjing Pei, Songcen Xu, Peng Dai 0002, Juwei Lu, Weimian Li, Youliang Yan |
ICCV | 2 |
| 2023 | Generalizing Event-Based Motion Deblurring in Real-World ScenariosabstractEvent-based motion deblurring has shown promising results by exploiting low-latency events. However, current approaches are limited in their practical usage, as they assume the same spatial resolution of inputs and specific blurriness distributions. This work addresses these limitations and aims to generalize the performance of event-based de-blurring in real-world scenarios. We propose a scale-aware network that allows flexible input spatial scales and enables learning from different temporal scales of motion blur. A two-stage self-supervised learning scheme is then developed to fit real-world data distribution. By utilizing the relativity of blurriness, our approach efficiently ensures the restored brightness and structure of latent images and further generalizes deblurring performance to handle varying spatial and temporal scales of motion blur in a self-distillation manner. Our method is extensively evaluated, demonstrating remarkable performance, and we also introduce a real-world dataset consisting of multi-scale blurry frames and events to facilitate research in event-based deblurring. Xiang Zhang 0022, Lei Yu 0006, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
ICCV | 4 |
| 2023 | ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency
Pengzhen Ren, Hang Xu 0004, Yi Zhu 0004, Guangrun Wang, Jianzhuang Liu, Xiaojun Chang, Xiaodan Liang |
ICLR | 6 |
| 2023 | Cross-Level Distillation and Feature Denoising for Cross-Domain Few-Shot Classification
Runqi Wang, Jianzhuang Liu, Asako Kanezaki |
ICLR | 3 |
| 2023 | Multiple Instance Differentiation Learning for Active Object DetectionabstractDespite the substantial progress of active learning for image recognition, there lacks a systematic investigation of instance-level active learning for object detection. In this paper, we propose to unify instance uncertainty calculation with image uncertainty estimation for informative image selection, creating a multiple instance differentiation learning (MIDL) method for instance-level active learning. MIDL consists of a classifier prediction differentiation module and a multiple instance differentiation module. The former leverages two adversarial instance classifiers trained on the labeled and unlabeled sets to estimate instance uncertainty of the unlabeled set. The latter treats unlabeled images as instance bags and re-estimates image-instance uncertainty using the instance classification model in a multiple instance learning fashion. Through weighting the instance uncertainty using instance class probability and instance objectness probability under the total probability formula, MIDL unifies the image uncertainty with instance uncertainty in the Bayesian theory framework. Extensive experiments validate that MIDL sets a solid baseline for instance-level active learning. On commonly used object detection datasets, it outperforms other state-of-the-art methods by significant margins, particularly when the labeled sets are small. Fang Wan 0001, Qixiang Ye, Tianning Yuan, Songcen Xu, Jianzhuang Liu, Xiangyang Ji, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Learning to Super-Resolve Blurry Images With EventsabstractSuper-Resolution from a single motion Blurred image (SRB) is a severely ill-posed problem due to the joint degradation of motion blurs and low spatial resolution. In this article, we employ events to alleviate the burden of SRB and propose an Event-enhanced SRB (E-SRB) algorithm, which can generate a sequence of sharp and clear images with High Resolution (HR) from a single blurry image with Low Resolution (LR). To achieve this end, we formulate an event-enhanced degeneration model to consider the low spatial resolution, motion blurs, and event noises simultaneously. We then build an event-enhanced Sparse Learning Network (eSL-Net++) upon a dual sparse learning scheme where both events and intensity frames are modeled with sparse representations. Furthermore, we propose an event shuffle-and-merge scheme to extend the single-frame SRB to the sequence-frame SRB without any additional training process. Experimental results on synthetic and real-world datasets show that the proposed eSL-Net++ outperforms state-of-the-art methods by a large margin. Datasets, codes, and more results are available at https://github.com/ShinyWang33/eSL-Net-Plusplus. Lei Yu 0006, Bishan Wang, Xiang Zhang 0022, Wen Yang 0001, Jianzhuang Liu, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Multi-Branch Distance-Sensitive Self-Attention Network for Image CaptioningabstractSelf-attention (SA) based networks have achieved great success in image captioning, constantly dominating the leaderboards of online benchmarks. However, existing SA networks still suffer from distance insensitivity and low-rank bottleneck. In this paper, we aim to optimize SA in terms of two aspects, thereby addressing the above issues. First, we introduce a Distance-sensitive Self-Attention (DSA), which considers the raw geometric distances between query-key pairs in the 2D images during SA modeling. Second, we present a simple yet effective approach, named Multi-branch Self-Attention (MSA) to compensate for the low-rank bottleneck. MSA treats a multi-head self-attention layer as a branch and duplicates it multiple times to increase the expressive power of SA. To validate the effectiveness of the two designs, we apply them to the standard self-attention network, and conduct extensive experiments on the highly competitive MS-COCO dataset. We achieve new state-of-the-art performance on both the local and online test sets,i.e., 135.1% CIDEr on the Karpathy split and 135.4% CIDEr on the official online split. Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou, Gen Luo, Liujuan Cao, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji |
IEEE Trans. Multim. | 7 |
| 2022 | Uncertainty-Driven Dehazing NetworkabstractDeep learning has made remarkable achievements for single image haze removal. However, existing deep dehazing models only give deterministic results without discussing the uncertainty of them. There exist two types of uncertainty in the dehazing models: aleatoric uncertainty that comes from noise inherent in the observations and epistemic uncertainty that accounts for uncertainty in the model. In this paper, we propose a novel uncertainty-driven dehazing network (UDN) that improves the dehazing results by exploiting the relationship between the uncertain and confident representations. We first introduce an Uncertainty Estimation Block (UEB) to predict the aleatoric and epistemic uncertainty together. Then, we propose an Uncertainty-aware Feature Modulation (UFM) block to adaptively enhance the learned features. UFM predicts a convolution kernel and channel-wise modulation cofficients conitioned on the uncertainty weighted representation. Moreover, we develop an uncertainty-driven self-distillation loss to improve the uncertain representation by transferring the knowledge from the confident one. Extensive experimental results on synthetic datasets and real-world images show that UDN achieves significant quantitative and qualitative improvements, outperforming the state-of-the-arts. Ming Hong, Jianzhuang Liu, Cuihua Li, Yanyun Qu |
AAAI | 2 |
| 2022 | SiamTrans: Zero-Shot Multi-Frame Image Restoration with Pre-trained Siamese TransformersabstractWe propose a novel zero-shot multi-frame image restoration method for removing unwanted obstruction elements (such as rains, snow, and moire patterns) that vary in successive frames. It has three stages: transformer pre-training, zero-shot restoration, and hard patch refinement. Using the pre-trained transformers, our model is able to tell the motion difference between the true image information and the obstructing elements. For zero-shot image restoration, we design a novel model, termed SiamTrans, which is constructed by Siamese transformers, encoders, and decoders. Each transformer has a temporal attention layer and several self-attention layers, to capture both temporal and spatial information of multiple frames. Only self-supervisedly pre-trained on the denoising task, SiamTrans is tested on three different low-level vision tasks (deraining, demoireing, and desnowing). Compared with related methods, SiamTrans achieves the best performances, even outperforming those with supervised learning. Lin Liu 0016, Shanxin Yuan, Jianzhuang Liu, Youliang Yan, Qi Tian 0001 |
AAAI | 3 |
| 2022 | Diversity Matters: Fully Exploiting Depth Clues for Reliable Monocular 3D Object DetectionabstractAs an inherently ill-posed problem, depth estimation from single images is the most challenging part of monocular 3D object detection (M3OD). Many existing methods rely on preconceived assumptions to bridge the missing spatial information in monocular images, and predict a sole depth value for every object of interest. However, these assumptions do not always hold in practical applications. To tackle this problem, we propose a depth solving system that fully explores the visual clues from the subtasks in M3OD and generates multiple estimations for the depth of each target. Since the depth estimations rely on different assumptions in essence, they present diverse distributions. Even if some assumptions collapse, the estimations established on the remaining assumptions are still reliable. In addition, we develop a depth selection and combination strategy. This strategy is able to remove abnormal estimations caused by collapsed assumptions, and adaptively combine the remaining estimations into a single one. In this way, our depth solving system becomes more precise and robust. Exploiting the clues from multiple subtasks of M3OD and without introducing any extra information, our method surpasses the current best method by more than 20% relatively on the Moderate level of test split in the KITTI 3D object detection benchmark, while still maintaining real-time efficiency. Zhuoling Li, Jianzhuang Liu, Haoqian Wang, Lihui Jiang |
CVPR | 4 |
| 2022 | ADAPT: Vision-Language Navigation with Modality-Aligned Action PromptsabstractVision-Language Navigation (VLN) is a challenging task that requires an embodied agent to perform action-level modality alignment, i.e., make instruction-asked actions sequentially in complex visual environments. Most existing VLN agents learn the instruction-path data directly and cannot sufficiently explore action-level alignment knowledge inside the multi-modal inputs. In this paper, we propose modAlity-aligneD Action PrompTs (ADAPT), which provides the VLN agent with action prompts to enable the explicit learning of action-level modality alignment to pursue successful navigation. Specifically, an action prompt is defined as a modality-aligned pair of an image sub-prompt and a text sub-prompt, where the former is a single-view observation and the latter is a phrase like “walk past the chair”. When starting navigation, the instruction-related action prompt set is retrieved from a prebuilt action prompt base and passed through a prompt encoder to obtain the prompt feature. Then the prompt feature is concatenated with the original instruction feature and fed to a multilayer transformer for action prediction. To collect high-quality action prompts into the prompt base, we use the Contrastive Language-Image Pretraining (CLIP) model which has powerful cross-modality alignment ability. A modality alignment loss and a sequential consistency loss are further introduced to enhance the alignment of the action prompt and enforce the agent to focus on the related prompt sequentially. Experimental results on both R2R and RxR show the superiority of ADAPT over state-of-the-art methods. Bingqian Lin, Yi Zhu 0004, Zicong Chen, Xiwen Liang, Jianzhuang Liu, Xiaodan Liang |
CVPR | 5 |
| 2022 | Prompt Distribution LearningabstractWe present prompt distribution learning for effectively adapting a pre-trained vision-language model to address downstream recognition tasks. Our method not only learns low-bias prompts from a few samples but also captures the distribution of diverse prompts to handle the varying visual representations. In this way, we provide high-quality task-related content for facilitating recognition. This prompt distribution learning is realized by an efficient approach that learns the output embeddings of prompts instead of the input embeddings. Thus, we can employ a Gaussian distribution to model them effectively and derive a surrogate loss for efficient training. Extensive experiments on 12 datasets demonstrate that our method consistently and significantly outperforms existing methods. For example, with 1 sample per category, it relatively improves the average result by 9.1% compared to human-crafted prompts. Yuning Lu, Jianzhuang Liu, Yonggang Zhang 0003, Xinmei Tian 0001 |
CVPR | 2 |
| 2022 | Uformer: A General U-Shaped Transformer for Image RestorationabstractIn this paper, we present Uformer, an effective and efficient Transformer-based architecture for image restoration, in which we build a hierarchical encoder-decoder network using the Transformer block. In Uformer, there are two core designs. First, we introduce a novel locally-enhanced window (LeWin) Transformer block, which performs non-overlapping window-based self-attention instead of global self-attention. It significantly reduces the computational complexity on high resolution feature map while capturing local context. Second, we propose a learnable multi-scale restoration modulator in the form of a multi-scale spatial bias to adjust features in multiple layers of the Uformer decoder. Our modulator demonstrates superior capability for restoring details for various image restoration tasks while introducing marginal extra parameters and computational cost. Powered by these two designs, Uformer enjoys a high capability for capturing both local and global dependencies for image restoration. To evaluate our approach, extensive experiments are conducted on several image restoration tasks, including image denoising, motion deblurring, defocus deblurring and deraining. Without bells and whistles, our Uformer achieves superior or comparable performance compared with the state-of-the-art algorithms. The code and models are available at https://github.com/ZhendongWang6/Uformer. Xiaodong Cun, Jianmin Bao, Wengang Zhou 0001, Jianzhuang Liu, Houqiang Li |
CVPR | 5 |
| 2022 | Neural Architecture Search with Representation Mutual InformationabstractPerformance evaluation strategy is one of the most important factors that determine the effectiveness and efficiency in Neural Architecture Search (NAS). Existing strategies, such as employing standard training or performance predictor, often suffer from high computational complexity and low generality. To address this issue, we propose to rank architectures by Representation Mutual Information (RMI). Specifically, given an arbitrary architecture that has decent accuracy, architectures that have high RMI with it always yield good accuracies. As an accurate performance indicator to facilitate NAS, RMI not only generalizes well to different search spaces, but is also efficient enough to evaluate architectures using only one batch of data. Building upon RMI, we further propose a new search algorithm termed RMI-NAS, facilitating with a theorem to guarantee the global optimal of the searched architecture. In particular, RMI-NAS first randomly samples architectures from the search space, which are then effectively classified as positive or negative samples by RMI. We then use these samples to train a random forest to explore new regions, while keeping track of the distribution of positive architectures. When the sample size is sufficient, the architecture with the largest probability from the aforementioned distribution is selected, which is theoretically proved to be the optimal solution. The architectures searched by our method achieve remarkable top-1 accuracies with the magnitude times faster search process. Besides, RMI-NAS also generalizes to different datasets and search spaces. Our code has been made available at https://git.openi.org.cn/PCL_AutoML/XNAS. Xiawu Zheng, Lei Zhang 0001, Chenglin Wu 0001, Fei Chao 0001, Jianzhuang Liu, Wei Zeng 0006, Yonghong Tian 0001, Rongrong Ji |
CVPR | 6 |
| 2022 | IntraQ: Learning Synthetic Images with Intra-Class Heterogeneity for Zero-Shot Network QuantizationabstractLearning to synthesize data has emerged as a promising direction in zero-shot quantization (ZSQ), which represents neural networks by low-bit integer without accessing any of the real data. In this paper, we observe an interesting phenomenon of intra-class heterogeneity in real data and show that existing methods fail to retain this property in their synthetic images, which causes a limited performance increase. To address this issue, we propose a novel zero-shot quantization method referred to as IntraQ. First, we propose a local object reinforcement that locates the target objects at different scales and positions of the synthetic images. Second, we introduce a marginal distance constraint to form class-related features distributed in a coarse area. Lastly, we devise a soft inception loss which injects a soft prior label to prevent the synthetic images from being over-fitting to a fixed object. Our IntraQ is demonstrated to well retain the intra-class heterogeneity in the synthetic images and also observed to perform state-of-the-art. For example, compared to the advanced ZSQ, our IntraQ obtains 9.17% increase of the top-1 accuracy on ImageNet when all layers of MobileNetV1 are quantized to 4-bit. Code is at https://github.com/zysxmu/IntraQ Yunshan Zhong, Mingbao Lin, Gongrui Nan, Jianzhuang Liu, Baochang Zhang 0001, Yonghong Tian 0001, Rongrong Ji |
CVPR | 4 |
| 2022 | CLIFF: Carrying Location Information in Full Frames into Human Pose and Shape Estimation
Zhihao Li 0002, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, Youliang Yan |
ECCV (5) | 2 |
| 2022 | Self-Supervision Can Be a Good Few-Shot Learner
Yuning Lu, Liangjian Wen, Jianzhuang Liu, Xinmei Tian 0001 |
ECCV (19) | 3 |
| 2022 | Anti-retroactive Interference for Lifelong Learning
Runqi Wang, Yuxiang Bao, Baochang Zhang 0001, Jianzhuang Liu, Wentao Zhu 0001, Guodong Guo |
ECCV (24) | 4 |
| 2022 | RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive LearningabstractConventional visual relationship detection models only use the numeric ids of relation labels for training, but ignore the semantic correlation between the labels, which leads to severe training biases and harms the generalization ability of representations.In this paper, we introduce compact language information of relation labels for regularizing the representation learning of visual relations.Specifically, we propose a simple yet effective visual Relationship prediction framework that transfers natural language knowledge learned from Contrastive Language-Image Pre-training (CLIP) models to enhance the relationship prediction, termed as RelCLIP.Benefiting from the powerful visual-semantic alignment ability of CLIP at image level, we introduce a novel Relational Contrastive Learning (RCL) approach that explores relation-level visual-semantic alignment via learning to match cross-modal relational embeddings.By collaboratively learning the semantic coherence and discrepancy from relation triplets, the model can generate more discriminative and robust representations.Experimental results on the Visual Genome dataset show that RelCLIP achieves significant improvements over strong baselines under full (providing accurate labels) and distant supervision (providing noise labels), demonstrating its powerful generalization ability in learning relationship representations. Yi Zhu 0004, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang, Feng Zhao 0004, Jianzhuang Liu |
EMNLP | 6 |
| 2022 | LayouTransformer: Generating Layout Patterns with Transformer via Sequential Pattern ModelingabstractGenerating legal and diverse layout patterns to establish large pattern libraries is fundamental for many lithography design applications. Existing pattern generation models typically regard the pattern generation problem as image generation of layout maps and learn to model the patterns via capturing pixel-level coherence, which is insufficient to achieve polygon-level modeling, e.g., shape and layout of patterns, thus leading to poor generation quality. In this paper, we regard the pattern generation problem as an unsupervised sequence generation problem, in order to learn the pattern design rules by explicitly modeling the shapes of polygons and the layouts among polygons. Specifically, we first propose a sequential pattern representation scheme that fully describes the geometric information of polygons by encoding the 2D layout patterns as sequences of tokens, i.e., vertexes and edges. Then we train a sequential generative model to capture the long-term dependency among tokens and thus learn the design rules from training examples. To generate a new pattern in sequence, each token is generated conditioned on the previously generated tokens that are from the same polygon or different polygons in the same layout map. Our framework, termed LayouTransformer, is based on the Transformer architecture due to its remarkable ability in sequence modeling. Comprehensive experiments show that our LayouTransformer not only generates a large amount of legal patterns but also maintains high generation diversity, demonstrating its superiority over existing pattern generative models. Liangjian Wen, Yi Zhu 0004, Guojin Chen, Bei Yu 0001, Jianzhuang Liu, Chunjing Xu |
ICCAD | 6 |
| 2022 | Structure-Preserving Graph Representation LearningabstractThough graph representation learning (GRL) has made significant progress, it is still a challenge to extract and embed the rich topological structure and feature information in an adequate way. Most existing methods focus on local structure and fail to fully incorporate the global topological structure. To this end, we propose a novel Structure-Preserving Graph Representation Learning (SPGRL) method, to fully capture the structure information of graphs. Specifically, to reduce the uncertainty and misinformation of the original graph, we construct a feature graph as a complementary view via k-Nearest Neighbor method. The feature graph can be used to contrast at node-level to capture the local relation. Besides, we retain the global topological structure information by maximizing the mutual information (MI) of the whole graph and feature embeddings, which is theoretically reduced to exchanging the feature embeddings of the feature and the original graphs to reconstruct themselves. Extensive experiments show that our method has quite superior performance on semi-supervised node classification task and excellent robustness under noise perturbation on graph structure or node features. The source code is available at https://github.com/uestc-lese/SPGRL. Ruiyi Fang, Liangjian Wen, Zhao Kang 0001, Jianzhuang Liu |
ICDM | 4 |
| 2022 | Refactoring ISP for High-Level Vision TasksabstractThe image signal processing (ISP) pipeline, which transforms raw sensor measurement to a color image, is composed of a sequence of processing modules. Traditionally, the ISP pipeline is manually tuned by experts for human perception. The resulting handcrafted ISP configuration does not necessarily benefit the downstream high-level vision tasks. To mitigate these problems, this paper presents a simple yet effective framework based on Evolutionary Algorithm to search for a set of compact ISP configurations for high-level vision tasks. In particular, we encode ISP structure into a binary string and ISP parameters into a set of float numbers. Then we jointly optimize them with task-specific loss and ISP computation budgets (e.g., running time) through solving a nonlinear multi-objective optimization problem. By mutating the configurations of the ISP pipeline, we are able to remove redundant modules and design an ISP with both low cost and high accuracy. We validate the proposed method on extreme noisy and low-light raw images, and experimental results show that our framework can help find effective and efficient ISP configurations for both object detection and semantic segmentation tasks. We further provide a detailed analysis on the importance of different modules in the ISP configurations, which benefits the design of ISP for downstream tasks in the future. Yongjie Shi, Songjiang Li, Xu Jia 0012, Jianzhuang Liu |
ICRA | 4 |
| 2022 | FNeVR: Neural Volume Rendering for Face AnimationabstractFace animation, one of the hottest topics in computer vision, has achieved a promising performance with the help of generative models. However, it remains a critical challenge to generate identity preserving and photo-realistic images due to the sophisticated motion deformation and complex facial detail modeling. To address these problems, we propose a Face Neural Volume Rendering (FNeVR) network to fully explore the potential of 2D motion warping and 3D volume rendering in a unified framework. In FNeVR, we design a 3D Face Volume Rendering (FVR) module to enhance the facial details for image rendering. Specifically, we first extract 3D information with a well designed architecture, and then introduce an orthogonal adaptive ray-sampling module for efficient rendering. We also design a lightweight pose editor, enabling FNeVR to edit the facial pose in a simple yet effective way. Extensive experiments show that our FNeVR obtains the best overall quality and performance on widely used talking-head benchmarks. Bohan Zeng, Hong Li 0016, Xuhui Liu, Jianzhuang Liu, Dapeng Chen, Wei Peng 0011, Baochang Zhang 0001 |
NeurIPS | 5 |
| 2022 | CoupAlign: Coupling Word-Pixel with Sentence-Mask Alignments for Referring Image SegmentationabstractReferring image segmentation aims at localizing all pixels of the visual objects described by a natural language sentence. Previous works learn to straightforwardly align the sentence embedding and pixel-level embedding for highlighting the referred objects, but ignore the semantic consistency of pixels within the same object, leading to incomplete masks and localization errors in predictions. To tackle this problem, we propose CoupAlign, a simple yet effective multi-level visual-semantic alignment method, to couple sentence-mask alignment with word-pixel alignment to enforce object mask constraint for achieving more accurate localization and segmentation. Specifically, the Word-Pixel Alignment (WPA) module performs early fusion of linguistic and pixel-level features in intermediate layers of the vision and language encoders. Based on the word-pixel aligned embedding, a set of mask proposals are generated to hypothesize possible objects. Then in the Sentence-Mask Alignment (SMA) module, the masks are weighted by the sentence embedding to localize the referred object, and finally projected back to aggregate the pixels for the target. To further enhance the learning of the two alignment modules, an auxiliary loss is designed to contrast the foreground and background pixels. By hierarchically aligning pixels and masks with linguistic features, our CoupAlign captures the pixel coherence at both visual and semantic levels, thus generating more accurate predictions. Extensive experiments on popular datasets (e.g., RefCOCO and G-Ref) show that our method achieves consistent improvements over state-of-the-art methods, e.g., about 2% oIoU increase on the validation and testing set of RefCOCO. Especially, CoupAlign has remarkable ability in distinguishing the target from multiple objects of the same class. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/CoupAlign. Yi Zhu 0004, Jianzhuang Liu, Xiaodan Liang, Wei Ke 0003 |
NeurIPS | 3 |
| 2022 | Neighbor2Neighbor: A Self-Supervised Framework for Deep Image DenoisingabstractIn recent years, image denoising has benefited a lot from deep neural networks. However, these models need large amounts of noisy-clean image pairs for supervision. Although there have been attempts in training denoising networks with only noisy images, existing self-supervised algorithms suffer from inefficient network training, heavy computational burden, or dependence on noise modeling. In this paper, we proposed a self-supervised framework named Neighbor2Neighbor for deep image denoising. We develop a theoretical motivation and prove that by designing specific samplers for training image pairs generation from only noisy images, we can train a self-supervised denoising network similar to the network trained with clean images supervision. Besides, we propose a regularizer in the perspective of optimization to narrow the optimization gap between the self-supervised denoiser and the supervised denoiser. We present a very simple yet effective self-supervised training scheme based on the theoretical understandings: training image pairs are generated by random neighbor sub-samplers, and denoising networks are trained with a regularized loss. Moreover, we propose a training strategy named BayerEnsemble to adapt the Neighbor2Neighbor framework in raw image denoising. The proposed Neighbor2Neighbor framework can enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. It also avoids heavy dependence on the assumption of the noise distribution. We evaluate the Neighbor2Neighbor framework through extensive experiments, including synthetic experiments with different noise distributions and real-world experiments under various scenarios. The code is available online: https://github.com/TaoHuang2018/Neighbor2Neighbor. Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu |
IEEE Trans. Image Process. | 5 |
| 2022 | Feature Calibration Network for Occluded Pedestrian DetectionabstractPedestrian detection in the wild remains a challenging problem especially for scenes containing serious occlusion. In this paper, we propose a novel feature learning method in the deep learning framework, referred to as Feature Calibration Network (FC-Net), to adaptively detect pedestrians under various occlusions. FC-Net is based on the observation that the visible parts of pedestrians are selective and decisive for detection, and is implemented as a self-paced feature learning framework with a self-activation (SA) module and a feature calibration (FC) module. In a new self-activated manner, FC-Net learns features which highlight the visible parts and suppress the occluded parts of pedestrians. The SA module estimates pedestrian activation maps by reusing classifier weights, without any additional parameter involved, therefore resulting in an extremely parsimony model to reinforce the semantics of features, while the FC module calibrates the convolutional features for adaptive pedestrian representation in both pixel-wise and region-based ways. Experiments on CityPersons and Caltech datasets demonstrate that FC-Net improves detection performance on occluded pedestrians up to 10% while maintaining excellent performance on non-occluded instances. Tianliang Zhang 0003, Qixiang Ye, Baochang Zhang 0001, Jianzhuang Liu, Xiaopeng Zhang 0008, Qi Tian 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Filter Sketch for Network PruningabstractWe propose a novel network pruning approach by information preserving of pretrained network weights (filters). Network pruning with the information preserving is formulated as a matrix sketch problem, which is efficiently solved by the off-the-shelf frequent direction method. Our approach, referred to as FilterSketch, encodes the second-order information of pretrained weights, which enables the representation capacity of pruned networks to be recovered with a simple fine-tuning procedure. FilterSketch requires neither training from scratch nor data-driven iterative optimization, leading to a several-orders-of-magnitude reduction of time cost in the optimization of pruning. Experiments on CIFAR-10 show that FilterSketch reduces 63.3% of floating-point operations (FLOPs) and prunes 59.9% of network parameters with negligible accuracy cost for ResNet-110. On ILSVRC-2012, it reduces 45.5% of FLOPs and removes 43.0% of parameters with only 0.69% accuracy drop for ResNet-50. Our code and pruned models can be found at https://github.com/lmbxmu/FilterSketch. Mingbao Lin, Liujuan Cao, Qixiang Ye, Yonghong Tian 0001, Jianzhuang Liu, Qi Tian 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Dual Distribution Alignment Network for Generalizable Person Re-IdentificationabstractDomain generalization (DG) offers a preferable real-world setting for Person Re-Identification (Re-ID), which trains a model using multiple source domain datasets and expects it to perform well in an unseen target domain without any model updating. Unfortunately, most DG approaches are designed explicitly for classification tasks, which fundamentally differs from the retrieval task Re-ID. Moreover, existing applications of DG in Re-ID cannot correctly handle the massive variation among Re-ID datasets. In this paper, we identify two fundamental challenges in DG for Person Re-ID: domain-wise variations and identity-wise similarities. To this end, we propose an end-to-end Dual Distribution Alignment Network (DDAN) to learn domain-invariant features with dual-level constraints: the domain-wise adversarial feature learning and the identity-wise similarity enhancement. These constraints effectively reduce the domain-shift among multiple source domains further while agreeing to real-world scenarios. We evaluate our method in a large-scale DG Re-ID benchmark and compare it with various cutting-edge DG approaches. Quantitative results show that DDAN achieves state-of-the-art performance. Peixian Chen, Pingyang Dai, Jianzhuang Liu, Feng Zheng 0001, Mingliang Xu 0001, Qi Tian 0001, Rongrong Ji |
AAAI | 3 |
| 2021 | Domain General Face Forgery Detection by Learning to WeightabstractIn this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making it challenging to detect fake faces in unseen domains. We argue that different faces contribute differently to a detection model trained on multiple domains, making the model likely to fit domain-specific biases. As such, we propose the LTW approach based on the meta-weight learning algorithm, which configures different weights for face images from different domains. The LTW network can balance the model's generalizability across multiple domains. Then, the meta-optimization calibrates the source domain's gradient enabling more discriminative features to be learned. The detection ability of the network is further improved by introducing an intra-class compact loss. Extensive experiments on several commonly used deepfake datasets to demonstrate the effectiveness of our method in detecting synthetic faces. Code and supplemental material are available at https://github.com/skJack/LTW. Ke Sun 0016, Hong Liu 0009, Qixiang Ye, Yue Gao 0002, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji |
AAAI | 5 |
| 2021 | Semi-Supervised Domain Adaptation Based on Dual-Level Domain Mixing for Semantic SegmentationabstractData-driven based approaches, in spite of great success in many tasks, have poor generalization when applied to unseen image domains, and require expensive cost of annotation especially for dense pixel prediction tasks such as semantic segmentation. Recently, both unsupervised domain adaptation (UDA) from large amounts of synthetic data and semi-supervised learning (SSL) with small set of labeled data have been studied to alleviate this issue. However, there is still a large gap on performance compared to their supervised counterparts. We focus on a more practical setting of semi-supervised domain adaptation (SSDA) where both a small set of labeled target data and large amounts of labeled source data are available. To address the task of SSDA, a novel framework based on dual-level domain mixing is proposed. The proposed framework consists of three stages. First, two kinds of data mixing methods are proposed to reduce domain gap in both region-level and sample-level respectively. We can obtain two complementary domain-mixed teachers based on dual-level mixed data from holistic and partial views respectively. Then, a student model is learned by distilling knowledge from these two teachers. Finally, pseudo labels of unlabeled data are generated in a self-training manner for another few rounds of teachers training. Extensive experimental results have demonstrated the effectiveness of our proposed framework on synthetic-to-real semantic segmentation benchmarks. Shuaijun Chen, Xu Jia 0012, Yongjie Shi, Jianzhuang Liu |
CVPR | 5 |
| 2021 | Multi-Source Domain Adaptation With Collaborative Learning for Semantic SegmentationabstractMulti-source unsupervised domain adaptation (MSDA) aims at adapting models trained on multiple labeled source domains to an unlabeled target domain. In this paper, we propose a novel multi-source domain adaptation framework based on collaborative learning for semantic segmentation. Firstly, a simple image translation method is introduced to align the pixel value distribution to reduce the gap between source domains and target domain to some extent. Then, to fully exploit the essential semantic information across source domains, we propose a collaborative learning method for domain adaptation without seeing any data from target domain. In addition, similar to the setting of unsupervised domain adaptation, unlabeled target domain data is leveraged to further improve the performance of domain adaptation. This is achieved by additionally constraining the outputs of multiple adaptation models with pseudo labels online generated by an ensembled model. Extensive experiments and ablation studies are conducted on the widely-used domain adaptation benchmark datasets in semantic segmentation. Our proposed method achieves 59.0% mIoU on the validation set of Cityscapes by training on the labeled Synscapes and GTA5 datasets and unlabeled training set of Cityscapes. It significantly outperforms all previous state-of-the-arts single-source and multi-source unsupervised domain adaptation methods. Xu Jia 0012, Shuaijun Chen, Jianzhuang Liu |
CVPR | 4 |
| 2021 | Neighbor2Neighbor: Self-Supervised Denoising From Single Noisy ImagesabstractIn the last few years, image denoising has benefited a lot from the fast development of neural networks. However, the requirement of large amounts of noisy-clean image pairs for supervision limits the wide use of these models. Although there have been a few attempts in training an image denoising model with only single noisy images, existing self-supervised denoising approaches suffer from inefficient network training, loss of useful information, or dependence on noise modeling. In this paper, we present a very simple yet effective method named Neighbor2Neighbor to train an effective image denoising model with only noisy images. Firstly, a random neighbor sub-sampler is proposed for the generation of training image pairs. In detail, input and target used to train a network are images sub-sampled from the same noisy image, satisfying the requirement that paired pixels of paired images are neighbors and have very similar appearance with each other. Secondly, a denoising network is trained on sub-sampled training pairs generated in the first stage, with a proposed regularizer as additional loss for better performance. The proposed Neighbor2Neighbor framework is able to enjoy the progress of state-of-the-art supervised denoising networks in network architecture design. Moreover, it avoids heavy dependence on the assumption of the noise distribution. We explain our approach from a theoretical perspective and further validate it through extensive experiments, including synthetic experiments with different noise distributions in sRGB space and real-world experiments on a denoising benchmark dataset in raw-RGB space. Tao Huang 0022, Songjiang Li, Xu Jia 0012, Huchuan Lu, Jianzhuang Liu |
CVPR | 5 |
| 2021 | Multi-Target Domain Adaptation With Collaborative Consistency LearningabstractRecently unsupervised domain adaptation for the semantic segmentation task has become more and more popular due to high-cost of pixel-level annotation on real-world images. However, most domain adaptation methods are only restricted to single-source-single-target pair, and can not be directly extended to multiple target domains. In this work, we propose a collaborative learning framework to achieve unsupervised multi-target domain adaptation. An unsupervised domain adaptation expert model is first trained for each source-target pair and is further encouraged to collaborate with each other through a bridge built between different target domains. These expert models are further improved by adding the regularization of making the consistent pixel-wise prediction for each sample with the same structured context. To obtain a single model that works across multiple target domains, we propose to simultaneously learn a student model which is trained to not only imitate the output of each expert on the corresponding target domain, but also to pull different expert close to each other with regularization on their weights. Extensive experiments demonstrate that the proposed method can effectively exploit rich structured information contained in both labeled source domain and multiple unlabeled target domains. Not only does it perform well across multiple target domains but also performs favorably against state-of-the-art unsupervised domain adaptation methods specially trained on a single source-target pair. Code is available at https://github.com/junpan19/MTDA. Takashi Isobe, Xu Jia 0012, Shuaijun Chen, Yongjie Shi, Jianzhuang Liu, Huchuan Lu, Shengjin Wang |
CVPR | 6 |
| 2021 | Towards Compact CNNs via Collaborative CompressionabstractChannel pruning and tensor decomposition have received extensive attention in convolutional neural network compression. However, these two techniques are traditionally deployed in an isolated manner, leading to significant accuracy drop when pursuing high compression rates. In this paper, we propose a Collaborative Compression (CC) scheme, which joints channel pruning and tensor decomposition to compress CNN models by simultaneously learning the model sparsity and low-rankness. Specifically, we first investigate the compression sensitivity of each layer in the network, and then propose a Global Compression Rate Optimization that transforms the decision problem of compression rate into an optimization problem. After that, we propose multi-step heuristic compression to remove redundant compression units step-by-step, which fully considers the effect of the remaining compression space (i.e., unremoved compression units). Our method demonstrates superior performance gains over previous ones on various datasets and backbone architectures. For example, we achieve 52.9% FLOPs reduction by removing 48.4% parameters on ResNet-50 with only a Top-1 accuracy drop of 0.56% on ImageNet 2012. Shaohui Lin, Jianzhuang Liu, Qixiang Ye, Mengdi Wang 0001, Fei Chao 0001, Fan Yang 0016, Jincheng Ma, Qi Tian 0001, Rongrong Ji |
CVPR | 3 |
| 2021 | Multiple Instance Active Learning for Object DetectionabstractDespite the substantial progress of active learning for image recognition, there still lacks an instance-level active learning method specified for object detection. In this paper, we propose Multiple Instance Active Object Detection (MI-AOD), to select the most informative images for detector training by observing instance-level uncertainty. MI-AOD defines an instance uncertainty learning module, which leverages the discrepancy of two adversarial instance classifiers trained on the labeled set to predict instance uncertainty of the unlabeled set. MI-AOD treats unlabeled images as instance bags and feature anchors in images as instances, and estimates the image uncertainty by re-weighting instances in a multiple instance learning (MIL) fashion. Iterative instance uncertainty learning and re-weighting facilitate suppressing noisy instances, toward bridging the gap between instance uncertainty and image-level uncertainty. Experiments validate that MI-AOD sets a solid baseline for instance-level active learning. On commonly used object detection datasets, MI-AOD outperforms state-of-the-art methods with significant margins, particularly when the labeled sets are small. Code is available at https://github.com/yuantn/MI-AOD. Tianning Yuan, Fang Wan 0001, Mengying Fu, Jianzhuang Liu, Songcen Xu, Xiangyang Ji, Qixiang Ye |
CVPR | 4 |
| 2021 | Occlude Them All: Occlusion-Aware Attention Network for Occluded Person Re-IDabstractPerson Re-Identification (ReID) has achieved remarkable performance along with the deep learning era. However, most approaches carry out ReID only based upon holistic pedestrian regions. In contrast, real-world scenarios involve occluded pedestrians, which provide partial visual appearances and destroy the ReID accuracy. A common strategy is to locate visible body parts by auxiliary model, which however suffers from significant domain gaps and data bias issues. To avoid such problematic models in occluded person ReID, we propose the Occlusion-Aware Mask Network (OAMN). In particular, we incorporate an attention-guided mask module, which requires guidance from labeled occlusion data. To this end, we propose a novel occlusion augmentation scheme that produces diverse and precisely labeled occlusion for any holistic dataset. The proposed scheme suits real-world scenarios better than existing schemes, which consider only limited types of occlusions. We also offer a novel occlusion unification scheme to tackle ambiguity information at the test phase. The above three components enable existing attention mechanisms to precisely capture body parts regardless of the occlusion. Comprehensive experiments on a variety of person ReID benchmarks demonstrate the superiority of OAMN over state-of-the-arts. Peixian Chen, Pingyang Dai, Jianzhuang Liu, Qixiang Ye, Mingliang Xu 0001, Qi'an Chen, Rongrong Ji |
ICCV | 4 |
| 2021 | Active Learning for Lane Detection: A Knowledge Distillation ApproachabstractLane detection is a key task for autonomous driving vehicles. Currently, lane detection relies on a huge amount of annotated images, which is a heavy burden. Active learning has been proposed to reduce annotation in many computer vision tasks, but no effort has been made for lane detection. Through experiments, we find that existing active learning methods perform poorly for lane detection, and the reasons are twofold. On one hand, most methods evaluate data uncertainties based on entropy, which is undesirable in lane detection because it encourages to select images with very few lanes or even no lane at all. On the other hand, existing methods are not aware of the noise of lane annotations, which is caused by heavy occlusion and unclear lane marks. In this paper, we build a novel knowledge distillation framework and evaluate the uncertainty of images based on the knowledge learnt by the student model. We show that the proposed uncertainty metric overcomes the above two problems. To reduce data redundancy, we explore the influence sets of image samples, and propose a new diversity metric for data selection. Finally we incorporate the uncertainty and diversity metrics, and develop a greedy algorithm for data selection. The experiments show that our method achieves new state-of-the-art on the lane detection benchmarks. In addition, we extend this method to common 2D object detection and the results show that it is also effective. Fengchao Peng, Jianzhuang Liu, Zhen Yang 0008 |
ICCV | 3 |
| 2021 | Motion Deblurring with Real EventsabstractIn this paper, we propose an end-to-end learning framework for event-based motion deblurring in a self-supervised manner, where real-world events are exploited to alleviate the performance degradation caused by data inconsistency. To achieve this end, optical flows are predicted from events, with which the blurry consistency and photometric consistency are exploited to enable self-supervision on the deblurring network with real-world data. Furthermore, a piecewise linear motion model is proposed to take into account motion non-linearities and thus leads to an accurate model for the physical formation of motion blurs in the real-world scenario. Extensive evaluation on both synthetic and real motion blur datasets demonstrates that the proposed algorithm bridges the gap between simulated and real-world motion blurs and shows remarkable performance for eventbased motion deblurring in real-world scenarios. Lei Yu 0006, Bishan Wang, Wen Yang 0001, Gui-Song Xia, Xu Jia 0012, Zhendong Qiao, Jianzhuang Liu |
ICCV | 8 |
| 2021 | ReCU: Reviving the Dead Weights in Binary Neural Networks
Mingbao Lin, Jianzhuang Liu, Jie Chen 0001, Ling Shao 0001, Yue Gao 0002, Yonghong Tian 0001, Rongrong Ji |
ICCV | 3 |
| 2021 | TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringabstractDue to the superior ability of global dependency modeling, Transformer and its variants have become the primary choice of many vision-and-language tasks. However, in tasks like Visual Question Answering (VQA) and Referring Expression Comprehension (REC), the multimodal prediction often requires visual information from macro- to micro-views. Therefore, how to dynamically schedule the global and local dependency modeling in Transformer has become an emerging issue. In this paper, we propose an example-dependent routing scheme called TRAnsformer Routing (TRAR) to address this issue1. Specifically, in TRAR, each visual Transformer layer is equipped with a routing module with different attention spans. The model can dynamically select the corresponding attentions based on the output of the previous inference step, so as to formulate the optimal routing path for each example. Notably, with careful designs, TRAR can reduce the additional computation and memory overhead to almost negligible. To validate TRAR, we conduct extensive experiments on five benchmark datasets of VQA and REC, and achieve superior performance gains than the standard Transformers and a bunch of state-of-the-art methods. Yiyi Zhou, Tianhe Ren, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu 0001, Rongrong Ji |
ICCV | 5 |
| 2021 | Binarized Neural Architecture Search for Efficient Object Recognition
Lian Zhuo, Baochang Zhang 0001, Xiawu Zheng, Jianzhuang Liu, Rongrong Ji, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 5 |
| 2021 | Rectified Binary Convolutional Networks with Generative Adversarial Learning
Chunlei Liu 0001, Wenrui Ding, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, David S. Doermann |
Int. J. Comput. Vis. | 5 |
| 2021 | Real-time semantic segmentation via sequential knowledge distillation
Jipeng Wu, Rongrong Ji, Jianzhuang Liu, Mingliang Xu 0001, Jiawen Zheng, Ling Shao 0001, Qi Tian 0001 |
Neurocomputing | 3 |
| 2021 | An End-to-End Foreground-Aware Network for Person Re-IdentificationabstractPerson re-identification is a crucial task of identifying pedestrians of interest across multiple surveillance camera views. For person re-identification, a pedestrian is usually represented with features extracted from a rectangular image region that inevitably contains the scene background, which incurs ambiguity to distinguish different pedestrians and degrades the accuracy. Thus, we propose an end-to-end foreground-aware network to discriminate against the foreground from the background by learning a soft mask for person re-identification. In our method, in addition to the pedestrian ID as supervision for the foreground, we introduce the camera ID of each pedestrian image for background modeling. The foreground branch and the background branch are optimized collaboratively. By presenting a target attention loss, the pedestrian features extracted from the foreground branch become more insensitive to backgrounds, which greatly reduces the negative impact of changing backgrounds on pedestrian matching across different camera views. Notably, in contrast to existing methods, our approach does not require an additional dataset to train a human landmark detector or a segmentation model for locating the background regions. The experimental results conducted on three challenging datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17, demonstrate the effectiveness of our approach. Wengang Zhou 0001, Jianzhuang Liu, Guo-Jun Qi, Qi Tian 0001, Houqiang Li |
IEEE Trans. Image Process. | 3 |
| 2021 | Light Field View Synthesis via Aperture Disparity and Warping Confidence MapabstractThis paper presents a learning-based approach to synthesize the view from an arbitrary camera position given a sparse set of images. A key challenge for this novel view synthesis arises from the reconstruction process, when the views from different input images may not be consistent due to obstruction in the light path. We overcome this by jointly modeling the epipolar property and occlusion in designing a convolutional neural network. We start by defining and computing the aperture disparity map, which approximates the parallax and measures the pixel-wise shift between two views. While this relates to free-space rendering and can fail near the object boundaries, we further develop a warping confidence map to address pixel occlusion in these challenging regions. The proposed method is evaluated on diverse real-world and synthetic light field scenes, and it shows better performance over several state-of-the-art techniques. Nan Meng, Jianzhuang Liu, Edmund Y. Lam |
IEEE Trans. Image Process. | 3 |
| 2021 | Interaction-Integrated Network for Natural Language Moment LocalizationabstractNatural language moment localization aims at localizing video clips according to a natural language description. The key to this challenging task lies in modeling the relationship between verbal descriptions and visual contents. Existing approaches often sample a number of clips from the video, and individually determine how each of them is related to the query sentence. However, this strategy can fail dramatically, in particular when the query sentence refers to some visual elements that appear outside of, or even are distant from, the target clip. In this paper, we address this issue by designing an Interaction-Integrated Network (I2N), which contains a few Interaction-Integrated Cells (I2Cs). The idea lies in the observation that the query sentence not only provides a description to the video clip, but also contains semantic cues on the structure of the entire video. Based on this, I2Cs go one step beyond modeling short-term contexts in the time domain by encoding long-term video content into every frame feature. By stacking a few I2Cs, the obtained network, I2N, enjoys an improved ability of inference, brought by both (I) multi-level correspondence between vision and language and (II) more accurate cross-modal alignment. When evaluated on a challenging video moment localization dataset named DiDeMo, I2N outperforms the state-of-the-art approach by a clear margin of 1.98%. On other two challenging datasets, Charades-STA and TACoS, I2N also reports competitive performance. Lingxi Xie, Jianzhuang Liu, Fei Wu 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Binarized Neural Architecture SearchabstractNeural architecture search (NAS) can have a significant impact in computer vision by automatically designing optimal neural network architectures for various tasks. A variant, binarized neural architecture search (BNAS), with a search space of binarized convolutions, can produce extremely compressed models. Unfortunately, this area remains largely unexplored. BNAS is more challenging than NAS due to the learning inefficiency caused by optimization requirements and the huge architecture space. To address these issues, we introduce channel sampling and operation space reduction into a differentiable NAS to significantly reduce the cost of searching. This is accomplished through a performance-based strategy used to abandon less potential operations. Two optimization methods for binarized neural networks are used to validate the effectiveness of our BNAS. Extensive experiments demonstrate that the proposed BNAS achieves a performance comparable to NAS on both CIFAR and ImageNet databases. An accuracy of 96.53% vs. 97.22% is achieved on the CIFAR-10 dataset, but with a significantly compressed model, and a 40% faster search than the state-of-the-art PC-DARTS. Lian Zhuo, Baochang Zhang 0001, Xiawu Zheng, Jianzhuang Liu, David S. Doermann, Rongrong Ji |
AAAI | 5 |
| 2020 | High-Order Residual Network for Light Field Super-ResolutionabstractPlenoptic cameras usually sacrifice the spatial resolution of their SAIs to acquire geometry information from different viewpoints. Several methods have been proposed to mitigate such spatio-angular trade-off, but seldom make use of the structural properties of the light field (LF) data efficiently. In this paper, we propose a novel high-order residual network to learn the geometric features hierarchically from the LF for reconstruction. An important component in the proposed network is the high-order residual block (HRB), which learns the local geometric features by considering the information from all input views. After fully obtaining the local features learned from each HRB, our model extracts the representative geometric features for spatio-angular upsampling through the global residual learning. Additionally, a refinement network is followed to further enhance the spatial details by minimizing a perceptual loss. Compared with previous work, our model is tailored to the rich structure inherent in the LF, and therefore can reduce the artifacts near non-Lambertian and occlusion regions. Experimental results show that our approach enables high-quality reconstruction even in challenging regions and outperforms state-of-the-art single image or LF reconstruction methods with both quantitative measurements and visual evaluation. Nan Meng, Jianzhuang Liu, Edmund Y. Lam |
AAAI | 3 |
| 2020 | Context-Transformer: Tackling Object Confusion for Few-Shot DetectionabstractFew-shot object detection is a challenging but realistic scenario, where only a few annotated training images are available for training detectors. A popular approach to handle this problem is transfer learning, i.e., fine-tuning a detector pretrained on a source-domain benchmark. However, such transferred detector often fails to recognize new objects in the target domain, due to low data diversity of training samples. To tackle this problem, we propose a novel Context-Transformer within a concise deep transfer framework. Specifically, Context-Transformer can effectively leverage source-domain object knowledge as guidance, and automatically exploit contexts from only a few training images in the target domain. Subsequently, it can adaptively integrate these relational clues to enhance the discriminative power of detector, in order to reduce object confusion in few-shot scenarios. Moreover, Context-Transformer is flexibly embedded in the popular SSD-style detectors, which makes it a plug-and-play module for end-to-end few-shot learning. Finally, we evaluate Context-Transformer on the challenging settings of few-shot detection and incremental few-shot detection. The experimental results show that, our framework outperforms the recent state-of-the-art approaches. Ze Yang 0002, Yali Wang 0001, Jianzhuang Liu, Yu Qiao 0001 |
AAAI | 4 |
| 2020 | SketchyCOCO: Image Generation From Freehand Scene SketchesabstractWe introduce the first method for automatic image generation from scene-level freehand sketches. Our model allows for controllable image generation by specifying the synthesis goal via freehand sketches. The key contribution is an attribute vector bridged Generative Adversarial Network called EdgeGAN, which supports high visual-quality object-level image content generation without using freehand sketches as training data. We have built a large-scale composite dataset called SketchyCOCO to support and evaluate the solution. We validate our approach on the tasks of both object-level and scene-level image generation on SketchyCOCO. Through quantitative, qualitative results, human evaluation and ablation studies, we demonstrate the method's capacity to generate realistic complex scene-level images from various freehand sketches. Chengying Gao, Limin Wang 0002, Jianzhuang Liu, Changqing Zou |
CVPR | 5 |
| 2020 | Multiple Anchor Learning for Visual Object DetectionabstractClassification and localization are two pillars of visual object detectors. However, in CNN-based detectors, these two modules are usually optimized under a fixed set of candidate (or anchor) bounding boxes. This configuration significantly limits the possibility to jointly optimize classification and localization. In this paper, we propose a Multiple Instance Learning (MIL) approach that selects anchors and jointly optimizes the two modules of a CNN-based object detector. Our approach, referred to as Multiple Anchor Learning (MAL), constructs anchor bags and selects the most representative anchors from each bag. Such an iterative selection process is potentially NP-hard to optimize. To address this issue, we solve MAL by repetitively depressing the confidence of selected anchors by perturbing their corresponding features. In an adversarial selection-depression manner, MAL not only pursues optimal solutions but also fully leverages multiple anchors/features to learn a detection model. Experiments show that MAL improves the baseline RetinaNet with significant margins on the commonly used MS-COCO object detection benchmark and achieves new state-of-the-art detection performance compared with recent methods. Wei Ke 0003, Tianliang Zhang 0003, Zeyi Huang, Qixiang Ye, Jianzhuang Liu |
CVPR | 5 |
| 2020 | Projection & Probability-Driven Black-Box AttackabstractGenerating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimensional space. In this paper, we propose Projection & Probability-driven Black-box Attack (PPBA) to tackle this problem by reducing the solution space and providing better optimization. For reducing the solution space, we first model the adversarial perturbation optimization problem as a process of recovering frequency-sparse perturbations with compressed sensing, under the setting that random noise in the low-frequency space is more likely to be adversarial. We then propose a simple method to construct a low-frequency constrained sensing matrix, which works as a plug-and-play projection matrix to reduce the dimensionality. Such a sensing matrix is shown to be flexible enough to be integrated into existing methods like NES and BanditsTD. For better optimization, we perform a random walk with a probability-driven strategy, which utilizes all queries over the whole progress to make full use of the sensing matrix for a less query budget. Extensive experiments show that our method requires at most 24% fewer queries with a higher attack success rate compared with state-of-the-art approaches. Finally, the attack method is evaluated on the real-world online service, i.e., Google Cloud Vision API, which further demonstrates our practical potentials. Jie Li 0052, Rongrong Ji, Hong Liu 0009, Jianzhuang Liu, Bineng Zhong 0001, Cheng Deng 0002, Qi Tian 0001 |
CVPR | 4 |
| 2020 | Joint Demosaicing and Denoising With Self GuidanceabstractUsually located at the very early stages of the computational photography pipeline, demosaicing and denoising play important parts in the modern camera image processing. Recently, some neural networks have shown the effectiveness in joint demosaicing and denoising (JDD). Most of them first decompose a Bayer raw image into a four-channel RGGB image and then feed it into a neural network. This practice ignores the fact that the green channels are sampled at a double rate compared to the red and the blue channels. In this paper, we propose a self-guidance network (SGNet), where the green channels are initially estimated and then works as a guidance to recover all missing values in the input image. In addition, as regions of different frequencies suffer different levels of degradation in image restoration. We propose a density-map guidance to help the model deal with a wide range of frequencies. Our model outperforms state-of-the-art joint demosaicing and denoising methods on four public datasets, including two real and two synthetic data sets. Finally, we also verify that our method obtains best results in joint demosaicing , denoising and super-resolution. Lin Liu 0016, Xu Jia 0012, Jianzhuang Liu, Qi Tian 0001 |
CVPR | 3 |
| 2020 | Noise-Aware Fully Webly Supervised Object DetectionabstractWe investigate the emerging task of learning object detectors with sole image-level labels on the web without requiring any other supervision like precise annotations or additional images from well-annotated benchmark datasets. Such a task, termed as fully webly supervised object detection, is extremely challenging, since image-level labels on the web are always noisy, leading to poor performance of the learned detectors. In this work, we propose an end-to-end framework to jointly learn webly supervised detectors and reduce the negative impact of noisy labels. Such noise is heterogeneous, which is further categorized into two types, namely background noise and foreground noise. Regarding the background noise, we propose a residual learning structure incorporated with weakly supervised detection, which decomposes background noise and models clean data. To explicitly learn the residual feature between clean data and noisy labels, we further propose a spatially-sensitive entropy criterion, which exploits the conditional distribution of detection results to estimate the confidence of background categories being noise. Regarding the foreground noise, a bagging-mixup learning is introduced, which suppresses foreground noisy signals from incorrectly labelled images, whilst maintaining the diversity of training data. We evaluate the proposed approach on popular benchmark datasets by training detectors on web images, which are retrieved by the corresponding category tags from photo-sharing sites. Extensive experiments show that our method achieves significant improvements over the state-of-the-art methods. Yunhang Shen, Rongrong Ji, Xiaopeng Hong, Feng Zheng 0001, Jianzhuang Liu, Mingliang Xu 0001, Qi Tian 0001 |
CVPR | 6 |
| 2020 | API-Net: Robust Generative Classifier via a Single Discriminator
Xinshuai Dong, Hong Liu 0009, Rongrong Ji, Liujuan Cao, Qixiang Ye, Jianzhuang Liu, Qi Tian 0001 |
ECCV (13) | 6 |
| 2020 | Wavelet-Based Dual-Branch Network for Image Demoiréing
Lin Liu 0016, Jianzhuang Liu, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis, Wengang Zhou 0001, Qi Tian 0001 |
ECCV (13) | 2 |
| 2020 | Large-Scale Few-Shot Learning via Multi-modal Knowledge Discovery
Shuo Wang 0008, Jun Yue 0004, Jianzhuang Liu, Qi Tian 0001, Meng Wang 0001 |
ECCV (10) | 3 |
| 2020 | Attacking Image Captioning Towards Accuracy-Preserving Target Words RemovalabstractIn this paper, we investigate the fragility of deep image captioning models against adversarial attacks. Different from existing works that generate common words and concepts, we focus on the adversarial attacks towards controllable image captioning, i.e., removing target words from captions by imposing adversarial noises to images while maintaining the captioning accuracy for the remaining visual content. We name this new task as Masked Image Captioning (MIC), which is expected to be training and labeling free for end-to-end captioning models. Meanwhile, we propose a novel adversarial learning approach for this new task, termed Show, Mask, and Tell (SMT), which crafts adversarial examples to mask the target concepts via minimizing an objective loss while training the noise generator. Concretely, three novel designs are introduced in this loss, i.e., word removal regularization, captioning accuracy regularization, and noise filtering regularization. For quantitative validation, we propose a benchmark dataset for MIC based on the MS COCO dataset, together with a new evaluation metric called Attack Quality. Experimental results show that the proposed approach achieves successful attacks by removing 93.8% and 91.9% target words while maintaining 97.3% and 97.4% accuracies on two cutting-edge captioning models, respectively. Jiayi Ji, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji, Fuhai Chen, Jianzhuang Liu, Qi Tian 0001 |
ACM Multimedia | 6 |
| 2020 | Self-Adaptively Learning to Demoiré from Focused and Defocused Image PairsabstractMoiré artifacts are common in digital photography, resulting from the interference between high-frequency scene content and the color filter array of the camera. Existing deep learning-based demoiréing methods trained on large scale datasets are limited in handling various complex moiré patterns, and mainly focus on demoiréing of photos taken of digital displays. Moreover, obtaining moiré-free ground-truth in natural scenes is difficult but needed for training. In this paper, we propose a self-adaptive learning method for demoiréing a high-frequency image, with the help of an additional defocused moiré-free blur image. Given an image degraded with moiré artifacts and a moiré-free blur image, our network predicts a moiré-free clean image and a blur kernel with a self-adaptive strategy that does not require an explicit training stage, instead performing test-time adaptation. Our model has two sub-networks and works iteratively. During each iteration, one sub-network takes the moiré image as input, removing moiré patterns and restoring image details, and the other sub-network estimates the blur kernel from the blur image. The two sub-networks are jointly optimized. Extensive experiments demonstrate that our method outperforms state-of-the-art methods and can produce high-quality demoiréd results. It can generalize well to the task of removing moiré artifacts caused by display screens. In addition, we build a new moiré dataset, including images with screen and texture moiré artifacts. As far as we know, this is the first dataset with real texture moiré patterns. Lin Liu 0016, Shanxin Yuan, Jianzhuang Liu, Liping Bao, Gregory Slabaugh, Qi Tian 0001 |
NeurIPS | 3 |
| 2020 | DID: Disentangling-Imprinting-Distilling for Continuous Low-Shot DetectionabstractPractical applications often face a challenging continuous low-shot detection scenario, where a target detection task only has a few annotated training images, and a number of such new tasks come in sequence. To address this challenge, we propose a generic detection scheme via Disentangling-Imprinting-Distilling (DID). DID can leverage delicate transfer insights into the main development flow of deep learning, i.e., architecture design (Disentangling), model initialization (Imprinting), and training methodology (Distilling). This allows DID to be a simple but effective solution for continuous low-shot detection. In addition, DID can integrate the supervision from different detection tasks into a progressive learning procedure. As a result, one can efficiently adapt the previous detector for a new low-shot task, while maintaining the learned detection knowledge in the history. Finally, we evaluate our DID on a number of challenging settings in continuous/incremental low-shot detection. All the results demonstrate that our DID outperforms the recent state-of-the-art approaches. The code and models are available at https://github.com/chenxy99/DID. Yali Wang 0001, Jianzhuang Liu, Yu Qiao 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back PropagationabstractThe advancement of deep convolutional neural networks (DCNNs) has driven significant improvement in the accuracy of recognition systems for many computer vision tasks. However, their practical applications are often restricted in resource-constrained environments. In this paper, we introduce projection convolutional neural networks (PCNNs) with a discrete back propagation via projection (DBPP) to improve the performance of binarized neural networks (BNNs). The contributions of our paper include: 1) for the first time, the projection function is exploited to efficiently solve the discrete back propagation problem, which leads to a new highly compressed CNNs (termed PCNNs); 2) by exploiting multiple projections, we learn a set of diverse quantized kernels that compress the full-precision kernels in a more efficient way than those proposed previously; 3) PCNNs achieve the best classification performance compared to other state-ofthe-art BNNs on the ImageNet and CIFAR datasets. Jiaxin Gu, Ce Li 0002, Baochang Zhang 0001, Jungong Han, Xianbin Cao 0001, Jianzhuang Liu, David S. Doermann |
AAAI | 6 |
| 2019 | Calibrated Stochastic Gradient Descent for Convolutional Neural NetworksabstractIn stochastic gradient descent (SGD) and its variants, the optimized gradient estimators may be as expensive to compute as the true gradient in many scenarios. This paper introduces a calibrated stochastic gradient descent (CSGD) algorithm for deep neural network optimization. A theorem is developed to prove that an unbiased estimator for the network variables can be obtained in a probabilistic way based on the Lipschitz hypothesis. Our work is significantly distinct from existing gradient optimization methods, by providing a theoretical framework for unbiased variable estimation in the deep learning paradigm to optimize the model parameter calculation. In particular, we develop a generic gradient calibration layer which can be easily used to build convolutional neural networks (CNNs). Experimental results demonstrate that CNNs with our CSGD optimization scheme can improve the stateof-the-art performance for natural image classification, digit recognition, ImageNet object classification, and object detection tasks. This work opens new research directions for developing more efficient SGD updates and analyzing the backpropagation algorithm. Lian Zhuo, Baochang Zhang 0001, Chen Chen 0001, Qixiang Ye, Jianzhuang Liu, David S. Doermann |
AAAI | 5 |
| 2019 | Exploiting Kernel Sparsity and Entropy for Interpretable CNN CompressionabstractCompressing convolutional neural networks (CNNs) has received ever-increasing research focus. However, most existing CNN compression methods do not interpret their inherent structures to distinguish the implicit redundancy. In this paper, we investigate the problem of CNN compression from a novel interpretable perspective. The relationship between the input feature maps and 2D kernels is revealed in a theoretical framework, based on which a kernel sparsity and entropy (KSE) indicator is proposed to quantitate the feature map importance in a feature-agnostic manner to guide model compression. Kernel clustering is further conducted based on the KSE indicator to accomplish high-precision CNN compression. KSE is capable of simultaneously compressing each layer in an efficient way, which is significantly faster compared to previous data-driven feature map pruning methods. We comprehensively evaluate the compression and speedup of the proposed method on CIFAR-10, SVHN and ImageNet 2012. Our method demonstrates superior performance gains over previous ones. In particular, it achieves 4.7× FLOPs reduction and 2.9× compression on ResNet-50 with only a top-5 accuracy drop of 0.35% on ImageNet 2012, which significantly outperforms state-of-the-art methods. Shaohui Lin, Baochang Zhang 0001, Jianzhuang Liu, David S. Doermann, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
CVPR | 4 |
| 2019 | Circulant Binary Convolutional Networks: Enhancing the Performance of 1-Bit DCNNs With Circulant Back PropagationabstractThe rapidly decreasing computation and memory cost has recently driven the success of many applications in the field of deep learning. Practical applications of deep learning in resource-limited hardware, such as embedded devices and smart phones, however, remain challenging. For binary convolutional networks, the reason lies in the degraded representation caused by binarizing full-precision filters. To address this problem, we propose new circulant filters (CiFs) and a circulant binary convolution (CBConv) to enhance the capacity of binarized convolutional features via our circulant back propagation (CBP). The CiFs can be easily incorporated into existing deep convolutional neural networks (DCNNs), which leads to new Circulant Binary Convolutional Networks (CBCNs). Extensive experiments confirm that the performance gap between the 1-bit and full-precision DCNNs is minimized by increasing the filter diversity, which further increases the representational ability in our networks. Our experiments on ImageNet show that CBCNs achieve 61.4% top-1 accuracy with ResNet18. Compared to the state-of-the-art such as XNOR, CBCNs can achieve up to 10% higher top-1 accuracy with more powerful representational ability. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jiaxin Gu, Jianzhuang Liu, Rongrong Ji, David S. Doermann |
CVPR | 6 |
| 2019 | Bayesian Optimized 1-Bit CNNsabstractDeep convolutional neural networks (DCNNs) have dominated the recent developments in computer vision through making various record-breaking models. However, it is still a great challenge to achieve powerful DCNNs in resource-limited environments, such as on embedded devices and smart phones. Researchers have realized that 1-bit CNNs can be one feasible solution to resolve the issue; however, they are baffled by the inferior performance compared to the full-precision DCNNs. In this paper, we propose a novel approach, called Bayesian optimized 1-bit CNNs (denoted as BONNs), taking the advantage of Bayesian learning, a well-established strategy for hard problems, to significantly improve the performance of extreme 1-bit CNNs. We incorporate the prior distributions of full-precision kernels and features into the Bayesian framework to construct 1-bit CNNs in an end-to-end manner, which have not been considered in any previous related methods. The Bayesian losses are achieved with a theoretical support to optimize the network simultaneously in both continuous and discrete spaces, aggregating different losses jointly to improve the model capacity. Extensive experiments on the ImageNet and CIFAR datasets show that BONNs achieve the best classification performance compared to state-of-the-art 1-bit CNNs. Jiaxin Gu, Junhe Zhao, Baochang Zhang 0001, Jianzhuang Liu, Guodong Guo, Rongrong Ji |
ICCV | 5 |
| 2019 | Multinomial Distribution Learning for Effective Neural Architecture SearchabstractArchitectures obtained by Neural Architecture Search (NAS) have achieved highly competitive performance in various computer vision tasks. However, the prohibitive computation demand of forward-backward propagation in deep neural networks and searching algorithms makes it difficult to apply NAS in practice. In this paper, we propose a Multinomial Distribution Learning for extremely effective NAS, which considers the search space as a joint multinomial distribution, i.e., the operation between two nodes is sampled from this distribution, and the optimal network structure is obtained by the operations with the most likely probability in this distribution. Therefore, NAS can be transformed to a multinomial distribution learning problem, i.e., the distribution is optimized to have a high expectation of the performance. Besides, a hypothesis that the performance ranking is consistent in every training epoch is proposed and demonstrated to further accelerate the learning process. Experiments on CIFAR-10 and ImageNet demonstrate the effectiveness of our method. On CIFAR-10, the structure searched by our method achieves 2.55% test error, while being 6.0× (only 4 GPU hours on GTX1080Ti) faster compared with state-of-the-art NAS algorithms. On ImageNet, our model achieves 74% top1 accuracy under MobileNet settings (MobileNet V1/V2), while being 1.2× faster with measured GPU latency. Test code with pre-trained models are available at https: //github.com/tanglang96/MDENAS. Xiawu Zheng, Rongrong Ji, Lang Tang, Baochang Zhang 0001, Jianzhuang Liu, Qi Tian 0001 |
ICCV | 5 |
| 2019 | Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNsabstractBinarized convolutional neural networks (BCNNs) are widely used to improve memory and computation efficiency of deep convolutional neural networks (DCNNs) for mobile and AI chips based applications. However, current BCNNs are not able to fully explore their corresponding full-precision models, causing a significant performance gap between them. In this paper, we propose rectified binary convolutional networks (RBCNs), towards optimized BCNNs, by combining full-precision kernels and feature maps to rectify the binarization process in a unified framework. In particular, we use a GAN to train the 1-bit binary network with the guidance of its corresponding full-precision model, which significantly improves the performance of BCNNs. The rectified convolutional layers are generic and flexible, and can be easily incorporated into existing DCNNs such as WideResNets and ResNets. Extensive experiments demonstrate the superior performance of the proposed RBCNs over state-of-the-art BCNNs. In particular, our method shows strong generalization on the object tracking task. Chunlei Liu 0001, Wenrui Ding, Xin Xia 0005, Baochang Zhang 0001, Jianzhuang Liu, Bohan Zhuang, Guodong Guo |
IJCAI | 6 |
| 2019 | Taylor Convolutional Networks for Image ClassificationabstractThis paper provides a new perspective to understand CNNs based on the Taylor expansion, leading to new Taylor Convolutional Networks (TaylorNets) for image classification. We introduce a principled combination of the high frequency information (i.e., detailed information) and low frequency information in the end-to-end TaylorNets, based on a nonlinear combination of the convolutional feature maps. The steerable module developed in TaylorNets is generic, which can be easily integrated into well-known deep architectures and learned within the same pipeline of the back propagation algorithm, yielding a higher representation capacity for CNNs. Extensive experimental results demonstrate the super capability of our TaylorNets which improve widely used CNNs architectures, such as conventional CNNs and ResNet, in terms of object classification accuracy on well-known benchmarks. The code will be publicly available. Ce Li 0002, Yipeng Mou, Baochang Zhang 0001, Jungong Han, Jianzhuang Liu |
WACV | 6 |
| 2018 | Modulated Convolutional NetworksabstractDespite great effectiveness of very deep and wide Convolutional Neural Networks (CNNs) in various computer vision tasks, the significant cost in terms of storage requirement of such networks impedes the deployment on computationally limited devices. In this paper, we propose new modulated convolutional networks (MCNs) to improve the portability of CNNs via binarized filters. In MCNs, we propose a new loss function which considers the filter loss, center loss and softmax loss in an end-to-end framework. We first introduce modulation filters (M-Filters) to recover the unbinarized filters, which leads to a new architecture to calculate the network model. The convolution operation is further approximated by considering intra-class compactness in the loss function. As a result, our MCNs can reduce the size of required storage space of convolutional filters by a factor of 32, in contrast to the full-precision model, while achieving much better performances than state-of-the-art binarized models. Most importantly, MCNs achieve a comparable performance to the full-precision Resnets and WideResnets. The code will be available publicly soon. Baochang Zhang 0001, Ce Li 0002, Rongrong Ji, Jungong Han, Xianbin Cao 0001, Jianzhuang Liu |
CVPR | 7 |
| 2018 | Memory Attention Networks for Skeleton-based Action RecognitionabstractSkeleton-based action recognition task is entangled with complex spatio-temporal variations of skeleton joints, and remains challenging for Recurrent Neural Networks (RNNs). In this work, we propose a temporal-then-spatial recalibration scheme to alleviate such complex variations, resulting in an end-to-end Memory Attention Networks (MANs) which consist of a Temporal Attention Recalibration Module (TARM) and a Spatio-Temporal Convolution Module (STCM). Specifically, the TARM is deployed in a residual learning module that employs a novel attention learning network to recalibrate the temporal attention of frames in a skeleton sequence. The STCM treats the attention calibrated skeleton joint sequences as images and leverages the Convolution Neural Networks (CNNs) to further model the spatial and temporal information of skeleton data. These two modules (TARM and STCM) seamlessly form a single network architecture that can be trained in an end-to-end fashion. MANs significantly boost the performance of skeleton-based action recognition and achieve the best results on four challenging benchmark datasets: NTU RGB+D, HDM05, SYSU-3D and UT-Kinect. Chunyu Xie, Ce Li 0002, Baochang Zhang 0001, Chen Chen 0001, Jungong Han, Jianzhuang Liu |
IJCAI | 6 |
| 2018 | Gabor Convolutional NetworksabstractSteerable properties dominate the design of traditional filters, e.g., Gabor filters, and endow features the capability of dealing with spatial transformations. However, such excellent properties have not been well explored in the popular deep convolutional neural networks (DCNNs). In this paper, we propose a new deep model, termed Gabor Convolutional Networks (GCNs or Gabor CNNs), which incorporates Gabor filters into DCNNs to enhance the resistance of deep learned features to the orientation and scale changes. By only manipulating the basic element of DCNNs based on Gabor filters, i.e., the convolution operator, GCNs can be easily implemented and are compatible with any popular deep learning architecture. Experimental results demonstrate the super capability of our algorithm in recognizing objects, where the scale and rotation changes occur frequently. The proposed GCNs have much fewer learnable network parameters, and thus is easier to train with an endtoend pipeline. The source code will be here1. Shangzhen Luan, Baochang Zhang 0001, Siyue Zhou, Chen Chen 0001, Jungong Han, Wankou Yang, Jianzhuang Liu |
WACV | 7 |
| 2018 | Gabor Convolutional NetworksabstractIn steerable filters, a filter of arbitrary orientation can be generated by a linear combination of a set of "basis filters." Steerable properties dominate the design of the traditional filters, e.g., Gabor filters and endow features the capability of handling spatial transformations. However, such properties have not yet been well explored in the deep convolutional neural networks (DCNNs). In this paper, we develop a new deep model, namely, Gabor convolutional networks (GCNs or Gabor CNNs), with Gabor filters incorporated into DCNNs such that the robustness of learned features against the orientation and scale changes can be reinforced. By manipulating the basic element of DCNNs, i.e., the convolution operator, based on Gabor filters, GCNs can be easily implemented and are readily compatible with any popular deep learning architecture. We carry out extensive experiments to demonstrate the promising performance of our GCNs framework, and the results show its superiority in recognizing objects, especially when the scale and rotation changes take place frequently. Moreover, the proposed GCNs have much fewer network parameters to be learned and can effectively reduce the training complexity of the network, leading to a more compact deep learning model while still maintaining a high feature representation capacity. The source code can be found at https://github.com/bczhangbczhang. Shangzhen Luan, Chen Chen 0001, Baochang Zhang 0001, Jungong Han, Jianzhuang Liu |
IEEE Trans. Image Process. | 5 |
| 2017 | Adaptive Local Movement Modeling for Robust Object TrackingabstractIn this paper, we present a new strategy for modeling the motion of local patches for single-object tracking that can be seamlessly applied to most part-based trackers in the literature. The proposed adaptive local movement modeling method is able to model the local movement distribution of the image patches defining the object to track and the reliability of each image patch. Given the output of a base tracking algorithm, a Gaussian mixture model (GMM) is first used to model the distribution of the movement of local patches relative to the center of gravity of the tracked object. Then, the GMM is combined with the chosen base tracker in a boosting framework, which gives an efficient integrated scheme for the tracking task. This provides a robust procedure to detect outliers in the local motion of the patches. The algorithm is highly configurable with the possibility to change the number of local patches used for tracking and to adapt to the variations of the tracked object. The extensive tracking results on standard data sets show that equipping state-of-the-art trackers with our technique remarkably improves their performance. Baochang Zhang 0001, Alessandro Perina, Alessio Del Bue, Vittorio Murino, Jianzhuang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Bounding Multiple Gaussians Uncertainty with Application to Object Tracking
Baochang Zhang 0001, Alessandro Perina, Vittorio Murino, Jianzhuang Liu, Rongrong Ji |
Int. J. Comput. Vis. | 5 |
| 2016 | An example-based approach to 3D man-made object reconstruction from line drawings
Changqing Zou, Tianfan Xue, Xiaojiang Peng, Honghua Li, Baochang Zhang 0001, Jianzhuang Liu |
Pattern Recognit. | 7 |
| 2015 | A maximum entropy feature descriptor for age invariant face recognitionabstractIn this paper, we propose a new approach to overcome the representation and matching problems in age invariant face recognition. First, a new maximum entropy feature descriptor (MEFD) is developed that encodes the microstructure of facial images into a set of discrete codes in terms of maximum entropy. By densely sampling the encoded face image, sufficient discriminatory and expressive information can be extracted for further analysis. A new matching method is also developed, called identity factor analysis (IFA), to estimate the probability that two faces have the same underlying identity. The effectiveness of the framework is confirmed by extensive experimentation on two face aging datasets, MORPH (the largest public-domain face aging dataset) and FGNET. We also conduct experiments on the famous LFW dataset to demonstrate the excellent generalizability of our new approach. Dihong Gong, Zhifeng Li 0001, Dacheng Tao, Jianzhuang Liu, Xuelong Li 0001 |
CVPR | 4 |
| 2015 | Sketch-based 3-D modeling for piecewise planar objects in single images
Changqing Zou, Xiaojiang Peng, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
Comput. Graph. | 6 |
| 2015 | A comparison of 3D shape retrieval methods based on a large-scale benchmark supporting multimodal queries
Bo Li 0013, Yijuan Lu, Chunyuan Li, Afzal Godil, Tobias Schreck, Masaki Aono, Martin Burtscher, Nihad Karim Chowdhury, Hongbo Fu 0001, Takahiko Furuya, Hai-Sheng Li 0002, Jianzhuang Liu, Henry Johan, Ryuichi Kosaka, Hitoshi Koyanagi, Ryutarou Ohbuchi, Atsushi Tatsuma, Yajuan Wan, Changqing Zou |
Comput. Vis. Image Underst. | 14 |
| 2015 | Hierarchical facial landmark localization via cascaded random binary patterns
Wei Zhang 0081, Huijun Ding, Jianzhuang Liu, Xiaoou Tang |
Pattern Recognit. | 4 |
| 2015 | Depth From Water ReflectionabstractThe scene in a water reflection image often exhibits bilateral symmetry. In this paper, we design a framework to reconstruct the depth from a single water reflection image. This problem can be regarded as a special case of two-view stereo vision. It is challenging to obtain correspondences from the real scene and the mirror scene due to their large appearance difference. We first propose an appearance adaptation method to transform the appearance of the mirror scene so that it is much closer to the real scene. We then present a stereo matching algorithm to obtain the disparity map of the real scene. Compared with other depth-from-symmetry work that deals with man-made objects, our algorithm can recover the depth maps of a variety of scenes, where both natural and man-made objects may exist. Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Image Process. | 2 |
| 2015 | Progressive 3D Reconstruction of Planar-Faced Manifold Objects with DRF-Based Line Drawing DecompositionabstractThis paper presents an approach for reconstructing polyhedral objects from single-view line drawings. Our approach separates a complex line drawing representing a manifold object into a series of simpler line drawings, based on the degree of reconstruction freedom (DRF). We then progressively reconstruct a complete 3D model from these simpler line drawings. Our experiments show that our decomposition algorithm is able to handle complex drawings which are challenging for the state of the art. The advantages of the presented progressive 3D reconstruction method over the existing reconstruction methods in terms of both robustness and efficiency are also demonstrated. Changqing Zou, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Transitive Distance Clustering with K-Means DualityabstractWe propose a very intuitive and simple approximation for the conventional spectral clustering methods. It effectively alleviates the computational burden of spectral clustering - reducing the time complexity from O(n3) to O(n2) - while capable of gaining better performance in our experiments. Specifically, by involving a more realistic and effective distance and the "k-means duality" property, our algorithm can handle datasets with complex cluster shapes, multi-scale clusters and noise. We also show its superiority in a series of its real applications on tasks including digit clustering as well as image segmentation. Zhiding Yu, Chunjing Xu, Deyu Meng, Zhuo Hui, Fanyi Xiao, Wenbo Liu 0002, Jianzhuang Liu |
CVPR | 7 |
| 2014 | Separation of Line Drawings Based on Split Faces for 3D Object ReconstructionabstractReconstructing 3D objects from single line drawings is often desirable in computer vision and graphics applications. If the line drawing of a complex 3D object is decomposed into primitives of simple shape, the object can be easily reconstructed. We propose an effective method to conduct the line drawing separation and turn a complex line drawing into parametric 3D models. This is achieved by recursively separating the line drawing using two types of split faces. Our experiments show that the proposed separation method can generate more basic and simple line drawings, and its combination with the example-based reconstruction can robustly recover wider range of complex parametric 3D objects than previous methods. Changqing Zou, Jianzhuang Liu |
CVPR | 3 |
| 2014 | Object Detection and Viewpoint Estimation with Auto-masking Neural Network
Jianzhuang Liu, Xiaoou Tang |
ECCV (3) | 2 |
| 2014 | Sketch-Based 3D Model Retrieval via Multi-feature FusionabstractSketch-based 3D model retrieval provides a convenient way for users to search for 3D models by sketches. Traditionally, this task is converted to a sketch-based 2D shape retrieval problem by projecting 3D models to 2D images. Local invariant features have been widely used to tackle this problem. However, it suffers from the lack of global context and easily fails when images of different 3D models share multiple similar regions. In this paper, we propose a joint description by fusing local statistical structures and global spatial features. Our description is invariant to scale, translate and rotation. An improved bag-of-features retrieval framework is applied to explore semantic visual word representations. Besides, a novel relevance feedback scheme which combines weight balancing and query modification is designed to further improve the retrieval performance. We conduct various experiments on the common sketch-based watertight model benchmark. The comparative results show that our approach significantly outperforms three state-of-the-art methods, demonstrating its effectiveness and robustness for sketch-based 3D model retrieval. Yafei Wen, Changqing Zou, Jianzhuang Liu, Shuze Du, Shifeng Chen |
ICPR | 3 |
| 2014 | Similarity Michaelis-Menten law pre-processing descriptor for face recognitionabstractThis paper presents a non-linear pre-processing method based on Similarity Michaelis-Menten law (SMML) for face recognition. Similarity Michaelis-Menten law can be used to explain visual sensitivity in the vertebrate retina. We preprocess input images using SMML, and then employ Local Binary Pattern (LBP) for face feature extraction. Advantages of SMML include improvement of light adaption, noise effect, detection right rate, robustness and efficiency, which inspire us exploit it for face pre-processing descriptor for the first time in the field of face recognition. And the parameters of SMML are spatiotemporally and locally estimated by the input image itself employing Sobel, which shows its advantages for face recognition. Extensive experiments clearly demonstrate the superiority of our method over the ones which only use LBP on FERET database in many aspects including the robustness against different facial expressions, lighting and aging of the subjects. Suli Ji, Baochang Zhang 0001, Dandan Du, Jianzhuang Liu |
IJCNN | 5 |
| 2014 | The BeiHang Keystroke Dynamics Systems, Databases and baselines
Baochang Zhang 0001, Haoran Zeng, LinLin Shen, Jianzhuang Liu, Jason Zhao |
Neurocomputing | 5 |
| 2014 | Viewpoint-Aware Representation for Sketch-Based 3D Model RetrievalabstractWe study the problem of sketch-based 3D model retrieval, and propose a solution powered by a new query-to-model distance metric and a powerful feature descriptor based on the bag-of-features framework. The main idea of the proposed query-to-model distance metric is to represent a query sketch using a compact set of sample views (called basic views) of each model, and to rank the models in ascending order of the representation errors. To better differentiate between relevant and irrelevant models, the representation is constrained to be essentially a combination of basic views with similar viewpoints. In another aspect, we propose a mid-level descriptor (called BOF-JESC) which robustly characterizes the edge information within junction-centered patches, to extract the salient shape features from sketches or model views. The combination of the query-to-model distance metric and the BOF-JESC descriptor achieves effective results on two latest benchmark datasets. Changqing Zou, Changhu Wang, Yafei Wen, Lei Zhang 0001, Jianzhuang Liu |
IEEE Signal Process. Lett. | 5 |
| 2014 | Multiview Facial Landmark Localization in RGB-D Images via Hierarchical Regression With Binary PatternsabstractIn this paper, we propose a real-time system of multiview facial landmark localization in RGB-D images. The facial landmark localization problem is formulated into a regression framework, which estimates both the head pose and the landmark positions. In this framework, we propose a coarse-to-fine approach to handle the high-dimensional regression output. At first, 3-D face position and rotation are estimated from the depth observation via a random regression forest. Afterward, the 3-D pose is refined by fusing the estimation from the RGB observation. Finally, the landmarks are located from the RGB observation with gradient boosted decision trees in a pose conditional model. The benefits of the proposed localization framework are twofold: the pose estimation and landmark localization are solved with hierarchical regression, which is different from previous approaches where the pose and landmark locations are iteratively optimized, which relies heavily on the initial pose estimation; due to the different characters of the RGB and depth cues, they are used for landmark localization at different stages and incorporated in a robust manner. In the experiments, we show that the proposed approach outperforms state-of-the-art algorithms on facial landmark localization with RGB-D input. Wei Zhang 0081, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Hidden Factor Analysis for Age Invariant Face RecognitionabstractAge invariant face recognition has received increasing attention due to its great potential in real world applications. In spite of the great progress in face recognition techniques, reliably recognizing faces across ages remains a difficult task. The facial appearance of a person changes substantially over time, resulting in significant intra-class variations. Hence, the key to tackle this problem is to separate the variation caused by aging from the person-specific features that are stable. Specifically, we propose a new method, called Hidden Factor Analysis (HFA). This method captures the intuition above through a probabilistic model with two latent factors: an identity factor that is age-invariant and an age factor affected by the aging process. Then, the observed appearance can be modeled as a combination of the components generated based on these factors. We also develop a learning algorithm that jointly estimates the latent factors and the model parameters using an EM procedure. Extensive experiments on two well-known public domain face aging datasets: MORPH (the largest public face aging database) and FGNET, clearly show that the proposed method achieves notable improvement over state-of-the-art algorithms. Dihong Gong, Zhifeng Li 0001, Dahua Lin, Jianzhuang Liu, Xiaoou Tang |
ICCV | 4 |
| 2013 | Complex 3D General Object Reconstruction from Line DrawingsabstractAn important topic in computer vision is 3D object reconstruction from line drawings. Previous algorithms either deal with simple general objects or are limited to only manifolds (a subset of solids). In this paper, we propose a novel approach to 3D reconstruction of complex general objects, including manifolds, non-manifold solids, and nonsolids. Through developing some 3D object properties, we use the degree of freedom of objects to decompose a complex line drawing into multiple simpler line drawings that represent meaningful building blocks of a complex object. After 3D objects are reconstructed from the decomposed line drawings, they are merged to form a complex object from their touching faces, edges, and vertices. Our experiments show a number of reconstruction examples from both complex line drawings and images with line drawings superimposed. Comparisons are also given to indicate that our algorithm can deal with much more complex line drawings of general objects than previous algorithms. Jianzhuang Liu, Xiaoou Tang |
ICCV | 2 |
| 2013 | Multi-feature canonical correlation analysis for face photo-sketch image retrievalabstractAutomatic face photo-sketch image retrieval has attracted great attention in recent years due to its important applications in real life. The major difficulty in automatic face photo-sketch image retrieval lies in the fact that there exists great discrepancy between the different image modalities (photo and sketch). In order to reduce such discrepancy and improve the performance of automatic face photo-sketch image retrieval, we propose a new framework called multi-feature canonical correlation analysis (MCCA) to effectively address this problem. The MCCA is an extension and improvement of the canonical correlation analysis (CCA) algorithmusing multiple features combined with two different random sampling methods in feature space and sample space. In this framework, we first represent each photo or sketch using a patch-based local feature representation scheme, in which histograms of oriented gradients (HOG) and multi-scale local binary pattern (MLBP) serve as the local descriptors. Canonical correlation analysis (CCA) is then performed on a collection of random subspaces to construct an ensemble of classifiers for photo-sketch image retrieval. Extensive experiments on two public-domain face photo-sketch datasets (CUFS and CUFSF) clearly show that the proposed approach obtains a substantial improvement over the state-of-the-art. Dihong Gong, Zhifeng Li 0001, Jianzhuang Liu, Yu Qiao 0001 |
ACM Multimedia | 3 |
| 2013 | Facial landmark localization based on hierarchical pose regression with cascaded random fernsabstractThe main challenge of facial landmark localization in real-world application is that the large changes of head pose and facial expressions cause substantial image appearance variations. To avoid high dimensional regression in the 3D and 2D facial pose spaces simultaneously, we propose a hierarchical pose regression approach, estimating the head rotation, facial components and landmarks hierarchically. The regression process works in a unified cascaded fern framework. We present generalized gradient boosted ferns (GBFs) for the regression framework, which give better performance than traditional ferns. The framework also achieves real time performance. We verify our method on the latest benchmark datasets. The results show that it outperforms state-of-the-art methods in both accuracy and speed. Wei Zhang 0081, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2013 | Weighted Margin Sparse Embedded classifier for brake cylinder detection
Yao Cao, Baochang Zhang 0001, Jianzhuang Liu, Jiangsha Ma |
Neurocomputing | 3 |
| 2013 | Edge preserving image denoising with a closed form solution
Shifeng Chen, Wei Zhang 0081, Jianzhuang Liu |
Pattern Recognit. | 4 |
| 2013 | Human Detection in Images via Piecewise Linear Support Vector MachinesabstractHuman detection in images is challenged by the view and posture variation problem. In this paper, we propose a piecewise linear support vector machine (PL-SVM) method to tackle this problem. The motivation is to exploit the piecewise discriminative function to construct a nonlinear classification boundary that can discriminate multiview and multiposture human bodies from the backgrounds in a high-dimensional feature space. A PL-SVM training is designed as an iterative procedure of feature space division and linear SVM training, aiming at the margin maximization of local linear SVMs. Each piecewise SVM model is responsible for a subspace, corresponding to a human cluster of a special view or posture. In the PL-SVM, a cascaded detector is proposed with block orientation features and a histogram of oriented gradient features. Extensive experiments show that compared with several recent SVM methods, our method reaches the state of the art in both detection accuracy and computational efficiency, and it performs best when dealing with low-resolution human regions in clutter backgrounds. Qixiang Ye, Zhenjun Han, Jianbin Jiao, Jianzhuang Liu |
IEEE Trans. Image Process. | 4 |
| 2013 | Learning Semantic Signatures for 3D Object RetrievalabstractIn this paper, we propose two kinds of semantic signatures for 3D object retrieval (3DOR). Humans are capable of describing an object using attribute terms like “symmetric” and “flyable”, or using its similarities to some known object classes. We convert such qualitative descriptions into attribute signature (AS) and reference set signature (RSS), respectively, and use them for 3DOR. We also show that AS and RSS can be understood as two different quantization methods of the same semantic space of human descriptions of objects. The advantages of the semantic signatures are threefold. First, they are much more compact than low-level shape features yet working with comparable retrieval accuracy. Therefore, the proposed semantic signatures require less storage space and computation cost in retrieval. Second, the high-level signatures are a good complement to low-level shape features. As a result, by incorporating the signatures we can improve the performance of state-of-the-art 3DOR methods by a large margin. To the best of our knowledge, we obtain the best results on two popular benchmarks. Third, the AS enables us to build a user-friendly interface, with which the user can trigger a search by simply clicking attribute bars instead of finding a 3D object as the query. This interface is of great significance in 3DOR considering the fact that while searching, the user usually does not have a 3D query at hand that is similar to his/her targeted objects in the database. Boqing Gong, Jianzhuang Liu, Xiaogang Wang 0001, Xiaoou Tang |
IEEE Trans. Multim. | 2 |
| 2013 | Style Transfer Via Image Component AnalysisabstractExample-based stylization provides an easy way of making artistic effects for images and videos. However, most existing methods do not consider the content and style separately. In this paper, we propose a style transfer algorithm via a novel component analysis approach, based on various image processing techniques. First, inspired by the steps of drawing a picture, an image is decomposed into three components: draft, paint and edge, which describe the content, main style, and strengthened strokes along the boundaries. Then the style is transferred from the template image to the source image in the paint and edge components. Style transfer is formulated as a global optimization problem by using Markov random fields, and a coarse-to-fine belief propagation algorithm is used to solve the optimization problem. To combine the draft component and the obtained style information, the final artistic result can be achieved via a reconstruction step. Compared to other algorithms, our method not only synthesizes the style, but also preserves the image content well. We also extend our algorithm from single image stylization to video personalization, by maintaining the temporal coherence and identifying faces in video sequences. The results indicate that our approach performs excellently in stylization and personalization for images and videos. Wayne Zhang 0001, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Multim. | 4 |
| 2012 | Precise 3D Reconstruction from a Single Image
Changqing Zou, Jianzhuang Liu |
ACCV (4) | 3 |
| 2012 | Example-based 3D object reconstruction from line drawingsabstractRecovering 3D geometry from a single 2D line drawing is an important and challenging problem in computer vision. It has wide applications in interactive 3D modeling from images, computer-aided design, and 3D object retrieval. Previous methods of 3D reconstruction from line drawings are mainly based on a set of heuristic rules. They are not robust to sketch errors and often fail for objects that do not satisfy the rules. In this paper, we propose a novel approach, called example-based 3D object reconstruction from line drawings, which is based on the observation that a natural or man-made complex 3D object normally consists of a set of basic 3D objects. Given a line drawing, a graphical model is built where each node denotes a basic object whose candidates are from a 3D model (example) database. The 3D reconstruction is solved using a maximum-a-posteriori (MAP) estimation such that the reconstructed result best fits the line drawing. Our experiments show that this approach achieves much better reconstruction accuracy and are more robust to imperfect line drawings than previous methods. Tianfan Xue, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2012 | Joint Example-Based Depth Map Super-ResolutionabstractThe fast development of time-of-flight (ToF) cameras in recent years enables capture of high frame-rate 3D depth maps of moving objects. However, the resolution of depth map captured by ToF is rather limited, and thus it cannot be directly used to build a high quality 3D model. In order to handle this problem, we propose a novel joint example-based depth map super-resolution method, which converts a low resolution depth map to a high resolution depth map, using a registered high resolution color image as a reference. Different from previous depth map SR methods without training stage, we learn a mapping function from a set of training samples and enhance the resolution of the depth map via sparse coding algorithm. We further use a reconstruction constraint to make object edges sharper. Experimental results show that our method outperforms state-of-the-art methods for depth map super-resolution. Tianfan Xue, Lifeng Sun, Jianzhuang Liu |
ICME | 4 |
| 2012 | Locating high-density clusters with noisy queries
Shifeng Chen, Changqing Zou, Jianzhuang Liu |
ICPR | 4 |
| 2012 | Online non-feedback image re-ranking via dominant data selectionabstractImage re-ranking aims at improving the precision of keyword-based image retrieval, mainly by introducing visual features to re-rank. Many existing approaches require offline training for every keyword, which are unsuitable for online image search. Other real-time approaches demand user interaction, which are inappropriate for large-scale image collection. To improve the accuracy of web image retrieval in a practicable manner, we propose a novel re-ranking algorithm to explore the cluster information of the image set. First, we build spectral graph on images that retrieved bysearch engine, and remove isolated nodes as noisy images. Then, we select positive samples from the most dominant cluster in initial top-ranked images, and the samples are used for semi-supervised learning and ranking. Our algorithm is online and non-feedback. Experiments on two public databases demonstrate that our algorithm outperforms the state-of-the-art approaches. Shifeng Chen, Jianzhuang Liu |
ACM Multimedia | 4 |
| 2012 | Optimal semi-supervised metric learning for image retrievalabstractIn a typical content-based image retrieval (CBIR) system, images are represented as vectors and similarities between images are measured by a specified distance metric. However, the traditional Euclidean distance cannot always deliver satisfactory performance, so an effective metric sensible to the input data is desired. Tremendous recent works on metric learning have exhibited promising performance, but most of them suffer from limited label information and expensive training costs. In this paper, we propose two novel metric learning approaches, Optimal Semi-Supervised Metric Learning and its kernelized version. In the proposed approaches, we incorporate information from both labeled and unlabeled data to design a convex and computationally tractable learning framework which results in a globally optimal solution to the target metric of much lower rank than the original data dimension. Experiments on several image benchmarks demonstrate that our approaches lead to consistently better distance metrics than the state-of-the-arts in terms of accuracy for image retrieval. Wei Liu 0005, Jianzhuang Liu |
ACM Multimedia | 3 |
| 2012 | 3-D Modeling From a Single View of a Symmetric Objectabstract3-D technologies are considered as the next generation of multimedia applications. Currently, one of the challenges faced by 3-D applications is the shortage of 3-D resources. To solve this problem, many 3-D modeling methods are proposed to directly recover 3-D geometry from 2-D images. However, these methods on single view modeling either require intensive user interaction, or are restricted to a specific kind of object. In this paper, we propose a novel 3-D modeling approach to recover 3-D geometry from a single image of a symmetric object with minimal user interaction. Symmetry is one of the most common properties of natural or manmade objects. Given a single view of a symmetric object, the user marks some symmetric lines and depth discontinuity regions on the image. Our algorithm first finds a set of planes to approximately fit to the object, and then a rough 3-D point cloud is generated by an optimization procedure. The occluded part of the object is further recovered using symmetry information. Experimental results on various indoor and outdoor objects show that the proposed system can obtain 3-D models from single images with only a little user interaction. Tianfan Xue, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Image Process. | 2 |
| 2011 | Symmetric piecewise planar object reconstruction from a single imageabstractRecovering 3D geometry from a single view of an object is an important and challenging problem in computer vision. Previous methods mainly focus on one specific class of objects without large topological changes, such as cars, faces, or human bodies. In this paper, we propose a novel single view reconstruction algorithm for symmetric piece-wise planar objects that are not restricted to some object classes. Symmetry is ubiquitous in manmade and natural objects and provides rich information for 3D reconstruction. Given a single view of a symmetric piecewise planar object, we first find out all the symmetric line pairs. The geometric properties of symmetric objects are used to narrow down the searching space. Then, based on the symmetric lines, a depth map is recovered through a Markov random field. Experimental results show that our algorithm can efficiently recover the 3D shapes of different objects with significant topological variations. Tianfan Xue, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2011 | Study on the BeiHang Keystroke Dynamics DatabaseabstractThis paper introduces a new BeiHang (BH) Keystroke Dynamics Database for testing and evaluation of biometric approaches. Different from the existing keystroke dynamics researches which solely rely on laboratory experiments, the developed database is collected from a real commercialized system and thus is more comprehensive and more faithful to human behavior. Moreover, our database comes with ready-to-use benchmark results of three keystroke dynamics methods, Nearest Neighbor classifier, Gaussian Model and One-Class Support Vector Machine. Both the database and benchmark results are open to the public and provide a significant experimental platform for international researchers in the keystroke dynamics area. Baochang Zhang 0001, Yao Cao, Sanqiang Zhao, Yongsheng Gao 0001, Jianzhuang Liu |
IJCB | 6 |
| 2011 | Sparse regression analysis for object recognitionabstractThis paper proposes a new method named Sparse Regression Analysis (SRA) for object representation and recognition. In SRA, ℓ1-norm minimization is combined with regression analysis to represent the input signal. The discriminative ability of SRA derives from the fact that the subset which most compactly expresses the input signal is activated in the regression analysis. To achieve a further improvement, Kernelized SRA (KSRA) is developed to make a nonlinear extension of SRA. The experiments are conducted on both palmprint and face recognition, which show that the proposed methods achieve a much better performance than sparse representation classifier, principal component analysis, and linear discriminant analysis. Baochang Zhang 0001, Shengping Zhang, Jianzhuang Liu |
ICIP | 3 |
| 2011 | 3D object retrieval with semantic attributesabstractHumans are capable of describing objects using attributes, such as "the object looks circular and is man-made". Motivated by these high-level descriptions, we build a user-friendly 3D object retrieval system, where the user can browse the database and search for targeted objects using semantic attributes. The main advantage of our system is that it does not require the user to find or sketch a 3D object as the query for 3D object retrieval. Besides, to the best of our knowledge, our system has obtained the best retrieval performance on three popular benchmarks. Boqing Gong, Jianzhuang Liu, Xiaogang Wang 0001, Xiaoou Tang |
ACM Multimedia | 2 |
| 2011 | Automatic object segmentation from large scale 3D urban point clouds through manifold embedded mode seekingabstractThis paper presents a system that can automatically segment objects in large scale 3D point clouds obtained from urban ranging images. The system consists of three steps: The first one involves a ground detection process that can detect relatively complex terrain and separate it from other objects. The second step superpixelizes the remaining objects to speed up the segmentation process. In the final step, a manifold embedded mode seeking method is adopted to segment the point clouds. Even though the segmentation of urban objects is a challenging problem in terms of accuracy and problem scale, our system can efficiently generate very good segmentation results. The proposed manifold learning effectively improves the segmentation performance due to the fact that continuous artificial objects often have manifold-like structures. Zhiding Yu, Chunjing Xu, Jianzhuang Liu, Oscar C. Au, Xiaoou Tang |
ACM Multimedia | 3 |
| 2011 | Edge-preserving single image super-resolutionabstractThis paper proposes a novel approach to single image super-resolution. First, an image up-sampling scheme is proposed which takes the advantages of both bilateral filtering and mean shift image segmentation. Then we use a shock filter to enhance strong edges in the initial up-sampling result and obtain an intermediate high-resolution image. Finally, we enforce a reconstruction constraint on the high-resolution image so that fine details can be inferred by back projection. Since strong edges in the intermediate result are enhanced, ringing artifacts can be suppressed in the back projection step. We compare our algorithm with several state-of-the-art image super-resolution algorithms. Qualitative and quantitative experimental results demonstrate that our approach performs the best. Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2011 | Decomposition of Complex Line Drawings with Hidden Lines for 3D Planar-Faced Manifold Object ReconstructionabstractThree-dimensional object reconstruction from a single 2D line drawing is an important problem in computer vision. Many methods have been presented to solve this problem, but they usually fail when the geometric structure of a 3D object becomes complex. In this paper, a novel approach based on a divide-and-conquer strategy is proposed to handle the 3D reconstruction of a planar-faced complex manifold object from its 2D line drawing with hidden lines visible. The approach consists of four steps: 1) identifying the internal faces of the line drawing, 2) decomposing the line drawing into multiple simpler ones based on the internal faces, 3) reconstructing the 3D shapes from these simpler line drawings, and 4) merging the 3D shapes into one complete object represented by the original line drawing. A number of examples are provided to show that our approach can handle 3D reconstruction of more complex objects than previous methods. Jianzhuang Liu, Yu Chen 0009, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Visual object tracking via sample-based Adaptive Sparse Representation (AdaSR)
Zhenjun Han, Jianbin Jiao, Baochang Zhang 0001, Qixiang Ye, Jianzhuang Liu |
Pattern Recognit. | 5 |
| 2011 | A deformation model to reduce the effect of expressions in 3D face recognition
Yueming Wang 0001, Gang Pan 0001, Jianzhuang Liu |
Vis. Comput. | 3 |
| 2010 | Constrained Metric Learning Via Distance Gap MaximizationabstractVectored data frequently occur in a variety of fields, which are easy to handle since they can be mathematically abstracted as points residing in a Euclidean space. An appropriate distance metric in the data space is quite demanding for a great number of applications. In this paper, we pose robust and tractable metric learning under pairwise constraints that are expressed as similarity judgements between data pairs. The major features of our approach include: 1) it maximizes the gap between the average squared distance among dissimilar pairs and the average squared distance among similar pairs; 2) it is capable of propagating similar constraints to all data pairs; and 3) it is easy to implement in contrast to the existing approaches using expensive optimization such as semidefinite programming. Our constrained metric learning approach has widespread applicability without being limited to particular backgrounds. Quantitative experiments are performed for classification and retrieval tasks, uncovering the effectiveness of the proposed approach. Wei Liu 0005, Xinmei Tian 0001, Dacheng Tao, Jianzhuang Liu |
AAAI | 4 |
| 2010 | Isoperimetric cut on a directed graphabstractIn this paper, we propose a novel probabilistic view of the spectral clustering algorithm. In our framework, the spectral clustering algorithm can be viewed as assigning class labels to samples to minimize the Bayes classification error rate by using a kernel density estimator (KDE). From this perspective, we propose to construct directed graphs using variable bandwidth KDEs. Such a variable bandwidth KDE based directed graph has the advantage that it encodes the local density information of the data in the graph edge weights. In order to cluster vertices of the directed graph, we develop a directed graph partitioning algorithm which optimizes a random walk isoperimetric ratio. The partitioning result can be obtained efficiently by solving a system of linear equations. We have applied our algorithm to several benchmark data sets and obtained promising results. Jianzhuang Liu, Xiaoou Tang |
CVPR | 3 |
| 2010 | Object cut: Complex 3D object reconstruction through line drawing separationabstractThis paper proposes an approach called object cut to tackle an important problem in computer vision, 3D object reconstruction from single line drawings. Given a complex line drawing representing a solid object, our algorithm finds the places, called cuts, to separate the line drawing into much simpler ones. The complex 3D object is obtained by first reconstructing the 3D objects from these simpler line drawings and then combining them together. Several propositions and criteria are presented for cut finding. A theorem is given to guarantee the existence and uniqueness of the separation of a line drawing along a cut. Our experiments show that the proposed approach can deal with more complex 3D object reconstruction than state-of-the-art methods. Tianfan Xue, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2010 | Continuous MRF based image denoising with a closed form solutionabstractIn this paper, we formulate the problem of image denoising as the maximum a posterior (MAP) estimation problem using Markov random fields (MRFs). Such an MAP estimation for MRFs is equivalent to a maximum likelihood estimation constrained on spatial homogeneity and is generally NP-hard in the discrete domain. To make it tractable, we convert it to a continuous label assignment problem based on a Gaussian MRF model and then obtain a closed form globally optimal solution. Since the Gaussian MRFs tend to over-smooth images and blur edges, we incorporate pre-estimated image edge information into the energy function to better preserve image structures. Patch similarity based pairwise interaction is also involved to better preserve image details and make the algorithm more robust to impulse noise. Both quantitative and qualitative comparative experimental results are given to demonstrate the better performance of our algorithm. Shifeng Chen, Jianzhuang Liu |
ICIP | 3 |
| 2010 | Clustering on Dependency Digraphs
Jianzhuang Liu |
ICIP | 3 |
| 2010 | Dimensionality Reduction via Tangential Learning
Jianzhuang Liu, Hwann-Tzong Chen |
ICIP | 3 |
| 2010 | Semi-supervised sparse metric learning using alternating linearization optimizationabstractIn plenty of scenarios, data can be represented as vectors and then mathematically abstracted as points in a Euclidean space. Because a great number of machine learning and data mining applications need proximity measures over data, a simple and universal distance metric is desirable, and metric learning methods have been explored to produce sensible distance measures consistent with data relationship. However, most existing methods suffer from limited labeled data and expensive training. In this paper, we address these two issues through employing abundant unlabeled data and pursuing sparsity of metrics, resulting in a novel metric learning approach called semi-supervised sparse metric learning. Two important contributions of our approach are: 1) it propagates scarce prior affinities between data to the global scope and incorporates the full affinities into the metric learning; and 2) it uses an efficient alternating linearization method to directly optimize the sparse metric. Compared with conventional methods, ours can effectively take advantage of semi-supervision and automatically discover the sparse metric structure underlying input data patterns. We demonstrate the efficacy of the proposed approach with extensive experiments carried out on six datasets, obtaining clear performance gains over the state-of-the-arts. Wei Liu 0005, Shiqian Ma, Dacheng Tao, Jianzhuang Liu |
KDD | 4 |
| 2010 | Fast image rearrangement via multi-scale patch copyingabstractIn this paper, we propose a simple interactive way for a novel type of image synthesis called image rearrangement whose goal is to construct a new image based on some objects cropped from source images. The synthesis results are obtained by copying patches from the source images in a globally consistent way. The patch copying problem is formulated with the Markov random field model, and belief propagation is used as the optimization tool. To speed up our algorithm, a two-step belief propagation and a multi-scale patch copying scheme are taken. Experimental results indicate that our algorithm obtains satisfactory results in both performance and efficiency. Jiayao Hu, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2010 | 3D object search through semantic componentabstractIn this paper, we present a novel concept named semantic component for 3D object search which describes a key component that semantically defines a 3D object. In most cases, the semantic component is intra-category stable and therefore can be used to construct an efficient 3D object retrieval scheme. By segmenting an object into segments and learning the similar segments shared by all the objects in the same category, we can summarise what human uses for object recognition, from the analysis of which we develop a method to find the semantic component of an object. In our experiments, the proposed method is justified and the effectiveness of our algorithm is also demonstrated. Chunjing Xu, Zhengwu Zhang, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2010 | Robust 3D Face Recognition by Local Shape Difference BoostingabstractThis paper proposes a new 3D face recognition approach, Collective Shape Difference Classifier (CSDC), to meet practical application requirements, i.e., high recognition performance, high computational efficiency, and easy implementation. We first present a fast posture alignment method which is self-dependent and avoids the registration between an input face against every face in the gallery. Then, a Signed Shape Difference Map (SSDM) is computed between two aligned 3D faces as a mediate representation for the shape comparison. Based on the SSDMs, three kinds of features are used to encode both the local similarity and the change characteristics between facial shapes. The most discriminative local features are selected optimally by boosting and trained as weak classifiers for assembling three collective strong classifiers, namely, CSDCs with respect to the three kinds of features. Different schemes are designed for verification and identification to pursue high performance in both recognition and computation. The experiments, carried out on FRGC v2 with the standard protocol, yield three verification rates all better than 97.9 percent with the FAR of 0.1 percent and rank-1 recognition rates above 98 percent. Each recognition against a gallery with 1,000 faces only takes about 3.6 seconds. These experimental results demonstrate that our algorithm is not only effective but also time efficient. Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Image Segmentation by MAP-ML EstimationsabstractImage segmentation plays an important role in computer vision and image analysis. In this paper, image segmentation is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs) and a graph cut algorithm is used to find the solution to the MAP estimation. The ML estimation is achieved by computing the means of region features in a Gaussian model. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. Its results match image edges very well and are consistent with human perception. Comparing to six state-of-the-art algorithms, extensive experiments have shown that our algorithm performs the best. Shifeng Chen, Liangliang Cao, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Image Process. | 4 |
| 2010 | Misalignment-Robust Face RecognitionabstractSubspace learning techniques for face recognition have been widely studied in the past three decades. In this paper, we study the problem of general subspace-based face recognition under the scenarios with spatial misalignments and/or image occlusions. For a given subspace derived from training data in a supervised, unsupervised, or semi-supervised manner, the embedding of a new datum and its underlying spatial misalignment parameters are simultaneously inferred by solving a constrained l1 norm optimization problem, which minimizes the l1 error between the misalignment-amended image and the image reconstructed from the given subspace along with its principal complementary subspace. A byproduct of this formulation is the capability to detect the underlying image occlusions. Extensive experiments on spatial misalignment estimation, image occlusion detection, and face recognition with spatial misalignments and/or image occlusions all validate the effectiveness of our proposed general formulation for misalignment-robust face recognition. Shuicheng Yan, Huan Wang 0001, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang |
IEEE Trans. Image Process. | 3 |
| 2010 | Local Derivative Pattern Versus Local Binary Pattern: Face Recognition With High-Order Local Pattern DescriptorabstractThis paper proposes a novel high-order local pattern descriptor, local derivative pattern (LDP), for face recognition. LDP is a general framework to encode directional pattern features based on local derivative variations. The n(th)-order LDP is proposed to encode the (n-1)(th) -order local derivative direction variations, which can capture more detailed information than the first-order local pattern used in local binary pattern (LBP). Different from LBP encoding the relationship between the central point and its neighbors, the LDP templates extract high-order local information by encoding various distinctive spatial relationships contained in a given local region. Both gray-level images and Gabor feature images are used to evaluate the comparative performances of LDP and LBP. Extensive experimental results on FERET, CAS-PEAL, CMU-PIE, Extended Yale B, and FRGC databases show that the high-order LDP consistently performs much better than LBP for both face identification and face verification under various conditions. Baochang Zhang 0001, Yongsheng Gao 0001, Sanqiang Zhao, Jianzhuang Liu |
IEEE Trans. Image Process. | 4 |
| 2009 | Constrained clustering via spectral regularizationabstractWe propose a novel framework for constrained spectral clustering with pairwise constraints which specify whether two objects belong to the same cluster or not. Unlike previous methods that modify the similarity matrix with pairwise constraints, we adapt the spectral embedding towards an ideal embedding as consistent with the pairwise constraints as possible. Our formulation leads to a small semidefinite program whose complexity is independent of the number of objects in the data set and the number of pairwise constraints, making it scalable to large-scale problems. The proposed approach is applicable directly to multi-class problems, handles both must-link and cannot-link constraints, and can effectively propagate pairwise constraints. Extensive experiments on real image data and UCI data have demonstrated the efficacy of our algorithm. Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2009 | 3D reconstruction of curved objects from single 2D line drawingsabstractAn important research area in computer vision is developing algorithms that can reconstruct the 3D surface of an object represented by a single 2D line drawing. Previous work on 3D reconstruction from single 2D line drawings focuses on objects with planar faces. In this paper, we propose a novel approach to the reconstruction of solid objects that have not only planar but also curved faces. Our approach consists of four steps: (1) identifying the curved faces and planar faces in a line drawing, (2) transforming the line drawing into one with straight edges only, (3) reconstructing the 3D wireframe of the curved object from the transformed line drawing and the original line drawing, and (4) generating the curved faces with Bezier patches and triangular meshes. With a number of experimental results, we demonstrate the ability of our approach to perform curved object reconstruction successfully. Yingze Wang, Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang |
CVPR | 3 |
| 2009 | Constrained clustering by spectral kernel learningabstractClustering performance can often be greatly improved by leveraging side information. In this paper, we consider constrained clustering with pairwise constraints, which specify some pairs of objects from the same cluster or not. The main idea is to design a kernel to respect both the proximity structure of the data and the given pairwise constraints. We propose a spectral kernel learning framework and formulate it as a convex quadratic program, which can be optimally solved efficiently. Our framework enjoys several desirable features: 1) it is applicable to multi-class problems; 2) it can handle both must-link and cannot-link constraints; 3) it can propagate pairwise constraints effectively; 4) it is scalable to large-scale problems; and 5) it can handle weighted pairwise constraints. Extensive experiments have demonstrated the superiority of the proposed approach. Zhenguo Li, Jianzhuang Liu |
ICCV | 2 |
| 2009 | Spectral Kernel Learning for Semi-Supervised Classification
Wei Liu 0005, Buyue Qian, Jingyu Cui, Jianzhuang Liu |
IJCAI | 4 |
| 2009 | Automatic facial expression recognition on a single 3D face by exploring shape deformationabstractFacial expression recognition has many applications in multimedia processing and the development of 3D data acquisition techniques makes it possible to identify expressions using 3D shape information. In this paper, we propose an automatic facial expression recognition approach based on a single 3D face. The shape of an expressional 3D face is approximated as the sum of two parts, a basic facial shape component (BFSC) and an expressional shape component (ESC). The BFSC represents the basic face structure and neutral-style shape and the ESC contains shape changes caused by facial expressions. To separate the BFSC and ESC, our method firstly builds a reference face for each input 3D non-neutral face by a learning method, which well represents the basic facial shape. Then, based on the BFSC and the original expressional face, a facial expression descriptor is designed. The surface depth changes are considered in the descriptor. Finally, the descriptor is input into an SVM to recognize the expression. Unlike previous methods which recognize a facial expression with the help of manually labeled key points and/or a neutral face, our method works on a single 3D face without any manual assistance. Extensive experiments are carried out on the BU-3DFE database and comparisons with existing methods are conducted. The experimental results show the effectiveness of our method. Boqing Gong, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2009 | Boosting 3D object retrieval by object flexibilityabstractIn this paper, we propose a novel feature, called object flexibility, at a point of a 3D object to describe how the neighborhood of this point is massively connected to the object. We show that this feature is stable to the deformation of objects' articulations, in addition to commonly concerned linear transforms, i.e., translation, scale, and rotation. A shape descriptor is obtained based on this feature using the bag-of-words model. As an application, the descriptor is used to perform 3D object retrieval. Extensive experiments demonstrate its superiority over a variety of existing 3D shape descriptors in the retrieval of articulated objects, as well as its enhancement of other shape descriptors to retrieve generic 3D objects. Boqing Gong, Chunjing Xu, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2009 | Video completion via motion guided spatial-temporal global optimizationabstractIn this paper, a novel global optimization based approach is proposed for video completion whose target is to restore the spatial-temporal missing regions of a video in a visually plausible way. Our algorithm consists of two stages: motion field completion and color completion via global optimization. First, local motions within the missing parts are completed patch-by-patch greedily using pre-computed available motions in the video. Then the missing regions are filled by sampling patches from available parts of the video. We formulate the video completion as a global energy minimization problem by Markov random fields (MRFs). Based on the completed motion field of the video, a well-defined energy function involving both spatial and temporal coherence relationship is constructed. A coarse-to-fine Belief Propagation (BP) is proposed to solve the optimization problem. Experimental results have demonstrated the good performance of our algorithm. Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2009 | Responses to the Comments on "What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines"abstractVarley (2009) made comments on our paper in (L. Cao et al., 2008) section by section. We answer them in this response paper. Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Responses to the Comments on "Plane-Based Optimization for 3D Object Reconstruction from Single Line Drawings"abstractWe disagree with the comments made by Varley [1] on our previous paper [2]. In this paper, we respond to his comments and show that they are not correct. Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | 2D Shape Matching by Contour FlexibilityabstractIn computer vision, shape matching is a challenging problem, especially when articulation and deformation of parts occur. These variations may be insignificant in terms of human recognition, but often cause a matching algorithm to give results that are inconsistent with our perception. In this paper, we propose a novel shape descriptor of planar contours, called contour flexibility, which represents the deformable potential at each point along a contour. With this descriptor, The local and global features can be obtained from the contour. We then present a shape matching scheme based on the features obtained. Experiments with comparisons to recently published algorithms show that our algorithm performs best. Chunjing Xu, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | A Theory of Phase Singularities for Image Representation and its Applications to Object Tracking and Image MatchingabstractThis paper studies phase singularities (PSs) for image representation. We show that PSs calculated with Laguerre-Gauss filters contain important information and provide a useful tool for image analysis. PSs are invariant to image translation and rotation. We introduce several invariant features to characterize the core structures around PSs and analyze the stability of PSs to noise addition and scale change. We also study the characteristics of PSs in a scale space, which lead to a method to select key scales along phase singularity curves. We demonstrate two applications of PSs: object tracking and image matching. In object tracking, we use the iterative closest point algorithm to determine the correspondences of PSs between two adjacent frames. The use of PSs allows us to precisely determine the motions of tracked objects. In image matching, we combine PSs and scale-invariant feature transform (SIFT) descriptor to deal with the variations between two images and examine the proposed method on a benchmark database. The results indicate that our method can find more correct matching pairs with higher repeatability rates than some well-known methods. Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Mitsou Takeda, Xiaoou Tang |
IEEE Trans. Image Process. | 4 |
| 2009 | Correspondence Propagation with Weak PriorsabstractFor the problem of image registration, the top few reliable correspondences are often relatively easy to obtain, while the overall matching accuracy may fall drastically as the desired correspondence number increases. In this paper, we present an efficient feature matching algorithm to employ sparse reliable correspondence priors for piloting the feature matching process. First, the feature geometric relationship within individual image is encoded as a spatial graph, and the pairwise feature similarity is expressed as a bipartite similarity graph between two feature sets; then the geometric neighborhood of the pairwise assignment is represented by a categorical product graph, along which the reliable correspondences are propagated; and finally a closed-form solution for feature matching is deduced by ensuring the feature geometric coherency as well as pairwise feature agreements. Furthermore, our algorithm is naturally applicable for incorporating manual correspondence priors for semi-supervised feature matching. Extensive experiments on both toy examples and real-world applications demonstrate the superiority of our algorithm over the state-of-the-art feature matching techniques. Huan Wang 0001, Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang |
IEEE Trans. Image Process. | 3 |
| 2008 | Clustering via Random Walk Hitting Time on Directed Graphs
Jianzhuang Liu, Xiaoou Tang |
AAAI | 2 |
| 2008 | MQSearch: image search by multi-class queryabstractImage search is becoming prevalent in web search as the number of digital photos grows exponentially on the internet. For a successful image search system, removing outliers in the top ranked results is a challenging task. Typical content based image search engines take an input image from one class as a query and compute relevance between the query and images in a database. The results often contain a large number of outliers, since these outliers may be similar to the query image in some way. In this paper we present a novel search scheme using query images from multiple classes. Instead of conducting query search for one image class at a time, we conduct multi-class query search jointly. By using several query classes that are similar to each other for multi-class query, we can utilize information across similar classes to fine tune the similarity measure to remove outliers. This strategy can be used for any information search application. In this work, we use content based image search to illustrate the concept. Yiwen Luo, Wei Liu 0005, Jianzhuang Liu, Xiaoou Tang |
CHI | 3 |
| 2008 | Sketching in the air: A vision-based system for 3D object designabstract3D object design has many applications including flexible 3D sketch input in CAD, computer game, webpage content design, image based object modeling, and 3D object retrieval. Most current 3D object design tools work on a 2D drawing plane such as computer screen or tablet, which is often inflexible with one dimension lost. On the other hand, virtual reality based methods have the drawbacks that there are awkward devices worn by the user and the virtual environment systems are expensive. In this paper, we propose a novel vision-based approach to 3D object design. Our system consists of a PC, a camera, and a mirror. We use the camera and mirror to track a wand so that the user can design 3D objects by sketching in 3D free space directly with out having to wear any cumbersome devices. A number of new techniques are developed for working in this system, including input of object wireframes, gestures for editing and drawing objects, and optimization-based planar and curved surface generation. Our system provides designers a new user interface for designing 3D objects conveniently. Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2008 | Misalignment-robust face recognitionabstractIn this paper, we study the problem of subspace-based face recognition under scenarios with spatial misalignments and/or image occlusions. For a given subspace, the embedding of a new datum and the underlying spatial misalignment parameters are simultaneously inferred by solving a constrained ℓ1norm optimization problem, which minimizes the error between the misalignment-amended image and the image reconstructed from the given subspace along with its principal complementary subspace. A byproduct of this formulation is the capability to detect the underlying image occlusions. Extensive experiments on spatial misalignment estimation, image occlusion detection, and face recognition with spatial misalignments and image occlusions all validate the effectiveness of our proposed general formulation. Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Jianzhuang Liu, Xiaoou Tang |
CVPR | 4 |
| 2008 | Output Regularized Metric Learning with Side Information
Wei Liu 0005, Steven C. H. Hoi, Jianzhuang Liu |
ECCV (3) | 3 |
| 2008 | 3D Face Recognition by Local Shape Difference Boosting
Yueming Wang 0001, Xiaoou Tang, Jianzhuang Liu, Gang Pan 0001, Rong Xiao 0003 |
ECCV (1) | 3 |
| 2008 | Phase singularities for image representation and matchingabstractPhase features are widely used in image processing and representation due to their stability to deformation and noise. However, phase singularities,where the signals vanish, are generally regarded as harmful and unreliable facts. In this paper, on the contrary, we will show that phase singularities calculated by Laguerre-Gauss filter contain important information of input image and can provide a reliable representation for image matching. We show that the positions of phase singularities are invariant to translation and rotation. Usually, it is possible to recover the input image up to a constant scaling only from the positions of phase singularities. We study phase singularities in scale space, which allows us to determine the "intrinsic scales" of key phase singularities. We introduce three physical measures of the local structures of phase singularities and combine these measures with SIFT descriptor for image matching. We execute experiments on benchmark database to examine the proposed methods. The results indicate that the proposed method can achieve comparable performance with certain well-known methods. Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Xiaoou Tang |
ICASSP | 4 |
| 2008 | Transductive Component AnalysisabstractIn this paper, we study semisupervised linear dimensionality reduction. Beyond conventional supervised methods which merely consider labeled instances, the semisupervised scheme allows to leverage abundant and ample unlabeled instances into learning so as to achieve better generalization performance. Under semisupervised settings, our objective is to learn a smooth as well as discriminative subspace and linear dimensionality reduction is thus achieved by mapping all samples into the subspace. Specifically, we present the transductive component analysis (TCA) algorithm to generate such a subspace founded on a graph-theoretic framework. Considering TCA is nonorthogonal, we further present the orthogonal transductive component analysis (OTCA) algorithm to iteratively produce a series of orthogonal basis vectors. OTCA has better discriminating power than TCA. Experiments carried out on synthetic and real-world datasets by OTCA show a clear improvement over the results of representative dimensionality reduction algorithms. Wei Liu 0005, Dacheng Tao, Jianzhuang Liu |
ICDM | 3 |
| 2008 | Pairwise constraint propagation by semidefinite programming for semi-supervised classificationabstractWe consider the general problem of learning from both pairwise constraints and unlabeled data. The pairwise constraints specify whether two objects belong to the same class or not, known as the must-link constraints and the cannot-link constraints. We propose to learn a mapping that is smooth over the data graph and maps the data onto a unit hypersphere, where two must-link objects are mapped to the same point while two cannot-link objects are mapped to be orthogonal. We show that such a mapping can be achieved by formulating a semidefinite programming problem, which is convex and can be solved globally. Our approach can effectively propagate pairwise constraints to the whole data set. It can be directly applied to multi-class classification and can handle data labels, pairwise constraints, or a mixture of them in a unified framework. Promising experimental results are presented for classification tasks on a variety of synthetic and real data sets. Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
ICML | 2 |
| 2008 | Precise object cutout from imagesabstractIn this paper we propose a novel approach to the problem of interactive foreground/background segmentation in images. With user provided strokes which indicate foreground and background seeds, we estimate two Gaussian mixture models, one for foreground and the other for background, and define two quantities to measure the initial probabilities of each pixel belonging to the foreground and the background respectively. An optimization function constructed based on the quantities and the boundary and coherent region information is proposed to solve the segmentation problem. By relaxing the hard binary segmentation to a soft labelling problem in the continuous domain, a closed form global optimal solution can be achieved, which directly results in the final binary segmentation output. Experimental results demonstrate the excellent performance of our algorithm. Shifeng Chen, Jianzhuang Liu |
ACM Multimedia | 3 |
| 2008 | What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden LinesabstractThe human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Many reconstruction methods have been proposed to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach to reconstructing a complete 3D object, including the shape of the back of the object, from a line drawing without hidden lines. First, we develop theoretical constraints and an algorithm for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We show a number of examples to demonstrate the success of our approach. Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Plane-Based Optimization for 3D Object Reconstruction from Single Line DrawingsabstractIn previous optimization-based methods of 3D planar-faced object reconstruction from single 2D line drawings, the missing depths of the vertices of a line drawing (and other parameters in some methods) are used as the variables of the objective functions. A 3D object with planar faces is derived by finding values for these variables that minimize the objective functions. These methods work well for simple objects with a small number N of variables. As N grows, however, it is very difficult for them to find expected objects. This is because with the nonlinear objective functions in a space of large dimension N, the search for optimal solutions can easily get trapped into local minima. In this paper, we use the parameters of the planes that pass through the planar faces of an object as the variables of the objective function. This leads to a set of linear constraints on the planes of the object, resulting in a much lower dimensional nullspace where optimization is easier to achieve. We prove that the dimension of this nullspace is exactly equal to the minimum number of vertex depths which define the 3D object. Since a practical line drawing is usually not an exact projection of a 3D object, we expand the nullspace to a larger space based on the singular value decomposition of the projection matrix of the line drawing. In this space, robust 3D reconstruction can be achieved. Compared with two most related methods, our method not only can reconstruct more complex 3D objects from 2D line drawings, but also is computationally more efficient. Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Regression From Uncertain Labels and Its Applications to Soft BiometricsabstractIn this paper, we investigate two soft-biometric problems: (1) age estimation and (2) pose estimation, within the scenario where uncertainties exist for the available labels of the training samples. These two tasks are generally formulated as the automatic design of a regressor from training samples with uncertain nonnegative labels. First, the nonnegative label is predicted as the Frobenius norm of a matrix, which is bilinearly transformed from the nonlinear mappings of a set of candidate kernels. Two transformation matrices are then learned for deriving such a matrix by solving two semidefinite programming (SDP) problems, in which the uncertain label of each sample is expressed as two inequality constraints. The objective function of SDP controls the ranks of these two matrices and, consequently, automatically determines the structure of the regressor. The whole framework for the automatic design of a regressor from samples with uncertain nonnegative labels has the following characteristics: (1) the SDP formulation makes full use of the uncertain labels, instead of using conventional fixed labels; (2) regression with the Frobenius norm of matrix naturally guarantees the nonnegativity of the labels, and greater prediction capability is achieved by integrating the squares of the matrix elements, which to some extent act as weak regressors; and (3) the regressor structure is automatically determined by the pursuit of simplicity, which potentially promotes the algorithmic generalization capability. Extensive experiments on two human age databases: (1) FG-NET and (2) Yamaha, and the Pointing'04 head pose database, demonstrate encouraging estimation accuracy improvements over conventional regression algorithms without taking the uncertainties within the labels into account. Shuicheng Yan, Huan Wang 0001, Xiaoou Tang, Jianzhuang Liu, Thomas S. Huang |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2007 | Iterative MAP and ML Estimations for Image SegmentationabstractImage segmentation plays an important role in computer vision and image analysis. In this paper, the segmentation problem is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum-likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs). A graph-cut algorithm is used to find the solution to the MAP-MRF estimation. The ML estimation is achieved by finding the means of region features. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. In addition, under the same framework, it can be extended to another algorithm that extracts objects of a particular class from a group of images. Extensive experiments have shown the effectiveness of our approach. Shifeng Chen, Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
CVPR | 3 |
| 2007 | A Closed-form Solution to 3D Reconstruction of Piecewise Planar Objects from Single ImagesabstractThis paper proposes a new approach to 3D reconstruction of piecewise planar objects based on two image regularities, connectivity and perspective symmetry. First, we formulate the whole shape of the objects in an image as a shape vector consisting of the normals of all the faces of the objects. Then, we impose several linear constraints on the shape vector using connectivity and perspective symmetry of the objects. Finally, we obtain a closed-form solution to the 3D reconstruction problem. We also develop an efficient algorithm to detect a face of perspective symmetry. Experimental results on real images are shown to demonstrate the effectiveness of our approach. Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2007 | Offline Signature Verification Using Online Handwriting RegistrationabstractThis paper proposes a novel framework for offline signature verification. Different from previous methods, our approach makes use of online handwriting instead of handwritten images for registration. The online registrations enable robust recovery of the writing trajectory from an input offline signature and thus allow effective shape matching between registration and verification signatures. In addition, we propose several new techniques to improve the performance of the new signature verification system: 1. we formulate and solve the recovery of writing trajectory within the framework of conditional random fields; 2. we propose a new shape descriptor, online context, for aligning signatures; 3. we develop a verification criterion which combines the duration and amplitude variances of handwriting. Experiments on a benchmark database show that the proposed method significantly outperforms the well-known offline signature verification methods and achieve comparable performance with online signature verification methods. Yu Qiao 0001, Jianzhuang Liu, Xiaoou Tang |
CVPR | 2 |
| 2007 | A Divide-and-Conquer Approach to 3D Object Reconstruction from Line Drawingsabstract3D object reconstruction from a single 2D line drawing is an important problem in both computer vision and graphics. Many methods have been put forward to solve this problem, but they usually fail when the geometric structure of a 3D object becomes complex. In this paper, a novel approach based on a divide-and-conquer strategy is proposed to handle 3D reconstruction of complex manifold objects from single 2D line drawings. The approach consists of three steps: 1) dividing a complex line drawing into multiple simpler line drawings based on the result efface identification; 2) reconstructing the 3D shapes from these simpler line drawings; and 3) merging the 3D shapes into one complete object represented by the original line drawing. A number of examples are given to show that our approach can handle 3D reconstruction of more complex objects than previous methods. Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang |
ICCV | 2 |
| 2007 | Noise Robust Spectral ClusteringabstractThis paper aims to introduce the robustness against noise into the spectral clustering algorithm. First, we propose a warping model to map the data into a new space on the basis of regularization. During the warping, each point spreads smoothly its spatial information to other points. After the warping, empirical studies show that the clusters become relatively compact and well separated, including the noise cluster that is formed by the noise points. In this new space, the number of clusters can be estimated by eigenvalue analysis. We further apply the spectral mapping to the data to obtain a low-dimensional data representation. Finally, the K-means algorithm is used to perform clustering. The proposed method is superior to previous spectral clustering methods in that (i) it is robust against noise because the noise points are grouped into one new cluster; (ii) the number of clusters and the parameters of the algorithm are determined automatically. Experimental results on synthetic and real data have demonstrated this superiority. Zhenguo Li, Jianzhuang Liu, Shifeng Chen, Xiaoou Tang |
ICCV | 2 |
| 2007 | Transductive regression piloted by inter-manifold relationsabstractIn this paper, we present a novel semisupervised regression algorithm working on multiclass data that may lie on multiple manifolds. Unlike conventional manifold regression algorithms that do not consider the class distinction of samples, our method introduces the class information to the regression process and tries to exploit the similar configurations shared by the label distribution of multi-class data. To utilize the correlations among data from different classes, we develop a cross-manifold label propagation process and employ labels from different classes to enhance the regression performance. The interclass relations are coded by a set of intermanifold graphs and a regularization item is introduced to impose inter-class smoothness on the possible solutions. In addition, the algorithm is further extended with the kernel trick for predicting labels of the out-of-sample data even without class information. Experiments on both synthesized data and real world problems validate the effectiveness of the proposed framework for semisupervised regression. Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Jianzhuang Liu, Xiaoou Tang |
ICML | 4 |
| 2007 | Bayesian Tensor Inference for Sketch-Based Facial Photo Hallucination
Wei Liu 0005, Xiaoou Tang, Jianzhuang Liu |
IJCAI | 3 |
| 2007 | Image matting using linear optimizationabstractAn image can be assumed to be a composite of the foreground and the background. The foreground and the background of each pixel are linearly combined in terms of this pixel's foreground opacity (called alpha). Image matting is the process of estimating the foreground, the background and the alpha for each pixel. In this paper, we transform the ill-posed image matting problem into two over-determined linear optimization problems by introducing two medium variables and imposing smoothness constraints. Closed form solutions can be obtained from the two problems. Extensive experimental results indicate that our algorithm can generate high-quality matting results. Shifeng Chen, Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2007 | Image inpainting by global structure and texture propagationabstractImage inpainting is a technique to repair damaged images or modify images in a non-detectable form. In this paper, a novel global algorithm for region filling is proposed for image inpainting. After removing objects from an image, our approach fills the regions using patches taken from the image. The filling process is formulated as an energy minimization problem by Markov random fields (MRFs) and the belief propagation (BP) is utilized to solve the problem. Our energy function includes structure and texture information obtained from the image. One challenge in using BP is that its computational complexity is the square of the number of label candidates. To reduce the large number of label candidates, we present a coarse-to-fine scheme where two BPs run with much smaller numbers of label candidates instead of one BP running with a large number of label candidates. Experimental results demonstrate the excellent performance of our algorithm over other related algorithms. Huang Ting, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2007 | Ranking with uncertain labels and its applications
Shuicheng Yan, Huan Wang 0001, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang |
Frontiers Comput. Sci. China | 3 |
| 2007 | A Parameter-Free Framework for General Supervised Subspace LearningabstractSupervised subspace learning techniques have been extensively studied in biometrics literature; however, there is little work dedicated to: 1) how to automatically determine the subspace dimension in the context of supervised learning, and 2) how to explicitly guarantee the classification performance on a training set. In this paper, by following our previous work on unified subspace learning framework in our earlier work, we present a general framework, called parameter-free graph embedding (PFGE) to solve the above two problems by posing a general supervised subspace learning task as a semidefinite programming problem. The semipositive feature Gram matrix, namely the product of the transformation matrix and its transpose, is derived by optimizing a trace difference form of an objective function extended from that in our earlier work with the constraints that guarantee the class homogeneity within the neighborhood of each datum. Then, the subspace dimension and the feature weights are simultaneously obtained via the singular value decomposition of the feature Gram matrix. In addition, to alleviate the computational complexity, the Kronecker product approximation of the feature Gram matrix is proposed by taking advantage of the essential matrix form of image pixels. The experiments on simulated data and real-world data demonstrate the capability of the new PFGE framework in estimating the subspace dimension for supervised learning as well as the superiority in classification performance over traditional algorithms for subspace learning Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | Formulating Face Verification With Semidefinite ProgrammingabstractThis paper presents a unified solution to three unsolved problems existing in face verification with subspace learning techniques: selection of verification threshold, automatic determination of subspace dimension, and deducing feature fusing weights. In contrast to previous algorithms which search for the projection matrix directly, our new algorithm investigates a similarity metric matrix (SMM). With a certain verification threshold, this matrix is learned by a semidefinite programming approach, along with the constraints of the kindred pairs with similarity larger than the threshold, and inhomogeneous pairs with similarity smaller than the threshold. Then, the subspace dimension and the feature fusing weights are simultaneously inferred from the singular value decomposition of the derived SMM. In addition, the weighted and tensor extensions are proposed to further improve the algorithmic effectiveness and efficiency, respectively. Essentially, the verification is conducted within an affine subspace in this new algorithm and is, hence, called the affine subspace for verification (ASV). Extensive experiments show that the ASV can achieve encouraging face verification accuracy in comparison to other subspace algorithms, even without the need to explore any parameters. Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang |
IEEE Trans. Image Process. | 2 |
| 2006 | Degen Generalized Cylinders and Their Properties
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
ECCV (1) | 2 |
| 2006 | 3D object retrieval using 2D line drawing and graph based relevance reedbackabstractThis paper aims to provide a user-friendly interface for 3D object retrieval. In previous 3D retrieval systems, the user mainly uses two methods to input a query: providing an existing 3D objects, or providing partial shape information of desired objects such as text and 2D shapes. The first method fails when the user does not have a similar 3D object in hand, and the second method cannot sufficiently describe 3D shapes of objects. We believe that the best way is to have a good interface that can convert a 2D sketch drawn by the user into a 3D object as the query. A 2D line drawing is easy to be drawn and is the simplest and most direct way of illustrating a 3D object. In this paper, we develop an interface of 3D object reconstruction from line drawings, which allows the user to draw line drawings of objects with both planar and curved surfaces. In addition, in order to refine the retrieved results, we develop a relevance feedback algorithm based on a novel graph discriminant analysis. Compared with recently published relevance feedback algorithms, our algorithm achieves better retrieval performance. Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2006 | Shape from regularities for interactive 3D reconstruction of piecewise planar objects from single imagesabstract3D object reconstruction from single 2D images has many applications in multimedia. This paper proposes an approach based on image regularities such as connectivity, parallelism, and orthogonality possessed by the objects with simple user interactions. It is assumed that the objects are piecewise planar. By representing the 3D objects as a shape vector consisting of the normals of the faces of the objects, we impose geometric constraints on this shape vector using the regularities of the objects. We derive a system of equations in terms of the shape vector and the focal length, which we can solve for the shape vector optimally. Experimental results on real images are shown to demonstrate the effectiveness of this method. Zhenguo Li, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2006 | Detecting irregularity in videos using kernel estimation and KD treesabstractAutomatic event understanding is the ultimate goal for many visual surveillance systems. In this paper, we propose a novel approach for on-line detecting unusual human activities in videos without the need to explicitly define all valid configurations. Within the framework of Bayesian inference, the detection process is formulated as an MAP estimation where we attempt to find whether activities in new video segments have similar activities in a video database. Our approach has three contributions: firstly, we build the statistical representation of normal behaviors in the database using nonparametric kernel density estimation; secondly, local feature descriptors are highly compressed using PCA and stored in a K-D tree structure, making the search for behavior-based similarity fast and effective; thirdly, the K-D trees are used to generate multiple hypotheses which compete for the optimal classification. The approach requires no tracking, no explicit motion estimation, and no predefined class-based templates. Experimental results have validated our approach in many real-world video sequences. Chunjing Xu, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 3 |
| 2005 | 3D Object Reconstruction from a Single 2D Line Drawing without Hidden LinesabstractThe human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Several reconstruction approaches have tried to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach for reconstructing complete 3D objects from line drawings without hidden lines. First, we develop some constraints and properties for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We give a number of examples to demonstrate the ability of our approach. Liangliang Cao, Jianzhuang Liu, Xiaoou Tang |
ICCV | 2 |
| 2005 | Evolutionary Search for Faces from Line DrawingsabstractSingle 2D line drawing is a straightforward method to illustrate 3D objects. The faces of an object depicted by a line drawing give very useful information for the reconstruction of its 3D geometry. Two recently proposed methods for face identification from line drawings are based on two steps: finding a set of circuits that may be faces and searching for real faces from the set according to some criteria. The two steps, however, involve two combinatorial problems. The number of the circuits generated in the first step grows exponentially with the number of edges of a line drawing. These circuits are then used as the input to the second combinatorial search step. When dealing with objects having more faces, the combinatorial explosion prevents these methods from finding solutions within feasible time. This paper proposes a new method to tackle the face identification problem by a variable-length genetic algorithm with a novel heuristic and geometric constraints incorporated for local search. The hybrid GA solves the two combinatorial problems simultaneously. Experimental results show that our algorithm can find the faces of a line drawing having more than 30 faces much more efficiently. In addition, simulated annealing for solving the face identification problem is also implemented for comparison. Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Video-based handwritten Chinese character recognition
Xiaoou Tang, Feng Lin 0002, Jianzhuang Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Insignificant shadow detection for video segmentationabstractTo prevent moving cast shadows from being misunderstood as part of moving objects in change detection based video segmentation, this paper proposes a novel approach to the cast shadow detection based on the edge and region information in multiple frames. First, an initial change detection mask containing moving objects and cast shadows is obtained. Then a Canny edge map is generated. After that, the shadow region is detected and removed through multiframe integration, edge matching, and region growing. Finally, a post processing procedure is used to eliminate noise and tune the boundaries of the objects. Our approach can be used for video segmentation in indoor environment. The experimental results demonstrate its good performance. Dong Xu 0001, Jianzhuang Liu, Xuelong Li 0001, Zhengkai Liu, Xiaoou Tang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Efficient Search of Faces from Complex Line Drawings
Jianzhuang Liu, Xiaoou Tang |
CVPR (2) | 1 |
| 2004 | Indoor shadow detection for video segmentationabstractTo prevent moving cast shadows from being misunderstood as part of moving objects in change detection based video segmentation, this paper proposes a novel approach to cast shadow detection based on the edge and region information in multiple frames. First, an initial change detection mask containing moving objects and cast shadows is obtained. Then, a Canny edge map is generated. After that, the shadow region is detected and removed through multi-frame integration, edge matching and region growing. Finally, a post-processing procedure is used to eliminate noise and tune the boundaries of the objects. Our approach can be used for removing moving cast shadows in indoor environments for better video segmentation. The experimental results demonstrate the good performance of our algorithm. Dong Xu 0001, Jianzhuang Liu, Zhengkai Liu, Xiaoou Tang |
ICME | 2 |
| 2003 | Video caption detection and extraction using temporal informationabstractVideo caption detection and extraction is an important step for information retrieval in video databases. In this paper, we extract text information in video by fully utilizing the temporal information contained in the video. First we create a binary abstract sequence from a video segment. By analyzing the statistical pixel changes in the sequence, we can effectively locate the (dis)appealing frames of captions. Finally we extract the captions to create a summary of the video segment. Bo Luo, Xiaoou Tang, Jianzhuang Liu, HongJiang Zhang |
ICIP (1) | 3 |
| 2002 | Identifying Faces in a 2D Line Drawing Representing a Manifold ObjectabstractA straightforward way to illustrate a 3D model is to use a line drawing. Faces in a 2D line drawing provide important information for reconstructing its 3D geometry. Manifold objects belong to a class of common solids and most solid systems are based on manifold geometry. In this paper, a new method is proposed for finding faces from single 2D line drawings representing manifolds. The face identification is formulated based on a property of manifolds: each edge of a manifold is shared exactly by two faces. The two main steps in our method are (1) searching for cycles from a line drawing and (2) searching for faces from the cycles. In order to speed up the face identification procedure, a number of properties, most of which relate to planar manifold geometry in line drawings, are presented to identify most of the cycles that are or are not real faces in a drawing, thus reducing the number of unknown cycles in the second searching. Schemes to deal with manifolds with curved faces and manifolds each represented by two or more disjoint graphs are also proposed. The experimental results show that our method can handle manifolds previous methods can handle, as well as those they cannot. Jianzhuang Liu, Yong Tsui Lee, Wai-kuen Cham |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | A spatial-temporal approach for video caption detection and recognitionabstractWe present a video caption detection and recognition system based on a fuzzy-clustering neural network (FCNN) classifier. Using a novel caption-transition detection scheme we locate both spatial and temporal positions of video captions with high precision and efficiency. Then employing several new character segmentation and binarization techniques, we improve the Chinese video-caption recognition accuracy from 13% to 86% on a set of news video captions. As the first attempt on Chinese video-caption recognition, our experiment results are very encouraging. Xiaoou Tang, Xinbo Gao 0001, Jianzhuang Liu, HongJiang Zhang |
IEEE Trans. Neural Networks | 3 |
| 2001 | A Graph-Based Method for Face Identification from a Single 2D Line DrawingabstractThe faces in a 2D fine drawing of an object provide important information for the reconstruction of its 3D geometry. In this paper, a graph-based optimization method is proposed for identifying the faces is a line drawing. The face identification is formulated as a maximum weight clique problem. This formulation is proven to be equivalent to the formulation proposed by Shpitalni and Upson (1996). The advantage of our formulation is that it enables one to develop a much faster algorithm to find the faces in a drawing. The significant improvement in speed is derived from two algorithms provided: the depth-first graph search for quickly generating possible faces from a drawing; and the maximum weight clique finding for obtaining the optimal face configurations of the drawing. The experimental results shown that our algorithm generates the same results of face identification as Shpitalni and Lipson's method, but is much faster when dealing with objects of more than 20 faces. Jianzhuang Liu, Yong Tsui Lee |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Handwritten Chinese Character Recognition Through a Video CameraabstractWe propose a handwritten Chinese character recognition system using a video camera. The system combines the advantages of both online and offline approaches. It allows users to write on any regular paper just like using an off-line system. At the same time, using a video camera attached on the computer, the system can capture the stroke temporal information similar to an online system. Xiaoou Tang, Cheung Yung Chan, Jianzhuang Liu |
ICIP | 3 |
| 1998 | Model-based edge reconstruction for low bit-rate wavelet-based image codingabstractLow bit-rate image coding brings about an obvious degradation to the compressed images, among which, the distortions at the edges are particular objectionable. A model-based edge reconstruction algorithm is proposed for wavelet-based image coding at low bit-rate. Our approach applies a general model to represent the variety of edges existing in an image. Based on this model, the problem of edge reconstruction is formulated as finding the original edge model parameters from the lossy image. The proposed method is able to improve the subjective visual quality and fidelity (PSNR) of images coded by wavelet-based coding using zerotree quantization. Wai-kuen Cham, Jianzhuang Liu |
ICASSP | 3 |
| 1997 | On-Line Chinese Character Recognition by Incorporating Human KnowledgeabstractA stroke order and number free method for on-line recognition of Chinese characters is proposed. Both input characters and the model characters are represented with complete attributed relational graphs (ARGs). The ARGs of the model base are built according to the human knowledge of the segment relation structure of Chinese characters. For the recognition purpose, an optimal matching measure between two ARGs is defined, and the graph matching is formulated as a search problem of finding the minimum cost path in a state space tree, using the A * algorithm. Moreover, to reduce the search time of the A *, besides a heuristic estimate, a novel strategy is employed which again uses the human knowledge of the segment position structure of Chinese characters to prune the tree. Our experimental results demonstrate the efficiency of the proposed method. Jianzhuang Liu, Wai-kuen Cham, Michael Ming Yuen Chang |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 1996 | Stroke order and stroke number free on-line Chinese character recognition using attributed relational graph matchingabstractA structural method for on-line recognition of Chinese characters is proposed, which is stroke order and stroke number free. Both input characters and the model characters are represented with complete attributed relational graphs (ARGs). A new optimal matching measure between two ARGs is defined. Classification of an input character can be implemented by matching its ARG against every ARG of the model base. The matching procedure is formulated as a search problem of finding the minimum cost path in a state space tree, using the A* algorithm. In order to speed up the search of the A*, besides a heuristic estimate, a novel strategy that utilizes the geometric position information of stroke segments of Chinese characters to prune the tree is employed. The efficiency of our method is demonstrated by the promising experimental results. Jianzhuang Liu, Wai-kuen Cham, Michael Ming Yuen Chang |
ICPR | 1 |
| 1994 | Fuzzy C-Means Clustering Algorithm with Two Layers and its Application to Image Segmentation Based on Two-Dimensional HistogramabstractThis paper presents a fast fuzzy c-means (FCM) clustering algorithm with two layers, which is a mergence of hard clustering and fuzzy clustering. The result of hard clustering is used to initialize the c cluster centers in fuzzy clustering, and then the number of iteration steps is reduced. The application of the proposed algorithm to image segmentation based on the two dimensional histogram is provided to show its computational efficience. Weixin Xie, Jianzhuang Liu |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |