EDBT 2026 Demo / reviewers in the wild / expert
Honggang Zhang 0002
dblp:82/1228-2
· DBLP profile ↗
88ranked-venue papers
4as first author
28since 2021 · last 2026
0000-0001-8287-6783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 68 · 1 first-author · 20 since 2021Artificial intelligence and machine learning · 45 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicabstractExisting state-of-the-art symbolic music generation models represent symbolic music as a sequence of attribute tokens with fixed unidirectional dependencies.However, from the perspective of music theory, the attributes of a musical note are inherently a set rather than a sequence.Building on this insight, we propose Amadeus, a novel symbolic music generation framework that adopts a two-level architecture: an autoregressive model for note sequences and a bidirectional discrete diffusion model for note attributes.This design enables flexible attribute control and adjustable decoding speed during inference.To further enhance sequential modeling, we introduce the Conditional Information Enhancement Module (CIEM).We also constructed AMD (Amadeus MIDI Dataset)-the largest open-source symbolic music dataset to date-supporting both pre-training and finetuning.We trained two models of different scales, Amadeus and Amadeus-M, and conducted extensive experiments, demonstrating substantial improvements over state-of-the-art methods across both objective and subjective metrics. Hongju Su, Ke Li 0004, Lan Yang 0014, Honggang Zhang 0002, Yi-Zhe Song |
ACL (1) | 4 |
| 2025 | VersaGen: Unleashing Versatile Visual Control for Text-to-Image SynthesisabstractDespite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results. Lan Yang 0014, Yonggang Qi, Honggang Zhang 0002, Kaiyue Pang, Ke Li 0004, Yi-Zhe Song |
AAAI | 4 |
| 2025 | V-Oracle: Making Progressive Reasoning in Deciphering Oracle Bones for You and MeabstractOracle Bone Script (OBS) is a vital treasure of human civilization, rich in insights from ancient societies. However, the evolution of written language over millennia complicates its decipherment. In this paper, we propose V-Oracle, an innovative framework that utilizes Large Multi-modal Models (LMMs) for interpreting OBS. V-Oracle applies principles of pictographic character formation and frames the task as a visual question-answering (VQA) problem, establishing a multi-step reasoning chain. It proposes a multi-dimensional data augmentation for synthesizing high-quality OBS samples, and also implements a multi-phase oracle alignment tuning to improve LMMs’ visual reasoning capabilities. Moreover, to bridge the evaluation gap in the OBS field, we further introduce Oracle-Bench, a comprehensive benchmark that emphasizes process-oriented assessment and incorporates both standard and out-of-distribution setups for realistic evaluation. Extensive experimental results can demonstrate the effectiveness of our method in providing quantitative analyses and superior deciphering capability. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, MinhuiWu MinhuiWu, Jiapeng Wang 0005, Zhuoma Gongque, Yadong Xue, Zhimin Bao, Lan Yang 0014, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 15 |
| 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?abstractRunqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, Minhui Wu, Xiaoshuai Song, Jiapeng Wang 0005, Zhuoma Gongque, Shanglin Lei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Peiqing Yang 0003, Zhimin Bao, Muxi Diao, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 20 |
| 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio GenerationabstractRecently, diffusion models have achieved great success in mono-channel audio generation.
However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions.
Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models.
To the best of our knowledge, this work represents the first attempt to address these issues.
We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources.
Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation.
Existing audio generation models tend to generate rather random and indistinct spatial audio.
To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.
By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt.
Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods.
The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules. Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 0006, Huadai Liu, Honggang Zhang 0002, Wei Xue 0002, Yike Guo |
ICLR | 6 |
| 2025 | OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical FeedbackabstractRecent advancements in Large Multimodal Models have demonstrated impressive performance in various tasks. However, their capabilities in error detection and resolution for Optical Character Recognition (OCR) remain underexplored. To address this gap, we construct the first visual instruction tuning dataset specifically for detailed OCR error analysis. Building on this foundation, we develop a universal, plug-and-play OCR-Critic model that incorporates three novel dynamic alignment strategies. These strategies systematically mitigate LMMs' weaknesses in OCR tasks by providing coarse-to-fine error feedback. To comprehensively evaluate these capabilities, we introduce OCR-ERROR, a benchmark designed to assess LMMs' ability to detect and categorize OCR errors, covering two task types, diverse error categories, and 2,400 rigorously validated samples. Experimental results show that OCR-Critic effectively identifies fine-grained OCR errors across multiple domains. With the integration of our dynamic alignment strategies, the LMM further achieves substantial performance gains on four prominent benchmarks, demonstrating both versatility and effectiveness. Qiuna Tan, Runqi Qiao, Guanting Dong 0001, Minhui Wu, Jiapeng Wang 0005, Miaoxuan Zhang, Chen Li 0031, Honggang Zhang 0002 |
ACM Multimedia | 11 |
| 2025 | Parameter-Efficient Adaptation of Vision-Language Models for Free-Hand Sketch RecognitionabstractHow to prompt a foundation model like CLIP towards a sketch expert is the question we seek to answer in this paper. Debates on the best way to prompt have been intense and divided, however converged on one particular point that of modelling prompt learning as context token optimisation. This paper scrutinises such technical route for sketch and argues the challenge is more than a stereotyped ask from context change. In particular, we pin down the problem to the dramatic cross-modality gap between sketch and the photo-centric visual world formed within CLIP. We first show through a pilot study that relocating a sketched object to a different spatial locality can significantly improve zero-shot CLIP performance on sketch. Our core contribution is then to regard spatial misalignment as the key to explaining poor sketch adaptation in CLIP prompts – that a sketched object does not reside in a place as if it were part of the scene compositions of photo. Methodologically, we leverage a lightweight network that explicitly allows differentiable spatial manipulation of sketch data and design regulatory self-supervised signals to encourage proper convergence. We showcase consistent complementary power of this simple approach by building on top of 10 existing contemporary prompting methods on the sketch recognition task. For example, we outperform the strong prompting baseline CoOp by 2.57%, MaPle by 4.83% and AdaptFormer by 5.07%. Notably, the latter two beat the traditional full parameter fine-tuning (82.98%83.39% vs. 81.51%), and does so with less than 1% of the total training parameters. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
VCIP | 4 |
| 2024 | Making Visual Sense of Oracle Bones for You and MeabstractVisual perception evolves over time. This is particularly the case of oracle bone scripts, where visual glyphs seem intuitive to people from distant past prove difficult to be understood in contemporary eyes. While semantic correspon-dence of an oracle can be found via a dictionary lookup, this proves to be not enough for public viewers to connect the dots, i.e., why does this oracle mean that? Common solution relies on a laborious curation process to collect visual guide for each oracle (Fig. 1), which hinges on the case-by-case effort and taste of curators. This paper delves into one natural follow-up question: can AI take over? Begin with a comprehensive human study, we show par-ticipants could indeed make better sense of an oracle glyph subjected to a proper visual guide and its efficacy can be approximated via a novel metric termed TransOV (Trans-ferable Oracle Visuals). We then define a new conditional visual generation task based on an oracle glyph and its se-mantic meaning and importantly approach it by circumventing any form of model training in the presence of fatal lack of oracle data. At its heart is to leverage foundation model like GPT-4V to reason about the visual cues hidden inside an oracle and take advantage of an existing text-to-image model for final visual guide generation. Extensive empirical evidence shows our AI-enabled visual guides achieve signif-icantly comparable TransOV performance compared with those collected under manual efforts. Finally, we demon-strate the versatility of our system under a more complex setting, where it is required to work alongside with an AI image denoiser to cope with raw oracle scan image inputs (cf processed clean oracle glyphs). Code is available at https://github.com/RQ-Lab/OBS-Visual. Runqi Qiao, Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002 |
CVPR | 4 |
| 2024 | Wired Perspectives: Multi-View Wire Art Embraces Generative AIabstractCreating multi-view wire art (MVWA), a static 3D sculpture with diverse interpretations from different viewpoints, is a complex task even for skilled artists. In response, we present DreamWire, an AI system enabling everyone to craft MVWA easily. Users express their vision through text prompts or scribbles, freeing them from intricate 3D wire organisation. Our approach synergises 3D Bézier curves, Prim's algorithm, and knowledge distillation from diffusion models or their variants (e.g., ControlNet). This blend enables the system to represent 3D wire art, ensuring spatial continuity and overcoming data scarcity. Extensive evaluation and analysis are conducted to shed insight on the inner workings of the proposed system, including the trade-off between connectivity and visual aesthetics. Zhiyu Qu, Lan Yang 0014, Honggang Zhang 0002, Tao Xiang 0002, Kaiyue Pang, Yi-Zhe Song |
CVPR | 3 |
| 2024 | Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 4 |
| 2024 | Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li 0001, Honggang Zhang 0002, Di Hu 0001 |
ECCV (74) | 5 |
| 2024 | Enhancing Few-shot Classification through Token Selection for Balanced LearningabstractIn recent years, patch-based approaches have shown promise in few-shot learning, with further improvements observed through the use of self-supervised learning. However, we observe that the mainstream object-oriented approach focuses mainly on the salient part of the subject and ignores the non-annotated part of the image. Based on the assumption that any patch of the image is beneficial to learning, we present an end-to-end learning framework, which reconsiders the whole image from a multi-level perspective. The learning of annotated subjects involves Direct Patch Learning (DPL) to promote balanced learning of different features and Gaussian Mixup (GMIX) to provide extra mixed patch-level labels. As for the non-annotated part, we utilize a cascading token selection strategy along with self-supervised learning to better utilize knowledge in the background in the current context by learning the consistent representation of different views from the same image. Finally, in inductive few-shot learning, our method outperforms many previous methods and achieves new state-of-the-art performance. Furthermore, it provides an insight that non-annotated parts are also favorable for few-shot learning. As an ablation study, the effectiveness of each designed component is verified. Wangding Zeng, Peiwen Sun, Honggang Zhang 0002 |
IJCNN | 3 |
| 2024 | Unveiling and Mitigating Bias in Audio Visual SegmentationabstractCommunity researchers have developed a range of advanced audio-visual segmentation models aimed at improving the quality of sounding objects' masks. While masks created by these models may initially appear plausible, they occasionally exhibit anomalies with incorrect grounding logic. We attribute this to real-world inherent preferences and distributions as a simpler signal for learning than the complex audio-visual grounding, which leads to the disregard of important modality information. Generally, the anomalous phenomena are often complex and cannot be directly observed systematically. In this study, we made a pioneering effort with the proper synthetic data to categorize and analyze phenomena as two types "audio priming bias" and "visual prior" according to the source of anomalies. For audio priming bias, to enhance audio sensitivity to different intensities and semantics, a perception module specifically for audio perceives the latent semantic information and incorporates information into a limited set of queries, namely active queries. Moreover, the interaction mechanism related to such active queries in the transformer decoder is customized to adapt to the need for interaction regulating among audio semantics. For visual prior, multiple contrastive training strategies are explored to optimize the model by incorporating a biased branch, without even changing the structure of the model. During experiments, observation demonstrates the presence and the impact that has been produced by the biases of the existing model. Finally, through experimental evaluation of AVS benchmarks, we demonstrate the effectiveness of our methods in handling both types of biases, achieving competitive performance across all three subsets. Peiwen Sun, Honggang Zhang 0002, Di Hu 0001 |
ACM Multimedia | 2 |
| 2024 | SketchAnimator: Animate Sketch via Motion Customization of Text-to-Video Diffusion ModelsabstractSketching is a uniquely human tool for expressing ideas and creativity. The animation of sketches infuses life into these static drawings, opening a new dimension for designers. Animating sketches is a time-consuming process that demands professional skills and extensive experience, often proving daunting for amateurs. In this paper, we propose a novel sketch animation model SketchAnimator, which enables adding creative motion to a given sketch, like "a jumping car". Namely, given an input sketch and a reference video, we divide the sketch animation into three stages: Appearance Learning, Motion Learning and Video Prior Distillation. In stages 1 and 2, we utilize LoRA to integrate sketch appearance information and motion dynamics from the reference video into the pre-trained T2V model. In the third stage, we utilize Score Distillation Sampling (SDS) to update the parameters of the Bézier curves in each sketch frame according to the acquired motion information. Consequently, our model produces a sketch video that not only retains the original appearance of the sketch but also mirrors the dynamic movements of the reference video. We compare our method with alternative approaches and demonstrate that it generates the desired sketch video under the challenge of one-shot motion customization. Ruolin Yang 0001, Da Li 0001, Honggang Zhang 0002, Yi-Zhe Song |
VCIP | 3 |
| 2024 | Annotation-Free Human Sketch Quality Assessment
Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
Int. J. Comput. Vis. | 3 |
| 2023 | Sketch-based Video Object Segmentation: Benchmark and Analysis
Ruolin Yang 0001, Da Li 0001, Conghui Hu, Timothy M. Hospedales, Honggang Zhang 0002, Yi-Zhe Song |
BMVC | 5 |
| 2023 | A Method of Audio-Visual Person Verification by Mining Connections between Time Series
Peiwen Sun, Zishan Liu, Yougen Yuan, Taotao Zhang, Honggang Zhang 0002, Pengfei Hu 0004 |
INTERSPEECH | 6 |
| 2022 | Finding Badly Drawn BunniesabstractAs lovely as bunnies are, your sketched version would probably not do it justice (Fig. 1). This paper recognises this very problem and studies sketch quality measurement for the first time - letting you find these badly drawn ones. Our key discovery lies in exploiting the magnitude ($L$2norm) of a sketch feature as a quantitative quality metric. We propose Geometry-Aware Classification Layer (GACL), a generic method that makes feature-magnitude-as-quality-metric possible and importantly does it without the need for specific quality annotations from humans. GACL sees feature magnitude and recognisability learning as a dual task, which can be simultaneously optimised under a neat crossentropy classification loss. GACL is lightweight with theoretic guarantees and enjoys a nice geometric interpretation to reason its success. We confirm consistent quality agreements between our GACL-induced metric and human perception through a carefully designed human study. Notably, we demonstrate three practical sketch applications enabled for the first time using our quantitative quality metric. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
CVPR | 3 |
| 2022 | Mining Diverse Clues with Transformers for Person Re-identification
Jin Feng, Tianming Du 0001, Honggang Zhang 0002 |
PRCV (3) | 4 |
| 2022 | ComLoss: A Novel Loss Towards More Compact Predictions for Pedestrian Detection
Jin Feng, Tianming Du 0001, Honggang Zhang 0002 |
PRCV (4) | 4 |
| 2022 | MCascade R-CNN: A Modified Cascade R-CNN for Detection of Calcified on Coronary Artery Angiography ImagesabstractAmong cardiovascular diseases, coronary artery calcification (CAC) is a high-risk factor for worsening protopathy and increased mortality. However, the coronary artery an-giogram, which is the main approach for CAC diagnosis, suffers from plenty of photographing noise. This brings difficulties to detect calcification from the background. In this paper, a modified Cascade R-CNN (MCascade R-CNN) network is proposed to deal with the problem of calcium detection in angiograms. In the proposed network, we propose an innovative balanced aggregation pyramid structure, integrating multi-level features of every depth in the feature map, based on enhanced propagation of strong semantic features. In addition, a new convolutional attention mechanism is designed to improve the performance of the detector. Experiments show that the proposed method enjoys better performance in detecting and marking CAC in angiograms, Wei Wang 0335, Honggang Zhang 0002, Lihua Xie 0003, Bo Xu 0002 |
VCIP | 4 |
| 2022 | A study of deep single sketch-based modeling: View/style invariance, sparsity and latent space disentanglement
Yulia Gryaditskaya, Honggang Zhang 0002, Yi-Zhe Song |
Comput. Graph. | 3 |
| 2021 | PIAP-DF: Pixel-Interested and Anti Person-Specific Facial Action Unit Detection Net with Discrete Feedback LearningabstractFacial Action Units (AUs) are of great significance in communication. Automatic AU detection can improve the understanding of psychological conditions and emotional status. Recently, several deep learning methods have been proposed to detect AUs automatically. However, several challenges, such as poor extraction of fine-grained and robust local AUs information, model overfitting on person-specific features, as well as the limitation of datasets with wrong labels, remain to be addressed. In this paper, we propose a joint strategy called PIAP-DF to solve these problems, which involves 1) a multi-stage Pixel-Interested learning method with pixel-level attention for each AU; 2) an Anti Person-Specific method aiming to eliminate features associated with any individual as much as possible; 3) a semi-supervised learning method with Discrete Feedback, designed to effectively utilize unlabeled data and mitigate the negative impacts of wrong labels. Experimental results on the two popular AU detection datasets BP4D and DISFA prove that PIAP-DF can be the new state-of-the-art method. Compared with the current best method, PIAP-DF improves the average F1 score by 3.2% on BP4D and by 0.5% on DISFA. All modules of PIAP-DF can be easily removed after training to obtain a lightweight model for practical application. Wangding Zeng, Dafei Zhao, Honggang Zhang 0002 |
ICCV | 4 |
| 2021 | SketchAA: Abstract Representation for Abstract SketchesabstractWhat makes free-hand sketches appealing for humans lies with its capability as a universal tool to depict the visual world. Such flexibility at human ease, however, introduces abstract renderings that pose unique challenges to computer vision models. In this paper, we propose a purpose-made sketch representation for human sketches. The key intuition is that such representation should be abstract at design, so to accommodate the abstract nature of sketches. This is achieved by interpreting sketch abstraction on two levels: appearance and structure. We abstract sketch structure as a pre-defined coarse-to-fine visual block hierarchy, and average visual features within each block to model appearance abstraction. We then discuss three general strategies on how to exploit feature synergy across different levels of this abstraction hierarchy. The superiority of explicitly abstracting sketch representation is empirically validated on a number of sketch analysis tasks, including sketch recognition, fine-grained sketch-based image retrieval, and generative sketch healing. Our simple design not only yields strong results on all said tasks, but also offers intuitive feature granularity control to tailor for various downstream tasks. Code will be made publicly available. Lan Yang 0014, Kaiyue Pang, Honggang Zhang 0002, Yi-Zhe Song |
ICCV | 3 |
| 2021 | Adaptive convolutional neural networks for accelerating magnetic resonance imaging via k-space data interpolation
Tianming Du 0001, Honggang Zhang 0002, Yuemeng Li, Stephen Pickup, Mark Rosen, Hee Kwon Song, Yong Fan 0001 |
Medical Image Anal. | 2 |
| 2021 | Towards Practical Sketch-Based 3D Shape Generation: The Role of Professional SketchesabstractIn this paper, for the first time, we investigate the problem of generating 3D shapes from professional 2D sketches via deep learning. We target sketches done by professional artists, as these sketches are likely to contain more details than the ones produced by novices, and thus the reconstruction from such sketches poses a higher demand on the level of detail in the reconstructed models. This is importantly different to previous work, where the training and testing was conducted on either synthetic sketches or sketches done by novices. Novices sketches often depict shapes that are physically unrealistic, while models trained with synthetic sketches could not cope with the level of abstraction and style found in real sketches. To address this problem, we collected the first large-scale dataset of professional sketches, where each sketch is paired with a reference 3D shape, with a total of 1,500 professional sketches collected across 500 3D shapes. The dataset is available at http://sketchx.ai/downloads/. We introduce two bespoke designs within a deep adversarial network to tackle the imprecision of human sketches and the unique figure/ground ambiguity problem inherent to sketch-based reconstruction. We show that existing 3D shapes generation methods designed for images fail to be naively applied to our problem, and demonstrate the effectiveness of our method both qualitatively and quantitatively. Yonggang Qi, Yulia Gryaditskaya, Honggang Zhang 0002, Yi-Zhe Song |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Learning Recurrent Memory Activation Networks for Visual TrackingabstractFacilitated by deep neural networks, numerous tracking methods have made significant advances. Existing deep trackers mainly utilize independent frames to model the target appearance, while paying less attention to its temporal coherence. In this paper, we propose a recurrent memory activation network (RMAN) to exploit the untapped temporal coherence of the target appearance for visual tracking. We build the RMAN on top of the long short-term memory network (LSTM) with an additional memory activation layer. Specifically, we first use the LSTM to model the temporal changes of the target appearance. Then we selectively activate the memory blocks via the activation layer to produce a temporally coherent representation. The recurrent memory activation layer enriches the target representations from independent frames and reduces the background interference through temporal consistency. The proposed RMAN is fully differentiable and can be optimized end-to-end. To facilitate network training, we propose a temporal coherence loss together with the original binary classification loss. Extensive experimental results on standard benchmarks demonstrate that our method performs favorably against the state-of-the-art approaches. Shi Pu 0002, Yibing Song, Chao Ma 0004, Honggang Zhang 0002, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Rainy Night Scene Understanding With Near Scene Semantic AdaptationabstractDeep networks have been used for semantic segmentation tasks on scenes of outdoor environments with increasing popularity. However, the majority of existing work centers on daytime scenes with favorable illumination and weather conditions, and relies on supervision with pixel-level annotations. This paper seeks to address the problem of semantic segmentation for rainy, night-time scenes without using pixel-level annotations. We introduce a near scene semantic approach that uses images of daytime scenes as a bridge for transferring knowledge from pre-trained segmentation models to rainy night images. Specifically, we first present near scene oriented Representation Adaptation (RA) to reduce the domain shift on the representation level. Next, we adapt the segmentation model from the daytime scenario, under varying weather conditions, to the rainy night scenario by using near scene oriented Segmentation Space Adaptation (SSA). Consequently, this further reduces the impact of the domain shift on the segmentation space level. For evaluation, we created a new dataset containing 7000 distinct daytime-night-time image pairs of near scenes obtained by a webcam, and 5266 daytime-rainy night image pairs collected by a car-mounted camera. In addition, we carefully annotated 226 rainy night images with classes defined in Cityscapes. The experimental results clearly demonstrate the advantage of the proposed algorithm. Shuai Di, Chun-Guang Li, Honggang Zhang 0002, Semir Elezovikj, Chiu C. Tan 0001, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Deep Sketch-Based Modeling: Tips and TricksabstractDeep image-based modeling received lots of attention in recent years, yet the parallel problem of sketch-based modeling has only been briefly studied, often as a potential application. In this work, for the first time, we identify the main differences between sketch and image inputs: (i) style variance, (ii) imprecise perspective, and (iii) sparsity. We discuss why each of these differences can pose a challenge, and even make a certain class of image-based methods inapplicable. We study alternative solutions to address each of the difference. By doing so, we drive out a few important insights: (i) sparsity commonly results in an incorrect prediction of foreground versus background, (ii) diversity of human styles, if not taken into account, can lead to very poor generalization properties, and finally (iii) unless a dedicated sketching interface is used, one can not expect sketches to match a perspective of a fixed viewpoint. Finally, we compare a set of representative deep single-image modeling solutions and show how their performance can be improved to tackle sketch input by taking into consideration the identified critical differences. Yulia Gryaditskaya, Honggang Zhang 0002, Yi-Zhe Song |
3DV | 3 |
| 2020 | Progressive Refinement Network for Occluded Pedestrian Detection
Kaili Zhao, Wen-Sheng Chu, Honggang Zhang 0002, Jun Guo 0002 |
ECCV (23) | 4 |
| 2020 | S3Net: Graph Representational Network For Sketch RecognitionabstractSketches are distinctly different to photos. They are highly abstract and exhibit a severe lack of visual cues. Prior works have therefore explored additional traits unique to sketches to help recognition such as stroke ordering. In this paper, we pioneer in studying the role of structure in sketches, for the task of sketch recognition. In particular, we propose a novel graph representation specifically designed for sketches, which follows the inherent hierarchical relationship (segment-stroke-sketch”) of sketching elements. By conforming to this hierarchy, we also introduce ajoint network that encapsulates both the structural and temporal traits of sketches for sketch recognition, termed S3Net.S3Netemploys a recurrent neural network (RNN) to extract segmentlevel features, followed by a graph convolutional network (GCN) to aggregate them into sketch-level features. The RNN first encodes temporal cues in sketches while its outputs are used as node embedding to construct a hierarchical sketch-graph. The GCN module then takes in this sketchgraph to produce a structure-aware embedding for sketches. Extensive experiments on the QuickDraw dataset, exhibit superior performance over state-of-the-arts, surpassing them by over 4%. Ablative studies further demonstrate the effectiveness of the proposed structural graph for both inter-class, and intra-class feature discrimination. Code is available at: https://github.com/yanglan0225/s3net;. Lan Yang 0014, Aneeshan Sain, Linpeng Li, Yonggang Qi, Honggang Zhang 0002, Yi-Zhe Song |
ICME | 5 |
| 2020 | Sketch-SNet: Deeper Subdivision of Temporal Cues for Sketch RecognitionabstractSketch recognition is essential in sketch-related researches. Different from the natural image, the sparse pixel distribution of sketch discards the visual texture which encourages researchers to explore the temporal information of sketch. Using of million-scale datasets, we explore the invariable structure and specific order of strokes in sketch. Prior works based on Recurrent Neural Network (RNN) output different features with changed stroke orders. In particular, we adopt a novel method by employing a Graph Convolutional Network (GCN) to extract invariable structural feature under any orders of strokes. Compared with traditional comprehension of sketch, we further split the temporal information of sketch into two types of feature, invariable structural feature (ISF) and drawing habits feature (DHF) with the aim of finer feature extraction in temporal information. We propose a two-branch GCN-RNN network, Sketch-SNet, to extract two types of feature respectively. The GCN branch is used to extract the ISF through receiving various shuffled strokes of an input sketch. The RNN branch takes the original order to extract DHF by learning the pattern of strokes' order. Extensive experiments on the Quick-Draw dataset demonstrate that our further subdivision of temporal information improves the performance of sketch recognition which surpasses state-of-the-art by a large margin. Yizhou Tan, Lan Yang 0014, Honggang Zhang 0002 |
ICPR | 3 |
| 2020 | MRP-Net: A Light Multiple Region Perception Neural Network for Multi-label AU DetectionabstractFacial Action Units (AUs) are of great significance in communication. Automatic AU detection can improve the understanding of psychological condition and emotional status. Recently, a number of deep learning methods have been proposed to take charge with problems in automatic AU detection. Several challenges, like unbalanced labels and ignorance of local information, remain to be addressed. In this paper, we propose a fast and light neural network called MRP-Net, which is an end-to-end trainable method for facial AU detection to solve these problems. First, we design a Multiple Region Perception (MRP) module aimed at capturing different locations and sizes of features in the deeper level of the network without facial landmark points. Then, in order to balance the positive and negative samples in the large dataset, a batch balanced method adjusting the weight of every sample in one batch in our loss function is suggested. Experimental results on two popular AU datasets, BP4D and DISFA prove that MRP-Net outperforms state-of-the-art methods. Compared with the best method, not only does MRP-Net have an average F1 score improvement of 2.95% on BP4D and 5.43% on DISFA, and it also decreases the number of network parameters by 54.62% and the number of network FLOPs by 19.6%. Honggang Zhang 0002, Gang Wang 0012 |
ICPR | 3 |
| 2020 | Robust Visual Tracking Via An Imbalance-Elimination MechanismabstractThe competitive performances in visual tracking are achieved mostly by tracking-by-detection based approaches, whose accuracy highly relies on a binary classifier that distinguishes targets from distractors in a set of candidates. However, severe class imbalance, with few positives (e.g., targets) relative to negatives (e.g., backgrounds), leads to degrade accuracy of classification or increase bias of tracking. In this paper, we propose an imbalance-elimination mechanism, which adopts a multi-class paradigm and utilizes a novel candidate generation strategy. Specifically, our multi-class model assigns samples into one positive class and four proposed negative classes, naturally alleviating class imbalance. We define negative classes by introducing proportions of targets in samples, which values explicitly reveal relative scales between targets and backgrounds. Further-more, during candidate generation, we exploit such scale-aware negative patterns to help adjust searching areas of candidates to incorporate larger target proportions, thus more accurate target candidates are obtained and more positive samples are included to ease class imbalance simultaneously. Extensive experiments on standard benchmarks show that our tracker achieves favorable performance against the state-of-the-art approaches, and offers robust discrimination of positive targets and negative patterns. Jin Feng, Kaili Zhao, Anxin Li, Honggang Zhang 0002 |
VCIP | 5 |
| 2020 | Learning Graph Topology Representation with Attention NetworksabstractContextualized neural language models have gained much attention in Information Retrieval (IR) with its ability to achieve better word understanding by capturing contextual structure on sentence level. However, to understand a document better, it is necessary to involve contextual structure from document level. Moreover, some words contributes more information to delivering the meaning of a document. Motivated by this, in this paper, we take the advantages of Graph Convolutional Networks (GCN) and Graph Attention Networks (GAN) to model global word-relation structure of a document with attention mechanism to improve context-aware document ranking. We propose to build a graph for a document to model the global contextual structure. The nodes and edges of the graph are constructed from contextual embeddings. We first apply graph convolution on the graph and then use attention networks to explore the influence of more informative words to obtain a new representation. This representation covers both local contextual and global structure information. The experimental results show that our method outperforms the state-of-the-art contextual language models, which demonstrate that incorporating contextual structure is useful for improving document ranking. Jiayue Zhang, Weiran Xu, Jun Guo 0002, Honggang Zhang 0002 |
VCIP | 5 |
| 2020 | Robust visual tracking by embedding combination and weighted-gradient optimization
Jin Feng, Shi Pu 0002, Kaili Zhao, Honggang Zhang 0002 |
Pattern Recognit. | 5 |
| 2019 | Generalising Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval (FG-SBIR) addresses matching specific photo instance using free-hand sketch as a query modality. Existing models aim to learn an embedding space in which sketch and photo can be directly compared. While successful, they require instance-level pairing within each coarse-grained category as annotated training data. Since the learned embedding space is domain-specific, these models do not generalise well across categories. This limits the practical applicability of FG-SBIR. In this paper, we identify cross-category generalisation for FG-SBIR as a domain generalisation problem, and propose the first solution. Our key contribution is a novel unsupervised learning approach to model a universal manifold of prototypical visual sketch traits. This manifold can then be used to paramaterise the learning of a sketch/photo representation. Model adaptation to novel categories then becomes automatic via embedding the novel sketch in the manifold and updating the representation and retrieval function accordingly. Experiments on the two largest FG-SBIR datasets, Sketchy and QMUL-Shoe-V2, demonstrate the efficacy of our approach in enabling cross-category generalisation of FG-SBIR. Kaiyue Pang, Ke Li 0004, Yongxin Yang, Honggang Zhang 0002, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 4 |
| 2019 | Self-Supervised Convolutional Subspace Clustering NetworkabstractSubspace clustering methods based on data self-expression have become very popular for learning from data that lie in a union of low-dimensional linear subspaces. However, the applicability of subspace clustering has been limited because practical visual data in raw form do not necessarily lie in such linear subspaces. On the other hand, while Convolutional Neural Network (ConvNet) has been demonstrated to be a powerful tool for extracting discriminative features from visual data, training such a ConvNet usually requires a large amount of labeled data, which are unavailable in subspace clustering applications. To achieve simultaneous feature learning and subspace clustering, we propose an end-to-end trainable framework, called Self-Supervised Convolutional Subspace Clustering Network (S$^2$ConvSCN), that combines a ConvNet module (for feature learning), a self-expression module (for subspace clustering) and a spectral clustering module (for self-supervision) into a joint optimization framework. Particularly, we introduce a dual self-supervision that exploits the output of spectral clustering to supervise the training of the feature learning module (via a classification loss) and the self-expression module (via a spectral clustering loss). Our experiments on four benchmark datasets show the effectiveness of the dual self-supervision and demonstrate superior performance of our proposed approach. Chun-Guang Li, Chong You, Xianbiao Qi, Honggang Zhang 0002, Jun Guo 0002, Zhouchen Lin |
CVPR | 5 |
| 2019 | Morphology Reconstruction of Obstructed Coronary Artery in Angiographic ImagesabstractIn cardiac arterial interventional therapy, coronary angiograms provides key information to physicians for treatment strategy selection. However, it is hard to extract the information about arteries with CTO lesion in absence of contrast opacification and visualization of the CTO coronary arteries in angiograms. In this paper, we present an algorithm, which can predict the extension direction of artery with CTO and reconstruct the morphology of blocked artery in coronary angiograms. First, our algorithm segment the CTO artery angiogram and extract the artery skeleton. Second, an iterative approach is preformed to reconstruct the skeleton of blocked artery. Finally, our algorithm generates the simulated postoperative angiogram according to the skeleton image. The results demonstrate that automatic morphology reconstructing of CTO coronary artery with high accuracy and reliability is feasible. Tianming Du 0001, Xiaotong Shi, Ruijia Wu, Honggang Zhang 0002, Jin Feng, Fuzhuo Sun |
VCIP | 4 |
| 2019 | Enhanced Initialization with Multi-Stage Learning for Robust Visual TrackingabstractVisual tracking is the task of estimating the trajectory of a target given its first location in a video. Existing tracking-by-detection approaches build trackers on binary classifiers. However, these approaches are sometimes not discriminative enough to distinguish the target from the background, which impedes the performance under challenging conditions. In this paper, we demonstrate that conventional initializing strategies make the model insufficiently trained. To make full use of the information contained in the first frame, we propose an enhanced scheme of model initialization. In the scheme, several kinds of negative samples are heuristically defined and a multi-stage learning mechanism is adopted to make the model initialization more efficient and stable. Validation experiment demonstrates the effectiveness of the mechanism. Experiments on benchmark datasets show that the proposed method improves the results of the baseline work and achieves state-of-the-art performance. Jin Feng, Shi Pu 0002, Kaili Zhao, Honggang Zhang 0002, Tianming Du 0001 |
VCIP | 4 |
| 2019 | Toward Deep Universal Sketch Perceptual GrouperabstractHuman free-hand sketches provide the useful data for studying human perceptual grouping, where the grouping principles such as the Gestalt laws of grouping are naturally in play during both the perception and sketching stages. In this paper, we make the first attempt to develop a universal sketch perceptual grouper. That is, a grouper that can be applied to sketches of any category created with any drawing style and ability, to group constituent strokes/segments into semantically meaningful object parts. The first obstacle to achieving this goal is the lack of large-scale datasets with grouping annotation. To overcome this, we contribute the largest sketch perceptual grouping dataset to date, consisting of 20 000 unique sketches evenly distributed over 25 object categories. Furthermore, we propose a novel deep perceptual grouping model learned with both generative and discriminative losses. The generative loss improves the generalization ability of the model, while the discriminative loss guarantees both local and global grouping consistency. Extensive experiments demonstrate that the proposed grouper significantly outperforms the state-of-the-art competitors. In addition, we show that our grouper is useful for a number of sketch analysis tasks, including sketch semantic segmentation, synthesis, and fine-grained sketch-based image retrieval. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 6 |
| 2018 | Dual Attention Matching Network for Context-Aware Feature Sequence Based Person Re-IdentificationabstractTypical person re-identification (ReID) methods usually describe each pedestrian with a single feature vector and match them in a task-specific metric space. However, the methods based on a single feature vector are not sufficient enough to overcome visual ambiguity, which frequently occurs in real scenario. In this paper, we propose a novel end-to-end trainable framework, called Dual ATtention Matching network (DuATM), to learn context-aware feature sequences and perform attentive sequence comparison simultaneously. The core component of our DuATM framework is a dual attention mechanism, in which both intrasequence and inter-sequence attention strategies are used for feature refinement and feature-pair alignment, respectively. Thus, detailed visual cues contained in the intermediate feature sequences can be automatically exploited and properly compared. We train the proposed DuATM network as a siamese network via a triplet loss assisted with a decorrelation loss and a cross-entropy loss. We conduct extensive experiments on both image and video based ReID benchmark datasets. Experimental results demonstrate the significant advantages of our approach compared to the state-of-the-art methods. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jason Kuen, Xiangfei Kong, Alex Chichung Kot, Gang Wang 0012 |
CVPR | 2 |
| 2018 | Universal Sketch Perceptual Grouping
Ke Li 0004, Kaiyue Pang, Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Honggang Zhang 0002 |
ECCV (8) | 7 |
| 2018 | Deep Attentive Tracking via Reciprocative LearningabstractVisual attention, derived from cognitive neuroscience, facilitates human perception on the most pertinent subset of the sensory data. Recently, significant efforts have been made to exploit attention schemes to advance computer vision systems. For visual tracking, it is often challenging to track target objects undergoing large appearance changes. Attention maps facilitate visual tracking by selectively paying attention to temporal robust features. Existing tracking-by-detection approaches mainly use additional attention modules to generate feature weights as the classifiers are not equipped with such mechanisms. In this paper, we propose a reciprocative learning algorithm to exploit visual attention for training deep classifiers. The proposed algorithm consists of feed-forward and backward operations to generate attention maps, which serve as regularization terms coupled with the original classification loss function for training. The deep classifier learns to attend to the regions of target objects robust to appearance changes. Extensive experiments on large-scale benchmark datasets show that the proposed attentive tracking method performs favorably against the state-of-the-art approaches. Shi Pu 0002, Yibing Song, Chao Ma 0004, Honggang Zhang 0002, Ming-Hsuan Yang 0001 |
NeurIPS | 4 |
| 2018 | Cross-Domain Traffic Scene Understanding: A Dense Correspondence-Based Transfer Learning ApproachabstractUnderstanding traffic scene images taken from vehicle mounted cameras is important for high-level tasks, such as advanced driver assistance systems and autonomous driving. It is a challenging problem due to large variations under different weather or illumination conditions. In this paper, we tackle the problem of traffic scene understanding from a cross-domain perspective. We attempt to understand the traffic scene from images taken from the same location but under different weather or illumination conditions (e.g., understanding the same traffic scene from images on a rainy night with the help of images taken on a sunny day). To this end, we propose a dense correspondence-based transfer learning (DCTL) approach, which consists of three main steps: 1) extracting deep representations of traffic scene images via a fine-tuned convolutional neural network; 2) constructing compact and effective representations via cross-domain metric learning and subspace alignment for cross-domain retrieval; and 3) transferring the annotations from the retrieved best matching image to the test image based on cross-domain dense correspondences and a probabilistic Markov random field. To verify the effectiveness of our DCTL approach, we conduct extensive experiments on a challenging data set, which contains 1828 images from six weather or illumination conditions. Shuai Di, Honggang Zhang 0002, Chun-Guang Li, Xue Mei, Danil V. Prokhorov, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Spatial Pyramid-Based Statistical Features for Person Re-Identification: A Comprehensive EvaluationabstractPerson re-identification (Re-Id) across nonoverlapping camera views is one of challenging problems in surveillance video analysis. The difficulties in person Re-Id mainly come from the large appearance variations caused by camera view angle, human pose, illumination, and occlusion. Recently, extensive efforts have been cast into addressing this problem by developing invariant features or discriminative distance metrics. However, there is still a lack of systematic evaluations on the pipeline for feature extraction and combination. In this paper, we propose a spatial pyramid-based statistical feature extraction framework as a unified pipeline of feature extraction and combination for person Re-Id, and systematically evaluate the configuration details in feature extraction and the fusion strategies in feature combination. Extensive experiments on benchmark datasets demonstrate the critical components in feature extraction. Moreover, by combining multiple features, our proposed approach can yield state-of-the-art performance. It should be mentioned that our approach achieves rank 1 matching rate of 45.8% on dataset VIPeR and 61.5% on dataset CUHK01, respectively. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li, Jun Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2017 | Synergistic Instance-Level Subspace Alignment for Fine-Grained Sketch-Based Image RetrievalabstractWe study the problem of fine-grained sketch-based image retrieval. By performing instance-level (rather than category-level) retrieval, it embodies a timely and practical application, particularly with the ubiquitous availability of touchscreens. Three factors contribute to the challenging nature of the problem: 1) free-hand sketches are inherently abstract and iconic, making visual comparisons with photos difficult; 2) sketches and photos are in two different visual domains, i.e., black and white lines versus color pixels; and 3) fine-grained distinctions are especially challenging when executed across domain and abstraction-level. To address these challenges, we propose to bridge the image-sketch gap both at the high level via parts and attributes, as well as at the low level via introducing a new domain alignment method. More specifically, first, we contribute a data set with 304 photos and 912 sketches, where each sketch and image is annotated with its semantic parts and associated part-level attributes. With the help of this data set, second, we investigate how strongly supervised deformable part-based models can be learned that subsequently enable automatic detection of part-level attributes, and provide pose-aligned sketch-image comparisons. To reduce the sketch-image gap when comparing low-level features, third, we also propose a novel method for instance-level domain-alignment that exploits both subspace and instance-level cues to better align the domains. Finally, fourth, these are combined in a matching framework integrating aligned low-level features, mid-level geometric structure, and high-level semantic attributes. Extensive experiments conducted on our new data set demonstrate effectiveness of the proposed method. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Timothy M. Hospedales, Tao Xiang 0002, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 6 |
| 2016 | Deep Region and Multi-label Learning for Facial Action Unit DetectionabstractRegion learning (RL) and multi-label learning (ML) have recently attracted increasing attentions in the field of facial Action Unit (AU) detection. Knowing that AUs are active on sparse facial regions, RL aims to identify these regions for a better specificity. On the other hand, a strong statistical evidence of AU correlations suggests that ML is a natural way to model the detection task. In this paper, we propose Deep Region and Multi-label Learning (DRML), a unified deep network that simultaneously addresses these two problems. One crucial aspect in DRML is a novel region layer that uses feed-forward functions to induce important facial regions, forcing the learned weights to capture structural information of the face. Our region layer serves as an alternative design between locally connected layers (i.e., confined kernels to individual pixels) and conventional convolution layers (i.e., shared kernels across an entire image). Unlike previous studies that solve RL and ML alternately, DRML by construction addresses both problems, allowing the two seemingly irrelevant problems to interact more directly. The complete network is end-to-end trainable, and automatically learns representations robust to variations inherent within a local region. Experiments on BP4D and DISFA benchmarks show that DRML performs the highest average F1-score and AUC within and across datasets in comparison with alternative methods. Kaili Zhao, Wen-Sheng Chu, Honggang Zhang 0002 |
CVPR | 3 |
| 2016 | Sketch-based image retrieval via Siamese convolutional neural networkabstractSketch-based image retrieval (SBIR) is a challenging task due to the ambiguity inherent in sketches when compared with photos. In this paper, we propose a novel convolutional neural network based on Siamese network for SBIR. The main idea is to pull output feature vectors closer for input sketch-image pairs that are labeled as similar, and push them away if irrelevant. This is achieved by jointly tuning two convolutional neural networks which linked by one loss function. Experimental results on Flickr15K demonstrate that the proposed method offers a better performance when compared with several state-of-the-art approaches. Yonggang Qi, Yi-Zhe Song, Honggang Zhang 0002, Jun Liu 0014 |
ICIP | 3 |
| 2016 | Structure and appearance preserving network flow for multi-object trackingabstractTracking-by-detection with temporal smoothness has recently attracted increasing attentions in the field of multi-object tracking. Occlusions and clutter are two key problems. To address these problems, this paper proposes a new structure and appearance preserving network flow (SAPNF) with tracking-by-detection, introducing spatial structural configuration and appearance overlapping constraint from frame to frame. One crucial aspect in SAPNF is to consider structure information and appearance smoothness simultaneously which benefits from each other. Unlike previous studies that only learn spatial information or appearance smoothness, a unified min-cost flow with the proposed new structure and appearance induces to track multi-object in crowded and cluttered scenes. Experiments on PETS and TUD benchmarks show that SAPNF performs the comparative results in comparison with alternative methods. Shi Pu 0002, Honggang Zhang 0002, Kaili Zhao |
ICPR | 2 |
| 2016 | Cross-modal face matching: Tackling visual abstraction using fine-grained attributesabstractDespite great strides made in facial verification, it remains challenging to match facial images across different modalities. This is mainly due to the cross-modal gap induced by feature heterogeneity. Much prior work had focused on bridging the feature gap, resulting in near-perfect matching accuracies for viewed sketches. Nonetheless, studies on matching unviewed (forensic) sketches and caricatures, a much harder problem due to the additional cross-modal gap introduced by visual abstraction, had only just commenced in recent years. In this paper, we focus on matching facial caricatures with photos by directly addressing the visual abstraction problem. We show that by synergizing a taxonomy of fine-grained visual attributes with part-aware low-level feature extraction, the visual abstraction gap can be effectively traversed, resulting in improved overall cross-modal matching accuracy. More specifically, (i) we propose a simple yet effective geometry-based attribute classifier to detect fine-grained attributes at part-level, and (ii) we demonstrate how meaningful facial regions can be reliably detected to enable localized feature extraction and attribute detection, and (iii) we show a common embedding can be learned using Canonical Correlation Analysis (CCA) that combines part-based low-level features and fine-grained visual attributes. We demonstrate the superiority of the proposed cross-modal strategy by evaluating on two recent photo-caricature datasets. Yichuan Hu, Ke Li 0004, Honggang Zhang 0002 |
VCIP | 3 |
| 2016 | Low-rank and structured sparse subspace clusteringabstractHigh dimensional data often lie approximately in low dimensional subspaces corresponding to multiple classes or categories. Segmenting the high dimensional data into their corresponding low dimensional subspaces is referred as subspace clustering. State of the art methods solve this problem in two steps. First, an affinity matrix is built from data based on self-expressiveness model, in which each data point is expressed as a linear combination of other data points. Second, the segmentation is obtained by spectral clustering. However, solving two dependent steps separately is still suboptimal. In this paper, we propose a joint affinity learning and spectral clustering approach for low-rank representation based subspace clustering, termed Low-Rank and Structured Sparse Subspace Clustering (LRS3C), where a subspace structured norm that depends on subspace clustering result is introduced into the objective of low-rank representation problem. We solve it efficiently via a combination of Linearized Alternation Direction Method (LADM) with spectral clustering. Experiments on Hopkins 155 motion segmentation database and Extended Yale B data set demonstrated the effectiveness of our method. Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002 |
VCIP | 3 |
| 2016 | Fine-grained sketch-based image retrieval: The role of part-aware attributesabstractWe study the problem of fine-grained sketch-based image retrieval. By performing instance-level (rather than category-level) retrieval, it embodies a timely and practical application, particularly with the ubiquitous availability of touchscreens. Three factors contribute to the challenging nature of the problem: (i) free-hand sketches are inherently abstract and iconic, making visual comparisons with photos more difficult, (ii) sketches and photos are in two different visual domains, i.e. black and white lines vs. color pixels, and (iii) fine-grained distinctions are especially challenging when executed across domain and abstraction-level. To address this, we propose to detect visual attributes at part-level, in order to build a new representation that not only captures fine-grained characteristics but also traverses across visual domains. More specifically, (i) we propose a dataset with 304 photos and 912 sketches, where each sketch and photo is annotated with its semantic parts and associated part-level attributes, and with the help of this dataset, we investigate (ii) how strongly-supervised deformable part-based models can be learned that subsequently enable automatic detection of part-level attributes, and (iii) a novel matching framework that synergistically integrates low-level features, mid-level geometric structure and high-level semantic attributes to boost retrieval performance. Extensive experiments conducted on our new dataset demonstrate value of the proposed method. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Timothy M. Hospedales, Honggang Zhang 0002, Yichuan Hu |
WACV | 5 |
| 2016 | Joint Patch and Multi-label Learning for Facial Action Unit and Holistic Expression RecognitionabstractMost action unit (AU) detection methods use one-versus-all classifiers without considering dependences between features or AUs. In this paper, we introduce a joint patch and multi-label learning (JPML) framework that models the structured joint dependence behind features, AUs, and their interplay. In particular, JPML leverages group sparsity to identify important facial patches, and learns a multi-label classifier constrained by the likelihood of co-occurring AUs. To describe such likelihood, we derive two AU relations, positive correlation and negative competition, by statistically analyzing more than 350,000 video frames annotated with multiple AUs. To the best of our knowledge, this is the first work that jointly addresses patch learning and multi-label learning for AU detection. In addition, we show that JPML can be extended to recognize holistic expressions by learning common and specific patches, which afford a more compact representation than the standard expression recognition methods. We evaluate JPML on three benchmark datasets CK+, BP4D, and GFT, using within-and cross-dataset scenarios. In four of five experiments, JPML achieved the highest averaged F1 scores in comparison with baseline and alternative methods that use either patch learning or multi-label learning alone. Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2015 | VecLP: A Realtime Video Recommendation System for Live TV ProgramsabstractWe propose VecLP, a novel Internet Video recommendation system working for Live TV Programs in this paper. Given little information on the live TV programs, our proposed VecLP system can effectively collect necessary information on both the programs and the subscribers as well as a large volume of related online videos, and then recommend the relevant Internet videos to the subscribers. For that, the key frames are firstly detected from the live TV programs, and then visual and textual features are extracted from these frames to enhance the understanding of the TV broadcasts. Furthermore, by utilizing the subscribers' profiles and their social relationships, a user preference model is constructed, which greatly improves the diversity of the recommendations in our system. The subscriber's browsing history is also recorded and used to make a further personalized recommendation. This work also illustrates how our proposed VecLP system makes it happen. Finally, we dispose some sort of new recommendation strategies in use at the system to meet special needs from diverse live TV programs and throw light upon how to fuse these strategies. Sheng Gao 0001, Honggang Zhang 0002, Jianxin Liao, Jun Guo 0002 |
AAAI | 3 |
| 2015 | Making better use of edges via perceptual groupingabstractWe propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of candidate edge pairs. In particular, RankSVM is employed for the first time to combine multiple Gestalt principles as cue for edge grouping. Afterwards, an edge grouping based object proposal measure is introduced that yields proposals comparable to state-of-the-art alternatives. We further show how human-like sketches can be generated from edge groupings and consequently used to deliver state-of-the-art sketch-based image retrieval performance. Last but not least, we tackle the problem of freehand human sketch segmentation by utilizing the proposed grouper to cluster strokes into semantic object parts. Yonggang Qi, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Timothy M. Hospedales, Yi Li 0004, Jun Guo 0002 |
CVPR | 4 |
| 2015 | Joint patch and multi-label learning for facial action unit detectionabstractThe face is one of the most powerful channel of nonverbal communication. The most commonly used taxonomy to describe facial behaviour is the Facial Action Coding System (FACS). FACS segments the visible effects of facial muscle activation into 30+ action units (AUs). AUs, which may occur alone and in thousands of combinations, can describe nearly all-possible facial expressions. Most existing methods for automatic AU detection treat the problem using one-vs-all classifiers and fail to exploit dependencies among AU and facial features. We introduce joint-patch and multi-label learning (JPML) to address these issues. JPML leverages group sparsity by selecting a sparse subset of facial patches while learning a multi-label classifier. In four of five comparisons on three diverse datasets, CK+, GFT, and BP4D, JPML produced the highest average F1 scores in comparison with state-of-the art. Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002 |
CVPR | 5 |
| 2015 | Learning Semi-Supervised Representation Towards a Unified Optimization Framework for Semi-Supervised LearningabstractState of the art approaches for Semi-Supervised Learning (SSL) usually follow a two-stage framework -- constructing an affinity matrix from the data and then propagating the partial labels on this affinity matrix to infer those unknown labels. While such a two-stage framework has been successful in many applications, solving two subproblems separately only once is still suboptimal because it does not fully exploit the correlation between the affinity and the labels. In this paper, we formulate the two stages of SSL into a unified optimization framework, which learns both the affinity matrix and the unknown labels simultaneously. In the unified framework, both the given labels and the estimated labels are used to learn the affinity matrix and to infer the unknown labels. We solve the unified optimization problem via an alternating direction method of multipliers combined with label propagation. Extensive experiments on a synthetic data set and several benchmark data sets demonstrate the effectiveness of our approach. Chun-Guang Li, Zhouchen Lin, Honggang Zhang 0002, Jun Guo 0002 |
ICCV | 3 |
| 2015 | Regularization in metric learning for person re-identificationabstractMetric learning plays a critical role in person re-identification problem. Unfortunately, due to the small size of training data, the metric learning used in this scenario suffers from over-fitting which leads to degenerated performance. In this paper, we investigate the effect of regularization in metric learning for person re-identification. Concretely we formulate the distance function from three perspectives and hence present four different regularized metric learning methods. Experiments on two popular benchmark data sets VIPeR and CUHK01 validate the effectiveness of our proposed regularization approaches. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
ICIP | 2 |
| 2015 | Improving tag matrix completion for image annotation and retrievalabstractImage annotation is a fundamental and challenging task in the field of semantic image retrieval. In this paper, we deal with image annotation via matrix completion. Concretely, we formulate the problem of annotating the tags of an image into a constrained optimization problem, in which the constraint is to keep the consistency with the given initial labels and the objective is to minimize the discrepancy between the correlation in visual content and the correlation in semantic tags. We solve the optimization problem with the linearized alternating direction method. Experimental results on benchmark data demonstrate the effectiveness of our proposals. Zhen Qin 0001, Chun-Guang Li, Honggang Zhang 0002, Jun Guo 0002 |
VCIP | 3 |
| 2015 | Im2Sketch: Sketch generation by unconflicted perceptual grouping
Yonggang Qi, Jun Guo 0002, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Zheng-Hua Tan |
Neurocomputing | 5 |
| 2015 | Multi-label learning with prior knowledge for facial expression analysis
Kaili Zhao, Honggang Zhang 0002, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
Neurocomputing | 2 |
| 2015 | Variational Bayesian Matrix Factorization for Bounded Support DataabstractA novel Bayesian matrix factorization method for bounded support data is presented. Each entry in the observation matrix is assumed to be beta distributed. As the beta distribution has two parameters, two parameter matrices can be obtained, which matrices contain only nonnegative values. In order to provide low-rank matrix factorization, the nonnegative matrix factorization (NMF) technique is applied. Furthermore, each entry in the factorized matrices, i.e., the basis and excitation matrices, is assigned with gamma prior. Therefore, we name this method as beta-gamma NMF (BG-NMF). Due to the integral expression of the gamma function, estimation of the posterior distribution in the BG-NMF model can not be presented by an analytically tractable solution. With the variational inference framework and the relative convexity property of the log-inverse-beta function, we propose a new lower-bound to approximate the objective function. With this new lower-bound, we derive an analytically tractable solution to approximately calculate the posterior distributions. Each of the approximated posterior distributions is also gamma distributed, which retains the conjugacy of the Bayesian estimation. In addition, a sparse BG-NMF can be obtained by including a sparseness constraint to the gamma prior. Evaluations with synthetic data and real life data demonstrate the good performance of the proposed method. Zhanyu Ma, Andrew E. Teschendorff, Arne Leijon, Yuanyuan Qiao 0002, Honggang Zhang 0002, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2014 | Nonlinear estimation of missing ΔLSF parameters by a mixture of Dirichlet distributionsabstractIn packet networks, a reliable scheme to handle packet loss during speech transmission is of great importance. As a common representation of the linear predictive coding (LPC) model, the line spectral frequency (LSF) parameters are widely used in speech quantization and transmission. In this paper, we propose a novel scheme to estimate the missing values occurring during LPC model transmission. In order to exploit the boundary and ordering properties of the LSF parameters, we utilize the ΔLSF representation and apply the Dirichlet mixture model (DMM) to capture the correlations among the elements in the ΔLSF vector. With the conditional distribution of the missing part given the received part, an optimal nonlinear minimum mean square error estimator for the missing values is proposed. Compared to the previously presented Gaussian mixture model based method, the proposed DMM based nonlinear estimator shows a convincing improvement. Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002, Honggang Zhang 0002 |
ICASSP | 4 |
| 2014 | An adaptive group lasso based multi-label regression approach for facial expression analysisabstractIn the realm of facial expression analysis, numerous attempts have been made to link each facial picture to one affective category. Nevertheless, in our daily life, few of the facial expressions are exactly one of the predefined affective states. Therefore, to analyze the facial expressions more effectively, this paper proposes an Adaptive Group Lasso based Multilabel Regression approach, which depicts each facial expression with multiple continuous values of predefined affective states. Adaptive Group Lasso is adopted to depict the relationship between different labels which different facial expressions share some same affective facial areas (patches). Moreover, to solve the multi-label regression problem, a convex optimization formulation is presented, which would guarantee a global optimal solution. The experiment results based on JAFFE dataset have verified the superior performance of our approach. Kaili Zhao, Honggang Zhang 0002, Jun Guo 0002 |
ICIP | 2 |
| 2014 | Person re-identification via region-of-interest based featuresabstractPerson re-identification is still a challenging task due to large visual appearance variations caused by illumination, background, viewpoints and poses in multi-camera surveillance. To address these challenges, many methods have been proposed. In this paper, we present an efficient method, called Region-of-Interest based Features (ROIF), via combining textural and chromatic features. It consists of two main phases - region-of-interest exploration from image and features extraction from ROI. Experimental results on the database VIPeR show that our method can yield promising accuracy with a quite cheap time cost. Jianlou Si, Honggang Zhang 0002, Chun-Guang Li |
VCIP | 2 |
| 2014 | A multimedia information fusion framework for web image categorization
Wenting Lu, Lei Li 0001, Tao Li 0001, Honggang Zhang 0002, Jun Guo 0002 |
Multim. Tools Appl. | 5 |
| 2013 | A Maximum K-Min Approach for Classification
Mingzhi Dong, Weihong Deng, Jun Guo 0002, Honggang Zhang 0002 |
AAAI | 6 |
| 2013 | Ordered histogram of shapemes: An ordered bag-of-features based shape descriptor for efficient shape matchingabstractIn this paper, we enhance the Shape Context-based descriptor, shapemes, by introducing an ordered bag-of-features model and dynamic programming. The proposed descriptor consists of a series of sub-histograms of shapemes, each of which represents a subset of sampled points. The division of the sampled points is based on their sequential positions on the contour of the shape, so the representation has intrinsic order and is therefore named ordered histogram of shapemes. Then dynamic programming is utilized for descriptor matching. The framework is effective and efficient owing to the following properties: 1) points division approach together with dynamic programming for invariance under the change of starting point, 2) Earth Mover's Distance for discriminative power, and 3) pre-caculated shapemes dissimilarity matrix for fast descriptor distance calculation. Experiments on standard shape database and real world application scenario demonstrate the effectiveness and efficiency of the descriptor and the matching framework. We make our code and experimental data publicly available for future reference. Lunshao Chai, Zhen Qin 0001, Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002 |
ICIP | 4 |
| 2013 | Representative reference-set and betweenness centrality for scene image categorizationabstractReference-based image classification approach introduces a reference-set for both image representation and dictionary learning. It significantly reduces the dimensionality of represented images and shows outstanding performance even with randomly selected reference images and simple distance measure. In this paper, we improve upon existing work with two major contributions. First, we show that a more representative reference-set contributes to better classification accuracy. To this end, we carefully adapt the K-means clustering algorithm in the feature space to select a distinguished reference-set. Second, in the image classification process, we propose to represent each image by measuring its betweenness centrality in a social network composed of the representative reference-set in each class, leading to a more coherent distance measure that considers the overall connectivity between the probe image and the reference-set. Extensive experiment results demonstrate that our proposed scheme achieves better performance than existing methods. Qun Li 0002, Zhen Qin 0001, Lunshao Chai, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
ICIP | 4 |
| 2013 | Sketching by perceptual groupingabstractSketch is used for rendering the visual world since prehistoric times, and has become ubiquitous nowadays with the increasing availability of touchscreens on portable devices. However, how to automatically map images to sketches, a problem that has profound implications on applications such as sketch-based image retrieval, still remains open. In this paper, we propose a novel method that draws a sketch automatically from a single natural image. Sketch extraction is posed within an unified contour grouping framework, where perceptual grouping is first used to form contour segment groups, followed by a group-based contour simplification method that generate the final sketches. In our experiment, for the first time we pose sketch evaluation as a sketch-based object recognition problem and the results validate the effectiveness of our system over the state-of-the-arts alternatives. Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song |
ICIP | 4 |
| 2013 | Perceptual grouping via untangling Gestalt principlesabstractGestalt principles, a set of conjoining rules derived from human visual studies, have been known to play an important role in computer vision. Many applications such as image segmentation, contour grouping and scene understanding often rely on such rules to work. However, the problem of Gestalt confliction, i.e., the relative importance of each rule compared with another, remains unsolved. In this paper, we investigate the problem of perceptual grouping by quantifying the confliction among three commonly used rules: similarity, continuity and proximity. More specifically, we propose to quantify the importance of Gestalt rules by solving a learning to rank problem, and formulate a multi-label graph-cuts algorithm to group image primitives while taking into account the learned Gestalt confliction. Our experiment results confirm the existence of Gestalt confliction in perceptual grouping and demonstrate an improved performance when such a confliction is accounted for via the proposed grouping algorithm. Finally, a novel cross domain image classification method is proposed by exploiting perceptual grouping as representation. Yonggang Qi, Jun Guo 0002, Yi Li 0004, Honggang Zhang 0002, Tao Xiang 0002, Yi-Zhe Song, Zheng-Hua Tan |
VCIP | 4 |
| 2013 | A multi-label classification approach for Facial Expression RecognitionabstractFacial Expression Recognition (FER) techniques have already been adopted in numerous multimedia systems. Plenty of previous research assumes that each facial picture should be linked to only one of the predefined affective labels. Nevertheless, in practical applications, few of the expressions are exactly one of the predefined affective states. Therefore, to depict the facial expressions more accurately, this paper proposes a multi-label classification approach for FER and each facial expression would be labeled with one or multiple affective states. Meanwhile, by modeling the relationship between labels via Group Lasso regularization term, a maximum margin multi-label classifier is presented and the convex optimization formulation guarantees a global optimal solution. To evaluate the performance of our classifier, the JAFFE dataset is extended into a multi-label facial expression dataset by setting threshold to its continuous labels marked in the original dataset and the labeling results have shown that multiple labels can output a far more accurate description of facial expression. At the same time, the classification results have verified the superior performance of our algorithm. Kaili Zhao, Honggang Zhang 0002, Mingzhi Dong, Jun Guo 0002, Yonggang Qi, Yi-Zhe Song |
VCIP | 2 |
| 2013 | Text extraction from natural scene image: A survey
Honggang Zhang 0002, Kaili Zhao, Yi-Zhe Song, Jun Guo 0002 |
Neurocomputing | 1 |
| 2013 | Reference-Based Scheme Combined With K-SVD for Scene Image CategorizationabstractA reference-based algorithm for scene image categorization is presented in this letter. In addition to using a reference-set for images representation, we also associate the reference-set with training data in sparse codes during the dictionary learning process. The reference-set is combined with the reconstruction error to form a unified objective function. The optimal solution is efficiently obtained using the K-SVD algorithm. After dictionaries are constructed, Locality-constrained Linear Coding (LLC) features of images are extracted. Then, we represent each image feature vector using the similarities between the image and the reference-set, leading to a significant reduction of the dimensionality in the feature space. Experimental results demonstrate that our method achieves outstanding performance. Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
IEEE Signal Process. Lett. | 2 |
| 2013 | Web Multimedia Object Classification Using Cross-Domain Correlation KnowledgeabstractGiven a collection of web images with the corresponding textual descriptions, in this paper, we propose a novel cross-domain learning method to classify these web multimedia objects by transferring the correlation knowledge among different information sources. Here, the knowledge is extracted from unlabeled objects through unsupervised learning and applied to perform supervised classification tasks. To mine more meaningful correlation knowledge, instead of using commonly used visual words in the traditional bag-of-visual-words (BoW) model, we discover higher level visual components (words and phrases) to incorporate the spatial and semantic information into our image representation model, i.e., bag-of-visual-phrases (BoP). By combining the enriched visual components with the textual words, we calculate the frequently co-occurring pairs among them to construct a cross-domain correlated graph in which the correlation knowledge is mined. After that, we investigate two different strategies to apply such knowledge to enrich the feature space where the supervised classification is performed. By transferring such knowledge, our cross-domain transfer learning method can not only handle large scale web multimedia objects, but also deal with the situation that the textual descriptions of a small portion of web images are missing. Empirical experiments on two different datasets of web multimedia objects are conducted to demonstrate the efficacy and effectiveness of our proposed cross-domain transfer learning method. Wenting Lu, Tao Li 0001, Weidong Guo, Honggang Zhang 0002, Jun Guo 0002 |
IEEE Trans. Multim. | 5 |
| 2012 | Re-ranking using compression-based distance measure for Content-based Commercial Product Image RetrievalabstractWith the prevalence of E-Commerce sites such as eBay, Content-based Commercial Product Image Retrieval (CBCPIR) has become an emerging application-oriented field of Content-based Image Retrieval (CBIR). Though a number of traditional CBIR techniques and evaluation criterions have been applied directly or with minor modifications, they tend to neglect one critical factor that greatly affects user experience: users usually care about the exact ranks of the results, especially few top ones, which should share very high similarity with the query image. In this work, we propose a novel two-stage retrieval framework that uses a compression-based re-ranking method and a new subjective retrieval evaluation criterion to address such a problem. More specifically, we extend the state-of-art texture descriptor Campana-Keogh (CK) method from data mining in several aspects and validate the superiority of our framework via extensive experiments and real-world user feedback. We also make our code and CBCPIR dataset publicly available. The number of images of the latter is much larger than current freely accessible ones and better represents real-world commercial product images. Lunshao Chai, Zhen Qin 0001, Honggang Zhang 0002, Jun Guo 0002, Christian R. Shelton |
ICIP | 3 |
| 2012 | Codebook optimization using word activation forces for scene categorizationabstractVisual codebook based quantization of robust appearance descriptors extracted from local image patches is an effective means of capturing image statistics for texture analysis and natural scene classification. In this paper, based on the newly proposed statistics of word activation forces (WAFs), we optimize the codebook. Currently, codebooks are typically created from a set of training images using a clustering algorithm. However, these codebooks are often functionally limited due to redundancy. We show that WAFs can remove the redundancy efficiently. In the experiment, the proposed method achieved the state-of-the-art performance on the Caltech-101, fifteen natural scene categories and VOC2007 databases. The optimization method also offers insights into the success of several recently proposed images classification approaches, including vector quantization (VQ) coding in the Spatial Pyramid Matching (SPM), sparse coding SPM (ScSPM), and Locality-constrained Linear Coding (LLC). Qun Li 0002, Honggang Zhang 0002, Jun Guo 0002, Bir Bhanu |
ICIP | 2 |
| 2011 | Web Multimedia Object Clustering via Information FusionabstractMultimedia information plays an increasingly important role in humans daily activities. Given a set of web multimedia objects (images with corresponding texts), a challenging problem is how to group these images into several clusters using the available information. Previous researches focus on either adopting individual information, or simply combining image and text information together for clustering. In this paper, we propose a novel approach (Dynamic Weighted Clustering) to separate images under the "supervision" of text descriptions, Also, we provide a comparative experimental investigation on utilizing text and image information to tackle web image clustering. Empirical experiments on a manually collected web multimedia object (related to the events after disasters) dataset are conducted to demonstrate the efficacy of our proposed method. Wenting Lu, Lei Li 0001, Tao Li 0001, Honggang Zhang 0002, Jun Guo 0002 |
ICDAR | 4 |
| 2011 | Weakly supervised locality sensitive hashing for duplicate image retrievalabstractLocality sensitive hashing (LSH) is quite popular in high dimensional data indexing. However, most of existing methods perform hashing in an unsupervised way, that is to say, hash functions are randomly generated without the prior information of the data. In this paper, we propose two improved LSH algorithms based on weakly supervised learning technique, which need only small quantities of labeled sample pairs. One is to select the most appropriate hash functions from a pool of functions using sample pairs labeled with “similar” or “dissimilar”. The other is to generate hash functions with positive sample pairs. The experiments show that the proposed algorithms reduce the search complexity compared with original LSH. Honggang Zhang 0002, Jun Guo 0002 |
ICIP | 2 |
| 2010 | Matching Image with Multiple Local FeaturesabstractIn this paper, we present the fusional feature composed of Affine-SIFT, MSER and color moment invariants. The fusional feature is more robust and distinctive than a single local feature. Instead of adding three local features together simply, an efficient two-level matching strategy is devised with the fusional feature, which speeds up the establishment of the local correspondences. To remove partial false positives, an affine transformation is estimated with the weighted RANSAC which decreases iteration times. The experimental results show that our approach can achieve more accurate correspondence. We prospect to apply the fusional feature and match strategy to image retrieval in the end. Honggang Zhang 0002, Jun Guo 0002 |
ICPR | 2 |
| 2010 | Local Sparse Representation Based ClassificationabstractIn this paper, we address the computational complexity issue in Sparse Representation based Classification (SRC). In SRC, it is time consuming to find a global sparse representation. To remedy this deficiency, we propose a Local Sparse Representation based Classification (LSRC) scheme, which performs sparse decomposition in local neighborhood. In LSRC, instead of solving the l1-norm constrained least square problem for all of training samples we solve a similar problem in a local neighborhood for each test sample. Experiments on face recognition data sets ORL and Extended Yale B demonstrated that the proposed LSRC algorithm can reduce the computational complexity and remain the comparative classification accuracy and robustness. Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002 |
ICPR | 3 |
| 2010 | Locality preserving and global discriminant projection with prior information
Honggang Zhang 0002, Weihong Deng, Jun Guo 0002, Jie Yang 0001 |
Mach. Vis. Appl. | 1 |
| 2009 | Learning Bundle Manifold by Double Neighborhood Graphs
Chun-Guang Li, Jun Guo 0002, Honggang Zhang 0002 |
ACCV (3) | 3 |
| 2009 | Detection of Vehicle Manufacture Logos Using Contextual Information
Wenting Lu, Honggang Zhang 0002, Kunyan Lan, Jun Guo 0002 |
ACCV (2) | 2 |
| 2009 | HCL2000 - A Large-scale Handwritten Chinese Character Database for Handwritten Character RecognitionabstractIn this paper, we present a large scale offline handwritten Chinese character database-HCL2000 which will be made public available for the research community. The database contains 3,755 frequently used simplified Chinese-characters written by 1,000 different subjects. The writerspsila information is incorporated in the database to facilitate testing on grouping writers with different background such as age, occupation, gender, and education etc. We investigate some characteristics of writing styles from different groups of writers. We evaluate HCL2000 database using three different algorithms as a baseline. We decide to publish the database along with this paper and make it free for a research purpose. Honggang Zhang 0002, Jun Guo 0002, Guang Chen 0003, Chun-Guang Li |
ICDAR | 1 |
| 2008 | Handwritten Chinese character recognition using Local Discriminant Projection with Prior InformationabstractIn this paper, we propose a new method to model the manifold of handwritten Chinese characters using the local discriminant projection. We utilize a cascade framework that combines global similarity with local discriminative cues to recognize Chinese characters. We find the similarity of different characters using a nearest-neighbor (NN) classifier, and followed by the Local Discriminant Projection with Prior Information (LDPPI) to map similar characters within a cluster to a low-dimensional space. We evaluate the proposed method on two large public datasets, ETL9B which contains 607,200 handwritten characters from 200 people, and HCL2000 which contains 3,755,000 characters written by 1,000 people. The experimental results demonstrate that the proposed method achieves 0.74% error rate on ETL9B database and 1.88% on HCL2000 database. Honggang Zhang 0002, Jie Yang 0001, Weihong Deng, Jun Guo 0002 |
ICPR | 1 |
| 2008 | Comments on "Globally Maximizing, Locally Minimizing: Unsupervised Discriminant Projection with Application to Face and Palm Biometrics"abstractIn [1], UDP is proposed to address the limitation of LPP for the clustering and classification tasks. In this communication, we show that the basic ideas of UDP and LPP are identical. In particular, UDP is just a simplified version of LPP on the assumption that the local density is uniform. Weihong Deng, Jiani Hu, Jun Guo 0002, Honggang Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |