EDBT 2026 Demo / reviewers in the wild / expert
Chao Zhang 0001
dblp:94/3019-1
· DBLP profile ↗
71ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0002-1018-4144ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
Jiahang Tu, Jiahua Dong 0001, Hanbin Zhao, Chao Zhang 0001, Nicu Sebe, Hui Qian 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware PromptingabstractRecent pre-trained vision-language models (PT-VLMs) often face a Multi-Domain Task Incremental Learning (MTIL) scenario in practice, where several classes and domains of multi-modal tasks are arrive incrementally. Without access to previously seen tasks and unseen tasks, memory-constrained MTIL suffers from forward and backward forgetting. To alleviate the above challenges, parameter-efficient fine-tuning techniques (PEFT), such as prompt tuning, are employed to adapt the PT-VLM to the diverse incrementally learned tasks. To achieve effective new task adaptation, existing methods only consider the effect of PEFT strategy selection, but neglect the influence of PEFT parameter setting (e.g., prompting). In this paper, we tackle the challenge of optimizing prompt designs for diverse tasks in MTIL and propose an Instance-Aware Prompting (IAP) framework. Specifically, our Instance-Aware Gated Prompting (IA-GP) strategy enhances adaptation to new tasks while mitigating forgetting by adaptively assigning prompts across transformer layers at the instance level. Our Instance-Aware Class-Distribution-Driven Prompting (IA-CDDP) improves the task adaptation process by determining an accurate task-label-related confidence score for each instance. Experimental evaluations across 11 datasets, using three performance metrics, demonstrate the effectiveness of our proposed method. The source codes are available at https://github.com/FerdinandZJU/IAP. Hao Fu 0023, Hanbin Zhao, Jiahua Dong 0001, Henghui Ding, Chao Zhang 0001, Hui Qian 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental Unlearning
Jiahang Tu, Mintong Kang, Hanbin Zhao, Chao Zhang 0001, Hui Qian 0001 |
ICCV | 5 |
| 2025 | EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestabstractThe sequential nature of modern LLMs makes them expensive and slow, and speculative sam- pling has proven to be an effective solution to this problem. Methods like EAGLE perform autoregression at the feature level, reusing top- layer features from the target model to achieve better results than vanilla speculative sampling. A growing trend in the LLM community is scaling up training data to improve model intelligence without increasing inference costs. However, we observe that scaling up data provides limited improvements for EAGLE. We identify that this limitation arises from EAGLE’s feature prediction constraints. In this paper, we introduce EAGLE-3, which abandons feature prediction in favor of direct token prediction and replaces reliance on top-layer features with multi-layer feature fusion via a technique named training-time test. These improvements significantly enhance performance and enable the draft model to fully benefit from scaling up training data. Our experiments include both chat models and reasoning models, evaluated on five tasks. The results show that EAGLE-3 achieves a speedup ratio up to 6.5x, with about 1.4x improvement over EAGLE-2. In the SGLang framework, EAGLE- 3 achieves a 1.38x throughput improvement at a batch size of 64. Fangyun Wei, Chao Zhang 0001, Hongyang Zhang 0001 |
NeurIPS | 3 |
| 2025 | Fighting Malicious Media Data: A Survey on Tampering Detection and Deepfake DetectionabstractOnline media data, in the form of images and videos, are becoming mainstream communication channels. However, recent advances in deep learning (DL), particularly deep generative models, open the doors for producing perceptually convincing images and videos at a low cost, which not only poses a serious threat to the trustworthiness of digital information but also has severe societal implications. This motivates a growing interest in research in media tampering detection (TD), i.e., using DL techniques to examine whether media data have been maliciously manipulated. Depending on the content of the targeted images, media forgery could be divided into image tampering and Deepfake techniques. The former typically moves or erases the visual elements in ordinary images, while the latter manipulates the expressions and even the identity of human faces. Accordingly, the means of defense include image TD and Deepfake detection (DFD), which share a wide variety of properties. In this article, we provide a comprehensive review of the current media TD approaches and discuss the challenges and trends in this field for future research. Zhenxin Li, Chao Zhang 0001, Jingjing Chen 0001, Zuxuan Wu, Larry Davis 0001, Yu-Gang Jiang 0001 |
Proc. IEEE | 3 |
| 2025 | PECTP: Parameter-Efficient Cross-Task Prompts for Incremental Vision TransformerabstractIncremental Learning (IL) aims to learn deep models on sequential tasks continually, where each new task includes a batch of new classes and deep models have no access to task ID information at the inference time. Recent vast pre-trained models (PTMs) have achieved outstanding performance by prompt technique in practical IL without the old samples (rehearsal-free) and with a memory constraint (memory-constrained): Prompt-extending and Prompt-fixed methods. However, prompt-extending methods need a large memory buffer to maintain an ever-expanding prompt pool and meet an extra challenging prompt selection problem. Prompt-fixed methods only learn a single set of prompts on one of the incremental tasks and can not handle all the incremental tasks effectively. To achieve a good balance between the memory cost and the performance on all the tasks, we propose a Parameter-Efficient Cross-Task Prompt (PECTP) framework with Prompt Retention Module (PRM) and classifier Head Retention Module (HRM). To make the final learned prompts effective on all incremental tasks, PRM constrains the evolution of cross-task prompts’ parameters from Outer Prompt Granularity and Inner Prompt Granularity. Besides, we employ HRM to inherit old knowledge in the previously learned classifier heads to facilitate the cross-task prompts’ generalization ability. Extensive experiments show the effectiveness of our method. The source codes are available at https://github.com/RAIAN08/PECTP. Hanbin Zhao, Chao Zhang 0001, Jiahua Dong 0001, Henghui Ding, Yu-Gang Jiang 0001, Hui Qian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | EAGLE-2: Faster Inference of Language Models with Dynamic Draft TreesabstractInference with modern Large Language Models (LLMs) is expensive and time-consuming, and speculative sampling has proven to be an effective solution.Most speculative sampling methods such as EAGLE use a static draft tree, implicitly assuming that the acceptance rate of draft tokens depends only on their position.Interestingly, we found that the acceptance rate of draft tokens is also contextdependent.In this paper, building upon EA-GLE, we propose EAGLE-2, which introduces a new technique of context-aware dynamic draft tree into drafting modeling.This improvement leverages the fact that the draft model of EAGLE is well-calibrated: the confidence scores from the draft model approximate acceptance rates with small errors.We conducted extensive evaluations on three series of LLMs and six tasks, with EAGLE-2 achieving speedup ratios 3.05x-4.26x,which is 20%-40% faster than EAGLE-1.EAGLE-2 also ensures that the distribution of the generated text remains unchanged, making it a lossless acceleration algorithm.The code is open sourced at https://github.com/SafeAILab/EAGLE. Fangyun Wei, Chao Zhang 0001, Hongyang Zhang 0001 |
EMNLP | 3 |
| 2024 | RAIN: Your Language Models Can Align Themselves without FinetuningabstractLarge language models (LLMs) often demonstrate inconsistencies with human preferences. Previous research typically gathered human preference data and then aligned the pre-trained models using reinforcement learning or instruction tuning, a.k.a. the finetuning step. In contrast, aligning frozen LLMs without requiring alignment data is more appealing. This work explores the potential of the latter setting. We discover that by integrating self-evaluation and rewind mechanisms, unaligned LLMs can directly produce responses consistent with human preferences via self-boosting. We introduce a novel inference method, Rewindable Auto-regressive INference (RAIN), that allows pre-trained LLMs to evaluate their own generation and use the evaluation results to guide rewind and generation for AI safety. Notably, RAIN operates without the need of extra data for model alignment and abstains from any training, gradient computation, or parameter updates. Experimental results evaluated by GPT-4 and humans demonstrate the effectiveness of RAIN: on the HH dataset, RAIN improves the harmlessness rate of LLaMA 30B from 82% of vanilla inference to 97%, while maintaining the helpfulness rate. On the TruthfulQA dataset, RAIN improves the truthfulness of the already-well-aligned LLaMA-2-chat 13B model by 5%. Fangyun Wei, Jinjing Zhao, Chao Zhang 0001, Hongyang Zhang 0001 |
ICLR | 4 |
| 2024 | GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerabstractCross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal features, and demonstrate exchange based methods underperform cross-attention mechanisms, while the computational demand of the latter inevitably restricts its use with longer sequences. To surmount the computational challenges, we propose *GeminiFusion*, a pixel-wise fusion approach that capitalizes on aligned cross-modal representations. *GeminiFusion* elegantly combines intra-modal and inter-modal attentions, dynamically integrating complementary information across modalities. We employ a layer-adaptive noise to adaptively control their interplay on a per-layer basis, thereby achieving a harmonized fusion process. Notably, *GeminiFusion* maintains linear complexity with respect to the number of input tokens, ensuring this multimodal framework operates with efficiency comparable to unimodal networks. Comprehensive evaluations across multimodal image-to-image translation, $3$D object detection and arbitrary-modal semantic segmentation tasks, including RGB, depth, LiDAR, event data, etc. demonstrate the superior performance of our *GeminiFusion* against leading-edge techniques. The PyTorch code is available [here](https://github.com/JiaDingCN/GeminiFusion). Ding Jia, Jianyuan Guo, Kai Han 0002, Han Wu 0009, Chao Zhang 0001, Chang Xu 0002, Xinghao Chen 0001 |
ICML | 5 |
| 2024 | EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyabstractAutoregressive decoding makes the inference of Large Language Models (LLMs) time-consuming. In this paper, we reconsider speculative sampling and derive two key observations. Firstly, autoregression at the feature (second-to-top-layer) level is more straightforward than at the token level. Secondly, the inherent uncertainty in feature (second-to-top-layer) level autoregression constrains its performance. Based on these insights, we introduce EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency), a simple yet highly efficient speculative sampling framework. By incorporating a token sequence advanced by one time step, EAGLE effectively resolves the uncertainty, enabling precise second-to-top-layer feature prediction with minimal overhead. We conducted comprehensive evaluations of EAGLE, including all models from the Vicuna and LLaMA2-Chat series, the MoE model Mixtral 8x7B Instruct, and tasks in dialogue, code generation, mathematical reasoning, and instruction following. For LLaMA2-Chat 70B, EAGLE achieved a latency speedup ratio of **2.7x-3.5x**, doubled throughput, while maintaining the distribution of the generated text. Fangyun Wei, Chao Zhang 0001, Hongyang Zhang 0001 |
ICML | 3 |
| 2024 | Direct-Effect Risk Minimization for Domain Generalization
Zejia Wu, Chao Zhang 0001, Hongyang Zhang 0001 |
ECML/PKDD (3) | 3 |
| 2024 | Self-Adaptive Training: Bridging Supervised and Self-Supervised LearningabstractWe propose self-adaptive training-a unified training algorithm that dynamically calibrates and enhances training processes by model predictions without incurring an extra computational cost-to advance both supervised and self-supervised learning of deep neural networks. We analyze the training dynamics of deep networks on training data that are corrupted by, e.g., random noise and adversarial examples. Our analysis shows that model predictions are able to magnify useful underlying information in data and this phenomenon occurs broadly even in the absence of any label information, highlighting that model predictions could substantially benefit the training processes: self-adaptive training improves the generalization of deep networks under noise and enhances the self-supervised representation learning. The analysis also sheds light on understanding deep learning, e.g., a potential explanation of the recently-discovered double-descent phenomenon in empirical risk minimization and the collapsing issue of the state-of-the-art self-supervised learning algorithms. Experiments on the CIFAR, STL, and ImageNet datasets verify the effectiveness of our approach in three applications: classification with label noise, selective classification, and linear evaluation. To facilitate future research, the code has been made publicly available at https://github.com/LayneH/self-adaptive-training. Lang Huang 0001, Chao Zhang 0001, Hongyang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Expediting Large-Scale Vision Transformer for Dense Prediction Without Fine-TuningabstractIn a wide range of dense prediction tasks, large-scale Vision Transformers have achieved state-of-the-art performance while requiring expensive computation. In contrast to most existing approaches accelerating Vision Transformers for image classification, we focus on accelerating Vision Transformers for dense prediction without any fine-tuning. We present two non-parametric operators specialized for dense prediction tasks, a token clustering layer to decrease the number of tokens for expediting and a token reconstruction layer to increase the number of tokens for recovering high-resolution. To accomplish this, the following steps are taken: i) token clustering layer is employed to cluster the neighboring tokens and yield low-resolution representations with spatial structures; ii) the following transformer layers are performed only to these clustered low-resolution tokens; and iii) reconstruction of high-resolution representations from refined low-resolution representations is accomplished using token reconstruction layer. The proposed approach shows promising results consistently on 6 dense prediction tasks, including object detection, semantic segmentation, panoptic segmentation, instance segmentation, depth estimation, and video instance segmentation. Additionally, we validate the effectiveness of the proposed approach on the very recent state-of-the-art open-vocabulary recognition methods. Furthermore, a number of recent representative approaches are benchmarked and compared on dense prediction tasks. Yuhui Yuan, Weicong Liang, Henghui Ding, Chao Zhang 0001, Han Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | DETRs with Hybrid MatchingabstractOne-to-one set matching is a key design for DETR to establish its end-to-end capability, so that object detection does not require a hand-crafted NMS (non-maximum suppression) to remove duplicate detections. This end-to-end signature is important for the versatility of DETR, and it has been generalized to broader vision tasks. However, we note that there are few queries assigned as positive samples and the one-to-one set matching significantly reduces the training efficacy of positive samples. We propose a simple yet effective method based on a hybrid matching scheme that combines the original one-to-one matching branch with an auxiliary one-to-many matching branch during training. Our hybrid strategy has been shown to significantly improve accuracy. In inference, only the original one-to-one match branch is used, thus maintaining the end-to-end merit and the same inference efficiency of DETR. The method is namedℋ-DETR, and it shows that a wide range of representative DETR methods can be consistently improved across a wide range of visual tasks, including Deformable-DETR, PETRv2, PETR, and TransTrack, among others. Code is available at: https://github.com/HDETR. Ding Jia, Yuhui Yuan, Haodi He, Xiaopei Wu, Haojun Yu, Weihong Lin, Lei Sun 0003, Chao Zhang 0001, Han Hu 0001 |
CVPR | 8 |
| 2023 | Rank-DETR for High Quality Object DetectionabstractModern detection transformers (DETRs) use a set of object queries to predict a list of bounding boxes, sort them by their classification confidence scores, and select the top-ranked predictions as the final detection results for the given input image. A highly performant object detector requires accurate ranking for the bounding box predictions. For DETR-based detectors, the top-ranked bounding boxes suffer from less accurate localization quality due to the misalignment between classification scores and localization accuracy, thus impeding the construction of high-quality detectors. In this work, we introduce a simple and highly performant DETR-based object detector by proposing a series of rank-oriented designs, combinedly called Rank-DETR. Our key contributions include: (i) a rank-oriented architecture design that can prompt positive predictions and suppress the negative ones to ensure lower false positive rates, as well as (ii) a rank-oriented loss function and matching cost design that prioritizes predictions of more accurate localization accuracy during ranking to boost the AP under high IoU thresholds. We apply our method to improve the recent SOTA methods (e.g., H-DETR and DINO-DETR) and report strong COCO object detection results when using different backbones such as ResNet-$50$, Swin-T, and Swin-L, demonstrating the effectiveness of our approach. Code is available at \url{https://github.com/LeapLabTHU/Rank-DETR}. Yifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan, Yukang Yang, Chao Zhang 0001, Han Hu 0001, Gao Huang 0001 |
NeurIPS | 6 |
| 2023 | DRRNets: Dynamic Recurrent Routing via Low-Rank Regularization in Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) continue to show outstanding performance in sequence learning tasks such as language modeling, but it remains difficult to train RNNs for long sequences. The main challenges lie in the complex dependencies, gradient vanishing or exploding, and low resource requirement in model deployment. In order to address these challenges, we propose dynamic recurrent routing neural networks (DRRNets), which can: 1) shorten the recurrent lengths by allocating recurrent routes dynamically for different dependencies and 2) reduce the number of parameters significantly by imposing low-rank constraints on the fully connected layers. A novel optimization algorithm via low-rank constraint and sparsity projection is developed to train the network. We verify the effectiveness of the proposed method by comparing it with multiple competitive approaches in several popular sequential learning tasks, such as language modeling and speaker recognition. The results in terms of different criteria demonstrate the superiority of our proposed method. Dongjing Shan, Yong Luo 0002, Xiongwei Zhang, Chao Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Learning Efficient Vision Transformers via Fine-Grained Manifold DistillationabstractIn the past few years, transformers have achieved promising performance on various computer vision tasks. Unfortunately, the immense inference overhead of most existing vision transformers withholds them from being deployed on edge devices such as cell phones and smart watches. Knowledge distillation is a widely used paradigm for compressing cumbersome architectures into compact students via transferring information. However, most of them are designed for convolutional neural networks (CNNs), which do not fully investigate the character of vision transformers. In this paper, we fully utilize the patch-level information and propose a fine-grained manifold distillation method for transformer-based networks. Specifically, we train a tiny student model to match a pre-trained teacher model in the patch-level manifold space. Then, we decouple the manifold matching loss into three terms with careful design to further reduce the computational costs for the patch relationship. Equipped with the proposed method, a DeiT-Tiny model containing 5M parameters achieves 76.5\% top-1 accuracy on ImageNet-1k, which is +2.0\% higher than previous distillation approaches. Transfer learning results on other classification benchmarks and downstream vision tasks also demonstrate the superiority of our method over the state-of-the-art algorithms. Zhiwei Hao 0001, Jianyuan Guo, Ding Jia, Kai Han 0002, Yehui Tang 0001, Chao Zhang 0001, Han Hu 0001, Yunhe Wang 0001 |
NeurIPS | 6 |
| 2022 | Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuningabstractVision transformers have recently achieved competitive results across various vision tasks but still suffer from heavy computation costs when processing a large number of tokens. Many advanced approaches have been developed to reduce the total number of tokens in the large-scale vision transformers, especially for image classification tasks. Typically, they select a small group of essential tokens according to their relevance with the [\texttt{class}] token, then fine-tune the weights of the vision transformer. Such fine-tuning is less practical for dense prediction due to the much heavier computation and GPU memory cost than image classification.In this paper, we focus on a more challenging problem, \ie, accelerating large-scale vision transformers for dense prediction without any additional re-training or fine-tuning. In response to the fact that high-resolution representations are necessary for dense prediction, we present two non-parametric operators, a \emph{token clustering layer} to decrease the number of tokens and a \emph{token reconstruction layer} to increase the number of tokens. The following steps are performed to achieve this: (i) we use the token clustering layer to cluster the neighboring tokens together, resulting in low-resolution representations that maintain the spatial structures; (ii) we apply the following transformer layers only to these low-resolution representations or clustered tokens; and (iii) we use the token reconstruction layer to re-create the high-resolution representations from the refined low-resolution representations. The results obtained by our method are promising on five dense prediction tasks including object detection, semantic segmentation, panoptic segmentation, instance segmentation, and depth estimation. Accordingly, our method accelerates $40\%\uparrow$ FPS and saves $30\%\downarrow$ GFLOPs of ``Segmenter+ViT-L/$16$'' while maintaining $99.5\%$ of the performance on ADE$20$K without fine-tuning the official weights. Weicong Liang, Yuhui Yuan, Henghui Ding, Xiao Luo 0001, Weihong Lin, Ding Jia, Zheng Zhang 0022, Chao Zhang 0001, Han Hu 0001 |
NeurIPS | 8 |
| 2021 | Positive-Unlabeled Data Purification in the Wild for Object DetectionabstractDeep learning based object detection approaches have achieved great progress with the benefit from large amount of labeled images. However, image annotation remains a laborious, time-consuming and error-prone process. To further improve the performance of detectors, we seek to exploit all available labeled data and excavate useful samples from massive unlabeled images in the wild, which is rarely discussed before. In this paper, we present a positive-unlabeled learning based scheme to expand training data by purifying valuable images from massive unlabeled ones, where the original training data are viewed as positive data and the unlabeled images in the wild are unlabeled data. To effectively utilized these purified data, we propose a self-distillation algorithm based on hint learning and ground truth bounded knowledge distillation. Experimental results verify that the proposed positive-unlabeled data purification can strengthen the original detector by mining the massive unlabeled data. In particular, our method boosts the mAP of FPN by +2.0% on COCO benchmark. Jianyuan Guo, Kai Han 0002, Han Wu 0009, Chao Zhang 0001, Xinghao Chen 0001, Chunjing Xu, Chang Xu 0002, Yunhe Wang 0001 |
CVPR | 4 |
| 2021 | F-Net: Fusion Neural Network for Vehicle Trajectory Prediction in Autonomous DrivingabstractRecent research has been remarkable in recurrent neural networks (RNNs) on sequence-to-sequence problems for image caption, and promising in convolutional neural networks (CNNs) on spatial analysis problems for image detection and sematic segmentation problems. In this paper, based on recurrent neural networks and convolutional neural networks, we propose a fusion neural network architecture named F-Net to deal with vehicle trajectory prediction on highway and urban scenarios in autonomous driving applications. The novelty of the proposed method is the attention mechanism that affects effectively in the progress of both RNN and CNN feature extraction. Besides, our sufficient usage of raw sensor data protects scene texture information of environment and interaction among surrounding vehicles. Experimental results on the nuScene dataset show that our proposed method outperforms the state-of-the-art methods. Jue Wang 0004, Ping Wang 0003, Chao Zhang 0001, Kuifeng Su, Jun Li 0010 |
ICASSP | 3 |
| 2021 | HRFormer: High-Resolution Vision Transformer for Dense PredictabstractWe present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and computational cost. We take advantage of the multi-resolution parallel design introduced in high-resolution convolutional networks (HRNet [45]), along with local-window self-attention that performs self-attention over small non-overlapping image windows [21], for improving the memory and computation efficiency. In addition, we introduce a convolution into the FFN to exchange information across the disconnected image windows. We demonstrate the effectiveness of the HighResolution Transformer on both human pose estimation and semantic segmentation tasks, e.g., HRFormer outperforms Swin transformer [27] by 1.3 AP on COCO pose estimation with 50% fewer parameters and 30% fewer FLOPs. Code is available at: https://github.com/HRNet/HRFormer Yuhui Yuan, Lang Huang 0001, Weihong Lin, Chao Zhang 0001, Xilin Chen 0001, Jingdong Wang 0001 |
NeurIPS | 5 |
| 2021 | OCNet: Object Context for Semantic Segmentation
Yuhui Yuan, Lang Huang 0001, Jianyuan Guo, Chao Zhang 0001, Xilin Chen 0001, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | An optimization scheme for segmented-memory neural network
Dongjing Shan, Chao Zhang 0001, Yongjian Nian |
Neurocomputing | 2 |
| 2020 | Hit-Detector: Hierarchical Trinity Architecture Search for Object DetectionabstractNeural Architecture Search (NAS) has achieved great success in image classification task. Some recent works have managed to explore the automatic design of efficient backbone or feature fusion layer for object detection. However, these methods focus on searching only one certain component of object detector while leaving others manually designed. We identify the inconsistency between searched component and manually designed ones would withhold the detector of stronger performance. To this end, we propose a hierarchical trinity search framework to simultaneously discover efficient architectures for all components (i.e. backbone, neck, and head) of object detector in an end-to-end manner. In addition, we empirically reveal that different parts of the detector prefer different operators. Motivated by this, we employ a novel scheme to automatically screen different sub search spaces for different components so as to perform the end-to-end search for each component on the corresponding sub search space efficiently. Without bells and whistles, our searched architecture, namely Hit-Detector, achieves 41.4% mAP on COCO minival set with 27M parameters. Our implementation is available at https://github.com/ggjy/HitDet.pytorch. Jianyuan Guo, Kai Han 0002, Yunhe Wang 0001, Chao Zhang 0001, Zhaohui Yang 0003, Han Wu 0009, Xinghao Chen 0001, Chang Xu 0002 |
CVPR | 4 |
| 2020 | Self-Adaptive Training: beyond Empirical Risk MinimizationabstractWe propose self-adaptive training---a new training algorithm that dynamically calibrates training process by model predictions without incurring extra computational cost---to improve generalization of deep learning for potentially corrupted training data. This problem is important to robustly learning from data that are corrupted by, e.g., random noises and adversarial examples. The standard empirical risk minimization (ERM) for such data, however, may easily overfit noises and thus suffers from sub-optimal performance. In this paper, we observe that model predictions can substantially benefit the training process: self-adaptive training significantly mitigates the overfitting issue and improves generalization over ERM under both random and adversarial noises. Besides, in sharp contrast to the recently-discovered double-descent phenomenon in ERM, self-adaptive training exhibits a single-descent error-capacity curve, indicating that such a phenomenon might be a result of overfitting of noises. Experiments on the CIFAR and ImageNet datasets verify the effectiveness of our approach in two applications: classification with label noise and selective classification. Lang Huang 0001, Chao Zhang 0001, Hongyang Zhang 0001 |
NeurIPS | 2 |
| 2019 | Low-resolution Visual Recognition via Deep Feature DistillationabstractHere we study the low-resolution visual recognition problem. Conventional methods are usually trained on images with large ROIs (regions of interest), while the regions and insider images are often small and blur in real-world applications. Therefore, deep neural networks learned on high-resolution images cannot be directly used for recognizing low-resolution objects. To overcome this challenging problem, we propose to use the teacher-student learning paradigm for distilling useful feature information from a pre-trained deep model on high-resolution visual data. In practice, a distillation loss is used to seek the perceptual consistency of low-resolution images and high-resolution images. By simultaneously optimizing the recognition loss and distillation loss, we formulate a novel low-resolution recognition approach. Experiments conducted on benchmarks demonstrate that the proposed method is capable to learn well-performed models for recognizing low-resolution objects, which is superior to the state-of-the-art methods. Mingjian Zhu, Kai Han 0002, Chao Zhang 0001, Jinlong Lin, Yunhe Wang 0001 |
ICASSP | 3 |
| 2019 | Beyond Human Parts: Dual Part-Aligned Representations for Person Re-IdentificationabstractPerson re-identification is a challenging task due to various complex factors. Recent studies have attempted to integrate human parsing results or externally defined attributes to help capture human parts or important object regions. On the other hand, there still exist many useful contextual cues that do not fall into the scope of predefined human parts or attributes. In this paper, we address the missed contextual cues by exploiting both the accurate human parts and the coarse non-human parts. In our implementation, we apply a human parsing model to extract the binary human part masks and a self-attention mechanism to capture the soft latent (non-human) part masks. We verify the effectiveness of our approach with new state-of-the-art performance on three challenging benchmarks: Market-1501, DukeMTMC-reID and CUHK03. Our implementation is available at https://github.com/ggjy/P2Net.pytorch. Jianyuan Guo, Yuhui Yuan, Lang Huang 0001, Chao Zhang 0001, Jin-Ge Yao, Kai Han 0002 |
ICCV | 4 |
| 2019 | Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalabstractCross-modal hashing encodes the multimedia data into a common binary hash space in which the correlations among the samples from different modalities can be effectively measured. Deep cross-modal hashing further improves the retrieval performance as the deep neural networks can generate more semantic relevant features and hash codes. In this paper, we study the unsupervised deep cross-modal hash coding and propose Deep Joint-Semantics Reconstructing Hashing (DJSRH), which has the following two main advantages. First, to learn binary codes that preserve the neighborhood structure of the original data, DJSRH constructs a novel joint-semantics affinity matrix which elaborately integrates the original neighborhood information from different modalities and accordingly is capable to capture the latent intrinsic semantic affinity for the input multi-modal instances. Second, DJSRH later trains the networks to generate binary codes that maximally reconstruct above joint-semantics relations via the proposed reconstructing framework, which is more competent for the batch-wise training as it reconstructs the specific similarity value unlike the common Laplacian constraint merely preserving the similarity order. Extensive experiments demonstrate the significant improvement by DJSRH in various cross-modal retrieval tasks. Shupeng Su, Zhisheng Zhong, Chao Zhang 0001 |
ICCV | 3 |
| 2019 | ADA-Tucker: Compressing deep neural networks via adaptive dimension adjustment tucker decomposition
Zhisheng Zhong, Fangyin Wei, Zhouchen Lin, Chao Zhang 0001 |
Neural Networks | 4 |
| 2018 | Autoencoder Inspired Unsupervised Feature SelectionabstractHigh-dimensional data in many areas such as computer vision and machine learning tasks brings in computational and analytical difficulty. Feature selection which selects a subset from observed features is a widely used approach for improving performance and effectiveness of machine learning models with high-dimensional data. In this paper, we propose a novel AutoEncoder Feature Selector (AEFS) for unsupervised feature selection which combines autoencoder regression and group lasso tasks. Compared to traditional feature selection methods, AEFS can select the most important features by excavating both linear and nonlinear information among features, which is more flexible than the conventional self-representation method for unsupervised feature selection with only linear assumptions. Experimental results on benchmark dataset show that the proposed method is superior to the state-of-the-art method. Kai Han 0002, Yunhe Wang 0001, Chao Zhang 0001, Chao Xu 0006 |
ICASSP | 3 |
| 2018 | Attribute-Aware Attention Model for Fine-grained Representation LearningabstractHow to learn a discriminative fine-grained representation is a key point in many computer vision applications, such as person re-identification, fine-grained classification, fine-grained image retrieval, etc. Most of the previous methods focus on learning metrics or ensemble to derive better global representation, which are usually lack of local information. Based on the considerations above, we propose a novel Attribute-Aware Attention Model ($A^3M$), which can learn local attribute representation and global category representation simultaneously in an end-to-end manner. The proposed model contains two attention models: attribute-guided attention module uses attribute information to help select category features in different regions, at the same time, category-guided attention module selects local features of different attributes with the help of category cues. Through this attribute-category reciprocal process, local and global features benefit from each other. Finally, the resulting feature contains more intrinsic information for image recognition instead of the noisy and irrelevant features. Extensive experiments conducted on Market-1501, CompCars, CUB-200-2011 and CARS196 demonstrate the effectiveness of our $A^3M$. Kai Han 0002, Jianyuan Guo, Chao Zhang 0001, Mingjian Zhu |
ACM Multimedia | 3 |
| 2018 | Greedy Hash: Towards Fast Optimization for Accurate Hash Coding in CNNabstractTo convert the input into binary code, hashing algorithm has been widely used for approximate nearest neighbor search on large-scale image sets due to its computation and storage efficiency. Deep hashing further improves the retrieval quality by combining the hash coding with deep neural network. However, a major difficulty in deep hashing lies in the discrete constraints imposed on the network output, which generally makes the optimization NP hard. In this work, we adopt the greedy principle to tackle this NP hard problem by iteratively updating the network toward the probable optimal discrete solution in each iteration. A hash coding layer is designed to implement our approach which strictly uses the sign function in forward propagation to maintain the discrete constraints, while in back propagation the gradients are transmitted intactly to the front layer to avoid the vanishing gradients. In addition to the theoretical derivation, we provide a new perspective to visualize and understand the effectiveness and efficiency of our algorithm. Experiments on benchmark datasets show that our scheme outperforms state-of-the-art hashing methods in both supervised and unsupervised tasks. Shupeng Su, Chao Zhang 0001, Kai Han 0002, Yonghong Tian 0001 |
NeurIPS | 2 |
| 2018 | Joint Sub-bands Learning with Clique Structures for Wavelet Domain Super-ResolutionabstractConvolutional neural networks (CNNs) have recently achieved great success in single-image super-resolution (SISR). However, these methods tend to produce over-smoothed outputs and miss some textural details. To solve these problems, we propose the Super-Resolution CliqueNet (SRCliqueNet) to reconstruct the high resolution (HR) image with better textural details in the wavelet domain. The proposed SRCliqueNet firstly extracts a set of feature maps from the low resolution (LR) image by the clique blocks group. Then we send the set of feature maps to the clique up-sampling module to reconstruct the HR image. The clique up-sampling module consists of four sub-nets which predict the high resolution wavelet coefficients of four sub-bands. Since we consider the edge feature properties of four sub-bands, the four sub-nets are connected to the others so that they can learn the coefficients of four sub-bands jointly. Finally we apply inverse discrete wavelet transform (IDWT) to the output of four sub-nets at the end of the clique up-sampling module to increase the resolution and reconstruct the HR image. Extensive quantitative and qualitative experiments on benchmark datasets show that our method achieves superior performance over the state-of-the-art methods. Zhisheng Zhong, Tiancheng Shen, Zhouchen Lin, Chao Zhang 0001 |
NeurIPS | 5 |
| 2018 | Dictionary learning with structured noise
Pan Zhou 0002, Cong Fang 0001, Zhouchen Lin, Chao Zhang 0001, Edward Y. Chang |
Neurocomputing | 4 |
| 2018 | Prior Fusion and Feature Transformation-Based Principal Component Analysis for Saliency DetectionabstractIn this paper, we propose a prior fusion and feature transformation-based principal component analysis (PCA) method for saliency detection. It relies on the inner statistics of the patches in the image for identifying unique patterns, and all the processes are done only once. First, three low-level priors are incorporated and act as guidance cues in the model; second, to ensure the validity of PCA distinctness model, a linear transform for the feature space is designed and needs to be trained; furthermore, an extended optimization framework is utilized to generate a smoothed saliency map based on the consistency of the adjacent patches. We compare three versions of our model with seven previous methods and test them on several benchmark datasets. Different kinds of strategies are adopted to evaluate the performance and the results demonstrate that our model achieves the state-of-the-art performance. Dongjing Shan, Chao Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2018 | Visual saliency based on extended manifold ranking and third-order optimization refinement
Dongjing Shan, Xiongwei Zhang, Chao Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Tensor Factorization for Low-Rank Tensor CompletionabstractRecently, a tensor nuclear norm (TNN) based method was proposed to solve the tensor completion problem, which has achieved state-of-the-art performance on image and video inpainting tasks. However, it requires computing tensor singular value decomposition (t-SVD), which costs much computation and thus cannot efficiently handle tensor data, due to its natural large scale. Motivated by TNN, we propose a novel low-rank tensor factorization method for efficiently solving the 3-way tensor completion problem. Our method preserves the low-rank structure of a tensor by factorizing it into the product of two tensors of smaller sizes. In the optimization process, our method only needs to update two smaller tensors, which can be more efficiently conducted than computing t-SVD. Furthermore, we prove that the proposed alternating minimization algorithm can converge to a Karush-Kuhn-Tucker point. Experimental results on the synthetic data recovery, image and video inpainting tasks clearly demonstrate the superior performance and efficiency of our developed method over state-of-the-arts including the TNN and matricization methods. Pan Zhou 0002, Canyi Lu, Zhouchen Lin, Chao Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | Hard-Aware Deeply Cascaded EmbeddingabstractRiding on the waves of deep neural networks, deep metric learning has achieved promising results in various tasks by using triplet network or Siamese network. Though the basic goal of making images from the same category closer than the ones from different categories is intuitive, it is hard to optimize the objective directly due to the quadratic or cubic sample size. Hard example mining is widely used to solve the problem, which spends the expensive computation on a subset of samples that are considered hard. However, hard is defined relative to a specific model. Then complex models will treat most samples as easy ones and vice versa for simple models, both of which are not good for training. It is difficult to define a model with the just right complexity and choose hard examples adequately as different samples are of diverse hard levels. This motivates us to propose the novel framework named Hard-Aware Deeply Cascaded Embedding(HDC) to ensemble a set of models with different complexities in cascaded manner to mine hard examples at multiple levels. A sample is judged by a series of models with increasing complexities and only updates models that consider the sample as a hard case. The HDC is evaluated on CARS196, CUB-200-2011, Stanford Online Products, VehicleID and DeepFashion datasets, and outperforms state-of-the-art methods by a large margin. Yuhui Yuan, Kuiyuan Yang, Chao Zhang 0001 |
ICCV | 3 |
| 2017 | Bilevel Model-Based Discriminative Dictionary Learning for RecognitionabstractMost supervised dictionary learning methods optimize the combinations of reconstruction error, sparsity prior, and discriminative terms. Thus, the learnt dictionaries may not be optimal for recognition tasks. Also, the sparse codes learning models in the training and the testing phases are inconsistent. Besides, without utilizing the intrinsic data structure, many dictionary learning methods only employ the$\ell _{0}$or$\ell _{1}$norm to encode each datum independently, limiting the performance of the learnt dictionaries. We present a novel bilevel model-based discriminative dictionary learning method for recognition tasks. The upper level directly minimizes the classification error, while the lower level uses the sparsity term and the Laplacian term to characterize the intrinsic data structure. The lower level is subordinate to the upper level. Therefore, our model achieves an overall optimality for recognition in that the learnt dictionary is directly tailored for recognition. Moreover, the sparse codes learning models in the training and the testing phases can be the same. We further propose a novel method to solve our bilevel optimization problem. It first replaces the lower level with its Karush-Kuhn–Tucker conditions and then applies the alternating direction method of multipliers to solve the equivalent problem. Extensive experiments demonstrate the effectiveness and robustness of our method. Pan Zhou 0002, Chao Zhang 0001, Zhouchen Lin |
IEEE Trans. Image Process. | 2 |
| 2016 | Universal Demosaicking of Color Filter ArraysabstractA large number of color filter arrays (CFAs), periodic or aperiodic, have been proposed. To reconstruct images from all different CFAs and compare their imaging quality, a universal demosaicking method is needed. This paper proposes a new universal demosaicking method based on inter-pixel chrominance capture and optimal demosaicking transformation. It skips the commonly used step to estimate the luminance component at each pixel, and thus, avoids the associated estimation error. Instead, we directly use the acquired CFA color intensity at each pixel as an input component. Two independent chrominance components are estimated at each pixel based on the inter-pixel chrominance in the window, which is captured with the difference of CFA color values between the pixel of interest and its neighbors. Two mechanisms are employed for the accurate estimation: distance-related and edge-sensing weighting to reflect the confidence levels of the inter-pixel chrominance components, and pseudoinverse-based estimation from the components in a window. Then from the acquired CFA color component and two estimated chrominance components, the three primary colors are reconstructed by a linear color transform, which is optimized for the least transform error. Our experiments show that the proposed method is much better than other published universal demosaicking methods. Chao Zhang 0001, Yan Li 0009, Jue Wang 0004, Pengwei Hao |
IEEE Trans. Image Process. | 1 |
| 2016 | Completing Low-Rank Matrices With Corrupted Samples From Few Coefficients in General BasisabstractSubspace recovery from the corrupted and missing data is crucial for various applications in signal processing and information theory. To complete missing values and detect column corruptions, the existing robust matrix completion (MC) methods mostly concentrate on recovering a low-rank matrix from a few corrupted coefficients with respect to standard basis, which, however, does not apply to more general basis, e.g., Fourier basis. In this paper, we prove that the range space of an m × n matrix with rank r can be exactly recovered from a few coefficients with respect to general basis, though r and the number of corrupted samples are both as high as O(min{m, n}/ log3(m + n)). Our model covers the previous ones as special cases, and robust MC can recover the intrinsic matrix with a higher rank. Moreover, we suggest a universal choice of the regularization parameter, which is λ = 1/√log n. By our ℓ2,1filtering algorithm, which has theoretical guarantees, we can further reduce the computational cost of our model. As an application, we also find that the solutions to extended robust low-rank representation and to our extended robust MC are mutually expressible, so both our theory and algorithm can be applied to the subspace clustering problem with missing values under certain conditions. The experiments verify our theories. Hongyang Zhang 0001, Zhouchen Lin, Chao Zhang 0001 |
IEEE Trans. Inf. Theory | 3 |
| 2016 | Integrated Low-Rank-Based Discriminative Feature Learning for RecognitionabstractFeature learning plays a central role in pattern recognition. In recent years, many representation-based feature learning methods have been proposed and have achieved great success in many applications. However, these methods perform feature learning and subsequent classification in two separate steps, which may not be optimal for recognition tasks. In this paper, we present a supervised low-rank-based approach for learning discriminative features. By integrating latent low-rank representation (LatLRR) with a ridge regression-based classifier, our approach combines feature learning with classification, so that the regulated classification error is minimized. In this way, the extracted features are more discriminative for the recognition tasks. Our approach benefits from a recent discovery on the closed-form solutions to noiseless LatLRR. When there is noise, a robust Principal Component Analysis (PCA)-based denoising step can be added as preprocessing. When the scale of a problem is large, we utilize a fast randomized algorithm to speed up the computation of robust PCA. Extensive experimental results demonstrate the effectiveness and robustness of our method. Pan Zhou 0002, Zhouchen Lin, Chao Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Exact Recoverability of Robust PCA via Outlier Pursuit with Tight Recovery BoundsabstractSubspace recovery from noisy or even corrupted data is critical for various applications in machine learning and data analysis. To detect outliers, Robust PCA (R PCA) via Outlier Pursuit was proposed and had found many successful applications. However, the current theoretical analysis on Outlier Pursuit only shows that it succeeds when the sparsity of the corruption matrix is of O(n/r), where n is the number of the samples and r is the rank of the intrinsic matrix which may be comparable to n. Moreover, the regularization parameter is suggested as 3/(7 squareroot gamma n}, where gamma is a parameter that is not known a priori. In this paper, with incoherence condition and proposed ambiguity condition we prove that Outlier Pursuit succeeds when the rank of the intrinsic matrix is of O(n log n) and the sparsity of the corruption matrix is of O(n). We further show that the orders of both bounds are tight. Thus R-PCA via Outlier Pursuit is able to recover intrinsic matrix of higher rank and identify much denser corruptions than what the existing results could predict. Moreover, we suggest that the regularization parameter be chosen as 1 squareroot{log n}, which is definite. Our analysis waives the necessity of tuning the regularization parameter and also significantly extends the working range of the Outlier Pursuit. Experiments on synthetic and real data verify our theories. Hongyang Zhang 0001, Zhouchen Lin, Chao Zhang 0001, Edward Y. Chang |
AAAI | 3 |
| 2015 | Manifold-Regularized Selectable Factor Extraction for Semi-supervised Image Classification
Chao Zhang 0001, Fangyun Wei, Hongyang Zhang 0001, Yiyuan She |
BMVC | 2 |
| 2015 | Generalised sampling theory with rational sampling factorsabstractThe authors consider the problem of reconstructing a signal from the outputs of several systems, which are sampled with periods related by rational factors. They present a unified framework in terms of which generalised sampling theories presented by previous researchers can be discussed. Using the spectral aliasing matrix (SAM), the modelling of generalised sampling can be studied in the cases of both bandlimited and non‐bandlimited input signals. In comparison with the generalised sampling theories by Papoulis, Brown and some other scientists, which were based on the same sampling rate, they extend the theory by introducing the SAM and formulating sampling in an elegant matrix representation and considering generalised sampling of different sampling rates with rational factors. Several solution methods of the derived system of linear equations are presented. These methods can also be applied in both the band‐limited and the non‐band‐limited input signal cases, and the easiest reconstruction implementation can be chosen in the temporal or the frequency domain. Finally, they present some experimental results with a one‐dimensional signal designed to test the performance of the proposed algorithm for various sampling rates and under different noise conditions. Chao Zhang 0001, Pengwei Hao |
IET Signal Process. | 1 |
| 2015 | Relations Among Some Low-Rank Subspace Recovery ModelsabstractRecovering intrinsic low-dimensional subspaces from data distributed on them is a key preprocessing step to many applications. In recent years, a lot of work has modeled subspace recovery as low-rank minimization problems. We find that some representative models, such as robust principal component analysis (R-PCA), robust low-rank representation (R-LRR), and robust latent low-rank representation (R-LatLRR), are actually deeply connected. More specifically, we discover that once a solution to one of the models is obtained, we can obtain the solutions to other models in closed-form formulations. Since R-PCA is the simplest, our discovery makes it the center of low-rank subspace recovery models. Our work has two important implications. First, R-PCA has a solid theoretical foundation. Under certain conditions, we could find globally optimal solutions to these low-rank models at an overwhelming probability, although these models are nonconvex. Second, we can obtain significantly faster algorithms for these models by solving R-PCA first. The computation cost can be further cut by applying low-complexity randomized algorithms, for example, our novel l2,1 filtering algorithm, to R-PCA. Although for the moment the formal proof of our l2,1 filtering algorithm is not yet available, experiments verify the advantages of our algorithm over other state-of-the-art methods based on the alternating direction method. Hongyang Zhang 0001, Zhouchen Lin, Chao Zhang 0001, Junbin Gao |
Neural Comput. | 3 |
| 2015 | Learning to Detect Anomalies in Surveillance VideoabstractDetecting anomalies in surveillance videos, that is, finding events or objects with low probability of occurrence, is a practical and challenging research topic in computer vision community. In this paper, we put forward a novel unsupervised learning framework for anomaly detection. At feature level, we propose a Sparse Semi-nonnegative Matrix Factorization (SSMF) to learn local patterns at each pixel, and a Histogram of Nonnegative Coefficients (HNC) can be constructed as local feature which is more expressive than previously used features like Histogram of Oriented Gradients (HOG). At model level, we learn a probability model which takes the spatial and temporal contextual information into consideration. Our framework is totally unsupervised requiring no human-labeled training data. With more expressive features and more complicated model, our framework can accurately detect and localize anomalies in surveillance video. We carried out extensive experiments on several benchmark video datasets for anomaly detection, and the results demonstrate the superiority of our framework to state-of-the-art approaches, validating the effectiveness of our framework. Tan Xiao, Chao Zhang 0001, Hongbin Zha |
IEEE Signal Process. Lett. | 2 |
| 2014 | Anomaly Detection via Local Coordinate Factorization and Spatio-Temporal Pyramid
Tan Xiao, Chao Zhang 0001, Hongbin Zha, Fangyun Wei |
ACCV (5) | 2 |
| 2014 | Robust latent low rank representation for subspace clustering
Hongyang Zhang 0001, Zhouchen Lin, Chao Zhang 0001, Junbin Gao |
Neurocomputing | 3 |
| 2013 | Robust real-time attention-based head-shoulder detection for video surveillanceabstractRobust head-shoulder detection has widely been utilized in video surveillance applications. However, most state-of-the-art approaches are very time-consuming and unable to effectively handle videos with radical pose and viewpoint changes. In this paper, we introduce a robust and rapid head-shoulder detection method, which is invariant to pose and viewpoint for video surveillance. The proposed method combines an attention-based foreground segmentation module and a multiview head-shoulder detection cascade to achieve high performance in both accuracy and speed. Specifically, the attention-based foreground segmentation module firstly detects not only active regions that involve motions but also static areas where people stand or sit still in video frames. These regions are then provided to a two-layer head-shoulder detection cascade which is composed of a preliminary linear classifier which eliminates most obvious non-head-shoulder windows rapidly, and a set of histogram intersection kernel (HIK) based multi-view classifiers that can precisely detect head-shoulders with different pose and viewpoint. When compared to the leading methods in the challenging PETS 2009 benchmark dataset, our approach obtains highly competitive results in terms of effectiveness and efficiency. Jinhui Tu, Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2013 | A Counterexample for the Validity of Using Nuclear Norm as a Convex Surrogate of Rank
Hongyang Zhang 0001, Zhouchen Lin, Chao Zhang 0001 |
ECML/PKDD (2) | 3 |
| 2013 | A comparison of typical ℓp minimization algorithms
Qin Lyu, Zhouchen Lin, Yiyuan She, Chao Zhang 0001 |
Neurocomputing | 4 |
| 2012 | Online anomaly detection in videos by clustering dynamic exemplarsabstractWe propose a non-parametric hierarchical event model to perform online anomaly detection in videos. A dynamic exemplar set is first used to represent observed event samples which updates itself every time when a new sample comes in. Upon this set, clusters are extracted to summarize the exemplars, offering a compact yet informative data structure for past event samples. Abnormal events are detected by both considering their dissimilarity with the model and low frequency. Experiments on real world crowd surveillance videos demonstrate the effectiveness and robustness of the proposed algorithm which shows reliable detection rates and low false alarms. Jie Feng 0012, Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2012 | Spatial consistency based selective reranking for content based object retrieval
Tan Xiao, Chao Zhang 0001, Hongbin Zha |
ICPR | 2 |
| 2011 | Salient object detection by compositionabstractConventional saliency analysis methods measure the saliency of individual pixels. The resulting saliency map inevitably loses information in the original image and finding salient objects in it is difficult. We propose to detect salient objects by directly measuring the saliency of an image window in the original image and adopt the well established sliding window based object detection paradigm. We present a simple definition for window saliency, i.e., the cost of composing the window using the remaining parts of the image. The definition uses the entire image as the context and agrees with human intuition. It no longer relies on idealistic assumptions usually used before (e.g., "back- ground is homogenous") and generalizes well to complex objects and backgrounds in real world images. To realize the definition, we illustrate how to incorporate different cues such as appearance, position, and size. Based on a segment-based representation, the window composition cost function can be efficiently evaluated by a greedy optimization algorithm. Extensive evaluation on challenging object detection datasets verifies better efficacy and efficiency of the proposed method comparing to the state-of-the-art, making it a good pre-processing tool for subsequent applications. Moreover, we hope to stimulate further work towards the challenging yet important problem of generic salient object detection. Jie Feng 0012, Litian Tao, Chao Zhang 0001, Jian Sun 0001 |
ICCV | 4 |
| 2011 | New color filter arrays of high light sensitivity and high demosaicking performanceabstractFor high light sensitivity, new CFA designs use panchromatic pixels, aka white pixels, that no visible spectrum energy is filtered. Kodak's CFA2.0 has 50% white pixels, but the demosaicking performance is not good. We present in this work a set of new color filter arrays (CFA) of high light sensitivity and high demosaicking performance which were obtained by using a CFA design methodology in the frequency domain. The new patterns are of size 5×5 and come from the same frequency structure, which has one luma in the base band at (0, 0) and four chromas (two conjugate pairs) placed at (4π/5, 2π/5), (- 4π/5, - 2π/5), (2π/5, - 4π/5) and (- 2π/5, 4π/5), respectively. The new patterns are optimized to have only white (panchromatic) and three primary color pixels and the pixels are found to be 40% white, 20% red, 20% green and 20% blue by pixel color constrained optimization. Our demosaicking experiments show that our new CFA patterns outperform Kodak CFA2.0 in both objective and subjective quality. Jue Wang 0004, Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2010 | A blind video watermark detection method based on 3D-DWT transformabstractIn this paper, we propose a new blind video watermark detection method which is based on 3D-DWT transform. We find that the coefficients in the high frequency band of temporal wavelet transform (TWT) are nearly orthogonal to the normally-distributed watermark, which is the basis of our proposed extraction method. The watermarking procedure is similar to the former 3D-DWT method, first 2D-DWT in the spatial domain, then 1D-DWT in the temporal domain (TWT). The difference is that we do not use the low frequency band in the 1D-DWT procedure, for the property of the TWT high frequency band that we have found in the detection procedure. We also propose to use two zero-mean normally-distributed watermarks in embedding to avoid block effects. Finally, an absolute value detection scheme is proposed. Using this scheme our scheme can resist frame dropping attack and temporal frame averaging attack, because the TWT high frequency band reflects the changing parts along the temporal axis. Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2010 | Online Learning with Self-Organizing Maps for Anomaly Detection in Crowd ScenesabstractDetecting abnormal behaviors in crowd scenes is quite important for public security and has been paid more and more attentions. Most previous methods use offline trained model to perform detection which can't handle the constantly changing crowd environment. In this paper, we propose a novel unsupervised algorithm to detect abnormal behavior patterns in crowd scenes with online learning. The crowd behavior pattern is extracted from the local spatio-temporal volume which consists of multiple motion patterns in temporal order. An online self-organizing map (SOM) is used to model the large number of behavior patterns in crowd. Each neuron can be updated by incrementally learning the new observations. To demonstrate the effectiveness of our proposed method, we have performed experiments on real-world crowd scenes. The online learning can efficiently reduce the false alarms while still be able to detect most of the anomalies. Jie Feng 0012, Chao Zhang 0001, Pengwei Hao |
ICPR | 2 |
| 2009 | Clustering-Based Descriptors for Fingerprint Indexing and Fast Retrieval
Shihua He, Chao Zhang 0001, Pengwei Hao |
ACCV (1) | 2 |
| 2009 | Heavy-Tailed Model for Visual Tracking via Robust Subspace Learning
Daojing Wang, Chao Zhang 0001, Pengwei Hao |
ACCV (2) | 2 |
| 2009 | Comparative study of features for fingerprint indexingabstractFor current fingerprint indexing schemes, global textures and minutiae structures are usually utilized. To extend the existing methods of feature extraction, we study the three most popular local descriptors, SIFT, SURF and DAISY, for fingerprint indexing and give a comparison of indexing performance for evaluation of these three features on public fingerprint databases. For index construction, the locality-sensitive hashing (LSH) is used to efficiently retrieve similarity queries in a small fraction of the database. Experiments show that SURF and DAISY are applicable for fingerprint indexing as SURF features perform equally well or better than SIFT features while DAISY improves not so significantly. Shihua He, Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2008 | Transferring Colours to Grayscale Images by Locally Linear EmbeddingabstractIn this paper, we propose a learning-based method for adding colours to grayscale images. In contrast to many previous computer-aided colourizing methods, which require intensive and accurate human intervention, our method needs only the user to provide a colourful image of the similar content as the grayscale image. We accept the “image manifold ” assumption and apply manifold learning methods to model the relations between the chromatic channels and the gray levels in the training images. Then we synthesize the objective chromatic channels using the learned relations. Experiments show that our method gives superior results to those of the previous work. 1 Jun Li 0010, Pengwei Hao, Chao Zhang 0001 |
BMVC | 3 |
| 2008 | Hallucinating faces from thermal infrared imagesabstractThis paper addresses the face hallucination problem of converting thermal infrared face images into photo-realistic ones. It is a challenging task because the two modalities are of dramatical difference, which makes many developed linear models inapplicable. We propose a learning-based framework synthesizing the normal face from the infrared input. Compared to the previous work, we further exploit the local linearity in not only the image spatial domain but also the image manifolds. We have also developed a measurement of the variance between an input and its prediction, thus we can apply the Markov random field model to the predicted normal face to improve the hallucination result. Experimental results show the advantage of our algorithm over the existing methods. Our algorithm can be readily generalized to solve other multi-modal image conversion problems as well. Jun Li 0010, Pengwei Hao, Chao Zhang 0001, Mingsong Dou |
ICIP | 3 |
| 2008 | Fingerprint indexing based on composite set of reduced SIFT featuresabstractMost of current fingerprint indexing schemes utilize features based on global textures and minutiae structures. To extend the existing technology of feature extraction, this paper proposes a new fingerprint indexing and retrieval scheme using scale invariant feature transformation (SIFT), which has been widely used in generic image retrieval. With slight loss in effectiveness, we reduce the number of features generated from one fingerprint for efficiency. To cope with the uncertainty of acquisition (e.g. partialness, distortion), we use a composite set of features to form multiple impressions for the fingerprint representation. In the index construction phase, the use of locality-sensitive hashing (LSH) allows us to perform similarity queries by only examining a small fraction of the database. Experiments on database FVC2000 and FVC2002 show the effectiveness of our proposed scheme. Xin Shuai, Chao Zhang 0001, Pengwei Hao |
ICPR | 2 |
| 2008 | Multivariate Laplace Filter: A heavy-tailed model for target trackingabstractVideo-based target tracking is a challenging task, because there always appears to be complex occlusion among the varying number of objects. Also, in practice, it is very common that the objects in a scene move irregularly with abrupt turns, which results in an interesting heavy-tailed phenomenon. As simulation has to run exceptionally long enough to capture the effect of the distribution tail, it is arduous to simulate heavy-tailed distribution. In this paper, we propose a new view to target tracking from a heavy-tailed perspective, establishing a simple but novel Multivariate Laplace Filter (MLF) tracking model, which efficiently and accurately describes the heavy-tailed issue and dramatically surmounts it. Some experimental results show the good performance of the proposed method. Daojing Wang, Chao Zhang 0001, Xuemin Zhao |
ICPR | 2 |
| 2007 | Converting Thermal Infrared Face Images into Normal Gray-Level Images
Mingsong Dou, Chao Zhang 0001, Pengwei Hao, Jun Li 0010 |
ACCV (2) | 2 |
| 2007 | Progressive Reversible Data Hiding by Symmetrical Histogram Expansion with Piecewise-Linear Haar TransformabstractIn this paper, we present a progressive reversible data hiding technique by symmetrical histogram expansion in the transform domain of piecewise-linear Haar (PLHaar). Data are embedded into the PLHaar coefficients of images progressively from the pivotal bin of a histogram of PLHaar coefficients to both sides of the pivot symmetrically. With PLHaar, no overflow or underflow occurs to the pixel values, and our data hiding method achieves the highest embedding capacity with PSNR around 50 dB compared to the previous methods in the literature. The method can also be applied to artificial images with an exactly flat histogram, which is impossible for a method to hide data in the spatial domain. The progressiveness of the proposed method also enables a rough but automatic capacity-PSNR control. The effectiveness of our method is demonstrated with a number of experiments. Lei Yang 0041, Pengwei Hao, Chao Zhang 0001 |
ICASSP (2) | 3 |
| 2007 | The Optimal ROS-Based Symmetric Phase-Only Filter for Fingerprint VerificationabstractSymmetric phase-only filter (SPOF) has been widely applied to image registration and recognition, and has been proved efficient for fingerprint verification. Fingerprint images have a characteristic that the dominant information is concentrated in an elliptic frequency band of their ridges in low frequency domain. However, the existent fingerprint verification methods based on SPOF do not take this characteristic into account. To improve the performance of SPOF for fingerprint recognition, an appropriate region of support (ROS) can be used to set the least significant frequency region to zero. By means of theoretical and experimental analysis, we have found the optimal ROS for SPOF that achieves the best discrimination power among all the possible ROSs. Experiments show that our optimal ROS-based SPOF is more efficient than the method of band-limited SPOF (BLPOC). Xin Shuai, Chao Zhang 0001, Pengwei Hao |
ICIP (2) | 2 |
| 2006 | Fingerprint Indexing Based on LAS RegistrationabstractFingerprint indexing is an efficient technique that greatly improves the performance of fingerprint based person authentication systems by reducing the number of comparison. In this paper, we propose an indexing method based on fingerprint registration with a novel feature called local axial symmetry (LAS). The location and direction estimation of reference point are achieved in a straightforward way after the LAS field is achieved. Then the registered orientation field is utilized as a feature vector to perform the following indexing. A new scheme of the experiment is introduced and satisfactory experimental results are achieved on FVC2000 DB2 that the average search space is only 2.34% of all fingers in the condition of equal-sized training set and testing set. Tong Liu 0018, Chao Zhang 0001, Pengwei Hao |
ICIP | 2 |
| 2005 | Fingerprint indexing based on singular point correlationabstractFingerprint indexing is an efficient technique that greatly improves the performance of automated fingerprint identification systems. We propose a continuous fingerprint indexing method based on location, direction estimation and correlation of fingerprint singular points. Location and direction estimation are achieved simultaneously by applying a T-shape model to directional field of fingerprint images. The T-shape model analyzes homocentric sectors around the candidate singular points to find lateral-axes and further main-axes. Then a distortion-tolerant filter of minimum average correlation energy is utilized to obtain a correlation-based similarity measure which gives the evidence of searching priority. The experiment is performed by 400-fingerprint retrieval from 10,000 templates and the mean search space is only 3.46% of the whole dataset. Tong Liu 0018, Guocai Zhu, Chao Zhang 0001, Pengwei Hao |
ICIP (3) | 3 |
| 2005 | Shear-resize factorizations for fast image registrationabstractOwing to its effectiveness and simplicity, intensity-based method works well for registration of images. However, it needs a large amount of computation for geometric transformation. In this paper, we present two shear-resize matrix factorizations to accelerate the transformation. A transform matrix can be factorized into two shears and a fixed non-uniform resize, or three shears and a customizable resize. A customizable resize can be uniform in all dimensions or scaling just in one dimension. Shears can be implemented very fast by memory-shift, and a resize can be done by simple axis-aligned interpolation. The factorizations can be applied to both rigid-body and affine transformations. Their efficiency is performed by experiments on some standard test images and fingerprint images. The methods are quite promising for hardware implementation, and can also be extended to 3D or higher dimensional fast geometric transformation. Chen Ying, Pengwei Hao, Chao Zhang 0001 |
ICIP (3) | 3 |