VLDB 2026 Research / reviewers in the wild / expert
Liqing Zhang 0001
dblp:20/4627-1
· DBLP profile ↗
221ranked-venue papers
7as first author
77since 2021 · last 2026
0000-0001-7597-8503ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 155 · 5 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 118 · 2 first-author · 56 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Quality Full-Head 3D Avatar Generation from Any Single Portrait ImageabstractIn this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modeling fine-grained facial textures and maintaining identity information. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring high-quality facial texture details. To further improve performance, we propose a novel multi-view diffusion named ID-TS diffusion model, which integrate identity and expression information into the two-stage multi-view diffusion process. The low-resolution stage ensures structural consistency of heads across multiple views, while the high-resolution stage preserves facial detail fidelity and coherence. Finally, we propose an enhanced feed-forward Gaussian avatar reconstruction method that optimizes the network on multi-view images of each single subject, significantly improving 3D facial texture details. Extensive experiments show that our method demonstrates robust performance across challenging scenarios, while showcasing broad applicability across numerous downstream tasks. Yujie Gao 0001, Chencheng Wang, Xianbing Sun, Jiahui Zhan, Wentao Wang 0009, Yiyi Zhang 0002, Haohua Zhao 0001, Liqing Zhang 0001, Jianfu Zhang 0003 |
AAAI | 8 |
| 2026 | HM-NVS: Hierarchical Multi-Modal Novel View Synthesis with Uncertainty-Aware Progressive RefinementabstractNovel view synthesis (NVS) from a single image is fundamental for multimedia applications including virtual/augmented reality and interactive 3D content retrieval. Despite notable progress, diffusion-based methods suffer from a persistent semantic gap: global textual conditioning lacks the granularity to preserve part-level attributes and inter-part spatial relationships under large viewpoint changes. We present Hierarchical Multi-Modal Novel View Synthesis (HM-NVS), a framework that bridges this gap through three synergistic contributions: (1) a hierarchical semantic extraction and adaptive fusion mechanism that integrates global, part-level, and geometric cues with view-dependent weighting; (2) a semantic-augmented 3D attention module that jointly enforces geometric and semantic consistency across generated views; and (3) an uncertainty-aware progressive refinement strategy that selectively allocates computation to difficult regions. Experiments on GSO, RTMV, and OmniObject3D demonstrate that HM-NVS improves PSNR by 15% and reduces LPIPS by 20% over state-of-the-art methods, with robust generalisation to non-rigid and articulated objects. Haohua Zhao 0001, Liqing Zhang 0001 |
ICMR | 3 |
| 2026 | Beyond Semantic Understanding: Physics-Informed Adaptation of Video Foundation Models for Precision Industrial MetrologyabstractVideo foundation models (VFMs) pretrained on internet-scale data have demonstrated remarkable generalization in semantic segmentation and object tracking. However, we identify a critical limitation when these models are deployed for quantitative industrial analysis: semantic understanding does not entail metrological precision. A VFM may correctly segment a rotating mechanical component while producing kinematic measurements that violate rigid-body constraints—a deficiency we term the semantic-precision gap. Through controlled experiments, we establish that this gap is structural: it persists under stronger backbones (SAM 3), parameter-efficient fine-tuning (LoRA, Adapters), and state-of-the-art video object segmentation (Cutie, DEVA), because none introduce the physical inductive biases necessary for geometric and kinematic reasoning. Haohua Zhao 0001, Liqing Zhang 0001 |
ICMR | 3 |
| 2026 | Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question AnsweringabstractVideo Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts. Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2025 | Hyperspectral Image Classification Based on Local Low-Rank RepresentationabstractHyperspectral image classification is an attractive and challenging task due to the difficulty to acquire large labeled datasets, and its susceptibility to natural environmental influences. Deep learning-based methods can effectively enhance the efficiency of hyperspectral feature learning. Tensor-based approaches can effectively help preserving the inherent structural information. This paper introduces a novel tensor-based framework for hyperspectral image classification. It leverages the ‘bot-tleneck’ structure of an AutoEncoder to extract low-dimensional representations in a self-supervised manner. Additionally, local low-rank constraints in the embedding space facilitate the distribution of features across a low-rank manifold. Experiments conducted on real hyperspectral datasets demonstrate that the proposed method yields superior classification performance. Jianting Wang, Haohua Zhao 0001, Jianfu Zhang 0003, Liqing Zhang 0001 |
CSCWD | 4 |
| 2025 | Warp Gait Across Ages: Cross-age Gait Video Translation with Part-aware Flow WarpingabstractCross-age gait video translation aims to translate a gait video of a subject captured at a certain age into another age while preserving the individual identity and realism. This has a wide range of applications, including age-invariant gait recognition. In this paper, we propose a method for cross-age gait video translation using a spatially smooth geometric warping field. More specifically, we employ a variant of the spatial transformer network (STN), called flow warping, which achieves accurate coarse-to-fine warping field inference. In addition, we extend the flow warping framework by introducing gait stance-dependent warping fields to better represent the temporal variations in the warping fields. We also develop an approach called Part-STN to infer part-dependent warping fields and then merge them into a unified warping field to retain consistency within each body part. We then train Part-STN with loss functions considering three aspects: aging effect (using an age group classification loss), identity preservation (using a cycle-consistency loss and triplet loss on the representation learned from a pretrained gait recognition model), and realism (using an adversarial loss and smoothness loss on estimated flow warping). Our framework outperforms state-of-the-art cross-age gait video translation methods quantitatively and qualitatively on the largest gait database OULP-Age, with respect to both age group classification, identity recognition and realism. Yiyi Zhang 0002, Hanchong Yan, Liqing Zhang 0001, Yasushi Yagi |
IJCB | 3 |
| 2025 | Rethinking Classifier Re-Training in Long-Tailed Recognition: Label Over-Smooth Can BalanceabstractIn the field of long-tailed recognition, the Decoupled Training paradigm has shown exceptional promise by dividing training into two stages: representation learning and classifier re-training. While previous work has tried to improve both stages simultaneously, this complicates isolating the effect of classifier re-training. Recent studies reveal that simple regularization can produce strong feature representations, highlighting the need to reassess classifier re-training methods. In this study, we revisit classifier re-training methods based on a unified feature representation and re-evaluate their performances.
We propose two new metrics, Logits Magnitude and Regularized Standard Deviation, to compare the differences and similarities between various methods.
Using these two newly proposed metrics, we demonstrate that when the Logits Magnitude across classes is nearly balanced, further reducing its overall value can effectively decrease errors and disturbances during training, leading to better model performance.
Based on our analysis using these metrics, we observe that adjusting the logits could improve model performance, leading us to develop a simple label over-smoothing approach to adjust the logits without requiring prior knowledge of class distribution.
This method softens the original one-hot labels by assigning a probability slightly higher than $\frac{1}{K}$ to the true class and slightly lower than $\frac{1}{K}$ to the other classes, where $K$ is the number of classes.
Our method achieves state-of-the-art performance on various imbalanced datasets, including CIFAR100-LT, ImageNet-LT, and iNaturalist2018. Han Lu 0004, Jiangtong Li, Yichen Xie 0002, Tianjiao Li 0001, Xiaokang Yang 0001, Liqing Zhang 0001, Junchi Yan |
ICLR | 7 |
| 2025 | Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and ModalitiesabstractReconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods. Yiyi Zhang 0002, Xinhao Hu, Li Niu 0002, Jianfu Zhang 0003, Yasushi Makihara, Yasushi Yagi, Wenlong Liao, Junchi Yan, Liqing Zhang 0001 |
ICLR | 12 |
| 2025 | Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringabstractVideo Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition. Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001 |
ICML | 11 |
| 2025 | Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsabstractProgress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. Yikun Ji, Yan Hong 0001, Jiahui Zhan, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Liqing Zhang 0001, Jianfu Zhang 0003 |
ACM Multimedia | 8 |
| 2025 | ID-MotionNet: Identity-Preserved 3D Skeleton Sequence Generation via Information Bottleneck Disentanglement
Jingyu Xie, Jianfu Zhang 0003, Liqing Zhang 0001, Yiyi Zhang 0002 |
PRCV (15) | 5 |
| 2025 | Defending adversarial attacks in Graph Neural Networks via tensor enhancement
Jianfu Zhang 0003, Yan Hong 0001, Dawei Cheng, Liqing Zhang 0001, Qibin Zhao |
Pattern Recognit. | 4 |
| 2024 | WeditGAN: Few-Shot Image Generation via Latent Space RelocationabstractIn few-shot image generation, directly training GAN models on just a handful of images faces the risk of overfitting. A popular solution is to transfer the models pretrained on large source domains to small target ones. In this work, we introduce WeditGAN, which realizes model transfer by editing the intermediate latent codes w in StyleGANs with learned constant offsets (delta w), discovering and constructing target latent spaces via simply relocating the distribution of source latent spaces. The established one-to-one mapping between latent spaces can naturally prevents mode collapse and overfitting. Besides, we also propose variants of WeditGAN to further enhance the relocation process by regularizing the direction or finetuning the intensity of delta w. Experiments on a collection of widely used source/target datasets manifest the capability of WeditGAN in generating realistic and diverse images, which is simple yet highly effective in the research area of few-shot image generation. Codes are available at https://github.com/Ldhlwh/WeditGAN. Yuxuan Duan, Li Niu 0002, Yan Hong 0001, Liqing Zhang 0001 |
AAAI | 4 |
| 2024 | Painterly Image Harmonization by Learning from Painterly ObjectsabstractGiven a composite image with photographic object and painterly background, painterly image harmonization targets at stylizing the composite object to be compatible with the background. Despite the competitive performance of existing painterly harmonization works, they did not fully leverage the painterly objects in artistic paintings. In this work, we explore learning from painterly objects for painterly image harmonization. In particular, we learn a mapping from background style and object information to object style based on painterly objects in artistic paintings. With the learnt mapping, we can hallucinate the target style of composite object, which is used to harmonize encoder feature maps to produce the harmonized image. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our proposed method. Li Niu 0002, Junyan Cao, Yan Hong 0001, Liqing Zhang 0001 |
AAAI | 4 |
| 2024 | Progressive Painterly Image Harmonization from Low-Level Styles to High-Level StylesabstractPainterly image harmonization aims to harmonize a photographic foreground object on the painterly background. Different from previous auto-encoder based harmonization networks, we develop a progressive multi-stage harmonization network, which harmonizes the composite foreground from low-level styles (e.g., color, simple texture) to high-level styles (e.g., complex texture). Our network has better interpretability and harmonization performance. Moreover, we design an early-exit strategy to automatically decide the proper stage to exit, which can skip the unnecessary and even harmful late stages. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our progressive harmonization network. Li Niu 0002, Yan Hong 0001, Junyan Cao, Liqing Zhang 0001 |
AAAI | 4 |
| 2024 | Assessing Image Inpainting via Re-Inpainting Self-Consistency EvaluationabstractImage inpainting, the task of reconstructing missing segments in corrupted images using available data, faces challenges in ensuring consistency and fidelity, especially under information-scarce conditions. Traditional evaluation methods, heavily dependent on the existence of unmasked reference images, inherently favor certain inpainting outcomes, introducing biases. Addressing this issue, we introduce an innovative evaluation paradigm that utilizes a self-supervised metric based on multiple re-inpainting passes. This approach, diverging from conventional reliance on direct comparisons in pixel or feature space with original images, emphasizes the principle of self-consistency to enable the exploration of various viable inpainting solutions, effectively reducing biases. Our extensive experiments across numerous benchmarks validate the alignment of our evaluation method with human judgment. Jianfu Zhang 0003, Yan Hong 0001, Yiyi Zhang 0002, Liqing Zhang 0001 |
CIKM | 5 |
| 2024 | Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringabstractDespite the recent progress made in Video Question-Answering (VideoQA), these methods typically function as black-boxes, making it difficult to understand their reasoning processes and perform consistent compositional reasoning. To address these challenges, we propose a model-agnostic Video Alignment and Answer Aggregation (VA3) framework, which is capable of enhancing both compositional consistency and accuracy of existing VidQA methods by integrating video aligner and answer aggregator modules. The video aligner hierarchically selects the relevant video clips based on the question, while the answer ag-gregator deduces the answer to the question based on its sub-questions, with compositional consistency ensured by the information flow along question decomposition graph and the contrastive learning strategy. We evaluate our framework on three settings of the AGQA-Decomp dataset with three baseline methods, and propose new metrics to measure the compositional consistency of VidQA methods more comprehensively. Moreover, we propose a large language model (LLM) based automatic question decomposition pipeline to apply our framework to any VidQA dataset. We extend MSVD and NExT-QA datasets with it to evaluate our VA3framework on broader scenarios. Extensive experiments show that our framework improves both compositional consistency and accuracy of existing methods, leading to more interpretable real-world VidQA models. Zhaohe Liao, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 4 |
| 2024 | COIN-Matting: Confounder Intervention for Image Matting
Zhaohe Liao, Jiangtong Li, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Li Niu 0002, Liqing Zhang 0001 |
ECCV (19) | 7 |
| 2024 | Hierarchical Attacks on Large-Scale Graph Neural NetworksabstractIn this paper, we present a novel hierarchical approach to adversarial attacks targeting Graph Neural Networks (GNNs), tailored to overcome the complexities inherent in large-scale poisoning attacks. Traditional global attack strategies often fail to yield effective results on extensive graph structures. Our innovative method implements a divide-and-conquer tactic, clustering nodes based on their embeddings and forming coarse-grained graphs from these clusters. We initiate perturbations at this coarse level, gradually honing them in more detailed, finer-grained graphs, while keeping non-essential nodes grouped. By employing meta-gradients derived from these refined graphs, we pinpoint critical edges for perturbation, thereby vastly simplifying the process and reducing the intricacy involved in manipulating large-scale graphs. This hierarchical strategy not only enhances the efficacy of the attacks but also maintains operational efficiency across expansive network structures. Jianfu Zhang 0003, Yan Hong 0001, Dawei Cheng, Liqing Zhang 0001, Qibin Zhao |
ICASSP | 4 |
| 2024 | DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric FinetuningabstractThe recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a specific domain either hard to describe or just unseen to the models. In this work, we propose DomainGallery, a few-shot domain-driven image generation method which aims at finetuning pretrained Stable Diffusion on few-shot target datasets in an attribute-centric manner. Specifically, DomainGallery features prior attribute erasure, attribute disentanglement, regularization and enhancement. These techniques are tailored to few-shot domain-driven generation in order to solve key issues that previous works have failed to settle. Extensive experiments are given to validate the superior performance of DomainGallery on a variety of domain-driven generation scenarios. Yuxuan Duan, Yan Hong 0001, Bo Zhang 0075, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 9 |
| 2024 | Painterly Image Harmonization via Adversarial Residual LearningabstractImage compositing plays a vital role in photo editing. After inserting a foreground object into another background image, the composite image may look unnatural and inharmonious. When the foreground is photorealistic and the background is an artistic painting, painterly image harmonization aims to transfer the style of background painting to the foreground object, which is a challenging task due to the large domain gap between foreground and background. In this work, we employ adversarial learning to bridge the domain gap between foreground feature map and background feature map. Specifically, we design a dual-encoder generator, in which the residual encoder produces the residual features added to the foreground feature map from main encoder. Then, a pixel-wise discriminator plays against the generator, encouraging the refined foreground feature map to be indistinguishable from background feature map. Extensive experiments demonstrate that our method could achieve more harmonious and visually appealing results than previous methods. Xudong Wang 0001, Li Niu 0002, Junyan Cao, Yan Hong 0001, Liqing Zhang 0001 |
WACV | 5 |
| 2024 | Multi-Model UNet: An Adversarial Defense Mechanism for Robust Visual TrackingabstractAbstract Currently, state-of-the-art object-tracking algorithms are facing a severe threat from adversarial attacks, which can significantly undermine their performance. In this research, we introduce MUNet, a novel defensive model designed for visual tracking. This model is capable of generating defensive images that can effectively counter attacks while maintaining a low computational overhead. To achieve this, we experiment with various configurations of MUNet models, finding that even a minimal three-layer setup significantly improves tracking robustness when the target tracker is under attack. Each model undergoes end-to-end training on randomly paired images, which include both clean and adversarial noise images. This training separately utilizes pixel-wise denoiser and feature-wise defender. Our proposed models significantly enhance tracking performance even when the target tracker is attacked or the target frame is clean. Additionally, MUNet can simultaneously share its parameters on both template and search regions. In experimental results, the proposed models successfully defend against top attackers on six benchmark datasets, including OTB100, LaSOT, UAV123, VOT2018, VOT2019, and GOT-10k. Performance results on all datasets show a significant improvement over all attackers, with a decline of less than 4.6% for every benchmark metric compared to the original tracker. Notably, our model demonstrates the ability to enhance tracking robustness in other blackbox trackers. Wattanapong Suttapak, Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
Neural Process. Lett. | 4 |
| 2023 | Amodal Instance Segmentation via Prior-Guided ExpansionabstractAmodal instance segmentation aims to infer the amodal mask, including both the visible part and occluded part of each object instance. Predicting the occluded parts is challenging. Existing methods often produce incomplete amodal boxes and amodal masks, probably due to lacking visual evidences to expand the boxes and masks. To this end, we propose a prior-guided expansion framework, which builds on a two-stage segmentation model (i.e., Mask R-CNN) and performs box-level (resp., pixel-level) expansion for amodal box (resp., mask) prediction, by retrieving regression (resp., flow) transformations from a memory bank of expansion prior. We conduct extensive experiments on KINS, D2SA, and COCOA cls datasets, which show the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
AAAI | 6 |
| 2023 | Few-Shot Defect Image Generation via Defect-Aware Feature ManipulationabstractThe performances of defect inspection have been severely hindered by insufficient defect images in industries, which can be alleviated by generating more samples as data augmentation. We propose the first defect image generation method in the challenging few-shot cases. Given just a handful of defect images and relatively more defect-free ones, our goal is to augment the dataset with new defect images. Our method consists of two training stages. First, we train a data-efficient StyleGAN2 on defect-free images as the backbone. Second, we attach defect-aware residual blocks to the backbone, which learn to produce reasonable defect masks and accordingly manipulate the features within the masked regions by training the added modules on limited defect images. Extensive experiments on MVTec AD dataset not only validate the effectiveness of our method in generating realistic and diverse defect images, but also manifest the benefits it brings to downstream defect inspection tasks. Codes are available at https://github.com/Ldhlwh/DFMGAN. Yuxuan Duan, Yan Hong 0001, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 4 |
| 2023 | Isometric Manifold Learning Using Hierarchical FlowabstractWe propose the Hierarchical Flow (HF) model constrained by isometric regularizations for manifold learning that combines manifold learning goals such as dimensionality reduction, inference, sampling, projection and density estimation into one unified framework. Our proposed HF model is regularized to not only produce embeddings preserving the geometric structure of the manifold, but also project samples onto the manifold in a manner conforming to the rigorous definition of projection. Theoretical guarantees are provided for our HF model to satisfy the two desired properties. In order to detect the real dimensionality of the manifold, we also propose a two-stage dimensionality reduction algorithm, which is a time-efficient algorithm thanks to the hierarchical architecture design of our HF model. Experimental results justify our theoretical analysis, demonstrate the superiority of our dimensionality reduction algorithm in cost of training time, and verify the effect of the aforementioned properties in improving performances on downstream tasks such as anomaly detection. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 4 |
| 2023 | Geometric Inductive Biases for Identifiable Unsupervised Learning of Disentangled Representations
Li Niu 0002, Liqing Zhang 0001 |
AAAI | 3 |
| 2023 | Image Cropping with Spatial-aware Feature and Rank ConsistencyabstractImage cropping aims to find visually appealing crops in an image. Despite the great progress made by previous methods, they are weak in capturing the spatial relationship between crops and aesthetic elements (e.g., salient objects, semantic edges). Besides, due to the high annotation cost of labeled data, the potential of unlabeled data awaits to be excavated. To address the first issue, we propose spatial-aware feature to encode the spatial relationship between candidate crops and aesthetic elements, by feeding the concatenation of crop mask and selectively aggregated feature maps to a light-weighted encoder. To address the second issue, we train a pair-wise ranking classifier on labeled images and transfer such knowledge to unlabeled images to enforce rank consistency. Experimental results on the benchmark datasets show that our proposed method performs favorably against state-of-the-art methods. Li Niu 0002, Bo Zhang 0075, Liqing Zhang 0001 |
CVPR | 4 |
| 2023 | Deep Image Harmonization with Learnable AugmentationabstractThe goal of image harmonization is adjusting the foreground appearance in a composite image to make the whole image harmonious. To construct paired training images, existing datasets adopt different ways to adjust the illumination statistics of foregrounds of real images to produce synthetic composite images. However, different datasets have considerable domain gap and the performances on small-scale datasets are limited by insufficient training data. In this work, we explore learnable augmentation to enrich the illumination diversity of small-scale datasets for better harmonization performance. In particular, our designed SYthetic COmposite Network (SycoNet) takes in a real image with foreground mask and a random vector to learn suitable color transformation, which is applied to the foreground of this real image to produce a synthetic composite image. Comprehensive experiments demonstrate the effectiveness of our proposed learnable augmentation for image harmonization. The code of SycoNet is released at https://github.com/bcmi/SycoNet-Adaptive-Image-Harmonization. Li Niu 0002, Junyan Cao, Wenyan Cong, Liqing Zhang 0001 |
ICCV | 4 |
| 2023 | Knowledge Proxy Intervention for Deconfounded Video Question AnsweringabstractRecently, Video Question-Answering (VideoQA) has drawn more and more attention from both the industry and the research community. Despite all the success achieved by recent works, dataset bias always harmfully misleads current methods focusing on spurious correlations in training data. To analyze the effects of dataset bias, we frame the VideoQA pipeline into a causal graph, which shows the causalities among video, question, aligned feature between video and question, answer, and underlying confounder. Through the causal graph, we prove that the confounder and the backdoor path lead to spurious causality. To tackle the challenge that the confounder in VideoQA is unobserved and non-enumerable in general, we propose a model-agnostic framework called Knowledge Proxy Intervention (KPI), which introduces an extra knowledge proxy variable in the causal graph to cut the backdoor path and remove the effect of confounder. Our KPI framework exploits the front-door adjustment, which requires no prior knowledge about the confounder. The effectiveness of our KPI framework is corroborated by three baseline methods on five benchmark datasets, including MSVD-QA, MSRVTT-QA, TGIF-QA, NExT-QA, and Causal-VidQA. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ICCV | 3 |
| 2023 | Deep Image Harmonization with Globally Guided Feature Transformation and Relation DistillationabstractGiven a composite image, image harmonization aims to adjust the foreground illumination to be consistent with background. Previous methods have explored transforming foreground features to achieve competitive performance. In this work, we show that using global information to guide foreground feature transformation could achieve significant improvement. Besides, we propose to transfer the foreground-background relation from real images to composite images, which can provide intermediate supervision for the transformed encoder features. Additionally, considering the drawbacks of existing harmonization datasets, we also contribute a ccHarmony dataset which simulates the natural illumination variation. Extensive experiments on iHarmony4 and our contributed dataset demonstrate the superiority of our method. Our ccHarmony dataset is released at https://github.com/bcmi/Image-HarmonizationDataset-ccHarmony. Li Niu 0002, Linfeng Tan, Xinhao Tao, Junyan Cao, Fengjun Guo, Liqing Zhang 0001 |
ICCV | 7 |
| 2023 | Fine-grained Visible Watermark RemovalabstractVisible watermark removal aims to erase the watermark from watermarked image and recover the background image, which is a challenging task due to the diverse watermarks. Previous works have designed dynamic network to handle various types of watermarks adaptively, but they ignore that even the watermarked region in a single image can be divided into multiple local parts with distinct visual appearances. In this work, we advance image-specific dynamic network towards part-specific dynamic network, which discovers multiple local parts within the watermarked region and handle them adaptively. Specifically, we propose a query-based multi-task framework, in which part query embeddings are jointly used in two branches to predict part masks and restore watermarked parts. Extensive experiments demonstrate the effectiveness of our fine-grained watermark removal network. Li Niu 0002, Xing Zhao 0010, Bo Zhang 0075, Liqing Zhang 0001 |
ICCV | 4 |
| 2023 | Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowabstractVirtual try-on is a critical image synthesis task that aims to transfer clothes from one image to another while preserving the details of both humans and clothes. While many existing methods rely on Generative Adversarial Networks (GANs) to achieve this, flaws can still occur, particularly at high resolutions. Recently, the diffusion model has emerged as a promising alternative for generating high-quality images in various applications. However, simply using clothes as a condition for guiding the diffusion model to inpaint is insufficient to maintain the details of the clothes. To overcome this challenge, we propose an exemplar-based inpainting approach that leverages a warping module to guide the diffusion model's generation effectively. The warping module performs initial processing on the clothes, which helps to preserve the local details of the clothes. We then combine the warped clothes with clothes-agnostic person image and add noise as the input of diffusion model. Additionally, the warped clothes is used as local conditions for each denoising process to ensure that the resulting output retains as much detail as possible. Our approach, namely Diffusion-based Conditional Inpainting for Virtual Try-ON(DCI-VTON), effectively utilizes the power of the diffusion model, and the incorporation of the warping module helps to produce high-quality and realistic virtual try-on results. Experimental results on VITON-HD demonstrate the effectiveness and superiority of our method. Source code and trained models will be publicly released at: https://github.com/bcmi/DCI-VTON-Virtual-Try-On. Junhong Gou, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 6 |
| 2023 | Painterly Image Harmonization using Diffusion ModelabstractPainterly image harmonization aims to insert photographic objects into paintings and obtain artistically coherent composite images. Previous methods for this task mainly rely on inference optimization or generative adversarial network, but they are either very time-consuming or struggling at fine control of the foreground objects (e.g., texture and content details). To address these issues, we propose a novel Painterly Harmonization stable Diffusion model (PHDiffusion), which includes a lightweight adaptive encoder and a Dual Encoder Fusion (DEF) module. Specifically, the adaptive encoder and the DEF module first stylize foreground features within each encoder. Then, the stylized foreground features from both encoders are combined to guide the harmonization process. During training, besides the noise loss in diffusion model, we additionally employ content loss and two style losses, i.e., AdaIN style loss and contrastive style loss, aiming to balance the trade-off between style migration and content preservation. Compared with the state-of-the-art models from related fields, our PHDiffusion can stylize the foreground more sufficiently and simultaneously retain finer content. Our code and model are available at https://github.com/bcmi/PHDiffusion-Painterly-Image-Harmonization Lingxiao Lu, Jiangtong Li, Junyan Cao, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 5 |
| 2023 | Deep Image Harmonization in Dual Color SpacesabstractImage harmonization is an essential step in image composition that adjusts the appearance of composite foreground to address the inconsistency between foreground and background. Existing methods primarily operate in correlated RGB color space, leading to entangled features and limited representation ability. In contrast, decorrelated color space (e.g., Lab) has decorrelated channels that provide disentangled color and illumination statistics. In this paper, we explore image harmonization in dual color spaces, which supplements entangled RGB features with disentangled L, a, b features to alleviate the workload in harmonization process. The network comprises a RGB harmonization backbone, an Lab encoding module, and an Lab control module. The backbone is a U-Net network translating composite image to harmonized image. Three encoders in Lab encoding module extract three control codes independently from L, a, b channels, which are used to manipulate the decoder features in harmonization backbone via Lab control module. Our code and model are available at https://github.com/bcmi/DucoNet-Image-Harmonization. Linfeng Tan, Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2023 | Natural Image Matting with Attended Global Context
Yiyi Zhang 0002, Li Niu 0002, Yasushi Makihara, Jianfu Zhang 0003, Weijie Zhao 0003, Yasushi Yagi, Liqing Zhang 0001 |
J. Comput. Sci. Technol. | 7 |
| 2023 | Diverse image inpainting with disentangled uncertainty
Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Haoyu Ling, Liqing Zhang 0001 |
Pattern Recognit. | 7 |
| 2023 | Delinquent Events Prediction in Temporal Networked-Guarantee LoansabstractUnder debt obligation promises, small- and medium-sized enterprises (SMEs) can guarantee each other to enhance their financial security to get loans from commercial banks. When the economy rises, the banks may reduce the threshold to some extent, which may introduce default risk during the economy down period, especially when many SMEs bind together and form complex networks. The risk may diffuse across the guarantee network and may result in a financial crisis. Macroprudential oversight of the guarantee network to eliminate any potential systematics financial risk is the central task of the regulatory commission and the commercial banks. Based on our observation, the delinquent probability of an SME depends not only on self-financial status but also highly related to its temporal behaviors and structural position in networks. The classic approach for loan assessment criteria face challenges in extracting temporal and structural patterns from dynamic networks. To address these issues, we propose a temporal delinquent event prediction (TDEP) framework that preserves temporal network structures and credit behavior sequences in an end-to-end model. In particular, we first employ a graph attention layer to learn the representation of nodes in temporal guarantee networks. We then design a recursive and self-attention mechanism to integrate both credit behavior and network structure information. The learned attentional weights are leveraged to uncover high-risk guarantee patterns that effectively accelerate the risk assessment process. Afterward, we conduct extensive experiments in a real-world guaranteed-loan data set to evaluate its performance. The results show the effectiveness of our proposed approach compared with the state-of-the-art baselines. Finally, we integrate the proposed model in a real-world loan risk management system. We present the implementation details of each subcomponent of the system and report out the performance after online deployment. Dawei Cheng, Zhibin Niu, Liqing Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | From Pixel to Patch: Synthesize Context-Aware Features for Zero-Shot Semantic SegmentationabstractZero-shot learning (ZSL) has been actively studied for image classification tasks to relieve the burden of annotating image labels. Interestingly, the semantic segmentation task requires more labor-intensive pixel-wise annotation, but zero-shot semantic segmentation has not attracted extensive research interest. Thus, we focus on zero-shot semantic segmentation that aims to segment unseen objects with only category-level semantic representations provided for unseen categories. In this article, we propose a novel context-aware feature generation network (CaGNet) that can synthesize context-aware pixel-wise visual features for unseen categories based on category-level semantic representations and pixel-wise contextual information. The synthesized features are used to fine-tune the classifier to enable segmenting of unseen objects. Furthermore, we extend pixel-wise feature generation and fine-tuning to patch-wise feature generation and fine-tuning, which additionally considers the interpixel relationship. Experimental results on Pascal-VOC, Pascal-context, and COCO-stuff show that our method significantly outperforms the existing zero-shot semantic segmentation methods. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Action-Aware Embedding Enhancement for Image-Text RetrievalabstractImage-text retrieval plays a central role in bridging vision and language, which aims to reduce the semantic discrepancy between images and texts. Most of existing works rely on refined words and objects representation through the data-oriented method to capture the word-object cooccurrence. Such approaches are prone to ignore the asymmetric action relation between images and texts, that is, the text has explicit action representation (i.e., verb phrase) while the image only contains implicit action information. In this paper, we propose Action-aware Memory-Enhanced embedding (AME) method for image-text retrieval, which aims to emphasize the action information when mapping the images and texts into a shared embedding space. Specifically, we integrate action prediction along with an action-aware memory bank to enrich the image and text features with action-similar text features. The effectiveness of our proposed AME method is verified by comprehensive experimental results on two benchmark datasets. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 3 |
| 2022 | Deep Image Harmonization by Bridging the Reality Gap
Junyan Cao, Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
BMVC | 5 |
| 2022 | Inharmonious Region Localization via Recurrent Self-Reasoning
Penghao Wu, Li Niu 0002, Jing Liang 0007, Liqing Zhang 0001 |
BMVC | 4 |
| 2022 | Inharmonious Region Localization with Auxiliary Style Feature
Penghao Wu, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 3 |
| 2022 | Visible Watermark Removal with Dynamic Kernel and Semantic-aware Propagation
Xing Zhao 0010, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 3 |
| 2022 | Weak-shot Semantic Segmentation by Transferring Semantic Affinity and Boundary
Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
BMVC | 5 |
| 2022 | High-Resolution Image Harmonization via Collaborative Dual TransformationsabstractGiven a composite image, image harmonization aims to adjust the foreground to make it compatible with the background. High-resolution image harmonization is in high demand, but still remains unexplored. Conventional image harmonization methods learn global RGB-to-RGB transformation which could effortlessly scale to high resolution, but ignore diverse local context. Recent deep learning methods learn the dense pixel-to-pixel transformation which could generate harmonious outputs, but are highly constrained in low resolution. In this work, we propose a high-resolution image harmonization network with Collaborative Dual Transformation (CDTNet) to combine pixel-to-pixel transformation and RGB-to-RGB transformation coherently in an end-to-end network. Our CDTNet consists of a low-resolution generator for pixel-to-pixel transformation, a color mapping module for RGB-to-RGB transformation, and a refinement module to take advantage of both. Extensive experiments on high-resolution bench-mark dataset and our created high-resolution real composite images demonstrate that our CDTNet strikes a good balance between efficiency and effectiveness. Our used datasets can be found in https://github.com/bcmi/CDTNet-High-Resolution-Image-Harmonization. Wenyan Cong, Xinhao Tao, Li Niu 0002, Jing Liang 0007, Xuesong Gao, Qihao Sun, Liqing Zhang 0001 |
CVPR | 7 |
| 2022 | XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingabstractRecently, various multimodal networks for Visually-Rich Document Understanding(VRDU) have been proposed, showing the promotion of transformers by integrating visual and layout information with the text embeddings. However, most existing approaches utilize the position embeddings to incorporate the sequence information, neglecting the noisy improper reading order obtained by OCR tools. In this paper, we propose a robust layout-aware multimodal network named XYLayoutLM to capture and leverage rich layout information from proper reading orders produced by our Augmented XY Cut. Moreover, a Dilated Conditional Position Encoding module is proposed to deal with the input sequence of variable lengths, and it additionally extracts local layout information from both textual and vi-sual modalities while generating position embeddings. Experiment results show that our XYLayoutLM achieves competitive results on document understanding tasks. Zhangxuan Gu, Changhua Meng, Ke Wang 0042, Jun Lan 0001, Weiqiang Wang 0002, Ming Gu 0011, Liqing Zhang 0001 |
CVPR | 7 |
| 2022 | From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringabstractVideo understanding has achieved great success in representation learning, such as video caption, video object grounding, and video descriptive question-answer. However, current methods still struggle on video reasoning, including evidence reasoning and commonsense reasoning. To facilitate deeper video understanding towards video reasoning, we present the task of Causal-VidQA, which includes four types of questions ranging from scene description (description) to evidence reasoning (explanation) and commonsense reasoning (prediction and counterfactual). For commonsense reasoning, we set up a two-step solution by answering the question and providing a proper reason. Through extensive experiments on existing VideoQA methods, we find that the state-of-the-art methods are strong in descriptions but weak in reasoning. We hope that Causal-VidQA can guide the research of video understanding from representation learning to deeper reasoning. The dataset and related resources are available at https://github.com/bcmi/Causal-VidQA.git. Jiangtong Li, Li Niu 0002, Liqing Zhang 0001 |
CVPR | 3 |
| 2022 | Dual-path Image Inpainting with Auxiliary GAN InversionabstractDeep image inpainting can inpaint a corrupted image using a feed-forward inference, but still fails to handle large missing area or complex semantics. Recently, GAN inversion based inpainting methods propose to leverage semantic information in pretrained generator (e.g., StyleGAN) to solve the above issues. Different from feed-forward methods, they seek for a closest latent code to the corrupted image and feed it to a pretrained generator. However, inferring the latent code is either time-consuming or inaccurate. In this paper, we develop a dual-path inpainting network with inversion path and feed-forward path, in which inversion path provides auxiliary information to help feed-forward path. We also design a novel deformable fusion module to align the feature maps in two paths. Experiments on FFHQ and LSUN demonstrate that our method is effective in solving the aforementioned problems while producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Xue Yang 0005, Liqing Zhang 0001 |
CVPR | 5 |
| 2022 | DeltaGAN: Towards Diverse Few-Shot Image Generation with Sample-Specific Delta
Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ECCV (16) | 4 |
| 2022 | Human-Centric Image Cropping with Partition-Aware and Content-Preserving Features
Bo Zhang 0075, Li Niu 0002, Xing Zhao 0010, Liqing Zhang 0001 |
ECCV (7) | 4 |
| 2022 | Learning Object Placement via Dual-Path Graph Completion
Liu Liu 0022, Li Niu 0002, Liqing Zhang 0001 |
ECCV (17) | 4 |
| 2022 | Deep Video Harmonization With Color Mapping ConsistencyabstractVideo harmonization aims to adjust the foreground of a composite video to make it compatible with the background. So far, video harmonization has only received limited attention and there is no public dataset for video harmonization. In this work, we construct a new video harmonization dataset HYouTube by adjusting the foreground of real videos to create synthetic composite videos. Moreover, we consider the temporal consistency in video harmonization task. Unlike previous works which establish the spatial correspondence, we design a novel framework based on the assumption of color mapping consistency, which leverages the color mapping of neighboring frames to refine the current frame. Extensive experiments on our HYouTube dataset prove the effectiveness of our proposed framework. Our dataset and code are available at https://github.com/bcmi/Video-Harmonization-Dataset-HYouTube. Xinyuan Lu, Shengyuan Huang, Li Niu 0002, Wenyan Cong, Liqing Zhang 0001 |
IJCAI | 5 |
| 2022 | Few-shot Image Generation Using Discrete Content RepresentationabstractFew-shot image generation and few-shot image translation are two related tasks, both of which aim to generate new images for an unseen category with only a few images. In this work, we make the first attempt to adapt few-shot image translation method to few-shot image generation task. Few-shot image translation disentangles an image into style vector and content map. An unseen style vector can be combined with different seen content maps to produce different images. However, it needs to store seen images to provide content maps and the unseen style vector may be incompatible with seen content maps. To adapt it to few-shot image generation task, we learn a compact dictionary of local content vectors via quantizing continuous content maps into discrete content maps instead of storing seen images. Furthermore, we model the autoregressive distribution of discrete content map conditioned on style vector, which can alleviate the incompatibility between content map and style vector. Qualitative and quantitative results on three real datasets demonstrate that our model can produce images of higher diversity and fidelity for unseen categories than previous methods. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2022 | Weak-shot Semantic Segmentation via Dual Similarity TransferabstractSemantic segmentation is a practical and active task, but severely suffers from the expensive cost of pixel-level labels when extending to more classes in wider applications. To this end, we focus on the problem named weak-shot semantic segmentation, where the novel classes are learnt from cheaper image-level labels with the support of base classes having off-the-shelf pixel-level labels. To tackle this problem, we propose a dual similarity transfer framework, which is built upon MaskFormer to disentangle the semantic segmentation task into single-label classification and binary segmentation for each proposal. Specifically, the binary segmentation sub-task allows proposal-pixel similarity transfer from base classes to novel classes, which enables the mask learning of novel classes. We also learn pixel-pixel similarity from base classes and distill such class-agnostic semantic similarity to the semantic masks of novel classes, which regularizes the segmentation model with pixel-level semantic relationship across images. In addition, we propose a complementary loss to facilitate the learning of novel classes. Comprehensive experiments on the challenging COCO-Stuff-10K and ADE20K datasets demonstrate the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
NeurIPS | 6 |
| 2022 | UniGAN: Reducing Mode Collapse in GANs using a Uniform GeneratorabstractDespite the significant progress that has been made in the training of Generative Adversarial Networks (GANs), the mode collapse problem remains a major challenge in training GANs, which refers to a lack of diversity in generative samples. In this paper, we propose a new type of generative diversity named uniform diversity, which relates to a newly proposed type of mode collapse named $u$-mode collapse where the generative samples distribute nonuniformly over the data manifold. From a geometric perspective, we show that the uniform diversity is closely related with the generator uniformity property, and the maximum uniform diversity is achieved if the generator is uniform. To learn a uniform generator, we propose UniGAN, a generative framework with a Normalizing Flow based generator and a simple yet sample efficient generator uniformity regularization, which can be easily adapted to any other generative framework. A new type of diversity metric named udiv is also proposed to estimate the uniform diversity given a set of generative samples in practice. Experimental results verify the effectiveness of our UniGAN in learning a uniform generator and improving uniform diversity. Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 3 |
| 2022 | Zero-shot sketch-based image retrieval with structure-aware asymmetric disentanglement
Jiangtong Li, Zhixin Ling, Li Niu 0002, Liqing Zhang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2022 | Hallucinating uncertain motion and future for static image action recognition
Li Niu 0002, Shengyuan Huang, Xing Zhao 0010, Liwei Kang, Yiyi Zhang 0002, Liqing Zhang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2022 | Diminishing-feature attack: The adversarial infiltration on visual tracking
Wattanapong Suttapak, Jianfu Zhang 0003, Liqing Zhang 0001 |
Neurocomputing | 3 |
| 2022 | Graph Neural Network for Fraud Detection via Spatial-Temporal AttentionabstractCard fraud is an important issue and incurs a considerable cost for both cardholders and issuing banks. Contemporary methods apply machine learning-based approaches to detect fraudulent behavior from transaction records. But manually generating features needs domain knowledge and may lay behind the modus operandi of fraud, which means we need to automatically focus on the most relevant fraudulent behavior patterns in the online detection system. Therefore, in this work, we propose a spatial-temporal attention-based graph network (STAGN) for credit card fraud detection. In particular, we learn the temporal and location-based transaction graph features by a graph neural network first. Afterwards, we employ the spatial-temporal attention on top of learned tensor representations, which are then fed into a 3D convolution network. The attentional weights are jointly learned in an end-to-end manner with 3D convolution and detection networks. After that, we conduct extensive experiments on the real-word card transaction dataset. The result shows that STAGN performs better than other state-of-the-art baselines in both AUC and precision-recall curves. Moreover, we conduct empirical studies with domain experts on the proposed method for fraud detection and knowledge discovery; the result demonstrates its superiority in detecting suspicious transactions, mining spatial and temporal fraud hotspots, and uncover fraud patterns. The effectiveness of the proposed method in other user behavior-based tasks is also demonstrated. Finally, in order to tackle the challenges of big data, we integrate our proposed STAGN into the fraud detection system as the predictive model and present the implementation detail of each module in the system. Dawei Cheng, Xiaoyang Wang 0002, Ying Zhang 0001, Liqing Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Activity Image-to-Video Retrieval by Disentangling Appearance and MotionabstractWith the rapid emergence of video data, image-to-video retrieval has attracted much attention. There are two types of image-to-video retrieval: instance-based and activity-based. The former task aims to retrieve videos containing the same main objects as the query image, while the latter focuses on finding the similar activity. Since dynamic information plays a significant role in the video, we pay attention to the latter task to explore the motion relation between images and videos. In this paper, we propose a Motion-assisted Activity Proposal-based Image-to-Video Retrieval (MAP-IVR) approach to disentangle the video features into motion features and appearance features and obtain appearance features from the images. Then, we perform image-to-video translation to improve the disentanglement quality. The retrieval is performed in both appearance and video feature spaces. Extensive experiments demonstrate that our MAP-IVR approach remarkably outperforms the state-of-the-art approaches on two benchmark activity-based video datasets. Liu Liu 0022, Jiangtong Li, Li Niu 0002, Ruicong Xu, Liqing Zhang 0001 |
AAAI | 5 |
| 2021 | Disentangled Information BottleneckabstractThe information bottleneck (IB) method is a technique for extracting information that is relevant for predicting the target random variable from the source random variable, which is typically implemented by optimizing the IB Lagrangian that balances the compression and prediction terms. However, the IB Lagrangian is hard to optimize, and multiple trials for tuning values of Lagrangian multiplier are required. Moreover, we show that the prediction performance strictly decreases as the compression gets stronger during optimizing the IB Lagrangian. In this paper, we implement the IB method from the perspective of supervised disentangling. Specifically, we introduce Disentangled Information Bottleneck (DisenIB) that is consistent on compressing source maximally without target prediction performance loss (maximum compression). Theoretical and experimental results demonstrate that our method is consistent on maximum compression, and performs well in terms of generalization, robustness to adversarial attack, out-of-distribution detection, and supervised disentangling. Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 4 |
| 2021 | Depth Privileged Object Detection in Indoor Scenes via Deformation HallucinationabstractRGB-D object detection has achieved significant advance, because depth provides complementary geometric information to RGB images. Considering depth images are unavailable in some scenarios, we focus on depth privileged object detection in indoor scenes, where the depth images are only available in the training phase. Under this setting, one prevalent research line is modality hallucination, in which depth image and depth feature are the common choices for hallucinating. In contrast, we choose to hallucinate depth deformation, which is explicit geometric information and efficient to hallucinate. Specifically, we employ the deformable convolution layer with augmented offsets as our deformation module and regard the offsets as geometric deformation, because the offsets enable flexibly sampling over the object and transforming to a canonical shape for ease of detection. In addition, we design a quality-based mechanism to avoid negative transfer of depth deformation. Experimental results and analyses on NYUDv2 and SUN RGB-D demonstrate the effectiveness of our method against the state-of-the-art methods for depth privileged object detection. Junjie Chen 0008, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 5 |
| 2021 | Image Composition Assessment with Saliency-augmented Multi-pattern Pooling
Bo Zhang 0075, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 3 |
| 2021 | Tensor Decomposition Via Core Tensor NetworksabstractTensor decomposition (TD) has shown promising performance in image completion and denoising. Existing methods always aim to decompose one tensor into latent factors or core tensors by optimizing a particular cost function based on a specific tensor model. These algorithms iteratively learn the optima from random initialization given any individual tensor, resulting in slow convergence and low efficiency. In this paper, we propose an efficient TD algorithm that aims to learn a global mapping from input tensors to latent core tensors, under the assumption that the mappings of multiple tensors might be shared or highly correlated. To this end, we train a deep neural network (DNN) to model the global mapping and then apply it to decompose a newly given tensor with high efficiency. Furthermore, the initial values of DNN are learned based on meta-learning methods. By leveraging the pretrained core tensor DNN, our proposed method enables us to perform TD efficiently and accurately. Experimental results demonstrate the significant improvements of our method over other TD methods in terms of speed and accuracy. Jianfu Zhang 0003, Zerui Tao, Liqing Zhang 0001, Qibin Zhao |
ICASSP | 3 |
| 2021 | Parallel Multi-Resolution Fusion Network for Image InpaintingabstractConventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow layers of the auto-encoder architecture can not be well integrated. Differing from the conventional image inpainting architecture, we design a parallel multi-resolution inpainting network with multi-resolution partial convolution, in which low-resolution branches focus on the global structure while high-resolution branches focus on the local texture details. All these high- and low-resolution streams are in parallel and fused repeatedly with multi-resolution masked representation fusion so that the reconstructed images are semantically robust and textually plausible. Experimental results show that our method can effectively fuse structure and texture information, producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Jianfu Zhang 0003, Li Niu 0002, Haoyu Ling, Xue Yang 0005, Liqing Zhang 0001 |
ICCV | 6 |
| 2021 | Inharmonious Region LocalizationabstractThe advance of image editing techniques allows users to create artistic works, but the manipulated regions may be incompatible with the background. Localizing the inharmonious region is an appealing yet challenging task. Realizing that this task requires effective aggregation of multi-scale contextual information and suppression of redundant information, we design novel Bi-directional Feature Integration (BFI) block and Global-context Guided Decoder (GGD) block to fuse multi-scale features in the encoder and decoder respectively. We also employ Mask-guided Dual Attention (MDA) block between the encoder and decoder to suppress the redundant information. Experiments on the image harmonization dataset demonstrate that our method achieves competitive performance for inharmonious region localization. The source code is available at https://github.com/bcmi/DIRL. Jing Liang 0007, Li Niu 0002, Liqing Zhang 0001 |
ICME | 3 |
| 2021 | Bargainnet: Background-Guided Domain Translation for Image HarmonizationabstractGiven a composite image with inharmonious foreground and background, image harmonization aims to adjust the foreground to make it compatible with the background. Previous image harmonization methods mainly focus on learning the mapping from composite image to real image, while ignoring the crucial guidance role that background plays. In this work, we formulate image harmonization task as background-guided domain translation. Specifically, we use a domain code extractor to capture the background domain information to guide the foreground harmonization, which is regulated by well-tailored triplet losses. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our proposed method. Code is available at https://github.com/bcmi/BargainNet. Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Jing Liang 0007, Liqing Zhang 0001 |
ICME | 5 |
| 2021 | Static Image Action Recognition with Hallucinated Fine-Grained Motion InformationabstractStatic image action recognition is a challenging task due to the lack of motion information in a static image. Some previous works have attempted to hallucinate the motion information in a static image using a generator learnt from freely available unlabeled videos. However, their hallucinated motion information is either low-level or coarse-grained, which may contain lots of noise or lose motion details. In contrast, we propose to hallucinate fine-grained high-level motion information, which is more robust and detail-preserving. Specifically, we hallucinate motion feature map which encodes the motion details of human body parts. We also hallucinate motion attention map to focus on motion-related regions. Our hallucinated motion information can greatly facilitate static image action recognition, which is confirmed by the experiments on two static action image datasets and two video datasets. Shengyuan Huang, Xing Zhao 0010, Li Niu 0002, Liqing Zhang 0001 |
ICME | 4 |
| 2021 | End-to-End Video Object Detection with Spatial-Temporal TransformersabstractRecently, DETR and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors. However, their performance on Video Object Detection (VOD) has not been well explored. In this paper, we present TransVOD, an end-to-end video object detection model based on a spatial-temporal Transformer architecture. The goal of this paper is to streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks. Besides, benefited from the object query design in DETR, our method does not need complicated post-processing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. In particular, we present temporal Transformer to aggregate both the spatial object queries and the feature memories of each frame. Our temporal Transformer consists of three components: Temporal Deformable Transformer Encoder (TDTE) to encode the multiple frame spatial details, Temporal Query Encoder (TQE) to fuse object queries, and Temporal Deformable Transformer Decoder (TDTD) to obtain current frame detection results. These designs boost the strong baseline deformable DETR by a significant margin (3%-4% mAP) on the ImageNet VID dataset. TransVOD yields comparable results performance on the benchmark of ImageNet VID. We hope our TransVOD can provide a new perspective for video object detection. Qianyu Zhou 0001, Xiangtai Li, Li Niu 0002, Yunhai Tong, Lizhuang Ma, Liqing Zhang 0001 |
ACM Multimedia | 10 |
| 2021 | Video Semantic Segmentation via Sparse Temporal TransformerabstractCurrently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed. Jiangtong Li, Wentao Wang 0009, Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 7 |
| 2021 | Visible Watermark Removal via Self-calibrated Localization and Background RefinementabstractSuperimposing visible watermarks on images provides a powerful weapon to cope with the copyright issue. Watermark removal techniques, which can strengthen the robustness of visible watermarks in an adversarial way, have attracted increasing research interest. Modern watermark removal methods perform watermark localization and background restoration simultaneously, which could be viewed as a multi-task learning problem. However, existing approaches suffer from incomplete detected watermark and degraded texture quality of restored background. Therefore, we design a two-stage multi-task network to address the above issues. The coarse stage consists of a watermark branch and a background branch, in which the watermark branch self-calibrates the roughly estimated mask and passes the calibrated mask to background branch to reconstruct the watermarked area. In the refinement stage, we integrate multi-level features to improve the texture quality of watermarked area. Extensive experiments on two datasets demonstrate the effectiveness of our proposed method. Jing Liang 0007, Li Niu 0002, Fengjun Guo, Liqing Zhang 0001 |
ACM Multimedia | 5 |
| 2021 | Weak-shot Fine-grained Classification via Similarity TransferabstractRecognizing fine-grained categories remains a challenging task, due to the subtle distinctions among different subordinate categories, which results in the need of abundant annotated samples. To alleviate the data-hungry problem, we consider the problem of learning novel categories from web data with the support of a clean set of base categories, which is referred to as weak-shot learning. In this setting, we propose a method called SimTrans to transfer pairwise semantic similarity from base categories to novel categories. Specifically, we firstly train a similarity net on clean data, and then leverage the transferred similarity to denoise web training data using two simple yet effective strategies. In addition, we apply adversarial loss on similarity net to enhance the transferability of similarity. Comprehensive experiments demonstrate the effectiveness of our weak-shot setting and our SimTrans method. Junjie Chen 0008, Li Niu 0002, Liu Liu 0022, Liqing Zhang 0001 |
NeurIPS | 4 |
| 2021 | Mixed Supervised Object Detection by Transferring Mask Prior and Semantic SimilarityabstractObject detection has achieved promising success, but requires large-scale fully-annotated data, which is time-consuming and labor-extensive. Therefore, we consider object detection with mixed supervision, which learns novel object categories using weak annotations with the help of full annotations of existing base object categories. Previous works using mixed supervision mainly learn the class-agnostic objectness from fully-annotated categories, which can be transferred to upgrade the weak annotations to pseudo full annotations for novel categories. In this paper, we further transfer mask prior and semantic similarity to bridge the gap between novel categories and base categories. Specifically, the ability of using mask prior to help detect objects is learned from base categories and transferred to novel categories. Moreover, the semantic similarity between objects learned from base categories is transferred to denoise the pseudo full annotations for novel categories. Experimental results on three benchmark datasets demonstrate the effectiveness of our method over existing methods. Codes are available at https://github.com/bcmi/TraMaS-Weak-Shot-Object-Detection. Li Niu 0002, Junjie Chen 0008, Liqing Zhang 0001 |
NeurIPS | 5 |
| 2021 | Depth Privileged Scene Recognition via Dual Attention HallucinationabstractRGB-D scene recognition has achieved promising performance because depth could provide complementary geometric information to RGB images. However, the inaccessibility of depth sensors severely limits RGB-D applications. In this paper, we focus on depth privileged setting, in which depth information is only available during training but not available during testing. Considering that the information obtained from RGB and depth images are complementary while attention is informative and transferable, our idea is using RGB input to hallucinate depth attention. We build our model upon modulated deformable convolutional layer and hallucinate dual attention: post-hoc importance weight and trainable spatial transformation. Specifically, we use modulation (resp., offset) learned from RGB to mimic Grad-CAM (resp., offset) learned from depth, to combine the strength of dual attention. We also design a weighted loss to avoid negative transfer according to the quality of depth attention. Extensive experiments on two benchmarks, i.e., SUN RGB-D and NYUDv2, demonstrate that our method outperforms the state-of-the-art methods for depth privileged scene recognition. Junjie Chen 0008, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Memorize, Associate and Match: Embedding Enhancement via Fine-Grained Alignment for Image-Text RetrievalabstractImage-text retrieval aims to capture the semantic correlation between images and texts. Existing image-text retrieval methods can be roughly categorized into embedding learning paradigm and pair-wise learning paradigm. The former paradigm fails to capture the fine-grained correspondence between images and texts. The latter paradigm achieves fine-grained alignment between regions and words, but the high cost of pair-wise computation leads to slow retrieval speed. In this paper, we propose a novel method named MEMBER by using Memory-based EMBedding Enhancement for image-text Retrieval (MEMBER), which introduces global memory banks to enable fine-grained alignment and fusion in embedding learning paradigm. Specifically, we enrich image (resp., text) features with relevant text (resp., image) features stored in the text (resp., image) memory bank. In this way, our model not only accomplishes mutual embedding enhancement across two modalities, but also maintains the retrieval efficiency. Extensive experiments demonstrate that our MEMBER remarkably outperforms state-of-the-art approaches on two large-scale benchmark datasets. Jiangtong Li, Liu Liu 0022, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Person Re-Identification With Reinforced Attribute Attention SelectionabstractPerson re-identification (Re-ID) aims to match pedestrian images across various scenes in video surveillance. There are a few works using attribute information to boost Re-ID performance. Specifically, those methods leverage attribute information to boost Re-ID performance by introducing auxiliary tasks like verifying the image level attribute information of two pedestrian images or recognizing identity level attributes. Identity level attribute annotations cost less manpower and are well-fitted for person re-identification task compared with image-level attribute annotations. However, the identity attribute information may be very noisy due to incorrect attribute annotation or lack of discriminativeness to distinguish different persons, which is probably unhelpful for the Re-ID task. In this paper, we propose a novel Attribute Attentional Block (AAB), which can be integrated into any backbone network or framework. Our AAB adopts reinforcement learning to drop noisy attributes based on our designed reward and then utilizes aggregated attribute attention of the remaining attributes to facilitate the Re-ID task. Experimental results demonstrate that our proposed method achieves state-of-the-art results on three benchmark datasets. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Hard Pixel Mining for Depth Privileged Semantic SegmentationabstractSemantic segmentation has achieved remarkable progress but remains challenging due to the complex scene, object occlusion, and so on. Some research works have attempted to use extra information such as a depth map to help RGB based semantic segmentation because the depth map could provide complementary geometric cues. However, due to the inaccessibility of depth sensors, depth information is usually unavailable for the test images. In this paper, we leverage only the depth of training images as the privileged information to mine the hard pixels in semantic segmentation, in which depth information is only available for training images but not available for test images. Specifically, we propose a novel Loss Weight Module, which outputs a loss weight map by employing two depth-related measurements of hard pixels: Depth Prediction Error and Depth-aware Segmentation Error. The loss weight map is then applied to segmentation loss, with the goal of learning a more robust model by paying more attention to the hard pixels. Besides, we also explore a curriculum learning strategy based on the loss weight map. Meanwhile, to fully mine the hard pixels on different scales, we apply our loss weight module to multi-scale side outputs. Our hard pixels mining method achieves the state-of-the-art results on three benchmark datasets, and even outperforms the methods which need depth input during testing. Zhangxuan Gu, Li Niu 0002, Haohua Zhao 0001, Liqing Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Spatio-Temporal Attention-Based Neural Network for Credit Card Fraud DetectionabstractCredit card fraud is an important issue and incurs a considerable cost for both cardholders and issuing institutions. Contemporary methods apply machine learning-based approaches to detect fraudulent behavior from transaction records. But manually generating features needs domain knowledge and may lay behind the modus operandi of fraud, which means we need to automatically focus on the most relevant patterns in fraudulent behavior. Therefore, in this work, we propose a spatial-temporal attention-based neural network (STAN) for fraud detection. In particular, transaction records are modeled by attention and 3D convolution mechanisms by integrating the corresponding information, including spatial and temporal behaviors. Attentional weights are jointly learned in an end-to-end manner with 3D convolution and detection networks. Afterward, we conduct extensive experiments on real-word fraud transaction dataset, the result shows that STAN performs better than other state-of-the-art baselines in both AUC and precision-recall curves. Moreover, we conduct empirical studies with domain experts on the proposed method for fraud post-analysis; the result demonstrates the effectiveness of our proposed method in both detecting suspicious transactions and mining fraud patterns. Dawei Cheng, Sheng Xiang 0001, Chencheng Shang, Yiyi Zhang 0002, Fangzhou Yang, Liqing Zhang 0001 |
AAAI | 6 |
| 2020 | Image Cropping with Composition and Saliency Aware Aesthetic Score MapabstractAesthetic image cropping is a practical but challenging task which aims at finding the best crops with the highest aesthetic quality in an image. Recently, many deep learning methods have been proposed to address this problem, but they did not reveal the intrinsic mechanism of aesthetic evaluation. In this paper, we propose an interpretable image cropping model to unveil the mystery. For each image, we use a fully convolutional network to produce an aesthetic score map, which is shared among all candidate crops during crop-level aesthetic evaluation. Then, we require the aesthetic score map to be both composition-aware and saliency-aware. In particular, the same region is assigned with different aesthetic scores based on its relative positions in different crops. Moreover, a visually salient region is supposed to have more sensitive aesthetic scores so that our network can learn to place salient objects at more proper positions. Such an aesthetic score map can be used to localize aesthetically important regions in an image, which sheds light on the composition rules learned by our model. We show the competitive performance of our model in the image cropping task on several benchmark datasets, and also demonstrate its generality in real-world applications. Li Niu 0002, Weijie Zhao 0003, Dawei Cheng, Liqing Zhang 0001 |
AAAI | 5 |
| 2020 | A Proposal-Based Approach for Activity Image-to-Video RetrievalabstractActivity image-to-video retrieval task aims to retrieve videos containing the similar activity as the query image, which is a challenging task because videos generally have many background segments irrelevant to the activity. In this paper, we utilize R-C3D model to represent a video by a bag of activity proposals, which can filter out background segments to some extent. However, there are still noisy proposals in each bag. Thus, we propose an Activity Proposal-based Image-to-Video Retrieval (APIVR) approach, which incorporates multi-instance learning into cross-modal retrieval framework to address the proposal noise issue. Specifically, we propose a Graph Multi-Instance Learning (GMIL) module with graph convolutional layer, and integrate this module with classification loss, adversarial loss, and triplet loss in our cross-modal retrieval framework. Moreover, we propose geometry-aware triplet loss based on point-to-subspace distance to preserve the structural information of activity proposals. Extensive experiments on three widely-used datasets verify the effectiveness of our approach. Ruicong Xu, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 4 |
| 2020 | Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionabstractStatic image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant unlabeled videos can be economically obtained. Therefore, several works have explored using unlabeled videos to facilitate image action recognition, which can be categorized into the following two groups: (a) enhance visual representations of action images with a designed proxy task on unlabeled videos, which falls into the scope of self-supervised learning; (b) generate auxiliary representations for action images with the generator learned from unlabeled videos. In this paper, we integrate the above two strategies in a unified framework, which consists of Visual Representation Enhancement (VRE) module and Motion Representation Augmentation (MRA) module. Specifically, the VRE module includes a proxy task which imposes pseudo motion label constraint and temporal coherence constraint on unlabeled videos, while the MRA module could predict the motion information of a static action image by exploiting unlabeled videos. We demonstrate the superiority of our framework based on four benchmark human action datasets with limited labeled data. Yiyi Zhang 0002, Li Niu 0002, Meichao Luo, Jianfu Zhang 0003, Dawei Cheng, Liqing Zhang 0001 |
AAAI | 7 |
| 2020 | DoveNet: Deep Image Harmonization via Domain VerificationabstractImage composition is an important operation in image processing, but the inconsistency between foreground and background significantly degrades the quality of composite image. Image harmonization, aiming to make the foreground compatible with the background, is a promising yet challenging task. However, the lack of high-quality publicly available dataset for image harmonization greatly hinders the development of image harmonization techniques. In this work, we contribute an image harmonization dataset iHarmony4 by generating synthesized composite images based on COCO (resp., Adobe5k, Flickr, day2night) dataset, leading to our HCOCO (resp., HAdobe5k, HFlickr, Hday2night) sub-dataset. Moreover, we propose a new deep image harmonization method DoveNet using a novel domain verification discriminator, with the insight that the foreground needs to be translated to the same domain as background. Extensive experiments on our constructed dataset demonstrate the effectiveness of our proposed method. Our dataset and code are available at https://github.com/bcmi/Image_Harmonization_Datasets. Wenyan Cong, Jianfu Zhang 0003, Li Niu 0002, Liu Liu 0022, Zhixin Ling, Weiyuan Li, Liqing Zhang 0001 |
CVPR | 7 |
| 2020 | Learning From Web Data With Self-Organizing Memory ModuleabstractLearning from web data has attracted lots of research interest in recent years. However, crawled web images usually have two types of noises, label noise and background noise, which induce extra difficulties in utilizing them effectively. Most existing methods either rely on human supervision or ignore the background noise. In this paper, we propose a novel method, which is capable of handling these two types of noises together, without the supervision of clean images in the training stage. Particularly, we formulate our method under the framework of multi-instance learning by grouping ROIs (i.e., images and their region proposals) from the same category into bags. ROIs in each bag are assigned with different weights based on the representative/discriminative scores of their nearest clusters, in which the clusters and their scores are obtained via our designed memory module. Our memory module could be naturally integrated with the classification module, leading to an end-to-end trainable system. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our method. Li Niu 0002, Junjie Chen 0008, Dawei Cheng, Liqing Zhang 0001 |
CVPR | 5 |
| 2020 | Beyond Without Forgetting: Multi-Task Learning for Classification with Disjoint DatasetsabstractMulti-task Learning (MTL) for classification with disjoint datasets aims to explore MTL when one task only has one labeled dataset. In existing methods, for each task, the unlabeled datasets are not fully exploited to facilitate this task. Inspired by semi-supervised learning, we use unlabeled datasets with pseudo labels to facilitate each task. However, there are two major issues: 1) the pseudo labels are very noisy; 2) the unlabeled datasets and the labeled dataset for each task has considerable data distribution mismatch. To address these issues, we propose our MTL with Selective Augmentation (MTL-SA) method to select the training samples in unlabeled datasets with confident pseudo labels and close data distribution to the labeled dataset. Then, we use the selected training samples to add information and use the remaining training samples to preserve information. Extensive experiments on face-centric and human-centric applications demonstrate the effectiveness of our MTL-SA method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 4 |
| 2020 | Matchinggan: Matching-Based Few-Shot Image GenerationabstractTo generate new images for a given category, most deep generative models require abundant training images from this category, which are often too expensive to acquire. To achieve the goal of generation based on only a few images, we propose matching-based Generative Adversarial Network (GAN) for few-shot generation, which includes a matching generator and a matching discriminator. Matching generator can match random vectors with a few conditional images from the same category and generate new images for this category based on the fused features. The matching discriminator extends conventional GAN discriminator by matching the feature of generated image with the fused feature of conditional images. Extensive experiments on three datasets demonstrate the effectiveness of our proposed method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 4 |
| 2020 | Risk Guarantee Prediction in Networked-LoansabstractThe guaranteed loan is a debt obligation promise that if one corporation gets trapped in risks, its guarantors will back the loan. When more and more companies involve, they subsequently form complex networks. Detecting and predicting risk guarantee in these networked-loans is important for the loan issuer. Therefore, in this paper, we propose a dynamic graph-based attention neural network for risk guarantee relationship prediction (DGANN). In particular, each guarantee is represented as an edge in dynamic loan networks, while companies are denoted as nodes. We present an attention-based graph neural network to encode the edges that preserve the financial status as well as network structures. The experimental result shows that DGANN could significantly improve the risk prediction accuracy in both the precision and recall compared with state-of-the-art baselines. We also conduct empirical studies to uncover the risk guarantee patterns from the learned attentional network features. The result provides an alternative way for loan risk management, which may inspire more work in the future. Dawei Cheng, Xiaoyang Wang 0002, Ying Zhang 0001, Liqing Zhang 0001 |
IJCAI | 4 |
| 2020 | Context-aware Feature Generation For Zero-shot Semantic SegmentationabstractExisting semantic segmentation models heavily rely on dense pixel-wise annotations. To reduce the annotation pressure, we focus on a challenging task named zero-shot semantic segmentation, which aims to segment unseen objects with zero annotations. This task can be accomplished by transferring knowledge across categories via semantic word embeddings. In this paper, we propose a novel context-aware feature generation method for zero-shot segmentation named CaGNet. In particular, with the observation that a pixel-wise feature highly depends on its contextual information, we insert a contextual module in a segmentation network to capture the pixel-wise contextual information, which guides the process of generating more diverse and context-aware features from semantic word embeddings. Our method achieves state-of-the-art results on three benchmark datasets for zero-shot segmentation. Zhangxuan Gu, Li Niu 0002, Zihan Zhao 0001, Liqing Zhang 0001 |
ACM Multimedia | 5 |
| 2020 | F2GAN: Fusing-and-Filling GAN for Few-shot Image GenerationabstractIn order to generate images for a given category, existing deep generative models generally rely on abundant training images. However, extensive data acquisition is expensive and fast learning ability from limited data is necessarily required in real-world applications. Also, these existing methods are not well-suited for fast adaptation to a new category. Few-shot image generation, aiming to generate images from only a few images for a new category, has attracted some research interest. In this paper, we propose a Fusing-and-Filling Generative Adversarial Network (F2GAN) to generate realistic and diverse images for a new category with only a few images. In our F2GAN, a fusion generator is designed to fuse the high-level features of conditional images with random interpolation coefficients, and then fills in attended low-level details with non-local attention module to produce a new image. Moreover, our discriminator can ensure the diversity of generated images by a mode seeking loss and an interpolation regression loss. Extensive experiments on five datasets demonstrate the effectiveness of our proposed method for few-shot image generation. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 6 |
| 2020 | Adversarial Query-by-Image Video Retrieval Based on Attention Mechanism
Ruicong Xu, Li Niu 0002, Liqing Zhang 0001 |
MMM (1) | 3 |
| 2020 | Knowledge Graph-based Event Embedding Framework for Financial Quantitative InvestmentsabstractEvent representative learning aims to embed news events into continuous space vectors for capturing syntactic and semantic information from text corpus, which is benefit to event-driven quantitative investments. However, the financial market reaction of events is also influenced by the lead-lag effect, which is driven by internal relationships. Therefore, in this paper, we present a knowledge graph-based event embedding framework for quantitative investments. In particular, we first extract structured events from raw texts, and construct the knowledge graph with the mentioned entities and relations simultaneously. Then, we leverage a joint model to merge the knowledge graph information into the objective function of an event embedding learning model. The learned representations are fed as inputs of downstream quantitative trading methods. Extensive experiments on real-world dataset demonstrate the effectiveness of the event embeddings learned from financial news and knowledge graphs. We also deploy the framework for quantitative algorithm trading. The accumulated portfolio return contributed by our method significantly outperforms other baselines. Dawei Cheng, Fangzhou Yang, Xiaoyang Wang 0002, Ying Zhang 0001, Liqing Zhang 0001 |
SIGIR | 5 |
| 2020 | Multi-mode neural network for human action recognitionabstractVideo data are of two different intrinsic modes, in‐frame and temporal. It is beneficial to incorporate static in‐frame features to acquire dynamic features for video applications. However, some existing methods such as recurrent neural networks do not have a good performance, and some other such as 3D convolutional neural networks (CNNs) are both memory consuming and time consuming. This study proposes an effective framework that takes the advantage of deep learning on the static image feature extraction to tackle the video data. After extracting in‐frame feature vectors using a pretrained deep network, the authors integrate them and form a multi‐mode feature matrix, which preserves the multi‐mode structure and high‐level representation. They propose two models for follow‐up classification. The authors first introduce a temporal CNN, which directly feeds the multi‐mode feature matrix into a CNN. However, they show that characteristics of the multi‐mode features differ significantly in distinct modes. The authors therefore further propose the multi‐mode neural network (MMNN), in which different modes deploy different types of layers. They evaluate their algorithm with the task of human action recognition. The experimental results show that the MMNN achieves a much better performance than the existing long short‐term memory‐based methods and consumes far fewer resources than the existing 3D end‐to‐end models. Haohua Zhao 0001, Weichen Xue, Zhangxuan Gu, Li Niu 0002, Liqing Zhang 0001 |
IET Comput. Vis. | 6 |
| 2020 | Image Editing via Segmentation Guided Self-Attention NetworkabstractImage editing is one of the most popular directions in computer vision. Recently, many methods have benefited from the advances in deep learning, showing promising performance in the image editing task by inpainting the editing areas. These methods take advantage of edge information as user guidance to generate the desired content. However, they are suffering from generating color discrepancy and inconsistent boundaries. In this letter, we propose a deep image editing method based on a self-attention network which copies information for each of the small patches from distant spatial locations. The proposed method smooths the image, computes segmentation maps, and utilizes the segmentation information for guiding the self-attention layers to explicitly leverage image features from surrounding areas with similar appearances. Experimental results show that the proposed method achieves better performance, is flexible for different purposes, and is fast for implementation. Jianfu Zhang 0003, Peiming Yang, Wentao Wang 0009, Yan Hong 0001, Liqing Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2019 | Multi-Attribute Transfer via Disentangled RepresentationabstractRecent studies show significant progress in image-to-image translation task, especially facilitated by Generative Adversarial Networks. They can synthesize highly realistic images and alter the attribute labels for the images. However, these works employ attribute vectors to specify the target domain which diminishes image-level attribute diversity. In this paper, we propose a novel model formulating disentangled representations by projecting images to latent units, grouped feature channels of Convolutional Neural Network, to disassemble the information between different attributes. Thanks to disentangled representation, we can transfer attributes according to the attribute labels and moreover retain the diversity beyond the labels, namely, the styles inside each image. This is achieved by specifying some attributes and swapping the corresponding latent units to “swap” the attributes appearance, or applying channel-wise interpolation to blend different attributes. To verify the motivation of our proposed model, we train and evaluate our model on face dataset CelebA. Furthermore, the evaluation of another facial expression dataset RaFD demonstrates the generalizability of our proposed model. Jianfu Zhang 0003, Yaoyi Li, Weijie Zhao 0003, Liqing Zhang 0001 |
AAAI | 5 |
| 2019 | A Dynamic Default Prediction Framework for Networked-guarantee LoansabstractCommercial banks normally require Small and Medium Enterprises (SMEs) to provide their warranties when applying for a loan. If the borrower defaults, the guarantor is obligated to repay its loan. Such a guarantee system is designed to reduce delinquent risks, but may introduce a new dimension risk if more and more SMEs involve and subsequently form complex temporal networks. Monitoring the financial status of SMEs in these networks, and preventing or reducing systematic loan risk, is an area of great concern for both the regulatory commission and the banks. To allow possible actions to be taken in advance, this paper studies the problem of predicting repayment delinquency in the networked-guarantee loans. We propose a dynamic default prediction framework (DDPF), which preserves temporal network structures and loan behavior sequences in an end-to-end model. In particular, we design a gated recursive and attention mechanism to integrate both the loan behavior and network information. Then, we uncover risky warrant patterns by the learned weights, which effectively accelerate risk evaluation process. Finally, we conduct extensive experiments in a real-world loan risk control system to evaluate its performance, the results demonstrate the effectiveness of our proposed approach compared with state-of-the-art baselines. Dawei Cheng, Yiyi Zhang 0002, Fangzhou Yang, Zhibin Niu, Liqing Zhang 0001 |
CIKM | 6 |
| 2019 | Clothes Keypoints Localization and Attribute Recognition via Prior KnowledgeabstractRich clothes datasets and high-quality annotations have driven recent advances in fashion clothes recognition. However, the existing approaches treat clothes as common images, ignoring the prior clothing knowledge such as spatial relations, symmetry, proportions, and key characteristics of clothes. In order to combine the semantic information with the advantages of deep learning, we propose Detection+, a model using the prior symmetric constraint to refine the keypoints located by any backbone detection networks. To deal with uncertainty in labelling clothing, we introduce a new loss to utilize all available data which contain "maybe" labels. Detection+ has reduced about 2.54% Normalized Error in FashionAI dataset and improved 3.2% AP in human keypoints dataset coco2017 compared to the Mask R-CNN baseline. A large number of experimental results show the proposed approach achieves better results in different recognition datasets (resp., FashionAI, and Deepfashion) with about (resp., 2.57% mAP, and 10% recall) improvements. Zhangxuan Gu, Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
ICME | 5 |
| 2019 | Multi-person 3D Pose Estimation from Monocular Image Sequences
Nayun Xu, Xutong Lu, Yucheng Xing, Haohua Zhao 0001, Li Niu 0002, Liqing Zhang 0001 |
ICONIP (2) | 7 |
| 2019 | Risk Assessment for Networked-guarantee Loans Using High-order Graph Attention RepresentationabstractAssessing and predicting the default risk of networked-guarantee loans is critical for the commercial banks and financial regulatory authorities. The guarantee relationships between the loan companies are usually modeled as directed networks. Learning the informative low-dimensional representation of the networks is important for the default risk prediction of loan companies, even for the assessment of systematic financial risk level. In this paper, we propose a high-order graph attention representation method (HGAR) to learn the embedding of guarantee networks. Because this financial network is different from other complex networks, such as social, language, or citation networks, we set the binary roles of vertices and define high-order adjacent measures based on financial domain characteristics. We design objective functions in addition to a graph attention layer to capture the importance of nodes. We implement a productive learning strategy and prove that the complexity is near-linear with the number of edges, which could scale to large datasets. Extensive experiments demonstrate the superiority of our model over state-of-the-art method. We also evaluate the model in a real-world loan risk control system, and the results validate the effectiveness of our proposed approaches. Dawei Cheng, Zhen-Wei Ma, Zhibin Niu, Liqing Zhang 0001 |
IJCAI | 5 |
| 2019 | GAIN: Gradient Augmented Inpainting Network for Irregular HolesabstractImage inpainting, which aims to fill the missing holes of the images, is a challenging task because the holes may contain complicated structures or different possible layouts. Deep learning methods have shown promising performance in image inpainting but still, suffer from generating poor-structured artifacts when the holes are large and irregular. Some existing methods use edge inpainting to help image inpainting, with binary edge map obtained from image gradient. However, by only using the binary edge map, these methods discard the rich information in image gradient and thus leave some critical issues (e.g. , color discrepancy) unattended. In this paper, we propose Gradient Augmented Inpainting Network (GAIN), which uses image gradient information instead of edge information to facilitate image inpainting. Specifically, we formulate a multi-task learning framework which performs image inpainting and gradient inpainting simultaneously. A novel GAI-Block is designed to encourage the information fusion between the image feature map and the gradient feature map. Moreover, gradient information is also used to determine the filling priority, which can guide the network to construct more plausible semantic structures for the holes. Experimental results on public datasets CelebA-HQ and Places2 show that our proposed method outperforms state-of-the-art methods quantitatively and qualitatively. Jianfu Zhang 0003, Li Niu 0002, Dexin Yang, Liwei Kang, Yaoyi Li, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 7 |
| 2019 | BHONEM: Binary High-Order Network Embedding Methods for Networked-Guarantee Loans
Dawei Cheng, Zhen-Wei Ma, Zhibin Niu, Liqing Zhang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2019 | Zero-Shot Learning via Category-Specific Visual-Semantic Mapping and Label RefinementabstractZero-Shot Learning (ZSL) aims to classify a test instance from an unseen category based on the training instances from seen categories, in which the gap between seen categories and unseen categories is generally bridged via visual-semantic mapping between the low-level visual feature space and the intermediate semantic space. However, the visual-semantic mapping (i.e., projection) learnt based on seen categories may not generalize well to unseen categories, which is known as the projection domain shift in ZSL. To address this projection domain shift issue, we propose a method named Adaptive Embedding ZSL (AEZSL) to learn an adaptive visual-semantic mapping for each unseen category, followed by progressive label refinement. Moreover, to avoid learning visual-semantic mapping for each unseen category in the large-scale classification task, we additionally propose a deep adaptive embedding model named Deep AEZSL (DAEZSL) sharing the similar idea (i.e., visual-semantic mapping should be category-specific and related to the semantic space) with AEZSL, which only needs to be trained once, but can be applied to arbitrary number of unseen categories. Extensive experiments demonstrate that our proposed methods achieve the state-of-theart results for image classification on three small-scale benchmark datasets and one large-scale benchmark dataset. Li Niu 0002, Jianfei Cai 0001, Ashok Veeraraghavan, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Visual Analytics for Networked-Guarantee Loans Risk ManagementabstractGroups of enterprises can guarantee each other and form complex networks in order to try to obtain loans from banks. Monitoring the financial status of a network, and preventing or reducing systematic risk in case of a crisis, is an area of great concern for the regulatory commission and for the banks. We set the ultimate goal of developing a visual analytic approach and tool for risk dissolving and decision-making. We have consolidated four main analysis tasks conducted by financial experts: i) Multi-faceted Default Risk Visualization, whereby a hybrid representation is devised to predict the default risk and an interface developed to visualize key indicators; ii) Risk Guarantee Patterns Discovery. We follow the Shneiderman mantra guidance for designing interactive visualization applications, whereby an interactive risk guarantee community detection and a motif detection based risk guarantee pattern discovery approach are described; iii) Network Evolution and Retrospective, whereby animation is used to help users to understand the guarantee dynamic; iv) Risk Communication Analysis. The temporal diffusion path analysis can be useful for the government and banks to monitor the spread of the default status. It also provides insight for taking precautionary measures to prevent and dissolve systematic financial risk. We implement the system with case studies using real-world bank loan data. Two financial experts are consulted to endorse the developed tool. To the best of our knowledge, this is the first visual analytics tool developed to explore networked-guarantee loan risks in a systematic manner. Zhibin Niu, Dawei Cheng, Liqing Zhang 0001, Jiawan Zhang |
PacificVis | 3 |
| 2018 | Multi-Shot Pedestrian Re-Identification via Sequential Decision MakingabstractMulti-shot pedestrian re-identification problem is at the core of surveillance video analysis. It matches two tracks of pedestrians from different cameras. In contrary to existing works that aggregate single frames features by time series model such as recurrent neural network, in this paper, we propose an interpretable reinforcement learning based approach to this problem. Particularly, we train an agent to verify a pair of images at each time. The agent could choose to output the result (same or different) or request another pair of images to verify (unsure). By this way, our model implicitly learns the difficulty of image pairs, and postpone the decision when the model does not accumulate enough evidence. Moreover, by adjusting the reward for unsure action, we can easily trade off between speed and accuracy. In three open benchmarks, our method are competitive with the state-of-the-art methods while only using 3% to 6% images. These promising results demonstrate that our method is favorable in both efficiency and performance. Jianfu Zhang 0003, Naiyan Wang, Liqing Zhang 0001 |
CVPR | 3 |
| 2018 | Exploring Motor Imagery Eeg Patterns for Stroke Patients with Deep Neural NetworksabstractStudies show that motor imagery based Brain-Computer Interface (BCI) systems can be utilized therapeutically in stroke rehabilitation. Efficient decoding of subjects' motor intentions is essential in BCI-based rehabilitation systems to manipulate a neural prosthesis or other devices for motor relearning. However, due to cortical reorganization, the desynchro-nization potential evoked by the motor imagery of patients with brain lesions is quite different from that evoked by the motor imagery of normal subjects. These differences can be attributed to active cortex regions, frequency bands and amplitude. In this paper, we use a deep learning method to explore the EEG patterns of key channels and the frequency band for stroke patients. The EEG data is bandpass filtered into multiple sub-bands split by a sliding window strategy. Under these sub-bands, diverse spatial-spectral features are extracted and fed into a deep neural network for classification, uncovering the spectral patterns (bandpass filters) and spatial patterns (spatial weights). Experimental results from five stroke patients show that our method has higher classification accuracy than several state-of-the-art approaches. By tracking gradual changes in EEG patterns during rehabilitation, we try to uncover the neurophysiological plasticity mechanism in the impaired cortexes of stroke patients. Dawei Cheng, Ye Liu 0008, Liqing Zhang 0001 |
ICASSP | 3 |
| 2018 | Learning Temporal Relationships Between Financial SignalsabstractPortfolio risk control is vital to financial institutions: investors seek to build equities with the highest return but with minimum risk. However, a general phenomenon is significant comovement among many financial signals, such as stocks and futures. One investment strategy is to choose less correlated assets. Classic approaches quantifying such relationships in real financial markets make it difficult to exclude factors such as market trends and autocorrelation. In this paper, we propose a signal process perspective for quantitative measurement. A machine learning based algorithm is designed to model returns, taking account of market sensitivity, autocorrelation, and relationships with other stocks. We then extend the model training algorithm using regularized least square and gradient descent to estimate parameters. A penalty factor is designed in the optimization function to address extreme large negative returns. After denoising common factors, the learned pure relationship parameters are applied to construct a relationship matrix. Finally, we use this matrix to build portfolios by constrained optimization. Empirical experiments on two stock datasets show that the proposed method outperforms several state-of-the-art methods in terms of mean average precision and cumulative returns. Dawei Cheng, Zhibin Niu, Liqing Zhang 0001 |
ICASSP | 4 |
| 2018 | Attention-Based Network for Cross-View Gait Recognition
Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (7) | 4 |
| 2018 | Recurrent RetinaNet: A Video Object Detection Model Based on Focal Loss
Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (4) | 3 |
| 2018 | Prediction Defaults for Networked-guarantee LoansabstractNetworked-guarantee loans may cause the systemic risk related concern of the government and banks in China. The prediction of default of enterprise loans is a typical extremely imbalanced prediction problem, and the networked-guarantee make this problem more difficult to solve. Since the guaranteed loan is a debt obligation promise, if one enterprise in the guarantee network falls into a financial crisis, the debt risk may spread like a virus across the guarantee network, even lead to a systemic financial crisis. In this paper, we propose an imbalanced network risk diffusion model to forecast the enterprise default risk in a short future. Positive weighted k-nearest neighbors (p-wkNN) algorithm is developed for the stand-alone case - when there is no default contagious; then a data-driven default diffusion model is integrated to further improve the prediction accuracy. We perform the empirical study on a real-world three-years loan record from a major commercial bank. The results show that our proposed method outperforms conventional credit risk methods in terms of AUC. In summary, our quantitative risk evaluation model shows promising prediction performance on real-world data, which could be useful to both regulators and stakeholders. Dawei Cheng, Zhibin Niu, Liqing Zhang 0001 |
ICPR | 4 |
| 2018 | Fashion Sensitive Clothing Recommendation Using Hierarchical Collocation ModelabstractAutomatic clothing recommendation grows dramatically due to the booming of apparel e-commerce. In this paper, we propose a novel clothing recommendation approach which is sensitive to the fashion trend. The proposed approach incorporates the expert knowledge into multiple dimensional information including purchase behaviors, image contents and product descriptions so as to provide recommendation of clothing in line with the forefront of fashion. Meanwhile, to meet with human visual aesthetics and user's collocation experience, we propose the integration of the convolutional neural network and the hierarchical collocation model (HCM) into our framework. The former is to extract effective visual features and attribute descriptors from the clothing items, while the latter embeds them into the concept of style topics which interpret the collocation pattern from a higher level of semantic knowledge. Such a data driven recommendation approach is able to learn clothing collocation metric from multi-dimensional clothing information. Experimental results show that our HCM method achieves better performance than other state-of-the-art baselines. Besides, it also ensures the fashion sensitivity of the recommended outfits. Zhengzhong Zhou, Xiu Di, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2018 | Extracting hierarchical spatial and temporal features for human action recognition
Keting Zhang, Liqing Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Supervised Dictionary Learning with Smooth Shrinkage for Image Denoising
Keting Zhang, Liqing Zhang 0001 |
Neural Process. Lett. | 2 |
| 2018 | Content-Based Image Retrieval Using Iterative Search
Zhengzhong Zhou, Liqing Zhang 0001 |
Neural Process. Lett. | 2 |
| 2017 | Non-blind image deconvolution using deep dual-pathway rectifier neural networkabstractRecently deep neural networks have been successfully used for natural image deconvolution. Whereas the existing methods usually involve an inversion of the blur followed by a denoising step. In this paper we propose a pure learning approach to learn a mapping from a blurred patch to a clean patch directly with a deep dual-pathway rectifier neural network. The experimental results show that our approach outperform the state-of-the-art methods on non-blind image deconvolution within reasonable training time. By analyzing the learned representations, we empirically show that our model works by efficiently detecting the blurry input patterns and then reconstructing the clean patch with the corresponding dictionary atoms. Keting Zhang, Weichen Xue, Liqing Zhang 0001 |
ICASSP | 3 |
| 2017 | Deep Encoding Features for Instance Retrieval
Zhiming Ding, Zhengzhong Zhou, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2017 | Pedestrian Counting System Based on Multiple Object Detection and Tracking
Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2017 | Deep Part-Based Image Feature for Clothing Retrieval
Laiping Zhou, Zhengzhong Zhou, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2016 | Higher-Order Correlation Coefficient Analysis for EEG-Based Brain-Computer InterfaceabstractElectroencephalogram (EEG) based brain-computer interface (BCI) has been proved to be an effective communication way between human brain and external devices. In order to effectively recover the cortical dynamics from the EEG signals and improve the classification performance, plenty of studies focused on constructing subject-specific spatial and spectral filters, achieving considerable improvement in classification accuracy. However, almost all the approaches aimed to find one common subspace for projection of all the samples in different classes. Studies have shown that active channels and frequency information were not only subject-dependent but also class-dependent. Thus the variety of class-dependent spatial and spectral characteristics can provide further discriminative information for classification. In this paper, we proposed a tensor-based method which attempted to seek individual spatial and spectral subspaces for each class by which each class was projected into its own subspace separately such that they were easily to be classified. Finally, we added a regularization term in this model to avoid overfitting. We evaluated the effectiveness and robustness of the proposed method on two different datasets including one widely-used benchmark EEG dataset collected from healthy subjects and one self-collected EEG dataset collected from stroke patients. The results demonstrated its superior performance. Ye Liu 0008, Qibin Zhao, Liqing Zhang 0001 |
ECAI | 4 |
| 2016 | Credit Card Fraud Detection Using Convolutional Neural Networks
Dawei Cheng, Liqing Zhang 0001 |
ICONIP (3) | 4 |
| 2016 | Encoding Multi-resolution Two-Stream CNNs for Action Recognition
Weichen Xue, Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2016 | Content-Based Image Retrieval Using Deep Search
Zhengzhong Zhou, Liqing Zhang 0001 |
ICONIP (2) | 2 |
| 2016 | Interactive Image Search for Clothing RecommendationabstractThis demo delivers a novel retrieval system which meets users' multi-dimensional requirements in clothing image search. In this system, users are able to use both image and keywords as query inputs. We employ the color, texture, shape and attributes as additional descriptors to further refine the requirements. We propose the Hybrid Topic (HT) model, a probabilistic network integrating the multi-channel descriptors into a unified framework, to learn the intricate semantic representation of the descriptors above. The proposed model provides an effective multi-modal representation of clothes. Our experiments show that the HT method significantly outperforms the CNN-based deep search methods. Zhengzhong Zhou, Jingjin Zhou, Liqing Zhang 0001 |
ACM Multimedia | 4 |
| 2016 | Demand-adaptive Clothing Image Retrieval Using Hybrid Topic ModelabstractThis paper proposes a novel approach to meet users' multi-dimensional requirements in clothing image retrieval. It enables users to add search conditions by modifying the color, texture, shape and attribute descriptors of the query images to further refine their requirements. We propose the Hybrid Topic (HT) model to learn the intricate semantic representation of the descriptors above. The model provides an effective multi-dimensional representation of clothes and is able to perform automatic image annotation by probabilistic reasoning from image search. Furthermore, we develop a demand-adaptive retrieval strategy which refines users' specific requirements and removes users' unwanted features. Our experiments show that the HT method significantly outperforms the deep neural network methods. The accuracy could be further improved in cooperation with image annotation and demand-adaptive retrieval strategy. Zhengzhong Zhou, Jingjin Zhou, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2016 | Bayesian Robust Tensor Factorization for Incomplete Multiway DataabstractWe propose a generative model for robust tensor factorization in the presence of both missing data and outliers. The objective is to explicitly infer the underlying low-CANDECOMP/PARAFAC (CP)-rank tensor capturing the global information and a sparse tensor capturing the local information (also considered as outliers), thus providing the robust predictive distribution over missing entries. The low-CP-rank tensor is modeled by multilinear interactions between multiple latent factors on which the column sparsity is enforced by a hierarchical prior, while the sparse tensor is modeled by a hierarchical view of Student-t distribution that associates an individual hyperparameter with each element independently. For model learning, we develop an efficient variational inference under a fully Bayesian treatment, which can effectively prevent the overfitting problem and scales linearly with data size. In contrast to existing related works, our method can perform model selection automatically and implicitly without the need of tuning parameters. More specifically, it can discover the groundtruth of CP rank and automatically adapt the sparsity inducing priors to various types of outliers. In addition, the tradeoff between the low-rank approximation and the sparse representation can be optimized in the sense of maximum model evidence. The extensive experiments and comparisons with many state-of-the-art algorithms on both synthetic and real-world data sets demonstrate the superiorities of our method from several perspectives. Qibin Zhao, Guoxu Zhou, Liqing Zhang 0001, Andrzej Cichocki, Shun-ichi Amari |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Object proposal by multi-branch hierarchical segmentationabstractHierarchical segmentation based object proposal methods have become an important step in modern object detection paradigm. However, standard single-way hierarchical methods are fundamentally flawed in that the errors in early steps cannot be corrected and accumulate. In this work, we propose a novel multi-branch hierarchical segmentation approach that alleviates such problems by learning multiple merging strategies in each step in a complementary manner, such that errors in one merging strategy could be corrected by the others. Our approach achieves the state-of-the-art performance for both object proposal and object detection tasks, comparing to previous object proposal methods. Chaoyang Wang 0001, Long Zhao 0003, Shuang Liang 0001, Liqing Zhang 0001, Jinyuan Jia 0002 |
CVPR | 4 |
| 2015 | Offline Sketch Parsing via Shapeness Estimation
Changhu Wang, Liqing Zhang 0001, Yong Rui |
IJCAI | 3 |
| 2015 | Sketch-based Image Retrieval via Shape WordsabstractThe explosive growth of touch screens has provided a good platform for sketch-based image retrieval. However, most previous works focused on low level descriptors of shapes and sketches. In this paper, we try to step forward and propose to leverage shape words descriptor for sketch-based image retrieval. First, the shape words are defined and an efficient algorithm is designed for shape words extraction. Then we generalize the classic Chamfer Matching algorithm to address the shape words matching problem. Finally, a novel inverted index structure is proposed to make shape words representation scalable to large scale image databases. Experimental results show that our method achieves competitive accuracy but requires much less memory, e.g., less than 3% of memory storage of MindFinder. Due to its competitive accuracy and low memory cost, our method can scale up to much larger database. Changcheng Xiao, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ICMR | 3 |
| 2015 | IdeaPanel: A Large Scale Interactive Sketch-based Image Search SystemabstractIn this work, we introduce the IdeaPanel system, an interactive sketch-based image search engine with millions of images. IdeaPanel enables users to sketch the target image in their minds and also supports tagging to describe their intentions. After a search is triggered, similar images will be returned in real time, based on which users can interactively refine their query sketches until ideal images are returned. Different from existing work, most of which requires a huge amount of memory for indexing and matching, IdeaPanel can achieve very competitive performance but requires much less memory storage. IdeaPanel needs only about 240MB memory to index 1.3M images (less than 3% of previous MindFinder system). Due to its high accuracy and low memory cost, IdealPanel can scale up to much larger database and thus has larger potential to return the most desired images for users. Changcheng Xiao, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ICMR | 3 |
| 2015 | PPTLens: Create Digital Objects with Sketch ImagesabstractIn this work, we introduce the PPTLens system to convert sketch images captured by smart phones to digital flowcharts in PowerPoint. Different from existing sketch recognition system, which is based on hand-drawn strokes, PPTLens enables users to use sketch images as inputs directly. It's more challenging since strokes extracted from sketch images might not only be very messy, but also without temporal information of the drawings. To implement the 'Image to Object' (I2O) scenario, we propose a novel sketch image recognition framework, including an effective stroke extraction strategy and a novel offline sketch parsing algorithm. By enabling sketch images as inputs, our system makes flowchart/diagram production much more convenient and easier. Changcheng Xiao, Changhu Wang, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2015 | Uncorrelated Multiway Discriminant Analysis for Motor Imagery EEG ClassificationabstractMotor imagery-based brain-computer interfaces (BCIs) training has been proved to be an effective communication system between human brain and external devices. A practical problem in BCI-based systems is how to correctly and efficiently identify and extract subject-specific features from the blurred scalp electroencephalography (EEG) and translate those features into device commands in order to control external devices. In real BCI-based applications, we usually define frequency bands and channels configuration that related to brain activities beforehand. However, a steady configuration usually loses effects due to individual variability among different subjects in practical applications. In this study, a robust tensor-based method is proposed for a multiway discriminative subspace extraction from tensor-represented EEG data, which performs well in motor imagery EEG classification without the prior neurophysiologic knowledge like channels configuration and active frequency bands. Motor imagery EEG patterns in spatial-spectral-temporal domain are detected directly from the multidimensional EEG, which may provide insights to the underlying cortical activity patterns. Extensive experiment comparisons have been performed on a benchmark dataset from the famous BCI competition III as well as self-acquired data from healthy subjects and stroke patients. The experimental results demonstrate the superior performance of the proposed method over the contemporary methods. Ye Liu 0008, Qibin Zhao, Liqing Zhang 0001 |
Int. J. Neural Syst. | 3 |
| 2015 | Feature learning from incomplete EEG with denoising autoencoder
Zbigniew R. Struzik, Liqing Zhang 0001, Andrzej Cichocki |
Neurocomputing | 3 |
| 2015 | Separation and Recognition of Electroencephalogram Patterns Using Temporal Independent Component AnalysisabstractA common problem in Electroencephalogram (EEG) analysis is how to separate EEG patterns from noisy recordings. Independent component analysis (ICA), which is an effective method to recover independent sources from sensor outputs without assuming any a priori knowledge, has been widely used in such biological signals analysis. However, when dealing with EEG signals, the mixing model usually does not satisfy the standard ICA assumptions due to the time-variable structures of source signals. In this case, EEG patterns should be precisely separated and recognized in a short time window. Another issue is that we usually over-separate the signals by ICA due to the over learning problem when the length of data is not sufficient. In order to tackle these problems mentioned above, we try to exploit both high order statistics and temporal structures of source signals under condition of short time windows. We utilize a temporal-independent component analysis (tICA) method to formulate the blind separation problem into a new framework of analyzing the mutual independence of the residual signals. Furthermore, in order to find better features for classification, both temporal and spatial features of EEG recordings are extracted by integrating tICA together with some other algorithm like Common Spatial Pattern (CSP) for feature extraction. Computer simulations are given to evaluate the efficiency and performance of tICA based on EEG data recorded not from the normal people but from some special populations suffering from neurophysiological diseases like stroke. To the best of our knowledge, this is the first time that EEG characteristics of stroke patients are explored and reported using ICA algorithm. Superior separation performance and high classification rate evidence that the tICA method is promising for EEG analysis. Ye Liu 0008, Mingfen Li, Liqing Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2015 | Statistically Adaptive Image Denoising Based on Overcomplete Topographic Sparse Coding
Haohua Zhao 0001, Zhiheng Huang, Takefumi Nagumo, Jun Murayama, Liqing Zhang 0001 |
Neural Process. Lett. | 6 |
| 2015 | Bayesian CP Factorization of Incomplete Tensors with Automatic Rank DeterminationabstractCANDECOMP/PARAFAC (CP) tensor factorization of incomplete data is a powerful technique for tensor completion through explicitly capturing the multilinear latent factors. The existing CP algorithms require the tensor rank to be manually specified, however, the determination of tensor rank remains a challenging problem especially for CP rank . In addition, existing approaches do not take into account uncertainty information of latent factors, as well as missing entries. To address these issues, we formulate CP factorization using a hierarchical probabilistic model and employ a fully Bayesian treatment by incorporating a sparsity-inducing prior over multiple latent factors and the appropriate hyperpriors over all hyperparameters, resulting in automatic rank determination. To learn the model, we develop an efficient deterministic Bayesian inference algorithm, which scales linearly with data size. Our method is characterized as a tuning parameter-free approach, which can effectively infer underlying multilinear factors with a low-rank constraint, while also providing predictive distributions over missing entries. Extensive simulations on synthetic data illustrate the intrinsic capability of our method to recover the ground-truth of CP rank and prevent the overfitting problem, even when a large amount of entries are missing. Moreover, the results from real-world applications, including image inpainting and facial image synthesis, demonstrate that our method outperforms state-of-the-art approaches for both tensor factorization and tensor completion in terms of predictive performance. Qibin Zhao, Liqing Zhang 0001, Andrzej Cichocki |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Sketch Recognition with Natural Correction and EditingabstractIn this paper, we target at the problem of sketch recognition. We systematically study how to incorporate users' correction and editing into isolated and full sketch recognition. This is a natural and necessary interaction in real systems such as Visio where very similar shapes exist. First, a novel algorithm is proposed to mine the prior shape knowledge for three editing modes. Second, to differentiate visually similar shapes, a novel symbol recognition algorithm is introduced by leveraging the learnt shape knowledge. Then, a novel editing detection algorithm is proposed to facilitate symbol recognition. Furthermore, both of the symbol recognizer and the editing detector are systematically incorporated into the full sketch recognition. Finally, based on the proposed algorithms, a real-time sketch recognition system is built to recognize hand-drawn flowcharts and diagrams with flexible interactions. Extensive experiments show the effectiveness of the proposed algorithms. Changhu Wang, Liqing Zhang 0001, Yong Rui |
AAAI | 3 |
| 2014 | Uncorrelated Multilinear Nearest Feature Line AnalysisabstractIn this paper, we propose a new subspace learning method, called uncorrelated multilinear nearest feature line analysis (UMNFLA), for the recognition of multidimensional objects, known as tensor objects. Motivated by the fact that existing nearest feature line (NFL) can effectively characterize the geometrical information of limited samples, and uncorrelated features are desirable for many pattern analysis applications since they contain minimum redundancy and ensure independence of features, we propose using the NFL metric to seek a feature subspace such that the within-class feature line (FL) distances are minimized and between-class FL distances are maximized simultaneously in the reduced subspace, and impose an uncorrelated constraint to extract statistically uncorrelated features directly from tensorial data. UMNFLA seeks a tensor-to-vector projection (TVP) that captures most of the variation in the original tensorial input, and employs sequential iterative steps based on the alternating projection method. Experimental results on the task of single trial electroencephalography (EEG) recognition suggest that UMNFLA is particularly effective in determining the low-dimensional projection space needed in such recognition tasks. Ye Liu 0008, Liqing Zhang 0001 |
ECAI | 2 |
| 2014 | Common Spatial-Spectral Boosting Pattern for Brain-Computer InterfaceabstractClassification of multichannel electroencephalogram (EEG) recordings during motor imagination has been exploited successfully for brain-computer interfaces (BCI). Frequency bands and channels configuration that relate to brain activities associated with BCI tasks are often pre-decided as default in EEG analysis without deliberations. However, a steady configuration usually loses effects due to individual variability across different subjects in practical applications. In this paper, we propose an adaptive boosting algorithm in a unifying theoretical framework to model the usually predetermined spatial-spectral configurations into variable preconditions, and further introduce a novel heuristic of stochastic gradient boost for training base learners under these preconditions. We evaluate the effectiveness and robustness of our proposed algorithm based on two data sets recorded from diverse populations including the healthy people and stroke patients. The results demonstrate its superior performance. Ye Liu 0008, Hao Zhang 0072, Qibin Zhao, Liqing Zhang 0001 |
ECAI | 4 |
| 2014 | Tensor-variate Gaussian processes regression and its application to video surveillanceabstractWe present a novel framework for tensor valued Gaussian processes (GP) regression, which exploits a covariance function defined on tensor representation of data inputs. In this way, we bring together the powerful GP methods supported by Bayesian inference and higher-order tensor analysis techniques into one framework. This enables us to account for the underlying structure of data within the model, providing a powerful framework for structural data analysis, such as 3D video sequences. To this end, we propose a new kernel function with tensor arguments under the assumption of generative models, in the form of product kernels where a symmetrical Kullback-Leibler divergence measure is exploited to define the covariance function for tensorial data. A fully Bayesian treatment is employed to estimate the hyperparameters and infer the predictive distributions. Simulation results on both the synthetic data and a real world application of estimating the crowd size from 3D videos demonstrate the effectiveness of the proposed framework. Qibin Zhao, Guoxu Zhou, Liqing Zhang 0001, Andrzej Cichocki |
ICASSP | 3 |
| 2014 | Classification of Stroke Patients' Motor Imagery EEG with Autoencoders in BCI-FES Rehabilitation Training System
Mushangshu Chen, Ye Liu 0008, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2014 | Celebrity Face Image Retrieval Using Multiple Features
Liqing Zhang 0001 |
ICONIP (3) | 2 |
| 2014 | Image Denoising with Rectified Linear Units
Yangwei Wu, Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (3) | 3 |
| 2014 | Sonification for EEG Frequency Spectrum and EEG-Based Emotion Features
Junwei Yue, Liqing Zhang 0001 |
ICONIP (3) | 4 |
| 2014 | Real Time Crowd Counting with Human Detection and Human Tracking
Xinjian Zhang, Liqing Zhang 0001 |
ICONIP (3) | 2 |
| 2014 | Reconstructable generalized maximum scatter difference discriminant analysisabstractDimensionality reduction is a key preprocessing step for many applications. Until our knowledge, unsupervised approaches such as PCA and ICA do not take label information of the original data into account, so a supervised approach such as Linear discriminant analysis (LDA) performs better on many classification tasks. Unfortunately, the classical LDA approach has shortcomings, such as the well-known small size problem, the heteroscedastic problem and the (C-1) low rank problem. The (C-1) low rank problem greatly limits the dimension of the extracted features. In addition, the calculation of the between-class and within-class scatter matrices in the classical LDA approach actually only takes account of the Mahalanobis distance like covariance distance of data centers and each data class, so if the dataset has very few classes or the data distribution of each class is not Gaussian-like but has some spatial structure in the feature space instead, classical LDA does not work well. In this paper we propose a dimensionality reduction approach which avoids the limitations of classical LDA and improves handling of the between-class scatter matrix. Our approach approach takes the distribution of data in each class into consideration to calculate the projection matrix. It does not assume that the data distribution of each class approximates Gaussian; each can have its own spatial structure. Experiments show that our method can obtain better projection directions than the classical LDA approach and greatly improve the classification accuracy. In addition, our approach is able to reconstruct the original signal well, while the classical LDA approach ignores the reconstruction property. Kai Huang 0005, Liqing Zhang 0001 |
IJCNN | 2 |
| 2014 | SmartVisio: Interactive Sketch Recognition with Natural Correction and EditingabstractIn this work, we introduce the SmartVisio system for interactive hand-drawn shape/diagram recognition. Different from existing work, SmartVisio is a real-time sketch recognition system based on Visio, to recognize hand-drawn flowchart/diagram with flexible interactions. This system enables a user to draw shapes or diagrams on the Visio interface, and then the hand-drawn shapes are automatically converted to formal shapes in real-time. To satisfy the interaction needs from common users, we propose an algorithm to detect a user's correction and editing during drawing, and then recognize in real time. We also propose a novel symbol recognition algorithm to better recognize or differentiate some visually similar shapes. By enabling users' natural correction/editing on various shapes, our system makes flowchart/diagram production much more natural and easier. Changhu Wang, Liqing Zhang 0001, Yong Rui |
ACM Multimedia | 3 |
| 2014 | Selected papers from the 2011 International Conference on Neural Information Processing (ICONIP 2011)
James T. Kwok, Liqing Zhang 0001 |
Neurocomputing | 2 |
| 2014 | Multifactor sparse feature extraction using Convolutive Nonnegative Tucker Decomposition
Qiang Wu 0009, Liqing Zhang 0001, Andrzej Cichocki |
Neurocomputing | 2 |
| 2014 | Semisupervised Sparse Multilinear Discriminant Analysis
Kai Huang 0005, Liqing Zhang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2013 | A Tensor-Variate Gaussian Process for Classification of Multidimensional Structured DataabstractAs tensors provide a natural and efficient representation of multidimensional structured data, in this paper, we consider probabilistic multinomial probit classification for tensor-variate inputs with Gaussian processes (GP) priors placed over the latent function. In order to take into account the underlying multimodes structure information within the model, we propose a framework of probabilistic product kernels for tensorial data based on a generative model assumption. More specifically, it can be interpreted as mapping tensors to probability density function space and measuring similarity by an information divergence. Since tensor kernels enable us to model input tensor observations, the proposed tensor-variate GP is considered as both a generative and discriminative model. Furthermore, a fully variational Bayesian treatment for multiclass GP classification with multinomial probit likelihood is employed to estimate the hyperparameters and infer the predictive distributions. Simulation results on both synthetic data and a real world application of human action recognition in videos demonstrate the effectiveness and advantages of the proposed approach for classification of multiway tensor data, especially in the case that the underlying structure information among multimodes is discriminative for the classification task. Qibin Zhao, Liqing Zhang 0001, Andrzej Cichocki |
AAAI | 2 |
| 2013 | Kernel-based tensor partial least squares for reconstruction of limb movementsabstractWe present a new supervised tensor regression method based on multi-way array decompositions and kernel machines. The main issue in the development of a kernel-based framework for tensorial data is that the kernel functions have to be defined on tensor-valued input, which here is defined based on multi-mode product kernels and probabilistic generative models. This strategy enables taking into account the underlying multilinear structure during the learning process. Based on the defined kernels for tensorial data, we develop a kernel-based tensor partial least squares approach for regression. The effectiveness of our method is demonstrated by a real-world application, i.e., the reconstruction of 3D movement trajectories from electrocorticography signals recorded from a monkey brain. Qibin Zhao, Guoxu Zhou, Tülay Adali, Liqing Zhang 0001, Andrzej Cichocki |
ICASSP | 4 |
| 2013 | Gestalt saliency: Salient region detection based on Gestalt principlesabstractSalient region detection is of great significance in computer vision such as object recognition, image segmentation and image retrieval. However, low-level saliency has certain limitations due to lack of object level information. In this paper, we propose a saliency detection method based on Gestalt principles in which we introduce mid-level Gestalt concepts for low-level saliency. We propose an algorithm based on Gestalt principles of similarity & anomaly to select and suppress the similar background regions, using variance of clusters of image regions. Moreover, we propose two smoothing procedures based on Gestalt principles of similarity & proximity to group near and similar regions and therefore uniformly highlight the salient object. Experimental results on public data set show that our method performs well compared with state-of-the-art approaches. Liqing Zhang 0001 |
ICIP | 2 |
| 2013 | Hidden Markov Model for Action Recognition Using Joint Angle Acceleration
Sha Huang, Liqing Zhang 0001 |
ICONIP (3) | 2 |
| 2013 | Spectral Power Estimation for Unevenly Spaced Motor Imagery Data
Zbigniew R. Struzik, Liqing Zhang 0001, Andrzej Cichocki |
ICONIP (1) | 3 |
| 2013 | Causal Neurofeedback Based BCI-FES Rehabilitation for Post-stroke Patients
Ye Liu 0008, Hao Zhang 0072, Liqing Zhang 0001 |
ICONIP (1) | 5 |
| 2013 | Image Denoising Based on Overcomplete Topographic Sparse Coding
Haohua Zhao 0001, Zhiheng Huang, Takefumi Nagumo, Jun Murayama, Liqing Zhang 0001 |
ICONIP (3) | 6 |
| 2013 | Motion Deblurring Using Super-Sparsity
Jingxiong Zhao, Haohua Zhao 0001, Keting Zhang, Liqing Zhang 0001 |
ICONIP (3) | 4 |
| 2013 | Optimal Calculation of Tensor Learning Approaches
Kai Huang 0005, Liqing Zhang 0001 |
ISNN (1) | 2 |
| 2013 | UMPCA Based Feature Extraction for ECG
Kai Huang 0005, Liqing Zhang 0001 |
ISNN (1) | 4 |
| 2013 | A Frequency Boosting Method for Motor Imagery EEG Classification in BCI-FES Rehabilitation Training System
Jianyi Liang, Hao Zhang 0072, Ye Liu 0008, Liqing Zhang 0001 |
ISNN (2) | 6 |
| 2013 | Design of assistive Wheelchair System directly Steered by Human ThoughtsabstractIntegration of brain-computer interface (BCI) technique and assistive device is one of chief and promising applications of BCI system. With BCI technique, people with disabilities do not have to communicate with external environment through traditional and natural pathways like peripheral nerves and muscles, and could achieve it only by their brain activities. In this paper, we designed an electroencephalogram (EEG)-based wheelchair which can be steered by users' own thoughts without any other involvements. We evaluated the feasibility of BCI-based wheelchair in terms of accuracies and real-world testing. The results demonstrate that our BCI wheelchair is of good performance not only in accuracy, but also in practical running testing in a real environment. This fact implies that people can steer wheelchair only by their thoughts, and may have a potential perspective in daily application for disabled people. Jianyi Liang, Qibin Zhao, Jie Li 0016, Kan Hong, Liqing Zhang 0001 |
Int. J. Neural Syst. | 6 |
| 2013 | Higher Order Partial Least Squares (HOPLS): A Generalized Multilinear Regression MethodabstractA new generalized multilinear regression model, termed the higher order partial least squares (HOPLS), is introduced with the aim to predict a tensor (multiway array) Y from a tensor X through projecting the data onto the latent space and performing regression on the corresponding latent variables. HOPLS differs substantially from other regression models in that it explains the data by a sum of orthogonal Tucker tensors, while the number of orthogonal loadings serves as a parameter to control model complexity and prevent overfitting. The low-dimensional latent space is optimized sequentially via a deflation operation, yielding the best joint subspace approximation for both X and Y. Instead of decomposing X and Y individually, higher order singular value decomposition on a newly defined generalized cross-covariance tensor is employed to optimize the orthogonal loadings. A systematic comparison on both synthetic data and real-world decoding of 3D movement trajectories from electrocorticogram signals demonstrate the advantages of HOPLS over the existing methods in terms of better predictive ability, suitability to handle small sample sizes, and robustness to noise. Qibin Zhao, Cesar F. Caiafa, Danilo P. Mandic, Zenas C. Chao, Yasuo Nagasaka, Naotaka Fujii, Liqing Zhang 0001, Andrzej Cichocki |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2012 | Free Hand-Drawn Sketch Segmentation
Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ECCV (1) | 3 |
| 2012 | ECG Classification Based on Non-cardiology Feature
Kai Huang 0005, Liqing Zhang 0001 |
ISNN (2) | 2 |
| 2012 | Query-adaptive shape topic mining for hand-drawn sketch recognitionabstractIn this work, we study the problem of hand-drawn sketch recognition. Due to large intra-class variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited pre-defined classes. Different from existing work, we target at developing a general sketch recognition system, to recognize any semantically meaningful object that a child can recognize. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. To alleviate the problems of intra-class shape variation and inter-class shape ambiguity in this unconstrained situation, a query-adaptive shape topic model is proposed to mine object topics and shape topics related to the sketch, in which, multiple layers of information such as sketch, object, shape, image, and semantic labels are modeled in a generative process. Besides sketch recognition, the proposed topic model can also be used for related applications such as sketch tagging, image tagging, and sketch-based image search. Extensive experiments on different applications show the effectiveness of the proposed topic model and the recognition system. Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 3 |
| 2012 | Sketch2Tag: automatic hand-drawn sketch recognitionabstractIn this work, we introduce the Sketch2Tag system for hand-drawn sketch recognition. Due to large variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited predefined classes. Different from existing work, Sketch2Tag is a general sketch recognition system, towards recognizing any semantically meaningful object that a child can recognize. This system enables a user to draw a sketch on the query panel, and then provides real-time recognition results. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. Better understanding a user's drawing will be of great value to a variety of applications, such as, improving the sketch-based image search by combining the recognition results as textual queries. Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 3 |
| 2012 | Sketch-based image retrieval on a large scale databaseabstractThe paper presents a simple and effective sketch-based algorithm for large scale image retrieval. One of the main challenges in image retrieval is to localize a region in an image which would be matched with the query image in contour. To tackle this problem, we use the human perception mechanism to identify two types of regions in one image: the first type of region (the main region) is defined by a weighted center of image features, suggesting that we could retrieve objects in images regardless of their sizes and positions. The second type of region, called region of interests (ROI), is to find the most salient part of an image, and is helpful to retrieve images with objects similar to the query in a complicated scene. So using the two types of regions as candidate regions for feature extraction, our algorithm could increase the retrieval rate dramatically. Besides, to accelerate the retrieval speed, we first extract orientation features and then organize them in a hierarchal way to generate global-to-local features. Based on this characteristic, a hierarchical database index structure could be built which makes it possible to retrieve images on a very large scale image database online. Finally a real-time image retrieval system on 4.5 million database is developed to verify the proposed algorithm. The experiment results show excellent retrieval performance of the proposed algorithm and comparisons with other algorithms are also given. Liuli Chen, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2012 | Guess what you draw: interactive contour-based image retrieval on a million-scale databaseabstractWe propose a real-time image retrieval system which allows users to search target images whose objects are similar to the query in contour, regardless of their sizes and positions appearing in the images. Even in a complicated scene, as long as the object's contour is most salient in the target image, the system is still able to capture it and lists the image in the retrieval results. Therefore, the system has better retrieval rate than existing systems and algorithms. One typical application of the proposed system is to help the computer understand what does the user draw or upload. It is based on the statistical distributions of tags of retrieved images, and the proposed system feeds back some candidate tags related to the query image. Such tags could be used for further retrieval to refine the result list. In addition, the system provides a friendly interactive interface with multiple queries. These queries are from different combinations of tags, a hand-drawn sketch and a natural image, and could help users search images flexibly and conveniently. The system runs on a database of 1.3 million images and could achieve a real-time retrieval speed. The results in the demonstration show excellent retrieval performance of the proposed system. Liuli Chen, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2012 | A brief introduction to the special issue for ISNN2010
Liqing Zhang 0001, James T. Kwok, Changshui Zhang |
Neurocomputing | 1 |
| 2012 | A hierarchical latent topic model based on sparse coding
Liqing Zhang 0001, Qianwei Bian |
Neurocomputing | 2 |
| 2011 | Edgel index for large-scale sketch-based image searchabstractRetrieving images to match with a hand-drawn sketch query is a highly desired feature, especially with the popularity of devices with touch screens. Although query-by-sketch has been extensively studied since 1990s, it is still very challenging to build a real-time sketch-based image search engine on a large-scale database due to the lack of effective and efficient matching/indexing solutions. The explosive growth of web images and the phenomenal success of search techniques have encouraged us to revisit this problem and target at solving the problem of web-scale sketch-based image retrieval. In this work, a novel index structure and the corresponding raw contour-based matching algorithm are proposed to calculate the similarity between a sketch query and natural images, and make sketch-based image retrieval scalable to millions of images. The proposed solution simultaneously considers storage cost, retrieval accuracy, and efficiency, based on which we have developed a real-time sketch-based image search engine by indexing more than 2 million images. Extensive experiments on various retrieval tasks (basic shape search, specific image search, and similar image search) show better accuracy and efficiency than state-of-the-art methods. Yang Cao 0008, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
CVPR | 3 |
| 2011 | ECG Classification Using ICA Features and Support Vector Machines
Liqing Zhang 0001 |
ICONIP (1) | 2 |
| 2011 | A Two Stage Algorithm for K-Mode Convolutive Nonnegative Tucker Decomposition
Qiang Wu 0009, Liqing Zhang 0001, Andrzej Cichocki |
ICONIP (2) | 2 |
| 2011 | A Novel Oddball Paradigm for Affective BCIs Using Emotional Faces as Stimuli
Qibin Zhao, Akinari Onishi, Yu Zhang 0009, Jianting Cao, Liqing Zhang 0001, Andrzej Cichocki |
ICONIP (1) | 5 |
| 2011 | Sparse Coding Image Denoising Based on Saliency Map Weight
Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (2) | 2 |
| 2011 | Contour-Based Large Scale Image Retrieval
Liqing Zhang 0001 |
ICONIP (3) | 2 |
| 2011 | Moving Object Detecting System with Phase Discrepancy
Liqing Zhang 0001 |
ISNN (2) | 3 |
| 2011 | Multilinear Subspace Regression: An Orthogonal Tensor Decomposition ApproachabstractA multilinear subspace regression model based on so called latent variable decomposition is introduced. Unlike standard regression methods which typically employ matrix (2D) data representations followed by vector subspace transformations, the proposed approach uses tensor subspace transformations to model common latent variables across both the independent and dependent data. The proposed approach aims to maximize the correlation between the so derived latent variables and is shown to be suitable for the prediction of multidimensional dependent data from multidimensional independent data, where for the estimation of the latent variables we introduce an algorithm based on Multilinear Singular Value Decomposition (MSVD) on a specially defined cross-covariance tensor. It is next shown that in this way we are also able to unify the existing Partial Least Squares (PLS) and N-way PLS regression algorithms within the same framework. Simulations on benchmark synthetic data confirm the advantages of the proposed approach, in terms of its predictive ability and robustness, especially for small sample sizes. The potential of the proposed technique is further illustrated on a real world task of the decoding of human intracranial electrocorticogram (ECoG) from a simultaneously recorded scalp electroencephalograph (EEG). Qibin Zhao, Cesar F. Caiafa, Danilo P. Mandic, Liqing Zhang 0001, Tonio Ball, Andreas Schulze-Bonhage, Andrzej Cichocki |
NIPS | 4 |
| 2011 | Robust Multifactor Speech Feature Extraction Based on Gabor AnalysisabstractThe performance of speech recognition systems relies on the consistency and adaptation of the speech feature in complex conditions during the training and testing stages. Traditional systems usually perform poorly under adverse noisy conditions and are not applicable to most real world problems. In this paper, we investigate the speech feature extraction problem in a noisy environment and propose a novel approach based on Gabor filtering and tensor factorization. Recent physiological and psychoacoustic experimental results suggest that the localized spectro-temporal features are essential for auditory perception. To explore this property, we represent the speech signal by using a general higher order tensor and employ two-dimensional Gabor functions with different scales and directions to analyze the localized patches of the power spectrogram. Then the Nonnegative Tensor PCA with sparse constraints is proposed to learn the projection matrices from multiple interrelated feature subspaces. The objective of the sparse constraints is to preserve the statistical characteristic of clean speech data by finding projection matrices of speech subspaces and reduce the noise components which have distributions different from those of clean speech. A multifactor analysis method is proposed to extract robust sparse features by processing the data samples in tensor structure. The simulation results indicate that our proposed method is able to improve the speech recognition performance, especially in noisy environments, compared with the traditional speech feature extraction methods. Qiang Wu 0009, Liqing Zhang 0001, Guangchuan Shi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | A Phase Discrepancy Analysis of Object Motion
Bolei Zhou, Liqing Zhang 0001 |
ACCV (3) | 3 |
| 2010 | Spatial-bag-of-featuresabstractIn this paper, we study the problem of large scale image retrieval by developing a new class of bag-of-features to encode geometric information of objects within an image. Beyond existing orderless bag-of-features, local features of an image are first projected to different directions or points to generate a series of ordered bag-of-features, based on which different families of spatial bag-of-features are designed to capture the invariance of object translation, rotation, and scaling. Then the most representative features are selected based on a boosting-like method to generate a new bag-of-features-like vector representation of an image. The proposed retrieval framework works well in image retrieval task owing to the following three properties: 1) the encoding of geometric information of objects for capturing objects' spatial transformation, 2) the supervised feature selection and combination strategy for enhancing the discriminative power, and 3) the representation of bag-of-features for effective image matching and indexing for large scale image retrieval. Extensive experiments on 5000 Oxford building images and 1 million Panoramio images show the effectiveness and efficiency of the proposed features as well as the retrieval framework. Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001 |
CVPR | 4 |
| 2010 | Multi-modal EEG Online Visualization and Neuro-Feedback
Kan Hong, Liqing Zhang 0001, Jie Li 0016 |
ISNN (2) | 2 |
| 2010 | Affine Invariant Topic Model for Generic Object Recognition
Zhenxiao Li, Liqing Zhang 0001 |
ISNN (2) | 2 |
| 2010 | MindFinder: interactive sketch-based image search on millions of imagesabstractIn this paper, we showcase the MindFinder system, which is an interactive sketch-based image search engine. Different from existing work, most of which is limited to a small scale database or only enables single modality input, MindFinder is a sketch-based multimodal search engine for million-level database. It enables users to sketch major curves of the target image in their mind, and also supports tagging and coloring operations to better express their search intentions. Owning to a friendly interface, our system supports multiple actions, which help users to flexibly design their queries. After each operation, top returned images are updated in real time, based on which users could interactively refine their initial thoughts until ideal images are returned. The novelty of the MindFinder system includes the following two aspects: 1) A multimodal searching scheme is proposed to retrieve images which meet users' requirements not only in structure, but also in semantic meaning and color tone. 2) An indexing framework is designed to make MindFinder scalable in terms of database size, memory cost, and response time. By scaling up the database to more than two million images, MindFinder not only helps users to easily present whatever they are imagining, but also has the potential to retrieve the most desired images in their mind. Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 5 |
| 2010 | Robust Feature Extraction for Speaker Recognition Based on Constrained Nonnegative Tensor Factorization
Qiang Wu 0009, Liqing Zhang 0001, Guangchuan Shi |
J. Comput. Sci. Technol. | 2 |
| 2010 | Regularized tensor discriminant analysis for single trial EEG classification in BCI
Jie Li 0016, Liqing Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2009 | A Novel Hierarchical Model of Attention: Maximizing Information Acquisition
Yang Cao 0008, Liqing Zhang 0001 |
ACCV (1) | 2 |
| 2009 | Scene Gist: A Holistic Generative Model of Natural Image
Bolei Zhou, Liqing Zhang 0001 |
ACCV (2) | 2 |
| 2009 | Robust speech feature extraction based on Gabor filtering and tensor factorizationabstractIn this paper, we investigate the speech feature extraction problem in the noisy environment. A novel approach based on Gabor filtering and tensor factorization is proposed. From recent physiological and psychoacoustic experimental results, localized spectro-temporal features are essential for auditory perception. We employ 2D-Gabor functions with different scales and directions to analyze the localized patches of power spectrogram, by which speech signal can be encoded as a general higher order tensor. Then nonnegative tensor PCA with sparse constraint is used to learn the projection matrices from multiple interrelated feature subspaces and extract the robust features. Experimental results confirm that our proposed method can improve the speech recognition performance, especially in noisy environment, compared with traditional speech feature extraction methods. Qiang Wu 0009, Liqing Zhang 0001, Guangchuan Shi |
ICASSP | 2 |
| 2009 | Multilinear generalization of Common Spatial PatternabstractThe Common Spatial Patterns (CSP) algorithm has been widely used in EEG classification and Brain Computer Interface (BCI). In this paper, we propose a multilinear formulation of the CSP, termed as TensorCSP or Common Tensor Discriminant Analysis (CTDA) for high-order tensor data. As a natural extension of CSP, the proposed algorithm uses the analogous optimization criteria in CSP and a new framework for simultaneous optimization of projection matrices on each mode based on tensor analysis theory is developed. Experimental results demonstrate that our proposed algorithm is able to improve classification accuracy of multi-class motor imagery EEG. Qibin Zhao, Liqing Zhang 0001, Andrzej Cichocki |
ICASSP | 2 |
| 2009 | Slice Oriented Tensor Decomposition of EEG Data for Feature Extraction in Space, Frequency and Time Domains
Qibin Zhao, Cesar F. Caiafa, Andrzej Cichocki, Liqing Zhang 0001, Anh Huy Phan 0001 |
ICONIP (1) | 4 |
| 2009 | Temporal competitive learning induced in neural networks by spike timing-dependent plasticityabstractIn this paper, we introduce a novel computational model of spike timing-dependent plasticity (STDP) that can induce competitive learning in neural networks, and hence can work as an efficient coding mechanism for temporal correlated neural activities. Most computational STDP models use either additive or multiplicative learning rules. Usually additive rules induce competition in many-to-one networks, yet they not only suffer from instability and slow converging speed, but also cannot be extended properly to many-to-many networks. Multiplicative rules on the other hand can reach stable results in a shorter time, but they do not cause competition in any kind of networks. So these models cannot readily explain complex phenomena in neural processing. Here we attack this problem by introducing a modified multiplicative STDP model with a mechanism called dasiaglobal depressionpsila, which induces competitive learning in many-to-many networks while preserves the virtues of original multiplicative models. Moreover, this model tends to group presynaptic neurons according to their firing patterns. Specifically, an ensemble of presynaptic neurons with correlated activities may collectively form strong connections with one postsynaptic neuron, while a different ensemble may connect with another postsynaptic neuron. Overall this model performs pattern grouping according to the input neural activities. We prove this point theoretically and experimentally in this paper. Liqing Zhang 0001 |
IJCNN | 2 |
| 2009 | Age Classification System with ICA Based Local Facial Features
Liqing Zhang 0001 |
ISNN (2) | 2 |
| 2008 | Object Recognition with Task Relevant Combined Local Features
Liqing Zhang 0001 |
ICIC (1) | 2 |
| 2008 | Spatiotemporal feature extraction based on invariance representationabstractThis paper investigates spatiotemporal feature extraction from temporal image sequences based on invariance representation. Invariance representation is one of important functions of the visual cortex. We propose a novel hierarchical model based on invariance and independent component analysis for spatiotemporal feature extraction. Training the model from patches sampled from natural scenes, we can obtain image basis with properties of translational, scaling, and rotational features. Further experiments on TV videos and facial image sequences show different characteristics of spatiotemporal features are achieved by training the proposed model. All these computer simulations verify that our proposed model is successful for spatiotemporal feature extraction. Wenlu Yang, Liqing Zhang 0001 |
IJCNN | 2 |
| 2008 | Incremental Common Spatial Pattern algorithm for BCIabstractA major challenge in applying machine learning methods to Brain-Computer Interfaces (BCIs) is to overcome the on-line non-stationarity of the data blocks. An effective BCI system should be adaptive to and robust against the dynamic variations in brain signals. One solution to it is to adapt the model parameters of BCI system online. However, CSP is poor at adaptability since it is a batch type algorithm. To overcome this, in this paper, we propose the Incremental Common Spatial Pattern (ICSP) algorithm which performs the adaptive feature extraction on-line. This method allows us to perform the online adjustment of spatial filter. This procedure helps the BCI system robust to possible non-stationarity of the EEG data. We test our method to data from BCI motor imagery experiments, and the results demonstrate the good performance of adaptation of the proposed algorithm. Qibin Zhao, Liqing Zhang 0001, Andrzej Cichocki, Jie Li 0016 |
IJCNN | 2 |
| 2008 | Robust Speaker Modeling Based on Constrained Nonnegative Tensor Factorization
Qiang Wu 0009, Liqing Zhang 0001, Guangchuan Shi |
ISNN (1) | 2 |
| 2008 | Dynamic visual attention: searching for coding length incrementsabstractA visual attention system should respond placidly when common stimuli are presented, while at the same time keep alert to anomalous visual inputs. In this paper, a dynamic visual attention model based on the rarity of features is proposed. We introduce the Incremental Coding Length (ICL) to measure the perspective entropy gain of each feature. The objective of our model is to maximize the entropy of the sampled visual features. In order to optimize energy consumption, the limit amount of energy of the system is re-distributed amongst features according to their Incremental Coding Length. By selecting features with large coding length increments, the computational system can achieve attention selectivity in both static and dynamic scenes. We demonstrate that the proposed model achieves superior accuracy in comparison to mainstream approaches in static saliency map generation. Moreover, we also show that our model captures several less-reported dynamic visual search behaviors, such as attentional swing and inhibition of return. Liqing Zhang 0001 |
NIPS | 2 |
| 2008 | Overcomplete topographic independent component analysis
Libo Ma, Liqing Zhang 0001 |
Neurocomputing | 2 |
| 2008 | A Note on Lewicki-Sejnowski Gradient for Learning Overcomplete RepresentationsabstractOvercomplete representations have greater robustness in noise environment and also have greater flexibility in matching structure in the data. Lewicki and Sejnowski (2000) proposed an efficient extended natural gradient for learning the overcomplete basis and developed an overcomplete representation approach. However, they derived their gradient by many approximations, and their proof is very complicated. To give a stronger theoretical basis, we provide a brief and more rigorous mathematical proof for this gradient in this note. In addition, we propose a more robust constrained Lewicki-Sejnowski gradient. Zhaoshui He, Shengli Xie 0001, Liqing Zhang 0001, Andrzej Cichocki |
Neural Comput. | 3 |
| 2007 | Saliency Detection: A Spectral Residual ApproachabstractThe ability of human visual system to detect visual saliency is extraordinarily fast and reliable. However, computational modeling of this basic intelligent behavior still remains a challenge. This paper presents a simple method for the visual saliency detection. Our model is independent of features, categories, or other forms of prior knowledge of the objects. By analyzing the log-spectrum of an input image, we extract the spectral residual of an image in spectral domain, and propose a fast method to construct the corresponding saliency map in spatial domain. We test this model on both natural pictures and artificial images such as psychological patterns. The result indicate fast and robust saliency detection of our method. Liqing Zhang 0001 |
CVPR | 2 |
| 2007 | An Auditory Neural Feature Extraction Method for Robust Speech RecognitionabstractThis paper proposes a neural mechanism motivated system to extract noise resistant features for robust speech recognition. We use nonnegative matrix factorization to construct two layers of auditory neurons which captures the essence of speech patterns. The responses of these neurons to speech are further processed to form an auditory neural cepstral coefficient (ANCC) representation for speech recognition. We test the robustness of ANCC feature on a 51-word corpus, with recognizers trained on clean speech in noisy conditions. Compared with MFCC, ANCC shows less performance degradation and achieves satisfactory recognition accuracies in both non-stationary noise and high noise level conditions. Liqing Zhang 0001, Bin Xia 0001 |
ICASSP (4) | 2 |
| 2007 | Flexible Component Analysis for Sparse, Smooth, Nonnegative Coding or Representation
Andrzej Cichocki, Anh Huy Phan 0001, Rafal Zdunek, Liqing Zhang 0001 |
ICONIP (1) | 4 |
| 2007 | Subject-Adaptive Real-Time BCI System
Liqing Zhang 0001 |
ICONIP (2) | 2 |
| 2007 | Head Pose Estimation Based on Tensor Factorization
Wenlu Yang, Liqing Zhang 0001 |
ICONIP (1) | 2 |
| 2007 | A Hierarchical Generative Model for Overcomplete Topographic Representations in Natural ImagesabstractIn this paper we propose a hierarchical generative model based on sparse coding and analysis of topographic energy dependencies. We further formulate the basic sparse coding into a hierarchical fashion by defining a higher-order topography on the coefficients of nearby basis functions. An algorithm for learning overcomplete topographic basis functions is derived from a direct approximation to the data likelihood. The basis functions learned by the algorithm demonstrate the topographic organization and the emergence of phase-and shift-invariant features—the similar properties of visual complex cells. Moreover, the proposed model yields overcomplete representations. We apply the model to the problem of image denoising. This task suits the model well since Gaussian additive noise is explicitly included in the model. The simulation results suggest that the proposed method outperforms conventional denoising algorithms. Our model is promising in a wide range of fields, such as signal processing and pattern recognition. Libo Ma, Liqing Zhang 0001 |
IJCNN | 2 |
| 2007 | Color conceptualizationabstractIn this paper, we propose a method to manipulate colors of an image. Based on a library of natural color images, our system evolves several prototypes of color distribution of the library, which we call "color concepts". By applying these color concepts on an input image, a user can easily change the mood of image colors in a global manner. Our results of photographs and paintings indicate that this method is capable of high-quality color manipulations. Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2006 | A Time-Dependent Model of Information Capacity of Visual Attention
Liqing Zhang 0001 |
ICONIP (1) | 2 |
| 2006 | Two-Stage Temporally Correlated Source Extraction Algorithm with Its Application in Extraction of Event-Related Potentials
Zhi-Lin Zhang, Liqing Zhang 0001, Xiu-Ling Wu, Jie Li 0016, Qibin Zhao |
ICONIP (2) | 2 |
| 2006 | Two-Stage Blind Deconvolution for V-BLAST OFDM System
Liqing Zhang 0001, Bin Xia 0001 |
ISNN (1) | 2 |
| 2006 | Local Independent Factorization of Natural Scenes
Libo Ma, Liqing Zhang 0001, Wenlu Yang |
ISNN (2) | 2 |
| 2006 | Multichannel Blind Deconvolution Using a Novel Filter Decomposition Method
Bin Xia 0001, Liqing Zhang 0001 |
ISNN (1) | 2 |
| 2005 | Stability Analysis of Multichannel Blind Deconvolution
Bin Xia 0001, Liqing Zhang 0001 |
ISNN (2) | 2 |
| 2004 | Multichannel Blind Deconvolution of Non-minimum Phase System Using Cascade Structure
Bin Xia 0001, Liqing Zhang 0001 |
ICONIP | 2 |
| 2004 | Temporal Independent Component Analysis for Separating Noisy Signals
Liqing Zhang 0001 |
ICONIP | 1 |
| 2004 | EEG Source Localization Using Independent Residual Analysis
Gang Tan, Liqing Zhang 0001 |
ISNN (2) | 2 |
| 2004 | Blind source estimation of FIR channels for binary sources: a grouping decision approach
Yuanqing Li 0001, Andrzej Cichocki, Liqing Zhang 0001 |
Signal Process. | 3 |
| 2004 | Self-adaptive blind source separation based on activation functions adaptationabstractIndependent component analysis is to extract independent signals from their linear mixtures without assuming prior knowledge of their mixing coefficients. As we know, a number of factors are likely to affect separation results in practical applications, such as the number of active sources, the distribution of source signals, and noise. The purpose of this paper to develop a general framework of blind separation from a practical point of view with special emphasis on the activation function adaptation. First, we propose the exponential generative model for probability density functions. A method of constructing an exponential generative model from the activation functions is discussed. Then, a learning algorithm is derived to update the parameters in the exponential generative model. The learning algorithm for the activation function adaptation is consistent with the one for training the demixing model. Stability analysis of the learning algorithm for the activation function is also discussed. Both theoretical analysis and simulations show that the proposed approach is universally convergent regardless of the distributions of sources. Finally, computer simulations are given to demonstrate the effectiveness and validity of the approach. Liqing Zhang 0001, Andrzej Cichocki, Shun-ichi Amari |
IEEE Trans. Neural Networks | 1 |
| 2003 | Blind deconvolution of FIR channels with binary sources: a grouping decision approachabstractThis paper proposes a novel grouping decision approach for blind deconvolution of FIR channels with binary sources. First, necessary and sufficient conditions for recoverability are derived. For single-input systems, a new deterministic algorithm based on grouping and decision is propose to recover the source up to a delay. Then the algorithm is extended to deal with high noise case and long decaying channel case. Furthermore blind deconvolution for multi-input systems also can be carried out as with the case of single input systems. All sources can be recovered sequentially. Finally, the validity and performance of the algorithms are illustrated by several simulation examples. Yuanqing Li 0001, Andrzej Cichocki, Liqing Zhang 0001 |
ICASSP (4) | 3 |
| 2001 | Semiparametric model and superefficiency in blind deconvolution
Liqing Zhang 0001, Shun-ichi Amari, Andrzej Cichocki |
Signal Process. | 1 |
| 1999 | Semiparametric Approach to Multichannel Blind Deconvolution of Nonminimum Phase Systems
Liqing Zhang 0001, Shun-ichi Amari, Andrzej Cichocki |
NIPS | 1 |
| 1999 | Natural gradient algorithm for blind separation of overdetermined mixture with additive noiseabstractWe study the natural gradient approach to blind separation of overdetermined mixtures. First we introduce a Lie group on the manifold of overdetermined mixtures, and endow a Riemannian metric on the manifold based on the property of the Lie group. Then we derive the natural gradient on the manifold using the isometry of the Riemannian metric. Using the natural gradient, we present a new learning algorithm based on the minimization of mutual information. Liqing Zhang 0001, Andrzej Cichocki, Shun-ichi Amari |
IEEE Signal Process. Lett. | 1 |
| 1998 | Two-stage Blind Deconvolution Using State-space Models
Andrzej Cichocki, Liqing Zhang 0001 |
ICONIP | 2 |
| 1998 | Blind Separation of Filtered Sources Using State-Space Approach
Liqing Zhang 0001, Andrzej Cichocki |
NIPS | 1 |