EDBT 2026 Demo / reviewers in the wild / expert
Chao Liang 0001
dblp:10/3072-1
· DBLP profile ↗
120ranked-venue papers
4as first author
49since 2021 · last 2026
0000-0002-8287-8655ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 92 · 4 first-author · 34 since 2021Artificial intelligence and machine learning · 35 · 1 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake DetectionabstractRecent deepfake detection studies often treat unseen sample detection as a ``zero-shot" task, training on images generated by known models but generalizing to unknown ones. A key real-world challenge arises when a model performs poorly on unknown samples, yet these samples remain available for analysis. This highlights that it should be approached as a ``few-shot" task, where effectively utilizing a small number of samples can lead to significant improvement. Unlike typical few-shot tasks focused on semantic understanding, deepfake detection prioritizes image realism, which closely mirrors real-world distributions. In this work, we propose the Few-shot Training-free Network (FTNet) for real-world few-shot deepfake detection. Simple yet effective, FTNet differs from traditional methods that rely on large-scale known data for training. Instead, FTNet uses only one fake sample from an evaluation set, mimicking the scenario where new samples emerge in the real world and can be gathered for use, without any training or parameter updates. During evaluation, each test sample is compared to the known fake and real samples, and it is classified based on the category of the nearest sample. We conduct a comprehensive analysis of AI-generated images from 29 different generative models and achieve a new SoTA performance, with an average improvement of 8.7% compared to existing methods. This work introduces a fresh perspective on real-world deepfake detection: when the model struggles to generalize on a few-shot sample, leveraging the failed samples leads to better performance. Shibo Yao, Renshuai Tao, Xiaolong Zheng 0001, Chao Liang 0001, Chunjie Zhang 0001 |
AAAI | 4 |
| 2026 | Dual-gated multi-behavior recommendation with variational graph autoencoder
Nana Huang, Pengfei Jiao, Zhidong Zhao, Chao Liang 0001, Fengjun Xiao |
Expert Syst. Appl. | 5 |
| 2026 | StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset
Zhengqian Wu, Zhixian Liu, Aodong Chen, Jingyang Zhang, Ruizhe Li 0004, Hanlin Ge, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
Int. J. Comput. Vis. | 9 |
| 2026 | Illumination Explorer: All-Frequency Illumination Estimation via HEALPix-Guided DiffusionabstractEstimating panoramic illumination from a single limited-FOV input image is a critical yet challenging task for rendering realistic objects with complex materials in augmented reality. Existing methods typically either estimate parameterized lighting models or directly generate panoramas in an end-to-end manner. However, both approaches present significant challenges: 1) Parameterized methods struggle to simultaneously capture both high-frequency and low-frequency information under real lighting conditions, and lack a unified model for indoor and outdoor scenes. 2) Direct generation methods often produce unpredictable results, making it difficult to control the position, color, and structure of light sources in the output panorama. In this paper, we propose a unified illumination estimation method based on pretrained diffusion models guided by Hierarchical Equal Area isoLatitude Pixelization (HEALPix). We introduce HEALPix as a novel representation for panoramic illumination, providing a discrete and structured parameterization that supports uniform spherical sampling and retains high-frequency lighting variations. Based on this representation, we construct a conditional illumination diffusion model to generate out-of-view illumination content in a perceptually compressed LDR space. To support direct HDR output, we propose a reversible HDR compression strategy compatible with diffusion model training. Extensive experiments demonstrate that our Illumination Explorer generates HDR panoramas with high illumination accuracy and rich textural detail, outperforming previous methods in realistic composition for 3D objects with different reflective materials. Code is available at https://github.com/nauyihsnehs/IllumiExp. Zhongyun Bao, Shiyuan Shen, Xiangqian Shen, Chao Liang 0001, Chunxia Xiao |
IEEE Trans. Image Process. | 4 |
| 2026 | IDRetracor: Towards Visual Forensics against Malicious Face SwappingabstractThe deepfake-based face swapping technique poses significant risks to personal identity security. Although many detection methods have been proposed to counter malicious face swapping, they typically provide only binary labels (Fake/Real), lacking reliable and interpretable evidence. To address this limitation, we introduce a novel task called face retracing, which aims to visually trace back the original target face from a given fake one through inverse mapping. This task is based on the observation that current face swapping methods are neither flawless nor entirely random, leaving recoverable traces of the original identity. To this end, we propose IDRetracor, a model designed to recover arbitrary original target identities from fake faces generated by various face swapping techniques. Specifically, we first employ a mapping resolver to estimate the possible solution space of the original face for inverse mapping. Then, we introduce Mapping-Aware Convolutions (MACs), which consist of multiple dynamically combined kernels guided by the mapping resolver to adaptively handle diverse face swapping patterns. Extensive experiments demonstrate that IDRetracor achieves strong performance in retracing original faces, validated by both quantitative metrics and qualitative assessments. Jikang Cheng, Jiaxin Ai, Zhen Han 0002, Chao Liang 0001, Qin Zou 0001, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story VideosabstractVideo question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos, the most significant feature of story videos is storylines, which are composed of complex interactions and long-range evolvement of core story topics including characters, actions and locations. Understanding these topics requires models to possess DVU capability. However, existing DVU datasets rarely organize questions according to these story topics, making them difficult to comprehensively assess VideoQA models' DVU capability of complex storylines. Additionally, the question quantity and video length of these dataset are limited by high labor costs of handcrafted dataset building method. In this paper, we devise a large language model based multi-agent collaboration framework, StoryMind, to automatically generate a new large-scale DVU dataset. The dataset, FriendsQA, derived from the renowned sitcom Friends with an average episode length of 1,358 seconds, contains 44.6K questions evenly distributed across 14 fine-grained topics. Finally, We conduct comprehensive experiments on 10 state-of-the-art VideoQA models using the FriendsQA dataset. Zhengqian Wu, Ruizhe Li 0004, Zijun Xu 0001, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
AAAI | 6 |
| 2025 | Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense GameabstractMulti-exit neural networks represent a promising approach to enhancing model inference efficiency, yet like common neural networks, they suffer from significantly reduced robustness against adversarial attacks. While some defense methods have been raised to strengthen the adversarial robustness of multi-exit neural networks, we identify a long-neglected flaw in the evaluation of previous studies: simply using a fixed set of exits for attack may lead to an overestimation of their defense capacity. Based on this finding, our work explores the following three key aspects in the adversarial robustness of multi-exit neural networks: (1) we discover that a mismatch of the network exits used by the attacker and defender is responsible for the overestimated robustness of previous defense methods; (2) by finding the best strategy in a two-player zero-sum game, we propose AIMER as an improved evaluation scheme to measure the intrinsic robustness of multi-exit neural networks; (3) going further, we introduce NEED defense method under the evaluation of AIMER that can optimize the defender’s strategy by finding a Nash equilibrium of the game. Experiments over 3 datasets, 7 architectures, 7 attacks and 4 baselines show that AIMER evaluates the robustness 13.52% lower than previous methods under AutoAttack, while the robust performance of NEED surpasses single-exit networks of the same backbones by 5.58% maximally. Keyizhi Xu, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
CVPR | 6 |
| 2025 | Link-based Contrastive Learning for One-Shot Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) transfers knowledge from a labeled source domain to an unlabeled target domain via distribution alignment. However, in real-world scenarios like public safety or access control, obtaining sufficient source data is challenging, limiting existing UDA methods. This paper investigates a realistic but rarely studied problem called one-shot unsupervised domain adaptation (OSUDA), where only one source example per category is available. OSUDA faces dual challenges in feature learning and domain alignment due to the extreme source data scarcity. To address these, we propose link-based contrastive learning (LCL), a simple yet effective approach for OSUDA. LCL leverages in-domain links to learn discriminative features from abundant unlabeled target data and cross-domain links to achieve precise domain alignment with only one source sample per category. Extensive experiments on four domain adaptation benchmarks (VisDA-2017, Office-31, Office-Home, and DomainNet) demonstrate LCL’s effectiveness under the OSUDA setting. Additionally, we construct a real-world OSUDA surveillance face recognition dataset, where LCL consistently improves recognition performance across various face recognition methods. Yue Zhang 0104, Mingyue Bin, Zhongyuan Wang 0001, Zhen Han 0002, Chao Liang 0001 |
CVPR | 6 |
| 2025 | Visual Relation Diffusion for Human-Object Interaction Detection
Yepeng Tang, Chunjie Zhang 0001, Xiaolong Zheng 0001, Chao Liang 0001, Yunchao Wei, Yao Zhao 0001 |
ICCV | 5 |
| 2025 | T2CV-Zero: Zero-shot Character Video Generation Via Text-to-Motion Model
Huasong Han, Shiyuan Shen, Chao Liang 0001, Chunxia Xiao |
IJCNN | 5 |
| 2025 | Quantum Interference-Inspired Who-What-Where Composite-Semantics Instance Search for Story VideosabstractThe Who-What-Where (3W) composite-semantics video Instance Search (INS) task aims to find video shots about a person doing an action in a location. The state-of-the-art (SOTA) methods decompose 3W INS into three 2W INS, i.e., who-what, what-where and where-who semantic correlation modeling, and directly multiply three 2W INS results to produce the final 3W INS result. Obviously, overlapping semantics exist among the above 2Ws, e.g., who-what and what-where share the action component. The semantic overlap indicates that the 2Ws are mutually interdependent rather than independent. According to probability theory, the product of interdependent variables cannot be directly multiplied to obtain an accurate result, and such a direct product would yield a suboptimal outcome. This interdependence exerts diverse influences on the 3W INS results. For instance, fusing two 2W INS results ''Dr. Kelleher-provide medical guidance'' and ''provide medical guidance-in the hospital'', ''provide medical guidance'' is a pivotal connection, of positively enhancing the rationality of both person and location. Conversely, while both ''Ross-lifts heavy objects'' and ''lift heavy objects-Ross'' are individually coherent, combining them by overlapping the shared element ''Ross'' creates a conflict between the hazardous setting and strenuous labor, ultimately undermining the overall plausibility. Inspired by quantum interference theory, we propose a Quantum Interference Partial Decomposition (QIPD) method to model the diverse influences of semantic overlap from 2W to 3W INS. Specifically, QIPD incorporates two core modules, i.e., semantic interference and temporal interference. The former derives the 3W amplitude by converting 2W samples into amplitudes and phases and performing interference, while the latter sets the current shot's phase as baseline, amplifying the influence of adjacent shots while attenuating distant shots. Extensive evaluations on three large-scale 3W INS datasets demonstrate that QIPD outperforms SOTA baselines. Zijun Xu 0001, Chunjie Zhang 0001, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
ACM Multimedia | 6 |
| 2025 | A Robust 3D CNN with Pyramidal Attention for Spatiotemporal Gait RecognitionabstractGait recognition has become an increasingly important biometric technique for identifying individuals from a distance without requiring their active cooperation. Since gait involves a sequence of motion patterns, effectively capturing temporal dynamics is essential for accurate recognition. Traditional methods that extract temporal features independently and fuse them at a later stage often fail to model the continuity and interdependence of motion across frames. To overcome this limitation, we propose a novel three-dimensional convolutional architecture named Robust Spatiotemporal 3D Convolutional Neural Network (RST3D), which jointly captures spatial and temporal correlations throughout gait sequences. The proposed architecture incorporates a comprehensive 3D convolutional block that operates along the temporal, height, and width dimensions, enabling the network to learn more expressive and coherent spatiotemporal representations. In addition, we introduce a Temporal Pyramidal Attention (TPA) block to enhance the network’s ability to model temporal dependencies by capturing discriminative motion patterns across multiple temporal scales. We evaluate our method on four large-scale gait recognition datasets: CASIA-B, OUMVLP, GREW, and Gait3D. Experimental results show that our approach consistently achieves superior performance compared to existing 3D CNN-based methods, particularly under challenging conditions such as view variation, clothing changes, and occlusion. Jianyu Chen 0008, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Zengmin Xu, Gang Wu 0010, Zhongyuan Wang 0001 |
MMAsia | 4 |
| 2025 | Multi-Modal Gait Recognition via Collaborative Feature Learning from Silhouettes and Skeletons
Jianyu Chen 0008, Zhongyuan Wang 0001, Qian Zhou 0001, Qin Zou 0001, Chao Liang 0001, Gang Wu 0010 |
PRCV (15) | 5 |
| 2025 | Multiple Pedestrian Tracking Under Occlusion: A Survey and OutlookabstractAs an intermediate task in computer vision, multiple pedestrian tracking (MPT) aiming at tracking the pedestrians from a given video, has attracted attention due to its potential academic and commercial value. However, pedestrians commonly suffer from occlusion due to diverse and complex scenarios, which increases the challenge of this task. This survey provides comprehensive review in terms of occlusion scenarios encountered during MPT, and investigates the model robustness of the existing methods in this scenarios. Firstly, this survey introduces the various and states of occlusion. Secondly, the related occlusion datasets are introduced. Subsequently, we categorize existing occlusion handling methods according to the tracking process and detail their pros and cons. In addition, occlusion handling precision (OHP) metric is proposed to evaluate the ability of a tracker in handling occlusion in this survey. Moreover, comprehensive analyzes and discussions in several public datasets are provided to verify the effectiveness of these methods. Finally, the existing issues and future directions for occlusion handling methods are discussed. In doing so, this work serves as a foundation for future research by providing researchers with information about the occlusion handling method of MPT. Guoheng Wei, Mang Ye, Kui Jiang, Chao Liang 0001, Mithun Mukherjee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | ED4: Explicit Data-Level Debiasing for Deepfake DetectionabstractLearning intrinsic bias from limited data has been considered the main reason for the failure of deepfake detection with generalizability. Apart from the discovered content and specific-forgery bias, we reveal a novel spatial bias, where detectors inertly anticipate observing structural forgery clues appearing at the image center, also can lead to the poor generalization of existing methods. We present ED4, a simple and effective strategy, to address aforementioned biases explicitly at the data level in a unified framework rather than implicit disentanglement via network design. In particular, we develop ClockMix to produce facial structure preserved mixtures with arbitrary samples, which allows the detector to learn from an exponentially extended data distribution with much more diverse identities, backgrounds, local manipulation traces, and the co-occurrence of multiple forgery artifacts. We further propose the Adversarial Spatial Consistency Module (AdvSCM) to prevent extracting features with spatial bias, which adversarially generates spatial-inconsistent images and constrains their extracted feature to be consistent. As a model-agnostic debiasing strategy, ED4 is plug-and-play: it can be integrated with various deepfake detectors to obtain significant benefits. We conduct extensive experiments to demonstrate its effectiveness and superiority over existing deepfake detection approaches. Code is available at https://github.com/beautyremain/ED4. Jikang Cheng, Ying Zhang 0021, Qin Zou 0001, Zhiyuan Yan 0002, Chao Liang 0001, Zhongyuan Wang 0001, Chen Li 0031 |
IEEE Trans. Image Process. | 5 |
| 2025 | Who, What, and Where: Composite-Semantics Instance Search for Story VideosabstractWho, What and Where (3W)are the three core elements of storytelling, and accurately identifying the 3W semantics is critical to understanding the story in a video. This paper studies the 3W composite-semantics video Instance Search (INS) problem, which aims to find video shots about a specific person doing a concrete action in a particular location. The popular Complete-Decomposition (CD) methods divide a composite-semantics query into multiple single-semantics queries, which are likely to yield inaccurate or incomplete retrieval results due to neglecting important semantic correlations. Recent Non-Decomposition (ND) methods utilize Vision Language Model (VLM) to directly measure the similarity between textual query and video content. However, the accuracy is limited by VLM's immature capability to recognize fine-grained objects. To address the above challenges, we propose a video structure-aware Partial-Decomposition (PD) method. Its core idea is to partially decompose the 3W INS problem into three semantic-correlated 2W INS problems i.e., person-action INS, action-location INS, and location-person INS. Thereafter, we respectively model the correlations between pairs of semantics at frames, shots and scenes of story videos. With the help of the spatial consistency and temporal continuity contained in the unique hierarchical structure of story videos, we can finally obtain identity-matching, logic-consistent, and content-coherent 3W INS results. To validate the effectiveness of the proposed method, we specifically build three large-scale 3W INS datasets based on three TV series Eastenders, Friends and The Big Bang Theory, totally comprising over 670K video shots spanning 700 hours. Extensive experiments show that the proposed PD method surpasses the current state-of-the-art CD and ND methods for 3W INS in story videos. Ankang Lu, Zhengqian Wu, Zhongyuan Wang 0001, Chao Liang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Portrait Shadow Removal Using Context-Aware Illumination Restoration NetworkabstractPortrait shadow removal is a challenging task due to the complex surface of the face. Although existing work in this field makes substantial progress, these methods tend to overlook information in the background areas. However, this background information not only contains some important illumination cues but also plays a pivotal role in achieving lighting harmony between the face and the background after shadow elimination. In this paper, we propose a Context-aware Illumination Restoration Network (CIRNet) for portrait shadow removal. Our CIRNet consists of three stages. First, the Coarse Shadow Removal Network (CSRNet) mitigates the illumination discrepancies between shadow and non-shadow areas. Next, the Area-aware Shadow Restoration Network (ASRNet) predicts the illumination characteristics of shadowed areas by utilizing background context and non-shadow portrait context as references. Lastly, we introduce a Global Fusion Network to adaptively merge contextual information from different areas and generate the final shadow removal result. This approach leverages the illumination information from the background region while ensuring a more consistent overall illumination in the generated images. Our approach can also be extended to high-resolution portrait shadow removal and portrait specular highlight removal. Besides, we construct the first real facial shadow dataset for portrait shadow removal, consisting of 6200 pairs of facial images. Qualitative and quantitative comparisons demonstrate the advantages of our proposed dataset as well as our method. Jiangjian Yu, Ling Zhang 0017, Qing Zhang 0006, Daiguo Zhou, Chao Liang 0001, Chunxia Xiao |
IEEE Trans. Image Process. | 6 |
| 2025 | Multi-Stage Statistical Texture-Guided GAN for Tilted Face FrontalizationabstractExisting pose-invariant face recognition mainly focuses on frontal or profile, whereas high-pitch angle face recognition, prevalent under surveillance videos, has yet to be investigated. More importantly, tilted faces significantly differ from frontal or profile faces in the potential feature space due to self-occlusion, thus seriously affecting key feature extraction for face recognition. In this paper, we asymptotically reshape challenging high-pitch angle faces into a series of small-angle approximate frontal faces and exploit a statistical approach to learn texture features to ensure accurate facial component generation. In particular, we design a statistical texture-guided GAN for tilted face frontalization (STG-GAN) consisting of three main components. First, the face encoder extracts shallow features, followed by the face statistical texture modeling module that learns multi-scale face texture features based on the statistical distributions of the shallow features. Then, the face decoder performs feature deformation guided by the face statistical texture features while highlighting the pose-invariant face discriminative information. With the addition of multi-scale content loss, identity loss and adversarial loss, we further develop a pose contrastive loss of potential spatial features to constrain pose consistency and make its face frontalization process more reliable. On this basis, we propose a divide-and-conquer strategy, using STG-GAN to progressively synthesize faces with small pitch angles in multiple stages to achieve frontalization gradually. A unified end-to-end training across multiple stages facilitates the generation of numerous intermediate results to achieve a reasonable approximation of the ground truth. Extensive qualitative and quantitative experiments on multiple-face datasets demonstrate the superiority of our approach. Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Chao Liang 0001, Zhen Han 0002 |
IEEE Trans. Image Process. | 5 |
| 2025 | Floor Plan Restoration: A Multimodal Method Under One SecondabstractFloor plan restoration aims to recover vector and semantic information from raster floor plan images, which is significant for advanced applications including interior design, interative walkthroughs, and layout planning. Existing methods generally adopt a two-stage paradigm: a parsing stage to extract semantic elements such as rooms, walls, doors, and windows from raster images; and then a vectorization stage to convert them into vector graphics. However, these methods are deficient in both accuracy and efficiency due to the neglect of the unique cues of floor plans compared to natural images. To address the above issues, we propose MMParseNet that yields accurate parsing results by incorporating multimodal cues unique to floor plans, such as room names, furniture icons, and room boundaries. Moreover, we implement an efficiency-optimized vectorization method based on PCA that avoids unnecessary iterative solutions. We conduct both quantitative and qualitative experiments on three public and one self-built dataset. The results exhibit consistent improvements in accuracy and sub-second overall restoration time across various datasets. You-Ming Fu, Chunxia Xiao, Hai-Ming Xiang, Chao Liang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | CFDiffusion: Controllable Foreground Relighting in Image Compositing via Diffusion ModelabstractInserting foreground objects into specific background scenes and eliminating the illumination inconsistency (eg., color, brightness) between them is an important and challenging task. It typically involves multiple processing tasks, such as image harmonization and shadow generation. In these two domains, there are already many mature solutions, but they often only focus on one of the tasks. Recently, some image composition methods have utilized diffusion models to address both of these issues simultaneously, but they cannot guarantee complete reconstruction of the foreground content. In this work, we propose CFDiffusion, which can simultaneously handle image harmonization and shadow generation. We first employ a shadow mask predictor to estimate the shadow mask of the foreground object. Next, we design a harmonization-shadow generator based on a diffusion model to harmonize the foreground and generate shadows concurrently. Additionally, we propose a foreground content enhancement module to ensure the complete preservation of foreground content at the insertion location, and we also develop an adaptive encoder to guide the harmonization process in the foreground area. The experimental results on the iHarmony4 dataset and the IH-SG dataset demonstrate the superiority of our CFDiffusion approach. Zhongyun Bao, Gang Fu 0003, Weilei He, Chao Liang 0001, Chunxia Xiao |
ACM Multimedia | 6 |
| 2024 | Foreground Harmonization and Shadow Generation for Composite Image
Zhongyun Bao, Gang Fu 0003, Weilei He, Chao Liang 0001, Chunxia Xiao |
ACM Multimedia | 6 |
| 2024 | TalkSee: Interactive Video Retrieval Engine Using Large Language Model
Guihe Gu, Zhengqian Wu, Jiangshan He, Zhongyuan Wang 0001, Chao Liang 0001 |
MMM (4) | 6 |
| 2024 | SATCount: A scale-aware transformer-based class-agnostic counting framework
Bin Yang 0026, Chao Liang 0001, Jun Chen 0001 |
Neural Networks | 4 |
| 2024 | Unlabeled Data Assistant: Improving Mask Robustness for Face RecognitionabstractThe existing masked face recognition algorithms almost tend to adopt synthetic masked face datasets for training. However, these models are limited as they rely on existing mask augmentation methods, which contain few mask patterns and cannot simulate shadows and textures in realistic scenes. To overcome this limitation, we propose a semi-supervised face recognition framework to fully exploit unlabeled real masked face samples, improving the mask robustness of the recognition model. More specifically, unlike the original face embedding network, we design a part-aware network to explore multi-region face representation based on the face structure. In this way, we obtain multiple face sub-embeddings, which correspond to different regions of the face, including the upper half, the lower half and the whole. Crucially, we use the norm of the sub-embedding to represent the activation state of the facial region features. For the input unlabeled masked face image, we restrict the sub-embedding norm of its lower half to weaken the face feature representation of the occluded area. For normal face samples, their partial features are kept activated by maintaining the sub-embedding norm, which guides the deep network does not ignore the available information. Moreover, we employ the margin-based recognition loss for normal samples to ensure that the model is sufficiently discriminative for normal facial features. Extensive experimental results on both normal and real masked face datasets show that our approach significantly outperforms the state-of-the-arts. Code is available at https://github.com/Baojin-Huang/UFace. Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Zhen Han 0002, Chao Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Toward Robust Adversarial Purification for Face Recognition Under Intensity-Unknown AttacksabstractRecent years have witnessed dramatic progress in adversarial attacks, which can easily mislead face recognition systems via the injection of imperceptible perturbations on the input image. Many defense methods have been proposed to mitigate the detrimental impact of adversarial attacks, including adversarial purification which intends to reconstruct clean images through a generative model. This paper studies a more practical and challenging problem: how to defend face recognition systems against intensity-unknown or even intensity-varying adversarial attacks? We attempt to crack this tough nut from the dimensionality of input resolutions. Looking into the performance of purification methods with various input resolutions, we reveal a phenomenon that, higher-resolution input images help better defend against weaker attacks, while lower-resolution ones are naturally defensive against stronger attacks. It inspires us to design an adaptive purification framework under intensity-unknown attacks, dubbed adversarial Intensity-guided Multi-scale Attention (IMA). Via the aggregation of information from different resolution scales and flexible adjustment according to an estimation of adversarial intensity, it leverages the respective advantages of different scales and constructs a robust ensemble against intensity-unknown attacks. We validate the superiority of IMA by defending against both face obfuscation and impersonation of 9 typical attack algorithms under gray-box, white-box and black-box evaluation, outperforming state-of-the-art defense methods on LFW and YTF datasets. Keyizhi Xu, Zhongyuan Wang 0001, Chunxia Xiao, Chao Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Person-action Instance Search in Story Videos: An Experimental StudyabstractPerson-Action instance search (P-A INS) aims to retrieve the instances of a specific person doing a specific action, which appears in the 2019–2021 INS tasks of the world-famous TREC Video Retrieval Evaluation (TRECVID). Most of the top-ranking solutions can be summarized with a Division-Fusion-Optimization (DFO) framework, in which person and action recognition scores are obtained separately, then fused, and, optionally, further optimized to generate the final ranking. However, TRECVID only evaluates the final ranking results, ignoring the effects of intermediate steps and their implementation methods. We argue that conducting the fine-grained evaluations of intermediate steps of DFO framework will (1) provide a quantitative analysis of the different methods’ performance in intermediate steps; (2) find out better design choices that contribute to improving retrieval performance; and (3) inspire new ideas for future research from the limitation analysis of current techniques. Particularly, we propose an indirect evaluation method motivated by the leave-one-out strategy, which finds an optimal solution surpassing the champion teams in 2020–2021 INS tasks. Moreover, to validate the generalizability and robustness of the proposed solution under various scenarios, we specifically construct a new large-scale P-A INS dataset and conduct comparative experiments with both the leading NIST TRECVID INS solution and the state-of-the-art P-A INS method. Finally, we discuss the limitations of our evaluation work and suggest future research directions. Yanrui Niu, Chao Liang 0001, Ankang Lu, Baojin Huang, Zhongyuan Wang 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Learning Degradation for Real-World Face Super-Resolution
Jun Chen 0001, Dongshu Xu, Chao Liang 0001, Zhen Han 0002 |
CGI | 5 |
| 2023 | Promoting adversarial transferability with enhanced loss flatnessabstractCarefully crafted small perturbations, when added to an image, can mislead the deep neural networks to give wrong outputs. Such mischievous images are called adversarial examples. Transfer-based black-box attacks use a surrogate white-box model to generate adversarial examples which can be transferred and attack black-box models with little known information. We propose to increase the transferability of adversarial examples by smoothing the geometric surface of loss function at the adversarial example point. By looking ahead the optimization path for a few steps, we define a future geometric vicinity using the integration of neighbourhood of those predicted data points. By sampling in this area and using the summation of gradients at those sampled data points for optimization, our method avoids local fluctuation of loss function. Experiments on ImageNet validation dataset show that our method outperforms state-of-the-art attacks by a large margin. Zhongyuan Wang 0001, Jikang Cheng, Chao Liang 0001 |
ICME | 5 |
| 2023 | Who, What and Where: Composite-semantic Instance Search for Story VideosabstractThis paper studies Who-What-Where (3W) composite-semantic video instance search (INS) problem, which aims to find a specific person doing a queried action in a particular place. Mainstream approaches adopt a complete decomposition strategy, which divides a composite-semantic query into multiple single-semantic queries. However, due to the lack of necessary correlation analysis among constituent semantics, these methods cannot always generate identity-matching and semantics-consistent 3W INS results. To address the above challenges, we propose a partial decomposition scheme with action as the link. Specifically, we selectively split the 3W INS as person-action INS and action-location INS. The former ensures the retrieved person and action share the same identity by modeling their relative spatial positions at the frame level, while the latter improves the semantic consistency between action and location with a cross-semantic attention mechanism at the shot level. Particularly, we build a large-scale 3W INS dataset, containing over 470k video shots, on basis of NIST TRECVID 2016-2021 INS tasks and verify the effectiveness of the proposed method with both quantitative and qualitative experiments. Chao Liang 0001, Zhongyuan Wang 0001 |
ICME | 2 |
| 2023 | A Hierarchical Deep Video Understanding Method with Shot-Based Instance Search and Large Language ModelabstractDeep video understanding (DVU) is often considered a challenge due to the aim of interpreting a video with storyline, which is designed to solve two levels of problems: predicting the human interaction in scene-level and identifying the relationship between two entities in movie-level. Based on our understanding of the movie characteristics and analysis of DVU tasks, in this paper, we propose a four-stage method to solve the task, which includes video structuring, shot based instance search, interaction & relation prediction and shot-scene summary & Question Answering (QA) with ChatGPT. In these four stages, shot based instance search allows accurate identification and tracking of characters at an appropriate video granularity. Using ChatGPT in QA, on the one hand, can narrow the answer space, on the other hand, with the help of the powerful text understanding ability, ChatGPT can help us answer the questions by giving background knowledge. We rank first in movie-level group 2 and scene-level group 1, second in movie-level group 1 and scene-level group 2 in ACM MM 2023 Grand Challenge. Ruizhe Li 0004, Zhengqian Wu, Chao Liang 0001 |
ACM Multimedia | 5 |
| 2023 | A Spatio-Temporal Identity Verification Method for Person-Action Instance Search in Movies
Yanrui Niu, Jingyao Yang, Chao Liang 0001, Baojin Huang, Zhongyuan Wang 0001 |
MMM (1) | 3 |
| 2023 | Floor Plan Analysis and Vectorization with Multimodal Information
Chao Liang 0001, You-Ming Fu, Chunxia Xiao, Hai-Ming Xiang |
MMM (1) | 2 |
| 2023 | A Survey of Personalized Interior DesignabstractAbstract Interior design is the core step of interior decoration, and it determines the overall layout and style of furniture. Traditional interior design is usually laborious and time‐consuming work carried out by professional designers and cannot always meet clients' personalized requirements. With the development of computer graphics, computer vision and machine learning, computer scientists have carried out much fruitful research work in computer‐aided personalized interior design (PID). In general, personalization research in interior design mainly focuses on furniture selection and floor plan preparation. In terms of the former, personalized furniture selection is achieved by selecting furniture that matches the resident's preference and style, while the latter allows the resident to personalize their floor plan design and planning. Finally, the automatic furniture layout task generates a stylistically matched and functionally complete furniture layout result based on the selected furniture and prepared floor plan. Therefore, the main challenge for PID is meeting residents' personalized requirements in terms of both furniture and floor plans. This paper answers the above question by reviewing recent progress in five separate but correlated areas, including furniture style analysis, furniture compatibility prediction, floor plan design, floor plan analysis and automatic furniture layout. For each topic, we review representative methods and compare and discuss their strengths and shortcomings. In addition, we collect and summarize public datasets related to PID and finally discuss its future research directions. Y. T. Wang, Chao Liang 0001, N. Huai, C. J. Zhang |
Comput. Graph. Forum | 2 |
| 2023 | Transferring fashion to surveillance with weak labels
Zheng He 0001, Chao Liang 0001, Jun Chen 0001, Chia-Wen Lin, Dapeng Tao |
Neural Comput. Appl. | 3 |
| 2023 | PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001 |
Pattern Recognit. | 7 |
| 2023 | HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002 |
Pattern Recognit. | 5 |
| 2023 | Confidence-Aware Active Feedback for Interactive Instance SearchabstractOnline relevance feedback (RF) is widely utilized in instance search (INS) tasks to further refine imperfect ranking results, but it often has low interaction efficiency. The active learning (AL) technique addresses this problem by selecting valuable feedback candidates. However, mainstream AL methods require an initial labeled set for a cold start and are often computationally complex to solve. Therefore, they cannot fully satisfy the requirements for online RF in interactive INS tasks. To address this issue, we propose a confidence-aware active feedback method (CAAF) that is specifically designed for online RF in interactive INS tasks. Inspired by the explicit difficulty modeling scheme in self-paced learning, CAAF utilizes a pairwise manifold ranking loss to evaluate the ranking confidence of each unlabeled sample. The ranking confidence improves not only the interaction efficiency by indicating valuable feedback candidates but also the ranking quality by modulating the diffusion weights in manifold ranking. In addition, we design two acceleration strategies, an approximate optimization scheme and a top-$K$search scheme, to reduce the computational complexity of CAAF. Extensive experiments on both image INS tasks and video INS tasks searching for buildings, landscapes, persons, and human behaviors demonstrate the effectiveness of the proposed method. Notably, in the real-world, large-scale video INS task of NIST TRECVID 2021, CAAF uses 25% fewer feedback samples to achieve a performance that is nearly equivalent to the champion solution. Moreover, with the same number of feedback samples, CAAF's mAP is 51.9%, significantly surpassing the champion solution by 5.9%. Code is available athttps://github.com/nercms-mmap/caaf. Yue Zhang 0104, Chao Liang 0001, Longxiang Jiang |
IEEE Trans. Multim. | 2 |
| 2022 | Ranking Aggregation with Interactive Feedback for Collaborative Person Re-identification
Chao Liang 0001, Yue Zhang 0104, Zhongyuan Wang 0001, Chunjie Zhang 0001 |
BMVC | 2 |
| 2022 | Face Super-Resolution with Better Semantics and More Efficient Guidance
Jun Chen 0001, Zheng Wang 0007, Chao Liang 0001, Zhen Han 0002, Chia-Wen Lin |
CGI | 4 |
| 2022 | Face hallucination based on degradation analysis for robust manifold
Ruimin Hu, Zheng He 0001, Chao Liang 0001, Zhongyuan Wang 0001 |
Neurocomputing | 4 |
| 2022 | Online multiple object tracking based on fusing global and partial features
Jun Chen 0001, Mithun Mukherjee 0001, Chao Liang 0001, Weijian Ruan |
Neurocomputing | 4 |
| 2022 | TICNet: A Target-Insight Correlation Network for Object TrackingabstractRecently, the correlation filter (CF) and Siamese network have become the two most popular frameworks in object tracking. Existing CF trackers, however, are limited by feature learning and context usage, making them sensitive to boundary effects. In contrast, Siamese trackers can easily suffer from the interference of semantic distractors. To address the above problems, we propose an end-to-end target-insight correlation network (TICNet) for object tracking, which aims at breaking the above limitations on top of a unified network. TICNet is an asymmetric dual-branch network involving a target-background awareness model (TBAM), a spatial-channel attention network (SCAN), and a distractor-aware filter (DAF) for end-to-end learning. Specifically, TBAM aims to distinguish a target from the background in the pixel level, yielding a target likelihood map based on color statistics to mine distractors for DAF learning. SCAN consists of a basic convolutional network, a channel-attention network, and a spatial-attention network, aiming to generate attentive weights to enhance the representation learning of the tracker. Especially, we formulate a differentiable DAF and employ it as a learnable layer in the network, thus helping suppress distracting regions in the background. During testing, DAF, together with TBAM, yields a response map for the final target estimation. Extensive experiments on seven benchmarks demonstrate that TICNet outperforms the state-of-the-art methods while running at real-time speed. Weijian Ruan, Mang Ye, Yi Wu 0001, Wu Liu 0005, Jun Chen 0001, Chao Liang 0001, Ge Li 0002, Chia-Wen Lin |
IEEE Trans. Cybern. | 6 |
| 2022 | Exemplar-Based, Semantic Guided Zero-Shot Visual RecognitionabstractZero-shot recognition has been a hot topic in recent years. Since no direct supervision is available, researchers use semantic information as the bridge instead. However, most zero-shot recognition methods jointly model images on the class level without considering the distinctive character of each image. To solve this problem, in this paper, we propose a novel exemplar-based, semantic guided zero-shot recognition method (EBSG). Both visual and semantic information of each image is used. We train visual sub-model to separate each image from the other images of different classes. We also train semantic sub-model to separate this image from the other images described with different semantics. We concatenate the outputs of visual and semantic sub-models to represent images. Image classification model is then learned by measuring visual similarity and semantic consistency of both source and target images. We conduct zero-shot recognition experiments on four widely used datasets. Experimental results show the effectiveness of the proposed EBSG method. Chunjie Zhang 0001, Chao Liang 0001, Yao Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Long-short Term Prediction for Occluded Multiple Object TrackingabstractOnline multiple object tracking (MOT) is a challenging problem in complex scenes due to frequent occlusions. Most of the existing MOT methods tend to focus on addressing an individual type of occlusion, which cannot meet the requirements of real complex scenes. In this paper, we propose a unified MOT framework that combines long- and short-term prediction models for online multiple object tracking. Basically, The short-term prediction model consists of an appearance-based model and a motion-based model, aiming at exploiting the appearance and motion of objects to handle different types of occlusions jointly. Furthermore, we adopt a cubic spline interpolation as a long-term prediction model to estimate the trajectory of the target in occluded frames. To handle different lengths of occlusions, an adaptive weighted fusion model is proposed to combine the short-term prediction model, and the long-term prediction model. Experimental results on several challenging datasets demonstrate that the proposed method outperforms state-of-the-art methods. Jun Chen 0001, Mithun Mukherjee 0001, Weijian Ruan, Chao Liang 0001, Yi Yu 0001 |
GLOBECOM | 5 |
| 2021 | Person Retrieval in Physical WorldabstractPerson re-identification (re-ID) gains plenty of achievements as a retrieval problem in constrained camera networks. However, most of the researches are concentrated on visual appearance, they still suffer from the complicated environments in unconstrained urban/campus surveillance scenario due to unreliable visual representations with extremely challenging problems as lack of training samples, amounts of irrelevant crowds, etc. Besides, most of the existing person re-ID datasets neglect the physical truth of realistic investigation application: 1) investigators search only few suspects among amounts of crowds. Moreover, he may not go through by every camera in the surveillance area and may appear in the same camera several times; and 2) the corresponding characteristic in multi-space of the same ID can be verified with each other. Therefore, we propose a person retrieval in physical world (PRPW) dataset with large-scale unconstrained surveillance scenario. It contains over 1.4 million bounding boxes, including 20 labeled IDs and numerous irrelevant crowds captured by 86 cameras. Furthermore, over 30,000 records of 20 mobile trajectories are collected in this dataset, and the 20 mobile trajectories are partially overlapped while passing by 86 cameras. Finally, based on two common senses and a verification experiment, we provide a proposal to tackle with PRPW task on the basis of trajectory association which utilizes global optimization to compensate for the errors caused by visual expression on local observation points. The comparison experiments with two typical unsupervised person re-ID methods are implemented on the constructed dataset. Wenxin Huang, Ruimin Hu, Chao Liang 0001, Xian Zhong |
ICME | 4 |
| 2021 | Occluded suspect search via channel-guided mechanism
Wenxin Huang, Ruimin Hu, Xiao Wang 0029, Chao Liang 0001, Jun Chen 0001 |
Neural Comput. Appl. | 4 |
| 2021 | A Survey of Multiple Pedestrian Tracking Based on Tracking-by-Detection FrameworkabstractMultiple pedestrian tracking (MPT) has gained significant attention due to its huge potential in a commercial application. It aims to predict multiple pedestrian trajectories and maintain their identities, given a video sequence. In the past decade, due to the advancement in pedestrian detection algorithms, Tracking-by-Detection (TBD) based algorithms have achieved tremendous successes. TBD has become the most popular MPT framework, and it has been actively studied in the past decade. In this paper, we give a comprehensive survey of recent advances in TBD-based MPT algorithms. We systematically analyze the existing TBD-based algorithms and organize the survey into four major parts. At first, this survey draws a timeline to introduce the milestones of TBD-based works which briefly reviews the development of the existing TBD-based methods. Second, the main procedures of the TBD framework are summarized, and each stage in the procedure is described in detail. Afterward, this survey analyzes the performance of existing TBD-based algorithms on MOT challenge datasets and discusses the factors that affect tracking performance. Finally, open issues and future directions in the TBD framework are discussed. Jun Chen 0001, Chao Liang 0001, Weijian Ruan, Mithun Mukherjee 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Expression-Aware Face Reconstruction via a Dual-Stream NetworkabstractRecently, 3D face reconstruction from a single image has achieved promising progress by adopting the 3D Morphable Model (3DMM). However, face images taken in-the-wild usually involve expressions with a large range of variety. This poses difficulty to use 3DMM to represent such various facial expressions owing to the limited expressive ability of its linear model, thereby resulting in distortion and ambiguity in local facial regions. To tackle this problem, we present a novel dual-stream network composed of a geometry stream and a texture stream to deal with expression variations. Specifically, in the geometry stream, we propose novel Attribute Spatial Maps (ASMs) to decompose a face into the identity and expression attributes and then separately record the essential spatial information of the two facial attributes in the 2D image space. This avoids the interaction between the two attributes, thus preserving the identity information and further improving the ability of coping with expression variations. In the texture stream, we propose to generate facial appearance with realistic texture and canonical layout by our Semantic Region Stylization Mechanism (SRSM), that transfers the style from an input face to a 3DMM albedo map in a region-adaptive manner. Moreover, we also propose a Shared Semantic Region Prediction Module (SSRPM) to explore the common correspondence of semantic regions between the above two face texture representations. Both quantitative and qualitative evaluations on public datasets demonstrate the effectiveness of our approach in face reconstruction under expression variations. Xiaoyu Chai, Jun Chen 0001, Chao Liang 0001, Dongshu Xu, Chia-Wen Lin |
IEEE Trans. Multim. | 3 |
| 2021 | Correlation Discrepancy Insight Network for Video Re-identificationabstractVideo-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods. Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Manifold Projection for Adversarial Defense on Face Recognition
Jianli Zhou, Chao Liang 0001, Jun Chen 0001 |
ECCV (30) | 2 |
| 2020 | Crowdsourcing-Based Ranking Aggregation for Person Re-IdentificationabstractPerson re-identification (re-ID) is widely applied in surveillance and criminal detection applications. The existing research focus on devising the stand-alone re-ID methods, ignoring their practical application in the multi-person collaboration scenario. To improve the search efficiency, a group of investigators are usually assigned the same task to re-identify a suspect from a shared gallery set. Due to their personalized viewpoints and search feedback operations, different investigators may obtain diverse search results of the same query target. In this case, merging different rankings and generating an improved result is of great importance. To this end, this paper proposes a crowdsourcing-based ranking aggregation to adaptively fuse multiple ranking lists for re-ID problem. The method estimates the reliability of individual investigators, with a specifically designed long tail distribution to fit the top ranking demand, and is feasible for human-machine interaction. Extensive experiments conducted on four dataset-s demonstrate the superiority of the proposed method. Yinxue Yu, Chao Liang 0001, Weijian Ruan, Longxiang Jiang |
ICASSP | 2 |
| 2020 | Expression-Aware Face Reconstruction Via A Dual-Stream NetworkabstractRecently, 3D face reconstruction from a single image has achieved promising results by adopting the 3D Morphable Model (3DMM). However, as face images in-the-wild have various expressions, it is difficult for 3DMM to handle diverse facial expressions with a large range of variations, due to the limited expressive ability of its linear model, thereby resulting in distortion and ambiguity on facial local regions. To tackle this issue, we present a novel dual-stream network to deal with expression variations. Specifically, in the geometry stream, we propose novel Attribute Spatial Maps to record the spatial information of facial identity and expression attributes in the 2D image space separately. This avoids the interaction between the two attributes, thus keeping the identity information and further improving the ability to cope with expression changes. In the texture stream, we utilize the 3DMM albedo map to a style transfer based method for synthesizing facial appearance, which results in expression-irrelevant as well as realistic face textures. Both quantitative and qualitative evaluations on public datasets demonstrate the ability of our approach to achieve comparable results in face reconstructions under expression variations. Xiaoyu Chai, Jun Chen 0001, Chao Liang 0001, Dongshu Xu, Chia-Wen Lin |
ICME | 3 |
| 2020 | When Pedestrian Detection Meets Nighttime Surveillance: A New BenchmarkabstractPedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising performances. In contrast, they often fail under unstable lighting conditions (e.g. nighttime). Night is a critical time for criminal suspects to act in the field of security. The existing nighttime pedestrian detection dataset is captured by a car camera, specially designed for autonomous driving scenarios. The dataset for nighttime surveillance scenario is still vacant. There are vast differences between autonomous driving and surveillance, including viewpoint and illumination. In this paper, we build a novel pedestrian detection dataset from the nighttime surveillance aspect: NightSurveillance1. As a benchmark dataset for pedestrian detection at nighttime, we compare the performances of state-of-the-art pedestrian detectors and the results reveal that the methods cannot solve all the challenging problems of NightSurveillance. We believe that NightSurveillance can further advance the research of pedestrian detection, especially in the field of surveillance security at nighttime. Xiao Wang 0029, Jun Chen 0001, Zheng Wang 0007, Wu Liu 0005, Shin'ichi Satoh 0001, Chao Liang 0001, Chia-Wen Lin |
IJCAI | 6 |
| 2020 | One-Shot Face Recognition with Feature Rectification via Adversarial Learning
Jianli Zhou, Jun Chen 0001, Chao Liang 0001 |
MMM (1) | 3 |
| 2020 | SIST: Online Scale-Adaptive Object tracking with Stepwise Insight
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Jun Chen 0001, Ruimin Hu |
Neurocomputing | 2 |
| 2020 | Single image de-raining via clique recursive feedback mechanism
Jun Chen 0001, Kui Jiang, Zhen Han 0002, Weijian Ruan, Zhongyuan Wang 0001, Chao Liang 0001 |
Neurocomputing | 7 |
| 2020 | Generating video animation from single still image in social media based on intelligent computing
Tao Hu 0012, Chao Liang 0001, Geyong Min, Keqin Li 0001, Chunxia Xiao |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Identity-Aware Face Super-Resolution for Low-Resolution Face RecognitionabstractAlthough deep learning-based face recognition techniques have achieved amazing performance in recent years, low-resolution (LR) face recognition remains challenging. In this letter, we address this problem by proposing an identity-aware face super-resolution network to recover identity information of LR faces. To learn identity-aware features effectively, the identity features are explicitly disentangled to two orthogonal components: the magnitude and angle of features that project identity features to a hypersphere space. We show that the magnitude of features is related to the quality of a face. The proposed approach shows its superiority on recovering identity-related textures which are beneficial to recover identity information for recognition. Extensive experiments demonstrate the effectiveness of the proposed algorithm in LR face recognition. Jun Chen 0001, Zheng Wang 0007, Chao Liang 0001, Chia-Wen Lin |
IEEE Signal Process. Lett. | 4 |
| 2020 | Modeling and Optimizing of the Multi-Layer Nearest Neighbor Network for Face Image Super-ResolutionabstractIn this paper, we propose a face super-resolution (FSR) method to handle the decreasing face recognition rate caused by low-quality images. To better model the input images, we build a nearest neighbor network (NNN) which consists of nodes and paths by introducing the second-layer nearest neighbors (SLNNs), where the paths of the network represent the distance between nodes. As the SLNN is trained in the high-resolution (HR) space and is exponentially supplementary to the traditional first-layer nearest neighbors (FLNNs), the neighbor inadequacy problem can be effectively solved by enriching the neighbor candidate set via NNN. Furthermore, we solve the NNN for the optimal weights of neighbors. Finally, we fuse the refined weights and neighbors for better reconstruction results. The effectiveness of this fusion strategy is validated by both quantitative and qualitative experimental results. The extensive experimental results on the public face datasets and real-world challenging low-resolution (LR) images demonstrate that the proposed method performs favorably against the state-of-the-art methods. Liang Chen 0026, Jinshan Pan, Ruimin Hu, Zhen Han 0002, Chao Liang 0001, Yi Wu 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | S3D: Scalable Pedestrian Detection via Score Scale Surface DiscriminationabstractPedestrian detection has remained an important research topic in both the computer vision and multimedia communities because of its importance in practical applications, such as driving assistance and video surveillance. Existing methods compare the response score with a fixed threshold to determine whether a candidate region contains pedestrians and produce dissatisfactory results that contain either missed detections or false detections, which are difficult to balance. This situation has a serious impact under the condition of variable scale. This paper investigates the functional relationship between the scores and scales of pedestrians. By designing experiments with multiple scales, we have found a discriminant surface in the score scale space. Pedestrians can be distinguished at various scale levels according to their locations on the discriminant surface. The proposed approach is evaluated using four challenging pedestrian detection datasets, including Caltech, INRIA, ETH, and KITTI, and the superior experimental results are achieved when compared with baseline methods. Xiao Wang 0029, Chao Liang 0001, Chen Chen 0001, Jun Chen 0001, Zheng Wang 0007, Zhen Han 0002, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Rain Streak Removal via Multi-scale Mixture Exponential Power ModelabstractRain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GMM compromises the ability of the model in fitting real noise, such as sparse noise, which is exactly the characteristic of the rain streaks. In this paper, a novel model named Mixture Exponential Power Model (MEPM) is exploited. It sets multiple Laplace noise components and expands the representation capability for the sparse noise. Moreover, considering that the rain streaks in a video occur in different distances from the camera, we encode rain streaks into Multi-scale Mixture Exponential Power Model. The model is opti-mized by expectation-maximization (EM) algorithm and La-grange multiplier strategy. Experiments are implemented on synthetic and real rain videos and verify the superiority of the proposed method, compared with state-of-the-art methods. Jun Chen 0001, Zhen Han 0002, Mingfu Xiong, Chao Liang 0001, Zhongyuan Wang 0001 |
ICASSP | 5 |
| 2019 | Cross-view Identical Part Area Alignment for Person Re-identificationabstractPerson re-identification aims to associate images captured by non-overlapping cameras. It is a challenging task because images are often in different conditions such as background clutter, illumination variation, viewpoint changes and different camera settings. Viewpoint changes and pose variations often cause body part self-occlusion and misalignment. To deal with the problem, local features from human body parts are extracted. However, with viewpoint changes, the body parts also rotate horizontally. It is inappropriate to extract feature from entire area of body parts directly because the visible surface of body parts would turn away if viewpoint changes. Comparing identical areas provides a new way to pay attention to the details of person images. In this paper, we propose a Rotation Invariant Network to find the identical areas in cross-view images to extract robust local features. Extensive experiment show the effectiveness of our method on public datasets including CUHK03, Market1501 and DukeMTMC. Dongshu Xu, Jun Chen 0001, Chao Liang 0001, Zheng Wang 0007, Ruimin Hu |
ICASSP | 3 |
| 2019 | Multi-Similarity Re-Ranking for Person Re-IdentificationabstractRe-ranking has been proved an effective method to boost the performance of person re-identification. Existing works focus on contextual or graph-based similarity to improve the initial ranking result. The former mainly concentrates on more accurate similarity description but neglects the manifold constraint. While, the later centers on solving similarities with manifold constraint, which acquires several accurate top ranks. In this paper, we propose a novel method which not only takes contextual similarity to generate top ranks accurately but also refines the ranks based on graph-based similarity. Specifically, given initial Euclidean distances between a probe and galleries, we mine contextual and graph-based similarities respectively and then re-rank all galleries with a diffusion procedure under constraints of both similarities. Experiments on two person re-ID datasets demonstrate that our method outperforms state-of-the-art re-ranking approaches in person re-identification. Longxiang Jiang, Chao Liang 0001, Dongshu Xu, Wenxin Huang |
ICIP | 2 |
| 2019 | Illumination animating and editing in a single picture using scene structure estimation
Bin Liao 0006, Chao Liang 0001, Fei Luo 0004, Chunxia Xiao |
Comput. Graph. | 3 |
| 2019 | Person re-identification with multiple similarity probabilities using deep metric learning for efficient smart security applications
Mingfu Xiong, Dan Chen 0001, Jun Chen 0001, Jingying Chen 0001, Benyun Shi, Chao Liang 0001, Ruimin Hu |
J. Parallel Distributed Comput. | 6 |
| 2019 | Multi-Correlation Filters With Triangle-Structure Constraints for Object TrackingabstractCorrelation filters (CFs) have been extensively used in tracking tasks due to their high efficiency although most of them regard the tracked target as a whole and are minimally effective in handling partial occlusion. In this study, we incorporate a part-based strategy into the framework of CFs and propose a novel multipart correlation tracker with triangle-structure constraints. Specifically, we train multiple CFs for the global object and local parts, which are then jointly applied to obtain the correlation response of any candidate during tracking. The tracker is robust in handling partial occlusion because of the use of part-based representation. The remaining global representation can contribute reliable cues in cases wherein several local filters drift away in a specific scene. We further propose a triangle-structure model to measure the structural similarity of candidates. The model employs multiple triangles to determine the spatial relationship among parts and helps constrain the location of the target. Moreover, we introduce an effective part selection scheme based on energy and integrity, which is generally applicable to part-tracking models. Extensive experiments on two public benchmarks demonstrate the superiority of the proposed method over the state-of-the-art approaches. Weijian Ruan, Jun Chen 0001, Yi Wu 0001, Jinqiao Wang, Chao Liang 0001, Ruimin Hu, Junjun Jiang |
IEEE Trans. Multim. | 5 |
| 2018 | Video-Based Person Re-Identification via Self Paced WeightingabstractPerson re-identification (re-id) is a fundamental technique to associate various person images, captured by differentsurveillance cameras, to the same person. Compared to the single image based person re-id methods, video-based personre-id has attracted widespread attentions because extra space-time information and more appearance cues that can beused to greatly improve the matching performance. However, most existing video-based person re-id methods equally treatall video frames, ignoring their quality discrepancy caused by object occlusion and motions, which is a common phenomenonin real surveillance scenario. Based on this finding, we propose a novel video-based person re-id method via self paced weighting (SPW). Firstly, we propose a self paced outlier detection method to evaluate the noise degree of video sub sequences. Thereafter, a weighted multi-pair distance metric learning approach is adopted to measure the distance of two person image sequences. Experimental results on two public datasets demonstrate the superiority of the proposed method over current state-of-the-art work. Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Weijian Ruan, Ruimin Hu |
AAAI | 2 |
| 2018 | Faster Seam Carving for Video RetargetingabstractVideo retargeting is to resize a video to a desired resolution or aspect ratio while preserving its salient content without visual distortion. The key to video retargeting is to reconcile spatio-temporal coherence of video frames, and most existing works use seam carving to achieve that by employing the dynamic programming to find optimal seams. However, these methods are too time-consuming due to high computational complexity of the dynamic programming. To this end, we propose a novel method which uses discontinuous and suboptimal seams for seam carving. Concretely, we obtain the discontinuous seams by allowing seams to move freely in homogeneous regions of the frame, which helps preserve the spatio-temporal coherence effectively. Then, the genetic algorithm is employed to find suboptimal seams, so as to reduce computational complexity. Finally, each frame can be retargeted to a new aspect ratio or size by repeatedly carving out seams. Compared to the-state-of-the-art methods, the proposed algorithm achieves comparable results at an average expense of only one third of their running time. Ruimin Hu, Chao Liang 0001, Chunxia Xiao, Weijian Ruan |
ICIP | 3 |
| 2018 | Equidistance constrained metric learning for person re-identification
Jin Wang 0019, Zheng Wang 0007, Chao Liang 0001, Changxin Gao, Nong Sang |
Pattern Recognit. | 3 |
| 2018 | Image Class Prediction by Joint Object, Context, and Background ModelingabstractState-of-the-art image classification methods often use spatial pyramid matching or its variants to make use of the spatial layout of visual features. However, objects may appear at various places with different scales and orientations. Besides, traditionally object-centric-based methods only consider objects and the background without fully exploring the context information. To solve these problems, in this paper we propose a novel image classification method by jointly modeling the object, context, and background information (OCB). OCB consists of three components: 1) locate the positions of objects; 2) determine the context areas of objects; and 3) treat the other areas as the background. We use objectness proposal techniques to select candidate bounding boxes. Boxes with high confidence scores are combined to determine objects' positions. To select the context areas, we use candidate boxes that have relatively lower confidence scores compared with boxes for object location selection. The other areas are viewed as the background. We jointly combine the object, context, and background for image representation and classification. Experiments on six data sets well demonstrate the superiority of the proposed OCB method over other spatial partition methods. Chunjie Zhang 0001, Guibo Zhu, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Person Reidentification via Discrepancy Matrix and Matrix MetricabstractPerson reidentification (re-id), as an important task in video surveillance and forensics applications, has been widely studied. Previous research efforts toward solving the person re-id problem have primarily focused on constructing robust vector description by exploiting appearance's characteristic, or learning discriminative distance metric by labeled vectors. Based on the cognition and identification process of human, we propose a new pattern, which transforms the feature description from characteristic vector to discrepancy matrix. In particular, in order to well identify a person, it converts the distance metric from vector metric to matrix metric, which consists of the intradiscrepancy projection and interdiscrepancy projection parts. We introduce a consistent term and a discriminative term to form the objective function. To solve it efficiently, we utilize a simple gradient-descent method under the alternating optimization process with respect to the two projections. Experimental results on public datasets demonstrate the effectiveness of the proposed pattern as compared with the state-of-the-art approaches. Zheng Wang 0007, Ruimin Hu, Chen Chen 0001, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Shin'ichi Satoh 0001 |
IEEE Trans. Cybern. | 6 |
| 2017 | Taichi distance for person re-identificationabstractMetric learning is an important issue in person re-identification, and Mahalanobis-distance based metric learning methods prevail in this field. All of these approaches can be considered as equivalently projecting all samples to a new metric space and calculating the Euclidean distance there. However, the performance of distinguishing similar samples from dissimilar ones via absolute distance is limited. In this paper, we suggest using relative distance instead. We adopt a bi-target perspective. The core idea is to construct a virtual opposite target for each original target. Then, the similarity between a sample and the others is judged by using both the original and opposite targets of the sample. In this way, we propose a bi-target metric method, named TAICHI distance. Considering simplicity and efficiency, we follow the KISSME metric in this paper. Extensive evaluations on challenging datasets confirm the effectiveness of the proposed method. Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Chen Chen 0001 |
ICASSP | 4 |
| 2017 | Transferring clothing parsing from fashion dataset to surveillanceabstractIn this paper we address the problem of automatic clothing parsing in surveillance video with the information from user-generated tags such as “jeans” and “T-shirt”. Although clothing parsing has achieved great success in fashion clothing, it is quite challenging to parse clothing in practical surveillance conditions due to complicated environmental interferences, such as illumination change, scale zooming, viewpoint variation and etc. Our method is developed to capture the clothing information from the fashion field and apply it to surveillance domain by weakly-supervised transfer learning. Most of attribute labels in surveillance images convey strong location information, which can be considered as weak labels to deal with the transfer method. Both quantitative and qualitative experiments conducted on practical surveillance datasets have shown the effectiveness of the proposed method. Jun Chen 0001, Chao Liang 0001, Wenhua Fang, Xiaoyuan Jing, Ruimin Hu |
ICASSP | 3 |
| 2017 | Object tracking via online trajectory optimization with multi-feature fusionabstractThe goal of object tracking is to estimate the state and trajectory of an interested target in a video sequence, thus both spatial and temporal information are of critical importance for tracking. However, most existing trackers usually determine targets just by the judgement like confidence from a single frame, which tend to treat tracking as a static detecting problem while neglecting the spatial-temporal relationship. In this paper, we propose a novel tracking method of online trajectory optimization with multi-feature fusion (TOFF). Considering the trajectory continuity, the accurate targets are determined by estimating the optimal short trajectories over all the video fragments, which can be obtained by the observation models that are iteratively updated based on the selected most reliable proposals. By employing the structured samples instead of binary-labeled samples, we construct a structured output model with representing descriptor of multi-feature fusion as the basic tracker. Extensive experiments on various challenging image sequences demonstrate the superiority of our method to several state-of-the-art methods. Weijian Ruan, Jun Chen 0001, Chao Liang 0001, Yi Wu 0001, Ruimin Hu |
ICME | 3 |
| 2017 | Low-resolution pedestrian detection via a novel resolution-score discriminative surfaceabstractPedestrian detection, as an important task in video surveillance and forensics applications, has been widely studied. However, its performance is unsatisfactory especially in the low resolution conditions. In realistic scenarios, the size of pedestrians in the images is often small, and detection can be challenging. To solve this problem, this paper proposes a novel resolution-score discriminative surface method to investigate the variation behaviors of detection scores under different pedestrian and non-pedestrian image resolutions. The discriminative surface consists of a series of positive and negative resolution-score lines, and each of them is a connected line to depict the variation relationship between pedestrian's detection scores under various image resolutions. On this basis, the resolution-score discriminative surface can classify a resolution-score line as a pedestrian or not according to whether it lies in the positive or the negative region. Experimental results on two public datasets and one campus surveillance dataset demonstrate the effectiveness of the proposed method. Xiao Wang 0029, Jun Chen 0001, Chao Liang 0001, Chen Chen 0001, Zheng Wang 0007, Ruimin Hu |
ICME | 3 |
| 2017 | Structural superpixel descriptor for visual trackingabstractObject representation is a major component in object tracking, however, most conventional patch-based methods just simply decompose the object into patches with grid or stochastic rectangles. This kind of decomposition ignores the intrinsic structure of object, leading to low discriminative power and weak representation effectiveness when similar objects appear or under background clutters. In this paper, we propose an effective object descriptor based on a hierarchical representation with superpixels for visual tracking, called Structural Superpixel Descriptor (SSD). The proposed SSD not only exploits the superpixels to capture the structural information of object, but also preserves the spatial layout structure among the superpixels inside each target candidate. Moreover, we propose an adaptive patch weighting method based on spatial constraint to alleviate various adverse impacts of background information, making the tracker more robust against background noises. We show that the proposed SSD makes full use of the intrinsic structure inside target candidates. Extensive experiments conducted on various challenging sequences demonstrate that the proposed tracker performs well against state-of-the-art algorithms. Ruimin Hu, Chao Liang 0001, Weijian Ruan, Bo Luo |
IJCNN | 3 |
| 2017 | A novel face super resolution approach for noisy images using contour feature and standard deviation prior
Liang Chen 0026, Ruimin Hu, Chao Liang 0001, Qing Li 0001, Zhen Han 0002 |
Multim. Tools Appl. | 3 |
| 2017 | Fine-Grained Image Classification via Low-Rank Sparse Coding With General and Class-Specific CodebooksabstractThis paper tries to separate fine-grained images by jointly learning the encoding parameters and codebooks through low-rank sparse coding (LRSC) with general and class-specific codebook generation. Instead of treating each local feature independently, we encode the local features within a spatial region jointly by LRSC. This ensures that the spatially nearby local features with similar visual characters are encoded by correlated parameters. In this way, we can make the encoded parameters more consistent for fine-grained image representation. Besides, we also learn a general codebook and a number of class-specific codebooks in combination with the encoding scheme. Since images of fine-grained classes are visually similar, the difference is relatively small between the general codebook and each class-specific codebook. We impose sparsity constraints to model this relationship. Moreover, the incoherences with different codebooks and class-specific codebooks are jointly considered. We evaluate the proposed method on several public image data sets. The experimental results show that by learning general and class-specific codebooks with the joint encoding of local features, we are able to model the differences among different fine-grained classes than many other fine-grained image classification methods. Chunjie Zhang 0001, Chao Liang 0001, Liang Li 0003, Jing Liu 0001, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | Person Re-Identification via Multiple Coarse-to-Fine Deep MetricsabstractPerson re-identification, aiming to identify images of the same person from various cameras views in different places, has attracted a lot of research interests in the field of artificial intelligence and multimedia. As one of its popular research directions, the metric learning method plays an important role for seeking a proper metric space to generate accurate feature comparison. However, the existing metric learning methods mainly aim to learn an optimal distance metric function through a single metric, making them difficult to consider multiple similar relationships between the samples. To solve this problem, this paper proposes a coarse-to-fine deep metric learning method equipped with multiple different Stacked Auto-Encoder (SAE) networks and classification networks. In the perspective of the human's visual mechanism, the multiple different levels of deep neural networks simulate the information processing of the brain's visual system, which employs different patterns to recognize the character of objects. In addition, a weighted assignment mechanism is presented to handle the different measure manners for final recognition accuracy. The experimental results conducted on two public datasets, i.e., VIPeR and CUHK have shown the prospective performance of the proposed method. Mingfu Xiong, Jun Chen 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu, Chao Liang 0001, Daming Shi 0001 |
ECAI | 6 |
| 2016 | Distance learning by treating negative samples differently and exploiting impostors with symmetric triplet constraint for person re-identificationabstractDistance learning (DL) is an effective technique for person reidentification (PR-ID). DL based methods learn the distance metric by exploiting the discriminative information contained in samples. In PR-ID, different types of negative samples own different amounts of discriminative information, and impostor samples usually own more than other well separable negative samples (WSN-samples). Therefore, how to make full use of the different discriminative information conveyed by all negative samples in the DL process is a critical issue to be investigated. In this paper, we propose a novel DL approach for PR-ID. Specifically, for each target sample, we divide its negative samples into impostors and WSN-samples. Then we learn the distance metric by utilizing impostors and WSN-samples differently. For impostors, we design a symmetric triplet constraint, which requires the impostor to be far away from both samples of its corresponding positive sample pair simultaneously; for WSN-samples, we require them to keep their favorable separability. Experimental results on three benchmark datasets demonstrate the effectiveness and efficiency of our approach. Xiaoke Zhu, Xiaoyuan Jing, Fei Wu 0004, Wei-Shi Zheng 0001, Ruimin Hu, Chunxia Xiao, Chao Liang 0001 |
ICME | 7 |
| 2016 | Scale-Adaptive Low-Resolution Person Re-Identification via Learning a Discriminating Surface
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Jinqiao Wang |
IJCAI | 5 |
| 2016 | Camera Network Based Person Re-identification by Leveraging Spatial-Temporal Constraint and Multiple Cameras Relations
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Xian Zhong, Chunjie Zhang 0001 |
MMM (1) | 3 |
| 2016 | Spatial Constrained Fine-Grained Color Name for Person Re-identification
Yang Yang 0062, Yuhong Yang 0001, Mang Ye, Wenxin Huang, Zheng Wang 0007, Chao Liang 0001, Chunjie Zhang 0001 |
MMM (1) | 6 |
| 2016 | Geometrically Based Linear Iterative Clustering for Quantitative Feature CorrespondenceabstractAbstract A major challenge in feature matching is the lack of objective criteria to determine corresponding points. Recent methods find match candidates first by exploring the proximity in descriptor space, and then rely on a ratio‐test strategy to determine final correspondences. However, these measurements are heuristic and subjectively excludes massive true positive correspondences that should be matched. In this paper, we propose a novel feature matching algorithm for image collections, which is capable of providing quantitative depiction to the plausibility of feature matches. We achieve this by exploring the epipolar consistency between feature points and their potential correspondences, and reformulate feature matching as an optimization problem in which the overall geometric inconsistency across the entire image set ought to be minimized. We derive the solution of the optimization problem in a simple linear iterative manner, where a k‐means‐type approach is designed to automatically generate consistent feature clusters. Experiments show that our method produces precise correspondences on a variety of image sets and retrieves many matches that are subjectively rejected by recent methods. We also demonstrate the usefulness of the framework in structure from motion task for denser point cloud reconstruction. Qingan Yan, Long Yang 0001, Chao Liang 0001, Huajun Liu, Ruimin Hu, Chunxia Xiao |
Comput. Graph. Forum | 3 |
| 2016 | Zero-Shot Person Re-identification via Cross-View ConsistencyabstractPerson re-identification, aiming to identify images of the same person from various cameras configured in different places, has attracted much attention in the multimedia retrieval community. In this problem, choosing a proper distance metric is a crucial aspect, and many classic methods utilize a uniform learnt metric. However, their performance is limited due to ignoring the zero-shot and fine-grained characteristics presented in real person re-identification applications. In this paper, we investigate two consistencies across two cameras, which are cross-view support consistency and cross-view projection consistency. The philosophy behind it is that, in spite of visual changes in two images of the same person under two camera views, the support sets in their respective views are highly consistent, and after being projected to the same view, their context sets are also highly consistent. Based on the above phenomena, we propose a data-driven distance metric (DDDM) method, re-exploiting the training data to adjust the metric for each query-gallery pair. Experiments conducted on three public data sets have validated the effectiveness of the proposed method, with a significant improvement over three baseline metric learning methods. In particular, on the public VIPeR dataset, the proposed method achieves an accuracy rate of 42.09% at rank-1, which outperforms the state-of-the-art methods by 4.29%. Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Junjun Jiang, Mang Ye, Jun Chen 0001, Qingming Leng |
IEEE Trans. Multim. | 3 |
| 2016 | Person Reidentification via Ranking Aggregation of Similarity Pulling and Dissimilarity PushingabstractPerson reidentification is a key technique to match different persons observed in nonoverlapping camera views. Many researchers treat it as a special object-retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods mainly utilize the similarity relationship between the probe and gallery images to optimize the original ranking list, but seldom consider the important dissimilarity relationship. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person reidentification. Its core idea is that the true match should not only be similar to those strongly similar galleries of the probe, but also be dissimilar to those strongly dissimilar galleries of the probe. Furthermore, motivated by the philosophy of multiview verification, a ranking aggregation algorithm is proposed to enhance the detection of similarity and dissimilarity based on the following assumption: the true match should be similar to the probe in different baseline methods. In other words, if a gallery blue image is strongly similar to the probe in one method, while simultaneously strongly dissimilar to the probe in another method, it will probably be a wrong match of the probe. Extensive experiments conducted on public benchmark datasets and comparisons with different baseline methods have shown the great superiority of the proposed ranking optimization method. Mang Ye, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Qingming Leng, Chunxia Xiao, Jun Chen 0001, Ruimin Hu |
IEEE Trans. Multim. | 2 |
| 2015 | Specific Person Retrieval via Incomplete Text DescriptionabstractSearching for specific persons from surveillance videos captured by different cameras, is a key yet under-addressed challenge in multimedia system. Related person retrieval works mainly focus on searching person by visual appearance, known as person re-identification. However, the initial visual image may not be available in some practical applications. For example, the criminal is described by a text description indirectly, "A young woman wearing a red casual with a backpack", the traditional methods can not conquer this issue. Based on a set of pre-defined attributes that the text description query can be transformed to an attribute vector, thus can be used to retrieval in the gallery set. And yet, the user-provided attributes are sometimes incomplete. This new issue is defined as Specific Person Retrieval via Incomplete Text Description. In this paper, we conduct a specific attribute completion to enrich the original text query and generate a more expressive attribute vector. Then, a pairwise-based metric learning is introduced for completed attribute vectors. Extensive experiments conducted on two benchmark datasets have shown our superior performance. Mang Ye, Chao Liang 0001, Zheng Wang 0007, Qingming Leng, Jun Chen 0001, Jun Liu 0036 |
ICMR | 2 |
| 2015 | A Unsupervised Person Re-identification Method Using Model Based Representation and RankingabstractAs a core technique supporting the multi-camera tracking task, person re-identification attracts increasing research interests in both academic and industrial communities. Its aim is to match individuals across a group of spatially non-overlapping surveillance cameras, which are usually interfered by various imaging conditions and object motions. Current methods mainly focus on robust feature representation and accurate distance measure, where intensive computations and expensive training samples prohibit their practical applications. To address the above problems, this paper proposes a new unsupervised person re-identification method featured by its competitive accuracy and high efficiency. Both merits stem from model based person image representation and ranking, with which, merely 4-dimension pixel-level features can achieve over 20% matching rate at Rank 1 on the challenging VIPeR dataset. Chao Liang 0001, Bingyue Huang, Ruimin Hu, Chunjie Zhang 0001, Xiaoyuan Jing, Jing Xiao 0004 |
ACM Multimedia | 1 |
| 2015 | Multi-Level Fusion for Person Re-identification with Incomplete MarksabstractMost video surveillance suspect investigation systems rely on the videos taken in different camera views. Actually, besides the videos, in the investigation process, investigators also manually label some marks, which, albeit incomplete, can be quite accurate and helpful in identifying persons. This paper studies the problem of Person Re-identification with Incomplete Marks (PRIM), aiming at ranking the persons in the gallery according to both the videos and incomplete marks. This problem is solved by a multi-step fusion algorithm, which consists of three key steps: (i) The early fusing step exploits both visual features and marked attributes to predict a complete and precise attribute vector. (ii) Based on the statistical attribute d ominance and saliency phenomena, a dominance-saliency matching model is suggested for measuring the distance between attribute vectors. (iii) The gallery is ranked separately by using visual features and attribute vectors, and the overall ranking list is the result of a late fusion. Experiments conducted on VIPeR dataset have validated the effectiveness of the proposed method in all the three key steps. The results also show that through introducing marks, the retrieval accuracy is significantly improved. Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Wenxin Huang |
ACM Multimedia | 4 |
| 2015 | Ranking Optimization for Person Re-identification via Similarity and DissimilarityabstractPerson re-identification is a key technique to match different persons observed in non-overlapping camera views.Many researchers treat it as a special object retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods utilize the similarity relationship between the probe and gallery images to optimize the original ranking list in which dissimilarity relationship is seldomly investigated. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person re-identification. Its core idea is based on the phenomenon that the true match should not only be similar to the strong similar samples of the probe but also dissimilar to the strong dissimilar samples. Extensive experiments have shown the great superiority of the proposed ranking optimization method. Mang Ye, Chao Liang 0001, Zheng Wang 0007, Qingming Leng, Jun Chen 0001 |
ACM Multimedia | 2 |
| 2015 | Object Detection in Low-Resolution Image via Sparse Representation
Wenhua Fang, Jun Chen 0001, Chao Liang 0001, Xiao Wang 0029, Yuanyuan Nan, Ruimin Hu |
MMM (1) | 3 |
| 2015 | Sparsity-Based Occlusion Handling Method for Person Re-identification
Bingyue Huang, Jun Chen 0001, Chao Liang 0001, Zheng Wang 0007, Kaimin Sun |
MMM (2) | 4 |
| 2015 | Person Re-identification Using Data-Driven Metric Adaptation
Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Junjun Jiang, Kaimin Sun, Qingming Leng, Bingyue Huang |
MMM (2) | 3 |
| 2015 | Coupled-View Based Ranking Optimization for Person Re-identification
Mang Ye, Jun Chen 0001, Qingming Leng, Chao Liang 0001, Zheng Wang 0007, Kaimin Sun |
MMM (1) | 4 |
| 2015 | Car re-identification from large scale images using semantic attributesabstractCar re-identification, searching a specific car object from a large-scale car image database, is investigated in this paper. Previous work mainly focuses on fixed pose and overlooks the special appearance. However, avoiding matching other poses would lead to coarse results of the car retrieval. And some special attributes like individual paintings which are greatly helpful for car retrieval have not drawn enough attention. This paper addresses these problems through multi-poses matching and re-ranking based on special attributes. Our core idea lies in query expansion method that can capture weighted attributes to build the retrieval model, which allows us to estimate invisible attributes by the visible ones to construct complete attributes vectors to car retrieval in any poses. Furthermore, we divide all attributes into two groups, special attributes and common attributes. Here special attributes represent the abnormal appearance like individual paintings or car damage while common attributes denote the intrinsic appearance of car. Using special attributes to re-rank results turns out to be beneficial to improve the retrieval performance. In the end, the experiments demonstrate the effectiveness of our approach on the car datasets. Chao Liang 0001, Wenhua Fang, Da Xiang, Chengping Ren, Jun Chen 0001 |
MMSP | 2 |
| 2015 | Image classification using boosted local features with random orientation and location selection
Chunjie Zhang 0001, Jian Cheng 0001, Yifan Zhang 0001, Jing Liu 0001, Chao Liang 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001 |
Inf. Sci. | 5 |
| 2015 | Person re-identification with content and context re-ranking
Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001 |
Multim. Tools Appl. | 3 |
| 2014 | Intra-View and Inter-View Supervised Correlation Analysis for Multi-View Feature LearningabstractMulti-view feature learning is an attractive research topic with great practical success. Canonical correlation analysis (CCA) has become an important technique in multi-view learning, since it can fully utilize the inter-view correlation. In this paper, we mainly study the CCA based multi-view supervised feature learning technique where the labels of training samples are known. Several supervised CCA based multi-view methods have been presented, which focus on investigating the supervised correlation across different views. However, they take no account of the intra-view correlation between samples. Researchers have also introduced the discriminant analysis technique into multi-view feature learning, such as multi-view discriminant analysis (MvDA). But they ignore the canonical correlation within each view and between all views. In this paper, we propose a novel multi-view feature learning approach based on intra-view and inter-view supervised correlation analysis (I2SCA), which can explore the useful correlation information of samples within each view and between all views. The objective function of I2SCA is designed to simultaneously extract the discriminatingly correlated features from both inter-view and intra-view. It can obtain an analytical solution without iterative calculation. And we provide a kernelized extension of I2SCA to tackle the linearly inseparable problem in the original feature space. Four widely-used datasets are employed as test data. Experimental results demonstrate that our proposed approaches outperform several representative multi-view supervised feature learning methods. Xiaoyuan Jing, Ruimin Hu, Yang-Ping Zhu, Chao Liang 0001, Jing-Yu Yang 0001 |
AAAI | 5 |
| 2014 | Robust tracking via saliency-based appearance modelabstractWe propose a novel local-based saliency measure (LBSM) method for object tracking problem. In LBSM method, salient patches are defined as the patches having great local changes. Then we apply the saliency information derived from LBSM to appearance model by giving weights to patches according to their saliency levels. The patches with higher saliency levels are given larger weights. As a result, the appearance model is improved owing to the use of saliency information. Extensive experiments conducted on various challenging sequences demonstrate the effectiveness of LB-SM in tracking procedure, and our saliency-based tracker performs well against state-of-the-art algorithms. Bo Luo, Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001 |
ICIP | 4 |
| 2014 | Pedestrian detection from salient regionsabstractClassic algorithms of pedestrian detection usually locate the latent position via sliding window techniques, which resize the matching window and/or original images at different scales and scan the image. However, this method has two main drawbacks. First, resizing at a fix rate cannot search through the whole scale space, resulting in the failure of accurate object location. Second, resizing and scanning at various scales is usually time-consuming, which is improper for practical applications. To conquer the above difficulties, a novel pedestrian detection method with salient information is proposed. In this paper, the salient detection model and the traditional covariance matrix descriptor are combined in a Bayesian framework to detect pedestrians in the still image. Finally, the efficiency of our approach compared with state-of-the-art results is demonstrated on the public INRIA dataset. Xiao Wang 0029, Jun Chen 0001, Wenhua Fang, Chao Liang 0001, Chunjie Zhang 0001, Ruimin Hu |
ICIP | 4 |
| 2014 | Image classification by non-negative sparse coding, correlation constrained low-rank and sparse decomposition
Chunjie Zhang 0001, Jing Liu 0001, Chao Liang 0001, Zhe Xue, Junbiao Pang, Qingming Huang |
Comput. Vis. Image Underst. | 3 |
| 2014 | Object categorization in sub-semantic space
Chunjie Zhang 0001, Jian Cheng 0001, Jing Liu 0001, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001 |
Neurocomputing | 5 |
| 2014 | Beyond visual word ambiguity: Weighted local feature encoding with governing region
Chunjie Zhang 0001, Xian Xiao, Junbiao Pang, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Undoing the codebook bias by linear transformation with sparsity and F-norm constraints for image classification
Chunjie Zhang 0001, Chao Liang 0001, Junbiao Pang, Yifan Zhang 0001, Jing Liu 0001, Qingming Huang |
Pattern Recognit. Lett. | 2 |
| 2014 | Face image super-resolution through locality-induced support regression
Junjun Jiang, Ruimin Hu, Chao Liang 0001, Zhen Han 0002, Chunjie Zhang 0001 |
Signal Process. | 3 |
| 2014 | Camera Compensation Using a Feature Projection Matrix for Person ReidentificationabstractMatching individuals within a group of spatially nonoverlapping surveillance cameras, also known as person reidentification, has recently attracted a lot of research interest. Current methods mainly focus on feature representation or distance measure, which directly compare person images captured by different cameras. However, it is still a problem because of various surveillance conditions; for example, view switching, lighting variations, and image scaling. Although the brightness transfer function was proposed to address the problem of illumination variation, it could not handle view and scale changes among various cameras. In this paper, we propose a new approach to compensate for the inconsistency of feature distributions of person images captured by different cameras. More precisely, a feature projection matrix (FPM) is learned to project image features of one camera to the feature space of another camera, from which the latent device difference can be effectively eliminated for the person reidentification task. In particular, we formulate the FPM learning as a smooth unconstrained convex optimization problem and use a simple gradient descent algorithm with stochastic samples to accelerate the solving process. Extensive comparative experiments conducted on three standard datasets have shown the promising prospect of the proposed method. Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001, Qingming Leng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Bidirectional ranking for person re-identificationabstractThis paper proposes a simple but efficient bidirectional ranking method to improve person re-identification results across non-overlapping cameras. Previous methods treat person reidentification as a special object retrieval problem, and compute the final rank result purely based on a unidirectional matching between the probe and all gallery images. However, the expected person image may be excluded from the probe's ??-nearest neighbor due to appearance changes caused by variations in illuminations, poses, viewpoints and occlusion. To solve the above problem, our method queries every gallery image in a new gallery composed of the original probe image and other gallery images, and revises the initial query result in accordance with both content and context similarities between bidirectional ranking lists. A latent assumption of our method is that images of the same person should not only have similar visual content, known as content similarity, but also possess similar k-nearest neighbors, known as context similarity. Extensive experiments conducted on a series of standard data sets have validated the effectiveness of our proposed method with an average improvement of 5-10% over original baseline methods. Qingming Leng, Ruimin Hu, Chao Liang 0001, Jun Chen 0001 |
ICME | 3 |
| 2013 | Camera compensation using feature projection matrix for person re-identificationabstractMatching individuals across a group of spatially non-overlapping surveillance cameras, also known as person re-identification, has recently attracted a lot of research interests. Current methods mainly focus on feature extraction or metric learning, which directly compare person images captured by different cameras, but seldom consider device differences caused by various surveillance conditions, e.g. view switching, scale zooming and illumination variation. Although brightness transfer function was proposed to address the problem of illumination variation, it could not handle view and scale changes among various cameras. In this paper, we propose an effective data-driven method to conquer device differences in the practical surveillance camera network. More precisely, with the help of a set of labelled pair-wise person images captured by two disjoint cameras, a feature projection matrix can be learned to project the person images of one camera to the feature space of the other camera, and thus images from these two different cameras can be accurately compared in a common feature space. Extensive comparative experiments conducted on three standard datasets have shown the promising prospect of our proposed methods. Ruimin Hu, Chao Liang 0001, Chunjie Zhang 0001, Qingming Leng |
ICME | 3 |
| 2013 | Beyond bag of words: image representation in sub-semantic spaceabstractDue to the semantic gap, the low-level features are not able to semantically represent images well. Besides, traditional semantic related image representation may not be able to cope with large inter class variations and are not very robust to noise. To solve these problems, in this paper, we propose a novel image representation method in the sub-semantic space. First, examplar classifiers are trained by separating each training image from the others and serve as the weak semantic similarity measurement. Then a graph is constructed by combining the visual similarity and weak semantic similarity of these training images. We partition this graph into visually and semantically similar sub-sets. Each sub-set of images are then used to train classifiers in order to separate this sub-set from the others. The learned sub-set classifiers are then used to construct a sub-semantic space based representation of images. This sub-semantic space is not only more semantically meaningful but also more reliable and resistant to noise. Finally, we make categorization of images using this sub-semantic space based representation on several public datasets to demonstrate the effectiveness of the proposed method. Chunjie Zhang 0001, Shuhui Wang, Chao Liang 0001, Jing Liu 0001, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2013 | Undo the codebook bias by linear transformation for visual applicationsabstractThe bag of visual words model (BoW) and its variants have demonstrate their effectiveness for visual applications and have been widely used by researchers. The BoW model first extracts local features and generates the corresponding codebook, the elements of a codebook are viewed as visual words. The local features within each image are then encoded to get the final histogram representation. However, the codebook is dataset dependent and has to be generated for each image dataset. This costs a lot of computational time and weakens the generalization power of the BoW model. To solve these problems, in this paper, we propose to undo the dataset bias by codebook linear transformation. To represent every points within the local feature space using Euclidean distance, the number of bases should be no less than the space dimensions. Hence, each codebook can be viewed as a linear transformation of these bases. In this way, we can transform the pre-learned codebooks for a new dataset. However, not all of the visual words are equally important for the new dataset, it would be more effective if we can make some selection using sparsity constraints and choose the most discriminative visual words for transformation. We propose an alternative optimization algorithm to jointly search for the optimal linear transformation matrixes and the encoding parameters. Image classification experimental results on several image datasets show the effectiveness of the proposed method. Chunjie Zhang 0001, Yifan Zhang 0001, Shuhui Wang, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 5 |
| 2013 | Beyond visual features: A weak semantic image representation using exemplar classifiers for classification
Chunjie Zhang 0001, Jing Liu 0001, Qi Tian 0001, Chao Liang 0001, Qingming Huang |
Neurocomputing | 4 |
| 2013 | Laplacian affine sparse coding with tilt and orientation consistency for image classification
Chunjie Zhang 0001, Shuhui Wang, Qingming Huang, Chao Liang 0001, Jing Liu 0001, Qi Tian 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Image classification using spatial pyramid robust sparse coding
Chunjie Zhang 0001, Shuhui Wang, Qingming Huang, Jing Liu 0001, Chao Liang 0001, Qi Tian 0001 |
Pattern Recognit. Lett. | 5 |
| 2013 | Image classification using Harr-like transformation of local features with coding residuals
Chunjie Zhang 0001, Jing Liu 0001, Chao Liang 0001, Qingming Huang, Qi Tian 0001 |
Signal Process. | 3 |
| 2013 | Script-to-Movie: A Computational Framework for Story Movie CompositionabstractTraditional movie production has always been a highly professional work that needs team collaboration, advanced devices and techniques, and vast time and money investment. These high threshold requirements not only prevent mass amateur enthusiasts entering this field, but also hinder professionals quickly previewing their conceived story plots. In this paper, we raise a novel application, named script-to-movie (S2M) composition, to automatically produce new movies from existing videos in accordance with user created script. Our motivation is to liberate producers from complex filming and editing operations, thereby people's story idea can be instantly converted into the vivid movie video. To support the novel “What You Dream Is What You See” (WYDIWYS) production mode, we first propose a hierarchical alignment method to automatically construct a video material database with detailed semantic description. Considering diverse story plots in user designed script, the database contains abundant video materials about different characters appearing in various time and places conditions. On this basis, the S2M composition is formulated as a constrained optimization problem, where semantic story plot and syntactic visual content are synthetically considered to identify a group of optimal video segments to narrate the user designed script story. Both quantitative and qualitative experimental results are reported to illustrate the effectiveness of the proposed S2M application. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Weiqing Min, Hanqing Lu |
IEEE Trans. Multim. | 1 |
| 2012 | Beyond local image features: Scene calssification using supervised semantic representationabstractThe use of local features for image representation has been proven very effective for a variety of visual tasks such as object localization and scene classification. However, local image features carry little semantic information which is potentially not enough for high level visual tasks. To solve this problem, in this paper, we propose to use a supervised semantic image representation for scene classification, where an image is represented as a response histogram. This response histogram is a combination of the prediction of pre-trained generic object classifiers and classifiers generated by supervised learning. Besides, the use of sparsity constraints makes the proposed representation more efficient and effective to compute. Performances on the UIUC-Sports dataset, the MIT Indoor scene dataset and the Scene-15 dataset demonstrate the effectiveness of the proposed method. Chunjie Zhang 0001, Jing Liu 0001, Chao Liang 0001, Jinhui Tang 0001, Hanqing Lu |
ICIP | 3 |
| 2011 | TVParser: An automatic TV video parsing methodabstractIn this paper, we propose an automatic approach to simultaneously name faces and discover scenes in TV shows. We follow the multi-modal idea of utilizing script to assist video content understanding, but without using timestamp (provided by script-subtitles alignment) as the connection. Instead, the temporal relation between faces in the video and names in the script is investigated in our approach, and an global optimal video-script alignment is inferred according to the character correspondence. The contribution of this paper is two-fold: (1) we propose a generative model, named TVParser, to depict the temporal character correspondence between video and script, from which face-name relationship can be automatically learned as a model parameter, and meanwhile, video scene structure can be effectively inferred as a hidden state sequence; (2) we find fast algorithms to accelerate both model parameter learning and state inference, resulting in an efficient and global optimal alignment. We conduct extensive comparative experiments on popular TV series and report comparable and even superior performance over existing methods. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu |
CVPR | 1 |
| 2011 | Robust movie character identification and the sensitivity analysisabstractAutomatic face identification of characters in movies has drawn significant research interests and led to various applications. It is a challenging problem due to the huge variation in the appearance of each character. Although existing methods demonstrate promising results in clean environment, the performances are limited in complex movie scenes due to the noises generated during the face tracking and face clustering process. In this paper we present a robust character identification approach by incorporating a noise insensitive relationship representation and a graph matching algorithm. Beyond existing character identification approaches, we further perform explicit sensitivity analysis on character identification by introducing two types of simulated noises. Experiments validate the advantage of the proposed method. Chao Liang 0001, Changsheng Xu, Jian Cheng 0001 |
ICME | 2 |
| 2011 | A Visualized Communication System Using Cross-Media Semantic Association
Yang Liu 0021, Chao Liang 0001, Changsheng Xu |
MMM (2) | 3 |
| 2010 | Personalized Sports Video Customization for Mobile Devices
Chao Liang 0001, Jian Cheng 0001, Changsheng Xu, Jinqiao Wang, Hanqing Lu, Jian Ma 0001 |
MMM | 1 |