VLDB 2026 Research / reviewers in the wild / expert
Cuiling Lan
dblp:95/8115
· DBLP profile ↗
76ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0001-9145-9957ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 8 first-author · 24 since 2021Artificial intelligence and machine learning · 44 · 29 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CSPO: Alleviating Reward Ambiguity for Structured Table-to-LaTeX GenerationabstractTables contain rich structured information, yet when stored as images their contents remain "locked" within pixels.Converting table images into LaTeX code enables faithful digitization and reuse, but current multimodal large language models (MLLMs) often fail to preserve structural, style, or content fidelity.Conventional post-training with reinforcement learning (RL) typically relies on a single aggregated reward, leading to reward ambiguity that conflates multiple behavioral aspects and hinders effective optimization.We propose Component-Specific Policy Optimization (CSPO), an RL framework that disentangles optimization across LaTeX tables components-structure, style, and content.In particular, CSPO assigns component-specific rewards and backpropagates each signal only through the tokens relevant to its component, alleviating reward ambiguity and enabling targeted component-wise optimization.To comprehensively assess performance, we introduce a set of hierarchical evaluation metrics.Extensive experiments demonstrate the effectiveness of CSPO, underscoring the importance of component-specific optimization for reliable structured generation. Yunfan Yang, Cuiling Lan, Yan Lu 0001 |
ACL (1) | 2 |
| 2026 | Long Video Understanding With Learnable Retrieval in Video-Language ModelsabstractThe remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understanding, utilizing video tokens as contextual input. However, employing LLMs for long video understanding presents significant challenges. The extensive number of video tokens leads to considerable computational costs for LLMs while using aggregated tokens results in loss of vision details. Moreover, the presence of abundant question-irrelevant tokens introduces noise to the video reasoning process. To address these issues, we introduce a simple yet effective learnable retrieval-based video-language model (R-VLM) for efficient long video understanding. Specifically, given a question and a long video, our model identifies the most relevant$K$video chunks and uses their associated visual tokens to serve as context for the LLM inference. This effectively reduces the number of video tokens, eliminates noise interference, and enhances system performance. We achieve this by incorporating a learnable lightweight MLP block to facilitate the efficient retrieval of question-relevant chunks, through the end-to-end training of our video-language model with a proposed soft matching loss. Experimental results on multiple zero-shot video question answering datasets validate the effectiveness of our framework for comprehending long videos. Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationabstractText-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and the textual description. (ii) how to improve the subjective quality of generated videos. To tackle the above challenges, we propose a new diffusion-based TI2V framework, termed TIV-Diffusion, via object-centric textual-visual alignment, intending to achieve precise control and high-quality video generation based on textual-described motion for different objects. Concretely, we enable our TIV-Diffuion model to perceive the textual-described objects and their motion trajectory by incorporating the fused textual and visual knowledge through scale-offset modulation. Moreover, to mitigate the problems of object disappearance and misaligned objects and motion, we introduce an object-centric textual-visual alignment module, which reduces the risk of misaligned objects/motion by decoupling the objects in the reference image and aligning textual features with each object individually. Based on the above innovations, our TIV-Diffusion achieves state-of-the-art high-quality video generation compared with existing TI2V methods. Xingrui Wang, Xin Li 0082, Yaosi Hu, Hanxin Zhu, Chen Hou, Cuiling Lan, Zhibo Chen 0001 |
AAAI | 6 |
| 2025 | Deciphering Functions of Neurons in Vision-Language ModelsabstractThe burgeoning growth of open-source vision-language models (VLMs) has catalyzed a plethora of applications across diverse domains. Ensuring the transparency and interpretability of these models is critical for fostering trustworthy and responsible AI systems. In this study, our objective is to delve into the internals of VLMs to interpret the functions of individual neurons. We observe the activations of neurons with respects to the input visual tokens and text tokens, and reveal some interesting findings. Particularly, we found that there are neurons responsible for only visual or text information, or both, respectively, which we refer to them as visual neurons, text neurons, and multi-modal neurons, respectively. We build a framework that automates the explanation of neurons with the assistant of GPT-4o. Meanwhile, for visual neurons, we propose an activation simulator to assess the reliability of the explanations for visual neurons. System statistical analyses on top of one representative VLM of LLaVA, uncover the behaviors/characteristics of different categories of neurons. Cuiling Lan, Yan Lu 0001 |
ACM Multimedia | 2 |
| 2025 | Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
Xin Li 0082, Yulin Ren, Xin Jin 0014, Cuiling Lan, Xingrui Wang, Wenjun Zeng 0001, Xinchao Wang, Zhibo Chen 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout AnalysisabstractSignificant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as paragraphs, has not kept pace. Previous works either treated text detection and grouping using separate models, or train a model from scratch while using a unified one. All of them have not yet made full use of the already well-trained text detectors and easily obtainable detection datasets. In this paper, we present Text Grouping Adapter (TGA), a module that can enable the utilization of various pretrained text detectors to learn layout analysis, allowing us to adopt a well-trained text detector right off the shelf or just fine-tune it efficiently. Designed to be compatible with various text detector architectures, TGA takes detected text regions and image features as universal inputs to as-semble text instance features. To capture broader contextual information for layout analysis, we propose to predict text group masks from text instance features by one-to-many assignment. Our comprehensive experiments demonstrate that, even with frozen pretrained models, incorporating our TGA into various pretrained text detectors and text spotters can achieve superior layout analysis performance, simultaneously inheriting generalized text detection ability from pretraining. In the case of full parameter fine-tuning, we can further improve layout analysis performance. Tianci Bi, Zhizheng Zhang 0004, Wenxuan Xie, Cuiling Lan, Yan Lu 0001, Nanning Zheng 0001 |
CVPR | 5 |
| 2024 | UCIP: A Universal Framework for Compressed Image Super-Resolution Using Dynamic Prompt
Xin Li 0082, Bingchen Li 0001, Yeying Jin, Cuiling Lan, Hanxin Zhu, Yulin Ren, Zhibo Chen 0001 |
ECCV (47) | 4 |
| 2024 | Slot-VLM: Object-Event Slots for Video-Language ModelingabstractVideo-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development of an effective method to encapsulate video content into a set of representative tokens to align with LLMs. In this work, we introduce Slot-VLM, a new framework designed to generate semantically decomposed video tokens, in terms of object-wise and event-wise visual representations, to facilitate LLM inference. Particularly, we design an Object-Event Slots module, i.e., OE-Slots, that adaptively aggregates the dense video tokens from the vision encoder to a set of representative slots. In order to take into account both the spatial object details and the varied temporal dynamics, we build OE-Slots with two branches: the Object-Slots branch and the Event-Slots branch. The Object-Slots branch focuses on extracting object-centric slots from features of high spatial resolution but low frame sample rate, emphasizing detailed object information. The Event-Slots branch is engineered to learn event-centric slots from high temporal sample rate but low spatial resolution features. These complementary slots are combined to form the vision context, serving as the input to the LLM for effective video reasoning. Our experimental results demonstrate the effectiveness of our Slot-VLM, which achieves the state-of-the-art performance on video question-answering. Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001 |
NeurIPS | 2 |
| 2024 | Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementabstractDisentangled representation learning strives to extract the intrinsic factors within the observed data. Factoring these representations in an unsupervised manner is notably challenging and usually requires tailored loss functions or specific structural designs. In this paper, we introduce a new perspective and framework, demonstrating that diffusion models with cross-attention itself can serve as a powerful inductive bias to facilitate the learning of disentangled representations. We propose to encode an image into a set of concept tokens and treat them as the condition of the latent diffusion model for image reconstruction, where cross attention over the concept tokens is used to bridge the encoder and the U-Net of the diffusion model. We analyze that the diffusion process inherently possesses the time-varying information bottlenecks. Such information bottlenecks and cross attention act as strong inductive biases for promoting disentanglement. Without any regularization term in the loss function, this framework achieves superior disentanglement performance on the benchmark datasets, surpassing all previous methods with intricate designs. We have conducted comprehensive ablation studies and visualization analyses, shedding a light on the functioning of this model. We anticipate that our findings will inspire more investigation on exploring diffusion model for disentangled representation learning towards more sophisticated data analysis and understanding. Tao Yang 0032, Cuiling Lan, Yan Lu 0001, Nanning Zheng 0001 |
NeurIPS | 2 |
| 2024 | ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided TransformerabstractOptical high-resolution imagery and OpenStreetMap (OSM) data are two important data sources of land-cover change detection (CD). Previous related studies focus on utilizing the information in OSM data to aid the CD on optical high-resolution images. This article pioneers the direct detection of land-cover changes utilizing paired OSM data and optical imagery, thereby expanding the scope of CD tasks. To this end, we propose an object-guided Transformer (ObjFormer) by naturally combining the object-based image analysis (OBIA) technique with the advanced vision Transformer architecture. This combination can significantly reduce the computational overhead in the self-attention module without adding extra parameters or layers. Specifically, ObjFormer has a hierarchical pseudo-Siamese encoder consisting of object-guided self-attention modules that extract multilevel heterogeneous features from OSM data and optical images; a decoder consisting of object-guided cross-attention modules can recover land-cover changes from the extracted heterogeneous features. Beyond basic binary CD (BCD), this article raises a new semi-supervised semantic CD (SCD) task that does not require any manually annotated land-cover labels to train semantic change detectors. Two lightweight semantic decoders are added to ObjFormer to accomplish this task efficiently. A converse cross-entropy (CCE) loss is designed to fully utilize negative samples, contributing to the great performance improvement in this task. A large-scale benchmark dataset called OpenMapCD containing 1287 map–image pairs covering 40 regions on six continents is constructed to conduct the detailed experiments. The results show the effectiveness of our methods in this new kind of CD task. In addition, case studies in two Japanese cities demonstrate the framework’s generalizability and practical potential. The code and dataset will be open-sourced inhttps://github.com/ChenHongruixuan/ObjFormer. Hongruixuan Chen, Cuiling Lan, Jian Song 0010, Clifford Broni-Bediako, Junshi Xia, Naoto Yokoya |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Semantic-Aware Message Broadcasting for Efficient Unsupervised Domain AdaptationabstractVision transformer has demonstrated great potential in abundant vision tasks. However, it also inevitably suffers from poor generalization capability when the distribution shift occurs in testing (i.e., out-of-distribution data). To mitigate this issue, we propose a novel method, Semantic-aware Message Broadcasting (SAMB), which enables more informative and flexible feature alignment for unsupervised domain adaptation (UDA). Particularly, we study the attention module in the vision transformer and notice that the alignment space using one global class token lacks enough flexibility, where it interacts information with all image tokens in the same manner but ignores the rich semantics of different regions. In this paper, we aim to improve the richness of the alignment features by enabling semantic-aware adaptive message broadcasting. Particularly, we introduce a group of learned group tokens as nodes to aggregate the global information from all image tokens, but encourage different group tokens to adaptively focus on the message broadcasting to different semantic regions. In this way, our message broadcasting encourages the group tokens to learn more informative and diverse information for effective domain alignment. Moreover, we systematically study the effects of adversarial-based feature alignment (ADA) and pseudo-label based self-training (PST) on UDA. We find that one simple two-stage training strategy with the cooperation of ADA and PST can further improve the adaptation capability of the vision transformer. Extensive experiments on DomainNet, OfficeHome, and VisDA-2017 demonstrate the effectiveness of our methods for UDA. Xin Li 0082, Cuiling Lan, Guoqiang Wei, Zhibo Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Domain Prompt Tuning via Meta Relabeling for Unsupervised Adversarial AdaptationabstractUnsupervised adversarial domain adaptation (ADA) aims to learn domain-invariant features by confusing a domain discriminator. As training goes on, the feature distributions of source and target samples are increasingly aligned/indistinguishable. The discrimination capability of the domain discriminator w.r.t. those aligned samples deteriorates due to the domain label of each sample is still fixed all through the learning process, which thus cannot effectively further drive the feature learning. A recently proposed method named Re-enforceable Adversarial Domain Adaptation (RADA) [1] tend to re-energize the domain discriminator during the training by using dynamic domain labels. Specifically, RADA sets up a heuristic criterion and uses it to relabel the well aligned target domain samples as source domain samples on the fly. In our study, we identify a critical problem of RADA: it is a kind of heuristic domain data re-partition solution without explicitly serving the adaptation task itself, suggesting that the criteria of RADA on which sample should be relabeled is hard to decide. To address the problem, we revisit domain relabeling process from a perspective of prompt tuning, and introduce a meta-optimized learnable prompts into RADA to replace some hand-craft designs in dynamic relabeling process, which scheme is named as RADA-prompt. Particularly, we employ a module of meta-prompter, which learns to adaptively relabel the samples based on the objective of serving UDA task. To train the meta-prompter, we leverage a domain alignment measurement and a classification measurement as the meta optimization objective. Extensive experiments on multiple unsupervised domain adaptation benchmarks demonstrate the effectiveness and superiority of RADA-prompt, this scheme also achieves state-of-the-art performance. Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Active Token MixerabstractThe three existing dominant network families, i.e., CNNs, Transformers and MLPs, differ from each other mainly in the ways of fusing spatial contextual information, leaving designing more effective token-mixing mechanisms at the core of backbone architecture development. In this work, we propose an innovative token-mixer, dubbed Active Token Mixer (ATM), to actively incorporate contextual information from other tokens in the global scope into the given query token. This fundamental operator actively predicts where to capture useful contexts and learns how to fuse the captured contexts with the query token at channel level. In this way, the spatial range of token-mixing can be expanded to a global scope with limited computational complexity, where the way of token-mixing is reformed. We take ATMs as the primary operators and assemble them into a cascade architecture, dubbed ATMNet. Extensive experiments demonstrate that ATMNet is generally applicable and comprehensively surpasses different families of SOTA vision backbones by a clear margin on a broad range of vision tasks, including visual recognition and dense prediction tasks. Code is available at https://github.com/microsoft/ActiveMLP. Guoqiang Wei, Zhizheng Zhang 0004, Cuiling Lan, Yan Lu 0001, Zhibo Chen 0001 |
AAAI | 3 |
| 2023 | Learning Distortion Invariant Representation for Image Restoration from a Causality PerspectiveabstractIn recent years, we have witnessed the great advancement of Deep neural networks (DNNs) in image restoration. However, a critical limitation is that they cannot generalize well to real-world degradations with different degrees or types. In this paper, we are the first to propose a novel training strategy for image restoration from the causality perspective, to improve the generalization ability of DNNs for unknown degradations. Our method, termed Distortion Invariant representation Learning (DIL), treats each distortion type and degree as one specific confounder, and learns the distortion-invariant representation by eliminating the harmful confounding effect of each degradation. We derive our DIL with the back-door criterion in causality by modeling the interventions of different distortions from the optimization perspective. Particularly, we introduce counterfactual distortion augmentation to simulate the virtual distortion types and degrees as the confounders. Then, we instantiate the intervention of each distortion with a virtual model updating based on corresponding distorted images, and eliminate them from the meta-learning perspective. Extensive experiments demonstrate the generalization capability of our DIL on unseen distortion types and degrees. Our code will be available at https://github.com/lixinustc/Causal-IR-DIL. Xin Li 0082, Bingchen Li 0001, Xin Jin 0014, Cuiling Lan, Zhibo Chen 0001 |
CVPR | 4 |
| 2023 | Deep Frequency Filtering for Domain GeneralizationabstractImproving the generalization ability of Deep Neural Networks (DNNs) is critical for their practical uses, which has been a longstanding challenge. Some theoretical studies have uncovered that DNNs have preferences for some frequency components in the learning process and indicated that this may affect the robustness of learned features. In this paper, we propose Deep Frequency Filtering (DFF)for learning domain-generalizable features, which is the first endeavour to explicitly modulate the frequency components of different transfer difficulties across domains in the latent space during training. To achieve this, we perform Fast Fourier Transform (FFT) for the feature maps at different layers, then adopt a light-weight module to learn attention masks from the frequency representations after FFT to enhance transferable components while suppressing the components not conducive to generalization. Further, we empirically compare the effectiveness of adopting different types of attention designs for implementing DFF. Extensive experiments demonstrate the effectiveness of our proposed DFF and show that applying our DFF on a plain baseline out-performs the state-of-the-art methods on different domain generalization tasks, including close-set classification and open-set retrieval. Shiqi Lin, Zhizheng Zhang 0004, Zhipeng Huang 0014, Yan Lu 0001, Cuiling Lan, Peng Chu, Quanzeng You, Jiang Wang 0012, Zicheng Liu 0001, Amey Parulkar, Viraj Navkal, Zhibo Chen 0001 |
CVPR | 5 |
| 2023 | Adaptive Frequency Filters As Efficient Global Token MixersabstractRecent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. However, their efficient deployments, especially on mobile devices, still suffer from noteworthy challenges due to the heavy computational costs of self-attention mechanisms, large kernels, or fully connected layers. In this work, we apply conventional convolution theorem to deep learning for addressing this and reveal that adaptive frequency filters can serve as efficient global token mixers. With this insight, we propose Adaptive Frequency Filtering (AFF) token mixer. This neural operator transfers a latent representation to the frequency domain via a Fourier transform and performs semantic-adaptive frequency filtering via an elementwise multiplication, which mathematically equals to a token mixing operation in the original latent space with a dynamic convolution kernel as large as the spatial resolution of this latent representation. We take AFF token mixers as primary neural operators to build a lightweight neural network, dubbed AFFNet. Extensive experiments demonstrate the effectiveness of our proposed AFF token mixer and show that AFFNet achieve superior accuracy and efficiency trade-offs compared to other lightweight network designs on broad visual tasks, including visual recognition and dense prediction tasks. Code is available at https://github.com/microsoft/TokenMixers. Zhipeng Huang 0014, Zhizheng Zhang 0004, Cuiling Lan, Zhengjun Zha, Yan Lu 0001, Baining Guo |
ICCV | 3 |
| 2023 | Template-guided Hierarchical Feature Restoration for Anomaly DetectionabstractTargeting for detecting anomalies of various sizes for complicated normal patterns, we propose a Template-guided Hierarchical Feature Restoration method, which introduces two key techniques, bottleneck compression and template-guided compensation, for anomaly-free feature restoration. Specially, our framework compresses hierarchical features of an image by bottleneck structure to preserve the most crucial features shared among normal samples. We design template-guided compensation to restore the distorted features towards anomaly-free features. Particularly, we choose the most similar normal sample as the template, and leverage hierarchical features from the template to compensate the distorted features. The bottleneck could partially filter out anomaly features, while the compensation further converts the reminding anomaly features towards normal with template guidance. Finally, anomalies are detected in terms of the cosine distance between the pre-trained features of an inference image and the corresponding restored anomaly-free features. Experimental results demonstrate the effectiveness of our approach, which achieves the state-of-the-art performance on the MVTec LOCO AD dataset. Hewei Guo, Liping Ren, Jingjing Fu, Yuwang Wang, Zhizheng Zhang 0004, Cuiling Lan, Haoqian Wang, Xinwen Hou |
ICCV | 6 |
| 2023 | Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionabstractReferring image segmentation, the task of segmenting any arbitrary entities described in free-form texts, opens up a variety of vision applications. However, manual labeling of training data for this task is prohibitively costly, leading to lack of labeled data for training. We address this issue by a weakly supervised learning approach using text descriptions of training images as the only source of supervision. To this end, we first present a new model that discovers semantic entities in input image and then combines such entities relevant to text query to predict the mask of the referent. We also present a new loss function that allows the model to be trained without any further supervision. Our method was evaluated on four public benchmarks for referring image segmentation, where it clearly outperformed the existing method for the same task and recent open-vocabulary segmentation models on all the benchmarks. Namyup Kim, Cuiling Lan, Suha Kwak |
ICCV | 3 |
| 2023 | Versatile Neural Processes for Learning Implicit Neural Representations
Zongyu Guo, Cuiling Lan, Zhizheng Zhang 0004, Yan Lu 0001, Zhibo Chen 0001 |
ICLR | 2 |
| 2023 | AReID: Rethinking Re-Identification and Occlusions for Multi-Object TrackingabstractRe-Identification has been widely leverage by tracking-by-detection Multi-Object Tracking methods to enhance the matching step. During training, Re-Identification is usually learned as a sub-task considering the projected centroid of the bounding boxes as the embedding tensor. This design is problematic for Re-Identification feature learning and leveraging because of the noise introduced by the big amount of occlusions. Specifically, the occlusion between objects of the target class causes embedding tensors to represent objects with wrong IDs and this specific scenario has been overlooked in previous literature. In this work we introduce an Adaptive use of Re-Identification features that aims to tackle this problem by using Re-Identification features only when they are reliable. Our method is generic and can be added to further boost performance to any Multi-Object Tracking method that uses Re-Identification. Results in multiple datasets using various base methods demonstrate the consistency of our method with state-of-the-art results. Rodolfo Quispe, Cuiling Lan, Zhizheng Zhang 0004, Hélio Pedrini |
ICMLA | 2 |
| 2023 | WEDGE: Web-Image Assisted Domain Generalization for Semantic SegmentationabstractDomain generalization for semantic segmentation is highly demanded in real applications, where a trained model is expected to work well in previously unseen domains. One challenge lies in the lack of data which could cover the diverse distributions of the possible unseen domains for training. In this paper, we propose a WEb-image assisted Domain GEneralization (WEDGE) scheme, which is the first to exploit the diversity of web-crawled images for generalizable semantic segmentation. To explore and exploit the real-world data distributions, we collect web-crawled images which present large diversity in terms of weather conditions, sites, lighting, camera styles, etc. We also present a method which injects styles of the web-crawled images into training images on-the-fly during training, which enables the network to experience images of diverse styles with reliable labels for effective training. Moreover, we use the web-crawled images with their predicted pseudo labels for training to further enhance the capability of the network. Extensive experiments demonstrate that our method clearly outperforms existing domain generalization techniques. Namyup Kim, Taeyoung Son, Jaehyun Pahk, Cuiling Lan, Wenjun Zeng 0001, Suha Kwak |
ICRA | 4 |
| 2023 | Generalizing to Unseen Domains: A Survey on Domain GeneralizationabstractMachine learning systems generally assume that the training and testing distributions are the same. To this end, a key requirement is to develop models that can generalize to unseen distributions. Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increasing interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. Great progress has been made in the area of domain generalization for years. This paper presents the first review of recent advances in this area. First, we provide a formal definition of domain generalization and discuss several related fields. We then thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. We categorize recent algorithms into three classes: data manipulation, representation learning, and learning strategy, and present several popular algorithms in detail for each category. Third, we introduce the commonly used datasets, applications, and our open-sourced codebase for fair evaluation. Finally, we summarize existing literature and present some potential research topics for the future. Jindong Wang 0001, Cuiling Lan, Chang Liu 0030, Yidong Ouyang, Tao Qin 0001, Wang Lu 0003, Yiqiang Chen 0001, Wenjun Zeng 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Skeleton-Based Mutually Assisted Interacted Object Localization and Human Action RecognitionabstractSkeleton data carries valuable motion information and is widely explored in human action recognition. However, not only the motion information but also the interaction with the environment provides discriminative cues to recognize the action of persons. In this paper, we propose a joint learning framework for mutually assisted “interacted object localization” and “human action recognition” based on skeleton data. The two tasks are serialized together and collaborate to promote each other, where preliminary action type derived from skeleton alone helps improve interacted object localization, which in turn provides valuable cues for the final human action recognition. Besides, we explore the temporal consistency of interacted object as constraint to better localize the interacted object with the absence of ground-truth labels. Extensive experiments on the datasets of SYSU-3D, NTU60 RGB+D, Northwestern-UCLA and UAV-Human show that our method achieves the best or competitive performance with the state-of-the-art methods for human action recognition. Visualization results show that our method can also provide reasonable interacted object localization results. Liang Xu 0012, Cuiling Lan, Wenjun Zeng 0001, Cewu Lu |
IEEE Trans. Multim. | 2 |
| 2022 | Lifelong Unsupervised Domain Adaptive Person Re-identification with Coordinated Anti-forgetting and AdaptationabstractUnsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to changing data statistics and sufficient exploitation of increasing samples. In this paper, to address more practical scenarios, we propose a new task, Lifelong Un-supervised Domain Adaptive (LUDA) person ReID. This is challenging because it requires the model to continuously adapt to unlabeled data in the target environments while alleviating catastrophic forgetting for such a fine-grained person retrieval task. We design an effective scheme for this task, dubbed CLUDA-ReID, where the anti-forgetting is harmoniously coordinated with the adaptation. Specifically, a meta-based Coordinated Data Replay strategy is proposed to replay old data and update the network with a coordinated optimization direction for both adaptation and memorization. Moreover, we propose Relational Consistency Learning for old knowledge distillation/inheritance in line with the objective of retrieval-based tasks. We set up two evaluation settings to simulate the practical application scenarios. Extensive experiments demonstrate the effectiveness of our CLUDA-ReID for both scenarios with stationary target streams and scenarios with dynamic target streams. Zhipeng Huang 0014, Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Peng Chu, Quanzeng You, Jiang Wang 0012, Zicheng Liu 0001, Zhengjun Zha |
CVPR | 3 |
| 2022 | ReSTR: Convolution-free Referring Image Segmentation Using TransformersabstractReferring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for this task rely heavily on convolutional neural networks, which however have trouble capturing long-range dependencies between entities in the language expression and are not flexible enough for modeling interactions between the two different modalities. To address these issues, we present the first convolution-free model for referring image segmentation using transformers, dubbed ReSTR. Since it extracts features of both modalities through transformer encoders, it can capture long-range dependencies between entities within each modality. Also, ReSTR fuses features of the two modalities by a self-attention encoder, which enables flexible and adaptive interactions between the two modalities in the fusion process. The fused features are fed to a segmentation module, which works adaptively according to the image and language expression in hand. ReSTR is evaluated and compared with previous work on all public benchmarks, where it outperforms all existing models. Namyup Kim, Suha Kwak, Cuiling Lan, Wenjun Zeng 0001 |
CVPR | 4 |
| 2022 | Mask-based Latent Reconstruction for Reinforcement LearningabstractFor deep reinforcement learning (RL) from pixels, learning effective state representations is crucial for achieving high performance. However, in practice, limited experience and high-dimensional inputs prevent effective representation learning. To address this, motivated by the success of mask-based modeling in other research fields, we introduce mask-based reconstruction to promote state representation learning in RL. Specifically, we propose a simple yet effective self-supervised method, Mask-based Latent Reconstruction (MLR), to predict complete state representations in the latent space from the observations with spatially and temporally masked pixels. MLR enables better use of context information when learning state representations to make them more informative, which facilitates the training of RL agents. Extensive experiments show that our MLR significantly improves the sample efficiency in RL and outperforms the state-of-the-art sample-efficient RL methods on multiple continuous and discrete control benchmarks. Our code is available at https://github.com/microsoft/Mask-based-Latent-Reconstruction. Tao Yu 0012, Zhizheng Zhang 0004, Cuiling Lan, Yan Lu 0001, Zhibo Chen 0001 |
NeurIPS | 3 |
| 2022 | FPCR-Net: Feature pyramidal correlation and residual reconstruction for optical flow estimation
Jing-Yu Yang 0002, Cuiling Lan, Wenjun Zeng 0001 |
Neurocomputing | 4 |
| 2022 | Style Normalization and Restitution for Domain Generalization and AdaptationabstractFor many computer vision applications, the learned models usually have high performance on the training datasets but suffer from significant performance degradation when deployed in new environments, where there are usually style differences between the training images and the testing images. For high-level vision tasks, an effective domain generalizable model is expected to be able to learn feature representations that are both generalizable and discriminative. In this paper, we design a novel Style Normalization and Restitution module (SNR) to simultaneously ensure high generalization and discrimination capability of the networks. In SNR, particularly, we filter out the style variations (e.g., illumination, color contrast) by performing Instance Normalization (IN) to obtain style normalized features, where the discrepancy among different samples/domains is reduced. However, such a process is task-ignorant and inevitably removes some task-relevant discriminative information, which may hurt the performance. To remedy this, we propose to distill task-relevant discriminative features from the residual (i.e., the difference between the original feature and the style normalized feature) and add them back to the network to ensure high discrimination. Moreover, for better disentanglement, we enforce a dual restitution loss constraint to encourage the better separation of task-relevant and task-irrelevant features. We validate the effectiveness of our SNR on different vision tasks, including classification, semantic segmentation, and object detection. Experiments demonstrate that our SNR is capable of improving the performance of networks for domain generalization (DG) and unsupervised domain adaptation (UDA). Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Beyond Triplet Loss: Meta Prototypical N-Tuple Loss for Person Re-identificationabstractPerson Re-identification (ReID) aims at matching a person of interest across images. In convolutional neural network (CNN) based approaches, loss design plays a vital role in pulling closer features of the same identity and pushing far apart features of different identities. In recent years, triplet loss achieves superior performance and is predominant in ReID. However, triplet loss considers only three instances of two classes in per-query optimization (with an anchor sample as query) and it is actually equivalent to a two-class classification. There is a lack of loss design which enables the joint optimization of multiple instances (of multiple classes) within per-query optimization for person ReID. In this paper, we introduce a multi-class classification loss,i.e., N-tuple loss, to jointly consider multiple ($N$) instances for per-query optimization. This in fact aligns better with the ReID test/inference process, which conducts the ranking/comparisons among multiple instances. Furthermore, for more efficient multi-class classification, we propose a new meta prototypical N-tuple loss. With the multi-class classification incorporated, our model achieves the state-of-the-art performance on the benchmark person ReID datasets Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang |
IEEE Trans. Multim. | 2 |
| 2021 | Exploiting Sample Uncertainty for Domain Adaptive Person Re-IdentificationabstractMany unsupervised domain adaptive (UDA) person ReID approaches combine clustering-based pseudo-label prediction with feature fine-tuning. However, because of domain gap, the pseudo-labels are not always reliable and there are noisy/incorrect labels. This would mislead the feature representation learning and deteriorate the performance. In this paper, we propose to estimate and exploit the credibility of the assigned pseudo-label of each sample to alleviate the influence of noisy labels, by suppressing the contribution of noisy samples. We build our baseline framework using the mean teacher method together with an additional contrastive loss. We have observed that a sample with a wrong pseudo-label through clustering in general has a weaker consistency between the output of the mean teacher model and the student model. Based on this finding, we propose to exploit the uncertainty (measured by consistency levels) to evaluate the reliability of the pseudo-label of a sample and incorporate the uncertainty to re-weight its contribution within various ReID losses, including the ID classification loss per sample, the triplet loss, and the contrastive loss. Our uncertainty-guided optimization brings significant improvement and achieves the state-of-the-art performance on benchmark datasets. Kecheng Zheng, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhengjun Zha |
AAAI | 2 |
| 2021 | MetaAlign: Coordinating Domain Alignment and Classification for Unsupervised Domain AdaptationabstractFor unsupervised domain adaptation (UDA), to alleviate the effect of domain shift, many approaches align the source and target domains in the feature space by adversarial learning or by explicitly aligning their statistics. However, the optimization objective of such domain alignment is generally not coordinated with that of the object classification task itself such that their descent directions for optimization may be inconsistent. This will reduce the effectiveness of domain alignment in improving the performance of UDA. In this paper, we aim to study and alleviate the optimization inconsistency problem between the domain alignment and classification tasks. We address this by proposing an effective meta-optimization based strategy dubbed MetaAlign, where we treat the domain alignment objective and the classification objective as the meta-train and meta-test tasks in a meta-learning scheme. MetaAlign encourages both tasks to be optimized in a coordinated way, which maximizes the inner product of the gradients of the two tasks during training. Experimental results demonstrate the effectiveness of our proposed method on top of various alignment-based baseline approaches, for tasks of object classification and object detection. MetaAlign helps achieve the state-of-the-art performance. Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
CVPR | 2 |
| 2021 | Re-energizing Domain Discriminator with Sample Relabeling for Adversarial Domain AdaptationabstractMany unsupervised domain adaptation (UDA) methods exploit domain adversarial training to align the features to reduce domain gap, where a feature extractor is trained to fool a domain discriminator in order to have aligned feature distributions. The discrimination capability of the domain classifier w.r.t. the increasingly aligned feature distributions deteriorates as training goes on, thus cannot effectively further drive the training of feature extractor. In this work, we propose an efficient optimization strategy named Re-enforceable Adversarial Domain Adaptation (RADA) which aims to re-energize the domain discriminator during the training by using dynamic domain labels. Particularly, we relabel the well aligned target domain samples as source domain samples on the fly. Such relabeling makes the less separable distributions more separable, and thus leads to a more powerful domain classifier w.r.t. the new data distributions, which in turn further drives feature alignment. Extensive experiments on multiple UDA benchmarks demonstrate the effectiveness and superiority of our RADA. Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
ICCV | 2 |
| 2021 | Generalizing to Unseen Domains: A Survey on Domain GeneralizationabstractDomain generalization (DG), i.e., out-of-distribution generalization, has attracted increased interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. For years, great progress has been achieved. This paper presents the first review for recent advances in domain generalization. First, we provide a formal definition of domain generalization and discuss several related fields. Then, we categorize recent algorithms into three classes and present them in detail: data manipulation, representation learning, and learning strategy, each of which contains several popular algorithms. Third, we introduce the commonly used datasets and applications. Finally, we summarize existing literature and present some potential research topics for the future. Jindong Wang 0001, Cuiling Lan, Chang Liu 0030, Yidong Ouyang, Tao Qin 0001 |
IJCAI | 2 |
| 2021 | Uncertainty-Aware Few-Shot Image ClassificationabstractFew-shot image classification learns to recognize new categories from limited labelled data. Metric learning based approaches have been widely investigated, where a query sample is classified by finding the nearest prototype from the support set based on their feature similarities. A neural network has different uncertainties on its calculated similarities of different pairs. Understanding and modeling the uncertainty on the similarity could promote the exploitation of limited samples in few-shot optimization. In this work, we propose Uncertainty-Aware Few-Shot framework for image classification by modeling uncertainty of the similarities of query-support pairs and performing uncertainty-aware optimization. Particularly, we exploit such uncertainty by converting observed similarities to probabilistic representations and incorporate them to the loss for more effective optimization. In order to jointly consider the similarities between a query and the prototypes in a support set, a graph-based model is utilized to estimate the uncertainty of the pairs. Extensive experiments show our proposed method brings significant improvements on top of a strong baseline and achieves the state-of-the-art performance. Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Shih-Fu Chang |
IJCAI | 2 |
| 2021 | Pose-Guided Feature Learning with Knowledge Distillation for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) aims to match person images with occlusion. It is fundamentally challenging because of the serious occlusion which aggravates the misalignment problem between images. At the cost of incorporating a pose estimator, many works introduce pose information to alleviate the misalignment in both training and testing. To achieve high accuracy while preserving low inference complexity, we propose a network named Pose-Guided Feature Learning with Knowledge Distillation (PGFL-KD), where the pose information is exploited to regularize the learning of semantics aligned features but is discarded in testing. PGFL-KD consists of a main branch (MB), and two pose-guided branches, e.g., a foreground-enhanced branch (FEB), and a body part semantics aligned branch (SAB). The FEB intends to emphasise the features of visible body parts while excluding the interference of obstructions and background (e.g., foreground feature alignment). The SAB encourages different channel groups to focus on different body parts to have body part semantics aligned representation. To get rid of the dependency on pose information when testing, we regularize the MB to learn the merits of the FEB and SAB through knowledge distillation and interaction-based training. Extensive experiments on occluded, partial, and holistic ReID tasks show the effectiveness of our proposed network. Kecheng Zheng, Cuiling Lan, Wenjun Zeng 0001, Jiawei Liu 0001, Zhizheng Zhang 0004, Zhengjun Zha |
ACM Multimedia | 2 |
| 2021 | ToAlign: Task-Oriented Alignment for Unsupervised Domain AdaptationabstractUnsupervised domain adaptive classifcation intends to improve the classifcation performance on unlabeled target domain. To alleviate the adverse effect of domain shift, many approaches align the source and target domains in the feature space. However, a feature is usually taken as a whole for alignment without explicitly making domain alignment proactively serve the classifcation task, leading to sub-optimal solution. In this paper, we propose an effective Task-oriented Alignment (ToAlign) for unsupervised domain adaptation (UDA). We study what features should be aligned across domains and propose to make the domain alignment proactively serve classifcation by performing feature decomposition and alignment under the guidance of the prior knowledge induced from the classifcation task itself. Particularly, we explicitly decompose a feature in the source domain into a task-related/discriminative feature that should be aligned, and a task-irrelevant feature that should be avoided/ignored, based on the classifcation meta-knowledge. Extensive experimental results on various benchmarks (e.g., Offce-Home, Visda-2017, and DomainNet) under different domain adaptation settings demonstrate the effectiveness of ToAlign which helps achieve the state-of-the-art performance. The code is publicly available at https://github.com/microsoft/UDA. Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001 |
NeurIPS | 2 |
| 2021 | PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement LearningabstractLearning good feature representations is important for deep reinforcement learning (RL). However, with limited experience, RL often suffers from data inefficiency for training. For un-experienced or less-experienced trajectories (i.e., state-action sequences), the lack of data limits the use of them for better feature learning. In this work, we propose a novel method, dubbed PlayVirtual, which augments cycle-consistent virtual trajectories to enhance the data efficiency for RL feature representation learning. Specifically, PlayVirtual predicts future states in a latent space based on the current state and action by a dynamics model and then predicts the previous states by a backward dynamics model, which forms a trajectory cycle. Based on this, we augment the actions to generate a large amount of virtual state-action trajectories. Being free of groudtruth state supervision, we enforce a trajectory to meet the cycle consistency constraint, which can significantly enhance the data efficiency. We validate the effectiveness of our designs on the Atari and DeepMind Control Suite benchmarks. Our method achieves the state-of-the-art performance on both benchmarks. Our code is available at https://github.com/microsoft/Playvirtual. Tao Yu 0012, Cuiling Lan, Wenjun Zeng 0001, Mingxiao Feng, Zhizheng Zhang 0004, Zhibo Chen 0001 |
NeurIPS | 2 |
| 2021 | CASINet: Content-Adaptive Scale Interaction Networks for scene parsing
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhizheng Zhang 0004, Zhibo Chen 0001 |
Neurocomputing | 2 |
| 2021 | AttributeNet: Attribute enhanced vehicle re-identification
Rodolfo Quispe, Cuiling Lan, Wenjun Zeng 0001, Hélio Pedrini |
Neurocomputing | 2 |
| 2020 | Uncertainty-Aware Multi-Shot Knowledge Distillation for Image-Based Object Re-IdentificationabstractObject re-identification (re-id) aims to identify a specific object across times or camera views, with the person re-id and vehicle re-id as the most widely studied applications. Re-id is challenging because of the variations in viewpoints, (human) poses, and occlusions. Multi-shots of the same object can cover diverse viewpoints/poses and thus provide more comprehensive information. In this paper, we propose exploiting the multi-shots of the same identity to guide the feature learning of each individual image. Specifically, we design an Uncertainty-aware Multi-shot Teacher-Student (UMTS) Network. It consists of a teacher network (T-net) that learns the comprehensive features from multiple images of the same object, and a student network (S-net) that takes a single image as input. In particular, we take into account the data dependent heteroscedastic uncertainty for effectively transferring the knowledge from the T-net to S-net. To the best of our knowledge, we are the first to make use of multi-shots of an object in a teacher-student learning manner for effectively boosting the single image based re-id. We validate the effectiveness of our approach on the popular vehicle re-id and person re-id datasets. In inference, the S-net alone significantly outperforms the baselines and achieves the state-of-the-art performance. Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
AAAI | 2 |
| 2020 | Semantics-Aligned Representation Learning for Person Re-IdentificationabstractPerson re-identification (reID) aims to match person images to retrieve the ones with the same identity. This is a challenging task, as the images to be matched are generally semantically misaligned due to the diversity of human poses and capture viewpoints, incompleteness of the visible bodies (due to occlusion), etc. In this paper, we propose a framework that drives the reID network to learn semantics-aligned feature representation through delicate supervision designs. Specifically, we build a Semantics Aligning Network (SAN) which consists of a base network as encoder (SA-Enc) for re-ID, and a decoder (SA-Dec) for reconstructing/regressing the densely semantics aligned full texture image. We jointly train the SAN under the supervisions of person re-identification and aligned texture generation. Moreover, at the decoder, besides the reconstruction loss, we add Triplet ReID constraints over the feature maps as the perceptual losses. The decoder is discarded in the inference and thus our scheme is computationally efficient. Ablation studies demonstrate the effectiveness of our design. We achieve the state-of-the-art performances on the benchmark datasets CUHK03, Market1501, MSMT17, and the partial person reID dataset Partial REID. Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Guoqiang Wei, Zhibo Chen 0001 |
AAAI | 2 |
| 2020 | Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action RecognitionabstractSkeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this paper, we propose a simple yet effective semantics-guided neural network (SGN) for skeleton-based action recognition. We explicitly introduce the high level semantics of joints (joint type and frame index) into the network to enhance the feature representation capability. In addition, we exploit the relationship of joints hierarchically through two modules, i.e., a joint-level module for modeling the correlations of joints in the same frame and a framelevel module for modeling the dependencies of frames by taking the joints in the same frame as a whole. A strong baseline is proposed to facilitate the study of this field. With an order of magnitude smaller model size than most previous works, SGN achieves the state-of-the-art performance on the NTU60, NTU120, and SYSU datasets. Pengfei Zhang 0005, Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Jianru Xue, Nanning Zheng 0001 |
CVPR | 2 |
| 2020 | Style Normalization and Restitution for Generalizable Person Re-IdentificationabstractExisting fully-supervised person re-identification (ReID) methods usually suffer from poor generalization capability caused by domain gaps. The key to solving this problem lies in filtering out identity-irrelevant interference and learning domain-invariant person representations. In this paper, we aim to design a generalizable person ReID framework which trains a model on source domains yet is able to generalize/perform well on target domains. To achieve this goal, we propose a simple yet effective Style Normalization and Restitution (SNR) module. Specifically, we filter out style variations (e.g., illumination, color contrast) by Instance Normalization (IN). However, such a process inevitably removes discriminative information. We propose to distill identity-relevant feature from the removed information and restitute it to the network to ensure high discrimination. For better disentanglement, we enforce a dual causal loss constraint in SNR to encourage the separation of identity-relevant features and identity-irrelevant features. Extensive experiments demonstrate the strong generalization capability of our framework. Our models empowered by the SNR modules significantly outperform the state-of-the-art domain generalization approaches on multiple widely-used person ReID benchmarks, and also show superiority on unsupervised domain adaptation. Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001, Li Zhang 0040 |
CVPR | 2 |
| 2020 | Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (reID) aims at matching the same person across video clips. It is a challenging task due to the existence of redundancy among frames, newly revealed appearance, occlusion, and motion blurs. In this paper, we propose an attentive feature aggregation module, namely Multi-Granularity Reference-aided Attentive Feature Aggregation (MG-RAFA), to delicately aggregate spatio-temporal features into a discriminative video-level feature representation. In order to determine the contribution/importance of a spatial-temporal feature node, we propose to learn the attention from a global view with convolutional operations. Specifically, we stack its relations, \ieno, pairwise correlations with respect to a representative set of reference feature nodes (S-RFNs) that represents global video information, together with the feature itself to infer the attention. Moreover, to exploit the semantics of different levels, we propose to learn multi-granularity attentions based on the relations captured at different granularities. Extensive ablation studies demonstrate the effectiveness of our attentive feature aggregation module MG-RAFA. Our framework achieves the state-of-the-art performance on three benchmark datasets. Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
CVPR | 2 |
| 2020 | Relation-Aware Global Attention for Person Re-IdentificationabstractFor person re-identification (re-id), attention mechanisms have become attractive as they aim at strengthening discriminative features and suppressing irrelevant ones, which matches well the key of re-id, i.e., discriminative feature learning. Previous approaches typically learn attention using local convolutions, ignoring the mining of knowledge from global structure patterns. Intuitively, the affinities among spatial positions/nodes in the feature map provide clustering-like information and are helpful for inferring semantics and thus attention, especially for person images where the feasible human poses are constrained. In this work, we propose an effective Relation-Aware Global Attention (RGA) module which captures the global structural information for better attention learning. Specifically, for each feature position, in order to compactly grasp the structural information of global scope and local appearance information, we propose to stack the relations, i.e., its pairwise correlations/affinities with all the feature positions (e.g., in raster scan order), and the feature itself together to learn the attention with a shallow convolutional model. Extensive ablation studies demonstrate that our RGA can significantly enhance the feature representation power and help achieve the state-of-the-art performance on several popular benchmarks. The source code is available at https://github.com/microsoft/Relation-Aware-Global-Attention-Networks. Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Xin Jin 0014, Zhibo Chen 0001 |
CVPR | 2 |
| 2020 | Global Distance-Distributions Separation for Unsupervised Person Re-identification
Xin Jin 0014, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
ECCV (7) | 2 |
| 2020 | Temporal-Spatial Mapping for Action RecognitionabstractDeep learning models have enjoyed great success for image related computer vision tasks such as image classification and object detection. For video related tasks such as human action recognition, however, the advancements are not as significant yet. The main challenge is the lack of effective and efficient models in modeling the rich temporal-spatial information in a video. We introduce a simple yet effective operation, termed temporal-spatial mapping, for capturing the temporal evolution of the frames by jointly analyzing all the frames of a video. We propose a video level 2D feature representation by transforming the convolutional features of all frames to a 2D feature map, referred to as VideoMap. With each row being the vectorized feature representation of a frame, the temporal-spatial features are compactly represented, while the temporal dynamic evolution is also well embedded. Based on the VideoMap representation, we further propose a temporal attention model within a shallow convolutional neural network to efficiently exploit the temporal-spatial dynamics. The experiment results show that the proposed scheme achieves state-of-the-art performance, with 4.2% accuracy gain over the temporal segment network, a competing baseline method, on the challenging human action benchmark dataset HMDB51. Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Xiaoyan Sun 0001, Jing-Yu Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | View Invariant 3D Human Pose EstimationabstractThe recent success of neural networks has significantly advanced the performance of 3D human pose estimation from 2D input images. However, the diversity of capturing viewpoints and the flexibility of the human poses remain some significant challenges. In this paper, we propose a view-invariant 3D human pose estimation module to alleviate the effects of viewpoint diversity. The proposed framework consists of a base network, which provides an initial estimation of a 3D pose, a view-invariant hierarchical correction network (VI-HC) on top of that to learn the 3D pose refinement under consistent views, and a view-invariant discriminative network (VID) to enforce high-level constraints over body configurations. In VI-HC, the initial 3D pose inputs are automatically transformed to consistent views for further refinements at the global body and local body parts level, respectively. For the VID, under consistent viewpoints, we use adversarial learning to differentiate between estimated 3D poses and real 3D poses to avoid implausible results. The experimental results demonstrate that the constraint on viewpoint consistency can dramatically enhance the performance of 3D human pose estimation. Our module shows robustness for different 3D pose base networks and achieves a significant improvement (about 9%) over a powerful baseline on the public 3D pose estimation benchmark Human3.6M. Guoqiang Wei, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) are capable of modeling temporal dependencies of complex sequential data. In general, current available structures of RNNs tend to concentrate on controlling the contributions of current and previous information. However, the exploration of different importance levels of different elements within an input vector is always ignored. We propose a simple yet effective Element-wise-Attention Gate (EleAttG), which can be easily added to an RNN block (e.g. all RNN neurons in an RNN layer), to empower the RNN neurons to have attentiveness capability. For an RNN block, an EleAttG is used for adaptively modulating the input by assigning different levels of importance, i.e., attention, to each element/dimension of the input. We refer to an RNN block equipped with an EleAttG as an EleAtt-RNN block. Instead of modulating the input as a whole, the EleAttG modulates the input at fine granularity, i.e., element-wise, and the modulation is content adaptive. The proposed EleAttG, as an additional fundamental unit, is general and can be applied to any RNN structures, e.g., standard RNN, Long Short-Term Memory (LSTM), or Gated Recurrent Unit (GRU). We demonstrate the effectiveness of the proposed EleAtt-RNN by applying it to different tasks including the action recognition, from both skeleton-based data and RGB videos, gesture recognition, and sequential MNIST classification. Experiments show that adding attentiveness through EleAttGs to RNN blocks significantly improves the power of RNNs. Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Densely Semantically Aligned Person Re-IdentificationabstractWe propose a densely semantically aligned person re-identification (re-ID) framework. It fundamentally addresses the body misalignment problem caused by pose/viewpoint variations, imperfect person detection, occlusion, etc.. By leveraging the estimation of the dense semantics of a person image, we construct a set of densely semantically aligned part images (DSAP-images), where the same spatial positions have the same semantics across different person images. We design a two-stream network that consists of a main full image stream (MF-Stream) and a densely semantically-aligned guiding stream (DSAG-Stream). The DSAG-Stream, with the DSAP-images as input, acts as a regulator to guide the MF-Stream to learn densely semantically aligned features from the original image. In the inference, the DSAG-Stream is discarded and only the MF-Stream is needed, which makes the inference system computationally efficient and robust. To our best knowledge, we are the first to make use of fine grained semantics for addressing misalignment problems for re-ID. Our method achieves rank-1 accuracy of 78.9% (new protocol) on the CUHK03 dataset, 90.4% on the CUHK01 dataset, and 95.7% on the Market1501 dataset, outperforming state-of-the-art methods. Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Zhibo Chen 0001 |
CVPR | 2 |
| 2019 | View Adaptive Neural Networks for High Performance Skeleton-Based Human Action RecognitionabstractSkeleton-based human action recognition has recently attracted increasing attention thanks to the accessibility and the popularity of 3D skeleton data. One of the key challenges in action recognition lies in the large variations of action representations when they are captured from different viewpoints. In order to alleviate the effects of view variations, this paper introduces a novel view adaptation scheme, which automatically determines the virtual observation viewpoints over the course of an action in a learning based data driven manner. Instead of re-positioning the skeletons using a fixed human-defined prior criterion, we design two view adaptive neural networks, i.e., VA-RNN and VA-CNN, which are respectively built based on the recurrent neural network (RNN) with the Long Short-term Memory (LSTM) and the convolutional neural network (CNN). For each network, a novel view adaptation module learns and determines the most suitable observation viewpoints, and transforms the skeletons to those viewpoints for the end-to-end recognition with a main classification network. Ablation studies find that the proposed view adaptive models are capable of transforming the skeletons of various views to much more consistent virtual viewpoints. Therefore, the models largely eliminate the influence of the viewpoints, enabling the networks to focus on the learning of action-specific features and thus resulting in superior performance. In addition, we design a two-stream scheme (referred to as VA-fusion) that fuses the scores of the two networks to provide the final prediction, obtaining enhanced performance. Moreover, random rotation of skeleton sequences is employed to improve the robustness of view adaptation models and alleviate overfitting during training. Extensive experimental evaluations on five challenging benchmarks demonstrate the effectiveness of the proposed view-adaptive networks and superior performance over state-of-the-art approaches. Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Skeleton-Based Action Recognition With Gated Convolutional Neural NetworksabstractFor skeleton-based action recognition, most of the existing works used recurrent neural networks. Using convolutional neural networks (CNNs) is another attractive solution considering their advantages in parallelization, effectiveness in feature learning, and model base sufficiency. Besides these, skeleton data are low-dimensional features. It is natural to arrange a sequence of skeleton features chronologically into an image, which retains the original information. Therefore, we solve the sequence learning problem as an image classification task using CNNs. For better learning ability, we build a classification network with stacked residual blocks and having a special design called linear skip gated connection which can benefit information propagation across multiple residual blocks. When arranging the coordinates of body joints in one frame into a skeleton feature, we systematically investigate the performance of part-based, chain-based, and traversal-based orders. Furthermore, a fully convolutional permutation network is designed to learn an optimized order for data rearrangement. Without any bells and whistles, our proposed model achieves state-of-the-art performance on two challenging benchmark datasets, outperforming existing methods significantly. Congqi Cao, Cuiling Lan, Yifan Zhang 0001, Wenjun Zeng 0001, Hanqing Lu, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Multi-Modality Multi-Task Recurrent Neural Network for Online Action DetectionabstractOnline action detection is a brand new challenge and plays a critical role in visual surveillance analytics. It goes one step further than a conventional action recognition task, which recognizes human actions from well-segmented clips. Online action detection is desired to identify the action type and localize action positions on the fly from the untrimmed stream data. In this paper, we propose a multi-modality multi-task recurrent neural network, which incorporates both RGB and Skeleton networks. We design different temporal modeling networks to capture specific characteristics from various modalities. Then, a deep long short-term memory subnetwork is utilized effectively to capture the complex long-range temporal dynamics, naturally avoiding the conventional sliding window design and thus ensuring high computational efficiency. Constrained by a multi-task objective function in the training phase, this network achieves superior detection performance and is capable of automatically localizing the start and end points of actions more accurately. Furthermore, embedding subtask of regression provides the ability to forecast the action prior to its occurrence. We evaluate the proposed method and several other methods in action detection and forecasting on the online action detection data set and gaming action data set datasets. Experimental results demonstrate that our model achieves the state-of-the-art performance on both tasks. Jiaying Liu 0001, Yanghao Li, Sijie Song, Junliang Xing, Cuiling Lan, Wenjun Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Adding Attentiveness to the Neurons in Recurrent Neural Networks
Pengfei Zhang 0005, Jianru Xue, Cuiling Lan, Wenjun Zeng 0001, Zhanning Gao, Nanning Zheng 0001 |
ECCV (9) | 3 |
| 2018 | Skeleton-Indexed Deep Multi-Modal Feature Learning for High Performance Human Action RecognitionabstractThis paper presents a new framework for action recognition with multi-modal data. A skeleton-indexed feature learning procedure is developed to further exploit the detailed local features from RGB and optical flow videos. In particular, the proposed framework is built based on a deep Convolutional Network (ConvNet) and a Recurrent Neural Network (RNN) with Long Short Term Memory (LSTM). A skeleton-indexed transform layer is designed to automatically extract visual features around key joints, and a part-aggregated pooling is developed to uniformly regulate the visual features from different body parts and actors. Besides, several fusion schemes are explored to take advantage of multi-modal data. The proposed deep architecture is end-to-end trainable and can better incorporate different modalities to learn effective feature representations. Quantitative experiment results on two datasets, the NTU RGB+D dataset and the MSR dataset, demonstrate the excellent performance of our scheme over other state-of-the-arts. To our knowledge, the performance obtained by the proposed framework is currently the best on the challenging NTU RGB+D dataset. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
ICME | 2 |
| 2018 | A Practical Hybrid Digital-Analog Scheme for Wireless Video TransmissionabstractWe propose a hybrid digital-analog framework for wireless video transmission, which benefits from both the high distortion-power performance of digital systems and the graceful performance degradation of analog systems. The proposed framework models video frames as a parallel Gaussian source, which is separated into digital and analog parts through scalar quantization. It features entropy coding and channel coding in digital transmission and power scaling in analog transmission. The key challenge in this framework is how to allocate the constrained power and bandwidth resources between and among digital and analog components to achieve minimal distortion at the receiver. Given the worst-case channel signal-to-noise ratio, we are able to derive a closed-form expression of the overall distortion. However, minimizing it is a mixed-integer non-linear programming problem, which is generally non-deterministic polynomial-time hard. By making reasonable and justified simplifications, we approach the optimal solution through a practical scheme. Evaluations show that the proposed scheme outperforms the state-of-the-art analog scheme SoftCast by a large margin. The gain in received video peak signal-to-noise ratio is up to 5.0 dB for various types of videos. Cuiling Lan, Chong Luo 0001, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Variable Block-Sized Signal-Dependent Transform for Video CodingabstractTransform, as one of the most important modules of mainstream video coding systems, seems very stable over the past several decades. However, recent developments indicate that bringing more options for transform can lead to coding efficiency benefits. In this paper, we go further to investigate how the coding efficiency can be improved over the state-of-the-art method by adapting a transform for each block. We present a variable block-sized signal-dependent transforms (SDTs) design based on the High Efficiency Video Coding (HEVC) framework. For a coding block ranged from $4\times4$ to $32\times32$ , we collect a quantity of similar blocks from the reconstructed area and use them to derive the Karhunen-Loève transform. We avoid sending overhead bits to denote the transform by performing the same procedure at the decoder. In this way, the transform for every block is tailored according to its statistics, to be signal-dependent. To make the large block-sized SDTs feasible, we present a fast algorithm for transform derivation. Experimental results show the effectiveness of the SDTs for different block sizes, which leads to up to 23.3% bit-saving. On average, we achieve BD-rate saving of 2.2%, 2.4%, 3.3%, and 7.1% under AI-Main10, RA-Main10, RA-Main10, and LP-Main10 configurations, respectively, compared with the test model HM-12 of HEVC. The proposed scheme has also been adopted into the joint exploration test model for the exploration of potential future video coding standard. Cuiling Lan, Jizheng Xu, Wenjun Zeng 0001, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Spatio-Temporal Attention-Based LSTM Networks for 3D Action Recognition and DetectionabstractHuman action analytics has attracted a lot of attention for decades in computer vision. It is important to extract discriminative spatio-temporal features to model the spatial and temporal evolutions of different actions. In this paper, we propose a spatial and temporal attention model to explore the spatial and temporal discriminative features for human action recognition and detection from skeleton data. We build our networks based on the recurrent neural networks with long short-term memory units. The learned model is capable of selectively focusing on discriminative joints of skeletons within each input frame and paying different levels of attention to the outputs of different frames. To ensure effective training of the network for action recognition, we propose a regularized cross-entropy loss to drive the learning process and develop a joint training strategy accordingly. Moreover, based on temporal attention, we develop a method to generate the action temporal proposals for action detection. We evaluate the proposed method on the SBU Kinect Interaction data set, the NTU RGB + D data set, and the PKU-MMD data set, respectively. Experiment results demonstrate the effectiveness of our proposed model on both action recognition and action detection. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton DataabstractHuman action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model, both on the small human action recognition dataset of SBU and the currently largest NTU dataset. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
AAAI | 2 |
| 2017 | Human Pose Estimation Using Global and Local NormalizationabstractIn this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on spatial configuration refinement by reducing variations of human poses statistically, which is motivated by the observation that the scattered distribution of the relative locations of joints (e.g., the left wrist is distributed nearly uniformly in a circular area around the left shoulder) makes the learning of convolutional spatial models hard. We present a two-stage normalization scheme, human body normalization and limb normalization, to make the distribution of the relative joint locations compact, resulting in easier learning of convolutional spatial models and more accurate pose estimation. In addition, our empirical results show that incorporating multi-scale supervision and multi-scale fusion into the joint detection network is beneficial. Experiment results demonstrate that our method consistently outperforms state-of-the-art methods on the benchmarks. Ke Sun 0009, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Dong Liu 0002, Jingdong Wang 0001 |
ICCV | 2 |
| 2017 | View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton DataabstractSkeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints during the occurrence of an action. Rather than re-positioning the skeletons based on a human defined prior criterion, we design a view adaptive recurrent neural network (RNN) with LSTM architecture, which enables the network itself to adapt to the most suitable observation viewpoints from end to end. Extensive experiment analyses show that the proposed view adaptive RNN model strives to (1) transform the skeletons of various views to much more consistent viewpoints and (2) maintain the continuity of the action rather than transforming every frame to the same position with the same body orientation. Our model achieves significant improvement over the state-of-the-art approaches on three benchmark datasets. Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001 |
ICCV | 2 |
| 2017 | Progressive Pseudo-analog Transmission for Mobile Video StreamingabstractWe propose a progressive pseudo-analog video transmission scheme that simultaneously handles SNR and bandwidth variations with graceful quality degradation for mobile video streaming. With the inherited SNR-adaptability from pseudo-analog transmission, the proposed progressive solution acquires bandwidth adaptability through an innovative scheduling algorithm with optimal power allocation. The basic idea is to aggressively transmit or retransmit important coefficients so that distortion is minimized at the receiver after each received packet. We derive the closed-form expression of reduced distortion for each packet under given transmission power and known channel conditions, and show that the optimal solution can be obtained with a water-filling algorithm. We also illustrate through analyses and simulations that a near-optimal solution can be found through approximation when only statistical channel information is available. Simulations show that our solution approaches the performance upper bound of pseudo-analog transmission in an additive white Gaussian noise channel and significantly outperforms existing pseudo-analog solutions in a fast Rayleigh fading channel. Trace-driven emulations are also carried out to demonstrate the advantage of the proposed solution over the state-of-the-art digital and pseudo-analog solutions under a real dramatically varying wireless environment. Dongliang He, Cuiling Lan, Chong Luo 0001, Enhong Chen, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Co-Occurrence Feature Learning for Skeleton Based Action Recognition Using Regularized Deep LSTM NetworksabstractSkeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term Memory (LSTM) can learn feature representations and model long-term temporal dependencies automatically, we propose an end-to-end fully connected deep LSTM network for skeleton based action recognition. Inspired by the observation that the co-occurrences of the joints intrinsically characterize human actions, we take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. To train the deep LSTM network effectively, we propose a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons. Experimental results on three human action recognition datasets consistently demonstrate the effectiveness of the proposed model. Wentao Zhu 0001, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Yanghao Li, Li Shen 0005, Xiaohui Xie |
AAAI | 2 |
| 2016 | Online Human Action Detection Using Joint Classification-Regression Recurrent Neural Networks
Yanghao Li, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Chunfeng Yuan, Jiaying Liu 0001 |
ECCV (7) | 2 |
| 2016 | OMP-based transform for inter coding in HEVCabstractDiscrete Cosine Transform (DCT) has been the commonly used transform for a few decades in image/video coding. However, DCT does not work well on the blocks having anisotropic correlations. In this paper, based on the adaptive dictionary, we propose a new online transform scheme using Orthogonal Matching Pursuit (OMP) for High Efficiency Video Coding (HEVC). For a coding block, we construct its dictionary by exploiting non-local correlations from the reconstructed regions. The OMP algorithm is implemented to obtain the sparse transform coefficients. Experimental results show that the BD-rate savings of the proposed scheme for the sequences with strong edges can be up to 19.9%. Cuiling Lan, Houqiang Li, Jizheng Xu, Feng Wu 0001 |
ISCAS | 2 |
| 2016 | Internal-video mode dependent directional transformabstractAs the projection of the real world, videos usually have many repeated patterns with similar structures cross regions, presenting strong non-local correlations. Moreover, different videos own different characteristics. Exploitation of the non-local correlations by off-line training of transforms has attracted considerable attention over the past years for compression. However, the samples used for training the transforms are usually collected from a predefined set of training videos to avoid the transmission of transform matrixes. There is no guarantee that the characteristics of those training videos and the corresponding transforms fit the current coding video very well. To address that, this paper proposes an internal-video mode dependent directional transform in HEVC for intra coding. In this scheme, for coding a video clip, based on different directions, a set of sample blocks is collected for each direction to train a Karhunen-Loeve transform at the video clip level. During encoding, the better transform between the proposed transform and original DCT/DST in HEVC is determined based on rate-distortion optimization. These transform matrixes are entropy coded to the bitstream. The proposed method can capture the statistical characteristics of the coded video clip and provide efficient transforms. Experimental results show that the proposed method achieves significant performance improvements in comparison with HEVC for intra coding, up to 13% BD-rate savings for the all intra configuration. Cuiling Lan, Yunhui Shi, Wenpeng Ding, Xiaoyan Sun 0001 |
VCIP | 2 |
| 2015 | Compound image compression using lossless and lossy LZMA in HEVCabstractWe present a compound image compression scheme based on the dictionary-based Lempel-Ziv-Markov chain algorithm (LZMA), under the framework of High Efficiency Video Coding (HEVC). Through matching strings from the sliding window dictionary, LZMA exploits the characteristics of the repeated patterns over the text and graphics regions of compound images, and represents them compactly. To obtain high compression efficiency even for noisy text and graphics contents, we have modified LZMA to support both lossless and lossy compression. We develop and treat it as a new intramode of HEVC. Experimental results show that the proposed scheme achieves significant coding gains for compound image compression. Thanks to the introduction of the lossy LZMA, the compression performance for noisy compound images is improved for more than 5dB in terms of PSNR in comparison with the lossless LZMA scheme. Cuiling Lan, Jizheng Xu, Wenjun Zeng 0001, Feng Wu 0001 |
ICME | 1 |
| 2015 | Progressive pseudo-analog transmission for mobile video live streamingabstractMobile video live streaming is facing great challenges in offering high quality of experience (QoE) under varying channel conditions. In this paper, we propose a progressive pseudo-analog transmission scheme in which the received video quality gracefully adapts to both SNR and bandwidth variations. Building upon the emerging pseudoanalog video transmission, the proposed scheme further adopts a greedy approach to improve the received video quality with each allocated bandwidth share. The optimal scheduling and power allocation are derived under the mean squared error (MSE) criterion. Testbed evaluations show that the proposed scheme outperforms the state-of-the-art digital and analog transmission schemes by a notable margin. Cuiling Lan, Dongliang He, Chong Luo 0001, Feng Wu 0001, Wenjun Zeng 0001 |
VCIP | 1 |
| 2015 | Structure-Preserving Hybrid Digital-Analog Video Delivery in Wireless NetworksabstractHybrid digital-analog (HDA) transmission has gained increasing attention recently in the context of wireless video delivery , for its ability to simultaneously achieve high transmission efficiency and smooth quality adaptation. However, previous systems are optimized solely based on the mean squared error criterion without taking the perceptual video quality into consideration. In this work, we propose a structure-preserving HDA video delivery system, named SharpCast, to improve both the objective and subjective visual quality. SharpCast decomposes a video into a content part and structure part. The latter is important to the human perception and therefore is protected with a robust digital transmission scheme. Then, the energy-intensive part in the content information is extracted and transmitted in digital for energy efficiency while the residual is transmitted in analog to achieve the desired smooth adaptation. We formulate the resource (power and bandwidth) allocation problem in SharpCast and solve the problem with a greedy strategy. Evaluations over nine standard 720p video sequences show that the proposed SharpCast system outperforms the state-of-the-art digital, analog, and HDA schemes by a notable margin in both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). Dongliang He, Chong Luo 0001, Cuiling Lan, Feng Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Energy-efficient design of real-time stream mining systemsabstractIn this paper, we propose an efficient solution for supporting real-time stream mining applications on heterogeneous systems operating at various processing speeds. Unlike the existing solutions that (1) rely on accurate knowledge or prediction of the service demand of each individual service request and (2) only consider a single type of delay constraint (e.g., typically, average or maximum delay), we propose an optimal algorithm, MinEnergy-MD, which determines the processing speeds for all classifiers based on the probability distribution of the service demand to minimize the average energy consumption while simultaneously satisfying multiple delay constraints. We conduct an extensive study to quantify the performance of MinEnergy-MD. Shaolei Ren, Cuiling Lan, Mihaela van der Schaar |
ICASSP | 2 |
| 2013 | Low-complexity reinforcement learning for delay-sensitive compression in networked video stream miningabstractIn networked video stream mining systems, real-time video contents are captured remotely and, subsequently, encoded and transmitted over bandwidth-constrained networks for classification at the receiver. One key task at the encoder is to adapt its compression on the fly based on time-varying network bandwidth and video characteristics — while attaining low delay and high classification accuracy. In this paper, we formalize the decision at the encoder side as an infinite horizon Markov Decision Process (MDP). We employ low-complexity, model-free reinforcement learning schemes to solve this problem efficiently under dynamic and unknown environment. Our proposed scheme adopts the technique of virtual experience (VE) update to drastically speed up convergence over conventional Q-learning, allowing the encoder to react to abrupt network changes on the order of minutes, instead of hours. In comparison to myopic optimization, it consistently achieves higher overall reward and lower sending delay under various network conditions. Cuiling Lan, Mihaela van der Schaar |
ICME | 2 |
| 2012 | Object-based coding for Kinect depth and color videosabstractSimultaneously capturing of color and depth videos, e.g. with Kinect, favors many applications and has become very popular. Efficient representation and compression of such data is important yet challenging. In this paper, we have designed an object-based coding system to compress Kinect-like depth and color videos. Segmentation is first conducted to obtain different object planes, where a mask image is utilized to identify them. We compress depth and color images respectively using the proposed object-based coding codec, which is designed based on High Efficiency Video Coding (HEVC). The mask image is losslessly compressed by adding a new context-based mode to HEVC. To assure the alignment of object boundaries on the depth image and those on the color image, a pre-processing is conducted over the depth image. The separate coding of the different object planes for the depth image can avoid the inefficiency coding of edges blocks at object boundaries and thus bring obvious coding gain. Moreover, the attractive functionality of “content-based” coding which permits the transmission of the interested object planes rather than an entire image provides a practical way to decrease the bitrate. Cuiling Lan, Jizheng Xu, Feng Wu 0001 |
VCIP | 1 |
| 2011 | Compression of compound images by combining several strategiesabstractCompound images are combinations of text, graphics and natural images. They possess characteristics different from those of natural images, such as a strong anisotropy, sparse color histograms and repeated patterns. Former research on compressing them has mainly focused on developing certain strategies based on some of these characteristics but has failed so far to fully exploit them simultaneously. In this paper, we investigate the combination of four up-to-date strategies to construct a comprehensive scheme for compound image compression. We have implemented these strategies as four types of modes with variable block sizes. Experimental results show that the proposed scheme achieves significant coding gains for compound image compression at all bitrates. Cuiling Lan, Jizheng Xu, Feng Wu 0001 |
MMSP | 1 |
| 2010 | Intra frame coding with template matching prediction and adaptive transformabstractFor natural images, there are usually repeating similar contents but hard to be well predicted locally. Prediction using template matching is an effective technology to exploit such a non-local correlation. In this paper, we propose an alternative scheme to further exploit the non-local correlation. In the proposed scheme, template matching is also used to search for probable similar references to the current block to be coded. We then use these references to train an adaptive transform, which most likely reflects the statistical characteristic of the current block. The proposed scheme can further exploit correlation between the current block and more possible references. Compared to the scheme that only integrates prediction by template matching, the proposed scheme shows improvement about 0.45dB PSNR increase or 9.3% bit saving on average, which leads to 1dB's gain or 19.5% bit saving on average compared to the state-of-the-art scheme without using template matching. Cuiling Lan, Jizheng Xu, Feng Wu 0001, Guangming Shi |
ICIP | 1 |
| 2010 | Compress Compound Images in H.264/MPGE-4 AVC by Exploiting Spatial CorrelationabstractCompound images are a combination of text, graphics and natural image. They present strong anisotropic features, especially on the text and graphics parts. These anisotropic features often render conventional compression inefficient. Thus, this paper proposes a novel coding scheme from the H.264 intraframe coding. In the scheme, two new intramodes are developed to better exploit spatial correlation in compound images. The first is the residual scalar quantization (RSQ) mode, where intrapredicted residues are directly quantized and coded without transform. The second is the base colors and index map (BCIM) mode that can be viewed as an adaptive color quantization. In this mode, an image block is represented by several representative colors, referred to as base colors, and an index map to compress. Every block selects its coding mode from two new modes and the previous intramodes in H.264 by rate-distortion optimization (RDO). Experimental results show that the proposed scheme improves the coding efficiency even more than 10 dB at most bit rates for compound images and keeps a comparable efficient performance to H.264 for natural images. Cuiling Lan, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | Compress Compound Images in H.264/MPEG-4 AVC by Fully Exploiting Spatial CorrelationabstractCompound images consist of text, graphics and natural images, which present strong anisotropic features. It makes existing image coding standards inefficient on compressing them. To solve the problem, this paper proposes a novel coding scheme based on the H.264 intra-frame coding. Two new intra modes are proposed to better exploit spatial correlations in compound images. The first is residual scalar quantization (RSQ) mode, where intra-predicted residues are directly quantized and entropy coded. The second is base colors and index map (BCIM) mode that can be viewed as an adaptive vector quantization. In this mode, an image block is represented by several representative colors, called as base colors, and an index map to compress. Two new modes as well as previous intra modes in H.264 are selected by the rate-distortion optimization (RDO) method in each block. Experimental results show that the proposed scheme not only improves the coding efficiency even more than 10 dB for compound images but also keeps the similar performance as H.264 for natural images. Cuiling Lan, Feng Wu 0001, Guangming Shi |
ISCAS | 1 |