Chi Su

dblp:28/8284 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
abstract
We propose SAT-HMR, a one-stage framework for real-time multi-person 3D human mesh estimation from a single RGB image. While current one-stage methods, which follow a DETR-style pipeline, achieve state-of-the-art (SOTA) performance with high-resolution inputs, we observe that this particularly benefits the estimation of individuals in smaller scales of the image (e.g., those of young age or far from the camera), but at the cost of significantly increased computation overhead. To address this, we introduce scale-adaptive tokens that are dynamically adjusted based on the relative scale of each individual in the image within the DETR framework. Specifically, individuals in smaller scales are processed at higher resolutions, larger ones at lower resolutions, and background regions are further distilled. These scale-adaptive tokens more efficiently encode the image features, facilitating subsequent decoding to regress the human mesh, while allowing the model to allocate computational resources more effectively and focus on more challenging cases. Experiments show that our method preserves the accuracy benefits of high-resolution processing while substantially reducing computational cost, achieving real-time inference with performance comparable to SOTA methods.
Chi Su, Xiaoxuan Ma 0001, Jiajun Su, Yizhou Wang 0001
CVPR1
2024 Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video Generation
Kenan Huang, Junbao Zhuo, Shuhui Wang, Chi Su, Qingming Huang, Huimin Ma 0001
ACM Multimedia4
2023 General Greedy De-Bias Learning
abstract
Neural networks often make predictions relying on the spurious correlations from the datasets rather than the intrinsic properties of the task of interest, facing with sharp degradation on out-of-distribution (OOD) test data. Existing de-bias learning frameworks try to capture specific dataset bias by annotations but they fail to handle complicated OOD scenarios. Others implicitly identify the dataset bias by special design low capability biased models or losses, but they degrade when the training and testing data are from the same distribution. In this paper, we propose a General Greedy De-bias learning framework (GGD), which greedily trains the biased models and base model. The base model is encouraged to focus on examples that are hard to solve with biased models, thus remaining robust against spurious correlations in the test stage. GGD largely improves models' OOD generalization ability on various tasks, but sometimes over-estimates the bias level and degrades on the in-distribution test. We further re-analyze the ensemble process of GGD and introduce the Curriculum Regularization inspired by curriculum learning, which achieves a good trade-off between in-distribution (ID) and out-of-distribution performance. Extensive experiments on image classification, adversarial question answering, and visual question answering demonstrate the effectiveness of our method. GGD can learn a more robust base model under the settings of both task-specific biased models with prior knowledge and self-ensemble biased model without prior knowledge. Codes are available at https://github.com/GeraldHan/GGD.
Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Self-Regulated Learning for Egocentric Video Activity Anticipation
abstract
Future activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a simple and effective Self-Regulated Learning framework, which aims to regulate the intermediate representation consecutively to produce representation that (a) emphasizes the novel information in the frame of the current time-stamp in contrast to previously observed content, and (b) reflects its correlation with previously observed frames. The former is achieved by minimizing a contrastive loss, and the latter can be achieved by a dynamic reweighing mechanism to attend to informative frames in the observed content with a similarity comparison between feature of the current frame and observed frames. The learned final video representation can be further enhanced by multi-task learning which performs joint feature learning on the target activity labels and the automatically detected action and object class tokens. SRL sharply outperforms existing state-of-the-art in most cases on two egocentric video datasets and two third-person video datasets. Its effectiveness is also verified by the experimental fact that the action and object concepts that support the activity semantics can be accurately identified.
Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Temporal Dynamic Concept Modeling Network for Explainable Video Event Recognition
abstract
Recently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field.
Weigang Zhang, Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Self-attention neural architecture search for semantic image segmentation
Zhenkun Fan, Guosheng Hu, Xin Sun 0003, Gaige Wang, Junyu Dong, Chi Su
Knowl. Based Syst.6
2022 On the Correlation Among Edge, Pose and Parsing
abstract
Semantic parsing, edge detection, and pose estimation of human are three closely-related tasks. They present human characteristics from three complementary aspects. Compared to learning them individually, solving these tasks jointly can explore the interaction of their contextual cues. However, prior works usually study the fusion of two of them, e.g., parsing and pose, parsing and edge. In this paper, we explore how pixel-level semantics, human boundaries and joint locations can be effectively learned in a unified model. Specifically, we propose an end-to-end trainable Human Task Correlation Machine (HTCorrM) to implement the three tasks. It is asymmetric in that it supports a main task using the other two as auxiliary tasks. We also introduce a Heterogeneous Non-Local module (HNL) to discover the correlations of the three heterogeneous domains. HNL fully explores the global dependency among tasks between any two positions in the feature map. Experimental results on human parsing, pose estimation and body edge detection demonstrate that HTCorrM achieves competitive performance. We show that when designated as the main task, the accuracy of each of the three tasks is improved. Importantly, comparative studies confirm the advantages of our proposed feature correlation strategy over feature concatenation or post processing.
Ziwei Zhang 0003, Chi Su, Liang Zheng 0001, Yuan Li 0014
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Toward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic Segmentation
abstract
Panoptic segmentation aims to partition an image to object instances and semantic content for thing and stuff categories, respectively. To date, learning weakly supervised panoptic segmentation (WSPS) with only image-level labels remains unexplored. In this paper, we propose an efficient jointly thing-and-stuff mining (JTSM) framework for WSPS. To this end, we design a novel mask of interest pooling (MoIPool) to extract fixed-size pixel-accurate feature maps of arbitrary-shape segmentations. MoIPool enables a panoptic mining branch to leverage multiple instance learning (MIL) to recognize things and stuff segmentation in a unified manner. We further refine segmentation masks with parallel instance and semantic segmentation branches via self-training, which collaborates the mined masks from panoptic mining with bottom-up object evidence as pseudo-ground-truth labels to improve spatial coherence and contour localization. Experimental results demonstrate the effectiveness of JTSM on PASCAL VOC and MS COCO. As a by-product, we achieve competitive results for weakly supervised object detection and instance segmentation. This work is a first step towards tackling challenge panoptic segmentation task with only image-level labels.
Yunhang Shen, Liujuan Cao, Feihong Lian, Baochang Zhang 0001, Chi Su, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji
CVPR6
2021 Greedy Gradient Ensemble for Robust Visual Question Answering
abstract
Language bias is a critical issue in Visual Question Answering (VQA), where models often exploit dataset biases for the final decision without considering the image information. As a result, they suffer from performance drop on out-of-distribution data and inadequate visual explanation. Based on experimental analysis for existing robust VQA methods, we stress the language bias in VQA that comes from two aspects, i.e., distribution bias and shortcut bias. We further propose a new de-bias framework, Greedy Gradient Ensemble (GGE), which combines multiple biased models for unbiased base model learning. With the greedy strategy, GGE forces the biased models to over-fit the biased data distribution in priority, thus makes the base model pay more attention to examples that are hard to solve by biased models. The experiments demonstrate that our method makes better use of visual information and achieves state-of-the-art performance on diagnosing dataset VQACP without using extra annotations.
Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, Qi Tian 0001
ICCV3
2021 Parallel Detection-and-Segmentation Learning for Weakly Supervised Instance Segmentation
abstract
Weakly supervised instance segmentation (WSIS) with only image-level labels has recently drawn much attention. To date, bottom-up WSIS methods refine discriminative cues from classifiers with sophisticated multi-stage training procedures, which also suffer from inconsistent object boundaries. And top-down WSIS methods are formulated as cascade detection-to-segmentation pipeline, in which the quality of segmentation learning heavily depends on pseudo masks generated from detectors. In this paper, we propose a unified parallel detection-and-segmentation learning (PDSL) framework to learn instance segmentation with only image-level labels, which draws inspiration from both top-down and bottom-up instance segmentation approaches. The detection module is the same as the typical design of any weakly supervised object detection, while the segmentation module leverages self-supervised learning to model class-agnostic foreground extraction, following by self-training to refine class-specific segmentation. We further design instance-activation correlation module to improve the coherence between detection and segmentation branches. Extensive experiments verify that the proposed method outperforms baselines and achieves the state-of-the-art results on PASCAL VOC and MS COCO.
Yunhang Shen, Liujuan Cao, Baochang Zhang 0001, Chi Su, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji
ICCV5
2021 Neural Architecture Search for Joint Human Parsing and Pose Estimation
abstract
Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process while pose estimation is usually a regression task, it is non-trivial to extract discriminative features for both tasks while modeling their correlation in the joint learning fashion. Recent studies have shown that Neural Architecture Search (NAS) has the ability to allocate efficient feature connections for specific tasks automatically. With the spirit of NAS, we propose to search for an efficient network architecture (NPPNet) to tackle two tasks at the same time. On the one hand, to extract task-specific features for the two tasks and lay the foundation for the further searching of feature interaction, we propose to search their encoder-decoder architectures, respectively. On the other hand, to ensure two tasks fully communicate with each other, we propose to embed NAS units in both multi-scale feature interaction and high-level feature fusion to establish optimal connections between two tasks. Experimental results on both parsing and pose estimation benchmark datasets have demonstrated that the searched model achieves state-of-the-art performances on both tasks.1
Dan Zeng 0001, Yuhang Huang 0006, Qian Bao, Junjie Zhang 0002, Chi Su, Wu Liu 0005
ICCV5
2020 Gradually Vanishing Bridge for Adversarial Domain Adaptation
abstract
In unsupervised domain adaptation, rich domain-specific characteristics bring great challenge to learn domain-invariant representations. However, domain discrepancy is considered to be directly minimized in existing solutions, which is difficult to achieve in practice. Some methods alleviate the difficulty by explicitly modeling domain-invariant and domain-specific parts in the representations, but the adverse influence of the explicit construction lies in the residual domain-specific characteristics in the constructed domain-invariant representations. In this paper, we equip adversarial domain adaptation with Gradually Vanishing Bridge (GVB) mechanism on both generator and discriminator. On the generator, GVB could not only reduce the overall transfer difficulty, but also reduce the influence of the residual domain-specific characteristics in domain-invariant representations. On the discriminator, GVB contributes to enhance the discriminating ability, and balance the adversarial training process. Experiments on three challenging datasets show that our GVB methods outperform strong competitors, and cooperate well with other adversarial methods. The code is available at https://github.com/cuishuhao/GVB.
Shuhao Cui, Shuhui Wang, Junbao Zhuo, Chi Su, Qingming Huang, Qi Tian 0001
CVPR4
2020 Label Decoupling Framework for Salient Object Detection
abstract
To get more accurate saliency maps, recent methods mainly focus on aggregating multi-level features from fully convolutional network (FCN) and introducing edge information as auxiliary supervision. Though remarkable progress has been achieved, we observe that the closer the pixel is to the edge, the more difficult it is to be predicted, because edge pixels have a very imbalance distribution. To address this problem, we propose a label decoupling framework (LDF) which consists of a label decoupling (LD) procedure and a feature interaction network (FIN). LD explicitly decomposes the original saliency map into body map and detail map, where body map concentrates on center areas of objects and detail map focuses on regions around edges. Detail map works better because it involves much more pixels than traditional edge supervision. Different from saliency map, body map discards edge pixels and only pays attention to center areas. This successfully avoids the distraction from edge pixels during training. Therefore, we employ two branches in FIN to deal with body map and detail map respectively. Feature interaction (FI) is designed to fuse the two complementary branches to predict the saliency map, which is then used to refine the two branches again. This iterative refinement is helpful for learning better representations and more precise saliency maps. Comprehensive experiments on six benchmark datasets demonstrate that LDF outperforms state-of-the-art approaches on different evaluation metrics.
Jun Wei 0006, Shuhui Wang, Zhe Wu 0006, Chi Su, Qingming Huang, Qi Tian 0001
CVPR4
2020 Correlating Edge, Pose With Parsing
abstract
According to existing studies, human body edge and pose are two beneficial factors to human parsing. The effectiveness of each of the high-level features (edge and pose) is confirmed through the concatenation of their features with the parsing features. Driven by the insights, this paper studies how human semantic boundaries and keypoint locations can jointly improve human parsing. Compared with the existing practice of feature concatenation, we find that uncovering the correlation among the three factors is a superior way of leveraging the pivotal contextual cues provided by edges and poses. To capture such correlations, we propose a Correlation Parsing Machine (CorrPM) employing a heterogeneous non-local block to discover the spatial affinity among feature maps from the edge, pose and parsing. The proposed CorrPM allows us to report new state-of-the-art accuracy on three human parsing datasets. Importantly, comparative studies confirm the advantages of feature correlation over the concatenation.
Ziwei Zhang 0003, Chi Su, Liang Zheng 0001
CVPR2
2020 Interpretable Visual Reasoning via Probabilistic Formulation Under Natural Supervision
Xinzhe Han, Shuhui Wang, Chi Su, Weigang Zhang, Qingming Huang, Qi Tian 0001
ECCV (9)3
2020 Towards More Explainability: Concept Knowledge Mining Network for Event Recognition
abstract
Event recognition of untrimmed video is a challenging task due to the big gap between low level visual features and event semantics. Beyond feature learning via deep neural networks, some recent works focus on analyzing event videos using concept-based representation. However, these methods simply aggregate the concept representation vectors of frames or segments, which inevitably introduces information loss on video-level concept knowledge. Moreover, the diversified relation between different concept domains (e.g., scene, object and action) has not been fully explored. To address the above issues, we propose a concept knowledge mining network (CKMN) for event recognition. CKMN is composed of an intra-domain concept knowledge mining subnetwork (IaCKM) and an inter-domain concept knowledge mining subnetwork~(IrCKM). IaCKM aims to obtain a complete concept representation by mining the existing pattern of each concept at different time granularities with dilated temporal pyramid convolution and temporal self-attention, while IrCKM explores the interaction between different types of concepts with co-attention style learning. We evaluate our method on FCVID and ActivityNet datasets. Experimental results show the effectiveness and better interpretability of our model on event analytics. Code is available at https://github.com/qzhb/CKMN.
Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001
ACM Multimedia3
2020 Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video Analysis
abstract
Event analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the spatial range covered by a single feature response, which is crucial in improving the image categorization accuracy. In video domain, video event semantics are actually described by complex interaction among different concepts, while their behaviors vary drastically from one video to another, leading to the difficulty in concept-based analytics for accurate event categorization. To model the concept behavior, we study temporal concept receptive field of concept-based event representation, which encodes the temporal occurrence pattern of different mid-level concepts. Accordingly, we introduce temporal dynamic convolution (TDC) to give stronger flexibility to concept-based event analytics. TDC can adjust the temporal concept receptive field size dynamically according to different inputs. Notably, a set of coefficients are learned to fuse the results of multiple convolutions with different kernel widths that provide various temporal concept receptive field sizes. Different coefficients can generate appropriate and accurate temporal concept receptive field size according to input videos and highlight crucial concepts. Based on TDC, we propose the temporal dynamic concept modeling network~(TDCMN) to learn an accurate and complete concept representation for efficient untrimmed video analysis. Experiment results on FCVID and ActivityNet show that TDCMN demonstrates adaptive event recognition ability conditioned on different inputs, and improve the event recognition performance of Concept-based methods by a large margin. Code is available at https://github.com/qzhb/TDCMN.
Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Weigang Zhang, Qingming Huang
ACM Multimedia3
2019 SIXray: A Large-Scale Security Inspection X-Ray Benchmark for Prohibited Item Discovery in Overlapping Images
abstract
In this paper, we present a large-scale dataset and establish a baseline for prohibited item discovery in Security Inspection X-ray images. Our dataset, named SIXray, consists of 1,059,231 X-ray images, in which 6 classes of 8,929 prohibited items are manually annotated. It raises a brand new challenge of overlapping image data, meanwhile shares the same properties with existing datasets, including complex yet meaningless contexts and class imbalance. We propose an approach named class-balanced hierarchical refinement (CHR) to deal with these difficulties. CHR assumes that each input image is sampled from a mixture distribution, and that deep networks require an iterative process to infer image contents accurately. To accelerate, we insert reversed connections to different network backbones, delivering high-level visual cues to assist mid-level features. In addition, a class-balanced loss function is designed to maximally alleviate the noise introduced by easy negative samples. We evaluate CHR on SIXray with different ratios of positive/negative samples. Compared to the baselines, CHR enjoys a better ability of discriminating objects especially using mid-level features, which offers the possibility of using a weakly-supervised approach towards accurate object localization. In particular, the advantage of CHR is more significant in the scenarios with fewer positive training samples, which demonstrates its potential application in real-world security inspection.
Caijing Miao, Lingxi Xie, Fang Wan 0001, Chi Su, Hongye Liu, Jianbin Jiao, Qixiang Ye
CVPR4
2019 Iterative Reorganization With Weak Spatial Constraints: Solving Arbitrary Jigsaw Puzzles for Unsupervised Representation Learning
abstract
Learning visual features from unlabeled image data is an important yet challenging task, which is often achieved by training a model on some annotation-free information. We consider spatial contexts, for which we solve so-called jigsaw puzzles, i.e., each image is cut into grids and then disordered, and the goal is to recover the correct configuration. Existing approaches formulated it as a classification task by defining a fixed mapping from a small subset of configurations to a class set, but these approaches ignore the underlying relationship between different configurations and also limit their applications to more complex scenarios. This paper presents a novel approach which applies to jigsaw puzzles with an arbitrary grid size and dimensionality. We provide a fundamental and generalized principle, that weaker cues are easier to be learned in an unsupervised manner and also transfer better. In the context of puzzle recognition, we use an iterative manner which, instead of solving the puzzle all at once, adjusts the order of the patches in each step until convergence. In each step, we combine both unary and binary features of each patch into a cost function judging the correctness of the current configuration. Our approach, by taking similarity between puzzles into consideration, enjoys a more efficient way of learning visual knowledge. We verify the effectiveness of our approach from two aspects. First, it solves arbitrarily complex puzzles, including high-dimensional puzzles, that prior methods are difficult to handle. Second, it serves as a reliable way of network initialization, which leads to better transfer performance in visual recognition tasks including classification, detection and segmentation.
Chen Wei 0005, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu 0001, Qi Tian 0001, Alan L. Yuille
CVPR5
2019 Snapshot Distillation: Teacher-Student Optimization in One Generation
abstract
Optimizing a deep neural network is a fundamental task in computer vision, yet direct training methods often suffer from over-fitting. Teacher-student optimization aims at providing complementary cues from a model trained previously, but these approaches are often considerably slow due to the pipeline of training a few generations in sequence, i.e., time complexity is increased by several times. This paper presents snapshot distillation (SD), the first framework which enables teacher-student optimization in one generation. The idea of SD is very simple: instead of borrowing supervision signals from previous generations, we extract such information from earlier epochs in the same generation, meanwhile make sure that the difference between teacher and student is sufficiently large so as to prevent under-fitting. To achieve this goal, we implement SD in a cyclic learning rate policy, in which the last snapshot of each cycle is used as the teacher for all iterations in the next cycle, and the teacher signal is smoothed to provide richer information. In standard image classification benchmarks such as CIFAR100 and ILSVRC2012, SD achieves consistent accuracy gain without heavy computational overheads. We also verify that models pre-trained with SD transfers well to object detection and semantic segmentation in the PascalVOC dataset.
Lingxi Xie, Chi Su, Alan L. Yuille
CVPR3
2019 Global and Local Deep Feature Representation Fusion for Vehicle Re-Identification
abstract
This paper introduces our submission to the Grand Challenges on Vehicle Re-Identification (ReID) held in the VCIP 2019. Vehicle Re-Identification, which aims to retrieve images of a query vehicle from a large-scale vehicle database, is of great significance to the urban security and city management. Although significant progress has been made in the last decade, vehicle ReID in the wild remains a very challenging problem due to the large intra-class variations of one vehicle instance from viewpoint, illumination, occlusion patterns, and the possibly small inter-class differentiation of different vehicle instances (e.g., two different vehicles of the same model and color). To deal with these challenges, we in this work present a vehicle ReID framework that integrates both the global visual cues along with the local part-based cues to learn discriminative feature representations. In addition, the proposed framework also makes use of the extra information such as brands, models and colors to further improve the performance. Experimental results are performed on the VCIP 2019 VehicleReID dataset and the proposed framework achieved the second place in the competition.
Xing Zhao Lee, Jiangtao Kong, Chi Su, Junliang Xing, Shengmei Shen
VCIP4
2018 Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age Estimation
abstract
Facial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estimation within each population separately. In practice, however, it is often very difficult to collect and label sufficient data for each population. Therefore, it would be helpful to exploit an existing large labeled dataset of one (source) population to improve the age estimation performance on another (target) population with only a small labeled dataset available. In this work, we propose a Deep Cross-Population (DCP) age estimation model to achieve this goal. In particular, our DCP model develops a two-stage training strategy. First, a novel cost-sensitive multitask loss function is designed to learn transferable aging features by training on the source population. Second, a novel order-preserving pair-wise loss function is designed to align the aging features of the two populations. By doing so, our DCP model can transfer the knowledge encoded in the source population to the target population. Extensive experiments on the two of the largest benchmark datasets show that our DCP model outperforms several strong baseline methods and many state-of-the-art methods.
Kai Li 0022, Junliang Xing, Chi Su, Weiming Hu 0004, Stephen J. Maybank
CVPR3
2018 Multi-Task Learning with Low Rank Attribute Embedding for Multi-Camera Person Re-Identification
abstract
We propose Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) to address the problem of person re-identification on multi-cameras. Re-identifications on different cameras are considered as related tasks, which allows the shared information among different tasks to be explored to improve the re-identification accuracy. The MTL-LORAE framework integrates low-level features with mid-level attributes as the descriptions for persons. To improve the accuracy of such description, we introduce the low-rank attribute embedding, which maps original binary attributes into a continuous space utilizing the correlative relationship between each pair of attributes. In this way, inaccurate attributes are rectified and missing attributes are recovered. The resulting objective function is constructed with an attribute embedding error and a quadratic loss concerning class labels. It is solved by an alternating optimization strategy. The proposed MTL-LORAE is tested on four datasets and is validated to outperform the existing methods with significant margins.
Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Multi-type attributes driven multi-camera person re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
Pattern Recognit.1
2017 Pose-Driven Deep Convolutional Model for Person Re-identification
abstract
Feature extraction and matching are two crucial components in person Re-Identification (ReID). The large pose deformations and the complex view variations exhibited by the captured person images significantly increase the difficulty of learning and matching of the features from person images. To overcome these difficulties, in this work we propose a Pose-driven Deep Convolutional (PDC) model to learn improved feature extraction and matching models from end to end. Our deep architecture explicitly leverages the human part cues to alleviate the pose variations and learn robust feature representations from both the global image and different local parts. To match the features from global human body and local body parts, a pose driven feature weighting sub-network is further designed to learn adaptive feature fusions. Extensive experimental analyses and results on three popular datasets demonstrate significant performance improvements of our model over all published state-of-the-art methods.
Chi Su, Jianing Li 0001, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
ICCV1
2017 Attributes driven tracklet-to-tracklet person re-identification using latent prototypes space mapping
Chi Su, Shiliang Zhang, Fan Yang 0016, Guangxiao Zhang, Qi Tian 0001, Wen Gao 0001, Larry Davis 0001
Pattern Recognit.1
2016 Deep Attributes Driven Multi-camera Person Re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
ECCV (2)1
2016 MARS: A Video Benchmark for Large-Scale Person Re-Identification
Liang Zheng 0001, Zhi Bie, Yifan Sun 0003, Jingdong Wang 0001, Chi Su, Shengjin Wang, Qi Tian 0001
ECCV (6)5
2015 Multi-Task Learning with Low Rank Attribute Embedding for Person Re-Identification
abstract
We propose a novel Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) framework for person re-identification. Re-identifications from multiple cameras are regarded as related tasks to exploit shared information to improve re-identification accuracy. Both low level features and semantic/data-driven attributes are utilized. Since attributes are generally correlated, we introduce a low rank attribute embedding into the MTL formulation to embed original binary attributes to a continuous attribute space, where incorrect and incomplete attributes are rectified and recovered to better describe people. The learning objective function consists of a quadratic loss regarding class labels and an attribute embedding error, which is solved by an alternating optimization procedure. Experiments on three person re-identification datasets have demonstrated that MTL-LORAE outperforms existing approaches by a large margin and produces state-of-the-art results.
Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001
ICCV1
2014 Active power dispatch method for a wind farm central controller considering wake effect
abstract
With the increasing integration of the wind power into the power system, wind farm are required to be controlled as a single unit and have all the same control tasks as conventional power plants. The wind farm central controller receives control orders from Transmission System Operator (TSO), then dispatch the wind power reference to each wind turbine. One of the most commonly used dispatch methods is to dispatch the wind power reference to each wind turbine proportional to each wind turbine's available wind power without the consideration of the wake effect. The wake which depends on the thrust efficient of upstream wind turbines in the wind farm influences the downstream wind speed which determines the available wind power of the downstream wind turbine. Optimize the wind power production of each wind turbine in the wind farm by the optimization of the pitch angle and tip-speed-ratio of each turbine can increase the total available wind power of the wind farm. This paper proposed a new dispatch method which dispatches the wind power reference to each wind turbine proportional to each wind turbine's optimal wind power to increase the available wind power of the whole wind farm. Particle Swarm Optimization (PSO) is used to obtain the optimal wind power for each wind turbine. A case study is carried out. The available wind power of the wind farm was compared between the traditional dispatch method and the proposed dispatch method with the consideration of the wake effect.
Chi Su, Mohsen Soltani, Zhe Chen 0007
IECON2
2013 A system based on sequence learning for event detection in surveillance video
abstract
Event detection in crowded surveillance videos is a challenging yet important problem. In this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVid'12 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) pair-wise events (e.g., PeopleMeet, PeopleSplitUp and Embrace); 2) action-like events (e.g., ObjectPut, CellToEar, PersonRuns and Pointing). In eSur system, we first employ people detection and tracking algorithms to locate target persons in 3D space-time domain. Then the video sequences in which target persons occur are partitioned into several spatio-temporal cubes. Visual features (i.e. cubic feature and MoSIFT) are computed over these cubes. After that, a sequence learning method, (namely SVM with dynamic time alignment kernel), is employed to infer the existence of an event for the video sequence. According to the TRECVid SED formal evaluation, eSur has yielded fairly encouraging results on TRECVid'12 dataset.
Xiaoyu Fang, Ziwei Xia, Chi Su, Teng Xu 0002, Yonghong Tian 0001, Yaowei Wang 0001, Tiejun Huang 0001
ICIP3
2013 Pair-wise event detection using cubic features and sequence discriminant learning
abstract
Event detection in crowded surveillance videos is a challenging yet important problem. This paper focuses on pair-wise events that involve the interaction of two persons (e.g., people embrace, meet or split) in crowded videos. To detect such an event accurately, we should build an effective representation model that can characterize the sequential properties of two persons' interaction. Towards this end, we propose a novel pair-wise event detection approach using cubic features and sequence discriminant learning. A video sequence is first partitioned into several spatio-temporal cubes, and multiple features (e.g., statistics of trajectories, bag of spatio-temporal interest points) are extracted on these cubes and then fused to form a cubic feature descriptor under multiple kernel learning (MKL) framework. After that, the SVM with dynamic time alignment kernel is used to infer the existence of an event in the video sequence. Experimental results show that the proposed approach achieves the encouraging performance on TRECVid SED dataset.
Xiaoyu Fang, Yonghong Tian 0001, Yaowei Wang 0001, Chi Su, Teng Xu 0002, Ziwei Xia, Wen Gao 0001
ICME4
2013 Comparison study of power system small signal stability improvement using SSSC and STATCOM
abstract
A static synchronous series compensator (SSSC) has the ability to emulate a reactance in series with the connected transmission line. A static synchronous compensator (STATCOM) is able to provide the reactive power to an electricity network. When fed with some supplementary signals from the connected power system, both SSSC and STATCOM are able to participate in the power system inter-area oscillation damping by changing the compensated reactance or the provided reactive power. This paper analyses the influence of SSSC and STATCOM on power system small signal stability. The damping controller schemes for SSSC and STATCOM are presented and discussed. The IEEE 39-bus New England system model as the test system is built in DIgSIELNT PowerFactory, in which the damping control strategies for both SSSC and STATCOM are validated by time domain simulations and modal analysis. Furthermore, comparison studies show that the SSSC is a better solution in term of equipment capabilities and costs.
Weihao Hu, Chi Su, Jiakun Fang, Zhe Chen 0007
IECON2
2013 Residue-based coordinated selection and parameter design of multiple power system stabilizers (PSSs)
abstract
Residue method is a commonly used approach to design the parameters of a power system stabilizer (PSS). In this paper, a residue identification method is adopted to obtain the system residues for different input-output pairs, using the system measurements data from time domain simulations. Then a coordinated approach for multiple PSS selection and parameter design based on residue method is proposed and formatted as an optimization problem. Particle swarm optimization (PSO) is adopted in this coordination process to find suitable parameters for PSSs so that the dominant oscillation modes can be well damped; while locations and input signals of PSSs are selected to keep PSS outputs small. The IEEE 39-bus New England system model as the test system is built in DIgSIELNT PowerFactory 14.0, in which the proposed coordination method is validated by time domain simulations and modal analysis.
Chi Su, Weihao Hu, Jiakun Fang, Zhe Chen 0007
IECON1
2013 Reactive power capability of the wind turbine with Doubly Fed Induction Generator
abstract
With the increasing integration into power grids, wind power plants play an important role in the power system. Many requirements for the wind power plants have been proposed in the grid codes. According to these grid codes, wind power plants should have the ability to perform voltage control and reactive power compensation at the point of common coupling (PCC). Besides the shunt flexible alternating current transmission system (FACTS) devices such as the static var compensator (SVC) and the static synchronous compensator (STATCOM), the wind turbine itself can also provide a certain amount of reactive power compensation, depending on the wind speed and the active power control strategy. This paper analyzes the reactive power capability of Doubly Fed Induction Generator (DFIG) based wind turbine, considering the rated stator current limit, the rated rotor current limit, the rated rotor voltage limit, and the reactive power capability of the grid side convertor (GSC). The boundaries of reactive power capability of DFIG based wind turbine are derived. The result was obtained using the software MATLAB.
Chi Su, Zhe Chen 0007
IECON2
2012 Single and Multiple View Detection, Tracking and Video Analysis in Crowded Environments
abstract
In this paper, we present our detection, tracking and event recognition methods and the results for PETS 2012. First, ROIs (Regions of Interest) based on geometric constraints are utilized in single view detection to eliminate the negative influence of clutter environment. Then, an optimized observation model is applied to address the ID switching or tracking drifting problem in single view tracking. Third, we introduce the multi-view Bayesian network (MBN) to reduce the "phantom" phenomena which frequently happen in general multi-view detection tasks. At last, a motion-based event recognition method is proposed to handle the event recognition task. Experimental results on the PETS 2012 dataset indicate that our methods are very promising.
Teng Xu 0002, Peixi Peng, Xiaoyu Fang, Chi Su, Yaowei Wang 0001, Yonghong Tian 0001, Wei Zeng 0006, Tiejun Huang 0001
AVSS4
2010 Learning classifier system using both labeled and unlabeled data
abstract
In this paper, we propose a Semi-UCS, which is an extension of the classical sUpervised Classifier System (UCS) [1] for semi-supervised learning tasks. A UCS works under a supervised learning scheme and uses only labeled data to train the system. In the Semi-UCS, we add an additional semi-supervised learning component to the original UCS, enablingthe LCS to learn from both labeled and unlabeled data. We provide three methods of how this semi-supervised learning component can be implemented: self-learning method, k-NN distance measure method and tri-training method. The Semi-UCS enlarges UCS's' application domains into semi-supervised settings and is a great addition to the LCS's model family. Experimental results on benchmark data sets of UCI repository have shown that Semi-UCS reaches a good performance for semi-supervised learning tasks.
Chi Su, Yang Gao 0001, Chun Cao
GECCO1