VLDB 2026 Research / reviewers in the wild / expert
Sheng Tang
dblp:62/1647
· DBLP profile ↗
97ranked-venue papers
10as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 33 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorSystems, architecture and hardware · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image DetectionabstractThe rapid evolution of generative technologies necessitates reliable methods for detecting AI-generated images. A critical limitation of current detectors is their failure to generalize to images from unseen generative models, as they often overfit to source-specific semantic cues rather than learning universal generative artifacts. To overcome this, we introduce a simple yet remarkably effective pixel-level mapping pre-processing step to disrupt the pixel value distribution of images and break the fragile, non-essential semantic patterns that detectors commonly exploit as shortcuts. This forces the detector to focus on more fundamental and generalizable high-frequency traces inherent to the image generation process. Through comprehensive experiments on GAN and diffusion-based generators, we show that our approach significantly boosts the cross-generator performance of state-of-the-art detectors. Extensive analysis further verifies our hypothesis that the disruption of semantic cues is the key to generalization. Chenming Zhou, Jiaan Wang, Yu Li 0016, Juan Cao 0001, Sheng Tang |
AAAI | 6 |
| 2026 | Dissecting Deepfake Artifacts via Multimodal Explanations
Yannan Bai, Danding Wang, Sheng Tang, Juan Cao 0001, Jintao Li 0001 |
MMM (1) | 3 |
| 2026 | IDFreq: Identity-preserved human video generation via frequency-based decomposition
Zhang Wan, Sheng Tang, Juan Cao 0001, Yu Li 0016 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | Semi-supervised adversarial training via disentangled contrastive learning
Ruize Zhang 0003, Yu Li 0016, Juan Cao 0001, Sheng Tang |
Multim. Syst. | 4 |
| 2026 | MoAnimate: Bridging the Motion-Oriented Latent Representation Gaps in Human Video AnimationabstractHuman animation strives to bring static characters to life. Existing methods produce high-quality outcomes for single-frame animation; however, they often fail to maintain satisfactory temporal consistency, especially in facial and hand movements. This limitation arises from commonly used motion modules that do not explicitly model inter-entity relationships. In this work, we introduce MoAnimate, a Motion-oriented Human Animation framework designed to improve inter-entity consistency. Specifically, we extract motion flows from driving videos and transfer them to align the shape of character. During initialization, we propose a motion-oriented latent refinement that optimizes low-frequency subbands to regulate the layout of visual objects along flow trajectories, while preserving random high-frequency subbands to accommodate appearance variations. During denoising, we further introduce a motion-oriented entity attention module to enable direct and efficient interaction among entities within a coordinated subspace. Extensive experiments demonstrate that our method significantly enhances temporal consistency, particularly the visual consistency of the entities. Haipeng Fang, Sheng Tang, Ziyao Huang 0002, Juan Cao 0001, Fan Tang, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT AccelerationabstractDiffusion transformers have shown exceptional performance in visual generation but incur high computational costs. Token reduction techniques that compress models by sharing the denoising process among similar tokens have been introduced. However, existing approaches neglect the denoising priors of the diffusion models, leading to suboptimal acceleration and diminished image quality. This study proposes a novel concept: attend to prune feature redundancies in areas not attended by the diffusion process. We analyze the location and degree of feature redundancies based on the structure-then-detail denoising priors. Subsequently, we introduce SDTM, a structure-then-detail token merging approach that dynamically compresses feature redundancies. Specifically, we design dynamic visual token merging, compression ratio adjusting, and prompt reweighting for different stages. Served in a post-training way, the proposed method can be integrated seamlessly into any DiT architecture. Extensive experiments across various backbones, schedulers, and datasets showcase the superiority of our method, for example, it achieves 1.55× acceleration with negligible impact on image quality. Project page: https://github.com/ICTMCG/SDTM. Haipeng Fang, Sheng Tang, Juan Cao 0001, Enshuo Zhang, Fan Tang, Tong-Yee Lee |
CVPR | 2 |
| 2025 | FR2ViT: Finetuning-free Token Reduction for Dense Prediction Through a Refinement-Reactivation ArchitectureabstractToken reduction is an efficient method for accelerating vision transformers. Techniques like token pruning and merging progressively decrease the number of active tokens to reduce the computation cost. However, when applied to dense prediction tasks, these techniques crudely cache low-level features or unfold non-self features to reconstruct the inactive tokens. This hard replacement results in significant information loss, especially at higher compression ratios. In this work, we propose FR2ViT, a finetuning-free refinement-reactivation architecture for efficient vision transformers, allowing tokens of varying informativeness to activate softly at differing costs. Specifically, we design a high-order informativeness assessment to partition tokens into superior and inferior categories accurately. We refine superior tokens through extensive attention interaction, while inferior tokens undergo effective reactivation, and we establish a lightweight interaction between two branches. Our approach ensures continuous and autonomous interaction of each token, and is finetuning-free to easily accommodate the large model era. Experiments across three dense prediction tasks demonstrate our method’s exceptional performance. For example, FR2ViT achieves a 1.95× FPS with only a 0.43% performance drop for Segmenter ViT-L. Haipeng Fang, Xinyi Zou, Jun Huang 0007, Juan Cao 0001, Sheng Tang |
ICASSP | 6 |
| 2025 | Latent inversion for consistent identity preservation in character animation
Sheng Tang, Zhang Wan, Juan Cao 0001, Jintao Li 0001 |
Vis. Comput. | 2 |
| 2024 | Topology-preserving Adversarial Training for Alleviating Natural Accuracy Degradation
Xiaoyue Mi, Fan Tang, Yepeng Weng, Danding Wang, Juan Cao 0001, Sheng Tang, Peng Li 0030, Yang Liu 0005 |
BMVC | 6 |
| 2024 | DragEntity: Trajectory Guided Video Generation using Entity and Positional RelationshipsabstractIn recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it's challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. Compared to previous methods, DragEntity offers two main advantages: 1) Our method is more user-friendly for interaction because it allows users to drag entities within the image rather than individual pixels. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its excellent performance in fine-grained control in video generation. Zhang Wan, Sheng Tang, Jiawei Wei, Ruize Zhang 0003, Juan Cao 0001 |
ACM Multimedia | 2 |
| 2024 | Self-Supervised Adversarial Training via Diverse Augmented Queries and Self-Supervised Double PerturbationabstractRecently, there have been some works studying self-supervised adversarial training, a learning paradigm that learns robust features without labels. While those works have narrowed the performance gap between self-supervised adversarial training (SAT) and supervised adversarial training (supervised AT), a well-established formulation of SAT and its connections with supervised AT are under-explored. Based on a simple SAT benchmark, we find that SAT still faces the problem of large robust generalization gap and degradation on natural samples. We hypothesize this is due to the lack of data complexity and model regularization and propose a method named as DAQ-SDP (Diverse Augmented Queries Self-supervised Double Perturbation). We first challenge the previous conclusion that complex data augmentations degrade robustness in SAT by using diversely augmented samples as queries to guide adversarial training. Inspired by previous works in supervised AT, we then incorporate a self-supervised double perturbation scheme to self-supervised learning (SSL), which promotes robustness transferable to downstream classification. Our work can be seamlessly combined with models pretrained by different SSL frameworks without revising the learning objectives and helps to bridge the gap between SAT and AT. Our method also improves both robust and natural accuracies across different SSL frameworks. Our code is available at https://github.com/rzzhang222/DAQ-SDP. Ruize Zhang 0003, Sheng Tang, Juan Cao 0001 |
NeurIPS | 2 |
| 2024 | Identity-Preserving Face Swapping via Dual Surrogate Generative ModelsabstractIn this study, we revisit the fundamental setting of face-swapping models and reveal that only using implicit supervision for training leads to the difficulty of advanced methods to preserve the source identity. We propose a novel reverse pseudo-input generation approach to offer supplemental data for training face-swapping models, which addresses the aforementioned issue. Unlike the traditional pseudo-label-based training strategy, we assume that arbitrary real facial images could serve as the ground-truth outputs for the face-swapping network and try to generate corresponding input pair data. Specifically, we involve a source-creating surrogate that alters the attributes of the real image while keeping the identity, and a target-creating surrogate intends to synthesize attribute-preserved target images with different identities. Our framework, which utilizes proxy-paired data as explicit supervision to direct the face-swapping training process, partially fulfills a credible and effective optimization direction to boost the identity-preserving capability. We design explicit and implicit adaption strategies to better approximate the explicit supervision for face swapping. Quantitative and qualitative experiments on FF++, FFHQ, and wild images show that our framework could improve the performance of various face-swapping pipelines in terms of visual fidelity and ID preserving. Furthermore, we display applications with our method on re-aging, swappable attribute customization, cross-domain, and video face swapping. Code is available under https://github.com/ ICTMCG/CSCS. Ziyao Huang 0002, Fan Tang, Yong Zhang 0034, Juan Cao 0001, Sheng Tang, Jintao Li 0001, Tong-Yee Lee |
ACM Trans. Graph. | 6 |
| 2023 | Progressive Open Space Expansion for Open-Set Model AttributionabstractDespite the remarkable progress in generative technology, the Janus-faced issues of intellectual property protection and malicious content supervision have arisen. Efforts have been paid to manage synthetic images by attributing them to a set of potential source models. However, the closed-set classification setting limits the application in real-world scenarios for handling contents generated by arbitrary models. In this study, we focus on a challenging task, namely Open-Set Model Attribution (OSMA), to simultaneously attribute images to known models and identify those from unknown ones. Compared to existing openset recognition (OSR) tasks focusing on semantic novelty, OSMA is more challenging as the distinction between images from known and unknown models may only lie in visually imperceptible traces. To this end, we propose a Progressive Open Space Expansion (POSE) solution, which simulates open-set samples that maintain the same semantics as closed-set samples but embedded with different imperceptible traces. Guided by a diversity constraint, the open space is simulated progressively by a set of lightweight augmentation models. We consider three real-world scenarios and construct an OSMA benchmark dataset, including unknown models trained with different random seeds, architectures, and datasets from known ones. Extensive experiments on the dataset demonstrate POSE is superior to both existing model attribution methods and off-the-shelf OSR methods. Github: https://github.com/ICTMCG/POSE Tianyun Yang, Danding Wang, Fan Tang, Juan Cao 0001, Sheng Tang |
CVPR | 6 |
| 2023 | VoxSeP: semi-positive voxels assist self-supervised 3D medical segmentation
Zijie Yang, Lingxi Xie, Xinyue Huo, Longhui Wei, Qi Tian 0001, Sheng Tang |
Multim. Syst. | 8 |
| 2022 | Finding the Host from the Lesion by Iteratively Mining the Registration GraphabstractVoxel-level annotation has always been a burden of training medical image segmentation models. This paper investigates an interesting problem that finds the host organ of a lesion without actually labeling the organ. To remedy the missing annotation, we construct a graph using an off-the-shelf registration algorithm, on which lesion labels over the training set are accumulated to obtain the pseudo organ for each case. These pseudo labels are used to train a deep network, whose predictions determine the affinity of each lesion on the registration graph. We iteratively update the pseudo labels with the affinity until the training convergence. Our method is evaluated on the MSD Liver and KiTS datasets, without seeing any organ annotation, we achieve the test Dice score of 93% for liver and 92% for kidney, and boosts the accuracy of tumor segmentation to a considerable degree, $3%$, which even surpasses the model trained with ground-truth of both organ and tumor. Zijie Yang, Lingxi Xie, Xinyue Huo, Sheng Tang, Qi Tian 0001, Yongdong Zhang 0001 |
ACM Multimedia | 4 |
| 2022 | Temporal Correlation-Diversity Representations for Video-Based Person Re-Identification
Litong Gong, Rui Zhang 0040, Sheng Tang, Juan Cao 0001 |
PRCV (1) | 3 |
| 2022 | Adaptive Spatial Location With Balanced Loss for Video CaptioningabstractMany pioneering approaches have verified the effectiveness of utilizing the global temporal and local object information for video understanding tasks and have achieved significant progress. However, existing methods utilize object detectors to extract all objects overall video frames. This may bring performance degradation due to the information redundancy both spatially and temporally. To address this problem, we propose an adaptive spatial location module for the video captioning task which dynamically predicts an important position of each video frame in the procedure of generating the description sentence. The proposed adaptive spatial location method not only makes our model focus on local object information, but also reduces time and memory consumption brought by the temporal redundancy in extensive video frames and improves the accuracy of generated description. Besides, we propose a balanced loss function to address the class imbalance problem existing in training data. The proposed balanced loss assigns different weight to each word of ground-truth sentence in the training process which can generate more diversified description sentences. Extensive experimental results on the MSVD and MSR-VTT dataset show that the proposed method achieves competitive performance compared to state-of-the-art methods. Linghui Li 0001, Yongdong Zhang 0001, Sheng Tang, Lingxi Xie, Xiaoyong Li 0003, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Mixed Dish Recognition With Contextual Relation and Domain AlignmentabstractMixed dish is a food category that contains different dishes mixed in one plate, and is popular in Eastern and Southeast Asia. Recognizing the individual dishes in a mixed dish image is important for health related applications, e.g. to calculate the nutrition values of the dish. However, most existing methods that focus on single dish classification are not applicable to the recognition of mixed dish images. The main challenge of mixed dish recognition comes from three aspects: a wide range of dish types, the complex dish combination with severe overlap between different dishes and the large visual variances of same dish type caused by different cooking/cutting methods applied in different canteens. In order to tackle these problems, we propose the contextual relation network that encodes the implicit and explicit contextual relations among multiple dishes from region-level features and label-level co-occurrence respectively. Besides, to address the visual variances of dish instances from different canteens, we introduce the domain adaption networks to align both local and global features, and eliminating domain gaps of dish features across different canteens. In addition, we collect a mixed dish image dataset containing 9254 mixed dish images from 6 canteens in Singapore. Extensive experiments on both our dataset and public one validate that our methods can achieve top performance for localizing and recognizing multiple dishes and solve the domain shift problem to a certain extent in mixed dish images. Lixi Deng, Jingjing Chen 0001, Chong-Wah Ngo, Qianru Sun, Sheng Tang, Yongdong Zhang 0001, Tat-Seng Chua |
IEEE Trans. Multim. | 5 |
| 2022 | Consensus Feature Network for Scene ParsingabstractScene parsing is challenging as it aims to assign one of the semantic categories to each pixel in scene images. Thus, pixel-level features are desired for scene parsing. However, classification networks are dominated by the discriminative portion, so directly applying classification networks to scene parsing will result in inconsistent parsing predictions within one instance and among instances of the same category. To address this problem, we propose two transform units to learn pixel-level consensus features. One is an Instance Consensus Transform (ICT) unit to learn the instance-level consensus features by aggregating features within the same instance. The other is a Category Consensus Transform (CCT) unit to pursue category-level consensus features through keeping the consensus of features among instances of the same category in scene images. The proposed ICT and CCT units are lightweight, data-driven and end-to-end trainable. The features learned by the two units are more coherent in both instance-level and category-level. Furthermore, we present the Consensus Feature Network (CFNet) based on the proposed ICT and CCT units, and demonstrate the effectiveness of each component in our method by performing extensive ablation experiments. Finally, our proposed CFNet achieves competitive performance on four datasets, including Cityscapes, Pascal Context, CamVid, and COCO Stuff. Sheng Tang, Rui Zhang 0040, Guodong Guo |
IEEE Trans. Multim. | 2 |
| 2021 | CGNet: A Light-Weight Context Guided Network for Semantic SegmentationabstractThe demand of applying semantic segmentation model on mobile devices has been increasing rapidly. Current state-of-the-art networks have enormous amount of parameters hence unsuitable for mobile devices, while other small memory footprint models follow the spirit of classification network and ignore the inherent characteristic of semantic segmentation. To tackle this problem, we propose a novel Context Guided Network (CGNet), which is a light-weight and efficient network for semantic segmentation. We first propose the Context Guided (CG) block, which learns the joint feature of both local feature and surrounding context effectively and efficiently, and further improves the joint feature with the global context. Based on the CG block, we develop CGNet which captures contextual information in all stages of the network. CGNet is specially tailored to exploit the inherent property of semantic segmentation and increase the segmentation accuracy. Moreover, CGNet is elaborately designed to reduce the number of parameters and save memory footprint. Under an equivalent number of parameters, the proposed CGNet significantly outperforms existing light-weight segmentation networks. Extensive experiments on Cityscapes and CamVid datasets verify the effectiveness of the proposed approach. Specifically, without any post-processing and multi-scale testing, the proposed CGNet achieves 64.8% mean IoU on Cityscapes with less than 0.5 M parameters. Sheng Tang, Rui Zhang 0040, Juan Cao 0001, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxabstractSolving long-tail large vocabulary object detection with deep learning based models is a challenging and demanding task, which is however under-explored. In this work, we provide the first systematic analysis on the underperformance of state-of-the-art models in front of long-tail distribution. We find existing detection methods are unable to model few-shot classes when the dataset is extremely skewed, which can result in classifier imbalance in terms of parameter magnitude. Directly adapting long-tail classification models to detection frameworks can not solve this problem due to the intrinsic difference between detection and classification. In this work, we propose a novel balanced group softmax (BAGS) module for balancing the classifiers within the detection frameworks through group-wise training. It implicitly modulates the training process for the head and tail classes and ensures they are both sufficiently trained, without requiring any extra sampling for the instances from the tail classes. Extensive experiments on the very recent long-tail large vocabulary object recognition benchmark LVIS show that our proposed BAGS significantly improves the performance of detectors with various backbones and frameworks on both object detection and instance segmentation. It beats all state-of-the-art methods transferred from long-tail image classification and establishes new state-of-the-art. Code is available at https://github.com/FishYuLi/BalancedGroupSoftmax. Yu Li 0016, Tao Wang 0053, Bingyi Kang, Sheng Tang, Jintao Li 0001, Jiashi Feng |
CVPR | 4 |
| 2020 | The Devil Is in Classification: A Simple Framework for Long-Tail Instance Segmentation
Tao Wang 0053, Yu Li 0016, Bingyi Kang, Junnan Li 0001, Jun Hao Liew, Sheng Tang, Steven C. H. Hoi, Jiashi Feng |
ECCV (14) | 6 |
| 2020 | Visual Relation Grounding in Videos
Junbin Xiao, Xindi Shang, Xun Yang 0001, Sheng Tang, Tat-Seng Chua |
ECCV (6) | 4 |
| 2020 | Ahff-Net: Adaptive Hierarchical Feature Fusion Network For Image InpaintingabstractGeneration-based image inpainting methods can capture semantic features but fail to generate consistent details and high image quality results due to highly abstract feature learning and the instability of GAN training. Current methods try to overcome these disadvantages but they either need additional marginal maps or are not suitable for different shapes of occlusion. In this paper, we introduce an adaptive hierarchical feature fusion network (AHFF-Net). Without additional maps, our method can obtain consistent edges and high-quality results with different occlusions. Specifically, to guarantee the consistency of low-level features, our hierarchical fusion generator captures and aggregates multi-scale and multi-level context features. To get the high-quality results, the conditional self-supervised discriminator pay more attention to the unknown area by conditional GAN loss and stabilize the training process by conditional rotation loss. The proposed network achieves the state-of-the-art consistently on the Paris StreetView and Places365-Standard datasets with three shapes of masks. Sheng Tang, Yu Li 0016, Rui Zhang 0040 |
ICIP | 2 |
| 2020 | Local and nonlocal constraints for compressed sensing video and multi-view image recovery
Yun Song, Dengyong Zhang, Qiang Tang 0006, Sheng Tang, Kun Yang 0001 |
Neurocomputing | 4 |
| 2020 | Perspective-Adaptive Convolutions for Scene ParsingabstractMany existing scene parsing methods adopt Convolutional Neural Networks with receptive fields of fixed sizes and shapes, which frequently results in inconsistent predictions of large objects and invisibility of small objects. To tackle this issue, we propose perspective-adaptive convolutions to acquire receptive fields of flexible sizes and shapes during scene parsing. Through adding a new perspective regression layer, we can dynamically infer the position-adaptive perspective coefficient vectors utilized to reshape the convolutional patches. Consequently, the receptive fields can be adjusted automatically according to the various sizes and perspective deformations of the objects in scene images. Our proposed convolutions are differentiable to learn the convolutional parameters and perspective coefficients in an end-to-end way without any extra training supervision of object sizes. Furthermore, considering that the standard convolutions lack contextual information and spatial dependencies, we propose a context adaptive bias to capture both local and global contextual information through average pooling on the local feature patches and global feature maps, followed by flexible attentive summing to the convolutional results. The attentive weights are position-adaptive and context-aware, and can be learned through adding an additional context regression layer. Experiments on Cityscapes and ADE20K datasets well demonstrate the effectiveness of the proposed methods. Rui Zhang 0040, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Tree-Structured Kronecker Convolutional Network for Semantic SegmentationabstractMost existing semantic segmentation methods employ atrous convolution to enlarge the receptive field of filters, but neglect partial information. To tackle this issue, we firstly propose a novel Kronecker convolution which adopts Kronecker product to expand the standard convolutional kernel for taking into account the partial feature neglected by atrous convolutions. Therefore, it can capture partial information and enlarge the receptive field of filters simultaneously without introducing extra parameters. Secondly, we propose a Tree-structured Feature Aggregation (TFA) module which follows a recursive rule to expand and forms a hierarchical structure. Thus, it can naturally learn representations of multi-scale objects and encode hierarchical contextual information in complex scenes. Finally, we design a Tree-structured Kronecker Convolutional Network (TKCN) which employs Kronecker convolution and TFA module. Extensive experiments on three datasets, PAS-CAL VOC 2012, PASCAL-Context and Cityscapes, verify the effectiveness of our proposed approach. Sheng Tang, Rui Zhang 0040, Juan Cao 0001, Jintao Li 0001 |
ICME | 2 |
| 2019 | Boundary Perception Guidance: A Scribble-Supervised Semantic Segmentation ApproachabstractSemantic segmentation suffers from the fact that densely annotated masks are expensive to obtain. To tackle this problem, we aim at learning to segment by only leveraging scribbles that are much easier to collect for supervision. To fully explore the limited pixel-level annotations from scribbles, we present a novel Boundary Perception Guidance (BPG) approach, which consists of two basic components, i.e., prediction refinement and boundary regression. Specifically, the prediction refinement progressively makes a better segmentation by adopting an iterative upsampling and a semantic feature enhancement strategy. In the boundary regression, we employ class-agnostic edge maps for supervision to effectively guide the segmentation network in localizing the boundaries between different semantic regions, leading to producing finer-grained representation of feature maps for semantic segmentation. The experiment results on the PASCAL VOC 2012 demonstrate the proposed BPG achieves mIoU of 73.2% without fully connected Conditional Random Field (CRF) and 76.0% with CRF, setting up the new state-of-the-art in literature. Bin Wang 0065, Guo-Jun Qi, Sheng Tang, Tianzhu Zhang 0001, Yunchao Wei, Yongdong Zhang 0001 |
IJCAI | 3 |
| 2019 | Spatiotemporal Breast Mass Detection Network (MD-Net) in 4D DCE-MRI Images
Lixi Deng, Sheng Tang, Huazhu Fu, Bin Wang 0065, Yongdong Zhang 0001 |
MICCAI (4) | 2 |
| 2019 | Mixed-dish Recognition with Contextual Relation NetworksabstractMixed dish is a food category that contains different dishes mixed in one plate, and is popular in Eastern and Southeast Asia. Recognizing individual dishes in a mixed dish image is important for health related applications, e.g. calculating the nutrition values. However, most existing methods that focus on single dish classification are not applicable to mixed-dish recognition. The new challenge in recognizing mixed-dish images are the complex ingredient combination and severe overlap among different dishes. In order to tackle these problems, we propose a novel approach called contextual relation networks (CR-Nets) that encodes the implicit and explicit contextual relations among multiple dishes using region-level features and label-level co-occurrence, respectively. This is inspired by the intuition that people are likely to choose dishes with common eating habits, e.g., with multiple nutrition but without repeating ingredients. In addition, we collect a large-scale dataset of mixed-dish images that contain $9,254$ mixed-dish images from $6$ school canteens in Singapore. Extensive experiments on both our dataset and a smaller-scale public dataset validate that our CR-Nets can achieve top performance for localizing the dishes and recognizing their food categories. Lixi Deng, Jingjing Chen 0001, Qianru Sun, Xiangnan He 0001, Sheng Tang, Zhaoyan Ming, Yongdong Zhang 0001, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2019 | Detection and tracking based tubelet generation for video object detection
Bin Wang 0065, Sheng Tang, Jun Bin Xiao, Quan-Feng Yan, Yongdong Zhang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Asymmetric GAN for Unpaired Image-to-Image TranslationabstractUnpaired image-to-image translation problem aims to model the mapping from one domain to another with unpaired training data. Current works like the well-acknowledged Cycle GAN provide a general solution for any two domains through modeling injective mappings with a symmetric structure. While in situations where two domains are asymmetric in complexity, i.e., the amount of information between two domains is different, these approaches pose problems of poor generation quality, mapping ambiguity, and model sensitivity. To address these issues, we propose Asymmetric GAN (AsymGAN) to adapt the asymmetric domains by introducing an auxiliary variable (aux) to learn the extra information for transferring from the information-poor domain to the information-rich domain, which improves the performance of state-of-the-art approaches in the following ways. First, aux better balances the information between two domains which benefits the quality of generation. Second, the imbalance of information commonly leads to mapping ambiguity, where we are able to model one-to-many mappings by tuning aux, and furthermore, our aux is controllable. Third, the training of Cycle GAN can easily make the generator pair sensitive to small disturbances and variations while our model decouples the ill-conditioned relevance of generators by injecting aux during training. We verify the effectiveness of our proposed method both qualitatively and quantitatively on asymmetric situation, label-photo task, on Cityscapes and Helen datasets, and show many applications of asymmetric image translations. In conclusion, our AsymGAN provides a better solution for unpaired image-to-image translation in asymmetric domains. Yu Li 0016, Sheng Tang, Rui Zhang 0040, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2018 | Auto-Balanced Filter Pruning for Efficient Convolutional Neural NetworksabstractIn recent years considerable research efforts have been devoted to compression techniques of convolutional neural networks (CNNs). Many works so far have focused on CNN connection pruning methods which produce sparse parameter tensors in convolutional or fully-connected layers. It has been demonstrated in several studies that even simple methods can effectively eliminate connections of a CNN. However, since these methods make parameter tensors just sparser but no smaller, the compression may not transfer directly to acceleration without support from specially designed hardware. In this paper, we propose an iterative approach named Auto-balanced Filter Pruning, where we pre-train the network in an innovative auto-balanced way to transfer the representational capacity of its convolutional layers to a fraction of the filters, prune the redundant ones, then re-train it to restore the accuracy. In this way, a smaller version of the original network is learned and the floating-point operations (FLOPs) are reduced. By applying this method on several common CNNs, we show that a large portion of the filters can be discarded without obvious accuracy drop, leading to significant reduction of computational burdens. Concretely, we reduce the inference cost of LeNet-5 on MNIST, VGG-16 and ResNet-56 on CIFAR-10 by 95.1%, 79.7% and 60.9%, respectively. Xiaohan Ding, Guiguang Ding, Jungong Han, Sheng Tang |
AAAI | 4 |
| 2018 | Zero-Shot Learning With Attribute SelectionabstractZero-shot learning (ZSL) is regarded as an effective way to construct classification models for target classes which have no labeled samples available. The basic framework is to transfer knowledge from (different) auxiliary source classes having sufficient labeled samples with some attributes shared by target and source classes as bridge. Attributes play an important role in ZSL but they have not gained sufficient attention in recent years. Previous works mostly assume attributes are perfect and treat each attribute equally. However, as shown in this paper, different attributes have different properties, such as their class distribution, variance, and entropy, which may have considerable impact on ZSL accuracy if treated equally. Based on this observation, in this paper we propose to use a subset of attributes, instead of the whole set, for building ZSL models. The attribute selection is conducted by considering the information amount and predictability under a novel joint optimization framework. To our knowledge, this is the first work that notices the influence of attributes themselves and proposes to use a refined attribute set for ZSL. Since our approach focuses on selecting good attributes for ZSL, it can be combined to any attribute based ZSL approaches so as to augment their performance. Experiments on four ZSL benchmarks demonstrate that our approach can improve zero-shot classification accuracy and yield state-of-the-art results. Guiguang Ding, Jungong Han, Sheng Tang |
AAAI | 4 |
| 2018 | Learning and Thinking Strategy for Training Sequence Generation Models
Yu Li 0016, Sheng Tang, Junbo Guo, Jintao Li 0001, Shuicheng Yan |
BMVC | 2 |
| 2018 | High Resolution Feature Recovering for Accelerating Urban Scene ParsingabstractBoth accuracy and speed are equally important in urban scene parsing. Most of the existing methods mainly focus on improving parsing accuracy, ignoring the problem of low inference speed due to large-sized input and high resolution feature maps. To tackle this issue, we propose a High Resolution Feature Recovering (HRFR) framework to accelerate a given parsing network. A Super-Resolution Recovering module is employed to recover features of large original-sized images from features of down-sampled input. Therefore, our framework can combine the advantages of (1) fast speed of networks with down-sampled input and (2) high accuracy of networks with large original-sized input. Additionally, we employ auxiliary intermediate supervision and boundary region re-weighting to facilitate the optimization of the network. Extensive experiments on the two challenging Cityscapes and CamVid datasets well demonstrate the effectiveness of the proposed HRFR framework, which can accelerate the scene parsing inference process by about 3.0x speedup from 1/2 down-sampled input with negligible accuracy reduction. Rui Zhang 0040, Sheng Tang, Luoqi Liu, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
IJCAI | 2 |
| 2018 | Automated Pulmonary Nodule Detection: High Sensitivity with Few Candidates
Bin Wang 0065, Guo-Jun Qi, Sheng Tang, Liheng Zhang, Lixi Deng, Yongdong Zhang 0001 |
MICCAI (2) | 3 |
| 2018 | Style Separation and Synthesis via Generative Adversarial NetworksabstractStyle synthesis attracts great interests recently, while few works focus on its dual problem "style separation". In this paper, we propose the Style Separation and Synthesis Generative Adversarial Network (S3-GAN) to simultaneously implement style separation and style synthesis on object photographs of specific categories. Based on the assumption that the object photographs lie on a manifold, and the contents and styles are independent, we employ S3-GAN to build mappings between the manifold and a latent vector space for separating and synthesizing the contents and styles. The S3-GAN consists of an encoder network, a generator network, and an adversarial network. The encoder network performs style separation by mapping an object photograph to a latent vector. Two halves of the latent vector represent the content and style, respectively. The generator network performs style synthesis by taking a concatenated vector as input. The concatenated vector contains the style half vector of the style target image and the content half vector of the content target image. Once obtaining the images from the generator network, an adversarial network is imposed to generate more photo-realistic images. Experiments on CelebA and UT Zappos 50K datasets demonstrate that the S3-GAN has the capacity of style separation and synthesis simultaneously, and could capture various styles in a single model. Rui Zhang 0040, Sheng Tang, Yu Li 0016, Junbo Guo, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
ACM Multimedia | 2 |
| 2018 | Hierarchical BoW with segmental sparse coding for large scale image classification and retrieval
Jianshe Zhou, Narentuya, Sheng Tang, Jie Liu 0022 |
Multim. Tools Appl. | 3 |
| 2018 | Implicit Negative Sub-Categorization and Sink Diversion for Object DetectionabstractIn this paper, we focus on improving the proposal classification stage in the object detection task and present implicit negative sub-categorization and sink diversion to lift the performance by strengthening loss function in this stage. First, based on the observation that the "background" class is generally very diverse and thus challenging to be handled as a single indiscriminative class in existing state-of-the-art methods, we propose to divide the background category into multiple implicit sub-categories to explicitly differentiate diverse patterns within it. Second, since the ground truth class inevitably has low-value probability scores for certain images, we propose to add a "sink" class and divert the probabilities of wrong classes to this class when necessary, such that the ground truth label will still have a higher probability than other wrong classes even though it has low probability output. Additionally, we propose to use dilated convolution, which is widely used in the semantic segmentation task, for efficient and valuable context information extraction. Extensive experiments on PASCAL VOC 2007 and 2012 data sets show that our proposed methods based on faster R-CNN implementation can achieve state-of-the-art mAPs, i.e., 84.1%, 82.6%, respectively, and obtain 2.5% improvement on ILSVRC DET compared with that of ResNet. Yu Li 0016, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2018 | GLA: Global-Local Attention for Image DescriptionabstractIn recent years, the task of automatically generating image description has attracted a lot of attention in the field of artificial intelligence. Benefitting from the development of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), many approaches based on the CNN-RNN framework have been proposed to solve this task and achieved remarkable process. However, two problems remain to be tackled in which the most existing methods use only the image-level representation. One problem is object missing, in which some important objects may be missing when generating the image description and the other is misprediction, when one object may be recognized in a wrong category. In this paper, to address these two problems, we propose a new method called global-local attention (GLA) for generating image description. The proposed GLA model utilizes an attention mechanism to integrate object-level features with image-level feature. Through this manner, our model can selectively pay attention to objects and context information concurrently. Therefore, our proposed GLA method can generate more relevant image description sentences and achieve the state-of-the-art performance on the well-known Microsoft COCO caption dataset with several popular evaluation metrics-CIDEr, METEOR, ROUGE-L and BLEU-1, 2, 3, 4. Sheng Tang, Yongdong Zhang 0001, Lixi Deng, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Image Caption with Global-Local AttentionabstractImage caption is becoming important in the field of artificial intelligence. Most existing methods based on CNN-RNN framework suffer from the problems of object missing and misprediction due to the mere use of global representation at image-level. To address these problems, in this paper, we propose a global-local attention (GLA) method by integrating local representation at object-level with global representation at image-level through attention mechanism. Thus, our proposed method can pay more attention to how to predict the salient objects more precisely with high recall while keeping context information at image-level cocurrently. Therefore, our proposed GLA method can generate more relevant sentences, and achieve the state-of-the-art performance on the well-known Microsoft COCO caption dataset with several popular metrics. Sheng Tang, Lixi Deng, Yongdong Zhang 0001, Qi Tian 0001 |
AAAI | 2 |
| 2017 | Scale-Adaptive Convolutions for Scene ParsingabstractMany existing scene parsing methods adopt Convolutional Neural Networks with fixed-size receptive fields, which frequently result in inconsistent predictions of large objects and invisibility of small objects. To tackle this issue, we propose a scale-adaptive convolution to acquire flexiblesize receptive fields during scene parsing. Through adding a new scale regression layer, we can dynamically infer the position-adaptive scale coefficients which are adopted to resize the convolutional patches. Consequently, the receptive fields can be adjusted automatically according to the various sizes of the objects in scene images. Thus, the problems of invisible small objects and inconsistent large-object predictions can be alleviated. Furthermore, our proposed scale-adaptive convolutions are not only differentiable to learn the convolutional parameters and scale coefficients in an end-to-end way, but also of high parallelizability for the convenience of GPU implementation. Additionally, since the new scale regression layers are learned implicitly, any extra training supervision of object sizes is unnecessary. Extensive experiments on Cityscapes and ADE20K datasets well demonstrate the effectiveness of the proposed scaleadaptive convolutions. Rui Zhang 0040, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001, Shuicheng Yan |
ICCV | 2 |
| 2017 | Global-residual and Local-boundary Refinement Networks for Rectifying Scene Parsing PredictionsabstractMost of existing scene parsing methods suffer from the serious problems of both inconsistent parsing results and object boundary shift. To tackle these problems, we first propose an iterative Global-residual Refinement Network (GRN) through exploiting global contextual information to predict the parsing residuals and iteratively smoothen the inconsistent parsing labels. Furthermore, we propose a Local-boundary Refinement Network (LRN) to learn the position-adaptive propagation coefficients so that local contextual information from neighbors can be optimally captured for refining object boundaries. Finally, we cascade the proposed two refinement networks after a fully residual convolutional neural network within a uniform framework. Extensive experiments on ADE20K and Cityscapes datasets well demonstrate the effectiveness of the two refinement methods for refining scene parsing predictions. Rui Zhang 0040, Sheng Tang, Jintao Li 0001, Shuicheng Yan |
IJCAI | 2 |
| 2017 | Collaborative Dictionary Learning and Soft Assignment for Sparse Coding of Image Features
Jie Liu 0022, Sheng Tang, Yu Li 0016 |
MMM (1) | 2 |
| 2017 | HDIdx: High-dimensional indexing for efficient approximate nearest neighbor search
Ji Wan, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001, Steven C. H. Hoi |
Neurocomputing | 2 |
| 2017 | Multi-modal tag localization for mobile video search
Rui Zhang 0040, Sheng Tang, Wu Liu 0005, Yongdong Zhang 0001, Jintao Li 0001 |
Multim. Syst. | 2 |
| 2017 | Object Localization Based on Proposal FusionabstractTraditional regression framework of object locali-zation such as Overfeat often suffers from the problem of inaccurate scoring due to the separate scoring of classification network and regression network upon inconsistent regions. To tackle this problem, in this paper, we propose a novel object localization framework based on multiple complementary region proposal methods from the view of classification rather than regression. On top of our framework, we first combine multiple complementary region proposals during both training and testing as a means of data augmentation to generate more dense and reliable proposals for fusion, then achieve optimal compromise between complexity and efficiency through category clustering for bounding box sharing among similar categories, and finally propose a dense proposal fusion approach to merge dense region proposals near true object for fine-tuning of the final bounding box's coordinates and updating the confidence of fused proposals for final decision. Extensive experiments on the well-known large scale ILSVRC 2015 LOC dataset verify the effectiveness of our object localization framework. Sheng Tang, Yu Li 0016, Lixi Deng, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Adaptive weighted imbalance learning with application to abnormal activity recognition
Xingyu Gao 0001, Zhenyu Chen 0003, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001 |
Neurocomputing | 3 |
| 2015 | Large visual words for large scale image classificationabstractRecently, using large visual vocabulary or codebooks to quantize and partition the set of local feature descriptors into large set of disjoint subsets termed visual words (or large visual words) has become an important research topic in solving many computer vision problems including near duplicate image retrieval, object retrieval, etc. Generally, large visual words means a heavy burden on the cost of time and memory space for both the construction of large vocabulary and the searching process, especially for large scale applications. In this paper, we present an efficient generation approach of large visual words with a very compact vocabulary, namely two dictionaries learned with sparse non-negative matrix factorization (NMF). After piecewise sparse decomposition of features with two learned dictionaries, we map a pair of indices of the dictionary's bases corresponding to the maximum elements of the two sparse codes to a large set of visual words upon the assumption that data with similar properties will share the same base with the largest sparse coefficient. With the help of an inverted file structure built through the large visual words, K-nearest neighbors (KNN) can be efficiently retrieved. Therefore, we can classify images very efficiently with the incorporation of our fast KNN search based on large visual words into SVM-KNN method. Experiments on the public Oxford dataset, and ACM Multimedia 2013 Yahoo! image classification challenge dataset show that our approach is both effective and efficient. Sheng Tang, Ke Lu 0002, Yongdong Zhang 0001 |
ICIP | 1 |
| 2015 | A Sparse Ensemble Learning System For Efficient Semantic IndexingabstractThis demo presents an extremely efficient concept detection system based on a novel bag of words extraction method and sparse ensemble learning. We will show that the presented system can efficiently build the concept detectors upon millions of images, and achieve real-time concept detection on unseen images with the state-of-the-arts accuracy. To do so, we first develop an efficient bag of visual words (BoW) construction method based on sparse non-negative matrix factorization (NMF) and GPU enabled SIFT feature extraction. We then develop a sparse ensemble learning method to build the detection model, which drastically reduces learning time in order of magnitude over traditional methods like Support Vector Machine. The demo video of the system is available at YouTube: http://youtu.be/57obnlCxqAs Sheng Tang, Yu Li 0016, Jun Bin Xiao, Jintao Li 0001 |
ICMR | 1 |
| 2015 | Scalable logo recognition based on compact sparse dictionary for mobile devicesabstractIn this paper, we present a novel scalable logo recognition system which can recognize a large number of logo categories locally on mobile devices. The system is unsupervised without any supervised training procedure, and very time efficient at low memory cost. It is also robust against challenging conditions such as noise addition, different image scale, rotation, etc. To achieve this goal, we propose an efficient segmental quantization approach for generation of large visual words over one million size with a very compact vocabulary. The vocabulary consists of two small dictionaries learned through sparse non-negative matrix factorization (NMF) of local SIFT descriptors. With an inverted index structure built through the large visual words, query images containing logos can be recognized through efficient retrieval of K-nearest neighbors (K-NN) of logo instances in the dataset. Our vocabulary size is very small, only one thousandth of that of traditional Approximate K-Means (AKM) method, which is of great importance for mobile devices with limited memory. Furthermore, based on the compact dictionary, we present a promising verification way of filtering false positives via sparse reconstruction of SIFT descriptors with a very few number of sparse codes due to the sparsity's property of lowest reconstruction error. Experiments on our dataset with 400 logo classes show that our system is very efficient and effective. Sheng Tang, Yongdong Zhang 0001 |
MMSP | 1 |
| 2015 | Pedestrian detection based on Region Proposal FusionabstractAlmost all existing state-of-the-art pedestrian detection methods use combination of hand-crafted features, which cannot well handle the particular challenges in real-world situation. In this paper, we take advantage of Regions with Convolution Neural Networks features (R-CNN) to extract more robust pedestrian features for effective pedestrian detection in complicated environments. To further improve the performance: 1) we propose a Region Proposal Fusion algorithm to get effective region proposals since after careful observation, we found that the quality of region proposals is crucially important for detection performance. 2) we exploit a pedestrian detection expansion method based on image retrieval with color moment features due to R-CNN's requirements of large number of training samples to avoid overfitting. Consequently, the final average miss rate is greatly reduced to 23% in the INRIA pedestrian detection dataset, which is much (23%) lower than that of original HOG (46%). Sheng Tang, Ruizhen Zhao, Yi-Gang Cen |
MMSP | 2 |
| 2015 | An efficient concept detection system via sparse ensemble learning
Sheng Tang, Yongdong Zhang 0001, Zuoxin Xu, Yantao Zheng, Jintao Li 0001 |
Neurocomputing | 1 |
| 2014 | A Representative Local Region Detector Based On Color-Contrast-MSERabstractIn order to extract representative local invariant regions in textured natural images, we propose a Color-Contrast-MSER (CCM) detector with color-contrast pixel ranking, which can reduce the number of meaningless regions extracted from backgrounds. The main contributions are threefold: (1) In contrast with the original MSER[3] which adopts intensity pixel ranking, we develop a new pixel ranking mechanism based on color contrast analysis. (2) In this paper, the pixel ranking value of each pixel is defined as the color contrast between a kernel-sized window and the background. Therefore we propose an adaptive background scale selection mechanism that simulates the background color distribution as the benchmark for color contrast. (3) The experimental results demonstrate that compared with the original MSER detector[3], our Color-Contrast-MSER (CCM) detector can extract more representative local regions with competitive repeatability score at only 50% computational time and 10% memory cost. Ke Gao 0012, Sheng Tang, Yongdong Zhang 0001 |
ICMR | 3 |
| 2014 | FSpH: Fitted spectral hashing for efficient similarity search
Yongdong Zhang 0001, Yu Wang 0009, Sheng Tang, Steven C. H. Hoi, Jintao Li 0001 |
Comput. Vis. Image Underst. | 3 |
| 2014 | Fusing audio vocabulary with visual features for pornographic video detection
Hongtao Xie 0001, Sheng Tang |
Future Gener. Comput. Syst. | 4 |
| 2014 | Representative selection based on sparse modeling
Yu Wang 0089, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001 |
Neurocomputing | 2 |
| 2014 | Semi-supervised learning via sparse model
Yu Wang 0089, Sheng Tang, Yantao Zheng, Yongdong Zhang 0001, Jintao Li 0001 |
Neurocomputing | 2 |
| 2014 | Pedestrian detection based on sparse coding and transfer learning
Feidie Liang, Sheng Tang, Yongdong Zhang 0001, Zuoxin Xu, Jintao Li 0001 |
Mach. Vis. Appl. | 2 |
| 2013 | Data driven multi-index hashingabstractBinary representation for large scale nearest neighbor search received more and more concern recently. Although binary codes can be directly used as indices of the hash tables, correlations between the bits may lead to non-uniform codes distribution and reduce the performance of the hash table. In this paper, we propose a data driven multi-index hashing method for exact nearest neighbor search in Hamming space. By exploring the statistics properties of the dataset, we can separate the correlated bits into different segments during the process of building multiple hash tables, and thus make binary codes distributed as uniformly as possible in each hash table. Experiments conducted on a huge amount of binary codes extracted from the UK Bench dataset show that our method can achieve significant acceleration in searching speed for large scale dataset. Ji Wan, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001 |
ICIP | 2 |
| 2013 | Fitted spectral hashingabstractSpectral hashing (SpH) is an efficient and simple binary hashing method, which assumes that data are sampled from a multidimensional uniform distribution. However, this assumption is too restrictive in practice. In this paper we propose an improved method, Fitted Spectral Hashing, to relax this distribution assumption. Our work is based on the fact that one-dimensional data of any distribution could be mapped to a uniform distribution without changing the local neighbor relations among data items. We have found that this mapping on each PCA direction has certain regular pattern, and could fit data well by S-Curve function, Sigmoid function. With more parameters Fourier function also fit data well. Thus with Sigmoid function and Fourier function, we propose two binary hashing methods. Experiments show that our methods are efficient and outperform state-of-the-art methods. Yu Wang 0089, Sheng Tang, Jintao Li 0001, DanYi Chen |
ACM Multimedia | 2 |
| 2013 | A Sparse Coding Based Transfer Learning Framework for Pedestrian Detection
Feidie Liang, Sheng Tang, Yu Wang 0089, Jintao Li 0001 |
MMM (2) | 2 |
| 2013 | Beyond Kmedoids: Sparse Model Based Medoids Algorithm for Representative Selection
Yu Wang 0089, Sheng Tang, Feidie Liang, Jintao Li 0001 |
MMM (2) | 2 |
| 2013 | Robust human body segmentation based on part appearance and spatial constraint
Sheng Tang, Yongdong Zhang 0001, Shiguo Lian, Shouxun Lin |
Neurocomputing | 2 |
| 2013 | Robust common visual pattern discovery using graph matching
Hongtao Xie 0001, Yongdong Zhang 0001, Ke Gao 0012, Sheng Tang, Kefu Xu, Li Guo 0001, Jintao Li 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2013 | Accurate Estimation of Human Body Orientation From RGB-D SensorsabstractAccurate estimation of human body orientation can significantly enhance the analysis of human behavior, which is a fundamental task in the field of computer vision. However, existing orientation estimation methods cannot handle the various body poses and appearances. In this paper, we propose an innovative RGB-D-based orientation estimation method to address these challenges. By utilizing the RGB-D information, which can be real time acquired by RGB-D sensors, our method is robust to cluttered environment, illumination change and partial occlusions. Specifically, efficient static and motion cue extraction methods are proposed based on the RGB-D superpixels to reduce the noise of depth data. Since it is hard to discriminate all the 360 (°) orientation using static cues or motion cues independently, we propose to utilize a dynamic Bayesian network system (DBNS) to effectively employ the complementary nature of both static and motion cues. In order to verify our proposed method, we build a RGB-D-based human body orientation dataset that covers a wide diversity of poses and appearances. Our intensive experimental evaluations on this dataset demonstrate the effectiveness and efficiency of the proposed method. Wu Liu 0005, Yongdong Zhang 0001, Sheng Tang, Jinhui Tang 0001, Richang Hong, Jintao Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2012 | Exploring multi-modality structure for cross domain adaptation in video concept annotation
Shao-Xi Xu, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001, Yantao Zheng |
Neurocomputing | 2 |
| 2012 | Exploring probabilistic localized video representation for human action recognition
Yan Song 0004, Sheng Tang, Yantao Zheng, Tat-Seng Chua, Yongdong Zhang 0001, Shouxun Lin |
Multim. Tools Appl. | 2 |
| 2012 | Sparse Ensemble Learning for Concept DetectionabstractThis work presents a novel sparse ensemble learning scheme for concept detection in videos. The proposed ensemble first exploits a sparse non-negative matrix factorization (NMF) process to represent data instances in parts and partition the data space into localities, and then coordinates the individual classifiers in each locality for final classification. In the sparse NMF, data exemplars are projected to a set of locality bases, in which the non-negative superposition of basis images reconstructs the original exemplars. This additive combination ensures that each locality captures the characteristics of data exemplars in part, thus enabling the local classifiers to hold reasonable diversity in their own regions of expertise. More importantly, the sparse NMF ensures that an exemplar is projected to only a few bases (localities) with non-zero coefficients. The resultant ensemble model is, therefore, sparse, in the way that only a small number of efficient classifiers in the ensemble will fire on a testing sample. Extensive tests on the TRECVid 08 and 09 datasets show that the proposed ensemble learning achieves promising results and outperforms existing approaches. The proposed scheme is feature-independent, and can be applied in many other large scale pattern recognition problems besides visual concept detection. Sheng Tang, Yantao Zheng, Yu Wang 0089, Tat-Seng Chua |
IEEE Trans. Multim. | 1 |
| 2011 | Fusing Audio-Words with Visual Features for Pornographic Video DetectionabstractThe traditional approach of filtering pornographic videos on the Internet is based on visual features of keyframes. However, it cannot meet users' needs owing to the proliferation of low-resolution videos. To improve the filtering performance, we propose a novel framework of fusing audio-words with visual features for pornographic video detection. Our intention is not only to fuse the two modalities of visual images and audio signals, but also to narrow down the semantic gap between low-level features and high-level concepts by using the mid-level feature "audio-words". To further improve the performance, we present the segmentation algorithm based on units of energy envelope and the decision algorithm based on periodic patterns. The results show that our approach outperforms the traditional one which is based on visual features and achieves satisfactory performance. Moreover, the proposed segmentation algorithm is better than the conventional one using the same length and the proposed decision algorithm exceeds the conventional one using thresholds. Yongdong Zhang 0001, Sheng Tang |
TrustCom | 4 |
| 2011 | Localized Multiple Kernel Learning for Realistic Human Action Recognition in VideosabstractRealistic human action recognition in videos has been a useful yet challenging task. Video shots of same actions may present huge intra-class variations in terms of visual appearance, kinetic patterns, video shooting, and editing styles. Heterogeneous feature representations of videos pose another challenge on how to effectively handle the redundancy, complementariness and disagreement in these features. This paper proposes a localized multiple kernel learning (L-MKL) algorithm to tackle the issues above. L-MKL integrates the localized classifier ensemble learning and multiple kernel learning in a unified framework to leverage the strengths of both. The basis of L-MKL is to build multiple kernel classifiers on diverse features at subspace localities of heterogeneous representations. L-MKL integrates the discriminability of complementary features locally and enables localized MKL classifiers to deliver better performance in its own region of expertise. Specifically, L-MKL develops a locality gating model to partition the input space of heterogeneous representations to a set of localities of simpler data structure. Each locality then learns its localized optimal combination of Mercer kernels of heterogeneous features. Finally, the gating model coordinates the localized multiple kernel classifiers globally to perform action recognition. Experiments on two datasets show that the proposed approach delivers promising performance. Yan Song 0004, Yantao Zheng, Sheng Tang, Yongdong Zhang 0001, Shouxun Lin, Tat-Seng Chua |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Efficient Feature Detection and Effective Post-Verification for Large Scale Near-Duplicate Image SearchabstractState-of-the-art near-duplicate image search systems mostly build on the bag-of-local features (BOF) representation. While favorable for simplicity and scalability, these systems have three shortcomings: 1) high time complexity of the local feature detection; 2) discriminability reduction of local descriptors due to BOF quantization; and 3) neglect of the geometric relationships among local features after BOF representation. To overcome these shortcomings, we propose a novel framework by using graphics processing units (GPU). The main contributions of our method are: 1) a new fast local feature detector coined Harris-Hessian (H-H) is designed according to the characteristics of GPU to accelerate the local feature detection; 2) the spatial information around each local feature is incorporated to improve its discriminability, supplying semi-local spatial coherent verification (LSC); and 3) a new pairwise weak geometric consistency constraint (P-WGC) algorithm is proposed to refine the search result. Additionally, part of the system is implemented on GPU to improve efficiency. Experiments conducted on reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of H-H, LSC, and P-WGC. Hongtao Xie 0001, Ke Gao 0012, Yongdong Zhang 0001, Sheng Tang, Jintao Li 0001 |
IEEE Trans. Multim. | 4 |
| 2010 | A distribution based video representation for human action recognitionabstractMost current research on human action recognition in videos uses the bag-of-words (BoW) representations based on vector quantization on local spatial temporal features, due to the simplicity and good performance of such representations. In contrast to the BoW schemes, this paper explores a localized, continuous and probabilistic video representation. Specifically, the proposed representation encodes the visual and motion information of an ensemble of local spatial temporal (ST) features of a video into a distribution estimated by a generative probabilistic model such as the Gaussian Mixture Model. Furthermore, this probabilistic video representation naturally gives rise to an information-theoretic distance metric of videos. This makes the representation readily applicable as input to most discriminative classifiers, such as the nearest neighbor schemes and the kernel methods. The experiments on two datasets, KTH and UCF sports, show that the proposed approach could deliver promising results. Yan Song 0004, Sheng Tang, Yantao Zheng, Tat-Seng Chua, Yongdong Zhang 0001, Shouxun Lin |
ICME | 2 |
| 2009 | Logo detection based on spatial-spectral saliency and partial spatial contextabstractLogo detection is important for brand advertising and surveillance applications. The central issues of this technology are fast localization and accurate matching. Based on key traits analysis of common logos, this paper presents a two-stage detection scheme based on spatialspectral saliency (SSS) and partial spatial context (PSC). SSS speeds up logo location and avoid the impact of cluttered background. PSC filters false matching using spatial consistency of local invariant points. The integration of SSS and PSC result in faster localization and increased accuracy. Experiments on a dataset of nearly 10,000 web images containing several popular logo types are presented. The results indicate that our method is applicable and precise for different logo detection scenarios. Ke Gao 0012, Shouxun Lin, Yongdong Zhang 0001, Sheng Tang, Dongming Zhang 0004 |
ICME | 4 |
| 2009 | Visual words based spatiotemporal sequence matching in video copy detectionabstractThis paper proposes a novel content-based copy retrieval scheme for video copy identification. Its goal is to detect matches between a doubtful video and the ones stored in the database of the legal holders of the videos. Due to various transformations the copy may has, we use visual words vector as a representation of a frame which is based on SIFT descriptor. Unlike traditional bag-of-words (BoW) based approach applied in semantic retrieval, in which the temporal variation during the video is always neglected, our matching algorithm takes into account spatial and temporal distances between a query clip and the one in database. Experiments show robustness and effectiveness of our approach according to various single and compound transformations. Huamin Ren, Shouxun Lin, Sheng Tang |
ICME | 4 |
| 2009 | Pseudo relevance feedback with incremental learning for high level feature detectionabstractPseudo Relevance Feedback (PRF) has shown effective performance in information retrieval, but it has seldom been applied in the area of high level feature detection (HLF). In this paper, we explicitly propose to introduce PRF into HLF. Our contributions mainly lie in two-fold: (1) proposing three novel PRF approaches to extract pseudo positive samples, i.e., Nearest-Neighbor (NN) based PRF, Score-Evaluation (SE) based PRF and Multi-Classifier Decision (MCD) based PRF; (2) utilizing incremental learning to reduce the re-training time. We evaluate our approaches on the benchmark of TRECVID2008. Reported results have shown that MCD based approach outperforms the other two and obtain an excellent gain in average precision with respect to the baseline without PRF. Shao-Xi Xu, Sheng Tang, Jintao Li 0001, Yongdong Zhang 0001 |
ICME | 2 |
| 2009 | Pornprobe: an LDA-SVM based pornography detection systemabstractWe present PornProbe, a pornography detection system that detects pornographic contents in videos. To build such a detection system, we leverage a large scale training data set with 65,827 positive training image samples out of a total of 420,615 training samples, and a novel detection scheme based on hierarchical LDA-SVM. The system combines the unsupervised clustering in Latent Dirichlet Allocation (LDA) and supervised learning in Support Vector Machine, so as to achieve both high precision and recall while ensuring efficiency in both training and testing. This demonstration shows how the system detects the pornographic scenes in restricted artistic (RA) movies. Sheng Tang, Jintao Li 0001, Yongdong Zhang 0001, Xiufeng Hua, Yantao Zheng, Jinhui Tang 0001, Tat-Seng Chua |
ACM Multimedia | 1 |
| 2009 | A density-based method for adaptive LDA model selection
Juan Cao 0001, Tian Xia 0002, Jintao Li 0001, Yongdong Zhang 0001, Sheng Tang |
Neurocomputing | 5 |
| 2008 | Personalized event-based news video retrieval with dynamic user-logabstractPersonalization especially in the domain of information retrieval is essentially important, as users might pose the same query even when they are searching for different information. It is thus necessary to create a retrieval engine which takes into consideration the dynamic information needs of different users. This paper presents our personalized news video retrieval engine, which exploits the individual userpsilas previous browsing history to customize and enhance their future search results. Specifically, the system utilizes the news topic hierarchy, a hierarchical news topic structure derived from unsupervised clustering on the news video corpus and event entities from news video and online news articles. We then dynamically project userpsilas browsing history onto this topic hierarchy to provide the basis for re-ranking relevant news videos. This system is tested on one month of TRECVID 2006 dataset consisting of 80 hours news video and found to return results in a more intuitive and personalized manner. Yantao Zheng, Shi-Yong Neo, Sheng Tang, Shouxun Lin |
ICME | 5 |
| 2008 | Document Clustering Based on Spectral Clustering and Non-negative Matrix Factorization
Sheng Tang, Jintao Li 0001, Yongdong Zhang 0001, Wei-ping Ye |
IEA/AIE | 2 |
| 2008 | Local Subspace-Based Denoising for Shot Boundary Detection
Xuefeng Pan, Yongdong Zhang 0001, Jintao Li 0001, Xiaoyuan Cao, Sheng Tang |
IEA/AIE | 5 |
| 2008 | A statistical framework for replay detection in soccer videoabstractA novel statistical framework for replay detection is presented in this paper. Unlike current methods, the proposed framework exploits both inherent characters and transition relations of replay and non-replay scenes based on annotation of the video, which realizes segments and classifies video stream into replay and non-replay shots simultaneously. After annotation, the detected replay segment is further verified and its boundaries are adjusted to get more accurate replay segment considering probability distribution of lengths of replay and non-replay shots. Experimental results on soccer video are promising, demonstrating the effectiveness of the proposed framework. Shouxun Lin, Yongdong Zhang 0001, Sheng Tang |
ISCAS | 4 |
| 2008 | A More Topologically Stable Locally Linear Embedding Algorithm Based on R*-Tree
Tian Xia 0002, Jintao Li 0001, Yongdong Zhang 0001, Sheng Tang |
PAKDD | 4 |
| 2008 | An Innovative Model of Tempo and Its Application in Action Scene Detection for Movie AnalysisabstractIn this paper, we present an innovative model of tempo and its application in action scene detection for movie analysis. For the first time, we clearly propose that tempo indicates the rhythm of both movie scenarios and human perception. By thoroughly analyzing both aspects, we classify the factors of tempo into two sorts. The first is based on the film grammar and we use the low level features of shot length and camera motion to describe filmmaking by directors. The second is based on the human perception and we originally propose the information measure for perception depending on the cognitive informatics, a newly emerging and significative subject. With the information in both visual and auditory modalities, the low level features of motion intensity, motion complexity, audio energy and audio pace are integrated for the formulation of information to describe the viewers' emotional changes to continuously developing storyline. With both aspects, tempo is defined and tempo flow plot is derived as the clue of storyline. On the basis of video structuralization and movie tempo analysis, we build a system for hierarchical browse and edit with action scene annotation. The large-scale experiments demonstrate the effectiveness and generality of tempo for action movie analysis.In this paper, we present an innovative model of tempo and its application in action scene detection for movie analysis. For the first time, we clearly propose that tempo indicates the rhythm of both movie scenarios and human perception. By thoroughly analyzing both aspects, we classify the factors of tempo into two sorts. The first is based on the film grammar and we use the low level features of Shot Length and Camera Motion to describe filmmaking by directors. The second is based on the human perception and we originally propose the information measure for perception depending on the cognitive informatics, a newly emerging and significative subject. With the information in both visual and auditory modalities, the low level features of Motion Intensity, Motion Complexity, Audio Energy and Audio Pace are integrated for the formulation of information to describe the viewers' emotional changes to continuously developing storyline. With both aspects, tempo is defined and tempo flow plot is derived as the clue of storyline. On the basis of video structuralization and movie tempo analysis, we build a system for hierarchical browse and edit with action scene annotation. The large-scale experiments demonstrate the effectiveness and generality of tempo for action movie analysis. Anan Liu, Jintao Li 0001, Yongdong Zhang 0001, Sheng Tang, Yan Song 0004, Zhaoxuan Yang |
WACV | 4 |
| 2008 | A Hierarchical Scheme for Rapid Video Copy DetectionabstractToday with the rapid increasing popularity of web video sharing, digital copyright protection encounters many troubles. Video copy detection schemes are emerging to cope with the digital video piracy and illegal distribution problems. But the large amount of video data and diversity of copy attacks pose difficulties on copy detection. This paper presents a hierarchical scheme to detect video copies, especially the temporal attacked and re-encoded ones. Our algorithm which is based on the ordinal signature of intra frames and effective R*-tree indexing structure archives real time performance. Comparison experiments are conducted on the benchmarked database of CIVR 2007 copy detection showcase and demonstrate the promising results of the proposed approach. Xiao Wu 0004, Yongdong Zhang 0001, Sheng Tang, Tian Xia 0002, Jintao Li 0001 |
WACV | 3 |
| 2008 | Personalized multimedia web summarizer for touristabstractIn this paper, we highlight the use of multimedia technology in generating intrinsic summaries of tourism related information. The system utilizes an automated process to gather, filter and classify information on various tourist spots on the Web. The end result present to the user is a personalized multimedia summary generated with respect to users queries filled with text, image, video and real-time news made retrievable for mobile devices. Preliminary experiments demonstrate the superiority of our presentation scheme to traditional methods. Xiao Wu 0004, Jintao Li 0001, Yongdong Zhang 0001, Sheng Tang, Shi-Yong Neo |
WWW | 4 |
| 2007 | Statistical Framework for Shot Segmentation and Classification in Sports Video
Shouxun Lin, Yongdong Zhang 0001, Sheng Tang |
ACCV (2) | 4 |
| 2007 | Retrieval Method for Video Content in Different Format Based on Spatiotemporal Features
Xuefeng Pan, Jintao Li 0001, Yongdong Zhang 0001, Sheng Tang, Juan Cao 0001 |
ECIR | 4 |
| 2007 | Interactive Spatio-Temporal Visual Map Model for Web Video RetrievalabstractThe massive amount of multimedia information especially video available on the Web requires a more precise and interactive retrieval. Current operational video retrieval systems do not make use of the implicit visual features but rely only on textual metadata supplied by the user during uploading. This greatly affects the retrieval performance as the metadata may not be comprehensive or consistent. In this paper, we describe the use of a spatio-temporal visual map (STVM) model to supplement Web video retrieval. This is done by employing the spatio-temporal visual similarity to rerank the text-retrieval results and find new results. Experimental results on a dynamic Web video corpus show significant improvement based on STVM model, with good usability scores based on human users. Huan-Bo Luan, Shouxun Lin, Sheng Tang, Shi-Yong Neo, Tat-Seng Chua |
ICME | 3 |
| 2007 | News Video Retrieval using Implicit Event SemanticsabstractCurrent state-of-the-art news video retrieval systems mainly focus on automated speech recognition (ASR) text to perform retrieval. This paradigm greatly affects retrieval performance as ASR text alone is not sufficient to provide an accurate representation of the entire news video. In this paper, we describe our automated retrieval framework which fuses the multimodal features and event structures present in news video to support precise news video retrieval. The contributions of this paper are: (a) we uncover and employ temporal event clusters to provide additional information during story level retrieval; and (b) we integrate other modality features with text features and incorporate event clusters for pseudo relevance feedback (PRF) in shot level re-ranking. Experiments performed on video search task using the TRECVID 2005/06 dataset show that the proposed approach is effective. Shi-Yong Neo, Yantao Zheng, Hai-Kiat Goh, Tat-Seng Chua, Sheng Tang |
ICME | 5 |
| 2007 | Visual Features Extraction Through Spatiotemporal Slice Analysis
Xuefeng Pan, Jintao Li 0001, Shan Ba, Yongdong Zhang 0001, Sheng Tang |
MMM (2) | 5 |
| 2007 | HTRDP evaluations on Chinese information processing and intelligent human-machine interface
Qun Liu 0001, Hong Liu 0007, Le Sun 0001, Sheng Tang, Deyi Xiong, Hongxu Hou, Yuanhua Lv, Shouxun Lin, Yueliang Qian |
Frontiers Comput. Sci. China | 5 |
| 2007 | Secure and Incidental Distortion Tolerant Digital Signature for Image Authentication
Yongdong Zhang 0001, Sheng Tang, Jintao Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2005 | Compact and Robust Image Hashing
Sheng Tang, Jintao Li 0001, Yongdong Zhang 0001 |
ICCSA (2) | 1 |
| 2005 | SSF fingerprint for image authentication: an incidental distortion resistant schemeabstractWe propose a novel method for image authentication which can distinguish incidental manipulations from malicious ones. The authentication fingerprint is based on the Hotelling's T-square statistic (HTS) via Principal Component Analysis (PCA) of block DCT coefficients. HTS values of all blocks construct an unique and stable "block-edge image", i.e., Structural and Statistical Fingerprint (SSF). The characteristic of the SSF is that it is short, and can tolerate content-preserving modifications while keeping sensitive to content-changing modifications, and can locate tampered blocks easily. Furthermore, we use Fisher criterion to obtain optimal threshold for distinguishing manipulations. The security of the SSF is also achieved by encryption of the DCT coefficients with chaotic sequences. Experiments show that the proposed method is effective for authentication. Sheng Tang, Jintao Li 0001, Yongdong Zhang 0001 |
ACM Multimedia | 1 |
| 2000 | Data fusion of multisensor dataabstractThe problem of data fusion of multisensor data for multitarget tracking is considered. A hierarchical fusion system is presented for fusion of numerical data from multiple local radar stations, and a fuzzy clustering technique is introduced. Simulation results are presented for a scenario having three local radar stations and three targets with trajectories containing some type of acceleration and varying noise levels on the measurements. Sheng Tang |
KES | 1 |