EDBT 2026 Demo / reviewers in the wild / expert
Li Su 0003
dblp:05/365-3
· DBLP profile ↗
60ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0003-4038-753XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 21 · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning ModelsabstractLarge Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLMs) unlearning approaches that typically focus on modifying only final answers are insufficient for LRMs, as they fail to remove sensitive content from intermediate steps, leading to persistent privacy leakage and degraded security. To address these challenges, we propose Sensitive Trajectory Regulation (STaR), a parameter-free, inference-time unlearning framework that achieves robust privacy protection throughout the reasoning process. Specifically, we first identify sensitive content via semantic-aware detection. Then, we inject global safety constraints through secure prompt encoder. Next, we perform trajectory-aware suppression to dynamically block sensitive content across the entire reasoning chain. Finally, we apply token-level adaptive filtering to prevent both exact and paraphrased sensitive tokens during generation. Furthermore, to overcome the inadequacies of existing evaluation protocols, we introduce two metrics: Multi-Decoding Consistency Assessment (MCS), which measures the consistency of unlearning across diverse decoding strategies, and Multi-Granularity Membership Inference Attack (MIA) Evaluation, which quantifies privacy protection at both answer and reasoning-chain levels. Experiments on the R-TOFU benchmark demonstrate that STaR achieves comprehensive and stable unlearning with minimal utility loss, setting a new standard for privacy-preserving reasoning in LRMs. Gaoxiang Cong 0001, Li Su 0003, Liang Li 0003 |
AAAI | 3 |
| 2026 | Cross-City Correlation Learning for Traffic ForecastingabstractTraffic forecasting is essential in city-level applications, where data-driven deep learning has become the most popular method. However, sufficient data in developing cities is not always accessible, posing a challenge for training effective models in scenarios with limited data. Recently, several works have promoted this issue through cross-city knowledge transfer and shown promising performances. However, existing methods can neither distinguish node divergence nor extract functional similarities between cities, which results in suboptimal performance. To overcome the limitations, we propose a Cross-city Correlation Learning (CCL) framework. Firstly, we construct a self-supervised learning model to infer accurate node-to-node and node-to-region cross-city correlations from multiple noisy labels without using any auxiliary information. Then, we achieve spatial knowledge transfer from a transfer-adaptive graph convolution network based on the learned correlations in two aspects: the learnable adjacency matrix and region-specific kernel parameters, which ensure the target models can transfer more and better utilize the knowledge from the source domain. The experiments are conducted on six real-world datasets and fully prove the effectiveness of the proposed framework. Zhe Wu 0006, Li Su 0003, Xinfeng Zhang 0001, Yaowei Wang 0001, Qingming Huang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-CaptioningabstractVideo has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval, and leverages it as a feature extractor for other three challenging tasks solved in a multi-task learning paradigm. Nevertheless, this work struggles to learn the comprehensive cognition of user-preferred content, due to disregarding the hierarchies and association relations across modalities. In this paper, guided by the shallow-to-deep principle, we propose a query-centric audio-visual cognition (QUAG) network to construct a reliable multi-modal representation for moment retrieval, segmentation and step-captioning. Specifically, we first design the modality-synergistic perception to obtain rich audio-visual content, by modeling global contrastive alignment and local fine-grained interaction between visual and audio modalities. Then, we devise the query-centric cognition that uses the deep-level query to perform the temporal-channel filtration on the shallow-level audio-visual representation. This can cognize user-preferred content and thus attain a query-centric audio-visual representation for three tasks. Extensive experiments show QUAG achieves the SOTA results on HIREST. Further, we test QUAG on the query-based video summarization task and verify its good generalization. Yunbin Tu, Liang Li 0003, Li Su 0003, Qingming Huang |
AAAI | 3 |
| 2025 | Extracting Global Temporal Patterns Within Short Look-Back Windows for Traffic ForecastingabstractWith the continuous expansion of urban areas, accurate and effective traffic forecasting has become essential for intelligent urban traffic management. As traffic data inherently exhibits temporal dynamics, modeling its temporal patterns is critical to improve prediction performance. However, constrained by computational complexity, existing methods rely primarily on short-term historical data, which is typically noisy and limits the ability to capture global temporal patterns. To address this issue, we propose a novel Dual-Stream Transformer model (DSformer) that effectively captures global temporal patterns through a time-index model. To mitigate the impact of noise in short look-back windows, DSformer explicitly learns a temporal matrix that encodes structured temporal dependencies. Furthermore, we design a time-index loss that encourages similar representations for adjacent time indices, thereby reducing error propagation across time steps. In parallel, a historical-value stream is employed to model local information. Finally, a self-adaptive learning module is constructed to flexibly and accurately fuse global and local information. Extensive experiments on real-world traffic forecasting tasks across ten diverse scenarios demonstrate that our method consistently outperforms state-of-the-art baselines while maintaining competitive efficiency. The code is available at https://github.com/sky836/DSFormer.git. Zhe Wu 0006, Li Su 0003 |
CIKM | 4 |
| 2025 | Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supervised paradigm. To address these limitations, we propose a novel paradigm: Single-Frame supervised VAD (SF-VAD), which uses a single annotated abnormal frame per abnormal video. SF-VAD ensures annotation efficiency while offering precise anomaly reference, facilitating robust anomaly modeling, and enhancing the detection of subtle anomalies in complex visual contexts. To validate its effectiveness, we construct three SF-VAD benchmarks by manually re-annotating the ShanghaiTech, UCF-Crime, and XD-Violence datasets in a practical procedure. Further, we devise Frame-guided Progressive Learning (FPL), to generalize sparse frame supervision to event-level anomaly understanding. FPL first leverages evidential learning to estimate anomaly relevance guided by annotated frames. Then it extends anomaly supervision by mining discrete abnormal events based on anomaly relevance and feature similarity. Meanwhile, FPL decouples normal patterns by isolating distinct normal frames outside abnormal events, reducing false alarms. Extensive experiments show SF-VAD achieves state-of-the-art detection results while offering a favorable trade-off between performance and annotation cost. Junxi Chen, Liang Li 0003, Yunbin Tu, Li Su 0003, Zhe Xue, Qingming Huang |
NeurIPS | 4 |
| 2024 | Context-aware Difference Distilling for Multi-change CaptioningabstractMulti-change captioning aims to describe complex and coupled changes within an image pair in natural language.Compared with singlechange captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes.In this paper, we propose a novel context-aware difference distilling (CARD) network to capture all genuine changes for yielding sentences.Given an image pair, CARD first decouples context features that aggregate all similar/dissimilar semantics, termed common/difference context features.Then, the consistency and independence constraints are designed to guarantee the alignment/discrepancy of common/difference context features.Further, the common context features guide the model to mine locally unchanged features, which are subtracted from the pair to distill locally difference features.Next, the difference context features augment the locally difference features to ensure that all changes are distilled.In this way, we obtain an omni-representation of all changes, which is translated into linguistic sentences by a transformer decoder.Extensive experiments on three public datasets show CARD performs favourably against state-of-the-art methods. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang |
ACL (1) | 3 |
| 2024 | Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionabstractWeakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels, Multi-Instance Learning (MIL) is prevailing in wVAD. However, MIL suffers from insufficiency of binary supervision to model diverse abnormal patterns. Besides, the coupling between abnormality and its context hinders the learning of clear abnormal event boundary. In this paper, we propose prompt-enhanced MIL to detect various abnormal events while ensuring clear event boundaries. Concretely, we design the abnormal-aware prompts by using abnormal class annotations together with learnable prompt, which can incorporate semantic priors into video features dynamically. The detector can utilize the semantic-rich features to capture diverse abnormal patterns. In addition, normal context prompt is introduced to amplify the distinction between abnormality and its context, facilitating the generation of clear boundary. With the mutual enhancement of abnormal-aware and normal context prompt, the model can construct discriminative representations to detect divergent anomalies without ambiguous event boundaries. Extensive experiments demonstrate our method achieves SOTA performance on three public benchmarks. The code is available at https://github.com/Junxi-Chen/PE-MIL. Junxi Chen, Liang Li 0003, Li Su 0003, Zhengjun Zha, Qingming Huang |
CVPR | 3 |
| 2024 | Distractors-Immune Representation Learning with Cross-Modal Contrastive Regularization for Change Captioning
Yunbin Tu, Liang Li 0003, Li Su 0003, Chenggang Yan 0001, Qingming Huang |
ECCV (43) | 3 |
| 2024 | Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious FinetuningabstractDiffusion models (DMs) have demonstrated remarkable proficiency in producing images based on textual prompts. Numerous methods have been proposed to ensure these models generate safe images. Early methods attempt to incorporate safety filters into models to mitigate the risk of generating harmful images but such external filters do not inherently detoxify the model and can be easily bypassed. Hence, model unlearning and data cleaning are the most essential methods for maintaining the safety of models, given their impact on model parameters.
However, malicious fine-tuning can still make models prone to generating harmful or undesirable images even with these methods.
Inspired by the phenomenon of catastrophic forgetting, we propose a training policy using contrastive learning to increase the latent space distance between clean and harmful data distribution, thereby protecting models from being fine-tuned to generate harmful images due to forgetting.
The experimental results demonstrate that our methods not only maintain clean image generation capabilities before malicious fine-tuning but also effectively prevent DMs from producing harmful images after malicious fine-tuning. Our method can also be combined with other safety methods to maintain their safety against malicious fine-tuning further. Jiadong Pan, Hongcheng Gao, Zongyu Wu 0001, Taihang Hu, Li Su 0003, Qingming Huang, Liang Li 0003 |
NeurIPS | 5 |
| 2024 | SMART: Syntax-Calibrated Multi-Aspect Relation Transformer for Change CaptioningabstractChange captioning aims to describe the semantic change between two similar images. In this process, as the most typical distractor, viewpoint change leads to the pseudo changes about appearance and position of objects, thereby overwhelming the real change. Besides, since the visual signal of change appears in a local region with weak feature, it is difficult for the model to directly translate the learned change features into the sentence. In this paper, we propose a syntax-calibrated multi-aspect relation transformer to learn effective change features under different scenes, and build reliable cross-modal alignment between the change features and linguistic words during caption generation. Specifically, a multi-aspect relation learning network is designed to 1) explore the fine-grained changes under irrelevant distractors (e.g., viewpoint change) by embedding the relations of semantics and relative position into the features of each image; 2) learn two view-invariant image representations by strengthening their global contrastive alignment relation, so as to help capture a stable difference representation; 3) provide the model with the prior knowledge about whether and where the semantic change happened by measuring the relation between the representations of captured difference and the image pair. Through the above manner, the model can learn effective change features for caption generation. Further, we introduce the syntax knowledge of Part-of-Speech (POS) and devise a POS-based visual switch to calibrate the transformer decoder. The POS-based visual switch dynamically utilizes visual information during different word generation based on the POS of words. This enables the decoder to build reliable cross-modal alignment, so as to generate a high-level linguistic sentence about change. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the three public datasets. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | SpikeODE: Image Reconstruction for Spike Camera With Neural Ordinary Differential EquationabstractThe recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. However, reconstructing high-quality images from the binary spike data remains a challenge due to the existence of noises in the camera. This paper proposes SpikeODE, a novel approach to reconstructing clear images by exploring temporal-spatial correlation to depress noises. The main idea of our method is to restore the continuous dynamic process of real scenes in a latent space and learn the temporal correlations in a fine-grained manner. Furthermore, to model the dynamic process more effectively, we design a conditional ODE where the latent state of each timestamp is conditioned on the observed spike data. Subsequently, forward and backward inferences are conducted through the ODE to investigate the correlations between the representation of the target timestamp and the information from both past and future contexts. Additionally, we incorporate a Unet structure with a pixel-wise attention mechanism at each level to learn spatial correlations. Experimental results demonstrate that our method outperforms state-of-the-art methods across several metrics. Chen Yang 0034, Guorong Li, Shuhui Wang, Li Su 0003, Laiyun Qing, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Self Supervised Progressive Network for High Performance Video Object SegmentationabstractRecently, self-supervised video object segmentation (VOS) has attracted much interest. However, most proxy tasks are proposed to train only a single backbone, which relies on a point-to-point correspondence strategy to propagate masks through a video sequence. Due to its simple pipeline, the performance of the single backbone paradigm is still unsatisfactory. Instead of following the previous literature, we propose our self-supervised progressive network (SSPNet) which consists of a memory retrieval module (MRM) and collaborative refinement module (CRM). The MRM can perform point-to-point correspondence and produce a propagated coarse mask for a query frame through self-supervised pixel-level and frame-level similarity learning. The CRM, which is trained via cycle consistency region tracking, aggregates the reference & query information and learns the collaborative relationship among them implicitly to refine the coarse mask. Furthermore, to learn semantic knowledge from unlabeled data, we also design two novel mask-generation strategies to provide the training data with meaningful semantic information for the CRM. Extensive experiments conducted on DAVIS-17, YouTube- VOS and SegTrack v2 demonstrate that our method surpasses the state-of-the-art self-supervised methods and narrows the gap with the fully supervised methods. Guorong Li, Dexiang Hong, Kai Xu 0013, Bineng Zhong 0001, Li Su 0003, Zhenjun Han, Qingming Huang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Self-supervised Cross-view Representation Reconstruction for Change CaptioningabstractChange captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruction (SCORER) network. Concretely, we first design a multi-head token-wise matching to model relationships between cross-view features from similar/dissimilar images. Then, by maximizing cross-view contrastive alignment of two similar images, SCORER learns two view-invariant image representations in a self-supervised way. Based on these, we reconstruct the representations of unchanged objects by cross-attention, thus learning a stable difference representation for caption generation. Further, we devise a cross-modal backward reasoning to improve the quality of caption. This module reversely models a "hallucination" representation with the caption and "before" representation. By pushing it closer to the "after" representation, we enforce the caption to be informative about the difference in a self-supervised manner. Extensive experiments show our method achieves the state-of-the-art results on four datasets. The code is available at https://github.com/tuyunbin/SCORER. Yunbin Tu, Liang Li 0003, Li Su 0003, Zhengjun Zha, Chenggang Yan 0001, Qingming Huang |
ICCV | 3 |
| 2023 | Self-Regulated Learning for Egocentric Video Activity AnticipationabstractFuture activity anticipation is a challenging problem in egocentric vision. As a standard future activity anticipation paradigm, recursive sequence prediction suffers from the accumulation of errors. To address this problem, we propose a simple and effective Self-Regulated Learning framework, which aims to regulate the intermediate representation consecutively to produce representation that (a) emphasizes the novel information in the frame of the current time-stamp in contrast to previously observed content, and (b) reflects its correlation with previously observed frames. The former is achieved by minimizing a contrastive loss, and the latter can be achieved by a dynamic reweighing mechanism to attend to informative frames in the observed content with a similarity comparison between feature of the current frame and observed frames. The learned final video representation can be further enhanced by multi-task learning which performs joint feature learning on the target activity labels and the automatically detected action and object class tokens. SRL sharply outperforms existing state-of-the-art in most cases on two egocentric video datasets and two third-person video datasets. Its effectiveness is also verified by the experimental fact that the action and object concepts that support the activity semantics can be accurately identified. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Viewpoint-Adaptive Representation Disentanglement Network for Change CaptioningabstractChange captioning is to describe the fine-grained change between a pair of images. The pseudo changes caused by viewpoint changes are the most typical distractors in this task, because they lead to the feature perturbation and shift for the same objects and thus overwhelm the real change representation. In this paper, we propose a viewpoint-adaptive representation disentanglement network to distinguish real and pseudo changes, and explicitly capture the features of change to generate accurate captions. Concretely, a position-embedded representation learning is devised to facilitate the model in adapting to viewpoint changes via mining the intrinsic properties of two image representations and modeling their position information. To learn a reliable change representation for decoding into a natural language sentence, an unchanged representation disentanglement is designed to identify and disentangle the unchanged features between the two position-embedded representations. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the four public datasets. The code is available at https://github.com/tuyunbin/VARD. Yunbin Tu, Liang Li 0003, Li Su 0003, Junping Du 0001, Ke Lu 0002, Qingming Huang |
IEEE Trans. Image Process. | 3 |
| 2023 | Neighborhood Contrastive Transformer for Change CaptioningabstractChange captioning is to describe the semantic change between a pair of similar images in natural language. It is more challenging than general image captioning, because it requires capturing fine-grained change information while being immune to irrelevant viewpoint changes, and solving syntax ambiguity in change descriptions. In this paper, we propose a neighborhood contrastive transformer to improve the model's perceiving ability for various changes under different scenes and cognition ability for complex syntax structure. Concretely, we first design a neighboring feature aggregating to integrate neighboring context into each feature, which helps quickly locate the inconspicuous changes under the guidance of conspicuous referents. Then, we devise a common feature distilling to compare two images at neighborhood level and extract common properties from each image, so as to learn effective contrastive information between them. Finally, we introduce the explicit dependencies between words to calibrate the transformer decoder, which helps better understand complex syntax structure during training. Extensive experimental results demonstrate that the proposed method achieves the state-of-the-art performance on three public datasets with different change scenarios. The code is available athttps://github.com/tuyunbin/NCT. Yunbin Tu, Liang Li 0003, Li Su 0003, Ke Lu 0002, Qingming Huang |
IEEE Trans. Multim. | 3 |
| 2023 | Temporal Dynamic Concept Modeling Network for Explainable Video Event RecognitionabstractRecently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field. Weigang Zhang, Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Enhanced Semantic Head for Cascade Instance SegmentationabstractRecently, cascade instance segmentation inspired by cascade object detection has achieved notable performance. Due to the lack of global information, many methods suffer from incomplete segmentation such as missing edge regions and discontinuities within instances. To solve this problem, we proposed an effective and flexible semantic head to extract enhanced spatial context information. A vision transformer is utilized to generate global context features, and a convolution network is adopted to generate spatial context features. After combining the two modules, we obtain enhanced semantic segmentation features for segmentation. Extensive experiments show that the enhanced semantic head achieves 40.6% and 42.3% mask AP for cascade predictor HTC and DSC, which surpass about 0.9 and 1.4 percentage points respectively. The enhanced semantic head is universal and effective to improve the performance of different cascade predictors. Xuerong Huang, Li Su 0003, Guorong Li, Xinfeng Zhang 0001, Laiyun Qing, Qingming Huang |
ICME | 2 |
| 2022 | A Sparse-Motif Ensemble Graph Convolutional Network against Over-smoothingabstractThe over-smoothing issue is a well-known challenge for Graph Convolutional Networks (GCN). Specifically, it is often observed that increasing the depth of GCN ends up in a trivial embedding subspace where the difference among node embeddings belonging to the same cluster tends to vanish. This paper believes that the main cause lies in the limited diversity along the message passing pipeline. Inspired by this, we propose a Sparse-Motif Ensemble Graph Convolutional Network (SMEGCN). We argue that merely employing the original graph Laplacian as the spectrum of the graph cannot capture the diversified local structure of complex graphs. Hence, to improve the diversity of the graph spectrum, we introduce local topological structures of complex graphs into GCN by employing the so-called graph motifs or the small network subgraphs. Moreover, we find that the motif connections are much denser than the edge connections, which might converge to an all-one matrix within a few times of message-passing. To fix this, we first propose the notion of sparse motif to avoid spurious motif connections. Subsequently, we propose a hierarchical motif aggregation mechanism to integrate the graph spectral information from a series of different sparse-motif message passing paths. Finally, we conduct a series of theoretical and experimental analyses to demonstrate the superiority of the proposed method. Zhiyong Yang 0001, Peisong Wen, Li Su 0003, Qingming Huang |
IJCAI | 4 |
| 2022 | I2Transformer: Intra- and Inter-Relation Embedding Transformer for TV Show CaptioningabstractTV show captioning aims to generate a linguistic sentence based on the video and its associated subtitle. Compared to purely video-based captioning, the subtitle can provide the captioning model with useful semantic clues such as actors’ sentiments and intentions. However, the effective use of subtitle is also very challenging, because it is the pieces of scrappy information and has semantic gap with visual modality. To organize the scrappy information together and yield a powerful omni-representation for all the modalities, an efficient captioning model requires understanding video contents, subtitle semantics, and the relations in between. In this paper, we propose an Intra- and Inter-relation Embedding Transformer (I2Transformer), consisting of an Intra-relation Embedding Block (IAE) and an Inter-relation Embedding Block (IEE) under the framework of a Transformer. First, the IAE captures the intra-relation in each modality via constructing the learnable graphs. Then, IEE learns the cross attention gates, and selects useful information from each modality based on their inter-relations, so as to derive the omni-representation as the input to the Transformer. Experimental results on the public dataset show that the I2Transformer achieves the state-of-the-art performance. We also evaluate the effectiveness of the IAE and IEE on two other relevant tasks of video with text inputs,i.e., TV show retrieval and video-guided machine translation. The encouraging performance further validates that the IAE and IEE blocks have a good generalization ability. The code is available athttps://github.com/tuyunbin/I2Transformer. Yunbin Tu, Liang Li 0003, Li Su 0003, Shengxiang Gao, Chenggang Yan 0001, Zhengjun Zha, Zhengtao Yu 0001, Qingming Huang |
IEEE Trans. Image Process. | 3 |
| 2022 | Weakly Supervised Anomaly Detection in Videos Considering the Openness of EventsabstractAlthough various weakly supervised anomaly detection methods have been proposed in recent years, generalization of anomaly detection is still not well-explored. Existing weakly supervised methods usually use normal and abnormal events to pose anomaly detection as a regression problem. However, defining concepts that encompass all possible normal and abnormal event patterns is nearly unrealistic, so the anomaly detection model is likely to face both open normal and abnormal events in practical applications. We find some weakly supervised anomaly detection methods suffer from performance degradation when faced with open events due to their poor generalization. To tackle this issue, we propose a two-branch weakly supervised approach, which can improve the anomaly detection performance of open events without affecting the performance of the seen events. Specifically, considering that the pattern of open events is different from that of seen events, we design a Test Data Analyzer (TDA) that determines whether the test video features belong to seen or open data and argue for separate treatment for them. For the seen data, a classifier trained by multiple instance learning is used to predict anomaly scores. For the open data, we design an anomaly detection model via meta-learning named Meta-Learning Anomaly Detection (MLAD), which can directly determine whether open data is abnormal without updating model parameters. In detail, MLAD synthesizes pseudo-seen data and pseudo-open data so that the model can learn to detect anomalies in open data by transferring the knowledge of seen data. Experimental results validate the effectiveness of our proposed method. Chen Zhang 0013, Guorong Li, Qianqian Xu 0001, Xinfeng Zhang 0001, Li Su 0003, Qingming Huang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Rethinking Graph Neural Architecture Search From Message-PassingabstractGraph neural networks (GNNs) emerged recently as a standard toolkit for learning from data on graphs. Current GNN designing works depend on immense human expertise to explore different message-passing mechanisms, and require manual enumeration to determine the proper message-passing depth. Inspired by the strong searching capability of neural architecture search (NAS) in CNN, this paper proposes Graph Neural Architecture Search (GNAS) with novel-designed search space. The GNAS can automatically learn better architecture with the optimal depth of message passing on the graph. Specifically, we design Graph Neural Architecture Paradigm (GAP) with tree-topology computation procedure and two types of fine-grained atomic operations (feature filtering & neighbor aggregation) from message-passing mechanism to construct powerful graph network search space. Feature filtering performs adaptive feature selection, and neighbor aggregation captures structural information and calculates neighbors’ statistics. Experiments show that our GNAS can search for better GNNs with multiple message-passing mechanisms and optimal message-passing depth. The searched network achieves remarkable improvement over state-of-the-art manual designed and search-based GNNs on five large-scale datasets at three classical graph tasks. Codes can be found at https://github.com/phython96/GNAS-MP. Shaofei Cai, Liang Li 0003, Jincan Deng, Beichen Zhang 0006, Zhengjun Zha, Li Su 0003, Qingming Huang |
CVPR | 6 |
| 2021 | Two-Stage Polishing Network for Camouflaged Object Detection
Zhe Wu 0006, Li Su 0003, Qingming Huang |
ICIG (1) | 4 |
| 2021 | Decomposition and Completion Network for Salient Object DetectionabstractRecently, fully convolutional networks (FCNs) have made great progress in the task of salient object detection and existing state-of-the-arts methods mainly focus on how to integrate edge information in deep aggregation models. In this paper, we propose a novel Decomposition and Completion Network (DCN), which integrates edge and skeleton as complementary information and models the integrity of salient objects in two stages. In the decomposition network, we propose a cross multi-branch decoder, which iteratively takes advantage of cross-task aggregation and cross-layer aggregation to integrate multi-level multi-task features and predict saliency, edge, and skeleton maps simultaneously. In the completion network, edge and skeleton maps are further utilized to fill flaws and suppress noises in saliency maps via hierarchical structure-aware feature learning and multi-scale feature completion. Through jointly learning with edge and skeleton information for localizing boundaries and interiors of salient objects respectively, the proposed network generates precise saliency maps with uniformly and completely segmented salient objects. Experiments conducted on five benchmark datasets demonstrate that the proposed model outperforms existing networks. Furthermore, we extend the proposed model to the task of RGB-D salient object detection, and it also achieves state-of-the-art performance. The code is available at https://github.com/wuzhe71/DCN. Zhe Wu 0006, Li Su 0003, Qingming Huang |
IEEE Trans. Image Process. | 2 |
| 2020 | Reverse Perspective Network for Perspective-Aware Object CountingabstractOne of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse perspective network explicitly evaluates the perspective distortions, and efficiently corrects the distortions by uniformly warping the input images. Then the proposed network delivers images with similar instance scales to the regressor. Thus the regression network doesn't need multi-scale receptive fields to match the various scales. Besides, to further solve the scale problem of more congested areas, we enhance the corresponding regions of ground-truth with the evaluation errors. Then we force the regressor to learn from the augmented ground-truth via an adversarial process. Furthermore, to verify the proposed model, we collected a vehicle counting dataset based on Unmanned Aerial Vehicles (UAVs). The proposed dataset has fierce scale variations. Extensive experimental results on four benchmark datasets show the improvements of our method against the state-of-the-arts. Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe |
CVPR | 4 |
| 2020 | Weakly-Supervised Crowd Counting Learns from Sorting Rather Than Locations
Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe |
ECCV (8) | 4 |
| 2020 | Siamese Dynamic Mask Estimation Network for Fast Video Object SegmentationabstractVideo object segmentation(VOS) has been a fundamental topic in recent years, and many deep learning-based methods have achieved state-of-the-art performance on multiple benchmarks. However, most of these methods rely on pixel-level matching between the template and the searched frames on the whole image while the targets only occupy a small region. Calculating on the entire image brings lots of additional computation cost. Besides, the whole image may contain some distracting information resulting in many false-positive matching points. To address this issue, motivated by one-stage instance object segmentation methods, we propose an efficient siamese dynamic mask estimation network for fast video object segmentation. The VOS is decoupled into two tasks, i.e., mask feature learning and dynamic kernel prediction. The former is responsible for learning high-quality features to preserve structural geometric information, and the latter learns a dynamic kernel that is used to convolve with the mask feature to generate a mask output. We use Siamese neural network as a feature extractor and directly predict masks after correlation. In this way, we can avoid using pixel-level matching, making our framework more simple and efficient. Experiment results on DAVIS 2016 /2017 datasets show that our proposed methods can run at 35 frames per second on NVIDIA RTX TITAN while preserving competitive accuracy. Dexiang Hong, Guorong Li, Kai Xu 0013, Li Su 0003, Qingming Huang |
ICPR | 4 |
| 2020 | Diverter-Guider Recurrent Network for Diverse Poems Generation from ImageabstractPoem generation from image aims to automatically generate the poetic sentences for presenting the image content or overtone. Previous works focused on 1-to-1 image-poem generation with the demands of poeticness and content relevance. This paper proposes the paradigm of multiple poems generation from one image, which is closer to human poetizing but more challenging. Its key problem is to simultaneously guarantee the diversity of multiple poems with poeticness and relevance. To this end, we propose an end-to-end probabilistic Diverter-Guider Recurrent Network (DG-Net), which is a context-based encoder-decoder generative model with the hierarchical stochastic variables. Specifically, the diverter-variable represents the decoding-context inferred from the input image to diversify the poem themes; the guider-variable is introduced as an attribute decoder to restricts the word-choice with supervised information. Extensive experiments on automatic evaluations and human judgments demonstrate the superior performance of DG-Net than existing poem generation methods. Qualitative study show that our model can generate diverse poems with the poeticness and relevance. Liang Li 0003, Li Su 0003, Shuhui Wang, Chenggang Yan 0001, Zhengjun Zha, Qingming Huang |
ACM Multimedia | 3 |
| 2020 | Towards More Explainability: Concept Knowledge Mining Network for Event RecognitionabstractEvent recognition of untrimmed video is a challenging task due to the big gap between low level visual features and event semantics. Beyond feature learning via deep neural networks, some recent works focus on analyzing event videos using concept-based representation. However, these methods simply aggregate the concept representation vectors of frames or segments, which inevitably introduces information loss on video-level concept knowledge. Moreover, the diversified relation between different concept domains (e.g., scene, object and action) has not been fully explored. To address the above issues, we propose a concept knowledge mining network (CKMN) for event recognition. CKMN is composed of an intra-domain concept knowledge mining subnetwork (IaCKM) and an inter-domain concept knowledge mining subnetwork~(IrCKM). IaCKM aims to obtain a complete concept representation by mining the existing pattern of each concept at different time granularities with dilated temporal pyramid convolution and temporal self-attention, while IrCKM explores the interaction between different types of concepts with co-attention style learning. We evaluate our method on FCVID and ActivityNet datasets. Experimental results show the effectiveness and better interpretability of our model on event analytics. Code is available at https://github.com/qzhb/CKMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang, Qi Tian 0001 |
ACM Multimedia | 4 |
| 2020 | Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video AnalysisabstractEvent analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the spatial range covered by a single feature response, which is crucial in improving the image categorization accuracy. In video domain, video event semantics are actually described by complex interaction among different concepts, while their behaviors vary drastically from one video to another, leading to the difficulty in concept-based analytics for accurate event categorization. To model the concept behavior, we study temporal concept receptive field of concept-based event representation, which encodes the temporal occurrence pattern of different mid-level concepts. Accordingly, we introduce temporal dynamic convolution (TDC) to give stronger flexibility to concept-based event analytics. TDC can adjust the temporal concept receptive field size dynamically according to different inputs. Notably, a set of coefficients are learned to fuse the results of multiple convolutions with different kernel widths that provide various temporal concept receptive field sizes. Different coefficients can generate appropriate and accurate temporal concept receptive field size according to input videos and highlight crucial concepts. Based on TDC, we propose the temporal dynamic concept modeling network~(TDCMN) to learn an accurate and complete concept representation for efficient untrimmed video analysis. Experiment results on FCVID and ActivityNet show that TDCMN demonstrates adaptive event recognition ability conditioned on different inputs, and improve the event recognition performance of Concept-based methods by a large margin. Code is available at https://github.com/qzhb/TDCMN. Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Weigang Zhang, Qingming Huang |
ACM Multimedia | 4 |
| 2020 | Structural Semantic Adversarial Active Learning for Image CaptioningabstractMost image captioning models achieve superior performances with the help of large-scale surprised training data, but it is prohibitively costly to label the image captions. To solve this problem, we propose a structural semantic adversarial active learning (SSAAL) model that leverages both visual and textual information for deriving the most representative samples while maximizing the image captioning performance. SSAAL consists of a semantic constructor, a snapshot& caption (SC) supervisor, and a labeled/unlabeled state discriminator. The constructor is designed to generate a structural semantic representation describing the objects, attributes and object relationships in the image. The SC supervisor is proposed to supervise this representation at the word-level and sentence-level in a multi-task learning manner, which directly relates the representation to ground-truth captions and updates it in the caption generating process. Finally, we introduce a state discriminator to predict the sample state and select images with sufficient semantic and fine-grained diversity. Extensive experiments on standard captioning dataset show that our model outperforms other active learning methods and achieves a competitive performance even though selecting a small amount of samples. Beichen Zhang 0006, Liang Li 0003, Li Su 0003, Shuhui Wang, Jincan Deng, Zhengjun Zha, Qingming Huang |
ACM Multimedia | 3 |
| 2020 | Fixation guided network for salient object detectionabstractConvolutional neural network (CNN) based salient object detection (SOD) has achieved great development in recent years. However, in some challenging cases, i.e. small-scale salient object, low contrast salient object and cluttered background, existing salient object detect methods are still not satisfying. In order to accurately detect salient objects, SOD networks need to fix the position of most salient part. Fixation prediction (FP) focuses on the most visual attractive regions, so we think it could assist in locating salient objects. As far as we know, there are few methods jointly consider SOD and FP tasks. In this paper, we propose a fixation guided salient object detection network (FGNet) to leverage the correlation between SOD and FP. FGNet consists of two branches to deal with fixation prediction and salient object detection respectively. Further, an effective feature cooperation module (FCM) is proposed to fuse complementary information between the two branches. Extensive experiments on four popular datasets and comparisons with twelve state-of-the-art methods show that the proposed FGNet well captures the main context of images and locates salient objects more accurately. Li Su 0003, Weigang Zhang, Qingming Huang |
MMAsia | 2 |
| 2020 | Video Anomaly Detection Using Open Data Filter and Domain AdaptationabstractVideo anomaly detection is a very challenging task because of the rarity, openness, and the definition of the anomalies. Researchers pay more attention to the characteristics of anomalies and have proposed a variety of anomaly detection models. However, most existing methods only use normal events to construct anomaly detection models and ignore the diversity and openness of normal events. Actually, because real-world video data often have an open-ended distribution, some normal patterns hardly ever appeared in the training data. In addition, analogous to human experience in identifying anomalies, rare abnormal events can play a certain role in the detection of similar abnormal events in the dataset. Therefore, assuming that a small number of abnormal events are known, we propose a novel supervised anomaly detection model which explicitly detects open normal events and open abnormal events in the dataset and treats open data and seen data with different classifiers. First, we use the training video to train an imbalanced classifier as the seen data classifier. Then, during the testing phase, an open data filter module isused to divide the test data into seen data and open data. Finally, we directly use the seen data classifier to generate anomaly scores for the seen test data. For the open test data, we adopt a domain adaptation method to reduce the distribution difference between it and the training data and train a new classifier to score for it. Extensive experimental results prove the effectiveness of our model. Chen Zhang 0013, Guorong Li, Li Su 0003, Weigang Zhang, Qingming Huang |
VCIP | 3 |
| 2020 | CSCNet: A Shallow Single Column Network for Crowd CountingabstractCrowd counting in complex scene is an important but challenge task. The scale variation of crowd makes the shallow network hard to extract effective features. In this paper, we propose a shallow single column network named CSCNet for crowd counting. The key component is complementary scale context block (CSCB). It is designed to capture complementary scale context and obtains a high accuracy with limited depth of the network. As far as we know, CSCNet is the shallowest single column network in existing works. We demonstrate our methods on three challenge benchmarks. Compared to state-of-the-art methods, CSCNet achieves comparable accuracy with much less complexity. CSCNet provides an alternative to achieve comparable or even better performance with about 30% of depth and 50% of width decrease. Besides, CSCNet performs more stably on both sparse and congested crowd scenes. Zhida Zhou, Li Su 0003, Guorong Li, Yifang Yang, Qingming Huang |
VCIP | 2 |
| 2020 | Conditional GAN based individual and global motion fusion for multiple object tracking in UAV videos
Hongyang Yu 0001, Guorong Li, Li Su 0003, Bineng Zhong 0001, Hongxun Yao, Qingming Huang |
Pattern Recognit. Lett. | 3 |
| 2019 | Learning Attribute-Specific Representations for Visual TrackingabstractIn recent years, convolutional neural networks (CNNs) have achieved great success in visual tracking. Most of existing methods train or fine-tune a binary classifier to distinguish the target from its background. However, they may suffer from the performance degradation due to insufficient training data. In this paper, we show that attribute information (e.g., illumination changes, occlusion and motion) in the context facilitates training an effective classifier for visual tracking. In particular, we design an attribute-based CNN with multiple branches, where each branch is responsible for classifying the target under a specific attribute. Such a design reduces the appearance diversity of the target under each attribute and thus requires less data to train the model. We combine all attributespecific features via ensemble layers to obtain more discriminative representations for the final target/background classification. The proposed method achieves favorable performance on the OTB100 dataset compared to state-of-the-art tracking methods. After being trained on the VOT datasets, the proposed network also shows a good generalization ability on the UAV-Traffic dataset, which has significantly different attributes and target appearances with the VOT datasets. Yuankai Qi, Shengping Zhang, Weigang Zhang, Li Su 0003, Qingming Huang, Ming-Hsuan Yang 0001 |
AAAI | 4 |
| 2019 | Cascaded Partial Decoder for Fast and Accurate Salient Object DetectionabstractExisting state-of-the-art salient object detection networks rely on aggregating multi-level features of pre-trained convolutional neural networks (CNNs). However, compared to high-level features, low-level features contribute less to performance. Meanwhile, they raise more computational cost because of their larger spatial resolutions. In this paper, we propose a novel Cascaded Partial Decoder (CPD) framework for fast and accurate salient object detection. On the one hand, the framework constructs partial decoder which discards larger resolution features of shallow layers for acceleration. On the other hand, we observe that integrating features of deep layers will obtain relatively precise saliency map. Therefore we directly utilize generated saliency map to recurrently optimize features of deep layers. This strategy efficiently suppresses distractors in the features and significantly improves their representation ability. Experiments conducted on five benchmark datasets exhibit that the proposed model not only achieves state-of-the-art but also runs much faster than existing models. Besides, we apply the proposed framework to optimize existing multi-level feature aggregation models and significantly improve their efficiency and accuracy. Zhe Wu 0006, Li Su 0003, Qingming Huang |
CVPR | 2 |
| 2019 | Stacked Cross Refinement Network for Edge-Aware Salient Object DetectionabstractSalient object detection is a fundamental computer vision task. The majority of existing algorithms focus on aggregating multi-level features of pre-trained convolutional neural networks. Moreover, some researchers attempt to utilize edge information for auxiliary training. However, existing edge-aware models design unidirectional frameworks which only use edge features to improve the segmentation features. Motivated by the logical interrelations between binary segmentation and edge maps, we propose a novel Stacked Cross Refinement Network (SCRN) for salient object detection in this paper. Our framework aims to simultaneously refine multi-level features of salient object detection and edge detection by stacking Cross Refinement Unit (CRU). According to the logical interrelations, the CRU designs two direction-specific integration operations, and bidirectionally passes messages between the two tasks. Incorporating the refined edge-preserving features with the typical U-Net, our model detects salient objects accurately. Extensive experiments conducted on six benchmark datasets demonstrate that our method outperforms existing state-of-the-art algorithms in both accuracy and efficiency. Besides, the attribute-based performance on the SOC dataset show that the proposed model ranks first in the majority of challenging scenes. Code can be found at https://github.com/wuzhe71/SCAN. Zhe Wu 0006, Li Su 0003, Qingming Huang |
ICCV | 2 |
| 2019 | Knowledge-guided Pairwise Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised referring expression grounding (REG) aims at localizing the referential entity in an image according to linguistic query, where the mapping between the image region (proposal) and the query is unknown in the training stage. In referring expressions, people usually describe a target entity in terms of its relationship with other contextual entities as well as visual attributes. However, previous weakly supervised REG methods rarely pay attention to the relationship between the entities. In this paper, we propose a knowledge-guided pairwise reconstruction network (KPRN), which models the relationship between the target entity (subject) and contextual entity (object) as well as grounds these two entities. Specifically, we first design a knowledge extraction module to guide the proposal selection of subject and object. The prior knowledge is obtained in a specific form of semantic similarities between each proposal and the subject/object. Second, guided by such knowledge, we design the subject and object attention module to construct the subject-object proposal pairs. The subject attention excludes the unrelated proposals from the candidate proposals. The object attention selects the most suitable proposal as the contextual proposal. Third, we introduce a pairwise attention and an adaptive weighting scheme to learn the correspondence between these proposal pairs and the query. Finally, a pairwise reconstruction module is used to measure the grounding for weakly supervised learning. Extensive experiments on four large-scale datasets show our method outperforms existing state-of-the-art methods by a large margin. Xuejing Liu, Liang Li 0003, Shuhui Wang, Zhengjun Zha, Li Su 0003, Qingming Huang |
ACM Multimedia | 5 |
| 2019 | Training Efficient Saliency Prediction Models with Knowledge DistillationabstractRecently, deep learning-based saliency prediction methods have achieved significant accuracy improvements. However, they are hard to embed in practical multimedia applications due to large memory consumption and running time caused by complicated architectures. In addition, most methods are fine-tuned from pre-trained models for classification tasks, and networks cannot flexibly be transferred for a new task. In this paper, a condensed and randomly initialized student network is employed to achieve higher efficiency by transferring knowledge from complicated and well-trained teacher networks. This is the first use of knowledge distillation for efficient pixel-wise saliency prediction. Instead of directly minimizing Euclidean distance between feature maps, we propose two statistical representations of feature maps (i.e., first-order and second-order statistics) as knowledge. We conduct experiments on three kinds of teacher networks and four benchmark datasets to verify the effectiveness of the proposed method. Compared with the teacher networks, the student networks achieve an acceleration ratio of 4.56-4.73. Compared with state-of-the-art approaches, the proposed model achieves competitive accuracy with faster running speed (up to 4.38 times) and smaller model size (up to 93.27% reduction). We further embedded the proposed saliency prediction model into a video captioning application. The saliency-embedded approaches improve video captioning on all test metrics with a small complexity cost. The student-model embedded approach achieves 25% time saving with similar performance to the teacher embedded one. Peng Zhang 0024, Li Su 0003, Liang Li 0003, Bing-Kun Bao, Pamela C. Cosman, Guorong Li, Qingming Huang |
ACM Multimedia | 2 |
| 2019 | Accelerating Topic Detection on Web for a Large-Scale Data Set via Stochastic Poisson Deconvolution
Jinzhong Lin, Junbiao Pang, Li Su 0003, Yugui Liu, Qingming Huang |
MMM (1) | 3 |
| 2019 | No-Reference Video Quality Assessment Based on Ensemble of Knowledge and Data-Driven Models
Li Su 0003, Pamela C. Cosman, Qihang Peng |
MMM (2) | 1 |
| 2019 | Learning Coupled Convolutional Networks Fusion for Video Saliency PredictionabstractVisual saliency provides important information for understanding scenes in many computer vision tasks. The existing video saliency algorithms mainly focus on predicting spatial and temporal saliency maps. However, these maps are simply fused without considering the complex dynamic scenes in videos. To overcome this drawback, we propose a deep convolutional fusion framework for video saliency prediction. The proposed model, which is based on coupled fully convolutional networks (FCNs), effectively encodes the spatiotemporal information by integrating spatial and temporal features. We demonstrate that this information is helpful for accurately fusing the spatial and temporal saliency maps according to changes in video scenes. In particular, we gradually design three different deep fusion architectures to investigate how to better utilize the spatiotemporal information. Moreover, we propose a reasonable sampling strategy for selecting suitable training sets for the coupled FCNs. Through extensive experiments, we demonstrate that our model outperforms the state-of-the-art algorithms on four public video saliency data sets. Zhe Wu 0006, Li Su 0003, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Semantic Manifold Alignment in Visual Feature Space for Zero-Shot LearningabstractZero-Shot Learning (ZSL) is getting more attention for its potential to solve a task without training examples, such as to recognize a category of unseen object in computer vision task. Most existing methods are suffered from hubness problem and semantic gap problem. In this paper, we propose a novel strategy based on Aligning Semantic Manifolds in Feature Space (ASMFS) to boost the performance of ZSL. Considering that the semantic representations must be predicted in the location of their corresponding visual instances, we adjust the predicted unseen semantic representations by the average of their K nearest neighbors (K-NN). The experimental results over two basic ZSL models and four public datasets demonstrate the universal enhancement performance of the proposed strategy. It significantly boosts the existing ZSL approaches with low over cost and outperforms eight state-of-the-art methods. Changsu Liao, Li Su 0003, Weigang Zhang, Qingming Huang |
ICME | 2 |
| 2017 | Saliency detection with two-level fully convolutional networksabstractThis paper proposes a deep architecture for saliency detection by fusing pixel-level and superpixel-level predictions. Different from the previous methods that either make dense pixellevel prediction with complex networks or region-level prediction for each region with fully-connected layers, this paper investigates an elegant route to make two-level predictions based on a same simple fully convolutional network via seamless transformation. In the transformation module, we integrate the low level features to model the similarities between pixels and superpixels as well as superpixels and superpixels. The pixel-level saliency map detects and highlights the salient object well and the superpixel-level saliency map preserves sharp boundary in a complementary way. A shallow fusion net is applied to learn to fuse the two saliency maps, followed by a CRF post-refinement module. Experiments on four benchmark data sets demonstrate that our method performs favorably against the state-of-art methods. Li Su 0003, Qingming Huang, Zhe Wu 0006 |
ICME | 2 |
| 2016 | Webpage saliency prediction with multi-features fusionabstractWe proposed a novel model to predict human's visual attention when free-viewing webpages. Compared with natural images, webpages are usually full of salient regions such as logos, text, and faces, while few of them attract human's attention in a short sight. Moreover, webpages perform distinct viewing patterns which are quite different from the natural images. In this paper, we introduced multi-features according to our observation on webpages characters and related eye-tracking data. Further, in order to achieve a flexible adaptation to various types of webpages, we employed a machine-learning framework based on our proposed features. Experimental results demonstrate that our model outperforms other state-of-the-art methods in webpage saliency prediction. Li Su 0003, Bo Wu 0016, Junbiao Pang, Zhe Wu 0006, Qingming Huang |
ICIP | 2 |
| 2016 | Accelerate convolutional neural networks for binary classification via cascading cost-sensitive featureabstractConvolutional Neural Networks (CNNs) have delivered impressive state-of-the-art performances for many vision tasks, while the computation costs of these networks during test-time are notorious. Empirical results have discovered that CNNs have learned the redundant representations both within and across different layers. When CNNs are applied for binary classification, we investigate a method to exploit this redundancy across layers, and construct a cascade of classifiers which explicitly balances classification accuracy and hierarchical feature extraction costs. Our method cost-sensitively selects feature points across several layers from trained networks and embeds non-expensive yet discriminative features into a cascade. Experiments on binary classification demonstrate that our framework leads to drastic test-time improvements, e.g., possible 47.2x speedup for TRECVID upper body detection, 2.82x speedup for Pascal VOC2007 People detection, 3.72x for INRIA Person detection with less than 0.5% drop in accuracies of the original networks. Junbiao Pang, Huihuang Lin, Li Su 0003, Chunjie Zhang 0001, Weigang Zhang, Lijuan Duan, Qingming Huang |
ICIP | 3 |
| 2016 | Robust latent poisson deconvolution from multiple imperfect features for web topic detectionabstractIn web topic detection, detecting “hot” topics from enormous User-Generated Content (UGC) on web data poses two main difficulties that conventional approaches can barely handle: 1) poor feature representations from noisy images and short texts; and 2) uncertain roles of modalities where visual content is either highly or weakly relevant to textual cues due to less-constrained data. In this paper, following the detection by ranking approach, we address the problem by learning a robust shared representation from multiple, noisy and complementary features, and integrating both textual and visual graphs into a k-Nearest Neighbor Similarity Graph (k-N2SG). Then Non-negative Matrix Factorization using Random walk (NMFR) is introduced to generate topic candidates. An efficient fusion of multiple graphs is then done by a Latent Poisson Deconvolution (LPD) which consists of a poisson deconvolution with sparse basis similarities for each edge. Experiments show significantly improved accuracy of the proposed approach in comparison with the state-of-the-art methods on two public data sets. Junbiao Pang, Chunjie Zhang 0001, Liang Li 0003, Li Su 0003, Weigang Zhang, Qingming Huang, Guiping Su |
ICME | 5 |
| 2016 | Compact and robust video fingerprinting using sparse represented featuresabstractIn this paper, we propose a compact and robust video fingerprinting scheme by using sparse represented features (SRF). The SRF are extracted by a two dimensional matching pursuit decomposition (2D-MPD) method. The motivation of using sparse features is that the sparse coding method can significantly reduce the data dimensionality and effectively retain the structure of the images. To further reduce the length of the fingerprint, a two-stage cascade SVD method is applied. The SVD-based feature extraction method can improve the robustness to certain attacks, such as geometric attack. Then, a locally adaptive quantization (AQ) method which considers the local probability distribution of the sample is applied. This method can quantize the real-valued fingerprints into binary bits without degrading the detection performance too much. At last, a secret key based interleaving method is applied to the binary fingerprints. The interleaving can enlarge the Hamming distance of the fingerprints, so the detection performance is improved. According to the experimental results, the proposed method offers a favourable robustness versus discriminability tradeoff over the state-of-the-art video fingerprint methods. Bo Wu 0016, Sridhar Krishnan 0001, Nan Zhang 0015, Li Su 0003 |
ICME | 4 |
| 2016 | Video saliency prediction with optimized optical flow and gravity center biasabstractDynamic videos are viewed fundamentally different from static images. Besides spatial features, motion feature also plays an important role as a temporal factor. Most existing video saliency models usually employ optical flow to represent the motion feature. However, optical flow often suffers from the discontinuity problem. And we also notice that human fixations in one single video frame are much sparser than that in an identical still picture. However, many spatial saliency models take each video frame as static image independently. In this paper, we predict the dynamic visual saliency by fusing spatial and temporal features. In order to construct the temporal relationships among a set of successive frames, we introduce a smoothness operator in optical flow field to obtain more accurate motion feature. Then, considering the sparse property of video saliency, we adapt the weights of the regions surrounding to the saliency gravity center in the final maps. The experiments show that our model is more consistent with humans eye-tracking benchmarks than the state-of-the-art models. Zhe Wu 0006, Li Su 0003, Qingming Huang, Bo Wu 0016, Guorong Li |
ICME | 2 |
| 2015 | Multiple layer parallel motion estimation on GPU for High Efficiency Video Coding (HEVC)abstractThis paper provides a multiple-layer parallel motion estimation (ME) scheme implemented on GPU for High Efficiency Video Coding (HEVC). The scheme is hierarchically structured, including four layers: coding tree unit (CTU), prediction unit (PU), motion vector (MV) selection and instruction optimization. In PU-layer, costs of various PU sizes were obtained through a SAD (sum of absolute differences) look-up table instead of progressive cost merging. And during MV selection, GPU's comparison instruction was used to avoid branches. At the same time, concurrent CTUs processing and SIMD (Single Instruction, Multiple Data) optimization also improve the performance significantly. Experimental results show that the proposed scheme can take full advantage of GPU and achieves over 90 times speedup compared with the HM10.0 using fast ME. Falei Luo, Siwei Ma 0001, Juncheng Ma, Honggang Qi, Li Su 0003, Wen Gao 0001 |
ISCAS | 5 |
| 2014 | Face Distortion Recovery Based on Online Learning Database for Conversational VideoabstractWith the real-time requirement for video conversation, the coding system needs to adopt low delay and low complexity strategy to encode conversational videos, which may result in a significant decline of the video quality under the constrained bandwidth of network. In conversational videos, the face region attracts most human attentions. Therefore, recovering the distortion of face region will effectively improve the visual quality of conversational video. Actually, the participants in a conversation are usually unchanged in a relative long period, and similar facial expressions of the participants would be often repetitive. However, conventional video coding methods just consider the correlation of several neighboring frames while the long-range correlation of similar face regions in the whole conversational video has not been fully used. In this paper, we propose a face distortion recovery system to improve the visual quality of decoded conversational video by online learning an own face feature database for each user. First, at the sender side, the face feature database is established and online updated to include different facial expressions of the person. Then, at the receiver side, the low quality face regions in decoded video are recovered with the face patches in the database. Experimental results show that, under low bits rates the proposed method achieves average 5.22 dB gain with small burden to update the database. Xi Wang 0014, Li Su 0003, Honggang Qi, Qingming Huang, Guorong Li |
IEEE Trans. Multim. | 2 |
| 2013 | Online Learning Based Face Distortion Recovery for Conversational Video CodingabstractIn a video conversation, the participants usually remain the same. As the conversation continues, similar facial expressions of the same person would occur intermittently. However, the correlation of similar face features has not been fully used since the conventional methods only focus on independent frames. We set up a face feature database and updated it online to include new facial expressions during the whole conversation. At the receiver side, the database is used to recover the face distortion and thus improve the visual quality. Additionally, the proposed method brings small burden to update the database and is generic to various CODEC. Xi Wang 0014, Li Su 0003, Qingming Huang, Guorong Li, Honggang Qi |
DCC | 2 |
| 2012 | Motion Based Perceptual Distortion and Rate Optimization for Video CodingabstractMost conventional distortion metrics regard a video frame as a static image, and seldom exploit using the motion information of video frames in succession. Moreover, these methods usually calculate the visual distortion based on the independent spatial pixels. Recently, many researches show that the way people perceive the video signals is similar to the way filters process signals in the frequency domain. Therefore, in order to achieve better visual quality, we introduce a novel distortion measurement into the video coding system, which is consistent with human visual perception, and establish a perception-based rate-distortion optimization model. In this paper, we adopt Gabor filter family to decompose the video signals into frequency domain, and combine the video motion information to measure the perceptual distortion. We call it Motion tuned Distortion metric For Video coding (MDFV). After that we set up an MDFV based rate-distortion optimization model to select the best encoding mode. The experimental results show that the proposed approach is effective. Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu, Ling-Yu Duan |
ICME | 2 |
| 2011 | Visual perception based Lagrangian rate distortion optimization for video codingabstractIn the conventional rate distortion optimization (RDO) video coding, the measure of distortion is mainly from the perspective of signal processing, while dose not fully take into account the characteristics of visual perception. People have concerns about not only the information of independent pixels, but also the temporal and spatial correlations between them. For different video content, human visual perception has different sensitivity. In this paper, in order to establish a RDO model which is more consistent with the human visual perception, we introduce the structural similarity and the content saliency information into the distortion metric. An adaptive Lagrange multiplier selection scheme is presented to allocate the bit resources more rationally by keeping the balance of the bit-rate and the visual quality. Experimental results show that the proposed method averagely reduces 10.14% bit-rate under the similar visual quality. Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu |
ICIP | 2 |
| 2011 | News video story sentiment classification and rankingabstractIn this paper, we present a novel approach for news video story sentiment analysis. Two research challenges are addressed: news video story sentiment classification and ranking. For classification, a graph based semi-supervised learning approach is utilized to classify the news stories into sentiment classes. Graph based semi-supervised learning is able to tackle the problem of lacking labeled data. After classification, two sentiment classes are obtained: positive and negative. In order to project the news videos into sentiment space, a multimodal approach by fusing the text sentiment and visual representation scores is adopted to rank the videos in each class. For sentiment representation, inter and intra sentiment class analysis is conducted based on affinity propagation clustering and PageRank algorithm. A user study is conducted to evaluate the video ranking performance. The experimental results on the selected topics are promising and demonstrate the proposed approach is effective. Chunxi Liu, Li Su 0003, Qingming Huang, Shuqiang Jiang |
ICME | 2 |
| 2010 | Bridging the gap between objective score and subjective preference in video quality assessmentabstractNowadays, the issue of objective video quality assessment has been extensively studied. However, the human visual system (HVS) is the ultimate receiver for videos thus leading to a gap between objective scores calculated by computers and subjective preferences given by observers. In this paper, we focus on bridging this gap by introducing a psychological criterion called contrast effect. That is, because of the impression about the quality of previous frame still remaining in observers' minds, they tend to underestimate or overestimate the quality of the current one. Noticing this fact, we propose a video quality assessment system with an additional revision module to bridge the gap mentioned above. Firstly, the video is described by several representative clips with large entropy values. Then, we present Quality Words (including luminance, contrast, structure and spatio-temporal texture) to evaluate the quality of distorted video. To characterize the spatio-temporal texture, a new descriptor called Rotation Sensitive 3D Texture Pattern (RS-3D) is proposed. Finally, we revise the result in the revision module motivated by contrast effect. Experiments on VQEG Phase I FR-TV test dataset verify the effectiveness of our method. Qianqian Xu 0001, Li Su 0003, Shuqiang Jiang, Qingming Huang |
ICME | 3 |
| 2009 | Complexity-Constrained H.264 Video EncodingabstractIn this paper, a joint complexity-distortion optimization approach is proposed for real-time H.264 video encoding under the power-constrained environment. The power consumption is first translated to the encoding computation costs measured by the number of scaled computation units consumed by basic operations. The solved problem is then specified to be the allocation and utilization of the computational resources. A computation allocation model (CAM) with virtual computation buffers is proposed to optimally allocate the computational resources to each video frame. In particular, the proposed CAM and the traditional hypothetical reference decoder model have the same temporal phase in operations. Further, to fully utilize the allocated computational resources, complexity-configurable motion estimation (CAME) and complexity-configurable mode decision (CAMD) algorithms are proposed for H.264 video encoding. In particular, the CAME is performed to select the path of motion search at the frame level, and the CAMD is performed to select the order of mode search at the macroblock level. Based on the hierarchical adjusting approach, the adaptive allocation of computational resources and the fine scalability of complexity control can be achieved. Li Su 0003, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Real-time video coding under power constraint based on H.264 codecabstractIn this paper, we propose a joint power-distortion optimization scheme for real-time H.264 video encoding under the power constraint. Firstly, the power constraint is translated to the complexity constraint based on DVS technology. Secondly, a computation allocation model (CAM) with virtual buffers is proposed to facilitate the optimal allocation of constrained computational resource for each frame. Thirdly, the complexity adjustable encoder based on optimal motion estimation and mode decision is proposed to meet the allocated resource. The proposed scheme takes the advantage of some new features of H.264/AVC video coding tools such as early termination strategy in fast ME. Moreover, it can avoid suffering from the high overhead of the parametric power control algorithms and achieve fine complexity scalability in a wide range with stable rate-distortion performance. The proposed scheme also shows the potential of a further reduction of computation and power consumption in the decoding without any change on the existing decoders. Li Su 0003, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001 |
VCIP | 1 |
| 2004 | Improved error concealment algorithms based on H.264/AVC non-normative decoderabstractWe propose several improved error concealment (EC) algorithms based on the H.264/AVC non-normative decoder. The major differences are that motion compensated EC is introduced for intra frames, whereas spatial EC is introduced for inter frames. As for the EC of intra frames, scene change detection, motion activity detection and MV retrieval are hierarchically performed to decide whether spatial or temporal information is to be used. As for the EC of inter frames, scene change is also detected to avoid merging the scenes from different video shots. Therefore, the main idea of the proposed algorithms is that both spatial and temporal correlations are utilized for the EC of intra and inter frames. Both subjective and objective simulations under Internet conditions show that the proposed algorithm greatly outperforms that in the H.264/AVC non-normative decoder. Li Su 0003, Yuan Zhang 0014, Wen Gao 0001, Qingming Huang, Yan Lu 0001 |
ICME | 1 |