EDBT 2026 Demo / reviewers in the wild / expert
Haichao Shi
dblp:180/1745
· DBLP profile ↗
30ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0001-8846-1853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoFAIR: Self-improving Post-training for Robust and Explainable AIGC Forensics with Lightweight Reward Distillation
Qingtong Liu, Xinyao Zhao, Haichao Shi, Weiguo Lin, Junjun Si, Zehua Ji |
ICIC (12) | 4 |
| 2025 | MUN: Image Forgery Localization Based on M³ Encoder and UN DecoderabstractImage forgeries can entirely change the semantic information of an image, and can be used for unscrupulous purposes. In this paper, we propose a novel image forgery localization network named as MUN, which consists of an M^3 encoder and a UN decoder. Firstly, the M^3 encoder is constructed based on a Multi-scale Max-pooling query module to extract Multi-clue forged features. Noiseprint++ is adopted to assist the RGB clue, and its deployment methodology is discussed. A Multi-scale Max-pooling Query (MMQ) module is proposed to integrate RGB and noise features. Secondly, a novel UN decoder is proposed to extract hierarchical features from both top-down and bottom-up directions, reconstructing both high-level and low-level features at the same time. Thirdly, we formulate an IoU-recalibrated Dynamic Cross-Entropy (IoUDCE) loss to dynamically adjust the weights on forged regions according to IoU which can adaptively balance the influence of authentic and forged regions. Last but not least, we propose a data augmentation method, i.e., Deviation Noise Augmentation (DNA), which acquires accessible prior knowledge of RGB distribution to improve the generalization ability. Extensive experiments on publicly available datasets show that MUN outperforms the state-of-the-art works. Shuhuan Chen, Haichao Shi, Xiaoyu Zhang 0002, Song Xiao 0001, Qiang Cai 0001 |
AAAI | 3 |
| 2025 | Generate First, Then Sample: Enhancing Fake News Detection with LLM-Augmented Reinforced SamplingabstractThe spread of fake news on online platforms has long been a pressing concern.Considering this, extensive efforts have been made to develop fake news detectors.However, a major drawback of these models is their relatively low performance -lagging by more than 20% -in identifying fake news compared to real news, making them less suitable for practical deployment.This gap is likely due to an imbalance in the dataset and the model's inadequate understanding of data distribution on the targeted platform.In this work, we focus on improving the model's effectiveness in detecting fake news.To achieve this, we first adopt an LLM to generate fake news in three different styles which are later incorporated into the training set, to augment the representation of fake news.Then, we apply Reinforcement Learning to dynamically sample fake news, allowing the model to learn the optimal real-to-fake news ratio for training an effective fake news detector on the targeted platform.This approach allows our model to perform effectively even with a limited amount of annotated news data and consistently improve detection accuracy across different platforms.Experimental results demonstrate that our approach achieves state-of-the-art performance on two benchmark datasets, improving fake news detection performance by 24.02% and 11.06% respectively. Yimeng Gu, Huidong Liu, Haichao Shi |
ACL (1) | 6 |
| 2025 | Concept-Centric Learning for Weakly-Supervised Temporal Sentence GroundingabstractWeakly-supervised temporal sentence grounding remains challenging when learning to temporally locate event boundaries in a video related to the given query. Conventional methods that rely on global query supervision suffer from key limitations (e.g. insufficient interactions between local video-query representations). To tackle it, in this paper, we propose the ConceptNet which achieves fine-grained alignments by leveraging concept-centric learning. Specifically, we extract the essential concepts (i.e., verbs and nouns) and design two corresponding networks: Temporal-Dynamic Network (TDNet) and Visual-Semantic Network (VSNet). The TDNet introduces a prompt-guided autoregressive task aimed at facilitating the learning of temporal dependencies, with the objective of enhancing the model more sensitive to event progression rather than static scenes. Besides, the VSNet is designed to answer the masked query templates from batch-wise concept pools for semantic alignments. Extensive evaluations on Charades-STA and ActivityNet Captions show the superiority of our ConceptNet when compared to previous state-of-the-arts. The code is available at https://github.com/rubyrosecraft/ConceptNet. Yaru Zhang, Haichao Shi |
ICME | 2 |
| 2025 | R2FND: Reinforced Rationale Learning for Fake News Detection with LLMsabstractThe widespread dissemination of fake news on social media has become a significant challenge, undermining public trust, and spreading misinformation worldwide. Detecting fake news effectively is critical to maintaining information integrity and mitigating its negative impacts on public health, politics, and societal stability. However, existing methods often lack interpretability, rely on shallow textual features, and fail to leverage contextual information effectively, limiting their practical applicability. To address these issues, we propose R2FND (Reinforced Rationale learning for Fake News Detection), a novel framework designed to improve both performance and interpretability in fake news detection. R2FND utilizes LLMs to generate news content and corresponding rationales, and employs reinforcement learning (RL) to select the most relevant rationales. These rationales, combined with the original news text, are then processed by a classification model to identify fake news. The framework integrates rationale learning and RL in an end-to-end framework, ensuring high-quality rationale selection while improving transparency by linking detection decisions to interpretable evidence. Extensive experiments on two benchmark datasets, Weibo21 and GossipCop, demonstrate the effectiveness of R2FND, achieving state-of-the-art detection accuracy and providing interpretable rationales for enhanced trustworthiness. Yimeng Gu, Haichao Shi |
IJCNN | 3 |
| 2025 | GELog: a GPT-Enhanced Log Representation Method for Anomaly DetectionabstractLog anomaly detection is a critical aspect of Artificial Intelligence for IT Operations (AIOps), as it enables the timely identification of system failures, thereby facilitating program understanding throughout the entire software maintenance and engineering life cycles. While existing methods only leverage raw log information for anomaly detection, they struggle to address challenges such as log differences due to log evolution, noise from log parsing, and stylistic differences between logs and natural language. To overcome these limitations, we propose GELog, an innovative log anomaly detection method. Specifically, GELog initially employs GPT to semantically enhance log templates. Subsequently, it extracts semantic vectors using pretrained sentence-bert and introduces an attention-based semantic fusion module that integrates the semantic representations of both the original and enhanced logs. Finally, GELog utilizes a Transformer-based model for anomaly detection. We evaluated the performance of GELog on four publicly available datasets, and the experimental results demonstrate that GELog significantly enhances the semantic representation of logs, achieving superior anomaly detection performance. Wenwu Xu, Haichao Shi, Guoqiao Zhou, Junliang Yao |
ICPC | 3 |
| 2025 | Knowledge Negative Distillation: Circumventing Overfitting to Unlock More Generalizable Deepfake Detection
Haichao Shi, Yaru Zhang |
ACM Multimedia | 2 |
| 2025 | Learning to Discriminate: Generalizable Deepfake Detection with Progressive Latent Space Noising and Feature Shift
Shuhuan Chen, Haichao Shi |
PRCV (15) | 2 |
| 2025 | Straighter Flow Matching via a Diffusion-Based Coupling Prior
Siyu Xing, Jie Cao 0002, Huaibo Huang, Haichao Shi, Xiaoyu Zhang 0002 |
PRCV (8) | 4 |
| 2025 | Global Cross-Entropy Loss for Deep Face RecognitionabstractContemporary deep face recognition techniques predominantly utilize the Softmax loss function, designed based on the similarities between sample features and class prototypes. These similarities can be categorized into four types: in-sample target similarity, in-sample non-target similarity, out-sample target similarity, and out-sample non-target similarity. When a sample feature from a specific class is designated as the anchor, the similarity between this sample and any class prototype is referred to as in-sample similarity. In contrast, the similarity between samples from other classes and any class prototype is known as out-sample similarity. The terms target and non-target indicate whether the sample and the class prototype used for similarity calculation belong to the same identity or not. The conventional Softmax loss function promotes higher in-sample target similarity than in-sample non-target similarity. However, it overlooks the relation between in-sample and out-sample similarity. In this paper, we propose Global Cross-Entropy loss (GCE), which promotes 1) greater in-sample target similarity over both the in-sample and out-sample non-target similarity, and 2) smaller in-sample non-target similarity to both in-sample and out-sample target similarity. In addition, we propose to establish a bilateral margin penalty for both in-sample target and non-target similarity, so that the discrimination and generalization of the deep face model are improved. To bridge the gap between training and testing of face recognition, we adapt the GCE loss into a pairwise framework by randomly replacing some class prototypes with sample features. We designate the model trained with the proposed Global Cross-Entropy loss as GFace. Extensive experiments on several public face benchmarks, including LFW, CALFW, CPLFW, CFP-FP, AgeDB, IJB-C, IJB-B, MFR-Ongoing, and MegaFace, demonstrate the superiority of GFace over other methods. Additionally, GFace exhibits robust performance in general visual recognition task. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Guoying Zhao 0001, Zhen Lei 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | AS-GCL: Asymmetric Spectral Augmentation on Graph Contrastive LearningabstractGraph Contrastive Learning (GCL) has emerged as the foremost approach for self-supervised learning on graph-structured data. GCL reduces reliance on labeled data by learning robust representations from various augmented views. However, existing GCL methods typically depend on consistent stochastic augmentations, which overlook their impact on the intrinsic structure of the spectral domain, thereby limiting the model's ability to generalize effectively. To address these limitations, we propose a novel paradigm called AS-GCL that incorporates asymmetric spectral augmentation for graph contrastive learning. A typical GCL framework consists of three key components: graph data augmentation, view encoding, and contrastive loss. Our method introduces significant enhancements to each of these components. Specifically, for data augmentation, we apply spectral-based augmentation to minimize spectral variations, strengthen structural invariance, and reduce noise. With respect to encoding, we employ parameter-sharing encoders with distinct diffusion operators to generate diverse, noise-resistant graph views. For contrastive loss, we introduce an upper-bound loss function that promotes generalization by maintaining a balanced distribution of intra- and inter-class distance. To our knowledge, we are the first to encode augmentation views of the spectral domain using asymmetric encoders. Extensive experiments on eight benchmark datasets across various node-level tasks demonstrate the advantages of the proposed method. Ruyue Liu, Rong Yin 0001, Yong Liu 0018, Xiaoshuai Hao, Haichao Shi, Can Ma, Weiping Wang 0005 |
IEEE Trans. Multim. | 5 |
| 2024 | Denoised Dual-Level Contrastive Network for Weakly-Supervised Temporal Sentence Grounding
Yaru Zhang, Haichao Shi |
CVM (2) | 3 |
| 2024 | Masked Face TransformerabstractThe COVID-19 pandemic makes wearing masks mandatory. Existing CNN-based face recognition (FR) systems suffer from severe performance degradation as masks occlude the vital facial regions. Recently, Vision Transformers have shown promising performance in various vision tasks with quadratic computation costs. Swin Transformer first proposes a successive window attention mechanism allowing the cross-window connection and more computational efficiency. Despite its potential, the deployment of Swin Transformer in masked face recognition encounters two challenges: 1) the attention range is insufficient to capture locally compatible face regions. 2) Masked face recognition can be defined as an occlusion-robust classification task with a known occlusion position, i.e., the position of the mask is minor-varying, which is overlooked but efficient in improving the model’s recognition accuracy. To alleviate the above problem, we propose a Masked Face Transformer (MFT) with Masked Face-compatible Attention (MFA). The proposed MFA 1) introduces two additional window partition configurations, e.g., row shift and column shift, to enlarge the attention range in Swin with invariant computation costs, and 2) suppresses the interaction between the masked and non-masked regions to retain their discrepancies. Additionally, as mask occlusion leads to a separation between the masked and non-masked samples of the same identity, we propose to explore the relationship between them by a ClassFormer module to enhance intra-class aggregation. Extensive experiments show that MFT outperforms state-of-the-art masked face recognition methods in both simulated and real masked face testing datasets. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Rumor Detection with Diverse Counterfactual EvidenceabstractThe growth in social media has exacerbated the threat of fake news to individuals and communities. This draws increasing attention to developing efficient and timely rumor detection methods. The prevailing approaches resort to graph neural networks (GNNs) to exploit the post-propagation patterns of the rumor-spreading process. However, these methods lack inherent interpretation of rumor detection due to the black-box nature of GNNs. Moreover, these methods suffer from less robust results as they employ all the propagation patterns for rumor detection. In this paper, we address the above issues with the proposed Diverse Counterfactual Evidence framework for Rumor Detection (DCE-RD). Our intuition is to exploit the diverse counterfactual evidence of an event graph to serve as multi-view interpretations, which are further aggregated for robust rumor detection results. Specifically, our method first designs a subgraph generation strategy to efficiently generate different subgraphs of the event graph. We constrain the removal of these subgraphs to cause the change in rumor detection results. Thus, these subgraphs naturally serve as counterfactual evidence for rumor detection. To achieve multi-view interpretation, we design a diversity loss inspired by Determinantal Point Processes (DPP) to encourage diversity among the counterfactual evidence. A GNN-based rumor detection model further aggregates the diverse counterfactual evidence discovered by the proposed DCE-RD to achieve interpretable and robust rumor detection results. Extensive experiments on two real-world datasets show the superior performance of our method. Our code is available at https://github.com/Vicinity111/DCE-RD. Kaiwei Zhang, Junchi Yu, Haichao Shi, Jian Liang 0001, Xiaoyu Zhang 0002 |
KDD | 3 |
| 2023 | OW-TAL: Learning Unknown Human Activities for Open-World Temporal Action Localization
Yaru Zhang, Xiaoyu Zhang 0002, Haichao Shi |
Pattern Recognit. | 3 |
| 2023 | Prism: Real-Time Privacy Protection Against Temporal Network Traffic AnalyzersabstractTraffic analysis is widely used in network monitoring. However, the attackers can sometimes infer sensitive information from the patterns of the encrypted network traffic, which poses a threat to network security. Most existing countermeasures are proposed to obfuscate traffic flows using adversarial examples. However, there are two challenges when adding perturbations to live network traffic. Firstly, the perturbations imposed on the feature space cannot be conveniently projected to original traffic flows in feature-space based methods. Secondly, it is laborious and impractical to apply symmetrical framework to encode/decode the adversarial traffic in traffic-space based approaches. To address the above issues, in this paper, we propose an asymmetric defending scheme, namelyPrism, to protect theliveconnection privacy against attacks of temporal network traffic analyzers. Specifically,Prismfirst extracts standardized temporal features via Power-Law Division (PLD) algorithm, and then employs Time-stacked State Transition Model (TSTM) to obtain the fingerprint of each application. Finally,Prismdefends against the analyzers with online traffic perturbation. Since thePrismis designed as a traffic-space based defender with asymmetric defending structure, the deployment is lightweight and efficient. Experimental results on two real-world datasets demonstrate the effectiveness and generalization of our adversarial perturbations. In particular, it is encouraging to see that our proposed defending scheme outperforms the advanced countermeasures, such as adversarial training and traffic filter. Wenhao Li 0005, Xiaoyu Zhang 0002, Huaifeng Bao, Zhaoxuan Li, Haichao Shi, Qiang Wang 0059 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | StochasticFormer: Stochastic Modeling for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WS-TAL) aims to identify the time intervals corresponding to actions of interest in untrimmed videos with video-level weak supervision. For most existing WS-TAL methods, two commonly encountered challenges are under-localization and over-localization, which inevitably bring about severe performance deterioration. To address the issues, this paper proposes a transformer-structured stochastic process modeling framework, namely StochasticFormer, to fully investigate finer-grained interactions among the intermediate predictions to achieve further refined localization. StochasticFormer is built on a standard attention-based pipeline to derive preliminary frame/snippet-level predictions. Then, the pseudo localization module generates variable-length pseudo action instances with the corresponding pseudo labels. Using the pseudo "action instance - action category" pairs as fine-grained pseudo supervision, the stochastic modeler aims to learn the underlying interaction among the intermediate predictions with an encoder-decoder network. The encoder consists of the deterministic and latent path to capture the local and global information, which are subsequently integrated by the decoder to obtain reliable predictions. The framework is optimized with three carefully designed losses, i.e. the video-level classification loss, the frame-level semantic coherence loss, and the ELBO loss. Extensive experiments on two benchmarks, i.e., THUMOS14 and ActivityNet1.2, have shown the efficacy of StochasticFormer compared with the state-of-the-art methods. Haichao Shi, Xiaoyu Zhang 0002 |
IEEE Trans. Image Process. | 1 |
| 2023 | AdapNet: Adaptability Decomposing Encoder-Decoder Network for Weakly Supervised Action Recognition and LocalizationabstractThe point process is a solid framework to model sequential data, such as videos, by exploring the underlying relevance. As a challenging problem for high-level video understanding, weakly supervised action recognition and localization in untrimmed videos have attracted intensive research attention. Knowledge transfer by leveraging the publicly available trimmed videos as external guidance is a promising attempt to make up for the coarse-grained video-level annotation and improve the generalization performance. However, unconstrained knowledge transfer may bring about irrelevant noise and jeopardize the learning model. This article proposes a novel adaptability decomposing encoder-decoder network to transfer reliable knowledge between the trimmed and untrimmed videos for action recognition and localization by bidirectional point process modeling, given only video-level annotations. By decomposing the original features into the domain-adaptable and domain-specific ones based on their adaptability, trimmed-untrimmed knowledge transfer can be safely confined within a more coherent subspace. An encoder-decoder-based structure is carefully designed and jointly optimized to facilitate effective action classification and temporal localization. Extensive experiments are conducted on two benchmark data sets (i.e., THUMOS14 and ActivityNet1.3), and the experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035, Jing Dong 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | ProGraph: Robust Network Traffic Identification With Graph PropagationabstractNetwork traffic identification is critical for effective network management. Existing methods mostly focus on invariant network environments with stable attribute distributions. Unfortunately, however, they can hardly be adaptive to the variation of practical networks and suffer from significant performance degradation. This problem largely stems from the over-dependence of existing methods on the vulnerable side-channel features. To address this issue, in this paper we propose a graph-based approach, namely ProGraph, to ensure robust network traffic classification among various network environments. The core idea of ProGraph is to construct a correlation graph with session clusters aggregated from different networks, based on which graph propagation can be effectively implemented to predict labels of testing nodes in an iterative manner. ProGraph enhances the correlation between clusters of the same class to provide reliable paths for label dissemination from the labeled clusters to the testing ones. It is encouraging to see that the proposed ProGraph achieves an accuracy of 92.25% in networks with constant attributes, while remaining stable with the accuracy of 90.89% when deployed in different networks, which significantly outperforms the state-of-the-art approaches. Meanwhile, ProGraph can accurately identify the novel classes which do not exist in the training dataset, with an AUC of 95.11. Last but not least, a carefully constructed dataset, namely CrossNet2021, containing network traffic of 20 classes of applications from two distinct networking scenarios, is made publicly available to support further research. Wenhao Li 0005, Xiaoyu Zhang 0002, Huaifeng Bao, Haichao Shi, Qiang Wang 0059 |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | Dynamic Graph Modeling for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised action localization is a challenging task that aims to localize action instances in untrimmed videos given only video-level supervision. Existing methods mostly distinguish action from background via attentive feature fusion with RGB and optical flow modalities. Unfortunately, this strategy fails to retain the distinct characteristics of each modality, leading to inaccurate localization under hard-to-discriminate cases such as action-context interference and in-action stationary period. As an action is typically comprised of multiple stages, an intuitive solution is to model the relation between the finer-grained action segments to obtain a more detailed analysis. In this paper, we propose a dynamic graph-based method, namely DGCNN, to explore the two-stream relation between action segments. To be specific, segments within a video which are likely to be actions are dynamically selected to construct an action graph. For each graph, a triplet adjacency matrix is devised to explore the temporal and contextual correlations between the pseudo action segments, which consists of three components, i.e., mutual importance, feature similarity, and high-level contextual similarity. The two-stream dynamic pseudo graphs, along with the pseudo background segments, are used to derive more detailed video representation. For action localization, a non-local based temporal refinement module is proposed to fully leverage the temporal consistency between consecutive segments. Experimental results on three datasets, i.e., THUMOS14, ActivityNet v1.2 and v1.3, demonstrate that our method is superior to the state-of-the-arts. Haichao Shi, Xiaoyu Zhang 0002, Lixing Gong, Yong Li 0034, Yongjun Bao |
ACM Multimedia | 1 |
| 2022 | Consistent Sub-Decision Network for Low-Quality Masked Face RecognitionabstractThe COVID-19 pandemic makes wearing masks mandatory in supermarkets, pharmacies, public transport, etc. Existing facial recognition systems encounter severe performance degradation as the masks occlude key facial regions. Recently, simulation-based methods are proposed to generate masked faces from unmasked faces. However, among simulated faces, there are low-quality samples with negative occlusion, which leads to ambiguous or absent facial features. In this paper, we propose a consistent sub-decision network to obtain sub-decisions that correspond to different facial regions and constrain sub-decisions by weighted bidirectional KL divergence to make the network concentrate on the upper faces without occlusion. In addition, we perform knowledge distillation to drive the masked face embeddings towards an approximation of the original data distribution to mitigate the information loss. Experiments show that the proposed method performs better than the baseline on public masked face recognition datasets, i.e., RMFD, MFR2, and MLFW. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Action Shuffling for Weakly Supervised Temporal LocalizationabstractWeakly supervised action localization is a challenging task with extensive applications, which aims to identify actions and the corresponding temporal intervals with only video-level annotations available. This paper analyzes the order-sensitive and location-insensitive properties of actions, and embodies them into a self-augmented learning framework to improve the weakly supervised action localization performance. To be specific, we propose a novel two-branch network architecture with intra/inter-action shuffling, referred to as ActShufNet. The intra-action shuffling branch lays out a self-supervised order prediction task to augment the video representation with inner-video relevance, whereas the inter-action shuffling branch imposes a reorganizing strategy on the existing action contents to augment the training set without resorting to any external resources. Furthermore, the global-local adversarial training is presented to enhance the model's robustness to irrelevant noises. Extensive experiments are conducted on three benchmark datasets, and the results clearly demonstrate the efficacy of the proposed method. Xiaoyu Zhang 0002, Haichao Shi, Xinchu Shi |
IEEE Trans. Image Process. | 2 |
| 2021 | Flexible Non-Autoregressive Extractive Summarization with Threshold: How to Extract a Non-Fixed Number of Summary SentencesabstractSentence-level extractive summarization is a fundamental yet challenging task, and recent powerful approaches prefer to pick sentences sorted by the predicted probabilities until the length limit is reached, a.k.a. ``Top-K Strategy''. This length limit is fixed based on the validation set, resulting in the lack of flexibility. In this work, we propose a more flexible and accurate non-autoregressive method for single document extractive summarization, extracting a non-fixed number of summary sentences without the sorting step. We call our approach ThresSum as it picks sentences simultaneously and individually from the source document when the predicted probabilities exceed a threshold. During training, the model enhances sentence representation through iterative refinement and the intermediate latent variables receive some weak supervision with soft labels, which are generated progressively by adjusting the temperature with a knowledge distillation algorithm. Specifically, the temperature is initialized with high value and drops along with the iteration until a temperature of 1. Experimental results on CNN/DM and NYT datasets have demonstrated the effectiveness of ThresSum, which significantly outperforms BERTSUMEXT with a substantial improvement of 0.74 ROUGE-1 score on CNN/DM. Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Pengfei Yin, Shi Wang 0002 |
AAAI | 3 |
| 2021 | SAPS: Self-Attentive Pathway Search for weakly-supervised action localization with background-action augmentation
Xiaoyu Zhang 0002, Yaru Zhang, Haichao Shi, Jing Dong 0003 |
Comput. Vis. Image Underst. | 3 |
| 2021 | Weakly-supervised action localization via embedding-modeling iterative optimization
Xiaoyu Zhang 0002, Haichao Shi, Peng Li 0035, Zekun Li 0001, Peng Ren 0001 |
Pattern Recognit. | 2 |
| 2020 | Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosabstractWeakly supervised action recognition and localization for untrimmed videos is a challenging problem with extensive applications. The overwhelming irrelevant background contents in untrimmed videos severely hamper effective identification of actions of interest. In this paper, we propose a novel multi-instance multi-label modeling network based on spatio-temporal pre-trimming to recognize actions and locate corresponding frames in untrimmed videos. Motivated by the fact that person is the key factor in a human action, we spatially and temporally segment each untrimmed video into person-centric clips with pose estimation and tracking techniques. Given the bag-of-instances structure associated with video-level labels, action recognition is naturally formulated as a multi-instance multi-label learning problem. The network is optimized iteratively with selective coarse-to-fine pre-trimming based on instance-label activation. After convergence, temporal localization is further achieved with local-global temporal class activation map. Extensive experiments are conducted on two benchmark datasets, i.e. THUMOS14 and ActivityNet1.3, and experimental results clearly corroborate the efficacy of our method when compared with the state-of-the-arts. Xiaoyu Zhang 0002, Haichao Shi, Peng Li 0035 |
AAAI | 2 |
| 2020 | DistilSum: : Distilling the Knowledge for Extractive SummarizationabstractA popular choice for extractive summarization is to conceptualize it as sentence-level classification, supervised by binary labels. While the common metric ROUGE prefers to measure the text similarity, instead of the performance of classifier. For example, BERTSUMEXT, the best extractive classifier so far, only achieves a precision of 32.9% at the top 3 extracted sentences ([email protected]) on CNN/DM dataset. It is obvious that current approaches cannot model the complex relationship of sentences exactly with 0/1 targets. In this paper, we introduce DistilSum, which contains teacher mechanism and student model. Teacher mechanism produces high entropy soft targets at a high temperature. Our student model is trained with the same temperature to match these informative soft targets and tested with temperature of 1 to distill for ground-truth labels. Compared with large version of BERTSUMEXT, our experimental result on CNN/DM achieves a substantial improvement of 0.99 ROUGE-L score (text similarity) and 3.95 [email protected] score (performance of classifier). Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Yanbing Liu 0007, Jianlong Tan |
CIKM | 3 |
| 2019 | Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak SupervisionabstractAction recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i.e., THUMOS14 and ActivityNet1.3), and experimental results clearly corroborate the efficacy of our method. Xiaoyu Zhang 0002, Haichao Shi, Kai Zheng 0001, Xiaobin Zhu 0001, Lixin Duan |
AAAI | 2 |
| 2019 | Weakly-Supervised Action Recognition and Localization via Knowledge Transfer
Haichao Shi, Xiaoyu Zhang 0002 |
PRCV (1) | 1 |
| 2019 | Active semi-supervised learning based on self-expressive correlation with generative adversarial networks
Xiaoyu Zhang 0002, Haichao Shi, Xiaobin Zhu 0001, Peng Li 0035 |
Neurocomputing | 2 |