Sihui Luo 0001

dblp:198/1554-1 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-2822-0446ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Illumination-aware softmask guided shadow removal
Lianmeng Wei, Sihui Luo 0001
J. Vis. Commun. Image Represent.2
2024 Multimodal Video Highlight Detection with Noise-Robust Learning
abstract
Video highlight detection aims to select the most interesting and attractive clips from lengthy videos, which is crucial for enhancing the video editing and viewing experience on social media platforms. Existing video highlight detection methods predominantly rely on visual modality information, and underutilize the abundant multimodality of videos. Furthermore, in supervised video analysis tasks, subjective judgments during label annotation can generate uncertain noise labels that negatively impact the learning process. To address these issues, we propose a noise-robust multimodal video highlight detection approach. Our approach first enhances feature representation by incorporating multimodal representations of a video’s visual and auditory information. This allows for the extraction of complementary information from different modalities. We then implement a noise-cleaning mechanism that utilizes multiple modalities to clean noise samples. This helps to suppress the negative impact of noise samples on the learning process, ensuring that the network learns more robust features from clean samples. We evaluate our approach on two public datasets, YouTube Highlights and TVSum, and demonstrate its efficacy in mitigating the impact of noise labels, while also improving the accuracy and robustness of video highlight detection.
Yinhui Jiang, Sihui Luo 0001, Lijun Guo
IJCNN2
2024 MCT-VHD: Multi-modal contrastive transformer for video highlight detection
Yinhui Jiang, Sihui Luo 0001, Lijun Guo
J. Vis. Commun. Image Represent.2
2023 Deep semantic image compression via cooperative network pruning
Sihui Luo 0001, Gongfan Fang, Mingli Song
J. Vis. Commun. Image Represent.1
2023 Feature Differentiation Reconstruction Network for Weakly-Supervised Video Anomaly Detection
abstract
Recent research into video anomaly detection under weakly supervised settings has made significant progress in identifying anomalies with only coarse-grained annotations. Mainstream weakly supervised methods improve detection performance by generating high-quality pseudo labels for video segments. However, these pseudo-label-based methods have been ordinarily hindered by manually-set constraint rules as the bottleneck. In this paper, we propose the Feature Differentiation Reconstruction Network (FDR-Net), which no longer relies on pseudo labels and instead uses a differential reconstruction strategy to improve the discriminability of the representation. Concretely, video features are first randomly masked out and then reconstructed with distinct targets for normal and abnormal videos during the differential reconstruction process. Besides, we also introduce a dense transformer-based encoder to refine spatial-temporal relationships among video segments. Comprehensive experiments on ShanghaiTech demonstrate the superior performance of our model.
Yiling Gong, Sihui Luo 0001, Chong Wang 0001
IEEE Signal Process. Lett.2
2021 Progressive Network Grafting for Few-Shot Knowledge Distillation
abstract
Knowledge distillation has demonstrated encouraging performances in deep model compression. Most existing approaches, however, require massive labeled data to accomplish the knowledge transfer, making the model compression a cumbersome and costly process. In this paper, we investigate the practical few-shot knowledge distillation scenario, where we assume only a few samples without human annotations are available for each category. To this end, we introduce a principled dual-stage distillation scheme tailored for few-shot data. In the first step, we graft the student blocks one by one onto the teacher, and learn the parameters of the grafted block intertwined with those of the other teacher blocks. In the second step, the trained student blocks are progressively connected and then together grafted onto the teacher network, allowing the learned student blocks to adapt themselves to each other and eventually replace the teacher network. Experiments demonstrate that our approach, with only a few unlabeled samples, achieves gratifying results on CIFAR10,CIFAR100, and ILSVRC-2012. On CIFAR10 and CIFAR100, our performances are even on par with those of knowledge distillation schemes that utilize the full datasets. The source code is available at https://github.com/zju-vipa/NetGraft.
Chengchao Shen, Xinchao Wang, Youtan Yin, Jie Song 0011, Sihui Luo 0001, Mingli Song
AAAI5
2020 Collaboration by Competition: Self-coordinated Knowledge Amalgamation for Multi-talent Student Learning
Sihui Luo 0001, Wenwen Pan 0003, Xinchao Wang, Dazhou Wang, Haihong Tang, Mingli Song
ECCV (6)1
2019 Knowledge Amalgamation from Heterogeneous Networks by Common Feature Learning
abstract
An increasing number of well-trained deep networks have been released online by researchers and developers, enabling the community to reuse them in a plug-and-play way without accessing the training annotations. However, due to the large number of network variants, such public-available trained models are often of different architectures, each of which being tailored for a specific task or dataset. In this paper, we study a deep-model reusing task, where we are given as input pre-trained networks of heterogeneous architectures specializing in distinct tasks, as teacher models. We aim to learn a multitalented and light-weight student model that is able to grasp the integrated knowledge from all such heterogeneous-structure teachers, again without accessing any human annotation. To this end, we propose a common feature learning scheme, in which the features of all teachers are transformed into a common space and the student is enforced to imitate them all so as to amalgamate the intact knowledge. We test the proposed approach on a list of benchmarks and demonstrate that the learned student is able to achieve very promising performance, superior to those of the teachers in their specialized tasks.
Sihui Luo 0001, Xinchao Wang, Gongfan Fang, Dapeng Tao, Mingli Song
IJCAI1
2019 Real-time intelligent big data processing: technology, platform, and applications
Tongya Zheng, Gang Chen 0001, Xinyu Wang 0001, Chun Chen 0001, Xingen Wang, Sihui Luo 0001
Sci. China Inf. Sci.6
2018 DeepSSH: Deep Semantic Structured Hashing for Explainable Person Re-Identification
abstract
For large collections of gallery images captured by sparsely distributed cameras, we often employ hashing based approaches to enhance the efficiency of person re-identification (re-id). However, these hashing based approaches fail to provide semantically explainable encoding in solving the re-id problem, which makes it infeasible to identify the correct matches in the collection by just using a semantic query. To overcome this limitation, we propose a new deep hashing network called Deep Semantic Structured Hashing (DeepSSH) to obtain the semantic structured representation of human. In the proposed DeepSSH framework, both the mid-level human attributes and the high-level ID labels are used to learn a deep hashing network. Then, based on the obtained semantic structured hash code and the attribute labels, we learn a decoder to find the partial hash code corresponding to the specified attributes. Finally, a new grain scalable re-id framework is constructed to support semantic query of a person by providing partial or full semantic description of a person instead of the whole photo. Experimental results show that DeepSSH is comparable with state-of-the-art hashing-based person re-id approaches, and the experiment in semantic analysis shows that our hash code owns semantic meaning indeed.
Sihui Luo 0001, Yezhou Yang, Mingli Song
ICIP2
2018 DeepSIC: Deep Semantic Image Compression
Sihui Luo 0001, Yezhou Yang, Yanling Yin, Chengchao Shen, Mingli Song
ICONIP (1)1
2018 Intra-class Structure Aware Networks for Screen Defect Detection
Chengchao Shen, Jie Song 0011, Shuyi Song, Sihui Luo 0001, Mingli Song
ICONIP (4)4
2017 Group Sparse-Based Mid-Level Representation for Action Recognition
abstract
Mid-level parts are shown to be effective for human action recognition in videos. Typically, these semantic parts are first mined with some heuristic rules, then videos are represented via volumetric max-pooling (VMP) method. However, these methods have two issues: 1) the VMP strategy divides videos by static grids. In this case, a semantic part may occur in different localizations in different videos. That means the VMP strategy loses the space-time invariance. To solve this problem, we propose to apply a saliency-driven max-pooling scheme to represent a video. We extract the video semantic cues by the saliency map, and dynamically pool the local maximum responses. This scheme can be considered as a semantic content-based feature alignment method and 2) the parts discovered by heuristic rules may be intuitive but not discriminative enough for action classification because they neglect the relations between the detectors. For this issue, we propose to apply a sparse classifier model to select discriminative parts. Moreover, to further improve the discriminative ability of the representation, we propose to conduct feature selection by the corresponding entry magnitude of the model coefficients. We conduct experiments on four challenging datasets-KTH, Olympic Sports, UCF50, and HMDB51. The results show that the proposed method significantly outperforms the state-of-the-art methods.
Shiwei Zhang 0001, Changxin Gao, Sihui Luo 0001, Nong Sang
IEEE Trans. Syst. Man Cybern. Syst.4