VLDB 2026 Research / reviewers in the wild / expert
Hong Lu 0001
dblp:47/2341-1
· DBLP profile ↗
88ranked-venue papers
9as first author
40since 2021 · last 2026
0000-0002-4572-2854ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?abstractIn multivariate long-term time series forecasting, the success of self-attention is commonly attributed to the attention matrix that encodes token interactions.In this paper, we provide evidence that challenges this view.Through extensive experiments on three classic and three latest Transformer models, we find that dotproduct attention can be replaced by elementwise operations without token interaction, such as the addition and Hadamard product, while maintaining or even improving accuracy.This motivates our central hypothesis: the effectiveness of self-attention in this task arises not from the dynamic attention matrix, but from the multi-branch feature extraction enabled by the parallel Query, Key, and Value projections and their fusion.To validate this hypothesis, we construct a minimalist multi-branch MLP that isolates the 'multi-branch mapping with element-wise operation' structure from the Transformer and show that it achieves competitive performance.Our findings indicate that the source of performance in self-attention is often misinterpreted, as its actual advantage stems from the architectural principle of multi-branch mapping and fusion, rather than the attention matrix. Xinyu Li 0014, Kexi Chen, Jiajie Shen, Ying Zheng 0004, Hong Lu 0001, Jin Zhao 0001, Xin Wang 0002 |
ACL (1) | 5 |
| 2026 | Modeling Point-to-Point Dependency for High-Dimensional Long-Term Series Forecasting
Xinyu Li 0014, Kexi Chen, Ying Zheng 0004, Zhiyi Yao, Yi Xie 0003, Jihan Dai, Lei Bai 0001, Jin Zhao 0001, Jiajie Shen, Yunqi Cai, Hong Lu 0001, Xin Wang 0002 |
WWW | 11 |
| 2026 | Align-then-generate: An effective cross-modal generation paradigm for multi-label zero-shot learning
Peirong Ma, Wu Ran, Yanhui Gu, Huaqiu Chen, Zhiquan He, Hong Lu 0001 |
Pattern Recognit. | 6 |
| 2025 | MoME: Mixture of Multi-Domain Experts for Multivariate Long-Term Series ForecastingabstractTime series forecasting is always important, with multivariate long-term series forecasting being its most challenging task. Here, the existing methods typically learn only in a single domain and focus on optimizing model structures, leading to incomplete information mining and imprecise predictions. To address this, we propose a generalized Mixture of Multi-Domain Experts (MoME) for multivariate long-term series forecasting. Unlike most existing methods, MoME focuses on multi-perspective information mining and fusing. To this end, MoME transforms time series into the frequency and spatial domains to learn their respective representations. MoME regards variates information as embedded features and applies fast Fourier transform to the time dimension. Then it learns embedded features in the frequency domain. In spatial domain learning, MoME applies self-attention mechanism on the variates dimension to efficiently capture dependencies among multiple variates. Finally, MoME fuses the outputs from all domains, reinterprets and integrates information across multiple domains, and predicts future time series. Extensive experiments prove that MoME outperforms state-of-the-art (SOTA) methods. Code is available at: https://github.com/lxy-PhD2022/MoME Xinyu Li 0014, Yunqi Cai, Hong Lu 0001, Xin Wang 0002, Jin Zhao 0001, Fenglin Qi, Jiajie Shen |
ICASSP | 5 |
| 2025 | General Compression Framework for Efficient Transformer Object TrackingabstractPrevious works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker. Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001 |
ICCV | 12 |
| 2025 | HMSformer: Hierarchical Multi-Scale Transformer for Multivariate Long-Term Series ForecastingabstractMulti-scale analysis, a classical and crucial methodology, is extensively employed in multivariate long-term time series forecasting. However, the existing methods struggle to model arbitrary scales, thus limiting their ability to delve into complex patterns, which in turn limits the forecasting precision. Therefore, we propose the HMSformer, a multi-scale Transformer that includes a shifted stacked embedding mechanism and a unified hierarchical framework, solving this problem in terms of the in-depth and comprehensiveness of multi-scale analysis. The mechanism conducts multi-scale modeling by sampling at different positions and sizes. This sampling reflects the arbitrariness of scales and allows for a thorough scan of any scale, supporting a deeper analysis of temporal relationships. Additionally, it converts input series into structured embeddings optimized for Transformer input, allowing the self-attention mechanism to effectively learn intricate cross-scale relationships. Moreover, the unified hierarchical framework integrates multi-scale modeling across different levels of abstraction by unifying global structures, local segments, and fine details into a consistent representation. Together, these two innovations ensure that HMSformer achieves a more comprehensive and in-depth analysis of temporal relationships. HMSformer demonstrates its effectiveness with a significant improvement, achieving an average mean squared error (MSE) that is 9.925% lower than the baseline across nine authoritative datasets. Source code is available at: https://github.com/lxy-PhD2022/HMSformer Xinyu Li 0014, Yunqi Cai, Hong Lu 0001, Xin Wang 0002, Jin Zhao 0001 |
ICME | 6 |
| 2025 | ForeNet: Unlocking Long-Term Series Forecasting in High-Dimensional Scenario via Forest StructureabstractFacing the key challenge in multivariate long-term series prediction, namely effectively modeling the long-term dependencies among high-dimensional variables, we propose an innovative forest network (ForeNet). Firstly, we construct a polytree, which takes variable as leaf node, convolution as edge, and progressive fusion as the root node. Polytree models the dependencies among variables from the bottom up, adopting a local-to-global progressive learning strategy. Moreover, polytree embeds the entire sequence as input channels for convolution, allowing interactions across arbitrary time steps between variables. Thereby, polytree could model long-term dependencies. Then, we construct multiple polytrees with varying branching factors, utilizing the self-attention mechanism to assign weights and combine polytrees into a forest. Forest adopts an ensemble learning strategy to capture complex patterns hidden under high-dimensional variables. Combining progressive and ensemble strategies, ForeNet could effectively address the key challenge mentioned at the beginning. Extensive experiments show that ForeNet could reduce the average MSE by up to 11.53% compared to the baseline, which verifies the effectiveness of ForeNet. Source code is available at: https://github.com/lxy-PhD2022/ForeNet Xinyu Li 0014, Hongxiang Zhou, Hong Lu 0001, Xin Wang 0002, Jin Zhao 0001 |
ICME | 5 |
| 2025 | Implicit Retinex Decomposition with Chromaticity Disentanglement for Low-Light Image Enhancement
Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma |
ACM Multimedia | 5 |
| 2025 | Low-Light Image Enhancement via Multi-Exposure Progressive Contrastive RegularizationabstractLow-light image enhancement (LLIE) aims to restore low-light images to their normal-light counterparts with optimal global illumination distribution and clear local details. With the advancement of deep learning, deep learning-based methods have become the mainstream in the LLIE community. However, most deep learning-based method cannot yet fully exploit the global and local contextual information in the low-light image. In this paper, we introduce a dual-branch module to simultaneously restore global and local features from spatial and frequency domain. To fuse these multi-level features, we propose a perception module to perform feature interaction between global and local features via cross attention and self-gating. By integrating the two developed modules into a U-Net backbone, we present a global-local interaction network for LLIE. Furthermore, recent studies have shown that contrastive learning can be an effective paradigm for the LLIE task. However, previous works typically use semantically-inconsistent under-/over-exposed images as negative samples. These images are very dissimilar to the ground-truth and cannot provide sufficient regularization in contrastive learning. To address this limitation, we explore a practical multi-exposure progressive contrastive regularization framework for LLIE. With a customized sample generation, sample selection, and progressive learning strategy, our proposed framework progressively narrows down the solution space around the optimum, and helps to improve the performance of LLIE methods without additional inference overhead. Combining the proposed network and contrastive regularization, our proposed method achieves favorable results compared to state-of-the-art LLIE methods on benchmark datasets. Extensive experiments further demonstrate the generalization ability of our proposed method. Zuojie Xie, Hao Ren 0002, Junjian Huang, Zhiquan He, Hong Lu 0001, Lvfan Yuan, Changyong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Unleashing the Potential of Hierarchical Region Clues for Open-Vocabulary Multi-Label ClassificationabstractOpen-vocabulary multi-label classification (OV-MLC) aims to leverage the rich multi-modal knowledge from Vision-language pre-training (VLP) models to further improve the recognition ability for unseen (novel) classes beyond the training set in multi-label scenarios. Existing OV-MLC methods only perform predictions on single hierarchical regions, and aggregate the prediction scores of these regions through simpletop-kmean pooling. This fails to unleash the potential of rich hierarchical region clues in multi-label images and does not fully exploit the discriminative information from all regions in the image, resulting in sub-optimal performance. In this work, we propose a novel OV-MLC framework to fully harness the power of multiple hierarchical region clues. Specifically, we first design a hierarchical clue gathering (HCG) module to gather different hierarchical clues, enabling more precise recognition of multiple object categories with different sizes in a multi-label image. Then, by viewing multi-label classification as single-label classification of each region within the image, we present a novel hierarchical score aggregation (HSA) approach, thereby better utilizing the predictions of each image region for each class. We also utilize a well-designed region selection strategy (RSS) to eliminate noise or background regions in an image that are irrelevant to classification, achieving higher multi-label classification accuracy. In addition, we propose a hybrid prompt learning (HPL) strategy to enhance visual-semantic consistency while preserving the generalization capability of label embeddings for unseen classes. Extensive experiments on public benchmark datasets demonstrate that our method significantly outperforms the current state-of-the-art. Peirong Ma, Wu Ran, Zhiquan He, Jian Pu, Hong Lu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate RainsabstractRecent advances in image deraining have focused on training powerful models on mixed multiple datasets comprising diverse rain types and backgrounds. However, this approach tends to overlook the inherent differences among rainy images, leading to suboptimal results. To overcome this limitation, we focus on addressing various rainy images by delving into meaningful representations that encapsulate both the rain and background components. Leveraging these representations as instructive guidance, we put forth a Context-based Instance-level Modulation (CoI-M) mechanism adept at efficiently modulating CNN- or Transformer-based models. Furthermore, we devise a rain-/detail-aware contrastive learning strategy to help extract joint rain-/detail-aware representations. By integrating CoI-M with the rain-/detail-aware Contrastive learning, we develop [CoIC](https://github.com/Schizophreni/CoIC), an innovative and potent algorithm tailored for training models on mixed datasets. Moreover, CoIC offers insight into modeling relationships of datasets, quantitatively assessing the impact of rain and details on restoration, and unveiling distinct behaviors of models given diverse inputs. Extensive experiments validate the efficacy of CoIC in boosting the deraining ability of CNN and Transformer models. CoIC also enhances the deraining prowess remarkably when real-world dataset is included. Wu Ran, Peirong Ma, Zhiquan He, Hao Ren 0002, Hong Lu 0001 |
ICLR | 5 |
| 2024 | Rainmer: Learning Multi-view Representations for Comprehensive Image Deraining and BeyondabstractWe address image deraining under complex backgrounds, diverse rain scenarios, and varying illumination conditions, representing a highly practical and challenging problem. Our approach utilizes synthetic, real-world, and nighttime datasets, wherein rich backgrounds, multiple degradation types, and diverse illumination conditions coexist. The primary challenge in training models on these datasets arises from the discrepancies among them, potentially leading to conflicts or competition during the training period. To address this issue, we first align the distribution of synthetic, real-world and nighttime datasets. Then we propose a novel contrastive learning strategy to extract multi-view (multiple) representations that effectively capture image details, degradations, and illuminations, thereby facilitating training across all datasets. Regarding multiple representations as profitable prompts for deraining, we devise a prompting strategy to integrate them into the decoding process. This contributes to a potent deraining model, dubbed Rainmer. Additionally, a spatial-channel interaction module is introduced to fully exploit cues when extracting multi-view representations. Extensive experiments on synthetic, real-world, and nighttime datasets demonstrate that Rainmer outperforms current representative methods. Moreover, Rainmer achieves superior performance on the All-in-One image restoration dataset, underscoring its effectiveness. Furthermore, quantitative results reveal that Rainmer significantly improves object detection performance on both daytime and nighttime rainy datasets. These observations substantiate the potential of Rainmer for practical applications. Wu Ran, Peirong Ma, Zhiquan He, Hong Lu 0001 |
ACM Multimedia | 4 |
| 2024 | Feature decoupling and reorganization network for single image deraining
Yunrui Cheng, Junjian Huang, Hao Ren 0002, Wu Ran, Hong Lu 0001 |
Multim. Syst. | 5 |
| 2024 | Context-aware coarse-to-fine network for single image desnowing
Yunrui Cheng, Hao Ren 0002, Hong Lu 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Low-Light Image Enhancement With Multi-Scale Attention and Frequency-Domain OptimizationabstractLow-light image enhancement aims to improve the perceptual quality of images captured in conditions of insufficient illumination. However, such images are often characterized by low visibility and noise, making the task challenging. Recently, significant progress has been made using deep learning-based approaches. Nonetheless, existing methods encounter difficulties in balancing global and local illumination enhancement and may fail to suppress noise in complex lighting conditions. To address these issues, we first propose a multi-scale illumination adjustment network to balance both global illumination and local contrast. Furthermore, to effectively suppress noise potentially amplified by the illumination adjustment, we introduce a wavelet-based attention network that efficiently perceives and removes noise in the frequency domain. We additionally incorporate a discrete wavelet transform loss to supervise the training process. Particularly, the proposed wavelet-based attention network has been shown to enhance the performance of existing low-light image enhancement methods. This observation indicates that the proposed wavelet-based attention network can be flexibly adapted to current approaches to yield superior enhancement results. Furthermore, extensive experiments conducted on benchmark datasets and downstream object detection task demonstrate that our proposed method achieves state-of-the-art performance and generalization ability. Zhiquan He, Wu Ran, Kehua Li, Chang-Yong Xie, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | A Transferable Generative Framework for Multi-Label Zero-Shot LearningabstractMulti-label zero-shot learning (MLZSL) is a more realistic and challenging task than single-label zero-shot learning (SLZSL), which aims to recognize multiple unseen classes in a single image. To adapt generative models to the MLZSL task and better recognize multiple unseen object categories in an image, this paper proposes a Transferable Generative Framework (TGF), which consists of a Multi-Label Semantic Embedding Autoencoders (SEAs), a Semantic-Related Multi-Label Feature Transformation Network (FTN) and a Multi-Label Feature Generation Networks (FGNs). First, SEAs adaptively encodes the class-level word vectors corresponding to each sample containing different number of classes into sample-level semantic embeddings with the same dimension. Then, FTN transforms global features extracted by a CNN pre-trained on single-label images into features that are semantic-related and more suitable for multi-label classification. Finally, FGNs generates both global and local features to better recognize the dominant and minor object categories in a multi-label image, respectively. Extensive experiments on three benchmark datasets show that TGF significantly outperforms state-of-the-arts. Specifically, compared with the previous best generative MLZSL method (i.e., Gen-MLZSL), TGF improves the mAP of the ZSL (GZSL) task by 5.4% (6.9%), 20.5% (27.9%), and 2.4% (3.9%) on NUS-WIDE, Open Images, and MS-COCO datasets, respectively. Peirong Ma, Zhiquan He, Wu Ran, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Fully Unsupervised Domain-Agnostic Image RetrievalabstractRecent research in cross-domain image retrieval has focused on addressing two challenging issues: handling domain variations in the data and dealing with the lack of sufficient training labels. However, these problems have often been studied separately, limiting the practicality and significance of the research outcomes. The existing cross-domain setting is also restricted to cases where domain labels are known during training, and all samples have semantic category information or instance correspondences. In this paper, we propose a novel approach to address a more general and practical problem:fully unsupervised domain-agnostic image retrievalunder the domain-unknown setting, where no annotations are provided. Our approach tackles both thedomain variationandmissing labelschallenges simultaneously. We introduce a new fully unsupervised One-Shot Synthesis-based Contrastive learning method (termed OSSCo) to project images from different data distributions into a shared feature space for similarity measurement. To handle the domain-unknown setting, we propose One-Shot unpaired image-to-image Translation (OST) between a randomly selected one-shot image and the rest of the training images. By minimizing the global distance between the original images and the generated images from OST, the model learns domain-agnostic representations. To address the label-unknown setting, we employ contrastive learning with a synthesis-based transform module from the OST training. This allows for effective representation learning without any annotations or external constraints. We evaluate our proposed method on diverse datasets, and the results demonstrate its effectiveness. Notably, our approach achieves comparable performance to current state-of-the-art supervised methods. Ziqiang Zheng, Hao Ren 0002, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | VPL-SLAM: A Vertical Line Supported Point Line Monocular SLAM SystemabstractTraditional monocular visual simultaneous localization and mapping (SLAM) systems rely on point features or line features to estimate and optimize the camera trajectory and build a map of the surrounding environment. However, in complex scenarios such as underground parking, the performance of traditional point-line SLAM systems tends to degrade due to mirror reflection, illumination change, poor texture, and other interference. This paper proposes VPL-SLAM, a structural vertical line supported point-line monocular SLAM system that works well in complex environments such as underground parking or campus. The proposed system leverages structural vertical lines at all instances of the process. With the assistance of the structural vertical lines and global vertical direction, our system can output a more accurate visual odometry result. Furthermore, the resulting map of our system is a more reasonable structural line feature map than the previous point-line-based monocular SLAM systems. Our system has been tested with the popular autonomous driving dataset Kitti Odometry. In addition, to fully test the proposed SLAM system, we also test our system using a self-collected dataset, including underground parking and campus scenarios. As a result, our proposal reveals a more accurate navigation result and a more reasonable structural resulting map compared to state-of-the-art point-line SLAM systems such as Structure PLP-SLAM. Qi Chen 0025, Yu Cao 0024, Guanghao Li 0001, Shoumeng Qiu, Xiangyang Xue 0001, Hong Lu 0001, Jian Pu |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | Real-Time Attentive Dilated U-Net for Extremely Dark Image EnhancementabstractImages taken under low-light conditions suffer from poor visibility, color distortion, and graininess, all of which degrade the image quality and hamper the performance of downstream vision tasks, such as object detection and instance segmentation in the field of autonomous driving, making low-light enhancement an indispensable basic component of high-level visual tasks. Low-light enhancement aims to mitigate these issues, and has garnered extensive attention and research over several decades. The primary challenge in low-light image enhancement arises from the low signal-to-noise ratio caused by insufficient lighting. This challenge becomes even more pronounced in near-zero lux conditions, where noise overwhelms the available image information. Both traditional image signal processing pipeline and conventional low-light image enhancement methods struggle in such scenarios. Recently, deep neural networks have been used to address this challenge. These networks take unmodified RAW images as input and produce the enhanced sRGB images, forming a deep learning based image signal processing pipeline. However, most of these networks are computationally expensive and thus far from practical use. In this article, we propose a lightweight model called attentive dilated U-Net (ADU-Net) to tackle this issue. Our model incorporates several innovative designs, including an asymmetric U-shape architecture, dilated residual modules for feature extraction, and attentive fusion modules for feature fusion. The dilated residual modules provide strong representative capability, whereas the attentive fusion modules effectively leverage low-level texture information and high-level semantic information within the network. Both modules employ a lightweight design but offer significant performance gains. Extensive experiments demonstrate that our method is highly effective, achieving an excellent balance between image quality and computational complexity—that is, taking less than 4ms for a high-definition 4K image on a single GTX 1080Ti GPU and yet maintaining competitive visual quality. Furthermore, our method exhibits pleasing scalability and generalizability, highlighting its potential for widespread applicability. Junjian Huang, Hao Ren 0002, Chuanlu Lv, Changyong Xie, Hong Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2023 | Instance-Aware Diffusion Implicit Process for Box-Based Instance SegmentationabstractThe diffusion model has demonstrated impressive performance in image generation, but its potential for discriminative tasks such as instance segmentation remains unexplored. In this paper, we propose an Instance-aware Diffusion Implicit Process (IDIP) framework for instance segmentation based on boxes. During training, IDIP diffuses ground-truth boxes across various time steps, extracting corresponding Region of Interest (RoI) features. Dynamic convolution is then used to predict boxes and categories for each RoI, and the mask head generates masks from these predictions. During inference, IDIP iteratively refines randomly generated boxes with the denoising diffusion implicit model, while the mask head derives final masks from RoIs based on the refined boxes. Our method surpasses existing approaches on the COCO benchmark, requiring fewer training steps and less memory resources due to its dynamic design and instance-aware characteristic. Hao Ren 0002, Xingsong Liu, Junjian Huang, Ru Wan, Jian Pu, Hong Lu 0001 |
ECAI | 6 |
| 2023 | GSFormer: Geometric-Spatial Transformer on Point Cloud CompletionabstractPoint cloud completion aims to complete the objects’ shape from incomplete 3D objects. Most works based on encoder-decoder lose the local geometric details of global features when encoding partial points. Besides, the decoder lacks the exploration of the correlation between global and local features. To solve these problems, 1) we propose a Geometric Transformer to learn the global and local geometric details of incomplete point clouds by learning their shape prior geometric information in the encoder, which is beneficial to generate geometric keypoints. The generated geometric keypoints contain the global structure information and local geometric details of the complete point cloud. 2) We propose a Spatial Transformer in the decoder, which can adaptively select neighborhood features to learn the long-distance geometric relationship between upsampling points. Experimental results show that our method achieves better performance on PCN and Shapenet-55/34 datasets. Yijun Long, Zhaoyu Chen 0001, Hong Lu 0001 |
ICME | 3 |
| 2023 | Weakly-supervised Temporal Action Localization with Adaptive Clustering and Refining NetworkabstractWeakly-supervised temporal action localization task aims to localize temporal boundaries of action instances by using only video-level labels. Existing methods primarily adopt Multi-Instance-Learning (MIL) scheme to handle this task. The effectiveness of MIL scheme depends heavily on the selection of top-k action snippets, which is unstable and requires manual tuning. To address these deficiencies, we propose an Adaptive Clustering and Refining Network (ACRNet). Specifically, we present an action-aware clustering strategy that is adaptable and requires no manual tuning to separate action and background snippets of diverse videos based on intra-class activation distribution. And a cluster refining step is included to eliminate false action snippets by considering inter-class activation distribution, which greatly improves robustness and localization accuracy. Extensive experiments on THUMOS14, ActivityNet 1.2&1.3 benchmarks show that our method achieves state-of-the-art performance. Hao Ren 0002, Wu Ran, Xingson Liu, Hong Lu 0001, Cheng Jin 0001 |
ICME | 5 |
| 2023 | SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (UVOS) aims at detecting the primary objects in a given video sequence without any human interposing. Most existing methods rely on two-stream architectures that separately encode the appearance and motion information before fusing them to identify the target and generate object masks. However, this pipeline is computationally expensive and can lead to suboptimal performance due to the difficulty of fusing the two modalities properly. In this paper, we propose a novel UVOS model called SimulFlow that simultaneously performs feature extraction and target identification, enabling efficient and effective unsupervised video object segmentation. Concretely, we design a novel SimulFlow Attention mechanism to bridege the image and motion by utilizing the flexibility of attention operation, where coarse masks predicted from fused feature at each stage are used to constrain the attention operation within the mask area and exclude the impact of noise. Because of the bidirectional information flow between visual and optical flow features in SimulFlow Attention, no extra hand-designed fusing module is required and we only adopt a light decoder to obtain the final prediction. We evaluate our method on several benchmark datasets and achieve state-of-the-art results. Our proposed approach not only outperforms existing methods but also addresses the computational complexity and fusion difficulties caused by two-stream architectures. Our models achieve 87.4 ℐ&F on DAVIS-16 with the highest speed (63.7 FPS on a 3090) and the lowest parameters (13.7 M). Our SimulFlow also obtains competitive results on video salient object detection datasets. Lingyi Hong, Wei Zhang 0016, Shuyong Gao, Hong Lu 0001 |
ACM Multimedia | 4 |
| 2023 | Weakly-Supervised Temporal Action Localization with Regional Similarity Consistency
Hao Ren 0002, Hong Lu 0001, Cheng Jin 0001 |
MMM (1) | 3 |
| 2023 | DaCo: domain-agnostic contrastive learning for visual place recognition
Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001 |
Appl. Intell. | 4 |
| 2023 | ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image RetrievalabstractThe huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset). Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | TRNR: Task-Driven Image Rain and Noise Removal With a Few Images Based on Patch AnalysisabstractThe recent success of learning-based image rain and noise removal can be attributed primarily to well-designed neural network architectures and large labeled datasets. However, we discover that current image rain and noise removal methods result in low utilization of images. To alleviate the reliance of deep models on large labeled datasets, we propose the task-driven image rain and noise removal (TRNR) based on a patch analysis strategy. The patch analysis strategy samples image patches with various spatial and statistical properties for training and can increase image utilization. Furthermore, the patch analysis strategy encourages us to introduce the N-frequency-K-shot learning task for the task-driven approach TRNR. TRNR allows neural networks to learn from numerous N-frequency-K-shot learning tasks, rather than from a large amount of data. To verify the effectiveness of TRNR, we build a Multi-Scale Residual Network (MSResNet) for both image rain removal and Gaussian noise removal. Specifically, we train MSResNet for image rain removal and noise removal with a few images (for example, 20.0% train-set of Rain100H). Experimental results demonstrate that TRNR enables MSResNet to learn more effectively when data is scarce. TRNR has also been shown in experiments to improve the performance of existing methods. Furthermore, MSResNet trained with a few images using TRNR outperforms most recent deep learning methods trained data-driven on large labeled datasets. These experimental results have confirmed the effectiveness and superiority of the proposed TRNR. The source code is available on https://github.com/Schizophreni/MSResNet-TRNR. Wu Ran, Bohong Yang, Peirong Ma, Hong Lu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | 3DCNN-Based Palpation Localization with Temporal Attention ModuleabstractPalpation is necessary in Traditional Chinese Medicine (TCM). In TCM, doctors need to touch patient’s wrist and exert pressure on it to obtain physiological signal. Locating the pulse position is an important step in palpation. Nowadays, researchers have proposed numerous methods to locate the pulse position accurately. In this paper, we propose an accurate and effective framework to locate the pulse position. Our framework is based on 3-dimensional convolutional neural network (3DCNN), and we propose a novel temporal attention module to further improve the performance. Our framework achieves superior results, the accuracy of locating the pulse position on a video with a resolution of 2048 × 1088 within 100 pixels is 97%. More and more cross-domain applications of computer science and medical science provides more ways for doctors to treat patients, such as telemedicine, computed tomography image analysis, and physiological signal analysis. Our research presents the great potential to be applied to palpation automation and intelligent medical treatment. Guanhao Huang, Hong Lu 0001, Jingjing Luo |
ICIP | 2 |
| 2022 | Semantic-Related Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) is a challenging task. Although no visual samples of unseen classes are provided during training, the classifier must learn to recognize all classes (i.e. both seen and unseen classes). Due to the ability to generate unseen classes samples, generative models have been widely used in GZSL. However, these generative models only learn from the seen classes, so the discriminability of the unseen class features they generate is usually poor, resulting in low unseen class classification accuracy. To solve this problem, this paper proposes a novel semantic-related feature generative (SRFG) model to improve visual-semantic consistency and alleviate seen-unseen bias effectively. SRFG can generate any number of semantic-related discriminative features for both seen and unseen classes. Extensive experiments on four benchmark datasets show that the proposed model significantly outperforms the state of the arts. Peirong Ma, Wu Ran, Hong Lu 0001 |
ICME | 3 |
| 2022 | Hybrid Uncalibrated Near-light Photometric Stereo in Realistic EnvironmentabstractPhotometric stereo aims at recovering the surface and shape of an object from a set of observations under different light conditions. Deep learning based methods have made substantial contributions to the surface normal estimation under complex surface reflections, various materials, near-light settings, etc. These deep learning based methods learn the mapping from observed images to the surface normal of an object directly via training deep neural networks on large labeled datasets. However, the shapes, materials, and reflectance properties of objects in these datasets are limited, leading to abridged performances in realistic environments. In this paper, we introduce a work-piece dataset for near-light photometric stereo under industrial application scenarios, which consists of observed images taken under at most 40 light conditions, and the ground truth surface normals. Based on this datasets, we propose the Hybrid Near-light Uncalibrated Photometric Stereo (HNUPS) for both unsupervised light calibration and surface normal estimation. Experimental results on the work-piece dataset demonstrate that HNUPS can obtain the least mean angular error when compared to recent photometric stereo methods, which have verified the effectiveness of the proposed HNUPS. Wu Ran, Xingsong Liu, Hong Lu 0001, Bohong Yang, Jingjing Luo |
ISCAS | 4 |
| 2022 | RARN: A Real-Time Skeleton-based Action Recognition Network for Auxiliary Rehabilitation TherapyabstractRehabilitation gymnastic training is an effective therapy for degeneration of spine disease in Traditional Chinese Medicine (TCM). In this paper, we propose a lightweight real-time Rehabilitation Action Recognition Network (RARN) using skeleton sequence obtained through 2d pose estimation and build an intelligent auxiliary rehabilitation therapy system. We design a set of training exercises consisting of 8 actions, and construct a dataset called Rehabilitation Action for Degenerative Spine Diseases (RDSD), containing 1012 skeleton sequences. We describe a demo application to conduct real-time action evaluation for rehabilitation therapy. The experimental results on RDSD shows that our system achieves high accuracy while still working under 6ms per frame averagely and the inference costs only about 0. 2ms per frame. Mengqi Shen, Hong Lu 0001 |
ISCAS | 2 |
| 2022 | SFCN: Spoon Fully Convolutional Networks for Pulse LocalizationabstractPulse localization is the basic task of the pulse diagnosis with the robot. Using neural network for localization can not only reduce the contact between the machine and the subject, relieve the discomfort of the process, but also reduce the preparation. Since the networks with the coordinate regression directly have large parameters and are not suited for input images with different size. In this paper, we propose a novel method, spoon fully convolutional networks (SFCN) with the landmark fitting method for pulse localization. SFCN includes the fully convolutional networks which are like the spoon while the landmark fitting method finds the pulse in the sub-pixels. The experiments show that our proposed method can locate the pulse with high accuracy and few parameters, which is suitable for application on robots. Bohong Yang, Hong Lu 0001, Jingjing Luo |
ISCAS | 3 |
| 2022 | Weakly-Supervised Temporal Action Localization with Multi-Head Cross-Modal Attention
Hao Ren 0002, Wu Ran, Hong Lu 0001, Cheng Jin 0001 |
PRICAI (3) | 4 |
| 2022 | Energy-Guided Feature Fusion for Zero-Shot Sketch-Based Image Retrieval
Hao Ren 0002, Ziqiang Zheng, Hong Lu 0001 |
Neural Process. Lett. | 3 |
| 2022 | GAN-MVAE: A discriminative latent feature generation framework for generalized zero-shot learning
Peirong Ma, Hong Lu 0001, Bohong Yang, Wu Ran |
Pattern Recognit. Lett. | 2 |
| 2022 | Compositional coding capsule network with k-means routing for text classification
Hao Ren 0002, Hong Lu 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Learning Cognitive Map Representations for Navigation by Sensory-Motor IntegrationabstractHow to transform a mixed flow of sensory and motor information into memory state of self-location and to build map representations of the environment are central questions in the navigation research. Studies in neuroscience have shown that place cells in the hippocampus of the rodent brains form dynamic cognitive representations of locations in the environment. We propose a neural-network model called sensory-motor integration network model (SeMINet) to learn cognitive map representations by integrating sensory and motor information while an agent is exploring a virtual environment. This biologically inspired model consists of a deep neural network representing visual features of the environment, a recurrent network of place units encoding spatial information by sensorimotor integration, and a secondary network to decode the locations of the agent from spatial representations. The recurrent connections between the place units sustain an activity bump in the network without the need of sensory inputs, and the asymmetry in the connections propagates the activity bump in the network, forming a dynamic memory state which matches the motion of the agent. A competitive learning process establishes the association between the sensory representations and the memory state of the place units, and is able to correct the cumulative path-integration errors. The simulation results demonstrate that the network forms neural codes that convey location information of the agent independent of its head direction. The decoding network reliably predicts the location even when the movement is subject to noise. The proposed SeMINet thus provides a brain-inspired neural-network model for cognitive map updated by both self-motion cues and visual cues. Dongye Zhao, Zheng Zhang 0001, Hong Lu 0001, Sen Cheng, Bailu Si, Xisheng Feng |
IEEE Trans. Cybern. | 3 |
| 2022 | Multi-Classes and Motion Properties for Concurrent Visual SLAM in Dynamic EnvironmentsabstractWorking in a dynamic environment is a challenging problem for visual simultaneous localization and mapping (visual SLAM). Most of the existing visual SLAM algorithms fail resulting in significant error or losing in tracking when moving objects dominate the scene. We found two reasons cause these issues: (i) Previous approaches use information from all regions in the image; (ii) Existing algorithms use just two groups and block all feature points from moveable objects. In this paper, we propose a novel Multi-classes and motion properties for Concurrent Visual SLAM (MCV-SLAM) algorithm, which defines classes into five categories and concurrently fuses prior knowledge and observation of moving objects with semantic segmentation to ensure visual SLAM works properly for dynamic environments in real time. We also propose an adaptive method to optimize camera pose by using more potential inlier feature points with continuous weights, while eliminating the impact of moving objects. Our experiments are performed on public datasets of both indoor and outdoor scenes with moving objects in dynamic environments. The experimental results demonstrate that our method outperforms previous works with greater robustness and smaller tracking errors, and our MCV-SLAM can deal with the situations (i.e., the dominance of moving objects, lack of matching points), which lead misestimating occurs in existing SLAMs. Bohong Yang, Wu Ran, Lin Wang 0033, Hong Lu 0001, Yi-Ping Phoebe Chen |
IEEE Trans. Multim. | 4 |
| 2021 | A Fast and Efficient Network for Single Image DerainingabstractRain streaks will degrade the visibility of images. To tackle this problem, we propose a novel Adaptive Dilated Network (ADN) to remove rain streaks from a single image while using less parameters and running faster than previous methods. Specifically, an Adaptive Dilated Block (ADB) is constructed as the sub-module of ADN. In ADB, we apply a shared dilated block to extract multi-scale features. Then a dilated selection block is added to leverage the importance of features in different scales. All the multi-scale features are fused together to obtain features with rich rain details. To further model the inter-dependencies of the fused features, a feature selection block is employed in ADB to assign different weights to each feature. Moreover, all the hierarchical features extracted by each ADB are concatenated together and fed into a rainy map generator to estimate rain layer. Experimental results demonstrate that the proposed method is superior to the state-of-the-art methods on performances and running time while using less parameters. The source code is available at https://github.com/nnUyi/ADN. Youzhao Yang, Hong Lu 0001 |
ICASSP | 2 |
| 2021 | Multi-Directional Convolution Networks with Spatial-Temporal Feature Pyramid Module for Action RecognitionabstractRecent attempts show that factorizing 3D convolutional filters into separate spatial and temporal components brings impressive improvement in action recognition. However, traditional temporal convolution operating along the temporal dimension will aggregate unrelated features, since the feature maps of fast-moving objects have shifted spatial positions. In this paper, we propose a novel and effective Multi-Directional Convolution (MDConv), which extracts features along different spatial-temporal orientations. Especially, MDConv has the same FLOPs and parameters as the traditional 1D temporal convolution. Also, we propose the Spatial-Temporal Feature Pyramid Module (STFPM) to fuse spatial semantics in different scales in a light-weight way. Our extensive experiments show that the models which integrate with MDConv achieve better accuracy on several large-scale action recognition benchmarks such as Kinetics, AVA and Something-Something V1&V2 datasets. Bohong Yang, Wu Ran, Hong Lu 0001, Yi-Ping Phoebe Chen |
ICASSP | 4 |
| 2020 | Single Image Rain Removal Boosting Via Directional GradientabstractImage rain removal has been widely studied with traditional methods and learning based methods for years. However, traditional methods like Gaussian mixture model and dictionary learning methods are time consuming and fail to well tackle images with heavy rain streaks since image patches are severely contaminated. By considering the line-like property and angle distribution of rain streaks, this problem can be well solved. In this paper, by introducing Directional b radient operator of arbitrary direction, we propose an efficient and robust Constraints based Model (DiG-CoM) for single image rain removal. Moreover, a density metric of rain streaks is applied to generalize the proposed model to light and heavy rain streak occasions. Extensive experiments on synthetic datasets demonstrate that the proposed model outperforms GMM and JCAS while requiring less time. Furthermore, on real-world occasions, the proposed method obtains better generalization ability compared with the stateof-the-art learning based methods. The source code is available at https://github.com/Schizophreni/Set-vanish-to-the-rain. Wu Ran, Youzhao Yang, Hong Lu 0001 |
ICME | 3 |
| 2020 | Rddan: A Residual Dense Dilated Aggregated Network For Single Image DerainingabstractRainy images contain rain streaks with different sizes, shapes, directions, and densities. To efficiently remove rain streaks from rainy images, it is necessary to capture rich rain details. In this paper, we propose a Residual Dense Dilated Aggregated Network (RDDAN) to focus on different types of rain steaks and efficiently model rain distribution from rainy images. Specifically, a Residual Dense Dilated Aggregated Block (RDDAB) is constructed to fully extract and exploit rain details hierarchically. In RDDAB, dilated aggregated module is applied to capture multi-scale rain details, dense connection is employed to fully exploit hierarchical features extracted by dilated aggregated module, and residual connection is introduced to keep flow of rain details among different blocks. Besides, all the features extracted by each RDDAB are fused progressively which allows the network to adaptively focus on significant hierarchical features inter blocks. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on synthetic and real-world datasets. The source code is available at https://github.com/nnUyi/RDDAN. Youzhao Yang, Wu Ran, Hong Lu 0001 |
ICME | 3 |
| 2020 | Pulse localization networks with infrared cameraabstractPulse localization is the basic task of the pulse diagnosis with robot. More accurate location can reduce the misdiagnosis caused by different types of pulse. Traditional works usually use a collection surface with a certain area for contact detection, and move the collection surface to collect changes of power for pulse localization. These methods often require the subjects place their wrist in a given position. In this paper, we propose a novel pulse localization method which uses the infrared camera as the input sensor, and locates the pulse on wrist with the neural network. This method can not only reduce the contact between the machine and the subject, reduce the discomfort of the process, but also reduce the preparation time for the test, which can improve the detection efficiency. The experiments show that our proposed method can locate the pulse with high accuracy. And we have applied this method to pulse diagnosis robot for pulse data collection. Bohong Yang, Hong Lu 0001, Xinyao Nie, Guanhao Huang, Jingjing Luo |
MMAsia | 3 |
| 2020 | In-classroom learning analytics based on student behavior, topic and teaching characteristic mining
Bohong Yang, Zeping Yao, Hong Lu 0001, Yaqian Zhou 0001, Jinkai Xu |
Pattern Recognit. Lett. | 3 |
| 2020 | Learning the Game of Go by Scalable Network Without Prior Knowledge of KomiabstractAlphaGo trains a value network to predict the win rate of the current state with 7.5 komi on a 19 × 19 board. The komi of most rectangular boards is unknown, so we do not know who the winner is at the end of the game. We need to use the human experience to guess a komi and then train the value network with this komi. Therefore, the accuracy of the value network is related to the accuracy of the guess. This article uses the board value network to calculate the score of the current state and tries to maximize the score. Then, we do not need to guess the komi. We also modify the network structure to support the board with arbitrary board size as input. Furthermore, we can transfer knowledge of the small board to the large board. We propose an algorithm that can adapt to the bonus rule. We have experimentally proved that our method is effective on a small board and has the ability to transfer knowledge to the large board. In order to better understand the learning process, we visualize the policy and score of some major branches. Finally, we show the solution that our program obtained on 6 × 6, 6 × 7, and 7 × 8 boards. Bohong Yang, Lin Wang 0033, Hong Lu 0001, Youzhao Yang |
IEEE Trans. Games | 3 |
| 2019 | Single Image Deraining using a Recurrent Multi-scale Aggregation and Enhancement NetworkabstractSingle image deraining is an ill-posed inverse problem due to the presence of non-uniform rain shapes, directions, and densities in images. In this paper, we propose a novel progressive single image deraining method named Recurrent Multi-scale Aggregation and Enhancement Network (ReMAEN). Differing from previous methods, ReMAEN contains a symmetric structure where recurrent blocks with shared channel attention are applied to select useful information collaboratively and remove rain streaks stage by stage. In ReMAEN, a Multi-scale Aggregation and Enhancement Block (MAEB) is constructed to detect multi-scale rain details. Moreover, to better leverage the rain details from rainy images, ReMAEN enables a symmetric skipping connection from low level to high level. Extensive experiments on synthetic and real-world datasets demonstrate that our method outperforms the state-of-the-art methods tremendously. The source code is available at https://github.com/nnUyi/ReMAEN. Youzhao Yang, Hong Lu 0001 |
ICME | 2 |
| 2019 | Weakly Supervised Image Retrieval via Coarse-scale Feature Fusion and Multi-level Attention BlocksabstractIn this paper, we propose an end-to-end Attention-Block network for image retrieval (ABIR), which greatly increases the retrieval accuracy without human annotations like bounding boxes. Specifically, our network utilizes coarse-scale feature fusion, which generates the attentive local features via combining the information from different intermediate layers. Detailed feature information is extracted with the application of two attention blocks. Extensive experiments show that our method outperforms the state-of-the-art by a significant margin on four public datasets for image retrieval tasks. Xinyao Nie, Hong Lu 0001, Zehua Guo 0003 |
ICMR | 2 |
| 2019 | Single Image Deraining via Recurrent Hierarchy Enhancement NetworkabstractSingle image deraining is an important problem in many computer vision tasks since rain streaks can severely hamper and degrade the visibility of images. In this paper, we propose a novel network named Recurrent Hierarchy Enhancement Network (ReHEN) to remove rain streaks from rainy images stage by stage. Unlike previous deep convolutional network methods, we adopt a Hierarchy Enhancement Unit (HEU) to fully extract local hierarchical features and generate effective features. Then a Recurrent Enhancement Unit (REU) is added to keep the useful information from HEU and benefit the rain removal in the later stages. To focus on different scales, shapes, and densities of rain streaks adaptively, Squeeze-and-Excitation (SE) block is applied in both HEU and REU to assign different scale factors to high-level features. Experiments on five synthetic datasets and a real-world rainy image set show that the proposed method outperforms the state-of-the-art methods considerably. The source code is available at https://github.com/nnUyi/ReHEN. Youzhao Yang, Hong Lu 0001 |
ACM Multimedia | 2 |
| 2019 | L0 Gradient Smoothing and Bimodal Histogram Analysis: A Robust Method for Sea-sky-line DetectionabstractSea-sky-line detection is an important research topic in the field of object detection and tracking on the sea. We propose an L0 gradient smoothing and bimodal histogram analysis based method to improve the robustness and accuracy of sea-sky-line detection. The proposed method mainly depends on the brightness difference between the sea region and the sky region in the image. First, we use L0 gradient smoothing to eliminate discrete noise in the image and achieve the modularity of brightness. Differing from previous methods, diagonal dividing is applied to obtain the brightness thresholds for the sky and sea regions. Then the thresholds are used for bimodal histogram analysis which helps to obtain the brightness near the sea-sky-line and narrow the detection region. After narrowing the detection region, the sea-sky-line in the image is extracted by a linear fitting method. To evaluate the performance of the proposed method, we manually construct an dataset which includes 40, 000 images taken in five scenes. Moreover, we also mark the corresponding ground-truth positions of sea-sky-line in each of the images. Extensive experiments on the dataset demonstrate that our method outperforms the state-of-the-art methods tremendously. Hong Lu 0001, Lizhe Qi |
MMAsia | 2 |
| 2016 | Wide line detection with water flowabstractLine detection plays a vital role in visual analysis tasks like Traditional Chinese Medicine (TCM) image analytics. However, most of the current methods ignore line thickness and perform poorly for the lines with different widths. This paper proposes a novel line detection method by using the water flow method. Unlike most edge-based and region-based line detectors, the water flow method is applied to obtaining the whole line response map by simply imitating the movement of water in the image smoothed by guided filter, which is viewed as a geomorphological map. In addition, this paper also proposes an adaptive parameter selection method so that the line detection can be more robust and accurate. Experimental results demonstrate the effectiveness of the proposed method on tongue crack images in comparison to the existing line extraction methods. Yangyang Hu, Hong Lu 0001, Fufeng Li, Weifei Zhang |
BIBM | 3 |
| 2015 | Teaching Video Analytics Based on Student Spatial and Temporal Behavior MiningabstractIn this paper, we propose to mine the class videos and analyze the behaviors of students to obtain the information on the focus of the students during one class. As a methodological contribution, we investigate face detection, tracking, and verification techniques to measure student engagement from the video records. We also analyze several behaviors of a single student or groups of students, such as temporal behaviors and spatial behaviors. Furthermore, we mine the relationship between students' behaviors and grades throughout the whole semester, including 16 weeks. This work can help teachers to improve the quality of class, find students who tend to get bad grades and give these students timely help. Experimental results on the test videos of different kinds of classes demonstrate the effectiveness of the proposed method. Jinxian Qin, Yaqian Zhou 0001, Hong Lu 0001, Heqing Ya |
ICMR | 3 |
| 2013 | A Segmentation and Graph-Based Video Sequence Matching Method for Video Copy DetectionabstractWe propose in this paper a segmentation and graph-based video sequence matching method for video copy detection. Specifically, due to the good stability and discriminative ability of local features, we use SIFT descriptor for video content description. However, matching based on SIFT descriptor is computationally expensive for large number of points and the high dimension. Thus, to reduce the computational complexity, we first use the dual-threshold method to segment the videos into segments with homogeneous content and extract keyframes from each segment. SIFT features are extracted from the keyframes of the segments. Then, we propose an SVD-based method to match two video frames with SIFT point set descriptors. To obtain the video sequence matching result, we propose a graph-based method. It can convert the video sequence matching into finding the longest path in the frame matching-result graph with time constraint. Experimental results demonstrate that the segmentation and graph-based video sequence matching method can detect video copies effectively. Also, the proposed method has advantages. Specifically, it can automatically find optimal sequence matching result from the disordered matching results based on spatial feature. It can also reduce the noise caused by spatial feature matching. And it is adaptive to video frame rate changes. Experimental results also demonstrate that the proposed method can obtain a better tradeoff between the effectiveness and the efficiency of video copy detection. Hong Lu 0001, Xiangyang Xue 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Gradient Ordinal Signature and Fixed-Point Embedding for Efficient Near-Duplicate Video DetectionabstractIn order to meet the requirement of large scale real-time near-duplicate video detection, this paper has achieved two goals. First, this paper proposes a more compact local image descriptor which is termed as gradient ordinal signature (GOS). GOS not only has the advantages of low dimension, simplicity in computation, and high discrimination but also is invariant to mirror reflection, rotation, and scale changes. Second, applying the characteristics of the proposed GOS and combining with the embedding theory of metric spaces, this paper proposes an efficient similarity search method based on the fixed-point embedding (FE). A main advantage of FE is that its parameters have good controllability, and its performance is stable and not sensitive to dataset changes. On the whole, the goal of our approach focuses on the speed rather than the accuracy of near-duplicate video detection. We have evaluated our method on four different settings to verify the two goals. Specifically, the tests include image and video datasets, respectively, to evaluate the performance of GOS. Experimental results demonstrate the effectiveness, efficiency, and lower memory usage of GOS. Furthermore, the third test compares FE with locality sensitivity hashing. FE also shows a speed improvement of about ten times and saves more than 60% in memory usage. The fourth test demonstrates that the combination of GOS and FE for near-duplicate video detection can achieve better overall efficiency than the state-of-the-art methods. Hong Lu 0001, Zhaohui Wen, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Salient Object Detection using concavity contextabstractConvexity (concavity) is a bottom-up cue to assign figure-ground relation in the perceptual organization [18]. It suggests that region on the convex side of a curved boundary tend to be figural. To explore the validity of this cue in the task of salient object detection, we segment the images in a test dataset into superpixels, and then locate the concave arcs and their bounding boxes along boundary of superpixels. Ecological statistics indicate that such bounding box contains salient object with a large probability. To utilize this spatial context information, i.e. concavity context, we follow the multi-scale analysis of human visual perception and design a hierarchical model. The model yields an affinity graph over candidate superpixels, in which weights between vertices are determined by the summation of concavity context on different scales in the hierarchy. Finally a graph-cut algorithm is performed to separate the salient and background objects. Evaluation on MSRA Salient Object Detection (SOD) dataset shows that concavity context is effective, and our approach provides improvement over state-of-the-art feature-based algorithms. Yao Lu 0028, Wei Zhang 0016, Hong Lu 0001, Xiangyang Xue 0001 |
ICCV | 3 |
| 2011 | Level influence of spatial pyramid matching in object classificationabstractIn this paper we propose to effectively consider the shape and size variations for object classification. Specifically, a novel image matching method is proposed to incorporate the image segmentation with Spatial Pyramid Matching (SPM), and test our method on flower classification. A Level Influence Factor (LIF) is introduced to represent weights of different pyramid levels based on the statistical information of each segmented image. Then the images are classified based on the LIF weighted spatial pyramid bag-of-visual-words feature, and some levels with weight values zeros are not needed to be compared further. Also, in SPM matching stage, the block in one image is compared with not only its corresponding block in another image, but also the spatially neighboring blocks of the corresponding blocks to find the best match. This fuzzy matching method can incorporate some translation of objects. Experiments are performed on a flower dataset containing 1360 images from 17 different categories. And experimental results demonstrate that our proposed method has better time efficiency than traditional SPM and outperforms the state-of-art flower classification methods. Hong Lu 0001, Renzhong Wei, Yanran Shen, Xiangyang Xue 0001 |
ACM Multimedia | 1 |
| 2011 | Refining local descriptors by embedding semantic information for visual categorizationabstractLocal descriptor extraction and vector quantization are the important components of widely-used Bag-of-Features (BoF) model for visual categorization. This paper proposes a simple and efficient approach to refine the local descriptors for vector quantization by embedding semantic information. The original local descriptors are integrated by a sequence of category-independent and category-dependent basis. Particularly, the category-dependent basis is learned by minimizing the joint loss minimization over local descriptors from different categories with a shared regularization penalty, which can be formulated as a linear programming problem. The transferred descriptors are further quantized and aggregated to the visual vocabulary. Experiments are performed on PASCAL VOC 2007 benchmark and the quantitative comparisons with several state-of-the-art approaches demonstrate the effectiveness of our proposed approach. Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001 |
ACM Multimedia | 3 |
| 2011 | Real-Time, Adaptive, and Locality-Based Graph Partitioning Method for Video Scene ClusteringabstractWe propose in this paper an efficient, adaptive, and locality-based graph partitioning method for video scene clustering. First, a graph partitioning method is proposed to group video shots into scenes, and a peer-group filtering (PGF) scheme is used to identify all the shots similar to each particular shot based on Fisher's discriminant analysis. To work with computable shot similarity measures that have only limited discriminating power, we develop a graph partitioning scheme to cluster the shots by maximizing the likeness of shots within the same cluster and minimizing that between different clusters. Second, considering that video data are normally obtained and viewed sequentially, we propose to perform a locality-based PGF and graph partitioning on video segments with 50 shots, 100 shots, and so on. This proposed locality-based method has the advantage that the number of scene clusters is not required to be known a priori, and it can achieve performance comparable to that processing on the whole video sequence. Experimental results are presented to demonstrate the effectiveness and efficiency of the proposed method. Hong Lu 0001, Yap-Peng Tan, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | SVD-SIFT for web near-duplicate image detectionabstractStable and high distinctive image features are the basis for web near-duplicate image detection. SIFT (scale invariant feature transform) not only has good scale and brightness invariance, also has a certain robustness to affine distortion, perspective change, and additive noise. However, to extract SIFT features to represent an image, hundreds or even thousands of SIFT key points need to be selected. And each key point needs to be described by using a 128-dimensional feature vector. Thus, the matching cost of detection method based on SIFT features is high. In this paper, we propose to apply the singular value decomposition (SVD) method for feature matching and extract the new features from the set of SIFT feature points. The extracted feature is termed as SVD-SIFT. Experimental results demonstrate that the method can obtain a better tradeoff between effectiveness and efficiency for detection. Hong Lu 0001, Xiangyang Xue 0001 |
ICIP | 2 |
| 2010 | How context helps: A discriminative codeword selection method for object detectionabstractWe first propose in this paper to localize objects in images based on the models learned from the weakly labeled images. This task is termed as region of interest (ROI) detection. Local features such as SIFT or HOG are extracted and the discriminative words from clustered codewords based on SIFT and HOG are selected to model the objects. Then how to find the discriminative words to model the object is important. Existing ROI detection methods consider the information from the foreground objects by selecting the words appearing more in the images belonging to one specific image class. Considering the information from background/context is also helpful for object detection and classification, we propose to select the discriminative words which appear more in the foreground/object and less in the background/context. Second, another task is to give the class label (object in this setting) for a given image and also give the position of the object appearing in the image. This task is termed as objection detection. A normal way for this task after ROI is to extract features from the detected regions and not from the whole image. Since the discriminative words extracted during ROI detection has good discriminative ability, we propose to use these words for object detection. Experimental results on PASCAL VOC 2006 dataset and a larger dataset containing 29 classes demonstrate the effectiveness of the proposed method. Renzhong Wei, Hong Lu 0001, Yingbin Zheng, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Weiguo Wu |
ICIP | 2 |
| 2010 | Semantic video indexing by fusing explicit and implicit context spacesabstractThis paper addresses the problem of context-based concept fusion (CBCF) for concept detection and semantic video indexing. We introduce a novel framework based on constructing context spaces of concepts, such that the contextual correlations are used to improve the performance of concept detectors. Different from traditional CBCF approach, we present two kinds of such context spaces: explicit context space for modeling the correlation of pairwise concepts, and implicit context space for representing latent themes trained from a set of concepts. The final concept detection scores are then directly fused from explicit and implicit context spaces. Experiments are presented on TRECVid 2006 benchmark and the comparisons with several state-of-the-art approaches demonstrate the effectiveness of proposed framework. Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001 |
ACM Multimedia | 3 |
| 2009 | Incorporating Spatial Correlogram into Bag-of-Features Model for Scene Categorization
Yingbin Zheng, Hong Lu 0001, Cheng Jin 0001, Xiangyang Xue 0001 |
ACCV (1) | 2 |
| 2008 | Scene segmentation based on video structure and spectral methodsabstractScene is an important semantic unit for video analysis, retrieval and browsing. However, due to the lack of a generic algorithm, many studies focus on specific methods for certain video genes, e.g., news, sports, etc. In this paper, we propose a general framework for scene segmentation. First, we construct a graph, in which the elements encode the shot-to-shot coherent characteristics of a video clip based on visual similarity and temporal relation between shots. In this step, we only exploit the inherent property of video itself and it is independent of video genres. Second, spectral method is applied on the graph to group shots into scenes. The proposed method is not only simple but also effective to deal with organized features. Experimental results validate the robustness of our method on different kinds of videos. Bin Li 0015, Hong Lu 0001, Xiangyang Xue 0001 |
ICARCV | 3 |
| 2008 | Multilayer in-place learning networks for modeling functional layers in the laminar cortex
Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
Neural Networks | 3 |
| 2008 | Metric learning by discriminant neighborhood embedding
Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 4 |
| 2007 | Salient Object Detection on Large-Scale Video DataabstractRecently more and more researches focus on the concept extraction from unstructured video data. To bridge the semantic gap between the low-level features and the high-level video concepts, a mid-level understanding of the video contents, i.e., salient object is detected based on the techniques of image segmentation and machine learning. Specifically, 21 salient object detectors are developed and tested on TRECVID 2005 development video corpus. In addition, a boosting method is proposed to select the most representative features to achieve a higher performance than only using single modality, and lower complexity than taking all features into account. Shile Zhang, Jianping Fan 0001, Hong Lu 0001, Xiangyang Xue 0001 |
CVPR | 3 |
| 2007 | Efficient Feature Extraction for Image ClassificationabstractIn many image classification applications, input feature space is often high-dimensional and dimensionality reduction is necessary to alleviate the curse of dimensionality or to reduce the cost of computation. In this paper, we extract discriminant features for image classification by learning a low-dimensional embedding from finite labeled samples. In the new feature space, intra-class compactness and extra-class separability are achieved simultaneously. Target dimensionality of the embedding is selected by spectral analysis. Our method is designed suitable for data with both uni- and multi-modal class distributions. We also develop its two-dimensional variant which makes use of the matrix representation of images. Experimental results on three real image datasets demonstrate the efficacy of our method compared to the state of the art. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Mingmin Chi, Hong Lu 0001 |
ICCV | 6 |
| 2007 | Optimal dimensionality of metric space for classificationabstractIn many real-world applications, Euclidean distance in the original space is not good due to the curse of dimensionality. In this paper, we propose a new method, called Discriminant Neighborhood Embedding (DNE), to learn an appropriate metric space for classification given finite training samples. We define a discriminant adjacent matrix in favor of classification task, i.e., neighboring samples in the same class are squeezed but those in different classes are separated as far as possible. The optimal dimensionality of the metric space can be estimated by spectral analysis in the proposed method, which is of great significance for high-dimensional patterns. Experiments with various datasets demonstrate the effectiveness of our method. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Hong Lu 0001 |
ICML | 5 |
| 2007 | The Multilayer In-Place Learning Network for the Development of General Invariances and Multi-Task LearningabstractCurrently, there is a lack of general-purpose in-place learning engines that incrementally learn multiple tasks, to develop "soft" multi-task-shared invariances in the intermediate internal representation while a developmental robot interacts with its environment. Computationally, biologically inspired in-place learning provides unusually efficient learning algorithms whose simplicity, low computational complexity, and generality are set apart from typical conventional learning algorithms. We present in this paper the multiple-layer in-place learning network (MILN) for this ambitious goal. As a key requirement for autonomous mental development, the network enables both unsupervised and supervised learning to occur concurrently, depending on whether motor supervision signals are available or not at the motor end (the last layer) during the agent's interactions with the environment. We present principles based on which MILN automatically develops invariant neurons in different layers and why such invariant neuronal clusters are important for learning later tasks in open-ended development. Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
IJCNN | 3 |
| 2007 | Incremetal Spatio-Temporal Feature Extraction and Retrieval for Large Video DatabaseabstractIn this paper we present a novel framework for semantic retrieval of video database. Each frame of video clips, characterized by its HSV (hue-saturation-value) color feature, is first projected onto the spatial principle components via CCIPCA (candid covariance-free incremental principal component analysis). Temporal Chebyshev polynomials for video clips of various lengths are captured subsequently. The similarity of two video clips is finally presented in a reasonable and computable form. The framework works incrementally and is suitable for videos of data streams in sequential order. Extensive experiments demonstrate that the framework can obtain promising results on video similarity comparison, and also with a comparably computational speedup. Bo Geng, Hong Lu 0001, Xiangyang Xue 0001 |
ISCAS | 2 |
| 2006 | An Efficient Early Termination Algorithm of Intra Prediction for H.264/AVCabstractWe propose in this paper an efficient early termination algorithm of intra prediction for H.264/AVC. It uses the spatial correlation after 16times16 inter prediction to make the judgement on whether to discard intra prediction or not. Experimental results demonstrate that the proposed algorithm can save the encoding time of H.264/AVC (JM98) between 25-45% with negligible degradation in the quality Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 2 |
| 2006 | Quotient Set-based Nonlinear Manifold for Image RestorationabstractIn this paper we propose a patch-wise coarse-to-fine algorithm for image restoration using the manifold way of visual perception. All undistorted image patches are supposed to lie on a quotient set-based nonlinear manifold, and restoration of each degraded image patch can be implemented by projecting it to a locally linear region of such nonlinear manifold. The details of the original image can be learned from the undistorted training samples. Moreover, there is no need for us to assume that the degradation function is linear or to estimate some parameters of the blurs and noises beforehand. Experimental results demonstrate the effectiveness of the proposed method Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
ICARCV | 4 |
| 2006 | In-Place Learning for Positional and Scale InvarianceabstractIn-place learning is a biologically inspired concept, meaning that the computational network is responsible for its own learning. With in-place learning, there is no need for a separate learning network. We present in this paper a multiple-layer in-place learning network (MILN) for learning positional and scale invariance. The network enables both unsupervised and supervised learning to occur concurrently. When supervision is available (e.g., from the environment during autonomous development), the network performs supervised learning through its multiple layers. When supervision is not available, the network practices while using its own practice motor signal as self-supervision (i.e., unsupervised per classical definition). We present principles based on which MILN automatically develops positional and scale invariant neurons in different layers. From sequentially sensed video streams, the proposed in-place learning algorithm develops a hierarchy of network representations. The global invariance was achieved through multi-layer quasi-invariances, with increasing invariance from early layers to the later layers. Experimental results are presented to show the effects of the principles. Juyang Weng, Hong Lu 0001, Tianyu Luwang, Xiangyang Xue 0001 |
IJCNN | 2 |
| 2006 | Null Foley-Sammon transform
Yue-Fei Guo, Lide Wu, Hong Lu 0001, Zhe Feng 0001, Xiangyang Xue 0001 |
Pattern Recognit. | 3 |
| 2006 | Discriminant neighborhood embedding for classification
Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 3 |
| 2006 | Hierarchical Indexing Structure for Efficient Similarity Search in Video RetrievalabstractWith the rapid increase in both centralized video archives and distributed WWW video resources, content-based video retrieval is gaining its importance. To support such applications efficiently, content-based video indexing must be addressed. Typically, each video is represented by a sequence of frames. Due to the high dimensionality of frame representation and the large number of frames, video indexing introduces an additional degree of complexity. In this paper, we address the problem of content-based video indexing and propose an efficient solution, called the Ordered VA-File (OVA-File) based on the VA-file. OVA-File is a hierarchical structure and has two novel features: 1) partitioning the whole file into slices such that only a small number of slices are accessed and checked during k Nearest Neighbor (kNN) search and 2) efficient handling of insertions of new vectors into the OVA-File, such that the average distance between the new vectors and those approximations near that position is minimized. To facilitate a search, we present an efficient approximate kNN algorithm named Ordered VA-LOW (OVA-LOW) based on the proposed OVA-File. OVA-LOW first chooses possible OVA-Slices by ranking the distances between their corresponding centers and the query vector, and then visits all approximations in the selected OVA-Slices to work out approximate kNN. The number of possible OVA-Slices is controlled by a user-defined parameter \delta. By adjusting \delta, OVA-LOW provides a trade-off between the query cost and the result quality. Query by video clip consisting of multiple frames is also discussed. Extensive experimental studies using real video data sets were conducted and the results showed that our methods can yield a significant speed-up over an existing VA-file-based method and iDistance with high query result quality. Furthermore, by incorporating temporal correlation of video content, our methods achieved much more efficient performance. Hong Lu 0001, Beng Chin Ooi, Heng Tao Shen, Xiangyang Xue 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2005 | Region-based Pornographic Image DetectionabstractRecent advantages in digital images and networks have made pornographic images more accessible than ever. Methods of detecting pornographic images have been proposed in some existing work. However, the skin detection modules adopted are all pixel-based and skin region shape features are rarely properly used. This paper presents a framework for pornographic image detection based on skin region information. Different from traditional works, our approach extracts color and texture features from arbitrary-shaped segmented regions. Then Gaussian mixture models are built for skin and non-skin region classification, and the skin map is produced based on the classification result. Finally, eigenregion features are used to describe the layout of skin regions on the whole image and pornographic images are detected according to the skin modality Bin Li 0015, Xiangyang Xue 0001, Hong Lu 0001 |
MMSP | 4 |
| 2005 | Efficient Video Clip Retrieval Using Index StructureabstractRetrieving similar video clips from large video database requires high query efficiency, precision and recall, which remains a challenging problem since the traditional query algorithms are inefficient and time-consuming. In this paper, we adopt the high-dimensional index structure vector-approximation file (VA-file) to organize the video database, and propose a new similarity measure which takes the temporal order among the video representations into account to improve the accuracy of query. Based on the VA-file and similarity measure, a new video clip retrieval algorithm is proposed in our method to achieve high query efficiency by using restricted sliding window to construct candidate video clips. Experimental results show that the proposed video retrieval method is efficient and effective Linjun Yang, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
MMSP | 3 |
| 2005 | An effective post-refinement method for shot boundary detectionabstractIn content-based video analysis, shot boundary detection (SBD) is a common first step which segments video data into elementary shots, each comprising a sequence of consecutive frames recording a video event or scene continuous in time and space. Many SBD methods have been proposed in the literature, and experimental results show that the existing methods work reasonably well for abrupt shot boundaries, but less effectively for gradual shot boundaries. In this paper, we propose an effective post-refinement method for identifying actual shot boundaries from the results obtained by existing SBD methods. The proposed method formulates the SBD problem as sequential detection of changes in the underlying feature distributions whose parameters are estimated from existing video shots. Specifically, the proposed post-refinement method enhances the performance of SBD by identifying as many false positives (false detections) and false negatives (miss detections) as possible. Experiments conducted on a large set of test videos, whose initial shot boundaries are obtained by four existing SBD methods, show that the proposed post-refinement method can improve markedly the detection recall and precision and is rather insensitive to the thresholds used by the existing methods in detecting the initial shot boundaries. Hong Lu 0001, Yap-Peng Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Efficient identification of speakers in news video based on shot segmentationabstractAn effective method for speaker identification in news video is presented in this paper, which is based on shot segmentation and exploits both audio and visual cues. Firstly, audio is segmented by shot segmentation based on the observation that there is only one speaker in a shot of news video in most cases. Furthermore, speech/non-speech discrimination is implemented on each shot. Finally, text-independent speaker identification is proposed using audio features on the discriminated speech shots. Experimental results show that our algorithm can obtain satisfactory performance in identifying speakers, so it can be used in real application. Xiangyang Xue 0001, Hong Lu 0001, You-san Nie |
ICARCV | 3 |
| 2004 | Improved shot boundary detection method based on text edgesabstractShot boundary detection is a pre-requisite technique for video indexing and retrieval. To avoid the influence of flashlight on abrupt shot detection, many edge-based techniques are studied thoroughly. However, these techniques are still susceptible to miss and mistake detecting the abrupt changes. Our observation shows that one of the reasons for these errors is the existence of superimposed text which has rich edges and is ever presented in video frames. To provide a solution, we present a novel method that utilizes the edge type, text edge (edge in text area) or non-text-edge (edge in other text area), reducing erroneous detection with the appearance of video text. Compared to other edge-based detection techniques, experimental results show that our proposed method achieves preferable performance. Liuhong Liang, Yang Liu 0246, Xiangyang Xue 0001, Hong Lu 0001, Yap-Peng Tan |
ICARCV | 4 |
| 2004 | Effective video text detection using line featuresabstractText superimposed on video frames provides synoptic or supplemental information on video semantics. In this paper, we propose a novel method to detect superimposed text effectively. First, we detect edges by an improved Canny edge detector. Then, a line-feature vector graph is generated based on the edge map and the stroke information is extracted. Finally text regions are generated and filtered according to line features. Experimental results show that, without much increasing the computational cost, our proposed method could suppress the false alarms notably. Furthermore, our method can be easily customized to applications with different tradeoffs in recall and precision. Yang Liu 0246, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 2 |
| 2004 | Video segmentation based on sequential change detectionabstractIn content-based video analysis, substantial research efforts have been focused on developing techniques to detect the boundaries between two successive shots, each comprising consecutive frames filmed with a single camera act. However, there is still room for further improvement in the detection performance. With the use of sequential change detection and the help of nonparametric density estimation principles, we propose A new shot boundary detection method that can maintain not only satisfactory detection accuracy, but also consistent detection performance based on the results of various test videos. Zhenyan Li, Hong Lu 0001, Yap-Peng Tan |
ICME | 2 |
| 2003 | An effective post-refinement method for shot boundary detectionabstractIn content-based video analysis, shot boundary detection is a common first step to segment video data into fundamental units of shots, each composing consecutive frames filmed with a single camera act. Many methods have been proposed in the literature for detection of shot boundaries. In this paper, we propose a new and effective post-refinement method on the detected shot boundaries by performing sequential detection of abrupt change in two underlying distributions. Experimental results show that the proposed method can eliminate most false detections and also recover many missed detections from the original detected shot boundaries, attaining better detection performance. Hong Lu 0001, Yap-Peng Tan |
ICIP (2) | 1 |
| 2003 | Unsupervised clustering of dominant scenes in sports video
Hong Lu 0001, Yap-Peng Tan |
Pattern Recognit. Lett. | 1 |
| 2002 | Content-based sports video analysis and modelingabstractWe propose in this paper some new methods for analyzing and modeling sports based on their low-level visual are first automatically identified by using the color features derived from its video are first automatically identified by using the color features derived from its video shots. To improve the performance on the identification of dominant scenes and reduce the dependency on a proper threshold, a comparison on different forms of shot color features, including shot color histograms, their principal components, and subspace linear discriminant representations, is performed. Second, the content compactness and motion attributes of clustered video shots are analyzed to differentiate dominant scene types. A scene transition diagram is then constructed to form a structural descriptor for sports video contents. Third, the video shots belonging to each dominant scene are processed, using customized schemes and domain specific knowledge, to identify interesting play events in the sports video. Experimental results on identification of dominant scenes, structural desprictors and high-level ball videos, are presented to demonstrate the possible applications of the proposed methods. Hong Lu 0001, Yap-Peng Tan |
ICARCV | 1 |
| 2002 | Model-based clustering and analysis of video scenesabstractWe make two contributions. First, we develop an unsupervised method to discover clusters of video scenes and summarize them with a concise Gaussian mixture model. To search for the best possible model, an effective procedure is devised to compare among models with different dimensions (i.e., numbers of mixture components) and, for a given dimension, among models with different parameters. Second, we propose a scene interference measure to characterize the interaction among different scenes of a video sequence. When applied to the clustered video scenes, the measure can reveal the dominant video segments of a class of videos without requiring much domain-specific knowledge. The proposed methods have been tested with a large number of sports videos and promising results are reported. Yap-Peng Tan, Hong Lu 0001 |
ICIP (1) | 2 |
| 2002 | On model-based clustering of video scenes using sceneletsabstractWe propose in this paper a model-based approach to clustering video scenes based on scenelets. We define a video scenelet as a short consecutive sample of frames of a video sequence. The approach makes use of an unsupervised method to represent scenelets of a video with a concise Gaussian mixture model and cluster them into different video scenes according to their visual similarities. In particular the expectation-maximization algorithm is employed to estimate the unknown model parameters, and Bayesian information criterion is used to determine the optimal number and model of scene clusters in a principled manner. This approach is fundamentally different from many existing video clustering methods, as it does not require explicit knowledge of shot boundaries. Instead, the shot boundaries can also be obtained as a by-product of the scene clustering process. The proposed methods have been tested with various types of sports videos and promising results are reported in this paper. Hong Lu 0001, Yap-Peng Tan |
ICME (1) | 1 |
| 2001 | Sports video analysis and structuringabstractWe propose a new method for structuring and analyzing sports video through clustering of video shots based on low-level visual content. The dominant scenes of a sports video are first automatically extracted by using the color features derived from its video shots. The video shots belonging to each dominant scene are then processed to identify interesting play events in the sports video. To reduce the influence of threshold selection on the results, different forms of shot color features, including shot color histogram and its principal component and subspace linear discriminant representations, are also examined. Experimental results on various kinds of sports videos, such as tennis, volleyball, basketball and football videos, are presented to demonstrate the effectiveness of the proposed method. Hong Lu 0001, Yap-Peng Tan |
MMSP | 1 |