VLDB 2026 Research / reviewers in the wild / expert
Weiyao Lin
dblp:42/6095
· DBLP profile ↗
161ranked-venue papers
21as first author
56since 2021 · last 2026
0000-0001-8307-7107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 12 first-author · 32 since 2021Artificial intelligence and machine learning · 70 · 5 first-author · 39 since 2021Systems, architecture and hardware · 12 · 5 first-author · 1 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PCGS: Progressive Compression of 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) achieves impressive rendering fidelity and speed for novel view synthesis. However, its substantial data size poses a significant challenge for practical applications. While many compression techniques have been proposed, they fail to efficiently utilize existing bitstreams in on-demand applications due to their lack of progressivity, leading to a waste of resource. To address this issue, we propose PCGS (Progressive Compression of 3D Gaussian Splatting), which adaptively controls both the quantity and quality of Gaussians (or anchors) to enable effective progressivity for on-demand applications. For quantity, we introduce a progressive masking strategy that incrementally incorporates new anchors while refining existing ones to enhance fidelity. For quality, we propose a progressive quantization approach that gradually reduces quantization step sizes to achieve finer modeling of Gaussian attributes. Furthermore, to compact the incremental bitstreams, we leverage existing quantization results to refine probability prediction, improving entropy coding efficiency across progressive levels. PCGS achieves progressivity while maintaining compression performance comparable to SoTA non-progressive methods. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
AAAI | 4 |
| 2026 | Head-Aware KV Cache Compression for Efficient Visual Autoregressive ModelingabstractVisual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n4) to O(Bn2). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster in- ference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B. Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo 0002, Zeren Zhang, Danping Zou, Weiyao Lin |
AAAI | 7 |
| 2026 | CogStream: Context-guided Streaming Video Question AnsweringabstractDespite advancements in Video Large Language Models (Vid-LLMs) improving multimodal understanding, challenges persist in streaming video reasoning due to its reliance on contextual information. Existing paradigms feed all available historical contextual information into Vid-LLMs, resulting in a significant computational burden for visual data processing. Furthermore, the inclusion of irrelevant context distracts models from key details. This paper introduces a challenging task called Context-guided Streaming Video Reasoning (CogStream), which simulates real-world streaming video scenarios, requiring models to identify the most relevant historical contextual information to deduce answers for questions about the current stream. To support CogStream, we present a densely annotated dataset featuring extensive and hierarchical question-answer pairs, generated by a semi-automatic pipeline. Additionally, we present CogReasoner as a baseline model. It effectively tackles this task by leveraging visual stream compression and historical dialogue retrieval. Extensive experiments prove the effectiveness of this method. Zicheng Zhao, Kangyu Wang, Rui Qian 0001, Weiyao Lin, Huabin Liu 0001 |
AAAI | 5 |
| 2026 | CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace CreditabstractKangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao, Lin Liu, Jianguo Li, Zhenzhong Lan, Weiyao Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao, Zhen-Zhong Lan, Weiyao Lin |
ACL (1) | 8 |
| 2026 | Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Chaofan Gan, Mingxi Lv, Ziran Qin, Li Shen 0008, Junhui Hou, Weiyao Lin |
Int. J. Comput. Vis. | 11 |
| 2026 | MECD+: Unlocking Event-Level Causal Graph Discovery for Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on brief video segments containing isolated events and basic causal relations, lacking comprehensive and structured causality analysis for videos with multiple interconnected events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relations between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD identifies the causal associations between these events to derive a comprehensive and structured event-level video causal graph explaining why and how the result event occurred. To address the challenges of MECD, we devise a novel framework inspired by the Granger Causality method, incorporating an efficient mask-based event prediction model to perform an Event Granger Test. It estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to mitigate challenges in MECD like causality confounding and illusory causality. Additionally, context chain reasoning is introduced to conduct more robust and generalized reasoning. Experiments validate the effectiveness of our framework in reasoning complete causal relations, outperforming GPT-4o and VideoChat2 by 5.77% and 2.70%, respectively. Further experiments demonstrate that causal relation graphs can also contribute to downstream video understanding tasks such as video question answering and video event prediction. Tieyuan Chen, Huabin Liu 0001, Yi Wang 0033, Yihang Chen 0002, Tianyao He, Chaofan Gan, Huanyu He, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Few-Shot Action Recognition via Intra- and Inter-Video Information MaximizationabstractCurrent few-shot action recognition involves two primary sources of information for classification: (1) intra-video information, determined by frame content within a single video clip, and (2) inter-video information, measured by relationships (e.g., feature similarity) among videos. However, existing methods inadequately exploit these two information sources. In terms of intra-video information, current sampling operations for input videos may omit critical action information, reducing the utilization efficiency of video data. For the inter-video information, the action misalignment among videos makes it challenging to calculate precise relationships. Moreover, how to jointly consider both inter- and intra-video information remains under-explored for few-shot action recognition. To this end, we propose a novel framework, Video Information Maximization (VIM), for few-shot video action recognition. VIM is equipped with an adaptive spatial-temporal video sampler and a spatial-temporal action alignment model to maximize intra- and inter-video information, respectively. The video sampler adaptively selects important frames and amplifies critical spatial regions for each input video based on the task at hand. This preserves and emphasizes informative parts of video clips while eliminating interference at the data level. The alignment model performs temporal and spatial action alignment sequentially at the feature level, leading to more precise measurements of inter-video similarity. Finally, based on the mutual information measurement, we introduce a new training objective into few-shot learning, which provides explicit guidance in jointly maximizing intra- and inter-video information in our VIM. Extensive experimental results on public datasets for few-shot action recognition demonstrate the effectiveness of our framework. Huabin Liu 0001, Tieyuan Chen, Yuxi Li 0009, Shuyuan Li, John See, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Feedforward Compression of Static and Streamable 3D Gaussian SplattingabstractRecent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis, yet their substantial storage cost remains a major barrier to practical deployment. Although several compression techniques have been explored, they share a common limitation:each existing 3DGS requires per-scene optimization to achieve compression, making the compressionslow and inefficient. In this work, we present Fast Compression of 3D Gaussian Splatting (FCGS), an optimization-free approach that compresses existing 3DGS in a single feed-forward pass, reducing compression time from minutes to seconds. To enhance compression efficiency, we design a multi-path entropy module that routes Gaussian attributes through separate entropy-constrained paths, achieving a better trade-off between size and fidelity. Furthermore, we introduce both inter- and intra-Gaussian context models to effectively remove redundancies for the unstructured Gaussian representation. Experimental results show that FCGS achieves over 20× compression while maintaining high fidelity, outperforming most State-of-The-Art (SoTA) per-scene optimization-based methods. Beyond static scenes, we further extend FCGS to a streamable setting which eliminates redundant temporal information, demonstrating its strong potential for compressing streamable 3DGS data. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Junhui Hou, Mehrtash Harandi, Jianfei Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Bridging the Gap Between Implicit and Explicit Representation for Efficient Image CompressionabstractNeural Image Compression (NIC) has achieved superior compression performance by modeling images as implicit feature representations, yet its practical deployment is severely hindered by computational overhead. Recently, GaussianImage was proposed as a computationally efficient explicit image representation paradigm, which renders images from 2D Gaussians. However, it suffers from inferior compression performance relative to mainstream NIC frameworks, mainly due to low compressibility and representation capability of explicit Gaussian parameters. To this end, we propose Pixel-Aligned Generalized 2D Splatting (PA-G2DS) as a computation-efficient and compression-friendly image representation format. Specifically, we deploy a learnable rendering function with implicit coefficients to enhance the reconstructed image quality and improve the compression ratio over explicit Gaussian coefficients. Incorporating the proposed PA-G2DS as a computationally efficient decoder, we further develop a suite of image codecs optimized for either compression ratios or flexible deployment scenarios. Experiments prove that the proposed codecs could achieve 30ms compression latency and millisecond-level decompression latency, reducing the performance and efficiency gap between implicit and explicit image representation. Furthermore, the proposed codec opens potential applications for NIC such as JPEG-like sequential decompression and random-access during decompression. Wenrui Dai, Chern Hong Lim, Carl J. Debono, Thittaporn Ganokratanaa, Junhui Hou, Weiyao Lin |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Neural Compression for 3D Geometry Sets
Junhui Hou, Weiyao Lin, Wenping Wang 0001 |
ICCV | 3 |
| 2025 | Fast Feedforward 3D Gaussian Splatting CompressionabstractWith 3D Gaussian Splatting (3DGS) advancing real-time and high-fidelity rendering for novel view synthesis, storage requirements pose challenges for their widespread adoption. Although various compression techniques have been proposed, previous art suffers from a common limitation: for any existing 3DGS, per-scene optimization is needed to achieve compression, making the compression sluggish and slow. To address this issue, we introduce Fast Compression of 3D Gaussian Splatting (FCGS), an optimization-free model that can compress 3DGS representations rapidly in a single feed-forward pass, which significantly reduces compression time from minutes to seconds. To enhance compression efficiency, we propose a multi-path entropy module that assigns Gaussian attributes to different entropy constraint paths for balance between size and fidelity. We also carefully design both inter- and intra-Gaussian context models to remove redundancies among the unstructured Gaussian blobs. Overall, FCGS achieves a compression ratio of over 20X while maintaining fidelity, surpassing most per-scene SOTA optimization-based methods. Code: github.com/YihangChen-ee/FCGS. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
ICLR | 4 |
| 2025 | Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient AttentionsabstractRecent advancements in Transformer-based large language models (LLMs) have set new standards in natural language processing. However, the classical softmax attention incurs significant computational costs, leading to a $O(T)$ complexity for per-token generation, where $T$ represents the context length. This work explores reducing LLMs' complexity while maintaining performance by introducing Rodimus and its enhanced version, Rodimus$+$. Rodimus employs an innovative data-dependent tempered selection (DDTS) mechanism within a linear attention-based, purely recurrent framework, achieving significant accuracy while drastically reducing the memory usage typically associated with recurrent models. This method exemplifies semantic compression by maintaining essential input information with fixed-size hidden states. Building on this, Rodimus$+$ combines Rodimus with the innovative Sliding Window Shared-Key Attention (SW-SKA) in a hybrid approach, effectively leveraging the complementary semantic, token, and head compression techniques. Our experiments demonstrate that Rodimus$+$-1.6B, trained on 1 trillion tokens, achieves superior downstream performance against models trained on more tokens, including Qwen2-1.5B and RWKV6-1.6B, underscoring its potential to redefine the accuracy-efficiency balance in LLMs. Model code and pre-trained checkpoints are open-sourced at https://github.com/codefuse-ai/rodimus. Hang Yu 0002, Zi Gong, Shizhan Liu, Weiyao Lin |
ICLR | 6 |
| 2025 | CAKE: Cascading and Adaptive KV Cache Eviction with Layer PreferencesabstractLarge language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference burden, they often fail to allocate resources rationally across layers with different attention patterns. In this paper, we introduce Cascading and Adaptive KV cache Eviction (CAKE), a novel approach that frames KV cache eviction as a ``cake-slicing problem.''
CAKE assesses layer-specific preferences by considering attention dynamics in both spatial and temporal dimensions, allocates rational cache size for layers accordingly, and manages memory constraints in a cascading manner. This approach enables a global view of cache allocation, adaptively distributing resources across diverse attention mechanisms while maintaining memory budgets.
CAKE also employs a new eviction indicator that considers the shifting importance of tokens over time, addressing limitations in existing methods that overlook temporal dynamics.
Comprehensive experiments on LongBench and NeedleBench show that CAKE maintains model performance with only 3.2\% of the KV cache and consistently outperforms current baselines across various models and memory constraints, particularly in low-memory settings. Additionally, CAKE achieves over 10$\times$ speedup in decoding latency compared to full cache when processing contexts of 128K tokens with FlashAttention-2. Our code is available at https://github.com/antgroup/cakekv. Ziran Qin, Yuchen Cao 0007, Mingbao Lin, Shixuan Fan, Weiyao Lin |
ICLR | 7 |
| 2025 | Efficient Lossless Compression with Distribution Quantized Finite-State Autoregressive ModelabstractLearned lossless image compression has been a popular topic in recent research. While outperforming many traditional compression techniques in terms of compression ratio, the high time and space complexity of such methods prevent them from wider applications. Recently, Finite-State AutoRegressive (FSAR) entropy coding was proposed, boosting autoregressive entropy coding for the quantized latent space utilizing a lookup table. However, such a method still consumes an extraordinary amount of memory, which prevents its application on low-end devices. In this work, we propose the Distribution Quantized FSAR (DQ-FSAR) model, building a two-step lookup table for the FSAR model that significantly reduces memory occupation. Furthermore, we design compact loss functions to optimize the DQ-FSAR model. Experiments show that our method reduces more than 99% memory usage compared to the baseline, with a similar compression ratio and speed. Carl J. Debono, Weiyao Lin |
ISCAS | 4 |
| 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsabstractPre-trained stable diffusion models (SD) have shown great advances in visual correspondence.
In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as massive activations, leading to uninformative representations and significant performance degradation for DiTs.
The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information.
We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs.
Building on these findings, we propose the Diffusion Transformer Feature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs.
Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation.
Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations.
Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (e.g., with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.). Chaofan Gan, Yuanpeng Tu, Tieyuan Chen, Yuxi Li 0009, Mehrtash Harandi, Weiyao Lin |
NeurIPS | 7 |
| 2025 | Low-rank Winograd transformation for 3D convolutional neural networks
Ziran Qin, Mingbao Lin, Huabin Liu 0001, John See, Gui Zou, Weiyao Lin |
Sci. China Inf. Sci. | 6 |
| 2025 | Achieving Procedure-Aware Instructional Video Correlation Learning Under Weak Supervision from a Collaborative Perspective
Tianyao He, Huabin Liu 0001, Zelin Ni, Yuxi Li 0009, Yang Zhang 0002, Weiyao Lin |
Int. J. Comput. Vis. | 9 |
| 2025 | Toward Accurate and Robust Pedestrian Detection via Variational Inference
Huanyu He, Weiyao Lin, Tianyao He, Yuxi Li 0009 |
Int. J. Comput. Vis. | 2 |
| 2025 | HAC++: Towards 100X Compression of 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a promising representation for novel view synthesis, boosting rapid rendering speed with high fidelity. However, the substantial Gaussians and their associated attributes necessitate effective compression techniques. Nevertheless, the sparse and unorganized nature of the point cloud of Gaussians (or anchors in our paper) presents challenges for compression. In this paper, we propose HAC++, which explicitly minimizes the representation's entropy during optimization, enabling efficient arithmetic coding after training for compressed storage. Specifically, to reduce entropy, HAC++ leverages the relationships between unorganized anchors and a structured hash grid, utilizing their mutual information for context modeling. Additionally, HAC++ captures intra-anchor contextual relationships to further enhance compression performance. To facilitate entropy coding, we utilize Gaussian distributions to precisely estimate the probability of each quantized attribute, where an adaptive quantization module is proposed to enable high-precision quantization of these attributes for improved fidelity restoration. Moreover, we incorporate an adaptive masking strategy to eliminate non-effective Gaussians and anchors. Overall, HAC++ achieves a remarkable size reduction of over $100\times$100× compared to vanilla 3DGS when averaged on all datasets, while simultaneously improving fidelity. It also delivers more than $20\times$20× size reduction compared to Scaffold-GS. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental LearningabstractContinual learning aims to acquire new knowledge while retaining past information. Class-incremental learning (CIL) presents a challenging scenario where classes are introduced sequentially. For video data, the task becomes more complex than image data because it requires learning and preserving both spatial appearance and temporal action involvement. To address this challenge, we propose a novel exemplar-free framework that equips separate spatiotemporal adapters to learn new class patterns, accommodating the incremental information representation requirements unique to each class. While separate adapters are proven to mitigate forgetting and fit unique requirements, naively applying them hinders the intrinsic connection between spatial and temporal information increments, affecting the efficiency of representing newly learned class information. Motivated by this, we introduce two key innovations from a causal perspective. First, a causal distillation module is devised to maintain the relation between spatial-temporal knowledge for a more efficient representation. Second, a causal compensation mechanism is proposed to reduce the conflicts during increment and memorization between different types of information. Extensive experiments conducted on benchmark datasets demonstrate that our framework can achieve new state-of-the-art results, surpassing current example-based methods by 4.2% in accuracy on average. The codes are accessible in https://github.com/tychen-SJTU/CSTA. Tieyuan Chen, Huabin Liu 0001, Chern Hong Lim, John See, Xing Gao 0005, Junhui Hou, Weiyao Lin |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisabstractVideo Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL to instructional videos is still quite challenging due to their intrinsic procedural temporal structure. Specifically, procedural knowledge is critical for accurate correlation analyses on instructional videos. Nevertheless, current procedure-learning methods heavily rely on step-level annotations, which are costly and not scalable. To address this problem, we introduce a weakly supervised framework called Collaborative Procedure Alignment (CPA) for procedure-aware correlation learning on instructional videos. Our framework comprises two core modules: collaborative step mining and frame-to-step alignment. The collaborative step mining module enables simultaneous and consistent step segmentation for paired videos, leveraging the semantic and temporal similarity between frames. Based on the identified steps, the frame-to-step alignment module performs alignment between the frames and steps across videos. The alignment result serves as a measurement of the correlation distance between two videos. We instantiate our framework in two distinct instructional video tasks: sequence verification and action quality assessment. Extensive experiments validate the effectiveness of our approach in providing accurate and interpretable correlation analyses for instructional videos. Tianyao He, Huabin Liu 0001, Yuxi Li 0009, Yang Zhang 0002, Weiyao Lin |
AAAI | 7 |
| 2024 | Density Matters: Improved Core-Set for Active Domain Adaptive SegmentationabstractActive domain adaptation has emerged as a solution to balance the expensive annotation cost and the performance of trained models in semantic segmentation. However, existing works usually ignore the correlation between selected samples and its local context in feature space, which leads to inferior usage of annotation budgets. In this work, we revisit the theoretical bound of the classical Core-set method and identify that the performance is closely related to the local sample distribution around selected samples. To estimate the density of local samples efficiently, we introduce a local proxy estimator with Dynamic Masked Convolution and develop a Density-aware Greedy algorithm to optimize the bound. Extensive experiments demonstrate the superiority of our approach. Moreover, with very few labels, our scheme achieves comparable performance to the fully supervised counterpart. Shizhan Liu, Zhengkai Jiang 0001, Yuxi Li 0009, Jinlong Peng, Yabiao Wang, Weiyao Lin |
AAAI | 6 |
| 2024 | HAC: Hash-Grid Assisted Context for 3D Gaussian Splatting Compression
Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
ECCV (7) | 3 |
| 2024 | TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional Reasoning
Huabin Liu 0001, Yang Zhang 0002, Weiyao Lin |
ECCV (5) | 5 |
| 2024 | BaSIC: BayesNet Structure Learning for Computational Scalable Neural Image Compression
Hang Yu 0002, Shizhan Liu, Wenrui Dai, Weiyao Lin |
ECCV (25) | 5 |
| 2024 | Finite-State Autoregressive Entropy Coding for Efficient Learned Lossless CompressionabstractLearned lossless data compression has garnered significant attention recently due to its superior compression ratios compared to traditional compressors. However, the computational efficiency of these models jeopardizes their practicality. This paper proposes a novel system for improving the compression ratio while maintaining computational efficiency for learned lossless data compression. Our approach incorporates two essential innovations. First, we propose the Finite-State AutoRegressive (FSAR) entropy coder, an efficient autoregressive Markov model based entropy coder that utilizes a lookup table to expedite autoregressive entropy coding. Next, we present a Straight-Through Hardmax Quantization (STHQ) scheme to enhance the optimization of discrete latent space. Our experiments show that the proposed lossless compression method could improve the compression ratio by up to 6\% compared to the baseline, with negligible extra computational time. Our work provides valuable insights into enhancing the computational efficiency of learned lossless data compression, which can have practical applications in various fields. Code is available at https://github.com/alipay/Finite_State_Autoregressive_Entropy_Coding. Hang Yu 0002, Weiyao Lin |
ICLR | 4 |
| 2024 | Visibility-guided Human Body Reconstruction from Uncalibrated Multi-view CamerasabstractWe present a novel method for 3D human body reconstruction with multi-view images from calibration-free cameras by multi-view fusion with explicit visibility modelling. Existing multi-view methods usually establish geometric constraints by using accurate camera intrinsic and extrinsic parameters. Despite remarkable performances, multi-view camera calibration often requires complex operations and additional maintenance to fix camera positions and angles, which restrict its applicability to real-world scenarios. In contrast, we leverage vertex-wise visibility prediction as calibration cues to guide the multi-view human body aggregation, which eliminates the need for camera calibration. Specifically, we estimate the UV position map and the vertex-wise visibility map of human body in each camera view, which allows us to align and aggregate multi-view information in a hierarchical manner. To further improve the alignment between human body and vertex-wise visual features, we propose an Occlusion-aware UV-pixel Refinement (OUVR) module, which takes the previous result of coarse alignment as input. The visible vertices are disentangled from the UV map and are reprojected on the image to describe the misalignment of current body estimation and image features. The UV map representation is adopted throughout the refinement process to avoid the potential error propagation brought by parametric representation. The effectiveness of our approach is validated on 3D human body reconstruction, as it surpasses current leading multi-view fusion methods, and showing comparable performance to methods that require accurate multi-view camera calibration. Zhenyu Xie, Huanyu He, Gui Zou, Weiyao Lin |
ICMR | 9 |
| 2024 | DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and CorrectionabstractWith the recent burst of 2D and 3D data, cross-modal retrieval has attracted increasing attention recently. However, manual labeling by non-experts will inevitably introduce corrupted annotations given ambiguous 2D/3D content. Though previous works have addressed this issue by designing a naive division strategy with hand-crafted thresholds, their performance generally exhibits great sensitivity to the threshold value. Besides, they fail to fully utilize the valuable supervisory signals within each divided subset. To tackle this problem, we propose a Divide-and-conquer 2D-3D cross-modal Alignment and Correction framework (DAC), which comprises Multimodal Dynamic Division (MDD) and Adaptive Alignment and Correction (AAC). Specifically, the former performs accurate sample division by adaptive credibility modeling for each sample based on the compensation information within multimodal loss distribution. Then in AAC, samples in distinct subsets are exploited with different alignment strategies to fully enhance the semantic compactness and meanwhile alleviate over-fitting to noisy labels, where a self-correction strategy is introduced to improve the quality of representation. Moreover. To evaluate the effectiveness in real-world scenarios, we introduce a challenging noisy benchmark, namely Objaverse-N200, which comprises 200k-level samples annotated with 1156 realistic noisy labels. Extensive experiments on both traditional and the newly proposed benchmarks demonstrate the generality and superiority of our DAC, where DAC outperforms state-of-the-art models by a large margin. (i.e., with +5.9% gain on ModelNet40 and +5.8% on Objaverse-N200). Chaofan Gan, Yuanpeng Tu, Yuxi Li 0009, Weiyao Lin |
ACM Multimedia | 4 |
| 2024 | MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningabstractVideo causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradigm and focusing on short videos containing only a single event and simple causal relationships, lacking comprehensive and structured causality analysis for videos with multiple events. To fill this gap, we introduce a new task and dataset, Multi-Event Causal Discovery (MECD). It aims to uncover the causal relationships between events distributed chronologically across long videos. Given visual segments and textual descriptions of events, MECD requires identifying the causal associations between these events to derive a comprehensive, structured event-level video causal diagram explaining why and how the final result event occurred. To address MECD, we devise a novel framework inspired by the Granger Causality method, using an efficient mask-based event prediction model to perform an Event Granger Test, which estimates causality by comparing the predicted result event when premise events are masked versus unmasked. Furthermore, we integrate causal inference techniques such as front-door adjustment and counterfactual inference to address challenges in MECD like causality confounding and illusory causality. Experiments validate the effectiveness of our framework in providing causal relationships in multi-event videos, outperforming GPT-4o and VideoLLaVA by 5.7% and 4.1%, respectively. Tieyuan Chen, Huabin Liu 0001, Tianyao He, Yihang Chen 0002, Chaofan Gan, Yang Zhang 0002, Weiyao Lin |
NeurIPS | 11 |
| 2024 | Scene Graph Lossless Compression with Adaptive Prediction for Objects and RelationsabstractThe scene graph is a novel data structure describing objects and their pairwise relationship within image scenes. As the size of scene graphs in vision and multimedia applications increases, the need for lossless storage and transmission of such data becomes more critical. However, the compression of scene graphs is less studied because of the complicated data structures involved and complex distributions. Existing solutions usually involve general-purpose compressors or graph structure compression methods, which are weak at reducing the redundancy in scene graph data. This article introduces a novel lossless compression framework with adaptive predictors for the joint compression of objects and relations in scene graph data. The proposed framework comprises a unified prior extractor and specialized element predictors to adapt to different data elements. Furthermore, to exploit the context information within and between graph elements, Graph Context Convolution is proposed to support different graph context modeling schemes for different graph elements. Finally, an overarching framework incorporates the learned distribution model to predict numerical data under complicated conditional constraints. Experiments conducted on labeled or generated scene graphs demonstrate the effectiveness of the proposed framework for scene graph lossless compression. Weiyao Lin, Wenrui Dai, Huabin Liu 0001, John See, Hongkai Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | A Unified Framework for Jointly Compressing Visual and Semantic DataabstractThe rapid advancement of multimedia and imaging technologies has resulted in increasingly diverse visual and semantic data. A large range of applications such as remote-assisted driving requires the amalgamated storage and transmission of various visual and semantic data. However, existing works suffer from the limitation of insufficiently exploiting the redundancy between different types of data. In this article, we propose a unified framework to jointly compress a diverse spectrum of visual and semantic data, including images, point clouds, segmentation maps, object attributes, and relations. We develop a unifying process that embeds the representations of these data into a joint embedding graph according to their categories, which enables flexible handling of joint compression tasks for various visual and semantic data. To fully leverage the redundancy between different data types, we further introduce an embedding-based adaptive joint encoding process and a Semantic Adaptation Module to efficiently encode diverse data based on the learned embeddings in the joint embedding graph. Experiments on the Cityscapes, MSCOCO, and KITTI datasets demonstrate the superiority of our framework, highlighting promising steps toward scalable multimedia processing. Shizhan Liu, Weiyao Lin, Yihang Chen 0002, Wenrui Dai, John See, Hongkai Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable BasisabstractBases have become an integral part of modern deep learning-based models for time series forecasting due to their ability to act as feature extractors or future references. To be effective, a basis must be tailored to the specific set of time series data and exhibit distinct correlation with each time series within the set. However, current state-of-the-art methods are limited in their ability to satisfy both of these requirements simultaneously. To address this challenge, we propose BasisFormer, an end-to-end time series forecasting architecture that leverages learnable and interpretable bases. This architecture comprises three components: First, we acquire bases through adaptive self-supervised learning, which treats the historical and future sections of the time series as two distinct views and employs contrastive learning. Next, we design a Coef module that calculates the similarity coefficients between the time series and bases in the historical view via bidirectional cross-attention. Finally, we present a Forecast module that selects and consolidates the bases in the future view based on the similarity coefficients, resulting in accurate future predictions. Through extensive experiments on six datasets, we demonstrate that BasisFormer outperforms previous state-of-the-art methods by 11.04% and 15.78% respectively for univariate and multivariate forecasting tasks. Code is
available at: https://github.com/nzl5116190/Basisformer. Zelin Ni, Hang Yu 0002, Shizhan Liu, Weiyao Lin |
NeurIPS | 5 |
| 2023 | HiEve: A Large-Scale Benchmark for Human-Centric Video Analysis in Complex Events
Weiyao Lin, Huabin Liu 0001, Shizhan Liu, Yuxi Li 0009, Hongkai Xiong, Guo-Jun Qi, Nicu Sebe |
Int. J. Comput. Vis. | 1 |
| 2023 | Spatio-Temporal Point Process for Multiple Object TrackingabstractMultiple object tracking (MOT) focuses on modeling the relationship of detected objects among consecutive frames and merge them into different trajectories. MOT remains a challenging task as noisy and confusing detection results often hinder the final performance. Furthermore, most existing research are focusing on improving detection algorithms and association strategies. As such, we propose a novel framework that can effectively predict and mask-out the noisy and confusing detection results before associating the objects into trajectories. In particular, we formulate such "bad" detection results as a sequence of events and adopt the spatio-temporal point process to model such events. Traditionally, the occurrence rate in a point process is characterized by an explicitly defined intensity function, which depends on the prior knowledge of some specific tasks. Thus, designing a proper model is expensive and time-consuming, with also limited ability to generalize well. To tackle this problem, we adopt the convolutional recurrent neural network (conv-RNN) to instantiate the point process, where its intensity function is automatically modeled by the training data. Furthermore, we show that our method captures both temporal and spatial evolution, which is essential in modeling events for MOT. Experimental results demonstrate notable improvements in addressing noisy and confusing detection results in MOT data sets. An improved state-of-the-art performance is achieved by incorporating our baseline MOT algorithm with the spatio-temporal point process model. Tao Wang 0002, Kean Chen, Weiyao Lin, John See, Zenghui Zhang, Xia Jia |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionabstractFew-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this similarity is not ideal since different action instances may show distinctive temporal distribution, resulting in severe misalignment issues across query and support videos. In this paper, we arrest this problem from two distinct aspects -- action duration misalignment and action evolution misalignment. We address them sequentially through a Two-stage Action Alignment Network (TA2N). The first stage locates the action by learning a temporal affine transform, which warps each video feature to its action duration while dismissing the action-irrelevant feature (e.g. background). Next, the second stage coordinates query feature to match the spatial-temporal action evolution of support by performing temporally rearrange and spatially offset prediction. Extensive experiments on benchmark datasets show the potential of the proposed method in achieving state-of-the-art performance for few-shot action recognition. Shuyuan Li, Huabin Liu 0001, Rui Qian 0001, Yuxi Li 0009, John See, Mengjuan Fei, Xiaoyuan Yu, Weiyao Lin |
AAAI | 8 |
| 2022 | Visual Sound Localization in the Wild by Cross-Modal Interference ErasingabstractThe task of audiovisual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real world scenarios, audios are usually contaminated by off screen sound and background noise. They will interfere with the procedure of identifying desired sources and building visual sound connections, making previous studies nonapplicable. In this work, we propose the Interference Eraser (IEr) framework, which tackles the problem of audiovisual sound source localization in the wild. The key idea is to eliminate the interference by redefining and carving discriminative audio representations. Specifically, we observe that the previous practice of learning only a single audio representation is insufficient due to the additive nature of audio signals. We thus extend the audio representation with our Audio Instance Identifier module, which clearly distinguishes sounding instances when audio signals of different volumes are unevenly mixed. Then we erase the influence of the audible but off screen sounds and the silent but visible objects by a Cross modal Referrer module with cross modality distillation. Quantitative and qualitative evaluations demonstrate that our framework achieves superior results on sound localization tasks, especially under real world scenarios. Rui Qian 0001, Hang Zhou 0009, Di Hu 0001, Weiyao Lin, Ziwei Liu 0002, Bolei Zhou, Xiaowei Zhou 0001 |
AAAI | 5 |
| 2022 | Speed up Object Detection on Gigapixel-level Images with Patch ArrangementabstractWith the appearance of super high-resolution (e.g., gigapixel-level) images, performing efficient object detection on such images becomes an important issue. Most ex-isting works for efficient object detection on high-resolution images focus on generating local patches where objects may exist, and then every patch is detected independently. How-ever, when the image resolution reaches gigapixel-level, they will suffer from a huge time cost for detecting numerous patches. Different from them, we devise a novel patch ar-rangement frameworkfor fast object detection on gigapixel-level images. Under this framework, a Patch Arrangement Network (PAN) is proposed to accelerate the detection by determining which patches could be packed together into a compact canvas. Specifically, PAN consists of (1) a Patch Filter Module (PFM) (2) a Patch Packing Module (PPM). PFM filters patch candidates by learning to select patches between two granularities. Subsequently, from the remaining patches, PPM determines how to pack these patches to-gether into a smaller number of canvases. Meanwhile, it generates an ideal layout of patches on canvas. These can-vases are fed to the detector to get final results. Experiments show that our method could improve the inference speed on gigapixel-level images by 5 x while maintaining great performance. Huabin Liu 0001, John See, Aixin Zhang, Weiyao Lin |
CVPR | 6 |
| 2022 | Pyraformer: Low-Complexity Pyramidal Attention for Long-Range Time Series Modeling and Forecasting
Shizhan Liu, Hang Yu 0002, Cong Liao, Weiyao Lin, Alex X. Liu, Schahram Dustdar |
ICLR | 5 |
| 2022 | Trace-Level Invisible Enhanced Network for 6D Pose EstimationabstractEstimating 6D pose of the object from a single image is es-sential for robotic manipulation. Many recent learning-based methods directly regress the pose from 2D-3D points corre-spondence. The problem is that, these methods only make use of visible information from the single-view image, resulting ambiguity for the network to solve pose from the limited cor-responding pairs. To overcome this problem, this paper intro-duces INVNet, integrating invisible information into the visi-ble 2D-3D correspondence to model geometry features of the 3D object. Instead of directly reconstruct the coordinate of in-visible points, we propose Trace-level Geometry Path, which estimates the trace-level depth of the object model for each image pixel. Specifically, our INVNet generates dense visible correspondence as well as Trace-level Geometry Path map, then learn to solve 6D pose from them. Meanwhile, each cam-era ray along with Trace-level Geometry Path is transformed to the object space by the predicted pose to compute invisi-ble correspondence loss from visible one, back to enhance its learning. Extensive experiments show that our approach out-performs state-of-the-art methods on the benchmark LM and LM-O datasets. Hanbo Sang, Zelin Ni, Huanyu He, Xuesong Gao, Qihao Sun, Sihai Zhang, Supavadee Aramvith, Weiyao Lin |
ICME | 10 |
| 2022 | Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action RecognitionabstractA primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algorithms at the feature level while little attention is paid to processing input video data. Moreover, existing frame sampling strategies may omit critical action information in temporal and spatial dimensions, which further impacts video utilization efficiency. In this paper, we propose a novel video frame sampler for few-shot action recognition to address this issue, where task-specific spatial-temporal frame sampling is achieved via a temporal selector (TS) and a spatial amplifier (SA). Specifically, our sampler first scans the whole video at a small computational cost to obtain a global perception of video frames. The TS plays its role in selecting top-T frames that contribute most significantly and subsequently. The SA emphasizes the discriminative information of each frame by amplifying critical regions with the guidance of saliency maps. We further adopt task-adaptive learning to dynamically adjust the sampling strategy according to the episode task at hand. Both the implementations of TS and SA are differentiable for end-to-end optimization, facilitating seamless integration of our proposed sampler with most few-shot action recognition methods. Extensive experiments show a significant boost in the performances on various benchmarks including long-term videos. Huabin Liu 0001, Weixian Lv, John See, Weiyao Lin |
ACM Multimedia | 4 |
| 2022 | Exploring the Semi-Supervised Video Object Segmentation Problem from a Cyclic Perspective
Yuxi Li 0009, Ning Xu 0007, John See, Weiyao Lin |
Int. J. Comput. Vis. | 5 |
| 2022 | Class-Aware Sounding Objects Localization via Audiovisual CorrespondenceabstractAudiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without category annotations, i.e., localizing the sounding object and recognizing its category. To address this problem, we propose a two-stage step-by-step learning framework to localize and recognize sounding objects in complex audiovisual scenarios using only the correspondence between audio and vision. First, we propose to determine the sounding area via coarse-grained audiovisual correspondence in the single source cases. Then visual features in the sounding area are leveraged as candidate object representations to establish a category-representation object dictionary for expressive visual character extraction. We generate class-aware object localization maps in cocktail-party scenarios and use audiovisual correspondence to suppress silent areas by referring to this dictionary. Finally, we employ category-level audiovisual consistency as the supervision to achieve fine-grained audio and sounding object distribution alignment. Experiments on both realistic and synthesized videos show that our model is superior in localizing and recognizing objects as well as filtering out silent ones. We also transfer the learned audiovisual network into the unsupervised object detection task, obtaining reasonable performance. Di Hu 0001, Yake Wei, Rui Qian 0001, Weiyao Lin, Ruihua Song, Ji-Rong Wen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Sequential Attention-Based Distinct Part Modeling for Balanced Pedestrian DetectionabstractDespite pedestrian detectors having made significant progress by introducing convolutional neural networks, their performance still suffers degradation, especially in occlusion scenes with more false positives (FPs) and false negatives (FNs). To alleviate the problem, we propose a novel Sequential Attention-based Distinct Part Modeling (SA-DPM) for balanced pedestrian detection. It takes one step further in constructing more robust representation that supports detection with fewer FNs and FPs. Specifically, the Sequential Attention serves as one internal perception process that captures several distinct part areas step by step from each pedestrian proposal (full-body). Different from the previous either-or feature selection, the following Joint Learning attempts to seek a reasonable trade-off between part and full-body features, and combines both features for more accurate classification and regression. Evaluation on the widely used pedestrian datasets including Caltech and Citypersons shows that the proposed SA-DPM achieves promising performance for both non-occluded and occluded pedestrian detection tasks, especially on Caltech Heavy Occlusion set, which yields a new state-of-the-art miss rate by 30.18% and outperforms the second best detector by 6.32%. Yan Luo 0003, Weiyao Lin, Xiaokang Yang 0001, Jun Sun 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Learning Scale-Consistent Attention Part Network for Fine-Grained Image RecognitionabstractDiscriminative region localization and feature learning are crucial for fine-grained visual recognition. Existing approaches solve this issue by attention mechanism or part based methods while neglecting consistency between attention and local parts, as well as the rich relation information among parts. This paper proposes a Scale-consistent Attention Part Network (SCAPNet) to address that issue, which seamlessly integrates three novel modules: grid gate attention unit (gGAU), scale-consistent attention part selection (SCAPS), and part relation modeling (PRM). The gGAU module represents the grid region at a certain fine-scale with middle layer CNN features and produces hard attention maps with the lightweight Gumbel-Max based gate. The SCAPS module utilizes attention to guide part selection across multi-scales and keep the selection scale-consistent. The PRM module utilizes the self-attention mechanism to build the relationship among parts based on their appearance and relative geo-positions. SCAPNet can be learned in an end-to-end way and demonstrates state-of-the-art accuracy on several publicly available fine-grained recognition datasets (CUB-200-2011, FGVC-Aircraft, Veg200, and Fru92). Huabin Liu 0001, John See, Weiyao Lin |
IEEE Trans. Multim. | 5 |
| 2022 | Uni-EDEN: Universal Encoder-Decoder Network by Multi-Granular Vision-Language Pre-trainingabstractVision-language pre-training has been an emerging and fast-developing research topic, which transfers multi-modal knowledge from rich-resource pre-training task to limited-resource downstream tasks. Unlike existing works that predominantly learn a single generic encoder, we present a pre-trainable Universal Encoder-DEcoder Network (Uni-EDEN) to facilitate both vision-language perception (e.g., visual question answering) and generation (e.g., image captioning). Uni-EDEN is a two-stream Transformer-based structure, consisting of three modules: object and sentence encoders that separately learns the representations of each modality and sentence decoder that enables both multi-modal reasoning and sentence generation via inter-modal interaction. Considering that the linguistic representations of each image can span different granularities in this hierarchy including, from simple to comprehensive, individual label, a phrase, and a natural sentence, we pre-train Uni-EDEN through multi-granular vision-language proxy tasks: Masked Object Classification, Masked Region Phrase Generation, Image-Sentence Matching, and Masked Sentence Generation. In this way, Uni-EDEN is endowed with the power of both multi-modal representation extraction and language modeling. Extensive experiments demonstrate the compelling generalizability of Uni-EDEN by fine-tuning it to four vision-language perception and generation downstream tasks. Yehao Li, Yingwei Pan, Ting Yao 0003, Weiyao Lin, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Temporal Alignment via Event Boundary for Few-shot Action Recongnition
Shuyuan Li, Huabin Liu 0001, Mengjuan Fei, Xiaoyuan Yu, Weiyao Lin |
BMVC | 5 |
| 2021 | Variational Pedestrian DetectionabstractPedestrian detection in a crowd is a challenging task due to a high number of mutually-occluding human instances, which brings ambiguity and optimization difficulties to the current IoU-based ground truth assignment procedure in classical object detection methods. In this paper, we develop a unique perspective of pedestrian detection as a variational inference problem. We formulate a novel and efficient algorithm for pedestrian detection by modeling the dense proposals as a latent variable while proposing a customized Auto-Encoding Variational Bayes (AEVB) algorithm. Through the optimization of our proposed algorithm, a classical detector can be fashioned into a variational pedestrian detector. Experiments conducted on CrowdHuman and CityPersons datasets show that the proposed algorithm serves as an efficient solution to handle the dense pedestrian detection problem for the case of single-stage detectors. Our method can also be flexibly applied to two-stage detectors, achieving notable performance enhancement. Huanyu He, Yuxi Li 0009, John See, Weiyao Lin |
CVPR | 6 |
| 2021 | Multi-Level Curriculum for Training A Distortion-Aware Barrel Distortion Rectification ModelabstractBarrel distortion rectification aims at removing the radial distortion in a distorted image captured by a wide-angle lens. Previous deep learning methods mainly solve this problem by learning the implicit distortion parameters or the nonlinear rectified mapping function in a direct manner. However, this type of manner results in an indistinct learning process of rectification and thus limits the deep perception of distortion. In this paper, inspired by the curriculum learning, we analyze the barrel distortion rectification task in a progressive and meaningful manner. By considering the relationship among different construction levels in an image, we design a multi-level curriculum that disassembles the rectification task into three levels, structure recovery, semantics embedding, and texture rendering. With the guidance of the curriculum that corresponds to the construction of images, the proposed hierarchical architecture enables a progressive rectification and achieves more accurate results. Moreover, we present a novel distortion-aware pre-training strategy to facilitate the initial learning of neural networks, promoting the model to converge faster and better. Experimental results on the synthesized and real-world distorted image datasets show that the proposed approach significantly outperforms other learning methods, both qualitatively and quantitatively. Kang Liao, Chunyu Lin, Lixin Liao, Yao Zhao 0001, Weiyao Lin |
ICCV | 5 |
| 2021 | Enhancing Self-supervised Video Representation Learning via Multi-level Feature OptimizationabstractThe crux of self-supervised video representation learning is to build general features from unlabeled videos. However, most recent works have mainly focused on high-level semantics and neglected lower-level representations and their temporal relationship which are crucial for general video understanding. To address these challenges, this paper proposes a multi-level feature optimization framework to improve the generalization and temporal modeling ability of learned video representations. Concretely, high-level features obtained from naive and prototypical contrastive learning are utilized to build distribution graphs, guiding the process of low-level and mid-level feature learning. We also devise a simple temporal modeling module from multi-level features to enhance motion pattern learning. Experiments demonstrate that multi-level feature optimization with the graph constraint and temporal modeling can greatly improve the representation ability in video understanding. Code is available$here$. Rui Qian 0001, Yuxi Li 0009, Huabin Liu 0001, John See, Shuangrui Ding, Weiyao Lin |
ICCV | 8 |
| 2021 | End-to-End Video Instance Segmentation via Spatial-Temporal Graph Neural NetworksabstractVideo instance segmentation is a challenging task that extends image instance segmentation to the video domain. Existing methods either rely only on single-frame information for the detection and segmentation subproblems or handle tracking as a separate post-processing step, which limit their capability to fully leverage and share useful spatial-temporal information for all the subproblems. In this paper, we propose a novel graph-neural-network (GNN) based method to handle the aforementioned limitation. Specifically, graph nodes representing instance features are used for detection and segmentation while graph edges representing instance relations are used for tracking. Both inter and intra-frame information is effectively propagated and shared via graph updates and all the subproblems (i.e. detection, segmentation and tracking) are jointly optimized in an unified framework. The performance of our method shows great improvement on the YoutubeVIS validation dataset compared to existing methods and achieves 36.5% AP with a ResNet-50 backbone, operating at 22 FPS. Tao Wang 0002, Ning Xu 0007, Kean Chen, Weiyao Lin |
ICCV | 4 |
| 2021 | SiamRCR: Reciprocal Classification and Regression for Visual Object TrackingabstractRecently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classification confidence as the final prediction. This strategy may miss the right result due to the accuracy misalignment between classification and regression. In this paper, we propose a novel siamese tracking algorithm called SiamRCR, addressing this problem with a simple, light and effective solution. It builds reciprocal links between classification and regression branches, which can dynamically re-weight their losses for each positive sample. In addition, we add a localization branch to predict the localization accuracy, so that it can work as the replacement of the regression assistance link during inference. This branch makes the training and inference more consistent. Extensive experimental results demonstrate the effectiveness of SiamRCR and its superiority over the state-of-the-art competitors on GOT-10k, LaSOT, TrackingNet, OTB-2015, VOT-2018 and VOT-2019. Moreover, our SiamRCR runs at 65 FPS, far above the real-time requirement. Jinlong Peng, Zhengkai Jiang 0001, Yueyang Gu, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Weiyao Lin |
IJCAI | 8 |
| 2021 | LSTC: Boosting Atomic Action Detection with Long-Short-Term ContextabstractIn this paper, we place the atomic action detection problem intoa Long-Short Term Context (LSTC) to analyze how the temporalreliance among video signals affect the action detection results. Todo this, we decompose the action recognition pipeline into short-term and long-term reliance, in terms of the hypothesis that the twokinds of context are conditionally independent given the objectiveaction instance. Within our design, a local aggregation branch isutilized to gather dense and informative short-term cues, while ahigh order long-term inference branch is designed to reason theobjective action class from high-order interaction between actor andother person or person pairs. Both branches independently predictthe context-specific actions and the results are merged in the end.We demonstrate that both temporal grains are beneficial to atomicaction recognition. On the mainstream benchmarks of atomic actiondetection, our design can bring significant performance gain fromthe existing state-of-the-art pipeline. Yuxi Li 0009, Boshen Zhang, Jian Li 0062, Yabiao Wang, Weiyao Lin, Chengjie Wang 0001, Feiyue Huang |
ACM Multimedia | 5 |
| 2021 | A regional distance regression network for monocular object distance estimation
Lianghui Ding, Yuxi Li 0009, Weiyao Lin, Mingbi Zhao, Xiaoyuan Yu, Yunlong Zhan |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | AP-Loss for Accurate One-Stage Object DetectionabstractOne-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework to replace the classification task in one-stage detectors with a ranking task, and adopting the average-precision loss (AP-loss) for the ranking problem. Due to its non-differentiability and non-convexity, the AP-loss cannot be optimized directly. For this purpose, we develop a novel optimization algorithm, which seamlessly combines the error-driven update scheme in perceptron learning and backpropagation algorithm in deep networks. We provide in-depth analyses on the good convergence property and computational complexity of the proposed algorithm, both theoretically and empirically. Experimental results demonstrate notable improvement in addressing the imbalance issue in object detection over existing AP-based optimization algorithms. An improved state-of-the-art performance is achieved in one-stage detectors based on AP-loss over detectors using classification-losses on various standard benchmarks. The proposed framework is also highly versatile in accommodating different network architectures. Code is available at https://github.com/cccorn/AP-loss. Kean Chen, Weiyao Lin, John See, Junni Zou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Group Reidentification with Multigrained Matching and IntegrationabstractThe task of reidentifying groups of people under different camera views is an important yet less-studied problem. Group reidentification (Re-ID) is a very challenging task since it is not only adversely affected by common issues in traditional single-object Re-ID problems, such as viewpoint and human pose variations, but also suffers from changes in group layout and group membership. In this paper, we propose a novel concept of group granularity by characterizing a group image by multigrained objects: individual people and subgroups of two and three people within a group. To achieve robust group Re-ID, we first introduce multigrained representations which can be extracted via the development of two separate schemes, that is, one with handcrafted descriptors and another with deep neural networks. The proposed representation seeks to characterize both appearance and spatial relations of multigrained objects, and is further equipped with importance weights which capture variations in intragroup dynamics. Optimal group-wise matching is facilitated by a multiorder matching process which, in turn, dynamically updates the importance weights in iterative fashion. We evaluated three multicamera group datasets containing complex scenarios and large dynamics, with experimental results demonstrating the effectiveness of our approach. Weiyao Lin, Yuxi Li 0009, John See, Junni Zou, Hongkai Xiong, Jingdong Wang 0001, Tao Mei 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Group Re-Identification With Group Context Graph Neural NetworksabstractGroup re-identification aims to match groups of people across disjoint cameras. In this task, the contextual information from neighbor individuals can be exploited for re-identifying each individual within the group as well as the entire group. However, compared with single person re-identification, it brings new challenges including group layout and group membership changes. Motivated by the observation that individuals who are close together are more likely to keep in the same group under different cameras than those who are far apart, we propose to model each group as a spatial K-nearest neighbor graph (SKNNG) and design a group context graph neural network (GCGNN) for graph representation learning. Specifically, for each node in the graph, the proposed GCGNN learns an embedding which aggregates the contextual information from neighbor nodes. We design multiple weighting kernels for neighborhood aggregation based on the graph properties including node in-degrees and spatial relationship attributes. We compute the similarity scores between node embeddings of two graphs for group member association and obtain the matching score between the two graphs by summing up the similarity scores of all linked node pairs. Experimental results on three public datasets show that our approach performs favorably against state-of-the-art methods and achieves high efficiency. Ji Zhu 0002, Hua Yang 0001, Weiyao Lin, Nian Liu 0002, Jia Wang 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Finding Action Tubes with a Sparse-to-Dense FrameworkabstractThe task of spatial-temporal action detection has attracted increasing researchers. Existing dominant methods solve this problem by relying on short-term information and dense serial-wise detection on each individual frames or clips. Despite their effectiveness, these methods showed inadequate use of long-term information and are prone to inefficiency. In this paper, we propose for the first time, an efficient framework that generates action tube proposals from video streams with a single forward pass in a sparse-to-dense manner. There are two key characteristics in this framework: (1) Both long-term and short-term sampled information are explicitly utilized in our spatio-temporal network, (2) A new dynamic feature sampling module (DTS) is designed to effectively approximate the tube output while keeping the system tractable. We evaluate the efficacy of our model on the UCF101-24, JHMDB-21 and UCFSports benchmark datasets, achieving promising results that are competitive to state-of-the-art methods. The proposed sparse-to-dense strategy rendered our framework about 7.6 times more efficient than the nearest competitor. Yuxi Li 0009, Weiyao Lin, Tao Wang 0002, John See, Rui Qian 0001, Ning Xu 0007, Limin Wang 0002, Shugong Xu |
AAAI | 2 |
| 2020 | PIoU Loss: Towards Accurate Oriented Object Detection in Complex Environments
Kean Chen, Weiyao Lin, John See, Yan Ke |
ECCV (5) | 3 |
| 2020 | CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization
Yuxi Li 0009, Weiyao Lin, John See, Ning Xu 0007, Shugong Xu, Yan Ke |
ECCV (16) | 2 |
| 2020 | Multiple Sound Sources Localization from Coarse to Fine
Rui Qian 0001, Di Hu 0001, Heinrich Dinkel, Mengyue Wu, Ning Xu 0007, Weiyao Lin |
ECCV (20) | 6 |
| 2020 | TRP: Trained Rank Pruning for Efficient Deep Neural NetworksabstractTo enable DNNs on edge devices like mobile phones, low-rank approximation has been widely adopted because of its solid theoretical rationale and efficient implementations. Several previous works attempted to directly approximate a pre-trained model by low-rank decomposition; however, small approximation errors in parameters can ripple over a large prediction loss. As a result, performance usually drops significantly and a sophisticated effort on fine-tuning is required to recover accuracy. Apparently, it is not optimal to separate low-rank approximation from training. Unlike previous works, this paper integrates low rank approximation and regularization into the training process. We propose Trained Rank Pruning (TRP), which alternates between low rank approximation and training. TRP maintains the capacity of the original network while imposing low-rank constraints during training. A nuclear regularization optimized by stochastic sub-gradient descent is utilized to further promote low rank in TRP. The TRP trained network inherently has a low-rank structure, and is approximated with negligible performance loss, thus eliminating the fine-tuning process after low rank decomposition. The proposed method is comprehensively evaluated on CIFAR-10 and ImageNet, outperforming previous compression methods using low rank approximation. Yuhui Xu 0002, Yuxi Li 0009, Shuai Zhang 0009, Wei Wen 0003, Yingyong Qi, Yiran Chen 0001, Weiyao Lin, Hongkai Xiong |
IJCAI | 8 |
| 2020 | ATRW: A Benchmark for Amur Tiger Re-identification in the WildabstractMonitoring the population and movements of endangered species is an important task to wildlife conversation. Traditional tagging methods do not scale to large populations, while applying computer vision methods to camera sensor data requires re-identification (re-ID) algorithms to obtain accurate counts and moving trajectory of wildlife. However, existing re-ID methods are largely targeted at persons and cars, which have limited pose variations and constrained capture environments. This paper tries to fill the gap by introducing a novel large-scale dataset, the Amur Tiger Re-identification in the Wild (ATRW) dataset. ATRW contains over 8,000 video clips from 92 Amur tigers, with bounding box, pose keypoint, and tiger identity annotations. In contrast to typical re-ID datasets, the tigers are captured in a diverse set of unconstrained poses and lighting conditions. We demonstrate with a set of baseline algorithms that ATRW is a challenging dataset for re-ID. Lastly, we propose a novel method for tiger re-identification, which introduces precise pose parts modeling in deep neural networks to handle large pose variation of tigers, and reaches notable performance improvement over existing re-ID methods. The ATRW dataset is public available at https://cvwc2019.github.io/challenge.html Shuyuan Li, Rui Qian 0001, Weiyao Lin |
ACM Multimedia | 5 |
| 2020 | Discriminative Sounding Objects Localization via Self-supervised Audiovisual MatchingabstractDiscriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised class-aware sounding object localization. First, we propose to learn robust object representations by aggregating the candidate sound localization results in the single source scenes. Then, class-aware object localization maps are generated in the cocktail-party scenarios by referring the pre-learned object knowledge, and the sounding objects are accordingly selected by matching audio and visual object category distributions, where the audiovisual consistency is viewed as the self-supervised signal. Experimental results in both realistic and synthesized cocktail-party videos demonstrate that our model is superior in filtering out silent objects and pointing out the location of sounding objects of different classes. Code is available at https://github.com/DTaoo/Discriminative-Sounding-Objects-Localization. Di Hu 0001, Rui Qian 0001, Minyue Jiang, Xiao Tan 0001, Shilei Wen, Errui Ding, Weiyao Lin, Dejing Dou |
NeurIPS | 7 |
| 2020 | Delving into the Cyclic Mechanism in Semi-supervised Video Object SegmentationabstractIn this paper, we take attempt to incorporate the cyclic mechanism with the vision task of semi-supervised video object segmentation. By resorting to the accurate reference mask of the first frame, we try to mitigate the error propagation problem in most of current video object segmentation pipelines. Firstly, we propose a cyclic scheme for offline training of segmentation networks. Then, we extend the offline pipeline to an online method by introducing a simple gradient correction module while keeping high efficiency as other offline methods. Finally we develop cycle effective receptive field (cycle-ERF) from gradient correction to provide a new perspective for analyzing object-specific regions of interests. We conduct comprehensive experiments on benchmarks of DAVIS17 and Youtube-VOS, demonstrating that our introduced cyclic mechanism is helpful to boost the segmentation quality. Yuxi Li 0009, Ning Xu 0007, Jinlong Peng, John See, Weiyao Lin |
NeurIPS | 5 |
| 2020 | TPM: Multiple object tracking with tracklet-plane matching
Jinlong Peng, Tao Wang 0002, Weiyao Lin, Jian Wang 0066, John See, Shilei Wen, Errui Ding |
Pattern Recognit. | 3 |
| 2020 | Adaptive lossless compression of skeleton sequences
Weiyao Lin, Tushar Shankar Shinde, Wenrui Dai, Mingzhou Liu 0001, Xiaoyi He, Anil Kumar Tiwari, Hongkai Xiong |
Signal Process. Image Commun. | 1 |
| 2020 | Background foreground boundary aware efficient motion search for surveillance videos
Tushar Shankar Shinde, Anil Kumar Tiwari, Weiyao Lin, Liquan Shen |
Signal Process. Image Commun. | 3 |
| 2020 | Partition-Aware Adaptive Switching Neural Networks for Post-Processing in HEVCabstractThis article addresses neural network based post-processing for the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC). We first propose a partition-aware convolution neural network (CNN) that utilizes the partition information produced by the encoder to assist in the post-processing. In contrast to existing CNN-based approaches, which only take the decoded frame as input, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the artifacts introduced by HEVC are efficiently reduced. We further introduce an adaptive-switching neural network (ASN) that consists of multiple independent CNNs to adaptively handle the variations in content and distortion within compressed-video frames, providing further reduction in visual artifacts. Additionally, an iterative training procedure is proposed to train these independent CNNs attentively on different local patch-wise classes. Experiments on benchmark sequences demonstrate the effectiveness of our partition-aware and adaptive-switching neural networks. Weiyao Lin, Xiaoyi He, Xintong Han, Dong Liu 0002, John See, Junni Zou, Hongkai Xiong, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Towards Accurate One-Stage Object Detection With AP-LossabstractOne-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework to replace the classification task in one-stage detectors with a ranking task, and adopting the Average-Precision loss (AP-loss) for the ranking problem. Due to its non-differentiability and non-convexity, the AP-loss cannot be optimized directly. For this purpose, we develop a novel optimization algorithm, which seamlessly combines the error-driven update scheme in perceptron learning and backpropagation algorithm in deep networks. We verify good convergence property of the proposed algorithm theoretically and empirically. Experimental results demonstrate notable performance improvement in state-of-the-art one-stage detectors based on AP-loss over different kinds of classification-losses on various benchmarks, without changing the network architectures. Kean Chen, Weiyao Lin, John See, Ling-Yu Duan, Zhibo Chen 0006, Changwei He, Junni Zou |
CVPR | 3 |
| 2019 | DNQ: Dynamic Network QuantizationabstractIn this paper, we propose a Dynamic Network Quantization (DNQ) framework. Unlike most existing quantization methods that use a universal quantization bit-width for the whole network, we utilize policy gradient [1] to train an agent to learn the bit-width of each layer by the bit-width controller. Yuhui Xu 0002, Shuai Zhang 0009, Yingyong Qi, Jiaxian Guo, Weiyao Lin, Hongkai Xiong |
DCC | 5 |
| 2019 | Dual-stream Shallow Networks for Facial Micro-expression RecognitionabstractMicro-expressions are spontaneous, brief and subtle facial muscle movements that exposes underlying emotions. Motivated by recent exploits into deep learning for micro-expression analysis, we propose a lightweight dual-stream shallow network in the form of a pair of truncated CNNs with heterogeneous input features. The merging of the convolutional features allows for discriminative learning of micro-expression classes stemming from both streams. Using activation heatmaps, we further demonstrate that salient facial areas are well emphasized, and correspond closely to relevant action units belonging to emotion classes. We empirically validate the proposed network on three benchmark databases, obtaining state-of-the-art performance on the CASME II and SAMM while remaining competitive on the SMIC. Further observations point towards the sufficiency of utilizing shallower deep networks for micro-expression recognition. Huai-Qian Khor, John See, Sze-Teng Liong, Raphael C.-W. Phan, Weiyao Lin |
ICIP | 5 |
| 2019 | Adaptive Hard Example Mining for Image CaptioningabstractReinforcement Learning (RL) based methods optimize evaluation metric directly in image captioning task. In these methods, metric scores of captions are regard as rewards for examples. However, existing methods suffer from inferior performances on hard examples. In this paper, we propose an adaptive hard example mining method with additional supervised training for image captioning. Beam search algorithm is leveraged to estimate score expectation for each example. Examples whose caption scores are lower than expectation are selected automatically. For the selected hard examples, we propose an additional reward policy for high-scoring captions to force model learning from them. The proposed method is hyper-parameter free without tuning. Experimental results on MSCOCO dataset validate effectiveness of the proposed method. Yongzhuang Wang, Yangmei Shen, Hongkai Xiong, Weiyao Lin |
ICIP | 4 |
| 2019 | Localization Guided Fight Action Detection in Surveillance VideosabstractAutomatic detection of fight behaviors in surveillance videos is an important task for surveillance systems. In this work, we propose a novel localization guided framework for detecting fight actions in surveillance videos. Specifically, we exploit optical flow maps to extract motion activation information, which indicates the location of active regions. Then, a detection guided alignment module is designed to adjust the localized active regions. This approach employs a two-stream based 3D convolution network as the backbone network with a novel motion acceleration representation on the temporal stream. While most existing methods are still evaluated on three benchmark datasets which were not originally collected from surveillance scenarios, we present a novel Fight Action Detection in Surveillance-videos (FADS) dataset for this purpose. With a total of 1,520 video clips, the FADS is the largest known dataset in terms of number of surveillance videos with fight scenes. Experimental results on both the benchmark datasets and the FADS show that our proposed localization guided method outperforms state-of-the-art techniques. Qichao Xu, John See, Weiyao Lin |
ICME | 3 |
| 2019 | ThiNet: Pruning CNN Filters for a Thinner NetabstractThis paper aims at accelerating and compressing deep neural networks to deploy CNN models into small devices like mobile phones or embedded gadgets. We focus on filter level pruning, i.e., the whole filter will be discarded if it is less important. An effective and unified framework, ThiNet (stands for "Thin Net"), is proposed in this paper. We formally establish filter pruning as an optimization problem, and reveal that we need to prune filters based on statistics computed from its next layer, not the current layer, which differentiates ThiNet from existing methods. We also propose "gcos" (Group COnvolution with Shuffling), a more accurate group convolution scheme, to further reduce the pruned model size. Experimental results demonstrate the effectiveness of our method, which has advanced the state-of-the-art. Moreover, we show that the original VGG-16 model can be compressed into a very small model (ThiNet-Tiny) with only 2.66 MB model size, but still preserve AlexNet level accuracy. This small model is evaluated on several benchmarks with different vision tasks (e.g., classification, detection, segmentation), and shows excellent generalization ability. Jian-Hao Luo, Hao Zhang 0038, Chen-Wei Xie, Jianxin Wu 0001, Weiyao Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Weak to Strong Detector Learning for Simultaneous Classification and LocalizationabstractThis paper aims at learning discriminative part detectors with only image-level labels. To this end, we need to develop effective technologies for both pattern mining and detection learning. Different from previous methods, which train part detectors in one step, we divide the detector learning process into two stages and formulate it as a weak to strong learning framework. In particular, we first learn exemplar detectors from the unaligned patterns and perform a detector-based spectral clustering to produce weak detectors that are only responsible for a few discriminative patterns. In this way, the weak detectors are able to offer right initial patterns for strong detector learning. Second, we learn strong detectors with patterns discovered from the weak detectors, which we formulate as a confidence-loss sparse multiple instance learning (cls-MIL) task. The cls-MIL considers the diversity of positive samples while avoiding drifting away from the well localized ones by assigning a confidence value to each positive sample. The responses of the learned detectors produce an effective mid-level image representation for both image classification and object localization. Experiments conducted on benchmark data sets well demonstrate the superiority of our method over existing approaches. Xiaopeng Zhang 0008, Hongkai Xiong, Weiyao Lin, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Deep Color Guided Coarse-to-Fine Convolutional Network Cascade for Depth Image Super-ResolutionabstractDepth image super-resolution is a significant yet challenging task. In this paper, we introduce a novel deep color guided coarse-to-fine convolutional neural network (CNN) framework to address this problem. First, we present a datadriven filter method to approximate the ideal filter for depth image super-resolution instead of hand-designed filters. Based on large data samples, the filter learned is more accurate and stable for upsampling depth image. Second, we introduce a coarse-to-fine CNN to learn different sizes of filter kernels. In coarse stage, larger filter kernels are learned by CNN to achieve crude high-resolution depth image. As to fine stage, the crude high-resolution depth image is used as the input so that smaller filter kernels are learned to gain more accurate results. Benefit from this network, we can progressively recover the high frequency details. Third, we construct a color guidance strategy that fuses color difference and spatial distance for depth image upsampling. We revise the interpolated high-resolution depth image according to the corresponding pixels in highresolution color maps. Guided by color information, the depth of high-resolution image obtained can alleviate texture copying artifacts and preserve edge details effectively. Quantitative and qualitative experimental results demonstrate our state-of-the-art performance for depth map super-resolution. Bin Sheng 0001, Ping Li 0016, Weiyao Lin, David Dagan Feng |
IEEE Trans. Image Process. | 4 |
| 2018 | Action Recognition With Coarse-to-Fine Deep Feature Integration and Asynchronous FusionabstractAction recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deep-based framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for representing actions, and 2) reducing the asynchrony between different information streams. We first introduce a coarse-to-fine network which extracts shared deep features at different action class granularities and progressively integrates them to obtain a more accurate feature representation for input actions. We further introduce an asynchronous fusion network. It fuses information from different streams by asynchronously integrating stream-wise features at different time points, hence better leveraging the complementary information in different streams. Experimental results on action recognition benchmarks demonstrate that our approach achieves the state-of-the-art performance. Weiyao Lin, Ke Lu 0002, Bin Sheng 0001, Jianxin Wu 0001, Bingbing Ni, Hongkai Xiong |
AAAI | 1 |
| 2018 | Deep Neural Network Compression With Single and Multiple Level QuantizationabstractNetwork quantization is an effective solution to compress deep neural networks for practical usage. Existing network quantization methods cannot sufficiently exploit the depth information to generate low-bit compressed network. In this paper, we propose two novel network quantization approaches, single-level network quantization (SLQ) for high-bit quantization and multi-level network quantization (MLQ) for extremely low-bit quantization (ternary). We are the first to consider the network quantization from both width and depth level. In the width level, parameters are divided into two parts: one for quantization and the other for re-training to eliminate the quantization loss. SLQ leverages the distribution of the parameters to improve the width level. In the depth level, we introduce incremental layer compensation to quantize layers iteratively which decreases the quantization loss in each iteration. The proposed approaches are validated with extensive experiments based on the state-of-the-art neural networks including AlexNet, VGG-16, GoogleNet and ResNet-18. Both SLQ and MLQ achieve impressive results. Yuhui Xu 0002, Yongzhuang Wang, Aojun Zhou, Weiyao Lin, Hongkai Xiong |
AAAI | 4 |
| 2018 | Network Decoupling: From Regular to Depthwise Separable Convolutions
Jianbo Guo, Yuxi Li 0009, Weiyao Lin, Yurong Chen 0001 |
BMVC | 3 |
| 2018 | Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
Yuxi Li 0009, Jiuwei Li, Weiyao Lin |
BMVC | 3 |
| 2018 | Enriched Long-Term Recurrent Convolutional Network for Facial Micro-Expression RecognitionabstractFacial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved superior performance in micro-expression recognition but at the cost of domain specificity and cumbersome parametric tunings. In this paper, we propose an Enriched Long-term Recurrent Convolutional Network (ELRCN) that first encodes each micro-expression frame into a feature vector through CNN module(s), then predicts the micro-expression by passing the feature vector through a Long Short-term Memory (LSTM) module. The framework contains 2 different network variants: (1) Channel-wise stacking of input data for spatial enrichment, (2) Feature-wise stacking of features for temporal enrichment. We demonstrate that the proposed approach is able to achieve reasonably good performance, without data augmentation. In addition, we also present ablation studies conducted on the framework and visualizations of what CNN "sees" when predicting the micro-expression classes. Huai-Qian Khor, John See, Raphael C.-W. Phan, Weiyao Lin |
FG | 4 |
| 2018 | Enhancing HEVC Compressed Videos with a Partition-Masked Convolutional Neural NetworkabstractIn this paper, we propose a partition-masked Convolution Neural Network (CNN) to achieve compressed-video enhancement for the state-of-the-art coding standard, High Efficiency Video Coding (HECV). More precisely, our method utilizes the partition information produced by the encoder to guide the quality enhancement process. In contrast to existing CNN-based approaches, which only take the decoded frame as the input to the CNN, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the degradation introduced by HEVC is reduced more efficiently. Experimental results show that our approach leads to over 9.76% BD-rate saving on benchmark sequences, which achieves the state-of-the-art performance. Xiaoyi He, Qiang Hu 0003, Xiaoyun Zhang 0001, Weiyao Lin, Xintong Han |
ICIP | 5 |
| 2018 | Key Joints Selection and Spatiotemporal Mining for Skeleton-Based Action RecognitionabstractTrajectories and spatiotemporal attention model have been successfully used in skeleton-based action recognition. Most existing methods focus more attention on temporal structure mining. However, only a few local joints and their position features (e.g., critical position changes of hand, head, leg etc.) are responsible for the action label. In this work, we introduce a novel action recognition framework using Key Joints Selection and Spatiotemporal Mining, which can identify both key joints and their position & velocity histogram as well as trajectory features for action classification. First, histogram of human joints position and velocity are developed to enhance the spatiotemporal structure representation of existing trajectory-based methods. Second, the key joints are selected according to their information gains, and then their position & velocity histograms are weighted and composed with trajectory features to form one richer representation for final action classification. Experiments on two widely-tested benchmark datasets show that by combining the strength of both richer features and key joints selecting, our method can achieve state-of-the-art or competitive performance compared with existing results using sophisticated models such as deep learning, with advantages regarding the recognition accuracy and robustness. Wu Luo, Weiyao Lin |
ICIP | 4 |
| 2018 | Group Re-Identification: Leveraging and Integrating Multi-Grain InformationabstractThis paper addresses an important yet less-studied problem: re-identifying groups of people in different camera views. Group re-identification (Re-ID) is very challenging since it is not only interfered by view-point and human pose variations in the traditional single-object Re-ID tasks, but also suffers from group layout and group member variations. To handle these issues, we propose to leverage the information of multi-grain objects: individual person and subgroups of two and three people inside a group image. We compute multi-grain representations to characterize the appearance and spatial features of multi-grain objects and evaluate the importance weight of each object for group Re-ID, so as to handle the interferences from group dynamics. We compute the optimal group-wise matching by using a multi-order matching process based on the multi-grain representation and importance weights. Furthermore, we dynamically update the importance weights according to the current matching results and then compute a new optimal group-wise matching. The two steps are iteratively conducted, yielding the final matching results.Experimental results on various datasets demonstrate the effectiveness of our approach. Weiyao Lin, Bin Sheng 0001, Ke Lu 0002, Junchi Yan, Jingdong Wang 0001, Errui Ding, Hongkai Xiong |
ACM Multimedia | 2 |
| 2018 | Multi-scale Spatiotemporal Information Fusion Network for Video Action RecognitionabstractTwo-stream convolutional networks have shown excellent performance in video action recognition in recent years. However, it remains unclear how to model the correlation between the temporal and spatial streams more effectively. First, the spatial stream and temporal stream pay attention to different aspects, which can lead to different recognition results. Second, the variety in the length of optical flow fields tends to have a great impact on the classification results. In this paper, we propose a novel multi-scale spatiotemporal information fusion network to fuse the spatial and temporal features. Specifically, our network takes advantage of multi-scale temporal information to better utilize the motion cues. Considering the complementary relationship between the spatial and temporal features, we take the hierarchical fusion strategies and asynchronous fusion method to fuse the two-stream features. Experimental results on two benchmark datasets (UCF101 and HMDB51) show that the proposed network achieves competitive performance. Yutong Cai, Weiyao Lin, John See, Ming-Ming Cheng, Guangcan Liu, Hongkai Xiong |
VCIP | 2 |
| 2018 | Tracklet Siamese Network with Constrained Clustering for Multiple Object TrackingabstractMultiple object tracking (MOT) is an important yet challenging task in video understanding and analysis. Basically, MOT aims to associate detected objects into trajectories based on their temporal relationships. The occlusion among moving objects poses a major challenge towards robust modeling of these relationships. In this paper, we propose a novel Tracklet Siamese Network (TSN) for learning similarities between track-lets characterized by appearance information, achieving superior performance on two MOTChallenge benchmark datasets. Our framework constructs short tracklets from highly-related object detections by excluding inaccurate object detections. We also adopt a constrained clustering technique to piece tracklets together into long trajectories, thus recovering many missing detections caused by original detector or the detection removing in the previous step. Comparisons against state-of-the-art methods were reported while ablation studies further substantiate the viability of components in our approach. Jinlong Peng, Fan Qiu, John See, Shaoshuai Huang, Ling-Yu Duan, Weiyao Lin |
VCIP | 7 |
| 2017 | Fractal Dimension Invariant Filtering and Its CNN-Based ImplementationabstractFractal analysis has been widely used in computer vision, especially in texture image processing and texture analysis. The key concept of fractal-based image model is the fractal dimension, which is invariant to bi-Lipschitz transformation of image, and thus capable of representing intrinsic structural information of image robustly. However, the invariance of fractal dimension generally does not hold after filtering, which limits the application of fractal-based image model. In this paper, we propose a novel fractal dimension invariant filtering (FDIF) method, extending the invariance of fractal dimension to filtering operations. Utilizing the notion of local self-similarity, we first develop a local fractal model for images. By adding a nonlinear post-processing step behind anisotropic filter banks, we demonstrate that the proposed filtering method is capable of preserving the local invariance of the fractal dimension of image. Meanwhile, we show that the FDIF method can be re-instantiated approximately via a CNN-based architecture, where the convolution layer extracts anisotropic structure of image and the nonlinear layer enhances the structure via preserving local fractal dimension of image. The proposed filtering method provides us with a novel geometric interpretation of CNN-based image model. Focusing on a challenging image processing task - detecting complicated curves from the texture-like images, the proposed method obtains superior results to the state-of-art approaches. Hongteng Xu, Junchi Yan, Nils Persson, Weiyao Lin, Hongyuan Zha |
CVPR | 4 |
| 2017 | ThiNet: A Filter Level Pruning Method for Deep Neural Network CompressionabstractWe propose an efficient and unified framework, namely ThiNet, to simultaneously accelerate and compress CNN models in both training and inference stages. We focus on the filter level pruning, i.e., the whole filter would be discarded if it is less important. Our method does not change the original network structure, thus it can be perfectly supported by any off-the-shelf deep learning libraries. We formally establish filter pruning as an optimization problem, and reveal that we need to prune filters based on statistics information computed from its next layer, not the current layer, which differentiates ThiNet from existing methods. Experimental results demonstrate the effectiveness of this strategy, which has advanced the state-of-the-art. We also show the performance of ThiNet on ILSVRC-12 benchmark. ThiNet achieves 3.31 x FLOPs reduction and 16.63× compression on VGG-16, with only 0.52% top-5 accuracy drop. Similar experiments with ResNet-50 reveal that even for a compact network, ThiNet can also reduce more than half of the parameters and FLOPs, at the cost of roughly 1% top-5 accuracy drop. Moreover, the original VGG-16 model can be further pruned into a very small model with only 5.05MB model size, preserving AlexNet level accuracy but showing much stronger generalization ability. Jian-Hao Luo, Jianxin Wu 0001, Weiyao Lin |
ICCV | 3 |
| 2017 | Towards Reversal-Invariant Image Representation
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2017 | Antialiased super-resolution with parallel high-frequency synthesis
Xudong Jiang 0003, Bin Sheng 0001, Weiyao Lin, Ping Li 0016, Lizhuang Ma, Ruimin Shen |
Multim. Tools Appl. | 3 |
| 2017 | A Tube-and-Droplet-Based Approach for Representing and Analyzing Motion TrajectoriesabstractTrajectory analysis is essential in many applications. In this paper, we address the problem of representing motion trajectories in a highly informative way, and consequently utilize it for analyzing trajectories. Our approach first leverages the complete information from given trajectories to construct a thermal transfer field which provides a context-rich way to describe the global motion pattern in a scene. Then, a 3D tube is derived which depicts an input trajectory by integrating its surrounding motion patterns contained in the thermal transfer field. The 3D tube effectively: 1) maintains the movement information of a trajectory, 2) embeds the complete contextual motion pattern around a trajectory, 3) visualizes information about a trajectory in a clear and unified way. We further introduce a droplet-based process. It derives a droplet vector from a 3D tube, so as to characterize the high-dimensional 3D tube information in a simple but effective way. Finally, we apply our tube-and-droplet representation to trajectory analysis applications including trajectory clustering, trajectory classification & abnormality detection, and 3D action recognition. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our approach. Weiyao Lin, Hongteng Xu, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Zicheng Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Learning Correspondence Structures for Person Re-IdentificationabstractThis paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure, which indicates the patchwise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global constraint-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Finally, we also extend our approach by introducing a multi-structure scheme, which learns a set of local correspondence structures to capture the spatial correspondence sub-patterns between a camera pair, so as to handle the spatial misalignments between individual images in a more precise way. Experimental results on various data sets demonstrate the effectiveness of our approach. Weiyao Lin, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Jingdong Wang 0001, Ke Lu 0002 |
IEEE Trans. Image Process. | 1 |
| 2017 | Picking Neural Activations for Fine-Grained RecognitionabstractIt is a challenging task to recognize fine-grained subcategories due to the highly localized and subtle differences among them. Different from most previous methods that rely on object/part annotations, this paper proposes an automatic fine-grained recognition approach, which is free of any object/part annotation at both training and testing stages. The key idea includes two steps of picking neural activations computed from the convolutional neural networks, one for localization, and the other for description. The first picking step is to find distinctive neurons that are sensitive to specific patterns significantly and consistently. Based on these picked neurons, we initialize positive samples and formulate the localization as a regularized multiple instance learning task, which aims at refining the detectors via iteratively alternating between new positive sample mining and part model retraining. The second picking step is to pool deep neural activations via a spatially weighted combination of Fisher Vectors coding. We conditionally select activations to encode them into the final representation, which considers the importance of each activation. Integrating the above techniques produces a powerful framework, and experiments conducted on several extensive fine-grained benchmarks demonstrate the superiority of our proposed algorithm over the existing methods. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Weiyao Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Picking Deep Filter Responses for Fine-Grained Image RecognitionabstractRecognizing fine-grained sub-categories such as birds and dogs is extremely challenging due to the highly localized and subtle differences in some specific parts. Most previous works rely on object / part level annotations to build part-based representation, which is demanding in practical applications. This paper proposes an automatic fine-grained recognition approach which is free of any object / part annotation at both training and testing stages. Our method explores a unified framework based on two steps of deep filter response picking. The first picking step is to find distinctive filters which respond to specific patterns significantly and consistently, and learn a set of part detectors via iteratively alternating between new positive sample mining and part model retraining. The second picking step is to pool deep filter responses via spatially weighted combination of Fisher Vectors. We conditionally pick deep filter responses to encode them into the final representation, which considers the importance of filter responses themselves. Integrating all these techniques produces a much more powerful framework, and experiments conducted on CUB-200-2011 and Stanford Dogs demonstrate the superiority of our proposed algorithm over the existing methods. Xiaopeng Zhang 0008, Hongkai Xiong, Wengang Zhou 0001, Weiyao Lin, Qi Tian 0001 |
CVPR | 4 |
| 2016 | Supervised-learning based face hallucination for enhancing face recognitionabstractThis paper presents a two-step supervised face hallucination framework based on class-specific dictionary learning. Since the performance of learning-based face hallucination relies on its training set, an inappropriate training set (e.g., an input face image is very different from the training set) can reduce the visual quality of reconstructed high-resolution (HR) face significantly. To address this problem, we propose to utilize supervised learning to learn a set of class-specific dictionaries so that one of the learned dictionaries can well fit the global and local characteristics of an input low-resolution (LR) face image. Besides, the representative coefficients of the input LR face image may be unreliable due to insufficient information contained in the LR input image. To resolve this issue, we propose a maximum a posteriori estimator to infer the global HR face. Experimental results demonstrate that our method cannot only effectively enhance the visual quality of a reconstructed HR face, but also significantly improves the accuracy of face recognition compared to existing hallucination methods. Weng-Tai Su, Chih-Chung Hsu, Chia-Wen Lin, Weiyao Lin |
ICASSP | 4 |
| 2016 | A Short Survey of Recent Advances in Graph MatchingabstractGraph matching, which refers to a class of computational problems of finding an optimal correspondence between the vertices of graphs to minimize (maximize) their node and edge disagreements (affinities), is a fundamental problem in computer science and relates to many areas such as combinatorics, pattern recognition, multimedia and computer vision. Compared with the exact graph (sub)isomorphism often considered in a theoretical setting, inexact weighted graph matching receives more attentions due to its flexibility and practical utility. A short review of the recent research activity concerning (inexact) weighted graph matching is presented, detailing the methodologies, formulations, and algorithms. It highlights the methods under several key bullets, e.g. how many graphs are involved, how the affinity is modeled, how the problem order is explored, and how the matching procedure is conducted etc. Moreover, the research activity at the forefront of graph matching applications especially in computer vision, multimedia and machine learning is reported. The aim is to provide a systematic and compact framework regarding the recent development and the current state-of-the-arts in graph matching. Junchi Yan, Xu-Cheng Yin, Weiyao Lin, Cheng Deng 0002, Hongyuan Zha, Xiaokang Yang 0001 |
ICMR | 3 |
| 2016 | Structure-aware image inpainting using patch scale optimization
Chao Dai, Bin Sheng 0001, Jing Zhang 0041, Weiyao Lin, Yubo Yuan 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2016 | Tree-Based Visualization and Optimization for Image CollectionabstractThe visualization of an image collection is the process of displaying a collection of images on a screen under some specific layout requirements. This paper focuses on an important problem that is not well addressed by the previous methods: visualizing image collections into arbitrary layout shapes while arranging images according to user-defined semantic or visual correlations (e.g., color or object category). To this end, we first propose a property-based tree construction scheme to organize images of a collection into a tree structure according to user-defined properties. In this way, images can be adaptively placed with the desired semantic or visual correlations in the final visualization layout. Then, we design a two-step visualization optimization scheme to further optimize image layouts. As a result, multiple layout effects including layout shape and image overlap ratio can be effectively controlled to guarantee a satisfactory visualization. Finally, we also propose a tree-transfer scheme such that visualization layouts can be adaptively changed when users select different "images of interest." We demonstrate the effectiveness of our proposed approach through the comparisons with state-of-the-art visualization techniques. Xintong Han, Weiyao Lin, Mingliang Xu 0001, Bin Sheng 0001, Tao Mei 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | Good Practices for Learning to Recognize Actions Using FV and VLADabstractHigh dimensional representations such as Fisher vectors (FV) and vectors of locally aggregated descriptors (VLAD) have shown state-of-the-art accuracy for action recognition in videos. The high dimensionality, on the other hand, also causes computational difficulties when scaling up to large-scale video data. This paper makes three lines of contributions to learning to recognize actions using high dimensional representations. First, we reviewed several existing techniques that improve upon FV or VLAD in image classification, and performed extensive empirical evaluations to assess their applicability for action recognition. Our analyses of these empirical results show that normality and bimodality are essential to achieve high accuracy. Second, we proposed a new pooling strategy for VLAD and three simple, efficient, and effective transformations for both FV and VLAD. Both proposed methods have shown higher accuracy than the original FV/VLAD method in extensive evaluations. Third, we proposed and evaluated new feature selection and compression methods for the FV and VLAD representations. This strategy uses only 4% of the storage of the original representation, but achieves comparable or even higher accuracy. Based on these contributions, we recommend a set of good practices for action recognition in videos for practitioners in this field. Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin |
IEEE Trans. Cybern. | 3 |
| 2016 | A Diffusion and Clustering-Based Approach for Finding Coherent Motions and Understanding Crowd ScenesabstractThis paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurrent activity mining. It processes input motion fields (e.g., optical flow fields) and produces a coherent motion field named thermal energy field. The thermal energy field is able to capture both motion correlation among particles and the motion trends of individual particles, which are helpful to discover coherency among them. We further introduce a two-step clustering process to construct stable semantic regions from the extracted time-varying coherent motions. These semantic regions can be used to recognize pre-defined activities in crowd scenes. Finally, we introduce a cluster-and-merge process, which automatically discovers recurrent activities in crowd scenes by clustering and merging the extracted coherent motions. Experiments on various videos demonstrate the effectiveness of our approach. Weiyao Lin, Yang Mi, Weiyue Wang 0002, Jianxin Wu 0001, Jingdong Wang 0001, Tao Mei 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Multitask Learning of Compact Semantic Codebooks for Context-Aware Scene ModelingabstractIn the past few decades, we have witnessed the success of bag-of-features (BoF) models in scene classification, object detection, and image segmentation. Whereas it is also well acknowledged that the limitation of BoF-based methods lies in the low-level feature encoding and coarse feature pooling. This paper proposes a novel scene classification method, which leverages several semantic codebooks learned in a multitask fashion for robust feature encoding, and designs a context-aware image representation for efficient feature pooling. Apart from conventional universal codebook learning approaches, the proposed method encodes each class of local features with a unique semantic codebook, which captures the distinct distribution of different semantic classes more effectively. Instead of learning each semantic codebook separately, we learn a compact global codebook, of which each semantic codebook is a sparse subset, with a two-stage iterative multitask learning algorithm. While minimizing the clustering divergence, the semantic codeword assignment is solved by submodular optimization simultaneously. Built upon the global and semantic codebooks, a context-aware image representation is further developed to encode both global and semantic features in image representation via contextual quantization, semantic response computation, and semantic pooling. Extensive experiments have been conducted to validate the effectiveness of the proposed method on various public benchmarks with several popular local features. Hongkai Xiong, Weiyao Lin, Junni Zou, Yuan F. Zheng |
IEEE Trans. Image Process. | 3 |
| 2015 | Real Time Learning Evaluation Based on Gaze TrackingabstractIn this paper, we present a system that extracts the information implied by eye movements and use this information to analyze students' learning behavior. Our system uses a common webcam to capture students' facial image sequences when they are learning in front of monitors. We then process these images and establish sequences of changing location of iris, which represent the movements of eyes. With the eye movement sequences, we train a HMM classifier that can analyze their pattern and generate learning status for any given moment in the lesson. These statuses could help the computers to get a better understanding about the students' intention and behavior during online learning. The status sequences of those who view the same lesson could also be used as a reference for teaching quality assessment. Jiayue Yi, Bin Sheng 0001, Ruimin Shen, Weiyao Lin, Enhua Wu |
CAD/Graphics | 4 |
| 2015 | Person Re-Identification with Correspondence Structure LearningabstractThis paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure which indicates the patch-wise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Experimental results on various datasets demonstrate the effectiveness of our approach. Weiyao Lin, Junchi Yan, Jianxin Wu 0001, Jingdong Wang 0001 |
ICCV | 2 |
| 2015 | RIDE: Reversal Invariant Descriptor EnhancementabstractIn many fine-grained object recognition datasets, image orientation (left/right) might vary from sample to sample. Since handcrafted descriptors such as SIFT are not reversal invariant, the stability of image representation based on them is consequently limited. A popular solution is to augment the datasets by adding a left-right reversed copy for each original image. This strategy improves recognition accuracy to some extent, but also brings the price of almost doubled time and memory consumptions. In this paper, we present RIDE (Reversal Invariant Descriptor Enhancement) for fine-grained object recognition. RIDE is a generalized algorithm which cancels out the impact of image reversal by estimating the orientation of local descriptors, and guarantees to produce the identical representation for an image and its left-right reversed copy. Experimental results reveal the consistent accuracy gain of RIDE with various types of descriptors. We also provide insightful discussions on the working mechanism of RIDE and its generalization to other applications. Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001 |
ICCV | 3 |
| 2015 | Unsupervised Trajectory Clustering via Adaptive Multi-kernel-Based ShrinkageabstractThis paper proposes a shrinkage-based framework for unsupervised trajectory clustering. Facing to the challenges of trajectory clustering, e.g., large variations within a cluster and ambiguities across clusters, we first introduce an adaptive multi-kernel-based estimation process to estimate the 'shrunk' positions and speeds of trajectories' points. This kernel-based estimation effectively leverages both multiple structural information within a trajectory and the local motion patterns across multiple trajectories, such that the discrimination of the shrunk point can be properly increased. We further introduce a speed-regularized optimization process, which utilizes the estimated speeds to regularize the optimal shrunk points, so as to guarantee the smoothness and the discriminative pattern of the final shrunk trajectory. Using our approach, the variations among similar trajectories can be reduced while the boundaries between different clusters are enlarged. Experimental results demonstrate that our approach is superior to the state-of-art approaches on both clustering accuracy and robustness. Besides, additional experiments further reveal the effectiveness of our approach when applied to trajectory analysis applications such as anomaly detection and route analysis. Hongteng Xu, Weiyao Lin, Hongyuan Zha |
ICCV | 3 |
| 2015 | Traffic flow matching with clique and triplet cuesabstractThis paper addresses a new problem of matching traffic-flow patterns from different scenes. We firstly introduce cliques to measure the topology similarity between traffic flow patterns. Based on the clique information, a matching cost function is formulated to find the optimal flow-pattern matching. In order to avoid wrong matches due to large variations in traffic flow distributions, we further introduce triplets to measure the flow-wise correlation in a scene and include them into the matching cost function. Thus, constraints of traffic flows' relative position can be suitably considered during the flow-pattern matching process. Finally, a random-walk-based graph matching method is also utilized to efficiently solve the matching cost function optimization problem. Experimental results on both simulated flow data and real flow data demonstrate the effectiveness of our approach. Lihang Liu, Weiyao Lin, Youping Zhong |
MMSP | 2 |
| 2015 | Summarizing surveillance videos with local-patch-learning-based abnormality detection, blob sequence optimization, and type-based synopsis
Weiyao Lin, Jiwen Lu, Bing Zhou 0003, Jinjun Wang, Yu Zhou 0015 |
Neurocomputing | 1 |
| 2015 | Unsupervised adaptive sign language recognition based on hypothesis comparison guided cross validation and linguistic prior filtering
Yu Zhou 0015, Xiaokang Yang 0001, Yongzheng Zhang 0002, Yipeng Wang 0001, Xiujuan Chai, Weiyao Lin |
Neurocomputing | 7 |
| 2015 | Discriminative and generative vocabulary tree: With application to vein image authentication and recognition
Jinjun Wang, Jing Xiao 0006, Weiyao Lin, Chuanfei Luo |
Image Vis. Comput. | 3 |
| 2015 | GPU-Accelerated Video Background Subtraction Using Gabor Detector
Lixia Qin, Bin Sheng 0001, Weiyao Lin, Wen Wu 0001, Ruimin Shen |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Towards Good Practices for Action Video EncodingabstractHigh dimensional representations such as VLAD or FV have shown excellent accuracy in action recognition. This paper shows that a proper encoding built upon VLAD can achieve further accuracy boost with only negligible computational cost. We empirically evaluated various VLAD improvement technologies to determine good practices in VLAD-based video encoding. Furthermore, we propose an interpretation that VLAD is a maximum entropy linear feature learning process. Combining this new perspective with observed VLAD data distribution properties, we propose a simple, lightweight, but powerful bimodal encoding method. Evaluated on 3 benchmark action recognition datasets (UCF101, HMDB51 and Youtube), the bimodal encoding improves VLAD by large margins in action recognition. Jianxin Wu 0001, Yu Zhang 0004, Weiyao Lin |
CVPR | 3 |
| 2014 | Finding Coherent Motions and Semantic Regions in Crowd Scenes: A Diffusion and Clustering Approach
Weiyue Wang 0002, Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Jingdong Wang 0001, Bin Sheng 0001 |
ECCV (1) | 2 |
| 2014 | Representing And Recognizing Motion Trajectories: A Tube And Droplet ApproachabstractThis paper addresses the problem of representing and recognizing motion trajectories. We first propose to derive scene-related equipotential lines for points in a motion trajectory and concatenate them to construct a 3D tube for representing the trajectory. Based on this 3D tube, a droplet-based method is further proposed which derives a "water droplet" from the 3D tube and recognizes trajectory activities accordingly. Our proposed 3D tube can effectively embed both motion and scene-related information of a motion trajectory while the proposed droplet- based method can suitably catch the characteristics of the 3D tube for activity recognition. Experimental results demonstrate the effectiveness of our approach. Weiyao Lin, Hang Su 0006, Jianxin Wu 0001, Jinjun Wang, Yu Zhou 0015 |
ACM Multimedia | 2 |
| 2014 | Visual Similarity Based Anti-phishing with the Combination of Local and Global FeaturesabstractPhishing uses a fake Web page to steal personal sensitive information such as credit card numbers and passwords. Generally, the fake Web page is visually similar to the legitimate target Web page. The phishers can obtain financial benefits through these information. Anti-phishing is very important for a variety of applications such as phishing attacks, online transaction security, and user privacy protection. In this paper, we propose a novel and effective visual similarity based phishing detection approach that compares the snapshot image pair of the suspected Web page and the protected Web page. The proposed approach is based on the key insight that both the local and the global features of the Web page image can be used to represent the visual characteristics of the Web page together. This approach is purely on the image level, and thus can effectively deal with the non-text phishing tricks including images or Flashes objects in the HTML contents. For the local feature, the existence of the target logo is detected. For the global feature, the similarity of the visible part of the Web page is considered. We implemented and evaluated the proposed approach on a large scale dataset consisting of 2,129 real world phishing Web pages and 1,367 irrelevant legitimate Web pages. The experimental results show that the proposed approach can achieve over 90.00% true positive rate and 97.00% true negative rate. Our approach has been applied in the anti-phishing project of a major Internet Service Provider and gives a periodical reports to the potential users. Yu Zhou 0015, Yongzheng Zhang 0002, Yipeng Wang 0001, Weiyao Lin |
TrustCom | 5 |
| 2014 | Early detection of all-zero 4×4 blocks in High Efficiency Video Coding
Hanli Wang, Weiyao Lin, Sam Kwong, Oscar C. Au, Jun Wu 0006, Zhihua Wei 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Facial expression cloning with elastic and muscle models
Weiyao Lin, Bing Zhou 0003, Zhenzhong Chen 0001, Bin Sheng 0001, Jianxin Wu 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Image anti-aliasing techniques for Internet visual media processing: a reviewabstractAnti-aliasing is a well-established technique in computer graphics that reduces the blocky or stair-wise appearance of pixels. This paper provides a comprehensive overview of the anti-aliasing techniques used in computer graphics, which can be classified into two categories: post-filtering based anti-aliasing and pre-filtering based anti-aliasing. We discuss post-filtering based anti-aliasing algorithms through classifying them into hardware anti-aliasing techniques and post-process techniques for deferred rendering. Comparisons are made among different methods to illustrate the strengths and weaknesses of every category. We also review the utilization of anti-aliasing techniques from the first category in different graphic processing units, i.e., different NVIDIA and AMD series. This review provides a guide that should allow researchers to position their work in this important research area, and new research problems are identified. Xudong Jiang 0003, Bin Sheng 0001, Weiyao Lin, Lizhuang Ma |
J. Zhejiang Univ. Sci. C | 3 |
| 2014 | A New Network-Based Algorithm for Human Activity Recognition in VideosabstractIn this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error-free network. In this network, each node corresponds to a patch of the scene and each edge represents the activity correlation between the corresponding patches. Based on this network, we further model people in the scene as packages, while human activities can be modeled as the process of package transmission in the network. By analyzing these specific package transmission processes, various activities can be effectively detected. The implementation of our NTB algorithm into abnormal activity detection and group activity recognition are described in detail in this paper. Experimental results demonstrate the effectiveness of our proposed algorithm. Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Hanli Wang, Bin Sheng 0001, Hongxiang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Inferring User Image-Search Goals Under the Implicit Guidance of UsersabstractThe analysis of user search goals for a query can be very useful in improving search engine relevance and user experience. Although the research on inferring user goals or intents for text search has received much attention, little has been proposed for image search. In this paper, we propose to leverage click session information, which indicates high correlations among the clicked images in a session in user click-through logs, and combine it with the clicked images' visual information for inferring user image-search goals. Since the click session information can serve as past users' implicit guidance for clustering the images, more precise user search goals can be obtained. Two strategies are proposed to combine image visual information with the click session information. Furthermore, a classification risk based approach is also proposed for automatically selecting the optimal number of search goals for a query. Experimental results based on a popular commercial search engine demonstrate the effectiveness of the proposed method. Zheng Lu 0003, Xiaokang Yang 0001, Weiyao Lin, Hongyuan Zha |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Flexible Image Similarity Computation Using Hyper-Spatial MatchingabstractSpatial pyramid matching (SPM) has been widely used to compute the similarity of two images in computer vision and image processing. While comparing images, SPM implicitly assumes that: in two images from the same category, similar objects will appear in similar locations. However, this is not always the case. In this paper, we propose hyper-spatial matching (HSM), a more flexible image similarity computing method, to alleviate the mis-matching problem in SPM. Besides the match between corresponding regions, HSM considers the relationship of all spatial pairs in two images, which includes more meaningful match than SPM. We propose two learning strategies to learn SVM models with the proposed HSM kernel in image classification, which are hundreds of times faster than a general purpose SVM solver applied to the HSM kernel (in both training and testing). We compare HSM and SPM on several challenging benchmarks, and show that HSM is better than SPM in describing image similarity. Yu Zhang 0004, Jianxin Wu 0001, Jianfei Cai 0001, Weiyao Lin |
IEEE Trans. Image Process. | 4 |
| 2013 | A new Local-Main-Gradient-Orientation HOG and contour differences based algorithm for object classificationabstractThis paper presents a new algorithm to better classify objects in videos. In our case, the objects are cars, vans, and people on the roads. First, in order to extract the moving objects more precisely, we have proposed a method for foreground extraction based on the contour differences between the video frame and the background image. Second, after we got the integrated moving object, we have proposed a new algorithm to extract better features from the object. The new algorithm is based on two extended Histogram of Oriented Gradient (HOG) descriptor. We have improved HOG in two aspects: (a) selecting the gradient information from the moving objects and discarding the background gradient; (b) weighting every bin of gradient orientation histogram according to their significance within predefined area, in order to emphasize the important gradient information. We obtained Contour-Difference HOG (CD-HOG) from the first extension and Local-Main-Gradient-Orientation HOG (LMGO-HOG) from the second extended HOG. These extensions can cope with the cluttered background and make the features more distinguishable. Each of the extended HOG descriptors can produce a satisfying performance separately and an even better one if they are applied in cascade. From extensive evaluations, we showed the wonderful performance of our algorithm, and the accuracy rate of 94.04% can be achieved in some cases. Xiaoqiong Su, Weiyao Lin, Xiaozhen Zheng, Xintong Han, Hang Chu, Xiaoyun Zhang 0001 |
ISCAS | 2 |
| 2013 | A New Network-Based Algorithm for Human Group Activity Recognition in Videos
Gaojian Li, Weiyao Lin, Jianxin Wu 0001, Yuanzhe Chen, Hui Wei 0001 |
MMM (1) | 2 |
| 2013 | Introduction to the Special Issue on "Recent advances on analysis and processing for distributed video systems"
Chia-Wen Lin, Weiyao Lin, Zhenzhong Chen 0001, Marco Tagliasacchi, Shantanu Rane |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Improved image deblurring based on salient-region segmentation
Weiyao Lin, Wei Li 0209, Bing Zhou 0003, Jijia Li |
Signal Process. Image Commun. | 2 |
| 2013 | Intra-and-Inter-Constraint-Based Video Enhancement Based on Piecewise Tone MappingabstractVideo enhancement plays an important role in various video applications. In this paper, we propose a new intra-and-inter-constraint-based video enhancement approach aiming to: 1) achieve high intraframe quality of the entire picture where multiple regions-of-interest (ROIs) can be adaptively and simultaneously enhanced, and 2) guarantee the interframe quality consistencies among video frames. We first analyze features from different ROIs and create a piecewise tone mapping curve for the entire frame such that the intraframe quality can be enhanced. We further introduce new interframe constraints to improve the temporal quality consistency. Experimental results show that the proposed algorithm obviously outperforms the state-of-the-art algorithms. Yuanzhe Chen, Weiyao Lin, Zhenzhong Chen 0001, Ning Xu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | A Heat-Map-Based Algorithm for Recognizing Group Activities in VideosabstractIn this paper, a new heat-map-based algorithm is proposed for group activity recognition. The proposed algorithm first models human trajectories as series of heat sources and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this HM, a new key-point-based (KPB) method is used for handling the alignments among HMs with different scales and rotations. A surface-fitting (SF) method is also proposed for recognizing group activities. Our proposed HM feature can efficiently embed the temporal motion information of the group activities while the proposed KPB and SF methods can effectively utilize the characteristics of the HM for activity recognition. Section IV demonstrates the effectiveness of our proposed algorithms. Weiyao Lin, Hang Chu, Jianxin Wu 0001, Bin Sheng 0001, Zhenzhong Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | A New Algorithm for Inferring User Search Goals with Feedback SessionsabstractFor a broad-topic and ambiguous query, different users may have different search goals when they submit it to a search engine. The inference and analysis of user search goals can be very useful in improving search engine relevance and user experience. In this paper, we propose a novel approach to infer user search goals by analyzing search engine query logs. First, we propose a framework to discover different user search goals for a query by clustering the proposed feedback sessions. Feedback sessions are constructed from user click-through logs and can efficiently reflect the information needs of users. Second, we propose a novel approach to generate pseudo-documents to better represent the feedback sessions for clustering. Finally, we propose a new criterion )“Classified Average Precision (CAP)” to evaluate the performance of inferring user search goals. Experimental results are presented using user click-through logs from a commercial search engine to validate the effectiveness of our proposed methods. Zheng Lu 0003, Hongyuan Zha, Xiaokang Yang 0001, Weiyao Lin, Zhaohui Zheng 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | Exclusive Visual Descriptor Quantization
Yu Zhang 0004, Jianxin Wu 0001, Weiyao Lin |
ACCV (1) | 3 |
| 2012 | Distributed optimal power control for multicarrier cognitive systemsabstractIn this paper, the power optimization of the multicarrier cognitive system underlying the primary network is investigated. We consider the interference coupled cognitive network under individual secondary user's power constraint and primary user's rate constraint. A multicarrier discrete distributed (MCDD) algorithm based on Gibbs sampler is proposed. Although the problem is nonconcave, MCDD is proved to converge to the global optimal solution. To reduce the computational complexity and convergence time, the Gibbs sampler based Lagrangian algorithm (GSLA) is proposed to get a near optimal solution. We also provide simulation results to show the effectiveness of the proposed algorithms. Guanying Ru, Hongxiang Li 0001, Thuan T. Tran, Weiyao Lin, Lingjia Liu 0001, Huasen Wu |
GLOBECOM | 4 |
| 2012 | Adaptive scheduling for multicasting hard deadline constrained prioritized data via network codingabstractNetwork coding offers a promising platform for multicast transmission by approaching its min-cut capacity. However, pushing the network throughput toward this upper bound comes with a sacrifice in delivery delay due to the decoding procedure that requires performing batch of coded packets. Further, in some transmission scenarios where the receivers experience deep fading or unable to collect a full set of the transmitted data, no useful information is recovered. The effect is more severe in the networks where the transmitted information has priority structure with hard deadline constraint due to the limited delivery time and data interdependencies. In this paper, we consider single-hop wireless networks where the transmitter wishes to multicast hard deadline constrained prioritized data to many receivers over lossy channels. We first study the network performance of a variety of transmission techniques, depending on how the transmitter schedules transmission in each time slot. We then propose an adaptive encoding and scheduling technique to maximize the network throughput. To find the optimal transmission scheduling at the presence of the network dynamics, we cast the problem in the framework of Markov Decision Processes (MDP) and use backward induction method to find an optimal solution. We further propose simulation-based algorithm and greedy scheduling technique that obtain high performance with much lower time complexity. Both analytical and simulation results have been provided to corroborate the effectiveness of the proposed techniques. Thuan T. Tran, Hongxiang Li 0001, Weiyao Lin, Lingjia Liu 0001, Samee Ullah Khan |
GLOBECOM | 3 |
| 2012 | An enhanced covariance spectrum sensing technique based on stochastic resonance in cognitive radio networksabstractIn this paper, a novel covariance spectrum sensing approach used in cognitive radio (CR) networks which is based on the dynamical stochastic resonance (SR) technique is proposed. When the optimal SR technique is introduced as the pre-processing method for the covariance-based detection and after it has been realized, it can increase the signal-to-noise ratio (SNR) of the primary user (PU) signal and accordingly increase the mean value of the decision statistic of the covariance-based detection, so that the detection probability of the proposed approach can be improved under constant false alarm rate (CFAR). Computer simulation results verify the effectiveness of the proposed approach compared with the traditional spectrum sensing methods. Di He 0002, Winston Li, Fusheng Zhu, Weiyao Lin |
ISCAS | 4 |
| 2012 | Facial expression mapping based on elastic and muscle-distribution-based modelsabstractIn this paper, a new algorithm is proposed for facial expression mapping. The proposed algorithm first introduces a new elastic model to balance the global and local warping effects such that the impacts from facial feature differences between people can be avoided, thus more reasonable geometric warping results can be created. Furthermore, a muscle-distribution-based (MD) model is also proposed. The proposed MD model utilizes the muscle distribution information of the human face to evaluate and strengthen the facial illumination details. By this way, the impacts from human face difference as well as the effects of unsuitable noise filtering can be effectively alleviated. Experimental results show that our proposed algorithm can create obviously better facial expression results than the existing methods. Weiyao Lin, Bin Sheng 0001, Jianxin Wu 0001, Hongxiang Li 0001 |
ISCAS | 2 |
| 2012 | A new heat-map-based algorithm for human group activity recognitionabstractIn this paper, a new heat-map-based (HMB) algorithm is proposed for human group activity recognition. The proposed algorithm first models people trajectories as series of "heat sources" and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this heat map, a new surface-fitting (SF) method is also proposed for recognizing human group activities. Our proposed HM feature can efficiently keep the temporal motion information of the group activities while the proposed SF method can effectively catch the characteristics of the heat map for activity recognition. Experimental results demonstrate the effectiveness of our proposed algorithm. Hang Chu, Weiyao Lin, Jianxin Wu 0001, Xingtong Zhou, Yuanzhe Chen, Hongxiang Li 0001 |
ACM Multimedia | 2 |
| 2012 | Parsing collective behaviors by hierarchical model with varying structureabstractCollective behaviors are usually composed of several groups. Considering the interactions among groups, this paper presents a novel framework to parse collective behaviors for video surveillance applications. We first propose a latent hierarchical model (LHM) with varying structure to represent the behavior with multiple groups. Furthermore, we also propose a multi-layer-based (MLB) inference method, where a sample-based heuristic search (SHS) is introduced to infer the group affiliation. And latent SVM is adopted to learn our model. With the proposed LHM, not only are the collective behaviors detected effectively, but also the group affiliation in the collective behaviors is figured out. Experiment results demonstrate the effectiveness of the proposed framework. Cong Zhang 0005, Xiaokang Yang 0001, Weiyao Lin |
ACM Multimedia | 4 |
| 2012 | Inferring user image-search goals by mining query logs with semi-supervised spectral clusteringabstractInferring user search goals for a query can be very useful in improving search engine relevance and user experience. Although the research on analyzing user goals or intents for text search has received much attention, little has been proposed for image search. In this paper, we propose a novel approach to infer user search goals in image search by mining search engine query logs with semi-supervised spectral clustering. We combine the visual information of the clicked images with user click information by using graph-based models and then cluster the images with spectral clustering to capture user image-search goals. Experimental results based on a popular commercial search engine demonstrate the effectiveness of the proposed method. Zheng Lu 0003, Xiaokang Yang 0001, Weiyao Lin, Hongyuan Zha |
VCIP | 3 |
| 2012 | A patch-based framework for detecting abnormal activities with a PTZ cameraabstractIn this paper, a novel patch-based (PB) framework is proposed for detecting abnormal activities using a Pan-Tilt-Zoom (PTZ) camera. We first propose a new scene-patch-based (SSB) algorithm which can efficiently extract the target object's global trajectory from the PTZ camera. Furthermore, we propose an extended network-based (ENB) algorithm for detecting abnormal activities. The proposed ENB algorithm models the entire scene as a network where each node in the network corresponds to a patch of the scene and each edge between nodes corresponds to the activity correlation between the scene patchs. Based on this network, a recursive training strategy is proposed to train the edge weights in the network such that abnormal activities can be effectively detected through these trained edge weights. Experimental results demonstrate the effectiveness of our proposed framework. Yisi Tao, Yuanzhe Chen, Weiyao Lin, Xintong Han, Hongxiang Li 0001, Zheng Lu 0003 |
VCIP | 3 |
| 2012 | A New Cooperative Spectrum Sensing Scheme for Cognitive Ad-Hoc Networks
Hongxiang Li 0001, Weiyao Lin, Lingjia Liu 0001, Samee Ullah Khan, Sentang Wu |
Mob. Networks Appl. | 3 |
| 2012 | Region-Based Rate Control for H.264/AVC for Low Bit-Rate ApplicationsabstractRate control plays an important role in video coding. However, in the conventional rate control algorithms, the number and position of macroblocks (MBs) inside one basic unit for rate control is inflexible and predetermined. The different characteristics of the MBs are not fully considered. Also, there is no overall optimization of the coding of basic units. This paper proposes a new region-based rate control scheme for H.264/advanced video coding to improve the coding efficiency. The inter-frame information is explored to objectively divide one frame into multiple regions based on their rate-distortion (R-D) behaviors. The MBs with similar characteristics are classified into the same region, and the entire region, instead of a single MB or a group of contiguous MBs, is treated as a basic unit for rate control. A linear rate-quantization stepsize model and a linear distortion-quantization stepsize model are proposed to accurately describe the R-D characteristics for the region-based basic units. Moreover, based on the above linear models, an overall optimization model is proposed to obtain suitable quantization parameters for the region-based basic units. Experimental results demonstrate that the proposed region-based rate control approach can achieve both better subjective and objective quality by performing the rate control adaptively with the content, compared to the conventional rate control approaches. Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Wei Li 0209, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Multiscale Semilocal Interpolation With AntialiasingabstractAliasing is a common artifact in low-resolution (LR) images generated by a downsampling process. Recovering the original high-resolution image from its LR counterpart while at the same time removing the aliasing artifacts is a challenging image interpolation problem. Since a natural image normally contains redundant similar patches, the values of missing pixels can be available at texture-relevant LR pixels. Based on this, we propose an iterative multiscale semilocal interpolation method that can effectively address the aliasing problem. The proposed method estimates each missing pixel from a set of texture-relevant semilocal LR pixels with the texture similarity iteratively measured from a sequence of patches of varying sizes. Specifically, in each iteration, top texture-relevant LR pixels are used to construct a data fidelity term in a maximum a posteriori estimation, and a bilateral total variation is used as the regularization term. Experimental results compared with existing interpolation methods demonstrate that our method can not only substantially alleviate the aliasing problem but also produce better results across a wide range of scenes both in terms of quantitative evaluation and subjective visual quality. Kai Guo 0001, Xiaokang Yang 0001, Hongyuan Zha, Weiyao Lin, Songyu Yu |
IEEE Trans. Image Process. | 4 |
| 2012 | Embedded I/O PAD Circuit Design for OTP Memory Power-Switch FunctionalityabstractAn additional high-voltage pad is generally applied for one-time-programming (OTP) memory product applications. This may increase the complexity of input/output (I/O) pad arrangement and the area penalty. In this paper, a novel approach of I/O circuit embedded with the power-switch function is proposed for multifunction integrations in one I/O pad. The capabilities of high-voltage programming, I/O signal handling, electrostatic discharge protection and latch-up prevention for this novel circuit are well examined from silicon verifications. Shao-Chang Huang, Ke-Horng Chen, Weiyao Lin, Zon-Lon Lee, Kun-Wei Chang, Erica Hsu, Wenson Lee, Lin-Fwu Chen, Chris Chun-Hung Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Hypothesis comparison guided cross validation for unsupervised signer adaptationabstractSigner adaptation is important to sign language recognition systems in that a one-size-fits-all model set can not perform well on all kinds of signers. Supervised signer adaptation must utilize the labeled adaptation data that are collected explicitly. To skip the data collecting process in signer adaptation, we propose an unsupervised adaptation method called hypothesis comparison guided cross validation (HC CV) algorithm. The algorithm not only addresses the problem of overlap between the data set to be labeled and the data set for adaptation, but also employs an additional hypothesis comparison step to decrease the noise rate of the adaptation data set. Experimental results show that the HC CV adaptation algorithm is superior to the CV adaptation algorithm and the conventional self-teaching algorithm. Though the algorithm is proposed for signer adaptation, it can also be applied to speaker adaptation and writer adaptation straightforwardly. Yu Zhou 0015, Xiaokang Yang 0001, Weiyao Lin, Yi Xu 0001, Long Xu 0001 |
ICME | 3 |
| 2011 | A new network-based algorithm for multi-camera abnormal activity detectionabstractIn this paper, a new abnormal activity detection algorithm is proposed for multi-camera surveillance applications. The proposed algorithm models the entire scene covered by the multi-camera system as a network. In this network, each node corresponds to a segmentation of the entire scene and each edge represents the activity correlation between the corresponding segmentations. Based on this network, the proposed algorithm further models human activities as the signal transmission process in the network. Thus, abnormal activities can be detected if their 'network transmission energy' is obviously larger than the normal case. Compared with the previous methods, the proposed algorithm is more general and is flexible to handle various multi-camera scenarios and configurations. Experimental results demonstrate the effectiveness of the proposed algorithm. Weiyao Lin, Xiaokang Yang 0001, Hongxiang Li 0001, Ning Xu 0007 |
ISCAS | 2 |
| 2011 | A new Temporal-Constraint-Based algorithm by handling temporal qualities for video enhancementabstractVideo enhancement has played very important roles in many applications. However, most existing enhancement methods only focus on the spatial quality within a frame while the temporal qualities of the enhanced video are often unguaranteed. In this paper, a new algorithm is proposed for video enhancement. The proposed algorithm introduces new temporal constraints and combines them with the spatial constraints such that both the spatial and temporal qualities of the video can be improved. Two strategies are proposed for including the temporal constraints. Experimental results demonstrate the effectiveness of the proposed algorithm. Weiyao Lin, Hongxiang Li 0001, Ning Xu 0007, Lining Zhang |
ISCAS | 2 |
| 2011 | Saliency-based visualization for image searchabstractIn this paper, we propose a novel algorithm for improving and visualizing image search results. The proposed algorithm improves user's image search experience by three steps: (1) re-rank the initial image search results by the random walk refinement based on visual consistency and saliency cues, (2) project the re-ranked images into a 2-dimentional panel according to their saliency information and correlations, (3) detect and extract the saliency regions in each image for final visualization. To evaluate the performance of our algorithm, user study has been conducted. Experimental results demonstrate that our visualization algorithm provides more pleasing image search experience than the conventional image search methods. Jiajie Hu, Bin Jin, Weiyao Lin, Hangzai Luo, Zhenzhong Chen 0001, Hongxiang Li 0001 |
MMSP | 3 |
| 2011 | A new package-group-transmission-based algorithm for human activity recognition in videosabstractIn this paper, a new package-group-transmission-based algorithm is proposed for human activity recognition in videos. The proposed algorithm first models the entire scene as a network where each node in the network corresponds to a segmentation of the scene. Based on this network, we further model people in the scene as groups of packages. Thus, various human activities can be modeled as the process of "package group transmission" in the network and these activities can be efficiently recognized by suitably analyzing the "package transmission" process. Our proposed algorithm can not only detect activities under the challenging multiple camera scenario, but also be able to recognize various complex group activities among people. Experimental results demonstrate the effectiveness of our proposed algorithm. Yuanzhe Chen, Weiyao Lin, Hongxiang Li 0001, Hangzai Luo, Yisi Tao, Donghua Liu |
VCIP | 2 |
| 2011 | Inferring users' image-search goals with pseudo-imagesabstractThe analysis of user search goals for a query can be very useful in improving search engine relevance and user experience. Although the research on inferring user goals or intents for text search has received much attention, little has been proposed for image search with visual information. In this paper, we propose a novel approach to capture user search goals in image search by exploring pseudo-images which are extracted by mining single sessions in user click-through logs to reflect user information needs. Moreover, we also propose a novel evaluation criterion to determine the number of user search goals for a query. Experimental results demonstrate the effectiveness of the proposed method. Zheng Lu 0003, Xiaokang Yang 0001, Weiyao Lin, Hongyuan Zha |
VCIP | 3 |
| 2011 | A new global-based video enhancement algorithm by fusing features of multiple region-of-interestsabstractVideo enhancement plays an important role in various video applications. It is desirable to achieve high visual quality of the entire picture where multiple region-of-interests (ROIs) within the frame can be adaptively and simultaneously enhanced. In this paper, a new global-based video enhancement algorithm is proposed. The proposed algorithm first analyzes features from different ROIs. Then, a 'global' tone mapping curve is created for the entire picture which can adaptively enhance different regions at the same time. According to the statistics of ROIs, two fusion strategies, i.e., piecewise-based and factor-based fusions, are proposed for creating the global tone mapping curve. Experimental results show that the proposed algorithm can obtain more appealing perceptual quality than the state-of-the-art algorithms. Ning Xu 0007, Weiyao Lin, Yu Zhou 0015, Yuanzhe Chen, Zhenzhong Chen 0001, Hongxiang Li 0001 |
VCIP | 2 |
| 2011 | Capacity of Multicarrier Multilayer Broadcast and Unicast Hybrid Cellular System with Independent Channel Coding over SubcarriersabstractIn this paper, we discuss the hybrid capacity region of a generic multicarrier multilayer broadcast and unicast cellular system with independent channel coding over subcarriers. In particular, we analytically derive the capacity region and provide conditions to achieve its boundary. The simulation results show that the hybrid capacity regions are considerably higher than those of the traditional time division multiplexing scheme. Siqian Liu, Hongxiang Li 0001, Guanying Ru, Weiyao Lin, Lingjia Liu 0001, Yang Yi 0002 |
VTC Fall | 4 |
| 2011 | Spectrum Optimization for OFDMA Based Hybrid Wireless NetworksabstractThis paper proposes a new scheme for cooperative hybrid network resource allocation. Different from the traditional cognitive radio networks that aim to utilize the spectrum holes (such as white space in TV spectrum) in an uncoordinated way, this paper proposes an OFDMA based collaborative hybrid network. We study the joint resource allocation problem for both the primary users and the secondary users. Results show that by cooperatively allocating resources in the primary network and the secondary network, we can achieve higher spectral efficiency while provide satisfactory admission control and QoS for all users. Guanying Ru, Hongxiang Li 0001, Siqian Liu, Weiyao Lin, Lingjia Liu 0001 |
VTC Fall | 4 |
| 2011 | A rate-control algorithm using inter-layer information for H.264/SVC for low-delay applications
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Ming-Ting Sun |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | A region-based rate-control scheme using inter-layer information for H.264/SVC
Hai-Miao Hu, Weiyao Lin, Bo Li 0006, Ming-Ting Sun |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Real-time control of individual agents for crowd simulation
Yunbo Rao, Leiting Chen, Qihe Liu, Weiyao Lin |
Multim. Tools Appl. | 4 |
| 2011 | A Fast Sub-Pixel Motion Estimation Algorithm for H.264/AVC Video CodingabstractMotion estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264/AVC makes it even more complicated when compared to ME in conventional video coding standards. It is important to develop fast and effective sub-pixel ME algorithms since: 1) the computation overhead by sub-pixel ME has become relatively significant while the complexity of integer-pixel search has been greatly reduced by fast algorithms, and 2) reducing sub-pixel search points can greatly save the computation for sub-pixel interpolation. In this letter, a novel fast sub-pixel ME algorithm is proposed which performs a “rough” sub-pixel search before the partition selection, and performs a “precise” sub-pixel search for the best partition. By reducing the searching load for the large number of non-best partitions, the computation complexity for sub-pixel search can be greatly decreased. Experimental results show that our method can reduce the sub-pixel search points by more than 50% compared to existing fast sub-pixel ME methods with negligible quality degradation. Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun, Zhenzhong Chen 0001, Hongxiang Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | A Computation Control Motion Estimation Method for Complexity-Scalable Video CodingabstractIn this paper, a new computation-control motion estimation (CCME) method is proposed which can perform motion estimation (ME) adaptively under different computation or power budgets while keeping high coding performance. We first propose a new class-based method to measure the macroblock (MB) importance where MBs are classified into different classes and their importance is measured by combining their class information as well as their initial matching cost information. Based on the new MB importance measure, a complete CCME framework is then proposed to allocate computation for ME. The proposed method performs ME in a one-pass flow. Experimental results demonstrate that the proposed method can allocate computation more accurately than previous methods and, thus, has better performance under the same computation budget. Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Group Event Detection With a Varying Number of Group Members for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with a varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between people. Furthermore, we propose a group activity detection algorithm which can handle both symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | A New Class-based Early Termination Method for Fast Motion Estimation in Video CodingabstractMotion Estimation (ME) is one of the most time-consuming parts in video coding. It is always desirable to develop fast ME algorithms to reduce the ME complexity. In this paper, a new early termination method is proposed for fast motion estimation. The proposed method first classifies each macroblock into one of three classes based on the estimation of the possible matching cost improvement from future search points. Different early termination strategies are then applied to different classes. Experimental results show that the proposed method can significantly reduce the search points with little quality degradation. Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun |
ISCAS | 1 |
| 2009 | A New One-pass Complexity-Scalable Computation-control Method for Video CodingabstractIn this paper, a new complexity-scalable computation control method is proposed which can perform motion estimation (ME) adaptively under different computation or power budgets while keeping high coding performance. We first propose a new class-based method to measure the macroblock (MB) importance where MBs are classified into different classes and their importance is measured by combining their class information as well as their initial matching cost information in the ME. Based on the new MB importance measure, a computation control framework is then proposed to allocate computation for ME. Experimental results demonstrate that the proposed method can allocate computation more accurately than previous methods and thus has better performance under the same computation budget. Weiyao Lin, Krit Panusopone, David M. Baylon, Ming-Ting Sun |
ISCAS | 1 |
| 2009 | Group Event Detection for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of group activities for video surveillance applications. We propose to use a group representative to handle the recognition with flexible or varying number of group members, and use an asynchronous hidden Markov model (AHMM) to model the relationship between two people. Furthermore, we propose a group activity detection algorithm which can handle symmetric and asymmetric group activities, and demonstrate that this approach enables the detection of hierarchical interactions between people. Experimental results show the effectiveness of our approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
ISCAS | 1 |
| 2008 | Fast sub-pixel motion estimation and mode decision for H.264abstractMotion Estimation (ME) is one of the most time-consuming parts in video coding. The use of multiple partition sizes in H.264 makes the ME even more complicated. It is important to develop fast sub-pixel ME algorithms due to (1) The computation overhead by sub-pixel ME has become relatively significant while the complexity of integer-pel search has been greatly reduced by fast algorithms and (2) Reducing sub-pel searching points can save the computation for interpolating sub-pixel values. In this paper, a new fast sub-pixel ME algorithm is proposed which performs a ‘rough’ sub-pel search before the partition selection and only performs the ‘precise’ sub-pel search for the best partition. Experimental results show that our method can reduce the sub-pel search points by more than 50% compared to existing fast sub-pel ME methods with little quality degradation. Weiyao Lin, David M. Baylon, Krit Panusopone, Ming-Ting Sun |
ISCAS | 1 |
| 2008 | Human activity recognition for video surveillanceabstractThis paper presents a novel approach for automatic recognition of human activities from video sequences. We first group features with high correlations into Category Feature Vectors (CFVs). Each activity is then described by a combination of GMMs (Gaussian Mixture Models) with each GMM representing the distribution of a CFV. We show that this approach offers flexibility to add new events and to deal with the problem of lacking training data for building models for unusual events. For improving the recognition accuracy, a Confident-Frame-based Recognizing algorithm (CFR) is proposed to recognize the human activity, where the video frames which have high confidence for recognition an activity (Confident-Frames) are used as a specialized model for classifying the rest of the video frames. Experimental results show the effectiveness of the proposed approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
ISCAS | 1 |
| 2008 | Activity Recognition Using a Combination of Category Components and Local Models for Video SurveillanceabstractThis paper presents a novel approach for automatic recognition of human activities for video surveillance applications. We propose to represent an activity by a combination of category components and demonstrate that this approach offers flexibility to add new activities to the system and an ability to deal with the problem of building models for activities lacking training data. For improving the recognition accuracy, a confident-frame-based recognition algorithm is also proposed, where the video frames with high confidence for recognizing an activity are used as a specialized local model to help classify the remainder of the video frames. Experimental results show the effectiveness of the proposed approach. Weiyao Lin, Ming-Ting Sun, Radha Poovendran, Zhengyou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |