VLDB 2026 Research / reviewers in the wild / expert
Jiajun Liu 0004
dblp:75/5729-4
· DBLP profile ↗
73ranked-venue papers
9as first author
44since 2021 · last 2026
0000-0001-8160-1796ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 1 first-author · 24 since 2021Databases, data management, data science and information retrieval · 24 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 16 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFedDG2: Distribution-Guided Gossip-Based Generalizable and Communication-Efficient Decentralized Federated LearningabstractTraditional Federated Learning (FL) focuses on collaborative global model training while ensuring privacy and personalization. Decentralized Federated Learning (DFL), a variant of FL, allows clients to independently manage and optimize local models without a central server. DFL reduces centralized communication bottlenecks and vulnerability to server failures or attacks. However, because the optimization dynamics change and there is no global model, generalization can suffer, making effective learning under data and model heterogeneity a critical challenge in DFL. Despite growing interest in DFL, the lack of distributional and uncertainty modeling in the literature limits reliability and effective generalization in non-IID settings. In this work, we propose DFedDG2, a personalized federated learning framework that operates within a peer-to-peer protocol. DFedDG2 offers the technical advantage of modeling each client’s local data distribution and exchanging this information with one-hop neighbors. In the decentralized network, clients perform distribution-aware gossip, where statistically similar clients exert greater influence to drive global alignment. This likelihood-weighted mixing fuses only a handful of vectors and scalars, significantly reducing communication costs while aligning semantic spaces across the network and enabling personalized training at the edge. In addition, theoretically, we prove that DFedDG2 achieves a sublinear convergence rate while the consensus error decays at a geometric rate under well-principled properties of gossip. Unlike the label-only non-IID experiments in DFL literature, we conduct extensive experiments on multiple data non-IID scenarios, topology variations, model heterogeneity, and uncertainty quantification, demonstrating the practical advantages of DFedDG2. Our results show that DFedDG2 not only achieves communication efficiency but also provides better generalization and improved reliability compared to state-of-the-art approaches. Biprodip Pal, Stanislav Funiak, Jiajun Liu 0004, Peyman Moghadam, Md. Saiful Islam 0003, Alan Wee-Chung Liew |
IEEE Internet Things J. | 3 |
| 2026 | Understanding the Effects of Projectors in Knowledge DistillationabstractConventionally, during the knowledge distillation process (e.g., feature distillation), an additional projector is often required to perform feature transformation due to the dimension mismatch between the teacher and the student networks. Interestingly, we discovered that even if the student and the teacher have the same feature dimensions, adding a projector still helps to improve the distillation performance. In addition, projectors even improve logit distillation if we add them to the architecture too. Inspired by these surprising findings and the general lack of understanding of the projectors in the knowledge distillation process from existing literature, this paper investigates the implicit role that projectors play, but so far been overlooked. Our empirical study shows that the student with a projector 1) obtains a better trade-off between the training accuracy and the testing accuracy compared to the student without a projector when it has the same feature dimensions as the teacher, 2) better preserves its similarity to the teacher beyond shallow and numeric resemblance, from the view of Centered Kernel Alignment (CKA) (Kornblith et al., 2019), and 3) avoids being over-confident (Guo et al., 2017) as the teacher does at the testing phase. Motivated by the positive effects of projectors, we propose a projector ensemble-based feature distillation method to further improve distillation performance. Despite the simplicity of the proposed strategy, empirical results from the evaluation of classification tasks on benchmark datasets demonstrate the superior classification performance of our method on a broad range of teacher-student pairs and verify, from the aspects of CKA and model calibration that the student's features are of improved quality with the projector ensemble design. Yudong Chen 0002, Sen Wang 0001, Jiajun Liu 0004, Xuwei Xu, Frank de Hoog, Branislav Kusy, Zi Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | A unified analysis on cross-architecture generalizability of coresetsabstractCoreset selection methods aim to identify a representative subset of training data that preserves competitive performance. However, mainstream coreset selection approaches are model-specific and assume they already have full information about the target model when the coreset is selected. This largely restricts the usefulness of coreset selection in practice. This work aims to fill that gap by formulating and investigating the problem of cross-architecture generalizability of coresets: we develop a unified theoretical framework that analyzes the upper bound of coreset selection objective functions, extend it to scenarios involving multiple downstream architectures, and provide an empirical analysis on cross-architecture coreset performance. Based on our findings, we propose a novel ensemble scoring method that aggregates multi-source knowledge to enhance cross-architecture generalizability. Our extensive experiments across thirteen architectures and six selection ratios provide comprehensive verification of our theoretical analysis. The source code is available at https://github.com/diqichen91/CACS.git . Diqi Chen, Jiajun Liu 0004, Frank de Hoog, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2026 | DBCore: Shaping generalizable decision boundaries for coreset selectionabstractCoreset selection for classification often relies on assessing individual sample difficulty or importance, leading to sample-wise or range-based selection, but this can overlook the collective impact on model decision boundaries. Realizing that the representative power a coreset possesses is tightly associated with the decision boundaries a model can form on it, we propose a novel approach that directly optimizes the Decision Boundary (DB) formed by the selected coreset. Specifically, we ask: How can we collectively select samples to create a DB that is globally smoothed yet locally detailed, ensuring maximum generalizability and noise-resilience to the original dataset? To address this, we define two key objectives: (1) Global shape retention – The selected coreset should form a smoothed version of the original DB, preserving its overall structure and preventing overfitting; (2) Local detail preservation – While smoothing prevents overfitting, excessive smoothing risks losing critical nuances. Thus, the selection must also retain key points near the original DB to capture local complexities. We formulate these objectives as a convex quadratic optimization problem with linear constraints and solve it efficiently. Extensive evaluations demonstrate the consistent and substantial advantages of our method over the state-of-the-art coreset selection strategies. The source code is available at https://github.com/diqichen91/DBCore.git . Diqi Chen, Jiajun Liu 0004, Frank de Hoog, Wangzhi Xing, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2026 | On learning denoisable student logitsabstractKnowledge Distillation (KD) aims to train a student model to mimic the behavior of a more powerful teacher model. In this paper, we reveal that through the lens of diffusion processes, student logits can be statistically treated as a noisy version of teacher logits, and KD helps reduce the noise level of student logits. This insight motivates us to design a framework leveraging KD to produce denoisable student logits that can be further recovered towards teacher logits via a reverse diffusion process. A key advantage of this approach is that the inference-diffusion process can occur in two physical locations and on separate devices, enabling a two-step and distributed inference process. The experimental results show that the derived denoisable student logits achieve comparable or even superior performance to standard KD’s, and the reverse diffusion process achieves a substantial improvement in accuracy, without needing the original image, thus preserving the privacy and security of the original data. Additionally, the logits can be further compressed before transmission, reducing the required bandwidth while achieving comparable overall performance. Diqi Chen, Yang Li 0184, Jiajun Liu 0004, Branislav Kusy, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2026 | SATE: Efficient knowledge distillation with implicit student-aware teacher ensemblesabstractRecent findings suggest that with the same teacher architecture, a fully converged or “stronger” checkpoint surprisingly leads to a worse student. This can be explained by the Information Bottleneck (IB) principle, as the features of a weaker teacher transfer more “dark” knowledge because they maintain higher mutual information with the inputs. Meanwhile, various works have shown that severe teacher-student structural disparity or capability mismatch often leads to worse student performance. To deal with these issues, we propose a generalizable and efficient Knowledge Distillation (KD) framework with implicit Student-Aware Teacher Ensembles (SATE). The SATE framework simultaneously trains a student network and a student-aware intermediate teacher as a learning companion. With the proposed co-training strategy, the intermediate teacher is trained gradually and forms implicit ensembles of weaker teachers along the learning process. Such a design enables the student model to retain more dark knowledge for better generalization ability. The proposed framework improves the training scheme in a plug-and-play way so that it can be applied to improve various classic and state-of-the-art KD methods on both intra-domain (up to 2.184 % ) and cross-domain (up to 7.358 % ) settings, under a diversified configurations on teacher-student architectures, and achieves a major efficient advantage over other generic frameworks. The code is available at https://github.com/diqichen91/SATE.git . Diqi Chen, Yang Li 0184, Jiajun Liu 0004, Jun Zhou 0001, Yongsheng Gao 0001 |
Pattern Recognit. | 3 |
| 2025 | Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty MinimizationabstractDespite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent through clarifying questions, current methods typically rely on heuristic or ad-hoc strategies without explicitly quantifying these uncertainties, limiting their effectiveness. Motivated by this gap, we propose UMIVR, an Uncertainty-Minimizing Interactive Text-to-Video Retrieval framework that explicitly quantifies three critical uncertainties-text ambiguity, mapping uncertainty, and frame uncertainty-via principled, training-free metrics: semantic entropy-based Text Ambiguity Score (TAS), Jensen-Shannon divergence-based Mapping Uncertainty Score (MUS), and a Temporal Quality-based Frame Sampler (TQFS). By adaptively generating targeted clarifying questions guided by these uncertainty measures, UMIVR iteratively refines user queries, significantly reducing retrieval ambiguity. Extensive experiments on multiple benchmarks validate UMIVR's effectiveness, achieving notable gains in Recall@1 (69.2\% after 10 interactive rounds) on the MSR-VTT-1k dataset, thereby establishing an uncertainty-minimizing foundation for interactive TVR. Bingqing Zhang, Heming Du, Yang Li 0184, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001 |
ICCV | 6 |
| 2025 | Beyond Static LLM Policies: Imitation-Enhanced Reinforcement Learning for RecommendationabstractRecommender systems (RecSys) have become critical tools for enhancing user engagement by delivering personalized content across diverse digital platforms. Recent advancements in large language models (LLMs) demonstrate significant potential for improving RecSys, primarily due to their exceptional generalization capabilities and sophisticated contextual understanding, which facilitate the generation of flexible and interpretable recommendations. However, the direct deployment of LLMs as primary recommendation policies presents notable challenges, including persistent latency issues stemming from frequent API calls and inherent model limitations such as hallucinations and biases. To address these issues, this paper proposes a novel offline reinforcement learning (RL) framework that leverages imitation learning from LLM-generated trajectories. Specifically, inverse reinforcement learning is employed to extract robust reward models from LLM demonstrations. This approach negates the need for LLM fine-tuning, thereby substantially reducing computational overhead. Simultaneously, the RL policy is guided by the cumulative rewards derived from these demonstrations, effectively transferring the semantic insights captured by the LLM. Comprehensive experiments conducted on two benchmark datasets validate the effectiveness of the proposed method, demonstrating superior performance when compared against state-of-the-art RL-based and in-context learning baselines. The code can be found at https://github.com/ArronDZhang/IL-Rec. Yi Zhang 0105, Lili Xie, Ruihong Qiu, Jiajun Liu 0004, Sen Wang 0001 |
ICDM | 4 |
| 2025 | Medium-Difficulty Samples Constitute Smoothed Decision Boundary for Knowledge Distillation on Pruned DatasetsabstractThis paper tackles a new problem of dataset pruning for Knowledge Distillation (KD), from a fresh perspective of Decision Boundary (DB) preservation and drifts. Existing dataset pruning methods generally assume that the post-pruning DB formed by the selected samples can be well-captured by future networks that use those samples for training. Therefore, they tend to preserve hard samples since hard samples are closer to the DB and better characterize the nuances in the distribution of the entire dataset. However, in KD, the limited learning capacity from the student network leads to imperfect preservation of the teacher's feature distribution, resulting in the drift of DB in the student space. Specifically, hard samples worsen such drifts as they are difficult for the student to learn, creating a situation where the student's DB can drift deeper into other classes and make incorrect classifications. Motivated by these findings, our method selects medium-difficulty samples for KD-based dataset pruning. We show that these samples constitute a smoothed version of the teacher's DB and are easier for the student to learn, obtaining a general feature distribution preservation for a class of samples and reasonable DB between different classes for the student. In addition, to reduce the distributional shift due to dataset pruning, we leverage the class-wise distributional information of the teacher's outputs to reshape the logits of the preserved samples. Experiments show that the proposed static pruning method can even perform better than the state-of-the-art dynamic pruning method which needs access to the entire dataset. In addition, our method halves the training times of KD and improves the student's accuracy by 0.4% on ImageNet with a 50% keep ratio. When the ratio further increases to 70%, our method achieves higher accuracy over the vanilla KD while reducing the training times by 30%. Code is available at https://github.com/chenyd7/MDSLR. Yudong Chen 0002, Xuwei Xu, Frank de Hoog, Jiajun Liu 0004, Sen Wang 0001 |
ICLR | 4 |
| 2025 | General Scene Adaptation for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in persistent environments with relatively consistent physical layouts, visual observations, and language styles from instructors. Such a gap in the task setting presents an opportunity to improve VLN agents by incorporating continuous adaptation to specific environments. To better reflect these real-world conditions, we introduce GSA-VLN (General Scene Adaptation for VLN), a novel task requiring agents to execute navigation instructions within a specific scene and simultaneously adapt to it for improved performance over time. To evaluate the proposed task, one has to address two challenges in existing VLN datasets: the lack of out-of-distribution (OOD) data, and the limited number and style diversity of instructions for each scene. Therefore, we propose a new dataset, GSA-R2R, which significantly expands the diversity and quantity of environments and instructions for the Room-to-Room (R2R) dataset to evaluate agent adaptability in both ID and OOD contexts. Furthermore, we design a three-stage instruction orchestration pipeline that leverages large language models (LLMs) to refine speaker-generated instructions and apply role-playing techniques to rephrase instructions into different speaking styles. This is motivated by the observation that each individual user often has consistent signatures or preferences in their instructions, taking the use case of home robotic assistants as an example. We conducted extensive experiments on GSA-R2R to thoroughly evaluate our dataset and benchmark various methods, revealing key factors enabling agents to adapt to specific environments. Based on our findings, we propose a novel method, Graph-Retained DUET (GR-DUET), which incorporates memory-based navigation graphs with an environment-specific training strategy, achieving state-of-the-art results on all GSA-R2R splits. Haodong Hong, Yanyuan Qiao, Sen Wang 0001, Jiajun Liu 0004, Qi Wu 0001 |
ICLR | 4 |
| 2025 | OmniRestore: Robust Universal Image Restoration from Combined and Unspecified DegradationsabstractConventional image restoration methods often implicitly assume that the degradation type in the input image is "seen" and "known" to the model, meaning it is trained and tested on the same type of degradation. More recent "all-in-one" models are designed to handle only one single degradation type in an image at a time, though the type can vary within a small, predefined set. This paper proposes OmniRestore, a novel approach to tackle a new and challenging task: "Omni Restoration", meaning restoring images with random, combined degradations of unspecified numbers and types. In this task, the restoration model must be able to restore images corrupted by multiple degradation types simultaneously, without prior knowledge of the exact types and the number of degradations in the input image. To address this, we devise a Mixture-of-Experts (MoE) architecture with a shared encoder and a group of type-sensitive decoder experts, alongside a two-stage training pipeline to expand the generalizability to various degradation types and their combinations. Extensive experiments demonstrate that our OmniRestore model consistently and significantly outperforms all state-of-the-art (SOTA) single-degradation models, vertical ensembles of those models, and "all-in-one" models on the Omni Restoration task. Our model also surpasses most of the competing models under a single-degradation setting with seen or unseen degradations. Our dataset and code are publicly available at https://github.com/anjusreekarnavar/OmniRestore. Anjusree Karnavar, Yang Li 0184, Jiajun Liu 0004, Jun Zhou 0001, Junhu Wang |
ICME | 3 |
| 2025 | RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network LayersabstractWe reveal that feedforward network (FFN) layers, rather than attention layers, are the primary contributors to Vision Transformer (ViT) inference latency, with their impact signifying as model size increases. This finding highlights a critical opportunity for optimizing the efficiency of large-scale ViTs by focusing on FFN layers. In this work, we propose a novel channel idle mechanism that facilitates post-training structural reparameterization for efficient FFN layers during testing. Specifically, a set of feature channels remains idle and bypasses the nonlinear activation function in each FFN layer, thereby forming a linear pathway that enables structural reparameterization during inference. This mechanism results in a family of **RePa**rameterizable **Vi**sion **T**ransformers (RePaViTs), which achieve remarkable latency reductions with acceptable sacrifices (sometimes gains) in accuracy across various ViTs. The benefits of our method scale consistently with model sizes, demonstrating greater speed improvements and progressively narrowing accuracy gaps or even higher accuracies on larger models. In particular, RePa-ViT-Large and RePa-ViT-Huge enjoy **66.8%** and **68.7%** speed-ups with **+1.7%** and **+1.1%** higher top-1 accuracies under the same training strategy, respectively. RePaViT is the first to employ structural reparameterization on FFN layers to expedite ViTs to our best knowledge, and we believe that it represents an auspicious direction for efficient ViTs. Source code is available at https://github.com/Ackesnal/RePaViT. Xuwei Xu, Yang Li 0184, Yudong Chen 0002, Jiajun Liu 0004, Sen Wang 0001 |
ICML | 4 |
| 2025 | Effective Tuning Strategies for Generalist Robot Manipulation PoliciesabstractGeneralist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to struggle with out-of-distribution scenarios due to the inherent difficulty of collecting sufficient action data to cover extensively diverse domains. While fine-tuning offers a practical way to quickly adapt a GMPs to novel domains and tasks with limited samples, we observe that the performance of the resulting GMPs differs significantly with respect to the design choices of fine-tuning strategies. In this work, we first conduct an indepth empirical study to investigate the effect of key factors in GMPs fine-tuning strategies, covering the action space, policy head, supervision signal and the choice of tunable parameters, where 2,500 rollouts are evaluated for a single configuration. We systematically discuss and summarize our findings and identify the key design choices, which we believe give a practical guideline for GMPs fine-tuning. We observe that in a lowdata regime, with carefully chosen fine-tuning strategies, a GMPs significantly outperforms the state-of-the-art imitation learning algorithms. The results presented in this work establish a new baseline for future studies on fine-tuned GMPs. Wenbo Zhang 0009, Yang Li 0184, Yanyuan Qiao, Siyuan Huang 0004, Jiajun Liu 0004, Feras Dayoub, Lingqiao Liu |
ICRA | 5 |
| 2025 | Building Efficient Segmentation Models from Large Open-Vocabulary Foundation Models Without Any LabelsabstractDespite the significant success of open-vocabulary large foundation models (LFM) for segmentation in recent years, most real-life applications remain closed-vocabulary tasks and efficiency remains a critical factor for usability. Leveraging the power of open-vocabulary LFMs to create efficient, accurate closed-vocabulary segmentation models without the burden of pixel annotations provides an essential tool for employing these models effectively. This work introduces a novel Open-vocabulary to Closed-vocabulary Segmentation (O2CSeg) framework, which builds a compact, closed-vocabulary segmentation model from a large open-vocabulary LFM without any annotations. Our method capitalises on AI-assisted text prompts and learnable prompts that correspond to target class names to fully unleash the potential of open-vocabulary LFMs, eliminating the expensive annotation process. To address the challenge of noisy pseudo-labels generated by the teacher during training, we propose a confidence margin-based re-weighting knowledge distillation scheme, ensuring the student captures high-quality knowledge from the teacher. We validate our framework across diverse datasets and student configurations, demonstrating its efficacy in achieving high efficiency and accuracy for segmentation without annotations. In many cases, the derived student network outperforms its open-vocabulary teacher significantly with a higher mIoU and 40 times faster speed. Our contributions offer a promising direction for fast prototyping of efficient semantic segmentation models in scenarios where annotations are lacking, or label sets are evolving. Yang Li 0184, Diqi Chen, Sen Wang 0001, Branislav Kusy, Jiajun Liu 0004 |
IJCNN | 5 |
| 2025 | Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task PlanningabstractTo enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This requires a smart map that fuses accurate geometric structure with rich, human-understandable semantics. To address this, we introduce the 3D Queryable Scene Representation (3D QSR), a novel framework built on multimedia data that unifies three complementary 3D representations: (1) 3D-consistent novel view rendering and segmentation from panoptic reconstruction, (2) precise geometry from 3D point clouds, and (3) structured, scalable organization via 3D scene graphs. Built on an object-centric design, the framework integrates with large vision-language models to enable semantic queryability by linking multimodal object embeddings, and supporting object-level retrieval of geometric, visual, and semantic information. The retrieved data are then loaded into a robotic task planner for downstream execution. Xun Li 0004, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang 0005, Madhawa Perera, Ziwei Wang 0003, Ahalya Ravendran, Brandon J. Matthews, Matt Adcock, Dadong Wang, Jiajun Liu 0004 |
ACM Multimedia | 12 |
| 2025 | Querying Autonomous Vehicle Point Clouds: Enhanced by 3D Object Counting with CounterNetabstractAutonomous vehicles generate massive volumes of point cloud data, yet only a subset is relevant for specific tasks such as collision detection, traffic analysis, or congestion monitoring. Effectively querying this data is essential to enable targeted analytics. In this work, we formalize point cloud querying by defining three core query types: RETRIEVAL, COUNT, and AGGREGATION, each aligned with distinct analytical scenarios. All these queries rely heavily on accurate object counts to produce meaningful results, making precise object counting a critical component of query execution. Prior work has focused on indexing techniques for 2D video data, assuming detection models provide accurate counting information. However, when applied to 3D point cloud data, state-of-the-art detection models often fail to generate reliable object counts, leading to substantial errors in query results. To address this limitation, we propose CounterNet, a heatmap-based network designed for accurate object counting in large-scale point cloud data. Rather than focusing on accurate object localization, CounterNet detects object presence by finding object centers to improve counting accuracy. We further enhance its performance with a feature map partitioning strategy using overlapping regions, enabling better handling of both small and large objects in complex traffic scenes. To adapt to varying frame characteristics, we introduce a per-frame dynamic model selection strategy that selects the most effective configuration for each input. Evaluations on three real-world autonomous vehicle datasets show that CounterNet improves counting accuracy by 5% to 20% across object categories, resulting in more reliable query outcomes across all supported query types. Zhifeng Bao, Hai Dong 0001, Ziwei Wang 0003, Jiajun Liu 0004 |
ACM Multimedia | 5 |
| 2025 | Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationabstractWe present Chain-of-Action (CoA), a novel visuomotor policy paradigm built upon Trajectory Autoregressive Modeling. Unlike conventional approaches that predict next step action(s) forward, CoA generates an entire trajectory by explicit backward reasoning with task-specific goals through an action-level Chain-of-Thought (CoT) process. This process is unified within a single autoregressive structure: (1) the first token corresponds to a stable keyframe action that encodes the task-specific goals; and (2) subsequent action tokens are generated autoregressively, conditioned on the initial keyframe and previously predicted actions. This backward action reasoning enforces a global-to-local structure, allowing each local action to be tightly constrained by the final goal. To further realize the action reasoning structure, CoA incorporates four complementary designs: continuous action token representation; dynamic stopping for variable-length trajectory generation; reverse temporal ensemble; and multi-token prediction to balance action chunk modeling with global structure. As a result, CoA gives strong spatial generalization capabilities while preserving the flexibility and simplicity of a visuomotor policy. Empirically, we observe that CoA outperforms representative imitation learning algorithms such as ACT and Diffusion Policy across 60 RLBench tasks and 8 real-world tasks. Wenbo Zhang 0009, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao, Yuchu Qin, Yang Li 0184, Jiajun Liu 0004, Tao Kong, Lingqiao Liu |
NeurIPS | 7 |
| 2025 | DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic RewardabstractModel-based offline reinforcement learning (RL) has emerged as a promising approach for recommender systems, enabling effective policy learning by interacting with frozen world models. However, the reward functions in these world models, trained on sparse offline logs, often suffer from inaccuracies. Specifically, existing methods face two major limitations in addressing this challenge: (1) deterministic use of reward functions as static look-up tables, which propagates inaccuracies during policy learning, and (2) static uncertainty designs that fail to effectively capture decision risks and mitigate the impact of these inaccuracies. In this work, a dual-agent framework, DARLR, is proposed to dynamically update world models to enhance recommendation policies. To achieve this, a selector is introduced to identify reference users by balancing similarity and diversity so that the recommender can aggregate information from these users and iteratively refine reward estimations for dynamic reward shaping. Further, the statistical features of the selected users guide the dynamic adaptation of an uncertainty penalty to better align with evolving recommendation requirements. Extensive experiments on four benchmark datasets demonstrate the superior performance of DARLR, validating its effectiveness. The code is available at this address. Yi Zhang 0105, Ruihong Qiu, Xuwei Xu, Jiajun Liu 0004, Sen Wang 0001 |
SIGIR | 4 |
| 2025 | TokenBinder: Text-Video Retrieval with One-to-Many Alignment ParadigmabstractText-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners. However, these frameworks predominantly employ a one(query)-to-one(candidate) alignment paradigm, which struggles to discern nuanced differences among candidates, leading to frequent mismatches. Inspired by Comparative Judgement in human cognitive science, where decisions are made by directly comparing items rather than evaluating them independently, we propose TokenBinder. This innovative two-stage TVR framework introduces a novel one-to-many coarse-to-fine alignment paradigm, imitating the human cognitive process of identifying specific items within a large collection. Our method employs a Focused-view Fusion Network with a sophisticated cross-attention mechanism, dynamically aligning and comparing features across multiple videos to capture finer nuances and contextual variations. Extensive experiments on six benchmark datasets confirm that TokenBinder substantially outperforms existing state-of-the-art methods. These results demonstrate its robustness and the effectiveness of its fine-grained alignment in bridging intra- and inter-modality information gaps in TVR tasks. Code is avaliable at https://github.com/bingqingzhang/TokenBinder. Bingqing Zhang, Heming Du, Xin Yu 0002, Xue Li 0001, Jiajun Liu 0004, Sen Wang 0001 |
WACV | 6 |
| 2025 | On-the-Fly Object-aware Representative Point Selection in Point CloudabstractPoint clouds are essential for object modeling and play a critical role in assisting driving tasks for autonomous vehicles (AVs). However, the significant volume of data generated by AVs creates challenges for storage, bandwidth, and processing cost. To tackle these challenges, we propose a representative point selection framework for point cloud downsampling, which preserves critical object-related information while effectively filtering out irrelevant background points. Our method involves two steps: (1) Object Presence Detection, where we introduce an unsupervised density peak-based classifier and a supervised Naïve Bayes classifier to handle diverse scenarios, and (2) Sampling Budget Allocation, where we propose a strategy that selects object-relevant points while maintaining a high retention rate of object information. Extensive experiments on the KITTI and nuScenes datasets demonstrate that our method consistently outperforms state-of-the-art baselines in both efficiency and effectiveness across varying sampling rates. As a model-agnostic solution, our approach integrates seamlessly with diverse downstream models, making it a valuable and scalable addition to the 3D point cloud downsampling toolkit for AV applications. Ziwei Wang 0003, Hai Dong 0001, Zhifeng Bao, Jiajun Liu 0004 |
WACV | 5 |
| 2024 | ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender SystemsabstractOffline reinforcement learning (RL) is an effective tool for real-world recommender systems with its capacity to model the dynamic interest of users and its interactive nature. Most existing offline RL recommender systems focus on model-based RL through learning a world model from offline data and building the recommendation policy by interacting with this model. Although these methods have made progress in the recommendation performance, the effectiveness of model-based offline RL methods is often constrained by the accuracy of the estimation of the reward model and the model uncertainties, primarily due to the extreme discrepancy between offline logged data and real-world data in user interactions with online platforms. To fill this gap, a more accurate reward model and uncertainty estimation are needed for the model-based RL methods. In this paper, a novel model-based Reward Shaping in Offline Reinforcement Learning for Recommender Systems, ROLeR, is proposed for reward and uncertainty estimation in recommendation systems. Specifically, a non-parametric reward shaping method is designed to refine the reward model. In addition, a flexible and more representative uncertainty penalty is designed to fit the needs of recommendation systems. Extensive experiments conducted on four benchmark datasets showcase that ROLeR achieves state-of-the-art performance compared with existing baselines. Source code can be downloaded at this address. Yi Zhang 0105, Ruihong Qiu, Jiajun Liu 0004, Sen Wang 0001 |
CIKM | 3 |
| 2024 | LLM as Copilot for Coarse-Grained Vision-and-Language Navigation
Yanyuan Qiao, Qianyi Liu, Jiajun Liu 0004, Jing Liu 0001, Qi Wu 0001 |
ECCV (5) | 3 |
| 2024 | Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
Haodong Hong, Sen Wang 0001, Zi Huang, Qi Wu 0001, Jiajun Liu 0004 |
IJCAI | 5 |
| 2024 | Real-time Multi-modal Object Detection and Tracking on Edge for Regulatory Compliance Monitoring
Jia Syuen Lim, Ziwei Wang 0003, Jiajun Liu 0004, Abdelwahed Khamis, Reza Arablouei, Robert Barlow, Ryan McAllister |
IJCAI | 3 |
| 2024 | Edge Deployable Online Domain Adaptation for Underwater Object DetectionabstractCollecting and curating data plays a crucial role in environmental surveying. In order to gather meaningful samples, it is often necessary to develop a real-time data curation system that processes data on-the-fly and enriches it with validated models from domain experts. One area where this is particularly important is underwater marine surveys, where the vastness of the sea requires human interaction in the curation process to focus on relevant areas for exploration. Additionally, ongoing surveys are susceptible to poor performance due to data drift, which hinders the ability to provide valuable feedback for guiding data collection. While recent advancements have shown promise in addressing these challenges, they often overlook the practical constraints associated with remote data collection, such as limited processing power and latency. To overcome these limitations, this paper proposes a real-time system that adapts to data drift and enables the recording of uncertain samples for further processing and analysis on shore. The results of our approach demonstrate a remarkable improvement in species recognition, achieving an almost 18% improvement compared to the best baseline method in unseen areas. Importantly, this improvement is achieved while meeting the real-time requirements of surveys and consuming only 15W of power. By effectively addressing the challenges of underwater data drift, our proposed approach provides an efficient and effective solution for environmental surveys. Djamahl Etchegaray, Yadan Luo, Yang Li 0184, Brendan Do, Jiajun Liu 0004, Zi Huang, Branislav Kusy |
IJCNN | 5 |
| 2024 | Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsabstractReal-world navigation often involves dealing with unexpected obstructions such as closed doors, moved objects, and unpredictable entities. However, mainstream Vision-and-Language Navigation (VLN) tasks typically assume instructions perfectly align with the fixed and predefined navigation graphs without any obstructions. This assumption overlooks potential discrepancies in actual navigation graphs and given instructions, which can cause major failures for both indoor and outdoor agents. To address this issue, we integrate diverse obstructions into the R2R dataset by modifying both the navigation graphs and visual observations, introducing an innovative dataset and task, R2R with UNexpected Obstructions (R2R-UNO). R2R-UNO contains various types and numbers of path obstructions to generate instruction-reality mismatches for VLN research. Experiments on R2R-UNO reveal that state-of-the-art VLN methods inevitably encounter significant challenges when facing such mismatches, indicating that they rigidly follow instructions rather than navigate adaptively. Therefore, we propose a novel method called ObVLN (Obstructed VLN), which includes a curriculum training strategy and virtual graph construction to help agents effectively adapt to obstructed environments. Empirical results show that ObVLN not only maintains robust performance in unobstructed scenarios but also achieves a substantial performance advantage with unexpected obstructions. The source code is available at https://github.com/honghd16/ObstructedVLN. Haodong Hong, Sen Wang 0001, Zi Huang, Qi Wu 0001, Jiajun Liu 0004 |
ACM Multimedia | 5 |
| 2024 | Towards Cost-Efficient Federated Multi-agent RL with Learnable Aggregation
Yi Zhang 0105, Sen Wang 0001, Zhi Chen 0010, Xuwei Xu, Stanislav Funiak, Jiajun Liu 0004 |
PAKDD (2) | 6 |
| 2024 | GTP-ViT: Efficient Vision Transformers via Graph-based Token PropagationabstractVision Transformers (ViTs) have revolutionized the field of computer vision, yet their deployments on resource-constrained devices remain challenging due to high computational demands. To expedite pre-trained ViTs, token pruning and token merging approaches have been developed, which aim at reducing the number of tokens involved in the computation. However, these methods still have some limitations, such as image information loss from pruned tokens and inefficiency in the token-matching process. In this paper, we introduce a novel Graph-based Token Propagation (GTP) method to resolve the challenge of balancing model efficiency and information preservation for efficient ViTs. Inspired by graph summarization algorithms, GTP meticulously propagates less significant tokens’ information to spatially and semantically connected tokens that are of greater importance. Consequently, the remaining few tokens serve as a summarization of the entire token graph, allowing the method to reduce computational complexity while preserving essential information of eliminated tokens. Combined with an innovative token selection strategy, GTP can efficiently identify image tokens to be propagated. Extensive experiments have validated GTP’s effectiveness, demonstrating both efficiency and performance improvements. Specifically, GTP decreases the computational complexity of both DeiT-S and DeiT-B by up to 26% with only a minimal 0.3% accuracy drop on ImageNet-1K without finetuning, and remarkably surpasses the state-of-the-art token merging method on various backbones at an even faster inference speed. The source code is available at https://github.com/Ackesnal/GTP-ViT. Xuwei Xu, Sen Wang 0001, Yudong Chen 0002, Yanping Zheng, Zhewei Wei, Jiajun Liu 0004 |
WACV | 6 |
| 2024 | Optimized Edge Node Allocation Considering User Delay Tolerance for Cost ReductionabstractWith the rise of 5G technology, Mobile (or Multi-Access) Edge Computing (MEC) has become crucial in modern network architecture. One key research area is the effective placement of edge nodes, which has attracted significant attention. Service providers strive to minimize deployment costs for these nodes within a network. Although many studies have explored optimal strategies for reducing these costs, most overlook the allocation of computational resources and the users’ tolerance for delays. These factors add complexity, making previous methods less adaptable. In this paper, we define the Cost Minimization in MEC Edge Node Placement problem. Our goal is to find the optimal strategy for deploying edge nodes that minimize costs while cater to users’ delay tolerance limits. We prove the NP-hardness of this problem and provide a range of solutions, including Cluster-based Mixed Integer Programming, Coverage First Search, and Distance-Aware Coverage First Search, to address this challenge effectively and efficiently. Additionally, we propose a fine-grained optimization approach for allocating computational resources to edge nodes based on user service requests, significantly lowering deployment costs. Extensive experiments on a large-scale real-world dataset show that our solutions outperform the state-of-the-art in efficiency, effectiveness, and scalability. Shixun Huang, Hai Dong 0001, Zhifeng Bao, Jiajun Liu 0004, Xun Yi |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationabstractVision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction tasks like semantic segmentation, as high-resolution inputs and outputs usually imply more tokens involved in computations. Directly removing the less attentive tokens has been discussed for the image classification task but can not be extended to semantic segmentation since a dense prediction is required for every patch. To this end, this work introduces a Dynamic Token Pruning (DToP) method based on the early exit of tokens for semantic segmentation. Motivated by the coarse-to-fine segmentation process by humans, we naturally split the widely adopted auxiliary-loss-based network architecture into several stages, where each auxiliary block grades every token’s difficulty level. We can finalize the prediction of easy tokens in advance without completing the entire forward pass. Moreover, we keep k highest confidence tokens for each semantic category to uphold the representative context information. Thus, computational complexity will change with the difficulty of the input, akin to the way humans do segmentation. Experiments suggest that the proposed DToP architecture reduces on average 20% ∼ 35% of computational cost for current semantic segmentation methods based on plain vision transformers without accuracy degradation. The code is available through the following link: https://github.com/zbwxp/Dynamic-Token-Pruning. Quan Tang 0001, Bowen Zhang 0009, Jiajun Liu 0004, Fagui Liu, Yifan Liu 0001 |
ICCV | 3 |
| 2023 | OCHID-Fi: Occlusion-Robust Hand Pose Estimation in 3D via RF-VisionabstractHand Pose Estimation (HPE) is crucial to many applications, but conventional cameras-based CM-HPE methods are completely subject to Line-of-Sight (LoS), as cameras cannot capture occluded objects. In this paper, we propose to exploit Radio-Frequency-Vision (RF-vision) capable of bypassing obstacles for achieving occluded HPE, and we introduce OCHID-Fi as the first RF-HPE method with 3D pose estimation capability. OCHID-Fi employs wideband RF sensors widely available on smart devices (e.g., iPhones) to probe 3D human hand pose and extract their skeletons behind obstacles. To overcome the challenge in labeling RF imaging given its human incomprehensible nature, OCHID-Fi employs a cross-modality and cross-domain training process. It uses a pre-trained CM-HPE network and a synchronized CM/RF dataset, to guide the training of its complex-valued RF-HPE network under LoS conditions. It further transfers knowledge learned from labeled LoS domain to unlabeled occluded domain via adversarial learning, enabling OCHID-Fi to generalize to unseen occluded scenarios. Experimental results demonstrate the superiority of OCHID-Fi: it achieves comparable accuracy to CM-HPE under normal conditions while maintaining such accuracy even in occluded scenarios, with empirical evidence for its generalizability to new domains. Tianyue Zheng, Zhe Chen 0015, Jingzhi Hu, Abdelwahed Khamis, Jiajun Liu 0004, Jun Luo 0001 |
ICCV | 6 |
| 2023 | Point-Syn2Real: Semi-Supervised Synthetic-to-Real Cross-Domain Learning for Object Classification in 3D Point CloudsabstractObject classification using LiDAR 3D point cloud data is critical for modern applications such as autonomous driving. However, labeling point cloud data is labor-intensive as it requires human annotators to visualize and inspect the 3D data from different perspectives. In this paper, we propose a semi-supervised cross-domain learning approach that does not rely on manual annotations of point clouds and performs similar to fully-supervised approaches. We utilize available 3D object models to train classifiers that can generalize to real-world point clouds. We simulate the acquisition of point clouds by sampling 3D object models from multiple viewpoints and with arbitrary partial occlusions. We then augment the resulting set of point clouds through random rotations and adding Gaussian noise to better emulate the real-world scenarios. We then train point cloud encoding models on the synthesized and augmented datasets and evaluate their cross-domain classification performance on corresponding real-world datasets. We also introduce PointSyn2Real, a new benchmark dataset for cross-domain learning on point clouds. The results of our extensive experiments with this dataset demonstrate that the proposed cross-domain learning approach for point clouds outperforms the related baseline and state-of-the-art approaches in both indoor and outdoor settings in terms of cross-domain generalizability.1 Ziwei Wang 0003, Reza Arablouei, Jiajun Liu 0004, Paulo Borges, Greg Bishop-Hurley, Nicholas Heaney |
ICME | 3 |
| 2023 | Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object DetectionabstractCurrent video object detection (VOD) models often encounter issues with over-aggregation due to redundant aggregation strategies, which perform feature aggregation on every frame. This results in suboptimal performance and increased computational complexity. In this work, we propose an image-level Object Detection Difficulty (ODD) metric to quantify the difficulty of detecting objects in a given image. The derived ODD scores can be used in the VOD process to mitigate over-aggregation. Specifically, we train an ODD predictor as an auxiliary head of a still-image object detector to compute the ODD score for each image based on the discrepancies between detection results and ground-truth bounding boxes. The ODD score enhances the VOD system in two ways: 1) it enables the VOD system to select superior global reference frames, thereby improving overall accuracy; and 2) it serves as an indicator in the newly designed ODD Scheduler to eliminate the aggregation of frames that are easy to detect, thus accelerating the VOD process. Comprehensive experiments demonstrate that, when utilized for selecting global reference frames, ODD-VOD consistently enhances the accuracy of Global-frame-based VOD models. When employed for acceleration, ODD-VOD consistently improves the frames per second (FPS) by an average of 73.3% across 8 different VOD models without sacrificing accuracy. When combined, ODD-VOD attains state-of-the-art performance when competing with many VOD methods in both accuracy and speed. Our work represents a significant advancement towards making VOD more practical for real-world applications. The code will be released at https://github.com/bingqingzhang/odd-vod. Bingqing Zhang, Sen Wang 0001, Yifan Liu 0001, Branislav Kusy, Xue Li 0001, Jiajun Liu 0004 |
ACM Multimedia | 6 |
| 2023 | Decoupled Graph Neural Networks for Large Dynamic GraphsabstractReal-world graphs, such as social networks, financial transactions, and recommendation systems, often demonstrate dynamic behavior. This phenomenon, known as graph stream, involves the dynamic changes of nodes and the emergence and disappearance of edges. To effectively capture both the structural and temporal aspects of these dynamic graphs, dynamic graph neural networks have been developed. However, existing methods are usually tailored to process either continuous-time or discrete-time dynamic graphs, and cannot be generalized from one to the other. In this paper, we propose a decoupled graph neural network for large dynamic graphs, including a unified dynamic propagation that supports efficient computation for both continuous and discrete dynamic graphs. Since graph structure-related computations are only performed during the propagation process, the prediction process for the downstream task can be trained separately without expensive graph computations, and therefore any sequence model can be plugged-in and used. As a result, our algorithm achieves exceptional scalability and expressiveness. We evaluate our algorithm on seven real-world datasets of both continuous-time and discrete-time dynamic graphs. The experimental results demonstrate that our algorithm achieves state-of-the-art performance in both kinds of dynamic graphs. Most notably, the scalability of our algorithm is well illustrated by its successful application to large graphs with up to over a billion temporal edges and over a hundred million nodes. Yanping Zheng, Zhewei Wei, Jiajun Liu 0004 |
Proc. VLDB Endow. | 3 |
| 2022 | EvAnGCN: Evolving Graph Deep Neural Network Based Anomaly Detection in Blockchain
Vatsal Patel, Sutharshan Rajasegarar, Lei Pan 0002, Jiajun Liu 0004, Liming Zhu 0001 |
ADMA (1) | 4 |
| 2022 | InvisibiliTee: Angle-Agnostic Cloaking from Person-Tracking Systems with a Tee
Yaxian Li, Bingqing Zhang, Guoping Zhao, Jiajun Liu 0004, Ziwei Wang 0003, Ji-Rong Wen |
ICANN (3) | 5 |
| 2022 | STAR-GNN: Spatial-Temporal Video Representation for Content-Based RetrievalabstractWe propose a video feature representation learning frame-work called STAR-GNN, which applies a pluggable graph neural network component on a multi-scale lattice feature graph. The essence of STAR-GNN is to exploit both the temporal dynamics and spatial contents as well as vi-sual connections between regions at different scales in the frames. It models a video with a lattice feature graph in which the nodes represent regions of different granularity, with weighted edges that represent the spatial and temporal links. The contextual nodes are aggregated simultaneously by graph neural networks with parameters trained with re-trieval triplet loss. In the experiments, we show that STAR-GNN effectively implements a dynamic attention mechanism on video frame sequences, resulting in the emphasis for dy-namic and semantically rich content in the video, and is robust to noise and redundancies. Empirical results show that STAR-GNN achieves state-of-the-art performance for Content-Based Video Retrieval. Guoping Zhao, Bingqing Zhang, Yaxian Li, Jiajun Liu 0004, Ji-Rong Wen |
ICME | 5 |
| 2022 | Instant Graph Neural Networks for Dynamic GraphsabstractGraph Neural Networks (GNNs) have been widely used for modeling graph-structured data. Recent breakthroughs have been made in improving the scalability of GNNs to work on graphs with millions of nodes. However, how to instantly represent continuous changes of large-scale dynamic graphs with GNNs is still an open problem. Existing dynamic GNNs focus on modeling the periodic evolution of graphs, often on a snapshot basis. Such methods suffer from two drawbacks: first, there is a substantial delay for the changes in the graph to be reflected in the graph representations, resulting in losses on the model's accuracy; second, repeatedly calculating the representation matrix on the entire graph in each snapshot is predominantly time-consuming and severely limits the scalability. In this paper, we propose Instant Graph Neural Network (InstantGNN), an incremental computation approach for the graph representation matrix of dynamic graphs. Set to work with dynamic graphs with the edge-arrival model, our method avoids time-consuming, repetitive computations and allows instant updates on the representation and instant predictions. Graphs with dynamic structures and dynamic attributes are both supported. The upper bounds of time complexity of those updates are also provided. Furthermore, our method provides an adaptive training strategy, which guides the model to retrain at moments when it can make the greatest performance gains. We conduct extensive experiments on several real-world and synthetic datasets. Empirical results demonstrate that our model achieves state-of-the-art accuracy while having orders-of-magnitude higher efficiency than existing methods. Yanping Zheng, Hanzhi Wang 0001, Zhewei Wei, Jiajun Liu 0004, Sibo Wang 0001 |
KDD | 4 |
| 2022 | In-situ data curation: a key to actionable AI at the edgeabstractMachine learning (ML) algorithms have shown great potential in edge-computing environments, however, the literature mainly focuses on model inference only. We investigate how ML can be operationalized and how in-situ curation can improve the quality of edge applications, in the context of ML-assisted environmental surveys. We show that camera-enabled ML systems deployed on edge devices can enable scientists to perform real-time monitoring of species of interest or characterization of natural habitats. However, the benefit of this new technology is only as good as the quality and accuracy of the edge ML model inferences. In this demonstration, we show that with small additional time investment, domain scientists can manually curate ML model outputs and thus obtain highly reliable scientific insights, leading to more effective and scalable environmental surveys. Branislav Kusy, Jiajun Liu 0004, Aninda Saha, Yang Li 0184, Ross Marchant, Jeremy Oorloff, Lachlan Tychsen-Smith, David Ahmedt-Aristizabal, Brendan Do, Joey Crosswell, Russ Babcock, Andrew D. L. Steven, Megha Malpani, Ard Oerlemans |
MobiCom | 2 |
| 2022 | A real-time edge-AI system for reef surveysabstractCrown-of-Thorn Starfish (COTS) outbreaks are a major cause of coral loss on the Great Barrier Reef (GBR) and substantial surveillance and control programs are ongoing to manage COTS populations to ecologically sustainable levels. In this paper, we present a comprehensive real-time machine learning-based underwater data collection and curation system on edge devices for COTS monitoring. In particular, we leverage the power of deep learning-based object detection techniques, and propose a resource-efficient COTS detector that performs detection inferences on the edge device to assist marine experts with COTS identification during the data collection phase. The preliminary results show that several strategies for improving computational efficiency (e.g., batch-wise processing, frame skipping, model input size) can be combined to run the proposed detection model on edge hardware with low resource consumption and low information loss. Yang Li 0184, Jiajun Liu 0004, Branislav Kusy, Ross Marchant, Brendan Do, Torsten Merz, Joey Crosswell, Andrew D. L. Steven, Lachlan Tychsen-Smith, David Ahmedt-Aristizabal, Jeremy Oorloff, Peyman Moghadam, Russ Babcock, Megha Malpani, Ard Oerlemans |
MobiCom | 2 |
| 2022 | Multi-modal sensing for behaviour recognitionabstractWe describe a multi-modal sensing system for reliable recognition of animal behaviour that can operate in varying environmental conditions. We present the system architecture including the utilised sensors, visualise some intermediate processes, and provide a quantitative performance evaluation in one of the target use cases. Ziwei Wang 0003, Jiajun Liu 0004, Reza Arablouei, Greg Bishop-Hurley, Melissa Matthews, Paulo Borges |
MobiCom | 2 |
| 2022 | Improved Feature Distillation via Projector EnsembleabstractIn knowledge distillation, previous feature distillation methods mainly focus on the design of loss functions and the selection of the distilled layers, while the effect of the feature projector between the student and the teacher remains under-explored. In this paper, we first discuss a plausible mechanism of the projector with empirical evidence and then propose a new feature distillation method based on a projector ensemble for further performance improvement. We observe that the student network benefits from a projector even if the feature dimensions of the student and the teacher are the same. Training a student backbone without a projector can be considered as a multi-task learning process, namely achieving discriminative feature extraction for classification and feature matching between the student and the teacher for distillation at the same time. We hypothesize and empirically verify that without a projector, the student network tends to overfit the teacher's feature distributions despite having different architecture and weights initialization. This leads to degradation on the quality of the student's deep features that are eventually used in classification. Adding a projector, on the other hand, disentangles the two learning tasks and helps the student network to focus better on the main feature extraction task while still being able to utilize teacher features as a guidance through the projector. Motivated by the positive effect of the projector in feature distillation, we propose an ensemble of projectors to further improve the quality of student features. Experimental results on different datasets with a series of teacher-student pairs illustrate the effectiveness of the proposed method. Code is available at https://github.com/chenyd7/PEFD. Yudong Chen 0002, Sen Wang 0001, Jiajun Liu 0004, Xuwei Xu, Frank de Hoog, Zi Huang |
NeurIPS | 3 |
| 2022 | AP-GAN: Adversarial patch attack on content-based image retrieval systems
Guoping Zhao, Jiajun Liu 0004, Yaxian Li, Ji-Rong Wen |
GeoInformatica | 3 |
| 2021 | Pyramid regional graph representation learning for content-based video retrieval
Guoping Zhao, Yaxian Li, Jiajun Liu 0004, Bingqing Zhang, Ji-Rong Wen |
Inf. Process. Manag. | 4 |
| 2020 | Geosocial Co-Clustering: A Novel Framework for Geosocial Community DetectionabstractAs location-based services using mobile devices have become globally popular these days, social network analysis (especially, community detection) increasingly benefits from combining social relationships with geographic preferences. In this regard, this article addresses the emerging problem of geosocial community detection. We first formalize the problem of geosocial co-clustering , which co-clusters the users in social networks and the locations they visited. Geosocial co-clustering detects higher-quality communities than existing approaches by improving the mapping clusterability , whereby users in the same community tend to visit locations in the same region. While geosocial co-clustering is soundly formalized as non-negative matrix tri-factorization , conventional matrix tri-factorization algorithms suffer from a significant computational overhead when handling large-scale datasets. Thus, we also develop an efficient framework for geosocial co-clustering, called GEOsocial COarsening and DEcomposition (GEOCODE) . To achieve efficient matrix tri-factorization, GEOCODE reduces the numbers of users and locations through coarsening and then decomposes the single whole matrix tri-factorization into a set of multiple smaller sub-matrix tri-factorizations. Thorough experiments conducted using real-world geosocial networks show that GEOCODE reduces the elapsed time by 19–69 times while achieving the accuracy of up to 94.8% compared with the state-of-the-art co-clustering algorithm. Furthermore, the benefit of the mapping clusterability is clearly demonstrated through a local expert recommendation application. Jungeun Kim, Jae-Gil Lee 0001, Byung Suk Lee 0001, Jiajun Liu 0004 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2019 | RUM: Network Representation Learning Using MotifsabstractWe bring the novel idea of exploiting motifs into network embedding, in a dual-level network representation learning model called RUM (network Representation learning Using Motifs). Towards the leveraging of graph motifs that constitute higher-order organizations in a network, we propose two strategies, namely MotifWalk and MotifRe-weighting for learning motif-aware network embeddings. Motif-based and node-based representations are simultaneously generated, so that both the high-order structures and each node's individual properties are preserved in the final embeddings. We demonstrate that RUM has strong and well-balanced capability of preserving lowerorder proximities while discovering and capturing higher-order network structures. In empirical evaluation, RUM is tested on multiple public datasets, that range from small to medium citation networks to a large social network with more than a million nodes. Results show that the use of motifs in the representation learning process brings substantial benefits in reallife tasks, resulting in up to 12% microF1 and 8% macroF1 relative gains for node classification performance over the bestperforming competing methods. Yanlei Yu, Zhiwu Lu 0001, Jiajun Liu 0004, Guoping Zhao, Ji-Rong Wen |
ICDE | 3 |
| 2018 | Skip-Connected Deep Convolutional Autoencoder for Restoration of Document ImagesabstractThe denoising and deblurring of images are the two essential restoration tasks in the document image processing task. As the preprocessing stages of the processing pipeline, the quality of denoising and deblurring heavily influences the result of subsequent tasks, such as character detection and recognition. In this paper, we propose a novel neural method for restoring document images. We named our network Skip-Connected Deep Convolutional Autoencoder (SCDCA), which is composed of multiple layers of convolution followed by a batch normalization layer and the leaky rectified linear unit (Leaky ReLU) activation function. Inspired by the idea of residual learning, we use two types of skip connections in the network. One is identity mapping between convolution layers and the other is used to connect the input and output. Through these connections, the network learns the residual between the noisy and clean images instead of learning an ordinary transformation function. We empirically evaluate our algorithm on an open and challenging document images dataset. We also assess our restoring results using the optical character recognition (OCR) test. Experimental results have demonstrated the effectiveness and efficiency of our proposed algorithm by comparing with several state-of-the-art methods. Guoping Zhao, Jiajun Liu 0004, Hua Guan, Ji-Rong Wen |
ICPR | 2 |
| 2018 | Improving Person Re-identification by Body Parts Segmentation Generated by GANabstractPerson re-identification(ReID) is a task of associating persons that cross the non-overlapping camera views at different locations and times. It is a challenging task due to the large variations in person pose, background, luminance, occlusion, low resolution, etc. How to extracting a powerful features representation is the prime problem in ReID and is still unsolved. In this paper, we propose a cascade network architecture combined with a generative adversarial networks(GANs) and a convolutional neural network(CNN) to improve the performance of person re-identification. The GANs first generates the person body parts segmentation from the person image, and then inputs the segmentation label into the connected CNN together with the original person image. Finally obtain a discriminative and robust feature representation for ReID task. The body parts segmentation partitioning the person image into multiple segments, such as background, head, face, arms, lags, etc. The body parts segmentation information contains accurate borders and category attributes for body parts, which makes the our model more accurate compared to other predefined rigid parts alignment models. Experiments are conduced on the CUHK03, Market1501, DukeMTMC-ReID datasets and the results demonstrate that this approach outperforms several existing state-of-the-art methods. Guoping Zhao, Jiajun Liu 0004, Yanlei Yu, Ji-Rong Wen |
IJCNN | 3 |
| 2018 | A deep cascade of neural networks for image inpainting, deblurring and denoising
Guoping Zhao, Jiajun Liu 0004, Weiying Wang |
Multim. Tools Appl. | 2 |
| 2017 | A Novel Framework for Online Sales Burst Prediction
Jiajun Liu 0004 |
ECML/PKDD (3) | 2 |
| 2017 | Learning in high-dimensional multimedia data: the state of the art
Lianli Gao, Jingkuan Song, Junming Shao, Jiajun Liu 0004, Jie Shao 0001 |
Multim. Syst. | 5 |
| 2016 | Learning abstract snippet detectors with Temporal embedding in convolutional neural NetworksabstractThe prediction of periodical time-series remains challenging due to various types of scaling, misalignments and distortion effects. Here, we propose a novel model called Temporal embedding-enhanced convolutional neural Network (TeNet) to learn repeatedly-occurring-yet-hidden structural elements in periodical time-series, called abstract snippet detectors, to predict future changes. Our model effectively learns a new feature space for a time-series dataset. In the new feature space, distorted time-series that have implicit similarity but substantial differences in value and sequence to regular patterns are re-aligned to the regular patterns in the dataset, and subsequently contribute to a robust prediction mode. The model is robust to various types of distortions and misalignments and demonstrates strong prediction power for periodical time-series. We conduct extensive experiments and discover that the proposed model shows significant and consistent advantages over existing methods on a variety of data modalities ranging from human mobility to household power consumption records, when evaluated under four metrics. The model is also robust to various factors such as number of samples, variance of data, numerical ranges of data etc. The experiments verify that the intuition behind the model can be generalized to multiple data types and applications and promises significant improvement in prediction performance across the datasets studied. Jiajun Liu 0004, Kun Zhao 0003, Branislav Kusy, Ji-Rong Wen, Kai Zheng 0001, Raja Jurdak |
ICDE | 1 |
| 2016 | Multi-view ensemble learning for dementia diagnosis from neuroimaging: An artificial neural network approach
Jiajun Liu 0004, Shuo Shang, Kai Zheng 0001, Ji-Rong Wen |
Neurocomputing | 1 |
| 2016 | Prediction-based Unobstructed Route Planning
Shuo Shang, Danhuai Guo, Jiajun Liu 0004, Ji-Rong Wen |
Neurocomputing | 3 |
| 2016 | Finding regions of interest using location based social media
Shuo Shang, Danhuai Guo, Jiajun Liu 0004, Kai Zheng 0001, Ji-Rong Wen |
Neurocomputing | 3 |
| 2016 | From the lab into the wild: Design and deployment methods for multi-modal tracking platforms
Philipp Sommer, Branislav Kusy, Raja Jurdak, Navinda Kottege, Jiajun Liu 0004, Kun Zhao 0003, Adam McKeown, David Westcott |
Pervasive Mob. Comput. | 5 |
| 2016 | A Novel Framework for Online Amnesic Trajectory Compression in Resource-Constrained EnvironmentsabstractState-of-the-art trajectory compression methods usually involve high space-time complexity or yield unsatisfactory compression rates, leading to rapid exhaustion of memory, computation, storage, and energy resources. Their ability is commonly limited when operating in a resource-constrained environment especially when the data volume (even when compressed) far exceeds the storage limit. Hence, we propose a novel online framework for error-bounded trajectory compression and ageing called the Amnesic Bounded Quadrant System (ABQS), whose core is the Bounded Quadrant System (BQS) algorithm family that includes a normal version (BQS), Fast version (FBQS), and a Progressive version (PBQS). ABQS intelligently manages a given storage and compresses the trajectories with different error tolerances subject to their ages. In the experiments, we conduct comprehensive evaluations for the BQS algorithm family and the ABQS framework. Using empirical GPS traces from flying foxes and cars, and synthetic data from simulation, we demonstrate the effectiveness of the standalone BQS algorithms in significantly reducing the time and space complexity of trajectory compression, while greatly improving the compression rates of the state-of-the-art algorithms (up to 45 percent). We also show that the operational time of the target resource-constrained hardware platform can be prolonged by up to 41 percent. We then verify that with ABQS, given data volumes that are far greater than storage space, ABQS is able to achieve 15 to 400 times smaller errors than the baselines. We also show that the algorithm is robust to extreme trajectory shapes. Jiajun Liu 0004, Kun Zhao 0003, Philipp Sommer, Shuo Shang, Branislav Kusy, Jae-Gil Lee 0001, Raja Jurdak |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Bounded Quadrant System: Error-bounded trajectory compression on the goabstractLong-term location tracking, where trajectory compression is commonly used, has gained high interest for many applications in transport, ecology, and wearable computing. However, state-of-the-art compression methods involve high space-time complexity or achieve unsatisfactory compression rate, leading to rapid exhaustion of memory, computation, storage and energy resources. We propose a novel online algorithm for error-bounded trajectory compression called the Bounded Quadrant System (BQS), which compresses trajectories with extremely small costs in space and time using convex-hulls. In this algorithm, we build a virtual coordinate system centered at a start point, and establish a rectangular bounding box as well as two bounding lines in each of its quadrants. In each quadrant, the points to be assessed are bounded by the convex-hull formed by the box and lines. Various compression error-bounds are therefore derived to quickly draw compression decisions without expensive error computations. In addition, we also propose a light version of the BQS version that achieves O(1) complexity in both time and space for processing each point to suit the most constrained computation environments. Furthermore, we briefly demonstrate how this algorithm can be naturally extended to the 3-D case. Using empirical GPS traces from flying foxes, cars and simulation, we demonstrate the effectiveness of our algorithm in significantly reducing the time and space complexity of trajectory compression, while greatly improving the compression rates of the state-of-the-art algorithms (up to 47%). We then show that with this algorithm, the operational time of the target resource-constrained hardware platform can be prolonged by up to 41%. Jiajun Liu 0004, Kun Zhao 0003, Philipp Sommer, Shuo Shang, Branislav Kusy, Raja Jurdak |
ICDE | 1 |
| 2015 | Interactive Top-k Spatial Keyword queriesabstractConventional top-k spatial keyword queries require users to explicitly specify their preferences between spatial proximity and keyword relevance. In this work we investigate how to eliminate this requirement by enhancing the conventional queries with interaction, resulting in Interactive Top-k Spatial Keyword (ITkSK) query. Having confirmed the feasibility by theoretical analysis, we propose a three-phase solution focusing on both effectiveness and efficiency. The first phase substantially narrows down the search space for subsequent phases by efficiently retrieving a set of geo-textual k-skyband objects as the initial candidates. In the second phase three practical strategies for selecting a subset of candidates are developed with the aim of maximizing the expected benefit for learning user preferences at each round of interaction. Finally we discuss how to determine the termination condition automatically and estimate the preference based on the user's feedback. Empirical study based on real PoI datasets verifies our theoretical observation that the quality of top-k results in spatial keyword queries can be greatly improved through only a few rounds of interactions. Kai Zheng 0001, Han Su 0001, Bolong Zheng, Shuo Shang, Jiajie Xu 0001, Jiajun Liu 0004, Xiaofang Zhou 0001 |
ICDE | 6 |
| 2015 | Planning unobstructed paths in traffic-aware spatial networks
Shuo Shang, Jiajun Liu 0004, Kai Zheng 0001, Hua Lu 0001, Torben Bach Pedersen, Ji-Rong Wen |
GeoInformatica | 2 |
| 2015 | Dimension reduction with meta object-groups for efficient image retrieval
Shuo Shang, Jiajun Liu 0004, Kun Zhao 0003, Mingrui Yang, Kai Zheng 0001, Ji-Rong Wen |
Neurocomputing | 2 |
| 2015 | VID Join: Mapping Trajectories to Points of Interest to Support Location-Based Services
Shuo Shang, Kexin Xie, Kai Zheng 0001, Jiajun Liu 0004, Ji-Rong Wen |
J. Comput. Sci. Technol. | 4 |
| 2014 | Human Mobility Prediction and Unobstructed Route Planning in Public Transport NetworksabstractWith the increasing availability of human-tracking data (e.g., Public transport IC card data, trajectory data, etc.), human mobility prediction is increasingly important. In this paper, we study a novel problem of using human-tracking data to predict human mobility and to detect over-crowded stations in public transport networks, and then finding unobstructed routes to go around these over-crowded stations. We believe that this study can bring significant benefits to users in many popular mobile applications such as route planning and recommendation, urban computing, and location based services in general. This problem is challenged by two difficulties: (1) how to detect crowded stations effectively, and (2) how to find unobstructed routes in public transport networks efficiently. To overcome these difficulties, we propose three human-mobility prediction methods based on uniform distribution, standard normal distribution, and priority ranking, respectively, to predict human mobility and to detect over-crowded stations. Then, we develop an efficient algorithm based on network expansion to find unobstructed routes in public transport networks. The performance of the developed algorithms has been verified by extensive experiments. Shuo Shang, Danhuai Guo, Jiajun Liu 0004, Kuien Liu |
MDM (2) | 3 |
| 2014 | Semi-supervised Feature Analysis for Multimedia Annotation by Mining Label Correlation
Xiaojun Chang, Haoquan Shen, Sen Wang 0001, Jiajun Liu 0004, Xue Li 0001 |
PAKDD (2) | 4 |
| 2014 | On the Influence Propagation of Web VideosabstractWe propose a novel approach to analyze how a popular video is propagated in the cyberspace, to identify if it originated from a certain sharing-site, and to identify how it reached the current popularity in its propagation. In addition, we also estimate their influences across different websites outside the major hosting website. Web video is gaining significance due to its rich and eye-ball grabbing content. This phenomenon is evidently amplified and accelerated by the advance of Web 2.0. When a video receives some degree of popularity, it tends to appear on various websites including not only video-sharing websites but also news websites, social networks or even Wikipedia. Numerous video-sharing websites have hosted videos that reached a phenomenal level of visibility and popularity in the entire cyberspace. As a result, it is becoming more difficult to determine how the propagation took place - was the video a piece of original work that was intentionally uploaded to its major hosting site by the authors, or did the video originate from some small site then reached the sharing site after already getting a good level of popularity, or did it originate from other places in the cyberspace but the sharing site made it popular. Existing study regarding this flow of influence is lacking. Literature that discuss the problem of estimating a video's influence in the whole cyberspace also remains rare. In this article we introduce a novel framework to identify the propagation of popular videos from its major hosting site's perspective, and to estimate its influence. We define a Unified Virtual Community Space (UVCS) to model the propagation and influence of a video, and devise a novel learning method called Noise-reductive Local-and-Global Learning (NLGL) to effectively estimate a video's origin and influence. Without losing generality, we conduct experiments on annotated dataset collected from a major video sharing site to evaluate the effectiveness of the framework. Surrounding the collected videos and their ranks, some interesting discussions regarding the propagation and influence of videos as well as user behavior are also presented. Jiajun Liu 0004, Yi Yang 0001, Zi Huang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Presenting diverse location views with real-time near-duplicate photo eliminationabstractSupported by the technical advances and the commercial success of GPS-enabled mobile devices, geo-tagged photos have drawn plenteous attention in research community. The explosive growth of geo-tagged photos enables many large-scale applications, such as location-based photo browsing, landmark recognition, etc. Meanwhile, as the number of geo-tagged photos continues to climb, new challenges are brought to various applications. The existence of massive near-duplicate geo-tagged photos jeopardizes the effective presentation for the above applications. A new dimension in the search and presentation of geo-tagged photos is urgently demanded. In this paper, we devise a location visualization framework to efficiently retrieve and present diverse views captured within a local proximity. Novel photos, in terms of capture locations and visual content, are identified and returned in response to a query location for diverse visualization. For real-time response and good scalability, a new Hybrid Index structure which integrates R-tree and Geographic Grid is proposed to quickly identify the Maximal Near-duplicate Photo Groups (MNPG) in the query proximity. The most novel photos from different groups are then returned to generate diverse views on the location. Extensive experiments on synthetic and real-life photo datasets prove the novelty and efficiency of our methods. Jiajun Liu 0004, Zi Huang, Hong Cheng 0001, Yueguo Chen, Heng Tao Shen, Yanchun Zhang |
ICDE | 1 |
| 2013 | Local image tagging via graph regularized joint group sparsity
Yang Yang 0002, Zi Huang, Yi Yang 0001, Jiajun Liu 0004, Heng Tao Shen, Jiebo Luo 0001 |
Pattern Recognit. | 4 |
| 2013 | A Gram-Based String Paradigm for Efficient Video Subsequence SearchabstractThe unprecedented increase in the generation and dissemination of video data has created an urgent demand for the large-scale video content management system to quickly retrieve videos of users' interests. Traditionally, video sequence data are managed by high-dimensional indexing structures, most of which suffer from the well-known “curse of dimensionality” and lack of support of subsequence retrieval. Inspired by the high efficiency of string indexing methods, in this paper, we present a string paradigm called VideoGram for large-scale video sequence indexing to achieve fast similarity search. In VideoGram, the feature space is modeled as a set of visual words. Each database video sequence is mapped into a string. A gram-based indexing structure is then built to tackle the effect of the “curse of dimensionality” and support video subsequence matching. Given a high-dimensional query video sequence, retrieval is performed by transforming the query into a string and then searching the matched strings from the index structure. By doing so, expensive high-dimensional similarity computations can be completely avoided. An efficient sequence search algorithm with upper bound pruning power is also presented. We conduct an extensive performance study on real-life video collections to validate the novelties of our proposal. Zi Huang, Jiajun Liu 0004, Bin Cui 0001, Xiaoyong Du 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Discovering areas of interest with geo-tagged images and check-insabstractGeo-tagged image is an ideal source for the discovery of popular travel places. However, the aspects of popular venues for daily-life purposes like dining and shopping are often missing in the mined locations from geo-tagged images. Fortunately check-in websites provide us a unique opportunity of analyzing people's preferences in their daily lives to complement the knowledge mined from geo-tagged images. This paper presents a novel approach for the discovery of Areas of Interest (AoI). By analyzing both geo-tagged images and check-ins, the approach exploits travelers' flavors as well as the preferences of daily-life activities of local residents to find AoI in a city. The proposed approach consists of two major steps. Firstly, we devise a density-based clustering method to discover AoI, mainly based on the image densities but also reinforced by the secondary densities from the images' neighboring venues. Then we propose a novel joint authority analysis framework to rank AoI. The framework simultaneously considers both the location-location transitions, and the user-location relations. An interactive presentation interface for visualizing AoI is also presented. The approach is tested with very large datasets for Shanghai city. They consist of 49,460 geo-tagged images from Panoramio.com, and 1,361,547 check-ins from the check-in website Qieke.com. By evaluating the ranking accuracy and quality of AoI, we demonstrate great improvements of our method over compared methods. Jiajun Liu 0004, Zi Huang, Lei Chen 0002, Heng Tao Shen, Zhixian Yan |
ACM Multimedia | 1 |
| 2012 | Robust cross-media transfer for visual event detectionabstractIn this paper, we present a novel approach, named Robust Cross-Media Transfer (RCMT), for visual event detection in social multimedia environments. Different from most existing methods, the proposed method can directly take different types of noisy social multimedia data as input and conduct robust event detection. More specifically, we build a robust model by employing an l2,1-norm regression model featuring noise tolerance, and also manage to integrate different types of social multimedia data by minimizing the distribution difference among them. Experimental results on real-life Flickr image dataset and YouTube video dataset demonstrate the effectiveness of our proposal, compared to state-of-the-art algorithms. Yang Yang 0002, Yi Yang 0001, Zi Huang, Jiajun Liu 0004, Zhigang Ma |
ACM Multimedia | 4 |
| 2011 | Efficient Histogram-Based Similarity Search in Ultra-High Dimensional Space
Jiajun Liu 0004, Zi Huang, Heng Tao Shen, Xiaofang Zhou 0001 |
DASFAA (2) | 1 |
| 2011 | Effective data co-reduction for multimedia similarity searchabstractMultimedia similarity search has been playing a critical role in many novel applications. Typically, multimedia objects are described by high-dimensional feature vectors (or points) which are organized in databases for retrieval. Although many high-dimensional indexing methods have been proposed to facilitate the search process, efficient retrieval over large, sparse and extremely high-dimensional databases remains challenging due to the continuous increases in data size and feature dimensionality. In this paper, we propose the first framework for Data Co-Reduction (DCR) on both data size and feature dimensionality. By utilizing recently developed co-clustering methods, DCR simultaneously reduces both size and dimensionality of the original data into a compact subspace, where lower bounds of the actual distances in the original space can be efficiently established to achieve fast and lossless similarity search in the filter-and refine approach. Particularly, DCR considers the duality between size and dimensionality, and achieves the optimal coreduction which generates the least number of candidates for actual distance computations. We conduct an extensive experimental study on large and real-life multimedia datasets, with dimensionality ranging from 432 to 1936. Our results demonstrate that DCR outperforms existing methods significantly for lossless retrieval, especially in the presence of extremely high dimensionality. Zi Huang, Heng Tao Shen, Jiajun Liu 0004, Xiaofang Zhou 0001 |
SIGMOD Conference | 3 |
| 2011 | Correlation-based retrieval for heavily changed near-duplicate videosabstractThe unprecedented and ever-growing number of Web videos nowadays leads to the massive existence of near-duplicate videos. Very often, some near-duplicate videos exhibit great content changes, while the user perceives little information change, for example, color features change significantly when transforming a color video with a blue filter. These feature changes contribute to low-level video similarity computations, making conventional similarity-based near-duplicate video retrieval techniques incapable of accurately capturing the implicit relationship between two near-duplicate videos with fairly large content modifications. In this paper, we introduce a new dimension for near-duplicate video retrieval. Different from existing near-duplicate video retrieval approaches which are based on video-content similarity, we explore the correlation between two videos. The intuition is that near-duplicate videos should preserve strong information correlation in spite of intensive content changes. More effective retrieval with stronger tolerance is achieved by replacing video-content similarity measures with information correlation analysis. Theoretical justification and experimental results prove the effectiveness of correlation-based near-duplicate retrieval. Jiajun Liu 0004, Zi Huang, Heng Tao Shen, Bin Cui 0001 |
ACM Trans. Inf. Syst. | 1 |